AI Frontier
← 浏览此厂商的报告

Tencent / Hunyuan / Technical Report

HunyuanVideo 1.5 Technical Report

HunyuanVideo 1.5 · 2025-11-24

概要原文

保留原文 · 保留原始语言

Abstract · 页码 1

We present HunyuanVideo 1.5, a lightweight yet powerful open-source video generation model that achieves state-of-the-art visual quality and motion coher- ence with only 8.3 billion parameters, enabling efficient inference on consumer- grade GPUs. This achievement is built upon several key components, including meticulous data curation, an advanced DiT architecture featuring selective and sliding tile attention (SSTA), enhanced bilingual understanding through glyph- aware text encoding, progressive pre-training and post-training, and an efficient video super-resolution network. Leveraging these designs, we developed a unified framework capable of high-quality text-to-video and image-to-video generation across multiple durations and resolutions. Extensive experiments demonstrate that this compact and proficient model establishes a new state-of-the-art among open-source video generation models. By releasing the code and model weights, we provide the community with a high-performance foundation that lowers the barrier to video creation and research, making advanced video generation ac- cessible to a broader audience. All open-source assets are publicly available at https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5.

核心图片

点击放大查看,下载原图获取完整细节。

图 1 · 模型架构页码 3
Figure 1: Caption Model Post-training Pipeline.
图 2 · 模型架构页码 4
Figure 2: Architecture of the Unified Diffusion Transformer.

点击图片切换缩放,按 Esc 关闭。每张图下方可下载高清文件。