AI Frontier
← 浏览此厂商的报告

Tencent / Hunyuan / Technical Report

HunyuanImage 3.0 Technical Report

HunyuanImage 3.0 · 2025-09-28

概要原文

保留原文 · 保留原始语言

Abstract · 页码 1

We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module publicly available. The achievement of HunyuanImage 3.0 relies on several key components, including meticulous data curation, advanced architecture design, a native Chain-of-Thoughts schema, progressive model pre- training, aggressive model post-training, and an efficient infrastructure that enables large-scale training and inference. With these advancements, we successfully trained a Mixture-of-Experts (MoE) model comprising over 80 billion parameters in total, with 13 billion parameters activated per token during inference, making it the largest and most powerful open-source image generative model to date. We conducted extensive experiments and the results of automatic and human evaluation of text-image alignment and visual quality demonstrate that HunyuanImage 3.0 rivals previous state-of-the-art models. By releasing the code and weights of HunyuanImage 3.0, we aim to enable the community to explore new ideas with a state-of-the-art foundation model, fostering a dynamic and vibrant multimodal ecosystem. All open source assets are publicly available at here.

核心图片

点击放大查看,下载原图获取完整细节。

图 3 · 概览页码 6
Figure 3: Illustration of HunyuanImage 3.0.
图 7 · 实验结果页码 12
Figure 7: GSB evaluation results.

点击图片切换缩放,按 Esc 关闭。每张图下方可下载高清文件。