AI Frontier
← 浏览此厂商的报告

NVIDIA / Technical Report

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Nemotron 3 Nano Omni · 2026-04-27

概要原文

保留原文 · 保留原始语言

Abstract. · 页码 1

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, across all modalities, enabled by advances in architecture, training data and recipes. In particular, Nemotron 3 delivers leading results in real-world document understanding, long audio-video comprehension, and agentic computer use. Built on the highly efficient Nemotron 3 Nano 30B-A3B backbone, Nemotron 3 Nano Omni further incorporates innovative multimodal token-reduction techniques to deliver substantially lower inference latency and higher throughput than other models of similar size. We are releasing model checkpoints in BF16, FP8, and FP4 formats, along with portions of the training data and codebase to facilitate further research and development.

核心图片

点击放大查看,下载原图获取完整细节。

图 1 · 模型架构页码 3
Figure 1 | Nemotron 3 Nano Omni architecture. For encoding images and videos we use dynamic resolution. Additionally, videos use Conv3D and optionally Efficient Video Sampling for higher throughput. Audio inputs are encoded using Parakeet v2 audio encoder. Visual, audio, and text tokens are concatenated and fed to the LLM.
图 2 · 模型架构页码 4
Figure 2 | Staged training recipe for the v3 omni-modal model. The pipeline first performs vision SFT, then joint omni SFT while progressively extending context length, followed by omni-modal RL training.

点击图片切换缩放,按 Esc 关闭。每张图下方可下载高清文件。