AI Frontier
← 浏览此厂商的报告

Google / DeepMind / Technical Report

Gemma 4 Technical Report

Gemma 4 · 2026-07-02

概要原文

保留原文 · 保留原始语言

开篇原文 · 页码 1

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

核心图片

点击放大查看,下载原图获取完整细节。

图 1 · 模型架构页码 3
Figure 1 | The autoregressive MTP drafter (blue blocks on the right) is fed activations and KV cache from the main model (gray blocks).

点击图片切换缩放,按 Esc 关闭。每张图下方可下载高清文件。