AI Frontier
← Browse this publisher

Google / DeepMind / Technical Report

Gemma 4 Technical Report

Gemma 4 · 2026-07-02

Source summary

Original wording · Original language

Opening passage · Page 1

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

Core figures

Enlarge to explore. Download the original for full detail.

Figure 1 · ArchitecturePage 3
Figure 1 | The autoregressive MTP drafter (blue blocks on the right) is fed activations and KV cache from the main model (gray blocks).

Click the image to zoom. Press Esc to close. Full-resolution files are available below each figure.