cs.CLJul 2, 2026

Gemma 4 Technical Report

Authors: Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, +293 more

Abstract

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

Explore similar work

CardsList
  1. DiffusionGemma Technical Report

    Jul 31, 2026DiffusionGemma Team, Adrien Ali Taïga, James Assiene +41Diffusion Language ModelsAutoregressive Generation

  2. Instella-MoE Technical Report

    Sep 1, 2026Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra +10Mixture-Of-Expert Language ModelsLarge Models