Full technical report tables

Gemma 4 Benchmarks

Every published score for all five model sizes - text, vision, audio, long context and human preference - with the evaluation settings spelled out, the competitive picture, and an honest section on where Gemma 4 falls short.

Source: Gemma 4 Technical Report (arXiv:2607.02770) · thinking mode enabled unless noted

Gemma 4 31B · headline scores
AIME 2026
89.2
MMLU Pro
85.2
MATH-Vision
85.6
GPQA Diamond
84.3
LiveCodeBench v6
80.0
MMMU Pro
76.9
IFEval 98.9 · Codeforces Elo 2150 · RULER @128k 96.4
Text & reasoning

The core results

All five sizes, with Gemma 3 27B included as a baseline so you can see how much changed in a generation. Thinking mode is enabled throughout.

AIME 2026

Competition mathematics, no tools
31B89.2
26B A4B88.3
12B77.5
E4B42.5
E2B37.5
Gemma 3 27B20.8

MMLU Pro

Broad knowledge and reasoning
31B85.2
26B A4B82.6
12B77.2
E4B69.4
E2B60.0
Gemma 3 27B67.6

GPQA Diamond

Graduate-level science
31B84.3
26B A4B82.3
12B78.8
E4B58.6
E2B43.4
Gemma 3 27B42.4

LiveCodeBench v6

Contamination-resistant coding
31B80.0
26B A4B77.1
12B72.0
E4B52.0
E2B44.0
Gemma 3 27B29.1

Complete text table

Benchmark31B26B A4B12B E4BE2BGemma 3 27B
MMLU Pro85.282.677.269.460.067.6
AIME 2026 no tools89.288.377.542.537.520.8
GPQA Diamond84.382.378.858.643.442.4
LiveCodeBench v680.077.172.052.044.029.1
Codeforces Elo215017181659940633110
BBH micro avg74.464.853.033.121.919.3
SciCode43.040.038.024.021.021.0
IFEval instruction following98.998.597.296.794.690.4
IFBench76.072.074.044.038.032.0
HLE Humanity's Last Exam19.58.75.2---
HLE with search26.517.2----
📈
The generational jump is the real story. Compare the last column. On AIME, Gemma 3 27B scored 20.8 and Gemma 4 31B scores 89.2. On Codeforces the Elo went from 110 to 2150. Even the 2-billion parameter E2B (37.5 on AIME) now beats the previous generation's 27B flagship. That's what a year of reasoning-focused post-training bought.
Vision

Multimodal results

Measured at maximum resolution - 1120 vision tokens. This matters more than it sounds: the default is lower, and the scores move with it.

Benchmark31B26B A4B12B E4BE2BGemma 3 27B
MMMU Pro76.973.869.152.644.249.7
MATH-Vision85.682.479.759.552.446.0
InfographicVQA92.089.388.470.063.970.6
MedXpertQA MM61.358.148.728.723.5-
OmniDocBench 1.5 lower is better ↓0.1310.1490.1640.1810.2900.365
🔍
Resolution changes the score - check which one you're running. At the low-resolution setting (280 tokens), the same 31B scores 75.8 on MMMU Pro instead of 76.9, and InfographicVQA drops from 92.0 all the way to 82.8. Document-heavy tasks are the most sensitive. The July 2026 refresh made 280 tokens the default with 1120 available, so if you're comparing your own results against the headline table, make sure you've raised the token budget.
Low resolution (280 tokens)31B26B A4B12BE4BE2B
MMMU Pro75.873.267.751.443.2
MATH-Vision83.480.376.759.253.0
InfographicVQA82.877.858.754.844.6
MedXpertQA MM60.755.747.428.722.5
Long context

Does 256K actually work?

Supporting a long context and being useful across it are different claims. These are measured without thinking mode.

BenchmarkContext31B26B A4B 12BE4BE2BGemma 3 27B
RULER32K96.897.396.495.283.091.1
RULER128K96.489.891.286.670.466.0
LOFT retrieval Recall@k128K79.566.366.458.550.58.6
GraphWalks F1<128K82.372.671.050.94.132.8
MTOB eng→kgv128K52.950.045.137.815.441.0
MTOB eng→kgv256K54.348.941.9---

The 31B holds up genuinely well

RULER barely moves from 32K to 128K - 96.8 to 96.4 - where Gemma 3 27B collapsed from 91.1 to 66.0. And MTOB is the only score here that improves at 256K over 128K, which suggests the extra context is being used rather than merely tolerated.

E2B is a different story

GraphWalks drops to 4.1 - effectively a failure. RULER falls to 70.4 at 128K. The small edge models nominally support 128K, but multi-hop reasoning across a long context is not something to rely on them for. Keep prompts short on E2B.

Audio

Speech recognition and translation

The edge models handle audio natively, and improved over Gemma 3n while shrinking the audio encoder substantially.

🎙️

FLEURS - transcription

Word error rate, averaged across languages. Lower is better.

  • Gemma 4 12B0.063
  • Gemma 4 E4B0.075
  • Gemma 4 E2B0.090
  • Gemma 3n E4B0.085
  • Gemma 3n E2B0.108
🌍

CoVoST - speech translation

CorpusBLEU into English, averaged. Higher is better.

  • Gemma 4 12B42.3
  • Gemma 4 E4B38.2
  • Gemma 4 E2B35.4
  • Gemma 3n E4B34.7
  • Gemma 3n E2B31.6
📉
Better and smaller at once. Against Gemma 3n at matching sizes, Gemma 4 improves translation by 12% (E2B) and 10% (E4B), and transcription by 17% and 12% - while the on-disk audio encoder shrank by 78%, from 680M parameters to 305M. That's the efficiency work that makes on-device audio practical.
Human preference

LMArena Elo

Blind side-by-side comparisons judged by human raters - the benchmark hardest to game, and the one that moves the most over time.

ModelEloRankCategoryAs of
Gemma 4 31B1451 ± 843 Leading open dense modelJune 19, 2026
Gemma 4 26B A4B1438 ± 861 Open MoEJune 19, 2026
Gemma 3 27B1366 ± 4157 Previous generationJune 19, 2026
📅
Arena rankings decay, and you'll see two different numbers quoted. At launch, LMArena placed Gemma 4 31B at #3 among open models and #27 overall, on par with Kimi-K2.5 and Qwen-3.5-397B - a jump of 87 Elo points over Gemma 3 27B. By the technical report's June 19 cutoff, the same model sat at rank 43 overall as newer models arrived. Both figures are accurate for their date; neither is the current standing. If a page quotes an Arena rank without a date, treat it as decorative.
Competitive picture

How it compares

Gemma 4's claim isn't that it beats frontier closed models. It's the performance you can get per gigabyte, and per dollar, from something you can download.

🏋️

Against open models

At launch the 31B was rated on par with Kimi-K2.5 and Qwen-3.5-397B on Arena - models roughly ten times its size. Among open dense models it led the leaderboard. That size-for-quality ratio is the strongest single claim in Gemma 4's favour.

🔒

Against frontier closed models

Gemini 3, GPT-5 and Claude Opus 4.x remain ahead on the hardest reasoning work - HLE is the clearest illustration, where 19.5 is a long way from frontier scores. For most everyday tasks the gap is far narrower than that number suggests.

📱

Against other small models

Where Gemma 4 is genuinely unmatched is the bottom end. E2B scoring 37.5 on AIME and 60.0 on MMLU Pro, in under a gigabyte of RAM on a phone, has no real equivalent in the open ecosystem.

⚖️
Why there's no head-to-head table here. Comparing published scores across vendors is mostly misleading - different harnesses, different shot counts, thinking budgets that aren't disclosed, and self-reported numbers that nobody independently reproduces. Where a like-for-like comparison exists, such as Arena's blind human ratings or the Gemma 3 baseline in the tables above, it's shown. Everything else would be a table that looks authoritative and isn't.
The honest part

Where Gemma 4 is weak

Benchmark pages usually stop at the flattering numbers. These are the ones that should change what you build with it.

Weak

Agentic and tool-use workflows

The single biggest gap. Independent aggregate testing ranks Gemma 4 31B near the bottom on agentic composites - #129 of 134, scoring 25.5/100 - despite its excellent MMLU Pro and GPQA results. Strong reasoning in a single turn does not translate to reliable multi-step tool use.

The July 2026 refresh improved this materially (31B gained 10.1% on Tau2 Telecom, and E4B went from effectively zero on TB2 agent benchmarks to a working score). Make sure you're on post-July weights - but still prototype agents carefully before committing.

Weak

Frontier-difficulty reasoning

HLE is designed to be brutally hard, and 19.5 for the 31B reflects that - dropping to 8.7 for the 26B A4B and 5.2 for the 12B. Search grounding raises the 31B to 26.5, which is the more useful configuration in practice, but this is where the gap to frontier closed models is widest.

Weak

Real-world software engineering

LiveCodeBench v6 at 80.0 looks superb, but that measures self-contained competitive-programming problems. On SWE-Rebench - patching actual repositories - independent testing puts the 31B at 41.6%. Both numbers are real; they measure very different jobs.

Weak

Small models on long or multi-hop context

E2B scores 4.1 on GraphWalks and 21.9 on BBH. The 128K context is real for retrieval but not for reasoning across a document. Treat E2B as excellent at short, well-scoped tasks and unreliable the moment a task needs chaining.

Methodology

How to read these numbers

Six things that determine whether the scores above will match what you measure yourself.

🧠

Thinking mode is on

Every text and vision score assumes thinking mode enabled. If you deploy with it minimised for latency, expect materially lower results - particularly on maths and code. The long-context table is the exception and is measured without thinking.

🖼️

Vision resolution is set to maximum

Headline vision numbers use 1120 vision tokens. The current default is 280, where InfographicVQA falls from 92.0 to 82.8. Raise the token budget before comparing.

🔬

Contamination varies by benchmark

Training data ends January 2025, so AIME 2026 is genuinely post-cutoff and LiveCodeBench v6 is designed to resist contamination. Older static benchmarks carry more risk of having leaked into training data.

📄

Most numbers are self-reported

The tables come from Google's own technical report. That's normal and they're generally reliable, but they aren't independent. The Arena Elo and the third-party agentic and SWE-Rebench figures cited above are the outside checks.

↕️

Two metrics run backwards

OmniDocBench and FLEURS word error rate are lower is better. A 0.131 beats a 0.365. It's an easy misread when scanning a table where everything else is a percentage.

📅

Weights changed mid-year

The July 15, 2026 refresh altered tool-calling behaviour, the chat template and vision defaults without changing the version number. Benchmarks run before and after that date are not strictly comparable.

🎯
The benchmark that matters is yours. None of these tables tell you whether Gemma 4 will work for your task. Assemble twenty or thirty representative examples from your actual workload, run them against two model sizes, and read the outputs. That takes an afternoon and tells you more than any leaderboard - especially since the smaller models are often good enough, which is the decision that actually saves money.

Test it on your own work.

Numbers on a page are a starting point. Run the model against your real prompts and judge for yourself.