Gemma 4 Benchmarks
Every published score for all five model sizes - text, vision, audio, long context and human preference - with the evaluation settings spelled out, the competitive picture, and an honest section on where Gemma 4 falls short.
Source: Gemma 4 Technical Report (arXiv:2607.02770) · thinking mode enabled unless noted
The core results
All five sizes, with Gemma 3 27B included as a baseline so you can see how much changed in a generation. Thinking mode is enabled throughout.
AIME 2026
MMLU Pro
GPQA Diamond
LiveCodeBench v6
Complete text table
| Benchmark | 31B | 26B A4B | 12B | E4B | E2B | Gemma 3 27B |
|---|---|---|---|---|---|---|
| MMLU Pro | 85.2 | 82.6 | 77.2 | 69.4 | 60.0 | 67.6 |
| AIME 2026 no tools | 89.2 | 88.3 | 77.5 | 42.5 | 37.5 | 20.8 |
| GPQA Diamond | 84.3 | 82.3 | 78.8 | 58.6 | 43.4 | 42.4 |
| LiveCodeBench v6 | 80.0 | 77.1 | 72.0 | 52.0 | 44.0 | 29.1 |
| Codeforces Elo | 2150 | 1718 | 1659 | 940 | 633 | 110 |
| BBH micro avg | 74.4 | 64.8 | 53.0 | 33.1 | 21.9 | 19.3 |
| SciCode | 43.0 | 40.0 | 38.0 | 24.0 | 21.0 | 21.0 |
| IFEval instruction following | 98.9 | 98.5 | 97.2 | 96.7 | 94.6 | 90.4 |
| IFBench | 76.0 | 72.0 | 74.0 | 44.0 | 38.0 | 32.0 |
| HLE Humanity's Last Exam | 19.5 | 8.7 | 5.2 | - | - | - |
| HLE with search | 26.5 | 17.2 | - | - | - | - |
Multimodal results
Measured at maximum resolution - 1120 vision tokens. This matters more than it sounds: the default is lower, and the scores move with it.
| Benchmark | 31B | 26B A4B | 12B | E4B | E2B | Gemma 3 27B |
|---|---|---|---|---|---|---|
| MMMU Pro | 76.9 | 73.8 | 69.1 | 52.6 | 44.2 | 49.7 |
| MATH-Vision | 85.6 | 82.4 | 79.7 | 59.5 | 52.4 | 46.0 |
| InfographicVQA | 92.0 | 89.3 | 88.4 | 70.0 | 63.9 | 70.6 |
| MedXpertQA MM | 61.3 | 58.1 | 48.7 | 28.7 | 23.5 | - |
| OmniDocBench 1.5 lower is better ↓ | 0.131 | 0.149 | 0.164 | 0.181 | 0.290 | 0.365 |
| Low resolution (280 tokens) | 31B | 26B A4B | 12B | E4B | E2B |
|---|---|---|---|---|---|
| MMMU Pro | 75.8 | 73.2 | 67.7 | 51.4 | 43.2 |
| MATH-Vision | 83.4 | 80.3 | 76.7 | 59.2 | 53.0 |
| InfographicVQA | 82.8 | 77.8 | 58.7 | 54.8 | 44.6 |
| MedXpertQA MM | 60.7 | 55.7 | 47.4 | 28.7 | 22.5 |
Does 256K actually work?
Supporting a long context and being useful across it are different claims. These are measured without thinking mode.
| Benchmark | Context | 31B | 26B A4B | 12B | E4B | E2B | Gemma 3 27B |
|---|---|---|---|---|---|---|---|
| RULER | 32K | 96.8 | 97.3 | 96.4 | 95.2 | 83.0 | 91.1 |
| RULER | 128K | 96.4 | 89.8 | 91.2 | 86.6 | 70.4 | 66.0 |
| LOFT retrieval Recall@k | 128K | 79.5 | 66.3 | 66.4 | 58.5 | 50.5 | 8.6 |
| GraphWalks F1 | <128K | 82.3 | 72.6 | 71.0 | 50.9 | 4.1 | 32.8 |
| MTOB eng→kgv | 128K | 52.9 | 50.0 | 45.1 | 37.8 | 15.4 | 41.0 |
| MTOB eng→kgv | 256K | 54.3 | 48.9 | 41.9 | - | - | - |
The 31B holds up genuinely well
RULER barely moves from 32K to 128K - 96.8 to 96.4 - where Gemma 3 27B collapsed from 91.1 to 66.0. And MTOB is the only score here that improves at 256K over 128K, which suggests the extra context is being used rather than merely tolerated.
E2B is a different story
GraphWalks drops to 4.1 - effectively a failure. RULER falls to 70.4 at 128K. The small edge models nominally support 128K, but multi-hop reasoning across a long context is not something to rely on them for. Keep prompts short on E2B.
Speech recognition and translation
The edge models handle audio natively, and improved over Gemma 3n while shrinking the audio encoder substantially.
FLEURS - transcription
Word error rate, averaged across languages. Lower is better.
- Gemma 4 12B0.063
- Gemma 4 E4B0.075
- Gemma 4 E2B0.090
- Gemma 3n E4B0.085
- Gemma 3n E2B0.108
CoVoST - speech translation
CorpusBLEU into English, averaged. Higher is better.
- Gemma 4 12B42.3
- Gemma 4 E4B38.2
- Gemma 4 E2B35.4
- Gemma 3n E4B34.7
- Gemma 3n E2B31.6
LMArena Elo
Blind side-by-side comparisons judged by human raters - the benchmark hardest to game, and the one that moves the most over time.
| Model | Elo | Rank | Category | As of |
|---|---|---|---|---|
| Gemma 4 31B | 1451 ± 8 | 43 | Leading open dense model | June 19, 2026 |
| Gemma 4 26B A4B | 1438 ± 8 | 61 | Open MoE | June 19, 2026 |
| Gemma 3 27B | 1366 ± 4 | 157 | Previous generation | June 19, 2026 |
How it compares
Gemma 4's claim isn't that it beats frontier closed models. It's the performance you can get per gigabyte, and per dollar, from something you can download.
Against open models
At launch the 31B was rated on par with Kimi-K2.5 and Qwen-3.5-397B on Arena - models roughly ten times its size. Among open dense models it led the leaderboard. That size-for-quality ratio is the strongest single claim in Gemma 4's favour.
Against frontier closed models
Gemini 3, GPT-5 and Claude Opus 4.x remain ahead on the hardest reasoning work - HLE is the clearest illustration, where 19.5 is a long way from frontier scores. For most everyday tasks the gap is far narrower than that number suggests.
Against other small models
Where Gemma 4 is genuinely unmatched is the bottom end. E2B scoring 37.5 on AIME and 60.0 on MMLU Pro, in under a gigabyte of RAM on a phone, has no real equivalent in the open ecosystem.
Where Gemma 4 is weak
Benchmark pages usually stop at the flattering numbers. These are the ones that should change what you build with it.
Agentic and tool-use workflows
The single biggest gap. Independent aggregate testing ranks Gemma 4 31B near the bottom on agentic composites - #129 of 134, scoring 25.5/100 - despite its excellent MMLU Pro and GPQA results. Strong reasoning in a single turn does not translate to reliable multi-step tool use.
The July 2026 refresh improved this materially (31B gained 10.1% on Tau2 Telecom, and E4B went from effectively zero on TB2 agent benchmarks to a working score). Make sure you're on post-July weights - but still prototype agents carefully before committing.
Frontier-difficulty reasoning
HLE is designed to be brutally hard, and 19.5 for the 31B reflects that - dropping to 8.7 for the 26B A4B and 5.2 for the 12B. Search grounding raises the 31B to 26.5, which is the more useful configuration in practice, but this is where the gap to frontier closed models is widest.
Real-world software engineering
LiveCodeBench v6 at 80.0 looks superb, but that measures self-contained competitive-programming problems. On SWE-Rebench - patching actual repositories - independent testing puts the 31B at 41.6%. Both numbers are real; they measure very different jobs.
Small models on long or multi-hop context
E2B scores 4.1 on GraphWalks and 21.9 on BBH. The 128K context is real for retrieval but not for reasoning across a document. Treat E2B as excellent at short, well-scoped tasks and unreliable the moment a task needs chaining.
How to read these numbers
Six things that determine whether the scores above will match what you measure yourself.
Thinking mode is on
Every text and vision score assumes thinking mode enabled. If you deploy with it minimised for latency, expect materially lower results - particularly on maths and code. The long-context table is the exception and is measured without thinking.
Vision resolution is set to maximum
Headline vision numbers use 1120 vision tokens. The current default is 280, where InfographicVQA falls from 92.0 to 82.8. Raise the token budget before comparing.
Contamination varies by benchmark
Training data ends January 2025, so AIME 2026 is genuinely post-cutoff and LiveCodeBench v6 is designed to resist contamination. Older static benchmarks carry more risk of having leaked into training data.
Most numbers are self-reported
The tables come from Google's own technical report. That's normal and they're generally reliable, but they aren't independent. The Arena Elo and the third-party agentic and SWE-Rebench figures cited above are the outside checks.
Two metrics run backwards
OmniDocBench and FLEURS word error rate are lower is better. A 0.131 beats a 0.365. It's an easy misread when scanning a table where everything else is a percentage.
Weights changed mid-year
The July 15, 2026 refresh altered tool-calling behaviour, the chat template and vision defaults without changing the version number. Benchmarks run before and after that date are not strictly comparable.
Test it on your own work.
Numbers on a page are a starting point. Run the model against your real prompts and judge for yourself.