Gemma Model
Comparison
Gemma against Gemma — five current sizes, four generations, and a dozen specialised variants. Put any two side by side, see what actually changed between generations, and check the licence before you build on one.
For comparisons against Llama, Qwen and closed models, see benchmarks
Compare any two models
Pick two and the differences are highlighted. Everything here is from Google's technical report or official model documentation.
The five Gemma 4 sizes
One table with everything that differs between them. Speed figures are single-stream on an H100 for the server models, and on-device for the edge ones — they aren't the same measurement, and the table says which is which.
| Spec | 31B | 26B A4B | 12B | E4B | E2B |
|---|---|---|---|---|---|
| Architecture | Dense | MoE | Unified | Edge | Edge |
| Active params / token | 31B | ~3.8B | 12B | 4.5B | 2.3B |
| Context window | 256K | 256K | 256K | 128K | 128K |
| Audio & video input | — | — | Yes | Yes | Yes |
| Memory · BF16 | 69.9 GB | 57.7 GB | 26.7 GB | 17.9 GB | 11.4 GB |
| Memory · Q4_0 | 17.5 GB | 14.4 GB | 6.7 GB | 4.5 GB | 2.9 GB |
| Runs on a phone | — | — | — | Yes | Yes |
| MMLU Pro | 85.2 | 82.6 | 77.2 | 69.4 | 60.0 |
| AIME 2026 | 89.2 | 88.3 | 77.5 | 42.5 | 37.5 |
| GPQA Diamond | 84.3 | 82.3 | 78.8 | 58.6 | 43.4 |
| LiveCodeBench v6 | 80.0 | 77.1 | 72.0 | 52.0 | 44.0 |
| MMMU Pro vision | 76.9 | 73.8 | 69.1 | 52.6 | 44.2 |
| RULER @128K | 96.4 | 89.8 | 91.2 | 86.6 | 70.4 |
| Speed H100 / device | 40 t/s | 177 t/s | — | 22 t/s | 52 t/s |
| With MTP | 125 t/s | 264 t/s | — | — | — |
Four generations of Gemma
Comparing flagship to flagship. The interesting part is that the jump from 3 to 4 is far larger than any before it.
| Generation | Released | Sizes | Context | Modalities | Licence |
|---|---|---|---|---|---|
| Gemma 4 | Mar 31, 2026 | E2B, E4B, 12B, 26B A4B, 31B | 256K | Text, image, audio, video | Apache 2.0 |
| Gemma 3 | Mar 10, 2025 | 270M, 1B, 4B, 12B, 27B | 128K | Text, image | Gemma Terms |
| Gemma 3n | Jun 26, 2025 | E2B, E4B | 32K | Text, image, audio | Gemma Terms |
| Gemma 2 | Jun 27, 2024 | 2B, 9B, 27B | 8K | Text | Gemma Terms |
| Gemma 1 | Feb 21, 2024 | 2B, 7B | 8K | Text | Gemma Terms |
Gemma 3 27B against Gemma 4 31B
| Benchmark | Gemma 3 27B | Gemma 4 31B | Change |
|---|---|---|---|
| AIME 2026 | 20.8 | 89.2 | +68.4 |
| Codeforces Elo | 110 | 2150 | +2040 |
| LiveCodeBench v6 | 29.1 | 80.0 | +50.9 |
| BBH | 19.3 | 74.4 | +55.1 |
| GPQA Diamond | 42.4 | 84.3 | +41.9 |
| MMLU Pro | 67.6 | 85.2 | +17.6 |
| MMMU Pro | 49.7 | 76.9 | +27.2 |
| MATH-Vision | 46.0 | 85.6 | +39.6 |
| RULER @128K | 66.0 | 96.4 | +30.4 |
| LOFT retrieval | 8.6 | 79.5 | +70.9 |
| IFEval | 90.4 | 98.9 | +8.5 |
| LMArena Elo | 1366 | 1451 | +85 |
Where the jump came from
Almost all of it is reasoning-focused post-training. AIME going from 20.8 to 89.2 and Codeforces from 110 to 2150 isn't a scaling result — the models are a similar size. It's thinking mode plus training that teaches the model to use it. The long-context numbers tell a similar story: LOFT retrieval went from a near-useless 8.6 to 79.5.
The small models moved most
Gemma 4 E2B — two billion parameters, running on a phone — scores 37.5 on AIME against Gemma 3 27B's 20.8, and 44.0 on LiveCodeBench against 29.1. A model you can fit in about a gigabyte now outperforms the previous generation's flagship on reasoning and code.
The licence divide
This is the single biggest practical difference between Gemma 4 and everything before it, and it doesn't show up in any benchmark.
Apache 2.0
- ✅ Commercial use, no conditions attached
- ✅ Modify, fine-tune and redistribute freely
- ✅ Build proprietary products on top
- ✅ No use restrictions to pass downstream
- ✅ Standard, well-understood terms your legal team already knows
- ✅ Google cannot unilaterally change the terms later
Applies to the Gemma 4 family. Google publishes it separately from the older Gemma Terms.
Gemma Terms of Use
- ⚠️ A Prohibited Use Policy applies, incorporated by reference
- ⚠️ You must pass the restrictions on to anyone you redistribute to
- ⚠️ You must supply recipients with a copy of the full agreement
- ⚠️ Google reserves the right to restrict usage remotely
- ⚠️ No use of Google trademarks or implied endorsement
- ⚠️ Terms can be updated by Google over time
Covers Gemma 1, 1.1, 2, 3 and 3n, plus CodeGemma, PaliGemma, ShieldGemma, EmbeddingGemma, FunctionGemma and others.
Should you move off Gemma 3?
Yes, clearly
If you're doing anything involving reasoning, maths, code or long context, the gap is enormous — and the Apache 2.0 licence removes downstream obligations you currently carry. For most projects this isn't a close call.
Weigh it up
If you have a heavily fine-tuned Gemma 3 that works, you'd be redoing that work. Budget for it — but note that a Gemma 4 model two sizes smaller may beat your tuned Gemma 3, which can cut serving costs enough to pay for the migration.
Maybe not yet
Some specialised variants — PaliGemma, ShieldGemma, CodeGemma — have no Gemma 4 equivalent. If you depend on one, you stay on the older terms until it's updated.
Variants, and which generation they're on
The wider family is large and unevenly updated. Most of these are still on the older licence.
| Variant | Released | Sizes | Purpose | Licence |
|---|---|---|---|---|
| DiffusionGemma | Jun 2026 | 26B A4B | Text diffusion — parallel generation, ~4× faster output | Apache 2.0 |
| TranslateGemma | Jan 15, 2026 | 4B, 12B, 27B | Dedicated machine translation | Gemma Terms |
| MedGemma 1.5 | Jan 13, 2026 | 4B | Medical imaging and clinical text | Gemma Terms |
| FunctionGemma | Dec 18, 2025 | 270M | Tiny tool-calling router | Gemma Terms |
| T5Gemma v2 | Dec 18, 2025 | 270M, 1B, 4B | Encoder-decoder tasks | Gemma Terms |
| VaultGemma | Sep 13, 2025 | 1B | Differentially private training | Gemma Terms |
| EmbeddingGemma | Sep 4, 2025 | 308M | Retrieval and RAG embeddings | Gemma Terms |
| PaliGemma 2 | Feb 19, 2025 | 3B, 10B, 28B | Vision-language | Gemma Terms |
| ShieldGemma 2 | Jul 31, 2024 | 4B | Content safety classification | Gemma Terms |
| CodeGemma | Apr 9, 2024 | 2B, 7B | Code completion — superseded by Gemma 4 | Gemma Terms |
| RecurrentGemma | Apr 9, 2024 | 2B, 9B | Griffin architecture, not a transformer | Gemma Terms |
Common questions
Which Gemma 4 model should I pick?
26B A4B for serving — near-flagship quality at 4.4× the speed. 31B when quality decides the outcome. 12B for a single consumer GPU, or for anything involving audio or video, since the larger models don't accept them. E4B and E2B for phones and edge devices.
Is the 31B actually worth it over the 26B?
Only sometimes. The accuracy gap is roughly two to three points on most benchmarks while the 31B runs at about a quarter of the speed at concurrency 1. If you're serving interactive users, the MoE is almost certainly the better trade. Reserve the 31B for offline batch work or cases where the last few points genuinely matter.
Are Gemma 3 models still supported?
They still work and remain downloadable, but they're on the older Gemma Terms of Use with its downstream obligations, and the capability gap is large. New projects should start on Gemma 4 unless they depend on a specialised variant that hasn't been updated.
What's the difference between Gemma 3n and Gemma 4 E2B/E4B?
Gemma 3n was the previous on-device family, and Gemma 4's E-series replaces it. Gemma 4 improves speech translation by 12% for E2B and 10% for E4B, and transcription by 17% and 12% — while shrinking the audio encoder from 680M to 305M parameters. Gemma 4 also brings the Apache 2.0 licence, which 3n doesn't have.
Can I still use CodeGemma or PaliGemma?
Yes, but think carefully. Both predate Gemma 4 by two generations, and plain Gemma 4 outperforms CodeGemma at coding and handles vision better than PaliGemma for most tasks. They also carry the older licence. The variants still worth using are those covering ground Gemma 4 doesn't — TranslateGemma, MedGemma, EmbeddingGemma, ShieldGemma.
Why does Gemma 4 not have a 1B or 270M model?
The E-series fills that role differently. Rather than a small dense model, E2B uses mixed 2/4/8-bit weights so its text-only footprint can drop to roughly 0.8 GB — comparable in memory to a tiny model, but far more capable. If you need something genuinely tiny for a narrow job, FunctionGemma at 270M or EmbeddingGemma at 308M are the specialised options.
Does the licence change apply retroactively to Gemma 3?
No. Google's Gemma Terms of Use still lists Gemma 1, 1.1, 2, 3 and 3n along with most specialised variants, and points to a separate Apache 2.0 licence for Gemma 4 only. Older models stay on the older terms.
Picked one?
Exact repository IDs and pull commands for every checkpoint in the family.