Gemma 4 Specs
Every published number for the Gemma 4 family in one place - architecture, memory, benchmarks, throughput, on-device performance and provenance. This is the reference the other pages link to rather than restate.
Every figure is attributed in provenance - self-reported and independent are marked separately
Model specifications
| Specification | 31B | 26B A4B | 12B | E4B | E2B |
|---|---|---|---|---|---|
| Total parameters | 31B | 26B | 11.95B | 4.5B | 2.3B |
| Active per token | 31B | ~3.8B | 11.95B | 4.5B | 2.3B |
| Architecture | Dense | MoE | Unified | Dense edge | Dense edge |
| Context window | 256K | 256K | 256K | 128K | 128K |
| Text input | β | β | β | β | β |
| Image input | β | β | β | β | β |
| Audio input | - | - | β | β | β |
| Video input | - | - | β | β | β |
| Output | Text only - no model in the family generates images or audio | ||||
| Thinking mode | β | β | β | β | β |
| Function calling | β | β | β | β | β |
| Runs on a phone | - | - | - | β | β |
| MTP drafter published | β | β | β | β | β |
Weights by precision
Weights only. Add roughly 20β30% for KV cache at moderate context; more as context and concurrency grow.
| Model | BF16 | SFP8 | Q4_0 | Mobile | Context |
|---|---|---|---|---|---|
| Gemma 4 31B | 69.9 GB | 34.9 GB | 17.5 GB | - | 256K |
| Gemma 4 26B A4B | 57.7 GB | 28.8 GB | 14.4 GB | - | 256K |
| Gemma 4 12B | 26.7 GB | 13.4 GB | 6.7 GB | - | 256K |
| Gemma 4 E4B | 17.9 GB | 8.9 GB | 4.5 GB | 2.5 GB | 128K |
| Gemma 4 E2B | 11.4 GB | 5.7 GB | 2.9 GB | 1.1 GB | 128K |
E2B's text-only footprint drops to roughly 0.8 GB with mixed 2/4/8-bit weights. Community Q4_K_M quants run about 10% larger than the official Q4_0 figures.
Implementation detail
Explained in context on the technical page - this is the reference table.
| Specification | 31B | 26B A4B | 12B | E4B | E2B |
|---|---|---|---|---|---|
| Base architecture | Decoder-only Transformer | ||||
| Normalisation | RMSNorm, pre + post Β· QKNorm on queries and keys | ||||
| Tokenizer | SentencePiece, 262,144 entries Β· split digits Β· preserved whitespace Β· byte fallback | ||||
| Local:global attention | 5:1 | 5:1 | 5:1 | 5:1 | 4:1 |
| Global positional | pp-RoPE, p = 0.25, base frequency 1M | ||||
| Local positional | Standard RoPE, base frequency 10k | ||||
| Global KV reduction | 37.5% - keys reused as values in global attention layers | ||||
| KV cache sharing | - | - | - | 18/42 | 20/35 |
| Vision path | 550M ViT | 550M ViT | 35M projection | 150M ViT p16 | 150M ViT p16 |
| Vision patch size | - | - | 48Γ48Γ3 RGB | 16 | 16 |
| Vision token budget | 70 Β· 140 Β· 280 (default) Β· 560 Β· 1120 | ||||
| Audio path | - | - | none - raw projection | 305M USM | 305M USM |
| Audio frame | - | - | 40 ms @ 16 kHz β 640-dim | 40 ms Mel | 40 ms Mel |
| MTP drafter layers | 4 | 4 | 4 | 4 | 4 |
| MTP drafter dim | 1024 | 1024 | ~0.4B params | 256 | 256 |
| MTP vocab projection | Clustered from 262k to 4,096 - roughly 64Γ reduction | ||||
| Training hardware | TPUv5p and TPUv6e Β· ZeRO-3 sharding Β· Pathways Β· JAX / GSPMD / MegaScale XLA | ||||
| Data cutoff | January 2025 | ||||
Reasoning and code
Instruction-tuned checkpoints with thinking mode enabled. Gemma 3 27B included as the generational baseline.
| Benchmark | 31B | 26B A4B | 12B | E4B | E2B | Gemma 3 27B |
|---|---|---|---|---|---|---|
| MMLU Pro | 85.2 | 82.6 | 77.2 | 69.4 | 60.0 | 67.6 |
| AIME 2026 no tools | 89.2 | 88.3 | 77.5 | 42.5 | 37.5 | 20.8 |
| GPQA Diamond | 84.3 | 82.3 | 78.8 | 58.6 | 43.4 | 42.4 |
| LiveCodeBench v6 | 80.0 | 77.1 | 72.0 | 52.0 | 44.0 | 29.1 |
| Codeforces Elo | 2150 | 1718 | 1659 | 940 | 633 | 110 |
| BBH micro avg | 74.4 | 64.8 | 53.0 | 33.1 | 21.9 | 19.3 |
| SciCode | 43.0 | 40.0 | 38.0 | 24.0 | 21.0 | 21.0 |
| IFEval | 98.9 | 98.5 | 97.2 | 96.7 | 94.6 | 90.4 |
| IFBench | 76.0 | 72.0 | 74.0 | 44.0 | 38.0 | 32.0 |
| HLE | 19.5 | 8.7 | 5.2 | - | - | - |
| HLE with search | 26.5 | 17.2 | - | - | - | - |
Human preference Β· LMArena text
| Model | Elo | Rank | Category | As of |
|---|---|---|---|---|
| Gemma 4 31B | 1451 Β± 8 | 43 | Leading open dense model | Jun 19, 2026 |
| Gemma 4 26B A4B | 1438 Β± 8 | 61 | Open MoE | Jun 19, 2026 |
| Gemma 3 27B | 1366 Β± 4 | 157 | Previous generation | Jun 19, 2026 |
At launch LMArena placed the 31B at #3 among open models and #27 overall. Arena ranks decay as new models arrive - always read them with a date attached.
Multimodal scores
Headline figures use the maximum 1120-token budget. The current default is 280 - the second table shows what that costs.
| Benchmark Β· 1120 tokens | 31B | 26B A4B | 12B | E4B | E2B | Gemma 3 27B |
|---|---|---|---|---|---|---|
| MMMU Pro | 76.9 | 73.8 | 69.1 | 52.6 | 44.2 | 49.7 |
| MATH-Vision | 85.6 | 82.4 | 79.7 | 59.5 | 52.4 | 46.0 |
| InfographicVQA | 92.0 | 89.3 | 88.4 | 70.0 | 63.9 | 70.6 |
| MedXpertQA MM | 61.3 | 58.1 | 48.7 | 28.7 | 23.5 | - |
| OmniDocBench 1.5 lower is better β | 0.131 | 0.149 | 0.164 | 0.181 | 0.290 | 0.365 |
| Benchmark Β· 280 tokens default | 31B | 26B A4B | 12B | E4B | E2B |
|---|---|---|---|---|---|
| MMMU Pro | 75.8 | 73.2 | 67.7 | 51.4 | 43.2 |
| MATH-Vision | 83.4 | 80.3 | 76.7 | 59.2 | 53.0 |
| InfographicVQA | 82.8 | 77.8 | 58.7 | 54.8 | 44.6 |
| MedXpertQA MM | 60.7 | 55.7 | 47.4 | 28.7 | 22.5 |
Document-heavy tasks are the most sensitive: InfographicVQA falls from 92.0 to 82.8 on the 31B, and from 88.4 to 58.7 on the 12B.
Retrieval and reasoning across the window
Measured without thinking mode.
| Benchmark | Length | 31B | 26B A4B | 12B | E4B | E2B | Gemma 3 27B |
|---|---|---|---|---|---|---|---|
| RULER | 32K | 96.8 | 97.3 | 96.4 | 95.2 | 83.0 | 91.1 |
| RULER | 128K | 96.4 | 89.8 | 91.2 | 86.6 | 70.4 | 66.0 |
| LOFT Recall@k | 128K | 79.5 | 66.3 | 66.4 | 58.5 | 50.5 | 8.6 |
| GraphWalks F1 | <128K | 82.3 | 72.6 | 71.0 | 50.9 | 4.1 | 32.8 |
| MTOB engβkgv | 128K | 52.9 | 50.0 | 45.1 | 37.8 | 15.4 | 41.0 |
| MTOB engβkgv | 256K | 54.3 | 48.9 | 41.9 | - | - | - |
Transcription and speech translation
| Benchmark | 12B | E4B | E2B | 3n E4B | 3n E2B |
|---|---|---|---|---|---|
| FLEURS word error rate, lower better β | 0.063 | 0.075 | 0.090 | 0.085 | 0.108 |
| CoVoST CorpusBLEU into English | 42.3 | 38.2 | 35.4 | 34.7 | 31.6 |
Against Gemma 3n at matching sizes: translation improved 12% (E2B) and 10% (E4B), transcription 17% and 12% - while the audio encoder shrank from 680M to 305M parameters (390 MB β 87 MB on disk). The 31B and 26B A4B accept no audio at all.
Throughput and latency
vLLM on H100 80GB. Independent measurement, not vendor-reported.
| Measurement | 31B dense | 26B A4B MoE | Conditions |
|---|---|---|---|
| Decode, baseline | 40.3 t/s | 177.1 t/s | Concurrency 1 |
| Decode, with MTP | 125.3 t/s | 264.2 t/s | Concurrency 1 |
| MTP speedup | 3.11Γ | 1.49Γ | Dense gains more - it was more bandwidth-bound |
| Aggregate throughput | 1,260 t/s | - | 1Γ H100, ISL 512 / OSL 256 |
| Aggregate, 8Γ H100 | 2,208 t/s | - | 1.75Γ for 8Γ hardware - sublinear |
| Time to first token | 279 ms | - | 1Γ H100, concurrency 1, 128-token prompt |
| TTFT, 8Γ H100 | 85 ms | - | Same prompt |
| TTFT stable to | ~concurrency 8 | - | Climbs after, as the prefill queue builds |
| Weights resident | 62 GB | - | BF16, leaving ~18 GB of an 80 GB card for KV cache |
FlashAttention 4 on Hopper GPUs, shipped July 2026: prefill throughput +25β70%, time-to-first-token down as much as 31%.
Mobile performance
LiteRT-LM builds. The GPU-versus-CPU gap is the single largest configuration effect on this page.
| Device Β· backend | Model | Prefill | Decode | First token | Memory |
|---|---|---|---|---|---|
| Galaxy S26 Ultra Β· GPU | E2B | 3,808 t/s | 52.1 t/s | 0.3 s | 676 MB |
| Galaxy S26 Ultra Β· CPU | E2B | 557 t/s | 46.9 t/s | 1.8 s | 1,733 MB |
| Galaxy S26 Ultra Β· GPU | E4B | 1,293 t/s | 22.1 t/s | 0.8 s | 710 MB |
| Galaxy S26 Ultra Β· CPU | E4B | 195 t/s | 17.7 t/s | 5.3 s | 3,283 MB |
| iPhone 17 Pro Β· GPU | E4B | 1,189 t/s | 25.1 t/s | 0.9 s | - |
| Dragonwing IQ8 Β· NPU | E2B | 3,700 t/s | 31 t/s | - | - |
| Raspberry Pi 5 Β· CPU | E2B | 133 t/s | 7.6 t/s | - | - |
Mobile file sizes: E2B 2.58 GB (web variant 2.0 GB, Qualcomm NPU build 2.97 GB); E4B 3.66 GB (2.24 GB decoder plus 0.67 GB memory-mapped embeddings).
| Checkpoint | Format | Available for | Purpose |
|---|---|---|---|
-it | Safetensors BF16 | All five | Instruction-tuned - the default |
| no suffix | Safetensors BF16 | All five | Base, pre-trained only |
-it-qat-q4_0-gguf | GGUF 4-bit | All five | Quantisation-aware trained 4-bit |
-it-assistant | Safetensors | All five | MTP drafter, ~0.4B - not a chat model |
-it-qat-q4_0-unquantized | Safetensors | Some sizes | QAT weights at full precision |
-it-litert-lm | .litertlm | E2B, E4B only | Mobile, via litert-community/ |
| Release | Date | Contents |
|---|---|---|
| Core release | March 31, 2026 | E2B, E4B, 26B A4B, 31B |
| Public announcement | April 2, 2026 | - |
| MTP checkpoints | April 16, 2026 | E2B, E4B, 26B A4B, 31B drafters |
| 12B Unified | June 3, 2026 | Encoder-free multimodal |
| Weights refresh | July 15, 2026 | FlashAttention 4, tool-calling fixes, vision defaults, chat template - no version bump |
Licence: Apache 2.0 for the entire Gemma 4 family. Gemma 3 and earlier, and most specialised variants including ShieldGemma, remain under the Gemma Terms of Use.
Where each number comes from
Worth knowing which figures are vendor-reported and which are independent, because they don't tell the same story.
| Data | Source | Type | Caveat |
|---|---|---|---|
| Architecture, memory, benchmarks | Gemma 4 Technical Report | Self-reported | Thinking mode on; vision at 1120 tokens |
| Modalities, context, model IDs | Google model documentation | Self-reported | - |
| LMArena Elo | LMArena | Independent | Ranks decay as new models arrive - dated June 19, 2026 |
| Throughput, TTFT, MTP speedups | Third-party vLLM benchmarking | Independent | One workload shape; yours will differ |
| On-device performance | LiteRT community builds | Independent | Flagship hardware - treat as a ceiling |
| Agentic and SWE-Rebench results | Third-party aggregation | Independent | Substantially less flattering - see limitations |
Numbers are one view.
See what the architecture actually does, or where the measurements stop being flattering.