Canonical datasheet Updated August 2026

Gemma 4 Specs

Every published number for the Gemma 4 family in one place - architecture, memory, benchmarks, throughput, on-device performance and provenance. This is the reference the other pages link to rather than restate.

Every figure is attributed in provenance - self-reported and independent are marked separately

Family at a glance
Models5 sizes
Parameters2.3B – 31B
Context128K / 256K
Vocabulary262,144
Languages140+
Data cutoffJanuary 2025
LicenceApache 2.0
Released March 31, 2026 Β· 12B added June 3 Β· refreshed July 15
Core

Model specifications

Specification31B26B A4B12B E4BE2B
Total parameters31B26B 11.95B4.5B2.3B
Active per token31B~3.8B 11.95B4.5B2.3B
ArchitectureDenseMoE UnifiedDense edgeDense edge
Context window256K256K 256K128K128K
Text inputβœ“βœ“ βœ“βœ“βœ“
Image inputβœ“βœ“ βœ“βœ“βœ“
Audio input-- βœ“βœ“βœ“
Video input-- βœ“βœ“βœ“
OutputText only - no model in the family generates images or audio
Thinking modeβœ“βœ“ βœ“βœ“βœ“
Function callingβœ“βœ“ βœ“βœ“βœ“
Runs on a phone-- -βœ“βœ“
MTP drafter publishedβœ“βœ“ βœ“βœ“βœ“
Memory

Weights by precision

Weights only. Add roughly 20–30% for KV cache at moderate context; more as context and concurrency grow.

ModelBF16SFP8Q4_0 MobileContext
Gemma 4 31B69.9 GB34.9 GB 17.5 GB-256K
Gemma 4 26B A4B57.7 GB28.8 GB 14.4 GB-256K
Gemma 4 12B26.7 GB 13.4 GB6.7 GB-256K
Gemma 4 E4B17.9 GB8.9 GB 4.5 GB2.5 GB128K
Gemma 4 E2B11.4 GB5.7 GB 2.9 GB1.1 GB128K

E2B's text-only footprint drops to roughly 0.8 GB with mixed 2/4/8-bit weights. Community Q4_K_M quants run about 10% larger than the official Q4_0 figures.

Architecture

Implementation detail

Explained in context on the technical page - this is the reference table.

Specification31B26B A4B12B E4BE2B
Base architecture Decoder-only Transformer
Normalisation RMSNorm, pre + post Β· QKNorm on queries and keys
Tokenizer SentencePiece, 262,144 entries Β· split digits Β· preserved whitespace Β· byte fallback
Local:global attention5:15:1 5:15:14:1
Global positional pp-RoPE, p = 0.25, base frequency 1M
Local positional Standard RoPE, base frequency 10k
Global KV reduction 37.5% - keys reused as values in global attention layers
KV cache sharing-- -18/4220/35
Vision path550M ViT550M ViT 35M projection150M ViT p16150M ViT p16
Vision patch size-- 48Γ—48Γ—3 RGB1616
Vision token budget 70 Β· 140 Β· 280 (default) Β· 560 Β· 1120
Audio path-- none - raw projection305M USM305M USM
Audio frame-- 40 ms @ 16 kHz β†’ 640-dim40 ms Mel40 ms Mel
MTP drafter layers44 444
MTP drafter dim10241024 ~0.4B params256256
MTP vocab projection Clustered from 262k to 4,096 - roughly 64Γ— reduction
Training hardware TPUv5p and TPUv6e Β· ZeRO-3 sharding Β· Pathways Β· JAX / GSPMD / MegaScale XLA
Data cutoffJanuary 2025
Text

Reasoning and code

Instruction-tuned checkpoints with thinking mode enabled. Gemma 3 27B included as the generational baseline.

Benchmark31B26B A4B12B E4BE2BGemma 3 27B
MMLU Pro85.282.677.269.460.067.6
AIME 2026 no tools89.288.377.542.537.520.8
GPQA Diamond84.382.378.858.643.442.4
LiveCodeBench v680.077.172.052.044.029.1
Codeforces Elo215017181659940633110
BBH micro avg74.464.853.033.121.919.3
SciCode43.040.038.024.021.021.0
IFEval98.998.597.296.794.690.4
IFBench76.072.074.044.038.032.0
HLE19.58.75.2---
HLE with search26.517.2----

Human preference Β· LMArena text

ModelEloRankCategoryAs of
Gemma 4 31B1451 Β± 843Leading open dense modelJun 19, 2026
Gemma 4 26B A4B1438 Β± 861Open MoEJun 19, 2026
Gemma 3 27B1366 Β± 4157Previous generationJun 19, 2026

At launch LMArena placed the 31B at #3 among open models and #27 overall. Arena ranks decay as new models arrive - always read them with a date attached.

Vision

Multimodal scores

Headline figures use the maximum 1120-token budget. The current default is 280 - the second table shows what that costs.

Benchmark Β· 1120 tokens31B26B A4B12B E4BE2BGemma 3 27B
MMMU Pro76.973.869.152.644.249.7
MATH-Vision85.682.479.759.552.446.0
InfographicVQA92.089.388.470.063.970.6
MedXpertQA MM61.358.148.728.723.5-
OmniDocBench 1.5 lower is better ↓0.1310.1490.1640.1810.2900.365
Benchmark Β· 280 tokens default31B26B A4B12BE4BE2B
MMMU Pro75.873.267.751.443.2
MATH-Vision83.480.376.759.253.0
InfographicVQA82.877.858.754.844.6
MedXpertQA MM60.755.747.428.722.5

Document-heavy tasks are the most sensitive: InfographicVQA falls from 92.0 to 82.8 on the 31B, and from 88.4 to 58.7 on the 12B.

Long context

Retrieval and reasoning across the window

Measured without thinking mode.

BenchmarkLength31B26B A4B 12BE4BE2BGemma 3 27B
RULER32K96.897.396.495.283.091.1
RULER128K96.489.891.286.670.466.0
LOFT Recall@k128K79.566.366.458.550.58.6
GraphWalks F1<128K82.372.671.050.94.132.8
MTOB eng→kgv128K52.950.045.137.815.441.0
MTOB eng→kgv256K54.348.941.9---
Audio

Transcription and speech translation

Benchmark12BE4BE2B 3n E4B3n E2B
FLEURS word error rate, lower better ↓ 0.0630.0750.090 0.0850.108
CoVoST CorpusBLEU into English 42.338.235.4 34.731.6

Against Gemma 3n at matching sizes: translation improved 12% (E2B) and 10% (E4B), transcription 17% and 12% - while the audio encoder shrank from 680M to 305M parameters (390 MB β†’ 87 MB on disk). The 31B and 26B A4B accept no audio at all.

Serving

Throughput and latency

vLLM on H100 80GB. Independent measurement, not vendor-reported.

Measurement31B dense26B A4B MoEConditions
Decode, baseline40.3 t/s177.1 t/sConcurrency 1
Decode, with MTP125.3 t/s264.2 t/sConcurrency 1
MTP speedup3.11Γ—1.49Γ—Dense gains more - it was more bandwidth-bound
Aggregate throughput1,260 t/s-1Γ— H100, ISL 512 / OSL 256
Aggregate, 8Γ— H1002,208 t/s-1.75Γ— for 8Γ— hardware - sublinear
Time to first token279 ms-1Γ— H100, concurrency 1, 128-token prompt
TTFT, 8Γ— H10085 ms-Same prompt
TTFT stable to~concurrency 8-Climbs after, as the prefill queue builds
Weights resident62 GB-BF16, leaving ~18 GB of an 80 GB card for KV cache

FlashAttention 4 on Hopper GPUs, shipped July 2026: prefill throughput +25–70%, time-to-first-token down as much as 31%.

On-device

Mobile performance

LiteRT-LM builds. The GPU-versus-CPU gap is the single largest configuration effect on this page.

Device Β· backendModelPrefillDecode First tokenMemory
Galaxy S26 Ultra Β· GPUE2B 3,808 t/s52.1 t/s0.3 s676 MB
Galaxy S26 Ultra Β· CPUE2B 557 t/s46.9 t/s1.8 s1,733 MB
Galaxy S26 Ultra Β· GPUE4B 1,293 t/s22.1 t/s0.8 s710 MB
Galaxy S26 Ultra Β· CPUE4B 195 t/s17.7 t/s5.3 s3,283 MB
iPhone 17 Pro Β· GPUE4B 1,189 t/s25.1 t/s0.9 s-
Dragonwing IQ8 Β· NPUE2B 3,700 t/s31 t/s--
Raspberry Pi 5 Β· CPUE2B 133 t/s7.6 t/s--

Mobile file sizes: E2B 2.58 GB (web variant 2.0 GB, Qualcomm NPU build 2.97 GB); E4B 3.66 GB (2.24 GB decoder plus 0.67 GB memory-mapped embeddings).

Distribution

Checkpoints and licensing

Each suffix explained on the variants page.

CheckpointFormatAvailable forPurpose
-itSafetensors BF16All fiveInstruction-tuned - the default
no suffixSafetensors BF16All fiveBase, pre-trained only
-it-qat-q4_0-ggufGGUF 4-bitAll fiveQuantisation-aware trained 4-bit
-it-assistantSafetensorsAll fiveMTP drafter, ~0.4B - not a chat model
-it-qat-q4_0-unquantizedSafetensorsSome sizesQAT weights at full precision
-it-litert-lm.litertlmE2B, E4B onlyMobile, via litert-community/
ReleaseDateContents
Core releaseMarch 31, 2026E2B, E4B, 26B A4B, 31B
Public announcementApril 2, 2026-
MTP checkpointsApril 16, 2026E2B, E4B, 26B A4B, 31B drafters
12B UnifiedJune 3, 2026Encoder-free multimodal
Weights refreshJuly 15, 2026 FlashAttention 4, tool-calling fixes, vision defaults, chat template - no version bump

Licence: Apache 2.0 for the entire Gemma 4 family. Gemma 3 and earlier, and most specialised variants including ShieldGemma, remain under the Gemma Terms of Use.

Provenance

Where each number comes from

Worth knowing which figures are vendor-reported and which are independent, because they don't tell the same story.

DataSourceTypeCaveat
Architecture, memory, benchmarksGemma 4 Technical Report Self-reportedThinking mode on; vision at 1120 tokens
Modalities, context, model IDsGoogle model documentation Self-reported-
LMArena EloLMArenaIndependent Ranks decay as new models arrive - dated June 19, 2026
Throughput, TTFT, MTP speedupsThird-party vLLM benchmarking IndependentOne workload shape; yours will differ
On-device performanceLiteRT community buildsIndependent Flagship hardware - treat as a ceiling
Agentic and SWE-Rebench resultsThird-party aggregation IndependentSubstantially less flattering - see limitations
πŸ“…
Two things that quietly invalidate comparisons. The July 15, 2026 refresh changed weights without changing the version number, so results measured before and after aren't strictly comparable. And every benchmark here assumes thinking mode enabled and vision at maximum resolution - neither of which is the default configuration you'll get out of the box.

Numbers are one view.

See what the architecture actually does, or where the measurements stop being flattering.