Capabilities &
Performance
What Gemma 4 can actually do - modalities, thinking, tool calling, agent skills, long context - and how fast it does it on real hardware. Benchmark scores live on their own page; this one is about behaviour and throughput.
For accuracy scores - MMLU, AIME, GPQA - see the benchmarks page
40.3 → 125.3 tokens/sec at concurrency 1
What each model can take in
A common misconception: the biggest model is not the most multimodal. Audio and video live in the smaller models, because that's where the encoder-free work landed first.
| Model | Text | Image | Audio | Video | Context | Output |
|---|---|---|---|---|---|---|
| Gemma 4 31B dense | ✓ | ✓ | - | - | 256K | Text |
| Gemma 4 26B A4B MoE | ✓ | ✓ | - | - | 256K | Text |
| Gemma 4 12B unified | ✓ | ✓ | ✓ | ✓ | 256K | Text |
| Gemma 4 E4B edge | ✓ | ✓ | ✓ | ✓ | 128K | Text |
| Gemma 4 E2B on-device | ✓ | ✓ | ✓ | ✓ | 128K | Text |
What it can do
Six capabilities that change how you'd design around the model, rather than a feature list.
Configurable thinking
The model can emit reasoning traces before answering, and the depth is a parameter rather than a fixed behaviour. Every accuracy figure Google publishes assumes thinking is on - so this is the single setting with the biggest effect on both quality and cost.
Design for it: minimal thinking for classification, routing and extraction where latency dominates. High for maths, code and anything multi-step. Don't leave it on high by default for chat - you'll pay for tokens the user never sees.
Tool calling with constrained decoding
Function calling is trained in, and the output is produced under constrained decoding - the sampler is restricted to tokens that keep the output schema-valid. That's a stronger guarantee than prompting a model to "respond in JSON" and hoping.
Caveat: structurally valid is not the same as semantically correct. The July 2026 refresh substantially improved which tools the model chooses; agentic multi-step reliability remains its weakest area, so keep chains short and verify.
Agent Skills
Skills are callable capabilities the model composes into workflows - querying a knowledge source, transforming content into summaries or visualisations, or handing off to a text-to-speech or image model. The notable part is that this runs locally: an edge device handled 4,000 input tokens across two distinct skills in under three seconds.
Long context that holds
Retrieval accuracy barely degrades from 32K to 128K on the 31B, and the MTOB score actually improves at 256K over 128K - the extra context is being used, not merely accepted.
Caveat: this holds for the large models. On E2B, multi-hop reasoning across long context fails badly. Long context for retrieval is fine there; long context for reasoning is not.
140+ languages
Multilingual coverage is built into the tokenizer rather than bolted on, which keeps token counts reasonable in non-Latin scripts - this directly affects your cost and effective context length. The edge models also transcribe and translate speech.
Adjustable vision resolution
Images are processed at a token budget you choose: 70, 140, 280, 560 or 1120. Low budgets are cheap and fine for "what's in this picture"; dense documents and screenshots need the high end. The default is 280, and OCR-style accuracy climbs steeply above it.
The four design choices that matter
Architecture is usually trivia. These four have direct, measurable consequences for what you can run and how fast.
MoE routing
The 26B activates only ~3.8B parameters per token. In practice that made it 4.4× faster than the dense 31B at concurrency 1 - 177 tokens/sec against 40 - at close to the same quality.
Encoder-free 12B
Image patches and raw audio project straight into the token space with no separate encoders. Fewer moving parts, less memory, and audio in a mid-sized model that fits a 16 GB laptop.
pp-RoPE positional encoding
Cuts KV cache footprint by up to 37.5%. That's what makes 256K context practical on hardware you can afford rather than merely supported on paper.
Multi-Token Prediction
A draft head that shares the model's own state proposes several tokens at once, verified in a single pass. No separate draft model to host - see the numbers below.
Real throughput on real hardware
Measured with vLLM on H100 80GB. Numbers move with your workload shape, but the relationships between them hold.
Multi-Token Prediction is the biggest single win
If you serve Gemma 4 and aren't using the MTP checkpoints, you're leaving a large amount of performance unclaimed - and it costs nothing in quality.
Gemma 4 31B dense
Gemma 4 26B A4B MoE
Bars are scaled to the same axis across both charts so the models are directly comparable.
Why the dense model gains more
Single-stream decoding on a dense model is memory-bandwidth bound - the GPU spends most of its time moving 62 GB of weights, not computing. Verifying several speculated tokens in one pass amortises that traffic, so the 31B nearly triples. The MoE was already reading far fewer weights per token, so there's less waste to reclaim, and it gains a more modest 1.49×.
What to actually do
Pull the MTP checkpoints for whichever size you serve. Acceptance rates fall off at deeper draft positions - most of the benefit comes from the first few - so there's no advantage in pushing the speculation depth high. And note these gains are measured at concurrency 1; under heavy batching the GPU is compute-bound and speculative decoding helps far less.
Other levers worth pulling
FlashAttention 4
Shipped in the July 2026 refresh for NVIDIA Hopper GPUs: prefill throughput up 25–70% and time to first token down as much as 31%. Most valuable when prompts are long - agent system prompts and tool definitions are exactly that.
QAT 4-bit weights
Google publishes quantisation-aware-trained 4-bit checkpoints, which lose far less quality than post-hoc quantisation. Roughly a quarter of the memory of BF16 - the difference between the 31B fitting on one 24 GB card and not.
Turn thinking down
Reasoning traces are generated tokens you pay for and the user never reads. For classification, routing and extraction, minimal thinking can cut latency dramatically with little accuracy cost on those task types.
Lower the vision budget
Dropping from 1120 to 280 vision tokens per image is a 4× reduction in prefill work. On general image questions the accuracy cost is about a point; on dense documents it's much larger, so tune per task rather than globally.
Prefer MoE for serving
The 26B A4B was 4.4× faster than the dense 31B at concurrency 1 for a small quality difference. For most production workloads that's the better trade - reserve the 31B for cases where quality genuinely decides the outcome.
Right-size the model
The cheapest optimisation is a smaller model that's still good enough. The 12B holds much of the 31B's quality at a fraction of the memory - and on many real workloads the gap is smaller than the benchmark spread implies.
Performance on a phone
The same capability set, four orders of magnitude less hardware.
Which model for which job
| If your workload is… | Use | Because |
|---|---|---|
| High-volume serving | 26B A4B + MTP |
264 tok/s at concurrency 1, close to 31B quality, ~3.8B active params |
| Hardest reasoning, quality decides | 31B + MTP |
Best scores in the family; MTP recovers most of the dense speed penalty |
| Anything with audio or video | 12B |
The largest model that accepts them at all |
| One consumer GPU or a laptop | 12B QAT |
6.7 GB at 4-bit, full 256K context, strong quality per gigabyte |
| Phone, offline, private | E2B GPU backend |
676 MB, 0.3 s to first token, no network at all |
| Long documents, retrieval | 31B or 26B A4B |
Retrieval holds to 128K+; small models degrade sharply on multi-hop |
| Multi-step agents | Post-July weights, short chains | Weakest area - verify each step rather than trusting long chains |
Common questions
Why is the 26B MoE faster than the 31B dense model?
A mixture-of-experts model routes each token to a subset of its parameters - about 3.8B of the 26B activate per token. Single-stream decoding is limited by how fast the GPU can read weights from memory, so reading far fewer weights per token makes it dramatically faster: 177 tokens/sec against the dense 31B's 40, at concurrency 1. The full 26B still has to fit in memory; you save on bandwidth, not capacity.
Should I always use the MTP checkpoints?
For low-concurrency and interactive workloads, yes - a 3.11× speedup on the 31B with no quality change is as close to free as optimisation gets. Under heavy batching the benefit shrinks, because the GPU becomes compute-bound rather than bandwidth-bound and there's less idle capacity for speculation to exploit.
Does thinking mode change speed a lot?
Substantially, because reasoning traces are real generated tokens with real cost. Every published accuracy figure assumes thinking is on, so turning it down trades measurable accuracy for latency - worth it for classification and extraction, rarely worth it for maths or code.
How much quality do I lose at 4-bit?
Less than you'd expect with Google's QAT checkpoints, because quantisation is part of training rather than applied afterwards. Community post-hoc 4-bit quants lose more. Below 4-bit the degradation becomes severe - a smaller model at 4-bit generally beats a larger one at 3-bit.
Can Gemma 4 generate images or audio?
No. Every model in the family outputs text only. Images and audio are inputs on the models that support them. For generation you'd pair Gemma with a separate model - which is one of the things Agent Skills is designed to coordinate.
Is the 256K context genuinely usable?
On the large models, yes for retrieval - the 31B's RULER accuracy barely moves from 32K to 128K, and MTOB improves at 256K over 128K. Long context also costs memory in the KV cache, though pp-RoPE reduces that by up to 37.5%. On E2B, treat long context as retrieval-only; multi-hop reasoning across it fails.
What's the fastest configuration overall?
For a single stream: 26B A4B with MTP checkpoints, FlashAttention 4 on a Hopper GPU, thinking minimised, and the smallest vision token budget your task tolerates. For total throughput across many users, batch aggressively and expect speculative decoding to matter less.
Capabilities are one half. Scores are the other.
See how these capabilities translate into measured accuracy across all five model sizes.