Capabilities &
Performance

What Gemma 4 can actually do - modalities, thinking, tool calling, agent skills, long context - and how fast it does it on real hardware. Benchmark scores live on their own page; this one is about behaviour and throughput.

For accuracy scores - MMLU, AIME, GPQA - see the benchmarks page

Speed at a glance · single H100
3.11×
faster decode on 31B with Multi-Token Prediction
40.3 → 125.3 tokens/sec at concurrency 1
26B MoE + MTP
264
26B MoE
177
31B + MTP
125
31B dense
40
tokens/sec, concurrency 1 · vLLM on H100 80GB
Inputs

What each model can take in

A common misconception: the biggest model is not the most multimodal. Audio and video live in the smaller models, because that's where the encoder-free work landed first.

ModelTextImageAudio VideoContextOutput
Gemma 4 31B dense ✓✓-- 256KText
Gemma 4 26B A4B MoE ✓✓-- 256KText
Gemma 4 12B unified ✓✓✓✓ 256KText
Gemma 4 E4B edge ✓✓✓✓ 128KText
Gemma 4 E2B on-device ✓✓✓✓ 128KText
🎧
If you need audio, the 12B is the ceiling. The 31B and 26B A4B take text and images only. So a task like "transcribe this meeting and summarise it" is a 12B job, not a 31B one - which is unusual enough to catch people out when they assume bigger means more capable. Every model outputs text only; none generate images or speech.
Capabilities

What it can do

Six capabilities that change how you'd design around the model, rather than a feature list.

🧠

Configurable thinking

Latency you can trade for accuracy

The model can emit reasoning traces before answering, and the depth is a parameter rather than a fixed behaviour. Every accuracy figure Google publishes assumes thinking is on - so this is the single setting with the biggest effect on both quality and cost.

Design for it: minimal thinking for classification, routing and extraction where latency dominates. High for maths, code and anything multi-step. Don't leave it on high by default for chat - you'll pay for tokens the user never sees.

🛠️

Tool calling with constrained decoding

Structurally valid, not just usually valid

Function calling is trained in, and the output is produced under constrained decoding - the sampler is restricted to tokens that keep the output schema-valid. That's a stronger guarantee than prompting a model to "respond in JSON" and hoping.

Caveat: structurally valid is not the same as semantically correct. The July 2026 refresh substantially improved which tools the model chooses; agentic multi-step reliability remains its weakest area, so keep chains short and verify.

🤖

Agent Skills

Multi-step workflows, on-device

Skills are callable capabilities the model composes into workflows - querying a knowledge source, transforming content into summaries or visualisations, or handing off to a text-to-speech or image model. The notable part is that this runs locally: an edge device handled 4,000 input tokens across two distinct skills in under three seconds.

📜

Long context that holds

256K on the large models, 128K on edge

Retrieval accuracy barely degrades from 32K to 128K on the 31B, and the MTOB score actually improves at 256K over 128K - the extra context is being used, not merely accepted.

Caveat: this holds for the large models. On E2B, multi-hop reasoning across long context fails badly. Long context for retrieval is fine there; long context for reasoning is not.

🌍

140+ languages

262K-entry vocabulary

Multilingual coverage is built into the tokenizer rather than bolted on, which keeps token counts reasonable in non-Latin scripts - this directly affects your cost and effective context length. The edge models also transcribe and translate speech.

🖼️

Adjustable vision resolution

70 to 1120 vision tokens

Images are processed at a token budget you choose: 70, 140, 280, 560 or 1120. Low budgets are cheap and fine for "what's in this picture"; dense documents and screenshots need the high end. The default is 280, and OCR-style accuracy climbs steeply above it.

Under the hood

The four design choices that matter

Architecture is usually trivia. These four have direct, measurable consequences for what you can run and how fast.

🧩

MoE routing

The 26B activates only ~3.8B parameters per token. In practice that made it 4.4× faster than the dense 31B at concurrency 1 - 177 tokens/sec against 40 - at close to the same quality.

🔗

Encoder-free 12B

Image patches and raw audio project straight into the token space with no separate encoders. Fewer moving parts, less memory, and audio in a mid-sized model that fits a 16 GB laptop.

📐

pp-RoPE positional encoding

Cuts KV cache footprint by up to 37.5%. That's what makes 256K context practical on hardware you can afford rather than merely supported on paper.

⚡

Multi-Token Prediction

A draft head that shares the model's own state proposes several tokens at once, verified in a single pass. No separate draft model to host - see the numbers below.

Serving performance

Real throughput on real hardware

Measured with vLLM on H100 80GB. Numbers move with your workload shape, but the relationships between them hold.

279 ms
Time to first token, 31B on one H100 at concurrency 1 with a 128-token prompt. Drops to 85 ms across 8 GPUs.
1,260 t/s
Aggregate throughput on a single H100 for a 512-in / 256-out workload. Eight GPUs reach 2,208 t/s sustained.
62 GB
31B weights in BF16, leaving roughly 18 GB of an 80 GB H100 for KV cache and activations.
📊
Scaling is sublinear, and that's normal. Eight H100s deliver 2,208 tokens/sec against a single card's 1,260 - roughly 1.75× for 8× the hardware. Multi-GPU buys you latency (85 ms versus 279 ms to first token) and capacity for concurrent users, not proportional throughput. If your goal is total tokens per dollar, more replicas of a smaller model usually beats tensor-parallelising a large one.
⚠️
TTFT stays flat until it doesn't. Time to first token holds steady up to about concurrency 8, then climbs as the prefill queue builds. If you're sizing for interactive use, that inflection is the number to design around - not peak throughput, which is measured at a concurrency where individual users are already waiting.
Free speed

Multi-Token Prediction is the biggest single win

If you serve Gemma 4 and aren't using the MTP checkpoints, you're leaving a large amount of performance unclaimed - and it costs nothing in quality.

Gemma 4 31B dense

Tokens/sec at concurrency 1 · H100
With MTP125.3
Baseline40.3
3.11× faster

Gemma 4 26B A4B MoE

Tokens/sec at concurrency 1 · H100
With MTP264.2
Baseline177.1
1.49× faster

Bars are scaled to the same axis across both charts so the models are directly comparable.

Why the dense model gains more

Single-stream decoding on a dense model is memory-bandwidth bound - the GPU spends most of its time moving 62 GB of weights, not computing. Verifying several speculated tokens in one pass amortises that traffic, so the 31B nearly triples. The MoE was already reading far fewer weights per token, so there's less waste to reclaim, and it gains a more modest 1.49×.

What to actually do

Pull the MTP checkpoints for whichever size you serve. Acceptance rates fall off at deeper draft positions - most of the benefit comes from the first few - so there's no advantage in pushing the speculation depth high. And note these gains are measured at concurrency 1; under heavy batching the GPU is compute-bound and speculative decoding helps far less.

Getting more

Other levers worth pulling

🚀

FlashAttention 4

Shipped in the July 2026 refresh for NVIDIA Hopper GPUs: prefill throughput up 25–70% and time to first token down as much as 31%. Most valuable when prompts are long - agent system prompts and tool definitions are exactly that.

📦

QAT 4-bit weights

Google publishes quantisation-aware-trained 4-bit checkpoints, which lose far less quality than post-hoc quantisation. Roughly a quarter of the memory of BF16 - the difference between the 31B fitting on one 24 GB card and not.

🎚️

Turn thinking down

Reasoning traces are generated tokens you pay for and the user never reads. For classification, routing and extraction, minimal thinking can cut latency dramatically with little accuracy cost on those task types.

🖼️

Lower the vision budget

Dropping from 1120 to 280 vision tokens per image is a 4× reduction in prefill work. On general image questions the accuracy cost is about a point; on dense documents it's much larger, so tune per task rather than globally.

🧩

Prefer MoE for serving

The 26B A4B was 4.4× faster than the dense 31B at concurrency 1 for a small quality difference. For most production workloads that's the better trade - reserve the 31B for cases where quality genuinely decides the outcome.

📉

Right-size the model

The cheapest optimisation is a smaller model that's still good enough. The 12B holds much of the 31B's quality at a fraction of the memory - and on many real workloads the gap is smaller than the benchmark spread implies.

The other end

Performance on a phone

The same capability set, four orders of magnitude less hardware.

0.3 s
Time to first token for E2B on a Galaxy S26 Ultra using the GPU backend - faster than the 31B on an H100.
676 MB
Working memory for E2B on GPU. On CPU the same model needs 1,733 MB and takes 1.8 s to first token.
52 t/s
Decode speed on-device - comfortably faster than reading speed, entirely offline.
📱
That 0.3 s is not a typo. E2B on a phone GPU reaches first token faster than the 31B on a datacentre GPU, because a 2B model in 676 MB has almost nothing to read before it can start. The tokens are less capable - but for classification, extraction and short summaries, on-device latency is genuinely better. Full detail on the Android page.
Putting it together

Which model for which job

If your workload is…UseBecause
High-volume serving26B A4B + MTP 264 tok/s at concurrency 1, close to 31B quality, ~3.8B active params
Hardest reasoning, quality decides31B + MTP Best scores in the family; MTP recovers most of the dense speed penalty
Anything with audio or video12B The largest model that accepts them at all
One consumer GPU or a laptop12B QAT 6.7 GB at 4-bit, full 256K context, strong quality per gigabyte
Phone, offline, privateE2B GPU backend 676 MB, 0.3 s to first token, no network at all
Long documents, retrieval31B or 26B A4B Retrieval holds to 128K+; small models degrade sharply on multi-hop
Multi-step agentsPost-July weights, short chains Weakest area - verify each step rather than trusting long chains
FAQ

Common questions

Why is the 26B MoE faster than the 31B dense model?

A mixture-of-experts model routes each token to a subset of its parameters - about 3.8B of the 26B activate per token. Single-stream decoding is limited by how fast the GPU can read weights from memory, so reading far fewer weights per token makes it dramatically faster: 177 tokens/sec against the dense 31B's 40, at concurrency 1. The full 26B still has to fit in memory; you save on bandwidth, not capacity.

Should I always use the MTP checkpoints?

For low-concurrency and interactive workloads, yes - a 3.11× speedup on the 31B with no quality change is as close to free as optimisation gets. Under heavy batching the benefit shrinks, because the GPU becomes compute-bound rather than bandwidth-bound and there's less idle capacity for speculation to exploit.

Does thinking mode change speed a lot?

Substantially, because reasoning traces are real generated tokens with real cost. Every published accuracy figure assumes thinking is on, so turning it down trades measurable accuracy for latency - worth it for classification and extraction, rarely worth it for maths or code.

How much quality do I lose at 4-bit?

Less than you'd expect with Google's QAT checkpoints, because quantisation is part of training rather than applied afterwards. Community post-hoc 4-bit quants lose more. Below 4-bit the degradation becomes severe - a smaller model at 4-bit generally beats a larger one at 3-bit.

Can Gemma 4 generate images or audio?

No. Every model in the family outputs text only. Images and audio are inputs on the models that support them. For generation you'd pair Gemma with a separate model - which is one of the things Agent Skills is designed to coordinate.

Is the 256K context genuinely usable?

On the large models, yes for retrieval - the 31B's RULER accuracy barely moves from 32K to 128K, and MTOB improves at 256K over 128K. Long context also costs memory in the KV cache, though pp-RoPE reduces that by up to 37.5%. On E2B, treat long context as retrieval-only; multi-hop reasoning across it fails.

What's the fastest configuration overall?

For a single stream: 26B A4B with MTP checkpoints, FlashAttention 4 on a Hopper GPU, thinking minimised, and the smallest vision token budget your task tolerates. For total throughput across many users, batch aggressively and expect speculative decoding to matter less.

Capabilities are one half. Scores are the other.

See how these capabilities translate into measured accuracy across all five model sizes.