Updated August 2026 Production deployment

How to Self-Host
Gemma 4

Running Gemma 4 as a service - sizing, containers, Kubernetes, monitoring and security. Starting with the arithmetic, because the most common reason people give for self-hosting turns out not to survive it.

Single-machine setup lives on the install page - this is the service version

1× H100 · Gemma 4 31B · vLLM
Aggregate throughput1,260 tok/s
Time to first token279 ms @ c1
TTFT stable to~concurrency 8
Weights (BF16)62 GB
Left for KV cache~18 GB
8× H100 gives2,208 tok/s
Measured at ISL 512 / OSL 256 - your workload will differ
Before anything else

Self-hosting is usually not the cheap option

People reach for self-hosting to save money. For Gemma 4 on rented GPUs, the arithmetic rarely supports that - so it's worth checking before you build infrastructure around the assumption.

GPU provider$/hour100% used 50%25%10%
Thunder Compute$2.19 $0.48$0.97$1.93$4.83
Runpod$2.89 $0.64$1.27$2.55$6.37
Lambda$3.99 $0.88$1.76$3.52$8.80
AWS p5$6.88 $1.52$3.03$6.07$15.17
Google Cloud$10.98 $2.42$4.84$9.68$24.21
OpenRouter API 31B, blended -$0.17 per 1M tokens, at any volume
Google Gemini API -$0.00 per token, rate-limited

Cost per 1M tokens, derived from 1,260 tok/s sustained on one H100. GPU rates checked August 2026.

🧮
Even at 100% utilisation, the cheapest rented H100 costs about 2.8× what the API charges. To match OpenRouter's blended rate for this workload you'd need roughly 284% utilisation on the cheapest provider, or 892% on AWS - which is to say it isn't reachable. And real deployments don't run at 100%; at a realistic 25% you're paying an order of magnitude more per token than the API, while Google's own API charges nothing at all.
⚖️
Where this calculation stops applying. It assumes rented H100s, one workload shape (512 in / 256 out) from one benchmark, and the 31B. Four things move it meaningfully: owned hardware, where the capital cost is already sunk and you're paying only for power; smaller models - a 12B on a cheaper GPU gives far more tokens per dollar; MTP drafters, worth up to 3.11× on the 31B; and API rate limits that simply don't accommodate your volume, in which case cost isn't the deciding variable anyway.

Run this arithmetic with your own numbers before committing. It takes ten minutes and occasionally saves a quarter of infrastructure work.
The real reasons

When self-hosting is right

All of these are good reasons. None of them are "it's cheaper".

🔒

Data cannot leave

Regulatory, contractual or policy constraints that rule out sending prompts to a third party. This is the strongest and most common legitimate reason, and no amount of API pricing changes it.

📴

Offline or air-gapped

Environments with no outbound internet at all. Self-hosting isn't the cheaper option here - it's the only option.

🎛️

Control over the stack

Pinning an exact checkpoint, custom quantisation, your own fine-tune, or behaviour that can't change under you. Hosted models get updated; a checkpoint you hold does not.

📊

Predictable cost

A fixed monthly GPU bill rather than a variable per-token one. More expensive on average, but easier to budget and immune to a traffic spike becoming an invoice.

🚦

Rate limits don't fit

When free-tier throughput is insufficient and paid tiers still cap below what you need. At that point the comparison isn't cost, it's feasibility.

⚡

Latency floor

Colocating the model with your application removes network round-trips. Matters for interactive or real-time workloads where tens of milliseconds are visible.

Capacity

Sizing the hardware

Weights are the floor, not the requirement. KV cache grows with context and concurrency, and it's what actually runs you out of memory.

ModelBF164-bitFits onNotes
26B A4B 57.7 GB14.4 GB1× H100 (BF16) · 1× 24 GB (4-bit) Best serving choice - ~3.8B active per token
31B69.9 GB17.5 GB 1× H100 (BF16, ~18 GB spare) · 1× 24 GB (4-bit) Highest quality; slowest single-stream
12B26.7 GB6.7 GB 1× A100 40 GB · 1× L4 (4-bit) Best tokens per dollar; only large model with audio
E4B17.9 GB4.5 GB Almost anythingFor high-volume simple tasks
📐

Budget for KV cache

The part people forget

On an 80 GB H100 the 31B leaves roughly 18 GB after weights. That headroom is shared between every concurrent request's context - so max context × concurrency is bounded by it, not by the model size.

Set --max-model-len to what your workload needs rather than the full 256K. vLLM pre-allocates based on it, and asking for the maximum starves concurrency.

📈

Plan around the TTFT knee

Not peak throughput

Time to first token holds steady up to roughly concurrency 8, then climbs as the prefill queue builds. For interactive workloads that inflection is your real capacity number.

Peak throughput figures are measured at concurrency levels where individual users are already waiting. Size for the experience, then check the throughput.

📉
Multi-GPU scaling is sublinear. Eight H100s deliver 2,208 tok/s against a single card's 1,260 - roughly 1.75× for 8× the hardware. What you buy is latency (85 ms to first token instead of 279 ms) and concurrent capacity, not proportional throughput. If total tokens per dollar is the goal, more replicas of a smaller model beats tensor-parallelising a large one almost every time.
Deployment

From container to cluster

# vLLM publishes an official image - no Python environment to manage
docker run --gpus all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -e HF_TOKEN=$HF_TOKEN \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai:latest \
  --model google/gemma-4-26B-A4B-it \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90 \
  --enable-auto-tool-choice \
  --tool-call-parser gemma

# --ipc=host matters: without it, shared memory limits cause
# obscure crashes under tensor parallelism.
# Mount the HF cache or every restart re-downloads ~58 GB.
# docker-compose.yml
services:
  gemma:
    image: vllm/vllm-openai:latest
    ipc: host
    ports: ["8000:8000"]
    volumes:
      - ~/.cache/huggingface:/root/.cache/huggingface
    environment:
      - HF_TOKEN=${HF_TOKEN}
    command: >
      --model google/gemma-4-26B-A4B-it
      --max-model-len 32768
      --gpu-memory-utilization 0.90
      --enable-auto-tool-choice
      --tool-call-parser gemma
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
      interval: 30s
      start_period: 300s   # model loading takes minutes
# Deployment - the parts that are Gemma-specific
spec:
  containers:
  - name: vllm
    image: vllm/vllm-openai:latest
    args:
      - --model=google/gemma-4-26B-A4B-it
      - --max-model-len=32768
      - --gpu-memory-utilization=0.90
      - --enable-auto-tool-choice
      - --tool-call-parser=gemma
    resources:
      limits:
        nvidia.com/gpu: 1
    volumeMounts:
      - name: model-cache
        mountPath: /root/.cache/huggingface
      - name: dshm                    # the --ipc=host equivalent
        mountPath: /dev/shm
    readinessProbe:
      httpGet: { path: /health, port: 8000 }
      initialDelaySeconds: 300    # do not send traffic during load
      periodSeconds: 10
  volumes:
    - name: dshm
      emptyDir: { medium: Memory, sizeLimit: 16Gi }
    - name: model-cache
      persistentVolumeClaim: { claimName: hf-cache }

# Use a shared PVC for the cache so replicas don't each pull 58 GB.
# Scale with replicas, not bigger GPUs - see the sizing note above.
# Flags that actually change behaviour for Gemma 4

--enable-auto-tool-choice        # required, with the parser below
--tool-call-parser gemma         # without both, tool calls silently no-op

--max-model-len 32768            # NOT 262144 unless you need it -
                                 # vLLM pre-allocates KV cache from this

--gpu-memory-utilization 0.90    # lower if you hit OOM under load

--quantization "fp8"              # halves weights, near-identical quality

--speculative-model google/gemma-4-26B-A4B-it-assistant
                                 # the MTP drafter - up to 3.11x on the 31B

--tensor-parallel-size 2         # only when one GPU cannot hold it

# And the dependency that stops most first deployments:
pip install -U "transformers>=5.5.0"
Running it

What to monitor

vLLM exposes Prometheus metrics at /metrics. These are the ones that predict trouble.

MetricWhy it mattersAct when
Time to first token The number users actually feel It climbs past your target - usually the prefill queue building
KV cache utilisation The real capacity ceilingSustained above ~90% - requests will start queueing or preempting
Requests waiting Queue depth ahead of the schedulerConsistently non-zero - you need another replica
Preemption count Requests evicted to free KV cacheAnything above occasional - reduce max-model-len or concurrency
Tokens/sec, in and out Actual throughput against your capacity modelIt diverges from your sizing assumptions
GPU memory and utilisation Headroom and whether you're compute or memory boundMemory near the limit, or utilisation low while requests queue
⏱️
Model load time changes how you deploy. Loading tens of gigabytes of weights takes minutes, not seconds - which means readiness probes need generous initialDelaySeconds, autoscaling reacts far too slowly to be your only defence against a spike, and rolling updates need surge capacity or you'll drop traffic. Keep a warm replica rather than scaling from zero.
Do not skip

Securing the endpoint

A vLLM server has no authentication, no rate limiting and no abuse controls. It assumes something else provides them.

🚨

Never expose the inference server directly

An open OpenAI-compatible endpoint on a public IP is an unmetered GPU for anyone who finds it - and they will be found by scanners.

Network Bind to localhost or a private subnet

The inference server should not have a public address at all. Put it behind a gateway you control and keep it on an internal network.

Gateway Authenticate and rate-limit in front

Your own auth, per-user quotas and request size limits. This is the layer that also gives you usage attribution and abuse detection.

Content Filter both directions

Open weights come with no provider moderation. Input and output classification is yours to add - see ethics & safety.

💸
The failure mode is expensive in an unusual way. A leaked API key costs you tokens until you rotate it. An exposed self-hosted endpoint costs you the full GPU bill while serving someone else's traffic - and because it's a fixed monthly cost, nothing spikes to alert you. Set alerts on request volume by source, not just on spend.
Before production

Deployment checklist

Infrastructure

Have these in place

  • ✅ Ran the cost arithmetic against the API, with your numbers
  • ✅ transformers >= 5.5.0 pinned in the image
  • ✅ Post-July-2026 weights, and the checkpoint pinned by revision
  • ✅ HF cache on a shared volume, not re-pulled per replica
  • ✅ --max-model-len sized to the workload, not the maximum
  • ✅ Both tool-calling flags set if you use tools
  • ✅ Readiness probe delay long enough for model load
  • ✅ Shared memory configured (--ipc=host or /dev/shm)
Operations

And these

  • ✅ Endpoint on a private network, gateway in front
  • ✅ Authentication and per-user rate limiting
  • ✅ Input and output content filtering
  • ✅ Prometheus scraping /metrics, alerts on TTFT and KV cache
  • ✅ A warm replica - autoscaling can't react in model-load time
  • ✅ Load tested at your real concurrency, not just smoke tested
  • ✅ A fallback path for when the cluster is unavailable
  • ✅ Alerting on request volume by source
FAQ

Common questions

Is self-hosting really more expensive than the API?

For Gemma 4 31B on rented GPUs, yes - by a wide margin. Even at 100% utilisation the cheapest H100 works out around $0.48 per million tokens against OpenRouter's $0.17 blended, and Google's API charges nothing per token at all.

It changes with owned hardware, smaller models, MTP drafters and very high sustained load. But if cost is your only reason, check the arithmetic before building.

Which model should I serve?

The 26B A4B for most workloads. Only about 3.8B parameters activate per token, making it roughly 4.4× faster than the dense 31B at low concurrency for a two-to-three point accuracy difference. Reserve the 31B for cases where quality genuinely decides the outcome, and consider the 12B if tokens per dollar is what you're optimising.

One big GPU or several small ones?

Several, usually. Multi-GPU scaling is sublinear - eight H100s give 1.75× a single card's throughput. Tensor parallelism is for models that cannot fit on one device, not a throughput strategy. Independent replicas behind a load balancer scale close to linearly and fail independently.

Can I autoscale it?

Partially. Loading tens of gigabytes takes minutes, so a scale-up event arrives long after the spike that triggered it. Keep a warm baseline sized for normal load, autoscale above that for sustained increases, and treat bursts as something to queue or shed rather than scale into. Scaling from zero is not viable for interactive traffic.

Do I need Kubernetes?

Not for one or two replicas - Docker Compose with a health check and a restart policy is genuinely sufficient, and much less to operate. Kubernetes earns its complexity when you need rolling updates without downtime, multiple models, or scheduling across a pool of GPU nodes.

Should I serve quantised weights in production?

FP8 is a good default on modern NVIDIA hardware - roughly half the memory with near-identical quality, which buys you KV cache headroom and therefore concurrency. Below 4-bit, quality degrades sharply enough that a smaller model at 4-bit is the better trade. Google's QAT checkpoints beat community quants at the same bit depth.

What breaks most often?

In rough order: the transformers version being too old, tool calling silently doing nothing because only one of the two flags was set, out-of-memory under load from an over-large --max-model-len, readiness probes failing during the minutes-long model load, and shared memory limits causing obscure crashes when --ipc=host was omitted.

Check the arithmetic first.

Self-host for control, residency and offline operation. Compare against the hosted options on cost.