How to Self-Host
Gemma 4
Running Gemma 4 as a service - sizing, containers, Kubernetes, monitoring and security. Starting with the arithmetic, because the most common reason people give for self-hosting turns out not to survive it.
Single-machine setup lives on the install page - this is the service version
Self-hosting is usually not the cheap option
People reach for self-hosting to save money. For Gemma 4 on rented GPUs, the arithmetic rarely supports that - so it's worth checking before you build infrastructure around the assumption.
| GPU provider | $/hour | 100% used | 50% | 25% | 10% |
|---|---|---|---|---|---|
| Thunder Compute | $2.19 | $0.48 | $0.97 | $1.93 | $4.83 |
| Runpod | $2.89 | $0.64 | $1.27 | $2.55 | $6.37 |
| Lambda | $3.99 | $0.88 | $1.76 | $3.52 | $8.80 |
| AWS p5 | $6.88 | $1.52 | $3.03 | $6.07 | $15.17 |
| Google Cloud | $10.98 | $2.42 | $4.84 | $9.68 | $24.21 |
| OpenRouter API 31B, blended | - | $0.17 per 1M tokens, at any volume | |||
| Google Gemini API | - | $0.00 per token, rate-limited | |||
Cost per 1M tokens, derived from 1,260 tok/s sustained on one H100. GPU rates checked August 2026.
Run this arithmetic with your own numbers before committing. It takes ten minutes and occasionally saves a quarter of infrastructure work.
When self-hosting is right
All of these are good reasons. None of them are "it's cheaper".
Data cannot leave
Regulatory, contractual or policy constraints that rule out sending prompts to a third party. This is the strongest and most common legitimate reason, and no amount of API pricing changes it.
Offline or air-gapped
Environments with no outbound internet at all. Self-hosting isn't the cheaper option here - it's the only option.
Control over the stack
Pinning an exact checkpoint, custom quantisation, your own fine-tune, or behaviour that can't change under you. Hosted models get updated; a checkpoint you hold does not.
Predictable cost
A fixed monthly GPU bill rather than a variable per-token one. More expensive on average, but easier to budget and immune to a traffic spike becoming an invoice.
Rate limits don't fit
When free-tier throughput is insufficient and paid tiers still cap below what you need. At that point the comparison isn't cost, it's feasibility.
Latency floor
Colocating the model with your application removes network round-trips. Matters for interactive or real-time workloads where tens of milliseconds are visible.
Sizing the hardware
Weights are the floor, not the requirement. KV cache grows with context and concurrency, and it's what actually runs you out of memory.
| Model | BF16 | 4-bit | Fits on | Notes |
|---|---|---|---|---|
| 26B A4B | 57.7 GB | 14.4 GB | 1× H100 (BF16) · 1× 24 GB (4-bit) | Best serving choice - ~3.8B active per token |
| 31B | 69.9 GB | 17.5 GB | 1× H100 (BF16, ~18 GB spare) · 1× 24 GB (4-bit) | Highest quality; slowest single-stream |
| 12B | 26.7 GB | 6.7 GB | 1× A100 40 GB · 1× L4 (4-bit) | Best tokens per dollar; only large model with audio |
| E4B | 17.9 GB | 4.5 GB | Almost anything | For high-volume simple tasks |
Budget for KV cache
On an 80 GB H100 the 31B leaves roughly 18 GB after weights. That headroom is shared between every concurrent request's context - so max context × concurrency is bounded by it, not by the model size.
Set --max-model-len to what your workload needs rather than the full
256K. vLLM pre-allocates based on it, and asking for the maximum starves concurrency.
Plan around the TTFT knee
Time to first token holds steady up to roughly concurrency 8, then climbs as the prefill queue builds. For interactive workloads that inflection is your real capacity number.
Peak throughput figures are measured at concurrency levels where individual users are already waiting. Size for the experience, then check the throughput.
From container to cluster
# vLLM publishes an official image - no Python environment to manage docker run --gpus all \ -v ~/.cache/huggingface:/root/.cache/huggingface \ -e HF_TOKEN=$HF_TOKEN \ -p 8000:8000 \ --ipc=host \ vllm/vllm-openai:latest \ --model google/gemma-4-26B-A4B-it \ --max-model-len 32768 \ --gpu-memory-utilization 0.90 \ --enable-auto-tool-choice \ --tool-call-parser gemma # --ipc=host matters: without it, shared memory limits cause # obscure crashes under tensor parallelism. # Mount the HF cache or every restart re-downloads ~58 GB.
# docker-compose.yml services: gemma: image: vllm/vllm-openai:latest ipc: host ports: ["8000:8000"] volumes: - ~/.cache/huggingface:/root/.cache/huggingface environment: - HF_TOKEN=${HF_TOKEN} command: > --model google/gemma-4-26B-A4B-it --max-model-len 32768 --gpu-memory-utilization 0.90 --enable-auto-tool-choice --tool-call-parser gemma deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu] healthcheck: test: ["CMD", "curl", "-f", "http://localhost:8000/health"] interval: 30s start_period: 300s # model loading takes minutes
# Deployment - the parts that are Gemma-specific spec: containers: - name: vllm image: vllm/vllm-openai:latest args: - --model=google/gemma-4-26B-A4B-it - --max-model-len=32768 - --gpu-memory-utilization=0.90 - --enable-auto-tool-choice - --tool-call-parser=gemma resources: limits: nvidia.com/gpu: 1 volumeMounts: - name: model-cache mountPath: /root/.cache/huggingface - name: dshm # the --ipc=host equivalent mountPath: /dev/shm readinessProbe: httpGet: { path: /health, port: 8000 } initialDelaySeconds: 300 # do not send traffic during load periodSeconds: 10 volumes: - name: dshm emptyDir: { medium: Memory, sizeLimit: 16Gi } - name: model-cache persistentVolumeClaim: { claimName: hf-cache } # Use a shared PVC for the cache so replicas don't each pull 58 GB. # Scale with replicas, not bigger GPUs - see the sizing note above.
# Flags that actually change behaviour for Gemma 4 --enable-auto-tool-choice # required, with the parser below --tool-call-parser gemma # without both, tool calls silently no-op --max-model-len 32768 # NOT 262144 unless you need it - # vLLM pre-allocates KV cache from this --gpu-memory-utilization 0.90 # lower if you hit OOM under load --quantization "fp8" # halves weights, near-identical quality --speculative-model google/gemma-4-26B-A4B-it-assistant # the MTP drafter - up to 3.11x on the 31B --tensor-parallel-size 2 # only when one GPU cannot hold it # And the dependency that stops most first deployments: pip install -U "transformers>=5.5.0"
What to monitor
vLLM exposes Prometheus metrics at /metrics. These are the ones that predict
trouble.
| Metric | Why it matters | Act when |
|---|---|---|
| Time to first token | The number users actually feel | It climbs past your target - usually the prefill queue building |
| KV cache utilisation | The real capacity ceiling | Sustained above ~90% - requests will start queueing or preempting |
| Requests waiting | Queue depth ahead of the scheduler | Consistently non-zero - you need another replica |
| Preemption count | Requests evicted to free KV cache | Anything above occasional - reduce max-model-len or concurrency |
| Tokens/sec, in and out | Actual throughput against your capacity model | It diverges from your sizing assumptions |
| GPU memory and utilisation | Headroom and whether you're compute or memory bound | Memory near the limit, or utilisation low while requests queue |
initialDelaySeconds,
autoscaling reacts far too slowly to be your only defence against a spike, and rolling updates need
surge capacity or you'll drop traffic. Keep a warm replica rather than scaling from zero.
Securing the endpoint
A vLLM server has no authentication, no rate limiting and no abuse controls. It assumes something else provides them.
Never expose the inference server directly
An open OpenAI-compatible endpoint on a public IP is an unmetered GPU for anyone who finds it - and they will be found by scanners.
The inference server should not have a public address at all. Put it behind a gateway you control and keep it on an internal network.
Your own auth, per-user quotas and request size limits. This is the layer that also gives you usage attribution and abuse detection.
Open weights come with no provider moderation. Input and output classification is yours to add - see ethics & safety.
Deployment checklist
Have these in place
- ✅ Ran the cost arithmetic against the API, with your numbers
- ✅
transformers >= 5.5.0pinned in the image - ✅ Post-July-2026 weights, and the checkpoint pinned by revision
- ✅ HF cache on a shared volume, not re-pulled per replica
- ✅
--max-model-lensized to the workload, not the maximum - ✅ Both tool-calling flags set if you use tools
- ✅ Readiness probe delay long enough for model load
- ✅ Shared memory configured (
--ipc=hostor/dev/shm)
And these
- ✅ Endpoint on a private network, gateway in front
- ✅ Authentication and per-user rate limiting
- ✅ Input and output content filtering
- ✅ Prometheus scraping
/metrics, alerts on TTFT and KV cache - ✅ A warm replica - autoscaling can't react in model-load time
- ✅ Load tested at your real concurrency, not just smoke tested
- ✅ A fallback path for when the cluster is unavailable
- ✅ Alerting on request volume by source
Common questions
Is self-hosting really more expensive than the API?
For Gemma 4 31B on rented GPUs, yes - by a wide margin. Even at 100% utilisation the cheapest H100 works out around $0.48 per million tokens against OpenRouter's $0.17 blended, and Google's API charges nothing per token at all.
It changes with owned hardware, smaller models, MTP drafters and very high sustained load. But if cost is your only reason, check the arithmetic before building.
Which model should I serve?
The 26B A4B for most workloads. Only about 3.8B parameters activate per token, making it roughly 4.4× faster than the dense 31B at low concurrency for a two-to-three point accuracy difference. Reserve the 31B for cases where quality genuinely decides the outcome, and consider the 12B if tokens per dollar is what you're optimising.
One big GPU or several small ones?
Several, usually. Multi-GPU scaling is sublinear - eight H100s give 1.75× a single card's throughput. Tensor parallelism is for models that cannot fit on one device, not a throughput strategy. Independent replicas behind a load balancer scale close to linearly and fail independently.
Can I autoscale it?
Partially. Loading tens of gigabytes takes minutes, so a scale-up event arrives long after the spike that triggered it. Keep a warm baseline sized for normal load, autoscale above that for sustained increases, and treat bursts as something to queue or shed rather than scale into. Scaling from zero is not viable for interactive traffic.
Do I need Kubernetes?
Not for one or two replicas - Docker Compose with a health check and a restart policy is genuinely sufficient, and much less to operate. Kubernetes earns its complexity when you need rolling updates without downtime, multiple models, or scheduling across a pool of GPU nodes.
Should I serve quantised weights in production?
FP8 is a good default on modern NVIDIA hardware - roughly half the memory with near-identical quality, which buys you KV cache headroom and therefore concurrency. Below 4-bit, quality degrades sharply enough that a smaller model at 4-bit is the better trade. Google's QAT checkpoints beat community quants at the same bit depth.
What breaks most often?
In rough order: the transformers version being too old, tool calling silently doing
nothing because only one of the two flags was set, out-of-memory under load from an over-large
--max-model-len, readiness probes failing during the minutes-long model load, and shared
memory limits causing obscure crashes when --ipc=host was omitted.
Check the arithmetic first.
Self-host for control, residency and offline operation. Compare against the hosted options on cost.