Gemma 4 Playground
A real chat playground that talks to Gemma 4 running locally through Ollama. Adjust sampling, switch thinking modes, and watch the model reason - all in your browser, with no server in between.
Connect a local model in three steps
This playground has no backend - it calls Ollama on your own machine, so your prompts never leave your computer. That also means nothing works until Ollama is running.
Download it from ollama.com and launch it. It runs quietly in the background on port 11434.
In a terminal. The 12B is the best all-rounder; use e4b if you're short on memory.
ollama pull gemma4:12b
Browsers block cross-origin requests by default, so Ollama has to be told this page is allowed. Set the variable, then restart Ollama.
# macOS launchctl setenv OLLAMA_ORIGINS "*" # Linux export OLLAMA_ORIGINS="*" # Windows (PowerShell) setx OLLAMA_ORIGINS "*"
Setting * lets any site in your browser reach your local
Ollama. On a shared or untrusted machine, use your specific origin instead of the wildcard.
Start a conversation
Type a message below, or pick something from the prompt library to try.
Things worth trying
Click any card to load it into the chat box. The ones marked with a brain are where thinking mode makes a visible difference - turn it up and watch the reasoning trace.
How to drive it well
Turn thinking up for hard things
Maths, multi-step logic and tricky debugging get substantially better with thinking on high - it's what drives Gemma 4's strong AIME and GPQA scores. Leave it off for chat and simple extraction, where it just costs latency.
Temperature is not a quality dial
Higher isn't smarter, just less predictable. Use 0–0.3 for factual work, code and extraction; 0.7–1.0 for brainstorming and prose. Above about 1.2 output degrades quickly.
Use the system prompt
Gemma 4 supports system prompts natively. Setting a role and output format there works better than repeating instructions in every message, and it survives the whole conversation.
Watch your context
Ollama defaults to a smaller context than the model supports. For long documents, raise it with
/set parameter num_ctx 32768 in the Ollama CLI, or the model will silently truncate.
Try a smaller model first
E4B answers most everyday prompts about as well as 31B and runs several times faster. Escalate to a bigger model when you can actually see the smaller one failing.
Compare like for like
If you're evaluating sizes, hold temperature, top-p and the system prompt fixed and set a seed. Otherwise you're measuring sampling noise, not the model.
About this playground
Where do my prompts go?
Nowhere. The page sends them directly from your browser to Ollama on your own machine. There's no backend here, no analytics on your messages, and no third-party API. If you disconnect from the internet after loading the page, it keeps working.
Why do I need to set OLLAMA_ORIGINS?
Browser security. A page served from one origin can't call a server on another unless that server
says it's allowed, and Ollama only permits localhost origins by default. Setting
OLLAMA_ORIGINS="*" tells it to accept requests from any page you have open.
That's a real trade-off: it means any website you visit could reach your local Ollama while it's running. On a machine you don't fully control, set the specific origin you're serving this page from rather than the wildcard.
Nothing happens when I hit send
Check the connection bar at the top. If it's red, Ollama either isn't running or isn't reachable -
confirm with ollama list in a terminal.
If the bar is green but sending fails, it's almost always CORS. Restart Ollama after setting
OLLAMA_ORIGINS; the variable is only read at startup.
It's very slow
The model is probably larger than your memory and spilling to disk. Drop to a
smaller size or a 4-bit quantisation - gemma4:e4b-it-qat is about 6 GB and much faster than
a 31B that doesn't fit. Turning thinking mode down also helps a lot, since reasoning traces are
generated tokens like any other.
Does it keep my conversation after I refresh?
No. Everything is held in memory for the current page only, so a refresh clears it. That's deliberate - nothing about your session is written to disk by this page.
Can I point it at something other than Ollama?
Any server exposing Ollama's /api/chat and /api/tags
endpoints will work - change the address in the connection bar. That includes Ollama running on
another machine on your network. It won't work with an OpenAI-style endpoint, since the streaming
format differs.
Don't have a model yet?
One command gets you running. The guide walks through every option if you'd rather choose deliberately.