Local models (Ollama)
Overdeck can point an agent at a model served locally by Ollama
instead of a cloud provider. Ollama 0.14 and newer serves the Anthropic Messages
API natively, which is the API Claude Code speaks — so the harness talks to your
GPU exactly the way it talks to Anthropic, the model traffic never leaves the host,
and the run records $0.
Model ids are ollama:<tag>, for example ollama:gemma4:12b. Overdeck strips the
ollama: prefix before the tag reaches the server; the prefix is what routes the
id to your local endpoint instead of a cloud provider.
Status: the plumbing is verified, the models are not. Overdeck’s launch path is
tested end to end — a real pan start reaches a local model, the agent’s prompt is
served by Ollama, and no model traffic leaves the host. But no local model has yet
completed an Overdeck work-agent task in testing. gemma4:12b specifically failed
as a work agent on a 24 GB RTX 3090: across three attempts it made zero tool
calls, answering the agent prompt with “I am ready. Please provide your
instructions” even when told which tool, which absolute path, and which git command
to use. In a much shorter prompt it did call a tool, but wrote to a directory it
invented and then reported success.So treat this page as a guide to wiring a local model up and experimenting with it,
not as a supported way to run autonomous work agents. Details and transcripts are in
the verification audit.
Requirements
- Ollama 0.14.0 or newer is the floor Overdeck enforces: earlier releases do not
serve the Anthropic Messages API, and the preflight refuses to launch against them
rather than failing mid-turn. Individual models need more. A model’s manifest
can require a newer Ollama than 0.14, and the pull then fails with
412: The model you are attempting to pull requires a newer version of Ollama —
gemma4:12b did exactly that on 0.19.0. Run the newest Ollama you can; upgrade if
a pull returns 412.
- About 24 GB of GPU memory for a 12B model at Q4. Verified on an RTX 3090
(24 GB, CUDA) on Linux with
gemma4:12b, which used 9.2 GB of VRAM at a 64K
window. Apple Silicon with 24 GB of unified memory (Metal) is a target shape, not
yet verified.
- Disk for the model.
gemma4:12b is roughly 8 GB at Q4 quantization.
- The claude-code harness. Other harnesses are refused for
ollama: models in
this release.
Install Ollama
Overdeck never runs an installer for you — a curl | sh fired from a setup script
is exactly the thing worth reading first. Run it yourself:
pan install detects Ollama. If it is missing, the install prints the command for
your platform and carries on. If it is present and you are on an interactive
terminal, it offers to pull gemma4:12b (defaulting to No, because 8 GB is not a
download to start by accident, and because it is not a working work-agent model). Pass
--skip-ollama to skip the step entirely.
Pull a model
Any tag your server has pulled works. gemma4:12b is the tag Overdeck’s install step
offers and the one this page’s examples use, because it is what the wiring was verified
against — not because it works as a work agent. It does not; see the warning above.
Overdeck never substitutes it for a model you asked for.
Set the context length
This is the setting that decides whether local agents work at all. Ollama
silently truncates a prompt longer than the model’s loaded context window: there
is no error, the agent simply stops seeing the beginning of its own instructions.
An Overdeck work agent’s first prompt alone runs to tens of thousands of tokens,
so the default window is far too small.
The window is the server’s to set, so you have to set it on the server. Overdeck
reads back whatever window the server gave the model and pins Claude Code to that
number, so the harness’s own compaction agrees with reality — but it cannot raise
the window for a server it did not start. Set OLLAMA_CONTEXT_LENGTH:
On a host where Overdeck starts the server itself (see pan up
below) it passes OLLAMA_CONTEXT_LENGTH from ollama.context_length, so the config
value is enough and the steps above are unnecessary.
pan doctor warns about any loaded model whose window is under 64K, which is the
check to run after changing this.
Getting this wrong is not a loud failure. Below about 32K a work agent’s first
prompt either gets silently truncated or is rejected outright with
Prompt is too long, and the agent dies on its first turn.
Point a role — or a workhorse slot — at the local model in ~/.overdeck/config.yaml.
Because no local model has completed a work-agent task yet, prefer --model on a single
pan start over pinning roles.work for real work:
The optional top-level ollama: block tunes the endpoint:
A base_url with no port gets Ollama’s default 11434.
A non-localhost base_url is refused at config load. The point of a local model is
that no prompt leaves the machine, so a remote endpoint is a configuration error
rather than a supported deployment.
Run an agent
Before the pane opens, Overdeck preflights the endpoint: it probes health, starts
ollama serve if nothing is listening, checks the version, checks the tag is
pulled, and warm-loads the model to read back its real context window. Every
failure names the fix rather than leaving you with a dead agent:
What pan doctor shows
pan doctor prints Ollama rows only when this host has a reason to care — a
configured ollama: model, or the binary installed. It reports the version and
base URL, warns per configured tag that is not pulled, and warns when a resident
model’s window is under 64K. It never warm-loads a model, so running it does not
pull 8 GB into VRAM as a side effect.
What pan up does
pan up starts ollama serve only when your config names an ollama: model. On
every other host it makes no network call at all. If the server will not start,
pan up prints a warning and continues — cloud-model agents are unaffected.
Limits
- claude-code only.
codex, opencode, kimi-code, acp, and muse are
refused for ollama: models.
- Localhost only. A non-localhost
base_url is a config error.
- Local workspaces only. The preflight checks this host, so remote (Fly.io)
workspaces cannot use local models.
- Cost records
$0. Local runs carry no pricing row, by design.
- Model traffic is local; the process is not.
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC
is exported, so Claude Code’s own telemetry is off, but any remote MCP servers you have
configured are still contacted at startup. If you need a fully offline run, unconfigure
them too.
- No local model has completed a work-agent task yet. See the warning at the top.
- The dashboard model picker and the Settings provider card do not list local tags
yet; configure them in
config.yaml or pass --model on the command line.