Skip to main content

Local models (Ollama)

Overdeck can point an agent at a model served locally by Ollama instead of a cloud provider. Ollama 0.14 and newer serves the Anthropic Messages API natively, which is the API Claude Code speaks — so the harness talks to your GPU exactly the way it talks to Anthropic, the model traffic never leaves the host, and the run records $0. Model ids are ollama:<tag>, for example ollama:gemma4:12b. Overdeck strips the ollama: prefix before the tag reaches the server; the prefix is what routes the id to your local endpoint instead of a cloud provider.
Status: the plumbing is verified, the models are not. Overdeck’s launch path is tested end to end — a real pan start reaches a local model, the agent’s prompt is served by Ollama, and no model traffic leaves the host. But no local model has yet completed an Overdeck work-agent task in testing. gemma4:12b specifically failed as a work agent on a 24 GB RTX 3090: across three attempts it made zero tool calls, answering the agent prompt with “I am ready. Please provide your instructions” even when told which tool, which absolute path, and which git command to use. In a much shorter prompt it did call a tool, but wrote to a directory it invented and then reported success.So treat this page as a guide to wiring a local model up and experimenting with it, not as a supported way to run autonomous work agents. Details and transcripts are in the verification audit.

Requirements

  • Ollama 0.14.0 or newer is the floor Overdeck enforces: earlier releases do not serve the Anthropic Messages API, and the preflight refuses to launch against them rather than failing mid-turn. Individual models need more. A model’s manifest can require a newer Ollama than 0.14, and the pull then fails with 412: The model you are attempting to pull requires a newer version of Ollama — gemma4:12b did exactly that on 0.19.0. Run the newest Ollama you can; upgrade if a pull returns 412.
  • About 24 GB of GPU memory for a 12B model at Q4. Verified on an RTX 3090 (24 GB, CUDA) on Linux with gemma4:12b, which used 9.2 GB of VRAM at a 64K window. Apple Silicon with 24 GB of unified memory (Metal) is a target shape, not yet verified.
  • Disk for the model. gemma4:12b is roughly 8 GB at Q4 quantization.
  • The claude-code harness. Other harnesses are refused for ollama: models in this release.

Install Ollama

Overdeck never runs an installer for you — a curl | sh fired from a setup script is exactly the thing worth reading first. Run it yourself:
pan install detects Ollama. If it is missing, the install prints the command for your platform and carries on. If it is present and you are on an interactive terminal, it offers to pull gemma4:12b (defaulting to No, because 8 GB is not a download to start by accident, and because it is not a working work-agent model). Pass --skip-ollama to skip the step entirely.

Pull a model

Any tag your server has pulled works. gemma4:12b is the tag Overdeck’s install step offers and the one this page’s examples use, because it is what the wiring was verified against — not because it works as a work agent. It does not; see the warning above. Overdeck never substitutes it for a model you asked for.

Set the context length

This is the setting that decides whether local agents work at all. Ollama silently truncates a prompt longer than the model’s loaded context window: there is no error, the agent simply stops seeing the beginning of its own instructions. An Overdeck work agent’s first prompt alone runs to tens of thousands of tokens, so the default window is far too small. The window is the server’s to set, so you have to set it on the server. Overdeck reads back whatever window the server gave the model and pins Claude Code to that number, so the harness’s own compaction agrees with reality — but it cannot raise the window for a server it did not start. Set OLLAMA_CONTEXT_LENGTH:
On a host where Overdeck starts the server itself (see pan up below) it passes OLLAMA_CONTEXT_LENGTH from ollama.context_length, so the config value is enough and the steps above are unnecessary. pan doctor warns about any loaded model whose window is under 64K, which is the check to run after changing this.
Getting this wrong is not a loud failure. Below about 32K a work agent’s first prompt either gets silently truncated or is rejected outright with Prompt is too long, and the agent dies on its first turn.

Configure Overdeck

Point a role — or a workhorse slot — at the local model in ~/.overdeck/config.yaml. Because no local model has completed a work-agent task yet, prefer --model on a single pan start over pinning roles.work for real work:
The optional top-level ollama: block tunes the endpoint:
A base_url with no port gets Ollama’s default 11434. A non-localhost base_url is refused at config load. The point of a local model is that no prompt leaves the machine, so a remote endpoint is a configuration error rather than a supported deployment.

Run an agent

Before the pane opens, Overdeck preflights the endpoint: it probes health, starts ollama serve if nothing is listening, checks the version, checks the tag is pulled, and warm-loads the model to read back its real context window. Every failure names the fix rather than leaving you with a dead agent:

What pan doctor shows

pan doctor prints Ollama rows only when this host has a reason to care — a configured ollama: model, or the binary installed. It reports the version and base URL, warns per configured tag that is not pulled, and warns when a resident model’s window is under 64K. It never warm-loads a model, so running it does not pull 8 GB into VRAM as a side effect.

What pan up does

pan up starts ollama serve only when your config names an ollama: model. On every other host it makes no network call at all. If the server will not start, pan up prints a warning and continues — cloud-model agents are unaffected.

Limits

  • claude-code only. codex, opencode, kimi-code, acp, and muse are refused for ollama: models.
  • Localhost only. A non-localhost base_url is a config error.
  • Local workspaces only. The preflight checks this host, so remote (Fly.io) workspaces cannot use local models.
  • Cost records $0. Local runs carry no pricing row, by design.
  • Model traffic is local; the process is not. CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC is exported, so Claude Code’s own telemetry is off, but any remote MCP servers you have configured are still contacted at startup. If you need a fully offline run, unconfigure them too.
  • No local model has completed a work-agent task yet. See the warning at the top.
  • The dashboard model picker and the Settings provider card do not list local tags yet; configure them in config.yaml or pass --model on the command line.