> ## Documentation Index
> Fetch the complete documentation index at: https://panopticon-cli.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Local models (Ollama)

> Wire an Overdeck agent to a GPU model served by Ollama, at zero cost and with model traffic staying on the host

# Local models (Ollama)

Overdeck can point an agent at a model served locally by [Ollama](https://ollama.com)
instead of a cloud provider. Ollama 0.14 and newer serves the Anthropic Messages
API natively, which is the API Claude Code speaks — so the harness talks to your
GPU exactly the way it talks to Anthropic, the model traffic never leaves the host,
and the run records `$0`.

Model ids are `ollama:<tag>`, for example `ollama:gemma4:12b`. Overdeck strips the
`ollama:` prefix before the tag reaches the server; the prefix is what routes the
id to your local endpoint instead of a cloud provider.

<Warning>
  **Status: the plumbing is verified, the models are not.** Overdeck's launch path is
  tested end to end — a real `pan start` reaches a local model, the agent's prompt is
  served by Ollama, and no model traffic leaves the host. **But no local model has yet
  completed an Overdeck work-agent task in testing.** `gemma4:12b` specifically failed
  as a work agent on a 24 GB RTX 3090: across three attempts it made **zero tool
  calls**, answering the agent prompt with "I am ready. Please provide your
  instructions" even when told which tool, which absolute path, and which git command
  to use. In a much shorter prompt it did call a tool, but wrote to a directory it
  invented and then reported success.

  So treat this page as a guide to wiring a local model up and experimenting with it,
  not as a supported way to run autonomous work agents. Details and transcripts are in
  [the verification audit](https://github.com/eltmon/overdeck/blob/main/docs/audits/pan-1641-local-ollama-verification.md).
</Warning>

## Requirements

* **Ollama 0.14.0 or newer** is the floor Overdeck enforces: earlier releases do not
  serve the Anthropic Messages API, and the preflight refuses to launch against them
  rather than failing mid-turn. **Individual models need more.** A model's manifest
  can require a newer Ollama than 0.14, and the pull then fails with
  `412: The model you are attempting to pull requires a newer version of Ollama` —
  `gemma4:12b` did exactly that on 0.19.0. Run the newest Ollama you can; upgrade if
  a pull returns 412.
* **About 24 GB of GPU memory** for a 12B model at Q4. Verified on an RTX 3090
  (24 GB, CUDA) on Linux with `gemma4:12b`, which used 9.2 GB of VRAM at a 64K
  window. Apple Silicon with 24 GB of unified memory (Metal) is a target shape, not
  yet verified.
* **Disk for the model.** `gemma4:12b` is roughly 8 GB at Q4 quantization.
* The **claude-code harness**. Other harnesses are refused for `ollama:` models in
  this release.

## Install Ollama

Overdeck never runs an installer for you — a `curl | sh` fired from a setup script
is exactly the thing worth reading first. Run it yourself:

```bash theme={null}
# Linux
curl -fsSL https://ollama.com/install.sh | sh

# macOS
brew install ollama   # or install the app from https://ollama.com/download/mac
```

`pan install` detects Ollama. If it is missing, the install prints the command for
your platform and carries on. If it is present and you are on an interactive
terminal, it offers to pull `gemma4:12b` (defaulting to **No**, because 8 GB is not a
download to start by accident, and because it is not a working work-agent model). Pass
`--skip-ollama` to skip the step entirely.

## Pull a model

Any tag your server has pulled works. `gemma4:12b` is the tag Overdeck's install step
offers and the one this page's examples use, because it is what the wiring was verified
against — **not because it works as a work agent.** It does not; see the warning above.
Overdeck never substitutes it for a model you asked for.

```bash theme={null}
ollama pull gemma4:12b
```

## Set the context length

**This is the setting that decides whether local agents work at all.** Ollama
silently truncates a prompt longer than the model's loaded context window: there
is no error, the agent simply stops seeing the beginning of its own instructions.
An Overdeck work agent's first prompt alone runs to tens of thousands of tokens,
so the default window is far too small.

**The window is the server's to set, so you have to set it on the server.** Overdeck
reads back whatever window the server gave the model and pins Claude Code to that
number, so the harness's own compaction agrees with reality — but it cannot raise
the window for a server it did not start. Set `OLLAMA_CONTEXT_LENGTH`:

```bash theme={null}
# Linux (systemd)
sudo systemctl edit ollama
# add:
#   [Service]
#   Environment="OLLAMA_CONTEXT_LENGTH=65536"
sudo systemctl restart ollama

# macOS (the Ollama app)
launchctl setenv OLLAMA_CONTEXT_LENGTH 65536
# then quit and reopen the Ollama app
```

On a host where Overdeck starts the server itself (see [pan up](#what-pan-up-does)
below) it passes `OLLAMA_CONTEXT_LENGTH` from `ollama.context_length`, so the config
value is enough and the steps above are unnecessary.

`pan doctor` warns about any loaded model whose window is under 64K, which is the
check to run after changing this.

<Warning>
  Getting this wrong is not a loud failure. Below about 32K a work agent's first
  prompt either gets silently truncated or is rejected outright with
  `Prompt is too long`, and the agent dies on its first turn.
</Warning>

## Configure Overdeck

Point a role — or a workhorse slot — at the local model in `~/.overdeck/config.yaml`.
Because no local model has completed a work-agent task yet, prefer `--model` on a single
`pan start` over pinning `roles.work` for real work:

```yaml theme={null}
# Experiment with one issue at a time instead:
#   pan start PAN-1234 --model ollama:<tag> --harness claude-code
#
# Pinning a role points every agent of that role at the local model. Only do this
# once you have confirmed your model actually drives an agent:
roles:
  work:
    model: ollama:<tag>
```

The optional top-level `ollama:` block tunes the endpoint:

```yaml theme={null}
ollama:
  base_url: http://localhost:11434   # default; must be a localhost address
  context_length: 65536              # default; used when Overdeck starts the server
```

A `base_url` with no port gets Ollama's default `11434`.

A non-localhost `base_url` is refused at config load. The point of a local model is
that no prompt leaves the machine, so a remote endpoint is a configuration error
rather than a supported deployment.

## Run an agent

```bash theme={null}
pan start PAN-1234 --model ollama:gemma4:12b --harness claude-code
```

Before the pane opens, Overdeck preflights the endpoint: it probes health, starts
`ollama serve` if nothing is listening, checks the version, checks the tag is
pulled, and warm-loads the model to read back its real context window. Every
failure names the fix rather than leaving you with a dead agent:

| What went wrong                          | What you see                                                                                   |
| ---------------------------------------- | ---------------------------------------------------------------------------------------------- |
| Nothing listening, and it will not start | ``Ollama did not become healthy at <url> within 30s. Start it with `ollama serve` and retry.`` |
| Ollama not installed                     | `Ollama could not be started at <url>. Install it from https://ollama.com/download`            |
| Server older than 0.14.0                 | A message naming `0.14.0` as the minimum                                                       |
| Tag not pulled                           | ``Ollama model <tag> is not pulled. Run `ollama pull <tag>`.``                                 |

### What `pan doctor` shows

`pan doctor` prints Ollama rows only when this host has a reason to care — a
configured `ollama:` model, or the binary installed. It reports the version and
base URL, warns per configured tag that is not pulled, and warns when a resident
model's window is under 64K. It never warm-loads a model, so running it does not
pull 8 GB into VRAM as a side effect.

### What `pan up` does

`pan up` starts `ollama serve` only when your config names an `ollama:` model. On
every other host it makes no network call at all. If the server will not start,
`pan up` prints a warning and continues — cloud-model agents are unaffected.

## Limits

* **claude-code only.** `codex`, `opencode`, `kimi-code`, `acp`, and `muse` are
  refused for `ollama:` models.
* **Localhost only.** A non-localhost `base_url` is a config error.
* **Local workspaces only.** The preflight checks this host, so remote (Fly.io)
  workspaces cannot use local models.
* **Cost records `$0`.** Local runs carry no pricing row, by design.
* **Model traffic is local; the process is not.** `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC`
  is exported, so Claude Code's own telemetry is off, but any remote MCP servers you have
  configured are still contacted at startup. If you need a fully offline run, unconfigure
  them too.
* **No local model has completed a work-agent task yet.** See the warning at the top.
* The dashboard model picker and the Settings provider card do not list local tags
  yet; configure them in `config.yaml` or pass `--model` on the command line.
