Use Ollama open models with crystl

Updated August 12, 2026

This tutorial takes you from a new Ollama installation to a Codex shard running on your Mac or another machine you control.

Ollama runs the model. crystl remembers where it is, checks that it is reachable, and gives each new shard the right URL and agent flags.

1. Install Ollama

On a Mac, download Ollama, drag it to Applications, and open it once. Ollama supports Apple silicon acceleration automatically. On a Linux model server, follow Ollama’s Linux installation guide.

Confirm that the CLI is available:

ollama -v

2. Download a coding model

This example uses Qwen3-Coder 30B:

ollama pull qwen3-coder:30b
ollama ls

The download is about 19 GB and the running model needs additional memory for context. If it does not fit comfortably, choose a smaller tool-capable model from the Ollama library and substitute its exact name throughout this tutorial.

Coding agents carry large prompts. Ollama recommends at least 64K context for coding tools when the hardware can support it. Set Context length in the Ollama app, or start a standalone server with:

OLLAMA_CONTEXT_LENGTH=64000 ollama serve

Do not start a second server if the Ollama app is already serving on port 11434.

3. Check Ollama directly

curl http://localhost:11434/api/tags

The response should include qwen3-coder:30b. You can also run ollama ps after the first request to see whether the model is on the GPU or CPU.

4. Add Ollama to crystl

Open Settings → agents → providers & keys, scroll to the open-model connections, and add:

FieldValue
Labelollama-local
KindOllama
Base URLhttp://localhost:11434
As seen from remoteleave empty
Keyleave empty
Context window64000
Enabledon

Save the connection. crystl supplies Ollama’s URL and the Codex --oss --local-provider ollama flags to new shards.

The context window field is optional, and it is worth filling in. Record the same number you passed as OLLAMA_CONTEXT_LENGTH above. crystl uses it to warn you when an endpoint is too small for the work you are about to start. Ollama’s own default is 4096, which is smaller than an agent’s standing prompt, so a slot left blank on a default server is the most common cause of garbled tool calls.

5. Start a Codex shard

crystl shard create --gem myapp --agent codex --local ollama-local \
  --model qwen3-coder:30b \
  --prompt "inspect this project and suggest the best first task"

crystl shard create is a Guild control command. You can create the same shard from the desktop app without Guild.

Give the shard a small real task that requires reading a file and running a command. If it can use its tools and report the result, the setup is complete.

Run Ollama on another machine

Run Ollama on the GPU machine and bind it to that machine’s private network address. For example:

OLLAMA_HOST=0.0.0.0:11434 OLLAMA_CONTEXT_LENGTH=64000 ollama serve

In crystl, use the machine’s LAN or Tailscale address, such as http://100.x.y.z:11434. Do not expose an unauthenticated Ollama server to the public internet. If the agent itself runs on that machine over SSH, set As seen from remote to http://localhost:11434.

Troubleshooting

crystl says unreachable: run the direct curl check from your Mac. For another machine, check its firewall and confirm Ollama is not bound only to 127.0.0.1.

The model is missing: ollama ls shows downloaded models. Model names and tags must match exactly.

The agent forgets instructions or loops: increase Ollama’s context length, then restart the server. Larger context consumes more memory.

Generation is extremely slow: run ollama ps. A model partly or entirely on CPU may be too large for the available GPU or unified memory.

Other setup tutorials

  • LM Studio: a graphical local runtime with an Anthropic-compatible server.
  • vLLM: high-throughput serving on a Linux GPU machine.
  • llama.cpp: a small native server with direct control over quantization.
  • z.ai: hosted GLM models, no local GPU.