Files
Codeman/docs/wiki/Custom-Model-Endpoints.md
T
DevvynandClaude Sonnet 5 5a9ff07f57 feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:

1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
   apply the endpoint's defaultModelId (or the first discovered model)
   silently. Now, via the new selectCustomModelEntry() (session-ui.js):
   - exactly one discovered model launches straight away, same as before
   - two or more open a new #customModelPickModal listing every discovered
     model; defaultModelId (if set) is marked but never auto-chosen, since
     the point of asking is letting ONE launch deliberately differ from
     the saved default, not just confirming it
   The endpoint is re-fetched at click time rather than trusting anything
   cached from the dropdown's own render, since the model list can have
   changed (the sweep below, or a settings-panel edit) since it opened.
   runCustomModelEntry() itself — the actual launch, routed through run()
   for the in-flight lock, snapshot-guarded against applying to the wrong
   session — is unchanged; it now just always receives an explicit model
   id from one of these two paths instead of computing one itself.

2. Periodic re-discovery. Every saved endpoint's models now refresh
   automatically every 5 minutes in the background
   (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
   as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
   off under testMode), so a model the server starts or stops serving shows
   up without another manual "Discover" click. The manual POST
   .../discover-models route and the new refreshAllCustomModelHosts()
   sweep (custom-model-routes.ts) now share one pure merge step
   (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
   that no longer appears) rather than two copies that could drift. The
   sweep is best-effort per host — one endpoint being unreachable on a
   cycle never blocks the others — and re-reads the store before each
   host's write, keyed by id, so a concurrent edit or delete from the
   settings panel always wins over a sweep that started before it.

Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.

Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 08:50:18 +08:00

5.5 KiB

Custom Model Endpoints

Point a harness at your own OpenAI-compatible server instead of its native cloud backend, for one session at a time. "Custom endpoint" covers local hardware (llama.cpp, Ollama, vLLM, a home GPU rig, DGX Spark, Strix Halo) and cloud services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a company gateway) alike, anything answering GET /v1/models and POST /v1/chat/completions in the standard shape.

Off by default. Turn it on in App Settings → Models → Custom model endpoints.

Adding an endpoint

Still in App Settings → Models → Custom model endpoints:

  1. + Add endpoint — give it an id, a label, and the base URL (http://192.168.1.50:8080, say). An API key is optional; most local servers don't check one.
  2. Discover — fetches the endpoint's own model list over GET /v1/models and stores it.
  3. Pick a default model from what was discovered. This is the model the Run-menu entry applies directly when only one model is discovered; with two or more, it's just the one pre-marked in the picker dialog described below, not a silent default.

Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker hosts — these are machine-level infra, not a per-user setting.

Model lists refresh themselves. Every saved endpoint is re-discovered automatically every 5 minutes in the background, so a model the server starts serving later — or stops serving — shows up without another manual click of Discover. One endpoint being unreachable on a given cycle (powered off, wrong network) never blocks the others from refreshing.

Running a session against one

With the setting on and at least one endpoint carrying a discovered model, the Run dropdown grows a Custom Endpoints section: one entry per harness that can redirect to a custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a session on that harness exactly the way its own entry would. It is a one-off "try this endpoint" action, not a sticky mode — the plain Run button still means "this harness, native cloud" afterward, and a fresh session never inherits whatever the last one was pointed at.

Which model it uses depends on how many the endpoint has discovered. With exactly one, the session launches straight away on that model — nothing to choose. With two or more, a small dialog asks which one to use for this launch before starting the session; the endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately use a different one without changing the saved default.

Applying a selection restarts the harness's process in place — same tab, same conversation where the harness supports resuming one, fresh environment. That restart is necessary, not incidental: every supported harness reads its endpoint config at process start, never per turn, so there is no live hot-swap while a turn is running.

Entries are hidden entirely for a session in a remote (SSH) or Docker case — support for redirecting those hasn't landed yet, see below. The picker also only appears in the desktop Run dropdown; the phone home screen builds its own run picker separately and does not currently offer these entries.

Which harnesses actually work

Harness Status
Claude Code, opencode, Pi, Grok, OMP Verified end-to-end against a real local server.
Codex Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug.
Gemini Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet.
DeepSeek Reaches the server but gets a consistent 404. Root cause not identified.
Antigravity No known custom-endpoint mechanism at all. Not offered.

Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry, not a fixed list here, so this table can go stale before this page does — a greyed-out or missing entry is the more current answer.

What it does not do

  • No remote or Docker sessions yet. Both restart their agent differently under the hood (reattaching a durable tmux session rather than relaunching the process), so redirecting them needs its own plumbing that hasn't been built.
  • No live hot-swap mid-conversation. Applying a selection always restarts the process.
  • No button to un-point a session from the UI yet. Clearing back to native cloud is an HTTP call (POST .../custom-model {"clear": true}) or deleting the session; the settings panel manages saved endpoints, not what a running session is currently pointed at.
  • Nothing is shared with your real cloud credentials. The endpoint's own key, if any, never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate, explicit choice per session.

Security

An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save time and against the address it actually resolves to), the same guard Web Tabs uses for saved dashboards. Endpoint records and any per-session config files a harness needs are written with owner-only permissions. See custom-model-endpoints-plan.md in the repository for the full design reasoning, including why this feature closed a pre-existing gap in how session environment overrides were guarded rather than opening a new one.