Files
Codeman/docs/custom-model-endpoints.md
T
DevvynandClaude Sonnet 5 f865f74a0f feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.

POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.

Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.

Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:04:10 +08:00

17 KiB

Custom Model Endpoint Profiles

Point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, or OMP — at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. "Custom endpoint" covers both local hardware (llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built boxes like NVIDIA DGX Spark or AMD Strix Halo mini-PCs) and cloud services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a company gateway) — anything answering GET /v1/models and POST /v1/chat/completions in the standard shape. Design doc, per-CLI recipe confidence table, and security reasoning: custom-model-endpoints-plan.md.

Status: fully wired end to end — registry capability, the injection engine, the endpoint store + discovery route, the session restart route, a settings-panel CRUD surface, and the Run-menu picker described below. Antigravity has no known custom-endpoint mechanism and is not supported. The HTTP API (examples below) still works directly and is what the picker itself calls under the hood.

Turning it on

App Settings → Models → Custom model endpoints (synced setting customModelEndpointsEnabled, default OFF). Turning it on does two things: it reveals the endpoint list/add/edit/discover panel in that same settings section, and it makes the Run menu offer a generated entry per (harness, endpoint) pair — see "The Run-menu picker" below. The API equivalent:

curl -sk -X PUT https://localhost:3000/api/settings \
  -H 'Content-Type: application/json' \
  -d '{"customModelEndpointsEnabled": true}'

Adding an endpoint

Via App Settings → Models → Custom model endpoints → + Add endpoint, or directly:

curl -sk -X POST https://localhost:3000/api/model-endpoints \
  -H 'Content-Type: application/json' \
  -d '{"id": "llama-box", "label": "Home llama.cpp", "baseUrl": "http://192.168.1.50:8080"}'

apiKey is optional (most local servers don't check it). authStyle (bearer | api-key, default bearer) controls which auth header convention discovery uses: bearer is Authorization: Bearer <key> (llama.cpp, OpenAI-compatible servers, most gateways), api-key is the api-key: <key> header Azure AI Foundry wants. There is deliberately no "send both" option: measured against a real llama-swap server, a request carrying both headers hung indefinitely. baseUrl must be http(s), carry no embedded credentials, and may not point at a link-local or cloud-metadata address; discovery re-checks the address the name actually resolves to.

Discover its available models:

curl -sk -X POST https://localhost:3000/api/model-endpoints/llama-box/discover-models

This calls the endpoint's own GET /v1/models and stores the returned list on the endpoint record; GET /api/model-endpoints lists everything configured, PUT/DELETE /api/model-endpoints/:id update or remove one. Endpoint management is admin-only in multi-user mode, same as remote/docker hosts — these are machine-level infra, not per-user settings.

Context length is discovered too, opportunistically and safely. The plain GET /v1/models response has no context-window field, but llama.cpp's llama-swap-proxied GET /props?model=<id> does (n_ctx). Discovery only ever calls it for a model llama-swap's own response already reports status.value === "loaded" for — never for an unloaded one, because llama-swap treats ?model= as a routing hint and asking about a model that isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side effect of what should be read-only discovery. A server with no status field on any entry at all (not llama-swap) gets no context-length enrichment, rather than guessing. A model's previously-learned context length survives a later cycle where it wasn't the loaded one; it's dropped only once the model disappears from the endpoint's list entirely. Stored per model in modelContextLengths and applied automatically (see "Applying a model to a session" below) so a CLI that would otherwise assume a large default context window for an unrecognized model id stops silently overflowing a much smaller real one.

defaultModelId names which discovered model the picker pre-marks for that endpoint — the settings panel's Edit form exposes it as a select populated from the endpoint's own discovered models, and the route refuses a value that isn't one of them. It is applied automatically only when the endpoint has exactly one discovered model (nothing to choose); with two or more it is a pre-selection in the model-picker dialog below, never a silent default. Re-discovering drops a default that no longer appears in the fresh list rather than carrying an invalid one forward.

Model lists refresh themselves. A background sweep (server.ts, CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, every 5 minutes) re-discovers every saved endpoint the same way the manual POST .../discover-models route does, best-effort per endpoint — one being unreachable on a given cycle never blocks the others. Off under npm test, same reasoning as the Codex plan-usage poll it sits beside: no real network to hit, no server instance to keep the timer alive for.

The Run-menu picker

With the setting on and at least one endpoint carrying a discovered model, the toolbar's Run dropdown grows a Custom Endpoints section: one entry per (harness that can redirect to a custom endpoint, saved endpoint) pair, e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI registry's own capabilities.customModelInjection at page render (window.__codemanCustomModelClis, server.ts) — never a hardcoded id list in the frontend — so a CLI whose injection recipe lands later shows up with no frontend change, and Antigravity (unsupported) never does.

Picking an entry re-fetches the endpoint (selectCustomModelEntry(), session-ui.js) rather than trusting anything cached from the dropdown's own render — the model list can have changed via the 5-minute sweep above or a settings-panel edit since the menu opened. With exactly one discovered model it runs straight away; with two or more, a small modal (#customModelPickModal) lists them and asks which one to use for this launch, with the endpoint's defaultModelId marked but not auto-chosen — the point of asking is letting one launch deliberately differ from the saved default, not just confirming it.

How the launch itself applies the endpoint depends on the harness. For opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP (runCustomModelEntry → _runCustomModelEntryOneShot), the endpoint/model is folded into the SAME POST /api/quick-start call that creates the session (customModel field), so the session launches directly on the endpoint — no restart, no visible relaunch. Claude (_runCustomModelEntryViaRestart) still uses the original two-step design: the launch runs a single native session exactly the way its own Run-menu entry would, then waits for the new session to go idle (GET .../wait?until=idle, bounded at 20s — a normal 200 either way, never an error, per the wait endpoint's own contract) before applying the endpoint via the restart route below. That wait exists because a freshly launched CLI reports itself as busy for its own startup (a boot spinner, a workspace-trust check) well before the apply call would otherwise reach it, and the apply route correctly refuses to restart a session mid-turn — a fresh boot looks exactly like one from the outside. A session still busy after the wait reaches the apply call anyway and gets that route's own honest SESSION_BUSY error, now visible as a sticky toast with a close button rather than a generic message that vanished in three seconds. Claude stays on this path because its own restart (--resume-based, keeping the conversation) is far less jarring than the other seven's, and runClaude()'s multi-tab launch and docker-config-drift confirm/retry loop make folding it into the one-shot path separate work. It is a one-off "try this endpoint" action, not a sticky mode: the plain Run button still means "this harness, native cloud" afterward. Entries are hidden entirely for a remote or Docker active case, since the apply route refuses both (see the next section).

Launching directly on an endpoint (no restart)

curl -sk -X POST https://localhost:3000/api/quick-start \
  -H 'Content-Type: application/json' \
  -d '{"caseName": "myapp", "mode": "codex", "customModel": {"endpointId": "llama-box", "modelId": "qwen3"}}'

POST /api/quick-start's customModel field ({endpointId, modelId, confirmed?}) computes the same injection the restart route below does, but BEFORE the session exists — the session is minted its own id up front (crypto.randomUUID()), the injection (env vars, and for a configDir-kind CLI, the written config file) targets that real id, and the session launches already pointed at the endpoint. No restart, because there was never a native-backend launch to restart away from. Runs the same llama-swap conflict check as the restart route (below) — a 409-shaped {requiresConfirmation, currentlyLoadedModel, affectedSessions} response with no session created, resolved by retrying with confirmed: true — and is refused the same way for a remote or Docker case. This is what the Run-menu picker uses for opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the restart route below (see "The Run-menu picker" above for why).

Applying a model to an ALREADY-RUNNING session

curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
  -H 'Content-Type: application/json' \
  -d '{"endpointId": "llama-box", "modelId": "qwen3"}'

This computes the CLI-specific env vars / config for that session's mode (see the recipe table in custom-model-endpoints-plan.md) and restarts the session's CLI process in place — same pane, same tmux session, fresh env. That restart is necessary, not incidental: every supported harness reads its endpoint config at process start, not per-turn, so there is no live hot-swap. A Claude session is relaunched with --resume <conversation> || --session-id <id>, so it continues the conversation it was on; pi, omp and grok are relaunched with the --model value that selects the injected provider (custom/<modelId> for pi and omp, codeman-custom for grok), since for those three the config file alone does not switch the model. Remote (SSH) and Docker sessions are refused (400) for now: their restart reattaches the durable remote/in-container tmux rather than relaunching the agent, so the selection would report success and change nothing.

Claude gets two more env vars when known/applicable, both declared on its registry entry (contextLengthVar/configDirVar), not hardcoded here:

  • CLAUDE_CODE_MAX_CONTEXT_TOKENS is set to modelId's discovered context length (see the discovery section above) whenever one is known. Without it, Claude Code assumes a large (200k) window for any unrecognized custom model id and never compacts, which reliably overflows a much smaller real local context — confirmed live: a stock ~33.7K-token system prompt against a 16384-token llama-swap model failed with exceeds the available context size. No entry for the model in modelContextLengths means the var is simply omitted, never a guess.
  • CLAUDE_CONFIG_DIR is pointed at the same isolated per-session directory the configDir-kind CLIs use (empty, no files written into it), so the injected ANTHROPIC_API_KEY never shares a directory with a stored claude.ai OAuth login. Claude Code still prints "Both claude.ai and ANTHROPIC_API_KEY set" when the two coexist in the same config directory — cosmetic (confirmed live: the API key wins for actual requests either way, visible in the terminal's own API Usage Billing line) but worth eliminating rather than living with. The directory's projects subdirectory is symlinked (a junction on Windows) back to the real ~/.claude/projects so the response viewer, subagent windows and Read My Mind keep working for that session — the same trade-off and fix documented for a manually-set CLAUDE_CONFIG_DIR in docs/wiki/Agent-CLIs.md, just applied automatically here. Best-effort: a platform that refuses the symlink keeps the pre-existing blind-response-viewer side effect rather than failing the whole custom-model apply over it.

That isolated directory needed one more fix to actually be usable non-interactively. An otherwise-empty CLAUDE_CONFIG_DIR has none of a real profile's prior "Detected a custom API key — use it?" approvals, so without more, Claude Code stops and asks that on every single launch — confirmed live, and with nobody at a TTY to answer, its own default answer ("No") silently refuses the very key this feature just injected, which looks like the endpoint being ignored entirely. customModelInjection's apiKeyTrustFile ({ relPath: '.claude.json', shape: 'claude-api-key-responses' } on claude's entry) pre-seeds that exact approval: the apply step merges customApiKeyResponses.approved: [apiKey] into <configDir>/.claude.json, the same field a real answered prompt itself writes to (confirmed against a real file after answering by hand once) — this answers the prompt in advance rather than bypassing it. The merge preserves whatever else the CLI already wrote into that file on an earlier launch in the same isolated directory (userID, numStartups, earlier approved keys), and a missing or corrupt file is treated as empty rather than failing the apply.

Clear back to the harness's native cloud default with:

curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
  -H 'Content-Type: application/json' -d '{"clear": true}'

Clearing also removes the env vars the selection injected from the tmux session (they persist there and would otherwise be inherited by the relaunched CLI) and deletes the per-session config directory (~/.codeman/custom-model-configs/<sessionId>, written 0600 because pi and omp embed the API key in it). That directory is also removed when the session is deleted. The selection survives a Codeman restart: the endpoint id, model and injected key NAMES are persisted, the values are re-derived from the endpoint store on recovery, and the pane keeps running against the endpoint in between because tmux retains its environment.

New sessions always default back to the harness's native backend. A custom-endpoint selection is a per-session choice, never a sticky global default — starting a fresh session doesn't inherit whatever the last one was pointed at.

Confidence per harness

Every harness except Antigravity has now been run end-to-end against a real llama-swap server via scripts/test-local-llm-harnesses.ts (a dynamic script that reads the live CLI registry, so a registry change is picked up automatically). Results:

  • Claude, opencode, Pi, Grok, OMP — verified: a real "hello world" reply came back through the endpoint.
  • Codex — the config is structurally correct, but Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement. This is a real protocol incompatibility, not a bug here; Codex support needs a Responses-API-compatible endpoint.
  • Gemini — fails with Invalid auth method selected, traced to an undocumented GATEWAY auth path gemini-cli selects once GOOGLE_GEMINI_BASE_URL is set. Unresolved after real investigation (several auth workarounds were tried and ruled out); do not rely on Gemini support yet.
  • DeepSeek — the request reaches the server (env vars are read) but gets a consistent HTTP_404. Root cause not identified; best-effort only.
  • Antigravity — no known custom-endpoint mechanism at all; unsupported.

See the confidence table in custom-model-endpoints-plan.md for the full detail behind each result. scripts/test-local-llm-harnesses.ts is the standalone script used to check a harness against a real endpoint outside the web UI entirely; see its own --help for usage.

Security note

Every env var this feature can set that redirects a session's traffic (ANTHROPIC_BASE_URL, GOOGLE_GEMINI_BASE_URL, CODEX_HOME, etc.) is listed in that CLI's privilegedEnvKeys in the CLI registry, so a non-granted multi-user owner cannot set one directly via the generic envOverrides API field — only through this feature's own route, which computes the value from an admin-configured, SSRF-guarded endpoint rather than trusting arbitrary client input. See the "Multi-user security hardening" section of custom-model-endpoints-plan.md for the full reasoning; several of these were reachable via the generic envOverrides field even before this feature existed, and building this surfaced and closed that gap.