Files
Codeman/docs/custom-model-endpoints.md
T
DevvynandClaude Sonnet 5 41416566aa feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).

- Registry: capabilities.customModelInjection per CLI entry (env /
  configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
  model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
  custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
  (POST /api/sessions/:id/custom-model), reusing the existing
  respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
  CLI's privilegedEnvKeys, closing a pre-existing gap where several were
  already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
  against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
  every CLI's injected values through a real HTTP shape

Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
  a top-level model string + [model_providers.custom]); fixing it then
  surfaced a genuine, documented protocol incompatibility (Codex only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement)
- Claude Code's async session-title-generation call validates
  ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
  hangs the whole -p invocation on an unrecognized name; documented for
  chunk 6, worked around in the standalone script only (--bare is NOT
  safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
  and api-key headers) reliably hung a real server; removed the option
  entirely rather than just changing the default

Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00

4.8 KiB

Custom Model Endpoint Profiles

Point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, or OMP — at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. "Custom endpoint" covers both local hardware (llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built boxes like NVIDIA DGX Spark or AMD Strix Halo mini-PCs) and cloud services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a company gateway) — anything answering GET /v1/models and POST /v1/chat/completions in the standard shape. Design doc, per-CLI recipe confidence table, and security reasoning: deployment_plan.md.

Status: backend is implemented and tested (registry capability, the injection engine, the endpoint store + discovery route, the session restart route). The toolbar picker / settings UI described below as the intended surface is not yet built — until it lands, use the HTTP API directly (examples below). Antigravity has no known custom-endpoint mechanism and is not supported.

Turning it on

App Settings → Agents & CLIs → Custom Model Endpoints (synced setting customModelEndpointsEnabled, default OFF). The API equivalent:

curl -sk -X PUT https://localhost:3000/api/settings \
  -H 'Content-Type: application/json' \
  -d '{"customModelEndpointsEnabled": true}'

Adding an endpoint

curl -sk -X POST https://localhost:3000/api/model-endpoints \
  -H 'Content-Type: application/json' \
  -d '{"id": "llama-box", "label": "Home llama.cpp", "baseUrl": "http://192.168.1.50:8080"}'

apiKey is optional (most local servers don't check it). authStyle (bearer | api-key | both, default both) controls which auth header convention discovery uses — both works whether the endpoint is llama.cpp (ignores the header) or a cloud gateway like Azure (wants api-key).

Discover its available models:

curl -sk -X POST https://localhost:3000/api/model-endpoints/llama-box/discover-models

This calls the endpoint's own GET /v1/models and stores the returned list on the endpoint record; GET /api/model-endpoints lists everything configured, PUT/DELETE /api/model-endpoints/:id update or remove one. Endpoint management is admin-only in multi-user mode, same as remote/docker hosts — these are machine-level infra, not per-user settings.

Applying a model to a session

curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
  -H 'Content-Type: application/json' \
  -d '{"endpointId": "llama-box", "modelId": "qwen3"}'

This computes the CLI-specific env vars / config for that session's mode (see the recipe table in deployment_plan.md) and restarts the session's CLI process in place — same pane, same tmux session, fresh env. That restart is necessary, not incidental: every supported harness reads its endpoint config at process start, not per-turn, so there is no live hot-swap. Clear back to the harness's native cloud default with:

curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
  -H 'Content-Type: application/json' -d '{"clear": true}'

New sessions always default back to the harness's native backend. A custom-endpoint selection is a per-session choice, never a sticky global default — starting a fresh session doesn't inherit whatever the last one was pointed at.

Confidence per harness

Only Claude, opencode, and Codex have been verified against a real llama.cpp server by hand. Gemini, Pi, Grok, DeepSeek, and OMP's recipes are correct on their one-shot invocation flags (confirmed against real installed binaries' own --help output) but their env-var/config conventions for a custom endpoint are still web-researched, not verified end-to-end — see the confidence table in deployment_plan.md before relying on one of those five in production. scripts/test-local-llm-harnesses.mjs is the standalone script used to check a harness against a real endpoint outside the web UI entirely; see its own --help for usage.

Security note

Every env var this feature can set that redirects a session's traffic (ANTHROPIC_BASE_URL, GOOGLE_GEMINI_BASE_URL, CODEX_HOME, etc.) is listed in that CLI's privilegedEnvKeys in the CLI registry, so a non-granted multi-user owner cannot set one directly via the generic envOverrides API field — only through this feature's own route, which computes the value from an admin-configured, SSRF-guarded endpoint rather than trusting arbitrary client input. See the "Multi-user security hardening" section of deployment_plan.md for the full reasoning; several of these were reachable via the generic envOverrides field even before this feature existed, and building this surfaced and closed that gap.