Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. Covers local hardware (llama.cpp, Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter). Off by default (customModelEndpointsEnabled, synced, default OFF). - Registry: capabilities.customModelInjection per CLI entry (env / configContentEnv / configDir / unsupported kinds) - Pure injection builder (custom-model-injection.ts) turning an endpoint + model id into the real env vars / config content per CLI - Endpoint store + CRUD routes (custom-model-hosts.ts, custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded - Session integration: Session.setCustomModel()/restartCli() (POST /api/sessions/:id/custom-model), reusing the existing respawn-pane -k primitive to restart the CLI process with new env - Multi-user hardening: every new redirect-capable env var added to its CLI's privilegedEnvKeys, closing a pre-existing gap where several were already reachable via the generic envOverrides field's prefix allowlist - Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries against a real endpoint outside the web UI, independent of tmux/sessions - Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying every CLI's injected values through a real HTTP shape Real end-to-end validation against a live llama-swap server (inside a codeman/agent:llm-test Docker image with all 9 CLI binaries) found and fixed three real bugs before they shipped: - Codex's config.toml schema was wrong ([model].default table instead of a top-level model string + [model_providers.custom]); fixing it then surfaced a genuine, documented protocol incompatibility (Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement) - Claude Code's async session-title-generation call validates ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and hangs the whole -p invocation on an unrecognized name; documented for chunk 6, worked around in the standalone script only (--bare is NOT safe for a real interactive session, which needs hooks) - The discovery route's authStyle: 'both' option (send both Authorization and api-key headers) reliably hung a real server; removed the option entirely rather than just changing the default Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see PR.md and deployment_plan.md for the full chunk breakdown and confidence table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
13 KiB
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses)
⭐ Shout-out up front: this feature was partly inspired by — and is a great fit for — Ark0N/Qwen5090, the maintainer's other project: a one-click Windows / one-command Linux installer that stands up Qwen3.8-27B locally on an RTX 5090 behind an OpenAI-compatible API (vLLM / NInfer / llama.cpp). Once this feature lands, pointing Codeman at a Qwen5090 box is just adding one endpoint entry — no extra code, no special-casing. Qwen5090 already wires up DeepSeek Harness and Claude Code as local coding agents itself, which is basically this feature's idea in miniature. 🙂
Status: draft / work-in-progress. This PR is not ready to merge — see Status below for exactly what's done and what's still open.
What
Adds a settings-gated (default OFF) way to point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP, or Antigravity — at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. "Custom endpoint" covers both:
- Local hardware: llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built on-prem boxes like NVIDIA DGX Spark, AMD Strix Halo (Ryzen AI Max) mini-PCs, or the Qwen5090 setup above.
- Cloud: Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a company's self-hosted gateway.
The user adds an endpoint by base URL (+ optional API key), Codeman
discovers its available models via GET /v1/models, and a new toolbar
picker lets them apply one of those models to a session — which then
restarts that session's CLI process pointed at the endpoint.
Why discovery instead of asking the user to type a model name: it turns
"go read your inference server's docs to find the exact model identifier it
expects" into "pick from a list Codeman already fetched" — one less place
for a user to get a name/casing wrong and have a harness fail with an
opaque "model not found." It also means this feature works unmodified
against multi-model hosting setups, not just a single-model server: a
gateway like llama-swap
(or vLLM/LiteLLM/Ollama serving several loaded/loadable models behind one
/v1/models list) already advertises every model it can hot-swap to, so
the toolbar picker becomes a live menu of everything that endpoint can
serve — no per-model endpoint entries, no separate configuration step,
just "add the gateway once, everything behind it shows up."
Why
The maintainer pays for a Claude Code subscription but also runs a capable
local model. Every harness Codeman drives already has its own mechanism
for pointing at a custom endpoint (env vars for Claude, a JSON config blob
for opencode, a TOML file for Codex, etc.) — Codeman just never exposed a
UI for it. Full motivation, the per-CLI recipe table, and the on-prem
hardware use cases are written up in deployment_plan.md.
How
src/config/cli-registry/{types,schema,stock}.ts— newcapabilities.customModelInjectionfield per CLI entry, one of four kinds:env(Claude, Gemini, Grok, DeepSeek),configContentEnv(opencode, reusing its existingOPENCODE_CONFIG_CONTENTmechanism),configDir(Codex/Pi/OMP — writes an isolated config file, never touches the user's real one), orunsupported(Antigravity — no known mechanism, toolbar entry stays disabled). Declared data-driven per the repo's existing "never branch on CLI id" rule.src/custom-model-injection.ts— pure function turning(CliEntry, endpoint, modelId)into the real env vars / config content. No IO; a caller writesconfigDirfiles to disk.src/custom-model-hosts.ts+src/web/routes/custom-model-routes.ts— read/write-array endpoint store (~/.codeman/custom-model-hosts.json, same shape asremote-hosts.ts) andGET/POST/PUT/DELETE /api/model-endpoints+POST /:id/discover-models, admin-gated in multi-user mode, SSRF-guarded via the sameisBlockedWebviewUrl()check web tabs use.src/web/schemas.ts—customModelEndpointsEnabled(synced, default OFF) + the endpoint payload schema.scripts/test-local-llm-harnesses.mjs— standalone smoke-test script that spawns each real CLI binary one-shot against a real endpoint and checks it can answer "hello world," independent of the web UI. Reads defaults from a gitignoredscripts/local-llm-test.config.json(see the committed.example.json) so real IPs/keys never land in git.
A finding along the way: multi-user privilege hardening
Building this surfaced that several env vars (GOOGLE_GEMINI_BASE_URL,
GROK_BASE_URL, CODEX_HOME, PI_CONFIG_DIR, OPENCODE_CONFIG_CONTENT,
etc.) were already reachable via the generic envOverrides API today,
pre-existing this PR, because Codeman's env allowlist is prefix-based and
global. A non-granted multi-user owner could already redirect a session's
endpoint/credentials via a plain envOverrides field. This PR adds all of
them to their CLI's privilegedEnvKeys (the existing clamp mechanism
DEEPSEEK_BASE_URL already used), closing that gap rather than widening it.
CODEX_HOME and PI_CONFIG_DIR are flagged as extra-sensitive: a
redirected config dir can restate approval/sandbox policy or, for Pi,
redirect to a dir Pi will execute .pi/extensions TypeScript from.
Claude is the deliberate exception: ANTHROPIC_* is not added to
Claude's allowed env prefixes at all, so it stays reachable only through
the dedicated, admin-configured, SSRF-guarded custom-model route — never
through a plain client-supplied envOverrides.
Status
Built in reviewable chunks; ✅ = done and verified (typecheck + lint + format + tests green), ⬜ = not started.
- ✅ 1. Registry types —
customModelInjectioncapability shape - ✅ 2. Pure injection builder —
custom-model-injection.ts+ 15 unit tests - ✅ 3. Endpoint store + CRUD routes —
custom-model-hosts.ts,custom-model-routes.ts, discovery + SSRF guard, 7 route tests - ✅ 4. Settings + security hardening —
customModelEndpointsEnabledflag,privilegedEnvKeysadditions across 7 CLI entries - ✅ 5. Session integration —
session.customModelstate field,session.setCustomModel()/session.restartCli()(a generalized, de-restrictedreattachRemote()reusing the existingrespawn-pane -kprimitive),POST /api/sessions/:id/custom-modelrestart route. 5 new route tests; the existingsession.test.ts/session-cleanup.test.tssuites can't run at all on this Windows dev box (no localtmux— confirmed identical on unmodifiedmaster, not a regression), which is exactly why the container test below matters. - ⬜ 6. Frontend — settings group, toolbar picker, tab badge
- ✅ 7. Mock-server contract tests —
test/fixtures/mock-openai-server.tsandtest/custom-model-injection-contract.test.ts, 10 tests replaying every CLI's real injected values through an HTTP call shaped the way that CLI sends it, against an in-process fake server - ✅ 8. Docs —
docs/custom-model-endpoints.md(user guide, HTTP-API-only until chunk 6 lands) + a CLAUDE.md pointer bullet
Also done outside the chunk list: the standalone
scripts/test-local-llm-harnesses.mjs smoke-test script + its gitignored
config file, the on-prem-hardware use-case writeup in deployment_plan.md
(DGX Spark, Strix Halo, Qwen5090), and a codeman/agent:llm-test Docker
image (all 9 CLI binaries, built from docker/agent.Dockerfile) for the
real end-to-end test against a live llama-swap server.
Testing performed so far
npm run typecheck— clean after every chunknpm run lint/npx prettier --check— cleannpm test -- test/cli-registry test/custom-model-injection.test.ts test/custom-model-injection-contract.test.ts test/routes/custom-model-routes.test.ts test/routes/session-custom-model.test.ts test/routes/external-cli-bypass-clamp.test.ts— 245+ tests passing, including the existing multi-user clamp suite (no regressions from theprivilegedEnvKeysadditions)node --check scripts/test-local-llm-harnesses.mjs+ manual--helprun- Real end-to-end run against the maintainer's live llama-swap server
(
http://10.10.11.241:8080), insidecodeman/agent:llm-test(all 9 CLI binaries, built viadocker/agent.Dockerfile), against the smallest available model (qwen3.5-0.8b-ud-q8_k_xl, 1.1GB — picked by parsing the server's own reported model sizes). Real findings, not simulated:- opencode: PASS. Genuinely round-tripped a "hello world" reply through the real endpoint.
- codex: real bug found and fixed. The recipe's TOML shape
(
[model].default) was rejected by a real codex binary ("invalid type: map, expected a string") — codex wants a top-levelmodelstring plus[model_providers.custom], and the API key rides as anenv_key-named env var, never a literal TOML field. Fixed incustom-model-injection.ts, the standalone script, and both test suites. Then a second, deeper finding: codex only speaks the Responses API now (wire_api = "responses", the only value it accepts since dropping"chat"support in Feb 2026) — a real run against the now-correctly-shaped config still failed (Reconnecting...× 5, then "high demand" errors) because llama-swap doesn't implement/v1/responses. This is a genuine, currently-unresolved protocol incompatibility, not a bug in this PR's code — documented prominently indeployment_plan.md's confidence table. - claude: PASS, after two real bugs found and fixed. (1) Claude
Code's async session-title-generation call also uses
ANTHROPIC_DEFAULT_HAIKU_MODELand validates it against Claude's own internal recognized-model list, printing[claude-code:unrecognized_model]and, in-pmode, hanging the whole invocation rather than just warning.--settings '{"autoTitle":false}'does NOT stop it (confirmed);--baredoes — the warning still prints, but the real prompt now runs and returns the real answer. ⚠️--bareis only safe for this standalone one-shot test script — it also disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it must NEVER be applied to a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc.). Whether an INTERACTIVE session with a custom model hits the same hang (vs. just a background warning, which would be harmless) is untested — flagged as an open item for chunk 5/6, not assumed either way. (2) A separate, genuinely nasty bug in the test script itself: aPOSTissued right after aGETin the same Node process reliably HUNG indefinitely against this real server (reproduced repeatedly: GET alone ~30ms, POST alone ~1-2s, GET-then-immediate-POST times out completely; a 2s pause between them fixed it every time) — looks like Node's fetch/undici reusing a pooled keep-alive connection the server doesn't handle cleanly for a second request right behind a first. Fixed with a 2s pause between the script's discovery GET and its baseline POST. This is a tooling-correctness fix (affects the script's own baseline check), not a claim about how any CLI's own HTTP client behaves. - Also found and fixed: an earlier design sent BOTH
Authorization: Bearerandapi-keyauth header conventions on every discovery/ baseline request, on the theory that an unused header is harmless. Live-tested against the real server, sending both reliably HUNG the request (reproduced 3×: either header alone ~500-600ms, both together no response inside 15s). Removed the'both'option entirely fromCustomModelAuthStyle(was'bearer' | 'api-key' | 'both', now just the first two, default'bearer') — in the schema, the store type, the discovery route, and the standalone script (--auth-styleflag added). This was a real, currently-shipped-in-this-PR bug fixed before it ever reached anyone, not a pre-existing one.
Not yet done / open questions for review
- Six of nine per-CLI recipes (Gemini, Pi, Grok, DeepSeek, OMP) are
web-researched, not verified against real binaries — see the
confidence table in
deployment_plan.md. Antigravity has no known mechanism at all and stays unsupported. - Chunk 5's session-restart design needs a careful look before implementation: switching a session's endpoint restarts its CLI process in place (confirmed acceptable with the maintainer — these harnesses read endpoint config at process start, not per-turn).
🤖 Generated with Claude Code