mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-04 06:29:42 +02:00
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.
Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
270 lines
15 KiB
Markdown
270 lines
15 KiB
Markdown
# Custom Model Endpoint Profiles
|
|
|
|
Point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi,
|
|
Grok, DeepSeek, or OMP — at a custom OpenAI-compatible endpoint instead of
|
|
its native cloud backend, for a given session. "Custom endpoint" covers both
|
|
**local** hardware (llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built
|
|
boxes like NVIDIA DGX Spark or AMD Strix Halo mini-PCs) and **cloud**
|
|
services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
|
|
company gateway) — anything answering `GET /v1/models` and
|
|
`POST /v1/chat/completions` in the standard shape. Design doc, per-CLI
|
|
recipe confidence table, and security reasoning:
|
|
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
|
|
|
|
> **Status**: fully wired end to end — registry capability, the injection
|
|
> engine, the endpoint store + discovery route, the session restart route,
|
|
> a settings-panel CRUD surface, and the Run-menu picker described below.
|
|
> Antigravity has no known custom-endpoint mechanism and is not supported.
|
|
> The HTTP API (examples below) still works directly and is what the picker
|
|
> itself calls under the hood.
|
|
|
|
## Turning it on
|
|
|
|
App Settings → Models → **Custom model endpoints** (synced setting
|
|
`customModelEndpointsEnabled`, default **OFF**). Turning it on does two
|
|
things: it reveals the endpoint list/add/edit/discover panel in that same
|
|
settings section, and it makes the Run menu offer a generated entry per
|
|
(harness, endpoint) pair — see "The Run-menu picker" below. The API
|
|
equivalent:
|
|
|
|
```bash
|
|
curl -sk -X PUT https://localhost:3000/api/settings \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"customModelEndpointsEnabled": true}'
|
|
```
|
|
|
|
## Adding an endpoint
|
|
|
|
Via App Settings → Models → Custom model endpoints → **+ Add endpoint**, or
|
|
directly:
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/model-endpoints \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"id": "llama-box", "label": "Home llama.cpp", "baseUrl": "http://192.168.1.50:8080"}'
|
|
```
|
|
|
|
`apiKey` is optional (most local servers don't check it). `authStyle`
|
|
(`bearer` | `api-key`, default `bearer`) controls which auth header
|
|
convention discovery uses: `bearer` is `Authorization: Bearer <key>`
|
|
(llama.cpp, OpenAI-compatible servers, most gateways), `api-key` is the
|
|
`api-key: <key>` header Azure AI Foundry wants. There is deliberately no
|
|
"send both" option: measured against a real llama-swap server, a request
|
|
carrying both headers hung indefinitely. `baseUrl` must be `http(s)`, carry
|
|
no embedded credentials, and may not point at a link-local or cloud-metadata
|
|
address; discovery re-checks the address the name actually resolves to.
|
|
|
|
Discover its available models:
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/model-endpoints/llama-box/discover-models
|
|
```
|
|
|
|
This calls the endpoint's own `GET /v1/models` and stores the returned list
|
|
on the endpoint record; `GET /api/model-endpoints` lists everything
|
|
configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
|
|
Endpoint management is admin-only in multi-user mode, same as remote/docker
|
|
hosts — these are machine-level infra, not per-user settings.
|
|
|
|
**Context length is discovered too, opportunistically and safely.** The plain
|
|
`GET /v1/models` response has no context-window field, but llama.cpp's
|
|
llama-swap-proxied `GET /props?model=<id>` does (`n_ctx`). Discovery only ever
|
|
calls it for a model llama-swap's own response already reports
|
|
`status.value === "loaded"` for — never for an unloaded one, because
|
|
llama-swap treats `?model=` as a routing hint and asking about a model that
|
|
isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side
|
|
effect of what should be read-only discovery. A server with no `status` field
|
|
on any entry at all (not llama-swap) gets no context-length enrichment,
|
|
rather than guessing. A model's previously-learned context length survives a
|
|
later cycle where it wasn't the loaded one; it's dropped only once the model
|
|
disappears from the endpoint's list entirely. Stored per model in
|
|
`modelContextLengths` and applied automatically (see "Applying a model to a
|
|
session" below) so a CLI that would otherwise assume a large default context
|
|
window for an unrecognized model id stops silently overflowing a much
|
|
smaller real one.
|
|
|
|
`defaultModelId` names which discovered model the picker pre-marks for that
|
|
endpoint — the settings panel's Edit form exposes it as a select populated
|
|
from the endpoint's own discovered `models`, and the route refuses a value
|
|
that isn't one of them. It is applied automatically only when the endpoint
|
|
has exactly one discovered model (nothing to choose); with two or more it
|
|
is a pre-selection in the model-picker dialog below, never a silent default.
|
|
Re-discovering drops a default that no longer appears in the fresh list
|
|
rather than carrying an invalid one forward.
|
|
|
|
**Model lists refresh themselves.** A background sweep (`server.ts`,
|
|
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
|
|
saved endpoint the same way the manual `POST .../discover-models` route
|
|
does, best-effort per endpoint — one being unreachable on a given cycle
|
|
never blocks the others. Off under `npm test`, same reasoning as the Codex
|
|
plan-usage poll it sits beside: no real network to hit, no server instance
|
|
to keep the timer alive for.
|
|
|
|
## The Run-menu picker
|
|
|
|
With the setting on and at least one endpoint carrying a discovered model,
|
|
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
|
|
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
|
|
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
|
|
registry's own `capabilities.customModelInjection` at page render
|
|
(`window.__codemanCustomModelClis`, `server.ts`) — never a hardcoded id list
|
|
in the frontend — so a CLI whose injection recipe lands later shows up with
|
|
no frontend change, and Antigravity (`unsupported`) never does.
|
|
|
|
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
|
|
`session-ui.js`) rather than trusting anything cached from the dropdown's
|
|
own render — the model list can have changed via the 5-minute sweep above
|
|
or a settings-panel edit since the menu opened. With exactly one discovered
|
|
model it runs straight away; with two or more, a small modal
|
|
(`#customModelPickModal`) lists them and asks which one to use for this
|
|
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
|
|
the point of asking is letting one launch deliberately differ from the
|
|
saved default, not just confirming it. Whichever way the model was decided,
|
|
the launch itself runs a single session on that harness exactly the way its
|
|
own Run-menu entry would (same case creation, env overrides, everything),
|
|
then **waits for the new session to go idle** (`GET .../wait?until=idle`,
|
|
bounded at 20s — a normal 200 either way, never an error, per the wait
|
|
endpoint's own contract) before applying the endpoint and model to it via
|
|
the route below. That wait exists because a freshly launched CLI reports
|
|
itself as `busy` for its own startup (a boot spinner, a workspace-trust
|
|
check) well before the apply call would otherwise reach it, and the apply
|
|
route correctly refuses to restart a session mid-turn — a fresh boot looks
|
|
exactly like one from the outside. A session still busy after the wait
|
|
reaches the apply call anyway and gets that route's own honest
|
|
`SESSION_BUSY` error, now visible as a sticky toast with a close button
|
|
rather than a generic message that vanished in three seconds. It is a
|
|
one-off "try this endpoint" action, not a sticky mode: the plain Run button
|
|
still means "this harness, native cloud" afterward. Entries are hidden
|
|
entirely for a remote or Docker active case, since the apply route refuses
|
|
both (see the next section).
|
|
|
|
## Applying a model to a session
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"endpointId": "llama-box", "modelId": "qwen3"}'
|
|
```
|
|
|
|
This computes the CLI-specific env vars / config for that session's mode
|
|
(see the recipe table in `custom-model-endpoints-plan.md`) and **restarts the session's
|
|
CLI process in place** — same pane, same tmux session, fresh env. That
|
|
restart is necessary, not incidental: every supported harness reads its
|
|
endpoint config at process start, not per-turn, so there is no live
|
|
hot-swap. A Claude session is relaunched with `--resume <conversation> ||
|
|
--session-id <id>`, so it continues the conversation it was on; pi, omp and
|
|
grok are relaunched with the `--model` value that selects the injected
|
|
provider (`custom/<modelId>` for pi and omp, `codeman-custom` for grok),
|
|
since for those three the config file alone does not switch the model.
|
|
**Remote (SSH) and Docker sessions are refused** (400) for now: their restart
|
|
reattaches the durable remote/in-container tmux rather than relaunching the
|
|
agent, so the selection would report success and change nothing.
|
|
|
|
**Claude gets two more env vars when known/applicable, both declared on its
|
|
registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
|
|
|
|
- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context
|
|
length (see the discovery section above) whenever one is known. Without
|
|
it, Claude Code assumes a large (200k) window for any unrecognized custom
|
|
model id and never compacts, which reliably overflows a much smaller real
|
|
local context — confirmed live: a stock ~33.7K-token system prompt against
|
|
a 16384-token llama-swap model failed with `exceeds the available context
|
|
size`. No entry for the model in `modelContextLengths` means the var is
|
|
simply omitted, never a guess.
|
|
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
|
|
the `configDir`-kind CLIs use (empty, no files written into it), so the
|
|
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
|
|
claude.ai OAuth login. Claude Code still prints "Both claude.ai and
|
|
ANTHROPIC_API_KEY set" when the two coexist in the same config directory —
|
|
cosmetic (confirmed live: the API key wins for actual requests either way,
|
|
visible in the terminal's own `API Usage Billing` line) but worth
|
|
eliminating rather than living with. The directory's `projects`
|
|
subdirectory is symlinked (a junction on Windows) back to the real
|
|
`~/.claude/projects` so the response viewer, subagent windows and Read My
|
|
Mind keep working for that session — the same trade-off and fix documented
|
|
for a manually-set `CLAUDE_CONFIG_DIR` in
|
|
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
|
|
automatically here. Best-effort: a platform that refuses the symlink keeps
|
|
the pre-existing blind-response-viewer side effect rather than failing the
|
|
whole custom-model apply over it.
|
|
|
|
**That isolated directory needed one more fix to actually be usable
|
|
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
|
|
real profile's prior "Detected a custom API key — use it?" approvals, so
|
|
without more, Claude Code stops and asks that on *every single launch* —
|
|
confirmed live, and with nobody at a TTY to answer, its own default answer
|
|
("No") silently refuses the very key this feature just injected, which
|
|
looks like the endpoint being ignored entirely. `customModelInjection`'s
|
|
`apiKeyTrustFile` (`{ relPath: '.claude.json', shape:
|
|
'claude-api-key-responses' }` on claude's entry) pre-seeds that exact
|
|
approval: the apply step merges `customApiKeyResponses.approved: [apiKey]`
|
|
into `<configDir>/.claude.json`, the same field a real answered prompt
|
|
itself writes to (confirmed against a real file after answering by hand
|
|
once) — this answers the prompt in advance rather than bypassing it. The
|
|
merge preserves whatever else the CLI already wrote into that file on an
|
|
earlier launch in the same isolated directory (`userID`, `numStartups`,
|
|
earlier approved keys), and a missing or corrupt file is treated as empty
|
|
rather than failing the apply.
|
|
|
|
Clear back to the harness's native cloud default with:
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
|
-H 'Content-Type: application/json' -d '{"clear": true}'
|
|
```
|
|
|
|
Clearing also removes the env vars the selection injected from the tmux
|
|
session (they persist there and would otherwise be inherited by the
|
|
relaunched CLI) and deletes the per-session config directory
|
|
(`~/.codeman/custom-model-configs/<sessionId>`, written 0600 because pi and
|
|
omp embed the API key in it). That directory is also removed when the
|
|
session is deleted. The selection survives a Codeman restart: the endpoint
|
|
id, model and injected key NAMES are persisted, the values are re-derived
|
|
from the endpoint store on recovery, and the pane keeps running against the
|
|
endpoint in between because tmux retains its environment.
|
|
|
|
**New sessions always default back to the harness's native backend.** A
|
|
custom-endpoint selection is a per-session choice, never a sticky global
|
|
default — starting a fresh session doesn't inherit whatever the last one was
|
|
pointed at.
|
|
|
|
## Confidence per harness
|
|
|
|
Every harness except Antigravity has now been run end-to-end against a real
|
|
llama-swap server via `scripts/test-local-llm-harnesses.ts` (a dynamic
|
|
script that reads the live CLI registry, so a registry change is picked up
|
|
automatically). Results:
|
|
|
|
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
|
|
came back through the endpoint.
|
|
- **Codex** — the config is structurally correct, but Codex only speaks the
|
|
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
|
|
This is a real protocol incompatibility, not a bug here; Codex support
|
|
needs a Responses-API-compatible endpoint.
|
|
- **Gemini** — fails with `Invalid auth method selected`, traced to an
|
|
undocumented `GATEWAY` auth path gemini-cli selects once
|
|
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
|
|
(several auth workarounds were tried and ruled out); do not rely on
|
|
Gemini support yet.
|
|
- **DeepSeek** — the request reaches the server (env vars are read) but
|
|
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
|
|
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
|
|
|
|
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
|
|
each result. `scripts/test-local-llm-harnesses.ts` is the standalone script
|
|
used to check a harness against a real endpoint outside the web UI
|
|
entirely; see its own `--help` for usage.
|
|
|
|
## Security note
|
|
|
|
Every env var this feature can set that redirects a session's traffic
|
|
(`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, etc.) is
|
|
listed in that CLI's `privilegedEnvKeys` in the CLI registry, so a
|
|
non-granted multi-user owner cannot set one directly via the generic
|
|
`envOverrides` API field — only through this feature's own route, which
|
|
computes the value from an admin-configured, SSRF-guarded endpoint rather
|
|
than trusting arbitrary client input. See the "Multi-user security
|
|
hardening" section of `custom-model-endpoints-plan.md` for the full reasoning; several
|
|
of these were reachable via the generic `envOverrides` field even before
|
|
this feature existed, and building this surfaced and closed that gap.
|