mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-04 14:39:42 +02:00
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.
Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.
Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
207 lines
11 KiB
Markdown
207 lines
11 KiB
Markdown
# Custom Model Endpoint Profiles
|
|
|
|
Point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi,
|
|
Grok, DeepSeek, or OMP — at a custom OpenAI-compatible endpoint instead of
|
|
its native cloud backend, for a given session. "Custom endpoint" covers both
|
|
**local** hardware (llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built
|
|
boxes like NVIDIA DGX Spark or AMD Strix Halo mini-PCs) and **cloud**
|
|
services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
|
|
company gateway) — anything answering `GET /v1/models` and
|
|
`POST /v1/chat/completions` in the standard shape. Design doc, per-CLI
|
|
recipe confidence table, and security reasoning:
|
|
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
|
|
|
|
> **Status**: fully wired end to end — registry capability, the injection
|
|
> engine, the endpoint store + discovery route, the session restart route,
|
|
> a settings-panel CRUD surface, and the Run-menu picker described below.
|
|
> Antigravity has no known custom-endpoint mechanism and is not supported.
|
|
> The HTTP API (examples below) still works directly and is what the picker
|
|
> itself calls under the hood.
|
|
|
|
## Turning it on
|
|
|
|
App Settings → Models → **Custom model endpoints** (synced setting
|
|
`customModelEndpointsEnabled`, default **OFF**). Turning it on does two
|
|
things: it reveals the endpoint list/add/edit/discover panel in that same
|
|
settings section, and it makes the Run menu offer a generated entry per
|
|
(harness, endpoint) pair — see "The Run-menu picker" below. The API
|
|
equivalent:
|
|
|
|
```bash
|
|
curl -sk -X PUT https://localhost:3000/api/settings \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"customModelEndpointsEnabled": true}'
|
|
```
|
|
|
|
## Adding an endpoint
|
|
|
|
Via App Settings → Models → Custom model endpoints → **+ Add endpoint**, or
|
|
directly:
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/model-endpoints \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"id": "llama-box", "label": "Home llama.cpp", "baseUrl": "http://192.168.1.50:8080"}'
|
|
```
|
|
|
|
`apiKey` is optional (most local servers don't check it). `authStyle`
|
|
(`bearer` | `api-key`, default `bearer`) controls which auth header
|
|
convention discovery uses: `bearer` is `Authorization: Bearer <key>`
|
|
(llama.cpp, OpenAI-compatible servers, most gateways), `api-key` is the
|
|
`api-key: <key>` header Azure AI Foundry wants. There is deliberately no
|
|
"send both" option: measured against a real llama-swap server, a request
|
|
carrying both headers hung indefinitely. `baseUrl` must be `http(s)`, carry
|
|
no embedded credentials, and may not point at a link-local or cloud-metadata
|
|
address; discovery re-checks the address the name actually resolves to.
|
|
|
|
Discover its available models:
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/model-endpoints/llama-box/discover-models
|
|
```
|
|
|
|
This calls the endpoint's own `GET /v1/models` and stores the returned list
|
|
on the endpoint record; `GET /api/model-endpoints` lists everything
|
|
configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
|
|
Endpoint management is admin-only in multi-user mode, same as remote/docker
|
|
hosts — these are machine-level infra, not per-user settings.
|
|
|
|
`defaultModelId` names which discovered model the picker pre-marks for that
|
|
endpoint — the settings panel's Edit form exposes it as a select populated
|
|
from the endpoint's own discovered `models`, and the route refuses a value
|
|
that isn't one of them. It is applied automatically only when the endpoint
|
|
has exactly one discovered model (nothing to choose); with two or more it
|
|
is a pre-selection in the model-picker dialog below, never a silent default.
|
|
Re-discovering drops a default that no longer appears in the fresh list
|
|
rather than carrying an invalid one forward.
|
|
|
|
**Model lists refresh themselves.** A background sweep (`server.ts`,
|
|
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
|
|
saved endpoint the same way the manual `POST .../discover-models` route
|
|
does, best-effort per endpoint — one being unreachable on a given cycle
|
|
never blocks the others. Off under `npm test`, same reasoning as the Codex
|
|
plan-usage poll it sits beside: no real network to hit, no server instance
|
|
to keep the timer alive for.
|
|
|
|
## The Run-menu picker
|
|
|
|
With the setting on and at least one endpoint carrying a discovered model,
|
|
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
|
|
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
|
|
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
|
|
registry's own `capabilities.customModelInjection` at page render
|
|
(`window.__codemanCustomModelClis`, `server.ts`) — never a hardcoded id list
|
|
in the frontend — so a CLI whose injection recipe lands later shows up with
|
|
no frontend change, and Antigravity (`unsupported`) never does.
|
|
|
|
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
|
|
`session-ui.js`) rather than trusting anything cached from the dropdown's
|
|
own render — the model list can have changed via the 5-minute sweep above
|
|
or a settings-panel edit since the menu opened. With exactly one discovered
|
|
model it runs straight away; with two or more, a small modal
|
|
(`#customModelPickModal`) lists them and asks which one to use for this
|
|
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
|
|
the point of asking is letting one launch deliberately differ from the
|
|
saved default, not just confirming it. Whichever way the model was decided,
|
|
the launch itself runs a single session on that harness exactly the way its
|
|
own Run-menu entry would (same case creation, env overrides, everything),
|
|
then **waits for the new session to go idle** (`GET .../wait?until=idle`,
|
|
bounded at 20s — a normal 200 either way, never an error, per the wait
|
|
endpoint's own contract) before applying the endpoint and model to it via
|
|
the route below. That wait exists because a freshly launched CLI reports
|
|
itself as `busy` for its own startup (a boot spinner, a workspace-trust
|
|
check) well before the apply call would otherwise reach it, and the apply
|
|
route correctly refuses to restart a session mid-turn — a fresh boot looks
|
|
exactly like one from the outside. A session still busy after the wait
|
|
reaches the apply call anyway and gets that route's own honest
|
|
`SESSION_BUSY` error, now visible as a sticky toast with a close button
|
|
rather than a generic message that vanished in three seconds. It is a
|
|
one-off "try this endpoint" action, not a sticky mode: the plain Run button
|
|
still means "this harness, native cloud" afterward. Entries are hidden
|
|
entirely for a remote or Docker active case, since the apply route refuses
|
|
both (see the next section).
|
|
|
|
## Applying a model to a session
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"endpointId": "llama-box", "modelId": "qwen3"}'
|
|
```
|
|
|
|
This computes the CLI-specific env vars / config for that session's mode
|
|
(see the recipe table in `custom-model-endpoints-plan.md`) and **restarts the session's
|
|
CLI process in place** — same pane, same tmux session, fresh env. That
|
|
restart is necessary, not incidental: every supported harness reads its
|
|
endpoint config at process start, not per-turn, so there is no live
|
|
hot-swap. A Claude session is relaunched with `--resume <conversation> ||
|
|
--session-id <id>`, so it continues the conversation it was on; pi, omp and
|
|
grok are relaunched with the `--model` value that selects the injected
|
|
provider (`custom/<modelId>` for pi and omp, `codeman-custom` for grok),
|
|
since for those three the config file alone does not switch the model.
|
|
**Remote (SSH) and Docker sessions are refused** (400) for now: their restart
|
|
reattaches the durable remote/in-container tmux rather than relaunching the
|
|
agent, so the selection would report success and change nothing.
|
|
|
|
Clear back to the harness's native cloud default with:
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
|
-H 'Content-Type: application/json' -d '{"clear": true}'
|
|
```
|
|
|
|
Clearing also removes the env vars the selection injected from the tmux
|
|
session (they persist there and would otherwise be inherited by the
|
|
relaunched CLI) and deletes the per-session config directory
|
|
(`~/.codeman/custom-model-configs/<sessionId>`, written 0600 because pi and
|
|
omp embed the API key in it). That directory is also removed when the
|
|
session is deleted. The selection survives a Codeman restart: the endpoint
|
|
id, model and injected key NAMES are persisted, the values are re-derived
|
|
from the endpoint store on recovery, and the pane keeps running against the
|
|
endpoint in between because tmux retains its environment.
|
|
|
|
**New sessions always default back to the harness's native backend.** A
|
|
custom-endpoint selection is a per-session choice, never a sticky global
|
|
default — starting a fresh session doesn't inherit whatever the last one was
|
|
pointed at.
|
|
|
|
## Confidence per harness
|
|
|
|
Every harness except Antigravity has now been run end-to-end against a real
|
|
llama-swap server via `scripts/test-local-llm-harnesses.ts` (a dynamic
|
|
script that reads the live CLI registry, so a registry change is picked up
|
|
automatically). Results:
|
|
|
|
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
|
|
came back through the endpoint.
|
|
- **Codex** — the config is structurally correct, but Codex only speaks the
|
|
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
|
|
This is a real protocol incompatibility, not a bug here; Codex support
|
|
needs a Responses-API-compatible endpoint.
|
|
- **Gemini** — fails with `Invalid auth method selected`, traced to an
|
|
undocumented `GATEWAY` auth path gemini-cli selects once
|
|
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
|
|
(several auth workarounds were tried and ruled out); do not rely on
|
|
Gemini support yet.
|
|
- **DeepSeek** — the request reaches the server (env vars are read) but
|
|
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
|
|
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
|
|
|
|
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
|
|
each result. `scripts/test-local-llm-harnesses.ts` is the standalone script
|
|
used to check a harness against a real endpoint outside the web UI
|
|
entirely; see its own `--help` for usage.
|
|
|
|
## Security note
|
|
|
|
Every env var this feature can set that redirects a session's traffic
|
|
(`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, etc.) is
|
|
listed in that CLI's `privilegedEnvKeys` in the CLI registry, so a
|
|
non-granted multi-user owner cannot set one directly via the generic
|
|
`envOverrides` API field — only through this feature's own route, which
|
|
computes the value from an admin-configured, SSRF-guarded endpoint rather
|
|
than trusting arbitrary client input. See the "Multi-user security
|
|
hardening" section of `custom-model-endpoints-plan.md` for the full reasoning; several
|
|
of these were reachable via the generic `envOverrides` field even before
|
|
this feature existed, and building this surfaced and closed that gap.
|