mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-03 22:19:42 +02:00
The loading banner now shows a live countdown against its own timeout (updated every poll, so every second by default) instead of a static "this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1-3 min) on llama-swap - 47s remaining". If the countdown reaches zero and the model still isn't ready, this is now treated as a real failure rather than a "keep waiting" shrug: - The banner turns into a sticky error (_showCenterStatus gains a `type` option - 'error' drops the spinner and adds a close button, since nothing is "in progress" anymore and a sticky message needs a way to dismiss it), naming the llama-swap server's own logs as where to look for detail. - The session that load was for is closed automatically (closeSession) - requested explicitly: a console left open and pointed at a model that never finished loading is worse than no console at all. Both apply paths now thread the new session's id through to _watchLlamaSwapLoading for this (new required 3rd parameter, after endpointId/modelId). _watchLlamaSwapGeneration's existing stale-call guard extends naturally to this: a superseded call's own eventual timeout recognises it no longer owns the banner and neither shows the error nor closes a session that may by then belong to a different, newer launch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
356 lines
20 KiB
Markdown
356 lines
20 KiB
Markdown
# Custom Model Endpoint Profiles
|
|
|
|
Point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi,
|
|
Grok, DeepSeek, or OMP — at a custom OpenAI-compatible endpoint instead of
|
|
its native cloud backend, for a given session. "Custom endpoint" covers both
|
|
**local** hardware (llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built
|
|
boxes like NVIDIA DGX Spark or AMD Strix Halo mini-PCs) and **cloud**
|
|
services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
|
|
company gateway) — anything answering `GET /v1/models` and
|
|
`POST /v1/chat/completions` in the standard shape. Design doc, per-CLI
|
|
recipe confidence table, and security reasoning:
|
|
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
|
|
|
|
> **Status**: fully wired end to end — registry capability, the injection
|
|
> engine, the endpoint store + discovery route, the session restart route,
|
|
> a settings-panel CRUD surface, and the Run-menu picker described below.
|
|
> Antigravity has no known custom-endpoint mechanism and is not supported.
|
|
> The HTTP API (examples below) still works directly and is what the picker
|
|
> itself calls under the hood.
|
|
|
|
## Turning it on
|
|
|
|
App Settings → Models → **Custom model endpoints** (synced setting
|
|
`customModelEndpointsEnabled`, default **OFF**). Turning it on does two
|
|
things: it reveals the endpoint list/add/edit/discover panel in that same
|
|
settings section, and it makes the Run menu offer a generated entry per
|
|
(harness, endpoint) pair — see "The Run-menu picker" below. The API
|
|
equivalent:
|
|
|
|
```bash
|
|
curl -sk -X PUT https://localhost:3000/api/settings \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"customModelEndpointsEnabled": true}'
|
|
```
|
|
|
|
## Adding an endpoint
|
|
|
|
Via App Settings → Models → Custom model endpoints → **+ Add endpoint**, or
|
|
directly:
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/model-endpoints \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"id": "llama-box", "label": "Home llama.cpp", "baseUrl": "http://192.168.1.50:8080"}'
|
|
```
|
|
|
|
`apiKey` is optional (most local servers don't check it). `authStyle`
|
|
(`bearer` | `api-key`, default `bearer`) controls which auth header
|
|
convention discovery uses: `bearer` is `Authorization: Bearer <key>`
|
|
(llama.cpp, OpenAI-compatible servers, most gateways), `api-key` is the
|
|
`api-key: <key>` header Azure AI Foundry wants. There is deliberately no
|
|
"send both" option: measured against a real llama-swap server, a request
|
|
carrying both headers hung indefinitely. `baseUrl` must be `http(s)`, carry
|
|
no embedded credentials, and may not point at a link-local or cloud-metadata
|
|
address; discovery re-checks the address the name actually resolves to.
|
|
|
|
Discover its available models:
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/model-endpoints/llama-box/discover-models
|
|
```
|
|
|
|
This calls the endpoint's own `GET /v1/models` and stores the returned list
|
|
on the endpoint record; `GET /api/model-endpoints` lists everything
|
|
configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
|
|
Endpoint management is admin-only in multi-user mode, same as remote/docker
|
|
hosts — these are machine-level infra, not per-user settings.
|
|
|
|
**Context length is discovered too, opportunistically and safely.** The plain
|
|
`GET /v1/models` response has no context-window field, but llama.cpp's
|
|
llama-swap-proxied `GET /props?model=<id>` does (`n_ctx`). Discovery only ever
|
|
calls it for a model llama-swap's own response already reports
|
|
`status.value === "loaded"` for — never for an unloaded one, because
|
|
llama-swap treats `?model=` as a routing hint and asking about a model that
|
|
isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side
|
|
effect of what should be read-only discovery. A server with no `status` field
|
|
on any entry at all (not llama-swap) gets no context-length enrichment,
|
|
rather than guessing. A model's previously-learned context length survives a
|
|
later cycle where it wasn't the loaded one; it's dropped only once the model
|
|
disappears from the endpoint's list entirely. Stored per model in
|
|
`modelContextLengths` and applied automatically (see "Applying a model to a
|
|
session" below) so a CLI that would otherwise assume a large default context
|
|
window for an unrecognized model id stops silently overflowing a much
|
|
smaller real one.
|
|
|
|
**File size is discovered too, when the server states one.** llama-swap
|
|
writes a GB figure into an auto-discovered model's own `description`
|
|
(`"Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"`), parsed
|
|
into `modelSizesGB` — unlike context length, this needs no `/props` probe
|
|
(the figure is right there in the `/v1/models` response) and so is populated
|
|
for every model regardless of loaded state. A hand-configured profile's own
|
|
description has no such figure and correctly gets no entry, never a guess.
|
|
Used only to label the Run-menu picker's "loading model" banner with a
|
|
rough, UNMEASURED expected-time estimate (`_estimateModelLoad()` in
|
|
session-ui.js, based on typical local NVMe/SSD throughput — not benchmarked
|
|
against any real endpoint's actual hardware/storage) and to scale that same
|
|
banner's own give-up timeout for a very large model; never anything a
|
|
server-side check relies on.
|
|
|
|
**The loading banner shows a live countdown against that same timeout, and
|
|
treats a real timeout as a failure, not a shrug.** It checks llama-swap's
|
|
own `/running` every second (`GET /api/model-endpoints/:id/running-status`)
|
|
and counts down against the size-scaled (or flat 5-minute) timeout live; if
|
|
the countdown reaches zero with the target model still not ready, the
|
|
banner turns into a sticky error naming the llama-swap server's own logs as
|
|
where to look, and the session the load was for is closed automatically —
|
|
a console left open and pointed at a model that never finished loading is
|
|
worse than no console at all.
|
|
|
|
`defaultModelId` names which discovered model the picker pre-marks for that
|
|
endpoint — the settings panel's Edit form exposes it as a select populated
|
|
from the endpoint's own discovered `models`, and the route refuses a value
|
|
that isn't one of them. It is applied automatically only when the endpoint
|
|
has exactly one discovered model (nothing to choose); with two or more it
|
|
is a pre-selection in the model-picker dialog below, never a silent default.
|
|
Re-discovering drops a default that no longer appears in the fresh list
|
|
rather than carrying an invalid one forward.
|
|
|
|
**Model lists refresh themselves.** A background sweep (`server.ts`,
|
|
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
|
|
saved endpoint the same way the manual `POST .../discover-models` route
|
|
does, best-effort per endpoint — one being unreachable on a given cycle
|
|
never blocks the others. Off under `npm test`, same reasoning as the Codex
|
|
plan-usage poll it sits beside: no real network to hit, no server instance
|
|
to keep the timer alive for.
|
|
|
|
## The Run-menu picker
|
|
|
|
With the setting on and at least one endpoint carrying a discovered model,
|
|
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
|
|
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
|
|
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
|
|
registry's own `capabilities.customModelInjection` at page render
|
|
(`window.__codemanCustomModelClis`, `server.ts`) — never a hardcoded id list
|
|
in the frontend — so a CLI whose injection recipe lands later shows up with
|
|
no frontend change, and Antigravity (`unsupported`) never does.
|
|
|
|
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
|
|
`session-ui.js`) rather than trusting anything cached from the dropdown's
|
|
own render — the model list can have changed via the 5-minute sweep above
|
|
or a settings-panel edit since the menu opened. With exactly one discovered
|
|
model it runs straight away; with two or more, a small modal
|
|
(`#customModelPickModal`) lists them and asks which one to use for this
|
|
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
|
|
the point of asking is letting one launch deliberately differ from the
|
|
saved default, not just confirming it.
|
|
|
|
**How the launch itself applies the endpoint depends on the harness.** For
|
|
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP (`runCustomModelEntry` →
|
|
`_runCustomModelEntryOneShot`), the endpoint/model is folded into the SAME
|
|
`POST /api/quick-start` call that creates the session (`customModel` field),
|
|
so the session launches directly on the endpoint — no restart, no visible
|
|
relaunch. Claude (`_runCustomModelEntryViaRestart`) still uses the original
|
|
two-step design: the launch runs a single native session exactly the way its
|
|
own Run-menu entry would, then **waits for the new session to go idle**
|
|
(`GET .../wait?until=idle`, bounded at 20s — a normal 200 either way, never
|
|
an error, per the wait endpoint's own contract) before applying the endpoint
|
|
via the restart route below. That wait exists because a freshly launched CLI
|
|
reports itself as `busy` for its own startup (a boot spinner, a
|
|
workspace-trust check) well before the apply call would otherwise reach it,
|
|
and the apply route correctly refuses to restart a session mid-turn — a
|
|
fresh boot looks exactly like one from the outside. A session still busy
|
|
after the wait reaches the apply call anyway and gets that route's own
|
|
honest `SESSION_BUSY` error, now visible as a sticky toast with a close
|
|
button rather than a generic message that vanished in three seconds. Claude
|
|
stays on this path because its own restart (`--resume`-based, keeping the
|
|
conversation) is far less jarring than the other seven's, and `runClaude()`'s
|
|
multi-tab launch and docker-config-drift confirm/retry loop make folding it
|
|
into the one-shot path separate work. It is a
|
|
one-off "try this endpoint" action, not a sticky mode: the plain Run button
|
|
still means "this harness, native cloud" afterward. Entries are hidden
|
|
entirely for a remote or Docker active case, since the apply route refuses
|
|
both (see the next section).
|
|
|
|
## Launching directly on an endpoint (no restart)
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/quick-start \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"caseName": "myapp", "mode": "codex", "customModel": {"endpointId": "llama-box", "modelId": "qwen3"}}'
|
|
```
|
|
|
|
`POST /api/quick-start`'s `customModel` field (`{endpointId, modelId,
|
|
confirmed?}`) computes the same injection the restart route below does, but
|
|
BEFORE the session exists — the session is minted its own id up front
|
|
(`crypto.randomUUID()`), the injection (env vars, and for a `configDir`-kind
|
|
CLI, the written config file) targets that real id, and the session launches
|
|
already pointed at the endpoint. No restart, because there was never a
|
|
native-backend launch to restart away from. Runs the same llama-swap
|
|
conflict check as the restart route (below) — a `409`-shaped
|
|
`{requiresConfirmation, currentlyLoadedModel, affectedSessions}` response
|
|
with no session created, resolved by retrying with `confirmed: true` — and
|
|
is refused the same way for a remote or Docker case. This is what the
|
|
Run-menu picker uses for opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP;
|
|
Claude still uses the restart route below (see "The Run-menu picker" above
|
|
for why).
|
|
|
|
## Applying a model to an ALREADY-RUNNING session
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"endpointId": "llama-box", "modelId": "qwen3"}'
|
|
```
|
|
|
|
This computes the CLI-specific env vars / config for that session's mode
|
|
(see the recipe table in `custom-model-endpoints-plan.md`) and **restarts the session's
|
|
CLI process in place** — same pane, same tmux session, fresh env. That
|
|
restart is necessary, not incidental: every supported harness reads its
|
|
endpoint config at process start, not per-turn, so there is no live
|
|
hot-swap. A Claude session is relaunched with `--resume <conversation> ||
|
|
--session-id <id>`, so it continues the conversation it was on; pi, omp and
|
|
grok are relaunched with the `--model` value that selects the injected
|
|
provider (`custom/<modelId>` for pi and omp, `codeman-custom` for grok),
|
|
since for those three the config file alone does not switch the model.
|
|
**Remote (SSH) and Docker sessions are refused** (400) for now: their restart
|
|
reattaches the durable remote/in-container tmux rather than relaunching the
|
|
agent, so the selection would report success and change nothing.
|
|
|
|
**Claude gets two more env vars when known/applicable, both declared on its
|
|
registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
|
|
|
|
- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context
|
|
length (see the discovery section above) whenever one is known. Without
|
|
it, Claude Code assumes a large (200k) window for any unrecognized custom
|
|
model id and never compacts, which reliably overflows a much smaller real
|
|
local context — confirmed live: a stock ~33.7K-token system prompt against
|
|
a 16384-token llama-swap model failed with `exceeds the available context
|
|
size`. No entry for the model in `modelContextLengths` means the var is
|
|
simply omitted, never a guess.
|
|
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
|
|
the `configDir`-kind CLIs use (empty, no files written into it), so the
|
|
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
|
|
claude.ai OAuth login. Claude Code still prints "Both claude.ai and
|
|
ANTHROPIC_API_KEY set" when the two coexist in the same config directory —
|
|
cosmetic (confirmed live: the API key wins for actual requests either way,
|
|
visible in the terminal's own `API Usage Billing` line) but worth
|
|
eliminating rather than living with. The directory's `projects`
|
|
subdirectory is symlinked (a junction on Windows) back to the real
|
|
`~/.claude/projects` so the response viewer, subagent windows and Read My
|
|
Mind keep working for that session — the same trade-off and fix documented
|
|
for a manually-set `CLAUDE_CONFIG_DIR` in
|
|
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
|
|
automatically here. Best-effort: a platform that refuses the symlink keeps
|
|
the pre-existing blind-response-viewer side effect rather than failing the
|
|
whole custom-model apply over it.
|
|
|
|
**That isolated directory needed one more fix to actually be usable
|
|
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
|
|
real profile's prior "Detected a custom API key — use it?" approvals, so
|
|
without more, Claude Code stops and asks that on *every single launch* —
|
|
confirmed live, and with nobody at a TTY to answer, its own default answer
|
|
("No") silently refuses the very key this feature just injected, which
|
|
looks like the endpoint being ignored entirely. `customModelInjection`'s
|
|
`apiKeyTrustFile` (`{ relPath: '.claude.json', shape:
|
|
'claude-api-key-responses' }` on claude's entry) pre-seeds that exact
|
|
approval: the apply step merges `customApiKeyResponses.approved: [apiKey]`
|
|
into `<configDir>/.claude.json`, the same field a real answered prompt
|
|
itself writes to (confirmed against a real file after answering by hand
|
|
once) — this answers the prompt in advance rather than bypassing it. The
|
|
merge preserves whatever else the CLI already wrote into that file on an
|
|
earlier launch in the same isolated directory (`userID`, `numStartups`,
|
|
earlier approved keys), and a missing or corrupt file is treated as empty
|
|
rather than failing the apply.
|
|
|
|
**llama-swap gets two more fixes on top of the context-length/config-dir
|
|
ones above, both from watching a real switch live.** llama.cpp only ever
|
|
runs one model at a time; llama-swap swaps the backing process on demand,
|
|
which can take anywhere from a few seconds to well over a minute:
|
|
|
|
- **The conflict check.** Both apply routes (the restart one here and the
|
|
one-shot `POST /api/quick-start` above) call llama-swap's own
|
|
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
|
|
compatible server (no such endpoint) is simply never checked. If a
|
|
*different* model is currently loaded and ready, and another **live
|
|
session's own selection** is using it, the apply returns
|
|
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
|
|
instead of silently switching — nothing is applied or created yet.
|
|
Retrying with `confirmed: true` skips the check. Switching with nothing
|
|
else affected proceeds immediately; this is a warning about disrupting
|
|
another session, never a gate on the switch itself.
|
|
- **Actually starting the load.** llama-swap has no "switch model" admin
|
|
call — the only thing that starts a swap is a real inference request
|
|
naming the model, and confirmed live: applying a selection alone never
|
|
reached llama-swap at all (nothing in its own server logs), since nothing
|
|
had actually asked it to load anything yet. Both apply routes now also
|
|
send the smallest real request that will — `POST <baseUrl>/v1/chat/
|
|
completions` with `max_tokens: 1` and one throwaway message — whenever the
|
|
target model isn't already the one loaded and ready, fire-and-forget (its
|
|
response is never read; `GET /api/model-endpoints/:id/running-status`,
|
|
polled client-side, is what actually confirms readiness). The response
|
|
also carries `modelSwapInProgress: true` in that case, which is what
|
|
drives the Run-menu picker's own "loading model" status banner.
|
|
|
|
Clear back to the harness's native cloud default with:
|
|
|
|
```bash
|
|
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
|
-H 'Content-Type: application/json' -d '{"clear": true}'
|
|
```
|
|
|
|
Clearing also removes the env vars the selection injected from the tmux
|
|
session (they persist there and would otherwise be inherited by the
|
|
relaunched CLI) and deletes the per-session config directory
|
|
(`~/.codeman/custom-model-configs/<sessionId>`, written 0600 because pi and
|
|
omp embed the API key in it). That directory is also removed when the
|
|
session is deleted. The selection survives a Codeman restart: the endpoint
|
|
id, model and injected key NAMES are persisted, the values are re-derived
|
|
from the endpoint store on recovery, and the pane keeps running against the
|
|
endpoint in between because tmux retains its environment.
|
|
|
|
**New sessions always default back to the harness's native backend.** A
|
|
custom-endpoint selection is a per-session choice, never a sticky global
|
|
default — starting a fresh session doesn't inherit whatever the last one was
|
|
pointed at.
|
|
|
|
## Confidence per harness
|
|
|
|
Every harness except Antigravity has now been run end-to-end against a real
|
|
llama-swap server via `scripts/test-local-llm-harnesses.ts` (a dynamic
|
|
script that reads the live CLI registry, so a registry change is picked up
|
|
automatically). Results:
|
|
|
|
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
|
|
came back through the endpoint.
|
|
- **Codex** — the config is structurally correct, but Codex only speaks the
|
|
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
|
|
This is a real protocol incompatibility, not a bug here; Codex support
|
|
needs a Responses-API-compatible endpoint.
|
|
- **Gemini** — fails with `Invalid auth method selected`, traced to an
|
|
undocumented `GATEWAY` auth path gemini-cli selects once
|
|
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
|
|
(several auth workarounds were tried and ruled out); do not rely on
|
|
Gemini support yet.
|
|
- **DeepSeek** — the request reaches the server (env vars are read) but
|
|
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
|
|
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
|
|
|
|
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
|
|
each result. `scripts/test-local-llm-harnesses.ts` is the standalone script
|
|
used to check a harness against a real endpoint outside the web UI
|
|
entirely; see its own `--help` for usage.
|
|
|
|
## Security note
|
|
|
|
Every env var this feature can set that redirects a session's traffic
|
|
(`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, etc.) is
|
|
listed in that CLI's `privilegedEnvKeys` in the CLI registry, so a
|
|
non-granted multi-user owner cannot set one directly via the generic
|
|
`envOverrides` API field — only through this feature's own route, which
|
|
computes the value from an admin-configured, SSRF-guarded endpoint rather
|
|
than trusting arbitrary client input. See the "Multi-user security
|
|
hardening" section of `custom-model-endpoints-plan.md` for the full reasoning; several
|
|
of these were reachable via the generic `envOverrides` field even before
|
|
this feature existed, and building this surfaced and closed that gap.
|