docs(custom-model): make the docs match the code, and trim the changeset

More from the review of 5fc391a4, all documentation rather than behaviour.

The changeset was 1602 words of development log, written as the PR grew, with bullets
and loose paragraphs interleaved. That text becomes CHANGELOG.md and the GitHub release
body verbatim, so it is now one user-facing account of what the feature does and what
the real-server work bought, at roughly a fifth the length.

docs/api-reference.md promised a `cmd` field on running-status that the route
deliberately strips (it carries model paths and can carry --api-key).

Two places claimed the apply routes validate `modelId` against the endpoint's
discovered models. Neither does. Dropped the claim rather than adding the check:
discovery can be up to five minutes stale, so a 400 there would refuse a launch that
actually works, and a typo'd id already fails on the CLI's own first request. CLAUDE.md
now says so explicitly, since the absence is the surprising part.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-09-19 12:34:24 +02:00
parent fe3bd0074c
commit 9af12afb57
4 changed files with 31 additions and 24 deletions
+27 -20
View File
@@ -2,27 +2,34 @@
'aicodeman': minor
---
**Custom model endpoints: Run-menu picker, and hardening from real llama-swap validation** (#430, follow-up to #393's HTTP-API-only cut). With **Custom model endpoints** on (App Settings → Models) and at least one saved endpoint carrying a discovered model, the Run dropdown grows a **Custom Endpoints** section generated live off the CLI registry's own `capabilities.customModelInjection` — one entry per (harness that can redirect to a custom endpoint, saved endpoint). Picking one launches that harness and applies the endpoint to it; with two or more discovered models a small, scrollable dialog asks which one first, the endpoint's `defaultModelId` marked but never auto-chosen. Endpoints also now re-discover themselves automatically every 5 minutes in the background, one unreachable endpoint never blocking the others.
feat(custom-model): pick a custom endpoint straight from the Run menu
Everything below was found and fixed against a **real llama-swap server**, not just unit tests:
#393 landed the backend for custom model endpoints and left it reachable only over the
HTTP API. This is the rest of it. Turn on Custom model endpoints in App Settings, save
an endpoint, and the Run dropdown grows a Custom Endpoints section built live off the
CLI registry, one entry per harness that can actually redirect plus each endpoint you
saved. Pick one and it launches that harness pointed at your server, asking which model
first when the endpoint has more than one. Endpoints re-discover themselves every five
minutes, and one unreachable endpoint never blocks the others. App Settings gains full
add, edit and delete for endpoints.
- **Session-busy false refusal.** A freshly launched CLI reports itself `busy` for its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying.
- **Errors and confirmations you can actually read.** A failed apply's real server-side reason (not a generic message) reaches the toast, and that specific message stays on screen with a close button instead of vanishing on the usual 3s timer.
- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on _every_ launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would.
- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS` — sourced primarily from llama-swap's own `GET /running`, whose `cmd` field carries the launch flags (`--fit-ctx`/`-c`/`--ctx-size`) actually in effect, since `GET /props`'s `n_ctx` was confirmed live to report the model's theoretical/trained maximum rather than the real `--fit-ctx`-shrunk runtime context (a 154112-vs-16384 discrepancy, caught only because the fixed value still overflowed) — `/props` is now a fallback for a plain llama.cpp server with no `/running` at all.
- **Context floor too small for Claude Code to even start.** Fixing the overflow above surfaced a second, unfixable-by-injection failure: Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on their own (confirmed live via an `in:0 out:0` failure on the very first message), which can exceed a small model's entire real context before any conversation history exists to trim — no `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since it only governs when history gets compacted. Applying such a model now returns a warning (gated on the CLI registry declaring a `contextLengthVar`, so it's a no-op for every other harness) instead of launching straight into a guaranteed first-message failure, and the Run-menu picker shows it as an in-app dialog naming the model, its discovered context and the ~40K safe floor, with the actual fix spelled out: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on auto-fit, which optimizes for the biggest model that fits rather than the biggest context. "Launch anyway" is still one click away.
- **The real root cause of "it still says opus, not my model."** llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on demand, which can take anywhere from a few seconds to well over a minute — long enough that a session mid-swap is indistinguishable from one that never left the native backend. Applying a selection now checks llama-swap's own `GET /running` first (feature-detected; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is never checked); if switching would unload a model **another live session is actively using**, the apply is refused with a warning naming that session instead of silently switching, and a confirmation retry proceeds anyway. Either way, a sticky "loading model…" toast now covers the actual swap window until llama-swap reports the target model ready, so a prompt sent mid-swap reads as "loading," never as silence or an answer from whatever was loaded a moment before.
Seven of the harnesses (opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP) now launch
directly onto the endpoint with no restart at all, where before you watched a native
boot followed immediately by a second one. Claude still launches and then restarts in
place, which its own resume makes far less jarring.
- **Claude's whole first-run sequence, on every single launch.** A fresh, otherwise-empty `CLAUDE_CONFIG_DIR` isn't just missing the API-key approval above — Claude Code treats it as a brand-new profile and replays the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running bypassed) a one-time permissions-bypass warning, every time, confirmed live. None of that shows up again for a real, already-onboarded profile. `customModelInjection`'s new `skipFirstRunPrompts` (claude's entry only) pre-seeds that same "already been through this" state — `hasCompletedOnboarding` and this session's own project trust into the same `.claude.json` the API-key approval merges into, `skipDangerousModePermissionPrompt` into `settings.json` — so a custom-model launch reaches the conversation exactly as fast as a native cloud one, with nobody there to click through a wizard.
Most of this release's work went into things that only show up against a real server,
and each was found that way rather than in tests: a freshly launched CLI reporting
itself busy for its own startup and getting refused; Claude Code assuming a large
context window for a model it does not recognise and silently overflowing a small one;
a model whose real context is below what Claude Code's own system prompt costs, which
no setting can fix and which now warns before launching into a certain failure; and the
big one, llama.cpp running exactly one model at a time, so applying a selection can
unload the model another session is using. That last case now asks first, tells you
which session it affects, and keeps a "loading model" notice on screen for the whole
swap window, so a prompt sent mid-swap reads as loading rather than as an answer from
whatever was loaded a moment ago. A background sweep also catches the reverse: your
session's model being evicted later by somebody else's ordinary use.
Two more, from actually clicking through the swap-confirm and context-warning dialogs live: their z-index sat under the centred status banner, so a dialog could render fully hidden behind "Claude started — switching to llama-swap…"; and their Cancel/confirm buttons stacked instead of sitting side by side (`.btn-toolbar`'s own `display: flex` needs a row-layout parent it never had). Both dialogs now clear the banner and lay their buttons out centred, side by side.
- **A session's model getting silently swapped out later, not just at launch.** The conflict check above only ever runs at the moment a session is created or a model applied — confirmed live: a second Codex session picking a different model launched with no warning at all, because nothing conflicted at that exact instant, yet it silently evicted the first session's model regardless (llama.cpp runs one model at a time). There was no mechanism to catch a swap caused by a DIFFERENT session's own later, ordinary use. A new periodic sweep (`detectCustomModelSwapDisplacements`, every 20s, one `GET /running` per distinct endpoint with a live custom-model session) now compares each such session's own model against what's actually loaded, and a new `custom-model:swapped-out` SSE event drives a global toast naming the displaced session and what's now loaded instead — so you find out before typing into a session that's about to trigger yet another reload. Notifies once per displacement, clearing once a session's own model is loaded and ready again so a later, genuinely new displacement notifies again.
- **The loading banner's second line is now the real backend log line, not just a countdown.** llama-swap's `GET /api/events` SSE stream carries the actual `llama-server` process's own stdout (`load_model: loading model '<path>'`, `llama_server: model loaded`, tokenizer warnings, all of it) tagged `source: "upstream"`, distinct from llama-swap's own `source: "proxy"` request-access lines — confirmed live end-to-end through a real forced swap, and it correctly stays on the last thing llama.cpp said once the load goes quiet rather than clearing to blank. ⚠️ This feature's own first cut targeted `GET /logs` instead (the name that suggested it) and shipped a live-tested implementation against it before this live check caught that `/logs` carries ONLY the proxy request log and never once showed a single backend line, even seconds after a real, confirmed swap — corrected before merge, not after.
Remote (SSH) and Docker sessions are refused for now (400) — their restart reattaches the durable remote/in-container tmux rather than relaunching the agent.
- **The loading banner's countdown is gone, replaced by a generic disclaimer and a Cancel button.** Its size-scaled expected-time estimate and matching auto-timeout were both a guess dressed up as a fact — real load time depends on hardware this feature has no way to know, and a fixed number could kill a genuinely slow load partway through. The banner now says "this can take a while depending on your hardware and the model size", polls indefinitely, and carries a **Cancel** button that ends the wait and closes the session on the user's own call rather than a guessed deadline.
**One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all.** Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). `POST /api/quick-start` now accepts a `customModel` field and computes the same injection _before_ the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own `--resume`-based restart is far less jarring, and `runClaude()`'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work).
Remote SSH and Docker sessions are refused for now, since their restart reattaches a
durable tmux rather than relaunching the agent.
+1 -1
View File
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -625,7 +625,7 @@ authStyle?, defaultModelId? }` creates one. `id` must match
`server.ts`), so there is no route for triggering "refresh all" — one
endpoint being unreachable on a cycle never blocks the others.
- `GET /api/v1/model-endpoints/:id/running-status` -> `{ isLlamaSwap,
running: [{model, state, cmd?}], logLine? }`, read-only, no admin gate
running: [{model, state}], logLine? }`, read-only, no admin gate
(any session owner who could already point a session at this endpoint can
equally ask what it currently has loaded). `isLlamaSwap` is
feature-detected via the endpoint's own `GET /running` — a plain
+2 -2
View File
@@ -920,8 +920,8 @@ Object.assign(CodemanApp.prototype, {
// _apiJson() (used everywhere else in this file) unwraps a success body to
// its `data`, but on failure it swallows the response entirely and returns
// null — exactly the `error` text a caller needs to tell "the endpoint is
// unreachable" apart from "the CLI can't be redirected", "not one of the
// discovered models", or "this is a Docker/remote session". Go through the
// unreachable" apart from "the CLI can't be redirected" or "this is a
// Docker/remote session". Go through the
// raw response here instead so a failure is diagnosable, not just present.
let { ok, data, res } = await this._applyCustomModelToSession(sessionId, endpointId, modelId);