Files
Codeman/.changeset/run-menu-custom-model-picker.md
T
DevvynandClaude Sonnet 5 5ddc028a2f feat(custom-model): detect and notify when a session's model gets swapped out later
The llama-swap conflict check on the apply/create routes only ever runs
at THAT session's own launch/apply moment, and cannot see a swap caused
by a DIFFERENT session's later, ordinary use. Confirmed live: a second
Codex session picking a different model launched with no warning at
all — nothing conflicted at that exact instant — yet it silently
evicted the first session's model regardless (llama.cpp runs one model
at a time). Reproduced and root-caused via direct API calls against a
live test-picker instance rather than guessing.

- detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups
  live sessions with a customModel by endpointId, checks each group's
  endpoint via GET /running once, and flags a session whose own modelId
  is no longer in the running list. Read-only, best-effort per endpoint
  like refreshAllCustomModelHosts's sibling sweep.
- Notifies once per displacement via a caller-owned de-dupe Set: a
  session id is added when displaced, removed once its own model is
  loaded/ready again, so a later genuinely-new displacement can notify
  again.
- New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
  20s — much shorter than the 5-minute model-list refresh, since this
  is time-sensitive) broadcasts a new custom-model:swapped-out SSE
  event per displacement. De-dupe Set cleared per-session on session
  cleanup to avoid an unbounded leak.
- Frontend: global toast (not tied to the displaced session's tab,
  since the point is warning before the user types into it) naming the
  session, its previous model, and what's currently loaded.

Chose the "detect after the fact" scope (vs. checking before every
message send, which would add a round-trip to every turn on every
custom-model session) per explicit user decision after being presented
the trade-off.

9 new tests for the detection logic (flag/clear/re-flag cycle,
unreachable/deleted endpoints, non-llama-swap servers, multiple
sessions on one endpoint). SSE registry bumped 158->159, parity test
passing. Typecheck/lint/frontend-syntax clean; full suite shows no new
regressions (9 more passing than baseline, matching the new tests;
same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 11:09:24 +08:00

8.8 KiB

aicodeman
aicodeman
minor

Custom model endpoints: Run-menu picker, and hardening from real llama-swap validation (#430, follow-up to #393's HTTP-API-only cut). With Custom model endpoints on (App Settings → Models) and at least one saved endpoint carrying a discovered model, the Run dropdown grows a Custom Endpoints section generated live off the CLI registry's own capabilities.customModelInjection — one entry per (harness that can redirect to a custom endpoint, saved endpoint). Picking one launches that harness and applies the endpoint to it; with two or more discovered models a small, scrollable dialog asks which one first, the endpoint's defaultModelId marked but never auto-chosen. Endpoints also now re-discover themselves automatically every 5 minutes in the background, one unreachable endpoint never blocking the others.

Everything below was found and fixed against a real llama-swap server, not just unit tests:

  • Session-busy false refusal. A freshly launched CLI reports itself busy for its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying.

  • Errors and confirmations you can actually read. Toasts now default to sticky with a close button (errors always were meant to stay, but a fixed 3s timer silently hid them); a failed apply's real server-side reason (not a generic message) reaches the toast.

  • "Both claude.ai and ANTHROPIC_API_KEY set" warning. A custom-model Claude session now runs with an isolated CLAUDE_CONFIG_DIR (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — projects is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on every launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would.

  • Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length and applies it as CLAUDE_CODE_MAX_CONTEXT_TOKENS — sourced primarily from llama-swap's own GET /running, whose cmd field carries the launch flags (--fit-ctx/-c/--ctx-size) actually in effect, since GET /props's n_ctx was confirmed live to report the model's theoretical/trained maximum rather than the real --fit-ctx-shrunk runtime context (a 154112-vs-16384 discrepancy, caught only because the fixed value still overflowed) — /props is now a fallback for a plain llama.cpp server with no /running at all.

  • Context floor too small for Claude Code to even start. Fixing the overflow above surfaced a second, unfixable-by-injection failure: Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on their own (confirmed live via an in:0 out:0 failure on the very first message), which can exceed a small model's entire real context before any conversation history exists to trim — no CLAUDE_CODE_MAX_CONTEXT_TOKENS value fixes that, since it only governs when history gets compacted. Applying such a model now returns a warning (gated on the CLI registry declaring a contextLengthVar, so it's a no-op for every other harness) instead of launching straight into a guaranteed first-message failure, and the Run-menu picker shows it as an in-app dialog naming the model, its discovered context and the ~40K safe floor, with the actual fix spelled out: give the model an explicit larger -c/--ctx-size in llama-swap's config instead of relying on auto-fit, which optimizes for the biggest model that fits rather than the biggest context. "Launch anyway" is still one click away.

  • The real root cause of "it still says opus, not my model." llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on demand, which can take anywhere from a few seconds to well over a minute — long enough that a session mid-swap is indistinguishable from one that never left the native backend. Applying a selection now checks llama-swap's own GET /running first (feature-detected; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is never checked); if switching would unload a model another live session is actively using, the apply is refused with a warning naming that session instead of silently switching, and a confirmation retry proceeds anyway. Either way, a sticky "loading model…" toast now covers the actual swap window until llama-swap reports the target model ready, so a prompt sent mid-swap reads as "loading," never as silence or an answer from whatever was loaded a moment before.

  • Claude's whole first-run sequence, on every single launch. A fresh, otherwise-empty CLAUDE_CONFIG_DIR isn't just missing the API-key approval above — Claude Code treats it as a brand-new profile and replays the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running bypassed) a one-time permissions-bypass warning, every time, confirmed live. None of that shows up again for a real, already-onboarded profile. customModelInjection's new skipFirstRunPrompts (claude's entry only) pre-seeds that same "already been through this" state — hasCompletedOnboarding and this session's own project trust into the same .claude.json the API-key approval merges into, skipDangerousModePermissionPrompt into settings.json — so a custom-model launch reaches the conversation exactly as fast as a native cloud one, with nobody there to click through a wizard.

Two more, from actually clicking through the swap-confirm and context-warning dialogs live: their z-index sat under the centred status banner, so a dialog could render fully hidden behind "Claude started — switching to llama-swap…"; and their Cancel/confirm buttons stacked instead of sitting side by side (.btn-toolbar's own display: flex needs a row-layout parent it never had). Both dialogs now clear the banner and lay their buttons out centred, side by side.

  • A session's model getting silently swapped out later, not just at launch. The conflict check above only ever runs at the moment a session is created or a model applied — confirmed live: a second Codex session picking a different model launched with no warning at all, because nothing conflicted at that exact instant, yet it silently evicted the first session's model regardless (llama.cpp runs one model at a time). There was no mechanism to catch a swap caused by a DIFFERENT session's own later, ordinary use. A new periodic sweep (detectCustomModelSwapDisplacements, every 20s, one GET /running per distinct endpoint with a live custom-model session) now compares each such session's own model against what's actually loaded, and a new custom-model:swapped-out SSE event drives a global toast naming the displaced session and what's now loaded instead — so you find out before typing into a session that's about to trigger yet another reload. Notifies once per displacement, clearing once a session's own model is loaded and ready again so a later, genuinely new displacement notifies again.

Remote (SSH) and Docker sessions are refused for now (400) — their restart reattaches the durable remote/in-container tmux rather than relaunching the agent.

One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all. Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). POST /api/quick-start now accepts a customModel field and computes the same injection before the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own --resume-based restart is far less jarring, and runClaude()'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work).