feat(custom-model): detect and notify when a session's model gets swapped out later

The llama-swap conflict check on the apply/create routes only ever runs
at THAT session's own launch/apply moment, and cannot see a swap caused
by a DIFFERENT session's later, ordinary use. Confirmed live: a second
Codex session picking a different model launched with no warning at
all — nothing conflicted at that exact instant — yet it silently
evicted the first session's model regardless (llama.cpp runs one model
at a time). Reproduced and root-caused via direct API calls against a
live test-picker instance rather than guessing.

- detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups
  live sessions with a customModel by endpointId, checks each group's
  endpoint via GET /running once, and flags a session whose own modelId
  is no longer in the running list. Read-only, best-effort per endpoint
  like refreshAllCustomModelHosts's sibling sweep.
- Notifies once per displacement via a caller-owned de-dupe Set: a
  session id is added when displaced, removed once its own model is
  loaded/ready again, so a later genuinely-new displacement can notify
  again.
- New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
  20s — much shorter than the 5-minute model-list refresh, since this
  is time-sensitive) broadcasts a new custom-model:swapped-out SSE
  event per displacement. De-dupe Set cleared per-session on session
  cleanup to avoid an unbounded leak.
- Frontend: global toast (not tied to the displaced session's tab,
  since the point is warning before the user types into it) naming the
  session, its previous model, and what's currently loaded.

Chose the "detect after the fact" scope (vs. checking before every
message send, which would add a round-trip to every turn on every
custom-model session) per explicit user decision after being presented
the trade-off.

9 new tests for the detection logic (flag/clear/re-flag cycle,
unreachable/deleted endpoints, non-llama-swap servers, multiple
sessions on one endpoint). SSE registry bumped 158->159, parity test
passing. Typecheck/lint/frontend-syntax clean; full suite shows no new
regressions (9 more passing than baseline, matching the new tests;
same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 11:09:24 +08:00
co-authored by Claude Sonnet 5
parent 470f75b08c
commit 5ddc028a2f
10 changed files with 385 additions and 2 deletions
+22
View File
@@ -335,6 +335,28 @@ completions` with `max_tokens: 1` and one throwaway message — whenever the
also carries `modelSwapInProgress: true` in that case, which is what
drives the Run-menu picker's own "loading model" status banner.
## Catching a swap after the fact
The conflict check above only runs at the moment a session is created or a
model is applied — it has no way to catch a swap that happens **later**.
Confirmed live: a session created while nothing else conflicted at that
exact instant can still get silently displaced afterward, once a
_different_ session's own normal use (or its own create-time load trigger)
asks llama-swap to load something else. llama-swap has no push
notification of its own for this, so a background sweep
(`detectCustomModelSwapDisplacements`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS`
= 20s in `server.ts`) polls `GET /running` once per distinct endpoint that
has at least one live custom-model session, and compares each such
session's own `modelId` against what is actually loaded. A session whose
model is no longer in that list gets a `custom-model:swapped-out` SSE event
(`{sessionId, sessionName, endpointId, previousModel, currentlyLoadedModel}`),
shown as a global toast — global rather than tied to that session's tab,
since the whole point is telling the user before they type into it
expecting the model they picked. Notifies **once per displacement**: the
same de-dupe `Set` clears a session's flag once its own model is loaded and
ready again, so a later, genuinely new displacement notifies again rather
than the session staying silently un-notified forever after the first one.
## Context-window floor warning
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
+10
View File
@@ -89,6 +89,16 @@ error telling you to check the llama-swap server's own logs, and **the session t
for is closed automatically** — a console left open and pointed at a model that never
finished loading would just be confusing to leave sitting there.
**You'll also be told if a session's model gets swapped out from under it later, not just
at launch.** The conflict warning above only fires at the moment you launch or apply a
model — llama.cpp only runs one model at a time, so if a DIFFERENT session using the same
endpoint later triggers its own load, whatever was loaded before (including a session you
already had running) gets silently evicted, with no warning at that instant since nothing
conflicted when it was first set up. A background check (every 20 seconds) catches this
after the fact and shows a toast naming which session lost its model and what's loaded now
— so you know before typing into that session that it's about to reload (and, in turn,
evict whatever displaced it).
**Claude Code specifically gets three extra fixes applied automatically:**
- Its discovered context length (see above) is passed through as