feat(custom-model): detect and notify when a session's model gets swapped out later

The llama-swap conflict check on the apply/create routes only ever runs
at THAT session's own launch/apply moment, and cannot see a swap caused
by a DIFFERENT session's later, ordinary use. Confirmed live: a second
Codex session picking a different model launched with no warning at
all — nothing conflicted at that exact instant — yet it silently
evicted the first session's model regardless (llama.cpp runs one model
at a time). Reproduced and root-caused via direct API calls against a
live test-picker instance rather than guessing.

- detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups
  live sessions with a customModel by endpointId, checks each group's
  endpoint via GET /running once, and flags a session whose own modelId
  is no longer in the running list. Read-only, best-effort per endpoint
  like refreshAllCustomModelHosts's sibling sweep.
- Notifies once per displacement via a caller-owned de-dupe Set: a
  session id is added when displaced, removed once its own model is
  loaded/ready again, so a later genuinely-new displacement can notify
  again.
- New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
  20s — much shorter than the 5-minute model-list refresh, since this
  is time-sensitive) broadcasts a new custom-model:swapped-out SSE
  event per displacement. De-dupe Set cleared per-session on session
  cleanup to avoid an unbounded leak.
- Frontend: global toast (not tied to the displaced session's tab,
  since the point is warning before the user types into it) naming the
  session, its previous model, and what's currently loaded.

Chose the "detect after the fact" scope (vs. checking before every
message send, which would add a round-trip to every turn on every
custom-model session) per explicit user decision after being presented
the trade-off.

9 new tests for the detection logic (flag/clear/re-flag cycle,
unreachable/deleted endpoints, non-llama-swap servers, multiple
sessions on one endpoint). SSE registry bumped 158->159, parity test
passing. Typecheck/lint/frontend-syntax clean; full suite shows no new
regressions (9 more passing than baseline, matching the new tests;
same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 11:09:24 +08:00
co-authored by Claude Sonnet 5
parent 470f75b08c
commit 5ddc028a2f
10 changed files with 385 additions and 2 deletions
+18 -1
View File
@@ -5,7 +5,7 @@
* and referenced by the frontend (`SSE_EVENTS` in `constants.js`).
* Both files MUST be kept in sync.
*
* 158 event constants organized by category:
* 159 event constants organized by category:
* - **Core** (1): init
* - **Transport** (1): sse:heartbeat
* - **Session lifecycle** (23): created, updated, deleted, terminal, idle, working, ...
@@ -28,6 +28,7 @@
* - **Hooks** (10): idle_prompt, permission_prompt, elicitation_dialog, elicitation_complete, elicitation_response, stop, agent_working, teammate_idle, task_completed, prompt_submitted
* (agent_working is the odd one out: reported by the DeepSeek Harness status bridge, not by a Claude Code hook)
* - **Approvals** (3): pending, updated, resolved (cross-session Approvals Inbox)
* - **Custom Model Endpoint Profiles** (1): swapped-out (a session's model got evicted by another session on the same llama-swap endpoint)
* - **Orchestrator** (12): stateChanged, planProgress, planReady, phase*, verification, task*, completed, error
* - **Clipboard** (1): write
* - **Cases** (4): created, linked, deleted, order-changed
@@ -384,6 +385,19 @@ export const ApprovalUpdated = 'approval:updated' as const;
/** A pending approval left the inbox (answered, superseded, expired, ...). */
export const ApprovalResolved = 'approval:resolved' as const;
// ─── Custom Model Endpoint Profiles ──────────────────────────────────────────
/**
* A session's own custom-model selection is no longer the model llama-swap has loaded —
* ANOTHER session's activity on the same endpoint evicted it (llama.cpp/llama-swap runs
* one model at a time). Detected after the fact by a periodic sweep (`server.ts`), never
* at the moment of eviction itself, since llama-swap has no push notification of its own;
* this session's next prompt will trigger reloading its model, evicting whatever displaced
* it in turn. Fires at most once per displacement (cleared once the sweep sees the
* session's own model loaded again), so it can't spam on every sweep interval.
*/
export const CustomModelSwappedOut = 'custom-model:swapped-out' as const;
// ─── Orchestrator ────────────────────────────────────────────────────────────
/** Orchestrator state machine transitioned. */
@@ -638,6 +652,9 @@ export const SseEvent = {
ApprovalUpdated,
ApprovalResolved,
// Custom Model Endpoint Profiles
CustomModelSwappedOut,
// Orchestrator
OrchestratorStateChanged,
OrchestratorPlanProgress,