feat(custom-model): detect and notify when a session's model gets swapped out later

The llama-swap conflict check on the apply/create routes only ever runs
at THAT session's own launch/apply moment, and cannot see a swap caused
by a DIFFERENT session's later, ordinary use. Confirmed live: a second
Codex session picking a different model launched with no warning at
all — nothing conflicted at that exact instant — yet it silently
evicted the first session's model regardless (llama.cpp runs one model
at a time). Reproduced and root-caused via direct API calls against a
live test-picker instance rather than guessing.

- detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups
  live sessions with a customModel by endpointId, checks each group's
  endpoint via GET /running once, and flags a session whose own modelId
  is no longer in the running list. Read-only, best-effort per endpoint
  like refreshAllCustomModelHosts's sibling sweep.
- Notifies once per displacement via a caller-owned de-dupe Set: a
  session id is added when displaced, removed once its own model is
  loaded/ready again, so a later genuinely-new displacement can notify
  again.
- New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
  20s — much shorter than the 5-minute model-list refresh, since this
  is time-sensitive) broadcasts a new custom-model:swapped-out SSE
  event per displacement. De-dupe Set cleared per-session on session
  cleanup to avoid an unbounded leak.
- Frontend: global toast (not tied to the displaced session's tab,
  since the point is warning before the user types into it) naming the
  session, its previous model, and what's currently loaded.

Chose the "detect after the fact" scope (vs. checking before every
message send, which would add a round-trip to every turn on every
custom-model session) per explicit user decision after being presented
the trade-off.

9 new tests for the detection logic (flag/clear/re-flag cycle,
unreachable/deleted endpoints, non-llama-swap servers, multiple
sessions on one endpoint). SSE registry bumped 158->159, parity test
passing. Typecheck/lint/frontend-syntax clean; full suite shows no new
regressions (9 more passing than baseline, matching the new tests;
same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 11:09:24 +08:00
co-authored by Claude Sonnet 5
parent 470f75b08c
commit 5ddc028a2f
10 changed files with 385 additions and 2 deletions
+19
View File
@@ -1709,6 +1709,25 @@ class CodemanApp {
console.error('[SSE] docker container recreated:', err);
}
});
// Custom Model Endpoint Profiles: a session's own model got evicted on llama-swap by
// another session's activity, detected AFTER the fact by a periodic server sweep (there
// is no push notification from llama-swap itself) — see detectCustomModelSwapDisplacements
// in custom-model-routes.ts. Global toast rather than a per-tab indicator: the displaced
// session need not be the one currently open, and the whole point is telling the user
// BEFORE they type into it expecting the model they picked.
addListener(SSE_EVENTS.CUSTOM_MODEL_SWAPPED_OUT, (e) => {
try {
const d = e.data ? JSON.parse(e.data) : {};
this.showToast(
`${d.sessionName || d.sessionId}'s model (${d.previousModel}) was swapped out on llama-swap by another ` +
`session — currently loaded: ${d.currentlyLoadedModel}. Sending a message there will reload it.`,
'warning',
{ duration: 0 }
);
} catch (err) {
console.error('[SSE] custom model swapped out:', err);
}
});
// Multi-user admin: live-refresh whichever admin views (panel/Users tab) are open.
addListener(SSE_EVENTS.ADMIN_USERS_CHANGED, () => {
window.codemanAdmin?.onUsersChanged?.();
+3
View File
@@ -1094,6 +1094,9 @@ const SSE_EVENTS = {
APPROVAL_UPDATED: 'approval:updated',
APPROVAL_RESOLVED: 'approval:resolved',
// Custom Model Endpoint Profiles
CUSTOM_MODEL_SWAPPED_OUT: 'custom-model:swapped-out',
// Subagents (Claude Code background agents)
SUBAGENT_DISCOVERED: 'subagent:discovered',
SUBAGENT_UPDATED: 'subagent:updated',
+85
View File
@@ -428,6 +428,91 @@ export async function refreshAllCustomModelHosts(): Promise<void> {
}
}
/** The subset of `Session` this sweep needs — kept minimal so a test can pass a plain object. */
export interface CustomModelSessionLike {
id: string;
name: string;
customModel?: { endpointId: string; modelId: string; label?: string };
}
/** One session whose model was just found evicted, ready to broadcast as `CustomModelSwappedOut`. */
export interface CustomModelSwapDisplacement {
sessionId: string;
sessionName: string;
endpointId: string;
previousModel: string;
currentlyLoadedModel: string;
}
/**
* Detects when a live session's own custom-model selection is no longer the model
* llama-swap actually has loaded — evicted by ANOTHER session's activity on the same
* endpoint, since llama.cpp/llama-swap runs one model at a time (the apply/create routes'
* own swap-conflict check only ever runs at THAT session's own launch/apply moment, so it
* cannot catch a later eviction triggered by a different session's normal use — confirmed
* live: a session created while nothing else had a live conflict at that instant can still
* get silently displaced afterward). Read-only, and best-effort per endpoint exactly like
* `refreshAllCustomModelHosts`'s sibling sweep — one endpoint's hiccup here never blocks
* checking the others.
*
* `notifiedSessionIds` is the caller's own de-dupe state (`server.ts` keeps one `Set` across
* sweeps), mutated in place: a session id is added once displaced and removed again once its
* own model is loaded and ready — so a LATER, genuinely new displacement can notify again
* rather than the session staying silently un-notified forever after the first one.
*/
export async function detectCustomModelSwapDisplacements(
sessions: Iterable<CustomModelSessionLike>,
notifiedSessionIds: Set<string>
): Promise<CustomModelSwapDisplacement[]> {
const byEndpoint = new Map<string, CustomModelSessionLike[]>();
for (const session of sessions) {
if (!session.customModel) continue;
const group = byEndpoint.get(session.customModel.endpointId);
if (group) group.push(session);
else byEndpoint.set(session.customModel.endpointId, [session]);
}
if (byEndpoint.size === 0) return [];
const hosts = await readCustomModelHosts(getDataDir());
const displacements: CustomModelSwapDisplacement[] = [];
for (const [endpointId, group] of byEndpoint) {
const host = hosts.find((h) => h.id === endpointId);
if (!host) continue; // endpoint deleted since these sessions were created — nothing to check
let status: LlamaSwapStatus;
try {
status = await getLlamaSwapStatus(host);
} catch {
continue; // unreachable this cycle — try again next tick, not fatal to the sweep
}
// Not llama-swap (feature-detected) or nothing loaded at all: nothing has been evicted,
// by construction — a plain llama.cpp/OpenAI-compatible server only ever runs the one
// model it was started with, so there is no "current model" to conflict with.
if (!status.isLlamaSwap || status.running.length === 0) continue;
const currentlyLoaded = status.running.find((r) => r.state === 'ready')?.model ?? status.running[0]?.model;
if (!currentlyLoaded) continue;
for (const session of group) {
const modelId = session.customModel!.modelId;
const stillLoaded = status.running.some((r) => r.model === modelId);
if (stillLoaded) {
notifiedSessionIds.delete(session.id); // back to normal — a future eviction can notify again
continue;
}
if (notifiedSessionIds.has(session.id)) continue; // already told them once for this displacement
notifiedSessionIds.add(session.id);
displacements.push({
sessionId: session.id,
sessionName: session.name,
endpointId,
previousModel: modelId,
currentlyLoadedModel: currentlyLoaded,
});
}
}
return displacements;
}
export function registerCustomModelRoutes(app: FastifyInstance): void {
app.get('/api/model-endpoints', async (req): Promise<RedactedHost[]> => {
if (isMultiUserMode() && !isAdmin(req)) return [];
+7 -1
View File
@@ -27,4 +27,10 @@ export { registerWsRoutes } from './ws-routes.js';
export { registerVoiceRoutes } from './voice-routes.js';
export { registerWebviewRoutes, tryWebviewRefererFallback } from './webview-routes.js';
export { registerTabLayoutRoutes } from './tab-layout-routes.js';
export { registerCustomModelRoutes, refreshAllCustomModelHosts } from './custom-model-routes.js';
export {
registerCustomModelRoutes,
refreshAllCustomModelHosts,
detectCustomModelSwapDisplacements,
type CustomModelSessionLike,
type CustomModelSwapDisplacement,
} from './custom-model-routes.js';
+39
View File
@@ -191,6 +191,7 @@ import {
registerTabLayoutRoutes,
registerCustomModelRoutes,
refreshAllCustomModelHosts,
detectCustomModelSwapDisplacements,
tryWebviewRefererFallback,
} from './routes/index.js';
import { isLostWebviewFrameNavigation } from './webview-proxy.js';
@@ -204,6 +205,12 @@ const __dirname = dirname(fileURLToPath(import.meta.url));
const SSE_CLIENT_ID_RE = /^[A-Za-z0-9_-]{8,64}$/;
const CODEX_USAGE_POLL_INTERVAL_MS = 5 * 60_000;
const CUSTOM_MODEL_REDISCOVER_INTERVAL_MS = 5 * 60_000;
// Much shorter than the model-LIST refresh above on purpose: this catches an actual
// eviction (a session's model no longer loaded, silently swapped out by another
// session's use), which the user wants to know about promptly, not once every 5
// minutes. Cheap either way — one /running GET per distinct endpoint with at least
// one live custom-model session, not per session.
const CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS = 20_000;
function escapeHtmlText(value: string): string {
return value.replaceAll('&', '&amp;').replaceAll('<', '&lt;').replaceAll('>', '&gt;');
@@ -281,6 +288,8 @@ export class WebServer extends EventEmitter {
// Store session listener references for explicit cleanup (prevents memory leaks)
private sessionListenerRefs: Map<string, SessionListenerRefs> = new Map();
private scheduledRuns: Map<string, ScheduledRun> = new Map();
/** De-dupe state for the swap-displacement sweep — see detectCustomModelSwapDisplacements. */
private _customModelDisplacedNotified: Set<string> = new Set();
/** Cron service (assigned in setupRoutes). */
private cronService!: CronService;
private sse: SseStreamManager;
@@ -1317,6 +1326,10 @@ export class WebServer extends EventEmitter {
session.ralphTracker.stopWatchingFixPlan();
}
// Custom Model Endpoint Profiles: drop this session's swap-displacement notify flag
// (see _checkCustomModelSwapDisplacements below) so it can't linger in that Set forever.
this._customModelDisplacedNotified.delete(sessionId);
// Kill all subagents spawned by this session (scoped to sessionId to avoid cross-session kills)
if (session && killMux) {
try {
@@ -2755,6 +2768,32 @@ export class WebServer extends EventEmitter {
);
}
// Custom Model Endpoint Profiles: the swap-conflict check on the apply/create routes
// only ever runs at THAT session's own launch/apply moment — it cannot catch a LATER
// eviction triggered by a different session's normal use, since llama-swap has no push
// notification of its own and only swaps in response to a real inference request
// (confirmed live: a session created while nothing else conflicted at that instant can
// still get silently displaced afterward). This periodic sweep is what catches that
// case after the fact and tells the displaced session's user, rather than leaving them
// to discover it only when their next prompt behaves unexpectedly.
if (!this.testMode) {
this.cleanup.setInterval(
() => {
detectCustomModelSwapDisplacements(this.sessions.values(), this._customModelDisplacedNotified)
.then((displacements) => {
for (const displacement of displacements) {
this.broadcast(SseEvent.CustomModelSwappedOut, displacement);
}
})
.catch((err) => {
console.error('[custom-model] swap-displacement check failed:', getErrorMessage(err));
});
},
CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
{ description: 'custom model swap-displacement check' }
);
}
// Start scheduled runs cleanup timer
this.cleanup.setInterval(
() => {
+18 -1
View File
@@ -5,7 +5,7 @@
* and referenced by the frontend (`SSE_EVENTS` in `constants.js`).
* Both files MUST be kept in sync.
*
* 158 event constants organized by category:
* 159 event constants organized by category:
* - **Core** (1): init
* - **Transport** (1): sse:heartbeat
* - **Session lifecycle** (23): created, updated, deleted, terminal, idle, working, ...
@@ -28,6 +28,7 @@
* - **Hooks** (10): idle_prompt, permission_prompt, elicitation_dialog, elicitation_complete, elicitation_response, stop, agent_working, teammate_idle, task_completed, prompt_submitted
* (agent_working is the odd one out: reported by the DeepSeek Harness status bridge, not by a Claude Code hook)
* - **Approvals** (3): pending, updated, resolved (cross-session Approvals Inbox)
* - **Custom Model Endpoint Profiles** (1): swapped-out (a session's model got evicted by another session on the same llama-swap endpoint)
* - **Orchestrator** (12): stateChanged, planProgress, planReady, phase*, verification, task*, completed, error
* - **Clipboard** (1): write
* - **Cases** (4): created, linked, deleted, order-changed
@@ -384,6 +385,19 @@ export const ApprovalUpdated = 'approval:updated' as const;
/** A pending approval left the inbox (answered, superseded, expired, ...). */
export const ApprovalResolved = 'approval:resolved' as const;
// ─── Custom Model Endpoint Profiles ──────────────────────────────────────────
/**
* A session's own custom-model selection is no longer the model llama-swap has loaded —
* ANOTHER session's activity on the same endpoint evicted it (llama.cpp/llama-swap runs
* one model at a time). Detected after the fact by a periodic sweep (`server.ts`), never
* at the moment of eviction itself, since llama-swap has no push notification of its own;
* this session's next prompt will trigger reloading its model, evicting whatever displaced
* it in turn. Fires at most once per displacement (cleared once the sweep sees the
* session's own model loaded again), so it can't spam on every sweep interval.
*/
export const CustomModelSwappedOut = 'custom-model:swapped-out' as const;
// ─── Orchestrator ────────────────────────────────────────────────────────────
/** Orchestrator state machine transitioned. */
@@ -638,6 +652,9 @@ export const SseEvent = {
ApprovalUpdated,
ApprovalResolved,
// Custom Model Endpoint Profiles
CustomModelSwappedOut,
// Orchestrator
OrchestratorStateChanged,
OrchestratorPlanProgress,