diff --git a/.changeset/run-menu-custom-model-picker.md b/.changeset/run-menu-custom-model-picker.md index f3962a7e..44613361 100644 --- a/.changeset/run-menu-custom-model-picker.md +++ b/.changeset/run-menu-custom-model-picker.md @@ -17,6 +17,8 @@ Everything below was found and fixed against a **real llama-swap server**, not j Two more, from actually clicking through the swap-confirm and context-warning dialogs live: their z-index sat under the centred status banner, so a dialog could render fully hidden behind "Claude started — switching to llama-swap…"; and their Cancel/confirm buttons stacked instead of sitting side by side (`.btn-toolbar`'s own `display: flex` needs a row-layout parent it never had). Both dialogs now clear the banner and lay their buttons out centred, side by side. +- **A session's model getting silently swapped out later, not just at launch.** The conflict check above only ever runs at the moment a session is created or a model applied — confirmed live: a second Codex session picking a different model launched with no warning at all, because nothing conflicted at that exact instant, yet it silently evicted the first session's model regardless (llama.cpp runs one model at a time). There was no mechanism to catch a swap caused by a DIFFERENT session's own later, ordinary use. A new periodic sweep (`detectCustomModelSwapDisplacements`, every 20s, one `GET /running` per distinct endpoint with a live custom-model session) now compares each such session's own model against what's actually loaded, and a new `custom-model:swapped-out` SSE event drives a global toast naming the displaced session and what's now loaded instead — so you find out before typing into a session that's about to trigger yet another reload. Notifies once per displacement, clearing once a session's own model is loaded and ready again so a later, genuinely new displacement notifies again. + Remote (SSH) and Docker sessions are refused for now (400) — their restart reattaches the durable remote/in-container tmux rather than relaunching the agent. **One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all.** Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). `POST /api/quick-start` now accepts a `customModel` field and computes the same injection _before_ the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own `--resume`-based restart is far less jarring, and `runClaude()`'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work). diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index f7aeb7e7..5c2fb10d 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -335,6 +335,28 @@ completions` with `max_tokens: 1` and one throwaway message — whenever the also carries `modelSwapInProgress: true` in that case, which is what drives the Run-menu picker's own "loading model" status banner. +## Catching a swap after the fact + +The conflict check above only runs at the moment a session is created or a +model is applied — it has no way to catch a swap that happens **later**. +Confirmed live: a session created while nothing else conflicted at that +exact instant can still get silently displaced afterward, once a +_different_ session's own normal use (or its own create-time load trigger) +asks llama-swap to load something else. llama-swap has no push +notification of its own for this, so a background sweep +(`detectCustomModelSwapDisplacements`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS` += 20s in `server.ts`) polls `GET /running` once per distinct endpoint that +has at least one live custom-model session, and compares each such +session's own `modelId` against what is actually loaded. A session whose +model is no longer in that list gets a `custom-model:swapped-out` SSE event +(`{sessionId, sessionName, endpointId, previousModel, currentlyLoadedModel}`), +shown as a global toast — global rather than tied to that session's tab, +since the whole point is telling the user before they type into it +expecting the model they picked. Notifies **once per displacement**: the +same de-dupe `Set` clears a session's flag once its own model is loaded and +ready again, so a later, genuinely new displacement notifies again rather +than the session staying silently un-notified forever after the first one. + ## Context-window floor warning Claude Code's own fixed per-turn overhead (system prompt + tool schemas, diff --git a/docs/wiki/Custom-Model-Endpoints.md b/docs/wiki/Custom-Model-Endpoints.md index 47727119..20eaa4c3 100644 --- a/docs/wiki/Custom-Model-Endpoints.md +++ b/docs/wiki/Custom-Model-Endpoints.md @@ -89,6 +89,16 @@ error telling you to check the llama-swap server's own logs, and **the session t for is closed automatically** — a console left open and pointed at a model that never finished loading would just be confusing to leave sitting there. +**You'll also be told if a session's model gets swapped out from under it later, not just +at launch.** The conflict warning above only fires at the moment you launch or apply a +model — llama.cpp only runs one model at a time, so if a DIFFERENT session using the same +endpoint later triggers its own load, whatever was loaded before (including a session you +already had running) gets silently evicted, with no warning at that instant since nothing +conflicted when it was first set up. A background check (every 20 seconds) catches this +after the fact and shows a toast naming which session lost its model and what's loaded now +— so you know before typing into that session that it's about to reload (and, in turn, +evict whatever displaced it). + **Claude Code specifically gets three extra fixes applied automatically:** - Its discovered context length (see above) is passed through as diff --git a/src/web/public/app.js b/src/web/public/app.js index 4c688899..44d24389 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -1709,6 +1709,25 @@ class CodemanApp { console.error('[SSE] docker container recreated:', err); } }); + // Custom Model Endpoint Profiles: a session's own model got evicted on llama-swap by + // another session's activity, detected AFTER the fact by a periodic server sweep (there + // is no push notification from llama-swap itself) — see detectCustomModelSwapDisplacements + // in custom-model-routes.ts. Global toast rather than a per-tab indicator: the displaced + // session need not be the one currently open, and the whole point is telling the user + // BEFORE they type into it expecting the model they picked. + addListener(SSE_EVENTS.CUSTOM_MODEL_SWAPPED_OUT, (e) => { + try { + const d = e.data ? JSON.parse(e.data) : {}; + this.showToast( + `${d.sessionName || d.sessionId}'s model (${d.previousModel}) was swapped out on llama-swap by another ` + + `session — currently loaded: ${d.currentlyLoadedModel}. Sending a message there will reload it.`, + 'warning', + { duration: 0 } + ); + } catch (err) { + console.error('[SSE] custom model swapped out:', err); + } + }); // Multi-user admin: live-refresh whichever admin views (panel/Users tab) are open. addListener(SSE_EVENTS.ADMIN_USERS_CHANGED, () => { window.codemanAdmin?.onUsersChanged?.(); diff --git a/src/web/public/constants.js b/src/web/public/constants.js index be9780b0..f661f13a 100644 --- a/src/web/public/constants.js +++ b/src/web/public/constants.js @@ -1094,6 +1094,9 @@ const SSE_EVENTS = { APPROVAL_UPDATED: 'approval:updated', APPROVAL_RESOLVED: 'approval:resolved', + // Custom Model Endpoint Profiles + CUSTOM_MODEL_SWAPPED_OUT: 'custom-model:swapped-out', + // Subagents (Claude Code background agents) SUBAGENT_DISCOVERED: 'subagent:discovered', SUBAGENT_UPDATED: 'subagent:updated', diff --git a/src/web/routes/custom-model-routes.ts b/src/web/routes/custom-model-routes.ts index 8258a8ef..bce9e921 100644 --- a/src/web/routes/custom-model-routes.ts +++ b/src/web/routes/custom-model-routes.ts @@ -428,6 +428,91 @@ export async function refreshAllCustomModelHosts(): Promise { } } +/** The subset of `Session` this sweep needs — kept minimal so a test can pass a plain object. */ +export interface CustomModelSessionLike { + id: string; + name: string; + customModel?: { endpointId: string; modelId: string; label?: string }; +} + +/** One session whose model was just found evicted, ready to broadcast as `CustomModelSwappedOut`. */ +export interface CustomModelSwapDisplacement { + sessionId: string; + sessionName: string; + endpointId: string; + previousModel: string; + currentlyLoadedModel: string; +} + +/** + * Detects when a live session's own custom-model selection is no longer the model + * llama-swap actually has loaded — evicted by ANOTHER session's activity on the same + * endpoint, since llama.cpp/llama-swap runs one model at a time (the apply/create routes' + * own swap-conflict check only ever runs at THAT session's own launch/apply moment, so it + * cannot catch a later eviction triggered by a different session's normal use — confirmed + * live: a session created while nothing else had a live conflict at that instant can still + * get silently displaced afterward). Read-only, and best-effort per endpoint exactly like + * `refreshAllCustomModelHosts`'s sibling sweep — one endpoint's hiccup here never blocks + * checking the others. + * + * `notifiedSessionIds` is the caller's own de-dupe state (`server.ts` keeps one `Set` across + * sweeps), mutated in place: a session id is added once displaced and removed again once its + * own model is loaded and ready — so a LATER, genuinely new displacement can notify again + * rather than the session staying silently un-notified forever after the first one. + */ +export async function detectCustomModelSwapDisplacements( + sessions: Iterable, + notifiedSessionIds: Set +): Promise { + const byEndpoint = new Map(); + for (const session of sessions) { + if (!session.customModel) continue; + const group = byEndpoint.get(session.customModel.endpointId); + if (group) group.push(session); + else byEndpoint.set(session.customModel.endpointId, [session]); + } + if (byEndpoint.size === 0) return []; + + const hosts = await readCustomModelHosts(getDataDir()); + const displacements: CustomModelSwapDisplacement[] = []; + + for (const [endpointId, group] of byEndpoint) { + const host = hosts.find((h) => h.id === endpointId); + if (!host) continue; // endpoint deleted since these sessions were created — nothing to check + let status: LlamaSwapStatus; + try { + status = await getLlamaSwapStatus(host); + } catch { + continue; // unreachable this cycle — try again next tick, not fatal to the sweep + } + // Not llama-swap (feature-detected) or nothing loaded at all: nothing has been evicted, + // by construction — a plain llama.cpp/OpenAI-compatible server only ever runs the one + // model it was started with, so there is no "current model" to conflict with. + if (!status.isLlamaSwap || status.running.length === 0) continue; + const currentlyLoaded = status.running.find((r) => r.state === 'ready')?.model ?? status.running[0]?.model; + if (!currentlyLoaded) continue; + + for (const session of group) { + const modelId = session.customModel!.modelId; + const stillLoaded = status.running.some((r) => r.model === modelId); + if (stillLoaded) { + notifiedSessionIds.delete(session.id); // back to normal — a future eviction can notify again + continue; + } + if (notifiedSessionIds.has(session.id)) continue; // already told them once for this displacement + notifiedSessionIds.add(session.id); + displacements.push({ + sessionId: session.id, + sessionName: session.name, + endpointId, + previousModel: modelId, + currentlyLoadedModel: currentlyLoaded, + }); + } + } + return displacements; +} + export function registerCustomModelRoutes(app: FastifyInstance): void { app.get('/api/model-endpoints', async (req): Promise => { if (isMultiUserMode() && !isAdmin(req)) return []; diff --git a/src/web/routes/index.ts b/src/web/routes/index.ts index df615f30..cd9dee34 100644 --- a/src/web/routes/index.ts +++ b/src/web/routes/index.ts @@ -27,4 +27,10 @@ export { registerWsRoutes } from './ws-routes.js'; export { registerVoiceRoutes } from './voice-routes.js'; export { registerWebviewRoutes, tryWebviewRefererFallback } from './webview-routes.js'; export { registerTabLayoutRoutes } from './tab-layout-routes.js'; -export { registerCustomModelRoutes, refreshAllCustomModelHosts } from './custom-model-routes.js'; +export { + registerCustomModelRoutes, + refreshAllCustomModelHosts, + detectCustomModelSwapDisplacements, + type CustomModelSessionLike, + type CustomModelSwapDisplacement, +} from './custom-model-routes.js'; diff --git a/src/web/server.ts b/src/web/server.ts index ad3b07c9..040ed677 100644 --- a/src/web/server.ts +++ b/src/web/server.ts @@ -191,6 +191,7 @@ import { registerTabLayoutRoutes, registerCustomModelRoutes, refreshAllCustomModelHosts, + detectCustomModelSwapDisplacements, tryWebviewRefererFallback, } from './routes/index.js'; import { isLostWebviewFrameNavigation } from './webview-proxy.js'; @@ -204,6 +205,12 @@ const __dirname = dirname(fileURLToPath(import.meta.url)); const SSE_CLIENT_ID_RE = /^[A-Za-z0-9_-]{8,64}$/; const CODEX_USAGE_POLL_INTERVAL_MS = 5 * 60_000; const CUSTOM_MODEL_REDISCOVER_INTERVAL_MS = 5 * 60_000; +// Much shorter than the model-LIST refresh above on purpose: this catches an actual +// eviction (a session's model no longer loaded, silently swapped out by another +// session's use), which the user wants to know about promptly, not once every 5 +// minutes. Cheap either way — one /running GET per distinct endpoint with at least +// one live custom-model session, not per session. +const CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS = 20_000; function escapeHtmlText(value: string): string { return value.replaceAll('&', '&').replaceAll('<', '<').replaceAll('>', '>'); @@ -281,6 +288,8 @@ export class WebServer extends EventEmitter { // Store session listener references for explicit cleanup (prevents memory leaks) private sessionListenerRefs: Map = new Map(); private scheduledRuns: Map = new Map(); + /** De-dupe state for the swap-displacement sweep — see detectCustomModelSwapDisplacements. */ + private _customModelDisplacedNotified: Set = new Set(); /** Cron service (assigned in setupRoutes). */ private cronService!: CronService; private sse: SseStreamManager; @@ -1317,6 +1326,10 @@ export class WebServer extends EventEmitter { session.ralphTracker.stopWatchingFixPlan(); } + // Custom Model Endpoint Profiles: drop this session's swap-displacement notify flag + // (see _checkCustomModelSwapDisplacements below) so it can't linger in that Set forever. + this._customModelDisplacedNotified.delete(sessionId); + // Kill all subagents spawned by this session (scoped to sessionId to avoid cross-session kills) if (session && killMux) { try { @@ -2755,6 +2768,32 @@ export class WebServer extends EventEmitter { ); } + // Custom Model Endpoint Profiles: the swap-conflict check on the apply/create routes + // only ever runs at THAT session's own launch/apply moment — it cannot catch a LATER + // eviction triggered by a different session's normal use, since llama-swap has no push + // notification of its own and only swaps in response to a real inference request + // (confirmed live: a session created while nothing else conflicted at that instant can + // still get silently displaced afterward). This periodic sweep is what catches that + // case after the fact and tells the displaced session's user, rather than leaving them + // to discover it only when their next prompt behaves unexpectedly. + if (!this.testMode) { + this.cleanup.setInterval( + () => { + detectCustomModelSwapDisplacements(this.sessions.values(), this._customModelDisplacedNotified) + .then((displacements) => { + for (const displacement of displacements) { + this.broadcast(SseEvent.CustomModelSwappedOut, displacement); + } + }) + .catch((err) => { + console.error('[custom-model] swap-displacement check failed:', getErrorMessage(err)); + }); + }, + CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS, + { description: 'custom model swap-displacement check' } + ); + } + // Start scheduled runs cleanup timer this.cleanup.setInterval( () => { diff --git a/src/web/sse-events.ts b/src/web/sse-events.ts index b89c3425..276cdd0e 100644 --- a/src/web/sse-events.ts +++ b/src/web/sse-events.ts @@ -5,7 +5,7 @@ * and referenced by the frontend (`SSE_EVENTS` in `constants.js`). * Both files MUST be kept in sync. * - * 158 event constants organized by category: + * 159 event constants organized by category: * - **Core** (1): init * - **Transport** (1): sse:heartbeat * - **Session lifecycle** (23): created, updated, deleted, terminal, idle, working, ... @@ -28,6 +28,7 @@ * - **Hooks** (10): idle_prompt, permission_prompt, elicitation_dialog, elicitation_complete, elicitation_response, stop, agent_working, teammate_idle, task_completed, prompt_submitted * (agent_working is the odd one out: reported by the DeepSeek Harness status bridge, not by a Claude Code hook) * - **Approvals** (3): pending, updated, resolved (cross-session Approvals Inbox) + * - **Custom Model Endpoint Profiles** (1): swapped-out (a session's model got evicted by another session on the same llama-swap endpoint) * - **Orchestrator** (12): stateChanged, planProgress, planReady, phase*, verification, task*, completed, error * - **Clipboard** (1): write * - **Cases** (4): created, linked, deleted, order-changed @@ -384,6 +385,19 @@ export const ApprovalUpdated = 'approval:updated' as const; /** A pending approval left the inbox (answered, superseded, expired, ...). */ export const ApprovalResolved = 'approval:resolved' as const; +// ─── Custom Model Endpoint Profiles ────────────────────────────────────────── + +/** + * A session's own custom-model selection is no longer the model llama-swap has loaded — + * ANOTHER session's activity on the same endpoint evicted it (llama.cpp/llama-swap runs + * one model at a time). Detected after the fact by a periodic sweep (`server.ts`), never + * at the moment of eviction itself, since llama-swap has no push notification of its own; + * this session's next prompt will trigger reloading its model, evicting whatever displaced + * it in turn. Fires at most once per displacement (cleared once the sweep sees the + * session's own model loaded again), so it can't spam on every sweep interval. + */ +export const CustomModelSwappedOut = 'custom-model:swapped-out' as const; + // ─── Orchestrator ──────────────────────────────────────────────────────────── /** Orchestrator state machine transitioned. */ @@ -638,6 +652,9 @@ export const SseEvent = { ApprovalUpdated, ApprovalResolved, + // Custom Model Endpoint Profiles + CustomModelSwappedOut, + // Orchestrator OrchestratorStateChanged, OrchestratorPlanProgress, diff --git a/test/custom-model-swap-displacement.test.ts b/test/custom-model-swap-displacement.test.ts new file mode 100644 index 00000000..a8034c44 --- /dev/null +++ b/test/custom-model-swap-displacement.test.ts @@ -0,0 +1,180 @@ +/** + * @fileoverview Tests for `detectCustomModelSwapDisplacements()`, the periodic sweep + * behind server.ts's "custom model swap-displacement check" timer + * (docs/custom-model-endpoints-plan.md). The apply/create routes' own swap-conflict check + * only ever runs at a session's own launch/apply moment — this sweep is what catches a + * LATER eviction triggered by a different session's normal use, which the launch-time + * check structurally cannot see. + * + * Kept in its own file for the same reason as `custom-model-endpoint-rediscovery.test.ts`: + * a sweep that walks every saved host would otherwise pick up hosts other tests in a + * shared file create, making an exact call-count assertion meaningless. + * + * Port: N/A (no server; drives readCustomModelHosts/writeCustomModelHosts directly plus + * the mocked webviewFetch dispatcher). + */ +import { describe, it, expect, vi, beforeEach } from 'vitest'; +import { getDataDir } from '../src/config/instance.js'; +import { writeCustomModelHosts, type CustomModelHost } from '../src/custom-model-hosts.js'; +import { + detectCustomModelSwapDisplacements, + type CustomModelSessionLike, +} from '../src/web/routes/custom-model-routes.js'; +import { webviewFetch } from '../src/web/webview-egress.js'; + +vi.mock('../src/web/webview-egress.js', async () => { + const actual = await vi.importActual('../src/web/webview-egress.js'); + return { ...actual, webviewFetch: vi.fn() }; +}); + +const fetchMock = vi.mocked(webviewFetch); + +const ENDPOINT: CustomModelHost = { + id: 'llama-swap', + label: 'llama-swap', + baseUrl: 'http://192.168.1.50:8080', + apiKey: 'k', +}; + +function session( + overrides: Partial & Pick +): CustomModelSessionLike { + return { name: overrides.id, ...overrides }; +} + +function mockRunning(running: Array<{ model: string; state: string }>) { + fetchMock.mockImplementation(async (url: URL) => { + if (url.pathname === '/running') return new Response(JSON.stringify({ running }), { status: 200 }); + throw new Error(`unexpected request in this test: ${url.href}`); + }); +} + +beforeEach(() => { + fetchMock.mockReset(); +}); + +describe('detectCustomModelSwapDisplacements', () => { + it('flags a session whose own model is no longer in the running list, naming what displaced it', async () => { + await writeCustomModelHosts(getDataDir(), [ENDPOINT]); + mockRunning([{ model: 'fast', state: 'ready' }]); + const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } }); + const notified = new Set(); + + const displacements = await detectCustomModelSwapDisplacements([w1], notified); + + expect(displacements).toEqual([ + { + sessionId: 'w1', + sessionName: 'w1', + endpointId: 'llama-swap', + previousModel: 'qwen3', + currentlyLoadedModel: 'fast', + }, + ]); + expect(notified.has('w1')).toBe(true); + }); + + it('does not flag a session whose own model is still the one loaded and ready', async () => { + await writeCustomModelHosts(getDataDir(), [ENDPOINT]); + mockRunning([{ model: 'qwen3', state: 'ready' }]); + const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } }); + + const displacements = await detectCustomModelSwapDisplacements([w1], new Set()); + + expect(displacements).toEqual([]); + }); + + it('notifies once per displacement — a repeat sweep with nothing changed does not re-flag it', async () => { + await writeCustomModelHosts(getDataDir(), [ENDPOINT]); + mockRunning([{ model: 'fast', state: 'ready' }]); + const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } }); + const notified = new Set(); + + const first = await detectCustomModelSwapDisplacements([w1], notified); + const second = await detectCustomModelSwapDisplacements([w1], notified); + + expect(first).toHaveLength(1); + expect(second).toEqual([]); + }); + + it('clears the notified flag once the session is back on its own model, so a later displacement flags again', async () => { + await writeCustomModelHosts(getDataDir(), [ENDPOINT]); + const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } }); + const notified = new Set(); + + mockRunning([{ model: 'fast', state: 'ready' }]); + await detectCustomModelSwapDisplacements([w1], notified); + expect(notified.has('w1')).toBe(true); + + mockRunning([{ model: 'qwen3', state: 'ready' }]); // back to normal + await detectCustomModelSwapDisplacements([w1], notified); + expect(notified.has('w1')).toBe(false); + + mockRunning([{ model: 'fast', state: 'ready' }]); // displaced again + const third = await detectCustomModelSwapDisplacements([w1], notified); + expect(third).toHaveLength(1); + }); + + it('skips a session on a non-llama-swap endpoint (no /running) — nothing to compare, never flagged', async () => { + await writeCustomModelHosts(getDataDir(), [ENDPOINT]); + fetchMock.mockResolvedValue(new Response('not found', { status: 404 })); + const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } }); + + const displacements = await detectCustomModelSwapDisplacements([w1], new Set()); + + expect(displacements).toEqual([]); + }); + + it('skips a session whose endpoint was deleted since it was created', async () => { + await writeCustomModelHosts(getDataDir(), []); // ENDPOINT never saved + const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } }); + + const displacements = await detectCustomModelSwapDisplacements([w1], new Set()); + + expect(displacements).toEqual([]); + expect(fetchMock).not.toHaveBeenCalled(); + }); + + it('ignores a plain session with no customModel selection at all', async () => { + const displacements = await detectCustomModelSwapDisplacements([session({ id: 'plain' })], new Set()); + expect(displacements).toEqual([]); + expect(fetchMock).not.toHaveBeenCalled(); + }); + + it('one endpoint failing (unreachable) never blocks checking sessions on another', async () => { + const DOWN: CustomModelHost = { id: 'down', label: 'down', baseUrl: 'http://192.168.1.60:8080' }; + await writeCustomModelHosts(getDataDir(), [ENDPOINT, DOWN]); + fetchMock.mockImplementation(async (url: URL) => { + if (url.href.includes('192.168.1.60')) throw new TypeError('fetch failed', { cause: new Error('ECONNREFUSED') }); + if (url.pathname === '/running') { + return new Response(JSON.stringify({ running: [{ model: 'fast', state: 'ready' }] }), { status: 200 }); + } + throw new Error(`unexpected request in this test: ${url.href}`); + }); + const onDown = session({ id: 'w-down', customModel: { endpointId: 'down', modelId: 'x' } }); + const onLlamaSwap = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } }); + + const displacements = await detectCustomModelSwapDisplacements([onDown, onLlamaSwap], new Set()); + + expect(displacements).toEqual([ + { + sessionId: 'w1', + sessionName: 'w1', + endpointId: 'llama-swap', + previousModel: 'qwen3', + currentlyLoadedModel: 'fast', + }, + ]); + }); + + it('multiple sessions on the same endpoint each get their own displacement entry', async () => { + await writeCustomModelHosts(getDataDir(), [ENDPOINT]); + mockRunning([{ model: 'gemma', state: 'ready' }]); + const w1 = session({ id: 'w1', name: 'w1-test2', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } }); + const w2 = session({ id: 'w2', name: 'w2-test2', customModel: { endpointId: 'llama-swap', modelId: 'fast' } }); + + const displacements = await detectCustomModelSwapDisplacements([w1, w2], new Set()); + + expect(displacements.map((d) => d.sessionId).sort()).toEqual(['w1', 'w2']); + }); +});