feat(custom-model): detect llama-swap model conflicts before switching

Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.

1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
   server) via its own GET /running, which plain llama.cpp has no concept of
   at all. New GET /api/model-endpoints/:id/running-status route exposes this
   read-only, for the frontend's polling loop below.

2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
   what llama-swap currently has loaded. If it differs from the requested
   model AND another live session's own customModel selection is actively
   using that loaded model, the apply is refused with a
   {requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
   instead of silently switching. A "confirmed: true" field on the retry
   skips the check. Switching with nothing else affected proceeds
   immediately, no confirmation asked, only ever when there is something to
   warn about.

3. The frontend (runCustomModelEntry) shows a native confirm() naming the
   affected session(s) and the model they'd lose, matching this codebase's
   existing convention for this class of decision (delete case, kill
   session, etc.) rather than a new modal. On a successful apply the response
   also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
   poll shows a sticky "Loading <model>..." toast via the new running-status
   route until llama-swap reports the target model ready (bounded at 2
   minutes), so a prompt sent mid-swap reads as "loading", never as silence
   or an answer from whatever was loaded a moment before.

Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 14:35:07 +08:00
co-authored by Claude Sonnet 5
parent 25f22b9839
commit bcebc81fcd
7 changed files with 493 additions and 8 deletions
+5
View File
@@ -5544,6 +5544,11 @@ Object.assign(CodemanApp.prototype, {
if (duration > 0) {
dismissTimer = setTimeout(dismiss, duration);
}
// Most callers ignore this — a handle exists for a long-running toast a caller needs
// to update or dismiss itself once its own condition resolves (e.g. a "loading model"
// toast a poll loop dismisses once the model reports ready).
return { dismiss, setMessage: (text) => { msgSpan.textContent = text; } };
},
+85 -6
View File
@@ -740,17 +740,96 @@ Object.assign(CodemanApp.prototype, {
// unreachable" apart from "the CLI can't be redirected", "not one of the
// discovered models", or "this is a Docker/remote session". Go through the
// raw response here instead so a failure is diagnosable, not just present.
const res = await this._api(`/api/sessions/${sessionId}/custom-model`, {
method: 'POST',
body: { endpointId, modelId },
});
const data = res ? await res.json().catch(() => null) : null;
if (!data || data.success === false) {
let { ok, data, res } = await this._applyCustomModelToSession(sessionId, endpointId, modelId);
// A success body comes back as {success:true, data:{...}} (server.ts's preSerialization
// envelope), but a route-level error is {success:false, error, errorCode} with no nested
// data — createErrorResponse() never wraps one. `payload` below is only ever meaningful
// once `data.success !== false`.
let payload = data?.success !== false ? data?.data : undefined;
// llama-swap runs one model at a time: switching would unload it out from under
// another session actively using it. The route only asks when that's actually true
// (never just because a swap is needed at all) — confirming re-sends the exact same
// call with `confirmed: true` so the route skips the check the second time.
if (ok && payload?.requiresConfirmation) {
const names = payload.affectedSessions.map((s) => s.name || s.id).join(', ');
const proceed = confirm(
`${names} ${payload.affectedSessions.length === 1 ? 'is' : 'are'} currently using ` +
`${payload.currentlyLoadedModel} on this endpoint. Switching to ${modelId} will unload it ` +
`for ${payload.affectedSessions.length === 1 ? 'that session' : 'those sessions'} too. Continue?`
);
if (!proceed) {
this.showToast('Kept the native backend — model switch cancelled', 'info');
return;
}
({ ok, data, res } = await this._applyCustomModelToSession(sessionId, endpointId, modelId, true));
payload = data?.success !== false ? data?.data : undefined;
}
if (!ok || !data || data.success === false) {
const detail = data?.error ? `: ${data.error}` : res ? ` (HTTP ${res.status})` : ' (request failed)';
this.showToast(`Session started on the native backend — could not apply the custom endpoint${detail}`, 'error');
return;
}
this.showToast(`Pointed at ${endpointId} — restarting the session...`, 'info');
// The apply above already succeeded — the session IS pointed at the endpoint — but
// llama-swap itself may still be unloading the old model and loading this one, which
// can take well over a minute. Without this, a prompt sent during that window either
// hangs silently or (the bug this whole feature exists to fix) gets answered by
// whatever was loaded a moment ago, reading as "it's still using the wrong model."
if (payload?.modelSwapInProgress) {
void this._watchLlamaSwapLoading(endpointId, modelId);
}
},
/** POST /api/sessions/:id/custom-model, returning {ok, data, res} rather than throwing —
* see runCustomModelEntry's own comment for why this goes through `_api()` (raw fetch)
* rather than `_apiJson()`: a failure's `error` detail must survive to the caller. */
async _applyCustomModelToSession(sessionId, endpointId, modelId, confirmed) {
const res = await this._api(`/api/sessions/${sessionId}/custom-model`, {
method: 'POST',
body: confirmed ? { endpointId, modelId, confirmed } : { endpointId, modelId },
});
const data = res ? await res.json().catch(() => null) : null;
return { ok: !!res, data, res };
},
/**
* Polls llama-swap's own `/running` (via the read-only running-status route) until
* `modelId` reports `state: 'ready'`, showing a sticky toast the whole time so a slow
* unload/reload (measured well over a minute for a large model) reads as "loading",
* never as silence or a wrong answer from whatever was loaded before. Bounded at 2
* minutes; still not ready by then gets a toast saying so rather than polling forever.
*
* `pollIntervalMs`/`maxWaitMs` exist to let a test drive this in milliseconds instead of
* minutes — real callers never pass them, which is what keeps the defaults live here
* rather than only in a test fixture.
*/
async _watchLlamaSwapLoading(endpointId, modelId, pollIntervalMs = 3000, maxWaitMs = 120000) {
const toast = this.showToast(`Loading ${modelId} on ${endpointId}… this can take a while`, 'info', {
duration: 0,
});
const deadline = Date.now() + maxWaitMs;
while (Date.now() < deadline) {
await new Promise((resolve) => setTimeout(resolve, pollIntervalMs));
const status = await this._apiJson(`/api/model-endpoints/${encodeURIComponent(endpointId)}/running-status`);
if (!status) continue; // transient failure — keep waiting rather than giving up early
if (!status.isLlamaSwap) {
// Endpoint changed under us, or wasn't llama-swap after all — nothing more to
// watch for, and not a failure worth a toast of its own.
toast?.dismiss();
return;
}
if (status.running.some((r) => r.model === modelId && r.state === 'ready')) {
toast?.dismiss();
this.showToast(`${modelId} is ready`, 'success', { duration: 2500 });
return;
}
}
toast?.dismiss();
this.showToast(`Still waiting for ${modelId} to finish loading on ${endpointId} — check the llama-swap server`, 'warning');
},
/**
+65
View File
@@ -177,6 +177,57 @@ type RedactedHost = ReturnType<typeof redactApiKey>;
* each do their own `discoverModels()` + error handling around one shared
* "how to apply a successful result" step.
*/
const RUNNING_TIMEOUT_MS = 5000;
export interface LlamaSwapRunningModel {
model: string;
state: string;
}
export interface LlamaSwapStatus {
/**
* Feature-detected via `GET /running`: true only when the server answered with
* llama-swap's own shape (`{ running: [...] }`). Plain llama.cpp (and any other
* OpenAI-compatible server) has no such endpoint and always runs the single model
* it was started with, so there is no "current model" to conflict with — every
* caller must treat `isLlamaSwap: false` as "nothing to check", never as an error.
*/
isLlamaSwap: boolean;
running: LlamaSwapRunningModel[];
}
/**
* Distinguishes llama-swap from a plain llama.cpp/OpenAI-compatible server, and reports
* what llama-swap currently has loaded — llama.cpp only ever runs one GGUF at a time, and
* llama-swap unloads/reloads it on demand when a request asks for a different one, which
* can take anywhere from a few seconds to over a minute. Read-only: this never triggers a
* swap itself (unlike `/props?model=`, `/running` takes no `model` parameter to route by).
* Best-effort like `discoverModels()`'s siblings: any failure (unreachable, non-2xx,
* unexpected shape) reads as "not llama-swap", never thrown.
*/
export async function getLlamaSwapStatus(
host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' | 'authStyle'>
): Promise<LlamaSwapStatus> {
try {
const res = await webviewFetch(new URL(`${host.baseUrl.replace(/\/+$/, '')}/running`), {
headers: authHeaders(host),
signal: AbortSignal.timeout(RUNNING_TIMEOUT_MS),
});
if (!res.ok) return { isLlamaSwap: false, running: [] };
const body = (await res.json()) as { running?: unknown };
if (!Array.isArray(body.running)) return { isLlamaSwap: false, running: [] };
const running = body.running
.filter(
(r): r is { model: string; state?: unknown } =>
!!r && typeof r === 'object' && typeof (r as { model?: unknown }).model === 'string'
)
.map((r) => ({ model: r.model, state: typeof r.state === 'string' ? r.state : 'unknown' }));
return { isLlamaSwap: true, running };
} catch {
return { isLlamaSwap: false, running: [] };
}
}
function applyDiscoveredModels(host: CustomModelHost, result: DiscoveryResult): CustomModelHost {
const { models, contextLengths } = result;
const defaultModelId = host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined;
@@ -303,4 +354,18 @@ export function registerCustomModelRoutes(app: FastifyInstance): void {
}
}
);
// Read-only, no admin gate: any session owner who can already point their own session
// at this endpoint (POST .../custom-model, ungated by design — see session-routes.ts)
// can equally ask what it currently has loaded, before or while that apply is pending.
app.get('/api/model-endpoints/:id/running-status', async (req): Promise<ApiResponse<LlamaSwapStatus>> => {
const { id } = req.params as { id: string };
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
const host = hosts.find((item) => item.id === id);
if (!host) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
if (isBlockedWebviewUrl(host.baseUrl)) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
return { success: true, data: await getLlamaSwapStatus(host) };
});
}
+28 -1
View File
@@ -55,6 +55,7 @@ import {
} from '../schemas.js';
import { readCustomModelHosts } from '../../custom-model-hosts.js';
import { applyCustomModelInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
import { getLlamaSwapStatus } from './custom-model-routes.js';
import { matchesPattern } from '../../config/cli-registry/patterns.js';
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
import { TabLayoutValidationError } from '../../tab-layout.js';
@@ -1209,6 +1210,32 @@ export function registerSessionRoutes(
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
}
// llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on
// demand, which can take anywhere from a few seconds to over a minute — long enough
// that a session mid-swap looks indistinguishable from one that never left the native
// backend. Feature-detected via llama-swap's own `GET /running` (a plain llama.cpp
// server has no such endpoint and reads as `isLlamaSwap: false` — nothing to check).
const swapStatus = await getLlamaSwapStatus(endpoint);
const currentlyLoaded = swapStatus.running.find((r) => r.state === 'ready')?.model ?? swapStatus.running[0]?.model;
const swapNeeded = swapStatus.isLlamaSwap && !!currentlyLoaded && currentlyLoaded !== body.modelId;
// Only ask when switching would actually take the model away from another session
// that is currently using it — never just because a swap is needed at all. `confirmed`
// (set by the caller after showing that warning once) skips asking again.
if (swapNeeded && !body.confirmed) {
const affectedSessions = [...ctx.sessions.values()]
.filter(
(s) =>
s.id !== session.id &&
s.customModel?.endpointId === endpoint.id &&
s.customModel?.modelId === currentlyLoaded
)
.map((s) => ({ id: s.id, name: s.name }));
if (affectedSessions.length > 0) {
return { requiresConfirmation: true, currentlyLoadedModel: currentlyLoaded, affectedSessions };
}
}
// A CLI whose config alone cannot select the model also gets its `model` launch param
// forced (pi/omp `custom/<id>`, grok's block name). The argv engine DROPS a token that
// fails its pattern rather than quoting it, which would silently launch the CLI on its
@@ -1250,7 +1277,7 @@ export function registerSessionRoutes(
const restarted = await session.restartCli();
persistAndBroadcastSession(ctx, session);
return { customModel: session.customModel, restarted };
return { customModel: session.customModel, restarted, modelSwapInProgress: swapNeeded };
});
// ========== Delete Session ==========
+4
View File
@@ -1934,6 +1934,10 @@ export const CustomModelSelectionSchema = z.union([
z.object({
endpointId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
modelId: z.string().min(1).max(200),
// Set once the caller has already shown the "this will unload <model> for session(s)
// X" warning (see session-routes.ts's llama-swap conflict check) and the user chose to
// proceed anyway — skips that check on this call instead of asking again.
confirmed: z.boolean().optional(),
}),
z.object({ clear: z.literal(true) }),
]);