feat(custom-model): remove loading-banner countdown, add manual Cancel

Replaces the size-scaled expected-time estimate + matching auto-timeout
with a generic hardware/model-size disclaimer and a user-driven Cancel
button, per explicit request. Real load time depends on hardware this
feature has no way to know (VRAM, storage speed, GPU contention), so
the old estimate/timeout was a guess dressed up as a fact — worse, one
that could kill a genuinely slow load partway through on slower
hardware.

- _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline
  entirely — polls indefinitely until ready or cancelled, no automatic
  give-up. Message is now "Loading <model> (<size>) on <endpoint> —
  this can take a while depending on your hardware and the model
  size.", with the real llama.cpp log line still on its own second
  line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/
  _formatRemaining (dead code once the countdown is gone) —
  _lookupModelSizeGB is kept, the GB figure still shows.
- _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real
  "Cancel" button (distinct from the error-type "×" close button,
  since Cancel has a real consequence) that calls it on click. Caller
  owns what cancelling actually means, same split as the swap-confirm
  modal's promise-resolving buttons.
- Cancelling dismisses the banner, shows an info toast (not an error —
  this was deliberate), and closes the session, mirroring what the old
  timeout used to do automatically but now on the user's own call.
- New .center-status-cancel CSS (bordered pill button, distinct from
  the plain "×" close glyph).

Test changes: removed the now-invalid timeout-auto-close/estimate
tests, added cancel-flow tests (dismiss/toast-type/session-close,
never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests
for the new Cancel button (bootAppWithRealCenterStatus, evaluating
panels-ui.js instead of stubbing _showCenterStatus, since this button
is worth verifying for real rather than just through the stub every
other test in the file uses). Typecheck/lint/frontend-syntax clean;
full suite shows no new regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 13:00:50 +08:00
co-authored by Claude Sonnet 5
parent 2d3fc65758
commit db9729e1fc
7 changed files with 253 additions and 179 deletions
@@ -23,4 +23,6 @@ Two more, from actually clicking through the swap-confirm and context-warning di
Remote (SSH) and Docker sessions are refused for now (400) — their restart reattaches the durable remote/in-container tmux rather than relaunching the agent.
- **The loading banner's countdown is gone, replaced by a generic disclaimer and a Cancel button.** Its size-scaled expected-time estimate and matching auto-timeout were both a guess dressed up as a fact — real load time depends on hardware this feature has no way to know, and a fixed number could kill a genuinely slow load partway through. The banner now says "this can take a while depending on your hardware and the model size", polls indefinitely, and carries a **Cancel** button that ends the wait and closes the session on the user's own call rather than a guessed deadline.
**One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all.** Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). `POST /api/quick-start` now accepts a `customModel` field and computes the same injection _before_ the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own `--resume`-based restart is far less jarring, and `runClaude()`'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work).
+32 -31
View File
@@ -103,22 +103,23 @@ into `modelSizesGB` — unlike context length, this needs no `/props` probe
(the figure is right there in the `/v1/models` response) and so is populated
for every model regardless of loaded state. A hand-configured profile's own
description has no such figure and correctly gets no entry, never a guess.
Used only to label the Run-menu picker's "loading model" banner with a
rough, UNMEASURED expected-time estimate (`_estimateModelLoad()` in
session-ui.js, based on typical local NVMe/SSD throughput — not benchmarked
against any real endpoint's actual hardware/storage) and to scale that same
banner's own give-up timeout for a very large model; never anything a
server-side check relies on.
Used only to label the Run-menu picker's "loading model" banner (e.g.
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB) on llama-swap..."); never anything
a server-side check relies on.
**The loading banner shows a live countdown against that same timeout, and
treats a real timeout as a failure, not a shrug.** It checks llama-swap's
own `/running` every second (`GET /api/model-endpoints/:id/running-status`)
and counts down against the size-scaled (or flat 5-minute) timeout live; if
the countdown reaches zero with the target model still not ready, the
banner turns into a sticky error naming the llama-swap server's own logs as
where to look, and the session the load was for is closed automatically —
a console left open and pointed at a model that never finished loading is
worse than no console at all.
**The loading banner is unbounded by design, and says so — no countdown, no
automatic give-up.** An earlier version scaled an expected-time estimate and
a timeout off the model's file size and auto-closed the session once that
elapsed, but a real load's actual duration depends on hardware this feature
has no way to know (VRAM, storage speed, whatever else is contending for the
GPU) — any fixed number was a guess dressed up as a fact, and a model that
genuinely takes 10+ minutes on slower hardware would just get killed
mid-load by its own display. The banner now says outright that it can take a
while depending on hardware and model size, polls
`GET /api/model-endpoints/:id/running-status` every second for as long as it
takes, and carries a **Cancel** button (rendered on the banner itself) that
ends the wait and closes the session the load was for — the user's own call
on when it's taking too long, not a fixed number baked into the client.
**The banner's second line is the real backend log line, not a guess.**
llama-swap's `GET /api/events` SSE stream carries the actual `llama-server`
@@ -126,22 +127,22 @@ process's own stdout — `load_model: loading model '<path>'`,
`llama_server: model loaded`, tokenizer warnings, all of it — tagged
`source: "upstream"`, distinct from llama-swap's own `source: "proxy"`
request-access lines. `running-status`'s response now includes `logLine`
(via `getLatestLlamaSwapLogLine`), and the banner shows it under the
countdown, e.g. "llama.cpp: load_model: loading model '...'" — confirmed
live end-to-end through a real forced swap, sequentially showing the model
path, a tokenizer warning, then staying on whatever llama.cpp last printed
once the load goes quiet (never cleared back to blank). ⚠️ **`GET /logs`
— the endpoint this feature's own first cut was built against — turns out
to carry ONLY llama-swap's own proxy request-access log.** Confirmed live
it never showed a single backend line, even seconds after a real, verified
model swap; `/api/events`'s `logData` frames are the only source that
actually has it, and its own `source` field (`upstream` vs `proxy`) is
what `getLatestLlamaSwapLogLine` filters on. One `/api/events` connection
is held open per endpoint and reused across every session watching a load
on it (confirmed live to stay open indefinitely, unlike `/logs`, which
closes after a fixed ~100KB), idle-closed after 30s of nobody polling it
(`pruneIdleLlamaSwapLogTails`, same 20s sweep as the swap-displacement
check below).
(via `getLatestLlamaSwapLogLine`), and the banner shows it on its own line
under the disclaimer, e.g. "llama.cpp: load_model: loading model '...'" —
confirmed live end-to-end through a real forced swap, sequentially showing
the model path, a tokenizer warning, then staying on whatever llama.cpp last
printed once the load goes quiet (never cleared back to blank). ⚠️
**`GET /logs` — the endpoint this feature's own first cut was built
against — turns out to carry ONLY llama-swap's own proxy request-access
log.** Confirmed live it never showed a single backend line, even seconds
after a real, verified model swap; `/api/events`'s `logData` frames are the
only source that actually has it, and its own `source` field (`upstream` vs
`proxy`) is what `getLatestLlamaSwapLogLine` filters on. One `/api/events`
connection is held open per endpoint and reused across every session
watching a load on it (confirmed live to stay open indefinitely, unlike
`/logs`, which closes after a fixed ~100KB), idle-closed after 30s of nobody
polling it (`pruneIdleLlamaSwapLogTails`, same 20s sweep as the
swap-displacement check below).
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
+9 -8
View File
@@ -80,14 +80,15 @@ just applying a selection never reached llama-swap's own logs at all until somet
it to load. Picking an entry now also sends the smallest real request that will trigger
that load, in the background, the moment the target model isn't already loaded and ready.
**The centred loading banner shows a live countdown, and a real timeout is an error, not a
shrug.** When it knows the model's discovered file size (its GB figure, when llama-swap
states one), it shows both a rough expected-time estimate and a live countdown against it —
e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1–3 min) on llama-swap — 47s remaining". If
the countdown reaches zero and the model still isn't ready, the banner turns into a sticky
error telling you to check the llama-swap server's own logs, and **the session that load was
for is closed automatically** — a console left open and pointed at a model that never
finished loading would just be confusing to leave sitting there.
**The centred loading banner has no countdown and no automatic timeout — it waits as long as
it takes, and tells you so.** When it knows the model's discovered file size (its GB figure,
when llama-swap states one) it's shown too, e.g. "Loading qwen3.8-27b (16.4 GB) on
llama-swap — this can take a while depending on your hardware and the model size." An
earlier version tried to estimate and enforce a time limit, but real load time depends on
hardware this feature has no way to know, so a fixed number was always a guess — worse, one
that could kill a genuinely slow load partway through. If it really is taking too long, a
**Cancel** button right on the banner ends the wait and **closes the session that load was
for**, on your own call rather than a guessed deadline.
**The banner also shows a real, live second line of what llama.cpp itself is doing** — not
a made-up progress phase, the actual next line the `llama-server` process printed, e.g.
+17 -1
View File
@@ -5567,9 +5567,16 @@ Object.assign(CodemanApp.prototype, {
* close button, since a sticky error the user cannot dismiss would just sit there). The
* DOM is rebuilt fresh each call rather than patched, since which children exist differs
* by type; `setMessage` still only ever touches the text node afterwards.
*
* `opts.onCancel` — when given (any type, but in practice only 'info': an 'error' banner
* already has its own close button), renders a "Cancel" button that calls it on click.
* The callback owns everything that follows (dismissing the banner, stopping whatever
* loop this was showing progress for, closing a session it was for) — this helper only
* renders the button and wires the click, the same "caller decides what cancel means"
* split as `_confirmModelSwap`'s promise-resolving buttons.
*/
_showCenterStatus(message, opts = {}) {
const { type = 'info' } = opts;
const { type = 'info', onCancel } = opts;
let el = document.getElementById('customModelCenterStatus');
if (!el) {
el = document.createElement('div');
@@ -5604,6 +5611,15 @@ Object.assign(CodemanApp.prototype, {
dismiss();
};
el.appendChild(closeBtn);
} else if (onCancel) {
const cancelBtn = document.createElement('button');
cancelBtn.className = 'center-status-cancel';
cancelBtn.textContent = 'Cancel';
cancelBtn.onclick = (e) => {
e.stopPropagation();
onCancel();
};
el.appendChild(cancelBtn);
}
el.hidden = false;
requestAnimationFrame(() => el.classList.add('show'));
+45 -74
View File
@@ -1009,40 +1009,6 @@ Object.assign(CodemanApp.prototype, {
return typeof size === 'number' && Number.isFinite(size) && size > 0 ? size : undefined;
},
/**
* Rough, UNMEASURED load-time brackets by model file size, for the loading banner's text
* and as a size-scaled fallback timeout (larger models get longer before
* _watchLlamaSwapLoading gives up and warns). Sourced from typical local NVMe/SSD
* throughput for llama.cpp's mmap-and-warm sequence — NOT benchmarked against any real
* endpoint's actual hardware/storage (network storage, spinning disks, or a GPU with
* less VRAM than the model needs would all be meaningfully slower), so the label is an
* expectation-setter, never a guarantee. `maxGB` is the bracket's own upper bound
* (inclusive); brackets are checked in order, so list them smallest first.
*/
_MODEL_LOAD_TIME_MATRIX: [
{ maxGB: 2, label: '~5–15s', waitMs: 60000 },
{ maxGB: 8, label: '~15–45s', waitMs: 120000 },
{ maxGB: 16, label: '~30–90s', waitMs: 180000 },
{ maxGB: 32, label: '~1–3 min', waitMs: 300000 },
{ maxGB: 64, label: '~2–5 min', waitMs: 480000 },
{ maxGB: Infinity, label: '~5+ min', waitMs: 900000 },
],
/** `sizeGB` -> `{label, waitMs}` from `_MODEL_LOAD_TIME_MATRIX`, or `null` when `sizeGB`
* is unknown (no estimate is always safer than a fabricated one). */
_estimateModelLoad(sizeGB) {
if (typeof sizeGB !== 'number' || !Number.isFinite(sizeGB) || sizeGB <= 0) return null;
return this._MODEL_LOAD_TIME_MATRIX.find((bracket) => sizeGB <= bracket.maxGB) ?? null;
},
/** `ms` -> `"1m 08s remaining"` / `"8s remaining"`, for the loading banner's live countdown. */
_formatRemaining(ms) {
const totalSec = Math.max(0, Math.ceil(ms / 1000));
const mins = Math.floor(totalSec / 60);
const secs = totalSec % 60;
return mins > 0 ? `${mins}m ${String(secs).padStart(2, '0')}s remaining` : `${secs}s remaining`;
},
/**
* Strips llama.cpp's own bootlog prefix (`<uptime> <I|W|E> <component> `, e.g.
* `0.31.428.568 I srv llama_server: model loaded`) for display, leaving just
@@ -1058,20 +1024,20 @@ Object.assign(CodemanApp.prototype, {
/**
* Polls llama-swap's own `/running` (via the read-only running-status route) until
* `modelId` reports `state: 'ready'`, showing a sticky banner with a live countdown the
* whole time so a slow unload/reload (measured well over a minute for a large model)
* reads as "loading, N seconds left", never as silence or a wrong answer from whatever
* was loaded before. Checks immediately (a fast load, or a re-apply onto an
* already-ready model, shouldn't wait a full interval to say so), then every
* `pollIntervalMs`. Bounded at `maxWaitMs` — defaults to a rough, size-scaled estimate
* (`_estimateModelLoad`) when the model's discovered size is known, falling back to a
* flat 5 minutes when it isn't.
* `modelId` reports `state: 'ready'`, showing a sticky banner the whole time so a slow
* unload/reload (measured well over a minute for a large model) reads as "loading,
* still working on it", never as silence or a wrong answer from whatever was loaded
* before. Checks immediately (a fast load, or a re-apply onto an already-ready model,
* shouldn't wait a full interval to say so), then every `pollIntervalMs`.
*
* If the countdown reaches zero with the model still not ready, this is a real failure,
* not a "keep waiting" — the banner turns into a sticky error naming the llama-swap
* server's own logs as where to look, and `sessionId` (the session this was launched
* for) is closed automatically: a console left open and pointed at a model that never
* finished loading is worse than no console at all.
* Deliberately UNBOUNDED — no estimate, no countdown, no automatic give-up. An earlier
* version scaled a timeout off the model's discovered file size and auto-closed the
* session when it elapsed, but a real load's actual duration depends on hardware this
* feature has no way to know (VRAM, storage speed, what else is contending for the
* GPU), so any fixed number was a guess dressed up as a fact — the banner now says so
* outright instead of pretending to a precision it doesn't have, and a Cancel button on
* the banner itself (`_showCenterStatus`'s `onCancel`) is how the user ends it if it's
* taking too long, closing `sessionId` the same way the old timeout used to.
*
* `_watchLlamaSwapGeneration` guards against two overlapping calls (a second launch
* started before the first one's loop finished) clobbering each other's banner:
@@ -1080,37 +1046,41 @@ Object.assign(CodemanApp.prototype, {
* or overwrite the WRONG one, or close the WRONG session. Each call claims the counter
* as its own "generation" and checks it still owns it before touching either.
*
* `pollIntervalMs`/`maxWaitMs` exist to let a test drive this in milliseconds instead of
* minutes — real callers never pass `maxWaitMs`, which is what keeps the size-scaled
* default live here rather than only in a test fixture.
* `pollIntervalMs` exists to let a test drive this in milliseconds instead of seconds —
* real callers never pass it.
*/
async _watchLlamaSwapLoading(endpointId, modelId, sessionId, pollIntervalMs = 1000, maxWaitMs) {
async _watchLlamaSwapLoading(endpointId, modelId, sessionId, pollIntervalMs = 1000) {
const generation = (this._watchLlamaSwapGeneration = (this._watchLlamaSwapGeneration || 0) + 1);
const isCurrent = () => this._watchLlamaSwapGeneration === generation;
const sizeGB = await this._lookupModelSizeGB(endpointId, modelId);
const estimate = this._estimateModelLoad(sizeGB);
const effectiveMaxWaitMs = maxWaitMs ?? estimate?.waitMs ?? 300000;
if (!isCurrent()) return; // a newer launch already took over before the lookup even finished
const sizeSuffix = sizeGB
? ` (${sizeGB.toFixed(1)} GB${estimate ? `, typically ${estimate.label}` : ''})`
: '';
const baseMessage = `Loading ${modelId}${sizeSuffix} on ${endpointId} —`;
// Second line, when llama-swap's /logs actually gives us one: the real backend
// llama-server process's own latest log line (load_model:/llama_server: ..., see
// getLatestLlamaSwapLogLine) — a countdown alone says "something is happening,
// trust me," this says what. Absent on the very first render (no poll has landed
// yet) and whenever the endpoint doesn't expose /logs at all — never fabricated.
const buildMessage = (remainingMs, logLine) => {
const sizeSuffix = sizeGB ? ` (${sizeGB.toFixed(1)} GB)` : '';
const baseMessage =
`Loading ${modelId}${sizeSuffix} on ${endpointId} — this can take a while depending on ` +
`your hardware and the model size.`;
// Second line, when llama-swap's own event feed actually gives us one: the real
// backend llama-server process's own latest log line (load_model:/llama_server: ...,
// see getLatestLlamaSwapLogLine) — a bare "please wait" says nothing is broken, this
// says what's actually happening. Absent on the very first render (no poll has
// landed yet) and whenever the endpoint doesn't expose it at all — never fabricated,
// and never cleared back to blank once seen (stays on the last real thing llama.cpp
// said if a later poll comes back with nothing new).
const buildMessage = (logLine) => {
const line = this._formatLlamaLogLine(logLine);
return `${baseMessage} ${this._formatRemaining(remainingMs)}` + (line ? `\nllama.cpp: ${line}` : '');
return baseMessage + (line ? `\nllama.cpp: ${line}` : '');
};
let cancelled = false;
// Prominent and screen-centred, not a corner toast — a real llama-swap model load can
// sit on screen for well over a minute, easy to mistake for nothing happening there.
const deadline = Date.now() + effectiveMaxWaitMs;
const toast = this._showCenterStatus(buildMessage(deadline - Date.now()));
while (Date.now() < deadline) {
const toast = this._showCenterStatus(buildMessage(), {
onCancel: () => {
cancelled = true;
},
});
while (!cancelled) {
const status = await this._apiJson(`/api/model-endpoints/${encodeURIComponent(endpointId)}/running-status`);
if (!isCurrent()) return; // a newer launch took over the banner — this loop is done
if (cancelled) break;
if (!status) {
// transient failure — keep waiting rather than giving up early
} else if (!status.isLlamaSwap) {
@@ -1123,16 +1093,17 @@ Object.assign(CodemanApp.prototype, {
this.showToast(`${modelId} is ready`, 'success', { duration: 2500 });
return;
}
if (!isCurrent()) return;
toast?.setMessage(buildMessage(deadline - Date.now(), status?.logLine));
if (!isCurrent() || cancelled) break;
toast?.setMessage(buildMessage(status?.logLine));
await new Promise((resolve) => setTimeout(resolve, pollIntervalMs));
}
if (!isCurrent()) return;
this._showCenterStatus(
`${modelId} did not finish loading on ${endpointId} within the expected time. ` +
`Check the llama-swap server logs for details.` +
(sessionId ? ' The session has been closed.' : ''),
{ type: 'error' }
// Cancelled by the user, not a timeout — an ordinary info toast, not a scary error
// banner, since this was deliberate rather than something going wrong.
toast?.dismiss();
this.showToast(
`Cancelled loading ${modelId} on ${endpointId}` + (sessionId ? ' — the session has been closed.' : '.'),
'info'
);
if (sessionId) {
try {
+22
View File
@@ -8636,6 +8636,28 @@ kbd {
opacity: 1;
}
/* The Cancel button on an 'info' banner (e.g. the model-loading banner) — a real button
rather than the bare "×" close glyph above, since "Cancel" is an action with a
consequence (the caller's onCancel closes a session), not a plain dismiss. */
.center-status-cancel {
flex-shrink: 0;
pointer-events: auto;
background: none;
border: 1px solid var(--border);
border-radius: 6px;
color: inherit;
opacity: 0.75;
font-size: 0.8rem;
font-weight: 500;
padding: 0.25rem 0.6rem;
cursor: pointer;
}
.center-status-cancel:hover {
opacity: 1;
border-color: var(--text-muted, var(--border));
}
.toast-success { border-color: rgba(34, 197, 94, 0.4); }
.toast-error { border-color: rgba(239, 68, 68, 0.4); }
.toast-warning { border-color: rgba(234, 179, 8, 0.4); }
+126 -65
View File
@@ -20,6 +20,9 @@ import { describe, expect, it } from 'vitest';
const CONSTANTS_JS = readFileSync(new URL('../src/web/public/constants.js', import.meta.url), 'utf-8');
const SESSION_UI_JS = readFileSync(new URL('../src/web/public/session-ui.js', import.meta.url), 'utf-8');
// Only for the real _showCenterStatus DOM tests below (`bootAppWithRealCenterStatus`) —
// every other test in this file stubs _showCenterStatus itself and has no need of it.
const PANELS_UI_JS = readFileSync(new URL('../src/web/public/panels-ui.js', import.meta.url), 'utf-8');
function resp(body: unknown, ok = true) {
return { ok, json: async () => body };
@@ -98,6 +101,28 @@ function bootApp(
return { dom, win, app };
}
/**
* Like `bootApp`, but also evaluates panels-ui.js so `_showCenterStatus` is the REAL
* implementation rather than the plain stub `bootApp` installs — for the Cancel-button
* rendering tests, which need to see actual DOM the app would produce.
*/
function bootAppWithRealCenterStatus() {
const dom = new JSDOM('<!doctype html><body></body>', { url: 'http://localhost/', runScripts: 'dangerously' });
const win = dom.window as unknown as Window & typeof globalThis & { CodemanApp: new () => any };
// jsdom doesn't polyfill requestAnimationFrame, and _showCenterStatus calls it to add
// the 'show' class — run it synchronously, which is all a non-visual test needs.
(win as unknown as { requestAnimationFrame: (cb: () => void) => number }).requestAnimationFrame = (cb) => {
cb();
return 0;
};
(win as unknown as { eval: (s: string) => void }).eval('window.CodemanApp = function CodemanApp() {};');
(win as unknown as { eval: (s: string) => void }).eval(CONSTANTS_JS);
(win as unknown as { eval: (s: string) => void }).eval(SESSION_UI_JS);
(win as unknown as { eval: (s: string) => void }).eval(PANELS_UI_JS);
const app = new win.CodemanApp();
return { win, app };
}
describe('Custom Model Endpoint Profiles: Run-menu picker generation', () => {
it('generates a real, clickable button per (capable CLI, endpoint) pair', async () => {
const { win, app } = bootApp({
@@ -619,7 +644,7 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
};
app._apiJson = async () => ({ isLlamaSwap: true, running: [{ model: 'qwen3', state: 'ready' }] });
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 200);
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5);
expect(bannerMessages[0]).toMatch(/loading qwen3/i);
expect(dismissed).toContain(bannerMessages[0]);
@@ -648,7 +673,7 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
return { isLlamaSwap: true, running: [{ model: 'qwen3', state: 'ready' }] };
};
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 200);
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5);
// First render (before any poll has landed) has no log line at all.
expect(bannerMessages[0]).not.toMatch(/llama\.cpp:/);
@@ -669,36 +694,70 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
app.showToast = () => {};
app._apiJson = async () => ({ isLlamaSwap: true, running: [{ model: 'qwen3', state: 'ready' }] });
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 200);
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5);
expect(bannerMessages.some((m) => m.includes('llama.cpp:'))).toBe(false);
});
it('gives up after the bounded wait, turns the banner into a sticky error, and closes the session', async () => {
it('is unbounded — never gives up on its own, even after many polls with no ready model', async () => {
// No countdown, no timeout: confirms the loop just keeps polling rather than
// eventually erroring out on its own after some fixed number of checks.
const { app } = bootApp({});
const banners: Array<{ message: string; opts: unknown }> = [];
app._showCenterStatus = (message: string, opts: unknown) => {
banners.push({ message, opts });
return { dismiss: () => {}, setMessage: () => {} };
};
app._showCenterStatus = () => ({ dismiss: () => {}, setMessage: () => {} });
app.showToast = () => {};
let calls = 0;
app._apiJson = async (path: string) => {
if (path === '/api/model-endpoints') return null;
calls += 1;
if (calls >= 20) return { isLlamaSwap: true, running: [{ model: 'qwen3', state: 'ready' }] };
return { isLlamaSwap: true, running: [{ model: 'something-else', state: 'ready' }] };
};
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 1);
expect(calls).toBe(20); // it really did keep polling past what the old bounded wait allowed
});
it('clicking Cancel on the banner dismisses it, shows an info toast (not an error), and closes the session', async () => {
const { app } = bootApp({});
let onCancel: (() => void) | undefined;
let dismissed = false;
app._showCenterStatus = (_message: string, opts?: { onCancel?: () => void }) => {
onCancel = opts?.onCancel;
return { dismiss: () => (dismissed = true), setMessage: () => {} };
};
const toastCalls: Array<{ message: string; type: string }> = [];
app.showToast = (message: string, type = 'info') => {
toastCalls.push({ message, type });
};
let closedSessionId: string | undefined;
app.closeSession = async (id: string) => {
closedSessionId = id;
};
app._apiJson = async () => ({ isLlamaSwap: true, running: [{ model: 'something-else', state: 'ready' }] });
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 30);
const watch = app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5);
// Give the loop a couple of ticks to actually be polling, then cancel it — a real
// click happens whenever the user gets around to it, not on the very first render.
await new Promise((resolve) => setTimeout(resolve, 15));
expect(onCancel).toBeTypeOf('function');
onCancel!();
await watch;
const errorBanner = banners.find((b) => (b.opts as { type?: string } | undefined)?.type === 'error');
expect(errorBanner?.message).toMatch(/did not finish loading/i);
expect(errorBanner?.message).toMatch(/llama-swap server logs/i);
expect(dismissed).toBe(true);
const cancelToast = toastCalls.find((t) => /cancelled/i.test(t.message));
expect(cancelToast?.type).toBe('info'); // not 'error' — this was deliberate, not a failure
expect(cancelToast?.message).toMatch(/session has been closed/i);
expect(closedSessionId).toBe('sess-1');
});
it('never closes anything when no sessionId was given (a caller that has none to close)', async () => {
const { app } = bootApp({});
app._showCenterStatus = () => ({ dismiss: () => {}, setMessage: () => {} });
let onCancel: (() => void) | undefined;
app._showCenterStatus = (_message: string, opts?: { onCancel?: () => void }) => {
onCancel = opts?.onCancel;
return { dismiss: () => {}, setMessage: () => {} };
};
app.showToast = () => {};
let closeCalled = false;
app.closeSession = async () => {
@@ -706,7 +765,10 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
};
app._apiJson = async () => ({ isLlamaSwap: true, running: [{ model: 'something-else', state: 'ready' }] });
await app._watchLlamaSwapLoading('llama-box', 'qwen3', undefined, 5, 30);
const watch = app._watchLlamaSwapLoading('llama-box', 'qwen3', undefined, 5);
await new Promise((resolve) => setTimeout(resolve, 15));
onCancel!();
await watch;
expect(closeCalled).toBe(false);
});
@@ -726,7 +788,7 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
};
app._apiJson = async () => ({ isLlamaSwap: false, running: [] });
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 200);
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5);
expect(bannerDismissed).toBe(true);
expect(toastCalls).toHaveLength(0); // no follow-up warning toast
@@ -746,7 +808,7 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
return { isLlamaSwap: true, running: [{ model: 'qwen3', state: 'ready' }] };
};
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 200);
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5);
expect(toastCalls.at(-1)).toMatch(/ready/i);
});
@@ -766,7 +828,7 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
// A huge interval that would time the test out if the function actually waited for
// it before the first check.
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 60000, 300000);
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 60000);
expect(calls).toBe(1);
});
@@ -779,48 +841,34 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
setMessage: () => {},
});
app.showToast = () => {};
// The FIRST call never sees its own target model ready, so left alone it would run all
// the way to its own timeout and (now) turn into an error + close its session — but no
// sessionId is passed, so there is nothing for it to close even if it does get there.
// The FIRST call never sees its own target model ready — left alone (unbounded, no
// timeout) it would poll forever, but being superseded below must still make it stop
// on its own very next isCurrent() check rather than needing a timeout to exit.
app._apiJson = async (path: string) => {
if (path === '/api/model-endpoints') return [];
return { isLlamaSwap: true, running: [] };
};
const firstCall = app._watchLlamaSwapLoading('llama-box', 'model-a', undefined, 5, 30);
const firstCall = app._watchLlamaSwapLoading('llama-box', 'model-a', undefined, 5);
// Second call, for a DIFFERENT model that IS ready right away, takes over the banner
// before the first call's own bounded wait has elapsed.
// Second call, for a DIFFERENT model that IS ready right away, takes over the banner.
app._apiJson = async (path: string) => {
if (path === '/api/model-endpoints') return [];
return { isLlamaSwap: true, running: [{ model: 'model-b', state: 'ready' }] };
};
await app._watchLlamaSwapLoading('llama-box', 'model-b', undefined, 5, 200);
await app._watchLlamaSwapLoading('llama-box', 'model-b', undefined, 5);
// Let the stale first call run out its own bounded wait and finish.
// Let the stale first call notice it's been superseded and return on its own.
await firstCall;
// Whatever the first call did or didn't show along the way, its own eventual
// completion (a timeout, in this case) must never touch a banner state that belongs
// to the newer, still-current call — exactly one dismiss, for model-b, is the tell.
// Whatever the first call did or didn't show along the way, being superseded must
// never touch a banner state that belongs to the newer, still-current call — exactly
// one dismiss, for model-b, is the tell.
expect(dismissCalls).toHaveLength(1);
expect(dismissCalls[0]).toContain('model-b');
});
});
describe('Custom Model Endpoint Profiles: model-size load-time estimate', () => {
it('_estimateModelLoad picks the smallest matching bracket, and returns null for an unknown size', () => {
const { app } = bootApp({});
expect(app._estimateModelLoad(1)).toMatchObject({ label: '~5–15s' });
expect(app._estimateModelLoad(2)).toMatchObject({ label: '~5–15s' }); // inclusive upper bound
expect(app._estimateModelLoad(2.1)).toMatchObject({ label: '~15–45s' });
expect(app._estimateModelLoad(16.35)).toMatchObject({ label: '~1–3 min' }); // just over the 16GB bracket
expect(app._estimateModelLoad(200)).toMatchObject({ label: '~5+ min' });
expect(app._estimateModelLoad(undefined)).toBeNull();
expect(app._estimateModelLoad(0)).toBeNull();
expect(app._estimateModelLoad(-5)).toBeNull();
expect(app._estimateModelLoad(NaN)).toBeNull();
});
describe('Custom Model Endpoint Profiles: model size lookup (no time estimate — see the unbounded-wait describe above)', () => {
it('_lookupModelSizeGB reads the size off the matching endpoint/model, ignoring one with no parseable size', async () => {
const { app } = bootApp({});
app._apiJson = async (path: string) => {
@@ -848,7 +896,7 @@ describe('Custom Model Endpoint Profiles: model-size load-time estimate', () =>
await expect(app._lookupModelSizeGB('llama-box', 'qwen3')).resolves.toBeUndefined();
});
it('the loading banner includes the size and estimate when the size is known', async () => {
it('the loading banner includes the size, and the generic hardware/model-size disclaimer, when the size is known', async () => {
const { app } = bootApp({});
const bannerMessages: string[] = [];
app._showCenterStatus = (message: string) => {
@@ -865,12 +913,12 @@ describe('Custom Model Endpoint Profiles: model-size load-time estimate', () =>
await app._watchLlamaSwapLoading('llama-box', 'qwen3.8-27b-ud-q4_k_xl', undefined, 5);
expect(bannerMessages[0]).toMatch(
/^Loading qwen3\.8-27b-ud-q4_k_xl \(16\.4 GB, typically ~1–3 min\) on llama-box — .+ remaining$/
expect(bannerMessages[0]).toBe(
'Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB) on llama-box — this can take a while depending on your hardware and the model size.'
);
});
it('the loading banner omits the size/estimate entirely when the size is unknown', async () => {
it('the loading banner omits the size but keeps the disclaimer when the size is unknown', async () => {
const { app } = bootApp({});
const bannerMessages: string[] = [];
app._showCenterStatus = (message: string) => {
@@ -885,28 +933,41 @@ describe('Custom Model Endpoint Profiles: model-size load-time estimate', () =>
await app._watchLlamaSwapLoading('llama-box', 'big', undefined, 5);
expect(bannerMessages[0]).toMatch(/^Loading big on llama-box — .+ remaining$/);
expect(bannerMessages[0]).toBe(
'Loading big on llama-box — this can take a while depending on your hardware and the model size.'
);
});
});
describe('Custom Model Endpoint Profiles: _showCenterStatus Cancel button (real DOM, not the stub)', () => {
it('renders a real, clickable Cancel button when onCancel is given, and wires it up', () => {
const { win, app } = bootAppWithRealCenterStatus();
let cancelled = false;
app._showCenterStatus('Loading qwen3 on llama-box…', { onCancel: () => (cancelled = true) });
const btn = win.document.querySelector('.center-status-cancel') as HTMLButtonElement | null;
expect(btn).not.toBeNull();
expect(btn!.textContent).toBe('Cancel');
btn!.onclick!(new (win as any).Event('click'));
expect(cancelled).toBe(true);
});
it('uses the size-scaled estimate as the default timeout when maxWaitMs is not passed', async () => {
// A 200GB model estimates to the top "~5+ min" bracket (900000ms); a huge poll interval
// would time the TEST out if the function only waited the flat, smaller previous
// default (300000ms) instead of the size-scaled one.
const { app } = bootApp({});
app._showCenterStatus = () => ({ dismiss: () => {}, setMessage: () => {} });
app.showToast = () => {};
let calls = 0;
app._apiJson = async (path: string) => {
if (path === '/api/model-endpoints') return [{ id: 'llama-box', modelSizesGB: { huge: 200 } }];
calls += 1;
if (calls < 3) return { isLlamaSwap: true, running: [] }; // not ready on the first couple of checks
return { isLlamaSwap: true, running: [{ model: 'huge', state: 'ready' }] };
};
it('renders no Cancel button at all when onCancel is not given', () => {
const { win, app } = bootAppWithRealCenterStatus();
// pollIntervalMs only — maxWaitMs omitted, so it must fall back to the size estimate.
await app._watchLlamaSwapLoading('llama-box', 'huge', undefined, 5);
app._showCenterStatus('Loading qwen3 on llama-box…');
expect(calls).toBe(3);
expect(win.document.querySelector('.center-status-cancel')).toBeNull();
});
it("an 'error' banner keeps its own × close button rather than growing a redundant Cancel, even if onCancel is passed", () => {
const { win, app } = bootAppWithRealCenterStatus();
app._showCenterStatus('Something went wrong', { type: 'error', onCancel: () => {} });
expect(win.document.querySelector('.center-status-close')).not.toBeNull();
expect(win.document.querySelector('.center-status-cancel')).toBeNull();
});
});