mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-08 16:39:42 +02:00
feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout (updated every poll, so every second by default) instead of a static "this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1-3 min) on llama-swap - 47s remaining". If the countdown reaches zero and the model still isn't ready, this is now treated as a real failure rather than a "keep waiting" shrug: - The banner turns into a sticky error (_showCenterStatus gains a `type` option - 'error' drops the spinner and adds a close button, since nothing is "in progress" anymore and a sticky message needs a way to dismiss it), naming the llama-swap server's own logs as where to look for detail. - The session that load was for is closed automatically (closeSession) - requested explicitly: a console left open and pointed at a model that never finished loading is worse than no console at all. Both apply paths now thread the new session's id through to _watchLlamaSwapLoading for this (new required 3rd parameter, after endpointId/modelId). _watchLlamaSwapGeneration's existing stale-call guard extends naturally to this: a superseded call's own eventual timeout recognises it no longer owns the banner and neither shows the error nor closes a session that may by then belong to a different, newer launch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
55dae31530
commit
7bbe408e44
@@ -97,6 +97,16 @@ against any real endpoint's actual hardware/storage) and to scale that same
|
|||||||
banner's own give-up timeout for a very large model; never anything a
|
banner's own give-up timeout for a very large model; never anything a
|
||||||
server-side check relies on.
|
server-side check relies on.
|
||||||
|
|
||||||
|
**The loading banner shows a live countdown against that same timeout, and
|
||||||
|
treats a real timeout as a failure, not a shrug.** It checks llama-swap's
|
||||||
|
own `/running` every second (`GET /api/model-endpoints/:id/running-status`)
|
||||||
|
and counts down against the size-scaled (or flat 5-minute) timeout live; if
|
||||||
|
the countdown reaches zero with the target model still not ready, the
|
||||||
|
banner turns into a sticky error naming the llama-swap server's own logs as
|
||||||
|
where to look, and the session the load was for is closed automatically —
|
||||||
|
a console left open and pointed at a model that never finished loading is
|
||||||
|
worse than no console at all.
|
||||||
|
|
||||||
`defaultModelId` names which discovered model the picker pre-marks for that
|
`defaultModelId` names which discovered model the picker pre-marks for that
|
||||||
endpoint — the settings panel's Edit form exposes it as a select populated
|
endpoint — the settings panel's Edit form exposes it as a select populated
|
||||||
from the endpoint's own discovered `models`, and the route refuses a value
|
from the endpoint's own discovered `models`, and the route refuses a value
|
||||||
|
|||||||
@@ -78,10 +78,16 @@ waiting on your first prompt to do it.** llama-swap has no "switch model" button
|
|||||||
— the only thing that starts a swap is a real request naming the model, and confirmed live:
|
— the only thing that starts a swap is a real request naming the model, and confirmed live:
|
||||||
just applying a selection never reached llama-swap's own logs at all until something asked
|
just applying a selection never reached llama-swap's own logs at all until something asked
|
||||||
it to load. Picking an entry now also sends the smallest real request that will trigger
|
it to load. Picking an entry now also sends the smallest real request that will trigger
|
||||||
that load, in the background, the moment the target model isn't already loaded and ready —
|
that load, in the background, the moment the target model isn't already loaded and ready.
|
||||||
which is what the prominent **"Loading `<model>`… this can take a while"** banner
|
|
||||||
(centred on screen, not a corner toast — a real load can take well over a minute) is
|
**The centred loading banner shows a live countdown, and a real timeout is an error, not a
|
||||||
actually watching for.
|
shrug.** When it knows the model's discovered file size (its GB figure, when llama-swap
|
||||||
|
states one), it shows both a rough expected-time estimate and a live countdown against it —
|
||||||
|
e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1–3 min) on llama-swap — 47s remaining". If
|
||||||
|
the countdown reaches zero and the model still isn't ready, the banner turns into a sticky
|
||||||
|
error telling you to check the llama-swap server's own logs, and **the session that load was
|
||||||
|
for is closed automatically** — a console left open and pointed at a model that never
|
||||||
|
finished loading would just be confusing to leave sitting there.
|
||||||
|
|
||||||
**Claude Code specifically gets two extra fixes applied automatically:**
|
**Claude Code specifically gets two extra fixes applied automatically:**
|
||||||
|
|
||||||
|
|||||||
+40
-18
@@ -5556,37 +5556,59 @@ Object.assign(CodemanApp.prototype, {
|
|||||||
* genuinely worth interrupting the eye for rather than living in the corner with every
|
* genuinely worth interrupting the eye for rather than living in the corner with every
|
||||||
* other toast (currently: a custom-model session's "switching backends" and "loading
|
* other toast (currently: a custom-model session's "switching backends" and "loading
|
||||||
* model" states, both of which can sit on screen for well over a minute and are easy to
|
* model" states, both of which can sit on screen for well over a minute and are easy to
|
||||||
* mistake for nothing happening). Non-blocking (`pointer-events: none`, no backdrop) —
|
* mistake for nothing happening). Non-blocking (`pointer-events: none` on the wrapper,
|
||||||
* this is informational, never a gate the user has to dismiss to keep working. Only one
|
* restored only on the card) — an info banner is never a gate the user has to dismiss to
|
||||||
* is ever shown at a time (the DOM node is created once and reused), which matches every
|
* keep working. Only one is ever shown at a time (the DOM node is created once and
|
||||||
* current caller: each hands off to the next rather than stacking.
|
* reused), which matches every current caller: each hands off to the next rather than
|
||||||
|
* stacking.
|
||||||
|
*
|
||||||
|
* `opts.type` — `'info'` (default, spinner, no close button — a caller ends it itself via
|
||||||
|
* `dismiss()`) or `'error'` (no spinner — nothing is in progress once this shows — with a
|
||||||
|
* close button, since a sticky error the user cannot dismiss would just sit there). The
|
||||||
|
* DOM is rebuilt fresh each call rather than patched, since which children exist differs
|
||||||
|
* by type; `setMessage` still only ever touches the text node afterwards.
|
||||||
*/
|
*/
|
||||||
_showCenterStatus(message) {
|
_showCenterStatus(message, opts = {}) {
|
||||||
|
const { type = 'info' } = opts;
|
||||||
let el = document.getElementById('customModelCenterStatus');
|
let el = document.getElementById('customModelCenterStatus');
|
||||||
if (!el) {
|
if (!el) {
|
||||||
el = document.createElement('div');
|
el = document.createElement('div');
|
||||||
el.id = 'customModelCenterStatus';
|
el.id = 'customModelCenterStatus';
|
||||||
el.className = 'center-status-banner';
|
document.body.appendChild(el);
|
||||||
|
}
|
||||||
|
el.className = `center-status-banner center-status-${type}`;
|
||||||
|
el.innerHTML = '';
|
||||||
|
const dismiss = () => {
|
||||||
|
el.classList.remove('show');
|
||||||
|
setTimeout(() => {
|
||||||
|
el.hidden = true;
|
||||||
|
}, 200);
|
||||||
|
};
|
||||||
|
if (type !== 'error') {
|
||||||
const spinner = document.createElement('span');
|
const spinner = document.createElement('span');
|
||||||
spinner.className = 'center-status-spinner';
|
spinner.className = 'center-status-spinner';
|
||||||
spinner.setAttribute('aria-hidden', 'true');
|
spinner.setAttribute('aria-hidden', 'true');
|
||||||
const text = document.createElement('span');
|
|
||||||
text.className = 'center-status-text';
|
|
||||||
el.appendChild(spinner);
|
el.appendChild(spinner);
|
||||||
el.appendChild(text);
|
|
||||||
document.body.appendChild(el);
|
|
||||||
}
|
}
|
||||||
const textEl = el.querySelector('.center-status-text');
|
const text = document.createElement('span');
|
||||||
if (textEl) textEl.textContent = message;
|
text.className = 'center-status-text';
|
||||||
|
text.textContent = message;
|
||||||
|
el.appendChild(text);
|
||||||
|
if (type === 'error') {
|
||||||
|
const closeBtn = document.createElement('button');
|
||||||
|
closeBtn.className = 'center-status-close';
|
||||||
|
closeBtn.textContent = '×';
|
||||||
|
closeBtn.setAttribute('aria-label', 'Dismiss');
|
||||||
|
closeBtn.onclick = (e) => {
|
||||||
|
e.stopPropagation();
|
||||||
|
dismiss();
|
||||||
|
};
|
||||||
|
el.appendChild(closeBtn);
|
||||||
|
}
|
||||||
el.hidden = false;
|
el.hidden = false;
|
||||||
requestAnimationFrame(() => el.classList.add('show'));
|
requestAnimationFrame(() => el.classList.add('show'));
|
||||||
return {
|
return {
|
||||||
dismiss: () => {
|
dismiss,
|
||||||
el.classList.remove('show');
|
|
||||||
setTimeout(() => {
|
|
||||||
el.hidden = true;
|
|
||||||
}, 200);
|
|
||||||
},
|
|
||||||
setMessage: (next) => {
|
setMessage: (next) => {
|
||||||
const t = el.querySelector('.center-status-text');
|
const t = el.querySelector('.center-status-text');
|
||||||
if (t) t.textContent = next;
|
if (t) t.textContent = next;
|
||||||
|
|||||||
@@ -765,7 +765,7 @@ Object.assign(CodemanApp.prototype, {
|
|||||||
const result = this._lastCustomModelLaunchResult;
|
const result = this._lastCustomModelLaunchResult;
|
||||||
this._lastCustomModelLaunchResult = undefined;
|
this._lastCustomModelLaunchResult = undefined;
|
||||||
if (result?.modelSwapInProgress) {
|
if (result?.modelSwapInProgress) {
|
||||||
void this._watchLlamaSwapLoading(endpointId, modelId);
|
void this._watchLlamaSwapLoading(endpointId, modelId, result.sessionId);
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
|
|
||||||
@@ -906,7 +906,7 @@ Object.assign(CodemanApp.prototype, {
|
|||||||
// Hand off to its own sticky toast rather than stacking a second one on top.
|
// Hand off to its own sticky toast rather than stacking a second one on top.
|
||||||
if (payload?.modelSwapInProgress) {
|
if (payload?.modelSwapInProgress) {
|
||||||
switchingToast?.dismiss();
|
switchingToast?.dismiss();
|
||||||
void this._watchLlamaSwapLoading(endpointId, modelId);
|
void this._watchLlamaSwapLoading(endpointId, modelId, sessionId);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -967,29 +967,43 @@ Object.assign(CodemanApp.prototype, {
|
|||||||
return this._MODEL_LOAD_TIME_MATRIX.find((bracket) => sizeGB <= bracket.maxGB) ?? null;
|
return this._MODEL_LOAD_TIME_MATRIX.find((bracket) => sizeGB <= bracket.maxGB) ?? null;
|
||||||
},
|
},
|
||||||
|
|
||||||
|
/** `ms` -> `"1m 08s remaining"` / `"8s remaining"`, for the loading banner's live countdown. */
|
||||||
|
_formatRemaining(ms) {
|
||||||
|
const totalSec = Math.max(0, Math.ceil(ms / 1000));
|
||||||
|
const mins = Math.floor(totalSec / 60);
|
||||||
|
const secs = totalSec % 60;
|
||||||
|
return mins > 0 ? `${mins}m ${String(secs).padStart(2, '0')}s remaining` : `${secs}s remaining`;
|
||||||
|
},
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Polls llama-swap's own `/running` (via the read-only running-status route) until
|
* Polls llama-swap's own `/running` (via the read-only running-status route) until
|
||||||
* `modelId` reports `state: 'ready'`, showing a sticky banner the whole time so a slow
|
* `modelId` reports `state: 'ready'`, showing a sticky banner with a live countdown the
|
||||||
* unload/reload (measured well over a minute for a large model) reads as "loading",
|
* whole time so a slow unload/reload (measured well over a minute for a large model)
|
||||||
* never as silence or a wrong answer from whatever was loaded before. Checks immediately
|
* reads as "loading, N seconds left", never as silence or a wrong answer from whatever
|
||||||
* (a fast load, or a re-apply onto an already-ready model, shouldn't wait a full interval
|
* was loaded before. Checks immediately (a fast load, or a re-apply onto an
|
||||||
* to say so), then every `pollIntervalMs`. Bounded at `maxWaitMs` — defaults to a rough,
|
* already-ready model, shouldn't wait a full interval to say so), then every
|
||||||
* size-scaled estimate (`_estimateModelLoad`) when the model's discovered size is known,
|
* `pollIntervalMs`. Bounded at `maxWaitMs` — defaults to a rough, size-scaled estimate
|
||||||
* falling back to a flat 5 minutes when it isn't; still not ready by then gets a toast
|
* (`_estimateModelLoad`) when the model's discovered size is known, falling back to a
|
||||||
* saying so rather than polling forever.
|
* flat 5 minutes when it isn't.
|
||||||
|
*
|
||||||
|
* If the countdown reaches zero with the model still not ready, this is a real failure,
|
||||||
|
* not a "keep waiting" — the banner turns into a sticky error naming the llama-swap
|
||||||
|
* server's own logs as where to look, and `sessionId` (the session this was launched
|
||||||
|
* for) is closed automatically: a console left open and pointed at a model that never
|
||||||
|
* finished loading is worse than no console at all.
|
||||||
*
|
*
|
||||||
* `_watchLlamaSwapGeneration` guards against two overlapping calls (a second launch
|
* `_watchLlamaSwapGeneration` guards against two overlapping calls (a second launch
|
||||||
* started before the first one's loop finished) clobbering each other's banner:
|
* started before the first one's loop finished) clobbering each other's banner:
|
||||||
* `_showCenterStatus` reuses one shared DOM node, so an older loop's `dismiss()`/message
|
* `_showCenterStatus` reuses one shared DOM node, so an older loop's `dismiss()`/message
|
||||||
* update firing after a newer one has already taken over the banner would otherwise hide
|
* update firing after a newer one has already taken over the banner would otherwise hide
|
||||||
* or overwrite the WRONG one. Each call claims the counter as its own "generation" and
|
* or overwrite the WRONG one, or close the WRONG session. Each call claims the counter
|
||||||
* checks it still owns it before touching the banner.
|
* as its own "generation" and checks it still owns it before touching either.
|
||||||
*
|
*
|
||||||
* `pollIntervalMs`/`maxWaitMs` exist to let a test drive this in milliseconds instead of
|
* `pollIntervalMs`/`maxWaitMs` exist to let a test drive this in milliseconds instead of
|
||||||
* minutes — real callers never pass `maxWaitMs`, which is what keeps the size-scaled
|
* minutes — real callers never pass `maxWaitMs`, which is what keeps the size-scaled
|
||||||
* default live here rather than only in a test fixture.
|
* default live here rather than only in a test fixture.
|
||||||
*/
|
*/
|
||||||
async _watchLlamaSwapLoading(endpointId, modelId, pollIntervalMs = 1000, maxWaitMs) {
|
async _watchLlamaSwapLoading(endpointId, modelId, sessionId, pollIntervalMs = 1000, maxWaitMs) {
|
||||||
const generation = (this._watchLlamaSwapGeneration = (this._watchLlamaSwapGeneration || 0) + 1);
|
const generation = (this._watchLlamaSwapGeneration = (this._watchLlamaSwapGeneration || 0) + 1);
|
||||||
const isCurrent = () => this._watchLlamaSwapGeneration === generation;
|
const isCurrent = () => this._watchLlamaSwapGeneration === generation;
|
||||||
const sizeGB = await this._lookupModelSizeGB(endpointId, modelId);
|
const sizeGB = await this._lookupModelSizeGB(endpointId, modelId);
|
||||||
@@ -999,10 +1013,11 @@ Object.assign(CodemanApp.prototype, {
|
|||||||
const sizeSuffix = sizeGB
|
const sizeSuffix = sizeGB
|
||||||
? ` (${sizeGB.toFixed(1)} GB${estimate ? `, typically ${estimate.label}` : ''})`
|
? ` (${sizeGB.toFixed(1)} GB${estimate ? `, typically ${estimate.label}` : ''})`
|
||||||
: '';
|
: '';
|
||||||
|
const baseMessage = `Loading ${modelId}${sizeSuffix} on ${endpointId} —`;
|
||||||
// Prominent and screen-centred, not a corner toast — a real llama-swap model load can
|
// Prominent and screen-centred, not a corner toast — a real llama-swap model load can
|
||||||
// sit on screen for well over a minute, easy to mistake for nothing happening there.
|
// sit on screen for well over a minute, easy to mistake for nothing happening there.
|
||||||
const toast = this._showCenterStatus(`Loading ${modelId}${sizeSuffix} on ${endpointId}… this can take a while`);
|
|
||||||
const deadline = Date.now() + effectiveMaxWaitMs;
|
const deadline = Date.now() + effectiveMaxWaitMs;
|
||||||
|
const toast = this._showCenterStatus(`${baseMessage} ${this._formatRemaining(deadline - Date.now())}`);
|
||||||
while (Date.now() < deadline) {
|
while (Date.now() < deadline) {
|
||||||
const status = await this._apiJson(`/api/model-endpoints/${encodeURIComponent(endpointId)}/running-status`);
|
const status = await this._apiJson(`/api/model-endpoints/${encodeURIComponent(endpointId)}/running-status`);
|
||||||
if (!isCurrent()) return; // a newer launch took over the banner — this loop is done
|
if (!isCurrent()) return; // a newer launch took over the banner — this loop is done
|
||||||
@@ -1018,11 +1033,24 @@ Object.assign(CodemanApp.prototype, {
|
|||||||
this.showToast(`${modelId} is ready`, 'success', { duration: 2500 });
|
this.showToast(`${modelId} is ready`, 'success', { duration: 2500 });
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
if (!isCurrent()) return;
|
||||||
|
toast?.setMessage(`${baseMessage} ${this._formatRemaining(deadline - Date.now())}`);
|
||||||
await new Promise((resolve) => setTimeout(resolve, pollIntervalMs));
|
await new Promise((resolve) => setTimeout(resolve, pollIntervalMs));
|
||||||
}
|
}
|
||||||
if (!isCurrent()) return;
|
if (!isCurrent()) return;
|
||||||
toast?.dismiss();
|
this._showCenterStatus(
|
||||||
this.showToast(`Still waiting for ${modelId} to finish loading on ${endpointId} — check the llama-swap server`, 'warning');
|
`${modelId} did not finish loading on ${endpointId} within the expected time. ` +
|
||||||
|
`Check the llama-swap server logs for details.` +
|
||||||
|
(sessionId ? ' The session has been closed.' : ''),
|
||||||
|
{ type: 'error' }
|
||||||
|
);
|
||||||
|
if (sessionId) {
|
||||||
|
try {
|
||||||
|
await this.closeSession(sessionId);
|
||||||
|
} catch {
|
||||||
|
// closeSession already reports its own failure via toast — nothing more to do here
|
||||||
|
}
|
||||||
|
}
|
||||||
},
|
},
|
||||||
|
|
||||||
/**
|
/**
|
||||||
|
|||||||
@@ -8577,11 +8577,37 @@ kbd {
|
|||||||
}
|
}
|
||||||
|
|
||||||
.center-status-text {
|
.center-status-text {
|
||||||
|
flex: 1;
|
||||||
pointer-events: auto;
|
pointer-events: auto;
|
||||||
white-space: pre-wrap;
|
white-space: pre-wrap;
|
||||||
word-break: break-word;
|
word-break: break-word;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* Error variant: the load didn't finish in time — nothing is "in progress" anymore (no
|
||||||
|
spinner), and since this one doesn't dismiss itself, it needs a close button the user
|
||||||
|
can actually click, so pointer-events is restored here too (see the wrapper's own
|
||||||
|
comment on why that's `none` by default). */
|
||||||
|
.center-status-error {
|
||||||
|
border-color: rgba(239, 68, 68, 0.5);
|
||||||
|
}
|
||||||
|
|
||||||
|
.center-status-close {
|
||||||
|
flex-shrink: 0;
|
||||||
|
pointer-events: auto;
|
||||||
|
background: none;
|
||||||
|
border: none;
|
||||||
|
color: inherit;
|
||||||
|
opacity: 0.6;
|
||||||
|
font-size: 1.2rem;
|
||||||
|
line-height: 1;
|
||||||
|
padding: 0 0.15rem;
|
||||||
|
cursor: pointer;
|
||||||
|
}
|
||||||
|
|
||||||
|
.center-status-close:hover {
|
||||||
|
opacity: 1;
|
||||||
|
}
|
||||||
|
|
||||||
.toast-success { border-color: rgba(34, 197, 94, 0.4); }
|
.toast-success { border-color: rgba(34, 197, 94, 0.4); }
|
||||||
.toast-error { border-color: rgba(239, 68, 68, 0.4); }
|
.toast-error { border-color: rgba(239, 68, 68, 0.4); }
|
||||||
.toast-warning { border-color: rgba(234, 179, 8, 0.4); }
|
.toast-warning { border-color: rgba(234, 179, 8, 0.4); }
|
||||||
|
|||||||
@@ -108,17 +108,17 @@ describe('_runCustomModelEntryOneShot', () => {
|
|||||||
expect(app._pendingCustomModelForLaunch).toBeUndefined();
|
expect(app._pendingCustomModelForLaunch).toBeUndefined();
|
||||||
});
|
});
|
||||||
|
|
||||||
it('starts the loading watcher when the launch reports modelSwapInProgress', async () => {
|
it('starts the loading watcher when the launch reports modelSwapInProgress, passing the new session id', async () => {
|
||||||
const { app } = bootApp();
|
const { app } = bootApp();
|
||||||
app.run = async () => {
|
app.run = async () => {
|
||||||
app._lastCustomModelLaunchResult = { modelSwapInProgress: true };
|
app._lastCustomModelLaunchResult = { modelSwapInProgress: true, sessionId: 'new-session' };
|
||||||
};
|
};
|
||||||
let watched: unknown[] | null = null;
|
let watched: unknown[] | null = null;
|
||||||
app._watchLlamaSwapLoading = async (...args: unknown[]) => {
|
app._watchLlamaSwapLoading = async (...args: unknown[]) => {
|
||||||
watched = args;
|
watched = args;
|
||||||
};
|
};
|
||||||
await app._runCustomModelEntryOneShot('codex', 'llama-box', 'qwen3');
|
await app._runCustomModelEntryOneShot('codex', 'llama-box', 'qwen3');
|
||||||
expect(watched).toEqual(['llama-box', 'qwen3']);
|
expect(watched).toEqual(['llama-box', 'qwen3', 'new-session']);
|
||||||
});
|
});
|
||||||
|
|
||||||
it('never starts the watcher when no swap was needed', async () => {
|
it('never starts the watcher when no swap was needed', async () => {
|
||||||
|
|||||||
@@ -577,7 +577,7 @@ describe('Custom Model Endpoint Profiles: llama-swap model-swap confirmation and
|
|||||||
|
|
||||||
await app.runCustomModelEntry('claude', 'llama-box', 'qwen3');
|
await app.runCustomModelEntry('claude', 'llama-box', 'qwen3');
|
||||||
|
|
||||||
expect(watched).toEqual(['llama-box', 'qwen3']);
|
expect(watched).toEqual(['llama-box', 'qwen3', 'new-session']);
|
||||||
});
|
});
|
||||||
|
|
||||||
it('a successful apply with no swap needed never starts the loading watcher', async () => {
|
it('a successful apply with no swap needed never starts the loading watcher', async () => {
|
||||||
@@ -616,25 +616,48 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
|
|||||||
};
|
};
|
||||||
app._apiJson = async () => ({ isLlamaSwap: true, running: [{ model: 'qwen3', state: 'ready' }] });
|
app._apiJson = async () => ({ isLlamaSwap: true, running: [{ model: 'qwen3', state: 'ready' }] });
|
||||||
|
|
||||||
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 5, 200);
|
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 200);
|
||||||
|
|
||||||
expect(bannerMessages[0]).toMatch(/loading qwen3/i);
|
expect(bannerMessages[0]).toMatch(/loading qwen3/i);
|
||||||
expect(dismissed).toContain(bannerMessages[0]);
|
expect(dismissed).toContain(bannerMessages[0]);
|
||||||
expect(toastCalls.at(-1)).toMatch(/ready/i);
|
expect(toastCalls.at(-1)).toMatch(/ready/i);
|
||||||
});
|
});
|
||||||
|
|
||||||
it('gives up after the bounded wait and warns instead of polling forever', async () => {
|
it('gives up after the bounded wait, turns the banner into a sticky error, and closes the session', async () => {
|
||||||
const { app } = bootApp({});
|
const { app } = bootApp({});
|
||||||
app._showCenterStatus = () => ({ dismiss: () => {}, setMessage: () => {} });
|
const banners: Array<{ message: string; opts: unknown }> = [];
|
||||||
const toastCalls: string[] = [];
|
app._showCenterStatus = (message: string, opts: unknown) => {
|
||||||
app.showToast = (message: string) => {
|
banners.push({ message, opts });
|
||||||
toastCalls.push(message);
|
return { dismiss: () => {}, setMessage: () => {} };
|
||||||
|
};
|
||||||
|
app.showToast = () => {};
|
||||||
|
let closedSessionId: string | undefined;
|
||||||
|
app.closeSession = async (id: string) => {
|
||||||
|
closedSessionId = id;
|
||||||
};
|
};
|
||||||
app._apiJson = async () => ({ isLlamaSwap: true, running: [{ model: 'something-else', state: 'ready' }] });
|
app._apiJson = async () => ({ isLlamaSwap: true, running: [{ model: 'something-else', state: 'ready' }] });
|
||||||
|
|
||||||
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 5, 30);
|
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 30);
|
||||||
|
|
||||||
expect(toastCalls.at(-1)).toMatch(/still waiting/i);
|
const errorBanner = banners.find((b) => (b.opts as { type?: string } | undefined)?.type === 'error');
|
||||||
|
expect(errorBanner?.message).toMatch(/did not finish loading/i);
|
||||||
|
expect(errorBanner?.message).toMatch(/llama-swap server logs/i);
|
||||||
|
expect(closedSessionId).toBe('sess-1');
|
||||||
|
});
|
||||||
|
|
||||||
|
it('never closes anything when no sessionId was given (a caller that has none to close)', async () => {
|
||||||
|
const { app } = bootApp({});
|
||||||
|
app._showCenterStatus = () => ({ dismiss: () => {}, setMessage: () => {} });
|
||||||
|
app.showToast = () => {};
|
||||||
|
let closeCalled = false;
|
||||||
|
app.closeSession = async () => {
|
||||||
|
closeCalled = true;
|
||||||
|
};
|
||||||
|
app._apiJson = async () => ({ isLlamaSwap: true, running: [{ model: 'something-else', state: 'ready' }] });
|
||||||
|
|
||||||
|
await app._watchLlamaSwapLoading('llama-box', 'qwen3', undefined, 5, 30);
|
||||||
|
|
||||||
|
expect(closeCalled).toBe(false);
|
||||||
});
|
});
|
||||||
|
|
||||||
it('stops polling (without a warning) once the endpoint no longer reads as llama-swap', async () => {
|
it('stops polling (without a warning) once the endpoint no longer reads as llama-swap', async () => {
|
||||||
@@ -652,7 +675,7 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
|
|||||||
};
|
};
|
||||||
app._apiJson = async () => ({ isLlamaSwap: false, running: [] });
|
app._apiJson = async () => ({ isLlamaSwap: false, running: [] });
|
||||||
|
|
||||||
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 5, 200);
|
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 200);
|
||||||
|
|
||||||
expect(bannerDismissed).toBe(true);
|
expect(bannerDismissed).toBe(true);
|
||||||
expect(toastCalls).toHaveLength(0); // no follow-up warning toast
|
expect(toastCalls).toHaveLength(0); // no follow-up warning toast
|
||||||
@@ -672,7 +695,7 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
|
|||||||
return { isLlamaSwap: true, running: [{ model: 'qwen3', state: 'ready' }] };
|
return { isLlamaSwap: true, running: [{ model: 'qwen3', state: 'ready' }] };
|
||||||
};
|
};
|
||||||
|
|
||||||
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 5, 200);
|
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 5, 200);
|
||||||
|
|
||||||
expect(toastCalls.at(-1)).toMatch(/ready/i);
|
expect(toastCalls.at(-1)).toMatch(/ready/i);
|
||||||
});
|
});
|
||||||
@@ -692,7 +715,7 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
|
|||||||
|
|
||||||
// A huge interval that would time the test out if the function actually waited for
|
// A huge interval that would time the test out if the function actually waited for
|
||||||
// it before the first check.
|
// it before the first check.
|
||||||
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 60000, 300000);
|
await app._watchLlamaSwapLoading('llama-box', 'qwen3', 'sess-1', 60000, 300000);
|
||||||
|
|
||||||
expect(calls).toBe(1);
|
expect(calls).toBe(1);
|
||||||
});
|
});
|
||||||
@@ -706,12 +729,13 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
|
|||||||
});
|
});
|
||||||
app.showToast = () => {};
|
app.showToast = () => {};
|
||||||
// The FIRST call never sees its own target model ready, so left alone it would run all
|
// The FIRST call never sees its own target model ready, so left alone it would run all
|
||||||
// the way to its own timeout and dismiss/warn.
|
// the way to its own timeout and (now) turn into an error + close its session — but no
|
||||||
|
// sessionId is passed, so there is nothing for it to close even if it does get there.
|
||||||
app._apiJson = async (path: string) => {
|
app._apiJson = async (path: string) => {
|
||||||
if (path === '/api/model-endpoints') return [];
|
if (path === '/api/model-endpoints') return [];
|
||||||
return { isLlamaSwap: true, running: [] };
|
return { isLlamaSwap: true, running: [] };
|
||||||
};
|
};
|
||||||
const firstCall = app._watchLlamaSwapLoading('llama-box', 'model-a', 5, 30);
|
const firstCall = app._watchLlamaSwapLoading('llama-box', 'model-a', undefined, 5, 30);
|
||||||
|
|
||||||
// Second call, for a DIFFERENT model that IS ready right away, takes over the banner
|
// Second call, for a DIFFERENT model that IS ready right away, takes over the banner
|
||||||
// before the first call's own bounded wait has elapsed.
|
// before the first call's own bounded wait has elapsed.
|
||||||
@@ -719,15 +743,16 @@ describe('Custom Model Endpoint Profiles: _watchLlamaSwapLoading polling', () =>
|
|||||||
if (path === '/api/model-endpoints') return [];
|
if (path === '/api/model-endpoints') return [];
|
||||||
return { isLlamaSwap: true, running: [{ model: 'model-b', state: 'ready' }] };
|
return { isLlamaSwap: true, running: [{ model: 'model-b', state: 'ready' }] };
|
||||||
};
|
};
|
||||||
await app._watchLlamaSwapLoading('llama-box', 'model-b', 5, 200);
|
await app._watchLlamaSwapLoading('llama-box', 'model-b', undefined, 5, 200);
|
||||||
|
|
||||||
// Let the stale first call run out its own bounded wait and finish.
|
// Let the stale first call run out its own bounded wait and finish.
|
||||||
await firstCall;
|
await firstCall;
|
||||||
|
|
||||||
// Whatever the first call did or didn't show along the way, its own eventual
|
// Whatever the first call did or didn't show along the way, its own eventual
|
||||||
// completion (a timeout, in this case) must never touch a banner state that belongs
|
// completion (a timeout, in this case) must never touch a banner state that belongs
|
||||||
// to the newer, still-current call — exactly one dismiss (model-b's own) is the tell.
|
// to the newer, still-current call — exactly one dismiss, for model-b, is the tell.
|
||||||
expect(dismissCalls).toEqual(['Loading model-b on llama-box… this can take a while']);
|
expect(dismissCalls).toHaveLength(1);
|
||||||
|
expect(dismissCalls[0]).toContain('model-b');
|
||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
@@ -787,10 +812,10 @@ describe('Custom Model Endpoint Profiles: model-size load-time estimate', () =>
|
|||||||
return { isLlamaSwap: true, running: [{ model: 'qwen3.8-27b-ud-q4_k_xl', state: 'ready' }] };
|
return { isLlamaSwap: true, running: [{ model: 'qwen3.8-27b-ud-q4_k_xl', state: 'ready' }] };
|
||||||
};
|
};
|
||||||
|
|
||||||
await app._watchLlamaSwapLoading('llama-box', 'qwen3.8-27b-ud-q4_k_xl', 5);
|
await app._watchLlamaSwapLoading('llama-box', 'qwen3.8-27b-ud-q4_k_xl', undefined, 5);
|
||||||
|
|
||||||
expect(bannerMessages[0]).toBe(
|
expect(bannerMessages[0]).toMatch(
|
||||||
'Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1–3 min) on llama-box… this can take a while'
|
/^Loading qwen3\.8-27b-ud-q4_k_xl \(16\.4 GB, typically ~1–3 min\) on llama-box — .+ remaining$/
|
||||||
);
|
);
|
||||||
});
|
});
|
||||||
|
|
||||||
@@ -807,9 +832,9 @@ describe('Custom Model Endpoint Profiles: model-size load-time estimate', () =>
|
|||||||
return { isLlamaSwap: true, running: [{ model: 'big', state: 'ready' }] };
|
return { isLlamaSwap: true, running: [{ model: 'big', state: 'ready' }] };
|
||||||
};
|
};
|
||||||
|
|
||||||
await app._watchLlamaSwapLoading('llama-box', 'big', 5);
|
await app._watchLlamaSwapLoading('llama-box', 'big', undefined, 5);
|
||||||
|
|
||||||
expect(bannerMessages[0]).toBe('Loading big on llama-box… this can take a while');
|
expect(bannerMessages[0]).toMatch(/^Loading big on llama-box — .+ remaining$/);
|
||||||
});
|
});
|
||||||
|
|
||||||
it('uses the size-scaled estimate as the default timeout when maxWaitMs is not passed', async () => {
|
it('uses the size-scaled estimate as the default timeout when maxWaitMs is not passed', async () => {
|
||||||
@@ -828,7 +853,7 @@ describe('Custom Model Endpoint Profiles: model-size load-time estimate', () =>
|
|||||||
};
|
};
|
||||||
|
|
||||||
// pollIntervalMs only — maxWaitMs omitted, so it must fall back to the size estimate.
|
// pollIntervalMs only — maxWaitMs omitted, so it must fall back to the size estimate.
|
||||||
await app._watchLlamaSwapLoading('llama-box', 'huge', 5);
|
await app._watchLlamaSwapLoading('llama-box', 'huge', undefined, 5);
|
||||||
|
|
||||||
expect(calls).toBe(3);
|
expect(calls).toBe(3);
|
||||||
});
|
});
|
||||||
|
|||||||
Reference in New Issue
Block a user