feat(custom-model): remove loading-banner countdown, add manual Cancel

Replaces the size-scaled expected-time estimate + matching auto-timeout
with a generic hardware/model-size disclaimer and a user-driven Cancel
button, per explicit request. Real load time depends on hardware this
feature has no way to know (VRAM, storage speed, GPU contention), so
the old estimate/timeout was a guess dressed up as a fact — worse, one
that could kill a genuinely slow load partway through on slower
hardware.

- _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline
  entirely — polls indefinitely until ready or cancelled, no automatic
  give-up. Message is now "Loading <model> (<size>) on <endpoint> —
  this can take a while depending on your hardware and the model
  size.", with the real llama.cpp log line still on its own second
  line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/
  _formatRemaining (dead code once the countdown is gone) —
  _lookupModelSizeGB is kept, the GB figure still shows.
- _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real
  "Cancel" button (distinct from the error-type "×" close button,
  since Cancel has a real consequence) that calls it on click. Caller
  owns what cancelling actually means, same split as the swap-confirm
  modal's promise-resolving buttons.
- Cancelling dismisses the banner, shows an info toast (not an error —
  this was deliberate), and closes the session, mirroring what the old
  timeout used to do automatically but now on the user's own call.
- New .center-status-cancel CSS (bordered pill button, distinct from
  the plain "×" close glyph).

Test changes: removed the now-invalid timeout-auto-close/estimate
tests, added cancel-flow tests (dismiss/toast-type/session-close,
never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests
for the new Cancel button (bootAppWithRealCenterStatus, evaluating
panels-ui.js instead of stubbing _showCenterStatus, since this button
is worth verifying for real rather than just through the stub every
other test in the file uses). Typecheck/lint/frontend-syntax clean;
full suite shows no new regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 13:00:50 +08:00
co-authored by Claude Sonnet 5
parent 2d3fc65758
commit db9729e1fc
7 changed files with 253 additions and 179 deletions
+32 -31
View File
@@ -103,22 +103,23 @@ into `modelSizesGB` — unlike context length, this needs no `/props` probe
(the figure is right there in the `/v1/models` response) and so is populated
for every model regardless of loaded state. A hand-configured profile's own
description has no such figure and correctly gets no entry, never a guess.
Used only to label the Run-menu picker's "loading model" banner with a
rough, UNMEASURED expected-time estimate (`_estimateModelLoad()` in
session-ui.js, based on typical local NVMe/SSD throughput — not benchmarked
against any real endpoint's actual hardware/storage) and to scale that same
banner's own give-up timeout for a very large model; never anything a
server-side check relies on.
Used only to label the Run-menu picker's "loading model" banner (e.g.
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB) on llama-swap..."); never anything
a server-side check relies on.
**The loading banner shows a live countdown against that same timeout, and
treats a real timeout as a failure, not a shrug.** It checks llama-swap's
own `/running` every second (`GET /api/model-endpoints/:id/running-status`)
and counts down against the size-scaled (or flat 5-minute) timeout live; if
the countdown reaches zero with the target model still not ready, the
banner turns into a sticky error naming the llama-swap server's own logs as
where to look, and the session the load was for is closed automatically —
a console left open and pointed at a model that never finished loading is
worse than no console at all.
**The loading banner is unbounded by design, and says so — no countdown, no
automatic give-up.** An earlier version scaled an expected-time estimate and
a timeout off the model's file size and auto-closed the session once that
elapsed, but a real load's actual duration depends on hardware this feature
has no way to know (VRAM, storage speed, whatever else is contending for the
GPU) — any fixed number was a guess dressed up as a fact, and a model that
genuinely takes 10+ minutes on slower hardware would just get killed
mid-load by its own display. The banner now says outright that it can take a
while depending on hardware and model size, polls
`GET /api/model-endpoints/:id/running-status` every second for as long as it
takes, and carries a **Cancel** button (rendered on the banner itself) that
ends the wait and closes the session the load was for — the user's own call
on when it's taking too long, not a fixed number baked into the client.
**The banner's second line is the real backend log line, not a guess.**
llama-swap's `GET /api/events` SSE stream carries the actual `llama-server`
@@ -126,22 +127,22 @@ process's own stdout — `load_model: loading model '<path>'`,
`llama_server: model loaded`, tokenizer warnings, all of it — tagged
`source: "upstream"`, distinct from llama-swap's own `source: "proxy"`
request-access lines. `running-status`'s response now includes `logLine`
(via `getLatestLlamaSwapLogLine`), and the banner shows it under the
countdown, e.g. "llama.cpp: load_model: loading model '...'" — confirmed
live end-to-end through a real forced swap, sequentially showing the model
path, a tokenizer warning, then staying on whatever llama.cpp last printed
once the load goes quiet (never cleared back to blank). ⚠️ **`GET /logs`
— the endpoint this feature's own first cut was built against — turns out
to carry ONLY llama-swap's own proxy request-access log.** Confirmed live
it never showed a single backend line, even seconds after a real, verified
model swap; `/api/events`'s `logData` frames are the only source that
actually has it, and its own `source` field (`upstream` vs `proxy`) is
what `getLatestLlamaSwapLogLine` filters on. One `/api/events` connection
is held open per endpoint and reused across every session watching a load
on it (confirmed live to stay open indefinitely, unlike `/logs`, which
closes after a fixed ~100KB), idle-closed after 30s of nobody polling it
(`pruneIdleLlamaSwapLogTails`, same 20s sweep as the swap-displacement
check below).
(via `getLatestLlamaSwapLogLine`), and the banner shows it on its own line
under the disclaimer, e.g. "llama.cpp: load_model: loading model '...'" —
confirmed live end-to-end through a real forced swap, sequentially showing
the model path, a tokenizer warning, then staying on whatever llama.cpp last
printed once the load goes quiet (never cleared back to blank). ⚠️
**`GET /logs` — the endpoint this feature's own first cut was built
against — turns out to carry ONLY llama-swap's own proxy request-access
log.** Confirmed live it never showed a single backend line, even seconds
after a real, verified model swap; `/api/events`'s `logData` frames are the
only source that actually has it, and its own `source` field (`upstream` vs
`proxy`) is what `getLatestLlamaSwapLogLine` filters on. One `/api/events`
connection is held open per endpoint and reused across every session
watching a load on it (confirmed live to stay open indefinitely, unlike
`/logs`, which closes after a fixed ~100KB), idle-closed after 30s of nobody
polling it (`pruneIdleLlamaSwapLogTails`, same 20s sweep as the
swap-displacement check below).
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
+9 -8
View File
@@ -80,14 +80,15 @@ just applying a selection never reached llama-swap's own logs at all until somet
it to load. Picking an entry now also sends the smallest real request that will trigger
that load, in the background, the moment the target model isn't already loaded and ready.
**The centred loading banner shows a live countdown, and a real timeout is an error, not a
shrug.** When it knows the model's discovered file size (its GB figure, when llama-swap
states one), it shows both a rough expected-time estimate and a live countdown against it —
e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1–3 min) on llama-swap — 47s remaining". If
the countdown reaches zero and the model still isn't ready, the banner turns into a sticky
error telling you to check the llama-swap server's own logs, and **the session that load was
for is closed automatically** — a console left open and pointed at a model that never
finished loading would just be confusing to leave sitting there.
**The centred loading banner has no countdown and no automatic timeout — it waits as long as
it takes, and tells you so.** When it knows the model's discovered file size (its GB figure,
when llama-swap states one) it's shown too, e.g. "Loading qwen3.8-27b (16.4 GB) on
llama-swap — this can take a while depending on your hardware and the model size." An
earlier version tried to estimate and enforce a time limit, but real load time depends on
hardware this feature has no way to know, so a fixed number was always a guess — worse, one
that could kill a genuinely slow load partway through. If it really is taking too long, a
**Cancel** button right on the banner ends the wait and **closes the session that load was
for**, on your own call rather than a guessed deadline.
**The banner also shows a real, live second line of what llama.cpp itself is doing** — not
a made-up progress phase, the actual next line the `llama-server` process printed, e.g.