feat(custom-model): live countdown on the loading banner; timeout is now an error

The loading banner now shows a live countdown against its own timeout
(updated every poll, so every second by default) instead of a static
"this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically
~1-3 min) on llama-swap - 47s remaining".

If the countdown reaches zero and the model still isn't ready, this is now
treated as a real failure rather than a "keep waiting" shrug:
- The banner turns into a sticky error (_showCenterStatus gains a `type`
  option - 'error' drops the spinner and adds a close button, since nothing
  is "in progress" anymore and a sticky message needs a way to dismiss it),
  naming the llama-swap server's own logs as where to look for detail.
- The session that load was for is closed automatically (closeSession) -
  requested explicitly: a console left open and pointed at a model that
  never finished loading is worse than no console at all. Both apply paths
  now thread the new session's id through to _watchLlamaSwapLoading for
  this (new required 3rd parameter, after endpointId/modelId).

_watchLlamaSwapGeneration's existing stale-call guard extends naturally to
this: a superseded call's own eventual timeout recognises it no longer owns
the banner and neither shows the error nor closes a session that may by
then belong to a different, newer launch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 20:22:02 +08:00
co-authored by Claude Sonnet 5
parent 55dae31530
commit 7bbe408e44
7 changed files with 181 additions and 64 deletions
+10
View File
@@ -97,6 +97,16 @@ against any real endpoint's actual hardware/storage) and to scale that same
banner's own give-up timeout for a very large model; never anything a
server-side check relies on.
**The loading banner shows a live countdown against that same timeout, and
treats a real timeout as a failure, not a shrug.** It checks llama-swap's
own `/running` every second (`GET /api/model-endpoints/:id/running-status`)
and counts down against the size-scaled (or flat 5-minute) timeout live; if
the countdown reaches zero with the target model still not ready, the
banner turns into a sticky error naming the llama-swap server's own logs as
where to look, and the session the load was for is closed automatically —
a console left open and pointed at a model that never finished loading is
worse than no console at all.
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
from the endpoint's own discovered `models`, and the route refuses a value
+10 -4
View File
@@ -78,10 +78,16 @@ waiting on your first prompt to do it.** llama-swap has no "switch model" button
— the only thing that starts a swap is a real request naming the model, and confirmed live:
just applying a selection never reached llama-swap's own logs at all until something asked
it to load. Picking an entry now also sends the smallest real request that will trigger
that load, in the background, the moment the target model isn't already loaded and ready —
which is what the prominent **"Loading `<model>`… this can take a while"** banner
(centred on screen, not a corner toast — a real load can take well over a minute) is
actually watching for.
that load, in the background, the moment the target model isn't already loaded and ready.
**The centred loading banner shows a live countdown, and a real timeout is an error, not a
shrug.** When it knows the model's discovered file size (its GB figure, when llama-swap
states one), it shows both a rough expected-time estimate and a live countdown against it —
e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1–3 min) on llama-swap — 47s remaining". If
the countdown reaches zero and the model still isn't ready, the banner turns into a sticky
error telling you to check the llama-swap server's own logs, and **the session that load was
for is closed automatically** — a console left open and pointed at a model that never
finished loading would just be confusing to leave sitting there.
**Claude Code specifically gets two extra fixes applied automatically:**