feat(custom-model): live countdown on the loading banner; timeout is now an error

The loading banner now shows a live countdown against its own timeout
(updated every poll, so every second by default) instead of a static
"this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically
~1-3 min) on llama-swap - 47s remaining".

If the countdown reaches zero and the model still isn't ready, this is now
treated as a real failure rather than a "keep waiting" shrug:
- The banner turns into a sticky error (_showCenterStatus gains a `type`
  option - 'error' drops the spinner and adds a close button, since nothing
  is "in progress" anymore and a sticky message needs a way to dismiss it),
  naming the llama-swap server's own logs as where to look for detail.
- The session that load was for is closed automatically (closeSession) -
  requested explicitly: a console left open and pointed at a model that
  never finished loading is worse than no console at all. Both apply paths
  now thread the new session's id through to _watchLlamaSwapLoading for
  this (new required 3rd parameter, after endpointId/modelId).

_watchLlamaSwapGeneration's existing stale-call guard extends naturally to
this: a superseded call's own eventual timeout recognises it no longer owns
the banner and neither shows the error nor closes a session that may by
then belong to a different, newer launch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 20:22:02 +08:00
co-authored by Claude Sonnet 5
parent 55dae31530
commit 7bbe408e44
7 changed files with 181 additions and 64 deletions
+10
View File
@@ -97,6 +97,16 @@ against any real endpoint's actual hardware/storage) and to scale that same
banner's own give-up timeout for a very large model; never anything a
server-side check relies on.
**The loading banner shows a live countdown against that same timeout, and
treats a real timeout as a failure, not a shrug.** It checks llama-swap's
own `/running` every second (`GET /api/model-endpoints/:id/running-status`)
and counts down against the size-scaled (or flat 5-minute) timeout live; if
the countdown reaches zero with the target model still not ready, the
banner turns into a sticky error naming the llama-swap server's own logs as
where to look, and the session the load was for is closed automatically —
a console left open and pointed at a model that never finished loading is
worse than no console at all.
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
from the endpoint's own discovered `models`, and the route refuses a value