The loading banner now shows a live countdown against its own timeout
(updated every poll, so every second by default) instead of a static
"this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically
~1-3 min) on llama-swap - 47s remaining".
If the countdown reaches zero and the model still isn't ready, this is now
treated as a real failure rather than a "keep waiting" shrug:
- The banner turns into a sticky error (_showCenterStatus gains a `type`
option - 'error' drops the spinner and adds a close button, since nothing
is "in progress" anymore and a sticky message needs a way to dismiss it),
naming the llama-swap server's own logs as where to look for detail.
- The session that load was for is closed automatically (closeSession) -
requested explicitly: a console left open and pointed at a model that
never finished loading is worse than no console at all. Both apply paths
now thread the new session's id through to _watchLlamaSwapLoading for
this (new required 3rd parameter, after endpointId/modelId).
_watchLlamaSwapGeneration's existing stale-call guard extends naturally to
this: a superseded call's own eventual timeout recognises it no longer owns
the banner and neither shows the error nor closes a session that may by
then belong to a different, newer launch.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG