mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout (updated every poll, so every second by default) instead of a static "this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1-3 min) on llama-swap - 47s remaining". If the countdown reaches zero and the model still isn't ready, this is now treated as a real failure rather than a "keep waiting" shrug: - The banner turns into a sticky error (_showCenterStatus gains a `type` option - 'error' drops the spinner and adds a close button, since nothing is "in progress" anymore and a sticky message needs a way to dismiss it), naming the llama-swap server's own logs as where to look for detail. - The session that load was for is closed automatically (closeSession) - requested explicitly: a console left open and pointed at a model that never finished loading is worse than no console at all. Both apply paths now thread the new session's id through to _watchLlamaSwapLoading for this (new required 3rd parameter, after endpointId/modelId). _watchLlamaSwapGeneration's existing stale-call guard extends naturally to this: a superseded call's own eventual timeout recognises it no longer owns the banner and neither shows the error nor closes a session that may by then belong to a different, newer launch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
55dae31530
commit
7bbe408e44
@@ -97,6 +97,16 @@ against any real endpoint's actual hardware/storage) and to scale that same
|
||||
banner's own give-up timeout for a very large model; never anything a
|
||||
server-side check relies on.
|
||||
|
||||
**The loading banner shows a live countdown against that same timeout, and
|
||||
treats a real timeout as a failure, not a shrug.** It checks llama-swap's
|
||||
own `/running` every second (`GET /api/model-endpoints/:id/running-status`)
|
||||
and counts down against the size-scaled (or flat 5-minute) timeout live; if
|
||||
the countdown reaches zero with the target model still not ready, the
|
||||
banner turns into a sticky error naming the llama-swap server's own logs as
|
||||
where to look, and the session the load was for is closed automatically —
|
||||
a console left open and pointed at a model that never finished loading is
|
||||
worse than no console at all.
|
||||
|
||||
`defaultModelId` names which discovered model the picker pre-marks for that
|
||||
endpoint — the settings panel's Edit form exposes it as a select populated
|
||||
from the endpoint's own discovered `models`, and the route refuses a value
|
||||
|
||||
@@ -78,10 +78,16 @@ waiting on your first prompt to do it.** llama-swap has no "switch model" button
|
||||
— the only thing that starts a swap is a real request naming the model, and confirmed live:
|
||||
just applying a selection never reached llama-swap's own logs at all until something asked
|
||||
it to load. Picking an entry now also sends the smallest real request that will trigger
|
||||
that load, in the background, the moment the target model isn't already loaded and ready —
|
||||
which is what the prominent **"Loading `<model>`… this can take a while"** banner
|
||||
(centred on screen, not a corner toast — a real load can take well over a minute) is
|
||||
actually watching for.
|
||||
that load, in the background, the moment the target model isn't already loaded and ready.
|
||||
|
||||
**The centred loading banner shows a live countdown, and a real timeout is an error, not a
|
||||
shrug.** When it knows the model's discovered file size (its GB figure, when llama-swap
|
||||
states one), it shows both a rough expected-time estimate and a live countdown against it —
|
||||
e.g. "Loading qwen3.8-27b (16.4 GB, typically ~1–3 min) on llama-swap — 47s remaining". If
|
||||
the countdown reaches zero and the model still isn't ready, the banner turns into a sticky
|
||||
error telling you to check the llama-swap server's own logs, and **the session that load was
|
||||
for is closed automatically** — a console left open and pointed at a model that never
|
||||
finished loading would just be confusing to leave sitting there.
|
||||
|
||||
**Claude Code specifically gets two extra fixes applied automatically:**
|
||||
|
||||
|
||||
Reference in New Issue
Block a user