feat(custom-model): show real-time llama.cpp backend status in the loading banner

Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.

- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
  (custom-model-routes.ts): one persistent GET /api/events (SSE)
  connection held open per endpoint, parsing logData frames and
  keeping the latest source:"upstream" (backend llama-server) line —
  filtering out llama-swap's own source:"proxy" request-access lines.
  Idle-closed after 30s of no polling, same 20s sweep as the existing
  swap-displacement check.
- running-status route now returns logLine alongside the existing
  isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
  ("llama.cpp: <line>", bootlog timestamp/level/component prefix
  stripped for display) that stays on the last real thing llama.cpp
  said rather than clearing to blank between polls.

⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).

12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 12:22:45 +08:00
co-authored by Claude Sonnet 5
parent 5ddc028a2f
commit 2d3fc65758
9 changed files with 520 additions and 11 deletions
+23
View File
@@ -120,6 +120,29 @@ where to look, and the session the load was for is closed automatically —
a console left open and pointed at a model that never finished loading is
worse than no console at all.
**The banner's second line is the real backend log line, not a guess.**
llama-swap's `GET /api/events` SSE stream carries the actual `llama-server`
process's own stdout — `load_model: loading model '<path>'`,
`llama_server: model loaded`, tokenizer warnings, all of it — tagged
`source: "upstream"`, distinct from llama-swap's own `source: "proxy"`
request-access lines. `running-status`'s response now includes `logLine`
(via `getLatestLlamaSwapLogLine`), and the banner shows it under the
countdown, e.g. "llama.cpp: load_model: loading model '...'" — confirmed
live end-to-end through a real forced swap, sequentially showing the model
path, a tokenizer warning, then staying on whatever llama.cpp last printed
once the load goes quiet (never cleared back to blank). ⚠️ **`GET /logs`
— the endpoint this feature's own first cut was built against — turns out
to carry ONLY llama-swap's own proxy request-access log.** Confirmed live
it never showed a single backend line, even seconds after a real, verified
model swap; `/api/events`'s `logData` frames are the only source that
actually has it, and its own `source` field (`upstream` vs `proxy`) is
what `getLatestLlamaSwapLogLine` filters on. One `/api/events` connection
is held open per endpoint and reused across every session watching a load
on it (confirmed live to stay open indefinitely, unlike `/logs`, which
closes after a fixed ~100KB), idle-closed after 30s of nobody polling it
(`pruneIdleLlamaSwapLogTails`, same 20s sweep as the swap-displacement
check below).
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
from the endpoint's own discovered `models`, and the route refuses a value
+8
View File
@@ -89,6 +89,14 @@ error telling you to check the llama-swap server's own logs, and **the session t
for is closed automatically** — a console left open and pointed at a model that never
finished loading would just be confusing to leave sitting there.
**The banner also shows a real, live second line of what llama.cpp itself is doing** — not
a made-up progress phase, the actual next line the `llama-server` process printed, e.g.
"llama.cpp: load_model: loading model '/models/.../Qwen3.8-27B.gguf'" then later
"llama.cpp: llama_server: model loaded". It comes straight from llama-swap's own event
feed, filtered down to just the backend process's own output (not llama-swap's own request
logging), and stays on whatever it last said once the load goes quiet, rather than
clearing back to nothing.
**You'll also be told if a session's model gets swapped out from under it later, not just
at launch.** The conflict warning above only fires at the moment you launch or apply a
model — llama.cpp only runs one model at a time, so if a DIFFERENT session using the same