feat(custom-model): show real-time llama.cpp backend status in the loading banner

Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.

- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
  (custom-model-routes.ts): one persistent GET /api/events (SSE)
  connection held open per endpoint, parsing logData frames and
  keeping the latest source:"upstream" (backend llama-server) line —
  filtering out llama-swap's own source:"proxy" request-access lines.
  Idle-closed after 30s of no polling, same 20s sweep as the existing
  swap-displacement check.
- running-status route now returns logLine alongside the existing
  isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
  ("llama.cpp: <line>", bootlog timestamp/level/component prefix
  stripped for display) that stays on the last real thing llama.cpp
  said rather than clearing to blank between polls.

⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).

12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 12:22:45 +08:00
co-authored by Claude Sonnet 5
parent 5ddc028a2f
commit 2d3fc65758
9 changed files with 520 additions and 11 deletions
+24 -2
View File
@@ -1043,6 +1043,19 @@ Object.assign(CodemanApp.prototype, {
return mins > 0 ? `${mins}m ${String(secs).padStart(2, '0')}s remaining` : `${secs}s remaining`;
},
/**
* Strips llama.cpp's own bootlog prefix (`<uptime> <I|W|E> <component> `, e.g.
* `0.31.428.568 I srv llama_server: model loaded`) for display, leaving just
* `llama_server: model loaded` — the raw line from the server is kept as-is
* (`GET .../running-status`'s `logLine` field), this trims it only for the loading
* banner's second line. Defensive: a line that doesn't match this shape (a different
* llama.cpp build, or llama-swap's own format changing) is shown verbatim rather than
* mangled or dropped.
*/
_formatLlamaLogLine(line) {
return typeof line === 'string' ? line.replace(/^[\d.]+\s+[IWE]\s+\S+\s+/, '') : line;
},
/**
* Polls llama-swap's own `/running` (via the read-only running-status route) until
* `modelId` reports `state: 'ready'`, showing a sticky banner with a live countdown the
@@ -1082,10 +1095,19 @@ Object.assign(CodemanApp.prototype, {
? ` (${sizeGB.toFixed(1)} GB${estimate ? `, typically ${estimate.label}` : ''})`
: '';
const baseMessage = `Loading ${modelId}${sizeSuffix} on ${endpointId} —`;
// Second line, when llama-swap's /logs actually gives us one: the real backend
// llama-server process's own latest log line (load_model:/llama_server: ..., see
// getLatestLlamaSwapLogLine) — a countdown alone says "something is happening,
// trust me," this says what. Absent on the very first render (no poll has landed
// yet) and whenever the endpoint doesn't expose /logs at all — never fabricated.
const buildMessage = (remainingMs, logLine) => {
const line = this._formatLlamaLogLine(logLine);
return `${baseMessage} ${this._formatRemaining(remainingMs)}` + (line ? `\nllama.cpp: ${line}` : '');
};
// Prominent and screen-centred, not a corner toast — a real llama-swap model load can
// sit on screen for well over a minute, easy to mistake for nothing happening there.
const deadline = Date.now() + effectiveMaxWaitMs;
const toast = this._showCenterStatus(`${baseMessage} ${this._formatRemaining(deadline - Date.now())}`);
const toast = this._showCenterStatus(buildMessage(deadline - Date.now()));
while (Date.now() < deadline) {
const status = await this._apiJson(`/api/model-endpoints/${encodeURIComponent(endpointId)}/running-status`);
if (!isCurrent()) return; // a newer launch took over the banner — this loop is done
@@ -1102,7 +1124,7 @@ Object.assign(CodemanApp.prototype, {
return;
}
if (!isCurrent()) return;
toast?.setMessage(`${baseMessage} ${this._formatRemaining(deadline - Date.now())}`);
toast?.setMessage(buildMessage(deadline - Date.now(), status?.logLine));
await new Promise((resolve) => setTimeout(resolve, pollIntervalMs));
}
if (!isCurrent()) return;