fix(custom-model): merge-time fixes for the Run-menu picker

Conflict resolution against the five PRs that landed while this was in review, plus
the items left for merge on the thread.

The real one was `session-ui.js`. #454 refactored all eight non-Claude `run*()`
functions to funnel through one `_launchQuickStartInstances()` helper that does the
POST itself, while this PR replaced that same POST in each of them with
`_quickStartWithCustomModelConfirm()`. Resolved in the helper rather than seven times
over: the helper now goes through the confirm path, and each body builder carries the
`customModel` spread. `runAntigravity` deliberately does NOT, since antigravity's
`customModelInjection` is `unsupported`; parity with this PR's own per-mode choices is
asserted rather than assumed.

That merge creates a question neither feature had alone: the confirm dialog now runs
inside a loop that can launch up to 20 instances. Both questions it can ask (context
window too small, and loading this will unload the model another session is using) are
decisions about the ENDPOINT, and every instance in a batch targets the same one, so
the answer is taken once and carried to the rest. Without that a 20-instance launch
asks the same question 20 times.

Also: `sse-events.ts` is 161 constants (master added two for remote wake, this adds
one, verified by counting rather than by arithmetic), `server.ts` keeps both new SSE
prefixes, the two comments pointing at code that no longer exists are corrected, and
CLAUDE.md's SSE and route counts move to 161 / ~236 / custom-model (6).

`pumpLlamaSwapLogTail`'s unparsed remainder is now capped at 64 KiB. It only shrank at
a `\n\n` frame boundary, so a backend that streams without one would grow it for the
life of a deliberately indefinite connection.

NOT changed, deliberately: the context warning and the swap-conflict warning still
share one `confirmed` flag with the context check first, so confirming "launch anyway"
on a too-small context also skips the "this unloads it for another session" ask. That
is the author's documented choice and the reviewer's own note calls it minor. Both
fixes are worse to make here than to defer: separate flags are new wire surface landed
unreviewed during a release, and reordering the checks adds a network round trip to a
path that currently short-circuits. Raised as a follow-up instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-09-19 12:27:01 +02:00
parent 1a99b5836c
commit 3b55957d79
5 changed files with 20 additions and 5 deletions
+1 -1
View File
@@ -68,7 +68,7 @@ export interface CustomModelHost {
* llama.cpp") — a hand-configured profile's own description has no such figure and
* correctly gets no entry, never a guess. Used only to label the Run-menu picker's
* "loading model" banner with a rough, unmeasured expected-time estimate
* (`estimateModelLoad()` in session-ui.js) — never a guarantee, and never anything a
* (the Run-menu picker's loading banner in session-ui.js) — never a guarantee, and never anything a
* server-side check relies on.
*/
modelSizesGB?: Record<string, number>;
+15 -1
View File
@@ -416,6 +416,14 @@ function parseBackendLogDataEvent(dataLine: string): string | undefined {
* line (`\n\n`), buffered the same way `/running`'s NDJSON-shaped siblings buffer partial
* chunks — a frame split across two `reader.read()` calls must not be parsed early.
*/
/**
* Cap on the unparsed remainder held between reads of the backend log stream. One
* SSE frame is a status line, so this is orders of magnitude more than a real frame
* needs; it exists so a server that never emits a frame boundary cannot grow the
* buffer without bound for the life of the connection.
*/
const MAX_LOG_TAIL_BUFFER_CHARS = 64 * 1024;
async function pumpLlamaSwapLogTail(
host: Pick<CustomModelHost, 'id' | 'baseUrl' | 'apiKey' | 'authStyle'>,
entry: LlamaSwapLogTail
@@ -435,6 +443,12 @@ async function pumpLlamaSwapLogTail(
buffer += decoder.decode(value, { stream: true });
const frames = buffer.split('\n\n');
buffer = frames.pop() ?? '';
// The remainder only shrinks at a frame boundary, so a server that streams
// without `\n\n` (or one very long frame) would grow it for as long as the
// connection is held, which is indefinitely by design. Past the cap the
// partial frame cannot become a useful log line anyway, so drop it and
// resynchronise on the next boundary rather than buffering forever.
if (buffer.length > MAX_LOG_TAIL_BUFFER_CHARS) buffer = '';
for (const frame of frames) {
const dataLine = frame.split('\n').find((l) => l.startsWith('data:'));
if (!dataLine) continue;
@@ -807,7 +821,7 @@ export function registerCustomModelRoutes(app: FastifyInstance): void {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
const status = await getLlamaSwapStatus(host);
// Only worth tailing /logs once llama-swap is actually confirmed — a plain
// Only worth tailing /api/events once llama-swap is actually confirmed — a plain
// llama.cpp/OpenAI-compatible server has no such endpoint at all.
const logLine = status.isLlamaSwap ? getLatestLlamaSwapLogLine(host) : undefined;
// `cmd` (the literal llama-server launch line, which can carry model paths and
+1 -1
View File
@@ -2837,7 +2837,7 @@ export class WebServer extends EventEmitter {
.catch((err) => {
console.error('[custom-model] swap-displacement check failed:', getErrorMessage(err));
});
// Same cadence, unrelated concern: close any /logs tail (see
// Same cadence, unrelated concern: close any /api/events tail (see
// getLatestLlamaSwapLogLine) nothing has polled in a while, so a loading banner
// that finished (or was abandoned) doesn't leave a connection open forever.
pruneIdleLlamaSwapLogTails();