mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-04 14:39:42 +02:00
fix(custom-model): merge-time fixes for the Run-menu picker
Conflict resolution against the five PRs that landed while this was in review, plus the items left for merge on the thread. The real one was `session-ui.js`. #454 refactored all eight non-Claude `run*()` functions to funnel through one `_launchQuickStartInstances()` helper that does the POST itself, while this PR replaced that same POST in each of them with `_quickStartWithCustomModelConfirm()`. Resolved in the helper rather than seven times over: the helper now goes through the confirm path, and each body builder carries the `customModel` spread. `runAntigravity` deliberately does NOT, since antigravity's `customModelInjection` is `unsupported`; parity with this PR's own per-mode choices is asserted rather than assumed. That merge creates a question neither feature had alone: the confirm dialog now runs inside a loop that can launch up to 20 instances. Both questions it can ask (context window too small, and loading this will unload the model another session is using) are decisions about the ENDPOINT, and every instance in a batch targets the same one, so the answer is taken once and carried to the rest. Without that a 20-instance launch asks the same question 20 times. Also: `sse-events.ts` is 161 constants (master added two for remote wake, this adds one, verified by counting rather than by arithmetic), `server.ts` keeps both new SSE prefixes, the two comments pointing at code that no longer exists are corrected, and CLAUDE.md's SSE and route counts move to 161 / ~236 / custom-model (6). `pumpLlamaSwapLogTail`'s unparsed remainder is now capped at 64 KiB. It only shrank at a `\n\n` frame boundary, so a backend that streams without one would grow it for the life of a deliberately indefinite connection. NOT changed, deliberately: the context warning and the swap-conflict warning still share one `confirmed` flag with the context check first, so confirming "launch anyway" on a too-small context also skips the "this unloads it for another session" ask. That is the author's documented choice and the reviewer's own note calls it minor. Both fixes are worse to make here than to defer: separate flags are new wire surface landed unreviewed during a release, and reordering the checks adds a network round trip to a path that currently short-circuits. Raised as a follow-up instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -416,6 +416,14 @@ function parseBackendLogDataEvent(dataLine: string): string | undefined {
|
||||
* line (`\n\n`), buffered the same way `/running`'s NDJSON-shaped siblings buffer partial
|
||||
* chunks — a frame split across two `reader.read()` calls must not be parsed early.
|
||||
*/
|
||||
/**
|
||||
* Cap on the unparsed remainder held between reads of the backend log stream. One
|
||||
* SSE frame is a status line, so this is orders of magnitude more than a real frame
|
||||
* needs; it exists so a server that never emits a frame boundary cannot grow the
|
||||
* buffer without bound for the life of the connection.
|
||||
*/
|
||||
const MAX_LOG_TAIL_BUFFER_CHARS = 64 * 1024;
|
||||
|
||||
async function pumpLlamaSwapLogTail(
|
||||
host: Pick<CustomModelHost, 'id' | 'baseUrl' | 'apiKey' | 'authStyle'>,
|
||||
entry: LlamaSwapLogTail
|
||||
@@ -435,6 +443,12 @@ async function pumpLlamaSwapLogTail(
|
||||
buffer += decoder.decode(value, { stream: true });
|
||||
const frames = buffer.split('\n\n');
|
||||
buffer = frames.pop() ?? '';
|
||||
// The remainder only shrinks at a frame boundary, so a server that streams
|
||||
// without `\n\n` (or one very long frame) would grow it for as long as the
|
||||
// connection is held, which is indefinitely by design. Past the cap the
|
||||
// partial frame cannot become a useful log line anyway, so drop it and
|
||||
// resynchronise on the next boundary rather than buffering forever.
|
||||
if (buffer.length > MAX_LOG_TAIL_BUFFER_CHARS) buffer = '';
|
||||
for (const frame of frames) {
|
||||
const dataLine = frame.split('\n').find((l) => l.startsWith('data:'));
|
||||
if (!dataLine) continue;
|
||||
@@ -807,7 +821,7 @@ export function registerCustomModelRoutes(app: FastifyInstance): void {
|
||||
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
|
||||
}
|
||||
const status = await getLlamaSwapStatus(host);
|
||||
// Only worth tailing /logs once llama-swap is actually confirmed — a plain
|
||||
// Only worth tailing /api/events once llama-swap is actually confirmed — a plain
|
||||
// llama.cpp/OpenAI-compatible server has no such endpoint at all.
|
||||
const logLine = status.isLlamaSwap ? getLatestLlamaSwapLogLine(host) : undefined;
|
||||
// `cmd` (the literal llama-server launch line, which can carry model paths and
|
||||
|
||||
+1
-1
@@ -2837,7 +2837,7 @@ export class WebServer extends EventEmitter {
|
||||
.catch((err) => {
|
||||
console.error('[custom-model] swap-displacement check failed:', getErrorMessage(err));
|
||||
});
|
||||
// Same cadence, unrelated concern: close any /logs tail (see
|
||||
// Same cadence, unrelated concern: close any /api/events tail (see
|
||||
// getLatestLlamaSwapLogLine) nothing has polled in a while, so a loading banner
|
||||
// that finished (or was abandoned) doesn't leave a connection open forever.
|
||||
pruneIdleLlamaSwapLogTails();
|
||||
|
||||
Reference in New Issue
Block a user