fix(custom-model): split the two confirmation questions, and seed the API key the way claude reads it

Two findings from the review of 5fc391a4, both fixed here rather than sent back.

**The API-key trust seed never matched a real key.** `seedApiKeyTrustFile()` wrote the
key verbatim into `customApiKeyResponses.approved`, but Claude Code stores and compares
only the last 20 characters (`key.trim().slice(-20)`, applied on both the write and the
lookup). For any real key the seed missed, so claude stopped at the interactive
"Detected a custom API key in your environment" prompt, whose default is
"No (recommended)": the launch hangs, or silently refuses the key this feature just
injected and falls through to an OAuth login the isolated config dir does not have. It
survived review because a keyless llama.cpp/llama-swap endpoint uses DEFAULT_API_KEY
('local-dummy-key', 15 chars), where slice(-20) returns the whole string and the seed
matches by accident, and every test used a key shorter than that. Now truncated through
`truncateApiKeyForTrustFile()`, with a test using a 57-character key that also asserts
the full credential never reaches that second file.

**One `confirmed` flag answered two different questions.** The context-floor warning
("this model's window is below what this CLI needs") and the swap-conflict warning
("loading this unloads the model another session is using") shared it, and the context
check runs first, so a user clicking "launch anyway" past the context warning silently
consented to evicting someone else's model. They are about different people, so an
answer to one is not consent to the other. Both routes now read `confirmedContext` and
`confirmedSwap` independently; the legacy `confirmed` still means both, because it
shipped in this feature's HTTP-API-only cut and an existing caller must keep working.
The frontend answers each question with its own flag and accumulates them, on the
one-shot path, the restart path and the batch carry-forward alike.

Also from the same review: the swap-confirm dialog no longer renders " are currently
using ..." when multi-user scoping leaves the affected-session list empty (the swap is
blocked regardless of ownership; only the NAMES are scoped), and the per-endpoint
llama-swap log tails are closed in `WebServer.stop()` instead of only by the idle sweep
whose interval that same teardown disposes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-09-19 12:32:50 +02:00
parent 3b55957d79
commit fe3bd0074c
9 changed files with 244 additions and 32 deletions
+12
View File
@@ -507,6 +507,18 @@ export function getLatestLlamaSwapLogLine(
return entry.latestLine;
}
/**
* Closes EVERY open log tail. The idle sweep above only runs on server.ts's periodic
* interval, and that interval is disposed on shutdown, so without this an outbound
* stream outlives `WebServer.stop()` against CLAUDE.md's "clear Maps in stop()" rule.
* Harmless today only because `cli.ts`'s shutdown handler reaches `process.exit(0)`,
* which is not a property to rely on: tests and any in-process restart do not.
*/
export function closeAllLlamaSwapLogTails(): void {
for (const entry of llamaSwapLogTails.values()) entry.controller.abort();
llamaSwapLogTails.clear();
}
/**
* Closes any log tail nothing has called `getLatestLlamaSwapLogLine` about in
* `LOG_TAIL_IDLE_MS` — a stream nobody is polling is an open connection with nothing to
+1
View File
@@ -32,6 +32,7 @@ export {
registerCustomModelRoutes,
refreshAllCustomModelHosts,
readCustomModelEndpointsEnabled,
closeAllLlamaSwapLogTails,
detectCustomModelSwapDisplacements,
pruneIdleLlamaSwapLogTails,
type CustomModelSessionLike,
+18 -8
View File
@@ -1245,9 +1245,11 @@ export function registerSessionRoutes(
// no matter what CLAUDE_CODE_MAX_CONTEXT_TOKENS says — confirmed live at ~36.4K tokens
// against a model configured with a real 16384-token context. Warn before committing
// to a restart that's certain to fail, rather than letting the user discover it via a
// cryptic 400 from the CLI itself. `confirmed` (already used for the swap-conflict
// warning below) skips this too — the user has already said "launch anyway" once.
if (!body.confirmed && exceedsSafeContextFloor(entry, contextLength)) {
// cryptic 400 from the CLI itself. Answered by `confirmedContext` (or the legacy
// `confirmed`, which still means both) — NOT by `confirmedSwap`: this warning is
// about the caller's own session, and the swap warning below is about someone
// else's, so an answer to one is not consent to the other.
if (!(body.confirmed || body.confirmedContext) && exceedsSafeContextFloor(entry, contextLength)) {
return {
requiresContextWarning: true,
modelId: body.modelId,
@@ -1275,9 +1277,12 @@ export function registerSessionRoutes(
const targetReady = swapStatus.running.some((r) => r.model === body.modelId && r.state === 'ready');
// Only ask when switching would actually take the model away from another session
// that is currently using it — never just because a swap is needed at all. `confirmed`
// (set by the caller after showing that warning once) skips asking again.
if (swapNeeded && !body.confirmed) {
// that is currently using it — never just because a swap is needed at all. Answered
// by `confirmedSwap` (or the legacy `confirmed`). ⚠ It must NOT read
// `confirmedContext`: this check runs second, and while the two shared one flag a
// user who clicked past a too-small-context warning had already, silently, agreed to
// evict another session's model.
if (swapNeeded && !(body.confirmed || body.confirmedSwap)) {
const conflicting = [...ctx.sessions.values()].filter(
(s) =>
s.id !== session.id && s.customModel?.endpointId === endpoint.id && s.customModel?.modelId === currentlyLoaded
@@ -3793,7 +3798,12 @@ export function registerSessionRoutes(
// overhead can exceed a small enough real context on the very first message,
// regardless of contextLengthVar. Warn before creating a session that's certain to
// fail immediately.
if (!customModel.confirmed && exceedsSafeContextFloor(cmEntry, cmContextLength)) {
// See the dedicated route above for why this reads `confirmedContext` and never
// `confirmedSwap`.
if (
!(customModel.confirmed || customModel.confirmedContext) &&
exceedsSafeContextFloor(cmEntry, cmContextLength)
) {
return {
requiresContextWarning: true,
modelId: customModel.modelId,
@@ -3816,7 +3826,7 @@ export function registerSessionRoutes(
// is loaded at all yet. Drives the actual load trigger below.
const cmTargetReady = cmSwapStatus.running.some((r) => r.model === customModel.modelId && r.state === 'ready');
qsCustomModelSwapInProgress = cmSwapStatus.isLlamaSwap && !cmTargetReady;
if (cmSwapNeeded && !customModel.confirmed) {
if (cmSwapNeeded && !(customModel.confirmed || customModel.confirmedSwap)) {
const cmConflicting = [...ctx.sessions.values()].filter(
(s) => s.customModel?.endpointId === cmEndpoint.id && s.customModel?.modelId === cmCurrentlyLoaded
);