feat(custom-model): warn before launching Claude on a model too small for its own overhead

Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.

- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
  registry entry declares contextLengthVar (currently only claude) and
  the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
  (40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
  quick-start customModel path) check this before the swap-conflict
  check and before launching/restarting anything, returning
  {requiresContextWarning, modelId, contextLength, minSafeContextTokens}
  — skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
  _resolveContextWarningConfirm (session-ui.js), wired into both
  _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
  (the path Claude actually uses) ahead of the swap-confirmation check.
  Explains the fix in-modal: give the model an explicit larger -c/
  --ctx-size in llama-swap instead of relying on --fit-ctx, which
  optimizes for the biggest model that fits rather than the biggest
  context.

Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 07:36:57 +08:00
co-authored by Claude Sonnet 5
parent 993710263d
commit b45a96358e
10 changed files with 500 additions and 20 deletions
+5 -4
View File
@@ -1,5 +1,5 @@
---
"aicodeman": minor
'aicodeman': minor
---
**Custom model endpoints: Run-menu picker, and hardening from real llama-swap validation** (#430, follow-up to #393's HTTP-API-only cut). With **Custom model endpoints** on (App Settings → Models) and at least one saved endpoint carrying a discovered model, the Run dropdown grows a **Custom Endpoints** section generated live off the CLI registry's own `capabilities.customModelInjection` — one entry per (harness that can redirect to a custom endpoint, saved endpoint). Picking one launches that harness and applies the endpoint to it; with two or more discovered models a small, scrollable dialog asks which one first, the endpoint's `defaultModelId` marked but never auto-chosen. Endpoints also now re-discover themselves automatically every 5 minutes in the background, one unreachable endpoint never blocking the others.
@@ -8,10 +8,11 @@ Everything below was found and fixed against a **real llama-swap server**, not j
- **Session-busy false refusal.** A freshly launched CLI reports itself `busy` for its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying.
- **Errors and confirmations you can actually read.** Toasts now default to sticky with a close button (errors always were meant to stay, but a fixed 3s timer silently hid them); a failed apply's real server-side reason (not a generic message) reaches the toast.
- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on *every* launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would.
- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length from llama.cpp/llama-swap's `GET /props?model=`, but **only** for a model llama-swap's own `/v1/models` response already reports loaded — never an unloaded one, since asking about one risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery — and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS`.
- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on _every_ launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would.
- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS` — sourced primarily from llama-swap's own `GET /running`, whose `cmd` field carries the launch flags (`--fit-ctx`/`-c`/`--ctx-size`) actually in effect, since `GET /props`'s `n_ctx` was confirmed live to report the model's theoretical/trained maximum rather than the real `--fit-ctx`-shrunk runtime context (a 154112-vs-16384 discrepancy, caught only because the fixed value still overflowed) — `/props` is now a fallback for a plain llama.cpp server with no `/running` at all.
- **Context floor too small for Claude Code to even start.** Fixing the overflow above surfaced a second, unfixable-by-injection failure: Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on their own (confirmed live via an `in:0 out:0` failure on the very first message), which can exceed a small model's entire real context before any conversation history exists to trim — no `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since it only governs when history gets compacted. Applying such a model now returns a warning (gated on the CLI registry declaring a `contextLengthVar`, so it's a no-op for every other harness) instead of launching straight into a guaranteed first-message failure, and the Run-menu picker shows it as an in-app dialog naming the model, its discovered context and the ~40K safe floor, with the actual fix spelled out: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on auto-fit, which optimizes for the biggest model that fits rather than the biggest context. "Launch anyway" is still one click away.
- **The real root cause of "it still says opus, not my model."** llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on demand, which can take anywhere from a few seconds to well over a minute — long enough that a session mid-swap is indistinguishable from one that never left the native backend. Applying a selection now checks llama-swap's own `GET /running` first (feature-detected; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is never checked); if switching would unload a model **another live session is actively using**, the apply is refused with a warning naming that session instead of silently switching, and a confirmation retry proceeds anyway. Either way, a sticky "loading model…" toast now covers the actual swap window until llama-swap reports the target model ready, so a prompt sent mid-swap reads as "loading," never as silence or an answer from whatever was loaded a moment before.
Remote (SSH) and Docker sessions are refused for now (400) — their restart reattaches the durable remote/in-container tmux rather than relaunching the agent.
**One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all.** Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). `POST /api/quick-start` now accepts a `customModel` field and computes the same injection *before* the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own `--resume`-based restart is far less jarring, and `runClaude()`'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work).
**One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all.** Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). `POST /api/quick-start` now accepts a `customModel` field and computes the same injection _before_ the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own `--resume`-based restart is far less jarring, and `runClaude()`'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work).
+52 -5
View File
@@ -239,8 +239,17 @@ registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
model id and never compacts, which reliably overflows a much smaller real
local context — confirmed live: a stock ~33.7K-token system prompt against
a 16384-token llama-swap model failed with `exceeds the available context
size`. No entry for the model in `modelContextLengths` means the var is
simply omitted, never a guess.
size`. No entry for the model in `modelContextLengths` means the var is
simply omitted, never a guess. ⚠️ **This var only affects when Claude
Code compacts conversation _history_ — it cannot fix a model whose real
context is smaller than Claude Code's own fixed per-turn overhead**
(system prompt + tool schemas, empirically ~36.4K tokens, confirmed live
via an `in:0 out:0` failure on the very first message, before any
history exists to compact). No context-length declaration changes that
fixed overhead, so a model below the safe floor fails outright on
message one regardless of what this var says. See "Context-window floor
warning" below for how Codeman catches this case before launching
instead of after.
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
the `configDir`-kind CLIs use (empty, no files written into it), so the
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
@@ -261,7 +270,7 @@ registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
**That isolated directory needed one more fix to actually be usable
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
real profile's prior "Detected a custom API key — use it?" approvals, so
without more, Claude Code stops and asks that on *every single launch* —
without more, Claude Code stops and asks that on _every single launch_ —
confirmed live, and with nobody at a TTY to answer, its own default answer
("No") silently refuses the very key this feature just injected, which
looks like the endpoint being ignored entirely. `customModelInjection`'s
@@ -285,7 +294,7 @@ which can take anywhere from a few seconds to well over a minute:
one-shot `POST /api/quick-start` above) call llama-swap's own
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
compatible server (no such endpoint) is simply never checked. If a
*different* model is currently loaded and ready, and another **live
_different_ model is currently loaded and ready, and another **live
session's own selection** is using it, the apply returns
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
instead of silently switching — nothing is applied or created yet.
@@ -298,13 +307,51 @@ which can take anywhere from a few seconds to well over a minute:
reached llama-swap at all (nothing in its own server logs), since nothing
had actually asked it to load anything yet. Both apply routes now also
send the smallest real request that will — `POST <baseUrl>/v1/chat/
completions` with `max_tokens: 1` and one throwaway message — whenever the
completions` with `max_tokens: 1` and one throwaway message — whenever the
target model isn't already the one loaded and ready, fire-and-forget (its
response is never read; `GET /api/model-endpoints/:id/running-status`,
polled client-side, is what actually confirms readiness). The response
also carries `modelSwapInProgress: true` in that case, which is what
drives the Run-menu picker's own "loading model" status banner.
## Context-window floor warning
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
empirically ~36.4K tokens) can exceed a small local model's _entire_ real
context on its own, before any conversation history exists to fill it —
confirmed live twice, both as an `in:0 out:0` failure on the very first
message sent. `CLAUDE_CODE_MAX_CONTEXT_TOKENS` (above) cannot fix this: it
only governs when Claude Code compacts conversation history, and there is
no history yet on message one. Applying such a model would look like the
endpoint being ignored, or the wrong model being used, when in fact the
endpoint applied correctly and the model is simply too small for this CLI.
Both apply routes (the restart route and the one-shot `POST
/api/quick-start`) now check for this **before** launching or restarting
anything, gated on the CLI's registry entry declaring a `contextLengthVar`
(currently only claude — the check is a no-op for every other CLI by
construction, never a hardcoded mode check). If the model's discovered
context (`modelContextLengths`, from discovery above) is below
`CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000, comfortably above the measured
~36.4K overhead), the response is `{requiresContextWarning: true, modelId,
contextLength, minSafeContextTokens}` instead of applying — nothing is
restarted or created yet. A context length that was never discovered at
all skips the check entirely (nothing to compare, so it fails open rather
than warning on every model an endpoint hasn't reported a size for).
Retrying with `confirmed: true` launches anyway.
The Run-menu picker shows this as an in-app modal
(`#customModelContextWarningModal`, matching the llama-swap conflict
modal's look) naming the model, its discovered context, and the safe
floor, and explaining the fix: reconfigure llama-swap to give that model
(or a smaller one) an explicit larger context instead of relying on
auto-fit (`--fit-ctx`), which optimizes for the biggest _model_ that fits
rather than the biggest _context_ — e.g. adding `-c 65536` (or as large a
`--ctx-size` as the hardware holds) to that model's llama-swap config
entry. A smaller model at a much larger explicit context often fits in
the same VRAM a bigger model's auto-fit context gets shrunk to make room
for.
Clear back to the harness's native cloud default with:
```bash
+15 -2
View File
@@ -28,7 +28,7 @@ shows up without another manual click of **Discover**. One endpoint being unreac
given cycle (powered off, wrong network) never blocks the others from refreshing.
**Context length is picked up automatically where it can be, safely.** Against a
llama.cpp/llama-swap server, discovery also learns each *currently loaded* model's real
llama.cpp/llama-swap server, discovery also learns each _currently loaded_ model's real
context window and applies it to the launched session (Claude Code today — see below), so
the harness stops assuming a large default window for a model name it doesn't recognise and
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
@@ -104,10 +104,23 @@ finished loading would just be confusing to leave sitting there.
"Detected a custom API key" prompt once would — without it, that prompt would otherwise
reappear on every single launch with nobody there to answer it.
**If a model's real context is too small for Claude Code to even get started, you get a
warning instead of a confusing failure.** Claude Code's own system prompt and tools take up
roughly 40K tokens on their own, before you've typed anything — a small local model with a
smaller real context than that fails outright on the very first message, no matter what
context size Codeman tells it to expect (raising the declared context only changes when
Claude Code trims _conversation history_, and there is none yet on message one). Picking
such a model now shows an in-app dialog naming the model, its discovered context and what's
needed, before anything launches or restarts, with the fix spelled out: reconfigure
llama-swap to give that model (or a smaller one) an explicit larger context instead of
relying on auto-fit (`--fit-ctx`), which sizes the context around fitting the biggest model
rather than the biggest context — for example adding `-c 65536` to that model's llama-swap
entry. "Launch anyway" is still there if you want to try regardless.
## Which harnesses actually work
| Harness | Status |
| ------- | ------ |
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. |
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
+23
View File
@@ -948,6 +948,29 @@
</div>
</div>
<!-- Custom Model Endpoint Profiles: context-window-too-small warning
(docs/custom-model-endpoints-plan.md) — shown before launching a CLI
whose own fixed system-prompt/tool-schema overhead exceeds the
model's real discovered context, which guarantees a first-message
failure regardless of CLAUDE_CODE_MAX_CONTEXT_TOKENS. See
_confirmContextWarning() in session-ui.js. -->
<div class="modal" id="customModelContextWarningModal">
<div class="modal-backdrop" onclick="app._resolveContextWarningConfirm(false)"></div>
<div class="modal-content modal-sm">
<div class="modal-header">
<h3>Context window too small</h3>
<button class="modal-close" onclick="app._resolveContextWarningConfirm(false)" aria-label="Cancel">&times;</button>
</div>
<div class="modal-body">
<p class="form-hint" id="customModelContextWarningMessage" style="white-space: pre-wrap;"></p>
</div>
<div class="modal-footer">
<button class="btn-toolbar" onclick="app._resolveContextWarningConfirm(false)">Cancel</button>
<button class="btn-toolbar btn-primary" onclick="app._resolveContextWarningConfirm(true)">Launch anyway</button>
</div>
</div>
</div>
<!-- Cron Jobs Modal -->
<div class="modal" id="cronModal">
<div class="modal-backdrop" onclick="app.closeCron()"></div>
+68
View File
@@ -693,6 +693,45 @@ Object.assign(CodemanApp.prototype, {
resolve?.(proceed);
},
/**
* In-app warning shown when the apply route reports `requiresContextWarning`: this
* model's real discovered context is smaller than the CLI's own fixed system-prompt/
* tool-schema overhead, which guarantees the very first message fails outright — no
* `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since there is no conversation
* history yet for compaction to trim. Same promise-based pattern as
* `_confirmModelSwap`; `_resolveContextWarningConfirm` settles it.
*/
_confirmContextWarning(modelId, contextLength, minSafeContextTokens) {
const modal = document.getElementById('customModelContextWarningModal');
const messageEl = document.getElementById('customModelContextWarningMessage');
if (messageEl) {
const known = typeof contextLength === 'number';
messageEl.textContent =
`${modelId} is configured with ` +
(known ? `only ${contextLength.toLocaleString()} tokens of` : 'an unknown (too small)') +
` context, but this CLI needs roughly ${minSafeContextTokens.toLocaleString()}+ tokens just for its own ` +
`system prompt and tools — before any conversation history. Its very first message will fail outright, ` +
`no matter what context size Codeman tells it to expect.\n\n` +
`To fix this, reconfigure llama-swap to give this model (or a smaller one) an explicit larger context ` +
`instead of relying on auto-fit (--fit-ctx), which optimizes for the biggest MODEL that fits, not the ` +
`biggest CONTEXT — e.g. add "-c 65536" (or as large a --ctx-size as your hardware holds) to its llama-swap ` +
`config entry. A smaller model at a much larger explicit context often fits in the same VRAM a bigger ` +
`model's auto-fit context gets shrunk to make room for.`;
}
modal?.classList.add('active');
return new Promise((resolve) => {
this._resolveContextWarningConfirmPromise = resolve;
});
},
/** Called by the modal's Cancel/Launch-anyway buttons and its backdrop click. */
_resolveContextWarningConfirm(proceed) {
document.getElementById('customModelContextWarningModal')?.classList.remove('active');
const resolve = this._resolveContextWarningConfirmPromise;
this._resolveContextWarningConfirmPromise = null;
resolve?.(proceed);
},
/** A model row in the picker modal was clicked: close it and launch with that choice. */
chooseCustomModelAndRun(modelId) {
const pending = this._pendingCustomModelPick;
@@ -789,6 +828,15 @@ Object.assign(CodemanApp.prototype, {
return res.json();
};
let data = await post(bodyObj);
if (data?.data?.requiresContextWarning) {
const { modelId, contextLength, minSafeContextTokens } = data.data;
const proceed = await this._confirmContextWarning(modelId, contextLength, minSafeContextTokens);
if (!proceed) {
this._lastCustomModelLaunchResult = undefined;
return { success: false, error: 'Launch cancelled — context window too small' };
}
data = await post({ ...bodyObj, customModel: { ...bodyObj.customModel, confirmed: true } });
}
if (data?.data?.requiresConfirmation) {
const { currentlyLoadedModel, affectedSessions } = data.data;
const names = affectedSessions.map((s) => s.name || s.id).join(', ');
@@ -871,6 +919,26 @@ Object.assign(CodemanApp.prototype, {
// once `data.success !== false`.
let payload = data?.success !== false ? data?.data : undefined;
// This CLI's own fixed overhead (system prompt + tool schemas) may exceed the
// model's real discovered context outright — no context-length declaration can
// fix that, since compaction only trims conversation history and there is none
// on message 1. Warn and let the user decide whether to launch anyway, same
// confirmed:true re-send pattern as the swap check below.
if (ok && payload?.requiresContextWarning) {
const proceed = await this._confirmContextWarning(
payload.modelId,
payload.contextLength,
payload.minSafeContextTokens
);
if (!proceed) {
switchingToast?.dismiss();
this.showToast('Kept the native backend — context window too small', 'info');
return;
}
({ ok, data, res } = await this._applyCustomModelToSession(sessionId, endpointId, modelId, true));
payload = data?.success !== false ? data?.data : undefined;
}
// llama-swap runs one model at a time: switching would unload it out from under
// another session actively using it. The route only asks when that's actually true
// (never just because a swap is needed at all) — confirming re-sends the exact same
+35
View File
@@ -23,11 +23,46 @@ import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
import { egressBlockedReason, webviewFetch } from '../webview-egress.js';
import { CustomModelHostSchema } from '../schemas.js';
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../../custom-model-hosts.js';
import type { CliEntry } from '../../config/cli-registry/types.js';
const CODEMAN_CONFIG_DIR = getDataDir();
const DISCOVER_TIMEOUT_MS = 8000;
const PROPS_TIMEOUT_MS = 5000;
/**
* Claude Code's own system prompt + tool schemas cost roughly this many tokens on EVERY
* request, before a single character of conversation history — confirmed live, twice, on
* requests reporting `in:0 out:0` (the very first exchange) failing at ~36.4K tokens. No
* `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes this: that setting only changes when Claude
* Code decides to COMPACT conversation history, and there is no history yet on the first
* message for it to trim. A model whose real context is below this floor will refuse
* Claude Code's very first message outright, unconditionally.
*
* Set well above the ~36.4K actually measured — CLAUDE.md size, active MCP servers, and
* enabled skills all add to a project's real baseline, so the observed figure is a floor
* for THAT one workspace, not a ceiling for every one. Erring conservative here means a
* borderline-safe model still gets warned about (the user can launch anyway), rather than
* this floor missing a genuinely-too-small one because a smaller test project happened to
* fit.
*/
export const CLAUDE_MIN_SAFE_CONTEXT_TOKENS = 40000;
/**
* True when applying this model to this CLI is heading for a guaranteed first-message
* failure per `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` above. Gated on `contextLengthVar` (today,
* only claude's registry entry declares one) rather than a hardcoded mode check: a CLI
* with a small enough baseline of its own to never trip this would have no reason to
* declare the field in the first place, so the check simply never applies to it.
*/
export function exceedsSafeContextFloor(
entry: Pick<CliEntry, 'capabilities'>,
contextLength: number | undefined
): boolean {
const cap = entry.capabilities.customModelInjection;
if (cap.kind !== 'env' || !cap.contextLengthVar) return false;
return typeof contextLength === 'number' && contextLength < CLAUDE_MIN_SAFE_CONTEXT_TOKENS;
}
function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse<never> | null {
if (!isMultiUserMode() || isAdmin(req)) return null;
reply.code(403);
+37 -3
View File
@@ -55,7 +55,12 @@ import {
} from '../schemas.js';
import { readCustomModelHosts } from '../../custom-model-hosts.js';
import { applyCustomModelInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
import { getLlamaSwapStatus, triggerLlamaSwapLoad } from './custom-model-routes.js';
import {
getLlamaSwapStatus,
triggerLlamaSwapLoad,
exceedsSafeContextFloor,
CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
} from './custom-model-routes.js';
import { matchesPattern } from '../../config/cli-registry/patterns.js';
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
import { TabLayoutValidationError } from '../../tab-layout.js';
@@ -1209,6 +1214,23 @@ export function registerSessionRoutes(
if (!endpoint) {
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
}
const contextLength = endpoint.modelContextLengths?.[body.modelId];
// Some CLIs (today: only claude) carry enough of their own fixed system-prompt/tool-
// schema overhead that a small enough real context guarantees a first-message failure
// no matter what CLAUDE_CODE_MAX_CONTEXT_TOKENS says — confirmed live at ~36.4K tokens
// against a model configured with a real 16384-token context. Warn before committing
// to a restart that's certain to fail, rather than letting the user discover it via a
// cryptic 400 from the CLI itself. `confirmed` (already used for the swap-conflict
// warning below) skips this too — the user has already said "launch anyway" once.
if (!body.confirmed && exceedsSafeContextFloor(entry, contextLength)) {
return {
requiresContextWarning: true,
modelId: body.modelId,
contextLength,
minSafeContextTokens: CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
};
}
// llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on
// demand, which can take anywhere from a few seconds to over a minute — long enough
@@ -1250,7 +1272,6 @@ export function registerSessionRoutes(
// fails its pattern rather than quoting it, which would silently launch the CLI on its
// own default provider again, so refuse an id the pattern cannot carry up front.
const modelSpec = entry.launch.params.model;
const contextLength = endpoint.modelContextLengths?.[body.modelId];
const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id, contextLength);
if (!applied) {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
@@ -3570,6 +3591,20 @@ export function registerSessionRoutes(
const cmHosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
const cmEndpoint = cmHosts.find((h) => h.id === customModel.endpointId);
if (!cmEndpoint) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
const cmContextLength = cmEndpoint.modelContextLengths?.[customModel.modelId];
// See the dedicated route's own comment for the full reasoning: some CLIs' own fixed
// overhead can exceed a small enough real context on the very first message,
// regardless of contextLengthVar. Warn before creating a session that's certain to
// fail immediately.
if (!customModel.confirmed && exceedsSafeContextFloor(cmEntry, cmContextLength)) {
return {
requiresContextWarning: true,
modelId: customModel.modelId,
contextLength: cmContextLength,
minSafeContextTokens: CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
};
}
// See the dedicated route's own comment for the full reasoning: llama.cpp runs one
// model at a time, llama-swap swaps on demand, and switching away from what another
@@ -3603,7 +3638,6 @@ export function registerSessionRoutes(
// with, not a placeholder: `new Session({ id: ... })` accepts an explicit id for
// exactly this reason.
qsCustomModelSessionId = randomUUID();
const cmContextLength = cmEndpoint.modelContextLengths?.[customModel.modelId];
const cmApplied = applyCustomModelInjection(
cmEntry,
cmEndpoint,
+99
View File
@@ -57,6 +57,9 @@ function bootApp(
<div class="modal" id="customModelSwapConfirmModal">
<p id="customModelSwapConfirmMessage"></p>
</div>
<div class="modal" id="customModelContextWarningModal">
<p id="customModelContextWarningMessage"></p>
</div>
</body>`,
{ url: 'http://localhost/', runScripts: 'dangerously' }
);
@@ -859,6 +862,64 @@ describe('Custom Model Endpoint Profiles: model-size load-time estimate', () =>
});
});
describe("Custom Model Endpoint Profiles: requiresContextWarning (this CLI's own overhead can exceed a small model's real context)", () => {
function launchHarness(applyResponses: Array<Record<string, unknown>>) {
const { win, app } = bootApp({});
app.activeSessionId = 'old-session';
app.run = async () => {
app.activeSessionId = 'new-session';
};
const applyBodies: unknown[] = [];
let call = 0;
app._api = async (path: string, opts?: { body?: unknown }) => {
if (path.endsWith('/custom-model')) {
applyBodies.push(opts?.body);
const data = applyResponses[Math.min(call, applyResponses.length - 1)];
call += 1;
return { ok: true, status: 200, json: async () => ({ success: true, data }) };
}
throw new Error(`unexpected _api call: ${path}`);
};
return { win, app, applyBodies };
}
it('confirming the in-app context-warning modal re-sends the apply with confirmed:true', async () => {
const { app, applyBodies } = launchHarness([
{ requiresContextWarning: true, modelId: 'qwen3', contextLength: 16384, minSafeContextTokens: 40000 },
{ customModel: { endpointId: 'llama-box' }, restarted: true, modelSwapInProgress: false },
]);
let confirmArgs: unknown[] | undefined;
app._confirmContextWarning = async (...args: unknown[]) => {
confirmArgs = args;
return true;
};
await app.runCustomModelEntry('claude', 'llama-box', 'qwen3');
expect(confirmArgs).toEqual(['qwen3', 16384, 40000]);
expect(applyBodies).toEqual([
{ endpointId: 'llama-box', modelId: 'qwen3' },
{ endpointId: 'llama-box', modelId: 'qwen3', confirmed: true },
]);
});
it('declining the in-app context-warning modal keeps the native backend and never re-sends the apply', async () => {
const { app, applyBodies } = launchHarness([
{ requiresContextWarning: true, modelId: 'qwen3', contextLength: 16384, minSafeContextTokens: 40000 },
]);
app._confirmContextWarning = async () => false;
let toastMessage: string | undefined;
app.showToast = (msg: string) => {
toastMessage = msg;
};
await app.runCustomModelEntry('claude', 'llama-box', 'qwen3');
expect(applyBodies).toHaveLength(1); // no second (confirmed) call
expect(toastMessage).toMatch(/context window too small/i);
});
});
describe('Custom Model Endpoint Profiles: _confirmModelSwap (in-app modal, replaces a native confirm() popup)', () => {
it('shows the message, activates the modal, and resolves true when "Switch anyway" is clicked', async () => {
const { win, app } = bootApp({});
@@ -884,3 +945,41 @@ describe('Custom Model Endpoint Profiles: _confirmModelSwap (in-app modal, repla
expect(win.document.getElementById('customModelSwapConfirmModal')!.classList.contains('active')).toBe(false);
});
});
describe('Custom Model Endpoint Profiles: _confirmContextWarning (in-app modal, native backend never restarted while it is up)', () => {
it('shows a message naming the model, the discovered context and the safe floor, activates the modal, and resolves true on "Launch anyway"', async () => {
const { win, app } = bootApp({});
const promise = app._confirmContextWarning('qwen3.8-27b-ud-q4_k_xl', 16384, 40000);
const modal = win.document.getElementById('customModelContextWarningModal')!;
expect(modal.classList.contains('active')).toBe(true);
const message = win.document.getElementById('customModelContextWarningMessage')!.textContent!;
expect(message).toContain('qwen3.8-27b-ud-q4_k_xl');
expect(message).toContain('16,384');
expect(message).toContain('40,000');
expect(message).toMatch(/llama-swap/i);
expect(message).toMatch(/fit-ctx/i);
app._resolveContextWarningConfirm(true);
expect(await promise).toBe(true);
expect(modal.classList.contains('active')).toBe(false);
});
it('resolves false when Cancel is clicked', async () => {
const { win, app } = bootApp({});
const promise = app._confirmContextWarning('qwen3', 16384, 40000);
app._resolveContextWarningConfirm(false);
expect(await promise).toBe(false);
expect(win.document.getElementById('customModelContextWarningModal')!.classList.contains('active')).toBe(false);
});
it('describes an unknown context length without printing a bogus number', async () => {
const { win, app } = bootApp({});
void app._confirmContextWarning('qwen3', undefined, 40000);
const message = win.document.getElementById('customModelContextWarningMessage')!.textContent!;
expect(message).not.toMatch(/undefined/);
expect(message).toMatch(/unknown/i);
app._resolveContextWarningConfirm(false);
});
});
@@ -275,6 +275,59 @@ describe('POST /api/quick-start: customModel (one-shot custom-model launch)', ()
});
});
describe("context-window floor warning (this CLI's own overhead can exceed a small model's real context)", () => {
const SMALL_CTX_ENDPOINT: CustomModelHost = {
id: 'ep-small',
label: 'tiny box',
baseUrl: 'http://192.168.1.51:8080',
apiKey: 'k',
modelContextLengths: { 'qwen3.8-27b-ud-q4_k_xl': 16384 },
};
it('warns instead of launching when the discovered context is below the safe floor', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT, SMALL_CTX_ENDPOINT]);
const res = await quickStart({
caseName: 'cm-small-ctx',
mode: 'claude',
customModel: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl' },
});
const body = res.json();
expect(body.requiresContextWarning).toBe(true);
expect(body.modelId).toBe('qwen3.8-27b-ud-q4_k_xl');
expect(body.contextLength).toBe(16384);
expect(body.minSafeContextTokens).toBe(40000);
// Nothing was actually created.
expect(ctx.sessions.size).toBe(1);
});
it('launches once confirmed, skipping the context check', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT, SMALL_CTX_ENDPOINT]);
const res = await quickStart({
caseName: 'cm-small-ctx-confirmed',
mode: 'claude',
customModel: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl', confirmed: true },
});
expect(res.statusCode).toBe(200);
expect(res.json().requiresContextWarning).toBeUndefined();
expect(ctx.sessions.size).toBe(2);
});
it('does not warn when nothing about context was discovered', async () => {
const res = await quickStart({
caseName: 'cm-no-ctx-data',
mode: 'claude',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.statusCode).toBe(200);
expect(res.json().requiresContextWarning).toBeUndefined();
});
});
describe('triggering the actual llama-swap load (not just watching for it)', () => {
it('sends a real inference request naming the target model, concurrently with launching the session', async () => {
const chatCalls: unknown[] = [];
+107
View File
@@ -351,6 +351,113 @@ describe('POST /api/sessions/:id/custom-model', () => {
});
});
describe("context-window floor warning (this CLI's own overhead can exceed a small model's real context)", () => {
const SMALL_CTX_ENDPOINT: CustomModelHost = {
id: 'ep-small',
label: 'tiny box',
baseUrl: 'http://192.168.1.51:8080',
apiKey: 'k',
modelContextLengths: { 'qwen3.8-27b-ud-q4_k_xl': 16384 },
};
it('warns instead of applying when the discovered context is below the safe floor', async () => {
const { app, ctx } = await setup();
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl' },
});
const body = res.json();
expect(body.success).not.toBe(false);
expect(body.requiresContextWarning).toBe(true);
expect(body.modelId).toBe('qwen3.8-27b-ud-q4_k_xl');
expect(body.contextLength).toBe(16384);
expect(body.minSafeContextTokens).toBe(40000);
// Nothing actually applied yet — this call only warned, it did not switch.
expect(session.setCustomModel).not.toHaveBeenCalled();
expect(session.restartCli).not.toHaveBeenCalled();
});
it('applies once confirmed, skipping the context check the second time', async () => {
const { app, ctx } = await setup();
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl', confirmed: true },
});
const body = res.json();
expect(body.requiresContextWarning).toBeUndefined();
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
expect(session.restartCli).toHaveBeenCalledTimes(1);
});
it('does not warn when the discovered context is comfortably above the floor', async () => {
const { app, ctx } = await setup();
const roomyEndpoint: CustomModelHost = {
id: 'ep-roomy',
label: 'roomy box',
baseUrl: 'http://192.168.1.52:8080',
apiKey: 'k',
modelContextLengths: { qwen3: 65536 },
};
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, roomyEndpoint]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-roomy', modelId: 'qwen3' },
});
expect(res.json().requiresContextWarning).toBeUndefined();
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
});
it('does not warn when the context length was never discovered (nothing to compare)', async () => {
const { app, ctx } = await setup();
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.json().requiresContextWarning).toBeUndefined();
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
});
it('does not warn for a CLI whose registry entry declares no contextLengthVar (opencode)', async () => {
// opencode's customModelInjection kind is configContentEnv, not env+contextLengthVar,
// so exceedsSafeContextFloor is false by construction regardless of context size.
const { app, ctx } = await setup();
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'opencode';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl' },
});
expect(res.json().requiresContextWarning).toBeUndefined();
});
});
describe('triggering the actual llama-swap load (not just watching for it)', () => {
it('sends a real inference request naming the target model when it is not already loaded and ready', async () => {
const { app, ctx } = await setup();