feat(custom-model): warn before launching Claude on a model too small for its own overhead

Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.

- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
  registry entry declares contextLengthVar (currently only claude) and
  the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
  (40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
  quick-start customModel path) check this before the swap-conflict
  check and before launching/restarting anything, returning
  {requiresContextWarning, modelId, contextLength, minSafeContextTokens}
  — skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
  _resolveContextWarningConfirm (session-ui.js), wired into both
  _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
  (the path Claude actually uses) ahead of the swap-confirmation check.
  Explains the fix in-modal: give the model an explicit larger -c/
  --ctx-size in llama-swap instead of relying on --fit-ctx, which
  optimizes for the biggest model that fits rather than the biggest
  context.

Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 07:36:57 +08:00
co-authored by Claude Sonnet 5
parent 993710263d
commit b45a96358e
10 changed files with 500 additions and 20 deletions
+35
View File
@@ -23,11 +23,46 @@ import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
import { egressBlockedReason, webviewFetch } from '../webview-egress.js';
import { CustomModelHostSchema } from '../schemas.js';
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../../custom-model-hosts.js';
import type { CliEntry } from '../../config/cli-registry/types.js';
const CODEMAN_CONFIG_DIR = getDataDir();
const DISCOVER_TIMEOUT_MS = 8000;
const PROPS_TIMEOUT_MS = 5000;
/**
* Claude Code's own system prompt + tool schemas cost roughly this many tokens on EVERY
* request, before a single character of conversation history — confirmed live, twice, on
* requests reporting `in:0 out:0` (the very first exchange) failing at ~36.4K tokens. No
* `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes this: that setting only changes when Claude
* Code decides to COMPACT conversation history, and there is no history yet on the first
* message for it to trim. A model whose real context is below this floor will refuse
* Claude Code's very first message outright, unconditionally.
*
* Set well above the ~36.4K actually measured — CLAUDE.md size, active MCP servers, and
* enabled skills all add to a project's real baseline, so the observed figure is a floor
* for THAT one workspace, not a ceiling for every one. Erring conservative here means a
* borderline-safe model still gets warned about (the user can launch anyway), rather than
* this floor missing a genuinely-too-small one because a smaller test project happened to
* fit.
*/
export const CLAUDE_MIN_SAFE_CONTEXT_TOKENS = 40000;
/**
* True when applying this model to this CLI is heading for a guaranteed first-message
* failure per `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` above. Gated on `contextLengthVar` (today,
* only claude's registry entry declares one) rather than a hardcoded mode check: a CLI
* with a small enough baseline of its own to never trip this would have no reason to
* declare the field in the first place, so the check simply never applies to it.
*/
export function exceedsSafeContextFloor(
entry: Pick<CliEntry, 'capabilities'>,
contextLength: number | undefined
): boolean {
const cap = entry.capabilities.customModelInjection;
if (cap.kind !== 'env' || !cap.contextLengthVar) return false;
return typeof contextLength === 'number' && contextLength < CLAUDE_MIN_SAFE_CONTEXT_TOKENS;
}
function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse<never> | null {
if (!isMultiUserMode() || isAdmin(req)) return null;
reply.code(403);
+37 -3
View File
@@ -55,7 +55,12 @@ import {
} from '../schemas.js';
import { readCustomModelHosts } from '../../custom-model-hosts.js';
import { applyCustomModelInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
import { getLlamaSwapStatus, triggerLlamaSwapLoad } from './custom-model-routes.js';
import {
getLlamaSwapStatus,
triggerLlamaSwapLoad,
exceedsSafeContextFloor,
CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
} from './custom-model-routes.js';
import { matchesPattern } from '../../config/cli-registry/patterns.js';
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
import { TabLayoutValidationError } from '../../tab-layout.js';
@@ -1209,6 +1214,23 @@ export function registerSessionRoutes(
if (!endpoint) {
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
}
const contextLength = endpoint.modelContextLengths?.[body.modelId];
// Some CLIs (today: only claude) carry enough of their own fixed system-prompt/tool-
// schema overhead that a small enough real context guarantees a first-message failure
// no matter what CLAUDE_CODE_MAX_CONTEXT_TOKENS says — confirmed live at ~36.4K tokens
// against a model configured with a real 16384-token context. Warn before committing
// to a restart that's certain to fail, rather than letting the user discover it via a
// cryptic 400 from the CLI itself. `confirmed` (already used for the swap-conflict
// warning below) skips this too — the user has already said "launch anyway" once.
if (!body.confirmed && exceedsSafeContextFloor(entry, contextLength)) {
return {
requiresContextWarning: true,
modelId: body.modelId,
contextLength,
minSafeContextTokens: CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
};
}
// llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on
// demand, which can take anywhere from a few seconds to over a minute — long enough
@@ -1250,7 +1272,6 @@ export function registerSessionRoutes(
// fails its pattern rather than quoting it, which would silently launch the CLI on its
// own default provider again, so refuse an id the pattern cannot carry up front.
const modelSpec = entry.launch.params.model;
const contextLength = endpoint.modelContextLengths?.[body.modelId];
const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id, contextLength);
if (!applied) {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
@@ -3570,6 +3591,20 @@ export function registerSessionRoutes(
const cmHosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
const cmEndpoint = cmHosts.find((h) => h.id === customModel.endpointId);
if (!cmEndpoint) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
const cmContextLength = cmEndpoint.modelContextLengths?.[customModel.modelId];
// See the dedicated route's own comment for the full reasoning: some CLIs' own fixed
// overhead can exceed a small enough real context on the very first message,
// regardless of contextLengthVar. Warn before creating a session that's certain to
// fail immediately.
if (!customModel.confirmed && exceedsSafeContextFloor(cmEntry, cmContextLength)) {
return {
requiresContextWarning: true,
modelId: customModel.modelId,
contextLength: cmContextLength,
minSafeContextTokens: CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
};
}
// See the dedicated route's own comment for the full reasoning: llama.cpp runs one
// model at a time, llama-swap swaps on demand, and switching away from what another
@@ -3603,7 +3638,6 @@ export function registerSessionRoutes(
// with, not a placeholder: `new Session({ id: ... })` accepts an explicit id for
// exactly this reason.
qsCustomModelSessionId = randomUUID();
const cmContextLength = cmEndpoint.modelContextLengths?.[customModel.modelId];
const cmApplied = applyCustomModelInjection(
cmEntry,
cmEndpoint,