feat(custom-model): warn before launching Claude on a model too small for its own overhead

Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.

- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
  registry entry declares contextLengthVar (currently only claude) and
  the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
  (40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
  quick-start customModel path) check this before the swap-conflict
  check and before launching/restarting anything, returning
  {requiresContextWarning, modelId, contextLength, minSafeContextTokens}
  — skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
  _resolveContextWarningConfirm (session-ui.js), wired into both
  _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
  (the path Claude actually uses) ahead of the swap-confirmation check.
  Explains the fix in-modal: give the model an explicit larger -c/
  --ctx-size in llama-swap instead of relying on --fit-ctx, which
  optimizes for the biggest model that fits rather than the biggest
  context.

Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-17 07:36:57 +08:00
co-authored by Claude Sonnet 5
parent 993710263d
commit b45a96358e
10 changed files with 500 additions and 20 deletions
+23
View File
@@ -948,6 +948,29 @@
</div>
</div>
<!-- Custom Model Endpoint Profiles: context-window-too-small warning
(docs/custom-model-endpoints-plan.md) — shown before launching a CLI
whose own fixed system-prompt/tool-schema overhead exceeds the
model's real discovered context, which guarantees a first-message
failure regardless of CLAUDE_CODE_MAX_CONTEXT_TOKENS. See
_confirmContextWarning() in session-ui.js. -->
<div class="modal" id="customModelContextWarningModal">
<div class="modal-backdrop" onclick="app._resolveContextWarningConfirm(false)"></div>
<div class="modal-content modal-sm">
<div class="modal-header">
<h3>Context window too small</h3>
<button class="modal-close" onclick="app._resolveContextWarningConfirm(false)" aria-label="Cancel">&times;</button>
</div>
<div class="modal-body">
<p class="form-hint" id="customModelContextWarningMessage" style="white-space: pre-wrap;"></p>
</div>
<div class="modal-footer">
<button class="btn-toolbar" onclick="app._resolveContextWarningConfirm(false)">Cancel</button>
<button class="btn-toolbar btn-primary" onclick="app._resolveContextWarningConfirm(true)">Launch anyway</button>
</div>
</div>
</div>
<!-- Cron Jobs Modal -->
<div class="modal" id="cronModal">
<div class="modal-backdrop" onclick="app.closeCron()"></div>
+68
View File
@@ -693,6 +693,45 @@ Object.assign(CodemanApp.prototype, {
resolve?.(proceed);
},
/**
* In-app warning shown when the apply route reports `requiresContextWarning`: this
* model's real discovered context is smaller than the CLI's own fixed system-prompt/
* tool-schema overhead, which guarantees the very first message fails outright — no
* `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since there is no conversation
* history yet for compaction to trim. Same promise-based pattern as
* `_confirmModelSwap`; `_resolveContextWarningConfirm` settles it.
*/
_confirmContextWarning(modelId, contextLength, minSafeContextTokens) {
const modal = document.getElementById('customModelContextWarningModal');
const messageEl = document.getElementById('customModelContextWarningMessage');
if (messageEl) {
const known = typeof contextLength === 'number';
messageEl.textContent =
`${modelId} is configured with ` +
(known ? `only ${contextLength.toLocaleString()} tokens of` : 'an unknown (too small)') +
` context, but this CLI needs roughly ${minSafeContextTokens.toLocaleString()}+ tokens just for its own ` +
`system prompt and tools — before any conversation history. Its very first message will fail outright, ` +
`no matter what context size Codeman tells it to expect.\n\n` +
`To fix this, reconfigure llama-swap to give this model (or a smaller one) an explicit larger context ` +
`instead of relying on auto-fit (--fit-ctx), which optimizes for the biggest MODEL that fits, not the ` +
`biggest CONTEXT — e.g. add "-c 65536" (or as large a --ctx-size as your hardware holds) to its llama-swap ` +
`config entry. A smaller model at a much larger explicit context often fits in the same VRAM a bigger ` +
`model's auto-fit context gets shrunk to make room for.`;
}
modal?.classList.add('active');
return new Promise((resolve) => {
this._resolveContextWarningConfirmPromise = resolve;
});
},
/** Called by the modal's Cancel/Launch-anyway buttons and its backdrop click. */
_resolveContextWarningConfirm(proceed) {
document.getElementById('customModelContextWarningModal')?.classList.remove('active');
const resolve = this._resolveContextWarningConfirmPromise;
this._resolveContextWarningConfirmPromise = null;
resolve?.(proceed);
},
/** A model row in the picker modal was clicked: close it and launch with that choice. */
chooseCustomModelAndRun(modelId) {
const pending = this._pendingCustomModelPick;
@@ -789,6 +828,15 @@ Object.assign(CodemanApp.prototype, {
return res.json();
};
let data = await post(bodyObj);
if (data?.data?.requiresContextWarning) {
const { modelId, contextLength, minSafeContextTokens } = data.data;
const proceed = await this._confirmContextWarning(modelId, contextLength, minSafeContextTokens);
if (!proceed) {
this._lastCustomModelLaunchResult = undefined;
return { success: false, error: 'Launch cancelled — context window too small' };
}
data = await post({ ...bodyObj, customModel: { ...bodyObj.customModel, confirmed: true } });
}
if (data?.data?.requiresConfirmation) {
const { currentlyLoadedModel, affectedSessions } = data.data;
const names = affectedSessions.map((s) => s.name || s.id).join(', ');
@@ -871,6 +919,26 @@ Object.assign(CodemanApp.prototype, {
// once `data.success !== false`.
let payload = data?.success !== false ? data?.data : undefined;
// This CLI's own fixed overhead (system prompt + tool schemas) may exceed the
// model's real discovered context outright — no context-length declaration can
// fix that, since compaction only trims conversation history and there is none
// on message 1. Warn and let the user decide whether to launch anyway, same
// confirmed:true re-send pattern as the swap check below.
if (ok && payload?.requiresContextWarning) {
const proceed = await this._confirmContextWarning(
payload.modelId,
payload.contextLength,
payload.minSafeContextTokens
);
if (!proceed) {
switchingToast?.dismiss();
this.showToast('Kept the native backend — context window too small', 'info');
return;
}
({ ok, data, res } = await this._applyCustomModelToSession(sessionId, endpointId, modelId, true));
payload = data?.success !== false ? data?.data : undefined;
}
// llama-swap runs one model at a time: switching would unload it out from under
// another session actively using it. The route only asks when that's actually true
// (never just because a swap is needed at all) — confirming re-sends the exact same