From b45a96358eb68683e8240e69df09c84223a54612 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Thu, 17 Sep 2026 07:36:57 +0800 Subject: [PATCH] feat(custom-model): warn before launching Claude on a model too small for its own overhead MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Claude Code's own fixed per-turn overhead (system prompt + tool schemas, ~36.4K tokens measured live) can exceed a small local model's entire real context before any conversation history exists to compact — confirmed live twice as an in:0 out:0 failure on the very first message sent. CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when history gets compacted, and there is none on message one. - exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's registry entry declares contextLengthVar (currently only claude) and the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS (40000). A no-op for every other CLI by construction. - Both apply routes (POST /api/sessions/:id/custom-model and the quick-start customModel path) check this before the swap-conflict check and before launching/restarting anything, returning {requiresContextWarning, modelId, contextLength, minSafeContextTokens} — skipped when confirmed:true. - Frontend: #customModelContextWarningModal + _confirmContextWarning/ _resolveContextWarningConfirm (session-ui.js), wired into both _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart (the path Claude actually uses) ahead of the swap-confirmation check. Explains the fix in-modal: give the model an explicit larger -c/ --ctx-size in llama-swap instead of relying on --fit-ctx, which optimizes for the biggest model that fits rather than the biggest context. Tests added for the route-level warning/confirm/skip cases and the frontend modal + launch-flow wiring. Docs updated (custom-model- endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running changeset extended. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG --- .changeset/run-menu-custom-model-picker.md | 9 +- docs/custom-model-endpoints.md | 57 +++++++++- docs/wiki/Custom-Model-Endpoints.md | 29 +++-- src/web/public/index.html | 23 ++++ src/web/public/session-ui.js | 68 ++++++++++++ src/web/routes/custom-model-routes.ts | 35 ++++++ src/web/routes/session-routes.ts | 40 ++++++- test/custom-model-run-menu-ui.test.ts | 99 +++++++++++++++++ test/routes/quick-start-custom-model.test.ts | 53 +++++++++ test/routes/session-custom-model.test.ts | 107 +++++++++++++++++++ 10 files changed, 500 insertions(+), 20 deletions(-) diff --git a/.changeset/run-menu-custom-model-picker.md b/.changeset/run-menu-custom-model-picker.md index bedbdd8c..315a9f24 100644 --- a/.changeset/run-menu-custom-model-picker.md +++ b/.changeset/run-menu-custom-model-picker.md @@ -1,5 +1,5 @@ --- -"aicodeman": minor +'aicodeman': minor --- **Custom model endpoints: Run-menu picker, and hardening from real llama-swap validation** (#430, follow-up to #393's HTTP-API-only cut). With **Custom model endpoints** on (App Settings → Models) and at least one saved endpoint carrying a discovered model, the Run dropdown grows a **Custom Endpoints** section generated live off the CLI registry's own `capabilities.customModelInjection` — one entry per (harness that can redirect to a custom endpoint, saved endpoint). Picking one launches that harness and applies the endpoint to it; with two or more discovered models a small, scrollable dialog asks which one first, the endpoint's `defaultModelId` marked but never auto-chosen. Endpoints also now re-discover themselves automatically every 5 minutes in the background, one unreachable endpoint never blocking the others. @@ -8,10 +8,11 @@ Everything below was found and fixed against a **real llama-swap server**, not j - **Session-busy false refusal.** A freshly launched CLI reports itself `busy` for its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying. - **Errors and confirmations you can actually read.** Toasts now default to sticky with a close button (errors always were meant to stay, but a fixed 3s timer silently hid them); a failed apply's real server-side reason (not a generic message) reaches the toast. -- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on *every* launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would. -- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length from llama.cpp/llama-swap's `GET /props?model=`, but **only** for a model llama-swap's own `/v1/models` response already reports loaded — never an unloaded one, since asking about one risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery — and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS`. +- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on _every_ launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would. +- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS` — sourced primarily from llama-swap's own `GET /running`, whose `cmd` field carries the launch flags (`--fit-ctx`/`-c`/`--ctx-size`) actually in effect, since `GET /props`'s `n_ctx` was confirmed live to report the model's theoretical/trained maximum rather than the real `--fit-ctx`-shrunk runtime context (a 154112-vs-16384 discrepancy, caught only because the fixed value still overflowed) — `/props` is now a fallback for a plain llama.cpp server with no `/running` at all. +- **Context floor too small for Claude Code to even start.** Fixing the overflow above surfaced a second, unfixable-by-injection failure: Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on their own (confirmed live via an `in:0 out:0` failure on the very first message), which can exceed a small model's entire real context before any conversation history exists to trim — no `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since it only governs when history gets compacted. Applying such a model now returns a warning (gated on the CLI registry declaring a `contextLengthVar`, so it's a no-op for every other harness) instead of launching straight into a guaranteed first-message failure, and the Run-menu picker shows it as an in-app dialog naming the model, its discovered context and the ~40K safe floor, with the actual fix spelled out: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on auto-fit, which optimizes for the biggest model that fits rather than the biggest context. "Launch anyway" is still one click away. - **The real root cause of "it still says opus, not my model."** llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on demand, which can take anywhere from a few seconds to well over a minute — long enough that a session mid-swap is indistinguishable from one that never left the native backend. Applying a selection now checks llama-swap's own `GET /running` first (feature-detected; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is never checked); if switching would unload a model **another live session is actively using**, the apply is refused with a warning naming that session instead of silently switching, and a confirmation retry proceeds anyway. Either way, a sticky "loading model…" toast now covers the actual swap window until llama-swap reports the target model ready, so a prompt sent mid-swap reads as "loading," never as silence or an answer from whatever was loaded a moment before. Remote (SSH) and Docker sessions are refused for now (400) — their restart reattaches the durable remote/in-container tmux rather than relaunching the agent. -**One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all.** Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). `POST /api/quick-start` now accepts a `customModel` field and computes the same injection *before* the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own `--resume`-based restart is far less jarring, and `runClaude()`'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work). +**One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all.** Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). `POST /api/quick-start` now accepts a `customModel` field and computes the same injection _before_ the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own `--resume`-based restart is far less jarring, and `runClaude()`'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work). diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index e3db6016..3f7b1571 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -239,8 +239,17 @@ registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:** model id and never compacts, which reliably overflows a much smaller real local context — confirmed live: a stock ~33.7K-token system prompt against a 16384-token llama-swap model failed with `exceeds the available context - size`. No entry for the model in `modelContextLengths` means the var is - simply omitted, never a guess. +size`. No entry for the model in `modelContextLengths` means the var is + simply omitted, never a guess. ⚠️ **This var only affects when Claude + Code compacts conversation _history_ — it cannot fix a model whose real + context is smaller than Claude Code's own fixed per-turn overhead** + (system prompt + tool schemas, empirically ~36.4K tokens, confirmed live + via an `in:0 out:0` failure on the very first message, before any + history exists to compact). No context-length declaration changes that + fixed overhead, so a model below the safe floor fails outright on + message one regardless of what this var says. See "Context-window floor + warning" below for how Codeman catches this case before launching + instead of after. - `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory the `configDir`-kind CLIs use (empty, no files written into it), so the injected `ANTHROPIC_API_KEY` never shares a directory with a stored @@ -261,7 +270,7 @@ registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:** **That isolated directory needed one more fix to actually be usable non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a real profile's prior "Detected a custom API key — use it?" approvals, so -without more, Claude Code stops and asks that on *every single launch* — +without more, Claude Code stops and asks that on _every single launch_ — confirmed live, and with nobody at a TTY to answer, its own default answer ("No") silently refuses the very key this feature just injected, which looks like the endpoint being ignored entirely. `customModelInjection`'s @@ -285,7 +294,7 @@ which can take anywhere from a few seconds to well over a minute: one-shot `POST /api/quick-start` above) call llama-swap's own `GET /running` first — feature-detected, so a plain llama.cpp/OpenAI- compatible server (no such endpoint) is simply never checked. If a - *different* model is currently loaded and ready, and another **live + _different_ model is currently loaded and ready, and another **live session's own selection** is using it, the apply returns `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` instead of silently switching — nothing is applied or created yet. @@ -298,13 +307,51 @@ which can take anywhere from a few seconds to well over a minute: reached llama-swap at all (nothing in its own server logs), since nothing had actually asked it to load anything yet. Both apply routes now also send the smallest real request that will — `POST /v1/chat/ - completions` with `max_tokens: 1` and one throwaway message — whenever the +completions` with `max_tokens: 1` and one throwaway message — whenever the target model isn't already the one loaded and ready, fire-and-forget (its response is never read; `GET /api/model-endpoints/:id/running-status`, polled client-side, is what actually confirms readiness). The response also carries `modelSwapInProgress: true` in that case, which is what drives the Run-menu picker's own "loading model" status banner. +## Context-window floor warning + +Claude Code's own fixed per-turn overhead (system prompt + tool schemas, +empirically ~36.4K tokens) can exceed a small local model's _entire_ real +context on its own, before any conversation history exists to fill it — +confirmed live twice, both as an `in:0 out:0` failure on the very first +message sent. `CLAUDE_CODE_MAX_CONTEXT_TOKENS` (above) cannot fix this: it +only governs when Claude Code compacts conversation history, and there is +no history yet on message one. Applying such a model would look like the +endpoint being ignored, or the wrong model being used, when in fact the +endpoint applied correctly and the model is simply too small for this CLI. + +Both apply routes (the restart route and the one-shot `POST +/api/quick-start`) now check for this **before** launching or restarting +anything, gated on the CLI's registry entry declaring a `contextLengthVar` +(currently only claude — the check is a no-op for every other CLI by +construction, never a hardcoded mode check). If the model's discovered +context (`modelContextLengths`, from discovery above) is below +`CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000, comfortably above the measured +~36.4K overhead), the response is `{requiresContextWarning: true, modelId, +contextLength, minSafeContextTokens}` instead of applying — nothing is +restarted or created yet. A context length that was never discovered at +all skips the check entirely (nothing to compare, so it fails open rather +than warning on every model an endpoint hasn't reported a size for). +Retrying with `confirmed: true` launches anyway. + +The Run-menu picker shows this as an in-app modal +(`#customModelContextWarningModal`, matching the llama-swap conflict +modal's look) naming the model, its discovered context, and the safe +floor, and explaining the fix: reconfigure llama-swap to give that model +(or a smaller one) an explicit larger context instead of relying on +auto-fit (`--fit-ctx`), which optimizes for the biggest _model_ that fits +rather than the biggest _context_ — e.g. adding `-c 65536` (or as large a +`--ctx-size` as the hardware holds) to that model's llama-swap config +entry. A smaller model at a much larger explicit context often fits in +the same VRAM a bigger model's auto-fit context gets shrunk to make room +for. + Clear back to the harness's native cloud default with: ```bash diff --git a/docs/wiki/Custom-Model-Endpoints.md b/docs/wiki/Custom-Model-Endpoints.md index a6eed35f..1b681582 100644 --- a/docs/wiki/Custom-Model-Endpoints.md +++ b/docs/wiki/Custom-Model-Endpoints.md @@ -28,7 +28,7 @@ shows up without another manual click of **Discover**. One endpoint being unreac given cycle (powered off, wrong network) never blocks the others from refreshing. **Context length is picked up automatically where it can be, safely.** Against a -llama.cpp/llama-swap server, discovery also learns each *currently loaded* model's real +llama.cpp/llama-swap server, discovery also learns each _currently loaded_ model's real context window and applies it to the launched session (Claude Code today — see below), so the harness stops assuming a large default window for a model name it doesn't recognise and overflowing a much smaller real one. It's deliberately never probed for a model that isn't @@ -104,15 +104,28 @@ finished loading would just be confusing to leave sitting there. "Detected a custom API key" prompt once would — without it, that prompt would otherwise reappear on every single launch with nobody there to answer it. +**If a model's real context is too small for Claude Code to even get started, you get a +warning instead of a confusing failure.** Claude Code's own system prompt and tools take up +roughly 40K tokens on their own, before you've typed anything — a small local model with a +smaller real context than that fails outright on the very first message, no matter what +context size Codeman tells it to expect (raising the declared context only changes when +Claude Code trims _conversation history_, and there is none yet on message one). Picking +such a model now shows an in-app dialog naming the model, its discovered context and what's +needed, before anything launches or restarts, with the fix spelled out: reconfigure +llama-swap to give that model (or a smaller one) an explicit larger context instead of +relying on auto-fit (`--fit-ctx`), which sizes the context around fitting the biggest model +rather than the biggest context — for example adding `-c 65536` to that model's llama-swap +entry. "Launch anyway" is still there if you want to try regardless. + ## Which harnesses actually work -| Harness | Status | -| ------- | ------ | -| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. | -| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. | -| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. | -| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. | -| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. | +| Harness | Status | +| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | +| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. | +| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. | +| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. | +| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. | +| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. | Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry, not a fixed list here, so this table can go stale before this page does — a greyed-out or diff --git a/src/web/public/index.html b/src/web/public/index.html index 9d2e0dc7..4e847eee 100644 --- a/src/web/public/index.html +++ b/src/web/public/index.html @@ -948,6 +948,29 @@ + + +