Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.
- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
registry entry declares contextLengthVar (currently only claude) and
the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
(40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
quick-start customModel path) check this before the swap-conflict
check and before launching/restarting anything, returning
{requiresContextWarning, modelId, contextLength, minSafeContextTokens}
— skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
_resolveContextWarningConfirm (session-ui.js), wired into both
_quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
(the path Claude actually uses) ahead of the swap-confirmation check.
Explains the fix in-modal: give the model an explicit larger -c/
--ctx-size in llama-swap instead of relying on --fit-ctx, which
optimizes for the biggest model that fits rather than the biggest
context.
Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
6.4 KiB
aicodeman
| aicodeman |
|---|
| minor |
Custom model endpoints: Run-menu picker, and hardening from real llama-swap validation (#430, follow-up to #393's HTTP-API-only cut). With Custom model endpoints on (App Settings → Models) and at least one saved endpoint carrying a discovered model, the Run dropdown grows a Custom Endpoints section generated live off the CLI registry's own capabilities.customModelInjection — one entry per (harness that can redirect to a custom endpoint, saved endpoint). Picking one launches that harness and applies the endpoint to it; with two or more discovered models a small, scrollable dialog asks which one first, the endpoint's defaultModelId marked but never auto-chosen. Endpoints also now re-discover themselves automatically every 5 minutes in the background, one unreachable endpoint never blocking the others.
Everything below was found and fixed against a real llama-swap server, not just unit tests:
- Session-busy false refusal. A freshly launched CLI reports itself
busyfor its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying. - Errors and confirmations you can actually read. Toasts now default to sticky with a close button (errors always were meant to stay, but a fixed 3s timer silently hid them); a failed apply's real server-side reason (not a generic message) reaches the toast.
- "Both claude.ai and ANTHROPIC_API_KEY set" warning. A custom-model Claude session now runs with an isolated
CLAUDE_CONFIG_DIR(empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login —projectsis symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on every launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would. - Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length and applies it as
CLAUDE_CODE_MAX_CONTEXT_TOKENS— sourced primarily from llama-swap's ownGET /running, whosecmdfield carries the launch flags (--fit-ctx/-c/--ctx-size) actually in effect, sinceGET /props'sn_ctxwas confirmed live to report the model's theoretical/trained maximum rather than the real--fit-ctx-shrunk runtime context (a 154112-vs-16384 discrepancy, caught only because the fixed value still overflowed) —/propsis now a fallback for a plain llama.cpp server with no/runningat all. - Context floor too small for Claude Code to even start. Fixing the overflow above surfaced a second, unfixable-by-injection failure: Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on their own (confirmed live via an
in:0 out:0failure on the very first message), which can exceed a small model's entire real context before any conversation history exists to trim — noCLAUDE_CODE_MAX_CONTEXT_TOKENSvalue fixes that, since it only governs when history gets compacted. Applying such a model now returns a warning (gated on the CLI registry declaring acontextLengthVar, so it's a no-op for every other harness) instead of launching straight into a guaranteed first-message failure, and the Run-menu picker shows it as an in-app dialog naming the model, its discovered context and the ~40K safe floor, with the actual fix spelled out: give the model an explicit larger-c/--ctx-sizein llama-swap's config instead of relying on auto-fit, which optimizes for the biggest model that fits rather than the biggest context. "Launch anyway" is still one click away. - The real root cause of "it still says opus, not my model." llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on demand, which can take anywhere from a few seconds to well over a minute — long enough that a session mid-swap is indistinguishable from one that never left the native backend. Applying a selection now checks llama-swap's own
GET /runningfirst (feature-detected; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is never checked); if switching would unload a model another live session is actively using, the apply is refused with a warning naming that session instead of silently switching, and a confirmation retry proceeds anyway. Either way, a sticky "loading model…" toast now covers the actual swap window until llama-swap reports the target model ready, so a prompt sent mid-swap reads as "loading," never as silence or an answer from whatever was loaded a moment before.
Remote (SSH) and Docker sessions are refused for now (400) — their restart reattaches the durable remote/in-container tmux rather than relaunching the agent.
One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all. Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). POST /api/quick-start now accepts a customModel field and computes the same injection before the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own --resume-based restart is far less jarring, and runClaude()'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work).