mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-04 22:49:41 +02:00
feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.
- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
registry entry declares contextLengthVar (currently only claude) and
the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
(40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
quick-start customModel path) check this before the swap-conflict
check and before launching/restarting anything, returning
{requiresContextWarning, modelId, contextLength, minSafeContextTokens}
— skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
_resolveContextWarningConfirm (session-ui.js), wired into both
_quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
(the path Claude actually uses) ahead of the swap-confirmation check.
Explains the fix in-modal: give the model an explicit larger -c/
--ctx-size in llama-swap instead of relying on --fit-ctx, which
optimizes for the biggest model that fits rather than the biggest
context.
Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
993710263d
commit
b45a96358e
@@ -239,8 +239,17 @@ registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
|
||||
model id and never compacts, which reliably overflows a much smaller real
|
||||
local context — confirmed live: a stock ~33.7K-token system prompt against
|
||||
a 16384-token llama-swap model failed with `exceeds the available context
|
||||
size`. No entry for the model in `modelContextLengths` means the var is
|
||||
simply omitted, never a guess.
|
||||
size`. No entry for the model in `modelContextLengths` means the var is
|
||||
simply omitted, never a guess. ⚠️ **This var only affects when Claude
|
||||
Code compacts conversation _history_ — it cannot fix a model whose real
|
||||
context is smaller than Claude Code's own fixed per-turn overhead**
|
||||
(system prompt + tool schemas, empirically ~36.4K tokens, confirmed live
|
||||
via an `in:0 out:0` failure on the very first message, before any
|
||||
history exists to compact). No context-length declaration changes that
|
||||
fixed overhead, so a model below the safe floor fails outright on
|
||||
message one regardless of what this var says. See "Context-window floor
|
||||
warning" below for how Codeman catches this case before launching
|
||||
instead of after.
|
||||
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
|
||||
the `configDir`-kind CLIs use (empty, no files written into it), so the
|
||||
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
|
||||
@@ -261,7 +270,7 @@ registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
|
||||
**That isolated directory needed one more fix to actually be usable
|
||||
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
|
||||
real profile's prior "Detected a custom API key — use it?" approvals, so
|
||||
without more, Claude Code stops and asks that on *every single launch* —
|
||||
without more, Claude Code stops and asks that on _every single launch_ —
|
||||
confirmed live, and with nobody at a TTY to answer, its own default answer
|
||||
("No") silently refuses the very key this feature just injected, which
|
||||
looks like the endpoint being ignored entirely. `customModelInjection`'s
|
||||
@@ -285,7 +294,7 @@ which can take anywhere from a few seconds to well over a minute:
|
||||
one-shot `POST /api/quick-start` above) call llama-swap's own
|
||||
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
|
||||
compatible server (no such endpoint) is simply never checked. If a
|
||||
*different* model is currently loaded and ready, and another **live
|
||||
_different_ model is currently loaded and ready, and another **live
|
||||
session's own selection** is using it, the apply returns
|
||||
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
|
||||
instead of silently switching — nothing is applied or created yet.
|
||||
@@ -298,13 +307,51 @@ which can take anywhere from a few seconds to well over a minute:
|
||||
reached llama-swap at all (nothing in its own server logs), since nothing
|
||||
had actually asked it to load anything yet. Both apply routes now also
|
||||
send the smallest real request that will — `POST <baseUrl>/v1/chat/
|
||||
completions` with `max_tokens: 1` and one throwaway message — whenever the
|
||||
completions` with `max_tokens: 1` and one throwaway message — whenever the
|
||||
target model isn't already the one loaded and ready, fire-and-forget (its
|
||||
response is never read; `GET /api/model-endpoints/:id/running-status`,
|
||||
polled client-side, is what actually confirms readiness). The response
|
||||
also carries `modelSwapInProgress: true` in that case, which is what
|
||||
drives the Run-menu picker's own "loading model" status banner.
|
||||
|
||||
## Context-window floor warning
|
||||
|
||||
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
|
||||
empirically ~36.4K tokens) can exceed a small local model's _entire_ real
|
||||
context on its own, before any conversation history exists to fill it —
|
||||
confirmed live twice, both as an `in:0 out:0` failure on the very first
|
||||
message sent. `CLAUDE_CODE_MAX_CONTEXT_TOKENS` (above) cannot fix this: it
|
||||
only governs when Claude Code compacts conversation history, and there is
|
||||
no history yet on message one. Applying such a model would look like the
|
||||
endpoint being ignored, or the wrong model being used, when in fact the
|
||||
endpoint applied correctly and the model is simply too small for this CLI.
|
||||
|
||||
Both apply routes (the restart route and the one-shot `POST
|
||||
/api/quick-start`) now check for this **before** launching or restarting
|
||||
anything, gated on the CLI's registry entry declaring a `contextLengthVar`
|
||||
(currently only claude — the check is a no-op for every other CLI by
|
||||
construction, never a hardcoded mode check). If the model's discovered
|
||||
context (`modelContextLengths`, from discovery above) is below
|
||||
`CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000, comfortably above the measured
|
||||
~36.4K overhead), the response is `{requiresContextWarning: true, modelId,
|
||||
contextLength, minSafeContextTokens}` instead of applying — nothing is
|
||||
restarted or created yet. A context length that was never discovered at
|
||||
all skips the check entirely (nothing to compare, so it fails open rather
|
||||
than warning on every model an endpoint hasn't reported a size for).
|
||||
Retrying with `confirmed: true` launches anyway.
|
||||
|
||||
The Run-menu picker shows this as an in-app modal
|
||||
(`#customModelContextWarningModal`, matching the llama-swap conflict
|
||||
modal's look) naming the model, its discovered context, and the safe
|
||||
floor, and explaining the fix: reconfigure llama-swap to give that model
|
||||
(or a smaller one) an explicit larger context instead of relying on
|
||||
auto-fit (`--fit-ctx`), which optimizes for the biggest _model_ that fits
|
||||
rather than the biggest _context_ — e.g. adding `-c 65536` (or as large a
|
||||
`--ctx-size` as the hardware holds) to that model's llama-swap config
|
||||
entry. A smaller model at a much larger explicit context often fits in
|
||||
the same VRAM a bigger model's auto-fit context gets shrunk to make room
|
||||
for.
|
||||
|
||||
Clear back to the harness's native cloud default with:
|
||||
|
||||
```bash
|
||||
|
||||
@@ -28,7 +28,7 @@ shows up without another manual click of **Discover**. One endpoint being unreac
|
||||
given cycle (powered off, wrong network) never blocks the others from refreshing.
|
||||
|
||||
**Context length is picked up automatically where it can be, safely.** Against a
|
||||
llama.cpp/llama-swap server, discovery also learns each *currently loaded* model's real
|
||||
llama.cpp/llama-swap server, discovery also learns each _currently loaded_ model's real
|
||||
context window and applies it to the launched session (Claude Code today — see below), so
|
||||
the harness stops assuming a large default window for a model name it doesn't recognise and
|
||||
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
|
||||
@@ -104,15 +104,28 @@ finished loading would just be confusing to leave sitting there.
|
||||
"Detected a custom API key" prompt once would — without it, that prompt would otherwise
|
||||
reappear on every single launch with nobody there to answer it.
|
||||
|
||||
**If a model's real context is too small for Claude Code to even get started, you get a
|
||||
warning instead of a confusing failure.** Claude Code's own system prompt and tools take up
|
||||
roughly 40K tokens on their own, before you've typed anything — a small local model with a
|
||||
smaller real context than that fails outright on the very first message, no matter what
|
||||
context size Codeman tells it to expect (raising the declared context only changes when
|
||||
Claude Code trims _conversation history_, and there is none yet on message one). Picking
|
||||
such a model now shows an in-app dialog naming the model, its discovered context and what's
|
||||
needed, before anything launches or restarts, with the fix spelled out: reconfigure
|
||||
llama-swap to give that model (or a smaller one) an explicit larger context instead of
|
||||
relying on auto-fit (`--fit-ctx`), which sizes the context around fitting the biggest model
|
||||
rather than the biggest context — for example adding `-c 65536` to that model's llama-swap
|
||||
entry. "Launch anyway" is still there if you want to try regardless.
|
||||
|
||||
## Which harnesses actually work
|
||||
|
||||
| Harness | Status |
|
||||
| ------- | ------ |
|
||||
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
|
||||
| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. |
|
||||
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
|
||||
| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. |
|
||||
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
|
||||
| Harness | Status |
|
||||
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
|
||||
| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. |
|
||||
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
|
||||
| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. |
|
||||
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
|
||||
|
||||
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
|
||||
not a fixed list here, so this table can go stale before this page does — a greyed-out or
|
||||
|
||||
Reference in New Issue
Block a user