mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-09 00:49:41 +02:00
feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.
- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
registry entry declares contextLengthVar (currently only claude) and
the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
(40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
quick-start customModel path) check this before the swap-conflict
check and before launching/restarting anything, returning
{requiresContextWarning, modelId, contextLength, minSafeContextTokens}
— skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
_resolveContextWarningConfirm (session-ui.js), wired into both
_quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
(the path Claude actually uses) ahead of the swap-confirmation check.
Explains the fix in-modal: give the model an explicit larger -c/
--ctx-size in llama-swap instead of relying on --fit-ctx, which
optimizes for the biggest model that fits rather than the biggest
context.
Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
993710263d
commit
b45a96358e
@@ -1,5 +1,5 @@
|
|||||||
---
|
---
|
||||||
"aicodeman": minor
|
'aicodeman': minor
|
||||||
---
|
---
|
||||||
|
|
||||||
**Custom model endpoints: Run-menu picker, and hardening from real llama-swap validation** (#430, follow-up to #393's HTTP-API-only cut). With **Custom model endpoints** on (App Settings → Models) and at least one saved endpoint carrying a discovered model, the Run dropdown grows a **Custom Endpoints** section generated live off the CLI registry's own `capabilities.customModelInjection` — one entry per (harness that can redirect to a custom endpoint, saved endpoint). Picking one launches that harness and applies the endpoint to it; with two or more discovered models a small, scrollable dialog asks which one first, the endpoint's `defaultModelId` marked but never auto-chosen. Endpoints also now re-discover themselves automatically every 5 minutes in the background, one unreachable endpoint never blocking the others.
|
**Custom model endpoints: Run-menu picker, and hardening from real llama-swap validation** (#430, follow-up to #393's HTTP-API-only cut). With **Custom model endpoints** on (App Settings → Models) and at least one saved endpoint carrying a discovered model, the Run dropdown grows a **Custom Endpoints** section generated live off the CLI registry's own `capabilities.customModelInjection` — one entry per (harness that can redirect to a custom endpoint, saved endpoint). Picking one launches that harness and applies the endpoint to it; with two or more discovered models a small, scrollable dialog asks which one first, the endpoint's `defaultModelId` marked but never auto-chosen. Endpoints also now re-discover themselves automatically every 5 minutes in the background, one unreachable endpoint never blocking the others.
|
||||||
@@ -8,10 +8,11 @@ Everything below was found and fixed against a **real llama-swap server**, not j
|
|||||||
|
|
||||||
- **Session-busy false refusal.** A freshly launched CLI reports itself `busy` for its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying.
|
- **Session-busy false refusal.** A freshly launched CLI reports itself `busy` for its own startup (spinner, workspace-trust check) well before the apply call would reach it, and the apply route correctly refuses to restart a session mid-turn — indistinguishable from a fresh boot. The picker now waits for the new session to go idle (bounded at 20s, never an error on timeout) before applying.
|
||||||
- **Errors and confirmations you can actually read.** Toasts now default to sticky with a close button (errors always were meant to stay, but a fixed 3s timer silently hid them); a failed apply's real server-side reason (not a generic message) reaches the toast.
|
- **Errors and confirmations you can actually read.** Toasts now default to sticky with a close button (errors always were meant to stay, but a fixed 3s timer silently hid them); a failed apply's real server-side reason (not a generic message) reaches the toast.
|
||||||
- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on *every* launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would.
|
- **"Both claude.ai and ANTHROPIC_API_KEY set" warning.** A custom-model Claude session now runs with an isolated `CLAUDE_CONFIG_DIR` (empty, no real credentials in it) so the injected API key never coexists with a stored OAuth login — `projects` is symlinked back to the real config dir so the response viewer/subagent windows/Read My Mind keep working. That isolated, otherwise-empty directory has none of a real profile's prior "Detected a custom API key — use it?" approvals either, which would otherwise re-ask on _every_ launch with nobody at a TTY to answer (and silently refuse the key on its own default); the apply step now pre-seeds that exact approval field the same way answering the prompt once by hand would.
|
||||||
- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length from llama.cpp/llama-swap's `GET /props?model=`, but **only** for a model llama-swap's own `/v1/models` response already reports loaded — never an unloaded one, since asking about one risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery — and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS`.
|
- **Context-window overflow.** Claude Code assumes a large default context window for a model id it doesn't recognize and never compacts, so a real local model's much smaller context silently overflowed (confirmed live: a stock ~33.7K-token system prompt against a 16384-token model). Discovery now also learns each model's real context length and applies it as `CLAUDE_CODE_MAX_CONTEXT_TOKENS` — sourced primarily from llama-swap's own `GET /running`, whose `cmd` field carries the launch flags (`--fit-ctx`/`-c`/`--ctx-size`) actually in effect, since `GET /props`'s `n_ctx` was confirmed live to report the model's theoretical/trained maximum rather than the real `--fit-ctx`-shrunk runtime context (a 154112-vs-16384 discrepancy, caught only because the fixed value still overflowed) — `/props` is now a fallback for a plain llama.cpp server with no `/running` at all.
|
||||||
|
- **Context floor too small for Claude Code to even start.** Fixing the overflow above surfaced a second, unfixable-by-injection failure: Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on their own (confirmed live via an `in:0 out:0` failure on the very first message), which can exceed a small model's entire real context before any conversation history exists to trim — no `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since it only governs when history gets compacted. Applying such a model now returns a warning (gated on the CLI registry declaring a `contextLengthVar`, so it's a no-op for every other harness) instead of launching straight into a guaranteed first-message failure, and the Run-menu picker shows it as an in-app dialog naming the model, its discovered context and the ~40K safe floor, with the actual fix spelled out: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on auto-fit, which optimizes for the biggest model that fits rather than the biggest context. "Launch anyway" is still one click away.
|
||||||
- **The real root cause of "it still says opus, not my model."** llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on demand, which can take anywhere from a few seconds to well over a minute — long enough that a session mid-swap is indistinguishable from one that never left the native backend. Applying a selection now checks llama-swap's own `GET /running` first (feature-detected; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is never checked); if switching would unload a model **another live session is actively using**, the apply is refused with a warning naming that session instead of silently switching, and a confirmation retry proceeds anyway. Either way, a sticky "loading model…" toast now covers the actual swap window until llama-swap reports the target model ready, so a prompt sent mid-swap reads as "loading," never as silence or an answer from whatever was loaded a moment before.
|
- **The real root cause of "it still says opus, not my model."** llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on demand, which can take anywhere from a few seconds to well over a minute — long enough that a session mid-swap is indistinguishable from one that never left the native backend. Applying a selection now checks llama-swap's own `GET /running` first (feature-detected; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is never checked); if switching would unload a model **another live session is actively using**, the apply is refused with a warning naming that session instead of silently switching, and a confirmation retry proceeds anyway. Either way, a sticky "loading model…" toast now covers the actual swap window until llama-swap reports the target model ready, so a prompt sent mid-swap reads as "loading," never as silence or an answer from whatever was loaded a moment before.
|
||||||
|
|
||||||
Remote (SSH) and Docker sessions are refused for now (400) — their restart reattaches the durable remote/in-container tmux rather than relaunching the agent.
|
Remote (SSH) and Docker sessions are refused for now (400) — their restart reattaches the durable remote/in-container tmux rather than relaunching the agent.
|
||||||
|
|
||||||
**One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all.** Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). `POST /api/quick-start` now accepts a `customModel` field and computes the same injection *before* the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own `--resume`-based restart is far less jarring, and `runClaude()`'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work).
|
**One more, from watching it launch live: opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP now launch directly on the endpoint, with no restart at all.** Picking one of these seven from the Run-menu picker used to launch natively first, wait for it to settle, then restart it in place with the endpoint applied — a deliberate two-step design, but visibly a native boot immediately followed by a second one, worst on a CLI whose TUI fully reinitializes on a restart (confirmed live on Codex). `POST /api/quick-start` now accepts a `customModel` field and computes the same injection _before_ the session exists, launching straight onto the endpoint the first time — no visible relaunch, and it also runs the same llama-swap conflict check (warns before unloading a model another live session is using) at create time. Claude still uses the original launch-then-restart path for now (its own `--resume`-based restart is far less jarring, and `runClaude()`'s multi-tab and docker-config-drift-retry logic make folding it into the one-shot path separate work).
|
||||||
|
|||||||
@@ -239,8 +239,17 @@ registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
|
|||||||
model id and never compacts, which reliably overflows a much smaller real
|
model id and never compacts, which reliably overflows a much smaller real
|
||||||
local context — confirmed live: a stock ~33.7K-token system prompt against
|
local context — confirmed live: a stock ~33.7K-token system prompt against
|
||||||
a 16384-token llama-swap model failed with `exceeds the available context
|
a 16384-token llama-swap model failed with `exceeds the available context
|
||||||
size`. No entry for the model in `modelContextLengths` means the var is
|
size`. No entry for the model in `modelContextLengths` means the var is
|
||||||
simply omitted, never a guess.
|
simply omitted, never a guess. ⚠️ **This var only affects when Claude
|
||||||
|
Code compacts conversation _history_ — it cannot fix a model whose real
|
||||||
|
context is smaller than Claude Code's own fixed per-turn overhead**
|
||||||
|
(system prompt + tool schemas, empirically ~36.4K tokens, confirmed live
|
||||||
|
via an `in:0 out:0` failure on the very first message, before any
|
||||||
|
history exists to compact). No context-length declaration changes that
|
||||||
|
fixed overhead, so a model below the safe floor fails outright on
|
||||||
|
message one regardless of what this var says. See "Context-window floor
|
||||||
|
warning" below for how Codeman catches this case before launching
|
||||||
|
instead of after.
|
||||||
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
|
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
|
||||||
the `configDir`-kind CLIs use (empty, no files written into it), so the
|
the `configDir`-kind CLIs use (empty, no files written into it), so the
|
||||||
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
|
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
|
||||||
@@ -261,7 +270,7 @@ registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
|
|||||||
**That isolated directory needed one more fix to actually be usable
|
**That isolated directory needed one more fix to actually be usable
|
||||||
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
|
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
|
||||||
real profile's prior "Detected a custom API key — use it?" approvals, so
|
real profile's prior "Detected a custom API key — use it?" approvals, so
|
||||||
without more, Claude Code stops and asks that on *every single launch* —
|
without more, Claude Code stops and asks that on _every single launch_ —
|
||||||
confirmed live, and with nobody at a TTY to answer, its own default answer
|
confirmed live, and with nobody at a TTY to answer, its own default answer
|
||||||
("No") silently refuses the very key this feature just injected, which
|
("No") silently refuses the very key this feature just injected, which
|
||||||
looks like the endpoint being ignored entirely. `customModelInjection`'s
|
looks like the endpoint being ignored entirely. `customModelInjection`'s
|
||||||
@@ -285,7 +294,7 @@ which can take anywhere from a few seconds to well over a minute:
|
|||||||
one-shot `POST /api/quick-start` above) call llama-swap's own
|
one-shot `POST /api/quick-start` above) call llama-swap's own
|
||||||
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
|
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
|
||||||
compatible server (no such endpoint) is simply never checked. If a
|
compatible server (no such endpoint) is simply never checked. If a
|
||||||
*different* model is currently loaded and ready, and another **live
|
_different_ model is currently loaded and ready, and another **live
|
||||||
session's own selection** is using it, the apply returns
|
session's own selection** is using it, the apply returns
|
||||||
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
|
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
|
||||||
instead of silently switching — nothing is applied or created yet.
|
instead of silently switching — nothing is applied or created yet.
|
||||||
@@ -298,13 +307,51 @@ which can take anywhere from a few seconds to well over a minute:
|
|||||||
reached llama-swap at all (nothing in its own server logs), since nothing
|
reached llama-swap at all (nothing in its own server logs), since nothing
|
||||||
had actually asked it to load anything yet. Both apply routes now also
|
had actually asked it to load anything yet. Both apply routes now also
|
||||||
send the smallest real request that will — `POST <baseUrl>/v1/chat/
|
send the smallest real request that will — `POST <baseUrl>/v1/chat/
|
||||||
completions` with `max_tokens: 1` and one throwaway message — whenever the
|
completions` with `max_tokens: 1` and one throwaway message — whenever the
|
||||||
target model isn't already the one loaded and ready, fire-and-forget (its
|
target model isn't already the one loaded and ready, fire-and-forget (its
|
||||||
response is never read; `GET /api/model-endpoints/:id/running-status`,
|
response is never read; `GET /api/model-endpoints/:id/running-status`,
|
||||||
polled client-side, is what actually confirms readiness). The response
|
polled client-side, is what actually confirms readiness). The response
|
||||||
also carries `modelSwapInProgress: true` in that case, which is what
|
also carries `modelSwapInProgress: true` in that case, which is what
|
||||||
drives the Run-menu picker's own "loading model" status banner.
|
drives the Run-menu picker's own "loading model" status banner.
|
||||||
|
|
||||||
|
## Context-window floor warning
|
||||||
|
|
||||||
|
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
|
||||||
|
empirically ~36.4K tokens) can exceed a small local model's _entire_ real
|
||||||
|
context on its own, before any conversation history exists to fill it —
|
||||||
|
confirmed live twice, both as an `in:0 out:0` failure on the very first
|
||||||
|
message sent. `CLAUDE_CODE_MAX_CONTEXT_TOKENS` (above) cannot fix this: it
|
||||||
|
only governs when Claude Code compacts conversation history, and there is
|
||||||
|
no history yet on message one. Applying such a model would look like the
|
||||||
|
endpoint being ignored, or the wrong model being used, when in fact the
|
||||||
|
endpoint applied correctly and the model is simply too small for this CLI.
|
||||||
|
|
||||||
|
Both apply routes (the restart route and the one-shot `POST
|
||||||
|
/api/quick-start`) now check for this **before** launching or restarting
|
||||||
|
anything, gated on the CLI's registry entry declaring a `contextLengthVar`
|
||||||
|
(currently only claude — the check is a no-op for every other CLI by
|
||||||
|
construction, never a hardcoded mode check). If the model's discovered
|
||||||
|
context (`modelContextLengths`, from discovery above) is below
|
||||||
|
`CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000, comfortably above the measured
|
||||||
|
~36.4K overhead), the response is `{requiresContextWarning: true, modelId,
|
||||||
|
contextLength, minSafeContextTokens}` instead of applying — nothing is
|
||||||
|
restarted or created yet. A context length that was never discovered at
|
||||||
|
all skips the check entirely (nothing to compare, so it fails open rather
|
||||||
|
than warning on every model an endpoint hasn't reported a size for).
|
||||||
|
Retrying with `confirmed: true` launches anyway.
|
||||||
|
|
||||||
|
The Run-menu picker shows this as an in-app modal
|
||||||
|
(`#customModelContextWarningModal`, matching the llama-swap conflict
|
||||||
|
modal's look) naming the model, its discovered context, and the safe
|
||||||
|
floor, and explaining the fix: reconfigure llama-swap to give that model
|
||||||
|
(or a smaller one) an explicit larger context instead of relying on
|
||||||
|
auto-fit (`--fit-ctx`), which optimizes for the biggest _model_ that fits
|
||||||
|
rather than the biggest _context_ — e.g. adding `-c 65536` (or as large a
|
||||||
|
`--ctx-size` as the hardware holds) to that model's llama-swap config
|
||||||
|
entry. A smaller model at a much larger explicit context often fits in
|
||||||
|
the same VRAM a bigger model's auto-fit context gets shrunk to make room
|
||||||
|
for.
|
||||||
|
|
||||||
Clear back to the harness's native cloud default with:
|
Clear back to the harness's native cloud default with:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
|
|||||||
@@ -28,7 +28,7 @@ shows up without another manual click of **Discover**. One endpoint being unreac
|
|||||||
given cycle (powered off, wrong network) never blocks the others from refreshing.
|
given cycle (powered off, wrong network) never blocks the others from refreshing.
|
||||||
|
|
||||||
**Context length is picked up automatically where it can be, safely.** Against a
|
**Context length is picked up automatically where it can be, safely.** Against a
|
||||||
llama.cpp/llama-swap server, discovery also learns each *currently loaded* model's real
|
llama.cpp/llama-swap server, discovery also learns each _currently loaded_ model's real
|
||||||
context window and applies it to the launched session (Claude Code today — see below), so
|
context window and applies it to the launched session (Claude Code today — see below), so
|
||||||
the harness stops assuming a large default window for a model name it doesn't recognise and
|
the harness stops assuming a large default window for a model name it doesn't recognise and
|
||||||
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
|
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
|
||||||
@@ -104,15 +104,28 @@ finished loading would just be confusing to leave sitting there.
|
|||||||
"Detected a custom API key" prompt once would — without it, that prompt would otherwise
|
"Detected a custom API key" prompt once would — without it, that prompt would otherwise
|
||||||
reappear on every single launch with nobody there to answer it.
|
reappear on every single launch with nobody there to answer it.
|
||||||
|
|
||||||
|
**If a model's real context is too small for Claude Code to even get started, you get a
|
||||||
|
warning instead of a confusing failure.** Claude Code's own system prompt and tools take up
|
||||||
|
roughly 40K tokens on their own, before you've typed anything — a small local model with a
|
||||||
|
smaller real context than that fails outright on the very first message, no matter what
|
||||||
|
context size Codeman tells it to expect (raising the declared context only changes when
|
||||||
|
Claude Code trims _conversation history_, and there is none yet on message one). Picking
|
||||||
|
such a model now shows an in-app dialog naming the model, its discovered context and what's
|
||||||
|
needed, before anything launches or restarts, with the fix spelled out: reconfigure
|
||||||
|
llama-swap to give that model (or a smaller one) an explicit larger context instead of
|
||||||
|
relying on auto-fit (`--fit-ctx`), which sizes the context around fitting the biggest model
|
||||||
|
rather than the biggest context — for example adding `-c 65536` to that model's llama-swap
|
||||||
|
entry. "Launch anyway" is still there if you want to try regardless.
|
||||||
|
|
||||||
## Which harnesses actually work
|
## Which harnesses actually work
|
||||||
|
|
||||||
| Harness | Status |
|
| Harness | Status |
|
||||||
| ------- | ------ |
|
| ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||||
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
|
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
|
||||||
| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. |
|
| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. |
|
||||||
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
|
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
|
||||||
| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. |
|
| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. |
|
||||||
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
|
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
|
||||||
|
|
||||||
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
|
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
|
||||||
not a fixed list here, so this table can go stale before this page does — a greyed-out or
|
not a fixed list here, so this table can go stale before this page does — a greyed-out or
|
||||||
|
|||||||
@@ -948,6 +948,29 @@
|
|||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
|
|
||||||
|
<!-- Custom Model Endpoint Profiles: context-window-too-small warning
|
||||||
|
(docs/custom-model-endpoints-plan.md) — shown before launching a CLI
|
||||||
|
whose own fixed system-prompt/tool-schema overhead exceeds the
|
||||||
|
model's real discovered context, which guarantees a first-message
|
||||||
|
failure regardless of CLAUDE_CODE_MAX_CONTEXT_TOKENS. See
|
||||||
|
_confirmContextWarning() in session-ui.js. -->
|
||||||
|
<div class="modal" id="customModelContextWarningModal">
|
||||||
|
<div class="modal-backdrop" onclick="app._resolveContextWarningConfirm(false)"></div>
|
||||||
|
<div class="modal-content modal-sm">
|
||||||
|
<div class="modal-header">
|
||||||
|
<h3>Context window too small</h3>
|
||||||
|
<button class="modal-close" onclick="app._resolveContextWarningConfirm(false)" aria-label="Cancel">×</button>
|
||||||
|
</div>
|
||||||
|
<div class="modal-body">
|
||||||
|
<p class="form-hint" id="customModelContextWarningMessage" style="white-space: pre-wrap;"></p>
|
||||||
|
</div>
|
||||||
|
<div class="modal-footer">
|
||||||
|
<button class="btn-toolbar" onclick="app._resolveContextWarningConfirm(false)">Cancel</button>
|
||||||
|
<button class="btn-toolbar btn-primary" onclick="app._resolveContextWarningConfirm(true)">Launch anyway</button>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
<!-- Cron Jobs Modal -->
|
<!-- Cron Jobs Modal -->
|
||||||
<div class="modal" id="cronModal">
|
<div class="modal" id="cronModal">
|
||||||
<div class="modal-backdrop" onclick="app.closeCron()"></div>
|
<div class="modal-backdrop" onclick="app.closeCron()"></div>
|
||||||
|
|||||||
@@ -693,6 +693,45 @@ Object.assign(CodemanApp.prototype, {
|
|||||||
resolve?.(proceed);
|
resolve?.(proceed);
|
||||||
},
|
},
|
||||||
|
|
||||||
|
/**
|
||||||
|
* In-app warning shown when the apply route reports `requiresContextWarning`: this
|
||||||
|
* model's real discovered context is smaller than the CLI's own fixed system-prompt/
|
||||||
|
* tool-schema overhead, which guarantees the very first message fails outright — no
|
||||||
|
* `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes that, since there is no conversation
|
||||||
|
* history yet for compaction to trim. Same promise-based pattern as
|
||||||
|
* `_confirmModelSwap`; `_resolveContextWarningConfirm` settles it.
|
||||||
|
*/
|
||||||
|
_confirmContextWarning(modelId, contextLength, minSafeContextTokens) {
|
||||||
|
const modal = document.getElementById('customModelContextWarningModal');
|
||||||
|
const messageEl = document.getElementById('customModelContextWarningMessage');
|
||||||
|
if (messageEl) {
|
||||||
|
const known = typeof contextLength === 'number';
|
||||||
|
messageEl.textContent =
|
||||||
|
`${modelId} is configured with ` +
|
||||||
|
(known ? `only ${contextLength.toLocaleString()} tokens of` : 'an unknown (too small)') +
|
||||||
|
` context, but this CLI needs roughly ${minSafeContextTokens.toLocaleString()}+ tokens just for its own ` +
|
||||||
|
`system prompt and tools — before any conversation history. Its very first message will fail outright, ` +
|
||||||
|
`no matter what context size Codeman tells it to expect.\n\n` +
|
||||||
|
`To fix this, reconfigure llama-swap to give this model (or a smaller one) an explicit larger context ` +
|
||||||
|
`instead of relying on auto-fit (--fit-ctx), which optimizes for the biggest MODEL that fits, not the ` +
|
||||||
|
`biggest CONTEXT — e.g. add "-c 65536" (or as large a --ctx-size as your hardware holds) to its llama-swap ` +
|
||||||
|
`config entry. A smaller model at a much larger explicit context often fits in the same VRAM a bigger ` +
|
||||||
|
`model's auto-fit context gets shrunk to make room for.`;
|
||||||
|
}
|
||||||
|
modal?.classList.add('active');
|
||||||
|
return new Promise((resolve) => {
|
||||||
|
this._resolveContextWarningConfirmPromise = resolve;
|
||||||
|
});
|
||||||
|
},
|
||||||
|
|
||||||
|
/** Called by the modal's Cancel/Launch-anyway buttons and its backdrop click. */
|
||||||
|
_resolveContextWarningConfirm(proceed) {
|
||||||
|
document.getElementById('customModelContextWarningModal')?.classList.remove('active');
|
||||||
|
const resolve = this._resolveContextWarningConfirmPromise;
|
||||||
|
this._resolveContextWarningConfirmPromise = null;
|
||||||
|
resolve?.(proceed);
|
||||||
|
},
|
||||||
|
|
||||||
/** A model row in the picker modal was clicked: close it and launch with that choice. */
|
/** A model row in the picker modal was clicked: close it and launch with that choice. */
|
||||||
chooseCustomModelAndRun(modelId) {
|
chooseCustomModelAndRun(modelId) {
|
||||||
const pending = this._pendingCustomModelPick;
|
const pending = this._pendingCustomModelPick;
|
||||||
@@ -789,6 +828,15 @@ Object.assign(CodemanApp.prototype, {
|
|||||||
return res.json();
|
return res.json();
|
||||||
};
|
};
|
||||||
let data = await post(bodyObj);
|
let data = await post(bodyObj);
|
||||||
|
if (data?.data?.requiresContextWarning) {
|
||||||
|
const { modelId, contextLength, minSafeContextTokens } = data.data;
|
||||||
|
const proceed = await this._confirmContextWarning(modelId, contextLength, minSafeContextTokens);
|
||||||
|
if (!proceed) {
|
||||||
|
this._lastCustomModelLaunchResult = undefined;
|
||||||
|
return { success: false, error: 'Launch cancelled — context window too small' };
|
||||||
|
}
|
||||||
|
data = await post({ ...bodyObj, customModel: { ...bodyObj.customModel, confirmed: true } });
|
||||||
|
}
|
||||||
if (data?.data?.requiresConfirmation) {
|
if (data?.data?.requiresConfirmation) {
|
||||||
const { currentlyLoadedModel, affectedSessions } = data.data;
|
const { currentlyLoadedModel, affectedSessions } = data.data;
|
||||||
const names = affectedSessions.map((s) => s.name || s.id).join(', ');
|
const names = affectedSessions.map((s) => s.name || s.id).join(', ');
|
||||||
@@ -871,6 +919,26 @@ Object.assign(CodemanApp.prototype, {
|
|||||||
// once `data.success !== false`.
|
// once `data.success !== false`.
|
||||||
let payload = data?.success !== false ? data?.data : undefined;
|
let payload = data?.success !== false ? data?.data : undefined;
|
||||||
|
|
||||||
|
// This CLI's own fixed overhead (system prompt + tool schemas) may exceed the
|
||||||
|
// model's real discovered context outright — no context-length declaration can
|
||||||
|
// fix that, since compaction only trims conversation history and there is none
|
||||||
|
// on message 1. Warn and let the user decide whether to launch anyway, same
|
||||||
|
// confirmed:true re-send pattern as the swap check below.
|
||||||
|
if (ok && payload?.requiresContextWarning) {
|
||||||
|
const proceed = await this._confirmContextWarning(
|
||||||
|
payload.modelId,
|
||||||
|
payload.contextLength,
|
||||||
|
payload.minSafeContextTokens
|
||||||
|
);
|
||||||
|
if (!proceed) {
|
||||||
|
switchingToast?.dismiss();
|
||||||
|
this.showToast('Kept the native backend — context window too small', 'info');
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
({ ok, data, res } = await this._applyCustomModelToSession(sessionId, endpointId, modelId, true));
|
||||||
|
payload = data?.success !== false ? data?.data : undefined;
|
||||||
|
}
|
||||||
|
|
||||||
// llama-swap runs one model at a time: switching would unload it out from under
|
// llama-swap runs one model at a time: switching would unload it out from under
|
||||||
// another session actively using it. The route only asks when that's actually true
|
// another session actively using it. The route only asks when that's actually true
|
||||||
// (never just because a swap is needed at all) — confirming re-sends the exact same
|
// (never just because a swap is needed at all) — confirming re-sends the exact same
|
||||||
|
|||||||
@@ -23,11 +23,46 @@ import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
|
|||||||
import { egressBlockedReason, webviewFetch } from '../webview-egress.js';
|
import { egressBlockedReason, webviewFetch } from '../webview-egress.js';
|
||||||
import { CustomModelHostSchema } from '../schemas.js';
|
import { CustomModelHostSchema } from '../schemas.js';
|
||||||
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../../custom-model-hosts.js';
|
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../../custom-model-hosts.js';
|
||||||
|
import type { CliEntry } from '../../config/cli-registry/types.js';
|
||||||
|
|
||||||
const CODEMAN_CONFIG_DIR = getDataDir();
|
const CODEMAN_CONFIG_DIR = getDataDir();
|
||||||
const DISCOVER_TIMEOUT_MS = 8000;
|
const DISCOVER_TIMEOUT_MS = 8000;
|
||||||
const PROPS_TIMEOUT_MS = 5000;
|
const PROPS_TIMEOUT_MS = 5000;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Claude Code's own system prompt + tool schemas cost roughly this many tokens on EVERY
|
||||||
|
* request, before a single character of conversation history — confirmed live, twice, on
|
||||||
|
* requests reporting `in:0 out:0` (the very first exchange) failing at ~36.4K tokens. No
|
||||||
|
* `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes this: that setting only changes when Claude
|
||||||
|
* Code decides to COMPACT conversation history, and there is no history yet on the first
|
||||||
|
* message for it to trim. A model whose real context is below this floor will refuse
|
||||||
|
* Claude Code's very first message outright, unconditionally.
|
||||||
|
*
|
||||||
|
* Set well above the ~36.4K actually measured — CLAUDE.md size, active MCP servers, and
|
||||||
|
* enabled skills all add to a project's real baseline, so the observed figure is a floor
|
||||||
|
* for THAT one workspace, not a ceiling for every one. Erring conservative here means a
|
||||||
|
* borderline-safe model still gets warned about (the user can launch anyway), rather than
|
||||||
|
* this floor missing a genuinely-too-small one because a smaller test project happened to
|
||||||
|
* fit.
|
||||||
|
*/
|
||||||
|
export const CLAUDE_MIN_SAFE_CONTEXT_TOKENS = 40000;
|
||||||
|
|
||||||
|
/**
|
||||||
|
* True when applying this model to this CLI is heading for a guaranteed first-message
|
||||||
|
* failure per `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` above. Gated on `contextLengthVar` (today,
|
||||||
|
* only claude's registry entry declares one) rather than a hardcoded mode check: a CLI
|
||||||
|
* with a small enough baseline of its own to never trip this would have no reason to
|
||||||
|
* declare the field in the first place, so the check simply never applies to it.
|
||||||
|
*/
|
||||||
|
export function exceedsSafeContextFloor(
|
||||||
|
entry: Pick<CliEntry, 'capabilities'>,
|
||||||
|
contextLength: number | undefined
|
||||||
|
): boolean {
|
||||||
|
const cap = entry.capabilities.customModelInjection;
|
||||||
|
if (cap.kind !== 'env' || !cap.contextLengthVar) return false;
|
||||||
|
return typeof contextLength === 'number' && contextLength < CLAUDE_MIN_SAFE_CONTEXT_TOKENS;
|
||||||
|
}
|
||||||
|
|
||||||
function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse<never> | null {
|
function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse<never> | null {
|
||||||
if (!isMultiUserMode() || isAdmin(req)) return null;
|
if (!isMultiUserMode() || isAdmin(req)) return null;
|
||||||
reply.code(403);
|
reply.code(403);
|
||||||
|
|||||||
@@ -55,7 +55,12 @@ import {
|
|||||||
} from '../schemas.js';
|
} from '../schemas.js';
|
||||||
import { readCustomModelHosts } from '../../custom-model-hosts.js';
|
import { readCustomModelHosts } from '../../custom-model-hosts.js';
|
||||||
import { applyCustomModelInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
|
import { applyCustomModelInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
|
||||||
import { getLlamaSwapStatus, triggerLlamaSwapLoad } from './custom-model-routes.js';
|
import {
|
||||||
|
getLlamaSwapStatus,
|
||||||
|
triggerLlamaSwapLoad,
|
||||||
|
exceedsSafeContextFloor,
|
||||||
|
CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
|
||||||
|
} from './custom-model-routes.js';
|
||||||
import { matchesPattern } from '../../config/cli-registry/patterns.js';
|
import { matchesPattern } from '../../config/cli-registry/patterns.js';
|
||||||
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
|
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
|
||||||
import { TabLayoutValidationError } from '../../tab-layout.js';
|
import { TabLayoutValidationError } from '../../tab-layout.js';
|
||||||
@@ -1209,6 +1214,23 @@ export function registerSessionRoutes(
|
|||||||
if (!endpoint) {
|
if (!endpoint) {
|
||||||
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
|
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
|
||||||
}
|
}
|
||||||
|
const contextLength = endpoint.modelContextLengths?.[body.modelId];
|
||||||
|
|
||||||
|
// Some CLIs (today: only claude) carry enough of their own fixed system-prompt/tool-
|
||||||
|
// schema overhead that a small enough real context guarantees a first-message failure
|
||||||
|
// no matter what CLAUDE_CODE_MAX_CONTEXT_TOKENS says — confirmed live at ~36.4K tokens
|
||||||
|
// against a model configured with a real 16384-token context. Warn before committing
|
||||||
|
// to a restart that's certain to fail, rather than letting the user discover it via a
|
||||||
|
// cryptic 400 from the CLI itself. `confirmed` (already used for the swap-conflict
|
||||||
|
// warning below) skips this too — the user has already said "launch anyway" once.
|
||||||
|
if (!body.confirmed && exceedsSafeContextFloor(entry, contextLength)) {
|
||||||
|
return {
|
||||||
|
requiresContextWarning: true,
|
||||||
|
modelId: body.modelId,
|
||||||
|
contextLength,
|
||||||
|
minSafeContextTokens: CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
// llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on
|
// llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on
|
||||||
// demand, which can take anywhere from a few seconds to over a minute — long enough
|
// demand, which can take anywhere from a few seconds to over a minute — long enough
|
||||||
@@ -1250,7 +1272,6 @@ export function registerSessionRoutes(
|
|||||||
// fails its pattern rather than quoting it, which would silently launch the CLI on its
|
// fails its pattern rather than quoting it, which would silently launch the CLI on its
|
||||||
// own default provider again, so refuse an id the pattern cannot carry up front.
|
// own default provider again, so refuse an id the pattern cannot carry up front.
|
||||||
const modelSpec = entry.launch.params.model;
|
const modelSpec = entry.launch.params.model;
|
||||||
const contextLength = endpoint.modelContextLengths?.[body.modelId];
|
|
||||||
const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id, contextLength);
|
const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id, contextLength);
|
||||||
if (!applied) {
|
if (!applied) {
|
||||||
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
|
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
|
||||||
@@ -3570,6 +3591,20 @@ export function registerSessionRoutes(
|
|||||||
const cmHosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
|
const cmHosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
|
||||||
const cmEndpoint = cmHosts.find((h) => h.id === customModel.endpointId);
|
const cmEndpoint = cmHosts.find((h) => h.id === customModel.endpointId);
|
||||||
if (!cmEndpoint) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
|
if (!cmEndpoint) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
|
||||||
|
const cmContextLength = cmEndpoint.modelContextLengths?.[customModel.modelId];
|
||||||
|
|
||||||
|
// See the dedicated route's own comment for the full reasoning: some CLIs' own fixed
|
||||||
|
// overhead can exceed a small enough real context on the very first message,
|
||||||
|
// regardless of contextLengthVar. Warn before creating a session that's certain to
|
||||||
|
// fail immediately.
|
||||||
|
if (!customModel.confirmed && exceedsSafeContextFloor(cmEntry, cmContextLength)) {
|
||||||
|
return {
|
||||||
|
requiresContextWarning: true,
|
||||||
|
modelId: customModel.modelId,
|
||||||
|
contextLength: cmContextLength,
|
||||||
|
minSafeContextTokens: CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
// See the dedicated route's own comment for the full reasoning: llama.cpp runs one
|
// See the dedicated route's own comment for the full reasoning: llama.cpp runs one
|
||||||
// model at a time, llama-swap swaps on demand, and switching away from what another
|
// model at a time, llama-swap swaps on demand, and switching away from what another
|
||||||
@@ -3603,7 +3638,6 @@ export function registerSessionRoutes(
|
|||||||
// with, not a placeholder: `new Session({ id: ... })` accepts an explicit id for
|
// with, not a placeholder: `new Session({ id: ... })` accepts an explicit id for
|
||||||
// exactly this reason.
|
// exactly this reason.
|
||||||
qsCustomModelSessionId = randomUUID();
|
qsCustomModelSessionId = randomUUID();
|
||||||
const cmContextLength = cmEndpoint.modelContextLengths?.[customModel.modelId];
|
|
||||||
const cmApplied = applyCustomModelInjection(
|
const cmApplied = applyCustomModelInjection(
|
||||||
cmEntry,
|
cmEntry,
|
||||||
cmEndpoint,
|
cmEndpoint,
|
||||||
|
|||||||
@@ -57,6 +57,9 @@ function bootApp(
|
|||||||
<div class="modal" id="customModelSwapConfirmModal">
|
<div class="modal" id="customModelSwapConfirmModal">
|
||||||
<p id="customModelSwapConfirmMessage"></p>
|
<p id="customModelSwapConfirmMessage"></p>
|
||||||
</div>
|
</div>
|
||||||
|
<div class="modal" id="customModelContextWarningModal">
|
||||||
|
<p id="customModelContextWarningMessage"></p>
|
||||||
|
</div>
|
||||||
</body>`,
|
</body>`,
|
||||||
{ url: 'http://localhost/', runScripts: 'dangerously' }
|
{ url: 'http://localhost/', runScripts: 'dangerously' }
|
||||||
);
|
);
|
||||||
@@ -859,6 +862,64 @@ describe('Custom Model Endpoint Profiles: model-size load-time estimate', () =>
|
|||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
|
describe("Custom Model Endpoint Profiles: requiresContextWarning (this CLI's own overhead can exceed a small model's real context)", () => {
|
||||||
|
function launchHarness(applyResponses: Array<Record<string, unknown>>) {
|
||||||
|
const { win, app } = bootApp({});
|
||||||
|
app.activeSessionId = 'old-session';
|
||||||
|
app.run = async () => {
|
||||||
|
app.activeSessionId = 'new-session';
|
||||||
|
};
|
||||||
|
const applyBodies: unknown[] = [];
|
||||||
|
let call = 0;
|
||||||
|
app._api = async (path: string, opts?: { body?: unknown }) => {
|
||||||
|
if (path.endsWith('/custom-model')) {
|
||||||
|
applyBodies.push(opts?.body);
|
||||||
|
const data = applyResponses[Math.min(call, applyResponses.length - 1)];
|
||||||
|
call += 1;
|
||||||
|
return { ok: true, status: 200, json: async () => ({ success: true, data }) };
|
||||||
|
}
|
||||||
|
throw new Error(`unexpected _api call: ${path}`);
|
||||||
|
};
|
||||||
|
return { win, app, applyBodies };
|
||||||
|
}
|
||||||
|
|
||||||
|
it('confirming the in-app context-warning modal re-sends the apply with confirmed:true', async () => {
|
||||||
|
const { app, applyBodies } = launchHarness([
|
||||||
|
{ requiresContextWarning: true, modelId: 'qwen3', contextLength: 16384, minSafeContextTokens: 40000 },
|
||||||
|
{ customModel: { endpointId: 'llama-box' }, restarted: true, modelSwapInProgress: false },
|
||||||
|
]);
|
||||||
|
let confirmArgs: unknown[] | undefined;
|
||||||
|
app._confirmContextWarning = async (...args: unknown[]) => {
|
||||||
|
confirmArgs = args;
|
||||||
|
return true;
|
||||||
|
};
|
||||||
|
|
||||||
|
await app.runCustomModelEntry('claude', 'llama-box', 'qwen3');
|
||||||
|
|
||||||
|
expect(confirmArgs).toEqual(['qwen3', 16384, 40000]);
|
||||||
|
expect(applyBodies).toEqual([
|
||||||
|
{ endpointId: 'llama-box', modelId: 'qwen3' },
|
||||||
|
{ endpointId: 'llama-box', modelId: 'qwen3', confirmed: true },
|
||||||
|
]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('declining the in-app context-warning modal keeps the native backend and never re-sends the apply', async () => {
|
||||||
|
const { app, applyBodies } = launchHarness([
|
||||||
|
{ requiresContextWarning: true, modelId: 'qwen3', contextLength: 16384, minSafeContextTokens: 40000 },
|
||||||
|
]);
|
||||||
|
app._confirmContextWarning = async () => false;
|
||||||
|
let toastMessage: string | undefined;
|
||||||
|
app.showToast = (msg: string) => {
|
||||||
|
toastMessage = msg;
|
||||||
|
};
|
||||||
|
|
||||||
|
await app.runCustomModelEntry('claude', 'llama-box', 'qwen3');
|
||||||
|
|
||||||
|
expect(applyBodies).toHaveLength(1); // no second (confirmed) call
|
||||||
|
expect(toastMessage).toMatch(/context window too small/i);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
describe('Custom Model Endpoint Profiles: _confirmModelSwap (in-app modal, replaces a native confirm() popup)', () => {
|
describe('Custom Model Endpoint Profiles: _confirmModelSwap (in-app modal, replaces a native confirm() popup)', () => {
|
||||||
it('shows the message, activates the modal, and resolves true when "Switch anyway" is clicked', async () => {
|
it('shows the message, activates the modal, and resolves true when "Switch anyway" is clicked', async () => {
|
||||||
const { win, app } = bootApp({});
|
const { win, app } = bootApp({});
|
||||||
@@ -884,3 +945,41 @@ describe('Custom Model Endpoint Profiles: _confirmModelSwap (in-app modal, repla
|
|||||||
expect(win.document.getElementById('customModelSwapConfirmModal')!.classList.contains('active')).toBe(false);
|
expect(win.document.getElementById('customModelSwapConfirmModal')!.classList.contains('active')).toBe(false);
|
||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
|
describe('Custom Model Endpoint Profiles: _confirmContextWarning (in-app modal, native backend never restarted while it is up)', () => {
|
||||||
|
it('shows a message naming the model, the discovered context and the safe floor, activates the modal, and resolves true on "Launch anyway"', async () => {
|
||||||
|
const { win, app } = bootApp({});
|
||||||
|
const promise = app._confirmContextWarning('qwen3.8-27b-ud-q4_k_xl', 16384, 40000);
|
||||||
|
|
||||||
|
const modal = win.document.getElementById('customModelContextWarningModal')!;
|
||||||
|
expect(modal.classList.contains('active')).toBe(true);
|
||||||
|
const message = win.document.getElementById('customModelContextWarningMessage')!.textContent!;
|
||||||
|
expect(message).toContain('qwen3.8-27b-ud-q4_k_xl');
|
||||||
|
expect(message).toContain('16,384');
|
||||||
|
expect(message).toContain('40,000');
|
||||||
|
expect(message).toMatch(/llama-swap/i);
|
||||||
|
expect(message).toMatch(/fit-ctx/i);
|
||||||
|
|
||||||
|
app._resolveContextWarningConfirm(true);
|
||||||
|
|
||||||
|
expect(await promise).toBe(true);
|
||||||
|
expect(modal.classList.contains('active')).toBe(false);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('resolves false when Cancel is clicked', async () => {
|
||||||
|
const { win, app } = bootApp({});
|
||||||
|
const promise = app._confirmContextWarning('qwen3', 16384, 40000);
|
||||||
|
app._resolveContextWarningConfirm(false);
|
||||||
|
expect(await promise).toBe(false);
|
||||||
|
expect(win.document.getElementById('customModelContextWarningModal')!.classList.contains('active')).toBe(false);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('describes an unknown context length without printing a bogus number', async () => {
|
||||||
|
const { win, app } = bootApp({});
|
||||||
|
void app._confirmContextWarning('qwen3', undefined, 40000);
|
||||||
|
const message = win.document.getElementById('customModelContextWarningMessage')!.textContent!;
|
||||||
|
expect(message).not.toMatch(/undefined/);
|
||||||
|
expect(message).toMatch(/unknown/i);
|
||||||
|
app._resolveContextWarningConfirm(false);
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|||||||
@@ -275,6 +275,59 @@ describe('POST /api/quick-start: customModel (one-shot custom-model launch)', ()
|
|||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
|
describe("context-window floor warning (this CLI's own overhead can exceed a small model's real context)", () => {
|
||||||
|
const SMALL_CTX_ENDPOINT: CustomModelHost = {
|
||||||
|
id: 'ep-small',
|
||||||
|
label: 'tiny box',
|
||||||
|
baseUrl: 'http://192.168.1.51:8080',
|
||||||
|
apiKey: 'k',
|
||||||
|
modelContextLengths: { 'qwen3.8-27b-ud-q4_k_xl': 16384 },
|
||||||
|
};
|
||||||
|
|
||||||
|
it('warns instead of launching when the discovered context is below the safe floor', async () => {
|
||||||
|
await writeCustomModelHosts(getDataDir(), [ENDPOINT, SMALL_CTX_ENDPOINT]);
|
||||||
|
|
||||||
|
const res = await quickStart({
|
||||||
|
caseName: 'cm-small-ctx',
|
||||||
|
mode: 'claude',
|
||||||
|
customModel: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl' },
|
||||||
|
});
|
||||||
|
|
||||||
|
const body = res.json();
|
||||||
|
expect(body.requiresContextWarning).toBe(true);
|
||||||
|
expect(body.modelId).toBe('qwen3.8-27b-ud-q4_k_xl');
|
||||||
|
expect(body.contextLength).toBe(16384);
|
||||||
|
expect(body.minSafeContextTokens).toBe(40000);
|
||||||
|
// Nothing was actually created.
|
||||||
|
expect(ctx.sessions.size).toBe(1);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('launches once confirmed, skipping the context check', async () => {
|
||||||
|
await writeCustomModelHosts(getDataDir(), [ENDPOINT, SMALL_CTX_ENDPOINT]);
|
||||||
|
|
||||||
|
const res = await quickStart({
|
||||||
|
caseName: 'cm-small-ctx-confirmed',
|
||||||
|
mode: 'claude',
|
||||||
|
customModel: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl', confirmed: true },
|
||||||
|
});
|
||||||
|
|
||||||
|
expect(res.statusCode).toBe(200);
|
||||||
|
expect(res.json().requiresContextWarning).toBeUndefined();
|
||||||
|
expect(ctx.sessions.size).toBe(2);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('does not warn when nothing about context was discovered', async () => {
|
||||||
|
const res = await quickStart({
|
||||||
|
caseName: 'cm-no-ctx-data',
|
||||||
|
mode: 'claude',
|
||||||
|
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
|
||||||
|
});
|
||||||
|
|
||||||
|
expect(res.statusCode).toBe(200);
|
||||||
|
expect(res.json().requiresContextWarning).toBeUndefined();
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
describe('triggering the actual llama-swap load (not just watching for it)', () => {
|
describe('triggering the actual llama-swap load (not just watching for it)', () => {
|
||||||
it('sends a real inference request naming the target model, concurrently with launching the session', async () => {
|
it('sends a real inference request naming the target model, concurrently with launching the session', async () => {
|
||||||
const chatCalls: unknown[] = [];
|
const chatCalls: unknown[] = [];
|
||||||
|
|||||||
@@ -351,6 +351,113 @@ describe('POST /api/sessions/:id/custom-model', () => {
|
|||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
|
describe("context-window floor warning (this CLI's own overhead can exceed a small model's real context)", () => {
|
||||||
|
const SMALL_CTX_ENDPOINT: CustomModelHost = {
|
||||||
|
id: 'ep-small',
|
||||||
|
label: 'tiny box',
|
||||||
|
baseUrl: 'http://192.168.1.51:8080',
|
||||||
|
apiKey: 'k',
|
||||||
|
modelContextLengths: { 'qwen3.8-27b-ud-q4_k_xl': 16384 },
|
||||||
|
};
|
||||||
|
|
||||||
|
it('warns instead of applying when the discovered context is below the safe floor', async () => {
|
||||||
|
const { app, ctx } = await setup();
|
||||||
|
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
|
||||||
|
const session = ctx.sessions.get('test-session-1')!;
|
||||||
|
session.mode = 'claude';
|
||||||
|
|
||||||
|
const res = await app.inject({
|
||||||
|
method: 'POST',
|
||||||
|
url: '/api/sessions/test-session-1/custom-model',
|
||||||
|
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl' },
|
||||||
|
});
|
||||||
|
|
||||||
|
const body = res.json();
|
||||||
|
expect(body.success).not.toBe(false);
|
||||||
|
expect(body.requiresContextWarning).toBe(true);
|
||||||
|
expect(body.modelId).toBe('qwen3.8-27b-ud-q4_k_xl');
|
||||||
|
expect(body.contextLength).toBe(16384);
|
||||||
|
expect(body.minSafeContextTokens).toBe(40000);
|
||||||
|
// Nothing actually applied yet — this call only warned, it did not switch.
|
||||||
|
expect(session.setCustomModel).not.toHaveBeenCalled();
|
||||||
|
expect(session.restartCli).not.toHaveBeenCalled();
|
||||||
|
});
|
||||||
|
|
||||||
|
it('applies once confirmed, skipping the context check the second time', async () => {
|
||||||
|
const { app, ctx } = await setup();
|
||||||
|
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
|
||||||
|
const session = ctx.sessions.get('test-session-1')!;
|
||||||
|
session.mode = 'claude';
|
||||||
|
|
||||||
|
const res = await app.inject({
|
||||||
|
method: 'POST',
|
||||||
|
url: '/api/sessions/test-session-1/custom-model',
|
||||||
|
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl', confirmed: true },
|
||||||
|
});
|
||||||
|
|
||||||
|
const body = res.json();
|
||||||
|
expect(body.requiresContextWarning).toBeUndefined();
|
||||||
|
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
|
||||||
|
expect(session.restartCli).toHaveBeenCalledTimes(1);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('does not warn when the discovered context is comfortably above the floor', async () => {
|
||||||
|
const { app, ctx } = await setup();
|
||||||
|
const roomyEndpoint: CustomModelHost = {
|
||||||
|
id: 'ep-roomy',
|
||||||
|
label: 'roomy box',
|
||||||
|
baseUrl: 'http://192.168.1.52:8080',
|
||||||
|
apiKey: 'k',
|
||||||
|
modelContextLengths: { qwen3: 65536 },
|
||||||
|
};
|
||||||
|
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, roomyEndpoint]);
|
||||||
|
const session = ctx.sessions.get('test-session-1')!;
|
||||||
|
session.mode = 'claude';
|
||||||
|
|
||||||
|
const res = await app.inject({
|
||||||
|
method: 'POST',
|
||||||
|
url: '/api/sessions/test-session-1/custom-model',
|
||||||
|
payload: { endpointId: 'ep-roomy', modelId: 'qwen3' },
|
||||||
|
});
|
||||||
|
|
||||||
|
expect(res.json().requiresContextWarning).toBeUndefined();
|
||||||
|
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('does not warn when the context length was never discovered (nothing to compare)', async () => {
|
||||||
|
const { app, ctx } = await setup();
|
||||||
|
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT]);
|
||||||
|
const session = ctx.sessions.get('test-session-1')!;
|
||||||
|
session.mode = 'claude';
|
||||||
|
|
||||||
|
const res = await app.inject({
|
||||||
|
method: 'POST',
|
||||||
|
url: '/api/sessions/test-session-1/custom-model',
|
||||||
|
payload: { endpointId: 'ep1', modelId: 'qwen3' },
|
||||||
|
});
|
||||||
|
|
||||||
|
expect(res.json().requiresContextWarning).toBeUndefined();
|
||||||
|
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('does not warn for a CLI whose registry entry declares no contextLengthVar (opencode)', async () => {
|
||||||
|
// opencode's customModelInjection kind is configContentEnv, not env+contextLengthVar,
|
||||||
|
// so exceedsSafeContextFloor is false by construction regardless of context size.
|
||||||
|
const { app, ctx } = await setup();
|
||||||
|
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
|
||||||
|
const session = ctx.sessions.get('test-session-1')!;
|
||||||
|
session.mode = 'opencode';
|
||||||
|
|
||||||
|
const res = await app.inject({
|
||||||
|
method: 'POST',
|
||||||
|
url: '/api/sessions/test-session-1/custom-model',
|
||||||
|
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl' },
|
||||||
|
});
|
||||||
|
|
||||||
|
expect(res.json().requiresContextWarning).toBeUndefined();
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
describe('triggering the actual llama-swap load (not just watching for it)', () => {
|
describe('triggering the actual llama-swap load (not just watching for it)', () => {
|
||||||
it('sends a real inference request naming the target model when it is not already loaded and ready', async () => {
|
it('sends a real inference request naming the target model when it is not already loaded and ready', async () => {
|
||||||
const { app, ctx } = await setup();
|
const { app, ctx } = await setup();
|
||||||
|
|||||||
Reference in New Issue
Block a user