mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-02 13:39:41 +02:00
fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker: 1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same config directory and warns about it (confirmed cosmetic - the API key wins for actual requests, verified via a real session's own API Usage Billing line). A custom-model claude session now gets an isolated CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty, no files written into it) so there is nothing to conflict with. projects is symlinked (junction on Windows) back into the real config dir so the response viewer, subagent windows and Read My Mind keep working for that session, best-effort. 2. Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognise and never compacts, so a custom endpoint's real, much smaller context (verified live: a 400 exceeding a 16384-token llama-swap model with a stock ~33.7K-token system prompt) silently overflows. Discovery now also learns each model's real context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx), but ONLY for a model llama-swap's own /v1/models response already marks status.value === 'loaded' - never an unloaded one, since llama-swap treats ?model= as a routing hint and probing an unloaded model risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery. A server with no status field at all gets no enrichment rather than a guess; a model not probed this round keeps its previously-learned value until it disappears from the list entirely. Stored per model (CustomModelHost.modelContextLengths) and applied via a new contextLengthVar registry field, set to CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude. Both new fields live on the existing env-kind customModelInjection capability shape, declared only on claude's registry entry - every other CLI's injection is unaffected (pinned by test). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
5c25a52f95
commit
0e8b1981af
@@ -66,6 +66,23 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
|
||||
Endpoint management is admin-only in multi-user mode, same as remote/docker
|
||||
hosts — these are machine-level infra, not per-user settings.
|
||||
|
||||
**Context length is discovered too, opportunistically and safely.** The plain
|
||||
`GET /v1/models` response has no context-window field, but llama.cpp's
|
||||
llama-swap-proxied `GET /props?model=<id>` does (`n_ctx`). Discovery only ever
|
||||
calls it for a model llama-swap's own response already reports
|
||||
`status.value === "loaded"` for — never for an unloaded one, because
|
||||
llama-swap treats `?model=` as a routing hint and asking about a model that
|
||||
isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side
|
||||
effect of what should be read-only discovery. A server with no `status` field
|
||||
on any entry at all (not llama-swap) gets no context-length enrichment,
|
||||
rather than guessing. A model's previously-learned context length survives a
|
||||
later cycle where it wasn't the loaded one; it's dropped only once the model
|
||||
disappears from the endpoint's list entirely. Stored per model in
|
||||
`modelContextLengths` and applied automatically (see "Applying a model to a
|
||||
session" below) so a CLI that would otherwise assume a large default context
|
||||
window for an unrecognized model id stops silently overflowing a much
|
||||
smaller real one.
|
||||
|
||||
`defaultModelId` names which discovered model the picker pre-marks for that
|
||||
endpoint — the settings panel's Edit form exposes it as a select populated
|
||||
from the endpoint's own discovered `models`, and the route refuses a value
|
||||
@@ -143,6 +160,34 @@ since for those three the config file alone does not switch the model.
|
||||
reattaches the durable remote/in-container tmux rather than relaunching the
|
||||
agent, so the selection would report success and change nothing.
|
||||
|
||||
**Claude gets two more env vars when known/applicable, both declared on its
|
||||
registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
|
||||
|
||||
- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context
|
||||
length (see the discovery section above) whenever one is known. Without
|
||||
it, Claude Code assumes a large (200k) window for any unrecognized custom
|
||||
model id and never compacts, which reliably overflows a much smaller real
|
||||
local context — confirmed live: a stock ~33.7K-token system prompt against
|
||||
a 16384-token llama-swap model failed with `exceeds the available context
|
||||
size`. No entry for the model in `modelContextLengths` means the var is
|
||||
simply omitted, never a guess.
|
||||
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
|
||||
the `configDir`-kind CLIs use (empty, no files written into it), so the
|
||||
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
|
||||
claude.ai OAuth login. Claude Code still prints "Both claude.ai and
|
||||
ANTHROPIC_API_KEY set" when the two coexist in the same config directory —
|
||||
cosmetic (confirmed live: the API key wins for actual requests either way,
|
||||
visible in the terminal's own `API Usage Billing` line) but worth
|
||||
eliminating rather than living with. The directory's `projects`
|
||||
subdirectory is symlinked (a junction on Windows) back to the real
|
||||
`~/.claude/projects` so the response viewer, subagent windows and Read My
|
||||
Mind keep working for that session — the same trade-off and fix documented
|
||||
for a manually-set `CLAUDE_CONFIG_DIR` in
|
||||
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
|
||||
automatically here. Best-effort: a platform that refuses the symlink keeps
|
||||
the pre-existing blind-response-viewer side effect rather than failing the
|
||||
whole custom-model apply over it.
|
||||
|
||||
Clear back to the harness's native cloud default with:
|
||||
|
||||
```bash
|
||||
|
||||
@@ -27,6 +27,15 @@ hosts — these are machine-level infra, not a per-user setting.
|
||||
shows up without another manual click of **Discover**. One endpoint being unreachable on a
|
||||
given cycle (powered off, wrong network) never blocks the others from refreshing.
|
||||
|
||||
**Context length is picked up automatically where it can be, safely.** Against a
|
||||
llama.cpp/llama-swap server, discovery also learns each *currently loaded* model's real
|
||||
context window and applies it to the launched session (Claude Code today — see below), so
|
||||
the harness stops assuming a large default window for a model name it doesn't recognise and
|
||||
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
|
||||
already loaded, since asking a llama-swap server about an unloaded model can trigger an
|
||||
actual, slow model swap as a side effect — a model just not currently loaded keeps whatever
|
||||
context length an earlier cycle already learned for it instead.
|
||||
|
||||
## Running a session against one
|
||||
|
||||
With the setting on and at least one endpoint carrying a discovered model, the **Run**
|
||||
@@ -61,6 +70,18 @@ redirecting those hasn't landed yet, see below. The picker also only appears in
|
||||
**Run** dropdown; the phone home screen builds its own run picker separately and does not
|
||||
currently offer these entries.
|
||||
|
||||
**Claude Code specifically gets two extra fixes applied automatically:**
|
||||
|
||||
- Its discovered context length (see above) is passed through as
|
||||
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`, so it doesn't send a full-size prompt against a much
|
||||
smaller real local context and overflow it.
|
||||
- Its session runs with an isolated `CLAUDE_CONFIG_DIR`, so the injected API key never sits
|
||||
in the same directory as a stored claude.ai login — that combination is harmless for actual
|
||||
requests (the API key wins) but the CLI still prints a "both claude.ai and
|
||||
ANTHROPIC_API_KEY set" warning about it, which this avoids entirely. The isolated directory
|
||||
keeps a link back to your real session history so the response viewer and similar features
|
||||
still work for that session.
|
||||
|
||||
## Which harnesses actually work
|
||||
|
||||
| Harness | Status |
|
||||
|
||||
Reference in New Issue
Block a user