fix(custom-model): stop trusting /props's n_ctx, parse the real context size from /running's cmd

Root cause of the context-overflow regression reported live: "API Error: 400
request (36437 tokens) exceeds the available context size (16384 tokens)".
Discovery had stored modelContextLengths.qwen3.8-27b-ud-q4_k_xl = 154112,
so CLAUDE_CODE_MAX_CONTEXT_TOKENS told Claude Code it had a huge window and
it never compacted - but the real llama-swap server was launched with
--fit-ctx 16384 (confirmed against /running's own cmd field) and refused
the request right at that real limit.

/props?model=<id>'s n_ctx (the field discovery read) is confirmed live to
be unreliable for a --fit-ctx-launched backend: it reported 154112 for the
same model /running says was launched with --fit-ctx 16384 - appears to
report the model's theoretical/trained maximum context, not the runtime-
configured one.

discoverModels() now parses the REAL configured size straight out of
llama-swap's own launch command instead (parseCtxFromCmd(), reading
/running's cmd field - --fit-ctx first, then the plain llama.cpp -c/
--ctx-size a hand-written command might use), and only falls back to the
old /props probe when cmd states no recognizable flag at all. One /running
call now covers every loaded model's context length in a single request,
same as it already did for the swap-conflict check and the load trigger.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 20:47:12 +08:00
co-authored by Claude Sonnet 5
parent 7bbe408e44
commit 993710263d
3 changed files with 149 additions and 8 deletions
+16 -3
View File
@@ -67,9 +67,8 @@ Endpoint management is admin-only in multi-user mode, same as remote/docker
hosts — these are machine-level infra, not per-user settings.
**Context length is discovered too, opportunistically and safely.** The plain
`GET /v1/models` response has no context-window field, but llama.cpp's
llama-swap-proxied `GET /props?model=<id>` does (`n_ctx`). Discovery only ever
calls it for a model llama-swap's own response already reports
`GET /v1/models` response has no context-window field. Discovery only ever
looks for one for a model llama-swap's own response already reports
`status.value === "loaded"` for — never for an unloaded one, because
llama-swap treats `?model=` as a routing hint and asking about a model that
isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side
@@ -83,6 +82,20 @@ session" below) so a CLI that would otherwise assume a large default context
window for an unrecognized model id stops silently overflowing a much
smaller real one.
**Where that number actually comes from matters, and got this wrong once
already.** The first cut read it from llama.cpp's own
`GET /props?model=<id>` (`n_ctx`) — plausible, and it worked in testing, but
confirmed live to be actively WRONG for a `--fit-ctx`-launched llama-swap
backend: `/props` reported `n_ctx: 154112` for a model llama-swap itself had
launched with `--fit-ctx 16384`, and the real server then refused a request
right at that real 16384-token limit — `/props`'s `n_ctx` appears to report
the model's theoretical/trained maximum there, not the runtime-configured
one. Discovery now parses the REAL configured size straight out of
llama-swap's own launch command instead (`GET /running`'s `cmd` field —
`--fit-ctx <N>` first, then the plain llama.cpp `-c`/`--ctx-size` a
hand-written command might use), and only falls back to the `/props` probe
when `cmd` states no recognizable flag at all.
**File size is discovered too, when the server states one.** llama-swap
writes a GB figure into an auto-discovered model's own `description`
(`"Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"`), parsed