mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-06 23:49:41 +02:00
fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker: 1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same config directory and warns about it (confirmed cosmetic - the API key wins for actual requests, verified via a real session's own API Usage Billing line). A custom-model claude session now gets an isolated CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty, no files written into it) so there is nothing to conflict with. projects is symlinked (junction on Windows) back into the real config dir so the response viewer, subagent windows and Read My Mind keep working for that session, best-effort. 2. Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognise and never compacts, so a custom endpoint's real, much smaller context (verified live: a 400 exceeding a 16384-token llama-swap model with a stock ~33.7K-token system prompt) silently overflows. Discovery now also learns each model's real context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx), but ONLY for a model llama-swap's own /v1/models response already marks status.value === 'loaded' - never an unloaded one, since llama-swap treats ?model= as a routing hint and probing an unloaded model risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery. A server with no status field at all gets no enrichment rather than a guess; a model not probed this round keeps its previously-learned value until it disappears from the list entirely. Stored per model (CustomModelHost.modelContextLengths) and applied via a new contextLengthVar registry field, set to CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude. Both new fields live on the existing env-kind customModelInjection capability shape, declared only on claude's registry entry - every other CLI's injection is unaffected (pinned by test). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
5c25a52f95
commit
0e8b1981af
@@ -50,6 +50,17 @@ export interface CustomModelHost {
|
||||
* after a re-discover is a property worth keeping even if the model list changes.
|
||||
*/
|
||||
defaultModelId?: string;
|
||||
/**
|
||||
* Discovered context-window size (tokens) per model id, keyed by the same strings as
|
||||
* `models`. Populated opportunistically during discovery (`custom-model-routes.ts`) from
|
||||
* llama.cpp/llama-swap's `GET /props?model=<id>` — the plain OpenAI-shaped `/v1/models`
|
||||
* response has no such field. Only ever probed for a model the server already reports as
|
||||
* loaded (llama-swap's `status.value === 'loaded'`); an unloaded one is deliberately never
|
||||
* probed, since llama-swap treats `/props?model=` as a routing hint that can trigger an
|
||||
* actual (slow, GPU-swapping) model load as a side effect of merely asking. A model this
|
||||
* has no entry for simply gets no context-length env override applied — never a guess.
|
||||
*/
|
||||
modelContextLengths?: Record<string, number>;
|
||||
}
|
||||
|
||||
export function customModelHostsPath(configDir: string): string {
|
||||
|
||||
Reference in New Issue
Block a user