From 0e8b1981af5969825a4883eff72dc1355dc13f43 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Wed, 16 Sep 2026 12:08:26 +0800 Subject: [PATCH] fix(custom-model): isolate Claude config dir and inject real context length Addresses two live-validation findings on the Run-menu custom-model picker: 1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same config directory and warns about it (confirmed cosmetic - the API key wins for actual requests, verified via a real session's own API Usage Billing line). A custom-model claude session now gets an isolated CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty, no files written into it) so there is nothing to conflict with. projects is symlinked (junction on Windows) back into the real config dir so the response viewer, subagent windows and Read My Mind keep working for that session, best-effort. 2. Context-window overflow. Claude Code assumes a large default context window for a model id it doesn't recognise and never compacts, so a custom endpoint's real, much smaller context (verified live: a 400 exceeding a 16384-token llama-swap model with a stock ~33.7K-token system prompt) silently overflows. Discovery now also learns each model's real context length from llama.cpp/llama-swap's GET /props?model= (n_ctx), but ONLY for a model llama-swap's own /v1/models response already marks status.value === 'loaded' - never an unloaded one, since llama-swap treats ?model= as a routing hint and probing an unloaded model risks triggering an actual, slow, GPU-swapping load as a side effect of read-only discovery. A server with no status field at all gets no enrichment rather than a guess; a model not probed this round keeps its previously-learned value until it disappears from the list entirely. Stored per model (CustomModelHost.modelContextLengths) and applied via a new contextLengthVar registry field, set to CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude. Both new fields live on the existing env-kind customModelInjection capability shape, declared only on claude's registry entry - every other CLI's injection is unaffected (pinned by test). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG --- docs/custom-model-endpoints.md | 45 +++++++ docs/wiki/Custom-Model-Endpoints.md | 21 +++ src/config/cli-registry/schema.ts | 7 + src/config/cli-registry/stock.ts | 17 +++ src/config/cli-registry/types.ts | 26 +++- src/custom-model-hosts.ts | 11 ++ src/custom-model-injection-apply.ts | 57 +++++++- src/custom-model-injection.ts | 18 ++- src/web/routes/custom-model-routes.ts | 89 +++++++++++-- src/web/routes/session-routes.ts | 3 +- src/web/schemas.ts | 3 + .../custom-model-endpoint-rediscovery.test.ts | 84 ++++++++++++ test/custom-model-injection-apply.test.ts | 125 ++++++++++++++++++ test/custom-model-injection.test.ts | 18 +++ 14 files changed, 504 insertions(+), 20 deletions(-) create mode 100644 test/custom-model-injection-apply.test.ts diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index df6832d8..d481cccb 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -66,6 +66,23 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one. Endpoint management is admin-only in multi-user mode, same as remote/docker hosts — these are machine-level infra, not per-user settings. +**Context length is discovered too, opportunistically and safely.** The plain +`GET /v1/models` response has no context-window field, but llama.cpp's +llama-swap-proxied `GET /props?model=` does (`n_ctx`). Discovery only ever +calls it for a model llama-swap's own response already reports +`status.value === "loaded"` for — never for an unloaded one, because +llama-swap treats `?model=` as a routing hint and asking about a model that +isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side +effect of what should be read-only discovery. A server with no `status` field +on any entry at all (not llama-swap) gets no context-length enrichment, +rather than guessing. A model's previously-learned context length survives a +later cycle where it wasn't the loaded one; it's dropped only once the model +disappears from the endpoint's list entirely. Stored per model in +`modelContextLengths` and applied automatically (see "Applying a model to a +session" below) so a CLI that would otherwise assume a large default context +window for an unrecognized model id stops silently overflowing a much +smaller real one. + `defaultModelId` names which discovered model the picker pre-marks for that endpoint — the settings panel's Edit form exposes it as a select populated from the endpoint's own discovered `models`, and the route refuses a value @@ -143,6 +160,34 @@ since for those three the config file alone does not switch the model. reattaches the durable remote/in-container tmux rather than relaunching the agent, so the selection would report success and change nothing. +**Claude gets two more env vars when known/applicable, both declared on its +registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:** + +- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context + length (see the discovery section above) whenever one is known. Without + it, Claude Code assumes a large (200k) window for any unrecognized custom + model id and never compacts, which reliably overflows a much smaller real + local context — confirmed live: a stock ~33.7K-token system prompt against + a 16384-token llama-swap model failed with `exceeds the available context + size`. No entry for the model in `modelContextLengths` means the var is + simply omitted, never a guess. +- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory + the `configDir`-kind CLIs use (empty, no files written into it), so the + injected `ANTHROPIC_API_KEY` never shares a directory with a stored + claude.ai OAuth login. Claude Code still prints "Both claude.ai and + ANTHROPIC_API_KEY set" when the two coexist in the same config directory — + cosmetic (confirmed live: the API key wins for actual requests either way, + visible in the terminal's own `API Usage Billing` line) but worth + eliminating rather than living with. The directory's `projects` + subdirectory is symlinked (a junction on Windows) back to the real + `~/.claude/projects` so the response viewer, subagent windows and Read My + Mind keep working for that session — the same trade-off and fix documented + for a manually-set `CLAUDE_CONFIG_DIR` in + [`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied + automatically here. Best-effort: a platform that refuses the symlink keeps + the pre-existing blind-response-viewer side effect rather than failing the + whole custom-model apply over it. + Clear back to the harness's native cloud default with: ```bash diff --git a/docs/wiki/Custom-Model-Endpoints.md b/docs/wiki/Custom-Model-Endpoints.md index 7ee45173..2127227a 100644 --- a/docs/wiki/Custom-Model-Endpoints.md +++ b/docs/wiki/Custom-Model-Endpoints.md @@ -27,6 +27,15 @@ hosts — these are machine-level infra, not a per-user setting. shows up without another manual click of **Discover**. One endpoint being unreachable on a given cycle (powered off, wrong network) never blocks the others from refreshing. +**Context length is picked up automatically where it can be, safely.** Against a +llama.cpp/llama-swap server, discovery also learns each *currently loaded* model's real +context window and applies it to the launched session (Claude Code today — see below), so +the harness stops assuming a large default window for a model name it doesn't recognise and +overflowing a much smaller real one. It's deliberately never probed for a model that isn't +already loaded, since asking a llama-swap server about an unloaded model can trigger an +actual, slow model swap as a side effect — a model just not currently loaded keeps whatever +context length an earlier cycle already learned for it instead. + ## Running a session against one With the setting on and at least one endpoint carrying a discovered model, the **Run** @@ -61,6 +70,18 @@ redirecting those hasn't landed yet, see below. The picker also only appears in **Run** dropdown; the phone home screen builds its own run picker separately and does not currently offer these entries. +**Claude Code specifically gets two extra fixes applied automatically:** + +- Its discovered context length (see above) is passed through as + `CLAUDE_CODE_MAX_CONTEXT_TOKENS`, so it doesn't send a full-size prompt against a much + smaller real local context and overflow it. +- Its session runs with an isolated `CLAUDE_CONFIG_DIR`, so the injected API key never sits + in the same directory as a stored claude.ai login — that combination is harmless for actual + requests (the API key wins) but the CLI still prints a "both claude.ai and + ANTHROPIC_API_KEY set" warning about it, which this avoids entirely. The isolated directory + keeps a link back to your real session history so the response viewer and similar features + still work for that session. + ## Which harnesses actually work | Harness | Status | diff --git a/src/config/cli-registry/schema.ts b/src/config/cli-registry/schema.ts index 52cdfbd2..1b66e069 100644 --- a/src/config/cli-registry/schema.ts +++ b/src/config/cli-registry/schema.ts @@ -340,6 +340,13 @@ const capabilitiesSchema = z // an env var, so it declares baseUrl/apiKey injection with no model var at all. modelVars: z.array(envName).max(8), launchModel: launchModelTemplate, + // Optional: the env var to carry a discovered per-model context-window size + // (claude's CLAUDE_CODE_MAX_CONTEXT_TOKENS), and/or the env var that isolates + // this session's config/credential directory from the user's real one (claude's + // CLAUDE_CONFIG_DIR) so an injected API key never collides with a stored OAuth + // session. See the customModelInjection doc comment in cli-registry/types.ts. + contextLengthVar: envName.optional(), + configDirVar: envName.optional(), }) .strict(), z diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index eb054082..f7d08bfc 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -237,6 +237,12 @@ const CLAUDE: CliEntry = { 'ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL', + // CLAUDE_CODE_MAX_CONTEXT_TOKENS already matches the CLAUDE_CODE_* allowedPrefix, and + // CLAUDE_CONFIG_DIR is already an allowed exact key (docs/wiki/Agent-CLIs.md), so both + // were already reachable via plain envOverrides before this pair existed — listed here + // only so the custom-model route clamps them the same way as every other injected var. + 'CLAUDE_CODE_MAX_CONTEXT_TOKENS', + 'CLAUDE_CONFIG_DIR', ], gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } }, // Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real @@ -247,6 +253,17 @@ const CLAUDE: CliEntry = { baseUrlVar: 'ANTHROPIC_BASE_URL', apiKeyVar: 'ANTHROPIC_API_KEY', modelVars: ['ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL'], + // Verified via Claude Code's own docs: CLAUDE_CODE_MAX_CONTEXT_TOKENS overrides the + // assumed context window and applies directly for a model name Claude Code doesn't + // recognize as one of its own — exactly the custom-model case. Without it, Claude Code + // assumes a large (200k) window for any unrecognized model id and never compacts, + // eventually overflowing a much smaller real local context (see plan doc reasoning + // above the interface for the confirmed failure). + contextLengthVar: 'CLAUDE_CODE_MAX_CONTEXT_TOKENS', + // Isolates this session's config/credential directory so an injected ANTHROPIC_API_KEY + // never shares a directory with a stored claude.ai OAuth login — see the doc comment on + // customModelInjection in cli-registry/types.ts for the traded-off side effect. + configDirVar: 'CLAUDE_CONFIG_DIR', }, }, overlays: { diff --git a/src/config/cli-registry/types.ts b/src/config/cli-registry/types.ts index 1ba07cb0..3489354d 100644 --- a/src/config/cli-registry/types.ts +++ b/src/config/cli-registry/types.ts @@ -496,9 +496,33 @@ export interface CliCapabilities { * declares). Absent = the config alone selects the model (claude's env vars, * opencode's blob, codex's top-level `model` key). Applied by the session's * respawn options through the entry's `legacyConfigField`, never by id. + * + * `contextLengthVar` (env kind only): the env var a discovered per-model context-window + * size is written to when known (claude's `CLAUDE_CODE_MAX_CONTEXT_TOKENS`) — without it, + * a CLI that assumes a large default window for an unrecognized model name keeps sending + * full-size prompts against a much smaller local server and eventually overflows its real + * context (verified: a 33.7K-token system prompt against a 16384-token llama-swap model). + * Absent when the CLI has no such override, or the value is unknown for this model. + * + * `configDirVar` (env kind only): the env var that redirects this session's config/ + * credential directory to an isolated, per-session one (claude's `CLAUDE_CONFIG_DIR`), so + * an injected API key never coexists with a stored claude.ai OAuth session in the same + * directory — the CLI still warns "both claude.ai and ANTHROPIC_API_KEY set" when they + * share a directory even though the API key wins for actual requests. Isolating it trades + * that cosmetic warning for a documented side effect: a relocated config directory writes + * transcripts outside `~/.claude/projects`, blinding the response viewer, subagent + * windows, and Read My Mind for that session (see docs/wiki/Agent-CLIs.md). */ customModelInjection: - | { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string } + | { + kind: 'env'; + baseUrlVar: string; + apiKeyVar: string; + modelVars: string[]; + launchModel?: string; + contextLengthVar?: string; + configDirVar?: string; + } | { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string } | { kind: 'configDir'; diff --git a/src/custom-model-hosts.ts b/src/custom-model-hosts.ts index 0130eee7..cea7e7e5 100644 --- a/src/custom-model-hosts.ts +++ b/src/custom-model-hosts.ts @@ -50,6 +50,17 @@ export interface CustomModelHost { * after a re-discover is a property worth keeping even if the model list changes. */ defaultModelId?: string; + /** + * Discovered context-window size (tokens) per model id, keyed by the same strings as + * `models`. Populated opportunistically during discovery (`custom-model-routes.ts`) from + * llama.cpp/llama-swap's `GET /props?model=` — the plain OpenAI-shaped `/v1/models` + * response has no such field. Only ever probed for a model the server already reports as + * loaded (llama-swap's `status.value === 'loaded'`); an unloaded one is deliberately never + * probed, since llama-swap treats `/props?model=` as a routing hint that can trigger an + * actual (slow, GPU-swapping) model load as a side effect of merely asking. A model this + * has no entry for simply gets no context-length env override applied — never a guess. + */ + modelContextLengths?: Record; } export function customModelHostsPath(configDir: string): string { diff --git a/src/custom-model-injection-apply.ts b/src/custom-model-injection-apply.ts index 2df6f57d..b9bbde3a 100644 --- a/src/custom-model-injection-apply.ts +++ b/src/custom-model-injection-apply.ts @@ -10,7 +10,8 @@ * cli-registry changes" requirement it was written against. */ -import { chmodSync, mkdirSync, writeFileSync, rmSync } from 'node:fs'; +import { chmodSync, existsSync, mkdirSync, writeFileSync, rmSync, symlinkSync } from 'node:fs'; +import { homedir, platform } from 'node:os'; import { join, dirname } from 'node:path'; import { dataPath } from './config/instance.js'; import type { CliEntry } from './config/cli-registry/types.js'; @@ -48,6 +49,36 @@ export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInj return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv }; } +/** + * Real, shared Claude config directory Codeman's own host process runs under — honors + * `CLAUDE_CONFIG_DIR` the same way `claude-credentials.ts`'s `claudeCredentialsPath()` + * does, so the symlink below points at wherever `~/.claude/projects` actually lives + * rather than assuming the plain default. + */ +function realClaudeConfigDir(): string { + const configured = typeof process.env.CLAUDE_CONFIG_DIR === 'string' && process.env.CLAUDE_CONFIG_DIR.trim(); + return configured || join(homedir(), '.claude'); +} + +/** + * Symlinks `/projects` back to the real, shared `~/.claude/projects`, so an + * isolated `CLAUDE_CONFIG_DIR` (used to keep an injected API key away from a stored OAuth + * session — see `configDirVar` on customModelInjection) doesn't also blind the response + * viewer, subagent windows, and Read My Mind for that session (docs/wiki/Agent-CLIs.md). + * Best-effort: a platform that refuses symlinks (unprivileged Windows without a junction + * fallback working, e.g.) just keeps the pre-existing documented side effect instead of + * failing the whole custom-model apply over a nice-to-have. + */ +function linkSharedProjectsDir(isolatedDir: string): void { + const link = join(isolatedDir, 'projects'); + if (existsSync(link)) return; // already linked (idempotent re-apply) or real dir wrote one + try { + symlinkSync(join(realClaudeConfigDir(), 'projects'), link, platform() === 'win32' ? 'junction' : 'dir'); + } catch { + // best-effort only — response viewer/subagent windows go blind for this session instead + } +} + /** Best-effort recursive removal of a previously-written configDir. Never throws. */ export function removeConfigDir(dir: string | undefined): void { if (!dir) return; @@ -79,14 +110,30 @@ export function applyCustomModelInjection( entry: Pick, endpoint: CustomModelEndpoint, modelId: string, - sessionId: string + sessionId: string, + /** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */ + contextLength?: number ): AppliedCustomModel | undefined { - const injection = buildCustomModelInjection(entry, endpoint, modelId); + const injection = buildCustomModelInjection(entry, endpoint, modelId, contextLength); if (injection.kind === 'unsupported') return undefined; if (injection.kind === 'env') { + // `configDirVar` (claude's CLAUDE_CONFIG_DIR): point it at the same isolated, + // per-session directory the `configDir` kind uses, but write no files into it — an + // empty directory has no stored OAuth credential to conflict with the injected API + // key, which is the whole point. Reusing the same path keyed by sessionId keeps this + // idempotent across a boot-recovery re-apply, same as the configDir kind below. + let envOverrides = injection.envOverrides; + let configDir: string | undefined; + if (injection.configDirVar) { + configDir = customModelConfigDir(sessionId); + mkdirSync(configDir, { recursive: true, mode: 0o700 }); + linkSharedProjectsDir(configDir); + envOverrides = { ...envOverrides, [injection.configDirVar]: configDir }; + } return { - envOverrides: injection.envOverrides, - envKeys: Object.keys(injection.envOverrides), + envOverrides, + envKeys: Object.keys(envOverrides), + configDir, launchModel: injection.launchModel, }; } diff --git a/src/custom-model-injection.ts b/src/custom-model-injection.ts index 5f46001c..381bb51e 100644 --- a/src/custom-model-injection.ts +++ b/src/custom-model-injection.ts @@ -44,6 +44,14 @@ export interface EnvInjection { envOverrides: Record; /** See {@link ConfigDirInjection.launchModel}. */ launchModel?: string; + /** + * Name of the env var the caller should point at an isolated, credential-free config + * directory for this session (claude's `CLAUDE_CONFIG_DIR`), from the registry entry's + * `customModelInjection.configDirVar`. The actual directory value isn't computed here — + * this module is pure and has no sessionId to derive one from — the IO wrapper + * (`custom-model-injection-apply.ts`) creates it and adds it to `envOverrides`. + */ + configDirVar?: string; } export interface ConfigDirInjection { @@ -90,7 +98,9 @@ function quoted(value: string): string { export function buildCustomModelInjection( entry: Pick, endpoint: CustomModelEndpoint, - modelId: string + modelId: string, + /** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */ + contextLength?: number ): CustomModelInjectionResult { const cap = entry.capabilities.customModelInjection; const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY; @@ -102,7 +112,11 @@ export function buildCustomModelInjection( [cap.apiKeyVar]: apiKey, }; for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId; - return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId); + if (cap.contextLengthVar && contextLength !== undefined && Number.isFinite(contextLength)) { + envOverrides[cap.contextLengthVar] = String(Math.trunc(contextLength)); + } + const result = withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId); + return cap.configDirVar ? { ...result, configDirVar: cap.configDirVar } : result; } case 'configContentEnv': { diff --git a/src/web/routes/custom-model-routes.ts b/src/web/routes/custom-model-routes.ts index f5ccdccd..eabd191e 100644 --- a/src/web/routes/custom-model-routes.ts +++ b/src/web/routes/custom-model-routes.ts @@ -26,6 +26,7 @@ import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } fro const CODEMAN_CONFIG_DIR = getDataDir(); const DISCOVER_TIMEOUT_MS = 8000; +const PROPS_TIMEOUT_MS = 5000; function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse | null { if (!isMultiUserMode() || isAdmin(req)) return null; @@ -72,7 +73,7 @@ function applyStoredApiKey(incoming: CustomModelHost, existing: CustomModelHost) return incoming.apiKey ? incoming : { ...incoming, apiKey: existing.apiKey }; } -async function discoverModels(host: Pick): Promise { +function authHeaders(host: Pick): Record { const headers: Record = {}; const apiKey = host.apiKey?.trim(); // Exactly ONE header, never both — see custom-model-hosts.ts's CustomModelAuthStyle @@ -80,14 +81,73 @@ async function discoverModels(host: Pick; +} + +/** + * Best-effort: fetches `GET /props?model=` (llama.cpp-native, llama-swap-proxied) for + * ONE already-loaded model and pulls its real `n_ctx` out. Never called for a model that + * isn't already loaded — see the caller and `CustomModelHost.modelContextLengths` for why + * that's a hard safety requirement, not just a nicety: llama-swap treats this endpoint's + * `?model=` as a routing hint, and asking it about an unloaded model risks triggering an + * actual (slow, GPU-swapping) load as a side effect of what should be read-only discovery. + * Any failure (unreachable, non-2xx, missing/malformed field) is swallowed — one model's + * context length is a nice-to-have, never worth failing the whole discovery pass over. + */ +async function fetchContextLength( + host: Pick, + modelId: string, + headers: Record +): Promise { + try { + const url = new URL(`${host.baseUrl.replace(/\/+$/, '')}/props`); + url.searchParams.set('model', modelId); + const res = await webviewFetch(url, { headers, signal: AbortSignal.timeout(PROPS_TIMEOUT_MS) }); + if (!res.ok) return undefined; + const body = (await res.json()) as { n_ctx?: unknown; default_generation_settings?: { n_ctx?: unknown } }; + const nCtx = body.n_ctx ?? body.default_generation_settings?.n_ctx; + return typeof nCtx === 'number' && Number.isFinite(nCtx) && nCtx > 0 ? nCtx : undefined; + } catch { + return undefined; + } +} + +async function discoverModels( + host: Pick +): Promise { + const headers = authHeaders(host); const res = await webviewFetch(new URL(`${host.baseUrl.replace(/\/+$/, '')}/v1/models`), { headers, signal: AbortSignal.timeout(DISCOVER_TIMEOUT_MS), }); if (!res.ok) throw new Error(`HTTP ${res.status}`); - const body = (await res.json()) as { data?: Array<{ id?: unknown }> }; - return (body.data ?? []).map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0); + const body = (await res.json()) as { data?: Array<{ id?: unknown; status?: { value?: unknown } }> }; + const entries = body.data ?? []; + const models = entries.map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0); + + // llama-swap-specific, feature-detected: a server that never mentions `status` on ANY + // entry gets no context-length enrichment at all, rather than treating "no status field" + // as "assume unloaded" — either reading is a guess, and skipping is the safe one, since + // fetchContextLength must only ever run against a model this server itself calls loaded. + const hasStatusField = entries.some((m) => m && typeof m === 'object' && 'status' in m); + const contextLengths: Record = {}; + if (hasStatusField) { + const loadedIds = entries + .filter((m) => m.status && typeof m.status === 'object' && (m.status as { value?: unknown }).value === 'loaded') + .map((m) => m.id) + .filter((id): id is string => typeof id === 'string' && id.length > 0); + for (const id of loadedIds) { + const ctx = await fetchContextLength(host, id, headers); + if (ctx !== undefined) contextLengths[id] = ctx; + } + } + return { models, contextLengths }; } /** @@ -117,9 +177,16 @@ type RedactedHost = ReturnType; * each do their own `discoverModels()` + error handling around one shared * "how to apply a successful result" step. */ -function applyDiscoveredModels(host: CustomModelHost, models: string[]): CustomModelHost { +function applyDiscoveredModels(host: CustomModelHost, result: DiscoveryResult): CustomModelHost { + const { models, contextLengths } = result; const defaultModelId = host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined; - return { ...host, models, defaultModelId, lastDiscoveredAt: new Date().toISOString() }; + // Merge onto what's already known rather than replacing: a model not probed this round + // (not currently loaded) keeps whatever context length an earlier round already learned + // for it, and one no longer in the fresh list is dropped, same reasoning as defaultModelId. + const merged = { ...host.modelContextLengths, ...contextLengths }; + const kept = Object.fromEntries(Object.entries(merged).filter(([id]) => models.includes(id))); + const modelContextLengths = Object.keys(kept).length > 0 ? kept : undefined; + return { ...host, models, defaultModelId, modelContextLengths, lastDiscoveredAt: new Date().toISOString() }; } /** @@ -135,9 +202,9 @@ export async function refreshAllCustomModelHosts(): Promise { const hosts = await readCustomModelHosts(dataDir); for (const host of hosts) { if (isBlockedWebviewUrl(host.baseUrl)) continue; - let models: string[]; + let result: DiscoveryResult; try { - models = await discoverModels(host); + result = await discoverModels(host); } catch { continue; // unreachable this cycle — try again next tick, not fatal to the sweep } @@ -147,7 +214,7 @@ export async function refreshAllCustomModelHosts(): Promise { const current = await readCustomModelHosts(dataDir); const index = current.findIndex((item) => item.id === host.id); if (index === -1) continue; // deleted mid-sweep - current[index] = applyDiscoveredModels(current[index], models); + current[index] = applyDiscoveredModels(current[index], result); await writeCustomModelHosts(dataDir, current); } } @@ -222,11 +289,11 @@ export function registerCustomModelRoutes(app: FastifyInstance): void { return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed'); } try { - const models = await discoverModels(host); + const result = await discoverModels(host); const next = [...hosts]; - next[index] = applyDiscoveredModels(host, models); + next[index] = applyDiscoveredModels(host, result); await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next); - return { success: true, data: { models } }; + return { success: true, data: { models: result.models } }; } catch (err) { const blocked = egressBlockedReason(err); return createErrorResponse( diff --git a/src/web/routes/session-routes.ts b/src/web/routes/session-routes.ts index 03f92e5d..44b47b7c 100644 --- a/src/web/routes/session-routes.ts +++ b/src/web/routes/session-routes.ts @@ -1214,7 +1214,8 @@ export function registerSessionRoutes( // fails its pattern rather than quoting it, which would silently launch the CLI on its // own default provider again, so refuse an id the pattern cannot carry up front. const modelSpec = entry.launch.params.model; - const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id); + const contextLength = endpoint.modelContextLengths?.[body.modelId]; + const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id, contextLength); if (!applied) { return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`); } diff --git a/src/web/schemas.ts b/src/web/schemas.ts index f1f68151..6407e054 100644 --- a/src/web/schemas.ts +++ b/src/web/schemas.ts @@ -1924,6 +1924,9 @@ export const CustomModelHostSchema = z.object({ // host to apply it, so the check belongs there once, not duplicated into a refine // that would run on every unrelated field edit too). defaultModelId: z.string().max(200).optional(), + // Server-populated by discovery (custom-model-routes.ts); accepted here only so a client + // round-tripping the GET response back through PUT (edit-save) doesn't drop it. + modelContextLengths: z.record(z.string().max(200), z.number().int().positive().max(100_000_000)).optional(), }); /** POST /api/sessions/:id/custom-model — apply or clear a session's custom-model selection. */ diff --git a/test/custom-model-endpoint-rediscovery.test.ts b/test/custom-model-endpoint-rediscovery.test.ts index a26e98d4..7736ef93 100644 --- a/test/custom-model-endpoint-rediscovery.test.ts +++ b/test/custom-model-endpoint-rediscovery.test.ts @@ -124,3 +124,87 @@ describe('refreshAllCustomModelHosts (the periodic re-discovery sweep)', () => { expect(fetchMock).not.toHaveBeenCalled(); }); }); + +describe('refreshAllCustomModelHosts: context-length enrichment (llama.cpp/llama-swap /props)', () => { + it('probes /props?model= only for a model reported loaded, and stores its n_ctx', async () => { + const dir = getDataDir(); + await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]); + fetchMock.mockImplementation(async (url: URL) => { + if (url.pathname === '/v1/models') { + return new Response( + JSON.stringify({ + data: [ + { id: 'loaded-model', status: { value: 'loaded' } }, + { id: 'unloaded-model', status: { value: 'unloaded' } }, + ], + }), + { status: 200 } + ); + } + if (url.pathname === '/props') { + // Must never be reached for the unloaded model — asserted below by call count. + expect(url.searchParams.get('model')).toBe('loaded-model'); + return new Response(JSON.stringify({ n_ctx: 16384 }), { status: 200 }); + } + throw new Error(`unexpected request: ${url.href}`); + }); + + await refreshAllCustomModelHosts(); + + const [updated] = await readCustomModelHosts(dir); + expect(updated.modelContextLengths).toEqual({ 'loaded-model': 16384 }); + const propsCalls = fetchMock.mock.calls.filter(([url]) => (url as URL).pathname === '/props'); + expect(propsCalls).toHaveLength(1); + }); + + it('never probes /props at all when no entry mentions status — feature-detected, not assumed unloaded', async () => { + const dir = getDataDir(); + await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]); + fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'qwen3' }] }), { status: 200 })); + + await refreshAllCustomModelHosts(); + + expect(fetchMock).toHaveBeenCalledTimes(1); // /v1/models only + const [updated] = await readCustomModelHosts(dir); + expect(updated.modelContextLengths).toBeUndefined(); + }); + + it('keeps a previously-learned context length for a model no longer loaded, drops it once the model disappears entirely', async () => { + const dir = getDataDir(); + await writeCustomModelHosts(dir, [ + host({ + id: 'ep', + baseUrl: 'http://localhost:8080', + models: ['a', 'b'], + modelContextLengths: { a: 8192, b: 4096 }, + }), + ]); + // This round: 'a' is loaded (re-confirmed), 'b' is gone from the list entirely. + fetchMock.mockImplementation(async (url: URL) => { + if (url.pathname === '/v1/models') { + return new Response(JSON.stringify({ data: [{ id: 'a', status: { value: 'loaded' } }] }), { status: 200 }); + } + return new Response(JSON.stringify({ n_ctx: 8192 }), { status: 200 }); + }); + + await refreshAllCustomModelHosts(); + + const [updated] = await readCustomModelHosts(dir); + expect(updated.modelContextLengths).toEqual({ a: 8192 }); + }); + + it('a failed /props probe for the loaded model is swallowed, leaving no context length rather than failing the sweep', async () => { + const dir = getDataDir(); + await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]); + fetchMock.mockImplementation(async (url: URL) => { + if (url.pathname === '/v1/models') { + return new Response(JSON.stringify({ data: [{ id: 'a', status: { value: 'loaded' } }] }), { status: 200 }); + } + return new Response('nope', { status: 500 }); + }); + + await expect(refreshAllCustomModelHosts()).resolves.toBeUndefined(); + const [updated] = await readCustomModelHosts(dir); + expect(updated.modelContextLengths).toBeUndefined(); + }); +}); diff --git a/test/custom-model-injection-apply.test.ts b/test/custom-model-injection-apply.test.ts new file mode 100644 index 00000000..c41fe80f --- /dev/null +++ b/test/custom-model-injection-apply.test.ts @@ -0,0 +1,125 @@ +/** + * @fileoverview Tests for the two custom-model IO-layer fixes on top of the pure builder + * (docs/custom-model-endpoints-plan.md): + * + * 1. `contextLengthVar` — a discovered per-model context length reaches the actual + * session env (CLAUDE_CODE_MAX_CONTEXT_TOKENS), so a CLI stops assuming a large + * default window for an unrecognized custom model id and overflowing a much + * smaller real one. + * 2. `configDirVar` — an isolated, empty config directory is created and pointed at + * (CLAUDE_CONFIG_DIR), so an injected API key never shares a directory with a + * stored claude.ai OAuth session; `projects` is symlinked back into the real + * config dir so the response viewer/subagent windows/Read My Mind keep working. + * + * Port: N/A (no server; filesystem-only, under a temp CODEMAN data dir from test/setup.ts). + */ +import { existsSync, lstatSync, readdirSync, rmSync } from 'node:fs'; +import { homedir } from 'node:os'; +import { join } from 'node:path'; +import { afterEach, describe, expect, it } from 'vitest'; +import { getCli } from '../src/config/cli-registry/index.js'; +import { applyCustomModelInjection, customModelConfigDir } from '../src/custom-model-injection-apply.js'; +import type { CustomModelEndpoint } from '../src/custom-model-injection.js'; + +const endpoint: CustomModelEndpoint = { + id: 'ep1', + label: 'llama.cpp box', + baseUrl: 'http://192.168.1.50:8080', + apiKey: 'my-key', +}; + +function entryOrThrow(id: string) { + const entry = getCli(id); + if (!entry) throw new Error(`missing CLI registry entry: ${id}`); + return entry; +} + +const sessionsToClean: string[] = []; +afterEach(() => { + for (const id of sessionsToClean.splice(0)) rmSync(customModelConfigDir(id), { recursive: true, force: true }); +}); + +describe('applyCustomModelInjection: context length', () => { + it('claude: passes a known context length through to CLAUDE_CODE_MAX_CONTEXT_TOKENS', () => { + sessionsToClean.push('sess-ctx-1'); + const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', 'sess-ctx-1', 16384); + expect(applied?.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe('16384'); + expect(applied?.envKeys).toContain('CLAUDE_CODE_MAX_CONTEXT_TOKENS'); + }); + + it('claude: omits the var entirely when the context length is unknown', () => { + sessionsToClean.push('sess-ctx-2'); + const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', 'sess-ctx-2'); + expect(applied?.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBeUndefined(); + }); + + it('deepseek: has no contextLengthVar declared, so a passed-in length is a no-op', () => { + const applied = applyCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3', 'sess-ctx-3', 16384); + expect(Object.keys(applied?.envOverrides ?? {}).sort()).toEqual(['DEEPSEEK_API_KEY', 'DEEPSEEK_BASE_URL']); + }); +}); + +describe('applyCustomModelInjection: CLAUDE_CONFIG_DIR isolation', () => { + it('claude: creates an isolated, empty config dir and points CLAUDE_CONFIG_DIR at it', () => { + const sessionId = 'sess-cfgdir-1'; + sessionsToClean.push(sessionId); + const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId); + const expectedDir = customModelConfigDir(sessionId); + expect(applied?.envOverrides.CLAUDE_CONFIG_DIR).toBe(expectedDir); + expect(applied?.configDir).toBe(expectedDir); + expect(existsSync(expectedDir)).toBe(true); + // No credential/config files written into it — isolation, not a real config copy. + const entries = readdirSync(expectedDir).filter((name) => name !== 'projects'); + expect(entries).toEqual([]); + }); + + it('claude: symlinks (or junctions) projects back to the real config dir so the response viewer keeps working', () => { + const sessionId = 'sess-cfgdir-2'; + sessionsToClean.push(sessionId); + const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId); + const link = join(applied!.configDir!, 'projects'); + // Best-effort: only assert the link exists if it was actually created (the real + // ~/.claude/projects may not exist on a bare CI box, in which case linking is skipped). + if (existsSync(join(homedir(), '.claude', 'projects'))) { + expect(existsSync(link)).toBe(true); + expect(lstatSync(link).isSymbolicLink() || lstatSync(link).isDirectory()).toBe(true); + } + }); + + it('claude: re-applying to the same session is idempotent (boot-recovery re-apply)', () => { + const sessionId = 'sess-cfgdir-3'; + sessionsToClean.push(sessionId); + const first = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId); + const second = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId); + expect(second?.configDir).toBe(first?.configDir); + expect(existsSync(first!.configDir!)).toBe(true); + }); + + it('pi: configDir-kind CLIs are unaffected — no configDirVar concept for them', () => { + const sessionId = 'sess-cfgdir-pi'; + sessionsToClean.push(sessionId); + const applied = applyCustomModelInjection(entryOrThrow('pi'), endpoint, 'qwen3', sessionId); + expect(applied?.envOverrides.HOME).toBe(customModelConfigDir(sessionId)); + }); + + it('deepseek: no configDirVar declared, so no config dir is created at all', () => { + const sessionId = 'sess-cfgdir-deepseek'; + const applied = applyCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3', sessionId); + expect(applied?.configDir).toBeUndefined(); + expect(existsSync(customModelConfigDir(sessionId))).toBe(false); + }); +}); + +describe('applyCustomModelInjection: pre-existing behavior unaffected', () => { + it('opencode: still returns a plain env-kind result with no configDir', () => { + const sessionId = 'sess-opencode-1'; + const applied = applyCustomModelInjection(entryOrThrow('opencode'), endpoint, 'qwen3', sessionId); + expect(applied?.configDir).toBeUndefined(); + expect(applied?.envOverrides.OPENCODE_CONFIG_CONTENT).toBeTruthy(); + }); + + it('antigravity: still undefined (unsupported)', () => { + const applied = applyCustomModelInjection(entryOrThrow('antigravity'), endpoint, 'qwen3', 'sess-agy-1'); + expect(applied).toBeUndefined(); + }); +}); diff --git a/test/custom-model-injection.test.ts b/test/custom-model-injection.test.ts index 2acd08ca..1989c0f1 100644 --- a/test/custom-model-injection.test.ts +++ b/test/custom-model-injection.test.ts @@ -58,6 +58,24 @@ describe('buildCustomModelInjection', () => { }); }); + it('claude: also declares configDirVar (CLAUDE_CONFIG_DIR isolation) on the env-kind result', () => { + const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3'); + if (result.kind !== 'env') throw new Error('unreachable'); + expect(result.configDirVar).toBe('CLAUDE_CONFIG_DIR'); + }); + + it('claude: injects CLAUDE_CODE_MAX_CONTEXT_TOKENS when a context length is known', () => { + const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', 16384); + if (result.kind !== 'env') throw new Error('unreachable'); + expect(result.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe('16384'); + }); + + it('claude: omits CLAUDE_CODE_MAX_CONTEXT_TOKENS when the context length is unknown', () => { + const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3'); + if (result.kind !== 'env') throw new Error('unreachable'); + expect(result.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBeUndefined(); + }); + it('claude: falls back to a dummy key when the endpoint has none', () => { const result = buildCustomModelInjection(entryOrThrow('claude'), { ...endpoint, apiKey: undefined }, 'qwen3'); if (result.kind !== 'env') throw new Error('unreachable');