diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index df6832d8..d481cccb 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -66,6 +66,23 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one. Endpoint management is admin-only in multi-user mode, same as remote/docker hosts — these are machine-level infra, not per-user settings. +**Context length is discovered too, opportunistically and safely.** The plain +`GET /v1/models` response has no context-window field, but llama.cpp's +llama-swap-proxied `GET /props?model=` does (`n_ctx`). Discovery only ever +calls it for a model llama-swap's own response already reports +`status.value === "loaded"` for — never for an unloaded one, because +llama-swap treats `?model=` as a routing hint and asking about a model that +isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side +effect of what should be read-only discovery. A server with no `status` field +on any entry at all (not llama-swap) gets no context-length enrichment, +rather than guessing. A model's previously-learned context length survives a +later cycle where it wasn't the loaded one; it's dropped only once the model +disappears from the endpoint's list entirely. Stored per model in +`modelContextLengths` and applied automatically (see "Applying a model to a +session" below) so a CLI that would otherwise assume a large default context +window for an unrecognized model id stops silently overflowing a much +smaller real one. + `defaultModelId` names which discovered model the picker pre-marks for that endpoint — the settings panel's Edit form exposes it as a select populated from the endpoint's own discovered `models`, and the route refuses a value @@ -143,6 +160,34 @@ since for those three the config file alone does not switch the model. reattaches the durable remote/in-container tmux rather than relaunching the agent, so the selection would report success and change nothing. +**Claude gets two more env vars when known/applicable, both declared on its +registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:** + +- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context + length (see the discovery section above) whenever one is known. Without + it, Claude Code assumes a large (200k) window for any unrecognized custom + model id and never compacts, which reliably overflows a much smaller real + local context — confirmed live: a stock ~33.7K-token system prompt against + a 16384-token llama-swap model failed with `exceeds the available context + size`. No entry for the model in `modelContextLengths` means the var is + simply omitted, never a guess. +- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory + the `configDir`-kind CLIs use (empty, no files written into it), so the + injected `ANTHROPIC_API_KEY` never shares a directory with a stored + claude.ai OAuth login. Claude Code still prints "Both claude.ai and + ANTHROPIC_API_KEY set" when the two coexist in the same config directory — + cosmetic (confirmed live: the API key wins for actual requests either way, + visible in the terminal's own `API Usage Billing` line) but worth + eliminating rather than living with. The directory's `projects` + subdirectory is symlinked (a junction on Windows) back to the real + `~/.claude/projects` so the response viewer, subagent windows and Read My + Mind keep working for that session — the same trade-off and fix documented + for a manually-set `CLAUDE_CONFIG_DIR` in + [`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied + automatically here. Best-effort: a platform that refuses the symlink keeps + the pre-existing blind-response-viewer side effect rather than failing the + whole custom-model apply over it. + Clear back to the harness's native cloud default with: ```bash diff --git a/docs/wiki/Custom-Model-Endpoints.md b/docs/wiki/Custom-Model-Endpoints.md index 7ee45173..2127227a 100644 --- a/docs/wiki/Custom-Model-Endpoints.md +++ b/docs/wiki/Custom-Model-Endpoints.md @@ -27,6 +27,15 @@ hosts — these are machine-level infra, not a per-user setting. shows up without another manual click of **Discover**. One endpoint being unreachable on a given cycle (powered off, wrong network) never blocks the others from refreshing. +**Context length is picked up automatically where it can be, safely.** Against a +llama.cpp/llama-swap server, discovery also learns each *currently loaded* model's real +context window and applies it to the launched session (Claude Code today — see below), so +the harness stops assuming a large default window for a model name it doesn't recognise and +overflowing a much smaller real one. It's deliberately never probed for a model that isn't +already loaded, since asking a llama-swap server about an unloaded model can trigger an +actual, slow model swap as a side effect — a model just not currently loaded keeps whatever +context length an earlier cycle already learned for it instead. + ## Running a session against one With the setting on and at least one endpoint carrying a discovered model, the **Run** @@ -61,6 +70,18 @@ redirecting those hasn't landed yet, see below. The picker also only appears in **Run** dropdown; the phone home screen builds its own run picker separately and does not currently offer these entries. +**Claude Code specifically gets two extra fixes applied automatically:** + +- Its discovered context length (see above) is passed through as + `CLAUDE_CODE_MAX_CONTEXT_TOKENS`, so it doesn't send a full-size prompt against a much + smaller real local context and overflow it. +- Its session runs with an isolated `CLAUDE_CONFIG_DIR`, so the injected API key never sits + in the same directory as a stored claude.ai login — that combination is harmless for actual + requests (the API key wins) but the CLI still prints a "both claude.ai and + ANTHROPIC_API_KEY set" warning about it, which this avoids entirely. The isolated directory + keeps a link back to your real session history so the response viewer and similar features + still work for that session. + ## Which harnesses actually work | Harness | Status | diff --git a/src/config/cli-registry/schema.ts b/src/config/cli-registry/schema.ts index 52cdfbd2..1b66e069 100644 --- a/src/config/cli-registry/schema.ts +++ b/src/config/cli-registry/schema.ts @@ -340,6 +340,13 @@ const capabilitiesSchema = z // an env var, so it declares baseUrl/apiKey injection with no model var at all. modelVars: z.array(envName).max(8), launchModel: launchModelTemplate, + // Optional: the env var to carry a discovered per-model context-window size + // (claude's CLAUDE_CODE_MAX_CONTEXT_TOKENS), and/or the env var that isolates + // this session's config/credential directory from the user's real one (claude's + // CLAUDE_CONFIG_DIR) so an injected API key never collides with a stored OAuth + // session. See the customModelInjection doc comment in cli-registry/types.ts. + contextLengthVar: envName.optional(), + configDirVar: envName.optional(), }) .strict(), z diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index eb054082..f7d08bfc 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -237,6 +237,12 @@ const CLAUDE: CliEntry = { 'ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL', + // CLAUDE_CODE_MAX_CONTEXT_TOKENS already matches the CLAUDE_CODE_* allowedPrefix, and + // CLAUDE_CONFIG_DIR is already an allowed exact key (docs/wiki/Agent-CLIs.md), so both + // were already reachable via plain envOverrides before this pair existed — listed here + // only so the custom-model route clamps them the same way as every other injected var. + 'CLAUDE_CODE_MAX_CONTEXT_TOKENS', + 'CLAUDE_CONFIG_DIR', ], gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } }, // Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real @@ -247,6 +253,17 @@ const CLAUDE: CliEntry = { baseUrlVar: 'ANTHROPIC_BASE_URL', apiKeyVar: 'ANTHROPIC_API_KEY', modelVars: ['ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL'], + // Verified via Claude Code's own docs: CLAUDE_CODE_MAX_CONTEXT_TOKENS overrides the + // assumed context window and applies directly for a model name Claude Code doesn't + // recognize as one of its own — exactly the custom-model case. Without it, Claude Code + // assumes a large (200k) window for any unrecognized model id and never compacts, + // eventually overflowing a much smaller real local context (see plan doc reasoning + // above the interface for the confirmed failure). + contextLengthVar: 'CLAUDE_CODE_MAX_CONTEXT_TOKENS', + // Isolates this session's config/credential directory so an injected ANTHROPIC_API_KEY + // never shares a directory with a stored claude.ai OAuth login — see the doc comment on + // customModelInjection in cli-registry/types.ts for the traded-off side effect. + configDirVar: 'CLAUDE_CONFIG_DIR', }, }, overlays: { diff --git a/src/config/cli-registry/types.ts b/src/config/cli-registry/types.ts index 1ba07cb0..3489354d 100644 --- a/src/config/cli-registry/types.ts +++ b/src/config/cli-registry/types.ts @@ -496,9 +496,33 @@ export interface CliCapabilities { * declares). Absent = the config alone selects the model (claude's env vars, * opencode's blob, codex's top-level `model` key). Applied by the session's * respawn options through the entry's `legacyConfigField`, never by id. + * + * `contextLengthVar` (env kind only): the env var a discovered per-model context-window + * size is written to when known (claude's `CLAUDE_CODE_MAX_CONTEXT_TOKENS`) — without it, + * a CLI that assumes a large default window for an unrecognized model name keeps sending + * full-size prompts against a much smaller local server and eventually overflows its real + * context (verified: a 33.7K-token system prompt against a 16384-token llama-swap model). + * Absent when the CLI has no such override, or the value is unknown for this model. + * + * `configDirVar` (env kind only): the env var that redirects this session's config/ + * credential directory to an isolated, per-session one (claude's `CLAUDE_CONFIG_DIR`), so + * an injected API key never coexists with a stored claude.ai OAuth session in the same + * directory — the CLI still warns "both claude.ai and ANTHROPIC_API_KEY set" when they + * share a directory even though the API key wins for actual requests. Isolating it trades + * that cosmetic warning for a documented side effect: a relocated config directory writes + * transcripts outside `~/.claude/projects`, blinding the response viewer, subagent + * windows, and Read My Mind for that session (see docs/wiki/Agent-CLIs.md). */ customModelInjection: - | { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string } + | { + kind: 'env'; + baseUrlVar: string; + apiKeyVar: string; + modelVars: string[]; + launchModel?: string; + contextLengthVar?: string; + configDirVar?: string; + } | { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string } | { kind: 'configDir'; diff --git a/src/custom-model-hosts.ts b/src/custom-model-hosts.ts index 0130eee7..cea7e7e5 100644 --- a/src/custom-model-hosts.ts +++ b/src/custom-model-hosts.ts @@ -50,6 +50,17 @@ export interface CustomModelHost { * after a re-discover is a property worth keeping even if the model list changes. */ defaultModelId?: string; + /** + * Discovered context-window size (tokens) per model id, keyed by the same strings as + * `models`. Populated opportunistically during discovery (`custom-model-routes.ts`) from + * llama.cpp/llama-swap's `GET /props?model=` — the plain OpenAI-shaped `/v1/models` + * response has no such field. Only ever probed for a model the server already reports as + * loaded (llama-swap's `status.value === 'loaded'`); an unloaded one is deliberately never + * probed, since llama-swap treats `/props?model=` as a routing hint that can trigger an + * actual (slow, GPU-swapping) model load as a side effect of merely asking. A model this + * has no entry for simply gets no context-length env override applied — never a guess. + */ + modelContextLengths?: Record; } export function customModelHostsPath(configDir: string): string { diff --git a/src/custom-model-injection-apply.ts b/src/custom-model-injection-apply.ts index 2df6f57d..b9bbde3a 100644 --- a/src/custom-model-injection-apply.ts +++ b/src/custom-model-injection-apply.ts @@ -10,7 +10,8 @@ * cli-registry changes" requirement it was written against. */ -import { chmodSync, mkdirSync, writeFileSync, rmSync } from 'node:fs'; +import { chmodSync, existsSync, mkdirSync, writeFileSync, rmSync, symlinkSync } from 'node:fs'; +import { homedir, platform } from 'node:os'; import { join, dirname } from 'node:path'; import { dataPath } from './config/instance.js'; import type { CliEntry } from './config/cli-registry/types.js'; @@ -48,6 +49,36 @@ export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInj return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv }; } +/** + * Real, shared Claude config directory Codeman's own host process runs under — honors + * `CLAUDE_CONFIG_DIR` the same way `claude-credentials.ts`'s `claudeCredentialsPath()` + * does, so the symlink below points at wherever `~/.claude/projects` actually lives + * rather than assuming the plain default. + */ +function realClaudeConfigDir(): string { + const configured = typeof process.env.CLAUDE_CONFIG_DIR === 'string' && process.env.CLAUDE_CONFIG_DIR.trim(); + return configured || join(homedir(), '.claude'); +} + +/** + * Symlinks `/projects` back to the real, shared `~/.claude/projects`, so an + * isolated `CLAUDE_CONFIG_DIR` (used to keep an injected API key away from a stored OAuth + * session — see `configDirVar` on customModelInjection) doesn't also blind the response + * viewer, subagent windows, and Read My Mind for that session (docs/wiki/Agent-CLIs.md). + * Best-effort: a platform that refuses symlinks (unprivileged Windows without a junction + * fallback working, e.g.) just keeps the pre-existing documented side effect instead of + * failing the whole custom-model apply over a nice-to-have. + */ +function linkSharedProjectsDir(isolatedDir: string): void { + const link = join(isolatedDir, 'projects'); + if (existsSync(link)) return; // already linked (idempotent re-apply) or real dir wrote one + try { + symlinkSync(join(realClaudeConfigDir(), 'projects'), link, platform() === 'win32' ? 'junction' : 'dir'); + } catch { + // best-effort only — response viewer/subagent windows go blind for this session instead + } +} + /** Best-effort recursive removal of a previously-written configDir. Never throws. */ export function removeConfigDir(dir: string | undefined): void { if (!dir) return; @@ -79,14 +110,30 @@ export function applyCustomModelInjection( entry: Pick, endpoint: CustomModelEndpoint, modelId: string, - sessionId: string + sessionId: string, + /** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */ + contextLength?: number ): AppliedCustomModel | undefined { - const injection = buildCustomModelInjection(entry, endpoint, modelId); + const injection = buildCustomModelInjection(entry, endpoint, modelId, contextLength); if (injection.kind === 'unsupported') return undefined; if (injection.kind === 'env') { + // `configDirVar` (claude's CLAUDE_CONFIG_DIR): point it at the same isolated, + // per-session directory the `configDir` kind uses, but write no files into it — an + // empty directory has no stored OAuth credential to conflict with the injected API + // key, which is the whole point. Reusing the same path keyed by sessionId keeps this + // idempotent across a boot-recovery re-apply, same as the configDir kind below. + let envOverrides = injection.envOverrides; + let configDir: string | undefined; + if (injection.configDirVar) { + configDir = customModelConfigDir(sessionId); + mkdirSync(configDir, { recursive: true, mode: 0o700 }); + linkSharedProjectsDir(configDir); + envOverrides = { ...envOverrides, [injection.configDirVar]: configDir }; + } return { - envOverrides: injection.envOverrides, - envKeys: Object.keys(injection.envOverrides), + envOverrides, + envKeys: Object.keys(envOverrides), + configDir, launchModel: injection.launchModel, }; } diff --git a/src/custom-model-injection.ts b/src/custom-model-injection.ts index 5f46001c..381bb51e 100644 --- a/src/custom-model-injection.ts +++ b/src/custom-model-injection.ts @@ -44,6 +44,14 @@ export interface EnvInjection { envOverrides: Record; /** See {@link ConfigDirInjection.launchModel}. */ launchModel?: string; + /** + * Name of the env var the caller should point at an isolated, credential-free config + * directory for this session (claude's `CLAUDE_CONFIG_DIR`), from the registry entry's + * `customModelInjection.configDirVar`. The actual directory value isn't computed here — + * this module is pure and has no sessionId to derive one from — the IO wrapper + * (`custom-model-injection-apply.ts`) creates it and adds it to `envOverrides`. + */ + configDirVar?: string; } export interface ConfigDirInjection { @@ -90,7 +98,9 @@ function quoted(value: string): string { export function buildCustomModelInjection( entry: Pick, endpoint: CustomModelEndpoint, - modelId: string + modelId: string, + /** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */ + contextLength?: number ): CustomModelInjectionResult { const cap = entry.capabilities.customModelInjection; const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY; @@ -102,7 +112,11 @@ export function buildCustomModelInjection( [cap.apiKeyVar]: apiKey, }; for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId; - return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId); + if (cap.contextLengthVar && contextLength !== undefined && Number.isFinite(contextLength)) { + envOverrides[cap.contextLengthVar] = String(Math.trunc(contextLength)); + } + const result = withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId); + return cap.configDirVar ? { ...result, configDirVar: cap.configDirVar } : result; } case 'configContentEnv': { diff --git a/src/web/routes/custom-model-routes.ts b/src/web/routes/custom-model-routes.ts index f5ccdccd..eabd191e 100644 --- a/src/web/routes/custom-model-routes.ts +++ b/src/web/routes/custom-model-routes.ts @@ -26,6 +26,7 @@ import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } fro const CODEMAN_CONFIG_DIR = getDataDir(); const DISCOVER_TIMEOUT_MS = 8000; +const PROPS_TIMEOUT_MS = 5000; function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse | null { if (!isMultiUserMode() || isAdmin(req)) return null; @@ -72,7 +73,7 @@ function applyStoredApiKey(incoming: CustomModelHost, existing: CustomModelHost) return incoming.apiKey ? incoming : { ...incoming, apiKey: existing.apiKey }; } -async function discoverModels(host: Pick): Promise { +function authHeaders(host: Pick): Record { const headers: Record = {}; const apiKey = host.apiKey?.trim(); // Exactly ONE header, never both — see custom-model-hosts.ts's CustomModelAuthStyle @@ -80,14 +81,73 @@ async function discoverModels(host: Pick; +} + +/** + * Best-effort: fetches `GET /props?model=` (llama.cpp-native, llama-swap-proxied) for + * ONE already-loaded model and pulls its real `n_ctx` out. Never called for a model that + * isn't already loaded — see the caller and `CustomModelHost.modelContextLengths` for why + * that's a hard safety requirement, not just a nicety: llama-swap treats this endpoint's + * `?model=` as a routing hint, and asking it about an unloaded model risks triggering an + * actual (slow, GPU-swapping) load as a side effect of what should be read-only discovery. + * Any failure (unreachable, non-2xx, missing/malformed field) is swallowed — one model's + * context length is a nice-to-have, never worth failing the whole discovery pass over. + */ +async function fetchContextLength( + host: Pick, + modelId: string, + headers: Record +): Promise { + try { + const url = new URL(`${host.baseUrl.replace(/\/+$/, '')}/props`); + url.searchParams.set('model', modelId); + const res = await webviewFetch(url, { headers, signal: AbortSignal.timeout(PROPS_TIMEOUT_MS) }); + if (!res.ok) return undefined; + const body = (await res.json()) as { n_ctx?: unknown; default_generation_settings?: { n_ctx?: unknown } }; + const nCtx = body.n_ctx ?? body.default_generation_settings?.n_ctx; + return typeof nCtx === 'number' && Number.isFinite(nCtx) && nCtx > 0 ? nCtx : undefined; + } catch { + return undefined; + } +} + +async function discoverModels( + host: Pick +): Promise { + const headers = authHeaders(host); const res = await webviewFetch(new URL(`${host.baseUrl.replace(/\/+$/, '')}/v1/models`), { headers, signal: AbortSignal.timeout(DISCOVER_TIMEOUT_MS), }); if (!res.ok) throw new Error(`HTTP ${res.status}`); - const body = (await res.json()) as { data?: Array<{ id?: unknown }> }; - return (body.data ?? []).map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0); + const body = (await res.json()) as { data?: Array<{ id?: unknown; status?: { value?: unknown } }> }; + const entries = body.data ?? []; + const models = entries.map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0); + + // llama-swap-specific, feature-detected: a server that never mentions `status` on ANY + // entry gets no context-length enrichment at all, rather than treating "no status field" + // as "assume unloaded" — either reading is a guess, and skipping is the safe one, since + // fetchContextLength must only ever run against a model this server itself calls loaded. + const hasStatusField = entries.some((m) => m && typeof m === 'object' && 'status' in m); + const contextLengths: Record = {}; + if (hasStatusField) { + const loadedIds = entries + .filter((m) => m.status && typeof m.status === 'object' && (m.status as { value?: unknown }).value === 'loaded') + .map((m) => m.id) + .filter((id): id is string => typeof id === 'string' && id.length > 0); + for (const id of loadedIds) { + const ctx = await fetchContextLength(host, id, headers); + if (ctx !== undefined) contextLengths[id] = ctx; + } + } + return { models, contextLengths }; } /** @@ -117,9 +177,16 @@ type RedactedHost = ReturnType; * each do their own `discoverModels()` + error handling around one shared * "how to apply a successful result" step. */ -function applyDiscoveredModels(host: CustomModelHost, models: string[]): CustomModelHost { +function applyDiscoveredModels(host: CustomModelHost, result: DiscoveryResult): CustomModelHost { + const { models, contextLengths } = result; const defaultModelId = host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined; - return { ...host, models, defaultModelId, lastDiscoveredAt: new Date().toISOString() }; + // Merge onto what's already known rather than replacing: a model not probed this round + // (not currently loaded) keeps whatever context length an earlier round already learned + // for it, and one no longer in the fresh list is dropped, same reasoning as defaultModelId. + const merged = { ...host.modelContextLengths, ...contextLengths }; + const kept = Object.fromEntries(Object.entries(merged).filter(([id]) => models.includes(id))); + const modelContextLengths = Object.keys(kept).length > 0 ? kept : undefined; + return { ...host, models, defaultModelId, modelContextLengths, lastDiscoveredAt: new Date().toISOString() }; } /** @@ -135,9 +202,9 @@ export async function refreshAllCustomModelHosts(): Promise { const hosts = await readCustomModelHosts(dataDir); for (const host of hosts) { if (isBlockedWebviewUrl(host.baseUrl)) continue; - let models: string[]; + let result: DiscoveryResult; try { - models = await discoverModels(host); + result = await discoverModels(host); } catch { continue; // unreachable this cycle — try again next tick, not fatal to the sweep } @@ -147,7 +214,7 @@ export async function refreshAllCustomModelHosts(): Promise { const current = await readCustomModelHosts(dataDir); const index = current.findIndex((item) => item.id === host.id); if (index === -1) continue; // deleted mid-sweep - current[index] = applyDiscoveredModels(current[index], models); + current[index] = applyDiscoveredModels(current[index], result); await writeCustomModelHosts(dataDir, current); } } @@ -222,11 +289,11 @@ export function registerCustomModelRoutes(app: FastifyInstance): void { return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed'); } try { - const models = await discoverModels(host); + const result = await discoverModels(host); const next = [...hosts]; - next[index] = applyDiscoveredModels(host, models); + next[index] = applyDiscoveredModels(host, result); await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next); - return { success: true, data: { models } }; + return { success: true, data: { models: result.models } }; } catch (err) { const blocked = egressBlockedReason(err); return createErrorResponse( diff --git a/src/web/routes/session-routes.ts b/src/web/routes/session-routes.ts index 03f92e5d..44b47b7c 100644 --- a/src/web/routes/session-routes.ts +++ b/src/web/routes/session-routes.ts @@ -1214,7 +1214,8 @@ export function registerSessionRoutes( // fails its pattern rather than quoting it, which would silently launch the CLI on its // own default provider again, so refuse an id the pattern cannot carry up front. const modelSpec = entry.launch.params.model; - const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id); + const contextLength = endpoint.modelContextLengths?.[body.modelId]; + const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id, contextLength); if (!applied) { return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`); } diff --git a/src/web/schemas.ts b/src/web/schemas.ts index f1f68151..6407e054 100644 --- a/src/web/schemas.ts +++ b/src/web/schemas.ts @@ -1924,6 +1924,9 @@ export const CustomModelHostSchema = z.object({ // host to apply it, so the check belongs there once, not duplicated into a refine // that would run on every unrelated field edit too). defaultModelId: z.string().max(200).optional(), + // Server-populated by discovery (custom-model-routes.ts); accepted here only so a client + // round-tripping the GET response back through PUT (edit-save) doesn't drop it. + modelContextLengths: z.record(z.string().max(200), z.number().int().positive().max(100_000_000)).optional(), }); /** POST /api/sessions/:id/custom-model — apply or clear a session's custom-model selection. */ diff --git a/test/custom-model-endpoint-rediscovery.test.ts b/test/custom-model-endpoint-rediscovery.test.ts index a26e98d4..7736ef93 100644 --- a/test/custom-model-endpoint-rediscovery.test.ts +++ b/test/custom-model-endpoint-rediscovery.test.ts @@ -124,3 +124,87 @@ describe('refreshAllCustomModelHosts (the periodic re-discovery sweep)', () => { expect(fetchMock).not.toHaveBeenCalled(); }); }); + +describe('refreshAllCustomModelHosts: context-length enrichment (llama.cpp/llama-swap /props)', () => { + it('probes /props?model= only for a model reported loaded, and stores its n_ctx', async () => { + const dir = getDataDir(); + await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]); + fetchMock.mockImplementation(async (url: URL) => { + if (url.pathname === '/v1/models') { + return new Response( + JSON.stringify({ + data: [ + { id: 'loaded-model', status: { value: 'loaded' } }, + { id: 'unloaded-model', status: { value: 'unloaded' } }, + ], + }), + { status: 200 } + ); + } + if (url.pathname === '/props') { + // Must never be reached for the unloaded model — asserted below by call count. + expect(url.searchParams.get('model')).toBe('loaded-model'); + return new Response(JSON.stringify({ n_ctx: 16384 }), { status: 200 }); + } + throw new Error(`unexpected request: ${url.href}`); + }); + + await refreshAllCustomModelHosts(); + + const [updated] = await readCustomModelHosts(dir); + expect(updated.modelContextLengths).toEqual({ 'loaded-model': 16384 }); + const propsCalls = fetchMock.mock.calls.filter(([url]) => (url as URL).pathname === '/props'); + expect(propsCalls).toHaveLength(1); + }); + + it('never probes /props at all when no entry mentions status — feature-detected, not assumed unloaded', async () => { + const dir = getDataDir(); + await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]); + fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'qwen3' }] }), { status: 200 })); + + await refreshAllCustomModelHosts(); + + expect(fetchMock).toHaveBeenCalledTimes(1); // /v1/models only + const [updated] = await readCustomModelHosts(dir); + expect(updated.modelContextLengths).toBeUndefined(); + }); + + it('keeps a previously-learned context length for a model no longer loaded, drops it once the model disappears entirely', async () => { + const dir = getDataDir(); + await writeCustomModelHosts(dir, [ + host({ + id: 'ep', + baseUrl: 'http://localhost:8080', + models: ['a', 'b'], + modelContextLengths: { a: 8192, b: 4096 }, + }), + ]); + // This round: 'a' is loaded (re-confirmed), 'b' is gone from the list entirely. + fetchMock.mockImplementation(async (url: URL) => { + if (url.pathname === '/v1/models') { + return new Response(JSON.stringify({ data: [{ id: 'a', status: { value: 'loaded' } }] }), { status: 200 }); + } + return new Response(JSON.stringify({ n_ctx: 8192 }), { status: 200 }); + }); + + await refreshAllCustomModelHosts(); + + const [updated] = await readCustomModelHosts(dir); + expect(updated.modelContextLengths).toEqual({ a: 8192 }); + }); + + it('a failed /props probe for the loaded model is swallowed, leaving no context length rather than failing the sweep', async () => { + const dir = getDataDir(); + await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]); + fetchMock.mockImplementation(async (url: URL) => { + if (url.pathname === '/v1/models') { + return new Response(JSON.stringify({ data: [{ id: 'a', status: { value: 'loaded' } }] }), { status: 200 }); + } + return new Response('nope', { status: 500 }); + }); + + await expect(refreshAllCustomModelHosts()).resolves.toBeUndefined(); + const [updated] = await readCustomModelHosts(dir); + expect(updated.modelContextLengths).toBeUndefined(); + }); +}); diff --git a/test/custom-model-injection-apply.test.ts b/test/custom-model-injection-apply.test.ts new file mode 100644 index 00000000..c41fe80f --- /dev/null +++ b/test/custom-model-injection-apply.test.ts @@ -0,0 +1,125 @@ +/** + * @fileoverview Tests for the two custom-model IO-layer fixes on top of the pure builder + * (docs/custom-model-endpoints-plan.md): + * + * 1. `contextLengthVar` — a discovered per-model context length reaches the actual + * session env (CLAUDE_CODE_MAX_CONTEXT_TOKENS), so a CLI stops assuming a large + * default window for an unrecognized custom model id and overflowing a much + * smaller real one. + * 2. `configDirVar` — an isolated, empty config directory is created and pointed at + * (CLAUDE_CONFIG_DIR), so an injected API key never shares a directory with a + * stored claude.ai OAuth session; `projects` is symlinked back into the real + * config dir so the response viewer/subagent windows/Read My Mind keep working. + * + * Port: N/A (no server; filesystem-only, under a temp CODEMAN data dir from test/setup.ts). + */ +import { existsSync, lstatSync, readdirSync, rmSync } from 'node:fs'; +import { homedir } from 'node:os'; +import { join } from 'node:path'; +import { afterEach, describe, expect, it } from 'vitest'; +import { getCli } from '../src/config/cli-registry/index.js'; +import { applyCustomModelInjection, customModelConfigDir } from '../src/custom-model-injection-apply.js'; +import type { CustomModelEndpoint } from '../src/custom-model-injection.js'; + +const endpoint: CustomModelEndpoint = { + id: 'ep1', + label: 'llama.cpp box', + baseUrl: 'http://192.168.1.50:8080', + apiKey: 'my-key', +}; + +function entryOrThrow(id: string) { + const entry = getCli(id); + if (!entry) throw new Error(`missing CLI registry entry: ${id}`); + return entry; +} + +const sessionsToClean: string[] = []; +afterEach(() => { + for (const id of sessionsToClean.splice(0)) rmSync(customModelConfigDir(id), { recursive: true, force: true }); +}); + +describe('applyCustomModelInjection: context length', () => { + it('claude: passes a known context length through to CLAUDE_CODE_MAX_CONTEXT_TOKENS', () => { + sessionsToClean.push('sess-ctx-1'); + const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', 'sess-ctx-1', 16384); + expect(applied?.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe('16384'); + expect(applied?.envKeys).toContain('CLAUDE_CODE_MAX_CONTEXT_TOKENS'); + }); + + it('claude: omits the var entirely when the context length is unknown', () => { + sessionsToClean.push('sess-ctx-2'); + const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', 'sess-ctx-2'); + expect(applied?.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBeUndefined(); + }); + + it('deepseek: has no contextLengthVar declared, so a passed-in length is a no-op', () => { + const applied = applyCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3', 'sess-ctx-3', 16384); + expect(Object.keys(applied?.envOverrides ?? {}).sort()).toEqual(['DEEPSEEK_API_KEY', 'DEEPSEEK_BASE_URL']); + }); +}); + +describe('applyCustomModelInjection: CLAUDE_CONFIG_DIR isolation', () => { + it('claude: creates an isolated, empty config dir and points CLAUDE_CONFIG_DIR at it', () => { + const sessionId = 'sess-cfgdir-1'; + sessionsToClean.push(sessionId); + const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId); + const expectedDir = customModelConfigDir(sessionId); + expect(applied?.envOverrides.CLAUDE_CONFIG_DIR).toBe(expectedDir); + expect(applied?.configDir).toBe(expectedDir); + expect(existsSync(expectedDir)).toBe(true); + // No credential/config files written into it — isolation, not a real config copy. + const entries = readdirSync(expectedDir).filter((name) => name !== 'projects'); + expect(entries).toEqual([]); + }); + + it('claude: symlinks (or junctions) projects back to the real config dir so the response viewer keeps working', () => { + const sessionId = 'sess-cfgdir-2'; + sessionsToClean.push(sessionId); + const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId); + const link = join(applied!.configDir!, 'projects'); + // Best-effort: only assert the link exists if it was actually created (the real + // ~/.claude/projects may not exist on a bare CI box, in which case linking is skipped). + if (existsSync(join(homedir(), '.claude', 'projects'))) { + expect(existsSync(link)).toBe(true); + expect(lstatSync(link).isSymbolicLink() || lstatSync(link).isDirectory()).toBe(true); + } + }); + + it('claude: re-applying to the same session is idempotent (boot-recovery re-apply)', () => { + const sessionId = 'sess-cfgdir-3'; + sessionsToClean.push(sessionId); + const first = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId); + const second = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId); + expect(second?.configDir).toBe(first?.configDir); + expect(existsSync(first!.configDir!)).toBe(true); + }); + + it('pi: configDir-kind CLIs are unaffected — no configDirVar concept for them', () => { + const sessionId = 'sess-cfgdir-pi'; + sessionsToClean.push(sessionId); + const applied = applyCustomModelInjection(entryOrThrow('pi'), endpoint, 'qwen3', sessionId); + expect(applied?.envOverrides.HOME).toBe(customModelConfigDir(sessionId)); + }); + + it('deepseek: no configDirVar declared, so no config dir is created at all', () => { + const sessionId = 'sess-cfgdir-deepseek'; + const applied = applyCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3', sessionId); + expect(applied?.configDir).toBeUndefined(); + expect(existsSync(customModelConfigDir(sessionId))).toBe(false); + }); +}); + +describe('applyCustomModelInjection: pre-existing behavior unaffected', () => { + it('opencode: still returns a plain env-kind result with no configDir', () => { + const sessionId = 'sess-opencode-1'; + const applied = applyCustomModelInjection(entryOrThrow('opencode'), endpoint, 'qwen3', sessionId); + expect(applied?.configDir).toBeUndefined(); + expect(applied?.envOverrides.OPENCODE_CONFIG_CONTENT).toBeTruthy(); + }); + + it('antigravity: still undefined (unsupported)', () => { + const applied = applyCustomModelInjection(entryOrThrow('antigravity'), endpoint, 'qwen3', 'sess-agy-1'); + expect(applied).toBeUndefined(); + }); +}); diff --git a/test/custom-model-injection.test.ts b/test/custom-model-injection.test.ts index 2acd08ca..1989c0f1 100644 --- a/test/custom-model-injection.test.ts +++ b/test/custom-model-injection.test.ts @@ -58,6 +58,24 @@ describe('buildCustomModelInjection', () => { }); }); + it('claude: also declares configDirVar (CLAUDE_CONFIG_DIR isolation) on the env-kind result', () => { + const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3'); + if (result.kind !== 'env') throw new Error('unreachable'); + expect(result.configDirVar).toBe('CLAUDE_CONFIG_DIR'); + }); + + it('claude: injects CLAUDE_CODE_MAX_CONTEXT_TOKENS when a context length is known', () => { + const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', 16384); + if (result.kind !== 'env') throw new Error('unreachable'); + expect(result.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe('16384'); + }); + + it('claude: omits CLAUDE_CODE_MAX_CONTEXT_TOKENS when the context length is unknown', () => { + const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3'); + if (result.kind !== 'env') throw new Error('unreachable'); + expect(result.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBeUndefined(); + }); + it('claude: falls back to a dummy key when the endpoint has none', () => { const result = buildCustomModelInjection(entryOrThrow('claude'), { ...endpoint, apiKey: undefined }, 'qwen3'); if (result.kind !== 'env') throw new Error('unreachable');