fix(custom-model): isolate Claude config dir and inject real context length

Addresses two live-validation findings on the Run-menu custom-model picker:

1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still
   coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same
   config directory and warns about it (confirmed cosmetic - the API key
   wins for actual requests, verified via a real session's own API Usage
   Billing line). A custom-model claude session now gets an isolated
   CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty,
   no files written into it) so there is nothing to conflict with. projects
   is symlinked (junction on Windows) back into the real config dir so the
   response viewer, subagent windows and Read My Mind keep working for that
   session, best-effort.

2. Context-window overflow. Claude Code assumes a large default context
   window for a model id it doesn't recognise and never compacts, so a
   custom endpoint's real, much smaller context (verified live: a 400
   exceeding a 16384-token llama-swap model with a stock ~33.7K-token system
   prompt) silently overflows. Discovery now also learns each model's real
   context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx),
   but ONLY for a model llama-swap's own /v1/models response already marks
   status.value === 'loaded' - never an unloaded one, since llama-swap
   treats ?model= as a routing hint and probing an unloaded model risks
   triggering an actual, slow, GPU-swapping load as a side effect of
   read-only discovery. A server with no status field at all gets no
   enrichment rather than a guess; a model not probed this round keeps its
   previously-learned value until it disappears from the list entirely.
   Stored per model (CustomModelHost.modelContextLengths) and applied via a
   new contextLengthVar registry field, set to
   CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude.

Both new fields live on the existing env-kind customModelInjection
capability shape, declared only on claude's registry entry - every other
CLI's injection is unaffected (pinned by test).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
Devvyn
2026-09-16 12:08:26 +08:00
co-authored by Claude Sonnet 5
parent 5c25a52f95
commit 0e8b1981af
14 changed files with 504 additions and 20 deletions
+7
View File
@@ -340,6 +340,13 @@ const capabilitiesSchema = z
// an env var, so it declares baseUrl/apiKey injection with no model var at all.
modelVars: z.array(envName).max(8),
launchModel: launchModelTemplate,
// Optional: the env var to carry a discovered per-model context-window size
// (claude's CLAUDE_CODE_MAX_CONTEXT_TOKENS), and/or the env var that isolates
// this session's config/credential directory from the user's real one (claude's
// CLAUDE_CONFIG_DIR) so an injected API key never collides with a stored OAuth
// session. See the customModelInjection doc comment in cli-registry/types.ts.
contextLengthVar: envName.optional(),
configDirVar: envName.optional(),
})
.strict(),
z
+17
View File
@@ -237,6 +237,12 @@ const CLAUDE: CliEntry = {
'ANTHROPIC_DEFAULT_SONNET_MODEL',
'ANTHROPIC_DEFAULT_HAIKU_MODEL',
'ANTHROPIC_DEFAULT_OPUS_MODEL',
// CLAUDE_CODE_MAX_CONTEXT_TOKENS already matches the CLAUDE_CODE_* allowedPrefix, and
// CLAUDE_CONFIG_DIR is already an allowed exact key (docs/wiki/Agent-CLIs.md), so both
// were already reachable via plain envOverrides before this pair existed — listed here
// only so the custom-model route clamps them the same way as every other injected var.
'CLAUDE_CODE_MAX_CONTEXT_TOKENS',
'CLAUDE_CONFIG_DIR',
],
gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } },
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real
@@ -247,6 +253,17 @@ const CLAUDE: CliEntry = {
baseUrlVar: 'ANTHROPIC_BASE_URL',
apiKeyVar: 'ANTHROPIC_API_KEY',
modelVars: ['ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL'],
// Verified via Claude Code's own docs: CLAUDE_CODE_MAX_CONTEXT_TOKENS overrides the
// assumed context window and applies directly for a model name Claude Code doesn't
// recognize as one of its own — exactly the custom-model case. Without it, Claude Code
// assumes a large (200k) window for any unrecognized model id and never compacts,
// eventually overflowing a much smaller real local context (see plan doc reasoning
// above the interface for the confirmed failure).
contextLengthVar: 'CLAUDE_CODE_MAX_CONTEXT_TOKENS',
// Isolates this session's config/credential directory so an injected ANTHROPIC_API_KEY
// never shares a directory with a stored claude.ai OAuth login — see the doc comment on
// customModelInjection in cli-registry/types.ts for the traded-off side effect.
configDirVar: 'CLAUDE_CONFIG_DIR',
},
},
overlays: {
+25 -1
View File
@@ -496,9 +496,33 @@ export interface CliCapabilities {
* declares). Absent = the config alone selects the model (claude's env vars,
* opencode's blob, codex's top-level `model` key). Applied by the session's
* respawn options through the entry's `legacyConfigField`, never by id.
*
* `contextLengthVar` (env kind only): the env var a discovered per-model context-window
* size is written to when known (claude's `CLAUDE_CODE_MAX_CONTEXT_TOKENS`) — without it,
* a CLI that assumes a large default window for an unrecognized model name keeps sending
* full-size prompts against a much smaller local server and eventually overflows its real
* context (verified: a 33.7K-token system prompt against a 16384-token llama-swap model).
* Absent when the CLI has no such override, or the value is unknown for this model.
*
* `configDirVar` (env kind only): the env var that redirects this session's config/
* credential directory to an isolated, per-session one (claude's `CLAUDE_CONFIG_DIR`), so
* an injected API key never coexists with a stored claude.ai OAuth session in the same
* directory — the CLI still warns "both claude.ai and ANTHROPIC_API_KEY set" when they
* share a directory even though the API key wins for actual requests. Isolating it trades
* that cosmetic warning for a documented side effect: a relocated config directory writes
* transcripts outside `~/.claude/projects`, blinding the response viewer, subagent
* windows, and Read My Mind for that session (see docs/wiki/Agent-CLIs.md).
*/
customModelInjection:
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string }
| {
kind: 'env';
baseUrlVar: string;
apiKeyVar: string;
modelVars: string[];
launchModel?: string;
contextLengthVar?: string;
configDirVar?: string;
}
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string }
| {
kind: 'configDir';
+11
View File
@@ -50,6 +50,17 @@ export interface CustomModelHost {
* after a re-discover is a property worth keeping even if the model list changes.
*/
defaultModelId?: string;
/**
* Discovered context-window size (tokens) per model id, keyed by the same strings as
* `models`. Populated opportunistically during discovery (`custom-model-routes.ts`) from
* llama.cpp/llama-swap's `GET /props?model=<id>` — the plain OpenAI-shaped `/v1/models`
* response has no such field. Only ever probed for a model the server already reports as
* loaded (llama-swap's `status.value === 'loaded'`); an unloaded one is deliberately never
* probed, since llama-swap treats `/props?model=` as a routing hint that can trigger an
* actual (slow, GPU-swapping) model load as a side effect of merely asking. A model this
* has no entry for simply gets no context-length env override applied — never a guess.
*/
modelContextLengths?: Record<string, number>;
}
export function customModelHostsPath(configDir: string): string {
+52 -5
View File
@@ -10,7 +10,8 @@
* cli-registry changes" requirement it was written against.
*/
import { chmodSync, mkdirSync, writeFileSync, rmSync } from 'node:fs';
import { chmodSync, existsSync, mkdirSync, writeFileSync, rmSync, symlinkSync } from 'node:fs';
import { homedir, platform } from 'node:os';
import { join, dirname } from 'node:path';
import { dataPath } from './config/instance.js';
import type { CliEntry } from './config/cli-registry/types.js';
@@ -48,6 +49,36 @@ export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInj
return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv };
}
/**
* Real, shared Claude config directory Codeman's own host process runs under — honors
* `CLAUDE_CONFIG_DIR` the same way `claude-credentials.ts`'s `claudeCredentialsPath()`
* does, so the symlink below points at wherever `~/.claude/projects` actually lives
* rather than assuming the plain default.
*/
function realClaudeConfigDir(): string {
const configured = typeof process.env.CLAUDE_CONFIG_DIR === 'string' && process.env.CLAUDE_CONFIG_DIR.trim();
return configured || join(homedir(), '.claude');
}
/**
* Symlinks `<isolatedDir>/projects` back to the real, shared `~/.claude/projects`, so an
* isolated `CLAUDE_CONFIG_DIR` (used to keep an injected API key away from a stored OAuth
* session — see `configDirVar` on customModelInjection) doesn't also blind the response
* viewer, subagent windows, and Read My Mind for that session (docs/wiki/Agent-CLIs.md).
* Best-effort: a platform that refuses symlinks (unprivileged Windows without a junction
* fallback working, e.g.) just keeps the pre-existing documented side effect instead of
* failing the whole custom-model apply over a nice-to-have.
*/
function linkSharedProjectsDir(isolatedDir: string): void {
const link = join(isolatedDir, 'projects');
if (existsSync(link)) return; // already linked (idempotent re-apply) or real dir wrote one
try {
symlinkSync(join(realClaudeConfigDir(), 'projects'), link, platform() === 'win32' ? 'junction' : 'dir');
} catch {
// best-effort only — response viewer/subagent windows go blind for this session instead
}
}
/** Best-effort recursive removal of a previously-written configDir. Never throws. */
export function removeConfigDir(dir: string | undefined): void {
if (!dir) return;
@@ -79,14 +110,30 @@ export function applyCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string,
sessionId: string
sessionId: string,
/** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */
contextLength?: number
): AppliedCustomModel | undefined {
const injection = buildCustomModelInjection(entry, endpoint, modelId);
const injection = buildCustomModelInjection(entry, endpoint, modelId, contextLength);
if (injection.kind === 'unsupported') return undefined;
if (injection.kind === 'env') {
// `configDirVar` (claude's CLAUDE_CONFIG_DIR): point it at the same isolated,
// per-session directory the `configDir` kind uses, but write no files into it — an
// empty directory has no stored OAuth credential to conflict with the injected API
// key, which is the whole point. Reusing the same path keyed by sessionId keeps this
// idempotent across a boot-recovery re-apply, same as the configDir kind below.
let envOverrides = injection.envOverrides;
let configDir: string | undefined;
if (injection.configDirVar) {
configDir = customModelConfigDir(sessionId);
mkdirSync(configDir, { recursive: true, mode: 0o700 });
linkSharedProjectsDir(configDir);
envOverrides = { ...envOverrides, [injection.configDirVar]: configDir };
}
return {
envOverrides: injection.envOverrides,
envKeys: Object.keys(injection.envOverrides),
envOverrides,
envKeys: Object.keys(envOverrides),
configDir,
launchModel: injection.launchModel,
};
}
+16 -2
View File
@@ -44,6 +44,14 @@ export interface EnvInjection {
envOverrides: Record<string, string>;
/** See {@link ConfigDirInjection.launchModel}. */
launchModel?: string;
/**
* Name of the env var the caller should point at an isolated, credential-free config
* directory for this session (claude's `CLAUDE_CONFIG_DIR`), from the registry entry's
* `customModelInjection.configDirVar`. The actual directory value isn't computed here —
* this module is pure and has no sessionId to derive one from — the IO wrapper
* (`custom-model-injection-apply.ts`) creates it and adds it to `envOverrides`.
*/
configDirVar?: string;
}
export interface ConfigDirInjection {
@@ -90,7 +98,9 @@ function quoted(value: string): string {
export function buildCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string
modelId: string,
/** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */
contextLength?: number
): CustomModelInjectionResult {
const cap = entry.capabilities.customModelInjection;
const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY;
@@ -102,7 +112,11 @@ export function buildCustomModelInjection(
[cap.apiKeyVar]: apiKey,
};
for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId;
return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
if (cap.contextLengthVar && contextLength !== undefined && Number.isFinite(contextLength)) {
envOverrides[cap.contextLengthVar] = String(Math.trunc(contextLength));
}
const result = withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
return cap.configDirVar ? { ...result, configDirVar: cap.configDirVar } : result;
}
case 'configContentEnv': {
+78 -11
View File
@@ -26,6 +26,7 @@ import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } fro
const CODEMAN_CONFIG_DIR = getDataDir();
const DISCOVER_TIMEOUT_MS = 8000;
const PROPS_TIMEOUT_MS = 5000;
function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse<never> | null {
if (!isMultiUserMode() || isAdmin(req)) return null;
@@ -72,7 +73,7 @@ function applyStoredApiKey(incoming: CustomModelHost, existing: CustomModelHost)
return incoming.apiKey ? incoming : { ...incoming, apiKey: existing.apiKey };
}
async function discoverModels(host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' | 'authStyle'>): Promise<string[]> {
function authHeaders(host: Pick<CustomModelHost, 'apiKey' | 'authStyle'>): Record<string, string> {
const headers: Record<string, string> = {};
const apiKey = host.apiKey?.trim();
// Exactly ONE header, never both — see custom-model-hosts.ts's CustomModelAuthStyle
@@ -80,14 +81,73 @@ async function discoverModels(host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' |
const style = host.authStyle ?? 'bearer';
if (apiKey && style === 'bearer') headers.Authorization = `Bearer ${apiKey}`;
if (apiKey && style === 'api-key') headers['api-key'] = apiKey;
return headers;
}
export interface DiscoveryResult {
models: string[];
/** See `CustomModelHost.modelContextLengths` — only ever populated for models already loaded. */
contextLengths: Record<string, number>;
}
/**
* Best-effort: fetches `GET /props?model=<id>` (llama.cpp-native, llama-swap-proxied) for
* ONE already-loaded model and pulls its real `n_ctx` out. Never called for a model that
* isn't already loaded — see the caller and `CustomModelHost.modelContextLengths` for why
* that's a hard safety requirement, not just a nicety: llama-swap treats this endpoint's
* `?model=` as a routing hint, and asking it about an unloaded model risks triggering an
* actual (slow, GPU-swapping) load as a side effect of what should be read-only discovery.
* Any failure (unreachable, non-2xx, missing/malformed field) is swallowed — one model's
* context length is a nice-to-have, never worth failing the whole discovery pass over.
*/
async function fetchContextLength(
host: Pick<CustomModelHost, 'baseUrl'>,
modelId: string,
headers: Record<string, string>
): Promise<number | undefined> {
try {
const url = new URL(`${host.baseUrl.replace(/\/+$/, '')}/props`);
url.searchParams.set('model', modelId);
const res = await webviewFetch(url, { headers, signal: AbortSignal.timeout(PROPS_TIMEOUT_MS) });
if (!res.ok) return undefined;
const body = (await res.json()) as { n_ctx?: unknown; default_generation_settings?: { n_ctx?: unknown } };
const nCtx = body.n_ctx ?? body.default_generation_settings?.n_ctx;
return typeof nCtx === 'number' && Number.isFinite(nCtx) && nCtx > 0 ? nCtx : undefined;
} catch {
return undefined;
}
}
async function discoverModels(
host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' | 'authStyle'>
): Promise<DiscoveryResult> {
const headers = authHeaders(host);
const res = await webviewFetch(new URL(`${host.baseUrl.replace(/\/+$/, '')}/v1/models`), {
headers,
signal: AbortSignal.timeout(DISCOVER_TIMEOUT_MS),
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = (await res.json()) as { data?: Array<{ id?: unknown }> };
return (body.data ?? []).map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0);
const body = (await res.json()) as { data?: Array<{ id?: unknown; status?: { value?: unknown } }> };
const entries = body.data ?? [];
const models = entries.map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0);
// llama-swap-specific, feature-detected: a server that never mentions `status` on ANY
// entry gets no context-length enrichment at all, rather than treating "no status field"
// as "assume unloaded" — either reading is a guess, and skipping is the safe one, since
// fetchContextLength must only ever run against a model this server itself calls loaded.
const hasStatusField = entries.some((m) => m && typeof m === 'object' && 'status' in m);
const contextLengths: Record<string, number> = {};
if (hasStatusField) {
const loadedIds = entries
.filter((m) => m.status && typeof m.status === 'object' && (m.status as { value?: unknown }).value === 'loaded')
.map((m) => m.id)
.filter((id): id is string => typeof id === 'string' && id.length > 0);
for (const id of loadedIds) {
const ctx = await fetchContextLength(host, id, headers);
if (ctx !== undefined) contextLengths[id] = ctx;
}
}
return { models, contextLengths };
}
/**
@@ -117,9 +177,16 @@ type RedactedHost = ReturnType<typeof redactApiKey>;
* each do their own `discoverModels()` + error handling around one shared
* "how to apply a successful result" step.
*/
function applyDiscoveredModels(host: CustomModelHost, models: string[]): CustomModelHost {
function applyDiscoveredModels(host: CustomModelHost, result: DiscoveryResult): CustomModelHost {
const { models, contextLengths } = result;
const defaultModelId = host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined;
return { ...host, models, defaultModelId, lastDiscoveredAt: new Date().toISOString() };
// Merge onto what's already known rather than replacing: a model not probed this round
// (not currently loaded) keeps whatever context length an earlier round already learned
// for it, and one no longer in the fresh list is dropped, same reasoning as defaultModelId.
const merged = { ...host.modelContextLengths, ...contextLengths };
const kept = Object.fromEntries(Object.entries(merged).filter(([id]) => models.includes(id)));
const modelContextLengths = Object.keys(kept).length > 0 ? kept : undefined;
return { ...host, models, defaultModelId, modelContextLengths, lastDiscoveredAt: new Date().toISOString() };
}
/**
@@ -135,9 +202,9 @@ export async function refreshAllCustomModelHosts(): Promise<void> {
const hosts = await readCustomModelHosts(dataDir);
for (const host of hosts) {
if (isBlockedWebviewUrl(host.baseUrl)) continue;
let models: string[];
let result: DiscoveryResult;
try {
models = await discoverModels(host);
result = await discoverModels(host);
} catch {
continue; // unreachable this cycle — try again next tick, not fatal to the sweep
}
@@ -147,7 +214,7 @@ export async function refreshAllCustomModelHosts(): Promise<void> {
const current = await readCustomModelHosts(dataDir);
const index = current.findIndex((item) => item.id === host.id);
if (index === -1) continue; // deleted mid-sweep
current[index] = applyDiscoveredModels(current[index], models);
current[index] = applyDiscoveredModels(current[index], result);
await writeCustomModelHosts(dataDir, current);
}
}
@@ -222,11 +289,11 @@ export function registerCustomModelRoutes(app: FastifyInstance): void {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
try {
const models = await discoverModels(host);
const result = await discoverModels(host);
const next = [...hosts];
next[index] = applyDiscoveredModels(host, models);
next[index] = applyDiscoveredModels(host, result);
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
return { success: true, data: { models } };
return { success: true, data: { models: result.models } };
} catch (err) {
const blocked = egressBlockedReason(err);
return createErrorResponse(
+2 -1
View File
@@ -1214,7 +1214,8 @@ export function registerSessionRoutes(
// fails its pattern rather than quoting it, which would silently launch the CLI on its
// own default provider again, so refuse an id the pattern cannot carry up front.
const modelSpec = entry.launch.params.model;
const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id);
const contextLength = endpoint.modelContextLengths?.[body.modelId];
const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id, contextLength);
if (!applied) {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
}
+3
View File
@@ -1924,6 +1924,9 @@ export const CustomModelHostSchema = z.object({
// host to apply it, so the check belongs there once, not duplicated into a refine
// that would run on every unrelated field edit too).
defaultModelId: z.string().max(200).optional(),
// Server-populated by discovery (custom-model-routes.ts); accepted here only so a client
// round-tripping the GET response back through PUT (edit-save) doesn't drop it.
modelContextLengths: z.record(z.string().max(200), z.number().int().positive().max(100_000_000)).optional(),
});
/** POST /api/sessions/:id/custom-model — apply or clear a session's custom-model selection. */