mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-06 15:39:41 +02:00
fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
@@ -261,6 +261,19 @@ const echoSchema = z
|
||||
})
|
||||
.strict();
|
||||
|
||||
/**
|
||||
* `capabilities.customModelInjection.launchModel`: the `model` launch-param value that
|
||||
* selects the injected provider, with `{modelId}` standing for the chosen id. Bounded to
|
||||
* the characters the `model`/`model-pi` token patterns accept plus the placeholder braces,
|
||||
* so a template can never smuggle a token the argv engine would have to quote.
|
||||
*/
|
||||
const launchModelTemplate = z
|
||||
.string()
|
||||
.min(1)
|
||||
.max(120)
|
||||
.regex(/^[a-zA-Z0-9._\-/:{}]+$/)
|
||||
.optional();
|
||||
|
||||
const capabilitiesSchema = z
|
||||
.object({
|
||||
external: z.boolean(),
|
||||
@@ -326,6 +339,7 @@ const capabilitiesSchema = z
|
||||
// Empty is valid: deepseek's model routing is a profile-composition concern, not
|
||||
// an env var, so it declares baseUrl/apiKey injection with no model var at all.
|
||||
modelVars: z.array(envName).max(8),
|
||||
launchModel: launchModelTemplate,
|
||||
})
|
||||
.strict(),
|
||||
z
|
||||
@@ -333,6 +347,7 @@ const capabilitiesSchema = z
|
||||
kind: z.literal('configContentEnv'),
|
||||
envVar: envName,
|
||||
template: z.literal('opencode-json'),
|
||||
launchModel: launchModelTemplate,
|
||||
})
|
||||
.strict(),
|
||||
z
|
||||
@@ -341,6 +356,7 @@ const capabilitiesSchema = z
|
||||
dirEnvVar: envName,
|
||||
fileName: z.string().min(1).max(80),
|
||||
template: z.enum(['codex-toml', 'pi-models-json', 'omp-models-yml', 'grok-toml']),
|
||||
launchModel: launchModelTemplate,
|
||||
})
|
||||
.strict(),
|
||||
z.object({ kind: z.literal('unsupported') }).strict(),
|
||||
|
||||
@@ -228,7 +228,7 @@ const CLAUDE: CliEntry = {
|
||||
privilegedParams: [],
|
||||
// ANTHROPIC_* is NOT in allowedPrefixes/allowedKeys above (deliberately — see the
|
||||
// allowedPrefixes comment nearby), so these are unreachable via plain envOverrides
|
||||
// today; listed here only so the dedicated custom-model route (deployment_plan.md
|
||||
// today; listed here only so the dedicated custom-model route (docs/custom-model-endpoints-plan.md
|
||||
// chunk 5) clamps them for a non-granted multi-user owner the same way every other
|
||||
// CLI's injection vars are clamped, the day that route widens who can set them.
|
||||
privilegedEnvKeys: [
|
||||
@@ -239,7 +239,7 @@ const CLAUDE: CliEntry = {
|
||||
'ANTHROPIC_DEFAULT_OPUS_MODEL',
|
||||
],
|
||||
gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } },
|
||||
// Custom Model Endpoint Profiles (deployment_plan.md) — verified by hand against a real
|
||||
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real
|
||||
// llama.cpp server. Claude reads these at process start only, so switching requires a
|
||||
// respawn, never a live hot-swap.
|
||||
customModelInjection: {
|
||||
@@ -781,6 +781,10 @@ const PI: CliEntry = {
|
||||
dirEnvVar: 'HOME',
|
||||
fileName: '.pi/agent/models.json',
|
||||
template: 'pi-models-json',
|
||||
// Writing models.json is not enough: without `--model custom/<id>` pi stays on its
|
||||
// own default provider and fails with "No API key found for the selected model"
|
||||
// (confirmed live). `custom` is the provider name pi-models-json declares.
|
||||
launchModel: 'custom/{modelId}',
|
||||
},
|
||||
// HOME is not `PI_`-prefixed, so unlike the old (wrong) PI_CONFIG_DIR guess this was
|
||||
// never reachable via the generic envOverrides allowlist at all — listed here anyway,
|
||||
@@ -895,6 +899,10 @@ const GROK: CliEntry = {
|
||||
dirEnvVar: 'GROK_HOME',
|
||||
fileName: 'config.toml',
|
||||
template: 'grok-toml',
|
||||
// The `[model.<name>]` block the grok-toml template writes; `--model <name>` is what
|
||||
// selects it (GROK_CUSTOM_MODEL_NAME in custom-model-injection.ts, pinned equal by
|
||||
// test/custom-model-injection.test.ts so the two cannot drift).
|
||||
launchModel: 'codeman-custom',
|
||||
},
|
||||
// GROK_HOME already matches the GROK_ allowedPrefix above, so it was ALREADY
|
||||
// reachable via plain envOverrides before this feature existed — same reasoning
|
||||
@@ -1189,6 +1197,9 @@ const OMP: CliEntry = {
|
||||
dirEnvVar: 'HOME',
|
||||
fileName: '.omp/agent/models.yml',
|
||||
template: 'omp-models-yml',
|
||||
// Same as pi: omp's own default model has no credential, so without an explicit
|
||||
// `--model custom/<id>` it never reaches the injected provider at all.
|
||||
launchModel: 'custom/{modelId}',
|
||||
},
|
||||
},
|
||||
overlays: {
|
||||
|
||||
@@ -460,7 +460,7 @@ export interface CliCapabilities {
|
||||
/**
|
||||
* How this CLI is pointed at a user-supplied custom OpenAI-compatible
|
||||
* endpoint (local, e.g. llama.cpp, or cloud, e.g. Azure AI Foundry) — the
|
||||
* Custom Model Endpoint Profiles feature (`deployment_plan.md`). Declared
|
||||
* Custom Model Endpoint Profiles feature (`docs/custom-model-endpoints-plan.md`). Declared
|
||||
* per entry, never branched on id, same as every other capability here.
|
||||
*
|
||||
* `env`: plain env vars (claude's `ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`/
|
||||
@@ -478,7 +478,7 @@ export interface CliCapabilities {
|
||||
* because those env vars are not grok's real custom-endpoint mechanism at
|
||||
* all. The real one is a `[model.<name>]` block in a `config.toml` under
|
||||
* `GROK_HOME` (verified against xAI's own docs), same shape as codex/pi/
|
||||
* omp — this is why the confidence table in deployment_plan.md exists:
|
||||
* omp — this is why the confidence table in docs/custom-model-endpoints-plan.md exists:
|
||||
* "researched" web docs can still be plausible-sounding and wrong.
|
||||
*
|
||||
* Every env var name this introduces that can redirect a session's
|
||||
@@ -486,15 +486,26 @@ export interface CliCapabilities {
|
||||
* `DEEPSEEK_BASE_URL` — a non-granted multi-user owner redirecting a
|
||||
* session to their own endpoint is a credential-exfiltration path, not
|
||||
* just a mischief redirect.
|
||||
*
|
||||
* `launchModel` is the value the entry's own `model` launch param must carry
|
||||
* for the CLI to SELECT the injected provider, as a template where
|
||||
* `{modelId}` is the chosen model id. Writing the config file is not enough
|
||||
* for pi and omp (`--model custom/<id>`, or the CLI stays on its own default
|
||||
* provider and reports "No API key found for the selected model") or for
|
||||
* grok (`--model codeman-custom`, the `[model.<name>]` block the config
|
||||
* declares). Absent = the config alone selects the model (claude's env vars,
|
||||
* opencode's blob, codex's top-level `model` key). Applied by the session's
|
||||
* respawn options through the entry's `legacyConfigField`, never by id.
|
||||
*/
|
||||
customModelInjection:
|
||||
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] }
|
||||
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' }
|
||||
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string }
|
||||
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string }
|
||||
| {
|
||||
kind: 'configDir';
|
||||
dirEnvVar: string;
|
||||
fileName: string;
|
||||
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' | 'grok-toml';
|
||||
launchModel?: string;
|
||||
}
|
||||
| { kind: 'unsupported' };
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user