mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
253 lines
11 KiB
TypeScript
253 lines
11 KiB
TypeScript
/**
|
|
* @fileoverview Pure builder for the Custom Model Endpoint Profiles feature
|
|
* (docs/custom-model-endpoints-plan.md): turns a CLI registry entry's
|
|
* `capabilities.customModelInjection` declaration, a configured endpoint,
|
|
* and a chosen model id into the concrete env vars / config-file content
|
|
* that would redirect that CLI's session at the endpoint.
|
|
*
|
|
* No IO here on purpose (mirrors `session-cli-builder.ts`) — a caller
|
|
* writes `ConfigDirInjection.files` to disk under an isolated per-session
|
|
* directory and points `dirEnvVar` at it; this module only computes what
|
|
* those files/env vars should contain.
|
|
*
|
|
* Confidence: `claude` and `opencode` are verified end-to-end against a real
|
|
* llama-swap server (a real "hello world" reply came back). `codex`'s
|
|
* config.toml STRUCTURE is now verified (an earlier `[model].default` table
|
|
* shape was rejected by a real codex binary with "invalid type: map,
|
|
* expected a string" — caught by `scripts/test-local-llm-harnesses.ts`),
|
|
* but `wire_api = "responses"` is the only value codex still accepts
|
|
* (support for `"chat"` was dropped in Feb 2026), and a plain OpenAI
|
|
* Chat-Completions server (llama.cpp, llama-swap, most local setups) does
|
|
* NOT implement the Responses API — so codex may still fail at the
|
|
* PROTOCOL level even with a correctly-shaped config file. That gap is
|
|
* real and current, not a stale warning; see docs/custom-model-endpoints-plan.md. The rest
|
|
* (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION flags
|
|
* confirmed against real installed binaries' own `--help` output, but
|
|
* their custom-endpoint env/config conventions remain web-researched,
|
|
* unverified.
|
|
*/
|
|
|
|
import type { CliEntry } from './config/cli-registry/types.js';
|
|
|
|
export interface CustomModelEndpoint {
|
|
id: string;
|
|
label: string;
|
|
/** Root URL, no trailing slash required — e.g. "http://192.168.1.50:8080" or an Azure AI Foundry URL. */
|
|
baseUrl: string;
|
|
/** Falls back to a harmless placeholder for endpoints (llama.cpp) that don't check it. */
|
|
apiKey?: string;
|
|
}
|
|
|
|
export interface EnvInjection {
|
|
kind: 'env';
|
|
/** Ready to merge into a session's envOverrides. */
|
|
envOverrides: Record<string, string>;
|
|
/** See {@link ConfigDirInjection.launchModel}. */
|
|
launchModel?: string;
|
|
}
|
|
|
|
export interface ConfigDirInjection {
|
|
kind: 'configDir';
|
|
/** Env var that must be set to the directory the caller writes `files` under. */
|
|
dirEnvVar: string;
|
|
files: Array<{ relPath: string; content: string }>;
|
|
/**
|
|
* Env vars the written config file REFERENCES by name rather than embedding a
|
|
* literal value (codex's `env_key = "..."` convention: config.toml never carries
|
|
* the API key itself, only the name of an env var codex reads it from). Merge
|
|
* these into the session's envOverrides alongside `dirEnvVar` — never skip them,
|
|
* or the config points at a credential that was never actually set.
|
|
*/
|
|
extraEnv?: Record<string, string>;
|
|
/**
|
|
* The value the CLI's `model` launch param must carry for it to SELECT the injected
|
|
* provider (pi/omp: `custom/<modelId>`; grok: the `[model.<name>]` block name). Absent
|
|
* when the config alone selects the model. Rendered from the registry entry's
|
|
* `customModelInjection.launchModel` template, never hand-built per CLI.
|
|
*/
|
|
launchModel?: string;
|
|
}
|
|
|
|
export interface UnsupportedInjection {
|
|
kind: 'unsupported';
|
|
}
|
|
|
|
export type CustomModelInjectionResult = EnvInjection | ConfigDirInjection | UnsupportedInjection;
|
|
|
|
const DEFAULT_API_KEY = 'local-dummy-key';
|
|
|
|
/** Normalizes a base URL to end in exactly one trailing `/v1`, for CLIs whose config expects the OpenAI-style suffix. */
|
|
export function withV1Suffix(baseUrl: string): string {
|
|
const trimmed = baseUrl.replace(/\/+$/, '');
|
|
return /\/v1$/.test(trimmed) ? trimmed : `${trimmed}/v1`;
|
|
}
|
|
|
|
/** JSON-escapes a string for embedding in a TOML/YAML double-quoted scalar — a safe superset of both grammars' basic escapes. */
|
|
function quoted(value: string): string {
|
|
return JSON.stringify(value);
|
|
}
|
|
|
|
export function buildCustomModelInjection(
|
|
entry: Pick<CliEntry, 'capabilities'>,
|
|
endpoint: CustomModelEndpoint,
|
|
modelId: string
|
|
): CustomModelInjectionResult {
|
|
const cap = entry.capabilities.customModelInjection;
|
|
const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY;
|
|
|
|
switch (cap.kind) {
|
|
case 'env': {
|
|
const envOverrides: Record<string, string> = {
|
|
[cap.baseUrlVar]: endpoint.baseUrl,
|
|
[cap.apiKeyVar]: apiKey,
|
|
};
|
|
for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId;
|
|
return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
|
|
}
|
|
|
|
case 'configContentEnv': {
|
|
const content = renderConfigContent(cap.template, endpoint, modelId, apiKey);
|
|
return withLaunchModel({ kind: 'env', envOverrides: { [cap.envVar]: content } }, cap.launchModel, modelId);
|
|
}
|
|
|
|
case 'configDir': {
|
|
const { content, extraEnv } = renderConfigFile(cap.template, endpoint, modelId, apiKey);
|
|
return withLaunchModel(
|
|
{ kind: 'configDir', dirEnvVar: cap.dirEnvVar, files: [{ relPath: cap.fileName, content }], extraEnv },
|
|
cap.launchModel,
|
|
modelId
|
|
);
|
|
}
|
|
|
|
case 'unsupported':
|
|
return { kind: 'unsupported' };
|
|
}
|
|
}
|
|
|
|
/** Render a `launchModel` template (`{modelId}` = the chosen id) onto an injection result. */
|
|
function withLaunchModel<T extends EnvInjection | ConfigDirInjection>(
|
|
result: T,
|
|
template: string | undefined,
|
|
modelId: string
|
|
): T {
|
|
if (!template) return result;
|
|
return { ...result, launchModel: template.split('{modelId}').join(modelId) };
|
|
}
|
|
|
|
function renderConfigContent(
|
|
template: 'opencode-json',
|
|
endpoint: CustomModelEndpoint,
|
|
modelId: string,
|
|
apiKey: string
|
|
): string {
|
|
switch (template) {
|
|
case 'opencode-json':
|
|
return JSON.stringify({
|
|
$schema: 'https://opencode.ai/config.json',
|
|
provider: {
|
|
custom: {
|
|
options: { baseURL: withV1Suffix(endpoint.baseUrl), apiKey },
|
|
models: { [modelId]: {} },
|
|
},
|
|
},
|
|
model: `custom/${modelId}`,
|
|
});
|
|
}
|
|
}
|
|
|
|
const CODEX_API_KEY_ENV_VAR = 'CODEMAN_CUSTOM_MODEL_API_KEY';
|
|
|
|
/** The `[model.<name>]` block name grok's config.toml uses for the injected model — also
|
|
* what `-m <name>` in the standalone script's ONE_SHOT argv must reference to select it. */
|
|
export const GROK_CUSTOM_MODEL_NAME = 'codeman-custom';
|
|
|
|
function renderConfigFile(
|
|
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' | 'grok-toml',
|
|
endpoint: CustomModelEndpoint,
|
|
modelId: string,
|
|
apiKey: string
|
|
): { content: string; extraEnv?: Record<string, string> } {
|
|
const baseUrl = withV1Suffix(endpoint.baseUrl);
|
|
switch (template) {
|
|
case 'codex-toml': {
|
|
// Verified against real codex (>= Feb 2026): `model` is a top-level STRING, never
|
|
// a `[model].default` table — codex rejects that with "invalid type: map, expected
|
|
// a string" (caught by scripts/test-local-llm-harnesses.ts against a real llama-swap
|
|
// server). The API key is NEVER a literal TOML field: codex's schema only supports
|
|
// `env_key`, the NAME of an env var it reads the credential from at runtime, so the
|
|
// actual value must ride along as an extra env var, never embedded in the file.
|
|
// ⚠️ `wire_api = "responses"` is the only value codex still accepts (it dropped
|
|
// `"chat"` support in Feb 2026) — a plain OpenAI Chat-Completions server (llama.cpp,
|
|
// llama-swap, most local setups) does NOT implement the Responses API, so this
|
|
// recipe may still fail at the PROTOCOL level even though the file now parses
|
|
// correctly. That is a real, currently-unresolved compatibility gap, not a syntax
|
|
// bug — track it before calling codex support done.
|
|
const content = [
|
|
`model = ${quoted(modelId)}`,
|
|
`model_provider = "custom"`,
|
|
'',
|
|
'[model_providers.custom]',
|
|
`name = "Custom Endpoint"`,
|
|
`base_url = ${quoted(baseUrl)}`,
|
|
`env_key = ${quoted(CODEX_API_KEY_ENV_VAR)}`,
|
|
`wire_api = "responses"`,
|
|
'',
|
|
].join('\n');
|
|
return { content, extraEnv: { [CODEX_API_KEY_ENV_VAR]: apiKey } };
|
|
}
|
|
case 'pi-models-json':
|
|
// Verified against pi's OWN bundled docs (models.md): `models` is an ARRAY of
|
|
// `{id: "..."}` objects, NOT an object keyed by model id — the earlier shape here
|
|
// silently loaded zero models ("No models available"), confirmed live. `authHeader:
|
|
// true` is required too: pi does not automatically send `Authorization: Bearer
|
|
// <apiKey>` just because `apiKey` is set (per the same doc) — without it, a real
|
|
// (non-llama.cpp) endpoint that actually checks the key would reject every request.
|
|
return {
|
|
content: JSON.stringify(
|
|
{
|
|
providers: {
|
|
custom: {
|
|
baseUrl,
|
|
apiKey,
|
|
api: 'openai-completions',
|
|
authHeader: true,
|
|
models: [{ id: modelId }],
|
|
},
|
|
},
|
|
},
|
|
null,
|
|
2
|
|
),
|
|
};
|
|
case 'omp-models-yml':
|
|
// Mirrors the pi-models-json fix above (omp shares pi's config lineage per
|
|
// CLAUDE.md — it reads several of pi's own env vars): a flat list of bare model
|
|
// name strings under `models` is UNCONFIRMED against real omp docs (none are
|
|
// bundled with the binary) — this now matches pi's `{id: "..."}` object-list
|
|
// shape and adds `authHeader: true` on the same reasoning, but has not itself
|
|
// been live-tested the way pi's fix was. Verify before raising its confidence.
|
|
return {
|
|
content: `providers:\n custom:\n baseUrl: ${quoted(baseUrl)}\n apiKey: ${quoted(apiKey)}\n api: openai-completions\n authHeader: true\n models:\n - id: ${quoted(modelId)}\n`,
|
|
};
|
|
case 'grok-toml': {
|
|
// Verified against xAI's own docs (docs.x.ai/build/settings/reference): a
|
|
// `[model.<name>]` block, NOT plain env vars — an earlier `env`-kind recipe for
|
|
// grok was wrong, not just unverified (see the customModelInjection doc comment
|
|
// in cli-registry/types.ts). `api_backend = "chat_completions"` is explicitly
|
|
// supported (unlike codex, which dropped it after Feb 2026), so this one CAN
|
|
// talk to a plain OpenAI-compatible server directly. `env_key` reuses grok's own
|
|
// documented fallback var name (XAI_API_KEY) rather than inventing a new one.
|
|
const content = [
|
|
`[model.${GROK_CUSTOM_MODEL_NAME}]`,
|
|
`model = ${quoted(modelId)}`,
|
|
`base_url = ${quoted(baseUrl)}`,
|
|
`name = "Custom Endpoint"`,
|
|
`env_key = "XAI_API_KEY"`,
|
|
`api_backend = "chat_completions"`,
|
|
'',
|
|
].join('\n');
|
|
return { content, extraEnv: { XAI_API_KEY: apiKey } };
|
|
}
|
|
}
|
|
}
|