diff --git a/CLAUDE.md b/CLAUDE.md index 4ef118e5..2583c903 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -227,7 +227,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. ⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. ⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. ⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs. -**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; backend + HTTP API only until the Run-menu picker lands, and the setting is read by nothing yet): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct but codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement — a confirmed protocol gap, not a bug; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. +**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct but codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement — a confirmed protocol gap, not a bug; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. ⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces, extended to the one frontend surface that needs to know which CLIs support this. One entry per (capable CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)"; picking one runs that CLI's own existing `run*()` function unmodified (case creation, env overrides, the works — forced to a single instance) and then calls `POST /api/sessions/:id/custom-model` on the session it selects, reusing the fact every `run*()` ends by selecting its new session rather than a parallel create path. `CustomModelHost.defaultModelId` is what the picker applies with no further choice — settings-ui.js's Edit form is a select populated from that endpoint's own discovered `models`, the route rejects a value that isn't a member, and re-discovery drops a stale one rather than carrying it forward invalid. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to default to). **Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w-` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization) diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index 975965ef..24acc171 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -11,20 +11,21 @@ company gateway) — anything answering `GET /v1/models` and recipe confidence table, and security reasoning: [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md). -> **Status**: backend is implemented and tested (registry capability, the -> injection engine, the endpoint store + discovery route, the session -> restart route). The toolbar picker / settings UI described below as the -> intended surface is **not yet built** — until it lands, use the HTTP API -> directly (examples below). Antigravity has no known custom-endpoint -> mechanism and is not supported. +> **Status**: fully wired end to end — registry capability, the injection +> engine, the endpoint store + discovery route, the session restart route, +> a settings-panel CRUD surface, and the Run-menu picker described below. +> Antigravity has no known custom-endpoint mechanism and is not supported. +> The HTTP API (examples below) still works directly and is what the picker +> itself calls under the hood. ## Turning it on -App Settings → Agents & CLIs → **Custom Model Endpoints** (synced setting -`customModelEndpointsEnabled`, default **OFF**). Until the toolbar picker -lands, nothing reads this setting: the HTTP routes below work whether it is -on or off, and it exists now only so the picker has a switch to hang off -when it ships. The API equivalent: +App Settings → Models → **Custom model endpoints** (synced setting +`customModelEndpointsEnabled`, default **OFF**). Turning it on does two +things: it reveals the endpoint list/add/edit/discover panel in that same +settings section, and it makes the Run menu offer a generated entry per +(harness, endpoint) pair — see "The Run-menu picker" below. The API +equivalent: ```bash curl -sk -X PUT https://localhost:3000/api/settings \ @@ -34,6 +35,9 @@ curl -sk -X PUT https://localhost:3000/api/settings \ ## Adding an endpoint +Via App Settings → Models → Custom model endpoints → **+ Add endpoint**, or +directly: + ```bash curl -sk -X POST https://localhost:3000/api/model-endpoints \ -H 'Content-Type: application/json' \ @@ -62,6 +66,34 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one. Endpoint management is admin-only in multi-user mode, same as remote/docker hosts — these are machine-level infra, not per-user settings. +`defaultModelId` names which discovered model the Run-menu picker applies +for that endpoint with no further choice — the settings panel's Edit form +exposes it as a select populated from the endpoint's own discovered +`models`, and the route refuses a value that isn't one of them. Leaving it +unset falls back to the first discovered model; re-discovering drops a +default that no longer appears in the fresh list rather than carrying an +invalid one forward. + +## The Run-menu picker + +With the setting on and at least one endpoint carrying a usable default +model (either an explicit `defaultModelId` or just one discovered model), +the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry +per (harness that can redirect to a custom endpoint, saved endpoint) pair, +e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI +registry's own `capabilities.customModelInjection` at page render +(`window.__codemanCustomModelClis`, `server.ts`) — never a hardcoded id list +in the frontend — so a CLI whose injection recipe lands later shows up with +no frontend change, and Antigravity (`unsupported`) never does. + +Picking an entry runs a single session on that harness exactly the way its +own Run-menu entry would (same case creation, env overrides, everything), +then immediately applies the endpoint's default model to it via the route +below. It is a one-off "try this endpoint" action, not a sticky mode: the +plain Run button still means "this harness, native cloud" afterward. +Entries are hidden entirely for a remote or Docker active case, since the +apply route refuses both (see the next section). + ## Applying a model to a session ```bash diff --git a/src/custom-model-hosts.ts b/src/custom-model-hosts.ts index 0cde1039..0130eee7 100644 --- a/src/custom-model-hosts.ts +++ b/src/custom-model-hosts.ts @@ -40,6 +40,16 @@ export interface CustomModelHost { authStyle?: CustomModelAuthStyle; models?: string[]; lastDiscoveredAt?: string; + /** + * The model the Run-menu picker (docs/custom-model-endpoints-plan.md) applies when + * this endpoint is picked with no further choice — one generated menu entry per + * (CLI, endpoint) pair, not per (CLI, endpoint, model), so it needs a single answer. + * Must be a member of `models` when set; the picker falls back to `models[0]` when + * this is unset, and disables the entry entirely when `models` is empty (nothing to + * default to). Never auto-set on discovery — the previous default staying valid + * after a re-discover is a property worth keeping even if the model list changes. + */ + defaultModelId?: string; } export function customModelHostsPath(configDir: string): string { diff --git a/src/web/public/index.html b/src/web/public/index.html index 2ba80243..2e6fded1 100644 --- a/src/web/public/index.html +++ b/src/web/public/index.html @@ -650,6 +650,14 @@ + + + +
+ + + diff --git a/src/web/public/session-ui.js b/src/web/public/session-ui.js index 67cbfad4..de783959 100644 --- a/src/web/public/session-ui.js +++ b/src/web/public/session-ui.js @@ -461,6 +461,7 @@ Object.assign(CodemanApp.prototype, { if (menu.classList.contains('active')) { this._loadRunModeHistory(); this._refreshRunModeAvailability(menu); + this._refreshCustomModelRunOptions(menu); const close = (ev) => { if (!menu.contains(ev.target)) { menu.classList.remove('active'); @@ -534,6 +535,124 @@ Object.assign(CodemanApp.prototype, { if (dsWeb) dsWeb.style.display = avail.deepseekBinary ? 'flex' : 'none'; }, + /** + * Generates the Run menu's Custom Model Endpoint entries + * (docs/custom-model-endpoints-plan.md): one button per (capable harness, saved + * endpoint) pair, e.g. "Claude Code (llama.cpp)". Hidden entirely when the + * feature is off, no endpoint has a usable default model, or the active case is + * remote/docker (the apply route refuses both — see session-routes.ts). + * + * `window.__codemanCustomModelClis` is server-injected at render time from the + * CLI registry's own `capabilities.customModelInjection` (never a hardcoded id + * list here), so a CLI gaining or losing the capability shows up with no + * frontend change. + */ + async _refreshCustomModelRunOptions(menu) { + const sep = menu.querySelector('#runModeCustomModelSep'); + const header = menu.querySelector('#runModeCustomModelHeader'); + const container = menu.querySelector('#runModeCustomModels'); + if (!container) return; + const hide = () => { + if (sep) sep.style.display = 'none'; + if (header) header.style.display = 'none'; + container.innerHTML = ''; + }; + + const settings = this.loadAppSettingsFromStorage(); + const capableClis = window.__codemanCustomModelClis || []; + if (!settings.customModelEndpointsEnabled || capableClis.length === 0) return hide(); + + const caseName = document.getElementById('quickStartCase')?.value; + const activeCase = caseName ? (this.cases || []).find((c) => c.name === caseName) : null; + if (activeCase?.location === 'remote' || activeCase?.location === 'docker') return hide(); + + let hosts; + try { + const res = await fetch('/api/model-endpoints'); + hosts = await res.json(); + } catch { + return hide(); + } + if (!Array.isArray(hosts) || hosts.length === 0) return hide(); + + const rows = []; + for (const host of hosts) { + const modelId = host.defaultModelId || (host.models || [])[0]; + if (!modelId) continue; // nothing discovered yet — the settings panel explains why + for (const cli of capableClis) { + rows.push(` + `); + } + } + if (rows.length === 0) return hide(); + if (sep) sep.style.display = ''; + if (header) header.style.display = ''; + container.innerHTML = rows.join(''); + }, + + /** + * Runs a session on `mode` and immediately applies `endpointId`/`modelId` to it + * via POST /api/sessions/:id/custom-model (see session-routes.ts) — the same + * restart-in-place apply path the (not-yet-built) endpoint-management surface + * would use for an already-running session. Reuses the existing per-mode run*() + * functions wholesale (case creation, env overrides, the works) rather than a + * parallel create path, forcing a single instance: a custom-model run is a + * one-off "try this endpoint" action, not a batch spawn. + */ + async runCustomModelEntry(mode, endpointId, modelId) { + document.getElementById('runModeMenu')?.classList.remove('active'); + const runners = { + claude: () => this.runClaude(), + opencode: () => this.runOpenCode(), + codex: () => this.runCodex(), + gemini: () => this.runGemini(), + pi: () => this.runPi(), + grok: () => this.runGrok(), + deepseek: () => this.runDeepSeek(), + omp: () => this.runOmp(), + }; + const runner = runners[mode]; + if (!runner) { + this.showToast(`No run function for mode ${mode}`, 'error'); + return; + } + + const tabCountEl = document.getElementById('tabCount'); + const prevTabCount = tabCountEl?.value; + if (tabCountEl) tabCountEl.value = '1'; + try { + await runner(); + } finally { + if (tabCountEl && prevTabCount !== undefined) tabCountEl.value = prevTabCount; + } + + // Every run*() ends by selecting the session it just created, so the active + // session at this point IS the new one — see runClaude/runShell's own comments + // on why selectSession must run before this reads activeSessionId. + const sessionId = this.activeSessionId; + if (!sessionId) return; // run() already reported its own error via toast + + try { + const res = await fetch(`/api/sessions/${sessionId}/custom-model`, { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify({ endpointId, modelId }), + }); + const data = await res.json(); + if (!data.success) { + this.showToast(`Session started on the native backend — could not apply the custom endpoint: ${data.error}`, 'warning'); + return; + } + this.showToast(`Pointed at ${endpointId} — restarting the session...`, 'info'); + } catch (err) { + this.showToast(`Session started, but applying the custom endpoint failed: ${err.message}`, 'warning'); + } + }, + /** * Start the DeepSeek Harness browser UI and open it as a Codeman web tab. * diff --git a/src/web/public/settings-ui.js b/src/web/public/settings-ui.js index 79bee3c0..19a2309b 100644 --- a/src/web/public/settings-ui.js +++ b/src/web/public/settings-ui.js @@ -395,6 +395,10 @@ Object.assign(CodemanApp.prototype, { document.getElementById('appSettingsShowUltracodeAgents').checked = settings.showUltracodeAgents ?? defaults.showUltracodeAgents ?? false; // Approvals Inbox: synced, default OFF (opt-in; only an explicit true enables). document.getElementById('appSettingsApprovalsInbox').checked = settings.approvalsInboxEnabled === true; + // Custom Model Endpoint Profiles: synced, default OFF. The toggle governs both + // the Run-menu picker's generated entries and this settings panel's visibility; + // the endpoint list itself is server state, loaded separately below. + document.getElementById('appSettingsCustomModelEndpoints').checked = settings.customModelEndpointsEnabled === true; // Read My Mind: synced, default OFF (opt-in; capture + prediction cost real tokens). document.getElementById('appSettingsReadMyMind').checked = settings.readMyMindEnabled === true; document.getElementById('appSettingsUltracodeFloatingWindows').checked = @@ -509,6 +513,10 @@ Object.assign(CodemanApp.prototype, { document.getElementById('appSettingsNiceValue').value = niceSettings.niceValue ?? 10; // Model configuration (loaded from server) this.loadModelConfigForSettings(); + // Custom Model Endpoint Profiles: server state, own load path (mirrors the + // model-config pair above) rather than the settings payload — endpoints are + // infra records (CRUD'd via /api/model-endpoints), not user preferences. + this.loadCustomModelEndpointsForSettings(); // Notification settings const notifPrefs = this.notificationManager?.preferences || {}; document.getElementById('appSettingsNotifEnabled').checked = notifPrefs.enabled ?? true; @@ -2106,6 +2114,7 @@ Object.assign(CodemanApp.prototype, { showSubagents: document.getElementById('appSettingsShowSubagents').checked, showUltracodeAgents: document.getElementById('appSettingsShowUltracodeAgents').checked, approvalsInboxEnabled: document.getElementById('appSettingsApprovalsInbox').checked, + customModelEndpointsEnabled: document.getElementById('appSettingsCustomModelEndpoints').checked, readMyMindEnabled: document.getElementById('appSettingsReadMyMind').checked, ultracodeFloatingWindows: document.getElementById('appSettingsUltracodeFloatingWindows').checked, showMultiMonitorButton: document.getElementById('appSettingsShowMultiMonitorButton').checked, @@ -2487,6 +2496,161 @@ Object.assign(CodemanApp.prototype, { } }, + // ═══════════════════════════════════════════════════════════════ + // Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) + // + // CRUD against /api/model-endpoints, rendered into the Models settings section. + // Deliberately its own load/save pair rather than folded into openAppSettings/ + // saveAppSettings: these are server-side infra records (like remote/docker + // hosts), not a settings-payload field, so the app-settings-structure guard's + // by-id contract does not apply to them — only the `customModelEndpointsEnabled` + // toggle itself goes through that path. + // ═══════════════════════════════════════════════════════════════ + + async loadCustomModelEndpointsForSettings() { + try { + const res = await fetch('/api/model-endpoints'); + const hosts = await res.json(); + this._customModelHosts = Array.isArray(hosts) ? hosts : []; + } catch (err) { + console.warn('Failed to load model endpoints:', err); + this._customModelHosts = this._customModelHosts || []; + } + this.renderCustomModelHostsList(); + }, + + renderCustomModelHostsList() { + const list = document.getElementById('customModelHostsList'); + if (!list) return; + const hosts = this._customModelHosts || []; + if (hosts.length === 0) { + list.innerHTML = '

No endpoints yet. Add one below to point a harness at a local or cloud OpenAI-compatible server.

'; + return; + } + list.innerHTML = hosts + .map((h) => { + const modelCount = (h.models || []).length; + const modelSummary = modelCount === 0 + ? 'No models discovered yet' + : `${modelCount} model${modelCount === 1 ? '' : 's'}${h.defaultModelId ? ` · default: ${escapeHtml(h.defaultModelId)}` : ' · no default set'}`; + return ` +
+
+ ${escapeHtml(h.label)} + ${escapeHtml(h.baseUrl)} — ${modelSummary} +
+
+ + + +
+
`; + }) + .join(''); + }, + + /** Opens the inline add/edit form. Pass no id to add a new endpoint. */ + openCustomModelHostEditor(hostId) { + const host = hostId ? (this._customModelHosts || []).find((h) => h.id === hostId) : null; + this._editingCustomModelHostId = host ? host.id : null; + document.getElementById('customModelHostEditorTitle').textContent = host ? `Edit ${host.label}` : 'Add endpoint'; + document.getElementById('customModelHostId').value = host?.id || ''; + document.getElementById('customModelHostId').disabled = !!host; // id is immutable once created + document.getElementById('customModelHostLabel').value = host?.label || ''; + document.getElementById('customModelHostBaseUrl').value = host?.baseUrl || ''; + document.getElementById('customModelHostApiKey').value = ''; // never round-tripped back into the field + document.getElementById('customModelHostApiKey').placeholder = host?.apiKey ? '•••••••• (unchanged if left blank)' : ''; + document.getElementById('customModelHostAuthStyle').value = host?.authStyle || 'bearer'; + this._populateCustomModelDefaultSelect(host); + document.getElementById('customModelHostEditor').style.display = ''; + }, + + closeCustomModelHostEditor() { + document.getElementById('customModelHostEditor').style.display = 'none'; + this._editingCustomModelHostId = null; + }, + + _populateCustomModelDefaultSelect(host) { + const select = document.getElementById('customModelHostDefaultModel'); + const models = host?.models || []; + select.innerHTML = + '' + + models.map((m) => ``).join(''); + select.value = host?.defaultModelId || ''; + select.disabled = models.length === 0; + }, + + async saveCustomModelHostFromEditor() { + const id = document.getElementById('customModelHostId').value.trim(); + const label = document.getElementById('customModelHostLabel').value.trim(); + const baseUrl = document.getElementById('customModelHostBaseUrl').value.trim(); + const apiKeyInput = document.getElementById('customModelHostApiKey').value; + const authStyle = document.getElementById('customModelHostAuthStyle').value; + const defaultModelId = document.getElementById('customModelHostDefaultModel').value || undefined; + if (!id || !label || !baseUrl) { + this.showToast('Id, label and base URL are all required', 'warning'); + return; + } + const editing = this._editingCustomModelHostId; + const existing = editing ? (this._customModelHosts || []).find((h) => h.id === editing) : null; + const body = { + id, + label, + baseUrl, + authStyle, + defaultModelId, + // A blank key on EDIT means "leave it alone", never "clear it" — the field + // is never pre-filled with the real value (see openCustomModelHostEditor), + // so an unedited save must not silently wipe a working credential. + apiKey: apiKeyInput ? apiKeyInput : existing?.apiKey, + models: existing?.models, + lastDiscoveredAt: existing?.lastDiscoveredAt, + }; + try { + const res = await fetch(editing ? `/api/model-endpoints/${encodeURIComponent(editing)}` : '/api/model-endpoints', { + method: editing ? 'PUT' : 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify(body), + }); + const data = await res.json(); + if (!data.success) { + this.showToast(data.error || 'Failed to save endpoint', 'error'); + return; + } + this.showToast(editing ? 'Endpoint updated' : 'Endpoint added', 'success'); + this.closeCustomModelHostEditor(); + await this.loadCustomModelEndpointsForSettings(); + } catch (err) { + this.showToast(`Failed to save endpoint: ${err.message}`, 'error'); + } + }, + + async discoverCustomModelHostModels(hostId) { + this.showToast('Discovering models…', 'info'); + try { + const res = await fetch(`/api/model-endpoints/${encodeURIComponent(hostId)}/discover-models`, { method: 'POST' }); + const data = await res.json(); + if (!data.success) { + this.showToast(data.error || 'Discovery failed', 'error'); + return; + } + this.showToast(`Found ${data.data.models.length} model${data.data.models.length === 1 ? '' : 's'}`, 'success'); + await this.loadCustomModelEndpointsForSettings(); + } catch (err) { + this.showToast(`Discovery failed: ${err.message}`, 'error'); + } + }, + + async deleteCustomModelHost(hostId) { + const host = (this._customModelHosts || []).find((h) => h.id === hostId); + if (!confirm(`Delete endpoint "${host?.label || hostId}"? Any session currently pointed at it keeps running until cleared.`)) return; + try { + await fetch(`/api/model-endpoints/${encodeURIComponent(hostId)}`, { method: 'DELETE' }); + await this.loadCustomModelEndpointsForSettings(); + } catch (err) { + this.showToast(`Failed to delete endpoint: ${err.message}`, 'error'); + } + }, // ═══════════════════════════════════════════════════════════════ // Visibility Settings & Device-Specific Defaults diff --git a/src/web/public/styles.css b/src/web/public/styles.css index 398eb6ce..2fe9fb65 100644 --- a/src/web/public/styles.css +++ b/src/web/public/styles.css @@ -16207,6 +16207,26 @@ html[data-tab-orientation='vertical'] .home-sessions { gap: 3px; } +/* Custom Model Endpoint Profiles' inline add/edit form: a nested panel rather + than a modal, so it needs its own border to read as a distinct sub-section + inside .set-group-body's flat row stack. */ +:is(#appSettingsModal, #sessionOptionsModal, #createCaseModal) .set-inline-form { + display: flex; + flex-direction: column; + gap: 3px; + margin-top: 6px; + padding: 10px 12px; + border: 1px solid var(--border); + border-radius: 8px; + background: rgba(0, 0, 0, 0.12); +} + +:is(#appSettingsModal, #sessionOptionsModal, #createCaseModal) .set-inline-form h5 { + margin: 0 0 4px; + font-size: 0.72rem; + color: var(--text-muted); +} + /* ── rows ─────────────────────────────────────────────────────────────── */ :is(#appSettingsModal, #sessionOptionsModal, #createCaseModal) .set-row { display: flex; diff --git a/src/web/routes/custom-model-routes.ts b/src/web/routes/custom-model-routes.ts index d5b1b1e8..0fc1b1f2 100644 --- a/src/web/routes/custom-model-routes.ts +++ b/src/web/routes/custom-model-routes.ts @@ -33,6 +33,21 @@ function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }) return createErrorResponse(ApiErrorCode.FORBIDDEN, 'Admin only in multi-user mode'); } +/** + * `defaultModelId` names the model the Run-menu picker applies for this endpoint with + * no further choice, so it must actually be one of the discovered `models` — a schema + * `.refine()` can't see across the two fields the way this can, and would also run on + * every unrelated field edit rather than only when either of these two changes. + */ +function invalidDefaultModel(host: Pick): ApiResponse | null { + if (host.defaultModelId === undefined) return null; + if ((host.models ?? []).includes(host.defaultModelId)) return null; + return createErrorResponse( + ApiErrorCode.INVALID_INPUT, + 'defaultModelId must be one of the endpoint’s discovered models' + ); +} + async function discoverModels(host: Pick): Promise { const headers: Record = {}; const apiKey = host.apiKey?.trim(); @@ -78,6 +93,8 @@ export function registerCustomModelRoutes(app: FastifyInstance): void { if (isBlockedWebviewUrl(host.baseUrl)) { return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed'); } + const badDefault = invalidDefaultModel(host); + if (badDefault) return badDefault; const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR); if (hosts.some((item) => item.id === host.id)) { return createErrorResponse(ApiErrorCode.ALREADY_EXISTS, 'Model endpoint already exists'); @@ -94,6 +111,8 @@ export function registerCustomModelRoutes(app: FastifyInstance): void { if (isBlockedWebviewUrl(host.baseUrl)) { return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed'); } + const badDefault = invalidDefaultModel(host); + if (badDefault) return badDefault; const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR); const index = hosts.findIndex((item) => item.id === id); if (index === -1) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found'); @@ -131,7 +150,12 @@ export function registerCustomModelRoutes(app: FastifyInstance): void { try { const models = await discoverModels(host); const next = [...hosts]; - next[index] = { ...host, models, lastDiscoveredAt: new Date().toISOString() }; + // A default that no longer appears in the fresh list would leave the Run-menu + // picker applying a model id the endpoint just told us it doesn't serve; drop + // it rather than carry it forward silently invalid. + const defaultModelId = + host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined; + next[index] = { ...host, models, defaultModelId, lastDiscoveredAt: new Date().toISOString() }; await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next); return { success: true, data: { models } }; } catch (err) { diff --git a/src/web/schemas.ts b/src/web/schemas.ts index b5dabdd0..f1f68151 100644 --- a/src/web/schemas.ts +++ b/src/web/schemas.ts @@ -1918,6 +1918,12 @@ export const CustomModelHostSchema = z.object({ authStyle: z.enum(['bearer', 'api-key']).optional(), models: z.array(z.string().max(200)).max(200).optional(), lastDiscoveredAt: z.string().max(64).optional(), + // The Run-menu picker's per-endpoint default; validated against `models` at the + // route layer (schema-level cross-field checks can't see the array narrowed the + // same way a `.refine()` closure could, and the route already re-reads the stored + // host to apply it, so the check belongs there once, not duplicated into a refine + // that would run on every unrelated field edit too). + defaultModelId: z.string().max(200).optional(), }); /** POST /api/sessions/:id/custom-model — apply or clear a session's custom-model selection. */ diff --git a/src/web/server.ts b/src/web/server.ts index cf5927c0..54c6b4ab 100644 --- a/src/web/server.ts +++ b/src/web/server.ts @@ -67,7 +67,7 @@ import { import { imageWatcher } from '../image-watcher.js'; import { workflowRunWatcher, summarizeRun } from '../workflow-run-watcher.js'; import { attachmentRegistry, buildFileThumbnailRoute, registerExternalAttachment } from '../attachment-registry.js'; -import { getCli } from '../config/cli-registry/registry.js'; +import { getCli, enabledClis } from '../config/cli-registry/registry.js'; import { readCustomModelHosts } from '../custom-model-hosts.js'; import { applyCustomModelInjection, customModelConfigDir, removeConfigDir } from '../custom-model-injection-apply.js'; import type { CustomModelBookkeeping } from '../types/session.js'; @@ -1596,6 +1596,18 @@ export class WebServer extends EventEmitter { '', `\n` ); + // Which run modes the Run-menu picker (docs/custom-model-endpoints-plan.md) may + // generate an entry for: read generically off the registry's `capabilities` + // (never an id list here) so a CLI whose customModelInjection lands later shows + // up in the picker with no frontend change, and one that ships `unsupported` + // (antigravity, and `shell`'s `kind !== 'agent'`) never does. + const customModelClis = enabledClis() + .filter((entry) => entry.kind === 'agent' && entry.capabilities.customModelInjection.kind !== 'unsupported') + .map((entry) => ({ id: entry.id, label: entry.label })); + html = html.replace( + '', + `\n` + ); } if (!soloSessionId && process.env.CODEMAN_GESTURE === '1') { html = html.replace('', `\n`); diff --git a/test/render-index-html.test.ts b/test/render-index-html.test.ts index 0cb6b334..df5fb417 100644 --- a/test/render-index-html.test.ts +++ b/test/render-index-html.test.ts @@ -187,6 +187,28 @@ describe('WebServer.renderIndexHtml', () => { }); }); + it('reports which run modes the custom-model Run-menu picker may generate an entry for', async () => { + // Read generically off the CLI registry's own capabilities, not a hardcoded id + // list — antigravity (`unsupported`) and shell (`kind !== 'agent'`) must be + // absent, and any enabled agent CLI with a real injection recipe must be + // present, with no mock needed since this reads the real stock registry. + const { server } = makeServer({}); + const html = await render(server); + expect(html).toContain('window.__codemanCustomModelClis='); + const clis = JSON.parse(html.match(/window\.__codemanCustomModelClis=(\[.*?\]);/)![1]) as Array<{ + id: string; + label: string; + }>; + const ids = clis.map((c) => c.id); + expect(ids).toContain('claude'); + expect(ids).not.toContain('antigravity'); + expect(ids).not.toContain('shell'); + for (const cli of clis) { + expect(typeof cli.id).toBe('string'); + expect(typeof cli.label).toBe('string'); + } + }); + it('still emits the object when nothing at all is installed', async () => { // The all-false case is the one that matters most and the easiest to get // wrong by only injecting when something resolves. @@ -218,6 +240,7 @@ describe('WebServer.renderIndexHtml', () => { const { server } = makeServer({}); const html = await render(server, 'sess-123'); expect(html).not.toContain('__codemanCliAvailable'); + expect(html).not.toContain('__codemanCustomModelClis'); }); it('does not expose gesture at all when CODEMAN_GESTURE is unset', async () => { diff --git a/test/routes/custom-model-routes.test.ts b/test/routes/custom-model-routes.test.ts index 7c864be5..059f210d 100644 --- a/test/routes/custom-model-routes.test.ts +++ b/test/routes/custom-model-routes.test.ts @@ -194,3 +194,104 @@ describe('custom model endpoint CRUD', () => { } }); }); + +describe('defaultModelId — the Run-menu picker’s per-endpoint default', () => { + afterEach(() => { + fetchMock.mockReset(); + }); + + it('rejects a defaultModelId that is not one of the endpoint’s discovered models, on both create and update', async () => { + const { app } = await setup(); + const create = await app.inject({ + method: 'POST', + url: '/api/model-endpoints', + payload: { + id: 'ep-default-reject', + label: 'A', + baseUrl: 'http://localhost:8080', + models: ['qwen3'], + defaultModelId: 'ghost', + }, + }); + expect(create.json().success).toBe(false); + expect(create.json().errorCode).toBe('INVALID_INPUT'); + + await app.inject({ + method: 'POST', + url: '/api/model-endpoints', + payload: { id: 'ep-default-reject', label: 'A', baseUrl: 'http://localhost:8080', models: ['qwen3'] }, + }); + const update = await app.inject({ + method: 'PUT', + url: '/api/model-endpoints/ep-default-reject', + payload: { label: 'A', baseUrl: 'http://localhost:8080', models: ['qwen3'], defaultModelId: 'ghost' }, + }); + expect(update.json().success).toBe(false); + expect(update.json().errorCode).toBe('INVALID_INPUT'); + }); + + it('accepts a defaultModelId that IS one of the discovered models', async () => { + const { app } = await setup(); + const res = await app.inject({ + method: 'POST', + url: '/api/model-endpoints', + payload: { + id: 'ep-default-accept', + label: 'A', + baseUrl: 'http://localhost:8080', + models: ['qwen3', 'llama3'], + defaultModelId: 'llama3', + }, + }); + expect(res.json().success).toBe(true); + expect(res.json().data.host.defaultModelId).toBe('llama3'); + }); + + it('drops a stale default that no longer appears in a fresh discovery, rather than carrying it forward invalid', async () => { + const { app } = await setup(); + await app.inject({ + method: 'POST', + url: '/api/model-endpoints', + payload: { + id: 'ep-default-drop', + label: 'A', + baseUrl: 'http://localhost:8080', + models: ['qwen3'], + defaultModelId: 'qwen3', + }, + }); + fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'llama3' }] }), { status: 200 })); + await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-default-drop/discover-models' }); + + const list = await app.inject({ method: 'GET', url: '/api/model-endpoints' }); + const stored = (list.json() as Array<{ id: string; defaultModelId?: string }>).find( + (h) => h.id === 'ep-default-drop' + ); + expect(stored?.defaultModelId).toBeUndefined(); + }); + + it('keeps a default that IS still present after a fresh discovery', async () => { + const { app } = await setup(); + await app.inject({ + method: 'POST', + url: '/api/model-endpoints', + payload: { + id: 'ep-default-keep', + label: 'A', + baseUrl: 'http://localhost:8080', + models: ['qwen3'], + defaultModelId: 'qwen3', + }, + }); + fetchMock.mockResolvedValue( + new Response(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }), { status: 200 }) + ); + await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-default-keep/discover-models' }); + + const list = await app.inject({ method: 'GET', url: '/api/model-endpoints' }); + const stored = (list.json() as Array<{ id: string; defaultModelId?: string }>).find( + (h) => h.id === 'ep-default-keep' + ); + expect(stored?.defaultModelId).toBe('qwen3'); + }); +});