diff --git a/CLAUDE.md b/CLAUDE.md index 2583c903..37653f21 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -227,7 +227,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. ⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. ⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. ⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs. -**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct but codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement — a confirmed protocol gap, not a bug; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. ⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces, extended to the one frontend surface that needs to know which CLIs support this. One entry per (capable CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)"; picking one runs that CLI's own existing `run*()` function unmodified (case creation, env overrides, the works — forced to a single instance) and then calls `POST /api/sessions/:id/custom-model` on the session it selects, reusing the fact every `run*()` ends by selecting its new session rather than a parallel create path. `CustomModelHost.defaultModelId` is what the picker applies with no further choice — settings-ui.js's Edit form is a select populated from that endpoint's own discovered `models`, the route rejects a value that isn't a member, and re-discovery drops a stale one rather than carrying it forward invalid. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to default to). +**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct but codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement — a confirmed protocol gap, not a bug; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. ⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`, JSON-escaped against a literal `` via the exported `escapeScriptJson()` since `label` is a user-`clis.json`-settable string unlike the neighbouring booleans-only `__codemanCliAvailable`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces. One entry per (capable, INSTALLED CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)", filtered through `isCliAvailable()` like the stock entries. Clicking one calls `selectCustomModelEntry(mode, endpointId)` (`session-ui.js`), which re-fetches the endpoint (never trusts anything cached from the dropdown's render — the 5-minute sweep below or a settings edit may have changed it since) and decides the model: exactly one discovered model launches straight away, two or more open `#customModelPickModal` to ask, with `defaultModelId` marked but never auto-chosen (asking exists so ONE launch can deliberately differ from the saved default). Either way the actual launch (`runCustomModelEntry`) routes through `run()` itself via a temporary `_runMode` swap — never `setRunMode()`, which would persist it as the user's new default — rather than a parallel dispatch table, which is what gives a custom-model launch the same `_runInFlight` lock every other Run click gets and means a CLI whose injection recipe lands later needs no update here. It then calls `POST /api/sessions/:id/custom-model` on the session `run()` produced, guarded by snapshotting `activeSessionId` before the call and requiring it to have actually changed after — every `run*()` handles its own failure internally and returns normally rather than throwing, so a declined/failed launch must not silently re-point and restart whatever session was already open. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to launch with). ⚠️ **Every saved endpoint's models also re-discover themselves automatically**, a `this.cleanup.setInterval` in `server.ts` (`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, 5 minutes, off under `testMode` like the Codex plan-usage poll beside it) calling the exported `refreshAllCustomModelHosts()` (`custom-model-routes.ts`) — one endpoint unreachable on a cycle never blocks the others, and a read-modify-write PER HOST (re-reading the store before each splice, keyed by id) means an admin's concurrent edit or delete wins over a sweep that started before it, never the reverse. **Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w-` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization) diff --git a/docs/api-reference.md b/docs/api-reference.md index 3a5dab66..9e274cb6 100644 --- a/docs/api-reference.md +++ b/docs/api-reference.md @@ -551,7 +551,11 @@ user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md). `lastDiscoveredAt`. A `defaultModelId` that no longer appears in the fresh list is dropped rather than carried forward invalid. Failures answer `502 OPERATION_FAILED` with the underlying connection error, or a named - egress refusal if the resolved address turned out to be blocked. + egress refusal if the resolved address turned out to be blocked. The same + refresh also runs automatically for every saved endpoint every 5 minutes + in the background (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, + started from `server.ts`), so there is no route for triggering "refresh + all" — one endpoint being unreachable on a cycle never blocks the others. - `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId } | { clear: true }` applies (or clears) the session's selection and **restarts the session's CLI process in place** — every supported harness diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index 24acc171..140f10dd 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -66,18 +66,26 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one. Endpoint management is admin-only in multi-user mode, same as remote/docker hosts — these are machine-level infra, not per-user settings. -`defaultModelId` names which discovered model the Run-menu picker applies -for that endpoint with no further choice — the settings panel's Edit form -exposes it as a select populated from the endpoint's own discovered -`models`, and the route refuses a value that isn't one of them. Leaving it -unset falls back to the first discovered model; re-discovering drops a -default that no longer appears in the fresh list rather than carrying an -invalid one forward. +`defaultModelId` names which discovered model the picker pre-marks for that +endpoint — the settings panel's Edit form exposes it as a select populated +from the endpoint's own discovered `models`, and the route refuses a value +that isn't one of them. It is applied automatically only when the endpoint +has exactly one discovered model (nothing to choose); with two or more it +is a pre-selection in the model-picker dialog below, never a silent default. +Re-discovering drops a default that no longer appears in the fresh list +rather than carrying an invalid one forward. + +**Model lists refresh themselves.** A background sweep (`server.ts`, +`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every +saved endpoint the same way the manual `POST .../discover-models` route +does, best-effort per endpoint — one being unreachable on a given cycle +never blocks the others. Off under `npm test`, same reasoning as the Codex +plan-usage poll it sits beside: no real network to hit, no server instance +to keep the timer alive for. ## The Run-menu picker -With the setting on and at least one endpoint carrying a usable default -model (either an explicit `defaultModelId` or just one discovered model), +With the setting on and at least one endpoint carrying a discovered model, the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry per (harness that can redirect to a custom endpoint, saved endpoint) pair, e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI @@ -86,13 +94,22 @@ registry's own `capabilities.customModelInjection` at page render in the frontend — so a CLI whose injection recipe lands later shows up with no frontend change, and Antigravity (`unsupported`) never does. -Picking an entry runs a single session on that harness exactly the way its +Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`, +`session-ui.js`) rather than trusting anything cached from the dropdown's +own render — the model list can have changed via the 5-minute sweep above +or a settings-panel edit since the menu opened. With exactly one discovered +model it runs straight away; with two or more, a small modal +(`#customModelPickModal`) lists them and asks which one to use for this +launch, with the endpoint's `defaultModelId` marked but not auto-chosen — +the point of asking is letting one launch deliberately differ from the +saved default, not just confirming it. Whichever way the model was decided, +the launch itself runs a single session on that harness exactly the way its own Run-menu entry would (same case creation, env overrides, everything), -then immediately applies the endpoint's default model to it via the route -below. It is a one-off "try this endpoint" action, not a sticky mode: the -plain Run button still means "this harness, native cloud" afterward. -Entries are hidden entirely for a remote or Docker active case, since the -apply route refuses both (see the next section). +then immediately applies the endpoint and model to it via the route below. +It is a one-off "try this endpoint" action, not a sticky mode: the plain +Run button still means "this harness, native cloud" afterward. Entries are +hidden entirely for a remote or Docker active case, since the apply route +refuses both (see the next section). ## Applying a model to a session diff --git a/docs/wiki/Custom-Model-Endpoints.md b/docs/wiki/Custom-Model-Endpoints.md index a989a858..8fccaab3 100644 --- a/docs/wiki/Custom-Model-Endpoints.md +++ b/docs/wiki/Custom-Model-Endpoints.md @@ -16,20 +16,32 @@ Still in App Settings → Models → Custom model endpoints: say). An API key is optional; most local servers don't check one. 2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it. 3. Pick a **default model** from what was discovered. This is the model the Run-menu entry - below applies with no further choice, so set it once you know which one you want. + applies directly when only one model is discovered; with two or more, it's just the one + pre-marked in the picker dialog described below, not a silent default. Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker hosts — these are machine-level infra, not a per-user setting. +**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every +5 minutes in the background, so a model the server starts serving later — or stops serving — +shows up without another manual click of **Discover**. One endpoint being unreachable on a +given cycle (powered off, wrong network) never blocks the others from refreshing. + ## Running a session against one -With the setting on and at least one endpoint carrying a usable default model, the **Run** +With the setting on and at least one endpoint carrying a discovered model, the **Run** dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a -session on that harness exactly the way its own entry would, then points it at the -endpoint's default model. It is a one-off "try this endpoint" action, not a sticky mode — the -plain **Run** button still means "this harness, native cloud" afterward, and a fresh session -never inherits whatever the last one was pointed at. +session on that harness exactly the way its own entry would. It is a one-off "try this +endpoint" action, not a sticky mode — the plain **Run** button still means "this harness, +native cloud" afterward, and a fresh session never inherits whatever the last one was +pointed at. + +**Which model it uses depends on how many the endpoint has discovered.** With exactly one, +the session launches straight away on that model — nothing to choose. With two or more, a +small dialog asks which one to use for this launch before starting the session; the +endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately +use a different one without changing the saved default. Applying a selection **restarts the harness's process in place** — same tab, same conversation where the harness supports resuming one, fresh environment. That restart is diff --git a/src/web/public/i18n.js b/src/web/public/i18n.js index d72cc33a..3a0fbc5b 100644 --- a/src/web/public/i18n.js +++ b/src/web/public/i18n.js @@ -311,6 +311,9 @@ 'What the Run-menu picker applies for this endpoint. Discover models first.': '运行菜单选择器会为此端点应用该模型。请先发现可用模型。', 'Custom Endpoints': '自定义端点', + 'Choose a model': '选择模型', + 'That endpoint no longer exists': '该端点已不存在', + 'No models discovered for this endpoint yet': '此端点尚未发现任何模型', 'Subagent Options': '子智能体选项', 'Enable Tracking': '启用跟踪', 'Active Tab Only': '仅活动标签页', diff --git a/src/web/public/index.html b/src/web/public/index.html index 65007ffc..244ae124 100644 --- a/src/web/public/index.html +++ b/src/web/public/index.html @@ -910,6 +910,23 @@ + + + + `, { url: 'http://localhost/', runScripts: 'dangerously' } ); @@ -68,6 +73,10 @@ function bootApp( app.loadAppSettingsFromStorage = () => ({ customModelEndpointsEnabled: options.settingsEnabled ?? true }); app.isCliAvailable = options.cliAvailable ?? (() => true); app.showToast = () => {}; + // Default no-op so a button's onclick (selectCustomModelEntry -> possibly + // straight to runCustomModelEntry for a single-model host) never rejects + // with "this.run is not a function"; tests of the launch itself override it. + app.run = async () => {}; // _apiJson unwraps the {success,data} envelope for real against a live // server; here it stands in for that, driven from a fixed `hosts` fixture // so these tests exercise the picker's OWN code, not the envelope helper. @@ -172,6 +181,122 @@ describe('Custom Model Endpoint Profiles: Run-menu picker generation', () => { }); }); +describe('Custom Model Endpoint Profiles: the "which model" picker', () => { + it('launches straight away for a host with exactly one discovered model, no dialog', async () => { + const { win, app } = bootApp({ + hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://localhost:8080', models: ['qwen3'] }], + }); + let launched: unknown[] | null = null; + app.runCustomModelEntry = async (...args: unknown[]) => { + launched = args; + }; + + await app.selectCustomModelEntry('claude', 'llama-box'); + + expect(launched).toEqual(['claude', 'llama-box', 'qwen3']); + expect(win.document.getElementById('customModelPickModal')!.classList.contains('active')).toBe(false); + }); + + it('opens the picker for a host with more than one discovered model, rather than launching directly', async () => { + const { win, app } = bootApp({ + hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://localhost:8080', models: ['qwen3', 'llama3'] }], + }); + let launched = false; + app.runCustomModelEntry = async () => { + launched = true; + }; + + await app.selectCustomModelEntry('claude', 'llama-box'); + + expect(launched).toBe(false); + const modal = win.document.getElementById('customModelPickModal')!; + expect(modal.classList.contains('active')).toBe(true); + const list = win.document.getElementById('customModelPickList')!; + expect(list.querySelectorAll('button').length).toBe(2); + expect(list.textContent).toContain('qwen3'); + expect(list.textContent).toContain('llama3'); + }); + + it('always asks with 2+ models, even when a defaultModelId is set — the point is letting this launch differ', async () => { + const { win, app } = bootApp({ + hosts: [ + { + id: 'llama-box', + label: 'llama.cpp', + baseUrl: 'http://localhost:8080', + models: ['qwen3', 'llama3'], + defaultModelId: 'qwen3', + }, + ], + }); + await app.selectCustomModelEntry('claude', 'llama-box'); + const modal = win.document.getElementById('customModelPickModal')!; + expect(modal.classList.contains('active')).toBe(true); + // The default is marked, not auto-chosen. + expect(win.document.getElementById('customModelPickList')!.textContent).toContain('Default'); + }); + + it('picking a row in the modal closes it and launches with that exact model', async () => { + const { win, app } = bootApp({ + hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://localhost:8080', models: ['qwen3', 'llama3'] }], + }); + let launched: unknown[] | null = null; + app.runCustomModelEntry = async (...args: unknown[]) => { + launched = args; + }; + win.app = app; + + await app.selectCustomModelEntry('claude', 'llama-box'); + const buttons = win.document.getElementById('customModelPickList')!.querySelectorAll('button'); + const llama3Btn = [...buttons].find((b) => b.textContent?.includes('llama3')) as unknown as HTMLButtonElement & { + onclick: (e: unknown) => void; + }; + expect(typeof llama3Btn.onclick).toBe('function'); + llama3Btn.onclick(new (win as any).Event('click')); + + expect(launched).toEqual(['claude', 'llama-box', 'llama3']); + expect(win.document.getElementById('customModelPickModal')!.classList.contains('active')).toBe(false); + }); + + it('re-fetches the endpoint at click time rather than trusting anything cached from the menu render', async () => { + // The background re-discovery sweep (server-side, every 5 minutes) or a + // settings-panel edit can change the model list between opening the + // dropdown and clicking a row — the picker must reflect what is current. + let fetchCount = 0; + const { win, app } = bootApp({}); + app._apiJson = async (path: string) => { + if (path !== '/api/model-endpoints') return null; + fetchCount += 1; + return [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://x', models: ['qwen3', 'llama3', 'phi4'] }]; + }; + await app.selectCustomModelEntry('claude', 'llama-box'); + expect(fetchCount).toBe(1); + expect(win.document.getElementById('customModelPickList')!.querySelectorAll('button').length).toBe(3); + }); + + it('toasts and does nothing when the endpoint has vanished by click time', async () => { + const { app } = bootApp({ hosts: [] }); + let toastMessage: string | null = null; + app.showToast = (msg: string) => { + toastMessage = msg; + }; + await app.selectCustomModelEntry('claude', 'ghost-endpoint'); + expect(toastMessage).toMatch(/no longer exists/i); + }); + + it('toasts and does nothing when the endpoint has zero discovered models by click time', async () => { + const { app } = bootApp({ + hosts: [{ id: 'llama-box', label: 'llama.cpp', baseUrl: 'http://x', models: [] }], + }); + let toastMessage: string | null = null; + app.showToast = (msg: string) => { + toastMessage = msg; + }; + await app.selectCustomModelEntry('claude', 'llama-box'); + expect(toastMessage).toMatch(/no models discovered/i); + }); +}); + describe('Custom Model Endpoint Profiles: applying a picked entry', () => { it('does not apply the endpoint to a session that was already open when the launch fails', async () => { const { app } = bootApp({});