diff --git a/CLAUDE.md b/CLAUDE.md index cc6ac628..991fefc0 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -227,7 +227,9 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. ⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. ⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. ⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs. -**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct but codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement — a confirmed protocol gap, not a bug; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. ⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`, JSON-escaped against a literal `` via the exported `escapeScriptJson()` since `label` is a user-`clis.json`-settable string unlike the neighbouring booleans-only `__codemanCliAvailable`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces. One entry per (capable, INSTALLED CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)", filtered through `isCliAvailable()` like the stock entries. Clicking one calls `selectCustomModelEntry(mode, endpointId)` (`session-ui.js`), which re-fetches the endpoint (never trusts anything cached from the dropdown's render — the 5-minute sweep below or a settings edit may have changed it since) and decides the model: exactly one discovered model launches straight away, two or more open `#customModelPickModal` to ask, with `defaultModelId` marked but never auto-chosen (asking exists so ONE launch can deliberately differ from the saved default). Either way the actual launch (`runCustomModelEntry`) routes through `run()` itself via a temporary `_runMode` swap — never `setRunMode()`, which would persist it as the user's new default — rather than a parallel dispatch table, which is what gives a custom-model launch the same `_runInFlight` lock every other Run click gets and means a CLI whose injection recipe lands later needs no update here. It then GETs `/api/sessions/:id/wait?until=idle&timeout=20000` on that session BEFORE applying — measured live, a freshly launched CLI reports itself `busy` for its own startup (boot spinner, workspace-trust check) well before the apply call would otherwise reach it, and the apply route's `isBusy()` guard correctly can't tell that apart from a real turn in progress, so every fresh launch failed with `SESSION_BUSY` until this wait was added. A timeout there is a normal 200 per the wait endpoint's own contract, never an error, so a session still busy after 20s just reaches the apply call anyway and gets that route's own honest error. It then calls `POST /api/sessions/:id/custom-model` on the session `run()` produced, guarded by snapshotting `activeSessionId` before the call and requiring it to have actually changed after — every `run*()` handles its own failure internally and returns normally rather than throwing, so a declined/failed launch must not silently re-point and restart whatever session was already open. ⚠️ The apply call reads the response body itself (`_api()`) rather than `_apiJson()`, which unwraps success but silently discards a failure body — losing the one thing (`error`) that distinguishes "still busy", "not a discovered model", "remote/Docker session" and everything else the route can report; the resulting toast is `type: 'error'`, which `showToast()` now defaults to STICKY (no auto-dismiss, an explicit close button) precisely so a message worth diagnosing survives long enough to be read — a 3s default hid the real reason behind every one of these failures until it was fixed. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to launch with). ⚠️ **Every saved endpoint's models also re-discover themselves automatically**, a `this.cleanup.setInterval` in `server.ts` (`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, 5 minutes, off under `testMode` like the Codex plan-usage poll beside it) calling the exported `refreshAllCustomModelHosts()` (`custom-model-routes.ts`) — one endpoint unreachable on a cycle never blocks the others, and a read-modify-write PER HOST (re-reading the store before each splice, keyed by id) means an admin's concurrent edit or delete wins over a sweep that started before it, never the reverse. +**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ **Two launch paths, chosen by mechanism, not preference — see the second paragraph below for why**: opencode/codex/gemini/pi/grok/deepseek/omp apply the selection ONE-SHOT, before the session/process ever exists, with no restart at all; claude alone still applies a selection by **restarting the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct, and codex only speaks the Responses API since Feb 2026 (`wire_api = "responses"`) — re-verified live against a llama-swap deployment that DOES answer `/v1/responses` (an earlier test's harder failure against a different deployment does not reproduce everywhere): a plain, no-tool-call chat turn gets a real reply, but a real tool-call attempt came back as `agent_message` TEXT (the tool-call JSON printed as the answer) rather than an executable `function_call` item — confirmed via `codex exec --json`'s raw event stream. Tool execution is what makes codex a coding agent, so it remains not usable for real work either way, just with a more precise failure mode than a flat protocol break; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. ⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`, JSON-escaped against a literal `` via the exported `escapeScriptJson()` since `label` is a user-`clis.json`-settable string unlike the neighbouring booleans-only `__codemanCliAvailable`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces. One entry per (capable, INSTALLED CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)", filtered through `isCliAvailable()` like the stock entries. Clicking one calls `selectCustomModelEntry(mode, endpointId)` (`session-ui.js`), which re-fetches the endpoint (never trusts anything cached from the dropdown's render — the 5-minute sweep below or a settings edit may have changed it since) and decides the model: exactly one discovered model launches straight away, two or more open `#customModelPickModal` to ask, with `defaultModelId` marked but never auto-chosen (asking exists so ONE launch can deliberately differ from the saved default). Either way the actual launch (`runCustomModelEntry`) routes through `run()` itself via a temporary `_runMode` swap — never `setRunMode()`, which would persist it as the user's new default — rather than a parallel dispatch table, which is what gives a custom-model launch the same `_runInFlight` lock every other Run click gets and means a CLI whose injection recipe lands later needs no update here. It then GETs `/api/sessions/:id/wait?until=idle&timeout=20000` on that session BEFORE applying — measured live, a freshly launched CLI reports itself `busy` for its own startup (boot spinner, workspace-trust check) well before the apply call would otherwise reach it, and the apply route's `isBusy()` guard correctly can't tell that apart from a real turn in progress, so every fresh launch failed with `SESSION_BUSY` until this wait was added. A timeout there is a normal 200 per the wait endpoint's own contract, never an error, so a session still busy after 20s just reaches the apply call anyway and gets that route's own honest error. It then calls `POST /api/sessions/:id/custom-model` on the session `run()` produced, guarded by snapshotting `activeSessionId` before the call and requiring it to have actually changed after — every `run*()` handles its own failure internally and returns normally rather than throwing, so a declined/failed launch must not silently re-point and restart whatever session was already open. ⚠️ The apply call reads the response body itself (`_api()`) rather than `_apiJson()`, which unwraps success but silently discards a failure body — losing the one thing (`error`) that distinguishes "still busy", "not a discovered model", "remote/Docker session" and everything else the route can report; the resulting toast is `type: 'error'`, which `showToast()` now defaults to STICKY (no auto-dismiss, an explicit close button) precisely so a message worth diagnosing survives long enough to be read — a 3s default hid the real reason behind every one of these failures until it was fixed. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to launch with). ⚠️ **Every saved endpoint's models also re-discover themselves automatically**, a `this.cleanup.setInterval` in `server.ts` (`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, 5 minutes, off under `testMode` like the Codex plan-usage poll beside it) calling the exported `refreshAllCustomModelHosts()` (`custom-model-routes.ts`) — one endpoint unreachable on a cycle never blocks the others, and a read-modify-write PER HOST (re-reading the store before each splice, keyed by id) means an admin's concurrent edit or delete wins over a sweep that started before it, never the reverse. + +**Everything below landed after the initial backend + picker cut, each confirmed live against a real llama-swap deployment.** ⚠️ **llama-swap runs one model at a time, and switching can disrupt ANOTHER live session** — before applying, both apply routes call llama-swap's own `GET /running` (feature-detected via `getLlamaSwapStatus()`, `custom-model-routes.ts`; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is simply never checked). If a different model is loaded and ready AND another live session's own selection is using it, the apply returns `{requiresConfirmation, currentlyLoadedModel, affectedSessions}` instead of switching silently; retrying with `confirmed: true` skips the check, and switching with nothing else affected proceeds immediately. llama-swap also has no dedicated "switch model" endpoint — the only thing that actually starts a swap is a real inference request naming the model (confirmed live: applying a selection alone never reached llama-swap's own logs, since nothing had asked it to load anything) — so both routes also fire `triggerLlamaSwapLoad()`, a fire-and-forget `POST /v1/chat/completions` with `max_tokens: 1`, whenever the target model isn't already loaded and ready. ⚠️ **That launch-time check cannot catch a swap caused by a DIFFERENT session's LATER, ordinary use** — confirmed live: a second Codex session picking a different model launched with no warning at all (nothing conflicted at that exact instant), yet it silently evicted the first session's model regardless, since llama-swap has no push notification of its own. `detectCustomModelSwapDisplacements()` (`custom-model-routes.ts`) is a separate periodic sweep (`server.ts`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS` = 20s) that compares each live custom-model session's own `modelId` against what `/running` actually reports loaded, broadcasting a `custom-model:swapped-out` SSE event — shown as a global toast, never tied to the displaced session's own tab, since the whole point is telling the user before they type into it — the first time a mismatch appears, via a caller-owned de-dupe `Set` cleared once that session's own model is loaded and ready again so a later, genuinely new displacement notifies again rather than staying silently un-notified forever after the first one. ⚠️ **Context length is read from the REAL launch command, never `/props`** — `/props?model=`'s `default_generation_settings.n_ctx` was confirmed live to report a `--fit-ctx`-launched backend's theoretical/trained maximum rather than the real runtime-configured size (a measured 154112-vs-16384 discrepancy, caught only because the unfixed value still overflowed), so discovery parses the actual configured size straight out of `/running`'s own `cmd` field instead (`parseCtxFromCmd`: `--fit-ctx ` first, then plain llama.cpp `-c`/`--ctx-size`), falling back to `/props` only when `cmd` states no recognizable flag at all. ⚠️ **Claude alone gets a context-window FLOOR check, on top of the ceiling `contextLengthVar` already fixes** — `exceedsSafeContextFloor()` (gated on the registry declaring `contextLengthVar`, so a no-op for every other CLI by construction) compares a model's discovered context against `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000): confirmed live, twice, that Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on the very first message, before any conversation history exists to compact, so a smaller real context fails outright regardless of what `CLAUDE_CODE_MAX_CONTEXT_TOKENS` says (that var only controls when HISTORY gets compacted, and there is none yet on message one). Below the floor, the apply returns `{requiresContextWarning, modelId, contextLength, minSafeContextTokens}` instead of launching, shown as an in-app dialog naming the actual fix: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on `--fit-ctx` auto-fit, which optimizes for the biggest MODEL that fits rather than the biggest CONTEXT. ⚠️ **A fresh, isolated `CLAUDE_CONFIG_DIR` looks like a brand-new Claude Code profile and replays its ENTIRE first-run sequence on every launch** — the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with a bypass-permissions flag) a one-time warning about it, confirmed live, none of which a real, already-onboarded profile shows again. `skipFirstRunPrompts` (claude's entry only, requires `apiKeyTrustFile` since it reuses the same file) pre-seeds that same "already been through this" state: `hasCompletedOnboarding` and this session's own `projects[workingDir].hasTrustDialogAccepted` merge into the same `.claude.json` the API-key trust file already writes to, and `skipDangerousModePermissionPrompt` merges into `settings.json` (a different file, same corrupt-tolerant merge). ⚠️ **The loading banner shows the REAL backend log line, not a guess, and has no countdown or auto-timeout at all.** `getLatestLlamaSwapLogLine()` holds one `GET /api/events` SSE connection open per endpoint (confirmed live to stay open indefinitely — read past 220KB over 8s with no `done`; idle-closed after 30s via `pruneIdleLlamaSwapLogTails`, same 20s sweep as the swap-displacement check above), parsing `logData` frames and keeping only `source: "upstream"` (the real `llama-server` process's own stdout) lines, never `source: "proxy"` (llama-swap's own request-access log). ⚠️ `GET /logs` — the endpoint this feature's own first cut targeted, since the name suggested it — was confirmed live to carry ONLY the proxy log and never a single backend line, even seconds after a real, verified model swap; caught and corrected by a live check before merge, not after. The banner itself dropped its size-scaled expected-time estimate and matching auto-timeout (a guess dressed up as a fact that could kill a genuinely slow load partway through on slower hardware) for a generic hardware/model-size disclaimer plus a user-driven **Cancel** button (`_showCenterStatus`'s `onCancel` option, a real button distinct from the plain "×" close glyph an `'error'`-type banner gets) that ends the wait and closes the session on the user's own call rather than a guessed deadline. **Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w-` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization) diff --git a/docs/api-reference.md b/docs/api-reference.md index 9e274cb6..00d60d8f 100644 --- a/docs/api-reference.md +++ b/docs/api-reference.md @@ -66,17 +66,17 @@ The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in `src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on the HTTP status. -| `errorCode` | HTTP | Meaning | -|-------------|------|---------| -| `INVALID_INPUT` | 400 | Malformed request / failed validation | -| `UNAUTHORIZED` | 401 | Authentication required or failed | -| `NOT_FOUND` | 404 | Resource does not exist | -| `SESSION_BUSY` | 409 | Session is busy | -| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) | -| `ALREADY_EXISTS` | 409 | Resource already exists | -| `OPERATION_FAILED` | 422 | Well-formed but could not be completed | -| `RATE_LIMITED` | 429 | Too many requests | -| `INTERNAL_ERROR` | 500 | Unexpected server error | +| `errorCode` | HTTP | Meaning | +| ------------------ | ---- | --------------------------------------------------- | +| `INVALID_INPUT` | 400 | Malformed request / failed validation | +| `UNAUTHORIZED` | 401 | Authentication required or failed | +| `NOT_FOUND` | 404 | Resource does not exist | +| `SESSION_BUSY` | 409 | Session is busy | +| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) | +| `ALREADY_EXISTS` | 409 | Resource already exists | +| `OPERATION_FAILED` | 422 | Well-formed but could not be completed | +| `RATE_LIMITED` | 429 | Too many requests | +| `INTERNAL_ERROR` | 500 | Unexpected server error | Adding a new error code is non-breaking; removing or renaming one is a major change. @@ -87,10 +87,10 @@ exist because SSE is Codeman's only other "tell me when" channel, and an agent driving the API from a shell tool cannot practically hold a stream and parse events inline. -| Call | Blocks until | -|------|--------------| -| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires | -| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output | +| Call | Blocks until | +| --------------------------------------------- | -------------------------------------------------- | +| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires | +| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output | | `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires | `POST .../input` with `wait` is not the same as a `POST` followed by a separate @@ -140,13 +140,13 @@ contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send ### Signals -| Signal | Source | Actually fires for | -|--------|--------|--------------------| -| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) | -| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) | -| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only | -| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below | -| `exit` | no process is behind the session | every mode | +| Signal | Source | Actually fires for | +| --------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) | +| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) | +| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only | +| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below | +| `exit` | no process is behind the session | every mode | `stop` is the signal to orchestrate on where it exists; `idle` is a heuristic fallback that can flap mid-turn when a spinner pauses. The default set when `until` @@ -156,12 +156,12 @@ can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit once the session is up: the default set's `idle` also resolves on a spinner pause, and on a fresh session the **startup** `idle` (emitted when the CLI first comes up) can land inside your first wait window and report a turn that never ran. Measured: -a session parked on the trust dialog emits no *further* `idle`, so it is the +a session parked on the trust dialog emits no _further_ `idle`, so it is the startup transition, not the dialog, that produces the false success below. ⚠️ **`exit` means "nothing is running", which includes "not started yet".** The server answers from `pid === null` plus a mux-layer pane-death probe, and that -covers a session that exited — including a worker that died *inside* its tmux pane +covers a session that exited — including a worker that died _inside_ its tmux pane while the local attach client (and therefore `pid`) lives on — one that was detached, and one that was **created but never started**. So the first wait after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in @@ -184,7 +184,7 @@ blocked, and polling `blocked` alone will sit at its timeout. ⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell session emits its one `idle` at startup and then stays `status: "idle"` forever, -whatever the pane is doing, so it never emits a *transition*. Since send-and-wait +whatever the pane is doing, so it never emits a _transition_. Since send-and-wait requires a transition (and so does `fresh=1`), both can only time out there: a documented default `wait` on a shell worker running `sleep 4` times out at the full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker @@ -218,11 +218,11 @@ with `from=buffer` keeps matching long after the dialog is gone. A worked versio ### `GET /api/v1/sessions/:id/wait` -| Param | Type | Default | Notes | -|-------|------|---------|-------| -| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback | -| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` | -| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time | +| Param | Type | Default | Notes | +| --------- | -------------------------------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback | +| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` | +| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time | ```bash curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" @@ -239,12 +239,12 @@ a plain signal wait, so check the endpoint path before blaming the parameters. ### `GET /api/v1/sessions/:id/wait-output` -| Param | Type | Default | Notes | -|-------|------|---------|-------| -| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found | -| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing | -| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking | -| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` | +| Param | Type | Default | Notes | +| --------- | ------------------------------- | -------- | ----------------------------------------------------------------------------------------------------------- | +| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found | +| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing | +| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking | +| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` | **Matching is literal, never a pattern.** A `regex` parameter is rejected with a `400` rather than ignored, so a caller that assumed otherwise finds out immediately @@ -296,10 +296,10 @@ hand-written query string decodes to a space. Two optional fields on the existing endpoint: -| Field | Type | Notes | -|-------|------|-------| -| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" | -| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` | +| Field | Type | Notes | +| ------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" | +| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` | Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as "absent" rather than failing validation. That is deliberate: `.optional()` would @@ -330,16 +330,24 @@ All three nest the wait result under `data.wait`, so one client helper works aga any of them: ```json -{ "success": true, "data": { - "sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464", - "status": "idle", - "limitPaused": false, - "wait": { - "signal": "stop", "until": ["stop", "idle", "exit"], - "timedOut": false, "immediate": false, "ended": false, "aborted": false, - "waitedMs": 8421, "timeoutMs": 60000 +{ + "success": true, + "data": { + "sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464", + "status": "idle", + "limitPaused": false, + "wait": { + "signal": "stop", + "until": ["stop", "idle", "exit"], + "timedOut": false, + "immediate": false, + "ended": false, + "aborted": false, + "waitedMs": 8421, + "timeoutMs": 60000 + } } -}} +} ``` `POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`, @@ -353,21 +361,21 @@ redelivery (harmless, the turn it refers to may be long over), while with client that reads `delivered === false` as "duplicate" silently treats a failed send as a success. -| Field | Type | Meaning | -|-------|------|---------| -| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) | -| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) | -| `wait.matched` | boolean | the string appeared (`/wait-output` only) | -| `wait.match` | string | the literal that was searched for (`/wait-output` only) | -| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) | -| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` | -| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) | -| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met | -| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome | -| `wait.waitedMs` | number | wall-clock ms actually spent waiting | -| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied | -| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand | -| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard | +| Field | Type | Meaning | +| ---------------- | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) | +| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) | +| `wait.matched` | boolean | the string appeared (`/wait-output` only) | +| `wait.match` | string | the literal that was searched for (`/wait-output` only) | +| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) | +| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` | +| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) | +| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met | +| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome | +| `wait.waitedMs` | number | wall-clock ms actually spent waiting | +| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied | +| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand | +| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard | Read the outcome by discriminator, in this order: @@ -390,12 +398,12 @@ read the timeout as "the worker is wedged" and kill a session that was working f ### Errors -| `errorCode` | HTTP | When | -|-------------|------|------| -| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` | -| `NOT_FOUND` | 404 | no such session, or one this caller does not own | -| `SESSION_BUSY` | 409 | this session's waiter cap is full | -| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem | +| `errorCode` | HTTP | When | +| --------------- | ---- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` | +| `NOT_FOUND` | 404 | no such session, or one this caller does not own | +| `SESSION_BUSY` | 409 | this session's waiter cap is full | +| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem | The two capacity codes are deliberately different. A process-wide cap reported as `SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The @@ -446,9 +454,9 @@ Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md). - `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first, ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId, - sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?, - toolSummary?, message?, cwd?, context?, options?: {n, label}[], - acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame; +sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?, +toolSummary?, message?, cwd?, context?, options?: {n, label}[], +acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame; `options` is present only when the dialog's numbered choices parsed confidently; `acknowledgedAt` marks an item a human has already looked at (see `/viewed` below) and tells clients not to re-arm its tab alert. Listing @@ -466,7 +474,7 @@ Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md). first, `422 OPERATION_FAILED` when the session refused input. - `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes. - `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId, - acknowledged: itemId | null }`. Marks the session's pending **idle** item as +acknowledged: itemId | null }`. Marks the session's pending **idle** item as seen by a human (the web UI calls it when you open the session's tab): the item stays pending and answerable, but stops arming the yellow tab alert on every client, including after a reload. Permission/question items are never @@ -491,7 +499,7 @@ user guide: [`readmymind.md`](readmymind.md). - `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals, - recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap +recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap 50, each <= 500 chars). A case with nothing recorded answers an empty profile with `updatedAt: 0`; nothing is persisted by reads. - `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict @@ -528,12 +536,12 @@ user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md). - `GET /api/v1/model-endpoints` -> `CustomModelHost[]`, an unwrapped bare array like every other list route (still riding the standard `{success, - data}` envelope on the wire — unwrap it the same way). Answers `[]` for a +data}` envelope on the wire — unwrap it the same way). Answers `[]` for a non-admin in multi-user mode. `apiKey` is never returned; `apiKeySet: - boolean` reports whether one is stored, so a client can render "unchanged +boolean` reports whether one is stored, so a client can render "unchanged if left blank" without ever holding the real value. - `POST /api/v1/model-endpoints` with `{ id, label, baseUrl, apiKey?, - authStyle?, defaultModelId? }` creates one. `id` must match +authStyle?, defaultModelId? }` creates one. `id` must match `^[a-zA-Z0-9_-]+$`; `authStyle` is `bearer` (default) or `api-key`, never both (a real server hung indefinitely when sent both headers on one request); `baseUrl` must be `http(s)`, carry no embedded credentials, and @@ -548,24 +556,72 @@ user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md). - `DELETE /api/v1/model-endpoints/:id` removes one. - `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's own `GET /v1/models` and stores the result as `models`, updating - `lastDiscoveredAt`. A `defaultModelId` that no longer appears in the fresh - list is dropped rather than carried forward invalid. Failures answer - `502 OPERATION_FAILED` with the underlying connection error, or a named - egress refusal if the resolved address turned out to be blocked. The same - refresh also runs automatically for every saved endpoint every 5 minutes - in the background (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, - started from `server.ts`), so there is no route for triggering "refresh - all" — one endpoint being unreachable on a cycle never blocks the others. -- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId } | - { clear: true }` applies (or clears) the session's selection and - **restarts the session's CLI process in place** — every supported harness - reads its endpoint config at process start, never per turn, so there is - no live hot-swap. A Claude session resumes its existing conversation - across the restart; pi/omp/grok additionally get a forced `--model`/`-m` - value, since for those three the config file alone does not select it. - `400 INVALID_INPUT` for a remote (SSH) or Docker session — both restart - their agent differently under the hood, and applying to one would report - success while changing nothing. + `lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own + response already reports loaded) `modelContextLengths` and `modelSizesGB`. + A `defaultModelId` that no longer appears in the fresh list is dropped + rather than carried forward invalid. Failures answer `502 OPERATION_FAILED` + with the underlying connection error, or a named egress refusal if the + resolved address turned out to be blocked. The same refresh also runs + automatically for every saved endpoint every 5 minutes in the background + (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, started from + `server.ts`), so there is no route for triggering "refresh all" — one + endpoint being unreachable on a cycle never blocks the others. +- `GET /api/v1/model-endpoints/:id/running-status` -> `{ isLlamaSwap, +running: [{model, state, cmd?}], logLine? }`, read-only, no admin gate + (any session owner who could already point a session at this endpoint can + equally ask what it currently has loaded). `isLlamaSwap` is + feature-detected via the endpoint's own `GET /running` — a plain + llama.cpp/OpenAI-compatible server has none and always answers `false`. + `logLine`, present only when `isLlamaSwap` is true, is the most recent + REAL backend `llama-server` process log line (`load_model: ...`, + `llama_server: model loaded`, etc.), sourced from the endpoint's own + `GET /api/events` SSE stream and filtered to `source: "upstream"` frames + only (never llama-swap's own `source: "proxy"` request-access log) — one + connection is held open per endpoint and reused across every poller, + idle-closed after 30s of nobody asking. This is what the Run-menu + picker's loading banner polls once a second while a model is loading. +- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId, +confirmed? } | { clear: true }` applies (or clears) the session's + selection and **restarts the session's CLI process in place** — every + supported harness reads its endpoint config at process start, never per + turn, so there is no live hot-swap. (`POST /api/v1/quick-start`'s own + `customModel: { endpointId, modelId, confirmed? }` field is the + no-restart equivalent for a session that doesn't exist yet — see below.) + A Claude session resumes its existing conversation across the restart; + pi/omp/grok additionally get a forced `--model`/`-m` value, since for + those three the config file alone does not select it. `400 INVALID_INPUT` + for a remote (SSH) or Docker session — both restart their agent + differently under the hood, and applying to one would report success + while changing nothing. Two more responses replace the normal + `{customModel, restarted}` shape, neither an error — both require + retrying the same call with `confirmed: true` to proceed anyway, and + neither restarts or creates anything on the first ask: + - `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` — + llama.cpp/llama-swap only runs one model at a time, and switching would + unload a model another **live session's own selection** is actively + using. Never returned for a plain (non-llama-swap) server, and never + just because a swap is needed at all — only when it would disrupt + someone else. + - `{requiresContextWarning: true, modelId, contextLength, +minSafeContextTokens}` — Claude Code's own fixed per-turn overhead + (system prompt + tool schemas) can exceed a small model's entire + discovered context on its own, before any conversation history exists + to compact, guaranteeing the very first message fails regardless of + `CLAUDE_CODE_MAX_CONTEXT_TOKENS`. Gated on the CLI registry declaring a + `contextLengthVar` (claude only today), so it never fires for another + harness. +- `POST /api/v1/quick-start`'s `customModel: { endpointId, modelId, +confirmed? }` field (alongside its normal `caseName`/`mode`/etc. body) + computes the same injection **before** the session exists and launches + directly on the endpoint — no restart, because there was never a + native-backend boot to restart away from. Runs the identical checks as + the dedicated route above (`requiresConfirmation`/`requiresContextWarning`, + same shapes, same `confirmed: true` retry), and is refused the same way + for a remote or Docker case. This is what the Run-menu picker uses for + opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the + dedicated restart route above (its `--resume`-based restart is far less + jarring than a full relaunch, and folding it into the one-shot path is + separate work — see `docs/custom-model-endpoints-plan.md`). ## Voice dictation @@ -575,7 +631,7 @@ same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synce [`claude-voice-plan.md`](claude-voice-plan.md). - `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?, - expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody +expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody signed in to Claude Code on the server), `expired` (the access token elapsed; running any Claude session refreshes it) or `malformed`. The OAuth token itself is never returned by this or any other endpoint. diff --git a/docs/custom-model-endpoints-plan.md b/docs/custom-model-endpoints-plan.md index d1d0f3c6..bb4ef289 100644 --- a/docs/custom-model-endpoints-plan.md +++ b/docs/custom-model-endpoints-plan.md @@ -352,8 +352,12 @@ pure unit tests and the live manual checks in Verification: up automatically with zero edits to the script). Already run to completion against the author's llama-swap server (a LAN address, inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries): - claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected** - (Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED** + claude/opencode/pi/grok/omp **PASS**, codex **partially works and still + isn't usable** (plain chat succeeds against a llama-swap deployment that + answers `/v1/responses`, but a real tool-call attempt comes back as + inert text rather than an executable `function_call` — see the + confidence table row for the full, re-verified picture), gemini/deepseek + **UNCONFIRMED** (reach the server, fail for undiagnosed reasons — see their table rows), antigravity **SKIP** (no mechanism). Re-run this against a real cloud endpoint (e.g. an Azure AI Foundry deployment) once one is available, to diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index 1740f3aa..cd42456e 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -12,11 +12,12 @@ recipe confidence table, and security reasoning: [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md). > **Status**: fully wired end to end — registry capability, the injection -> engine, the endpoint store + discovery route, the session restart route, -> a settings-panel CRUD surface, and the Run-menu picker described below. -> Antigravity has no known custom-endpoint mechanism and is not supported. -> The HTTP API (examples below) still works directly and is what the picker -> itself calls under the hood. +> engine, the endpoint store + discovery route, both the restart-in-place +> apply route (Claude) and the one-shot quick-start launch path (every +> other supported harness), a settings-panel CRUD surface, and the Run-menu +> picker described below. Antigravity has no known custom-endpoint +> mechanism and is not supported. The HTTP API (examples below) still works +> directly and is what the picker itself calls under the hood. ## Turning it on @@ -351,8 +352,9 @@ which can take anywhere from a few seconds to well over a minute: naming the model, and confirmed live: applying a selection alone never reached llama-swap at all (nothing in its own server logs), since nothing had actually asked it to load anything yet. Both apply routes now also - send the smallest real request that will — `POST /v1/chat/ -completions` with `max_tokens: 1` and one throwaway message — whenever the + send the smallest real request that will — + `POST /v1/chat/completions` with `max_tokens: 1` and one + throwaway message — whenever the target model isn't already the one loaded and ready, fire-and-forget (its response is never read; `GET /api/model-endpoints/:id/running-status`, polled client-side, is what actually confirms readiness). The response