diff --git a/.changeset/fix-replay-output-that-arrived-after-the-capture.md b/.changeset/fix-replay-output-that-arrived-after-the-capture.md deleted file mode 100644 index 023623ef..00000000 --- a/.changeset/fix-replay-output-that-arrived-after-the-capture.md +++ /dev/null @@ -1,5 +0,0 @@ ---- -"aicodeman": patch ---- - -fix(terminal): keep the output a pane capture could not contain. Opening a session, a backpressure refresh, a clear-terminal reload and a full-history re-pull all load the screen from a tmux pane capture, and anything the CLI printed between that capture and the end of the load used to be dropped, so its next partial redraw landed on a frame the terminal had never seen: missing or garbled output right after a tab switch or a refresh, plainest in a shell session. Each load now replays exactly the output that arrived after the capture, through one shared rule for all four paths, and a refresh that restores your scroll position no longer snaps back to the bottom afterwards. diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index 7402d884..ffaacfe8 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -10,7 +10,7 @@ "name": "codeman", "source": "./plugins/codeman", "description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.", - "version": "1.30.0", + "version": "1.31.0", "author": { "name": "Ark0N", "url": "https://github.com/Ark0N" diff --git a/CHANGELOG.md b/CHANGELOG.md index 10fdcd21..79b8a1c0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,113 @@ # aicodeman +## 1.31.0 + +### Minor Changes + +- 035bfbc: feat(remote): wake a sleeping remote host from Codeman + + A remote SSH case pointing at a machine that suspends used to fail the same way every + time: the session was there, the host was not, and typing into it went nowhere. A host + can now carry a wake target, either a MAC address for Wake-on-LAN (Codeman builds the + magic packet itself, so nothing reaches a shell) or a wake command of your own, and + Codeman uses it when you ask for the host: when you type into a sleeping session, when + you press the wake button on the banner, or when you start or attach a session on that + host. Input you type while it wakes is buffered and flushed once it is back, up to 4 KB, + and a chunk over that is refused outright rather than delivered as a fragment. + + Waking only ever happens because you asked. No watcher, dropped-session handler or + boot-recovery path can reach it, since a machine woken by a reconnect watcher would come + back seconds after every suspend. + +- fbee1b2: feat(custom-model): pick a custom endpoint straight from the Run menu + + #393 landed the backend for custom model endpoints and left it reachable only over the + HTTP API. This is the rest of it. Turn on Custom model endpoints in App Settings, save + an endpoint, and the Run dropdown grows a Custom Endpoints section built live off the + CLI registry, one entry per harness that can actually redirect plus each endpoint you + saved. Pick one and it launches that harness pointed at your server, asking which model + first when the endpoint has more than one. Endpoints re-discover themselves every five + minutes, and one unreachable endpoint never blocks the others. App Settings gains full + add, edit and delete for endpoints. + + Seven of the harnesses (opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP) now launch + directly onto the endpoint with no restart at all, where before you watched a native + boot followed immediately by a second one. Claude still launches and then restarts in + place, which its own resume makes far less jarring. + + Most of this release's work went into things that only show up against a real server, + and each was found that way rather than in tests: a freshly launched CLI reporting + itself busy for its own startup and getting refused; Claude Code assuming a large + context window for a model it does not recognise and silently overflowing a small one; + a model whose real context is below what Claude Code's own system prompt costs, which + no setting can fix and which now warns before launching into a certain failure; and the + big one, llama.cpp running exactly one model at a time, so applying a selection can + unload the model another session is using. That last case now asks first, tells you + which session it affects, and keeps a "loading model" notice on screen for the whole + swap window, so a prompt sent mid-swap reads as loading rather than as an answer from + whatever was loaded a moment ago. A background sweep also catches the reverse: your + session's model being evicted later by somebody else's ordinary use. + + Two things worth knowing if you drive this over the HTTP API or run multi-user. The two + questions an apply can ask (the model's context window is too small, and loading it will + unload the model another session is using) are now answered by separate + `confirmedContext` and `confirmedSwap` fields rather than one `confirmed`. They shared a + flag until now, and since the context check runs first, confirming that one silently + agreed to evict another session's model as well. The old `confirmed` still means both. + And `CLAUDE_CONFIG_DIR` is now admin-only in multi-user mode: it joined claude's + privileged env keys, so a non-granted owner can no longer set it through `envOverrides`, + and an already-persisted one is dropped on reboot-restore, which returns that session to + the default Claude account rather than the per-client one it was pointed at. Single-user + installs are unaffected. + + Remote SSH and Docker sessions are refused for now, since their restart reattaches a + durable tmux rather than relaunching the agent. + +### Patch Changes + +- c9515b1: fix(terminal): keep the output a pane capture could not contain. Opening a session, a backpressure refresh, a clear-terminal reload and a full-history re-pull all load the screen from a tmux pane capture, and anything the CLI printed between that capture and the end of the load used to be dropped, so its next partial redraw landed on a frame the terminal had never seen: missing or garbled output right after a tab switch or a refresh, plainest in a shell session. Each load now replays exactly the output that arrived after the capture, through one shared rule for all four paths, and a refresh that restores your scroll position no longer snaps back to the bottom afterwards. +- 3edf9aa: fix(terminal): replay a pane capture at the geometry it was taken at + + Opening a session could draw a frame built for a pane bigger than your terminal. A + taller pane wrote its overflow rows onto the last line and lost the rows underneath + (against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew + the survivors twice), and a wider one wrapped every row and scrolled the whole frame + up by one. The terminal response now reports the geometry the capture was really + taken at, so the browser can see the mismatch and replay once at the size that stuck. + A pane that cannot be sized to fit is diagnosed once per session instead of on every + tab switch. + +- 035bfbc: ### Thanks + - @irisitymichaelgrundberg for three terminal fixes in one release: keeping the output a pane capture could not contain (#436), replaying a capture at the geometry it was taken at (#435, five rounds and a Playwright suite that fails against the merge base), and trimming the padding out of a copied selection (#451), where the scan-instead-of-regex call avoided a 2.9s freeze nobody would have traced back to a copy. + - @timkjr for a first contribution that found a real silent failure: the Instance count stepper next to the Run button had only ever applied to Claude, so on the other eight run modes it launched one session and said nothing (#454). + - @Randalix for Wake-on-LAN on remote hosts (#439), built and live-tested against a real sleeping machine, and for reading the whole diff again between rounds rather than only the parts that were asked about. + - @opticon454 for turning #393's backend-only custom model endpoints into the whole feature (#430), and for validating it against a real llama-swap box rather than against the tests: the `/props` versus `/running` context discrepancy and the DeepSeek `/v1` root cause were both tracked down to the SDK source instead of guessed at. + +- c376534: fix(run): make the Instance count stepper work for every non-Claude mode + + The Instance count stepper next to the Run button only ever applied to Claude. + Setting it to 3 and launching OpenCode, Codex, Gemini, Antigravity, Pi, OMP, Grok or + DeepSeek started exactly one session, with no error and no hint that the control had + done nothing. All eight now launch the count you asked for, and the opening banner + says how many are starting. The one exception is a launch started from the Custom + Endpoints section of the Run menu, which always starts a single session. + +- 19ffe9b: fix(input): make sure a prompt sent through the API actually leaves the composer. Claude Code 2.1.277 started ignoring Enter for the first 30 to 50 seconds after the composer paints while still accepting the typed text, so a prompt sent right after a session came up sat unsent in the pane and every waiter (send-and-wait, the agent skill, cron, the maintainer bot) burned its whole timeout on a turn that never started. The server now reads the pane after every programmatic write that carried Enter and presses Enter again, on a 2 to 60 second schedule, only while the composer verifiably still holds the text it sent; an empty composer, other text, or a pane with no composer at all ends it. The agent skill's `sendwait` gets the same loop for servers that predate this, and its preamble version moves to 1.30.1 so an already-seeded agent picks up the fresh copy. +- f9edb33: fix(terminal): trim the padding out of a copied selection + + Copying out of a pane put a wall of spaces on the clipboard. xterm hands back + whole screen rows and trims only the cells that were never written to, so the + real spaces a full-screen program paints across the unused part of a row count + as content: measured against Claude Code in a 282-column pane, single lines + arrived carrying 138 trailing spaces. Pasting that into a chat client or an + editor meant deleting the whitespace by hand, while Windows Terminal, iTerm2 and + GNOME Terminal all trim it for you. A copy now drops the trailing run from every + line, on all four paths (the Ctrl+C chord, right-click, the phone selection + button and Auto Copy), while leading indentation is left exactly as it is. An + Alt+drag rectangular selection is copied verbatim, because its columns lining up + is the point of that gesture. A selection holding nothing but padding is refused + rather than copied as bare line breaks. + ## 1.30.0 ### Minor Changes diff --git a/CLAUDE.md b/CLAUDE.md index 196ad22e..19833f7f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -77,7 +77,7 @@ When user says "COM": CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed. -**Version**: 1.30.0 (must match `package.json`) +**Version**: 1.31.0 (must match `package.json`) ## Project Overview @@ -128,11 +128,11 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph ## Common Gotchas -- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink. ⚠️ **Input must END with `\r` or Enter is never sent**: `sendInput()` only issues `send-keys Enter` when the payload contains a carriage return, a `\r`-less `POST /api/sessions/:id/input` still succeeds (send-and-wait even reports `delivered:true`) while the text sits unsubmitted on the composer, and any `wait` burns its whole timeout on a turn that never started. Embedded newlines are stripped, not rejected, so `"echo A\necho B\r"` runs the joined `echo Aecho B` +- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink. ⚠️ **Input must END with `\r` or Enter is never sent**: `sendInput()` only issues `send-keys Enter` when the payload contains a carriage return, a `\r`-less `POST /api/sessions/:id/input` still succeeds (send-and-wait even reports `delivered:true`) while the text sits unsubmitted on the composer, and any `wait` burns its whole timeout on a turn that never started. Embedded newlines are stripped, not rejected, so `"echo A\necho B\r"` runs the joined `echo Aecho B`. ⚠️ **Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints** while still taking the typed text (measured 2026-09-19: an Enter at 28 s stranded the prompt, one at 51 s submitted it), so text+`\r` sent at readiness sits unsent with `0 tokens` and a `wait` burns its timeout. So the SERVER verifies every programmatic write that carried a `\r`: `SubmitVerifier` (`session-submit-verifier.ts`, armed from `writeViaMux`) reads the pane on a 2 s to 60 s schedule and re-sends Enter only while the LAST composer line (the CLI's own `promptGlyph`) verifiably still holds the head of what was sent; an empty composer, other text, or no composer line at all (a shell, a direct-PTY session) ends it, and a newer write replaces the schedule. The skill's `sendwait` keeps its own copy of the loop (`_composer_text` in `skills/codeman/preamble.sh`) for servers that predate this. The `shift+tab` footer only means the composer painted, never that Enter is accepted - **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks - **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly. Both `aicodeman` and `codeman` bin aliases are installed (`package.json` `bin`) - **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically) -- **`envOverrides` flow `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` / `ANTIGRAVITY_*` / `PI_*` / `GROK_*` / `XAI_*` / `DSH_*` / `DEEPSEEK_*` env vars, plus exact-key `CLAUDE_CONFIG_DIR`** — Set via `POST /api/sessions { envOverrides }`, stored on `Session._envOverrides`, exported by `tmux-manager.buildEnvExports()` at spawn time, persisted in `SessionState.envOverrides`. **Do NOT** write these to `/.claude/settings.local.json` — that's the old path and creates UI/disk drift. (`GOOGLE_*` is the deliberately-broad Vertex-AI namespace for Gemini — see Multi-CLI prefix discipline.) `CLAUDE_CONFIG_DIR` (#255, exact match via `ALLOWED_ENV_KEYS` in `schemas.ts`) points a session at a separate Claude account/config dir for per-client subscriptions; it persists to state.json (a path, not a secret; losing it on restart would silently switch accounts). ⚠️ A relocated config dir writes transcripts outside `~/.claude/projects`, so the response viewer, subagent windows, ultracode panel and Read My Mind capture go blind for that session unless the user symlinks `projects` back into the shared tree (`ln -s ~/.claude/projects /projects`). → [architecture-invariants#per-session-env-overrides-exact-key-allowlist-and-claude_config_dir](docs/architecture-invariants.md#per-session-env-overrides-exact-key-allowlist-and-claude_config_dir) +- **`envOverrides` flow `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` / `ANTIGRAVITY_*` / `PI_*` / `GROK_*` / `XAI_*` / `DSH_*` / `DEEPSEEK_*` env vars, plus exact-key `CLAUDE_CONFIG_DIR`** — Set via `POST /api/sessions { envOverrides }`, stored on `Session._envOverrides`, exported by `tmux-manager.buildEnvExports()` at spawn time, persisted in `SessionState.envOverrides`. **Do NOT** write these to `/.claude/settings.local.json` — that's the old path and creates UI/disk drift. (`GOOGLE_*` is the deliberately-broad Vertex-AI namespace for Gemini — see Multi-CLI prefix discipline.) `CLAUDE_CONFIG_DIR` (#255, exact match via `ALLOWED_ENV_KEYS` in `schemas.ts`) points a session at a separate Claude account/config dir for per-client subscriptions; it persists to state.json (a path, not a secret; losing it on restart would silently switch accounts). ⚠️ A relocated config dir writes transcripts outside `~/.claude/projects`, so the response viewer, subagent windows, ultracode panel and Read My Mind capture go blind for that session unless the user symlinks `projects` back into the shared tree (`ln -s ~/.claude/projects /projects`). ⚠️ It is also one of claude's `privilegedEnvKeys` (Custom Model Endpoint Profiles, since it can redirect a session's traffic same as any other injected var), so in multi-user mode setting it via `envOverrides` is admin-only, and a non-granted owner's already-persisted `CLAUDE_CONFIG_DIR` is stripped on reboot-restore — silently returning that session to the default Claude account rather than the one it was pointed at (see `session-env-clamp.ts`). → [architecture-invariants#per-session-env-overrides-exact-key-allowlist-and-claude_config_dir](docs/architecture-invariants.md#per-session-env-overrides-exact-key-allowlist-and-claude_config_dir) - **Effort is NOT an env var** — never carry effort as `CLAUDE_CODE_EFFORT_LEVEL`: the env var hard-locks effort and blocks in-session `/effort` switching (incl. ultracode). It flows as the dedicated `effort` payload field → `Session._effort` → `claude --effort ` for regular levels incl. `max` (the settings `effortLevel` key is `enum(["low","medium","high","xhigh"]).catch(undefined)` — `max` gets SILENTLY dropped there), or `claude --settings '{"ultracode":true}'` for ultracode (rejected by `--effort`). Both are soft defaults the user can override anytime. Legacy env-var entries are auto-migrated by the Session constructor and unset from tmux sessions in `applyEnvOverrides()`. See `buildEffortCliArgs()` in `session-cli-builder.ts`, tests in `test/effort-injection.test.ts` - **Model choice flows via `settings.local.json`, NOT `--model` or env** — the App Settings **Claude Model** picker (`claudeModel` in `settings.json`) is read by `session-ui.js` at session create (wins over the legacy 1M-Opus toggles `opusContext1m`/`opusContext1mEnabled`), sent as the `modelOverride` payload field, and `updateCaseModel()` (`hooks-config.ts`) writes/deletes the `model` key in `/.claude/settings.local.json`. This is the intended exception to the envOverrides rule above: model legitimately lives in `settings.local.json` (a soft default — in-session `/model` still works); env vars do not - **Multi-CLI prefix discipline** — env-var prefix is CLI-specific (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*` vs `ANTIGRAVITY_*` vs `PI_*` vs `GROK_*` vs `DSH_*`) and the `ALLOWED_ENV_PREFIXES` allowlist in `schemas.ts` enforces this; non-prefix exceptions are exact keys in `ALLOWED_ENV_KEYS` (currently only `CLAUDE_CONFIG_DIR`), never a widened prefix. Gemini additionally allowlists the **broad `GOOGLE_*`** namespace (intentional: Vertex AI auth needs `GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI`; it is the loosest allowlist entry, affecting only the user's own spawned CLI), and Grok allowlists **`XAI_*`** for the same vendor-namespace reason (`XAI_API_KEY` is grok's documented auth var). When adding a setting, decide which CLI(s) it applies to and gate the env export accordingly. Never blanket-forward all prefixes. ⚠️ Pi is the case that proves the rule: its ~34 provider keys (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `HF_TOKEN`, …) share NO prefix, and the allowlist is one GLOBAL list applied by a refine with no mode context, so admitting them for pi would widen it for every mode at once — they stay out, and pi users authenticate via `/login` or the server process's own env. ⚠️ DeepSeek repeats pi's lesson exactly: a dsh `settings.yaml` can nominate ANY env var as a provider credential (`apiKeyEnv`), so only the vendor namespaces `DSH_*` (launcher inputs incl. `DSH_PERMISSION_MODE`) and `DEEPSEEK_*` (`DEEPSEEK_API_KEY`/`DEEPSEEK_BASE_URL`) are admitted; foreign provider keys authenticate from dsh's own files or the server env. Resolver design pattern: `docs/opencode-integration.md`, `docs/pi-integration.md`, `docs/grok-integration.md`, `docs/deepseek-integration.md` @@ -165,13 +165,13 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph | **AI** | `src/ai-checker-base.ts`, `ai-idle-checker.ts`, `ai-plan-checker.ts` | | | **Tasks** | `src/task.ts`, `task-queue.ts`, `task-tracker.ts` | | | **State** | `src/state-store.ts`, `run-summary.ts`, `session-lifecycle-log.ts`, `intent-store.ts`, `tab-layout.ts` (pure model) + `-service` (sole mutation boundary) + `-persistence` + `-legacy-order` | | -| **Infra** | `src/hooks-config.ts`, `push-store`, `tunnel-manager`, `image-watcher`, `file-stream-manager`, `remote-hosts` + `remote-reconnect` (pure), `docker-hosts` + `docker-export` | Remote/docker case overlays; see Key Patterns | +| **Infra** | `src/hooks-config.ts`, `push-store`, `tunnel-manager`, `image-watcher`, `file-stream-manager`, `remote-hosts` + `remote-reconnect` + `remote-wake` (IO: `dgram`/`net`/`child_process`), `docker-hosts` + `docker-export` | Remote/docker case overlays; see Key Patterns | | **Web tabs** | `src/webview-store.ts`, `webview-capabilities.ts`, `src/web/webview-proxy.ts` (pure), `src/web/routes/webview-routes.ts` | Dashboard URLs as tabs; NOT a SessionMode | | **Search** | `src/search-service.ts` | Pure in-memory core for `GET /api/search` | | **Attachments** | `src/attachment-registry.ts`, `attachment-magic`, `generated-artifact-attachments`, `session-attachment-history`, `document-preview-cache`, `document-thumbnailer`, `document-conversion-limiter`, `config/attachment-guard` | See Key Patterns | | **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/` (`claude-md.ts` + `case-template.md`) | `templates/` holds the CLAUDE.md scaffold generated into new cases | | **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (27 modules + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | | -| **Frontend** | `src/web/public/app.js` (~6.9K lines, core) + 33 modules + `sw.js` (+ `voice-pcm-worklet.js`, fetched from JS, not in the load order) | See Frontend section for the load order, which is authoritative | +| **Frontend** | `src/web/public/app.js` (~6.9K lines, core) + 34 modules + `sw.js` (+ `voice-pcm-worklet.js`, fetched from JS, not in the load order) | See Frontend section for the load order, which is authoritative | | **Types** | `src/types/index.ts` (barrel) → 22 domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts | ★ = Large, central file (>50KB) — read its `@fileoverview` first. All files have `@fileoverview` JSDoc — read that before diving in. Discovery aid: `grep -l '@fileoverview' src/web/routes/*.ts` lists all route modules; same grep works for `src/types/`, `src/web/public/*.js`. @@ -217,6 +217,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Remote sessions + remote SSH cases**: a case can point at a remote host. The agent runs inside a durable remote `tmux -L codeman-remote` (session name `codeman-ssh-`, deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so an instance on the target host never adopts it), fronted by a LOCAL tmux pane running `ssh`. Attached (`owned:false`) sessions **detach, never kill** on tab close; owned ones propagate `kill-session`. A bounded-backoff watcher auto-reconnects dropped sessions (`remoteAutoReconnect`, default ON). ⚠️ **It revives ONLY when the durable remote tmux session is verifiably still alive** (`remoteTmuxSessionAlive()`, a `has-session` probe over ssh, #355): a clean agent exit (Ctrl-C, Ctrl-D, `exit`) tears that session down, and `isPaneDead()` cannot tell it from a transport drop, so the watcher used to relaunch a FRESH agent after every clean exit (claude only looked fine because its `|| --resume` fallback masked it). An unreachable host answers `undefined`, which also means do not revive. ⚠️ `has-session` prints NOTHING on success, so the probe is classified by EXIT STATUS (`classifyRemoteAliveExit`: 0 alive, ssh's 255 or a timeout unknown, anything else gone); reading stdout classified every live session as gone and silently disabled transport-drop reconnects. The answer is cached per session and forgotten whenever the pane is seen alive again, or a stale `true` from one transport drop would revive the next clean exit. ⚠️ **File reads in a remote case are the second ssh surface** (#415, `src/remote-files.ts`): they go through `buildSshConnectionArgs()` as well, a browser-supplied path is only ever a `shellescape`d token, an unreachable host answers 502 (never 404), the size cap uses the REMOTE size, and no remote file is ever copied onto the server's disk — which is why writes, office previews and thumbnails are deliberately unsupported over ssh (the `PUT` guard sits BEFORE the local path validation, or a same-named local directory such as an sshfs mount takes the write). The probe's symlink resolution FAILS CLOSED (a path it cannot canonicalize is a 404, never its own unresolved string: the directory-only fallback let a `notes.txt -> ~/.ssh/id_rsa` link pass containment), and ssh children are BOUNDED by `src/remote-ssh-limiter.ts` plus one batched probe per attachment-history listing, because terminal output in a remote session is written on the remote host and a prompt-injected agent can print hundreds of `codeman://attach` links. The ATTACHMENT routes (a clicked path outside the case dir) go through the same layer, and which host a record is read from follows the SESSION, never the path string. ⚠️ **Command-injection surface: every ssh command line must flow through `buildSshConnectionArgs()`**, which `shellescape`s every user field. Never hand-build an ssh line elsewhere. ⚠️ Run flows must route remote cases through `POST /api/quick-start`, not `POST /api/sessions` (which stat-validates `workingDir` locally and has no `caseName`). → [architecture-invariants#remote-sessions-over-ssh](docs/architecture-invariants.md#remote-sessions-over-ssh), [#remote-ssh-cases](docs/architecture-invariants.md#remote-ssh-cases), `docs/remote-sessions.md` +**Wake-on-LAN (`remote-wake.ts`)**: an optional `RemoteHost.wakeMac` (Codeman builds the magic packet itself) or `RemoteHost.wakeCommand` (single executable path, run without a shell, takes precedence) lets the INPUT route, `POST /api/sessions/:id/wake`, and the user's own create/attach request (`POST /api/quick-start`, `POST /api/sessions` with `attachRemoteSession`, via `ensureHostAwake`) wake a sleeping host instead of writing into a stalled ssh pane. ⚠️ An explicit request — input, the wake button, or the user pressing Run/Attach — and NOTHING else may wake: the auto-reconnect watcher, `handleRemoteSessionDropped`, boot recovery and `cron-service.ts` have no access to the registry (a wake there would re-wake the host seconds after every suspend, and the create wake is wired in the route rather than the shared session service for exactly that reason), which `test/remote-wake.test.ts` asserts as two wiring guards — the second also pins that `server.ts` holds the registry for its LIFETIME only (`drop` on cleanup, `stop` on shutdown) and never calls a waking method. `GET /api/sessions/:id/reachability` merely probes and never wakes. Detection is a throttled bare TCP probe — deliberately no `ServerAliveInterval`, because keepalives move bytes into an idle connection every interval and that is what a byte-threshold idle detector must not read as activity. ⚠️ A host behind `jumpHost`/`socksProxy`/a `ProxyCommand` option is reachability-UNKNOWN (`isProbeable()`): the probe connects to `host:port`, which such a host does not answer even while ssh works, so the registry never buffers for it, never gates create/attach on it (`'unprobeable'`), and `/reachability` answers `reachable: null, probeable: false` — the banner keys on a PROVEN `false`, and the banner's 30 s poller runs only for a host with a wake target (a timer connecting to a host Codeman cannot wake is the same timer-driven traffic the keepalive rule forbids). Input arriving during a wake is buffered (a chunk over 4 KB is dropped whole, never delivered as a fragment; the route answers `{buffered:true}` / `{buffered:true, dropped:true}` so the caller can tell) and flushed in order after `reattachRemote()` with `fromUser` — a flush write that fails drops the rest (logged) rather than retaining it for a wake hours later; send-and-wait blocks instead and answers `OPERATION_FAILED` when the host never returns. ⚠️ In multi-user mode the attach path 403s a non-admin BEFORE the host is looked up: the wake runs an executable, and remote hosts are admin-only infra everywhere else. ⚠️ Browser keystrokes travel over the WebSocket, which deliberately does NOT pass through the registry (that is the hot path), so only the HTTP input path ever queues anything — the banner must not promise queued input for the Wake button. A request that waits on the wake (create/attach, and the button) uses the 40 s `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS`, not the 90 s session default, because the dashboard's reverse proxy cuts a request at its own 60 s `proxy_read_timeout`. The wake fields are re-read from `remote-hosts.json` on recovery and, throttled+cached via `RemoteWakeDeps.resolveRemote`, for a LIVE session, since the persisted `remote` snapshot never sees a field added later. UI: the amber `#hostWakeBanner` (`host-wake-ui.js`) with Wake / "Configure WoL" → `#wakeConfigModal`. The `remote:` SSE family is session-scoped in multi-user mode; a create/attach wake names its requester (`username`) since it has no session yet. `remote-wake.ts` refuses real IO under `VITEST` like `remote-files.ts`. + **Docker cases**: a case can point at a **container**, with any of the CLI run modes running inside it. Like remote-SSH this is a **LOCATION OVERLAY on cases, never a `SessionMode` of its own**. Exactly one long-lived container **per case**, shared by all its sessions, so killing a session kills only that session's in-container tmux and **never** `docker stop` while siblings remain. The workspace is a real host dir bind-mounted at the **same absolute path**, which is what keeps file-routes/watchers on real host bytes and makes the in-container transcript projHash match the host. Credentials are **seeded** (RO mount, copied into the container once) rather than shared RW, so in-container CLIs never write refreshed tokens back to the host, and bind mounts are excluded from `docker commit` so exports stay secret-free. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** Config drift is detected via a label hash and a drifted launch is REFUSED rather than silently launched with stale config. ⚠️ A case may instead **ADOPT** a container the user already runs (`DockerCase.owned === false`, mirror of remote-SSH's `owned:false`): Codeman only `exec`s into it and never creates, starts, stops, restarts or removes it, so a missing or stopped container FAILS CLOSED with an actionable message instead of being fixed. Absent = owned, so existing cases are byte-identical. ⚠️ An ADOPTED container may back SEVERAL cases at different in-container directories (`classifyAdoptContainerConflict` in `docker-hosts.ts`: an exact twin on the same container AND directory is refused, an owned container still backs exactly one case, and a container another user adopted is refused), which is what the Add Case panel's "copy an existing case" picker relies on; the wire carries `CaseInfo.docker.owned` ONLY when false, so the picker tests `=== false`, never truthiness. The guarantee is enforced at four independent layers because it cannot be observed by using the feature: `buildDockerStopCommand`/`buildDockerRemoveCommand` throw during pure STRING CONSTRUCTION, `removeDockerContainer` refuses again, drift reports "none" (an adopted container carries no `codeman.confighash` label, so a real comparison would 409 the launch forever), and the boot reaper skips it. ⚠️ Two lifecycle touches the original design missed and that are easy to re-introduce: the full-image export `docker commit`s the container (refused for an adopted case) and the workspace export `docker pause`s it first (skipped — it freezes the owner's processes for the length of the tar). ⚠️ `owned` is applied AFTER `dockerConfigHash`, which takes an explicit field list, or every pre-existing case would trip the drift gate at once. ⚠️ Run modes for a container case come from the CONTAINER (`availableModes`, live-probed): gating the run menu on HOST CLIs (#201) is right for local sessions and wrong here, since a host with no `claude` may run a container that ships one. ⚠️ **A failed probe means opposite things per ownership** — for an ADOPTED case it is a fault worth reporting, for an OWNED one it is the NORMAL state before the first session (the launch chain creates the container), so treating it as a fault hid every agent mode on every freshly linked Docker case behind "start it yourself first". That is why `CaseInfo.docker.owned` is on the wire. ⚠️ Claude is launched WITHOUT `--dangerously-skip-permissions` when the container's exec user is root (Claude Code refuses the flag as root and the refusal is visible only inside the container); which flag to drop is a per-CLI fact, so it is the registry's `overlays.docker.rootCommand`, never a branch. ⚠️ Adoption is **admin-only in multi-user mode**, unlike `docker-link`: linking creates OUR container, whose one bind mount `isWorkingDirAllowed` has already confined, while an adopted container's mounts belong to its owner and one mounting `/` hands the adopter the host. The same reasoning admin-gates the container listing and the in-container directory browser; the preflight instead admits a non-admin for a container already linked to a case they own, because the run menu probes it for every docker case. ⚠️ On the loopback-only prod bind a container cannot reach 127.0.0.1, so in-container hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; otherwise idle detection falls back to output-based. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` (user guide), `docs/docker-cases-plan.md` (design) **Docker Compose deployment** (`docker/`, contributed): Codeman itself runs in a container and spawns Docker cases as **SIBLING** containers through the mounted host socket (Docker-outside-of-Docker), never nested. That inverts one assumption the bare-host path takes for granted: the daemon no longer shares Codeman's filesystem, so a bind source valid *inside* Codeman means nothing to it. `resolveDockerDaemonMountSource()` translates sources under HOME into the daemon's namespace via `CODEMAN_DOCKER_HOST_HOME`, and `CODEMAN_CASES_PATH` points the cases dir at a host-absolute bind mount so a workspace resolves to the SAME absolute path on both sides (which is what keeps the transcript projHash matching, per Docker cases above). ⚠️ **`CODEMAN_CASES_PATH` must move every consumer or none**: it is resolved once in `config/cases-dir.ts` because `src/cli.ts` resolves case paths too, and when only the server's `CASES_DIR` learned the override, `codeman skill install --case ` reported "Case not found" on exactly the deployment the override exists for. ⚠️ **`.dockerignore` patterns match the WHOLE context-relative path**, so a bare `.env` line excludes only the ROOT file: `docker/.env` (which holds `CODEMAN_PASSWORD` and any provider keys) rode `COPY . .` into the image until `**/.env` was added — verified in both directions with a real build context. ⚠️ A Compose LONG-form bind (`type: bind`) **creates a missing host source directory ROOT-OWNED** rather than refusing. `Start-Codeman.sh` pre-creates both `CODEMAN_APPDATA_PATH` and `CODEMAN_CASES_PATH` on the host before `up`, which is what keeps the daemon from ever having to materialise either as root in the first place; the container ALSO starts as root (`cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the base `cap_drop: ALL`; `test/docker-entrypoint.test.ts` pins that list) so `docker/entrypoint.sh` can correct a bind source that turns up root-owned anyway (a restored backup, a cleared directory, plain `docker compose up` run without the script) before dropping to `PUID:PGID` via `setpriv` — a directory owned by neither root nor `PUID:PGID` is never re-owned, since that ownership is not this container's to reassign; it is PROBED for writability as the runtime account (`setpriv ... test -w`, so ACLs, group-writable trees and CIFS/NFS mounts pass) and refused with a message naming path, owner and PUID:PGID if that fails. ⚠️ `KILL` is in that list for tini, not the entrypoint: `init: true` keeps tini as root while the server runs as PUID, and without CAP_KILL its SIGTERM forward fails and the server is SIGKILLed on every `compose down`/`restart` instead of flushing state. ⚠️ `/opt/codeman-cli` (the runtime-owned CLI prefix) is APPENDED to `PATH`, never prepended, and the entrypoint pins its own `PATH` to the system dirs: the root part of the start resolves `setpriv` by bare name, and a prefix ahead of `/usr/bin` let a planted `setpriv` run as uid 0 (measured). `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` drops `--memory-swap` (and filters only that one kernel warning) for hosts without swap accounting; `--memory` still applies. ⚠️ The deployment ALSO self-updates in place (the repo bind mount at `/opt/codeman` + a restart-by-exiting supervisor) — see Self-update below and `docs/docker-self-update.md` before touching `server.Dockerfile`, the compose file or `.env.example`, since each is an input to the updater's environment gate. `docs/docker-compose.md` + `docker/README.md` (user guides) @@ -227,7 +229,9 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. ⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. ⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. ⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs. -**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; backend + HTTP API only until the Run-menu picker lands, and the setting is read by nothing yet): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ Applying a selection **restarts the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct but codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement — a confirmed protocol gap, not a bug; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek reaches the server but gets a consistent `HTTP_404` (root cause not identified); antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. +**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ **Two launch paths, chosen by mechanism, not preference — see the second paragraph below for why**: opencode/codex/gemini/pi/grok/deepseek/omp apply the selection ONE-SHOT, before the session/process ever exists, with no restart at all; claude alone still applies a selection by **restarting the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` kills a WORKING pane, so a CLI whose launch declares a `fallback` chain (claude) gets the live conversation id pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct, and codex only speaks the Responses API since Feb 2026 (`wire_api = "responses"`) — re-verified live against a llama-swap deployment that DOES answer `/v1/responses` (an earlier test's harder failure against a different deployment does not reproduce everywhere): a plain, no-tool-call chat turn gets a real reply, but a real tool-call attempt came back as `agent_message` TEXT (the tool-call JSON printed as the answer) rather than an executable `function_call` item — confirmed via `codex exec --json`'s raw event stream. Tool execution is what makes codex a coding agent, so it remains not usable for real work either way, just with a more precise failure mode than a flat protocol break; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek's originally-reported `HTTP_404` is root-caused and fixed — its bundled `@deepseek-ai/dsh-llm-deepseek` module builds `${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` of its own (confirmed by installing the real package and reading its source), so a new `appendV1Suffix` flag on its registry entry (alone — claude/gemini must not get it) runs `endpoint.baseUrl` through `withV1Suffix()` before writing it, live-confirmed against llama-swap (`.../chat/completions` 404s, `.../v1/chat/completions` succeeds) though not yet re-run through an actual `dsh` binary, which isn't installable in this environment; antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. ⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`, JSON-escaped against a literal `` via the exported `escapeScriptJson()` since `label` is a user-`clis.json`-settable string unlike the neighbouring booleans-only `__codemanCliAvailable`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces. One entry per (capable, INSTALLED CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)", filtered through `isCliAvailable()` like the stock entries. Clicking one calls `selectCustomModelEntry(mode, endpointId)` (`session-ui.js`), which re-fetches the endpoint (never trusts anything cached from the dropdown's render — the 5-minute sweep below or a settings edit may have changed it since) and decides the model: exactly one discovered model launches straight away, two or more open `#customModelPickModal` to ask, with `defaultModelId` marked but never auto-chosen (asking exists so ONE launch can deliberately differ from the saved default). Either way the actual launch (`runCustomModelEntry`) routes through `run()` itself via a temporary `_runMode` swap — never `setRunMode()`, which would persist it as the user's new default — rather than a parallel dispatch table, which is what gives a custom-model launch the same `_runInFlight` lock every other Run click gets and means a CLI whose injection recipe lands later needs no update here. It then GETs `/api/sessions/:id/wait?until=idle&timeout=20000` on that session BEFORE applying — measured live, a freshly launched CLI reports itself `busy` for its own startup (boot spinner, workspace-trust check) well before the apply call would otherwise reach it, and the apply route's `isBusy()` guard correctly can't tell that apart from a real turn in progress, so every fresh launch failed with `SESSION_BUSY` until this wait was added. A timeout there is a normal 200 per the wait endpoint's own contract, never an error, so a session still busy after 20s just reaches the apply call anyway and gets that route's own honest error. It then calls `POST /api/sessions/:id/custom-model` on the session `run()` produced, guarded by snapshotting `activeSessionId` before the call and requiring it to have actually changed after — every `run*()` handles its own failure internally and returns normally rather than throwing, so a declined/failed launch must not silently re-point and restart whatever session was already open. ⚠️ The apply call reads the response body itself (`_api()`) rather than `_apiJson()`, which unwraps success but silently discards a failure body — losing the one thing (`error`) that distinguishes "still busy", "remote/Docker session" and everything else the route can report (⚠️ neither route validates `modelId` against the endpoint's discovered list, deliberately: discovery can be up to 5 minutes stale, so a 400 there would refuse a launch that works — a typo'd id fails on the CLI's own first request instead); the resulting toast is `type: 'error'` with an explicit `duration: 0` (no auto-dismiss, an explicit close button) at that one call site — not a blanket sticky-error default, which stacked unbounded on `.toast-container` with no cap or eviction — precisely so a message worth diagnosing survives long enough to be read instead of vanishing on the usual 3s timer. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to launch with). ⚠️ **Every saved endpoint's models also re-discover themselves automatically**, a `this.cleanup.setInterval` in `server.ts` (`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, 5 minutes, off under `testMode` like the Codex plan-usage poll beside it) calling the exported `refreshAllCustomModelHosts()` (`custom-model-routes.ts`) — one endpoint unreachable on a cycle never blocks the others, and a read-modify-write PER HOST (re-reading the store before each splice, keyed by id) means an admin's concurrent edit or delete wins over a sweep that started before it, never the reverse. + +**Everything below landed after the initial backend + picker cut, each confirmed live against a real llama-swap deployment.** ⚠️ **llama-swap runs one model at a time, and switching can disrupt ANOTHER live session** — before applying, both apply routes call llama-swap's own `GET /running` (feature-detected via `getLlamaSwapStatus()`, `custom-model-routes.ts`; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is simply never checked). If a different model is loaded and ready AND another live session's own selection is using it, the apply returns `{requiresConfirmation, currentlyLoadedModel, affectedSessions}` instead of switching silently; retrying with `confirmedSwap: true` skips the check, and switching with nothing else affected proceeds immediately. ⚠️ **The swap question and the context-floor question below have SEPARATE flags** (`confirmedSwap`, `confirmedContext`), because the context check runs first and while both shared one `confirmed` a user who clicked past a too-small context silently consented to evicting another session's model too; the legacy `confirmed` still means both, since it shipped in the HTTP-API-only cut. llama-swap also has no dedicated "switch model" endpoint — the only thing that actually starts a swap is a real inference request naming the model (confirmed live: applying a selection alone never reached llama-swap's own logs, since nothing had asked it to load anything) — so both routes also fire `triggerLlamaSwapLoad()`, a fire-and-forget `POST /v1/chat/completions` with `max_tokens: 1`, whenever the target model isn't already loaded and ready. ⚠️ **That launch-time check cannot catch a swap caused by a DIFFERENT session's LATER, ordinary use** — confirmed live: a second Codex session picking a different model launched with no warning at all (nothing conflicted at that exact instant), yet it silently evicted the first session's model regardless, since llama-swap has no push notification of its own. `detectCustomModelSwapDisplacements()` (`custom-model-routes.ts`) is a separate periodic sweep (`server.ts`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS` = 20s) that compares each live custom-model session's own `modelId` against what `/running` actually reports loaded, broadcasting a `custom-model:swapped-out` SSE event — shown as a global toast, never tied to the displaced session's own tab, since the whole point is telling the user before they type into it — the first time a mismatch appears, via a caller-owned de-dupe `Set` cleared once that session's own model is loaded and ready again so a later, genuinely new displacement notifies again rather than staying silently un-notified forever after the first one. ⚠️ **Context length is read from the REAL launch command, never `/props`** — `/props?model=`'s `default_generation_settings.n_ctx` was confirmed live to report a `--fit-ctx`-launched backend's theoretical/trained maximum rather than the real runtime-configured size (a measured 154112-vs-16384 discrepancy, caught only because the unfixed value still overflowed), so discovery parses the actual configured size straight out of `/running`'s own `cmd` field instead (`parseCtxFromCmd`: `--fit-ctx ` first, then plain llama.cpp `-c`/`--ctx-size`), falling back to `/props` only when `cmd` states no recognizable flag at all. ⚠️ **Claude alone gets a context-window FLOOR check, on top of the ceiling `contextLengthVar` already fixes** — `exceedsSafeContextFloor()` (gated on the registry declaring `contextLengthVar`, so a no-op for every other CLI by construction) compares a model's discovered context against `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000): confirmed live, twice, that Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on the very first message, before any conversation history exists to compact, so a smaller real context fails outright regardless of what `CLAUDE_CODE_MAX_CONTEXT_TOKENS` says (that var only controls when HISTORY gets compacted, and there is none yet on message one). Below the floor, the apply returns `{requiresContextWarning, modelId, contextLength, minSafeContextTokens}` instead of launching, shown as an in-app dialog naming the actual fix: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on `--fit-ctx` auto-fit, which optimizes for the biggest MODEL that fits rather than the biggest CONTEXT. ⚠️ **A fresh, isolated `CLAUDE_CONFIG_DIR` looks like a brand-new Claude Code profile and replays its ENTIRE first-run sequence on every launch** — the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with a bypass-permissions flag) a one-time warning about it, confirmed live, none of which a real, already-onboarded profile shows again. `skipFirstRunPrompts` (claude's entry only, requires `apiKeyTrustFile` since it reuses the same file) pre-seeds that same "already been through this" state: `hasCompletedOnboarding` and this session's own `projects[workingDir].hasTrustDialogAccepted` merge into the same `.claude.json` the API-key trust file already writes to, and `skipDangerousModePermissionPrompt` merges into `settings.json` (a different file, same corrupt-tolerant merge). ⚠️ **The loading banner shows the REAL backend log line, not a guess, and has no countdown or auto-timeout at all.** `getLatestLlamaSwapLogLine()` holds one `GET /api/events` SSE connection open per endpoint (confirmed live to stay open indefinitely — read past 220KB over 8s with no `done`; idle-closed after 30s via `pruneIdleLlamaSwapLogTails`, same 20s sweep as the swap-displacement check above), parsing `logData` frames and keeping only `source: "upstream"` (the real `llama-server` process's own stdout) lines, never `source: "proxy"` (llama-swap's own request-access log). ⚠️ `GET /logs` — the endpoint this feature's own first cut targeted, since the name suggested it — was confirmed live to carry ONLY the proxy log and never a single backend line, even seconds after a real, verified model swap; caught and corrected by a live check before merge, not after. The banner itself dropped its size-scaled expected-time estimate and matching auto-timeout (a guess dressed up as a fact that could kill a genuinely slow load partway through on slower hardware) for a generic hardware/model-size disclaimer plus a user-driven **Cancel** button (`_showCenterStatus`'s `onCancel` option, a real button distinct from the plain "×" close glyph an `'error'`-type banner gets) that ends the wait and closes the session on the user's own call rather than a guessed deadline. **Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w-` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization) @@ -255,7 +259,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) -**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set); Shell selection and automatic drop recovery always use a bounded 1 MiB `?tail=` window. Shell loads the rest only when **Load full history** is pressed; ordinary scrolling must not trigger a multi-megabyte reset+replay on xterm's main thread. Other modes may re-pull at the TOP (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). Live writes are one-chunk-in-flight, released by xterm's parse callback, so xterm's private queue cannot bypass the browser's 128 KiB render cap. While WebSocket owns terminal I/O, duplicate SSE terminal events are dropped before JSON parsing, and recovery is single-flight per active session. ⚠️ **A `full=1` capture ENDS with a cursor move back to the pane's own caret position**, counted UP from the last replayed row — without it the caret stays where the last character landed, which for an agent CLI is the status line, and every cursor-relative update the CLI sends afterwards is measured from the wrong row. The move is relative, not `CUP`: absolute row addressing is only right while the browser's rows equal the pane's, and `resizeWindow` does not wait for tmux, so a capture can be taken before a requested resize applies. That makes row alignment load-bearing on this path: no transform that can DELETE A LINE may run over the capture, so it keeps its trailing blank rows and skips redraw-bloat stripping, the banner trim and the leading-whitespace strip. ⚠️ Those three skips key on whether a capture actually CAME BACK (`isFullCapture`), never on `?full=1` alone — the fallback to the byte history is a stream of successive frames that must still be stripped, and a session with no mux takes it on every load. A capture holding nothing visible returns '' so the byte history survives instead of a blank screen replacing it. ⚠️ A full re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) +**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set); Shell selection and automatic drop recovery always use a bounded 1 MiB `?tail=` window. Shell loads the rest only when **Load full history** is pressed; ordinary scrolling must not trigger a multi-megabyte reset+replay on xterm's main thread. Other modes may re-pull at the TOP (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). Live writes are one-chunk-in-flight, released by xterm's parse callback, so xterm's private queue cannot bypass the browser's 128 KiB render cap. While WebSocket owns terminal I/O, duplicate SSE terminal events are dropped before JSON parsing, and recovery is single-flight per active session. ⚠️ **A `full=1` capture ENDS with a cursor move back to the pane's own caret position**, counted UP from the last replayed row — without it the caret stays where the last character landed, which for an agent CLI is the status line, and every cursor-relative update the CLI sends afterwards is measured from the wrong row. The move is relative, not `CUP`: absolute row addressing is only right while the browser's rows equal the pane's, and `resizeWindow` does not wait for tmux, so a capture can be taken before a requested resize applies. That makes row alignment load-bearing on this path: no transform that can DELETE A LINE may run over the capture, so it keeps its trailing blank rows and skips redraw-bloat stripping, the banner trim and the leading-whitespace strip. ⚠️ Those three skips key on whether a capture actually CAME BACK (`isFullCapture`), never on `?full=1` alone — the fallback to the byte history is a stream of successive frames that must still be stripped, and a session with no mux takes it on every load. A capture holding nothing visible returns '' so the byte history survives instead of a blank screen replacing it. ⚠️ A full re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. ⚠️ **A visible capture now REPORTS the geometry it was taken at** (`captureCols`/`captureRows`, #435), because a frame built for a pane taller or wider than the browser is damaged two ways at once (overflow rows clamp onto the last line; a narrower browser wraps every painted row) and nothing in the response used to say so. Both fields are ABSENT when no frame was positioned, so every consumer tests `Number.isFinite`, never truthiness: a `display-message` cursor query that fails makes `capturePaneBuffer` return the raw capture while the route still labels it `mux-visible`. The comparison runs on `mux-visible` ONLY, the replay is capped at one attempt, and a pane that cannot be sized to fit latches in `_geometryRetryUseless` so it is diagnosed once per session rather than on every tab switch. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) **Terminal touch gestures: link taps and text selection**: on a touch device xterm's own handlers see neither — `touch-action: none` plus touchstart's preventDefault suppress the browser's compatibility mouse events, `_installMobileTapMouseGuard` drops the trusted ones that still arrive, and the synthetic `mousedown`/`mouseup` pair dispatched for mouse REPORTING goes to the `.xterm` root, an ANCESTOR of the screen element the linkifier and SelectionService listen on. So both gestures are driven explicitly. ⚠️ **A tap activates the link under it** through the SAME provider that feeds the hover linkifier (`_terminalLinkAtPoint`, containment mirroring xterm's `_linkAtPosition`), synchronously inside `touchend` — that is what keeps the user gesture `window.open` needs — and BEFORE any mouse report, mirroring `_handleDesktopTerminalClick`'s skip for a hovered link. Two rows keep their meaning: the caret's logical line (`_tapIsOnCaretLine`, where a tap places the cursor in text the USER typed) and TUI-owned rows (`_isActionableMobileTerminalTap`, answering a dialog). ⚠️ The caret line is the boundary rather than the tap INTENT, because a shell classifies every tap as `'input'` and gating on that would leave every URL in shell output inert. ⚠️ **Long-press selects** by driving xterm's public `select()` (renderer-independent — under WebGL the glyphs are pixels and native selection cannot exist), drag or a further tap extends, and Copy goes through `copyTerminalSelection()` for its execCommand fallback on plain-HTTP installs. Three guards are load-bearing and each came from a real phone: the compat mouse pair after `touchend` (xterm focuses on mousedown and SelectionService resets the model there, so the keyboard sprang up and the selection vanished on lift), the platform's own ~500ms long-press (Android Chrome focuses the nearest editable element — the helper textarea — through no event a handler can preventDefault, so a bounded focus guard blurs it and `contextmenu` is suppressed for the gesture window), and `copyTerminalSelection()`'s closing `terminal.focus()` (right on desktop, wrong on a phone). Tests: `test/terminal-touch-tap.test.ts`. @@ -300,7 +304,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph ### Frontend -Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `i18n.js`(1.5) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `terminal-keycode229-recovery.js`(5.55) → `sanitize-html.js`(5.6) → `app.js`(6) → `tab-rail-resize.js`(6.5) → `terminal-ui.js`(7) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `cron-ui.js`(9.7) → `settings-ui.js`(10) → `panels-ui.js`(11) → `readmymind-ui.js`(11.3) → `ultracode-panel.js`(11.5) → `approvals-ui.js`(11.6) → `reboot-restore-ui.js`(11.65) → `admin-ui.js`(11.7) → `session-ui.js`(12) → `webview-tabs.js`(12.5) → `mobile-overview.js`(12.55) → `home-sessions.js`(12.56) → `entrance-animations.js`(12.6) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15) → `ultracode-windows.js`(15.5) → `session-lineage.js`(15.6) → `image-input.js`(16). `i18n.js` translates static + newly inserted application DOM while skipping terminal/response/file/user-name surfaces; `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData). `terminal-keycode229-recovery.js` forwards a committed `input` event that xterm's `_inputEvent` guard drops (Chrome-on-Android soft keyboards send `composed: true` after a keydown), and only when xterm emitted no canonical data for that keystroke. ⚠️ **That decision is settled at the NEXT keydown as well as on its own zero-delay timer** (#441): the drain runs from xterm's custom key handler, which fires BEFORE xterm processes that key, so a soft keyboard that commits the last character and sends Enter in one InputConnection transaction puts the character on the wire ahead of the `\r`. On the timer alone that character is not merely late, it is LOST: xterm emits the `\r` first and bumps the canonical counter past the candidate's snapshot, so the candidate stands down (measured, `hell\r` where the user typed `hello`). The trade is that a keydown decides with less evidence than the timer did, since xterm's own keyCode-229 rescue has not run yet; that is safe for Enter, which clears the textarea so the pending diff emits nothing. Ordering is pinned by `test/terminal-keycode229-recovery.browser.test.ts`, which the CI gate does NOT run. +Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `i18n.js`(1.5) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `terminal-keycode229-recovery.js`(5.55) → `sanitize-html.js`(5.6) → `app.js`(6) → `tab-rail-resize.js`(6.5) → `terminal-ui.js`(7) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `cron-ui.js`(9.7) → `settings-ui.js`(10) → `panels-ui.js`(11) → `readmymind-ui.js`(11.3) → `ultracode-panel.js`(11.5) → `approvals-ui.js`(11.6) → `reboot-restore-ui.js`(11.65) → `admin-ui.js`(11.7) → `session-ui.js`(12) → `host-wake-ui.js`(12.2) → `webview-tabs.js`(12.5) → `mobile-overview.js`(12.55) → `home-sessions.js`(12.56) → `entrance-animations.js`(12.6) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15) → `ultracode-windows.js`(15.5) → `session-lineage.js`(15.6) → `image-input.js`(16). `i18n.js` translates static + newly inserted application DOM while skipping terminal/response/file/user-name surfaces; `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData). `terminal-keycode229-recovery.js` forwards a committed `input` event that xterm's `_inputEvent` guard drops (Chrome-on-Android soft keyboards send `composed: true` after a keydown), and only when xterm emitted no canonical data for that keystroke. ⚠️ **That decision is settled at the NEXT keydown as well as on its own zero-delay timer** (#441): the drain runs from xterm's custom key handler, which fires BEFORE xterm processes that key, so a soft keyboard that commits the last character and sends Enter in one InputConnection transaction puts the character on the wire ahead of the `\r`. On the timer alone that character is not merely late, it is LOST: xterm emits the `\r` first and bumps the canonical counter past the candidate's snapshot, so the candidate stands down (measured, `hell\r` where the user typed `hello`). The trade is that a keydown decides with less evidence than the timer did, since xterm's own keyCode-229 rescue has not run yet; that is safe for Enter, which clears the textarea so the pending diff emits nothing. Ordering is pinned by `test/terminal-keycode229-recovery.browser.test.ts`, which the CI gate does NOT run. **Entrance animations** (`entrance-animations.js`, all OFF by default): opt-in animations for the four things that appear when work starts, chosen per surface via `data-tab-anim` / `data-term-anim` / `data-win-anim` / `data-line-anim` on ``. Defaults are the `legacy` theme, so an untouched install behaves exactly as before and every hook short-circuits on its first line. ⚠️ Tabs and connection lines are **destroyed mid-animation** on every re-render (`_fullRenderSessionTabs()` replaces the strip's innerHTML; `_updateConnectionLinesImmediate()` does `svg.innerHTML = ''`), so both are tracked by id and re-applied to the fresh element with a **negative `animation-delay`** to resume rather than restart. ⚠️ The terminal-pane styles may animate **transform / opacity / clip-path only**, xterm's FitAddon derives rows+cols from `getComputedStyle(parent).width/height`, so animating width/height/padding there would resize the PTY; `test/entrance-animations.test.ts` pins that property allowlist, plus the rule→keyframes→theme-option chain a style silently does nothing without. ⚠️ **`blur` is the ONE style that puts a `filter` on the terminal container**, against the standing rule, because every alternative was measured against a live xterm and does not work: a `backdrop-filter` veil on `::before` blurs perfectly while STATIC and Chrome silently drops the backdrop the moment ANY animation runs on that pseudo-element (the veil computes `blur(15.3px)` and the text behind it stays razor sharp), and driving the radius from rAF buys the same full-screen blur per frame plus main-thread work. The cost the rule exists to avoid is inherent to blurring a terminal, so the style buys it knowingly: opt-in, OFF by default, one ~520ms run per session open, class straight back off, `will-change` still unset. Worst-case price, headless SwiftShader with no GPU: frame deltas 16.7ms → 33.3ms for the run, against 16.7ms flat for `fade`. Do not generalise it — a second filtered terminal style needs its own measurement. ⚠️ The `blur` connection line animates `filter` too, so both kinds of line hold their glow in **`--line-glow`** and both of its keyframes say `blur(N) var(--line-glow)`: the function lists then match and interpolate, instead of the glow vanishing for the run and popping back (a lineage line's glow is a different colour entirely, set per element). Its 100% frame deliberately omits `opacity` so the endpoint comes from the element's own resting value — 0.9 subagent, 0.72 lineage, 0.95 working — which is what `line-enter-fade`'s hardcoded 0.9 gets wrong. ⚠️ Window styles other than `beam` transform the window, which moves the rect its connection line is aimed at; `beam` deliberately animates opacity/filter only so its line can draw toward a stable target. Persisted to its own `codeman:*Anim` localStorage keys (per-device, deliberately NOT in the `.strict()` `SettingsUpdateSchema`); picker in App Settings → Appearance, full per-surface lab at `?animlab=1`. @@ -316,7 +320,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L **Welcome "Resume Conversation" list** (terminal-ui.js): `loadHistorySessions()` fetches once and caches the corpus on `_historyAll`/`_historyCases`; every subsequent view (filter box, sort select, expand, the periodic refresh in panels-ui.js) goes through `_renderHistoryList()`, so never append rows to `#historyList` directly or re-fetch to re-sort. ⚠️ The box height is **class-driven**: expanding the list without `.history-list.expanded` leaves the collapsed `max-height` in place and just deepens a scroll well, which is the bug #260 reported (35 sessions in a ~4-row box). ⚠️ The A–Z sort keys off `_historyRowLabel()`, the SAME string the row renders (`name || firstPrompt || path`), most rows are transcript-backed and have no session name, so sorting on `name` alone silently does nothing. ⚠️ A filter implies expansion, and `_renderSearch()` hides `#historyHeader` (title + controls) as one unit while a search is active. Tests: `test/history-list-controls.test.ts`. -**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)** lives in that same handler: with a selection it copies, with none it must `return true` **without** `preventDefault()` or the interrupt is lost. `copyTerminalSelection` is deliberately absent from `SHORTCUT_ACTIONS` because the generic capture loop preventDefaults every match it dispatches. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry) +**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)** lives in that same handler: with a selection it copies, with none it must `return true` **without** `preventDefault()` or the interrupt is lost. `copyTerminalSelection` is deliberately absent from `SHORTCUT_ACTIONS` because the generic capture loop preventDefaults every match it dispatches. ⚠️ **The gate tests the CLEANED selection, not the raw one** (`CodemanCopySelection.clean` in constants.js, pure; `cleanedTerminalSelection()` reads the live terminal): xterm returns whole screen ROWS and trims only never-written cells, so the spaces a full-screen TUI paints across the rest of a row are content and reach the clipboard (138 of them per line, measured in a 282-column pane). The clean drops each line's TRAILING run and nothing else. ⚠️ **A shared LEADING indent is deliberately NOT stripped**, and that is a decision, not a gap: it was built, measured and dropped before #451 merged, because it fired on 73% of real non-TUI text (401,445 windows sampled) and no width threshold separates a TUI margin from content (a claude pane's own margins are 2 and 5 columns; the commonest non-TUI shared run is 4). Do not re-add it without reading the rule in architecture-invariants. ⚠️ Alt+drag COLUMN selections are returned untouched (`_activeSelectionMode === 3`), and the padding-only clear is FEEDBACK rather than interrupt protection, since a padding-only selection now cleans to `''` and falls through to the PTY on its own. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry) **Per-device vs synced settings**: the `displayKeys` set in settings-ui.js is a **client-side merge policy**, not a wire filter. A display key seeds from the server only when localStorage has no value for it, which is what prevents one device overwriting another; `showPlanUsageLimits` is additionally `delete`d from the incoming payload outright. Separately, `SettingsUpdateSchema` is `.strict()` and simply **does not declare** `skin`, `showFileViewerButton`, `showCronButton`, `webglRendererEnabled`, `localEchoEnabled`, `cjkInputEnabled`, or `extendedKeyboardBar`, so sending one of those is a validation error. The rest (`showResponseViewer`, `showPlanUsageLimits`, `language`, and most `show*` keys) ARE in the schema and do persist server-side; they are per-device by client policy only. ⚠️ Adding a new per-device setting means deciding **both** questions: membership in `displayKeys`, and presence in the schema. @@ -350,7 +354,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L **SSE staleness watchdog** (`computeSseStale()` in constants.js, `_checkSseStale()` + a 5s interval in app.js): an `EventSource` that stops delivering does not always error, so `onerror` never fires, the header dot stays green, and every SSE-driven surface (tab status dots, sessions created on another device, renames) freezes until the user reloads. ⚠️ The 15s server keepalive was an SSE **comment** (`:keepalive`), and comments are **invisible to `EventSource` by spec**, so there was nothing a client could observe: it is now the named `sse:heartbeat` event (`cleanupDeadClients()`, sse-stream-manager.ts), which is exactly why the frame had to change type. ⚠️ Staleness is judged **only while the status is `connected`** and the device is online; that guard is the loop breaker, since a forced `connectSSE()` leaves `connected` immediately and cannot re-fire while a reconnect is in flight. ⚠️ The liveness stamp is applied inside `addListener` itself, so every registered handler (the `_SSE_HANDLER_MAP` wrappers AND the directly-registered ones) feeds it from one place; the heartbeat's own listener is a no-op that exists **only** to be registered, since `EventSource` drops named events nobody listens for. ⚠️ The watchdog interval is cleared at the top of `connectSSE()` and nowhere else (its only teardown path); clearing it elsewhere stacks intervals. Recovery needs no new sync path: the reconnect re-runs `handleInit` → `_resetAllAppState()`. The forced reconnect logs one diagnostic line, because a middlebox that strips heartbeats presents as "silently reconnects every 45s". -**Z-index layers**: subagent windows (1000), plan agents (1100), mobile/tablet fixed header (1200, `mobile.css`), modals on ≤768px (1300 — must beat the fixed header or the modal close button is buried), log viewers (2000), connection-loss overlay (2500, above the fixed header and modals), image popups (3000), response viewer (5000, backdrop 4999), file-preview overlay (5100 — must outrank the response viewer, which can launch it; at its old 2000 a path clicked in the chat opened BEHIND the chat), toasts/path picker (10000+, deliberately above the preview), terminal touch-selection bar (900 — above terminal content and the local-echo overlay, deliberately BELOW floating agent windows so it can never cover their controls), local echo overlay (7). +**Z-index layers**: subagent windows (1000), plan agents (1100), mobile/tablet fixed header (1200, `mobile.css`), modals on ≤768px (1300 — must beat the fixed header or the modal close button is buried), log viewers (2000), connection-loss overlay (2500, above the fixed header and modals), image popups (3000), response viewer (5000, backdrop 4999), file-preview overlay (5100 — must outrank the response viewer, which can launch it; at its old 2000 a path clicked in the chat opened BEHIND the chat), toasts/path picker (10000+, deliberately above the preview), the custom-model center-status banner (10001, `.center-status-banner` — `[hidden]` must re-assert `display: none` over its own `display: flex`, same trap as `.home-sessions[hidden]`, or `dismiss()` leaves an invisible click-blocker dead centre on screen), the swap-confirm and context-warning modals (10010, `#customModelSwapConfirmModal`/`#customModelContextWarningModal` — must clear both the plain `.modal` z-index of 1000 and the center-status banner it can appear over), terminal touch-selection bar (900 — above terminal content and the local-echo overlay, deliberately BELOW floating agent windows so it can never cover their controls), local echo overlay (7). **Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min). @@ -379,11 +383,11 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L ### SSE Event Registry -158 event constants in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). **Both must be kept in sync**, and `test/sse-registry-parity.test.ts` is the guard that pins it (currently exactly in sync, 158 = 158, no drift either direction). ⚠️ `hook:agent_working` is the one hook event with no Claude Code hook behind it — the DeepSeek status bridge reports it (see External CLI modes). The backend file's `@fileoverview` carries the per-category breakdown, including the two Web tab events. +161 event constants in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). **Both must be kept in sync**, and `test/sse-registry-parity.test.ts` is the guard that pins it (currently exactly in sync, 161 = 161, no drift either direction). ⚠️ `hook:agent_working` is the one hook event with no Claude Code hook behind it — the DeepSeek status bridge reports it (see External CLI modes). The backend file's `@fileoverview` carries the per-category breakdown, including the two Web tab events. ### API Routes -~233 handlers across 27 route files in `src/web/routes/`: system (56), sessions (34), cases (34), files (17), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), approvals (4), readmymind (4), custom-model (5), reboot-restore (3), me (2), teams (2), tab-layout (2), search (1), hooks (1), clipboard (1), status-telemetry (1), voice (1 + the `/ws/voice/stream` relay), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details. +~236 handlers across 27 route files in `src/web/routes/`: system (56), sessions (37), cases (34), files (17), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), approvals (4), readmymind (4), custom-model (6), reboot-restore (3), me (2), teams (2), tab-layout (2), search (1), hooks (1), clipboard (1), status-telemetry (1), voice (1 + the `/ws/voice/stream` relay), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details. **HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`). diff --git a/config/test-suites.ts b/config/test-suites.ts index 233e0e7d..c0c444af 100644 --- a/config/test-suites.ts +++ b/config/test-suites.ts @@ -28,6 +28,7 @@ export const BROWSER_TEST_GLOBS = [ 'test/terminal-copy-shortcut.test.ts', 'test/terminal-keycode229-recovery.browser.test.ts', 'test/capture-load-window.browser.test.ts', + 'test/capture-geometry-retry.browser.test.ts', 'test/codex-predictive-echo.test.ts', // also needs a real codex binary ]; diff --git a/docs/api-reference.md b/docs/api-reference.md index 53b38e8a..c85b4ff7 100644 --- a/docs/api-reference.md +++ b/docs/api-reference.md @@ -324,6 +324,30 @@ from the session's current state rather than requiring a new transition: the original turn may be long over. It comes back as `"delivered": false, "duplicate": true`. +**Wake-on-LAN hosts** (`docs/remote-sessions.md` §Wake-on-LAN): when the session's +remote host has a wake target and is asleep, the non-wait form answers `200` with +`{"buffered": true}` — the bytes are held and flushed after the host is back — or +`{"buffered": true, "dropped": true}` for a chunk over the 4 KB wake buffer, which +is gone (never delivered as a fragment). Both fields are additive to the historical +bare `{}`. With `wait`, the route blocks on the wake instead and answers +`422 OPERATION_FAILED` ("did not come back after a wake-on-LAN request — nothing was +sent") when the host never returns, rather than writing into the stalled pane and +reporting `delivered:true` plus a timeout. + +Two endpoints back that flow directly, both scoped to one session's remote host and +both refusing a session that is not remote (`400 INVALID_INPUT`): + +| Method | Path | Purpose | +| --- | --- | --- | +| `GET` | `/api/sessions/:id/reachability` | Whether the session's remote host answers SSH right now, plus whether a wake target is configured. Read-only: it never wakes. `{"reachable": true\|false\|null, "wakeConfigured": "mac"\|"command"\|"none"}`, where `null` means the answer is unknown (a proxied host, where a TCP probe proves nothing). | +| `POST` | `/api/sessions/:id/wake` | Wake the host and wait for it to accept SSH again, bounded by the request budget. `422 OPERATION_FAILED` when it does not come back; `400 INVALID_INPUT` with "No wake-on-LAN target configured for this host" when nothing is set. | + +⚠️ Waking is deliberately reachable only from an explicit user action (this route, a +session create/attach, or typing into a sleeping session). No watcher, dropped-session +handler or boot-recovery path may wake a host, or a suspended machine would be woken +again seconds after every suspend; `test/remote-wake.test.ts` pins that as an import +fence around `src/remote-wake.ts`. + ### Response All three nest the wait result under `data.wait`, so one client helper works against @@ -558,6 +582,115 @@ All four enforce session ownership in multi-user mode; a foreign session id answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the same directory are distinct by construction. +## Custom Model Endpoints + +Points a session's harness at a user-configured OpenAI-compatible endpoint — +local (llama.cpp, vLLM, DGX Spark) or cloud (Azure AI Foundry, OpenRouter) — +instead of its native cloud backend, gated by the opt-in +`customModelEndpointsEnabled` setting (default OFF). Endpoints are +machine-level infra, like remote/docker hosts: writes are admin-only in +multi-user mode. Design: [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md); +user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md). + +- `GET /api/v1/model-endpoints` -> `CustomModelHost[]`, an unwrapped bare + array like every other list route (still riding the standard `{success, +data}` envelope on the wire — unwrap it the same way). Answers `[]` for a + non-admin in multi-user mode. `apiKey` is never returned; `apiKeySet: +boolean` reports whether one is stored, so a client can render "unchanged + if left blank" without ever holding the real value. +- `POST /api/v1/model-endpoints` with `{ id, label, baseUrl, apiKey?, +authStyle?, defaultModelId? }` creates one. `id` must match + `^[a-zA-Z0-9_-]+$`; `authStyle` is `bearer` (default) or `api-key`, never + both (a real server hung indefinitely when sent both headers on one + request); `baseUrl` must be `http(s)`, carry no embedded credentials, and + is refused if it points at (or resolves to) a link-local or + cloud-metadata address. `409 ALREADY_EXISTS` on a duplicate id. +- `PUT /api/v1/model-endpoints/:id` updates one. An **absent** `apiKey` + keeps the stored one rather than clearing it — the client never receives + the real value to resend deliberately unchanged, so omission is the only + way to say "leave it alone"; there is no way to clear a key back to unset + this way. `defaultModelId`, when set, must be one of that endpoint's own + `models` (`400 INVALID_INPUT` otherwise). +- `DELETE /api/v1/model-endpoints/:id` removes one. +- `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's + own `GET /v1/models` and stores the result as `models`, updating + `lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own + response already reports loaded) `modelContextLengths` and `modelSizesGB`. + A `defaultModelId` that no longer appears in the fresh list is dropped + rather than carried forward invalid. Failures answer `422 OPERATION_FAILED` + with the underlying connection error, or a named egress refusal if the + resolved address turned out to be blocked. The same refresh also runs + automatically for every saved endpoint every 5 minutes in the background + (`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, started from + `server.ts`), so there is no route for triggering "refresh all" — one + endpoint being unreachable on a cycle never blocks the others. +- `GET /api/v1/model-endpoints/:id/running-status` -> `{ isLlamaSwap, +running: [{model, state}], logLine? }`, read-only, no admin gate + (any session owner who could already point a session at this endpoint can + equally ask what it currently has loaded). `isLlamaSwap` is + feature-detected via the endpoint's own `GET /running` — a plain + llama.cpp/OpenAI-compatible server has none and always answers `false`. + `logLine`, present only when `isLlamaSwap` is true, is the most recent + REAL backend `llama-server` process log line (`load_model: ...`, + `llama_server: model loaded`, etc.), sourced from the endpoint's own + `GET /api/events` SSE stream and filtered to `source: "upstream"` frames + only (never llama-swap's own `source: "proxy"` request-access log) — one + connection is held open per endpoint and reused across every poller, + idle-closed after 30s of nobody asking. This is what the Run-menu + picker's loading banner polls once a second while a model is loading. +- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId, +confirmed? } | { clear: true }` applies (or clears) the session's + selection and **restarts the session's CLI process in place** — every + supported harness reads its endpoint config at process start, never per + turn, so there is no live hot-swap. (`POST /api/v1/quick-start`'s own + `customModel: { endpointId, modelId, confirmed? }` field is the + no-restart equivalent for a session that doesn't exist yet — see below.) + A Claude session resumes its existing conversation across the restart; + pi/omp/grok additionally get a forced `--model`/`-m` value, since for + those three the config file alone does not select it. `400 INVALID_INPUT` + for a remote (SSH) or Docker session — both restart their agent + differently under the hood, and applying to one would report success + while changing nothing. Two more responses replace the normal + `{customModel, restarted}` shape, neither an error, and neither restarts + or creates anything on the first ask. ⚠️ **Each is answered by its OWN + flag on the retry, and answering one is not consent to the other**: they + are questions about different people, and while they shared a single flag + a caller who confirmed the context warning silently agreed to evict + another session's model as well. Send `confirmedContext: true` to proceed + past the context warning, `confirmedSwap: true` past the swap conflict, + and both when both were asked (they accumulate, so the second retry still + carries the first answer). The original `confirmed: true` still means + BOTH and is still accepted, because it shipped in this feature's + HTTP-API-only cut; new callers should send the specific one: + - `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` — + llama.cpp/llama-swap only runs one model at a time, and switching would + unload a model another **live session's own selection** is actively + using. Never returned for a plain (non-llama-swap) server, and never + just because a swap is needed at all — only when it would disrupt + someone else. + - `{requiresContextWarning: true, modelId, contextLength, +minSafeContextTokens}` — Claude Code's own fixed per-turn overhead + (system prompt + tool schemas) can exceed a small model's entire + discovered context on its own, before any conversation history exists + to compact, guaranteeing the very first message fails regardless of + `CLAUDE_CODE_MAX_CONTEXT_TOKENS`. Gated on the CLI registry declaring a + `contextLengthVar` (claude only today), so it never fires for another + harness. +- `POST /api/v1/quick-start`'s `customModel: { endpointId, modelId, +confirmed?, confirmedContext?, confirmedSwap? }` field (alongside its +normal `caseName`/`mode`/etc. body) + computes the same injection **before** the session exists and launches + directly on the endpoint — no restart, because there was never a + native-backend boot to restart away from. Runs the identical checks as + the dedicated route above (`requiresConfirmation`/`requiresContextWarning`, + same shapes, same per-question `confirmedContext`/`confirmedSwap` retry), + and is refused the same way + for a remote or Docker case. This is what the Run-menu picker uses for + opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the + dedicated restart route above (its `--resume`-based restart is far less + jarring than a full relaunch, and folding it into the one-shot path is + separate work — see `docs/custom-model-endpoints-plan.md`). + ## Voice dictation Browser dictation transcribed through this server's Claude Code login, i.e. the diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index fb4c727f..9de5be90 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -54,6 +54,8 @@ Model is NOT a session field: it is a composition entry in the profile's config ### Remote SSH cases +**Remote host wake-on-LAN from user input**: an optional `RemoteHost.wakeMac` (magic packet built and broadcast by Codeman) or `RemoteHost.wakeCommand` (a single executable path, run WITHOUT a shell, and the explicit override) lets the input route — and an explicit `POST /api/sessions/:id/wake` — wake a SLEEPING host instead of writing into a stalled ssh pane; `tmux send-keys` succeeds against a stalled pane, so the bytes used to vanish silently. The wake flow lives in `src/remote-wake.ts` and is reachable **only** from an EXPLICIT user request: `POST /api/sessions/:id/input`, that explicit wake route, and the create/attach path (`POST /api/quick-start` for a remote case, `POST /api/sessions` with `attachRemoteSession`, via `ensureHostAwake`), because "the user pressed Run on a sleeping host" is the same kind of request and the tmux probe would otherwise fail with a misleading "needs tmux installed". Everything TIMER-driven must never wake a host: the COD-108 auto-reconnect watcher, `Server.handleRemoteSessionDropped` and boot recovery have no access to the registry, or a host would be re-woken seconds after each suspend and could never stay asleep (asserted by wiring guards in `test/remote-wake.test.ts`, not just documented — including that `ensureHostAwake` is called from the HTTP route only, since `cron-service.ts` builds sessions through the shared service with nobody waiting on the answer). `GET /api/sessions/:id/reachability` only ASKS — it never wakes — and feeds the amber "host unreachable" banner (`host-wake-ui.js`) whose action is either Wake or, with no target configured, "Configure WoL" → `#wakeConfigModal` (saved via `PUT /api/remote-hosts/:id`). Detection is a throttled bare TCP probe (no ssh, no `ServerAliveInterval` — keepalives would move bytes into an idle connection every interval; and a host behind a jump host/SOCKS proxy is reachability-UNKNOWN, never "asleep": `isProbeable()` keeps the registry from buffering, gating or bannering on a probe that cannot reach it), input is buffered and flushed in order after `reattachRemote()` (the send-and-wait path blocks instead, as does the create path, with a shorter request budget), and the wake fields are re-read from `remote-hosts.json` on recovery AND (throttled, cached) live for a running session, because the persisted `remote` snapshot would never see a field added later (`rehydrateRemoteHostFields` + `RemoteWakeDeps.resolveRemote`). Design + invariants: `docs/remote-sessions.md` §Wake-on-LAN from user input. + **Remote SSH cases** (COD-94/#145): cases can point at a **remote host** (`~/.codeman/remote-hosts.json` + `remote-cases.json` via `src/remote-hosts.ts`; CRUD under `/api/cases` — cases route file). A remote session launches a LOCAL tmux pane running `ssh ` that creates a durable REMOTE tmux session on a **dedicated socket** `-L codeman-remote` with name `codeman-ssh-` — deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so a Codeman instance on the target host never adopts it; no `-g` global tmux options are set remotely. `remotePath`/`identityFile` are schema-guarded against shell injection (backticks/`$` rejected — same approach as `extraSshOptions`); remote tmux availability is probed via `checkRemoteTmuxAvailable()` in quick-start (ssh args carry `-o ConnectTimeout=10`). Remote claude defaults to an idempotent `claude --session-id || claude --resume ` pair under a login shell, so a respawn or reattach continues the SAME conversation rather than starting a fresh one (remote omp gets the same treatment via `--continue`; ⚠️ because the claude arm is an `a || b` pair under `-c`, that pane's PID is the login shell, not the agent); per-host `commands.*` override. Session kill best-effort kills the remote tmux too. `SessionState.remote`/`MuxSession.remote` round-trip through recovery (`restoreMuxSessions` passes `remote` back into the Session constructor). ⚠️ Run flows must route remote cases through `POST /api/quick-start` (which resolves the remote case and skips LOCAL CLI availability gates) — `POST /api/sessions` stat-validates `workingDir` locally and has no `caseName`. `envOverrides`/`effort`/`modelOverride`/`codexConfig`/`geminiConfig` are rejected for remote quick-starts (not silently dropped). UI: Create Case modal → Remote tab. Tests: `test/remote-hosts.test.ts`, `test/remote-ssh-options.test.ts`. ⚠️ **Reading a file in a remote case goes over ssh too** (#415): `src/remote-files.ts` is the single remote-READ layer (`buildRemoteFileCommand` = `buildSshConnectionArgs` + one shellescaped remote command; `remoteProbePaths` returns remote realpath + stat; `remoteCreateReadStream` streams a `Range` via `tail -c +N | head -c L` and its `close()` must be wired to the response's `close` or the ssh child outlives an aborted download). The guard order matches the local path exactly (`validateSessionFilePathLexical` → remote realpath of BOTH file and workspace root → containment → sensitive-path → size cap on the REMOTE size), a request path arrives from the browser and is only ever interpolated as a `shellescape`d token, and an unreachable host answers **502**, never a 404. ⚠️ The probe's symlink resolution FAILS CLOSED: `readlink -f` where it exists, otherwise a `cd -P`/`pwd -P` directory walk plus a bounded plain-`readlink` loop over the last component, and anything it cannot fully resolve is reported unresolvable (404), never as the unresolved string — the first version resolved the directory chain only, so on a host without `readlink -f` a `ws/notes.txt -> ~/.ssh/id_rsa` link passed containment under its own path while `cat` served the key. Records are NUL-separated and index-keyed so a newline in a filename cannot shift the mapping. ⚠️ ssh children are BOUNDED: probes and buffered reads go through `src/remote-ssh-limiter.ts` (a `document-conversion-limiter`-shaped semaphore, default 4), the attachment-history list probes its whole history in ONE batched call (`probeRemoteAttachmentHistory`, threaded into `registerExternalAttachment({remoteProbes})`), and probes chunk at 40 paths — a prompt-injected agent printing `codeman://attach` links in a remote session used to fork one `ssh` per link. `describeExecError` never returns Node's `Command failed: ` message (identity path + probe script in a 502 body). The `PUT /file-content` guard sits AHEAD of `validateSessionFilePath`, which resolves LOCALLY, or a same-named local directory (an sshfs mount) takes the write. Under `VITEST` the three IO functions refuse rather than connect. This covers the ATTACHMENT routes too, which is the half a clicked path needs when the file is OUTSIDE the case directory (`_isExternalPreviewPath` sends it to `POST …/attachments`): registration, by-id `raw`, metadata and the history list all resolve over ssh (`registerExternalAttachment({remote})`, `resolveServableRemoteAttachment`), and what decides the host is the SESSION, never the path string — the same absolute path means a different file on each host. Deliberately NOT supported over ssh: writes (`edit=1`/`PUT` answer 400, `editable` is always false), office previews/thumbnails, the file tree/picker, `tail-file`. Tests: `test/remote-files.test.ts`, `test/routes/file-routes-remote.test.ts`. ### Docker cases @@ -72,6 +74,8 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do **The env allowlist has two tiers, and exceptions go in the exact-key tier, never a widened prefix** (#255): `ALLOWED_ENV_PREFIXES` in `src/web/schemas.ts` carries the CLI-namespace prefixes, and `ALLOWED_ENV_KEYS` carries exact keys (currently only `CLAUDE_CONFIG_DIR`). `CLAUDE_CONFIG_DIR` relocates the Claude CLI's user config (credentials, settings, stats), which is how one machine runs sessions on separate Claude subscriptions: point a case's sessions at e.g. `~/.claude-clients/acme` via `envOverrides` and run `/login` there once against the client's account. The exact match matters: `CLAUDE_` as a prefix would open every future Claude CLI variable unreviewed, and near-misses (`CLAUDE_CONFIG_DIR_EXTRA`) stay rejected (`test/env-overrides-schema.test.ts`). No new security boundary is crossed: sessions already run as the server's OS account, and `applyEnvOverrides()` shellescapes values into socket-scoped `tmux setenv`. Two carry rules: **(1)** the key must survive `getEnvOverridesForPersist()` in `session.ts` (it is a path, not a secret; dropping it from state.json would silently move a rebuilt-after-reboot session back to the default account); **(2)** ⚠️ a relocated config dir writes transcripts outside `homedir()/.claude/projects`, which `subagent-watcher.ts`, `workflow-run-watcher.ts`, the response-viewer routes and Read My Mind capture all hardcode — those surfaces go blind for such a session. Documented workaround: symlink the transcripts back into the shared tree (`ln -s ~/.claude/projects /projects`), keeping credentials separate while the watchers keep working. +⚠️ **As of the custom-model endpoint feature, `CLAUDE_CONFIG_DIR` is ALSO admin-only in multi-user mode**, which is a change to the above rather than a restatement of it. It joined claude's `privilegedEnvKeys` (`stock.ts`) alongside `CLAUDE_CODE_MAX_CONTEXT_TOKENS`, because the rule that every traffic-redirecting var that feature can inject must be listed there is worth keeping literally true. `privilegedEnvKeys` has exactly one consumer, `ownerClampedEnvKeys()` in `src/session-env-clamp.ts`, which feeds the generic `envOverrides` clamp on `POST /api/sessions`, `POST /api/quick-start` and reboot-restore. So for a non-granted owner two things now follow: the key cannot be set through `envOverrides` at all, and an ALREADY-PERSISTED one is stripped on reboot-restore, which silently moves that session back to the default Claude account. That second consequence is the one to watch, since it turns a working per-client setup into a wrong-account one across a host reboot with no error anywhere. `session-env-clamp.ts`'s own fileoverview used to state the opposite invariant (that claude's privileged keys are the five `ANTHROPIC_*` names, so a persisted record cannot carry a clamped key) and was corrected when this landed; a claude clamp test now pins the behaviour next to the deepseek and omp ones. + ### Agent wait primitives **Agent wait primitives** (`GET /api/sessions/:id/wait`, `GET /api/sessions/:id/wait-output`, and the `wait`/`waitTimeout` fields on `POST /api/sessions/:id/input`): bounded long-polls that let an agent driving Codeman from a shell tool block until something happens. They exist because SSE was the only "tell me when" channel Codeman had, and a curl-driven caller cannot practically hold a stream and parse events inline. The blocking core is `src/web/session-wait-registry.ts` (no IO, no `Session` reference, so it unit-tests in isolation), bounds live in `src/config/agent-wait.ts`, and the wiring is three `notifySignal()` calls next to existing broadcasts (`session-listener-wiring.ts` for `working`/`idle`/`exit`, `hook-event-routes.ts` for `stop`/`blocked`) plus `notifyOutput()` riding the already-attached `terminal` listener. Design: `docs/agent-control-plan.md` §3; wire contract: `docs/api-reference.md`. @@ -132,6 +136,8 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do **Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell full history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[;H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. +**A capture reports the geometry it was taken at** (#435): a visible frame repaints each row at an absolute position, counting up to the pane's height and out to the pane's width, so a terminal smaller than that pane damages it two ways at once. Too short and every address past the browser's own height clamps onto the last line, overwriting the rows underneath (measured: against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew the survivors twice). Too narrow and each row is painted out to the pane's width, so the browser wraps every painted row and the wrap on the last one scrolls the whole frame up by one. Nothing in the response used to say what geometry the frame was built for, so the client could not see either case. `PaneCaptureOptions.capturedGeometry` carries it out, and the terminal response publishes it as `captureCols`/`captureRows`. ⚠️ **Both fields are ABSENT unless a frame was really positioned**, and every consumer must test `Number.isFinite` rather than truthiness: `mux-visible` is necessary but not sufficient, because when the `display-message` cursor query fails `capturePaneBuffer` skips the snapshot repaint and returns the raw capture, and the route still labels that non-empty body `mux-visible`. A body that positioned nothing has no geometry to describe and nothing to repair, so a comparison that fires there buys a second capture, a reset plus chunked rewrite, a dropped and reopened WebSocket and a discarded xterm snapshot for no gain. ⚠️ **The comparison runs on `mux-visible` ONLY.** A full-history body is linear scrollback closed by a RELATIVE cursor move, which is relative precisely so the browser's row count need not match the pane's, and a byte-history body carries no row alignment at all, so a size mismatch damages neither and a replay repairs neither. That gate matters because the first select of every non-shell session per page takes the full-history path, where an ungated comparison would fire most often on the one response it cannot help, at the price of a second whole-scrollback capture. ⚠️ **The replay is capped at one attempt and latches per session when it cannot converge.** `resizeRetry` stops two competing fits trading replays forever; a pane already drawing at the size just requested is left alone, which is the signature of a clamp rather than a race (`getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does not, so a terminal under 40 columns or 10 rows reports a pane permanently bigger than itself and would replay on every tab switch); and a pass that still does not converge joins `_geometryRetryUseless`, so the case `Session.resize` declines outright (a small viewport while a desktop viewport's size claim is live, where the retry re-sends the same declined resize and captures the same pane) costs one attempt per session per page load instead of one per select. ⚠️ **A retry pass must not re-arm `_fullHistoryLoaded`**: it did not consume the full-history pull, and re-arming it would spend a whole-scrollback capture on the next select. That branch is currently unreachable by construction, since reaching it needs `source === 'mux-visible'` while a `full=1` pass is answered `mux-full-history` or `history`; a static test over the source is the habit this repo uses for an invariant nothing can execute. Tests: `test/capture-geometry-retry.browser.test.ts` (eight cases, five of which fail against the merge base), `test/tmux-capture-full-history.test.ts`, `test/routes/session-routes.test.ts`. + ### Terminal scrollback: strip flavors and wheel/touch forwarding **Two strip flavors, one carry** (#205, `session.ts:_handleTerminalOutput`): the FULL strip (`isAltScreenStripMode` = codex/claude/gemini) removes alt-screen toggles, `3J`, and mouse-tracking DECSETs. Every other mode (shell/opencode/antigravity/pi) gets the NARROW strip (`isMuxAltScreenOnlyStripMode`) — alt-screen toggles ONLY — and only when tmux-backed (`useMux`). Rationale: the tmux CLIENT emits `smcup` as its first bytes at attach, before any program runs, parking xterm in the scrollback-less alternate buffer for the whole session (touch scrolling no-ops; xterm's own wheel handler converts the wheel to Up/Down arrows = readline history cycling — both #205 symptoms). tmux never forwards a pane program's alt-screen toggles to its client (it repaints instead; measured — vim/less inside a pane emit zero to the client), so the only thing the narrow strip ever removes is tmux's own smcup. It keeps `3J` (a user's `clear` is a deliberate scrollback wipe) and the mouse DECSETs (tmux passes those through even with `mouse off`; stripping them would break htop/vim mouse support). ⚠️ The `useMux` gate is load-bearing: `startShell()`/`startInteractive()` fall back to a DIRECT PTY when mux creation fails, and there the inner program's own `?1049h` really does reach xterm — stripping it would break vim/less/htop for real. The replay path (`session-routes.ts`, via `session.usesMux`) applies the same narrow branch; the frontend mirror (`_shouldReportMouseToCli()`) stays claude/codex/gemini because only the FULL strip touches mouse DECSETs. The chunk-boundary carry (`_altScreenSeqCarry`) runs for both flavors. Tests: `test/claude-scrollback-strip.test.ts`. @@ -302,6 +308,14 @@ Invariants: Copy goes through `_copyText()` (Clipboard API, then hidden-textarea + `execCommand`), not raw `navigator.clipboard`, because `install.sh`'s LAN option serves plain HTTP where `navigator.clipboard` is undefined; the fallback steals focus, so the terminal is refocused afterwards. Related: xterm registers its own `copy` listener on the terminal element gated on `hasSelection()`, which is why right-click → Copy has always worked. Selection itself is unavailable on touch devices by design (`user-select: none` on the terminal subtree), and in `shell`/`opencode`/`antigravity` tabs the TUI owns the mouse, so selecting there needs Shift+drag. Tests: `test/terminal-copy-selection.test.ts` (gate + wiring invariants), `test/terminal-copy-shortcut.test.ts` (browser, real key presses). +**The main terminal's four copy paths clean the selection first** (`CodemanCopySelection.clean` in constants.js, pure; `cleanedTerminalSelection()` in terminal-ui.js is the half that reads the live terminal). Those four are the `Ctrl+C` chord, right-click, the phone selection button and Auto Copy. ⚠️ Three routes still copy the RAW padded rows, all of them predating the clean: the browser's own Edit → Copy, which xterm's own `copy` listener on the terminal element serves with `selectionText` directly; a `copy-selection` shortcut the user disabled in App Settings, where nothing calls `preventDefault()` and that native listener runs; and the subagent/teammate windows, which build their own `Terminal` in panels-ui.js with no copy wiring at all. xterm hands back whole screen ROWS and its own trim drops only cells that were never written to, so the real spaces a full-screen TUI paints across the unused part of a row count as content and reach the clipboard. Measured against Claude Code in a 282-column pane, single lines arrived carrying 138 trailing spaces on top of the two-space transcript indent. The clean drops each line's trailing run. Three rules keep it honest: + +1. **Trailing padding only. A shared LEADING indent is deliberately NOT stripped**, and that is a decision rather than an omission: it was built, measured and dropped before #451 merged. It looks like the mirror image of the trailing trim and is not, because no native terminal does it and the transform cannot tell a TUI's margin from content that is genuinely indented. Measured over 401,445 three-row windows across 1,010 tracked files in this repo it fired on **73%** of them (92% inside a YAML workflow, 76% over `git log` output, 48% in a TypeScript source), and no width threshold separates the two because they are the same widths: a live Claude Code pane's own margins measure 2 and 5 columns while the most common non-TUI shared run is 4, sitting between them. The failure modes are what settle it. A wrong trailing trim costs nothing; a wrong dedent silently deletes information that was on screen, with no signal and nothing in the clipboard to hint at it, and it is wrong on `git log` bodies, on indented code read out of `cat` (semantic in Python), on `git diff` context rows where the leading space is the marker, and on stack traces. ⚠️ It also could not be made self-consistent cheaply: whether the first row joined the measurement depended on the mousedown COLUMN, which the user never sees, so one block of three rows produced three different clipboard results, and the flag read `getSelectionPosition().start`, which is the mousedown anchor xterm never normalises, so dragging UP through a block read it off the bottom row (the PR's test stub hardcoded a downward drag, so its suite could not express the case). If it is ever revisited, the one qualification that measured clean is **painted trailing padding** (a full-screen TUI writes real spaces across every row, while a shell pane leaves those cells never-written for xterm to trim): zero false positives over all 401,445 windows, no new plumbing. It still mangles a `git log` body sitting inside an agent's own gutter, which is why it was not taken. +2. **A COLUMN selection is returned untouched.** Alt+drag makes one (xterm's `shouldColumnSelect` keys on `altKey` alone, and neither `Terminal` Codeman builds passes the one option, `macOptionClickForcesSelection`, that would disable it), and a rectangle's rows lining up is the whole point of the gesture. xterm exposes the mode nowhere public, so the check reads `terminal._core._selectionService._activeSelectionMode` (`SelectionMode.COLUMN` is 3) and cleans normally if a future xterm renames it. +3. **The emptiness gate is `trim()`, not truthiness, and it still clears the selection.** A multi-row drag across padding cleans to line breaks alone, which are truthy, and a bare newline pasted into a chat composer submits it. ⚠️ The clear is FEEDBACK, not protection for the interrupt, and the comments that said otherwise were describing the pre-clean code: the `Ctrl+C` gate above now tests the CLEANED selection, so a padding-only selection left set cleans to `''` on every later press and falls through to the PTY as `0x03` anyway. What the clear buys is that a highlight which copied nothing does not linger unexplained, which is also what the toast is for. + +Tests: `test/terminal-copy-clean.test.ts`. + ### Auto Copy (copy-on-select) **Auto Copy** (`autoCopySelection`, per-device, default OFF) puts a finished terminal selection on the clipboard without a keystroke. It is a thin layer over the smart-copy machinery above and shares `_copyText()` with it, but the two paths differ in every decision that matters: @@ -312,6 +326,8 @@ Copy goes through `_copyText()` (Clipboard API, then hidden-textarea + `execComm - **Touch has its own entry point.** `_endTouchSelectionGesture()` and `_selectTouchSelectionLine()` call the flush directly, because the touch path `preventDefault()`s its touchend (that is what stops the compat mouse pair from stealing the selection back), so no mouseup ever reaches the document there. Without those two calls the toggle is simply dead on a phone. - **It must NOT do what `copyTerminalSelection()` does.** That one clears the selection (so a second `Ctrl+C` is an interrupt) and focuses the terminal. Clearing would make text vanish from under the cursor that just highlighted it, and focusing opens the on-screen keyboard over it on a phone. Focus is instead RESTORED to whatever held it before the copy, which only matters for the `execCommand` fallback (it focuses a temp textarea on the way through); the Clipboard API path never moves focus at all. +⚠️ **The toggle is read before the selection is.** `_flushAutoCopySelection()` resolves `_autoCopySelectionEnabled()` first and only then reads and cleans, because Auto Copy is OFF by default and a selection can run to the 50 000-row scrollback ceiling; cleaning ahead of the check would spend that work on every mouseup on the page. The clean runs before `decideAutoCopy()` so its dedupe and its cap both measure the text that actually reaches the clipboard. + `decideAutoCopy()` (constants.js, pure) holds the guards: setting off, blank or whitespace-only text (what a drag across empty cells produces), and a `AUTO_COPY_MAX_CHARS` (1M) cap. ⚠️ The cap is not decoration: a drag off the top of the viewport autoscrolls, so one gesture can sweep the whole 50k-line scrollback. Past it the copy is REFUSED rather than truncated, with a toast pointing at `Ctrl+C`, which still copies everything through the explicit path. ⚠️ **Two dedupe rules, and both earn their place.** A genuine selection change (`pending`) always copies, so re-selecting the same text after copying something else in between still works. Otherwise only text differing from the last auto-copy does, which is what stops an unrelated mouseup from re-copying a stale selection AND what makes the first copy of a drag work at all: xterm fires `onSelectionChange` from its own document `mouseup` handler, and listener order between the two is registration order, not something this code controls. Gating on `pending` alone silently drops that first copy. diff --git a/docs/custom-model-endpoints-plan.md b/docs/custom-model-endpoints-plan.md index 85407adf..5078f17b 100644 --- a/docs/custom-model-endpoints-plan.md +++ b/docs/custom-model-endpoints-plan.md @@ -104,17 +104,17 @@ declared capability, never an `if (mode === 'claude')` branch. ## Per-CLI injection recipes (confidence-ranked) -| CLI | Mechanism | Confidence | -| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude | -| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"":{}}}},"model":"custom/"}` | **Verified by user** | -| `codex` | TOML `config.toml`: top-level `model = ""` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box | -| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done | -| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" | -| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly | -| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working | -| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live | -| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism | +| CLI | Mechanism | Confidence | +| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude | +| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"":{}}}},"model":"custom/"}` | **Verified by user** | +| `codex` | TOML `config.toml`: top-level `model = ""` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol picture more nuanced than a flat break, re-verified live twice on 2026-09-17 against a llama-swap deployment that DOES answer `/v1/responses`** (an earlier test's `Reconnecting...`/`high demand` failure does not reproduce against every llama-swap setup): a plain, no-tool-call chat turn (`codex exec 'reply with just OK'`) returned a real reply. But a real tool-call attempt (`run the shell command: echo hello`) came back as an `agent_message` TEXT item — the tool-call JSON printed as the model's answer, not a `function_call` item codex would actually execute (confirmed via `codex exec --json`'s raw event stream: `item.completed`/`agent_message`, never `function_call`). Since tool execution is what makes codex a coding agent at all, this remains **not usable for real work**, just with a different, more specific failure mode than previously documented — still do not present this as working. Separately, EVERY custom-endpoint codex session also prints `warning: Model metadata for '' not found. Defaulting to fallback metadata...` on launch (confirmed harmless — the successful plain-text reply above still had it): codex's per-model metadata (reasoning tiers, system-prompt templates, context-window figures) comes from `models_cache.json`, a LOCAL CACHE of OpenAI's own hosted model catalog that a custom model can never appear in by construction. No config.toml override exists for it, and the isolated `CODEX_HOME` never gets a `models_cache.json` written into it at all (confirmed: inspected a live, actively-used isolated dir — codex evidently can't reach OpenAI's catalog endpoint for this session and just falls back silently every time, with no file left behind to fix or clean up). Fabricating a fake catalog entry to suppress the warning would mean copying the _shape_ of OpenAI's own proprietary schema — including their real per-model system-prompt content, visible in a genuine `models_cache.json` — for a warning confirmed to have no effect on the actual (broken) tool-calling outcome; not worth building | +| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done | +| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" | +| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly | +| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`), now with `appendV1Suffix: true` (see confidence). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Root cause of the original `HTTP_404` found and fixed, by reading dsh's own bundled source — the same bar pi/grok's fixes were held to.** Installed `@deepseek-ai/dsh` (all its real published dependencies) into a scratch directory purely to read `@deepseek-ai/dsh-llm-deepseek/lib/index.js`: it builds its request as `fetch(\`${connection.baseURL}/chat/completions\`, ...)`with`baseURL`read straight from`DEEPSEEK_BASE_URL`(or defaulting to DeepSeek's real public API root,`https://api.deepseek.com`, which also carries no `/v1`) — no `/v1` insertion of dsh's own, unlike the OpenAI-SDK convention this recipe originally assumed. llama-swap/llama.cpp only ever serves the OpenAI-conventional `/v1/chat/completions`. Confirmed live: `POST /chat/completions` → `404`, `POST /v1/chat/completions` → `200`, on the exact same endpoint — and dsh's own error-message template, `DeepSeek API error (HTTP ${status})`, reproduces the originally reported `dsh: HTTP_404: DeepSeek API error (HTTP 404)` precisely. Fixed by adding `appendV1Suffix` (env kind only, deepseek's entry alone — claude/gemini must NOT get it, since claude was already confirmed working against the unmodified `baseUrl`), which runs `endpoint.baseUrl` through the same `withV1Suffix()` helper `configDir`-kind CLIs already use. ⚠️ Not yet re-run end-to-end with a real `dsh` binary — no install available in this environment (no npm-installed CLI binary in `PATH`, and the `codeman-test-picker` container doesn't bundle it either); the fix is source-confirmed and live-verified at the HTTP level, but a genuine "hello world" reply through `dsh` itself is the remaining step before promoting this to **verified** alongside claude/opencode/pi/grok/omp | +| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live | +| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism | Everything web-researched-but-unverified gets implemented but must be smoke-tested against real installs of those CLIs before being called done — @@ -208,6 +208,14 @@ extra per-model configuration on Codeman's side at all. ### 4. Toolbar UI +> **Superseded.** This section describes the toolbar-button design as originally +> planned. What actually shipped is a Run-menu picker instead: one generated entry +> per (capable harness, saved endpoint) pair directly in the existing `#runModeMenu` +> dropdown, rather than a separate `#customModelBtn`/`#customModelMenu` surface. See +> [`docs/custom-model-endpoints.md`](custom-model-endpoints.md#the-run-menu-picker) +> for the current design; the sections below (session-restart mechanics, security) +> remain accurate regardless of which UI calls the underlying route. + - New header/toolbar button (e.g. `#customModelBtn`, `btn-toolbar btn-custom-model`), marker-hidden by default (`btn-custom-model--hidden`) and revealed by `applyHeaderVisibilitySettings()` only when @@ -344,8 +352,12 @@ pure unit tests and the live manual checks in Verification: up automatically with zero edits to the script). Already run to completion against the author's llama-swap server (a LAN address, inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries): - claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected** - (Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED** + claude/opencode/pi/grok/omp **PASS**, codex **partially works and still + isn't usable** (plain chat succeeds against a llama-swap deployment that + answers `/v1/responses`, but a real tool-call attempt comes back as + inert text rather than an executable `function_call` — see the + confidence table row for the full, re-verified picture), gemini/deepseek + **UNCONFIRMED** (reach the server, fail for undiagnosed reasons — see their table rows), antigravity **SKIP** (no mechanism). Re-run this against a real cloud endpoint (e.g. an Azure AI Foundry deployment) once one is available, to diff --git a/docs/custom-model-endpoints.md b/docs/custom-model-endpoints.md index 975965ef..fecc8f8a 100644 --- a/docs/custom-model-endpoints.md +++ b/docs/custom-model-endpoints.md @@ -11,20 +11,22 @@ company gateway) — anything answering `GET /v1/models` and recipe confidence table, and security reasoning: [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md). -> **Status**: backend is implemented and tested (registry capability, the -> injection engine, the endpoint store + discovery route, the session -> restart route). The toolbar picker / settings UI described below as the -> intended surface is **not yet built** — until it lands, use the HTTP API -> directly (examples below). Antigravity has no known custom-endpoint -> mechanism and is not supported. +> **Status**: fully wired end to end — registry capability, the injection +> engine, the endpoint store + discovery route, both the restart-in-place +> apply route (Claude) and the one-shot quick-start launch path (every +> other supported harness), a settings-panel CRUD surface, and the Run-menu +> picker described below. Antigravity has no known custom-endpoint +> mechanism and is not supported. The HTTP API (examples below) still works +> directly and is what the picker itself calls under the hood. ## Turning it on -App Settings → Agents & CLIs → **Custom Model Endpoints** (synced setting -`customModelEndpointsEnabled`, default **OFF**). Until the toolbar picker -lands, nothing reads this setting: the HTTP routes below work whether it is -on or off, and it exists now only so the picker has a switch to hang off -when it ships. The API equivalent: +App Settings → Models → **Custom model endpoints** (synced setting +`customModelEndpointsEnabled`, default **OFF**). Turning it on does two +things: it reveals the endpoint list/add/edit/discover panel in that same +settings section, and it makes the Run menu offer a generated entry per +(harness, endpoint) pair — see "The Run-menu picker" below. The API +equivalent: ```bash curl -sk -X PUT https://localhost:3000/api/settings \ @@ -34,6 +36,9 @@ curl -sk -X PUT https://localhost:3000/api/settings \ ## Adding an endpoint +Via App Settings → Models → Custom model endpoints → **+ Add endpoint**, or +directly: + ```bash curl -sk -X POST https://localhost:3000/api/model-endpoints \ -H 'Content-Type: application/json' \ @@ -62,7 +67,173 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one. Endpoint management is admin-only in multi-user mode, same as remote/docker hosts — these are machine-level infra, not per-user settings. -## Applying a model to a session +**Context length is discovered too, opportunistically and safely.** The plain +`GET /v1/models` response has no context-window field. Discovery only ever +looks for one for a model llama-swap's own response already reports +`status.value === "loaded"` for — never for an unloaded one, because +llama-swap treats `?model=` as a routing hint and asking about a model that +isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side +effect of what should be read-only discovery. A server with no `status` field +on any entry at all (not llama-swap) gets no context-length enrichment, +rather than guessing. A model's previously-learned context length survives a +later cycle where it wasn't the loaded one; it's dropped only once the model +disappears from the endpoint's list entirely. Stored per model in +`modelContextLengths` and applied automatically (see "Applying a model to a +session" below) so a CLI that would otherwise assume a large default context +window for an unrecognized model id stops silently overflowing a much +smaller real one. + +**Where that number actually comes from matters, and got this wrong once +already.** The first cut read it from llama.cpp's own +`GET /props?model=` (`n_ctx`) — plausible, and it worked in testing, but +confirmed live to be actively WRONG for a `--fit-ctx`-launched llama-swap +backend: `/props` reported `n_ctx: 154112` for a model llama-swap itself had +launched with `--fit-ctx 16384`, and the real server then refused a request +right at that real 16384-token limit — `/props`'s `n_ctx` appears to report +the model's theoretical/trained maximum there, not the runtime-configured +one. Discovery now parses the REAL configured size straight out of +llama-swap's own launch command instead (`GET /running`'s `cmd` field — +`--fit-ctx ` first, then the plain llama.cpp `-c`/`--ctx-size` a +hand-written command might use), and only falls back to the `/props` probe +when `cmd` states no recognizable flag at all. + +**File size is discovered too, when the server states one.** llama-swap +writes a GB figure into an auto-discovered model's own `description` +(`"Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"`), parsed +into `modelSizesGB` — unlike context length, this needs no `/props` probe +(the figure is right there in the `/v1/models` response) and so is populated +for every model regardless of loaded state. A hand-configured profile's own +description has no such figure and correctly gets no entry, never a guess. +Used only to label the Run-menu picker's "loading model" banner (e.g. +"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB) on llama-swap..."); never anything +a server-side check relies on. + +**The loading banner is unbounded by design, and says so — no countdown, no +automatic give-up.** An earlier version scaled an expected-time estimate and +a timeout off the model's file size and auto-closed the session once that +elapsed, but a real load's actual duration depends on hardware this feature +has no way to know (VRAM, storage speed, whatever else is contending for the +GPU) — any fixed number was a guess dressed up as a fact, and a model that +genuinely takes 10+ minutes on slower hardware would just get killed +mid-load by its own display. The banner now says outright that it can take a +while depending on hardware and model size, polls +`GET /api/model-endpoints/:id/running-status` every second for as long as it +takes, and carries a **Cancel** button (rendered on the banner itself) that +ends the wait and closes the session the load was for — the user's own call +on when it's taking too long, not a fixed number baked into the client. + +**The banner's second line is the real backend log line, not a guess.** +llama-swap's `GET /api/events` SSE stream carries the actual `llama-server` +process's own stdout — `load_model: loading model ''`, +`llama_server: model loaded`, tokenizer warnings, all of it — tagged +`source: "upstream"`, distinct from llama-swap's own `source: "proxy"` +request-access lines. `running-status`'s response now includes `logLine` +(via `getLatestLlamaSwapLogLine`), and the banner shows it on its own line +under the disclaimer, e.g. "llama.cpp: load_model: loading model '...'" — +confirmed live end-to-end through a real forced swap, sequentially showing +the model path, a tokenizer warning, then staying on whatever llama.cpp last +printed once the load goes quiet (never cleared back to blank). ⚠️ +**`GET /logs` — the endpoint this feature's own first cut was built +against — turns out to carry ONLY llama-swap's own proxy request-access +log.** Confirmed live it never showed a single backend line, even seconds +after a real, verified model swap; `/api/events`'s `logData` frames are the +only source that actually has it, and its own `source` field (`upstream` vs +`proxy`) is what `getLatestLlamaSwapLogLine` filters on. One `/api/events` +connection is held open per endpoint and reused across every session +watching a load on it (confirmed live to stay open indefinitely, unlike +`/logs`, which closes after a fixed ~100KB), idle-closed after 30s of nobody +polling it (`pruneIdleLlamaSwapLogTails`, same 20s sweep as the +swap-displacement check below). + +`defaultModelId` names which discovered model the picker pre-marks for that +endpoint — the settings panel's Edit form exposes it as a select populated +from the endpoint's own discovered `models`, and the route refuses a value +that isn't one of them. It is applied automatically only when the endpoint +has exactly one discovered model (nothing to choose); with two or more it +is a pre-selection in the model-picker dialog below, never a silent default. +Re-discovering drops a default that no longer appears in the fresh list +rather than carrying an invalid one forward. + +**Model lists refresh themselves.** A background sweep (`server.ts`, +`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every +saved endpoint the same way the manual `POST .../discover-models` route +does, best-effort per endpoint — one being unreachable on a given cycle +never blocks the others. Off under `npm test`, same reasoning as the Codex +plan-usage poll it sits beside: no real network to hit, no server instance +to keep the timer alive for. + +## The Run-menu picker + +With the setting on and at least one endpoint carrying a discovered model, +the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry +per (harness that can redirect to a custom endpoint, saved endpoint) pair, +e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI +registry's own `capabilities.customModelInjection` at page render +(`window.__codemanCustomModelClis`, `server.ts`) — never a hardcoded id list +in the frontend — so a CLI whose injection recipe lands later shows up with +no frontend change, and Antigravity (`unsupported`) never does. + +Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`, +`session-ui.js`) rather than trusting anything cached from the dropdown's +own render — the model list can have changed via the 5-minute sweep above +or a settings-panel edit since the menu opened. With exactly one discovered +model it runs straight away; with two or more, a small modal +(`#customModelPickModal`) lists them and asks which one to use for this +launch, with the endpoint's `defaultModelId` marked but not auto-chosen — +the point of asking is letting one launch deliberately differ from the +saved default, not just confirming it. + +**How the launch itself applies the endpoint depends on the harness.** For +opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP (`runCustomModelEntry` → +`_runCustomModelEntryOneShot`), the endpoint/model is folded into the SAME +`POST /api/quick-start` call that creates the session (`customModel` field), +so the session launches directly on the endpoint — no restart, no visible +relaunch. Claude (`_runCustomModelEntryViaRestart`) still uses the original +two-step design: the launch runs a single native session exactly the way its +own Run-menu entry would, then **waits for the new session to go idle** +(`GET .../wait?until=idle`, bounded at 20s — a normal 200 either way, never +an error, per the wait endpoint's own contract) before applying the endpoint +via the restart route below. That wait exists because a freshly launched CLI +reports itself as `busy` for its own startup (a boot spinner, a +workspace-trust check) well before the apply call would otherwise reach it, +and the apply route correctly refuses to restart a session mid-turn — a +fresh boot looks exactly like one from the outside. A session still busy +after the wait reaches the apply call anyway and gets that route's own +honest `SESSION_BUSY` error, now visible as a sticky toast with a close +button rather than a generic message that vanished in three seconds. Claude +stays on this path because its own restart (`--resume`-based, keeping the +conversation) is far less jarring than the other seven's, and `runClaude()`'s +multi-tab launch and docker-config-drift confirm/retry loop make folding it +into the one-shot path separate work. It is a +one-off "try this endpoint" action, not a sticky mode: the plain Run button +still means "this harness, native cloud" afterward. Entries are hidden +entirely for a remote or Docker active case, since the apply route refuses +both (see the next section). + +## Launching directly on an endpoint (no restart) + +```bash +curl -sk -X POST https://localhost:3000/api/quick-start \ + -H 'Content-Type: application/json' \ + -d '{"caseName": "myapp", "mode": "codex", "customModel": {"endpointId": "llama-box", "modelId": "qwen3"}}' +``` + +`POST /api/quick-start`'s `customModel` field (`{endpointId, modelId, +confirmed?}`) computes the same injection the restart route below does, but +BEFORE the session exists — the session is minted its own id up front +(`crypto.randomUUID()`), the injection (env vars, and for a `configDir`-kind +CLI, the written config file) targets that real id, and the session launches +already pointed at the endpoint. No restart, because there was never a +native-backend launch to restart away from. Runs the same llama-swap +conflict check as the restart route (below) — a `409`-shaped +`{requiresConfirmation, currentlyLoadedModel, affectedSessions}` response +with no session created, resolved by retrying with `confirmedSwap: true` — and +is refused the same way for a remote or Docker case. This is what the +Run-menu picker uses for opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; +Claude still uses the restart route below (see "The Run-menu picker" above +for why). + +## Applying a model to an ALREADY-RUNNING session ```bash curl -sk -X POST https://localhost:3000/api/sessions//custom-model \ @@ -84,6 +255,179 @@ since for those three the config file alone does not switch the model. reattaches the durable remote/in-container tmux rather than relaunching the agent, so the selection would report success and change nothing. +**Claude gets two more env vars when known/applicable, both declared on its +registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:** + +- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context + length (see the discovery section above) whenever one is known. Without + it, Claude Code assumes a large (200k) window for any unrecognized custom + model id and never compacts, which reliably overflows a much smaller real + local context — confirmed live: a stock ~33.7K-token system prompt against + a 16384-token llama-swap model failed with `exceeds the available context +size`. No entry for the model in `modelContextLengths` means the var is + simply omitted, never a guess. ⚠️ **This var only affects when Claude + Code compacts conversation _history_ — it cannot fix a model whose real + context is smaller than Claude Code's own fixed per-turn overhead** + (system prompt + tool schemas, empirically ~36.4K tokens, confirmed live + via an `in:0 out:0` failure on the very first message, before any + history exists to compact). No context-length declaration changes that + fixed overhead, so a model below the safe floor fails outright on + message one regardless of what this var says. See "Context-window floor + warning" below for how Codeman catches this case before launching + instead of after. +- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory + the `configDir`-kind CLIs use (empty, no files written into it), so the + injected `ANTHROPIC_API_KEY` never shares a directory with a stored + claude.ai OAuth login. Claude Code still prints "Both claude.ai and + ANTHROPIC_API_KEY set" when the two coexist in the same config directory — + cosmetic (confirmed live: the API key wins for actual requests either way, + visible in the terminal's own `API Usage Billing` line) but worth + eliminating rather than living with. The directory's `projects` + subdirectory is symlinked (a junction on Windows) back to the real + `~/.claude/projects` so the response viewer, subagent windows and Read My + Mind keep working for that session — the same trade-off and fix documented + for a manually-set `CLAUDE_CONFIG_DIR` in + [`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied + automatically here. Best-effort: a platform that refuses the symlink keeps + the pre-existing blind-response-viewer side effect rather than failing the + whole custom-model apply over it. ⚠️ **This relocates the whole `.claude` + tree, not just transcripts**: a custom-model Claude session also loses the + user's global `settings.json`, user-level skills (the codeman agent skill + included), user-level agents and commands, and the MCP servers configured + in `~/.claude.json` — none of those are symlinked back, only `projects` is. + A fine trade for "point this session at my local llama.cpp," but worth + knowing before it surprises you mid-session. + +**That isolated directory needed one more fix to actually be usable +non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a +real profile's prior "Detected a custom API key — use it?" approvals, so +without more, Claude Code stops and asks that on _every single launch_ — +confirmed live, and with nobody at a TTY to answer, its own default answer +("No") silently refuses the very key this feature just injected, which +looks like the endpoint being ignored entirely. `customModelInjection`'s +`apiKeyTrustFile` (`{ relPath: '.claude.json', shape: +'claude-api-key-responses' }` on claude's entry) pre-seeds that exact +approval: the apply step merges `customApiKeyResponses.approved: [apiKey]` +into `/.claude.json`, the same field a real answered prompt +itself writes to (confirmed against a real file after answering by hand +once) — this answers the prompt in advance rather than bypassing it. The +merge preserves whatever else the CLI already wrote into that file on an +earlier launch in the same isolated directory (`userID`, `numStartups`, +earlier approved keys), and a missing or corrupt file is treated as empty +rather than failing the apply. + +**A fresh `CLAUDE_CONFIG_DIR` isn't just missing that one approval — Claude +Code treats it as a brand-new profile and replays its ENTIRE first-run +sequence on every launch: the theme picker, the security-notes screen, the +per-project "trust this folder?" dialog, and (running with +`--dangerously-skip-permissions`) a one-time warning about bypassing +permissions.** Confirmed live: none of these show up again for a real, +already-onboarded profile, but every custom-model session gets a fresh, +otherwise-empty isolated directory, so it saw all four every single time. +`customModelInjection`'s `skipFirstRunPrompts` (`true` on claude's entry, +requires `apiKeyTrustFile` since it reuses the same file) pre-seeds the +state a real profile accumulates from answering all of that once: +`hasCompletedOnboarding: true` and the launching session's own +`projects[workingDir].hasTrustDialogAccepted: true` go into the same +`/.claude.json` the API-key approval above already merges into +(other projects, and other fields on this session's own project entry, are +left untouched), and `skipDangerousModePermissionPrompt: true` goes into +`/settings.json` — a different file, merged the same +corrupt-tolerant way. `workingDir` is used exactly as the session was +launched with as its cwd, never realpath'd or slash-normalized, since +that's the literal string Claude Code itself uses as the project key. + +**llama-swap gets two more fixes on top of the context-length/config-dir +ones above, both from watching a real switch live.** llama.cpp only ever +runs one model at a time; llama-swap swaps the backing process on demand, +which can take anywhere from a few seconds to well over a minute: + +- **The conflict check.** Both apply routes (the restart one here and the + one-shot `POST /api/quick-start` above) call llama-swap's own + `GET /running` first — feature-detected, so a plain llama.cpp/OpenAI- + compatible server (no such endpoint) is simply never checked. If a + _different_ model is currently loaded and ready, and another **live + session's own selection** is using it, the apply returns + `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` + instead of silently switching — nothing is applied or created yet. + Retrying with `confirmedSwap: true` skips the check (the legacy `confirmed: true` + still means both questions). Switching with nothing + else affected proceeds immediately; this is a warning about disrupting + another session, never a gate on the switch itself. +- **Actually starting the load.** llama-swap has no "switch model" admin + call — the only thing that starts a swap is a real inference request + naming the model, and confirmed live: applying a selection alone never + reached llama-swap at all (nothing in its own server logs), since nothing + had actually asked it to load anything yet. Both apply routes now also + send the smallest real request that will — + `POST /v1/chat/completions` with `max_tokens: 1` and one + throwaway message — whenever the + target model isn't already the one loaded and ready, fire-and-forget (its + response is never read; `GET /api/model-endpoints/:id/running-status`, + polled client-side, is what actually confirms readiness). The response + also carries `modelSwapInProgress: true` in that case, which is what + drives the Run-menu picker's own "loading model" status banner. + +## Catching a swap after the fact + +The conflict check above only runs at the moment a session is created or a +model is applied — it has no way to catch a swap that happens **later**. +Confirmed live: a session created while nothing else conflicted at that +exact instant can still get silently displaced afterward, once a +_different_ session's own normal use (or its own create-time load trigger) +asks llama-swap to load something else. llama-swap has no push +notification of its own for this, so a background sweep +(`detectCustomModelSwapDisplacements`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS` += 20s in `server.ts`) polls `GET /running` once per distinct endpoint that +has at least one live custom-model session, and compares each such +session's own `modelId` against what is actually loaded. A session whose +model is no longer in that list gets a `custom-model:swapped-out` SSE event +(`{sessionId, sessionName, endpointId, previousModel, currentlyLoadedModel}`), +shown as a global toast — global rather than tied to that session's tab, +since the whole point is telling the user before they type into it +expecting the model they picked. Notifies **once per displacement**: the +same de-dupe `Set` clears a session's flag once its own model is loaded and +ready again, so a later, genuinely new displacement notifies again rather +than the session staying silently un-notified forever after the first one. + +## Context-window floor warning + +Claude Code's own fixed per-turn overhead (system prompt + tool schemas, +empirically ~36.4K tokens) can exceed a small local model's _entire_ real +context on its own, before any conversation history exists to fill it — +confirmed live twice, both as an `in:0 out:0` failure on the very first +message sent. `CLAUDE_CODE_MAX_CONTEXT_TOKENS` (above) cannot fix this: it +only governs when Claude Code compacts conversation history, and there is +no history yet on message one. Applying such a model would look like the +endpoint being ignored, or the wrong model being used, when in fact the +endpoint applied correctly and the model is simply too small for this CLI. + +Both apply routes (the restart route and the one-shot `POST +/api/quick-start`) now check for this **before** launching or restarting +anything, gated on the CLI's registry entry declaring a `contextLengthVar` +(currently only claude — the check is a no-op for every other CLI by +construction, never a hardcoded mode check). If the model's discovered +context (`modelContextLengths`, from discovery above) is below +`CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000, comfortably above the measured +~36.4K overhead), the response is `{requiresContextWarning: true, modelId, +contextLength, minSafeContextTokens}` instead of applying — nothing is +restarted or created yet. A context length that was never discovered at +all skips the check entirely (nothing to compare, so it fails open rather +than warning on every model an endpoint hasn't reported a size for). +Retrying with `confirmedContext: true` launches anyway (the legacy `confirmed: true` still means both questions). + +The Run-menu picker shows this as an in-app modal +(`#customModelContextWarningModal`, matching the llama-swap conflict +modal's look) naming the model, its discovered context, and the safe +floor, and explaining the fix: reconfigure llama-swap to give that model +(or a smaller one) an explicit larger context instead of relying on +auto-fit (`--fit-ctx`), which optimizes for the biggest _model_ that fits +rather than the biggest _context_ — e.g. adding `-c 65536` (or as large a +`--ctx-size` as the hardware holds) to that model's llama-swap config +entry. A smaller model at a much larger explicit context often fits in +the same VRAM a bigger model's auto-fit context gets shrunk to make room +for. + Clear back to the harness's native cloud default with: ```bash @@ -101,6 +445,14 @@ id, model and injected key NAMES are persisted, the values are re-derived from the endpoint store on recovery, and the pane keeps running against the endpoint in between because tmux retains its environment. +⚠️ Clearing removes injected keys **by name**, and `CLAUDE_CONFIG_DIR` is one +of the names claude's selection injects — so a session that ALSO had +`CLAUDE_CONFIG_DIR` set through the generic `envOverrides` field (the +per-client-account case) loses that override on clear too, and silently +falls back to the server's default Claude account. If you route a session +to a specific account this way, re-apply the override after clearing a +custom-model selection from it. + **New sessions always default back to the harness's native backend.** A custom-endpoint selection is a per-session choice, never a sticky global default — starting a fresh session doesn't inherit whatever the last one was @@ -115,17 +467,53 @@ automatically). Results: - **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply came back through the endpoint. -- **Codex** — the config is structurally correct, but Codex only speaks the - Responses API since Feb 2026, which llama.cpp/llama-swap don't implement. - This is a real protocol incompatibility, not a bug here; Codex support - needs a Responses-API-compatible endpoint. +- **Codex** — the config is structurally correct, and against a llama-swap + server that DOES answer `/v1/responses` (confirmed live: a plain, + no-tool-call chat turn returned a real reply), the picture is more + nuanced than a flat failure. A real tool-call attempt (`run the shell +command: echo hello`) came back as `agent_message` TEXT — literally the + tool-call JSON printed as the model's answer — instead of a + `function_call` item Codex would actually execute (confirmed via `codex +exec --json`'s raw event stream). So plain chat can work while the thing + that makes Codex a coding agent — actually running commands and editing + files — does not; treat Codex as still unreliable for real work against a + llama.cpp/llama-swap endpoint, tool-calling gap included, not just the + earlier-documented `wire_api` mismatch (which not every deployment hits + the same way — some legitimately have no `/v1/responses` route at all). + Separately, EVERY custom-endpoint Codex session prints `Model metadata +for '' not found. Defaulting to fallback metadata...` on launch — + confirmed harmless (the reply above still came back correctly): Codex's + model metadata (reasoning-tier options, per-model system-prompt + templates, context-window figures) comes from `models_cache.json`, a + local cache of OpenAI's own hosted model catalog that a custom local + model can never appear in by construction, since it isn't one of + OpenAI's models. There's no config.toml override for a model's metadata, + and fabricating a fake catalog entry would mean copying the _shape_ of + OpenAI's own proprietary schema (their per-model system-prompt content + included) for a warning that doesn't otherwise affect behavior — not + something to build into discovery. - **Gemini** — fails with `Invalid auth method selected`, traced to an undocumented `GATEWAY` auth path gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation (several auth workarounds were tried and ruled out); do not rely on Gemini support yet. -- **DeepSeek** — the request reaches the server (env vars are read) but - gets a consistent `HTTP_404`. Root cause not identified; best-effort only. +- **DeepSeek** — root cause of the `HTTP_404` found and fixed. DeepSeek + Harness's own bundled provider module (`@deepseek-ai/dsh-llm-deepseek`) + builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with no + `/v1` insertion of its own (its real public API, `https://api.deepseek.com`, + expects the caller's base URL to already carry any needed prefix) — + confirmed by reading its own source and, live, that + `POST /chat/completions` 404s against llama-swap while + `POST /v1/chat/completions` succeeds; the harness's own error + template (`DeepSeek API error (HTTP ${status})`) matches the originally + reported symptom exactly. `customModelInjection`'s new `appendV1Suffix` + (deepseek's entry only — claude/gemini must NOT get it, since claude was + already confirmed working against the raw `baseUrl`) fixes it by writing + `DEEPSEEK_BASE_URL` with `/v1` appended. Not yet re-run end-to-end with a + real `dsh` binary (no install available in this environment) — the fix + is source-confirmed and live-verified at the HTTP level, but a real + "hello world" reply through `dsh` itself is still outstanding before + calling this fully verified like the harnesses above. - **Antigravity** — no known custom-endpoint mechanism at all; unsupported. See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind diff --git a/docs/remote-sessions.md b/docs/remote-sessions.md index 2c7f6b35..c10c6f07 100644 --- a/docs/remote-sessions.md +++ b/docs/remote-sessions.md @@ -352,6 +352,177 @@ path but the SESSION (`session.remote`): a remote session never falls back to lo `fs`, and a local session never opens an ssh connection — including for attachment records, which are keyed to the session that registered them. +## Wake-on-LAN from user input + +A durable remote session survives an SSH drop (COD-104/108), but nothing brought the +HOST back. When the remote machine suspended, the local pane's `ssh` child **stalled** +rather than exited: `tmux send-keys` SUCCEEDS against a stalled pane, so typed input +vanished with no error anywhere, and without a keepalive the pane could look alive for +the OS TCP timeout. The only recovery was waiting for the reconnect watcher, which +gave up after ~13 minutes and, once exhausted, never retried. + +An **optional** `wakeMac` (one or more MAC addresses, comma-separated) or `wakeCommand` on a +remote host closes that: on user input, `POST /api/sessions/:id/input` probes the host, and if +it is unreachable it wakes it, polls until the host answers, reattaches the pane +(`Session.reattachRemote()`, which idempotently attaches the still-running remote tmux — the +agent conversation is not restarted), and flushes the input that arrived meanwhile. +Implementation: `src/remote-wake.ts`. + +The same wake path also serves **opening** a session, which is where a sleeping host used to +be a dead end: pressing Run on a remote case (`POST /api/quick-start`) or Attach on a +discovered remote tmux session (`POST /api/sessions` + `attachRemoteSession`) probes the host +first, and on a sleeping one wakes it, waits for SSH and only then runs the tmux prereq probe. +Without that the run failed with `could not verify tmux on remote host …` — an ssh error that +blames tmux for a machine that is merely suspended. The wait is **blocking** (the caller gets +the session or the error) but bounded by `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s) rather +than the 90 s session default, because the dashboard sits behind a reverse proxy whose default +`proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while the session was +still being created. The budget covers the whole request, not just the wait (40 s wake + 1.5 s +probe + the tmux prereq probe's own 15 s timeout = 56.5 s worst case). A host with no wake target is not even probed on this path, so nothing +changes for it, and `remote:hostWaking` is broadcast without a `sessionId` (the toast then reads +"the session starts when it is back" — there is no session yet, and no input queued behind it). + +Two wake paths, `wakeCommand` first because it is the explicit override: + +- **`wakeMac`** — Codeman builds the magic packet itself (`buildMagicPacket`, six `0xFF` + bytes then the MAC repeated 16×; the shape is asserted byte-for-byte) and broadcasts it + over UDP port 9 (`sendWakePackets`). This is the normal case: no external script, and one + MAC list per host instead of one per consumer. +- **`wakeCommand`** — a single executable path, run WITHOUT a shell. For hosts that need a + router/another machine to send the packet. + +**UI**: a banner (`#hostWakeBanner`, `host-wake-ui.js`) appears while the ACTIVE remote +session's host is unreachable — amber, since the Codeman session is healthy and only the +machine is asleep. With a wake target the action is **Wake** (`POST /api/sessions/:id/wake`); +with none it is **Configure WoL** and opens `#wakeConfigModal`, a small form for that host's +`wakeMac`/`wakeCommand` that saves with `PUT /api/remote-hosts/:id` (in multi-user mode that +GET is admin-only, so a non-admin is told the setting is admin-only instead of "host not +found"). Reachability for the banner comes from `GET /api/sessions/:id/reachability`: once +when the remote tab is activated (a user action), and every 30 s while the tab is visible +**only for a host with a wake target** — each poll is a TCP connect to the host, and a timer +that connects to a host Codeman could not wake anyway is exactly the timer-driven traffic +the keepalive rule below rejects (it cannot wake a host, but it can keep an activity-based +suspend timer from firing). A host the probe cannot reach (see the next section) is never +polled. ⚠️ The button is pressed from the SAME +dashboard as Run/Attach, so it holds its request open under the same proxy and uses the same +40 s budget — and it **queues nothing**: browser keystrokes travel over the WebSocket, which +deliberately does not pass through the registry (that is the hot path this feature keeps its +hands off), so the banner says "waiting for the host to come back" for the button and only +claims "input is queued" when the HTTP input path actually buffered bytes +(`queuedInput` on the two SSE events). + +**Hosts behind a jump host or SOCKS proxy are reachability-UNKNOWN.** The probe is a bare +TCP connect to `host:port`, and a host reached through `jumpHost`, `socksProxy` or a +`ProxyCommand`/`ProxyJump` in `extraSshOptions` does not answer that even while ssh works — +the direct address may not route at all (the cloudflared case). Acting on the resulting +"unreachable" verdict was wrong three times over: a permanent banner over a healthy session, +a create-path error that replaced a genuine "needs tmux" with "not reachable", and — with a +wake target configured — every HTTP input buffered for the life of the session, because the +readiness poll could never succeed. `isProbeable()` (`remote-wake.ts`) decides from the +proxy fields, which travel on `WakeableRemote`; for such a host the registry delivers input +unchanged, `GET …/reachability` answers `reachable: null, probeable: false` (unknown is not +`false`, and only a proven `false` raises the banner), the create/attach path is not gated +(`ensureHostAwake` → `'unprobeable'`, handled like `'no-target'`), and the quick-start +"not reachable" message is reserved for a **proven** unreachable host (`=== false`). A wake +target can still be fired for it through `POST /api/sessions/:id/wake`, blind: the packet or +command goes out and the response says only whether it did — no readiness poll, no reattach +(the COD-108 watcher owns the pane once ssh works again), no "waking" toast. + +The invariants worth keeping: + +- **Authorization comes before the wake.** In multi-user mode the attach path + (`POST /api/sessions` + `attachRemoteSession`) answers `403` to a non-admin BEFORE the + host is looked up or probed: remote hosts are admin-only infrastructure everywhere else + (the list is `[]` for a non-admin, write and discovery routes are `adminOnly`), and the + wake spawns the host's `wakeCommand` or broadcasts a packet — a gate that came after the + wake handed an unprivileged account a way to run that executable for any configured + `hostId`, hold the request for the wake budget, and only then be refused for the + workingDir. The quick-start path resolves its remote case through `canAccessOwned` + first. Pinned in `test/routes/session-remote-wake.test.ts` (wake spy stays empty). +- **The caller is told what happened to its bytes.** The non-wait input route answers + `{buffered:true}` when the registry took the chunk and `{buffered:true, dropped:true}` + when it was over the cap and is gone; the send-and-wait route answers `OPERATION_FAILED` + when the host never comes back, like the create and attach paths, instead of writing + into the stalled pane and reporting `delivered:true` plus a timeout. Flushed chunks are + written with `fromUser`, so a first prompt that was buffered through a wake can still + name the tab. +- **Only an EXPLICIT request may wake a host:** user input on an established session, the wake + button, or the user's own session create/attach request (`ensureHostAwake`). Everything that + runs on a TIMER must never wake one — the COD-108 watcher, the server's dropped-session + handler, boot recovery and session discovery have no access to the wake registry, and neither + has the shared session service, because `cron-service.ts` builds sessions there with nobody + waiting on the answer; a wake on such a path would re-wake the host seconds after every + suspend, so it could never stay asleep (the same failure `hufflepuff-mcp-lazy` exists to + prevent for MCP keepalives). A reachability check, a discovery listing and the tmux prereq + probe never wake: they are questions, not actions. All of it is enforced by tests in + `test/remote-wake.test.ts` (two wiring guards: one pins the importers — the route module and + `server.ts`, which holds the registry for its LIFETIME only, `drop()` on session cleanup and + `stop()` on shutdown — and one asserts `server.ts` calls nothing but those two, while + `ensureHostAwake` has exactly one caller file) and `test/routes/session-remote-wake.test.ts`, + not by comments. +- **Detection is a bare TCP connect** to the SSH port (then the configured `port`, else 22), + throttled per session, and only for wake-enabled hosts. No `ServerAliveInterval` is added to + the launch command: keepalives push bytes into an otherwise idle connection every interval, + which is exactly what a byte-threshold idle detector must not count as activity. A probe is + ~200 bytes per 30 s, orders of magnitude below any such threshold, and the SYN alone cannot + wake a host. +- **Input is buffered while a wake is in flight** (`REMOTE_WAKE_PENDING_MAX_BYTES`, + oldest whole chunks dropped, bounded so user input cannot grow memory) and flushed in + order after the reattach, with a settle delay so bytes cannot land in a still-connecting + pane. ⚠️ A chunk LARGER than the cap (one big paste is one `input` value) is dropped + **outright**, never trimmed: it was never typed character by character, so its tail is not + "what the user just typed" but a fragment of a command they never sent — the drop is logged + instead. ⚠️ Only the HTTP input route reaches the registry; the **WebSocket keystroke path + is deliberately NOT wake-aware**, so typing into a sleeping host sends nothing and queues + nothing (the banner's Wake button is the recovery for that case, which is why it must not + promise queued input). The **send-and-wait** path blocks on the wake instead — its response + is open anyway, and buffering would break the wait contract. ⚠️ A flush write that FAILS + drops the whole remaining buffer (logged) rather than retaining it: the wake still resolves + and marks the host reachable, so the next input takes the deliver path while a retained + chunk would wait for the NEXT wake — replayed hours later, after everything typed since, + possibly ending in a carriage return. Same policy as the oversized paste. +- **The command runs without a shell** (`spawn(path, [], { stdio: 'ignore' })` — `shell` + defaults to `false`), the schema + requires a single executable path (no arguments, no `$`/backtick), and `wakeMac` is a + structural hex-pair allowlist. A broken or missing wake target fails the wake, never the + input route. +- **`wakeMac`/`wakeCommand` are host-level config, refreshed on recovery AND live** + (`rehydrateRemoteHostFields` in `src/remote-hosts.ts` plus `RemoteWakeDeps.resolveRemote`). + A session's `remote` block is persisted at launch time, so a field added to + `remote-hosts.json` later would otherwise never reach an already-running session — not even + across a Codeman restart, and certainly not right after saving the banner's config dialog. + Recovery rehydration covers restarts, the (throttled, cache-backed) resolver covers the live + session; the host config is authoritative for both (removing the field disables the feature + again). Other host-level fields deliberately stay as persisted, so neither path can + silently re-point an existing pane's SSH options. +- **UI/SSE**: `remote:hostWaking` and `remote:hostWakeFailed` (plus the reused + `remote:sessionReconnected`) drive the banner and toasts, all from `host-wake-ui.js` — + its handlers are the ONLY definitions, since a second one in another mixin would be + silently shadowed by script order. Both carry `queuedInput`, which is true only when the + server actually holds bytes for that session — the wording keys off that, not off "a wake + is running", so the button path never claims input is queued. In multi-user mode the + whole `remote:` family is **session-scoped** (`deriveSseHint`, `server.ts`): an event with + a `sessionId` reaches that session's owner, and the create/attach wake — which has no + session yet — carries the requesting `username` instead (`ensureHostAwake({ requestedBy })`), + since its payload names a `hostId`/`label` that `GET /api/remote-hosts` withholds from + non-admins. With neither, it reaches admins only. +- **No real IO under vitest.** `probeRemoteHostReachable`, `runRemoteWakeCommand` and the + default UDP socket of `sendWakePackets` throw under `VITEST` (as `remote-files.ts` does), + so a test that reaches the defaults fails loudly instead of connecting, spawning or + broadcasting from CI. Every consumer injects its IO (`RemoteWakeDeps`, the socket + factory); `createDefaultRemoteWakeDeps({ probe })` also polls readiness with THAT probe, + which is the leak the guard found. + +Tests: `test/remote-wake.test.ts` (decision/throttle table, single-flight registry, +buffering + flush order, MAC parsing/magic packet, live host-config resolution, the proxied +host, SSE payload routing, the vitest IO guard, and the wiring guard), +`test/routes/session-remote-wake.test.ts` (the input route buffers instead of writing into a +sleeping host — and writes straight into a proxied one —, the reachability route never wakes +and reports a proxied host as unknown, and the wake route reports the no-target case the UI +turns into "configure WoL"), `test/sse-routing-remote.test.ts` (multi-user routing of the +`remote:` family) and `test/host-wake-banner.test.ts` (banner visibility and when the poller +may connect). + ## API Routes are registered in `src/web/routes/case-routes.ts`: @@ -365,6 +536,10 @@ Routes are registered in `src/web/routes/case-routes.ts`: | `GET` | `/api/remote-hosts/:hostId/sessions` | Discover `codeman-*` sessions on the host (COD-105; `listRemoteCodemanSessions`, never errors) | | `POST` | `/api/cases/remote-link` | Link a case to a remote host (creates the `RemoteCase`) | +`RemoteHost` accepts the optional `wakeMac` (magic packet, sent by Codeman) and `wakeCommand` +(single executable path, run without a shell, takes precedence) — see **Wake-on-LAN from user +input** above. + Attaching to a discovered session is a **session-create** path, not a host route: `POST /api/sessions` accepts `attachRemoteSession: { hostId, remoteSessionName }` (schema in `schemas.ts`; `remoteSessionName` must match `^codeman-[a-zA-Z0-9._-]+$`), diff --git a/docs/wiki/Agent-CLIs.md b/docs/wiki/Agent-CLIs.md index eb013b29..d7ad4b22 100644 --- a/docs/wiki/Agent-CLIs.md +++ b/docs/wiki/Agent-CLIs.md @@ -273,9 +273,16 @@ into the case's `.claude/settings.local.json` so that `/model` keeps working. - **Shell** for the times you want a terminal on your phone with no agent at all. It is a genuinely useful mode, not a fallback. +## Pointing one at your own server + +Most of these harnesses can also run against a custom OpenAI-compatible endpoint instead of +their native cloud backend, for one session at a time, an opt-in feature covered in full on +[Custom Model Endpoints](Custom-Model-Endpoints). + ## Read next - [Core Concepts](Core-Concepts) - run modes versus location overlays. +- [Custom Model Endpoints](Custom-Model-Endpoints) - run a harness against your own server. - [Settings Reference](Settings-Reference) - model, effort, and permission-mode settings. - [Keeping Agents Running](Keeping-Agents-Running) - what idle detection does per mode. - [Security](Security) - what skipping permission prompts actually means. diff --git a/docs/wiki/Custom-Model-Endpoints.md b/docs/wiki/Custom-Model-Endpoints.md new file mode 100644 index 00000000..ebe05fee --- /dev/null +++ b/docs/wiki/Custom-Model-Endpoints.md @@ -0,0 +1,183 @@ +# Custom Model Endpoints + +Point a harness at your own OpenAI-compatible server instead of its native cloud backend, for +one session at a time. "Custom endpoint" covers **local** hardware (llama.cpp, Ollama, vLLM, +a home GPU rig, DGX Spark, Strix Halo) and **cloud** services (Azure AI Foundry's +OpenAI-compatible endpoint, OpenRouter, a company gateway) alike, anything answering +`GET /v1/models` and `POST /v1/chat/completions` in the standard shape. + +**Off by default.** Turn it on in App Settings → Models → **Custom model endpoints**. + +## Adding an endpoint + +Still in App Settings → Models → Custom model endpoints: + +1. **+ Add endpoint** — give it an id, a label, and the base URL (`http://192.168.1.50:8080`, + say). An API key is optional; most local servers don't check one. +2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it. +3. Pick a **default model** from what was discovered. This is the model the Run-menu entry + applies directly when only one model is discovered; with two or more, it's just the one + pre-marked in the picker dialog described below, not a silent default. + +Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker +hosts — these are machine-level infra, not a per-user setting. + +**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every +5 minutes in the background, so a model the server starts serving later — or stops serving — +shows up without another manual click of **Discover**. One endpoint being unreachable on a +given cycle (powered off, wrong network) never blocks the others from refreshing. + +**Context length is picked up automatically where it can be, safely.** Against a +llama.cpp/llama-swap server, discovery also learns each _currently loaded_ model's real +context window and applies it to the launched session (Claude Code today — see below), so +the harness stops assuming a large default window for a model name it doesn't recognise and +overflowing a much smaller real one. It's deliberately never probed for a model that isn't +already loaded, since asking a llama-swap server about an unloaded model can trigger an +actual, slow model swap as a side effect — a model just not currently loaded keeps whatever +context length an earlier cycle already learned for it instead. + +## Running a session against one + +With the setting on and at least one endpoint carrying a discovered model, the **Run** +dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a +custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a +session on that harness exactly the way its own entry would. It is a one-off "try this +endpoint" action, not a sticky mode — the plain **Run** button still means "this harness, +native cloud" afterward, and a fresh session never inherits whatever the last one was +pointed at. + +**Which model it uses depends on how many the endpoint has discovered.** With exactly one, +the session launches straight away on that model — nothing to choose. With two or more, a +small dialog asks which one to use for this launch before starting the session; the +endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately +use a different one without changing the saved default. + +**For opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP, picking an entry launches +straight onto the endpoint** — no restart, because the endpoint is applied before the +session's process ever starts. **Claude still restarts the harness's process in place** — +same tab, same conversation (`--resume`) — after a normal native launch, since that restart +is far less jarring for Claude than for the other seven, whose own TUI can fully +reinitialize on a restart. Either way, every supported harness reads its endpoint config at +process start, never per turn, so there is no live hot-swap while a turn is running. + +Picking an entry that launches a **brand-new** Claude session waits (up to 20 seconds) for it to +finish its own startup before applying — a freshly started CLI reports itself as busy for its +boot sequence, and applying to a genuinely busy session is refused so a real, in-progress +turn is never interrupted out from under you. A session that is still busy after that wait +(a very slow-starting CLI, or one you started typing into right away) surfaces that refusal +as an ordinary error, which now stays on screen with a close button instead of vanishing +after a few seconds — read it, it names the actual reason rather than a generic failure. + +Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for +redirecting those hasn't landed yet, see below. The picker also only appears in the desktop +**Run** dropdown; the phone home screen builds its own run picker separately and does not +currently offer these entries. + +**Against llama-swap, applying a selection also starts the actual model load, rather than +waiting on your first prompt to do it.** llama-swap has no "switch model" button of its own +— the only thing that starts a swap is a real request naming the model, and confirmed live: +just applying a selection never reached llama-swap's own logs at all until something asked +it to load. Picking an entry now also sends the smallest real request that will trigger +that load, in the background, the moment the target model isn't already loaded and ready. + +**The centred loading banner has no countdown and no automatic timeout — it waits as long as +it takes, and tells you so.** When it knows the model's discovered file size (its GB figure, +when llama-swap states one) it's shown too, e.g. "Loading qwen3.8-27b (16.4 GB) on +llama-swap — this can take a while depending on your hardware and the model size." An +earlier version tried to estimate and enforce a time limit, but real load time depends on +hardware this feature has no way to know, so a fixed number was always a guess — worse, one +that could kill a genuinely slow load partway through. If it really is taking too long, a +**Cancel** button right on the banner ends the wait and **closes the session that load was +for**, on your own call rather than a guessed deadline. + +**The banner also shows a real, live second line of what llama.cpp itself is doing** — not +a made-up progress phase, the actual next line the `llama-server` process printed, e.g. +"llama.cpp: load_model: loading model '/models/.../Qwen3.8-27B.gguf'" then later +"llama.cpp: llama_server: model loaded". It comes straight from llama-swap's own event +feed, filtered down to just the backend process's own output (not llama-swap's own request +logging), and stays on whatever it last said once the load goes quiet, rather than +clearing back to nothing. + +**You'll also be told if a session's model gets swapped out from under it later, not just +at launch.** The conflict warning above only fires at the moment you launch or apply a +model — llama.cpp only runs one model at a time, so if a DIFFERENT session using the same +endpoint later triggers its own load, whatever was loaded before (including a session you +already had running) gets silently evicted, with no warning at that instant since nothing +conflicted when it was first set up. A background check (every 20 seconds) catches this +after the fact and shows a toast naming which session lost its model and what's loaded now +— so you know before typing into that session that it's about to reload (and, in turn, +evict whatever displaced it). + +**Claude Code specifically gets three extra fixes applied automatically:** + +- Its discovered context length (see above) is passed through as + `CLAUDE_CODE_MAX_CONTEXT_TOKENS`, so it doesn't send a full-size prompt against a much + smaller real local context and overflow it. +- Its session runs with an isolated `CLAUDE_CONFIG_DIR`, so the injected API key never sits + in the same directory as a stored claude.ai login — that combination is harmless for actual + requests (the API key wins) but the CLI still prints a "both claude.ai and + ANTHROPIC_API_KEY set" warning about it, which this avoids entirely. The isolated directory + keeps a link back to your real session history so the response viewer and similar features + still work for that session. That isolated directory starts with no prior approvals of its + own, so Codeman also pre-approves the injected key the same way answering Claude Code's own + "Detected a custom API key" prompt once would — without it, that prompt would otherwise + reappear on every single launch with nobody there to answer it. +- **That same fresh isolated directory also looks like a brand-new Claude Code profile**, so + without this fix it replayed the WHOLE first-run sequence every single launch: the theme + picker, the security-notes screen, the "trust this folder?" dialog, and a one-time warning + about running with permissions bypassed — none of which a real, already-used profile shows + again. Codeman now pre-seeds that same "already been through this once" state (onboarding + completed, this session's own project marked trusted, the bypass-permissions warning + acknowledged) so a custom-model launch reaches the actual conversation exactly as fast as a + native cloud one does, instead of stopping at a wizard with nobody there to click through it. + +**If a model's real context is too small for Claude Code to even get started, you get a +warning instead of a confusing failure.** Claude Code's own system prompt and tools take up +roughly 40K tokens on their own, before you've typed anything — a small local model with a +smaller real context than that fails outright on the very first message, no matter what +context size Codeman tells it to expect (raising the declared context only changes when +Claude Code trims _conversation history_, and there is none yet on message one). Picking +such a model now shows an in-app dialog naming the model, its discovered context and what's +needed, before anything launches or restarts, with the fix spelled out: reconfigure +llama-swap to give that model (or a smaller one) an explicit larger context instead of +relying on auto-fit (`--fit-ctx`), which sizes the context around fitting the biggest model +rather than the biggest context — for example adding `-c 65536` to that model's llama-swap +entry. "Launch anyway" is still there if you want to try regardless. + +## Which harnesses actually work + +| Harness | Status | +| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. | +| **Codex** | Config is correct, and plain chat can work against a server that speaks the Responses API — but a real tool-call attempt comes back as inert text instead of running, so it's still not usable for real coding work. | +| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. | +| **DeepSeek** | The original 404 is root-caused and fixed (DeepSeek Harness's own code was missing a `/v1` most local servers require) — not yet re-run against a real `dsh` install to confirm end-to-end. | +| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. | + +Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry, +not a fixed list here, so this table can go stale before this page does — a greyed-out or +missing entry is the more current answer. + +## What it does not do + +- **No remote or Docker sessions yet.** Both restart their agent differently under the hood + (reattaching a durable tmux session rather than relaunching the process), so redirecting + them needs its own plumbing that hasn't been built. +- **No live hot-swap mid-conversation.** Applying a selection always restarts the process. +- **No button to un-point a session from the UI yet.** Clearing back to native cloud is an + HTTP call (`POST .../custom-model {"clear": true}`) or deleting the session; the settings + panel manages saved endpoints, not what a running session is currently pointed at. +- **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any, + never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate, + explicit choice per session. + +## Security + +An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save +time and against the address it actually resolves to), the same guard Web Tabs uses for +saved dashboards. Endpoint records and any per-session config files a harness needs are +written with owner-only permissions. See +[custom-model-endpoints-plan.md](https://github.com/Ark0N/Codeman/blob/master/docs/custom-model-endpoints-plan.md) +in the repository for the full design reasoning, including why this feature closed a +pre-existing gap in how session environment overrides were guarded rather than opening a new +one. diff --git a/docs/wiki/Keyboard-Shortcuts.md b/docs/wiki/Keyboard-Shortcuts.md index bfa32d5d..098f7e05 100644 --- a/docs/wiki/Keyboard-Shortcuts.md +++ b/docs/wiki/Keyboard-Shortcuts.md @@ -34,6 +34,8 @@ Press `Ctrl+?` in the app for the same list in a floating overlay. | Right-click | Copy the selection. With nothing selected the native menu is left alone. | | `Ctrl+Z` | Swallowed in agent sessions so a running CLI cannot be suspended. Normal job control in a shell. | +Anything you copy is cleaned on the way to the clipboard: each line loses the padding spaces a full-screen program paints across the rest of the row. Leading indentation is left exactly as it is, so indented code, a `git log` message body and `git diff` context lines paste back the way they looked on screen. An `Alt+drag` rectangular selection is copied exactly as it looks, so its columns stay lined up. + ## Everything else | Shortcut | Action | diff --git a/docs/wiki/Settings-Reference.md b/docs/wiki/Settings-Reference.md index 3255f56f..14b04056 100644 --- a/docs/wiki/Settings-Reference.md +++ b/docs/wiki/Settings-Reference.md @@ -92,6 +92,10 @@ Model and effort are both **soft defaults**: the model is written into the case' `.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort` inside a session override them at any time. +**Custom model endpoints** (off by default) adds a saved-endpoint list plus a matching +section to the Run dropdown, for pointing a harness at your own OpenAI-compatible server +instead of its native cloud backend. See [Custom Model Endpoints](Custom-Model-Endpoints). + ### Agents & CLIs | Setting | Notes | diff --git a/docs/wiki/_Sidebar.md b/docs/wiki/_Sidebar.md index af84a756..2e6d54b5 100644 --- a/docs/wiki/_Sidebar.md +++ b/docs/wiki/_Sidebar.md @@ -12,6 +12,7 @@ - [The Dashboard](The-Dashboard) - [Agent CLIs](Agent-CLIs) +- [Custom Model Endpoints](Custom-Model-Endpoints) - [Working With Files](Working-With-Files) - [Input And Voice](Input-And-Voice) - [Mobile Guide](Mobile-Guide) diff --git a/package-lock.json b/package-lock.json index f9640246..1263ccb9 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "aicodeman", - "version": "1.30.0", + "version": "1.31.0", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "aicodeman", - "version": "1.30.0", + "version": "1.31.0", "hasInstallScript": true, "license": "MIT", "workspaces": [ diff --git a/package.json b/package.json index 48a185f6..46f130ce 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "aicodeman", - "version": "1.30.0", + "version": "1.31.0", "description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence", "type": "module", "main": "dist/index.js", diff --git a/plugins/codeman/.claude-plugin/plugin.json b/plugins/codeman/.claude-plugin/plugin.json index c2def406..841cc8c2 100644 --- a/plugins/codeman/.claude-plugin/plugin.json +++ b/plugins/codeman/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "codeman", "description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.", - "version": "1.30.0", + "version": "1.31.0", "author": { "name": "Ark0N", "url": "https://github.com/Ark0N" diff --git a/plugins/codeman/skills/codeman/SKILL.md b/plugins/codeman/skills/codeman/SKILL.md index 9a4ae5f3..c2aac7da 100644 --- a/plugins/codeman/skills/codeman/SKILL.md +++ b/plugins/codeman/skills/codeman/SKILL.md @@ -47,7 +47,7 @@ later call opens with, and your first REAL call performs them anyway: ```bash . "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null -[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; } +[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; } ``` ⚠️ **Never spend a Bash call on this check alone.** §1's block opens with this same @@ -75,8 +75,8 @@ PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" mkdir -p "$(dirname "$PRE")" # Rewrite unless the file already ends with THIS version's stamp, so a stale or a # half-written file self-heals here instead of costing you a round trip to rm it. -grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE' -# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ---- +grep -qs '^CODEMAN_PREAMBLE=1.30.1$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE' +# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ---- API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}" SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}" # Credentials, cheapest first. Your session has usually INHERITED the server's @@ -150,6 +150,27 @@ _trust_key() { # -> "confirm" | "move" | "" (nothing safe to press) | tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \ | sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/' } +# ---- the composer: is the prompt still sitting there, unsent? ---- +# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the +# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that +# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt +# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through +# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS +# the composer and keeps pressing Enter while the prompt is still there. +_composer_text() { # -> the composer's text with ALL whitespace removed: "" once + # the prompt was taken, "?" when the pane shows no composer at all. The composer is + # the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher + # up in the transcript, so only the last one says whether the text was taken. + local t + t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \ + | jq -r '.data.terminalBuffer // empty' \ + | sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \ + | tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1) + [ -n "$t" ] || { printf '?'; return 0; } + # Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does + # not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH). + printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" +} _accept_trust() { # -> 0 once it has answered the dialog, 1 if it could not local sid="$1" k i=1 while [ "$i" -le 6 ]; do @@ -262,12 +283,16 @@ spawn_workers() { # worker a silent no-op that still "succeeds" and reports the previous turn's state. # Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a # deliberate duplicate, at the SAME number (§5.3). -# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the -# typed prompt stranded on the composer while a long wait runs its whole timeout -# (observed live). So the first wait is short; on its timeout a bare \r goes out (the -# missing Enter when the prompt is stranded, a no-op when the turn is genuinely -# running), then the ORIGINAL frame is resent unchanged, which the server takes as a -# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker +# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude +# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints), +# leaving the typed prompt stranded on the composer while a long wait runs its whole +# timeout (observed live, twelve reviews in a row). So the first wait is short; on its +# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate: +# the server re-waits without retyping, §5.3) and kept open in the background, while +# the composer is READ (_composer_text) and, as long as the prompt is still sitting +# there, a bare \r goes out about every ten seconds, up to twelve times. An empty +# composer ends the loop, so a prompt that was taken is never nudged again, and the +# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker # spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) -- # and for those only. Hook-less workspaces and the other modes resolve on flapping # idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not @@ -275,7 +300,7 @@ spawn_workers() { # it accepts the send and then burns both waits. One timeout on a dsh worker whose # pane clearly finished means that profile, so switch that worker to markers. sendwait() { - local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r + local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i # `wait:"stop,exit"`, never the `wait:true` default set: that set also carries # `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a # dsh worker whose TUI repaints rarely the session reads `idle` while the model @@ -290,16 +315,38 @@ sendwait() { r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ -H 'Content-Type: application/json' --data-binary "$body") if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then - "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \ - -d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \ - '{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null - # The resend is a tagged DUPLICATE, so the server skips the write and reports - # `delivered:false` for it -- truthfully, but about the wrong send. The first - # one delivered, so carry that forward, or §1's cleanup reads a completed turn - # as an undelivered one and keeps a finished worker forever. - r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ - -H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \ - | jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end') + # ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call, + # in the background, while the Enter loop below works the composer. Signals have + # no history: a `stop` that fires while no wait is open (during a composer read + # between two short waits, measured) is lost, and the next wait then runs its + # whole timeout on a turn that already ended. The resend is a tagged DUPLICATE, + # so the server skips the write and re-waits without retyping (§5.3). + tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1 + "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ + -H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" & + bg=$! + # The prompt's head with whitespace removed, matched literally (the "$head" + # quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob). + head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24) + while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended + c=$(_composer_text "$sid") + if [ "$c" = '?' ]; then + [ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it + elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then + break # composer empty (taken) or holding other text + fi + n=$((n+1)) + "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \ + -d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \ + '{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null + i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done + done + wait "$bg" + # The duplicate reports `delivered:false` -- truthfully, but about the wrong send. + # The first one delivered, so carry that forward, or §1's cleanup reads a completed + # turn as an undelivered one and keeps a finished worker forever. + r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp") + rm -f "$tmp" fi printf '%s\n' "$r" } @@ -325,10 +372,10 @@ last_text() { # The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept # bare on purpose: the write condition above anchors on it with $, so an inline comment # here would fail that match and rewrite this file on every single bootstrap. -CODEMAN_PREAMBLE=1.22.0 +CODEMAN_PREAMBLE=1.30.1 PREAMBLE ) -. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; } +. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; } ``` Every later Bash call that touches the API starts with the same two loader lines from @@ -379,7 +426,7 @@ and no per-call body to hand-build. ```bash . "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader -[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; } +[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; } N=(alpha beta) # INVENT one fresh case name per worker; never list cases first # (a name may carry a mode: `beta:deepseek`, see below) T=('reply with one line: the absolute path of your working directory' @@ -440,9 +487,11 @@ Four things this block leans on, each one link away, no detour needed to run it: skill: §5.1. Those workspaces do get hooks now, unless the operator disabled it. - `sendwait` supplies the `\r`, picks a fresh `seq`, and self-heals a stranded Enter. A prompt without the `\r` is never submitted (§3), a reused `seq` is silently - swallowed as an already-applied duplicate, and an Enter eaten by an Ink repaint - strands the prompt on the composer until a bare `\r` follows: all three are reasons - to let `sendwait` build the call rather than hand-rolling it. + swallowed as an already-applied duplicate, and a lost Enter strands the prompt on the + composer until a bare `\r` follows: Claude Code 2.1.277 and later ignore Enter for the + first 30 to 50 seconds after the composer paints while still taking the text, so + `sendwait` reads the composer and keeps pressing Enter until the prompt has left it. + All three are reasons to let `sendwait` build the call rather than hand-rolling it. - Each `sendwait` costs that worker one billed turn, as does every prompt you send it. - Deleting the sessions does **not** remove the case directories. They are marked as agent-created, so `GET /api/v1/cases/agent-created` lists them for cleanup: §5.14. diff --git a/plugins/codeman/skills/codeman/preamble.sh b/plugins/codeman/skills/codeman/preamble.sh index cadc4a78..11982e34 100644 --- a/plugins/codeman/skills/codeman/preamble.sh +++ b/plugins/codeman/skills/codeman/preamble.sh @@ -1,4 +1,4 @@ -# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ---- +# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ---- API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}" SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}" # Credentials, cheapest first. Your session has usually INHERITED the server's @@ -72,6 +72,27 @@ _trust_key() { # -> "confirm" | "move" | "" (nothing safe to press) | tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \ | sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/' } +# ---- the composer: is the prompt still sitting there, unsent? ---- +# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the +# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that +# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt +# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through +# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS +# the composer and keeps pressing Enter while the prompt is still there. +_composer_text() { # -> the composer's text with ALL whitespace removed: "" once + # the prompt was taken, "?" when the pane shows no composer at all. The composer is + # the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher + # up in the transcript, so only the last one says whether the text was taken. + local t + t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \ + | jq -r '.data.terminalBuffer // empty' \ + | sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \ + | tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1) + [ -n "$t" ] || { printf '?'; return 0; } + # Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does + # not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH). + printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" +} _accept_trust() { # -> 0 once it has answered the dialog, 1 if it could not local sid="$1" k i=1 while [ "$i" -le 6 ]; do @@ -184,12 +205,16 @@ spawn_workers() { # worker a silent no-op that still "succeeds" and reports the previous turn's state. # Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a # deliberate duplicate, at the SAME number (§5.3). -# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the -# typed prompt stranded on the composer while a long wait runs its whole timeout -# (observed live). So the first wait is short; on its timeout a bare \r goes out (the -# missing Enter when the prompt is stranded, a no-op when the turn is genuinely -# running), then the ORIGINAL frame is resent unchanged, which the server takes as a -# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker +# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude +# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints), +# leaving the typed prompt stranded on the composer while a long wait runs its whole +# timeout (observed live, twelve reviews in a row). So the first wait is short; on its +# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate: +# the server re-waits without retyping, §5.3) and kept open in the background, while +# the composer is READ (_composer_text) and, as long as the prompt is still sitting +# there, a bare \r goes out about every ten seconds, up to twelve times. An empty +# composer ends the loop, so a prompt that was taken is never nudged again, and the +# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker # spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) -- # and for those only. Hook-less workspaces and the other modes resolve on flapping # idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not @@ -197,7 +222,7 @@ spawn_workers() { # it accepts the send and then burns both waits. One timeout on a dsh worker whose # pane clearly finished means that profile, so switch that worker to markers. sendwait() { - local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r + local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i # `wait:"stop,exit"`, never the `wait:true` default set: that set also carries # `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a # dsh worker whose TUI repaints rarely the session reads `idle` while the model @@ -212,16 +237,38 @@ sendwait() { r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ -H 'Content-Type: application/json' --data-binary "$body") if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then - "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \ - -d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \ - '{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null - # The resend is a tagged DUPLICATE, so the server skips the write and reports - # `delivered:false` for it -- truthfully, but about the wrong send. The first - # one delivered, so carry that forward, or §1's cleanup reads a completed turn - # as an undelivered one and keeps a finished worker forever. - r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ - -H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \ - | jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end') + # ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call, + # in the background, while the Enter loop below works the composer. Signals have + # no history: a `stop` that fires while no wait is open (during a composer read + # between two short waits, measured) is lost, and the next wait then runs its + # whole timeout on a turn that already ended. The resend is a tagged DUPLICATE, + # so the server skips the write and re-waits without retyping (§5.3). + tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1 + "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ + -H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" & + bg=$! + # The prompt's head with whitespace removed, matched literally (the "$head" + # quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob). + head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24) + while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended + c=$(_composer_text "$sid") + if [ "$c" = '?' ]; then + [ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it + elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then + break # composer empty (taken) or holding other text + fi + n=$((n+1)) + "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \ + -d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \ + '{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null + i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done + done + wait "$bg" + # The duplicate reports `delivered:false` -- truthfully, but about the wrong send. + # The first one delivered, so carry that forward, or §1's cleanup reads a completed + # turn as an undelivered one and keeps a finished worker forever. + r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp") + rm -f "$tmp" fi printf '%s\n' "$r" } @@ -247,4 +294,4 @@ last_text() { # The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept # bare on purpose: the write condition above anchors on it with $, so an inline comment # here would fail that match and rewrite this file on every single bootstrap. -CODEMAN_PREAMBLE=1.22.0 +CODEMAN_PREAMBLE=1.30.1 diff --git a/plugins/codeman/skills/codeman/reference/recipes.md b/plugins/codeman/skills/codeman/reference/recipes.md index 94e0264f..af814653 100644 --- a/plugins/codeman/skills/codeman/reference/recipes.md +++ b/plugins/codeman/skills/codeman/reference/recipes.md @@ -21,7 +21,7 @@ by sourcing the preamble file the §0 bootstrap wrote, and checking its version ```bash . "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null -[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; } +[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; } ``` Do **not** re-paste the preamble body into each call. Sourcing it is what retires the diff --git a/skills/codeman/SKILL.md b/skills/codeman/SKILL.md index 9a4ae5f3..c2aac7da 100644 --- a/skills/codeman/SKILL.md +++ b/skills/codeman/SKILL.md @@ -47,7 +47,7 @@ later call opens with, and your first REAL call performs them anyway: ```bash . "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null -[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; } +[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; } ``` ⚠️ **Never spend a Bash call on this check alone.** §1's block opens with this same @@ -75,8 +75,8 @@ PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" mkdir -p "$(dirname "$PRE")" # Rewrite unless the file already ends with THIS version's stamp, so a stale or a # half-written file self-heals here instead of costing you a round trip to rm it. -grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE' -# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ---- +grep -qs '^CODEMAN_PREAMBLE=1.30.1$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE' +# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ---- API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}" SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}" # Credentials, cheapest first. Your session has usually INHERITED the server's @@ -150,6 +150,27 @@ _trust_key() { # -> "confirm" | "move" | "" (nothing safe to press) | tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \ | sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/' } +# ---- the composer: is the prompt still sitting there, unsent? ---- +# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the +# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that +# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt +# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through +# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS +# the composer and keeps pressing Enter while the prompt is still there. +_composer_text() { # -> the composer's text with ALL whitespace removed: "" once + # the prompt was taken, "?" when the pane shows no composer at all. The composer is + # the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher + # up in the transcript, so only the last one says whether the text was taken. + local t + t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \ + | jq -r '.data.terminalBuffer // empty' \ + | sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \ + | tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1) + [ -n "$t" ] || { printf '?'; return 0; } + # Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does + # not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH). + printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" +} _accept_trust() { # -> 0 once it has answered the dialog, 1 if it could not local sid="$1" k i=1 while [ "$i" -le 6 ]; do @@ -262,12 +283,16 @@ spawn_workers() { # worker a silent no-op that still "succeeds" and reports the previous turn's state. # Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a # deliberate duplicate, at the SAME number (§5.3). -# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the -# typed prompt stranded on the composer while a long wait runs its whole timeout -# (observed live). So the first wait is short; on its timeout a bare \r goes out (the -# missing Enter when the prompt is stranded, a no-op when the turn is genuinely -# running), then the ORIGINAL frame is resent unchanged, which the server takes as a -# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker +# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude +# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints), +# leaving the typed prompt stranded on the composer while a long wait runs its whole +# timeout (observed live, twelve reviews in a row). So the first wait is short; on its +# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate: +# the server re-waits without retyping, §5.3) and kept open in the background, while +# the composer is READ (_composer_text) and, as long as the prompt is still sitting +# there, a bare \r goes out about every ten seconds, up to twelve times. An empty +# composer ends the loop, so a prompt that was taken is never nudged again, and the +# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker # spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) -- # and for those only. Hook-less workspaces and the other modes resolve on flapping # idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not @@ -275,7 +300,7 @@ spawn_workers() { # it accepts the send and then burns both waits. One timeout on a dsh worker whose # pane clearly finished means that profile, so switch that worker to markers. sendwait() { - local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r + local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i # `wait:"stop,exit"`, never the `wait:true` default set: that set also carries # `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a # dsh worker whose TUI repaints rarely the session reads `idle` while the model @@ -290,16 +315,38 @@ sendwait() { r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ -H 'Content-Type: application/json' --data-binary "$body") if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then - "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \ - -d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \ - '{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null - # The resend is a tagged DUPLICATE, so the server skips the write and reports - # `delivered:false` for it -- truthfully, but about the wrong send. The first - # one delivered, so carry that forward, or §1's cleanup reads a completed turn - # as an undelivered one and keeps a finished worker forever. - r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ - -H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \ - | jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end') + # ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call, + # in the background, while the Enter loop below works the composer. Signals have + # no history: a `stop` that fires while no wait is open (during a composer read + # between two short waits, measured) is lost, and the next wait then runs its + # whole timeout on a turn that already ended. The resend is a tagged DUPLICATE, + # so the server skips the write and re-waits without retyping (§5.3). + tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1 + "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ + -H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" & + bg=$! + # The prompt's head with whitespace removed, matched literally (the "$head" + # quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob). + head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24) + while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended + c=$(_composer_text "$sid") + if [ "$c" = '?' ]; then + [ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it + elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then + break # composer empty (taken) or holding other text + fi + n=$((n+1)) + "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \ + -d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \ + '{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null + i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done + done + wait "$bg" + # The duplicate reports `delivered:false` -- truthfully, but about the wrong send. + # The first one delivered, so carry that forward, or §1's cleanup reads a completed + # turn as an undelivered one and keeps a finished worker forever. + r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp") + rm -f "$tmp" fi printf '%s\n' "$r" } @@ -325,10 +372,10 @@ last_text() { # The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept # bare on purpose: the write condition above anchors on it with $, so an inline comment # here would fail that match and rewrite this file on every single bootstrap. -CODEMAN_PREAMBLE=1.22.0 +CODEMAN_PREAMBLE=1.30.1 PREAMBLE ) -. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; } +. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; } ``` Every later Bash call that touches the API starts with the same two loader lines from @@ -379,7 +426,7 @@ and no per-call body to hand-build. ```bash . "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader -[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; } +[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; } N=(alpha beta) # INVENT one fresh case name per worker; never list cases first # (a name may carry a mode: `beta:deepseek`, see below) T=('reply with one line: the absolute path of your working directory' @@ -440,9 +487,11 @@ Four things this block leans on, each one link away, no detour needed to run it: skill: §5.1. Those workspaces do get hooks now, unless the operator disabled it. - `sendwait` supplies the `\r`, picks a fresh `seq`, and self-heals a stranded Enter. A prompt without the `\r` is never submitted (§3), a reused `seq` is silently - swallowed as an already-applied duplicate, and an Enter eaten by an Ink repaint - strands the prompt on the composer until a bare `\r` follows: all three are reasons - to let `sendwait` build the call rather than hand-rolling it. + swallowed as an already-applied duplicate, and a lost Enter strands the prompt on the + composer until a bare `\r` follows: Claude Code 2.1.277 and later ignore Enter for the + first 30 to 50 seconds after the composer paints while still taking the text, so + `sendwait` reads the composer and keeps pressing Enter until the prompt has left it. + All three are reasons to let `sendwait` build the call rather than hand-rolling it. - Each `sendwait` costs that worker one billed turn, as does every prompt you send it. - Deleting the sessions does **not** remove the case directories. They are marked as agent-created, so `GET /api/v1/cases/agent-created` lists them for cleanup: §5.14. diff --git a/skills/codeman/preamble.sh b/skills/codeman/preamble.sh index cadc4a78..11982e34 100644 --- a/skills/codeman/preamble.sh +++ b/skills/codeman/preamble.sh @@ -1,4 +1,4 @@ -# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ---- +# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ---- API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}" SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}" # Credentials, cheapest first. Your session has usually INHERITED the server's @@ -72,6 +72,27 @@ _trust_key() { # -> "confirm" | "move" | "" (nothing safe to press) | tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \ | sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/' } +# ---- the composer: is the prompt still sitting there, unsent? ---- +# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the +# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that +# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt +# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through +# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS +# the composer and keeps pressing Enter while the prompt is still there. +_composer_text() { # -> the composer's text with ALL whitespace removed: "" once + # the prompt was taken, "?" when the pane shows no composer at all. The composer is + # the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher + # up in the transcript, so only the last one says whether the text was taken. + local t + t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \ + | jq -r '.data.terminalBuffer // empty' \ + | sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \ + | tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1) + [ -n "$t" ] || { printf '?'; return 0; } + # Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does + # not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH). + printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" +} _accept_trust() { # -> 0 once it has answered the dialog, 1 if it could not local sid="$1" k i=1 while [ "$i" -le 6 ]; do @@ -184,12 +205,16 @@ spawn_workers() { # worker a silent no-op that still "succeeds" and reports the previous turn's state. # Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a # deliberate duplicate, at the SAME number (§5.3). -# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the -# typed prompt stranded on the composer while a long wait runs its whole timeout -# (observed live). So the first wait is short; on its timeout a bare \r goes out (the -# missing Enter when the prompt is stranded, a no-op when the turn is genuinely -# running), then the ORIGINAL frame is resent unchanged, which the server takes as a -# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker +# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude +# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints), +# leaving the typed prompt stranded on the composer while a long wait runs its whole +# timeout (observed live, twelve reviews in a row). So the first wait is short; on its +# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate: +# the server re-waits without retyping, §5.3) and kept open in the background, while +# the composer is READ (_composer_text) and, as long as the prompt is still sitting +# there, a bare \r goes out about every ten seconds, up to twelve times. An empty +# composer ends the loop, so a prompt that was taken is never nudged again, and the +# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker # spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) -- # and for those only. Hook-less workspaces and the other modes resolve on flapping # idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not @@ -197,7 +222,7 @@ spawn_workers() { # it accepts the send and then burns both waits. One timeout on a dsh worker whose # pane clearly finished means that profile, so switch that worker to markers. sendwait() { - local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r + local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i # `wait:"stop,exit"`, never the `wait:true` default set: that set also carries # `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a # dsh worker whose TUI repaints rarely the session reads `idle` while the model @@ -212,16 +237,38 @@ sendwait() { r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ -H 'Content-Type: application/json' --data-binary "$body") if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then - "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \ - -d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \ - '{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null - # The resend is a tagged DUPLICATE, so the server skips the write and reports - # `delivered:false` for it -- truthfully, but about the wrong send. The first - # one delivered, so carry that forward, or §1's cleanup reads a completed turn - # as an undelivered one and keeps a finished worker forever. - r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ - -H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \ - | jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end') + # ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call, + # in the background, while the Enter loop below works the composer. Signals have + # no history: a `stop` that fires while no wait is open (during a composer read + # between two short waits, measured) is lost, and the next wait then runs its + # whole timeout on a turn that already ended. The resend is a tagged DUPLICATE, + # so the server skips the write and re-waits without retyping (§5.3). + tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1 + "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \ + -H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" & + bg=$! + # The prompt's head with whitespace removed, matched literally (the "$head" + # quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob). + head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24) + while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended + c=$(_composer_text "$sid") + if [ "$c" = '?' ]; then + [ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it + elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then + break # composer empty (taken) or holding other text + fi + n=$((n+1)) + "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \ + -d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \ + '{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null + i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done + done + wait "$bg" + # The duplicate reports `delivered:false` -- truthfully, but about the wrong send. + # The first one delivered, so carry that forward, or §1's cleanup reads a completed + # turn as an undelivered one and keeps a finished worker forever. + r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp") + rm -f "$tmp" fi printf '%s\n' "$r" } @@ -247,4 +294,4 @@ last_text() { # The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept # bare on purpose: the write condition above anchors on it with $, so an inline comment # here would fail that match and rewrite this file on every single bootstrap. -CODEMAN_PREAMBLE=1.22.0 +CODEMAN_PREAMBLE=1.30.1 diff --git a/skills/codeman/reference/recipes.md b/skills/codeman/reference/recipes.md index 94e0264f..af814653 100644 --- a/skills/codeman/reference/recipes.md +++ b/skills/codeman/reference/recipes.md @@ -21,7 +21,7 @@ by sourcing the preamble file the §0 bootstrap wrote, and checking its version ```bash . "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null -[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; } +[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; } ``` Do **not** re-paste the preamble body into each call. Sourcing it is what retires the diff --git a/src/config/cli-registry/schema.ts b/src/config/cli-registry/schema.ts index 52cdfbd2..d1249384 100644 --- a/src/config/cli-registry/schema.ts +++ b/src/config/cli-registry/schema.ts @@ -340,6 +340,35 @@ const capabilitiesSchema = z // an env var, so it declares baseUrl/apiKey injection with no model var at all. modelVars: z.array(envName).max(8), launchModel: launchModelTemplate, + // Optional: the env var to carry a discovered per-model context-window size + // (claude's CLAUDE_CODE_MAX_CONTEXT_TOKENS), and/or the env var that isolates + // this session's config/credential directory from the user's real one (claude's + // CLAUDE_CONFIG_DIR) so an injected API key never collides with a stored OAuth + // session. See the customModelInjection doc comment in cli-registry/types.ts. + contextLengthVar: envName.optional(), + configDirVar: envName.optional(), + // Relative path, WITHIN the isolated configDirVar directory, of a trust-dialog + // seed file the CLI itself owns the shape of — claude's `.claude.json` + // `customApiKeyResponses.approved` list, the same field an interactive "Detected + // a custom API key — use it?" prompt writes to on a real terminal. Only makes + // sense alongside configDirVar (an isolated, otherwise-empty directory has none + // of a real profile's prior approvals), and only implemented for the + // 'claude-api-key-responses' shape today — see custom-model-injection-apply.ts. + apiKeyTrustFile: z + .object({ relPath: z.string().min(1).max(80), shape: z.literal('claude-api-key-responses') }) + .strict() + .optional(), + // An isolated config directory replays the CLI's whole first-run sequence (theme + // picker, security notes, per-project trust dialog, bypass-permissions warning) + // on every launch, same root cause as apiKeyTrustFile above — this reuses that + // same file to pre-seed the state a real, already-onboarded profile carries. See + // the customModelInjection doc comment in cli-registry/types.ts. + skipFirstRunPrompts: z.boolean().optional(), + // DeepSeek-only, confirmed by reading its own bundled SDK source: it concatenates + // "/chat/completions" onto baseUrlVar's value with no "/v1" of its own, while + // llama-swap/llama.cpp only serves the "/v1/..." path — claude/gemini must NOT + // get this. See the customModelInjection doc comment in cli-registry/types.ts. + appendV1Suffix: z.boolean().optional(), }) .strict(), z diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index eb054082..1e455192 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -228,15 +228,31 @@ const CLAUDE: CliEntry = { privilegedParams: [], // ANTHROPIC_* is NOT in allowedPrefixes/allowedKeys above (deliberately — see the // allowedPrefixes comment nearby), so these are unreachable via plain envOverrides - // today; listed here only so the dedicated custom-model route (docs/custom-model-endpoints-plan.md - // chunk 5) clamps them for a non-granted multi-user owner the same way every other - // CLI's injection vars are clamped, the day that route widens who can set them. + // today. privilegedEnvKeys has exactly one consumer, ownerClampedEnvKeys() in + // session-env-clamp.ts, which feeds the generic envOverrides clamp on + // POST /api/sessions, POST /api/quick-start and reboot-restore — no custom-model + // route reads this field at all, and the values it injects are merged in AFTER + // that clamp runs regardless of what's listed here. privilegedEnvKeys: [ 'ANTHROPIC_BASE_URL', 'ANTHROPIC_API_KEY', 'ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL', + // CLAUDE_CODE_MAX_CONTEXT_TOKENS already matches the CLAUDE_CODE_* allowedPrefix, and + // CLAUDE_CONFIG_DIR is already an allowed exact key (docs/wiki/Agent-CLIs.md), so both + // were already reachable via plain envOverrides before this pair existed and this + // feature does not strictly need either listed. They stay listed anyway, because + // types.ts's rule ("every traffic-redirecting var this feature introduces MUST also + // appear in privilegedEnvKeys") is meant to hold literally, not with an exception + // carved out for the two vars that happen not to need it today. The real + // consequence lands on the GENERIC envOverrides clamp above, not on this feature: + // a non-granted multi-user owner can no longer set CLAUDE_CONFIG_DIR through + // envOverrides at all (the per-client-account override, #255), and a PERSISTED one + // is now stripped on reboot-restore for such an owner too — see + // session-env-clamp.ts's own fileoverview. + 'CLAUDE_CODE_MAX_CONTEXT_TOKENS', + 'CLAUDE_CONFIG_DIR', ], gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } }, // Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real @@ -247,6 +263,32 @@ const CLAUDE: CliEntry = { baseUrlVar: 'ANTHROPIC_BASE_URL', apiKeyVar: 'ANTHROPIC_API_KEY', modelVars: ['ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL'], + // Verified via Claude Code's own docs: CLAUDE_CODE_MAX_CONTEXT_TOKENS overrides the + // assumed context window and applies directly for a model name Claude Code doesn't + // recognize as one of its own — exactly the custom-model case. Without it, Claude Code + // assumes a large (200k) window for any unrecognized model id and never compacts, + // eventually overflowing a much smaller real local context (see plan doc reasoning + // above the interface for the confirmed failure). + contextLengthVar: 'CLAUDE_CODE_MAX_CONTEXT_TOKENS', + // Isolates this session's config/credential directory so an injected ANTHROPIC_API_KEY + // never shares a directory with a stored claude.ai OAuth login — see the doc comment on + // customModelInjection in cli-registry/types.ts for the traded-off side effect. + configDirVar: 'CLAUDE_CONFIG_DIR', + // ⚠️ Required alongside configDirVar, not optional in practice: verified live that an + // isolated, otherwise-empty config directory makes claude stop at an interactive + // "Detected a custom API key — use it?" prompt on EVERY launch, defaulting to "No" with + // no one at the TTY to answer — silently refusing the very key this feature injected. + // Pre-seeding this file's customApiKeyResponses.approved list (verified against a real + // ~/.claude.json after answering the prompt once by hand) answers it in advance instead. + apiKeyTrustFile: { relPath: '.claude.json', shape: 'claude-api-key-responses' }, + // ⚠️ Same isolated-directory root cause, one step further: verified live that on top + // of the API-key prompt above, a fresh CLAUDE_CONFIG_DIR also replays claude's ENTIRE + // first-run sequence on every launch — the theme picker, the security-notes screen, + // the per-project "trust this folder?" dialog, and (running with + // --dangerously-skip-permissions) a one-time bypass-permissions warning — none of + // which a real, already-onboarded profile shows again. Pre-seeds that same + // already-onboarded state instead of leaving a human to click through it. + skipFirstRunPrompts: true, }, }, overlays: { @@ -1071,15 +1113,28 @@ const DEEPSEEK: CliEntry = { // privilege rather than granting it, and clamping it here was a real regression // (test/deepseek-mode.test.ts) fixed before this shipped. privilegedEnvKeys: ['DSH_PERMISSION_MODE', 'DSH_HOME', 'DEEPSEEK_BASE_URL'], - // Web-researched, unverified, partial: reuses the already-existing DEEPSEEK_BASE_URL/ - // DEEPSEEK_API_KEY keys above. No modelVars — dsh's model is a profile-composition - // entry (see `model: { source: 'none' }` above), not an env var, so forcing a specific - // model name may not fully work; verify against a real profile before shipping. + // Reuses the already-existing DEEPSEEK_BASE_URL/DEEPSEEK_API_KEY keys above. No + // modelVars — dsh's model is a profile-composition entry (see `model: { source: 'none' + // }` above), not an env var, so forcing a specific model name may not fully work; + // verify against a real profile before shipping. + // + // ⚠️ appendV1Suffix is REQUIRED, not optional-nice-to-have: without it every request + // 404s. Confirmed live and by reading dsh's own bundled source + // (@deepseek-ai/dsh-llm-deepseek): it builds the request URL as + // `${DEEPSEEK_BASE_URL}/chat/completions` with no "/v1" of its own (its real public + // API, https://api.deepseek.com, expects the caller's base URL to already carry any + // needed prefix), while llama-swap/llama.cpp only serves the OpenAI-conventional + // "/v1/chat/completions" — a bare POST to ".../chat/completions" 404s live, and the + // 404 reported here originally ("dsh: HTTP_404: DeepSeek API error (HTTP 404)") + // matches dsh's own error-message template for exactly this failure. See the + // customModelInjection doc comment in cli-registry/types.ts for the full reasoning, + // including why claude/gemini must NOT get this. customModelInjection: { kind: 'env', baseUrlVar: 'DEEPSEEK_BASE_URL', apiKeyVar: 'DEEPSEEK_API_KEY', modelVars: [], + appendV1Suffix: true, }, }, overlays: { diff --git a/src/config/cli-registry/types.ts b/src/config/cli-registry/types.ts index 1ba07cb0..871a3762 100644 --- a/src/config/cli-registry/types.ts +++ b/src/config/cli-registry/types.ts @@ -496,9 +496,74 @@ export interface CliCapabilities { * declares). Absent = the config alone selects the model (claude's env vars, * opencode's blob, codex's top-level `model` key). Applied by the session's * respawn options through the entry's `legacyConfigField`, never by id. + * + * `contextLengthVar` (env kind only): the env var a discovered per-model context-window + * size is written to when known (claude's `CLAUDE_CODE_MAX_CONTEXT_TOKENS`) — without it, + * a CLI that assumes a large default window for an unrecognized model name keeps sending + * full-size prompts against a much smaller local server and eventually overflows its real + * context (verified: a 33.7K-token system prompt against a 16384-token llama-swap model). + * Absent when the CLI has no such override, or the value is unknown for this model. + * + * `configDirVar` (env kind only): the env var that redirects this session's config/ + * credential directory to an isolated, per-session one (claude's `CLAUDE_CONFIG_DIR`), so + * an injected API key never coexists with a stored claude.ai OAuth session in the same + * directory — the CLI still warns "both claude.ai and ANTHROPIC_API_KEY set" when they + * share a directory even though the API key wins for actual requests. Isolating it trades + * that cosmetic warning for a documented side effect: a relocated config directory writes + * transcripts outside `~/.claude/projects`, blinding the response viewer, subagent + * windows, and Read My Mind for that session (see docs/wiki/Agent-CLIs.md). + * + * `apiKeyTrustFile` (env kind only, alongside configDirVar): an isolated config directory + * has none of a real profile's prior "detected a custom API key, use it?" approvals, so + * without this the CLI stops and asks interactively on every single launch — with no one + * at a TTY to answer, that's a hang, not a warning (confirmed live: claude's own default + * answer, "No", would silently refuse to use the very key this feature just injected). + * `relPath`/`shape` name the file (claude's `.claude.json`) and its + * `customApiKeyResponses.approved` field this pre-seeds — the exact field a real answered + * prompt itself writes to, so this isn't bypassing the check, just answering it the same + * way a one-off prior approval on a shared profile already would. + * + * `skipFirstRunPrompts` (env kind only, alongside apiKeyTrustFile): an isolated config + * directory is not just missing API-key approvals — it is a brand-new profile as far as + * the CLI is concerned, so it also replays its ENTIRE first-run sequence on every launch: + * the theme picker, the security-notes screen, the per-project "trust this folder?" + * dialog, and (running with a bypass-permissions flag) a one-time warning about it — + * confirmed live, none of which a real, long-used profile ever shows again. `true` + * pre-seeds the same state a real profile accumulates from having answered all of that + * once: `hasCompletedOnboarding` and the launching session's own project entry in the + * `apiKeyTrustFile` (claude's `.claude.json`), plus `skipDangerousModePermissionPrompt` + * in claude's `settings.json` — see `seedFirstRunState`/`seedSkipBypassPermissionsPrompt` + * in custom-model-injection-apply.ts. Requires `apiKeyTrustFile` to be set too, since it + * reuses that file. + * + * `appendV1Suffix` (env kind only): the raw `endpoint.baseUrl` gets `withV1Suffix()` + * applied before being written to `baseUrlVar`, instead of being used verbatim. + * DeepSeek needs this and claude/gemini must NOT get it — a per-CLI asymmetry confirmed + * by reading each SDK's own request-building source, not assumed: DeepSeek Harness's + * bundled `@deepseek-ai/dsh-llm-deepseek` concatenates `${connection.baseURL}/chat/ + * completions` with no `/v1` insertion of its own (its real public API base, + * `https://api.deepseek.com`, expects the caller's base URL to already carry any + * needed prefix), while llama-swap/llama.cpp only ever serves the OpenAI-conventional + * `/v1/chat/completions` — confirmed live: a bare `POST /chat/completions` + * 404s, `POST /v1/chat/completions` succeeds, and the harness's own error + * message template (`DeepSeek API error (HTTP ${status})`) reproduces the exact + * `HTTP_404` this feature originally shipped with unexplained. Claude Code's own SDK, + * by contrast, was already confirmed working end-to-end against the RAW `baseUrl` with + * no suffix — appending one there would be wrong, not just redundant. */ customModelInjection: - | { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string } + | { + kind: 'env'; + baseUrlVar: string; + apiKeyVar: string; + modelVars: string[]; + launchModel?: string; + contextLengthVar?: string; + apiKeyTrustFile?: { relPath: string; shape: 'claude-api-key-responses' }; + configDirVar?: string; + skipFirstRunPrompts?: boolean; + appendV1Suffix?: boolean; + } | { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string } | { kind: 'configDir'; diff --git a/src/config/remote-wake-limits.ts b/src/config/remote-wake-limits.ts new file mode 100644 index 00000000..14b840de --- /dev/null +++ b/src/config/remote-wake-limits.ts @@ -0,0 +1,23 @@ +/** + * @fileoverview Limits shared between Wake-on-LAN parsing and its request schema. + * + * Its own module because `src/remote-wake.ts` is import-fenced: only + * `web/routes/session-routes.ts` and `web/server.ts` may import it, so that no + * watcher or boot-recovery path can WAKE a host (pinned by the wiring guard in + * `test/remote-wake.test.ts`). `web/schemas.ts` needs the same MAC-count limit and + * must not become a third importer, and it would drag `dgram`/`net`/`child_process` + * into every module that validates a request body. A plain constant satisfies both. + */ + +/** + * How many comma-separated MACs one `wakeMac` may carry. + * + * ⚠ Single source for `parseMacList()` and `RemoteHostSchema.wakeMac`. The two used + * to disagree: the schema's 128-character cap admits seven MACs while the parser + * rejected more than four all-or-nothing, so a five-MAC value validated, persisted to + * `remote-hosts.json`, and then resolved to NO wake target. The host read as + * unconfigured and the banner offered "Configure WoL" for a host the user had just + * configured, which is the worst shape a validation gap can take: accepted, stored, + * silently inert. + */ +export const MAX_WAKE_MACS = 4; diff --git a/src/custom-model-hosts.ts b/src/custom-model-hosts.ts index 0cde1039..bd90cef7 100644 --- a/src/custom-model-hosts.ts +++ b/src/custom-model-hosts.ts @@ -40,6 +40,38 @@ export interface CustomModelHost { authStyle?: CustomModelAuthStyle; models?: string[]; lastDiscoveredAt?: string; + /** + * The model the Run-menu picker (docs/custom-model-endpoints-plan.md) applies when + * this endpoint is picked with no further choice — one generated menu entry per + * (CLI, endpoint) pair, not per (CLI, endpoint, model), so it needs a single answer. + * Must be a member of `models` when set; the picker falls back to `models[0]` when + * this is unset, and disables the entry entirely when `models` is empty (nothing to + * default to). Never auto-set on discovery — the previous default staying valid + * after a re-discover is a property worth keeping even if the model list changes. + */ + defaultModelId?: string; + /** + * Discovered context-window size (tokens) per model id, keyed by the same strings as + * `models`. Populated opportunistically during discovery (`custom-model-routes.ts`) from + * llama.cpp/llama-swap's `GET /props?model=` — the plain OpenAI-shaped `/v1/models` + * response has no such field. Only ever probed for a model the server already reports as + * loaded (llama-swap's `status.value === 'loaded'`); an unloaded one is deliberately never + * probed, since llama-swap treats `/props?model=` as a routing hint that can trigger an + * actual (slow, GPU-swapping) model load as a side effect of merely asking. A model this + * has no entry for simply gets no context-length env override applied — never a guess. + */ + modelContextLengths?: Record; + /** + * Discovered file size (GB) per model id, keyed by the same strings as `models`. + * Populated during discovery by parsing llama-swap's own `description` field for an + * auto-discovered model ("Auto-discovered 16.35 GB - parameters auto-fitted by + * llama.cpp") — a hand-configured profile's own description has no such figure and + * correctly gets no entry, never a guess. Used only to label the Run-menu picker's + * "loading model" banner with a rough, unmeasured expected-time estimate + * (the Run-menu picker's loading banner in session-ui.js) — never a guarantee, and never anything a + * server-side check relies on. + */ + modelSizesGB?: Record; } export function customModelHostsPath(configDir: string): string { diff --git a/src/custom-model-injection-apply.ts b/src/custom-model-injection-apply.ts index 2df6f57d..43cfd5ad 100644 --- a/src/custom-model-injection-apply.ts +++ b/src/custom-model-injection-apply.ts @@ -10,7 +10,8 @@ * cli-registry changes" requirement it was written against. */ -import { chmodSync, mkdirSync, writeFileSync, rmSync } from 'node:fs'; +import { chmodSync, existsSync, mkdirSync, readFileSync, writeFileSync, rmSync, symlinkSync } from 'node:fs'; +import { homedir, platform } from 'node:os'; import { join, dirname } from 'node:path'; import { dataPath } from './config/instance.js'; import type { CliEntry } from './config/cli-registry/types.js'; @@ -48,6 +49,168 @@ export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInj return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv }; } +/** + * Real, shared Claude config directory Codeman's own host process runs under — honors + * `CLAUDE_CONFIG_DIR` the same way `claude-credentials.ts`'s `claudeCredentialsPath()` + * does, so the symlink below points at wherever `~/.claude/projects` actually lives + * rather than assuming the plain default. + */ +function realClaudeConfigDir(): string { + const configured = typeof process.env.CLAUDE_CONFIG_DIR === 'string' && process.env.CLAUDE_CONFIG_DIR.trim(); + return configured || join(homedir(), '.claude'); +} + +/** + * Symlinks `/projects` back to the real, shared `~/.claude/projects`, so an + * isolated `CLAUDE_CONFIG_DIR` (used to keep an injected API key away from a stored OAuth + * session — see `configDirVar` on customModelInjection) doesn't also blind the response + * viewer, subagent windows, and Read My Mind for that session (docs/wiki/Agent-CLIs.md). + * Best-effort: a platform that refuses symlinks (unprivileged Windows without a junction + * fallback working, e.g.) just keeps the pre-existing documented side effect instead of + * failing the whole custom-model apply over a nice-to-have. + */ +function linkSharedProjectsDir(isolatedDir: string): void { + const link = join(isolatedDir, 'projects'); + if (existsSync(link)) return; // already linked (idempotent re-apply) or real dir wrote one + try { + symlinkSync(join(realClaudeConfigDir(), 'projects'), link, platform() === 'win32' ? 'junction' : 'dir'); + } catch { + // best-effort only — response viewer/subagent windows go blind for this session instead + } +} + +/** + * Pre-approves the injected API key in an isolated config directory's trust-dialog state + * (`customModelInjection.apiKeyTrustFile`), so an otherwise-empty directory doesn't make the + * CLI stop at an interactive "Detected a custom API key — use it?" prompt on every single + * launch. Confirmed live: with nobody at the TTY to answer, that prompt's own default + * ("No") silently refuses the very key this feature just injected — this isn't bypassing + * the check, it's answering it the same field a real answered prompt itself writes to + * (verified against a real `~/.claude.json` after answering by hand once). + * + * Merges rather than overwrites: the file may already carry fields the CLI itself wrote on + * an earlier launch in this same isolated directory (machineID, userID, other approved + * keys), and a corrupt or partially-written file (a crash mid-write) is treated as absent + * rather than failing the whole apply over a nice-to-have. + */ +/** + * The form Claude Code actually stores an approved key in: the trimmed last 20 + * characters. Mirrors the CLI's own `e.trim().slice(-20)`, which is applied on BOTH + * the write and the lookup, so anything else never matches. + */ +export function truncateApiKeyForTrustFile(apiKey: string): string { + return apiKey.trim().slice(-20); +} + +function seedApiKeyTrustFile( + configDir: string, + trustFile: { relPath: string; shape: 'claude-api-key-responses' }, + apiKey: string +): void { + const filePath = join(configDir, trustFile.relPath); + let existing: Record = {}; + try { + existing = JSON.parse(readFileSync(filePath, 'utf8')) as Record; + } catch { + existing = {}; + } + const responses = (existing.customApiKeyResponses ?? {}) as { approved?: unknown; rejected?: unknown }; + const approved = new Set(Array.isArray(responses.approved) ? (responses.approved as string[]) : []); + // ⚠ Claude Code stores and compares only the LAST 20 CHARACTERS of a key, never the + // whole thing: its lookup is `approved.includes(key.trim().slice(-20))` (decompiled + // from the 2.1.278 bundle, and corroborated by real `~/.claude.json` files, whose + // customApiKeyResponses entries are all exactly 20 characters). Seeding the full key + // therefore never matches for a REAL key, and claude stops at the interactive + // "Detected a custom API key in your environment" prompt, whose default is + // "No (recommended)" — so the launch hangs or silently refuses the key this feature + // just injected. It went unnoticed because a keyless llama.cpp/llama-swap endpoint + // uses DEFAULT_API_KEY ('local-dummy-key', 15 chars), where slice(-20) is the whole + // string and the seed matches by accident. Truncating here also keeps a full + // third-party credential from being written into a second file on disk. + approved.add(truncateApiKeyForTrustFile(apiKey)); + const rejected = Array.isArray(responses.rejected) ? responses.rejected : []; + existing.customApiKeyResponses = { approved: [...approved], rejected }; + try { + writeFileSync(filePath, JSON.stringify(existing, null, 2), { encoding: 'utf8', mode: 0o600 }); + chmodSync(filePath, 0o600); + } catch { + // best-effort only — the interactive prompt returns instead of a hard failure here + } +} + +/** + * Pre-seeds the two remaining pieces of "already been onboarded" state a fresh + * `CLAUDE_CONFIG_DIR` has none of (`customModelInjection.skipFirstRunPrompts`, alongside + * apiKeyTrustFile): claude replays its whole first-run sequence — the theme picker, the + * security-notes screen, and (per-project) the "trust this folder?" dialog — against ANY + * config directory that has never completed it, confirmed live against a genuinely fresh + * isolated directory. `hasCompletedOnboarding` skips the theme/security-notes screens + * outright; `projects[workingDir].hasTrustDialogAccepted` answers the trust dialog for + * THIS session's own working directory the same way a real profile's own prior approval + * would — other projects in the file are left alone, and `workingDir` is used verbatim + * (never realpath'd or slash-normalized) since that's the literal string claude itself + * uses as the project key, being whatever string the session was actually launched with + * as its cwd. + * + * Same merge-not-overwrite and corrupt-file-tolerant behavior as `seedApiKeyTrustFile` + * (same file, so a second sequential read-modify-write here is deliberate rather than + * folding both into one pass — keeps each seed independently testable and optional). + */ +function seedFirstRunOnboardingState( + configDir: string, + trustFile: { relPath: string; shape: 'claude-api-key-responses' }, + workingDir: string +): void { + const filePath = join(configDir, trustFile.relPath); + let existing: Record = {}; + try { + existing = JSON.parse(readFileSync(filePath, 'utf8')) as Record; + } catch { + existing = {}; + } + existing.hasCompletedOnboarding = true; + const projects = + existing.projects && typeof existing.projects === 'object' && !Array.isArray(existing.projects) + ? (existing.projects as Record>) + : {}; + const existingProject = projects[workingDir] && typeof projects[workingDir] === 'object' ? projects[workingDir] : {}; + projects[workingDir] = { ...existingProject, hasTrustDialogAccepted: true }; + existing.projects = projects; + try { + writeFileSync(filePath, JSON.stringify(existing, null, 2), { encoding: 'utf8', mode: 0o600 }); + chmodSync(filePath, 0o600); + } catch { + // best-effort only — the interactive dialogs return instead of a hard failure here + } +} + +/** + * Pre-seeds the "skip the bypass-permissions warning" setting (`customModelInjection. + * skipFirstRunPrompts`, alongside apiKeyTrustFile) into an isolated config directory's + * `settings.json` — a real, already-onboarded profile answers claude's one-time warning + * about running with a bypass-permissions flag once and never sees it again, but every + * custom-model session launches with a fresh, otherwise-empty CLAUDE_CONFIG_DIR that + * carries none of that (confirmed live). A different file from apiKeyTrustFile's + * `.claude.json` — this is claude's own global `settings.json`, not project-keyed — + * so it gets its own merge-not-overwrite read-modify-write. + */ +function seedSkipBypassPermissionsPrompt(configDir: string): void { + const filePath = join(configDir, 'settings.json'); + let existing: Record = {}; + try { + existing = JSON.parse(readFileSync(filePath, 'utf8')) as Record; + } catch { + existing = {}; + } + existing.skipDangerousModePermissionPrompt = true; + try { + writeFileSync(filePath, JSON.stringify(existing, null, 2), { encoding: 'utf8', mode: 0o600 }); + chmodSync(filePath, 0o600); + } catch { + // best-effort only — the interactive warning returns instead of a hard failure here + } +} + /** Best-effort recursive removal of a previously-written configDir. Never throws. */ export function removeConfigDir(dir: string | undefined): void { if (!dir) return; @@ -79,14 +242,43 @@ export function applyCustomModelInjection( entry: Pick, endpoint: CustomModelEndpoint, modelId: string, - sessionId: string + sessionId: string, + /** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */ + contextLength?: number, + /** + * The session's own working directory — only used for `skipFirstRunPrompts`'s per-project + * trust-dialog seed, and only when provided (boot recovery, which has no reason to + * re-answer a dialog that already fired once, omits it rather than re-deriving it). + */ + workingDir?: string ): AppliedCustomModel | undefined { - const injection = buildCustomModelInjection(entry, endpoint, modelId); + const injection = buildCustomModelInjection(entry, endpoint, modelId, contextLength); if (injection.kind === 'unsupported') return undefined; if (injection.kind === 'env') { + // `configDirVar` (claude's CLAUDE_CONFIG_DIR): point it at the same isolated, + // per-session directory the `configDir` kind uses, but write no files into it — an + // empty directory has no stored OAuth credential to conflict with the injected API + // key, which is the whole point. Reusing the same path keyed by sessionId keeps this + // idempotent across a boot-recovery re-apply, same as the configDir kind below. + let envOverrides = injection.envOverrides; + let configDir: string | undefined; + if (injection.configDirVar) { + configDir = customModelConfigDir(sessionId); + mkdirSync(configDir, { recursive: true, mode: 0o700 }); + linkSharedProjectsDir(configDir); + if (injection.apiKeyTrustFile && injection.apiKey) { + seedApiKeyTrustFile(configDir, injection.apiKeyTrustFile, injection.apiKey); + } + if (injection.skipFirstRunPrompts && injection.apiKeyTrustFile) { + if (workingDir) seedFirstRunOnboardingState(configDir, injection.apiKeyTrustFile, workingDir); + seedSkipBypassPermissionsPrompt(configDir); + } + envOverrides = { ...envOverrides, [injection.configDirVar]: configDir }; + } return { - envOverrides: injection.envOverrides, - envKeys: Object.keys(injection.envOverrides), + envOverrides, + envKeys: Object.keys(envOverrides), + configDir, launchModel: injection.launchModel, }; } diff --git a/src/custom-model-injection.ts b/src/custom-model-injection.ts index 5f46001c..9aab9422 100644 --- a/src/custom-model-injection.ts +++ b/src/custom-model-injection.ts @@ -16,14 +16,21 @@ * shape was rejected by a real codex binary with "invalid type: map, * expected a string" — caught by `scripts/test-local-llm-harnesses.ts`), * but `wire_api = "responses"` is the only value codex still accepts - * (support for `"chat"` was dropped in Feb 2026), and a plain OpenAI - * Chat-Completions server (llama.cpp, llama-swap, most local setups) does - * NOT implement the Responses API — so codex may still fail at the - * PROTOCOL level even with a correctly-shaped config file. That gap is - * real and current, not a stale warning; see docs/custom-model-endpoints-plan.md. The rest - * (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION flags - * confirmed against real installed binaries' own `--help` output, but - * their custom-endpoint env/config conventions remain web-researched, + * (support for `"chat"` was dropped in Feb 2026). ⚠️ Re-verified live + * against a llama-swap deployment that DOES answer `/v1/responses`: a + * plain, no-tool-call turn gets a real reply, but a real tool-call attempt + * comes back as `agent_message` TEXT (the tool-call JSON printed as the + * answer) rather than a `function_call` item codex would execute — + * confirmed via `codex exec --json`'s raw event stream. Tool execution is + * what makes codex a coding agent, so this remains not usable for real + * work even where plain chat succeeds; see docs/custom-model-endpoints-plan.md + * for the full picture (including the harmless `Model metadata ... not + * found` warning every custom-endpoint codex session prints — sourced from + * a local cache of OpenAI's OWN hosted model catalog that a custom model + * can never appear in, confirmed to have no effect on the outcome above). + * The rest (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION + * flags confirmed against real installed binaries' own `--help` output, + * but their custom-endpoint env/config conventions remain web-researched, * unverified. */ @@ -44,6 +51,20 @@ export interface EnvInjection { envOverrides: Record; /** See {@link ConfigDirInjection.launchModel}. */ launchModel?: string; + /** + * Name of the env var the caller should point at an isolated, credential-free config + * directory for this session (claude's `CLAUDE_CONFIG_DIR`), from the registry entry's + * `customModelInjection.configDirVar`. The actual directory value isn't computed here — + * this module is pure and has no sessionId to derive one from — the IO wrapper + * (`custom-model-injection-apply.ts`) creates it and adds it to `envOverrides`. + */ + configDirVar?: string; + /** See `customModelInjection.apiKeyTrustFile` — carried through so the IO wrapper can seed it. */ + apiKeyTrustFile?: { relPath: string; shape: 'claude-api-key-responses' }; + /** The literal API key value this injection used, for `apiKeyTrustFile` to pre-approve. */ + apiKey?: string; + /** See `customModelInjection.skipFirstRunPrompts` — carried through so the IO wrapper can seed it. */ + skipFirstRunPrompts?: boolean; } export interface ConfigDirInjection { @@ -90,7 +111,9 @@ function quoted(value: string): string { export function buildCustomModelInjection( entry: Pick, endpoint: CustomModelEndpoint, - modelId: string + modelId: string, + /** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */ + contextLength?: number ): CustomModelInjectionResult { const cap = entry.capabilities.customModelInjection; const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY; @@ -98,11 +121,18 @@ export function buildCustomModelInjection( switch (cap.kind) { case 'env': { const envOverrides: Record = { - [cap.baseUrlVar]: endpoint.baseUrl, + [cap.baseUrlVar]: cap.appendV1Suffix ? withV1Suffix(endpoint.baseUrl) : endpoint.baseUrl, [cap.apiKeyVar]: apiKey, }; for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId; - return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId); + if (cap.contextLengthVar && contextLength !== undefined && Number.isFinite(contextLength)) { + envOverrides[cap.contextLengthVar] = String(Math.trunc(contextLength)); + } + let result: EnvInjection = withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId); + if (cap.configDirVar) result = { ...result, configDirVar: cap.configDirVar }; + if (cap.apiKeyTrustFile) result = { ...result, apiKeyTrustFile: cap.apiKeyTrustFile, apiKey }; + if (cap.skipFirstRunPrompts) result = { ...result, skipFirstRunPrompts: true }; + return result; } case 'configContentEnv': { @@ -177,11 +207,13 @@ function renderConfigFile( // `env_key`, the NAME of an env var it reads the credential from at runtime, so the // actual value must ride along as an extra env var, never embedded in the file. // ⚠️ `wire_api = "responses"` is the only value codex still accepts (it dropped - // `"chat"` support in Feb 2026) — a plain OpenAI Chat-Completions server (llama.cpp, - // llama-swap, most local setups) does NOT implement the Responses API, so this - // recipe may still fail at the PROTOCOL level even though the file now parses - // correctly. That is a real, currently-unresolved compatibility gap, not a syntax - // bug — track it before calling codex support done. + // `"chat"` support in Feb 2026). Even against a llama-swap deployment that DOES + // answer `/v1/responses`, a real tool-call attempt came back as plain TEXT (the + // tool-call JSON printed as the model's answer) rather than an executable + // `function_call` item — confirmed live via `codex exec --json`. Tool execution is + // what makes codex a coding agent, so this remains not usable for real work even + // where plain chat succeeds — see the confidence table in + // docs/custom-model-endpoints-plan.md, not a syntax bug in this file. const content = [ `model = ${quoted(modelId)}`, `model_provider = "custom"`, diff --git a/src/mux-interface.ts b/src/mux-interface.ts index e1104d15..f76ca47a 100644 --- a/src/mux-interface.ts +++ b/src/mux-interface.ts @@ -159,6 +159,15 @@ export interface PaneCaptureOptions { * the 1MB execSync default (ENOBUFS). */ maxCaptureBytes?: number; + /** + * Filled in by the implementation with the pane geometry the capture was + * really taken at, which is not always the geometry the caller last asked + * for: a resize and a capture can race, and a pane whose size a desktop + * viewport has claimed ignores a smaller client's resize outright. A + * visible-frame capture addresses every row absolutely, so a consumer + * rendering it needs the real height to know the frame fits. + */ + capturedGeometry?: { cols: number; rows: number }; } /** diff --git a/src/remote-hosts.ts b/src/remote-hosts.ts index 37689ac2..3dea3bde 100644 --- a/src/remote-hosts.ts +++ b/src/remote-hosts.ts @@ -540,6 +540,32 @@ export function remoteDisplayPath( return `${remote.username}@${remote.host}:${path}`; } +/** + * Refresh HOST-level config on a RESTORED `SessionRemote`. + * + * A session's `remote` block is persisted at launch time (mux-sessions.json / + * state.json) and recovery uses that snapshot, so a field ADDED to the host config + * later never reaches an already-running session — not even across a Codeman + * restart. That is exactly how a `wakeCommand` added to `remote-hosts.json` would + * silently do nothing until the session is relaunched (which for an owned remote + * session means killing the remote tmux). + * + * Deliberately narrow: ONLY `wakeCommand`/`wakeMac` are taken from the host config, + * and the host is authoritative for them (removing one in the config turns that + * wake path off again). The other host-level fields (`commands`, ssh options) stay as + * persisted so this cannot silently change how an existing pane connects. + */ +export function rehydrateRemoteHostFields( + remote: T | undefined, + hostsById: ReadonlyMap +): T | undefined { + if (!remote) return remote; + const host = hostsById.get(remote.hostId); + if (!host) return remote; + if (remote.wakeCommand === host.wakeCommand && remote.wakeMac === host.wakeMac) return remote; + return { ...remote, wakeCommand: host.wakeCommand, wakeMac: host.wakeMac }; +} + export function toSessionRemote(host: RemoteHost, remoteCase: RemoteCase): SessionRemote { return { hostId: host.id, @@ -549,6 +575,10 @@ export function toSessionRemote(host: RemoteHost, remoteCase: RemoteCase): Sessi port: host.port, remotePath: remoteCase.remotePath, commands: host.commands, + // Wake-on-LAN command/MAC travel with the session so the input route can wake a + // sleeping host without a second config read (see remote-wake.ts). + wakeCommand: host.wakeCommand, + wakeMac: host.wakeMac, // COD-105 — the COD-104 launch path creates the remote session, so we own it // (an explicit kill may propagate a remote kill-session). Discovered+attached // sessions go through `toAttachedSessionRemote` with `owned: false`. @@ -587,6 +617,10 @@ export function toAttachedSessionRemote( port: host.port, remotePath, commands: host.commands, + // An attached session can be woken exactly the same way — the identity of the + // creator does not change whether the host is asleep. + wakeCommand: host.wakeCommand, + wakeMac: host.wakeMac, // Discovered + attached — another Codeman created it. Detach-not-kill. owned: false, remoteSessionName, diff --git a/src/remote-wake.ts b/src/remote-wake.ts new file mode 100644 index 00000000..6226c1fc --- /dev/null +++ b/src/remote-wake.ts @@ -0,0 +1,1035 @@ +/** + * @fileoverview Wake a SLEEPING remote host from user input (user-triggered Wake-on-LAN). + * + * A durable remote session survives SSH drops (COD-104) and auto-reconnects + * (COD-108), but nothing brings the HOST back: if the remote machine suspended, + * the local tmux pane's `ssh` child stalls silently. `tmux send-keys` then + * SUCCEEDS against a pane that will never deliver the bytes, so typed input is + * lost with no error anywhere — the failure this module exists to close. + * + * Design (deliberately narrow, see docs/remote-sessions.md §Wake-on-LAN): + * - An EXPLICIT request wakes a host, and nothing else: user input on an + * established session (`handleInput`), the wake button (`ensureAwake`), or the + * user's own session create/attach request (`ensureHostAwake`, wired in the HTTP + * routes). Everything that runs on a TIMER — the auto-reconnect watcher, boot + * recovery, the reachability probe, session discovery — must never wake one, or + * a host would be re-woken ~45 s after each suspend and could never stay asleep + * (the "keepalive pings a sleeping host" failure already solved for a different + * consumer by `hufflepuff-mcp-lazy`). The create path is deliberately wired in + * `session-routes.ts` and NOT in the shared session service, because + * `cron-service.ts` builds sessions there without a user waiting on the answer. + * - Detection is a cheap TCP connect to the SSH port (no auth, no ssh client, + * a few hundred bytes — below any meaningful activity threshold), throttled + * per session. No SSH keepalive is added to the launch command: keepalives + * would move bytes into an otherwise idle connection every interval, which is + * exactly the "an open pipe keeps the host awake" bug the remote-side idle + * detector was rewritten to avoid. + * - While a wake is in flight, input is BUFFERED and flushed in order once the + * pane is reattached, so the user's first characters after a long pause are + * not the ones that get eaten. + * + * The pure decisions and the IO are separated so the decision table can be + * unit-tested without tmux, ssh, or a real host. + * + * @module remote-wake + */ + +import { spawn } from 'node:child_process'; +import dgram from 'node:dgram'; +import net from 'node:net'; +import { MAX_WAKE_MACS } from './config/remote-wake-limits.js'; + +/** Minimum spacing between two reachability probes for the same session. */ +export const REMOTE_WAKE_PROBE_MIN_INTERVAL_MS = 30_000; +/** TCP-connect timeout for a reachability probe (host awake ≈ a few ms). */ +export const REMOTE_WAKE_PROBE_TIMEOUT_MS = 1_500; +/** Poll spacing while waiting for a woken host to accept SSH again. */ +export const REMOTE_WAKE_READY_INTERVAL_MS = 1_500; +/** Bounded wait for the host to come back after the wake command ran. */ +export const REMOTE_WAKE_READY_TIMEOUT_MS = 90_000; +/** + * Budget for a wake that an HTTP REQUEST is waiting on (session create/attach). + * Deliberately shorter than {@link REMOTE_WAKE_READY_TIMEOUT_MS}: the dashboard is + * served through a reverse proxy whose default `proxy_read_timeout` is 60 s, so a + * 90 s wait would be cut off AT THE PROXY while the session was still being built — + * the browser reports a failure for a session that exists. The budget has to cover + * the WHOLE request, not just the wait: 40 s here + the 1.5 s reachability probe + + * the tmux prereq probe's own 15 s timeout = 56.5 s worst case, still under 60 s. + * ⚠ The wake ITSELF counts against this, which the original arithmetic omitted: a + * `command` target can spend REMOTE_WAKE_COMMAND_TIMEOUT_MS before the readiness poll + * begins, which would have made the real worst case ~68 s. `_wakeAndWait` therefore + * subtracts the wake's measured elapsed time from this budget rather than adding to it. + * A warm S3 resume measures ~12 s, so 40 s is >3× the observed wake. + */ +export const REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS = 40_000; +/** The wake command itself must not hang the wake flow. */ +export const REMOTE_WAKE_COMMAND_TIMEOUT_MS = 10_000; +/** + * Settle time between respawning the ssh pane and flushing buffered input: the + * respawned `ssh` needs a moment to run `tmux -L codeman-remote … -A` and attach, + * and bytes written into a still-connecting pane land in nothing. + */ +export const REMOTE_WAKE_ATTACH_SETTLE_MS = 1_500; +/** + * Cap on buffered input per session while a host is being woken. 4 KB is a lot + * of typing for a ~10 s wake; beyond it the OLDEST bytes are dropped (keeping the + * tail preserves what the user just typed, and a silently unbounded buffer would + * be a memory leak keyed on user input). + */ +export const REMOTE_WAKE_PENDING_MAX_BYTES = 4096; +/** Default SSH port used when the host config has no explicit `port`. */ +export const DEFAULT_SSH_PORT = 22; + +/** What the input path should do with a chunk of user input. Pure. */ +export type RemoteInputAction = 'deliver' | 'probe' | 'buffer'; +/** + * The caller-facing outcome of {@link RemoteWakeRegistry.handleInput}: either the + * caller writes the bytes as usual, or the registry took ownership of them. + */ +export type RemoteInputOutcome = 'deliver' | 'buffered' | 'dropped'; + +/** + * Decide what to do with an input chunk on an input route. Mirrors + * {@link RemoteWakeRegistry.handleInput} so the throttle table has exactly ONE + * definition and is unit-testable: + * + * - a wake already in flight → buffer (the flush owns delivery), + * - no wake command configured → deliver (feature off, today's behavior), + * - the last probe said "down" → buffer (no second probe; re-probing a known + * sleeping host on every keystroke would add seconds of latency per character), + * - never probed / throttle window elapsed → probe, + * - probed "up" inside the window → deliver. + * + * Pure — no clock, no IO. + */ +export function decideRemoteInputAction(args: { + hasWakeTarget: boolean; + waking: boolean; + probeAgeMs: number; + lastReachable?: boolean; + minProbeIntervalMs?: number; +}): RemoteInputAction { + if (args.waking) return 'buffer'; + if (!args.hasWakeTarget) return 'deliver'; + if (args.lastReachable === false) return 'buffer'; + const interval = args.minProbeIntervalMs ?? REMOTE_WAKE_PROBE_MIN_INTERVAL_MS; + if (args.probeAgeMs >= interval) return 'probe'; + return 'deliver'; +} + +/** + * Append `data` to the pending buffer, dropping the OLDEST whole chunks when the cap + * is exceeded. Returns the resulting buffer — the SAME array reference when the chunk + * was rejected, so the caller can tell the two apart. Pure. + * + * A chunk LARGER than the cap is dropped outright rather than trimmed: one paste is + * one `input` value, and it was never typed character by character, so delivering its + * tail would execute a fragment of it (with the trailing carriage return, if the paste + * had one) — a partial command the user never sent. Keeping the tail is right for + * typing, where the newest bytes are the ones the user just produced, and wrong for a + * chunk that arrived whole. + */ +export function appendBoundedPending( + pending: string[], + data: string, + maxBytes = REMOTE_WAKE_PENDING_MAX_BYTES +): string[] { + if (Buffer.byteLength(data) > maxBytes) return pending; + const next = [...pending, data]; + let total = next.reduce((sum, chunk) => sum + Buffer.byteLength(chunk), 0); + while (next.length > 1 && total > maxBytes) { + total -= Buffer.byteLength(next[0]); + next.shift(); + } + return next; +} + +/** The remote fields the wake flow needs. Structurally satisfied by `SessionRemote`. */ +export interface WakeableRemote { + wakeCommand?: string; + wakeMac?: string; + hostId: string; + label: string; + host: string; + port?: number; + /** SSH jump host (`-J`): the host is reached THROUGH it, never directly. */ + jumpHost?: string; + /** SOCKS5 proxy (`ProxyCommand=nc -X 5 …`): same, the direct address may not even route. */ + socksProxy?: string; + /** Extra `-o KEY=VALUE` options; a `ProxyCommand`/`ProxyJump` in here proxies the host too. */ + extraSshOptions?: string[]; +} + +/** + * Whether the bare TCP probe can answer for this host at all. Pure. + * + * The probe connects straight to `host:port`. A host behind a jump host or a SOCKS + * proxy (the cloudflared case) is reachable ONLY through that proxy, so the direct + * connect fails while ssh works — and every consumer of the verdict would then act on + * a "sleeping" host that is fine: a permanent banner, a create-path gate that hides the + * real ssh error, and (with a wake target) input buffered for the life of the session + * because the readiness poll can never succeed. Such a host is reachability-UNKNOWN: + * the registry never buffers for it, never gates on it, and reports `null` rather than + * `false`. A wake target can still be fired for it, blind. + */ +export function isProbeable(remote: WakeableRemote): boolean { + if (remote.jumpHost || remote.socksProxy) return false; + return !(remote.extraSshOptions ?? []).some((option) => /^\s*proxy(command|jump)\s*=/i.test(option)); +} + +/** + * A resolved wake path for a host. `command` wins over `mac` (an explicit override + * beats the default path), and `null` means the host cannot be woken at all — which + * is what the UI turns into "configure WoL" instead of "wake". + */ +export type WakeTarget = { kind: 'command'; command: string } | { kind: 'mac'; macs: number[][] } | null; + +/** + * Resolve the wake target from host config. Pure. + * + * A malformed `wakeMac` resolves to `null` rather than throwing: the schema + * already rejects one at config time, so this can only be reached with a config + * written by hand, and a broken MAC must not break the input route. + */ +export function resolveWakeTarget(remote: WakeableRemote | undefined): WakeTarget { + if (!remote) return null; + if (remote.wakeCommand) return { kind: 'command', command: remote.wakeCommand }; + if (remote.wakeMac) { + const macs = parseMacList(remote.wakeMac); + if (macs && macs.length > 0) return { kind: 'mac', macs }; + } + return null; +} + +/** + * Parse a comma-separated MAC list into byte arrays. Pure; returns null when any + * entry is malformed (all-or-nothing, so a typo cannot half-arm a host). + */ +export function parseMacList(value: string, maxMacs = MAX_WAKE_MACS): number[][] | null { + const parts = value + .split(',') + .map((part) => part.trim()) + .filter((part) => part.length > 0); + if (parts.length === 0 || parts.length > maxMacs) return null; + const macs: number[][] = []; + for (const part of parts) { + const match = + /^([0-9a-fA-F]{2})[:-]([0-9a-fA-F]{2})[:-]([0-9a-fA-F]{2})[:-]([0-9a-fA-F]{2})[:-]([0-9a-fA-F]{2})[:-]([0-9a-fA-F]{2})$/.exec( + part + ); + if (!match) return null; + macs.push(match.slice(1).map((hex) => Number.parseInt(hex, 16))); + } + return macs; +} + +/** + * Build a Wake-on-LAN "magic packet": six `0xFF` bytes then the MAC repeated 16 + * times. Pure — the shape is asserted byte-for-byte in the tests because a packet + * that is off by one byte simply never wakes anything. + */ +export function buildMagicPacket(mac: number[]): Buffer { + const packet = Buffer.alloc(6 + 16 * 6, 0xff); + for (let repeat = 0; repeat < 16; repeat++) { + Buffer.from(mac).copy(packet, 6 + repeat * 6); + } + return packet; +} + +/** + * The slice of `Session` the wake flow uses — an interface rather than the + * concrete class so the registry is testable without a tmux server. + */ +export interface WakeableSession { + readonly id: string; + readonly remote: WakeableRemote | undefined; + /** COD-108 reattach: respawns the local ssh pane, idempotently attaching the durable remote tmux. */ + reattachRemote(): Promise; + /** Write bytes to the session's pane. `fromUser` marks input a person typed (or an agent sent for them). */ + writeViaMux(data: string, options?: { fromUser?: boolean }): Promise; +} + +/** Injected IO so the registry holds no direct dependency on ssh/net/child_process in tests. */ +export interface RemoteWakeDeps { + /** Cheap reachability probe. Must resolve false (never throw) for a sleeping host. */ + probe(remote: WakeableRemote): Promise; + /** Run the resolved wake target (magic packet or host command). Resolves false on failure. */ + wake(target: NonNullable): Promise; + /** Poll until the woken host accepts connections again. */ + waitUntilReady(remote: WakeableRemote, opts?: { timeoutMs?: number; signal?: AbortSignal }): Promise; + /** Sleep helper (injected for tests). */ + delay(ms: number): Promise; + /** Notify the COD-108 watcher so an exhausted backoff is reset. */ + noteReconnected?(sessionId: string, success: boolean): void; + /** SSE broadcast. */ + broadcast?( + event: 'remote:hostWaking' | 'remote:hostWakeFailed' | 'remote:sessionReconnected', + payload: Record + ): void; + /** Structured diagnostics. */ + log?(message: string): void; + /** + * Resolve the host's CURRENT wake config for a session whose persisted `remote` + * snapshot predates it (or was configured after launch). Called at most once per + * `REMOTE_WAKE_RESOLVE_TTL_MS` per session, and only when the session's own copy + * has no wake target — so a config saved in the UI works without restarting the + * session, without a per-keystroke config read. + */ + resolveRemote?(session: WakeableSession): Promise; +} + +/** Probe freshness for the UI's reachability check (a tab switch is not a hammer). */ +export const REMOTE_WAKE_REACHABILITY_TTL_MS = 5_000; +/** How long a resolved host config is trusted before asking the resolver again. */ +export const REMOTE_WAKE_RESOLVE_TTL_MS = 30_000; + +/** + * What the UI is allowed to offer for a host: how it can be woken, if at all. The + * `'none'` case is what the banner turns into "configure WoL" instead of "wake". + */ +export type WakeConfigured = 'command' | 'mac' | 'none'; + +/** Which wake path a host config provides (mirrors {@link resolveWakeTarget}). Pure. */ +export function wakeConfigured(remote: WakeableRemote | undefined): WakeConfigured { + const target = resolveWakeTarget(remote); + if (!target) return 'none'; + return target.kind; +} + +/** + * Outcome of waking a host for a caller that has NO session yet (the create/attach + * routes). A union rather than a boolean because the cases need different handling: + * `'no-target'` must leave the caller's behavior byte-identical (no probe, no extra + * latency for a host without WoL), `'unprobeable'` likewise (a proxied host, see + * {@link isProbeable} — the probe cannot tell asleep from awake, so nothing is gated on + * it), and only `'failed'` is an error that deserves its own message instead of the + * caller's usual one. + */ +export type HostWakeOutcome = 'no-target' | 'unprobeable' | 'ready' | 'failed'; + +/** + * State key for a host-scoped wake. Prefixed so it can never collide with a session + * id, and keyed on the HOST rather than the case: two cases on one host share a + * single in-flight wake and one probe verdict. Such an entry is tiny (no input + * buffer) and bounded by the number of configured hosts, so it is never dropped. + */ +function hostWakeKey(hostId: string): string { + return `host:${hostId}`; +} + +/** Per-session wake bookkeeping. */ +interface WakeState { + probedAt: number; + reachable?: boolean; + waking: Promise | null; + pending: string[]; + /** Host config resolved after launch (see `RemoteWakeDeps.resolveRemote`). */ + resolvedRemote?: WakeableRemote; + resolvedAt: number; +} + +/** + * Per-session wake state + single-flight wake flow. + * + * One instance per web server (module singleton in the routes file, like the + * signal-wait registry). State is keyed by session id and dropped with the + * session. + */ +export class RemoteWakeRegistry { + private readonly states = new Map(); + /** + * Aborted by {@link stop} on shutdown. Every in-flight readiness poll is holding an + * HTTP request open (the wake route blocks on it), and Fastify's `close()` waits for + * in-flight requests — so without this a restart during a wake sits out the full 90 s + * budget. The same problem `sessionWaits.cancelEverything()` exists for. + */ + private readonly shutdown = new AbortController(); + private stopped = false; + + constructor(private readonly deps: RemoteWakeDeps) {} + + /** Drop a session's state (session closed/killed). The pending buffer goes with it. */ + drop(sessionId: string): void { + this.states.delete(sessionId); + } + + /** + * Resolve every in-flight wake as failed and refuse new ones (server shutdown). + * + * Called from `WebServer.stop()`: an in-flight wake is awaited by a request, and the + * server's own `app.close()` does not abort in-flight requests, so shutdown would wait + * out the poll. Nothing is lost by failing them — the state flush happens earlier in + * `stop()`, and the process is going away. + */ + stop(): void { + this.stopped = true; + this.shutdown.abort(); + } + + /** Whether a wake is currently in flight (diagnostics/tests). */ + isWaking(sessionId: string): boolean { + return this.states.get(sessionId)?.waking != null; + } + + /** Buffered input bytes for a session (diagnostics/tests). */ + pendingBytes(sessionId: string): number { + const state = this.states.get(sessionId); + if (!state) return 0; + return state.pending.reduce((sum, chunk) => sum + Buffer.byteLength(chunk), 0); + } + + /** + * Number of keys with wake state (diagnostics/tests). Pins that a LOCAL session never + * gets an entry: the input gate runs on every keystroke, so an entry per local session + * would be a map the size of the session list, swept only on cleanup. + */ + stateCount(): number { + return this.states.size; + } + + /** Whether this session's host has any wake path configured at all. */ + async hasWakeTarget(session: WakeableSession): Promise { + return resolveWakeTarget(await this._effectiveRemote(session)) !== null; + } + + /** Which wake path is configured (`'none'` when the UI should offer configuration). */ + async wakeConfigured(session: WakeableSession): Promise { + return wakeConfigured(await this._effectiveRemote(session)); + } + + /** + * Reachability for the UI: probe unless a recent result is still fresh. `null` for a + * host the probe cannot reach (see {@link isProbeable}): unknown is not unreachable. + * + * Shares the per-session probe state with the input path on purpose — a fresh + * answer is exactly what the input ladder wants, and an `unreachable` verdict here + * makes the next keystroke buffer + wake instead of vanishing into a stalled pane. + */ + async checkReachable( + session: WakeableSession, + opts: { force?: boolean; ttlMs?: number } = {} + ): Promise { + const remote = await this._effectiveRemote(session); + if (!remote) return true; + // `null`, never `false`: the UI keys the banner on a PROVEN unreachable host. + if (!isProbeable(remote)) return null; + const state = this._state(session.id); + const ttl = opts.force ? 0 : (opts.ttlMs ?? REMOTE_WAKE_REACHABILITY_TTL_MS); + if (Date.now() - state.probedAt >= ttl) { + state.probedAt = Date.now(); + state.reachable = await this.deps.probe(remote); + } + return state.reachable === true; + } + + /** + * Decide + act for one input chunk. + * + * `'deliver'` means the caller writes it as usual (today's path, zero added + * cost). `'buffered'` means the registry took ownership of the bytes: it either + * queued them behind an in-flight wake or started a wake, and will flush them + * in order once the pane is reattached. + */ + async handleInput(session: WakeableSession, data: string): Promise { + const remote = await this._effectiveRemote(session); + // A proxied host can never pass the readiness poll, so buffering for it would hold + // the bytes for the life of the session (reproduced upstream: three inputs, nothing + // written, no reattach). Deliver, as if the feature were off. + if (remote && !isProbeable(remote)) return 'deliver'; + const state = this._state(session.id); + const target = resolveWakeTarget(remote); + const action = decideRemoteInputAction({ + hasWakeTarget: target !== null, + waking: state.waking != null, + probeAgeMs: Date.now() - state.probedAt, + lastReachable: state.reachable, + }); + + if (action === 'deliver') return 'deliver'; + if (action === 'buffer') { + const queued = this._enqueue(session.id, data); + // A buffered verdict with no wake in flight still has to DRIVE a wake (the + // previous one failed and reset the probe state, or the ladder landed here + // directly) — otherwise the bytes would sit in the buffer forever. + if (state.waking == null && target) void this.wake(session); + return queued; + } + + // action === 'probe' — the throttle window elapsed, so one TCP connect is owed. + state.probedAt = Date.now(); + state.reachable = remote ? await this.deps.probe(remote) : true; + if (state.reachable) return 'deliver'; + + const queued = this._enqueue(session.id, data); + void this.wake(session); + return queued; + } + + /** + * Block until the host is reachable and the pane is reattached — the + * send-and-wait path, where the HTTP response stays open anyway and buffering + * would break the wait contract. + */ + async ensureAwake(session: WakeableSession, opts: { force?: boolean; timeoutMs?: number } = {}): Promise { + if (this.stopped) return false; + const remote = await this._effectiveRemote(session); + const target = resolveWakeTarget(remote); + if (!remote || !target) return true; + // A proxied host: the send-and-wait path has nothing to gate on (unknown is not + // asleep), so it delivers; the manual button still wakes, blind (see `wake`). + if (!isProbeable(remote)) return opts.force ? this.wake(session, opts) : true; + const state = this._state(session.id); + // `force` is the manual path (a user pressed "wake"): a cached "reachable" from + // seconds ago must not talk the button out of waking a host that just slept. + if (opts.force || (state.reachable !== false && Date.now() - state.probedAt >= REMOTE_WAKE_PROBE_MIN_INTERVAL_MS)) { + state.probedAt = Date.now(); + state.reachable = await this.deps.probe(remote); + } + if (state.reachable) return true; + // The manual button is pressed from the SAME dashboard the create/attach paths are, + // so it holds its request open under the same reverse proxy — it needs the request + // budget, not the 90 s session default (see REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS). + return this.wake(session, { timeoutMs: opts.timeoutMs }); + } + + /** + * Host-scoped reachability, for a caller that has no session yet (create/attach). + * Shares the per-HOST probe state with {@link ensureHostAwake}, so the probe the + * wake flow just paid for also answers "was that ssh failure really a sleeping + * machine?". Never wakes anything — it is a question, not an action. `null` when the + * question cannot be answered (see {@link isProbeable}). + */ + async checkHostReachable( + remote: WakeableRemote, + opts: { force?: boolean; ttlMs?: number } = {} + ): Promise { + // `null` for a proxied host: callers gate on `=== false` (proven unreachable), so an + // unknown verdict leaves their ordinary error path — "needs tmux" — intact. + if (!isProbeable(remote)) return null; + const state = this._state(hostWakeKey(remote.hostId)); + const ttl = opts.force ? 0 : (opts.ttlMs ?? REMOTE_WAKE_REACHABILITY_TTL_MS); + if (Date.now() - state.probedAt >= ttl) { + state.probedAt = Date.now(); + state.reachable = await this.deps.probe(remote); + } + return state.reachable === true; + } + + /** + * Wake a host for a REQUEST that is waiting on it — the session create/attach + * routes, where there is no session to reattach and no input to buffer yet. + * + * `'no-target'` returns without probing, so a host without WoL config costs + * nothing and behaves exactly as before. Single-flight per host, so a double click + * (or two cases on the same host) sends one packet and shares one readiness poll. + */ + async ensureHostAwake( + remote: WakeableRemote, + opts: { timeoutMs?: number; requestedBy?: string } = {} + ): Promise { + if (!resolveWakeTarget(remote)) return 'no-target'; + // The probe cannot tell a proxied host asleep from awake, and a wake that cannot + // verify readiness would only delay the request by its whole budget. Not gated. + if (!isProbeable(remote)) return 'unprobeable'; + if (this.stopped) return 'failed'; + const state = this._state(hostWakeKey(remote.hostId)); + if (state.waking) return (await state.waking) ? 'ready' : 'failed'; + + state.probedAt = Date.now(); + state.reachable = await this.deps.probe(remote); + if (state.reachable) return 'ready'; + this.deps.log?.(`[RemoteWake] ${remote.label} (${remote.host}) is unreachable — waking it for a new session`); + return (await this.wakeHost(remote, opts)) ? 'ready' : 'failed'; + } + + /** + * Single-flight wake for a host with no session (see {@link ensureHostAwake}). Uses the + * same single-flight `waking` slot the session flow uses — but a DIFFERENT key + * (`host:` vs the session id), so a session wake and a create-path wake for the same + * host are two independent flows rather than one shared poll. Harmless (both are + * user-initiated and the host only wakes once), and keying them together would mean a + * create request joining an unrelated session's wake and inheriting its budget. + */ + private async wakeHost(remote: WakeableRemote, opts: { timeoutMs?: number; requestedBy?: string }): Promise { + const state = this._state(hostWakeKey(remote.hostId)); + if (state.waking) return state.waking; + state.waking = (async (): Promise => { + try { + return await this._wakeAndWait(remote, state, { + timeoutMs: opts.timeoutMs, + forNewSession: true, + requestedBy: opts.requestedBy, + }); + } catch (err) { + // Injected IO is documented not to throw, but a rejected promise here would + // surface as an unhandled rejection AND take the route down with it (the + // session path catches for exactly this reason). A broken wake target must + // fail the wake, never the create route beyond its own error response. + this.deps.log?.(`[RemoteWake] unexpected failure: ${err instanceof Error ? err.message : String(err)}`); + return false; + } finally { + state.waking = null; + } + })(); + return state.waking; + } + + /** + * Single-flight wake: probe-free (the caller already knows the host is down), + * run the wake command, poll for readiness, reattach the pane, flush the buffer. + */ + async wake(session: WakeableSession, opts: { timeoutMs?: number } = {}): Promise { + if (this.stopped) return false; + const remote = await this._effectiveRemote(session); + const target = resolveWakeTarget(remote); + if (!remote || !target) return true; + if (!isProbeable(remote)) return this._wakeBlind(remote, target); + const state = this._state(session.id); + if (state.waking) return state.waking; + + state.waking = (async (): Promise => { + const id = session.id; + try { + const ready = await this._wakeAndWait(remote, state, { sessionId: id, timeoutMs: opts.timeoutMs }); + if (!ready) return false; + + const reattached = await session.reattachRemote(); + if (!reattached) { + this.deps.log?.(`[RemoteWake] ${remote.label} is up but the pane could not be reattached`); + return false; + } + // The reset also clears an EXHAUSTED COD-108 backoff, which otherwise + // never fires again for this session (see remote-reconnect.ts). + this.deps.noteReconnected?.(id, true); + this.deps.broadcast?.('remote:sessionReconnected', { sessionId: id }); + this.deps.log?.(`[RemoteWake] ${remote.label} reattached for session ${id}`); + + await this.deps.delay(REMOTE_WAKE_ATTACH_SETTLE_MS); + await this._flush(state, session); + return true; + } catch (err) { + this.deps.log?.(`[RemoteWake] unexpected failure: ${err instanceof Error ? err.message : String(err)}`); + return false; + } finally { + state.waking = null; + } + })(); + + return state.waking; + } + + /** + * Fire the wake target for a host whose readiness cannot be verified (see + * {@link isProbeable}): no readiness poll (it could never succeed), no reattach (the + * COD-108 watcher owns the pane once ssh works again), no `hostWaking` broadcast (its + * toast promises a wait that does not happen). The caller learns only whether the + * packet/command went out — and a wake IO that throws is a failed wake, never a + * rejected route. + */ + private async _wakeBlind(remote: WakeableRemote, target: NonNullable): Promise { + this.deps.log?.(`[RemoteWake] waking ${remote.label} (${remote.host}) via ${target.kind}, blind: proxied host`); + try { + return await this.deps.wake(target); + } catch (err) { + this.deps.log?.(`[RemoteWake] unexpected failure: ${err instanceof Error ? err.message : String(err)}`); + return false; + } + } + + /** + * Broadcast + run the wake target + wait for SSH. Shared by the session flow (which + * then reattaches and flushes the buffer) and the create/attach flow (which has no + * pane yet). On failure the probe state is reset so the NEXT attempt probes and + * retries instead of trusting a stale "down" verdict forever. + */ + private async _wakeAndWait( + remote: WakeableRemote, + state: WakeState, + opts: { sessionId?: string; timeoutMs?: number; forNewSession?: boolean; requestedBy?: string } = {} + ): Promise { + const target = resolveWakeTarget(remote); + if (!target) return true; + const forWhat = opts.sessionId ? `for session ${opts.sessionId}` : 'for a new session'; + // Routing for multi-user mode (server.ts `deriveSseHint`): a session-scoped event + // reaches its owner, and the create/attach wake has no session yet — so it names + // the requesting user instead, or it would reach admins only. The payload carries + // `hostId`/`label`, which non-admins are not shown elsewhere, so it must not go global. + const scope = opts.sessionId + ? { sessionId: opts.sessionId } + : { forNewSession: true, ...(opts.requestedBy ? { username: opts.requestedBy } : {}) }; + // No `sessionId` for a create-path wake: the toast handler is then the only one + // that acts (a banner for a session that does not exist yet would have no target), + // which is exactly the `forNewSession` distinction the UI renders. + // + // `queuedInput` is true only on the typing path, where bytes are actually held for + // this session. The wake BUTTON and the send-and-wait path hold nothing, so a UI + // that keyed "input is queued until it is back" off "a wake is running" would promise + // something the user can disprove by typing (browser keystrokes go over the + // WebSocket, which never passes through this registry). + this.deps.broadcast?.('remote:hostWaking', { + ...scope, + hostId: remote.hostId, + label: remote.label, + queuedInput: state.pending.length > 0, + }); + this.deps.log?.(`[RemoteWake] waking ${remote.label} (${remote.host}) via ${target.kind} ${forWhat}`); + + const wakeStartedAt = Date.now(); + const woke = await this.deps.wake(target); + if (!woke) { + this.deps.log?.( + `[RemoteWake] wake failed for ${remote.label}: ${target.kind === 'command' ? target.command : 'magic packet'}` + ); + } + + // The request-scoped budget has to cover the WHOLE request, and the wake is part + // of it. A `command` target is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS, so a slow + // one burned 10 s before the readiness poll even started and pushed a wakeCommand + // host's worst case to ~68 s, past the 60 s proxy_read_timeout this budget exists to + // stay under. A magic packet is effectively instant, so this subtracts nothing there. + // Floored at one poll interval so a wake that ate the whole budget still gets one + // probe rather than being declared unreachable without asking. + const wakeElapsedMs = Date.now() - wakeStartedAt; + const readyTimeoutMs = + opts.timeoutMs === undefined + ? undefined + : Math.max(REMOTE_WAKE_READY_INTERVAL_MS, opts.timeoutMs - wakeElapsedMs); + + const ready = await this.deps.waitUntilReady(remote, { + timeoutMs: readyTimeoutMs, + signal: this.shutdown.signal, + }); + if (!ready) { + this.deps.log?.( + `[RemoteWake] ${remote.label} did not come back — ${opts.forNewSession ? 'the session was not started' : 'input stays buffered'}` + ); + this.deps.broadcast?.('remote:hostWakeFailed', { + ...scope, + hostId: remote.hostId, + label: remote.label, + queuedInput: state.pending.length > 0, + }); + state.probedAt = 0; + state.reachable = undefined; + return false; + } + + state.reachable = true; + state.probedAt = Date.now(); + return true; + } + + private _state(sessionId: string): WakeState { + let state = this.states.get(sessionId); + if (!state) { + state = { probedAt: 0, reachable: undefined, waking: null, pending: [], resolvedAt: 0 }; + this.states.set(sessionId, state); + } + return state; + } + + /** + * The host config to act on: the session's own `remote` when it is fresh enough, else a + * freshly resolved one. + * + * The persisted `remote` snapshot is taken at launch, so a wake target configured AFTER + * the session started (e.g. through the banner's config dialog, or by adding `wakeMac` + * to `remote-hosts.json`) is invisible to it. Recovery rehydration (server.ts) covers + * restarts; this covers the live session, and it is why saving the dialog takes effect + * without restarting anything. + * + * ⚠️ The host config wins in BOTH directions, so the resolver is consulted on the TTL + * regardless of whether the session already carries a target. Preferring the snapshot + * whenever it HAD one meant removing a MAC/command in the config (or the dialog) never + * took effect for a running session — the feature stayed on with a target nobody could + * see in the config any more, which is exactly the "host config is authoritative" + * promise failing in the one direction a user can observe. + */ + private async _effectiveRemote(session: WakeableSession): Promise { + // The local-session return comes FIRST, before `_state`: this runs on every input + // chunk (`hasWakeTarget` gates the route), so allocating state here would put an + // entry in the map for every local session the user types in — sessions the feature + // can never apply to, and whose pending buffers would then have to be swept. + if (!session.remote) return undefined; + const state = this._state(session.id); + if (!this.deps.resolveRemote) return state.resolvedRemote ?? session.remote; + if (state.resolvedAt !== 0 && Date.now() - state.resolvedAt < REMOTE_WAKE_RESOLVE_TTL_MS) { + return state.resolvedRemote ?? session.remote; + } + state.resolvedAt = Date.now(); + try { + const resolved = await this.deps.resolveRemote(session); + if (resolved) state.resolvedRemote = resolved; + } catch (err) { + this.deps.log?.( + `[RemoteWake] host config lookup failed for session ${session.id}: ${err instanceof Error ? err.message : String(err)}` + ); + } + return state.resolvedRemote ?? session.remote; + } + + /** Queue a chunk; `'dropped'` when it was over the cap and never entered the buffer. */ + private _enqueue(sessionId: string, data: string): 'buffered' | 'dropped' { + const state = this._state(sessionId); + const next = appendBoundedPending(state.pending, data); + if (next === state.pending) { + // Oversized chunk: dropped whole (see `appendBoundedPending`), so the buffer is + // untouched and nothing is delivered as a fragment. Logged, and reported to the + // route, which answers `dropped:true` — the user's paste is gone and a bare 200 + // could not say so. + this.deps.log?.( + `[RemoteWake] dropped a ${Buffer.byteLength(data)}-byte input chunk for session ${sessionId} — over the ${REMOTE_WAKE_PENDING_MAX_BYTES}-byte wake buffer, and a truncated paste must not be delivered as a fragment` + ); + return 'dropped'; + } + const before = state.pending.reduce((sum, chunk) => sum + Buffer.byteLength(chunk), 0); + const after = next.reduce((sum, chunk) => sum + Buffer.byteLength(chunk), 0); + if (before + Buffer.byteLength(data) > after) { + this.deps.log?.(`[RemoteWake] pending buffer cap reached for session ${sessionId} — oldest input dropped`); + } + state.pending = next; + return 'buffered'; + } + + private async _flush(state: WakeState, session: WakeableSession): Promise { + while (state.pending.length > 0) { + const chunk = state.pending[0]; + // Take the chunk OUT before awaiting the write. Input arriving during the await is + // enqueued by `handleInput` (a wake is still in flight, so it takes the buffer + // path), and `appendBoundedPending` may then drop the OLDEST chunk to stay under + // the cap — which would be this one, already on its way to the pane. Shifting + // afterwards removed the NEXT chunk instead, so the drop-oldest bookkeeping lost a + // chunk that was never written while the log line blamed the one that was. + state.pending = state.pending.slice(1); + // `fromUser`: these bytes came through the input route as a person's prompt, so + // they may name the tab — without it a session whose FIRST prompt was buffered + // through a wake could never be auto-named. + const ok = await session.writeViaMux(chunk, { fromUser: true }).catch(() => false); + if (!ok) { + // Drop the rest, and say so. Retaining it looked safer but was worse: the wake + // still resolves and marks the host reachable, so the NEXT input takes the + // deliver path while the old chunks sit here — to be replayed by the next wake, + // possibly hours later, after everything typed since, and maybe ending in a + // carriage return. Same policy as the oversized paste: gone, with a log line. + const dropped = state.pending.length + 1; + state.pending = []; + this.deps.log?.( + `[RemoteWake] flush failed for session ${session.id} — ${dropped} buffered chunk(s) dropped rather than replayed on a later wake` + ); + return; + } + } + } +} + +// ========== Default IO ========== + +/** + * Under vitest none of this may do real IO (a TCP connect, a child process, a UDP + * broadcast) — mirrors `remote-files.ts`. Every consumer injects its deps + * (`RemoteWakeDeps`, the socket factory); this is what makes that seam non-optional + * instead of a convention the next test can forget. + */ +function assertNotUnderTest(what: string): void { + if (process.env.VITEST) { + throw new Error(`remote-wake: ${what} is disabled under test — inject a fake (RemoteWakeDeps / WakeSocketFactory)`); + } +} + +/** + * Cheap reachability probe: a bare TCP connect to the SSH port. Only meaningful for a + * host the registry deems probeable (see {@link isProbeable}); the registry never asks + * it about a proxied host. + * + * Deliberately NOT an `ssh … true` probe: that opens a full session (auth, + * remote log, process) every throttle window for a question a SYN already + * answers. Any byte count it does move is a few hundred bytes per probe, far + * below the remote idle detector's traffic threshold, so probing cannot keep a + * host awake. + */ +export function probeRemoteHostReachable( + remote: WakeableRemote, + timeoutMs = REMOTE_WAKE_PROBE_TIMEOUT_MS +): Promise { + assertNotUnderTest('the TCP probe'); + const port = remote.port ?? DEFAULT_SSH_PORT; + return new Promise((resolve) => { + let settled = false; + const finish = (value: boolean) => { + if (settled) return; + settled = true; + socket.destroy(); + resolve(value); + }; + const socket = net.connect({ host: remote.host, port }); + socket.setTimeout(timeoutMs, () => finish(false)); + socket.once('connect', () => finish(true)); + socket.once('error', () => finish(false)); + }); +} + +/** + * Run a host's wake command (e.g. a Wake-on-LAN wrapper script). No shell — the + * value is a single executable path, so nothing in it can be interpreted. + * Resolves false on any failure (missing binary, non-zero exit, timeout) rather + * than throwing: a broken wake command must not break the input route. + */ +export function runRemoteWakeCommand(command: string, timeoutMs = REMOTE_WAKE_COMMAND_TIMEOUT_MS): Promise { + assertNotUnderTest('the wake command'); + return new Promise((resolve) => { + let settled = false; + const finish = (value: boolean) => { + if (settled) return; + settled = true; + resolve(value); + }; + let child: ReturnType; + try { + child = spawn(command, [], { stdio: 'ignore' }); + } catch { + finish(false); + return; + } + const timer = setTimeout(() => { + child.kill('SIGKILL'); + finish(false); + }, timeoutMs); + child.once('error', () => { + clearTimeout(timer); + finish(false); + }); + child.once('exit', (code) => { + clearTimeout(timer); + finish(code === 0); + }); + }); +} + +/** + * Send Wake-on-LAN magic packets for every MAC, over UDP to the broadcast address. + * + * This is the whole reason `wakeMac` exists: the common case needs no external + * script. Broadcast on 255.255.255.255 is what the CLI `wakeonlan` does and what the + * NICs here answer to; the socket is closed as soon as the packets are queued, so a + * sleeping host cannot leave a handle behind. Resolves false on any failure (no + * interface to broadcast on, permission) rather than throwing — a broken network + * must not break the wake flow, which reports the failure itself. + */ +export function sendWakePackets( + addresses: number[][], + port = 9, + createSocket: WakeSocketFactory = () => { + assertNotUnderTest('the UDP broadcast'); + return dgram.createSocket('udp4'); + } +): Promise { + if (addresses.length === 0) return Promise.resolve(false); + return new Promise((resolve) => { + const socket = createSocket(); + let settled = false; + const finish = (value: boolean) => { + if (settled) return; + settled = true; + try { + socket.close(); + } catch { + /* already closed */ + } + resolve(value); + }; + socket.once('error', () => finish(false)); + // ⚠️ `setBroadcast` BEFORE the socket is bound fails with EBADF on Linux, and the + // send that follows fails with EACCES — i.e. the packet silently never leaves the + // machine. So the broadcast flag is set in the bind callback, always. (Found by + // the live test: macOS/BSD tolerate the wrong order, Linux does not.) + socket.bind(() => { + try { + socket.setBroadcast(true); + } catch { + finish(false); + return; + } + let pending = addresses.length; + let failed = false; + for (const mac of addresses) { + socket.send(buildMagicPacket(mac), port, '255.255.255.255', (err?: Error | null) => { + if (err) failed = true; + pending--; + if (pending === 0) finish(!failed); + }); + } + }); + }); +} + +/** The `dgram` surface {@link sendWakePackets} uses — injectable so the bind/setBroadcast ORDER is testable. */ +export interface WakeSocket { + bind(callback: () => void): void; + setBroadcast(flag: boolean): void; + send(msg: Buffer, port: number, address: string, callback: (err?: Error | null) => void): void; + close(): void; + once(event: 'error', listener: (err: Error) => void): void; +} + +export type WakeSocketFactory = () => WakeSocket; + +/** Poll the host until it accepts connections again, or the bound is hit. */ +export async function waitUntilRemoteReady( + remote: WakeableRemote, + opts: { + intervalMs?: number; + timeoutMs?: number; + signal?: AbortSignal; + probe?: (remote: WakeableRemote) => Promise; + } = {} +): Promise { + const intervalMs = opts.intervalMs ?? REMOTE_WAKE_READY_INTERVAL_MS; + const timeoutMs = opts.timeoutMs ?? REMOTE_WAKE_READY_TIMEOUT_MS; + const probe = opts.probe ?? probeRemoteHostReachable; + const deadline = Date.now() + timeoutMs; + // Probe immediately: WoL from a warm S3 is fast (~7.5 s measured on this setup), + // and the first poll is what turns "just woke" into a sub-interval response. + for (;;) { + if (opts.signal?.aborted) return false; + if (await probe(remote)) return true; + if (Date.now() + intervalMs > deadline) return false; + // Abortable sleep, so a shutdown does not wait out the current interval either. + await delayOrAbort(intervalMs, opts.signal); + } +} + +/** `delay`, but it also ends the moment `signal` aborts (so cancellation is immediate). */ +function delayOrAbort(ms: number, signal?: AbortSignal): Promise { + if (!signal) return delay(ms); + if (signal.aborted) return Promise.resolve(); + return new Promise((resolve) => { + const done = (): void => { + clearTimeout(timer); + signal.removeEventListener('abort', done); + resolve(); + }; + const timer = setTimeout(done, ms); + signal.addEventListener('abort', done, { once: true }); + }); +} + +const delay = (ms: number): Promise => new Promise((resolve) => setTimeout(resolve, ms)); + +/** + * Production wiring: all IO defaults, overridable for tests. + * + * The readiness poll uses the SAME probe as the rest of the deps, overridden or not. + * Wiring it to the module default instead let a caller that injected `probe` still + * poll the real host during the wait — under vitest, a TCP connect to a production + * address on every shutdown test (which the vitest guard is what finally caught). + */ +export function createDefaultRemoteWakeDeps(overrides: Partial = {}): RemoteWakeDeps { + const probe = overrides.probe ?? probeRemoteHostReachable; + return { + probe, + wake: (target) => (target.kind === 'command' ? runRemoteWakeCommand(target.command) : sendWakePackets(target.macs)), + waitUntilReady: (remote, opts) => waitUntilRemoteReady(remote, { ...opts, probe }), + delay, + ...overrides, + }; +} diff --git a/src/session-env-clamp.ts b/src/session-env-clamp.ts index b9abf3a6..788020a3 100644 --- a/src/session-env-clamp.ts +++ b/src/session-env-clamp.ts @@ -6,13 +6,18 @@ * stripped before the session is built. The create and resume routes are what * this bites on: they clamp what a request asked for. * - * The reboot-restore route calls it as defence in depth, and today it can strip - * nothing. `Session.getEnvOverridesForPersist()` keeps only `CLAUDE_CODE_*` and - * `CLAUDE_CONFIG_DIR` out of a session's overrides, claude's `privilegedEnvKeys` - * are the five `ANTHROPIC_*` names, and that pass admits claude alone — so a - * persisted record cannot carry a clamped key. The call is there for the day the - * persisted set widens. The grant re-resolution that does bite on that path is - * `resolveClaudeModeForUsername`, which recomputes the permission mode. + * The reboot-restore route calls it as defence in depth, and it CAN strip + * something today: `Session.getEnvOverridesForPersist()` keeps only + * `CLAUDE_CODE_*` and `CLAUDE_CONFIG_DIR` out of a session's overrides, and + * claude's `privilegedEnvKeys` now includes both `CLAUDE_CODE_MAX_CONTEXT_TOKENS` + * and `CLAUDE_CONFIG_DIR` (Custom Model Endpoint Profiles, since both can + * redirect a claude session's traffic — see stock.ts's own comment on why they + * are listed despite not needing the clamp for that feature). So a non-granted + * owner's persisted `CLAUDE_CONFIG_DIR` (the per-client-account override, #255) + * is now stripped on reboot-restore, silently returning that session to the + * default Claude account rather than the account it was pointed at. The grant + * re-resolution that ALSO bites on that path is `resolveClaudeModeForUsername`, + * which recomputes the permission mode. * * This lives outside `web/routes` on purpose. The question it answers is about * session privilege rather than about HTTP, and `cron/cron-service.ts` sets the diff --git a/src/session-submit-verifier.ts b/src/session-submit-verifier.ts new file mode 100644 index 00000000..eb39edc8 --- /dev/null +++ b/src/session-submit-verifier.ts @@ -0,0 +1,133 @@ +/** + * @fileoverview Verify that a programmatically sent prompt actually LEFT the composer, + * and press Enter again while it has not. + * + * Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment its + * composer paints but ignores Enter for the first 30 to 50 seconds after it, so the + * `send-keys -l ` + `send-keys Enter` pair `TmuxManager.sendInput()` sends 50 ms + * apart leaves the prompt sitting on the composer with `0 tokens`, and every caller + * that then waits for the turn (send-and-wait, the agent skill, the maintainer bot, + * cron, Ralph) burns its whole timeout on a turn that never started. Measured through + * the input route on 2026-09-19: an Enter at 28 s stranded, one at 51 s submitted. + * + * The rule: after a write that carried a carriage return, read the pane on a short + * schedule; while the LAST composer line (the CLI's own prompt glyph) still holds the + * head of what was sent, send Enter again. An empty composer ends it, and so does a + * composer holding anything else, because that text is the user's or the CLI's, never + * ours. A pane with no composer line at all (a shell, a CLI whose glyph is not + * declared, a direct-PTY session with no pane to read) does nothing: this runs for + * EVERY programmatic sender, so a blind Enter here could confirm a dialog nobody asked + * about. The composer is the last glyph line on purpose: Claude Code echoes a submitted + * prompt with the same glyph higher up in the transcript, so only the last one says + * whether the text was taken. + * + * Pure apart from the injected capture, send and log, so the schedule, the cap and + * every stop condition are unit-tested with fake timers (test/session-submit-verifier.test.ts). + */ +import { stripAnsi } from './utils/index.js'; + +/** + * When to look, counted from the write: 2 s catches the common case (taken) with one + * capture, and the tail reaches 60 s, past twice the longest window measured. Enter is + * re-sent at every check that still finds the prompt, so a 50 s window costs about + * seven Enters and one capture each; a taken prompt costs one capture. + */ +export const SUBMIT_VERIFY_DELAYS_MS: readonly number[] = [ + 2_000, 3_000, 5_000, 5_000, 5_000, 10_000, 10_000, 10_000, 10_000, +]; + +/** How many leading characters of the prompt have to match, whitespace removed. */ +const PROMPT_HEAD_CHARS = 24; + +const compact = (s: string): string => s.replace(/\s+/g, ''); + +/** + * Whether `prompt` is still sitting unsubmitted in the composer of `screen`. + * + * - `true`: the last `glyph` line holds the prompt's head. + * - `false`: the composer is empty (the prompt was taken) or holds other text. + * - `undefined`: no composer line at all; nothing can be said, so nothing is sent. + * + * Whitespace is removed on both sides before comparing, because the composer wraps a + * long prompt onto indented continuation lines and Claude Code draws a no-break space + * after the glyph; `\s` covers that one in JavaScript. + */ +export function promptStillInComposer(screen: string, prompt: string, glyph: string): boolean | undefined { + if (!glyph) return undefined; + const composerLines = stripAnsi(screen) + .split('\n') + .map((l) => l.trim()) + .filter((l) => l.startsWith(glyph)); + if (composerLines.length === 0) return undefined; + const composer = compact(composerLines[composerLines.length - 1].slice(glyph.length)); + if (!composer) return false; + const head = compact(prompt).slice(0, PROMPT_HEAD_CHARS); + return head.length > 0 && composer.startsWith(head); +} + +export interface SubmitVerifierDeps { + /** The rendered pane, or null when there is none to read. */ + capture: () => string | null | undefined; + /** Press Enter once. Failures are swallowed; the next check decides again. */ + sendEnter: () => Promise | unknown; + /** The CLI's composer glyph, resolved at check time (the registry can change). */ + glyph: () => string; + log?: (message: string) => void; + /** Test seam; production uses SUBMIT_VERIFY_DELAYS_MS. */ + delaysMs?: readonly number[]; +} + +/** + * One per session. `arm(text)` starts the schedule for the prompt just sent and + * cancels any earlier one: a newer write owns the composer now, and re-sending Enter + * for an older prompt could submit the newer one early. `cancel()` is for teardown. + */ +export class SubmitVerifier { + private timer: NodeJS.Timeout | null = null; + private generation = 0; + + constructor(private readonly deps: SubmitVerifierDeps) {} + + arm(text: string): void { + this.cancel(); + const gen = this.generation; + const delays = this.deps.delaysMs ?? SUBMIT_VERIFY_DELAYS_MS; + let step = 0; + let elapsed = 0; + let resent = 0; + + const schedule = (): void => { + if (step >= delays.length) return; + const delay = delays[step++]; + elapsed += delay; + this.timer = setTimeout(() => void check(), delay); + this.timer.unref?.(); + }; + const check = async (): Promise => { + this.timer = null; + if (gen !== this.generation) return; + const screen = this.deps.capture(); + if (promptStillInComposer(screen ?? '', text, this.deps.glyph()) !== true) return; + resent++; + this.deps.log?.( + `prompt still in the composer after ${Math.round(elapsed / 1000)}s, re-sending Enter (${resent}/${delays.length})` + ); + try { + await this.deps.sendEnter(); + } catch { + // The next check re-reads the screen and decides again. + } + if (gen !== this.generation) return; + schedule(); + }; + schedule(); + } + + cancel(): void { + this.generation++; + if (this.timer) { + clearTimeout(this.timer); + this.timer = null; + } + } +} diff --git a/src/session.ts b/src/session.ts index 16720e59..fdad22f8 100644 --- a/src/session.ts +++ b/src/session.ts @@ -110,6 +110,7 @@ import { import { DEFAULT_TMUX_HISTORY_LIMIT } from './config/terminal-history.js'; import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js'; import { getCli } from './config/cli-registry/registry.js'; +import { SubmitVerifier } from './session-submit-verifier.js'; import { compileVersionRegex } from './config/cli-registry/patterns.js'; import { resolveSessionCliVersion } from './utils/cli-resolver.js'; import { @@ -502,6 +503,8 @@ export class Session extends EventEmitter { private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read private _trustDialogTimer: NodeJS.Timeout | null = null; // Re-read after a keystroke (see below) + /** Re-sends Enter while a programmatic prompt still sits in the composer (session-submit-verifier.ts). */ + private _submitVerifier: SubmitVerifier | null = null; private _interactiveStartedAt = 0; // When the interactive pane launched (bounds that scan) private _taskTracker: TaskTracker; @@ -3190,6 +3193,9 @@ export class Session extends EventEmitter { } private _clearAllTimers(): void { + // Stop re-sending Enter for a prompt this session will never take now + this._submitVerifier?.cancel(); + this._submitVerifier = null; // Clear the workspace-trust follow-up read if (this._trustDialogTimer) { clearTimeout(this._trustDialogTimer); @@ -3721,7 +3727,10 @@ export class Session extends EventEmitter { const submittedPrompt = this._trackSubmit(data, options); if (this._mux && this._muxSession) { const sent = await this._mux.sendInput(this.id, data); - if (sent) this._emitSubmittedPrompt(submittedPrompt); + if (sent) { + this._emitSubmittedPrompt(submittedPrompt); + this._verifySubmitted(data); + } return sent; } // Fallback to PTY write @@ -3733,6 +3742,39 @@ export class Session extends EventEmitter { return false; } + /** + * Arm the composer check for a write that carried Enter (session-submit-verifier.ts): + * Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints, + * so the pair `sendInput` just sent can leave the text stranded. Only a mux session + * can read its pane, only text can be stranded, and the glyph is the CLI's own. + */ + private _verifySubmitted(data: string): void { + if (!data.includes('\r') || !this._mux?.capturePaneText || !this._muxSession) return; + const text = data.replace(/[\r\n]/g, '').trimEnd(); + if (!text) return; + this._submitVerifier ??= new SubmitVerifier({ + capture: () => + this._isStopped || !this._mux || !this._muxSession + ? null + : this._mux.capturePaneText?.(this._muxSession.muxName), + sendEnter: () => this._mux?.sendInput(this.id, '\r'), + // ⚠ NO fallback glyph here, unlike the screen-reading probe elsewhere in this file. + // Only claude and codex declare a promptGlyph; the other eight modes would fall back + // to claude's `❯`, which is ALSO starship's default shell prompt (and pure's, and + // spaceship's, and p10k lean's). On a shell session the line `❯ npm run build` sits + // on screen for as long as the command runs, promptStillInComposer() reads that as + // "still unsubmitted", and the verifier presses Enter into the running program's + // stdin on its 2s..60s schedule. Mostly a stray blank line; not harmless against a + // y/N prompt, `read -p`, an installer or a pager, where it takes the default. + // promptStillInComposer() returns undefined for an empty glyph, so this makes the + // verifier inert for every CLI that does not declare one, which is what the Claude + // Code 2.1.277 defect it exists for actually calls for. + glyph: () => getCli(this.mode)?.capabilities.workDetect?.promptGlyph ?? '', + log: (m) => console.log(`[Session ${this.id.slice(0, 8)}] ${m}`), + }); + this._submitVerifier.arm(text); + } + /** Current PTY dimensions — used to skip no-op resizes that trigger Ink redraws */ private _ptyCols = 120; private _ptyRows = 40; diff --git a/src/tmux-manager.ts b/src/tmux-manager.ts index c7040a0d..bc02fac9 100644 --- a/src/tmux-manager.ts +++ b/src/tmux-manager.ts @@ -3482,6 +3482,14 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer { { encoding: 'utf-8', timeout: EXEC_TIMEOUT_MS } ) ); + // Report the size the pane was really drawing at. The visible-frame path + // below addresses every row absolutely, so a consumer whose terminal is + // shorter than this piles the overflow rows onto its last line and loses + // the rows it overwrote. The full-history path instead ends in a RELATIVE + // cursor move, which costs it nothing when the two sizes disagree, so the + // geometry is reported there for diagnosis rather than for repair. Only + // the caller can see both sizes, so hand it this one. + if (opts && geometry) opts.capturedGeometry = { cols: geometry.cols, rows: geometry.rows }; if (fullHistory) { // Without geometry there is no cursor move, so fall back to the old trim. diff --git a/src/types/session.ts b/src/types/session.ts index c993666f..20d788b2 100644 --- a/src/types/session.ts +++ b/src/types/session.ts @@ -115,6 +115,25 @@ export interface RemoteHost extends RemoteSshOptions { username: string; port?: number; commands?: Partial>; + /** + * Optional Wake-on-LAN MAC address(es), comma-separated (e.g. + * `04:d9:f5:80:c6:58`). Codeman sends the magic packet itself (UDP port 9 + * broadcast), so the common case needs no external script. A SLEEPING host's + * port-22 probe still fails, which is what triggers the wake — this only + * controls HOW the host is woken. + */ + wakeMac?: string; + /** + * Optional Wake-on-LAN command that powers this host on from SLEEP (e.g. a + * wrapper script like `/home/joe/bin/whuff`). TAKES PRECEDENCE over `wakeMac` + * (an explicit override for hosts that need a router/other-host wake). Absent + * = no wake support and today's behavior exactly. Executed WITHOUT a shell (a + * single executable path, never a command line), only from user input or an + * explicit wake request on a session whose host is unreachable — never from + * the auto-reconnect/boot-recovery path, which would re-wake a host seconds + * after each suspend. + */ + wakeCommand?: string; } export interface RemoteCase { @@ -155,6 +174,13 @@ export interface SessionRemote extends RemoteSshOptions { * session was created elsewhere. Only meaningful when `owned === false`. */ remoteSessionName?: string; + /** + * Wake-on-LAN command carried over from the host config (see `RemoteHost.wakeCommand`) + * so the input route can wake a sleeping host without re-reading the host list. + */ + wakeCommand?: string; + /** Wake-on-LAN MAC address(es) from the host config (see `RemoteHost.wakeMac`). */ + wakeMac?: string; } /** diff --git a/src/web/public/app.js b/src/web/public/app.js index bb40152d..64d29ace 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -219,7 +219,8 @@ const _SSE_HANDLER_MAP = [ // Remote auto-reconnect (COD-108) [SSE_EVENTS.REMOTE_SESSION_RECONNECTED, '_onRemoteSessionReconnected'], [SSE_EVENTS.REMOTE_RECONNECT_EXHAUSTED, '_onRemoteReconnectExhausted'], - + [SSE_EVENTS.REMOTE_HOST_WAKING, '_onRemoteHostWaking'], + [SSE_EVENTS.REMOTE_HOST_WAKE_FAILED, '_onRemoteHostWakeFailed'], // Ralph [SSE_EVENTS.SESSION_RALPH_LOOP_UPDATE, '_onRalphLoopUpdate'], [SSE_EVENTS.SESSION_RALPH_TODO_UPDATE, '_onRalphTodoUpdate'], @@ -549,6 +550,12 @@ class CodemanApp { // repaint-mode CLI pane, where tmux keeps no history of its own). The pull is // refused for those and retried far more slowly — see _maybeRefetchFullHistory. this._fullHistoryRepullUseless = new Set(); + // Sessions where the geometry replay has already been tried and did NOT + // converge, so the pane is one this browser cannot size. Mirrors the Set + // above: `resizeRetry` caps the recursion inside one select, and this is + // what stops a fresh select from paying for the same answer again — see + // the geometry gate in selectSession. + this._geometryRetryUseless = new Set(); this.terminalLoadStates = new Map(); // Map this.respawnStatus = {}; this.respawnTimers = {}; // Track timed respawn timers @@ -1713,6 +1720,25 @@ class CodemanApp { console.error('[SSE] docker container recreated:', err); } }); + // Custom Model Endpoint Profiles: a session's own model got evicted on llama-swap by + // another session's activity, detected AFTER the fact by a periodic server sweep (there + // is no push notification from llama-swap itself) — see detectCustomModelSwapDisplacements + // in custom-model-routes.ts. Global toast rather than a per-tab indicator: the displaced + // session need not be the one currently open, and the whole point is telling the user + // BEFORE they type into it expecting the model they picked. + addListener(SSE_EVENTS.CUSTOM_MODEL_SWAPPED_OUT, (e) => { + try { + const d = e.data ? JSON.parse(e.data) : {}; + this.showToast( + `${d.sessionName || d.sessionId}'s model (${d.previousModel}) was swapped out on llama-swap by another ` + + `session — currently loaded: ${d.currentlyLoadedModel}. Sending a message there will reload it.`, + 'warning', + { duration: 0 } + ); + } catch (err) { + console.error('[SSE] custom model swapped out:', err); + } + }); // Multi-user admin: live-refresh whichever admin views (panel/Users tab) are open. addListener(SSE_EVENTS.ADMIN_USERS_CHANGED, () => { window.codemanAdmin?.onUsersChanged?.(); @@ -1818,6 +1844,9 @@ class CodemanApp { _onInit(data) { _crashDiag.log(`INIT: ${data.sessions?.length || 0} sessions`); this.handleInit(data); + // Start the remote-host reachability poller even if no session switch follows + // (a page loaded with the remote tab already active) — see host-wake-ui.js. + this._ensureHostWakePoller?.(); } _onSessionCreated(data) { @@ -5783,28 +5812,7 @@ class CodemanApp { if (ta) ta.dispatchEvent(new CompositionEvent('compositionend', { data: '' })); } } catch {} - // Flush local echo text to PTY before switching tabs. - // Send as a single batch (no Enter) so it lands in the session's readline - // input buffer — avoids "old text resent on Enter" and overlay render bugs. - // Track flushed length so _render() offsets the overlay correctly even before - // the PTY echo arrives in the terminal buffer. - if (this.activeSessionId) { - const echoText = this._localEchoOverlay?.pendingText || ''; - // Include buffer-detected flushed text (from Tab completion, etc.) - // so it's preserved across tab switches. - const existingFlushed = this._localEchoOverlay?.getFlushed()?.count || 0; - const existingFlushedText = this._localEchoOverlay?.getFlushed()?.text || ''; - if (echoText) { - this._sendInputAsync(this.activeSessionId, echoText); - } - const totalOffset = existingFlushed + echoText.length; - if (totalOffset > 0) { - if (!this._flushedOffsets) this._flushedOffsets = new Map(); - if (!this._flushedTexts) this._flushedTexts = new Map(); - this._flushedOffsets.set(this.activeSessionId, totalOffset); - this._flushedTexts.set(this.activeSessionId, existingFlushedText + echoText); - } - } + this._flushLocalEchoTo(this.activeSessionId); this._localEchoOverlay?.clear(); // Predictions are ephemeral + already sent: nothing to save/restore // across a tab switch (unlike the buffer overlay's setFlushed machinery) @@ -5819,6 +5827,45 @@ class CodemanApp { } } + /** + * Hand the local-echo overlay's unsent text to `sessionId` before anything + * clears it, and record what has now been flushed so `_render()` offsets the + * overlay correctly even before the PTY echo comes back. + * + * On a touch device the characters the user has typed live ONLY here until + * Enter — they have never reached the PTY — so whoever clears the overlay + * owes them a flush first. It is sent as one batch with no Enter, so it lands + * in the session's readline buffer rather than submitting a line the user has + * not finished. + * + * ⚠️ The session is a PARAMETER because the two callers are looking at + * different ones. `_cleanupPreviousSession` flushes to the tab being left, + * which is still `activeSessionId` when it runs. The `forceReload` branch in + * `selectSession` flushes to the tab being RELOADED, and must do it before it + * nulls `activeSessionId`: reading the field after that null is what silently + * dropped the text, since the guard here then saw no session and the + * unconditional `clear()` that follows took the characters with it. + * @param {string|null} sessionId + */ + _flushLocalEchoTo(sessionId) { + if (!sessionId) return; + const echoText = this._localEchoOverlay?.pendingText || ''; + // Include buffer-detected flushed text (from Tab completion, etc.) + // so it's preserved across tab switches. + const existingFlushed = this._localEchoOverlay?.getFlushed()?.count || 0; + const existingFlushedText = this._localEchoOverlay?.getFlushed()?.text || ''; + if (echoText) { + this._sendInputAsync(sessionId, echoText); + } + const totalOffset = existingFlushed + echoText.length; + if (totalOffset > 0) { + if (!this._flushedOffsets) this._flushedOffsets = new Map(); + if (!this._flushedTexts) this._flushedTexts = new Map(); + this._flushedOffsets.set(sessionId, totalOffset); + this._flushedTexts.set(sessionId, existingFlushedText + echoText); + } + } + _resetTerminalForReplay() { this.terminal.reset(); this.terminal.write('\x1b[3J\x1b[H\x1b[2J'); @@ -6093,6 +6140,13 @@ class CodemanApp { this._loadBufferQueue = null; this._terminalRefreshOwner = null; this._chunkedWriteGen = (this._chunkedWriteGen || 0) + 1; + // Anything typed but not yet submitted lives in the local-echo overlay and + // has never reached the PTY. `_cleanupPreviousSession` below flushes it, + // but only for a session it can still see, and the null on the next line + // hides this one from it. Flush first or the characters are cleared + // unread. The geometry replay re-enters here with no gesture behind it, + // so on a touch device this fires while the user is still typing. + this._flushLocalEchoTo(sessionId); this.activeSessionId = null; } // Focus terminal SYNCHRONOUSLY before any await — iOS Safari only honors @@ -6159,6 +6213,9 @@ class CodemanApp { // bar (issue #262). Also disarms a one-shot Ctrl left over from the tab we // just left, so it can never fire against the session we just opened. if (typeof KeyboardAccessoryBar !== 'undefined') KeyboardAccessoryBar.refreshForActiveSession(); + // Remote-host reachability banner: only meaningful for a remote session, so this + // also clears it when the newly active tab is local. + this.refreshHostWakeBanner?.(sessionId); // Restore flushed offset AND text IMMEDIATELY so backspace/typing work during // the async buffer load. Without this, the offset is 0 during the @@ -6272,6 +6329,10 @@ class CodemanApp { // sendResize is a no-op on the server when dims haven't changed, so // calling it every tab switch is cheap. const dimsChanged = await this.sendResize(sessionId, { forceHttp: true }).catch(() => false); + // The size the capture below will be taken against. The debounced resize + // handler can move the terminal again while the load runs, so this is a + // recorded value rather than a later read of `_lastResizeDims`. + const dimsAtCapture = this.getTerminalDimensions?.(); if (this._isStaleSelect(selectGen)) { this._clearTerminalLoadState(sessionId, selectGen); return; @@ -6520,6 +6581,84 @@ class CodemanApp { // annoyance that disappear on the user's next keypress; data loss is not // acceptable. Do NOT re-introduce Ctrl+L here. this.sendResize(sessionId); + // sendResize fits synchronously before its first await, so this reads the + // size that survived the load rather than the one the capture was taken + // at. The two differ whenever the terminal was still settling. + const dimsAfterLoad = this.getTerminalDimensions?.(); + // Only a visible-frame capture positions its rows absolutely, and only + // that frame can be damaged by a terminal of the wrong size. A `full=1` + // body is linear scrollback closed by a RELATIVE cursor move + // (`formatCursorRestore`), which is relative precisely so the browser's + // row count need not match the pane's, and a `history` body is the byte + // stream, which carries no row alignment to protect. Replaying either at + // a different size repairs nothing, and the full-history replay costs a + // second whole-scrollback capture to learn that. Since the first select + // of every non-shell session per page takes the full-history path, an + // ungated comparison fires most often on the one response it cannot help. + const framePositionsRowsAbsolutely = data.source === 'mux-visible'; + // `mux-visible` is necessary but not sufficient: when the `display-message` + // cursor query fails, `capturePaneBuffer` skips the snapshot repaint and + // returns the raw capture, and the route still labels a non-empty body + // `mux-visible`. That body positions nothing and reports no geometry, so a + // size that moved during such a load has nothing to repair, and replaying + // would buy a second capture, a reset plus chunked rewrite, a dropped + // WebSocket and a discarded xterm snapshot for it. The two comparisons + // below already stand down on an absent field; this one has to as well. + const sizeMovedUnderLoad = + framePositionsRowsAbsolutely && + Number.isFinite(data.captureRows) && + !!dimsAtCapture && + !!dimsAfterLoad && + (dimsAfterLoad.cols !== dimsAtCapture.cols || dimsAfterLoad.rows !== dimsAtCapture.rows); + // A capture positions every row absolutely, so a pane taller than this + // terminal writes its overflow rows onto the last line and loses the rows + // it overwrote. A pane WIDER than this terminal damages the same frame a + // second way: `formatPaneSnapshot` paints each row out to the pane's own + // width, so a narrower browser wraps every painted row, and the wrap on + // the last one scrolls the whole frame up by a row. Both happen when the + // capture wins a race against the resize meant to precede it, which is + // what the retry below repairs. + // + // It also happens when `Session.resize` DECLINED the resize, which it does + // for a small viewport while a desktop viewport's size claim is live. The + // retry cannot repair that one: it re-sends the same declined resize and + // captures the same too-tall pane. `resizeRetry` stops it after the one + // extra attempt, and the frame is shown as-is. Repairing that case means + // changing who owns the pane size, which is a policy question this does + // not touch. What the flag does buy there is that the client can SEE the + // mismatch at all, which it previously could not. + // + // An ABSENT field is not a fit. It means the capture reported no geometry + // at all, so nothing was positioned and there is nothing to repair. + const capturedTallerThanTerminal = + framePositionsRowsAbsolutely && + Number.isFinite(data.captureRows) && + data.captureRows > (this.terminal?.rows || 0); + const capturedWiderThanTerminal = + framePositionsRowsAbsolutely && + Number.isFinite(data.captureCols) && + data.captureCols > (this.terminal?.cols || 0); + // The retry replays at `dimsAfterLoad`, so it can only change what is on + // screen if the pane was drawing at some OTHER size. When the reported + // geometry already IS that size, the second pass captures the identical + // frame and pays a full reload to do it: another fetch, another + // `_resetTerminalForReplay()` and chunked rewrite (a visible re-flash), + // and, because it goes through `forceReload`, a dropped and reopened + // WebSocket plus a deleted xterm snapshot. + // + // That equality is the signature of a CLAMP rather than a race. + // `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does + // not, so a terminal narrower than 40 columns or shorter than 10 rows + // reports a pane permanently bigger than itself, and every select would + // retry without ever converging. A race never produces this equality: its + // whole premise is that the pane was still at the size we asked it to + // leave. The other non-converging case, `Session.resize` declining a + // small viewport while a desktop claim is live, does not produce it + // either — that pane sits at the DESKTOP's size — so it still costs the + // one capped attempt, and stopping it needs the pane-ownership policy + // this does not touch. + const captureMatchesRequestedSize = + !!dimsAfterLoad && data.captureCols === dimsAfterLoad.cols && data.captureRows === dimsAfterLoad.rows; // Defer secondary panel updates so they don't block the main thread // after terminal content is already visible. @@ -6620,6 +6759,67 @@ class CodemanApp { this._clearTerminalLoadState(sessionId, selectGen); _crashDiag.log(`SELECT_DONE: ${selectDoneMs.toFixed(0)}ms`); console.log(`[CRASH-DIAG] selectSession DONE: ${sessionId.slice(0,8)} in ${selectDoneMs.toFixed(0)}ms`); + // Remember whether the replay was worth it, because `resizeRetry` only + // caps the recursion INSIDE one select and says nothing about the next + // one. A pane this browser cannot size — one whose resize `Session.resize` + // declines while a desktop claim is live, or one a second tmux client is + // also holding — reports the same mismatch on every select, so without a + // memo the diagnosis is paid for again on every tab switch, forever: two + // fetches per select rather than one. Each extra pass costs a second + // `capture-pane`, which is `execSync` and blocks the server's event loop, + // plus a reset and chunked rewrite, a discarded snapshot and cache entry, + // and a dropped and reopened WebSocket. + // + // A retry pass that STILL does not fit is the proof, since the retry ran + // at the size that stuck and the pane ignored it. Geometry that fits + // clears the memo, so a pane that becomes sizeable again (the desktop tab + // closes, the claim goes idle) is repaired on the next select. The race + // case is untouched: it converges on its first attempt, so it never + // reaches the branch that latches. + const capturedGeometryFits = + framePositionsRowsAbsolutely && + Number.isFinite(data.captureRows) && + !capturedTallerThanTerminal && + !capturedWiderThanTerminal; + if (capturedGeometryFits) { + this._geometryRetryUseless?.delete(sessionId); + } else if (options?.resizeRetry && (capturedTallerThanTerminal || capturedWiderThanTerminal)) { + (this._geometryRetryUseless ||= new Set()).add(sessionId); + } + // What is on screen was drawn for a geometry this terminal does not have. + // Replaying once against the size that stuck is the only thing that + // repairs it: SIGWINCH reaches the CLI only on a real size change, and + // the pane is already at its final size, so no redraw is coming. + // `resizeRetry` caps this at one attempt, so two competing fits cannot + // trade replays forever. + if ( + (sizeMovedUnderLoad || capturedTallerThanTerminal || capturedWiderThanTerminal) && + !captureMatchesRequestedSize && + !this._geometryRetryUseless?.has(sessionId) && + !options?.resizeRetry && + !this._isStaleSelect(selectGen) + ) { + _crashDiag.log( + `RESIZE_RETRY: capture ${data.captureCols}x${data.captureRows} vs terminal ` + + `${this.terminal?.cols}x${this.terminal?.rows}` + + (sizeMovedUnderLoad ? ' (size moved under load)' : '') + ); + // Re-arm the full-history pull ONLY if this pass actually used one, so + // the retry replays the same content at the geometry that stuck. A pass + // that took the bounded tail must retry on the tail too: clearing the + // flag unconditionally would UPGRADE a tab switch into a fresh + // multi-megabyte scrollback capture it never asked for. + // + // UNREACHABLE as written, and kept for the invariant rather than the + // branch. A `useFullHistory` pass sends `full=1`, and the route answers + // `full=1` with `mux-full-history` or `history`, never `mux-visible` + // (see the source ladder in session-routes.ts), so the gate above + // already rules out every pass that consumed the flag. Do not read this + // line as evidence that a page load retries: it does not, and the test + // suite pins that it does not. + if (useFullHistory) this._fullHistoryLoaded.delete(sessionId); + await this.selectSession(sessionId, { auto: true, forceReload: true, resizeRetry: true }); + } } catch (err) { if (this._isLoadingBuffer) this._finishBufferLoad(bufferLoadOwner); this._restoringFlushedState = false; diff --git a/src/web/public/constants.js b/src/web/public/constants.js index be9780b0..6521ccc1 100644 --- a/src/web/public/constants.js +++ b/src/web/public/constants.js @@ -806,6 +806,71 @@ function decideAutoCopy({ enabled, text, lastCopied, pending } = {}) { return 'copy'; } +// The text a copy should put on the clipboard, given xterm's raw selection. +// Pure: the caller reads the selection and decides the mode, this transforms. +// +// xterm hands back whole screen ROWS, and its own trim only drops cells that +// were never written to. A full-screen TUI writes real spaces across the part +// of a row it is not using, so that padding counts as content and rides along +// to the clipboard: measured against Claude Code in a 282-column pane, single +// lines arrived carrying 138 trailing spaces. Native terminals trim it on copy +// (Windows Terminal, iTerm2 and GNOME Terminal all do), decideAutoCopy above +// already calls a wall of spaces "never what the gesture meant", and +// _selectTouchSelectionLine already treats those cells as padding. This is that +// same rule for the mouse and keyboard paths, which never had it. +// +// ⚠ Trailing padding ONLY. A shared LEADING indent is deliberately left alone, +// and this note is here so the idea is not re-derived: it was built, measured +// and dropped before merge. Removing the longest leading run every selected row +// shares looks like the mirror image of the trailing trim and is not, because +// no native terminal does it and the transform cannot tell a TUI's margin from +// content that is genuinely indented. Measured over 401 445 three-row windows +// across 1 010 tracked files in this repo, it fired on 73% of them: 92% inside +// a YAML workflow, 76% over `git log` output, 48% in a TypeScript source file. +// No width threshold separates the two, because they are the same widths: a +// live Claude Code pane's own margins measure 2 and 5 columns while the most +// common non-TUI shared run is 4, sitting between them. +// +// The asymmetry that settles it is in the failure modes. A wrong trailing trim +// costs nothing. A wrong dedent silently deletes information that was on the +// screen, with no signal to the user and nothing in the clipboard to hint at +// it, and it is wrong on `git log` bodies, on indented code read out of `cat` +// (semantic in Python), on `git diff` context rows where the leading space is +// the marker, and on stack traces. +// +// ⚠ It also cannot be made consistent cheaply. Whether the first row joins the +// measurement depended on the mousedown COLUMN, which the user never sees, so +// one block of three rows produced three different clipboard results; and the +// flag read `getSelectionPosition().start`, which is the mousedown anchor that +// xterm never normalises, so dragging UP through a block read it off the bottom +// row. If it is ever revisited, the one qualification that measured clean is +// painted trailing padding (a full-screen TUI writes real spaces across every +// row; a shell pane leaves those cells never-written, so xterm trims them): +// zero false positives over all 401 445 windows. It still mangles a `git log` +// body sitting inside an agent's own gutter, which is why it was not taken now. +function cleanCopiedSelection(text) { + if (typeof text !== 'string' || !text) return ''; + // Split on \n and leave any \r in place: xterm joins rows with \r\n on + // Windows, and the clipboard should keep the endings xterm chose. + // Scanned rather than matched. A selection can run to the 50 000-row + // scrollback ceiling, and `/[ \t]+(\r?)$/` is QUADRATIC on a line whose spaces + // are followed by any non-space character, which is what right-aligned or + // centred TUI content looks like: the engine retries the run from every + // whitespace position and backtracks over it. Measured over 50 000 rows with a + // 280-column run, that regex took 2.9s against 1.3ms for the scan below, and a + // 2 000-column run took 16s. It is also the faster of the two on an ordinary + // padded row. A length is returned rather than a trimmed string so a + // \r-terminated line costs no substring either. + const trimEnd = (line) => { + let end = line.length; + if (end > 0 && line[end - 1] === '\r') end--; + let cut = end; + while (cut > 0 && (line[cut - 1] === ' ' || line[cut - 1] === '\t')) cut--; + return cut === end ? line : line.slice(0, cut) + line.slice(end); + }; + return text.split('\n').map(trimEnd).join('\n'); +} + if (typeof window !== 'undefined') { window.WEBGL_FALLBACK = WEBGL_FALLBACK; window.evaluateWebGLLongTaskTrip = evaluateWebGLLongTaskTrip; @@ -854,6 +919,9 @@ if (typeof window !== 'undefined') { decide: decideAutoCopy, MAX_CHARS: AUTO_COPY_MAX_CHARS, }; + window.CodemanCopySelection = { + clean: cleanCopiedSelection, + }; window.CodemanTerminalFont = { DEFAULT_STACK: TERMINAL_FONT_DEFAULT_STACK, resolve: resolveTerminalFontFamily, @@ -1057,6 +1125,9 @@ const SSE_EVENTS = { REMOTE_SESSION_DROPPED: 'remote:sessionDropped', REMOTE_SESSION_RECONNECTED: 'remote:sessionReconnected', REMOTE_RECONNECT_EXHAUSTED: 'remote:reconnectExhausted', + // Wake-on-LAN from user input on a sleeping remote host + REMOTE_HOST_WAKING: 'remote:hostWaking', + REMOTE_HOST_WAKE_FAILED: 'remote:hostWakeFailed', // Ralph SESSION_RALPH_LOOP_UPDATE: 'session:ralphLoopUpdate', @@ -1094,6 +1165,9 @@ const SSE_EVENTS = { APPROVAL_UPDATED: 'approval:updated', APPROVAL_RESOLVED: 'approval:resolved', + // Custom Model Endpoint Profiles + CUSTOM_MODEL_SWAPPED_OUT: 'custom-model:swapped-out', + // Subagents (Claude Code background agents) SUBAGENT_DISCOVERED: 'subagent:discovered', SUBAGENT_UPDATED: 'subagent:updated', diff --git a/src/web/public/host-wake-ui.js b/src/web/public/host-wake-ui.js new file mode 100644 index 00000000..0a1f76e2 --- /dev/null +++ b/src/web/public/host-wake-ui.js @@ -0,0 +1,440 @@ +/** + * @fileoverview Remote-host wake-on-LAN: the "host unreachable" banner + its config dialog. + * + * A sleeping remote host does not fail loudly. The local tmux pane runs `ssh`, and when + * the machine suspends, that ssh child stalls: `tmux send-keys` still SUCCEEDS, so typed + * input disappears with no error and the pane looks alive. The server side + * (`src/remote-wake.ts`) buffers input and wakes the host when the user types; this + * module makes the state VISIBLE and gives it a button, which is what turns "why is + * nothing happening" into one click. + * + * Behavior: + * - Asks `GET /api/sessions/:id/reachability` for the ACTIVE remote session only: + * once when the tab is activated (a user action), and every `POLL_MS` while the tab + * is visible ONLY for a host with a wake target. The timer is the one thing here that + * is not user-driven, and each poll is a TCP connect to the host — the same + * timer-driven traffic invariant #2 rejects keepalives for: it cannot wake a host, + * but it can keep an activity-based suspend timer from firing. So a host Codeman + * could not wake anyway is never polled on a timer. A host behind a jump host or + * SOCKS proxy (`probeable: false`) is never polled at all: the probe cannot reach + * it, so its answer would only ever be a false "asleep". The endpoint shares the + * server's probe cache with the input path, so opening the tab also primes the + * wake path. + * - Unreachable + a configured wake target → "Wake" button → `POST /api/sessions/:id/wake` + * (which wakes, waits, reattaches the pane and flushes buffered input). + * - Unreachable + NO wake target → "Configure WoL" → `#wakeConfigModal`, a small form + * for this host's MAC/command that saves via `PUT /api/remote-hosts/:id`. The server + * re-resolves host config while the session is live, so saving takes effect without + * restarting the session. + * - SSE (`remote:hostWaking`, `remote:hostWakeFailed`, `remote:sessionReconnected`) + * keeps the banner in sync while a wake is running. + * + * @mixin Extends CodemanApp.prototype via Object.assign + * @dependency app.js (CodemanApp class, this.sessions, this.activeSessionId, showToast) + * @dependency constants.js (SSE_EVENTS — the remote:hostWaking / remote:hostWakeFailed names) + * @loadorder 12.2 — loaded after session-ui.js, before webview-tabs.js + */ + +const HOST_WAKE_POLL_MS = 30_000; + +Object.assign(CodemanApp.prototype, { + /** Per-tab banner state (single active session at a time). */ + _hostWake: null, + /** The page-wide poller interval (created once, see `_ensureHostWakePoller`). */ + _hostWakeTimer: null, + + /** Fresh state for a session we just switched to. */ + _hostWakeState() { + return { + sessionId: null, + /** Last reachability answer, or null before the first poll. */ + reachable: null, + /** 'command' | 'mac' | 'none' — what the banner action should do. */ + wakeConfigured: 'none', + host: '', + label: '', + /** + * False for a host the server's probe cannot reach (behind a jump host or SOCKS + * proxy): its reachability is unknown, so there is no banner and no polling. + */ + probeable: true, + /** True between clicking Wake and the answer coming back. */ + waking: false, + /** + * True only when the server is actually holding bytes for this session (the typing + * path buffers them). Browser keystrokes go over the WebSocket, which never passes + * through the wake registry — so the Wake BUTTON must not claim input is queued. + */ + queuedInput: false, + /** Set when the last wake attempt or poll failed. */ + error: '', + }; + }, + + /** + * Entry point from the session switcher — called for every active session, remote or + * not, so it must be cheap and must clear the banner for local sessions. + * + * ⚠️ The POLLER is page-wide and independent of this call on purpose: a session + * switch is not the only way the active tab changes (boot restore, a page loaded with + * the tab already active, and `selectSession`'s own early return for the tab you are + * already on), and the banner must not depend on any single one of those paths + * running — that is exactly how it could silently never appear. + */ + refreshHostWakeBanner(sessionId) { + this._ensureHostWakePoller(); + const state = this._hostWake; + if (state && state.sessionId && state.sessionId !== sessionId) this._hostWake = null; + this._hostWakeTick(); + }, + + /** Create the page-wide poller once (interval + a visibility wake-up). */ + _ensureHostWakePoller() { + if (this._hostWakeTimer) return; + this._hostWakeTimer = setInterval(() => this._hostWakeTick({ periodic: true }), HOST_WAKE_POLL_MS); + document.addEventListener('visibilitychange', () => { + if (document.visibilityState === 'visible') this._hostWakeTick({ periodic: true }); + }); + }, + + /** + * One poller tick: resolve the ACTIVE session, reset the banner when it changed, and + * ask the server. No-op while the page is hidden (a background tab must not poll). + * + * `periodic` marks the timer (and the visibility wake-up) as opposed to a tab + * activation: a periodic tick polls only a host with a wake target, see the module + * comment. The activation poll is what still offers "Configure WoL" for a sleeping + * host that has none — one connect, on a user action. + */ + _hostWakeTick({ periodic = false } = {}) { + if (typeof document !== 'undefined' && document.visibilityState === 'hidden') return; + const sessionId = this.activeSessionId; + const session = sessionId && this.sessions ? this.sessions.get(sessionId) : null; + if (!sessionId || !session || !session.remote) { + // Render unconditionally: `refreshHostWakeBanner` clears `_hostWake` BEFORE + // calling this tick, so a guard here would skip the repaint and leave the + // banner up on every chat (the clear and the repaint must not be coupled to + // whoever cleared the state). Idempotent — with a null state it just hides. + this._hostWake = null; + this._renderHostWakeBanner(); + return; + } + let state = this._hostWake; + let fresh = false; + if (!state || state.sessionId !== sessionId) { + fresh = true; + state = this._hostWake = this._hostWakeState(); + state.sessionId = sessionId; + state.host = session.remote.host || ''; + state.label = session.remote.label || 'Remote host'; + // Text from the session payload first (instant, no round trip), corrected by the + // poll — a session whose wake config was added after launch only knows it after + // the server resolves host config. The kind matters: the payload can say WHICH + // path is configured, so a command-only host is not mislabelled 'mac' until the + // first poll lands. + state.wakeConfigured = session.remote.wakeMac ? 'mac' : session.remote.wakeCommand ? 'command' : 'none'; + // Known from the payload already: a proxied host is not probeable (the server + // says so too, on every answer), so not even the activation poll is worth a + // round trip whose verdict could only be a wrong "asleep". + state.probeable = !(session.remote.jumpHost || session.remote.socksProxy); + this._renderHostWakeBanner(); + } + if (!state.probeable) return; + if (periodic && !fresh && state.wakeConfigured === 'none') return; + this._pollHostReachability(); + }, + + /** One reachability check for the active remote session. */ + async _pollHostReachability(force = false) { + const state = this._hostWake; + if (!state || !state.sessionId) return; + const sessionId = state.sessionId; + try { + const res = await fetch(`/api/sessions/${encodeURIComponent(sessionId)}/reachability${force ? '?force=1' : ''}`); + const data = await res.json(); + if (!data.success) return; + // The tab may have changed while this was in flight. + if (this._hostWake !== state || state.sessionId !== sessionId) return; + // `reachable` is `null` (unknown, not unreachable) for a host the probe cannot + // reach — only a PROVEN `false` may raise the banner. + state.reachable = data.data.reachable !== false; + if (data.data.probeable === false) state.probeable = false; + state.wakeConfigured = data.data.wakeConfigured || 'none'; + if (data.data.host) state.host = data.data.host; + if (data.data.label) state.label = data.data.label; + if (state.reachable) { + state.waking = false; + state.error = ''; + } + this._renderHostWakeBanner(); + } catch { + /* A failed poll is not a state change: leave the banner as it was. */ + } + }, + + /** Draw the banner from `_hostWake`. */ + _renderHostWakeBanner() { + const state = this._hostWake; + const banner = this.$('hostWakeBanner'); + const text = this.$('hostWakeBannerText'); + const detail = this.$('hostWakeBannerDetail'); + const action = this.$('hostWakeBannerAction'); + if (!banner || !text || !action) return; + + const visible = Boolean(state && state.sessionId && state.reachable === false); + banner.hidden = !visible; + if (!visible) return; + + const hasTarget = state.wakeConfigured !== 'none'; + const target = state.label || state.host || 'Remote host'; + if (state.waking) { + text.textContent = `Waking ${target} …`; + } else if (state.error) { + text.textContent = `${target} did not wake up`; + } else { + text.textContent = `${target} is not reachable`; + } + if (detail) { + detail.textContent = state.waking + ? state.queuedInput + ? 'input is queued until it is back' + : 'waiting for the host to come back' + : hasTarget + ? `ssh ${state.host}` + : 'no wake-on-LAN configured'; + } + // After a FAILED wake the only useful next step is fixing the target (wrong MAC, + // host moved NIC, command gone) — otherwise a configured-but-broken host would be + // stuck behind a button that keeps failing with no way to edit it. + const offerConfig = !hasTarget || Boolean(state.error); + action.textContent = state.waking ? 'Waking …' : offerConfig ? 'Configure WoL' : 'Wake'; + action.disabled = state.waking; + }, + + /** Banner button: wake the host, or open the setup dialog when nothing is configured. */ + hostWakeAction() { + const state = this._hostWake; + if (!state || !state.sessionId || state.waking) return; + if (state.wakeConfigured === 'none' || state.error) { + this.openWakeConfigDialog(); + return; + } + this.wakeRemoteHost(); + }, + + /** POST the manual wake for the active session and follow the result. */ + async wakeRemoteHost() { + const state = this._hostWake; + if (!state || !state.sessionId) return; + const sessionId = state.sessionId; + state.waking = true; + // The button path holds nothing: whatever the user typed went into the stalled pane + // over the WebSocket and is gone. Saying otherwise is a promise the next keystroke + // disproves. + state.queuedInput = false; + state.error = ''; + this._renderHostWakeBanner(); + try { + const res = await fetch(`/api/sessions/${encodeURIComponent(sessionId)}/wake`, { method: 'POST' }); + const data = await res.json(); + if (this._hostWake !== state || state.sessionId !== sessionId) return; + state.waking = false; + if (!data.success) { + // The ROUTE is the authority on whether a target is configured, so ask it again + // (`/reachability` reports `wakeConfigured`) rather than pattern-matching the + // error message: the message is prose, and the code is generic (`INVALID_INPUT` + // covers "Not a remote session" too). + state.error = data.error || 'Wake failed'; + this._renderHostWakeBanner(); + await this._pollHostReachability(true); + return; + } + state.reachable = data.data.reachable !== false; + state.wakeConfigured = data.data.wakeConfigured || state.wakeConfigured; + if (state.reachable) { + this.showToast(`${state.label || 'Remote host'} is awake`, 'success'); + } else { + state.error = 'timeout'; + } + this._renderHostWakeBanner(); + } catch (err) { + if (this._hostWake !== state) return; + state.waking = false; + state.error = err && err.message ? err.message : 'Wake failed'; + this._renderHostWakeBanner(); + } + }, + + /** + * Why the host could not be read. In multi-user mode `GET /api/remote-hosts` returns + * `[]` to a non-admin, so "Remote host not found" would blame a config the user simply + * is not allowed to see — the save is admin-only, and that is what it should say. + */ + _wakeConfigUnavailableMessage() { + const me = window.__codemanUser || {}; + return me.multiUser && me.role !== 'admin' ? 'Wake-on-LAN configuration is admin-only' : 'Remote host not found'; + }, + + /** Open the small WoL dialog for the banner's host, pre-filled from the host config. */ + async openWakeConfigDialog() { + const state = this._hostWake; + const session = state && state.sessionId && this.sessions ? this.sessions.get(state.sessionId) : null; + if (!session || !session.remote) return; + const hostId = session.remote.hostId; + const label = this.$('wakeConfigHostLabel'); + const mac = this.$('wakeConfigMac'); + const command = this.$('wakeConfigCommand'); + const status = this.$('wakeConfigStatus'); + if (!mac || !command) return; + + mac.value = session.remote.wakeMac || ''; + command.value = session.remote.wakeCommand || ''; + if (label) label.textContent = session.remote.label || hostId; + if (status) status.textContent = ''; + this._wakeConfigHostId = hostId; + const modal = this.$('wakeConfigModal'); + if (modal) modal.classList.add('active'); + + // Read the saved host so the dialog shows what is actually persisted (the session + // payload may predate a change made in another tab). + try { + const res = await fetch('/api/remote-hosts'); + const data = await res.json(); + const hosts = data.success ? data.data : []; + const host = Array.isArray(hosts) ? hosts.find((item) => item.id === hostId) : null; + if (host && this._wakeConfigHostId === hostId) { + mac.value = host.wakeMac || ''; + command.value = host.wakeCommand || ''; + } else if (!host && this._wakeConfigHostId === hostId && status) { + // Say it up front rather than only when Save fails. + status.textContent = this._wakeConfigUnavailableMessage(); + } + } catch { + /* The form is already usable from the session payload. */ + } + }, + + closeWakeConfigDialog() { + const modal = this.$('wakeConfigModal'); + if (modal) modal.classList.remove('active'); + this._wakeConfigHostId = null; + }, + + /** Save MAC/command for the host, then re-check whether the session can wake now. */ + async saveWakeConfig() { + const hostId = this._wakeConfigHostId; + const mac = this.$('wakeConfigMac'); + const command = this.$('wakeConfigCommand'); + const status = this.$('wakeConfigStatus'); + const save = this.$('wakeConfigSave'); + if (!hostId || !mac || !command) return; + + const macValue = mac.value.trim(); + const commandValue = command.value.trim(); + if ( + macValue && + !/^[0-9a-fA-F]{2}([:-][0-9a-fA-F]{2}){5}(\s*,\s*[0-9a-fA-F]{2}([:-][0-9a-fA-F]{2}){5})*$/.test(macValue) + ) { + if (status) status.textContent = 'MAC must look like 04:d9:f5:80:c6:58 (comma-separated for several).'; + return; + } + if (commandValue && /\s/.test(commandValue)) { + if (status) status.textContent = 'The wake command must be a single executable path (no arguments).'; + return; + } + + if (save) save.disabled = true; + if (status) status.textContent = 'Saving …'; + try { + const listRes = await fetch('/api/remote-hosts'); + const listData = await listRes.json(); + const hosts = listData.success ? listData.data : []; + const host = Array.isArray(hosts) ? hosts.find((item) => item.id === hostId) : null; + if (!host) throw new Error(this._wakeConfigUnavailableMessage()); + // PUT takes the whole host (schema-validated), so send back everything we know and + // only replace the wake fields. `undefined` drops the key entirely. + const payload = { + ...host, + wakeMac: macValue || undefined, + wakeCommand: commandValue || undefined, + }; + const res = await fetch(`/api/remote-hosts/${encodeURIComponent(hostId)}`, { + method: 'PUT', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify(payload), + }); + const data = await res.json(); + if (!data.success) throw new Error(data.error || 'Save failed'); + this.showToast('Wake settings saved', 'success'); + this.closeWakeConfigDialog(); + // The server re-resolves host config for live sessions, so the banner can offer + // the wake right away — probe fresh instead of waiting out the poll interval. + await this._pollHostReachability(true); + } catch (err) { + if (status) status.textContent = err && err.message ? err.message : 'Save failed'; + } finally { + if (save) save.disabled = false; + } + }, + + /** + * SSE `remote:hostWaking` — a wake is running (ours or one started by typing). + * + * ⚠️ The ONLY definition of this handler: `panels-ui.js` must not define it too. + * Both mix into `Codeman.prototype` and this file loads later, so a second copy + * would be silently shadowed (the guard in `sse-dispatch-table.test.ts` sees that a + * handler exists, not that two modules claim the same name). The toast is + * deliberately UNCONDITIONAL — a wake can start for a background session (input on + * a non-active tab) where there is no banner to update. + */ + _onRemoteHostWaking(data) { + const label = data && data.label ? data.label : 'Remote host'; + // A create-path wake (the user pressed Run / Attach) has no session yet, so + // nothing is queued behind it — the wording has to say what actually happens. + const forNewSession = Boolean(data && data.forNewSession); + // Only the typing path buffers bytes; the wake button and the send-and-wait path + // hold none, and a browser keystroke never reaches the registry at all. + const queuedInput = Boolean(data && data.queuedInput); + // Long enough to cover the wake + attach (~10s measured on a warm S3), and it + // is replaced by `remote:sessionReconnected` the moment the pane is back. + this.showToast( + forNewSession + ? `Waking ${label} … the session starts when it is back` + : queuedInput + ? `Waking ${label} … input is queued` + : `Waking ${label} … waiting for it to come back`, + 'info', + { duration: 12000 } + ); + const state = this._hostWake; + if (!state || !data || state.sessionId !== data.sessionId) return; + state.waking = true; + state.queuedInput = queuedInput; + state.error = ''; + if (data.label) state.label = data.label; + this._renderHostWakeBanner(); + }, + + /** SSE `remote:hostWakeFailed` — the host did not come back in time. */ + _onRemoteHostWakeFailed(data) { + const label = data && data.label ? data.label : 'Remote host'; + const forNewSession = Boolean(data && data.forNewSession); + const queuedInput = Boolean(data && data.queuedInput); + this.showToast( + forNewSession + ? `${label} did not wake up — no session was started` + : queuedInput + ? `${label} did not wake up — queued input is still held` + : `${label} did not wake up`, + 'error', + { duration: 15000 } + ); + const state = this._hostWake; + if (!state || !data || state.sessionId !== data.sessionId) return; + state.waking = false; + state.queuedInput = queuedInput; + state.error = 'timeout'; + state.reachable = false; + this._renderHostWakeBanner(); + }, +}); diff --git a/src/web/public/i18n.js b/src/web/public/i18n.js index bee8fcab..af1554ff 100644 --- a/src/web/public/i18n.js +++ b/src/web/public/i18n.js @@ -286,6 +286,34 @@ 'Prompt sent': '提示已发送', 'Inserted, press Enter in the terminal to send': '已插入,在终端中按 Enter 发送', 'Could not reach the session': '无法连接到会话', + 'Custom model endpoints': '自定义模型端点', + 'Point a harness at your own OpenAI-compatible server (llama.cpp, vLLM, DGX Spark, Azure AI Foundry, OpenRouter) instead of its native cloud backend. When on, the Run menu offers an extra entry per harness that supports it, per saved endpoint.': + '让工具指向您自己的兼容 OpenAI 服务器(llama.cpp、vLLM、DGX Spark、Azure AI Foundry、OpenRouter),而非其原生云端后端。开启后,"运行"菜单会为每个支持此功能的工具、每个已保存的端点新增一个条目。', + 'Enable custom model endpoints': '启用自定义模型端点', + 'Adds a per-endpoint entry to the Run menu for every harness that can redirect to one.': + '为每个可重定向到端点的工具,在"运行"菜单中添加对应条目。', + 'No endpoints yet. Add one below to point a harness at a local or cloud OpenAI-compatible server.': + '暂无端点。请在下方添加一个,以便将工具指向本地或云端的兼容 OpenAI 服务器。', + Discover: '发现模型', + '+ Add endpoint': '+ 添加端点', + 'Add endpoint': '添加端点', + Id: 'ID', + 'Short, stable — used in URLs, never shown to the CLI.': '简短且固定 — 用于 URL,不会展示给 CLI。', + Label: '标签', + 'Base URL': '基础 URL', + 'API key': 'API 密钥', + 'Optional. Left blank on edit keeps the existing key.': '可选。编辑时留空将保留现有密钥。', + 'Auth header': '认证请求头', + 'Never send both — some servers hang indefinitely.': '切勿同时发送两者 — 部分服务器会因此无限期挂起。', + 'Authorization: Bearer (default)': 'Authorization: Bearer(默认)', + 'api-key header (Azure)': 'api-key 请求头(Azure)', + 'Default model': '默认模型', + 'What the Run-menu picker applies for this endpoint. Discover models first.': + '运行菜单选择器会为此端点应用该模型。请先发现可用模型。', + 'Custom Endpoints': '自定义端点', + 'Choose a model': '选择模型', + 'That endpoint no longer exists': '该端点已不存在', + 'No models discovered for this endpoint yet': '此端点尚未发现任何模型', 'Subagent Options': '子智能体选项', 'Enable Tracking': '启用跟踪', 'Active Tab Only': '仅活动标签页', @@ -521,6 +549,7 @@ 'Respawn Blocked': '重生已阻止', 'Task Complete': '任务完成', 'Copied to clipboard': '已复制到剪贴板', + 'Nothing to copy': '没有可复制的内容', // Terminal touch-selection bar (long-press to select). The bar is a sibling of // `.xterm`, not a descendant, so SKIP_SELECTOR does not cover it and these apply. Copy: '复制', diff --git a/src/web/public/index.html b/src/web/public/index.html index cbe9a480..bb53fef7 100644 --- a/src/web/public/index.html +++ b/src/web/public/index.html @@ -213,6 +213,18 @@ + + + @@ -668,6 +680,14 @@ + + + +
+ + + + + + + + + + + + +
+

Custom model endpoints

synced
+

Point a harness at your own OpenAI-compatible server (llama.cpp, vLLM, DGX Spark, Azure AI Foundry, OpenRouter) instead of its native cloud backend. When on, the Run menu offers an extra entry per harness that supports it, per saved endpoint.

+
+
+
+ Enable custom model endpoints + Adds a per-endpoint entry to the Run menu for every harness that can redirect to one. +
+ +
+ + +
+
@@ -2878,6 +3013,11 @@ Optional. Leave blank for the default port 22. +
+ + + Optional. Comma-separated for several NICs. Codeman sends the magic packet itself so a sleeping host can be woken from the session banner. +
@@ -2886,6 +3026,11 @@
Advanced SSH
+
+ + + Optional override for the MAC above (takes precedence). A single executable path, run without a shell — use it when the host needs a router/other machine to send the packet. +
@@ -3481,6 +3626,36 @@ text is set via value/textContent only: predictor output derives from observable (injectable) content, and the explicit click here is the security boundary (nothing is ever auto-sent). --> + + +