# Architecture invariants Implementation detail extracted from `CLAUDE.md` so that file stays small enough to load into every session cheaply. Nothing here was rewritten: these are the original paragraphs, verbatim, including the version history and PR references that explain *why* each rule exists. `CLAUDE.md` keeps the short form of each rule plus a pointer to the section here. Read the pointer first; come here when you need the mechanism, the file names, or the history behind a constraint. --- ## Network binding and instance isolation ### Default bind, and the non-loopback warning path **Default bind is loopback-only; non-loopback without a password starts but warns** — since COD-29 (PR #107) the web server defaults to `--host 127.0.0.1` (was `0.0.0.0`). As of **0.9.0** binding a non-loopback host (`--host`/`-H`/`CODEMAN_HOST`) without `CODEMAN_PASSWORD` **no longer refuses to start — it starts and prints a loud warning** listing the fixes (set `CODEMAN_PASSWORD`, bind loopback + tunnel/`tailscale serve`, or `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` to acknowledge → terser note). Host classification is `isLoopbackBindHost()` in `network-auth-policy.ts`; the warn-vs-start logic is in `server.ts` `start()`; flags wired in `cli.ts`. ⚠️ Operational note: the production systemd unit runs `node dist/index.js web --https` with no `--host`, so it binds **localhost only** — reach it remotely via `tailscale serve`/tunnel to `127.0.0.1`, or add `Environment=CODEMAN_HOST=0.0.0.0` + `Environment=CODEMAN_PASSWORD=…` to `~/.config/systemd/user/codeman-web.service`. A loopback bind is reachable through a same-host tunnel (cloudflared/tailscale → `127.0.0.1`) but NOT by a browser hitting the box's LAN IP. Auth user defaults to `admin`. **Installer note** (1.8.x, `install.sh`): interactive installs now PROMPT for the binding, defaulting to LAN (`0.0.0.0`) with a required password prompt (skipping the password needs an explicit confirm and prints a loud warning); non-interactive installs keep loopback unless `CODEMAN_HOST` is preset, and re-runs/updates preserve the EXISTING binding (`read_existing_binding()` parses the current systemd unit / launchd plist). The server binary's own default is unchanged. **Full model: `docs/security-architecture.md`.** ### Instance isolation and the multi-instance attach danger **Instance isolation / multi-instance attach danger** — data dir (`~/.codeman`) and tmux socket (`tmux -L codeman`) are PROCESS-WIDE and shared by every Codeman on the machine, derived from `CODEMAN_INSTANCE` via `src/config/instance.ts` (`getDataDir()`/`dataPath()`/`DEFAULT_TMUX_SOCKET`). ⚠️ A 2nd instance on the SAME socket **discovers and attaches PTYs to the first instance's live sessions** (`tmux -L codeman attach-session …`), resizing/mutating them — `$HOME` isolation is NOT enough (tmux is system-global). To run two instances, give each a distinct `CODEMAN_INSTANCE` (scopes BOTH dir+socket: `~/.codeman-` + `-L codeman-`), or set `CODEMAN_TMUX_SOCKET` + `CODEMAN_DATA_DIR` individually. **`CODEMAN_INSTANCE` defaults to empty = the production layout (`~/.codeman`, `-L codeman`, port 3000)**, so this branch is safe to ship to master without disturbing existing installs. To run THIS beta alongside prod, launch with `scripts/run-beta.sh` (`CODEMAN_INSTANCE=beta` + `CODEMAN_PORT=5000`) — it never collides with prod's data dir/socket/port. Any new `~/.codeman/...` path MUST go through `dataPath()`, never `join(homedir(), '.codeman', …)`. ## Session launch modes ### External CLI modes (OpenCode, Codex, Gemini, Antigravity) **External CLI modes (OpenCode, Codex, Gemini, Antigravity)**: `isExternalCliMode()` in `session.ts` (`mode === 'opencode' || 'codex' || 'gemini' || 'antigravity'`) gates Claude-specific behavior — Ralph tracker, BashToolParser, token/CLI-info parsing, and ❯-prompt readiness detection are all skipped (these CLIs render their own TUIs; readiness = output stabilization instead). All four modes **require tmux — no direct PTY fallback** — because secrets are injected via `tmux setenv` (socket-scoped `${this.tmux()} setenv`, never on the spawn command line): OpenCode gets `OPENCODE_CONFIG_CONTENT` etc., Codex gets `OPENAI_API_KEY`/`CODEX_API_KEY`/`CODEX_HOME` (`setCodexEnvVars`), Gemini gets `GEMINI_API_KEY`/`GOOGLE_API_KEY`/`GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI` etc. (`setGeminiEnvVars`, all in `tmux-manager.ts`). Codex specifics: command built by `buildCodexCommand()` (`--model`, `resume `, `--dangerously-bypass-approvals-and-sandbox` from the `codexConfig` payload / `codexDangerouslyBypassApprovals` app setting; `renderMode` is schema-coerced to `'hybrid'`, the only supported mode). Gemini specifics: command built by `buildGeminiCommand()` (`--skip-trust` always, `--approval-mode ` defaulting to `yolo` for parity with Claude's `--dangerously-skip-permissions`, `--model`, `--resume` from the `geminiConfig` payload); availability via `GET /api/gemini/status` — session/quick-start routes fail with `OPERATION_FAILED` + install hint (`npm install -g @google/gemini-cli`) when missing. Codex AND Gemini export `COLORTERM=truecolor` + unset `NO_COLOR` (other modes unset `COLORTERM`); Gemini joins `isAltScreenStripMode()` (Codex/Claude/Gemini are Ink TUIs that repaint inline → strip alt-screen/`3J` so scrollback survives). Codex availability via `GET /api/codex/status`. Antigravity specifics: command built by `buildAntigravityCommand()` (`--model`, `--conversation ` resume, `--dangerously-skip-permissions` from the `antigravityConfig` payload); availability via `GET /api/antigravity/status` — routes fail with `OPERATION_FAILED` + install hint (`curl -fsSL https://antigravity.google/cli/install.sh | bash`) when missing. Unlike the other three it is NOT an npm package (standalone binary, `~/.local/bin/agy`), which is why `docker/agent.Dockerfile` installs it with its own `--dir /usr/local/bin` step rather than in the `npm install -g` line, and why it does NOT join `isAltScreenStripMode()`. Frontend: run-mode dropdown → `runCodex()`/`runGemini()` in `session-ui.js` ("Run CX"/"Run GM" labels), App Settings → Codex CLI tab; Respawn/Ralph options are Claude-only, so session options open on the Summary tab for external CLI sessions. ⚠️ `run*()` MUST unwrap the `{success,data}` envelope (`(await res.json()).data.available` / `data.data.sessionId`) — reading the raw shape silently breaks the run. Tests: `test/run-mode-ui.test.ts` + `test/gemini-mode.test.ts` (vm-sandbox harness, no real DOM). **Codex input path (issues #218/#219/#220/#222)**: the local-echo overlay is **DISABLED for codex-mode sessions** (`_updateLocalEchoState` in terminal-ui.js, same branch as shell). Codex's composer is interactive per keystroke: typing "/" pops a live-filtering command picker (#222 was "picker never appears" because the "/" sat in the overlay until Enter), the composer grows/rewraps as it fills (#220: a long typed prompt existed ONLY in the overlay DOM, so codex never grew the composer), arrows and Ctrl+Backspace edit server-side state (#218: arrows were forwarded to an EMPTY composer while the typed text sat pending; the `\x08` control-char flush then left the overlay stateless so `\x7f` was swallowed as "nothing to remove"), and pastes arrive bracketed (#219: `terminal.paste()` wraps in `\x1b[200~..201~`, which the multi-byte-ESC branch forwarded WITHOUT flushing pending text, so the paste landed before it). The shared overlay branch (claude/gemini/opencode still buffer) gained three fixes: bracketed pastes flush pending text first, composer nav keys (`isComposerNavKey` allowlist in `CodemanTerminalInput` — arrows/Home/End/Delete/PgUp/PgDn incl. modifiers, deliberately excluding DA/CPR/DSR query responses) flush and hand the session to **pass-through** (plain PTY echo until Enter/Ctrl+C, because after cursor movement the append-only overlay cannot track edits), and a backspace that finds no overlay state is FORWARDED instead of swallowed. ⚠️ **Codex drops keystrokes that arrive in the same PTY read as a bracketed paste** (upstream `bottom_pane/paste_burst.rs` holds rapid chars for paste classification; verified against codex 0.147.0 by writing `hello\x1b[200~PASTED\x1b[201~` into the tmux client PTY in one write → composer shows only `PASTED`, while a 100ms gap yields `helloPASTED`), so the flush sends the typed text immediately and delays the paste sequence by 80ms — the same two-phase shape as the Enter branch's delayed `\r`. Related protocol fact: xterm.js sends `0x08` for Ctrl+Backspace, which codex's keymap binds to delete-ONE-char (`ctrl(Char('h'))`); real word-delete needs the kitty CSI-u encoding (`\x1b[127;5u`), which xterm.js 6.0.0 cannot emit (kitty support lands in 6.1.0-beta) — an upstream limitation, not a Codeman bug. E2E technique: codex 0.147 reaches its composer with any dummy key in `$CODEX_HOME/auth.json` (`{"OPENAI_API_KEY":"sk-test-..."}`), so a real TUI can be driven headlessly (envOverrides `CODEX_HOME` rides the `CODEX_*` allowlist) without real credentials. Tests: `test/local-echo-codex-gating.test.ts` (vm harness: nav-key classifier truth table, per-mode gating, flush helper, pass-through routing). ### Remote sessions over SSH **Remote sessions (SSH)**: Sessions can run the agent inside a durable `tmux -L codeman-remote new-session -A` **on a remote host** so it survives the SSH drop (COD-104), and can also **discover + attach** to `codeman-*` sessions another Codeman launched there — attached (`owned:false`) sessions **detach, never kill** on tab close (COD-105). **Shared/collaborative** (COD-106): remote set-options are scoped per-session (never `-g`) and `window-size latest` lets multiple clients attach the same session at different viewports without clamping to the smallest; a client count surfaces a "shared · N" badge. **Auto-reconnect** (COD-108): a bounded-backoff watcher re-establishes a dropped remote session's local ssh pane and reattaches the still-running durable remote tmux (kill-switch `remoteAutoReconnect`, default ON); the pure pieces (backoff schedule, per-session reconnect state, `decideReconnect` eligibility) live in `src/remote-reconnect.ts` (tests: `test/remote-auto-reconnect.test.ts`), while `tmux-manager.ts` owns the live pane probe + timers. Owned sessions propagate `kill-session` to the remote on close; non-owned never do. ⚠️ Command-injection surface (COD-107): all ssh command lines flow through the single shell-safe `buildSshConnectionArgs()` — every user field (`-J jumpHost`, `-i identity`, `-o`) is `shellescape`d; never hand-build an ssh line elsewhere. Full design: `docs/remote-sessions.md`. ### Remote SSH cases **Remote SSH cases** (COD-94/#145): cases can point at a **remote host** (`~/.codeman/remote-hosts.json` + `remote-cases.json` via `src/remote-hosts.ts`; CRUD under `/api/cases` — cases route file). A remote session launches a LOCAL tmux pane running `ssh ` that creates a durable REMOTE tmux session on a **dedicated socket** `-L codeman-remote` with name `codeman-ssh-` — deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so a Codeman instance on the target host never adopts it; no `-g` global tmux options are set remotely. `remotePath`/`identityFile` are schema-guarded against shell injection (backticks/`$` rejected — same approach as `extraSshOptions`); remote tmux availability is probed via `checkRemoteTmuxAvailable()` in quick-start (ssh args carry `-o ConnectTimeout=10`). Remote claude defaults to `exec claude --dangerously-skip-permissions`; per-host `commands.*` override. Session kill best-effort kills the remote tmux too. `SessionState.remote`/`MuxSession.remote` round-trip through recovery (`restoreMuxSessions` passes `remote` back into the Session constructor). ⚠️ Run flows must route remote cases through `POST /api/quick-start` (which resolves the remote case and skips LOCAL CLI availability gates) — `POST /api/sessions` stat-validates `workingDir` locally and has no `caseName`. `envOverrides`/`effort`/`modelOverride`/`codexConfig`/`geminiConfig` are rejected for remote quick-starts (not silently dropped). UI: Create Case modal → Remote tab. Tests: `test/remote-hosts.test.ts`, `test/remote-ssh-options.test.ts`. ### Docker cases **Docker cases** (shipped 1.4.0; user guide `docs/docker-cases.md`, design `docs/docker-cases-plan.md`): a case can point at a **container** instead of a local/remote path, and any of the five CLI backends runs INSIDE it. Like remote-SSH, it is a **LOCATION OVERLAY on cases, never a sixth `SessionMode`** (`SessionMode` is unchanged). Storage `~/.codeman/docker-hosts.json` + `docker-cases.json` via `src/docker-hosts.ts` (direct mirror of `remote-hosts.ts`: `readDockerHosts`/`readDockerCases`, `toSessionDocker`, `dockerDisplayPath`, and the PURE builders `buildDockerBaseArgs`/`buildDockerCreateArgs`/`containerApiUrl`/`hostGatewayAlias`/`dockerConfigHash`). CRUD `/api/docker-hosts` + `/api/cases/docker-link`, plus **one-click** `/api/cases/docker-quickcreate` (Create New "Run in Docker" checkbox → case folder in `CASES_DIR` + auto-provisioned shared `default` host + auto-start a session inside; an expandable Template picker Small/Medium/Large/GPU or any override creates a per-case `q-` host), and export/import (`/api/docker-cases/:name/export`, `/api/docker-cases/import`, `GET/DELETE /api/docker-exports`) — all in `case-routes.ts`. Run flows route through `POST /api/quick-start` like remote (session-routes.ts docker branch, skips LOCAL CLI-availability gates). **Launch model**: exactly one long-lived container **per case** (`codeman-case-`, PID1 `sleep infinity` under `--init`); a LOCAL tmux pane runs `docker exec -it` into a **durable in-container tmux** on dedicated socket `-L codeman-docker`, session `codeman-dkr-` (deliberately fails `SAFE_MUX_NAME_PATTERN` so a Codeman running INSIDE the container never adopts it, exactly like remote's `codeman-ssh-`). Builders `buildDockerLaunchCommand`/`buildDockerKillCommand` in `tmux-manager.ts` (image-check → `docker inspect||create` → start → exec, all idempotent). The container is **shared by all sessions of the case**: `buildDockerKillCommand` kills ONLY that session's in-container tmux session, NEVER `docker stop` while siblings remain; `docker rm -f` happens only on case-delete (plus an instance-scoped boot reaper keyed on the `codeman.instance` label). **Two-layer durability/resume** (the central design point): (1) Codeman-PROCESS restart with the container still up → `tmux new-session -A` reattaches the SAME live agent (paneCommand ignored); (2) container stop/reboot/OOM → inner tmux is gone, so the re-run pane command resumes the conversation from the bind-mounted transcript: claude mode pins a DETERMINISTIC conversation id via `claudeDockerPaneCommand()` (`tmux-manager.ts`) — fresh launch `claude --session-id || claude --resume ` (a duplicate `--session-id` exits 1 "already in use", so the fallback RESUMES after a container stop; verified CLI behavior), explicit resume `--resume || --session-id ` so a stale id never dead-panes (leading `exec ` is stripped — an exec'd first branch could never fall back); codex `resume ` / gemini `--resume` keep `appendResumeFlag`. The resume id rides `resumeSessionId` through create/respawn options and persists on `DockerCase.lastClaudeSessionId` via `persistDockerCaseClaudeSessionId()` (written at quick-start launch, and again on hook/last-response conversation-id adoption so post-`/clear` switches track; seeded back when `resumeOnStart`, default true); `-A` makes the pane command self-selecting (inert on reattach, active only when tmux was re-created). **Config drift** (`dockerConfigHash` → `codeman.confighash` label): quick-start compares via `checkDockerConfigDrift()` and REFUSES a drifted launch with `CONFLICT`; the UI confirm calls `POST /api/docker-cases/:name/recreate` (refused while case sessions are live) which `docker rm -f`s so the next launch recreates with the new config — host config edits actually take effect. **Workspace** is a REAL host dir bind-mounted at the SAME absolute path (mirror, `dst==src`), so `Session.workingDir = hostWorkspacePath` keeps file-routes/attachments/watchers on real host bytes AND the in-container transcript projHash matches the host so subagent/workflow correlation (and thus resume-id capture) works; `resolveMuxAttachCwd` returns `/tmp` for docker (the local pane only runs `docker exec`). **Creds** arrive commit-safe and ISOLATED (1.4.1; replaced the whole-dir RW mounts that let in-container CLIs write refreshed tokens/state back to the host): shared RW across the boundary is ONLY what host-side reads/resume need (`~/.claude/projects` transcripts; codex `sessions/` + `history.jsonl` for response-viewer/`codex resume`); everything else is SEEDED (RO mount, copied into container HOME once at launch via `[ -e ] || cp`; the container refreshes its own copy and never writes back): `~/.claude.json` is merged through `buildSeamlessClaudeConfig()` (forces `hasCompletedOnboarding` + theme + workspace trust, so no login wizard/theme picker/trust prompt inside the container), plus `.claude/{.credentials.json,settings.json,stats-cache.json}`, plus whole-dir seeds for `~/.gemini`/`~/.config/{gcloud,opencode}` (`resolveDockerClaudeArtifacts`/`resolveDockerCredentialArtifacts` in `docker-hosts.ts`). Bind mounts are physically excluded from `docker commit`, so exports stay secret-free; API-key CLIs get exec-time NAME-ONLY `--env OPENAI_API_KEY` (no `=value`); the SEALED profile is `mountCredentials:false` + `network:none`. NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket. **Hardening** on every create: `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit`, `--memory`==`--memory-swap`, non-root via `--user :0` (Linux, GID 0 for writable HOME) / `--userns=keep-id` (podman rootless) / baked uid (Docker Desktop), `--pull=never`, `--init`. Base image `codeman/agent:base` is BUILT LOCALLY from `docker/agent.Dockerfile` (node22 + tmux + claude/codex/gemini/opencode, OpenShift arbitrary-uid HOME, `C.UTF-8` locale so tmux/Ink render real box-drawing glyphs; Codeman also sets `LANG`/`LC_ALL` at run time for containers built before that line) via `scripts/build-agent-image.mjs` OR **auto-built on first use** (1.4.1: `ensureAgentBaseImage()` in `docker-hosts.ts`; idempotent + concurrency-safe, only the DEFAULT image ref is ever auto-built, `--pull=never` stays absolute; build output streams over SSE `docker:imageBuildStarted`/`imageBuildProgress`/`imageBuildComplete`/`imageBuildFailed`, and quick-create returns `imageBuilding:true` while the first launch awaits the gate); tmux-in-image is a HARD gated prerequisite (`checkDockerTmuxAvailable`), never a silent bare-exec fallback. **Hooks + model**: the workspace-scaffolding block DOES run for docker (writes `.claude/settings.local.json` + the CLAUDE.md scaffold into the real host dir), so `modelOverride` works via `settings.local.json` — it is a `QuickStartSchema` field applied for local AND docker quick-starts (`updateCaseModel`), sent by the frontend docker run path (the one deliberate difference from remote, which rejects it); `effort`/`envOverrides`/`codexConfig`/`geminiConfig`/`openCodeConfig` stay rejected. In-container hook curls hit `containerApiUrl(process.env.CODEMAN_API_URL, engine)` (swaps ONLY the hostname to the gateway alias, preserving scheme+port so prod HTTPS still works); the host guard allowlists both `host.docker.internal`/`host.containers.internal` (`DOCKER_HOST_GATEWAY_ALIASES` in `network-auth-policy.ts`). ⚠️ On a **loopback-only** bind (the prod default) a container cannot reach 127.0.0.1, so in-container hooks fire ONLY when `CODEMAN_DOCKER_BRIDGE_HOOKS=1` — an opt-in SECOND listener on the docker bridge gateway (`_startDockerBridgeHooksListener` in `server.ts`; gateway auto-detected via `detectDockerBridgeGateway`, or set `CODEMAN_DOCKER_BRIDGE_HOST`) that serves ONLY the hook endpoints (403 for any other path) into the same secret-gated pipeline; otherwise idle detection falls back to output-based through the docker-exec PTY. Container-set `CLAUDE_CODE_TMPDIR` keeps claude launching regardless of workspace path. `SessionState.docker`/`MuxSession.docker` round-trip through recovery. Every docker IO path is `IS_TEST_MODE` (VITEST) no-op'd; the pure builders are unit-tested. **Export/import** (`src/docker-export.ts`): full-image (`docker commit` + `save | gzip` + workspace tar + manifest) or workspace-only → one portable `~/.codeman/docker-exports/-.codeman-container.tgz`; import validates per-member sha256, traversal-guards the workspace tar, `docker load`s + quarantine-retags the image (`codeman/imported-:`, never overwriting a local tag); a `saveImageToTar` stream `pipeline` avoids truncation. **GPU** passthrough (`gpus` → `--gpus`, needs the NVIDIA container toolkit) and **elastic disk** (no `--storage-opt` cap, so container storage grows with data). SSE `docker:exportComplete`/`exportFailed`/`importComplete` (both registries). **UI** in `session-ui.js`: Create Case **Docker** tab (collapsed/compact form since 1.4.1), the one-click checkbox + Template picker, short `(docker)` case-menu tags, and a Manage-tab Export button; docker AND remote sessions name their tabs `w-` via the shared `_nextCaseSessionStartNumber()` so all tabs follow one naming convention. Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/docker-export.test.ts`, `test/network-host-guard.test.ts`. ## Session data and lifecycle ### Input delivery and WS resilience **Input**: `session.writeViaMux()` for programmatic/curl input — tmux `send-keys -l` (literal) + `send-keys Enter`. Single-line only (fire-and-once). Interactive **browser** input goes through a durable **exactly-once** layer: each frame carries a stable `clientId` + monotonic per-session `seq`, persisted to localStorage until the server ACKs (`{t:'ia',seq}` over WS, or HTTP 2xx), so a dropped link/reconnect can't lose or double-deliver a prompt. **WS resilience** (#149): the upgrade URL carries `cid = clientId + ':' + perTabNonce`, and `ws-connection-registry.ts` supersedes only same-TAB reconnects (two tabs on one session coexist; input frames keep the bare `clientId` for seq dedup); reconnects back off exponentially (attempts preserved across `_connectWs`), and the header connection chip renders from a real `_wsState` lifecycle (`connecting`/`connected`/`fallback`/`reconnecting`/`disconnected`). ### Agent wait primitives **Agent wait primitives** (`GET /api/sessions/:id/wait`, `GET /api/sessions/:id/wait-output`, and the `wait`/`waitTimeout` fields on `POST /api/sessions/:id/input`): bounded long-polls that let an agent driving Codeman from a shell tool block until something happens. They exist because SSE was the only "tell me when" channel Codeman had, and a curl-driven caller cannot practically hold a stream and parse events inline. The blocking core is `src/web/session-wait-registry.ts` (no IO, no `Session` reference, so it unit-tests in isolation), bounds live in `src/config/agent-wait.ts`, and the wiring is three `notifySignal()` calls next to existing broadcasts (`session-listener-wiring.ts` for `working`/`idle`/`exit`, `hook-event-routes.ts` for `stop`/`blocked`) plus `notifyOutput()` riding the already-attached `terminal` listener. Design: `docs/agent-control-plan.md` §3; wire contract: `docs/api-reference.md`. **Ordering rules, both load-bearing and both invisible to a reader of either side alone.** ⚠️ **exit-before-cancel**: `_doCleanupSession` in `server.ts` must call `sessionWaits.notifySignal(id, 'exit')` BEFORE `sessionWaits.cancelAll(id)`, because that method detaches the session's listeners before `session.stop()`, so on a delete the PTY's own exit event never reaches the registry and an `until=exit` caller would get a bare `ended: true` instead of the signal it asked for. Found by live-testing the delete path, not by the unit tests. ⚠️ **registered-before-write**: the send-and-wait path on `POST .../input` registers the waiter BEFORE writing to the PTY, and that ordering is the entire reason the combined endpoint exists rather than documenting "POST, then GET .../wait": between the write and the session flipping to `working` there is a window in which a separate wait sees the session still idle and instantly reports the PREVIOUS turn as this turn's answer. Two consequences hang off it: `useMux` delivery is awaited on that path (the response is staying open anyway, so a `writeViaMux` failure becomes observable for the first time) while the non-wait path keeps its fire-and-forget shape byte for byte, and because `shouldApplyInput()` MUTATES (it records the seq) before registration, a registration that fails on a full pool must `forgetInputSeq()` before returning, or the caller's retry is rejected as a duplicate and the input is lost by the very mechanism reliable delivery exists for. **A timeout is a 200, deliberately.** `{"timedOut": true, "signal": null}` with HTTP 200 is the long-poll succeeding at answering "did this happen within N ms?" with "no". The documented client pattern is a loop over short waits (`DEFAULT_WAIT_MS` is 60s precisely because prod is reached through `tailscale serve` and users run cloudflared, both of which cut idle connections), and turning every poll boundary into a 4xx would make that loop indistinguishable from a real failure. The alternatives are all worse: `408` is auto-retried by several clients and proxies, silently doubling the polling load; `504` is what a genuine tunnel failure looks like, so reusing it destroys the caller's ability to tell the two apart; `204` cannot carry `waitedMs`/`status`/`limitPaused` and breaks the uniform envelope the versioning policy makes a stable promise. Errors are reserved for `INVALID_INPUT` (400), `NOT_FOUND` (404, also the multi-user ownership answer via `findSessionOrFail`), `SESSION_BUSY` (409, this session's cap) and `RATE_LIMITED` (429, the per-owner or process-wide cap: a global cap reported as `SESSION_BUSY` tells the caller to switch sessions, which cannot help). ⚠️ The clamp is silent, so the EFFECTIVE timeout is echoed back as `wait.timeoutMs`: a caller that asked for 30 minutes, got the 600s ceiling, and could not see it would read the timeout as "the worker is wedged" and kill a session that was working fine. All three endpoints nest the result under `data.wait` for the same reason, so one client helper works against any of them instead of an agent's `is_done()` reading `undefined` off the shape it did not expect. **Matching is literal, and that is a language constraint, not a missing feature.** `search-service.ts` already avoids regex so there is no ReDoS surface, and this endpoint is more exposed still: the pattern is caller-supplied and the input is a live stream. herdr can offer `--regex` on its equivalent because Rust's regex crate is linear-time with no backtracking; JavaScript's `RegExp` backtracks, so the same feature here is a denial-of-service primitive. A `regex` parameter is therefore REJECTED with a 400 rather than ignored, since an agent that assumed otherwise would silently wait on the wrong thing. `match` is capped at 200 chars (`MAX_MATCH_LENGTH`), which is also what keeps the per-waiter carry buffer small: a match can straddle two PTY chunks, so each waiter carries `match.length - 1` characters of the previous chunk and tests `carry + chunk`. ⚠️ **`from=now` does not mean "printed after you asked"**: tmux repaints the visible screen on attach, on resize, and on any TUI redraw, and a repaint arrives as ordinary `terminal` data, so text already on screen can satisfy a fresh wait (observed live: a marker echoed a minute earlier matched instantly). This is inherent to running the agent under a multiplexer and is not fixable in the registry, so the contract is a marker unique per call (`echo DONE_$RANDOM`), never a generic one like `BUILD OK`, and every recipe must show that. ⚠️ **There is exactly ONE definition of the matched stream, `normalizeForMatch()`**: `stripAnsi()` (CSI, OSC, `ESC =`/`ESC >`) plus `ANSI_ESCAPE_RESIDUE` for what that helper leaves behind — above all the `ESC ( B` charset switch a stock bash prompt emits on every line, which in the first build survived into the matched text and made `match=tnode:` silently fail against a prompt that plainly renders `tnode:`. The haystack, the carry and the snippet window all derive from that one function, so matching and the snippet cannot drift apart again. ⚠️ **`splitTrailingEscape()` holds back a partial escape at a chunk boundary** until its tail arrives; its `INCOMPLETE_ANSI_TAIL` pattern must stay in lockstep with `ANSI_ESCAPE_RESIDUE` (a chunk cut between the `(` and the `B` is otherwise a fresh way to smuggle an escape into the haystack), and it is deliberately non-global (the repo-wide `lastIndex` hazard). With the carry, a match may straddle PTY chunks: `printf STRAD; sleep 1; printf DLEQQ` is matchable as `STRADDLEQQ` (measured live, both `from` modes). ⚠️ **The matcher still sees the byte stream, not the rendered pane** (`GET .../terminal` is a tmux capture, `source:'mux-visible'`): linear output agrees once escapes are stripped, but Claude Code positions words with cursor moves instead of spaces, so TUI text can arrive space-less (`Quicksafetycheck:Isthis...` — measured; some phrases keep their spaces depending on how the TUI drew them), which is why the documented advice remains ONE short space-free token the caller printed itself. The returned `snippet` is a RENDERING of the matched window, not a quotation: `SNIPPET_CONTROL_BYTES` drops bare control bytes that carry no ESC (BEL, NUL, backspace — `normalizeForMatch` removes escape SEQUENCES only) and blank runs are collapsed, because the consumer is an agent piping it through `jq` into its OWN pane, where a worker's raw bytes could otherwise reset or garble the orchestrator's display. **Lifetime discipline, per the 24-hour-session rules.** Every waiter owns exactly one timer, cleared on resolve; per-session waiter sets are deleted when they empty; `cancelAll()` runs on session teardown and `cancelEverything()` in `stop()`. ⚠️ **Waiter timers are deliberately NOT unref'd**, the opposite of the usual advice: an unref'd timer would let the process exit mid-wait and strand the HTTP response, so shutdown resolves waiters explicitly instead. All three routes also release the waiter when the client disconnects (`abortOnClientHangUp()` in `session-routes.ts`), which the caps make load-bearing rather than tidy: the documented loop-over-short-waits pattern is naturally written as `curl --max-time 30 ".../wait?timeout=60000"`, and without it every iteration abandons a waiter that lives out its full timeout, so the seventeenth call gets a 409 for a session nobody else is waiting on. ⚠️ **That listener goes on `reply.raw`, guarded by `writableFinished`, NOT on `req.raw`** (the obvious choice, and the one the SSE route in `server.ts` can afford because it only ever serves a GET). `req.raw` emits `close` as soon as the REQUEST BODY has finished streaming, which on a POST happens before the handler blocks: measured at +1ms with `aborted: false`, indistinguishable from a real hang-up, so wiring it there cancels every send-and-wait instantly and silently kills the feature, while GET keeps working because a GET has no body to finish. `reply.raw` emits `close` both on a completed response and on a dead socket, and `writableFinished` is the only thing that separates them, so the guard is load-bearing rather than defensive. ⚠️ `app.inject()` never emits `close` at all, so none of this is observable in a route test: the regression test has to bind a real port. `MAX_WAITERS_PER_SESSION` (16) is a COMBINED signal-plus-output budget (`waiterCount()` sums both maps), not 16 of each; `MAX_WAITERS_PER_OWNER` (48) applies only when an owner is passed, so single-user mode behaves exactly as before; `MAX_WAITERS_TOTAL` (128) mirrors `MAX_SSE_CLIENTS` in `map-limits.ts`, since each pending waiter costs an open HTTP response plus a timer. Capacity is asserted BEFORE the expensive work on both sides: `waitForOutput()` checks before the `initialText` scan, and the `wait-output` route checks before reading `session.terminalBuffer`, whose getter joins the whole 32MB accumulator, so a request that is going to be rejected never pays for a buffer materialization. `from=buffer` scans only the tail (`MAX_BUFFER_SCAN_BYTES`, 256KB) because the question it answers is "did this appear recently", not "ever", and the tail is continuous with the live stream (`append()` and `emit('terminal')` receive the same bytes). **Signal availability is decided by MODE, not by `isExternalCliMode()`.** `stop` and `blocked` come from Claude Code hooks, so `hooksAvailableForMode()` is true only for `claude`: `shell` is not an external-CLI mode but installs no hooks either, so `until=stop` on a shell session is a guaranteed unresolvable wait dressed up as a timeout. The behavior split is deliberate and must survive refactors: an EXPLICIT request for an unavailable signal is a 400 naming the mode, while the DEFAULT set (`stop,idle,exit`) silently drops them and echoes the narrowed set back as `wait.until`, because omitting the parameter must never 400. Signal quality is not uniform either: `stop` is definitive (Claude Code says the turn is over), `idle` is inferred from output stabilization plus prompt detection and can flap mid-turn when a spinner pauses, which is why `stop` is the documented default to orchestrate on and `idle` is the fallback for sessions that emit no hooks. ⚠️ **`idle` being ACCEPTED for a mode does not mean it ever FIRES there.** `startShell()` emits exactly one `idle` on a 500ms readiness timer and nothing afterwards, so a shell session sits at `status:'idle'` no matter what its pane is doing; since send-and-wait and `fresh=1` both require a TRANSITION, both can only time out on a shell worker (measured: a default `wait` on `sleep 4` burned its full 25s). The ❯-prompt and spinner detection that drives the real `working`/`idle` cycle is Claude's output format, so hook-less modes synchronize with `wait-output` markers or `exit`, and the docs must say so rather than listing `idle` as "available" and letting the reader infer it is usable. ⚠️ **Signals are edge-triggered with no history**, and that is a real orchestration limit: a signal that fires while no waiter is registered is gone, unobservable by any later wait variant (`until=stop` after the turn ended just times out, `fresh` or not — measured, R2-A). Documented client patterns must therefore register the waiter before the event can fire (send-and-wait) or gather on latched `wait-output` markers with `from=buffer`; the skill's fan-out flow was rewritten accordingly, and "fire-and-forget N prompts, then gather signal-waits sequentially" must never be documented again. The durable fix, a server-side latched last-signal-per-turn, is deferred with Part 3 of `docs/agent-control-plan.md`. ⚠️ Relatedly, the route corrects liveness that `SessionStatus` cannot express: `currentSignalFor()` answers `exit` whenever `pid === null` (exited, detached, or created-and-never-started) **or the mux pane is dead**, because `Session` parks a DEAD PTY at `_status = 'idle'` and trusting the status would answer the default wait `{signal:'idle', immediate:true}` for a crashed worker while `until=exit` blocked forever on an event that already happened. ⚠️ **`pid` alone cannot carry liveness for a tmux-backed session**: that pid is the local `tmux attach` CLIENT, so a worker exiting inside its pane leaves `pane_dead=1` with the client alive and `pid` never goes null — the `pid === null` branch is unreachable in the normal configuration (unit tests exercise it because `MockSession` sets `pid` by hand; only a live instance showed the gap). Liveness is therefore probed at the mux layer: `workerIsDead()` consults `mux.isPaneDead(muxName)` with a ~750 ms per-pane cache, ONLY on blocking waits (measured: 0 tmux execs across 100 non-wait input POSTs — the browser hot path pays nothing), plus a refcounted 3 s `watchForDeadWorker` interval so a worker dying while a wait is parked resolves it in ~3 s instead of burning the timeout. The probe fails SAFE by construction ("cannot tell" is never "dead": non-mux sessions, a missing or throwing `isPaneDead`, all return false). On send-and-wait, a "successful" `send-keys` into a dead pane additionally overrides `delivered` to `false` and rolls the dedup seq back (`undoOnFailure`), because the bytes went nowhere and a retry against a restarted worker must not be refused as a duplicate. The cost is that a just-created session reads as `exit`, which the wire docs must spell out as "not started yet"; the fix lives at the route rather than in `signalForStatus()` because the registry deliberately holds no `Session` reference. ⚠️ Two `claude`-mode cases still lose hooks for reasons outside the registry: a Docker case cannot reach a loopback-bound Codeman without `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, and a remote-SSH case runs the agent on another host whose hooks may never reach this server. Bounds are env-overridable (`CODEMAN_WAIT_MAX_MS`, `CODEMAN_WAIT_DEFAULT_MS`, `CODEMAN_WAIT_MAX_PER_SESSION`, `CODEMAN_WAIT_MAX_PER_OWNER`, `CODEMAN_WAIT_MAX_TOTAL`, `CODEMAN_WAIT_BUFFER_SCAN_BYTES`) and each is clamped to a hard bound so a typo degrades to the default instead of disabling the protection; they are internal tuning knobs like the rest of `src/config/`, NOT part of the SemVer-covered env-var surface in `versioning-policy.md`. Tests: `test/session-wait-registry.test.ts`, `test/routes/session-wait-routes.test.ts`, `test/routes/session-wait-output-routes.test.ts`, `test/routes/session-input-wait.test.ts`. ### Auto-resume on usage limit **Auto-resume on usage limit** ("token pause" control, opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit ("5-hour limit reached ∙ resets 8pm" and all 1.0.x–2.1.x variants), `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time from cleaned output; `SessionAutoOps` arms a timer for reset+2min, then sends Esc (dismisses the rate-limit dialog) + `continue`. Still-limited responses re-arm the loop (5-min retry on stale times); a `working` transition cancels it. Claude-mode only (detection rides `_processExpensiveParsers`). Persists/recovers via `SessionState.autoResumeEnabled`/`autoResumeAt`; respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected` — prevents `/clear` from wiping the paused conversation). Endpoint: `POST /api/sessions/:id/auto-resume`; SSE: `session:limitPauseScheduled`/`limitResume`/`limitResumeCancelled`. Tests: `test/usage-limit-patterns.test.ts`, `test/session-auto-resume.test.ts`. ### Plan-usage chip (statusLine telemetry) **Plan-usage chip** (statusLine telemetry, `showPlanUsageLimits`, per-device: desktop default **ON** since 1.9.3, handhelds OFF): Claude Code (v2.1.80+) pipes a JSON blob to a configured `statusLine.command` on each render; on Pro/Max it carries a `rate_limits` object (`five_hour`/`seven_day` windows only — no Opus weekly field — each `{used_percentage 0-100, resets_at epoch-SECONDS}`). Codeman injects its OWN statusLine exporter (`generateStatusLineCommand()` in `hooks-config.ts`, identified by the `/api/status-telemetry` marker — it only ever adds/updates/removes a statusLine that is _ours_, never a user's hand-authored one) that POSTs the blob to `POST /api/status-telemetry`. That route (auth-exempt like `/api/hook-event` — localhost-only, hook-secret-gated whenever auth is active, COD-91) parses via `usage-telemetry.ts` (pure, unit-tested), broadcasts SSE `session:statusTelemetry` (de-duped per session by `telemetrySignature` since the statusline fires on every assistant message), and returns a compact plain-text footer for the exporter to **print-through** (so injecting our statusLine doesn't blank the in-terminal footer). `plan-usage-latest.ts` holds the process-wide last value, replayed in the SSE init snapshot (`getLightState`) so the header chip (`#planUsageChip`, revealed by `planUsageChipEnabled()` in settings-ui.js, the single resolver behind the checkbox, the chip and the create-time `statusLineTelemetry` flag) renders immediately on page load / reconnect without per-browser localStorage. Claude-mode only. **Distinct from auto-resume** (which reacts to the limit _message_; this proactively shows the live %). Design: `docs/usage-limits-display-plan.md`. Tests: `test/usage-telemetry.test.ts`. ### Cron jobs **Cron (cron-style `CronJob`s)**: saved, named jobs with a recurring schedule (`once`/`interval`/`daily`/`weekly`), enable/disable, Run Now, next-run calc, and per-job run history (`CronJobRun`). ⚠️ **Distinct from the legacy `ScheduledRun`** (`/api/scheduled`, a run-now duration-bounded autonomous loop) — the two never interact; the legacy concept keeps the `Scheduled*` names, the recurring-job feature is `Cron*`. `CronService` (`src/cron/cron-service.ts`) owns CRUD + the 30s background due-tick (`tickDueJobs`, registered via `cleanup.setInterval` in `server.ts`; `init()` recomputes nextRunAt on boot) and **reuses the existing session layer** (create → `addSession` → `setupSessionListeners` → `startInteractive`/`startShell` → prompt via `writeViaMux`/`write`) rather than rebuilding tmux logic. Next-run math is pure/unit-tested in `cron-time.ts` (SERVER-LOCAL timezone for daily/weekly). Dup-launch guard = `lastDueKey` (jobId:fireTime); schedule is advanced BEFORE launch so a slow launch can't re-trigger. `once` jobs self-disable after firing (`completedOnce`). Persisted via `AppState.cronJobs`/`cronJobRuns` (StateStore accessors). Routes `/api/cron/jobs*` + `/api/cron/runs` (`cron-routes.ts`, `CronPort`); schema `CronJobSchema` (cross-field `superRefine`; the `.partial()` update schema does NOT re-run it); SSE `cron:*`. Frontend `cron-ui.js` (#cronModal). Claude/shell/opencode/codex/gemini agent types. Tests: `test/cron-time.test.ts`, `test/cron-service.test.ts`. Design: `docs/cron-discovery.md`. ### Unified session list and Session Manager **Unified session list** (COD-160/#139): `GET /api/sessions/unified?limit=&q=` merges live sessions, persisted state, lifecycle-log history, and Claude transcript files into one deduped list (pure core in `src/services/unified-session-service.ts`). Transcript rows are keyed by conversation UUID and folded into their owning session via a `claudeSessionId → Codeman id` alias map (resumed//clear-respawned sessions must not appear twice); lifecycle name/mode resolution is first-seen-wins (the log returns entries NEWEST-first). No terminal buffers in the response (unlike `/api/sessions`). Consumed by the Cmd+K Session Manager (#146). Session Manager polish (COD-162/#157, 1.6.0): **pinning** via `POST /api/sessions/:id/pin` (`session:pinned` SSE; killing a pinned session demotes it to a lightweight stopped record that stays visible/resumable, and cleanup skips pinned records); **cross-device tab order** via `PUT /api/session-order` (`session:orderChanged` SSE, persisted in `state.json`; pure `normalizeSessionOrder`/`mergeSessionOrder` in `src/session-order.ts`: pushing device wins, server-only ids fall to the end, never dropped); resume from the manager keeps the original session name (COD-143); `firstPrompt` is backfilled for sessions whose id != transcript UUID and the most recent prompt (`lastPrompt`) is shown + searched (COD-140/145). ### Full-scrollback replay **Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load OF EACH SESSION per page load requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other tab one frame of history); later switches keep the cheap `?tail=` visible-frame path. On top of that, scrolling up while already at the TOP of the buffer re-pulls `full=1` on demand (`_maybeRefetchFullHistory`, 4s per-session cooldown, in-flight + tab-switch guards, viewport position held across the replay). The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`. ### Terminal scrollback: strip flavors and wheel/touch forwarding **Two strip flavors, one carry** (#205, `session.ts:_handleTerminalOutput`): the FULL strip (`isAltScreenStripMode` = codex/claude/gemini) removes alt-screen toggles, `3J`, and mouse-tracking DECSETs. Every other mode (shell/opencode/antigravity) gets the NARROW strip (`isMuxAltScreenOnlyStripMode`) — alt-screen toggles ONLY — and only when tmux-backed (`useMux`). Rationale: the tmux CLIENT emits `smcup` as its first bytes at attach, before any program runs, parking xterm in the scrollback-less alternate buffer for the whole session (touch scrolling no-ops; xterm's own wheel handler converts the wheel to Up/Down arrows = readline history cycling — both #205 symptoms). tmux never forwards a pane program's alt-screen toggles to its client (it repaints instead; measured — vim/less inside a pane emit zero to the client), so the only thing the narrow strip ever removes is tmux's own smcup. It keeps `3J` (a user's `clear` is a deliberate scrollback wipe) and the mouse DECSETs (tmux passes those through even with `mouse off`; stripping them would break htop/vim mouse support). ⚠️ The `useMux` gate is load-bearing: `startShell()`/`startInteractive()` fall back to a DIRECT PTY when mux creation fails, and there the inner program's own `?1049h` really does reach xterm — stripping it would break vim/less/htop for real. The replay path (`session-routes.ts`, via `session.usesMux`) applies the same narrow branch; the frontend `_sessionUsesServerMouseStrip()` mirror stays claude/codex/gemini because only the FULL strip touches mouse DECSETs. The chunk-boundary carry (`_altScreenSeqCarry`) runs for both flavors. Tests: `test/claude-scrollback-strip.test.ts`. **Wheel/touch forwarding is NOT gated on viewport-at-bottom** (#205, `terminal-ui.js:_shouldForwardWheelToApp`): for sessions verified to scroll their own transcript on SGR wheel reports (codex, claude ≥ 2.1.187 — version via the local/docker/remote `--version` probes), the plain wheel AND touch drags forward as coalesced SGR reports (`_forwardScrollToApp` → `_sendSyntheticSgrWheel`, 40ms batches, 5-tick cap, 512-byte queue bound). It used to gate on the viewport being at the bottom so both scrollbacks stayed reachable, but a repaint-mode CLI keeps NO terminal scrollback of its own — xterm's buffer holds only replayed repaint frames, so local scrolling drags the CLI's pinned prompt box up the screen over stale frames; and `scrollToLastNonEmptyLine()` routinely parked the viewport off-bottom, silently pinning the wheel to local. Forwarding now snaps the viewport home first (SGR coordinates address the LIVE screen — a report computed from a scrolled-up viewport would hit-test the wrong row). Local scrollback remains on Shift+wheel and the `terminalWheelLocalScrollback` opt-out (both also cover touch via the shared gate; touch has no Shift, so the setting is its only local pin). `_wheelScrollLines()` normalizes `deltaMode` (Firefox fires LINE deltas ≈3/notch — read as pixels that rounded to 0 and fell to the ±1 fallback, ~4× too slow; PAGE deltas scale by `terminal.rows`) while keeping the #154 Shift-axis trap (macOS trackpads put Shift+scroll magnitude on deltaX). Tests: `test/terminal-touch-tap.test.ts`. **A false gate on a Claude session must not mean a DEAD gesture** (#205 round 2, `_maybePageCliTranscript`): every way `_shouldForwardWheelToApp()` returns false leaves a repaint-mode pane scrolling a buffer that has nothing in it (`baseY === 0`) — the version probe came back empty, the CLI really is older than 2.1.187, or the user turned on `terminalWheelLocalScrollback`. The 1.12.0 retest reported exactly that: a wheel that did nothing at all while Fn+Up (PageUp) paged back through intact text, which is the proof that the CLI's own history and the PTY input path were both fine. So under the triple guard (claude mode + gate false + `baseY === 0`) wheel and touch travel is translated into coalesced `\x1b[5~` / `\x1b[6~` through the same 40ms queue as the SGR reports, at half a screen of travel per page key (the key jumps a whole screen; a 1:1 mapping was unusably slow with a discrete wheel). ⚠️ Shift is excluded on purpose — it is the explicit "give me local scrollback" gesture and must keep that meaning. ⚠️ `terminalWheelLocalScrollback` is deliberately NOT scoped away from repaint-mode CLIs even though it is a footgun there: that would silently override an explicit user choice, so the fallback catches it instead. **Server-side counterpart**: `getClaudeCliVersion()` caches SUCCESS for the process lifetime but must never cache FAILURE — it used to, so one timed-out or PATH-starved probe at the first Claude session start disabled wheel-forwarding for every Claude session until the server restarted (a dead wheel on phone, tablet and laptop at once, the signature of a server-side cause). Failures now retry with a 1/2/4…15min backoff; the policy is the pure `resolveClaudeCliVersion()`. Tests: `test/terminal-scroll-routing.test.ts`, `test/claude-cli-version-cache.test.ts`. **Why the wheel went where it went is LOGGED** (`_logScrollRouting`): one console line per session per distinct decision — `[scroll] → forward-sgr|page-keys|local-scrollback|repull-refused-downgrade (mode=…, cliVersion=…, localScrollbackOptOut=…, mouseTracking=…, localScrollbackRows=…)`. #205 ran two rounds of remote guesswork over questions this line answers directly; keep it when touching the routing. **The wheel listener is CAPTURE-phase and Codeman owns the scroll** (#205 follow-up, measured on the live instance): xterm's viewport is a vscode-style ScrollableElement that consumes wheel events itself (preventDefault + stopPropagation) whenever it believes a scrollbar exists, ignores `attachCustomWheelEventHandler`, and goes DEAF after `terminal.reset()` — a tab switch or full-history replay leaves its scroll dimensions stale, after which wheel events neither scroll nor propagate reliably. A bubble-phase container listener therefore never fired once local scrollback existed (forwarding, deltaMode and the top-of-buffer re-pull all silently dead exactly on sessions WITH history), and after a tab switch nothing scrolled at all ("works at first, breaks after a tab switch"). The container wheel listener is `{capture: true}`, stops propagation, and scrolls locally via buffer-level `terminal.scrollLines()` (immune to the stale scroller). ⚠️ Two cases are deliberately passed through untouched, in this order BEFORE preventDefault: `mouseTrackingMode !== 'none'` (xterm's encoder forwards the wheel to the PTY — htop/vim with mouse on) and `buffer.active.type === 'alternate'` (direct-PTY vim/less: xterm's alt-scroll converts the wheel to cursor keys). Do not "simplify" this back to a bubble listener or re-delegate local scrolling to xterm's viewport. E2E guard: the reload → tab-switch → wheel matrix in the #205 verification scripts. ### Run launch synchronization **Run launch synchronization**: the main Run entrypoint in `session-ui.js` holds an in-flight lock and disables `#runBtn` for the whole launch (at least 500ms), so a double click cannot create duplicate sessions with the same `w-` name. A successful create/quick-start also calls `_ensureCreatedSessionVisible()` before `selectSession()`: local creates use the response's full session snapshot; quick-start modes fetch `GET /api/sessions/:id` only when `session:created` SSE has not already populated the map. The normal `_onSessionCreated()` handler remains the idempotent upsert, so POST-first and SSE-first ordering both produce one immediately-rendered tab. Tests: `test/run-mode-ui.test.ts`. ### Circuit breakers: Ralph and PTY-exit **Circuit breaker**: Prevents respawn thrashing. States: `CLOSED` → `HALF_OPEN` → `OPEN`. Reset: `/api/sessions/:id/ralph-circuit-breaker/reset`. **Distinct: PTY-exit breaker** (COD-115/118/#147, `session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits (crash loops on attach), blocks further auto-restarts, broadcasts SSE `session:respawnBreakerTripped` + push (in `PUSH_EVENT_MAP`). Reset ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive` (sent by the user-facing restart control) — the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. Sessions also scrub inherited `TMUX`/`TMUX_PANE` env so Codeman-in-tmux doesn't nest. Tests: `test/respawn-pty-breaker.test.ts`. ## Features ### Attachments **Attachments** (live external document references; COD-37/#119 core, COD-38/#120 previews, COD-39/#121 history): all wiring in `file-routes.ts`. **Registry** (`attachment-registry.ts`): an **in-memory** map of a stable `attachmentId` → an absolute, `realpath`-resolved, extension-allowlisted file path, so browser requests (`GET /api/sessions/:id/attachments/:attachmentId/raw`) never carry arbitrary absolute paths; `POST /api/sessions/:id/attachments` registers one. **Magic links** (`attachment-magic.ts`): parses `codeman://attach?...` out of terminal output — ⚠️ this scanner is prompt-injectable, so the scan path is **force-confined to the session workspace** (a hostile prompt could otherwise make it read arbitrary host files over SSE); emits the `attachment:detected` SSE event. Security gate is an extension **allowlist** (`isSupportedAttachmentExtension`, in the registry/magic modules), not a blocklist; a separate path layer (`config/attachment-guard.ts`) confines reads to the workspace (`attachmentConfineToWorkspace`) and blocks sensitive trees (`/root`, `/etc`). **Previews + thumbnails** (COD-38): `:attachmentId/preview` + `:attachmentId/thumbnail` (and the workspace-file equivalents `file-preview`/`file-thumbnail`) render Office docs/PDFs via external converters (`pdftoppm` / LibreOffice `soffice` / Word-COM `powershell`); `document-preview-cache.ts` is a shared disk cache (de-dups _identical_ in-flight inputs), `document-thumbnailer.ts` does best-effort first-page images, and `document-conversion-limiter.ts` is a **global converter-spawn concurrency cap** (`runWithConversionLimit`) — without it, N distinct large docs detected at once fork N multi-minute converter processes = a localhost fork-bomb-shaped resource-exhaustion vector. **History drawer** (COD-39): `session-attachment-history.ts` tracks the last `ATTACHMENT_HISTORY_LIMIT` (100) attachments per session (`Session._attachmentHistory`, persisted via `SessionState.attachmentHistory`, replayed so externals re-register on reconnect); `GET /api/sessions/:id/attachments` is the list endpoint. ⚠️ The history drawer's launcher button is desktop-only — hidden on phones (regression-guarded; see `mobile-header-buttons-policy` test). Session-local files keep using the existing workspace-scoped `file-routes` paths; the registry is only for explicit live externals. **Codex generated artifacts** (COD-166/#150, `generated-artifact-attachments.ts`): codex-mode sessions ALSO scan (ANSI-stripped) output for `Saved to: file:///…` lines and surface those files as attachment cards with a relaxed trust policy — the allow decision runs on the **realpath-resolved** path against `os.homedir()`-anchored `~/.codex` marker dirs (symlink escapes fall back to force-confinement); gated to `mode === 'codex'` only (`source` is a REQUIRED param through the listener-deps chain — a dropped arg here silently kills the feature). Image thumbnails pass through jpg/jpeg/gif/webp. ### Filesystem path picker **Filesystem path picker** (Link Existing "Browse" button + the extended mobile keyboard's `📁 Path` key): a lazy one-directory-at-a-time browser over `GET /api/filesystem/browse`, with `GET /api/filesystem/preview` serving the tapped file. It starts at the active session's working directory (falling back to `/mnt/d`), hides dot entries, and inserts the chosen path **without** Enter so the prompt is not submitted. The companion `⌫ All` key clears only the current unsent prompt buffer and must never emit the agent's `/clear` command. ⚠️ **This is a second file-serving surface, so it carries the same confinement burden as [Attachments](#attachments) and does not inherit it automatically.** Traversal is allowlisted to Home, `CASES_DIR`, `/mnt/d`, or extra roots explicitly configured via `CODEMAN_FILE_PICKER_ROOTS`; sensitive trees are blocked and symlink escapes are rejected after `realpath` resolution rather than before. Without the realpath step a symlink inside an allowed root would walk straight out of it. `preview` reuses the shared conversion cache and the **global** `document-conversion-limiter`, which is what stops N concurrent large-document previews from forking N multi-minute converter processes. Content types are pinned: images and PDF inline, DOCX/PPTX through the converters, and Markdown/TXT/JSON as inert `text/plain` (never `text/html`, which would be stored XSS on our own origin). Size caps are 2MB for text and 50MB for binary/document previews. ⚠️ **Ownership scoping is also not inherited, and both endpoints must do it themselves.** Two separate holes shipped in the original version and are now regression-guarded in `test/routes/file-routes.test.ts`: 1. The optional `sessionId` param adds that session's `workingDir` as a "Current Folder" root. It is looked up directly off `ctx.sessions`/`ctx.store` rather than through `findSessionOrFail`, so the `canAccessOwned` check has to be written out by hand. Without it a multi-user caller pins **another user's** working directory as a browse root just by passing their session id. It reports 404 rather than 403 so the endpoint does not confirm that a session id exists. 2. `Home` and `CASES_DIR` were unconditional roots. Per-user spaces live at `/`, which is **inside `homedir()`**, so a `Home` root alone let any authenticated user browse and preview every other user's workspace. In multi-user mode a non-admin now gets only `My Space` (their own `userSpacePath`) plus anything in `CODEMAN_FILE_PICKER_ROOTS`; `/mnt/d` is dropped too, since a broad host mount should be an explicit operator decision in a multi-user deployment. Admins and single-user mode keep the host-wide set unchanged. The general rule: **any new endpoint that turns a caller-supplied `sessionId` into a filesystem path is an ownership boundary**, whether or not it goes through `findSessionOrFail`. ### File Viewer edit mode **File Viewer edit mode** (issue #212, design in `docs/file-viewer-edit-plan.md`): the file-preview overlay can edit workspace text files in place — `GET /api/sessions/:id/file-content?edit=1` (read-for-edit) + `PUT /api/sessions/:id/file-content` (save), policy in `src/config/file-editing.ts`, UI in `panels-ui.js`. This is the **only file surface that writes**, so it carries every rule the read surfaces have plus its own: - **Confinement is the read path's, plus write-only gates.** `findSessionOrFail` (ownership) → `validateSessionFilePath` (realpath + workspace boundary; escapes report as 404, same as reads) → sensitive-path + attachment-guard blocklists (403) → `.git/` subtree deny (403 — `.git/hooks/*` is code execution) → extension **allowlist** (400; `svg` and `env` deliberately excluded). ⚠️ **There is no `O_CREAT` anywhere in the handler** — that absence is what makes "edit-in-place only, never create" a structural property instead of a convention. Do not add a create path without treating it as a new security surface. - **A truncated buffer must never become an edit buffer.** The plain preview truncates to `lines` (default 500); saving such a buffer would silently delete everything past the cut, and the hash check cannot catch it (the loaded prefix hashes differently from the full file, which reads as an ordinary conflict at best). `edit=1` therefore never truncates — it 413s over `MAX_EDITABLE_BYTES` (512KB) instead — and the frontend always re-fetches with `edit=1` before swapping in the textarea, even though the preview already holds content. - **Concurrency is optimistic by content hash, not mtime.** The client echoes the sha256 it loaded (`baseHash`); mismatch → 409 CONFLICT (plain envelope — the error arm carries no data; the client re-fetches `edit=1` for fresh state) unless `force:true`. mtime alone is wrong: agents rewrite files within one timestamp tick. - **Writes are `wx` temp + `fchmod` + `fsync` + `rename` in the target's directory.** `wx` cannot follow a pre-existing symlink and `rename()` replaces (not follows) a symlink final component, which closes the validate-then-write TOCTOU window; `fchmod` because `open()`'s mode argument is masked by the umask; a symlink whose target is *inside* the workspace is deliberately written through (validation returns the realpath). Trade-off (same as vim): the inode changes, so hardlinks keep old content. - **Corruption guards**: NUL-sniff + UTF-8 **round-trip compare** (`Buffer.from(buf.toString('utf8'), 'utf8').equals(buf)`) refuse binary and non-UTF-8 files — decoding latin-1 yields U+FFFD replacements and writing those back destroys the original bytes. EOL is detected server-side and re-applied on save because a `