# Architecture invariants Implementation detail extracted from `CLAUDE.md` so that file stays small enough to load into every session cheaply. Most sections are the original paragraphs, verbatim, including the version history and PR references that explain *why* each rule exists; newer ones are written here first and summarized back into `CLAUDE.md` as a short rule plus a pointer. `CLAUDE.md` keeps the short form of each rule plus a pointer to the section here. Read the pointer first; come here when you need the mechanism, the file names, or the history behind a constraint. --- ## Network binding and instance isolation ### Default bind, and the non-loopback warning path **Default bind is loopback-only; non-loopback without a password starts but warns** — since COD-29 (PR #107) the web server defaults to `--host 127.0.0.1` (was `0.0.0.0`). As of **0.9.0** binding a non-loopback host (`--host`/`-H`/`CODEMAN_HOST`) without `CODEMAN_PASSWORD` **no longer refuses to start — it starts and prints a loud warning** listing the fixes (set `CODEMAN_PASSWORD`, bind loopback + tunnel/`tailscale serve`, or `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` to acknowledge → terser note). Host classification is `isLoopbackBindHost()` in `network-auth-policy.ts`; the warn-vs-start logic is in `server.ts` `start()`; flags wired in `cli.ts`. ⚠️ Operational note: the production systemd unit runs `node dist/index.js web --https` with no `--host`, so it binds **localhost only** — reach it remotely via `tailscale serve`/tunnel to `127.0.0.1`, or add `Environment=CODEMAN_HOST=0.0.0.0` + `Environment=CODEMAN_PASSWORD=…` to `~/.config/systemd/user/codeman-web.service`. A loopback bind is reachable through a same-host tunnel (cloudflared/tailscale → `127.0.0.1`) but NOT by a browser hitting the box's LAN IP. Auth user defaults to `admin`. **Installer note** (1.8.x, `install.sh`): interactive installs now PROMPT for the binding, defaulting to LAN (`0.0.0.0`) with a required password prompt (skipping the password needs an explicit confirm and prints a loud warning); non-interactive installs keep loopback unless `CODEMAN_HOST` is preset, and re-runs/updates preserve the EXISTING binding (`read_existing_binding()` parses the current systemd unit / launchd plist). The server binary's own default is unchanged. **Full model: `docs/security-architecture.md`.** ### Instance isolation and the multi-instance attach danger **Instance isolation / multi-instance attach danger** — data dir (`~/.codeman`) and tmux socket (`tmux -L codeman`) are PROCESS-WIDE and shared by every Codeman on the machine, derived from `CODEMAN_INSTANCE` via `src/config/instance.ts` (`getDataDir()`/`dataPath()`/`DEFAULT_TMUX_SOCKET`). ⚠️ A 2nd instance on the SAME socket **discovers and attaches PTYs to the first instance's live sessions** (`tmux -L codeman attach-session …`), resizing/mutating them — `$HOME` isolation is NOT enough (tmux is system-global). To run two instances, give each a distinct `CODEMAN_INSTANCE` (scopes BOTH dir+socket: `~/.codeman-` + `-L codeman-`), or set `CODEMAN_TMUX_SOCKET` + `CODEMAN_DATA_DIR` individually. **`CODEMAN_INSTANCE` defaults to empty = the production layout (`~/.codeman`, `-L codeman`, port 3000)**, so this branch is safe to ship to master without disturbing existing installs. To run THIS beta alongside prod, launch with `scripts/run-beta.sh` (`CODEMAN_INSTANCE=beta` + `CODEMAN_PORT=5000`) — it never collides with prod's data dir/socket/port. Any new `~/.codeman/...` path MUST go through `dataPath()`, never `join(homedir(), '.codeman', …)`, and any new `tmux -L` caller through `resolveTmuxSocketName()` (same module): it applies the `CODEMAN_TMUX_SOCKET` override only when the name is safe and falls back to `DEFAULT_TMUX_SOCKET` otherwise. `TmuxManager` was the only such caller until `codeman tui` shelled out to tmux from a SECOND process for its degraded-mode listing and its attach handoff; a hardcoded `codeman` there would have pointed a beta instance straight at prod's panes. ## Session launch modes **Terminal colour env** — the stock registry decides each CLI's colour vars. Claude, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek and OMP export `COLORTERM=truecolor`; `shell` and `opencode` unset it. All of those except Claude also unset `NO_COLOR`, so a user who exports `NO_COLOR` globally keeps monochrome Claude panes. The variable matters because a CLI inheriting no `COLORTERM` quantizes every RGB color it draws down to whatever palette `TERM` alone implies, and the pane's `TERM` is not a constant. ⚠️ Codeman sets no `default-terminal` and passes tmux no `-f`, so it is tmux's default (`tmux-256color` since 3.2) unless the user's own `~/.tmux.conf` says otherwise, which Codeman's server DOES read. At `tmux-256color` supports-color reports 256 colors and `rgb(55, 55, 55)` lands on `ESC[48;5;237m`: visible, but not the color the theme named. Where `TERM` resolves to a 16-color entry instead (tmux older than 3.2, or a conf setting `default-terminal screen`) every dark background collapses to `ESC[40m`, the terminal's own black, and the block disappears entirely. That is why the same Claude theme looks different on two machines, and why a bug report here is worth pairing with the reporter's `tmux -V` and their pane's real `TERM`. ⚠️ These declarations reach the local tmux pane via `buildEnvExports()`, its attach client via `cliExportsTruecolor()`, and the direct-PTY fallback via `buildClaudeEnv()` — they do NOT reach a remote pane, which `buildRemoteLaunchCommand()` builds with no env exports at all. A Docker pane takes `COLORTERM=truecolor` from the hardcoded `envCreate`/`execEnv` in `tmux-manager.ts`, which apply to every mode including the two the registry says must unset it. A `~/.codeman/clis.json` override replaces these arrays wholesale (`deepMerge`), so a custom entry can drop either list. That is why both consumers apply the entry BEFORE Codeman's own variables rather than after: `buildEnvExports()` emits `...cliEnv` ahead of `export CODEMAN_MUX=1`, and `buildClaudeEnv()` assigns `PATH`/`TERM`/`CODEMAN_*` after its unset/export loop. Reversed, a config-supplied `unset` naming `CODEMAN_HOOK_SECRET_FILE` would strip it on one path and not the other. ### External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP) **External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP)**: `isExternalCliMode()` in `session.ts` (`mode === 'opencode' || 'codex' || 'gemini' || 'antigravity' || 'pi' || 'grok' || 'deepseek'`) gates Claude-specific behavior — Ralph tracker, BashToolParser, token/CLI-info parsing, and ❯-prompt readiness detection are all skipped (these CLIs render their own TUIs; readiness = output stabilization instead). ⚠️ **Work detection left this gate in #385** and is now per-CLI `capabilities.workDetect` data (`promptGlyph` + `workingLine`), because gating it on the mode left every Codex session reporting `idle` for its entire life; a CLI declaring neither falls back to Claude's pair, which is logic-identical to the pre-registry behaviour. All seven modes **require tmux — no direct PTY fallback** — because secrets are injected via `tmux setenv` (socket-scoped `${this.tmux()} setenv`, never on the spawn command line): OpenCode gets `OPENCODE_CONFIG_CONTENT` etc., Codex gets `OPENAI_API_KEY`/`CODEX_API_KEY`/`CODEX_HOME` (`setCodexEnvVars`), Gemini gets `GEMINI_API_KEY`/`GOOGLE_API_KEY`/`GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI` etc. (`setGeminiEnvVars`, all in `tmux-manager.ts`). Codex specifics: command built by `buildCodexCommand()` (`--model`, `resume `, `--dangerously-bypass-approvals-and-sandbox` from the `codexConfig` payload / `codexDangerouslyBypassApprovals` app setting; `renderMode` is schema-coerced to `'hybrid'`, the only supported mode). Gemini specifics: command built by `buildGeminiCommand()` (`--skip-trust` always, `--approval-mode ` defaulting to `yolo` for parity with Claude's `--dangerously-skip-permissions`, `--model`, `--resume` from the `geminiConfig` payload); availability via `GET /api/gemini/status` — session/quick-start routes fail with `OPERATION_FAILED` + install hint (`npm install -g @google/gemini-cli`) when missing. Codex, Gemini, Antigravity, Pi, Grok, DeepSeek and OMP export `COLORTERM=truecolor` and unset `NO_COLOR`; `opencode` unsets `COLORTERM`. **Terminal colour env** under Session launch modes covers Claude and says which panes those declarations actually reach. Gemini joins `isAltScreenStripMode()` (Codex/Claude/Gemini are Ink TUIs that repaint inline → strip alt-screen/`3J` so scrollback survives). Codex availability via `GET /api/codex/status`. Antigravity specifics: command built by `buildAntigravityCommand()` (`--model`, `--conversation ` resume, `--dangerously-skip-permissions` from the `antigravityConfig` payload); availability via `GET /api/antigravity/status` — routes fail with `OPERATION_FAILED` + install hint (`curl -fsSL https://antigravity.google/cli/install.sh | bash`) when missing. Unlike the other three it is NOT an npm package (standalone binary, `~/.local/bin/agy`), which is why `docker/agent.Dockerfile` installs it with its own `--dir /usr/local/bin` step rather than in the `npm install -g` line, and why it does NOT join `isAltScreenStripMode()`. Frontend: run-mode dropdown → `runCodex()`/`runGemini()` in `session-ui.js` ("Run CX"/"Run GM" labels), App Settings → Agents & CLIs → Codex; Respawn/Ralph options are Claude-only, so session options open on the Session tab for external CLI sessions. ⚠️ `run*()` MUST unwrap the `{success,data}` envelope (`(await res.json()).data.available` / `data.data.sessionId`) — reading the raw shape silently breaks the run. Tests: `test/run-mode-ui.test.ts` + `test/gemini-mode.test.ts` (vm-sandbox harness, no real DOM). Grok specifics: command built by `buildGrokCommand()` (`--always-approve` from `grokConfig.alwaysApprove` — grok's `bypassPermissions` permission mode, deny rules still apply; `--model`; `--resume ` / `--continue`, id-regexed so grok's resume-by-TITLE feature can never put an arbitrary string on the spawn line); availability via `GET /api/grok/status`, which carries `version` because the resolver version-probes candidates (`grok` has npm squatters, e.g. @vibe-kit/grok-cli — `GROK_VERSION_REGEX` is shared with the dependency registry so doctor and run mode agree). Like antigravity it is a standalone binary (xAI installer → `~/.grok/bin`, symlinked into `~/.local/bin`), so `docker/agent.Dockerfile` installs it in its own step (copy to `/usr/local/bin`, drop root's `~/.grok` in the same layer) and it stays OUT of `isAltScreenStripMode()` (fullscreen alt-screen TUI with mouse support — the opencode case, not the Ink case). Env allowlist: `GROK_*` plus the vendor namespace `XAI_*` (`XAI_API_KEY` is grok's documented headless auth var — the same narrow-vendor-namespace reasoning as `GOOGLE_*` for gemini). Docker cred seeding is per-file (`auth.json`, `config.toml`, `pager.toml` from `~/.grok` — the dir also holds `sessions/`, `memory/`, and the ~160MB binary under `downloads/`). Grok tests: `test/grok-mode.test.ts`, `test/grok-cli-resolver.test.ts`. **DeepSeek Harness (`dsh`) specifics** — the mode that breaks three of the assumptions the six above share, so read this before changing anything about it. ⚠️ **The agent is a PROFILE, not the binary.** `dsh` is a launcher over `$DSH_HOME/profiles/` (an ordered stack of plugin-bundle patch layers), and DeepSeek ships only `web` (browser UI), `headless` (one-shot) and `base` (no app). The interactive terminal front door is ALWAYS third-party. So availability is TWO questions, not one, and `isDeepSeekRunnable()` (binary AND a pane-capable profile) is what the Run button gates on while `isDeepSeekAvailable()` (binary only) gates the "add a profile" affordance and the web-UI shortcut. Reporting only the binary would let Run spawn a pane that dies on arrival, which is this mode's single most confusing failure. `buildDeepSeekCommand()` emits `dsh --profile [--resume [id]]`; an absent profile resolves through `resolveDefaultDeepSeekProfile()`, which prefers a recognized TUI, then an UNRECOGNIZED profile (third-party by construction — a classifier that has not heard of a bundle must not hide it), and refuses `web`/`headless`, which cannot drive a pane. ⚠️ **The permission switch is an ENV VAR, not a flag.** The harness has no `--dangerously-skip-permissions` equivalent; its sandbox/approval rows read `DSH_PERMISSION_MODE` with three presets (`read-only` / `workspace-write` / `danger-full-access`; measured from `dsh --dump-default-config`). It is exported via `tmux setenv` in `_configureCliEnv()`, never on the command line, and `test/deepseek-mode.test.ts` pins that nothing permission-shaped ever reaches the spawn line. This is the ONE place a Codeman env export is the right mechanism rather than the forbidden one: unlike `CLAUDE_CODE_EFFORT_LEVEL` (which hard-locks in-session `/effort`), the harness reads it with `??` as a boot-time DEFAULT, so it stays soft. Absent = `workspace-write`, which still asks, so the multi-user clamp is the only-if-sent branch (codex/antigravity/grok shape, not pi's materialize) — and it clamps down to `workspace-write`, NOT `read-only`, because the clamp removes privilege without breaking a session's ability to edit its own workspace. ⚠️ **Clamping the config is only HALF the gate here, and this is the only CLI where that is true.** Every sibling's bypass is a command-line flag, reachable only through the per-CLI config `clampExternalCliBypassForOwner()` already owns. DeepSeek's is an env var, `DSH_*` is an allowlisted `envOverrides` prefix (it must be — that is also how the harness's ordinary knobs are set), and `applyEnvOverrides()` runs AFTER `_configureCliEnv()` in tmux-manager, so `envOverrides: {DSH_PERMISSION_MODE: 'danger-full-access'}` sent on the SAME request as a clamped config lands last and wins. `clampEnvOverridesForOwner()` (session-routes.ts, exported as `_clampEnvOverridesForOwner` for tests) DROPS `DSH_PERMISSION_MODE`, `DSH_HOME` and `DEEPSEEK_BASE_URL` for a non-granted owner (the last because `_configureCliEnv()` forwards the SERVER's own `DEEPSEEK_API_KEY` into the pane, so a redirected base URL would send it to a foreign host) rather than rewriting them, since dropping falls through to what `_configureCliEnv()` exports, which is already the clamped value. `DSH_HOME` is on that list because it aims the launcher at a profile tree and a profile's plugin code executes at BOOT, before any approval row can apply — the wider of the two holes. No-op in single-user mode and for a granted owner, like every other clamp. ⚠️ **It is the only non-claude mode that passes `hooksAvailableForMode()`, and it earned that.** The community terminal front door reports its own lifecycle to a supervising process through a generic env-var-gated contract inherited from Herdr: with `HERDR_ENV=1` + `HERDR_BIN_PATH` + `HERDR_PANE_ID` set it shells out ` pane report-agent --state idle|working|blocked …` on every state change and treats exit 0 as delivered. `deepseek-status-shim.ts` GENERATES a small script into the data dir (like `self-update-runner.sh`, so npm installs and git clones behave alike) and points `HERDR_BIN_PATH` at it; it forwards to `POST /api/hook-event` as `idle→stop`, `blocked→permission_prompt`, `working→agent_working`. So a dsh session gets real respawn triggers, real `wait` stop/blocked signals and real Approvals Inbox items instead of output-stabilization guesswork. This is an interface implementation, not an impersonation — no real `herdr` binary is ever executed. A TUI that does not implement the contract simply never calls the shim and falls back to stabilization, so the feature is inert rather than harmful there. ⚠️ **For deepseek alone, `hooksAvailableForMode()` is a per-SESSION question**, which is why it takes a `HookCapabilityOptions` second argument and every call site passes `sessionHookOptions(session)`: `deepSeekConfig.statusReporting: false` skips the `HERDR_*` export, and that triple is the only reason a dsh session posts anything, so answering from the mode alone would accept `until=stop` on a session where nothing can ever send one — the infinite-wait-dressed-as-a-timeout the predicate exists to prevent. The option defaults permissive (`!== false`), so a call site that forgets it degrades to the old behaviour instead of 400ing a working session. ⚠️ Profile conformance is the LIMIT of what is knowable at request time: `resolveDefaultDeepSeekProfile()` deliberately treats an unrecognized profile as launchable, so a non-conforming TUI still answers true and still times out on an explicit `stop` — which is why the DEFAULT signal set keeps `idle`/`exit`. ⚠️ **The predicate is not a stand-in for "is this a claude session"**, though it read like one while `claude` was the only true answer: Read My Mind (`POST /api/sessions/:id/readmymind`) and intent capture (`captureIntentPrompt`) read Claude's own transcript and were silently widened to deepseek by this change, so both compare `mode === 'claude'` directly and a static check in `test/deepseek-mode.test.ts` keeps them there. ⚠️ **`agent_working` is a hook event with no Claude Code hook behind it** (157th SSE constant). It exists because a harness turn cannot run while one of its own modal approvals is on screen, so "the agent started working" proves a dialog was answered in the terminal. It joins `APPROVAL_RESOLVING_EVENTS`; without it a dsh session's red alert would survive until the next `stop`, the exact stuck-alert bug the claude path already had to fix once — and the pane-capture staleness sweep that fixed it there is Claude-dialog-shaped and cannot help here. ⚠️ **Answers are READ FROM DISK for this mode, not scraped off the pane** (`src/deepseek-transcript.ts`, behind `GET /api/sessions/:id/last-response`). dsh writes a structured JSONL transcript at `$DSH_HOME/sessions///session.jsonl.zstd`, so it belongs with claude and codex rather than with the pane-segmented modes — and for dsh specifically the segmenter was not merely coarse but WRONG: dsh-TUI paints a full-screen splash, so a `last-response` call on a fresh dsh session returned its ASCII-art logo, which anything polling for a worker's first answer reads as an answer. Four mechanisms in that file are load-bearing. **(1)** dsh appends **one zstd FRAME per write**, and Node's `zlib` zstd decoder — one-shot AND streaming — stops at the first frame end: a real 56-line transcript decoded as 1 line / 158 bytes, i.e. the session header alone, so every call would have reported "nothing said yet" forever. `zstdFrameRanges()` walks frame and block headers (no decompression) to find exact boundaries and decompresses each frame; splitting on the 4-byte magic instead would corrupt everything after a magic sequence that happens to occur inside compressed data. zstd itself is resolved at RUNTIME (`zstdSupported()`), because it landed in Node 22.15 while the project floor is 22.0 and `@types/node` still does not declare it; without it the mode falls back to the pane exactly as before, which is why `readDeepSeekLastResponse()` distinguishes `null` ("this reader cannot run here") from an empty result ("read fine, nothing said yet"). **(2)** Every turn also records a plugin-sourced `user/message` (dsh's runtime-context snapshot: sandbox policy, approval policy, cwd), so only `source.kind === 'user'` is a real prompt. **(3)** A turn that ends in `reason.kind: 'error'` is surfaced as `Turn error: …` (and a non-error early stop such as `max-tokens` as `Turn ended: …`) rather than as an empty string, which an agent reads as "still thinking" through fifteen polls. **(4)** Reply text is assembled per (turn, step): a finalized `assistant/message` wins, and the streamed `assistant/chunk` / `text-chunks` deltas are consulted ONLY for a step that never finalized (so a partial answer is readable mid-turn without ever being appended twice) — ⚠️ and "finalized" is tracked as a SET of steps, not as non-empty text, because a step whose whole reply was reasoning strips to `''` at the `` boundary and would otherwise resurrect the raw, unstripped deltas in its place (measured on a real conversation). ⚠️ Session→transcript pairing is by the transcript's own header `cwd` plus a ±60 s boot window against the Codeman session's `createdAt`, never by reproducing dsh's directory mangling (which already has two forms on disk, `` and `session-`) and never by newest-mtime alone: mtime alone handed a freshly spawned worker its PREDECESSOR's answer in the same case directory, which is worse than saying nothing because an agent cannot tell a stale answer from a fresh one. The `DSH_HOME` override reaches the reader through the narrow `Session.deepSeekHomeOverride` getter rather than an `envOverrides` accessor, since that map can hold provider credentials; it is ephemeral by design (never persisted), so a session that overrode it and outlived a server restart resolves the default tree and reads as "nothing said yet". Tests: `test/deepseek-transcript.test.ts`. ⚠️ **This is also what makes dsh the one non-claude mode the `codeman` agent skill drives like claude** (`skills/codeman/preamble.sh`, preamble 1.20.0): with a real end-of-turn signal AND a real transcript, `spawn_workers alpha beta:deepseek` is a mixed fleet in one call and `sendwait`/`last_text` need no per-mode variant. Two traps are handled in the preamble rather than left to the agent. **Readiness is not the stop signal**: the harness reports `idle` at BOOT roughly 300 ms before its composer paints (measured 2.26 s vs 2.56 s after spawn, twice), so a send-and-wait fired straight after `quick-start` resolves on the boot edge, reports a turn that never ran, and strands the prompt in a pane that was not yet accepting input — `spawn_worker`'s dsh branch gates on the composer glyph (`❯`, overridable via `DSH_READY_MARK`) instead, after which the boot edge is spent and unobservable. And `sendwait` asks for `wait:"stop,exit"` rather than the `wait:true` default set, because that set also carries `idle`, which for any external CLI is inferred from output stabilization: on a dsh worker whose TUI repaints rarely, a re-wait resolved in 0 ms with `signal:"idle"` on a turn that had three minutes left to run. The skill also sends `deepSeekConfig.permissionMode: 'danger-full-access'` for its own workers, matching what the Run button sends, because the harness default still asks and a worker parked on an approval row cannot finish a fan-out (the multi-user clamp still applies). ⚠️ **The resolver needs the strictest identity probe of any CLI**, because `dsh` is not merely a squattable npm name: Debian ships an unrelated `dsh` (dancer's shell, `apt install dsh`) that would answer a version probe convincingly. `probeDeepSeekVersion()` therefore checks `dsh --help` against `DEEPSEEK_IDENTITY_REGEX` (`DeepSeek Harness`) FIRST and only then reads a version, and `test/deepseek-cli-resolver.test.ts` pins both the rejection and the VITEST hermeticity gate with a real executable fixture. `DEEPSEEK_VERSION_REGEX` keeps the prerelease tail (`0.1.1-rc.2`), since truncating it would report an rc as a release; it is shared with the `dsh` dependency-registry entry so doctor and run mode agree about the version even though the resolver is stricter about identity. Model is NOT a session field: it is a composition entry in the profile's config tree (`agent-default-model`), configured in `~/.dsh/settings.yaml` + `cordis.patch.yml`, so both create paths deliberately resolve no model for this mode. Env allowlist: `DSH_*` + `DEEPSEEK_*`; provider keys named by a settings-file `apiKeyEnv` stay OUT, which is pi's 34-provider-key problem in a new shape and gets the same answer. Docker seeds `~/.dsh` per-file (`.env`, `settings.yaml`, `cordis.patch.yml`) and the image installs its OWN profile, because `profiles/` is a per-profile `node_modules` tree — host-arch-specific and far too large to copy per container start. Stays OUT of `isAltScreenStripMode()` (third-party fullscreen TUI — the opencode case). ⚠️ `classifyProfile()` reads the profile's BUNDLES, and "unknown means launchable" is deliberate (anyone can publish an app bundle), but it has one knowably-wrong case: `readProfile()` returns an empty bundle list for a `package.json` with no `dsh.profile.bundles`, which made the SHIPPED `web`/`headless` profiles look third-party and launchable. The directory name is therefore consulted as a LAST resort (`STOCK_NON_INTERACTIVE_PROFILES`), after the bundle patterns, so real bundle evidence always wins over a name the user chose. The loose `tui` arm carries word boundaries for the same reason: it decides which profile boots by default, and matching the middle of `intuition` is not a rule anyone could predict. ⚠️ The generated shim is written **temp + rename**, not in place: the TUI can be exec'ing that exact path while an upgraded Codeman refreshes it, and a half-written file is a syntax error the caller then retries four times per state change forever. Bump `SHIM_VERSION` whenever `SHIM_SOURCE` changes, or an existing shim keeps matching the embedded marker and is never refreshed. Availability via `GET /api/deepseek/status`, the widest per-CLI status shape (`available`/`runnable`/`path`/`version`/`dshHome`/`defaultProfile`/`profiles`); `POST /api/deepseek/install-profile` bootstraps a profile and is the only endpoint in Codeman that installs third-party code — regex-confined specifier, argv-array spawn, privileged grant required in multi-user mode, and the held-open request is bounded by a HAND-ROLLED timeout over a `detached: true` process group (negative-pid SIGTERM→SIGKILL, as `runGit()` does in git-clone.ts). ⚠️ Node's own `spawn` `timeout` is NOT enough: a plugin install fans out into package-manager children, the built-in timeout signals only the direct child, and the survivors hold the inherited stdio pipes open so `close` never fires and the request leaks forever. User guide: `docs/deepseek-integration.md`. Tests: `test/deepseek-mode.test.ts`, `test/deepseek-cli-resolver.test.ts`. **Pi specifics** (#206, `docs/pi-integration.md`): command built by `buildPiCommand()` (`--model` — the only builder whose model regex admits `:` and `/`, for `sonnet:high` and `openai/gpt-4o` — plus `--provider`, `--thinking`, `--session ` / `-c`, and the TRI-STATE `--approve`/`--no-approve`). ⚠️ **Pi has no permission prompts and no sandbox**, so there is no `--dangerously-skip-permissions` analog and Codeman must not invent one; the privilege-shaped knob is `approveProjectTrust`, which makes pi LOAD AND EXECUTE repo-local `.pi/extensions` TypeScript and npm-install missing project packages. It therefore joins `clampExternalCliBypassForOwner()`'s **materialize** branch (gemini's, not codex/antigravity's only-if-sent one): an absent config still yields `--no-approve` for a non-granted owner, because pi's own default is an interactive prompt the session user could answer themselves. ⚠️ `--api-key` is NEVER wired — it would put a provider secret on the spawn command line. ⚠️ Pi stays **out** of `isAltScreenStripMode()`: its default TUI renders into the main screen with terminal-owned scrollback (nothing to strip), and since 0.84.0 the user can flip to a fullscreen TUI at runtime via `/settings`, where the alt screen is load-bearing — being out of the list is exactly what makes that switch safe. ⚠️ Only the `PI_*` env prefix was added; pi's ~34 provider keys share no prefix and `ALLOWED_ENV_PREFIXES` is a single GLOBAL list with no mode context, so admitting them would widen the allowlist for every mode at once (a mode-aware allowlist is the tracked follow-up). ⚠️ `pi` is a short, GENERIC binary name, so unlike the sibling resolvers `pi-cli-resolver.ts` sanity-probes `pi --version` (cached, vitest-skipped) and requires semver-shaped output; `GET /api/pi/status` carries `version` on top of the sibling `{available, path}` shape so a misresolution is diagnosable. Local echo: pi lands on the `'buffer'` overlay via the fallthrough in `_updateLocalEchoState` (pinned in `test/local-echo-codex-gating.test.ts`); if pi's live composer turns out to fight it the way codex's did, the fallback is one `'off'` branch. Tests: `test/pi-mode.test.ts`, `test/routes/external-cli-bypass-clamp.test.ts` (first-ever coverage of the clamp). **OMP (`omp`, Oh My Pi) specifics** (`docs/omp-integration.md`): architecturally the simplest of the family — omp owns its own auth, provider routing, and trust decisions entirely in `~/.omp` config files (default `tools.approvalMode: yolo`), so `buildOmpCommand()` only ever emits `--model`/`--resume `/`--continue`, and there is no bypass-permissions flag for Codeman to wire or clamp. ⚠️ **That does NOT make the multi-user clamp a no-op**: `OMP_*` is an allowlisted `envOverrides` prefix and admits `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN` (where omp resolves credentials from), both dropped for a non-granted owner in `clampEnvOverridesForOwner()` — the same shape as `DEEPSEEK_BASE_URL` — even though, unlike DeepSeek, Codeman forwards no operator-held key into an omp pane today (found in Ark0N/Codeman#353 review). `--continue` alone is ambiguous the moment any other omp conversation has touched the same working directory more recently, since it just picks the newest session file on disk — `resolveAndClaimOmpSessionId()` (`src/utils/omp-session-resolver.ts`) resolves and PINS the real id instead, verifying each candidate's own file header (`{"type":"session","id",cwd"}`, not just the mangled-directory match) and tracking already-claimed ids in a process-wide registry so two omp tabs in the same case dir can't alias onto each other's conversation. ⚠️ Resolution/pinning happens ONLY at the point a respawn is actually confirmed (`_pinOmpRespawnId()`, called from `_setupOrAttachMuxSession()`'s dead-pane branch and `reattachRemote()`) — earlier code resolved eagerly while merely building respawn options, which could mis-pin a still-ALIVE session's id purely from boot-recovery timing. `src/omp-transcript.ts` independently scans `~/.omp/agent/sessions/**/*.jsonl` for Past Sessions history, the omp analog of Claude's own transcript scan, so a conversation survives even a full "Kill Tmux". ⚠️ omp's own env knobs are mostly `PI_*`, not `OMP_*` (`PI_CONFIG_DIR`, `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`, `PI_SUBPROCESS_CMD`, `PI_SHELL_PREFIX`), and `PI_*` is already allowlisted globally for pi — so a redirected `PI_CONFIG_DIR` silently moves the `~/.omp` tree the resolver and transcript scanner hardcode, degrading pinning/history with no error; a known gap shared with pi, not fixed here. ⚠️ Docker: `appendResumeFlag()`'s `case 'omp'` keys off the top-level `resumeSessionId`, which Docker panes never receive for omp (built from `defaultDockerCommandForMode`, with no `ompConfig` threaded through) — host-side history recovery and pinning work through the shared `sessions/` mount, but `--resume` does not currently reach an in-container omp process on respawn (flagged in review, not yet fixed). Stays out of `isAltScreenStripMode()` (narrow scrollback strip, alt-screen toggles only) and lands on the `'buffer'` local-echo policy via the `_updateLocalEchoState` fallthrough, same as grok and pi. Resolver: `omp-cli-resolver.ts` version-probes like pi/grok (`omp` is a short, generic name) and requires `omp/`-shaped output; `OMP_SEARCH_DIRS` leads with `~/.local/bin` (omp.sh's installer targets `$HOME/.local/bin` with no `--dir` override — verified against a real `--no-cache` Docker build, `~/.omp/bin` was the wrong first guess). Tests: `test/omp-mode.test.ts`, `test/omp-cli-resolver.test.ts`, `test/omp-session-resolver.test.ts`, `test/omp-fresh-run-no-resume.test.ts`. **Codex input path (issues #218/#219/#220/#222)**: codex-mode sessions use **predictive write-through echo, never the buffer overlay**. The buffer overlay stays disabled exactly as 1.12.2 left it (`_updateLocalEchoState` in terminal-ui.js, same branch as shell; `_localEchoEnabled` remains false for codex), and the additive `_localEchoPolicy` field selects `'predict'` for codex when `localEchoEnabled` is on. Codex's composer is interactive per keystroke: typing "/" pops a live-filtering command picker (#222 was "picker never appears" because the "/" sat in the overlay until Enter), the composer grows/rewraps as it fills (#220: a long typed prompt existed ONLY in the overlay DOM, so codex never grew the composer), arrows and Ctrl+Backspace edit server-side state (#218: arrows were forwarded to an EMPTY composer while the typed text sat pending; the `\x08` control-char flush then left the overlay stateless so `\x7f` was swallowed as "nothing to remove"), and pastes arrive bracketed (#219: `terminal.paste()` wraps in `\x1b[200~..201~`, which the multi-byte-ESC branch forwarded WITHOUT flushing pending text, so the paste landed before it). The shared overlay branch (claude/gemini/opencode still buffer) gained three fixes: bracketed pastes flush pending text first, composer nav keys (`isComposerNavKey` allowlist in `CodemanTerminalInput` — arrows/Home/End/Delete/PgUp/PgDn incl. modifiers, deliberately excluding DA/CPR/DSR query responses) flush and hand the session to **pass-through** (plain PTY echo until Enter/Ctrl+C, because after cursor movement the append-only overlay cannot track edits), and a backspace that finds no overlay state is FORWARDED instead of swallowed. ⚠️ **Codex drops keystrokes that arrive in the same PTY read as a bracketed paste** (upstream `bottom_pane/paste_burst.rs` holds rapid chars for paste classification; verified against codex 0.147.0 by writing `hello\x1b[200~PASTED\x1b[201~` into the tmux client PTY in one write → composer shows only `PASTED`, while a 100ms gap yields `helloPASTED`), so the flush sends the typed text immediately and delays the paste sequence by 80ms — the same two-phase shape as the Enter branch's delayed `\r`. Related protocol fact: xterm.js sends `0x08` for Ctrl+Backspace, which codex's keymap binds to delete-ONE-char (`ctrl(Char('h'))`); real word-delete needs the kitty CSI-u encoding (`\x1b[127;5u`), which xterm.js 6.0.0 cannot emit (kitty support lands in 6.1.0-beta) — an upstream limitation, not a Codeman bug. E2E technique: codex 0.147 reaches its composer with any dummy key in `$CODEX_HOME/auth.json` (`{"OPENAI_API_KEY":"sk-test-..."}`), so a real TUI can be driven headlessly (envOverrides `CODEX_HOME` rides the `CODEX_*` allowlist) without real credentials. **Predictive write-through echo invariants** (the codex echo mode, `PredictiveEchoAddon` in `packages/xterm-zerolag-input`): (1) the onData hook `_predictHookOnData` is a PLAIN STATEMENT between the buffer block and Normal Mode — no `return`, try/catch-wrapped, never touches `_pendingInput` — so the wire path is byte-identical with the predictor active, absent or throwing (pinned at vm level and by an end-to-end trace-equality E2E); (2) it ships as a SEPARATE bundle `vendor/xterm-predictive-echo.js` so the zerolag bundle stays byte-identical, and a missing/broken bundle degrades codex to plain 1.12.2 echo (`typeof PredictiveEchoOverlay !== 'undefined'` guard); (3) predictions paint only while the cursor sits on the measured composer row (`isCodexComposerRow`, `CODEX_COMPOSER_ROW_RE = /^› /` — matches the empty-composer placeholder, typing, and the slash picker; rejects modal rows and 2-space wrapped continuation rows, the #220 ghost zone, which deliberately fall back to real echo); (4) reconciliation reads the PARSED buffer with `baseY + row` (xterm's `cursorY` is baseY-relative; `viewportY` only coincides while scrolled to bottom), confirms prefix-only on cell match PLUS cursor advance, cascades only on TWO consecutive foreign NON-BLANK passes (blanks are neutral: codex clears its placeholder on first echo), and TTL-bounds the rest; (5) after an UNPREDICTED wire edit (backspace into echoed text, any 'clear'-classified input, an IME/plain-paste 'text' commit, or every bypass send incl. `_handleCjkInput`) the addon holds new predictions until the next PARSED write: the displayed cursor is stale for one RTT and anchoring on it paints ghosts one cell off; (6) the per-device `localEchoEnabled` toggle is the kill switch returning exact 1.12.2 behavior. Measured constants + fixtures: `docs/predictive-echo-plan.md`, recorded via `scripts/dev/record-codex-frames.mjs` through the production tmux+strip pipeline. Tests: `test/local-echo-codex-gating.test.ts` (vm harness: nav-key + predict classifier truth tables, policy matrix, wire-neutrality pins), `packages/xterm-zerolag-input/test/` (addon laws, real-fixture replay, seeded fuzz), `test/codex-predictive-echo.test.ts` (E2E vs real codex incl. byte-identity + 300ms-RTT). ### Remote sessions over SSH **Remote sessions (SSH)**: Sessions can run the agent inside a durable `tmux -L codeman-remote new-session -A` **on a remote host** so it survives the SSH drop (COD-104), and can also **discover + attach** to `codeman-*` sessions another Codeman launched there — attached (`owned:false`) sessions **detach, never kill** on tab close (COD-105). **Shared/collaborative** (COD-106): remote set-options are scoped per-session (never `-g`) and `window-size latest` lets multiple clients attach the same session at different viewports without clamping to the smallest; a client count surfaces a "shared · N" badge. **Auto-reconnect** (COD-108): a bounded-backoff watcher re-establishes a dropped remote session's local ssh pane and reattaches the still-running durable remote tmux (kill-switch `remoteAutoReconnect`, default ON); the pure pieces (backoff schedule, per-session reconnect state, `decideReconnect` eligibility) live in `src/remote-reconnect.ts` (tests: `test/remote-auto-reconnect.test.ts`), while `tmux-manager.ts` owns the live pane probe + timers. Owned sessions propagate `kill-session` to the remote on close; non-owned never do. ⚠️ Command-injection surface (COD-107): all ssh command lines flow through the single shell-safe `buildSshConnectionArgs()` — every user field (`-J jumpHost`, `-i identity`, `-o`) is `shellescape`d; never hand-build an ssh line elsewhere. Full design: `docs/remote-sessions.md`. ### Remote SSH cases **Remote SSH cases** (COD-94/#145): cases can point at a **remote host** (`~/.codeman/remote-hosts.json` + `remote-cases.json` via `src/remote-hosts.ts`; CRUD under `/api/cases` — cases route file). A remote session launches a LOCAL tmux pane running `ssh ` that creates a durable REMOTE tmux session on a **dedicated socket** `-L codeman-remote` with name `codeman-ssh-` — deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so a Codeman instance on the target host never adopts it; no `-g` global tmux options are set remotely. `remotePath`/`identityFile` are schema-guarded against shell injection (backticks/`$` rejected — same approach as `extraSshOptions`); remote tmux availability is probed via `checkRemoteTmuxAvailable()` in quick-start (ssh args carry `-o ConnectTimeout=10`). Remote claude defaults to an idempotent `claude --session-id || claude --resume ` pair under a login shell, so a respawn or reattach continues the SAME conversation rather than starting a fresh one (remote omp gets the same treatment via `--continue`; ⚠️ because the claude arm is an `a || b` pair under `-c`, that pane's PID is the login shell, not the agent); per-host `commands.*` override. Session kill best-effort kills the remote tmux too. `SessionState.remote`/`MuxSession.remote` round-trip through recovery (`restoreMuxSessions` passes `remote` back into the Session constructor). ⚠️ Run flows must route remote cases through `POST /api/quick-start` (which resolves the remote case and skips LOCAL CLI availability gates) — `POST /api/sessions` stat-validates `workingDir` locally and has no `caseName`. `envOverrides`/`effort`/`modelOverride`/`codexConfig`/`geminiConfig` are rejected for remote quick-starts (not silently dropped). UI: Create Case modal → Remote tab. Tests: `test/remote-hosts.test.ts`, `test/remote-ssh-options.test.ts`. ⚠️ **Reading a file in a remote case goes over ssh too** (#415): `src/remote-files.ts` is the single remote-READ layer (`buildRemoteFileCommand` = `buildSshConnectionArgs` + one shellescaped remote command; `remoteProbePaths` returns remote realpath + stat; `remoteCreateReadStream` streams a `Range` via `tail -c +N | head -c L` and its `close()` must be wired to the response's `close` or the ssh child outlives an aborted download). The guard order matches the local path exactly (`validateSessionFilePathLexical` → remote realpath of BOTH file and workspace root → containment → sensitive-path → size cap on the REMOTE size), a request path arrives from the browser and is only ever interpolated as a `shellescape`d token, and an unreachable host answers **502**, never a 404. ⚠️ The probe's symlink resolution FAILS CLOSED: `readlink -f` where it exists, otherwise a `cd -P`/`pwd -P` directory walk plus a bounded plain-`readlink` loop over the last component, and anything it cannot fully resolve is reported unresolvable (404), never as the unresolved string — the first version resolved the directory chain only, so on a host without `readlink -f` a `ws/notes.txt -> ~/.ssh/id_rsa` link passed containment under its own path while `cat` served the key. Records are NUL-separated and index-keyed so a newline in a filename cannot shift the mapping. ⚠️ ssh children are BOUNDED: probes and buffered reads go through `src/remote-ssh-limiter.ts` (a `document-conversion-limiter`-shaped semaphore, default 4), the attachment-history list probes its whole history in ONE batched call (`probeRemoteAttachmentHistory`, threaded into `registerExternalAttachment({remoteProbes})`), and probes chunk at 40 paths — a prompt-injected agent printing `codeman://attach` links in a remote session used to fork one `ssh` per link. `describeExecError` never returns Node's `Command failed: ` message (identity path + probe script in a 502 body). The `PUT /file-content` guard sits AHEAD of `validateSessionFilePath`, which resolves LOCALLY, or a same-named local directory (an sshfs mount) takes the write. Under `VITEST` the three IO functions refuse rather than connect. This covers the ATTACHMENT routes too, which is the half a clicked path needs when the file is OUTSIDE the case directory (`_isExternalPreviewPath` sends it to `POST …/attachments`): registration, by-id `raw`, metadata and the history list all resolve over ssh (`registerExternalAttachment({remote})`, `resolveServableRemoteAttachment`), and what decides the host is the SESSION, never the path string — the same absolute path means a different file on each host. Deliberately NOT supported over ssh: writes (`edit=1`/`PUT` answer 400, `editable` is always false), office previews/thumbnails, the file tree/picker, `tail-file`. Tests: `test/remote-files.test.ts`, `test/routes/file-routes-remote.test.ts`. ### Docker cases **Docker cases** (shipped 1.4.0; user guide `docs/docker-cases.md`, design `docs/docker-cases-plan.md`): a case can point at a **container** instead of a local/remote path, and any of the CLI run modes runs INSIDE it. Like remote-SSH, it is a **LOCATION OVERLAY on cases, never a `SessionMode` of its own** (`SessionMode` is unchanged). Storage `~/.codeman/docker-hosts.json` + `docker-cases.json` via `src/docker-hosts.ts` (direct mirror of `remote-hosts.ts`: `readDockerHosts`/`readDockerCases`, `toSessionDocker`, `dockerDisplayPath`, and the PURE builders `buildDockerBaseArgs`/`buildDockerCreateArgs`/`containerApiUrl`/`hostGatewayAlias`/`dockerConfigHash`). CRUD `/api/docker-hosts` + `/api/cases/docker-link`, plus **one-click** `/api/cases/docker-quickcreate` (Create New "Run in Docker" checkbox → case folder in `CASES_DIR` + auto-provisioned shared `default` host + auto-start a session inside; an expandable Template picker Small/Medium/Large/GPU or any override creates a per-case `q-` host), and export/import (`/api/docker-cases/:name/export`, `/api/docker-cases/import`, `GET/DELETE /api/docker-exports`) — all in `case-routes.ts`. Run flows route through `POST /api/quick-start` like remote (session-routes.ts docker branch, skips LOCAL CLI-availability gates). **Launch model**: exactly one long-lived container **per case** (`codeman-case-`, PID1 `sleep infinity` under `--init`); a LOCAL tmux pane runs `docker exec -it` into a **durable in-container tmux** on dedicated socket `-L codeman-docker`, session `codeman-dkr-` (deliberately fails `SAFE_MUX_NAME_PATTERN` so a Codeman running INSIDE the container never adopts it, exactly like remote's `codeman-ssh-`). Builders `buildDockerLaunchCommand`/`buildDockerKillCommand` in `tmux-manager.ts` (image-check → `docker inspect||create` → start → exec, all idempotent). The container is **shared by all sessions of the case**: `buildDockerKillCommand` kills ONLY that session's in-container tmux session, NEVER `docker stop` while siblings remain; `docker rm -f` happens only on case-delete (plus an instance-scoped boot reaper keyed on the `codeman.instance` label). **Two-layer durability/resume** (the central design point): (1) Codeman-PROCESS restart with the container still up → `tmux new-session -A` reattaches the SAME live agent (paneCommand ignored); (2) container stop/reboot/OOM → inner tmux is gone, so the re-run pane command resumes the conversation from the bind-mounted transcript: claude mode pins a DETERMINISTIC conversation id via `claudeDockerPaneCommand()` (`tmux-manager.ts`) — fresh launch `claude --session-id || claude --resume ` (a duplicate `--session-id` exits 1 "already in use", so the fallback RESUMES after a container stop; verified CLI behavior), explicit resume `--resume || --session-id ` so a stale id never dead-panes (leading `exec ` is stripped — an exec'd first branch could never fall back); codex `resume ` / gemini `--resume` keep `appendResumeFlag`. The resume id rides `resumeSessionId` through create/respawn options and persists on `DockerCase.lastClaudeSessionId` via `persistDockerCaseClaudeSessionId()` (written at quick-start launch, and again on hook/last-response conversation-id adoption so post-`/clear` switches track; seeded back when `resumeOnStart`, default true); `-A` makes the pane command self-selecting (inert on reattach, active only when tmux was re-created). **Config drift** (`dockerConfigHash` → `codeman.confighash` label): quick-start compares via `checkDockerConfigDrift()` and REFUSES a drifted launch with `CONFLICT`; the UI confirm calls `POST /api/docker-cases/:name/recreate` (refused while case sessions are live) which `docker rm -f`s so the next launch recreates with the new config — host config edits actually take effect. **Workspace** is a REAL host dir bind-mounted at the SAME absolute path (mirror, `dst==src`), so `Session.workingDir = hostWorkspacePath` keeps file-routes/attachments/watchers on real host bytes AND the in-container transcript projHash matches the host so subagent/workflow correlation (and thus resume-id capture) works; `resolveMuxAttachCwd` returns `/tmp` for docker (the local pane only runs `docker exec`). **Creds** arrive commit-safe and ISOLATED (1.4.1; replaced the whole-dir RW mounts that let in-container CLIs write refreshed tokens/state back to the host): shared RW across the boundary is ONLY what host-side reads/resume need (`~/.claude/projects` transcripts; codex `sessions/` + `history.jsonl` for response-viewer/`codex resume`); everything else is SEEDED (RO mount, copied into container HOME once at launch via `[ -e ] || cp`; the container refreshes its own copy and never writes back): `~/.claude.json` is merged through `buildSeamlessClaudeConfig()` (forces `hasCompletedOnboarding` + theme + workspace trust, so no login wizard/theme picker/trust prompt inside the container), plus `.claude/{.credentials.json,settings.json,stats-cache.json}`, plus whole-dir seeds for `~/.gemini`/`~/.config/{gcloud,opencode}` (`resolveDockerClaudeArtifacts`/`resolveDockerCredentialArtifacts` in `docker-hosts.ts`). Bind mounts are physically excluded from `docker commit`, so exports stay secret-free; API-key CLIs get exec-time NAME-ONLY `--env OPENAI_API_KEY` (no `=value`); the SEALED profile is `mountCredentials:false` + `network:none`. NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket. **Hardening** on every create: `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit`, `--memory`==`--memory-swap`, non-root via `--user :0` (Linux, GID 0 for writable HOME) / `--userns=keep-id` (podman rootless) / baked uid (Docker Desktop), `--pull=never`, `--init`. Base image `codeman/agent:base` is BUILT LOCALLY from `docker/agent.Dockerfile` (node22 + tmux + every enabled npm CLI from the registry, since the `CLI_NPM_PACKAGES` build arg is generated from `stock.ts` and entries carrying `agentImageLayer` or no `npmPackage` get their own layers, see `docs/docker-cases.md`; OpenShift arbitrary-uid HOME, `C.UTF-8` locale so tmux/Ink render real box-drawing glyphs; Codeman also sets `LANG`/`LC_ALL` at run time for containers built before that line) via `scripts/build-agent-image.mjs` OR **auto-built on first use** (1.4.1: `ensureAgentBaseImage()` in `docker-hosts.ts`; idempotent + concurrency-safe, only the DEFAULT image ref is ever auto-built, `--pull=never` stays absolute; build output streams over SSE `docker:imageBuildStarted`/`imageBuildProgress`/`imageBuildComplete`/`imageBuildFailed`, and quick-create returns `imageBuilding:true` while the first launch awaits the gate); tmux-in-image is a HARD gated prerequisite (`checkDockerTmuxAvailable`), never a silent bare-exec fallback. **Hooks + model**: the workspace-scaffolding block DOES run for docker (writes `.claude/settings.local.json` + the CLAUDE.md scaffold into the real host dir), so `modelOverride` works via `settings.local.json` — it is a `QuickStartSchema` field applied for local AND docker quick-starts (`updateCaseModel`), sent by the frontend docker run path (the one deliberate difference from remote, which rejects it); `effort`/`envOverrides`/`codexConfig`/`geminiConfig`/`openCodeConfig` stay rejected. In-container hook curls hit `containerApiUrl(process.env.CODEMAN_API_URL, engine)` (swaps ONLY the hostname to the gateway alias, preserving scheme+port so prod HTTPS still works); the host guard allowlists both `host.docker.internal`/`host.containers.internal` (`DOCKER_HOST_GATEWAY_ALIASES` in `network-auth-policy.ts`). ⚠️ On a **loopback-only** bind (the prod default) a container cannot reach 127.0.0.1, so in-container hooks fire ONLY when `CODEMAN_DOCKER_BRIDGE_HOOKS=1` — an opt-in SECOND listener on the docker bridge gateway (`_startDockerBridgeHooksListener` in `server.ts`; gateway auto-detected via `detectDockerBridgeGateway`, or set `CODEMAN_DOCKER_BRIDGE_HOST`) that serves ONLY the hook endpoints (403 for any other path) into the same secret-gated pipeline; otherwise idle detection falls back to output-based through the docker-exec PTY. Container-set `CLAUDE_CODE_TMPDIR` keeps claude launching regardless of workspace path. `SessionState.docker`/`MuxSession.docker` round-trip through recovery. Every docker IO path is `IS_TEST_MODE` (VITEST) no-op'd; the pure builders are unit-tested. **Export/import** (`src/docker-export.ts`): full-image (`docker commit` + `save | gzip` + workspace tar + manifest) or workspace-only → one portable `~/.codeman/docker-exports/-.codeman-container.tgz`; import validates per-member sha256, traversal-guards the workspace tar, `docker load`s + quarantine-retags the image (`codeman/imported-:`, never overwriting a local tag); a `saveImageToTar` stream `pipeline` avoids truncation. **GPU** passthrough (`gpus` → `--gpus`, needs the NVIDIA container toolkit) and **elastic disk** (no `--storage-opt` cap, so container storage grows with data). SSE `docker:exportComplete`/`exportFailed`/`importComplete` (both registries). **UI** in `session-ui.js`: Create Case **Docker** tab (collapsed/compact form since 1.4.1), the one-click checkbox + Template picker, short `(docker)` case-menu tags, and a Manage-tab Export button; docker AND remote sessions name their tabs `w-` via the shared `_nextCaseSessionStartNumber()` so all tabs follow one naming convention. **Adopting an already-running container** (`DockerCase.owned === false`, `POST /api/cases/docker-adopt` + the read-only `POST /api/docker-cases/adopt-preflight`, `GET /api/docker-hosts/:hostId/containers`, `POST /api/docker-cases/browse`): the mirror of remote-SSH's `owned:false` attach. For an adopted container the launch chain only LOOKS and then execs — no image gate (the image is theirs), no create, and above all no `start`, since starting a container we do not own is precisely the mutation adoption promises never to perform; a missing or stopped container fails closed with an actionable message. Credential seeding is skipped too (those copies read from create-time read-only mounts that do not exist here, and writing host credentials into someone's container is not ours to do), so its CLIs must already be authenticated inside it. Absent `owned` = owned, so every pre-existing case is byte-identical. ⚠️ The guarantee is NEGATIVE, so it cannot be observed by using the feature — only by asserting the mutating verbs are absent — and it is therefore enforced at four deliberately independent layers: `buildDockerStopCommand`/`buildDockerRemoveCommand` throw during pure STRING CONSTRUCTION (no shape of caller bug can produce a `docker stop`/`rm` for a container we do not own), `removeDockerContainer` refuses again at the lowest layer, `checkDockerConfigDrift` reports "none" (an adopted container carries no `codeman.confighash` label, so a real comparison would always report drift and the launch gate would 409 forever, offering a recreate we may not perform), and the orphan reaper skips it through a check independent of the two conditions that already cover it. ⚠️ **The export path is the one place that still touched the container** and both halves had to be closed: a full export `docker commit`s it (refused for an adopted case — it packages someone else's container, with their logins, into a bundle Codeman hands out) and even a workspace-only export `docker pause`d it first for snapshot consistency (skipped: the freeze stops the owner's processes for as long as the tar takes). ⚠️ `owned` is applied AFTER the config hash; `dockerConfigHash` takes an explicit field list, so ownership can never shift an existing case's hash and mass-trip the drift gate, whose only remedy is "recreate the container". ⚠️ The container workdir is verified INSIDE the container: it defaults to `hostWorkspacePath` for an OWNED case only because the create-time bind mount puts the host directory at that exact path, and adoption mounts nothing, so the two are independent facts — without the check `docker exec --workdir ` fails with an OCI chdir error the pane surfaces as a bare `execvp failed`. ⚠️ Run-mode availability comes from the CONTAINER (`availableModes`), live-probed rather than trusted from attach time: gating the dropdown on host CLI availability (#201) is right for local sessions and wrong here. The probe modes and the BINARY each mode looks for both come from the CLI registry (`enabledCliIds()` / `discovery.binaries[0]`), never a local table — a hand-written list silently froze once already, missing `omp` and hiding that mode on every docker case; `antigravity` ships as `agy` and `deepseek` as `dsh`, so a mode-name probe reports both missing on a container that has them, and a mode with no binary (`shell`) is reported available without a lookup. ⚠️ **A FAILED probe means opposite things per ownership.** For an adopted case it is a real fault (only the user can start that container). For an owned case it is the NORMAL state before the first session — the launch chain creates the container on demand — so recording it as an error hid every agent mode on every freshly linked Docker case behind "start it yourself first", for a container Codeman was about to create itself; `CaseInfo.docker.owned` exists on the wire so the frontend can tell the two apart. ⚠️ Claude launches WITHOUT `--dangerously-skip-permissions` when the container's exec user is root: Claude Code refuses the flag as root ("cannot be used with root/sudo privileges", still true in 2.1.261) and the refusal is visible only inside the container, so the pane just dies. Our base image runs a non-root user and never hits it; an adopted container's user belongs to its owner and is frequently root. Which flag to drop is a per-CLI fact, so it is `overlays.docker.rootCommand` in the registry rather than an id branch. ⚠️ **Admin-only in multi-user mode**, unlike `docker-link` right next to it: linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has already confined to the caller's space, while an adopted container's mounts are whatever its owner gave it — one mounting `/` hands the adopter a shell over the whole host, defeating exactly the workspace scoping that mode exists to enforce. The container listing and the in-container directory browser are gated with it (both are machine-level reads over containers belonging to anyone); the preflight is NOT, because the run menu probes it for every docker case, so it admits a non-admin only for a container already linked to a case they can access. Tests: `test/docker-adopted-container.test.ts`. Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/docker-export.test.ts`, `test/network-host-guard.test.ts`. ## Session data and lifecycle ### Input delivery and WS resilience **Input**: `session.writeViaMux()` for programmatic/curl input — tmux `send-keys -l` (literal) + `send-keys Enter`. Single-line only (fire-and-once). Interactive **browser** input goes through a durable **exactly-once** layer: each frame carries a stable `clientId` + monotonic per-session `seq`, persisted to localStorage until the server ACKs (`{t:'ia',seq}` over WS, or HTTP 2xx), so a dropped link/reconnect can't lose or double-deliver a prompt. **WS resilience** (#149): the upgrade URL carries `cid = clientId + ':' + perTabNonce`, and `ws-connection-registry.ts` supersedes only same-TAB reconnects (two tabs on one session coexist; input frames keep the bare `clientId` for seq dedup); reconnects back off exponentially (attempts preserved across `_connectWs`), and the header connection chip renders from a real `_wsState` lifecycle (`connecting`/`connected`/`fallback`/`reconnecting`/`disconnected`). ### Per-session env overrides: exact-key allowlist and CLAUDE_CONFIG_DIR **The env allowlist has two tiers, and exceptions go in the exact-key tier, never a widened prefix** (#255): `ALLOWED_ENV_PREFIXES` in `src/web/schemas.ts` carries the CLI-namespace prefixes, and `ALLOWED_ENV_KEYS` carries exact keys (currently only `CLAUDE_CONFIG_DIR`). `CLAUDE_CONFIG_DIR` relocates the Claude CLI's user config (credentials, settings, stats), which is how one machine runs sessions on separate Claude subscriptions: point a case's sessions at e.g. `~/.claude-clients/acme` via `envOverrides` and run `/login` there once against the client's account. The exact match matters: `CLAUDE_` as a prefix would open every future Claude CLI variable unreviewed, and near-misses (`CLAUDE_CONFIG_DIR_EXTRA`) stay rejected (`test/env-overrides-schema.test.ts`). No new security boundary is crossed: sessions already run as the server's OS account, and `applyEnvOverrides()` shellescapes values into socket-scoped `tmux setenv`. Two carry rules: **(1)** the key must survive `getEnvOverridesForPersist()` in `session.ts` (it is a path, not a secret; dropping it from state.json would silently move a rebuilt-after-reboot session back to the default account); **(2)** ⚠️ a relocated config dir writes transcripts outside `homedir()/.claude/projects`, which `subagent-watcher.ts`, `workflow-run-watcher.ts`, the response-viewer routes and Read My Mind capture all hardcode — those surfaces go blind for such a session. Documented workaround: symlink the transcripts back into the shared tree (`ln -s ~/.claude/projects /projects`), keeping credentials separate while the watchers keep working. ### Agent wait primitives **Agent wait primitives** (`GET /api/sessions/:id/wait`, `GET /api/sessions/:id/wait-output`, and the `wait`/`waitTimeout` fields on `POST /api/sessions/:id/input`): bounded long-polls that let an agent driving Codeman from a shell tool block until something happens. They exist because SSE was the only "tell me when" channel Codeman had, and a curl-driven caller cannot practically hold a stream and parse events inline. The blocking core is `src/web/session-wait-registry.ts` (no IO, no `Session` reference, so it unit-tests in isolation), bounds live in `src/config/agent-wait.ts`, and the wiring is three `notifySignal()` calls next to existing broadcasts (`session-listener-wiring.ts` for `working`/`idle`/`exit`, `hook-event-routes.ts` for `stop`/`blocked`) plus `notifyOutput()` riding the already-attached `terminal` listener. Design: `docs/agent-control-plan.md` §3; wire contract: `docs/api-reference.md`. **Ordering rules, both load-bearing and both invisible to a reader of either side alone.** ⚠️ **exit-before-cancel**: `_doCleanupSession` in `server.ts` must call `sessionWaits.notifySignal(id, 'exit')` BEFORE `sessionWaits.cancelAll(id)`, because that method detaches the session's listeners before `session.stop()`, so on a delete the PTY's own exit event never reaches the registry and an `until=exit` caller would get a bare `ended: true` instead of the signal it asked for. Found by live-testing the delete path, not by the unit tests. ⚠️ **registered-before-write**: the send-and-wait path on `POST .../input` registers the waiter BEFORE writing to the PTY, and that ordering is the entire reason the combined endpoint exists rather than documenting "POST, then GET .../wait": between the write and the session flipping to `working` there is a window in which a separate wait sees the session still idle and instantly reports the PREVIOUS turn as this turn's answer. Two consequences hang off it: `useMux` delivery is awaited on that path (the response is staying open anyway, so a `writeViaMux` failure becomes observable for the first time) while the non-wait path keeps its fire-and-forget shape byte for byte, and because `shouldApplyInput()` MUTATES (it records the seq) before registration, a registration that fails on a full pool must `forgetInputSeq()` before returning, or the caller's retry is rejected as a duplicate and the input is lost by the very mechanism reliable delivery exists for. **A timeout is a 200, deliberately.** `{"timedOut": true, "signal": null}` with HTTP 200 is the long-poll succeeding at answering "did this happen within N ms?" with "no". The documented client pattern is a loop over short waits (`DEFAULT_WAIT_MS` is 60s precisely because prod is reached through `tailscale serve` and users run cloudflared, both of which cut idle connections), and turning every poll boundary into a 4xx would make that loop indistinguishable from a real failure. The alternatives are all worse: `408` is auto-retried by several clients and proxies, silently doubling the polling load; `504` is what a genuine tunnel failure looks like, so reusing it destroys the caller's ability to tell the two apart; `204` cannot carry `waitedMs`/`status`/`limitPaused` and breaks the uniform envelope the versioning policy makes a stable promise. Errors are reserved for `INVALID_INPUT` (400), `NOT_FOUND` (404, also the multi-user ownership answer via `findSessionOrFail`), `SESSION_BUSY` (409, this session's cap) and `RATE_LIMITED` (429, the per-owner or process-wide cap: a global cap reported as `SESSION_BUSY` tells the caller to switch sessions, which cannot help). ⚠️ The clamp is silent, so the EFFECTIVE timeout is echoed back as `wait.timeoutMs`: a caller that asked for 30 minutes, got the 600s ceiling, and could not see it would read the timeout as "the worker is wedged" and kill a session that was working fine. All three endpoints nest the result under `data.wait` for the same reason, so one client helper works against any of them instead of an agent's `is_done()` reading `undefined` off the shape it did not expect. **Matching is literal, and that is a language constraint, not a missing feature.** `search-service.ts` already avoids regex so there is no ReDoS surface, and this endpoint is more exposed still: the pattern is caller-supplied and the input is a live stream. herdr can offer `--regex` on its equivalent because Rust's regex crate is linear-time with no backtracking; JavaScript's `RegExp` backtracks, so the same feature here is a denial-of-service primitive. A `regex` parameter is therefore REJECTED with a 400 rather than ignored, since an agent that assumed otherwise would silently wait on the wrong thing. `match` is capped at 200 chars (`MAX_MATCH_LENGTH`), which is also what keeps the per-waiter carry buffer small: a match can straddle two PTY chunks, so each waiter carries `match.length - 1` characters of the previous chunk and tests `carry + chunk`. ⚠️ **`from=now` does not mean "printed after you asked"**: tmux repaints the visible screen on attach, on resize, and on any TUI redraw, and a repaint arrives as ordinary `terminal` data, so text already on screen can satisfy a fresh wait (observed live: a marker echoed a minute earlier matched instantly). This is inherent to running the agent under a multiplexer and is not fixable in the registry, so the contract is a marker unique per call (`echo DONE_$RANDOM`), never a generic one like `BUILD OK`, and every recipe must show that. ⚠️ **There is exactly ONE definition of the matched stream, `normalizeForMatch()`**: `stripAnsi()` (CSI, OSC, `ESC =`/`ESC >`) plus `ANSI_ESCAPE_RESIDUE` for what that helper leaves behind — above all the `ESC ( B` charset switch a stock bash prompt emits on every line, which in the first build survived into the matched text and made `match=tnode:` silently fail against a prompt that plainly renders `tnode:`. The haystack, the carry and the snippet window all derive from that one function, so matching and the snippet cannot drift apart again. ⚠️ **`splitTrailingEscape()` holds back a partial escape at a chunk boundary** until its tail arrives; its `INCOMPLETE_ANSI_TAIL` pattern must stay in lockstep with `ANSI_ESCAPE_RESIDUE` (a chunk cut between the `(` and the `B` is otherwise a fresh way to smuggle an escape into the haystack), and it is deliberately non-global (the repo-wide `lastIndex` hazard). With the carry, a match may straddle PTY chunks: `printf STRAD; sleep 1; printf DLEQQ` is matchable as `STRADDLEQQ` (measured live, both `from` modes). ⚠️ **The matcher still sees the byte stream, not the rendered pane** (`GET .../terminal` is a tmux capture, `source:'mux-visible'`): linear output agrees once escapes are stripped, but Claude Code positions words with cursor moves instead of spaces, so TUI text can arrive space-less (`Quicksafetycheck:Isthis...` — measured; some phrases keep their spaces depending on how the TUI drew them), which is why the documented advice remains ONE short space-free token the caller printed itself. The returned `snippet` is a RENDERING of the matched window, not a quotation: `SNIPPET_CONTROL_BYTES` drops bare control bytes that carry no ESC (BEL, NUL, backspace — `normalizeForMatch` removes escape SEQUENCES only) and blank runs are collapsed, because the consumer is an agent piping it through `jq` into its OWN pane, where a worker's raw bytes could otherwise reset or garble the orchestrator's display. **Lifetime discipline, per the 24-hour-session rules.** Every waiter owns exactly one timer, cleared on resolve; per-session waiter sets are deleted when they empty; `cancelAll()` runs on session teardown and `cancelEverything()` in `stop()`. ⚠️ **Waiter timers are deliberately NOT unref'd**, the opposite of the usual advice: an unref'd timer would let the process exit mid-wait and strand the HTTP response, so shutdown resolves waiters explicitly instead. All three routes also release the waiter when the client disconnects (`abortOnClientHangUp()` in `session-routes.ts`), which the caps make load-bearing rather than tidy: the documented loop-over-short-waits pattern is naturally written as `curl --max-time 30 ".../wait?timeout=60000"`, and without it every iteration abandons a waiter that lives out its full timeout, so the seventeenth call gets a 409 for a session nobody else is waiting on. ⚠️ **That listener goes on `reply.raw`, guarded by `writableFinished`, NOT on `req.raw`** (the obvious choice, and the one the SSE route in `server.ts` can afford because it only ever serves a GET). `req.raw` emits `close` as soon as the REQUEST BODY has finished streaming, which on a POST happens before the handler blocks: measured at +1ms with `aborted: false`, indistinguishable from a real hang-up, so wiring it there cancels every send-and-wait instantly and silently kills the feature, while GET keeps working because a GET has no body to finish. `reply.raw` emits `close` both on a completed response and on a dead socket, and `writableFinished` is the only thing that separates them, so the guard is load-bearing rather than defensive. ⚠️ `app.inject()` never emits `close` at all, so none of this is observable in a route test: the regression test has to bind a real port. `MAX_WAITERS_PER_SESSION` (16) is a COMBINED signal-plus-output budget (`waiterCount()` sums both maps), not 16 of each; `MAX_WAITERS_PER_OWNER` (48) applies only when an owner is passed, so single-user mode behaves exactly as before; `MAX_WAITERS_TOTAL` (128) mirrors `MAX_SSE_CLIENTS` in `map-limits.ts`, since each pending waiter costs an open HTTP response plus a timer. Capacity is asserted BEFORE the expensive work on both sides: `waitForOutput()` checks before the `initialText` scan, and the `wait-output` route checks before reading `session.terminalBuffer`, whose getter joins the whole 32MB accumulator, so a request that is going to be rejected never pays for a buffer materialization. `from=buffer` scans only the tail (`MAX_BUFFER_SCAN_BYTES`, 256KB) because the question it answers is "did this appear recently", not "ever", and the tail is continuous with the live stream (`append()` and `emit('terminal')` receive the same bytes). **Signal availability is decided by MODE, not by `isExternalCliMode()`.** `stop` and `blocked` come from Claude Code hooks, so `hooksAvailableForMode()` is true only for `claude`: `shell` is not an external-CLI mode but installs no hooks either, so `until=stop` on a shell session is a guaranteed unresolvable wait dressed up as a timeout. The behavior split is deliberate and must survive refactors: an EXPLICIT request for an unavailable signal is a 400 naming the mode, while the DEFAULT set (`stop,idle,exit`) silently drops them and echoes the narrowed set back as `wait.until`, because omitting the parameter must never 400. Signal quality is not uniform either: `stop` is definitive (Claude Code says the turn is over), `idle` is inferred from output stabilization plus prompt detection and can flap mid-turn when a spinner pauses, which is why `stop` is the documented default to orchestrate on and `idle` is the fallback for sessions that emit no hooks. ⚠️ **`idle` being ACCEPTED for a mode does not mean it ever FIRES there.** `startShell()` emits exactly one `idle` on a 500ms readiness timer and nothing afterwards, so a shell session sits at `status:'idle'` no matter what its pane is doing; since send-and-wait and `fresh=1` both require a TRANSITION, both can only time out on a shell worker (measured: a default `wait` on `sleep 4` burned its full 25s). The ❯-prompt and spinner detection that drives the real `working`/`idle` cycle is Claude's output format, so hook-less modes synchronize with `wait-output` markers or `exit`, and the docs must say so rather than listing `idle` as "available" and letting the reader infer it is usable. ⚠️ **Signals are edge-triggered with no history**, and that is a real orchestration limit: a signal that fires while no waiter is registered is gone, unobservable by any later wait variant (`until=stop` after the turn ended just times out, `fresh` or not — measured, R2-A). Documented client patterns must therefore register the waiter before the event can fire (send-and-wait) or gather on latched `wait-output` markers with `from=buffer`; the skill's fan-out flow was rewritten accordingly, and "fire-and-forget N prompts, then gather signal-waits sequentially" must never be documented again. The durable fix, a server-side latched last-signal-per-turn, is deferred with Part 3 of `docs/agent-control-plan.md`. ⚠️ Relatedly, the route corrects liveness that `SessionStatus` cannot express: `currentSignalFor()` answers `exit` whenever `pid === null` (exited, detached, or created-and-never-started) **or the mux pane is dead**, because `Session` parks a DEAD PTY at `_status = 'idle'` and trusting the status would answer the default wait `{signal:'idle', immediate:true}` for a crashed worker while `until=exit` blocked forever on an event that already happened. ⚠️ **`pid` alone cannot carry liveness for a tmux-backed session**: that pid is the local `tmux attach` CLIENT, so a worker exiting inside its pane leaves `pane_dead=1` with the client alive and `pid` never goes null — the `pid === null` branch is unreachable in the normal configuration (unit tests exercise it because `MockSession` sets `pid` by hand; only a live instance showed the gap). Liveness is therefore probed at the mux layer: `workerIsDead()` consults `mux.isPaneDead(muxName)` with a ~750 ms per-pane cache, ONLY on blocking waits (measured: 0 tmux execs across 100 non-wait input POSTs — the browser hot path pays nothing), plus a refcounted 3 s `watchForDeadWorker` interval so a worker dying while a wait is parked resolves it in ~3 s instead of burning the timeout. The probe fails SAFE by construction ("cannot tell" is never "dead": non-mux sessions, a missing or throwing `isPaneDead`, all return false). On send-and-wait, a "successful" `send-keys` into a dead pane additionally overrides `delivered` to `false` and rolls the dedup seq back (`undoOnFailure`), because the bytes went nowhere and a retry against a restarted worker must not be refused as a duplicate. The cost is that a just-created session reads as `exit`, which the wire docs must spell out as "not started yet"; the fix lives at the route rather than in `signalForStatus()` because the registry deliberately holds no `Session` reference. ⚠️ Two `claude`-mode cases still lose hooks for reasons outside the registry: a Docker case cannot reach a loopback-bound Codeman without `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, and a remote-SSH case runs the agent on another host whose hooks may never reach this server. Bounds are env-overridable (`CODEMAN_WAIT_MAX_MS`, `CODEMAN_WAIT_DEFAULT_MS`, `CODEMAN_WAIT_MAX_PER_SESSION`, `CODEMAN_WAIT_MAX_PER_OWNER`, `CODEMAN_WAIT_MAX_TOTAL`, `CODEMAN_WAIT_BUFFER_SCAN_BYTES`) and each is clamped to a hard bound so a typo degrades to the default instead of disabling the protection; they are internal tuning knobs like the rest of `src/config/`, NOT part of the SemVer-covered env-var surface in `versioning-policy.md`. Tests: `test/session-wait-registry.test.ts`, `test/routes/session-wait-routes.test.ts`, `test/routes/session-wait-output-routes.test.ts`, `test/routes/session-input-wait.test.ts`. ### Auto-resume on usage limit **Auto-resume on usage limit** ("token pause" control, opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit ("5-hour limit reached ∙ resets 8pm" and all 1.0.x–2.1.x variants), `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time from cleaned output; `SessionAutoOps` arms a timer for reset+2min, then sends Esc (dismisses the rate-limit dialog) + `continue`. Still-limited responses re-arm the loop (5-min retry on stale times); a `working` transition cancels it. Claude-mode only (detection rides `_processExpensiveParsers`). Persists/recovers via `SessionState.autoResumeEnabled`/`autoResumeAt`; respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected` — prevents `/clear` from wiping the paused conversation). Endpoint: `POST /api/sessions/:id/auto-resume`; SSE: `session:limitPauseScheduled`/`limitResume`/`limitResumeCancelled`. Tests: `test/usage-limit-patterns.test.ts`, `test/session-auto-resume.test.ts`. ### Plan-usage chip (statusLine telemetry) **Plan-usage chip** (`showPlanUsageLimits`, per-device: desktop default **ON** since 1.9.3, handhelds OFF) renders compact Claude and Codex provider rows. Claude Code (v2.1.80+) pipes a JSON blob to a configured `statusLine.command` on each render; on Pro/Max it carries a `rate_limits` object (`five_hour`/`seven_day` windows only — no Opus weekly field — each `{used_percentage 0-100, resets_at epoch-SECONDS}`). ⚠️ **Injected as an EPHEMERAL `claude --settings` CLI flag at spawn (2026-09-07), never written to disk** — `resolveStatusLineCliCommand()`/`ensureStatusLineExporterScript()` in `hooks-config.ts` (`generateStatusLineCommand()`/`applyStatusLineConfig()` remain, but only as the legacy disk-write self-heal path: a workspace an older Codeman build touched gets its stale `.claude/settings.local.json` entry stripped the first time a session starts there again). The exporter WRAPS a user's own real statusline (`findEffectiveUserStatusLineCommand()`, walking Claude Code's own settings precedence) rather than replacing it, and POSTs the `rate_limits` blob to `POST /api/status-telemetry`. That route (auth-exempt like `/api/hook-event` — localhost-only, hook-secret-gated whenever auth is active, COD-91) parses via `usage-telemetry.ts` (pure, unit-tested), broadcasts SSE `session:statusTelemetry` (de-duped per session by `telemetrySignature` since the statusline fires on every assistant message), and returns a compact plain-text footer for the exporter to **print-through** (foreground POST in the no-wrap branch so its own stdout becomes the footer, printing NOTHING on failure, `curl -sfk` plus `|| true`, since a bare brand word on the statusline is what discussion #405 opened with; backgrounded — `>/dev/null 2>&1 ` for every pane Codeman spawns); lifecycle name/mode resolution is first-seen-wins (the log returns entries NEWEST-first). No terminal buffers in the response (unlike `/api/sessions`). Consumed by the Cmd+K Session Manager (#146). Session Manager polish (COD-162/#157, 1.6.0): **pinning** via `POST /api/sessions/:id/pin` (`session:pinned` SSE; killing a pinned session demotes it to a lightweight stopped record that stays visible/resumable, and cleanup skips pinned records); **cross-device tab order** via `PUT /api/session-order` (`session:orderChanged` SSE, persisted in `state.json`; pure `normalizeSessionOrder`/`mergeSessionOrder` in `src/session-order.ts`: pushing device wins, server-only ids fall to the end, never dropped); resume from the manager keeps the original session name (COD-143); `firstPrompt` is backfilled for sessions whose id != transcript UUID and the most recent prompt (`lastPrompt`) is shown + searched (COD-140/145). ### Session lineage lines (tab → tab it spawned) **The relationship did not exist before this** (1.17.0): `SessionState` had no `parentSessionId`, `quick-start` recorded only the multi-user *human* owner, and an agent's spawn call is plain `curl` from a tmux pane, so nothing in the request identifies the caller (`SO_PEERCRED` needs a unix socket; the API is TCP). The caller therefore supplies it — every managed pane already gets `CODEMAN_SESSION_ID` from `session-cli-builder.ts`. Two equivalent inputs, body wins: a `parentSessionId` field on `POST /api/sessions` / `POST /api/quick-start`, or the `X-Codeman-Parent-Session` header, which exists so the agent skill can set it ONCE on its shared curl invocation and have every present and future spawn recipe carry it. **Resolved, not trusted** (`resolveParentSessionId()`, route-helpers.ts): exact id first, then a UNIQUE prefix of ≥8 chars (ids reach agents truncated — mux names and a Docker export's `$CODEMAN_SESSION_ID` both carry 8), and an ambiguous prefix resolves to NOTHING rather than to a guess. The parent must be a live session the caller can already see (`canAccessOwned`) AND carry the same owner as the session being created, so a multi-user caller cannot staple their session under someone else's tab. ⚠️ **Everything unresolvable is DROPPED, never a 400**: a stale id from a cached skill preamble must cost a decorative line, not a worker. ⚠️ It is decoration at every layer — never an ownership, permission or lifecycle signal; a child outlives its parent, and the Session ctor refuses a self-parent (reachable only via recovery, where both values come off disk). It rides `toState()` into `session_created` / `session_updated`, so there is **no new SSE event**, and `server.ts`'s recovery path restores it so lineage survives a restart. **Rendering is an additional LAYER, not a second pass** (`session-lineage.js`, loadorder 15.6): `_updateConnectionLinesImmediate()` (subagent-windows.js) calls `_appendLineageConnectionLines(svg, rects)` at its tail, exactly like ultracode's two layers, so all of them share ONE batched read→write reflow and the same `tab:` rect cache. Geometry is pure and unit-tested in `computeLineagePath()` (constants.js): both endpoints live in one horizontal strip, so the subagent shape (tab-bottom → window-top) has nothing to aim at, and every pair gets a U-bridge HANGING BELOW the strip, anchored on both tabs' BOTTOM edges (dip scales with distance, plus a per-sibling step so several children of one parent nest instead of overprinting, plus the row offset when the strip has wrapped). ⚠️ **A wrapped strip used to get its own shape, and that shape was the bug** (fixed 2026-08-14): `tabs-two-rows`/`tabs-auto-wrap` put a parent on row 1 ~14px above its child on row 2, so the old parent-bottom → child-TOP bezier had 14px to bend in and drew a flat line inside the row gap, siblings overprinting. Hanging the control points below the LOWER row gives the wrapped case the same bracket as the flat one and deletes the branch. The same pass raised the dip clamp (44 → 104, 0.06 → 0.085/px) because a skill worker is appended to the END of the strip, where the old cap flattened an 800-1500px span into a straight thread across the terminal, and traded weight for a second, wider glow (2 → 2.5px, `4 4` → `5 5` dashes at `-20`, opacity .55 → .72 / .95 working) because the original styling vanished into terminal text at 1:1. ⚠️ **Desktop only, for a z-index reason**: the overlay is `z-index: 999` and the desktop header is 100, so arcs paint OVER it — which is exactly what lets them touch tab bottoms. Under 1024px mobile.css makes the header `position: fixed; z-index: 1200` and would bury them, and the phone strip is a scroller where both endpoints are rarely on screen at once. Raising the SVG to ~1250 (above the fixed header, below modals at 1300) is the phase-2 option, and needs a real check against the mobile overview and the drawer. ⚠️ **`data-agent-id="lineage:"` is load-bearing**, not a label: `_applyLineEntrances()` queries paths by that attribute, so tagging them this way is the whole reason the arcs get the draw-in animation AND its negative-`animation-delay` resume across `svg.innerHTML = ''` with zero new animation code. ⚠️ `.session-tabs` is `overflow-x: auto`, so a tab scrolled out of the strip still HAS a rect — one lying over the logo or the header buttons; edges with an endpoint outside the strip are SKIPPED (clamping would point at a tab that is not there), and a passive `scroll` listener re-anchors the rest, since a scroll moves both endpoints without firing any render. The incremental tab render also redraws when `_lineageEdgeCount > 0`: a badge appearing widens a tab and shifts every tab after it. Setting: `sessionLineageLines`, per-device (in `displayKeys`, absent from the `.strict()` `SettingsUpdateSchema`), desktop default ON. Tests: `test/session-lineage-lines.test.ts` (geometry), `test/routes/session-routes-parent-lineage.test.ts` (resolution + reject paths). ### Full-scrollback replay **Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell full history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, while the marker callback supplies accurate parse timing without extending the pre-existing queued-event discard window. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[;H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. ### Terminal scrollback: strip flavors and wheel/touch forwarding **Two strip flavors, one carry** (#205, `session.ts:_handleTerminalOutput`): the FULL strip (`isAltScreenStripMode` = codex/claude/gemini) removes alt-screen toggles, `3J`, and mouse-tracking DECSETs. Every other mode (shell/opencode/antigravity/pi) gets the NARROW strip (`isMuxAltScreenOnlyStripMode`) — alt-screen toggles ONLY — and only when tmux-backed (`useMux`). Rationale: the tmux CLIENT emits `smcup` as its first bytes at attach, before any program runs, parking xterm in the scrollback-less alternate buffer for the whole session (touch scrolling no-ops; xterm's own wheel handler converts the wheel to Up/Down arrows = readline history cycling — both #205 symptoms). tmux never forwards a pane program's alt-screen toggles to its client (it repaints instead; measured — vim/less inside a pane emit zero to the client), so the only thing the narrow strip ever removes is tmux's own smcup. It keeps `3J` (a user's `clear` is a deliberate scrollback wipe) and the mouse DECSETs (tmux passes those through even with `mouse off`; stripping them would break htop/vim mouse support). ⚠️ The `useMux` gate is load-bearing: `startShell()`/`startInteractive()` fall back to a DIRECT PTY when mux creation fails, and there the inner program's own `?1049h` really does reach xterm — stripping it would break vim/less/htop for real. The replay path (`session-routes.ts`, via `session.usesMux`) applies the same narrow branch; the frontend mirror (`_shouldReportMouseToCli()`) stays claude/codex/gemini because only the FULL strip touches mouse DECSETs. The chunk-boundary carry (`_altScreenSeqCarry`) runs for both flavors. Tests: `test/claude-scrollback-strip.test.ts`. ⚠️ **What the full strip removes, it must REMEMBER.** Stripping the mouse DECSETs means xterm's `modes.mouseTrackingMode` is permanently `'none'` for those modes, so the browser hand-encodes click reports to compensate (`_sendSyntheticSgrTap`). With no state to consult it did that on EVERY click, which delivered mouse reports to programs that never asked for them: the same pane runs a plain shell whenever the CLI has exited or a `shell` was started inside a claude-mode session, and a shell prints the report as literal text (`[<0;88;20M`), garbling the next line typed. `_recordStrippedMouseMode()` therefore records each stripped sequence as it goes and publishes `cliMouseTracking` through `toState()`, and `_shouldReportMouseToCli()` requires it. ⚠️ Only the TRACKING modes count (1000/1001/1002/1003): 1005/1006 select an ENCODING and 1007 is alt-scroll, and counting those would put the stray reports straight back. ⚠️ The change broadcasts IMMEDIATELY rather than through `broadcastSessionStateDebounced`, because the flag flips when a dialog opens and the user can click that dialog inside the 500ms debounce window. Measured on a live claude 2.x: the CLI holds a tracking mode on continuously (so clicks keep being reported exactly as before), while a bash prompt in the same stripped mode reports nothing. Fails toward silence: after a server restart the flag is false until the CLI re-emits, which tmux does at client attach. **Only claude ≥ 2.1.187 forwards the wheel; every other mode scrolls local scrollback** (#227 follow-up, `terminal-ui.js:_shouldForwardWheelToApp`). Codex was in the forward list until a reporter hit a completely dead wheel in codex tabs while the scrollbar drag worked. Measured against codex-cli 0.147.0 in a bare tmux: it never enables mouse tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change nothing on screen, because it runs an INLINE viewport (`alternate_on=0`) and pushes its transcript into the terminal's own scrollback (tmux `history_size` grows) instead of paging in-app. So for codex, local scrollback IS the transcript and forwarding swallowed every tick. ⚠️ "The TUI is a strip mode" is NOT evidence that it consumes wheel reports — verify with a real `\x1b[<64;c;rM` write into a live pane before adding a mode here. Hand-encoded SGR TAPS are gated by `_shouldReportMouseToCli()` (strip mode AND the server-observed `cliMouseTracking` flag, recorded by `_recordStrippedMouseMode` in session.ts as it strips): codex never enables mouse tracking, so since #325 no tap report is sent there at all — click-to-position was already a measured no-op in codex, and a pane that has fallen back to a shell no longer receives `[<0;88;20M` junk. **Wheel/touch forwarding is NOT gated on viewport-at-bottom** (#205, `terminal-ui.js:_shouldForwardWheelToApp`): for sessions verified to scroll their own transcript on SGR wheel reports (claude ≥ 2.1.187 — version via the local/docker/remote `--version` probes), the plain wheel AND touch drags forward as coalesced SGR reports (`_forwardScrollToApp` → `_sendSyntheticSgrWheel`, 40ms batches, 5-tick cap, 512-byte queue bound). It used to gate on the viewport being at the bottom so both scrollbacks stayed reachable, but a repaint-mode CLI keeps NO terminal scrollback of its own — xterm's buffer holds only replayed repaint frames, so local scrolling drags the CLI's pinned prompt box up the screen over stale frames; and `scrollToLastNonEmptyLine()` routinely parked the viewport off-bottom, silently pinning the wheel to local. Forwarding now snaps the viewport home first (SGR coordinates address the LIVE screen — a report computed from a scrolled-up viewport would hit-test the wrong row). Local scrollback remains on Shift+wheel and the `terminalWheelLocalScrollback` opt-out (both also cover touch via the shared gate; touch has no Shift, so the setting is its only local pin). `_wheelScrollLines()` normalizes `deltaMode` (Firefox fires LINE deltas ≈3/notch — read as pixels that rounded to 0 and fell to the ±1 fallback, ~4× too slow; PAGE deltas scale by `terminal.rows`) while keeping the #154 Shift-axis trap (macOS trackpads put Shift+scroll magnitude on deltaX). Tests: `test/terminal-touch-tap.test.ts`. **A false gate on a Claude session must not mean a DEAD gesture** (#205 round 2, `_maybePageCliTranscript`): every way `_shouldForwardWheelToApp()` returns false leaves a repaint-mode pane scrolling a buffer that has nothing in it (`baseY === 0`) — the version probe came back empty, the CLI really is older than 2.1.187, or the user turned on `terminalWheelLocalScrollback`. The 1.12.0 retest reported exactly that: a wheel that did nothing at all while Fn+Up (PageUp) paged back through intact text, which is the proof that the CLI's own history and the PTY input path were both fine. So under the triple guard (claude mode + gate false + `baseY === 0`) wheel and touch travel is translated into coalesced `\x1b[5~` / `\x1b[6~` through the same 40ms queue as the SGR reports, at half a screen of travel per page key (the key jumps a whole screen; a 1:1 mapping was unusably slow with a discrete wheel). ⚠️ Shift is excluded on purpose — it is the explicit "give me local scrollback" gesture and must keep that meaning. ⚠️ `terminalWheelLocalScrollback` is deliberately NOT scoped away from repaint-mode CLIs even though it is a footgun there: that would silently override an explicit user choice, so the fallback catches it instead. **Server-side counterpart**: `getClaudeCliVersion()` caches SUCCESS for the process lifetime but must never cache FAILURE — it used to, so one timed-out or PATH-starved probe at the first Claude session start disabled wheel-forwarding for every Claude session until the server restarted (a dead wheel on phone, tablet and laptop at once, the signature of a server-side cause). Failures now retry with a 1/2/4…15min backoff; the policy is the pure `resolveClaudeCliVersion()`. Tests: `test/terminal-scroll-routing.test.ts`, `test/claude-cli-version-cache.test.ts`. **Why the wheel went where it went is LOGGED** (`_logScrollRouting`): one console line per session per distinct decision — `[scroll] → forward-sgr|page-keys|local-scrollback|repull-refused-downgrade (mode=…, cliVersion=…, localScrollbackOptOut=…, mouseTracking=…, localScrollbackRows=…)`. #205 ran two rounds of remote guesswork over questions this line answers directly; keep it when touching the routing. **The wheel listener is CAPTURE-phase and Codeman owns the scroll** (#205 follow-up, measured on the live instance): xterm's viewport is a vscode-style ScrollableElement that consumes wheel events itself (preventDefault + stopPropagation) whenever it believes a scrollbar exists, ignores `attachCustomWheelEventHandler`, and goes DEAF after `terminal.reset()` — a tab switch or full-history replay leaves its scroll dimensions stale, after which wheel events neither scroll nor propagate reliably. A bubble-phase container listener therefore never fired once local scrollback existed (forwarding, deltaMode and the top-of-buffer re-pull all silently dead exactly on sessions WITH history), and after a tab switch nothing scrolled at all ("works at first, breaks after a tab switch"). The container wheel listener is `{capture: true}`, stops propagation, and scrolls locally via buffer-level `terminal.scrollLines()` (immune to the stale scroller). ⚠️ Two cases are deliberately passed through untouched, in this order BEFORE preventDefault: `mouseTrackingMode !== 'none'` (xterm's encoder forwards the wheel to the PTY — htop/vim with mouse on) and `buffer.active.type === 'alternate'` (direct-PTY vim/less: xterm's alt-scroll converts the wheel to cursor keys). Do not "simplify" this back to a bubble listener or re-delegate local scrolling to xterm's viewport. E2E guard: the reload → tab-switch → wheel matrix in the #205 verification scripts. ### Run launch synchronization **Run launch synchronization**: the main Run entrypoint in `session-ui.js` holds an in-flight lock and disables `#runBtn` for the whole launch (at least 500ms), so a double click cannot create duplicate sessions with the same `w-` name. A successful create/quick-start also calls `_ensureCreatedSessionVisible()` before `selectSession()`: local creates use the response's full session snapshot; quick-start modes fetch `GET /api/sessions/:id` only when `session:created` SSE has not already populated the map. The normal `_onSessionCreated()` handler remains the idempotent upsert, so POST-first and SSE-first ordering both produce one immediately-rendered tab. Tests: `test/run-mode-ui.test.ts`. ### Circuit breakers: Ralph and PTY-exit **Circuit breaker**: Prevents respawn thrashing. States: `CLOSED` → `HALF_OPEN` → `OPEN`. Reset: `/api/sessions/:id/ralph-circuit-breaker/reset`. **Distinct: PTY-exit breaker** (COD-115/118/#147, `session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits (crash loops on attach), blocks further auto-restarts, broadcasts SSE `session:respawnBreakerTripped` + push (in `PUSH_EVENT_MAP`). Reset ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive` (sent by the user-facing restart control) — the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. Sessions also scrub inherited `TMUX`/`TMUX_PANE` env so Codeman-in-tmux doesn't nest. Tests: `test/respawn-pty-breaker.test.ts`. ## Features ### Attachments **Attachments** (live external document references; COD-37/#119 core, COD-38/#120 previews, COD-39/#121 history): all wiring in `file-routes.ts`. **Registry** (`attachment-registry.ts`): an **in-memory** map of a stable `attachmentId` → an absolute, `realpath`-resolved, extension-allowlisted file path, so browser requests (`GET /api/sessions/:id/attachments/:attachmentId/raw`) never carry arbitrary absolute paths; `POST /api/sessions/:id/attachments` registers one. **Magic links** (`attachment-magic.ts`): parses `codeman://attach?...` out of terminal output — ⚠️ this scanner is prompt-injectable, so the scan path is **force-confined to the session workspace** (a hostile prompt could otherwise make it read arbitrary host files over SSE); emits the `attachment:detected` SSE event. Security gate is an extension **allowlist** (`isSupportedAttachmentExtension`, in the registry/magic modules), not a blocklist; a separate path layer (`config/attachment-guard.ts`) confines reads to the workspace (`attachmentConfineToWorkspace`) and blocks sensitive trees (`/root`, `/etc`). **Previews + thumbnails** (COD-38): `:attachmentId/preview` + `:attachmentId/thumbnail` (and the workspace-file equivalents `file-preview`/`file-thumbnail`) render Office docs/PDFs via external converters (`pdftoppm` / LibreOffice `soffice` / Word-COM `powershell`); `document-preview-cache.ts` is a shared disk cache (de-dups _identical_ in-flight inputs), `document-thumbnailer.ts` does best-effort first-page images, and `document-conversion-limiter.ts` is a **global converter-spawn concurrency cap** (`runWithConversionLimit`) — without it, N distinct large docs detected at once fork N multi-minute converter processes = a localhost fork-bomb-shaped resource-exhaustion vector. **History drawer** (COD-39): `session-attachment-history.ts` tracks the last `ATTACHMENT_HISTORY_LIMIT` (100) attachments per session (`Session._attachmentHistory`, persisted via `SessionState.attachmentHistory`, replayed so externals re-register on reconnect); `GET /api/sessions/:id/attachments` is the list endpoint. ⚠️ The history drawer's launcher button is desktop-only — hidden on phones (regression-guarded; see `mobile-header-buttons-policy` test). Session-local files keep using the existing workspace-scoped `file-routes` paths; the registry is only for explicit live externals. **Codex generated artifacts** (COD-166/#150, `generated-artifact-attachments.ts`): codex-mode sessions ALSO scan (ANSI-stripped) output for `Saved to: file:///…` lines and surface those files as attachment cards with a relaxed trust policy — the allow decision runs on the **realpath-resolved** path against `os.homedir()`-anchored `~/.codex` marker dirs (symlink escapes fall back to force-confinement); gated to `mode === 'codex'` only (`source` is a REQUIRED param through the listener-deps chain — a dropped arg here silently kills the feature). Image thumbnails pass through jpg/jpeg/gif/webp. ### File-path links (terminal + response viewer) A file path an agent prints is a link on both surfaces it can appear on, and clicking it opens the file-preview overlay. Three things make that work and each has bitten: **One pattern, two consumers.** `FILE_PATH_LINK_PATTERN` / `absoluteFilePathPattern()` live in `constants.js`; the xterm link provider (`registerFilePathLinkProvider`, terminal-ui.js) and the response viewer's `_linkifyFilePaths()` (app.js) both build a fresh instance from it. ⚠️ Fresh per call, never one shared object: `lastIndex` is per-object state on a `/g` regex. The pattern is anchored on a known absolute root and terminated by a known extension, so a fraction (`3/4`) or a date can't match and trailing punctuation stays out. Roots include `Users` and `mnt`, without which nothing was clickable on macOS or WSL. The linear-time guard and the "terminal-ui builds from the factory" structural check are in `test/link-provider-regex.test.ts`. **The chat linkifier walks text nodes.** `_linkifyFilePaths()` builds anchors with `createElement`/`textContent` on the rendered subtree, never by rebuilding sanitized markup as a string — the source is model output. Subtrees already inside an `` are skipped (marked autolinks URLs; a nested anchor would swallow the click), and the anchor's text is the path verbatim so "copy code" still yields what the agent printed. `test/response-viewer-file-links.test.ts` pins both properties. **Out-of-workspace paths go through the attachment routes, not the file routes.** `file-content`/`file-raw` resolve against `workingDir` and 404 anything that escapes it, which is correct and unchanged — but the paths agents most often print (a `/tmp` capture, Claude's own scratchpad, another checkout) are exactly that, so clicking one used to report "File not found" for a file sitting on disk. `openFilePreview()` now detects the case (`_isExternalPreviewPath`, a string compare for ROUTING only; the real decision stays server-side) and registers the path via `POST /api/sessions/:id/attachments` first, rendering by id. ⚠️ That registration passes `notify: false`, which suppresses ONLY the `attachment:detected` broadcast — the guard, the registry entry and the by-id routes are identical either way. Without it every click also popped an attachment card announcing the file already filling the screen. ⚠️ The click is an explicit user action on the **explicit, Origin-guarded** registration route, which is why it may cross the workspace boundary at all; the passive magic-link scanner stays force-confined. A type outside `SUPPORTED_ATTACHMENT_EXTENSIONS` (`.svg`, `.bmp`) is refused with a message naming what IS previewable, rather than the registry's own policy term. ⚠️ **The terminal routes an out-of-workspace path to the preview, not the log viewer.** The log viewer spawns `tail -f` and allows only the workspace, `/var/log` and `~/logs`, so an external `.log`/`.json`/code path answered `Path must be within working directory or allowed log directories` while the SAME path clicked in the response viewer previewed fine. `activate()` now checks `_isExternalPreviewPath` alongside `previewsInFileViewer`. In-workspace text keeps the tail viewer, which is the point of it (live follow); nothing widened `file-stream-manager`'s allowlist, so no `tail -f` is spawned on an arbitrary host path. **Text reuses the edit-mode allowlist; markup stays download-only.** `TEXT_ATTACHMENT_EXTENSIONS` IS `EDITABLE_EXTENSIONS` (`config/file-editing.ts`) rather than a second curated list that would drift from it: if the viewer would open a file for editing inside the workspace, the same file outside it can be read. The justification for widening is that the agent in the session can already `cat` any of these and the picker already previews them, so the suffix was never the confidentiality gate; the path guard is (sensitive-file blocklist, `/root` and `/etc` trees, realpath first). ⚠️ Two consequences had to be handled at the same time: `~/.codeman*/state.json` joined `isSensitivePath` (it persists `SessionState.envOverrides`, and the env allowlist admits key-shaped names like `GEMINI_API_KEY`, so it can hold a live credential), and `html`/`htm` joined `svg` in `serveRawFile`'s **download-only** branch so that widening what can be READ never widens what can RUN on our own origin. Text with no dedicated MIME entry goes out as inert `text/plain; charset=utf-8` + `nosniff`, matching the picker. The by-id text preview is bounded like the workspace one: a `Range` request for the first 512KB (a real partial read, not a discarded 50MB download) plus a 500-line cap, with the footer saying so. **Media is single-sourced across the two preview paths.** `VIDEO_ATTACHMENT_EXTENSIONS` / `AUDIO_ATTACHMENT_EXTENSIONS` live in `attachment-registry.ts` and are imported by `file-content`'s media classification, so a clip plays identically whether it is in the workspace or reached by id from outside it. They diverged first: the workspace path had its own inline sets and the registry allowlist had no media at all, so a video an agent wrote to `/tmp` was refused as an unsupported type while the same file inside the repo played. ⚠️ Three things have to line up for a player rather than a dead frame: the extension in the allowlist, a **real MIME entry** in `MIME_TYPES` (a `