diff --git a/CLAUDE.md b/CLAUDE.md index 30a4fcbd..16af0804 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -170,15 +170,15 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph | **Search** | `src/search-service.ts` | Pure in-memory core for `GET /api/search` | | **Attachments** | `src/attachment-registry.ts`, `attachment-magic`, `generated-artifact-attachments`, `session-attachment-history`, `document-preview-cache`, `document-thumbnailer`, `document-conversion-limiter`, `config/attachment-guard` | See Key Patterns | | **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/` (`claude-md.ts` + `case-template.md`) | `templates/` holds the CLAUDE.md scaffold generated into new cases | -| **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (27 modules + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | | -| **Frontend** | `src/web/public/app.js` (~6.9K lines, core) + 35 modules + `sw.js` (+ `voice-pcm-worklet.js`, fetched from JS, not in the load order) | See Frontend section for the load order, which is authoritative | -| **Types** | `src/types/index.ts` (barrel) → 22 domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts | +| **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (one module per domain + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | | +| **Frontend** | `src/web/public/app.js` (core) + the modules listed in the Frontend load order + `sw.js` (+ `voice-pcm-worklet.js`, fetched from JS, not in the load order) | See Frontend section for the load order, which is authoritative | +| **Types** | `src/types/index.ts` (barrel) → domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts | ★ = Large, central file (>50KB) — read its `@fileoverview` first. All files have `@fileoverview` JSDoc — read that before diving in. Discovery aid: `grep -l '@fileoverview' src/web/routes/*.ts` lists all route modules; same grep works for `src/types/`, `src/web/public/*.js`. **Local packages**: `packages/xterm-zerolag-input/` (local echo overlay, single-source, see Gotchas). `packages/gesture-control/` (`codeman-gesture-control`, hand-tracking overlay source, built via `npm run build:gesture`). -**Config**: `src/config/` — 23 files plus the `cli-registry/` subdir, no barrel (`index.ts`) exists; import from the specific file. ⚠️ There are TWO `config/` directories: the repo-root `config/` holds tooling only (ESLint, knip, the vitest configs, `test-suites.ts`), while runtime config lives in `src/config/`. Throughout this file a bare `config/.ts` in a code context means `src/config/.ts`. +**Config**: `src/config/` — flat files plus the `cli-registry/` subdir, no barrel (`index.ts`) exists; import from the specific file. ⚠️ There are TWO `config/` directories: the repo-root `config/` holds tooling only (ESLint, knip, the vitest configs, `test-suites.ts`), while runtime config lives in `src/config/`. Throughout this file a bare `config/.ts` in a code context means `src/config/.ts`. **Utilities**: `src/utils/` — re-exported via index. Key: `CleanupManager`, `LRUMap` (⚠ NOT in the barrel — import from `./utils/lru-map.js` directly), `StaleExpirationMap`, `BufferAccumulator`, `stripAnsi`, `Debouncer`, `KeyedDebouncer`. Also: `claude-cli-resolver`/`opencode-cli-resolver`/`codex-cli-resolver`/`gemini-cli-resolver`/`antigravity-cli-resolver`/`pi-cli-resolver`/`grok-cli-resolver`/`deepseek-cli-resolver`/`omp-cli-resolver` (CLI path resolution, one per `SessionMode`, all nine sharing the lookup chain in `cli-executable-resolver`: server PATH, then that CLI's install dirs, then an interactive login shell LAST, since it is the only step that spawns anything and it is what finds nvm/Homebrew installs under a service manager's minimal PATH; ⚠ `pi-`, `grok-` and `deepseek-cli-resolver` additionally probe the binary's identity, since `pi` is a generic name, `grok` has npm squatters, and Debian ships an unrelated `dsh`), `file-query` (⚠ Files-panel search matcher, glob-by-two-pointer, never RegExp), `string-similarity` (fuzzy matching), `regex-patterns` (ANSI/token/spinner patterns), `assertNever` (exhaustive checks), `token-validation` (auth tokens), `nice-wrapper` (process priority), `shell-resolver` (⚠ resolves a real login shell for `mode: 'shell'`; the literal string `$SHELL` used to be expanded by the SERVER's shell, which is empty in a container), `event-loop-monitor` (a sync `execSync` freezes the port while the process stays alive, leaving no trace), `dependency-checker` + `dependency-report` (the `codeman doctor` probe engine, registry in `config/dependency-registry.ts`). @@ -193,110 +193,114 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Input**: `session.writeViaMux()` for programmatic/curl input via tmux `send-keys -l` + `send-keys Enter`, single-line only. Interactive **browser** input goes through a durable **exactly-once** layer: a stable `clientId` + monotonic per-session `seq` persisted to localStorage until the server ACKs, so a dropped link cannot lose or double-deliver a prompt. `ws-connection-registry.ts` supersedes only same-TAB reconnects, so two tabs on one session coexist. → [architecture-invariants#input-delivery-and-ws-resilience](docs/architecture-invariants.md#input-delivery-and-ws-resilience) -**Agent wait primitives**: bounded long-polls so an agent driving Codeman from a shell can block instead of poll: `GET /api/sessions/:id/wait` (lifecycle signal), `GET /api/sessions/:id/wait-output` (literal substring, **never** regex) and `wait`/`waitTimeout` on `POST /api/sessions/:id/input`. Registry in `session-wait-registry.ts` (pure, no `Session` reference), bounds in `config/agent-wait.ts`. ⚠️ **A timeout is a 200** (`wait.timedOut`), never an error, so callers loop over short waits. ⚠️ `stop`/`blocked` are hook-driven and fire for **`claude` and `deepseek` ONLY** (`shell` installs none either); asking for one explicitly on any other mode is a 400, the default set silently drops them. `deepseek` qualifies because the DeepSeek Harness TUI REPORTS idle/working/blocked to its supervisor and Codeman is that supervisor (`deepseek-status-shim.ts`), so its signals are definitive rather than inferred — `hooksAvailableForMode()` in `session-wait-registry.ts` is the one place that rule lives. ⚠️ Send-and-wait registers the waiter BEFORE the write (a separate POST-then-wait races and reports the PREVIOUS turn), and both teardown paths must `notifySignal('exit')` BEFORE `cancelAll()`. ⚠️ Client-hangup abort listens on **`reply.raw`** guarded by `writableFinished`: on `req.raw`, `close` fires when the request BODY ends, which on a POST killed every send-and-wait instantly and no `app.inject()` test could see it. ⚠️ Worker liveness cannot come from `session.pid` — for a tmux session that is the local attach client, which outlives a worker dying inside its pane — so it is probed at the mux layer (`isPaneDead`, ~750 ms cache) on blocking waits only, never on the input hot path. ⚠️ Signals are edge-triggered with no history: one that fires with no waiter registered is unobservable afterwards, so gather fan-outs with send-and-wait or latched `wait-output` markers, never fire-and-forget-then-sequential-signal-waits. ⚠️ **`deepseek` is therefore the one non-claude mode the skill drives like claude** — `spawn_workers alpha beta:deepseek` is a mixed fleet in one call, and `sendwait`/`last_text` need no variant. Two traps are baked into the preamble rather than left to the agent: the harness's boot `idle` report lands ~300 ms BEFORE its composer paints (2.26 s vs 2.56 s, measured), so readiness must come from the composer and never from the signal, or a send-and-wait resolves on the boot edge and reports a turn that never ran; and `sendwait` asks for `wait:"stop,exit"` rather than the default set, because that set also carries `idle`, which for an external CLI is inferred from output stabilization — on a dsh worker whose TUI repaints rarely, a re-wait resolved in 0 ms with `signal:"idle"` on a turn with minutes left to run. The primitives are packaged as the **`skills/codeman` agent skill**: installable via `codeman skill install [--case ]` / `skill uninstall`, as a Claude Code plugin from the repo's own marketplace (`/plugin marketplace add Ark0N/Codeman` + `/plugin install codeman@codeman`, see the root-files note above; ⚠️ plugin skills are namespaced, so a Claude Code holding BOTH the plugin and a user-level or per-case copy lists the skill twice, `codeman` and `codeman:codeman`, measured 2026-09-14: neither shadows the other, both work, the docs tell users to pick one route), or auto-injected into a case's `.claude/skills/` on Claude session create behind `agentSkillEnabled` (SYNCED, default OFF). Injection is ADD-ONLY at create, marker-owned (`applyAgentSkill` in `hooks-config.ts` never touches an unmarked user copy) and refuses symlinks (this repo's own `.claude/skills/codeman` is a symlink to the source, which the injector must never write through). ⚠️ Claude Code loads a same-named USER-LEVEL skill (`~/.claude/skills/codeman`, written once by `codeman skill install` with no `--case`) over the per-case copy, and nothing used to refresh it: a stale Aug-9 user copy shadowed every fresh injection (2026-08-14: agents ran the old recipes, spawned workers serially and lost their lineage arcs), so session create now also refreshes a marker-owned user copy (`refreshUserAgentSkill`; refresh-only, never installs, foreign/symlink refused). Session create additionally pre-seeds the skill's §0 preamble cache (`seedAgentSessionPreamble` → `${XDG_CACHE_HOME:-~/.cache}/codeman-agent-.sh`, local claude sessions only), single-sourced from `skills/codeman/preamble.sh` and pinned byte-identical to SKILL.md's §0 heredoc by `test/agent-skill.test.ts`, so the skill's bootstrap is a two-line loader instead of a ~150-line paste the model types out (~47 s of generation, measured live). → [architecture-invariants#agent-wait-primitives](docs/architecture-invariants.md#agent-wait-primitives), `docs/api-reference.md` +**Agent wait primitives**: bounded long-polls: `GET /api/sessions/:id/wait`, `GET .../wait-output` (literal substring, **never** regex) and `wait`/`waitTimeout` on `POST .../input`. Registry `session-wait-registry.ts` (pure), bounds `config/agent-wait.ts`. ⚠️ A timeout is a 200 (`wait.timedOut`). ⚠️ `stop`/`blocked` exist for `claude` and `deepseek` ONLY (rule lives in `hooksAvailableForMode()`): explicit request elsewhere is a 400. ⚠️ Send-and-wait registers the waiter BEFORE the write; teardown must `notifySignal('exit')` BEFORE `cancelAll()`; hangup abort listens on `reply.raw` (guarded by `writableFinished`), never `req.raw`; liveness comes from `isPaneDead`, never `session.pid`. ⚠️ Signals are edge-triggered with no history: gather fan-outs via send-and-wait or `wait-output` markers. Packaged as the `skills/codeman` skill (`codeman skill install`, plugin marketplace, or injection behind `agentSkillEnabled`, SYNCED, default OFF): injection is add-only, marker-owned (`applyAgentSkill`), refuses symlinks, and refreshes a marker-owned user-level copy (`refreshUserAgentSkill`). → [architecture-invariants#agent-wait-primitives](docs/architecture-invariants.md#agent-wait-primitives), `docs/api-reference.md` -**Agent-created case marker** (`src/agent-case-marker.ts`): a case directory `POST /api/quick-start` **creates** for an agent-driven spawn gets a `.codeman-agent-case.json` marker, so the scratch workspaces a long orchestration leaves behind (one per worker, and deleting the session does not remove them) can still be told apart from the user's real projects months later. `GET /api/cases` publishes it as `agentCreated`; `GET /api/cases/agent-created` is the read-only cleanup listing, adding `inUse` (a live session's `workingDir` is that case) and `modifiedAt`; Add Case → Manage badges each one and offers a review-then-delete sweep. The signal is the skill preamble's `X-Codeman-Agent-Origin` header (or an `agentOrigin` body field), falling back to a RESOLVED `parentSessionId` — nothing in the browser UI sets lineage, so a create request naming its spawning session came from an agent by construction, and that fallback is what still labels workers spawned by a stale skill copy. ⚠️ **Only the branch that CREATES the directory may write it.** A linked case, a cloned repo or any pre-existing path must never be labelled: the label drives a recursive-delete affordance, and mislabelling someone's repo there is the one failure mode that costs real work. `POST /api/sessions` takes an existing `workingDir`, so it writes no marker at all, by construction. ⚠️ Reading is TOTAL: anything that is not a well-formed version-1 marker (truncated write, hand-edited junk) reads as *not* agent-created rather than as a half-trusted entry, and deleting the file is the supported way to adopt a scratch case as a real one — which is what the `note` written into it tells whoever finds it. ⚠️ Removal stays on the existing `DELETE /api/cases/:name`, one name at a time, so there is exactly ONE recursive-delete path; the UI's sweep names every directory in its confirm and EXCLUDES an `inUse` case outright rather than confirming it away. ⚠️ Marker in the case dir rather than a registry under `~/.codeman`: it survives a wiped data dir or a different instance, is removed by the same `rm -rf` that removes the case (so no stale-entry pruning), and a user who runs `ls -a` can see what labelled their directory. Adding the header changed the preamble, so `CODEMAN_PREAMBLE` was bumped (1.22.0) — a cached copy is version-checked, and forgetting the bump leaves every already-seeded agent sending the old headers. Tests: `test/agent-case-marker.test.ts`, `test/routes/agent-case-marker-routes.test.ts`. +**Agent-created case marker** (`src/agent-case-marker.ts`): a case dir that `POST /api/quick-start` CREATES for an agent-driven spawn (signal: the preamble's `X-Codeman-Agent-Origin` header / `agentOrigin` field, else a resolved `parentSessionId`) gets `.codeman-agent-case.json`, published as `agentCreated` on `GET /api/cases`; `GET /api/cases/agent-created` is the cleanup listing (`inUse`, `modifiedAt`) behind Add Case → Manage. ⚠️ Only the branch that CREATES the directory may write it: never label a linked case, cloned repo or pre-existing path (it drives a recursive delete). ⚠️ Reading is total: anything but a well-formed v1 marker reads as not agent-created. ⚠️ Removal stays on `DELETE /api/cases/:name` (the ONE recursive-delete path), and the sweep excludes `inUse` cases. ⚠️ Changing the preamble's headers requires bumping `CODEMAN_PREAMBLE`. → [architecture-invariants#agent-created-case-marker](docs/architecture-invariants.md#agent-created-case-marker) **Agent preamble cache GC**: the §0 preamble seeded per claude session (`$XDG_CACHE_HOME/codeman-agent-.sh`) is now REMOVED with the session (`removeAgentSessionPreamble` from `_doCleanupSession`, `killMux` only — a detach leaves the session recoverable and its agent would come back to a loader whose file we deleted) and swept at boot (`pruneAgentSessionPreambles(this.sessions.keys())`, once, after restore, so every session this instance owns is in the keep set). Nothing removed them before: 236 leftovers measured on a working machine, the oldest three weeks old. ⚠️ The sweep needs BOTH guards — never a live session's file at any age (the two-line loader reads it mid-run), and `AGENT_PREAMBLE_MAX_AGE_MS` (7d) of age on top, which is what keeps ANOTHER instance's sessions (whose ids this process cannot see) out of the blast radius. Losing one is degradation, not breakage: the §0 fallback block rewrites it. Tests live with the seed's in `test/agent-skill.test.ts`. **Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`. -⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer (`❯`) about once a second all through a turn, so the old "saw a ❯, wait 2s → idle" rule flipped every working session to idle two seconds in (measured: a session mid-tool-call at 17 minutes reporting `status:"idle"`). Its working indicator is `✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`: the glyph animates through `· ✢ ✳ ∗ ✻ ✽`, the gerund is randomized, and the finished line (`✻ Cooked for 2m 49s`) carries the same glyph, so neither `SPINNER_PATTERN` (braille, not what current versions draw) nor a keyword list can see it. Matching the new line in the STREAM does not work either: tmux ships partial repaints, so the whole line reaches the PTY only every few tens of seconds. So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks the SCREEN via `capturePaneText()` + `CLAUDE_WORKING_LINE_PATTERN` before believing it; a sustained run of repaints (`session-activity.ts`, pure + unit tested) is what marks a turn as started, with the same screen probe vetoing keystroke echo. Idle now lands ~3-5s after a turn ends instead of 2s into one. ⚠️ **The composer glyph and the working line are per-CLI registry DATA** (`capabilities.workDetect`, #385), not Claude constants: claude declares `❯` plus the pattern above, codex declares `›` plus `[Ee]sc to interrupt`, and a CLI that declares neither falls back to Claude's pair, which is what every session used before the registry carried one. Before that, this whole mechanism was gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` for its entire life. ⚠️ `workingLine` is config-supplied (a user `clis.json` can set it) and the compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()`: a nested quantifier there is a ReDoS against the event loop, and the helper returns null rather than throwing so the fallback is structural. ⚠️ **A quiet pane is not always a pane that wants you.** An agent that arms a monitor, backgrounds a shell or hands work to a cloud session is told to end its turn, so Claude Code's `idle_prompt` notification lands a minute later on a session that wants nothing. The CLI states what it is still running on its own screen (optional `capabilities.workDetect.watchingLine`, with `watchingLines` for how far up the screen it sits), the idle probe reads it into `Session.watching`, and `notePrompt()` then opens that idle item ALREADY acknowledged so no surface alerts. ⚠️ **The fix is the alert that does not fire, not the `watching` badge**, and only `idle` is eligible — but a prose question is not a dialog, so a question asked in plain text while background work runs is silenced with the false alarms. ⚠️ **The label is pane-derived and therefore prompt-injectable**: the window must cover only rows the CLI draws and each pattern must anchor on chrome only that CLI can produce, or an agent silences its own alert by printing the words. → [architecture-invariants#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you](docs/architecture-invariants.md#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you). Tests: `test/session-watching.test.ts`, `test/watching-no-alert.test.ts`. +⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer all through a turn, and its working line (`✻ Actualizing… (13m 23s · …)`) is invisible to `SPINNER_PATTERN`, keyword lists and the raw stream. `_confirmIdle()` (session.ts) requires the pane to go quiet AND the SCREEN (`capturePaneText()` + the working-line pattern) to agree; a sustained run of repaints (`session-activity.ts`) marks a turn as started. ⚠️ The composer glyph and working line are per-CLI registry DATA (`capabilities.workDetect`), never Claude constants; a CLI declaring neither falls back to Claude's pair. ⚠️ `workingLine` is config-supplied and runs on the PTY hot path, so it must compile through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()` (ReDoS guard; null, not throw). → [architecture-invariants#idle-detection-composer-glyph-and-working-line](docs/architecture-invariants.md#idle-detection-composer-glyph-and-working-line) -**An exited agent in a live pane** (`paneExit`, Ark0N/Codeman#446): Codeman creates every pane with `remain-on-exit on`, so `/exit` ends the CLI while tmux keeps the pane, the tmux session and the `tmux attach-session` process Codeman records as `Session.pid`. No PTY exit handler fires, so a session whose agent is gone reads as a live idle one. `TmuxManager.startPaneExitWatcher()` reads `#{pane_dead}`/`#{pane_dead_status}`/`#{pane_dead_signal}` from one batched `list-panes -a` on its OWN always-on interval, and `SessionState.paneExit` rides the existing `session:updated` broadcast. ⚠️ **Never set `status: 'error'` for an exited pane** — that value belongs to the PTY-exit breaker and makes the browser offer a restart — and **never null the `pid`**, which is what makes `selectSession()` re-attach and launch a fresh CLI. ⚠️ **The field is TRI-STATE and its third state is absence**, meaning UNKNOWN, which must never render as alive; `Session.paneExitApplies` is the one place that scoping lives and it fails closed for a direct-PTY session, a remote SSH session, a docker case and a record rebuilt from the socket. ⚠️ **An absent `#{pane_dead_status}` is not 0** (measured on tmux 3.2a, a SIGKILLed pane reports neither a status nor a signal), so never write `status ?? 0`: absent-stays-absent is what will keep a later clean-exit sweep off crashed agents. ⚠️ **A path that starts a command in a pane must clear the record AND persist**, since the watcher's next tick sees the field already cleared and writes nothing. → [architecture-invariants#an-exited-agent-in-a-live-pane-paneexit](docs/architecture-invariants.md#an-exited-agent-in-a-live-pane-paneexit) +⚠️ **A quiet pane is not always a pane that wants you.** A CLI can declare an optional `capabilities.workDetect.watchingLine` (a monitor, background shell or cloud hand-off it is still running); the idle probe reads it into `Session.watching` and `notePrompt()` opens that idle item ALREADY acknowledged, so no surface alerts. Only `idle` is eligible, and the label is pane-derived and prompt-injectable, so a pattern must anchor on chrome only that CLI draws. → [architecture-invariants#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you](docs/architecture-invariants.md#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you). Tests: `test/session-watching.test.ts`, `test/watching-no-alert.test.ts`. -**Workspace-trust dialog auto-accept** (`session-trust-dialog.ts`, pure + unit tested): Claude Code asks once per directory ("Is this a project you created or one you trust?") before it will read or edit anything, and since Codeman sessions run permission-skipping or classifier-guarded modes the answer is always yes, so a session parked on that dialog is simply stuck. ⚠️ **Match the compacted SCREEN, never the stream.** tmux repaints a row by writing each word and then a cursor-forward (`\x1b[C`) instead of a space, and Ink colours each word separately, so the wire carries `I\x1b[Ctrust\x1b[Cthis\x1b[Cfolder`; stripping the escapes leaves `Itrustthisfolder`, because the spaces are not there to strip, they were never sent. A plain `includes('trust this folder')` therefore never matched a single chunk and the auto-accept was silently DEAD for every session that hit the dialog. `compactScreenText()` removes ALL whitespace instead (plus the `ESC ( B` charset selects that `stripAnsi` does not cover, which would otherwise land inside a phrase as a literal `(B`), which survives both that repaint style and the spaced full-screen redraw. ⚠️ **Never answer it with a blind `\r`.** The layout has changed under us at least twice, and Claude Code 2.1.252 dropped the option numbers, put "No, exit" FIRST and highlights IT by default, so the Enter that answered the old dialog now picks *exit* and the pane dies (`Pane is dead (status 1)`) seconds after the session starts. `trustDialogNextKey()` reads the `❯` marker and returns ONE step at a time (an arrow while the cursor is on the wrong option, Enter only once the screen shows it on the trust option), with the pane re-read between steps, so a dropped arrow costs a repaint instead of the session; a frame that does not say which option is highlighted returns null and waits for the next repaint. ⚠️ The LAST marked option in the text wins, because the direct-PTY fallback reads an append-only buffer where every repaint since launch is still present and an older frame must not out-vote the freshest one. ⚠️ Answering types into a live session, so THREE guards must all hold and none is redundant: a **startup-only window** (`TRUST_DIALOG_WINDOW_MS`, 90s, since the dialog renders before the main UI and leaving it open forever would let an agent transcript that merely QUOTES the dialog trigger an Enter, this file being an example), a **two-marker match** requiring a trust phrase AND one of the dialog's own confirm affordances (`isTrustDialogScreen`), and an **attempt cap** (`TRUST_DIALOG_MAX_ATTEMPTS`, 6: a keystroke can land while Ink is still mounting the widget and be dropped, which is the other half of why sessions got stuck here, but retrying forever would hammer keys into whatever came next; it was 3 while one Enter answered the dialog, and answering now costs at least two keystrokes). ⚠️ It reads `capturePaneText()` and falls back to a deliberately SHORT tail of the terminal buffer only on a direct-PTY session, which has no pane: that buffer is append-only, so a longer tail would keep re-matching a dialog answered minutes ago. ⚠️ **The scan must schedule its own next read** (`_trustDialogTimer`, cleared in `_clearAllTimers()`): it runs from the PTY `onData` handler, which was enough while one Enter answered the dialog, but the arrow that moves the cursor is the LAST output the pane produces, so a two-keystroke answer waiting on more output parks forever with the cursor sitting on the right option (measured on a live 2.1.252 spawn: cursor moved at 6 s, then nothing). +**An exited agent in a live pane** (`paneExit`, #446): panes use `remain-on-exit on`, so `/exit` leaves a pane, session and pid that look alive; `TmuxManager.startPaneExitWatcher()` publishes `SessionState.paneExit` via `session:updated`. ⚠️ Never set `status: 'error'` or null the `pid` for it; the field is TRI-STATE (absent = UNKNOWN, never alive, scoped by `Session.paneExitApplies`); an absent `#{pane_dead_status}` is not 0; a path that starts a command in a pane must clear the record AND persist. → [architecture-invariants#an-exited-agent-in-a-live-pane-paneexit](docs/architecture-invariants.md#an-exited-agent-in-a-live-pane-paneexit) -**Process-tree walks are bounded** (`proc-tree.ts`, pure + unit tested): `collectDescendants(pid, byParent)` is the ONE descendant traversal, fed by a single cached `ps -eo pid=,ppid=` snapshot (`refreshProcSnapshot()` in tmux-manager.ts: in-flight-shared, async because `execSync`'s timeout cannot return at all while spawnSync waits on an unkillable child, and ANY error discards the result rather than caching a truncated `ps`, which would make whole subtrees invisible to the kill path). ⚠️ **The unbounded version took a machine down** (2026-07-30): it ran `pgrep -P ` once per node and recursed with no visited set, no depth limit and no node cap, so across ~28 adopted tmux trees the fan-out exploded while each `pgrep` blocked in the WSL kernel reading `/proc//cgroup`, ending at ~13,000 `pgrep` processes in D-state, a load average above 13,000, and a machine recoverable only by restarting WSL, which cost every running session. Three properties make that impossible and each has a test: a cycle terminates (a real tree has none, a stale snapshot can still produce one), depth is capped (`PROC_WALK_MAX_DEPTH`), node count is capped (`PROC_WALK_MAX_NODES`). The fourth is structural: the function takes a snapshot and cannot spawn anything at all. ⚠️ It lives in its own module because as a private method of `tmux-manager.ts` the regression test had to keep its own COPY of the algorithm, which is a test that passes while the shipped code rots. ⚠️ Truncation is reported through `onTruncated` rather than silently, with BOTH caps named: a silent depth cap hides a deep tree exactly as effectively as a silent node cap hides a wide one. +**Dead-pane respawn resume pin** (`_buildRespawnPaneOptionsWithResumePin()`, session.ts): recovering a dead pane, like a custom-model `restartCli()`, must pin the conversation or claude refuses the reused `--session-id`. The pin takes the first transcript-backed candidate (chain tail, launch seed, own id), never `_claudeSessionId`, adds nothing when none is backed, and is never applied to remote or docker sessions. → [architecture-invariants#dead-pane-respawn-the-resume-pin](docs/architecture-invariants.md#dead-pane-respawn-the-resume-pin) + +**Workspace-trust dialog auto-accept** (`session-trust-dialog.ts`, pure): Claude Code's per-directory trust dialog is always answered yes, or the session is stuck. ⚠️ Match the compacted SCREEN (`compactScreenText()`, all whitespace removed), never the stream (tmux sends words joined by cursor-forwards, not spaces). ⚠️ Never answer with a blind `\r` (newer versions highlight "No, exit" first): `trustDialogNextKey()` returns ONE key per re-read frame (arrow, then Enter only once `❯` is on the trust option), and the LAST marked option wins. ⚠️ All three guards must hold: startup window `TRUST_DIALOG_WINDOW_MS` (90s), two-marker match (`isTrustDialogScreen`), attempt cap `TRUST_DIALOG_MAX_ATTEMPTS` (6). ⚠️ Read `capturePaneText()`; only a direct-PTY session falls back to a SHORT buffer tail. ⚠️ The scan must schedule its own next read (`_trustDialogTimer`, cleared in `_clearAllTimers()`), not rely on PTY output. → [architecture-invariants#workspace-trust-dialog-auto-accept](docs/architecture-invariants.md#workspace-trust-dialog-auto-accept) + +**Process-tree walks are bounded** (`proc-tree.ts`, pure): `collectDescendants(pid, byParent)` is the ONE descendant traversal, fed by one cached `ps -eo pid=,ppid=` snapshot (`refreshProcSnapshot()` in tmux-manager.ts: async, in-flight-shared, and ANY error discards the result rather than caching a truncated `ps`). ⚠️ Never walk a process tree with per-node `pgrep` or unbounded recursion (the unbounded version took a machine down): the walk must terminate on cycles, cap depth (`PROC_WALK_MAX_DEPTH`) and node count (`PROC_WALK_MAX_NODES`), and never spawn anything. ⚠️ Keep it in its own module so the test exercises the shipped code, and report truncation through `onTruncated` naming both caps, never silently. → [architecture-invariants#process-tree-walks-are-bounded](docs/architecture-invariants.md#process-tree-walks-are-bounded) **Auto-resume on usage limit** (opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit, `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time and `SessionAutoOps` arms a timer for reset+2min, then sends Esc + `continue`. ⚠️ Respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected`), which is what prevents `/clear` from wiping the paused conversation. Claude-mode only. → [architecture-invariants#auto-resume-on-usage-limit](docs/architecture-invariants.md#auto-resume-on-usage-limit) -**Plan-usage chip** (`showPlanUsageLimits`, per-device: desktop default **ON**, handhelds OFF via the mobile block in `getDefaultSettings()`): resolve DISPLAY ONLY through `planUsageChipEnabled()` in settings-ui.js, which backs the two call sites that must never disagree (the App Settings checkbox, the chip's visibility). The SAME persisted setting also doubles as the server-side telemetry COLLECTION switch — `readPlanUsageTelemetryEnabled()` (hooks-config.ts) reads it fresh from `settings.json` at every claude session create/respawn (`TmuxManager.createSession`/`respawnPane`), so it applies uniformly to every claude-creation path (interactive Run, cron, Ralph Loop API, quick-start) with no per-session state and no per-request field — a Codeman restart cannot silently kill it (there is nothing per-session to lose). ⚠️ An ABSENT key reads as ON (mirroring `readWorkspaceHooksEnabled()`), so the default is resolved by the READER and `GET /api/settings` stays a plain read that never writes: a reconcile write there ran on every page load and could replace an unreadable `settings.json` with a one-key file. ⚠️ A save sends `showPlanUsageLimits` ONLY when it FLIPS the chip on that device (`planUsageCollectionFlip()` in settings-ui.js): the chip defaults OFF on handhelds, so a phone saving its font size used to persist `false` and switch collection off for every desktop. Claude data comes from Codeman's marked `statusLine.command` exporter (injected as an EPHEMERAL `claude --settings` CLI flag, never written to disk — see `resolveStatusLineCliCommand`), which POSTs `rate_limits` to `POST /api/status-telemetry`, never overwrites a user's hand-authored statusLine (it WRAPS it instead — `findEffectiveUserStatusLineCommand`), and prints the footer through. Main Codex usage comes from a read-only host `account/rateLimits/read` app-server poll at startup and every 5 minutes; exclude model-specific buckets such as Spark, and omit the Codex row when no signed-in limit is available. Distinct from auto-resume, which reacts to Claude's limit *message* rather than showing live %. → [architecture-invariants#plan-usage-chip-statusline-telemetry](docs/architecture-invariants.md#plan-usage-chip-statusline-telemetry), `docs/usage-limits-display-plan.md` +**Plan-usage chip** (`showPlanUsageLimits`, per-device: desktop default **ON**, handhelds OFF): resolve DISPLAY only through `planUsageChipEnabled()` (settings-ui.js). The same setting is the server-side COLLECTION switch, read fresh by `readPlanUsageTelemetryEnabled()` (hooks-config.ts) at every claude create/respawn. ⚠️ An ABSENT key reads as ON in the reader; `GET /api/settings` must never write. ⚠️ A save sends `showPlanUsageLimits` ONLY when it flips the chip on that device (`planUsageCollectionFlip()`), or a phone switches collection off for every desktop. Claude data comes from the statusLine exporter, injected as an EPHEMERAL `claude --settings` flag (`resolveStatusLineCliCommand`), never written to disk, WRAPPING a user's own statusLine, posting to `POST /api/status-telemetry`. Codex comes from a read-only `account/rateLimits/read` poll (main bucket only). → [architecture-invariants#plan-usage-chip-statusline-telemetry](docs/architecture-invariants.md#plan-usage-chip-statusline-telemetry), `docs/usage-limits-display-plan.md` **Orchestrator**: State machine that turns a user goal into a phased plan and drives it to completion: `idle → planning → approval → executing → verifying → (replanning) → completed/failed`. `OrchestratorLoop` (engine) delegates plan generation to `orchestrator-planner` and per-phase verification gates to `orchestrator-verifier`, executing phases via team agents/`task-queue`. State persists under the `orchestrator` key in `state.json`. Distinct from Ralph (single-session autonomous loop) — orchestrator coordinates multi-phase, multi-agent execution. See `docs/orchestrator-loop-architecture.md`. **Cron (`CronJob`s)**: saved, named jobs on a recurring schedule (`once`/`interval`/`daily`/`weekly`) with per-job run history. ⚠️ **Distinct from the legacy `ScheduledRun`** (`/api/scheduled`, a run-now duration-bounded loop); the two never interact and keep separate `Scheduled*` / `Cron*` names. `CronService` **reuses the existing session layer** rather than rebuilding tmux logic. Next-run math is pure and unit-tested in `cron-time.ts` (server-local timezone). The schedule is advanced BEFORE launch so a slow launch cannot re-trigger. → [architecture-invariants#cron-jobs](docs/architecture-invariants.md#cron-jobs), `docs/cron-discovery.md` -**Remote sessions + remote SSH cases**: a case can point at a remote host. The agent runs inside a durable remote `tmux -L codeman-remote` (session name `codeman-ssh-`, deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so an instance on the target host never adopts it), fronted by a LOCAL tmux pane running `ssh`. Attached (`owned:false`) sessions **detach, never kill** on tab close; owned ones propagate `kill-session`. A bounded-backoff watcher auto-reconnects dropped sessions (`remoteAutoReconnect`, default ON). ⚠️ **It revives ONLY when the durable remote tmux session is verifiably still alive** (`remoteTmuxSessionAlive()`, a `has-session` probe over ssh, #355): a clean agent exit (Ctrl-C, Ctrl-D, `exit`) tears that session down, and `isPaneDead()` cannot tell it from a transport drop, so the watcher used to relaunch a FRESH agent after every clean exit (claude only looked fine because its `|| --resume` fallback masked it). An unreachable host answers `undefined`, which also means do not revive. ⚠️ `has-session` prints NOTHING on success, so the probe is classified by EXIT STATUS (`classifyRemoteAliveExit`: 0 alive, ssh's 255 or a timeout unknown, anything else gone); reading stdout classified every live session as gone and silently disabled transport-drop reconnects. The answer is cached per session and forgotten whenever the pane is seen alive again, or a stale `true` from one transport drop would revive the next clean exit. ⚠️ **File reads in a remote case are the second ssh surface** (#415, `src/remote-files.ts`): they go through `buildSshConnectionArgs()` as well, a browser-supplied path is only ever a `shellescape`d token, an unreachable host answers 502 (never 404), the size cap uses the REMOTE size, and no remote file is ever copied onto the server's disk — which is why writes, office previews and thumbnails are deliberately unsupported over ssh (the `PUT` guard sits BEFORE the local path validation, or a same-named local directory such as an sshfs mount takes the write). The probe's symlink resolution FAILS CLOSED (a path it cannot canonicalize is a 404, never its own unresolved string: the directory-only fallback let a `notes.txt -> ~/.ssh/id_rsa` link pass containment), and ssh children are BOUNDED by `src/remote-ssh-limiter.ts` plus one batched probe per attachment-history listing, because terminal output in a remote session is written on the remote host and a prompt-injected agent can print hundreds of `codeman://attach` links. The ATTACHMENT routes (a clicked path outside the case dir) go through the same layer, and which host a record is read from follows the SESSION, never the path string. ⚠️ **Command-injection surface: every ssh command line must flow through `buildSshConnectionArgs()`**, which `shellescape`s every user field. Never hand-build an ssh line elsewhere. ⚠️ Run flows must route remote cases through `POST /api/quick-start`, not `POST /api/sessions` (which stat-validates `workingDir` locally and has no `caseName`). → [architecture-invariants#remote-sessions-over-ssh](docs/architecture-invariants.md#remote-sessions-over-ssh), [#remote-ssh-cases](docs/architecture-invariants.md#remote-ssh-cases), `docs/remote-sessions.md` +**Remote sessions + remote SSH cases**: a case can point at a remote host. The agent runs in a durable remote `tmux -L codeman-remote`, fronted by a LOCAL tmux pane running `ssh`. Attached (`owned:false`) sessions **detach, never kill**; owned ones propagate `kill-session`. Auto-reconnect (`remoteAutoReconnect`, default ON) revives ONLY when `remoteTmuxSessionAlive()` proves the remote session alive. ⚠️ Classify that probe by EXIT STATUS (`classifyRemoteAliveExit`), never stdout. ⚠️ **Command-injection surface: every ssh command line must flow through `buildSshConnectionArgs()`**; never hand-build one. ⚠️ Remote file reads (`src/remote-files.ts`, attachment routes too) take browser paths only as `shellescape`d tokens, resolve symlinks fail-closed, cap on the REMOTE size, never copy to local disk, are bounded by `src/remote-ssh-limiter.ts`, and pick the host from the SESSION, never the path; no writes over ssh (the `PUT` guard must precede local path validation). ⚠️ Route remote cases through `POST /api/quick-start`, not `POST /api/sessions`. → [architecture-invariants#remote-sessions-over-ssh](docs/architecture-invariants.md#remote-sessions-over-ssh), [#remote-ssh-cases](docs/architecture-invariants.md#remote-ssh-cases), `docs/remote-sessions.md` -**Wake-on-LAN (`remote-wake.ts`)**: an optional `RemoteHost.wakeMac` (Codeman builds the magic packet itself) or `RemoteHost.wakeCommand` (single executable path, run without a shell, takes precedence) lets the INPUT route, `POST /api/sessions/:id/wake`, and the user's own create/attach request (`POST /api/quick-start`, `POST /api/sessions` with `attachRemoteSession`, via `ensureHostAwake`) wake a sleeping host instead of writing into a stalled ssh pane. ⚠️ An explicit request — input, the wake button, or the user pressing Run/Attach — and NOTHING else may wake: the auto-reconnect watcher, `handleRemoteSessionDropped`, boot recovery and `cron-service.ts` have no access to the registry (a wake there would re-wake the host seconds after every suspend, and the create wake is wired in the route rather than the shared session service for exactly that reason), which `test/remote-wake.test.ts` asserts as two wiring guards — the second also pins that `server.ts` holds the registry for its LIFETIME only (`drop` on cleanup, `stop` on shutdown) and never calls a waking method. `GET /api/sessions/:id/reachability` merely probes and never wakes. Detection is a throttled bare TCP probe — deliberately no `ServerAliveInterval`, because keepalives move bytes into an idle connection every interval and that is what a byte-threshold idle detector must not read as activity. ⚠️ A host behind `jumpHost`/`socksProxy`/a `ProxyCommand` option is reachability-UNKNOWN (`isProbeable()`): the probe connects to `host:port`, which such a host does not answer even while ssh works, so the registry never buffers for it, never gates create/attach on it (`'unprobeable'`), and `/reachability` answers `reachable: null, probeable: false` — the banner keys on a PROVEN `false`, and the banner's 30 s poller runs only for a host with a wake target (a timer connecting to a host Codeman cannot wake is the same timer-driven traffic the keepalive rule forbids). Input arriving during a wake is buffered (a chunk over 4 KB is dropped whole, never delivered as a fragment; the route answers `{buffered:true}` / `{buffered:true, dropped:true}` so the caller can tell) and flushed in order after `reattachRemote()` with `fromUser` — a flush write that fails drops the rest (logged) rather than retaining it for a wake hours later; send-and-wait blocks instead and answers `OPERATION_FAILED` when the host never returns. ⚠️ In multi-user mode the attach path 403s a non-admin BEFORE the host is looked up: the wake runs an executable, and remote hosts are admin-only infra everywhere else. ⚠️ Browser keystrokes travel over the WebSocket, which deliberately does NOT pass through the registry (that is the hot path), so only the HTTP input path ever queues anything — the banner must not promise queued input for the Wake button. A request that waits on the wake (create/attach, and the button) uses the 40 s `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS`, not the 90 s session default, because the dashboard's reverse proxy cuts a request at its own 60 s `proxy_read_timeout`. The wake fields are re-read from `remote-hosts.json` on recovery and, throttled+cached via `RemoteWakeDeps.resolveRemote`, for a LIVE session, since the persisted `remote` snapshot never sees a field added later. UI: the amber `#hostWakeBanner` (`host-wake-ui.js`) with Wake / "Configure WoL" → `#wakeConfigModal`. The `remote:` SSE family is session-scoped in multi-user mode; a create/attach wake names its requester (`username`) since it has no session yet. `remote-wake.ts` refuses real IO under `VITEST` like `remote-files.ts`. +**Wake-on-LAN (`remote-wake.ts`)**: optional `RemoteHost.wakeMac` (magic packet) or `RemoteHost.wakeCommand` (single executable, no shell, takes precedence) lets HTTP input, `POST /api/sessions/:id/wake` and the user's create/attach (`ensureHostAwake`) wake a sleeping host. ⚠️ Only an explicit user request may wake: never give the registry to the auto-reconnect watcher, `handleRemoteSessionDropped`, boot recovery or `cron-service.ts`, and `GET /api/sessions/:id/reachability` must never wake. ⚠️ Detection is a throttled bare TCP probe; never add `ServerAliveInterval`, and a `jumpHost`/`socksProxy`/`ProxyCommand` host is reachability-UNKNOWN (`isProbeable()`): never buffer, gate or banner on it. ⚠️ In multi-user mode a non-admin attach 403s BEFORE host lookup. ⚠️ WS keystrokes bypass the registry, so the banner (`host-wake-ui.js`) must not promise queued input. Waiting requests use the 40 s `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS`. → [architecture-invariants#remote-ssh-cases](docs/architecture-invariants.md#remote-ssh-cases) -**Docker cases**: a case can point at a **container**, with any of the CLI run modes running inside it. Like remote-SSH this is a **LOCATION OVERLAY on cases, never a `SessionMode` of its own**. Exactly one long-lived container **per case**, shared by all its sessions, so killing a session kills only that session's in-container tmux and **never** `docker stop` while siblings remain. The workspace is a real host dir bind-mounted at the **same absolute path**, which is what keeps file-routes/watchers on real host bytes and makes the in-container transcript projHash match the host. Credentials are **seeded** (RO mount, copied into the container once) rather than shared RW, so in-container CLIs never write refreshed tokens back to the host, and bind mounts are excluded from `docker commit` so exports stay secret-free. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** Config drift is detected via a label hash and a drifted launch is REFUSED rather than silently launched with stale config. ⚠️ A case may instead **ADOPT** a container the user already runs (`DockerCase.owned === false`, mirror of remote-SSH's `owned:false`): Codeman only `exec`s into it and never creates, starts, stops, restarts or removes it, so a missing or stopped container FAILS CLOSED with an actionable message instead of being fixed. Absent = owned, so existing cases are byte-identical. ⚠️ An ADOPTED container may back SEVERAL cases at different in-container directories (`classifyAdoptContainerConflict` in `docker-hosts.ts`: an exact twin on the same container AND directory is refused, an owned container still backs exactly one case, and a container another user adopted is refused), which is what the Add Case panel's "copy an existing case" picker relies on; the wire carries `CaseInfo.docker.owned` ONLY when false, so the picker tests `=== false`, never truthiness. The guarantee is enforced at four independent layers because it cannot be observed by using the feature: `buildDockerStopCommand`/`buildDockerRemoveCommand` throw during pure STRING CONSTRUCTION, `removeDockerContainer` refuses again, drift reports "none" (an adopted container carries no `codeman.confighash` label, so a real comparison would 409 the launch forever), and the boot reaper skips it. ⚠️ Two lifecycle touches the original design missed and that are easy to re-introduce: the full-image export `docker commit`s the container (refused for an adopted case) and the workspace export `docker pause`s it first (skipped — it freezes the owner's processes for the length of the tar). ⚠️ `owned` is applied AFTER `dockerConfigHash`, which takes an explicit field list, or every pre-existing case would trip the drift gate at once. ⚠️ Run modes for a container case come from the CONTAINER (`availableModes`, live-probed): gating the run menu on HOST CLIs (#201) is right for local sessions and wrong here, since a host with no `claude` may run a container that ships one. ⚠️ **A failed probe means opposite things per ownership** — for an ADOPTED case it is a fault worth reporting, for an OWNED one it is the NORMAL state before the first session (the launch chain creates the container), so treating it as a fault hid every agent mode on every freshly linked Docker case behind "start it yourself first". That is why `CaseInfo.docker.owned` is on the wire. ⚠️ Claude is launched WITHOUT `--dangerously-skip-permissions` when the container's exec user is root (Claude Code refuses the flag as root and the refusal is visible only inside the container); which flag to drop is a per-CLI fact, so it is the registry's `overlays.docker.rootCommand`, never a branch. ⚠️ Adoption is **admin-only in multi-user mode**, unlike `docker-link`: linking creates OUR container, whose one bind mount `isWorkingDirAllowed` has already confined, while an adopted container's mounts belong to its owner and one mounting `/` hands the adopter the host. The same reasoning admin-gates the container listing and the in-container directory browser; the preflight instead admits a non-admin for a container already linked to a case they own, because the run menu probes it for every docker case. ⚠️ On the loopback-only prod bind a container cannot reach 127.0.0.1, so in-container hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; otherwise idle detection falls back to output-based. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` (user guide), `docs/docker-cases-plan.md` (design) +**Docker cases**: a case can point at a **container** running any CLI mode inside it, a **LOCATION OVERLAY on cases, never a `SessionMode`**. One long-lived container **per case**, shared by its sessions: killing a session kills only its in-container tmux, **never** `docker stop` while siblings remain. The workspace is bind-mounted at the **same absolute path**. Credentials are **seeded**, never shared RW. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** A drifted config (label hash) REFUSES the launch. ⚠️ An **adopted** container (`DockerCase.owned === false`) is only `exec`ed into: never create, start, stop, restart, remove, `docker commit` or `docker pause` it; fail closed. Test `owned === false`, never truthiness. ⚠️ Apply `owned` AFTER `dockerConfigHash`. ⚠️ Run modes come from the CONTAINER (`availableModes`), and a failed probe is normal for an OWNED case. ⚠️ Root exec user drops the bypass flag via the registry's `overlays.docker.rootCommand`, never a branch. ⚠️ Adoption is admin-only in multi-user mode. ⚠️ Loopback prod needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for in-container hooks. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` -**Docker Compose deployment** (`docker/`, contributed): Codeman itself runs in a container and spawns Docker cases as **SIBLING** containers through the mounted host socket (Docker-outside-of-Docker), never nested. That inverts one assumption the bare-host path takes for granted: the daemon no longer shares Codeman's filesystem, so a bind source valid *inside* Codeman means nothing to it. `resolveDockerDaemonMountSource()` translates sources under HOME into the daemon's namespace via `CODEMAN_DOCKER_HOST_HOME`, and `CODEMAN_CASES_PATH` points the cases dir at a host-absolute bind mount so a workspace resolves to the SAME absolute path on both sides (which is what keeps the transcript projHash matching, per Docker cases above). ⚠️ **`CODEMAN_CASES_PATH` must move every consumer or none**: it is resolved once in `config/cases-dir.ts` because `src/cli.ts` resolves case paths too, and when only the server's `CASES_DIR` learned the override, `codeman skill install --case ` reported "Case not found" on exactly the deployment the override exists for. ⚠️ **`.dockerignore` patterns match the WHOLE context-relative path**, so a bare `.env` line excludes only the ROOT file: `docker/.env` (which holds `CODEMAN_PASSWORD` and any provider keys) rode `COPY . .` into the image until `**/.env` was added — verified in both directions with a real build context. ⚠️ A Compose LONG-form bind (`type: bind`) **creates a missing host source directory ROOT-OWNED** rather than refusing. `Start-Codeman.sh` pre-creates both `CODEMAN_APPDATA_PATH` and `CODEMAN_CASES_PATH` on the host before `up`, which is what keeps the daemon from ever having to materialise either as root in the first place; the container ALSO starts as root (`cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the base `cap_drop: ALL`; `test/docker-entrypoint.test.ts` pins that list) so `docker/entrypoint.sh` can correct a bind source that turns up root-owned anyway (a restored backup, a cleared directory, plain `docker compose up` run without the script) before dropping to `PUID:PGID` via `setpriv` — a directory owned by neither root nor `PUID:PGID` is never re-owned, since that ownership is not this container's to reassign; it is PROBED for writability as the runtime account (`setpriv ... test -w`, so ACLs, group-writable trees and CIFS/NFS mounts pass) and refused with a message naming path, owner and PUID:PGID if that fails. ⚠️ `KILL` is in that list for tini, not the entrypoint: `init: true` keeps tini as root while the server runs as PUID, and without CAP_KILL its SIGTERM forward fails and the server is SIGKILLed on every `compose down`/`restart` instead of flushing state. ⚠️ `/opt/codeman-cli` (the runtime-owned CLI prefix) is APPENDED to `PATH`, never prepended, and the entrypoint pins its own `PATH` to the system dirs: the root part of the start resolves `setpriv` by bare name, and a prefix ahead of `/usr/bin` let a planted `setpriv` run as uid 0 (measured). `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` drops `--memory-swap` (and filters only that one kernel warning) for hosts without swap accounting; `--memory` still applies. ⚠️ The deployment ALSO self-updates in place (the repo bind mount at `/opt/codeman` + a restart-by-exiting supervisor) — see Self-update below and `docs/docker-self-update.md` before touching `server.Dockerfile`, the compose file or `.env.example`, since each is an input to the updater's environment gate. `docs/docker-compose.md` + `docker/README.md` (user guides) +**Docker Compose deployment** (`docker/`): Codeman runs in a container and spawns Docker cases as **SIBLING** containers via the host socket, never nested. `resolveDockerDaemonMountSource()` maps HOME bind sources into the daemon's namespace (`CODEMAN_DOCKER_HOST_HOME`); `CODEMAN_CASES_PATH` makes workspaces resolve to the same absolute path on both sides. ⚠️ `CODEMAN_CASES_PATH` must move every consumer: resolve it only via `config/cases-dir.ts`. ⚠️ `.dockerignore` matches whole paths: keep `**/.env` or `docker/.env` secrets ship in the image. ⚠️ Long-form binds create missing sources ROOT-OWNED: `Start-Codeman.sh` pre-creates them, and `docker/entrypoint.sh` (root, `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against `cap_drop: ALL`, `KILL` for tini; pinned by the test) fixes ownership then drops to `PUID:PGID` via `setpriv`, never re-owning foreign dirs. ⚠️ Append `/opt/codeman-cli` to `PATH`, never prepend. ⚠️ `server.Dockerfile`, the compose file and `.env.example` feed the self-updater's environment gate (`docs/docker-self-update.md`). → [architecture-invariants#docker-compose-deployment](docs/architecture-invariants.md#docker-compose-deployment), `docs/docker-compose.md` -**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. **No code outside `stock.ts` may branch on a CLI id**; behaviour that genuinely differs is either a capability field or a NAMED PROFILE selected by one (`profiles.ts`), and `test/cli-registry-no-id-branching.test.ts` fails the build if an id check reappears — it matches `===`, `!==`, `case '':` and `[...].includes(mode)`, because an earlier `===`-only version let 36 negated branches survive the conversion (including a seven-mode Ralph chain whose own comment asked the next person to keep it in step with `isExternalCliMode()` by hand). A second CI-gated guard, `test/frontend-cli-no-id-branching.test.ts`, covers the two frontend files the run-menu consolidation (#458) touched, `session-ui.js` and `mobile-overview.js`: its allowlist is keyed by expression rather than line number, with an occurrence count per entry, so a new branch reusing an already-approved expression fails as a count mismatch instead of riding in on the old approval. ⚠️ Config contains no shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. ⚠️ `external`, `hooks` and `altScreen` are three INDEPENDENT capabilities on purpose; deriving one from another shipped the `until=stop`-hangs-on-shell bug. ⚠️ Three capability fields carry a REGEX from config (`discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`) and all three must compile through `compileVersionRegex()`, which caps length and refuses nested quantifiers; `workingLine` is the one that runs on the PTY hot path. ⚠️ **`param` is TWO namespaces.** `launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a LAUNCH PARAM; the legacy `Config` wire field is a separate namespace, bridged only by `launch.legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no load error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. Codex is the entry where the two names differ (`bypassApprovals` vs `dangerouslyBypassApprovals`) and therefore the one that catches a regression. ⚠️ Six fields are DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured (except `accent`, measured against styles.css on 2026-09-21), so re-measure before wiring one up; the list is pinned so it cannot quietly grow. Spawn commands are pinned as literal strings in `test/cli-registry-spawn-golden.test.ts`, remote/docker pane commands in `test/location-overlay-commands.test.ts`. ⚠️ That second golden no longer covers **remote claude or remote omp**: both now have their own arm in `buildRemoteLaunchCommand` (a `--session-id || --resume` pair, and `--continue`, so a respawn continues the same conversation) and never reach `defaultRemoteCommandForMode`, which is what that test asserts. Their real pins are `toContain` substrings in `test/tmux-manager.test.ts` and `test/remote-shared-sessions.test.ts`; changing either arm will NOT fail the golden. ⚠️ Anything reading the registry resolves it AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks) — a module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. `~/.codeman/clis.json` overrides any entry (read-only in this release; nothing writes it, so importing the registry has no filesystem side effects). → `docs/cli-registry.md` +**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` (discovery, launch argv template, env handling, `capabilities`, and the `overlays` behind remote/docker pane commands). **No code outside `stock.ts` may branch on a CLI id**: use a capability field or a NAMED PROFILE (`profiles.ts`); `test/cli-registry-no-id-branching.test.ts` and `test/frontend-cli-no-id-branching.test.ts` enforce it. ⚠️ Config holds typed argv tokens, never shell text; literals are validated at LOAD time and a bad one rejects the whole entry. ⚠️ Keep `external`, `hooks` and `altScreen` independent; never derive one from another. ⚠️ Config regexes (`discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine`) must compile through `compileVersionRegex()`. ⚠️ `privilegedParams[].param` names a LAUNCH PARAM, not the legacy `Config` field (bridged only by `launch.legacyConfigAliases`); a wrong name silently clamps nothing. ⚠️ Resolve the registry AT CALL TIME, never in a module-level const. ⚠️ Remote claude/omp arms of `buildRemoteLaunchCommand` are not covered by the pane-command golden. `~/.codeman/clis.json` overrides entries (read-only). → [architecture-invariants#cli-registry](docs/architecture-invariants.md#cli-registry), `docs/cli-registry.md` -**External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. ⚠️ **Work detection is no longer part of that gate**: it is per-CLI `capabilities.workDetect` data (see the ❯ note above), so a CLI that declares its own glyph and working line gets the same screen-probed idle confirmation claude gets, and one that declares neither keeps the output-stabilization behaviour. All eight **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. ⚠️ **Codex sessions use PREDICTIVE WRITE-THROUGH echo, never the buffer overlay** (`_localEchoPolicy` in `_updateLocalEchoState`, terminal-ui.js): codex's composer reacts per keystroke ("/" pops a live-filtering picker, arrows edit server-side state, the composer grows as it wraps), so buffer-until-Enter starved it into issues #218/#219/#220/#222 and stays disabled (`_localEchoEnabled` remains false for codex). Instead, `PredictiveEchoAddon` (separate `vendor/xterm-predictive-echo.js` bundle) paints each keystroke at the predicted cell while the wire path stays BYTE-IDENTICAL: the onData hook (`_predictHookOnData`) is a plain statement with no `return`, so control always falls through into the untouched send path — pinned by vm and E2E byte-identity tests. Predictions reconcile against the parsed buffer and only while the cursor sits on the measured composer row (`isCodexComposerRow`, `/^› /`). Codex also **drops keystrokes that share a PTY read with a bracketed paste**, so flushed text and the paste sequence must go out as separate delayed writes (mirroring the Enter branch's delayed `\r`). Tests: `test/local-echo-codex-gating.test.ts`, `test/codex-predictive-echo.test.ts` (E2E vs real codex), `packages/xterm-zerolag-input/test/codex-replay.test.ts`. ⚠️ **Pi is the opposite kind of CLI and needs the opposite instincts**: it has NO permission prompts and no sandbox, so there is no bypass flag to send and Codeman must not invent one; its privileged knob is the tri-state `approveProjectTrust` (`--approve`/`--no-approve`), which makes pi EXECUTE repo-local `.pi/extensions` TypeScript, so the multi-user clamp puts pi in the **materialize** branch (an absent config still yields `--no-approve` for a non-granted owner) and `--api-key` is never wired. Pi stays OUT of `isAltScreenStripMode()` (main-screen TUI, and its 0.84.0 fullscreen mode is runtime-switchable via `/settings`, where the alt screen is load-bearing), and lands on the `'buffer'` echo policy via the `_updateLocalEchoState` fallthrough. Pi's own tests: `test/pi-mode.test.ts`, `test/routes/external-cli-bypass-clamp.test.ts`; user guide `docs/pi-integration.md`. ⚠️ **Grok is codex-shaped on permissions but opencode-shaped on rendering**: its bypass switch is `alwaysApprove` (`--always-approve`, grok's `bypassPermissions` mode — the Run button sends it `true` like antigravity's, and the clamp's only-if-sent branch strips it for non-granted owners), while its fullscreen alt-screen TUI keeps it OUT of `isAltScreenStripMode()`; the resolver version-probes `grok --version` like pi's (npm squatters exist for the name — `GET /api/grok/status` surfaces path + version), and grok lands on the `'buffer'` echo policy via the fallthrough (UNMEASURED against a live authenticated session; if its composer turns out per-keystroke-reactive like codex, flip it to the `'off'` branch). Grok's own tests: `test/grok-mode.test.ts`, `test/grok-cli-resolver.test.ts`; user guide `docs/grok-integration.md`. ⚠️ **DeepSeek breaks three of this family's assumptions, so do not pattern-match it onto its siblings.** (1) The agent is a **PROFILE, not the binary**: `dsh` is a launcher over `$DSH_HOME/profiles/` and DeepSeek ships only `web`/`headless`/`base`, so the terminal front door is ALWAYS third-party and "installed" ≠ "runnable" — the Run button gates on `isDeepSeekRunnable()` (binary AND a pane-capable profile) while `isDeepSeekAvailable()` gates the "add a profile" affordance; a `web`/`headless` profile is refused at spawn because it cannot drive a pane. (2) The permission switch is the **`DSH_PERMISSION_MODE` env export, not a flag** (`read-only`/`workspace-write`/`danger-full-access`) — the harness has none, and this is the one legitimate exception to the effort-style env-var ban because it is read with `??` as a boot-time default, so it stays soft; absent = `workspace-write`, which asks, hence the only-if-sent clamp branch, clamping to `workspace-write` (never `read-only`, which would break the workspace). ⚠️ **That clamp needs a second half no other CLI needs**, because the switch is an env var and `DSH_*` is an allowlisted `envOverrides` prefix: `applyEnvOverrides()` runs AFTER `_configureCliEnv()` in tmux-manager, so a non-granted owner sending `DSH_PERMISSION_MODE` on the SAME request would land last and hand back exactly the privilege the config clamp removed. `clampEnvOverridesForOwner()` (session-routes.ts) DROPS `DSH_PERMISSION_MODE`, `DSH_HOME` and `DEEPSEEK_BASE_URL` for a non-granted owner (the last because `_configureCliEnv()` forwards the SERVER's own `DEEPSEEK_API_KEY` into the pane, so a redirected base URL would send it to a foreign host) (dropping falls through to what `_configureCliEnv()` exports, which is the clamped value); `DSH_HOME` is there because it points the launcher at a profile tree whose plugin code runs at BOOT, before any approval row applies. Every OTHER CLI's bypass is a command-line flag reachable only through its config, which is why the config clamp alone is the whole gate for them. (3) It is the **only non-claude mode that passes `hooksAvailableForMode()`**, and for it alone that predicate is a per-SESSION question rather than a per-mode one (`deepSeekConfig.statusReporting: false` disarms the bridge, so every call site passes `sessionHookOptions(session)`; answering from the mode there re-creates the infinite-wait-dressed-as-a-timeout the guard exists to prevent). It passes because the terminal front door reports idle/working/blocked to a supervisor over a generic env-gated contract and `deepseek-status-shim.ts` makes Codeman that supervisor — real `stop`/`blocked` signals, real Approvals Inbox items, plus the `agent_working` event that clears an alert answered in the terminal. ⚠️ The resolver needs the strictest identity probe of the family (`dsh --help` must say `DeepSeek Harness`) because Debian ships an unrelated `dsh` (dancer's shell) that would pass a version probe. Model is NOT a session field (it is a profile composition entry). ⚠️ `hooksAvailableForMode()` is about hook SIGNALS and is not a stand-in for "is this a claude session": Read My Mind and intent capture read Claude's own transcript and compare `mode === 'claude'` directly, because when `deepseek` earned a yes the shared predicate silently widened both to a mode with no transcript to read (pinned by a static check in `test/deepseek-mode.test.ts`). ⚠️ **It is also the only external CLI whose answers are READ FROM DISK rather than scraped off the pane**: `deepseek-transcript.ts` reads `$DSH_HOME/sessions///session.jsonl.zstd` and backs the `last-response` route for dsh, because the pane segmenter served dsh-TUI's ASCII-art SPLASH as the worker's answer (measured), which anything polling for a first answer reads as an answer. Three traps live in that file: dsh appends **one zstd FRAME per write** and Node's `zlib` zstd decoder stops at the first (a real 56-line transcript decoded as 1 line, so the module walks frame headers itself; a Node older than 22.15 has no zstd and falls back to the pane); every turn also records a **plugin-sourced `user/message`** (the runtime-context snapshot) that must not render as the user's words; and a failed `turn/end` is surfaced as `Turn error: …` rather than as an empty string that reads as "still thinking". ⚠️ Session→transcript pairing is by the header's own `cwd` plus a ±60 s boot window, never by reproducing dsh's directory mangling (which has already changed form once) — and NEVER by newest-mtime alone, which handed a fresh worker its predecessor's answer in the same case dir. DeepSeek's own tests: `test/deepseek-mode.test.ts`, `test/deepseek-cli-resolver.test.ts`, `test/deepseek-transcript.test.ts`; user guide `docs/deepseek-integration.md`. OMP (`omp`) needs no bypass flag (the CLI's own `~/.omp` config governs trust/model routing, defaulting to `tools.approvalMode: yolo`), so its registry entry declares `privilegedParams: []` and a launch spec that only ever passes `--model`/`--resume`/`--continue` — but the multi-user clamp is NOT a no-op for it: `OMP_*` is an allowlisted `envOverrides` prefix, and the two credential-resolution keys it admits, `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN`, are clamped in `clampEnvOverridesForOwner()` for a non-granted owner, the same shape as `DEEPSEEK_BASE_URL`. Separately, `PI_*` is already allowlisted (pi needs it) and omp reads several of its knobs too (`PI_CONFIG_DIR`, `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`, `PI_SUBPROCESS_CMD`, `PI_SHELL_PREFIX`) — a redirected `PI_CONFIG_DIR` moves the `~/.omp` tree `omp-session-resolver.ts`/`omp-transcript.ts` hardcode, silently breaking pinning/history; this is a known gap shared with pi, not fixed here. → [architecture-invariants#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp) +**External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token parsing, ❯ readiness; readiness is output stabilization); work detection is per-CLI `capabilities.workDetect` data, not this gate. All eight **require tmux, no direct PTY fallback** (secrets go via socket-scoped `tmux setenv`, never the command line). ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope. ⚠️ **Codex uses predictive write-through echo, never the buffer overlay**: `_predictHookOnData` must never `return` (wire path stays byte-identical), and flushed text and a bracketed paste must go out as separate delayed writes. ⚠️ **Pi**: no bypass flag, never invent one; `approveProjectTrust` executes repo code, so it is in the clamp's **materialize** branch; never wire `--api-key`. ⚠️ **Grok**: `alwaysApprove` is stripped for non-granted owners (only-if-sent). ⚠️ **DeepSeek**: the agent is a PROFILE (Run gates on `isDeepSeekRunnable()`); the permission switch is the `DSH_PERMISSION_MODE` env var, so `clampEnvOverridesForOwner()` must DROP `DSH_PERMISSION_MODE`, `DSH_HOME` and `DEEPSEEK_BASE_URL` for non-granted owners; `hooksAvailableForMode()` is per-SESSION for it (pass `sessionHookOptions(session)`) and is never a stand-in for `mode === 'claude'`; answers come from `deepseek-transcript.ts`, paired by header `cwd` + boot window, never newest-mtime. ⚠️ **OMP**: `OMP_AUTH_BROKER_URL`/`_TOKEN` are clamped the same way. → [architecture-invariants#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp) -**DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. ⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. ⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. ⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs. +**DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. ⚠️ Every piece is load-bearing: exactly one server (a second click REUSES it), restarted when the browser authority changes (`--trusted-host`, last asker wins), killed on shutdown via `stopDeepSeekWeb()` (the detached child would otherwise outlive Codeman and hold its port), and failures returned to the caller. ⚠️ Never hardcode the port: search from 3080 across 40, detect free ports by BINDING, then wait for the server to really answer. ⚠️ Both `POST` and `DELETE` must stay behind `canUsernameRunPrivilegedCommands` (booting a profile runs its plugin code; the server is shared). → [architecture-invariants#deepseek-web-ui](docs/architecture-invariants.md#deepseek-web-ui) -**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET /v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. ⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. ⚠️ **Two launch paths, chosen by mechanism, not preference — see the second paragraph below for why**: opencode/codex/gemini/pi/grok/deepseek/omp apply the selection ONE-SHOT, before the session/process ever exists, with no restart at all; claude alone still applies a selection by **restarting the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` relaunches a CLI in an existing pane, so a CLI whose launch declares a `fallback` chain (claude) gets a conversation pinned as `resumeSessionId` for that one respawn: `--session-id ` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume || --session-id ` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. ⚠️ **The dead-pane respawn needs the same pin and shares it** (`_buildRespawnPaneOptionsWithResumePin()` in `session.ts`, used by `restartCli()`, the dead-pane respawn in `_setupOrAttachMuxSession()`, and its create path when that path RELAUNCHES a CLI: after a failed respawn, or when tmux lost the whole session rather than the pane): a pane whose agent EXITED owns a transcript too, so recovering one with the bare launch line hit the same refusal and the conversation was stranded behind a tab that looked merely idle. The pin walks three candidates in order — the conversation chain's tail, the launch seed, then the session's own id — and takes the first one a transcript backs, never `_claudeSessionId` (which also holds history-correlated GUESSES keyed on the working directory, and launching from one would open and write to a conversation that was never this pane's). ⚠️ The create-path pin is written to `_resumeSessionId` as well, so unlike `restartCli()`'s one-respawn pin it PERSISTS through `toState()` as `resumeSessionId`: that field means "what the user asked to resume at creation, or what recovery pinned", and after a dead-pane respawn `_claudeSessionId` names whatever the walk actually pinned rather than the chain tail. ⚠️ A candidate no transcript backs is passed over, and falling off the end of the walk ADDS no pin (the options keep whatever launch seed they already carried): a divergent pin leaves `--session-id ` in the fallback branch, where a failed resume collides all over again, while pinning an id with no transcript prints claude's "No conversation found" into a brand-new session's scrollback and costs the running branch its `nice` priority (`wrapWithNice()` prefixes only the first branch of an `a || b`). ⚠️ Remote and docker sessions are never pinned: their pane commands are already self-healing, the conversation lives on the far side, and a local id resolves to nothing there. ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. ⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. **Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct, and codex only speaks the Responses API since Feb 2026 (`wire_api = "responses"`) — re-verified live against a llama-swap deployment that DOES answer `/v1/responses` (an earlier test's harder failure against a different deployment does not reproduce everywhere): a plain, no-tool-call chat turn gets a real reply, but a real tool-call attempt came back as `agent_message` TEXT (the tool-call JSON printed as the answer) rather than an executable `function_call` item — confirmed via `codex exec --json`'s raw event stream. Tool execution is what makes codex a coding agent, so it remains not usable for real work either way, just with a more precise failure mode than a flat protocol break; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek's originally-reported `HTTP_404` is root-caused and fixed — its bundled `@deepseek-ai/dsh-llm-deepseek` module builds `${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` of its own (confirmed by installing the real package and reading its source), so a new `appendV1Suffix` flag on its registry entry (alone — claude/gemini must not get it) runs `endpoint.baseUrl` through `withV1Suffix()` before writing it, live-confirmed against llama-swap (`.../chat/completions` 404s, `.../v1/chat/completions` succeeds) though not yet re-run through an actual `dsh` binary, which isn't installable in this environment; antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. ⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`, JSON-escaped against a literal `` via the exported `escapeScriptJson()` since `label` is a user-`clis.json`-settable string unlike the neighbouring booleans-only `__codemanCliAvailable`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces. One entry per (capable, INSTALLED CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)", filtered through `isCliAvailable()` like the stock entries. Clicking one calls `selectCustomModelEntry(mode, endpointId)` (`session-ui.js`), which re-fetches the endpoint (never trusts anything cached from the dropdown's render — the 5-minute sweep below or a settings edit may have changed it since) and decides the model: exactly one discovered model launches straight away, two or more open `#customModelPickModal` to ask, with `defaultModelId` marked but never auto-chosen (asking exists so ONE launch can deliberately differ from the saved default). ⚠️ **The picker promotes exactly one row to the top**: whichever model llama-swap reports `ready` right now (tagged "Currently loaded", queried via `GET /api/model-endpoints/:id/running-status`, client-side bounded to ~800ms via `Promise.race` so a sleeping/firewalled endpoint cannot leave the modal invisible for the route's own 5s server-side timeout) beats a merely remembered choice, and — only when nothing is currently loaded — the model actually launched last for this exact (harness, endpoint) pair (tagged "Last used", read from the per-device `codeman:customModelLastUsed::` localStorage key). Neither tag reorders past the top, and the "Default" pill is a SEPARATE span rather than a third value of the same slot, so a promoted row that is also the endpoint's `defaultModelId` shows both (on a single-purpose GPU box that is the common case; an exclusive slot silently dropped the Default marking for exactly that row). "Last used" is written by `_runCustomModelEntryViaRestart` (claude) and `_quickStartWithCustomModelConfirm` (every one-shot launch; `runCustomModelEntry` itself only dispatches between the two) only once the model is actually applied — never on the mere click — because a context-window-warning decline means this exact model cannot work with this CLI at all, and promoting a model that cannot launch would be actively wrong, not just premature. Either way the actual launch (`runCustomModelEntry`) routes through `run()` itself via a temporary `_runMode` swap — never `setRunMode()`, which would persist it as the user's new default — rather than a parallel dispatch table, which is what gives a custom-model launch the same `_runInFlight` lock every other Run click gets and means a CLI whose injection recipe lands later needs no update here. It then GETs `/api/sessions/:id/wait?until=idle&timeout=20000` on that session BEFORE applying — measured live, a freshly launched CLI reports itself `busy` for its own startup (boot spinner, workspace-trust check) well before the apply call would otherwise reach it, and the apply route's `isBusy()` guard correctly can't tell that apart from a real turn in progress, so every fresh launch failed with `SESSION_BUSY` until this wait was added. A timeout there is a normal 200 per the wait endpoint's own contract, never an error, so a session still busy after 20s just reaches the apply call anyway and gets that route's own honest error. It then calls `POST /api/sessions/:id/custom-model` on the session `run()` produced, guarded by snapshotting `activeSessionId` before the call and requiring it to have actually changed after — every `run*()` handles its own failure internally and returns normally rather than throwing, so a declined/failed launch must not silently re-point and restart whatever session was already open. ⚠️ The apply call reads the response body itself (`_api()`) rather than `_apiJson()`, which unwraps success but silently discards a failure body — losing the one thing (`error`) that distinguishes "still busy", "remote/Docker session" and everything else the route can report (⚠️ neither route validates `modelId` against the endpoint's discovered list, deliberately: discovery can be up to 5 minutes stale, so a 400 there would refuse a launch that works — a typo'd id fails on the CLI's own first request instead); the resulting toast is `type: 'error'` with an explicit `duration: 0` (no auto-dismiss, an explicit close button) at that one call site — not a blanket sticky-error default, which stacked unbounded on `.toast-container` with no cap or eviction — precisely so a message worth diagnosing survives long enough to be read instead of vanishing on the usual 3s timer. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to launch with). ⚠️ **Every saved endpoint's models also re-discover themselves automatically**, a `this.cleanup.setInterval` in `server.ts` (`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, 5 minutes, off under `testMode` like the Codex plan-usage poll beside it) calling the exported `refreshAllCustomModelHosts()` (`custom-model-routes.ts`) — one endpoint unreachable on a cycle never blocks the others, and a read-modify-write PER HOST (re-reading the store before each splice, keyed by id) means an admin's concurrent edit or delete wins over a sweep that started before it, never the reverse. +**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`): points a session at a user-configured OpenAI-compatible endpoint (store `custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`, 0600). `authStyle` is `bearer` or `api-key`, never both headers (hangs the server). The per-CLI redirect is registry data, `capabilities.customModelInjection` (`env` / `configContentEnv` / `configDir` / `unsupported`), computed by the pure `custom-model-injection.ts`; ⚠️ `configDir` writes an isolated per-session config, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config. ⚠️ Claude applies via `Session.restartCli()` (a relaunch in the existing pane, so the conversation is pinned as `resumeSessionId`), the rest one-shot at launch; retired env keys must also be `setenv -u`'d (`_pendingEnvUnsets`), since tmux env survives `respawn-pane`. ⚠️ Remote/Docker sessions are refused (400). ⚠️ Persist only KEYS (`__customModel`), never values (they carry the API key). ⚠️ Every redirectable var must be in that CLI's `privilegedEnvKeys`, and `ANTHROPIC_*` stays out of claude's `allowedPrefixes`. ⚠️ The Run-menu picker builds entries from `window.__codemanCustomModelClis` (escaped via `escapeScriptJson()`), never a hardcoded CLI id list, and launches through `run()` via a temporary `_runMode` swap, never `setRunMode()`. → [architecture-invariants#custom-model-endpoint-profiles](docs/architecture-invariants.md#custom-model-endpoint-profiles) -**Everything below landed after the initial backend + picker cut, each confirmed live against a real llama-swap deployment.** ⚠️ **llama-swap runs one model at a time, and switching can disrupt ANOTHER live session** — before applying, both apply routes call llama-swap's own `GET /running` (feature-detected via `getLlamaSwapStatus()`, `custom-model-routes.ts`; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is simply never checked). If a different model is loaded and ready AND another live session's own selection is using it, the apply returns `{requiresConfirmation, currentlyLoadedModel, affectedSessions}` instead of switching silently; retrying with `confirmedSwap: true` skips the check, and switching with nothing else affected proceeds immediately. ⚠️ **The swap question and the context-floor question below have SEPARATE flags** (`confirmedSwap`, `confirmedContext`), because the context check runs first and while both shared one `confirmed` a user who clicked past a too-small context silently consented to evicting another session's model too; the legacy `confirmed` still means both, since it shipped in the HTTP-API-only cut. llama-swap also has no dedicated "switch model" endpoint — the only thing that actually starts a swap is a real inference request naming the model (confirmed live: applying a selection alone never reached llama-swap's own logs, since nothing had asked it to load anything) — so both routes also fire `triggerLlamaSwapLoad()`, a fire-and-forget `POST /v1/chat/completions` with `max_tokens: 1`, whenever the target model isn't already loaded and ready. ⚠️ **That launch-time check cannot catch a swap caused by a DIFFERENT session's LATER, ordinary use** — confirmed live: a second Codex session picking a different model launched with no warning at all (nothing conflicted at that exact instant), yet it silently evicted the first session's model regardless, since llama-swap has no push notification of its own. `detectCustomModelSwapDisplacements()` (`custom-model-routes.ts`) is a separate periodic sweep (`server.ts`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS` = 20s) that compares each live custom-model session's own `modelId` against what `/running` actually reports loaded, broadcasting a `custom-model:swapped-out` SSE event — shown as a global toast, never tied to the displaced session's own tab, since the whole point is telling the user before they type into it — the first time a mismatch appears, via a caller-owned de-dupe `Set` cleared once that session's own model is loaded and ready again so a later, genuinely new displacement notifies again rather than staying silently un-notified forever after the first one. ⚠️ **Context length is read from the REAL launch command, never `/props`** — `/props?model=`'s `default_generation_settings.n_ctx` was confirmed live to report a `--fit-ctx`-launched backend's theoretical/trained maximum rather than the real runtime-configured size (a measured 154112-vs-16384 discrepancy, caught only because the unfixed value still overflowed), so discovery parses the actual configured size straight out of `/running`'s own `cmd` field instead (`parseCtxFromCmd`: `--fit-ctx ` first, then plain llama.cpp `-c`/`--ctx-size`), falling back to `/props` only when `cmd` states no recognizable flag at all. ⚠️ **Claude alone gets a context-window FLOOR check, on top of the ceiling `contextLengthVar` already fixes** — `exceedsSafeContextFloor()` (gated on the registry declaring `contextLengthVar`, so a no-op for every other CLI by construction) compares a model's discovered context against `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000): confirmed live, twice, that Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on the very first message, before any conversation history exists to compact, so a smaller real context fails outright regardless of what `CLAUDE_CODE_MAX_CONTEXT_TOKENS` says (that var only controls when HISTORY gets compacted, and there is none yet on message one). Below the floor, the apply returns `{requiresContextWarning, modelId, contextLength, minSafeContextTokens}` instead of launching, shown as an in-app dialog naming the actual fix: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on `--fit-ctx` auto-fit, which optimizes for the biggest MODEL that fits rather than the biggest CONTEXT. ⚠️ **A fresh, isolated `CLAUDE_CONFIG_DIR` looks like a brand-new Claude Code profile and replays its ENTIRE first-run sequence on every launch** — the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with a bypass-permissions flag) a one-time warning about it, confirmed live, none of which a real, already-onboarded profile shows again. `skipFirstRunPrompts` (claude's entry only, requires `apiKeyTrustFile` since it reuses the same file) pre-seeds that same "already been through this" state: `hasCompletedOnboarding` and this session's own `projects[workingDir].hasTrustDialogAccepted` merge into the same `.claude.json` the API-key trust file already writes to, and `skipDangerousModePermissionPrompt` merges into `settings.json` (a different file, same corrupt-tolerant merge). ⚠️ **The loading banner shows the REAL backend log line, not a guess, and has no countdown or auto-timeout at all.** `getLatestLlamaSwapLogLine()` holds one `GET /api/events` SSE connection open per endpoint (confirmed live to stay open indefinitely — read past 220KB over 8s with no `done`; idle-closed after 30s via `pruneIdleLlamaSwapLogTails`, same 20s sweep as the swap-displacement check above), parsing `logData` frames and keeping only `source: "upstream"` (the real `llama-server` process's own stdout) lines, never `source: "proxy"` (llama-swap's own request-access log). ⚠️ `GET /logs` — the endpoint this feature's own first cut targeted, since the name suggested it — was confirmed live to carry ONLY the proxy log and never a single backend line, even seconds after a real, verified model swap; caught and corrected by a live check before merge, not after. The banner itself dropped its size-scaled expected-time estimate and matching auto-timeout (a guess dressed up as a fact that could kill a genuinely slow load partway through on slower hardware) for a generic hardware/model-size disclaimer plus a user-driven **Cancel** button (`_showCenterStatus`'s `onCancel` option, a real button distinct from the plain "×" close glyph an `'error'`-type banner gets) that ends the wait and closes the session on the user's own call rather than a guessed deadline. +⚠️ **llama-swap endpoints** (one model at a time): the apply routes check `GET /running` and return `requiresConfirmation` before evicting a model another live session uses; `confirmedSwap` and `confirmedContext` are SEPARATE flags and must stay so. Claude alone gets a context floor (`CLAUDE_MIN_SAFE_CONTEXT_TOKENS`); context is parsed from `/running`'s `cmd`, never trusted from `/props`. Backend log lines come from llama-swap's `/api/events` `upstream` source, never `/logs`. → [architecture-invariants#custom-model-endpoint-profiles](docs/architecture-invariants.md#custom-model-endpoint-profiles) -**Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w-` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization) +**Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms) so a double click cannot create duplicate `w-` sessions; `_ensureCreatedSessionVisible()` runs before `selectSession()` and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first both render exactly one tab. ⚠️ **Closing has the mirror-image race**: `closeSession()` must read `wasActive` BEFORE its `await` and announce the delete via `_closingSessions`, and `_onSessionDeleted` skips the active-session handoff for ids in that set; never read `activeSessionId` after the fact. The fallback picks the first `sessionOrder` entry still in `sessions`. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization) -**Session lineage lines** (tab → tab it spawned, `sessionLineageLines`, per-device, desktop default ON): a create request may name the session that spawned it, as a `parentSessionId` body field on `POST /api/sessions` / `POST /api/quick-start` or the `X-Codeman-Parent-Session` header (the agent skill sets that once on its shared curl invocation, so every spawn recipe carries it). `resolveParentSessionId()` (route-helpers.ts) **resolves rather than trusts** it: exact id, else a UNIQUE ≥8-char prefix (ids reach agents truncated), it must be a live session the caller can see AND carry the same owner, and **anything unresolvable is DROPPED, never a 400** — a cosmetic field must not be able to fail a worker spawn. It rides `toState()` into `session_created`, so there is no new SSE event. ⚠️ Rendering is an ADDITIONAL LAYER on the existing SVG pass (`_appendLineageConnectionLines` called at the tail of `_updateConnectionLinesImmediate()`, exactly like ultracode), sharing one batched read→write reflow and the `tab:` rect cache; geometry is pure in `computeLineagePath()` (constants.js). ⚠️ **ONE shape, and the second one was the bug**: every pair (flat strip or wrapped) gets a U-bridge hanging below the strip, anchored on both tabs' BOTTOM edges. A wrapped strip used to get a parent-bottom → child-TOP bezier with a ~14px row gap to bend in, which drew a flat line hidden in the gap with siblings overprinting. ⚠️ The dip is a **mis-tuned-in-both-directions corridor** (44px cap = straight thread at strip-wide spans, #285; 104px cap + full row offset = ~106px over-bow into the terminal, 2026-08-15): it now hangs from the **STRIP's bottom edge** (fallback: lower tab bottom), capped at 64px, with NO per-row offsets stacked on top — the strip-bottom baseline is also what keeps a row-1 pair's arc from drawing through row 2's tab labels. ⚠️ **Colors are keyed on the SPAWNING tab, not per child**: every arc leaving one tab is the same color however many workers it spawns, so the strip reads as "these five came from w1, those two came from w2" — per-child coloring gave one tab's own children a different color each, which is the distinction the colors exist to make. A child that spawns in turn is a parent in its own right and gets its own color for the arcs below it, so a chain changes color at each generation while each generation's fan-out stays uniform. Assignment cycles `CodemanLineage.COLORS` in first-seen order per parent id (first entry empty = the skin-tuned `--session-blue`, so the first spawning tab keeps it; the rest vivid fixed hexes), memoized rather than derived from draw index (the SVG is wiped and rebuilt constantly, so an index-based color would flicker), and set inline as `--lineage-color` so styles.css keeps owning opacity/glow/dash. `test/session-lineage-lines.test.ts` drives the real `_appendLineageConnectionLines()` and asserts the painted property, since testing the color function alone would pass just as happily with the child id passed back in. ⚠️ **Desktop only**: the overlay is `z-index: 999` and the desktop header is 100 (arcs paint over it, which is what lets them touch tab bottoms), but under 1024px mobile.css makes the header `fixed; z-index: 1200` and would bury them. ⚠️ Paths carry `data-agent-id="lineage:"` because that is what `_applyLineEntrances()` queries — that one attribute is what gives them the entrance animation and its negative-`animation-delay` resume across `svg.innerHTML=''`. ⚠️ `.session-tabs` is `overflow-x: auto`, so a scrolled-out tab still HAS a rect (over the logo); edges with an endpoint outside the strip are skipped, and a passive `scroll` listener re-anchors the rest. +**Session lineage lines** (tab → tab it spawned, `sessionLineageLines`, per-device, desktop default ON): a create request may name its spawner via a `parentSessionId` body field or the `X-Codeman-Parent-Session` header; `resolveParentSessionId()` (route-helpers.ts) resolves it (exact id or unique ≥8-char prefix, live, visible, same owner) and ⚠️ anything unresolvable is DROPPED, never a 400. Rides `toState()`, no new SSE event. ⚠️ Rendering is a LAYER on the existing SVG pass (`_appendLineageConnectionLines` at the tail of `_updateConnectionLinesImmediate()`), geometry pure in `computeLineagePath()`: one U-bridge shape hanging from the strip bottom, colors keyed on the SPAWNING tab and memoized (never by draw index). ⚠️ Desktop only (z-index vs the fixed mobile header). ⚠️ Paths must keep `data-agent-id="lineage:"` (the entrance animation queries it); skip edges whose endpoint is scrolled out of the strip. → [architecture-invariants#session-lineage-lines-tab--tab-it-spawned](docs/architecture-invariants.md#session-lineage-lines-tab--tab-it-spawned) -**Auto-named sessions** (`autoNameSessions`, SYNCED, default OFF; #376): a placeholder tab (`w3-myapp`) takes its first real prompt as a title, in the `: ` form (`w3-myapp: fix the login redirect`) that `parseSessionPrefix()` (app.js, #232) already renders as the title alone with the prefix in the tooltip and that `_nextCaseSessionStartNumber()` still counts, so the case identity and the `w<n>` counter survive. Ownership is the tri-state `SessionState.nameSource`: `placeholder` (Codeman's own `w<n>-<case>` or no name, inferred by `isGeneratedSessionName()` when a persisted state predates the field), `auto` (titled once), `manual` (the `name` setter, i.e. `PUT /api/sessions/:id/name`, which auto-naming never touches again). ⚠️ **First prompt means the FIRST**: `applyAutoName()` flips a placeholder to `auto` whether or not the string changed, so a later "1" cannot rename the tab; a prompt that yields no title (`/clear`, a `!` shell escape) leaves the session eligible for the next one. ⚠️ **Only user-originated input counts** (`SessionWriteOptions.fromUser`, set by the browser WS path and `POST /api/sessions/:id/input` ONLY, so a forgotten flag on a new path fails toward not naming): Ralph kick-starts, respawn `/clear`s, cron launches, approval answers and the trust-dialog keys write through the same `write()`/`writeViaMux()` and used to name every Ralph tab "Read @ralph_prompt.md…". A `startMode: 'shell'` CLI never feeds the tracker (a capability, not an id check; a shell tab was renamed after every `ls`), and the send-key route's Shift+Enter line feed bypasses the session entirely, so it calls `trackUserInput()` or the two lines join with no separator. ⚠️ **The tracker (`session-auto-name.ts`, pure) sits on the raw keystroke stream**, so every key has an explicit rule: a bare Esc is resolved at the END of the chunk it arrives in (it used to stay in escape mode and eat the next prompt's first character, or a whole CJK prompt); SGR mouse reports, Tab, cursor keys and Shift+Tab leave the draft alone (a wheel tick mid-word used to drop the first half); Up/Down and Ctrl+P/N/R TAINT the draft so Enter submits nothing rather than a fragment; bracketed-paste newlines are newlines IN the composer, never Enter. The title is the first sentence past a minimum length ("e.g. fix this now" is not "e.g."), capped at 72 code points, and the composed name honours `MAX_SESSION_NAME_LENGTH`. The listener (`session-listener-wiring.ts`) checks eligibility BEFORE reading the setting, so an already-named session costs no settings read per prompt. Opt-in because the prompt lands in the tab name, `mux-sessions.json`, every `session:updated` and `/api/search` (Read My Mind keeps prompts 0600 for the same reason). Tests: `test/session-auto-name.test.ts`, `test/session-listener-wiring.test.ts`, `test/routes/session-name-routes.test.ts`. → [architecture-invariants#auto-named-sessions-first-prompt--tab-title](docs/architecture-invariants.md#auto-named-sessions-first-prompt--tab-title) +**Auto-named sessions** (`autoNameSessions`, SYNCED, default OFF): a placeholder tab (`w3-myapp`) takes its first real prompt as a title in the `<prefix>: <title>` form, so the case identity and `w<n>` counter survive. Ownership is `SessionState.nameSource` (`placeholder` | `auto` | `manual`; the `name` setter / `PUT /api/sessions/:id/name` makes it `manual`, never touched again). ⚠️ `applyAutoName()` flips to `auto` even if the string is unchanged, so only the FIRST titled prompt names the tab. ⚠️ Only user input counts: `SessionWriteOptions.fromUser` is set by the browser WS path and `POST /api/sessions/:id/input` ONLY; any new user-input path must set it (and the send-key Shift+Enter path must call `trackUserInput()`). ⚠️ The pure tracker (`session-auto-name.ts`) sits on the raw keystroke stream with an explicit rule per key; add a rule for any new key class. Tests: `test/session-auto-name.test.ts`. → [architecture-invariants#auto-named-sessions-first-prompt--tab-title](docs/architecture-invariants.md#auto-named-sessions-first-prompt--tab-title) **Maintainer bot (external)**: the Telegram bot that reviews open PRs and triages discussion threads in Codeman sessions used to live at `scripts/pr-bot/`. It moved OUT of this repository on 2026-09-14, to `~/codeman-cases/prbot/` (its own private git repo, systemd unit `codeman-pr-bot`, guide + agent rules in its own `README.md` and `CLAUDE.md`). It is a CLIENT of Codeman's HTTP API like any other, so nothing here depends on it and it is not part of the server, the CLI or the npm package. ⚠️ It spawns real sessions named `prbot-<n>` / `dscbot-<n>` on the local Codeman and holds clones under `~/.codeman/pr-bot/`, so those session names and that data dir are taken; it also fetches PR heads into `refs/pr-bot/*` of this checkout and must never check out, reset or clean it. The CHANGELOG entries for 1.25.0 and earlier still describe it, which is history rather than drift. -**Unified session list**: `GET /api/sessions/unified` merges live sessions, persisted state, lifecycle-log history, and transcript files into one deduped list (pure core in `src/services/unified-session-service.ts`). ⚠️ **Transcript history is THREE stores, not one**, because each CLI keeps its conversations in its own: Claude's `~/.claude/projects`, omp's `~/.omp/agent/sessions` and codex's `~/.codex/sessions` (#386). Rows fold into their owning session via the `claudeSessionId → Codeman id` alias map, so resumed and `/clear`-respawned sessions do not appear twice; that field is named for Claude and carries whatever id the CLI names its conversation with, which for every non-Claude row diverges from the Codeman id by construction. ⚠️ **`resumeId` is set by a SCANNER row only, never by a live session**, and that is what makes it safe to resume on: a row carrying one is a conversation already on disk, so `resumeHistorySession()` sends `codexConfig.resumeSessionId` and a row without one is a genuinely fresh session. Every surface that re-projects these rows has to carry the field through, the phone overview included, or a tap on that surface silently starts a second conversation. No terminal buffers in the response, unlike `/api/sessions`. Backs the Cmd+K Session Manager, plus pinning and cross-device tab order (`PUT /api/session-order`; pure merge helpers in `src/session-order.ts`, pushing device wins and server-only ids are never dropped). → [architecture-invariants#unified-session-list-and-session-manager](docs/architecture-invariants.md#unified-session-list-and-session-manager) +**Unified session list**: `GET /api/sessions/unified` merges live sessions, persisted state, lifecycle-log history and transcript files into one deduped list (pure core `src/services/unified-session-service.ts`), backing the Cmd+K Session Manager, pinning and cross-device tab order (`PUT /api/session-order`, `src/session-order.ts`). ⚠️ Transcript history is THREE stores (`~/.claude/projects`, `~/.omp/agent/sessions`, `~/.codex/sessions`), folded via the `claudeSessionId → Codeman id` alias map (not Claude-only despite the name). ⚠️ `resumeId` is set by a SCANNER row only, never a live session; every surface that re-projects these rows (phone overview included) must carry it through, or a tap silently starts a second conversation. → [architecture-invariants#unified-session-list-and-session-manager](docs/architecture-invariants.md#unified-session-list-and-session-manager) -**Owner tab layouts** (COD-359, `tab-layout*.ts` + `GET`/`PUT /api/tab-layout`): named tab GROUPS over the flat tab strip, scoped per owner (`SINGLE_USER_LAYOUT_OWNER` = `@single` when multi-user is off), persisted under the `tabLayouts` key in state.json. A layout is `{version, groups[], ungrouped[], updatedAt}` whose refs point at either a session or a saved webview (`TabRefKind`), capped at 32 groups / 512 refs. ⚠️ **BACKEND ONLY as of 1.24.1**: nothing in `src/web/public/` calls these routes yet, so a UI built on top is new frontend work, not a rewiring job. ⚠️ **`TabLayoutService` is the single mutation boundary** and every lifecycle caller (session created/removed, webview created/deleted, a legacy order PUT) describes ONE completed server action and gets AT MOST ONE versioned write; writing layout state from a route or a manager directly is what the service exists to prevent. ⚠️ The layout does not replace `PUT /api/session-order`, it PROJECTS onto it: `tab-layout-legacy-order.ts` is the pure translation both ways (`putLegacyOrder()` recomposes a global order from the owner's groups), so changing one side without the other silently desyncs the tab strip from the stored layout. ⚠️ **Reconciliation is gated on a SUCCESSFUL restore** (`markRestorationComplete` / `markRestorationFailed` / `markRestorationSkipped`, plus `assertDeletionReady()`): pruning refs against a session list that failed to load would delete live tabs, so a failed restore must leave the layout untouched. `PUT` takes exactly `{baseVersion, layout}` (any other key shape is a validation error), answers a stale `baseVersion` with the current layout rather than clobbering, and is capped at 128 KiB. Broadcasts `tab:layoutChanged`, owner-routed via `deriveTabLayoutSseHint`. +**Owner tab layouts** (`tab-layout*.ts` + `GET`/`PUT /api/tab-layout`): named tab GROUPS over the flat strip, scoped per owner (`@single` when multi-user is off), persisted as `tabLayouts` in state.json. BACKEND ONLY: no frontend calls these routes yet. ⚠️ `TabLayoutService` is the single mutation boundary (one completed server action = at most one versioned write); never write layout state from a route or manager directly. ⚠️ The layout PROJECTS onto `PUT /api/session-order` via `tab-layout-legacy-order.ts`; change both sides together. ⚠️ Reconciliation is gated on a SUCCESSFUL restore (`markRestorationComplete`/`assertDeletionReady()`): a failed restore must leave the layout untouched or live tabs get pruned. → [architecture-invariants#owner-tab-layouts](docs/architecture-invariants.md#owner-tab-layouts) -**Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `elicitation_complete`, `elicitation_response`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`, `prompt_submitted` (UserPromptSubmit, #367: a Claude pane reports its live conversation id first-hand). See `src/hooks-config.ts`; upstream hook semantics mirrored in `docs/claude-code-hooks-reference.md`. ⚠️ **Every claude session INSTALLS the hooks block into its workspace** (`applyWorkspaceHooks` in hooks-config.ts → `ensureCodemanHooks`, an add-only merge that keeps a user's own handlers), from EVERY claude create path — both interactive routes, cron fires, legacy scheduled runs, the plan-orchestrator one-shots — and from `restoreMuxSessions()` for sessions recovered on server start (that boot sweep skips a workspace that no longer exists, so a deleted repo with a surviving tmux session is never resurrected as an empty dir). Before 2026-08-15 hooks were written ONLY when Codeman created the case DIRECTORY, so a linked case / cloned repo — where most sessions actually run — had no hooks at all and every hook-driven surface was silently dead there: an AskUserQuestion dialog blocked the pane while the tab and the phone overview both read a calm `idle`, with no Approvals Inbox item, no push, no definitive `stop`/`idle_prompt` for respawn and no `stop`/`blocked` for the wait endpoints. The escape hatch is the synced `workspaceHooksEnabled` setting (App Settings → Agents & CLIs → Claude, **default ON**); OFF restores the old behavior, where a Codeman block that is already there is still refreshed when stale (COD-91) but one is never added. ⚠️ Route the decision through `applyWorkspaceHooks` rather than calling `ensureCodemanHooks` at a new site, or the setting silently stops applying to that path. ⚠️ Claude Code RE-READS `settings.local.json`, so an already-running session starts firing hooks without a restart (measured 2026-08-15) — and the notification for a blocking dialog is delayed by Claude Code (~30s), so the alert trails the dialog. ⚠️ An AskUserQuestion / plan-selection dialog arrives as **`permission_prompt`**, not `elicitation_dialog` (that one is MCP elicitation), so it renders as the RED "needs you" alert, not the yellow idle one. +**Hook events**: Claude Code hooks trigger via `/api/hook-event` (`permission_prompt`, `elicitation_dialog`, `elicitation_complete`, `elicitation_response`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`, `prompt_submitted`); see `src/hooks-config.ts` and `docs/claude-code-hooks-reference.md`. ⚠️ Every claude session installs the hooks block into its workspace (add-only merge) from every create path and from `restoreMuxSessions()`, gated by `workspaceHooksEnabled` (SYNCED, default ON). ⚠️ Route that decision through `applyWorkspaceHooks`, never call `ensureCodemanHooks` at a new site, or the setting silently stops applying. ⚠️ An AskUserQuestion / plan-selection dialog arrives as `permission_prompt` (RED alert), not `elicitation_dialog` (MCP elicitation). → [architecture-invariants#hook-events-and-workspace-hook-installation](docs/architecture-invariants.md#hook-events-and-workspace-hook-installation) -**Reboot restore** (#411/#442, `src/reboot-restore.ts` pure + `web/reboot-restore-registry.ts` + `routes/reboot-restore-routes.ts` + `reboot-restore-ui.js`): a host reboot takes the tmux server with it, so every pane dies and the board comes up empty with no explanation. At boot Codeman works out which sessions that reboot destroyed, holds the plan IN MEMORY (no new state file, and a server restart simply drops the offer), and the banner asks. ⚠️ **The heuristic decides whether to ASK, never whether to act**: two signals have to agree (the socket holds no panes at all while state still lists sessions, AND the host booted after the newest persisted activity), and a wrong yes costs one dismissable line rather than N CLI processes nobody asked for. ⚠️ Rebuilding is TAKE-then-build: entries leave the plan synchronously before the first `await` and the route is single-flighted per owner, so a double-click or a second device cannot put two panes on one conversation. Anything that never became a pane goes BACK on offer, with one deliberate exception, `already-live`, which unlike a missing workspace or a withdrawn grant cannot stop being true. ⚠️ Three things are re-checked at click time rather than trusted from boot (the owner's privilege grant, the workspace still being on disk, and the conversation not already being live), and the already-live sets are read FRESH per iteration rather than snapshotted: the loop awaits a real `startInteractive()` per entry, so a snapshot taken before it is tens of seconds stale by the tenth entry and would miss a conversation the user resumed by hand in that window. The confinement re-check is keyed on the entry's OWNER, never the caller, or an admin spending another user's entry is waved through by `isWorkingDirAllowed`. ⚠️ A rebuilt session comes back **attached, idle and disarmed**: the pane is NEW, so terminal scrollback is gone (the banner says so) while the conversation continues, respawn controllers and Ralph loops are never re-armed, and `reapplyPersistedSessionState(..., { rearmAutoResumeSchedule: false })` keeps auto-resume ENABLED but drops the pre-reboot `autoResumeAt` stamp, or one click has every restored session type `continue` into itself a minute later, unattended. That option exists only for this path; a Codeman restart still re-arms, because the limit footer will not reprint on its own. ⚠️ The rebuild passes `nameSource` through, or the constructor re-infers it from the name and a hand-renamed session shaped like `w<n>-<case>` comes back as `placeholder` for auto-naming to overwrite. ⚠️ Claude-mode only (others carry their conversation id in their own config object), and remote/docker sessions are never offered (`remote-or-docker`), because both need another host or container to be up. ⚠️ A failed rebuild is undone with `discardPartiallyBuiltSession()`, deliberately NOT `cleanupSession()`: the delete path would count the session's tokens into the lifetime totals, demote a pinned record to `stopped` (which this pass reads as an intentional kill, making the session permanently unrestorable) and recursively remove the WORKSPACE's `.claude-images`. ⚠️ `os.uptime()` reports the HOST's uptime, which a container shares, and that cuts both ways: after a genuine host reboot a containerized Codeman does see a short uptime and the banner works, but a container-only restart is invisible to it, which is the case where this would help most. Tests: `test/reboot-restore.test.ts`, `test/routes/reboot-restore-routes.test.ts`, `test/routes/reboot-restore-rebuild-failure.test.ts`, `test/discard-partially-built-session.test.ts`. +**Reboot restore** (`src/reboot-restore.ts` pure + `web/reboot-restore-registry.ts` + `routes/reboot-restore-routes.ts` + `reboot-restore-ui.js`): after a host reboot kills every pane, Codeman holds an IN-MEMORY plan of the destroyed sessions and a banner offers to rebuild them. ⚠️ The heuristic only decides whether to ASK. ⚠️ Rebuild is TAKE-then-build (entries leave the plan before the first `await`, single-flighted per owner). ⚠️ Re-check grant, workspace and already-live at click time, reading already-live FRESH per entry; key confinement on the entry's OWNER, never the caller. ⚠️ Rebuilt sessions come back disarmed: no respawn/Ralph, `rearmAutoResumeSchedule: false`, and pass `nameSource` through. ⚠️ Undo a failed rebuild with `discardPartiallyBuiltSession()`, NEVER `cleanupSession()`. Claude-mode only, never remote/docker. Tests: `test/reboot-restore.test.ts`. → [architecture-invariants#reboot-restore](docs/architecture-invariants.md#reboot-restore) -**Approvals Inbox** (cross-session queue of prompts waiting on a human; `approvalsInboxEnabled`, SYNCED, default OFF: every surface is opt-in; only the store and answer endpoints run regardless, so flipping it ON shows anything already pending): `web/approval-inbox.ts` is a `sessionWaits`-style singleton fed by `/api/hook-event`, holding at most ONE item per session (a new prompt supersedes), claude-mode only, in-memory. Cards are answered via `POST /api/approvals/:id/answer`, which sends a digit / Esc / idle-prompt text through `writeViaMux` (menu answers never carry `\r`). ⚠️ `option` digits are accepted ONLY when they match options parsed from the captured pane frame, and the answer path RE-CAPTURES the pane first (a dialog that no longer parses on screen means the keystroke would land in the composer, so refuse with 409). ⚠️ Resolution on the heuristic `working` signal ALONE is restricted to `idle` items; a permission/question item gets the pane-VERIFIED variant on that same signal (`resolveIfDialogGone()` → `verifyStillAnswerable()`), so the heuristic only decides when to LOOK and the screen decides the outcome. That is what clears a dialog answered in the terminal mid-turn; the other definitive signals are `stop`, `elicitation_complete`/`elicitation_response`, exit/delete, answer, supersede and the 12h TTL. ⚠️ **Viewing a session ACKNOWLEDGES its idle item, it does not resolve it** (`POST /api/approvals/session/:sessionId/viewed` → `acknowledgedAt` → `approval:updated`): the item stays pending (still answerable, still Read My Mind context) and only stops arming the yellow tab alert. That flag is what makes the clear durable, since the view-clears-idle rule used to live in one browser's memory and `seedApprovals()` re-armed the alert on the next reload while other devices never heard about it at all; the local half is `markIdleAlertSeen()` (app.js), called from BOTH `selectSession` paths, including the already-active early return, where a click could otherwise never clear the alert. ⚠️ **Only a HUMAN opening a session acknowledges**: `selectSession(id, { auto: true })` marks the three selections the APP makes (boot restore, a solo window opening its target, the fallback after the active session is closed) and skips the acknowledgement, so a page load cannot silently spend an alert the user never saw. The flag defaults to user-initiated, so an untagged call site fails toward acknowledging rather than toward an alert nothing can clear; `test/session-select-ack-gate.test.ts` pins both the gate and the tagged call sites. Idle-only by construction (`acknowledge()` defaults to `['idle']`): looking at a permission/question dialog does not answer it. ⚠️ Same rule on the input path: `_ackDelivery` (app.js) spends the IDLE alert only, via that same `markIdleAlertSeen()`. It used to `clearPendingHooks(sessionId)` with no kind, so one keystroke wiped a RED alert on that device while the dialog was still up, the other devices stayed red, and a reload re-seeded it. ⚠️ Claude Code fires no "permission answered" hook (only `elicitation_complete`/`elicitation_response`, i.e. the question flavor), so an answered-in-the-terminal dialog would otherwise sit pending until `stop`: `GET /api/approvals` therefore runs a **staleness sweep** over the caller's own items via `verifyStillAnswerable()`, which is deliberately the conservative check the answer path uses (only an item whose ORIGINAL frame parsed options can be dropped, so an unreadable capture keeps the alert rather than losing a live one). ⚠️ **`applyCapture()` is therefore ADD-ONLY for `options`**: a re-capture that parses nothing must never erase a parse an earlier one found. Claude Code delays the Notification hook behind the dialog (measured 6s, documented ~30s), so the 600ms re-capture routinely lands on a frame the user has ALREADY answered; clearing the field there made the item permanently unsweepable, because `verifyStillAnswerable()` reads a MISSING `options` as "we never could read this dialog" and keeps such items answerable by design. The red "needs you" then survived every sweep AND every page reload, went away only on `stop` (2026-08-20: a confirmed question left a tab flowing red for ~8 minutes while the turn ran on), and the stale card still accepted an answer, typing a bare `1` into a composer with no dialog under it. Pinned by `test/approval-inbox.test.ts`. ⚠️ A frame that parses no options is CONCLUSIVE in exactly two cases, and the second one closes the late-hook hole: the item once parsed options (they cannot vanish while the dialog is up), or the frame shows Claude actively running a turn. A modal dialog BLOCKS the turn, so the two cannot coexist — measured on v2.1.237, a live-dialog frame carries neither the `… (13s` timer NOR the `esc to interrupt` footer, which the dialog replaces with `Enter to select · ↑/↓ to navigate · Esc to cancel`. Anything else stays answerable, so an unreadable capture still keeps the alert. That second signal is reached by a delayed staleness pass (`STALE_CHECK_DELAY_MS`, 3s) scheduled alongside the re-capture, because a prompt answered BEFORE the hook lands creates an item whose FIRST capture already has no dialog in it: nothing ever parsed, `stop` may have fired already, and the alert then outlived reloads until the 12h TTL. ⚠️ That pass must stay comfortably LATER than `RECAPTURE_DELAY_MS`, whose whole reason for existing is that the hook can beat Ink to the screen — resolving inside the paint window would clear the alert for a dialog that was about to appear. The frontend seeds from `GET /api/approvals` in `handleInit` **regardless of the setting**: the seed re-arms the tab-alert state machine (`setPendingHook`) unconditionally, and only populating `this.approvals` (the inbox surfaces) is gated — seeding used to be gated wholesale, which left a reloaded page with NO red tab while a permission dialog sat blocking a session (2026-08-15); `_onApprovalResolved` clears the pending-hook alert unconditionally for the same reason. ⚠️ The red/yellow tab alert itself is a STEADY border/background/dot with a pulse on top: the original keyframes swung to transparent at 0%/100%, so half of every cycle looked like a normal tab. Push Approve/Deny buttons stay gated on the setting (`sendPushNotifications` strips `actions`/`approvalId` when OFF) and are answered from `sw.js` directly so they work with no tab open. Surfaces (all gated on the setting): header bell (marker-hidden until count > 0, phones never show it) + drawer (`approvals-ui.js`), phone overview NEEDS YOU answer strips (`mobile-overview.js`). Design: `docs/approvals-inbox-plan.md`. +**Approvals Inbox** (`approvalsInboxEnabled`, SYNCED, default OFF; the store and answer endpoints run regardless): `web/approval-inbox.ts` is an in-memory, claude-only queue fed by `/api/hook-event`, at most ONE item per session, answered via `POST /api/approvals/:id/answer` through `writeViaMux` (menu answers never carry `\r`). ⚠️ Accept `option` digits ONLY if they match options parsed from a fresh RE-CAPTURE of the pane; a dialog no longer on screen is a 409. ⚠️ Resolve permission/question items only via the pane-verified `verifyStillAnswerable()`; the heuristic `working` signal alone may resolve `idle` items only. ⚠️ `applyCapture()` is ADD-ONLY for `options`. ⚠️ Viewing ACKNOWLEDGES an idle item (never resolves it) and only a human selection does: app-made selections pass `selectSession(id, { auto: true })`; `_ackDelivery` spends the IDLE alert only. ⚠️ `handleInit` seeds tab alerts from `GET /api/approvals` REGARDLESS of the setting. → [architecture-invariants#approvals-inbox](docs/architecture-invariants.md#approvals-inbox) -**Read My Mind intent profiles** (phase 1 of `docs/readmymind-plan.md`; `readMyMindEnabled`, SYNCED, default OFF): per-CASE profiles (user-stated `goals` + the user's recent real prompts), keyed by owner + realpath(workingDir) so they survive `/clear`/respawns and multi-user scoping is structural. Capture rides the transcript (`transcript:user_prompt` from `transcript-watcher.ts`), NOT the input paths: `POST /input` sees only programmatic prompts and the WS channel is raw keystrokes. The listener lives inside `startTranscriptWatcher()`'s `if (!watcher)` block (outside it would duplicate per hook event) and is claude-only + gated on the setting per event. Store: `src/intent-store.ts` singleton, `intents.json` written 0600 tmp+rename (prompts can contain secrets; never fed to `/api/search`). Endpoints: GET/PUT/DELETE `/api/sessions/:id/intent` + POST `/api/sessions/:id/readmymind` (`readmymind-routes.ts`, ownership via `findSessionOrFail` WITH `req`; registrations stay the bare `app.<method>('path')` shape, the endpoints.md drift scanner cannot see generics). **Phase 2 (predictor + 🧠 button)**: `readmymind-context.ts` is the PURE budgeted assembler (9 ranked sources, drop order siblings→away→workspace→tools, sections 1-4 truncate only); IO lives in `readmymind-collectors.ts` (transcript TAIL read — the live watcher keeps only a 500-char snippet — + git signals, skipped for remote-SSH cases) and the route; `readmymind-predictor.ts` reuses the AiCheckerBase spawn mechanics standalone (verdict-shaped base vs freeform JSON) as a mutable singleton routes call and tests stub. Claude-mode only (400), one in flight per session (409 CONFLICT), model = `readMyMindModel` setting defaulting to `AI_CHECK_MODEL` (opus, decided). Frontend `readmymind-ui.js`: header 🧠 marker-hidden (`btn-readmymind--hidden`) until the setting is ON; phones hide it in mobile.css and get a keyboard-accessory 🧠 key instead (ships in BOTH bar templates, revealed by the `rmm-enabled` class on the BAR element — setMode() rebuilds button innerHTML, so per-key state would be wiped; synced at init + every `applyHeaderVisibilitySettings()`). Alternate suggestions render as tappable rows that swap into the editable field without losing edits; Rethink rejects the whole shown set and carries the optional steer note (`#readMyMindSteer`, sent as `steer`, shown in ready + empty-result phases, cleared on each open). Suggestions render via value/`textContent` ONLY and Send/Insert go through `POST /input` (server-side, so the sendEnterKey/local-echo trap does not apply) — nothing auto-sends, ever. User guide: `docs/readmymind.md`. +**Read My Mind intent profiles** (`readMyMindEnabled`, SYNCED, default OFF; `docs/readmymind-plan.md`): per-CASE profiles (goals + recent prompts) keyed by owner + realpath(workingDir), captured from the transcript (`transcript:user_prompt`), never the input paths; the listener must stay inside `startTranscriptWatcher()`'s `if (!watcher)` block. Store `src/intent-store.ts` → `intents.json`, ⚠️ written 0600 tmp+rename and never fed to `/api/search` (prompts carry secrets). Routes in `readmymind-routes.ts` (ownership via `findSessionOrFail` WITH `req`); predictor = pure `readmymind-context.ts` + IO in `readmymind-collectors.ts` + `readmymind-predictor.ts`, claude-only, one in flight per session (409). ⚠️ Suggestions render via value/`textContent` ONLY and nothing auto-sends, ever. Frontend `readmymind-ui.js`. → [architecture-invariants#read-my-mind-intent-profiles](docs/architecture-invariants.md#read-my-mind-intent-profiles) -**Voice dictation via Claude** (`claudeVoiceEnabled`, SYNCED, default OFF): the mic button can transcribe through this machine's Claude Code login instead of a Deepgram key, using the same speech-to-text service the CLI's own `/voice` mode uses. ⚠️ **Claude Code's voice mode itself is unusable here**: it opens the HOST's microphone (`sox`/`arecord`), and the CLI runs in a headless tmux pane while the human is in a browser elsewhere. So Codeman captures in the browser and borrows only the backend. Audio goes browser → Codeman → Anthropic (`src/web/voice-stream.ts`): the OAuth token never reaches the page, and the browser only sends PCM and receives text. ⚠️ Credentials are **read-only** (`src/claude-credentials.ts`) and Codeman never refreshes them — a refresh rotates the refresh token and could sign the user out of their own CLI; an elapsed token reports `expired` instead. ⚠️ Capture MUST be linear16/16 kHz/mono, so it uses an **AudioWorklet**, not MediaRecorder (which cannot emit raw PCM); `voice-pcm-worklet.js` is fetched from JS, so it is invisible to `cacheBustAssets` and borrows voice-input.js's `?v=` token — **edit the two together**. ⚠️ Transcript frames carry the WHOLE running transcript, not deltas: the Claude path replaces where the Deepgram path appends. Provider choice is `voiceSettings.provider` (`auto` prefers Claude → Deepgram → Web Speech). → `docs/claude-voice-plan.md` +**Voice dictation via Claude** (`claudeVoiceEnabled`, SYNCED, default OFF): the mic transcribes through this machine's Claude Code login (the CLI `/voice` backend) instead of Deepgram; the browser captures, `src/web/voice-stream.ts` relays to Anthropic. ⚠️ The OAuth token never reaches the page. ⚠️ Credentials are READ-ONLY (`src/claude-credentials.ts`): never refresh them (it rotates the refresh token and can sign the user out of their CLI). ⚠️ Capture must be linear16/16 kHz/mono via an AudioWorklet; `voice-pcm-worklet.js` borrows voice-input.js's `?v=` token, so edit the two together. ⚠️ Claude transcript frames are cumulative: replace, never append. → [architecture-invariants#voice-dictation-via-claude](docs/architecture-invariants.md#voice-dictation-via-claude) **Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`. **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) -**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set); Shell selection and automatic drop recovery always use a bounded 1 MiB `?tail=` window. Shell loads the rest only when **Load full history** is pressed; ordinary scrolling must not trigger a multi-megabyte reset+replay on xterm's main thread. Other modes may re-pull at the TOP (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). Live writes are one-chunk-in-flight, released by xterm's parse callback, so xterm's private queue cannot bypass the browser's 128 KiB render cap. ⚠️ **A frame dropped at that cap MUST be recovered, and the recovery must verify itself** (`_scheduleDroppedOutputRecovery`, app.js): a hole in a TUI byte stream is a desynced cursor, which is muffled text (#464). It was a fire-and-forget 2s timer that nulled its own handle and then called `_onSessionNeedsRefresh()` — whose early returns (a buffer load in flight, a refresh already owning the session) are MOST likely to be true during exactly the burst that caused the drop, so the recovery was lost silently and the bytes were never replayed. `_onSessionNeedsRefresh` now returns whether it actually repainted, and the scheduler re-arms while it has not, bounded by `DROP_RECOVERY_MAX_ATTEMPTS` because the early returns it retries past are transient contention. A refresh that died at the capture fetch DEADLINE (it returns `'deadline'`) is a stalled link, not contention, and is NOT retried: each retry would be another `?full=1` capture waiting out a deadline of up to two minutes. The same debounce still collapses a burst into one attempt, and the `TERMINAL DROP` crash-trail line sits behind it, since one line per dropped frame evicted the whole 50-entry trail in under a second. Tests: `test/dropped-output-recovery.test.ts`, whose retry case and no-retry case only pin the fix as a pair. While WebSocket owns terminal I/O, duplicate SSE terminal events are dropped before JSON parsing, and recovery is single-flight per active session. ⚠️ **A `full=1` capture ENDS with a cursor move back to the pane's own caret position**, counted UP from the last replayed row — without it the caret stays where the last character landed, which for an agent CLI is the status line, and every cursor-relative update the CLI sends afterwards is measured from the wrong row. The move is relative, not `CUP`: absolute row addressing is only right while the browser's rows equal the pane's, and `resizeWindow` does not wait for tmux, so a capture can be taken before a requested resize applies. That makes row alignment load-bearing on this path: no transform that can DELETE A LINE may run over the capture, so it keeps its trailing blank rows and skips redraw-bloat stripping, the banner trim and the leading-whitespace strip. ⚠️ Those three skips key on whether a capture actually CAME BACK (`isFullCapture`), never on `?full=1` alone — the fallback to the byte history is a stream of successive frames that must still be stripped, and a session with no mux takes it on every load. A capture holding nothing visible returns '' so the byte history survives instead of a blank screen replacing it. ⚠️ A full re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. ⚠️ **A visible capture now REPORTS the geometry it was taken at** (`captureCols`/`captureRows`, #435), because a frame built for a pane taller or wider than the browser is damaged two ways at once (overflow rows clamp onto the last line; a narrower browser wraps every painted row) and nothing in the response used to say so. Both fields are ABSENT when no frame was positioned, so every consumer tests `Number.isFinite`, never truthiness: a `display-message` cursor query that fails makes `capturePaneBuffer` return the raw capture while the route still labels it `mux-visible`. The comparison runs on `mux-visible` ONLY, the replay is capped at one attempt, and a pane that cannot be sized to fit latches in `_geometryRetryUseless` so it is diagnosed once per session rather than on every tab switch. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) +**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and Shell loads the rest only via **Load full history**, never on ordinary scroll. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) **Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions) -**Terminal touch gestures: link taps and text selection**: on a touch device xterm's own handlers see neither — `touch-action: none` plus touchstart's preventDefault suppress the browser's compatibility mouse events, `_installMobileTapMouseGuard` drops the trusted ones that still arrive, and the synthetic `mousedown`/`mouseup` pair dispatched for mouse REPORTING goes to the `.xterm` root, an ANCESTOR of the screen element the linkifier and SelectionService listen on. So both gestures are driven explicitly. ⚠️ **A tap activates the link under it** through the SAME provider that feeds the hover linkifier (`_terminalLinkAtPoint`, containment mirroring xterm's `_linkAtPosition`), synchronously inside `touchend` — that is what keeps the user gesture `window.open` needs — and BEFORE any mouse report, mirroring `_handleDesktopTerminalClick`'s skip for a hovered link. Two rows keep their meaning: the caret's logical line (`_tapIsOnCaretLine`, where a tap places the cursor in text the USER typed) and TUI-owned rows (`_isActionableMobileTerminalTap`, answering a dialog). ⚠️ The caret line is the boundary rather than the tap INTENT, because a shell classifies every tap as `'input'` and gating on that would leave every URL in shell output inert. ⚠️ **Long-press selects** by driving xterm's public `select()` (renderer-independent — under WebGL the glyphs are pixels and native selection cannot exist), drag or a further tap extends, and Copy goes through `copyTerminalSelection()` for its execCommand fallback on plain-HTTP installs. Three guards are load-bearing and each came from a real phone: the compat mouse pair after `touchend` (xterm focuses on mousedown and SelectionService resets the model there, so the keyboard sprang up and the selection vanished on lift), the platform's own ~500ms long-press (Android Chrome focuses the nearest editable element — the helper textarea — through no event a handler can preventDefault, so a bounded focus guard blurs it and `contextmenu` is suppressed for the gesture window), and `copyTerminalSelection()`'s closing `terminal.focus()` (right on desktop, wrong on a phone). Tests: `test/terminal-touch-tap.test.ts`. +**Terminal touch gestures: link taps and text selection**: on touch devices xterm's linkifier and SelectionService never see the gesture, so both are driven explicitly (terminal-ui.js). ⚠️ A tap activates the link under it through the SAME provider as the hover linkifier (`_terminalLinkAtPoint`), synchronously inside `touchend` (keeps the user gesture `window.open` needs) and BEFORE any mouse report; the caret's logical line (`_tapIsOnCaretLine`) and TUI-owned rows (`_isActionableMobileTerminalTap`) keep their meaning. ⚠️ Gate on the caret line, never on tap intent (a shell calls every tap `'input'`). ⚠️ Long-press selects via xterm's public `select()`; keep the three guards: suppress the compat mouse pair after `touchend`, the bounded focus guard + `contextmenu` suppression for the platform long-press, and no closing `terminal.focus()` on phones. Tests: `test/terminal-touch-tap.test.ts`. → [architecture-invariants#terminal-touch-gestures-link-taps-and-text-selection](docs/architecture-invariants.md#terminal-touch-gestures-link-taps-and-text-selection) -**Auto Copy (copy-on-select)** (`autoCopySelection`, per-device, default OFF): a finished terminal selection lands on the clipboard with no keystroke. ⚠️ It fires at the END of a gesture, never in `onSelectionChange` (that callback runs per cell crossed, so copying there is one clipboard write per mouse move); it only ARMS `_autoCopyPending`, and a document-level `mouseup` listener flushes. ⚠️ The flush is SYNCHRONOUS inside the handler because both clipboard paths need user activation (Firefox gates `navigator.clipboard.writeText` on it, and the plain-HTTP `execCommand` fallback must run in the gesture's own task); a timer or a wait for `onSelectionChange` loses it, invisibly in Chrome. ⚠️ Touch needs its OWN calls from `_endTouchSelectionGesture()`/`_selectTouchSelectionLine()`: that path `preventDefault()`s its touchend, so no mouseup ever arrives and the toggle would be dead on phones. ⚠️ Unlike `copyTerminalSelection()` it must NOT clear the selection (the text would vanish under the cursor that highlighted it) and must NOT focus the terminal (that opens the on-screen keyboard over it); focus is RESTORED to whatever held it, which only matters for the `execCommand` fallback. Guards are pure in `decideAutoCopy()` (constants.js): off, blank/whitespace-only, and a 1M-char cap (an autoscrolling drag can sweep the whole 50k-line scrollback), refused rather than truncated with a toast pointing at Ctrl+C. Silent on success except once per page load; failures toast, throttled 10s. Tests: `test/terminal-auto-copy.test.ts`. +**Auto Copy (copy-on-select)** (`autoCopySelection`, per-device, default OFF): a finished terminal selection lands on the clipboard with no keystroke. ⚠️ Copy at the END of a gesture, never in `onSelectionChange` (per-cell); it only arms `_autoCopyPending` and a document-level `mouseup` flushes. ⚠️ The flush must be SYNCHRONOUS in the handler (both clipboard paths need user activation); never defer it to a timer. ⚠️ Touch needs its own calls from `_endTouchSelectionGesture()`/`_selectTouchSelectionLine()` (no mouseup arrives). ⚠️ Unlike `copyTerminalSelection()`, never clear the selection or focus the terminal; restore prior focus. Guards are pure in `decideAutoCopy()` (constants.js, 1M-char cap, refused not truncated). Tests: `test/terminal-auto-copy.test.ts`. → [architecture-invariants#auto-copy-copy-on-select](docs/architecture-invariants.md#auto-copy-copy-on-select) -**Ctrl+V paste trap** (`image-input.js`): `Ctrl+V` routes through `_handleImagePaste()`, which appends a hidden `contenteditable` trap, focuses it, and reads the clipboard out of the paste event that lands there. Images upload and their saved paths are typed into the session; text goes through `terminal.paste()` so bracketed-paste markers survive. ⚠️ **The trap must consume exactly ONE paste event.** Two routes deliver one for a single keypress and Firefox fires both: `document.execCommand('paste')` dispatches a trusted event and still returns `false`, because the trap cancels it, while Chromium refuses that command and dispatches nothing; separately, the keydown's own default action delivers a paste to the now-focused trap, because returning `false` from the custom key handler never cancels the DOM event (see smart copy above). Measured on a live install: Firefox two events per keypress, Chromium and WebKit one. Handling both wrote the clipboard to the PTY twice, and right-click → Paste stayed correct because it carries no keydown. ⚠️ Removing the `execCommand('paste')` call would also end the doubling, and all three engines still deliver one event without it, but it stays for the mobile engines a desktop measurement cannot reach: where a browser aims the key's default action at the element focused when the keydown began, the command is the only route into the trap, and the trap is the only place image blobs are read. The one-shot flag lives on the trap rather than on a browser check, so any count produces one insert. Tests: `test/image-paste-trap.test.ts`. → [architecture-invariants#terminal-paste-ctrlv](docs/architecture-invariants.md#terminal-paste-ctrlv) +**Ctrl+V paste trap** (`image-input.js`): `Ctrl+V` routes through `_handleImagePaste()`, which focuses a hidden `contenteditable` trap and reads the clipboard from the paste event landing there; images upload and their paths are typed in, text goes through `terminal.paste()` so bracketed-paste markers survive. ⚠️ **The trap must consume exactly ONE paste event** (Firefox delivers two per keypress: the `execCommand('paste')` event and the keydown's default action); the one-shot flag lives on the trap, never on a browser check. ⚠️ Do not remove the `execCommand('paste')` call: on some mobile engines it is the only route into the trap, and the trap is the only place image blobs are read. Tests: `test/image-paste-trap.test.ts`. → [architecture-invariants#terminal-paste-ctrlv](docs/architecture-invariants.md#terminal-paste-ctrlv) -**Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity/omp get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for **claude ≥ 2.1.187 ONLY** at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. ⚠️ Codex was in that list and must never go back without a fresh measurement: codex-cli 0.147.0 ignores SGR wheel reports entirely (`mouse_any_flag=0`, inline viewport, transcript pushed into terminal scrollback), so forwarding produced a dead wheel (#227 follow-up). `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). ⚠️ When that gate is FALSE on a claude session whose local buffer is hollow (`baseY === 0`), the gesture becomes coalesced PageUp/PageDown key sends (`_maybePageCliTranscript`) instead of a no-op; ⚠️ and `getClaudeCliVersion()` must never cache a FAILED probe (one timeout used to disable forwarding process-wide until restart). ⚠️ **A click is hand-reported to the CLI only while the CLI actually has mouse tracking on.** The full strip removes the mouse DECSETs, so xterm's `mouseTrackingMode` is permanently `none` there and the browser hand-encodes SGR reports (`_sendSyntheticSgrTap`); without state it did that on EVERY click, so a stripped-mode pane running a plain shell (CLI exited, or a shell started inside a claude-mode session) received reports it never asked for and printed them as literal text (`[<0;88;20M`), garbling the next typed line. `_recordStrippedMouseMode()` (session.ts) records what the strip removes, `toState()` publishes `cliMouseTracking`, and `_shouldReportMouseToCli()` gates all three report sites on it. Only 1000/1001/1002/1003 count (1005/1006 are encodings, 1007 is alt-scroll), and the change broadcasts UNdebounced since a dialog can be clicked inside the 500ms window. `_logScrollRouting()` prints the routing decision and its inputs once per session — read it before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding) -**Detached start + service install** (issue #231): `codeman web -d` relaunches the SAME entry script with `detached:true` (setsid), so there is no controlling terminal and no shell job entry. ⚠️ `nohup` is NOT what makes this work: Node re-arms SIGHUP to its default disposition even when it inherits "ignore", and `cli.ts` handles SIGHUP with a graceful shutdown, so a delivered HUP still stops the server. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile check + `/api/status` probe): a second instance on the shared tmux socket attaches PTYs to the first one's live sessions. ⚠️ Neither may report success it has not observed — the parent polls `/api/status` until the child answers or dies, since `launchctl load` and a clean spawn are both silent about a server that starts and immediately exits. `--stop` verifies the pid still LOOKS like a Codeman server (`ps -o command=`) before signalling, because pids get recycled. Unit/label names live in `config/service-names.ts` so install.sh, `detectSupervisor()` and `service install` cannot drift into supervising two copies; they are instance-scoped, and identical to the historical names for the default instance. `service install` bakes the installing shell's PATH into the unit (launchd gives a job `/usr/bin:/bin:/usr/sbin:/sbin`, which finds neither a Homebrew/nvm `node` nor `tmux`/`claude`) and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install) +**Terminal scrollback strip + wheel/touch forwarding**: codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity/omp get a NARROW strip (alt-screen toggles only). ⚠️ Gated on `useMux`: direct-PTY sessions must keep the alt screen. Wheel and touch forward to the CLI for **claude ≥ 2.1.187 ONLY**; ⚠️ never re-add codex without a fresh measurement (it ignores SGR wheel reports). ⚠️ `getClaudeCliVersion()` must never cache a FAILED probe. ⚠️ Hand-report clicks only while the CLI has mouse tracking on: `_shouldReportMouseToCli()` gates all three report sites on `cliMouseTracking` (from `_recordStrippedMouseMode()`, session.ts), or a plain shell prints the reports as literal text. Read `_logScrollRouting()` before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding) +**Detached start + service install**: `codeman web -d` relaunches the same entry script `detached:true` (setsid); `nohup` is not what makes it survive. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile + `/api/status` probe), or a second instance attaches to the first one's live sessions. ⚠️ Never report success not observed: poll `/api/status` until the child answers or dies. `--stop` must verify the pid still looks like Codeman (`ps -o command=`) before signalling. Unit/label names live only in `config/service-names.ts`. `service install` bakes the installing shell's PATH into the unit and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install) -**Self-update** (App Settings → System → Updates): in-app updater for git-clone installs supervised by systemd/launchd (`systemd`, `launchd`, `launchd-daemon`, `docker-compose`, else `none` → "restart manually"). The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` that outlives the restart and writes progress to `update-status.json`, which the browser polls across the connection drop. `src/web/self-update.ts` splits pure helpers (unit-tested) from IO wrappers. npm installs report as non-updatable. ⚠️ **The Compose deployment is the one supervisor that does NOT outlive the restart**: there the restart IS the container exiting (`restart: unless-stopped` relaunches it), which kills the script too — safe only because the terminal `restarting` marker is written BEFORE the kill, so nothing may be appended after it. Two config facts make it work at all and both are load-bearing: the repo is a HOST BIND MOUNT over `/opt/codeman` (a pull into the baked image copy would land in the writable layer and be silently discarded by the next `up`), and the runtime image keeps devDependencies + a build toolchain (`npm run build` is tsc+esbuild, and node-pty has no Linux prebuild), which is why `npm prune --omit=dev` is gone and the updater passes `--include=dev` against `NODE_ENV=production`. ⚠️ An in-place container update applies CODE ONLY — a restart reuses the existing image and config — so `evaluateEnvironmentGate()` REFUSES a release that changes `server.Dockerfile`/`docker-compose.yaml` (sha256 vs the baseline `Start-Codeman.sh` writes to `docker-env-applied.json` on every start) or adds `.env.example` keys the user's `.env` lacks, and refuses when the restart policy would not bring the container back. That third check exists because **Compose resolves an unset `${VAR}` to the EMPTY STRING and starts anyway**, so a new required setting otherwise arrives as a silently blank env var. Every unknown fails OPEN in the gate (no baseline, unreadable `.env`, no socket): failing closed would permanently block containers created before the fingerprint file existed. ⚠️ The KILL does not: the server exits only when `--restart-by-exit 1` was passed, i.e. the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (set ONLY there, since that file is what sets `restart: unless-stopped`; the image ENV deliberately does not) or the daemon reported an auto-restart policy; otherwise the build lands as `completed-needs-manual-restart`, because exiting blind takes a `docker run` container with no restart policy down with no UI left to recover it. The gate is re-evaluated on `POST /api/system/update`, so hiding the button is UX, not the control. ⚠️ The four global agent CLIs in `server.Dockerfile` are PINNED on purpose — unpinned, a user's CLI versions are a function of when their image was built rather than of any commit, which is the one environment change no diff-derived gate can see; pinning turns it into a Dockerfile change the gate already catches. `test/docker-compose-env-parity.test.ts` is the merge-side guard (every compose `${VAR}` ↔ an `.env.example` entry). → [docs/docker-self-update.md](docs/docker-self-update.md), [architecture-invariants#self-update](docs/architecture-invariants.md#self-update) +**Self-update** (App Settings → System → Updates): in-app updater for git-clone installs under a supervisor (`systemd`, `launchd`, `launchd-daemon`, `docker-compose`, else `none`). The work runs in a DETACHED `scripts/self-update.sh` writing `update-status.json`, polled across the restart; pure helpers in `src/web/self-update.ts`. ⚠️ Compose: the restart kills the script, so nothing may be appended after the `restarting` marker; the repo must stay a host bind mount over `/opt/codeman` and the image must keep devDependencies + toolchain. ⚠️ `evaluateEnvironmentGate()` refuses releases that change `server.Dockerfile`/`docker-compose.yaml` or add `.env.example` keys, re-evaluated on `POST /api/system/update`; unknowns fail OPEN, but the exit-to-restart needs `--restart-by-exit 1` (`CODEMAN_RESTART_BY_EXIT=1` only in the Compose file). ⚠️ Keep the agent CLIs in `server.Dockerfile` pinned. → [docs/docker-self-update.md](docs/docker-self-update.md), [architecture-invariants#self-update](docs/architecture-invariants.md#self-update) -**Reverse-proxy base path** (`--base-url` / `CODEMAN_BASE_URL`, default `/`; `src/config/base-path.ts` is the pure single-source, normalized to `''` for root or `/foo`): lets Codeman be mounted under a sub-path behind a proxy that **forwards the prefix unchanged** (does NOT strip it). Deliberately few choke points, mirrored ingress/egress: **(server ingress)** `stripBasePath()` runs inside Fastify's `rewriteUrl` so ALL routes stay declared prefix-agnostic (`/api/...`, `/ws/...`) — and a request arriving WITHOUT the prefix (hooks, health checks, docker bridge, all hitting the raw port) is left untouched, so the server answers at both; **(server egress)** one `onSend` hook prepends the base to every root-absolute `Location` header, covering all redirects; **(HTML)** `renderIndexHtml` rewrites the shipped `<base href="/">` to the mount and injects `window.__CODEMAN_BASE__` — the template's asset refs are all RELATIVE so `<base>` handles them for free; **(frontend runtime URLs)** root-absolute URLs ignore `<base>`, so `CodemanBase.url()` (constants.js) is the route builder, applied transparently by a `fetch` wrapper and explicitly at the few EventSource/WebSocket/`window.open`/`<img|iframe|a>`-src sites; **(sw.js/manifest)** the worker derives its base from `self.location`, the manifest uses relative `start_url`/`scope`; **(web-tab proxy)** `proxyPrefixFor(cap, basePath)` is the single base-aware root that cascades to the injected `<base>`, root-absolute HTML rewrites, the `runtimeUrlShim`, `Set-Cookie` Path and `Location` rebasing — while the INGRESS parsers (`capabilityFromProxyPath`, `resolveUpstreamUrl`) stay base-agnostic because `rewriteUrl` strips the prefix before routing, and `capabilityFromReferer(referer, basePath)` strips it from the browser-supplied Referer. ⚠️ `--base-url` rides the daemon relaunch via `buildWebArgs` and the service unit via `resolveServicePlan`. Pure helpers unit-tested in `test/base-path.test.ts` + `test/webview-proxy.test.ts`; HTML injection in `test/render-index-html.test.ts`. +**Reverse-proxy base path** (`--base-url` / `CODEMAN_BASE_URL`, default `/`; pure single source `src/config/base-path.ts`, normalized to `''` or `/foo`): mounts Codeman under a sub-path behind a proxy that forwards the prefix unchanged. Few choke points: `stripBasePath()` in Fastify's `rewriteUrl` (routes stay prefix-agnostic; unprefixed requests still answer), one `onSend` hook rebasing `Location`, `renderIndexHtml` rewriting `<base href>` + injecting `window.__CODEMAN_BASE__`, and `CodemanBase.url()` (constants.js) for runtime URLs. ⚠️ Keep template asset refs RELATIVE, and route every root-absolute frontend URL (EventSource/WebSocket/`window.open`/src) through `CodemanBase.url()`. ⚠️ Web-tab proxy egress goes through `proxyPrefixFor(cap, basePath)`; ingress parsers stay base-agnostic. ⚠️ `--base-url` must ride `buildWebArgs` and `resolveServicePlan`. Tests: `test/base-path.test.ts`. → [architecture-invariants#reverse-proxy-base-path](docs/architecture-invariants.md#reverse-proxy-base-path) **Attachments** (live external document references; all wiring in `file-routes.ts`): a **registry** maps a stable `attachmentId` to a realpath-resolved, extension-allowlisted absolute path, so browser requests never carry arbitrary absolute paths. ⚠️ The **magic-link scanner** (`codeman://attach?...` in terminal output) is **prompt-injectable**, so its scan path is force-confined to the session workspace; a hostile prompt could otherwise exfiltrate arbitrary host files over SSE. The security gate is an extension **allowlist**, not a blocklist. `document-conversion-limiter.ts` caps converter spawns globally: without it, N large docs detected at once fork N multi-minute processes, which is a resource-exhaustion vector. → [architecture-invariants#attachments](docs/architecture-invariants.md#attachments) -**File-path links (terminal + chat)**: a path an agent prints is clickable on BOTH surfaces and opens the file-preview overlay. ⚠️ ONE pattern (`FILE_PATH_LINK_PATTERN` / `absoluteFilePathPattern()` in constants.js) feeds the xterm link provider AND the response viewer's `_linkifyFilePaths()`; a fresh instance per call, since `lastIndex` is per-object state. The chat linkifier walks TEXT NODES with DOM APIs (the source is model output; never rebuild sanitized markup as a string) and skips subtrees already inside an `<a>`. ⚠️ **An out-of-workspace path is served through the ATTACHMENT routes, not the file routes** — `file-content`/`file-raw` are workspace-confined and 404 exactly the paths agents print most (a `/tmp` capture, Claude's scratchpad), so `openFilePreview()` registers such a path via `POST /api/sessions/:id/attachments` with **`notify: false`** (suppresses only the `attachment:detected` broadcast — same guard, same routes; without it every click also popped a card announcing the file already on screen) and renders by id. The click is an explicit action on the explicit, Origin-guarded route, which is what distinguishes it from the force-confined magic-link scanner. ⚠️ **Media extensions are single-sourced** (`VIDEO_ATTACHMENT_EXTENSIONS`/`AUDIO_ATTACHMENT_EXTENSIONS` in `attachment-registry.ts`, imported by `file-content`'s classification) so a clip plays the same in or out of the workspace; a player needs all THREE of allowlist + a real `MIME_TYPES` entry (octet-stream renders a dead player) + the range-aware body. ⚠️ **`TEXT_ATTACHMENT_EXTENSIONS` IS `EDITABLE_EXTENSIONS`** (never a second list): if the viewer would edit it inside the workspace, it can be read outside. Widening READ must never widen RUN, so `html`/`htm` joined `svg` in `serveRawFile`'s download-only branch, other text goes out as inert `text/plain`+`nosniff`, and `~/.codeman*/state.json` joined `isSensitivePath` (it persists `envOverrides`, which can hold `GEMINI_API_KEY`). ⚠️ The terminal sends an **out-of-workspace** path to the preview instead of the log viewer (that one spawns `tail -f` and reaches only workspace + `/var/log` + `~/logs`); in-workspace text keeps the tail viewer and `file-stream-manager`'s allowlist is untouched. The image-watcher keeps its own narrow detection list, so none of this cards every file an agent writes. → [architecture-invariants#file-path-links-terminal--response-viewer](docs/architecture-invariants.md#file-path-links-terminal--response-viewer) +**File-path links (terminal + chat)**: a path an agent prints is clickable on BOTH surfaces and opens the file-preview overlay. ⚠️ ONE pattern (`FILE_PATH_LINK_PATTERN` / `absoluteFilePathPattern()` in constants.js) feeds the xterm link provider AND `_linkifyFilePaths()`, a fresh instance per call (`lastIndex`). The chat linkifier walks TEXT NODES with DOM APIs, never rebuilds sanitized markup as a string. ⚠️ An out-of-workspace path goes through the ATTACHMENT routes (`POST /api/sessions/:id/attachments` with `notify: false`), never by widening `file-content`/`file-raw` or `file-stream-manager`'s `tail -f` allowlist. ⚠️ `TEXT_ATTACHMENT_EXTENSIONS` IS `EDITABLE_EXTENSIONS` (never a second list), and widening READ must never widen RUN: `html`/`htm`/`svg` stay download-only, other text is inert `text/plain`+`nosniff`. Media extensions are single-sourced in `attachment-registry.ts`. → [architecture-invariants#file-path-links-terminal--response-viewer](docs/architecture-invariants.md#file-path-links-terminal--response-viewer) **Filesystem path picker** (Link Existing "Browse" + the mobile keyboard's `📁 Path` key): lazy one-directory browsing via `GET /api/filesystem/browse`, with `GET /api/filesystem/preview` for the tapped file. Inserts the path **without** Enter, so the prompt is never submitted; the sibling `⌫ All` key clears only the unsent prompt and must never send the agent's `/clear`. ⚠️ This is a **second file-serving surface and inherits neither the attachment confinement nor its ownership scoping** — it allowlists Home, `CASES_DIR`, `/mnt/d` and `CODEMAN_FILE_PICKER_ROOTS`, blocks sensitive trees, and rejects symlink escapes **after** `realpath`. ⚠️ The optional `sessionId` is an ownership boundary that must be `canAccessOwned`-checked by hand (it does not go through `findSessionOrFail`), and in multi-user mode a non-admin gets only their own `userSpacePath` as a root: per-user spaces live INSIDE `homedir()`, so a `Home` root exposes every other user's workspace. Previews go through the same global conversion limiter, and Markdown/TXT/JSON are served as inert `text/plain`. → [architecture-invariants#filesystem-path-picker](docs/architecture-invariants.md#filesystem-path-picker) **File Viewer edit mode** (issue #212): the file-preview overlay edits workspace text files in place — `GET .../file-content?edit=1` + `PUT /api/sessions/:id/file-content`, policy in `src/config/file-editing.ts`. This is a **third file surface and the only one that WRITES**: read-path confinement (realpath + workspace + ownership) plus sensitive/blocked/`.git` denies and an extension **allowlist**; writes are `wx`-temp + rename (no `O_CREAT` anywhere = edit-in-place is structural); optimistic concurrency via sha256 `baseHash` → 409. ⚠️ `edit=1` never truncates and the client must never save a plain-preview buffer (the 500-line truncation would silently delete the rest). ⚠️ CRLF/UTF-8 guards: EOL re-applied server-side, non-UTF-8 refused via round-trip compare. → [architecture-invariants#file-viewer-edit-mode](docs/architecture-invariants.md#file-viewer-edit-mode), `docs/file-viewer-edit-plan.md` -**Files panel search** (COD-236, the `q` param on `GET /api/sessions/:id/files`): `compileFileQuery()` (`utils/file-query.ts`, pure, no IO, so it unit-tests directly) compiles the query into a reusable predicate, which is what lets the server-side walk prune instead of streaming the whole tree. ⚠️ **A query turns that endpoint into a FLAT match list rather than a nested tree**, and the walk deliberately recurses past non-matching directories, since the whole point of searching is to reach a file whose ancestors do not match. An empty or whitespace-only query compiles to `null`, which is what keeps the default tree response byte-identical when no search is requested. ⚠️ **Globs are never compiled into a RegExp**: `*a*a*a…` translated to `^.*a.*a.*a…$` is a classic backtracking blowup evaluated synchronously against every walked path, so one pathological query would freeze the event loop for the whole server (the same reason `search-service.ts` is regex-free). `globMatch()` is a two-pointer wildcard walk instead, O(text · pattern) with both operands short by construction, and an overlong query (`MAX_QUERY_LENGTH`, 256) also compiles to `null` rather than running. A query containing `/` matches the relative path, otherwise the bare entry name; globs match anchored and case-insensitively (`*` spans any run, slashes included, `?` exactly one character), everything else is a plain case-insensitive substring. +**Files panel search** (COD-236, the `q` param on `GET /api/sessions/:id/files`): `compileFileQuery()` (`utils/file-query.ts`, pure) compiles the query into a predicate the server-side walk prunes with; a query returns a FLAT match list and the walk recurses past non-matching directories. An empty, whitespace-only or overlong (`MAX_QUERY_LENGTH`, 256) query compiles to `null`, keeping the default tree response byte-identical. ⚠️ **Never compile a glob into a RegExp** (`*a*a*a…` backtracks and freezes the event loop for the whole server): `globMatch()` is a two-pointer wildcard walk. → [architecture-invariants#files-panel-search](docs/architecture-invariants.md#files-panel-search) -**Raw file bodies are streamed and range-aware**: `file-raw`, the attachments `/raw` route and `GET /api/download` always advertise `Accept-Ranges: bytes` and answer a `Range` header with `206` + `Content-Range` (single-range only; parser is pure + unit-tested in `src/web/http-range.ts`, a malformed spec is ignored → 200 while an out-of-bounds one is a 416). ⚠️ **The size cap on all three is a sanity bound, not memory protection** (`MAX_FILE_DOWNLOAD_BYTES` in `config/buffer-limits.ts`, default 2GB, env `CODEMAN_MAX_DOWNLOAD_BYTES`, `0` = unlimited): the bodies stream, so size costs a read stream and not RSS (measured: a 600MB download moved peak RSS by ~37MB). Its predecessor was a hardcoded 50MB whose comment still said "prevent memory exhaustion" long after the `readFile()` it described was replaced by `sendFileBody()`, so all it did was refuse legitimate downloads of build artifacts, videos and archives. `/api/download` was the last route that really did buffer the whole file, and now shares `sendFileBody()` with the other two. ⚠️ A 200-only response is what made the File Viewer's `<video>` unseekable: Chrome then reports `video.seekable` as `[0, 0]`, the scrub bar is inert and `currentTime = x` silently reverts (measured on an 18MB mp4), and Safari refuses to start the media at all. ⚠️ These bodies go out through `reply.hijack()`, which bypasses Fastify's status handling — `sendRawStream` must copy the status onto `reply.raw` by hand or a partial body ships labelled `200` and the browser treats a slice as the whole file. ⚠️ Closing the preview must **pause and unload** the media (`_stopFilePreviewMedia` in panels-ui.js): dropping the overlay's `visible` class is `display:none` and nothing else, and a DETACHED `HTMLMediaElement` keeps playing, which is how the X button used to leave a video audible with no player to pause. +**Raw file bodies are streamed and range-aware**: `file-raw`, the attachments `/raw` route and `GET /api/download` share `sendFileBody()`, advertise `Accept-Ranges: bytes` and answer `Range` with `206` + `Content-Range` (single-range, parser in `src/web/http-range.ts`); without it `<video>` cannot seek. The size cap (`MAX_FILE_DOWNLOAD_BYTES`, default 2GB, env `CODEMAN_MAX_DOWNLOAD_BYTES`, `0` = unlimited) is a sanity bound, not memory protection; never reintroduce a whole-file buffer. ⚠️ Bodies go out via `reply.hijack()`, so `sendRawStream` must copy the status onto `reply.raw` by hand or a partial body ships as `200`. ⚠️ Closing the preview must pause and unload media (`_stopFilePreviewMedia`), since a detached `HTMLMediaElement` keeps playing. → [architecture-invariants#raw-file-bodies-streamed-and-range-aware](docs/architecture-invariants.md#raw-file-bodies-streamed-and-range-aware) **Ultracode / workflow-run visualization** (opt-in, default OFF): the Workflow tool writes a completion artifact only at run *end*, so live in-flight runs exist solely as transcript dirs. `workflow-run-watcher.ts` therefore synthesizes ACTIVE runs from transcripts until the completion artifact appears and supersedes them. It is **STANDALONE** and deliberately never imports or touches `subagent-watcher.ts`, despite reading the same tree. Two independent toggles: `showUltracodeAgents` (docked panel) and `ultracodeFloatingWindows` (floating windows); the watcher starts if **either** is on. → [architecture-invariants#ultracode--workflow-run-visualization](docs/architecture-invariants.md#ultracode-and-workflow-run-visualization) -**Clone a repository as a case** (issue #236, Add Case → **Clone Repo**): `POST /api/cases/clone` clones a public repo into the caller's case space synchronously (request held open, bounded by `GIT_CLONE_TIMEOUT_MS`, no job store); `POST /api/cases/clone-preflight` reports whether the URL can be cloned anonymously plus its real branches/tags. Core in `src/git-clone.ts`. ⚠️ **The URL is a code-execution surface**: `ext::sh -c <cmd>` (and ANY `<name>::<payload>` helper) makes git run a command, so every `::` form is refused, a leading `-` is refused, and every spawn is an argv array with `--` before the operands. ⚠️ **Non-interactive or the open request hangs** — `gitNonInteractiveEnv()` closes the terminal/askpass/ssh/GCM prompt paths; `HOME`/`PATH` stay inherited, so a user's OWN credential helper may authenticate (Codeman still never collects or stores credentials, and refuses a `user:password@` URL). ⚠️ Timeout kills the process GROUP (clone fans out into child processes), the destination is removed only if this attempt created it, and repository contents win over scaffolding (existing `CLAUDE.md` kept, hooks merged, repo-shipped `.claude/settings*` reported as a warning since its hooks run locally). The **Brain** picker sets the toolbar run mode on success. → [architecture-invariants#clone-a-repository-as-a-case](docs/architecture-invariants.md#clone-a-repository-as-a-case) +**Clone a repository as a case** (issue #236, Add Case → **Clone Repo**): `POST /api/cases/clone` clones synchronously into the caller's case space (bounded by `GIT_CLONE_TIMEOUT_MS`, no job store); `POST /api/cases/clone-preflight` checks anonymous cloneability and lists refs. Core in `src/git-clone.ts`. ⚠️ **The URL is a code-execution surface**: refuse every `::` form and a leading `-`, spawn only argv arrays with `--` before operands. ⚠️ Stay non-interactive (`gitNonInteractiveEnv()`) or the open request hangs; never collect credentials, refuse `user:password@` URLs. ⚠️ Timeout kills the process GROUP, remove the destination only if this attempt created it, and repo contents win over scaffolding (existing `CLAUDE.md` kept, hooks merged, repo `.claude/settings*` warned about). → [architecture-invariants#clone-a-repository-as-a-case](docs/architecture-invariants.md#clone-a-repository-as-a-case) **Cross-session search**: `GET /api/search` federates an in-memory search over session metadata, run-summary events, and attachment-history entries. The pure core `searchSources()` does substring matching with hard per-type caps: **no regex (so no ReDoS) and no filesystem reads (so no traversal)**. The server-private `externalPath` is never read. PAST sessions (#261) come from `session-history-index.ts`, a capped snapshot of the unified list filled **outside** the request path (`/api/sessions/unified` publishes it; a stale one is rebuilt fire-and-forget), that indirection is what keeps the no-fs property. ⚠️ The snapshot is stored UNSCOPED with a per-row owner and MUST be re-filtered through `canAccessOwned()` on read; history rows carry `jumpTo.kind:'resume-session'`, since a closed session has no tab to select. → [architecture-invariants#cross-session-search](docs/architecture-invariants.md#cross-session-search) -**Web tabs** (dashboard URLs as tabs): a saved URL renders as a tab beside agent sessions. **NOT a `SessionMode` of its own** (no PTY, no tmux, no respawn), same reasoning that keeps Docker/remote-SSH as case overlays. Dashboards are **proxied through Codeman's own origin** by default, because a direct iframe fails three ways at once: prod is HTTPS so `http://` targets are blocked as mixed content, many dashboards send `X-Frame-Options: DENY`, and our own `default-src 'self'` CSP blocks cross-origin frames. Proxying leaves the prod CSP unchanged (`/webview/...` is `'self'`). ⚠️ The proxy is **NOT an API surface**: it authenticates on an in-memory capability in the path and is correspondingly exempt from the cookie + Origin checks; that exemption is fenced to safe methods and non-`/api` paths and is pinned by `test/webview-auth-exemption.test.ts`. ⚠️ Iframes omit `allow-same-origin` unless a dashboard is explicitly marked `trusted`, and `Authorization`/`codeman_session` are stripped upstream in **both** modes so `CODEMAN_PASSWORD` cannot leak. ⚠️ A sandboxed frame is **opaque-origin**, which breaks two things `curl` can never reproduce: its runtime-built root-absolute URLs escape `<base>` (fixed by an injected `runtimeUrlShim()`), and its same-host `fetch`/XHR are CORS-checked with `Origin: null` (fixed by `buildProxyCorsHeaders()` plus exempting the proxy from the global `OPTIONS`-204 short-circuit in `registerSecurityHeaders`). Both present as the dashboard's own "Failed to fetch" while the page renders fine. ⚠️ **Egress guard**: link-local and cloud-metadata targets (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`, the Azure/Alibaba fixed addresses, `metadata.google.internal`) are refused at save time AND on the RESOLVED address at connect time (`webview-egress-policy.ts`, pure, plus `webview-egress.ts`: a `lookup` hook on the undici Agent behind `webviewFetch()` and on the `ws` client). An IP literal never reaches a lookup hook (`net.connect` skips DNS for it), so the synchronous hostname check at each connect site is NOT redundant. Loopback and RFC1918 stay allowed on purpose: a `localhost` Grafana is the feature. The proxy uses the `undici` PACKAGE's own `fetch` + `Agent`, never Node's global fetch with a foreign dispatcher (Node bundles its own copy; a protocol mismatch fails silently). ⚠️ **A loopback link in agent output opens as a proxied web tab** (`openLinkThroughWebTabIfLoopback` in webview-tabs.js, called from the terminal link provider and the response viewer): a phone cannot reach the box's `localhost:5173`, so the tap routes through Codeman's origin instead, reusing a saved same-origin dashboard or saving one under `host:port`. Two rules that look cosmetic and are not. **`*.localhost` is deliberately NOT in the auto-route set** even though it IS loopback to a browser: the link source is agent-written terminal output and prompt-injectable, every other member of that set is an address literal that can only mean this box, and a `*.localhost` DNS name is not one (a resolver with a search domain retries `evil.localhost` as `evil.localhost.<search domain>`, which an attacker can control, turning agent output plus one tap into a persisted server-side fetch of an agent-chosen origin). A user who really runs `api.localhost` saves it by hand, which is an explicit action. The PAGE-side test is deliberately broader (`isOnBoxHostname`), since a false positive there only declines to proxy. And **`this.webviews` being set does NOT mean it is loaded**: `initWebviews()` assigns a truthy EMPTY map synchronously and only then awaits the list, so a tap during page load must join the in-flight refresh (`_webviewsRefresh`/`_webviewsLoaded`) or it finds nothing to reuse and POSTs a duplicate record for an origin that already exists. Capabilities are revoked on logout / admin logout / user deletion (`revokeOwner`, which had NO caller for two releases while its docstring said otherwise), and proxied responses carry `Referrer-Policy: same-origin` so a dashboard cannot hand the capability-bearing URL to a third party. → [architecture-invariants#web-tabs](docs/architecture-invariants.md#web-tabs), `docs/web-tabs.md` +**Web tabs** (dashboard URLs as tabs): a saved URL renders as a tab beside sessions, **NOT a `SessionMode`**, and is proxied through Codeman's own origin (`/webview/<cap>/...`). ⚠️ The proxy is NOT an API surface: its capability-based auth exemption stays fenced to safe methods and non-route paths (`test/webview-auth-exemption.test.ts`). ⚠️ Iframes omit `allow-same-origin` unless `trusted`, and `Authorization`/`codeman_session` are stripped upstream in both modes. ⚠️ **Egress guard**: link-local and cloud-metadata targets are refused at save time, by a sync hostname check at each connect (IP literals skip DNS), AND on the resolved address (`webview-egress.ts`); use the `undici` package's own `fetch` + `Agent`, never Node's global fetch. ⚠️ Loopback links in agent output auto-open as proxied web tabs (`openLinkThroughWebTabIfLoopback`), but never auto-route `*.localhost` (prompt-injectable DNS). Capabilities are revoked on logout (`revokeOwner`). → [architecture-invariants#web-tabs](docs/architecture-invariants.md#web-tabs), `docs/web-tabs.md` **Multi-user mode** (opt-in `--multiuser` / `CODEMAN_MULTIUSER=1`, OFF by default): named users with scrypt-hashed passwords in `~/.codeman/users.json`. Gated everywhere by `isMultiUserMode()`; when OFF, behavior is byte-identical to single-user because every scoping helper short-circuits. ⚠️ **Not a security boundary at the agent layer**: every session still runs as the SAME OS account. This separates WORKSPACES; it does not sandbox users (Docker cases are the isolation story). Ownership threads through `Session.owner` and is enforced in `findSessionOrFail`, list endpoints, SSE routing (fail-closed), WS, search, and file-preview. → [architecture-invariants#multi-user-mode](docs/architecture-invariants.md#multi-user-mode), `docs/multi-user-plan.md` @@ -310,53 +314,53 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `i18n.js`(1.5) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `terminal-keycode229-recovery.js`(5.55) → `sanitize-html.js`(5.6) → `app.js`(6) → `tab-rail-resize.js`(6.5) → `terminal-ui.js`(7) → `terminal-split.js`(7.5) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `cron-ui.js`(9.7) → `settings-ui.js`(10) → `panels-ui.js`(11) → `readmymind-ui.js`(11.3) → `ultracode-panel.js`(11.5) → `approvals-ui.js`(11.6) → `reboot-restore-ui.js`(11.65) → `admin-ui.js`(11.7) → `session-ui.js`(12) → `host-wake-ui.js`(12.2) → `webview-tabs.js`(12.5) → `mobile-overview.js`(12.55) → `home-sessions.js`(12.56) → `entrance-animations.js`(12.6) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15) → `ultracode-windows.js`(15.5) → `session-lineage.js`(15.6) → `image-input.js`(16). `i18n.js` translates static + newly inserted application DOM while skipping terminal/response/file/user-name surfaces; `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData). `terminal-keycode229-recovery.js` forwards a committed `input` event that xterm's `_inputEvent` guard drops (Chrome-on-Android soft keyboards send `composed: true` after a keydown), and only when xterm emitted no canonical data for that keystroke. ⚠️ **That decision is settled at the NEXT keydown as well as on its own zero-delay timer** (#441): the drain runs from xterm's custom key handler, which fires BEFORE xterm processes that key, so a soft keyboard that commits the last character and sends Enter in one InputConnection transaction puts the character on the wire ahead of the `\r`. On the timer alone that character is not merely late, it is LOST: xterm emits the `\r` first and bumps the canonical counter past the candidate's snapshot, so the candidate stands down (measured, `hell\r` where the user typed `hello`). The trade is that a keydown decides with less evidence than the timer did, since xterm's own keyCode-229 rescue has not run yet; that is safe for Enter, which clears the textarea so the pending diff emits nothing. Ordering is pinned by `test/terminal-keycode229-recovery.browser.test.ts`, which the CI gate does NOT run. -**Entrance animations** (`entrance-animations.js`, all OFF by default): opt-in animations for the four things that appear when work starts, chosen per surface via `data-tab-anim` / `data-term-anim` / `data-win-anim` / `data-line-anim` on `<html>`. Defaults are the `legacy` theme, so an untouched install behaves exactly as before and every hook short-circuits on its first line. ⚠️ Tabs and connection lines are **destroyed mid-animation** on every re-render (`_fullRenderSessionTabs()` replaces the strip's innerHTML; `_updateConnectionLinesImmediate()` does `svg.innerHTML = ''`), so both are tracked by id and re-applied to the fresh element with a **negative `animation-delay`** to resume rather than restart. ⚠️ The terminal-pane styles may animate **transform / opacity / clip-path only**, xterm's FitAddon derives rows+cols from `getComputedStyle(parent).width/height`, so animating width/height/padding there would resize the PTY; `test/entrance-animations.test.ts` pins that property allowlist, plus the rule→keyframes→theme-option chain a style silently does nothing without. ⚠️ **`blur` is the ONE style that puts a `filter` on the terminal container**, against the standing rule, because every alternative was measured against a live xterm and does not work: a `backdrop-filter` veil on `::before` blurs perfectly while STATIC and Chrome silently drops the backdrop the moment ANY animation runs on that pseudo-element (the veil computes `blur(15.3px)` and the text behind it stays razor sharp), and driving the radius from rAF buys the same full-screen blur per frame plus main-thread work. The cost the rule exists to avoid is inherent to blurring a terminal, so the style buys it knowingly: opt-in, OFF by default, one ~520ms run per session open, class straight back off, `will-change` still unset. Worst-case price, headless SwiftShader with no GPU: frame deltas 16.7ms → 33.3ms for the run, against 16.7ms flat for `fade`. Do not generalise it — a second filtered terminal style needs its own measurement. ⚠️ The `blur` connection line animates `filter` too, so both kinds of line hold their glow in **`--line-glow`** and both of its keyframes say `blur(N) var(--line-glow)`: the function lists then match and interpolate, instead of the glow vanishing for the run and popping back (a lineage line's glow is a different colour entirely, set per element). Its 100% frame deliberately omits `opacity` so the endpoint comes from the element's own resting value — 0.9 subagent, 0.72 lineage, 0.95 working — which is what `line-enter-fade`'s hardcoded 0.9 gets wrong. ⚠️ Window styles other than `beam` transform the window, which moves the rect its connection line is aimed at; `beam` deliberately animates opacity/filter only so its line can draw toward a stable target. Persisted to its own `codeman:*Anim` localStorage keys (per-device, deliberately NOT in the `.strict()` `SettingsUpdateSchema`); picker in App Settings → Appearance, full per-surface lab at `?animlab=1`. +**Entrance animations** (`entrance-animations.js`, all OFF by default): opt-in animations for tabs, terminal, windows and connection lines, chosen via `data-tab-anim` / `data-term-anim` / `data-win-anim` / `data-line-anim` on `<html>`; the default `legacy` theme short-circuits every hook. ⚠️ Tabs and lines are destroyed mid-animation on re-render, so re-apply to the fresh element by id with a negative `animation-delay` (resume, never restart). ⚠️ Terminal-pane styles may animate only transform / opacity / clip-path (anything else resizes the PTY via FitAddon); `blur` is the ONE sanctioned `filter` exception, do not generalise it. ⚠️ Line glow lives in `--line-glow` so blur keyframes interpolate. Persisted per-device in `codeman:*Anim` localStorage keys, never in `SettingsUpdateSchema`; lab at `?animlab=1`. Test: `test/entrance-animations.test.ts`. → [architecture-invariants#entrance-animations](docs/architecture-invariants.md#entrance-animations) -**Mobile tab strip scrolling** (issue #257): under 768px the tab strip is a horizontal scroller (desktop wraps to a second row instead), so the active tab can sit off-screen. Three rules keep it reachable and they only work together: `_updateActiveTabImmediate()` scrolls the selected tab into view via `computeTabScrollLeft()` (pure, in constants.js) using **rect math on the strip's own `scrollLeft`**, never `scrollIntoView()`, which would also scroll the document under a fixed header; `_fullRenderSessionTabs()` **restores `scrollLeft`** across the `innerHTML` rebuild, since ambient rebuilds (a task badge appearing, a session created elsewhere) otherwise snap a mid-swipe strip back to 0; and it re-reveals the active tab **only when it changed** (`_lastRenderedActiveTabId`), so browsing the far end of the strip is not undone by background renders. ⚠️ **The ACTIVE tab is the only one with action icons, and on a phone they can eat it**: `.session-tab.active .tab-name` reserves `min-width: 44px` in the phone block (under 600px), because a short session name rendered a 13px label against a 50px gear+close cluster, putting the tab's geometric CENTRE on the gear, so a thumb aiming at the tab opened Session Options instead of switching (measured at 360/393/430px; only long names cleared it). ⚠️ **The floor is set by the 10th tab onward, not by the tabs you can see**: `.tab-number` renders only for `_tabIdx < 9`, so tab 10 loses 16px + a gap off its left and its centre sits 10px further right. The centre clears the icons when `reserved > icons + rightEdge - leftRunUp - gap` (= 50 + 9 - 17 - 4 = **38px**), hit-testing snaps to whole pixels so 39px still lands on the gear, and the practical floor is 40px — a NUMBERED tab clears it at 20px, which is exactly why reasoning from the tabs on screen would put the centre back on the gear. `test/mobile-tab-tap-zones.test.ts` recomputes that inequality from the stylesheet, so widening the gear or the padding fails there rather than on a phone. The guarantee is centre-off-the-ICONS, not centre-inside-the-label (on a numberless tab it lands in the gap between them, which still switches). Non-active tabs keep their icons hidden and stay tappable end to end. ⚠️ Mobile no longer hoists the active session to the front of the strip: that reordering ran on full renders only, so tab order flipped depending on which render path fired, and it renumbered the Alt+N badges. Scroll-into-view replaces it; do not reintroduce it. +**Mobile tab strip scrolling** (issue #257): under 768px the tab strip scrolls horizontally, so the active tab must be kept reachable. `_updateActiveTabImmediate()` reveals it via `computeTabScrollLeft()` (constants.js, rect math on the strip's own `scrollLeft`, never `scrollIntoView()`, which scrolls the document under the fixed header); `_fullRenderSessionTabs()` must restore `scrollLeft` across rebuilds and re-reveal only when the active tab changed (`_lastRenderedActiveTabId`). ⚠️ The phone-block `min-width` on `.session-tab.active .tab-name` keeps the tab's centre off the gear/close icons, sized for numberless tabs 10+ (floor 40px); do not shrink it. ⚠️ Never reintroduce hoisting the active session to the front of the strip. Test: `test/mobile-tab-tap-zones.test.ts`. → [architecture-invariants#mobile-tab-strip-scrolling](docs/architecture-invariants.md#mobile-tab-strip-scrolling) -**Session list layout: header strip or left sidebar** (`sessionListLayout`, App Settings → Appearance → Tabs, default `header`; per-device policy — it IS in `SettingsUpdateSchema` and persists server-side, but `displayKeys` makes a device keep its own value): with many sessions the horizontal strip stops being scannable, so the list can move into a vertical `<aside>` with a filter box and a live count, collapsible to a 44px rail (`--sidebar-width` 260 / `--sidebar-width-collapsed` 44) via **Alt+B** (`toggleSessionSidebar`; Alt, not Ctrl+B, which must reach tmux/readline in the terminal). ⚠️ **There is ONE `#sessionTabs` element and it is MOVED between hosts** (`#sessionTabsHost` in the header, `#sessionSidebarList` in the aside, `#tabRail` for the vertical rail below), never a second list — so every render path, drag-reorder handler and Alt+N index keeps working unchanged. Exactly TWO functions reparent it and they must run in this order: `applySessionListLayout()` first (sidebar wins), then `applyTabOrientation()` (settings-ui.js), which moves the tabs into `#tabRail` only when the sidebar does not own them. ⚠️ It sets `data-session-list` / `data-sidebar` on `<html>` and must run BEFORE `applyTabWrapSettings()`, which is the one owner of `tabs-two-rows`/`tabs-show-folder` and reads those attributes. ⚠️ **The vertical tab rail** (`tabOrientation`/`tabRailWidth`/`sessionSidebarFontSize`, all per-device display keys that ARE in the schema, like `sessionListLayout`) is a SECOND vertical list next to the sidebar: the orientation setting is silently ignored while the sidebar layout is chosen, desktop/tablet only (`resolveTabOrientation` forces horizontal on mobile), resizable via `tab-rail-resize.js` (which owns terminal refits during the drag). ⚠️ **Detailed rows are a property of a vertical LIST, not of one surface** (`sessionListLayout: 'sidebar-rich'` for the sidebar, `tabRailDetail: 'rich'|'simple'` for the rail, rail default **rich**): both draw the home screen's per-session line (`created 3d ago · working 12m`) plus a status pill, from the SAME row model (`_sidebarRichRow`/`_sidebarRichMetaHTML` in app.js, classified by `_mobileOverviewState`/`_mobileOverviewSince`), and the render paths ask ONE gate, `isRichTabRows()` (= `isSessionSidebarRich() || isTabRailRich()`). ⚠️ Detail rides on its own attribute (`data-sidebar-detail` / `data-tab-rail-detail`) so every existing `[data-session-list="sidebar"]` / `[data-tab-orientation='vertical']` rule keeps matching both variants untouched; a flip of detail ALONE still needs a full render (the stamps line is emitted by the row template, not toggled by CSS) and must re-run `applyTabWrapSettings()`, which owns the folder line and is now rail-aware. ⚠️ The rich CSS rules carry a rail twin as a COMMA-GROUPED selector, never `:is()` (an `:is()` list takes its most specific argument, which would lift the sidebar arm from (0,3,1) to the rail's (0,5,1)). ⚠️ Width is the whole reason there are thresholds: the rich sidebar is 300px (`--sidebar-width-rich`) and a rail that has never been sized defaults to **320** (`RICH_DEFAULT_WIDTH`, the existing Wide preset) instead of 256, because at 256 the stamps line ellipsizes mid-word; a user-narrowed rail drops the created stamp below 288 (`tab-rail-tight`, CSS only) and drops rich rows entirely below 240 (`tab-rail-compact`, which re-renders). A stored width is never overridden. ⚠️ Only detailed rows carry stamps that go stale with no event behind them, so `_startSidebarRichClock()` (20s, rewrites text in place — a re-render would restart every row's animation) must be armed and disarmed by BOTH `applySessionListLayout()` and `applyTabOrientation()`. ⚠️ **Axis decisions must use `_isVerticalTabList()`** (sidebar OR rail), never `isSessionSidebarActive()` alone: the rail leaves `data-session-list` at `header`, and the sidebar-only predicate shipped four rail bugs at once (drag insertion side read from clientX, active tab never scrolled into view, floating windows anchored below tabs instead of beside them, connector redraws skipped on rail scroll). The pre-paint script stamps `data-tab-orientation` (+ `--tab-rail-width`) like it stamps the sidebar keys, or vertical mode flashes through the header strip; the name font size defaults to 12px, the sidebar's historical size, so untouched installs are never restyled. ⚠️ Leaving sidebar mode **clears `_sidebarFilter`**: the filter box only exists in the aside, so a stale filter would hide sessions from the header strip with no reachable control to clear it. ⚠️ On handhelds the aside is an off-canvas overlay rather than a docked rail, and a closed drawer keeps `display: flex`, so it is marked `inert` + `aria-hidden` (`_isSessionSidebarOverlay()`) or its filter box and ~4 tab stops per session stay in the tab order; the DOCKED desktop rail must never be inerted, its rows are still clickable. The desktop home rail (`home-sessions.js`) defers to it, since both dock the session list flush left. ⚠️ **The detailed RAIL additionally wears the home rail's CARD, and the detailed SIDEBAR deliberately does not**: the rail is an occasional, resizable list you scan, the sidebar is a permanently-docked nav column where 20 stacked cards read as a wall — so the card rules are RAIL-SCOPED and were NOT added to the comma-grouped selectors above, which is the one-line change that would silently restyle the sidebar. Card state accents reuse `home-sessions-blink-red`/`-yellow` rather than declaring a second pair, and every state dot rule excludes `.tab-alert-action`/`.tab-alert-idle` by hand, because those alert rules are only (0,3,0) and the rail's are (0,5,1)+. There is no `.active` rule in that block on purpose: `.session-tab.active` already paints border/background/box-shadow `!important`. ⚠️ **Row ORDER is a third rail attribute** (`tabRailSort: 'activity'|'manual'`, `data-tab-rail-sort`, default **activity**): a sorted rail answers the home screens' question with the home screens' answer, `CodemanSessionOrder` over rows classified by `_mobileOverviewState` (`_tabRailSortOrder` in app.js). ⚠️ **It is applied as the flex `order` property, never by reordering the DOM**: `#sessionTabs` stays in `sessionOrder`, so the Alt+N badge (`_tabIdx`, which therefore does NOT run 1,2,3 down a sorted rail — it names a shortcut, not a position), drag-and-drop, the arrow-key walk, the sidebar filter and `_scrollActiveTabIntoView()` all keep reading the list they always read, and a session changing state moves ONE inline style instead of forcing the full rebuild that would restart every card's animation on every SSE tick. The incremental render path therefore has to re-apply it (a state flip adds no tab, so the full rebuild is never reached) and an empty string is what clears it when sorting stops. ⚠️ Web tabs are pinned past the cards by a CSS `order: 9999` rather than an inline one, since `renderWebviewTabs()` emits the same markup for every layout; the flex default of 0 would interleave them among the sorted sessions. ⚠️ `setupTabDragHandlers()` sets `draggable="false"` and returns while sorting is on: the drop rewrites `sessionOrder` correctly and the sort then puts the card straight back, so the affordance would be a lie — `'manual'` is the way back to drag-reordering. ⚠️ **The arrow-key walk is the one place that must follow the SORT rather than the DOM**: `_tabKeydownHandler` steps `querySelectorAll` order, which is `sessionOrder`, so on a sorted rail ArrowDown from the top card landed wherever that session sat in the tab order instead of on the card below it. It now sorts its node list by the COMPUTED `order` first (computed, not inline: web tabs get their `order: 9999` from CSS and would otherwise read as 0 and lead the walk). This is the exception that proves the DOM-stays-in-sessionOrder rule: the badge, drag model and Alt+N index all deliberately keep reading the DOM, and only the thing the user steps with their eyes follows the paint. +**Session list layout: header strip or left sidebar** (`sessionListLayout`, default `header`; per-device via `displayKeys`, also in `SettingsUpdateSchema`): the list can move into a collapsible `<aside>` (Alt+B, `toggleSessionSidebar`) or, via `tabOrientation`, a resizable vertical `#tabRail` (desktop/tablet only). ⚠️ There is ONE `#sessionTabs`, MOVED between hosts, never a second list: `applySessionListLayout()` runs first, then `applyTabOrientation()`, both BEFORE `applyTabWrapSettings()`, and both arm/disarm `_startSidebarRichClock()`. ⚠️ Axis decisions use `_isVerticalTabList()`, never `isSessionSidebarActive()` alone. ⚠️ Rich rows share one gate, `isRichTabRows()`; rich CSS pairs sidebar+rail with comma-grouped selectors, never `:is()`, and card rules stay rail-scoped. ⚠️ Rail sort (`tabRailSort`, default `activity`) is the flex `order` property only, never a DOM reorder; the arrow-key walk alone follows computed `order`. ⚠️ Leaving sidebar mode clears `_sidebarFilter`; the handheld overlay drawer is `inert` when closed, the docked rail never. → [architecture-invariants#session-list-layout-header-strip-vs-left-sidebar](docs/architecture-invariants.md#session-list-layout-header-strip-vs-left-sidebar) -**Phone overview home screen** (`mobile-overview.js`, phones only, per-device `mobileOverviewEnabled`, default ON): under 600px the "C" logo shows a session overview (NEEDS YOU / CURRENT SESSIONS / PAST SESSIONS) instead of the welcome overlay; tablet and desktop are unchanged. The branch lives in `showWelcome()`/`hideWelcome()` (terminal-ui.js) behind `shouldUseMobileOverview()`, which is **width-driven** (`getDeviceType() === 'mobile'`) because this is a layout decision, unlike the settings namespace which stays handheld-based. ⚠️ The container ships with the `hidden` attribute and only this module removes it: never give `.mobile-overview` a bare `display` rule, since desktop does not load `mobile.css` (`media="(max-width: 1023px)"`) and would then render it unstyled. Live re-renders ride on the tail of `_renderSessionTabsImmediate()` (every state change it needs already funnels there); PAST rows come from one `_fetchUnifiedSessions(60)` per home-screen visit and resume through the shared `resumeHistorySession()`, so they behave exactly like the welcome screen's Resume list. ⚠️ Two things must stay in lockstep with surfaces outside this module, because divergence reads as a bug rather than a style: the split Run button carries the **toolbar's own classes** (`btn-toolbar btn-run mode-<backend>` / `btn-run-gear`) so the per-backend gradient and the light-skin overrides apply unchanged (mobile.css must therefore set no `background`/`color` on it), and row status uses the **session-tab language** (green dot when fine, `pulse` while working, yellow blinking row when waiting for input, red blinking row when a question is pending, mirroring `tab-alert-idle`/`tab-alert-action`). The picker mirrors the toolbar run-mode menu (`setRunMode()` + `run()`, `openWebviewFromMenu()` for saved dashboards) and deliberately omits its Recent-Sessions block, since PAST SESSIONS is that. Status pills carry `data-i18n-skip` (generic words like "idle" collide with state strings elsewhere). +**Phone overview home screen** (`mobile-overview.js`, per-device `mobileOverviewEnabled`, default ON): under 600px the "C" logo shows NEEDS YOU / CURRENT / PAST SESSIONS instead of the welcome overlay, branched in `showWelcome()`/`hideWelcome()` via width-driven `shouldUseMobileOverview()`. ⚠️ The container ships `hidden` and only this module removes it: never give `.mobile-overview` a bare `display` rule (desktop does not load `mobile.css`). ⚠️ The split Run button must carry the toolbar's own classes (`btn-toolbar btn-run mode-<backend>` / `btn-run-gear`) and mobile.css must set no `background`/`color` on it; row status must mirror the session-tab alert language. PAST rows resume through the shared `resumeHistorySession()`. Status pills carry `data-i18n-skip`. → [architecture-invariants#phone-overview-home-screen](docs/architecture-invariants.md#phone-overview-home-screen) -**Desktop home tab rail** (`home-sessions.js`, desktop only): the welcome overlay centers ~560px of content in a ~1400px window, so its left gutter is dead space; it carries the open tabs as a rail **docked flush to the left edge, full height** (a vertically centered card floating mid-gutter read as debris). Rows are in **overview order** (see below), and each carries a **created** stamp plus the **state duration** the order is computed from (`created 3d ago · working 12m`, word and anchor from `_mobileOverviewSince()` so both home screens say the same thing). A rail sorted by a number it does not show reads as arbitrarily shuffled, and a working row's plain last-active stamp always says "just now". ⚠️ The number badge is the **Alt+1..9 index**, i.e. the position in the TAB STRIP, so on a sorted rail it deliberately does NOT run 1,2,3 downward: it names a shortcut, not a row position, and renumbering it to look tidy would make every badge lie. State classification is REUSED from mobile-overview.js (`_mobileOverviewState`/`_mobileOverviewCaseFor`), which is why the module loads after it. ⚠️ The rail is `position: absolute` so the centered content never moves, which is exactly why it needs a **width gate in two places** — `HOME_SESSIONS_MIN_WIDTH` (1180) in the JS plus a `max-width: 1179px` media query as the backstop for a resize that outruns the matchMedia listener; drift between them means a rail overlapping the search panel, and `test/home-sessions.test.ts` pins them equal. ⚠️ `.home-sessions` is `display: flex`, so `[hidden]` must be re-asserted as `display: none` or the module's only visibility lever does nothing. ⚠️ Size scales with the viewport off **one knob**: `width: clamp(250px, 19vw, 430px)` plus a fluid `font-size` on `.home-sessions`, with every child sized in `em` — reintroducing `rem`/px type inside the block silently breaks the scaling, and widening the clamp past the gutter reintroduces the overlap the gate exists to prevent. The age stamps are refreshed **in place** by a 20s clock (`_tickHomeSessionsTimes()`, disarmed in `hideHomeSessions()`), never by re-rendering, which would restart every row's blink and working ring. Working state is deliberately byte-identical to the phone's: pulsing green dot + the `tab-load-spin` ring reused from the tab strip + the same green halo (added to `.mobile-overview-dot--working` at the same time), so "working" reads the same on every surface; **idle** is deliberately NOT that green — dot and pill mix toward `--text-muted` so a glance separates running from sitting. Live re-renders ride the tail of `_renderSessionTabsImmediate()` alongside the phone overview. +**Desktop home tab rail** (`home-sessions.js`, desktop only): the welcome overlay's left gutter carries the open tabs as a rail docked flush left, full height, in overview order, each row showing `created … · <state> <duration>` from `_mobileOverviewSince()`; state classification is reused from mobile-overview.js (so it loads after it). ⚠️ The number badge is the Alt+1..9 tab-strip index, never renumber it to row position. ⚠️ The width gate lives in two places that must stay equal: `HOME_SESSIONS_MIN_WIDTH` (1180) and a `max-width: 1179px` media query. ⚠️ `.home-sessions[hidden]` must re-assert `display: none`. ⚠️ Size all children in `em` off the one `clamp()` knob, never `rem`/px. Age stamps tick in place (`_tickHomeSessionsTimes()`), never by re-render. Test: `test/home-sessions.test.ts`. → [architecture-invariants#desktop-home-tab-rail](docs/architecture-invariants.md#desktop-home-tab-rail) -**Home-screen session order** (`CodemanSessionOrder` in constants.js, pure + unit-tested in `test/session-overview-order.test.ts`): BOTH home screens (phone overview and desktop rail) order rows through this ONE comparator, because they list the same sessions and must answer "which of these wants me next?" the same way. Rank is `needs` → `error` → `waiting` → `working` → `idle` → `done`, and ⚠️ **the tiebreak flips direction halfway down**: states a session is still IN sort **oldest-first** (blocked longest / running longest = most urgent), states it has STOPPED in sort **newest-first** (the session that just went quiet is the one you came back for). ⚠️ The running group keys off **`lastSubmitAt`** (the pane's last Enter), never `lastActivityAt`: a working Claude pane repaints about once a second, so its last-activity stamp is always "now" and would rank every running turn as freshly started. A working pane with no submit stamp falls back to last activity, which lands it at the SHORT end of the group rather than falsely leading it. ⚠️ A **0 stamp means "unknown", not "the epoch"**, and it sorts last within its state either way, or a brand-new session would head every oldest-first group. Final tiebreak is the user's tab order (`orderIndex`), so the list is deterministic and cannot shuffle between renders. The tab strip itself is NOT sorted by this; it stays user-ordered and drag-reorderable. +**Home-screen session order** (`CodemanSessionOrder` in constants.js, pure): BOTH home screens (phone overview, desktop rail) must order rows through this ONE comparator. Rank `needs` → `error` → `waiting` → `working` → `idle` → `done`. ⚠️ The tiebreak flips: states a session is still IN sort oldest-first, states it has STOPPED sort newest-first. ⚠️ The running group keys off `lastSubmitAt`, never `lastActivityAt` (a working pane repaints constantly). ⚠️ A 0 stamp means unknown and sorts last within its state. Final tiebreak is `orderIndex`, so the list never shuffles. The tab strip itself is NOT sorted by this. Test: `test/session-overview-order.test.ts`. → [architecture-invariants#home-screen-session-order](docs/architecture-invariants.md#home-screen-session-order) **Welcome "Resume Conversation" list** (terminal-ui.js): `loadHistorySessions()` fetches once and caches the corpus on `_historyAll`/`_historyCases`; every subsequent view (filter box, sort select, expand, the periodic refresh in panels-ui.js) goes through `_renderHistoryList()`, so never append rows to `#historyList` directly or re-fetch to re-sort. ⚠️ The box height is **class-driven**: expanding the list without `.history-list.expanded` leaves the collapsed `max-height` in place and just deepens a scroll well, which is the bug #260 reported (35 sessions in a ~4-row box). ⚠️ The A–Z sort keys off `_historyRowLabel()`, the SAME string the row renders (`name || firstPrompt || path`), most rows are transcript-backed and have no session name, so sorting on `name` alone silently does nothing. ⚠️ A filter implies expansion, and `_renderSearch()` hides `#historyHeader` (title + controls) as one unit while a search is active. Tests: `test/history-list-controls.test.ts`. -**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)** lives in that same handler: with a selection it copies, with none it must `return true` **without** `preventDefault()` or the interrupt is lost. `copyTerminalSelection` is deliberately absent from `SHORTCUT_ACTIONS` because the generic capture loop preventDefaults every match it dispatches. ⚠️ **The gate tests the CLEANED selection, not the raw one** (`CodemanCopySelection.clean` in constants.js, pure; `cleanedTerminalSelection()` reads the live terminal): xterm returns whole screen ROWS and trims only never-written cells, so the spaces a full-screen TUI paints across the rest of a row are content and reach the clipboard (138 of them per line, measured in a 282-column pane). The clean drops each line's trailing run of spaces and tabs, and it takes a LEADING run only under the rule below. ⚠️ **A LEADING margin is stripped only when the CLI DECLARES one** (`capabilities.transcriptGutter`, a bounded integer; claude and codex each declare 2, measured on live panes, and nothing else declares any). The server publishes the map as `window.__codemanTranscriptGutter` off the capability, never as an id list, and `_activeCliGutterColumns()` looks the session's mode up in it. ⚠️ **Deriving the width from the pane is what fails, and it failed twice.** The selection's own shared indent fired on 73% of ordinary indented text (401,445 windows), because a three-row window of nested YAML shares an indent for the same reason a margin does. Painted trailing padding — a TUI writes real spaces across the unused part of a row, a shell leaves them never-written — has no false positives but is a function of pane WIDTH, since that padding exists only while a rendered line stops short of the CLI's own layout width and Claude's prose wraps to fill it: measured on one live transcript the padded-row share ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the strip silently did nothing at every ordinary window size. Taking the narrowest indent on the surrounding rows fires at every width and over-strips ~1%, because a file listing inside the transcript can be the narrowest thing on screen. ⚠️ The declared width is a **CEILING**: `clean()` strips the lesser of it and the run every selected line shares, so a block only ever shifts as a unit and a selection reaching column 0 loses nothing. Over 1,392,281 selections across six pane widths of real Claude screens it over-strips none and leaves every relative indent intact. Behind the per-device `copyStripMargin` (default **ON**, in `displayKeys`, deliberately NOT in the `.strict()` `SettingsUpdateSchema`), read as `!== false` because the desktop branch of `getDefaultSettings()` returns `{}`. Nothing reads the terminal buffer on this path. ⚠️ **The margin strip is NOT idempotent, so no caller may clean twice**: it takes the lesser of the declared width and the shared run, so a second pass takes up to `margin` columns more. The trailing trim alone is a fixed point, and `copyTerminalSelection()` leaned on that by re-cleaning whatever it was handed — so the Ctrl+C branch cleans to decide whether to copy and then passes the RAW selection on, and every other copy path already hands over the raw text or reads it live. Both panes of a split resolve their own width, since `_cliGutterColumns()` and `_normalisedSelectionRange()` take the session and the terminal to read. ⚠️ Alt+drag COLUMN selections are returned untouched (`_activeSelectionMode === 3`), and the padding-only clear is FEEDBACK rather than interrupt protection, since a padding-only selection now cleans to `''` and falls through to the PTY on its own. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry) +**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)**: with no selection it must `return true` without `preventDefault()` or the interrupt is lost; keep `copyTerminalSelection` out of `SHORTCUT_ACTIONS`. The gate tests the CLEANED selection (`CodemanCopySelection.clean`: trailing padding, plus a LEADING margin only up to the width the CLI declares in `capabilities.transcriptGutter`, never one derived from the pane); the strip is not idempotent, so clean once and pass the RAW selection on, and leave Alt+drag column selections untouched. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry) **Per-device vs synced settings**: the `displayKeys` set in settings-ui.js is a **client-side merge policy**, not a wire filter. A display key seeds from the server only when localStorage has no value for it, which is what prevents one device overwriting another; `showPlanUsageLimits` is additionally `delete`d from the incoming payload outright. Separately, `SettingsUpdateSchema` is `.strict()` and simply **does not declare** `skin`, `showFileViewerButton`, `showCronButton`, `webglRendererEnabled`, `localEchoEnabled`, `cjkInputEnabled`, or `extendedKeyboardBar`, so sending one of those is a validation error. The rest (`showResponseViewer`, `showPlanUsageLimits`, `language`, and most `show*` keys) ARE in the schema and do persist server-side; they are per-device by client policy only. ⚠️ Adding a new per-device setting means deciding **both** questions: membership in `displayKeys`, and presence in the schema. -**Settings surface** (`#appSettingsModal` + `#sessionOptionsModal` + `#createCaseModal`): the `set-*` language (left rail, groups of rows, control pinned right) is shared by all three modals through ONE `:is(#appSettingsModal, #sessionOptionsModal, #createCaseModal)` scope in styles.css: an `:is()` list takes its most specific argument's specificity, so every rule keeps the id weight it had and nothing downstream shifts. **App Settings** is a rail that is a **table of contents over ONE scrolling document**, not a tab switcher: every section stays mounted (`.set-section`, ids `settings-updates|terminal|layout|appearance|models|clis|notifications|voice|shortcuts|system`, in that order, the version and the updater leading and the rest of the system settings tailing), and `switchSettingsTab(id)` keeps its historical name but SCROLLS instead of hiding. **Session Options** and **Add Case** use the same surface with a rail that really SWITCHES (`switchOptionsTab` / `switchCaseModalTab` show one `.set-section` and `.hidden` the rest, since Summary owns its own scroller, Respawn is long, and Add Case is six independent forms). ⚠️ They also take a deliberate **size-up** that App Settings does not (900px shell, 236px rail, `height:auto` between `min(560px,80vh)` and 88vh, vs App Settings' tight 760×620): they are short task panels, not a document you scan, and at scanning density they read as a few fields marooned in an empty frame. Those per-modal blocks are the design, not drift. Phones (≤860px) give App Settings the sticky `#appSettingsJump` pill and give the other two a horizontal rail strip, which neither has a pill for. ⚠️ The Session Options rail entry labelled **Session** still keys off `context` (`data-tab="context"`, `#context-tab`, `switchOptionsTab('context')`), the rename is label-only. Add Case keeps its legacy `.form-row` markup (six panels of it, every id read back by session-ui.js) and is mapped onto the look by an adapter block scoped to `#createCaseModal .set-doc`. Do not restructure those forms just to reach the row classes. ⚠️ That adapter's `summary { display:flex }` **kills the native disclosure triangle**, so every `<details>` there needs the explicit `.set-adv-chev` and both marker suppressions (`list-style` + `::-webkit-details-marker`); without it five collapsed blocks render as plain headings nobody clicks. ⚠️ **The load/save contract is `getElementById` by id**: `openAppSettings()`/`saveAppSettings()`/`openSessionOptions()` read every control by a fixed id, so moving a control between sections is free but renaming or dropping one silently stops it loading or saving. Static guards: `test/app-settings-structure.test.ts` + `test/session-options-structure.test.ts` (rail↔section pairing, one-visible-section, the `data-claude-only` entries external CLIs drop). ⚠️ Model cards (`#appSettingsModelCards`) and the effort segment are **views over hidden `<select>`s** that remain the source of truth; the cards hold the BASE model and the "1M context window" switch composes `base + [1m]` back into `claudeModel`, which is what retires the old "takes precedence over the toggle below" trap. ⚠️ `.modal-tabs`/`.modal-tab-btn`/`.modal-tab-content` are RETIRED: no modal uses them and their CSS is deleted, and a reappearance means a modal drifted off the shared surface. ⚠️ The **Header & Panels live preview** is a scale model rebuilt from the chips (`_syncLayoutPreview`); it owns NO icons, it CLONES `.set-chip-ico` out of the chip, so each icon has exactly one copy in index.html. A chip joins it via `data-preview` (slot) + `data-preview-order`, or `data-preview-text` for readouts that are not buttons. Its frame is painted from skin tokens only (hardcoded black alphas turned it into a grey slab on the light skins) and is `data-i18n-skip`. ⚠️ In Session Options → Respawn, auto-resume is a `.set-callout` whose `<label>` **wraps its own switch with no `for=`** (nesting associates them; the label+`for` pair has historically double-fired), and the cycle steps are real checkboxes (`.set-checks`), not chips. ⚠️ `admin-ui.js` injects the multi-user Users entry into `.set-rail-items` + `.set-doc`, so those hooks must survive any restructure. → [architecture-invariants#settings-surface-app-settings-session-options-add-case](docs/architecture-invariants.md#settings-surface-app-settings-session-options-add-case) +**Settings surface** (`#appSettingsModal` + `#sessionOptionsModal` + `#createCaseModal`): one `set-*` language shared through a single `:is(...)` id scope in styles.css. App Settings' rail is a table of contents over ONE scrolling document (`switchSettingsTab` scrolls); Session Options and Add Case really switch (`switchOptionsTab` / `switchCaseModalTab`), and their larger per-modal size blocks are the design, not drift. ⚠️ **The load/save contract is `getElementById` by id**: renaming or dropping a control id silently stops it loading or saving. ⚠️ The Session Options "Session" entry still keys off `context` (label-only rename). ⚠️ Add Case keeps its legacy `.form-row` markup via an adapter; every `<details>` there needs `.set-adv-chev` plus both marker suppressions. ⚠️ Model cards and the effort segment are views over hidden `<select>`s, which stay the source of truth. ⚠️ `.modal-tabs*` classes are retired; `admin-ui.js` needs `.set-rail-items` + `.set-doc` to survive any restructure. Guard: `test/app-settings-structure.test.ts`. → [architecture-invariants#settings-surface-app-settings-session-options-add-case](docs/architecture-invariants.md#settings-surface-app-settings-session-options-add-case) **Header button visibility**: most header controls are opt-in and hidden by a marker class (`btn-multimonitor--hidden`, `btn-response-viewer-header--hidden`, `btn-file-viewer--hidden`, `btn-cron--hidden`) that `applyHeaderVisibilitySettings()` (settings-ui.js) toggles after settings load; the multi-monitor button is instead stripped at render by `renderIndexHtml`. ⚠️ Hiding must go through the marker class: the base rules are `display:inline-flex !important`, so an inline style cannot override them. Current desktop default is WS/CPU/MEM + File Viewer + gear, with the token chip and lifecycle-log button OFF. ⚠️ New header controls must not leak onto phones; `test/mobile-header-buttons-policy.test.ts` is the static guard. → [architecture-invariants#header-button-visibility-multi-monitor-response-viewer-file-viewer-cron](docs/architecture-invariants.md#header-button-visibility-multi-monitor-response-viewer-file-viewer-cron) **Gesture control** (camera hand-tracking overlay, opt-in, default OFF): `CODEMAN_GESTURE=1` makes the feature *available*; `gestureControlEnabled` turns it on. The bundle is injected by `renderIndexHtml` only when enabled, which is why that method is `async` and reads settings with `readSettings(true)` (a fresh read: a post-save reload lands inside the 2s cache TTL and would otherwise render the pre-toggle state). **Source lives in `packages/gesture-control/`; edit there, run `npm run build:gesture`, and commit the regenerated bundle** because dev serves the committed bundle with no runtime bundler. The MediaPipe wasm + model are fetched separately and gitignored. ⚠️ Keep `MP_VERSION` in `fetch-gesture-assets.mjs` in sync with `@mediapipe/tasks-vision`. → [architecture-invariants#gesture-control-the-source-package](docs/architecture-invariants.md#gesture-control-the-source-package) -**Terminal font weight** (`terminalFontWeight` / `terminalFontWeightBold`, per-device, default = xterm's own `normal`/`bold`): bold text on the theme's default foreground carries exactly ONE cue, the weight step. Claude Code marks its markdown bold with a bare `ESC[1m` and no colour change, and xterm substitutes a bright colour for bold only when the foreground is a palette index 0-7, so the substitution never fires there. A two-face family keeps that step small and 400 stays 400 whatever family is chosen, which is why the NORMAL slot is settable at all. `CodemanTerminalFont.resolveWeights()` (constants.js, pure) resolves both slots, each against **its own** xterm default, so an unset bold weight can never inherit `normal`. ⚠️ **The `@font-face` descriptor, not the file, is what the browser synthesizes from**: `fonts/jetbrains-mono-variable.woff2` carries a `wght` axis of 100-800, and while `styles.css` declared it `400 700` every weight below 400 rendered identically to 400 and 800 identically to 700 — measured — so the setting was a no-op for anyone without Fira Code or Cascadia Code installed, which is most installs. It is declared `100 800`; re-narrowing it silently guts the feature (`test/terminal-font-weight.test.ts` pins the range). ⚠️ A live save must reach **both echo overlays** (`refreshFont()` — they cache `terminal.options.fontWeight` and paint it into their spans, so typed characters otherwise keep the old weight, most visible on a phone) **and open Agent Teams panes** (they read their options at construction, exactly like `applyTerminalSkin()` propagates). ⚠️ `_awaitTerminalFont()` is deliberately untouched: `CharSizeService` measures through the CSS `font` shorthand, which RESETS the weight, so the measured face is always the 400 one and a weighted descriptor would ask for nothing new. +**Terminal font weight** (`terminalFontWeight` / `terminalFontWeightBold`, per-device, default = xterm's own `normal`/`bold`): Claude Code's markdown bold is a bare `ESC[1m`, so the weight step is its only cue. `CodemanTerminalFont.resolveWeights()` (constants.js, pure) resolves each slot against **its own** xterm default. ⚠️ The `@font-face` for `fonts/jetbrains-mono-variable.woff2` must stay declared `100 800` (the browser synthesizes from the descriptor, not the file); narrowing it silently makes the setting a no-op. ⚠️ A live save must reach both echo overlays (`refreshFont()`) and open Agent Teams panes. ⚠️ Leave `_awaitTerminalFont()` untouched. Test: `test/terminal-font-weight.test.ts`. → [architecture-invariants#terminal-font-weight](docs/architecture-invariants.md#terminal-font-weight) **Theme skins / branding / i18n**: `skin` selects a palette via `data-skin` on `<html>`, applied by an **inline pre-paint script** in `index.html` reading `localStorage['codeman:skin']` to avoid a flash of wrong theme. ⚠️ A skin is **four things that must stay in sync**, and missing any one degrades silently: the `html[data-skin="…"]` token block in `styles.css`, the xterm ANSI palette in `terminal-ui.js`, the pre-paint allowlist, and the Settings picker (both in `index.html`). `test/skin-themes.test.ts` is the static guard. Light skins additionally need `color-scheme: light` and xterm `minimumContrastRatio: 4.5`, and `applyTerminalSkin()` must call the local-echo overlay's `refreshFont()` because it caches the terminal fg/bg. `displayName` changes user-facing browser branding only and must NEVER rename npm package, CLI, API, storage, CSS, or protocol identifiers. `language` (`en`/`zh-CN`) keeps English as the canonical source so live switching stays reversible. User display names flow through `textContent`/attribute APIs and the server title's HTML escaper, never `innerHTML`. → [architecture-invariants#theme-skins](docs/architecture-invariants.md#theme-skins) **Foldable settings identity**: responsive layout is width-driven via `MobileDetection.getDeviceType()`, but the localStorage namespace uses `MobileDetection.isHandheldDevice()` so an unfolded Android foldable keeps `codeman-app-settings-mobile`. ⚠️ Do not switch per-device settings namespaces from instantaneous viewport width: a posture-triggered WebView reload would lose opt-in UI. Regression profile: `OPPO Find N5 (unfolded)` in `test/mobile/devices.ts`. → [architecture-invariants#foldable-settings-identity](docs/architecture-invariants.md#foldable-settings-identity) -**Folding devices: a fold is not a keyboard, and dialogs avoid the hinge** (Apple's [Designing for iPhone Duo](https://developer.apple.com/design/human-interface-guidelines/designing-for-iphone-duo)): ⚠️ **A visual-viewport resize that changes the WIDTH is the device changing shape** (a rotation, or a foldable opening or closing) **and is never the virtual keyboard**, which only ever takes height. `handleViewportResize()` read any height drop over 150px as the keyboard appearing, so closing an iPhone Duo (626→466pt wide, 890→678pt tall) latched `keyboardVisible` with no keyboard on screen: the accessory bar appeared, `main` grew 84px of dead padding, and `updateAppHeight()` (which bails while the keyboard is up) stopped refreshing `--app-height`. The latch is STICKY, since clearing it needs the height back within 100px of a baseline belonging to a display the user is no longer looking at, so it survived until the device was opened again, and rotating any phone hit it too. The shape branch re-baselines instead, which is also what lets a keyboard opened AFTER the fold be detected. ⚠️ `init()` must seed `lastViewportWidth`, or the very first resize reads as a shape change and swallows a real keyboard. ⚠️ **The hinge is a reserved region.** The CSS Viewport Segments media features report two segments only while a foldable is actually bent, and `--fold-inline-end`/`--fold-block-end` (styles.css) measure the strip to keep clear: `0px` on everything else, so the rules are inert by construction rather than by a branch. Codeman's seven centred overlays are all `position: fixed; inset: 0` flex boxes, and each shrinks its CONTENT box with padding rather than the box itself, so the backdrop still covers the far side of the fold and still swallows taps there. ⚠️ Each rule RE-STATES the overlay's own gutter (a later `padding-right` longhand beats the earlier `padding` shorthand it composes with), and the palette needs the compound `.modal.command-palette-modal` because mobile.css loads later and pads it with a shorthand under 768px. `test/foldable-layout.test.ts` DERIVES the overlay list from the stylesheet and compares both numbers, so a new overlay or a moved gutter fails there instead of on hardware nobody has. ⚠️ Physical sides, not logical ones: dialogs sit in the LEFT segment (the TOP one in tabletop pose) in every language, because the HIG keeps Duo's side controls on the same physical edge in RTL. ⚠️ **With the keyboard up, a shape change baselines to `window.innerHeight`** (the layout viewport, keyboard-free on both engines), never the shrunk visual height: the shrunk baseline made the settle event that follows every rotation or fold read as the keyboard closing, and the layout could not recover, since no further drop could re-arm the show branch. ⚠️ **A base gutter that a LATER `@media` block overrides needs its own fold restatement in that block, on a ZERO base**: the phone-width path picker and path preview drop to `padding: 0` under 600px, and the unconditional fold rules at the end of the file put 16px and 18px back on every phone (measured at 393 and 500). The palette's compound rule lives INSIDE the 600-768px band whose mobile.css shorthand it composes with, because outside it there is no side gutter and the addition pushed the shell 6px off centre; the response viewer's tabletop cap has a twin at the end of mobile.css, whose phone block (under 600px) otherwise outranks it. `test/foldable-layout.test.ts` simulates the cascade across BOTH stylesheets at every breakpoint, with and without the fold rules, so a moved gutter or an unscoped composition fails there rather than on hardware nobody has. Profiles: `iPhone Duo (outer)` / `iPhone Duo (inner)` in `test/mobile/devices.ts`; unit coverage in `test/viewport-shape-change.test.ts`. → [architecture-invariants#folding-devices](docs/architecture-invariants.md#folding-devices) +**Folding devices: a fold is not a keyboard, and dialogs avoid the hinge**: ⚠️ in `handleViewportResize()`, a visual-viewport resize that changes the WIDTH is a shape change (rotation, fold) and must never be read as the keyboard; it re-baselines instead, or `keyboardVisible` latches with no keyboard. ⚠️ `init()` must seed `lastViewportWidth`. ⚠️ With the keyboard up, a shape change baselines to `window.innerHeight`, never the shrunk visual height. ⚠️ The hinge is reserved via `--fold-inline-end`/`--fold-block-end` (0px when unfolded): each overlay fold rule must re-state its own gutter, a base gutter overridden by a later `@media` block needs its own fold restatement there on a zero base, and dialogs use physical sides (left/top segment) in every language. Guard: `test/foldable-layout.test.ts`. → [architecture-invariants#folding-devices](docs/architecture-invariants.md#folding-devices) **WebGL renderer toggle** (`webglRendererEnabled`, per-device): the GPU-stall watchdog's sticky `codeman-webgl-disabled` marker survives page loads and is cleared only by an explicit OFF→ON save or `?webgl=force`. `?nowebgl` forces the DOM renderer per-load. → [architecture-invariants#webgl-renderer-toggle](docs/architecture-invariants.md#webgl-renderer-toggle) -**Shell keyboard accessory bar + one-shot Ctrl** (issue #262, `keyboard-accessory.js`): a **shell**-mode session automatically swaps the mobile accessory bar for terminal controls (Ctrl, Esc, Tab, four arrows, paste, dismiss); every other mode keeps the agent bar. `setMode()` now records the user's `extendedKeyboardBar` preference as the **base** layout and `refreshForActiveSession()` (called from `selectSession`) resolves base-vs-shell, so a settings save during a shell session cannot yank the bar away and switching back restores the user's choice. ⚠️ **Ctrl is a ONE-SHOT modifier applied in `terminal.onData`, not in a keydown handler**: a virtual keyboard emits no usable key events, so the character only exists as onData text. The hook sits AFTER `shouldSuppressTerminalQueryResponse` (xterm answers DA/CPR through onData too, and one of those would silently spend the modifier) and BEFORE every send path, so the control byte follows the normal control-char route. ⚠️ **Not every onData chunk is a keystroke**, and the query filter is not enough on its own: xterm ALSO emits mouse and focus reports on its own initiative, so the hook skips them via `isTerminalFocusOrMouseReport()` (they still reach the PTY, they just don't count as the next key). The mouse half is live — a shell session keeps the NARROW strip, so mouse DECSETs reach the browser and one tap while vim/htop runs spent the armed modifier silently (measured). The focus half is defense in depth: `FOCUS_ESCAPE_FILTER` in `session.ts` strips `\x1b[?1004h` from every PTY read, so `sendFocusMode` never turns on today; if it ever did, the bar's own post-key refocus would emit `\x1b[I` and eat the modifier before the user typed. ⚠️ It must disarm on ALL of: use, second tap, any other accessory key, session switch, keyboard dismissal, and a layout swap; a modifier left armed turns the next innocent keystroke into a control byte. ⚠️ **onData is not the only input path** — with `cjkInputEnabled` on, the CJK textarea owns the keyboard (onData returns early for everything it swallows, and the focus router sends `terminal.focus()` there, which is where the bar refocuses after every key), so `_handleCjkInput()` applies the modifier too. It is that module's single choke point to the PTY, so one call covers typed characters, IME flushes, Enter, backspace and arrows. Without it an armed modifier could neither fire NOR be spent, and survived to a later keystroke. Mapping is `ctrlByteFor()` (`code & 0x1f` over @A-Z[\]^_ and a-z, plus Ctrl+Space=NUL / Ctrl+?=DEL); characters with no control equivalent pass through unchanged, like a hardware keyboard. ⚠️ The armed style is `.accessory-btn.accessory-btn-ctrl.armed` (0,3,0) in BOTH stylesheets, and it cannot outrank mobile.css's light-skin repaint at **(0,3,1)** (`:is()` inherits its most specific argument, and that list holds `.btn-toolbar.btn-shell`) — so that rule excludes the state by hand as `.accessory-btn:not(.armed)`. Without the exclusion the armed button renders identically to a resting one on all four light skins, which is worse than no armed style at all. +**Shell keyboard accessory bar + one-shot Ctrl** (`keyboard-accessory.js`): a shell-mode session swaps the mobile accessory bar for terminal controls; `setMode()` records `extendedKeyboardBar` as the base layout and `refreshForActiveSession()` resolves base-vs-shell. ⚠️ Ctrl is a one-shot modifier applied in `terminal.onData` (after `shouldSuppressTerminalQueryResponse`, before every send path), and must skip `isTerminalFocusOrMouseReport()` chunks. ⚠️ It must disarm on use, second tap, any other accessory key, session switch, keyboard dismissal and layout swap. ⚠️ `_handleCjkInput()` must apply it too (the CJK textarea bypasses onData). Mapping: `ctrlByteFor()`. ⚠️ mobile.css's light-skin repaint must keep excluding `.accessory-btn:not(.armed)` or the armed state is invisible. → [architecture-invariants#shell-keyboard-accessory-bar-and-one-shot-ctrl](docs/architecture-invariants.md#shell-keyboard-accessory-bar-and-one-shot-ctrl) -**Mobile prompt composer** (PR #444, the first slice of #359, `keyboard-accessory.js`): the agent bars' Paste key is now **Compose**, a dialog with a native multiline textarea (autocorrect, autocapitalize, spellcheck) where Enter adds a line and only **Send** submits; the shell bar keeps the direct Paste dialog, since shell input is not an agent prompt. Opening it ADOPTS the whole editable terminal prompt (`_takePendingLocalEcho`): the local-echo overlay's pending text has never reached the PTY, but the flushed prefix has, so that prefix is erased with backspaces counted in CODE POINTS (`Array.from(text).length`; measured on Claude Code 2.1.278, `a` + emoji + `b` takes three, and the UTF-16 count sent four and ate the neighbour; `clearTerminalInput()` in terminal-ui.js moved with it). ⚠️ **Drafts are per-session and in memory only** (`_composerDrafts`, never persisted: prompts routinely carry secrets, and persisting them would need the 0600 treatment the intent store gets). Every non-Send exit (Cancel, backdrop, Escape, "Use terminal keyboard") leaves the taken text ONLY in the draft, with the dot on the key (`has-draft`, kept in step by `_syncComposerDraftIndicator`) as the signal that the terminal prompt is empty on purpose, and `_cleanupSessionData` discards the draft with the session. ⚠️ **Delivery is a hand-built bracketed-paste frame** (`\x1b[200~` + the text with newlines mapped to `\r` + `\x1b[201~`, byte-identical to what `terminal.paste()` would emit) through `_sendInputAsync` WITHOUT `{ useMux: true }`, plus a SEPARATE Enter 120 ms later WITH it (codex drops keys that share a PTY read with a bracketed paste). Not `terminal.paste()`, for two reasons: xterm's `bracketedPasteMode` mirror is false for every session after a tab switch or reload (`terminal.reset()` in the replay re-clones the DEC modes, the tmux capture carries no `?2004h`, and tmux never forwards the pane's DECSET to a client after attach), so a paste through xterm would go out unbracketed and the CLI would submit at the first `\r`; and xterm's onData is where the local-echo paste branch flushes pending overlay text AHEAD of the block, text the composer has already taken and erased from the PTY, so the frame goes straight to the wire with the composer as the prompt's only owner. The unconditional frame is safe for a CLI that never enabled DECSET 2004 because tmux does the gating (measured against a live pane: markers stripped for a `cat -v` pane, forwarded intact to a process that had emitted `?2004h`). ⚠️ The frame must never take the mux fallback: `TmuxManager.sendInput()` strips every `\r` and `\n`, which welds the lines together and submits them. ⚠️ The size guard sits on the `MAX_INPUT_LENGTH` boundary (64 KiB of UTF-16 code units): `ws-routes.ts` drops a longer frame WITHOUT an ACK, which would wedge the durable queue, so `_composerMaxLength` is derived from that limit minus both markers and an oversized prompt stays a draft with a toast. ⚠️ `.prompt-composer-overlay` is a `.paste-overlay` with a gutter of its own (a `padding` shorthand whose bottom is 12px plus the safe area), and the unconditional `.paste-overlay` fold rule at the end of styles.css is a later longhand at the same specificity, so it ERASED that gutter (measured at 393x852: `padding-bottom: 0px` flat, and the hinge strip REPLACING the gutter with the fold variables set): the composer has its own restatement after the fold rules, the dialog's `max-height` subtracts `--fold-block-end`, and `test/foldable-layout.test.ts` lists the composer in `ELEMENTS` by hand, because its derived overlay list keys on rules that declare `position: fixed; inset: 0` themselves. Tests: `test/mobile-prompt-composer.test.ts` (in the CI gate, deliberately not under `test/mobile/**`). +**Mobile prompt composer** (`keyboard-accessory.js`): the agent bars' Paste key is **Compose**, a native multiline dialog where only **Send** submits (the shell bar keeps plain Paste). Opening it adopts the whole terminal prompt (`_takePendingLocalEcho`), erasing the flushed prefix with backspaces counted in code points. ⚠️ Drafts are per-session and in memory only (`_composerDrafts`), never persisted (prompts carry secrets). ⚠️ Delivery is a hand-built bracketed-paste frame via `_sendInputAsync` WITHOUT `useMux`, then a separate delayed Enter WITH it; never `terminal.paste()`, and the frame must never take the mux fallback (it strips newlines). ⚠️ `_composerMaxLength` must stay derived from `MAX_INPUT_LENGTH` minus the markers, or an oversized frame wedges the durable queue. ⚠️ The composer overlay needs its own gutter restatement after the fold rules. Test: `test/mobile-prompt-composer.test.ts`. → [architecture-invariants#mobile-prompt-composer](docs/architecture-invariants.md#mobile-prompt-composer) -**The PTY and the browser terminal must never disagree about size** (issue #464, `syncTerminalGeometry()` in terminal-ui.js): Claude Code's TUI wraps its frame at the width the PTY reported and erases the previous frame by walking the cursor up the number of rows it BELIEVES that frame occupied. A browser terminal of a different width makes each logical line take more physical rows than Ink counted, so `eraseLines(n)` clears too few and the new frame paints over rows nothing erased — the doubled lines and half-overwritten prose in #464, **measured** against this repo's xterm (a 120-column PTY against a 62-column terminal renders every wrapped line twice; `test/terminal-pty-geometry.test.ts` pins it, and pins the clean render at matching widths so the assertion cannot pass against code that fixes nothing). ⚠️ **`fitAddon.fit()` is NOT the way to resize this terminal.** It resizes xterm to `proposeDimensions()` RAW while every server-facing path reports those floored at 40x10, so whenever the floor bit the two diverged silently — **measured in Chrome at 430px**: font size 44 proposed 13 columns, the server was told 40, and xterm stayed at 13. `syncTerminalGeometry()` fits, floors and applies as one step and is the ONE function that may change the size; `test/terminal-pty-geometry.test.ts` sweeps the seven modules that touch the main terminal (terminal-ui, mobile-handlers, app, ralph-panel, settings-ui, tab-rail-resize, notification-manager) for a bare fit; the split pane and teammate terminals own their own sizing and are out of scope. ⚠️ **Withhold the fit wherever you withhold the SIGWINCH.** `throttledResize` (virtual keyboard up) and `sendResize` (session detached into its own window) used to reflow locally and skip only the server write, which is the one combination that cannot be right. ⚠️ **A font change is a geometry change**: `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size and told the server nothing, so raising the font on a phone left the CLI wrapping at the old column count (`_refitAfterCellSizeChange`). ⚠️ **Resize is no longer write-only.** `Session.resize` DECLINES a small-viewport request while a desktop connection holds an active sizing claim and says nothing, so both transports now answer with `Session.ptyGeometry` (`{"t":"zc"}` on the socket, the body of the resize POST) and `_onPtyGeometryReport` adopts it — a terminal that keeps a WIDTH the PTY refused renders GARBLED, not merely wrong-sized. ⚠️ While that refusal stands (`_paneWidthRefused`), a resize ASKS for the container's width without applying it (`_geometryForResizeRequest`: rows follow the container, columns stay at the PTY's): fitting first re-wrapped the whole buffer to the container and back on every 30s mobile retry, and ran the scrollback clear for a resize that brings no redraw. `selectSession` clears the flag, since it belongs to the previous pane. ⚠️ **`ptyGeometry` is null without a live pane**, never the field values: they are seeded at spawn (`_notePtySpawnGeometry`, so a reattached pane reports its tmux window's real size) and moved by `resize()`, but a dead-pane session still holds the constructor defaults of 120x40 or a gone pane's size — reporting those made a client adopt a size no process was ever told and claim another device owned the pane when none existed. ⚠️ **COLUMNS ONLY.** Adopting the PTY's ROWS was a regression: a phone taking a desktop's 43 rows into a viewport with room for 18 painted an `.xterm-screen` far taller than its container, xterm's own viewport then had nothing to scroll, and the CLI's input line sat below the container with no gesture able to reach it — output visible, typing invisible, for as long as the claim stayed hot. Width is the axis the wrap arithmetic depends on; rows only decide how much is on screen, and keeping the local count keeps the composer at the bottom of a viewport that scrolls. ⚠️ **`.term-overflows-x` keys on what does not FIT, not on a PTY mismatch**, because the 40-column floor widens the terminal past a narrow container with the PTY agreeing throughout (measured at 360px: font 18 paints 433px, font 24 paints 578px — 38% unreachable, and `increaseFontSize` reaches 24 in two taps). `_syncTerminalOverflowAffordance()` MEASURES `.xterm-screen` against the container on the next frame rather than deriving it from cell arithmetic. ⚠️ That rule sets **both** overflow axes — mobile.css loads later with `.terminal-container { overflow: visible }`, and a bare `overflow-x` would leave overflow-y computing to `auto`, handing the browser a vertical scroll container the terminal's touch handler does not know about — and sets **no `touch-action`**: the terminal's own touchmove handler pans it (`canPanHorizontally`), because `touchstart` preventDefault()s every 'content' tap and that cancels a native `pan-x` before it starts (measured: a 140px swipe reached scrollLeft 141 without that preventDefault and 0 with it). Removing the class returns `scrollLeft` to 0, so a resolved mismatch cannot leave the pane parked off-screen. +**PTY and browser terminal geometry** (#464): a browser terminal whose width differs from the PTY's garbles Claude's redraws, so `syncTerminalGeometry()` (terminal-ui.js) is the ONE function that may resize the main terminal (never a bare `fitAddon.fit()`), a font change is a geometry change, and the fit is withheld wherever the SIGWINCH is. ⚠️ Resize is answered with `Session.ptyGeometry`, and a client adopts its COLUMNS only (never rows); `ptyGeometry` is null without a live pane. → [architecture-invariants#pty-and-browser-terminal-geometry](docs/architecture-invariants.md#pty-and-browser-terminal-geometry) -**Terminal resilience: replay clears, renderer liveness, fetch deadlines**: three rules that each close a way the terminal silently stops being correct. ⚠️ **A replay clear MUST be in-stream, never `reset()`/`clear()`.** xterm's `write()` is asynchronously queued while `Terminal.reset()` is synchronous and, per upstream, "does not clear input buffers and does not reset the parser" — so bytes queued just before a reset are parsed AFTER it and fuse into the snapshot written next. **Measured** against the real xterm in this repo: `write('p8'); reset(); write('rmissions')` renders `p8rmissions`; the queued `\x1bc` renders `rmissions` and clears scrollback. `_resetTerminalForReplay()` (app.js) is the ONE clear, a single queued `\x1bc` (RIS), and all three replay paths go through it; RIS rather than `\x1b[3J\x1b[H\x1b[2J` because the erase leaves modes, charsets, scroll regions and SGR state alone. Callers may still chunk the content — ordering in the queue is what matters, not writing it in one call. ⚠️ **The renderer watchdog reads xterm privates and CANNOT be covered by the gate.** `_kickRenderer()` (terminal-ui.js) cancels a stale `_core._renderService._renderDebouncer._animationFrame` and forces a repaint. **Verified against xterm 6.0.0** (jsdom, after `open()`): the field path resolves, a forced stale handle genuinely makes `refreshRows` a no-op, and the kick schedules a fresh frame. **Reasoned, not reproduced here**: the premise that iOS discards scheduled rAF callbacks when a PWA backgrounds, which is what leaves the handle stale — that half wants a real-device pass. Codeman has exactly ONE xterm for the whole page load, so one backgrounding would wedge it until a reload. `_renderService` only exists after `open()`, which needs a real DOM, and the gate runs in node — so `test/xterm-private-api.test.ts` pins the RESOLVED lockfile version (not the `^6.0.0` range, which a real upgrade slips through) and a bump means re-verifying by hand. Every access is optional-chained on purpose: a renamed field must degrade to a no-op, never throw on a 2s timer. ⚠️ **Every terminal capture carries a deadline, and the helper reads the BODY** (`_fetchTerminalCapture`, app.js). `await fetch()` settles on response HEADERS, so clearing the timer there leaves the body — the multi-megabyte `?full=1` capture this exists for — unbounded: **measured** at 4026ms under a 1000ms deadline before the fix. The helper therefore returns `{json, headers, headersAt}` rather than a `Response`, and `_terminalCaptureInflight` is scoped the same way so a body still streaming counts toward a capture starting beside it. It degrades to a plain fetch where `AbortController` is missing — the deadline is a safety net, not a dependency. Tests: `test/terminal-resilience.test.ts` (pure decisions), `test/xterm-private-api.test.ts`. +**Terminal resilience**: a replay clear is the queued in-stream `\x1bc` in `_resetTerminalForReplay()`, never `reset()`/`clear()` (queued bytes fuse into the snapshot); the renderer watchdog `_kickRenderer()` reads xterm privates, pinned by `test/xterm-private-api.test.ts` against the resolved lockfile version; every terminal capture fetch has a deadline that covers the BODY (`_fetchTerminalCapture`). → [architecture-invariants#terminal-resilience-replay-clears-renderer-liveness-fetch-deadlines](docs/architecture-invariants.md#terminal-resilience-replay-clears-renderer-liveness-fetch-deadlines) -**WebSocket output-gap reconcile** (`_wsOutputGapSession`, app.js): terminal OUTPUT frames carry no sequence number (input frames do — `seq`+`cid`, at-most-once, ACKed), so a dropped socket leaves a hole nothing replays. ⚠️ **The gap is narrower than "the device went offline"**: if the network drops, SSE drops with it and `handleInit`'s keepTerminal branch already calls `_onSessionNeedsRefresh`. The uncovered case is the WS dying while SSE stays up (half-open socket, proxy idle-timeout, ping timeout), because `_onSSETerminal` discards every SSE terminal frame while `_wsReady` is true and `_wsReady` only flips in `ws.onclose`. Reaching `onclose` at all means the drop was unintentional (`_disconnectWs` nulls the handler first), so the session is marked and the next successful open reconciles. ⚠️ **The marker must be cleared by every path that repaints that session's buffer, and ONLY once one actually has.** `_markTerminalBufferReconciled()` is called from `selectSession` after its load, from `_cleanupSessionData`, and from `_onSessionNeedsRefresh` **at the repaint itself, not in its `finally`** — clearing on every exit meant a reconcile that threw, or hit the fetch deadline (the flaky link the marker exists for), dropped the gap with nothing to retry it. `ws.onopen` no longer clears it up front either, so a reconcile that is skipped or fails is tried again on the next open; re-entry is safe because `_terminalRefreshOwner` makes a second reconcile for the same session a no-op. `selectSession` loads the buffer and only THEN calls `_connectWs`, so without that clear the socket opening afterwards replays the whole buffer a second time on top of the one just written. Sequencing the output frames is the real fix and is not done. This is reasoned from the code path, not observed on a device. +**WebSocket output-gap reconcile** (`_wsOutputGapSession`, app.js): output frames carry no sequence number, so an unintentional WS close while SSE stays up marks the session and the next open reconciles. ⚠️ The marker is cleared only once a repaint actually happened (`_markTerminalBufferReconciled()`, never in a `finally`). → [architecture-invariants#websocket-output-gap-reconcile](docs/architecture-invariants.md#websocket-output-gap-reconcile) -**Service worker: precache and cache key are BUILD-GENERATED** (`sw.js` + `scripts/build.mjs`): the build content-hashes assets and rewrites two exact declarations in `sw.js` — `const BUILD_ID = 'dev';` and `const HASHED_ASSETS = [];`. ⚠️ **Each must appear exactly once or the build THROWS**, which is deliberate: the list used to be hand-maintained with PRE-hash names, so every entry 404'd in production and `cache.add().catch(() => {})` hid it (15 of 23 verified failing against a running instance). The dev literals are valid on their own, so dev serves an unrewritten worker with an empty precache. ⚠️ **`caches.match` must pass `ignoreSearch: true`**: `renderIndexHtml` runs `cacheBustAssets`, which appends `?v=<mtime>` to every same-origin `.js`/`.css` reference INCLUDING content-hashed names, so the page requests `/app.<hash>.js?v=<mtime>` while the cache holds `/app.<hash>.js`. Without it no precached entry is reachable and the install downloads ~1.3MB that can never be served — once per deploy, since `CACHE_NAME` now carries the build id. That per-build key is what makes `activate`'s cleanup actually delete anything; it used to be the constant `'codeman-v1'`, so assets from every past release accumulated forever. Contract pinned by `test/sw-precache-manifest.test.ts`, which PARSES the `HASHABLE` list out of `build.mjs` rather than copying it. +**Service worker precache** (`sw.js` + `scripts/build.mjs`): `BUILD_ID` and `HASHED_ASSETS` are build-generated and the build THROWS unless each declaration appears exactly once; `caches.match` must pass `ignoreSearch: true` because `cacheBustAssets` appends `?v=` to hashed names. → [architecture-invariants#service-worker-precache-and-cache-key](docs/architecture-invariants.md#service-worker-precache-and-cache-key) -**Dismissing the on-screen keyboard** (PRs #279/#280, `terminal-ui.js`): the terminal parks focus on a hidden textarea that nothing used to release, so TWO gestures now blur it, and they own different regions. **(1)** `_installMobileKeyboardDismiss()` — a document-level `touchend` that fires only while the terminal input actually holds focus, **never inside `#terminalContainer`** (tap classification owns that) and **never on a control** (`MOBILE_KEYBOARD_DISMISS_EXEMPT_SELECTOR`, matched with `closest()` so an icon inside a button counts). Session tabs are covered by the selector's `[tabindex]:not([tabindex="-1"])` arm, which is what stops a tab tap from blurring and then being re-focused by `selectSession()`. **(2)** In `_handleMobileTerminalTap`, a second tap on **inert `content`** (`startedWithTerminalFocus`) blurs instead of re-focusing. ⚠️ Scoped to `content` on purpose: the prompt row (`input`) keeps focus-then-position so a second tap still places the caret, and actionable rows blur earlier via `_isActionableMobileTerminalTap`. ⚠️ **A scroll ends in `touchend` too** — dismissing there closes the keyboard and drops the composer mid-read, so travel is tracked from `touchstart` and multi-touch is never a tap. Both classifiers MUST share one threshold: `initTerminal`'s `TAP_THRESHOLD` reads `MOBILE_KEYBOARD_DISMISS_TAP_SLOP`, since a gesture the terminal calls a scroll and the dismiss handler calls a tap is exactly that bug. ⚠️ **The gate excludes `test/mobile/**`, so CI cannot see the only test covering (1)** — run `npm run test:mobile -- test/mobile/keyboard.test.ts` by hand and diff the FAIL list against master. (Not `npm test --`: the gate's config excludes that path, so a file filter pointing into it matches nothing and exits green having run zero tests.) That blind spot is why merging the two PRs, which conflicted semantically but not textually, produced a red suite with two green CI checks. +**Dismissing the on-screen keyboard** (`terminal-ui.js`): two gestures blur the terminal's hidden textarea. (1) `_installMobileKeyboardDismiss()`, a document `touchend` that must never fire inside `#terminalContainer` or on a control (`MOBILE_KEYBOARD_DISMISS_EXEMPT_SELECTOR`, via `closest()`). (2) In `_handleMobileTerminalTap`, a second tap on inert `content` blurs; the prompt row keeps focus-then-position. ⚠️ A scroll also ends in `touchend`: both classifiers must share one threshold (`TAP_THRESHOLD` reads `MOBILE_KEYBOARD_DISMISS_TAP_SLOP`), and multi-touch is never a tap. ⚠️ CI cannot see the only test for (1): run `npm run test:mobile -- test/mobile/keyboard.test.ts` by hand and diff the FAIL list against master. → [architecture-invariants#dismissing-the-on-screen-keyboard](docs/architecture-invariants.md#dismissing-the-on-screen-keyboard) **Phone toolbar: Enter replaces Shell** (post-1.8.0): inside `@media (max-width: 599px)` `btn-shell` is `display:none` and `btn-enter` takes its slot (`order: 4`); starting a shell moved into the Run dropdown (`Terminal / Shell` → `setRunMode('shell')` → `run()` → `runShell()`, button label "Run SH"). `runMode` is `z.string().max(20)` server-side, so new modes need no schema change. Desktop and tablet keep the green Run Shell button unchanged. @@ -366,9 +370,9 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L **Connection-loss UI** (`computeConnectionLossUi()` in constants.js, writer `_updateConnectionLossUi()` in app.js): the service worker serves the cached app shell, so an unreachable server (phone off the tailnet, VPN down, server stopped) used to render a normal-looking empty dashboard whose only tell was the 8px header dot, which reads as "no sessions", not "no connection". Two surfaces now: a full-screen **overlay** while no server state has loaded this page load (nothing behind it is worth preserving), and a non-blocking **banner** once it has (the terminal scrollback stays readable). ⚠️ A **2.5s grace** is load-bearing: a COM deploy restarts the server and SSE is back in ~200ms, and a banner on every deploy trains the user to ignore it. `navigator.onLine === false` skips the grace, since that is never a blip. Retry re-arms SSE **and** the terminal WS (`planWsReconnect` can 'give-up', and the SSE backoff caps at 30s). -**SSE staleness watchdog** (`computeSseStale()` in constants.js, `_checkSseStale()` + a 5s interval in app.js): an `EventSource` that stops delivering does not always error, so `onerror` never fires, the header dot stays green, and every SSE-driven surface (tab status dots, sessions created on another device, renames) freezes until the user reloads. ⚠️ The 15s server keepalive was an SSE **comment** (`:keepalive`), and comments are **invisible to `EventSource` by spec**, so there was nothing a client could observe: it is now the named `sse:heartbeat` event (`cleanupDeadClients()`, sse-stream-manager.ts), which is exactly why the frame had to change type. ⚠️ Staleness is judged **only while the status is `connected`** and the device is online; that guard is the loop breaker, since a forced `connectSSE()` leaves `connected` immediately and cannot re-fire while a reconnect is in flight. ⚠️ The liveness stamp is applied inside `addListener` itself, so every registered handler (the `_SSE_HANDLER_MAP` wrappers AND the directly-registered ones) feeds it from one place; the heartbeat's own listener is a no-op that exists **only** to be registered, since `EventSource` drops named events nobody listens for. ⚠️ The watchdog interval is cleared at the top of `connectSSE()` and nowhere else (its only teardown path); clearing it elsewhere stacks intervals. Recovery needs no new sync path: the reconnect re-runs `handleInit` → `_resetAllAppState()`. The forced reconnect logs one diagnostic line, because a middlebox that strips heartbeats presents as "silently reconnects every 45s". +**SSE staleness watchdog** (`computeSseStale()` in constants.js, `_checkSseStale()` + a 5s interval in app.js): an `EventSource` can stop delivering without erroring, so the client forces a reconnect when nothing arrives. ⚠️ The server keepalive must stay the named `sse:heartbeat` event (`cleanupDeadClients()`, sse-stream-manager.ts), never an SSE comment, which `EventSource` cannot observe; its no-op client listener must stay registered. ⚠️ Judge staleness only while `connected` and online (the loop breaker). ⚠️ The liveness stamp lives inside `addListener`. ⚠️ Clear the interval only at the top of `connectSSE()`, or intervals stack. → [architecture-invariants#sse-staleness-watchdog](docs/architecture-invariants.md#sse-staleness-watchdog) -**Z-index layers**: subagent windows (1000), split picker menu (1000, `.split-picker-menu`), plan agents (1100), mobile/tablet fixed header (1200, `mobile.css`), modals on ≤768px (1300 — must beat the fixed header or the modal close button is buried), log viewers (2000), connection-loss overlay (2500, above the fixed header and modals), image popups (3000), response viewer (5000, backdrop 4999), file-preview overlay (5100 — must outrank the response viewer, which can launch it; at its old 2000 a path clicked in the chat opened BEHIND the chat), toasts/path picker (10000+, deliberately above the preview), the custom-model center-status banner (10001, `.center-status-banner` — `[hidden]` must re-assert `display: none` over its own `display: flex`, same trap as `.home-sessions[hidden]`, or `dismiss()` leaves an invisible click-blocker dead centre on screen), the swap-confirm and context-warning modals (10010, `#customModelSwapConfirmModal`/`#customModelContextWarningModal` — must clear both the plain `.modal` z-index of 1000 and the center-status banner it can appear over), terminal touch-selection bar (900 — above terminal content and the local-echo overlay, deliberately BELOW floating agent windows so it can never cover their controls), local echo overlay (7). +**Z-index layers** (keep new overlays consistent with this stack): local echo overlay (7), terminal touch-selection bar (900, below floating agent windows), subagent windows + split picker menu (1000), plan agents (1100), mobile/tablet fixed header (1200), modals on ≤768px (1300, must beat the fixed header), log viewers (2000), connection-loss overlay (2500), image popups (3000), response viewer (5000, backdrop 4999), file-preview overlay (5100, must outrank the response viewer that launches it), toasts/path picker (10000+), custom-model center-status banner (10001; its `[hidden]` must re-assert `display: none` or `dismiss()` leaves an invisible click-blocker), custom-model swap-confirm/context-warning modals (10010). → [architecture-invariants#z-index-layers](docs/architecture-invariants.md#z-index-layers) **Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min). @@ -397,11 +401,11 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L ### SSE Event Registry -161 event constants in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). **Both must be kept in sync**, and `test/sse-registry-parity.test.ts` is the guard that pins it (currently exactly in sync, 161 = 161, no drift either direction). ⚠️ `hook:agent_working` is the one hook event with no Claude Code hook behind it — the DeepSeek status bridge reports it (see External CLI modes). The backend file's `@fileoverview` carries the per-category breakdown, including the two Web tab events. +Event constants live in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). **Both must be kept in sync**; `test/sse-registry-parity.test.ts` pins it. ⚠️ `hook:agent_working` is the one hook event with no Claude Code hook behind it — the DeepSeek status bridge reports it (see External CLI modes). The backend file's `@fileoverview` carries the per-category breakdown, including the two Web tab events. ### API Routes -~236 handlers across 27 route files in `src/web/routes/`: system (56), sessions (37), cases (34), files (17), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), approvals (4), readmymind (4), custom-model (6), reboot-restore (3), me (2), teams (2), tab-layout (2), search (1), hooks (1), clipboard (1), status-telemetry (1), voice (1 + the `/ws/voice/stream` relay), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details. +One module per domain in `src/web/routes/` (plus a barrel; `ls src/web/routes/` for the current list). Beyond the `/api` routes: the `/webview/:cap/*` proxy, the `/ws/voice/stream` relay and the terminal WebSocket. Each file has `@fileoverview` with endpoint details. **HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`). @@ -450,7 +454,7 @@ Raw `npx vitest` skips the config (and with it `setup.ts`); always use `npm test **Config**: Vitest with `globals: true`, `fileParallelism: false`. Timeout 30s, teardown 60s. `config/vitest.config.ts` is the everything-config behind `test:all`; `config/vitest.ci.config.ts` is the gate and derives its excludes from `config/test-suites.ts`, which is also what `vitest.browser.config.ts` and `vitest.perf.config.ts` derive their includes from — so the exclusions and the runners cannot drift apart. Keep shared options in sync across them. -**Tmux safety**: under vitest (`VITEST` env var, set automatically), `TmuxManager` no-ops ALL shell commands and becomes a pure in-memory mock — tests physically cannot create/kill/attach real tmux sessions (`IS_TEST_MODE` in `src/tmux-manager.ts`). Every docker IO path is no-op'd the same way. `Session` is test-gated too: instead of attaching a real tmux client, it spawns a raw-mode echo PTY (`TEST_PTY_SCRIPT` in `src/session.ts`), so integration tests get a live input/output loop that echoes each byte exactly once. `test/setup.ts` gives every test file a temporary `HOME`/`USERPROFILE` (all `homedir()`-derived state, `~/.codeman` and `~/codeman-cases` included, resolves into a per-file fixture; the Playwright browser cache path is preserved), and additionally strips `CODEMAN_PASSWORD`/`CODEMAN_USERNAME` (so auth state from the running instance can't leak into tests) and `CODEMAN_GESTURE` (a shell-exported gesture flag would flip render-injection assertions), and strips the three instance-selection vars `CODEMAN_INSTANCE`/`CODEMAN_DATA_DIR`/`CODEMAN_TMUX_SOCKET` (#356/#371; `test/test-env-isolation.test.ts` pins the list, and its STATIC half reads setup.ts so a dropped `delete` fails everywhere rather than only on a box that exports the var). ⚠️ `CODEMAN_DATA_DIR` is the one that matters: `getDataDir()` reads it as an ABSOLUTE override before it ever looks at `homedir()`, so one inherited from the shell (a second instance, a beta run, a shell left over from `codeman web -d`) bypasses the temp HOME entirely, and a bare suite run once overwrote the real `remote-hosts.json` with a route test's fixture. `os.homedir()` itself DOES follow `$HOME`, so the temp HOME is what redirects everything else; `CODEMAN_INSTANCE` must be stripped in the setup file and never in a hook, because `config/instance.ts` captures it into a module-level const on first import. Tests that delete case trees go through `safeRmHomeTree()` (`test/mocks`), which refuses any path outside the temp HOME, so a wrong anchor leaves a temp dir behind instead of deleting `~/codeman-cases`. ⚠️ Raw `npx vitest` without `--config` skips `setup.ts` and with it the temp-HOME isolation. +**Tmux safety**: under vitest (`VITEST`), `TmuxManager` no-ops ALL shell commands (`IS_TEST_MODE` in `src/tmux-manager.ts`), docker IO is no-op'd likewise, and `Session` spawns an echo PTY (`TEST_PTY_SCRIPT`) instead of attaching tmux. `test/setup.ts` gives each file a temp `HOME`/`USERPROFILE` and strips `CODEMAN_PASSWORD`/`CODEMAN_USERNAME`, `CODEMAN_GESTURE` and `CODEMAN_INSTANCE`/`CODEMAN_DATA_DIR`/`CODEMAN_TMUX_SOCKET` (pinned by `test/test-env-isolation.test.ts`). ⚠️ `CODEMAN_DATA_DIR` overrides the temp HOME, so never drop its strip; strip `CODEMAN_INSTANCE` in the setup file, never in a hook (captured at first import). ⚠️ Delete case trees only via `safeRmHomeTree()`. ⚠️ Raw `npx vitest` without `--config` skips `setup.ts` and its isolation. → [architecture-invariants#test-isolation-tmux-docker-and-home](docs/architecture-invariants.md#test-isolation-tmux-docker-and-home) **Ports**: Pick unique ports manually, 3150+. Search `const PORT =` before adding new tests. Never 3000 (the live instance). @@ -458,15 +462,15 @@ Raw `npx vitest` skips the config (and with it `setup.ts`); always use `npm test **Testing against the live instance**: prod is HTTPS-only on :3000 (`curl -sk https://localhost:3000/...`). ⚠️ `w1`/`w2`/`w3` are the user's REAL sessions — never send input to them. Create your own throwaway session (`POST /api/sessions` then `POST /api/sessions/:id/shell`; creation alone leaves `pid: null` and no pane), test against that, and `DELETE` it by exact id when done. -**Respawn tests**: Use `MockSession` from `test/mocks/index.ts` (defined in `test/mocks/mock-session.ts`). **Route tests**: `app.inject({ method, url, payload })` in `test/routes/` — no live port needed. **Mobile tests**: Playwright suite in `test/mobile/` (138 device profiles). Browser-testing infra and practices: `docs/browser-testing-guide.md`. +**Respawn tests**: Use `MockSession` from `test/mocks/index.ts` (defined in `test/mocks/mock-session.ts`). **Route tests**: `app.inject({ method, url, payload })` in `test/routes/` — no live port needed. **Mobile tests**: Playwright suite in `test/mobile/` (device profiles in `test/mobile/devices.ts`). Browser-testing infra and practices: `docs/browser-testing-guide.md`. ## Debugging ```bash -tmux list-sessions # List tmux sessions -curl localhost:3000/api/sessions | jq # Check sessions -curl localhost:3000/api/status | jq # Full app state -curl localhost:3000/api/subagents | jq # Background agents +tmux -L codeman list-sessions # Codeman's own socket (bare `tmux` shows the default one) +curl -sk https://localhost:3000/api/sessions | jq # Check sessions (prod is HTTPS-only; dev on :3000 is plain http) +curl -sk https://localhost:3000/api/status | jq # Full app state +curl -sk https://localhost:3000/api/subagents | jq # Background agents cat ~/.codeman/state.json | jq # Persisted state ``` @@ -482,6 +486,6 @@ Two constraints worth knowing before you touch them: the env-derived PTY buffer ## Scripts & Tunnel -**`install.sh`** (repo root, ~140KB) is the public entry point: `curl -fsSL <raw url> | bash` installs Node/tmux/git/build tools if missing, clones to `~/.codeman/app`, builds, and offers a systemd/launchd service. Since installer v2 (2026-09-20) it is **look, ask, work, done**: `preflight_detect` prints what is on the machine, every human step runs BEFORE the build (one consent for all missing packages, ONE sudo prompt kept warm by `sudo_session_start`, the AI CLI menu, the Tailscale login/operator/HTTPS-toggle preflight), then the clone/`npm install`/build/service/serve run unattended behind `run_step` spinners (output in `~/.codeman/install.log`, tail shown on failure), and `print_done_screen` ends on the URL with a terminal QR code (the `qrcode` package Codeman already ships). The network-access question is 3-way: **Tailscale** (loopback bind + `tailscale serve`, curl-verified end-to-end), **LAN** (0.0.0.0 + password prompt), or **local-only**; it preserves the existing binding on re-runs via `read_existing_binding()`, which also reads back `CODEMAN_BASE_URL`/`CODEMAN_PORT` from the unit, and the flag/env PRESET paths keep an existing password too (`${CODEMAN_PASSWORD:-$EXISTING_PASSWORD}`: `--lan --service` on a password-protected unit used to rewrite it open with the unauthenticated ack, found in review 2026-09-21); `--password`/`--port` flip `RECONFIGURE` so they reach the unit instead of taking the quiet update path. ⚠️ The env a hand-started `codeman web` needs (host, password, ack, base URL, port) is composed in ONE place, `start_command_hint` for the done screen's Start line and `export_bind_env` for main()'s `exec` branch, so the two cannot disagree; `stop_background_helpers` runs right before that `exec`, because exec skips the EXIT trap and the sudo keepalive (keyed on `$$`, which becomes the server's pid) would otherwise refresh the sudo timestamp for the server's whole life. Ctrl+C in the HTTPS-toggle poll is trapped for the poll only and skips Tailscale for the run rather than killing the installer. ⚠️ **Tailscale is two halves on purpose**: `tailscale_prepare` (question phase: preflight, the opt-in rename, the serve SHAPE) and `tailscale_apply` (after the build: the one serve command). The shape is decided up front because it can change the service unit: when `:443` root already belongs to another app, the default is a **sub-path** (`tailscale serve --bg --set-path /codeman <port>` + `CODEMAN_BASE_URL=/codeman` in the unit; measured 2026-09-20: serve STRIPS the mount prefix before proxying, Codeman's `stripBasePath` tolerates unprefixed requests, and `--base-url` is what makes the emitted URLs carry it, verified live through the maintainer's tailnet incl. hashed assets and SSE), else a second port (8443+), replace, or skip. `detect_tailscale_serve_url` recognises all three shapes (`ts_serve_find_port_mapping`). ⚠️ **Rename is opt-in and defaults to NO everywhere** (owner decision 2026-09-20: the tailnet name is the machine's SSH identity); `--name`/`install.sh name` do it, `--yes` and non-interactive never do. Serve config is keyed by the DNS name it was written under, so `tailscale_rename_node` takes OUR mapping down first (`tailscale_remove_our_mapping`, `TS_MAPPING_REMOVED_BY_RENAME`) and it is re-added under the new name; the previous name is recorded in `~/.codeman/tailscale-rename` so uninstall can offer it back. Tailscale state is detected dynamically from `tailscale serve status --json` (no marker files); the installer must NEVER `tailscale serve reset`, touch a mapping it did not create, run `tailscale funnel` or advertise a Tailscale Service (a Service needs a TAGGED node + admin approval, so it is a docs hint only); `test/install-sh-invariants.test.ts` pins all of those plus the rename-before-shape order and the flag/header parity. A foreign `/Library/LaunchDaemons/com.codeman.web.plist` (the Mac mini's headless setup) is now LEFT ALONE rather than replaced by a LaunchAgent. Subcommands: `update`, `uninstall`, `tailscale` (retrofit), `name [<n>]`, `status` (the done screen again), `cloudflared` (the tunnel client, no longer a question in the main flow); flags `--tailscale|--lan|--local`, `--name|--no-rename`, `--service|--run|--no-start`, `--yes`, `--password`, `--port` pipe through `bash -s --` and set the same variables as their env twins (`CODEMAN_NONINTERACTIVE=1` approves system changes for automation and never installs Tailscale, renames or starts a service; `CODEMAN_TAILSCALE=1` presets the Tailscale choice). Design + verification record: `docs/installer-v2-plan.md`. Its CLI knowledge is a GENERATED block (`npm run generate:cli-catalog`, markers in the file), not a hand-written list: detection, the install menu and the closing reminder all read it, which is what stops the class of bug upstream `b6d0f1fa` fixed by hand (a user with only omp installed being told no AI CLI was found). ⚠️ It must stay **bash 3.2** clean — macOS ships it and the documented install is `curl | bash` under `set -euo pipefail`, so `declare -A`, `mapfile`, namerefs, `${x,,}` and here-strings are all fatal there; CI runs `bash -n` plus a real `bash:3.2` container, since expanding an EMPTY array under `set -u` is a runtime abort `bash -n` cannot see. ⚠️ It executes ONLY commands from the embedded block (`CLI_INSTALL_CMD_TRUSTED`) — there is no network fetch of the catalogue at install time to worry about at all. +**`install.sh`** (repo root) is the public `curl | bash` installer: it installs Node/tmux/git/build tools, clones to `~/.codeman/app`, builds, and offers a systemd/launchd service, asking every question BEFORE the unattended build (log in `~/.codeman/install.log`). Network access is Tailscale / LAN / local-only; re-runs preserve the existing binding AND password (`read_existing_binding()`). ⚠️ Compose the hand-start env only in `start_command_hint`/`export_bind_env`, and call `stop_background_helpers` before `exec`. ⚠️ Tailscale rename is opt-in (default NO, never under `--yes`/non-interactive); NEVER `tailscale serve reset`, touch a mapping it did not create, run `tailscale funnel` or advertise a Service (`test/install-sh-invariants.test.ts`). ⚠️ Stay **bash 3.2** clean (no `declare -A`, `mapfile`, namerefs, `${x,,}`, here-strings, empty-array expansion under `set -u`). ⚠️ Execute only commands from the generated CLI block (`CLI_INSTALL_CMD_TRUSTED`, `npm run generate:cli-catalog`). → [architecture-invariants#installsh-the-public-installer](docs/architecture-invariants.md#installsh-the-public-installer) Other key scripts: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh [quick|named] start|stop|status|url` (quick = random trycloudflare URL, default; `named setup|enable` = fixed-hostname tunnel via `scripts/codeman-tunnel-named.service`; bare `start|stop|url` still means quick), `scripts/run-beta.sh` (isolated beta instance), `scripts/build-agent-image.mjs` (docker base image), `scripts/self-update.sh` (detached updater). Production services: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel. diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index fba1243b..02412f19 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -16,6 +16,21 @@ Implementation detail extracted from `CLAUDE.md` so that file stays small enough **Instance isolation / multi-instance attach danger** — data dir (`~/.codeman`) and tmux socket (`tmux -L codeman`) are PROCESS-WIDE and shared by every Codeman on the machine, derived from `CODEMAN_INSTANCE` via `src/config/instance.ts` (`getDataDir()`/`dataPath()`/`DEFAULT_TMUX_SOCKET`). ⚠️ A 2nd instance on the SAME socket **discovers and attaches PTYs to the first instance's live sessions** (`tmux -L codeman attach-session …`), resizing/mutating them — `$HOME` isolation is NOT enough (tmux is system-global). To run two instances, give each a distinct `CODEMAN_INSTANCE` (scopes BOTH dir+socket: `~/.codeman-<name>` + `-L codeman-<name>`), or set `CODEMAN_TMUX_SOCKET` + `CODEMAN_DATA_DIR` individually. **`CODEMAN_INSTANCE` defaults to empty = the production layout (`~/.codeman`, `-L codeman`, port 3000)**, so this branch is safe to ship to master without disturbing existing installs. To run THIS beta alongside prod, launch with `scripts/run-beta.sh` (`CODEMAN_INSTANCE=beta` + `CODEMAN_PORT=5000`) — it never collides with prod's data dir/socket/port. Any new `~/.codeman/...` path MUST go through `dataPath()`, never `join(homedir(), '.codeman', …)`, and any new `tmux -L` caller through `resolveTmuxSocketName()` (same module): it applies the `CODEMAN_TMUX_SOCKET` override only when the name is safe and falls back to `DEFAULT_TMUX_SOCKET` otherwise. `TmuxManager` was the only such caller until `codeman tui` shelled out to tmux from a SECOND process for its degraded-mode listing and its attach handoff; a hardcoded `codeman` there would have pointed a beta instance straight at prod's panes. +### Reverse-proxy base path + +**Reverse-proxy base path** (`--base-url` / `CODEMAN_BASE_URL`, default `/`; `src/config/base-path.ts` is the pure single-source, normalized to `''` for root or `/foo`): lets Codeman be mounted under a sub-path behind a proxy that **forwards the prefix unchanged** (does NOT strip it). Deliberately few choke points, mirrored ingress/egress: + +- **(server ingress)** `stripBasePath()` runs inside Fastify's `rewriteUrl` so ALL routes stay declared prefix-agnostic (`/api/...`, `/ws/...`) — and a request arriving WITHOUT the prefix (hooks, health checks, docker bridge, all hitting the raw port) is left untouched, so the server answers at both. +- **(server egress)** one `onSend` hook prepends the base to every root-absolute `Location` header, covering all redirects. +- **(HTML)** `renderIndexHtml` rewrites the shipped `<base href="/">` to the mount and injects `window.__CODEMAN_BASE__` — the template's asset refs are all RELATIVE so `<base>` handles them for free. +- **(frontend runtime URLs)** root-absolute URLs ignore `<base>`, so `CodemanBase.url()` (constants.js) is the route builder, applied transparently by a `fetch` wrapper and explicitly at the few EventSource/WebSocket/`window.open`/`<img|iframe|a>`-src sites. +- **(sw.js/manifest)** the worker derives its base from `self.location`, the manifest uses relative `start_url`/`scope`. +- **(web-tab proxy)** `proxyPrefixFor(cap, basePath)` is the single base-aware root that cascades to the injected `<base>`, root-absolute HTML rewrites, the `runtimeUrlShim`, `Set-Cookie` Path and `Location` rebasing — while the INGRESS parsers (`capabilityFromProxyPath`, `resolveUpstreamUrl`) stay base-agnostic because `rewriteUrl` strips the prefix before routing, and `capabilityFromReferer(referer, basePath)` strips it from the browser-supplied Referer. + +⚠️ `--base-url` rides the daemon relaunch via `buildWebArgs` and the service unit via `resolveServicePlan`. + +Pure helpers unit-tested in `test/base-path.test.ts` + `test/webview-proxy.test.ts`; HTML injection in `test/render-index-html.test.ts`. + ## Session launch modes **Terminal colour env** — the stock registry decides each CLI's colour vars. Claude, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek and OMP export `COLORTERM=truecolor`; `shell` and `opencode` unset it. All of those except Claude also unset `NO_COLOR`, so a user who exports `NO_COLOR` globally keeps monochrome Claude panes. The variable matters because a CLI inheriting no `COLORTERM` quantizes every RGB color it draws down to whatever palette `TERM` alone implies, and the pane's `TERM` is not a constant. ⚠️ Codeman sets no `default-terminal` and passes tmux no `-f`, so it is tmux's default (`tmux-256color` since 3.2) unless the user's own `~/.tmux.conf` says otherwise, which Codeman's server DOES read. At `tmux-256color` supports-color reports 256 colors and `rgb(55, 55, 55)` lands on `ESC[48;5;237m`: visible, but not the color the theme named. Where `TERM` resolves to a 16-color entry instead (tmux older than 3.2, or a conf setting `default-terminal screen`) every dark background collapses to `ESC[40m`, the terminal's own black, and the block disappears entirely. That is why the same Claude theme looks different on two machines, and why a bug report here is worth pairing with the reporter's `tmux -V` and their pane's real `TERM`. ⚠️ These declarations reach the local tmux pane via `buildEnvExports()`, its attach client via `cliExportsTruecolor()`, and the direct-PTY fallback via `buildClaudeEnv()` — they do NOT reach a remote pane, which `buildRemoteLaunchCommand()` builds with no env exports at all. A Docker pane takes `COLORTERM=truecolor` from the hardcoded `envCreate`/`execEnv` in `tmux-manager.ts`, which apply to every mode including the two the registry says must unset it. A `~/.codeman/clis.json` override replaces these arrays wholesale (`deepMerge`), so a custom entry can drop either list. That is why both consumers apply the entry BEFORE Codeman's own variables rather than after: `buildEnvExports()` emits `...cliEnv` ahead of `export CODEMAN_MUX=1`, and `buildClaudeEnv()` assigns `PATH`/`TERM`/`CODEMAN_*` after its unset/export loop. Reversed, a config-supplied `unset` naming `CODEMAN_HOOK_SECRET_FILE` would strip it on one path and not the other. @@ -48,22 +63,58 @@ Model is NOT a session field: it is a composition entry in the profile's config **Codex input path (issues #218/#219/#220/#222)**: codex-mode sessions use **predictive write-through echo, never the buffer overlay**. The buffer overlay stays disabled exactly as 1.12.2 left it (`_updateLocalEchoState` in terminal-ui.js, same branch as shell; `_localEchoEnabled` remains false for codex), and the additive `_localEchoPolicy` field selects `'predict'` for codex when `localEchoEnabled` is on. Codex's composer is interactive per keystroke: typing "/" pops a live-filtering command picker (#222 was "picker never appears" because the "/" sat in the overlay until Enter), the composer grows/rewraps as it fills (#220: a long typed prompt existed ONLY in the overlay DOM, so codex never grew the composer), arrows and Ctrl+Backspace edit server-side state (#218: arrows were forwarded to an EMPTY composer while the typed text sat pending; the `\x08` control-char flush then left the overlay stateless so `\x7f` was swallowed as "nothing to remove"), and pastes arrive bracketed (#219: `terminal.paste()` wraps in `\x1b[200~..201~`, which the multi-byte-ESC branch forwarded WITHOUT flushing pending text, so the paste landed before it). The shared overlay branch (claude/gemini/opencode still buffer) gained three fixes: bracketed pastes flush pending text first, composer nav keys (`isComposerNavKey` allowlist in `CodemanTerminalInput` — arrows/Home/End/Delete/PgUp/PgDn incl. modifiers, deliberately excluding DA/CPR/DSR query responses) flush and hand the session to **pass-through** (plain PTY echo until Enter/Ctrl+C, because after cursor movement the append-only overlay cannot track edits), and a backspace that finds no overlay state is FORWARDED instead of swallowed. ⚠️ **Codex drops keystrokes that arrive in the same PTY read as a bracketed paste** (upstream `bottom_pane/paste_burst.rs` holds rapid chars for paste classification; verified against codex 0.147.0 by writing `hello\x1b[200~PASTED\x1b[201~` into the tmux client PTY in one write → composer shows only `PASTED`, while a 100ms gap yields `helloPASTED`), so the flush sends the typed text immediately and delays the paste sequence by 80ms — the same two-phase shape as the Enter branch's delayed `\r`. Related protocol fact: xterm.js sends `0x08` for Ctrl+Backspace, which codex's keymap binds to delete-ONE-char (`ctrl(Char('h'))`); real word-delete needs the kitty CSI-u encoding (`\x1b[127;5u`), which xterm.js 6.0.0 cannot emit (kitty support lands in 6.1.0-beta) — an upstream limitation, not a Codeman bug. E2E technique: codex 0.147 reaches its composer with any dummy key in `$CODEX_HOME/auth.json` (`{"OPENAI_API_KEY":"sk-test-..."}`), so a real TUI can be driven headlessly (envOverrides `CODEX_HOME` rides the `CODEX_*` allowlist) without real credentials. **Predictive write-through echo invariants** (the codex echo mode, `PredictiveEchoAddon` in `packages/xterm-zerolag-input`): (1) the onData hook `_predictHookOnData` is a PLAIN STATEMENT between the buffer block and Normal Mode — no `return`, try/catch-wrapped, never touches `_pendingInput` — so the wire path is byte-identical with the predictor active, absent or throwing (pinned at vm level and by an end-to-end trace-equality E2E); (2) it ships as a SEPARATE bundle `vendor/xterm-predictive-echo.js` so the zerolag bundle stays byte-identical, and a missing/broken bundle degrades codex to plain 1.12.2 echo (`typeof PredictiveEchoOverlay !== 'undefined'` guard); (3) predictions paint only while the cursor sits on the measured composer row (`isCodexComposerRow`, `CODEX_COMPOSER_ROW_RE = /^› /` — matches the empty-composer placeholder, typing, and the slash picker; rejects modal rows and 2-space wrapped continuation rows, the #220 ghost zone, which deliberately fall back to real echo); (4) reconciliation reads the PARSED buffer with `baseY + row` (xterm's `cursorY` is baseY-relative; `viewportY` only coincides while scrolled to bottom), confirms prefix-only on cell match PLUS cursor advance, cascades only on TWO consecutive foreign NON-BLANK passes (blanks are neutral: codex clears its placeholder on first echo), and TTL-bounds the rest; (5) after an UNPREDICTED wire edit (backspace into echoed text, any 'clear'-classified input, an IME/plain-paste 'text' commit, or every bypass send incl. `_handleCjkInput`) the addon holds new predictions until the next PARSED write: the displayed cursor is stale for one RTT and anchoring on it paints ghosts one cell off; (6) the per-device `localEchoEnabled` toggle is the kill switch returning exact 1.12.2 behavior. Measured constants + fixtures: `docs/predictive-echo-plan.md`, recorded via `scripts/dev/record-codex-frames.mjs` through the production tmux+strip pipeline. Tests: `test/local-echo-codex-gating.test.ts` (vm harness: nav-key + predict classifier truth tables, policy matrix, wire-neutrality pins), `packages/xterm-zerolag-input/test/` (addon laws, real-fixture replay, seeded fuzz), `test/codex-predictive-echo.test.ts` (E2E vs real codex incl. byte-identity + 300ms-RTT). +Further detail: with the registry, a CLI that declares its own composer glyph and working line in `capabilities.workDetect` gets the same screen-probed idle confirmation claude gets (see the ❯ note in CLAUDE.md's idle-detection paragraph), and one that declares neither keeps the output-stabilization behaviour. The tmux-only rule now covers all EIGHT external modes (OMP included): none has a direct PTY fallback, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. + +Grok's permission clamp: the Run button sends `grokConfig.alwaysApprove: true` (like antigravity's bypass), and the clamp's only-if-sent branch strips it for non-granted owners in multi-user mode. Grok lands on the `'buffer'` echo policy via the `_updateLocalEchoState` fallthrough, which is UNMEASURED against a live authenticated session; if its composer turns out per-keystroke-reactive like codex, flip it to the `'off'` branch. Grok's user guide is `docs/grok-integration.md`; the resolver's `grok --version` probe is surfaced (path + version) by `GET /api/grok/status`. + +DeepSeek: the Run button gates on `isDeepSeekRunnable()` and a `web`/`headless` profile is refused at spawn because it cannot drive a pane. The status bridge gives dsh real `stop`/`blocked` signals, real Approvals Inbox items, and the `agent_working` event that clears an alert answered in the terminal. `clampEnvOverridesForOwner()` drops (rather than rewrites) the three keys, and dropping falls through to what `_configureCliEnv()` exports, which is the clamped value; every OTHER CLI's bypass is a command-line flag reachable only through its config, which is why the config clamp alone is the whole gate for them. + +OMP's registry entry declares `privilegedParams: []` and a launch spec that only ever passes `--model`/`--resume`/`--continue`. + ### Remote sessions over SSH **Remote sessions (SSH)**: Sessions can run the agent inside a durable `tmux -L codeman-remote new-session -A` **on a remote host** so it survives the SSH drop (COD-104), and can also **discover + attach** to `codeman-*` sessions another Codeman launched there — attached (`owned:false`) sessions **detach, never kill** on tab close (COD-105). **Shared/collaborative** (COD-106): remote set-options are scoped per-session (never `-g`) and `window-size latest` lets multiple clients attach the same session at different viewports without clamping to the smallest; a client count surfaces a "shared · N" badge. **Auto-reconnect** (COD-108): a bounded-backoff watcher re-establishes a dropped remote session's local ssh pane and reattaches the still-running durable remote tmux (kill-switch `remoteAutoReconnect`, default ON); the pure pieces (backoff schedule, per-session reconnect state, `decideReconnect` eligibility) live in `src/remote-reconnect.ts` (tests: `test/remote-auto-reconnect.test.ts`), while `tmux-manager.ts` owns the live pane probe + timers. Owned sessions propagate `kill-session` to the remote on close; non-owned never do. ⚠️ Command-injection surface (COD-107): all ssh command lines flow through the single shell-safe `buildSshConnectionArgs()` — every user field (`-J jumpHost`, `-i identity`, `-o`) is `shellescape`d; never hand-build an ssh line elsewhere. Full design: `docs/remote-sessions.md`. +Further detail: ⚠️ **Auto-reconnect revives ONLY when the durable remote tmux session is verifiably still alive** (`remoteTmuxSessionAlive()`, a `has-session` probe over ssh, #355): a clean agent exit (Ctrl-C, Ctrl-D, `exit`) tears that session down, and `isPaneDead()` cannot tell it from a transport drop, so the watcher used to relaunch a FRESH agent after every clean exit (claude only looked fine because its `|| --resume` fallback masked it). An unreachable host answers `undefined`, which also means do not revive. ⚠️ `has-session` prints NOTHING on success, so the probe is classified by EXIT STATUS (`classifyRemoteAliveExit`: 0 alive, ssh's 255 or a timeout unknown, anything else gone); reading stdout classified every live session as gone and silently disabled transport-drop reconnects. The answer is cached per session and forgotten whenever the pane is seen alive again, or a stale `true` from one transport drop would revive the next clean exit. + +File reads in a remote case are the second ssh surface (see Remote SSH cases below). The reason writes, office previews and thumbnails are deliberately unsupported over ssh is that no remote file is ever copied onto the server's disk. ssh children are bounded because terminal output in a remote session is written on the remote host, so a prompt-injected agent there can print hundreds of `codeman://attach` links. + ### Remote SSH cases **Remote host wake-on-LAN from user input**: an optional `RemoteHost.wakeMac` (magic packet built and broadcast by Codeman) or `RemoteHost.wakeCommand` (a single executable path, run WITHOUT a shell, and the explicit override) lets the input route — and an explicit `POST /api/sessions/:id/wake` — wake a SLEEPING host instead of writing into a stalled ssh pane; `tmux send-keys` succeeds against a stalled pane, so the bytes used to vanish silently. The wake flow lives in `src/remote-wake.ts` and is reachable **only** from an EXPLICIT user request: `POST /api/sessions/:id/input`, that explicit wake route, and the create/attach path (`POST /api/quick-start` for a remote case, `POST /api/sessions` with `attachRemoteSession`, via `ensureHostAwake`), because "the user pressed Run on a sleeping host" is the same kind of request and the tmux probe would otherwise fail with a misleading "needs tmux installed". Everything TIMER-driven must never wake a host: the COD-108 auto-reconnect watcher, `Server.handleRemoteSessionDropped` and boot recovery have no access to the registry, or a host would be re-woken seconds after each suspend and could never stay asleep (asserted by wiring guards in `test/remote-wake.test.ts`, not just documented — including that `ensureHostAwake` is called from the HTTP route only, since `cron-service.ts` builds sessions through the shared service with nobody waiting on the answer). `GET /api/sessions/:id/reachability` only ASKS — it never wakes — and feeds the amber "host unreachable" banner (`host-wake-ui.js`) whose action is either Wake or, with no target configured, "Configure WoL" → `#wakeConfigModal` (saved via `PUT /api/remote-hosts/:id`). Detection is a throttled bare TCP probe (no ssh, no `ServerAliveInterval` — keepalives would move bytes into an idle connection every interval; and a host behind a jump host/SOCKS proxy is reachability-UNKNOWN, never "asleep": `isProbeable()` keeps the registry from buffering, gating or bannering on a probe that cannot reach it), input is buffered and flushed in order after `reattachRemote()` (the send-and-wait path blocks instead, as does the create path, with a shorter request budget), and the wake fields are re-read from `remote-hosts.json` on recovery AND (throttled, cached) live for a running session, because the persisted `remote` snapshot would never see a field added later (`rehydrateRemoteHostFields` + `RemoteWakeDeps.resolveRemote`). Design + invariants: `docs/remote-sessions.md` §Wake-on-LAN from user input. **Remote SSH cases** (COD-94/#145): cases can point at a **remote host** (`~/.codeman/remote-hosts.json` + `remote-cases.json` via `src/remote-hosts.ts`; CRUD under `/api/cases` — cases route file). A remote session launches a LOCAL tmux pane running `ssh <host>` that creates a durable REMOTE tmux session on a **dedicated socket** `-L codeman-remote` with name `codeman-ssh-<id>` — deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so a Codeman instance on the target host never adopts it; no `-g` global tmux options are set remotely. `remotePath`/`identityFile` are schema-guarded against shell injection (backticks/`$` rejected — same approach as `extraSshOptions`); remote tmux availability is probed via `checkRemoteTmuxAvailable()` in quick-start (ssh args carry `-o ConnectTimeout=10`). Remote claude defaults to an idempotent `claude --session-id <id> || claude --resume <id>` pair under a login shell, so a respawn or reattach continues the SAME conversation rather than starting a fresh one (remote omp gets the same treatment via `--continue`; ⚠️ because the claude arm is an `a || b` pair under `-c`, that pane's PID is the login shell, not the agent); per-host `commands.*` override. Session kill best-effort kills the remote tmux too. `SessionState.remote`/`MuxSession.remote` round-trip through recovery (`restoreMuxSessions` passes `remote` back into the Session constructor). ⚠️ Run flows must route remote cases through `POST /api/quick-start` (which resolves the remote case and skips LOCAL CLI availability gates) — `POST /api/sessions` stat-validates `workingDir` locally and has no `caseName`. `envOverrides`/`effort`/`modelOverride`/`codexConfig`/`geminiConfig` are rejected for remote quick-starts (not silently dropped). UI: Create Case modal → Remote tab. Tests: `test/remote-hosts.test.ts`, `test/remote-ssh-options.test.ts`. ⚠️ **Reading a file in a remote case goes over ssh too** (#415): `src/remote-files.ts` is the single remote-READ layer (`buildRemoteFileCommand` = `buildSshConnectionArgs` + one shellescaped remote command; `remoteProbePaths` returns remote realpath + stat; `remoteCreateReadStream` streams a `Range` via `tail -c +N | head -c L` and its `close()` must be wired to the response's `close` or the ssh child outlives an aborted download). The guard order matches the local path exactly (`validateSessionFilePathLexical` → remote realpath of BOTH file and workspace root → containment → sensitive-path → size cap on the REMOTE size), a request path arrives from the browser and is only ever interpolated as a `shellescape`d token, and an unreachable host answers **502**, never a 404. ⚠️ The probe's symlink resolution FAILS CLOSED: `readlink -f` where it exists, otherwise a `cd -P`/`pwd -P` directory walk plus a bounded plain-`readlink` loop over the last component, and anything it cannot fully resolve is reported unresolvable (404), never as the unresolved string — the first version resolved the directory chain only, so on a host without `readlink -f` a `ws/notes.txt -> ~/.ssh/id_rsa` link passed containment under its own path while `cat` served the key. Records are NUL-separated and index-keyed so a newline in a filename cannot shift the mapping. ⚠️ ssh children are BOUNDED: probes and buffered reads go through `src/remote-ssh-limiter.ts` (a `document-conversion-limiter`-shaped semaphore, default 4), the attachment-history list probes its whole history in ONE batched call (`probeRemoteAttachmentHistory`, threaded into `registerExternalAttachment({remoteProbes})`), and probes chunk at 40 paths — a prompt-injected agent printing `codeman://attach` links in a remote session used to fork one `ssh` per link. `describeExecError` never returns Node's `Command failed: <ssh line>` message (identity path + probe script in a 502 body). The `PUT /file-content` guard sits AHEAD of `validateSessionFilePath`, which resolves LOCALLY, or a same-named local directory (an sshfs mount) takes the write. Under `VITEST` the three IO functions refuse rather than connect. This covers the ATTACHMENT routes too, which is the half a clicked path needs when the file is OUTSIDE the case directory (`_isExternalPreviewPath` sends it to `POST …/attachments`): registration, by-id `raw`, metadata and the history list all resolve over ssh (`registerExternalAttachment({remote})`, `resolveServableRemoteAttachment`), and what decides the host is the SESSION, never the path string — the same absolute path means a different file on each host. Deliberately NOT supported over ssh: writes (`edit=1`/`PUT` answer 400, `editable` is always false), office previews/thumbnails, the file tree/picker, `tail-file`. Tests: `test/remote-files.test.ts`, `test/routes/file-routes-remote.test.ts`. +Further wake-on-LAN detail: the wiring guards in `test/remote-wake.test.ts` are two; the second also pins that `server.ts` holds the registry for its LIFETIME only (`drop` on cleanup, `stop` on shutdown) and never calls a waking method. The create wake is wired in the route rather than the shared session service for exactly the no-timer-wakes reason. No `ServerAliveInterval`, because keepalives move bytes into an idle connection every interval and that is what a byte-threshold idle detector must not read as activity. For a host behind `jumpHost`/`socksProxy`/a `ProxyCommand` option, the probe connects to `host:port`, which such a host does not answer even while ssh works; the registry never gates create/attach on it (`'unprobeable'`), and `/reachability` answers `reachable: null, probeable: false`. The banner keys on a PROVEN `false`, and the banner's 30 s poller runs only for a host with a wake target (a timer connecting to a host Codeman cannot wake is the same timer-driven traffic the keepalive rule forbids). + +Input arriving during a wake is buffered: a chunk over 4 KB is dropped whole, never delivered as a fragment, and the route answers `{buffered:true}` / `{buffered:true, dropped:true}` so the caller can tell. It is flushed in order after `reattachRemote()` with `fromUser`; a flush write that fails drops the rest (logged) rather than retaining it for a wake hours later. Send-and-wait blocks instead and answers `OPERATION_FAILED` when the host never returns. ⚠️ In multi-user mode the attach path 403s a non-admin BEFORE the host is looked up: the wake runs an executable, and remote hosts are admin-only infra everywhere else. ⚠️ Browser keystrokes travel over the WebSocket, which deliberately does NOT pass through the registry (that is the hot path), so only the HTTP input path ever queues anything; the banner must not promise queued input for the Wake button. A request that waits on the wake (create/attach, and the button) uses the 40 s `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS`, not the 90 s session default, because the dashboard's reverse proxy cuts a request at its own 60 s `proxy_read_timeout`. The `remote:` SSE family is session-scoped in multi-user mode; a create/attach wake names its requester (`username`) since it has no session yet. `remote-wake.ts` refuses real IO under `VITEST` like `remote-files.ts`. + ### Docker cases **Docker cases** (shipped 1.4.0; user guide `docs/docker-cases.md`, design `docs/docker-cases-plan.md`): a case can point at a **container** instead of a local/remote path, and any of the CLI run modes runs INSIDE it. Like remote-SSH, it is a **LOCATION OVERLAY on cases, never a `SessionMode` of its own** (`SessionMode` is unchanged). Storage `~/.codeman/docker-hosts.json` + `docker-cases.json` via `src/docker-hosts.ts` (direct mirror of `remote-hosts.ts`: `readDockerHosts`/`readDockerCases`, `toSessionDocker`, `dockerDisplayPath`, and the PURE builders `buildDockerBaseArgs`/`buildDockerCreateArgs`/`containerApiUrl`/`hostGatewayAlias`/`dockerConfigHash`). CRUD `/api/docker-hosts` + `/api/cases/docker-link`, plus **one-click** `/api/cases/docker-quickcreate` (Create New "Run in Docker" checkbox → case folder in `CASES_DIR` + auto-provisioned shared `default` host + auto-start a session inside; an expandable Template picker Small/Medium/Large/GPU or any override creates a per-case `q-<name>` host), and export/import (`/api/docker-cases/:name/export`, `/api/docker-cases/import`, `GET/DELETE /api/docker-exports`) — all in `case-routes.ts`. Run flows route through `POST /api/quick-start` like remote (session-routes.ts docker branch, skips LOCAL CLI-availability gates). **Launch model**: exactly one long-lived container **per case** (`codeman-case-<slug>`, PID1 `sleep infinity` under `--init`); a LOCAL tmux pane runs `docker exec -it` into a **durable in-container tmux** on dedicated socket `-L codeman-docker`, session `codeman-dkr-<id8>` (deliberately fails `SAFE_MUX_NAME_PATTERN` so a Codeman running INSIDE the container never adopts it, exactly like remote's `codeman-ssh-<id8>`). Builders `buildDockerLaunchCommand`/`buildDockerKillCommand` in `tmux-manager.ts` (image-check → `docker inspect||create` → start → exec, all idempotent). The container is **shared by all sessions of the case**: `buildDockerKillCommand` kills ONLY that session's in-container tmux session, NEVER `docker stop` while siblings remain; `docker rm -f` happens only on case-delete (plus an instance-scoped boot reaper keyed on the `codeman.instance` label). **Two-layer durability/resume** (the central design point): (1) Codeman-PROCESS restart with the container still up → `tmux new-session -A` reattaches the SAME live agent (paneCommand ignored); (2) container stop/reboot/OOM → inner tmux is gone, so the re-run pane command resumes the conversation from the bind-mounted transcript: claude mode pins a DETERMINISTIC conversation id via `claudeDockerPaneCommand()` (`tmux-manager.ts`) — fresh launch `claude --session-id <sessionId> || claude --resume <sessionId>` (a duplicate `--session-id` exits 1 "already in use", so the fallback RESUMES after a container stop; verified CLI behavior), explicit resume `--resume <rid> || --session-id <sid>` so a stale id never dead-panes (leading `exec ` is stripped — an exec'd first branch could never fall back); codex `resume <id>` / gemini `--resume` keep `appendResumeFlag`. The resume id rides `resumeSessionId` through create/respawn options and persists on `DockerCase.lastClaudeSessionId` via `persistDockerCaseClaudeSessionId()` (written at quick-start launch, and again on hook/last-response conversation-id adoption so post-`/clear` switches track; seeded back when `resumeOnStart`, default true); `-A` makes the pane command self-selecting (inert on reattach, active only when tmux was re-created). **Config drift** (`dockerConfigHash` → `codeman.confighash` label): quick-start compares via `checkDockerConfigDrift()` and REFUSES a drifted launch with `CONFLICT`; the UI confirm calls `POST /api/docker-cases/:name/recreate` (refused while case sessions are live) which `docker rm -f`s so the next launch recreates with the new config — host config edits actually take effect. **Workspace** is a REAL host dir bind-mounted at the SAME absolute path (mirror, `dst==src`), so `Session.workingDir = hostWorkspacePath` keeps file-routes/attachments/watchers on real host bytes AND the in-container transcript projHash matches the host so subagent/workflow correlation (and thus resume-id capture) works; `resolveMuxAttachCwd` returns `/tmp` for docker (the local pane only runs `docker exec`). **Creds** arrive commit-safe and ISOLATED (1.4.1; replaced the whole-dir RW mounts that let in-container CLIs write refreshed tokens/state back to the host): shared RW across the boundary is ONLY what host-side reads/resume need (`~/.claude/projects` transcripts; codex `sessions/` + `history.jsonl` for response-viewer/`codex resume`); everything else is SEEDED (RO mount, copied into container HOME once at launch via `[ -e ] || cp`; the container refreshes its own copy and never writes back): `~/.claude.json` is merged through `buildSeamlessClaudeConfig()` (forces `hasCompletedOnboarding` + theme + workspace trust, so no login wizard/theme picker/trust prompt inside the container), plus `.claude/{.credentials.json,settings.json,stats-cache.json}`, plus whole-dir seeds for `~/.gemini`/`~/.config/{gcloud,opencode}`, plus, ONLY when `CODEMAN_AGENT_IMAGE_INSTALL_GH` / `_AZ` is `1` (`enabledByEnv`, read at container create), per-file seeds of the `gh` and `az` sign-ins (`~/.config/gh/{hosts.yml,config.yml}`, five sign-in files from `~/.azure`, never its logs or extensions), which the agent image's system git credential helpers read for github.com / Azure DevOps (`resolveDockerClaudeArtifacts`/`resolveDockerCredentialArtifacts` in `docker-hosts.ts`). Bind mounts are physically excluded from `docker commit`, so exports stay secret-free; API-key CLIs get exec-time NAME-ONLY `--env OPENAI_API_KEY` (no `=value`); the SEALED profile is `mountCredentials:false` + `network:none`. NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket. **Hardening** on every create: `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit`, `--memory`==`--memory-swap`, non-root via `--user <hostUid>:0` (Linux, GID 0 for writable HOME) / `--userns=keep-id` (podman rootless) / baked uid (Docker Desktop), `--pull=never`, `--init`. Base image `codeman/agent:base` is BUILT LOCALLY from `docker/agent.Dockerfile` (node22 + tmux + every enabled npm CLI from the registry, since the `CLI_NPM_PACKAGES` build arg is generated from `stock.ts` and entries carrying `agentImageLayer` or no `npmPackage` get their own layers, see `docs/docker-cases.md`; OpenShift arbitrary-uid HOME, `C.UTF-8` locale so tmux/Ink render real box-drawing glyphs; Codeman also sets `LANG`/`LC_ALL` at run time for containers built before that line) via `scripts/build-agent-image.mjs` OR **auto-built on first use** (1.4.1: `ensureAgentBaseImage()` in `docker-hosts.ts`; idempotent + concurrency-safe, only the DEFAULT image ref is ever auto-built, `--pull=never` stays absolute; build output streams over SSE `docker:imageBuildStarted`/`imageBuildProgress`/`imageBuildComplete`/`imageBuildFailed`, and quick-create returns `imageBuilding:true` while the first launch awaits the gate); tmux-in-image is a HARD gated prerequisite (`checkDockerTmuxAvailable`), never a silent bare-exec fallback. **Hooks + model**: the workspace-scaffolding block DOES run for docker (writes `.claude/settings.local.json` + the CLAUDE.md scaffold into the real host dir), so `modelOverride` works via `settings.local.json` — it is a `QuickStartSchema` field applied for local AND docker quick-starts (`updateCaseModel`), sent by the frontend docker run path (the one deliberate difference from remote, which rejects it); `effort`/`envOverrides`/`codexConfig`/`geminiConfig`/`openCodeConfig` stay rejected. In-container hook curls hit `containerApiUrl(process.env.CODEMAN_API_URL, engine)` (swaps ONLY the hostname to the gateway alias, preserving scheme+port so prod HTTPS still works); the host guard allowlists both `host.docker.internal`/`host.containers.internal` (`DOCKER_HOST_GATEWAY_ALIASES` in `network-auth-policy.ts`). ⚠️ On a **loopback-only** bind (the prod default) a container cannot reach 127.0.0.1, so in-container hooks fire ONLY when `CODEMAN_DOCKER_BRIDGE_HOOKS=1` — an opt-in SECOND listener on the docker bridge gateway (`_startDockerBridgeHooksListener` in `server.ts`; gateway auto-detected via `detectDockerBridgeGateway`, or set `CODEMAN_DOCKER_BRIDGE_HOST`) that serves ONLY the hook endpoints (403 for any other path) into the same secret-gated pipeline; otherwise idle detection falls back to output-based through the docker-exec PTY. Container-set `CLAUDE_CODE_TMPDIR` keeps claude launching regardless of workspace path. `SessionState.docker`/`MuxSession.docker` round-trip through recovery. Every docker IO path is `IS_TEST_MODE` (VITEST) no-op'd; the pure builders are unit-tested. **Export/import** (`src/docker-export.ts`): full-image (`docker commit` + `save | gzip` + workspace tar + manifest) or workspace-only → one portable `~/.codeman/docker-exports/<case>-<ts>.codeman-container.tgz`; import validates per-member sha256, traversal-guards the workspace tar, `docker load`s + quarantine-retags the image (`codeman/imported-<case>:<ts>`, never overwriting a local tag); a `saveImageToTar` stream `pipeline` avoids truncation. **GPU** passthrough (`gpus` → `--gpus`, needs the NVIDIA container toolkit) and **elastic disk** (no `--storage-opt` cap, so container storage grows with data). SSE `docker:exportComplete`/`exportFailed`/`importComplete` (both registries). **UI** in `session-ui.js`: Create Case **Docker** tab (collapsed/compact form since 1.4.1), the one-click checkbox + Template picker, short `(docker)` case-menu tags, and a Manage-tab Export button; docker AND remote sessions name their tabs `w<n>-<case>` via the shared `_nextCaseSessionStartNumber()` so all tabs follow one naming convention. **Adopting an already-running container** (`DockerCase.owned === false`, `POST /api/cases/docker-adopt` + the read-only `POST /api/docker-cases/adopt-preflight`, `GET /api/docker-hosts/:hostId/containers`, `POST /api/docker-cases/browse`): the mirror of remote-SSH's `owned:false` attach. For an adopted container the launch chain only LOOKS and then execs — no image gate (the image is theirs), no create, and above all no `start`, since starting a container we do not own is precisely the mutation adoption promises never to perform; a missing or stopped container fails closed with an actionable message. Credential seeding is skipped too (those copies read from create-time read-only mounts that do not exist here, and writing host credentials into someone's container is not ours to do), so its CLIs must already be authenticated inside it. Absent `owned` = owned, so every pre-existing case is byte-identical. ⚠️ The guarantee is NEGATIVE, so it cannot be observed by using the feature — only by asserting the mutating verbs are absent — and it is therefore enforced at four deliberately independent layers: `buildDockerStopCommand`/`buildDockerRemoveCommand` throw during pure STRING CONSTRUCTION (no shape of caller bug can produce a `docker stop`/`rm` for a container we do not own), `removeDockerContainer` refuses again at the lowest layer, `checkDockerConfigDrift` reports "none" (an adopted container carries no `codeman.confighash` label, so a real comparison would always report drift and the launch gate would 409 forever, offering a recreate we may not perform), and the orphan reaper skips it through a check independent of the two conditions that already cover it. ⚠️ **The export path is the one place that still touched the container** and both halves had to be closed: a full export `docker commit`s it (refused for an adopted case — it packages someone else's container, with their logins, into a bundle Codeman hands out) and even a workspace-only export `docker pause`d it first for snapshot consistency (skipped: the freeze stops the owner's processes for as long as the tar takes). ⚠️ `owned` is applied AFTER the config hash; `dockerConfigHash` takes an explicit field list, so ownership can never shift an existing case's hash and mass-trip the drift gate, whose only remedy is "recreate the container". ⚠️ The container workdir is verified INSIDE the container: it defaults to `hostWorkspacePath` for an OWNED case only because the create-time bind mount puts the host directory at that exact path, and adoption mounts nothing, so the two are independent facts — without the check `docker exec --workdir <missing>` fails with an OCI chdir error the pane surfaces as a bare `execvp failed`. ⚠️ Run-mode availability comes from the CONTAINER (`availableModes`), live-probed rather than trusted from attach time: gating the dropdown on host CLI availability (#201) is right for local sessions and wrong here. The probe modes and the BINARY each mode looks for both come from the CLI registry (`enabledCliIds()` / `discovery.binaries[0]`), never a local table — a hand-written list silently froze once already, missing `omp` and hiding that mode on every docker case; `antigravity` ships as `agy` and `deepseek` as `dsh`, so a mode-name probe reports both missing on a container that has them, and a mode with no binary (`shell`) is reported available without a lookup. ⚠️ **A FAILED probe means opposite things per ownership.** For an adopted case it is a real fault (only the user can start that container). For an owned case it is the NORMAL state before the first session — the launch chain creates the container on demand — so recording it as an error hid every agent mode on every freshly linked Docker case behind "start it yourself first", for a container Codeman was about to create itself; `CaseInfo.docker.owned` exists on the wire so the frontend can tell the two apart. ⚠️ Claude launches WITHOUT `--dangerously-skip-permissions` when the container's exec user is root: Claude Code refuses the flag as root ("cannot be used with root/sudo privileges", still true in 2.1.261) and the refusal is visible only inside the container, so the pane just dies. Our base image runs a non-root user and never hits it; an adopted container's user belongs to its owner and is frequently root. Which flag to drop is a per-CLI fact, so it is `overlays.docker.rootCommand` in the registry rather than an id branch. ⚠️ **Admin-only in multi-user mode**, unlike `docker-link` right next to it: linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has already confined to the caller's space, while an adopted container's mounts are whatever its owner gave it — one mounting `/` hands the adopter a shell over the whole host, defeating exactly the workspace scoping that mode exists to enforce. The container listing and the in-container directory browser are gated with it (both are machine-level reads over containers belonging to anyone); the preflight is NOT, because the run menu probes it for every docker case, so it admits a non-admin only for a container already linked to a case they can access. Tests: `test/docker-adopted-container.test.ts`. Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/docker-export.test.ts`, `test/network-host-guard.test.ts`. +Further detail: ⚠️ An ADOPTED container may back SEVERAL cases at different in-container directories (`classifyAdoptContainerConflict` in `docker-hosts.ts`: an exact twin on the same container AND directory is refused, an owned container still backs exactly one case, and a container another user adopted is refused), which is what the Add Case panel's "copy an existing case" picker relies on. The wire carries `CaseInfo.docker.owned` ONLY when false, so the picker tests `=== false`, never truthiness. The export-path closures (refusing `docker commit`, skipping `docker pause`) are two lifecycle touches the original design missed and that are easy to re-introduce. + +### CLI registry + +**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. **No code outside `stock.ts` may branch on a CLI id**; behaviour that genuinely differs is either a capability field or a NAMED PROFILE selected by one (`profiles.ts`), and `test/cli-registry-no-id-branching.test.ts` fails the build if an id check reappears — it matches `===`, `!==`, `case '<id>':` and `[...].includes(mode)`, because an earlier `===`-only version let 36 negated branches survive the conversion (including a seven-mode Ralph chain whose own comment asked the next person to keep it in step with `isExternalCliMode()` by hand). A second CI-gated guard, `test/frontend-cli-no-id-branching.test.ts`, covers the two frontend files the run-menu consolidation (#458) touched, `session-ui.js` and `mobile-overview.js`: its allowlist is keyed by expression rather than line number, with an occurrence count per entry, so a new branch reusing an already-approved expression fails as a count mismatch instead of riding in on the old approval. + +⚠️ Config contains no shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. + +⚠️ `external`, `hooks` and `altScreen` are three INDEPENDENT capabilities on purpose; deriving one from another shipped the `until=stop`-hangs-on-shell bug. + +⚠️ Three capability fields carry a REGEX from config (`discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`) and all three must compile through `compileVersionRegex()`, which caps length and refuses nested quantifiers; `workingLine` is the one that runs on the PTY hot path. + +⚠️ **`param` is TWO namespaces.** `launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a LAUNCH PARAM; the legacy `<Mode>Config` wire field is a separate namespace, bridged only by `launch.legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no load error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. Codex is the entry where the two names differ (`bypassApprovals` vs `dangerouslyBypassApprovals`) and therefore the one that catches a regression. + +⚠️ Six fields are DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured (except `accent`, measured against styles.css on 2026-09-21), so re-measure before wiring one up; the list is pinned so it cannot quietly grow. + +Spawn commands are pinned as literal strings in `test/cli-registry-spawn-golden.test.ts`, remote/docker pane commands in `test/location-overlay-commands.test.ts`. ⚠️ That second golden no longer covers **remote claude or remote omp**: both now have their own arm in `buildRemoteLaunchCommand` (a `--session-id || --resume` pair, and `--continue`, so a respawn continues the same conversation) and never reach `defaultRemoteCommandForMode`, which is what that test asserts. Their real pins are `toContain` substrings in `test/tmux-manager.test.ts` and `test/remote-shared-sessions.test.ts`; changing either arm will NOT fail the golden. + +⚠️ Anything reading the registry resolves it AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks) — a module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. `~/.codeman/clis.json` overrides any entry (read-only in this release; nothing writes it, so importing the registry has no filesystem side effects). User guide: `docs/cli-registry.md`. + ## Session data and lifecycle ### Input delivery and WS resilience @@ -90,13 +141,13 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do **Signal availability is decided by MODE, not by `isExternalCliMode()`.** `stop` and `blocked` come from Claude Code hooks, so `hooksAvailableForMode()` is true only for `claude`: `shell` is not an external-CLI mode but installs no hooks either, so `until=stop` on a shell session is a guaranteed unresolvable wait dressed up as a timeout. The behavior split is deliberate and must survive refactors: an EXPLICIT request for an unavailable signal is a 400 naming the mode, while the DEFAULT set (`stop,idle,exit`) silently drops them and echoes the narrowed set back as `wait.until`, because omitting the parameter must never 400. Signal quality is not uniform either: `stop` is definitive (Claude Code says the turn is over), `idle` is inferred from output stabilization plus prompt detection and can flap mid-turn when a spinner pauses, which is why `stop` is the documented default to orchestrate on and `idle` is the fallback for sessions that emit no hooks. ⚠️ **`idle` being ACCEPTED for a mode does not mean it ever FIRES there.** `startShell()` emits exactly one `idle` on a 500ms readiness timer and nothing afterwards, so a shell session sits at `status:'idle'` no matter what its pane is doing; since send-and-wait and `fresh=1` both require a TRANSITION, both can only time out on a shell worker (measured: a default `wait` on `sleep 4` burned its full 25s). The ❯-prompt and spinner detection that drives the real `working`/`idle` cycle is Claude's output format, so hook-less modes synchronize with `wait-output` markers or `exit`, and the docs must say so rather than listing `idle` as "available" and letting the reader infer it is usable. ⚠️ **Signals are edge-triggered with no history**, and that is a real orchestration limit: a signal that fires while no waiter is registered is gone, unobservable by any later wait variant (`until=stop` after the turn ended just times out, `fresh` or not — measured, R2-A). Documented client patterns must therefore register the waiter before the event can fire (send-and-wait) or gather on latched `wait-output` markers with `from=buffer`; the skill's fan-out flow was rewritten accordingly, and "fire-and-forget N prompts, then gather signal-waits sequentially" must never be documented again. The durable fix, a server-side latched last-signal-per-turn, is deferred with Part 3 of `docs/agent-control-plan.md`. ⚠️ Relatedly, the route corrects liveness that `SessionStatus` cannot express: `currentSignalFor()` answers `exit` whenever `pid === null` (exited, detached, or created-and-never-started) **or the mux pane is dead**, because `Session` parks a DEAD PTY at `_status = 'idle'` and trusting the status would answer the default wait `{signal:'idle', immediate:true}` for a crashed worker while `until=exit` blocked forever on an event that already happened. ⚠️ **`pid` alone cannot carry liveness for a tmux-backed session**: that pid is the local `tmux attach` CLIENT, so a worker exiting inside its pane leaves `pane_dead=1` with the client alive and `pid` never goes null — the `pid === null` branch is unreachable in the normal configuration (unit tests exercise it because `MockSession` sets `pid` by hand; only a live instance showed the gap). Liveness is therefore probed at the mux layer: `workerIsDead()` consults `mux.isPaneDead(muxName)` with a ~750 ms per-pane cache, ONLY on blocking waits (measured: 0 tmux execs across 100 non-wait input POSTs — the browser hot path pays nothing), plus a refcounted 3 s `watchForDeadWorker` interval so a worker dying while a wait is parked resolves it in ~3 s instead of burning the timeout. The probe fails SAFE by construction ("cannot tell" is never "dead": non-mux sessions, a missing or throwing `isPaneDead`, all return false). On send-and-wait, a "successful" `send-keys` into a dead pane additionally overrides `delivered` to `false` and rolls the dedup seq back (`undoOnFailure`), because the bytes went nowhere and a retry against a restarted worker must not be refused as a duplicate. The cost is that a just-created session reads as `exit`, which the wire docs must spell out as "not started yet"; the fix lives at the route rather than in `signalForStatus()` because the registry deliberately holds no `Session` reference. ⚠️ Two `claude`-mode cases still lose hooks for reasons outside the registry: a Docker case cannot reach a loopback-bound Codeman without `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, and a remote-SSH case runs the agent on another host whose hooks may never reach this server. Bounds are env-overridable (`CODEMAN_WAIT_MAX_MS`, `CODEMAN_WAIT_DEFAULT_MS`, `CODEMAN_WAIT_MAX_PER_SESSION`, `CODEMAN_WAIT_MAX_PER_OWNER`, `CODEMAN_WAIT_MAX_TOTAL`, `CODEMAN_WAIT_BUFFER_SCAN_BYTES`) and each is clamped to a hard bound so a typo degrades to the default instead of disabling the protection; they are internal tuning knobs like the rest of `src/config/`, NOT part of the SemVer-covered env-var surface in `versioning-policy.md`. Tests: `test/session-wait-registry.test.ts`, `test/routes/session-wait-routes.test.ts`, `test/routes/session-wait-output-routes.test.ts`, `test/routes/session-input-wait.test.ts`. -### The watching signal (a quiet pane that is not waiting for you) +Further detail: `stop`/`blocked` now fire for **`claude` and `deepseek`** (`shell` installs no hooks either); asking for one explicitly on any other mode is a 400 and the default set silently drops them. `deepseek` qualifies because the DeepSeek Harness TUI REPORTS idle/working/blocked to its supervisor and Codeman is that supervisor (`deepseek-status-shim.ts`), so its signals are definitive rather than inferred, and `hooksAvailableForMode()` in `session-wait-registry.ts` is the one place that rule lives (see the External CLI modes section for the per-session `statusReporting` wrinkle and the two dsh traps baked into the skill preamble: boot `idle` landing ~300 ms before the composer paints, and `sendwait` asking for `wait:"stop,exit"`). That is also why `deepseek` is the one non-claude mode the skill drives like claude: `spawn_workers alpha beta:deepseek` is a mixed fleet in one call, and `sendwait`/`last_text` need no variant. -**A session that armed a monitor, backgrounded a shell or started a background terminal ends its turn and goes quiet, and a minute later Claude Code's idle notification arrives.** Before this existed, that prompt became an approval item like any other, so every surface filed the session under NEEDS YOU with nothing for a human to answer. The CLI says which kind of quiet it is on its own screen, and reading that row is the whole mechanism: `capabilities.workDetect.watchingLine` (optional, per CLI) plus `watchingLines` (how many non-blank rows at the foot of the screen may hold it, default `WATCHING_TAIL_LINES` = 1). `_confirmIdle()` already captures the pane at the moment a turn ends, so `_readWatching()` runs `watchingLabel()` (pure, `session-activity.ts`) over that same capture; the label lands on `Session.watching` and rides `toLightDetailedState()` out to every payload. ⚠️ It is cached BESIDE `_lastPaneProbeWorking` and goes stale with it, because the probe returns its cached boolean without re-capturing inside `PANE_PROBE_MIN_INTERVAL_MS` and a label from a capture nobody took is a guess. ⚠️ A capture that FAILS clears the label (and announces the change) rather than keeping the last one: a stale label opens the next idle prompt already acknowledged, so keeping it would turn a failed `capture-pane` into a missed alert, while clearing it costs at most an alert the next readable capture takes back. ⚠️ It then FREEZES once `_confirmIdle()` concludes — nothing looks at the pane again until it produces output — which is correct rather than tolerable, since work ending repaints the pane either way (a monitor firing wakes the agent; codex drops its background-terminal row by itself); a timer to keep it fresh would spend a `capture-pane` per idle session per tick to learn nothing. ⚠️ A server restart looks like a hole in that and is not one: the field is live state and starts empty, but reconciliation re-attaches the pane and the attach repaint arms the idle confirmation, which probes and re-reads the label with no input from anyone (measured 2026-09-23, back within ~20 s). A restored session showing no label has no chip on its screen. +**Packaging as the `skills/codeman` agent skill.** The primitives are packaged as the `skills/codeman` agent skill, installable three ways: via `codeman skill install [--case <name>]` / `skill uninstall`; as a Claude Code plugin from the repo's own marketplace (`/plugin marketplace add Ark0N/Codeman` + `/plugin install codeman@codeman`, see the root-files note in CLAUDE.md); or auto-injected into a case's `.claude/skills/` on Claude session create behind `agentSkillEnabled` (SYNCED, default OFF). ⚠️ Plugin skills are namespaced, so a Claude Code holding BOTH the plugin and a user-level or per-case copy lists the skill twice, `codeman` and `codeman:codeman` (measured 2026-09-14: neither shadows the other, both work); the docs tell users to pick one route. -**The fix is the alert that does not fire; the badge is cosmetic.** `hook-event-routes` passes the label to `notePrompt()`, which opens the idle item ALREADY acknowledged (`acknowledgedAt` + `acknowledgedReason`). Nothing new suppresses anything: `acknowledge()` has always meant "the alert this prompt armed is spent", so the item stays pending, answerable and available as Read My Mind context, and a wrong label costs a card that does not blink rather than an alert that was never created. Every surface follows from that one flag — the broadcast carries `acknowledgedReason` so a live page declines to arm (`_onHookIdlePrompt`, settings-ui.js), the push is skipped, a reloading page reads `acknowledgedAt` in `seedApprovals()` as it always did, `classifySession()` and `pendingApprovalCount()` (tui-model.ts, tui-render.ts) ignore an acknowledged item, and the TUI card drops to the `info` tone and says why. It re-arms for free: the next idle prompt supersedes the item and is built fresh. ⚠️ Only `idle` is eligible, so a permission or question dialog still goes red whatever else the agent started — but a prose question is NOT a dialog, so an agent that arms a monitor and then asks "which branch?" in plain text is silenced along with the false alarms. That is the accepted cost of the design and the reason the kind gate sits at the single place items are created. +Injection is ADD-ONLY at create, marker-owned (`applyAgentSkill` in `hooks-config.ts` never touches an unmarked user copy) and refuses symlinks (this repo's own `.claude/skills/codeman` is a symlink to the source, which the injector must never write through). ⚠️ Claude Code loads a same-named USER-LEVEL skill (`~/.claude/skills/codeman`, written once by `codeman skill install` with no `--case`) over the per-case copy, and nothing used to refresh it: a stale Aug-9 user copy shadowed every fresh injection (2026-08-14: agents ran the old recipes, spawned workers serially and lost their lineage arcs), so session create now also refreshes a marker-owned user copy (`refreshUserAgentSkill`; refresh-only, never installs, foreign/symlink refused). -**⚠️ The label is pane-derived, so the window and the anchor are a trust boundary, not formatting.** An agent that gets its own text matched silences its own alert. Two things prevent it, and BOTH belong to whoever adds a pattern for a new CLI: the window must cover only rows that CLI draws, and the pattern must anchor on chrome only that CLI can produce. Claude satisfies both — its chip is the LAST row, so the default window of one row excludes even the status line directly above it, whose content comes from a `statusLine` command a bypassed session can write into its own `.claude/settings.json`. Codex does not: its row sits above the composer, and the slot it occupies holds the last row of the TRANSCRIPT whenever no terminal is running, so matching the complete row raises the bar without closing it. What contains that is `hooks: 'none'` — no hook event from a codex session reaches `notePrompt()`, so a forged label costs a wrong badge and nothing else, and a CLI that gains hook signals must not keep a pattern that soft. The label is also ANSI-stripped and capped (`MAX_WATCHING_LABEL_CHARS`) at the source, and every interpolation of it into markup goes through `escapeHtml()`, since a config-supplied capture group decides what it holds. Tests: `test/session-watching.test.ts` (the label and both CLI patterns), `test/watching-no-alert.test.ts` (the negative claim across all four surfaces). +Session create additionally pre-seeds the skill's §0 preamble cache (`seedAgentSessionPreamble` → `${XDG_CACHE_HOME:-~/.cache}/codeman-agent-<id>.sh`, local claude sessions only), single-sourced from `skills/codeman/preamble.sh` and pinned byte-identical to SKILL.md's §0 heredoc by `test/agent-skill.test.ts`, so the skill's bootstrap is a two-line loader instead of a ~150-line paste the model types out (~47 s of generation, measured live). ### Auto-resume on usage limit @@ -106,6 +157,8 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do **Plan-usage chip** (`showPlanUsageLimits`, per-device: desktop default **ON** since 1.9.3, handhelds OFF) renders compact Claude and Codex provider rows. Claude Code (v2.1.80+) pipes a JSON blob to a configured `statusLine.command` on each render; on Pro/Max it carries a `rate_limits` object (`five_hour`/`seven_day` windows only — no Opus weekly field — each `{used_percentage 0-100, resets_at epoch-SECONDS}`). ⚠️ **Injected as an EPHEMERAL `claude --settings` CLI flag at spawn (2026-09-07), never written to disk** — `resolveStatusLineCliCommand()`/`ensureStatusLineExporterScript()` in `hooks-config.ts` (`generateStatusLineCommand()`/`applyStatusLineConfig()` remain, but only as the legacy disk-write self-heal path: a workspace an older Codeman build touched gets its stale `.claude/settings.local.json` entry stripped the first time a session starts there again). The exporter WRAPS a user's own real statusline (`findEffectiveUserStatusLineCommand()`, walking Claude Code's own settings precedence) rather than replacing it, and POSTs the `rate_limits` blob to `POST /api/status-telemetry`. That route (auth-exempt like `/api/hook-event` — localhost-only, hook-secret-gated whenever auth is active, COD-91) parses via `usage-telemetry.ts` (pure, unit-tested), broadcasts SSE `session:statusTelemetry` (de-duped per session by `telemetrySignature` since the statusline fires on every assistant message), and returns a compact plain-text footer for the exporter to **print-through** (foreground POST in the no-wrap branch so its own stdout becomes the footer, printing NOTHING on failure, `curl -sfk` plus `|| true`, since a bare brand word on the statusline is what discussion #405 opened with; backgrounded — `>/dev/null 2>&1 </dev/null &`, closing stdin too — only in the wrap branch, where the user's own command owns the footer; `curl --max-time 5` bounds a hung, not just refused, Codeman). Main Codex subscription usage comes from the signed-in host CLI's read-only app-server `account/rateLimits/read` request at startup and every 5 minutes; `usage-telemetry.ts` selects only the main `codex` bucket (never model-specific buckets such as Spark), maps whatever 5-hour/7-day windows it supplies, and omits the provider row when unavailable. Credentials stay inside the CLI and no auth material is sent to the browser. `plan-usage-latest.ts` merges both process-wide sources and replays them in the SSE init snapshot (`getLightState`) so `#planUsageChip` renders immediately on page load/reconnect. `planUsageChipEnabled()` remains the single resolver behind the checkbox and chip visibility (DISPLAY only) — the SAME `showPlanUsageLimits` setting also doubles as the server-side telemetry COLLECTION switch, read FRESH from `settings.json` by `readPlanUsageTelemetryEnabled()` at every claude session create/respawn (`TmuxManager.createSession`/`respawnPane`), never cached, with no per-session field and no per-request wire field — applies uniformly across every claude-creation path (interactive Run, cron, Ralph Loop API, quick-start) and survives a Codeman restart by construction (nothing per-session to lose). ⚠️ An ABSENT key reads as ON, the same way an absent `workspaceHooksEnabled` does: the desktop chip already shows as on for an install that never touched the setting, and the exporter posts only to this Codeman over loopback. Resolving the default in the reader is what keeps `GET /api/settings` a plain read. It briefly reconciled the key on first read (persisting `true` when absent), but `readJsonConfig()` answers `{}` for ANY read failure, not only ENOENT, and every page load hits that route, so one unlucky read replaced the whole settings file with a one-key file; pinned by `test/routes/system-routes-settings-get-plan-usage-default.test.ts`. ⚠️ The client sends `showPlanUsageLimits` in a settings save ONLY when that save FLIPS the chip relative to what the device had (`planUsageCollectionFlip()` in settings-ui.js): the chip defaults OFF on handhelds, so sending it on every save let a phone saving its font size persist `false` and switch collection off for every desktop, whose chip then went stale with no error anywhere. An explicit toggle on any device still writes the switch. Injection covers LOCAL tmux-spawned claude sessions only: the non-tmux direct-PTY fallback (`Session.startInteractive` when tmux is unavailable) and the remote/docker pane builders do not carry the flag. Registry-gated on `getCli(mode)?.capabilities.statusLineTelemetry` rather than a hardcoded mode string. **Distinct from auto-resume** (which reacts to the Claude limit _message_; this proactively shows live percentages). Design: `docs/usage-limits-display-plan.md`. Tests: `test/usage-telemetry.test.ts`, `test/codex-plan-usage.test.ts`, `test/plan-usage-chip.test.ts`, `test/plan-usage-latest.test.ts`, `test/hooks-config.test.ts` (statusline exporter script + `readPlanUsageTelemetryEnabled`), `test/statusline-cli-flag.test.ts`. +Further detail: the handheld OFF default comes from the mobile block in `getDefaultSettings()`. `planUsageChipEnabled()` backs the two call sites that must never disagree (the App Settings checkbox and the chip's visibility). The absent-key-reads-as-ON rule mirrors `readWorkspaceHooksEnabled()`. + ### Cron jobs **Cron (cron-style `CronJob`s)**: saved, named jobs with a recurring schedule (`once`/`interval`/`daily`/`weekly`), enable/disable, Run Now, next-run calc, and per-job run history (`CronJobRun`). ⚠️ **Distinct from the legacy `ScheduledRun`** (`/api/scheduled`, a run-now duration-bounded autonomous loop) — the two never interact; the legacy concept keeps the `Scheduled*` names, the recurring-job feature is `Cron*`. `CronService` (`src/cron/cron-service.ts`) owns CRUD + the 30s background due-tick (`tickDueJobs`, registered via `cleanup.setInterval` in `server.ts`; `init()` recomputes nextRunAt on boot) and **reuses the existing session layer** (create → `addSession` → `setupSessionListeners` → `startInteractive`/`startShell` → prompt via `writeViaMux`/`write`) rather than rebuilding tmux logic. Next-run math is pure/unit-tested in `cron-time.ts` (SERVER-LOCAL timezone for daily/weekly). Dup-launch guard = `lastDueKey` (jobId:fireTime); schedule is advanced BEFORE launch so a slow launch can't re-trigger. `once` jobs self-disable after firing (`completedOnce`). Persisted via `AppState.cronJobs`/`cronJobRuns` (StateStore accessors). Routes `/api/cron/jobs*` + `/api/cron/runs` (`cron-routes.ts`, `CronPort`); schema `CronJobSchema` (cross-field `superRefine`; the `.partial()` update schema does NOT re-run it); SSE `cron:*`. Frontend `cron-ui.js` (#cronModal). Claude/shell/opencode/codex/gemini/antigravity/pi agent types. Tests: `test/cron-time.test.ts`, `test/cron-service.test.ts`. Design: `docs/cron-discovery.md`. @@ -114,6 +167,8 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do **Unified session list** (COD-160/#139): `GET /api/sessions/unified?limit=&q=` merges live sessions, persisted state, lifecycle-log history, and transcript files into one deduped list (pure core in `src/services/unified-session-service.ts`). ⚠️ **Transcript history is three stores** (#386): Claude's `~/.claude/projects`, omp's `~/.omp/agent/sessions` and codex's `~/.codex/sessions`. Rows are keyed by whatever id that CLI names the conversation with and folded into their owning session via the `claudeSessionId → Codeman id` alias map (resumed//clear-respawned sessions must not appear twice; the field keeps its Claude-era name and is not Claude-only). A codex row additionally carries `resumeId`, the rollout's own thread id, set by the scanner and never by a live session, which is what lets a row be resumed through `codexConfig.resumeSessionId` while a row without one stays a fresh session; the alias chain therefore includes `config.codexConfig?.resumeSessionId`, and a fresh codex pane is matched by `session_meta.originator` (`codeman_<sessionId>` for every pane Codeman spawns); lifecycle name/mode resolution is first-seen-wins (the log returns entries NEWEST-first). No terminal buffers in the response (unlike `/api/sessions`). Consumed by the Cmd+K Session Manager (#146). Session Manager polish (COD-162/#157, 1.6.0): **pinning** via `POST /api/sessions/:id/pin` (`session:pinned` SSE; killing a pinned session demotes it to a lightweight stopped record that stays visible/resumable, and cleanup skips pinned records); **cross-device tab order** via `PUT /api/session-order` (`session:orderChanged` SSE, persisted in `state.json`; pure `normalizeSessionOrder`/`mergeSessionOrder` in `src/session-order.ts`: pushing device wins, server-only ids fall to the end, never dropped); resume from the manager keeps the original session name (COD-143); `firstPrompt` is backfilled for sessions whose id != transcript UUID and the most recent prompt (`lastPrompt`) is shown + searched (COD-140/145). +Further detail: `resumeHistorySession()` is what sends `codexConfig.resumeSessionId` for a row carrying `resumeId` (a conversation already on disk), and a row without one is a genuinely fresh session. Every surface that re-projects these rows has to carry `resumeId` through, the phone overview included, or a tap on that surface silently starts a second conversation. + ### Session lineage lines (tab → tab it spawned) **The relationship did not exist before this** (1.17.0): `SessionState` had no `parentSessionId`, `quick-start` recorded only the multi-user *human* owner, and an agent's spawn call is plain `curl` from a tmux pane, so nothing in the request identifies the caller (`SO_PEERCRED` needs a unix socket; the API is TCP). The caller therefore supplies it — every managed pane already gets `CODEMAN_SESSION_ID` from `session-cli-builder.ts`. Two equivalent inputs, body wins: a `parentSessionId` field on `POST /api/sessions` / `POST /api/quick-start`, or the `X-Codeman-Parent-Session` header, which exists so the agent skill can set it ONCE on its shared curl invocation and have every present and future spawn recipe carry it. @@ -126,6 +181,10 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do ⚠️ **`data-agent-id="lineage:<childId>"` is load-bearing**, not a label: `_applyLineEntrances()` queries paths by that attribute, so tagging them this way is the whole reason the arcs get the draw-in animation AND its negative-`animation-delay` resume across `svg.innerHTML = ''` with zero new animation code. ⚠️ `.session-tabs` is `overflow-x: auto`, so a tab scrolled out of the strip still HAS a rect — one lying over the logo or the header buttons; edges with an endpoint outside the strip are SKIPPED (clamping would point at a tab that is not there), and a passive `scroll` listener re-anchors the rest, since a scroll moves both endpoints without firing any render. The incremental tab render also redraws when `_lineageEdgeCount > 0`: a badge appearing widens a tab and shifts every tab after it. Setting: `sessionLineageLines`, per-device (in `displayKeys`, absent from the `.strict()` `SettingsUpdateSchema`), desktop default ON. Tests: `test/session-lineage-lines.test.ts` (geometry), `test/routes/session-routes-parent-lineage.test.ts` (resolution + reject paths). +Further detail (current geometry and colors, superseding the dip numbers above where they differ): ⚠️ **ONE shape, and the second one was the bug**: every pair (flat strip or wrapped) gets a U-bridge hanging below the strip, anchored on both tabs' BOTTOM edges. ⚠️ The dip is a **mis-tuned-in-both-directions corridor** (44px cap = straight thread at strip-wide spans, #285; 104px cap + full row offset = ~106px over-bow into the terminal, 2026-08-15): it now hangs from the **STRIP's bottom edge** (fallback: lower tab bottom), capped at 64px, with NO per-row offsets stacked on top — the strip-bottom baseline is also what keeps a row-1 pair's arc from drawing through row 2's tab labels. + +⚠️ **Colors are keyed on the SPAWNING tab, not per child**: every arc leaving one tab is the same color however many workers it spawns, so the strip reads as "these five came from w1, those two came from w2" — per-child coloring gave one tab's own children a different color each, which is the distinction the colors exist to make. A child that spawns in turn is a parent in its own right and gets its own color for the arcs below it, so a chain changes color at each generation while each generation's fan-out stays uniform. Assignment cycles `CodemanLineage.COLORS` in first-seen order per parent id (first entry empty = the skin-tuned `--session-blue`, so the first spawning tab keeps it; the rest vivid fixed hexes), memoized rather than derived from draw index (the SVG is wiped and rebuilt constantly, so an index-based color would flicker), and set inline as `--lineage-color` so styles.css keeps owning opacity/glow/dash. `test/session-lineage-lines.test.ts` drives the real `_appendLineageConnectionLines()` and asserts the painted property, since testing the color function alone would pass just as happily with the child id passed back in. + ### Auto-named sessions (first prompt → tab title) **Shipped opt-in, in the prefix form, after a review round that found five ways the first cut named a tab wrong** (#376, 1.30.0). The contributed version renamed on EVERY prompt (`applyAutoName` never left the eligible state, so "fix the login bug" then "1" left the tab named **1**), fed its tracker from every write path (shell tabs renamed after each command, every Ralph and respawn tab named "Read @ralph_prompt.md and follow the instructions."), stayed in escape mode after a bare Esc until a byte in `0x40-0x7e` arrived (Esc then "fix the login bug" submitted **ix the login bug**, Esc then a CJK prompt submitted nothing and the following prompt lost its first character, Esc then digits grew the escape buffer to 19001 characters), treated the newlines inside a bracketed paste as Enter, cleared the draft on ANY CSI (including the SGR wheel reports Codeman forwards to claude ≥ 2.1.187, so "fix the " + wheel + "login bug" gave **login bug**) and on Tab (the `@` completer), and replaced the whole name, which dropped the case from the tab and reset `_nextCaseSessionStartNumber()` so every new session in the case became `w1-<case>` again. Each of those is a named rule in `session-auto-name.ts` with a test. @@ -140,12 +199,16 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do **Opt-in, and the listener orders its checks for cost.** The prompt lands in the tab name, `mux-sessions.json`, every `session:updated` broadcast, the TUI, both home screens and `/api/search` (which matches on `sessionName`), while Read My Mind deliberately keeps prompts 0600 and out of search because prompts can carry secrets; so `autoNameSessions` is synced and default OFF, like `agentSkillEnabled`, `approvalsInboxEnabled` and `readMyMindEnabled`. The listener checks `nameSource` and derives the title BEFORE reading `settings.json`, so an already-named session costs nothing per prompt. +Further detail: the `<prefix>: <title>` form (`w3-myapp: fix the login redirect`) is the one `parseSessionPrefix()` (app.js, #232) already renders as the title alone with the prefix in the tooltip, and that `_nextCaseSessionStartNumber()` still counts. The `name` setter is reached via `PUT /api/sessions/:id/name`. The listener lives in `session-listener-wiring.ts`. Tests: `test/session-auto-name.test.ts`, `test/session-listener-wiring.test.ts`, `test/routes/session-name-routes.test.ts`. + ### Full-scrollback replay **Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell full history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. **A capture reports the geometry it was taken at** (#435): a visible frame repaints each row at an absolute position, counting up to the pane's height and out to the pane's width, so a terminal smaller than that pane damages it two ways at once. Too short and every address past the browser's own height clamps onto the last line, overwriting the rows underneath (measured: against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew the survivors twice). Too narrow and each row is painted out to the pane's width, so the browser wraps every painted row and the wrap on the last one scrolls the whole frame up by one. Nothing in the response used to say what geometry the frame was built for, so the client could not see either case. `PaneCaptureOptions.capturedGeometry` carries it out, and the terminal response publishes it as `captureCols`/`captureRows`. ⚠️ **Both fields are ABSENT unless a frame was really positioned**, and every consumer must test `Number.isFinite` rather than truthiness: `mux-visible` is necessary but not sufficient, because when the `display-message` cursor query fails `capturePaneBuffer` skips the snapshot repaint and returns the raw capture, and the route still labels that non-empty body `mux-visible`. A body that positioned nothing has no geometry to describe and nothing to repair, so a comparison that fires there buys a second capture, a reset plus chunked rewrite, a dropped and reopened WebSocket and a discarded xterm snapshot for no gain. ⚠️ **The comparison runs on `mux-visible` ONLY.** A full-history body is linear scrollback closed by a RELATIVE cursor move, which is relative precisely so the browser's row count need not match the pane's, and a byte-history body carries no row alignment at all, so a size mismatch damages neither and a replay repairs neither. That gate matters because the first select of every non-shell session per page takes the full-history path, where an ungated comparison would fire most often on the one response it cannot help, at the price of a second whole-scrollback capture. ⚠️ **The replay is capped at one attempt and latches per session when it cannot converge.** `resizeRetry` stops two competing fits trading replays forever; a pane already drawing at the size just requested is left alone, because a retry would capture the identical frame; and a pass that still does not converge joins `_geometryRetryUseless`, so the case `Session.resize` declines outright (a small viewport while a desktop viewport's size claim is live, where the retry re-sends the same declined resize and captures the same pane) costs one attempt per session per page load instead of one per select. ⚠️ **The clamp used to manufacture that equality, and no longer can** (#464). This paragraph previously explained it as the signature of a clamp: `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` did not, so a terminal under 40 columns or 10 rows reported a pane permanently bigger than itself and would replay on every tab switch. That divergence is fixed at the source — `syncTerminalGeometry()` (terminal-ui.js) fits, floors and APPLIES in one step, so the browser terminal IS the size it reports. The only remaining reason the two can differ is a resize the server declined, which the server now reports back (`Session.ptyGeometry`, the `{"t":"zc"}` frame and the resize response) for the client to adopt by COLUMNS. The equality guard stays, for the plain case of a pane already at the requested size. ⚠️ **A retry pass must not re-arm `_fullHistoryLoaded`**: it did not consume the full-history pull, and re-arming it would spend a whole-scrollback capture on the next select. That branch is currently unreachable by construction, since reaching it needs `source === 'mux-visible'` while a `full=1` pass is answered `mux-full-history` or `history`; a static test over the source is the habit this repo uses for an invariant nothing can execute. Tests: `test/capture-geometry-retry.browser.test.ts` (eight cases, five of which fail against the merge base), `test/tmux-capture-full-history.test.ts`, `test/routes/session-routes.test.ts`. +⚠️ **A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery must verify itself** (`_scheduleDroppedOutputRecovery`, app.js): a hole in a TUI byte stream is a desynced cursor, which is muffled text (#464). It was a fire-and-forget 2s timer that nulled its own handle and then called `_onSessionNeedsRefresh()` — whose early returns (a buffer load in flight, a refresh already owning the session) are MOST likely to be true during exactly the burst that caused the drop, so the recovery was lost silently and the bytes were never replayed. `_onSessionNeedsRefresh` now returns whether it actually repainted, and the scheduler re-arms while it has not, bounded by `DROP_RECOVERY_MAX_ATTEMPTS` because the early returns it retries past are transient contention. A refresh that died at the capture fetch DEADLINE (it returns `'deadline'`) is a stalled link, not contention, and is NOT retried: each retry would be another `?full=1` capture waiting out a deadline of up to two minutes. The same debounce still collapses a burst into one attempt, and the `TERMINAL DROP` crash-trail line sits behind it, since one line per dropped frame evicted the whole 50-entry trail in under a second. Tests: `test/dropped-output-recovery.test.ts`, whose retry case and no-retry case only pin the fix as a pair. + ### Terminal scrollback: strip flavors and wheel/touch forwarding **Two strip flavors, one carry** (#205, `session.ts:_handleTerminalOutput`): the FULL strip (`isAltScreenStripMode` = codex/claude/gemini) removes alt-screen toggles, `3J`, and mouse-tracking DECSETs. Every other mode (shell/opencode/antigravity/pi) gets the NARROW strip (`isMuxAltScreenOnlyStripMode`) — alt-screen toggles ONLY — and only when tmux-backed (`useMux`). Rationale: the tmux CLIENT emits `smcup` as its first bytes at attach, before any program runs, parking xterm in the scrollback-less alternate buffer for the whole session (touch scrolling no-ops; xterm's own wheel handler converts the wheel to Up/Down arrows = readline history cycling — both #205 symptoms). tmux never forwards a pane program's alt-screen toggles to its client (it repaints instead; measured — vim/less inside a pane emit zero to the client), so the only thing the narrow strip ever removes is tmux's own smcup. It keeps `3J` (a user's `clear` is a deliberate scrollback wipe) and the mouse DECSETs (tmux passes those through even with `mouse off`; stripping them would break htop/vim mouse support). ⚠️ The `useMux` gate is load-bearing: `startShell()`/`startInteractive()` fall back to a DIRECT PTY when mux creation fails, and there the inner program's own `?1049h` really does reach xterm — stripping it would break vim/less/htop for real. The replay path (`session-routes.ts`, via `session.usesMux`) applies the same narrow branch; the frontend mirror (`_shouldReportMouseToCli()`) stays claude/codex/gemini because only the FULL strip touches mouse DECSETs. The chunk-boundary carry (`_altScreenSeqCarry`) runs for both flavors. Tests: `test/claude-scrollback-strip.test.ts`. @@ -166,13 +229,99 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do **Run launch synchronization**: the main Run entrypoint in `session-ui.js` holds an in-flight lock and disables `#runBtn` for the whole launch (at least 500ms), so a double click cannot create duplicate sessions with the same `w<n>-<case>` name. A successful create/quick-start also calls `_ensureCreatedSessionVisible()` before `selectSession()`: local creates use the response's full session snapshot; quick-start modes fetch `GET /api/sessions/:id` only when `session:created` SSE has not already populated the map. The normal `_onSessionCreated()` handler remains the idempotent upsert, so POST-first and SSE-first ordering both produce one immediately-rendered tab. Tests: `test/run-mode-ui.test.ts`. +Further detail, closing: ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. + ### Circuit breakers: Ralph and PTY-exit **Circuit breaker**: Prevents respawn thrashing. States: `CLOSED` → `HALF_OPEN` → `OPEN`. Reset: `/api/sessions/:id/ralph-circuit-breaker/reset`. **Distinct: PTY-exit breaker** (COD-115/118/#147, `session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits (crash loops on attach), blocks further auto-restarts, broadcasts SSE `session:respawnBreakerTripped` + push (in `PUSH_EVENT_MAP`). Reset ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive` (sent by the user-facing restart control) — the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. Sessions also scrub inherited `TMUX`/`TMUX_PANE` env so Codeman-in-tmux doesn't nest. Tests: `test/respawn-pty-breaker.test.ts`. ### An exited agent in a live pane (`paneExit`) -**Codeman creates every pane with `remain-on-exit on`, so a session whose agent exited still looks alive.** `/exit` ends the CLI, tmux keeps the pane and the tmux session, and the `tmux attach-session` process Codeman records as `Session.pid` runs on, so no PTY exit handler fires and the record keeps its pid and `status: 'idle'` (Ark0N/Codeman#446). `SessionState.paneExit` (`{status?, signal?, at}`) is the fact tmux already knows, published through `toState()` so it rides `session:updated` and lands in `state.json` on the same persist — there is no SSE event for it. One batched `tmux list-panes -a` per tick fills it, from `TmuxManager.startPaneExitWatcher()`, which has its OWN always-on interval: the stats collector cannot carry it, because the browser arms and disarms that one with the Monitor panel (`panels-ui.js`) and boot skips it entirely when no session was recovered. ⚠️ **The field is TRI-STATE and its third state is absence**, meaning UNKNOWN, which renders as nothing and must NEVER read as alive; it covers a running pane, a session the read did not list, a failed probe, and every session shape a dead local pane does not describe. `Session.paneExitApplies` is the single place that scoping lives, and it fails closed for four shapes: a direct-PTY session (no pane), a remote SSH session (the local pane is the ssh client, whose death is a transport drop OR an exit — the whole of #355), a docker case (the local pane is a `docker exec` into the container's own tmux), and a session rebuilt from the socket (`MuxSession.discovered`: its synthetic `restored-<fragment>` id matches no `state.json` entry, so a remote session rediscovered after `mux-sessions.json` was lost would arrive looking local). ⚠️ **Never set `status: 'error'`** for an exited pane — that value is the PTY-exit breaker's and the browser answers it with a "restart it?" confirm — and **never null the `pid`**, which is what makes `selectSession()` re-attach and launch a fresh CLI. Local panes keep `remain-on-exit on`; flipping them to `failed` ends the tmux session, nulls the pid and reintroduces the auto-revive #355 removed. ⚠️ **An absent `#{pane_dead_status}` is not 0**: measured on tmux 3.2a a SIGKILLed pane reports neither a status nor a signal (`#{pane_dead_signal}` did not exist before tmux 3.4), so folding it into 0 would turn an unexplained death into a clean exit. A session answers only when the read listed EXACTLY ONE pane for it, since Codeman never splits a pane and a session the user split by hand has none that speaks for the agent. The three synchronous `isPaneDead()` callers (the `/wait` route, the TUI, the attach path) keep their own probes — this watcher is never fresh enough for them. ⚠️ **The always-on timer gates the READ, never the tick.** `hasObservablePaneSession()` (`tmux-manager.ts`) skips the tmux exec while every session on the manager is one of the shapes `paneExitApplies` forces to UNKNOWN, so an instance running only remote or Docker work keeps ticking and costs nothing; the two predicates are two copies of one rule, and `test/session-pane-exit.test.ts` pins them against each other for the four session shapes that exist today — a FIFTH condition added to one and not the other still fails nothing, so change them together. Skipping retracts nothing, for the same reason a failed read does not. ⚠️ **The muted status dot is a specificity fight, and it is fought on three surfaces.** The tab renders `status` as before, and `tab-agent-exited` only quiets the dot, so the rule excludes three states BY HAND: `.tab-alert-action` and `.tab-alert-idle` on the tab, and `.tab-status.error` on the dot itself. Each of those colours means "this needs you" — the two alerts because a human is blocked, `error` because the browser answers it with a "restart it?" confirm — and each must survive the exit. The rich tab rail needs a SECOND copy of the rule, because its own `tab-state-*` dot rules are (0,9,1) against the strip's (0,5,0) — measured, an exited session on a detailed rail kept a full green dot and the working halo beside a badge reading "exited". Its twin matches that specificity exactly and therefore must stay BELOW those rules in source order. mobile.css needs a THIRD copy, with `!important`, because the phone block enlarges a `busy` dot and gives it a green glow that way, and `status` stays `busy` for a pane whose agent died mid-turn — without it a phone renders a grey dot still wearing the green halo. `test/session-pane-exit-ui.test.ts` resolves the real stylesheets in jsdom rather than matching selector text — styles.css for the desktop cases and both files for the phone ones — so the ordering, the hand-written exclusions and a missing phone rule all fail there. Tests: `test/session-pane-exit.test.ts`, `test/tmux-manager.test.ts`, `test/session-pane-exit-ui.test.ts`. +**Codeman creates every pane with `remain-on-exit on`, so a session whose agent exited still looks alive.** `/exit` ends the CLI, tmux keeps the pane and the tmux session, and the `tmux attach-session` process Codeman records as `Session.pid` runs on, so no PTY exit handler fires and the record keeps its pid and `status: 'idle'` (Ark0N/Codeman#446). `SessionState.paneExit` (`{status?, signal?, at}`) is the fact tmux already knows, published through `toState()` so it rides `session:updated` and lands in `state.json` on the same persist — there is no SSE event for it. One batched `tmux list-panes -a` per tick fills it, from `TmuxManager.startPaneExitWatcher()`, which has its OWN always-on interval: the stats collector cannot carry it, because the browser arms and disarms that one with the Monitor panel (`panels-ui.js`) and boot skips it entirely when no session was recovered. ⚠️ **The field is TRI-STATE and its third state is absence**, meaning UNKNOWN, which renders as nothing and must NEVER read as alive; it covers a running pane, a session the read did not list, a failed probe, and every session shape a dead local pane does not describe. `Session.paneExitApplies` is the single place that scoping lives, and it fails closed for four shapes: a direct-PTY session (no pane), a remote SSH session (the local pane is the ssh client, whose death is a transport drop OR an exit — the whole of #355), a docker case (the local pane is a `docker exec` into the container's own tmux), and a session rebuilt from the socket (`MuxSession.discovered`: its synthetic `restored-<fragment>` id matches no `state.json` entry, so a remote session rediscovered after `mux-sessions.json` was lost would arrive looking local). ⚠️ **Never set `status: 'error'`** for an exited pane — that value is the PTY-exit breaker's and the browser answers it with a "restart it?" confirm — and **never null the `pid`**, which is what makes `selectSession()` re-attach and launch a fresh CLI. Local panes keep `remain-on-exit on`; flipping them to `failed` ends the tmux session, nulls the pid and reintroduces the auto-revive #355 removed. ⚠️ **An absent `#{pane_dead_status}` is not 0**: measured on tmux 3.2a a SIGKILLed pane reports neither a status nor a signal (`#{pane_dead_signal}` did not exist before tmux 3.4), so folding it into 0 would turn an unexplained death into a clean exit. A session answers only when the read listed EXACTLY ONE pane for it, since Codeman never splits a pane and a session the user split by hand has none that speaks for the agent. The three synchronous `isPaneDead()` callers (the `/wait` route, the TUI, the attach path) keep their own probes — this watcher is never fresh enough for them. ⚠️ **A path that starts a command in a pane must clear the record AND persist**, since the watcher's next tick sees the field already cleared and writes nothing. ⚠️ **The always-on timer gates the READ, never the tick.** `hasObservablePaneSession()` (`tmux-manager.ts`) skips the tmux exec while every session on the manager is one of the shapes `paneExitApplies` forces to UNKNOWN, so an instance running only remote or Docker work keeps ticking and costs nothing; the two predicates are two copies of one rule, and `test/session-pane-exit.test.ts` pins them against each other for the four session shapes that exist today — a FIFTH condition added to one and not the other still fails nothing, so change them together. Skipping retracts nothing, for the same reason a failed read does not. ⚠️ **The muted status dot is a specificity fight, and it is fought on three surfaces.** The tab renders `status` as before, and `tab-agent-exited` only quiets the dot, so the rule excludes three states BY HAND: `.tab-alert-action` and `.tab-alert-idle` on the tab, and `.tab-status.error` on the dot itself. Each of those colours means "this needs you" — the two alerts because a human is blocked, `error` because the browser answers it with a "restart it?" confirm — and each must survive the exit. The rich tab rail needs a SECOND copy of the rule, because its own `tab-state-*` dot rules are (0,9,1) against the strip's (0,5,0) — measured, an exited session on a detailed rail kept a full green dot and the working halo beside a badge reading "exited". Its twin matches that specificity exactly and therefore must stay BELOW those rules in source order. mobile.css needs a THIRD copy, with `!important`, because the phone block enlarges a `busy` dot and gives it a green glow that way, and `status` stays `busy` for a pane whose agent died mid-turn — without it a phone renders a grey dot still wearing the green halo. `test/session-pane-exit-ui.test.ts` resolves the real stylesheets in jsdom rather than matching selector text — styles.css for the desktop cases and both files for the phone ones — so the ordering, the hand-written exclusions and a missing phone rule all fail there. Tests: `test/session-pane-exit.test.ts`, `test/tmux-manager.test.ts`, `test/session-pane-exit-ui.test.ts`. + +### Dead-pane respawn: the resume pin + +**The dead-pane respawn needs the same resume pin as a custom-model `restartCli()` and shares it** (`_buildRespawnPaneOptionsWithResumePin()` in `session.ts`, used by `restartCli()`, the dead-pane respawn in `_setupOrAttachMuxSession()`, and its create path when that path RELAUNCHES a CLI: after a failed respawn, or when tmux lost the whole session rather than the pane): a pane whose agent EXITED owns a transcript too, so recovering one with the bare launch line hit the same refusal and the conversation was stranded behind a tab that looked merely idle. The pin walks three candidates in order — the conversation chain's tail, the launch seed, then the session's own id — and takes the first one a transcript backs, never `_claudeSessionId` (which also holds history-correlated GUESSES keyed on the working directory, and launching from one would open and write to a conversation that was never this pane's). ⚠️ The create-path pin is written to `_resumeSessionId` as well, so unlike `restartCli()`'s one-respawn pin it PERSISTS through `toState()` as `resumeSessionId`: that field means "what the user asked to resume at creation, or what recovery pinned", and after a dead-pane respawn `_claudeSessionId` names whatever the walk actually pinned rather than the chain tail. ⚠️ A candidate no transcript backs is passed over, and falling off the end of the walk ADDS no pin (the options keep whatever launch seed they already carried): a divergent pin leaves `--session-id <this.id>` in the fallback branch, where a failed resume collides all over again, while pinning an id with no transcript prints claude's "No conversation found" into a brand-new session's scrollback and costs the running branch its `nice` priority (`wrapWithNice()` prefixes only the first branch of an `a || b`). ⚠️ Remote and docker sessions are never pinned: their pane commands are already self-healing, the conversation lives on the far side, and a local id resolves to nothing there. + +### Idle detection: composer glyph and working line + +⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer (`❯`) about once a second all through a turn, so the old "saw a ❯, wait 2s → idle" rule flipped every working session to idle two seconds in (measured: a session mid-tool-call at 17 minutes reporting `status:"idle"`). Its working indicator is `✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`: the glyph animates through `· ✢ ✳ ∗ ✻ ✽`, the gerund is randomized, and the finished line (`✻ Cooked for 2m 49s`) carries the same glyph, so neither `SPINNER_PATTERN` (braille, not what current versions draw) nor a keyword list can see it. Matching the new line in the STREAM does not work either: tmux ships partial repaints, so the whole line reaches the PTY only every few tens of seconds. + +So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks the SCREEN via `capturePaneText()` + `CLAUDE_WORKING_LINE_PATTERN` before believing it; a sustained run of repaints (`session-activity.ts`, pure + unit tested) is what marks a turn as started, with the same screen probe vetoing keystroke echo. Idle now lands ~3-5s after a turn ends instead of 2s into one. + +⚠️ **The composer glyph and the working line are per-CLI registry DATA** (`capabilities.workDetect`, #385), not Claude constants: claude declares `❯` plus the pattern above, codex declares `›` plus `[Ee]sc to interrupt`, and a CLI that declares neither falls back to Claude's pair, which is what every session used before the registry carried one. Before that, this whole mechanism was gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` for its entire life. + +⚠️ `workingLine` is config-supplied (a user `clis.json` can set it) and the compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()`: a nested quantifier there is a ReDoS against the event loop, and the helper returns null rather than throwing so the fallback is structural. + +⚠️ A third, optional `workDetect` field, `watchingLine` (plus `watchingLines`), tells a quiet pane that is still RUNNING something apart from one that wants a human; it is read from the same idle-confirmation capture. See [The watching signal](#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you). + +### The watching signal (a quiet pane that is not waiting for you) + +**A session that armed a monitor, backgrounded a shell or started a background terminal ends its turn and goes quiet, and a minute later Claude Code's idle notification arrives.** Before this existed, that prompt became an approval item like any other, so every surface filed the session under NEEDS YOU with nothing for a human to answer. The CLI says which kind of quiet it is on its own screen, and reading that row is the whole mechanism: `capabilities.workDetect.watchingLine` (optional, per CLI) plus `watchingLines` (how many non-blank rows at the foot of the screen may hold it, default `WATCHING_TAIL_LINES` = 1). `_confirmIdle()` already captures the pane at the moment a turn ends, so `_readWatching()` runs `watchingLabel()` (pure, `session-activity.ts`) over that same capture; the label lands on `Session.watching` and rides `toLightDetailedState()` out to every payload. ⚠️ It is cached BESIDE `_lastPaneProbeWorking` and goes stale with it, because the probe returns its cached boolean without re-capturing inside `PANE_PROBE_MIN_INTERVAL_MS` and a label from a capture nobody took is a guess. ⚠️ A capture that FAILS clears the label (and announces the change) rather than keeping the last one: a stale label opens the next idle prompt already acknowledged, so keeping it would turn a failed `capture-pane` into a missed alert, while clearing it costs at most an alert the next readable capture takes back. ⚠️ It then FREEZES once `_confirmIdle()` concludes — nothing looks at the pane again until it produces output — which is correct rather than tolerable, since work ending repaints the pane either way (a monitor firing wakes the agent; codex drops its background-terminal row by itself); a timer to keep it fresh would spend a `capture-pane` per idle session per tick to learn nothing. ⚠️ A server restart looks like a hole in that and is not one: the field is live state and starts empty, but reconciliation re-attaches the pane and the attach repaint arms the idle confirmation, which probes and re-reads the label with no input from anyone (measured 2026-09-23, back within ~20 s). A restored session showing no label has no chip on its screen. + +**The fix is the alert that does not fire; the badge is cosmetic.** `hook-event-routes` passes the label to `notePrompt()`, which opens the idle item ALREADY acknowledged (`acknowledgedAt` + `acknowledgedReason`). Nothing new suppresses anything: `acknowledge()` has always meant "the alert this prompt armed is spent", so the item stays pending, answerable and available as Read My Mind context, and a wrong label costs a card that does not blink rather than an alert that was never created. Every surface follows from that one flag — the broadcast carries `acknowledgedReason` so a live page declines to arm (`_onHookIdlePrompt`, settings-ui.js), the push is skipped, a reloading page reads `acknowledgedAt` in `seedApprovals()` as it always did, `classifySession()` and `pendingApprovalCount()` (tui-model.ts, tui-render.ts) ignore an acknowledged item, and the TUI card drops to the `info` tone and says why. It re-arms for free: the next idle prompt supersedes the item and is built fresh. ⚠️ Only `idle` is eligible, so a permission or question dialog still goes red whatever else the agent started — but a prose question is NOT a dialog, so an agent that arms a monitor and then asks "which branch?" in plain text is silenced along with the false alarms. That is the accepted cost of the design and the reason the kind gate sits at the single place items are created. + +**⚠️ The label is pane-derived, so the window and the anchor are a trust boundary, not formatting.** An agent that gets its own text matched silences its own alert. Two things prevent it, and BOTH belong to whoever adds a pattern for a new CLI: the window must cover only rows that CLI draws, and the pattern must anchor on chrome only that CLI can produce. Claude satisfies both — its chip is the LAST row, so the default window of one row excludes even the status line directly above it, whose content comes from a `statusLine` command a bypassed session can write into its own `.claude/settings.json`. Codex does not: its row sits above the composer, and the slot it occupies holds the last row of the TRANSCRIPT whenever no terminal is running, so matching the complete row raises the bar without closing it. What contains that is `hooks: 'none'` — no hook event from a codex session reaches `notePrompt()`, so a forged label costs a wrong badge and nothing else, and a CLI that gains hook signals must not keep a pattern that soft. The label is also ANSI-stripped and capped (`MAX_WATCHING_LABEL_CHARS`) at the source, and every interpolation of it into markup goes through `escapeHtml()`, since a config-supplied capture group decides what it holds. Tests: `test/session-watching.test.ts` (the label and both CLI patterns), `test/watching-no-alert.test.ts` (the negative claim across all four surfaces). + +### Workspace-trust dialog auto-accept + +**Workspace-trust dialog auto-accept** (`session-trust-dialog.ts`, pure + unit tested): Claude Code asks once per directory ("Is this a project you created or one you trust?") before it will read or edit anything, and since Codeman sessions run permission-skipping or classifier-guarded modes the answer is always yes, so a session parked on that dialog is simply stuck. + +⚠️ **Match the compacted SCREEN, never the stream.** tmux repaints a row by writing each word and then a cursor-forward (`\x1b[C`) instead of a space, and Ink colours each word separately, so the wire carries `I\x1b[Ctrust\x1b[Cthis\x1b[Cfolder`; stripping the escapes leaves `Itrustthisfolder`, because the spaces are not there to strip, they were never sent. A plain `includes('trust this folder')` therefore never matched a single chunk and the auto-accept was silently DEAD for every session that hit the dialog. `compactScreenText()` removes ALL whitespace instead (plus the `ESC ( B` charset selects that `stripAnsi` does not cover, which would otherwise land inside a phrase as a literal `(B`), which survives both that repaint style and the spaced full-screen redraw. + +⚠️ **Never answer it with a blind `\r`.** The layout has changed under us at least twice, and Claude Code 2.1.252 dropped the option numbers, put "No, exit" FIRST and highlights IT by default, so the Enter that answered the old dialog now picks *exit* and the pane dies (`Pane is dead (status 1)`) seconds after the session starts. `trustDialogNextKey()` reads the `❯` marker and returns ONE step at a time (an arrow while the cursor is on the wrong option, Enter only once the screen shows it on the trust option), with the pane re-read between steps, so a dropped arrow costs a repaint instead of the session; a frame that does not say which option is highlighted returns null and waits for the next repaint. ⚠️ The LAST marked option in the text wins, because the direct-PTY fallback reads an append-only buffer where every repaint since launch is still present and an older frame must not out-vote the freshest one. + +⚠️ Answering types into a live session, so THREE guards must all hold and none is redundant: a **startup-only window** (`TRUST_DIALOG_WINDOW_MS`, 90s, since the dialog renders before the main UI and leaving it open forever would let an agent transcript that merely QUOTES the dialog trigger an Enter, this file being an example), a **two-marker match** requiring a trust phrase AND one of the dialog's own confirm affordances (`isTrustDialogScreen`), and an **attempt cap** (`TRUST_DIALOG_MAX_ATTEMPTS`, 6: a keystroke can land while Ink is still mounting the widget and be dropped, which is the other half of why sessions got stuck here, but retrying forever would hammer keys into whatever came next; it was 3 while one Enter answered the dialog, and answering now costs at least two keystrokes). + +⚠️ It reads `capturePaneText()` and falls back to a deliberately SHORT tail of the terminal buffer only on a direct-PTY session, which has no pane: that buffer is append-only, so a longer tail would keep re-matching a dialog answered minutes ago. + +⚠️ **The scan must schedule its own next read** (`_trustDialogTimer`, cleared in `_clearAllTimers()`): it runs from the PTY `onData` handler, which was enough while one Enter answered the dialog, but the arrow that moves the cursor is the LAST output the pane produces, so a two-keystroke answer waiting on more output parks forever with the cursor sitting on the right option (measured on a live 2.1.252 spawn: cursor moved at 6 s, then nothing). + +### Owner tab layouts + +**Owner tab layouts** (COD-359, `tab-layout*.ts` + `GET`/`PUT /api/tab-layout`): named tab GROUPS over the flat tab strip, scoped per owner (`SINGLE_USER_LAYOUT_OWNER` = `@single` when multi-user is off), persisted under the `tabLayouts` key in state.json. The pure model is `tab-layout.ts`, `tab-layout-service.ts` is the sole mutation boundary, plus `tab-layout-persistence.ts` and `tab-layout-legacy-order.ts`. A layout is `{version, groups[], ungrouped[], updatedAt}` whose refs point at either a session or a saved webview (`TabRefKind`), capped at 32 groups / 512 refs. + +⚠️ **BACKEND ONLY as of 1.24.1**: nothing in `src/web/public/` calls these routes yet, so a UI built on top is new frontend work, not a rewiring job. + +⚠️ **`TabLayoutService` is the single mutation boundary** and every lifecycle caller (session created/removed, webview created/deleted, a legacy order PUT) describes ONE completed server action and gets AT MOST ONE versioned write; writing layout state from a route or a manager directly is what the service exists to prevent. + +⚠️ The layout does not replace `PUT /api/session-order`, it PROJECTS onto it: `tab-layout-legacy-order.ts` is the pure translation both ways (`putLegacyOrder()` recomposes a global order from the owner's groups), so changing one side without the other silently desyncs the tab strip from the stored layout. + +⚠️ **Reconciliation is gated on a SUCCESSFUL restore** (`markRestorationComplete` / `markRestorationFailed` / `markRestorationSkipped`, plus `assertDeletionReady()`): pruning refs against a session list that failed to load would delete live tabs, so a failed restore must leave the layout untouched. + +`PUT` takes exactly `{baseVersion, layout}` (any other key shape is a validation error), answers a stale `baseVersion` with the current layout rather than clobbering, and is capped at 128 KiB. Broadcasts `tab:layoutChanged`, owner-routed via `deriveTabLayoutSseHint`. + +### Hook events and workspace hook installation + +**Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `elicitation_complete`, `elicitation_response`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`, `prompt_submitted` (UserPromptSubmit, #367: a Claude pane reports its live conversation id first-hand). See `src/hooks-config.ts`; upstream hook semantics mirrored in `docs/claude-code-hooks-reference.md`. + +⚠️ **Every claude session INSTALLS the hooks block into its workspace** (`applyWorkspaceHooks` in hooks-config.ts → `ensureCodemanHooks`, an add-only merge that keeps a user's own handlers), from EVERY claude create path — both interactive routes, cron fires, legacy scheduled runs, the plan-orchestrator one-shots — and from `restoreMuxSessions()` for sessions recovered on server start (that boot sweep skips a workspace that no longer exists, so a deleted repo with a surviving tmux session is never resurrected as an empty dir). Before 2026-08-15 hooks were written ONLY when Codeman created the case DIRECTORY, so a linked case / cloned repo — where most sessions actually run — had no hooks at all and every hook-driven surface was silently dead there: an AskUserQuestion dialog blocked the pane while the tab and the phone overview both read a calm `idle`, with no Approvals Inbox item, no push, no definitive `stop`/`idle_prompt` for respawn and no `stop`/`blocked` for the wait endpoints. + +The escape hatch is the synced `workspaceHooksEnabled` setting (App Settings → Agents & CLIs → Claude, **default ON**); OFF restores the old behavior, where a Codeman block that is already there is still refreshed when stale (COD-91) but one is never added. ⚠️ Route the decision through `applyWorkspaceHooks` rather than calling `ensureCodemanHooks` at a new site, or the setting silently stops applying to that path. + +⚠️ Claude Code RE-READS `settings.local.json`, so an already-running session starts firing hooks without a restart (measured 2026-08-15) — and the notification for a blocking dialog is delayed by Claude Code (~30s), so the alert trails the dialog. + +⚠️ An AskUserQuestion / plan-selection dialog arrives as **`permission_prompt`**, not `elicitation_dialog` (that one is MCP elicitation), so it renders as the RED "needs you" alert, not the yellow idle one. + +### Reboot restore + +**Reboot restore** (#411/#442, `src/reboot-restore.ts` pure + `web/reboot-restore-registry.ts` + `routes/reboot-restore-routes.ts` + `reboot-restore-ui.js`): a host reboot takes the tmux server with it, so every pane dies and the board comes up empty with no explanation. At boot Codeman works out which sessions that reboot destroyed, holds the plan IN MEMORY (no new state file, and a server restart simply drops the offer), and the banner asks. + +⚠️ **The heuristic decides whether to ASK, never whether to act**: two signals have to agree (the socket holds no panes at all while state still lists sessions, AND the host booted after the newest persisted activity), and a wrong yes costs one dismissable line rather than N CLI processes nobody asked for. + +⚠️ Rebuilding is TAKE-then-build: entries leave the plan synchronously before the first `await` and the route is single-flighted per owner, so a double-click or a second device cannot put two panes on one conversation. Anything that never became a pane goes BACK on offer, with one deliberate exception, `already-live`, which unlike a missing workspace or a withdrawn grant cannot stop being true. + +⚠️ Three things are re-checked at click time rather than trusted from boot (the owner's privilege grant, the workspace still being on disk, and the conversation not already being live), and the already-live sets are read FRESH per iteration rather than snapshotted: the loop awaits a real `startInteractive()` per entry, so a snapshot taken before it is tens of seconds stale by the tenth entry and would miss a conversation the user resumed by hand in that window. The confinement re-check is keyed on the entry's OWNER, never the caller, or an admin spending another user's entry is waved through by `isWorkingDirAllowed`. + +⚠️ A rebuilt session comes back **attached, idle and disarmed**: the pane is NEW, so terminal scrollback is gone (the banner says so) while the conversation continues, respawn controllers and Ralph loops are never re-armed, and `reapplyPersistedSessionState(..., { rearmAutoResumeSchedule: false })` keeps auto-resume ENABLED but drops the pre-reboot `autoResumeAt` stamp, or one click has every restored session type `continue` into itself a minute later, unattended. That option exists only for this path; a Codeman restart still re-arms, because the limit footer will not reprint on its own. + +⚠️ The rebuild passes `nameSource` through, or the constructor re-infers it from the name and a hand-renamed session shaped like `w<n>-<case>` comes back as `placeholder` for auto-naming to overwrite. + +⚠️ Claude-mode only (others carry their conversation id in their own config object), and remote/docker sessions are never offered (`remote-or-docker`), because both need another host or container to be up. + +⚠️ A failed rebuild is undone with `discardPartiallyBuiltSession()`, deliberately NOT `cleanupSession()`: the delete path would count the session's tokens into the lifetime totals, demote a pinned record to `stopped` (which this pass reads as an intentional kill, making the session permanently unrestorable) and recursively remove the WORKSPACE's `.claude-images`. + +⚠️ `os.uptime()` reports the HOST's uptime, which a container shares, and that cuts both ways: after a genuine host reboot a containerized Codeman does see a short uptime and the banner works, but a container-only restart is invisible to it, which is the case where this would help most. Tests: `test/reboot-restore.test.ts`, `test/routes/reboot-restore-routes.test.ts`, `test/routes/reboot-restore-rebuild-failure.test.ts`, `test/discard-partially-built-session.test.ts`. ## Features @@ -239,6 +388,8 @@ Tests: `test/file-editing-policy.test.ts` (pure policy), `test/routes/file-write Tests: `test/git-clone.test.ts` (pure half exhaustively, plus REAL git against a REAL local bare repo for clone/ref/timeout/cleanup), `test/routes/case-clone-routes.test.ts` (deliberately **unmocked fs**, real clone through the endpoint). +Further detail: the synchronous clone request is bounded by `GIT_CLONE_TIMEOUT_MS`. + ### Ultracode and workflow-run visualization **Ultracode / Workflow-run visualization** (opt-in `showUltracodeAgents`, default OFF; released 1.1.2): the Workflow tool ("ultracode") writes a COMPLETION artifact per run at `~/.claude/projects/<projHash>/<sessionUuid>/workflows/wf_*.json` (written only at run end); LIVE in-flight runs exist only as transcript dirs at `…/subagents/workflows/wf_<id>/` (journal.jsonl + `agent-*.jsonl`). `workflow-run-watcher.ts` (STANDALONE — deliberately never imports/touches `subagent-watcher.ts`; separate singleton, though it independently reads the same `subagents/workflows/` tree) scans BOTH sources via periodic poll + per-directory chokidar watchers with per-source mtime skip (LRU agentStatCache + journalCache), synthesizing ACTIVE runs (live per-agent tokens/tools/state from transcripts, title/phases from the workflow script) until the completion `wf_*.json` appears and supersedes, and broadcasts SSE `workflow:run_discovered`/`run_updated`/`run_removed`. The watcher is started when **either** `showUltracodeAgents` **or** `ultracodeFloatingWindows` is on (`server.ts` `isWorkflowAgentTrackingEnabled()` returns `(showUltracodeAgents ?? false) || (ultracodeFloatingWindows ?? false)`). Served via `GET /api/workflows` (optional `?minutes=` filter) and `GET /api/workflows/:runId`. Frontend `ultracode-panel.js` renders a docked master-detail view (LEFT: runs + phases; RIGHT: per-agent tokens + tool-calls; click an agent card → its live transcript via client-side `agentId` join). **Additionally**, `ultracode-windows.js` auto-pops a draggable **floating window per active run** (gated on a **DEDICATED** `ultracodeFloatingWindows` toggle, default OFF — independent of the dock panel's `showUltracodeAgents`; see `_ultracodeFloatingEnabled()`), connected by a glowing line to the originating session tab (resolved by `session.claudeSessionId === run.sessionUuid`) — same line idiom as subagent windows, drawn into the shared `#connectionLines` SVG from the tail of `_updateConnectionLinesImmediate`. The window auto-closes ~8s after its run finishes; explicit dismissals are remembered. Clicking an agent card opens an **in-page** connected transcript window (not a browser popup); both run and transcript windows minimize **into** the originating session tab as a merged `ULTRA` badge (🧬 runs / 📄 transcripts) with a restore/dismiss dropdown — minimized runs are skipped by auto-pop. Gesture beta: floating subagent/ultracode windows are pinch-draggable (a `window` grab kind in `entry.ts`). Types: `src/types/workflow-run.ts`. Config: `src/config/workflow-config.ts`. @@ -273,6 +424,16 @@ Invariants: **Self-update** (App Settings → System → Updates): in-app updater for **git-clone installs** supervised by systemd/launchd. Supervisors: `systemd` (user unit), `launchd` (GUI LaunchAgent, gui-domain kickstart), `launchd-daemon` (KeepAlive system LaunchDaemon on headless Macs — restarts rootlessly by killing the server PID and letting launchd respawn it; detected only when the daemon is bootstrapped AND KeepAlive), else `none` → "restart manually" message; on next boot a manual-restart status auto-completes when the running version matches the target. The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` (`git checkout <release tag> && npm install && npm run build && restart`) that outlives the restart; it writes progress to `dataPath('update-status.json')`, which the browser polls across the connection drop. Channel = latest `codeman@X.Y.Z` release tag; dirty trees are auto-stashed. `src/web/self-update.ts` splits PURE helpers (semver/tag parsing, reconcile decision — unit-tested) from IO wrappers (`getInstallInfo`/`checkForUpdate`/`startUpdate`/`reconcileUpdateOnBoot`). Routes: `GET /api/system/update/check`, `POST /api/system/update`, `GET /api/system/update/status`. Types: `src/types/update.ts`. npm installs report as non-updatable. +Further detail on the Docker Compose supervisor (`docker-compose`), see also `docs/docker-self-update.md`: + +⚠️ **The Compose deployment is the one supervisor that does NOT outlive the restart**: there the restart IS the container exiting (`restart: unless-stopped` relaunches it), which kills the script too — safe only because the terminal `restarting` marker is written BEFORE the kill, so nothing may be appended after it. Two config facts make it work at all and both are load-bearing: the repo is a HOST BIND MOUNT over `/opt/codeman` (a pull into the baked image copy would land in the writable layer and be silently discarded by the next `up`), and the runtime image keeps devDependencies + a build toolchain (`npm run build` is tsc+esbuild, and node-pty has no Linux prebuild), which is why `npm prune --omit=dev` is gone and the updater passes `--include=dev` against `NODE_ENV=production`. + +⚠️ An in-place container update applies CODE ONLY — a restart reuses the existing image and config — so `evaluateEnvironmentGate()` REFUSES a release that changes `server.Dockerfile`/`docker-compose.yaml` (sha256 vs the baseline `Start-Codeman.sh` writes to `docker-env-applied.json` on every start) or adds `.env.example` keys the user's `.env` lacks, and refuses when the restart policy would not bring the container back. That third check exists because **Compose resolves an unset `${VAR}` to the EMPTY STRING and starts anyway**, so a new required setting otherwise arrives as a silently blank env var. Every unknown fails OPEN in the gate (no baseline, unreadable `.env`, no socket): failing closed would permanently block containers created before the fingerprint file existed. + +⚠️ The KILL does not: the server exits only when `--restart-by-exit 1` was passed, i.e. the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (set ONLY there, since that file is what sets `restart: unless-stopped`; the image ENV deliberately does not) or the daemon reported an auto-restart policy; otherwise the build lands as `completed-needs-manual-restart`, because exiting blind takes a `docker run` container with no restart policy down with no UI left to recover it. The gate is re-evaluated on `POST /api/system/update`, so hiding the button is UX, not the control. + +⚠️ The four global agent CLIs in `server.Dockerfile` are PINNED on purpose — unpinned, a user's CLI versions are a function of when their image was built rather than of any commit, which is the one environment change no diff-derived gate can see; pinning turns it into a Dockerfile change the gate already catches. `test/docker-compose-env-parity.test.ts` is the merge-side guard (every compose `${VAR}` ↔ an `.env.example` entry). + ### Web tabs **Web tabs** (saved dashboard URLs rendered as tabs beside agent sessions; user guide `docs/web-tabs.md`). A webview is **NOT a `SessionMode` of its own**: no PTY, no tmux, no respawn, no idle detection. It is a separate resource (`~/.codeman/webviews.json` via `src/webview-store.ts`, types in `src/types/webview.ts`, limits in `src/config/webview-limits.ts`) that shares only the tab strip and the main content area, exactly as Docker and remote-SSH are case overlays rather than modes. @@ -302,10 +463,164 @@ Invariants: **Pre-existing bug fixed alongside**: `.toolbar` has `backdrop-filter`, which makes it a stacking context and TRAPS `.run-mode-menu`'s `z-index: 1000` inside it. With the toolbar itself at `z-index: auto`, `.welcome-overlay` (z-index 10, inside `<main>`) painted over the popped-up Run menu, making **every** item in it unclickable whenever no session was open. `.toolbar` now carries `z-index: 20` (must stay below `.modal`'s 1000). +Further detail: the refused egress targets are `169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`, the Azure/Alibaba fixed metadata addresses and `metadata.google.internal`. The sandboxed-frame breakages (runtime root-absolute URLs escaping `<base>`, fixed by `runtimeUrlShim()`; same-host `fetch`/XHR CORS-checked with `Origin: null`, fixed by `buildProxyCorsHeaders()` plus the `OPTIONS`-204 exemption in `registerSecurityHeaders`) are two things `curl` can never reproduce. + +⚠️ **A loopback link in agent output opens as a proxied web tab** (`openLinkThroughWebTabIfLoopback` in webview-tabs.js, called from the terminal link provider and the response viewer): a phone cannot reach the box's `localhost:5173`, so the tap routes through Codeman's origin instead, reusing a saved same-origin dashboard or saving one under `host:port`. Two rules there look cosmetic and are not. + +**`*.localhost` is deliberately NOT in the auto-route set** even though it IS loopback to a browser: the link source is agent-written terminal output and prompt-injectable, every other member of that set is an address literal that can only mean this box, and a `*.localhost` DNS name is not one (a resolver with a search domain retries `evil.localhost` as `evil.localhost.<search domain>`, which an attacker can control, turning agent output plus one tap into a persisted server-side fetch of an agent-chosen origin). A user who really runs `api.localhost` saves it by hand, which is an explicit action. The PAGE-side test is deliberately broader (`isOnBoxHostname`), since a false positive there only declines to proxy. + +**`this.webviews` being set does NOT mean it is loaded**: `initWebviews()` assigns a truthy EMPTY map synchronously and only then awaits the list, so a tap during page load must join the in-flight refresh (`_webviewsRefresh`/`_webviewsLoaded`) or it finds nothing to reuse and POSTs a duplicate record for an origin that already exists. + ### Multi-user mode **Multi-user mode** (opt-in `--multiuser` / `CODEMAN_MULTIUSER=1`, OFF by default; shipped 1.5.0 via PR #161, design `docs/multi-user-plan.md`): named users with individually scrypt-hashed passwords in `~/.codeman/users.json` (via `src/user-store.ts`: atomic 0600 write, short-TTL cache, SERIALIZED read-modify-write so a fire-and-forget `touchLastLogin` can't clobber a concurrent route write, last-admin invariants). Gated everywhere by `isMultiUserMode()` (`src/config/multiuser.ts`); when OFF, behavior is byte-identical to single-user (all scoping helpers short-circuit). ⚠️ **Not a security boundary at the agent layer** — every session still runs as the SAME OS account; this separates WORKSPACES, it does not sandbox users (Docker cases are the isolation story). Auth: a PARALLEL async branch in `middleware/auth.ts` (single-user branch untouched) verifies `username:password` against the store, mints identity-carrying cookies (`AuthSessionRecord` gains `username`/`role`/`mustChangePassword`), decorates `req.authUser` (Fastify augmentation; single-user leaves it undefined and the ownership helpers default to a synthetic admin), enforces a per-username failure bucket + the `mustChangePassword` lockbox. Ownership threads through `Session.owner` (stamped from `req.authUser`/`job.owner` at every `new Session()`, round-tripped via `MuxSession.owner` on recovery); `findSessionOrFail(ctx,id,req)` does a NOT_FOUND owner check; list endpoints + `getLightState` + SSE (`deriveSseHint` routes session-scoped events by owner, fail-closed; machine-level + host-plan telemetry admin-only) + WS + search + file-preview all filter by owner. §6.3 permission policy: non-granted users are forced to `--permission-mode auto` (via `resolveClaudeModeForUser` at all spawn sites, incl. one-shots because `buildPromptArgs` now respects the session mode), and shell mode / cron `launchCommand` require the `canBypassPermissions` grant. Cases live in per-user `~/codeman-users/<name>/cases` (`resolveCasesDir`); a non-admin's `workingDir` is realpath-confined there; host CRUD is admin-only. Admin API `src/web/routes/admin-routes.ts` (`/api/admin/users*`, one-time passwords, audit log `admin-audit.jsonl`) + self-service `/api/me` + `/api/me/password` (`me-routes.ts`); frontend `public/admin-ui.js` (identity boot, change-password modal + interceptor, admin Users tab, and the header **Admin Panel** button `#adminPanelBtn`: ships `btn-admin-panel--hidden`, revealed for admins in multi-user mode, phone-hidden via mobile.css; opens the full Admin Panel modal with user CRUD, per-user permission toggles, and case-folder list/delete via `GET/DELETE /api/admin/users/:username/cases[/:caseName]`; live-refreshes on SSE `admin:usersChanged`, wired in app.js). CLI `codeman users add|passwd|list|rm`. Per-user session cap via `sessionCapacityState`/`sessionCapacityMessage`. Tests: `test/user-store.test.ts`, `test/multiuser-auth.test.ts`, `test/ownership-scoping.test.ts`, `test/admin-routes.test.ts`, `test/admin-ui.test.ts`. +### Agent-created case marker + +**Agent-created case marker** (`src/agent-case-marker.ts`): a case directory `POST /api/quick-start` **creates** for an agent-driven spawn gets a `.codeman-agent-case.json` marker, so the scratch workspaces a long orchestration leaves behind (one per worker, and deleting the session does not remove them) can still be told apart from the user's real projects months later. `GET /api/cases` publishes it as `agentCreated`; `GET /api/cases/agent-created` is the read-only cleanup listing, adding `inUse` (a live session's `workingDir` is that case) and `modifiedAt`; Add Case → Manage badges each one and offers a review-then-delete sweep. The signal is the skill preamble's `X-Codeman-Agent-Origin` header (or an `agentOrigin` body field), falling back to a RESOLVED `parentSessionId` — nothing in the browser UI sets lineage, so a create request naming its spawning session came from an agent by construction, and that fallback is what still labels workers spawned by a stale skill copy. + +⚠️ **Only the branch that CREATES the directory may write it.** A linked case, a cloned repo or any pre-existing path must never be labelled: the label drives a recursive-delete affordance, and mislabelling someone's repo there is the one failure mode that costs real work. `POST /api/sessions` takes an existing `workingDir`, so it writes no marker at all, by construction. + +⚠️ Reading is TOTAL: anything that is not a well-formed version-1 marker (truncated write, hand-edited junk) reads as *not* agent-created rather than as a half-trusted entry, and deleting the file is the supported way to adopt a scratch case as a real one — which is what the `note` written into it tells whoever finds it. + +⚠️ Removal stays on the existing `DELETE /api/cases/:name`, one name at a time, so there is exactly ONE recursive-delete path; the UI's sweep names every directory in its confirm and EXCLUDES an `inUse` case outright rather than confirming it away. + +⚠️ Marker in the case dir rather than a registry under `~/.codeman`: it survives a wiped data dir or a different instance, is removed by the same `rm -rf` that removes the case (so no stale-entry pruning), and a user who runs `ls -a` can see what labelled their directory. + +Adding the header changed the preamble, so `CODEMAN_PREAMBLE` was bumped (1.22.0) — a cached copy is version-checked, and forgetting the bump leaves every already-seeded agent sending the old headers. Tests: `test/agent-case-marker.test.ts`, `test/routes/agent-case-marker-routes.test.ts`. + +### Docker Compose deployment + +**Docker Compose deployment** (`docker/`, contributed): Codeman itself runs in a container and spawns Docker cases as **SIBLING** containers through the mounted host socket (Docker-outside-of-Docker), never nested. That inverts one assumption the bare-host path takes for granted: the daemon no longer shares Codeman's filesystem, so a bind source valid *inside* Codeman means nothing to it. `resolveDockerDaemonMountSource()` translates sources under HOME into the daemon's namespace via `CODEMAN_DOCKER_HOST_HOME`, and `CODEMAN_CASES_PATH` points the cases dir at a host-absolute bind mount so a workspace resolves to the SAME absolute path on both sides (which is what keeps the transcript projHash matching, per Docker cases). + +⚠️ **`CODEMAN_CASES_PATH` must move every consumer or none**: it is resolved once in `config/cases-dir.ts` because `src/cli.ts` resolves case paths too, and when only the server's `CASES_DIR` learned the override, `codeman skill install --case <name>` reported "Case not found" on exactly the deployment the override exists for. + +⚠️ **`.dockerignore` patterns match the WHOLE context-relative path**, so a bare `.env` line excludes only the ROOT file: `docker/.env` (which holds `CODEMAN_PASSWORD` and any provider keys) rode `COPY . .` into the image until `**/.env` was added — verified in both directions with a real build context. + +⚠️ A Compose LONG-form bind (`type: bind`) **creates a missing host source directory ROOT-OWNED** rather than refusing. `Start-Codeman.sh` pre-creates both `CODEMAN_APPDATA_PATH` and `CODEMAN_CASES_PATH` on the host before `up`, which is what keeps the daemon from ever having to materialise either as root in the first place; the container ALSO starts as root (`cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the base `cap_drop: ALL`; `test/docker-entrypoint.test.ts` pins that list) so `docker/entrypoint.sh` can correct a bind source that turns up root-owned anyway (a restored backup, a cleared directory, plain `docker compose up` run without the script) before dropping to `PUID:PGID` via `setpriv` — a directory owned by neither root nor `PUID:PGID` is never re-owned, since that ownership is not this container's to reassign; it is PROBED for writability as the runtime account (`setpriv ... test -w`, so ACLs, group-writable trees and CIFS/NFS mounts pass) and refused with a message naming path, owner and PUID:PGID if that fails. + +⚠️ `KILL` is in that list for tini, not the entrypoint: `init: true` keeps tini as root while the server runs as PUID, and without CAP_KILL its SIGTERM forward fails and the server is SIGKILLed on every `compose down`/`restart` instead of flushing state. + +⚠️ `/opt/codeman-cli` (the runtime-owned CLI prefix) is APPENDED to `PATH`, never prepended, and the entrypoint pins its own `PATH` to the system dirs: the root part of the start resolves `setpriv` by bare name, and a prefix ahead of `/usr/bin` let a planted `setpriv` run as uid 0 (measured). + +`CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` drops `--memory-swap` (and filters only that one kernel warning) for hosts without swap accounting; `--memory` still applies. + +⚠️ The deployment ALSO self-updates in place (the repo bind mount at `/opt/codeman` + a restart-by-exiting supervisor) — see Self-update and `docs/docker-self-update.md` before touching `server.Dockerfile`, the compose file or `.env.example`, since each is an input to the updater's environment gate. User guides: `docs/docker-compose.md` + `docker/README.md`. + +### DeepSeek web UI + +**DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. The session version worked and was still wrong in use: it put a terminal tab on screen next to the web tab the user actually asked for, every single time, and nothing about a long-lived HTTP server needs to be a tab. + +⚠️ What a session gave for free now has to be paid for explicitly, and every piece is load-bearing: **exactly one** server (a second click REUSES it rather than racing it for a port, which two sessions structurally could not do), **restarted when the browser authority changes** (`--trusted-host` fences dsh's `/api` against the browser authority, and a Codeman reachable at both loopback and a tailnet name has two, so whoever asks last wins: the asker is by definition the origin about to load the page), **killed on shutdown** (`stopDeepSeekWeb()` in the server teardown, because the child is detached so its whole plugin tree can be signalled at once, which also means it would OUTLIVE Codeman and hold its port against the next start), and **failures returned to the caller**, since with no tab there is nowhere for a stack trace to land. + +⚠️ The port search starts at dsh's own default 3080 and walks 40, never fixed: that default is precisely the port most likely to be taken already by the user's own `dsh web`, and hardcoding it killed this feature with EADDRINUSE once. Free-port detection BINDS rather than connects (a connect probe cannot tell "free" from "listening but not answering yet"), so it is racy by nature and the caller still waits for the server to really answer before reporting success. + +⚠️ Both `POST` and `DELETE` sit at the **same privilege bar as the profile installer** (`canUsernameRunPrivilegedCommands`) even though the action reads as "open a page": booting a dsh profile executes the plugin code in it, and the server is a single shared instance, so stopping it in multi-user mode takes it out from under other users' tabs. + +### Custom Model Endpoint Profiles + +**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`, design doc `docs/custom-model-endpoints-plan.md`; full stack — settings-panel CRUD + the Run-menu picker, on top of the backend below): points a session at a user-configured custom OpenAI-compatible endpoint — local (llama.cpp, DGX Spark, Strix Halo) or cloud (Azure AI Foundry, OpenRouter) — instead of its harness's native cloud backend. Endpoints are a read/write-array store (`custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`) discovered via `GET <baseUrl>/v1/models`; `CustomModelHost.authStyle` is `'bearer'` (default, `Authorization: Bearer`) or `'api-key'` (Azure's convention) — **never both**, live-tested against a real server: sending both headers on one request reliably hangs it indefinitely, reproduced 3×. + +⚠️ The actual per-CLI redirect is `capabilities.customModelInjection` on the CLI registry (four kinds: `env` for claude/gemini/deepseek, `configContentEnv` reusing opencode's existing `OPENCODE_CONFIG_CONTENT`, `configDir` for codex/pi/grok/omp — writes an isolated per-session config file, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config — and `unsupported` for antigravity, which has no known mechanism), computed by the pure `custom-model-injection.ts` (mirrors `session-cli-builder.ts`'s no-IO discipline). ⚠️ `PI_CONFIG_DIR` does NOTHING for pi or omp (grepped pi's entire bundled JS source — the string appears nowhere); both hardcode `~/.pi/agent/models.json` / `~/.omp/agent/models.yml` with no dedicated override, so the real redirect for both is the child process's own **`HOME`**, and both need `models` as an ARRAY of `{id}` objects (an object keyed by id silently loads zero models). Grok's real mechanism turned out to be a `config.toml` `[model.<name>]` block redirected via `GROK_HOME` — its original env-var-based recipe was flat-out wrong (produced "Not signed in" against a real binary), not just unverified. + +⚠️ **Two launch paths, chosen by mechanism, not preference**: opencode/codex/gemini/pi/grok/deepseek/omp apply the selection ONE-SHOT, before the session/process ever exists, with no restart at all; claude alone still applies a selection by **restarting the session's CLI process in place** via `Session.restartCli()` — a de-restricted `reattachRemote()` reusing the same `respawn-pane -k` primitive local/remote respawns already share — because every one of these harnesses reads its endpoint config at process start, never per-turn, so there is no live hot-swap; `Session.setCustomModel()` undoes the PREVIOUS selection's env keys (and deletes its old `configDir`) before merging the new ones in, so switching endpoints or clearing back to native cloud never leaves a stale key behind. ⚠️ Deleting a key from `_envOverrides` is NOT enough on its own: `tmux setenv` persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so the retired keys are queued (`_pendingEnvUnsets`) and ride `RespawnPaneOptions.unsetEnvKeys` into `applyEnvOverrides()`, which `setenv -u`s them BEFORE re-applying the live overrides. ⚠️ `restartCli()` relaunches a CLI in an existing pane, so a CLI whose launch declares a `fallback` chain (claude) gets a conversation pinned as `resumeSessionId` for that one respawn: `--session-id <id>` refuses an id that already has a transcript (`Session ID ... is already in use`), and without the `--resume <id> || --session-id <id>` shape the docker/remote pane commands already use, applying a model killed the pane and lost the session. The dead-pane respawn shares this pin, see [Dead-pane respawn: the resume pin](#dead-pane-respawn-the-resume-pin). ⚠️ pi, omp and grok need the config file AND a `model` launch param (`custom/<id>` for pi/omp, grok's `[model.codeman-custom]` block name): that is the registry's `customModelInjection.launchModel` template, applied onto the respawn options through `legacyConfigField` by `_withCustomModelLaunchModel()`, never by id, and a model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. ⚠️ Remote (SSH) and Docker sessions are REFUSED (400): their `restartCli()` reattaches a durable tmux rather than restarting the agent and the env lands on the local pane, so they used to report `restarted:true` and change nothing. The selection survives a Codeman restart as the disk-only `__customModel` (bookkeeping: env KEYS, config dir, launch model; never the values, which carry the API key and are re-derived from the endpoint store on recovery), the config dir is removed with the session, and every secret-bearing file (`custom-model-hosts.json`, the per-session config dir) is written 0600. + +⚠️ **Security**: every env var this feature can redirect (`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, `GROK_HOME`, `HOME` for pi/omp, `OPENCODE_CONFIG_CONTENT`, etc.) is in that CLI's `privilegedEnvKeys` — several of these were reachable via the generic `envOverrides` field's prefix allowlist BEFORE this feature existed (the env allowlist is global and prefix-based, not per-CLI-scoped), so building this surfaced and closed a pre-existing gap rather than opening a new one. `ANTHROPIC_*` is deliberately NOT in claude's `allowedPrefixes` at all — Anthropic-traffic redirection can only happen through this feature's own admin-configured, SSRF-guarded route, never a plain client-supplied `envOverrides`. + +**Confidence, verified end-to-end against a real llama-swap server via the DYNAMIC `scripts/test-local-llm-harnesses.ts`** (reads the live CLI registry, so a registry change needs zero script edits): claude/opencode/pi/grok/omp **PASS**; codex config structure is correct, and codex only speaks the Responses API since Feb 2026 (`wire_api = "responses"`) — re-verified live against a llama-swap deployment that DOES answer `/v1/responses` (an earlier test's harder failure against a different deployment does not reproduce everywhere): a plain, no-tool-call chat turn gets a real reply, but a real tool-call attempt came back as `agent_message` TEXT (the tool-call JSON printed as the answer) rather than an executable `function_call` item — confirmed via `codex exec --json`'s raw event stream. Tool execution is what makes codex a coding agent, so it remains not usable for real work either way, just with a more precise failure mode than a flat protocol break; gemini fails with `Invalid auth method selected` (an undocumented `GATEWAY` AuthType gemini-cli selects once `GOOGLE_GEMINI_BASE_URL` is set — unresolved after real investigation); deepseek's originally-reported `HTTP_404` is root-caused and fixed — its bundled `@deepseek-ai/dsh-llm-deepseek` module builds `${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` of its own (confirmed by installing the real package and reading its source), so a new `appendV1Suffix` flag on its registry entry (alone — claude/gemini must not get it) runs `endpoint.baseUrl` through `withV1Suffix()` before writing it, live-confirmed against llama-swap (`.../chat/completions` 404s, `.../v1/chat/completions` succeeds) though not yet re-run through an actual `dsh` binary, which isn't installable in this environment; antigravity has no known mechanism at all. See the confidence table in `docs/custom-model-endpoints-plan.md` for the full detail on each. + +⚠️ **The Run-menu picker generates entries from `window.__codemanCustomModelClis`** (`server.ts`, injected at page render from `enabledClis().filter(kind==='agent' && customModelInjection.kind!=='unsupported')`, JSON-escaped against a literal `</script>` via the exported `escapeScriptJson()` since `label` is a user-`clis.json`-settable string unlike the neighbouring booleans-only `__codemanCliAvailable`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces. One entry per (capable, INSTALLED CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)", filtered through `isCliAvailable()` like the stock entries. Clicking one calls `selectCustomModelEntry(mode, endpointId)` (`session-ui.js`), which re-fetches the endpoint (never trusts anything cached from the dropdown's render — the 5-minute sweep below or a settings edit may have changed it since) and decides the model: exactly one discovered model launches straight away, two or more open `#customModelPickModal` to ask, with `defaultModelId` marked but never auto-chosen (asking exists so ONE launch can deliberately differ from the saved default). + +⚠️ **The picker promotes exactly one row to the top**: whichever model llama-swap reports `ready` right now (tagged "Currently loaded", queried via `GET /api/model-endpoints/:id/running-status`, client-side bounded to ~800ms via `Promise.race` so a sleeping/firewalled endpoint cannot leave the modal invisible for the route's own 5s server-side timeout) beats a merely remembered choice, and — only when nothing is currently loaded — the model actually launched last for this exact (harness, endpoint) pair (tagged "Last used", read from the per-device `codeman:customModelLastUsed:<mode>:<endpointId>` localStorage key). Neither tag reorders past the top, and the "Default" pill is a SEPARATE span rather than a third value of the same slot, so a promoted row that is also the endpoint's `defaultModelId` shows both (on a single-purpose GPU box that is the common case; an exclusive slot silently dropped the Default marking for exactly that row). "Last used" is written by `_runCustomModelEntryViaRestart` (claude) and `_quickStartWithCustomModelConfirm` (every one-shot launch; `runCustomModelEntry` itself only dispatches between the two) only once the model is actually applied — never on the mere click — because a context-window-warning decline means this exact model cannot work with this CLI at all, and promoting a model that cannot launch would be actively wrong, not just premature. + +Either way the actual launch (`runCustomModelEntry`) routes through `run()` itself via a temporary `_runMode` swap — never `setRunMode()`, which would persist it as the user's new default — rather than a parallel dispatch table, which is what gives a custom-model launch the same `_runInFlight` lock every other Run click gets and means a CLI whose injection recipe lands later needs no update here. It then GETs `/api/sessions/:id/wait?until=idle&timeout=20000` on that session BEFORE applying — measured live, a freshly launched CLI reports itself `busy` for its own startup (boot spinner, workspace-trust check) well before the apply call would otherwise reach it, and the apply route's `isBusy()` guard correctly can't tell that apart from a real turn in progress, so every fresh launch failed with `SESSION_BUSY` until this wait was added. A timeout there is a normal 200 per the wait endpoint's own contract, never an error, so a session still busy after 20s just reaches the apply call anyway and gets that route's own honest error. It then calls `POST /api/sessions/:id/custom-model` on the session `run()` produced, guarded by snapshotting `activeSessionId` before the call and requiring it to have actually changed after — every `run*()` handles its own failure internally and returns normally rather than throwing, so a declined/failed launch must not silently re-point and restart whatever session was already open. + +⚠️ The apply call reads the response body itself (`_api()`) rather than `_apiJson()`, which unwraps success but silently discards a failure body — losing the one thing (`error`) that distinguishes "still busy", "remote/Docker session" and everything else the route can report (⚠️ neither route validates `modelId` against the endpoint's discovered list, deliberately: discovery can be up to 5 minutes stale, so a 400 there would refuse a launch that works — a typo'd id fails on the CLI's own first request instead); the resulting toast is `type: 'error'` with an explicit `duration: 0` (no auto-dismiss, an explicit close button) at that one call site — not a blanket sticky-error default, which stacked unbounded on `.toast-container` with no cap or eviction — precisely so a message worth diagnosing survives long enough to be read instead of vanishing on the usual 3s timer. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to launch with). + +⚠️ **Every saved endpoint's models also re-discover themselves automatically**, a `this.cleanup.setInterval` in `server.ts` (`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, 5 minutes, off under `testMode` like the Codex plan-usage poll beside it) calling the exported `refreshAllCustomModelHosts()` (`custom-model-routes.ts`) — one endpoint unreachable on a cycle never blocks the others, and a read-modify-write PER HOST (re-reading the store before each splice, keyed by id) means an admin's concurrent edit or delete wins over a sweep that started before it, never the reverse. + +**Later additions (llama-swap), each confirmed live against a real llama-swap deployment.** + +⚠️ **llama-swap runs one model at a time, and switching can disrupt ANOTHER live session** — before applying, both apply routes call llama-swap's own `GET /running` (feature-detected via `getLlamaSwapStatus()`, `custom-model-routes.ts`; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is simply never checked). If a different model is loaded and ready AND another live session's own selection is using it, the apply returns `{requiresConfirmation, currentlyLoadedModel, affectedSessions}` instead of switching silently; retrying with `confirmedSwap: true` skips the check, and switching with nothing else affected proceeds immediately. + +⚠️ **The swap question and the context-floor question below have SEPARATE flags** (`confirmedSwap`, `confirmedContext`), because the context check runs first and while both shared one `confirmed` a user who clicked past a too-small context silently consented to evicting another session's model too; the legacy `confirmed` still means both, since it shipped in the HTTP-API-only cut. llama-swap also has no dedicated "switch model" endpoint — the only thing that actually starts a swap is a real inference request naming the model (confirmed live: applying a selection alone never reached llama-swap's own logs, since nothing had asked it to load anything) — so both routes also fire `triggerLlamaSwapLoad()`, a fire-and-forget `POST <baseUrl>/v1/chat/completions` with `max_tokens: 1`, whenever the target model isn't already loaded and ready. + +⚠️ **That launch-time check cannot catch a swap caused by a DIFFERENT session's LATER, ordinary use** — confirmed live: a second Codex session picking a different model launched with no warning at all (nothing conflicted at that exact instant), yet it silently evicted the first session's model regardless, since llama-swap has no push notification of its own. `detectCustomModelSwapDisplacements()` (`custom-model-routes.ts`) is a separate periodic sweep (`server.ts`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS` = 20s) that compares each live custom-model session's own `modelId` against what `/running` actually reports loaded, broadcasting a `custom-model:swapped-out` SSE event — shown as a global toast, never tied to the displaced session's own tab, since the whole point is telling the user before they type into it — the first time a mismatch appears, via a caller-owned de-dupe `Set` cleared once that session's own model is loaded and ready again so a later, genuinely new displacement notifies again rather than staying silently un-notified forever after the first one. + +⚠️ **Context length is read from the REAL launch command, never `/props`** — `/props?model=`'s `default_generation_settings.n_ctx` was confirmed live to report a `--fit-ctx`-launched backend's theoretical/trained maximum rather than the real runtime-configured size (a measured 154112-vs-16384 discrepancy, caught only because the unfixed value still overflowed), so discovery parses the actual configured size straight out of `/running`'s own `cmd` field instead (`parseCtxFromCmd`: `--fit-ctx <N>` first, then plain llama.cpp `-c`/`--ctx-size`), falling back to `/props` only when `cmd` states no recognizable flag at all. + +⚠️ **Claude alone gets a context-window FLOOR check, on top of the ceiling `contextLengthVar` already fixes** — `exceedsSafeContextFloor()` (gated on the registry declaring `contextLengthVar`, so a no-op for every other CLI by construction) compares a model's discovered context against `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000): confirmed live, twice, that Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on the very first message, before any conversation history exists to compact, so a smaller real context fails outright regardless of what `CLAUDE_CODE_MAX_CONTEXT_TOKENS` says (that var only controls when HISTORY gets compacted, and there is none yet on message one). Below the floor, the apply returns `{requiresContextWarning, modelId, contextLength, minSafeContextTokens}` instead of launching, shown as an in-app dialog naming the actual fix: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on `--fit-ctx` auto-fit, which optimizes for the biggest MODEL that fits rather than the biggest CONTEXT. + +⚠️ **A fresh, isolated `CLAUDE_CONFIG_DIR` looks like a brand-new Claude Code profile and replays its ENTIRE first-run sequence on every launch** — the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with a bypass-permissions flag) a one-time warning about it, confirmed live, none of which a real, already-onboarded profile shows again. `skipFirstRunPrompts` (claude's entry only, requires `apiKeyTrustFile` since it reuses the same file) pre-seeds that same "already been through this" state: `hasCompletedOnboarding` and this session's own `projects[workingDir].hasTrustDialogAccepted` merge into the same `.claude.json` the API-key trust file already writes to, and `skipDangerousModePermissionPrompt` merges into `settings.json` (a different file, same corrupt-tolerant merge). + +⚠️ **The loading banner shows the REAL backend log line, not a guess, and has no countdown or auto-timeout at all.** `getLatestLlamaSwapLogLine()` holds one `GET /api/events` SSE connection open per endpoint (confirmed live to stay open indefinitely — read past 220KB over 8s with no `done`; idle-closed after 30s via `pruneIdleLlamaSwapLogTails`, same 20s sweep as the swap-displacement check above), parsing `logData` frames and keeping only `source: "upstream"` (the real `llama-server` process's own stdout) lines, never `source: "proxy"` (llama-swap's own request-access log). ⚠️ `GET /logs` — the endpoint this feature's own first cut targeted, since the name suggested it — was confirmed live to carry ONLY the proxy log and never a single backend line, even seconds after a real, verified model swap; caught and corrected by a live check before merge, not after. The banner itself dropped its size-scaled expected-time estimate and matching auto-timeout (a guess dressed up as a fact that could kill a genuinely slow load partway through on slower hardware) for a generic hardware/model-size disclaimer plus a user-driven **Cancel** button (`_showCenterStatus`'s `onCancel` option, a real button distinct from the plain "×" close glyph an `'error'`-type banner gets) that ends the wait and closes the session on the user's own call rather than a guessed deadline. + +### Approvals Inbox + +**Approvals Inbox** (cross-session queue of prompts waiting on a human; `approvalsInboxEnabled`, SYNCED, default OFF: every surface is opt-in; only the store and answer endpoints run regardless, so flipping it ON shows anything already pending): `web/approval-inbox.ts` is a `sessionWaits`-style singleton fed by `/api/hook-event`, holding at most ONE item per session (a new prompt supersedes), claude-mode only, in-memory. Cards are answered via `POST /api/approvals/:id/answer`, which sends a digit / Esc / idle-prompt text through `writeViaMux` (menu answers never carry `\r`). + +⚠️ `option` digits are accepted ONLY when they match options parsed from the captured pane frame, and the answer path RE-CAPTURES the pane first (a dialog that no longer parses on screen means the keystroke would land in the composer, so refuse with 409). + +⚠️ Resolution on the heuristic `working` signal ALONE is restricted to `idle` items; a permission/question item gets the pane-VERIFIED variant on that same signal (`resolveIfDialogGone()` → `verifyStillAnswerable()`), so the heuristic only decides when to LOOK and the screen decides the outcome. That is what clears a dialog answered in the terminal mid-turn; the other definitive signals are `stop`, `elicitation_complete`/`elicitation_response`, exit/delete, answer, supersede and the 12h TTL. + +⚠️ **Viewing a session ACKNOWLEDGES its idle item, it does not resolve it** (`POST /api/approvals/session/:sessionId/viewed` → `acknowledgedAt` → `approval:updated`): the item stays pending (still answerable, still Read My Mind context) and only stops arming the yellow tab alert. That flag is what makes the clear durable, since the view-clears-idle rule used to live in one browser's memory and `seedApprovals()` re-armed the alert on the next reload while other devices never heard about it at all; the local half is `markIdleAlertSeen()` (app.js), called from BOTH `selectSession` paths, including the already-active early return, where a click could otherwise never clear the alert. + +⚠️ **Only a HUMAN opening a session acknowledges**: `selectSession(id, { auto: true })` marks the three selections the APP makes (boot restore, a solo window opening its target, the fallback after the active session is closed) and skips the acknowledgement, so a page load cannot silently spend an alert the user never saw. The flag defaults to user-initiated, so an untagged call site fails toward acknowledging rather than toward an alert nothing can clear; `test/session-select-ack-gate.test.ts` pins both the gate and the tagged call sites. Idle-only by construction (`acknowledge()` defaults to `['idle']`): looking at a permission/question dialog does not answer it. + +⚠️ Same rule on the input path: `_ackDelivery` (app.js) spends the IDLE alert only, via that same `markIdleAlertSeen()`. It used to `clearPendingHooks(sessionId)` with no kind, so one keystroke wiped a RED alert on that device while the dialog was still up, the other devices stayed red, and a reload re-seeded it. + +⚠️ Claude Code fires no "permission answered" hook (only `elicitation_complete`/`elicitation_response`, i.e. the question flavor), so an answered-in-the-terminal dialog would otherwise sit pending until `stop`: `GET /api/approvals` therefore runs a **staleness sweep** over the caller's own items via `verifyStillAnswerable()`, which is deliberately the conservative check the answer path uses (only an item whose ORIGINAL frame parsed options can be dropped, so an unreadable capture keeps the alert rather than losing a live one). + +⚠️ **`applyCapture()` is therefore ADD-ONLY for `options`**: a re-capture that parses nothing must never erase a parse an earlier one found. Claude Code delays the Notification hook behind the dialog (measured 6s, documented ~30s), so the 600ms re-capture routinely lands on a frame the user has ALREADY answered; clearing the field there made the item permanently unsweepable, because `verifyStillAnswerable()` reads a MISSING `options` as "we never could read this dialog" and keeps such items answerable by design. The red "needs you" then survived every sweep AND every page reload, went away only on `stop` (2026-08-20: a confirmed question left a tab flowing red for ~8 minutes while the turn ran on), and the stale card still accepted an answer, typing a bare `1` into a composer with no dialog under it. Pinned by `test/approval-inbox.test.ts`. + +⚠️ A frame that parses no options is CONCLUSIVE in exactly two cases, and the second one closes the late-hook hole: the item once parsed options (they cannot vanish while the dialog is up), or the frame shows Claude actively running a turn. A modal dialog BLOCKS the turn, so the two cannot coexist — measured on v2.1.237, a live-dialog frame carries neither the `… (13s` timer NOR the `esc to interrupt` footer, which the dialog replaces with `Enter to select · ↑/↓ to navigate · Esc to cancel`. Anything else stays answerable, so an unreadable capture still keeps the alert. That second signal is reached by a delayed staleness pass (`STALE_CHECK_DELAY_MS`, 3s) scheduled alongside the re-capture, because a prompt answered BEFORE the hook lands creates an item whose FIRST capture already has no dialog in it: nothing ever parsed, `stop` may have fired already, and the alert then outlived reloads until the 12h TTL. + +⚠️ That pass must stay comfortably LATER than `RECAPTURE_DELAY_MS`, whose whole reason for existing is that the hook can beat Ink to the screen — resolving inside the paint window would clear the alert for a dialog that was about to appear. The frontend seeds from `GET /api/approvals` in `handleInit` **regardless of the setting**: the seed re-arms the tab-alert state machine (`setPendingHook`) unconditionally, and only populating `this.approvals` (the inbox surfaces) is gated — seeding used to be gated wholesale, which left a reloaded page with NO red tab while a permission dialog sat blocking a session (2026-08-15); `_onApprovalResolved` clears the pending-hook alert unconditionally for the same reason. + +⚠️ The red/yellow tab alert itself is a STEADY border/background/dot with a pulse on top: the original keyframes swung to transparent at 0%/100%, so half of every cycle looked like a normal tab. Push Approve/Deny buttons stay gated on the setting (`sendPushNotifications` strips `actions`/`approvalId` when OFF) and are answered from `sw.js` directly so they work with no tab open. Surfaces (all gated on the setting): header bell (marker-hidden until count > 0, phones never show it) + drawer (`approvals-ui.js`), phone overview NEEDS YOU answer strips (`mobile-overview.js`). Design: `docs/approvals-inbox-plan.md`. + +### Read My Mind intent profiles + +**Read My Mind intent profiles** (phase 1 of `docs/readmymind-plan.md`; `readMyMindEnabled`, SYNCED, default OFF): per-CASE profiles (user-stated `goals` + the user's recent real prompts), keyed by owner + realpath(workingDir) so they survive `/clear`/respawns and multi-user scoping is structural. Capture rides the transcript (`transcript:user_prompt` from `transcript-watcher.ts`), NOT the input paths: `POST /input` sees only programmatic prompts and the WS channel is raw keystrokes. The listener lives inside `startTranscriptWatcher()`'s `if (!watcher)` block (outside it would duplicate per hook event) and is claude-only + gated on the setting per event. Store: `src/intent-store.ts` singleton, `intents.json` written 0600 tmp+rename (prompts can contain secrets; never fed to `/api/search`). Endpoints: GET/PUT/DELETE `/api/sessions/:id/intent` + POST `/api/sessions/:id/readmymind` (`readmymind-routes.ts`, ownership via `findSessionOrFail` WITH `req`; registrations stay the bare `app.<method>('path')` shape, the endpoints.md drift scanner cannot see generics). + +**Phase 2 (predictor + 🧠 button)**: `readmymind-context.ts` is the PURE budgeted assembler (9 ranked sources, drop order siblings→away→workspace→tools, sections 1-4 truncate only); IO lives in `readmymind-collectors.ts` (transcript TAIL read — the live watcher keeps only a 500-char snippet — + git signals, skipped for remote-SSH cases) and the route; `readmymind-predictor.ts` reuses the AiCheckerBase spawn mechanics standalone (verdict-shaped base vs freeform JSON) as a mutable singleton routes call and tests stub. Claude-mode only (400), one in flight per session (409 CONFLICT), model = `readMyMindModel` setting defaulting to `AI_CHECK_MODEL` (opus, decided). + +Frontend `readmymind-ui.js`: header 🧠 marker-hidden (`btn-readmymind--hidden`) until the setting is ON; phones hide it in mobile.css and get a keyboard-accessory 🧠 key instead (ships in BOTH bar templates, revealed by the `rmm-enabled` class on the BAR element — setMode() rebuilds button innerHTML, so per-key state would be wiped; synced at init + every `applyHeaderVisibilitySettings()`). Alternate suggestions render as tappable rows that swap into the editable field without losing edits; Rethink rejects the whole shown set and carries the optional steer note (`#readMyMindSteer`, sent as `steer`, shown in ready + empty-result phases, cleared on each open). Suggestions render via value/`textContent` ONLY and Send/Insert go through `POST /input` (server-side, so the sendEnterKey/local-echo trap does not apply) — nothing auto-sends, ever. User guide: `docs/readmymind.md`. + +### Voice dictation via Claude + +**Voice dictation via Claude** (`claudeVoiceEnabled`, SYNCED, default OFF): the mic button can transcribe through this machine's Claude Code login instead of a Deepgram key, using the same speech-to-text service the CLI's own `/voice` mode uses. + +⚠️ **Claude Code's voice mode itself is unusable here**: it opens the HOST's microphone (`sox`/`arecord`), and the CLI runs in a headless tmux pane while the human is in a browser elsewhere. So Codeman captures in the browser and borrows only the backend. Audio goes browser → Codeman → Anthropic (`src/web/voice-stream.ts`): the OAuth token never reaches the page, and the browser only sends PCM and receives text. + +⚠️ Credentials are **read-only** (`src/claude-credentials.ts`) and Codeman never refreshes them — a refresh rotates the refresh token and could sign the user out of their own CLI; an elapsed token reports `expired` instead. + +⚠️ Capture MUST be linear16/16 kHz/mono, so it uses an **AudioWorklet**, not MediaRecorder (which cannot emit raw PCM); `voice-pcm-worklet.js` is fetched from JS, so it is invisible to `cacheBustAssets` and borrows voice-input.js's `?v=` token — **edit the two together**. + +⚠️ Transcript frames carry the WHOLE running transcript, not deltas: the Claude path replaces where the Deepgram path appends. Provider choice is `voiceSettings.provider` (`auto` prefers Claude → Deepgram → Web Speech). → `docs/claude-voice-plan.md` + +### Files panel search + +**Files panel search** (COD-236, the `q` param on `GET /api/sessions/:id/files`): `compileFileQuery()` (`utils/file-query.ts`, pure, no IO, so it unit-tests directly) compiles the query into a reusable predicate, which is what lets the server-side walk prune instead of streaming the whole tree. + +⚠️ **A query turns that endpoint into a FLAT match list rather than a nested tree**, and the walk deliberately recurses past non-matching directories, since the whole point of searching is to reach a file whose ancestors do not match. An empty or whitespace-only query compiles to `null`, which is what keeps the default tree response byte-identical when no search is requested. + +⚠️ **Globs are never compiled into a RegExp**: `*a*a*a…` translated to `^.*a.*a.*a…$` is a classic backtracking blowup evaluated synchronously against every walked path, so one pathological query would freeze the event loop for the whole server (the same reason `search-service.ts` is regex-free). `globMatch()` is a two-pointer wildcard walk instead, O(text · pattern) with both operands short by construction, and an overlong query (`MAX_QUERY_LENGTH`, 256) also compiles to `null` rather than running. + +Matching semantics: a query containing `/` matches the relative path, otherwise the bare entry name; globs match anchored and case-insensitively (`*` spans any run, slashes included, `?` exactly one character), everything else is a plain case-insensitive substring. + +### Raw file bodies: streamed and range-aware + +**Raw file bodies are streamed and range-aware**: `file-raw`, the attachments `/raw` route and `GET /api/download` always advertise `Accept-Ranges: bytes` and answer a `Range` header with `206` + `Content-Range` (single-range only; parser is pure + unit-tested in `src/web/http-range.ts`, a malformed spec is ignored → 200 while an out-of-bounds one is a 416). + +⚠️ **The size cap on all three is a sanity bound, not memory protection** (`MAX_FILE_DOWNLOAD_BYTES` in `config/buffer-limits.ts`, default 2GB, env `CODEMAN_MAX_DOWNLOAD_BYTES`, `0` = unlimited): the bodies stream, so size costs a read stream and not RSS (measured: a 600MB download moved peak RSS by ~37MB). Its predecessor was a hardcoded 50MB whose comment still said "prevent memory exhaustion" long after the `readFile()` it described was replaced by `sendFileBody()`, so all it did was refuse legitimate downloads of build artifacts, videos and archives. `/api/download` was the last route that really did buffer the whole file, and now shares `sendFileBody()` with the other two. + +⚠️ A 200-only response is what made the File Viewer's `<video>` unseekable: Chrome then reports `video.seekable` as `[0, 0]`, the scrub bar is inert and `currentTime = x` silently reverts (measured on an 18MB mp4), and Safari refuses to start the media at all. + +⚠️ These bodies go out through `reply.hijack()`, which bypasses Fastify's status handling — `sendRawStream` must copy the status onto `reply.raw` by hand or a partial body ships labelled `200` and the browser treats a slice as the whole file. + +⚠️ Closing the preview must **pause and unload** the media (`_stopFilePreviewMedia` in panels-ui.js): dropping the overlay's `visible` class is `display:none` and nothing else, and a DETACHED `HTMLMediaElement` keeps playing, which is how the X button used to leave a video audible with no player to pause. + ## Frontend ### Command palette and shortcut registry @@ -391,6 +706,8 @@ Anatomy: `.set-shell` → `.set-shell-head` (title + `.set-head-actions`) + `.se `data-claude-only` lives on the **rail entries**, so external-CLI sessions lose Respawn and Ralph and open on Session (`switchOptionsTab('context')`, verified in the browser). `admin-ui.js` injects the multi-user Users entry as a `data-section="settings-users"` rail button plus a `#settings-users` section appended to `.set-rail-items` / `.set-doc`, so those two hooks must survive any restructure. +Further detail: in the Settings section order, the version and the updater lead the document and the rest of the system settings tail it. ⚠️ Model cards (`#appSettingsModelCards`) and the effort segment are **views over hidden `<select>`s** that remain the source of truth; the cards hold the BASE model and the "1M context window" switch composes `base + [1m]` back into `claudeModel`, which is what retires the old "takes precedence over the toggle below" trap. Without the explicit chevron on the Add Case adapter's `<details>`, the five collapsed blocks render as plain headings nobody clicks. + ### WebGL renderer toggle **WebGL renderer toggle** (#140, `webglRendererEnabled`): per-device (`displayKeys` set, stripped from the server payload — NOT in `SettingsUpdateSchema`, which is `.strict()`). The GPU-stall watchdog's sticky `codeman-webgl-disabled` marker survives page loads; it's cleared only by an explicit OFF→ON save transition or `?webgl=force` (`shouldSkipWebGL` in constants.js). `?nowebgl` still forces the DOM renderer per-load. @@ -415,6 +732,24 @@ Anatomy: `.set-shell` → `.set-shell-head` (title + `.set-head-actions`) + `.se ⚠️ Collapsed means **different things per viewport**: at 1024px and up the sidebar keeps a 44px icon rail so the ambient signal (status dot, task/subagent/ultracode badges) survives — the Alt+N number, the name/folder and the `sh`/`oc`/`cx`/`gm` mode chip do NOT, because 44px minus paddings and borders is ~34px of content box and the chip lives inside `.tab-info`; below 1024px `mobile.css` turns the sidebar into an off-canvas overlay where collapsed == drawer closed (mirrored into an `.open` class plus `inert`/`aria-hidden`, since `translateX(-100%)` alone leaves every row in the Tab order), it defaults to CLOSED when the user has made no choice, and picking a session or web tab dismisses it. ⚠️ **That 1024px breakpoint is the only handheld test the sidebar may use** (`_isSessionSidebarOverlay()`, mirrored in the pre-paint script): `MobileDetection.getDeviceType()` calls everything from 768px up `'desktop'`, so using it gave 768-1023px the overlay CSS with docked-sidebar logic — drawer opening itself on load, immune to selection and Escape. The toggle chord (default Alt+B) also needs its gate in `terminal-ui.js`'s `attachCustomKeyEventHandler`, or `preventDefault()` in the capture handler still lets xterm write ESC b into the live PTY (same trap as COD-153). The sidebar filter only applies while its input is on screen — `applySidebarFilter()` strips the class in the header strip, the collapsed rail and the closed drawer, because a filter with no reachable control hides sessions permanently. Collapse state lives in its OWN `codeman-sidebar-collapsed` key, **not** in the settings blob — `saveAppSettings()` rebuilds that blob from DOM controls, so a key without a control is wiped on every Save. Solo (`/session/:id`) windows never get a sidebar (three guards: `getSessionListLayout()`, the pre-paint script, and `body.solo-mode`), because `#sessionTabs` parked in a `display:none` subtree measures 0/0 for tab overflow and inline rename. The sidebar CSS block sits at the END of `styles.css`, **after** the `html:not([data-skin="og"])` nesting block, and is layout-only — any colour on `.session-tab` there would render correctly on the `og` skin only. Same for the `mobile.css` block: it must stay at the end of the file or the earlier compact-strip rules clip the list to a 36px sliver. Two surfaces DEFER to the sidebar rather than adapt: **lineage arcs are skipped** in sidebar layout (`_appendLineageConnectionLines` early-returns — `computeLineagePath()`'s whole geometry hangs a U-bridge from the horizontal STRIP's bottom edge, so against a vertical list every arc would loop to the foot of the sidebar; a sideways lineage shape needs its own visual tuning, it is not a by-product of re-parenting), and the **desktop home tab rail** (`shouldShowHomeSessions()`) stays hidden while the sidebar is active, because both dock the session list flush left and the rail would render the same list next to it, z-ordered UNDER it. The subagent/ultracode connectors DO adapt (`_tabAnchor()`/`_tabConnectorPath()` in app.js: right-edge anchor, horizontal bezier), and the lineage strip-scroll listener redraws them on the sidebar's vertical scroll. `_scrollActiveTabIntoView()` owns active-row reveal on BOTH axes: sidebar mode branches to `scrollIntoView({block:'nearest'})` because the horizontal `computeTabScrollLeft` math no-ops against a vertical scroller, and `_fullRenderSessionTabs()` restores `scrollTop` alongside the #257 `scrollLeft` restore or ambient rebuilds yank a mid-scroll sidebar back to the top. Tests: `test/session-list-layout.test.ts`. +Further detail: with many sessions the horizontal strip stops being scannable, which is why the list can move into the vertical `<aside>` with a filter box and a live count, collapsible to a 44px rail (`--sidebar-width` 260 / `--sidebar-width-collapsed` 44) via **Alt+B** (`toggleSessionSidebar`; Alt, not Ctrl+B, which must reach tmux/readline in the terminal). The setting lives in App Settings → Appearance → Tabs. There is a THIRD host for `#sessionTabs` besides `#sessionTabsHost` and `#sessionSidebarList`: `#tabRail`, the vertical rail. Exactly TWO functions reparent it and they must run in this order: `applySessionListLayout()` first (sidebar wins), then `applyTabOrientation()` (settings-ui.js), which moves the tabs into `#tabRail` only when the sidebar does not own them. ⚠️ `applySessionListLayout()` sets `data-session-list` / `data-sidebar` on `<html>` and must run BEFORE `applyTabWrapSettings()`, which is the one owner of `tabs-two-rows`/`tabs-show-folder` and reads those attributes. + +⚠️ **The vertical tab rail** (`tabOrientation`/`tabRailWidth`/`sessionSidebarFontSize`, all per-device display keys that ARE in the schema, like `sessionListLayout`) is a SECOND vertical list next to the sidebar: the orientation setting is silently ignored while the sidebar layout is chosen, desktop/tablet only (`resolveTabOrientation` forces horizontal on mobile), resizable via `tab-rail-resize.js` (which owns terminal refits during the drag). The pre-paint script stamps `data-tab-orientation` (+ `--tab-rail-width`) like it stamps the sidebar keys, or vertical mode flashes through the header strip; the name font size defaults to 12px, the sidebar's historical size, so untouched installs are never restyled. + +⚠️ **Detailed rows are a property of a vertical LIST, not of one surface** (`sessionListLayout: 'sidebar-rich'` for the sidebar, `tabRailDetail: 'rich'|'simple'` for the rail, rail default **rich**): both draw the home screen's per-session line (`created 3d ago · working 12m`) plus a status pill, from the SAME row model (`_sidebarRichRow`/`_sidebarRichMetaHTML` in app.js, classified by `_mobileOverviewState`/`_mobileOverviewSince`), and the render paths ask ONE gate, `isRichTabRows()` (= `isSessionSidebarRich() || isTabRailRich()`). ⚠️ Detail rides on its own attribute (`data-sidebar-detail` / `data-tab-rail-detail`) so every existing `[data-session-list="sidebar"]` / `[data-tab-orientation='vertical']` rule keeps matching both variants untouched; a flip of detail ALONE still needs a full render (the stamps line is emitted by the row template, not toggled by CSS) and must re-run `applyTabWrapSettings()`, which owns the folder line and is now rail-aware. ⚠️ The rich CSS rules carry a rail twin as a COMMA-GROUPED selector, never `:is()` (an `:is()` list takes its most specific argument, which would lift the sidebar arm from (0,3,1) to the rail's (0,5,1)). + +⚠️ Width is the whole reason there are thresholds: the rich sidebar is 300px (`--sidebar-width-rich`) and a rail that has never been sized defaults to **320** (`RICH_DEFAULT_WIDTH`, the existing Wide preset) instead of 256, because at 256 the stamps line ellipsizes mid-word; a user-narrowed rail drops the created stamp below 288 (`tab-rail-tight`, CSS only) and drops rich rows entirely below 240 (`tab-rail-compact`, which re-renders). A stored width is never overridden. ⚠️ Only detailed rows carry stamps that go stale with no event behind them, so `_startSidebarRichClock()` (20s, rewrites text in place — a re-render would restart every row's animation) must be armed and disarmed by BOTH `applySessionListLayout()` and `applyTabOrientation()`. + +⚠️ **Axis decisions must use `_isVerticalTabList()`** (sidebar OR rail), never `isSessionSidebarActive()` alone: the rail leaves `data-session-list` at `header`, and the sidebar-only predicate shipped four rail bugs at once (drag insertion side read from clientX, active tab never scrolled into view, floating windows anchored below tabs instead of beside them, connector redraws skipped on rail scroll). + +⚠️ Leaving sidebar mode **clears `_sidebarFilter`**: the filter box only exists in the aside, so a stale filter would hide sessions from the header strip with no reachable control to clear it. ⚠️ On handhelds the closed drawer keeps `display: flex`, so without `inert` + `aria-hidden` (`_isSessionSidebarOverlay()`) its filter box and ~4 tab stops per session stay in the tab order; the DOCKED desktop rail must never be inerted, its rows are still clickable. + +⚠️ **The detailed RAIL additionally wears the home rail's CARD, and the detailed SIDEBAR deliberately does not**: the rail is an occasional, resizable list you scan, the sidebar is a permanently-docked nav column where 20 stacked cards read as a wall — so the card rules are RAIL-SCOPED and were NOT added to the comma-grouped selectors above, which is the one-line change that would silently restyle the sidebar. Card state accents reuse `home-sessions-blink-red`/`-yellow` rather than declaring a second pair, and every state dot rule excludes `.tab-alert-action`/`.tab-alert-idle` by hand, because those alert rules are only (0,3,0) and the rail's are (0,5,1)+. There is no `.active` rule in that block on purpose: `.session-tab.active` already paints border/background/box-shadow `!important`. + +⚠️ **Row ORDER is a third rail attribute** (`tabRailSort: 'activity'|'manual'`, `data-tab-rail-sort`, default **activity**): a sorted rail answers the home screens' question with the home screens' answer, `CodemanSessionOrder` over rows classified by `_mobileOverviewState` (`_tabRailSortOrder` in app.js). ⚠️ **It is applied as the flex `order` property, never by reordering the DOM**: `#sessionTabs` stays in `sessionOrder`, so the Alt+N badge (`_tabIdx`, which therefore does NOT run 1,2,3 down a sorted rail — it names a shortcut, not a position), drag-and-drop, the arrow-key walk, the sidebar filter and `_scrollActiveTabIntoView()` all keep reading the list they always read, and a session changing state moves ONE inline style instead of forcing the full rebuild that would restart every card's animation on every SSE tick. The incremental render path therefore has to re-apply it (a state flip adds no tab, so the full rebuild is never reached) and an empty string is what clears it when sorting stops. ⚠️ Web tabs are pinned past the cards by a CSS `order: 9999` rather than an inline one, since `renderWebviewTabs()` emits the same markup for every layout; the flex default of 0 would interleave them among the sorted sessions. ⚠️ `setupTabDragHandlers()` sets `draggable="false"` and returns while sorting is on: the drop rewrites `sessionOrder` correctly and the sort then puts the card straight back, so the affordance would be a lie — `'manual'` is the way back to drag-reordering. + +⚠️ **The arrow-key walk is the one place that must follow the SORT rather than the DOM**: `_tabKeydownHandler` steps `querySelectorAll` order, which is `sessionOrder`, so on a sorted rail ArrowDown from the top card landed wherever that session sat in the tab order instead of on the card below it. It now sorts its node list by the COMPUTED `order` first (computed, not inline: web tabs get their `order: 9999` from CSS and would otherwise read as 0 and lead the walk). This is the exception that proves the DOM-stays-in-sessionOrder rule: the badge, drag model and Alt+N index all deliberately keep reading the DOM, and only the thing the user steps with their eyes follows the paint. + ### Split-pane sessions **Split-pane sessions** (`showSplitButton`, header button, default OFF, per-device like `showFileViewerButton` — not in `SettingsUpdateSchema`, `displayKeys` in settings-ui.js): shows two live sessions side-by-side in one Codeman window. Pane A is the untouched, existing singleton terminal (`this.terminal`/`this._ws` in terminal-ui.js); Pane B is a new, independent `SplitTerminalPane` (terminal-split.js) with its own xterm instance and its own `/ws/sessions/:id/terminal` WebSocket. ⚠️ **Pane B is deliberately plainer than Pane A** — no local-echo overlay, no CJK IME, no touch/mobile handlers, no keyboard accessory bar — since this is a desktop-only feature (a split view needs a wide viewport) and those features exist for mobile/touch input; `.btn-split` is hard-hidden below 1180px regardless of the setting by the `@media (max-width: 1179px)` rule in styles.css (mobile.css only carries a comment pointing at it: that file loads up to 1023px, so it cannot cover the 1024-1179px tablet range the feature also needs to stay off), and the per-device setting means turning it on at a desk can never sync it onto a phone in the first place. No persistence: closing the browser tab or reloading always returns to the normal single-pane view; there is no localStorage key for split state. ⚠️ Splitting a session against itself is disallowed (the picker excludes the active session), **as is splitting against a popped-out (detached) session** — `buildSplitPickerSessions()` excludes `detachedSessions` because a detached session's own window is already claiming its PTY size, and `MAX_WS_PER_SESSION` needs no change since Pane A/B are always two different sessions. ⚠️ Either pane's session ending (deleted locally or from another client) auto-collapses the split — Pane A's session ending promotes Pane B to the new single pane via `selectSession(id, { auto: true })` (an app-driven selection, so it must not spend the session's idle alert — see the Approvals Inbox note above), never by trying to hot-swap the lightweight `SplitTerminalPane` object into the primary singleton state. ⚠️ Pane B refits on every window/sidebar/tab-rail resize via the SAME trailing-edge `ResizeObserver` callback that resizes Pane A (`throttledResize` in terminal-ui.js) — it only ever measured Pane A's own container, so without an explicit `this._splitPane?.fit()` call there Pane B silently kept its stale PTY size through every resize that did not happen to be a divider drag. ⚠️ A dropped WebSocket leaves Pane B visibly dead (a message written into its own xterm buffer) rather than silently swallowing keystrokes with nothing on screen to explain why — there is no reconnect logic for v1, matching the "deliberately plainer than Pane A" design. Related but distinct: `detachSession()` already opens one session in a separate OS-level browser window (`isSoloWindow`) — that is prior art for "two sessions visible at once" but not for one window with a draggable in-page divider, which is what this feature adds. ⚠️ Pane B installs its own `attachCustomKeyEventHandler` gating the same app-level chords Pane A's own handler gates (command palette, Alt+1-9/[/] tab nav, Alt+B sidebar toggle, Ctrl+Z suspend, Shift/Ctrl+Enter newline, smart-copy Ctrl+C) — without it the document capture-phase handler's `preventDefault()` (which never stops xterm) let each chord ALSO write its raw byte/escape sequence into Pane B's live PTY on top of whatever the app action did to Pane A (COD-153). Ctrl+Z is swallowed unless Pane B's own session is `mode === 'shell'`, mirroring terminal-ui.js's reasoning: in a plain shell it is the user's own job-control tool, everywhere else it silently suspends an unattended agent loop. Shift/Ctrl+Enter POSTs to `/api/sessions/:id/send-key` (`{key:'S-Enter'|'C-Enter'}`, tmux `send-keys -H` for a real 0x0a) targeting THIS pane's own `sessionId` rather than the primary pane's `activeSessionId` — without it xterm's plain `\r` would submit an incomplete prompt instead of adding a line to it. Smart-copy Ctrl+C/Ctrl+Shift+C is re-implemented against `this.terminal` (Pane B's own) rather than reusing `app.copyTerminalSelection()`, which reads Pane A's terminal and would copy the wrong pane's selection; Ctrl+Shift+C never falls through even with nothing to copy, mirroring terminal-ui.js's own `ev.shiftKey` branch. ⚠️ **This is a UX-parity fix, not an interrupt-safety one** — verified live in a real browser: xterm's `evaluateKeyboardEvent` routes a shifted ctrl-letter into a branch that assigns `c.key` only for two special cases (`_`→US, `@`→NUL), so it emits no data for Ctrl+Shift+C at all regardless of any application gate; a synthetic keydown with the gate removed produces zero WS frames, proving no accidental interrupt reaches the PTY either way. What gating the whole copy block on `hasSelection()` (an earlier draft) actually cost: with no selection, a selection-less Ctrl+Shift+C fell straight to `return true`, silently ceding the keystroke to the BROWSER's own handling (e.g. Chrome's Inspect-Element binding) with no feedback and no copy attempt — Pane A always intercepts it. Ctrl+V stays on xterm's own default paste, since Pane B has no image-paste trap to route it to. ⚠️ `buildSplitPickerSessions()` also excludes any session with `pid === null` (an exited CLI, a crash-looped session whose breaker tripped, a restore that never re-attached): Pane B has no equivalent of `selectSession()`'s auto re-attach POST, so a pane opened onto one has nothing reading its tmux pane — no `terminal` events ever arrive, and `Session.write()` silently drops every keystroke with no ack either way, so the loss is invisible behind a socket that reports healthy. Pane B's input frames deliberately carry no `cid`/`seq` (`ws-routes.ts` supports that), matching the no-overlay/no-IME "deliberately plainer" list above, since it has no exactly-once delivery layer to key them against. Design: `docs/split-pane-sessions-plan.md`. @@ -453,6 +788,158 @@ Anatomy: `.set-shell` → `.set-shell-head` (title + `.set-head-actions`) + `.se **The hinge is a reserved region, and the CSS is inert by construction.** The CSS Viewport Segments media features report two segments only while a foldable is actually bent; `--fold-inline-end` / `--fold-block-end` (end of styles.css) measure the strip to keep clear from the LEADING segment (`env(viewport-segment-right 0 0)` and `env(viewport-segment-bottom 0 0)`, physical sides in every text direction) and are `0px` on everything else. The seven centred overlays are all `position: fixed; inset: 0` flex boxes; each shrinks its CONTENT box with padding rather than the box itself, so the backdrop still covers the far side of the fold and still swallows taps there. The cascade traps, each measured in headless Chromium against styles.css + mobile.css in index.html link order: (1) a later `padding-right` longhand beats the earlier `padding` shorthand it composes with, so each fold rule re-states the overlay's own gutter; (2) a base gutter that a LATER `@media` block overrides needs its own fold restatement in that block on a ZERO base, since the unconditional rules at the end of the file otherwise put the gutter back (the phone path picker and path preview drop to `padding: 0` under 600px and came back at 16px and 18px on every phone); (3) a compound rule written to outrank a mobile.css shorthand must be scoped to the band where that shorthand applies, because outside it there is no gutter to compose with (the palette's `.modal.command-palette-modal` added 0.75rem at 393, 900 and 1400px and pushed the shell 6px off centre, and inside the band it lost its bottom gutter to the shorthand until it restated that side too); (4) mobile.css loads AFTER styles.css, so a same-specificity rule there wins under its own breakpoint (the response viewer's `max-height: 92dvh` under 600px beat the tabletop cap, which now has a twin at the end of mobile.css). `test/foldable-layout.test.ts` DERIVES the overlay list from the stylesheet, simulates the cascade across both files at every breakpoint with and without the fold rules, and requires the two results to differ by exactly the fold strip, so each of the four fails there instead of on hardware nobody has. +Further detail: the design follows Apple's [Designing for iPhone Duo](https://developer.apple.com/design/human-interface-guidelines/designing-for-iphone-duo). The sticky keyboard latch survived until the device was opened again, and rotating any phone hit it too. ⚠️ Physical sides, not logical ones: dialogs sit in the LEFT segment (the TOP one in tabletop pose) in every language, because the HIG keeps Duo's side controls on the same physical edge in RTL. The palette needs the compound `.modal.command-palette-modal` because mobile.css loads later and pads it with a shorthand under 768px; its compound rule lives INSIDE the 600-768px band whose mobile.css shorthand it composes with. The fold measurements were taken at 393 and 500px widths. Device profiles: `iPhone Duo (outer)` / `iPhone Duo (inner)` in `test/mobile/devices.ts`; unit coverage in `test/viewport-shape-change.test.ts`. + +### Terminal touch gestures: link taps and text selection + +**Terminal touch gestures: link taps and text selection**: on a touch device xterm's own handlers see neither — `touch-action: none` plus touchstart's preventDefault suppress the browser's compatibility mouse events, `_installMobileTapMouseGuard` drops the trusted ones that still arrive, and the synthetic `mousedown`/`mouseup` pair dispatched for mouse REPORTING goes to the `.xterm` root, an ANCESTOR of the screen element the linkifier and SelectionService listen on. So both gestures are driven explicitly. + +⚠️ **A tap activates the link under it** through the SAME provider that feeds the hover linkifier (`_terminalLinkAtPoint`, containment mirroring xterm's `_linkAtPosition`), synchronously inside `touchend` — that is what keeps the user gesture `window.open` needs — and BEFORE any mouse report, mirroring `_handleDesktopTerminalClick`'s skip for a hovered link. Two rows keep their meaning: the caret's logical line (`_tapIsOnCaretLine`, where a tap places the cursor in text the USER typed) and TUI-owned rows (`_isActionableMobileTerminalTap`, answering a dialog). + +⚠️ The caret line is the boundary rather than the tap INTENT, because a shell classifies every tap as `'input'` and gating on that would leave every URL in shell output inert. + +⚠️ **Long-press selects** by driving xterm's public `select()` (renderer-independent — under WebGL the glyphs are pixels and native selection cannot exist), drag or a further tap extends, and Copy goes through `copyTerminalSelection()` for its execCommand fallback on plain-HTTP installs. Three guards are load-bearing and each came from a real phone: the compat mouse pair after `touchend` (xterm focuses on mousedown and SelectionService resets the model there, so the keyboard sprang up and the selection vanished on lift), the platform's own ~500ms long-press (Android Chrome focuses the nearest editable element — the helper textarea — through no event a handler can preventDefault, so a bounded focus guard blurs it and `contextmenu` is suppressed for the gesture window), and `copyTerminalSelection()`'s closing `terminal.focus()` (right on desktop, wrong on a phone). + +Tests: `test/terminal-touch-tap.test.ts`. + +### Entrance animations + +**Entrance animations** (`entrance-animations.js`, all OFF by default): opt-in animations for the four things that appear when work starts, chosen per surface via `data-tab-anim` / `data-term-anim` / `data-win-anim` / `data-line-anim` on `<html>`. Defaults are the `legacy` theme, so an untouched install behaves exactly as before and every hook short-circuits on its first line. Persisted to its own `codeman:*Anim` localStorage keys (per-device, deliberately NOT in the `.strict()` `SettingsUpdateSchema`); picker in App Settings → Appearance, full per-surface lab at `?animlab=1`. + +⚠️ Tabs and connection lines are **destroyed mid-animation** on every re-render (`_fullRenderSessionTabs()` replaces the strip's innerHTML; `_updateConnectionLinesImmediate()` does `svg.innerHTML = ''`), so both are tracked by id and re-applied to the fresh element with a **negative `animation-delay`** to resume rather than restart. + +⚠️ The terminal-pane styles may animate **transform / opacity / clip-path only**, xterm's FitAddon derives rows+cols from `getComputedStyle(parent).width/height`, so animating width/height/padding there would resize the PTY; `test/entrance-animations.test.ts` pins that property allowlist, plus the rule→keyframes→theme-option chain a style silently does nothing without. + +⚠️ **`blur` is the ONE style that puts a `filter` on the terminal container**, against the standing rule, because every alternative was measured against a live xterm and does not work: a `backdrop-filter` veil on `::before` blurs perfectly while STATIC and Chrome silently drops the backdrop the moment ANY animation runs on that pseudo-element (the veil computes `blur(15.3px)` and the text behind it stays razor sharp), and driving the radius from rAF buys the same full-screen blur per frame plus main-thread work. The cost the rule exists to avoid is inherent to blurring a terminal, so the style buys it knowingly: opt-in, OFF by default, one ~520ms run per session open, class straight back off, `will-change` still unset. Worst-case price, headless SwiftShader with no GPU: frame deltas 16.7ms → 33.3ms for the run, against 16.7ms flat for `fade`. Do not generalise it — a second filtered terminal style needs its own measurement. + +⚠️ The `blur` connection line animates `filter` too, so both kinds of line hold their glow in **`--line-glow`** and both of its keyframes say `blur(N) var(--line-glow)`: the function lists then match and interpolate, instead of the glow vanishing for the run and popping back (a lineage line's glow is a different colour entirely, set per element). Its 100% frame deliberately omits `opacity` so the endpoint comes from the element's own resting value — 0.9 subagent, 0.72 lineage, 0.95 working — which is what `line-enter-fade`'s hardcoded 0.9 gets wrong. + +⚠️ Window styles other than `beam` transform the window, which moves the rect its connection line is aimed at; `beam` deliberately animates opacity/filter only so its line can draw toward a stable target. + +### Mobile tab strip scrolling + +**Mobile tab strip scrolling** (issue #257): under 768px the tab strip is a horizontal scroller (desktop wraps to a second row instead), so the active tab can sit off-screen. Three rules keep it reachable and they only work together: `_updateActiveTabImmediate()` scrolls the selected tab into view via `computeTabScrollLeft()` (pure, in constants.js) using **rect math on the strip's own `scrollLeft`**, never `scrollIntoView()`, which would also scroll the document under a fixed header; `_fullRenderSessionTabs()` **restores `scrollLeft`** across the `innerHTML` rebuild, since ambient rebuilds (a task badge appearing, a session created elsewhere) otherwise snap a mid-swipe strip back to 0; and it re-reveals the active tab **only when it changed** (`_lastRenderedActiveTabId`), so browsing the far end of the strip is not undone by background renders. + +⚠️ **The ACTIVE tab is the only one with action icons, and on a phone they can eat it**: `.session-tab.active .tab-name` reserves `min-width: 44px` in the phone block (under 600px), because a short session name rendered a 13px label against a 50px gear+close cluster, putting the tab's geometric CENTRE on the gear, so a thumb aiming at the tab opened Session Options instead of switching (measured at 360/393/430px; only long names cleared it). + +⚠️ **The floor is set by the 10th tab onward, not by the tabs you can see**: `.tab-number` renders only for `_tabIdx < 9`, so tab 10 loses 16px + a gap off its left and its centre sits 10px further right. The centre clears the icons when `reserved > icons + rightEdge - leftRunUp - gap` (= 50 + 9 - 17 - 4 = **38px**), hit-testing snaps to whole pixels so 39px still lands on the gear, and the practical floor is 40px — a NUMBERED tab clears it at 20px, which is exactly why reasoning from the tabs on screen would put the centre back on the gear. `test/mobile-tab-tap-zones.test.ts` recomputes that inequality from the stylesheet, so widening the gear or the padding fails there rather than on a phone. The guarantee is centre-off-the-ICONS, not centre-inside-the-label (on a numberless tab it lands in the gap between them, which still switches). Non-active tabs keep their icons hidden and stay tappable end to end. + +⚠️ Mobile no longer hoists the active session to the front of the strip: that reordering ran on full renders only, so tab order flipped depending on which render path fired, and it renumbered the Alt+N badges. Scroll-into-view replaces it; do not reintroduce it. + +### Phone overview home screen + +**Phone overview home screen** (`mobile-overview.js`, phones only, per-device `mobileOverviewEnabled`, default ON): under 600px the "C" logo shows a session overview (NEEDS YOU / CURRENT SESSIONS / PAST SESSIONS) instead of the welcome overlay; tablet and desktop are unchanged. The branch lives in `showWelcome()`/`hideWelcome()` (terminal-ui.js) behind `shouldUseMobileOverview()`, which is **width-driven** (`getDeviceType() === 'mobile'`) because this is a layout decision, unlike the settings namespace which stays handheld-based. + +⚠️ The container ships with the `hidden` attribute and only this module removes it: never give `.mobile-overview` a bare `display` rule, since desktop does not load `mobile.css` (`media="(max-width: 1023px)"`) and would then render it unstyled. Live re-renders ride on the tail of `_renderSessionTabsImmediate()` (every state change it needs already funnels there); PAST rows come from one `_fetchUnifiedSessions(60)` per home-screen visit and resume through the shared `resumeHistorySession()`, so they behave exactly like the welcome screen's Resume list. + +⚠️ Two things must stay in lockstep with surfaces outside this module, because divergence reads as a bug rather than a style: the split Run button carries the **toolbar's own classes** (`btn-toolbar btn-run mode-<backend>` / `btn-run-gear`) so the per-backend gradient and the light-skin overrides apply unchanged (mobile.css must therefore set no `background`/`color` on it), and row status uses the **session-tab language** (green dot when fine, `pulse` while working, yellow blinking row when waiting for input, red blinking row when a question is pending, mirroring `tab-alert-idle`/`tab-alert-action`). The picker mirrors the toolbar run-mode menu (`setRunMode()` + `run()`, `openWebviewFromMenu()` for saved dashboards) and deliberately omits its Recent-Sessions block, since PAST SESSIONS is that. Status pills carry `data-i18n-skip` (generic words like "idle" collide with state strings elsewhere). + +### Desktop home tab rail + +**Desktop home tab rail** (`home-sessions.js`, desktop only): the welcome overlay centers ~560px of content in a ~1400px window, so its left gutter is dead space; it carries the open tabs as a rail **docked flush to the left edge, full height** (a vertically centered card floating mid-gutter read as debris). Rows are in **overview order** (see Home-screen session order), and each carries a **created** stamp plus the **state duration** the order is computed from (`created 3d ago · working 12m`, word and anchor from `_mobileOverviewSince()` so both home screens say the same thing). A rail sorted by a number it does not show reads as arbitrarily shuffled, and a working row's plain last-active stamp always says "just now". + +⚠️ The number badge is the **Alt+1..9 index**, i.e. the position in the TAB STRIP, so on a sorted rail it deliberately does NOT run 1,2,3 downward: it names a shortcut, not a row position, and renumbering it to look tidy would make every badge lie. State classification is REUSED from mobile-overview.js (`_mobileOverviewState`/`_mobileOverviewCaseFor`), which is why the module loads after it. + +⚠️ The rail is `position: absolute` so the centered content never moves, which is exactly why it needs a **width gate in two places** — `HOME_SESSIONS_MIN_WIDTH` (1180) in the JS plus a `max-width: 1179px` media query as the backstop for a resize that outruns the matchMedia listener; drift between them means a rail overlapping the search panel, and `test/home-sessions.test.ts` pins them equal. ⚠️ `.home-sessions` is `display: flex`, so `[hidden]` must be re-asserted as `display: none` or the module's only visibility lever does nothing. ⚠️ Size scales with the viewport off **one knob**: `width: clamp(250px, 19vw, 430px)` plus a fluid `font-size` on `.home-sessions`, with every child sized in `em` — reintroducing `rem`/px type inside the block silently breaks the scaling, and widening the clamp past the gutter reintroduces the overlap the gate exists to prevent. + +The age stamps are refreshed **in place** by a 20s clock (`_tickHomeSessionsTimes()`, disarmed in `hideHomeSessions()`), never by re-rendering, which would restart every row's blink and working ring. Working state is deliberately byte-identical to the phone's: pulsing green dot + the `tab-load-spin` ring reused from the tab strip + the same green halo (added to `.mobile-overview-dot--working` at the same time), so "working" reads the same on every surface; **idle** is deliberately NOT that green — dot and pill mix toward `--text-muted` so a glance separates running from sitting. Live re-renders ride the tail of `_renderSessionTabsImmediate()` alongside the phone overview. + +### Home-screen session order + +**Home-screen session order** (`CodemanSessionOrder` in constants.js, pure + unit-tested in `test/session-overview-order.test.ts`): BOTH home screens (phone overview and desktop rail) order rows through this ONE comparator, because they list the same sessions and must answer "which of these wants me next?" the same way. Rank is `needs` → `error` → `waiting` → `working` → `idle` → `done`. + +⚠️ **The tiebreak flips direction halfway down**: states a session is still IN sort **oldest-first** (blocked longest / running longest = most urgent), states it has STOPPED in sort **newest-first** (the session that just went quiet is the one you came back for). + +⚠️ The running group keys off **`lastSubmitAt`** (the pane's last Enter), never `lastActivityAt`: a working Claude pane repaints about once a second, so its last-activity stamp is always "now" and would rank every running turn as freshly started. A working pane with no submit stamp falls back to last activity, which lands it at the SHORT end of the group rather than falsely leading it. + +⚠️ A **0 stamp means "unknown", not "the epoch"**, and it sorts last within its state either way, or a brand-new session would head every oldest-first group. Final tiebreak is the user's tab order (`orderIndex`), so the list is deterministic and cannot shuffle between renders. The tab strip itself is NOT sorted by this; it stays user-ordered and drag-reorderable. + +### Terminal font weight + +**Terminal font weight** (`terminalFontWeight` / `terminalFontWeightBold`, per-device, default = xterm's own `normal`/`bold`): bold text on the theme's default foreground carries exactly ONE cue, the weight step. Claude Code marks its markdown bold with a bare `ESC[1m` and no colour change, and xterm substitutes a bright colour for bold only when the foreground is a palette index 0-7, so the substitution never fires there. A two-face family keeps that step small and 400 stays 400 whatever family is chosen, which is why the NORMAL slot is settable at all. `CodemanTerminalFont.resolveWeights()` (constants.js, pure) resolves both slots, each against **its own** xterm default, so an unset bold weight can never inherit `normal`. + +⚠️ **The `@font-face` descriptor, not the file, is what the browser synthesizes from**: `fonts/jetbrains-mono-variable.woff2` carries a `wght` axis of 100-800, and while `styles.css` declared it `400 700` every weight below 400 rendered identically to 400 and 800 identically to 700 — measured — so the setting was a no-op for anyone without Fira Code or Cascadia Code installed, which is most installs. It is declared `100 800`; re-narrowing it silently guts the feature (`test/terminal-font-weight.test.ts` pins the range). + +⚠️ A live save must reach **both echo overlays** (`refreshFont()` — they cache `terminal.options.fontWeight` and paint it into their spans, so typed characters otherwise keep the old weight, most visible on a phone) **and open Agent Teams panes** (they read their options at construction, exactly like `applyTerminalSkin()` propagates). + +⚠️ `_awaitTerminalFont()` is deliberately untouched: `CharSizeService` measures through the CSS `font` shorthand, which RESETS the weight, so the measured face is always the 400 one and a weighted descriptor would ask for nothing new. + +### Shell keyboard accessory bar and one-shot Ctrl + +**Shell keyboard accessory bar + one-shot Ctrl** (issue #262, `keyboard-accessory.js`): a **shell**-mode session automatically swaps the mobile accessory bar for terminal controls (Ctrl, Esc, Tab, four arrows, paste, dismiss); every other mode keeps the agent bar. `setMode()` now records the user's `extendedKeyboardBar` preference as the **base** layout and `refreshForActiveSession()` (called from `selectSession`) resolves base-vs-shell, so a settings save during a shell session cannot yank the bar away and switching back restores the user's choice. + +⚠️ **Ctrl is a ONE-SHOT modifier applied in `terminal.onData`, not in a keydown handler**: a virtual keyboard emits no usable key events, so the character only exists as onData text. The hook sits AFTER `shouldSuppressTerminalQueryResponse` (xterm answers DA/CPR through onData too, and one of those would silently spend the modifier) and BEFORE every send path, so the control byte follows the normal control-char route. + +⚠️ **Not every onData chunk is a keystroke**, and the query filter is not enough on its own: xterm ALSO emits mouse and focus reports on its own initiative, so the hook skips them via `isTerminalFocusOrMouseReport()` (they still reach the PTY, they just don't count as the next key). The mouse half is live — a shell session keeps the NARROW strip, so mouse DECSETs reach the browser and one tap while vim/htop runs spent the armed modifier silently (measured). The focus half is defense in depth: `FOCUS_ESCAPE_FILTER` in `session.ts` strips `\x1b[?1004h` from every PTY read, so `sendFocusMode` never turns on today; if it ever did, the bar's own post-key refocus would emit `\x1b[I` and eat the modifier before the user typed. + +⚠️ It must disarm on ALL of: use, second tap, any other accessory key, session switch, keyboard dismissal, and a layout swap; a modifier left armed turns the next innocent keystroke into a control byte. + +⚠️ **onData is not the only input path** — with `cjkInputEnabled` on, the CJK textarea owns the keyboard (onData returns early for everything it swallows, and the focus router sends `terminal.focus()` there, which is where the bar refocuses after every key), so `_handleCjkInput()` applies the modifier too. It is that module's single choke point to the PTY, so one call covers typed characters, IME flushes, Enter, backspace and arrows. Without it an armed modifier could neither fire NOR be spent, and survived to a later keystroke. + +Mapping is `ctrlByteFor()` (`code & 0x1f` over @A-Z[\]^_ and a-z, plus Ctrl+Space=NUL / Ctrl+?=DEL); characters with no control equivalent pass through unchanged, like a hardware keyboard. + +⚠️ The armed style is `.accessory-btn.accessory-btn-ctrl.armed` (0,3,0) in BOTH stylesheets, and it cannot outrank mobile.css's light-skin repaint at **(0,3,1)** (`:is()` inherits its most specific argument, and that list holds `.btn-toolbar.btn-shell`) — so that rule excludes the state by hand as `.accessory-btn:not(.armed)`. Without the exclusion the armed button renders identically to a resting one on all four light skins, which is worse than no armed style at all. + +### Mobile prompt composer + +**Mobile prompt composer** (PR #444, the first slice of #359, `keyboard-accessory.js`): the agent bars' Paste key is now **Compose**, a dialog with a native multiline textarea (autocorrect, autocapitalize, spellcheck) where Enter adds a line and only **Send** submits; the shell bar keeps the direct Paste dialog, since shell input is not an agent prompt. Opening it ADOPTS the whole editable terminal prompt (`_takePendingLocalEcho`): the local-echo overlay's pending text has never reached the PTY, but the flushed prefix has, so that prefix is erased with backspaces counted in CODE POINTS (`Array.from(text).length`; measured on Claude Code 2.1.278, `a` + emoji + `b` takes three, and the UTF-16 count sent four and ate the neighbour; `clearTerminalInput()` in terminal-ui.js moved with it). + +⚠️ **Drafts are per-session and in memory only** (`_composerDrafts`, never persisted: prompts routinely carry secrets, and persisting them would need the 0600 treatment the intent store gets). Every non-Send exit (Cancel, backdrop, Escape, "Use terminal keyboard") leaves the taken text ONLY in the draft, with the dot on the key (`has-draft`, kept in step by `_syncComposerDraftIndicator`) as the signal that the terminal prompt is empty on purpose, and `_cleanupSessionData` discards the draft with the session. + +⚠️ **Delivery is a hand-built bracketed-paste frame** (`\x1b[200~` + the text with newlines mapped to `\r` + `\x1b[201~`, byte-identical to what `terminal.paste()` would emit) through `_sendInputAsync` WITHOUT `{ useMux: true }`, plus a SEPARATE Enter 120 ms later WITH it (codex drops keys that share a PTY read with a bracketed paste). Not `terminal.paste()`, for two reasons: xterm's `bracketedPasteMode` mirror is false for every session after a tab switch or reload (`terminal.reset()` in the replay re-clones the DEC modes, the tmux capture carries no `?2004h`, and tmux never forwards the pane's DECSET to a client after attach), so a paste through xterm would go out unbracketed and the CLI would submit at the first `\r`; and xterm's onData is where the local-echo paste branch flushes pending overlay text AHEAD of the block, text the composer has already taken and erased from the PTY, so the frame goes straight to the wire with the composer as the prompt's only owner. The unconditional frame is safe for a CLI that never enabled DECSET 2004 because tmux does the gating (measured against a live pane: markers stripped for a `cat -v` pane, forwarded intact to a process that had emitted `?2004h`). + +⚠️ The frame must never take the mux fallback: `TmuxManager.sendInput()` strips every `\r` and `\n`, which welds the lines together and submits them. + +⚠️ The size guard sits on the `MAX_INPUT_LENGTH` boundary (64 KiB of UTF-16 code units): `ws-routes.ts` drops a longer frame WITHOUT an ACK, which would wedge the durable queue, so `_composerMaxLength` is derived from that limit minus both markers and an oversized prompt stays a draft with a toast. + +⚠️ `.prompt-composer-overlay` is a `.paste-overlay` with a gutter of its own (a `padding` shorthand whose bottom is 12px plus the safe area), and the unconditional `.paste-overlay` fold rule at the end of styles.css is a later longhand at the same specificity, so it ERASED that gutter (measured at 393x852: `padding-bottom: 0px` flat, and the hinge strip REPLACING the gutter with the fold variables set): the composer has its own restatement after the fold rules, the dialog's `max-height` subtracts `--fold-block-end`, and `test/foldable-layout.test.ts` lists the composer in `ELEMENTS` by hand, because its derived overlay list keys on rules that declare `position: fixed; inset: 0` themselves. + +Tests: `test/mobile-prompt-composer.test.ts` (in the CI gate, deliberately not under `test/mobile/**`). + +### Dismissing the on-screen keyboard + +**Dismissing the on-screen keyboard** (PRs #279/#280, `terminal-ui.js`): the terminal parks focus on a hidden textarea that nothing used to release, so TWO gestures now blur it, and they own different regions. + +**(1)** `_installMobileKeyboardDismiss()` — a document-level `touchend` that fires only while the terminal input actually holds focus, **never inside `#terminalContainer`** (tap classification owns that) and **never on a control** (`MOBILE_KEYBOARD_DISMISS_EXEMPT_SELECTOR`, matched with `closest()` so an icon inside a button counts). Session tabs are covered by the selector's `[tabindex]:not([tabindex="-1"])` arm, which is what stops a tab tap from blurring and then being re-focused by `selectSession()`. + +**(2)** In `_handleMobileTerminalTap`, a second tap on **inert `content`** (`startedWithTerminalFocus`) blurs instead of re-focusing. ⚠️ Scoped to `content` on purpose: the prompt row (`input`) keeps focus-then-position so a second tap still places the caret, and actionable rows blur earlier via `_isActionableMobileTerminalTap`. + +⚠️ **A scroll ends in `touchend` too** — dismissing there closes the keyboard and drops the composer mid-read, so travel is tracked from `touchstart` and multi-touch is never a tap. Both classifiers MUST share one threshold: `initTerminal`'s `TAP_THRESHOLD` reads `MOBILE_KEYBOARD_DISMISS_TAP_SLOP`, since a gesture the terminal calls a scroll and the dismiss handler calls a tap is exactly that bug. + +⚠️ **The gate excludes `test/mobile/**`, so CI cannot see the only test covering (1)** — run `npm run test:mobile -- test/mobile/keyboard.test.ts` by hand and diff the FAIL list against master. (Not `npm test --`: the gate's config excludes that path, so a file filter pointing into it matches nothing and exits green having run zero tests.) That blind spot is why merging the two PRs, which conflicted semantically but not textually, produced a red suite with two green CI checks. + +### SSE staleness watchdog + +**SSE staleness watchdog** (`computeSseStale()` in constants.js, `_checkSseStale()` + a 5s interval in app.js): an `EventSource` that stops delivering does not always error, so `onerror` never fires, the header dot stays green, and every SSE-driven surface (tab status dots, sessions created on another device, renames) freezes until the user reloads. + +⚠️ The 15s server keepalive was an SSE **comment** (`:keepalive`), and comments are **invisible to `EventSource` by spec**, so there was nothing a client could observe: it is now the named `sse:heartbeat` event (`cleanupDeadClients()`, sse-stream-manager.ts), which is exactly why the frame had to change type. + +⚠️ Staleness is judged **only while the status is `connected`** and the device is online; that guard is the loop breaker, since a forced `connectSSE()` leaves `connected` immediately and cannot re-fire while a reconnect is in flight. + +⚠️ The liveness stamp is applied inside `addListener` itself, so every registered handler (the `_SSE_HANDLER_MAP` wrappers AND the directly-registered ones) feeds it from one place; the heartbeat's own listener is a no-op that exists **only** to be registered, since `EventSource` drops named events nobody listens for. + +⚠️ The watchdog interval is cleared at the top of `connectSSE()` and nowhere else (its only teardown path); clearing it elsewhere stacks intervals. Recovery needs no new sync path: the reconnect re-runs `handleInit` → `_resetAllAppState()`. The forced reconnect logs one diagnostic line, because a middlebox that strips heartbeats presents as "silently reconnects every 45s". + +### PTY and browser terminal geometry + +**The PTY and the browser terminal must never disagree about size** (issue #464, `syncTerminalGeometry()` in terminal-ui.js): Claude Code's TUI wraps its frame at the width the PTY reported and erases the previous frame by walking the cursor up the number of rows it BELIEVES that frame occupied. A browser terminal of a different width makes each logical line take more physical rows than Ink counted, so `eraseLines(n)` clears too few and the new frame paints over rows nothing erased — the doubled lines and half-overwritten prose in #464, **measured** against this repo's xterm (a 120-column PTY against a 62-column terminal renders every wrapped line twice; `test/terminal-pty-geometry.test.ts` pins it, and pins the clean render at matching widths so the assertion cannot pass against code that fixes nothing). ⚠️ **`fitAddon.fit()` is NOT the way to resize this terminal.** It resizes xterm to `proposeDimensions()` RAW while every server-facing path reports those floored at 40x10, so whenever the floor bit the two diverged silently — **measured in Chrome at 430px**: font size 44 proposed 13 columns, the server was told 40, and xterm stayed at 13. `syncTerminalGeometry()` fits, floors and applies as one step and is the ONE function that may change the size; `test/terminal-pty-geometry.test.ts` sweeps the seven modules that touch the main terminal (terminal-ui, mobile-handlers, app, ralph-panel, settings-ui, tab-rail-resize, notification-manager) for a bare fit; the split pane and teammate terminals own their own sizing and are out of scope. ⚠️ **Withhold the fit wherever you withhold the SIGWINCH.** `throttledResize` (virtual keyboard up) and `sendResize` (session detached into its own window) used to reflow locally and skip only the server write, which is the one combination that cannot be right. ⚠️ **A font change is a geometry change**: `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size and told the server nothing, so raising the font on a phone left the CLI wrapping at the old column count (`_refitAfterCellSizeChange`). ⚠️ **Resize is no longer write-only.** `Session.resize` DECLINES a small-viewport request while a desktop connection holds an active sizing claim and says nothing, so both transports now answer with `Session.ptyGeometry` (`{"t":"zc"}` on the socket, the body of the resize POST) and `_onPtyGeometryReport` adopts it — a terminal that keeps a WIDTH the PTY refused renders GARBLED, not merely wrong-sized. ⚠️ While that refusal stands (`_paneWidthRefused`), a resize ASKS for the container's width without applying it (`_geometryForResizeRequest`: rows follow the container, columns stay at the PTY's): fitting first re-wrapped the whole buffer to the container and back on every 30s mobile retry, and ran the scrollback clear for a resize that brings no redraw. `selectSession` clears the flag, since it belongs to the previous pane. ⚠️ **`ptyGeometry` is null without a live pane**, never the field values: they are seeded at spawn (`_notePtySpawnGeometry`, so a reattached pane reports its tmux window's real size) and moved by `resize()`, but a dead-pane session still holds the constructor defaults of 120x40 or a gone pane's size — reporting those made a client adopt a size no process was ever told and claim another device owned the pane when none existed. ⚠️ **COLUMNS ONLY.** Adopting the PTY's ROWS was a regression: a phone taking a desktop's 43 rows into a viewport with room for 18 painted an `.xterm-screen` far taller than its container, xterm's own viewport then had nothing to scroll, and the CLI's input line sat below the container with no gesture able to reach it — output visible, typing invisible, for as long as the claim stayed hot. Width is the axis the wrap arithmetic depends on; rows only decide how much is on screen, and keeping the local count keeps the composer at the bottom of a viewport that scrolls. ⚠️ **`.term-overflows-x` keys on what does not FIT, not on a PTY mismatch**, because the 40-column floor widens the terminal past a narrow container with the PTY agreeing throughout (measured at 360px: font 18 paints 433px, font 24 paints 578px — 38% unreachable, and `increaseFontSize` reaches 24 in two taps). `_syncTerminalOverflowAffordance()` MEASURES `.xterm-screen` against the container on the next frame rather than deriving it from cell arithmetic. ⚠️ That rule sets **both** overflow axes — mobile.css loads later with `.terminal-container { overflow: visible }`, and a bare `overflow-x` would leave overflow-y computing to `auto`, handing the browser a vertical scroll container the terminal's touch handler does not know about — and sets **no `touch-action`**: the terminal's own touchmove handler pans it (`canPanHorizontally`), because `touchstart` preventDefault()s every 'content' tap and that cancels a native `pan-x` before it starts (measured: a 140px swipe reached scrollLeft 141 without that preventDefault and 0 with it). Removing the class returns `scrollLeft` to 0, so a resolved mismatch cannot leave the pane parked off-screen. + +### Terminal resilience: replay clears, renderer liveness, fetch deadlines + +**Terminal resilience: replay clears, renderer liveness, fetch deadlines**: three rules that each close a way the terminal silently stops being correct. ⚠️ **A replay clear MUST be in-stream, never `reset()`/`clear()`.** xterm's `write()` is asynchronously queued while `Terminal.reset()` is synchronous and, per upstream, "does not clear input buffers and does not reset the parser" — so bytes queued just before a reset are parsed AFTER it and fuse into the snapshot written next. **Measured** against the real xterm in this repo: `write('p8'); reset(); write('rmissions')` renders `p8rmissions`; the queued `\x1bc` renders `rmissions` and clears scrollback. `_resetTerminalForReplay()` (app.js) is the ONE clear, a single queued `\x1bc` (RIS), and all three replay paths go through it; RIS rather than `\x1b[3J\x1b[H\x1b[2J` because the erase leaves modes, charsets, scroll regions and SGR state alone. Callers may still chunk the content — ordering in the queue is what matters, not writing it in one call. ⚠️ **The renderer watchdog reads xterm privates and CANNOT be covered by the gate.** `_kickRenderer()` (terminal-ui.js) cancels a stale `_core._renderService._renderDebouncer._animationFrame` and forces a repaint. **Verified against xterm 6.0.0** (jsdom, after `open()`): the field path resolves, a forced stale handle genuinely makes `refreshRows` a no-op, and the kick schedules a fresh frame. **Reasoned, not reproduced here**: the premise that iOS discards scheduled rAF callbacks when a PWA backgrounds, which is what leaves the handle stale — that half wants a real-device pass. Codeman has exactly ONE xterm for the whole page load, so one backgrounding would wedge it until a reload. `_renderService` only exists after `open()`, which needs a real DOM, and the gate runs in node — so `test/xterm-private-api.test.ts` pins the RESOLVED lockfile version (not the `^6.0.0` range, which a real upgrade slips through) and a bump means re-verifying by hand. Every access is optional-chained on purpose: a renamed field must degrade to a no-op, never throw on a 2s timer. ⚠️ **Every terminal capture carries a deadline, and the helper reads the BODY** (`_fetchTerminalCapture`, app.js). `await fetch()` settles on response HEADERS, so clearing the timer there leaves the body — the multi-megabyte `?full=1` capture this exists for — unbounded: **measured** at 4026ms under a 1000ms deadline before the fix. The helper therefore returns `{json, headers, headersAt}` rather than a `Response`, and `_terminalCaptureInflight` is scoped the same way so a body still streaming counts toward a capture starting beside it. It degrades to a plain fetch where `AbortController` is missing — the deadline is a safety net, not a dependency. Tests: `test/terminal-resilience.test.ts` (pure decisions), `test/xterm-private-api.test.ts`. + +### WebSocket output-gap reconcile + +**WebSocket output-gap reconcile** (`_wsOutputGapSession`, app.js): terminal OUTPUT frames carry no sequence number (input frames do — `seq`+`cid`, at-most-once, ACKed), so a dropped socket leaves a hole nothing replays. ⚠️ **The gap is narrower than "the device went offline"**: if the network drops, SSE drops with it and `handleInit`'s keepTerminal branch already calls `_onSessionNeedsRefresh`. The uncovered case is the WS dying while SSE stays up (half-open socket, proxy idle-timeout, ping timeout), because `_onSSETerminal` discards every SSE terminal frame while `_wsReady` is true and `_wsReady` only flips in `ws.onclose`. Reaching `onclose` at all means the drop was unintentional (`_disconnectWs` nulls the handler first), so the session is marked and the next successful open reconciles. ⚠️ **The marker must be cleared by every path that repaints that session's buffer, and ONLY once one actually has.** `_markTerminalBufferReconciled()` is called from `selectSession` after its load, from `_cleanupSessionData`, and from `_onSessionNeedsRefresh` **at the repaint itself, not in its `finally`** — clearing on every exit meant a reconcile that threw, or hit the fetch deadline (the flaky link the marker exists for), dropped the gap with nothing to retry it. `ws.onopen` no longer clears it up front either, so a reconcile that is skipped or fails is tried again on the next open; re-entry is safe because `_terminalRefreshOwner` makes a second reconcile for the same session a no-op. `selectSession` loads the buffer and only THEN calls `_connectWs`, so without that clear the socket opening afterwards replays the whole buffer a second time on top of the one just written. Sequencing the output frames is the real fix and is not done. This is reasoned from the code path, not observed on a device. + +### Service worker precache and cache key + +**Service worker: precache and cache key are BUILD-GENERATED** (`sw.js` + `scripts/build.mjs`): the build content-hashes assets and rewrites two exact declarations in `sw.js` — `const BUILD_ID = 'dev';` and `const HASHED_ASSETS = [];`. ⚠️ **Each must appear exactly once or the build THROWS**, which is deliberate: the list used to be hand-maintained with PRE-hash names, so every entry 404'd in production and `cache.add().catch(() => {})` hid it (15 of 23 verified failing against a running instance). The dev literals are valid on their own, so dev serves an unrewritten worker with an empty precache. ⚠️ **`caches.match` must pass `ignoreSearch: true`**: `renderIndexHtml` runs `cacheBustAssets`, which appends `?v=<mtime>` to every same-origin `.js`/`.css` reference INCLUDING content-hashed names, so the page requests `/app.<hash>.js?v=<mtime>` while the cache holds `/app.<hash>.js`. Without it no precached entry is reachable and the install downloads ~1.3MB that can never be served — once per deploy, since `CACHE_NAME` now carries the build id. That per-build key is what makes `activate`'s cleanup actually delete anything; it used to be the constant `'codeman-v1'`, so assets from every past release accumulated forever. Contract pinned by `test/sw-precache-manifest.test.ts`, which PARSES the `HASHABLE` list out of `build.mjs` rather than copying it. + +### Z-index layers + +**Z-index layers**: subagent windows (1000), split picker menu (1000, `.split-picker-menu`), plan agents (1100), mobile/tablet fixed header (1200, `mobile.css`), modals on ≤768px (1300 — must beat the fixed header or the modal close button is buried), log viewers (2000), connection-loss overlay (2500, above the fixed header and modals), image popups (3000), response viewer (5000, backdrop 4999), file-preview overlay (5100 — must outrank the response viewer, which can launch it; at its old 2000 a path clicked in the chat opened BEHIND the chat), toasts/path picker (10000+, deliberately above the preview), the custom-model center-status banner (10001, `.center-status-banner` — `[hidden]` must re-assert `display: none` over its own `display: flex`, same trap as `.home-sessions[hidden]`, or `dismiss()` leaves an invisible click-blocker dead centre on screen), the swap-confirm and context-warning modals (10010, `#customModelSwapConfirmModal`/`#customModelContextWarningModal` — must clear both the plain `.modal` z-index of 1000 and the center-status banner it can appear over), terminal touch-selection bar (900 — above terminal content and the local-echo overlay, deliberately BELOW floating agent windows so it can never cover their controls), local echo overlay (7). + ## Security layers ### Layer-by-layer detail @@ -477,6 +964,16 @@ Anatomy: `.set-shell` → `.set-shell-head` (title + `.set-head-actions`) + `.se Target: 20 sessions, 50 agent windows at 60fps. Limits in `src/config/`: terminal 32MB (see below), text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100. **Terminal history** (`src/config/terminal-history.ts`, COD-80): tmux history-limit 100k lines, PTY buffer 32MB max / 24MB trim (env `CODEMAN_MAX_TERMINAL_BUFFER`/`CODEMAN_TRIM_TERMINAL_TO`; the env-derived trim is clamped ≤75% of max — trim ≥ max would disable `BufferAccumulator` trimming entirely = unbounded memory); browser xterm scrollback stays a separate hardcoded 50k (`DEFAULT_SCROLLBACK` in constants.js — 100k/tab is a mobile-memory hazard). tmux <3.7 allocates history at pane creation, so `createSession()` sets the global default in the same command queue immediately before `new-session`; tmux 3.7+ instead creates the session and targets only that pane, because changing the global option can resize and trim unrelated live panes. A settings change resizes tracked panes only on 3.7+ and otherwise affects future panes; no version can recover lines already evicted. Settings keys `terminalScrollbackLines`/`terminalBufferMaxBytes`/`terminalBufferTrimBytes` are schema-validated but inert (only `tmuxHistoryLimit` is wired); `buffer-limits.ts` re-exports the defaults. Text/message limits are env-overridable too (`CODEMAN_MAX_TEXT_OUTPUT`/`CODEMAN_TRIM_TEXT_TO`/`CODEMAN_MAX_MESSAGES`). **Image upload** (`image-input.js` / `config/buffer-limits.ts`): up to `_maxBatchImages` 20 images/batch (bounded concurrency 3), per-file `MAX_PASTE_IMAGE_BYTES` 50MB (env `CODEMAN_MAX_PASTE_IMAGE_BYTES`); the mobile camera-roll picker auto-downscales to fit before upload. **HEIC paste uploads** (#151): converted server-side to JPEG in a `worker_threads` worker (`web/heic-jpeg-worker.ts`, resourceLimits + 30s timeout) gated by `runWithConversionLimit()`; detection is magic-byte based (covers Android/MIUI HEIFs mislabeled as JPEG); headers declaring > 64MP are rejected 415 BEFORE decode (decompression-bomb guard). Deps: `heic-decode` + `jpeg-js`. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Anti-flicker pipeline: `docs/terminal-anti-flicker.md`. +### Process-tree walks are bounded + +**Process-tree walks are bounded** (`proc-tree.ts`, pure + unit tested): `collectDescendants(pid, byParent)` is the ONE descendant traversal, fed by a single cached `ps -eo pid=,ppid=` snapshot (`refreshProcSnapshot()` in tmux-manager.ts: in-flight-shared, async because `execSync`'s timeout cannot return at all while spawnSync waits on an unkillable child, and ANY error discards the result rather than caching a truncated `ps`, which would make whole subtrees invisible to the kill path). + +⚠️ **The unbounded version took a machine down** (2026-07-30): it ran `pgrep -P <pid>` once per node and recursed with no visited set, no depth limit and no node cap, so across ~28 adopted tmux trees the fan-out exploded while each `pgrep` blocked in the WSL kernel reading `/proc/<pid>/cgroup`, ending at ~13,000 `pgrep` processes in D-state, a load average above 13,000, and a machine recoverable only by restarting WSL, which cost every running session. Three properties make that impossible and each has a test: a cycle terminates (a real tree has none, a stale snapshot can still produce one), depth is capped (`PROC_WALK_MAX_DEPTH`), node count is capped (`PROC_WALK_MAX_NODES`). The fourth is structural: the function takes a snapshot and cannot spawn anything at all. + +⚠️ It lives in its own module because as a private method of `tmux-manager.ts` the regression test had to keep its own COPY of the algorithm, which is a test that passes while the shipped code rots. + +⚠️ Truncation is reported through `onTruncated` rather than silently, with BOTH caps named: a silent depth cap hides a deep tree exactly as effectively as a silent node cap hides a wide one. + ## Local packages and build artifacts ### xterm-zerolag-input is single-source @@ -498,3 +995,35 @@ The repair is a **chmod, not a rebuild**: the prebuilt binary is fine. `scripts/ ### Headless screenshot capture **Headless screenshots: `deviceScaleFactor` MUST be 1, and write unique filenames** — `scripts/capture-real-overview.mjs` (drives a live session in headless Chromium → overview PNG). Two traps, both observed 2026-06-14: **(1) DSF=2 doubles the console font.** xterm's WebGL renderer draws terminal glyphs at ~2× their nominal size under `deviceScaleFactor: 2`, while STILL reporting nominal cell dims (`terminal.cols`/`_renderService.dimensions.css.cell` say 8px/187cols — they lie), so it's invisible to any internal measurement and only the pixels reveal it. The HTML chrome (header/toolbar) is unaffected → ONLY the console font looks comically large. Default to **DSF=1** (script does); the image is 1× res but the font is true-to-browser. **(2) Stable filenames → stale renders.** Overwriting a fixed path (`claude-overview.png`) in place leaves OS image viewers (eog/feh) — and any HTTP client behind a long/`immutable` cache — showing the OLD render; the user reads it as "the fix didn't work". The script now mints a timestamped `claude-overview-<ts>.png` per run. ⚠️ This was a LOCAL image-viewer cache, NOT a Codeman serving bug: `file-routes` previews send `Cache-Control: no-cache` and `/api/screenshots/:name` sends none. The one real Codeman-side footgun: `server.ts` serves non-content-hashed static assets `public, max-age=31536000, immutable`, and `cacheBustAssets()` only rewrites `.js`/`.css` refs — a stable-named **image** referenced from public/ would go stale on overwrite. Reflect the per-device UI to match a real device when capturing: seed `localStorage` `codeman:skin`, `codeman-font-size`, and the desktop `codeman-app-settings` blob (the plan-usage chip is a per-device display key deleted from the server payload — a fresh browser hides it unless seeded; close side panels for a full-width terminal). + +### Test isolation: tmux, docker and HOME + +**Tmux safety**: under vitest (`VITEST` env var, set automatically), `TmuxManager` no-ops ALL shell commands and becomes a pure in-memory mock — tests physically cannot create/kill/attach real tmux sessions (`IS_TEST_MODE` in `src/tmux-manager.ts`). Every docker IO path is no-op'd the same way. `Session` is test-gated too: instead of attaching a real tmux client, it spawns a raw-mode echo PTY (`TEST_PTY_SCRIPT` in `src/session.ts`), so integration tests get a live input/output loop that echoes each byte exactly once. + +`test/setup.ts` gives every test file a temporary `HOME`/`USERPROFILE` (all `homedir()`-derived state, `~/.codeman` and `~/codeman-cases` included, resolves into a per-file fixture; the Playwright browser cache path is preserved), and additionally strips `CODEMAN_PASSWORD`/`CODEMAN_USERNAME` (so auth state from the running instance can't leak into tests) and `CODEMAN_GESTURE` (a shell-exported gesture flag would flip render-injection assertions), and strips the three instance-selection vars `CODEMAN_INSTANCE`/`CODEMAN_DATA_DIR`/`CODEMAN_TMUX_SOCKET` (#356/#371; `test/test-env-isolation.test.ts` pins the list, and its STATIC half reads setup.ts so a dropped `delete` fails everywhere rather than only on a box that exports the var). + +⚠️ `CODEMAN_DATA_DIR` is the one that matters: `getDataDir()` reads it as an ABSOLUTE override before it ever looks at `homedir()`, so one inherited from the shell (a second instance, a beta run, a shell left over from `codeman web -d`) bypasses the temp HOME entirely, and a bare suite run once overwrote the real `remote-hosts.json` with a route test's fixture. `os.homedir()` itself DOES follow `$HOME`, so the temp HOME is what redirects everything else; `CODEMAN_INSTANCE` must be stripped in the setup file and never in a hook, because `config/instance.ts` captures it into a module-level const on first import. + +Tests that delete case trees go through `safeRmHomeTree()` (`test/mocks`), which refuses any path outside the temp HOME, so a wrong anchor leaves a temp dir behind instead of deleting `~/codeman-cases`. + +⚠️ Raw `npx vitest` without `--config` skips `setup.ts` and with it the temp-HOME isolation. + +### install.sh: the public installer + +**`install.sh`** (repo root) is the public entry point: `curl -fsSL <raw url> | bash` installs Node/tmux/git/build tools if missing, clones to `~/.codeman/app`, builds, and offers a systemd/launchd service. Since installer v2 (2026-09-20) it is **look, ask, work, done**: `preflight_detect` prints what is on the machine, every human step runs BEFORE the build (one consent for all missing packages, ONE sudo prompt kept warm by `sudo_session_start`, the AI CLI menu, the Tailscale login/operator/HTTPS-toggle preflight), then the clone/`npm install`/build/service/serve run unattended behind `run_step` spinners (output in `~/.codeman/install.log`, tail shown on failure), and `print_done_screen` ends on the URL with a terminal QR code (the `qrcode` package Codeman already ships). Design + verification record: `docs/installer-v2-plan.md`. + +The network-access question is 3-way: **Tailscale** (loopback bind + `tailscale serve`, curl-verified end-to-end), **LAN** (0.0.0.0 + password prompt), or **local-only**; it preserves the existing binding on re-runs via `read_existing_binding()`, which also reads back `CODEMAN_BASE_URL`/`CODEMAN_PORT` from the unit, and the flag/env PRESET paths keep an existing password too (`${CODEMAN_PASSWORD:-$EXISTING_PASSWORD}`: `--lan --service` on a password-protected unit used to rewrite it open with the unauthenticated ack, found in review 2026-09-21); `--password`/`--port` flip `RECONFIGURE` so they reach the unit instead of taking the quiet update path. + +⚠️ The env a hand-started `codeman web` needs (host, password, ack, base URL, port) is composed in ONE place, `start_command_hint` for the done screen's Start line and `export_bind_env` for main()'s `exec` branch, so the two cannot disagree; `stop_background_helpers` runs right before that `exec`, because exec skips the EXIT trap and the sudo keepalive (keyed on `$$`, which becomes the server's pid) would otherwise refresh the sudo timestamp for the server's whole life. Ctrl+C in the HTTPS-toggle poll is trapped for the poll only and skips Tailscale for the run rather than killing the installer. + +⚠️ **Tailscale is two halves on purpose**: `tailscale_prepare` (question phase: preflight, the opt-in rename, the serve SHAPE) and `tailscale_apply` (after the build: the one serve command). The shape is decided up front because it can change the service unit: when `:443` root already belongs to another app, the default is a **sub-path** (`tailscale serve --bg --set-path /codeman <port>` + `CODEMAN_BASE_URL=/codeman` in the unit; measured 2026-09-20: serve STRIPS the mount prefix before proxying, Codeman's `stripBasePath` tolerates unprefixed requests, and `--base-url` is what makes the emitted URLs carry it, verified live through the maintainer's tailnet incl. hashed assets and SSE), else a second port (8443+), replace, or skip. `detect_tailscale_serve_url` recognises all three shapes (`ts_serve_find_port_mapping`). + +⚠️ **Rename is opt-in and defaults to NO everywhere** (owner decision 2026-09-20: the tailnet name is the machine's SSH identity); `--name`/`install.sh name` do it, `--yes` and non-interactive never do. Serve config is keyed by the DNS name it was written under, so `tailscale_rename_node` takes OUR mapping down first (`tailscale_remove_our_mapping`, `TS_MAPPING_REMOVED_BY_RENAME`) and it is re-added under the new name; the previous name is recorded in `~/.codeman/tailscale-rename` so uninstall can offer it back. Tailscale state is detected dynamically from `tailscale serve status --json` (no marker files); the installer must NEVER `tailscale serve reset`, touch a mapping it did not create, run `tailscale funnel` or advertise a Tailscale Service (a Service needs a TAGGED node + admin approval, so it is a docs hint only); `test/install-sh-invariants.test.ts` pins all of those plus the rename-before-shape order and the flag/header parity. + +A foreign `/Library/LaunchDaemons/com.codeman.web.plist` (the Mac mini's headless setup) is now LEFT ALONE rather than replaced by a LaunchAgent. Subcommands: `update`, `uninstall`, `tailscale` (retrofit), `name [<n>]`, `status` (the done screen again), `cloudflared` (the tunnel client, no longer a question in the main flow); flags `--tailscale|--lan|--local`, `--name|--no-rename`, `--service|--run|--no-start`, `--yes`, `--password`, `--port` pipe through `bash -s --` and set the same variables as their env twins (`CODEMAN_NONINTERACTIVE=1` approves system changes for automation and never installs Tailscale, renames or starts a service; `CODEMAN_TAILSCALE=1` presets the Tailscale choice). + +Its CLI knowledge is a GENERATED block (`npm run generate:cli-catalog`, markers in the file), not a hand-written list: detection, the install menu and the closing reminder all read it, which is what stops the class of bug upstream `b6d0f1fa` fixed by hand (a user with only omp installed being told no AI CLI was found). + +⚠️ It must stay **bash 3.2** clean — macOS ships it and the documented install is `curl | bash` under `set -euo pipefail`, so `declare -A`, `mapfile`, namerefs, `${x,,}` and here-strings are all fatal there; CI runs `bash -n` plus a real `bash:3.2` container, since expanding an EMPTY array under `set -u` is a runtime abort `bash -n` cannot see. + +⚠️ It executes ONLY commands from the embedded block (`CLI_INSTALL_CMD_TRUSTED`) — there is no network fetch of the catalogue at install time to worry about at all.