# CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. > Deep implementation detail lives in [`docs/architecture-invariants.md`](docs/architecture-invariants.md). This file holds the rules that prevent mistakes; that file holds the mechanisms, file inventories, and the history behind each rule. Pointers below are written as `→ architecture-invariants#anchor`. When the goal is raw throughput, [`docs/SPEEDRUN.md`](docs/SPEEDRUN.md) is the fast-execution protocol (it removes ceremony, never the safety rules here). The user-facing manual is `docs/wiki/` (mirrored to the GitHub wiki by CI; see the CI note under Additional Commands), and `AGENTS.md` deliberately just points here. > > **This file is in `.prettierignore` on purpose.** Prettier's markdown printer escapes underscores inside the glob-heavy paths used throughout (`agent-*.jsonl` became `agent-\_.jsonl`, collapsing backtick spans and corrupting a whole paragraph). Do not remove the ignore entry, and do not run `prettier --write` on it. > > **Repo root is kept short on purpose** (the README sits below the file listing on GitHub). Config lives in `config/` (`eslint.config.js`, `knip.json`, the vitest configs), Prettier's config is the `"prettier"` key in `package.json`, and `SECURITY.md` is under `.github/`. Root-only files are the ones tools genuinely require there: `CLAUDE.md` + `AGENTS.md` (loaded from the root by Claude Code / Codex), `CHANGELOG.md` (changesets writes it next to `package.json`), `tsconfig.json`, `.editorconfig`, `.nvmrc`/`.npmrc`, `.prettierignore` (resolved relative to cwd), `LICENSE` (GitHub detection), `.dockerignore` (the build context is the repo root, so Docker resolves it there and nowhere else) and `install.sh` (its raw URL is the published install one-liner). `.claude-plugin/marketplace.json` is root-only for the same reason: `/plugin marketplace add Ark0N/Codeman` reads it from the repo root and nowhere else, which makes the repo its own plugin marketplace. The one plugin it lists is `plugins/codeman/` (manifest + README + a MIRROR of `skills/codeman/`), and ⚠️ the plugin is a small separate directory on purpose: `claude plugin install` copies the plugin root into its cache, and a plugin root that carries a `package.json` gets an **npm install** at install time (measured with the repo root as plugin root: 832 MB, 511 packages and this repo's postinstall on every installer's machine), while a symlink to `skills/codeman` would dangle in the copy. `skills/codeman/` stays the single source; `scripts/sync-plugin.mjs` mirrors it and writes `package.json`'s version into both manifests inside `version-packages`, and `test/plugin-manifest.test.ts` pins the byte-identity, the versions, the absence of a `package.json` in the plugin root and that no other component (`commands/`, `agents/`, `hooks/`, `.mcp.json`, `settings.json`) rides along. `npm run check:plugin` runs that drift check plus both strict validations with EXPLICIT paths (`plugins/codeman`, `.claude-plugin/marketplace.json`): a bare `.` argument copied out of prose reads as a full stop and gets dropped, which surfaces as `missing required argument 'path'`. Needs the `claude` CLI, so it is a local check, not a CI step. Don't relocate those. ## Quick Reference | Task | Command | | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Dev server | `npm run dev` (or `npx tsx src/index.ts web`) | | Type check | `npm run typecheck` (= `tsc --noEmit`) | | Lint | `npm run lint` (fix: `npm run lint:fix`) | | Format | `npm run format` (check: `npm run format:check`) | | Tests | `npm test` (the CI gate — safe to run bare) · one file: `npm test -- test/.test.ts` · see Testing for the excluded suites | | Build | `npm run build` (esbuild via `scripts/build.mjs`, NOT tsc — `tsc --noEmit` is type-check only) | | Production | `npm run build && systemctl --user restart codeman-web` | ## CRITICAL: Session Safety **You may be running inside a Codeman-managed tmux session.** Before killing ANY tmux or Claude process: 1. Check: `echo $CODEMAN_MUX` - if `1`, you're in a managed session 2. **NEVER** run `tmux kill-session`, `pkill tmux`, or `pkill claude` without confirming 3. Use the web UI or `./scripts/tmux-manager.sh` instead of direct kill commands **The working tree is shared with other agent sessions.** Several Codeman sessions run against THIS one checkout, so another session can `git checkout` a different branch, or leave half-finished untracked files, while you are mid-task. - **Always `git branch --show-current` immediately before committing.** Observed 2026-07-27: another session ran `git checkout -b feat/web-tabs`, a commit silently landed there instead of master, and the follow-up `git push origin master` cheerfully reported "Everything up-to-date". - To land a commit on master **without** switching branches (which would yank the tree out from under the other session): `git push origin HEAD:master` then `git branch -f master HEAD`. Never `git checkout master` to "fix" it. - **Never `git add -A`/`git add .`** — stage explicit paths. A sweep will pick up another session's WIP. - Another session's broken WIP can block `npm run build`, since `tsc` is the first step and the build gates on it. That is not your bug to fix. ⚠️ `tsc` still EMITS on type errors, so a failed `npm run build` leaves a rebuilt `dist/index.js` compiled from their tree; check what it pulled in before restarting the service. To deploy frontend-only changes past a blocked `tsc`, run the asset stage of `scripts/build.mjs` (everything after the `tsc`/`chmod` lines is independent of it). ## CRITICAL: Always Test Before Deploying **NEVER COM without verifying your changes actually work.** For every fix: 1. **Backend changes**: Hit the API endpoint with `curl` and verify the response 2. **Frontend changes**: Use Playwright to load the page and assert the UI renders correctly. Use `waitUntil: 'domcontentloaded'` (not `networkidle` — SSE keeps the connection open). Wait 3-4s for polling/async data to populate, then check element visibility, text content, and CSS values 3. **Only after verification passes**, proceed with COM The production server caches static files for 1 year, `immutable` (`maxAge: '1y'` in `server.ts`). To avoid stale frontend after a deploy, `renderIndexHtml` runs `cacheBustAssets(html)` — it appends `?v=` to **every same-origin `.js`/`.css`** reference (mtime memoized ~1s so a burst of renders is cheap; external/already-versioned/missing refs untouched). Because `index.html` is served `no-cache`, a **normal reload now picks up edited modules/styles — no hard refresh needed** (the gesture bundle is injected separately with its own `?v=`). If you add an asset referenced by an _absolute_ URL or from JS rather than a `` via the exported `escapeScriptJson()` since `label` is a user-`clis.json`-settable string unlike the neighbouring booleans-only `__codemanCliAvailable`), never a hardcoded per-CLI id list in the frontend — the same "no branching on CLI id outside stock.ts" discipline the registry itself enforces. One entry per (capable, INSTALLED CLI, saved endpoint) pair, e.g. "Claude Code (llama.cpp)", filtered through `isCliAvailable()` like the stock entries. Clicking one calls `selectCustomModelEntry(mode, endpointId)` (`session-ui.js`), which re-fetches the endpoint (never trusts anything cached from the dropdown's render — the 5-minute sweep below or a settings edit may have changed it since) and decides the model: exactly one discovered model launches straight away, two or more open `#customModelPickModal` to ask, with `defaultModelId` marked but never auto-chosen (asking exists so ONE launch can deliberately differ from the saved default). ⚠️ **The picker promotes exactly one row to the top**: whichever model llama-swap reports `ready` right now (tagged "Currently loaded", queried via `GET /api/model-endpoints/:id/running-status`, client-side bounded to ~800ms via `Promise.race` so a sleeping/firewalled endpoint cannot leave the modal invisible for the route's own 5s server-side timeout) beats a merely remembered choice, and — only when nothing is currently loaded — the model actually launched last for this exact (harness, endpoint) pair (tagged "Last used", read from the per-device `codeman:customModelLastUsed::` localStorage key). Neither tag reorders past the top, and the "Default" pill is a SEPARATE span rather than a third value of the same slot, so a promoted row that is also the endpoint's `defaultModelId` shows both (on a single-purpose GPU box that is the common case; an exclusive slot silently dropped the Default marking for exactly that row). "Last used" is written by `_runCustomModelEntryViaRestart` (claude) and `_quickStartWithCustomModelConfirm` (every one-shot launch; `runCustomModelEntry` itself only dispatches between the two) only once the model is actually applied — never on the mere click — because a context-window-warning decline means this exact model cannot work with this CLI at all, and promoting a model that cannot launch would be actively wrong, not just premature. Either way the actual launch (`runCustomModelEntry`) routes through `run()` itself via a temporary `_runMode` swap — never `setRunMode()`, which would persist it as the user's new default — rather than a parallel dispatch table, which is what gives a custom-model launch the same `_runInFlight` lock every other Run click gets and means a CLI whose injection recipe lands later needs no update here. It then GETs `/api/sessions/:id/wait?until=idle&timeout=20000` on that session BEFORE applying — measured live, a freshly launched CLI reports itself `busy` for its own startup (boot spinner, workspace-trust check) well before the apply call would otherwise reach it, and the apply route's `isBusy()` guard correctly can't tell that apart from a real turn in progress, so every fresh launch failed with `SESSION_BUSY` until this wait was added. A timeout there is a normal 200 per the wait endpoint's own contract, never an error, so a session still busy after 20s just reaches the apply call anyway and gets that route's own honest error. It then calls `POST /api/sessions/:id/custom-model` on the session `run()` produced, guarded by snapshotting `activeSessionId` before the call and requiring it to have actually changed after — every `run*()` handles its own failure internally and returns normally rather than throwing, so a declined/failed launch must not silently re-point and restart whatever session was already open. ⚠️ The apply call reads the response body itself (`_api()`) rather than `_apiJson()`, which unwraps success but silently discards a failure body — losing the one thing (`error`) that distinguishes "still busy", "remote/Docker session" and everything else the route can report (⚠️ neither route validates `modelId` against the endpoint's discovered list, deliberately: discovery can be up to 5 minutes stale, so a 400 there would refuse a launch that works — a typo'd id fails on the CLI's own first request instead); the resulting toast is `type: 'error'` with an explicit `duration: 0` (no auto-dismiss, an explicit close button) at that one call site — not a blanket sticky-error default, which stacked unbounded on `.toast-container` with no cap or eviction — precisely so a message worth diagnosing survives long enough to be read instead of vanishing on the usual 3s timer. Entries are hidden for a remote/docker active case (the apply route refuses both) and for an endpoint with no discovered models at all (nothing to launch with). ⚠️ **Every saved endpoint's models also re-discover themselves automatically**, a `this.cleanup.setInterval` in `server.ts` (`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, 5 minutes, off under `testMode` like the Codex plan-usage poll beside it) calling the exported `refreshAllCustomModelHosts()` (`custom-model-routes.ts`) — one endpoint unreachable on a cycle never blocks the others, and a read-modify-write PER HOST (re-reading the store before each splice, keyed by id) means an admin's concurrent edit or delete wins over a sweep that started before it, never the reverse. **Everything below landed after the initial backend + picker cut, each confirmed live against a real llama-swap deployment.** ⚠️ **llama-swap runs one model at a time, and switching can disrupt ANOTHER live session** — before applying, both apply routes call llama-swap's own `GET /running` (feature-detected via `getLlamaSwapStatus()`, `custom-model-routes.ts`; a plain llama.cpp/OpenAI-compatible server has no such endpoint and is simply never checked). If a different model is loaded and ready AND another live session's own selection is using it, the apply returns `{requiresConfirmation, currentlyLoadedModel, affectedSessions}` instead of switching silently; retrying with `confirmedSwap: true` skips the check, and switching with nothing else affected proceeds immediately. ⚠️ **The swap question and the context-floor question below have SEPARATE flags** (`confirmedSwap`, `confirmedContext`), because the context check runs first and while both shared one `confirmed` a user who clicked past a too-small context silently consented to evicting another session's model too; the legacy `confirmed` still means both, since it shipped in the HTTP-API-only cut. llama-swap also has no dedicated "switch model" endpoint — the only thing that actually starts a swap is a real inference request naming the model (confirmed live: applying a selection alone never reached llama-swap's own logs, since nothing had asked it to load anything) — so both routes also fire `triggerLlamaSwapLoad()`, a fire-and-forget `POST /v1/chat/completions` with `max_tokens: 1`, whenever the target model isn't already loaded and ready. ⚠️ **That launch-time check cannot catch a swap caused by a DIFFERENT session's LATER, ordinary use** — confirmed live: a second Codex session picking a different model launched with no warning at all (nothing conflicted at that exact instant), yet it silently evicted the first session's model regardless, since llama-swap has no push notification of its own. `detectCustomModelSwapDisplacements()` (`custom-model-routes.ts`) is a separate periodic sweep (`server.ts`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS` = 20s) that compares each live custom-model session's own `modelId` against what `/running` actually reports loaded, broadcasting a `custom-model:swapped-out` SSE event — shown as a global toast, never tied to the displaced session's own tab, since the whole point is telling the user before they type into it — the first time a mismatch appears, via a caller-owned de-dupe `Set` cleared once that session's own model is loaded and ready again so a later, genuinely new displacement notifies again rather than staying silently un-notified forever after the first one. ⚠️ **Context length is read from the REAL launch command, never `/props`** — `/props?model=`'s `default_generation_settings.n_ctx` was confirmed live to report a `--fit-ctx`-launched backend's theoretical/trained maximum rather than the real runtime-configured size (a measured 154112-vs-16384 discrepancy, caught only because the unfixed value still overflowed), so discovery parses the actual configured size straight out of `/running`'s own `cmd` field instead (`parseCtxFromCmd`: `--fit-ctx ` first, then plain llama.cpp `-c`/`--ctx-size`), falling back to `/props` only when `cmd` states no recognizable flag at all. ⚠️ **Claude alone gets a context-window FLOOR check, on top of the ceiling `contextLengthVar` already fixes** — `exceedsSafeContextFloor()` (gated on the registry declaring `contextLengthVar`, so a no-op for every other CLI by construction) compares a model's discovered context against `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000): confirmed live, twice, that Claude Code's own system prompt and tool schemas cost roughly 36.4K tokens on the very first message, before any conversation history exists to compact, so a smaller real context fails outright regardless of what `CLAUDE_CODE_MAX_CONTEXT_TOKENS` says (that var only controls when HISTORY gets compacted, and there is none yet on message one). Below the floor, the apply returns `{requiresContextWarning, modelId, contextLength, minSafeContextTokens}` instead of launching, shown as an in-app dialog naming the actual fix: give the model an explicit larger `-c`/`--ctx-size` in llama-swap's config instead of relying on `--fit-ctx` auto-fit, which optimizes for the biggest MODEL that fits rather than the biggest CONTEXT. ⚠️ **A fresh, isolated `CLAUDE_CONFIG_DIR` looks like a brand-new Claude Code profile and replays its ENTIRE first-run sequence on every launch** — the theme picker, the security-notes screen, the per-project "trust this folder?" dialog, and (running with a bypass-permissions flag) a one-time warning about it, confirmed live, none of which a real, already-onboarded profile shows again. `skipFirstRunPrompts` (claude's entry only, requires `apiKeyTrustFile` since it reuses the same file) pre-seeds that same "already been through this" state: `hasCompletedOnboarding` and this session's own `projects[workingDir].hasTrustDialogAccepted` merge into the same `.claude.json` the API-key trust file already writes to, and `skipDangerousModePermissionPrompt` merges into `settings.json` (a different file, same corrupt-tolerant merge). ⚠️ **The loading banner shows the REAL backend log line, not a guess, and has no countdown or auto-timeout at all.** `getLatestLlamaSwapLogLine()` holds one `GET /api/events` SSE connection open per endpoint (confirmed live to stay open indefinitely — read past 220KB over 8s with no `done`; idle-closed after 30s via `pruneIdleLlamaSwapLogTails`, same 20s sweep as the swap-displacement check above), parsing `logData` frames and keeping only `source: "upstream"` (the real `llama-server` process's own stdout) lines, never `source: "proxy"` (llama-swap's own request-access log). ⚠️ `GET /logs` — the endpoint this feature's own first cut targeted, since the name suggested it — was confirmed live to carry ONLY the proxy log and never a single backend line, even seconds after a real, verified model swap; caught and corrected by a live check before merge, not after. The banner itself dropped its size-scaled expected-time estimate and matching auto-timeout (a guess dressed up as a fact that could kill a genuinely slow load partway through on slower hardware) for a generic hardware/model-size disclaimer plus a user-driven **Cancel** button (`_showCenterStatus`'s `onCancel` option, a real button distinct from the plain "×" close glyph an `'error'`-type banner gets) that ends the wait and closes the session on the user's own call rather than a guessed deadline. **Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w-` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. ⚠️ **Closing has the mirror-image race and one owner**: `closeSession()` reads `wasActive` BEFORE its `await` and announces the delete via `_closingSessions`, while `_onSessionDeleted` skips the active-session handoff for an id in that set. Both used to read `activeSessionId` after the fact, so the `session_deleted` broadcast for your own delete could null it first and closing the tab you were on landed on the welcome screen instead of the next session, on the same build, depending on timing. The fallback also picks the first order entry that is still in `sessions` (a dead id can linger in `sessionOrder`, same reason Alt+N indexes a live-filtered list). A delete from ANOTHER client still shows the welcome screen, which is the honest answer when what you were looking at was taken away. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization) **Session lineage lines** (tab → tab it spawned, `sessionLineageLines`, per-device, desktop default ON): a create request may name the session that spawned it, as a `parentSessionId` body field on `POST /api/sessions` / `POST /api/quick-start` or the `X-Codeman-Parent-Session` header (the agent skill sets that once on its shared curl invocation, so every spawn recipe carries it). `resolveParentSessionId()` (route-helpers.ts) **resolves rather than trusts** it: exact id, else a UNIQUE ≥8-char prefix (ids reach agents truncated), it must be a live session the caller can see AND carry the same owner, and **anything unresolvable is DROPPED, never a 400** — a cosmetic field must not be able to fail a worker spawn. It rides `toState()` into `session_created`, so there is no new SSE event. ⚠️ Rendering is an ADDITIONAL LAYER on the existing SVG pass (`_appendLineageConnectionLines` called at the tail of `_updateConnectionLinesImmediate()`, exactly like ultracode), sharing one batched read→write reflow and the `tab:` rect cache; geometry is pure in `computeLineagePath()` (constants.js). ⚠️ **ONE shape, and the second one was the bug**: every pair (flat strip or wrapped) gets a U-bridge hanging below the strip, anchored on both tabs' BOTTOM edges. A wrapped strip used to get a parent-bottom → child-TOP bezier with a ~14px row gap to bend in, which drew a flat line hidden in the gap with siblings overprinting. ⚠️ The dip is a **mis-tuned-in-both-directions corridor** (44px cap = straight thread at strip-wide spans, #285; 104px cap + full row offset = ~106px over-bow into the terminal, 2026-08-15): it now hangs from the **STRIP's bottom edge** (fallback: lower tab bottom), capped at 64px, with NO per-row offsets stacked on top — the strip-bottom baseline is also what keeps a row-1 pair's arc from drawing through row 2's tab labels. ⚠️ **Colors are keyed on the SPAWNING tab, not per child**: every arc leaving one tab is the same color however many workers it spawns, so the strip reads as "these five came from w1, those two came from w2" — per-child coloring gave one tab's own children a different color each, which is the distinction the colors exist to make. A child that spawns in turn is a parent in its own right and gets its own color for the arcs below it, so a chain changes color at each generation while each generation's fan-out stays uniform. Assignment cycles `CodemanLineage.COLORS` in first-seen order per parent id (first entry empty = the skin-tuned `--session-blue`, so the first spawning tab keeps it; the rest vivid fixed hexes), memoized rather than derived from draw index (the SVG is wiped and rebuilt constantly, so an index-based color would flicker), and set inline as `--lineage-color` so styles.css keeps owning opacity/glow/dash. `test/session-lineage-lines.test.ts` drives the real `_appendLineageConnectionLines()` and asserts the painted property, since testing the color function alone would pass just as happily with the child id passed back in. ⚠️ **Desktop only**: the overlay is `z-index: 999` and the desktop header is 100 (arcs paint over it, which is what lets them touch tab bottoms), but under 1024px mobile.css makes the header `fixed; z-index: 1200` and would bury them. ⚠️ Paths carry `data-agent-id="lineage:"` because that is what `_applyLineEntrances()` queries — that one attribute is what gives them the entrance animation and its negative-`animation-delay` resume across `svg.innerHTML=''`. ⚠️ `.session-tabs` is `overflow-x: auto`, so a scrolled-out tab still HAS a rect (over the logo); edges with an endpoint outside the strip are skipped, and a passive `scroll` listener re-anchors the rest. **Auto-named sessions** (`autoNameSessions`, SYNCED, default OFF; #376): a placeholder tab (`w3-myapp`) takes its first real prompt as a title, in the `: ` form (`w3-myapp: fix the login redirect`) that `parseSessionPrefix()` (app.js, #232) already renders as the title alone with the prefix in the tooltip and that `_nextCaseSessionStartNumber()` still counts, so the case identity and the `w<n>` counter survive. Ownership is the tri-state `SessionState.nameSource`: `placeholder` (Codeman's own `w<n>-<case>` or no name, inferred by `isGeneratedSessionName()` when a persisted state predates the field), `auto` (titled once), `manual` (the `name` setter, i.e. `PUT /api/sessions/:id/name`, which auto-naming never touches again). ⚠️ **First prompt means the FIRST**: `applyAutoName()` flips a placeholder to `auto` whether or not the string changed, so a later "1" cannot rename the tab; a prompt that yields no title (`/clear`, a `!` shell escape) leaves the session eligible for the next one. ⚠️ **Only user-originated input counts** (`SessionWriteOptions.fromUser`, set by the browser WS path and `POST /api/sessions/:id/input` ONLY, so a forgotten flag on a new path fails toward not naming): Ralph kick-starts, respawn `/clear`s, cron launches, approval answers and the trust-dialog keys write through the same `write()`/`writeViaMux()` and used to name every Ralph tab "Read @ralph_prompt.md…". A `startMode: 'shell'` CLI never feeds the tracker (a capability, not an id check; a shell tab was renamed after every `ls`), and the send-key route's Shift+Enter line feed bypasses the session entirely, so it calls `trackUserInput()` or the two lines join with no separator. ⚠️ **The tracker (`session-auto-name.ts`, pure) sits on the raw keystroke stream**, so every key has an explicit rule: a bare Esc is resolved at the END of the chunk it arrives in (it used to stay in escape mode and eat the next prompt's first character, or a whole CJK prompt); SGR mouse reports, Tab, cursor keys and Shift+Tab leave the draft alone (a wheel tick mid-word used to drop the first half); Up/Down and Ctrl+P/N/R TAINT the draft so Enter submits nothing rather than a fragment; bracketed-paste newlines are newlines IN the composer, never Enter. The title is the first sentence past a minimum length ("e.g. fix this now" is not "e.g."), capped at 72 code points, and the composed name honours `MAX_SESSION_NAME_LENGTH`. The listener (`session-listener-wiring.ts`) checks eligibility BEFORE reading the setting, so an already-named session costs no settings read per prompt. Opt-in because the prompt lands in the tab name, `mux-sessions.json`, every `session:updated` and `/api/search` (Read My Mind keeps prompts 0600 for the same reason). Tests: `test/session-auto-name.test.ts`, `test/session-listener-wiring.test.ts`, `test/routes/session-name-routes.test.ts`. → [architecture-invariants#auto-named-sessions-first-prompt--tab-title](docs/architecture-invariants.md#auto-named-sessions-first-prompt--tab-title) **Maintainer bot (external)**: the Telegram bot that reviews open PRs and triages discussion threads in Codeman sessions used to live at `scripts/pr-bot/`. It moved OUT of this repository on 2026-09-14, to `~/codeman-cases/prbot/` (its own private git repo, systemd unit `codeman-pr-bot`, guide + agent rules in its own `README.md` and `CLAUDE.md`). It is a CLIENT of Codeman's HTTP API like any other, so nothing here depends on it and it is not part of the server, the CLI or the npm package. ⚠️ It spawns real sessions named `prbot-<n>` / `dscbot-<n>` on the local Codeman and holds clones under `~/.codeman/pr-bot/`, so those session names and that data dir are taken; it also fetches PR heads into `refs/pr-bot/*` of this checkout and must never check out, reset or clean it. The CHANGELOG entries for 1.25.0 and earlier still describe it, which is history rather than drift. **Unified session list**: `GET /api/sessions/unified` merges live sessions, persisted state, lifecycle-log history, and transcript files into one deduped list (pure core in `src/services/unified-session-service.ts`). ⚠️ **Transcript history is THREE stores, not one**, because each CLI keeps its conversations in its own: Claude's `~/.claude/projects`, omp's `~/.omp/agent/sessions` and codex's `~/.codex/sessions` (#386). Rows fold into their owning session via the `claudeSessionId → Codeman id` alias map, so resumed and `/clear`-respawned sessions do not appear twice; that field is named for Claude and carries whatever id the CLI names its conversation with, which for every non-Claude row diverges from the Codeman id by construction. ⚠️ **`resumeId` is set by a SCANNER row only, never by a live session**, and that is what makes it safe to resume on: a row carrying one is a conversation already on disk, so `resumeHistorySession()` sends `codexConfig.resumeSessionId` and a row without one is a genuinely fresh session. Every surface that re-projects these rows has to carry the field through, the phone overview included, or a tap on that surface silently starts a second conversation. No terminal buffers in the response, unlike `/api/sessions`. Backs the Cmd+K Session Manager, plus pinning and cross-device tab order (`PUT /api/session-order`; pure merge helpers in `src/session-order.ts`, pushing device wins and server-only ids are never dropped). → [architecture-invariants#unified-session-list-and-session-manager](docs/architecture-invariants.md#unified-session-list-and-session-manager) **Owner tab layouts** (COD-359, `tab-layout*.ts` + `GET`/`PUT /api/tab-layout`): named tab GROUPS over the flat tab strip, scoped per owner (`SINGLE_USER_LAYOUT_OWNER` = `@single` when multi-user is off), persisted under the `tabLayouts` key in state.json. A layout is `{version, groups[], ungrouped[], updatedAt}` whose refs point at either a session or a saved webview (`TabRefKind`), capped at 32 groups / 512 refs. ⚠️ **BACKEND ONLY as of 1.24.1**: nothing in `src/web/public/` calls these routes yet, so a UI built on top is new frontend work, not a rewiring job. ⚠️ **`TabLayoutService` is the single mutation boundary** and every lifecycle caller (session created/removed, webview created/deleted, a legacy order PUT) describes ONE completed server action and gets AT MOST ONE versioned write; writing layout state from a route or a manager directly is what the service exists to prevent. ⚠️ The layout does not replace `PUT /api/session-order`, it PROJECTS onto it: `tab-layout-legacy-order.ts` is the pure translation both ways (`putLegacyOrder()` recomposes a global order from the owner's groups), so changing one side without the other silently desyncs the tab strip from the stored layout. ⚠️ **Reconciliation is gated on a SUCCESSFUL restore** (`markRestorationComplete` / `markRestorationFailed` / `markRestorationSkipped`, plus `assertDeletionReady()`): pruning refs against a session list that failed to load would delete live tabs, so a failed restore must leave the layout untouched. `PUT` takes exactly `{baseVersion, layout}` (any other key shape is a validation error), answers a stale `baseVersion` with the current layout rather than clobbering, and is capped at 128 KiB. Broadcasts `tab:layoutChanged`, owner-routed via `deriveTabLayoutSseHint`. **Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `elicitation_complete`, `elicitation_response`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`, `prompt_submitted` (UserPromptSubmit, #367: a Claude pane reports its live conversation id first-hand). See `src/hooks-config.ts`; upstream hook semantics mirrored in `docs/claude-code-hooks-reference.md`. ⚠️ **Every claude session INSTALLS the hooks block into its workspace** (`applyWorkspaceHooks` in hooks-config.ts → `ensureCodemanHooks`, an add-only merge that keeps a user's own handlers), from EVERY claude create path — both interactive routes, cron fires, legacy scheduled runs, the plan-orchestrator one-shots — and from `restoreMuxSessions()` for sessions recovered on server start (that boot sweep skips a workspace that no longer exists, so a deleted repo with a surviving tmux session is never resurrected as an empty dir). Before 2026-08-15 hooks were written ONLY when Codeman created the case DIRECTORY, so a linked case / cloned repo — where most sessions actually run — had no hooks at all and every hook-driven surface was silently dead there: an AskUserQuestion dialog blocked the pane while the tab and the phone overview both read a calm `idle`, with no Approvals Inbox item, no push, no definitive `stop`/`idle_prompt` for respawn and no `stop`/`blocked` for the wait endpoints. The escape hatch is the synced `workspaceHooksEnabled` setting (App Settings → Agents & CLIs → Claude, **default ON**); OFF restores the old behavior, where a Codeman block that is already there is still refreshed when stale (COD-91) but one is never added. ⚠️ Route the decision through `applyWorkspaceHooks` rather than calling `ensureCodemanHooks` at a new site, or the setting silently stops applying to that path. ⚠️ Claude Code RE-READS `settings.local.json`, so an already-running session starts firing hooks without a restart (measured 2026-08-15) — and the notification for a blocking dialog is delayed by Claude Code (~30s), so the alert trails the dialog. ⚠️ An AskUserQuestion / plan-selection dialog arrives as **`permission_prompt`**, not `elicitation_dialog` (that one is MCP elicitation), so it renders as the RED "needs you" alert, not the yellow idle one. **Reboot restore** (#411/#442, `src/reboot-restore.ts` pure + `web/reboot-restore-registry.ts` + `routes/reboot-restore-routes.ts` + `reboot-restore-ui.js`): a host reboot takes the tmux server with it, so every pane dies and the board comes up empty with no explanation. At boot Codeman works out which sessions that reboot destroyed, holds the plan IN MEMORY (no new state file, and a server restart simply drops the offer), and the banner asks. ⚠️ **The heuristic decides whether to ASK, never whether to act**: two signals have to agree (the socket holds no panes at all while state still lists sessions, AND the host booted after the newest persisted activity), and a wrong yes costs one dismissable line rather than N CLI processes nobody asked for. ⚠️ Rebuilding is TAKE-then-build: entries leave the plan synchronously before the first `await` and the route is single-flighted per owner, so a double-click or a second device cannot put two panes on one conversation. Anything that never became a pane goes BACK on offer, with one deliberate exception, `already-live`, which unlike a missing workspace or a withdrawn grant cannot stop being true. ⚠️ Three things are re-checked at click time rather than trusted from boot (the owner's privilege grant, the workspace still being on disk, and the conversation not already being live), and the already-live sets are read FRESH per iteration rather than snapshotted: the loop awaits a real `startInteractive()` per entry, so a snapshot taken before it is tens of seconds stale by the tenth entry and would miss a conversation the user resumed by hand in that window. The confinement re-check is keyed on the entry's OWNER, never the caller, or an admin spending another user's entry is waved through by `isWorkingDirAllowed`. ⚠️ A rebuilt session comes back **attached, idle and disarmed**: the pane is NEW, so terminal scrollback is gone (the banner says so) while the conversation continues, respawn controllers and Ralph loops are never re-armed, and `reapplyPersistedSessionState(..., { rearmAutoResumeSchedule: false })` keeps auto-resume ENABLED but drops the pre-reboot `autoResumeAt` stamp, or one click has every restored session type `continue` into itself a minute later, unattended. That option exists only for this path; a Codeman restart still re-arms, because the limit footer will not reprint on its own. ⚠️ The rebuild passes `nameSource` through, or the constructor re-infers it from the name and a hand-renamed session shaped like `w<n>-<case>` comes back as `placeholder` for auto-naming to overwrite. ⚠️ Claude-mode only (others carry their conversation id in their own config object), and remote/docker sessions are never offered (`remote-or-docker`), because both need another host or container to be up. ⚠️ A failed rebuild is undone with `discardPartiallyBuiltSession()`, deliberately NOT `cleanupSession()`: the delete path would count the session's tokens into the lifetime totals, demote a pinned record to `stopped` (which this pass reads as an intentional kill, making the session permanently unrestorable) and recursively remove the WORKSPACE's `.claude-images`. ⚠️ `os.uptime()` reports the HOST's uptime, which a container shares, and that cuts both ways: after a genuine host reboot a containerized Codeman does see a short uptime and the banner works, but a container-only restart is invisible to it, which is the case where this would help most. Tests: `test/reboot-restore.test.ts`, `test/routes/reboot-restore-routes.test.ts`, `test/routes/reboot-restore-rebuild-failure.test.ts`, `test/discard-partially-built-session.test.ts`. **Approvals Inbox** (cross-session queue of prompts waiting on a human; `approvalsInboxEnabled`, SYNCED, default OFF: every surface is opt-in; only the store and answer endpoints run regardless, so flipping it ON shows anything already pending): `web/approval-inbox.ts` is a `sessionWaits`-style singleton fed by `/api/hook-event`, holding at most ONE item per session (a new prompt supersedes), claude-mode only, in-memory. Cards are answered via `POST /api/approvals/:id/answer`, which sends a digit / Esc / idle-prompt text through `writeViaMux` (menu answers never carry `\r`). ⚠️ `option` digits are accepted ONLY when they match options parsed from the captured pane frame, and the answer path RE-CAPTURES the pane first (a dialog that no longer parses on screen means the keystroke would land in the composer, so refuse with 409). ⚠️ Resolution on the heuristic `working` signal ALONE is restricted to `idle` items; a permission/question item gets the pane-VERIFIED variant on that same signal (`resolveIfDialogGone()` → `verifyStillAnswerable()`), so the heuristic only decides when to LOOK and the screen decides the outcome. That is what clears a dialog answered in the terminal mid-turn; the other definitive signals are `stop`, `elicitation_complete`/`elicitation_response`, exit/delete, answer, supersede and the 12h TTL. ⚠️ **Viewing a session ACKNOWLEDGES its idle item, it does not resolve it** (`POST /api/approvals/session/:sessionId/viewed` → `acknowledgedAt` → `approval:updated`): the item stays pending (still answerable, still Read My Mind context) and only stops arming the yellow tab alert. That flag is what makes the clear durable, since the view-clears-idle rule used to live in one browser's memory and `seedApprovals()` re-armed the alert on the next reload while other devices never heard about it at all; the local half is `markIdleAlertSeen()` (app.js), called from BOTH `selectSession` paths, including the already-active early return, where a click could otherwise never clear the alert. ⚠️ **Only a HUMAN opening a session acknowledges**: `selectSession(id, { auto: true })` marks the three selections the APP makes (boot restore, a solo window opening its target, the fallback after the active session is closed) and skips the acknowledgement, so a page load cannot silently spend an alert the user never saw. The flag defaults to user-initiated, so an untagged call site fails toward acknowledging rather than toward an alert nothing can clear; `test/session-select-ack-gate.test.ts` pins both the gate and the tagged call sites. Idle-only by construction (`acknowledge()` defaults to `['idle']`): looking at a permission/question dialog does not answer it. ⚠️ Same rule on the input path: `_ackDelivery` (app.js) spends the IDLE alert only, via that same `markIdleAlertSeen()`. It used to `clearPendingHooks(sessionId)` with no kind, so one keystroke wiped a RED alert on that device while the dialog was still up, the other devices stayed red, and a reload re-seeded it. ⚠️ Claude Code fires no "permission answered" hook (only `elicitation_complete`/`elicitation_response`, i.e. the question flavor), so an answered-in-the-terminal dialog would otherwise sit pending until `stop`: `GET /api/approvals` therefore runs a **staleness sweep** over the caller's own items via `verifyStillAnswerable()`, which is deliberately the conservative check the answer path uses (only an item whose ORIGINAL frame parsed options can be dropped, so an unreadable capture keeps the alert rather than losing a live one). ⚠️ **`applyCapture()` is therefore ADD-ONLY for `options`**: a re-capture that parses nothing must never erase a parse an earlier one found. Claude Code delays the Notification hook behind the dialog (measured 6s, documented ~30s), so the 600ms re-capture routinely lands on a frame the user has ALREADY answered; clearing the field there made the item permanently unsweepable, because `verifyStillAnswerable()` reads a MISSING `options` as "we never could read this dialog" and keeps such items answerable by design. The red "needs you" then survived every sweep AND every page reload, went away only on `stop` (2026-08-20: a confirmed question left a tab flowing red for ~8 minutes while the turn ran on), and the stale card still accepted an answer, typing a bare `1` into a composer with no dialog under it. Pinned by `test/approval-inbox.test.ts`. ⚠️ A frame that parses no options is CONCLUSIVE in exactly two cases, and the second one closes the late-hook hole: the item once parsed options (they cannot vanish while the dialog is up), or the frame shows Claude actively running a turn. A modal dialog BLOCKS the turn, so the two cannot coexist — measured on v2.1.237, a live-dialog frame carries neither the `… (13s` timer NOR the `esc to interrupt` footer, which the dialog replaces with `Enter to select · ↑/↓ to navigate · Esc to cancel`. Anything else stays answerable, so an unreadable capture still keeps the alert. That second signal is reached by a delayed staleness pass (`STALE_CHECK_DELAY_MS`, 3s) scheduled alongside the re-capture, because a prompt answered BEFORE the hook lands creates an item whose FIRST capture already has no dialog in it: nothing ever parsed, `stop` may have fired already, and the alert then outlived reloads until the 12h TTL. ⚠️ That pass must stay comfortably LATER than `RECAPTURE_DELAY_MS`, whose whole reason for existing is that the hook can beat Ink to the screen — resolving inside the paint window would clear the alert for a dialog that was about to appear. The frontend seeds from `GET /api/approvals` in `handleInit` **regardless of the setting**: the seed re-arms the tab-alert state machine (`setPendingHook`) unconditionally, and only populating `this.approvals` (the inbox surfaces) is gated — seeding used to be gated wholesale, which left a reloaded page with NO red tab while a permission dialog sat blocking a session (2026-08-15); `_onApprovalResolved` clears the pending-hook alert unconditionally for the same reason. ⚠️ The red/yellow tab alert itself is a STEADY border/background/dot with a pulse on top: the original keyframes swung to transparent at 0%/100%, so half of every cycle looked like a normal tab. Push Approve/Deny buttons stay gated on the setting (`sendPushNotifications` strips `actions`/`approvalId` when OFF) and are answered from `sw.js` directly so they work with no tab open. Surfaces (all gated on the setting): header bell (marker-hidden until count > 0, phones never show it) + drawer (`approvals-ui.js`), phone overview NEEDS YOU answer strips (`mobile-overview.js`). Design: `docs/approvals-inbox-plan.md`. **Read My Mind intent profiles** (phase 1 of `docs/readmymind-plan.md`; `readMyMindEnabled`, SYNCED, default OFF): per-CASE profiles (user-stated `goals` + the user's recent real prompts), keyed by owner + realpath(workingDir) so they survive `/clear`/respawns and multi-user scoping is structural. Capture rides the transcript (`transcript:user_prompt` from `transcript-watcher.ts`), NOT the input paths: `POST /input` sees only programmatic prompts and the WS channel is raw keystrokes. The listener lives inside `startTranscriptWatcher()`'s `if (!watcher)` block (outside it would duplicate per hook event) and is claude-only + gated on the setting per event. Store: `src/intent-store.ts` singleton, `intents.json` written 0600 tmp+rename (prompts can contain secrets; never fed to `/api/search`). Endpoints: GET/PUT/DELETE `/api/sessions/:id/intent` + POST `/api/sessions/:id/readmymind` (`readmymind-routes.ts`, ownership via `findSessionOrFail` WITH `req`; registrations stay the bare `app.<method>('path')` shape, the endpoints.md drift scanner cannot see generics). **Phase 2 (predictor + 🧠 button)**: `readmymind-context.ts` is the PURE budgeted assembler (9 ranked sources, drop order siblings→away→workspace→tools, sections 1-4 truncate only); IO lives in `readmymind-collectors.ts` (transcript TAIL read — the live watcher keeps only a 500-char snippet — + git signals, skipped for remote-SSH cases) and the route; `readmymind-predictor.ts` reuses the AiCheckerBase spawn mechanics standalone (verdict-shaped base vs freeform JSON) as a mutable singleton routes call and tests stub. Claude-mode only (400), one in flight per session (409 CONFLICT), model = `readMyMindModel` setting defaulting to `AI_CHECK_MODEL` (opus, decided). Frontend `readmymind-ui.js`: header 🧠 marker-hidden (`btn-readmymind--hidden`) until the setting is ON; phones hide it in mobile.css and get a keyboard-accessory 🧠 key instead (ships in BOTH bar templates, revealed by the `rmm-enabled` class on the BAR element — setMode() rebuilds button innerHTML, so per-key state would be wiped; synced at init + every `applyHeaderVisibilitySettings()`). Alternate suggestions render as tappable rows that swap into the editable field without losing edits; Rethink rejects the whole shown set and carries the optional steer note (`#readMyMindSteer`, sent as `steer`, shown in ready + empty-result phases, cleared on each open). Suggestions render via value/`textContent` ONLY and Send/Insert go through `POST /input` (server-side, so the sendEnterKey/local-echo trap does not apply) — nothing auto-sends, ever. User guide: `docs/readmymind.md`. **Voice dictation via Claude** (`claudeVoiceEnabled`, SYNCED, default OFF): the mic button can transcribe through this machine's Claude Code login instead of a Deepgram key, using the same speech-to-text service the CLI's own `/voice` mode uses. ⚠️ **Claude Code's voice mode itself is unusable here**: it opens the HOST's microphone (`sox`/`arecord`), and the CLI runs in a headless tmux pane while the human is in a browser elsewhere. So Codeman captures in the browser and borrows only the backend. Audio goes browser → Codeman → Anthropic (`src/web/voice-stream.ts`): the OAuth token never reaches the page, and the browser only sends PCM and receives text. ⚠️ Credentials are **read-only** (`src/claude-credentials.ts`) and Codeman never refreshes them — a refresh rotates the refresh token and could sign the user out of their own CLI; an elapsed token reports `expired` instead. ⚠️ Capture MUST be linear16/16 kHz/mono, so it uses an **AudioWorklet**, not MediaRecorder (which cannot emit raw PCM); `voice-pcm-worklet.js` is fetched from JS, so it is invisible to `cacheBustAssets` and borrows voice-input.js's `?v=` token — **edit the two together**. ⚠️ Transcript frames carry the WHOLE running transcript, not deltas: the Claude path replaces where the Deepgram path appends. Provider choice is `voiceSettings.provider` (`auto` prefers Claude → Deepgram → Web Speech). → `docs/claude-voice-plan.md` **Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`. **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) **Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set); Shell selection and automatic drop recovery always use a bounded 1 MiB `?tail=` window. Shell loads the rest only when **Load full history** is pressed; ordinary scrolling must not trigger a multi-megabyte reset+replay on xterm's main thread. Other modes may re-pull at the TOP (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). Live writes are one-chunk-in-flight, released by xterm's parse callback, so xterm's private queue cannot bypass the browser's 128 KiB render cap. While WebSocket owns terminal I/O, duplicate SSE terminal events are dropped before JSON parsing, and recovery is single-flight per active session. ⚠️ **A `full=1` capture ENDS with a cursor move back to the pane's own caret position**, counted UP from the last replayed row — without it the caret stays where the last character landed, which for an agent CLI is the status line, and every cursor-relative update the CLI sends afterwards is measured from the wrong row. The move is relative, not `CUP`: absolute row addressing is only right while the browser's rows equal the pane's, and `resizeWindow` does not wait for tmux, so a capture can be taken before a requested resize applies. That makes row alignment load-bearing on this path: no transform that can DELETE A LINE may run over the capture, so it keeps its trailing blank rows and skips redraw-bloat stripping, the banner trim and the leading-whitespace strip. ⚠️ Those three skips key on whether a capture actually CAME BACK (`isFullCapture`), never on `?full=1` alone — the fallback to the byte history is a stream of successive frames that must still be stripped, and a session with no mux takes it on every load. A capture holding nothing visible returns '' so the byte history survives instead of a blank screen replacing it. ⚠️ A full re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. ⚠️ **A visible capture now REPORTS the geometry it was taken at** (`captureCols`/`captureRows`, #435), because a frame built for a pane taller or wider than the browser is damaged two ways at once (overflow rows clamp onto the last line; a narrower browser wraps every painted row) and nothing in the response used to say so. Both fields are ABSENT when no frame was positioned, so every consumer tests `Number.isFinite`, never truthiness: a `display-message` cursor query that fails makes `capturePaneBuffer` return the raw capture while the route still labels it `mux-visible`. The comparison runs on `mux-visible` ONLY, the replay is capped at one attempt, and a pane that cannot be sized to fit latches in `_geometryRetryUseless` so it is diagnosed once per session rather than on every tab switch. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) **Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions) **Terminal touch gestures: link taps and text selection**: on a touch device xterm's own handlers see neither — `touch-action: none` plus touchstart's preventDefault suppress the browser's compatibility mouse events, `_installMobileTapMouseGuard` drops the trusted ones that still arrive, and the synthetic `mousedown`/`mouseup` pair dispatched for mouse REPORTING goes to the `.xterm` root, an ANCESTOR of the screen element the linkifier and SelectionService listen on. So both gestures are driven explicitly. ⚠️ **A tap activates the link under it** through the SAME provider that feeds the hover linkifier (`_terminalLinkAtPoint`, containment mirroring xterm's `_linkAtPosition`), synchronously inside `touchend` — that is what keeps the user gesture `window.open` needs — and BEFORE any mouse report, mirroring `_handleDesktopTerminalClick`'s skip for a hovered link. Two rows keep their meaning: the caret's logical line (`_tapIsOnCaretLine`, where a tap places the cursor in text the USER typed) and TUI-owned rows (`_isActionableMobileTerminalTap`, answering a dialog). ⚠️ The caret line is the boundary rather than the tap INTENT, because a shell classifies every tap as `'input'` and gating on that would leave every URL in shell output inert. ⚠️ **Long-press selects** by driving xterm's public `select()` (renderer-independent — under WebGL the glyphs are pixels and native selection cannot exist), drag or a further tap extends, and Copy goes through `copyTerminalSelection()` for its execCommand fallback on plain-HTTP installs. Three guards are load-bearing and each came from a real phone: the compat mouse pair after `touchend` (xterm focuses on mousedown and SelectionService resets the model there, so the keyboard sprang up and the selection vanished on lift), the platform's own ~500ms long-press (Android Chrome focuses the nearest editable element — the helper textarea — through no event a handler can preventDefault, so a bounded focus guard blurs it and `contextmenu` is suppressed for the gesture window), and `copyTerminalSelection()`'s closing `terminal.focus()` (right on desktop, wrong on a phone). Tests: `test/terminal-touch-tap.test.ts`. **Auto Copy (copy-on-select)** (`autoCopySelection`, per-device, default OFF): a finished terminal selection lands on the clipboard with no keystroke. ⚠️ It fires at the END of a gesture, never in `onSelectionChange` (that callback runs per cell crossed, so copying there is one clipboard write per mouse move); it only ARMS `_autoCopyPending`, and a document-level `mouseup` listener flushes. ⚠️ The flush is SYNCHRONOUS inside the handler because both clipboard paths need user activation (Firefox gates `navigator.clipboard.writeText` on it, and the plain-HTTP `execCommand` fallback must run in the gesture's own task); a timer or a wait for `onSelectionChange` loses it, invisibly in Chrome. ⚠️ Touch needs its OWN calls from `_endTouchSelectionGesture()`/`_selectTouchSelectionLine()`: that path `preventDefault()`s its touchend, so no mouseup ever arrives and the toggle would be dead on phones. ⚠️ Unlike `copyTerminalSelection()` it must NOT clear the selection (the text would vanish under the cursor that highlighted it) and must NOT focus the terminal (that opens the on-screen keyboard over it); focus is RESTORED to whatever held it, which only matters for the `execCommand` fallback. Guards are pure in `decideAutoCopy()` (constants.js): off, blank/whitespace-only, and a 1M-char cap (an autoscrolling drag can sweep the whole 50k-line scrollback), refused rather than truncated with a toast pointing at Ctrl+C. Silent on success except once per page load; failures toast, throttled 10s. Tests: `test/terminal-auto-copy.test.ts`. **Ctrl+V paste trap** (`image-input.js`): `Ctrl+V` routes through `_handleImagePaste()`, which appends a hidden `contenteditable` trap, focuses it, and reads the clipboard out of the paste event that lands there. Images upload and their saved paths are typed into the session; text goes through `terminal.paste()` so bracketed-paste markers survive. ⚠️ **The trap must consume exactly ONE paste event.** Two routes deliver one for a single keypress and Firefox fires both: `document.execCommand('paste')` dispatches a trusted event and still returns `false`, because the trap cancels it, while Chromium refuses that command and dispatches nothing; separately, the keydown's own default action delivers a paste to the now-focused trap, because returning `false` from the custom key handler never cancels the DOM event (see smart copy above). Measured on a live install: Firefox two events per keypress, Chromium and WebKit one. Handling both wrote the clipboard to the PTY twice, and right-click → Paste stayed correct because it carries no keydown. ⚠️ Removing the `execCommand('paste')` call would also end the doubling, and all three engines still deliver one event without it, but it stays for the mobile engines a desktop measurement cannot reach: where a browser aims the key's default action at the element focused when the keydown began, the command is the only route into the trap, and the trap is the only place image blobs are read. The one-shot flag lives on the trap rather than on a browser check, so any count produces one insert. Tests: `test/image-paste-trap.test.ts`. → [architecture-invariants#terminal-paste-ctrlv](docs/architecture-invariants.md#terminal-paste-ctrlv) **Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity/omp get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for **claude ≥ 2.1.187 ONLY** at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. ⚠️ Codex was in that list and must never go back without a fresh measurement: codex-cli 0.147.0 ignores SGR wheel reports entirely (`mouse_any_flag=0`, inline viewport, transcript pushed into terminal scrollback), so forwarding produced a dead wheel (#227 follow-up). `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). ⚠️ When that gate is FALSE on a claude session whose local buffer is hollow (`baseY === 0`), the gesture becomes coalesced PageUp/PageDown key sends (`_maybePageCliTranscript`) instead of a no-op; ⚠️ and `getClaudeCliVersion()` must never cache a FAILED probe (one timeout used to disable forwarding process-wide until restart). ⚠️ **A click is hand-reported to the CLI only while the CLI actually has mouse tracking on.** The full strip removes the mouse DECSETs, so xterm's `mouseTrackingMode` is permanently `none` there and the browser hand-encodes SGR reports (`_sendSyntheticSgrTap`); without state it did that on EVERY click, so a stripped-mode pane running a plain shell (CLI exited, or a shell started inside a claude-mode session) received reports it never asked for and printed them as literal text (`[<0;88;20M`), garbling the next typed line. `_recordStrippedMouseMode()` (session.ts) records what the strip removes, `toState()` publishes `cliMouseTracking`, and `_shouldReportMouseToCli()` gates all three report sites on it. Only 1000/1001/1002/1003 count (1005/1006 are encodings, 1007 is alt-scroll), and the change broadcasts UNdebounced since a dialog can be clicked inside the 500ms window. `_logScrollRouting()` prints the routing decision and its inputs once per session — read it before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding) **Detached start + service install** (issue #231): `codeman web -d` relaunches the SAME entry script with `detached:true` (setsid), so there is no controlling terminal and no shell job entry. ⚠️ `nohup` is NOT what makes this work: Node re-arms SIGHUP to its default disposition even when it inherits "ignore", and `cli.ts` handles SIGHUP with a graceful shutdown, so a delivered HUP still stops the server. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile check + `/api/status` probe): a second instance on the shared tmux socket attaches PTYs to the first one's live sessions. ⚠️ Neither may report success it has not observed — the parent polls `/api/status` until the child answers or dies, since `launchctl load` and a clean spawn are both silent about a server that starts and immediately exits. `--stop` verifies the pid still LOOKS like a Codeman server (`ps -o command=`) before signalling, because pids get recycled. Unit/label names live in `config/service-names.ts` so install.sh, `detectSupervisor()` and `service install` cannot drift into supervising two copies; they are instance-scoped, and identical to the historical names for the default instance. `service install` bakes the installing shell's PATH into the unit (launchd gives a job `/usr/bin:/bin:/usr/sbin:/sbin`, which finds neither a Homebrew/nvm `node` nor `tmux`/`claude`) and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install) **Self-update** (App Settings → System → Updates): in-app updater for git-clone installs supervised by systemd/launchd (`systemd`, `launchd`, `launchd-daemon`, `docker-compose`, else `none` → "restart manually"). The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` that outlives the restart and writes progress to `update-status.json`, which the browser polls across the connection drop. `src/web/self-update.ts` splits pure helpers (unit-tested) from IO wrappers. npm installs report as non-updatable. ⚠️ **The Compose deployment is the one supervisor that does NOT outlive the restart**: there the restart IS the container exiting (`restart: unless-stopped` relaunches it), which kills the script too — safe only because the terminal `restarting` marker is written BEFORE the kill, so nothing may be appended after it. Two config facts make it work at all and both are load-bearing: the repo is a HOST BIND MOUNT over `/opt/codeman` (a pull into the baked image copy would land in the writable layer and be silently discarded by the next `up`), and the runtime image keeps devDependencies + a build toolchain (`npm run build` is tsc+esbuild, and node-pty has no Linux prebuild), which is why `npm prune --omit=dev` is gone and the updater passes `--include=dev` against `NODE_ENV=production`. ⚠️ An in-place container update applies CODE ONLY — a restart reuses the existing image and config — so `evaluateEnvironmentGate()` REFUSES a release that changes `server.Dockerfile`/`docker-compose.yaml` (sha256 vs the baseline `Start-Codeman.sh` writes to `docker-env-applied.json` on every start) or adds `.env.example` keys the user's `.env` lacks, and refuses when the restart policy would not bring the container back. That third check exists because **Compose resolves an unset `${VAR}` to the EMPTY STRING and starts anyway**, so a new required setting otherwise arrives as a silently blank env var. Every unknown fails OPEN in the gate (no baseline, unreadable `.env`, no socket): failing closed would permanently block containers created before the fingerprint file existed. ⚠️ The KILL does not: the server exits only when `--restart-by-exit 1` was passed, i.e. the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (set ONLY there, since that file is what sets `restart: unless-stopped`; the image ENV deliberately does not) or the daemon reported an auto-restart policy; otherwise the build lands as `completed-needs-manual-restart`, because exiting blind takes a `docker run` container with no restart policy down with no UI left to recover it. The gate is re-evaluated on `POST /api/system/update`, so hiding the button is UX, not the control. ⚠️ The four global agent CLIs in `server.Dockerfile` are PINNED on purpose — unpinned, a user's CLI versions are a function of when their image was built rather than of any commit, which is the one environment change no diff-derived gate can see; pinning turns it into a Dockerfile change the gate already catches. `test/docker-compose-env-parity.test.ts` is the merge-side guard (every compose `${VAR}` ↔ an `.env.example` entry). → [docs/docker-self-update.md](docs/docker-self-update.md), [architecture-invariants#self-update](docs/architecture-invariants.md#self-update) **Reverse-proxy base path** (`--base-url` / `CODEMAN_BASE_URL`, default `/`; `src/config/base-path.ts` is the pure single-source, normalized to `''` for root or `/foo`): lets Codeman be mounted under a sub-path behind a proxy that **forwards the prefix unchanged** (does NOT strip it). Deliberately few choke points, mirrored ingress/egress: **(server ingress)** `stripBasePath()` runs inside Fastify's `rewriteUrl` so ALL routes stay declared prefix-agnostic (`/api/...`, `/ws/...`) — and a request arriving WITHOUT the prefix (hooks, health checks, docker bridge, all hitting the raw port) is left untouched, so the server answers at both; **(server egress)** one `onSend` hook prepends the base to every root-absolute `Location` header, covering all redirects; **(HTML)** `renderIndexHtml` rewrites the shipped `<base href="/">` to the mount and injects `window.__CODEMAN_BASE__` — the template's asset refs are all RELATIVE so `<base>` handles them for free; **(frontend runtime URLs)** root-absolute URLs ignore `<base>`, so `CodemanBase.url()` (constants.js) is the route builder, applied transparently by a `fetch` wrapper and explicitly at the few EventSource/WebSocket/`window.open`/`<img|iframe|a>`-src sites; **(sw.js/manifest)** the worker derives its base from `self.location`, the manifest uses relative `start_url`/`scope`; **(web-tab proxy)** `proxyPrefixFor(cap, basePath)` is the single base-aware root that cascades to the injected `<base>`, root-absolute HTML rewrites, the `runtimeUrlShim`, `Set-Cookie` Path and `Location` rebasing — while the INGRESS parsers (`capabilityFromProxyPath`, `resolveUpstreamUrl`) stay base-agnostic because `rewriteUrl` strips the prefix before routing, and `capabilityFromReferer(referer, basePath)` strips it from the browser-supplied Referer. ⚠️ `--base-url` rides the daemon relaunch via `buildWebArgs` and the service unit via `resolveServicePlan`. Pure helpers unit-tested in `test/base-path.test.ts` + `test/webview-proxy.test.ts`; HTML injection in `test/render-index-html.test.ts`. **Attachments** (live external document references; all wiring in `file-routes.ts`): a **registry** maps a stable `attachmentId` to a realpath-resolved, extension-allowlisted absolute path, so browser requests never carry arbitrary absolute paths. ⚠️ The **magic-link scanner** (`codeman://attach?...` in terminal output) is **prompt-injectable**, so its scan path is force-confined to the session workspace; a hostile prompt could otherwise exfiltrate arbitrary host files over SSE. The security gate is an extension **allowlist**, not a blocklist. `document-conversion-limiter.ts` caps converter spawns globally: without it, N large docs detected at once fork N multi-minute processes, which is a resource-exhaustion vector. → [architecture-invariants#attachments](docs/architecture-invariants.md#attachments) **File-path links (terminal + chat)**: a path an agent prints is clickable on BOTH surfaces and opens the file-preview overlay. ⚠️ ONE pattern (`FILE_PATH_LINK_PATTERN` / `absoluteFilePathPattern()` in constants.js) feeds the xterm link provider AND the response viewer's `_linkifyFilePaths()`; a fresh instance per call, since `lastIndex` is per-object state. The chat linkifier walks TEXT NODES with DOM APIs (the source is model output; never rebuild sanitized markup as a string) and skips subtrees already inside an `<a>`. ⚠️ **An out-of-workspace path is served through the ATTACHMENT routes, not the file routes** — `file-content`/`file-raw` are workspace-confined and 404 exactly the paths agents print most (a `/tmp` capture, Claude's scratchpad), so `openFilePreview()` registers such a path via `POST /api/sessions/:id/attachments` with **`notify: false`** (suppresses only the `attachment:detected` broadcast — same guard, same routes; without it every click also popped a card announcing the file already on screen) and renders by id. The click is an explicit action on the explicit, Origin-guarded route, which is what distinguishes it from the force-confined magic-link scanner. ⚠️ **Media extensions are single-sourced** (`VIDEO_ATTACHMENT_EXTENSIONS`/`AUDIO_ATTACHMENT_EXTENSIONS` in `attachment-registry.ts`, imported by `file-content`'s classification) so a clip plays the same in or out of the workspace; a player needs all THREE of allowlist + a real `MIME_TYPES` entry (octet-stream renders a dead player) + the range-aware body. ⚠️ **`TEXT_ATTACHMENT_EXTENSIONS` IS `EDITABLE_EXTENSIONS`** (never a second list): if the viewer would edit it inside the workspace, it can be read outside. Widening READ must never widen RUN, so `html`/`htm` joined `svg` in `serveRawFile`'s download-only branch, other text goes out as inert `text/plain`+`nosniff`, and `~/.codeman*/state.json` joined `isSensitivePath` (it persists `envOverrides`, which can hold `GEMINI_API_KEY`). ⚠️ The terminal sends an **out-of-workspace** path to the preview instead of the log viewer (that one spawns `tail -f` and reaches only workspace + `/var/log` + `~/logs`); in-workspace text keeps the tail viewer and `file-stream-manager`'s allowlist is untouched. The image-watcher keeps its own narrow detection list, so none of this cards every file an agent writes. → [architecture-invariants#file-path-links-terminal--response-viewer](docs/architecture-invariants.md#file-path-links-terminal--response-viewer) **Filesystem path picker** (Link Existing "Browse" + the mobile keyboard's `📁 Path` key): lazy one-directory browsing via `GET /api/filesystem/browse`, with `GET /api/filesystem/preview` for the tapped file. Inserts the path **without** Enter, so the prompt is never submitted; the sibling `⌫ All` key clears only the unsent prompt and must never send the agent's `/clear`. ⚠️ This is a **second file-serving surface and inherits neither the attachment confinement nor its ownership scoping** — it allowlists Home, `CASES_DIR`, `/mnt/d` and `CODEMAN_FILE_PICKER_ROOTS`, blocks sensitive trees, and rejects symlink escapes **after** `realpath`. ⚠️ The optional `sessionId` is an ownership boundary that must be `canAccessOwned`-checked by hand (it does not go through `findSessionOrFail`), and in multi-user mode a non-admin gets only their own `userSpacePath` as a root: per-user spaces live INSIDE `homedir()`, so a `Home` root exposes every other user's workspace. Previews go through the same global conversion limiter, and Markdown/TXT/JSON are served as inert `text/plain`. → [architecture-invariants#filesystem-path-picker](docs/architecture-invariants.md#filesystem-path-picker) **File Viewer edit mode** (issue #212): the file-preview overlay edits workspace text files in place — `GET .../file-content?edit=1` + `PUT /api/sessions/:id/file-content`, policy in `src/config/file-editing.ts`. This is a **third file surface and the only one that WRITES**: read-path confinement (realpath + workspace + ownership) plus sensitive/blocked/`.git` denies and an extension **allowlist**; writes are `wx`-temp + rename (no `O_CREAT` anywhere = edit-in-place is structural); optimistic concurrency via sha256 `baseHash` → 409. ⚠️ `edit=1` never truncates and the client must never save a plain-preview buffer (the 500-line truncation would silently delete the rest). ⚠️ CRLF/UTF-8 guards: EOL re-applied server-side, non-UTF-8 refused via round-trip compare. → [architecture-invariants#file-viewer-edit-mode](docs/architecture-invariants.md#file-viewer-edit-mode), `docs/file-viewer-edit-plan.md` **Files panel search** (COD-236, the `q` param on `GET /api/sessions/:id/files`): `compileFileQuery()` (`utils/file-query.ts`, pure, no IO, so it unit-tests directly) compiles the query into a reusable predicate, which is what lets the server-side walk prune instead of streaming the whole tree. ⚠️ **A query turns that endpoint into a FLAT match list rather than a nested tree**, and the walk deliberately recurses past non-matching directories, since the whole point of searching is to reach a file whose ancestors do not match. An empty or whitespace-only query compiles to `null`, which is what keeps the default tree response byte-identical when no search is requested. ⚠️ **Globs are never compiled into a RegExp**: `*a*a*a…` translated to `^.*a.*a.*a…$` is a classic backtracking blowup evaluated synchronously against every walked path, so one pathological query would freeze the event loop for the whole server (the same reason `search-service.ts` is regex-free). `globMatch()` is a two-pointer wildcard walk instead, O(text · pattern) with both operands short by construction, and an overlong query (`MAX_QUERY_LENGTH`, 256) also compiles to `null` rather than running. A query containing `/` matches the relative path, otherwise the bare entry name; globs match anchored and case-insensitively (`*` spans any run, slashes included, `?` exactly one character), everything else is a plain case-insensitive substring. **Raw file bodies are streamed and range-aware**: `file-raw`, the attachments `/raw` route and `GET /api/download` always advertise `Accept-Ranges: bytes` and answer a `Range` header with `206` + `Content-Range` (single-range only; parser is pure + unit-tested in `src/web/http-range.ts`, a malformed spec is ignored → 200 while an out-of-bounds one is a 416). ⚠️ **The size cap on all three is a sanity bound, not memory protection** (`MAX_FILE_DOWNLOAD_BYTES` in `config/buffer-limits.ts`, default 2GB, env `CODEMAN_MAX_DOWNLOAD_BYTES`, `0` = unlimited): the bodies stream, so size costs a read stream and not RSS (measured: a 600MB download moved peak RSS by ~37MB). Its predecessor was a hardcoded 50MB whose comment still said "prevent memory exhaustion" long after the `readFile()` it described was replaced by `sendFileBody()`, so all it did was refuse legitimate downloads of build artifacts, videos and archives. `/api/download` was the last route that really did buffer the whole file, and now shares `sendFileBody()` with the other two. ⚠️ A 200-only response is what made the File Viewer's `<video>` unseekable: Chrome then reports `video.seekable` as `[0, 0]`, the scrub bar is inert and `currentTime = x` silently reverts (measured on an 18MB mp4), and Safari refuses to start the media at all. ⚠️ These bodies go out through `reply.hijack()`, which bypasses Fastify's status handling — `sendRawStream` must copy the status onto `reply.raw` by hand or a partial body ships labelled `200` and the browser treats a slice as the whole file. ⚠️ Closing the preview must **pause and unload** the media (`_stopFilePreviewMedia` in panels-ui.js): dropping the overlay's `visible` class is `display:none` and nothing else, and a DETACHED `HTMLMediaElement` keeps playing, which is how the X button used to leave a video audible with no player to pause. **Ultracode / workflow-run visualization** (opt-in, default OFF): the Workflow tool writes a completion artifact only at run *end*, so live in-flight runs exist solely as transcript dirs. `workflow-run-watcher.ts` therefore synthesizes ACTIVE runs from transcripts until the completion artifact appears and supersedes them. It is **STANDALONE** and deliberately never imports or touches `subagent-watcher.ts`, despite reading the same tree. Two independent toggles: `showUltracodeAgents` (docked panel) and `ultracodeFloatingWindows` (floating windows); the watcher starts if **either** is on. → [architecture-invariants#ultracode--workflow-run-visualization](docs/architecture-invariants.md#ultracode-and-workflow-run-visualization) **Clone a repository as a case** (issue #236, Add Case → **Clone Repo**): `POST /api/cases/clone` clones a public repo into the caller's case space synchronously (request held open, bounded by `GIT_CLONE_TIMEOUT_MS`, no job store); `POST /api/cases/clone-preflight` reports whether the URL can be cloned anonymously plus its real branches/tags. Core in `src/git-clone.ts`. ⚠️ **The URL is a code-execution surface**: `ext::sh -c <cmd>` (and ANY `<name>::<payload>` helper) makes git run a command, so every `::` form is refused, a leading `-` is refused, and every spawn is an argv array with `--` before the operands. ⚠️ **Non-interactive or the open request hangs** — `gitNonInteractiveEnv()` closes the terminal/askpass/ssh/GCM prompt paths; `HOME`/`PATH` stay inherited, so a user's OWN credential helper may authenticate (Codeman still never collects or stores credentials, and refuses a `user:password@` URL). ⚠️ Timeout kills the process GROUP (clone fans out into child processes), the destination is removed only if this attempt created it, and repository contents win over scaffolding (existing `CLAUDE.md` kept, hooks merged, repo-shipped `.claude/settings*` reported as a warning since its hooks run locally). The **Brain** picker sets the toolbar run mode on success. → [architecture-invariants#clone-a-repository-as-a-case](docs/architecture-invariants.md#clone-a-repository-as-a-case) **Cross-session search**: `GET /api/search` federates an in-memory search over session metadata, run-summary events, and attachment-history entries. The pure core `searchSources()` does substring matching with hard per-type caps: **no regex (so no ReDoS) and no filesystem reads (so no traversal)**. The server-private `externalPath` is never read. PAST sessions (#261) come from `session-history-index.ts`, a capped snapshot of the unified list filled **outside** the request path (`/api/sessions/unified` publishes it; a stale one is rebuilt fire-and-forget), that indirection is what keeps the no-fs property. ⚠️ The snapshot is stored UNSCOPED with a per-row owner and MUST be re-filtered through `canAccessOwned()` on read; history rows carry `jumpTo.kind:'resume-session'`, since a closed session has no tab to select. → [architecture-invariants#cross-session-search](docs/architecture-invariants.md#cross-session-search) **Web tabs** (dashboard URLs as tabs): a saved URL renders as a tab beside agent sessions. **NOT a `SessionMode` of its own** (no PTY, no tmux, no respawn), same reasoning that keeps Docker/remote-SSH as case overlays. Dashboards are **proxied through Codeman's own origin** by default, because a direct iframe fails three ways at once: prod is HTTPS so `http://` targets are blocked as mixed content, many dashboards send `X-Frame-Options: DENY`, and our own `default-src 'self'` CSP blocks cross-origin frames. Proxying leaves the prod CSP unchanged (`/webview/...` is `'self'`). ⚠️ The proxy is **NOT an API surface**: it authenticates on an in-memory capability in the path and is correspondingly exempt from the cookie + Origin checks; that exemption is fenced to safe methods and non-`/api` paths and is pinned by `test/webview-auth-exemption.test.ts`. ⚠️ Iframes omit `allow-same-origin` unless a dashboard is explicitly marked `trusted`, and `Authorization`/`codeman_session` are stripped upstream in **both** modes so `CODEMAN_PASSWORD` cannot leak. ⚠️ A sandboxed frame is **opaque-origin**, which breaks two things `curl` can never reproduce: its runtime-built root-absolute URLs escape `<base>` (fixed by an injected `runtimeUrlShim()`), and its same-host `fetch`/XHR are CORS-checked with `Origin: null` (fixed by `buildProxyCorsHeaders()` plus exempting the proxy from the global `OPTIONS`-204 short-circuit in `registerSecurityHeaders`). Both present as the dashboard's own "Failed to fetch" while the page renders fine. ⚠️ **Egress guard**: link-local and cloud-metadata targets (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`, the Azure/Alibaba fixed addresses, `metadata.google.internal`) are refused at save time AND on the RESOLVED address at connect time (`webview-egress-policy.ts`, pure, plus `webview-egress.ts`: a `lookup` hook on the undici Agent behind `webviewFetch()` and on the `ws` client). An IP literal never reaches a lookup hook (`net.connect` skips DNS for it), so the synchronous hostname check at each connect site is NOT redundant. Loopback and RFC1918 stay allowed on purpose: a `localhost` Grafana is the feature. The proxy uses the `undici` PACKAGE's own `fetch` + `Agent`, never Node's global fetch with a foreign dispatcher (Node bundles its own copy; a protocol mismatch fails silently). ⚠️ **A loopback link in agent output opens as a proxied web tab** (`openLinkThroughWebTabIfLoopback` in webview-tabs.js, called from the terminal link provider and the response viewer): a phone cannot reach the box's `localhost:5173`, so the tap routes through Codeman's origin instead, reusing a saved same-origin dashboard or saving one under `host:port`. Two rules that look cosmetic and are not. **`*.localhost` is deliberately NOT in the auto-route set** even though it IS loopback to a browser: the link source is agent-written terminal output and prompt-injectable, every other member of that set is an address literal that can only mean this box, and a `*.localhost` DNS name is not one (a resolver with a search domain retries `evil.localhost` as `evil.localhost.<search domain>`, which an attacker can control, turning agent output plus one tap into a persisted server-side fetch of an agent-chosen origin). A user who really runs `api.localhost` saves it by hand, which is an explicit action. The PAGE-side test is deliberately broader (`isOnBoxHostname`), since a false positive there only declines to proxy. And **`this.webviews` being set does NOT mean it is loaded**: `initWebviews()` assigns a truthy EMPTY map synchronously and only then awaits the list, so a tap during page load must join the in-flight refresh (`_webviewsRefresh`/`_webviewsLoaded`) or it finds nothing to reuse and POSTs a duplicate record for an origin that already exists. Capabilities are revoked on logout / admin logout / user deletion (`revokeOwner`, which had NO caller for two releases while its docstring said otherwise), and proxied responses carry `Referrer-Policy: same-origin` so a dashboard cannot hand the capability-bearing URL to a third party. → [architecture-invariants#web-tabs](docs/architecture-invariants.md#web-tabs), `docs/web-tabs.md` **Multi-user mode** (opt-in `--multiuser` / `CODEMAN_MULTIUSER=1`, OFF by default): named users with scrypt-hashed passwords in `~/.codeman/users.json`. Gated everywhere by `isMultiUserMode()`; when OFF, behavior is byte-identical to single-user because every scoping helper short-circuits. ⚠️ **Not a security boundary at the agent layer**: every session still runs as the SAME OS account. This separates WORKSPACES; it does not sandbox users (Docker cases are the isolation story). Ownership threads through `Session.owner` and is enforced in `findSessionOrFail`, list endpoints, SSE routing (fail-closed), WS, search, and file-preview. → [architecture-invariants#multi-user-mode](docs/architecture-invariants.md#multi-user-mode), `docs/multi-user-plan.md` **Away digest**: `GET /api/away-digest` aggregates what happened while you were away from the lifecycle log, run-summary events, live sessions, token stats, and recent subagents. Pure aggregator in `web/away-digest.ts`. ⚠️ Returns `{success:true,digest}`, a legacy raw-ish shape consistent with the other raw GET handlers in `system-routes.ts`; frontend and tests read `.digest`. → [architecture-invariants#away-digest](docs/architecture-invariants.md#away-digest) **Ralph todo-config**: per-session `maxTodos` (FIFO-eviction cap, default 500 = `MAX_TODOS_PER_SESSION`) + `todoExpirationMinutes` (auto-expiry, default 60) set via `POST /api/sessions/:id/ralph-config` (`RalphConfigSchema`, both `.int().positive()`). Stored on the tracker (`setMaxTodos`/`setTodoExpirationMinutes`) and **persisted/read-back via `RalphTrackerState`** (surfaced in the `loopState` getter → `toState()` + SSE broadcast → modal `populateRalphForm`), mirroring how `maxIterations` round-trips. Claude-only (skipped by `isExternalCliMode`). **Port interfaces**: Routes declare dependencies via port interfaces (`src/web/ports/`). Routes use intersection types (e.g., `SessionPort & EventPort`). ### Frontend Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `i18n.js`(1.5) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `terminal-keycode229-recovery.js`(5.55) → `sanitize-html.js`(5.6) → `app.js`(6) → `tab-rail-resize.js`(6.5) → `terminal-ui.js`(7) → `terminal-split.js`(7.5) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `cron-ui.js`(9.7) → `settings-ui.js`(10) → `panels-ui.js`(11) → `readmymind-ui.js`(11.3) → `ultracode-panel.js`(11.5) → `approvals-ui.js`(11.6) → `reboot-restore-ui.js`(11.65) → `admin-ui.js`(11.7) → `session-ui.js`(12) → `host-wake-ui.js`(12.2) → `webview-tabs.js`(12.5) → `mobile-overview.js`(12.55) → `home-sessions.js`(12.56) → `entrance-animations.js`(12.6) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15) → `ultracode-windows.js`(15.5) → `session-lineage.js`(15.6) → `image-input.js`(16). `i18n.js` translates static + newly inserted application DOM while skipping terminal/response/file/user-name surfaces; `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData). `terminal-keycode229-recovery.js` forwards a committed `input` event that xterm's `_inputEvent` guard drops (Chrome-on-Android soft keyboards send `composed: true` after a keydown), and only when xterm emitted no canonical data for that keystroke. ⚠️ **That decision is settled at the NEXT keydown as well as on its own zero-delay timer** (#441): the drain runs from xterm's custom key handler, which fires BEFORE xterm processes that key, so a soft keyboard that commits the last character and sends Enter in one InputConnection transaction puts the character on the wire ahead of the `\r`. On the timer alone that character is not merely late, it is LOST: xterm emits the `\r` first and bumps the canonical counter past the candidate's snapshot, so the candidate stands down (measured, `hell\r` where the user typed `hello`). The trade is that a keydown decides with less evidence than the timer did, since xterm's own keyCode-229 rescue has not run yet; that is safe for Enter, which clears the textarea so the pending diff emits nothing. Ordering is pinned by `test/terminal-keycode229-recovery.browser.test.ts`, which the CI gate does NOT run. **Entrance animations** (`entrance-animations.js`, all OFF by default): opt-in animations for the four things that appear when work starts, chosen per surface via `data-tab-anim` / `data-term-anim` / `data-win-anim` / `data-line-anim` on `<html>`. Defaults are the `legacy` theme, so an untouched install behaves exactly as before and every hook short-circuits on its first line. ⚠️ Tabs and connection lines are **destroyed mid-animation** on every re-render (`_fullRenderSessionTabs()` replaces the strip's innerHTML; `_updateConnectionLinesImmediate()` does `svg.innerHTML = ''`), so both are tracked by id and re-applied to the fresh element with a **negative `animation-delay`** to resume rather than restart. ⚠️ The terminal-pane styles may animate **transform / opacity / clip-path only**, xterm's FitAddon derives rows+cols from `getComputedStyle(parent).width/height`, so animating width/height/padding there would resize the PTY; `test/entrance-animations.test.ts` pins that property allowlist, plus the rule→keyframes→theme-option chain a style silently does nothing without. ⚠️ **`blur` is the ONE style that puts a `filter` on the terminal container**, against the standing rule, because every alternative was measured against a live xterm and does not work: a `backdrop-filter` veil on `::before` blurs perfectly while STATIC and Chrome silently drops the backdrop the moment ANY animation runs on that pseudo-element (the veil computes `blur(15.3px)` and the text behind it stays razor sharp), and driving the radius from rAF buys the same full-screen blur per frame plus main-thread work. The cost the rule exists to avoid is inherent to blurring a terminal, so the style buys it knowingly: opt-in, OFF by default, one ~520ms run per session open, class straight back off, `will-change` still unset. Worst-case price, headless SwiftShader with no GPU: frame deltas 16.7ms → 33.3ms for the run, against 16.7ms flat for `fade`. Do not generalise it — a second filtered terminal style needs its own measurement. ⚠️ The `blur` connection line animates `filter` too, so both kinds of line hold their glow in **`--line-glow`** and both of its keyframes say `blur(N) var(--line-glow)`: the function lists then match and interpolate, instead of the glow vanishing for the run and popping back (a lineage line's glow is a different colour entirely, set per element). Its 100% frame deliberately omits `opacity` so the endpoint comes from the element's own resting value — 0.9 subagent, 0.72 lineage, 0.95 working — which is what `line-enter-fade`'s hardcoded 0.9 gets wrong. ⚠️ Window styles other than `beam` transform the window, which moves the rect its connection line is aimed at; `beam` deliberately animates opacity/filter only so its line can draw toward a stable target. Persisted to its own `codeman:*Anim` localStorage keys (per-device, deliberately NOT in the `.strict()` `SettingsUpdateSchema`); picker in App Settings → Appearance, full per-surface lab at `?animlab=1`. **Mobile tab strip scrolling** (issue #257): under 768px the tab strip is a horizontal scroller (desktop wraps to a second row instead), so the active tab can sit off-screen. Three rules keep it reachable and they only work together: `_updateActiveTabImmediate()` scrolls the selected tab into view via `computeTabScrollLeft()` (pure, in constants.js) using **rect math on the strip's own `scrollLeft`**, never `scrollIntoView()`, which would also scroll the document under a fixed header; `_fullRenderSessionTabs()` **restores `scrollLeft`** across the `innerHTML` rebuild, since ambient rebuilds (a task badge appearing, a session created elsewhere) otherwise snap a mid-swipe strip back to 0; and it re-reveals the active tab **only when it changed** (`_lastRenderedActiveTabId`), so browsing the far end of the strip is not undone by background renders. ⚠️ **The ACTIVE tab is the only one with action icons, and on a phone they can eat it**: `.session-tab.active .tab-name` reserves `min-width: 44px` in the phone block (under 600px), because a short session name rendered a 13px label against a 50px gear+close cluster, putting the tab's geometric CENTRE on the gear, so a thumb aiming at the tab opened Session Options instead of switching (measured at 360/393/430px; only long names cleared it). ⚠️ **The floor is set by the 10th tab onward, not by the tabs you can see**: `.tab-number` renders only for `_tabIdx < 9`, so tab 10 loses 16px + a gap off its left and its centre sits 10px further right. The centre clears the icons when `reserved > icons + rightEdge - leftRunUp - gap` (= 50 + 9 - 17 - 4 = **38px**), hit-testing snaps to whole pixels so 39px still lands on the gear, and the practical floor is 40px — a NUMBERED tab clears it at 20px, which is exactly why reasoning from the tabs on screen would put the centre back on the gear. `test/mobile-tab-tap-zones.test.ts` recomputes that inequality from the stylesheet, so widening the gear or the padding fails there rather than on a phone. The guarantee is centre-off-the-ICONS, not centre-inside-the-label (on a numberless tab it lands in the gap between them, which still switches). Non-active tabs keep their icons hidden and stay tappable end to end. ⚠️ Mobile no longer hoists the active session to the front of the strip: that reordering ran on full renders only, so tab order flipped depending on which render path fired, and it renumbered the Alt+N badges. Scroll-into-view replaces it; do not reintroduce it. **Session list layout: header strip or left sidebar** (`sessionListLayout`, App Settings → Appearance → Tabs, default `header`; per-device policy — it IS in `SettingsUpdateSchema` and persists server-side, but `displayKeys` makes a device keep its own value): with many sessions the horizontal strip stops being scannable, so the list can move into a vertical `<aside>` with a filter box and a live count, collapsible to a 44px rail (`--sidebar-width` 260 / `--sidebar-width-collapsed` 44) via **Alt+B** (`toggleSessionSidebar`; Alt, not Ctrl+B, which must reach tmux/readline in the terminal). ⚠️ **There is ONE `#sessionTabs` element and it is MOVED between hosts** (`#sessionTabsHost` in the header, `#sessionSidebarList` in the aside, `#tabRail` for the vertical rail below), never a second list — so every render path, drag-reorder handler and Alt+N index keeps working unchanged. Exactly TWO functions reparent it and they must run in this order: `applySessionListLayout()` first (sidebar wins), then `applyTabOrientation()` (settings-ui.js), which moves the tabs into `#tabRail` only when the sidebar does not own them. ⚠️ It sets `data-session-list` / `data-sidebar` on `<html>` and must run BEFORE `applyTabWrapSettings()`, which is the one owner of `tabs-two-rows`/`tabs-show-folder` and reads those attributes. ⚠️ **The vertical tab rail** (`tabOrientation`/`tabRailWidth`/`sessionSidebarFontSize`, all per-device display keys that ARE in the schema, like `sessionListLayout`) is a SECOND vertical list next to the sidebar: the orientation setting is silently ignored while the sidebar layout is chosen, desktop/tablet only (`resolveTabOrientation` forces horizontal on mobile), resizable via `tab-rail-resize.js` (which owns terminal refits during the drag). ⚠️ **Detailed rows are a property of a vertical LIST, not of one surface** (`sessionListLayout: 'sidebar-rich'` for the sidebar, `tabRailDetail: 'rich'|'simple'` for the rail, rail default **rich**): both draw the home screen's per-session line (`created 3d ago · working 12m`) plus a status pill, from the SAME row model (`_sidebarRichRow`/`_sidebarRichMetaHTML` in app.js, classified by `_mobileOverviewState`/`_mobileOverviewSince`), and the render paths ask ONE gate, `isRichTabRows()` (= `isSessionSidebarRich() || isTabRailRich()`). ⚠️ Detail rides on its own attribute (`data-sidebar-detail` / `data-tab-rail-detail`) so every existing `[data-session-list="sidebar"]` / `[data-tab-orientation='vertical']` rule keeps matching both variants untouched; a flip of detail ALONE still needs a full render (the stamps line is emitted by the row template, not toggled by CSS) and must re-run `applyTabWrapSettings()`, which owns the folder line and is now rail-aware. ⚠️ The rich CSS rules carry a rail twin as a COMMA-GROUPED selector, never `:is()` (an `:is()` list takes its most specific argument, which would lift the sidebar arm from (0,3,1) to the rail's (0,5,1)). ⚠️ Width is the whole reason there are thresholds: the rich sidebar is 300px (`--sidebar-width-rich`) and a rail that has never been sized defaults to **320** (`RICH_DEFAULT_WIDTH`, the existing Wide preset) instead of 256, because at 256 the stamps line ellipsizes mid-word; a user-narrowed rail drops the created stamp below 288 (`tab-rail-tight`, CSS only) and drops rich rows entirely below 240 (`tab-rail-compact`, which re-renders). A stored width is never overridden. ⚠️ Only detailed rows carry stamps that go stale with no event behind them, so `_startSidebarRichClock()` (20s, rewrites text in place — a re-render would restart every row's animation) must be armed and disarmed by BOTH `applySessionListLayout()` and `applyTabOrientation()`. ⚠️ **Axis decisions must use `_isVerticalTabList()`** (sidebar OR rail), never `isSessionSidebarActive()` alone: the rail leaves `data-session-list` at `header`, and the sidebar-only predicate shipped four rail bugs at once (drag insertion side read from clientX, active tab never scrolled into view, floating windows anchored below tabs instead of beside them, connector redraws skipped on rail scroll). The pre-paint script stamps `data-tab-orientation` (+ `--tab-rail-width`) like it stamps the sidebar keys, or vertical mode flashes through the header strip; the name font size defaults to 12px, the sidebar's historical size, so untouched installs are never restyled. ⚠️ Leaving sidebar mode **clears `_sidebarFilter`**: the filter box only exists in the aside, so a stale filter would hide sessions from the header strip with no reachable control to clear it. ⚠️ On handhelds the aside is an off-canvas overlay rather than a docked rail, and a closed drawer keeps `display: flex`, so it is marked `inert` + `aria-hidden` (`_isSessionSidebarOverlay()`) or its filter box and ~4 tab stops per session stay in the tab order; the DOCKED desktop rail must never be inerted, its rows are still clickable. The desktop home rail (`home-sessions.js`) defers to it, since both dock the session list flush left. ⚠️ **The detailed RAIL additionally wears the home rail's CARD, and the detailed SIDEBAR deliberately does not**: the rail is an occasional, resizable list you scan, the sidebar is a permanently-docked nav column where 20 stacked cards read as a wall — so the card rules are RAIL-SCOPED and were NOT added to the comma-grouped selectors above, which is the one-line change that would silently restyle the sidebar. Card state accents reuse `home-sessions-blink-red`/`-yellow` rather than declaring a second pair, and every state dot rule excludes `.tab-alert-action`/`.tab-alert-idle` by hand, because those alert rules are only (0,3,0) and the rail's are (0,5,1)+. There is no `.active` rule in that block on purpose: `.session-tab.active` already paints border/background/box-shadow `!important`. ⚠️ **Row ORDER is a third rail attribute** (`tabRailSort: 'activity'|'manual'`, `data-tab-rail-sort`, default **activity**): a sorted rail answers the home screens' question with the home screens' answer, `CodemanSessionOrder` over rows classified by `_mobileOverviewState` (`_tabRailSortOrder` in app.js). ⚠️ **It is applied as the flex `order` property, never by reordering the DOM**: `#sessionTabs` stays in `sessionOrder`, so the Alt+N badge (`_tabIdx`, which therefore does NOT run 1,2,3 down a sorted rail — it names a shortcut, not a position), drag-and-drop, the arrow-key walk, the sidebar filter and `_scrollActiveTabIntoView()` all keep reading the list they always read, and a session changing state moves ONE inline style instead of forcing the full rebuild that would restart every card's animation on every SSE tick. The incremental render path therefore has to re-apply it (a state flip adds no tab, so the full rebuild is never reached) and an empty string is what clears it when sorting stops. ⚠️ Web tabs are pinned past the cards by a CSS `order: 9999` rather than an inline one, since `renderWebviewTabs()` emits the same markup for every layout; the flex default of 0 would interleave them among the sorted sessions. ⚠️ `setupTabDragHandlers()` sets `draggable="false"` and returns while sorting is on: the drop rewrites `sessionOrder` correctly and the sort then puts the card straight back, so the affordance would be a lie — `'manual'` is the way back to drag-reordering. ⚠️ **The arrow-key walk is the one place that must follow the SORT rather than the DOM**: `_tabKeydownHandler` steps `querySelectorAll` order, which is `sessionOrder`, so on a sorted rail ArrowDown from the top card landed wherever that session sat in the tab order instead of on the card below it. It now sorts its node list by the COMPUTED `order` first (computed, not inline: web tabs get their `order: 9999` from CSS and would otherwise read as 0 and lead the walk). This is the exception that proves the DOM-stays-in-sessionOrder rule: the badge, drag model and Alt+N index all deliberately keep reading the DOM, and only the thing the user steps with their eyes follows the paint. **Phone overview home screen** (`mobile-overview.js`, phones only, per-device `mobileOverviewEnabled`, default ON): under 600px the "C" logo shows a session overview (NEEDS YOU / CURRENT SESSIONS / PAST SESSIONS) instead of the welcome overlay; tablet and desktop are unchanged. The branch lives in `showWelcome()`/`hideWelcome()` (terminal-ui.js) behind `shouldUseMobileOverview()`, which is **width-driven** (`getDeviceType() === 'mobile'`) because this is a layout decision, unlike the settings namespace which stays handheld-based. ⚠️ The container ships with the `hidden` attribute and only this module removes it: never give `.mobile-overview` a bare `display` rule, since desktop does not load `mobile.css` (`media="(max-width: 1023px)"`) and would then render it unstyled. Live re-renders ride on the tail of `_renderSessionTabsImmediate()` (every state change it needs already funnels there); PAST rows come from one `_fetchUnifiedSessions(60)` per home-screen visit and resume through the shared `resumeHistorySession()`, so they behave exactly like the welcome screen's Resume list. ⚠️ Two things must stay in lockstep with surfaces outside this module, because divergence reads as a bug rather than a style: the split Run button carries the **toolbar's own classes** (`btn-toolbar btn-run mode-<backend>` / `btn-run-gear`) so the per-backend gradient and the light-skin overrides apply unchanged (mobile.css must therefore set no `background`/`color` on it), and row status uses the **session-tab language** (green dot when fine, `pulse` while working, yellow blinking row when waiting for input, red blinking row when a question is pending, mirroring `tab-alert-idle`/`tab-alert-action`). The picker mirrors the toolbar run-mode menu (`setRunMode()` + `run()`, `openWebviewFromMenu()` for saved dashboards) and deliberately omits its Recent-Sessions block, since PAST SESSIONS is that. Status pills carry `data-i18n-skip` (generic words like "idle" collide with state strings elsewhere). **Desktop home tab rail** (`home-sessions.js`, desktop only): the welcome overlay centers ~560px of content in a ~1400px window, so its left gutter is dead space; it carries the open tabs as a rail **docked flush to the left edge, full height** (a vertically centered card floating mid-gutter read as debris). Rows are in **overview order** (see below), and each carries a **created** stamp plus the **state duration** the order is computed from (`created 3d ago · working 12m`, word and anchor from `_mobileOverviewSince()` so both home screens say the same thing). A rail sorted by a number it does not show reads as arbitrarily shuffled, and a working row's plain last-active stamp always says "just now". ⚠️ The number badge is the **Alt+1..9 index**, i.e. the position in the TAB STRIP, so on a sorted rail it deliberately does NOT run 1,2,3 downward: it names a shortcut, not a row position, and renumbering it to look tidy would make every badge lie. State classification is REUSED from mobile-overview.js (`_mobileOverviewState`/`_mobileOverviewCaseFor`), which is why the module loads after it. ⚠️ The rail is `position: absolute` so the centered content never moves, which is exactly why it needs a **width gate in two places** — `HOME_SESSIONS_MIN_WIDTH` (1180) in the JS plus a `max-width: 1179px` media query as the backstop for a resize that outruns the matchMedia listener; drift between them means a rail overlapping the search panel, and `test/home-sessions.test.ts` pins them equal. ⚠️ `.home-sessions` is `display: flex`, so `[hidden]` must be re-asserted as `display: none` or the module's only visibility lever does nothing. ⚠️ Size scales with the viewport off **one knob**: `width: clamp(250px, 19vw, 430px)` plus a fluid `font-size` on `.home-sessions`, with every child sized in `em` — reintroducing `rem`/px type inside the block silently breaks the scaling, and widening the clamp past the gutter reintroduces the overlap the gate exists to prevent. The age stamps are refreshed **in place** by a 20s clock (`_tickHomeSessionsTimes()`, disarmed in `hideHomeSessions()`), never by re-rendering, which would restart every row's blink and working ring. Working state is deliberately byte-identical to the phone's: pulsing green dot + the `tab-load-spin` ring reused from the tab strip + the same green halo (added to `.mobile-overview-dot--working` at the same time), so "working" reads the same on every surface; **idle** is deliberately NOT that green — dot and pill mix toward `--text-muted` so a glance separates running from sitting. Live re-renders ride the tail of `_renderSessionTabsImmediate()` alongside the phone overview. **Home-screen session order** (`CodemanSessionOrder` in constants.js, pure + unit-tested in `test/session-overview-order.test.ts`): BOTH home screens (phone overview and desktop rail) order rows through this ONE comparator, because they list the same sessions and must answer "which of these wants me next?" the same way. Rank is `needs` → `error` → `waiting` → `working` → `idle` → `done`, and ⚠️ **the tiebreak flips direction halfway down**: states a session is still IN sort **oldest-first** (blocked longest / running longest = most urgent), states it has STOPPED in sort **newest-first** (the session that just went quiet is the one you came back for). ⚠️ The running group keys off **`lastSubmitAt`** (the pane's last Enter), never `lastActivityAt`: a working Claude pane repaints about once a second, so its last-activity stamp is always "now" and would rank every running turn as freshly started. A working pane with no submit stamp falls back to last activity, which lands it at the SHORT end of the group rather than falsely leading it. ⚠️ A **0 stamp means "unknown", not "the epoch"**, and it sorts last within its state either way, or a brand-new session would head every oldest-first group. Final tiebreak is the user's tab order (`orderIndex`), so the list is deterministic and cannot shuffle between renders. The tab strip itself is NOT sorted by this; it stays user-ordered and drag-reorderable. **Welcome "Resume Conversation" list** (terminal-ui.js): `loadHistorySessions()` fetches once and caches the corpus on `_historyAll`/`_historyCases`; every subsequent view (filter box, sort select, expand, the periodic refresh in panels-ui.js) goes through `_renderHistoryList()`, so never append rows to `#historyList` directly or re-fetch to re-sort. ⚠️ The box height is **class-driven**: expanding the list without `.history-list.expanded` leaves the collapsed `max-height` in place and just deepens a scroll well, which is the bug #260 reported (35 sessions in a ~4-row box). ⚠️ The A–Z sort keys off `_historyRowLabel()`, the SAME string the row renders (`name || firstPrompt || path`), most rows are transcript-backed and have no session name, so sorting on `name` alone silently does nothing. ⚠️ A filter implies expansion, and `_renderSearch()` hides `#historyHeader` (title + controls) as one unit while a search is active. Tests: `test/history-list-controls.test.ts`. **Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)** lives in that same handler: with a selection it copies, with none it must `return true` **without** `preventDefault()` or the interrupt is lost. `copyTerminalSelection` is deliberately absent from `SHORTCUT_ACTIONS` because the generic capture loop preventDefaults every match it dispatches. ⚠️ **The gate tests the CLEANED selection, not the raw one** (`CodemanCopySelection.clean` in constants.js, pure; `cleanedTerminalSelection()` reads the live terminal): xterm returns whole screen ROWS and trims only never-written cells, so the spaces a full-screen TUI paints across the rest of a row are content and reach the clipboard (138 of them per line, measured in a 282-column pane). The clean drops each line's TRAILING run and nothing else. ⚠️ **A shared LEADING indent is deliberately NOT stripped**, and that is a decision, not a gap: it was built, measured and dropped before #451 merged, because it fired on 73% of real non-TUI text (401,445 windows sampled) and no width threshold separates a TUI margin from content (a claude pane's own margins are 2 and 5 columns; the commonest non-TUI shared run is 4). Do not re-add it without reading the rule in architecture-invariants. ⚠️ Alt+drag COLUMN selections are returned untouched (`_activeSelectionMode === 3`), and the padding-only clear is FEEDBACK rather than interrupt protection, since a padding-only selection now cleans to `''` and falls through to the PTY on its own. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry) **Per-device vs synced settings**: the `displayKeys` set in settings-ui.js is a **client-side merge policy**, not a wire filter. A display key seeds from the server only when localStorage has no value for it, which is what prevents one device overwriting another; `showPlanUsageLimits` is additionally `delete`d from the incoming payload outright. Separately, `SettingsUpdateSchema` is `.strict()` and simply **does not declare** `skin`, `showFileViewerButton`, `showCronButton`, `webglRendererEnabled`, `localEchoEnabled`, `cjkInputEnabled`, or `extendedKeyboardBar`, so sending one of those is a validation error. The rest (`showResponseViewer`, `showPlanUsageLimits`, `language`, and most `show*` keys) ARE in the schema and do persist server-side; they are per-device by client policy only. ⚠️ Adding a new per-device setting means deciding **both** questions: membership in `displayKeys`, and presence in the schema. **Settings surface** (`#appSettingsModal` + `#sessionOptionsModal` + `#createCaseModal`): the `set-*` language (left rail, groups of rows, control pinned right) is shared by all three modals through ONE `:is(#appSettingsModal, #sessionOptionsModal, #createCaseModal)` scope in styles.css: an `:is()` list takes its most specific argument's specificity, so every rule keeps the id weight it had and nothing downstream shifts. **App Settings** is a rail that is a **table of contents over ONE scrolling document**, not a tab switcher: every section stays mounted (`.set-section`, ids `settings-updates|terminal|layout|appearance|models|clis|notifications|voice|shortcuts|system`, in that order, the version and the updater leading and the rest of the system settings tailing), and `switchSettingsTab(id)` keeps its historical name but SCROLLS instead of hiding. **Session Options** and **Add Case** use the same surface with a rail that really SWITCHES (`switchOptionsTab` / `switchCaseModalTab` show one `.set-section` and `.hidden` the rest, since Summary owns its own scroller, Respawn is long, and Add Case is six independent forms). ⚠️ They also take a deliberate **size-up** that App Settings does not (900px shell, 236px rail, `height:auto` between `min(560px,80vh)` and 88vh, vs App Settings' tight 760×620): they are short task panels, not a document you scan, and at scanning density they read as a few fields marooned in an empty frame. Those per-modal blocks are the design, not drift. Phones (≤860px) give App Settings the sticky `#appSettingsJump` pill and give the other two a horizontal rail strip, which neither has a pill for. ⚠️ The Session Options rail entry labelled **Session** still keys off `context` (`data-tab="context"`, `#context-tab`, `switchOptionsTab('context')`), the rename is label-only. Add Case keeps its legacy `.form-row` markup (six panels of it, every id read back by session-ui.js) and is mapped onto the look by an adapter block scoped to `#createCaseModal .set-doc`. Do not restructure those forms just to reach the row classes. ⚠️ That adapter's `summary { display:flex }` **kills the native disclosure triangle**, so every `<details>` there needs the explicit `.set-adv-chev` and both marker suppressions (`list-style` + `::-webkit-details-marker`); without it five collapsed blocks render as plain headings nobody clicks. ⚠️ **The load/save contract is `getElementById` by id**: `openAppSettings()`/`saveAppSettings()`/`openSessionOptions()` read every control by a fixed id, so moving a control between sections is free but renaming or dropping one silently stops it loading or saving. Static guards: `test/app-settings-structure.test.ts` + `test/session-options-structure.test.ts` (rail↔section pairing, one-visible-section, the `data-claude-only` entries external CLIs drop). ⚠️ Model cards (`#appSettingsModelCards`) and the effort segment are **views over hidden `<select>`s** that remain the source of truth; the cards hold the BASE model and the "1M context window" switch composes `base + [1m]` back into `claudeModel`, which is what retires the old "takes precedence over the toggle below" trap. ⚠️ `.modal-tabs`/`.modal-tab-btn`/`.modal-tab-content` are RETIRED: no modal uses them and their CSS is deleted, and a reappearance means a modal drifted off the shared surface. ⚠️ The **Header & Panels live preview** is a scale model rebuilt from the chips (`_syncLayoutPreview`); it owns NO icons, it CLONES `.set-chip-ico` out of the chip, so each icon has exactly one copy in index.html. A chip joins it via `data-preview` (slot) + `data-preview-order`, or `data-preview-text` for readouts that are not buttons. Its frame is painted from skin tokens only (hardcoded black alphas turned it into a grey slab on the light skins) and is `data-i18n-skip`. ⚠️ In Session Options → Respawn, auto-resume is a `.set-callout` whose `<label>` **wraps its own switch with no `for=`** (nesting associates them; the label+`for` pair has historically double-fired), and the cycle steps are real checkboxes (`.set-checks`), not chips. ⚠️ `admin-ui.js` injects the multi-user Users entry into `.set-rail-items` + `.set-doc`, so those hooks must survive any restructure. → [architecture-invariants#settings-surface-app-settings-session-options-add-case](docs/architecture-invariants.md#settings-surface-app-settings-session-options-add-case) **Header button visibility**: most header controls are opt-in and hidden by a marker class (`btn-multimonitor--hidden`, `btn-response-viewer-header--hidden`, `btn-file-viewer--hidden`, `btn-cron--hidden`) that `applyHeaderVisibilitySettings()` (settings-ui.js) toggles after settings load; the multi-monitor button is instead stripped at render by `renderIndexHtml`. ⚠️ Hiding must go through the marker class: the base rules are `display:inline-flex !important`, so an inline style cannot override them. Current desktop default is WS/CPU/MEM + File Viewer + gear, with the token chip and lifecycle-log button OFF. ⚠️ New header controls must not leak onto phones; `test/mobile-header-buttons-policy.test.ts` is the static guard. → [architecture-invariants#header-button-visibility-multi-monitor-response-viewer-file-viewer-cron](docs/architecture-invariants.md#header-button-visibility-multi-monitor-response-viewer-file-viewer-cron) **Gesture control** (camera hand-tracking overlay, opt-in, default OFF): `CODEMAN_GESTURE=1` makes the feature *available*; `gestureControlEnabled` turns it on. The bundle is injected by `renderIndexHtml` only when enabled, which is why that method is `async` and reads settings with `readSettings(true)` (a fresh read: a post-save reload lands inside the 2s cache TTL and would otherwise render the pre-toggle state). **Source lives in `packages/gesture-control/`; edit there, run `npm run build:gesture`, and commit the regenerated bundle** because dev serves the committed bundle with no runtime bundler. The MediaPipe wasm + model are fetched separately and gitignored. ⚠️ Keep `MP_VERSION` in `fetch-gesture-assets.mjs` in sync with `@mediapipe/tasks-vision`. → [architecture-invariants#gesture-control-the-source-package](docs/architecture-invariants.md#gesture-control-the-source-package) **Terminal font weight** (`terminalFontWeight` / `terminalFontWeightBold`, per-device, default = xterm's own `normal`/`bold`): bold text on the theme's default foreground carries exactly ONE cue, the weight step. Claude Code marks its markdown bold with a bare `ESC[1m` and no colour change, and xterm substitutes a bright colour for bold only when the foreground is a palette index 0-7, so the substitution never fires there. A two-face family keeps that step small and 400 stays 400 whatever family is chosen, which is why the NORMAL slot is settable at all. `CodemanTerminalFont.resolveWeights()` (constants.js, pure) resolves both slots, each against **its own** xterm default, so an unset bold weight can never inherit `normal`. ⚠️ **The `@font-face` descriptor, not the file, is what the browser synthesizes from**: `fonts/jetbrains-mono-variable.woff2` carries a `wght` axis of 100-800, and while `styles.css` declared it `400 700` every weight below 400 rendered identically to 400 and 800 identically to 700 — measured — so the setting was a no-op for anyone without Fira Code or Cascadia Code installed, which is most installs. It is declared `100 800`; re-narrowing it silently guts the feature (`test/terminal-font-weight.test.ts` pins the range). ⚠️ A live save must reach **both echo overlays** (`refreshFont()` — they cache `terminal.options.fontWeight` and paint it into their spans, so typed characters otherwise keep the old weight, most visible on a phone) **and open Agent Teams panes** (they read their options at construction, exactly like `applyTerminalSkin()` propagates). ⚠️ `_awaitTerminalFont()` is deliberately untouched: `CharSizeService` measures through the CSS `font` shorthand, which RESETS the weight, so the measured face is always the 400 one and a weighted descriptor would ask for nothing new. **Theme skins / branding / i18n**: `skin` selects a palette via `data-skin` on `<html>`, applied by an **inline pre-paint script** in `index.html` reading `localStorage['codeman:skin']` to avoid a flash of wrong theme. ⚠️ A skin is **four things that must stay in sync**, and missing any one degrades silently: the `html[data-skin="…"]` token block in `styles.css`, the xterm ANSI palette in `terminal-ui.js`, the pre-paint allowlist, and the Settings picker (both in `index.html`). `test/skin-themes.test.ts` is the static guard. Light skins additionally need `color-scheme: light` and xterm `minimumContrastRatio: 4.5`, and `applyTerminalSkin()` must call the local-echo overlay's `refreshFont()` because it caches the terminal fg/bg. `displayName` changes user-facing browser branding only and must NEVER rename npm package, CLI, API, storage, CSS, or protocol identifiers. `language` (`en`/`zh-CN`) keeps English as the canonical source so live switching stays reversible. User display names flow through `textContent`/attribute APIs and the server title's HTML escaper, never `innerHTML`. → [architecture-invariants#theme-skins](docs/architecture-invariants.md#theme-skins) **Foldable settings identity**: responsive layout is width-driven via `MobileDetection.getDeviceType()`, but the localStorage namespace uses `MobileDetection.isHandheldDevice()` so an unfolded Android foldable keeps `codeman-app-settings-mobile`. ⚠️ Do not switch per-device settings namespaces from instantaneous viewport width: a posture-triggered WebView reload would lose opt-in UI. Regression profile: `OPPO Find N5 (unfolded)` in `test/mobile/devices.ts`. → [architecture-invariants#foldable-settings-identity](docs/architecture-invariants.md#foldable-settings-identity) **Folding devices: a fold is not a keyboard, and dialogs avoid the hinge** (Apple's [Designing for iPhone Duo](https://developer.apple.com/design/human-interface-guidelines/designing-for-iphone-duo)): ⚠️ **A visual-viewport resize that changes the WIDTH is the device changing shape** (a rotation, or a foldable opening or closing) **and is never the virtual keyboard**, which only ever takes height. `handleViewportResize()` read any height drop over 150px as the keyboard appearing, so closing an iPhone Duo (626→466pt wide, 890→678pt tall) latched `keyboardVisible` with no keyboard on screen: the accessory bar appeared, `main` grew 84px of dead padding, and `updateAppHeight()` (which bails while the keyboard is up) stopped refreshing `--app-height`. The latch is STICKY, since clearing it needs the height back within 100px of a baseline belonging to a display the user is no longer looking at, so it survived until the device was opened again, and rotating any phone hit it too. The shape branch re-baselines instead, which is also what lets a keyboard opened AFTER the fold be detected. ⚠️ `init()` must seed `lastViewportWidth`, or the very first resize reads as a shape change and swallows a real keyboard. ⚠️ **The hinge is a reserved region.** The CSS Viewport Segments media features report two segments only while a foldable is actually bent, and `--fold-inline-end`/`--fold-block-end` (styles.css) measure the strip to keep clear: `0px` on everything else, so the rules are inert by construction rather than by a branch. Codeman's seven centred overlays are all `position: fixed; inset: 0` flex boxes, and each shrinks its CONTENT box with padding rather than the box itself, so the backdrop still covers the far side of the fold and still swallows taps there. ⚠️ Each rule RE-STATES the overlay's own gutter (a later `padding-right` longhand beats the earlier `padding` shorthand it composes with), and the palette needs the compound `.modal.command-palette-modal` because mobile.css loads later and pads it with a shorthand under 768px. `test/foldable-layout.test.ts` DERIVES the overlay list from the stylesheet and compares both numbers, so a new overlay or a moved gutter fails there instead of on hardware nobody has. ⚠️ Physical sides, not logical ones: dialogs sit in the LEFT segment (the TOP one in tabletop pose) in every language, because the HIG keeps Duo's side controls on the same physical edge in RTL. ⚠️ **With the keyboard up, a shape change baselines to `window.innerHeight`** (the layout viewport, keyboard-free on both engines), never the shrunk visual height: the shrunk baseline made the settle event that follows every rotation or fold read as the keyboard closing, and the layout could not recover, since no further drop could re-arm the show branch. ⚠️ **A base gutter that a LATER `@media` block overrides needs its own fold restatement in that block, on a ZERO base**: the phone-width path picker and path preview drop to `padding: 0` under 600px, and the unconditional fold rules at the end of the file put 16px and 18px back on every phone (measured at 393 and 500). The palette's compound rule lives INSIDE the 600-768px band whose mobile.css shorthand it composes with, because outside it there is no side gutter and the addition pushed the shell 6px off centre; the response viewer's tabletop cap has a twin at the end of mobile.css, whose phone block (under 600px) otherwise outranks it. `test/foldable-layout.test.ts` simulates the cascade across BOTH stylesheets at every breakpoint, with and without the fold rules, so a moved gutter or an unscoped composition fails there rather than on hardware nobody has. Profiles: `iPhone Duo (outer)` / `iPhone Duo (inner)` in `test/mobile/devices.ts`; unit coverage in `test/viewport-shape-change.test.ts`. → [architecture-invariants#folding-devices](docs/architecture-invariants.md#folding-devices) **WebGL renderer toggle** (`webglRendererEnabled`, per-device): the GPU-stall watchdog's sticky `codeman-webgl-disabled` marker survives page loads and is cleared only by an explicit OFF→ON save or `?webgl=force`. `?nowebgl` forces the DOM renderer per-load. → [architecture-invariants#webgl-renderer-toggle](docs/architecture-invariants.md#webgl-renderer-toggle) **Shell keyboard accessory bar + one-shot Ctrl** (issue #262, `keyboard-accessory.js`): a **shell**-mode session automatically swaps the mobile accessory bar for terminal controls (Ctrl, Esc, Tab, four arrows, paste, dismiss); every other mode keeps the agent bar. `setMode()` now records the user's `extendedKeyboardBar` preference as the **base** layout and `refreshForActiveSession()` (called from `selectSession`) resolves base-vs-shell, so a settings save during a shell session cannot yank the bar away and switching back restores the user's choice. ⚠️ **Ctrl is a ONE-SHOT modifier applied in `terminal.onData`, not in a keydown handler**: a virtual keyboard emits no usable key events, so the character only exists as onData text. The hook sits AFTER `shouldSuppressTerminalQueryResponse` (xterm answers DA/CPR through onData too, and one of those would silently spend the modifier) and BEFORE every send path, so the control byte follows the normal control-char route. ⚠️ **Not every onData chunk is a keystroke**, and the query filter is not enough on its own: xterm ALSO emits mouse and focus reports on its own initiative, so the hook skips them via `isTerminalFocusOrMouseReport()` (they still reach the PTY, they just don't count as the next key). The mouse half is live — a shell session keeps the NARROW strip, so mouse DECSETs reach the browser and one tap while vim/htop runs spent the armed modifier silently (measured). The focus half is defense in depth: `FOCUS_ESCAPE_FILTER` in `session.ts` strips `\x1b[?1004h` from every PTY read, so `sendFocusMode` never turns on today; if it ever did, the bar's own post-key refocus would emit `\x1b[I` and eat the modifier before the user typed. ⚠️ It must disarm on ALL of: use, second tap, any other accessory key, session switch, keyboard dismissal, and a layout swap; a modifier left armed turns the next innocent keystroke into a control byte. ⚠️ **onData is not the only input path** — with `cjkInputEnabled` on, the CJK textarea owns the keyboard (onData returns early for everything it swallows, and the focus router sends `terminal.focus()` there, which is where the bar refocuses after every key), so `_handleCjkInput()` applies the modifier too. It is that module's single choke point to the PTY, so one call covers typed characters, IME flushes, Enter, backspace and arrows. Without it an armed modifier could neither fire NOR be spent, and survived to a later keystroke. Mapping is `ctrlByteFor()` (`code & 0x1f` over @A-Z[\]^_ and a-z, plus Ctrl+Space=NUL / Ctrl+?=DEL); characters with no control equivalent pass through unchanged, like a hardware keyboard. ⚠️ The armed style is `.accessory-btn.accessory-btn-ctrl.armed` (0,3,0) in BOTH stylesheets, and it cannot outrank mobile.css's light-skin repaint at **(0,3,1)** (`:is()` inherits its most specific argument, and that list holds `.btn-toolbar.btn-shell`) — so that rule excludes the state by hand as `.accessory-btn:not(.armed)`. Without the exclusion the armed button renders identically to a resting one on all four light skins, which is worse than no armed style at all. **Mobile prompt composer** (PR #444, the first slice of #359, `keyboard-accessory.js`): the agent bars' Paste key is now **Compose**, a dialog with a native multiline textarea (autocorrect, autocapitalize, spellcheck) where Enter adds a line and only **Send** submits; the shell bar keeps the direct Paste dialog, since shell input is not an agent prompt. Opening it ADOPTS the whole editable terminal prompt (`_takePendingLocalEcho`): the local-echo overlay's pending text has never reached the PTY, but the flushed prefix has, so that prefix is erased with backspaces counted in CODE POINTS (`Array.from(text).length`; measured on Claude Code 2.1.278, `a` + emoji + `b` takes three, and the UTF-16 count sent four and ate the neighbour; `clearTerminalInput()` in terminal-ui.js moved with it). ⚠️ **Drafts are per-session and in memory only** (`_composerDrafts`, never persisted: prompts routinely carry secrets, and persisting them would need the 0600 treatment the intent store gets). Every non-Send exit (Cancel, backdrop, Escape, "Use terminal keyboard") leaves the taken text ONLY in the draft, with the dot on the key (`has-draft`, kept in step by `_syncComposerDraftIndicator`) as the signal that the terminal prompt is empty on purpose, and `_cleanupSessionData` discards the draft with the session. ⚠️ **Delivery is a hand-built bracketed-paste frame** (`\x1b[200~` + the text with newlines mapped to `\r` + `\x1b[201~`, byte-identical to what `terminal.paste()` would emit) through `_sendInputAsync` WITHOUT `{ useMux: true }`, plus a SEPARATE Enter 120 ms later WITH it (codex drops keys that share a PTY read with a bracketed paste). Not `terminal.paste()`, for two reasons: xterm's `bracketedPasteMode` mirror is false for every session after a tab switch or reload (`terminal.reset()` in the replay re-clones the DEC modes, the tmux capture carries no `?2004h`, and tmux never forwards the pane's DECSET to a client after attach), so a paste through xterm would go out unbracketed and the CLI would submit at the first `\r`; and xterm's onData is where the local-echo paste branch flushes pending overlay text AHEAD of the block, text the composer has already taken and erased from the PTY, so the frame goes straight to the wire with the composer as the prompt's only owner. The unconditional frame is safe for a CLI that never enabled DECSET 2004 because tmux does the gating (measured against a live pane: markers stripped for a `cat -v` pane, forwarded intact to a process that had emitted `?2004h`). ⚠️ The frame must never take the mux fallback: `TmuxManager.sendInput()` strips every `\r` and `\n`, which welds the lines together and submits them. ⚠️ The size guard sits on the `MAX_INPUT_LENGTH` boundary (64 KiB of UTF-16 code units): `ws-routes.ts` drops a longer frame WITHOUT an ACK, which would wedge the durable queue, so `_composerMaxLength` is derived from that limit minus both markers and an oversized prompt stays a draft with a toast. ⚠️ `.prompt-composer-overlay` is a `.paste-overlay` with a gutter of its own (a `padding` shorthand whose bottom is 12px plus the safe area), and the unconditional `.paste-overlay` fold rule at the end of styles.css is a later longhand at the same specificity, so it ERASED that gutter (measured at 393x852: `padding-bottom: 0px` flat, and the hinge strip REPLACING the gutter with the fold variables set): the composer has its own restatement after the fold rules, the dialog's `max-height` subtracts `--fold-block-end`, and `test/foldable-layout.test.ts` lists the composer in `ELEMENTS` by hand, because its derived overlay list keys on rules that declare `position: fixed; inset: 0` themselves. Tests: `test/mobile-prompt-composer.test.ts` (in the CI gate, deliberately not under `test/mobile/**`). **The PTY and the browser terminal must never disagree about size** (issue #464, `syncTerminalGeometry()` in terminal-ui.js): Claude Code's TUI wraps its frame at the width the PTY reported and erases the previous frame by walking the cursor up the number of rows it BELIEVES that frame occupied. A browser terminal of a different width makes each logical line take more physical rows than Ink counted, so `eraseLines(n)` clears too few and the new frame paints over rows nothing erased — the doubled lines and half-overwritten prose in #464, **measured** against this repo's xterm (a 120-column PTY against a 62-column terminal renders every wrapped line twice; `test/terminal-pty-geometry.test.ts` pins it, and pins the clean render at matching widths so the assertion cannot pass against code that fixes nothing). ⚠️ **`fitAddon.fit()` is NOT the way to resize this terminal.** It resizes xterm to `proposeDimensions()` RAW while every server-facing path reports those floored at 40x10, so whenever the floor bit the two diverged silently — **measured in Chrome at 430px**: font size 44 proposed 13 columns, the server was told 40, and xterm stayed at 13. `syncTerminalGeometry()` fits, floors and applies as one step and is the ONE function that may change the size; `test/terminal-pty-geometry.test.ts` sweeps every module for a bare fit. ⚠️ **Withhold the fit wherever you withhold the SIGWINCH.** `throttledResize` (virtual keyboard up) and `sendResize` (session detached into its own window) used to reflow locally and skip only the server write, which is the one combination that cannot be right. ⚠️ **A font change is a geometry change**: `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size and told the server nothing, so raising the font on a phone left the CLI wrapping at the old column count (`_refitAfterCellSizeChange`). ⚠️ **Resize is no longer write-only.** `Session.resize` DECLINES a small-viewport request while a desktop connection holds an active sizing claim and says nothing, so both transports now answer with `session.ptyCols`/`ptyRows` (`{"t":"zc"}` on the socket, the body of the resize POST) and `_onPtyGeometryReport` adopts them — a terminal that keeps a shape the PTY refused renders GARBLED, not merely wrong-sized. Adopting can leave the pane wider than the screen, and `.terminal-container` is `overflow: hidden`, so `.pty-oversized` grants horizontal reach for exactly as long as the mismatch lasts. ⚠️ That rule sets **both** overflow axes and its own `touch-action`: mobile.css loads later and sets `.terminal-container { overflow: visible; touch-action: none }`, and a bare `overflow-x` would leave overflow-y computing to `auto`, handing the browser a vertical scroll container the terminal's touch handler does not know about. **Verified in Chrome at 430px** against a live server with a desktop client holding the claim: the phone adopts 198x43, `overflow-x: auto` / `overflow-y: hidden` / `touch-action: pan-x`, the full pane width reachable by scrolling (how many pixels depends on the content, so no figure is pinned here), and the terminal's own vertical scroll still working under `pan-x`. Removing the class returns `scrollLeft` to 0, so a resolved mismatch cannot leave the pane parked off-screen. **Terminal resilience: replay clears, renderer liveness, fetch deadlines**: three rules that each close a way the terminal silently stops being correct. ⚠️ **A replay clear MUST be in-stream, never `reset()`/`clear()`.** xterm's `write()` is asynchronously queued while `Terminal.reset()` is synchronous and, per upstream, "does not clear input buffers and does not reset the parser" — so bytes queued just before a reset are parsed AFTER it and fuse into the snapshot written next. **Measured** against the real xterm in this repo: `write('p8'); reset(); write('rmissions')` renders `p8rmissions`; the queued `\x1bc` renders `rmissions` and clears scrollback. `_resetTerminalForReplay()` (app.js) is the ONE clear, a single queued `\x1bc` (RIS), and all three replay paths go through it; RIS rather than `\x1b[3J\x1b[H\x1b[2J` because the erase leaves modes, charsets, scroll regions and SGR state alone. Callers may still chunk the content — ordering in the queue is what matters, not writing it in one call. ⚠️ **The renderer watchdog reads xterm privates and CANNOT be covered by the gate.** `_kickRenderer()` (terminal-ui.js) cancels a stale `_core._renderService._renderDebouncer._animationFrame` and forces a repaint. **Verified against xterm 6.0.0** (jsdom, after `open()`): the field path resolves, a forced stale handle genuinely makes `refreshRows` a no-op, and the kick schedules a fresh frame. **Reasoned, not reproduced here**: the premise that iOS discards scheduled rAF callbacks when a PWA backgrounds, which is what leaves the handle stale — that half wants a real-device pass. Codeman has exactly ONE xterm for the whole page load, so one backgrounding would wedge it until a reload. `_renderService` only exists after `open()`, which needs a real DOM, and the gate runs in node — so `test/xterm-private-api.test.ts` pins the RESOLVED lockfile version (not the `^6.0.0` range, which a real upgrade slips through) and a bump means re-verifying by hand. Every access is optional-chained on purpose: a renamed field must degrade to a no-op, never throw on a 2s timer. ⚠️ **Every terminal capture carries a deadline, and the helper reads the BODY** (`_fetchTerminalCapture`, app.js). `await fetch()` settles on response HEADERS, so clearing the timer there leaves the body — the multi-megabyte `?full=1` capture this exists for — unbounded: **measured** at 4026ms under a 1000ms deadline before the fix. The helper therefore returns `{json, headers, headersAt}` rather than a `Response`, and `_terminalCaptureInflight` is scoped the same way so a body still streaming counts toward a capture starting beside it. It degrades to a plain fetch where `AbortController` is missing — the deadline is a safety net, not a dependency. Tests: `test/terminal-resilience.test.ts` (pure decisions), `test/xterm-private-api.test.ts`. **WebSocket output-gap reconcile** (`_wsOutputGapSession`, app.js): terminal OUTPUT frames carry no sequence number (input frames do — `seq`+`cid`, at-most-once, ACKed), so a dropped socket leaves a hole nothing replays. ⚠️ **The gap is narrower than "the device went offline"**: if the network drops, SSE drops with it and `handleInit`'s keepTerminal branch already calls `_onSessionNeedsRefresh`. The uncovered case is the WS dying while SSE stays up (half-open socket, proxy idle-timeout, ping timeout), because `_onSSETerminal` discards every SSE terminal frame while `_wsReady` is true and `_wsReady` only flips in `ws.onclose`. Reaching `onclose` at all means the drop was unintentional (`_disconnectWs` nulls the handler first), so the session is marked and the next successful open reconciles. ⚠️ **The marker must be cleared by every path that repaints that session's buffer, and ONLY once one actually has.** `_markTerminalBufferReconciled()` is called from `selectSession` after its load, from `_cleanupSessionData`, and from `_onSessionNeedsRefresh` **at the repaint itself, not in its `finally`** — clearing on every exit meant a reconcile that threw, or hit the fetch deadline (the flaky link the marker exists for), dropped the gap with nothing to retry it. `ws.onopen` no longer clears it up front either, so a reconcile that is skipped or fails is tried again on the next open; re-entry is safe because `_terminalRefreshOwner` makes a second reconcile for the same session a no-op. `selectSession` loads the buffer and only THEN calls `_connectWs`, so without that clear the socket opening afterwards replays the whole buffer a second time on top of the one just written. Sequencing the output frames is the real fix and is not done. This is reasoned from the code path, not observed on a device. **Service worker: precache and cache key are BUILD-GENERATED** (`sw.js` + `scripts/build.mjs`): the build content-hashes assets and rewrites two exact declarations in `sw.js` — `const BUILD_ID = 'dev';` and `const HASHED_ASSETS = [];`. ⚠️ **Each must appear exactly once or the build THROWS**, which is deliberate: the list used to be hand-maintained with PRE-hash names, so every entry 404'd in production and `cache.add().catch(() => {})` hid it (15 of 23 verified failing against a running instance). The dev literals are valid on their own, so dev serves an unrewritten worker with an empty precache. ⚠️ **`caches.match` must pass `ignoreSearch: true`**: `renderIndexHtml` runs `cacheBustAssets`, which appends `?v=<mtime>` to every same-origin `.js`/`.css` reference INCLUDING content-hashed names, so the page requests `/app.<hash>.js?v=<mtime>` while the cache holds `/app.<hash>.js`. Without it no precached entry is reachable and the install downloads ~1.3MB that can never be served — once per deploy, since `CACHE_NAME` now carries the build id. That per-build key is what makes `activate`'s cleanup actually delete anything; it used to be the constant `'codeman-v1'`, so assets from every past release accumulated forever. Contract pinned by `test/sw-precache-manifest.test.ts`, which PARSES the `HASHABLE` list out of `build.mjs` rather than copying it. **Dismissing the on-screen keyboard** (PRs #279/#280, `terminal-ui.js`): the terminal parks focus on a hidden textarea that nothing used to release, so TWO gestures now blur it, and they own different regions. **(1)** `_installMobileKeyboardDismiss()` — a document-level `touchend` that fires only while the terminal input actually holds focus, **never inside `#terminalContainer`** (tap classification owns that) and **never on a control** (`MOBILE_KEYBOARD_DISMISS_EXEMPT_SELECTOR`, matched with `closest()` so an icon inside a button counts). Session tabs are covered by the selector's `[tabindex]:not([tabindex="-1"])` arm, which is what stops a tab tap from blurring and then being re-focused by `selectSession()`. **(2)** In `_handleMobileTerminalTap`, a second tap on **inert `content`** (`startedWithTerminalFocus`) blurs instead of re-focusing. ⚠️ Scoped to `content` on purpose: the prompt row (`input`) keeps focus-then-position so a second tap still places the caret, and actionable rows blur earlier via `_isActionableMobileTerminalTap`. ⚠️ **A scroll ends in `touchend` too** — dismissing there closes the keyboard and drops the composer mid-read, so travel is tracked from `touchstart` and multi-touch is never a tap. Both classifiers MUST share one threshold: `initTerminal`'s `TAP_THRESHOLD` reads `MOBILE_KEYBOARD_DISMISS_TAP_SLOP`, since a gesture the terminal calls a scroll and the dismiss handler calls a tap is exactly that bug. ⚠️ **The gate excludes `test/mobile/**`, so CI cannot see the only test covering (1)** — run `npm run test:mobile -- test/mobile/keyboard.test.ts` by hand and diff the FAIL list against master. (Not `npm test --`: the gate's config excludes that path, so a file filter pointing into it matches nothing and exits green having run zero tests.) That blind spot is why merging the two PRs, which conflicted semantically but not textually, produced a red suite with two green CI checks. **Phone toolbar: Enter replaces Shell** (post-1.8.0): inside `@media (max-width: 599px)` `btn-shell` is `display:none` and `btn-enter` takes its slot (`order: 4`); starting a shell moved into the Run dropdown (`Terminal / Shell` → `setRunMode('shell')` → `run()` → `runShell()`, button label "Run SH"). `runMode` is `z.string().max(20)` server-side, so new modes need no schema change. Desktop and tablet keep the green Run Shell button unchanged. ⚠️ **`sendEnterKey()` MUST go through `terminal._core.coreService.triggerDataEvent('\r', true)`** — not `sendInput()`, and never a raw POST to `/api/sessions/:id/input`. `localEchoEnabled` defaults to `MobileDetection.isTouchDevice()`, so on every phone the characters you type are buffered in the `LocalEchoOverlay` and have **never reached the PTY**; the `onData` Enter branch in terminal-ui.js is what flushes `pendingText` first and only then sends `\r` (after an 80ms delay so text lands first). Sending a bare `\r` submits an empty line and strands the typed text on screen, so the button looks dead. Replaying the keypress reuses the overlay flush, the flushed-offset cleanup and the ordering instead of reimplementing them. `KeyboardAccessory.sendKey()` is for escape sequences (arrows/Esc) and is the WRONG template to copy for input. ⚠️ **Skin overrides outrank plain class rules.** `styles.css` nests its skin block inside `html:not([data-skin="og"]) { … }`, so a bare `.btn-toolbar` rule in there resolves to specificity **(0,2,1)** and beats a `.btn-toolbar.btn-x` rule **(0,2,0)** in `mobile.css` regardless of load order. Toolbar-button colors set from mobile.css therefore need `!important` — that is why mobile.css leans on it so heavily. Symptom: only your `!important` properties land and everything else silently renders in generic toolbar grey. **Connection-loss UI** (`computeConnectionLossUi()` in constants.js, writer `_updateConnectionLossUi()` in app.js): the service worker serves the cached app shell, so an unreachable server (phone off the tailnet, VPN down, server stopped) used to render a normal-looking empty dashboard whose only tell was the 8px header dot, which reads as "no sessions", not "no connection". Two surfaces now: a full-screen **overlay** while no server state has loaded this page load (nothing behind it is worth preserving), and a non-blocking **banner** once it has (the terminal scrollback stays readable). ⚠️ A **2.5s grace** is load-bearing: a COM deploy restarts the server and SSE is back in ~200ms, and a banner on every deploy trains the user to ignore it. `navigator.onLine === false` skips the grace, since that is never a blip. Retry re-arms SSE **and** the terminal WS (`planWsReconnect` can 'give-up', and the SSE backoff caps at 30s). **SSE staleness watchdog** (`computeSseStale()` in constants.js, `_checkSseStale()` + a 5s interval in app.js): an `EventSource` that stops delivering does not always error, so `onerror` never fires, the header dot stays green, and every SSE-driven surface (tab status dots, sessions created on another device, renames) freezes until the user reloads. ⚠️ The 15s server keepalive was an SSE **comment** (`:keepalive`), and comments are **invisible to `EventSource` by spec**, so there was nothing a client could observe: it is now the named `sse:heartbeat` event (`cleanupDeadClients()`, sse-stream-manager.ts), which is exactly why the frame had to change type. ⚠️ Staleness is judged **only while the status is `connected`** and the device is online; that guard is the loop breaker, since a forced `connectSSE()` leaves `connected` immediately and cannot re-fire while a reconnect is in flight. ⚠️ The liveness stamp is applied inside `addListener` itself, so every registered handler (the `_SSE_HANDLER_MAP` wrappers AND the directly-registered ones) feeds it from one place; the heartbeat's own listener is a no-op that exists **only** to be registered, since `EventSource` drops named events nobody listens for. ⚠️ The watchdog interval is cleared at the top of `connectSSE()` and nowhere else (its only teardown path); clearing it elsewhere stacks intervals. Recovery needs no new sync path: the reconnect re-runs `handleInit` → `_resetAllAppState()`. The forced reconnect logs one diagnostic line, because a middlebox that strips heartbeats presents as "silently reconnects every 45s". **Z-index layers**: subagent windows (1000), split picker menu (1000, `.split-picker-menu`), plan agents (1100), mobile/tablet fixed header (1200, `mobile.css`), modals on ≤768px (1300 — must beat the fixed header or the modal close button is buried), log viewers (2000), connection-loss overlay (2500, above the fixed header and modals), image popups (3000), response viewer (5000, backdrop 4999), file-preview overlay (5100 — must outrank the response viewer, which can launch it; at its old 2000 a path clicked in the chat opened BEHIND the chat), toasts/path picker (10000+, deliberately above the preview), the custom-model center-status banner (10001, `.center-status-banner` — `[hidden]` must re-assert `display: none` over its own `display: flex`, same trap as `.home-sessions[hidden]`, or `dismiss()` leaves an invisible click-blocker dead centre on screen), the swap-confirm and context-warning modals (10010, `#customModelSwapConfirmModal`/`#customModelContextWarningModal` — must clear both the plain `.modal` z-index of 1000 and the center-status banner it can appear over), terminal touch-selection bar (900 — above terminal content and the local-echo overlay, deliberately BELOW floating agent windows so it can never cover their controls), local echo overlay (7). **Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min). **Keyboard shortcuts**: Escape (close), Ctrl+? (shortcut overlay), Ctrl/Cmd/Alt+K (session palette), Ctrl+W (kill), Ctrl+Tab (next), Alt+[/] (prev/next tab), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter or Ctrl+Enter (newline), Ctrl+C (copy selection, else interrupt) / Ctrl+Shift+C (copy, never interrupts), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font), Shift+Wheel (local scrollback when mouse passthrough is active), Shift+drag (start a selection in a stripped-DECSET pane, where xterm's own Shift branch is unreachable and a Shift+drag used to select nothing; `_installShiftDragSelection`), right-click (copy the selection, the mintty/PuTTY convention, since xterm paints into a canvas and the native menu has no Copy for it; with nothing selected the native menu is left alone). Rebindable via the registry. ### Security **Full model: [`docs/security-architecture.md`](docs/security-architecture.md)** (network binding, auth pipeline, the tunnel caveat, file-serving hardening, supply-chain, instance isolation, recommended setups). **Layer-by-layer detail with the history behind each: [architecture-invariants#security-layers](docs/architecture-invariants.md#security-layers).** | Layer | The rule | | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Auth** | Optional HTTP Basic via `CODEMAN_USERNAME` (default `admin`) / `CODEMAN_PASSWORD`. Active only when `CODEMAN_PASSWORD` is set (`middleware/auth.ts`) | | **Network bind** | Defaults to loopback. Non-loopback without a password starts but warns loudly. Classifier: `network-auth-policy.ts` | | **Host guard** | Always-on Host-header allowlist blocking DNS rebinding. ⚠️ **Custom reverse-proxy domains are rejected** unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix` | | **CSRF / Origin** | Always-on cross-site Origin guard on state-changing requests. **A missing Origin is allowed** so curl/CLI and hooks keep working. ⚠️ The body parser keeps `text/plain` RAW; auto-JSON-parsing it enabled simple-request CSRF | | **QR Auth** | Single-use 6-char tokens (60s TTL) for tunnel login. See `docs/qr-auth-plan.md` | | **Sessions** | 24h cookie (`codeman_session`), auto-extend, device context audit | | **Rate limit** | 10 failed auth/IP → 429 (15min decay). QR and hook-secret have separate buckets, so neither can lock out login | | **Hook bypass** | `/api/hook-event` + `/api/status-telemetry` skip Basic auth (localhost-only, schema-validated), but when auth is active the loopback bypass requires `X-Codeman-Hook-Secret` **unconditionally** (Codeman cannot detect a user's own loopback reverse proxy) | | **Lost-frame page** | The THIRD unauthenticated 200, beside the two hook routes, and the only one decided by request headers alone: a `GET`/`HEAD` carrying `Sec-Fetch-Dest: iframe\|frame`, `Accept: text/html` and mode `navigate` (or none), for a path that is NOT a registered route (never `/api/`, `/ws/`, `/q/`), is answered BEFORE the credential checks with the static web-tab recovery page (`lostWebviewFramePage`: no reflected input, `default-src 'none'` plus its own script hash, `no-store`). `/` is the one registered route also admitted, only when the request carries neither `codeman_session` nor `Authorization` (nothing in Codeman frames its own root; a sandboxed frame has neither), since the landing page masks to exactly `/` and its reload otherwise rendered Codeman inside the web tab. ⚠️ A non-browser client can set those headers, so an unauthenticated caller can tell a registered route (401) from a non-route (200) and enumerate the route table; accepted, the routes are public in `docs/api-reference.md`. Pinned by `test/webview-auth-exemption.test.ts` + `test/webview-lost-root-frame.test.ts` | | **Tunnel** | Enabling a tunnel **refuses** without `CODEMAN_PASSWORD` unless exposure is acknowledged via `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` or the per-request `acknowledgeUnauthTunnel:true` action field (never persisted) | | **Validation** | Zod schemas, Unicode-aware path allowlist regex, env prefix allowlist (`CLAUDE_CODE_*`/`OPENCODE_*`/`CODEX_*`/`GEMINI_*`/`GOOGLE_*`/`ANTIGRAVITY_*`/`PI_*`/`GROK_*`/`XAI_*`/`DSH_*`/`DEEPSEEK_*`) | | **Headers** | CORS localhost-only, CSP, X-Frame-Options, HSTS if HTTPS | **Security-relevant env vars**: `CODEMAN_MUX` (managed session), `CODEMAN_API_URL` (auto-set for hooks), `CODEMAN_ALLOWED_HOSTS` (extra Host/Origin allowlist entries for reverse proxies; bare `.suffix` matches subdomains), `CODEMAN_DOCKER_BRIDGE_HOOKS=1` (opt-in hooks-only listener on the docker bridge gateway). ### SSE Event Registry 161 event constants in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). **Both must be kept in sync**, and `test/sse-registry-parity.test.ts` is the guard that pins it (currently exactly in sync, 161 = 161, no drift either direction). ⚠️ `hook:agent_working` is the one hook event with no Claude Code hook behind it — the DeepSeek status bridge reports it (see External CLI modes). The backend file's `@fileoverview` carries the per-category breakdown, including the two Web tab events. ### API Routes ~236 handlers across 27 route files in `src/web/routes/`: system (56), sessions (37), cases (34), files (17), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), approvals (4), readmymind (4), custom-model (6), reboot-restore (3), me (2), teams (2), tab-layout (2), search (1), hooks (1), clipboard (1), status-telemetry (1), voice (1 + the `/ws/voice/stream` relay), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details. **HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`). ## Adding Features - **API endpoint**: Types in `src/types/` domain file, route in `src/web/routes/*-routes.ts`. Return the `ApiResponse` envelope (`{ success: true, data }`; errors via `createErrorResponse()` with proper status code). Validate with Zod schemas in `schemas.ts`. - **SSE event**: Add to `src/web/sse-events.ts` + `SSE_EVENTS` in `constants.js`, emit via `broadcast()`, handle in `app.js` (`addListener(`) - **Session setting**: Add to `SessionState`, include in `session.toState()`, call `persistSessionState()` - **App setting**: decide per-device vs synced first. Per-device keys go in the `displayKeys` set in settings-ui.js and must NOT be added to `SettingsUpdateSchema` (it is `.strict()`). ⚠️ Anything in `PUT /api/settings` that acts on a setting (the `toggleService` watcher calls) must resolve from **`merged`** (persisted + incoming), never from the raw request body: a partial PUT omits keys it doesn't intend to change, and `body.x ?? default` turns every omission into "apply the default" and silently resets live services. Pinned by `test/routes/system-routes-settings-partial-put.test.ts`. - **Hook event**: Add to `HookEventType`, add hook in `hooks-config.ts:generateHooksConfig()`, update `HookEventSchema` - **Mobile feature**: Add to relevant singleton, guard with `MobileDetection.isMobile()`. New header buttons must stay off phones (`test/mobile-header-buttons-policy.test.ts`). - **New test**: Pick unique port (search `const PORT =`). Route tests use `app.inject()` (no port needed) — see `test/routes/_route-test-utils.ts`. **Validation**: Zod v4 (different API from v3). Define schemas in `schemas.ts`, use `.parse()`/`.safeParse()`. ## State Files All in `~/.codeman/`: `state.json` (sessions, settings, respawn, orchestrator, cron jobs/runs, owner tab layouts), `mux-sessions.json` (tmux recovery), `settings.json` (user prefs), `push-keys.json` + `push-subscriptions.json`, `session-lifecycle.jsonl` (audit log), `update-status.json` (self-updater progress, polled across the service restart), `docker-env-applied.json` (Compose deployment only: sha256 of the Dockerfile + compose file the running container was built from, written by `Start-Codeman.sh`, read by the self-updater's environment gate), `docker-build-source.json` (Compose deployment only: the checkout's HEAD commit and `package-lock.json` hash the `codeman-node-modules`/`codeman-dist` volumes currently reflect, written by both `Start-Codeman.sh` and a successful in-place self-update, compared to detect and refresh a volume left stale by an externally-triggered rebuild), `linked-cases.json`, `webviews.json` (saved web-tab dashboard URLs), `remote-hosts.json` + `remote-cases.json`, `docker-hosts.json` + `docker-cases.json` + `docker-exports/`, `subagent-window-states.json` + `subagent-parents.json` (subagent window layout, GET/PUT `/api/subagent-window-states`/`-parents`), `hook-secret` (per-instance), `users.json` (multi-user, mode 0600) + `admin-audit.jsonl`, `intents.json` (Read My Mind intent profiles, mode 0600), `certs/` (self-signed TLS for `--https`), `.env` (CODEMAN_USERNAME/PASSWORD fallback for the `codeman attach` CLI), `install.log` (installer step output, written by `install.sh`'s `run_step`) and `tailscale-rename` (the node name before `install.sh` renamed it, so uninstall can offer it back; both installer-route only). Transient: `self-update-runner.sh`. Multi-user case spaces live OUTSIDE the data dir at `~/codeman-users/<username>/cases` (shared across instances like `~/codeman-cases`, override `CODEMAN_USER_SPACES_DIR`). **Generated top-level dirs** (all gitignored — don't edit or commit): `dist/` (esbuild output), `out/`, `coverage/`, `test-results/`, `tmp/`, `screenshots-echo-diag/`. The committed gesture bundle (`src/web/public/gesture/gesture-codeman.js`) IS tracked, but its runtime wasm/model assets (`src/web/public/gesture/wasm/`, `*.task`) are fetched and gitignored. ## Testing **`npm test` is the gate and is safe to run bare** — it runs `config/vitest.ci.config.ts`, exactly what CI runs, so local green means CI green. ```bash npm test # The gate — what CI runs npm test -- test/<specific-file>.test.ts # Single file npm test -- -t "pattern" # By name ``` Three suites are deliberately left out, because they cannot pass on an arbitrary machine. Each has its own runner, and a failure there means "not runnable here", not a regression: ```bash npm run test:browser # Playwright + chromium, live server; codex-predictive-echo also needs a real codex binary npm run test:mobile # the above plus environment-specific PNG baselines (own config, own pretest vendor step) npm run test:perf # wall-clock benchmarks — need an otherwise idle machine npm run test:all # literally everything; fails ~87 tests on a clean master here, which is why it is not the default ``` ⚠️ **`npm test` cannot see those suites**, so a change touching mobile/gesture/terminal-render behaviour needs the matching runner by hand — diff its FAIL list against master rather than reading it as pass/fail. That blind spot is what let two semantically-conflicting PRs merge green (see the on-screen-keyboard note above). ⚠️ **A file filter must match the runner.** `npm test -- test/mobile/keyboard.test.ts` matches nothing and exits GREEN having run zero tests, because the gate's config excludes that path — an excluded file needs its own runner (`npm run test:mobile -- <file>`, `npm run test:browser -- <file>`, `npm run test:perf -- <file>`). Vitest treats "no files matched a filter" as success, so read the file count, not just the colour. Raw `npx vitest` skips the config (and with it `setup.ts`); always use `npm test --` or pass `--config`. **Config**: Vitest with `globals: true`, `fileParallelism: false`. Timeout 30s, teardown 60s. `config/vitest.config.ts` is the everything-config behind `test:all`; `config/vitest.ci.config.ts` is the gate and derives its excludes from `config/test-suites.ts`, which is also what `vitest.browser.config.ts` and `vitest.perf.config.ts` derive their includes from — so the exclusions and the runners cannot drift apart. Keep shared options in sync across them. **Tmux safety**: under vitest (`VITEST` env var, set automatically), `TmuxManager` no-ops ALL shell commands and becomes a pure in-memory mock — tests physically cannot create/kill/attach real tmux sessions (`IS_TEST_MODE` in `src/tmux-manager.ts`). Every docker IO path is no-op'd the same way. `Session` is test-gated too: instead of attaching a real tmux client, it spawns a raw-mode echo PTY (`TEST_PTY_SCRIPT` in `src/session.ts`), so integration tests get a live input/output loop that echoes each byte exactly once. `test/setup.ts` gives every test file a temporary `HOME`/`USERPROFILE` (all `homedir()`-derived state, `~/.codeman` and `~/codeman-cases` included, resolves into a per-file fixture; the Playwright browser cache path is preserved), and additionally strips `CODEMAN_PASSWORD`/`CODEMAN_USERNAME` (so auth state from the running instance can't leak into tests) and `CODEMAN_GESTURE` (a shell-exported gesture flag would flip render-injection assertions), and strips the three instance-selection vars `CODEMAN_INSTANCE`/`CODEMAN_DATA_DIR`/`CODEMAN_TMUX_SOCKET` (#356/#371; `test/test-env-isolation.test.ts` pins the list, and its STATIC half reads setup.ts so a dropped `delete` fails everywhere rather than only on a box that exports the var). ⚠️ `CODEMAN_DATA_DIR` is the one that matters: `getDataDir()` reads it as an ABSOLUTE override before it ever looks at `homedir()`, so one inherited from the shell (a second instance, a beta run, a shell left over from `codeman web -d`) bypasses the temp HOME entirely, and a bare suite run once overwrote the real `remote-hosts.json` with a route test's fixture. `os.homedir()` itself DOES follow `$HOME`, so the temp HOME is what redirects everything else; `CODEMAN_INSTANCE` must be stripped in the setup file and never in a hook, because `config/instance.ts` captures it into a module-level const on first import. Tests that delete case trees go through `safeRmHomeTree()` (`test/mocks`), which refuses any path outside the temp HOME, so a wrong anchor leaves a temp dir behind instead of deleting `~/codeman-cases`. ⚠️ Raw `npx vitest` without `--config` skips `setup.ts` and with it the temp-HOME isolation. **Ports**: Pick unique ports manually, 3150+. Search `const PORT =` before adding new tests. Never 3000 (the live instance). ⚠️ **Browser tests can pass vacuously on mobile input paths.** Two traps, both hit on 2026-07-27 while fixing the phone Enter button: **(1)** driving input with `app.sendInput('…')` writes PAST the `LocalEchoOverlay`, so `pendingText` stays empty and any overlay bug is invisible — type with `page.keyboard.type()` instead; **(2)** headless Chromium reports `MobileDetection.isTouchDevice()` **false even with `hasTouch: true`**, so `_localEchoEnabled` is off and the local-echo branch never executes. Force it (`app._localEchoEnabled = true`) or the test proves nothing. Assert on real state (`app._localEchoOverlay.pendingText`, plus `tmux -L codeman capture-pane -p -t <pane>` for what actually reached the PTY), not on HTTP 200. **Testing against the live instance**: prod is HTTPS-only on :3000 (`curl -sk https://localhost:3000/...`). ⚠️ `w1`/`w2`/`w3` are the user's REAL sessions — never send input to them. Create your own throwaway session (`POST /api/sessions` then `POST /api/sessions/:id/shell`; creation alone leaves `pid: null` and no pane), test against that, and `DELETE` it by exact id when done. **Respawn tests**: Use `MockSession` from `test/mocks/index.ts` (defined in `test/mocks/mock-session.ts`). **Route tests**: `app.inject({ method, url, payload })` in `test/routes/` — no live port needed. **Mobile tests**: Playwright suite in `test/mobile/` (138 device profiles). Browser-testing infra and practices: `docs/browser-testing-guide.md`. ## Debugging ```bash tmux list-sessions # List tmux sessions curl localhost:3000/api/sessions | jq # Check sessions curl localhost:3000/api/status | jq # Full app state curl localhost:3000/api/subagents | jq # Background agents cat ~/.codeman/state.json | jq # Persisted state ``` Mobile screenshots: `~/.codeman/screenshots/`, accessed via `GET/POST /api/screenshots`. ## Performance & Limits Target: 20 sessions, 50 agent windows at 60fps. Limits live in `src/config/` (terminal 32MB, text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100), most env-overridable. Two constraints worth knowing before you touch them: the env-derived PTY buffer trim is **clamped to ≤75% of max**, because a trim ≥ max would disable `BufferAccumulator` trimming entirely and make memory unbounded; and browser xterm scrollback is a **separate hardcoded 50k** (`DEFAULT_SCROLLBACK` in constants.js), deliberately lower than tmux's 100k history because 100k per tab is a mobile-memory hazard. tmux <3.7 allocates `history-limit` at pane creation, while tmux 3.7+ can resize live panes (lowering the value can discard retained lines); already-evicted lines never return. The settings keys `terminalScrollbackLines`/`terminalBufferMaxBytes`/`terminalBufferTrimBytes` are schema-validated but **inert**; only `tmuxHistoryLimit` is wired. → [architecture-invariants#buffers-uploads-and-terminal-history](docs/architecture-invariants.md#buffers-uploads-and-terminal-history), `docs/terminal-anti-flicker.md` **Memory leaks (24+ hour sessions)**: use `CleanupManager`, clear Maps in `stop()`, guard async with `if (this.cleanup.isStopped) return`. Frontend: store handler refs, clean in `close*()`. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Verify: `npm test -- test/memory-leak-prevention.test.ts`. ## Scripts & Tunnel **`install.sh`** (repo root, ~140KB) is the public entry point: `curl -fsSL <raw url> | bash` installs Node/tmux/git/build tools if missing, clones to `~/.codeman/app`, builds, and offers a systemd/launchd service. Since installer v2 (2026-09-20) it is **look, ask, work, done**: `preflight_detect` prints what is on the machine, every human step runs BEFORE the build (one consent for all missing packages, ONE sudo prompt kept warm by `sudo_session_start`, the AI CLI menu, the Tailscale login/operator/HTTPS-toggle preflight), then the clone/`npm install`/build/service/serve run unattended behind `run_step` spinners (output in `~/.codeman/install.log`, tail shown on failure), and `print_done_screen` ends on the URL with a terminal QR code (the `qrcode` package Codeman already ships). The network-access question is 3-way: **Tailscale** (loopback bind + `tailscale serve`, curl-verified end-to-end), **LAN** (0.0.0.0 + password prompt), or **local-only**; it preserves the existing binding on re-runs via `read_existing_binding()`, which also reads back `CODEMAN_BASE_URL`/`CODEMAN_PORT` from the unit, and the flag/env PRESET paths keep an existing password too (`${CODEMAN_PASSWORD:-$EXISTING_PASSWORD}`: `--lan --service` on a password-protected unit used to rewrite it open with the unauthenticated ack, found in review 2026-09-21); `--password`/`--port` flip `RECONFIGURE` so they reach the unit instead of taking the quiet update path. ⚠️ The env a hand-started `codeman web` needs (host, password, ack, base URL, port) is composed in ONE place, `start_command_hint` for the done screen's Start line and `export_bind_env` for main()'s `exec` branch, so the two cannot disagree; `stop_background_helpers` runs right before that `exec`, because exec skips the EXIT trap and the sudo keepalive (keyed on `$$`, which becomes the server's pid) would otherwise refresh the sudo timestamp for the server's whole life. Ctrl+C in the HTTPS-toggle poll is trapped for the poll only and skips Tailscale for the run rather than killing the installer. ⚠️ **Tailscale is two halves on purpose**: `tailscale_prepare` (question phase: preflight, the opt-in rename, the serve SHAPE) and `tailscale_apply` (after the build: the one serve command). The shape is decided up front because it can change the service unit: when `:443` root already belongs to another app, the default is a **sub-path** (`tailscale serve --bg --set-path /codeman <port>` + `CODEMAN_BASE_URL=/codeman` in the unit; measured 2026-09-20: serve STRIPS the mount prefix before proxying, Codeman's `stripBasePath` tolerates unprefixed requests, and `--base-url` is what makes the emitted URLs carry it, verified live through the maintainer's tailnet incl. hashed assets and SSE), else a second port (8443+), replace, or skip. `detect_tailscale_serve_url` recognises all three shapes (`ts_serve_find_port_mapping`). ⚠️ **Rename is opt-in and defaults to NO everywhere** (owner decision 2026-09-20: the tailnet name is the machine's SSH identity); `--name`/`install.sh name` do it, `--yes` and non-interactive never do. Serve config is keyed by the DNS name it was written under, so `tailscale_rename_node` takes OUR mapping down first (`tailscale_remove_our_mapping`, `TS_MAPPING_REMOVED_BY_RENAME`) and it is re-added under the new name; the previous name is recorded in `~/.codeman/tailscale-rename` so uninstall can offer it back. Tailscale state is detected dynamically from `tailscale serve status --json` (no marker files); the installer must NEVER `tailscale serve reset`, touch a mapping it did not create, run `tailscale funnel` or advertise a Tailscale Service (a Service needs a TAGGED node + admin approval, so it is a docs hint only); `test/install-sh-invariants.test.ts` pins all of those plus the rename-before-shape order and the flag/header parity. A foreign `/Library/LaunchDaemons/com.codeman.web.plist` (the Mac mini's headless setup) is now LEFT ALONE rather than replaced by a LaunchAgent. Subcommands: `update`, `uninstall`, `tailscale` (retrofit), `name [<n>]`, `status` (the done screen again), `cloudflared` (the tunnel client, no longer a question in the main flow); flags `--tailscale|--lan|--local`, `--name|--no-rename`, `--service|--run|--no-start`, `--yes`, `--password`, `--port` pipe through `bash -s --` and set the same variables as their env twins (`CODEMAN_NONINTERACTIVE=1` approves system changes for automation and never installs Tailscale, renames or starts a service; `CODEMAN_TAILSCALE=1` presets the Tailscale choice). Design + verification record: `docs/installer-v2-plan.md`. Its CLI knowledge is a GENERATED block (`npm run generate:cli-catalog`, markers in the file), not a hand-written list: detection, the install menu and the closing reminder all read it, which is what stops the class of bug upstream `b6d0f1fa` fixed by hand (a user with only omp installed being told no AI CLI was found). ⚠️ It must stay **bash 3.2** clean — macOS ships it and the documented install is `curl | bash` under `set -euo pipefail`, so `declare -A`, `mapfile`, namerefs, `${x,,}` and here-strings are all fatal there; CI runs `bash -n` plus a real `bash:3.2` container, since expanding an EMPTY array under `set -u` is a runtime abort `bash -n` cannot see. ⚠️ It executes ONLY commands from the embedded block (`CLI_INSTALL_CMD_TRUSTED`) — there is no network fetch of the catalogue at install time to worry about at all. Other key scripts: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh [quick|named] start|stop|status|url` (quick = random trycloudflare URL, default; `named setup|enable` = fixed-hostname tunnel via `scripts/codeman-tunnel-named.service`; bare `start|stop|url` still means quick), `scripts/run-beta.sh` (isolated beta instance), `scripts/build-agent-image.mjs` (docker base image), `scripts/self-update.sh` (detached updater). Production services: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.