diff --git a/CLAUDE.md b/CLAUDE.md index 522490e2..b6a53c41 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -201,7 +201,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`. -⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer (`❯`) about once a second all through a turn, so the old "saw a ❯, wait 2s → idle" rule flipped every working session to idle two seconds in (measured: a session mid-tool-call at 17 minutes reporting `status:"idle"`). Its working indicator is `✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`: the glyph animates through `· ✢ ✳ ∗ ✻ ✽`, the gerund is randomized, and the finished line (`✻ Cooked for 2m 49s`) carries the same glyph, so neither `SPINNER_PATTERN` (braille, not what current versions draw) nor a keyword list can see it. Matching the new line in the STREAM does not work either: tmux ships partial repaints, so the whole line reaches the PTY only every few tens of seconds. So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks the SCREEN via `capturePaneText()` + `CLAUDE_WORKING_LINE_PATTERN` before believing it; a sustained run of repaints (`session-activity.ts`, pure + unit tested) is what marks a turn as started, with the same screen probe vetoing keystroke echo. Idle now lands ~3-5s after a turn ends instead of 2s into one. ⚠️ **The composer glyph and the working line are per-CLI registry DATA** (`capabilities.workDetect`, #385), not Claude constants: claude declares `❯` plus the pattern above, codex declares `›` plus `[Ee]sc to interrupt`, and a CLI that declares neither falls back to Claude's pair, which is what every session used before the registry carried one. Before that, this whole mechanism was gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` for its entire life. ⚠️ `workingLine` is config-supplied (a user `clis.json` can set it) and the compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()`: a nested quantifier there is a ReDoS against the event loop, and the helper returns null rather than throwing so the fallback is structural. +⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer (`❯`) about once a second all through a turn, so the old "saw a ❯, wait 2s → idle" rule flipped every working session to idle two seconds in (measured: a session mid-tool-call at 17 minutes reporting `status:"idle"`). Its working indicator is `✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`: the glyph animates through `· ✢ ✳ ∗ ✻ ✽`, the gerund is randomized, and the finished line (`✻ Cooked for 2m 49s`) carries the same glyph, so neither `SPINNER_PATTERN` (braille, not what current versions draw) nor a keyword list can see it. Matching the new line in the STREAM does not work either: tmux ships partial repaints, so the whole line reaches the PTY only every few tens of seconds. So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks the SCREEN via `capturePaneText()` + `CLAUDE_WORKING_LINE_PATTERN` before believing it; a sustained run of repaints (`session-activity.ts`, pure + unit tested) is what marks a turn as started, with the same screen probe vetoing keystroke echo. Idle now lands ~3-5s after a turn ends instead of 2s into one. ⚠️ **The composer glyph and the working line are per-CLI registry DATA** (`capabilities.workDetect`, #385), not Claude constants: claude declares `❯` plus the pattern above, codex declares `›` plus `[Ee]sc to interrupt`, and a CLI that declares neither falls back to Claude's pair, which is what every session used before the registry carried one. Before that, this whole mechanism was gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` for its entire life. ⚠️ `workingLine` is config-supplied (a user `clis.json` can set it) and the compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()`: a nested quantifier there is a ReDoS against the event loop, and the helper returns null rather than throwing so the fallback is structural. ⚠️ **A quiet pane is not always a pane that wants you.** An agent that arms a monitor, backgrounds a shell or hands work to a cloud session is told to end its turn, so Claude Code's `idle_prompt` notification lands a minute later on a session that wants nothing. The CLI states what it is still running on its own screen (optional `capabilities.workDetect.watchingLine`, with `watchingLines` for how far up the screen it sits), the idle probe reads it into `Session.watching`, and `notePrompt()` then opens that idle item ALREADY acknowledged so no surface alerts. ⚠️ **The fix is the alert that does not fire, not the `watching` badge**, and only `idle` is eligible — but a prose question is not a dialog, so a question asked in plain text while background work runs is silenced with the false alarms. ⚠️ **The label is pane-derived and therefore prompt-injectable**: the window must cover only rows the CLI draws and each pattern must anchor on chrome only that CLI can produce, or an agent silences its own alert by printing the words. → [architecture-invariants#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you](docs/architecture-invariants.md#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you). Tests: `test/session-watching.test.ts`, `test/watching-no-alert.test.ts`. **An exited agent in a live pane** (`paneExit`, Ark0N/Codeman#446): Codeman creates every pane with `remain-on-exit on`, so `/exit` ends the CLI while tmux keeps the pane, the tmux session and the `tmux attach-session` process Codeman records as `Session.pid`. No PTY exit handler fires, so a session whose agent is gone reads as a live idle one. `TmuxManager.startPaneExitWatcher()` reads `#{pane_dead}`/`#{pane_dead_status}`/`#{pane_dead_signal}` from one batched `list-panes -a` on its OWN always-on interval, and `SessionState.paneExit` rides the existing `session:updated` broadcast. ⚠️ **Never set `status: 'error'` for an exited pane** — that value belongs to the PTY-exit breaker and makes the browser offer a restart — and **never null the `pid`**, which is what makes `selectSession()` re-attach and launch a fresh CLI. ⚠️ **The field is TRI-STATE and its third state is absence**, meaning UNKNOWN, which must never render as alive; `Session.paneExitApplies` is the one place that scoping lives and it fails closed for a direct-PTY session, a remote SSH session, a docker case and a record rebuilt from the socket. ⚠️ **An absent `#{pane_dead_status}` is not 0** (measured on tmux 3.2a, a SIGKILLed pane reports neither a status nor a signal), so never write `status ?? 0`: absent-stays-absent is what will keep a later clean-exit sweep off crashed agents. ⚠️ **A path that starts a command in a pane must clear the record AND persist**, since the watcher's next tick sees the field already cleared and writes nothing. → [architecture-invariants#an-exited-agent-in-a-live-pane-paneexit](docs/architecture-invariants.md#an-exited-agent-in-a-live-pane-paneexit) @@ -225,7 +225,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Docker Compose deployment** (`docker/`, contributed): Codeman itself runs in a container and spawns Docker cases as **SIBLING** containers through the mounted host socket (Docker-outside-of-Docker), never nested. That inverts one assumption the bare-host path takes for granted: the daemon no longer shares Codeman's filesystem, so a bind source valid *inside* Codeman means nothing to it. `resolveDockerDaemonMountSource()` translates sources under HOME into the daemon's namespace via `CODEMAN_DOCKER_HOST_HOME`, and `CODEMAN_CASES_PATH` points the cases dir at a host-absolute bind mount so a workspace resolves to the SAME absolute path on both sides (which is what keeps the transcript projHash matching, per Docker cases above). ⚠️ **`CODEMAN_CASES_PATH` must move every consumer or none**: it is resolved once in `config/cases-dir.ts` because `src/cli.ts` resolves case paths too, and when only the server's `CASES_DIR` learned the override, `codeman skill install --case ` reported "Case not found" on exactly the deployment the override exists for. ⚠️ **`.dockerignore` patterns match the WHOLE context-relative path**, so a bare `.env` line excludes only the ROOT file: `docker/.env` (which holds `CODEMAN_PASSWORD` and any provider keys) rode `COPY . .` into the image until `**/.env` was added — verified in both directions with a real build context. ⚠️ A Compose LONG-form bind (`type: bind`) **creates a missing host source directory ROOT-OWNED** rather than refusing. `Start-Codeman.sh` pre-creates both `CODEMAN_APPDATA_PATH` and `CODEMAN_CASES_PATH` on the host before `up`, which is what keeps the daemon from ever having to materialise either as root in the first place; the container ALSO starts as root (`cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the base `cap_drop: ALL`; `test/docker-entrypoint.test.ts` pins that list) so `docker/entrypoint.sh` can correct a bind source that turns up root-owned anyway (a restored backup, a cleared directory, plain `docker compose up` run without the script) before dropping to `PUID:PGID` via `setpriv` — a directory owned by neither root nor `PUID:PGID` is never re-owned, since that ownership is not this container's to reassign; it is PROBED for writability as the runtime account (`setpriv ... test -w`, so ACLs, group-writable trees and CIFS/NFS mounts pass) and refused with a message naming path, owner and PUID:PGID if that fails. ⚠️ `KILL` is in that list for tini, not the entrypoint: `init: true` keeps tini as root while the server runs as PUID, and without CAP_KILL its SIGTERM forward fails and the server is SIGKILLed on every `compose down`/`restart` instead of flushing state. ⚠️ `/opt/codeman-cli` (the runtime-owned CLI prefix) is APPENDED to `PATH`, never prepended, and the entrypoint pins its own `PATH` to the system dirs: the root part of the start resolves `setpriv` by bare name, and a prefix ahead of `/usr/bin` let a planted `setpriv` run as uid 0 (measured). `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` drops `--memory-swap` (and filters only that one kernel warning) for hosts without swap accounting; `--memory` still applies. ⚠️ The deployment ALSO self-updates in place (the repo bind mount at `/opt/codeman` + a restart-by-exiting supervisor) — see Self-update below and `docs/docker-self-update.md` before touching `server.Dockerfile`, the compose file or `.env.example`, since each is an input to the updater's environment gate. `docs/docker-compose.md` + `docker/README.md` (user guides) -**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. **No code outside `stock.ts` may branch on a CLI id**; behaviour that genuinely differs is either a capability field or a NAMED PROFILE selected by one (`profiles.ts`), and `test/cli-registry-no-id-branching.test.ts` fails the build if an id check reappears — it matches `===`, `!==`, `case '':` and `[...].includes(mode)`, because an earlier `===`-only version let 36 negated branches survive the conversion (including a seven-mode Ralph chain whose own comment asked the next person to keep it in step with `isExternalCliMode()` by hand). A second CI-gated guard, `test/frontend-cli-no-id-branching.test.ts`, covers the two frontend files the run-menu consolidation (#458) touched, `session-ui.js` and `mobile-overview.js`: its allowlist is keyed by expression rather than line number, with an occurrence count per entry, so a new branch reusing an already-approved expression fails as a count mismatch instead of riding in on the old approval. ⚠️ Config contains no shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. ⚠️ `external`, `hooks` and `altScreen` are three INDEPENDENT capabilities on purpose; deriving one from another shipped the `until=stop`-hangs-on-shell bug. ⚠️ Two capability fields carry a REGEX from config (`discovery.version.regex` and `capabilities.workDetect.workingLine`) and both must compile through `compileVersionRegex()`, which caps length and refuses nested quantifiers; `workingLine` is the one that runs on the PTY hot path. ⚠️ **`param` is TWO namespaces.** `launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a LAUNCH PARAM; the legacy `Config` wire field is a separate namespace, bridged only by `launch.legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no load error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. Codex is the entry where the two names differ (`bypassApprovals` vs `dangerouslyBypassApprovals`) and therefore the one that catches a regression. ⚠️ Six fields are DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured, so re-measure before wiring one up; the list is pinned so it cannot quietly grow. Spawn commands are pinned as literal strings in `test/cli-registry-spawn-golden.test.ts`, remote/docker pane commands in `test/location-overlay-commands.test.ts`. ⚠️ That second golden no longer covers **remote claude or remote omp**: both now have their own arm in `buildRemoteLaunchCommand` (a `--session-id || --resume` pair, and `--continue`, so a respawn continues the same conversation) and never reach `defaultRemoteCommandForMode`, which is what that test asserts. Their real pins are `toContain` substrings in `test/tmux-manager.test.ts` and `test/remote-shared-sessions.test.ts`; changing either arm will NOT fail the golden. ⚠️ Anything reading the registry resolves it AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks) — a module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. `~/.codeman/clis.json` overrides any entry (read-only in this release; nothing writes it, so importing the registry has no filesystem side effects). → `docs/cli-registry.md` +**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. **No code outside `stock.ts` may branch on a CLI id**; behaviour that genuinely differs is either a capability field or a NAMED PROFILE selected by one (`profiles.ts`), and `test/cli-registry-no-id-branching.test.ts` fails the build if an id check reappears — it matches `===`, `!==`, `case '':` and `[...].includes(mode)`, because an earlier `===`-only version let 36 negated branches survive the conversion (including a seven-mode Ralph chain whose own comment asked the next person to keep it in step with `isExternalCliMode()` by hand). A second CI-gated guard, `test/frontend-cli-no-id-branching.test.ts`, covers the two frontend files the run-menu consolidation (#458) touched, `session-ui.js` and `mobile-overview.js`: its allowlist is keyed by expression rather than line number, with an occurrence count per entry, so a new branch reusing an already-approved expression fails as a count mismatch instead of riding in on the old approval. ⚠️ Config contains no shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. ⚠️ `external`, `hooks` and `altScreen` are three INDEPENDENT capabilities on purpose; deriving one from another shipped the `until=stop`-hangs-on-shell bug. ⚠️ Three capability fields carry a REGEX from config (`discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`) and all three must compile through `compileVersionRegex()`, which caps length and refuses nested quantifiers; `workingLine` is the one that runs on the PTY hot path. ⚠️ **`param` is TWO namespaces.** `launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a LAUNCH PARAM; the legacy `Config` wire field is a separate namespace, bridged only by `launch.legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no load error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. Codex is the entry where the two names differ (`bypassApprovals` vs `dangerouslyBypassApprovals`) and therefore the one that catches a regression. ⚠️ Six fields are DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured, so re-measure before wiring one up; the list is pinned so it cannot quietly grow. Spawn commands are pinned as literal strings in `test/cli-registry-spawn-golden.test.ts`, remote/docker pane commands in `test/location-overlay-commands.test.ts`. ⚠️ That second golden no longer covers **remote claude or remote omp**: both now have their own arm in `buildRemoteLaunchCommand` (a `--session-id || --resume` pair, and `--continue`, so a respawn continues the same conversation) and never reach `defaultRemoteCommandForMode`, which is what that test asserts. Their real pins are `toContain` substrings in `test/tmux-manager.test.ts` and `test/remote-shared-sessions.test.ts`; changing either arm will NOT fail the golden. ⚠️ Anything reading the registry resolves it AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks) — a module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. `~/.codeman/clis.json` overrides any entry (read-only in this release; nothing writes it, so importing the registry has no filesystem side effects). → `docs/cli-registry.md` **External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. ⚠️ **Work detection is no longer part of that gate**: it is per-CLI `capabilities.workDetect` data (see the ❯ note above), so a CLI that declares its own glyph and working line gets the same screen-probed idle confirmation claude gets, and one that declares neither keeps the output-stabilization behaviour. All eight **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. ⚠️ **Codex sessions use PREDICTIVE WRITE-THROUGH echo, never the buffer overlay** (`_localEchoPolicy` in `_updateLocalEchoState`, terminal-ui.js): codex's composer reacts per keystroke ("/" pops a live-filtering picker, arrows edit server-side state, the composer grows as it wraps), so buffer-until-Enter starved it into issues #218/#219/#220/#222 and stays disabled (`_localEchoEnabled` remains false for codex). Instead, `PredictiveEchoAddon` (separate `vendor/xterm-predictive-echo.js` bundle) paints each keystroke at the predicted cell while the wire path stays BYTE-IDENTICAL: the onData hook (`_predictHookOnData`) is a plain statement with no `return`, so control always falls through into the untouched send path — pinned by vm and E2E byte-identity tests. Predictions reconcile against the parsed buffer and only while the cursor sits on the measured composer row (`isCodexComposerRow`, `/^› /`). Codex also **drops keystrokes that share a PTY read with a bracketed paste**, so flushed text and the paste sequence must go out as separate delayed writes (mirroring the Enter branch's delayed `\r`). Tests: `test/local-echo-codex-gating.test.ts`, `test/codex-predictive-echo.test.ts` (E2E vs real codex), `packages/xterm-zerolag-input/test/codex-replay.test.ts`. ⚠️ **Pi is the opposite kind of CLI and needs the opposite instincts**: it has NO permission prompts and no sandbox, so there is no bypass flag to send and Codeman must not invent one; its privileged knob is the tri-state `approveProjectTrust` (`--approve`/`--no-approve`), which makes pi EXECUTE repo-local `.pi/extensions` TypeScript, so the multi-user clamp puts pi in the **materialize** branch (an absent config still yields `--no-approve` for a non-granted owner) and `--api-key` is never wired. Pi stays OUT of `isAltScreenStripMode()` (main-screen TUI, and its 0.84.0 fullscreen mode is runtime-switchable via `/settings`, where the alt screen is load-bearing), and lands on the `'buffer'` echo policy via the `_updateLocalEchoState` fallthrough. Pi's own tests: `test/pi-mode.test.ts`, `test/routes/external-cli-bypass-clamp.test.ts`; user guide `docs/pi-integration.md`. ⚠️ **Grok is codex-shaped on permissions but opencode-shaped on rendering**: its bypass switch is `alwaysApprove` (`--always-approve`, grok's `bypassPermissions` mode — the Run button sends it `true` like antigravity's, and the clamp's only-if-sent branch strips it for non-granted owners), while its fullscreen alt-screen TUI keeps it OUT of `isAltScreenStripMode()`; the resolver version-probes `grok --version` like pi's (npm squatters exist for the name — `GET /api/grok/status` surfaces path + version), and grok lands on the `'buffer'` echo policy via the fallthrough (UNMEASURED against a live authenticated session; if its composer turns out per-keystroke-reactive like codex, flip it to the `'off'` branch). Grok's own tests: `test/grok-mode.test.ts`, `test/grok-cli-resolver.test.ts`; user guide `docs/grok-integration.md`. ⚠️ **DeepSeek breaks three of this family's assumptions, so do not pattern-match it onto its siblings.** (1) The agent is a **PROFILE, not the binary**: `dsh` is a launcher over `$DSH_HOME/profiles/` and DeepSeek ships only `web`/`headless`/`base`, so the terminal front door is ALWAYS third-party and "installed" ≠ "runnable" — the Run button gates on `isDeepSeekRunnable()` (binary AND a pane-capable profile) while `isDeepSeekAvailable()` gates the "add a profile" affordance; a `web`/`headless` profile is refused at spawn because it cannot drive a pane. (2) The permission switch is the **`DSH_PERMISSION_MODE` env export, not a flag** (`read-only`/`workspace-write`/`danger-full-access`) — the harness has none, and this is the one legitimate exception to the effort-style env-var ban because it is read with `??` as a boot-time default, so it stays soft; absent = `workspace-write`, which asks, hence the only-if-sent clamp branch, clamping to `workspace-write` (never `read-only`, which would break the workspace). ⚠️ **That clamp needs a second half no other CLI needs**, because the switch is an env var and `DSH_*` is an allowlisted `envOverrides` prefix: `applyEnvOverrides()` runs AFTER `_configureCliEnv()` in tmux-manager, so a non-granted owner sending `DSH_PERMISSION_MODE` on the SAME request would land last and hand back exactly the privilege the config clamp removed. `clampEnvOverridesForOwner()` (session-routes.ts) DROPS `DSH_PERMISSION_MODE`, `DSH_HOME` and `DEEPSEEK_BASE_URL` for a non-granted owner (the last because `_configureCliEnv()` forwards the SERVER's own `DEEPSEEK_API_KEY` into the pane, so a redirected base URL would send it to a foreign host) (dropping falls through to what `_configureCliEnv()` exports, which is the clamped value); `DSH_HOME` is there because it points the launcher at a profile tree whose plugin code runs at BOOT, before any approval row applies. Every OTHER CLI's bypass is a command-line flag reachable only through its config, which is why the config clamp alone is the whole gate for them. (3) It is the **only non-claude mode that passes `hooksAvailableForMode()`**, and for it alone that predicate is a per-SESSION question rather than a per-mode one (`deepSeekConfig.statusReporting: false` disarms the bridge, so every call site passes `sessionHookOptions(session)`; answering from the mode there re-creates the infinite-wait-dressed-as-a-timeout the guard exists to prevent). It passes because the terminal front door reports idle/working/blocked to a supervisor over a generic env-gated contract and `deepseek-status-shim.ts` makes Codeman that supervisor — real `stop`/`blocked` signals, real Approvals Inbox items, plus the `agent_working` event that clears an alert answered in the terminal. ⚠️ The resolver needs the strictest identity probe of the family (`dsh --help` must say `DeepSeek Harness`) because Debian ships an unrelated `dsh` (dancer's shell) that would pass a version probe. Model is NOT a session field (it is a profile composition entry). ⚠️ `hooksAvailableForMode()` is about hook SIGNALS and is not a stand-in for "is this a claude session": Read My Mind and intent capture read Claude's own transcript and compare `mode === 'claude'` directly, because when `deepseek` earned a yes the shared predicate silently widened both to a mode with no transcript to read (pinned by a static check in `test/deepseek-mode.test.ts`). ⚠️ **It is also the only external CLI whose answers are READ FROM DISK rather than scraped off the pane**: `deepseek-transcript.ts` reads `$DSH_HOME/sessions///session.jsonl.zstd` and backs the `last-response` route for dsh, because the pane segmenter served dsh-TUI's ASCII-art SPLASH as the worker's answer (measured), which anything polling for a first answer reads as an answer. Three traps live in that file: dsh appends **one zstd FRAME per write** and Node's `zlib` zstd decoder stops at the first (a real 56-line transcript decoded as 1 line, so the module walks frame headers itself; a Node older than 22.15 has no zstd and falls back to the pane); every turn also records a **plugin-sourced `user/message`** (the runtime-context snapshot) that must not render as the user's words; and a failed `turn/end` is surfaced as `Turn error: …` rather than as an empty string that reads as "still thinking". ⚠️ Session→transcript pairing is by the header's own `cwd` plus a ±60 s boot window, never by reproducing dsh's directory mangling (which has already changed form once) — and NEVER by newest-mtime alone, which handed a fresh worker its predecessor's answer in the same case dir. DeepSeek's own tests: `test/deepseek-mode.test.ts`, `test/deepseek-cli-resolver.test.ts`, `test/deepseek-transcript.test.ts`; user guide `docs/deepseek-integration.md`. OMP (`omp`) needs no bypass flag (the CLI's own `~/.omp` config governs trust/model routing, defaulting to `tools.approvalMode: yolo`), so its registry entry declares `privilegedParams: []` and a launch spec that only ever passes `--model`/`--resume`/`--continue` — but the multi-user clamp is NOT a no-op for it: `OMP_*` is an allowlisted `envOverrides` prefix, and the two credential-resolution keys it admits, `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN`, are clamped in `clampEnvOverridesForOwner()` for a non-granted owner, the same shape as `DEEPSEEK_BASE_URL`. Separately, `PI_*` is already allowlisted (pi needs it) and omp reads several of its knobs too (`PI_CONFIG_DIR`, `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`, `PI_SUBPROCESS_CMD`, `PI_SHELL_PREFIX`) — a redirected `PI_CONFIG_DIR` moves the `~/.omp` tree `omp-session-resolver.ts`/`omp-transcript.ts` hardcode, silently breaking pinning/history; this is a known gap shared with pi, not fixed here. → [architecture-invariants#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index f35c624c..2e1ea17b 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -90,6 +90,14 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do **Signal availability is decided by MODE, not by `isExternalCliMode()`.** `stop` and `blocked` come from Claude Code hooks, so `hooksAvailableForMode()` is true only for `claude`: `shell` is not an external-CLI mode but installs no hooks either, so `until=stop` on a shell session is a guaranteed unresolvable wait dressed up as a timeout. The behavior split is deliberate and must survive refactors: an EXPLICIT request for an unavailable signal is a 400 naming the mode, while the DEFAULT set (`stop,idle,exit`) silently drops them and echoes the narrowed set back as `wait.until`, because omitting the parameter must never 400. Signal quality is not uniform either: `stop` is definitive (Claude Code says the turn is over), `idle` is inferred from output stabilization plus prompt detection and can flap mid-turn when a spinner pauses, which is why `stop` is the documented default to orchestrate on and `idle` is the fallback for sessions that emit no hooks. ⚠️ **`idle` being ACCEPTED for a mode does not mean it ever FIRES there.** `startShell()` emits exactly one `idle` on a 500ms readiness timer and nothing afterwards, so a shell session sits at `status:'idle'` no matter what its pane is doing; since send-and-wait and `fresh=1` both require a TRANSITION, both can only time out on a shell worker (measured: a default `wait` on `sleep 4` burned its full 25s). The ❯-prompt and spinner detection that drives the real `working`/`idle` cycle is Claude's output format, so hook-less modes synchronize with `wait-output` markers or `exit`, and the docs must say so rather than listing `idle` as "available" and letting the reader infer it is usable. ⚠️ **Signals are edge-triggered with no history**, and that is a real orchestration limit: a signal that fires while no waiter is registered is gone, unobservable by any later wait variant (`until=stop` after the turn ended just times out, `fresh` or not — measured, R2-A). Documented client patterns must therefore register the waiter before the event can fire (send-and-wait) or gather on latched `wait-output` markers with `from=buffer`; the skill's fan-out flow was rewritten accordingly, and "fire-and-forget N prompts, then gather signal-waits sequentially" must never be documented again. The durable fix, a server-side latched last-signal-per-turn, is deferred with Part 3 of `docs/agent-control-plan.md`. ⚠️ Relatedly, the route corrects liveness that `SessionStatus` cannot express: `currentSignalFor()` answers `exit` whenever `pid === null` (exited, detached, or created-and-never-started) **or the mux pane is dead**, because `Session` parks a DEAD PTY at `_status = 'idle'` and trusting the status would answer the default wait `{signal:'idle', immediate:true}` for a crashed worker while `until=exit` blocked forever on an event that already happened. ⚠️ **`pid` alone cannot carry liveness for a tmux-backed session**: that pid is the local `tmux attach` CLIENT, so a worker exiting inside its pane leaves `pane_dead=1` with the client alive and `pid` never goes null — the `pid === null` branch is unreachable in the normal configuration (unit tests exercise it because `MockSession` sets `pid` by hand; only a live instance showed the gap). Liveness is therefore probed at the mux layer: `workerIsDead()` consults `mux.isPaneDead(muxName)` with a ~750 ms per-pane cache, ONLY on blocking waits (measured: 0 tmux execs across 100 non-wait input POSTs — the browser hot path pays nothing), plus a refcounted 3 s `watchForDeadWorker` interval so a worker dying while a wait is parked resolves it in ~3 s instead of burning the timeout. The probe fails SAFE by construction ("cannot tell" is never "dead": non-mux sessions, a missing or throwing `isPaneDead`, all return false). On send-and-wait, a "successful" `send-keys` into a dead pane additionally overrides `delivered` to `false` and rolls the dedup seq back (`undoOnFailure`), because the bytes went nowhere and a retry against a restarted worker must not be refused as a duplicate. The cost is that a just-created session reads as `exit`, which the wire docs must spell out as "not started yet"; the fix lives at the route rather than in `signalForStatus()` because the registry deliberately holds no `Session` reference. ⚠️ Two `claude`-mode cases still lose hooks for reasons outside the registry: a Docker case cannot reach a loopback-bound Codeman without `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, and a remote-SSH case runs the agent on another host whose hooks may never reach this server. Bounds are env-overridable (`CODEMAN_WAIT_MAX_MS`, `CODEMAN_WAIT_DEFAULT_MS`, `CODEMAN_WAIT_MAX_PER_SESSION`, `CODEMAN_WAIT_MAX_PER_OWNER`, `CODEMAN_WAIT_MAX_TOTAL`, `CODEMAN_WAIT_BUFFER_SCAN_BYTES`) and each is clamped to a hard bound so a typo degrades to the default instead of disabling the protection; they are internal tuning knobs like the rest of `src/config/`, NOT part of the SemVer-covered env-var surface in `versioning-policy.md`. Tests: `test/session-wait-registry.test.ts`, `test/routes/session-wait-routes.test.ts`, `test/routes/session-wait-output-routes.test.ts`, `test/routes/session-input-wait.test.ts`. +### The watching signal (a quiet pane that is not waiting for you) + +**A session that armed a monitor, backgrounded a shell or started a background terminal ends its turn and goes quiet, and a minute later Claude Code's idle notification arrives.** Before this existed, that prompt became an approval item like any other, so every surface filed the session under NEEDS YOU with nothing for a human to answer. The CLI says which kind of quiet it is on its own screen, and reading that row is the whole mechanism: `capabilities.workDetect.watchingLine` (optional, per CLI) plus `watchingLines` (how many non-blank rows at the foot of the screen may hold it, default `WATCHING_TAIL_LINES` = 1). `_confirmIdle()` already captures the pane at the moment a turn ends, so `_readWatching()` runs `watchingLabel()` (pure, `session-activity.ts`) over that same capture; the label lands on `Session.watching` and rides `toLightDetailedState()` out to every payload. ⚠️ It is cached BESIDE `_lastPaneProbeWorking` and goes stale with it, because the probe returns its cached boolean without re-capturing inside `PANE_PROBE_MIN_INTERVAL_MS` and a label from a capture nobody took is a guess. ⚠️ It then FREEZES once `_confirmIdle()` concludes — nothing looks at the pane again until it produces output — which is correct rather than tolerable, since work ending repaints the pane either way (a monitor firing wakes the agent; codex drops its background-terminal row by itself); a timer to keep it fresh would spend a `capture-pane` per idle session per tick to learn nothing. ⚠️ A server restart looks like a hole in that and is not one: the field is live state and starts empty, but reconciliation re-attaches the pane and the attach repaint arms the idle confirmation, which probes and re-reads the label with no input from anyone (measured 2026-09-23, back within ~20 s). A restored session showing no label has no chip on its screen. + +**The fix is the alert that does not fire; the badge is cosmetic.** `hook-event-routes` passes the label to `notePrompt()`, which opens the idle item ALREADY acknowledged (`acknowledgedAt` + `acknowledgedReason`). Nothing new suppresses anything: `acknowledge()` has always meant "the alert this prompt armed is spent", so the item stays pending, answerable and available as Read My Mind context, and a wrong label costs a card that does not blink rather than an alert that was never created. Every surface follows from that one flag — the broadcast carries `acknowledgedReason` so a live page declines to arm (`_onHookIdlePrompt`, settings-ui.js), the push is skipped, a reloading page reads `acknowledgedAt` in `seedApprovals()` as it always did, `classifySession()` and `pendingApprovalCount()` (tui-model.ts, tui-render.ts) ignore an acknowledged item, and the TUI card drops to the `info` tone and says why. It re-arms for free: the next idle prompt supersedes the item and is built fresh. ⚠️ Only `idle` is eligible, so a permission or question dialog still goes red whatever else the agent started — but a prose question is NOT a dialog, so an agent that arms a monitor and then asks "which branch?" in plain text is silenced along with the false alarms. That is the accepted cost of the design and the reason the kind gate sits at the single place items are created. + +**⚠️ The label is pane-derived, so the window and the anchor are a trust boundary, not formatting.** An agent that gets its own text matched silences its own alert. Two things prevent it, and BOTH belong to whoever adds a pattern for a new CLI: the window must cover only rows that CLI draws, and the pattern must anchor on chrome only that CLI can produce. Claude satisfies both — its chip is the LAST row, so the default window of one row excludes even the status line directly above it, whose content comes from a `statusLine` command a bypassed session can write into its own `.claude/settings.json`. Codex does not: its row sits above the composer, and the slot it occupies holds the last row of the TRANSCRIPT whenever no terminal is running, so matching the complete row raises the bar without closing it. What contains that is `hooks: 'none'` — no hook event from a codex session reaches `notePrompt()`, so a forged label costs a wrong badge and nothing else, and a CLI that gains hook signals must not keep a pattern that soft. The label is also ANSI-stripped and capped (`MAX_WATCHING_LABEL_CHARS`) at the source, and every interpolation of it into markup goes through `escapeHtml()`, since a config-supplied capture group decides what it holds. Tests: `test/session-watching.test.ts` (the label and both CLI patterns), `test/watching-no-alert.test.ts` (the negative claim across all four surfaces). + ### Auto-resume on usage limit **Auto-resume on usage limit** ("token pause" control, opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit ("5-hour limit reached ∙ resets 8pm" and all 1.0.x–2.1.x variants), `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time from cleaned output; `SessionAutoOps` arms a timer for reset+2min, then sends Esc (dismisses the rate-limit dialog) + `continue`. Still-limited responses re-arm the loop (5-min retry on stale times); a `working` transition cancels it. Claude-mode only (detection rides `_processExpensiveParsers`). Persists/recovers via `SessionState.autoResumeEnabled`/`autoResumeAt`; respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected` — prevents `/clear` from wiping the paused conversation). Endpoint: `POST /api/sessions/:id/auto-resume`; SSE: `session:limitPauseScheduled`/`limitResume`/`limitResumeCancelled`. Tests: `test/usage-limit-patterns.test.ts`, `test/session-auto-resume.test.ts`. diff --git a/docs/cli-registry.md b/docs/cli-registry.md index 22168cea..dbeb24c6 100644 --- a/docs/cli-registry.md +++ b/docs/cli-registry.md @@ -36,7 +36,8 @@ interface CliEntry { launch: CliLaunch; // the structured argv template env: CliEnv; // exports, tmux setenv keys, the env-override allowlist capabilities: CliCapabilities; // what every call site reads instead of the id - // .workDetect?: { promptGlyph, workingLine } — how this CLI's pane shows work + // .workDetect?: { promptGlyph, workingLine, watchingLine?, watchingLines? } — how + // this CLI's pane shows work, and how it shows work it started in the background overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store } ``` @@ -45,10 +46,46 @@ interface CliEntry { ### Regexes that come from config -Two capability fields carry a regular expression an override file can set: `discovery.version.regex` and `capabilities.workDetect.workingLine`. Both go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing. +Three capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`. All three go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing. `workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused. +`watchingLine` reads a different row of the same screen. A CLI draws it while work the agent +itself started is still running — Claude prints `⏵⏵ bypass permissions on · 1 monitor · ← for +agents` while a monitor, a backgrounded shell or a cloud session is live. Codeman turns that +into `Session.watching`, and an idle prompt from such a session opens already acknowledged, +so a pane waiting for its own background work never raises an alert a human cannot answer. +Group 1 is the label, and a CLI that declares no pattern reports no background work. + +Two CLIs declare such a row today, and they put it in different places. Claude writes its +chip on the last row of the screen, so it keeps the default one-row window and anchors on +the `·` its footer joins items with. Codex pins +`1 background terminal running · /ps to view · /stop to close` ABOVE its composer, which +puts the row third from the bottom once the status line and the composer are counted, so its +entry declares `watchingLines: 3` and matches that row end to end. Both were measured +against live panes rather than read out of a binary, which is the standard for adding a +third. + +That label is the one value in the registry that an AGENT can influence, because it comes off +the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern +for a new CLI. `watchingLabel()` in `session-activity.ts` searches only the last few +non-blank rows, which should be the part of the screen the CLI draws rather than the agent, +and the pattern should anchor on chrome only that CLI can produce. Keep the window as small +as the layout allows, since every row it adds is another row the agent may be able to write. +The label is also ANSI-stripped and length-capped at the source, and every interpolation of +it into markup goes through `escapeHtml()`, since it ends up on a badge and in an approval +card. + +The two shipped entries do not sit equally well behind that rule, and the difference decides +what a pattern is allowed to do. Claude's chip is the last row, so its one-row window holds +nothing the agent can write — not even the status line above it, whose command a session +running with permissions bypassed can write into its own `.claude/settings.json`. Codex's row +shares its slot with the last row of the transcript whenever no terminal is running, so a +message ending in that exact line is matched. What keeps that harmless is `hooks: 'none'`: no +hook event from a codex session reaches the approvals inbox, so a forged label costs a wrong +badge and cannot silence an alert. Before giving a CLI both hook signals and a pattern, make +sure its row is one the agent cannot write. + ### Three capabilities that must stay independent `external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks. diff --git a/docs/wiki/Notifications-And-Approvals.md b/docs/wiki/Notifications-And-Approvals.md index c096cd50..41472d1a 100644 --- a/docs/wiki/Notifications-And-Approvals.md +++ b/docs/wiki/Notifications-And-Approvals.md @@ -103,6 +103,30 @@ locked phone and the agent continues. With the inbox off, the buttons are stripped from the notification payload entirely rather than being shown and failing. +## When a session is watching its own work + +An agent that starts a monitor, puts a shell in the background or hands a task to a cloud +session is told by its CLI to end the turn and wait to be notified. The pane then goes +quiet, and the CLI's idle notification arrives about a minute later — for a session that +wants nothing from you. + +Codeman reads what the CLI prints about its own background work and treats that prompt +differently. It raises no tab alert, no desktop notification and no push, the session stays +out of NEEDS YOU on every surface, and the row wears a blue **watching** badge instead. Hover +it, or read it on a phone through your screen reader, and it says what is running: "1 +monitor", "2 shells", "1 background terminal". + +The prompt itself is not thrown away. It sits in the Approvals drawer as an ordinary card, +still answerable, with a line reading "quiet, watching 1 monitor" where a card you had +already looked at would say nothing. The next time that session goes quiet for an ordinary +reason, it alerts you exactly as before. + +Two limits are worth knowing. A permission prompt or a question dialog still goes red +whatever else the agent started, because that one blocks it outright. A question asked in +plain prose is not a dialog, so an agent that starts a monitor and then writes "which branch +should I target?" is quiet along with the rest — check a watching session yourself if it has +been quiet longer than the work it is waiting for should take. + ## The phone overview On phones, tapping the "C" logo gives a session overview with **NEEDS YOU** first, then diff --git a/src/config/cli-registry/schema.ts b/src/config/cli-registry/schema.ts index 68dee0d8..9ff73a59 100644 --- a/src/config/cli-registry/schema.ts +++ b/src/config/cli-registry/schema.ts @@ -325,8 +325,29 @@ const capabilitiesSchema = z (src) => compileVersionRegex(src) !== null, 'workingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers' ), + // Same guard, same reasons: this one runs over the foot of a pane capture every + // time a session settles, and ~/.codeman/clis.json can set it. + watchingLine: z + .string() + .min(1) + .refine( + (src) => compileVersionRegex(src) !== null, + 'watchingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers' + ) + .optional(), + // Bounded hard: this is how far up the screen a config file may push the search, + // and every row it adds is one more row the agent itself may be able to write. + watchingLines: z.number().int().min(1).max(8).optional(), }) .strict() + // A window with nothing to search is a typo, not a configuration. Refused at LOAD + // time for the same reason `privilegedParams[].param` is checked against the params + // the entry declares: the failure is otherwise silent and looks like a feature that + // simply never fires. + .refine( + (v) => v.watchingLines === undefined || v.watchingLine !== undefined, + 'watchingLines has nothing to bound without a watchingLine' + ) .optional(), model: z .object({ source: z.enum(['flag', 'claude-settings-file', 'none']), param: z.string().optional() }) diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index c0b7b58d..1b96e457 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -220,6 +220,21 @@ const CLAUDE: CliEntry = { workDetect: { promptGlyph: '❯', workingLine: String.raw`…\s*\((?:\d+h\s+)?(?:\d+m\s+)?\d+s\b|esc to interrupt`, + // Claude prints what it started in the background on the footer row beneath its + // composer, as `⏵⏵ bypass permissions on · 1 monitor · ← for agents`. The labels are + // the CLI's own words for each kind of background task, and group 1 is the one + // Codeman badges the session with. Verified against a live 2.1.278 pane on + // 2026-09-21. + // ⚠️ Two things keep an agent from writing its own label here, and both matter. + // The footer is the LAST row, so the default one-row window (`WATCHING_TAIL_LINES`) + // holds nothing but Ink's own chrome — in particular it leaves out the status line + // directly above, whose content comes from a `statusLine` command a bypassed + // session can write into its own `.claude/settings.json`. And the leading `·` keeps + // the match on the footer's own item list rather than on any text that happens to + // carry a count. A footer that ever drew the chip as its only item would report no + // watching rather than open that door. See `watchingLabel()` in + // `session-activity.ts`. + watchingLine: String.raw`·\s*(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?|Artifact comment monitors?))`, }, requiresMux: false, // Claude installs Codeman's own hooks block into every workspace it runs in, so its @@ -532,7 +547,37 @@ const CODEX: CliEntry = { // `Working (2m 49s • esc to interrupt)` above it while a turn runs. It animates no // braille spinner, and it never prints `esc to interrupt` at rest, so that phrase // alone separates a running turn from an idle one. - workDetect: { promptGlyph: '›', workingLine: '[Ee]sc to interrupt' }, + // Codex pins a row of its own while a background terminal it started is still + // running: ` 1 background terminal running · /ps to view · /stop to close`. Unlike + // Claude's footer chip that row sits ABOVE the composer, which puts it third from the + // bottom once the status line and the composer are counted, hence `watchingLines`. + // Measured against a live codex-cli 0.154.0 pane on 2026-09-22: the row appears when + // the terminal starts, follows the composer down as the conversation grows, and is + // gone after `/stop`. + // ⚠️ This entry CANNOT promise what Claude's does, and the difference is Codex's + // layout rather than its pattern. The third row from the bottom is the chip only + // while a terminal runs; with none running it is the last row of the transcript, + // which the agent writes. Matching the complete row raises the bar — an assistant + // message has to end with this exact line, to the character — but nothing here makes + // forging it impossible, so do not read the Claude comment above as applying here. + // What contains it is that codex declares `hooks: 'none'`: no hook event from a codex + // session ever reaches `notePrompt()`, so there is no idle item to pre-acknowledge + // and a forged label costs a wrong badge and nothing else. A CLI that gains hook + // signals must not keep a pattern this soft. + // ⚠️ Background TERMINALS are the only background work codex advertises on screen. + // A sub-agent started without waiting outlives the turn just as a terminal does — + // measured 2026-09-22, the sandboxed process was still running — and the pane shows + // nothing at all for it: the last rows are the composer and the status line, and + // `Sub-agents running` lives in the on-demand `/subagents` panel, not above the + // composer. So a codex session waiting on a sub-agent reads as plainly idle here. + // Nothing is misfiled by that (codex raises no idle prompts), and there is no row to + // match until codex pins one. + workDetect: { + promptGlyph: '›', + workingLine: '[Ee]sc to interrupt', + watchingLine: String.raw`^\s{0,4}(\d+ background terminals?) running · /ps to view · /stop to close$`, + watchingLines: 3, + }, // Two columns, like claude's, measured on a live 0.154.0 answer: the `•`/`›`/`⚠` // markers sit in the gutter, prose continuations sit at 2, and a nested YAML block // the model wrote rendered at 2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, diff --git a/src/config/cli-registry/types.ts b/src/config/cli-registry/types.ts index abe77677..34bb1851 100644 --- a/src/config/cli-registry/types.ts +++ b/src/config/cli-registry/types.ts @@ -344,6 +344,24 @@ export interface CliCapabilities { promptGlyph: string; /** Source of a regex matching the status line this CLI draws while a turn runs. */ workingLine: string; + /** + * Source of a regex matching the row this CLI draws while work it started in the + * background is still running, e.g. Claude's `· 1 monitor ·` footer chip or Codex's + * `1 background terminal running · /ps to view`. Capture group 1 is the label Codeman + * shows, and the whole match stands in when the pattern declares no group. A CLI that + * omits this reports no background work, which is what every CLI did before the field + * existed. + */ + watchingLine?: string; + /** + * How many rows at the FOOT of the screen that row can appear in, counting non-blank + * rows only. Claude writes its chip on the last row and keeps the default; Codex pins + * its own above the composer, which puts it third from the bottom, so it declares + * more. Keep each number as small as that CLI's layout allows: every extra row is + * another row an agent might be able to write, and the label is what silences an + * alert. See `watchingLabel()` in `session-activity.ts`. + */ + watchingLines?: number; }; /** * How many columns this CLI indents its transcript body by, so a copy taken from its diff --git a/src/session-activity.ts b/src/session-activity.ts index 324aff1b..c49cf4e3 100644 --- a/src/session-activity.ts +++ b/src/session-activity.ts @@ -21,6 +21,8 @@ * in 12/12 windows and the four idle ones in 0/12. */ +import { stripAnsi } from './utils/regex-patterns.js'; + /** * A gap longer than this ends a run of continuous output. Claude repaints at * least once a second while working, so this leaves generous headroom. @@ -91,3 +93,68 @@ export function isSustainedActivity(streak: ActivityStreak | null, streakMs: num export function isPaneQuiet(lastActivityAt: number, now: number, silenceMs: number = IDLE_SILENCE_MS): boolean { return now - lastActivityAt >= silenceMs; } + +/** + * How many rows at the foot of a pane capture may hold the background-work row, for a + * CLI that declares no number of its own (`capabilities.workDetect.watchingLines`). + * + * One, because the tightest window is the right default and Claude Code needs no more: + * it draws its chip on the LAST row of the screen. Blank rows are dropped before the + * window is taken, so a trailing blank costs nothing, and a CLI that ever prints a row + * BELOW its chip loses the badge rather than gaining a hole. + * + * ⚠️ The size of this window is a trust boundary, not a tidiness measure, and the row + * it excludes first is the one that taught us so: Claude's status line sits directly + * above the footer, its content comes from a `statusLine` command, and a session running + * with permissions bypassed can write that command into `.claude/settings.json` in its + * own workspace. A window of two therefore let an agent print `· 1 monitor ·` onto a row + * of its own and silence its own idle alert. Every row added here is another row + * somebody may be able to write, so widen this only for a CLI whose layout forces it, + * and never to a whole-pane search. + */ +export const WATCHING_TAIL_LINES = 1; + +/** Longest label a badge will carry. A footer chip is a handful of words. */ +export const MAX_WATCHING_LABEL_CHARS = 40; + +/** + * What a pane says is still running in the background, e.g. `1 monitor` or `2 shells`. + * + * The CLI writes that chip while a monitor, a backgrounded shell or a cloud session it + * started is still going, which is exactly the case where the agent has ended its turn + * without wanting anything from the user. `pattern` comes from the CLI's own registry + * entry (`capabilities.workDetect.watchingLine`); group 1 is the label when the pattern + * declares one, and the whole match stands in when it does not. + * + * Each candidate row is tested on its own, bottom row first, so a pattern can anchor + * itself with `^` or `$` against a single row rather than against a joined block. Blank + * rows are dropped before the window is taken, because a CLI that leaves a blank line + * between its chrome rows would otherwise spend the window on nothing. The answer is + * stripped of ANSI and capped, because it ends up on a badge and in an approval card. + * + * @param tailLines how many non-blank rows from the bottom to look at, defaulting to + * `WATCHING_TAIL_LINES`; a CLI declares its own when its row is not the last one + * @returns the label, or null when the pane shows no background work + */ +export function watchingLabel( + paneText: string | null | undefined, + pattern: RegExp, + tailLines: number = WATCHING_TAIL_LINES +): string | null { + if (!paneText) return null; + const lines = stripAnsi(paneText) + .split('\n') + .map((line) => line.trimEnd()) + .filter((line) => line !== ''); + for (const line of lines.slice(-Math.max(1, tailLines)).reverse()) { + // A pattern compiled by compileVersionRegex() never carries the `g` flag, but a + // caller reaching in from a test or a config reload might, and a stale lastIndex + // would make the same screen match every other call. + pattern.lastIndex = 0; + const match = pattern.exec(line); + if (!match) continue; + const label = (match[1] ?? match[0]).trim().slice(0, MAX_WATCHING_LABEL_CHARS); + if (label) return label; + } + return null; +} diff --git a/src/session.ts b/src/session.ts index 70ba4abb..ebed29ed 100644 --- a/src/session.ts +++ b/src/session.ts @@ -84,6 +84,8 @@ import { trackActivityStreak, isSustainedActivity, isPaneQuiet, + watchingLabel, + WATCHING_TAIL_LINES, IDLE_RECHECK_MS, PANE_PROBE_MIN_INTERVAL_MS, PANE_PROBE_RECHECK_MS, @@ -500,8 +502,35 @@ export class Session extends EventEmitter { private _activityStreak: ActivityStreak | null = null; // Unbroken run of PTY repaints (working detection) private _lastPaneProbeAt = 0; // Throttle for the tmux screen probe private _lastPaneProbeWorking: boolean | null = null; // Its last verdict (null = could not read) + /** + * Background work the pane's own footer reports, e.g. `1 monitor`; null for none. + * + * Cached BESIDE `_lastPaneProbeWorking` and refreshed only by a capture that really + * happened, so it goes stale exactly as that verdict does. The probe returns its + * cached boolean without re-capturing inside `PANE_PROBE_MIN_INTERVAL_MS`, and a + * label derived from a capture nobody took would be a guess wearing a fact's clothes. + * + * ⚠️ It then FREEZES once `_confirmIdle()` concludes: `activityTimeout` is null from + * there, and nothing looks at the pane again until it produces output. That is + * correct rather than merely tolerable, because work ending repaints the pane either + * way — a monitor firing wakes the agent, and codex drops its background-terminal row + * on its own. Do not add a timer to keep this fresh; it would spend a `capture-pane` + * per idle session per tick to learn nothing. + * + * A server restart is not a hole in that either, though it looks like one: this field + * is live state and starts empty. Reconciliation re-attaches the pane, the attach + * repaint carries the composer glyph, and the idle confirmation that arms on it probes + * and re-reads the label with no input from anyone — measured 2026-09-23 on a restarted + * instance, back within ~20 s for a session whose background terminal was still + * running. A session that comes back with no label has no chip on its screen. + */ + private _watching: string | null = null; /** Lazily compiled `capabilities.workDetect.workingLine`. See _workingLinePattern(). */ private _workingLineRe: RegExp | undefined = undefined; + /** Lazily compiled `capabilities.workDetect.watchingLine`. See _watchingLinePattern(). */ + private _watchingLineRe: RegExp | null | undefined = undefined; + /** Resolved with the pattern above: how many rows at the foot of the screen to search. */ + private _watchingWindow = WATCHING_TAIL_LINES; private _trustDialogAccepted: boolean = false; // Stops the trust-dialog scan (answered, or given up) private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read @@ -1218,6 +1247,16 @@ export class Session extends EventEmitter { return this._isWorking; } + /** + * What the pane says is still running in the background, e.g. `1 monitor`, or null when + * nothing is. A session with a label here has ended its turn without wanting anything + * from the user, so a surface that would otherwise file it under "needs you" can say + * what it is waiting for instead. + */ + get watching(): string | null { + return this._watching; + } + /** * Check if the session's process tree has active child processes beyond Claude itself. * Detects running bash tools, test suites, builds, servers, etc. that Claude spawned. @@ -1779,6 +1818,7 @@ export class Session extends EventEmitter { totalCost: this._totalCost, messageCount: this._messages.length, isWorking: this._isWorking, + watching: this._watching, lastPromptTime: this._lastPromptTime, // Buffer statistics for monitoring long-running sessions bufferStats: { @@ -2937,9 +2977,60 @@ export class Session extends EventEmitter { this._lastPaneProbeAt = now; const text = this._mux.capturePaneText?.(this._muxSession.muxName) ?? null; this._lastPaneProbeWorking = text === null ? null : this._workingLinePattern().test(text); + this._readWatching(text); return this._lastPaneProbeWorking; } + /** + * Read the background-work chip off the same capture the working probe just took. + * + * The two questions are different. A turn that is running is work the user is waiting + * for; a monitor, a backgrounded shell or a cloud session the agent started is work + * the AGENT is waiting for, and it is the reason a pane can sit at its composer with + * nothing to say and still not want anything from the user. `_confirmIdle` takes this + * capture at exactly the moment the turn ends, which is the moment the answer starts + * mattering. + * + * A capture that could not be read leaves the last answer standing, the way the + * working probe treats its own null: no evidence is not evidence of none. + */ + private _readWatching(paneText: string | null): void { + const pattern = this._watchingLinePattern(); + // Called only from the probe, and only with what a capture returned: `null` is + // "the screen could not be read", which is not evidence that nothing is running. + if (!pattern || paneText === null) return; + const label = watchingLabel(paneText, pattern, this._watchingWindow); + if (label === this._watching) return; + this._watching = label; + // ⚠️ This CHANGES while the session's status does not, so it needs an event of its + // own. The label is usually set on the idle transition, which broadcasts anyway, but + // it CLEARS when the work ends — and for a CLI whose background work ends without + // taking a turn (measured on codex: a background terminal finishing repaints the row + // away and nothing else happens) the session is idle before and after. Without this, + // the server knew the badge was gone and every open page went on drawing it until + // some unrelated event arrived. + this.emit('watchingChanged'); + } + + /** + * The regex matching this CLI's background-work chip, or null for a CLI whose registry + * entry declares none. Compiled once per session, like the working-line pattern, and + * null rather than a fallback: no other CLI has been measured drawing such a chip, and + * guessing one would badge sessions on the strength of an unread screen. + */ + private _watchingLinePattern(): RegExp | null { + if (this._watchingLineRe === undefined) { + // The pattern and the window it runs over are one decision, so they are resolved + // together: how far up the screen a CLI's row can sit is as much a property of its + // layout as the row itself. Claude writes on the last row and keeps the default, + // Codex pins one above its composer and declares more. + const detect = getCli(this.mode)?.capabilities.workDetect; + this._watchingLineRe = detect?.watchingLine ? compileVersionRegex(detect.watchingLine) : null; + this._watchingWindow = detect?.watchingLines ?? WATCHING_TAIL_LINES; + } + return this._watchingLineRe; + } + /** * The regex matching this CLI's "a turn is running" status line. * @@ -4191,6 +4282,9 @@ export class Session extends EventEmitter { async stop(killMux: boolean = true): Promise { // Set stopped flag first to prevent new timers from being created this._isStopped = true; + // A pane that is gone is watching nothing. Nothing probes a stopped session, so + // without this the last chip it drew would ride along on its row forever. + this._watching = null; this._clearAllTimers(); diff --git a/src/tui/tui-approvals.ts b/src/tui/tui-approvals.ts index 2cb10dea..000c6342 100644 --- a/src/tui/tui-approvals.ts +++ b/src/tui/tui-approvals.ts @@ -28,8 +28,11 @@ import type { ApprovalItem, ApprovalOption } from '../web/approval-inbox.js'; import type { TuiApprovalAnswer } from './tui-client.js'; -/** Card severity, in the same red/yellow vocabulary the web inbox uses. */ -export type TuiApprovalTone = 'err' | 'warn'; +/** + * Card severity, in the same red/yellow vocabulary the web inbox uses, plus the quiet + * third case: an item that opened acknowledged asks for nothing and reads grey. + */ +export type TuiApprovalTone = 'err' | 'warn' | 'info'; export interface TuiApprovalCard { tone: TuiApprovalTone; @@ -50,8 +53,14 @@ function clean(text: string | undefined): string { return (text ?? '').replace(/\s+/g, ' ').trim().slice(0, MAX_CARD_TEXT); } +/** + * How loud the card is. An idle prompt the inbox opened ALREADY acknowledged is not + * asking for anything — the session is watching work it started itself — so it drops to + * `info` and out of the warning vocabulary the other two share with the web inbox. + */ export function approvalTone(item: ApprovalItem): TuiApprovalTone { - return item.kind === 'idle' ? 'warn' : 'err'; + if (item.kind !== 'idle') return 'err'; + return item.acknowledgedReason ? 'info' : 'warn'; } /** @@ -65,9 +74,13 @@ export function approvalCard(item: ApprovalItem): TuiApprovalCard { const summary = clean(item.toolSummary) || clean(item.toolName); if (item.kind === 'idle') { + // Say what it is waiting for rather than asking for a reply, in the same words the + // web drawer uses for the same item. The prompt is still answerable, so the hint + // stays either way. + const quiet = clean(item.acknowledgedReason); return { - tone: 'warn', - title: message || 'waiting for your reply', + tone: approvalTone(item), + title: quiet ? `quiet, ${quiet}` : message || 'waiting for your reply', detail: [], options: [], hint: 'p to reply', diff --git a/src/tui/tui-model.ts b/src/tui/tui-model.ts index d8db248a..eae6dcb7 100644 --- a/src/tui/tui-model.ts +++ b/src/tui/tui-model.ts @@ -74,13 +74,25 @@ export function isLiveRow(session: TuiSessionRow): boolean { * outranks a stale `busy` status because the hook is the newer signal. An * errored session has no state of its own here and joins the waiting tier, * since it is equally something only a human can clear. + * + * ⚠️ An ACKNOWLEDGED item no longer decides the row. `acknowledgedAt` means the + * alert this prompt armed has been spent, either because somebody opened the + * session on another device or because the inbox opened the item that way for a + * session watching its own background work. The item itself stays pending and + * answerable, so the row keeps carrying it and the approval card still renders; + * it simply stops dragging the session into NEEDS YOU. The web has honoured + * that since acknowledgement existed (`approvals-ui.js` clears the pending hook + * that `_mobileOverviewState` reads), and this gate is where the TUI had been + * reading past it: acknowledging on a phone cleared the alert everywhere except + * here. Only `idle` can be acknowledged, so a permission or question dialog is + * unaffected by construction, and both are checked ahead of the flag anyway. */ export function classifySession(session: TuiSessionRow, approval?: ApprovalItem): TuiSessionState { if (!isLiveRow(session)) return 'recent'; if (approval) { if (approval.kind === 'permission') return 'blocked-permission'; if (approval.kind === 'question') return 'blocked-question'; - return 'waiting'; + if (!approval.acknowledgedAt) return 'waiting'; } if (session.status === 'error') return 'waiting'; if (session.isWorking === true || session.status === 'busy') return 'working'; @@ -96,7 +108,11 @@ export function classifySession(session: TuiSessionRow, approval?: ApprovalItem) * turn's own start is the pane's last Enter. */ export function stateSince(state: TuiSessionState, session: TuiSessionRow, approval?: ApprovalItem): number { - if (approval) return approval.createdAt; + // The prompt's own age measures the state only while the prompt is what put the + // row in that state. An acknowledged item still rides along on a row that is + // plainly idle or working, and dating such a row from it would report how long + // ago the prompt arrived as though it were how long the session has been quiet. + if (approval && STATE_GROUP[state] === 'needs-you') return approval.createdAt; if (state === 'working') return session.lastSubmitAt ?? session.createdAt ?? 0; return session.lastActivityAt ?? session.createdAt ?? 0; } diff --git a/src/tui/tui-render.ts b/src/tui/tui-render.ts index 8033104c..1893e59d 100644 --- a/src/tui/tui-render.ts +++ b/src/tui/tui-render.ts @@ -462,7 +462,8 @@ export function computeListWindow( * The pending dialog, drawn above the tail: the question, the parsed options * with their digits, and the keys that answer them. Red for a permission or * question prompt, yellow for an idle one, the same severity vocabulary the web - * inbox uses. + * inbox uses. An idle prompt that opened acknowledged carries neither: it reads + * grey with the idle glyph, because nothing about it wants the reader. */ export function renderApprovalCard( item: ApprovalItem, @@ -472,8 +473,8 @@ export function renderApprovalCard( ): string[] { const paint = painterFor(opts.color); const card = approvalCard(item); - const color = card.tone === 'err' ? SGR.red : SGR.yellow; - const glyph = card.tone === 'err' ? glyphs.blockedPermission : glyphs.waiting; + const color = card.tone === 'err' ? SGR.red : card.tone === 'warn' ? SGR.yellow : SGR.gray; + const glyph = card.tone === 'err' ? glyphs.blockedPermission : card.tone === 'warn' ? glyphs.waiting : glyphs.idle; const lines: string[] = []; const push = (text: string, style: string): void => { lines.push(padDisplay(paint(clipStyledLine(text, width), style), width)); @@ -571,10 +572,19 @@ function previewBody( // Chrome // ───────────────────────────────────────────────────────────────────────────── -/** Sessions with a prompt waiting on a human, which is what the badge counts. */ +/** + * Sessions with a prompt waiting on a human, which is what the badge counts. + * + * An ACKNOWLEDGED item is not one of them. Its alert has been spent, either by somebody + * opening the session elsewhere or because the inbox opened it that way for a session + * watching its own background work, and the row has already left NEEDS YOU by the same + * flag (`classifySession`). Counting it here would put a number in the header for a + * group the reader can see is empty. + */ export function pendingApprovalCount(model: TuiRenderModel): number { let count = 0; - for (const group of model.groups()) for (const row of group.rows) if (row.approval) count++; + for (const group of model.groups()) + for (const row of group.rows) if (row.approval && !row.approval.acknowledgedAt) count++; return count; } diff --git a/src/web/approval-inbox.ts b/src/web/approval-inbox.ts index 7712bb30..4330f42c 100644 --- a/src/web/approval-inbox.ts +++ b/src/web/approval-inbox.ts @@ -71,6 +71,15 @@ export interface ApprovalItem { * and reach the user's other devices. See `acknowledge()`. */ acknowledgedAt?: number; + /** + * Why the item arrived already acknowledged, for display only: the inbox + * writes `watching 1 monitor` for a session that went quiet because work it + * started itself is still running. A human acknowledgement leaves this unset, + * so a card can say "quiet, watching 1 monitor" rather than implying somebody + * looked. ⚠️ Pane-derived text, so it is bounded at the source and must not + * reach the DOM as markup — see `watchingLabel()` in `session-activity.ts`. + */ + acknowledgedReason?: string; /** * Present only when the frame parsed confidently. Gates which digits the * answer endpoint accepts; absent → only approve('1')/deny(Esc) are allowed. @@ -93,6 +102,12 @@ interface NotePromptArgs { toolSummary?: string; message?: string; cwd?: string; + /** + * What the session's pane says is still running in the background + * (`Session.watching`, e.g. `1 monitor`). An idle prompt from such a session + * opens ALREADY acknowledged: see `notePrompt()`. + */ + watching?: string | null; /** Returns the raw (ANSI-bearing) pane frame, or null when unavailable. */ capture?: () => string | null; } @@ -217,6 +232,22 @@ export class ApprovalInbox { * Record a prompt for a session, superseding any previous item, and return * the new item. Captures context immediately and once more after a short * delay (see RECAPTURE_DELAY_MS). + * + * ⚠️ An idle prompt from a session that is WATCHING its own background work + * opens already acknowledged (`args.watching`). Claude Code ends the turn + * after arming a monitor or backgrounding a shell and then reports the pane + * idle a minute later, so the alert that follows asks a human to look at a + * session that wants nothing from them. Acknowledging is deliberately what + * happens here rather than skipping the item: the prompt is real and stays + * pending, answerable and available as Read My Mind context, and only the + * alert it would have armed is spent. A wrong label therefore costs a card + * that does not blink, never an alert that was never created. + * + * It re-arms by itself. The next idle prompt supersedes this item and builds + * a fresh one, so once the background work ends and the session goes quiet + * for an ordinary reason, that item carries no acknowledgement and alerts + * normally. Only `idle` is eligible: a permission or question dialog blocks + * the agent whatever else it started, so its alert must survive. */ notePrompt(args: NotePromptArgs): ApprovalItem { this.resolveForSession(args.sessionId, 'superseded'); @@ -231,6 +262,10 @@ export class ApprovalInbox { message: args.message, cwd: args.cwd, }; + if (args.kind === 'idle' && args.watching) { + item.acknowledgedAt = item.createdAt; + item.acknowledgedReason = `watching ${args.watching}`; + } this.applyCapture(item, args.capture); this.items.set(args.sessionId, item); if (args.capture) this.captures.set(args.sessionId, args.capture); diff --git a/src/web/public/app.js b/src/web/public/app.js index 4762d978..ea54eed1 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -4623,6 +4623,11 @@ class CodemanApp { return { state, pill: this._sidebarRichPillLabel(state), + // What the pane's own footer says is still running in the background ("1 monitor", + // "2 shells"). A row that has one went quiet because the agent is waiting for that, + // which is a different thing from waiting for the user — so it rides BESIDE the + // state pill and never replaces it. + watching: typeof session.watching === 'string' ? session.watching : '', createdAt: Number(session.createdAt) || 0, since: this._mobileOverviewSince ? this._mobileOverviewSince(state, session) : null, }; @@ -4658,6 +4663,17 @@ class CodemanApp { parts.push(stamp(row.since.key, row.since.at, 'for', 'tab-meta-since')); } parts.push(`${escapeHtml(row.pill)}`); + // The word is duplicated from mobile-overview.js for the same reason the pill labels + // above are: it is one word, and this file must render a complete row even when a + // stale cached mobile-overview.js has arrived without it. + // The visible text is that constant. The pane-derived label appears only in the + // tooltip, where escapeHtml() (which escapes both quote characters) is what this file + // already relies on for every untrusted string it puts in an attribute, and where the + // source caps it at MAX_WATCHING_LABEL_CHARS before it ever gets here. + if (row.watching) { + const title = escapeHtml(`Still running in the background: ${row.watching}`); + parts.push(`watching`); + } // Both absolute stamps ALSO on the line itself, not only on the two items. // Below 288px the rail hides `.tab-meta-created` (the `tab-rail-tight` // rule), and a tooltip on a `display: none` element has no hover target — @@ -4695,7 +4711,11 @@ class CodemanApp { const prev = tab.dataset.tabState; // The since ANCHOR moves without the state changing (each new turn re-stamps // lastSubmitAt), so key the compare on both. - const sig = `${row.state}:${row.since ? row.since.at : 0}:${row.createdAt}`; + // Unescaped on purpose, and it still matches the attribute the initial render wrote: + // that one goes through escapeHtml() because it is interpolated into markup, and the + // browser hands the decoded string back through `dataset`. `watching` is the only + // pane-derived value in this signature, which is why it is the only one escaped there. + const sig = `${row.state}:${row.since ? row.since.at : 0}:${row.createdAt}:${row.watching}`; if (tab.dataset.tabMetaSig === sig) return; tab.dataset.tabMetaSig = sig; tab.dataset.tabState = row.state; @@ -5319,7 +5339,7 @@ class CodemanApp { const richMeta = this._sidebarRichMetaHTML(richRow); const richClass = richRow ? ` tab-state-${richRow.state}` : ''; const richData = richRow - ? ` data-tab-state="${richRow.state}" data-tab-meta-sig="${richRow.state}:${richRow.since ? richRow.since.at : 0}:${richRow.createdAt}"` + ? ` data-tab-state="${richRow.state}" data-tab-meta-sig="${richRow.state}:${richRow.since ? richRow.since.at : 0}:${richRow.createdAt}:${escapeHtml(richRow.watching)}"` : ''; // '' whenever the server said nothing about this pane's agent, which covers diff --git a/src/web/public/approvals-ui.js b/src/web/public/approvals-ui.js index 502cce7a..fc5746f5 100644 --- a/src/web/public/approvals-ui.js +++ b/src/web/public/approvals-ui.js @@ -222,6 +222,15 @@ Object.assign(CodemanApp.prototype, { return; } list.innerHTML = items.map((item) => this._approvalCardHtml(item)).join(''); + // The quiet reason is the only pane-derived string on a card, and it is the one + // an agent could write itself (it prints its own footer row), so it reaches the + // DOM as text and never as markup. The card leaves an empty span for it. + for (const item of items) { + if (!item.acknowledgedReason) continue; + const card = list.querySelector(`[data-approval-id="${CSS.escape(item.id)}"]`); + const slot = card && card.querySelector('.approval-quiet'); + if (slot) slot.textContent = 'quiet, ' + item.acknowledgedReason; + } }, _approvalCardHtml(item) { @@ -260,6 +269,9 @@ Object.assign(CodemanApp.prototype, { `${escapeHtml(item.sessionName || item.sessionId.slice(0, 8))}` + `${age}` + `` + + // Filled by renderApprovalsDrawer through textContent, never here: see the note + // there. An item a human acknowledged carries no reason and gets no line. + (item.acknowledgedReason ? `
` : '') + (summary ? `
${escapeHtml(summary)}
` : '') + (item.context ? `
${escapeHtml(item.context)}
` : '') + `
${actions}
` + diff --git a/src/web/public/home-sessions.js b/src/web/public/home-sessions.js index 453eec10..66430d9a 100644 --- a/src/web/public/home-sessions.js +++ b/src/web/public/home-sessions.js @@ -212,6 +212,9 @@ Object.assign(CodemanApp.prototype, { dir: this._shortenHomePath ? this._shortenHomePath(session.workingDir) : session.workingDir || '', state, pill: HOME_SESSIONS_PILL_LABEL[state] || state, + // What the pane's footer says is still running in the background, straight off + // the session payload. Same field, same meaning as on the phone overview. + watching: typeof session.watching === 'string' ? session.watching : '', // Epoch ms, straight off the session payload; formatting happens at // render time so the clock below can redo it without a re-render. createdAt: Number(session.createdAt) || 0, @@ -450,6 +453,12 @@ Object.assign(CodemanApp.prototype, { // what stops it ellipsizing. const meta = this._buildHomeSessionsMeta(row); meta.appendChild(pill); + // Built by the phone overview so both home screens word the badge identically. + // Guarded like every other cross-file call here: a stale cached mobile-overview.js + // must cost the badge, not the rail. + if (row.watching && typeof this._buildWatchingBadge === 'function') { + meta.appendChild(this._buildWatchingBadge(row.watching, 'home-sessions-pill')); + } item.appendChild(meta); return item; diff --git a/src/web/public/mobile-overview.js b/src/web/public/mobile-overview.js index 606bc1f1..940a87e8 100644 --- a/src/web/public/mobile-overview.js +++ b/src/web/public/mobile-overview.js @@ -61,6 +61,14 @@ const MOBILE_OVERVIEW_RUN_MODES = [ ]; /** Pill copy per state. Kept short: a phone row has ~90px for it. */ +/** + * The one word every surface puts on the watching badge, and the tooltip that says + * what the pane actually reported. Both live here so the phone overview, the desktop + * home rail and the rich sidebar rows cannot word the same badge three ways. + */ +const WATCHING_BADGE_TEXT = 'watching'; +const watchingBadgeTitle = (label) => 'Still running in the background: ' + label; + const MOBILE_OVERVIEW_PILL_LABEL = { needs: 'needs you', error: 'error', @@ -185,6 +193,10 @@ Object.assign(CodemanApp.prototype, { dir: this._shortenHomePath ? this._shortenHomePath(session.workingDir) : session.workingDir || '', state, pill: MOBILE_OVERVIEW_PILL_LABEL[state] || state, + // What the pane's own footer says is still running in the background ("1 monitor", + // "2 shells"), straight off the session payload. A row that has one is quiet + // because the agent is waiting for that, not because it is waiting for you. + watching: typeof session.watching === 'string' ? session.watching : '', // Epoch ms, straight off the session payload; formatting happens at // render time so the clock can redo it without a re-render. createdAt: Number(session.createdAt) || 0, @@ -713,6 +725,8 @@ Object.assign(CodemanApp.prototype, { pill.textContent = row.pill; item.appendChild(pill); + if (row.watching) item.appendChild(this._buildWatchingBadge(row.watching, 'mobile-overview-pill')); + const chevron = document.createElement('span'); chevron.className = 'mobile-overview-chevron'; chevron.setAttribute('aria-hidden', 'true'); @@ -734,6 +748,40 @@ Object.assign(CodemanApp.prototype, { return item; }, + // ═══════════════════════════════════════════════════════════════ + // Watching badge + // ═══════════════════════════════════════════════════════════════ + + /** + * The badge a session wears while work it started in the background is still + * running: a monitor, a backgrounded shell, a cloud session. + * + * It says one word, and the label the pane itself printed ("1 monitor") rides in + * the tooltip, because the badge shares a row with the state pill on the narrowest + * screen this app renders on. It does NOT replace that pill: an agent can arm a + * monitor and ask the user a question in the same breath, so the row still says + * "needs you" and this says what else is going on. + * + * Shared with the desktop home rail (home-sessions.js), for the same reason + * `_mobileOverviewState` is: one badge, one wording, one place to change it. The + * caller names its own pill class, because each surface styles its pills itself and + * the phone's live inside a media query the desktop never enters. + */ + _buildWatchingBadge(label, baseClass) { + const badge = document.createElement('span'); + const base = baseClass || 'mobile-overview-pill'; + badge.className = base + ' ' + base + '--watching'; + badge.setAttribute('data-i18n-skip', ''); + badge.textContent = WATCHING_BADGE_TEXT; + // The label rides in BOTH, because a tooltip is desktop-only: a phone has no hover + // target, and a screen reader gets the one word either way. This is the surface the + // badge was built for first, so "watching" with no way to learn what would be the + // wrong place to save a line. + badge.title = watchingBadgeTitle(label); + badge.setAttribute('aria-label', watchingBadgeTitle(label)); + return badge; + }, + // ═══════════════════════════════════════════════════════════════ // Age stamps: started / how long in this state // ═══════════════════════════════════════════════════════════════ diff --git a/src/web/public/mobile.css b/src/web/public/mobile.css index 5e2fafdd..3c528fed 100644 --- a/src/web/public/mobile.css +++ b/src/web/public/mobile.css @@ -2977,6 +2977,13 @@ html.mobile-init .file-browser-panel { color: var(--green); } + /* Accent, and none of the three above: a session watching work it started itself + is not asking the user for anything, and red and yellow are what say it is. */ + .mobile-overview-pill--watching { + border-color: var(--accent); + color: var(--accent); + } + .mobile-overview-chevron { flex-shrink: 0; color: var(--text-muted); diff --git a/src/web/public/settings-ui.js b/src/web/public/settings-ui.js index 2dfbcaf9..351c0b82 100644 --- a/src/web/public/settings-ui.js +++ b/src/web/public/settings-ui.js @@ -14,6 +14,14 @@ Object.assign(CodemanApp.prototype, { // Hooks (Claude Code hook events) _onHookIdlePrompt(data) { + // A prompt the server opened ALREADY acknowledged raises no alert here. Today that + // means the session is watching work it started itself (`acknowledgedReason` reads + // "watching 1 monitor"), so the pane is quiet because the agent is waiting for its + // own monitor, not for you. The item still exists and still shows in the drawer; + // only the tab alert and the desktop notification are declined. A page that reloads + // instead of receiving this event reaches the same conclusion from `acknowledgedAt` + // in seedApprovals (approvals-ui.js). + if (data.acknowledgedReason) return; // Always track pending hook - alert will show when switching away from session if (data.sessionId) { this.setPendingHook(data.sessionId, 'idle_prompt'); diff --git a/src/web/public/styles.css b/src/web/public/styles.css index abbe8f3c..494aa59d 100644 --- a/src/web/public/styles.css +++ b/src/web/public/styles.css @@ -12427,6 +12427,17 @@ kbd { color: var(--text-dim); font-size: 0.68rem; } +/* Why this card is not blinking at anyone: the session is watching work it + started itself. Accent, like the watching badge on a session row, and never + the red or yellow that mean a human is needed. */ +.approval-quiet { + margin-bottom: 6px; + color: var(--accent); + font-size: 0.68rem; + overflow: hidden; + text-overflow: ellipsis; + white-space: nowrap; +} .approval-summary { color: var(--text); font-size: 0.76rem; @@ -16314,6 +16325,17 @@ html[data-tab-orientation='vertical'] .home-sessions { color: color-mix(in srgb, var(--green) 45%, var(--text-muted)); } +/* Accent, deliberately none of the three above: a session that is watching something + it started (a monitor, a backgrounded shell, a cloud session) is not asking for + anything, so it must not borrow the red or the yellow that mean it is. This badge + sits BESIDE the state pill rather than replacing it, since an agent can arm a + monitor and ask a question in the same breath. */ +.home-sessions-pill--watching { + background: color-mix(in srgb, var(--accent) 14%, transparent); + border-color: color-mix(in srgb, var(--accent) 40%, transparent); + color: var(--accent); +} + /* Row accents: same language as the session tabs and the phone overview — red means a question is pending, yellow means it wants input, green means work is happening. Nothing else on this screen may reuse these colors. */ @@ -18376,6 +18398,16 @@ html[data-tab-orientation='vertical'][data-tab-rail-detail='rich']:not(.tab-rail color: color-mix(in srgb, var(--green) 45%, var(--text-muted)); } +/* Accent, deliberately none of the three above: a watching session is waiting for + work it started itself, not for the user, and the red and yellow here are spoken + for by sessions that ARE waiting for the user. */ +html[data-sidebar-detail="rich"] .session-sidebar .tab-pill--watching, +html[data-tab-orientation='vertical'][data-tab-rail-detail='rich']:not(.tab-rail-compact) .tab-rail .tab-pill--watching { + background: color-mix(in srgb, var(--accent) 14%, transparent); + border-color: color-mix(in srgb, var(--accent) 40%, transparent); + color: var(--accent); +} + /* Three lines of content per row instead of two, so give them room to breathe and stop the row actions crowding the pill. diff --git a/src/web/routes/hook-event-routes.ts b/src/web/routes/hook-event-routes.ts index fe96a0f7..e4d8cd65 100644 --- a/src/web/routes/hook-event-routes.ts +++ b/src/web/routes/hook-event-routes.ts @@ -176,6 +176,12 @@ export function registerHookEventRoutes( // identity beyond the shared per-instance secret, so a prompt claimed for a // session that can never show one must not create an answerable item). let approvalId: string | undefined; + // Set when the item opened ALREADY acknowledged, which today means the session is + // watching work it started itself. It rides the broadcast so a live page declines to + // arm the alert (a reloading page learns the same thing from `acknowledgedAt` when it + // seeds from /api/approvals), and it suppresses the push: an alert nobody can answer + // is worth even less on a phone than in a tab. + let acknowledgedReason: string | undefined; const approvalKind = APPROVAL_KIND_BY_EVENT[event]; if (session && hooksAvailableForMode(session.mode, sessionHookOptions(session))) { if (approvalKind) { @@ -194,6 +200,10 @@ export function registerHookEventRoutes( toolSummary: typeof toolSummary === 'string' ? toolSummary : undefined, message: typeof safeData.message === 'string' ? safeData.message : undefined, cwd: typeof safeData.cwd === 'string' ? safeData.cwd : undefined, + // What the pane says is still running in the background. An idle prompt from a + // session that is watching its own work opens acknowledged, so it never arms an + // alert nobody can answer; notePrompt() carries the whole reasoning. + watching: session.watching, // Visible tmux frame first (it IS the dialog); raw byte-buffer tail as // the fallback for direct-PTY sessions and the no-op test mux. capture: () => { @@ -203,6 +213,7 @@ export function registerHookEventRoutes( }, }); approvalId = item.id; + acknowledgedReason = item.acknowledgedReason; } else if (APPROVAL_RESOLVING_EVENTS.has(event)) { approvalInbox.resolveForSession(sessionId, 'resolved_in_terminal'); } @@ -213,6 +224,7 @@ export function registerHookEventRoutes( timestamp: Date.now(), ...safeData, ...(approvalId && { approvalId }), + ...(acknowledgedReason && { acknowledgedReason }), }); // Full state ride-along, same shape as the working/idle handlers: the home // screens rank the blocked group on lastActivityAt, and without this a @@ -224,12 +236,17 @@ export function registerHookEventRoutes( // on approvalId, and the answer route refuses keystrokes for dsh dialogs // (third-party TUI, unmeasured contract) — so a dsh push stays a plain // notification instead of offering buttons whose answer would be refused. - ctx.sendPushNotifications(`hook:${event}`, { - sessionId, - sessionName, - ...safeData, - ...(approvalId && session?.mode !== 'deepseek' && { approvalId }), - }); + // Nothing to push for a prompt that opened acknowledged: the agent is waiting for its + // own monitor or backgrounded shell, and a phone buzzing about it is the same false + // alarm as the tab alert, delivered where it is hardest to ignore. + if (!acknowledgedReason) { + ctx.sendPushNotifications(`hook:${event}`, { + sessionId, + sessionName, + ...safeData, + ...(approvalId && session?.mode !== 'deepseek' && { approvalId }), + }); + } // Track in run summary. `prompt_submitted` fires on EVERY prompt of every // Claude pane; only the ones where the conversation actually moved (a /clear diff --git a/src/web/session-listener-wiring.ts b/src/web/session-listener-wiring.ts index 6522c01a..55bb0152 100644 --- a/src/web/session-listener-wiring.ts +++ b/src/web/session-listener-wiring.ts @@ -43,6 +43,7 @@ export interface SessionListenerRefs { exit: (code: number | null) => void; working: () => void; idle: () => void; + watchingChanged: () => void; taskCreated: (task: BackgroundTask) => void; taskUpdated: (task: BackgroundTask) => void; taskCompleted: (task: BackgroundTask) => void; @@ -263,6 +264,17 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe } }, + /** + * Pushes the session state when `Session.watching` changes without the status + * changing with it. That is the badge appearing or, more often, going away: a CLI can + * finish its background work without taking a turn, so the row is idle before and + * after and no other broadcast fires. There is no SSE event of its own, because the + * badge reads off the session payload every surface already has. + */ + watchingChanged: () => { + deps.broadcastSessionStateDebounced(session.id); + }, + // ─── Background Task Events ────────────────────────────── /** Broadcasts `task:created` — new background task discovered */ @@ -495,6 +507,7 @@ export function attachSessionListeners(session: Session, refs: SessionListenerRe session.on('exit', refs.exit); session.on('working', refs.working); session.on('idle', refs.idle); + session.on('watchingChanged', refs.watchingChanged); session.on('taskCreated', refs.taskCreated); session.on('taskUpdated', refs.taskUpdated); session.on('taskCompleted', refs.taskCompleted); @@ -531,6 +544,7 @@ export function detachSessionListeners(session: Session, refs: SessionListenerRe session.off('exit', refs.exit); session.off('working', refs.working); session.off('idle', refs.idle); + session.off('watchingChanged', refs.watchingChanged); session.off('taskCreated', refs.taskCreated); session.off('taskUpdated', refs.taskUpdated); session.off('taskCompleted', refs.taskCompleted); diff --git a/test/approval-inbox.test.ts b/test/approval-inbox.test.ts index e5524159..42dcf615 100644 --- a/test/approval-inbox.test.ts +++ b/test/approval-inbox.test.ts @@ -330,6 +330,58 @@ describe('ApprovalInbox', () => { expect(inbox.getById(second.id)).toBeDefined(); }); + describe('a session watching its own background work', () => { + it('opens its idle prompt already acknowledged, and says why', () => { + const item = inbox.notePrompt({ sessionId: 's1', sessionName: 'w1', kind: 'idle', watching: '1 monitor' }); + expect(item.acknowledgedAt).toBe(item.createdAt); + expect(item.acknowledgedReason).toBe('watching 1 monitor'); + }); + + it('keeps the prompt pending and answerable: only its alert is spent', () => { + // Acknowledging rather than skipping creation is what makes a wrong label cheap. + // The prompt is real either way, and this way it is still in the drawer, still + // answerable and still Read My Mind context. + const item = inbox.notePrompt({ sessionId: 's1', sessionName: 'w1', kind: 'idle', watching: '2 shells' }); + expect(inbox.listPending().map((i) => i.id)).toEqual([item.id]); + expect(inbox.getForSession('s1')?.id).toBe(item.id); + expect(inbox.verifyStillAnswerable(item.id)).toBe(true); + }); + + it('never pre-acknowledges a dialog that blocks the agent', () => { + // A permission or question dialog blocks the turn whatever else the agent started, + // so watching says nothing about whether a human is needed. + for (const kind of ['permission', 'question'] as const) { + const item = inbox.notePrompt({ sessionId: `s-${kind}`, sessionName: 'w1', kind, watching: '1 monitor' }); + expect(item.acknowledgedAt).toBeUndefined(); + expect(item.acknowledgedReason).toBeUndefined(); + } + }); + + it('leaves an ordinary idle prompt alone', () => { + for (const watching of [undefined, null, '']) { + const item = inbox.notePrompt({ sessionId: 's1', sessionName: 'w1', kind: 'idle', watching }); + expect(item.acknowledgedAt).toBeUndefined(); + } + }); + + it('re-arms by itself once the background work is over', () => { + // The next prompt supersedes this one and is built fresh, so nothing has to + // remember to clear the flag when the monitor ends. + const watched = inbox.notePrompt({ sessionId: 's1', sessionName: 'w1', kind: 'idle', watching: '1 monitor' }); + expect(watched.acknowledgedAt).toBeDefined(); + const after = inbox.notePrompt({ sessionId: 's1', sessionName: 'w1', kind: 'idle' }); + expect(after.id).not.toBe(watched.id); + expect(after.acknowledgedAt).toBeUndefined(); + expect(after.acknowledgedReason).toBeUndefined(); + }); + + it('cannot be acknowledged a second time by a human opening the session', () => { + const item = inbox.notePrompt({ sessionId: 's1', sessionName: 'w1', kind: 'idle', watching: '1 monitor' }); + expect(inbox.acknowledge('s1')).toBeUndefined(); + expect(inbox.getById(item.id)?.acknowledgedReason).toBe('watching 1 monitor'); + }); + }); + describe('verifyStillAnswerable', () => { it('resolves the item and refuses when a parsed dialog left the screen', () => { const { resolved } = collect(inbox); diff --git a/test/cli-registry-schema.test.ts b/test/cli-registry-schema.test.ts index 762eff92..abd31ff3 100644 --- a/test/cli-registry-schema.test.ts +++ b/test/cli-registry-schema.test.ts @@ -122,6 +122,49 @@ describe('workDetect.workingLine is guarded like every other config regex', () = expect(compileVersionRegex(src), `${entry.id} declares a workingLine the guard refuses`).not.toBeNull(); } }); + + it('holds the optional watchingLine to the same guard', () => { + expectRejected((e) => { + (e.capabilities as Record).workDetect = { + promptGlyph: '>', + workingLine: 'working', + watchingLine: '(a+)+b', + }; + }, 'this one runs over a pane capture every time a session settles, so it can freeze the event loop the same way'); + }); + + it('bounds how far up the screen a config file may search', () => { + // The window is the injection guard: every row it adds is another row the agent + // itself may be able to write, and the label is what silences an idle alert. + for (const lines of [0, 9, 2.5]) { + expectRejected((e) => { + (e.capabilities as Record).workDetect = { + promptGlyph: '>', + workingLine: 'working', + watchingLine: 'chip (\\d+)', + watchingLines: lines, + }; + }, 'a config file must not be able to widen the search to the whole pane'); + } + }); + + it('refuses a window with no pattern to bound', () => { + expectRejected((e) => { + (e.capabilities as Record).workDetect = { + promptGlyph: '>', + workingLine: 'working', + watchingLines: 3, + }; + }, 'a window with nothing to search is a typo whose failure is otherwise silent'); + }); + + it('accepts every shipped watchingLine', () => { + for (const entry of STOCK_CLIS) { + const src = entry.capabilities.workDetect?.watchingLine; + if (!src) continue; + expect(compileVersionRegex(src), `${entry.id} declares a watchingLine the guard refuses`).not.toBeNull(); + } + }); }); describe('no shell text can reach the command line', () => { diff --git a/test/home-sessions.test.ts b/test/home-sessions.test.ts index f4bcccaf..67da3e55 100644 --- a/test/home-sessions.test.ts +++ b/test/home-sessions.test.ts @@ -356,3 +356,50 @@ describe('home screens: one order, one numbering', () => { expect(branch).not.toContain('this.sessionOrder[idx]'); }); }); + +describe('home sessions column: watching badge', () => { + it('carries the label off the session payload', () => { + const app = loadHomeSessionsApp({ + sessions: sessionMap([{ id: 'watcher', watching: '1 monitor' }, { id: 'plain' }]), + sessionOrder: ['watcher', 'plain'], + cases: CASES, + }); + + const rows = Object.fromEntries(app.buildHomeSessionRows().map((r: any) => [r.id, r.watching])); + expect(rows).toEqual({ watcher: '1 monitor', plain: '' }); + }); + + it("builds the badge with the phone overview, in the rail's own pill class", () => { + // Both home screens must word this badge identically, so the rail borrows the + // builder rather than writing a second one — the same arrangement it already has + // for state classification and for the stamp wording. + const app = loadHomeSessionsApp({ + sessions: sessionMap([{ id: 'watcher', watching: '2 shells' }]), + sessionOrder: ['watcher'], + cases: CASES, + }); + + const row = app._buildHomeSessionRow(app.buildHomeSessionRows()[0]); + const badges = collect(row).filter((el: any) => String(el.className).includes('--watching')); + expect(badges).toHaveLength(1); + expect(badges[0].className).toBe('home-sessions-pill home-sessions-pill--watching'); + expect(badges[0].textContent).toBe('watching'); + expect(badges[0].title).toBe('Still running in the background: 2 shells'); + }); + + it('draws no badge for a session running nothing in the background', () => { + const app = loadHomeSessionsApp({ + sessions: sessionMap([{ id: 'plain' }]), + sessionOrder: ['plain'], + cases: CASES, + }); + + const row = app._buildHomeSessionRow(app.buildHomeSessionRows()[0]); + expect(collect(row).filter((el: any) => String(el.className).includes('--watching'))).toHaveLength(0); + }); +}); + +/** Every node under one the builders made, the row itself included. */ +function collect(node: any): any[] { + return [node, ...(node.children || []).flatMap((child: any) => collect(child))]; +} diff --git a/test/mobile-overview.test.ts b/test/mobile-overview.test.ts index 232b3d35..2269b585 100644 --- a/test/mobile-overview.test.ts +++ b/test/mobile-overview.test.ts @@ -19,8 +19,11 @@ function fakeElement(): any { type: '', dataset: {}, style: {}, + attrs: {} as Record, children: [] as any[], - setAttribute() {}, + setAttribute(name: string, value: string) { + el.attrs[name] = value; + }, appendChild(child: any) { el.children.push(child); return child; @@ -485,3 +488,53 @@ describe('mobile overview run picker (CLI availability gating)', () => { expect(gate).toContain('isCliAvailable'); }); }); + +describe('mobile overview watching badge', () => { + it('carries what the pane says is running in the background', () => { + const app = loadOverviewApp(); + const model = app.buildMobileOverviewModel({ + sessions: [session({ id: 'a', status: 'idle', watching: '1 monitor' }), session({ id: 'b', status: 'idle' })], + cases: CASES, + }); + + const rows = Object.fromEntries(model.current.map((r: any) => [r.id, r.watching])); + expect(rows).toEqual({ a: '1 monitor', b: '' }); + }); + + it('leaves a session that is ALSO blocked on a dialog in NEEDS YOU', () => { + // An agent can arm a monitor and ask the user a question in the same breath, so the + // badge adds a fact to the row and never moves it out of the group that says a human + // is needed. Only the row's own pill decides that. + const app = loadOverviewApp(); + const model = app.buildMobileOverviewModel({ + sessions: [session({ id: 'a', status: 'idle', watching: '2 shells' })], + cases: CASES, + pendingHooks: new Map([['a', new Set(['permission_prompt'])]]), + }); + + expect(model.needsYou.map((r: any) => r.id)).toEqual(['a']); + expect(model.needsYou[0].pill).toBe('needs you'); + expect(model.needsYou[0].watching).toBe('2 shells'); + }); + + it('says one word and puts the detail where every surface can reach it', () => { + const app = loadOverviewApp(); + const badge = app._buildWatchingBadge('1 monitor', 'mobile-overview-pill'); + + expect(badge.className).toBe('mobile-overview-pill mobile-overview-pill--watching'); + expect(badge.textContent).toBe('watching'); + expect(badge.title).toBe('Still running in the background: 1 monitor'); + // A phone has no hover target and a screen reader reads neither the class nor the + // tooltip, so the label has to be here too or this surface says only "watching". + expect(badge.attrs['aria-label']).toBe('Still running in the background: 1 monitor'); + }); + + it('takes the pill class of whichever surface asks for it', () => { + // The phone's pill styles live inside a media query the desktop rail never enters, + // so the rail passes its own base class and gets the same badge in its own clothes. + const app = loadOverviewApp(); + expect(app._buildWatchingBadge('1 shell', 'home-sessions-pill').className).toBe( + 'home-sessions-pill home-sessions-pill--watching' + ); + }); +}); diff --git a/test/mocks/mock-session.ts b/test/mocks/mock-session.ts index 5520c3a5..5f04d97e 100644 --- a/test/mocks/mock-session.ts +++ b/test/mocks/mock-session.ts @@ -35,6 +35,12 @@ export class MockSession extends EventEmitter { /** `null` once the PTY is gone (or before it has ever started) — see `pid` in Session. */ pid: number | null = 12345; isWorking: boolean = false; + /** + * Mirrors Session.watching — what the pane's footer says is still running in the + * background. An idle prompt from such a session opens acknowledged, so the routes + * need to be able to set it. + */ + watching: string | null = null; private _activeChildProcesses: { pid: number; command: string }[] = []; ralphTracker: null = null; writeBuffer: string[] = []; diff --git a/test/routes/approval-routes.test.ts b/test/routes/approval-routes.test.ts index 8fa48544..439459a0 100644 --- a/test/routes/approval-routes.test.ts +++ b/test/routes/approval-routes.test.ts @@ -325,6 +325,65 @@ describe('approval routes', () => { }); }); + describe('an idle prompt from a session watching its own background work', () => { + beforeEach(() => { + session.terminalBuffer = 'claude> waiting at the composer'; + session.watching = '1 monitor'; + }); + + it('opens acknowledged, so no surface has an alert to raise', async () => { + await postHook(harness, 'idle_prompt', { message: 'Claude is waiting for your input' }); + const [item] = await listApprovals(harness); + expect(item).toMatchObject({ kind: 'idle', acknowledgedReason: 'watching 1 monitor' }); + expect(item.acknowledgedAt).toEqual(expect.any(Number)); + }); + + it('tells a live page why, so it declines to arm the alert', async () => { + await postHook(harness, 'idle_prompt', {}); + const broadcast = harness.ctx.broadcast.mock.calls.find((c) => c[0] === 'hook:idle_prompt'); + expect(broadcast?.[1]).toMatchObject({ acknowledgedReason: 'watching 1 monitor' }); + }); + + it('sends no push', async () => { + // The loudest surface, and the one a false alarm is hardest to ignore on. + await postHook(harness, 'idle_prompt', {}); + expect(harness.ctx.sendPushNotifications.mock.calls.find((c) => c[0] === 'hook:idle_prompt')).toBeUndefined(); + }); + + it('stays answerable from the drawer', async () => { + await postHook(harness, 'idle_prompt', {}); + const [item] = await listApprovals(harness); + const res = await harness.app.inject({ + method: 'POST', + url: `/api/approvals/${item.id}/answer`, + payload: { action: 'text', text: 'carry on' }, + }); + expect(res.statusCode).toBe(200); + expect(session.writeBuffer.join('')).toContain('carry on'); + }); + + it('still raises a permission dialog from the same session', async () => { + // Watching says nothing about a dialog: that one blocks the agent outright. + session.terminalBuffer = PERMISSION_DIALOG; + await postHook(harness, 'permission_prompt', { tool_name: 'Bash' }); + const [item] = await listApprovals(harness); + expect(item.acknowledgedAt).toBeUndefined(); + expect(item.acknowledgedReason).toBeUndefined(); + const broadcast = harness.ctx.broadcast.mock.calls.find((c) => c[0] === 'hook:permission_prompt'); + expect(broadcast?.[1]).not.toMatchObject({ acknowledgedReason: expect.any(String) }); + expect(harness.ctx.sendPushNotifications.mock.calls.find((c) => c[0] === 'hook:permission_prompt')).toBeDefined(); + }); + + it('alerts normally again once the background work is over', async () => { + await postHook(harness, 'idle_prompt', {}); + session.watching = null; + await postHook(harness, 'idle_prompt', {}); + const [item] = await listApprovals(harness); + expect(item.acknowledgedAt).toBeUndefined(); + expect(harness.ctx.sendPushNotifications.mock.calls.filter((c) => c[0] === 'hook:idle_prompt')).toHaveLength(1); + }); + }); + it('viewing a session acknowledges its idle prompt (item stays pending) and broadcasts it', async () => { session.terminalBuffer = 'claude> waiting at the composer'; await postHook(harness, 'idle_prompt', {}); diff --git a/test/session-listener-wiring.test.ts b/test/session-listener-wiring.test.ts index 1173fecb..3345c25d 100644 --- a/test/session-listener-wiring.test.ts +++ b/test/session-listener-wiring.test.ts @@ -22,6 +22,21 @@ describe('session listener wiring', () => { expect(registerAttachment).toHaveBeenNthCalledWith(2, 'wiring-attach-source-test', '/tmp/report.pdf', 'external'); }); + it('pushes the session state when the watching label changes on its own', () => { + // The badge appears on the idle transition, which broadcasts anyway. It goes AWAY + // when the background work ends, and a CLI can do that without taking a turn — codex + // repaints its background-terminal row away and stays idle — so nothing else fires + // and every open page would keep drawing a badge the server had already dropped. + const session = new Session({ id: 'wiring-watching-test', workingDir: '/tmp', mode: 'codex' }); + const broadcastSessionStateDebounced = vi.fn(); + const deps = { broadcastSessionStateDebounced } as unknown as Parameters[1]; + + const refs = createSessionListeners(session, deps); + refs.watchingChanged(); + + expect(broadcastSessionStateDebounced).toHaveBeenCalledWith('wiring-watching-test'); + }); + /** The listener reads the setting asynchronously; let its promise chain settle. */ const flush = () => new Promise((resolve) => setTimeout(resolve, 5)); diff --git a/test/session-sidebar-ux.test.ts b/test/session-sidebar-ux.test.ts index 8ea2d0d6..6b2ed047 100644 --- a/test/session-sidebar-ux.test.ts +++ b/test/session-sidebar-ux.test.ts @@ -73,3 +73,53 @@ describe('vertical session navigation UX contract', () => { expect(i18n).toContain("'Adjust only session names in the vertical sidebar.':"); }); }); + +describe('watching badge on a rich session row', () => { + it('reads the label off the session payload', () => { + expect(app).toContain("watching: typeof session.watching === 'string' ? session.watching : ''"); + }); + + it('renders it beside the state pill rather than in place of it', () => { + // A session can be watching a monitor AND holding a question for the user, so the + // pill that says which one still decides the row; this badge only adds a fact. + const meta = app.slice(app.indexOf('_sidebarRichMetaHTML(row) {')); + const body = meta.slice(0, meta.indexOf('_sidebarRichStampText(timestamp, format) {')); + expect(body).toContain('tab-pill tab-pill--${escapeHtml(row.state)}'); + expect(body).toContain('tab-pill tab-pill--watching'); + expect(body).toContain('Still running in the background:'); + }); + + it('re-renders the row when the background work changes', () => { + // The meta line is rebuilt only when this signature moves, so a badge left out of + // it would appear and disappear a render late, or not at all. + expect(app).toContain( + 'const sig = `${row.state}:${row.since ? row.since.at : 0}:${row.createdAt}:${row.watching}`' + ); + }); + + it('escapes the label everywhere it reaches markup', () => { + // `watching` is pane-derived and a config-supplied pattern decides what its capture + // group holds, so every interpolation of it into HTML has to go through escapeHtml(). + // The row is installed with innerHTML, which makes an unescaped quote in that + // attribute an injection rather than a cosmetic bug. + expect(app).toContain('${richRow.createdAt}:${escapeHtml(richRow.watching)}"'); + expect(app).not.toContain('${richRow.createdAt}:${richRow.watching}"'); + }); + + it('words the tooltip exactly as the phone overview does', () => { + // Both files build this sentence themselves, deliberately, so that a stale cached + // module still renders a complete row. Substring-matching the prefix would let the + // two drift; the whole sentence is what has to agree. + const overview = readFileSync(resolve(publicDir, 'mobile-overview.js'), 'utf8'); + expect(overview).toContain("'Still running in the background: ' + label"); + expect(app).toContain('`Still running in the background: ${row.watching}`'); + }); + + it('colours it with the accent, never with the two colours that mean a human is needed', () => { + const rule = styles.slice(styles.indexOf('.tab-pill--watching')); + const block = rule.slice(0, rule.indexOf('}')); + expect(block).toContain('var(--accent)'); + expect(block).not.toContain('var(--red)'); + expect(block).not.toContain('var(--yellow)'); + }); +}); diff --git a/test/session-watching.test.ts b/test/session-watching.test.ts new file mode 100644 index 00000000..b28f9bbb --- /dev/null +++ b/test/session-watching.test.ts @@ -0,0 +1,396 @@ +/** + * The badge a session wears while work it started in the background is still running. + * + * The bug this pins: an agent that arms a monitor, backgrounds a shell or hands work to a + * cloud session is told to end its turn, so the pane falls quiet, Claude Code's idle + * notification arrives a minute later, and every Codeman surface files the session under + * NEEDS YOU. Nothing wants the user there. The CLI itself says so on the last row of its + * screen (`⏵⏵ bypass permissions on · 1 monitor · ← for agents`), and reading that row is + * what tells a session waiting for its own background work from one waiting for a human. + * + * The pane fixtures below are verbatim captures (`tmux -L codeman capture-pane -p`) from a + * live Claude Code 2.1.278 session on 2026-09-21. + */ +import { describe, expect, it, vi, afterEach } from 'vitest'; +import { Session } from '../src/session.js'; +import { getCli } from '../src/config/cli-registry/index.js'; +import { compileVersionRegex } from '../src/config/cli-registry/patterns.js'; +import { + watchingLabel, + WATCHING_TAIL_LINES, + MAX_WATCHING_LABEL_CHARS, + IDLE_SILENCE_MS, +} from '../src/session-activity.js'; + +/** The registry's own patterns, which are what every consumer runs. */ +const CLAUDE_WATCHING = compileVersionRegex(getCli('claude')!.capabilities.workDetect!.watchingLine!)!; +const CODEX_WATCHING = compileVersionRegex(getCli('codex')!.capabilities.workDetect!.watchingLine!)!; +const CODEX_TAIL = getCli('codex')!.capabilities.workDetect!.watchingLines!; + +/** + * The foot of a Codex pane, verbatim (codex-cli 0.154.0, 2026-09-22). Codex does not + * write on its last row: the status line is there, the composer above it, and the + * background-terminal row above that, which is why codex declares its own window. + */ +const CODEX_STATUS = + ' gpt-5.6-sol medium · Context 98% left · ~/codeman-cases/codex-probe · 5h 99% left · weekly 94% left'; +const CODEX_WITH_TERMINAL = [ + '• OK', + '', + ' 1 background terminal running · /ps to view · /stop to close', + '', + '', + '› Ask Codex to do anything', + '', + CODEX_STATUS, + '', +].join('\n'); +const CODEX_STOPPED = [ + '• OK', + '', + '• Stopping all background terminals.', + '', + '', + '› Ask Codex to do anything', + '', + CODEX_STATUS, + '', +].join('\n'); + +/** The bottom of a Claude pane: composer, the user's status line, the footer row. */ +function pane(footer: string, body = ''): string { + return ( + body + + '╭──────────────────────────────────────╮\n' + + '│ ❯ │\n' + + '╰──────────────────────────────────────╯\n' + + ' ~/innovi/gtd-board [main] Opus 5 ctx: 11%\n' + + ` ${footer}\n` + ); +} + +const WITH_MONITOR = pane('⏵⏵ bypass permissions on · 1 monitor · ← for agents'); +const WITH_SHELL = pane('⏵⏵ bypass permissions on · 1 shell · ← for agents'); +const NOTHING_RUNNING = pane('⏵⏵ bypass permissions on (shift+tab to cycle) · ← for agents'); + +/** A composer repaint: the frame Claude ships roughly once a second while working. */ +const COMPOSER_REPAINT = + '\x1b[31;1H\x1b[38;5;246m❯\xa0\x1b[39m\x1b[0m\x1b[33;1H \x1b[38;5;246mOpus 5 in:143,699 out:669 ctx:14%\x1b[39m'; + +type SessionInternals = { + _handleTerminalOutput(data: string): void; + _detectInteractiveActivity(data: string): void; +}; + +/** One PTY chunk, exactly as the interactive handler sees it. */ +function feed(session: Session, data: string): void { + const internals = session as unknown as SessionInternals; + internals._handleTerminalOutput(data); + internals._detectInteractiveActivity(data); +} + +/** A session whose mux reports a fixed (or scripted) screen for the pane probe to read. */ +function withFakePane(screen: string | (() => string), mode: 'claude' | 'codex' = 'claude'): Session { + const read = typeof screen === 'function' ? screen : () => screen; + const mux = { + isAvailable: () => true, + capturePaneText: () => read(), + } as unknown as NonNullable[0]>['mux']; + return new Session({ + workingDir: '/tmp', + mode, + mux, + muxSession: { muxName: 'codeman-test', sessionId: 'test', createdAt: Date.now() }, + } as ConstructorParameters[0]); +} + +/** Codex's own composer repaint, the frame that arms its idle confirmation. */ +const CODEX_COMPOSER_REPAINT = '\x1b[31;1H\x1b[38;5;246m›\xa0\x1b[39m\x1b[0m'; + +/** Run one turn and let it end, which is when the probe reads the screen. */ +function runAndSettle(session: Session, repaint: string = COMPOSER_REPAINT): void { + for (let i = 0; i < 3; i++) { + feed(session, repaint); + vi.advanceTimersByTime(1000); + } + vi.advanceTimersByTime(IDLE_SILENCE_MS + 2000); +} + +describe('watchingLabel', () => { + it('reads the label off the footer row', () => { + expect(watchingLabel(WITH_MONITOR, CLAUDE_WATCHING)).toBe('1 monitor'); + expect(watchingLabel(WITH_SHELL, CLAUDE_WATCHING)).toBe('1 shell'); + }); + + it('reads every kind of background work the CLI names', () => { + const labels = [ + '2 monitors', + '3 shells', + '1 cloud session', + '2 cloud sessions', + '1 local agent', + '4 background tasks', + '1 MCP task', + '1 background dynamic workflow', + '2 remote dynamic workflows', + '1 Artifact comment monitor', + '2 teams', + ]; + for (const label of labels) { + expect(watchingLabel(pane(`⏵⏵ bypass permissions on · ${label} · ← for agents`), CLAUDE_WATCHING)).toBe(label); + } + }); + + it('says nothing about a pane that is running nothing', () => { + expect(watchingLabel(NOTHING_RUNNING, CLAUDE_WATCHING)).toBeNull(); + expect(watchingLabel('', CLAUDE_WATCHING)).toBeNull(); + expect(watchingLabel(null, CLAUDE_WATCHING)).toBeNull(); + }); + + it('ignores the same words in the transcript above the composer', () => { + // The whole reason the search is confined to the foot of the screen: a session that + // PRINTS "1 monitor" (this one has been discussing exactly that) is not running one. + const transcript = + '> does Codeman know about watching?\n' + + '⏺ The footer says · 1 monitor · while a monitor is armed, and · 2 shells · for\n' + + ' backgrounded commands. Codeman reads neither today.\n' + + ' Nothing else on the screen means background work is running.\n'; + expect(watchingLabel(pane('⏵⏵ bypass permissions on · ← for agents', transcript), CLAUDE_WATCHING)).toBeNull(); + }); + + it('looks no further up the screen than the tail it declares', () => { + const chip = '⏵⏵ bypass permissions on · 1 monitor · ← for agents'; + const below = Array(WATCHING_TAIL_LINES).fill(' still here').join('\n'); + // Blank lines are dropped before the tail is taken, so a pane padded with them must + // still read its own footer. + expect(watchingLabel(`${chip}\n\n\n\n\n\n`, CLAUDE_WATCHING)).toBe('1 monitor'); + expect(watchingLabel(`${chip}\n${below}\n`, CLAUDE_WATCHING)).toBeNull(); + }); + + it('refuses a chip on the row above the footer, which the agent can write', () => { + // The status line is one row up, its text comes from a `statusLine` command, and a + // session running with permissions bypassed can write that command into + // `.claude/settings.json` in its own workspace. The window is what keeps that row + // out, so this is the test that would fail if somebody widened it. + const forged = pane('⏵⏵ bypass permissions on (shift+tab to cycle) · ← for agents').replace( + ' ~/innovi/gtd-board [main] Opus 5 ctx: 11%', + ' ~/innovi/gtd-board [main] Opus 5 ctx: 11% · 1 monitor' + ); + expect(watchingLabel(forged, CLAUDE_WATCHING)).toBeNull(); + // And with the window widened by one, the same screen does match — which is the + // whole reason the default is one row. + expect(watchingLabel(forged, CLAUDE_WATCHING, 2)).toBe('1 monitor'); + }); + + it('keeps Claude on the default window, because its chip is the last row', () => { + expect(getCli('claude')?.capabilities.workDetect?.watchingLines).toBeUndefined(); + expect(WATCHING_TAIL_LINES).toBe(1); + }); + + it('refuses a label the footer did not separate, which is the injection guard', () => { + // The pattern anchors on the `·` the footer joins its items with. Without that + // anchor an agent could silence its own idle alert by printing the words, since the + // only rows it cannot write are the footer and the status line. + expect(watchingLabel(pane('1 monitor'), CLAUDE_WATCHING)).toBeNull(); + expect(watchingLabel(pane('running 2 shells for the build'), CLAUDE_WATCHING)).toBeNull(); + expect(watchingLabel(pane('⏵⏵ bypass permissions on · 1 monitor'), CLAUDE_WATCHING)).toBe('1 monitor'); + }); + + it('reads a coloured footer, because a capture may carry ANSI', () => { + const coloured = pane('\u001b[2m⏵⏵ bypass permissions on\u001b[0m · \u001b[36m1 monitor\u001b[0m · ← for agents'); + expect(watchingLabel(coloured, CLAUDE_WATCHING)).toBe('1 monitor'); + }); + + it('caps the label, because it ends up on a badge and in an approval card', () => { + const long = `· ${'9'.repeat(MAX_WATCHING_LABEL_CHARS * 2)} monitors`; + const label = watchingLabel(pane(`⏵⏵ bypass permissions on ${long} · ← for agents`), CLAUDE_WATCHING); + expect(label?.length).toBe(MAX_WATCHING_LABEL_CHARS); + }); + + it('survives a pattern handed to it with the global flag set', () => { + // compileVersionRegex() never sets `g`, but a test or a reloaded config might, and a + // sticky lastIndex would make the same screen match every other call. + const global = new RegExp(CLAUDE_WATCHING.source, 'g'); + expect(watchingLabel(WITH_MONITOR, global)).toBe('1 monitor'); + expect(watchingLabel(WITH_MONITOR, global)).toBe('1 monitor'); + }); +}); + +describe('Session.watching', () => { + afterEach(() => { + vi.useRealTimers(); + }); + + it('carries what the pane reported once the turn ends', () => { + vi.useFakeTimers(); + const session = withFakePane(WITH_MONITOR); + expect(session.watching).toBeNull(); + + runAndSettle(session); + + expect(session.status).toBe('idle'); + expect(session.watching).toBe('1 monitor'); + }); + + it('lets the badge go when the background work is over', () => { + vi.useFakeTimers(); + const screen = { text: WITH_MONITOR }; + const session = withFakePane(() => screen.text); + + runAndSettle(session); + expect(session.watching).toBe('1 monitor'); + + screen.text = NOTHING_RUNNING; + runAndSettle(session); + expect(session.watching).toBeNull(); + }); + + it('announces the change, because the session status does not move with it', () => { + // Measured on codex: a background terminal finishing repaints the row away and the + // session is idle before and after, so no other event fires. Without this one the + // server drops the label and every open page goes on drawing the badge. + vi.useFakeTimers(); + const screen = { text: WITH_MONITOR }; + const session = withFakePane(() => screen.text); + const changes: (string | null)[] = []; + session.on('watchingChanged', () => changes.push(session.watching)); + + runAndSettle(session); + expect(changes).toEqual(['1 monitor']); + + // A repaint that carries the composer glyph but no chip: the pane went quiet again + // without a turn, which is exactly the case the event exists for. + screen.text = NOTHING_RUNNING; + feed(session, COMPOSER_REPAINT); + vi.advanceTimersByTime(IDLE_SILENCE_MS + 2000); + expect(changes).toEqual(['1 monitor', null]); + expect(session.status).toBe('idle'); + }); + + it('says nothing while the answer stays the same', () => { + vi.useFakeTimers(); + const session = withFakePane(WITH_MONITOR); + const changes: (string | null)[] = []; + session.on('watchingChanged', () => changes.push(session.watching)); + + runAndSettle(session); + runAndSettle(session); + runAndSettle(session); + expect(changes).toEqual(['1 monitor']); + }); + + it('keeps its last answer when the screen cannot be read', () => { + vi.useFakeTimers(); + const screen: { text: string | null } = { text: WITH_MONITOR }; + const session = withFakePane(() => screen.text as string); + + runAndSettle(session); + expect(session.watching).toBe('1 monitor'); + + // A capture that fails is not evidence that nothing is running, which is the same + // rule the working probe applies to its own null. + screen.text = null; + runAndSettle(session); + expect(session.watching).toBe('1 monitor'); + }); + + it('reads Codex own row, three up from the bottom of its screen', () => { + vi.useFakeTimers(); + const session = withFakePane(CODEX_WITH_TERMINAL, 'codex'); + runAndSettle(session, CODEX_COMPOSER_REPAINT); + expect(session.status).toBe('idle'); + expect(session.watching).toBe('1 background terminal'); + }); + + it('reports nothing for a CLI whose screen nobody has characterised', () => { + vi.useFakeTimers(); + expect(getCli('gemini')?.capabilities.workDetect).toBeUndefined(); + const session = withFakePane(WITH_MONITOR, 'gemini'); + runAndSettle(session); + expect(session.watching).toBeNull(); + }); + + it('never reads another CLI screen', () => { + vi.useFakeTimers(); + // Each pattern is anchored on chrome its own CLI draws, so neither can fire on the + // other's pane. A shared fallback would have both reading a screen nobody measured. + const codexOnClaudeScreen = withFakePane(WITH_MONITOR, 'codex'); + runAndSettle(codexOnClaudeScreen, CODEX_COMPOSER_REPAINT); + expect(codexOnClaudeScreen.watching).toBeNull(); + + const claudeOnCodexScreen = withFakePane(CODEX_WITH_TERMINAL, 'claude'); + runAndSettle(claudeOnCodexScreen); + expect(claudeOnCodexScreen.watching).toBeNull(); + }); + + it('rides along on the payload every session surface reads', () => { + vi.useFakeTimers(); + const session = withFakePane(WITH_SHELL); + runAndSettle(session); + + expect(session.toLightDetailedState().watching).toBe('1 shell'); + }); +}); + +describe('the row Codex draws', () => { + it('reads the label, and only while a terminal is running', () => { + expect(watchingLabel(CODEX_WITH_TERMINAL, CODEX_WATCHING, CODEX_TAIL)).toBe('1 background terminal'); + expect(watchingLabel(CODEX_STOPPED, CODEX_WATCHING, CODEX_TAIL)).toBeNull(); + }); + + it('counts terminals', () => { + const three = CODEX_WITH_TERMINAL.replace('1 background terminal running', '3 background terminals running'); + expect(watchingLabel(three, CODEX_WATCHING, CODEX_TAIL)).toBe('3 background terminals'); + }); + + it('needs the window Codex declares: its row is not the last one', () => { + // Pins WHY `watchingLines` exists. Claude's default of two rows reaches the status + // line and the composer, and Codex's row sits one further up. + expect(watchingLabel(CODEX_WITH_TERMINAL, CODEX_WATCHING, 2)).toBeNull(); + expect(CODEX_TAIL).toBeGreaterThanOrEqual(3); + }); + + it('refuses a mention that is not the whole row', () => { + // The pattern matches Codex's row end to end, so prose about background terminals — + // including prose quoting part of the row — is not enough. + for (const line of [ + '• I left 1 background terminal running for you.', + ' 1 background terminal running · /ps to view', + ' see: 1 background terminal running · /ps to view · /stop to close', + ]) { + const claim = CODEX_STOPPED.replace('• Stopping all background terminals.', line); + expect(watchingLabel(claim, CODEX_WATCHING, CODEX_TAIL)).toBeNull(); + } + }); + + it('CAN be forged by Codex own output, and is contained by Codex having no hooks', () => { + // Codex's row is third from the bottom only while a terminal runs; with none running + // that slot is the last row of the transcript, which the agent writes. Matching the + // complete row raises the bar but closes nothing, so this test states the limitation + // rather than a protection the code does not have. + const forged = CODEX_STOPPED.replace( + '• Stopping all background terminals.', + ' 1 background terminal running · /ps to view · /stop to close' + ); + expect(watchingLabel(forged, CODEX_WATCHING, CODEX_TAIL)).toBe('1 background terminal'); + + // What makes that cost a wrong badge and nothing more: no hook event from a codex + // session reaches the approvals inbox, so there is no idle item to pre-acknowledge + // and no alert to silence. A CLI that gains hook signals needs a harder anchor first. + expect(getCli('codex')?.capabilities.hooks).toBe('none'); + }); +}); + +describe('the registry pattern Claude declares', () => { + it('is one the config-regex guard accepts', () => { + // Same guard as `workingLine`: ~/.codeman/clis.json can set this field, and the + // compiled pattern runs over a pane capture on a timer. + expect(compileVersionRegex(getCli('claude')!.capabilities.workDetect!.watchingLine!)).not.toBeNull(); + }); + + it('does not fire on the status line a user configured', () => { + // Plan-usage and context figures live one row above the footer and carry numbers. + const statusLine = ' ~/innovi/gtd-board [main] Opus 5 (1M context) high ctx: 10% 5h: 48% (32m) 7d: 15% (6d10h)'; + expect(CLAUDE_WATCHING.test(statusLine)).toBe(false); + }); +}); diff --git a/test/tui/tui-approvals.test.ts b/test/tui/tui-approvals.test.ts index f9072f80..81d0585c 100644 --- a/test/tui/tui-approvals.test.ts +++ b/test/tui/tui-approvals.test.ts @@ -59,6 +59,24 @@ describe('approvalCard', () => { expect(card.hint).toBe('p to reply'); }); + it('says why an acknowledged idle prompt is quiet, in the drawer own words', () => { + // The session is watching work it started itself. Asking for a reply would be the + // same false alarm the acknowledgement exists to remove, so the card states the + // reason and drops out of the warning vocabulary — while staying answerable. + const card = approvalCard( + item({ + kind: 'idle', + message: 'Claude is waiting for your input', + options: undefined, + acknowledgedAt: 1_700_000_000_000, + acknowledgedReason: 'watching 1 monitor', + }) + ); + expect(card.tone).toBe('info'); + expect(card.title).toBe('quiet, watching 1 monitor'); + expect(card.hint).toBe('p to reply'); + }); + it('drops the approve/deny-only hint when the frame did not parse', () => { const card = approvalCard(item({ options: undefined })); expect(card.options).toEqual([]); @@ -81,6 +99,12 @@ describe('approvalTone', () => { expect(approvalTone(item({ kind: 'question' }))).toBe('err'); expect(approvalTone(item({ kind: 'idle' }))).toBe('warn'); }); + + it('is neither for a prompt that opened acknowledged', () => { + expect(approvalTone(item({ kind: 'idle', acknowledgedReason: 'watching 2 shells' }))).toBe('info'); + // A dialog stays red whatever else the session started. + expect(approvalTone(item({ kind: 'permission', acknowledgedReason: 'watching 2 shells' }))).toBe('err'); + }); }); describe('approvalAnswerForKey', () => { diff --git a/test/tui/tui-model.test.ts b/test/tui/tui-model.test.ts index 2c3b34ef..c081b9cf 100644 --- a/test/tui/tui-model.test.ts +++ b/test/tui/tui-model.test.ts @@ -71,6 +71,32 @@ describe('classifySession', () => { expect(classifySession(row, approval({ sessionId: 'a', kind: 'idle' }))).toBe('waiting'); }); + it('stops an ACKNOWLEDGED prompt deciding the row', () => { + // Acknowledgement means the alert this prompt armed has been spent, either because + // a human opened the session elsewhere or because the inbox opened the item that way + // for a session watching its own background work. The web has honoured that since + // acknowledgement existed; this gate used to read past it, so an alert cleared on a + // phone stayed lit here alone. + const quiet = session({ sessionId: 'a', status: 'idle' }); + const seen = approval({ sessionId: 'a', kind: 'idle', acknowledgedAt: NOW - 1_000 }); + expect(classifySession(quiet, seen)).toBe('idle'); + expect(classifySession(session({ sessionId: 'a', status: 'busy' }), seen)).toBe('working'); + }); + + it('keeps a blocking dialog lit whatever its acknowledgement says', () => { + // `acknowledge()` is idle-only by construction, so this cannot happen through the + // routes. It is pinned because the cost of the two being wired together later is a + // permission dialog that stops asking. + const row = session({ sessionId: 'a' }); + const at = NOW - 1_000; + expect(classifySession(row, approval({ sessionId: 'a', kind: 'permission', acknowledgedAt: at }))).toBe( + 'blocked-permission' + ); + expect(classifySession(row, approval({ sessionId: 'a', kind: 'question', acknowledgedAt: at }))).toBe( + 'blocked-question' + ); + }); + it('classifies a row the server no longer has live as history', () => { expect(classifySession(session({ sessionId: 'a', sources: ['history'], status: 'busy' }))).toBe('recent'); expect(classifySession(session({ sessionId: 'a', sources: ['persisted', 'lifecycle'] }))).toBe('recent'); @@ -78,6 +104,46 @@ describe('classifySession', () => { }); }); +describe('an acknowledged prompt on a watching session', () => { + const sessions = [session({ sessionId: 'watcher', status: 'idle', lastActivityAt: NOW - 120_000 })]; + const approvals = approvalMap([ + approval({ + sessionId: 'watcher', + kind: 'idle', + createdAt: NOW - 30_000, + acknowledgedAt: NOW - 30_000, + acknowledgedReason: 'watching 1 monitor', + }), + ]); + + it('raises no alert: the row leaves NEEDS YOU entirely', () => { + const groups = groupSessions(buildRows(sessions, approvals)); + const byKey = Object.fromEntries(groups.map((group) => [group.key, group.rows.map((r) => r.session.sessionId)])); + expect(byKey['needs-you']).toEqual([]); + expect(byKey['idle']).toEqual(['watcher']); + }); + + it('keeps the item on the row, because it is still pending and still answerable', () => { + const [row] = buildRows(sessions, approvals); + expect(row.approval?.acknowledgedReason).toBe('watching 1 monitor'); + }); + + it('dates the row from the pane going quiet, not from the prompt', () => { + // The prompt's age measures the state only while the prompt is what put the row in + // it. Reading it here would report "idle 30s" for a session quiet for two minutes. + const [row] = buildRows(sessions, approvals); + expect(row.since).toBe(NOW - 120_000); + }); + + it('alerts again once the same session goes quiet for an ordinary reason', () => { + // The inbox supersedes the acknowledged item and builds a fresh one, so this is the + // next prompt rather than the same one changing its mind. + const fresh = approvalMap([approval({ sessionId: 'watcher', kind: 'idle', createdAt: NOW - 1_000 })]); + const groups = groupSessions(buildRows(sessions, fresh)); + expect(groups[0].rows.map((row) => row.session.sessionId)).toEqual(['watcher']); + }); +}); + describe('groupSessions', () => { it('always returns the four groups in display order', () => { expect(groupSessions([]).map((group) => group.key)).toEqual(['needs-you', 'working', 'idle', 'recent']); diff --git a/test/tui/tui-render.test.ts b/test/tui/tui-render.test.ts index 43e5f2c9..fcb3df27 100644 --- a/test/tui/tui-render.test.ts +++ b/test/tui/tui-render.test.ts @@ -11,6 +11,7 @@ import { charWidth, stripStyles, toDisplayLines, visibleWidth } from '../../src/ import { composerMove, createComposer } from '../../src/tui/tui-composer.js'; import { computeLayout, needsBanner } from '../../src/tui/tui-layout.js'; import { createTuiModel, type TuiModelStore } from '../../src/tui/tui-model.js'; +import type { ApprovalItem } from '../../src/web/approval-inbox.js'; import { composerCursorCell, detectGlyphTier, @@ -19,6 +20,7 @@ import { formatPlanUsage, formatTokens, glyphsFor, + pendingApprovalCount, renderFrame, rowLabel, type TuiRenderOptions, @@ -682,3 +684,33 @@ describe('the unicode glyph set is safe to render', () => { expect(every.join('')).not.toContain('\u270B'); }); }); + +describe('pendingApprovalCount', () => { + const prompt = (over: Partial): ApprovalItem => ({ + id: 'bbb2:1', + sessionId: 'bbb2', + sessionName: 'w6-docs', + kind: 'idle', + createdAt: NOW - 30_000, + ...over, + }); + + it('counts a prompt nobody has seen', () => { + const model = fixture(); + model.setApprovals([prompt({})]); + expect(pendingApprovalCount(model)).toBe(1); + }); + + it('does not count one whose alert is already spent', () => { + // Both ways an item gets acknowledged: a human opening the session elsewhere, and the + // inbox opening it that way for a session watching its own background work. The row + // has left NEEDS YOU by the same flag, so a number in the header would point at a + // group the reader can see is empty. + const model = fixture(); + model.setApprovals([prompt({ acknowledgedAt: NOW - 20_000 })]); + expect(pendingApprovalCount(model)).toBe(0); + + model.setApprovals([prompt({ acknowledgedAt: NOW - 20_000, acknowledgedReason: 'watching 1 monitor' })]); + expect(pendingApprovalCount(model)).toBe(0); + }); +}); diff --git a/test/watching-no-alert.test.ts b/test/watching-no-alert.test.ts new file mode 100644 index 00000000..9c3ce8a7 --- /dev/null +++ b/test/watching-no-alert.test.ts @@ -0,0 +1,226 @@ +// Port: none (pure classifiers + vm-loaded frontend modules — no browser, no server). +// +// The point of the watching signal is an alert that does NOT fire, so the test that +// matters is the negative one. A session that ended its turn because it armed a monitor +// or backgrounded a shell gets an idle prompt from Claude Code about a minute later, and +// that prompt must reach every surface as a card nobody has to look at rather than as an +// alert. The same session holding a permission dialog must still go red everywhere. +// +// Each surface is exercised through the code it really runs: the TUI classifier, the +// live SSE handler in settings-ui.js, the reload seed in approvals-ui.js, and the state +// classifier both home screens share. +import { readFileSync } from 'node:fs'; +import { resolve } from 'node:path'; +import vm from 'node:vm'; +import { describe, expect, it, beforeEach } from 'vitest'; +import { ApprovalInbox, type ApprovalItem } from '../src/web/approval-inbox.js'; +import { buildRows, groupSessions } from '../src/tui/tui-model.js'; +import type { TuiSessionRow } from '../src/tui/tui-types.js'; + +const PUBLIC = resolve(import.meta.dirname, '../src/web/public'); +const SESSION = 'watcher-session'; + +/** A live unified row for the session under test, quiet at its composer. */ +function row(overrides: Partial = {}): TuiSessionRow { + return { + sessionId: SESSION, + sources: ['live'], + name: 'watch-probe', + mode: 'claude', + status: 'idle', + workingDir: '/home/dev/case', + createdAt: 1_000, + lastActivityAt: 2_000, + ...overrides, + }; +} + +/** + * The item the server really produces for this case, built by the real inbox rather + * than by hand, so a change to how `watching` is honoured breaks this file too. + */ +function itemFor(watching: string | null): ApprovalItem { + const inbox = new ApprovalInbox(); + const item = inbox.notePrompt({ sessionId: SESSION, sessionName: 'watch-probe', kind: 'idle', watching }); + inbox.stop(); + return item; +} + +/** Minimal fake DOM node, enough for the handlers these tests drive. */ +function fakeElement(): Record { + const el: Record = { + className: '', + textContent: '', + title: '', + dataset: {}, + style: {}, + children: [] as unknown[], + hidden: false, + classList: { add() {}, remove() {}, toggle() {}, contains: () => false }, + setAttribute() {}, + querySelector: () => null, + querySelectorAll: () => [], + appendChild(child: unknown) { + (el.children as unknown[]).push(child); + return child; + }, + }; + return el; +} + +interface FrontendApp { + pendingHooks: Map>; + notifications: string[]; + approvals: Map; + seeded: ApprovalItem[]; + setPendingHook(sessionId: string, hook: string): void; + clearPendingHooks(sessionId: string, hook?: string): void; + _onHookIdlePrompt(data: Record): void; + _onHookPermissionPrompt(data: Record): void; + seedApprovals(): Promise; + _mobileOverviewState(session: Record, hooks?: Set): string; +} + +/** + * The three frontend modules that decide whether a prompt becomes an alert, loaded into + * one context the way the page loads them. Everything they call that belongs to app.js + * is stubbed to record rather than to render. + */ +function loadFrontend(seed: ApprovalItem[] = []): FrontendApp { + const CodemanApp = function CodemanApp(this: unknown) {} as unknown as { prototype: Record }; + const context = vm.createContext({ + CodemanApp, + console, + window: {}, + CSS: { escape: (s: string) => s }, + document: { + documentElement: { getAttribute: () => null, dataset: {} }, + getElementById: () => null, + querySelector: () => null, + createElement: () => fakeElement(), + createElementNS: () => fakeElement(), + }, + MobileDetection: { getDeviceType: () => 'desktop' }, + }); + for (const file of ['constants.js', 'mobile-overview.js', 'approvals-ui.js', 'settings-ui.js']) { + vm.runInContext(readFileSync(resolve(PUBLIC, file), 'utf8'), context, { filename: file }); + } + + const app = Object.create(CodemanApp.prototype) as FrontendApp & Record; + app.pendingHooks = new Map(); + app.notifications = []; + app.approvals = new Map(); + app.seeded = seed; + app.setPendingHook = (sessionId: string, hook: string) => { + if (!app.pendingHooks.has(sessionId)) app.pendingHooks.set(sessionId, new Set()); + app.pendingHooks.get(sessionId)!.add(hook); + }; + app.clearPendingHooks = (sessionId: string, hook?: string) => { + if (!hook) app.pendingHooks.delete(sessionId); + else app.pendingHooks.get(sessionId)?.delete(hook); + }; + Object.assign(app, { + _notifySession: (_id: string, _level: string, kind: string) => app.notifications.push(kind), + _apiJson: async () => ({ approvals: app.seeded }), + approvalsInboxEnabled: () => true, + renderApprovals: () => {}, + renderSessionTabs: () => {}, + loadAppSettingsFromStorage: () => ({}), + }); + return app; +} + +describe('a watching session raises no alert on any surface', () => { + const watched = itemFor('1 monitor'); + + it('the store opens the prompt acknowledged, which is what every surface reads', () => { + expect(watched.acknowledgedAt).toEqual(expect.any(Number)); + expect(watched.acknowledgedReason).toBe('watching 1 monitor'); + }); + + it('codeman tui leaves the row out of NEEDS YOU', () => { + const groups = groupSessions(buildRows([row()], new Map([[SESSION, watched]]))); + const needsYou = groups.find((group) => group.key === 'needs-you')!; + expect(needsYou.rows).toEqual([]); + }); + + it('a live page declines to arm the tab alert and raises no desktop notification', () => { + const app = loadFrontend(); + app._onHookIdlePrompt({ sessionId: SESSION, acknowledgedReason: watched.acknowledgedReason }); + expect(app.pendingHooks.get(SESSION)).toBeUndefined(); + expect(app.notifications).toEqual([]); + }); + + it('a reloading page does not arm it either', async () => { + const app = loadFrontend([watched]); + await app.seedApprovals(); + expect(app.pendingHooks.get(SESSION)).toBeUndefined(); + // The card itself is still there to answer, which is the whole point of + // acknowledging the prompt rather than never creating it. + expect(app.approvals.get(watched.id)?.acknowledgedReason).toBe('watching 1 monitor'); + }); + + it('so both home screens classify the session as plainly idle', () => { + const app = loadFrontend(); + app._onHookIdlePrompt({ sessionId: SESSION, acknowledgedReason: watched.acknowledgedReason }); + const state = app._mobileOverviewState({ status: 'idle' }, app.pendingHooks.get(SESSION)); + expect(state).toBe('idle'); + }); +}); + +describe('an ordinary idle prompt still alerts everywhere', () => { + const plain = itemFor(null); + + it('the store leaves it unacknowledged', () => { + expect(plain.acknowledgedAt).toBeUndefined(); + }); + + it('codeman tui puts the row in NEEDS YOU', () => { + const groups = groupSessions(buildRows([row()], new Map([[SESSION, plain]]))); + expect(groups[0].rows.map((r) => r.session.sessionId)).toEqual([SESSION]); + }); + + it('a live page arms the tab alert and notifies', () => { + const app = loadFrontend(); + app._onHookIdlePrompt({ sessionId: SESSION, message: 'Claude is waiting for your input' }); + expect([...(app.pendingHooks.get(SESSION) ?? [])]).toEqual(['idle_prompt']); + expect(app.notifications).toEqual(['hook-idle']); + }); + + it('a reloading page arms it from the seed', async () => { + const app = loadFrontend([plain]); + await app.seedApprovals(); + expect([...(app.pendingHooks.get(SESSION) ?? [])]).toEqual(['idle_prompt']); + }); + + it('so both home screens put the session in NEEDS YOU', () => { + const app = loadFrontend(); + app._onHookIdlePrompt({ sessionId: SESSION }); + expect(app._mobileOverviewState({ status: 'idle' }, app.pendingHooks.get(SESSION))).toBe('waiting'); + }); +}); + +describe('a dialog blocking the agent alerts even while it watches', () => { + it('a live page arms and notifies whatever else the session started', () => { + // A permission prompt never opens acknowledged (the inbox gates on kind), so the + // handler never sees a reason and this is the ordinary path. Pinned because the two + // handlers sit side by side and the guard belongs on exactly one of them. + const app = loadFrontend(); + app._onHookPermissionPrompt({ sessionId: SESSION, tool: 'Bash' }); + expect([...(app.pendingHooks.get(SESSION) ?? [])]).toEqual(['permission_prompt']); + expect(app.notifications).toEqual(['hook-permission']); + }); + + it('codeman tui shows it as blocked', () => { + const inbox = new ApprovalInbox(); + const item = inbox.notePrompt({ + sessionId: SESSION, + sessionName: 'watch-probe', + kind: 'permission', + watching: '1 monitor', + }); + inbox.stop(); + const [built] = buildRows([row()], new Map([[SESSION, item]])); + expect(built.state).toBe('blocked-permission'); + }); +});