diff --git a/.changeset/watching-badge.md b/.changeset/watching-badge.md new file mode 100644 index 00000000..12a8d384 --- /dev/null +++ b/.changeset/watching-badge.md @@ -0,0 +1,5 @@ +--- +'aicodeman': minor +--- + +Sessions now say when they are watching work they started themselves. An agent that arms a monitor, backgrounds a shell or hands a task to a cloud session is told to end its turn, so the pane goes quiet, Claude Code's idle notification lands a minute later and the session shows up under NEEDS YOU with nothing for anyone to answer. Claude prints what it is still running on the last row of its screen, and the CLI registry now carries that row as `capabilities.workDetect.watchingLine`, so the idle probe reads the label ("1 monitor", "2 shells") along with the working line it already reads. The label reaches every session payload as `watching`, and the phone overview, the desktop home rail and the rich sidebar rows wear it as a `watching` badge in the accent colour. It sits beside the state pill and never replaces it, because an agent can arm a monitor and ask you a question in the same breath. diff --git a/CLAUDE.md b/CLAUDE.md index 6dd34e47..ee58931f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -201,7 +201,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`. -⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer (`❯`) about once a second all through a turn, so the old "saw a ❯, wait 2s → idle" rule flipped every working session to idle two seconds in (measured: a session mid-tool-call at 17 minutes reporting `status:"idle"`). Its working indicator is `✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`: the glyph animates through `· ✢ ✳ ∗ ✻ ✽`, the gerund is randomized, and the finished line (`✻ Cooked for 2m 49s`) carries the same glyph, so neither `SPINNER_PATTERN` (braille, not what current versions draw) nor a keyword list can see it. Matching the new line in the STREAM does not work either: tmux ships partial repaints, so the whole line reaches the PTY only every few tens of seconds. So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks the SCREEN via `capturePaneText()` + `CLAUDE_WORKING_LINE_PATTERN` before believing it; a sustained run of repaints (`session-activity.ts`, pure + unit tested) is what marks a turn as started, with the same screen probe vetoing keystroke echo. Idle now lands ~3-5s after a turn ends instead of 2s into one. ⚠️ **The composer glyph and the working line are per-CLI registry DATA** (`capabilities.workDetect`, #385), not Claude constants: claude declares `❯` plus the pattern above, codex declares `›` plus `[Ee]sc to interrupt`, and a CLI that declares neither falls back to Claude's pair, which is what every session used before the registry carried one. Before that, this whole mechanism was gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` for its entire life. ⚠️ `workingLine` is config-supplied (a user `clis.json` can set it) and the compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()`: a nested quantifier there is a ReDoS against the event loop, and the helper returns null rather than throwing so the fallback is structural. +⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer (`❯`) about once a second all through a turn, so the old "saw a ❯, wait 2s → idle" rule flipped every working session to idle two seconds in (measured: a session mid-tool-call at 17 minutes reporting `status:"idle"`). Its working indicator is `✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`: the glyph animates through `· ✢ ✳ ∗ ✻ ✽`, the gerund is randomized, and the finished line (`✻ Cooked for 2m 49s`) carries the same glyph, so neither `SPINNER_PATTERN` (braille, not what current versions draw) nor a keyword list can see it. Matching the new line in the STREAM does not work either: tmux ships partial repaints, so the whole line reaches the PTY only every few tens of seconds. So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks the SCREEN via `capturePaneText()` + `CLAUDE_WORKING_LINE_PATTERN` before believing it; a sustained run of repaints (`session-activity.ts`, pure + unit tested) is what marks a turn as started, with the same screen probe vetoing keystroke echo. Idle now lands ~3-5s after a turn ends instead of 2s into one. ⚠️ **The composer glyph and the working line are per-CLI registry DATA** (`capabilities.workDetect`, #385), not Claude constants: claude declares `❯` plus the pattern above, codex declares `›` plus `[Ee]sc to interrupt`, and a CLI that declares neither falls back to Claude's pair, which is what every session used before the registry carried one. Before that, this whole mechanism was gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` for its entire life. ⚠️ `workingLine` is config-supplied (a user `clis.json` can set it) and the compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()`: a nested quantifier there is a ReDoS against the event loop, and the helper returns null rather than throwing so the fallback is structural. ⚠️ **A quiet pane is not always a pane that wants you.** An agent that arms a monitor, backgrounds a shell or hands work to a cloud session is told to end its turn, so the pane goes quiet, Claude Code's `idle_prompt` notification lands a minute later and every surface files the session under NEEDS YOU with nothing to answer. The CLI states what it is still running on the last row of its screen (`⏵⏵ bypass permissions on · 1 monitor · ← for agents`), which is the optional `capabilities.workDetect.watchingLine`: `_confirmIdle()`'s own capture feeds `watchingLabel()` (`session-activity.ts`, pure), the label lands on `Session.watching` and rides `toLightDetailedState()` out to every surface as a `watching` badge. ⚠️ That badge NEVER replaces the state pill and never moves a row out of NEEDS YOU: an agent can arm a monitor and ask a question in the same breath, and only the pill says which. The search is confined to the last `WATCHING_TAIL_LINES` lines of the capture, because the transcript above the composer quotes arbitrary text and a session that PRINTS "1 monitor" is not running one. Tests: `test/session-watching.test.ts`. **Workspace-trust dialog auto-accept** (`session-trust-dialog.ts`, pure + unit tested): Claude Code asks once per directory ("Is this a project you created or one you trust?") before it will read or edit anything, and since Codeman sessions run permission-skipping or classifier-guarded modes the answer is always yes, so a session parked on that dialog is simply stuck. ⚠️ **Match the compacted SCREEN, never the stream.** tmux repaints a row by writing each word and then a cursor-forward (`\x1b[C`) instead of a space, and Ink colours each word separately, so the wire carries `I\x1b[Ctrust\x1b[Cthis\x1b[Cfolder`; stripping the escapes leaves `Itrustthisfolder`, because the spaces are not there to strip, they were never sent. A plain `includes('trust this folder')` therefore never matched a single chunk and the auto-accept was silently DEAD for every session that hit the dialog. `compactScreenText()` removes ALL whitespace instead (plus the `ESC ( B` charset selects that `stripAnsi` does not cover, which would otherwise land inside a phrase as a literal `(B`), which survives both that repaint style and the spaced full-screen redraw. ⚠️ **Never answer it with a blind `\r`.** The layout has changed under us at least twice, and Claude Code 2.1.252 dropped the option numbers, put "No, exit" FIRST and highlights IT by default, so the Enter that answered the old dialog now picks *exit* and the pane dies (`Pane is dead (status 1)`) seconds after the session starts. `trustDialogNextKey()` reads the `❯` marker and returns ONE step at a time (an arrow while the cursor is on the wrong option, Enter only once the screen shows it on the trust option), with the pane re-read between steps, so a dropped arrow costs a repaint instead of the session; a frame that does not say which option is highlighted returns null and waits for the next repaint. ⚠️ The LAST marked option in the text wins, because the direct-PTY fallback reads an append-only buffer where every repaint since launch is still present and an older frame must not out-vote the freshest one. ⚠️ Answering types into a live session, so THREE guards must all hold and none is redundant: a **startup-only window** (`TRUST_DIALOG_WINDOW_MS`, 90s, since the dialog renders before the main UI and leaving it open forever would let an agent transcript that merely QUOTES the dialog trigger an Enter, this file being an example), a **two-marker match** requiring a trust phrase AND one of the dialog's own confirm affordances (`isTrustDialogScreen`), and an **attempt cap** (`TRUST_DIALOG_MAX_ATTEMPTS`, 6: a keystroke can land while Ink is still mounting the widget and be dropped, which is the other half of why sessions got stuck here, but retrying forever would hammer keys into whatever came next; it was 3 while one Enter answered the dialog, and answering now costs at least two keystrokes). ⚠️ It reads `capturePaneText()` and falls back to a deliberately SHORT tail of the terminal buffer only on a direct-PTY session, which has no pane: that buffer is append-only, so a longer tail would keep re-matching a dialog answered minutes ago. ⚠️ **The scan must schedule its own next read** (`_trustDialogTimer`, cleared in `_clearAllTimers()`): it runs from the PTY `onData` handler, which was enough while one Enter answered the dialog, but the arrow that moves the cursor is the LAST output the pane produces, so a two-keystroke answer waiting on more output parks forever with the cursor sitting on the right option (measured on a live 2.1.252 spawn: cursor moved at 6 s, then nothing). @@ -223,7 +223,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Docker Compose deployment** (`docker/`, contributed): Codeman itself runs in a container and spawns Docker cases as **SIBLING** containers through the mounted host socket (Docker-outside-of-Docker), never nested. That inverts one assumption the bare-host path takes for granted: the daemon no longer shares Codeman's filesystem, so a bind source valid *inside* Codeman means nothing to it. `resolveDockerDaemonMountSource()` translates sources under HOME into the daemon's namespace via `CODEMAN_DOCKER_HOST_HOME`, and `CODEMAN_CASES_PATH` points the cases dir at a host-absolute bind mount so a workspace resolves to the SAME absolute path on both sides (which is what keeps the transcript projHash matching, per Docker cases above). ⚠️ **`CODEMAN_CASES_PATH` must move every consumer or none**: it is resolved once in `config/cases-dir.ts` because `src/cli.ts` resolves case paths too, and when only the server's `CASES_DIR` learned the override, `codeman skill install --case ` reported "Case not found" on exactly the deployment the override exists for. ⚠️ **`.dockerignore` patterns match the WHOLE context-relative path**, so a bare `.env` line excludes only the ROOT file: `docker/.env` (which holds `CODEMAN_PASSWORD` and any provider keys) rode `COPY . .` into the image until `**/.env` was added — verified in both directions with a real build context. ⚠️ A Compose LONG-form bind (`type: bind`) **creates a missing host source directory ROOT-OWNED** rather than refusing. `Start-Codeman.sh` pre-creates both `CODEMAN_APPDATA_PATH` and `CODEMAN_CASES_PATH` on the host before `up`, which is what keeps the daemon from ever having to materialise either as root in the first place; the container ALSO starts as root (`cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the base `cap_drop: ALL`; `test/docker-entrypoint.test.ts` pins that list) so `docker/entrypoint.sh` can correct a bind source that turns up root-owned anyway (a restored backup, a cleared directory, plain `docker compose up` run without the script) before dropping to `PUID:PGID` via `setpriv` — a directory owned by neither root nor `PUID:PGID` is never re-owned, since that ownership is not this container's to reassign; it is PROBED for writability as the runtime account (`setpriv ... test -w`, so ACLs, group-writable trees and CIFS/NFS mounts pass) and refused with a message naming path, owner and PUID:PGID if that fails. ⚠️ `KILL` is in that list for tini, not the entrypoint: `init: true` keeps tini as root while the server runs as PUID, and without CAP_KILL its SIGTERM forward fails and the server is SIGKILLed on every `compose down`/`restart` instead of flushing state. ⚠️ `/opt/codeman-cli` (the runtime-owned CLI prefix) is APPENDED to `PATH`, never prepended, and the entrypoint pins its own `PATH` to the system dirs: the root part of the start resolves `setpriv` by bare name, and a prefix ahead of `/usr/bin` let a planted `setpriv` run as uid 0 (measured). `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` drops `--memory-swap` (and filters only that one kernel warning) for hosts without swap accounting; `--memory` still applies. ⚠️ The deployment ALSO self-updates in place (the repo bind mount at `/opt/codeman` + a restart-by-exiting supervisor) — see Self-update below and `docs/docker-self-update.md` before touching `server.Dockerfile`, the compose file or `.env.example`, since each is an input to the updater's environment gate. `docs/docker-compose.md` + `docker/README.md` (user guides) -**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. **No code outside `stock.ts` may branch on a CLI id**; behaviour that genuinely differs is either a capability field or a NAMED PROFILE selected by one (`profiles.ts`), and `test/cli-registry-no-id-branching.test.ts` fails the build if an id check reappears — it matches `===`, `!==`, `case '':` and `[...].includes(mode)`, because an earlier `===`-only version let 36 negated branches survive the conversion (including a seven-mode Ralph chain whose own comment asked the next person to keep it in step with `isExternalCliMode()` by hand). A second CI-gated guard, `test/frontend-cli-no-id-branching.test.ts`, covers the two frontend files the run-menu consolidation (#458) touched, `session-ui.js` and `mobile-overview.js`: its allowlist is keyed by expression rather than line number, with an occurrence count per entry, so a new branch reusing an already-approved expression fails as a count mismatch instead of riding in on the old approval. ⚠️ Config contains no shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. ⚠️ `external`, `hooks` and `altScreen` are three INDEPENDENT capabilities on purpose; deriving one from another shipped the `until=stop`-hangs-on-shell bug. ⚠️ Two capability fields carry a REGEX from config (`discovery.version.regex` and `capabilities.workDetect.workingLine`) and both must compile through `compileVersionRegex()`, which caps length and refuses nested quantifiers; `workingLine` is the one that runs on the PTY hot path. ⚠️ **`param` is TWO namespaces.** `launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a LAUNCH PARAM; the legacy `Config` wire field is a separate namespace, bridged only by `launch.legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no load error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. Codex is the entry where the two names differ (`bypassApprovals` vs `dangerouslyBypassApprovals`) and therefore the one that catches a regression. ⚠️ Six fields are DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured, so re-measure before wiring one up; the list is pinned so it cannot quietly grow. Spawn commands are pinned as literal strings in `test/cli-registry-spawn-golden.test.ts`, remote/docker pane commands in `test/location-overlay-commands.test.ts`. ⚠️ That second golden no longer covers **remote claude or remote omp**: both now have their own arm in `buildRemoteLaunchCommand` (a `--session-id || --resume` pair, and `--continue`, so a respawn continues the same conversation) and never reach `defaultRemoteCommandForMode`, which is what that test asserts. Their real pins are `toContain` substrings in `test/tmux-manager.test.ts` and `test/remote-shared-sessions.test.ts`; changing either arm will NOT fail the golden. ⚠️ Anything reading the registry resolves it AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks) — a module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. `~/.codeman/clis.json` overrides any entry (read-only in this release; nothing writes it, so importing the registry has no filesystem side effects). → `docs/cli-registry.md` +**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. **No code outside `stock.ts` may branch on a CLI id**; behaviour that genuinely differs is either a capability field or a NAMED PROFILE selected by one (`profiles.ts`), and `test/cli-registry-no-id-branching.test.ts` fails the build if an id check reappears — it matches `===`, `!==`, `case '':` and `[...].includes(mode)`, because an earlier `===`-only version let 36 negated branches survive the conversion (including a seven-mode Ralph chain whose own comment asked the next person to keep it in step with `isExternalCliMode()` by hand). A second CI-gated guard, `test/frontend-cli-no-id-branching.test.ts`, covers the two frontend files the run-menu consolidation (#458) touched, `session-ui.js` and `mobile-overview.js`: its allowlist is keyed by expression rather than line number, with an occurrence count per entry, so a new branch reusing an already-approved expression fails as a count mismatch instead of riding in on the old approval. ⚠️ Config contains no shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. ⚠️ `external`, `hooks` and `altScreen` are three INDEPENDENT capabilities on purpose; deriving one from another shipped the `until=stop`-hangs-on-shell bug. ⚠️ Three capability fields carry a REGEX from config (`discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`) and all three must compile through `compileVersionRegex()`, which caps length and refuses nested quantifiers; `workingLine` is the one that runs on the PTY hot path. ⚠️ **`param` is TWO namespaces.** `launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a LAUNCH PARAM; the legacy `Config` wire field is a separate namespace, bridged only by `launch.legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no load error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. Codex is the entry where the two names differ (`bypassApprovals` vs `dangerouslyBypassApprovals`) and therefore the one that catches a regression. ⚠️ Six fields are DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured, so re-measure before wiring one up; the list is pinned so it cannot quietly grow. Spawn commands are pinned as literal strings in `test/cli-registry-spawn-golden.test.ts`, remote/docker pane commands in `test/location-overlay-commands.test.ts`. ⚠️ That second golden no longer covers **remote claude or remote omp**: both now have their own arm in `buildRemoteLaunchCommand` (a `--session-id || --resume` pair, and `--continue`, so a respawn continues the same conversation) and never reach `defaultRemoteCommandForMode`, which is what that test asserts. Their real pins are `toContain` substrings in `test/tmux-manager.test.ts` and `test/remote-shared-sessions.test.ts`; changing either arm will NOT fail the golden. ⚠️ Anything reading the registry resolves it AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks) — a module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. `~/.codeman/clis.json` overrides any entry (read-only in this release; nothing writes it, so importing the registry has no filesystem side effects). → `docs/cli-registry.md` **External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. ⚠️ **Work detection is no longer part of that gate**: it is per-CLI `capabilities.workDetect` data (see the ❯ note above), so a CLI that declares its own glyph and working line gets the same screen-probed idle confirmation claude gets, and one that declares neither keeps the output-stabilization behaviour. All eight **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. ⚠️ **Codex sessions use PREDICTIVE WRITE-THROUGH echo, never the buffer overlay** (`_localEchoPolicy` in `_updateLocalEchoState`, terminal-ui.js): codex's composer reacts per keystroke ("/" pops a live-filtering picker, arrows edit server-side state, the composer grows as it wraps), so buffer-until-Enter starved it into issues #218/#219/#220/#222 and stays disabled (`_localEchoEnabled` remains false for codex). Instead, `PredictiveEchoAddon` (separate `vendor/xterm-predictive-echo.js` bundle) paints each keystroke at the predicted cell while the wire path stays BYTE-IDENTICAL: the onData hook (`_predictHookOnData`) is a plain statement with no `return`, so control always falls through into the untouched send path — pinned by vm and E2E byte-identity tests. Predictions reconcile against the parsed buffer and only while the cursor sits on the measured composer row (`isCodexComposerRow`, `/^› /`). Codex also **drops keystrokes that share a PTY read with a bracketed paste**, so flushed text and the paste sequence must go out as separate delayed writes (mirroring the Enter branch's delayed `\r`). Tests: `test/local-echo-codex-gating.test.ts`, `test/codex-predictive-echo.test.ts` (E2E vs real codex), `packages/xterm-zerolag-input/test/codex-replay.test.ts`. ⚠️ **Pi is the opposite kind of CLI and needs the opposite instincts**: it has NO permission prompts and no sandbox, so there is no bypass flag to send and Codeman must not invent one; its privileged knob is the tri-state `approveProjectTrust` (`--approve`/`--no-approve`), which makes pi EXECUTE repo-local `.pi/extensions` TypeScript, so the multi-user clamp puts pi in the **materialize** branch (an absent config still yields `--no-approve` for a non-granted owner) and `--api-key` is never wired. Pi stays OUT of `isAltScreenStripMode()` (main-screen TUI, and its 0.84.0 fullscreen mode is runtime-switchable via `/settings`, where the alt screen is load-bearing), and lands on the `'buffer'` echo policy via the `_updateLocalEchoState` fallthrough. Pi's own tests: `test/pi-mode.test.ts`, `test/routes/external-cli-bypass-clamp.test.ts`; user guide `docs/pi-integration.md`. ⚠️ **Grok is codex-shaped on permissions but opencode-shaped on rendering**: its bypass switch is `alwaysApprove` (`--always-approve`, grok's `bypassPermissions` mode — the Run button sends it `true` like antigravity's, and the clamp's only-if-sent branch strips it for non-granted owners), while its fullscreen alt-screen TUI keeps it OUT of `isAltScreenStripMode()`; the resolver version-probes `grok --version` like pi's (npm squatters exist for the name — `GET /api/grok/status` surfaces path + version), and grok lands on the `'buffer'` echo policy via the fallthrough (UNMEASURED against a live authenticated session; if its composer turns out per-keystroke-reactive like codex, flip it to the `'off'` branch). Grok's own tests: `test/grok-mode.test.ts`, `test/grok-cli-resolver.test.ts`; user guide `docs/grok-integration.md`. ⚠️ **DeepSeek breaks three of this family's assumptions, so do not pattern-match it onto its siblings.** (1) The agent is a **PROFILE, not the binary**: `dsh` is a launcher over `$DSH_HOME/profiles/` and DeepSeek ships only `web`/`headless`/`base`, so the terminal front door is ALWAYS third-party and "installed" ≠ "runnable" — the Run button gates on `isDeepSeekRunnable()` (binary AND a pane-capable profile) while `isDeepSeekAvailable()` gates the "add a profile" affordance; a `web`/`headless` profile is refused at spawn because it cannot drive a pane. (2) The permission switch is the **`DSH_PERMISSION_MODE` env export, not a flag** (`read-only`/`workspace-write`/`danger-full-access`) — the harness has none, and this is the one legitimate exception to the effort-style env-var ban because it is read with `??` as a boot-time default, so it stays soft; absent = `workspace-write`, which asks, hence the only-if-sent clamp branch, clamping to `workspace-write` (never `read-only`, which would break the workspace). ⚠️ **That clamp needs a second half no other CLI needs**, because the switch is an env var and `DSH_*` is an allowlisted `envOverrides` prefix: `applyEnvOverrides()` runs AFTER `_configureCliEnv()` in tmux-manager, so a non-granted owner sending `DSH_PERMISSION_MODE` on the SAME request would land last and hand back exactly the privilege the config clamp removed. `clampEnvOverridesForOwner()` (session-routes.ts) DROPS `DSH_PERMISSION_MODE`, `DSH_HOME` and `DEEPSEEK_BASE_URL` for a non-granted owner (the last because `_configureCliEnv()` forwards the SERVER's own `DEEPSEEK_API_KEY` into the pane, so a redirected base URL would send it to a foreign host) (dropping falls through to what `_configureCliEnv()` exports, which is the clamped value); `DSH_HOME` is there because it points the launcher at a profile tree whose plugin code runs at BOOT, before any approval row applies. Every OTHER CLI's bypass is a command-line flag reachable only through its config, which is why the config clamp alone is the whole gate for them. (3) It is the **only non-claude mode that passes `hooksAvailableForMode()`**, and for it alone that predicate is a per-SESSION question rather than a per-mode one (`deepSeekConfig.statusReporting: false` disarms the bridge, so every call site passes `sessionHookOptions(session)`; answering from the mode there re-creates the infinite-wait-dressed-as-a-timeout the guard exists to prevent). It passes because the terminal front door reports idle/working/blocked to a supervisor over a generic env-gated contract and `deepseek-status-shim.ts` makes Codeman that supervisor — real `stop`/`blocked` signals, real Approvals Inbox items, plus the `agent_working` event that clears an alert answered in the terminal. ⚠️ The resolver needs the strictest identity probe of the family (`dsh --help` must say `DeepSeek Harness`) because Debian ships an unrelated `dsh` (dancer's shell) that would pass a version probe. Model is NOT a session field (it is a profile composition entry). ⚠️ `hooksAvailableForMode()` is about hook SIGNALS and is not a stand-in for "is this a claude session": Read My Mind and intent capture read Claude's own transcript and compare `mode === 'claude'` directly, because when `deepseek` earned a yes the shared predicate silently widened both to a mode with no transcript to read (pinned by a static check in `test/deepseek-mode.test.ts`). ⚠️ **It is also the only external CLI whose answers are READ FROM DISK rather than scraped off the pane**: `deepseek-transcript.ts` reads `$DSH_HOME/sessions///session.jsonl.zstd` and backs the `last-response` route for dsh, because the pane segmenter served dsh-TUI's ASCII-art SPLASH as the worker's answer (measured), which anything polling for a first answer reads as an answer. Three traps live in that file: dsh appends **one zstd FRAME per write** and Node's `zlib` zstd decoder stops at the first (a real 56-line transcript decoded as 1 line, so the module walks frame headers itself; a Node older than 22.15 has no zstd and falls back to the pane); every turn also records a **plugin-sourced `user/message`** (the runtime-context snapshot) that must not render as the user's words; and a failed `turn/end` is surfaced as `Turn error: …` rather than as an empty string that reads as "still thinking". ⚠️ Session→transcript pairing is by the header's own `cwd` plus a ±60 s boot window, never by reproducing dsh's directory mangling (which has already changed form once) — and NEVER by newest-mtime alone, which handed a fresh worker its predecessor's answer in the same case dir. DeepSeek's own tests: `test/deepseek-mode.test.ts`, `test/deepseek-cli-resolver.test.ts`, `test/deepseek-transcript.test.ts`; user guide `docs/deepseek-integration.md`. OMP (`omp`) needs no bypass flag (the CLI's own `~/.omp` config governs trust/model routing, defaulting to `tools.approvalMode: yolo`), so its registry entry declares `privilegedParams: []` and a launch spec that only ever passes `--model`/`--resume`/`--continue` — but the multi-user clamp is NOT a no-op for it: `OMP_*` is an allowlisted `envOverrides` prefix, and the two credential-resolution keys it admits, `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN`, are clamped in `clampEnvOverridesForOwner()` for a non-granted owner, the same shape as `DEEPSEEK_BASE_URL`. Separately, `PI_*` is already allowlisted (pi needs it) and omp reads several of its knobs too (`PI_CONFIG_DIR`, `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`, `PI_SUBPROCESS_CMD`, `PI_SHELL_PREFIX`) — a redirected `PI_CONFIG_DIR` moves the `~/.omp` tree `omp-session-resolver.ts`/`omp-transcript.ts` hardcode, silently breaking pinning/history; this is a known gap shared with pi, not fixed here. → [architecture-invariants#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp) diff --git a/docs/cli-registry.md b/docs/cli-registry.md index 22168cea..7e15e525 100644 --- a/docs/cli-registry.md +++ b/docs/cli-registry.md @@ -36,7 +36,8 @@ interface CliEntry { launch: CliLaunch; // the structured argv template env: CliEnv; // exports, tmux setenv keys, the env-override allowlist capabilities: CliCapabilities; // what every call site reads instead of the id - // .workDetect?: { promptGlyph, workingLine } — how this CLI's pane shows work + // .workDetect?: { promptGlyph, workingLine, watchingLine? } — how this CLI's pane + // shows work, and how it shows work it started in the background overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store } ``` @@ -45,10 +46,19 @@ interface CliEntry { ### Regexes that come from config -Two capability fields carry a regular expression an override file can set: `discovery.version.regex` and `capabilities.workDetect.workingLine`. Both go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing. +Three capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`. All three go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing. `workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused. +`watchingLine` reads a different row of the same screen. A CLI draws it while work the agent +itself started is still running — Claude prints `⏵⏵ bypass permissions on · 1 monitor · ← for +agents` while a monitor, a backgrounded shell or a cloud session is live — and Codeman shows +that as the session's watching badge, so a quiet pane waiting for its own background work +does not read as a pane waiting for a human. `watchingLabel()` in `session-activity.ts` runs +the pattern over the last few lines of a capture only, because the transcript above the +composer quotes arbitrary text and a session that PRINTS "1 monitor" is not running one. +Group 1 is the label, and a CLI that declares no pattern reports no background work. + ### Three capabilities that must stay independent `external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks. diff --git a/src/config/cli-registry/schema.ts b/src/config/cli-registry/schema.ts index d1249384..21761c8d 100644 --- a/src/config/cli-registry/schema.ts +++ b/src/config/cli-registry/schema.ts @@ -308,6 +308,16 @@ const capabilitiesSchema = z (src) => compileVersionRegex(src) !== null, 'workingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers' ), + // Same guard, same reasons: this one runs over the foot of a pane capture every + // time a session settles, and ~/.codeman/clis.json can set it. + watchingLine: z + .string() + .min(1) + .refine( + (src) => compileVersionRegex(src) !== null, + 'watchingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers' + ) + .optional(), }) .strict() .optional(), diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index 1e455192..2c2de234 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -206,6 +206,11 @@ const CLAUDE: CliEntry = { workDetect: { promptGlyph: '❯', workingLine: String.raw`…\s*\((?:\d+h\s+)?(?:\d+m\s+)?\d+s\b|esc to interrupt`, + // Claude prints what it started in the background on the footer row beneath its + // composer, as `⏵⏵ bypass permissions on · 1 monitor · ← for agents`. The labels are the + // CLI's own words for each kind of background task, and group 1 is the one Codeman + // badges the session with. Verified against a live 2.1.278 pane on 2026-09-21. + watchingLine: String.raw`(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?|Artifact comment monitors?))`, }, requiresMux: false, // Claude installs Codeman's own hooks block into every workspace it runs in, so its diff --git a/src/config/cli-registry/types.ts b/src/config/cli-registry/types.ts index 871a3762..e5f29db8 100644 --- a/src/config/cli-registry/types.ts +++ b/src/config/cli-registry/types.ts @@ -344,6 +344,14 @@ export interface CliCapabilities { promptGlyph: string; /** Source of a regex matching the status line this CLI draws while a turn runs. */ workingLine: string; + /** + * Source of a regex matching the chip this CLI draws at the foot of its screen while + * work it started in the background is still running, e.g. Claude's `· 1 monitor ·`. + * Capture group 1 is the label Codeman shows, and the whole match stands in when the + * pattern declares no group. A CLI that omits this reports no background work, which + * is what every CLI did before the field existed. + */ + watchingLine?: string; }; /** No direct-PTY fallback: the CLI must run inside tmux (secrets ride tmux setenv). */ requiresMux: boolean; diff --git a/src/session-activity.ts b/src/session-activity.ts index 324aff1b..e593aacb 100644 --- a/src/session-activity.ts +++ b/src/session-activity.ts @@ -91,3 +91,42 @@ export function isSustainedActivity(streak: ActivityStreak | null, streakMs: num export function isPaneQuiet(lastActivityAt: number, now: number, silenceMs: number = IDLE_SILENCE_MS): boolean { return now - lastActivityAt >= silenceMs; } + +/** + * How many lines at the foot of a pane capture may hold the background-work chip. + * + * Claude Code draws that chip on the last row of the screen, under its composer box + * and under whatever status line the user configured, so five lines reach it with + * room to spare. The ceiling is the point of the constant: the transcript above the + * composer quotes arbitrary text, and a session that PRINTS the words "1 monitor" + * must not be read as running one. + */ +export const WATCHING_TAIL_LINES = 5; + +/** + * What a pane says is still running in the background, e.g. `1 monitor` or `2 shells`. + * + * The CLI writes that chip while a monitor, a backgrounded shell or a cloud session it + * started is still going, which is exactly the case where the agent has ended its turn + * without wanting anything from the user. `pattern` comes from the CLI's own registry + * entry (`capabilities.workDetect.watchingLine`); group 1 is the label when the pattern + * declares one, and the whole match stands in when it does not. + * + * @returns the label, or null when the pane shows no background work + */ +export function watchingLabel(paneText: string | null | undefined, pattern: RegExp): string | null { + if (!paneText) return null; + const lines = paneText + .split('\n') + .map((line) => line.trimEnd()) + .filter((line) => line !== ''); + if (lines.length === 0) return null; + // A pattern compiled by compileVersionRegex() never carries the `g` flag, but a caller + // reaching in from a test or a config reload might, and a stale lastIndex would make + // the same screen match every other call. + pattern.lastIndex = 0; + const match = pattern.exec(lines.slice(-WATCHING_TAIL_LINES).join('\n')); + if (!match) return null; + const label = (match[1] ?? match[0]).trim(); + return label === '' ? null : label; +} diff --git a/src/session.ts b/src/session.ts index fdad22f8..3ab12908 100644 --- a/src/session.ts +++ b/src/session.ts @@ -81,6 +81,7 @@ import { trackActivityStreak, isSustainedActivity, isPaneQuiet, + watchingLabel, IDLE_RECHECK_MS, PANE_PROBE_MIN_INTERVAL_MS, PANE_PROBE_RECHECK_MS, @@ -497,8 +498,11 @@ export class Session extends EventEmitter { private _activityStreak: ActivityStreak | null = null; // Unbroken run of PTY repaints (working detection) private _lastPaneProbeAt = 0; // Throttle for the tmux screen probe private _lastPaneProbeWorking: boolean | null = null; // Its last verdict (null = could not read) + private _watching: string | null = null; // Background work the pane's own footer reports /** Lazily compiled `capabilities.workDetect.workingLine`. See _workingLinePattern(). */ private _workingLineRe: RegExp | undefined = undefined; + /** Lazily compiled `capabilities.workDetect.watchingLine`. See _watchingLinePattern(). */ + private _watchingLineRe: RegExp | null | undefined = undefined; private _trustDialogAccepted: boolean = false; // Stops the trust-dialog scan (answered, or given up) private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read @@ -1118,6 +1122,16 @@ export class Session extends EventEmitter { return this._isWorking; } + /** + * What the pane says is still running in the background, e.g. `1 monitor`, or null when + * nothing is. A session with a label here has ended its turn without wanting anything + * from the user, so a surface that would otherwise file it under "needs you" can say + * what it is waiting for instead. + */ + get watching(): string | null { + return this._watching; + } + /** * Check if the session's process tree has active child processes beyond Claude itself. * Detects running bash tools, test suites, builds, servers, etc. that Claude spawned. @@ -1675,6 +1689,7 @@ export class Session extends EventEmitter { totalCost: this._totalCost, messageCount: this._messages.length, isWorking: this._isWorking, + watching: this._watching, lastPromptTime: this._lastPromptTime, // Buffer statistics for monitoring long-running sessions bufferStats: { @@ -2709,9 +2724,43 @@ export class Session extends EventEmitter { this._lastPaneProbeAt = now; const text = this._mux.capturePaneText?.(this._muxSession.muxName) ?? null; this._lastPaneProbeWorking = text === null ? null : this._workingLinePattern().test(text); + this._readWatching(text); return this._lastPaneProbeWorking; } + /** + * Read the background-work chip off the same capture the working probe just took. + * + * The two questions are different. A turn that is running is work the user is waiting + * for; a monitor, a backgrounded shell or a cloud session the agent started is work + * the AGENT is waiting for, and it is the reason a pane can sit at its composer with + * nothing to say and still not want anything from the user. `_confirmIdle` takes this + * capture at exactly the moment the turn ends, which is the moment the answer starts + * mattering. + * + * A capture that could not be read leaves the last answer standing, the way the + * working probe treats its own null: no evidence is not evidence of none. + */ + private _readWatching(paneText: string | null): void { + const pattern = this._watchingLinePattern(); + if (!pattern || paneText === null) return; + this._watching = watchingLabel(paneText, pattern); + } + + /** + * The regex matching this CLI's background-work chip, or null for a CLI whose registry + * entry declares none. Compiled once per session, like the working-line pattern, and + * null rather than a fallback: no other CLI has been measured drawing such a chip, and + * guessing one would badge sessions on the strength of an unread screen. + */ + private _watchingLinePattern(): RegExp | null { + if (this._watchingLineRe === undefined) { + const src = getCli(this.mode)?.capabilities.workDetect?.watchingLine; + this._watchingLineRe = src ? compileVersionRegex(src) : null; + } + return this._watchingLineRe; + } + /** * The regex matching this CLI's "a turn is running" status line. * @@ -3963,6 +4012,9 @@ export class Session extends EventEmitter { async stop(killMux: boolean = true): Promise { // Set stopped flag first to prevent new timers from being created this._isStopped = true; + // A pane that is gone is watching nothing. Nothing probes a stopped session, so + // without this the last chip it drew would ride along on its row forever. + this._watching = null; this._clearAllTimers(); diff --git a/src/web/public/app.js b/src/web/public/app.js index 4b27df17..2797023a 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -4574,6 +4574,11 @@ class CodemanApp { return { state, pill: this._sidebarRichPillLabel(state), + // What the pane's own footer says is still running in the background ("1 monitor", + // "2 shells"). A row that has one went quiet because the agent is waiting for that, + // which is a different thing from waiting for the user — so it rides BESIDE the + // state pill and never replaces it. + watching: typeof session.watching === 'string' ? session.watching : '', createdAt: Number(session.createdAt) || 0, since: this._mobileOverviewSince ? this._mobileOverviewSince(state, session) : null, }; @@ -4609,6 +4614,13 @@ class CodemanApp { parts.push(stamp(row.since.key, row.since.at, 'for', 'tab-meta-since')); } parts.push(`${escapeHtml(row.pill)}`); + // The word is duplicated from mobile-overview.js for the same reason the pill labels + // above are: it is one word, and this file must render a complete row even when a + // stale cached mobile-overview.js has arrived without it. + if (row.watching) { + const title = escapeHtml(`Still running in the background: ${row.watching}`); + parts.push(`watching`); + } // Both absolute stamps ALSO on the line itself, not only on the two items. // Below 288px the rail hides `.tab-meta-created` (the `tab-rail-tight` // rule), and a tooltip on a `display: none` element has no hover target — @@ -4646,7 +4658,7 @@ class CodemanApp { const prev = tab.dataset.tabState; // The since ANCHOR moves without the state changing (each new turn re-stamps // lastSubmitAt), so key the compare on both. - const sig = `${row.state}:${row.since ? row.since.at : 0}:${row.createdAt}`; + const sig = `${row.state}:${row.since ? row.since.at : 0}:${row.createdAt}:${row.watching}`; if (tab.dataset.tabMetaSig === sig) return; tab.dataset.tabMetaSig = sig; tab.dataset.tabState = row.state; diff --git a/src/web/public/home-sessions.js b/src/web/public/home-sessions.js index 453eec10..66430d9a 100644 --- a/src/web/public/home-sessions.js +++ b/src/web/public/home-sessions.js @@ -212,6 +212,9 @@ Object.assign(CodemanApp.prototype, { dir: this._shortenHomePath ? this._shortenHomePath(session.workingDir) : session.workingDir || '', state, pill: HOME_SESSIONS_PILL_LABEL[state] || state, + // What the pane's footer says is still running in the background, straight off + // the session payload. Same field, same meaning as on the phone overview. + watching: typeof session.watching === 'string' ? session.watching : '', // Epoch ms, straight off the session payload; formatting happens at // render time so the clock below can redo it without a re-render. createdAt: Number(session.createdAt) || 0, @@ -450,6 +453,12 @@ Object.assign(CodemanApp.prototype, { // what stops it ellipsizing. const meta = this._buildHomeSessionsMeta(row); meta.appendChild(pill); + // Built by the phone overview so both home screens word the badge identically. + // Guarded like every other cross-file call here: a stale cached mobile-overview.js + // must cost the badge, not the rail. + if (row.watching && typeof this._buildWatchingBadge === 'function') { + meta.appendChild(this._buildWatchingBadge(row.watching, 'home-sessions-pill')); + } item.appendChild(meta); return item; diff --git a/src/web/public/mobile-overview.js b/src/web/public/mobile-overview.js index 606bc1f1..49b5ed1c 100644 --- a/src/web/public/mobile-overview.js +++ b/src/web/public/mobile-overview.js @@ -61,6 +61,14 @@ const MOBILE_OVERVIEW_RUN_MODES = [ ]; /** Pill copy per state. Kept short: a phone row has ~90px for it. */ +/** + * The one word every surface puts on the watching badge, and the tooltip that says + * what the pane actually reported. Both live here so the phone overview, the desktop + * home rail and the rich sidebar rows cannot word the same badge three ways. + */ +const WATCHING_BADGE_TEXT = 'watching'; +const watchingBadgeTitle = (label) => 'Still running in the background: ' + label; + const MOBILE_OVERVIEW_PILL_LABEL = { needs: 'needs you', error: 'error', @@ -185,6 +193,10 @@ Object.assign(CodemanApp.prototype, { dir: this._shortenHomePath ? this._shortenHomePath(session.workingDir) : session.workingDir || '', state, pill: MOBILE_OVERVIEW_PILL_LABEL[state] || state, + // What the pane's own footer says is still running in the background ("1 monitor", + // "2 shells"), straight off the session payload. A row that has one is quiet + // because the agent is waiting for that, not because it is waiting for you. + watching: typeof session.watching === 'string' ? session.watching : '', // Epoch ms, straight off the session payload; formatting happens at // render time so the clock can redo it without a re-render. createdAt: Number(session.createdAt) || 0, @@ -713,6 +725,8 @@ Object.assign(CodemanApp.prototype, { pill.textContent = row.pill; item.appendChild(pill); + if (row.watching) item.appendChild(this._buildWatchingBadge(row.watching, 'mobile-overview-pill')); + const chevron = document.createElement('span'); chevron.className = 'mobile-overview-chevron'; chevron.setAttribute('aria-hidden', 'true'); @@ -734,6 +748,35 @@ Object.assign(CodemanApp.prototype, { return item; }, + // ═══════════════════════════════════════════════════════════════ + // Watching badge + // ═══════════════════════════════════════════════════════════════ + + /** + * The badge a session wears while work it started in the background is still + * running: a monitor, a backgrounded shell, a cloud session. + * + * It says one word, and the label the pane itself printed ("1 monitor") rides in + * the tooltip, because the badge shares a row with the state pill on the narrowest + * screen this app renders on. It does NOT replace that pill: an agent can arm a + * monitor and ask the user a question in the same breath, so the row still says + * "needs you" and this says what else is going on. + * + * Shared with the desktop home rail (home-sessions.js), for the same reason + * `_mobileOverviewState` is: one badge, one wording, one place to change it. The + * caller names its own pill class, because each surface styles its pills itself and + * the phone's live inside a media query the desktop never enters. + */ + _buildWatchingBadge(label, baseClass) { + const badge = document.createElement('span'); + const base = baseClass || 'mobile-overview-pill'; + badge.className = base + ' ' + base + '--watching'; + badge.setAttribute('data-i18n-skip', ''); + badge.textContent = WATCHING_BADGE_TEXT; + badge.title = watchingBadgeTitle(label); + return badge; + }, + // ═══════════════════════════════════════════════════════════════ // Age stamps: started / how long in this state // ═══════════════════════════════════════════════════════════════ diff --git a/src/web/public/mobile.css b/src/web/public/mobile.css index 125e376b..ad582b2a 100644 --- a/src/web/public/mobile.css +++ b/src/web/public/mobile.css @@ -2961,6 +2961,13 @@ html.mobile-init .file-browser-panel { color: var(--green); } + /* Accent, and none of the three above: a session watching work it started itself + is not asking the user for anything, and red and yellow are what say it is. */ + .mobile-overview-pill--watching { + border-color: var(--accent); + color: var(--accent); + } + .mobile-overview-chevron { flex-shrink: 0; color: var(--text-muted); diff --git a/src/web/public/styles.css b/src/web/public/styles.css index e0cc8e4a..250b81c7 100644 --- a/src/web/public/styles.css +++ b/src/web/public/styles.css @@ -16249,6 +16249,17 @@ html[data-tab-orientation='vertical'] .home-sessions { color: color-mix(in srgb, var(--green) 45%, var(--text-muted)); } +/* Accent, deliberately none of the three above: a session that is watching something + it started (a monitor, a backgrounded shell, a cloud session) is not asking for + anything, so it must not borrow the red or the yellow that mean it is. This badge + sits BESIDE the state pill rather than replacing it, since an agent can arm a + monitor and ask a question in the same breath. */ +.home-sessions-pill--watching { + background: color-mix(in srgb, var(--accent) 14%, transparent); + border-color: color-mix(in srgb, var(--accent) 40%, transparent); + color: var(--accent); +} + /* Row accents: same language as the session tabs and the phone overview — red means a question is pending, yellow means it wants input, green means work is happening. Nothing else on this screen may reuse these colors. */ @@ -18311,6 +18322,16 @@ html[data-tab-orientation='vertical'][data-tab-rail-detail='rich']:not(.tab-rail color: color-mix(in srgb, var(--green) 45%, var(--text-muted)); } +/* Accent, deliberately none of the three above: a watching session is waiting for + work it started itself, not for the user, and the red and yellow here are spoken + for by sessions that ARE waiting for the user. */ +html[data-sidebar-detail="rich"] .session-sidebar .tab-pill--watching, +html[data-tab-orientation='vertical'][data-tab-rail-detail='rich']:not(.tab-rail-compact) .tab-rail .tab-pill--watching { + background: color-mix(in srgb, var(--accent) 14%, transparent); + border-color: color-mix(in srgb, var(--accent) 40%, transparent); + color: var(--accent); +} + /* Three lines of content per row instead of two, so give them room to breathe and stop the row actions crowding the pill. diff --git a/test/cli-registry-schema.test.ts b/test/cli-registry-schema.test.ts index 762eff92..fa7c70b5 100644 --- a/test/cli-registry-schema.test.ts +++ b/test/cli-registry-schema.test.ts @@ -122,6 +122,24 @@ describe('workDetect.workingLine is guarded like every other config regex', () = expect(compileVersionRegex(src), `${entry.id} declares a workingLine the guard refuses`).not.toBeNull(); } }); + + it('holds the optional watchingLine to the same guard', () => { + expectRejected((e) => { + (e.capabilities as Record).workDetect = { + promptGlyph: '>', + workingLine: 'working', + watchingLine: '(a+)+b', + }; + }, 'this one runs over a pane capture every time a session settles, so it can freeze the event loop the same way'); + }); + + it('accepts every shipped watchingLine', () => { + for (const entry of STOCK_CLIS) { + const src = entry.capabilities.workDetect?.watchingLine; + if (!src) continue; + expect(compileVersionRegex(src), `${entry.id} declares a watchingLine the guard refuses`).not.toBeNull(); + } + }); }); describe('no shell text can reach the command line', () => { diff --git a/test/home-sessions.test.ts b/test/home-sessions.test.ts index f4bcccaf..67da3e55 100644 --- a/test/home-sessions.test.ts +++ b/test/home-sessions.test.ts @@ -356,3 +356,50 @@ describe('home screens: one order, one numbering', () => { expect(branch).not.toContain('this.sessionOrder[idx]'); }); }); + +describe('home sessions column: watching badge', () => { + it('carries the label off the session payload', () => { + const app = loadHomeSessionsApp({ + sessions: sessionMap([{ id: 'watcher', watching: '1 monitor' }, { id: 'plain' }]), + sessionOrder: ['watcher', 'plain'], + cases: CASES, + }); + + const rows = Object.fromEntries(app.buildHomeSessionRows().map((r: any) => [r.id, r.watching])); + expect(rows).toEqual({ watcher: '1 monitor', plain: '' }); + }); + + it("builds the badge with the phone overview, in the rail's own pill class", () => { + // Both home screens must word this badge identically, so the rail borrows the + // builder rather than writing a second one — the same arrangement it already has + // for state classification and for the stamp wording. + const app = loadHomeSessionsApp({ + sessions: sessionMap([{ id: 'watcher', watching: '2 shells' }]), + sessionOrder: ['watcher'], + cases: CASES, + }); + + const row = app._buildHomeSessionRow(app.buildHomeSessionRows()[0]); + const badges = collect(row).filter((el: any) => String(el.className).includes('--watching')); + expect(badges).toHaveLength(1); + expect(badges[0].className).toBe('home-sessions-pill home-sessions-pill--watching'); + expect(badges[0].textContent).toBe('watching'); + expect(badges[0].title).toBe('Still running in the background: 2 shells'); + }); + + it('draws no badge for a session running nothing in the background', () => { + const app = loadHomeSessionsApp({ + sessions: sessionMap([{ id: 'plain' }]), + sessionOrder: ['plain'], + cases: CASES, + }); + + const row = app._buildHomeSessionRow(app.buildHomeSessionRows()[0]); + expect(collect(row).filter((el: any) => String(el.className).includes('--watching'))).toHaveLength(0); + }); +}); + +/** Every node under one the builders made, the row itself included. */ +function collect(node: any): any[] { + return [node, ...(node.children || []).flatMap((child: any) => collect(child))]; +} diff --git a/test/mobile-overview.test.ts b/test/mobile-overview.test.ts index 232b3d35..6c6420c8 100644 --- a/test/mobile-overview.test.ts +++ b/test/mobile-overview.test.ts @@ -485,3 +485,50 @@ describe('mobile overview run picker (CLI availability gating)', () => { expect(gate).toContain('isCliAvailable'); }); }); + +describe('mobile overview watching badge', () => { + it('carries what the pane says is running in the background', () => { + const app = loadOverviewApp(); + const model = app.buildMobileOverviewModel({ + sessions: [session({ id: 'a', status: 'idle', watching: '1 monitor' }), session({ id: 'b', status: 'idle' })], + cases: CASES, + }); + + const rows = Object.fromEntries(model.current.map((r: any) => [r.id, r.watching])); + expect(rows).toEqual({ a: '1 monitor', b: '' }); + }); + + it('leaves a session that is ALSO blocked on a dialog in NEEDS YOU', () => { + // An agent can arm a monitor and ask the user a question in the same breath, so the + // badge adds a fact to the row and never moves it out of the group that says a human + // is needed. Only the row's own pill decides that. + const app = loadOverviewApp(); + const model = app.buildMobileOverviewModel({ + sessions: [session({ id: 'a', status: 'idle', watching: '2 shells' })], + cases: CASES, + pendingHooks: new Map([['a', new Set(['permission_prompt'])]]), + }); + + expect(model.needsYou.map((r: any) => r.id)).toEqual(['a']); + expect(model.needsYou[0].pill).toBe('needs you'); + expect(model.needsYou[0].watching).toBe('2 shells'); + }); + + it('says one word and puts the detail in the tooltip', () => { + const app = loadOverviewApp(); + const badge = app._buildWatchingBadge('1 monitor', 'mobile-overview-pill'); + + expect(badge.className).toBe('mobile-overview-pill mobile-overview-pill--watching'); + expect(badge.textContent).toBe('watching'); + expect(badge.title).toBe('Still running in the background: 1 monitor'); + }); + + it('takes the pill class of whichever surface asks for it', () => { + // The phone's pill styles live inside a media query the desktop rail never enters, + // so the rail passes its own base class and gets the same badge in its own clothes. + const app = loadOverviewApp(); + expect(app._buildWatchingBadge('1 shell', 'home-sessions-pill').className).toBe( + 'home-sessions-pill home-sessions-pill--watching' + ); + }); +}); diff --git a/test/session-sidebar-ux.test.ts b/test/session-sidebar-ux.test.ts index 8ea2d0d6..f74c1ace 100644 --- a/test/session-sidebar-ux.test.ts +++ b/test/session-sidebar-ux.test.ts @@ -73,3 +73,35 @@ describe('vertical session navigation UX contract', () => { expect(i18n).toContain("'Adjust only session names in the vertical sidebar.':"); }); }); + +describe('watching badge on a rich session row', () => { + it('reads the label off the session payload', () => { + expect(app).toContain("watching: typeof session.watching === 'string' ? session.watching : ''"); + }); + + it('renders it beside the state pill rather than in place of it', () => { + // A session can be watching a monitor AND holding a question for the user, so the + // pill that says which one still decides the row; this badge only adds a fact. + const meta = app.slice(app.indexOf('_sidebarRichMetaHTML(row) {')); + const body = meta.slice(0, meta.indexOf('_sidebarRichStampText(timestamp, format) {')); + expect(body).toContain('tab-pill tab-pill--${escapeHtml(row.state)}'); + expect(body).toContain('tab-pill tab-pill--watching'); + expect(body).toContain('Still running in the background:'); + }); + + it('re-renders the row when the background work changes', () => { + // The meta line is rebuilt only when this signature moves, so a badge left out of + // it would appear and disappear a render late, or not at all. + expect(app).toContain( + 'const sig = `${row.state}:${row.since ? row.since.at : 0}:${row.createdAt}:${row.watching}`' + ); + }); + + it('colours it with the accent, never with the two colours that mean a human is needed', () => { + const rule = styles.slice(styles.indexOf('.tab-pill--watching')); + const block = rule.slice(0, rule.indexOf('}')); + expect(block).toContain('var(--accent)'); + expect(block).not.toContain('var(--red)'); + expect(block).not.toContain('var(--yellow)'); + }); +}); diff --git a/test/session-watching.test.ts b/test/session-watching.test.ts new file mode 100644 index 00000000..e465a063 --- /dev/null +++ b/test/session-watching.test.ts @@ -0,0 +1,220 @@ +/** + * The badge a session wears while work it started in the background is still running. + * + * The bug this pins: an agent that arms a monitor, backgrounds a shell or hands work to a + * cloud session is told to end its turn, so the pane falls quiet, Claude Code's idle + * notification arrives a minute later, and every Codeman surface files the session under + * NEEDS YOU. Nothing wants the user there. The CLI itself says so on the last row of its + * screen (`⏵⏵ bypass permissions on · 1 monitor · ← for agents`), and reading that row is + * what tells a session waiting for its own background work from one waiting for a human. + * + * The pane fixtures below are verbatim captures (`tmux -L codeman capture-pane -p`) from a + * live Claude Code 2.1.278 session on 2026-09-21. + */ +import { describe, expect, it, vi, afterEach } from 'vitest'; +import { Session } from '../src/session.js'; +import { getCli } from '../src/config/cli-registry/index.js'; +import { compileVersionRegex } from '../src/config/cli-registry/patterns.js'; +import { watchingLabel, WATCHING_TAIL_LINES, IDLE_SILENCE_MS } from '../src/session-activity.js'; + +/** The registry's own pattern for Claude, which is what every consumer runs. */ +const CLAUDE_WATCHING = compileVersionRegex(getCli('claude')!.capabilities.workDetect!.watchingLine!)!; + +/** The bottom of a Claude pane: composer, the user's status line, the footer row. */ +function pane(footer: string, body = ''): string { + return ( + body + + '╭──────────────────────────────────────╮\n' + + '│ ❯ │\n' + + '╰──────────────────────────────────────╯\n' + + ' ~/innovi/gtd-board [main] Opus 5 ctx: 11%\n' + + ` ${footer}\n` + ); +} + +const WITH_MONITOR = pane('⏵⏵ bypass permissions on · 1 monitor · ← for agents'); +const WITH_SHELL = pane('⏵⏵ bypass permissions on · 1 shell · ← for agents'); +const NOTHING_RUNNING = pane('⏵⏵ bypass permissions on (shift+tab to cycle) · ← for agents'); + +/** A composer repaint: the frame Claude ships roughly once a second while working. */ +const COMPOSER_REPAINT = + '\x1b[31;1H\x1b[38;5;246m❯\xa0\x1b[39m\x1b[0m\x1b[33;1H \x1b[38;5;246mOpus 5 in:143,699 out:669 ctx:14%\x1b[39m'; + +type SessionInternals = { + _handleTerminalOutput(data: string): void; + _detectInteractiveActivity(data: string): void; +}; + +/** One PTY chunk, exactly as the interactive handler sees it. */ +function feed(session: Session, data: string): void { + const internals = session as unknown as SessionInternals; + internals._handleTerminalOutput(data); + internals._detectInteractiveActivity(data); +} + +/** A session whose mux reports a fixed (or scripted) screen for the pane probe to read. */ +function withFakePane(screen: string | (() => string), mode: 'claude' | 'codex' = 'claude'): Session { + const read = typeof screen === 'function' ? screen : () => screen; + const mux = { + isAvailable: () => true, + capturePaneText: () => read(), + } as unknown as NonNullable[0]>['mux']; + return new Session({ + workingDir: '/tmp', + mode, + mux, + muxSession: { muxName: 'codeman-test', sessionId: 'test', createdAt: Date.now() }, + } as ConstructorParameters[0]); +} + +/** Run one turn and let it end, which is when the probe reads the screen. */ +function runAndSettle(session: Session): void { + for (let i = 0; i < 3; i++) { + feed(session, COMPOSER_REPAINT); + vi.advanceTimersByTime(1000); + } + vi.advanceTimersByTime(IDLE_SILENCE_MS + 2000); +} + +describe('watchingLabel', () => { + it('reads the label off the footer row', () => { + expect(watchingLabel(WITH_MONITOR, CLAUDE_WATCHING)).toBe('1 monitor'); + expect(watchingLabel(WITH_SHELL, CLAUDE_WATCHING)).toBe('1 shell'); + }); + + it('reads every kind of background work the CLI names', () => { + const labels = [ + '2 monitors', + '3 shells', + '1 cloud session', + '2 cloud sessions', + '1 local agent', + '4 background tasks', + '1 MCP task', + '1 background dynamic workflow', + '2 remote dynamic workflows', + '1 Artifact comment monitor', + '2 teams', + ]; + for (const label of labels) { + expect(watchingLabel(pane(`⏵⏵ bypass permissions on · ${label} · ← for agents`), CLAUDE_WATCHING)).toBe(label); + } + }); + + it('says nothing about a pane that is running nothing', () => { + expect(watchingLabel(NOTHING_RUNNING, CLAUDE_WATCHING)).toBeNull(); + expect(watchingLabel('', CLAUDE_WATCHING)).toBeNull(); + expect(watchingLabel(null, CLAUDE_WATCHING)).toBeNull(); + }); + + it('ignores the same words in the transcript above the composer', () => { + // The whole reason the search is confined to the foot of the screen: a session that + // PRINTS "1 monitor" (this one has been discussing exactly that) is not running one. + const transcript = + '> does Codeman know about watching?\n' + + '⏺ The footer says 1 monitor while a monitor is armed, and 2 shells for\n' + + ' backgrounded commands. Codeman reads neither today.\n' + + ' Nothing else on the screen means background work is running.\n'; + expect(watchingLabel(pane('⏵⏵ bypass permissions on · ← for agents', transcript), CLAUDE_WATCHING)).toBeNull(); + }); + + it('looks no further up the screen than the tail it declares', () => { + const chip = '⏵⏵ bypass permissions on · 1 monitor · ← for agents'; + const blanks = Array(WATCHING_TAIL_LINES).fill(' still here').join('\n'); + // Blank lines are dropped before the tail is taken, so padding with them must not + // push the footer out of range. + expect(watchingLabel(`${chip}\n\n\n\n\n\n`, CLAUDE_WATCHING)).toBe('1 monitor'); + expect(watchingLabel(`${chip}\n${blanks}\n`, CLAUDE_WATCHING)).toBeNull(); + }); + + it('survives a pattern handed to it with the global flag set', () => { + // compileVersionRegex() never sets `g`, but a test or a reloaded config might, and a + // sticky lastIndex would make the same screen match every other call. + const global = new RegExp(CLAUDE_WATCHING.source, 'g'); + expect(watchingLabel(WITH_MONITOR, global)).toBe('1 monitor'); + expect(watchingLabel(WITH_MONITOR, global)).toBe('1 monitor'); + }); +}); + +describe('Session.watching', () => { + afterEach(() => { + vi.useRealTimers(); + }); + + it('carries what the pane reported once the turn ends', () => { + vi.useFakeTimers(); + const session = withFakePane(WITH_MONITOR); + expect(session.watching).toBeNull(); + + runAndSettle(session); + + expect(session.status).toBe('idle'); + expect(session.watching).toBe('1 monitor'); + }); + + it('lets the badge go when the background work is over', () => { + vi.useFakeTimers(); + const screen = { text: WITH_MONITOR }; + const session = withFakePane(() => screen.text); + + runAndSettle(session); + expect(session.watching).toBe('1 monitor'); + + screen.text = NOTHING_RUNNING; + runAndSettle(session); + expect(session.watching).toBeNull(); + }); + + it('keeps its last answer when the screen cannot be read', () => { + vi.useFakeTimers(); + const screen: { text: string | null } = { text: WITH_MONITOR }; + const session = withFakePane(() => screen.text as string); + + runAndSettle(session); + expect(session.watching).toBe('1 monitor'); + + // A capture that fails is not evidence that nothing is running, which is the same + // rule the working probe applies to its own null. + screen.text = null; + runAndSettle(session); + expect(session.watching).toBe('1 monitor'); + }); + + it('reports nothing for a CLI whose footer nobody has characterised', () => { + vi.useFakeTimers(); + expect(getCli('codex')?.capabilities.workDetect?.watchingLine).toBeUndefined(); + // Codex draws its own composer glyph, so this session settles the same way; what it + // must not do is read Claude's footer on a screen that is not Claude's. + const session = withFakePane(WITH_MONITOR, 'codex'); + + for (let i = 0; i < 3; i++) { + feed(session, '\x1b[31;1H\x1b[38;5;246m›\xa0\x1b[39m\x1b[0m'); + vi.advanceTimersByTime(1000); + } + vi.advanceTimersByTime(IDLE_SILENCE_MS + 2000); + + expect(session.watching).toBeNull(); + }); + + it('rides along on the payload every session surface reads', () => { + vi.useFakeTimers(); + const session = withFakePane(WITH_SHELL); + runAndSettle(session); + + expect(session.toLightDetailedState().watching).toBe('1 shell'); + }); +}); + +describe('the registry pattern Claude declares', () => { + it('is one the config-regex guard accepts', () => { + // Same guard as `workingLine`: ~/.codeman/clis.json can set this field, and the + // compiled pattern runs over a pane capture on a timer. + expect(compileVersionRegex(getCli('claude')!.capabilities.workDetect!.watchingLine!)).not.toBeNull(); + }); + + it('does not fire on the status line a user configured', () => { + // Plan-usage and context figures live one row above the footer and carry numbers. + const statusLine = ' ~/innovi/gtd-board [main] Opus 5 (1M context) high ctx: 10% 5h: 48% (32m) 7d: 15% (6d10h)'; + expect(CLAUDE_WATCHING.test(statusLine)).toBe(false); + }); +});