feat(codex): read Codex's own background-terminal row

Codex states background work too, and it says so in a different place. Claude
writes `· 1 monitor ·` on the last row of the screen; Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its
composer, which puts that row third from the bottom once the status line and
the composer are counted.

So how far up the screen to look is now per-CLI data as well:
`capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and
defaulting to Claude's two. That bound is the point. The window is half the
injection guard, since every row it adds is another row the agent itself may be
able to write, and the label is what silences an idle alert. The other half is
the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only
the CLI can offer, so a session that writes "I left 1 background terminal
running for you" into its own output matches nothing.

Measured against a live codex-cli 0.154.0 pane rather than read out of a
binary. The row appears when the terminal starts, follows the composer down as
the conversation grows, and is gone after `/stop`. Verified end to end on an
isolated beta: the session payload carried `watching: "1 background terminal"`
and the badge rendered with it, and both cleared when the terminal stopped. The
fixtures in the tests are that capture verbatim.

Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a
Codex session this is the badge alone, which is the case the maintainer said a
registry field could cover and a hook never could. Cross-CLI tests pin that
neither pattern fires on the other's screen, and that a CLI declaring nothing
still reports nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Michael Grundberg
2026-09-22 08:58:52 +02:00
co-authored by Claude Opus 5
parent 74884a20eb
commit 64c288a683
10 changed files with 185 additions and 45 deletions
+1 -1
View File
@@ -201,7 +201,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer (`❯`) about once a second all through a turn, so the old "saw a ❯, wait 2s → idle" rule flipped every working session to idle two seconds in (measured: a session mid-tool-call at 17 minutes reporting `status:"idle"`). Its working indicator is `✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`: the glyph animates through `· ✢ ✳ ∗ ✻ ✽`, the gerund is randomized, and the finished line (`✻ Cooked for 2m 49s`) carries the same glyph, so neither `SPINNER_PATTERN` (braille, not what current versions draw) nor a keyword list can see it. Matching the new line in the STREAM does not work either: tmux ships partial repaints, so the whole line reaches the PTY only every few tens of seconds. So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks the SCREEN via `capturePaneText()` + `CLAUDE_WORKING_LINE_PATTERN` before believing it; a sustained run of repaints (`session-activity.ts`, pure + unit tested) is what marks a turn as started, with the same screen probe vetoing keystroke echo. Idle now lands ~3-5s after a turn ends instead of 2s into one. ⚠️ **The composer glyph and the working line are per-CLI registry DATA** (`capabilities.workDetect`, #385), not Claude constants: claude declares `❯` plus the pattern above, codex declares `›` plus `[Ee]sc to interrupt`, and a CLI that declares neither falls back to Claude's pair, which is what every session used before the registry carried one. Before that, this whole mechanism was gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` for its entire life. ⚠️ `workingLine` is config-supplied (a user `clis.json` can set it) and the compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()`: a nested quantifier there is a ReDoS against the event loop, and the helper returns null rather than throwing so the fallback is structural. ⚠️ **A quiet pane is not always a pane that wants you.** An agent that arms a monitor, backgrounds a shell or hands work to a cloud session is told to end its turn, so the pane goes quiet, Claude Code's `idle_prompt` notification lands a minute later and every surface files the session under NEEDS YOU with nothing to answer. The CLI states what it is still running on the last row of its screen (`⏵⏵ bypass permissions on · 1 monitor · ← for agents`), which is the optional `capabilities.workDetect.watchingLine`: `_confirmIdle()`'s own capture feeds `watchingLabel()` (`session-activity.ts`, pure), and the label lands on `Session.watching`. ⚠️ **The fix is the alert that does not fire, not the badge.** `hook-event-routes` passes that label to `notePrompt()`, which opens the idle item ALREADY acknowledged (`acknowledgedAt` + `acknowledgedReason`), so the prompt stays pending, answerable and Read-My-Mind context while the alert it would have armed is spent — the same `acknowledge()` semantics a human viewing the session has always produced. The broadcast carries `acknowledgedReason` so a live page declines to arm (`_onHookIdlePrompt`), the push is skipped, a reloading page reads `acknowledgedAt` in `seedApprovals()`, and `classifySession()` ignores an acknowledged item so `codeman tui` agrees. It re-arms for free: the next idle prompt supersedes this item and is built fresh. ⚠️ Only `idle` is eligible, so a permission or question dialog still goes red whatever else the agent started. The badge is the cosmetic half, and it rides BESIDE the state pill on the surfaces that have one, never in place of it. ⚠️ The label is **pane-derived and therefore prompt-injectable**: the search is confined to the last `WATCHING_TAIL_LINES` lines and each CLI's pattern anchors on its footer's own separator, because an agent that got a bare `1 monitor` matched would silence its own alert by printing it. Tests: `test/session-watching.test.ts` and `test/watching-no-alert.test.ts` (the negative one).
⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer (`❯`) about once a second all through a turn, so the old "saw a ❯, wait 2s → idle" rule flipped every working session to idle two seconds in (measured: a session mid-tool-call at 17 minutes reporting `status:"idle"`). Its working indicator is `✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`: the glyph animates through `· ✢ ✳ ∗ ✻ ✽`, the gerund is randomized, and the finished line (`✻ Cooked for 2m 49s`) carries the same glyph, so neither `SPINNER_PATTERN` (braille, not what current versions draw) nor a keyword list can see it. Matching the new line in the STREAM does not work either: tmux ships partial repaints, so the whole line reaches the PTY only every few tens of seconds. So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks the SCREEN via `capturePaneText()` + `CLAUDE_WORKING_LINE_PATTERN` before believing it; a sustained run of repaints (`session-activity.ts`, pure + unit tested) is what marks a turn as started, with the same screen probe vetoing keystroke echo. Idle now lands ~3-5s after a turn ends instead of 2s into one. ⚠️ **The composer glyph and the working line are per-CLI registry DATA** (`capabilities.workDetect`, #385), not Claude constants: claude declares `❯` plus the pattern above, codex declares `›` plus `[Ee]sc to interrupt`, and a CLI that declares neither falls back to Claude's pair, which is what every session used before the registry carried one. Before that, this whole mechanism was gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` for its entire life. ⚠️ `workingLine` is config-supplied (a user `clis.json` can set it) and the compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()`: a nested quantifier there is a ReDoS against the event loop, and the helper returns null rather than throwing so the fallback is structural. ⚠️ **A quiet pane is not always a pane that wants you.** An agent that arms a monitor, backgrounds a shell or hands work to a cloud session is told to end its turn, so the pane goes quiet, Claude Code's `idle_prompt` notification lands a minute later and every surface files the session under NEEDS YOU with nothing to answer. The CLI states what it is still running on the last row of its screen (`⏵⏵ bypass permissions on · 1 monitor · ← for agents`), which is the optional `capabilities.workDetect.watchingLine`: `_confirmIdle()`'s own capture feeds `watchingLabel()` (`session-activity.ts`, pure), and the label lands on `Session.watching`. ⚠️ **The fix is the alert that does not fire, not the badge.** `hook-event-routes` passes that label to `notePrompt()`, which opens the idle item ALREADY acknowledged (`acknowledgedAt` + `acknowledgedReason`), so the prompt stays pending, answerable and Read-My-Mind context while the alert it would have armed is spent — the same `acknowledge()` semantics a human viewing the session has always produced. The broadcast carries `acknowledgedReason` so a live page declines to arm (`_onHookIdlePrompt`), the push is skipped, a reloading page reads `acknowledgedAt` in `seedApprovals()`, and `classifySession()` ignores an acknowledged item so `codeman tui` agrees. It re-arms for free: the next idle prompt supersedes this item and is built fresh. ⚠️ Only `idle` is eligible, so a permission or question dialog still goes red whatever else the agent started. The badge is the cosmetic half, and it rides BESIDE the state pill on the surfaces that have one, never in place of it. ⚠️ The label is **pane-derived and therefore prompt-injectable**: the search is confined to the last few non-blank rows (`WATCHING_TAIL_LINES`, or the CLI's own `watchingLines` — claude writes its chip on the last row, codex pins one above its composer) and each pattern anchors on chrome only that CLI draws, because an agent that got a bare `1 monitor` matched would silence its own alert by printing it. Tests: `test/session-watching.test.ts` and `test/watching-no-alert.test.ts` (the negative one).
**Workspace-trust dialog auto-accept** (`session-trust-dialog.ts`, pure + unit tested): Claude Code asks once per directory ("Is this a project you created or one you trust?") before it will read or edit anything, and since Codeman sessions run permission-skipping or classifier-guarded modes the answer is always yes, so a session parked on that dialog is simply stuck. ⚠️ **Match the compacted SCREEN, never the stream.** tmux repaints a row by writing each word and then a cursor-forward (`\x1b[C`) instead of a space, and Ink colours each word separately, so the wire carries `I\x1b[Ctrust\x1b[Cthis\x1b[Cfolder`; stripping the escapes leaves `Itrustthisfolder`, because the spaces are not there to strip, they were never sent. A plain `includes('trust this folder')` therefore never matched a single chunk and the auto-accept was silently DEAD for every session that hit the dialog. `compactScreenText()` removes ALL whitespace instead (plus the `ESC ( B` charset selects that `stripAnsi` does not cover, which would otherwise land inside a phrase as a literal `(B`), which survives both that repaint style and the spaced full-screen redraw. ⚠️ **Never answer it with a blind `\r`.** The layout has changed under us at least twice, and Claude Code 2.1.252 dropped the option numbers, put "No, exit" FIRST and highlights IT by default, so the Enter that answered the old dialog now picks *exit* and the pane dies (`Pane is dead (status 1)`) seconds after the session starts. `trustDialogNextKey()` reads the `❯` marker and returns ONE step at a time (an arrow while the cursor is on the wrong option, Enter only once the screen shows it on the trust option), with the pane re-read between steps, so a dropped arrow costs a repaint instead of the session; a frame that does not say which option is highlighted returns null and waits for the next repaint. ⚠️ The LAST marked option in the text wins, because the direct-PTY fallback reads an append-only buffer where every repaint since launch is still present and an older frame must not out-vote the freshest one. ⚠️ Answering types into a live session, so THREE guards must all hold and none is redundant: a **startup-only window** (`TRUST_DIALOG_WINDOW_MS`, 90s, since the dialog renders before the main UI and leaving it open forever would let an agent transcript that merely QUOTES the dialog trigger an Enter, this file being an example), a **two-marker match** requiring a trust phrase AND one of the dialog's own confirm affordances (`isTrustDialogScreen`), and an **attempt cap** (`TRUST_DIALOG_MAX_ATTEMPTS`, 6: a keystroke can land while Ink is still mounting the widget and be dropped, which is the other half of why sessions got stuck here, but retrying forever would hammer keys into whatever came next; it was 3 while one Enter answered the dialog, and answering now costs at least two keystrokes). ⚠️ It reads `capturePaneText()` and falls back to a deliberately SHORT tail of the terminal buffer only on a direct-PTY session, which has no pane: that buffer is append-only, so a longer tail would keep re-matching a dialog answered minutes ago. ⚠️ **The scan must schedule its own next read** (`_trustDialogTimer`, cleared in `_clearAllTimers()`): it runs from the PTY `onData` handler, which was enough while one Enter answered the dialog, but the arrow that moves the cursor is the LAST output the pane produces, so a two-keystroke answer waiting on more output parks forever with the cursor sitting on the right option (measured on a live 2.1.252 spawn: cursor moved at 6 s, then nothing).