fix(watching): close the review findings on the label and its window

A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.

The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.

The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.

Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.

Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.

Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Michael Grundberg
2026-09-22 17:11:33 +02:00
co-authored by Claude Opus 5
parent 64c288a683
commit 05c788ce9d
19 changed files with 292 additions and 59 deletions
+4 -2
View File
@@ -4,9 +4,11 @@
A session watching work it started itself no longer asks you to look at it. An agent that arms a monitor, backgrounds a shell or hands a task to a cloud session is told to end its turn, so the pane goes quiet, Claude Code's idle notification lands a minute later, and the session shows up under NEEDS YOU with nothing for anyone to answer.
A CLI states what it is still running on its own screen, and the CLI registry now carries that row as an optional `capabilities.workDetect.watchingLine`. The idle probe reads the label off the capture it already takes when a turn ends, and it reaches the session payload as `watching`. Claude writes its chip on the last row (`· 1 monitor ·`, also shells, cloud sessions and background tasks); Codex pins `1 background terminal running · /ps to view` above its composer and declares the deeper window that needs. Both patterns were measured against live panes, and a CLI that declares none reports no background work, as every CLI did before.
A CLI states what it is still running on its own screen, and the CLI registry now carries that row as an optional `capabilities.workDetect.watchingLine`. The idle probe reads the label off the capture it already takes when a turn ends, and it reaches the session payload as `watching`. Claude writes its chip on the last row (`· 1 monitor ·`, also shells, cloud sessions and background tasks); Codex pins `1 background terminal running · /ps to view · /stop to close` above its composer and declares the deeper window that needs. Both patterns were measured against live panes, and a CLI that declares none reports no background work, as every CLI did before. The label is read from the foot of the screen only, because it comes off the agent's own pane: `docs/cli-registry.md` sets out what a new pattern has to satisfy before it may be trusted to quiet an alert.
The prompt that follows then opens already acknowledged. It stays in the Approvals drawer, stays answerable and stays Read My Mind context, and only the alert it would have armed is spent: no tab alert, no desktop notification, no push, and no NEEDS YOU row on any surface, `codeman tui` included. The card says why, reading "quiet, watching 1 monitor" rather than implying somebody looked. It re-arms by itself, because the next idle prompt supersedes this item and is built fresh. A permission or question dialog still goes red whatever else the agent started.
The prompt that follows then opens already acknowledged. It stays in the Approvals drawer, stays answerable and stays Read My Mind context, and only the alert it would have armed is spent: no tab alert, no desktop notification, no push, and no NEEDS YOU row on any surface, `codeman tui` included. The card says why on both surfaces, reading "quiet, watching 1 monitor" rather than implying somebody looked. It re-arms by itself, because the next idle prompt supersedes this item and is built fresh. A permission prompt or a question dialog still goes red whatever else the agent started.
One limit worth knowing: a question asked in plain prose is not a dialog, so an agent that starts a monitor and then writes "which branch should I target?" goes quiet along with the false alarms until the background work ends. `docs/wiki/Notifications-And-Approvals.md` says so where users will meet it.
Sessions also wear a `watching` badge beside their state pill on the phone overview, the desktop home rail and the rich sidebar and rail rows.