fix(watching): close the review findings on the label and its window

A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.

The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.

The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.

Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.

Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.

Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Michael Grundberg
2026-09-22 17:11:33 +02:00
co-authored by Claude Opus 5
parent 64c288a683
commit 05c788ce9d
19 changed files with 292 additions and 59 deletions
File diff suppressed because one or more lines are too long
+21 -11
View File
@@ -58,23 +58,33 @@ so a pane waiting for its own background work never raises an alert a human cann
Group 1 is the label, and a CLI that declares no pattern reports no background work.
Two CLIs declare such a row today, and they put it in different places. Claude writes its
chip on the last row of the screen, so it keeps the default window of `WATCHING_TAIL_LINES`
rows and anchors on the `·` its footer joins items with. Codex pins
chip on the last row of the screen, so it keeps the default one-row window and anchors on
the `·` its footer joins items with. Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its composer, which
puts the row third from the bottom once the status line and the composer are counted, so its
entry declares `watchingLines: 4` and anchors on the ` · /ps to view` tail. Both were
measured against live panes rather than read out of a binary, which is the standard for
adding a third.
entry declares `watchingLines: 3` and matches that row end to end. Both were measured
against live panes rather than read out of a binary, which is the standard for adding a
third.
That label is the one value in the registry that an AGENT can influence, because it comes off
the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern
for a new CLI. `watchingLabel()` in `session-activity.ts` searches only the last few
non-blank rows, which is the part of the screen the CLI draws rather than the agent, and the
pattern anchors on chrome only that CLI can produce. Without both, an agent could silence its
own idle alert by printing the words into its output. Keep the window as small as the layout
allows, since every row it adds is another row the agent may be able to write. The label is
also ANSI-stripped and length-capped at the source, since it ends up on a badge and in an
approval card.
non-blank rows, which should be the part of the screen the CLI draws rather than the agent,
and the pattern should anchor on chrome only that CLI can produce. Keep the window as small
as the layout allows, since every row it adds is another row the agent may be able to write.
The label is also ANSI-stripped and length-capped at the source, and every interpolation of
it into markup goes through `escapeHtml()`, since it ends up on a badge and in an approval
card.
The two shipped entries do not sit equally well behind that rule, and the difference decides
what a pattern is allowed to do. Claude's chip is the last row, so its one-row window holds
nothing the agent can write — not even the status line above it, whose command a session
running with permissions bypassed can write into its own `.claude/settings.json`. Codex's row
shares its slot with the last row of the transcript whenever no terminal is running, so a
message ending in that exact line is matched. What keeps that harmless is `hooks: 'none'`: no
hook event from a codex session reaches the approvals inbox, so a forged label costs a wrong
badge and cannot silence an alert. Before giving a CLI both hook signals and a pattern, make
sure its row is one the agent cannot write.
### Three capabilities that must stay independent
+24
View File
@@ -103,6 +103,30 @@ locked phone and the agent continues.
With the inbox off, the buttons are stripped from the notification payload entirely rather
than being shown and failing.
## When a session is watching its own work
An agent that starts a monitor, puts a shell in the background or hands a task to a cloud
session is told by its CLI to end the turn and wait to be notified. The pane then goes
quiet, and the CLI's idle notification arrives about a minute later — for a session that
wants nothing from you.
Codeman reads what the CLI prints about its own background work and treats that prompt
differently. It raises no tab alert, no desktop notification and no push, the session stays
out of NEEDS YOU on every surface, and the row wears a blue **watching** badge instead. Hover
it, or read it on a phone through your screen reader, and it says what is running: "1
monitor", "2 shells", "1 background terminal".
The prompt itself is not thrown away. It sits in the Approvals drawer as an ordinary card,
still answerable, with a line reading "quiet, watching 1 monitor" where a card you had
already looked at would say nothing. The next time that session goes quiet for an ordinary
reason, it alerts you exactly as before.
Two limits are worth knowing. A permission prompt or a question dialog still goes red
whatever else the agent started, because that one blocks it outright. A question asked in
plain prose is not a dialog, so an agent that starts a monitor and then writes "which branch
should I target?" is quiet along with the rest — check a watching session yourself if it has
been quiet longer than the work it is waiting for should take.
## The phone overview
On phones, tapping the "C" logo gives a session overview with **NEEDS YOU** first, then