fix(watching): close the review findings on the label and its window

A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.

The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.

The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.

Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.

Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.

Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Michael Grundberg
2026-09-22 17:11:33 +02:00
co-authored by Claude Opus 5
parent 64c288a683
commit 05c788ce9d
19 changed files with 292 additions and 59 deletions
+10
View File
@@ -148,6 +148,16 @@ describe('workDetect.workingLine is guarded like every other config regex', () =
}
});
it('refuses a window with no pattern to bound', () => {
expectRejected((e) => {
(e.capabilities as Record<string, unknown>).workDetect = {
promptGlyph: '>',
workingLine: 'working',
watchingLines: 3,
};
}, 'a window with nothing to search is a typo whose failure is otherwise silent');
});
it('accepts every shipped watchingLine', () => {
for (const entry of STOCK_CLIS) {
const src = entry.capabilities.workDetect?.watchingLine;
+8 -2
View File
@@ -19,8 +19,11 @@ function fakeElement(): any {
type: '',
dataset: {},
style: {},
attrs: {} as Record<string, string>,
children: [] as any[],
setAttribute() {},
setAttribute(name: string, value: string) {
el.attrs[name] = value;
},
appendChild(child: any) {
el.children.push(child);
return child;
@@ -514,13 +517,16 @@ describe('mobile overview watching badge', () => {
expect(model.needsYou[0].watching).toBe('2 shells');
});
it('says one word and puts the detail in the tooltip', () => {
it('says one word and puts the detail where every surface can reach it', () => {
const app = loadOverviewApp();
const badge = app._buildWatchingBadge('1 monitor', 'mobile-overview-pill');
expect(badge.className).toBe('mobile-overview-pill mobile-overview-pill--watching');
expect(badge.textContent).toBe('watching');
expect(badge.title).toBe('Still running in the background: 1 monitor');
// A phone has no hover target and a screen reader reads neither the class nor the
// tooltip, so the label has to be here too or this surface says only "watching".
expect(badge.attrs['aria-label']).toBe('Still running in the background: 1 monitor');
});
it('takes the pill class of whichever surface asks for it', () => {
+18
View File
@@ -97,6 +97,24 @@ describe('watching badge on a rich session row', () => {
);
});
it('escapes the label everywhere it reaches markup', () => {
// `watching` is pane-derived and a config-supplied pattern decides what its capture
// group holds, so every interpolation of it into HTML has to go through escapeHtml().
// The row is installed with innerHTML, which makes an unescaped quote in that
// attribute an injection rather than a cosmetic bug.
expect(app).toContain('${richRow.createdAt}:${escapeHtml(richRow.watching)}"');
expect(app).not.toContain('${richRow.createdAt}:${richRow.watching}"');
});
it('words the tooltip exactly as the phone overview does', () => {
// Both files build this sentence themselves, deliberately, so that a stale cached
// module still renders a complete row. Substring-matching the prefix would let the
// two drift; the whole sentence is what has to agree.
const overview = readFileSync(resolve(publicDir, 'mobile-overview.js'), 'utf8');
expect(overview).toContain("'Still running in the background: ' + label");
expect(app).toContain('`Still running in the background: ${row.watching}`');
});
it('colours it with the accent, never with the two colours that mean a human is needed', () => {
const rule = styles.slice(styles.indexOf('.tab-pill--watching'));
const block = rule.slice(0, rule.indexOf('}'));
+46 -6
View File
@@ -167,6 +167,26 @@ describe('watchingLabel', () => {
expect(watchingLabel(`${chip}\n${below}\n`, CLAUDE_WATCHING)).toBeNull();
});
it('refuses a chip on the row above the footer, which the agent can write', () => {
// The status line is one row up, its text comes from a `statusLine` command, and a
// session running with permissions bypassed can write that command into
// `.claude/settings.json` in its own workspace. The window is what keeps that row
// out, so this is the test that would fail if somebody widened it.
const forged = pane('⏵⏵ bypass permissions on (shift+tab to cycle) · ← for agents').replace(
' ~/innovi/gtd-board [main] Opus 5 ctx: 11%',
' ~/innovi/gtd-board [main] Opus 5 ctx: 11% · 1 monitor'
);
expect(watchingLabel(forged, CLAUDE_WATCHING)).toBeNull();
// And with the window widened by one, the same screen does match — which is the
// whole reason the default is one row.
expect(watchingLabel(forged, CLAUDE_WATCHING, 2)).toBe('1 monitor');
});
it('keeps Claude on the default window, because its chip is the last row', () => {
expect(getCli('claude')?.capabilities.workDetect?.watchingLines).toBeUndefined();
expect(WATCHING_TAIL_LINES).toBe(1);
});
it('refuses a label the footer did not separate, which is the injection guard', () => {
// The pattern anchors on the `·` the footer joins its items with. Without that
// anchor an agent could silence its own idle alert by printing the words, since the
@@ -296,14 +316,34 @@ describe('the row Codex draws', () => {
expect(CODEX_TAIL).toBeGreaterThanOrEqual(3);
});
it('refuses the same words in the transcript, which is the injection guard', () => {
// ` · /ps to view` is chrome: only the CLI offers that slash command. Without the
// anchor an agent could print the sentence and silence itself.
const claim = CODEX_STOPPED.replace(
it('refuses a mention that is not the whole row', () => {
// The pattern matches Codex's row end to end, so prose about background terminals —
// including prose quoting part of the row — is not enough.
for (const line of [
'• I left 1 background terminal running for you.',
' 1 background terminal running · /ps to view',
' see: 1 background terminal running · /ps to view · /stop to close',
]) {
const claim = CODEX_STOPPED.replace('• Stopping all background terminals.', line);
expect(watchingLabel(claim, CODEX_WATCHING, CODEX_TAIL)).toBeNull();
}
});
it('CAN be forged by Codex own output, and is contained by Codex having no hooks', () => {
// Codex's row is third from the bottom only while a terminal runs; with none running
// that slot is the last row of the transcript, which the agent writes. Matching the
// complete row raises the bar but closes nothing, so this test states the limitation
// rather than a protection the code does not have.
const forged = CODEX_STOPPED.replace(
'• Stopping all background terminals.',
'• I left 1 background terminal running for you.'
' 1 background terminal running · /ps to view · /stop to close'
);
expect(watchingLabel(claim, CODEX_WATCHING, CODEX_TAIL)).toBeNull();
expect(watchingLabel(forged, CODEX_WATCHING, CODEX_TAIL)).toBe('1 background terminal');
// What makes that cost a wrong badge and nothing more: no hook event from a codex
// session reaches the approvals inbox, so there is no idle item to pre-acknowledge
// and no alert to silence. A CLI that gains hook signals needs a harder anchor first.
expect(getCli('codex')?.capabilities.hooks).toBe('none');
});
});
+24
View File
@@ -59,6 +59,24 @@ describe('approvalCard', () => {
expect(card.hint).toBe('p to reply');
});
it('says why an acknowledged idle prompt is quiet, in the drawer own words', () => {
// The session is watching work it started itself. Asking for a reply would be the
// same false alarm the acknowledgement exists to remove, so the card states the
// reason and drops out of the warning vocabulary — while staying answerable.
const card = approvalCard(
item({
kind: 'idle',
message: 'Claude is waiting for your input',
options: undefined,
acknowledgedAt: 1_700_000_000_000,
acknowledgedReason: 'watching 1 monitor',
})
);
expect(card.tone).toBe('info');
expect(card.title).toBe('quiet, watching 1 monitor');
expect(card.hint).toBe('p to reply');
});
it('drops the approve/deny-only hint when the frame did not parse', () => {
const card = approvalCard(item({ options: undefined }));
expect(card.options).toEqual([]);
@@ -81,6 +99,12 @@ describe('approvalTone', () => {
expect(approvalTone(item({ kind: 'question' }))).toBe('err');
expect(approvalTone(item({ kind: 'idle' }))).toBe('warn');
});
it('is neither for a prompt that opened acknowledged', () => {
expect(approvalTone(item({ kind: 'idle', acknowledgedReason: 'watching 2 shells' }))).toBe('info');
// A dialog stays red whatever else the session started.
expect(approvalTone(item({ kind: 'permission', acknowledgedReason: 'watching 2 shells' }))).toBe('err');
});
});
describe('approvalAnswerForKey', () => {
+32
View File
@@ -11,6 +11,7 @@ import { charWidth, stripStyles, toDisplayLines, visibleWidth } from '../../src/
import { composerMove, createComposer } from '../../src/tui/tui-composer.js';
import { computeLayout, needsBanner } from '../../src/tui/tui-layout.js';
import { createTuiModel, type TuiModelStore } from '../../src/tui/tui-model.js';
import type { ApprovalItem } from '../../src/web/approval-inbox.js';
import {
composerCursorCell,
detectGlyphTier,
@@ -19,6 +20,7 @@ import {
formatPlanUsage,
formatTokens,
glyphsFor,
pendingApprovalCount,
renderFrame,
rowLabel,
type TuiRenderOptions,
@@ -682,3 +684,33 @@ describe('the unicode glyph set is safe to render', () => {
expect(every.join('')).not.toContain('\u270B');
});
});
describe('pendingApprovalCount', () => {
const prompt = (over: Partial<ApprovalItem>): ApprovalItem => ({
id: 'bbb2:1',
sessionId: 'bbb2',
sessionName: 'w6-docs',
kind: 'idle',
createdAt: NOW - 30_000,
...over,
});
it('counts a prompt nobody has seen', () => {
const model = fixture();
model.setApprovals([prompt({})]);
expect(pendingApprovalCount(model)).toBe(1);
});
it('does not count one whose alert is already spent', () => {
// Both ways an item gets acknowledged: a human opening the session elsewhere, and the
// inbox opening it that way for a session watching its own background work. The row
// has left NEEDS YOU by the same flag, so a number in the header would point at a
// group the reader can see is empty.
const model = fixture();
model.setApprovals([prompt({ acknowledgedAt: NOW - 20_000 })]);
expect(pendingApprovalCount(model)).toBe(0);
model.setApprovals([prompt({ acknowledgedAt: NOW - 20_000, acknowledgedReason: 'watching 1 monitor' })]);
expect(pendingApprovalCount(model)).toBe(0);
});
});