fix(terminal): recover a dropped output frame, do not merely schedule it

`_onSessionTerminal` drops an incoming frame once the app-owned render queues
already hold 128KB. That is the right call — the alternative is an unbounded
backlog — but a hole in a TUI byte stream is a desynced cursor, and a desynced
cursor is muffled text (#464). The drop was only half of it.

The recovery was a fire-and-forget timer: it nulled its own handle and then
called `_onSessionNeedsRefresh()`, which opens with four early returns. Two of
them — a buffer load already in flight, a refresh already owning this session —
are MOST likely to be true during exactly the output burst that caused the
drop. So the recovery was skipped precisely when it was needed, with nothing
left to retry it, and the dropped bytes were never replayed.

`_onSessionNeedsRefresh` reports whether it actually repainted now, and
`_scheduleDroppedOutputRecovery` re-arms while it has not. Bounded by
`DROP_RECOVERY_MAX_ATTEMPTS`, because every reason the refresh can be skipped is
transient contention that clears in seconds and a permanently failing refresh
must not become a loop against the API; giving up at the cap leaves exactly what
the old code left, so the floor is no worse than before. The same 2s debounce
still collapses a burst of drops into one attempt.

This is the principle Ark0N established reviewing #431 for the WebSocket
output-gap marker — only a repaint that actually happened settles the recovery —
applied to the one recovery path that still trusted a timer having fired.

The retry decision is a pure function in constants.js so the gate can reach it,
and the scheduler itself is driven from app.js under a fake clock. The retry
case and the no-retry case only pin the fix AS A PAIR: either alone passes
against something wrong, one against the old fire-and-forget timer and the other
against retrying forever. Checked by reverting app.js to the old shape, where
three of the twelve fail.

Two harness details that would otherwise have made the tests lie. The vm context
baked in the real `setTimeout`, so `vi.useFakeTimers()` could not reach the
scheduler and every case reported zero calls; it delegates lazily now. And
app.js reached `CodemanDroppedOutput` as a bare global, which resolves in a
browser but not in the vm — worth fixing beyond the test, because that call sits
inside a timer where a ReferenceError is swallowed and would take the recovery
with it. It reads through `window.` like terminal-ui.js does with its own
constants.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Rounak Datta
2026-09-22 21:34:03 +05:30
co-authored by Claude Opus 5
parent e1e7dc5bd8
commit 00f022ccf8
5 changed files with 360 additions and 14 deletions
+1 -1
View File
@@ -259,7 +259,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set); Shell selection and automatic drop recovery always use a bounded 1 MiB `?tail=` window. Shell loads the rest only when **Load full history** is pressed; ordinary scrolling must not trigger a multi-megabyte reset+replay on xterm's main thread. Other modes may re-pull at the TOP (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). Live writes are one-chunk-in-flight, released by xterm's parse callback, so xterm's private queue cannot bypass the browser's 128 KiB render cap. While WebSocket owns terminal I/O, duplicate SSE terminal events are dropped before JSON parsing, and recovery is single-flight per active session. ⚠️ **A `full=1` capture ENDS with a cursor move back to the pane's own caret position**, counted UP from the last replayed row — without it the caret stays where the last character landed, which for an agent CLI is the status line, and every cursor-relative update the CLI sends afterwards is measured from the wrong row. The move is relative, not `CUP`: absolute row addressing is only right while the browser's rows equal the pane's, and `resizeWindow` does not wait for tmux, so a capture can be taken before a requested resize applies. That makes row alignment load-bearing on this path: no transform that can DELETE A LINE may run over the capture, so it keeps its trailing blank rows and skips redraw-bloat stripping, the banner trim and the leading-whitespace strip. ⚠️ Those three skips key on whether a capture actually CAME BACK (`isFullCapture`), never on `?full=1` alone — the fallback to the byte history is a stream of successive frames that must still be stripped, and a session with no mux takes it on every load. A capture holding nothing visible returns '' so the byte history survives instead of a blank screen replacing it. ⚠️ A full re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. ⚠️ **A visible capture now REPORTS the geometry it was taken at** (`captureCols`/`captureRows`, #435), because a frame built for a pane taller or wider than the browser is damaged two ways at once (overflow rows clamp onto the last line; a narrower browser wraps every painted row) and nothing in the response used to say so. Both fields are ABSENT when no frame was positioned, so every consumer tests `Number.isFinite`, never truthiness: a `display-message` cursor query that fails makes `capturePaneBuffer` return the raw capture while the route still labels it `mux-visible`. The comparison runs on `mux-visible` ONLY, the replay is capped at one attempt, and a pane that cannot be sized to fit latches in `_geometryRetryUseless` so it is diagnosed once per session rather than on every tab switch. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set); Shell selection and automatic drop recovery always use a bounded 1 MiB `?tail=` window. Shell loads the rest only when **Load full history** is pressed; ordinary scrolling must not trigger a multi-megabyte reset+replay on xterm's main thread. Other modes may re-pull at the TOP (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). Live writes are one-chunk-in-flight, released by xterm's parse callback, so xterm's private queue cannot bypass the browser's 128 KiB render cap. ⚠️ **A frame dropped at that cap MUST be recovered, and the recovery must verify itself** (`_scheduleDroppedOutputRecovery`, app.js): a hole in a TUI byte stream is a desynced cursor, which is muffled text (#464). It was a fire-and-forget 2s timer that nulled its own handle and then called `_onSessionNeedsRefresh()` — whose early returns (a buffer load in flight, a refresh already owning the session) are MOST likely to be true during exactly the burst that caused the drop, so the recovery was lost silently and the bytes were never replayed. `_onSessionNeedsRefresh` now returns whether it actually repainted, and the scheduler re-arms while it has not, bounded by `DROP_RECOVERY_MAX_ATTEMPTS` because every reason it can be skipped is transient contention. The same debounce still collapses a burst into one attempt. Tests: `test/dropped-output-recovery.test.ts`, whose retry case and no-retry case only pin the fix as a pair. While WebSocket owns terminal I/O, duplicate SSE terminal events are dropped before JSON parsing, and recovery is single-flight per active session. ⚠️ **A `full=1` capture ENDS with a cursor move back to the pane's own caret position**, counted UP from the last replayed row — without it the caret stays where the last character landed, which for an agent CLI is the status line, and every cursor-relative update the CLI sends afterwards is measured from the wrong row. The move is relative, not `CUP`: absolute row addressing is only right while the browser's rows equal the pane's, and `resizeWindow` does not wait for tmux, so a capture can be taken before a requested resize applies. That makes row alignment load-bearing on this path: no transform that can DELETE A LINE may run over the capture, so it keeps its trailing blank rows and skips redraw-bloat stripping, the banner trim and the leading-whitespace strip. ⚠️ Those three skips key on whether a capture actually CAME BACK (`isFullCapture`), never on `?full=1` alone — the fallback to the byte history is a stream of successive frames that must still be stripped, and a session with no mux takes it on every load. A capture holding nothing visible returns '' so the byte history survives instead of a blank screen replacing it. ⚠️ A full re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. ⚠️ **A visible capture now REPORTS the geometry it was taken at** (`captureCols`/`captureRows`, #435), because a frame built for a pane taller or wider than the browser is damaged two ways at once (overflow rows clamp onto the last line; a narrower browser wraps every painted row) and nothing in the response used to say so. Both fields are ABSENT when no frame was positioned, so every consumer tests `Number.isFinite`, never truthiness: a `display-message` cursor query that fails makes `capturePaneBuffer` return the raw capture while the route still labels it `mux-visible`. The comparison runs on `mux-visible` ONLY, the replay is capped at one attempt, and a pane that cannot be sized to fit latches in `_geometryRetryUseless` so it is diagnosed once per session rather than on every tab switch. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions)