From 9676e90133deebbc4fc3b3ad8426aa2accb9d7d5 Mon Sep 17 00:00:00 2001 From: timkjr Date: Fri, 25 Sep 2026 19:05:16 -0500 Subject: [PATCH 1/2] fix(terminal): let a Shell pane's scroll-up reach tmux history A burst of output leaves a Shell pane with about one screen of browser scrollback, because tmux repaints the burst instead of scrolling it, while tmux itself keeps every line. Shell declined the scroll-to-top re-pull other modes use, and the Load full history button renders only once a replay was truncated, so a Shell tab under 1 MiB could not scroll back at all. The scroll gesture now pulls ?full=1&tail=TERMINAL_TAIL_SIZE, the same bound a tab switch loads; the route's existing tail cut marks longer histories 'tail', so the banner still offers the unbounded pull. A window no longer than the browser's buffer is not rewritten. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/architecture-invariants.md | 2 +- docs/wiki/The-Dashboard.md | 5 +- src/web/public/app.js | 34 +++++- src/web/public/terminal-ui.js | 6 +- test/history-truncation-notice.test.ts | 9 +- test/shell-scroll-history-pull.test.ts | 142 +++++++++++++++++++++++++ 6 files changed, 184 insertions(+), 14 deletions(-) create mode 100644 test/shell-scroll-history-pull.test.ts diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 6a25c5b3..742419f6 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -205,7 +205,7 @@ Further detail: the `: ` form (`w3-myapp: fix the login redirect` ### Full-scrollback replay -**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell full history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. +**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell FULL history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. That gesture instead pulls a BOUNDED window of it (`?full=1&tail=<TERMINAL_TAIL_SIZE>`, so a capture beyond 1 MiB comes back `truncationReason: 'tail'` and the banner offers the rest), and skips the replay when the window holds no more rows than the browser already has. Declining the gesture outright was a dead end: after any burst a shell pane held about one screen of browser scrollback, and the button renders only once a replay was truncated, so a young shell tab had no way back to history tmux still held. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. **A capture reports the geometry it was taken at** (#435): a visible frame repaints each row at an absolute position, counting up to the pane's height and out to the pane's width, so a terminal smaller than that pane damages it two ways at once. Too short and every address past the browser's own height clamps onto the last line, overwriting the rows underneath (measured: against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew the survivors twice). Too narrow and each row is painted out to the pane's width, so the browser wraps every painted row and the wrap on the last one scrolls the whole frame up by one. Nothing in the response used to say what geometry the frame was built for, so the client could not see either case. `PaneCaptureOptions.capturedGeometry` carries it out, and the terminal response publishes it as `captureCols`/`captureRows`. ⚠️ **Both fields are ABSENT unless a frame was really positioned**, and every consumer must test `Number.isFinite` rather than truthiness: `mux-visible` is necessary but not sufficient, because when the `display-message` cursor query fails `capturePaneBuffer` skips the snapshot repaint and returns the raw capture, and the route still labels that non-empty body `mux-visible`. A body that positioned nothing has no geometry to describe and nothing to repair, so a comparison that fires there buys a second capture, a reset plus chunked rewrite, a dropped and reopened WebSocket and a discarded xterm snapshot for no gain. ⚠️ **The comparison runs on `mux-visible` ONLY.** A full-history body is linear scrollback closed by a RELATIVE cursor move, which is relative precisely so the browser's row count need not match the pane's, and a byte-history body carries no row alignment at all, so a size mismatch damages neither and a replay repairs neither. That gate matters because the first select of every non-shell session per page takes the full-history path, where an ungated comparison would fire most often on the one response it cannot help, at the price of a second whole-scrollback capture. ⚠️ **The replay is capped at one attempt and latches per session when it cannot converge.** `resizeRetry` stops two competing fits trading replays forever; a pane already drawing at the size just requested is left alone, because a retry would capture the identical frame; and a pass that still does not converge joins `_geometryRetryUseless`, so the case `Session.resize` declines outright (a small viewport while a desktop viewport's size claim is live, where the retry re-sends the same declined resize and captures the same pane) costs one attempt per session per page load instead of one per select. ⚠️ **The clamp used to manufacture that equality, and no longer can** (#464). This paragraph previously explained it as the signature of a clamp: `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` did not, so a terminal under 40 columns or 10 rows reported a pane permanently bigger than itself and would replay on every tab switch. That divergence is fixed at the source — `syncTerminalGeometry()` (terminal-ui.js) fits, floors and APPLIES in one step, so the browser terminal IS the size it reports. The only remaining reason the two can differ is a resize the server declined, which the server now reports back (`Session.ptyGeometry`, the `{"t":"zc"}` frame and the resize response) for the client to adopt by COLUMNS. The equality guard stays, for the plain case of a pane already at the requested size. ⚠️ **A retry pass must not re-arm `_fullHistoryLoaded`**: it did not consume the full-history pull, and re-arming it would spend a whole-scrollback capture on the next select. That branch is currently unreachable by construction, since reaching it needs `source === 'mux-visible'` while a `full=1` pass is answered `mux-full-history` or `history`; a static test over the source is the habit this repo uses for an invariant nothing can execute. Tests: `test/capture-geometry-retry.browser.test.ts` (eight cases, five of which fail against the merge base), `test/tmux-capture-full-history.test.ts`, `test/routes/session-routes.test.ts`. diff --git a/docs/wiki/The-Dashboard.md b/docs/wiki/The-Dashboard.md index cf1f415e..323299fe 100644 --- a/docs/wiki/The-Dashboard.md +++ b/docs/wiki/The-Dashboard.md @@ -151,8 +151,9 @@ Worth knowing: - **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open. Shell sessions open from a bounded recent tail so a large transcript cannot stall tab - switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling - and automatic output recovery stay within the bounded browser buffer. + switching. Scrolling to the top of a Shell pane pulls the most recent 1 MiB of its tmux + history; press **Load full history** to pull the rest explicitly. Automatic output + recovery stays within the bounded browser buffer. - **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is always local scrollback. Other CLIs scroll locally. diff --git a/src/web/public/app.js b/src/web/public/app.js index 05577c34..1da813be 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -6320,10 +6320,15 @@ class CodemanApp { if (!sessionId || this._fullHistoryRepullInFlight || this._isLoadingBuffer) return; if (this.detachedSessions?.has(sessionId)) return; const session = this.sessions.get(sessionId); - // A shell's full capture can be many megabytes. Replaying it from an - // ordinary scroll gesture blocks xterm's main thread, so keep that cost - // behind the explicit "Load full history" button. - if (!force && session?.mode === 'shell') return; + // A shell's full capture can be many megabytes, and replaying all of it from + // an ordinary scroll gesture blocks xterm's main thread. So a shell scroll + // pulls a BOUNDED window of tmux's full history (the same 1 MiB a tab switch + // loads, but of the scrollback rather than the visible frame) and the + // unbounded pull stays behind the "Load full history" button. Declining + // outright left a shell pane about one screen of browser scrollback after any + // burst, and the button only renders once a replay was truncated, so a young + // shell tab had no way back to output tmux was still holding. + const boundedShellPull = !force && session?.mode === 'shell'; const now = Date.now(); // Momentum scrolling fires this dozens of times per flick, and a burst of new // output is the normal reason to want a re-pull, so cooldown rather than latch. @@ -6336,7 +6341,12 @@ class CodemanApp { this._fullHistoryRepullInFlight = true; try { const requestStartedAt = performance.now(); - const capture = await this._fetchTerminalCapture(`/api/sessions/${sessionId}/terminal?full=1`, { full: true }); + const capture = await this._fetchTerminalCapture( + boundedShellPull + ? `/api/sessions/${sessionId}/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}` + : `/api/sessions/${sessionId}/terminal?full=1`, + { full: true } + ); const headersReceivedAt = capture.headersAt; const payload = capture.json?.data ?? {}; const bodyParsedAt = performance.now(); @@ -6368,6 +6378,20 @@ class CodemanApp { this._setHistoryTruncation(sessionId, { ...payload, exhausted: true }); return; } + // A bounded window no longer than the browser's buffer buys nothing, and + // resetting to rewrite it would jump the viewport on every scroll that + // outlasts the cooldown at the top. An untruncated window IS all of tmux's + // history, so nothing is missing; a window cut at the tail size would trade + // more old rows than it recovers, and its 'tail' truncation keeps the + // banner offering the unbounded pull. Not latched as useless: the next + // burst of output can put more history in tmux than the browser has. + if ( + boundedShellPull && + this._estimateReplayRows(buffer, this.terminal.cols) <= this.terminal.buffer.active.length + ) { + this._setHistoryTruncation(sessionId, payload); + return; + } this._setHistoryTruncation(sessionId, payload); this._fullHistoryRepullUseless?.delete(sessionId); const rowsBefore = this.terminal.buffer.active.length; diff --git a/src/web/public/terminal-ui.js b/src/web/public/terminal-ui.js index 903661e5..413fe99b 100644 --- a/src/web/public/terminal-ui.js +++ b/src/web/public/terminal-ui.js @@ -3304,9 +3304,9 @@ Object.assign(CodemanApp.prototype, { /** * Post-scroll companion to _noteTerminalUserScroll: hitting the TOP of the * buffer while scrolling up gives the app a chance to pull the rest of tmux's - * scrollback (issue #205, see _maybeRefetchFullHistory). Shell sessions decline - * automatic pulls because their captures can be large; their banner button is - * the explicit path. Must be called AFTER scrollLines(), since the check is on + * scrollback (issue #205, see _maybeRefetchFullHistory). Shell sessions pull a + * bounded window because their captures can be large; their banner button is + * the unbounded path. Must be called AFTER scrollLines(), since the check is on * the resulting position, and it is deliberately not folded into * _noteTerminalUserScroll for exactly that reason. */ diff --git a/test/history-truncation-notice.test.ts b/test/history-truncation-notice.test.ts index b24fe3b6..50eb5735 100644 --- a/test/history-truncation-notice.test.ts +++ b/test/history-truncation-notice.test.ts @@ -122,7 +122,7 @@ describe('the in-terminal truncation line is gone (static guard)', () => { expect(app).not.toContain('earlier output truncated for performance'); }); - it('loads a bounded shell tail first and keeps full history user-triggered', () => { + it('loads a bounded shell tail first and keeps unbounded full history user-triggered', () => { const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8'); expect(app).toContain("session?.mode !== 'shell' && !this._fullHistoryLoaded.has(sessionId)"); expect(app).toContain("!restoredSnapshot && session?.mode !== 'shell'"); @@ -131,10 +131,13 @@ describe('the in-terminal truncation line is gone (static guard)', () => { // an abort deadline (a `?full=1` body can be megabytes and used to hang // indefinitely on a stalled mobile link). The URL and the full-vs-tail // decision this guard exists to pin are unchanged. - expect(app).toContain('this._fetchTerminalCapture(`/api/sessions/${sessionId}/terminal?full=1`, { full: true })'); + expect(app).toContain(': `/api/sessions/${sessionId}/terminal?full=1`,\n { full: true }'); expect(app).toContain("if (this.sessions.get(sessionId)?.mode !== 'shell')"); expect(app).toContain("if (session?.mode === 'shell')"); - expect(app).toContain("if (!force && session?.mode === 'shell') return;"); + // A shell scroll gesture pulls a BOUNDED window of full history; only the + // button pulls all of it (behaviour pinned in shell-scroll-history-pull.test.ts). + expect(app).toContain("const boundedShellPull = !force && session?.mode === 'shell';"); + expect(app).toContain('`/api/sessions/${sessionId}/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`'); expect(app).toContain("trigger: force ? 'full-history-button' : 'full-history-scroll'"); }); diff --git a/test/shell-scroll-history-pull.test.ts b/test/shell-scroll-history-pull.test.ts new file mode 100644 index 00000000..229d8a0c --- /dev/null +++ b/test/shell-scroll-history-pull.test.ts @@ -0,0 +1,142 @@ +/** + * @fileoverview A shell pane's scroll-up must reach the history tmux still holds. + * + * tmux repaints a burst of output instead of scrolling it, so after `cat` of a + * file longer than the screen the browser holds about one screen of scrollback + * while tmux holds all of it. Other modes recover it by re-pulling `?full=1` + * when the wheel reaches the top (`_maybeRefetchFullHistory`, issue #205). Shell + * declined that gesture outright to keep a multi-megabyte capture off xterm's + * main thread, leaving only the "Load full history" button, and that button + * renders only once a replay was truncated. A young shell tab therefore had no + * way to scroll back at all. + * + * The gesture now pulls a BOUNDED window (`?full=1&tail=TERMINAL_TAIL_SIZE`), + * the button stays the unbounded path, and a window the browser already holds + * in full is not rewritten. + * + * The method is extracted from app.js and run in a `vm` against stubs (no jsdom + * on this box; see connection-indicator.test.ts), with the REAL row estimators + * from terminal-ui.js, which decide both the downgrade and the no-gain skip. + */ +import { readFileSync } from 'node:fs'; +import { performance } from 'node:perf_hooks'; +import { resolve } from 'node:path'; +import vm from 'node:vm'; +import { describe, expect, it, vi } from 'vitest'; + +const PUBLIC = resolve(import.meta.dirname, '../src/web/public'); +const TERMINAL_TAIL_SIZE = 1024 * 1024; + +function methodSource(source: string, method: string): string { + const start = source.search(new RegExp(`^ {2}(?:async )?${method}\\(`, 'm')); + expect(start, `${method} not found`).toBeGreaterThan(-1); + const next = /^ {2}(?:async )?[A-Za-z_$][\w$]*\(/m.exec(source.slice(start + 1)); + return next ? source.slice(start, start + 1 + next.index) : source.slice(start); +} + +/** Real terminal-ui.js mixin, for `_estimateReplayRows` / `_replayWouldShrinkBuffer`. */ +function loadTerminalMixin(): Record<string, unknown> { + const source = readFileSync(resolve(PUBLIC, 'terminal-ui.js'), 'utf8'); + const FakeCodemanApp = function () {} as unknown as { prototype: Record<string, unknown> }; + const context = vm.createContext({ + console, + performance, + setTimeout, + clearTimeout, + setInterval: vi.fn(), + clearInterval: vi.fn(), + requestAnimationFrame: vi.fn(), + CodemanApp: FakeCodemanApp, + window: { addEventListener: vi.fn(), removeEventListener: vi.fn() }, + document: { addEventListener: vi.fn() }, + }); + vm.runInContext(source, context); + return FakeCodemanApp.prototype; +} + +function loadRefetch(): (this: unknown, opts?: { force?: boolean }) => Promise<void> { + const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8'); + const body = methodSource(app, '_maybeRefetchFullHistory'); + const context = vm.createContext({ performance, TERMINAL_TAIL_SIZE, TERMINAL_CHUNK_SIZE: 32 * 1024 }); + return vm.runInContext(`({ ${body} })._maybeRefetchFullHistory`, context); +} + +const mixin = loadTerminalMixin(); +const refetch = loadRefetch(); +const lines = (n: number) => Array.from({ length: n }, (_, i) => `line ${i}`).join('\r\n'); + +function makeApp(mode: string, { bufferRows, capture }: { bufferRows: number; capture: string }) { + const urls: string[] = []; + const app = { + activeSessionId: 's1', + sessions: new Map([['s1', { mode }]]), + detachedSessions: new Set<string>(), + _fullHistoryRepullInFlight: false, + _isLoadingBuffer: false, + _fullHistoryRepullAt: new Map<string, number>(), + _fullHistoryRepullUseless: new Set<string>(), + terminalBufferCache: new Map<string, string>(), + terminal: { + cols: 80, + rows: 30, + buffer: { active: { length: bufferRows } }, + scrollToLine: vi.fn(), + scrollToTop: vi.fn(), + }, + _estimateReplayRows: mixin._estimateReplayRows, + _replayWouldShrinkBuffer: mixin._replayWouldShrinkBuffer, + _fetchTerminalCapture: vi.fn(async (url: string) => { + urls.push(url); + return { + headersAt: performance.now(), + headers: { get: () => '' }, + json: { data: { terminalBuffer: capture, source: 'mux-full-history' } }, + }; + }), + _recordTerminalLoadTiming: vi.fn(), + _logScrollRouting: vi.fn(), + _setHistoryTruncation: vi.fn(), + _resetTerminalForReplay: vi.fn(), + _bufferLoadFinishOpts: vi.fn(() => ({})), + chunkedTerminalWrite: vi.fn(async () => ({ parsedAt: performance.now(), bufferLength: 400, completed: true })), + _syncStickyScrollBaseline: vi.fn(), + }; + return { app, urls }; +} + +describe('shell scroll-up pulls a bounded window of tmux history', () => { + it('a shell scroll gesture requests full history bounded by the tail size', async () => { + const { app, urls } = makeApp('shell', { bufferRows: 40, capture: lines(300) }); + await refetch.call(app); + expect(urls).toEqual([`/api/sessions/s1/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`]); + // It then actually replays the recovered history. + expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1); + expect(app.chunkedTerminalWrite).toHaveBeenCalledTimes(1); + }); + + it('the Load full history button stays unbounded for a shell', async () => { + const { app, urls } = makeApp('shell', { bufferRows: 40, capture: lines(300) }); + await refetch.call(app, { force: true }); + expect(urls).toEqual(['/api/sessions/s1/terminal?full=1']); + }); + + it('other modes keep the unbounded scroll pull', async () => { + const { app, urls } = makeApp('claude', { bufferRows: 40, capture: lines(300) }); + await refetch.call(app); + expect(urls).toEqual(['/api/sessions/s1/terminal?full=1']); + }); + + it('a bounded window the browser already holds is not rewritten', async () => { + // Browser already has every row the window carries: resetting to rewrite + // it would jump the viewport on every scroll that outlasts the cooldown. + const { app } = makeApp('shell', { bufferRows: 320, capture: lines(300) }); + await refetch.call(app); + expect(app._resetTerminalForReplay).not.toHaveBeenCalled(); + expect(app.chunkedTerminalWrite).not.toHaveBeenCalled(); + // Not latched as useless: more output can put more history in tmux. + expect(app._fullHistoryRepullUseless.has('s1')).toBe(false); + // The truncation state is still recorded, so a window capped at the tail + // size keeps offering the button. + expect(app._setHistoryTruncation).toHaveBeenCalledTimes(1); + }); +}); From f6aa50239fd44426d47c713750fcd402b0ebee92 Mon Sep 17 00:00:00 2001 From: timkjr <timkjr@k-lab.lan> Date: Fri, 25 Sep 2026 20:17:22 -0500 Subject: [PATCH 2/2] fix(terminal): skip a bounded Shell window before the downgrade guard A window cut at the tail size can be smaller than the browser's buffer while tmux still holds more. The downgrade guard reads that as "tmux has nothing more to give", which is true of an unbounded capture only, so a bounded window reaching it marked the session exhausted and removed Load full history from the banner. The bounded skip now runs first, so such a window never reaches the exhausted path, and it no longer writes banner state: relabelling it from the bounded payload would call a terminal holding all of a Load full history pull "the most recent 1 MiB". A skipped window that came back truncated cannot reach anything older than the browser shows, and every ask costs the server a synchronous capture-pane of the whole history (tail is applied after the capture), so it puts the session on the 60 s cooldown. An untruncated one keeps 4 s. _replayWouldShrinkBuffer takes optional pre-estimated rows so a megabyte capture is not scanned twice. CLAUDE.md's Full-scrollback replay entry no longer says Shell never pulls on ordinary scroll. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 2 +- src/web/public/app.js | 42 ++++--- src/web/public/terminal-ui.js | 8 +- test/shell-scroll-history-pull.test.ts | 167 ++++++++++++++++++++++++- test/terminal-scroll-routing.test.ts | 10 +- 6 files changed, 205 insertions(+), 26 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index c4a910cc..dac96cc7 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -265,7 +265,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) -**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and Shell loads the rest only via **Load full history**, never on ordinary scroll. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) +**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and a Shell scroll-to-top pulls a bounded `?full=1&tail=` window (a window no longer than the browser's buffer is skipped before the downgrade guard, so it never marks the session exhausted); the unbounded pull stays behind **Load full history**. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) **Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 742419f6..e23dc232 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -205,7 +205,7 @@ Further detail: the `<prefix>: <title>` form (`w3-myapp: fix the login redirect` ### Full-scrollback replay -**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell FULL history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. That gesture instead pulls a BOUNDED window of it (`?full=1&tail=<TERMINAL_TAIL_SIZE>`, so a capture beyond 1 MiB comes back `truncationReason: 'tail'` and the banner offers the rest), and skips the replay when the window holds no more rows than the browser already has. Declining the gesture outright was a dead end: after any burst a shell pane held about one screen of browser scrollback, and the button renders only once a replay was truncated, so a young shell tab had no way back to history tmux still held. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. +**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell FULL history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. That gesture instead pulls a BOUNDED window of it (`?full=1&tail=<TERMINAL_TAIL_SIZE>`, so a capture beyond 1 MiB comes back `truncationReason: 'tail'` and the banner offers the rest), and skips the replay when the window holds no more rows than the browser already has. ⚠️ **That skip runs BEFORE the downgrade guard, and it never touches the banner state**: the guard reads "smaller than the browser" as "tmux has nothing more to give", true of an unbounded capture and false of a window cut at the tail size, so a bounded window that reached it marked the session `exhausted` and removed **Load full history** while tmux still held the rest; and re-labelling the banner from a skipped payload would call a terminal that holds all of a Load full history pull "the most recent 1 MiB". A skipped window that came back truncated puts the session on the 60 s cooldown (`_fullHistoryRepullUseless`), because it can never reach anything older than the browser shows and `tail` is applied AFTER the server's synchronous `capture-pane` of the whole history, so each ask costs every client an event-loop stall (about 0.3 s at 100k lines); an untruncated one keeps the 4 s cooldown, since it is all of tmux's history and the next burst can add to it. Declining the gesture outright was a dead end: after any burst a shell pane held about one screen of browser scrollback, and the button renders only once a replay was truncated, so a young shell tab had no way back to history tmux still held. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. **A capture reports the geometry it was taken at** (#435): a visible frame repaints each row at an absolute position, counting up to the pane's height and out to the pane's width, so a terminal smaller than that pane damages it two ways at once. Too short and every address past the browser's own height clamps onto the last line, overwriting the rows underneath (measured: against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew the survivors twice). Too narrow and each row is painted out to the pane's width, so the browser wraps every painted row and the wrap on the last one scrolls the whole frame up by one. Nothing in the response used to say what geometry the frame was built for, so the client could not see either case. `PaneCaptureOptions.capturedGeometry` carries it out, and the terminal response publishes it as `captureCols`/`captureRows`. ⚠️ **Both fields are ABSENT unless a frame was really positioned**, and every consumer must test `Number.isFinite` rather than truthiness: `mux-visible` is necessary but not sufficient, because when the `display-message` cursor query fails `capturePaneBuffer` skips the snapshot repaint and returns the raw capture, and the route still labels that non-empty body `mux-visible`. A body that positioned nothing has no geometry to describe and nothing to repair, so a comparison that fires there buys a second capture, a reset plus chunked rewrite, a dropped and reopened WebSocket and a discarded xterm snapshot for no gain. ⚠️ **The comparison runs on `mux-visible` ONLY.** A full-history body is linear scrollback closed by a RELATIVE cursor move, which is relative precisely so the browser's row count need not match the pane's, and a byte-history body carries no row alignment at all, so a size mismatch damages neither and a replay repairs neither. That gate matters because the first select of every non-shell session per page takes the full-history path, where an ungated comparison would fire most often on the one response it cannot help, at the price of a second whole-scrollback capture. ⚠️ **The replay is capped at one attempt and latches per session when it cannot converge.** `resizeRetry` stops two competing fits trading replays forever; a pane already drawing at the size just requested is left alone, because a retry would capture the identical frame; and a pass that still does not converge joins `_geometryRetryUseless`, so the case `Session.resize` declines outright (a small viewport while a desktop viewport's size claim is live, where the retry re-sends the same declined resize and captures the same pane) costs one attempt per session per page load instead of one per select. ⚠️ **The clamp used to manufacture that equality, and no longer can** (#464). This paragraph previously explained it as the signature of a clamp: `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` did not, so a terminal under 40 columns or 10 rows reported a pane permanently bigger than itself and would replay on every tab switch. That divergence is fixed at the source — `syncTerminalGeometry()` (terminal-ui.js) fits, floors and APPLIES in one step, so the browser terminal IS the size it reports. The only remaining reason the two can differ is a resize the server declined, which the server now reports back (`Session.ptyGeometry`, the `{"t":"zc"}` frame and the resize response) for the client to adopt by COLUMNS. The equality guard stays, for the plain case of a pane already at the requested size. ⚠️ **A retry pass must not re-arm `_fullHistoryLoaded`**: it did not consume the full-history pull, and re-arming it would spend a whole-scrollback capture on the next select. That branch is currently unreachable by construction, since reaching it needs `source === 'mux-visible'` while a `full=1` pass is answered `mux-full-history` or `history`; a static test over the source is the habit this repo uses for an invariant nothing can execute. Tests: `test/capture-geometry-retry.browser.test.ts` (eight cases, five of which fail against the merge base), `test/tmux-capture-full-history.test.ts`, `test/routes/session-routes.test.ts`. diff --git a/src/web/public/app.js b/src/web/public/app.js index 1da813be..a8bda99f 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -6367,7 +6367,33 @@ class CodemanApp { // Bail on a tab switch mid-fetch: writing here would paint another session's // history into the terminal the user is now looking at. if (!buffer || this.activeSessionId !== sessionId) return; - if (this._replayWouldShrinkBuffer(buffer)) { + const windowRows = this._estimateReplayRows(buffer, this.terminal.cols); + // A bounded window no longer than the browser's buffer buys nothing, and + // resetting to rewrite it would jump the viewport on every scroll that + // outlasts the cooldown at the top. This runs BEFORE the downgrade guard + // on purpose: that guard reads "smaller than the browser" as "tmux has + // nothing more to give", which is true of an unbounded capture but not of a + // window cut at the tail size, so a bounded window must never reach the + // exhausted path, which would take Load full history off the banner while + // tmux still holds the rest. Nothing was written here, so the banner state + // is left as the load that produced it set it: re-labelling it from this + // payload would call a terminal that holds ALL of a Load full history pull + // "the most recent 1 MiB". + if (boundedShellPull && windowRows <= this.terminal.buffer.active.length) { + // An untruncated window IS all of tmux's history, so nothing is missing, + // and the next burst of output can put more in tmux than the browser has: + // keep the normal 4 s cooldown. A truncated one is the opposite case, since + // the gesture can never reach anything older than what the browser already + // shows, and every ask costs the server a synchronous capture-pane of the + // whole history (`tail` is applied after the capture): back off to 60 s. + // Trade-off: only a successful replay clears that latch, so a tab switch or + // burst that shrinks the browser's buffer below the window can leave a + // scroll-to-top inert for up to a minute. Load full history (`force`) + // bypasses the cooldown, and the latch is bounded, never permanent. + if (payload.truncated) (this._fullHistoryRepullUseless ||= new Set()).add(sessionId); + return; + } + if (this._replayWouldShrinkBuffer(buffer, windowRows)) { timing.refused = true; timing.totalMs = performance.now() - requestStartedAt; this._recordTerminalLoadTiming(timing); @@ -6378,20 +6404,6 @@ class CodemanApp { this._setHistoryTruncation(sessionId, { ...payload, exhausted: true }); return; } - // A bounded window no longer than the browser's buffer buys nothing, and - // resetting to rewrite it would jump the viewport on every scroll that - // outlasts the cooldown at the top. An untruncated window IS all of tmux's - // history, so nothing is missing; a window cut at the tail size would trade - // more old rows than it recovers, and its 'tail' truncation keeps the - // banner offering the unbounded pull. Not latched as useless: the next - // burst of output can put more history in tmux than the browser has. - if ( - boundedShellPull && - this._estimateReplayRows(buffer, this.terminal.cols) <= this.terminal.buffer.active.length - ) { - this._setHistoryTruncation(sessionId, payload); - return; - } this._setHistoryTruncation(sessionId, payload); this._fullHistoryRepullUseless?.delete(sessionId); const rowsBefore = this.terminal.buffer.active.length; diff --git a/src/web/public/terminal-ui.js b/src/web/public/terminal-ui.js index 413fe99b..e969e1e3 100644 --- a/src/web/public/terminal-ui.js +++ b/src/web/public/terminal-ui.js @@ -3354,13 +3354,17 @@ Object.assign(CodemanApp.prototype, { * below the last line, and _estimateReplayRows can only approximate wrapping. * Only a capture that is worse by more than a full screen counts as a * downgrade, which leaves every genuine recovery case untouched. + * + * A caller that already estimated the capture's rows passes them as + * `estimatedRows`, so a megabyte capture is not scanned twice. */ - _replayWouldShrinkBuffer(capture) { + _replayWouldShrinkBuffer(capture, estimatedRows) { const term = this.terminal; const rowsNow = term?.buffer?.active?.length || 0; if (!rowsNow) return false; const screen = term?.rows || 24; - return this._estimateReplayRows(capture, term?.cols) + screen < rowsNow; + const rows = estimatedRows ?? this._estimateReplayRows(capture, term?.cols); + return rows + screen < rowsNow; }, /** diff --git a/test/shell-scroll-history-pull.test.ts b/test/shell-scroll-history-pull.test.ts index 229d8a0c..ac9de9b0 100644 --- a/test/shell-scroll-history-pull.test.ts +++ b/test/shell-scroll-history-pull.test.ts @@ -14,6 +14,14 @@ * the button stays the unbounded path, and a window the browser already holds * in full is not rewritten. * + * ORDER MATTERS: that skip must run BEFORE the downgrade guard. The guard reads + * "smaller than the browser" as "tmux has nothing more to give", which is true of + * an unbounded capture and false of a window cut at the tail size, so a bounded + * window that reached it marked the session exhausted and took Load full history + * off the banner while tmux still held the rest. The second block below drives + * the real `_setHistoryTruncation` and the real `computeHistoryTruncationNotice` + * to pin what the user is actually told. + * * The method is extracted from app.js and run in a `vm` against stubs (no jsdom * on this box; see connection-indicator.test.ts), with the REAL row estimators * from terminal-ui.js, which decide both the downgrade and the no-gain skip. @@ -61,11 +69,43 @@ function loadRefetch(): (this: unknown, opts?: { force?: boolean }) => Promise<v return vm.runInContext(`({ ${body} })._maybeRefetchFullHistory`, context); } +/** The REAL `_setHistoryTruncation`, so the banner state a pull leaves behind is what production would hold. */ +function loadSetHistoryTruncation(): (this: unknown, sessionId: string, payload?: Record<string, unknown>) => void { + const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8'); + const body = methodSource(app, '_setHistoryTruncation'); + return vm.runInContext(`({ ${body} })._setHistoryTruncation`, vm.createContext({})); +} + +/** The REAL banner decision from constants.js: what the user is told, and whether Load full history is offered. */ +function loadNotice() { + const context = vm.createContext({ console, window: {}, document: {}, navigator: { userAgent: 'test' } }); + vm.runInContext( + `${readFileSync(resolve(PUBLIC, 'constants.js'), 'utf8')}\n;globalThis.__notice = computeHistoryTruncationNotice;`, + context, + { filename: 'constants.js' } + ); + return (context as { __notice: (s: Record<string, unknown>) => { visible: boolean; canLoadMore: boolean } }).__notice; +} + const mixin = loadTerminalMixin(); const refetch = loadRefetch(); +const setHistoryTruncation = loadSetHistoryTruncation(); +const computeNotice = loadNotice(); const lines = (n: number) => Array.from({ length: n }, (_, i) => `line ${i}`).join('\r\n'); -function makeApp(mode: string, { bufferRows, capture }: { bufferRows: number; capture: string }) { +/** A `?full=1&tail=` answer whose window was CUT at the tail size: tmux holds ~3 MiB, the window carries 1 MiB. */ +const TAIL_CUT = { + truncated: true, + truncationReason: 'tail', + fullSize: 3 * 1024 * 1024, + retainedBytes: TERMINAL_TAIL_SIZE, + source: 'mux-full-history', +}; + +function makeApp( + mode: string, + { bufferRows, capture, payload = {} }: { bufferRows: number; capture: string; payload?: Record<string, unknown> } +) { const urls: string[] = []; const app = { activeSessionId: 's1', @@ -90,12 +130,18 @@ function makeApp(mode: string, { bufferRows, capture }: { bufferRows: number; ca return { headersAt: performance.now(), headers: { get: () => '' }, - json: { data: { terminalBuffer: capture, source: 'mux-full-history' } }, + json: { data: { terminalBuffer: capture, source: 'mux-full-history', ...payload } }, }; }), _recordTerminalLoadTiming: vi.fn(), _logScrollRouting: vi.fn(), - _setHistoryTruncation: vi.fn(), + // The real method behind a spy, so a test sees both what it was called with + // and the banner state (`_historyTruncation`) it leaves behind. + _historyTruncation: new Map<string, unknown>(), + _renderHistoryTruncationBanner: vi.fn(), + _setHistoryTruncation: vi.fn((sessionId: string, p?: Record<string, unknown>): void => { + setHistoryTruncation.call(app, sessionId, p); + }), _resetTerminalForReplay: vi.fn(), _bufferLoadFinishOpts: vi.fn(() => ({})), chunkedTerminalWrite: vi.fn(async () => ({ parsedAt: performance.now(), bufferLength: 400, completed: true })), @@ -135,8 +181,117 @@ describe('shell scroll-up pulls a bounded window of tmux history', () => { expect(app.chunkedTerminalWrite).not.toHaveBeenCalled(); // Not latched as useless: more output can put more history in tmux. expect(app._fullHistoryRepullUseless.has('s1')).toBe(false); - // The truncation state is still recorded, so a window capped at the tail - // size keeps offering the button. - expect(app._setHistoryTruncation).toHaveBeenCalledTimes(1); + // Nothing was written, so the banner state is left exactly as it was. + expect(app._setHistoryTruncation).not.toHaveBeenCalled(); + }); +}); + +describe('a skipped bounded window never damages the Load full history banner', () => { + it('a tail-cut window smaller than the browser is not replayed and never marks the session exhausted', async () => { + // The browser holds far more rows than a 1 MiB window carries, and tmux holds + // ~3 MiB. The downgrade guard reads that as "tmux has nothing more to give", + // which is true of an unbounded capture and false of a window cut at the tail. + const { app } = makeApp('shell', { bufferRows: 5000, capture: lines(300), payload: TAIL_CUT }); + // The tab load that put this session on screen left it truncated and recoverable. + app._setHistoryTruncation('s1', TAIL_CUT); + app._setHistoryTruncation.mockClear(); + + await refetch.call(app); + + expect(app._resetTerminalForReplay).not.toHaveBeenCalled(); + expect(app.chunkedTerminalWrite).not.toHaveBeenCalled(); + // Not even a relabel: a skipped window writes nothing, banner state included. + expect(app._setHistoryTruncation).not.toHaveBeenCalled(); + const notice = computeNotice(app._historyTruncation.get('s1') as Record<string, unknown>); + expect(notice.visible).toBe(true); + // "Earlier output is no longer kept" would be a lie: tmux still holds ~2 MiB more. + expect(notice.canLoadMore).toBe(true); + }); + + it('a skip right after Load full history leaves the banner as that load set it', async () => { + // Load full history replayed everything, so nothing is truncated any more. + const afterLoadFullHistory = { + truncated: false, + fullSize: 3 * 1024 * 1024, + retainedBytes: 3 * 1024 * 1024, + source: 'mux-full-history', + }; + const { app } = makeApp('shell', { bufferRows: 5000, capture: lines(300), payload: TAIL_CUT }); + app._setHistoryTruncation('s1', afterLoadFullHistory); + const before = structuredClone(app._historyTruncation.get('s1')); + app._setHistoryTruncation.mockClear(); + + await refetch.call(app); + + // Relabelling it from the bounded payload would call a terminal that holds ALL + // of the history "the most recent 1.0 MB". + expect(app._setHistoryTruncation).not.toHaveBeenCalled(); + expect(app._historyTruncation.get('s1')).toEqual(before); + expect(computeNotice(app._historyTruncation.get('s1') as Record<string, unknown>).visible).toBe(false); + }); + + it('backs off for a minute after a truncated skip, and keeps the 4 s cooldown after an untruncated one', async () => { + // Truncated: the gesture cannot reach anything older than the browser shows, and + // every ask costs the server a synchronous capture of the whole history. + // Only just larger than the window, so the downgrade guard does not fire here: + // the back-off has to come from the skip itself. + const cut = makeApp('shell', { bufferRows: 320, capture: lines(300), payload: TAIL_CUT }); + await refetch.call(cut.app); + expect(cut.app._fullHistoryRepullUseless.has('s1')).toBe(true); + expect(cut.app._fetchTerminalCapture).toHaveBeenCalledTimes(1); + + // Well past 4 s, still inside the minute: no second capture. + cut.app._fullHistoryRepullAt.set('s1', Date.now() - 10_000); + await refetch.call(cut.app); + expect(cut.app._fetchTerminalCapture).toHaveBeenCalledTimes(1); + + cut.app._fullHistoryRepullAt.set('s1', Date.now() - 61_000); + await refetch.call(cut.app); + expect(cut.app._fetchTerminalCapture).toHaveBeenCalledTimes(2); + + // Untruncated: it IS all of tmux's history, and the next burst can add to it. + const whole = makeApp('shell', { bufferRows: 320, capture: lines(300) }); + await refetch.call(whole.app); + expect(whole.app._fullHistoryRepullUseless.has('s1')).toBe(false); + whole.app._fullHistoryRepullAt.set('s1', Date.now() - 5000); + await refetch.call(whole.app); + expect(whole.app._fetchTerminalCapture).toHaveBeenCalledTimes(2); + }); + + it('the downgrade guard still refuses an unbounded capture smaller than the browser, and still marks it exhausted', async () => { + // Reordering must not weaken the guard it moved above: a repaint-mode pane's + // capture really is one frame, and rewriting with it would destroy history. + const oneFrame = makeApp('claude', { bufferRows: 300, capture: lines(36) }); + await refetch.call(oneFrame.app); + expect(oneFrame.app._resetTerminalForReplay).not.toHaveBeenCalled(); + expect(oneFrame.app._setHistoryTruncation).toHaveBeenCalledWith('s1', expect.objectContaining({ exhausted: true })); + expect(oneFrame.app._fullHistoryRepullUseless.has('s1')).toBe(true); + + // The button is unbounded too, so the same guard governs it for a shell. + const button = makeApp('shell', { bufferRows: 5000, capture: lines(300) }); + await refetch.call(button.app, { force: true }); + expect(button.app._resetTerminalForReplay).not.toHaveBeenCalled(); + expect(button.app._setHistoryTruncation).toHaveBeenCalledWith('s1', expect.objectContaining({ exhausted: true })); + }); +}); + +describe('_replayWouldShrinkBuffer takes rows the caller already estimated', () => { + const shrink = mixin._replayWouldShrinkBuffer as (this: unknown, capture: string, rows?: number) => boolean; + const make = () => ({ + terminal: { cols: 80, rows: 30, buffer: { active: { length: 200 } } }, + _estimateReplayRows: vi.fn(mixin._estimateReplayRows as (t: string, c: number) => number), + }); + + it('does not scan the capture again when handed the estimate', () => { + const ctx = make(); + expect(shrink.call(ctx, lines(300), 300)).toBe(false); + expect(shrink.call(ctx, lines(300), 5)).toBe(true); + expect(ctx._estimateReplayRows).not.toHaveBeenCalled(); + }); + + it('still estimates for itself when called the old way', () => { + const ctx = make(); + expect(shrink.call(ctx, lines(300))).toBe(false); + expect(ctx._estimateReplayRows).toHaveBeenCalledTimes(1); }); }); diff --git a/test/terminal-scroll-routing.test.ts b/test/terminal-scroll-routing.test.ts index b3cb888b..d2dd206a 100644 --- a/test/terminal-scroll-routing.test.ts +++ b/test/terminal-scroll-routing.test.ts @@ -111,12 +111,20 @@ describe('full-history re-pull downgrade guard (issue #205 round 2)', () => { // Anchor on the open paren, not the full empty signature: the method takes // options since #258 ({ force }) and this guard is about ORDER, not arity. const start = source.indexOf('async _maybeRefetchFullHistory('); - const guard = source.indexOf('this._replayWouldShrinkBuffer(buffer)', start); + // Also anchored on the open paren: the guard is handed the rows the caller + // already estimated, and this test is about ORDER, not the argument list. + const guard = source.indexOf('this._replayWouldShrinkBuffer(buffer', start); + const boundedSkip = source.indexOf('boundedShellPull && windowRows <=', start); const reset = source.indexOf('this._resetTerminalForReplay()', start); expect(start).toBeGreaterThan(-1); expect(guard).toBeGreaterThan(start); expect(guard).toBeLessThan(reset); // refuse first, only then reset+rewrite + // A bounded shell window is skipped BEFORE the guard sees it: the guard reads + // "smaller than the browser" as "tmux has nothing more", which a window cut at + // the tail size does not mean (see shell-scroll-history-pull.test.ts). + expect(boundedSkip).toBeGreaterThan(start); + expect(boundedSkip).toBeLessThan(guard); // A hollow pane must also stop re-fetching megabytes on every scroll-up. expect(source).toContain('this._fullHistoryRepullUseless'); expect(source).toContain('this._fullHistoryRepullUseless?.has(sessionId) ? 60000 : 4000');