diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 6a25c5b3..742419f6 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -205,7 +205,7 @@ Further detail: the `: ` form (`w3-myapp: fix the login redirect` ### Full-scrollback replay -**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell full history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. +**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell FULL history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. That gesture instead pulls a BOUNDED window of it (`?full=1&tail=<TERMINAL_TAIL_SIZE>`, so a capture beyond 1 MiB comes back `truncationReason: 'tail'` and the banner offers the rest), and skips the replay when the window holds no more rows than the browser already has. Declining the gesture outright was a dead end: after any burst a shell pane held about one screen of browser scrollback, and the button renders only once a replay was truncated, so a young shell tab had no way back to history tmux still held. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. **A capture reports the geometry it was taken at** (#435): a visible frame repaints each row at an absolute position, counting up to the pane's height and out to the pane's width, so a terminal smaller than that pane damages it two ways at once. Too short and every address past the browser's own height clamps onto the last line, overwriting the rows underneath (measured: against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew the survivors twice). Too narrow and each row is painted out to the pane's width, so the browser wraps every painted row and the wrap on the last one scrolls the whole frame up by one. Nothing in the response used to say what geometry the frame was built for, so the client could not see either case. `PaneCaptureOptions.capturedGeometry` carries it out, and the terminal response publishes it as `captureCols`/`captureRows`. ⚠️ **Both fields are ABSENT unless a frame was really positioned**, and every consumer must test `Number.isFinite` rather than truthiness: `mux-visible` is necessary but not sufficient, because when the `display-message` cursor query fails `capturePaneBuffer` skips the snapshot repaint and returns the raw capture, and the route still labels that non-empty body `mux-visible`. A body that positioned nothing has no geometry to describe and nothing to repair, so a comparison that fires there buys a second capture, a reset plus chunked rewrite, a dropped and reopened WebSocket and a discarded xterm snapshot for no gain. ⚠️ **The comparison runs on `mux-visible` ONLY.** A full-history body is linear scrollback closed by a RELATIVE cursor move, which is relative precisely so the browser's row count need not match the pane's, and a byte-history body carries no row alignment at all, so a size mismatch damages neither and a replay repairs neither. That gate matters because the first select of every non-shell session per page takes the full-history path, where an ungated comparison would fire most often on the one response it cannot help, at the price of a second whole-scrollback capture. ⚠️ **The replay is capped at one attempt and latches per session when it cannot converge.** `resizeRetry` stops two competing fits trading replays forever; a pane already drawing at the size just requested is left alone, because a retry would capture the identical frame; and a pass that still does not converge joins `_geometryRetryUseless`, so the case `Session.resize` declines outright (a small viewport while a desktop viewport's size claim is live, where the retry re-sends the same declined resize and captures the same pane) costs one attempt per session per page load instead of one per select. ⚠️ **The clamp used to manufacture that equality, and no longer can** (#464). This paragraph previously explained it as the signature of a clamp: `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` did not, so a terminal under 40 columns or 10 rows reported a pane permanently bigger than itself and would replay on every tab switch. That divergence is fixed at the source — `syncTerminalGeometry()` (terminal-ui.js) fits, floors and APPLIES in one step, so the browser terminal IS the size it reports. The only remaining reason the two can differ is a resize the server declined, which the server now reports back (`Session.ptyGeometry`, the `{"t":"zc"}` frame and the resize response) for the client to adopt by COLUMNS. The equality guard stays, for the plain case of a pane already at the requested size. ⚠️ **A retry pass must not re-arm `_fullHistoryLoaded`**: it did not consume the full-history pull, and re-arming it would spend a whole-scrollback capture on the next select. That branch is currently unreachable by construction, since reaching it needs `source === 'mux-visible'` while a `full=1` pass is answered `mux-full-history` or `history`; a static test over the source is the habit this repo uses for an invariant nothing can execute. Tests: `test/capture-geometry-retry.browser.test.ts` (eight cases, five of which fail against the merge base), `test/tmux-capture-full-history.test.ts`, `test/routes/session-routes.test.ts`. diff --git a/docs/wiki/The-Dashboard.md b/docs/wiki/The-Dashboard.md index cf1f415e..323299fe 100644 --- a/docs/wiki/The-Dashboard.md +++ b/docs/wiki/The-Dashboard.md @@ -151,8 +151,9 @@ Worth knowing: - **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open. Shell sessions open from a bounded recent tail so a large transcript cannot stall tab - switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling - and automatic output recovery stay within the bounded browser buffer. + switching. Scrolling to the top of a Shell pane pulls the most recent 1 MiB of its tmux + history; press **Load full history** to pull the rest explicitly. Automatic output + recovery stays within the bounded browser buffer. - **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is always local scrollback. Other CLIs scroll locally. diff --git a/src/web/public/app.js b/src/web/public/app.js index 05577c34..1da813be 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -6320,10 +6320,15 @@ class CodemanApp { if (!sessionId || this._fullHistoryRepullInFlight || this._isLoadingBuffer) return; if (this.detachedSessions?.has(sessionId)) return; const session = this.sessions.get(sessionId); - // A shell's full capture can be many megabytes. Replaying it from an - // ordinary scroll gesture blocks xterm's main thread, so keep that cost - // behind the explicit "Load full history" button. - if (!force && session?.mode === 'shell') return; + // A shell's full capture can be many megabytes, and replaying all of it from + // an ordinary scroll gesture blocks xterm's main thread. So a shell scroll + // pulls a BOUNDED window of tmux's full history (the same 1 MiB a tab switch + // loads, but of the scrollback rather than the visible frame) and the + // unbounded pull stays behind the "Load full history" button. Declining + // outright left a shell pane about one screen of browser scrollback after any + // burst, and the button only renders once a replay was truncated, so a young + // shell tab had no way back to output tmux was still holding. + const boundedShellPull = !force && session?.mode === 'shell'; const now = Date.now(); // Momentum scrolling fires this dozens of times per flick, and a burst of new // output is the normal reason to want a re-pull, so cooldown rather than latch. @@ -6336,7 +6341,12 @@ class CodemanApp { this._fullHistoryRepullInFlight = true; try { const requestStartedAt = performance.now(); - const capture = await this._fetchTerminalCapture(`/api/sessions/${sessionId}/terminal?full=1`, { full: true }); + const capture = await this._fetchTerminalCapture( + boundedShellPull + ? `/api/sessions/${sessionId}/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}` + : `/api/sessions/${sessionId}/terminal?full=1`, + { full: true } + ); const headersReceivedAt = capture.headersAt; const payload = capture.json?.data ?? {}; const bodyParsedAt = performance.now(); @@ -6368,6 +6378,20 @@ class CodemanApp { this._setHistoryTruncation(sessionId, { ...payload, exhausted: true }); return; } + // A bounded window no longer than the browser's buffer buys nothing, and + // resetting to rewrite it would jump the viewport on every scroll that + // outlasts the cooldown at the top. An untruncated window IS all of tmux's + // history, so nothing is missing; a window cut at the tail size would trade + // more old rows than it recovers, and its 'tail' truncation keeps the + // banner offering the unbounded pull. Not latched as useless: the next + // burst of output can put more history in tmux than the browser has. + if ( + boundedShellPull && + this._estimateReplayRows(buffer, this.terminal.cols) <= this.terminal.buffer.active.length + ) { + this._setHistoryTruncation(sessionId, payload); + return; + } this._setHistoryTruncation(sessionId, payload); this._fullHistoryRepullUseless?.delete(sessionId); const rowsBefore = this.terminal.buffer.active.length; diff --git a/src/web/public/terminal-ui.js b/src/web/public/terminal-ui.js index 903661e5..413fe99b 100644 --- a/src/web/public/terminal-ui.js +++ b/src/web/public/terminal-ui.js @@ -3304,9 +3304,9 @@ Object.assign(CodemanApp.prototype, { /** * Post-scroll companion to _noteTerminalUserScroll: hitting the TOP of the * buffer while scrolling up gives the app a chance to pull the rest of tmux's - * scrollback (issue #205, see _maybeRefetchFullHistory). Shell sessions decline - * automatic pulls because their captures can be large; their banner button is - * the explicit path. Must be called AFTER scrollLines(), since the check is on + * scrollback (issue #205, see _maybeRefetchFullHistory). Shell sessions pull a + * bounded window because their captures can be large; their banner button is + * the unbounded path. Must be called AFTER scrollLines(), since the check is on * the resulting position, and it is deliberately not folded into * _noteTerminalUserScroll for exactly that reason. */ diff --git a/test/history-truncation-notice.test.ts b/test/history-truncation-notice.test.ts index b24fe3b6..50eb5735 100644 --- a/test/history-truncation-notice.test.ts +++ b/test/history-truncation-notice.test.ts @@ -122,7 +122,7 @@ describe('the in-terminal truncation line is gone (static guard)', () => { expect(app).not.toContain('earlier output truncated for performance'); }); - it('loads a bounded shell tail first and keeps full history user-triggered', () => { + it('loads a bounded shell tail first and keeps unbounded full history user-triggered', () => { const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8'); expect(app).toContain("session?.mode !== 'shell' && !this._fullHistoryLoaded.has(sessionId)"); expect(app).toContain("!restoredSnapshot && session?.mode !== 'shell'"); @@ -131,10 +131,13 @@ describe('the in-terminal truncation line is gone (static guard)', () => { // an abort deadline (a `?full=1` body can be megabytes and used to hang // indefinitely on a stalled mobile link). The URL and the full-vs-tail // decision this guard exists to pin are unchanged. - expect(app).toContain('this._fetchTerminalCapture(`/api/sessions/${sessionId}/terminal?full=1`, { full: true })'); + expect(app).toContain(': `/api/sessions/${sessionId}/terminal?full=1`,\n { full: true }'); expect(app).toContain("if (this.sessions.get(sessionId)?.mode !== 'shell')"); expect(app).toContain("if (session?.mode === 'shell')"); - expect(app).toContain("if (!force && session?.mode === 'shell') return;"); + // A shell scroll gesture pulls a BOUNDED window of full history; only the + // button pulls all of it (behaviour pinned in shell-scroll-history-pull.test.ts). + expect(app).toContain("const boundedShellPull = !force && session?.mode === 'shell';"); + expect(app).toContain('`/api/sessions/${sessionId}/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`'); expect(app).toContain("trigger: force ? 'full-history-button' : 'full-history-scroll'"); }); diff --git a/test/shell-scroll-history-pull.test.ts b/test/shell-scroll-history-pull.test.ts new file mode 100644 index 00000000..229d8a0c --- /dev/null +++ b/test/shell-scroll-history-pull.test.ts @@ -0,0 +1,142 @@ +/** + * @fileoverview A shell pane's scroll-up must reach the history tmux still holds. + * + * tmux repaints a burst of output instead of scrolling it, so after `cat` of a + * file longer than the screen the browser holds about one screen of scrollback + * while tmux holds all of it. Other modes recover it by re-pulling `?full=1` + * when the wheel reaches the top (`_maybeRefetchFullHistory`, issue #205). Shell + * declined that gesture outright to keep a multi-megabyte capture off xterm's + * main thread, leaving only the "Load full history" button, and that button + * renders only once a replay was truncated. A young shell tab therefore had no + * way to scroll back at all. + * + * The gesture now pulls a BOUNDED window (`?full=1&tail=TERMINAL_TAIL_SIZE`), + * the button stays the unbounded path, and a window the browser already holds + * in full is not rewritten. + * + * The method is extracted from app.js and run in a `vm` against stubs (no jsdom + * on this box; see connection-indicator.test.ts), with the REAL row estimators + * from terminal-ui.js, which decide both the downgrade and the no-gain skip. + */ +import { readFileSync } from 'node:fs'; +import { performance } from 'node:perf_hooks'; +import { resolve } from 'node:path'; +import vm from 'node:vm'; +import { describe, expect, it, vi } from 'vitest'; + +const PUBLIC = resolve(import.meta.dirname, '../src/web/public'); +const TERMINAL_TAIL_SIZE = 1024 * 1024; + +function methodSource(source: string, method: string): string { + const start = source.search(new RegExp(`^ {2}(?:async )?${method}\\(`, 'm')); + expect(start, `${method} not found`).toBeGreaterThan(-1); + const next = /^ {2}(?:async )?[A-Za-z_$][\w$]*\(/m.exec(source.slice(start + 1)); + return next ? source.slice(start, start + 1 + next.index) : source.slice(start); +} + +/** Real terminal-ui.js mixin, for `_estimateReplayRows` / `_replayWouldShrinkBuffer`. */ +function loadTerminalMixin(): Record<string, unknown> { + const source = readFileSync(resolve(PUBLIC, 'terminal-ui.js'), 'utf8'); + const FakeCodemanApp = function () {} as unknown as { prototype: Record<string, unknown> }; + const context = vm.createContext({ + console, + performance, + setTimeout, + clearTimeout, + setInterval: vi.fn(), + clearInterval: vi.fn(), + requestAnimationFrame: vi.fn(), + CodemanApp: FakeCodemanApp, + window: { addEventListener: vi.fn(), removeEventListener: vi.fn() }, + document: { addEventListener: vi.fn() }, + }); + vm.runInContext(source, context); + return FakeCodemanApp.prototype; +} + +function loadRefetch(): (this: unknown, opts?: { force?: boolean }) => Promise<void> { + const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8'); + const body = methodSource(app, '_maybeRefetchFullHistory'); + const context = vm.createContext({ performance, TERMINAL_TAIL_SIZE, TERMINAL_CHUNK_SIZE: 32 * 1024 }); + return vm.runInContext(`({ ${body} })._maybeRefetchFullHistory`, context); +} + +const mixin = loadTerminalMixin(); +const refetch = loadRefetch(); +const lines = (n: number) => Array.from({ length: n }, (_, i) => `line ${i}`).join('\r\n'); + +function makeApp(mode: string, { bufferRows, capture }: { bufferRows: number; capture: string }) { + const urls: string[] = []; + const app = { + activeSessionId: 's1', + sessions: new Map([['s1', { mode }]]), + detachedSessions: new Set<string>(), + _fullHistoryRepullInFlight: false, + _isLoadingBuffer: false, + _fullHistoryRepullAt: new Map<string, number>(), + _fullHistoryRepullUseless: new Set<string>(), + terminalBufferCache: new Map<string, string>(), + terminal: { + cols: 80, + rows: 30, + buffer: { active: { length: bufferRows } }, + scrollToLine: vi.fn(), + scrollToTop: vi.fn(), + }, + _estimateReplayRows: mixin._estimateReplayRows, + _replayWouldShrinkBuffer: mixin._replayWouldShrinkBuffer, + _fetchTerminalCapture: vi.fn(async (url: string) => { + urls.push(url); + return { + headersAt: performance.now(), + headers: { get: () => '' }, + json: { data: { terminalBuffer: capture, source: 'mux-full-history' } }, + }; + }), + _recordTerminalLoadTiming: vi.fn(), + _logScrollRouting: vi.fn(), + _setHistoryTruncation: vi.fn(), + _resetTerminalForReplay: vi.fn(), + _bufferLoadFinishOpts: vi.fn(() => ({})), + chunkedTerminalWrite: vi.fn(async () => ({ parsedAt: performance.now(), bufferLength: 400, completed: true })), + _syncStickyScrollBaseline: vi.fn(), + }; + return { app, urls }; +} + +describe('shell scroll-up pulls a bounded window of tmux history', () => { + it('a shell scroll gesture requests full history bounded by the tail size', async () => { + const { app, urls } = makeApp('shell', { bufferRows: 40, capture: lines(300) }); + await refetch.call(app); + expect(urls).toEqual([`/api/sessions/s1/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`]); + // It then actually replays the recovered history. + expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1); + expect(app.chunkedTerminalWrite).toHaveBeenCalledTimes(1); + }); + + it('the Load full history button stays unbounded for a shell', async () => { + const { app, urls } = makeApp('shell', { bufferRows: 40, capture: lines(300) }); + await refetch.call(app, { force: true }); + expect(urls).toEqual(['/api/sessions/s1/terminal?full=1']); + }); + + it('other modes keep the unbounded scroll pull', async () => { + const { app, urls } = makeApp('claude', { bufferRows: 40, capture: lines(300) }); + await refetch.call(app); + expect(urls).toEqual(['/api/sessions/s1/terminal?full=1']); + }); + + it('a bounded window the browser already holds is not rewritten', async () => { + // Browser already has every row the window carries: resetting to rewrite + // it would jump the viewport on every scroll that outlasts the cooldown. + const { app } = makeApp('shell', { bufferRows: 320, capture: lines(300) }); + await refetch.call(app); + expect(app._resetTerminalForReplay).not.toHaveBeenCalled(); + expect(app.chunkedTerminalWrite).not.toHaveBeenCalled(); + // Not latched as useless: more output can put more history in tmux. + expect(app._fullHistoryRepullUseless.has('s1')).toBe(false); + // The truncation state is still recorded, so a window capped at the tail + // size keeps offering the button. + expect(app._setHistoryTruncation).toHaveBeenCalledTimes(1); + }); +});