2 Commits
Author SHA1 Message Date
Rounak DattaandClaude Opus 5 abd39318e6 fix(terminal): deadline must cover the body, precache must ignore the cache-bust query
Review fixes. Two of these are defects in the previous commit.

1. The fetch deadline only covered time-to-headers. `await fetch()` settles on
   response headers, so clearing the abort timer in a finally around it left the
   body — the multi-megabyte `?full=1` capture the deadline exists for —
   completely unbounded; it only ever bounded a server that accepts a connection
   and never replies. Measured against a server that sends headers immediately
   and stalls the body 4s under a 1s deadline: fetch resolved at 30ms, timer
   cleared there, body completed at 4026ms unaborted. Now the body is read
   inside `_fetchTerminalCapture`, which returns {json, headers, headersAt} —
   headers because two callers read server-timing, headersAt because those same
   callers measure header-vs-body time and can no longer observe that moment.
   `_terminalCaptureInflight` is scoped the same way, so a body still streaming
   counts toward a capture starting beside it. Same test now aborts at 1005ms.

2. The precache could never be hit, and the previous commit made that expensive
   rather than free. `renderIndexHtml` runs `cacheBustAssets`, which appends
   `?v=<mtime>` to every same-origin .js/.css reference INCLUDING content-hashed
   names — confirmed against a running instance:
   `vendor/xterm-zerolag-input.6fee72f2.js?v=1789402869101`. `caches.match` is
   query-sensitive, so entries keyed on the bare hashed path were unreachable;
   deriving the list from the manifest turned cheap 404s into ~1.3MB downloaded
   at every install that nothing could read back, once per deploy now that
   CACHE_NAME rotates. The fallback match takes `{ ignoreSearch: true }`, which
   also lets runtime-cached entries survive an mtime change.

3. `_wsOutputGapSession` was only cleared in ws.onopen, so paths that already
   repaint the buffer left it set and the socket replayed everything a second
   time. `selectSession` loads the buffer and only THEN calls `_connectWs`, so
   neither the _isLoadingBuffer nor the _terminalRefreshOwner guard applied.
   `_markTerminalBufferReconciled()` is now called from _onSessionNeedsRefresh's
   finally, from selectSession after its load, and from _cleanupSessionData.

   The scope claim was also wrong and is corrected in the comment: when the
   network drops, SSE drops with it and handleInit's keepTerminal branch already
   reconciles. The genuinely uncovered case is the WS dying while SSE stays up,
   where _onSSETerminal discards SSE terminal frames until _wsReady flips in
   onclose — up to the ping+pong window of output nothing writes.

4. CLAUDE.md said "all of them measured rather than reasoned", which the PR's
   own "not verified" section contradicted. Split explicitly: the replay race is
   measured, the watchdog mechanism is verified against xterm 6.0.0 under jsdom
   (field path resolves, a forced stale handle makes refreshRows a no-op, the
   kick schedules a fresh frame), and the iOS rAF-discard premise is reasoned
   and still wants a device. Adds the two missing entries — the WebSocket
   reconcile and the sw.js/build.mjs "keep these in sync or the build throws"
   contract.

Also: test/xterm-private-api.test.ts pins the RESOLVED lockfile version instead
of the declared `^6.0.0` range, which was the wrong assertion in both directions
— a real upgrade to 6.4.0 can rename a private field while resolving inside the
range, and an innocuous range edit failed while changing nothing installed. And
test/sw-precache-manifest.test.ts now parses HASHABLE out of scripts/build.mjs
rather than hand-copying it, which was the same drift this PR exists to fix; the
parse is guarded against silently matching nothing.

The deadline fix has a behavioural test against a real socket plus a source
guard asserting `await res.json()` precedes the finally — verified to fail when
the helper is reverted to the old shape, so it is not vacuous.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 12:25:36 +05:30
Rounak DattaandClaude Opus 5 c0422c4e21 feat(terminal): renderer watchdog, atomic replay clear, fetch deadlines, reconnect recovery
Four ways the terminal can silently stop being correct — in each case the
buffer keeps updating, nothing throws, and the only recourse is a reload.

1. Renderer freeze after backgrounding. iOS DISCARDS scheduled rAF callbacks
   when a PWA backgrounds, and xterm's RenderDebouncer only clears its
   `_animationFrame` handle from inside that callback — so one drop leaves it
   permanently set and every later refresh() early-returns. Parsing is
   decoupled from rendering, so bytes keep filling the buffer correctly while
   nothing paints. Codeman has exactly ONE xterm for the whole page load, so a
   single backgrounding wedges it until a reload. Adds a 2s liveness poll and
   `_kickRenderer()`, which does what the dropped `_innerRefresh` would have.

2. Replay clears raced live output. xterm's write() is async-queued while
   reset() is synchronous and, per upstream, "does not clear input buffers and
   does not reset the parser" — so bytes queued before a reset are parsed after
   it and fuse into the snapshot. Verified against the real xterm 6 here:
   write('p8'); reset(); write('rmissions') renders "p8rmissions". The main
   path was already safe via a queued erase; the needsRefresh and clearTerminal
   paths were not. All three now share one queued `\x1bc` (RIS), which unlike
   3J/H/2J also resets modes, charsets, scroll regions and SGR state.

3. Output lost on WebSocket reconnect. Input frames carry seq+cid and are
   delivered exactly once; output frames carry nothing. ws.onopen re-sends dims
   and flushes queued input, and needsRefresh only fires on external-CLI
   startup and SSE backpressure drain — never on reconnect. Output produced
   while offline was simply absent afterwards. Interim fix: reaching onclose
   means the drop was unintentional, so the session is marked and the next open
   reconciles from the server buffer. Sequencing output is the follow-up.

4. Terminal captures had no deadline. No AbortController anywhere in the
   frontend, including `?full=1`, which the code itself calls "unbounded-ish
   work: at the default history limit it can be megabytes". Adds a budget that
   scales with full-vs-tail and with captures in flight, degrading to a plain
   fetch where AbortController is missing.

Also: the service-worker precache was dead — the build content-hashes assets
but sw.js listed pre-hash names, so 15 of 23 entries 404'd (verified against a
running instance) and cache.add().catch() hid it. Offline still worked via
runtime caching, but CACHE_NAME was a constant so activate's cleanup never
deleted anything and every past release's assets accumulated. Both are now
derived from the build manifest. Crash-trail entries are flattened and capped,
since they are joined with \n into one value and one call site interpolates a
server-controlled WS close reason.

The watchdog reads xterm privates — there is no public API. Every access is
optional-chained so a shape change degrades to a no-op. `_renderService` only
exists after open(), which needs a real DOM, so the gate cannot assert the
field path; test/xterm-private-api.test.ts pins the dependency range instead.

Tests: 23 new (terminal-resilience, sw-precache-manifest, xterm-private-api),
all pure/static so they run in the gate, which excludes the mobile suite. One
static source guard in history-truncation-notice updated for the renamed call;
the behaviour it pins is unchanged.

Not verified: no browser available, so no runtime reproduction of the freeze
and no real-device test of the reconnect path. Both warrant a device pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 12:24:52 +05:30