feat(sse): heal a stalled SSE stream with a heartbeat + client watchdog

An EventSource that stops delivering does not always error. A proxy that
idle-closed the connection, a laptop resumed from sleep, a tailnet reconnect:
`onerror` never fires, the header dot stays green, and every SSE-driven surface
(tab status dots, sessions created on another device, renames) freezes until the
user reloads. Nothing on the client tracked stream liveness at all.

The server already wrote a keepalive every 15s, but as an SSE `:keepalive`
COMMENT, and comments are invisible to `EventSource` by spec, so there was
nothing a client could observe.

Server:
- `sse:heartbeat` under a new Transport category in the event registry
  (155 constants now, both counts updated).
- `cleanupDeadClients()` writes that named frame (`{"t":<epoch ms>}`) instead of
  the comment. Interval, tunnel padding and dead-socket eviction are unchanged.
  The write stays per-client rather than going through `broadcast()`: the frame
  carries no session data, so it needs no multi-user owner routing.

Client:
- `computeSseStale()` in constants.js, a pure policy beside
  `computeConnectionLossUi`. Stale only when the transport believes it is
  `connected`, the device is online, and no frame has arrived for 45s (three
  missed heartbeats). The `connected`-only guard is also the loop breaker: a
  forced reconnect leaves that state immediately, so the watchdog cannot re-fire
  while one is in flight.
- The liveness stamp is applied inside `addListener` itself, so the
  `_SSE_HANDLER_MAP` wrappers and the directly-registered listeners all feed it
  from one place instead of three that can drift. The heartbeat's own listener
  is a no-op that exists only to be registered, since `EventSource` drops named
  events nobody listens for.
- A 5s watchdog forces `connectSSE()` when the policy says stale, and is cleared
  at the top of `connectSSE()` and nowhere else (its only teardown path).
  Recovery needs no new sync path: the reconnect re-runs `handleInit`, which
  already rebuilds from the server. `visibilitychange` -> visible checks too,
  riding the existing listener, since a background tab's timers are throttled
  and a wake is exactly when a stream comes back zombie.
- The forced reconnect logs one diagnostic line: if a middlebox ever strips or
  delays heartbeats, the failure mode is "silently reconnects every 45s", which
  is undebuggable from a field report without it.

Tests: `test/sse-staleness.test.ts` (node VM over constants.js, threshold
boundaries and every not-stale guard) and `test/sse-heartbeat.test.ts` (drives
`cleanupDeadClients()` with fake replies: named frame not a comment, parseable
payload, padding only with a tunnel, dead clients still evicted).

Verified end to end on an isolated instance: with the stream closed client-side
(no `onerror`), a rename sticks, an out-of-band session stays invisible, then
the watchdog reconnects on its own and it appears without a reload.

Event names are part of the stable API contract, so this is a MINOR bump.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-08-13 17:28:35 +02:00
parent d19895651d
commit c790166564
8 changed files with 434 additions and 11 deletions
+42
View File
@@ -379,6 +379,41 @@ function computeConnectionLossUi(input) {
};
}
// SSE staleness policy: is this stream a zombie?
//
// An EventSource that stops delivering does not always error. A proxy that
// idle-closed the connection, a laptop resumed from sleep, a tailnet
// reconnect: `onerror` never fires, the header dot stays green, and every
// SSE-driven surface (tab status dots, sessions created on another device,
// renames) freezes until the user reloads. The server writes a
// `sse:heartbeat` frame every 15s, so silence longer than three of them means
// the stream is dead even though the transport still claims otherwise.
//
// Stale ONLY when the transport believes it is 'connected': the other states
// already have the reconnect/backoff machinery running, and re-firing on top
// of them would stack reconnects. That guard is also the loop breaker: a
// forced reconnect leaves 'connected' immediately, so the watchdog cannot
// fire again while one is in flight. `navigator.onLine === false` is not
// staleness either; there is nothing to reconnect to yet.
//
// Pure: no DOM, no timers, no side effects. `now` is passed in.
const SSE_STALE_TIMEOUT_MS = 45000; // three missed 15s heartbeats
function computeSseStale(input) {
const {
lastMessageAt = null,
now = 0,
status = 'connected',
isOnline = true,
timeoutMs = SSE_STALE_TIMEOUT_MS,
} = input || {};
if (!isOnline || status !== 'connected') return false;
// No frame has ever arrived: `init` lands on connect, so this is a stream
// that has not opened yet rather than one that went quiet.
if (typeof lastMessageAt !== 'number' || !(lastMessageAt > 0)) return false;
return now - lastMessageAt >= timeoutMs;
}
if (typeof window !== 'undefined') {
window.WEBGL_FALLBACK = WEBGL_FALLBACK;
window.evaluateWebGLLongTaskTrip = evaluateWebGLLongTaskTrip;
@@ -401,6 +436,10 @@ if (typeof window !== 'undefined') {
compute: computeConnectionLossUi,
GRACE_MS: CONNECTION_LOSS_GRACE_MS,
};
window.CodemanSseStale = {
compute: computeSseStale,
TIMEOUT_MS: SSE_STALE_TIMEOUT_MS,
};
}
// Scheduler API — prioritize terminal writes over background UI updates.
@@ -514,6 +553,9 @@ const SSE_EVENTS = {
// Core
INIT: 'init',
// Transport
HEARTBEAT: 'sse:heartbeat',
// Session lifecycle
SESSION_CREATED: 'session:created',
SESSION_UPDATED: 'session:updated',