feat(sse): heal a stalled SSE stream with a heartbeat + client watchdog

An EventSource that stops delivering does not always error. A proxy that
idle-closed the connection, a laptop resumed from sleep, a tailnet reconnect:
`onerror` never fires, the header dot stays green, and every SSE-driven surface
(tab status dots, sessions created on another device, renames) freezes until the
user reloads. Nothing on the client tracked stream liveness at all.

The server already wrote a keepalive every 15s, but as an SSE `:keepalive`
COMMENT, and comments are invisible to `EventSource` by spec, so there was
nothing a client could observe.

Server:
- `sse:heartbeat` under a new Transport category in the event registry
  (155 constants now, both counts updated).
- `cleanupDeadClients()` writes that named frame (`{"t":<epoch ms>}`) instead of
  the comment. Interval, tunnel padding and dead-socket eviction are unchanged.
  The write stays per-client rather than going through `broadcast()`: the frame
  carries no session data, so it needs no multi-user owner routing.

Client:
- `computeSseStale()` in constants.js, a pure policy beside
  `computeConnectionLossUi`. Stale only when the transport believes it is
  `connected`, the device is online, and no frame has arrived for 45s (three
  missed heartbeats). The `connected`-only guard is also the loop breaker: a
  forced reconnect leaves that state immediately, so the watchdog cannot re-fire
  while one is in flight.
- The liveness stamp is applied inside `addListener` itself, so the
  `_SSE_HANDLER_MAP` wrappers and the directly-registered listeners all feed it
  from one place instead of three that can drift. The heartbeat's own listener
  is a no-op that exists only to be registered, since `EventSource` drops named
  events nobody listens for.
- A 5s watchdog forces `connectSSE()` when the policy says stale, and is cleared
  at the top of `connectSSE()` and nowhere else (its only teardown path).
  Recovery needs no new sync path: the reconnect re-runs `handleInit`, which
  already rebuilds from the server. `visibilitychange` -> visible checks too,
  riding the existing listener, since a background tab's timers are throttled
  and a wake is exactly when a stream comes back zombie.
- The forced reconnect logs one diagnostic line: if a middlebox ever strips or
  delays heartbeats, the failure mode is "silently reconnects every 45s", which
  is undebuggable from a field report without it.

Tests: `test/sse-staleness.test.ts` (node VM over constants.js, threshold
boundaries and every not-stale guard) and `test/sse-heartbeat.test.ts` (drives
`cleanupDeadClients()` with fake replies: named frame not a comment, parseable
payload, padding only with a tunnel, dead clients still evicted).

Verified end to end on an isolated instance: with the stream closed client-side
(no `onerror`), a rename sticks, an out-of-band session stays invisible, then
the watchdog reconnects on its own and it appears without a reload.

Event names are part of the stable API contract, so this is a MINOR bump.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-08-13 17:28:35 +02:00
parent d19895651d
commit c790166564
8 changed files with 434 additions and 11 deletions
+12 -6
View File
@@ -470,12 +470,20 @@ export class SseStreamManager {
// ========== Client Health ==========
/**
* Clean up dead SSE clients and send keep-alive comments.
* Clean up dead SSE clients and send the liveness heartbeat.
* Keep-alive prevents proxy/load-balancer timeouts on idle connections.
* Dead client cleanup prevents memory leaks from abruptly terminated connections.
*
* The heartbeat is a NAMED event, not the `:keepalive` comment it used to be:
* comments are invisible to `EventSource` by spec, so a stream that stopped
* delivering without erroring was undetectable to the client (see
* `SseEvent.Heartbeat`). Written per-client rather than through `broadcast()`
* deliberately: the frame carries no session data, so it needs no owner
* routing, and this loop is already walking every client to check its socket.
*/
cleanupDeadClients(): void {
const deadClients: FastifyReply[] = [];
const heartbeat = `event: ${SseEvent.Heartbeat}\ndata: ${JSON.stringify({ t: Date.now() })}\n\n`;
for (const [client] of this.sseClients) {
try {
@@ -484,11 +492,9 @@ export class SseStreamManager {
if (!socket || socket.destroyed || !socket.writable) {
deadClients.push(client);
} else {
// Send SSE comment as keep-alive. Only add padding when tunnel is
// active — it flushes Cloudflare proxy buffers but wastes bandwidth
// for direct/Tailscale connections.
const ka = this._isTunnelActive ? ':keepalive\n' + SSE_PADDING : ':keepalive\n\n';
client.raw.write(ka);
// Only add padding when tunnel is active: it flushes Cloudflare
// proxy buffers but wastes bandwidth for direct/Tailscale connections.
client.raw.write(this._isTunnelActive ? heartbeat + SSE_PADDING : heartbeat);
}
} catch {
// Error accessing socket means client is dead