fix(terminal): merge-time fixes for dropped-output recovery (#470)

- The TERMINAL DROP crash-trail line moves behind the scheduler's debounce
  guard, so it is written once per window rather than once per dropped
  frame. At the server's 8ms batching, one second of drops evicted the whole
  50-entry trail, including the recovery lines that explain it.
- A refresh that failed at the capture fetch deadline now returns
  'deadline', and the scheduler does not retry it: that is a stalled link,
  not contention, and each retry was another ?full=1 capture waiting out a
  deadline of up to two minutes. The early-return retries are unchanged.
  CLAUDE.md and the code comments no longer claim every skip reason is
  transient contention.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-09-23 11:40:03 +02:00
parent 7a30a31430
commit f4d1ee8027
4 changed files with 77 additions and 20 deletions
+11 -6
View File
@@ -1713,27 +1713,32 @@ function sanitizeDiagEntry(msg) {
// so a skipped refresh lost the recovery silently and the dropped bytes were
// never replayed.
//
// Bounded, because every reason the refresh can be skipped is transient
// contention that clears in seconds, and a permanently failing refresh must not
// become a forever-loop against the API. Giving up after the cap leaves exactly
// the garbled frames the old code left, so the floor is no worse than before.
// Bounded, because the early returns it retries past are transient contention
// that clears in seconds, and a permanently failing refresh must not become a
// forever-loop against the API. A refresh that hit the capture fetch DEADLINE
// is not contention but a stalled link, and is not retried at all: each retry
// would be another `?full=1` capture waiting out a deadline of up to two
// minutes, where the old code cost exactly one. Giving up after the cap leaves
// exactly the garbled frames the old code left, so the floor is no worse.
const DROP_RECOVERY_DELAY_MS = 2000;
const DROP_RECOVERY_MAX_ATTEMPTS = 5;
/**
* Should a dropped-output recovery run again?
*
* @param {{repainted: boolean, attempt: number, stillActive: boolean}} state
* @param {{repainted: boolean, timedOut?: boolean, attempt: number, stillActive: boolean}} state
* `repainted` — whether `_onSessionNeedsRefresh` actually rewrote the buffer.
* `timedOut` - whether it failed at the capture fetch deadline.
* `attempt` — how many have already run, zero-based.
* `stillActive` — whether the dropped session is still the one on screen.
* @returns {boolean}
*/
function shouldRetryDroppedOutputRecovery({ repainted, attempt, stillActive }) {
function shouldRetryDroppedOutputRecovery({ repainted, timedOut = false, attempt, stillActive }) {
// Switched away: `selectSession` repaints from the server on its own, so a
// retry here would be a second replay of a buffer that is about to be written.
if (!stillActive) return false;
if (repainted) return false;
if (timedOut) return false;
return attempt + 1 < DROP_RECOVERY_MAX_ATTEMPTS;
}