Files
Codeman/test/claude-cli-version-cache.test.ts
T
Codeman maintainer 9dc4620f03 fix(terminal): stop the scroll-to-top re-pull from deleting history, page the CLI when local scrollback is hollow (#205)
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.

1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
   terminal and rewrites it from the capture, which is a win when tmux holds
   more than the browser, but for a repaint-mode pane that capture is roughly
   ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
   pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
   all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
   (escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
   rewrite when it is more than one screen short; a refused session's cooldown
   goes from 4s to 60s so a hollow pane stops re-fetching megabytes.

2. A false forwarding gate on a Claude session no longer means a dead gesture.
   Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
   travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
   SGR reports, at half a screen of travel per page key. Shift is excluded: it
   keeps meaning "local scrollback".

3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
   exception and guarded on !== undefined, so one timed-out or PATH-starved
   probe at the first Claude session start disabled wheel-forwarding for every
   Claude session until the server restarted, which fits a report of breakage on
   phone, tablet and laptop at once. Success is still cached for the process
   lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
   a pure function so the semantics are testable without spawning claude.

4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
   scoping: the setting keeps meaning exactly what it says, and fix 2 catches
   the case where "local" is empty. The App Settings tooltip now says to leave
   it off for Claude/Codex sessions.

5. _logScrollRouting() prints one line per session per distinct decision:
   forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
   mode, cliVersion, the opt-out state, mouse tracking and local scrollback
   depth. #205 ran two rounds of remote guesswork over questions that line
   answers directly.

Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:41:39 +02:00

110 lines
4.8 KiB
TypeScript

/**
* Issue #205, round 2: `getClaudeCliVersion()` used to cache FAILURE forever.
*
* It stored `null` on any exception and guarded on `!== undefined`, so a single
* failed probe — the 5s exec timeout, a PATH-starved systemd/launchd
* environment, a transient fs hiccup — at the first Claude session start left
* `cliVersion` undefined for every Claude session until the server restarted.
* An undefined `cliVersion` silently disables wheel-forwarding to Claude's own
* transcript (`_shouldForwardWheelToApp`), which is the only route to history
* for a repaint-mode pane: a dead wheel on every device at once, which is what
* the reporter described (phone + iPad + laptop all broken together points at a
* SERVER-side cause, not a browser one).
*
* The probe itself can't run under vitest (it would spawn a real `claude`), so
* these drive the cache policy directly with an injected probe and clock.
*/
import { describe, expect, it, vi } from 'vitest';
import {
claudeVersionRetryDelayMs,
getClaudeCliVersion,
resolveClaudeCliVersion,
type ClaudeVersionProbeState,
} from '../src/utils/claude-cli-resolver.js';
const freshState = (): ClaudeVersionProbeState => ({ failures: 0, lastFailureAt: 0 });
describe('claude --version probe caching', () => {
it('probes once on success and never spawns again', () => {
const state = freshState();
const probe = vi.fn(() => '2.1.223');
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBe('2.1.223');
expect(resolveClaudeCliVersion(state, 2_000, probe)).toBe('2.1.223');
expect(resolveClaudeCliVersion(state, 9_999_999, probe)).toBe('2.1.223');
expect(probe).toHaveBeenCalledTimes(1);
});
it('RETRIES after a failed probe instead of poisoning the process', () => {
const state = freshState();
const probe = vi
.fn<() => string | null>()
.mockImplementationOnce(() => {
throw new Error('spawn claude ETIMEDOUT'); // the shipped failure mode
})
.mockImplementationOnce(() => '2.1.223');
// First session start: probe blows up, no version.
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBeNull();
// Immediately after, the negative cache holds — no probe storm.
expect(resolveClaudeCliVersion(state, 30_000, probe)).toBeNull();
expect(probe).toHaveBeenCalledTimes(1);
// Once the retry window elapses, the next session start probes again and
// wheel-forwarding comes back without a server restart.
expect(resolveClaudeCliVersion(state, 61_000, probe)).toBe('2.1.223');
expect(probe).toHaveBeenCalledTimes(2);
});
it('treats an unparseable version like a failure (retryable, not cached)', () => {
const state = freshState();
const probe = vi.fn<() => string | null>(() => null); // e.g. output without a x.y.z
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBeNull();
expect(resolveClaudeCliVersion(state, 61_000, probe)).toBeNull();
expect(probe).toHaveBeenCalledTimes(2);
expect(state.version).toBeUndefined(); // nothing cached as "known bad"
});
it('clears the failure streak once a probe succeeds', () => {
const state = freshState();
const probe = vi
.fn<() => string | null>()
.mockImplementationOnce(() => null)
.mockImplementationOnce(() => '2.1.223');
resolveClaudeCliVersion(state, 1_000, probe);
expect(state.failures).toBe(1);
resolveClaudeCliVersion(state, 61_000, probe);
expect(state.failures).toBe(0);
expect(state.lastFailureAt).toBe(0);
});
it('backs off so a genuinely missing binary cannot probe on every session start', () => {
expect(claudeVersionRetryDelayMs(0)).toBe(0);
expect(claudeVersionRetryDelayMs(1)).toBe(60_000);
expect(claudeVersionRetryDelayMs(2)).toBe(120_000);
expect(claudeVersionRetryDelayMs(3)).toBe(240_000);
// Capped, so it keeps retrying forever without ever spinning.
expect(claudeVersionRetryDelayMs(50)).toBe(15 * 60_000);
const state = freshState();
const probe = vi.fn<() => string | null>(() => null);
resolveClaudeCliVersion(state, 0, probe); // failure 1 → retry at 60s
resolveClaudeCliVersion(state, 30_000, probe); // still inside the window
expect(probe).toHaveBeenCalledTimes(1);
resolveClaudeCliVersion(state, 60_000, probe); // failure 2 → retry at 120s
resolveClaudeCliVersion(state, 119_000, probe);
expect(probe).toHaveBeenCalledTimes(2);
resolveClaudeCliVersion(state, 180_001, probe);
expect(probe).toHaveBeenCalledTimes(3);
});
it('stays hermetic under vitest without recording a phantom failure', () => {
// The guard returns before the probe, and — unlike the old code, which wrote
// null into the cache here — leaves the cache untouched.
expect(getClaudeCliVersion()).toBeNull();
expect(getClaudeCliVersion()).toBeNull();
});
});