Commit Graph
3 Commits
Author SHA1 Message Date
Codeman maintainer 28b44237ae fix(remote): classify the has-session probe by exit status, and forget it once the pane is back
#355 made the remote auto-reconnect watcher revive a dead pane only when the
durable remote tmux session is verifiably still alive, which is the right rule:
a clean Ctrl-C / Ctrl-D / exit tears that session down and must never relaunch
a fresh agent. Its probe, though, read `has-session`'s stdout and treated an
empty string as "gone". `tmux has-session` prints NOTHING on success (measured
on a scratch socket: exit 0, empty stdout, the failure message goes to stderr),
so every live remote session classified as gone and transport-drop reconnects
were silently disabled along with the clean-exit revives.

The probe now goes by exit status through a pure, unit-tested mapping
(`classifyRemoteAliveExit`): 0 is alive; ssh's own 255, a timeout (`killed`,
no numeric code) and a spawn failure are unknown, which the watcher already
treats as do-not-revive; any other status is the remote command's and means
gone (tmux's 1 for a missing session, 127 when tmux is not installed there).

Two smaller things in the same area:

- The cached answer was never invalidated, so after one successful reattach a
  stale `true` would have revived the NEXT clean exit (the original bug back
  after the first transport drop), and a cached `false` from a clean exit would
  have left a manually restarted session with auto-reconnect permanently off.
  The tick now forgets the cache entry whenever the pane is seen alive.
- The fire-and-forget probe has a 15s timeout against a 5s tick, so an
  unreachable host stacked up to three ssh processes per dead session. An
  in-flight set caps it at one.

The probe command is pinned as a literal string, and the reattach-then-clean-exit
sequence is driven through the watcher in the tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:50:22 +02:00
timkjr da5f5447d0 fix(remote): never auto-revive a remote session after a clean agent exit
The COD-108 reconnect watcher treated any dead local pane as a dropped
transport and re-ran the pane command — so a normal ctrl-c/ctrl-d on a
remote claude/opencode/omp auto-spawned a FRESH agent (claude only
looked correct because its '--session-id || --resume' fallback resumed,
with a loud 'already in use' error first).

Distinguish a transport drop from an intentional exit: only reconnect
when the durable remote tmux session (codeman-ssh-*) is verifiably
still alive on the remote host. A clean exit tears that session down;
the watcher now probes it via ssh has-session and skips (remote-gone)
when it is gone OR unknown (fail closed). The probe is cached
per-session and fired async so the 5s tick never blocks on ssh.

Tests: 3 new cases pinning remote-gone / unknown / alive decisions.
Verified live: all remote CLIs stay dead after ctrl-c/ctrl-d.
2026-08-29 17:43:07 -05:00
Aamer Akhter 6dba8b5227 COD-108 auto-reconnect remote tmux sessions on SSH drop
Continuous remote-only reconnect watcher closing the COD-104 durability
arc: when a remote session's local ssh pane dies mid-run, re-establish it
automatically instead of leaving a dead pane until the user pokes it.

Design decisions (per cod108 design doc):
- D1 event->owner: TmuxManager watcher DETECTS a dead remote pane and emits
  `remoteSessionDropped`; the session owner (server) reassembles the same
  RespawnPaneOptions and calls Session.reattachRemote() -> respawnPane, which
  re-runs the idempotent remote command (owned new-session -A / non-owned
  attach) and REJOINS the still-running durable remote tmux session. The
  watcher never reassembles options itself, and never routes through the
  Claude-idle respawn-controller.
- D2 bounded backoff: per-session exponential backoff [5s,15s,45s,2m,5m,5m],
  reset on a successful reattach, `remoteReconnectExhausted` emitted once after
  the cap. Pure, unit-tested schedule + eligibility decision.
- D3 always-on + kill-switch: `remoteAutoReconnect` app setting (default ON),
  read each tick; when false the watcher does nothing.

Guards: killSession() (incl. the non-owned DETACH early-return) and shutdown
add the session to an intentional-teardown guard set + clear its backoff
BEFORE teardown, so a closed/killed tab is never auto-revived. Exactly one
reconnect in flight per session (inFlight guard prevents stacked respawns).
Per-session reconnect/guard state cleared on session removal.

New: src/remote-reconnect.ts (pure backoff + decideReconnect), TmuxManager
startRemoteReconnectWatcher/stop + runRemoteReconnectTick + noteRemoteReconnect
+ guardRemoteReconnect + clearRemoteReconnectState; Session.reattachRemote()
(+ extracted _buildRespawnPaneOptions, shared with interactive start); server
wiring + watcher start; 3 SSE events (sse-events.ts + constants.js in sync,
broadcast + app.js exhausted "Reconnect" affordance); remoteAutoReconnect
schema + settings-ui toggle.

Tests: test/remote-auto-reconnect.test.ts (21) - pure schedule, eligibility
(guarded never reconnects, non-remote/pane-alive/not-due skip, over-cap
exhaust), and manager-level integration (dead remote pane -> dropped ->
backoff -> exhausted; guarded emits nothing; reset-on-success; kill-switch
off; state-cleared-on-remove). Verified real-remote against aa-desktop: drop
local ssh pane -> watcher emitted -> respawnPane reattached the SAME remote
session (remote pane_pid unchanged 3939->3939); test session cleaned up, the
real host sessions left untouched.

Checks: tsc, eslint, check:frontend-syntax, check:public-assets, prettier
--check, build all green; tmux-manager/session-routes/session-manager/
sse-registry-parity suites pass.

(cherry picked from commit d13d58b1994eb6594fd2eadea208104d36204f9d)
2026-07-17 15:59:36 -04:00