feat(tmux): report a dead pane's exit from the batched pane list

Codeman creates every tmux pane with `remain-on-exit on`. When the agent exits,
tmux keeps the pane, the tmux session, and the `tmux attach-session` process
Codeman records as the session's pid, so no PTY exit handler fires and nothing
writes the exit down. tmux itself knows: it marks the pane dead and reports the
exit status. This reads that.

`PANE_LIST_FORMAT` gains `#{pane_dead}`, `#{pane_dead_status}` and
`#{pane_dead_signal}`, and `startPaneExitWatcher()` refreshes a
muxName-to-observation map from ONE batched `tmux list-panes -a` per tick. Boot
reconciliation already ran that same call, so it now fills the map too and
recovery starts with a reading.

The watcher owns its own interval rather than riding `startStatsCollection()`,
which the issue suggested. That collector is armed when a browser opens the
Monitor panel and DISARMED when it closes it, and boot skips it entirely unless
recovery found a live session, so a session created on a freshly booted server
would publish nothing and one browser could turn detection off for every other.
Measured on an isolated instance: a dead pane with status 0 reported nothing
until `POST /api/mux-sessions/stats/start` was called by hand. It is still one
batched read per tick; only the timer changed.

Three rules keep a positive answer trustworthy. A session answers only when
tmux listed exactly one pane for it, because Codeman never splits a pane and a
session the user split by hand has none that speaks for the agent. A pane
answers only when `#{pane_dead}` said 1 or 0, because an empty field is a tmux
that did not answer. An absent status stays absent rather than becoming 0:
measured on tmux 3.2a, a SIGKILLed pane reports neither a status nor a signal,
and calling that a clean exit would be wrong in the direction that matters.

Two guards stop a slow read undoing a fast one. `EXEC_TIMEOUT_MS` is 5000 ms
against a 2000 ms interval, so a read can outlive two ticks: one already in
flight suppresses the next, and a generation counter that every
`clearPaneExit()` bumps discards a read that started before a respawn or a
kill. An observation also carries its pane pid, so a second command in the same
pane that exits the same way starts a new timestamp rather than inheriting the
first death's.

A non-empty read of `list-panes -a` is authoritative for the whole socket, so
sessions missing from it are pruned, which also bounds the map as tmux sessions
come and go outside `killSession()`. A failed or empty read retracts nothing.

The manager reports the raw pane reading and applies no session-shape scoping,
because the remote-reconnect watcher beside it needs exactly that raw reading.

`parsePaneList` becomes `parsePaneRows`, returning one row per pane instead of
a name-to-pid map; reconciliation builds its map from the rows. The parser's
existing cases carry over unchanged, including the launchd/systemd literal-tab
regression from PR #71.

Refs Ark0N/Codeman#446.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Michael Grundberg
2026-09-21 10:00:54 +02:00
co-authored by Claude Opus 5
parent 9466acfc1a
commit 02dc46dcd7
4 changed files with 546 additions and 32 deletions
+30
View File
@@ -24,6 +24,7 @@ import type {
OmpConfig,
SessionRemote,
SessionDocker,
PaneExit,
} from './types.js';
/**
@@ -56,6 +57,16 @@ export interface MuxSession {
respawnConfig?: PersistedRespawnConfig;
/** Whether Ralph / Todo tracking is enabled */
ralphEnabled?: boolean;
/**
* This record was rebuilt from the tmux socket rather than from Codeman's own
* bookkeeping, so everything on it but the name and the pid is a guess. Its
* synthetic `restored-<fragment>` id cannot find the session's `state.json`
* entry either, which means a remote or docker session rediscovered this way
* arrives with no `remote`/`docker` metadata and looks local. Anything that
* would be WRONG about such a session rather than merely vague must fail
* closed on this flag.
*/
discovered?: boolean;
}
/**
@@ -180,6 +191,7 @@ export interface PaneCaptureOptions {
* - `sessionKilled` (data: { sessionId: string }) - Session terminated
* - `sessionDied` (data: { sessionId: string }) - Session died unexpectedly
* - `statsUpdated` (sessions: MuxSessionWithStats[]) - Stats refreshed
* - `paneExitsUpdated` () - A pane read finished; ask `getPaneExit()` per session
*/
export interface TerminalMultiplexer extends EventEmitter {
/** Which backend this instance uses */
@@ -294,6 +306,24 @@ export interface TerminalMultiplexer extends EventEmitter {
/** Check if the pane in a session is dead (command exited but remain-on-exit keeps it alive) */
isPaneDead(muxName: string): boolean;
/**
* What the last pane read saw of this session's agent, or `undefined` for
* UNKNOWN (Ark0N/Codeman#446). Unlike `isPaneDead()` this costs nothing: it
* reads a map the batched watcher fills, so it answers no fresher than that
* watcher's interval and the three synchronous `isPaneDead()` callers still
* need their own probe. See {@link PaneExit}.
*/
getPaneExit?(muxName: string): PaneExit | undefined;
/** Forget a session's exit observation, e.g. once its pane has been respawned. */
clearPaneExit?(muxName: string): void;
/** Start polling every pane on the socket for an exited agent. */
startPaneExitWatcher?(intervalMs?: number): void;
/** Stop the pane-exit watcher. */
stopPaneExitWatcher?(): void;
/** Respawn a dead pane with a fresh command. Returns the new PID or null on failure. */
respawnPane(options: RespawnPaneOptions): Promise<number | null>;