feat(sessions): offer to rebuild the sessions a host reboot destroyed

A host reboot takes the tmux server down with it, so every pane dies,
reconciliation finds nothing to attach to, and the board comes up empty.
Picking yesterday's work back up meant finding each conversation in history
and resuming it by hand, one at a time.

The boot pass now works out what the reboot killed and leaves it on offer.
It runs inside restoreMuxSessions(), in the window where reconciliation has
reported the dead sessions and cleanupStaleSessions() has not pruned their
records yet, which is the only place the records can still be read. The
board shows a banner, and nothing is created until the user clicks it.

A click rather than an automatic restore is what makes the reboot heuristic
acceptable. The heuristic cannot tell a reboot from a crash that took tmux
down inside the same window, so it decides whether to ASK, never whether to
act: a wrong yes costs a line of text the user dismisses instead of N CLI
processes nobody asked for.

Four things are re-checked when the click arrives rather than trusted from
boot, because hours can pass and the board moves on. The owner's privilege
grant re-resolves through the env clamp. The workspace must still be on
disk. A conversation the user already resumed by hand from the Resume list
is skipped, since two panes running --resume on one conversation would
fight over the same transcript. Entries leave the plan synchronously before
the first await, and the route is single-flighted, so a double-click or two
devices cannot both reach the same entry.

A restored session comes back attached, idle and disarmed. Respawn
controllers and Ralph loops are deliberately not re-armed: a machine that
just came up is the worst moment to turn an autonomous run loose. Its
workspace hooks are installed by the restore route itself, because the
boot-time sweep sits behind a gate that is false after a reboot and has
finished long before the click; without them a session goes silently blind,
with no stop or idle events for respawn, no Approvals Inbox item and no red
tab on a blocking dialog. Stats collection starts the same way.

The pane is new, so the conversation continues and the terminal scrollback
does not. The banner says so rather than letting an empty pane read as a
broken restore.

The plan lives in memory only. A server restart drops it, which costs the
convenience this adds and never the conversation: the conversation is the
transcript under ~/.claude/projects, which the Welcome screen's Resume list
and the Session Manager already read, so a dropped plan returns the user to
resuming by hand.

clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the
question it answers is about session privilege rather than about HTTP and
it now has a caller outside the route layer. Its test hook stays re-exported
from session-routes.ts.

Claude sessions only for this pass. The other CLIs name their thread in
their own config object, which this does not thread through yet. Remote and
docker sessions are skipped on purpose, because both need another host or a
container to be up and a freshly booted machine cannot promise either.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Michael Grundberg
2026-09-16 08:05:55 +02:00
co-authored by Claude Opus 5
parent bd286bf502
commit da933d70be
15 changed files with 1462 additions and 67 deletions
+225
View File
@@ -0,0 +1,225 @@
/**
* @fileoverview Decide which sessions a host reboot destroyed and may be rebuilt.
*
* A server restart and a host reboot both leave `reconcileSessions()` reporting
* dead sessions, and they need opposite handling. A server restart leaves the
* tmux panes running, so recovery ATTACHES to them. A host reboot takes the tmux
* server down with it, so there is nothing to attach to and the pane has to be
* created again. This module holds the decision half of that second case, kept
* free of tmux and disk access so it can be unit tested without either. Every
* observation it reads is gathered by the caller and passed in.
*
* "Eligible" here means a session the user did not end on purpose. The rule that
* an intentional kill or detach is never auto-revived is enforced at runtime by
* an in-memory guard in `TmuxManager`, and memory does not survive a reboot. The
* durable equivalent is the record `cleanupSession()` leaves behind. An unpinned
* kill deletes the record outright, so it is already absent here. A pinned kill
* goes through `demoteOrRemoveSession()` and lands as `status: 'stopped'`, which
* is the marker this module refuses. Pruning keeps a pinned record WITHOUT
* touching its status, so a pinned session a reboot killed still reads `idle` or
* `busy` and stays eligible.
*
* @dependencies types (SessionState), config/cli-registry
* @consumedby web/server (plan build at boot), web/routes/reboot-restore-routes
*
* @module reboot-restore
*/
import type { SessionState } from './types.js';
import { getCli } from './config/cli-registry/registry.js';
/** Session statuses a reboot restore may rebuild. `stopped` is the kill marker. */
const RESTORABLE_STATUSES: ReadonlySet<string> = new Set(['idle', 'busy', 'error']);
/** Observations the reboot heuristic reads. Gathered by the caller, never here. */
export interface RebootEvidence {
/** Sessions that still had a live pane during reconciliation. */
livePaneCount: number;
/** Sessions reconciliation just marked dead. */
deadSessionCount: number;
/** `os.uptime()`, in seconds. */
uptimeSeconds: number;
/** Newest `lastActivityAt` across the persisted records, in ms since the epoch. */
newestPersistedActivityAt: number;
/** `Date.now()` when the evidence was gathered, in ms. */
now: number;
}
/**
* Decide whether the machine plausibly rebooted rather than the server restarting.
*
* Two signals have to agree. The socket must hold no panes at all while state
* still lists sessions, which rules out an ordinary server restart. The host
* must also have booted after the newest persisted session activity, which is
* the corroboration `os.uptime()` provides cheaply. A wiped tmux socket on a
* long-uptime host fails the second test, so a user who killed the tmux server
* by hand does not get every session offered back to them.
*
* This heuristic decides whether to ASK, never whether to act. A wrong yes costs
* the user a banner they dismiss, because the restore itself waits for a click.
*/
export function looksLikeHostReboot(evidence: RebootEvidence): boolean {
if (evidence.deadSessionCount === 0) return false;
if (evidence.livePaneCount > 0) return false;
if (evidence.newestPersistedActivityAt <= 0) return false;
const bootedAt = evidence.now - evidence.uptimeSeconds * 1000;
return bootedAt > evidence.newestPersistedActivityAt;
}
/**
* Pick the conversation the rebuilt pane should resume.
*
* The chain's tail is the newest conversation the session was holding, which is
* what a compact or a clear leaves behind; `resumeSessionId` covers a session
* that was itself started as a resume, and the session id is the original
* conversation for everything else.
*/
export function resolveResumeConversationId(state: SessionState): string {
const chain = state.claudeSessionChain;
const chainTail = Array.isArray(chain) && chain.length > 0 ? chain[chain.length - 1] : undefined;
return chainTail || state.resumeSessionId || state.id;
}
/** Why one session was passed over. Reported for logging and assertions. */
export interface RebootRestoreRejection {
sessionId: string;
reason:
| 'no-persisted-record'
| 'intentionally-ended'
| 'respawn-blocked'
| 'remote-or-docker'
| 'unsupported-mode'
| 'no-working-dir'
| 'workspace-missing'
| 'already-live';
}
/** One restorable session, as the banner shows it and the rebuild replays it. */
export interface RebootRestoreEntry {
sessionId: string;
name?: string;
workingDir: string;
owner?: string;
mode: string;
/** The conversation the rebuilt pane resumes. */
resumeConversationId: string;
/**
* The persisted record, kept whole so the rebuild can replay what it held.
* Read at boot, before pruning deletes it, and held in memory until the click.
*/
state: SessionState;
}
export interface RebootRestorePlan {
restore: RebootRestoreEntry[];
skipped: RebootRestoreRejection[];
}
/**
* Split the sessions reconciliation just killed into the ones a reboot restore
* may offer and the ones it must leave alone.
*
* @param deadSessionIds Session ids `reconcileSessions()` reported as dead.
* @param persisted The `state.json` session records, which `cleanupStaleSessions()`
* has not pruned yet at the point this runs.
* @param workspaceExists Whether a working directory is still on disk. A tmux
* session can outlive its deleted repo, and rebuilding one there would scaffold
* an empty tree. The caller owns the disk access; the click re-checks, because
* a repo can be deleted between the boot and the click.
*/
export function planRebootRestore(
deadSessionIds: readonly string[],
persisted: Readonly<Record<string, SessionState>>,
workspaceExists: (workingDir: string) => boolean
): RebootRestorePlan {
const restore: RebootRestoreEntry[] = [];
const skipped: RebootRestoreRejection[] = [];
for (const sessionId of deadSessionIds) {
const state = persisted[sessionId];
if (!state) {
// An unpinned kill already deleted the record, so absence IS the guard.
skipped.push({ sessionId, reason: 'no-persisted-record' });
continue;
}
if (!RESTORABLE_STATUSES.has(state.status)) {
// A pinned kill was demoted to `stopped`. Reviving it would undo the kill.
skipped.push({ sessionId, reason: 'intentionally-ended' });
continue;
}
if (state.respawnBlocked === true) {
// The crash-loop breaker tripped on this pane. Re-creating it restarts the loop.
skipped.push({ sessionId, reason: 'respawn-blocked' });
continue;
}
if (state.remote || state.docker) {
// Both need another host or a container to be up, which a just-booted machine
// cannot promise. The remote reconnect watcher owns the remote case already.
skipped.push({ sessionId, reason: 'remote-or-docker' });
continue;
}
// Capability, not a CLI id: this pass resumes by handing the CLI a conversation
// id through the top-level `resumeSessionId`, which only a CLI whose history the
// claude-jsonl reader understands can consume that way. Others carry their thread
// id in their own `<Mode>Config`, which this pass does not thread through.
if (getCli(state.mode ?? 'claude')?.capabilities.transcript !== 'claude-jsonl') {
skipped.push({ sessionId, reason: 'unsupported-mode' });
continue;
}
if (!state.workingDir) {
skipped.push({ sessionId, reason: 'no-working-dir' });
continue;
}
if (!workspaceExists(state.workingDir)) {
skipped.push({ sessionId, reason: 'workspace-missing' });
continue;
}
restore.push({
sessionId,
name: state.name,
workingDir: state.workingDir,
owner: state.owner,
mode: state.mode ?? 'claude',
resumeConversationId: resolveResumeConversationId(state),
state,
});
}
return { restore, skipped };
}
/**
* Drop the entries whose conversation is already on screen.
*
* Hours can pass between the boot that built the plan and the click that spends
* it, and the Resume list can reach the same conversation in the meantime. Two
* panes running `claude --resume` on one conversation is the failure this
* prevents, so a match on either the session id or the conversation id is enough
* to skip the entry.
*/
export function rejectAlreadyLive(
entries: readonly RebootRestoreEntry[],
liveSessionIds: ReadonlySet<string>,
liveConversationIds: ReadonlySet<string>
): RebootRestorePlan {
const restore: RebootRestoreEntry[] = [];
const skipped: RebootRestoreRejection[] = [];
for (const entry of entries) {
if (liveSessionIds.has(entry.sessionId) || liveConversationIds.has(entry.resumeConversationId)) {
skipped.push({ sessionId: entry.sessionId, reason: 'already-live' });
continue;
}
restore.push(entry);
}
return { restore, skipped };
}
/** Newest `lastActivityAt` across persisted records, or 0 when there are none. */
export function newestPersistedActivity(persisted: Readonly<Record<string, SessionState>>): number {
let newest = 0;
for (const state of Object.values(persisted)) {
const stamp = state.lastActivityAt ?? state.createdAt ?? 0;
if (stamp > newest) newest = stamp;
}
return newest;
}