feat(sessions): offer to rebuild the sessions a host reboot destroyed

A host reboot takes the tmux server down with it, so every pane dies,
reconciliation finds nothing to attach to, and the board comes up empty.
Picking yesterday's work back up meant finding each conversation in history
and resuming it by hand, one at a time.

The boot pass now works out what the reboot killed and leaves it on offer.
It runs inside restoreMuxSessions(), in the window where reconciliation has
reported the dead sessions and cleanupStaleSessions() has not pruned their
records yet, which is the only place the records can still be read. The
board shows a banner, and nothing is created until the user clicks it.

A click rather than an automatic restore is what makes the reboot heuristic
acceptable. The heuristic cannot tell a reboot from a crash that took tmux
down inside the same window, so it decides whether to ASK, never whether to
act: a wrong yes costs a line of text the user dismisses instead of N CLI
processes nobody asked for.

Four things are re-checked when the click arrives rather than trusted from
boot, because hours can pass and the board moves on. The owner's privilege
grant re-resolves through the env clamp. The workspace must still be on
disk. A conversation the user already resumed by hand from the Resume list
is skipped, since two panes running --resume on one conversation would
fight over the same transcript. Entries leave the plan synchronously before
the first await, and the route is single-flighted, so a double-click or two
devices cannot both reach the same entry.

A restored session comes back attached, idle and disarmed. Respawn
controllers and Ralph loops are deliberately not re-armed: a machine that
just came up is the worst moment to turn an autonomous run loose. Its
workspace hooks are installed by the restore route itself, because the
boot-time sweep sits behind a gate that is false after a reboot and has
finished long before the click; without them a session goes silently blind,
with no stop or idle events for respawn, no Approvals Inbox item and no red
tab on a blocking dialog. Stats collection starts the same way.

The pane is new, so the conversation continues and the terminal scrollback
does not. The banner says so rather than letting an empty pane read as a
broken restore.

The plan lives in memory only. A server restart drops it, which costs the
convenience this adds and never the conversation: the conversation is the
transcript under ~/.claude/projects, which the Welcome screen's Resume list
and the Session Manager already read, so a dropped plan returns the user to
resuming by hand.

clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the
question it answers is about session privilege rather than about HTTP and
it now has a caller outside the route layer. Its test hook stays re-exported
from session-routes.ts.

Claude sessions only for this pass. The other CLIs name their thread in
their own config object, which this does not thread through yet. Remote and
docker sessions are skipped on purpose, because both need another host or a
container to be up and a freshly booted machine cannot promise either.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Michael Grundberg
2026-09-16 08:05:55 +02:00
co-authored by Claude Opus 5
parent bd286bf502
commit da933d70be
15 changed files with 1462 additions and 67 deletions
+5
View File
@@ -0,0 +1,5 @@
---
'aicodeman': minor
---
Offer to rebuild the sessions a host reboot destroyed. A reboot takes the tmux server down with it, so every pane dies and the board comes up empty. Codeman now works out what was running, and the board offers to restore it behind a click. The conversations come back; the terminal scrollback does not, and the banner says so.
+225
View File
@@ -0,0 +1,225 @@
/**
* @fileoverview Decide which sessions a host reboot destroyed and may be rebuilt.
*
* A server restart and a host reboot both leave `reconcileSessions()` reporting
* dead sessions, and they need opposite handling. A server restart leaves the
* tmux panes running, so recovery ATTACHES to them. A host reboot takes the tmux
* server down with it, so there is nothing to attach to and the pane has to be
* created again. This module holds the decision half of that second case, kept
* free of tmux and disk access so it can be unit tested without either. Every
* observation it reads is gathered by the caller and passed in.
*
* "Eligible" here means a session the user did not end on purpose. The rule that
* an intentional kill or detach is never auto-revived is enforced at runtime by
* an in-memory guard in `TmuxManager`, and memory does not survive a reboot. The
* durable equivalent is the record `cleanupSession()` leaves behind. An unpinned
* kill deletes the record outright, so it is already absent here. A pinned kill
* goes through `demoteOrRemoveSession()` and lands as `status: 'stopped'`, which
* is the marker this module refuses. Pruning keeps a pinned record WITHOUT
* touching its status, so a pinned session a reboot killed still reads `idle` or
* `busy` and stays eligible.
*
* @dependencies types (SessionState), config/cli-registry
* @consumedby web/server (plan build at boot), web/routes/reboot-restore-routes
*
* @module reboot-restore
*/
import type { SessionState } from './types.js';
import { getCli } from './config/cli-registry/registry.js';
/** Session statuses a reboot restore may rebuild. `stopped` is the kill marker. */
const RESTORABLE_STATUSES: ReadonlySet<string> = new Set(['idle', 'busy', 'error']);
/** Observations the reboot heuristic reads. Gathered by the caller, never here. */
export interface RebootEvidence {
/** Sessions that still had a live pane during reconciliation. */
livePaneCount: number;
/** Sessions reconciliation just marked dead. */
deadSessionCount: number;
/** `os.uptime()`, in seconds. */
uptimeSeconds: number;
/** Newest `lastActivityAt` across the persisted records, in ms since the epoch. */
newestPersistedActivityAt: number;
/** `Date.now()` when the evidence was gathered, in ms. */
now: number;
}
/**
* Decide whether the machine plausibly rebooted rather than the server restarting.
*
* Two signals have to agree. The socket must hold no panes at all while state
* still lists sessions, which rules out an ordinary server restart. The host
* must also have booted after the newest persisted session activity, which is
* the corroboration `os.uptime()` provides cheaply. A wiped tmux socket on a
* long-uptime host fails the second test, so a user who killed the tmux server
* by hand does not get every session offered back to them.
*
* This heuristic decides whether to ASK, never whether to act. A wrong yes costs
* the user a banner they dismiss, because the restore itself waits for a click.
*/
export function looksLikeHostReboot(evidence: RebootEvidence): boolean {
if (evidence.deadSessionCount === 0) return false;
if (evidence.livePaneCount > 0) return false;
if (evidence.newestPersistedActivityAt <= 0) return false;
const bootedAt = evidence.now - evidence.uptimeSeconds * 1000;
return bootedAt > evidence.newestPersistedActivityAt;
}
/**
* Pick the conversation the rebuilt pane should resume.
*
* The chain's tail is the newest conversation the session was holding, which is
* what a compact or a clear leaves behind; `resumeSessionId` covers a session
* that was itself started as a resume, and the session id is the original
* conversation for everything else.
*/
export function resolveResumeConversationId(state: SessionState): string {
const chain = state.claudeSessionChain;
const chainTail = Array.isArray(chain) && chain.length > 0 ? chain[chain.length - 1] : undefined;
return chainTail || state.resumeSessionId || state.id;
}
/** Why one session was passed over. Reported for logging and assertions. */
export interface RebootRestoreRejection {
sessionId: string;
reason:
| 'no-persisted-record'
| 'intentionally-ended'
| 'respawn-blocked'
| 'remote-or-docker'
| 'unsupported-mode'
| 'no-working-dir'
| 'workspace-missing'
| 'already-live';
}
/** One restorable session, as the banner shows it and the rebuild replays it. */
export interface RebootRestoreEntry {
sessionId: string;
name?: string;
workingDir: string;
owner?: string;
mode: string;
/** The conversation the rebuilt pane resumes. */
resumeConversationId: string;
/**
* The persisted record, kept whole so the rebuild can replay what it held.
* Read at boot, before pruning deletes it, and held in memory until the click.
*/
state: SessionState;
}
export interface RebootRestorePlan {
restore: RebootRestoreEntry[];
skipped: RebootRestoreRejection[];
}
/**
* Split the sessions reconciliation just killed into the ones a reboot restore
* may offer and the ones it must leave alone.
*
* @param deadSessionIds Session ids `reconcileSessions()` reported as dead.
* @param persisted The `state.json` session records, which `cleanupStaleSessions()`
* has not pruned yet at the point this runs.
* @param workspaceExists Whether a working directory is still on disk. A tmux
* session can outlive its deleted repo, and rebuilding one there would scaffold
* an empty tree. The caller owns the disk access; the click re-checks, because
* a repo can be deleted between the boot and the click.
*/
export function planRebootRestore(
deadSessionIds: readonly string[],
persisted: Readonly<Record<string, SessionState>>,
workspaceExists: (workingDir: string) => boolean
): RebootRestorePlan {
const restore: RebootRestoreEntry[] = [];
const skipped: RebootRestoreRejection[] = [];
for (const sessionId of deadSessionIds) {
const state = persisted[sessionId];
if (!state) {
// An unpinned kill already deleted the record, so absence IS the guard.
skipped.push({ sessionId, reason: 'no-persisted-record' });
continue;
}
if (!RESTORABLE_STATUSES.has(state.status)) {
// A pinned kill was demoted to `stopped`. Reviving it would undo the kill.
skipped.push({ sessionId, reason: 'intentionally-ended' });
continue;
}
if (state.respawnBlocked === true) {
// The crash-loop breaker tripped on this pane. Re-creating it restarts the loop.
skipped.push({ sessionId, reason: 'respawn-blocked' });
continue;
}
if (state.remote || state.docker) {
// Both need another host or a container to be up, which a just-booted machine
// cannot promise. The remote reconnect watcher owns the remote case already.
skipped.push({ sessionId, reason: 'remote-or-docker' });
continue;
}
// Capability, not a CLI id: this pass resumes by handing the CLI a conversation
// id through the top-level `resumeSessionId`, which only a CLI whose history the
// claude-jsonl reader understands can consume that way. Others carry their thread
// id in their own `<Mode>Config`, which this pass does not thread through.
if (getCli(state.mode ?? 'claude')?.capabilities.transcript !== 'claude-jsonl') {
skipped.push({ sessionId, reason: 'unsupported-mode' });
continue;
}
if (!state.workingDir) {
skipped.push({ sessionId, reason: 'no-working-dir' });
continue;
}
if (!workspaceExists(state.workingDir)) {
skipped.push({ sessionId, reason: 'workspace-missing' });
continue;
}
restore.push({
sessionId,
name: state.name,
workingDir: state.workingDir,
owner: state.owner,
mode: state.mode ?? 'claude',
resumeConversationId: resolveResumeConversationId(state),
state,
});
}
return { restore, skipped };
}
/**
* Drop the entries whose conversation is already on screen.
*
* Hours can pass between the boot that built the plan and the click that spends
* it, and the Resume list can reach the same conversation in the meantime. Two
* panes running `claude --resume` on one conversation is the failure this
* prevents, so a match on either the session id or the conversation id is enough
* to skip the entry.
*/
export function rejectAlreadyLive(
entries: readonly RebootRestoreEntry[],
liveSessionIds: ReadonlySet<string>,
liveConversationIds: ReadonlySet<string>
): RebootRestorePlan {
const restore: RebootRestoreEntry[] = [];
const skipped: RebootRestoreRejection[] = [];
for (const entry of entries) {
if (liveSessionIds.has(entry.sessionId) || liveConversationIds.has(entry.resumeConversationId)) {
skipped.push({ sessionId: entry.sessionId, reason: 'already-live' });
continue;
}
restore.push(entry);
}
return { restore, skipped };
}
/** Newest `lastActivityAt` across persisted records, or 0 when there are none. */
export function newestPersistedActivity(persisted: Readonly<Record<string, SessionState>>): number {
let newest = 0;
for (const state of Object.values(persisted)) {
const stamp = state.lastActivityAt ?? state.createdAt ?? 0;
if (stamp > newest) newest = stamp;
}
return newest;
}
+91
View File
@@ -0,0 +1,91 @@
/**
* @fileoverview The env-var half of the multi-user privilege clamp.
*
* A session's `envOverrides` can hand back privilege that the per-CLI config
* clamp removed, so a non-granted owner's overrides get the privileged keys
* stripped before the session is built. Two callers need that today. The create
* and resume routes clamp what a request asked for, and the reboot-restore route
* clamps what a persisted record carried, because a record written while its
* owner held a grant must not replay that grant after the grant is gone.
*
* This lives outside `web/routes` on purpose. The question it answers is about
* session privilege rather than about HTTP, and `cron/cron-service.ts` sets the
* precedent by importing `canUsernameRunPrivilegedCommands` from `user-store.ts`
* directly and re-resolving the owner's grant when a job fires. Every caller here
* re-resolves the grant at the moment it builds a session, for the same reason.
*
* @dependencies user-store (canUsernameRunPrivilegedCommands), config/cli-registry
* @consumedby web/routes/session-routes, web/routes/reboot-restore-routes
*
* @module session-env-clamp
*/
import { canUsernameRunPrivilegedCommands } from './user-store.js';
import { enabledClis } from './config/cli-registry/registry.js';
/**
* Env-var keys a non-granted owner must not be able to set, because each one
* hands back privilege `clampExternalCliBypassForOwner()` just removed, or redirects a
* credential-resolution endpoint.
*
* The DeepSeek three are reachable because `DSH_*` and `DEEPSEEK_*` are
* allowlisted `envOverrides` prefixes (schemas.ts) — which they have to be, since
* that is also how a user configures the harness's non-privileged knobs.
*
* - `DSH_PERMISSION_MODE` IS the harness's permission switch. Every other CLI's
* bypass is a command-line FLAG, reachable only through the per-CLI config the
* clamp already owns; this one is an env var, so the config clamp alone is
* half a gate.
* - `DSH_HOME` points the launcher at a profile tree, and a profile's plugin code
* executes at BOOT, before any approval row can apply. A user who can write a
* workspace can put a profile in it, so this is the wider of the two.
* - `DEEPSEEK_BASE_URL` aims the provider endpoint, and `_configureCliEnv()`
* forwards the SERVER's own `DEEPSEEK_API_KEY` into every dsh pane before
* `applyEnvOverrides()` runs — so a non-granted owner who could set the base
* URL would have the operator's API key sent as a bearer credential to a host
* of their choosing. (`DEEPSEEK_API_KEY` itself stays overridable: supplying
* your OWN key removes privilege rather than granting it.)
* - `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN` are where omp resolves
* credentials from — the same shape as `DEEPSEEK_BASE_URL` above, reachable
* because `OMP_*` is an allowlisted prefix. Unlike DeepSeek, Codeman does not
* forward any operator-held key into an omp pane today (omp's provider
* credentials live in `~/.omp` config files, not env vars), so there is no
* known concrete exfiltration path yet — clamped defensively anyway, since a
* non-granted owner redirecting where a shared multi-tenant deployment
* resolves auth from is not something to allow silently (found in
* Ark0N/Codeman#353 review; omp's own knobs are otherwise mostly `PI_*`,
* already allowlisted for pi and not addressed here — see resolveOmpHome()).
*/
export function ownerClampedEnvKeys(): string[] {
return enabledClis().flatMap((entry) => entry.capabilities.privilegedEnvKeys);
}
/**
* Env-var half of the multi-user bypass clamp.
*
* `clampExternalCliBypassForOwner()` in `web/routes/session-routes.ts` clamps the
* per-CLI CONFIG, and for every CLI
* but DeepSeek that is the whole story. Here it is not: `applyEnvOverrides()` runs
* AFTER `_configureCliEnv()` in tmux-manager, so an override sent on the SAME
* request lands last and wins, and a non-granted owner could restore
* `danger-full-access` on the very request the config clamp downgraded.
*
* Keys are DROPPED rather than rewritten: dropping falls through to what
* `_configureCliEnv()` exports, which is the clamped config and the server's own
* `DSH_HOME`, i.e. exactly the intended state. No-op in single-user mode and for a
* granted owner, like every other clamp here
* (`canUsernameRunPrivilegedCommands()` returns true when `!isMultiUserMode()`),
* and it returns the caller's own object untouched when there is nothing to strip.
*/
export async function clampEnvOverridesForOwner(
owner: string | undefined,
envOverrides: Record<string, string> | undefined
): Promise<Record<string, string> | undefined> {
if (!envOverrides) return envOverrides;
const keys = ownerClampedEnvKeys();
if (!keys.some((key) => key in envOverrides)) return envOverrides;
if (await canUsernameRunPrivilegedCommands(owner)) return envOverrides;
const clamped = { ...envOverrides };
for (const key of keys) delete clamped[key];
return clamped;
}
+2
View File
@@ -957,6 +957,8 @@ class CodemanApp {
this.registerServiceWorker();
// Fetch tunnel status for header indicator (desktop only)
this.loadTunnelStatus();
// Ask whether a host reboot left sessions worth rebuilding (banner, never automatic)
this.initRebootRestoreBanner?.();
// Share a single settings fetch between both consumers
const settingsPromise = fetch('/api/settings').then(r => r.ok ? r.json() : null).then(env => env?.data ?? null).catch(() => null);
this.loadQuickStartCases(null, settingsPromise);
+19
View File
@@ -213,6 +213,24 @@
<button class="offline-banner-retry" id="offlineBannerRetry" onclick="app.retryConnection()">Retry now</button>
</div>
<!-- Reboot-restore offer: shown when the server found sessions a host reboot
killed and is asking whether to rebuild them. Populated by
reboot-restore-ui.js; nothing is created until the user clicks. -->
<div class="reboot-restore-banner" id="rebootRestoreBanner" role="status" hidden>
<span class="reboot-restore-banner-icon" aria-hidden="true">↺</span>
<span class="reboot-restore-banner-text" id="rebootRestoreBannerText"></span>
<span class="reboot-restore-banner-detail" id="rebootRestoreBannerDetail"></span>
<span class="reboot-restore-banner-note">Conversations return; terminal history does not.</span>
<button
class="reboot-restore-banner-accept"
id="rebootRestoreBannerAccept"
onclick="app.restoreRebootSessions()"
>
Restore
</button>
<button class="reboot-restore-banner-dismiss" onclick="app.dismissRebootRestore()">Dismiss</button>
</div>
<!-- Timer Banner (shown when timed run is active) -->
<div class="timer-banner" id="timerBanner" style="display: none;">
<div class="timer-content">
@@ -3535,6 +3553,7 @@
<script defer src="readmymind-ui.js"></script>
<script defer src="ultracode-panel.js"></script>
<script defer src="approvals-ui.js"></script>
<script defer src="reboot-restore-ui.js"></script>
<script defer src="admin-ui.js"></script>
<script defer src="session-ui.js"></script>
<script defer src="webview-tabs.js"></script>
+95
View File
@@ -0,0 +1,95 @@
/**
* @fileoverview Reboot-restore banner: offer back the sessions a host reboot destroyed.
*
* A host reboot takes the tmux server down with it, so every session's pane dies
* and the board comes up empty. The server works out what was running from the
* records it still holds at boot, and this banner asks the user whether to
* rebuild them. Nothing is created until they click, because the server's
* reboot guess is a heuristic and a wrong automatic restore would spawn CLI
* processes nobody asked for.
*
* Seeded once from `GET /api/reboot-restore` on init. Restore posts to
* `POST /api/reboot-restore/restore`, Dismiss posts to
* `POST /api/reboot-restore/dismiss`, and either way the banner goes away. The
* restored sessions arrive as ordinary `session:created` events, so no extra
* rendering is needed here.
*
* The banner says that terminal history did not survive, because a restored
* session is a new pane: the conversation continues and the scrollback does not.
* Saying so is what keeps an empty pane from reading as a broken restore.
* Backend: src/web/reboot-restore-registry.ts, src/web/routes/reboot-restore-routes.ts.
*
* @mixin Extends CodemanApp.prototype via Object.assign
* @dependency app.js (CodemanApp class, showToast)
* @dependency api-client.js at runtime (this._apiJson / this._apiPost)
* @loadorder 11.7 of 17, after approvals-ui.js
*/
Object.assign(CodemanApp.prototype, {
/** Ask the server whether a reboot left anything on offer, and show the banner if so. */
async initRebootRestoreBanner() {
const data = await this._apiJson('/api/reboot-restore');
const sessions = data?.sessions ?? [];
if (sessions.length === 0) return;
this._rebootRestoreSessions = sessions;
this.renderRebootRestoreBanner();
},
renderRebootRestoreBanner() {
const banner = this.$('rebootRestoreBanner');
if (!banner) return;
const sessions = this._rebootRestoreSessions ?? [];
if (sessions.length === 0) {
banner.hidden = true;
return;
}
const count = sessions.length;
const text = this.$('rebootRestoreBannerText');
if (text) {
const noun = count === 1 ? 'session' : 'sessions';
text.textContent = `Restore ${count} ${noun} from before the reboot`;
}
const detail = this.$('rebootRestoreBannerDetail');
if (detail) {
// Names, so the user can tell what they are about to relaunch.
const names = sessions
.map((s) => s.name || s.workingDir?.split('/').pop() || s.id.slice(0, 8))
.slice(0, 4)
.join(', ');
detail.textContent = count > 4 ? `${names}, …` : names;
detail.title = sessions.map((s) => `${s.name || s.id}\n${s.workingDir}`).join('\n\n');
}
banner.hidden = false;
},
/** Rebuild everything on offer. The panes are new, so scrollback does not come back. */
async restoreRebootSessions() {
const button = this.$('rebootRestoreBannerAccept');
if (button) button.disabled = true;
const res = await this._apiPost('/api/reboot-restore/restore', {});
const body = res && res.ok ? await res.json().catch(() => null) : null;
if (!body) {
if (button) button.disabled = false;
this.showToast?.('Could not restore the sessions', 'error');
return;
}
const restored = body.restored?.length ?? 0;
const skipped = body.skipped?.length ?? 0;
this._rebootRestoreSessions = [];
this.renderRebootRestoreBanner();
if (restored > 0) {
const noun = restored === 1 ? 'conversation' : 'conversations';
this.showToast?.(`Restored ${restored} ${noun}. Terminal history did not survive the reboot.`, 'success');
}
if (skipped > 0) {
this.showToast?.(`${skipped} could not be restored (workspace gone, or already open)`, 'warning');
}
},
/** Drop the offer. The Resume list still reaches every one of these conversations. */
async dismissRebootRestore() {
this._rebootRestoreSessions = [];
this.renderRebootRestoreBanner();
await this._apiPost('/api/reboot-restore/dismiss', {});
},
});
+75
View File
@@ -15243,6 +15243,81 @@ html[data-skin="daylight-blue"] .welcome-btn-tunnel.active:hover {
skin, including the light ones. Visibility is driven by the `hidden`
attribute, so the display rules need !important to lose to it. */
/* Reboot-restore offer. Amber rather than red: nothing is wrong, the board is
asking a question, and the user can ignore it. See reboot-restore-ui.js. */
.reboot-restore-banner {
display: flex;
align-items: center;
gap: 0.6rem;
padding: 0.45rem 1rem;
background: linear-gradient(90deg, #b45309, #92400e);
border-bottom: 1px solid rgba(0, 0, 0, 0.35);
color: #fff;
font-size: 0.78rem;
font-weight: 600;
letter-spacing: 0.01em;
flex-shrink: 0;
z-index: 1250;
}
.reboot-restore-banner[hidden] {
display: none !important;
}
.reboot-restore-banner-icon {
flex-shrink: 0;
font-size: 0.95rem;
line-height: 1;
}
.reboot-restore-banner-text {
white-space: nowrap;
}
.reboot-restore-banner-detail {
color: rgba(255, 255, 255, 0.8);
font-weight: 500;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.reboot-restore-banner-note {
color: rgba(255, 255, 255, 0.75);
font-weight: 500;
white-space: nowrap;
margin-left: auto;
}
.reboot-restore-banner-accept,
.reboot-restore-banner-dismiss {
flex-shrink: 0;
padding: 0.2rem 0.6rem;
border-radius: 5px;
border: 1px solid rgba(255, 255, 255, 0.55);
background: rgba(255, 255, 255, 0.12);
color: #fff;
font-size: 0.72rem;
font-weight: 600;
cursor: pointer;
}
.reboot-restore-banner-accept:hover,
.reboot-restore-banner-dismiss:hover {
background: rgba(255, 255, 255, 0.24);
}
.reboot-restore-banner-accept:disabled {
opacity: 0.6;
cursor: default;
}
.reboot-restore-banner-dismiss {
border-color: rgba(255, 255, 255, 0.3);
background: transparent;
font-weight: 500;
}
.offline-banner {
display: flex;
align-items: center;
+142
View File
@@ -0,0 +1,142 @@
/**
* @fileoverview The pending restore plan: what a host reboot destroyed, waiting on a click.
*
* The boot pass builds this plan inside `restoreMuxSessions()`, in the window
* where reconciliation has reported the dead sessions and `cleanupStaleSessions()`
* has not pruned their records yet. The board then offers "restore N sessions
* from before the reboot", and `web/routes/reboot-restore-routes` spends the plan
* when the user clicks.
*
* Invariants:
* - Entries are in-memory only. A server restart drops the plan, and nothing
* re-builds it, because the records it was built from are pruned by then.
* That costs the convenience this feature adds and never the conversation:
* the conversation IS the transcript under `~/.claude/projects`, which
* `services/unified-session-service.ts` reads for the Welcome screen's Resume
* list and the Session Manager, and `resumeHistorySession()` in
* `web/public/terminal-ui.js` resumes from a row there with no persisted
* session record involved. A dropped plan therefore returns the user to
* resuming by hand, one at a time, which is where they are without this
* feature. What the plan held that a transcript does not is the owner, the
* name, the env overrides, the effort and the lineage.
* - Module-level singleton in the style of `web/approval-inbox.ts`: no `Session`
* import and no IO, which keeps it unit-testable and cycle-free.
* - Spending is take-then-build: `take()` removes entries synchronously, before
* the route's first `await`, so a double-click or two devices cannot both
* reach the same entry and put two panes on one conversation.
* - One restore runs at a time. `beginSpending()` single-flights the route, so
* two concurrent clicks cannot interleave pane creation.
*
* @dependencies reboot-restore (RebootRestoreEntry)
* @consumedby web/server (plan build at boot), web/routes/reboot-restore-routes
*
* @module web/reboot-restore-registry
*/
import type { RebootRestoreEntry } from '../reboot-restore.js';
/**
* A plan older than this is dropped on read. A machine that rebooted yesterday
* has moved on, and an offer nobody took by then is noise rather than a rescue.
*/
const PLAN_TTL_MS = 24 * 60 * 60 * 1000;
export class RebootRestoreRegistry {
/** Keyed by session id, in the order the boot pass found them. */
private entries = new Map<string, RebootRestoreEntry>();
/** When the boot pass built the plan, in ms since the epoch. */
private builtAt = 0;
/** True while a restore route call is between its take and its last pane. */
private spending = false;
/** Replace the plan with what the boot pass found. An empty list clears it. */
set(entries: readonly RebootRestoreEntry[]): void {
this.entries = new Map(entries.map((entry) => [entry.sessionId, entry]));
this.builtAt = entries.length > 0 ? Date.now() : 0;
}
/**
* The entries a viewer may see, newest plan first-come order preserved.
*
* @param canAccess Ownership predicate, so a user sees their own entries and
* an admin sees all. Applied here rather than in the route so the count the
* banner shows and the entries a click spends come from one filter.
*/
list(canAccess: (owner: string | undefined) => boolean): RebootRestoreEntry[] {
this.dropIfExpired();
return [...this.entries.values()].filter((entry) => canAccess(entry.owner));
}
/**
* Remove and return the entries a click is about to spend.
*
* Synchronous and total: an entry leaves the plan here, before any pane is
* created, so a second click finds nothing to spend. Entries a caller may not
* access are left in place, and unknown ids are ignored.
*
* @param sessionIds The ids to spend, or undefined for every visible entry.
*/
take(canAccess: (owner: string | undefined) => boolean, sessionIds?: readonly string[]): RebootRestoreEntry[] {
this.dropIfExpired();
const wanted = sessionIds ? new Set(sessionIds) : undefined;
const taken: RebootRestoreEntry[] = [];
for (const entry of [...this.entries.values()]) {
if (wanted && !wanted.has(entry.sessionId)) continue;
if (!canAccess(entry.owner)) continue;
this.entries.delete(entry.sessionId);
taken.push(entry);
}
return taken;
}
/**
* Put entries back after a rebuild never got as far as creating a pane.
*
* Used for the click-time rejections, so a conversation the user resumed by
* hand meanwhile does not silently vanish from the banner while a workspace
* that came back stays offered.
*/
restore(entries: readonly RebootRestoreEntry[]): void {
for (const entry of entries) this.entries.set(entry.sessionId, entry);
if (entries.length > 0 && this.builtAt === 0) this.builtAt = Date.now();
}
/** Drop the entries a viewer can see. Returns how many went. */
clear(canAccess: (owner: string | undefined) => boolean): number {
const removable = [...this.entries.values()].filter((entry) => canAccess(entry.owner));
for (const entry of removable) this.entries.delete(entry.sessionId);
if (this.entries.size === 0) this.builtAt = 0;
return removable.length;
}
/**
* Claim the right to run a restore, or report that one is already running.
* Callers that get `true` must call `endSpending()` in a `finally`.
*/
beginSpending(): boolean {
if (this.spending) return false;
this.spending = true;
return true;
}
endSpending(): void {
this.spending = false;
}
/** Test hook: forget everything, including the single-flight claim. */
reset(): void {
this.entries.clear();
this.builtAt = 0;
this.spending = false;
}
private dropIfExpired(): void {
if (this.builtAt > 0 && Date.now() - this.builtAt > PLAN_TTL_MS) {
this.entries.clear();
this.builtAt = 0;
}
}
}
/** Process-wide singleton, mirroring `approvalInbox`. */
export const rebootRestoreRegistry = new RebootRestoreRegistry();
+1
View File
@@ -11,6 +11,7 @@ export { registerCronRoutes } from './cron-routes.js';
export { registerSystemRoutes } from './system-routes.js';
export { registerHookEventRoutes } from './hook-event-routes.js';
export { registerApprovalRoutes } from './approval-routes.js';
export { registerRebootRestoreRoutes } from './reboot-restore-routes.js';
export { registerReadMyMindRoutes } from './readmymind-routes.js';
export { registerStatusTelemetryRoutes } from './status-telemetry-routes.js';
export { registerCaseRoutes } from './case-routes.js';
+194
View File
@@ -0,0 +1,194 @@
/**
* @fileoverview Reboot-restore routes: offer back the sessions a host reboot destroyed.
*
* The boot pass leaves a plan in `web/reboot-restore-registry` when the machine
* plausibly rebooted. The board reads it, shows a banner, and the user decides:
* - `GET /api/reboot-restore`: what is on offer, ownership-scoped
* - `POST /api/reboot-restore/restore`: rebuild some or all of it
* - `POST /api/reboot-restore/dismiss`: drop the offer
*
* A click, not the heuristic, is what creates panes. The heuristic only decides
* whether the banner appears, so a wrong yes costs a line of text the user
* dismisses rather than N CLI processes nobody asked for.
*
* Rebuilding is take-then-build: entries leave the plan synchronously at the top
* of the route, before the first `await`, and the whole route is single-flighted,
* so a double-click or two devices cannot put two panes on one conversation.
* Three things are re-checked at click time rather than trusted from boot: the
* owner's privilege grant, the workspace still being on disk, and the
* conversation not already being live because the user resumed it by hand.
*
* A rebuilt session comes back attached, idle and disarmed. Respawn controllers
* and Ralph loops are deliberately not re-armed, and its terminal scrollback is
* gone, because the pane is new. The banner says so.
*/
import { FastifyInstance } from 'fastify';
import { existsSync } from 'node:fs';
import { ApiErrorCode, createErrorResponse, getErrorMessage } from '../../types.js';
import { RebootRestoreRequestSchema } from '../schemas.js';
import { parseBody, getAuthUser, canAccessOwned } from '../route-helpers.js';
import { rebootRestoreRegistry } from '../reboot-restore-registry.js';
import { rejectAlreadyLive, type RebootRestoreEntry, type RebootRestoreRejection } from '../../reboot-restore.js';
import { clampEnvOverridesForOwner } from '../../session-env-clamp.js';
import { Session } from '../../session.js';
import { resolveClaudeModeForUsername } from '../../user-store.js';
import { getCli } from '../../config/cli-registry/registry.js';
import { applyWorkspaceHooks } from '../../hooks-config.js';
import { getLifecycleLog } from '../../session-lifecycle-log.js';
import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js';
import { SseEvent } from '../sse-events.js';
import type { SessionAttachmentHistoryItem } from '../../types.js';
import type { SessionPort, EventPort, ConfigPort, InfraPort } from '../ports/index.js';
type RebootRestoreCtx = SessionPort & EventPort & ConfigPort & InfraPort;
/** The banner's view of one restorable session. The record itself never leaves the server. */
function toBannerItem(entry: RebootRestoreEntry) {
return {
id: entry.sessionId,
name: entry.name,
workingDir: entry.workingDir,
mode: entry.mode,
owner: entry.owner,
};
}
export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRestoreCtx): void {
const accessorFor = (req: Parameters<typeof getAuthUser>[0]) => {
const user = getAuthUser(req);
return (owner: string | undefined) => canAccessOwned(user, owner);
};
// ========== What is on offer ==========
app.get('/api/reboot-restore', async (req) => {
const entries = rebootRestoreRegistry.list(accessorFor(req));
return {
sessions: entries.map(toBannerItem),
// Said plainly here so the banner never implies a full restore: the pane is
// new, so the conversation continues and the terminal history does not.
scrollbackRestored: false,
};
});
// ========== Spend it ==========
app.post('/api/reboot-restore/restore', async (req, reply) => {
const body = parseBody(RebootRestoreRequestSchema, req.body, 'Invalid reboot restore request');
const canAccess = accessorFor(req);
// Take BEFORE the first await: a second click must find nothing to spend.
if (!rebootRestoreRegistry.beginSpending()) {
return reply.code(409).send(createErrorResponse(ApiErrorCode.CONFLICT, 'A reboot restore is already running'));
}
const taken = rebootRestoreRegistry.take(canAccess, body.sessionIds);
try {
if (taken.length === 0) return { restored: [], skipped: [] };
// The plan was built at boot and the board has moved on since. A conversation
// the user resumed by hand from the Resume list is already on screen, and a
// second pane on it would fight the first for the same transcript.
const liveSessionIds = new Set(ctx.sessions.keys());
const liveConversationIds = new Set(
[...ctx.sessions.values()].map((session) => session.claudeSessionId).filter((id): id is string => !!id)
);
const { restore, skipped } = rejectAlreadyLive(taken, liveSessionIds, liveConversationIds);
// An entry nothing rebuilt stays on offer rather than disappearing silently.
rebootRestoreRegistry.restore(skipped.map((s) => taken.find((e) => e.sessionId === s.sessionId)!));
const restored: ReturnType<typeof toBannerItem>[] = [];
const failures: RebootRestoreRejection[] = [...skipped];
const workspaceHooksEnabled = await ctx.getWorkspaceHooksEnabled();
for (const entry of restore) {
// A repo can be deleted between the boot that planned this and the click.
if (!existsSync(entry.workingDir)) {
failures.push({ sessionId: entry.sessionId, reason: 'workspace-missing' });
continue;
}
try {
const saved = entry.state;
const claudeModeConfig = await ctx.getClaudeModeConfig();
const session = new Session({
// The old id is reused on purpose: a pinned record, subagent parents,
// window states and the lifecycle log all key off it, and the unpinned
// record is gone, so there is nothing to collide with.
id: saved.id,
workingDir: saved.workingDir,
mode: saved.mode,
name: saved.name,
createdAt: saved.createdAt,
mux: ctx.mux,
useMux: true,
// No `muxSession`: the reboot took the pane with it, so `startInteractive()`
// takes its create branch and makes a fresh one.
claudeMode: await resolveClaudeModeForUsername(claudeModeConfig.claudeMode, saved.owner),
allowedTools: claudeModeConfig.allowedTools,
resumeSessionId: entry.resumeConversationId,
// Re-resolved against the owner's CURRENT grant, never replayed from the
// record: a grant held when the record was written may be gone now.
envOverrides: await clampEnvOverridesForOwner(
saved.owner,
(saved as { __envOverrides?: Record<string, string> }).__envOverrides
),
effort: saved.effort,
attachmentHistory:
(saved as { __attachmentHistory?: SessionAttachmentHistoryItem[] }).__attachmentHistory ??
saved.attachmentHistory,
lastSubmitAt: saved.lastSubmitAt,
claudeSessionChain: saved.claudeSessionChain,
lastActivityAt: saved.lastActivityAt,
owner: saved.owner,
parentSessionId: saved.parentSessionId,
});
await ctx.addSession(session);
ctx.persistSessionState(session);
await ctx.setupSessionListeners(session);
await session.startInteractive();
// A session without its workspace hooks goes silently blind: no stop or
// idle events for respawn, no Approvals Inbox item, no red tab on a
// blocking dialog. The boot-time sweep finished hours ago, so the click
// path installs them itself. `hooks: 'always'` is the capability that says
// this CLI installs Codeman's hooks into the workspace.
if (workspaceHooksEnabled && getCli(session.mode)?.capabilities.hooks === 'always') {
await applyWorkspaceHooks(session.workingDir, true).catch((err: unknown) =>
console.warn(`[reboot-restore] hook install failed for ${session.workingDir}: ${getErrorMessage(err)}`)
);
}
getLifecycleLog().log({ event: 'recovered', sessionId: session.id, name: session.name });
// Every other open tab and phone needs this; the clicking tab already has
// the response, and the client's handler is an idempotent upsert.
ctx.broadcast(SseEvent.SessionCreated, ctx.getSessionStateWithRespawn(session));
restored.push(toBannerItem(entry));
} catch (err) {
// One workspace that has gone missing must not stop the rest of the pass.
console.error(`[reboot-restore] failed to rebuild ${entry.sessionId}:`, err);
failures.push({ sessionId: entry.sessionId, reason: 'workspace-missing' });
}
}
if (restored.length > 0) {
// A reboot leaves recovery with nothing alive to find, so its own block never
// started the stats collector. This clears and re-arms its interval, so it is
// safe to call whether or not the collector is already running.
ctx.mux.startStatsCollection(STATS_COLLECTION_INTERVAL_MS);
}
return { restored, skipped: failures };
} finally {
rebootRestoreRegistry.endSpending();
}
});
// ========== Drop it ==========
app.post('/api/reboot-restore/dismiss', async (req) => {
const dismissed = rebootRestoreRegistry.clear(accessorFor(req));
return { dismissed };
});
}
+1 -66
View File
@@ -89,6 +89,7 @@ import {
} from '../route-helpers.js';
import { buildAgentCaseMarker, writeAgentCaseMarker } from '../../agent-case-marker.js';
import { canUsernameRunPrivilegedCommands, resolveClaudeModeForUsername } from '../../user-store.js';
import { clampEnvOverridesForOwner } from '../../session-env-clamp.js';
import { enabledClis, getCli } from '../../config/cli-registry/registry.js';
import { resolveCliLaunchError } from '../../utils/cli-launcher.js';
import { legacyConfigForMode } from '../../session-cli-registry-bridge.js';
@@ -442,72 +443,6 @@ export async function _clampExternalCliBypassForOwner(
};
}
/**
* Env-var keys a non-granted owner must not be able to set, because each one
* hands back privilege the config clamp above just removed, or redirects a
* credential-resolution endpoint.
*
* The DeepSeek three are reachable because `DSH_*` and `DEEPSEEK_*` are
* allowlisted `envOverrides` prefixes (schemas.ts) — which they have to be, since
* that is also how a user configures the harness's non-privileged knobs.
*
* - `DSH_PERMISSION_MODE` IS the harness's permission switch. Every other CLI's
* bypass is a command-line FLAG, reachable only through the per-CLI config the
* clamp already owns; this one is an env var, so the config clamp alone is
* half a gate.
* - `DSH_HOME` points the launcher at a profile tree, and a profile's plugin code
* executes at BOOT, before any approval row can apply. A user who can write a
* workspace can put a profile in it, so this is the wider of the two.
* - `DEEPSEEK_BASE_URL` aims the provider endpoint, and `_configureCliEnv()`
* forwards the SERVER's own `DEEPSEEK_API_KEY` into every dsh pane before
* `applyEnvOverrides()` runs — so a non-granted owner who could set the base
* URL would have the operator's API key sent as a bearer credential to a host
* of their choosing. (`DEEPSEEK_API_KEY` itself stays overridable: supplying
* your OWN key removes privilege rather than granting it.)
* - `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN` are where omp resolves
* credentials from — the same shape as `DEEPSEEK_BASE_URL` above, reachable
* because `OMP_*` is an allowlisted prefix. Unlike DeepSeek, Codeman does not
* forward any operator-held key into an omp pane today (omp's provider
* credentials live in `~/.omp` config files, not env vars), so there is no
* known concrete exfiltration path yet — clamped defensively anyway, since a
* non-granted owner redirecting where a shared multi-tenant deployment
* resolves auth from is not something to allow silently (found in
* Ark0N/Codeman#353 review; omp's own knobs are otherwise mostly `PI_*`,
* already allowlisted for pi and not addressed here — see resolveOmpHome()).
*/
function ownerClampedEnvKeys(): string[] {
return enabledClis().flatMap((entry) => entry.capabilities.privilegedEnvKeys);
}
/**
* Env-var half of the multi-user bypass clamp.
*
* `clampExternalCliBypassForOwner()` clamps the per-CLI CONFIG, and for every CLI
* but DeepSeek that is the whole story. Here it is not: `applyEnvOverrides()` runs
* AFTER `_configureCliEnv()` in tmux-manager, so an override sent on the SAME
* request lands last and wins, and a non-granted owner could restore
* `danger-full-access` on the very request the config clamp downgraded.
*
* Keys are DROPPED rather than rewritten: dropping falls through to what
* `_configureCliEnv()` exports, which is the clamped config and the server's own
* `DSH_HOME`, i.e. exactly the intended state. No-op in single-user mode and for a
* granted owner, like every other clamp here
* (`canUsernameRunPrivilegedCommands()` returns true when `!isMultiUserMode()`),
* and it returns the caller's own object untouched when there is nothing to strip.
*/
async function clampEnvOverridesForOwner(
owner: string | undefined,
envOverrides: Record<string, string> | undefined
): Promise<Record<string, string> | undefined> {
if (!envOverrides) return envOverrides;
const keys = ownerClampedEnvKeys();
if (!keys.some((key) => key in envOverrides)) return envOverrides;
if (await canUsernameRunPrivilegedCommands(owner)) return envOverrides;
const clamped = { ...envOverrides };
for (const key of keys) delete clamped[key];
return clamped;
}
/** Test hook: the env-var half of the same multi-user safety gate. */
export const _clampEnvOverridesForOwner = clampEnvOverridesForOwner;
+14
View File
@@ -1161,6 +1161,20 @@ const NotificationEventSchema = z
})
.optional();
/**
* Body of `POST /api/reboot-restore/restore`.
*
* `sessionIds` restores a subset, and omitting it restores everything the caller
* can see. The ids are session ids from `GET /api/reboot-restore`, and an id the
* caller does not own is ignored rather than refused, matching how the session
* list scopes rather than 403s.
*/
export const RebootRestoreRequestSchema = z
.object({
sessionIds: z.array(z.string().max(128)).max(200).optional(),
})
.strict();
export const SettingsUpdateSchema = z
.object({
// User-facing product branding. This changes browser/UI copy only; package,
+56 -1
View File
@@ -39,7 +39,9 @@ import { fileURLToPath } from 'node:url';
import { existsSync, mkdirSync, readFileSync, chmodSync, rmSync, statSync } from 'node:fs';
import fs from 'node:fs/promises';
import { execSync } from 'node:child_process';
import { hostname as getHostname } from 'node:os';
import { hostname as getHostname, uptime as osUptime } from 'node:os';
import { looksLikeHostReboot, newestPersistedActivity, planRebootRestore } from '../reboot-restore.js';
import { rebootRestoreRegistry } from './reboot-restore-registry.js';
import { dataPath, getDataDir, CODEMAN_INSTANCE } from '../config/instance.js';
import { normalizeBasePath, stripBasePath, joinBasePath } from '../config/base-path.js';
import { GLYPH, palette } from '../cli-style.js';
@@ -171,6 +173,7 @@ import {
registerScheduledRoutes,
registerHookEventRoutes,
registerApprovalRoutes,
registerRebootRestoreRoutes,
registerReadMyMindRoutes,
registerStatusTelemetryRoutes,
registerSystemRoutes,
@@ -1060,6 +1063,7 @@ export class WebServer extends EventEmitter {
registerScheduledRoutes(this.app, ctx);
registerHookEventRoutes(this.app, ctx);
registerApprovalRoutes(this.app, ctx);
registerRebootRestoreRoutes(this.app, ctx);
registerReadMyMindRoutes(this.app, ctx);
registerStatusTelemetryRoutes(this.app, ctx);
registerSystemRoutes(this.app, ctx);
@@ -2857,6 +2861,52 @@ export class WebServer extends EventEmitter {
return false;
}
/**
* Work out what a host reboot destroyed, and leave it on offer for the board.
*
* Runs inside `restoreMuxSessions()`, in the window after `reconcileSessions()`
* has reported the dead sessions and before `finalizeRestoredState()` prunes
* their records, so `state.json` is still the full picture here. That window is
* the only place the plan can be built, which is why the boot pass builds it
* even though nothing is rebuilt until a user clicks.
*
* Nothing is created here. The plan goes to `rebootRestoreRegistry`, the board
* offers it as a banner, and `web/routes/reboot-restore-routes` rebuilds what
* the user asks for. A wrong reboot guess therefore costs a line of text the
* user dismisses, not N CLI processes nobody asked for.
*
* @returns how many sessions are on offer.
*/
private planRebootRestoreOffer(dead: string[], livePaneCount: number): number {
if (dead.length === 0) return 0;
const persisted = this.store.getSessions();
if (
!looksLikeHostReboot({
livePaneCount,
deadSessionCount: dead.length,
uptimeSeconds: osUptime(),
newestPersistedActivityAt: newestPersistedActivity(persisted),
now: Date.now(),
})
) {
return 0;
}
const { restore, skipped } = planRebootRestore(dead, persisted, (workingDir) => existsSync(workingDir));
if (skipped.length > 0) {
console.log(`[Server] Reboot restore is passing over ${skipped.length} dead session(s):`);
for (const rejection of skipped) {
console.log(`[Server] ${rejection.sessionId}: ${rejection.reason}`);
}
}
rebootRestoreRegistry.set(restore);
if (restore.length > 0) {
console.log(`[Server] Host reboot detected; offering ${restore.length} session(s) for restore`);
}
return restore.length;
}
private async restoreMuxSessions(): Promise<boolean> {
try {
// Reconcile mux sessions to find which ones are still alive (also discovers unknown ones)
@@ -2866,6 +2916,11 @@ export class WebServer extends EventEmitter {
console.log(`[Server] Discovered ${discovered.length} unknown mux session(s)`);
}
// Build the reboot-restore offer HERE: `dead` is only known after
// reconciliation, and the records it reads are pruned by
// `cleanupStaleSessions()` as soon as `finalizeRestoredState()` runs.
this.planRebootRestoreOffer(dead, alive.length);
if (alive.length > 0 || discovered.length > 0) {
console.log(`[Server] Found ${alive.length + discovered.length} alive mux session(s) from previous run`);
+367
View File
@@ -0,0 +1,367 @@
/**
* @fileoverview The decision half of reboot restore, and proof that the existing
* recovery construction path can CREATE a resumed pane.
*
* Three things are under test. `src/reboot-restore.ts` decides whether the
* machine rebooted and which dead sessions may be offered back. The plan
* registry in `src/web/reboot-restore-registry.ts` holds that offer between the
* boot that builds it and the click that spends it. The third is the claim the
* whole feature rests on: a `Session` built the way `restoreMuxSessions()`
* already builds one, but given no `muxSession` and a `resumeSessionId`, creates
* a fresh pane that resumes the old conversation. If that holds, the restore
* needs no new session-creation service.
*
* `reconcileSessions()` reports every session ALIVE under vitest, so the
* server's own boot pass cannot be reached from here. The decision logic is
* therefore driven directly, and the construction claim is driven through a real
* `Session` against the in-memory tmux layer vitest substitutes.
*/
import { mkdirSync, rmSync } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
import { afterEach, describe, expect, it } from 'vitest';
import { Session } from '../src/session.js';
import { TmuxManager } from '../src/tmux-manager.js';
import type { SessionState } from '../src/types.js';
import {
looksLikeHostReboot,
newestPersistedActivity,
planRebootRestore,
rejectAlreadyLive,
resolveResumeConversationId,
type RebootRestoreEntry,
} from '../src/reboot-restore.js';
import { RebootRestoreRegistry } from '../src/web/reboot-restore-registry.js';
const HOUR = 60 * 60 * 1000;
const NOW = 1_760_000_000_000;
function persistedSession(overrides: Partial<SessionState> & { id: string }): SessionState {
return {
pid: 99999,
status: 'idle',
workingDir: '/tmp/spike',
currentTaskId: null,
createdAt: NOW - 4 * HOUR,
lastActivityAt: NOW - 2 * HOUR,
mode: 'claude',
...overrides,
} as SessionState;
}
describe('reboot detection', () => {
const base = {
livePaneCount: 0,
deadSessionCount: 2,
// The host came up 10 minutes ago, well after the sessions were last active.
uptimeSeconds: 600,
newestPersistedActivityAt: NOW - 2 * HOUR,
now: NOW,
};
it('calls it a reboot when the socket is empty and the host booted after the last activity', () => {
expect(looksLikeHostReboot(base)).toBe(true);
});
it('refuses when some panes survived, which is an ordinary server restart', () => {
expect(looksLikeHostReboot({ ...base, livePaneCount: 3 })).toBe(false);
});
it('refuses on a long-uptime host, where someone wiped the tmux socket by hand', () => {
// Up for 30 days: the sessions were active long AFTER this boot, so the panes
// went away for some reason other than the machine restarting.
expect(looksLikeHostReboot({ ...base, uptimeSeconds: 30 * 24 * 60 * 60 })).toBe(false);
});
it('refuses when nothing died', () => {
expect(looksLikeHostReboot({ ...base, deadSessionCount: 0 })).toBe(false);
});
it('reads the newest activity stamp across the persisted records', () => {
const persisted = {
a: persistedSession({ id: 'a', lastActivityAt: NOW - 5 * HOUR }),
b: persistedSession({ id: 'b', lastActivityAt: NOW - 1 * HOUR }),
};
expect(newestPersistedActivity(persisted)).toBe(NOW - 1 * HOUR);
});
});
describe('which dead sessions may be rebuilt', () => {
it('rebuilds a session that was simply running when the power went out', () => {
const persisted = { live: persistedSession({ id: 'live', status: 'busy' }) };
const plan = planRebootRestore(['live'], persisted, () => true);
expect(plan.restore.map((s) => s.sessionId)).toEqual(['live']);
});
it('never revives a session the user killed while pinned (COD-142 demotes it to stopped)', () => {
const persisted = { killed: persistedSession({ id: 'killed', status: 'stopped', pinned: true }) };
const plan = planRebootRestore(['killed'], persisted, () => true);
expect(plan.restore).toEqual([]);
expect(plan.skipped).toEqual([{ sessionId: 'killed', reason: 'intentionally-ended' }]);
});
it('never revives a session whose record an unpinned kill already deleted', () => {
const plan = planRebootRestore(['gone'], {}, () => true);
expect(plan.restore).toEqual([]);
expect(plan.skipped).toEqual([{ sessionId: 'gone', reason: 'no-persisted-record' }]);
});
it('never revives a pane whose PTY-exit breaker had tripped', () => {
const persisted = { crashy: persistedSession({ id: 'crashy', respawnBlocked: true }) };
expect(planRebootRestore(['crashy'], persisted, () => true).skipped[0].reason).toBe('respawn-blocked');
});
it('leaves remote sessions to the COD-108 reconnect watcher', () => {
const persisted = {
r: persistedSession({
id: 'r',
remote: { hostId: 'h', host: 'example.test', username: 'u', sessionName: 'n', owned: true },
} as Partial<SessionState> & { id: string }),
};
expect(planRebootRestore(['r'], persisted, () => true).skipped[0].reason).toBe('remote-or-docker');
});
it('leaves docker sessions alone, since the container may not be up', () => {
const persisted = {
d: persistedSession({ id: 'd', docker: { containerId: 'abc', caseId: 'c' } } as Partial<SessionState> & {
id: string;
}),
};
expect(planRebootRestore(['d'], persisted, () => true).skipped[0].reason).toBe('remote-or-docker');
});
it('skips a CLI whose history the claude transcript reader does not understand', () => {
const persisted = { c: persistedSession({ id: 'c', mode: 'codex' }) };
expect(planRebootRestore(['c'], persisted, () => true).skipped[0].reason).toBe('unsupported-mode');
});
});
describe('a workspace that is no longer on disk', () => {
it('is kept out of the offer, so a click cannot scaffold a deleted repo', () => {
const persisted = { gone: persistedSession({ id: 'gone', workingDir: '/tmp/deleted-repo' }) };
const plan = planRebootRestore(['gone'], persisted, () => false);
expect(plan.restore).toEqual([]);
expect(plan.skipped).toEqual([{ sessionId: 'gone', reason: 'workspace-missing' }]);
});
it('is judged per session, not for the batch', () => {
const persisted = {
kept: persistedSession({ id: 'kept', workingDir: '/tmp/still-here' }),
gone: persistedSession({ id: 'gone', workingDir: '/tmp/deleted-repo' }),
};
const plan = planRebootRestore(['kept', 'gone'], persisted, (dir) => dir === '/tmp/still-here');
expect(plan.restore.map((entry) => entry.sessionId)).toEqual(['kept']);
expect(plan.skipped.map((s) => s.reason)).toEqual(['workspace-missing']);
});
});
describe('a conversation that came back on its own before the click', () => {
const entry: RebootRestoreEntry = {
sessionId: 'abc',
workingDir: '/tmp/spike',
mode: 'claude',
resumeConversationId: 'conv-1',
state: persistedSession({ id: 'abc' }),
};
it('is skipped when the user resumed it by hand from the Resume list', () => {
// Same conversation, different session id: the Resume list creates a NEW id.
const result = rejectAlreadyLive([entry], new Set(['other']), new Set(['conv-1']));
expect(result.restore).toEqual([]);
expect(result.skipped).toEqual([{ sessionId: 'abc', reason: 'already-live' }]);
});
it('is skipped when a session with that id is already on the board', () => {
const result = rejectAlreadyLive([entry], new Set(['abc']), new Set());
expect(result.skipped).toEqual([{ sessionId: 'abc', reason: 'already-live' }]);
});
it('is rebuilt when neither its id nor its conversation is live', () => {
const result = rejectAlreadyLive([entry], new Set(['other']), new Set(['conv-other']));
expect(result.restore.map((e) => e.sessionId)).toEqual(['abc']);
expect(result.skipped).toEqual([]);
});
});
describe('the plan the banner spends', () => {
const all = () => true;
const entryFor = (sessionId: string, owner?: string): RebootRestoreEntry => ({
sessionId,
owner,
workingDir: '/tmp/spike',
mode: 'claude',
resumeConversationId: `conv-${sessionId}`,
state: persistedSession({ id: sessionId, owner }),
});
it('hands an entry to the first caller and nothing to the second', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a'), entryFor('b')]);
expect(registry.take(all).map((e) => e.sessionId)).toEqual(['a', 'b']);
// The double-click: two panes on one conversation is what this prevents.
expect(registry.take(all)).toEqual([]);
});
it('spends only the ids a caller asked for', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a'), entryFor('b')]);
expect(registry.take(all, ['b']).map((e) => e.sessionId)).toEqual(['b']);
expect(registry.list(all).map((e) => e.sessionId)).toEqual(['a']);
});
it("shows a user their own sessions and leaves another owner's alone", () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('mine', 'alice'), entryFor('theirs', 'bob')]);
const asAlice = (owner: string | undefined) => owner === 'alice';
expect(registry.list(asAlice).map((e) => e.sessionId)).toEqual(['mine']);
expect(registry.take(asAlice).map((e) => e.sessionId)).toEqual(['mine']);
// Bob's entry is still on offer for Bob.
expect(registry.list(() => true).map((e) => e.sessionId)).toEqual(['theirs']);
});
it('puts back an entry that no pane was created for', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a')]);
const taken = registry.take(all);
registry.restore(taken);
expect(registry.list(all).map((e) => e.sessionId)).toEqual(['a']);
});
it('runs one restore at a time', () => {
const registry = new RebootRestoreRegistry();
expect(registry.beginSpending()).toBe(true);
expect(registry.beginSpending()).toBe(false);
registry.endSpending();
expect(registry.beginSpending()).toBe(true);
});
it('drops what a dismiss cleared', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a'), entryFor('b')]);
expect(registry.clear(all)).toBe(2);
expect(registry.list(all)).toEqual([]);
});
it('forgets a plan nobody took for a day', () => {
const registry = new RebootRestoreRegistry();
registry.set([entryFor('a')]);
const dayLater = Date.now() + 25 * HOUR;
const realNow = Date.now;
Date.now = () => dayLater;
try {
expect(registry.list(all)).toEqual([]);
} finally {
Date.now = realNow;
}
});
});
describe('which conversation a rebuilt pane resumes', () => {
it('prefers the chain tail, the conversation the CLI reported last', () => {
const state = persistedSession({
id: 'sess-1',
resumeSessionId: 'launch-id',
claudeSessionChain: ['launch-id', 'after-clear'],
});
expect(resolveResumeConversationId(state)).toBe('after-clear');
});
it('falls back to the id the session originally resumed', () => {
const state = persistedSession({ id: 'sess-1', resumeSessionId: 'resumed-id' });
expect(resolveResumeConversationId(state)).toBe('resumed-id');
});
it('falls back to the session id, which is what Claude was launched with', () => {
expect(resolveResumeConversationId(persistedSession({ id: 'sess-1' }))).toBe('sess-1');
});
});
describe('the recovery construction path can create a resumed pane', () => {
const workingDir = join(homedir(), 'codeman-cases', 'reboot-restore-spike');
const sessions: Session[] = [];
afterEach(() => {
for (const s of sessions.splice(0)) s.stop();
rmSync(workingDir, { recursive: true, force: true });
});
/** Built exactly as the reboot pass builds one: no `muxSession`, plus a resume id. */
function rebuildFromPersistedState(state: SessionState, mux: TmuxManager): Session {
mkdirSync(workingDir, { recursive: true });
const session = new Session({
id: state.id,
workingDir,
mode: state.mode,
name: state.name,
createdAt: state.createdAt,
mux,
useMux: true,
resumeSessionId: resolveResumeConversationId(state),
owner: state.owner,
lastActivityAt: state.lastActivityAt,
claudeSessionChain: state.claudeSessionChain,
});
sessions.push(session);
return session;
}
it('creates a NEW mux session rather than needing one to attach to', async () => {
const mux = new TmuxManager();
const state = persistedSession({ id: 'aaaaaaa1-1111-4111-8111-111111111111', name: 'w1-spike' });
const session = rebuildFromPersistedState(state, mux);
expect(mux.getSessions()).toHaveLength(0);
await session.startInteractive();
const created = mux.getSessions();
expect(created).toHaveLength(1);
expect(created[0].sessionId).toBe('aaaaaaa1-1111-4111-8111-111111111111');
expect(created[0].workingDir).toBe(workingDir);
});
it('comes back pointed at the conversation the pane was holding', async () => {
const mux = new TmuxManager();
const state = persistedSession({
id: 'aaaaaaa2-2222-4222-8222-222222222222',
resumeSessionId: 'launch-id',
claudeSessionChain: ['launch-id', 'after-clear'],
});
const session = rebuildFromPersistedState(state, mux);
await session.startInteractive();
// The chain tail wins: a `/clear` before the reboot moved the CLI off the launch id.
expect(session.claudeSessionId).toBe('after-clear');
});
it('comes back idle, with no prompt sent and no autonomous loop armed', async () => {
const mux = new TmuxManager();
const state = persistedSession({
id: 'aaaaaaa3-3333-4333-8333-333333333333',
ralphEnabled: true,
respawnEnabled: true,
});
const session = rebuildFromPersistedState(state, mux);
await session.startInteractive();
// No prompt was queued: nothing is waiting on a task. The status itself is not
// assertable here, because the test PTY echoes and the activity detector reads
// that echo as work; in production the pane settles once the CLI finishes booting.
expect(session.currentTaskId).toBeNull();
// The pass never touches the tracker, so a persisted Ralph loop stays cold.
expect(session.ralphTracker.enabled).toBe(false);
});
it('keeps the owner it was persisted with, there being no request to read one from', async () => {
const mux = new TmuxManager();
const state = persistedSession({ id: 'aaaaaaa4-4444-4444-8444-444444444444', owner: 'alice' });
const session = rebuildFromPersistedState(state, mux);
await session.startInteractive();
expect(session.owner).toBe('alice');
expect(mux.getSessions()[0].owner).toBe('alice');
});
});
+175
View File
@@ -0,0 +1,175 @@
/**
* Reboot-restore route tests (src/web/routes/reboot-restore-routes.ts) via
* app.inject(), no live port.
*
* Every entry these tests put on offer names a workspace that does not exist, so
* the route's click-time workspace check rejects it before any `Session` is
* constructed. That keeps the tests on the route's own guards — taking, scoping,
* single-flighting and re-checking — and leaves pane creation to
* test/reboot-restore.test.ts, which drives a real `Session` for it.
*
* The routes read the process-wide `rebootRestoreRegistry` singleton, so every
* test resets it; a leaked entry would bleed into the next one.
*/
import { describe, it, expect, afterEach } from 'vitest';
import Fastify, { type FastifyInstance } from 'fastify';
import fastifyCookie from '@fastify/cookie';
import { registerRebootRestoreRoutes } from '../../src/web/routes/reboot-restore-routes.js';
import { rebootRestoreRegistry } from '../../src/web/reboot-restore-registry.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { httpStatusForErrorCode, type ApiErrorCode } from '../../src/types.js';
import { createMockRouteContext } from '../mocks/index.js';
import type { RebootRestoreEntry } from '../../src/reboot-restore.js';
import type { SessionState } from '../../src/types.js';
async function createHarness(authUser?: { username: string; role: 'admin' | 'user' }): Promise<FastifyInstance> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
if (authUser) {
app.addHook('onRequest', async (req) => {
(req as unknown as { authUser: typeof authUser }).authUser = authUser;
});
}
registerRebootRestoreRoutes(app, createMockRouteContext() as never);
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
if (!req.url.startsWith('/api')) return done(null, payload);
if (payload === null || typeof payload !== 'object') return done(null, payload);
const p = payload as { success?: unknown; errorCode?: unknown };
if (p.success === false) {
if (reply.statusCode === 200 && typeof p.errorCode === 'string') {
reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
}
return done(null, payload);
}
if (p.success === true) return done(null, payload);
return done(null, { success: true, data: payload });
});
installRouteErrorHandler(app);
await app.ready();
return app;
}
/** An entry whose workspace is deliberately absent, so no pane is ever created. */
function offerEntry(sessionId: string, owner?: string): RebootRestoreEntry {
return {
sessionId,
name: `session ${sessionId}`,
workingDir: `/tmp/codeman-reboot-restore-missing/${sessionId}`,
owner,
mode: 'claude',
resumeConversationId: `conv-${sessionId}`,
state: {
id: sessionId,
pid: null,
status: 'idle',
workingDir: `/tmp/codeman-reboot-restore-missing/${sessionId}`,
currentTaskId: null,
createdAt: 1_760_000_000_000,
mode: 'claude',
owner,
} as SessionState,
};
}
afterEach(() => {
rebootRestoreRegistry.reset();
});
describe('GET /api/reboot-restore', () => {
it('reports nothing when no reboot left anything behind', async () => {
const app = await createHarness();
const res = await app.inject({ method: 'GET', url: '/api/reboot-restore' });
expect(res.statusCode).toBe(200);
expect(res.json().data.sessions).toEqual([]);
await app.close();
});
it('names what is on offer, and says the scrollback is not coming back', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
const app = await createHarness();
const body = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(body.sessions.map((s: { id: string }) => s.id)).toEqual(['a', 'b']);
expect(body.scrollbackRestored).toBe(false);
await app.close();
});
it('never carries the persisted record itself to the browser', async () => {
rebootRestoreRegistry.set([offerEntry('a', 'alice')]);
const app = await createHarness({ username: 'alice', role: 'admin' });
const body = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(Object.keys(body.sessions[0]).sort()).toEqual(['id', 'mode', 'name', 'owner', 'workingDir']);
expect(body.sessions[0].state).toBeUndefined();
await app.close();
});
});
describe('POST /api/reboot-restore/restore', () => {
it('spends the offer, so a second click finds nothing left to spend', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
const app = await createHarness();
const first = (await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} })).json().data;
// The workspace is gone, so nothing was rebuilt — but the entry was taken.
expect(first.restored).toEqual([]);
expect(first.skipped).toEqual([{ sessionId: 'a', reason: 'workspace-missing' }]);
const second = (await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} })).json().data;
expect(second.restored).toEqual([]);
expect(second.skipped).toEqual([]);
await app.close();
});
it('spends only the sessions the click named', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
const app = await createHarness();
const res = await app.inject({
method: 'POST',
url: '/api/reboot-restore/restore',
payload: { sessionIds: ['b'] },
});
expect(res.json().data.skipped).toEqual([{ sessionId: 'b', reason: 'workspace-missing' }]);
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions.map((s: { id: string }) => s.id)).toEqual(['a']);
await app.close();
});
it('refuses a body it does not recognise rather than guessing', async () => {
const app = await createHarness();
const res = await app.inject({
method: 'POST',
url: '/api/reboot-restore/restore',
payload: { sessionIds: 'not-an-array' },
});
expect(res.statusCode).toBeGreaterThanOrEqual(400);
await app.close();
});
it('turns a second concurrent restore away rather than interleaving it', async () => {
rebootRestoreRegistry.set([offerEntry('a')]);
// Claimed by a restore already in flight.
expect(rebootRestoreRegistry.beginSpending()).toBe(true);
const app = await createHarness();
const res = await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
expect(res.statusCode).toBe(409);
rebootRestoreRegistry.endSpending();
await app.close();
});
});
describe('POST /api/reboot-restore/dismiss', () => {
it('drops the offer and leaves the banner with nothing to show', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
const app = await createHarness();
const res = await app.inject({ method: 'POST', url: '/api/reboot-restore/dismiss', payload: {} });
expect(res.json().data.dismissed).toBe(2);
const after = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(after.sessions).toEqual([]);
await app.close();
});
});