fix(sessions): undo a failed rebuild without deleting the user's data

A second review of the previous commit found that its own repair for the
session leak introduced three defects, all from reaching for
cleanupSession() to undo a half-built session. That function is the
user-initiated delete, not an undo.

It banked the session's historical token and cost totals into the lifetime
figures, and a reboot never runs cleanup, so those totals had never been
counted before; every failed rebuild added them again. It saw the pin that
had just been restored and demoted the record to `stopped`, which this pass
reads as the durable marker of a deliberate kill, so a pinned session whose
rebuild failed became permanently unrestorable. And it recursively removed
`.claude-images` from the working directory, which belongs to the workspace
rather than to the session, so a failed rebuild destroyed the pasted images
of any other live session in that repo.

discardPartiallyBuiltSession() now undoes only what the construction did:
the map entry, the tab-layout slot, the listeners and any pane the launch
created before throwing. The persisted record, the lifetime totals, the
Ralph state and the workspace's files are left alone.

Re-applying the persisted state also splits in two, which removes the first
two defects at the root rather than only at the call site. The half that
shapes the pane, the custom-model environment and the nice priority, still
runs before the spawn. The half that is the session's own history now runs
after it, so a session whose pane never started carries no totals and no pin
for anything downstream to misread.

The rest of that review. The multi-user workspace confinement re-check read
the requesting user's grant, and returns true for an admin, so the case its
own comment described was the one it missed; it now resolves the entry
owner's grant through isWorkingDirAllowedForUsername, the way cron does. A
forbidden workspace goes back on offer, matching both the registry's stated
contract and the API reference. The client re-reads the plan after a restore
instead of blanking the banner, so entries the server put back stay
reachable, and a 409 now says a restore is already running rather than
reporting a failure. A dismiss arriving mid-restore wins, through a
generation counter the route carries across its take. The re-application
also restores the tab colour, the image-watcher flag and the original
pinnedAt, via a new Session.restorePin that does not re-stamp the pin time.
The phone breakpoint gains min-width: 0, without which a nowrap flex item
never shrinks and the buttons still overflow, and it folds into the existing
phone block.

Ralph's loop configuration still does not survive a restore, because
toState() reads it off a live tracker and there is no way to keep it without
arming the loop. The method now says so rather than leaving it implied.

Tests. The capacity test could not fail on the property it existed for: it
filled the board past the cap before the loop, so a single pre-loop check
would have passed it. It now leaves one seat, so only a per-iteration check
restores exactly one entry. New tests cover the ordering around the spawn,
a throw before the loop returning the whole plan and releasing the flight,
the dismiss-during-restore race, and that the failure path calls the narrow
discard rather than the delete. The shared mock context gains the port
method it was missing, which is what made the first run of these tests fail
for the wrong reason.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Michael Grundberg
2026-09-16 15:04:35 +02:00
co-authored by Claude Opus 5
parent fbede5cd2a
commit fa52753e8b
9 changed files with 273 additions and 47 deletions
+29 -14
View File
@@ -32,7 +32,7 @@ import {
getAuthUser,
canAccessOwned,
ownerFor,
isWorkingDirAllowed,
isWorkingDirAllowedForUsername,
sessionCapacityMessage,
} from '../route-helpers.js';
import { rebootRestoreRegistry } from '../reboot-restore-registry.js';
@@ -83,7 +83,6 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
app.post('/api/reboot-restore/restore', async (req, reply) => {
const body = parseBody(RebootRestoreRequestSchema, req.body, 'Invalid reboot restore request');
const user = getAuthUser(req);
const canAccess = accessorFor(req);
const owner = ownerFor(req);
@@ -93,6 +92,7 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
if (!rebootRestoreRegistry.beginSpending(owner)) {
return reply.code(409).send(createErrorResponse(ApiErrorCode.CONFLICT, 'A reboot restore is already running'));
}
const generation = rebootRestoreRegistry.currentGeneration();
const taken = rebootRestoreRegistry.take(canAccess, body.sessionIds);
// Entries nothing built a pane for, returned to the plan on every exit path
// including a throw. Without this a failure between here and the loop would
@@ -136,10 +136,15 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
// Multi-user workspace separation: the create route confines a non-admin's
// workingDir to their own case space, and a grant can be withdrawn between
// the session's creation and this restore, so the confinement is re-run
// rather than inherited from the record.
if (!isWorkingDirAllowed(user, entry.workingDir)) {
// rather than inherited from the record. Keyed on the OWNER, not on the
// caller: an admin spending another user's entry must be held to that
// user's confinement, and `isWorkingDirAllowed` would wave an admin
// through. The same reason the two grant re-checks below read
// `saved.owner`.
if (!(await isWorkingDirAllowedForUsername(entry.owner, entry.workingDir))) {
// Left on offer: a withdrawn grant can be restored, unlike an already-open
// conversation, so this is not the permanent kind of refusal.
failures.push({ sessionId: entry.sessionId, reason: 'workspace-forbidden' });
unspent.delete(entry);
continue;
}
try {
@@ -180,12 +185,16 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
await ctx.addSession(session);
await ctx.setupSessionListeners(session);
// Before the pane spawns: the custom-model selection reaches it through
// the environment. Before the first persist: a constructed session holds
// none of this, so persisting it first would replace the fuller record
// with the reduced one and drop the pin that keeps it from being pruned.
await ctx.reapplyPersistedSessionState(session, saved);
// Shapes the pane, so it has to land before the CLI process starts.
await ctx.reapplyPersistedSessionState(session, saved, 'before-spawn');
await session.startInteractive();
// The session's own history, applied only once the pane exists: on a
// failed start these totals would belong to a session that never ran.
// Both halves precede the first persist, because a constructed session
// carries none of this and `toState()` is written wholesale, so
// persisting first would replace the fuller record with the reduced one
// and drop the pin that keeps it from being pruned.
await ctx.reapplyPersistedSessionState(session, saved, 'after-spawn');
ctx.persistSessionState(session);
// A session without its workspace hooks goes silently blind: no stop or
@@ -211,10 +220,15 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
// listeners, and the commonest cause is a CLI binary that is not on the
// PATH of a freshly booted machine.
console.error(`[reboot-restore] failed to rebuild ${entry.sessionId}:`, err);
// Not cleanupSession(): that is the user-initiated delete, and it would
// count this session's historical tokens into the lifetime totals, demote
// a pinned record to `stopped` (which this pass reads as an intentional
// kill, making the session permanently unrestorable) and delete the
// workspace's `.claude-images`. This undoes only the construction.
await ctx
.cleanupSession(entry.sessionId, true, 'reboot restore failed to start the session')
.catch((cleanupErr: unknown) =>
console.error(`[reboot-restore] cleanup after a failed rebuild failed: ${getErrorMessage(cleanupErr)}`)
.discardPartiallyBuiltSession(entry.sessionId)
.catch((discardErr: unknown) =>
console.error(`[reboot-restore] discarding a failed rebuild failed: ${getErrorMessage(discardErr)}`)
);
failures.push({ sessionId: entry.sessionId, reason: 'rebuild-failed' });
// Left on offer: the user can put the binary back and click again.
@@ -234,7 +248,8 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
} finally {
// Anything that never became a pane goes back on offer, including after a
// throw, so a transient failure costs a retry rather than the whole plan.
rebootRestoreRegistry.restore([...unspent]);
// Passing the generation makes a Dismiss that landed mid-restore win.
rebootRestoreRegistry.restore([...unspent], generation);
rebootRestoreRegistry.endSpending(owner);
}
});