mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-06 15:39:41 +02:00
fix(sessions): undo a failed rebuild without deleting the user's data
A second review of the previous commit found that its own repair for the session leak introduced three defects, all from reaching for cleanupSession() to undo a half-built session. That function is the user-initiated delete, not an undo. It banked the session's historical token and cost totals into the lifetime figures, and a reboot never runs cleanup, so those totals had never been counted before; every failed rebuild added them again. It saw the pin that had just been restored and demoted the record to `stopped`, which this pass reads as the durable marker of a deliberate kill, so a pinned session whose rebuild failed became permanently unrestorable. And it recursively removed `.claude-images` from the working directory, which belongs to the workspace rather than to the session, so a failed rebuild destroyed the pasted images of any other live session in that repo. discardPartiallyBuiltSession() now undoes only what the construction did: the map entry, the tab-layout slot, the listeners and any pane the launch created before throwing. The persisted record, the lifetime totals, the Ralph state and the workspace's files are left alone. Re-applying the persisted state also splits in two, which removes the first two defects at the root rather than only at the call site. The half that shapes the pane, the custom-model environment and the nice priority, still runs before the spawn. The half that is the session's own history now runs after it, so a session whose pane never started carries no totals and no pin for anything downstream to misread. The rest of that review. The multi-user workspace confinement re-check read the requesting user's grant, and returns true for an admin, so the case its own comment described was the one it missed; it now resolves the entry owner's grant through isWorkingDirAllowedForUsername, the way cron does. A forbidden workspace goes back on offer, matching both the registry's stated contract and the API reference. The client re-reads the plan after a restore instead of blanking the banner, so entries the server put back stay reachable, and a 409 now says a restore is already running rather than reporting a failure. A dismiss arriving mid-restore wins, through a generation counter the route carries across its take. The re-application also restores the tab colour, the image-watcher flag and the original pinnedAt, via a new Session.restorePin that does not re-stamp the pin time. The phone breakpoint gains min-width: 0, without which a nowrap flex item never shrinks and the buttons still overflow, and it folds into the existing phone block. Ralph's loop configuration still does not survive a restore, because toState() reads it off a live tracker and there is no way to keep it without arming the loop. The method now says so rather than leaving it implied. Tests. The capacity test could not fail on the property it existed for: it filled the board past the cap before the loop, so a single pre-loop check would have passed it. It now leaves one seat, so only a per-iteration check restores exactly one entry. New tests cover the ordering around the spawn, a throw before the loop returning the whole plan and releasing the flight, the dismiss-during-restore race, and that the failure path calls the narrow discard rather than the delete. The shared mock context gains the port method it was missing, which is what made the first run of these tests fail for the wrong reason. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
fbede5cd2a
commit
fa52753e8b
@@ -32,7 +32,7 @@ import {
|
||||
getAuthUser,
|
||||
canAccessOwned,
|
||||
ownerFor,
|
||||
isWorkingDirAllowed,
|
||||
isWorkingDirAllowedForUsername,
|
||||
sessionCapacityMessage,
|
||||
} from '../route-helpers.js';
|
||||
import { rebootRestoreRegistry } from '../reboot-restore-registry.js';
|
||||
@@ -83,7 +83,6 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
|
||||
app.post('/api/reboot-restore/restore', async (req, reply) => {
|
||||
const body = parseBody(RebootRestoreRequestSchema, req.body, 'Invalid reboot restore request');
|
||||
const user = getAuthUser(req);
|
||||
const canAccess = accessorFor(req);
|
||||
const owner = ownerFor(req);
|
||||
|
||||
@@ -93,6 +92,7 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
if (!rebootRestoreRegistry.beginSpending(owner)) {
|
||||
return reply.code(409).send(createErrorResponse(ApiErrorCode.CONFLICT, 'A reboot restore is already running'));
|
||||
}
|
||||
const generation = rebootRestoreRegistry.currentGeneration();
|
||||
const taken = rebootRestoreRegistry.take(canAccess, body.sessionIds);
|
||||
// Entries nothing built a pane for, returned to the plan on every exit path
|
||||
// including a throw. Without this a failure between here and the loop would
|
||||
@@ -136,10 +136,15 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
// Multi-user workspace separation: the create route confines a non-admin's
|
||||
// workingDir to their own case space, and a grant can be withdrawn between
|
||||
// the session's creation and this restore, so the confinement is re-run
|
||||
// rather than inherited from the record.
|
||||
if (!isWorkingDirAllowed(user, entry.workingDir)) {
|
||||
// rather than inherited from the record. Keyed on the OWNER, not on the
|
||||
// caller: an admin spending another user's entry must be held to that
|
||||
// user's confinement, and `isWorkingDirAllowed` would wave an admin
|
||||
// through. The same reason the two grant re-checks below read
|
||||
// `saved.owner`.
|
||||
if (!(await isWorkingDirAllowedForUsername(entry.owner, entry.workingDir))) {
|
||||
// Left on offer: a withdrawn grant can be restored, unlike an already-open
|
||||
// conversation, so this is not the permanent kind of refusal.
|
||||
failures.push({ sessionId: entry.sessionId, reason: 'workspace-forbidden' });
|
||||
unspent.delete(entry);
|
||||
continue;
|
||||
}
|
||||
try {
|
||||
@@ -180,12 +185,16 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
|
||||
await ctx.addSession(session);
|
||||
await ctx.setupSessionListeners(session);
|
||||
// Before the pane spawns: the custom-model selection reaches it through
|
||||
// the environment. Before the first persist: a constructed session holds
|
||||
// none of this, so persisting it first would replace the fuller record
|
||||
// with the reduced one and drop the pin that keeps it from being pruned.
|
||||
await ctx.reapplyPersistedSessionState(session, saved);
|
||||
// Shapes the pane, so it has to land before the CLI process starts.
|
||||
await ctx.reapplyPersistedSessionState(session, saved, 'before-spawn');
|
||||
await session.startInteractive();
|
||||
// The session's own history, applied only once the pane exists: on a
|
||||
// failed start these totals would belong to a session that never ran.
|
||||
// Both halves precede the first persist, because a constructed session
|
||||
// carries none of this and `toState()` is written wholesale, so
|
||||
// persisting first would replace the fuller record with the reduced one
|
||||
// and drop the pin that keeps it from being pruned.
|
||||
await ctx.reapplyPersistedSessionState(session, saved, 'after-spawn');
|
||||
ctx.persistSessionState(session);
|
||||
|
||||
// A session without its workspace hooks goes silently blind: no stop or
|
||||
@@ -211,10 +220,15 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
// listeners, and the commonest cause is a CLI binary that is not on the
|
||||
// PATH of a freshly booted machine.
|
||||
console.error(`[reboot-restore] failed to rebuild ${entry.sessionId}:`, err);
|
||||
// Not cleanupSession(): that is the user-initiated delete, and it would
|
||||
// count this session's historical tokens into the lifetime totals, demote
|
||||
// a pinned record to `stopped` (which this pass reads as an intentional
|
||||
// kill, making the session permanently unrestorable) and delete the
|
||||
// workspace's `.claude-images`. This undoes only the construction.
|
||||
await ctx
|
||||
.cleanupSession(entry.sessionId, true, 'reboot restore failed to start the session')
|
||||
.catch((cleanupErr: unknown) =>
|
||||
console.error(`[reboot-restore] cleanup after a failed rebuild failed: ${getErrorMessage(cleanupErr)}`)
|
||||
.discardPartiallyBuiltSession(entry.sessionId)
|
||||
.catch((discardErr: unknown) =>
|
||||
console.error(`[reboot-restore] discarding a failed rebuild failed: ${getErrorMessage(discardErr)}`)
|
||||
);
|
||||
failures.push({ sessionId: entry.sessionId, reason: 'rebuild-failed' });
|
||||
// Left on offer: the user can put the binary back and click again.
|
||||
@@ -234,7 +248,8 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
} finally {
|
||||
// Anything that never became a pane goes back on offer, including after a
|
||||
// throw, so a transient failure costs a retry rather than the whole plan.
|
||||
rebootRestoreRegistry.restore([...unspent]);
|
||||
// Passing the generation makes a Dismiss that landed mid-restore win.
|
||||
rebootRestoreRegistry.restore([...unspent], generation);
|
||||
rebootRestoreRegistry.endSpending(owner);
|
||||
}
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user