fix(sessions): act on the dual review of the reboot-restore route

Fifteen findings from two independent reviews of #442, three of them
blocking. Every one is addressed here.

The three blockers all sat in the restore route. A rebuild that threw after
addSession left a registered session with no pane behind it, visible on the
board, holding a layout slot and written to state.json, with its plan entry
already spent; the catch now cleans the session up and puts the entry back.
The loop checked neither the global nor the per-user session cap, so one
click could take a board past a documented limit; capacity is now re-checked
per iteration, because the loop is itself creating the sessions it counts.
Worst of the three, a rebuilt session carried none of the state its
constructor has no parameter for and then persisted itself over the record
that held it, zeroing token and cost totals and dropping the pin. The pin
matters most: pruning keeps a record only while it is pinned, so discarding
it handed the record to the next stale sweep. A new
reapplyPersistedSessionState() on the session port restores the pin, the
token totals, auto-compact, auto-clear, auto-resume, nice priority, the
flicker filter and the custom-model selection, and it runs before both
startInteractive and the first persist.

The rest, in the order they bite a user. Every rebuild failure was reported
as workspace-missing, so the banner told users their repo was gone when the
agent had simply failed to start; there are now distinct reasons, and the
toast names each one. The client read restored and skipped off the outer
response object rather than through the uniform envelope, so every count
came back zero and neither toast ever fired. A board left open across the
reboot never learned an offer existed, because the banner was seeded only on
the page-load path; it now re-reads on every SSE init. The workspace check
was existence-only, skipping the multi-user confinement that the create
route applies, so a withdrawn grant would not be noticed. The banner had no
phone breakpoint while its text was nowrap and its buttons could not shrink.

Smaller: a missing workspace is now re-offered rather than dropped, while an
already-open conversation is dropped rather than re-offered forever; a throw
anywhere in the route returns the unspent entries instead of discarding the
plan; the single flight is keyed by owner, since take() already stops two
callers receiving one entry; the env clamp's header no longer claims a
protection it cannot provide on this path today, and names the check that
does bite; the three endpoints are documented in docs/api-reference.md; and
the module header now says that os.uptime() reads the host's clock, so the
feature is effectively off inside a container.

The review also explained why the tests missed all of this: they proved the
construction claim through their own copy of the construction rather than
through the route, and the route tests used workspaces that did not exist,
so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts
mocks the Session module to drive the route's real path, and covers the
cleanup, the reason reported, the re-application ordering, the broadcast and
the caps. The mock route context gains the port method and the mux call the
route needs.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Michael Grundberg
2026-09-16 11:44:37 +02:00
co-authored by Claude Opus 5
parent da933d70be
commit fbede5cd2a
13 changed files with 632 additions and 121 deletions
+24 -15
View File
@@ -24,8 +24,9 @@
* - Spending is take-then-build: `take()` removes entries synchronously, before
* the route's first `await`, so a double-click or two devices cannot both
* reach the same entry and put two panes on one conversation.
* - One restore runs at a time. `beginSpending()` single-flights the route, so
* two concurrent clicks cannot interleave pane creation.
* - One restore runs at a time per owner. `beginSpending()` single-flights the
* route, so two concurrent clicks cannot interleave pane creation for the same
* user, while two different users never block each other.
*
* @dependencies reboot-restore (RebootRestoreEntry)
* @consumedby web/server (plan build at boot), web/routes/reboot-restore-routes
@@ -46,8 +47,13 @@ export class RebootRestoreRegistry {
private entries = new Map<string, RebootRestoreEntry>();
/** When the boot pass built the plan, in ms since the epoch. */
private builtAt = 0;
/** True while a restore route call is between its take and its last pane. */
private spending = false;
/**
* Owners with a restore in flight, between its take and its last pane.
* Keyed by owner so one user's restore does not turn another user's click into
* a conflict; `take()` already guarantees no two callers get the same entry.
* Single-user mode has one key, `undefined`, so it behaves as one global flight.
*/
private spending = new Set<string | undefined>();
/** Replace the plan with what the boot pass found. An empty list clears it. */
set(entries: readonly RebootRestoreEntry[]): void {
@@ -92,9 +98,11 @@ export class RebootRestoreRegistry {
/**
* Put entries back after a rebuild never got as far as creating a pane.
*
* Used for the click-time rejections, so a conversation the user resumed by
* hand meanwhile does not silently vanish from the banner while a workspace
* that came back stays offered.
* Used for the click-time rejections that may resolve themselves: a workspace
* that comes back, a capacity limit the user makes room under, a CLI that
* starts once its binary is on the PATH. A conversation the user resumed by
* hand is NOT put back, because that one cannot stop being true, and an entry
* the banner keeps re-offering forever is noise only Dismiss can clear.
*/
restore(entries: readonly RebootRestoreEntry[]): void {
for (const entry of entries) this.entries.set(entry.sessionId, entry);
@@ -110,24 +118,25 @@ export class RebootRestoreRegistry {
}
/**
* Claim the right to run a restore, or report that one is already running.
* Callers that get `true` must call `endSpending()` in a `finally`.
* Claim the right to run a restore for one owner, or report that owner already
* has one running. Callers that get `true` must call `endSpending()` in a
* `finally` with the same owner.
*/
beginSpending(): boolean {
if (this.spending) return false;
this.spending = true;
beginSpending(owner?: string): boolean {
if (this.spending.has(owner)) return false;
this.spending.add(owner);
return true;
}
endSpending(): void {
this.spending = false;
endSpending(owner?: string): void {
this.spending.delete(owner);
}
/** Test hook: forget everything, including the single-flight claim. */
reset(): void {
this.entries.clear();
this.builtAt = 0;
this.spending = false;
this.spending.clear();
}
private dropIfExpired(): void {