Commit Graph
5 Commits
Author SHA1 Message Date
Codeman maintainer bb8ada7e5f fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does.

1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred
   it from the name: a session the user renamed by hand to something shaped
   like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the
   next prompt overwrote their name. The route persists right after, so the
   loss went to disk. `restoreMuxSessions()` already passes it.

2. The already-live sets were snapshotted once before a loop that awaits a
   real `startInteractive()` per entry, so by the tenth entry the snapshot
   was tens of seconds old and a conversation resumed by hand from the
   Resume list in that window was invisible to it: two panes on one
   transcript, the exact thing the check exists to prevent. Both sets are
   now read per iteration, and the late case is spent rather than re-offered
   for the same reason the batch case is.

3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path.
   The stamp predates the reboot and the pane is new, so honouring it meant
   one click had every restored session type `continue` into itself about a
   minute later, unattended, against the route header's own promise that a
   restored session comes back idle and disarmed. The setting stays ENABLED,
   so it re-arms on the next real limit message. A Codeman restart still
   re-arms from the stamp, because the limit footer will not reprint on its
   own; the new option exists only to tell the two paths apart.

4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()`
   and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession`
   performs that it was missing. Cosmetic, but a run left open reads as
   still going in the away digest.

5. A restored claude session gets `seedAgentSessionPreamble()` like both
   create paths, so the agent skill's bootstrap stays a two-line loader.

6. The heuristic's container comment was wrong in one direction and quiet
   about the real gap: after a genuine host reboot a containerized Codeman
   sees the host's short uptime and the banner does appear. What it cannot
   see is a container-only restart, which is where this would help most.

7. The banner is hidden in a solo window, which shows one session and has
   no tab strip to put restored ones in.

Also reverts 17 of the 18 hunks in docs/api-reference.md, which were
Prettier reformatting of prose the PR does not otherwise touch (docs/ is
outside the format glob), keeping only the Reboot restore section and
repairing the two continuation lines that reformat de-indented; renumbers
reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which
loads after it; and gives the feature its CLAUDE.md entry plus a route
test for the multi-user workspace-forbidden branch, the only new rule that
had nothing behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:46:04 +02:00
Michael GrundbergandClaude Opus 5 71ed7b127c fix(sessions): make the discard a real inverse of the construction
Third review of the reboot-restore branch. The narrow discard the previous
commit introduced avoided everything cleanupSession() did wrongly, and in
dropping so much of it also dropped four things it had to keep.

The worst broke the retry the whole design rests on. setupSessionListeners()
returns early while sessionListenerRefs still holds the session id, and the
discard never cleared that entry. So the advertised flow — a rebuild fails
because the agent binary is missing, the user fixes their PATH and clicks
again — reused the same id, wired no listeners at all, and produced a tab
that never showed output, never updated its status and never persisted. That
is worse than the leak the discard was added to prevent. Three more
registrations leaked with it: a RunSummaryTracker and its interval, an image
watcher on the workspace, and the Ralph fix-plan watcher. The discard now
undoes each registration setupSessionListeners() makes, in its order, and
the per-session custom-model config directory, which holds the endpoint's
API key literally and which nothing else would ever remove.

The image-watcher flag was restored after the code that reads it, so a
session came back reporting the feature as on with nothing watching. It
moves to the before-spawn phase, and that phase now runs before the
listeners rather than after them.

The generation counter that lets a mid-restore dismiss win was global while
clear() is ownership-scoped, so one user's dismiss discarded another user's
unspent entries, permanently, because nothing rebuilds an in-memory plan. It
is now per owner. Bumping only the owners of entries the dismiss removed was
not enough either: take() has already emptied the plan by then, so a dismiss
landing mid-restore saw nothing of that owner's to remove and invalidated
nothing. The owners that matter are those with a restore in flight, filtered
by what the dismissing user may access, and that is what clear() now bumps.
Plan expiry bumps too, so a restore straddling the 24-hour boundary cannot
hand entries back and give an expired plan another full day.

Tests. discardPartiallyBuiltSession had no test at all: the only
implementation any test ran was the mock's one-line stub, which is why every
defect above was invisible. test/discard-partially-built-session.ts drives
the real WebServer, and the retry assertion fails if the listener refs are
left behind — verified by reverting the fix. The dismiss-race test drove the
registry by hand, so deleting the route's generation argument left it green;
it now goes through the route, and two further tests cover the multi-user
cases.

The mock context has now gone stale twice, because route tests pass it as
`ctx as never` and tsconfig.json includes only src, so nothing ever compares
it to the ports. A type-level guard is therefore inert — I wrote one and
confirmed it never fires. test/mocks/mock-route-context-completeness.ts
compares the mock's keys against WebServer.createRouteContext() at runtime
instead, and names what is missing.

Also: the API reference now says workspace-forbidden is judged against the
owner's grant, the banner's module header no longer claims Restore always
dismisses it, and the detail span gets the same min-width: 0 the phone rule
already needed.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:34:47 +02:00
Michael GrundbergandClaude Opus 5 fa52753e8b fix(sessions): undo a failed rebuild without deleting the user's data
A second review of the previous commit found that its own repair for the
session leak introduced three defects, all from reaching for
cleanupSession() to undo a half-built session. That function is the
user-initiated delete, not an undo.

It banked the session's historical token and cost totals into the lifetime
figures, and a reboot never runs cleanup, so those totals had never been
counted before; every failed rebuild added them again. It saw the pin that
had just been restored and demoted the record to `stopped`, which this pass
reads as the durable marker of a deliberate kill, so a pinned session whose
rebuild failed became permanently unrestorable. And it recursively removed
`.claude-images` from the working directory, which belongs to the workspace
rather than to the session, so a failed rebuild destroyed the pasted images
of any other live session in that repo.

discardPartiallyBuiltSession() now undoes only what the construction did:
the map entry, the tab-layout slot, the listeners and any pane the launch
created before throwing. The persisted record, the lifetime totals, the
Ralph state and the workspace's files are left alone.

Re-applying the persisted state also splits in two, which removes the first
two defects at the root rather than only at the call site. The half that
shapes the pane, the custom-model environment and the nice priority, still
runs before the spawn. The half that is the session's own history now runs
after it, so a session whose pane never started carries no totals and no pin
for anything downstream to misread.

The rest of that review. The multi-user workspace confinement re-check read
the requesting user's grant, and returns true for an admin, so the case its
own comment described was the one it missed; it now resolves the entry
owner's grant through isWorkingDirAllowedForUsername, the way cron does. A
forbidden workspace goes back on offer, matching both the registry's stated
contract and the API reference. The client re-reads the plan after a restore
instead of blanking the banner, so entries the server put back stay
reachable, and a 409 now says a restore is already running rather than
reporting a failure. A dismiss arriving mid-restore wins, through a
generation counter the route carries across its take. The re-application
also restores the tab colour, the image-watcher flag and the original
pinnedAt, via a new Session.restorePin that does not re-stamp the pin time.
The phone breakpoint gains min-width: 0, without which a nowrap flex item
never shrinks and the buttons still overflow, and it folds into the existing
phone block.

Ralph's loop configuration still does not survive a restore, because
toState() reads it off a live tracker and there is no way to keep it without
arming the loop. The method now says so rather than leaving it implied.

Tests. The capacity test could not fail on the property it existed for: it
filled the board past the cap before the loop, so a single pre-loop check
would have passed it. It now leaves one seat, so only a per-iteration check
restores exactly one entry. New tests cover the ordering around the spawn,
a throw before the loop returning the whole plan and releasing the flight,
the dismiss-during-restore race, and that the failure path calls the narrow
discard rather than the delete. The shared mock context gains the port
method it was missing, which is what made the first run of these tests fail
for the wrong reason.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:04:35 +02:00
Michael GrundbergandClaude Opus 5 fbede5cd2a fix(sessions): act on the dual review of the reboot-restore route
Fifteen findings from two independent reviews of #442, three of them
blocking. Every one is addressed here.

The three blockers all sat in the restore route. A rebuild that threw after
addSession left a registered session with no pane behind it, visible on the
board, holding a layout slot and written to state.json, with its plan entry
already spent; the catch now cleans the session up and puts the entry back.
The loop checked neither the global nor the per-user session cap, so one
click could take a board past a documented limit; capacity is now re-checked
per iteration, because the loop is itself creating the sessions it counts.
Worst of the three, a rebuilt session carried none of the state its
constructor has no parameter for and then persisted itself over the record
that held it, zeroing token and cost totals and dropping the pin. The pin
matters most: pruning keeps a record only while it is pinned, so discarding
it handed the record to the next stale sweep. A new
reapplyPersistedSessionState() on the session port restores the pin, the
token totals, auto-compact, auto-clear, auto-resume, nice priority, the
flicker filter and the custom-model selection, and it runs before both
startInteractive and the first persist.

The rest, in the order they bite a user. Every rebuild failure was reported
as workspace-missing, so the banner told users their repo was gone when the
agent had simply failed to start; there are now distinct reasons, and the
toast names each one. The client read restored and skipped off the outer
response object rather than through the uniform envelope, so every count
came back zero and neither toast ever fired. A board left open across the
reboot never learned an offer existed, because the banner was seeded only on
the page-load path; it now re-reads on every SSE init. The workspace check
was existence-only, skipping the multi-user confinement that the create
route applies, so a withdrawn grant would not be noticed. The banner had no
phone breakpoint while its text was nowrap and its buttons could not shrink.

Smaller: a missing workspace is now re-offered rather than dropped, while an
already-open conversation is dropped rather than re-offered forever; a throw
anywhere in the route returns the unspent entries instead of discarding the
plan; the single flight is keyed by owner, since take() already stops two
callers receiving one entry; the env clamp's header no longer claims a
protection it cannot provide on this path today, and names the check that
does bite; the three endpoints are documented in docs/api-reference.md; and
the module header now says that os.uptime() reads the host's clock, so the
feature is effectively off inside a container.

The review also explained why the tests missed all of this: they proved the
construction claim through their own copy of the construction rather than
through the route, and the route tests used workspaces that did not exist,
so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts
mocks the Session module to drive the route's real path, and covers the
cleanup, the reason reported, the re-application ordering, the broadcast and
the caps. The mock route context gains the port method and the mux call the
route needs.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 11:44:37 +02:00
Michael GrundbergandClaude Opus 5 da933d70be feat(sessions): offer to rebuild the sessions a host reboot destroyed
A host reboot takes the tmux server down with it, so every pane dies,
reconciliation finds nothing to attach to, and the board comes up empty.
Picking yesterday's work back up meant finding each conversation in history
and resuming it by hand, one at a time.

The boot pass now works out what the reboot killed and leaves it on offer.
It runs inside restoreMuxSessions(), in the window where reconciliation has
reported the dead sessions and cleanupStaleSessions() has not pruned their
records yet, which is the only place the records can still be read. The
board shows a banner, and nothing is created until the user clicks it.

A click rather than an automatic restore is what makes the reboot heuristic
acceptable. The heuristic cannot tell a reboot from a crash that took tmux
down inside the same window, so it decides whether to ASK, never whether to
act: a wrong yes costs a line of text the user dismisses instead of N CLI
processes nobody asked for.

Four things are re-checked when the click arrives rather than trusted from
boot, because hours can pass and the board moves on. The owner's privilege
grant re-resolves through the env clamp. The workspace must still be on
disk. A conversation the user already resumed by hand from the Resume list
is skipped, since two panes running --resume on one conversation would
fight over the same transcript. Entries leave the plan synchronously before
the first await, and the route is single-flighted, so a double-click or two
devices cannot both reach the same entry.

A restored session comes back attached, idle and disarmed. Respawn
controllers and Ralph loops are deliberately not re-armed: a machine that
just came up is the worst moment to turn an autonomous run loose. Its
workspace hooks are installed by the restore route itself, because the
boot-time sweep sits behind a gate that is false after a reboot and has
finished long before the click; without them a session goes silently blind,
with no stop or idle events for respawn, no Approvals Inbox item and no red
tab on a blocking dialog. Stats collection starts the same way.

The pane is new, so the conversation continues and the terminal scrollback
does not. The banner says so rather than letting an empty pane read as a
broken restore.

The plan lives in memory only. A server restart drops it, which costs the
convenience this adds and never the conversation: the conversation is the
transcript under ~/.claude/projects, which the Welcome screen's Resume list
and the Session Manager already read, so a dropped plan returns the user to
resuming by hand.

clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the
question it answers is about session privilege rather than about HTTP and
it now has a caller outside the route layer. Its test hook stays re-exported
from session-routes.ts.

Claude sessions only for this pass. The other CLIs name their thread in
their own config object, which this does not thread through yet. Remote and
docker sessions are skipped on purpose, because both need another host or a
container to be up and a freshly booted machine cannot promise either.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 08:05:55 +02:00