Commit Graph
4 Commits
Author SHA1 Message Date
Michael GrundbergandClaude Opus 5 62ceb4e87b fix(sessions): correct what the missing-pid rule actually recognises
A second real reboot disproved the mechanism the previous commit was built
on. Typing `/exit` does not persist `pid: null`, and the session was
restored anyway.

The pid a session record carries is its `tmux attach-session` process, not
the agent. `/exit` ends the CLI inside the pane, `remain-on-exit` keeps the
pane, and the attach process stays alive throughout — so Codeman's PTY never
exits, no exit handler runs, and the record keeps both its pid and
`status: 'idle'`. The lifecycle log for the session that came back shows
created, started, stale_cleaned and recovered, with no exit event at all,
which is the proof: Codeman never learned the agent was gone.

So nothing durable distinguishes an exited agent from a session that was
idle when the power went, and this pass restores both. Ark0N/Codeman#446 is
about making Codeman notice the dead pane; contrary to what the previous
commit's message claimed, this genuinely does wait on that. Until a record
can say the agent is gone, the user dismisses or closes those sessions.

The rule itself is kept, because a record with no attach process does
describe a session that never started or whose pane died outright, and
refusing it is right. Only its documentation was wrong. The module header,
the branch comment and the test names now say what it recognises instead of
claiming the case it cannot see.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 19:49:06 +02:00
Michael GrundbergandClaude Opus 5 5108a24bf0 fix(sessions): never restore a session whose agent was already exited
Found by a real reboot, which is the first thing to catch it. Typing `/exit`
ends the CLI process and leaves the session record behind, and the
process-exit handler persists `pid: null` with `status: 'idle'` before
anything else runs. By status alone that is indistinguishable from a session
sitting idle when the power went, so the boot pass offered those sessions
back and a click spawned the agents the user had deliberately closed — the
exact case the eligibility rule exists to exclude.

The absent pid is what tells the two apart, and the plan step now refuses a
record without one, under its own `not-running` reason so the boot log says
why. On a healthy board every running session carries a pid; a record with
none describes an agent that is already gone.

Deliberately the conservative direction. A session that somehow persisted no
pid while genuinely running is not offered, and its conversation stays
reachable from the Resume list, which is where every session would be
without this feature. The opposite error spawns processes nobody asked for.

Ark0N/Codeman#446 covers the dead panes those exits leave behind, but this
does not wait on it: the rule belongs here whether or not the record's shape
changes later.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 18:46:15 +02:00
Michael GrundbergandClaude Opus 5 39976041e0 fix(sessions): let a dismiss reach the entries a restore is holding
Fourth review of the reboot-restore branch, and the third to find a defect
in the previous round's fix. This one is the same shape as its predecessor:
a counter keyed on one thing, compared against a set keyed on another.

The generation counter was indexed by the entry's owner, while the in-flight
set holds the caller doing the restoring. Those are the same person exactly
when a user restores their own sessions, which is every case the tests
covered. The route deliberately supports the other case: an admin may spend
another user's entries. So when an admin restored Bob's sessions and Bob
dismissed the banner, nothing matched, the entries came back, and a plan Bob
had explicitly dismissed was re-armed for another twenty-four hours.

Rather than reconcile the two key spaces, the counter is gone. `take()` now
parks the entries it hands out, remembering which caller is spending them,
and they stay parked until that restore ends. A dismiss filters the parked
entries by `canAccess(entry.owner)` — the same predicate it already applies
to the plan — so it reaches them wherever they are. `releaseFlight()` puts
back only what is still parked. Expiry and a fresh boot plan unpark
everything, for the same reason. There is one key space now, the entry's
owner, and the spender is only ever used to tell two concurrent flights
apart. That removes `generations`, `snapshotGenerations()`, `bump()`,
`bumpAll()` and the argument threaded through the route.

The discard grew the teardown it still lacked. A rebuild can fail after
startInteractive() resolved, and a restored workspace still carries
Codeman's hooks, so the CLI can post a hook event within milliseconds; the
transcript watcher that starts from it, the attachment registry, the wait
registry and the approvals inbox all outlive the listeners and would meet
the retry, which reuses the session id by design. Its steps also run in
reverse order now, so no live listener can reach a tracker that has already
stopped, and the mux kill has its own guard, because stop() kills the pane
in its last block after destroying four trackers.

Tests. The run-summary test named an interval and asserted a map entry, so
dropping stop() left it green; it now spies on stop(). Nothing pinned that
before-spawn must precede setupSessionListeners, which reads the flag that
phase restores, so swapping the two lines was silent; the ordering test now
includes the listener setup. The retry assertion was a tautology and now
asserts a different refs object. Both strengthened tests were verified by
reverting their fix. Two new tests cover the admin-restores-another-owner
cases this round was about. The server in the discard test is built once and
stopped, since its constructor registers handlers on module-level watchers,
and the workspace is removed through safeRmHomeTree.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 16:51:00 +02:00
Michael GrundbergandClaude Opus 5 da933d70be feat(sessions): offer to rebuild the sessions a host reboot destroyed
A host reboot takes the tmux server down with it, so every pane dies,
reconciliation finds nothing to attach to, and the board comes up empty.
Picking yesterday's work back up meant finding each conversation in history
and resuming it by hand, one at a time.

The boot pass now works out what the reboot killed and leaves it on offer.
It runs inside restoreMuxSessions(), in the window where reconciliation has
reported the dead sessions and cleanupStaleSessions() has not pruned their
records yet, which is the only place the records can still be read. The
board shows a banner, and nothing is created until the user clicks it.

A click rather than an automatic restore is what makes the reboot heuristic
acceptable. The heuristic cannot tell a reboot from a crash that took tmux
down inside the same window, so it decides whether to ASK, never whether to
act: a wrong yes costs a line of text the user dismisses instead of N CLI
processes nobody asked for.

Four things are re-checked when the click arrives rather than trusted from
boot, because hours can pass and the board moves on. The owner's privilege
grant re-resolves through the env clamp. The workspace must still be on
disk. A conversation the user already resumed by hand from the Resume list
is skipped, since two panes running --resume on one conversation would
fight over the same transcript. Entries leave the plan synchronously before
the first await, and the route is single-flighted, so a double-click or two
devices cannot both reach the same entry.

A restored session comes back attached, idle and disarmed. Respawn
controllers and Ralph loops are deliberately not re-armed: a machine that
just came up is the worst moment to turn an autonomous run loose. Its
workspace hooks are installed by the restore route itself, because the
boot-time sweep sits behind a gate that is false after a reboot and has
finished long before the click; without them a session goes silently blind,
with no stop or idle events for respawn, no Approvals Inbox item and no red
tab on a blocking dialog. Stats collection starts the same way.

The pane is new, so the conversation continues and the terminal scrollback
does not. The banner says so rather than letting an empty pane read as a
broken restore.

The plan lives in memory only. A server restart drops it, which costs the
convenience this adds and never the conversation: the conversation is the
transcript under ~/.claude/projects, which the Welcome screen's Resume list
and the Session Manager already read, so a dropped plan returns the user to
resuming by hand.

clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the
question it answers is about session privilege rather than about HTTP and
it now has a caller outside the route layer. Its test hook stays re-exported
from session-routes.ts.

Claude sessions only for this pass. The other CLIs name their thread in
their own config object, which this does not thread through yet. Remote and
docker sessions are skipped on purpose, because both need another host or a
container to be up and a freshly booted machine cannot promise either.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 08:05:55 +02:00