mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-06 15:39:41 +02:00
fix(sessions): act on the dual review of the reboot-restore route
Fifteen findings from two independent reviews of #442, three of them blocking. Every one is addressed here. The three blockers all sat in the restore route. A rebuild that threw after addSession left a registered session with no pane behind it, visible on the board, holding a layout slot and written to state.json, with its plan entry already spent; the catch now cleans the session up and puts the entry back. The loop checked neither the global nor the per-user session cap, so one click could take a board past a documented limit; capacity is now re-checked per iteration, because the loop is itself creating the sessions it counts. Worst of the three, a rebuilt session carried none of the state its constructor has no parameter for and then persisted itself over the record that held it, zeroing token and cost totals and dropping the pin. The pin matters most: pruning keeps a record only while it is pinned, so discarding it handed the record to the next stale sweep. A new reapplyPersistedSessionState() on the session port restores the pin, the token totals, auto-compact, auto-clear, auto-resume, nice priority, the flicker filter and the custom-model selection, and it runs before both startInteractive and the first persist. The rest, in the order they bite a user. Every rebuild failure was reported as workspace-missing, so the banner told users their repo was gone when the agent had simply failed to start; there are now distinct reasons, and the toast names each one. The client read restored and skipped off the outer response object rather than through the uniform envelope, so every count came back zero and neither toast ever fired. A board left open across the reboot never learned an offer existed, because the banner was seeded only on the page-load path; it now re-reads on every SSE init. The workspace check was existence-only, skipping the multi-user confinement that the create route applies, so a withdrawn grant would not be noticed. The banner had no phone breakpoint while its text was nowrap and its buttons could not shrink. Smaller: a missing workspace is now re-offered rather than dropped, while an already-open conversation is dropped rather than re-offered forever; a throw anywhere in the route returns the unspent entries instead of discarding the plan; the single flight is keyed by owner, since take() already stops two callers receiving one entry; the env clamp's header no longer claims a protection it cannot provide on this path today, and names the check that does bite; the three endpoints are documented in docs/api-reference.md; and the module header now says that os.uptime() reads the host's clock, so the feature is effectively off inside a container. The review also explained why the tests missed all of this: they proved the construction claim through their own copy of the construction rather than through the route, and the route tests used workspaces that did not exist, so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts mocks the Session module to drive the route's real path, and covers the cleanup, the reason reported, the re-application ordering, the broadcast and the caps. The mock route context gains the port method and the mux call the route needs. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
da933d70be
commit
fbede5cd2a
+18
-2
@@ -57,6 +57,12 @@ export interface RebootEvidence {
|
||||
*
|
||||
* This heuristic decides whether to ASK, never whether to act. A wrong yes costs
|
||||
* the user a banner they dismiss, because the restore itself waits for a click.
|
||||
*
|
||||
* ⚠️ `os.uptime()` reports the HOST's uptime, which a container shares. A Codeman
|
||||
* running in Docker therefore sees a long uptime after its own container restarts,
|
||||
* the boot test fails, and no banner appears. The feature is effectively off for
|
||||
* containerized installs. That is the safe direction to fail in, and fixing it
|
||||
* needs a boot signal the container actually owns rather than a wider heuristic.
|
||||
*/
|
||||
export function looksLikeHostReboot(evidence: RebootEvidence): boolean {
|
||||
if (evidence.deadSessionCount === 0) return false;
|
||||
@@ -80,7 +86,14 @@ export function resolveResumeConversationId(state: SessionState): string {
|
||||
return chainTail || state.resumeSessionId || state.id;
|
||||
}
|
||||
|
||||
/** Why one session was passed over. Reported for logging and assertions. */
|
||||
/**
|
||||
* Why one session was passed over. Reported for logging and shown to the user.
|
||||
*
|
||||
* The first six are decided before anything is built. `capacity-reached` and
|
||||
* `rebuild-failed` can only happen once a click is spending the plan, and they
|
||||
* are the two the banner must not confuse with a missing workspace: one means
|
||||
* "try again after closing something", the other means the CLI would not start.
|
||||
*/
|
||||
export interface RebootRestoreRejection {
|
||||
sessionId: string;
|
||||
reason:
|
||||
@@ -91,7 +104,10 @@ export interface RebootRestoreRejection {
|
||||
| 'unsupported-mode'
|
||||
| 'no-working-dir'
|
||||
| 'workspace-missing'
|
||||
| 'already-live';
|
||||
| 'workspace-forbidden'
|
||||
| 'already-live'
|
||||
| 'capacity-reached'
|
||||
| 'rebuild-failed';
|
||||
}
|
||||
|
||||
/** One restorable session, as the banner shows it and the rebuild replays it. */
|
||||
|
||||
Reference in New Issue
Block a user