fix(reboot-restore): the merge-time items from the #442 review

Seven things, none of which changes what the feature does.

1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred
   it from the name: a session the user renamed by hand to something shaped
   like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the
   next prompt overwrote their name. The route persists right after, so the
   loss went to disk. `restoreMuxSessions()` already passes it.

2. The already-live sets were snapshotted once before a loop that awaits a
   real `startInteractive()` per entry, so by the tenth entry the snapshot
   was tens of seconds old and a conversation resumed by hand from the
   Resume list in that window was invisible to it: two panes on one
   transcript, the exact thing the check exists to prevent. Both sets are
   now read per iteration, and the late case is spent rather than re-offered
   for the same reason the batch case is.

3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path.
   The stamp predates the reboot and the pane is new, so honouring it meant
   one click had every restored session type `continue` into itself about a
   minute later, unattended, against the route header's own promise that a
   restored session comes back idle and disarmed. The setting stays ENABLED,
   so it re-arms on the next real limit message. A Codeman restart still
   re-arms from the stamp, because the limit footer will not reprint on its
   own; the new option exists only to tell the two paths apart.

4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()`
   and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession`
   performs that it was missing. Cosmetic, but a run left open reads as
   still going in the away digest.

5. A restored claude session gets `seedAgentSessionPreamble()` like both
   create paths, so the agent skill's bootstrap stays a two-line loader.

6. The heuristic's container comment was wrong in one direction and quiet
   about the real gap: after a genuine host reboot a containerized Codeman
   sees the host's short uptime and the banner does appear. What it cannot
   see is a container-only restart, which is where this would help most.

7. The banner is hidden in a solo window, which shows one session and has
   no tab strip to put restored ones in.

Also reverts 17 of the 18 hunks in docs/api-reference.md, which were
Prettier reformatting of prose the PR does not otherwise touch (docs/ is
outside the format glob), keeping only the Reboot restore section and
repairing the two continuation lines that reformat de-indented; renumbers
reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which
loads after it; and gives the feature its CLAUDE.md entry plus a route
test for the multi-user workspace-forbidden branch, the only new rule that
had nothing behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-09-18 13:46:04 +02:00
parent ea5323d990
commit bb8ada7e5f
9 changed files with 228 additions and 109 deletions
+6 -4
View File
File diff suppressed because one or more lines are too long
+80 -88
View File
@@ -66,17 +66,17 @@ The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on `src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
the HTTP status. the HTTP status.
| `errorCode` | HTTP | Meaning | | `errorCode` | HTTP | Meaning |
| ------------------ | ---- | --------------------------------------------------- | |-------------|------|---------|
| `INVALID_INPUT` | 400 | Malformed request / failed validation | | `INVALID_INPUT` | 400 | Malformed request / failed validation |
| `UNAUTHORIZED` | 401 | Authentication required or failed | | `UNAUTHORIZED` | 401 | Authentication required or failed |
| `NOT_FOUND` | 404 | Resource does not exist | | `NOT_FOUND` | 404 | Resource does not exist |
| `SESSION_BUSY` | 409 | Session is busy | | `SESSION_BUSY` | 409 | Session is busy |
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) | | `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
| `ALREADY_EXISTS` | 409 | Resource already exists | | `ALREADY_EXISTS` | 409 | Resource already exists |
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed | | `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
| `RATE_LIMITED` | 429 | Too many requests | | `RATE_LIMITED` | 429 | Too many requests |
| `INTERNAL_ERROR` | 500 | Unexpected server error | | `INTERNAL_ERROR` | 500 | Unexpected server error |
Adding a new error code is non-breaking; removing or renaming one is a major change. Adding a new error code is non-breaking; removing or renaming one is a major change.
@@ -87,10 +87,10 @@ exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse driving the API from a shell tool cannot practically hold a stream and parse
events inline. events inline.
| Call | Blocks until | | Call | Blocks until |
| --------------------------------------------- | -------------------------------------------------- | |------|--------------|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires | | `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output | | `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires | | `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate `POST .../input` with `wait` is not the same as a `POST` followed by a separate
@@ -140,13 +140,13 @@ contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
### Signals ### Signals
| Signal | Source | Actually fires for | | Signal | Source | Actually fires for |
| --------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | |--------|--------|--------------------|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) | | `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) | | `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only | | `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below | | `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode | | `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic `stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until` fallback that can flap mid-turn when a spinner pauses. The default set when `until`
@@ -156,12 +156,12 @@ can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit
once the session is up: the default set's `idle` also resolves on a spinner pause, once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up) and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured: can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no _further_ `idle`, so it is the a session parked on the trust dialog emits no *further* `idle`, so it is the
startup transition, not the dialog, that produces the false success below. startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The ⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died _inside_ its tmux pane covers a session that exited — including a worker that died *inside* its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
@@ -184,7 +184,7 @@ blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell ⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever, session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a _transition_. Since send-and-wait whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there: requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
@@ -218,11 +218,11 @@ with `from=buffer` keeps matching long after the dialog is gone. A worked versio
### `GET /api/v1/sessions/:id/wait` ### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes | | Param | Type | Default | Notes |
| --------- | -------------------------------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | |-------|------|---------|-------|
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback | | `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` | | `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time | | `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
```bash ```bash
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
@@ -239,12 +239,12 @@ a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output` ### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes | | Param | Type | Default | Notes |
| --------- | ------------------------------- | -------- | ----------------------------------------------------------------------------------------------------------- | |-------|------|---------|-------|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found | | `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing | | `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking | | `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` | | `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a **Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately `400` rather than ignored, so a caller that assumed otherwise finds out immediately
@@ -296,10 +296,10 @@ hand-written query string decodes to a space.
Two optional fields on the existing endpoint: Two optional fields on the existing endpoint:
| Field | Type | Notes | | Field | Type | Notes |
| ------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | |-------|------|-------|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" | | `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` | | `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would "absent" rather than failing validation. That is deliberate: `.optional()` would
@@ -330,24 +330,16 @@ All three nest the wait result under `data.wait`, so one client helper works aga
any of them: any of them:
```json ```json
{ { "success": true, "data": {
"success": true, "sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"data": { "status": "idle",
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464", "limitPaused": false,
"status": "idle", "wait": {
"limitPaused": false, "signal": "stop", "until": ["stop", "idle", "exit"],
"wait": { "timedOut": false, "immediate": false, "ended": false, "aborted": false,
"signal": "stop", "waitedMs": 8421, "timeoutMs": 60000
"until": ["stop", "idle", "exit"],
"timedOut": false,
"immediate": false,
"ended": false,
"aborted": false,
"waitedMs": 8421,
"timeoutMs": 60000
}
} }
} }}
``` ```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`, `POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
@@ -361,21 +353,21 @@ redelivery (harmless, the turn it refers to may be long over), while with
client that reads `delivered === false` as "duplicate" silently treats a failed send client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success. as a success.
| Field | Type | Meaning | | Field | Type | Meaning |
| ---------------- | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | |-------|------|---------|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) | | `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) | | `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) | | `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) | | `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) | | `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` | | `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) | | `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met | | `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome | | `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting | | `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied | | `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand | | `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard | | `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order: Read the outcome by discriminator, in this order:
@@ -398,12 +390,12 @@ read the timeout as "the worker is wedged" and kill a session that was working f
### Errors ### Errors
| `errorCode` | HTTP | When | | `errorCode` | HTTP | When |
| --------------- | ---- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | |-------------|------|------|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` | | `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own | | `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full | | `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem | | `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The `SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
@@ -454,9 +446,9 @@ Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first, - `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId, ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?, sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
toolSummary?, message?, cwd?, context?, options?: {n, label}[], toolSummary?, message?, cwd?, context?, options?: {n, label}[],
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame; acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
`options` is present only when the dialog's numbered choices parsed `options` is present only when the dialog's numbered choices parsed
confidently; `acknowledgedAt` marks an item a human has already looked at confidently; `acknowledgedAt` marks an item a human has already looked at
(see `/viewed` below) and tells clients not to re-arm its tab alert. Listing (see `/viewed` below) and tells clients not to re-arm its tab alert. Listing
@@ -474,7 +466,7 @@ acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
first, `422 OPERATION_FAILED` when the session refused input. first, `422 OPERATION_FAILED` when the session refused input.
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes. - `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
- `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId, - `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId,
acknowledged: itemId | null }`. Marks the session's pending **idle** item as acknowledged: itemId | null }`. Marks the session's pending **idle** item as
seen by a human (the web UI calls it when you open the session's tab): the seen by a human (the web UI calls it when you open the session's tab): the
item stays pending and answerable, but stops arming the yellow tab alert on item stays pending and answerable, but stops arming the yellow tab alert on
every client, including after a reload. Permission/question items are never every client, including after a reload. Permission/question items are never
@@ -498,19 +490,19 @@ heuristic decides whether to ASK, never whether to act.
Claude-mode sessions only (others carry their conversation id in their own Claude-mode sessions only (others carry their conversation id in their own
config object); remote and docker sessions are never offered, because both need config object); remote and docker sessions are never offered, because both need
another host or container to be up. The plan is in-memory, so a server restart another host or container to be up. The plan is in-memory, so a server restart
drops it and the offer is gone — the conversations themselves are unaffected, drops it and the offer is gone; the conversations themselves are unaffected,
since they live in the CLI's own transcript store and stay reachable from the since they live in the CLI's own transcript store and stay reachable from the
Resume list. A plan nobody spends expires after 24 hours. Resume list. A plan nobody spends expires after 24 hours.
- `GET /api/v1/reboot-restore` → `{ sessions: RestorableSession[], - `GET /api/v1/reboot-restore` → `{ sessions: RestorableSession[],
scrollbackRestored: false }`, ownership-scoped in multi-user mode. scrollbackRestored: false }`, ownership-scoped in multi-user mode.
`RestorableSession`: `{ id, name?, workingDir, mode, owner? }`. The persisted `RestorableSession`: `{ id, name?, workingDir, mode, owner? }`. The persisted
record itself is never sent. `scrollbackRestored` is always `false` and exists record itself is never sent. `scrollbackRestored` is always `false` and exists
so a client states it: a restored session is a NEW pane, so the conversation so a client states it: a restored session is a NEW pane, so the conversation
continues and the terminal history does not. continues and the terminal history does not.
- `POST /api/v1/reboot-restore/restore` with `{ sessionIds?: string[] }` (omit - `POST /api/v1/reboot-restore/restore` with `{ sessionIds?: string[] }` (omit
to restore everything the caller can see) → `{ restored: RestorableSession[], to restore everything the caller can see) → `{ restored: RestorableSession[],
skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing` skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing`
(the directory is gone), `workspace-forbidden` (in multi-user mode it is (the directory is gone), `workspace-forbidden` (in multi-user mode it is
outside the workspace of the user the session belongs to, re-checked against outside the workspace of the user the session belongs to, re-checked against
that owner's current grant rather than the caller's), `already-live` (the conversation is already that owner's current grant rather than the caller's), `already-live` (the conversation is already
@@ -521,7 +513,7 @@ skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing`
removed from the plan before any pane is built, so a double-click cannot put removed from the plan before any pane is built, so a double-click cannot put
two panes on one conversation; anything that never became a pane goes back on two panes on one conversation; anything that never became a pane goes back on
offer, except `already-live`, which cannot stop being true. A restored session offer, except `already-live`, which cannot stop being true. A restored session
comes back attached, idle and disarmed — respawn controllers and Ralph loops comes back attached, idle and disarmed: respawn controllers and Ralph loops
are never re-armed automatically. are never re-armed automatically.
- `POST /api/v1/reboot-restore/dismiss` → `{ dismissed: n }`. Drops the offer - `POST /api/v1/reboot-restore/dismiss` → `{ dismissed: n }`. Drops the offer
for everything the caller can see. for everything the caller can see.
@@ -541,7 +533,7 @@ user guide: [`readmymind.md`](readmymind.md).
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the - `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals, session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
50, each <= 500 chars). A case with nothing recorded answers an empty 50, each <= 500 chars). A case with nothing recorded answers an empty
profile with `updatedAt: 0`; nothing is persisted by reads. profile with `updatedAt: 0`; nothing is persisted by reads.
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict - `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
@@ -574,7 +566,7 @@ same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synce
[`claude-voice-plan.md`](claude-voice-plan.md). [`claude-voice-plan.md`](claude-voice-plan.md).
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?, - `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
signed in to Claude Code on the server), `expired` (the access token elapsed; signed in to Claude Code on the server), `expired` (the access token elapsed;
running any Claude session refreshes it) or `malformed`. The OAuth token running any Claude session refreshes it) or `malformed`. The OAuth token
itself is never returned by this or any other endpoint. itself is never returned by this or any other endpoint.
+11 -5
View File
@@ -72,11 +72,17 @@ export interface RebootEvidence {
* This heuristic decides whether to ASK, never whether to act. A wrong yes costs * This heuristic decides whether to ASK, never whether to act. A wrong yes costs
* the user a banner they dismiss, because the restore itself waits for a click. * the user a banner they dismiss, because the restore itself waits for a click.
* *
* ⚠️ `os.uptime()` reports the HOST's uptime, which a container shares. A Codeman * ⚠️ `os.uptime()` reports the HOST's uptime, which a container shares, and that
* running in Docker therefore sees a long uptime after its own container restarts, * cuts BOTH ways rather than simply switching the feature off in Docker. After a
* the boot test fails, and no banner appears. The feature is effectively off for * genuine host reboot a containerized Codeman sees the host's short uptime, so the
* containerized installs. That is the safe direction to fail in, and fixing it * banner DOES appear and the feature works. What it cannot see is a container-only
* needs a boot signal the container actually owns rather than a wider heuristic. * restart: the host uptime is long, the boot test fails, and no banner appears
* although every in-container pane is gone (`docker/server.Dockerfile` installs
* tmux inside the Codeman container, and the self-updater restarts the Compose
* deployment by exiting the container, so that is the case where this would help
* most). Failing quiet is the safe direction, and closing the gap needs a boot
* signal the container owns (PID 1's start time, gated on the existing
* `isRunningInContainer()`) rather than a wider heuristic.
*/ */
export function looksLikeHostReboot(evidence: RebootEvidence): boolean { export function looksLikeHostReboot(evidence: RebootEvidence): boolean {
if (evidence.deadSessionCount === 0) return false; if (evidence.deadSessionCount === 0) return false;
+13 -1
View File
@@ -27,7 +27,19 @@ export interface SessionPort {
reapplyPersistedSessionState( reapplyPersistedSessionState(
session: Session, session: Session,
saved: SessionState, saved: SessionState,
phase: 'before-spawn' | 'after-spawn' phase: 'before-spawn' | 'after-spawn',
options?: {
/**
* Re-arm a PENDING auto-resume schedule from the record's `autoResumeAt`.
* Default true, which is what a Codeman restart wants: the limit footer
* will not reprint on its own, so dropping the stamp there strands the
* pause. A reboot restore passes false: the stamp predates the reboot,
* the pane is new, and re-arming means every restored session types
* `continue` into itself about a minute after one click. Auto-resume
* stays ENABLED either way, so it re-arms on fresh evidence.
*/
rearmAutoResumeSchedule?: boolean;
}
): Promise<void>; ): Promise<void>;
/** /**
* Undo a session that was registered but never got a working pane: the map * Undo a session that was registered but never got a working pane: the map
+1 -1
View File
@@ -26,7 +26,7 @@
* @mixin Extends CodemanApp.prototype via Object.assign * @mixin Extends CodemanApp.prototype via Object.assign
* @dependency app.js (CodemanApp class, showToast) * @dependency app.js (CodemanApp class, showToast)
* @dependency api-client.js at runtime (this._api / this._apiJson) * @dependency api-client.js at runtime (this._api / this._apiJson)
* @loadorder 11.7 of 17, after approvals-ui.js * @loadorder 11.65, after approvals-ui.js and before admin-ui.js (11.7)
*/ */
/** Plain-language wording for one skip reason, for the toast after a restore. */ /** Plain-language wording for one skip reason, for the toast after a restore. */
+4
View File
@@ -2548,6 +2548,10 @@ body.solo-mode .header-tokens,
body.solo-mode .btn-notifications, body.solo-mode .btn-notifications,
body.solo-mode .btn-multimonitor, body.solo-mode .btn-multimonitor,
body.solo-mode .header-plan-usage, body.solo-mode .header-plan-usage,
/* A solo window shows ONE session and has no tab strip to put restored ones in,
so offering to rebuild a list of them there is an offer it cannot show the
result of. The dashboard that spawned this window carries the banner. */
body.solo-mode .reboot-restore-banner,
body.solo-mode .btn-lifecycle-log { body.solo-mode .btn-lifecycle-log {
display: none !important; display: none !important;
} }
+47 -7
View File
@@ -41,7 +41,7 @@ import { clampEnvOverridesForOwner } from '../../session-env-clamp.js';
import { Session } from '../../session.js'; import { Session } from '../../session.js';
import { resolveClaudeModeForUsername } from '../../user-store.js'; import { resolveClaudeModeForUsername } from '../../user-store.js';
import { getCli } from '../../config/cli-registry/registry.js'; import { getCli } from '../../config/cli-registry/registry.js';
import { applyWorkspaceHooks } from '../../hooks-config.js'; import { applyWorkspaceHooks, seedAgentSessionPreamble } from '../../hooks-config.js';
import { getLifecycleLog } from '../../session-lifecycle-log.js'; import { getLifecycleLog } from '../../session-lifecycle-log.js';
import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js'; import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js';
import { SseEvent } from '../sse-events.js'; import { SseEvent } from '../sse-events.js';
@@ -105,11 +105,16 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
// the user resumed by hand from the Resume list is already on screen, and a // the user resumed by hand from the Resume list is already on screen, and a
// second pane on it would fight the first for the same transcript. This one // second pane on it would fight the first for the same transcript. This one
// is never re-offered: unlike a missing workspace, it cannot stop being true. // is never re-offered: unlike a missing workspace, it cannot stop being true.
const liveSessionIds = new Set(ctx.sessions.keys()); // Read fresh each time rather than snapshotted once: the loop below awaits a
const liveConversationIds = new Set( // real `startInteractive()` per entry, so by the tenth entry a snapshot taken
[...ctx.sessions.values()].map((session) => session.claudeSessionId).filter((id): id is string => !!id) // here is tens of seconds old, and a conversation the user resumed by hand in
); // that window would be invisible to it.
const { restore, skipped } = rejectAlreadyLive(taken, liveSessionIds, liveConversationIds); const liveSessionIds = () => new Set(ctx.sessions.keys());
const liveConversationIds = () =>
new Set(
[...ctx.sessions.values()].map((session) => session.claudeSessionId).filter((id): id is string => !!id)
);
const { restore, skipped } = rejectAlreadyLive(taken, liveSessionIds(), liveConversationIds());
for (const entry of taken) { for (const entry of taken) {
if (skipped.some((s) => s.sessionId === entry.sessionId)) unspent.delete(entry); if (skipped.some((s) => s.sessionId === entry.sessionId)) unspent.delete(entry);
} }
@@ -119,6 +124,18 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
const workspaceHooksEnabled = await ctx.getWorkspaceHooksEnabled(); const workspaceHooksEnabled = await ctx.getWorkspaceHooksEnabled();
for (const entry of restore) { for (const entry of restore) {
// The already-live check, re-run against the board as it is NOW. The pass
// above decided the batch; this catches a conversation that went live while
// an earlier entry in this same batch was starting. Spent rather than
// returned to the plan, for the same reason as the batch pass: unlike a
// missing workspace or a withdrawn grant, an open conversation is not a
// condition that stops being true.
const [lateLive] = rejectAlreadyLive([entry], liveSessionIds(), liveConversationIds()).skipped;
if (lateLive) {
failures.push(lateLive);
unspent.delete(entry);
continue;
}
// Capacity is re-checked per iteration, because this loop is itself // Capacity is re-checked per iteration, because this loop is itself
// creating the sessions it counts. The offer can be a day old, so the // creating the sessions it counts. The offer can be a day old, so the
// board may be fuller now than the plan assumed. // board may be fuller now than the plan assumed.
@@ -157,6 +174,12 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
workingDir: saved.workingDir, workingDir: saved.workingDir,
mode: saved.mode, mode: saved.mode,
name: saved.name, name: saved.name,
// Without this the constructor re-infers ownership from the name, so a
// session the user renamed by hand to something shaped like `w<n>-<case>`
// comes back as `placeholder` and auto-naming overwrites their name on
// the next prompt. The route persists below, so the loss would go to
// disk. `restoreMuxSessions()` passes it for the same reason.
nameSource: saved.nameSource,
createdAt: saved.createdAt, createdAt: saved.createdAt,
mux: ctx.mux, mux: ctx.mux,
useMux: true, useMux: true,
@@ -197,7 +220,14 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
// the reduced one and drop the pin that keeps it from being pruned. A // the reduced one and drop the pin that keeps it from being pruned. A
// listener-driven persist can still land inside the debounce window // listener-driven persist can still land inside the debounce window
// while the pane starts; the write below repairs the record. // while the pane starts; the write below repairs the record.
await ctx.reapplyPersistedSessionState(session, saved, 'after-spawn'); // `rearmAutoResumeSchedule: false`: the saved stamp predates the reboot and
// the pane is new, so honouring it would have every restored session type
// `continue` into itself about a minute after one click. Auto-resume stays
// enabled and re-arms on the next real limit message. This is also what the
// module header promises ("comes back attached, idle and disarmed").
await ctx.reapplyPersistedSessionState(session, saved, 'after-spawn', {
rearmAutoResumeSchedule: false,
});
ctx.persistSessionState(session); ctx.persistSessionState(session);
// A session without its workspace hooks goes silently blind: no stop or // A session without its workspace hooks goes silently blind: no stop or
@@ -211,6 +241,16 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
); );
} }
// Both create paths seed this; without it a restored claude session's agent
// skill falls back to writing out the whole ~150-line §0 preamble. Remote and
// docker sessions never reach here (the plan rejects them as
// `remote-or-docker`), so the local-only condition is structural.
if (getCli(session.mode)?.capabilities.agentSkillInjection && (await ctx.getAgentSkillEnabled())) {
await seedAgentSessionPreamble(session.id).catch((err: unknown) =>
console.warn(`[agent-skill] preamble seed failed for ${session.id}: ${getErrorMessage(err)}`)
);
}
getLifecycleLog().log({ event: 'recovered', sessionId: session.id, name: session.name }); getLifecycleLog().log({ event: 'recovered', sessionId: session.id, name: session.name });
// Every other open tab and phone needs this; the clicking tab already has // Every other open tab and phone needs this; the clicking tab already has
// the response, and the client's handler is an idempotent upsert. // the response, and the client's handler is an idempotent upsert.
+19 -2
View File
@@ -2940,7 +2940,8 @@ export class WebServer extends EventEmitter {
async reapplyPersistedSessionState( async reapplyPersistedSessionState(
session: Session, session: Session,
saved: SessionState, saved: SessionState,
phase: 'before-spawn' | 'after-spawn' phase: 'before-spawn' | 'after-spawn',
options?: { rearmAutoResumeSchedule?: boolean }
): Promise<void> { ): Promise<void> {
if (phase === 'before-spawn') { if (phase === 'before-spawn') {
// The custom-model env has to be rebuilt from the endpoint store: the persist // The custom-model env has to be rebuilt from the endpoint store: the persist
@@ -2968,7 +2969,14 @@ export class WebServer extends EventEmitter {
session.setAutoClear(saved.autoClearEnabled ?? false, saved.autoClearThreshold); session.setAutoClear(saved.autoClearEnabled ?? false, saved.autoClearThreshold);
} }
if (saved.autoResumeEnabled) { if (saved.autoResumeEnabled) {
session.restoreAutoResume(true, saved.autoResumeAt); // The stamp is re-armed by default, because a Codeman restart leaves the
// limit footer un-reprinted and dropping it there would strand the pause.
// A reboot restore opts out: that stamp predates the reboot, the pane is
// new, and honouring it means every session the user restored types
// `continue` into itself about a minute later, unattended. The setting
// itself stays on either way, so it re-arms on the next limit message.
const rearm = options?.rearmAutoResumeSchedule !== false;
session.restoreAutoResume(true, rearm ? saved.autoResumeAt : undefined);
} }
if (saved.inputTokens !== undefined || saved.outputTokens !== undefined || saved.totalCost !== undefined) { if (saved.inputTokens !== undefined || saved.outputTokens !== undefined || saved.totalCost !== undefined) {
session.restoreTokens(saved.inputTokens ?? 0, saved.outputTokens ?? 0, saved.totalCost ?? 0); session.restoreTokens(saved.inputTokens ?? 0, saved.outputTokens ?? 0, saved.totalCost ?? 0);
@@ -3024,9 +3032,18 @@ export class WebServer extends EventEmitter {
session.ralphTracker.stopWatchingFixPlan(); session.ralphTracker.stopWatchingFixPlan();
const summaryTracker = this.runSummaryTrackers.get(sessionId); const summaryTracker = this.runSummaryTrackers.get(sessionId);
if (summaryTracker) { if (summaryTracker) {
// Closes the run's own record before the tracker goes, the way
// `_doCleanupSession()` does. Cosmetic rather than load-bearing, but a
// run left open reads as still going in the away digest.
summaryTracker.recordSessionStopped();
summaryTracker.stop(); summaryTracker.stop();
this.runSummaryTrackers.delete(sessionId); this.runSummaryTrackers.delete(sessionId);
} }
// Also mirrors `_doCleanupSession()`. The PERSISTED Ralph state is left
// alone on purpose (that is one of the things separating this from
// cleanupSession); this only clears the in-memory tracker the failed
// construction built, which the retry reuses the id of.
session.ralphTracker.fullReset();
// --- what anything else may have attached to this id in the meantime --- // --- what anything else may have attached to this id in the meantime ---
// A rebuild can fail AFTER startInteractive() resolved, and a restored // A rebuild can fail AFTER startInteractive() resolved, and a restored
+47 -1
View File
@@ -11,7 +11,10 @@
* The routes read the process-wide `rebootRestoreRegistry` singleton, so every * The routes read the process-wide `rebootRestoreRegistry` singleton, so every
* test resets it; a leaked entry would bleed into the next one. * test resets it; a leaked entry would bleed into the next one.
*/ */
import { describe, it, expect, afterEach } from 'vitest'; import { describe, it, expect, afterEach, beforeEach } from 'vitest';
import { mkdtempSync, rmSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import Fastify, { type FastifyInstance } from 'fastify'; import Fastify, { type FastifyInstance } from 'fastify';
import fastifyCookie from '@fastify/cookie'; import fastifyCookie from '@fastify/cookie';
import { registerRebootRestoreRoutes } from '../../src/web/routes/reboot-restore-routes.js'; import { registerRebootRestoreRoutes } from '../../src/web/routes/reboot-restore-routes.js';
@@ -187,6 +190,49 @@ describe('POST /api/reboot-restore/restore', () => {
}); });
}); });
describe('POST /api/reboot-restore/restore: multi-user workspace confinement', () => {
const saved: Record<string, string | undefined> = {};
let realDir: string;
beforeEach(() => {
saved.CODEMAN_MULTIUSER = process.env.CODEMAN_MULTIUSER;
process.env.CODEMAN_MULTIUSER = '1';
// This branch sits AFTER the existsSync check, so the workspace has to be
// real for the confinement rule to be the thing that rejects the entry.
realDir = mkdtempSync(join(tmpdir(), 'codeman-reboot-restore-real-'));
});
afterEach(() => {
if (saved.CODEMAN_MULTIUSER === undefined) delete process.env.CODEMAN_MULTIUSER;
else process.env.CODEMAN_MULTIUSER = saved.CODEMAN_MULTIUSER;
rmSync(realDir, { recursive: true, force: true });
});
it("refuses a workspace outside the OWNER's case space, and leaves it on offer", async () => {
const entry = offerEntry('a', 'alice');
entry.workingDir = realDir;
(entry.state as { workingDir: string }).workingDir = realDir;
rebootRestoreRegistry.set([entry]);
// An admin does the clicking. The confinement is still resolved against
// alice, the entry's OWNER: `isWorkingDirAllowed` waves an admin through, so
// reading the caller here would hand an admin the power to rebuild another
// user's session anywhere on the box.
const app = await createHarness({ username: 'root-user', role: 'admin' });
const res = await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
expect(res.statusCode).toBe(200);
expect(res.json().data.restored).toEqual([]);
expect(res.json().data.skipped).toEqual([{ sessionId: 'a', reason: 'workspace-forbidden' }]);
// A withdrawn grant can be given back, so unlike `already-live` this is not
// the permanent kind of refusal and the entry stays claimable.
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
expect(left.sessions.map((s: { id: string }) => s.id)).toEqual(['a']);
await app.close();
});
});
describe('POST /api/reboot-restore/dismiss', () => { describe('POST /api/reboot-restore/dismiss', () => {
it('drops the offer and leaves the banner with nothing to show', async () => { it('drops the offer and leaves the banner with nothing to show', async () => {
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]); rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);