mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does. 1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred it from the name: a session the user renamed by hand to something shaped like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the next prompt overwrote their name. The route persists right after, so the loss went to disk. `restoreMuxSessions()` already passes it. 2. The already-live sets were snapshotted once before a loop that awaits a real `startInteractive()` per entry, so by the tenth entry the snapshot was tens of seconds old and a conversation resumed by hand from the Resume list in that window was invisible to it: two panes on one transcript, the exact thing the check exists to prevent. Both sets are now read per iteration, and the late case is spent rather than re-offered for the same reason the batch case is. 3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path. The stamp predates the reboot and the pane is new, so honouring it meant one click had every restored session type `continue` into itself about a minute later, unattended, against the route header's own promise that a restored session comes back idle and disarmed. The setting stays ENABLED, so it re-arms on the next real limit message. A Codeman restart still re-arms from the stamp, because the limit footer will not reprint on its own; the new option exists only to tell the two paths apart. 4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()` and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession` performs that it was missing. Cosmetic, but a run left open reads as still going in the away digest. 5. A restored claude session gets `seedAgentSessionPreamble()` like both create paths, so the agent skill's bootstrap stays a two-line loader. 6. The heuristic's container comment was wrong in one direction and quiet about the real gap: after a genuine host reboot a containerized Codeman sees the host's short uptime and the banner does appear. What it cannot see is a container-only restart, which is where this would help most. 7. The banner is hidden in a solo window, which shows one session and has no tab strip to put restored ones in. Also reverts 17 of the 18 hunks in docs/api-reference.md, which were Prettier reformatting of prose the PR does not otherwise touch (docs/ is outside the format glob), keeping only the Reboot restore section and repairing the two continuation lines that reformat de-indented; renumbers reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which loads after it; and gives the feature its CLAUDE.md entry plus a route test for the multi-user workspace-forbidden branch, the only new rule that had nothing behind it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+80
-88
@@ -66,17 +66,17 @@ The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
|
||||
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
|
||||
the HTTP status.
|
||||
|
||||
| `errorCode` | HTTP | Meaning |
|
||||
| ------------------ | ---- | --------------------------------------------------- |
|
||||
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
|
||||
| `UNAUTHORIZED` | 401 | Authentication required or failed |
|
||||
| `NOT_FOUND` | 404 | Resource does not exist |
|
||||
| `SESSION_BUSY` | 409 | Session is busy |
|
||||
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
|
||||
| `ALREADY_EXISTS` | 409 | Resource already exists |
|
||||
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
|
||||
| `RATE_LIMITED` | 429 | Too many requests |
|
||||
| `INTERNAL_ERROR` | 500 | Unexpected server error |
|
||||
| `errorCode` | HTTP | Meaning |
|
||||
|-------------|------|---------|
|
||||
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
|
||||
| `UNAUTHORIZED` | 401 | Authentication required or failed |
|
||||
| `NOT_FOUND` | 404 | Resource does not exist |
|
||||
| `SESSION_BUSY` | 409 | Session is busy |
|
||||
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
|
||||
| `ALREADY_EXISTS` | 409 | Resource already exists |
|
||||
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
|
||||
| `RATE_LIMITED` | 429 | Too many requests |
|
||||
| `INTERNAL_ERROR` | 500 | Unexpected server error |
|
||||
|
||||
Adding a new error code is non-breaking; removing or renaming one is a major change.
|
||||
|
||||
@@ -87,10 +87,10 @@ exist because SSE is Codeman's only other "tell me when" channel, and an agent
|
||||
driving the API from a shell tool cannot practically hold a stream and parse
|
||||
events inline.
|
||||
|
||||
| Call | Blocks until |
|
||||
| --------------------------------------------- | -------------------------------------------------- |
|
||||
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
|
||||
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
|
||||
| Call | Blocks until |
|
||||
|------|--------------|
|
||||
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
|
||||
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
|
||||
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
|
||||
|
||||
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
|
||||
@@ -140,13 +140,13 @@ contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
|
||||
|
||||
### Signals
|
||||
|
||||
| Signal | Source | Actually fires for |
|
||||
| --------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
|
||||
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
|
||||
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
|
||||
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
|
||||
| `exit` | no process is behind the session | every mode |
|
||||
| Signal | Source | Actually fires for |
|
||||
|--------|--------|--------------------|
|
||||
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
|
||||
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
|
||||
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
|
||||
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
|
||||
| `exit` | no process is behind the session | every mode |
|
||||
|
||||
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
|
||||
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
|
||||
@@ -156,12 +156,12 @@ can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit
|
||||
once the session is up: the default set's `idle` also resolves on a spinner pause,
|
||||
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
|
||||
can land inside your first wait window and report a turn that never ran. Measured:
|
||||
a session parked on the trust dialog emits no _further_ `idle`, so it is the
|
||||
a session parked on the trust dialog emits no *further* `idle`, so it is the
|
||||
startup transition, not the dialog, that produces the false success below.
|
||||
|
||||
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
|
||||
server answers from `pid === null` plus a mux-layer pane-death probe, and that
|
||||
covers a session that exited — including a worker that died _inside_ its tmux pane
|
||||
covers a session that exited — including a worker that died *inside* its tmux pane
|
||||
while the local attach client (and therefore `pid`) lives on — one that was
|
||||
detached, and one that was **created but never started**. So the first wait
|
||||
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
|
||||
@@ -184,7 +184,7 @@ blocked, and polling `blocked` alone will sit at its timeout.
|
||||
|
||||
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
|
||||
session emits its one `idle` at startup and then stays `status: "idle"` forever,
|
||||
whatever the pane is doing, so it never emits a _transition_. Since send-and-wait
|
||||
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
|
||||
requires a transition (and so does `fresh=1`), both can only time out there:
|
||||
a documented default `wait` on a shell worker running `sleep 4` times out at the
|
||||
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
|
||||
@@ -218,11 +218,11 @@ with `from=buffer` keeps matching long after the dialog is gone. A worked versio
|
||||
|
||||
### `GET /api/v1/sessions/:id/wait`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
| --------- | -------------------------------------------------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
|
||||
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
|
||||
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
|
||||
| Param | Type | Default | Notes |
|
||||
|-------|------|---------|-------|
|
||||
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
|
||||
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
|
||||
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
|
||||
|
||||
```bash
|
||||
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
|
||||
@@ -239,12 +239,12 @@ a plain signal wait, so check the endpoint path before blaming the parameters.
|
||||
|
||||
### `GET /api/v1/sessions/:id/wait-output`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
| --------- | ------------------------------- | -------- | ----------------------------------------------------------------------------------------------------------- |
|
||||
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
|
||||
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
|
||||
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
|
||||
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
|
||||
| Param | Type | Default | Notes |
|
||||
|-------|------|---------|-------|
|
||||
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
|
||||
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
|
||||
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
|
||||
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
|
||||
|
||||
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
|
||||
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
|
||||
@@ -296,10 +296,10 @@ hand-written query string decodes to a space.
|
||||
|
||||
Two optional fields on the existing endpoint:
|
||||
|
||||
| Field | Type | Notes |
|
||||
| ------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
|
||||
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
|
||||
| Field | Type | Notes |
|
||||
|-------|------|-------|
|
||||
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
|
||||
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
|
||||
|
||||
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
|
||||
"absent" rather than failing validation. That is deliberate: `.optional()` would
|
||||
@@ -330,24 +330,16 @@ All three nest the wait result under `data.wait`, so one client helper works aga
|
||||
any of them:
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"data": {
|
||||
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
|
||||
"status": "idle",
|
||||
"limitPaused": false,
|
||||
"wait": {
|
||||
"signal": "stop",
|
||||
"until": ["stop", "idle", "exit"],
|
||||
"timedOut": false,
|
||||
"immediate": false,
|
||||
"ended": false,
|
||||
"aborted": false,
|
||||
"waitedMs": 8421,
|
||||
"timeoutMs": 60000
|
||||
}
|
||||
{ "success": true, "data": {
|
||||
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
|
||||
"status": "idle",
|
||||
"limitPaused": false,
|
||||
"wait": {
|
||||
"signal": "stop", "until": ["stop", "idle", "exit"],
|
||||
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
|
||||
"waitedMs": 8421, "timeoutMs": 60000
|
||||
}
|
||||
}
|
||||
}}
|
||||
```
|
||||
|
||||
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
|
||||
@@ -361,21 +353,21 @@ redelivery (harmless, the turn it refers to may be long over), while with
|
||||
client that reads `delivered === false` as "duplicate" silently treats a failed send
|
||||
as a success.
|
||||
|
||||
| Field | Type | Meaning |
|
||||
| ---------------- | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
|
||||
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
|
||||
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
|
||||
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
|
||||
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
|
||||
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
|
||||
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
|
||||
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
|
||||
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
|
||||
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
|
||||
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
|
||||
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
|
||||
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
|
||||
| Field | Type | Meaning |
|
||||
|-------|------|---------|
|
||||
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
|
||||
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
|
||||
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
|
||||
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
|
||||
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
|
||||
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
|
||||
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
|
||||
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
|
||||
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
|
||||
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
|
||||
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
|
||||
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
|
||||
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
|
||||
|
||||
Read the outcome by discriminator, in this order:
|
||||
|
||||
@@ -398,12 +390,12 @@ read the timeout as "the worker is wedged" and kill a session that was working f
|
||||
|
||||
### Errors
|
||||
|
||||
| `errorCode` | HTTP | When |
|
||||
| --------------- | ---- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
|
||||
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
|
||||
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
|
||||
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
|
||||
| `errorCode` | HTTP | When |
|
||||
|-------------|------|------|
|
||||
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
|
||||
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
|
||||
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
|
||||
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
|
||||
|
||||
The two capacity codes are deliberately different. A process-wide cap reported as
|
||||
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
|
||||
@@ -454,9 +446,9 @@ Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
|
||||
|
||||
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
|
||||
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
|
||||
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
|
||||
toolSummary?, message?, cwd?, context?, options?: {n, label}[],
|
||||
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
|
||||
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
|
||||
toolSummary?, message?, cwd?, context?, options?: {n, label}[],
|
||||
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
|
||||
`options` is present only when the dialog's numbered choices parsed
|
||||
confidently; `acknowledgedAt` marks an item a human has already looked at
|
||||
(see `/viewed` below) and tells clients not to re-arm its tab alert. Listing
|
||||
@@ -474,7 +466,7 @@ acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
|
||||
first, `422 OPERATION_FAILED` when the session refused input.
|
||||
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
|
||||
- `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId,
|
||||
acknowledged: itemId | null }`. Marks the session's pending **idle** item as
|
||||
acknowledged: itemId | null }`. Marks the session's pending **idle** item as
|
||||
seen by a human (the web UI calls it when you open the session's tab): the
|
||||
item stays pending and answerable, but stops arming the yellow tab alert on
|
||||
every client, including after a reload. Permission/question items are never
|
||||
@@ -498,19 +490,19 @@ heuristic decides whether to ASK, never whether to act.
|
||||
Claude-mode sessions only (others carry their conversation id in their own
|
||||
config object); remote and docker sessions are never offered, because both need
|
||||
another host or container to be up. The plan is in-memory, so a server restart
|
||||
drops it and the offer is gone — the conversations themselves are unaffected,
|
||||
drops it and the offer is gone; the conversations themselves are unaffected,
|
||||
since they live in the CLI's own transcript store and stay reachable from the
|
||||
Resume list. A plan nobody spends expires after 24 hours.
|
||||
|
||||
- `GET /api/v1/reboot-restore` → `{ sessions: RestorableSession[],
|
||||
scrollbackRestored: false }`, ownership-scoped in multi-user mode.
|
||||
scrollbackRestored: false }`, ownership-scoped in multi-user mode.
|
||||
`RestorableSession`: `{ id, name?, workingDir, mode, owner? }`. The persisted
|
||||
record itself is never sent. `scrollbackRestored` is always `false` and exists
|
||||
so a client states it: a restored session is a NEW pane, so the conversation
|
||||
continues and the terminal history does not.
|
||||
- `POST /api/v1/reboot-restore/restore` with `{ sessionIds?: string[] }` (omit
|
||||
to restore everything the caller can see) → `{ restored: RestorableSession[],
|
||||
skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing`
|
||||
skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing`
|
||||
(the directory is gone), `workspace-forbidden` (in multi-user mode it is
|
||||
outside the workspace of the user the session belongs to, re-checked against
|
||||
that owner's current grant rather than the caller's), `already-live` (the conversation is already
|
||||
@@ -521,7 +513,7 @@ skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing`
|
||||
removed from the plan before any pane is built, so a double-click cannot put
|
||||
two panes on one conversation; anything that never became a pane goes back on
|
||||
offer, except `already-live`, which cannot stop being true. A restored session
|
||||
comes back attached, idle and disarmed — respawn controllers and Ralph loops
|
||||
comes back attached, idle and disarmed: respawn controllers and Ralph loops
|
||||
are never re-armed automatically.
|
||||
- `POST /api/v1/reboot-restore/dismiss` → `{ dismissed: n }`. Drops the offer
|
||||
for everything the caller can see.
|
||||
@@ -541,7 +533,7 @@ user guide: [`readmymind.md`](readmymind.md).
|
||||
|
||||
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
|
||||
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
|
||||
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
|
||||
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
|
||||
50, each <= 500 chars). A case with nothing recorded answers an empty
|
||||
profile with `updatedAt: 0`; nothing is persisted by reads.
|
||||
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
|
||||
@@ -574,7 +566,7 @@ same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synce
|
||||
[`claude-voice-plan.md`](claude-voice-plan.md).
|
||||
|
||||
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
|
||||
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
|
||||
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
|
||||
signed in to Claude Code on the server), `expired` (the access token elapsed;
|
||||
running any Claude session refreshes it) or `malformed`. The OAuth token
|
||||
itself is never returned by this or any other endpoint.
|
||||
|
||||
+11
-5
@@ -72,11 +72,17 @@ export interface RebootEvidence {
|
||||
* This heuristic decides whether to ASK, never whether to act. A wrong yes costs
|
||||
* the user a banner they dismiss, because the restore itself waits for a click.
|
||||
*
|
||||
* ⚠️ `os.uptime()` reports the HOST's uptime, which a container shares. A Codeman
|
||||
* running in Docker therefore sees a long uptime after its own container restarts,
|
||||
* the boot test fails, and no banner appears. The feature is effectively off for
|
||||
* containerized installs. That is the safe direction to fail in, and fixing it
|
||||
* needs a boot signal the container actually owns rather than a wider heuristic.
|
||||
* ⚠️ `os.uptime()` reports the HOST's uptime, which a container shares, and that
|
||||
* cuts BOTH ways rather than simply switching the feature off in Docker. After a
|
||||
* genuine host reboot a containerized Codeman sees the host's short uptime, so the
|
||||
* banner DOES appear and the feature works. What it cannot see is a container-only
|
||||
* restart: the host uptime is long, the boot test fails, and no banner appears
|
||||
* although every in-container pane is gone (`docker/server.Dockerfile` installs
|
||||
* tmux inside the Codeman container, and the self-updater restarts the Compose
|
||||
* deployment by exiting the container, so that is the case where this would help
|
||||
* most). Failing quiet is the safe direction, and closing the gap needs a boot
|
||||
* signal the container owns (PID 1's start time, gated on the existing
|
||||
* `isRunningInContainer()`) rather than a wider heuristic.
|
||||
*/
|
||||
export function looksLikeHostReboot(evidence: RebootEvidence): boolean {
|
||||
if (evidence.deadSessionCount === 0) return false;
|
||||
|
||||
@@ -27,7 +27,19 @@ export interface SessionPort {
|
||||
reapplyPersistedSessionState(
|
||||
session: Session,
|
||||
saved: SessionState,
|
||||
phase: 'before-spawn' | 'after-spawn'
|
||||
phase: 'before-spawn' | 'after-spawn',
|
||||
options?: {
|
||||
/**
|
||||
* Re-arm a PENDING auto-resume schedule from the record's `autoResumeAt`.
|
||||
* Default true, which is what a Codeman restart wants: the limit footer
|
||||
* will not reprint on its own, so dropping the stamp there strands the
|
||||
* pause. A reboot restore passes false: the stamp predates the reboot,
|
||||
* the pane is new, and re-arming means every restored session types
|
||||
* `continue` into itself about a minute after one click. Auto-resume
|
||||
* stays ENABLED either way, so it re-arms on fresh evidence.
|
||||
*/
|
||||
rearmAutoResumeSchedule?: boolean;
|
||||
}
|
||||
): Promise<void>;
|
||||
/**
|
||||
* Undo a session that was registered but never got a working pane: the map
|
||||
|
||||
@@ -26,7 +26,7 @@
|
||||
* @mixin Extends CodemanApp.prototype via Object.assign
|
||||
* @dependency app.js (CodemanApp class, showToast)
|
||||
* @dependency api-client.js at runtime (this._api / this._apiJson)
|
||||
* @loadorder 11.7 of 17, after approvals-ui.js
|
||||
* @loadorder 11.65, after approvals-ui.js and before admin-ui.js (11.7)
|
||||
*/
|
||||
|
||||
/** Plain-language wording for one skip reason, for the toast after a restore. */
|
||||
|
||||
@@ -2548,6 +2548,10 @@ body.solo-mode .header-tokens,
|
||||
body.solo-mode .btn-notifications,
|
||||
body.solo-mode .btn-multimonitor,
|
||||
body.solo-mode .header-plan-usage,
|
||||
/* A solo window shows ONE session and has no tab strip to put restored ones in,
|
||||
so offering to rebuild a list of them there is an offer it cannot show the
|
||||
result of. The dashboard that spawned this window carries the banner. */
|
||||
body.solo-mode .reboot-restore-banner,
|
||||
body.solo-mode .btn-lifecycle-log {
|
||||
display: none !important;
|
||||
}
|
||||
|
||||
@@ -41,7 +41,7 @@ import { clampEnvOverridesForOwner } from '../../session-env-clamp.js';
|
||||
import { Session } from '../../session.js';
|
||||
import { resolveClaudeModeForUsername } from '../../user-store.js';
|
||||
import { getCli } from '../../config/cli-registry/registry.js';
|
||||
import { applyWorkspaceHooks } from '../../hooks-config.js';
|
||||
import { applyWorkspaceHooks, seedAgentSessionPreamble } from '../../hooks-config.js';
|
||||
import { getLifecycleLog } from '../../session-lifecycle-log.js';
|
||||
import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js';
|
||||
import { SseEvent } from '../sse-events.js';
|
||||
@@ -105,11 +105,16 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
// the user resumed by hand from the Resume list is already on screen, and a
|
||||
// second pane on it would fight the first for the same transcript. This one
|
||||
// is never re-offered: unlike a missing workspace, it cannot stop being true.
|
||||
const liveSessionIds = new Set(ctx.sessions.keys());
|
||||
const liveConversationIds = new Set(
|
||||
[...ctx.sessions.values()].map((session) => session.claudeSessionId).filter((id): id is string => !!id)
|
||||
);
|
||||
const { restore, skipped } = rejectAlreadyLive(taken, liveSessionIds, liveConversationIds);
|
||||
// Read fresh each time rather than snapshotted once: the loop below awaits a
|
||||
// real `startInteractive()` per entry, so by the tenth entry a snapshot taken
|
||||
// here is tens of seconds old, and a conversation the user resumed by hand in
|
||||
// that window would be invisible to it.
|
||||
const liveSessionIds = () => new Set(ctx.sessions.keys());
|
||||
const liveConversationIds = () =>
|
||||
new Set(
|
||||
[...ctx.sessions.values()].map((session) => session.claudeSessionId).filter((id): id is string => !!id)
|
||||
);
|
||||
const { restore, skipped } = rejectAlreadyLive(taken, liveSessionIds(), liveConversationIds());
|
||||
for (const entry of taken) {
|
||||
if (skipped.some((s) => s.sessionId === entry.sessionId)) unspent.delete(entry);
|
||||
}
|
||||
@@ -119,6 +124,18 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
const workspaceHooksEnabled = await ctx.getWorkspaceHooksEnabled();
|
||||
|
||||
for (const entry of restore) {
|
||||
// The already-live check, re-run against the board as it is NOW. The pass
|
||||
// above decided the batch; this catches a conversation that went live while
|
||||
// an earlier entry in this same batch was starting. Spent rather than
|
||||
// returned to the plan, for the same reason as the batch pass: unlike a
|
||||
// missing workspace or a withdrawn grant, an open conversation is not a
|
||||
// condition that stops being true.
|
||||
const [lateLive] = rejectAlreadyLive([entry], liveSessionIds(), liveConversationIds()).skipped;
|
||||
if (lateLive) {
|
||||
failures.push(lateLive);
|
||||
unspent.delete(entry);
|
||||
continue;
|
||||
}
|
||||
// Capacity is re-checked per iteration, because this loop is itself
|
||||
// creating the sessions it counts. The offer can be a day old, so the
|
||||
// board may be fuller now than the plan assumed.
|
||||
@@ -157,6 +174,12 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
workingDir: saved.workingDir,
|
||||
mode: saved.mode,
|
||||
name: saved.name,
|
||||
// Without this the constructor re-infers ownership from the name, so a
|
||||
// session the user renamed by hand to something shaped like `w<n>-<case>`
|
||||
// comes back as `placeholder` and auto-naming overwrites their name on
|
||||
// the next prompt. The route persists below, so the loss would go to
|
||||
// disk. `restoreMuxSessions()` passes it for the same reason.
|
||||
nameSource: saved.nameSource,
|
||||
createdAt: saved.createdAt,
|
||||
mux: ctx.mux,
|
||||
useMux: true,
|
||||
@@ -197,7 +220,14 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
// the reduced one and drop the pin that keeps it from being pruned. A
|
||||
// listener-driven persist can still land inside the debounce window
|
||||
// while the pane starts; the write below repairs the record.
|
||||
await ctx.reapplyPersistedSessionState(session, saved, 'after-spawn');
|
||||
// `rearmAutoResumeSchedule: false`: the saved stamp predates the reboot and
|
||||
// the pane is new, so honouring it would have every restored session type
|
||||
// `continue` into itself about a minute after one click. Auto-resume stays
|
||||
// enabled and re-arms on the next real limit message. This is also what the
|
||||
// module header promises ("comes back attached, idle and disarmed").
|
||||
await ctx.reapplyPersistedSessionState(session, saved, 'after-spawn', {
|
||||
rearmAutoResumeSchedule: false,
|
||||
});
|
||||
ctx.persistSessionState(session);
|
||||
|
||||
// A session without its workspace hooks goes silently blind: no stop or
|
||||
@@ -211,6 +241,16 @@ export function registerRebootRestoreRoutes(app: FastifyInstance, ctx: RebootRes
|
||||
);
|
||||
}
|
||||
|
||||
// Both create paths seed this; without it a restored claude session's agent
|
||||
// skill falls back to writing out the whole ~150-line §0 preamble. Remote and
|
||||
// docker sessions never reach here (the plan rejects them as
|
||||
// `remote-or-docker`), so the local-only condition is structural.
|
||||
if (getCli(session.mode)?.capabilities.agentSkillInjection && (await ctx.getAgentSkillEnabled())) {
|
||||
await seedAgentSessionPreamble(session.id).catch((err: unknown) =>
|
||||
console.warn(`[agent-skill] preamble seed failed for ${session.id}: ${getErrorMessage(err)}`)
|
||||
);
|
||||
}
|
||||
|
||||
getLifecycleLog().log({ event: 'recovered', sessionId: session.id, name: session.name });
|
||||
// Every other open tab and phone needs this; the clicking tab already has
|
||||
// the response, and the client's handler is an idempotent upsert.
|
||||
|
||||
+19
-2
@@ -2940,7 +2940,8 @@ export class WebServer extends EventEmitter {
|
||||
async reapplyPersistedSessionState(
|
||||
session: Session,
|
||||
saved: SessionState,
|
||||
phase: 'before-spawn' | 'after-spawn'
|
||||
phase: 'before-spawn' | 'after-spawn',
|
||||
options?: { rearmAutoResumeSchedule?: boolean }
|
||||
): Promise<void> {
|
||||
if (phase === 'before-spawn') {
|
||||
// The custom-model env has to be rebuilt from the endpoint store: the persist
|
||||
@@ -2968,7 +2969,14 @@ export class WebServer extends EventEmitter {
|
||||
session.setAutoClear(saved.autoClearEnabled ?? false, saved.autoClearThreshold);
|
||||
}
|
||||
if (saved.autoResumeEnabled) {
|
||||
session.restoreAutoResume(true, saved.autoResumeAt);
|
||||
// The stamp is re-armed by default, because a Codeman restart leaves the
|
||||
// limit footer un-reprinted and dropping it there would strand the pause.
|
||||
// A reboot restore opts out: that stamp predates the reboot, the pane is
|
||||
// new, and honouring it means every session the user restored types
|
||||
// `continue` into itself about a minute later, unattended. The setting
|
||||
// itself stays on either way, so it re-arms on the next limit message.
|
||||
const rearm = options?.rearmAutoResumeSchedule !== false;
|
||||
session.restoreAutoResume(true, rearm ? saved.autoResumeAt : undefined);
|
||||
}
|
||||
if (saved.inputTokens !== undefined || saved.outputTokens !== undefined || saved.totalCost !== undefined) {
|
||||
session.restoreTokens(saved.inputTokens ?? 0, saved.outputTokens ?? 0, saved.totalCost ?? 0);
|
||||
@@ -3024,9 +3032,18 @@ export class WebServer extends EventEmitter {
|
||||
session.ralphTracker.stopWatchingFixPlan();
|
||||
const summaryTracker = this.runSummaryTrackers.get(sessionId);
|
||||
if (summaryTracker) {
|
||||
// Closes the run's own record before the tracker goes, the way
|
||||
// `_doCleanupSession()` does. Cosmetic rather than load-bearing, but a
|
||||
// run left open reads as still going in the away digest.
|
||||
summaryTracker.recordSessionStopped();
|
||||
summaryTracker.stop();
|
||||
this.runSummaryTrackers.delete(sessionId);
|
||||
}
|
||||
// Also mirrors `_doCleanupSession()`. The PERSISTED Ralph state is left
|
||||
// alone on purpose (that is one of the things separating this from
|
||||
// cleanupSession); this only clears the in-memory tracker the failed
|
||||
// construction built, which the retry reuses the id of.
|
||||
session.ralphTracker.fullReset();
|
||||
|
||||
// --- what anything else may have attached to this id in the meantime ---
|
||||
// A rebuild can fail AFTER startInteractive() resolved, and a restored
|
||||
|
||||
@@ -11,7 +11,10 @@
|
||||
* The routes read the process-wide `rebootRestoreRegistry` singleton, so every
|
||||
* test resets it; a leaked entry would bleed into the next one.
|
||||
*/
|
||||
import { describe, it, expect, afterEach } from 'vitest';
|
||||
import { describe, it, expect, afterEach, beforeEach } from 'vitest';
|
||||
import { mkdtempSync, rmSync } from 'node:fs';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { join } from 'node:path';
|
||||
import Fastify, { type FastifyInstance } from 'fastify';
|
||||
import fastifyCookie from '@fastify/cookie';
|
||||
import { registerRebootRestoreRoutes } from '../../src/web/routes/reboot-restore-routes.js';
|
||||
@@ -187,6 +190,49 @@ describe('POST /api/reboot-restore/restore', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/reboot-restore/restore: multi-user workspace confinement', () => {
|
||||
const saved: Record<string, string | undefined> = {};
|
||||
let realDir: string;
|
||||
|
||||
beforeEach(() => {
|
||||
saved.CODEMAN_MULTIUSER = process.env.CODEMAN_MULTIUSER;
|
||||
process.env.CODEMAN_MULTIUSER = '1';
|
||||
// This branch sits AFTER the existsSync check, so the workspace has to be
|
||||
// real for the confinement rule to be the thing that rejects the entry.
|
||||
realDir = mkdtempSync(join(tmpdir(), 'codeman-reboot-restore-real-'));
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
if (saved.CODEMAN_MULTIUSER === undefined) delete process.env.CODEMAN_MULTIUSER;
|
||||
else process.env.CODEMAN_MULTIUSER = saved.CODEMAN_MULTIUSER;
|
||||
rmSync(realDir, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
it("refuses a workspace outside the OWNER's case space, and leaves it on offer", async () => {
|
||||
const entry = offerEntry('a', 'alice');
|
||||
entry.workingDir = realDir;
|
||||
(entry.state as { workingDir: string }).workingDir = realDir;
|
||||
rebootRestoreRegistry.set([entry]);
|
||||
|
||||
// An admin does the clicking. The confinement is still resolved against
|
||||
// alice, the entry's OWNER: `isWorkingDirAllowed` waves an admin through, so
|
||||
// reading the caller here would hand an admin the power to rebuild another
|
||||
// user's session anywhere on the box.
|
||||
const app = await createHarness({ username: 'root-user', role: 'admin' });
|
||||
const res = await app.inject({ method: 'POST', url: '/api/reboot-restore/restore', payload: {} });
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json().data.restored).toEqual([]);
|
||||
expect(res.json().data.skipped).toEqual([{ sessionId: 'a', reason: 'workspace-forbidden' }]);
|
||||
|
||||
// A withdrawn grant can be given back, so unlike `already-live` this is not
|
||||
// the permanent kind of refusal and the entry stays claimable.
|
||||
const left = (await app.inject({ method: 'GET', url: '/api/reboot-restore' })).json().data;
|
||||
expect(left.sessions.map((s: { id: string }) => s.id)).toEqual(['a']);
|
||||
await app.close();
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/reboot-restore/dismiss', () => {
|
||||
it('drops the offer and leaves the banner with nothing to show', async () => {
|
||||
rebootRestoreRegistry.set([offerEntry('a'), offerEntry('b')]);
|
||||
|
||||
Reference in New Issue
Block a user