docs(skill): never branch on .status, it is wrong in both directions

Measured on a live claude worker: `GET /api/v1/sessions/:id` reported
`status: "idle"` while the worker was mid-turn and actively producing output, with
`lastActivityAt` equal to the moment of the call. The skill already warned that a
worker which dies inside its pane also reads `idle`, so the field is unreliable in
both directions and nothing an agent does should depend on it.

Synchronize on `stop` via send-and-wait or on an output marker. To judge from
outside, sample `terminal?tail=` twice a few seconds apart: a changing buffer is the
only cheap positive proof a worker is still working. `wait?until=exit` stays the
death check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Codeman maintainer
2026-08-09 12:29:15 +02:00
parent 4ed86aa0cd
commit 0aa16cd4d3
2 changed files with 24 additions and 1 deletions
+7
View File
@@ -161,6 +161,13 @@ You are yourself a session on this server, and the API has **no undo**.
**clamped** (ceiling 600 s): read back `wait.timeoutMs` for what was applied. The **clamped** (ceiling 600 s): read back `wait.timeoutMs` for what was applied. The
clamp covers positive integers only: `0`, a negative, a fraction or `30s` is a 400, clamp covers positive integers only: `0`, a negative, a fraction or `30s` is a 400,
so round any computed remainder and drop it entirely rather than sending zero. so round any computed remainder and drop it entirely rather than sending zero.
- **Never branch on `.data.status`.** It is a heuristic and is often wrong in both
directions: measured on a live claude worker reading `idle` while it was mid-turn
and actively producing output (`lastActivityAt` equal to the moment of the call),
and a worker that died inside its pane also reads `idle`. Synchronize on `stop` via
send-and-wait, or on an output marker. To judge from outside, sample
`terminal?tail=` twice a few seconds apart: a changing buffer is the only cheap
positive proof a worker is still working. `wait?until=exit` is the death check.
- **`stop` and `blocked` fire for `claude` sessions only** (Claude Code hooks). On - **`stop` and `blocked` fire for `claude` sessions only** (Claude Code hooks). On
`shell`/`opencode`/`codex`/`gemini`/`antigravity`, requesting them explicitly is a `shell`/`opencode`/`codex`/`gemini`/`antigravity`, requesting them explicitly is a
400 — and lifecycle transitions there are coarse (a short shell command may emit 400 — and lifecycle transitions there are coarse (a short shell command may emit
+17 -1
View File
@@ -38,7 +38,7 @@ read the status with `-w '%{http_code}'` and the raw body before assuming a bug.
| Task | Call | | Task | Call |
|------|------| |------|------|
| list sessions (metadata only, ~1.5 KB each, safe to poll) | `GET /api/v1/sessions` | | list sessions (metadata only, ~1.5 KB each, safe to poll) | `GET /api/v1/sessions` |
| one session (has `.data.pid`, `null` until the PTY spawns) | `GET /api/v1/sessions/:id` — ⚠️ **not a liveness check**: a worker that dies inside its pane keeps `status:"idle"` and a pid (the tmux attach client); `wait?until=exit` is the death check | | one session (has `.data.pid`, `null` until the PTY spawns) | `GET /api/v1/sessions/:id` — ⚠️ **neither a liveness nor a busy check**, see below |
| unified list incl. history | `GET /api/v1/sessions/unified` → `.data.sessions[]` (NOT `.data[]`), and it folds in transcript history from the whole machine — never use it to verify cleanup; `GET /api/v1/sessions` is the cleanup check | | unified list incl. history | `GET /api/v1/sessions/unified` → `.data.sessions[]` (NOT `.data[]`), and it folds in transcript history from the whole machine — never use it to verify cleanup; `GET /api/v1/sessions` is the cleanup check |
| start case + session in one call | `POST /api/v1/quick-start` | | start case + session in one call | `POST /api/v1/quick-start` |
| send input | `POST /api/v1/sessions/:id/input` | | send input | `POST /api/v1/sessions/:id/input` |
@@ -60,6 +60,22 @@ cleanup: your worker keeps burning tokens where neither you nor the user can see
and the list you would check to confirm cleanup shows it gone. Delete plainly, and let and the list you would check to confirm cleanup shows it gone. Delete plainly, and let
`killMux` default. `killMux` default.
⚠️ **`.data.status` is a heuristic and is often simply wrong. Never branch on it.**
Measured on a live claude worker: `status` read `idle` while the worker was mid-turn
and actively producing output, with `lastActivityAt` equal to the moment of the call.
It is wrong in both directions, so neither value tells you anything you can act on:
- **`idle` does not mean finished.** Use `stop` (the definitive end-of-turn hook) via
send-and-wait, or an output marker. If you must judge from outside, sample
`terminal?tail=` twice a few seconds apart and compare: a changing buffer is the
only cheap positive proof that a worker is still working.
- **`idle` does not mean alive.** A worker that dies inside its pane keeps
`status:"idle"` and a pid (that pid is the local tmux attach client, not the
worker). `wait?until=exit` is the death check.
Treat `status` as a UI hint. Every synchronization decision in these recipes is built
on signals and markers for exactly this reason.
⚠️ `GET /api/v1/sessions/:id/output` → `.data.textOutput` looks like the obvious read ⚠️ `GET /api/v1/sessions/:id/output` → `.data.textOutput` looks like the obvious read
but stays **empty for interactive tmux-backed sessions** (it is fed only by the legacy but stays **empty for interactive tmux-backed sessions** (it is fed only by the legacy
JSON-stream path). Verified empty on live claude and shell sessions. Use JSON-stream path). Verified empty on live claude and shell sessions. Use