`GET /api/pi/status` shipped undocumented in the agent skill, and only a human
reading the doc noticed. Turns out none of its five siblings were documented
either, so this adds the whole family in one place: spawning with a mode whose
CLI is absent fails with OPERATION_FAILED rather than falling back, which is
exactly what an agent picking a backend it did not choose needs to know. Pi's
extra `.data.version` is called out, since a false `available:false` there means
an unrelated `pi` is in front on PATH.
On whether the endpoint scanner should also check registered-to-documented:
measured, and NO for the general case. The skill documents 34 of 217 registered
endpoints deliberately (it is an agent guide, not an API reference), so a blanket
reverse check needs a 183-entry allowlist that would fail CI on unrelated route
work and get appended to mechanically, which is worse than the gap it closes.
Grouping by path shape does not save it either: the families that yields are
things like `DELETE /api/<any>/:id`, lumping cases, webviews and docker hosts
together, and it would not have caught this gap anyway (the family had zero
documented members).
What IS cheap is a family the schema can enumerate with no allowlist: the new
assertion derives the agent modes from the Zod enum and requires each one's
`/api/<mode>/status` to be documented, so a seventh backend fails here until it
is. The sibling scanner still proves the other direction, that nothing documented
is a 404. Both mutation-checked: dropping pi's probe fails the new guard, and
documenting a nonexistent probe fails the old one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Second review pass on #282, the three items left open after f4dcfbe.
1. `codeman doctor` and the run mode disagreed about pi. The registry entry
accepted a bare `which pi` hit while pi-cli-resolver demanded semver-shaped
`--version` output, so the Dependencies panel could report an installed Pi CLI
on a box where Run Pi stays hidden, which reads as a broken mode rather than a
missing install. Both sides now share one exported PI_VERSION_REGEX, and
PathResolver gains an opt-in `requireVersionMatch` so a binary that fails the
shape check is reported MISSING instead of installed-with-unknown-version.
Only pi sets it; every other tool keeps its current behaviour.
2. The isAltScreenStripMode comment justified excluding pi with "the alt screen
is load-bearing for its fullscreen TUI". That is not what exclusion does: pi
is tmux-backed, so it falls through to isMuxAltScreenOnlyStripMode, which
strips the alt-screen toggles anyway. What exclusion actually preserves is
`\x1b[3J` and the mouse DECSETs, which is the real reason (pi renders into the
main screen and is mouse-aware). Comment and changeset now say that, and state
the consequence: fullscreen pi paints into the main buffer, like vim in a tmux
shell session.
3. skills/codeman still enumerated the five pre-pi modes in nine places, telling
agents a backend does not exist and understating class-wide caveats by one
mode. All updated, plus stale session.ts line references refreshed.
Tests: a new static guard derives the mode set from the Zod schema (not a copy)
and fails when a skill enumeration lists a partial set of external CLIs, verified
by mutation. It also documents the one legitimate exception it found: the "writes
no transcript" lists drop codex, which does write a rollout Codeman reads back.
Plus doctor cases for an unrelated `pi` on PATH and registry/resolver regex parity.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.
Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
sources: pending approval dialog, user goals, last assistant turn tail,
recent prompts, tool activity, git workspace signals, away context,
sibling sessions, rethink state; 30 KB budget, whole-section drop from
the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
keeps only a 500-char snippet) and git signal collection (execFile,
2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
session, opus by default (readMyMindModel setting), strict JSON
contract with 1-3 suggestions (continue / verify / redirect), newline
stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
prediction in flight per session (409 CONFLICT), rethink body
{ steer, rejected }; ownership via findSessionOrFail
Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
until readMyMindEnabled is ON, desktop only (phone key is phase 3);
modal with editable suggestion + rationale and Send / Insert /
Rethink / Dismiss; suggestion text rendered via value/textContent only
and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
strings
Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals
plus the user's recently submitted prompts, captured from the Claude session
transcript behind the new synced readMyMindEnabled setting (default OFF).
- intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps,
consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json
- transcript-watcher.ts: new transcript:user_prompt event for typed user turns
(tool_result-only entries stay silent); capture wiring in server.ts is
claude-only and gated on the setting per event
- readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership
via findSessionOrFail, strict Zod schema
- agent skill: SKILL.md recipe + endpoints.md rows so agents can read and
record intent (PUT replaces: read + merge; never delete unprompted)
- groundwork for the phase-2 predictor button; nothing is ever auto-sent
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Version-gated fail-closed at 2.1.224 (the cross-session-messaging release,
flag presence verified against that binary): an unknown or older CLI yields
a spawn command byte-identical to before, because claude aborts startup on
an unknown option and that would kill every session spawn. The value is
allowlist-sanitized ahead of the double-quoted interpolation, and only the
local command carries the flag; docker/remote builders never see it since
their CLI is not the probed binary. Verified E2E on an isolated instance:
cmdline shows --name, ListAgents lists the session name, replies arrive
tagged from-name.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude Code v2.1.224+ gives sessions ListAgents/SendMessage and a per-session
inbox socket. Codeman's claude workers are ordinary local Claude Code sessions,
so the agent skill now teaches task delivery and result collection over
messaging where available (multi-line exactly-once messages, mid-turn steering,
latched replies), with the HTTP primitives keeping spawn, readiness,
synchronization, liveness and delete, and a bounded fallback to the HTTP
recipes whenever the feature is absent. All mechanics verified live against
claude-cli 2.1.226.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured on a live claude worker: `GET /api/v1/sessions/:id` reported
`status: "idle"` while the worker was mid-turn and actively producing output, with
`lastActivityAt` equal to the moment of the call. The skill already warned that a
worker which dies inside its pane also reads `idle`, so the field is unreliable in
both directions and nothing an agent does should depend on it.
Synchronize on `stop` via send-and-wait or on an output marker. To judge from
outside, sample `terminal?tail=` twice a few seconds apart: a changing buffer is the
only cheap positive proof a worker is still working. `wait?until=exit` stays the
death check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The readiness gate matched `bypass`, which is the status bar of ONE permission mode.
`buildPermissionArgs()` also spawns `--permission-mode auto`, `--allowedTools` and
plain `normal`, and the mode is not exposed on `GET /api/v1/sessions/:id`, so an agent
cannot know which token to expect. A non-default worker was therefore reported broken
after burning the whole ladder.
Measured one pane per mode against claude-cli 2.1.226:
--dangerously-skip-permissions -> "bypass permissions on"
--permission-mode auto -> "auto mode on"
--allowedTools Read,Grep -> "don't ask on"
(none, normal) -> "don't ask on"
--permission-mode plan -> "plan mode on"
Every one ends `(shift+tab to cycle)`, so `shift+tab` is the single space-free token
that means "the composer is up" in every mode, and it is what the ladder matches now.
Verified live end to end on a virgin case: stage 1 misses while the trust dialog is up,
stage 2 accepts it, stage 3 matches in 623ms.
⚠️ `shift+tab` contains a `+`, so it only works through `--data-urlencode`. In a
hand-built query the `+` decodes to a space and the server searches for `shift tab`,
which never appears; the response echoes `match: "shift tab"`, which is how to spot it.
Measured both ways. The stage-4 fallback (make the worker echo a split token, proving
readiness by answering rather than by chrome) stays as the last resort, and is now also
verified live: it matched in 2.5s, with the token surviving the space-less TUI intact.
Also portable ANSI stripping: the read pipelines used `sed 's/\x1b...'`, and BSD sed
(the macOS default) has no `\xHH` escape, so on macOS the strip silently removed
nothing and handed the agent raw ANSI. They now build a real ESC with `printf`.
And endpoints.md gaps: the `FORBIDDEN` 403 row and which auth responses are plain text
rather than the JSON envelope, the input size cap, the undocumented `killMux` parameter
on DELETE, and the fact that zero/negative/non-integer timeouts are rejected with a 400
rather than clamped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.
Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
Case names resolve through linked-cases.json first, mirroring the
server's resolveCasePath(), so a case linked in from outside
~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
skill is never touched, and a symlinked skill dir is refused (this
repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
ports/config-port.ts, server.ts, session-routes.ts (add-only injection
on Claude session create and quick-start), plus the App Settings toggle.
Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
Shell state does not survive between agent tool calls, and an undefined
is_self exited 127, firing the `||` branch and deleting the caller's own
session with the one guard bypassed. The request now lives inside the
guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
tool call, so the documented resend-identical-request loop stopped being
a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
workers. It returns clean transcript text; the terminal scrape it
replaces returns a wall of TUI repaint noise. Its transcript flush lags
the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
yielded the literal session id "null" and burned the whole readiness
budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
corrected the hooks-config comment that claimed a toggle-off sweep
exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
and that caseName resolves linked cases, so a generic name can land a
worker in a real repo.
Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>