Claude Code 2.1.277 takes typed text the moment its composer paints but
ignores Enter for the first 30 to 50 seconds after it (measured 2026-09-19
through the input route: an Enter at 28 s stranded the prompt, one at 51 s
submitted it). The text+Enter pair `sendInput` sends 50 ms apart therefore
left every programmatic prompt sitting unsent, and every waiter burned its
timeout on a turn that never started.
Server: `SubmitVerifier` (session-submit-verifier.ts), armed from
`writeViaMux` for every mux write that carried a carriage return, reads the
pane on a 2 s to 60 s schedule and re-sends Enter only while the last
composer line (the CLI's own prompt glyph) still holds the head of what was
sent. An empty composer, other text, or no composer line at all ends it; a
newer write replaces the schedule.
Skill: `sendwait` gets the same loop (`_composer_text`, no-break space
stripped by its bytes for BSD sed) for servers that predate this, and the
preamble version moves to 1.30.1 so seeded agents pick up the fresh copy.
SKILL.md's heredoc and the plugin mirror are regenerated.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The eye button rendered `data.text`, which is one row: the last assistant
row of the transcript. A Claude answer is a median of 3 model messages
(p90 11) split around tool calls, so the brief view usually showed the tail
of an answer ("Done.", "Let me look.") and the substance appeared only after
More. The full view was fine, which is why the brief one read as broken by
comparison.
The brief view now asks `?context=turn`. The reader answers with the
assistant messages of the last ANSWERED turn (`selectLastAnsweredTurn`: the
highest `turn` that has an assistant row, so a prompt queued after the
answer does not blank the view) and the frontend renders them exactly as
the full view renders that turn: one badge, then continuation segments,
gated on the numeric `turn` as before.
`data.text` is unchanged in every context — still the last assistant row,
never `messages.at(-1)` — because agent pollers hash it. Readers that emit
no turns (Codex, the pane parser, DeepSeek, an older server) return `text`
only for `context=turn`, and the brief view keeps its single card for them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
A long orchestration creates one case directory per worker and deleting the
sessions never removed them, so ~/codeman-cases accumulated scratch folders
that were indistinguishable from real projects. They are now labelled and
have a cleanup path.
- src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven
spawn gets a .codeman-agent-case.json marker (when, by whom, parent session,
mode). Only the create branch writes it, so a linked case, a cloned repo or
any pre-existing path is never labelled; reading is total, so a malformed
marker means "not agent-created" rather than a half-trusted entry.
- The signal is the new X-Codeman-Agent-Origin header the skill preamble sets
on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body
field, falling back to a resolved parentSessionId so a worker spawned by a
stale skill copy is still labelled.
- GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is
a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage
badges each case and offers a review-then-delete sweep that names every
directory in its confirm and skips any case a live session is working in.
Removal stays on the existing DELETE /api/cases/:name.
- Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was
written per claude session and never removed (236 leftovers measured on a
working machine). Now deleted with the session and swept at boot, guarded by
a live-session keep set plus a 7-day age floor.
Verified end to end on an isolated instance: marker written for header, body
and lineage-only spawns, absent with no agent signal and for a pre-existing
directory; inUse flipping on session end; badge, sticky bar, confirm and sweep
driven in a browser; preamble seeded on create, removed on delete, boot sweep
taking only the aged orphans.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Claude reader concatenated every assistant row between two human prompts
into one card, fusing up to 74 distinct model messages into a single card, and
it never read the attachment rows that hold a prompt typed while the agent was
working. Measured over 57 real transcripts on 2026-09-01, the viewer shows
1,806 messages instead of 356 and 353 user cards instead of 178, with the
assistant text sequence unchanged row for row and the response without
?context=full byte-identical on all 57 files.
One assistant row IS one whole model message: in that corpus no assistant row
carries more than one content block and no message id carries more than one
text block, so there was nothing to reassemble. Each row becomes its own
message carrying an additive {kind, label, turn}, and the frontend renders a
same-role run inside one turn as badge-less continuation segments — which is
what keeps a p90 of 11 messages per turn from reading as card spam. A numeric
turn gates that rendering, so Codex, the external-CLI pane parser and an older
server keep one badge per card.
A prompt typed while Claude is working is recorded ONLY as an
attachment/queued_command row. Taking it when origin.kind is 'human' and
commandMode is 'prompt' recovers 162 user cards from 163 such rows — one is a
verbatim repeat inside an unanswered user run and is collapsed by the existing
dedup guard — and restores the turn boundary whose absence let the assistant
runs fuse. The CLI's own queue entries are cleanly separable: of 322
queued_command rows, 159 are commandMode 'task-notification' and not one of
them carries an origin key.
This narrows #169 rather than reverting it: sidechain exclusion, the
restored-<uuid8> rebind, replayed-snapshot dedup and synthetic-row filtering
are all unchanged and still asserted.
Every claude session in a directory claude had not seen before died about six
seconds after it started (`Pane is dead (status 1)`), before the agent drew a
composer. Reproduced on a fresh case and measured.
Claude Code 2.1.252 rewrote the dialog. It used to be
❯ 1. Yes, I trust this folder
2. No, exit
and is now unnumbered, reversed, and highlights the option that quits:
❯ No, exit
Yes, I trust this folder
Detection still worked (the confirm affordance carries the match once the
numbered option text is gone), so the failure was entirely in the answer: the
auto-accept pressed Enter on the highlighted default, which is now exit.
trustDialogNextKey() reads the ❯ marker off the rendered pane and returns ONE
keystroke at a time: an arrow while the cursor is on the wrong option, Enter
only once the screen shows it on the trust option, and null for a frame that
does not say. Both layouts are handled, and which way the trust option lies is
read from the frame rather than assumed, so a further reordering costs a
repaint instead of a session. The last marked option wins, because the
direct-PTY fallback reads an append-only buffer where an older frame must not
out-vote the freshest one.
Two things only a live pane showed:
- The scan ran solely from the PTY onData handler. The arrow that moves the
cursor is the last output the pane produces, so the first fix parked every
session with the cursor sitting on the right option and no Enter ever sent.
It now schedules its own follow-up read (_trustDialogTimer, cleared in
_clearAllTimers()), offset past the scan throttle so the chain cannot break
on a boundary.
- The keystroke cap goes 3 -> 6, since answering is no longer one press.
The bundled codeman skill had the same blind \r as its bounded fallback, so
preamble 1.21.0 replaces it with _trust_key/_accept_trust: read
terminal?full=1, steer onto the trust option, re-read, then confirm. Those
keystrokes go out under their own clientId, because input sequence numbers are
monotonic per client and spending prompt numbers on dialog keys would make the
next send-and-wait look like a stale duplicate and vanish while reporting
success. The readiness recipes in docs/extending-codeman.md,
docs/api-reference.md and the skill's own reference carry the corrected answer,
plus a symptom-table entry for a worker whose pane is dead seconds after spawn.
Verified live on an isolated instance (own data dir and tmux socket): fresh
case -> arrow at 5 s -> Enter at 7 s -> composer, with hasTrustDialogAccepted
recorded. With the server-side auto-accept disabled in a throwaway copy, the
skill's fallback cleared a genuinely parked dialog in 1.1 s and spawn_worker
took a brand-new case to a live composer in 7.2 s; spawn_workers + sendwait +
last_text then ran end to end.
Two prose lists in skills/codeman/ named some but not all external CLI
modes after the omp-mode rebase, which is exactly the drift
test/agent-skill-mode-lists.test.ts exists to catch: SKILL.md's
no-hook-signals list was missing omp, and endpoints.md's version-probe
sentence named pi/grok/omp as a bare 3-mode run with no matching class.
The agent skill could spawn a worker in any mode, but it could only
DRIVE a claude one: every other CLI has neither a real end-of-turn
signal nor an answer to read, so the recipes route them through output
markers.
dsh has both halves now -- its harness reports idle/working/blocked to
Codeman, and the previous commit reads its transcript -- so it joins
claude as a mode the four verbs work on unchanged. `spawn_workers alpha
beta:deepseek` is a mixed fleet in one call, and `sendwait` / `last_text`
/ `delete_session` need no per-mode variant.
Preamble 1.20.0 (SKILL.md's §0 heredoc regenerated from it):
- `spawn_worker` grows a deepseek branch that gates on the harness
composer. ⚠️ Readiness there is NOT the stop signal: the harness
reports idle at BOOT ~300 ms before its composer paints (measured
2.26 s vs 2.56 s after spawn), so a send-and-wait fired straight after
quick-start resolves on the boot edge, reports a turn that never ran,
and strands the prompt in a pane not yet taking input. Waiting for the
composer also spends that edge, since signals are edge-triggered.
- `spawn_workers` takes `name[:mode]`, so a mixed fleet stays one
concurrent call. Case names still have to be unique -- the mode never
disambiguates two workers that would share a directory.
- `sendwait` asks for `wait:"stop,exit"` instead of the `wait:true`
default set. That set also carries `idle`, which for an external CLI is
inferred from output stabilization: on a dsh worker whose TUI repaints
rarely, the re-wait resolved in 0 ms with `signal:"idle"` on a turn
with three minutes left to run. It also makes a wrong mode loud -- the
modes that cannot deliver `stop` answer 400 before writing anything,
instead of resolving on a flap.
- The self-heal resend carries `delivered:true` forward. The resend is a
tagged duplicate, so the server truthfully reports `delivered:false`
about a write it skipped, and §1's cleanup then read a completed turn
as an undelivered one and kept a finished worker forever.
- dsh workers spawn with the permission posture the Run button sends,
because the harness default still asks and a worker parked on an
approval row cannot finish a fan-out. The multi-user clamp still
applies.
Docs: a worked dsh flow in recipes.md, readiness and the signal rules in
verbs.md, and the corrections this makes necessary -- `stop`/`blocked`
are no longer claude-only, and `last-response` is no longer permanently
empty for deepseek. The integration guide gains a section on reading a
session back and driving one as a worker; its web-UI section was also
stale (that server moved out of a shell session).
The static guard that keeps those lists from naming some external CLIs but
not others is extended rather than exempted: it now knows the three real
classes inside that family (no transcript, no hook signals, and the
positive twin -- the modes whose answers can be read), with the hook class
derived from `hooksAvailableForMode()` so the predicate and the prose
cannot drift apart. Any other partial list still fails, and a new backend
belongs to none of the classes until someone says so.
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.
DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:
1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
$DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
`base` -- the interactive terminal front door is always a third-party
plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
button gates on the latter, because reporting only the binary would spawn a
pane that dies on arrival. When the binary is present but no profile is,
the run menu offers to install one (POST /api/deepseek/install-profile).
2. The permission switch is an env var, not a flag. The harness has no
command-line permission option; its sandbox/approval rows read
DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
own workspace-write, which still asks, so the multi-user clamp is the
only-if-sent branch and clamps to workspace-write, never read-only.
3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
earned that. The terminal front door reports idle/working/blocked to a
supervising process over a generic env-gated contract; a generated shim
(deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
report to /api/hook-event as stop / agent_working / permission_prompt. So a
dsh session gets definitive respawn triggers, real wait-endpoint signals and
real Approvals Inbox items instead of output-stabilization guesswork.
`agent_working` is new (157th SSE constant) and joins
APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
alert at once.
The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.
Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.
Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.
Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The agent skill's endpoints.md is what other agents read as ground truth, and
three of its facts went stale when grok landed:
- `/api/v1/grok/status` was added to the probe list, but the sentence after it
still said only Pi's response carries `.data.version`. Grok's carries it for
the same reason (a squatted binary name), and an agent that trusts the old
wording has no way to tell a misresolved grok from an absent one.
- the `active-tools` bullet listed grok among the modes it stays empty for, then
claimed in the same breath that `isExternalCliMode` "lists only those five".
- its three source line refs had all drifted: `isExternalCliMode` is now
session.ts:174-183 (it was already wrong before this branch), the external-CLI
early return is session.ts:2261, and TEXT_COMMAND_PATTERN is
bash-tool-parser.ts:89.
CLAUDE.md and architecture-invariants.md counted modes in their Docker-cases and
Web-tabs paragraphs ("any of the five CLI backends", "never a sixth
SessionMode"). Both numbers were already stale before grok (antigravity and pi
had made it seven) and grok is now in the agent image, so the counts are gone
rather than incremented: the invariant those sentences carry is that Docker and
web tabs are not modes at all, which no number has ever helped state. The two
plan docs keep their original wording, being historical design records.
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
endpoints.md named 1.18.x as the version where workspace hooks became a
setting, but 1.18.x servers do NOT have this behavior — an agent driving
one would falsely conclude its workspace has hooks. The feature ships in
1.19.0. And preamble.sh changed content this PR without bumping its
CODEMAN_PREAMBLE stamp, so a cache stamped 1.18.3 would pass the
staleness check while holding old bytes; stamp bumped to 1.19.0 in
preamble.sh and the SKILL.md heredoc together (byte-identity pin).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The workspace-hooks install makes the skill's central hooks rule wrong in the
cautious direction. Six places told a worker that a linked case or a raw
workingDir has no `stop`/`blocked` and that send-and-wait cannot be trusted
there, so an agent would hand-roll output-marker synchronization in exactly the
workspaces where `wait:true` now works.
Rewritten against the setting rather than directory provenance:
- verbs.md §5.1: the where-to-spawn table, the rule paragraph (now naming
`workspaceHooksEnabled`, default ON, the add-only merge, and the boot sweep of
recovered sessions), and the silent-failure warning. The three cases that stay
hook-less regardless are called out: remote SSH sessions, docker cases that
opted out, and a workspace Codeman cannot write to.
- verbs.md §5.3: the send-and-wait precondition is "the workspace has the hooks
block", not "a case Codeman created".
- endpoints.md: the Signals-by-mode table is now keyed on the setting, with rows
for OFF, for remote/docker-opt-out, and for a session from an older server.
The old create-path grep list becomes a "before 1.18.x" note.
- SKILL.md §2 + the cost list, recipes.md Flow-1 contrast, messaging.md step 1.
"Check, do not assume" is kept and promoted to the load-bearing habit, because
the setting is not visible from the call and a session created by an older server
that has not restarted still has nothing.
The `spawn_worker` hooks grep STAYS: it guards the setting being off, remote
sessions, and older servers. Only its diagnostic changes, since "pick an unused
name" is no longer the fix. That text lives in both the §0 heredoc and
`preamble.sh`, which `test/agent-skill.test.ts` pins byte-identical, so both are
patched with the same bytes.
Docs only, no behavior change. 23 skill tests green, full test:ci 5109 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two live failures from one root cause: Claude Code loads a same-named
user-level skill (~/.claude/skills/codeman, written once by `codeman skill
install`) over the fresh per-case copy, and nothing ever refreshed it. A
stale Aug-9 copy (pre fast-path, pre lineage header) made every agent-driven
spawn run the old recipes: workers spawned serially with pid polls and
without X-Codeman-Parent-Session, so the web UI drew no lineage arcs.
- refreshUserAgentSkill(): session create now refreshes a marker-owned
user-level copy (refresh-only: absent copies are not installed,
foreign/symlink copies stay untouched).
- seedAgentSessionPreamble(): local claude session create pre-seeds the
skill's preamble into ${XDG_CACHE_HOME:-~/.cache}/codeman-agent-<id>.sh,
single-sourced from the new skills/codeman/preamble.sh, so the skill's §0
bootstrap collapses to a two-line loader instead of a ~150-line paste the
model has to type out (measured ~47s of generation per run).
- SKILL.md: §0 now leads with the loader and keeps the full block as the
stale/missing fallback; explicit verbatim-paste warning (a hand-assembled
preamble is how the header and the fast-path functions got lost);
spawn_worker also sends parentSessionId in the body as defense in depth;
preamble stamp bumped to 1.18.3 so pre-fix cached preambles self-heal.
- test/agent-skill.test.ts pins preamble.sh byte-identical to the SKILL.md
heredoc and covers seeding (XDG + HOME fallback, 0600) and the user-level
refresh (absent/stale/foreign).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The changeset ships this as 1.18.2, so the stamp, the bootstrap's grep/write
condition, both re-source guards and the recipes guard all carry 1.18.2 now
instead of a version that would never exist.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fifteen review findings on the fast-path rewrite plus one caught live, all
verified against a real 1.18.1 server before landing:
- sendwait picks a fresh seq (the epoch second) instead of a fixed 2, so a
second prompt to the same worker is typed instead of silently swallowed as
an already-applied duplicate; explicit seq remains for deliberate resends
- sendwait self-heals stranded delivery: an Ink repaint occasionally eats the
Enter (observed live), so a timed-out short first wait sends one bare \r and
re-waits by resending the identical frame as a tagged duplicate
- spawn_worker verifies the resolved casePath carries Codeman hooks (the same
/api/hook-event marker the server checks), refusing names that resolve to
linked or pre-existing hook-less directories instead of running the job in
what may be the user's real repo
- spawn_worker probes the trust dialog after a short 5s composer wait, not the
full 45s, restoring the ladder staging verbs.md documents; on a readiness
miss it deletes the half-spawned session and returns 1 with empty stdout,
so a prompt can never be typed blind into a trust dialog
- spawn_workers refuses duplicate case names and empty argument lists, and
keys result files by index
- section 1 is bash 3.2 compatible (indexed arrays, no declare -A), prints the
full delivered/timedOut/signal tuple per worker with an explicit line for a
missing result, deletes only workers whose turn really ended (a timeout
means still working), cleans up spawned siblings when any spawn fails, and
guards its mktemp
- last_text takes the previous answer as an optional second argument for
consecutive-turn reads (the transcript briefly serves the prior answer
after a stop, observed live)
- the stale duplicate bullets in section 1's closing list are gone
- reference/verbs.md joins the mode-list drift guard's file list
- README's skill inventory covers verbs.md and the new SKILL.md shape
- the changeset is minor so the shipped release matches the 1.19.0 stamp
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured against a live 1.18.1 server, the API does the whole job in about ten
seconds: two cold claude workers spawned and ready in 6.3s, both tasked and both
answers read in 4.0s more. The slowness users reported was agent-side.
Three causes, all of them things the skill taught:
- It taught serial spawning. Nothing in the main document showed `&`/`wait`, so
"spawn two workers" read as "do the readiness ladder twice", which is one model
turn per worker.
- It had no spawn primitive. The happy path had to be reassembled on every run from
where-to-spawn, a four-stage readiness ladder, send-and-wait, the fan-out caveats
and a recipe with two variants. Each is a decision, and most carry a warning.
- It cost ~16k tokens before the first call, at 3.6:1 prose to code, with 25 warning
glyphs and 55 occurrences of "never". A document that is mostly failure modes
teaches caution, and caution bills as thinking tokens.
The preamble now defines the verbs rather than describing them: spawn_worker,
spawn_workers (concurrent), sendwait, last_text. Section 1 composes them into the
whole job in one Bash call and says to stop reading there.
Two ceremonies the measurements retired: the pid poll (one iteration, 33ms, and
wait-output already blocks on the composer) and reading settings.local.json to check
hooks for a case quick-start creates, which always has them. That check stays
required for linked cases and raw paths, where its absence silently breaks
send-and-wait.
The bootstrap's write condition now greps the version stamp, so a stale or truncated
preamble self-heals rather than failing and asking for a manual rm. The stamp line is
kept bare because the grep anchors on it with $; an inline comment there would rewrite
the file on every bootstrap.
Section 5 moved to reference/verbs.md behind an index, cutting the always-paid
SKILL.md from ~16.4k to ~7.6k tokens. Section numbers and anchor slugs are unchanged,
so existing references still resolve; all 201 anchors across the five files were
checked, with the checker positive-controlled against an injected bad link.
Verified by extracting the code blocks from the shipped file and running them against
the live server: bootstrap plus full fast path, two workers resolving on the
definitive stop signal, answers read and sessions deleted, in 6.8s.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four documentation defects found while analysing the agent skill against the
code it drives.
The lineage section attributed "deletes its session as soon as the one-shot
prompt returns" to `POST /api/v1/sessions/:id/run`. That is true of
`POST /api/v1/run`, which creates a throwaway session and calls cleanupSession
on both the success and the error path; the per-session route deletes nothing.
Name the right endpoint, and give the real reason the per-session one carries
no lineage: it is not a create call.
While verifying that, the per-session route turned out to be a sharper trap
than documented. `runPrompt()` rejects whenever a PTY already exists, which is
every interactive session, but the route has already returned `{}` with HTTP
200 by then and routes the rejection only to SSE. An agent calling it against
a live worker reads the 200 as delivery. Document it.
`Flow 3b` never existed in recipes.md. The real mapping is Flow 3 = shell
fan-out, Flow 4 = claude fan-out, Flow 5 = worker blocked on a prompt, so the
same sentence was also mislabelling Flow 4. Fixed in SKILL.md and in the
endpoints.md reference to it; every other Flow reference audited and correct.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`GET /api/pi/status` shipped undocumented in the agent skill, and only a human
reading the doc noticed. Turns out none of its five siblings were documented
either, so this adds the whole family in one place: spawning with a mode whose
CLI is absent fails with OPERATION_FAILED rather than falling back, which is
exactly what an agent picking a backend it did not choose needs to know. Pi's
extra `.data.version` is called out, since a false `available:false` there means
an unrelated `pi` is in front on PATH.
On whether the endpoint scanner should also check registered-to-documented:
measured, and NO for the general case. The skill documents 34 of 217 registered
endpoints deliberately (it is an agent guide, not an API reference), so a blanket
reverse check needs a 183-entry allowlist that would fail CI on unrelated route
work and get appended to mechanically, which is worse than the gap it closes.
Grouping by path shape does not save it either: the families that yields are
things like `DELETE /api/<any>/:id`, lumping cases, webviews and docker hosts
together, and it would not have caught this gap anyway (the family had zero
documented members).
What IS cheap is a family the schema can enumerate with no allowlist: the new
assertion derives the agent modes from the Zod enum and requires each one's
`/api/<mode>/status` to be documented, so a seventh backend fails here until it
is. The sibling scanner still proves the other direction, that nothing documented
is a 404. Both mutation-checked: dropping pi's probe fails the new guard, and
documenting a nonexistent probe fails the old one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Second review pass on #282, the three items left open after f4dcfbe.
1. `codeman doctor` and the run mode disagreed about pi. The registry entry
accepted a bare `which pi` hit while pi-cli-resolver demanded semver-shaped
`--version` output, so the Dependencies panel could report an installed Pi CLI
on a box where Run Pi stays hidden, which reads as a broken mode rather than a
missing install. Both sides now share one exported PI_VERSION_REGEX, and
PathResolver gains an opt-in `requireVersionMatch` so a binary that fails the
shape check is reported MISSING instead of installed-with-unknown-version.
Only pi sets it; every other tool keeps its current behaviour.
2. The isAltScreenStripMode comment justified excluding pi with "the alt screen
is load-bearing for its fullscreen TUI". That is not what exclusion does: pi
is tmux-backed, so it falls through to isMuxAltScreenOnlyStripMode, which
strips the alt-screen toggles anyway. What exclusion actually preserves is
`\x1b[3J` and the mouse DECSETs, which is the real reason (pi renders into the
main screen and is mouse-aware). Comment and changeset now say that, and state
the consequence: fullscreen pi paints into the main buffer, like vim in a tmux
shell session.
3. skills/codeman still enumerated the five pre-pi modes in nine places, telling
agents a backend does not exist and understating class-wide caveats by one
mode. All updated, plus stale session.ts line references refreshed.
Tests: a new static guard derives the mode set from the Zod schema (not a copy)
and fails when a skill enumeration lists a partial set of external CLIs, verified
by mutation. It also documents the one legitimate exception it found: the "writes
no transcript" lists drop codex, which does write a rollout Codeman reads back.
Plus doctor cases for an unrelated `pi` on PATH and registry/resolver regex parity.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.
Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
sources: pending approval dialog, user goals, last assistant turn tail,
recent prompts, tool activity, git workspace signals, away context,
sibling sessions, rethink state; 30 KB budget, whole-section drop from
the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
keeps only a 500-char snippet) and git signal collection (execFile,
2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
session, opus by default (readMyMindModel setting), strict JSON
contract with 1-3 suggestions (continue / verify / redirect), newline
stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
prediction in flight per session (409 CONFLICT), rethink body
{ steer, rejected }; ownership via findSessionOrFail
Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
until readMyMindEnabled is ON, desktop only (phone key is phase 3);
modal with editable suggestion + rationale and Send / Insert /
Rethink / Dismiss; suggestion text rendered via value/textContent only
and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
strings
Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals
plus the user's recently submitted prompts, captured from the Claude session
transcript behind the new synced readMyMindEnabled setting (default OFF).
- intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps,
consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json
- transcript-watcher.ts: new transcript:user_prompt event for typed user turns
(tool_result-only entries stay silent); capture wiring in server.ts is
claude-only and gated on the setting per event
- readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership
via findSessionOrFail, strict Zod schema
- agent skill: SKILL.md recipe + endpoints.md rows so agents can read and
record intent (PUT replaces: read + merge; never delete unprompted)
- groundwork for the phase-2 predictor button; nothing is ever auto-sent
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Version-gated fail-closed at 2.1.224 (the cross-session-messaging release,
flag presence verified against that binary): an unknown or older CLI yields
a spawn command byte-identical to before, because claude aborts startup on
an unknown option and that would kill every session spawn. The value is
allowlist-sanitized ahead of the double-quoted interpolation, and only the
local command carries the flag; docker/remote builders never see it since
their CLI is not the probed binary. Verified E2E on an isolated instance:
cmdline shows --name, ListAgents lists the session name, replies arrive
tagged from-name.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude Code v2.1.224+ gives sessions ListAgents/SendMessage and a per-session
inbox socket. Codeman's claude workers are ordinary local Claude Code sessions,
so the agent skill now teaches task delivery and result collection over
messaging where available (multi-line exactly-once messages, mid-turn steering,
latched replies), with the HTTP primitives keeping spawn, readiness,
synchronization, liveness and delete, and a bounded fallback to the HTTP
recipes whenever the feature is absent. All mechanics verified live against
claude-cli 2.1.226.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured on a live claude worker: `GET /api/v1/sessions/:id` reported
`status: "idle"` while the worker was mid-turn and actively producing output, with
`lastActivityAt` equal to the moment of the call. The skill already warned that a
worker which dies inside its pane also reads `idle`, so the field is unreliable in
both directions and nothing an agent does should depend on it.
Synchronize on `stop` via send-and-wait or on an output marker. To judge from
outside, sample `terminal?tail=` twice a few seconds apart: a changing buffer is the
only cheap positive proof a worker is still working. `wait?until=exit` stays the
death check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The readiness gate matched `bypass`, which is the status bar of ONE permission mode.
`buildPermissionArgs()` also spawns `--permission-mode auto`, `--allowedTools` and
plain `normal`, and the mode is not exposed on `GET /api/v1/sessions/:id`, so an agent
cannot know which token to expect. A non-default worker was therefore reported broken
after burning the whole ladder.
Measured one pane per mode against claude-cli 2.1.226:
--dangerously-skip-permissions -> "bypass permissions on"
--permission-mode auto -> "auto mode on"
--allowedTools Read,Grep -> "don't ask on"
(none, normal) -> "don't ask on"
--permission-mode plan -> "plan mode on"
Every one ends `(shift+tab to cycle)`, so `shift+tab` is the single space-free token
that means "the composer is up" in every mode, and it is what the ladder matches now.
Verified live end to end on a virgin case: stage 1 misses while the trust dialog is up,
stage 2 accepts it, stage 3 matches in 623ms.
⚠️ `shift+tab` contains a `+`, so it only works through `--data-urlencode`. In a
hand-built query the `+` decodes to a space and the server searches for `shift tab`,
which never appears; the response echoes `match: "shift tab"`, which is how to spot it.
Measured both ways. The stage-4 fallback (make the worker echo a split token, proving
readiness by answering rather than by chrome) stays as the last resort, and is now also
verified live: it matched in 2.5s, with the token surviving the space-less TUI intact.
Also portable ANSI stripping: the read pipelines used `sed 's/\x1b...'`, and BSD sed
(the macOS default) has no `\xHH` escape, so on macOS the strip silently removed
nothing and handed the agent raw ANSI. They now build a real ESC with `printf`.
And endpoints.md gaps: the `FORBIDDEN` 403 row and which auth responses are plain text
rather than the JSON envelope, the input size cap, the undocumented `killMux` parameter
on DELETE, and the fact that zero/negative/non-integer timeouts are rejected with a 400
rather than clamped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.
Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
Case names resolve through linked-cases.json first, mirroring the
server's resolveCasePath(), so a case linked in from outside
~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
skill is never touched, and a symlinked skill dir is refused (this
repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
ports/config-port.ts, server.ts, session-routes.ts (add-only injection
on Claude session create and quick-start), plus the App Settings toggle.
Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
Shell state does not survive between agent tool calls, and an undefined
is_self exited 127, firing the `||` branch and deleting the caller's own
session with the one guard bypassed. The request now lives inside the
guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
tool call, so the documented resend-identical-request loop stopped being
a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
workers. It returns clean transcript text; the terminal scrape it
replaces returns a wall of TUI repaint noise. Its transcript flush lags
the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
yielded the literal session id "null" and burned the whole readiness
budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
corrected the hooks-config comment that claimed a toggle-off sweep
exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
and that caseName resolves linked cases, so a generic name can land a
worker in a real repo.
Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>