Measured on a live claude worker: `GET /api/v1/sessions/:id` reported `status: "idle"` while the worker was mid-turn and actively producing output, with `lastActivityAt` equal to the moment of the call. The skill already warned that a worker which dies inside its pane also reads `idle`, so the field is unreliable in both directions and nothing an agent does should depend on it. Synchronize on `stop` via send-and-wait or on an output marker. To judge from outside, sample `terminal?tail=` twice a few seconds apart: a changing buffer is the only cheap positive proof a worker is still working. `wait?until=exit` stays the death check. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
24 KiB
name, description
| name | description |
|---|---|
| codeman | Drive Codeman, the session manager this agent is running inside, over its HTTP API: list sessions, start worker sessions, send them prompts, block until they finish (wait / wait-output / send-and-wait), read their output, and clean up. Use when asked to orchestrate or parallelize work across Codeman sessions, watch another session, or start and manage workers. Only usable inside a Codeman-managed session (CODEMAN_MUX=1); refuse to act otherwise. |
Driving Codeman from inside a session
You are an agent running inside a Codeman-managed terminal session. Codeman is the server that spawned you; its HTTP API can start, prompt, watch, and delete other sessions. Every recipe below was verified live. Full endpoint tables and troubleshooting: reference/endpoints.md. Worked multi-worker flows: reference/recipes.md.
0. Guard, and the one thing that breaks every recipe below
⚠️ Your shell state does not survive between tool calls. Each Bash call starts a
fresh shell, so $API, $SELF, the CURL array and delete_session are all gone by
the next call, and $$ is a different pid. Three consequences, all of which have
teeth:
- Re-run this entire preamble at the top of every Bash call that touches the API. Running it once and assuming it stuck is the single most likely way to break a run.
- Never re-paste only half of it. The delete guard below is written so that a
missing definition deletes nothing, but that only holds if you never hand-roll a
DELETEof your own. - Never put
$$in aclientId. It changes per call, so the "resend the identical request" loop in §3 would stop being a duplicate and would retype the prompt, submitting the turn twice. Use a fixed literal (codeman-agent-1below).
Only real environment variables (CODEMAN_*) survive, which is why this preamble
rebuilds everything else from them.
test "${CODEMAN_MUX:-}" = 1 || { echo "Not inside a Codeman-managed session; refusing to act."; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Codeman does NOT hand a session the server password. If one is set, the two
# in-reach copies are the data dir's .env (the same fallback `codeman attach`
# uses — hand-authored; nothing ever writes it) and the supervisor definition
# that install.sh wrote the password into, which is where a stock
# password-protected install actually keeps it. The data dir is wherever the
# hook-secret file lives. Values may be quoted or `export`-prefixed.
ENV_FILE="${CODEMAN_HOOK_SECRET_FILE:+${CODEMAN_HOOK_SECRET_FILE%hook-secret}.env}"
envval() { sed -n "s/^\(export \)\{0,1\}$1=//p" "$ENV_FILE" | tail -1 | sed 's/^"\(.*\)"$/\1/; s/^'\''\(.*\)'\''$/\1/'; }
if [ -z "${CODEMAN_PASSWORD:-}" ] && [ -n "$ENV_FILE" ] && [ -f "$ENV_FILE" ]; then
CODEMAN_USERNAME=$(envval CODEMAN_USERNAME)
CODEMAN_PASSWORD=$(envval CODEMAN_PASSWORD)
fi
if [ -z "${CODEMAN_PASSWORD:-}" ]; then # stock installs: install.sh puts it in the service definition
UNIT="$HOME/.config/systemd/user/codeman-web.service"
PLIST="$HOME/Library/LaunchAgents/com.codeman.web.plist"
if [ -f "$UNIT" ]; then
# install.sh backslash-escapes " and \ in the unit value; undo it or a password
# containing either recovers wrong and auth fails.
CODEMAN_PASSWORD=$(sed -n 's/^Environment="CODEMAN_PASSWORD=\(.*\)"$/\1/p' "$UNIT" | head -1 | sed 's/\\\(["\\]\)/\1/g')
elif [ -f "$PLIST" ]; then
# install.sh XML-escapes the plist value; undo it (& LAST, mirroring escape order).
CODEMAN_PASSWORD=$(awk '/<key>CODEMAN_PASSWORD<\/key>/{getline; print}' "$PLIST" | sed -n 's/.*<string>\(.*\)<\/string>.*/\1/p' \
| sed -e 's/</</g' -e 's/>/>/g' -e 's/&/\&/g')
fi
fi
AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:$CODEMAN_PASSWORD")
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on https (self-signed cert)
# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
# `is_self "$SID" || curl -X DELETE ...` shape failed OPEN, because an undefined
# is_self exits 127 and the `||` branch then ran the delete completely unguarded.
# Undefined delete_session is "command not found", which deletes nothing.
delete_session() {
local id="${1:-}"
[ -n "$id" ] || { echo "refusing: empty session id"; return 1; }
[ "${#SELF}" -ge 8 ] || { echo "refusing: \$SELF unset or too short to prove this is not me"; return 1; }
# ids appear in full AND 8-char form (Docker exports a truncated $SELF; mux names and
# UI surfaces carry 8-char ids), so compare by prefix in BOTH directions. Equality or
# a one-directional check each miss a real combination, and the miss deletes you.
case "$id" in "$SELF"*) echo "refusing: $id is me"; return 1 ;; esac
case "$SELF" in "$id"*) echo "refusing: $id is me"; return 1 ;; esac
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$id"
}
CID=codeman-agent-1 # FIXED literal, never "agent-$$" (see §0)
- If
CODEMAN_MUXis not1, stop and say so. Do not guess an API URL; a server you are not part of is not yours to drive. - A 401 is plain text, not the JSON envelope, so on a password-protected server
every
jqin these recipes dies withjq: parse errorinstead of showingUNAUTHORIZED. If that happens, check the status with-w '%{http_code}'; if it is 401 and neither fallback above found a credential, stop and tell the user you need credentials. The hook-secret bypass covers only/api/hook-eventand/api/status-telemetry, never session control. - These endpoints first ship in Codeman 1.13.0, but do not gate on the version
number: a dev build can serve them while reporting an older version. Probe
instead:
GET .../waiton a real session id answering 404 with an.errorstartingRoutemeans the server predates the wait endpoints (fall back to pollingGET .../terminal?tail=and say so);Session ... not foundmeans your session id is wrong, not the server.
1. Safety rules — read before any mutating call
You are yourself a session on this server, and the API has no undo.
- Never act on your own session, and know that
delete_sessionis the ONLY guard. The server has no self-protection: a session that DELETEs its own id succeeds and dies silently (verified live). Always delete throughdelete_session "$SID"from §0; never write a barecurl -X DELETEand never reintroduce theis_self … || curl -X DELETE …shape. That older form failed open: with the function undefined (a half-re-pasted preamble, see §0) bash returns 127, the||branch fires, and the delete runs with no self-check at all. Wrapping the request inside the guard is what makes a lost preamble delete nothing instead of deleting you. Apply the same prefix-both-directions reasoning before any kill, respawn, or input call you write by hand. - Mutating calls you may make unprompted (this is an allowlist):
POST /api/v1/quick-start,POST /api/v1/sessions/:id/input, andDELETE /api/v1/sessions/:idonly for a session you created in this conversation, by exact id. Keep a list of the ids you create. Everything else mutating needs the user to have asked for it. - Never call these unless the user explicitly asked, naming the target:
DELETE /api/cases/:name— recursively deletes a real directory of the user's code from disk. One wrong case name destroys work that was never yours.DELETE /api/sessions(no id) andDELETE /api/subagents(no id) — bulk kills.- respawn / ralph / orchestrator / cron mutations — respawn runs
/clear(wipes a conversation), orchestrator state is a single global slot, cron jobs outlive you. PUT /api/settings,POST /api/system/update— global UI settings; server restart.
- Never
tmux kill-session,pkill tmux,pkill claude. The API is the only interface. - Sessions count against a 50-session cap and case creation is uncapped: clean up every
session you start, and don't retry
quick-startin a loop.
2. Rules of the road
-
End every input with
\r— literally the two characters\rinside the JSON string. Codeman types the text and sends Enter only when the input contains a carriage return; without it your command sits unsubmitted on the worker's prompt and everything downstream times out.{"input":"run the tests\r",...}. No response field catches this:delivered:truemeans "written to the pane", not "submitted" — a\r-less send still reportsdelivered:trueand then every wait times out, which is why the loops below are bounded and check the terminal. -
Single-line input only. Newlines are stripped; one line per call.
-
Build request bodies with
jq -nfor any prompt you did not author as a literal. The inline-d '{"input":"'"$P"'\r"}'pattern breaks on the first double quote, backslash, or$in a real prompt:BODY=$(jq -n --arg p "$PROMPT" '{input:($p+"\r"),useMux:true,clientId:"agent-1",seq:1,wait:true,waitTimeout:60000}') "${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' --data-binary "$BODY" -
Exactly-once delivery: always send a stable
clientIdand a monotonic per-sessionseqonPOST .../input. A retry after a dropped connection then cannot double-type the prompt. Incrementseqfor each NEW input; reuse the same pair only to re-ask about the same delivery. -
Envelope: success is
{"success":true,"data":…}, errors are{"success":false,"error","errorCode"}. Read.data. Use/api/v1/*paths. -
A wait timeout is HTTP 200,
{wait:{timedOut:true,signal:null}}— not an error. Loop over short waits (60 s); proxies cut long-idle connections. Timeouts are clamped (ceiling 600 s): read backwait.timeoutMsfor what was applied. The clamp covers positive integers only:0, a negative, a fraction or30sis a 400, so round any computed remainder and drop it entirely rather than sending zero. -
Never branch on
.data.status. It is a heuristic and is often wrong in both directions: measured on a live claude worker readingidlewhile it was mid-turn and actively producing output (lastActivityAtequal to the moment of the call), and a worker that died inside its pane also readsidle. Synchronize onstopvia send-and-wait, or on an output marker. To judge from outside, sampleterminal?tail=twice a few seconds apart: a changing buffer is the only cheap positive proof a worker is still working.wait?until=exitis the death check. -
stopandblockedfire forclaudesessions only (Claude Code hooks). Onshell/opencode/codex/gemini/antigravity, requesting them explicitly is a 400 — and lifecycle transitions there are coarse (a short shell command may emit noidletransition at all, verified live), so synchronize those modes with output markers, not signals. -
Your typed command echoes into the output stream, so a marker that appears verbatim in the input line matches before the command runs. Always split the marker (recipe below), keep it unique per call, and use
from=bufferso a marker that printed before your wait landed is still found. Matching is literal — no regex. -
Match single space-free tokens against TUI output. A full-screen TUI (claude, codex, …) positions text with cursor movements, not literal spaces, so the stripped stream can read
Yes,Itrustthisfolderand a multi-word match is unreliable there — whether a phrase keeps its spaces depends on how the TUI happened to draw it (observed live: some match, some never fire). Plain command output (shell workers,echolines) keeps real spaces.
3. Recipes (each verified live)
List sessions / find yourself — metadata only, safe to poll:
"${CURL[@]}" "$API/api/v1/sessions" | jq '.data[] | {id, name, mode, status}'
"${CURL[@]}" "$API/api/v1/sessions" | jq --arg s "$SELF" '.data[] | select(.id | startswith($s))'
Start a claude worker and wait until it is actually ready. A new session reports
idle before its CLI has spawned, and a brand-new case shows a trust dialog
first, so neither "wait for idle" nor "wait for ❯" means ready (the trust dialog
contains ❯ too — observed live). Codeman can auto-accept that dialog itself, but
the accept rides a stream match that misses on some runs (both outcomes seen live),
so wait for the composer first and handle the dialog only as the bounded fallback —
never send a blind Enter up front (if auto-accept already fired, it lands in the
composer). Stage 1 is short on purpose: an already-trusted case matches shift+tab in
under a second, while a virgin case can never pass stage 1 (the dialog is up, so
the composer is not) and always pays it in full before the fallback runs — the long
budget belongs to stage 3, after the dialog is answered.
⚠️ Match shift+tab, never bypass. The permission mode is a server-side setting
(claudeMode) that is not exposed on GET /api/v1/sessions/:id, so you cannot read
which mode a worker runs. bypass permissions on is only the DEFAULT mode's statusline.
Measured against claude-cli 2.1.226, one pane per mode:
| how Codeman spawned it | statusline reads | shift+tab |
bypass |
|---|---|---|---|
--dangerously-skip-permissions (default) |
bypass permissions on |
yes | yes |
--permission-mode auto |
auto mode on |
yes | no |
--allowedTools … |
don't ask on |
yes | no |
neither (normal) |
don't ask on |
yes | no |
Every mode ends its status bar with (shift+tab to cycle), so shift+tab is the one
token that means "the composer is up" regardless of mode, and it is space-free, which is
what makes it survive the TUI stream. Matching bypass instead reports a perfectly
healthy non-default worker as broken after burning the full ladder.
⚠️ shift+tab contains a +, so it MUST go through --data-urlencode. In a
hand-built query the + decodes to a space and the server searches for shift tab,
which never appears (measured: matched:false, and the response echoes back
match: "shift tab", which is how you spot it).
Stage 4 stays as the last resort for the case where even that misses: a worker that answers a trivial prompt is ready, whatever its statusline reads.
# ALWAYS check .success: on failure `.data.sessionId` is null, jq -r prints the string
# "null", and the flow below then burns its full readiness budget against
# /api/v1/sessions/null before reporting jq noise instead of the actual cause.
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
if [ -z "$SID" ]; then
# SESSION_BUSY here is the 50-session cap, not the waiter cap; FORBIDDEN/CONFLICT/
# OPERATION_FAILED/INVALID_INPUT are the others. None are retryable in a loop.
jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed; stopping."
exit 1
fi
for _ in $(seq 1 30); do # bounded: a bad SID would otherwise poll forever
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# ⚠️ pid != null proves STARTUP only, never life: a worker that later dies inside
# its pane keeps status "idle" and a pid (the local tmux attach client, not the
# worker). The death check is wait?until=exit, below.
SEQ=1 # $CID came from the §0 preamble; do NOT rebuild it from $$
# stage 1-3: `shift+tab` is the composer's status bar in EVERY permission mode (see the
# table above), so this works whatever `claudeMode` the server runs. Single-token
# matches only: TUI text is space-less. The `+` needs --data-urlencode.
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# composer never appeared → the trust dialog is probably still up; accept it once
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
fi
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# stage 4, last resort: the composer never appeared at all. A miss is still not proof
# of a broken worker, and answering is proof that it works. Split the token (your keystrokes echo
# into the stream) and keep it unique per call. This costs the worker one turn, so
# it runs only after the fast path missed. It must stay AFTER stage 2, which is the
# only thing that clears the trust dialog: free text plus \r into a dialog still up
# answers it blind, which is the same footgun as the up-front Enter.
TOK="${RANDOM}_$$"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"reply with the word READY immediately followed by _'"$TOK"' and nothing else\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=READY_$TOK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=60000' \
| jq -e '.data.wait.matched' >/dev/null \
|| echo "worker $SID never became ready; inspect terminal?tail="
fi
Send a prompt and wait for the turn to finish (claude workers — the call to
prefer). It registers the waiter before typing, closing the race where a separate
wait sees the previous turn's idle state. Loop by resending the identical request:
the repeat is a tagged duplicate (same clientId+seq) that does not retype but
answers from the session's current state. Verified: the stop hook resolves this in
seconds; a duplicate resend answers in ~20 ms without retyping.
for TRY in $(seq 1 10); do # BOUNDED: a \r-less send never produces a signal and resends are no-op duplicates
R=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"run the tests, then summarize in one line\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ',"wait":true,"waitTimeout":60000}')
if jq -e '.data.wait.timedOut' <<<"$R" >/dev/null; then
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # two straight timeouts: prompt sitting unsubmitted?
continue
fi
# Resolved — but a duplicate answering immediately reports the session's CURRENT
# state ("it is idle now"), NOT that a new turn ran. A \r-less send lands exactly
# here on try 2 (verified live), so check the terminal before believing it:
if jq -e '.data.duplicate and .data.wait.immediate' <<<"$R" >/dev/null; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" | jq -r '.data.terminalBuffer' | tail -5
# your prompt still on the ❯ composer line = never submitted (missing \r);
# submit it with {"input":"\r"} (the only recovery), then loop again
fi
break
done
SEQ=$((SEQ+1)); jq '.data.wait.signal, .data.status' <<<"$R"
Read the outcome in this order: wait.signal != null → done (stop is definitive;
idle is heuristic) — unless it arrived as duplicate:true + immediate:true,
which only says the session is idle now and must be confirmed from the terminal
(above); wait.timedOut → loop again (bounded); wait.ended → session gone, stop.
If the loop exhausts its cap, do not keep looping: read the terminal, report what
you see, and remember that a still-typed-but-unsubmitted prompt (missing \r) can
only be recovered by submitting it with {"input":"\r"}.
Shell worker + completion marker — the pattern for shell mode (no hooks there).
The typed line must not contain the marker verbatim (the input echo would match
instantly — observed live), so build it with a variable the worker's shell expands:
N="${RANDOM}_$$"; MARK="DONE_$N" # unique per call: tmux repaints replay old text
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}'
SEQ=$((SEQ+1))
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=120000' \
| jq -r '.data.wait | {matched, snippet}'
The typed line shows ${M}_…, the real output shows DONE_… rc=<exit code>, and the
snippet carries the exit code back to you.
Read a worker's answer. For claude and codex workers this is the read path:
last-response returns the agent's final message as clean text, taken from the
transcript rather than the screen, so it carries none of the TUI's box-drawing or
repaint noise.
for _ in $(seq 1 10); do # the transcript write LAGS the stop signal
TXT=$("${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
.data is {text, timestamp}. ⚠️ Poll it, do not read it once. text is written
from the transcript file, which is flushed slightly after the stop hook fires, so a
single read taken the instant send-and-wait returns comes back "" even though the
turn finished (verified live: empty on the first call, full text seconds later). text
is also "" before the worker's first completed turn, and always "" for modes with
no transcript (shell, opencode, gemini, antigravity, verified live), which is
why the loop above is bounded rather than open-ended. Fall back to the terminal buffer there, tail in bytes
(textOutput in GET .../output stays empty for interactive sessions; don't use it):
# \x1b is a GNU-sed extension: BSD sed (macOS) matches it as a literal "x1b", so the
# same one-liner strips NOTHING there and hands you raw ANSI. Feed sed a real ESC.
ESC=$(printf '\033')
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=3000" | jq -r '.data.terminalBuffer' \
| sed -e "s/${ESC}\[[0-9;?]*[a-zA-Z]//g" -e "s/${ESC}([B0]//g" | grep -v '^[[:space:]]*$' | tail -30
⚠️ Do not use that pipeline to read a claude/codex answer. A full-screen TUI draws
with cursor moves, so the stripped buffer is largely one long line: tail -30 has
almost nothing to split on and you get a wall of repaint noise with the answer buried
in it (verified live, side by side with last-response returning the exact prose).
The terminal buffer is for diagnosis (is my prompt sitting unsubmitted?), not for
reading answers. Avoid ?full=1 (entire tmux scrollback, a context bomb) unless doing
a post-mortem.
Detect a dead worker cheaply: GET .../wait?until=exit&timeout=60000 answers
immediately (signal:"exit", immediate:true) if the PTY is gone — including a
worker that exited inside its pane, which GET .../sessions/:id keeps reporting
as status:"idle" with a pid (that pid is the local tmux attach client, not the
worker). The wait routes are the only liveness check; a worker dying while a wait
is parked resolves it within ~3 s. A session deleted mid-wait resolves in ~1 s.
Clean up — only ids you created, one at a time, always through the §0 helper:
delete_session "$SID"
Everything else (endpoint tables, per-mode signal table, error codes, capacity limits, Docker/remote caveats): reference/endpoints.md. Fan-out orchestration and blocked-worker handling: reference/recipes.md.