Compare commits

..
Author SHA1 Message Date
Codeman maintainer 94aa53c65b chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:38:40 +02:00
Codeman maintainer e88b971bb7 feat(skill): add the agent-skill install layer and harden the packaged skill
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.

Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
  Case names resolve through linked-cases.json first, mirroring the
  server's resolveCasePath(), so a case linked in from outside
  ~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
  hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
  skill is never touched, and a symlinked skill dir is refused (this
  repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
  ports/config-port.ts, server.ts, session-routes.ts (add-only injection
  on Claude session create and quick-start), plus the App Settings toggle.

Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
  Shell state does not survive between agent tool calls, and an undefined
  is_self exited 127, firing the `||` branch and deleting the caller's own
  session with the one guard bypassed. The request now lives inside the
  guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
  tool call, so the documented resend-identical-request loop stopped being
  a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
  workers. It returns clean transcript text; the terminal scrape it
  replaces returns a wall of TUI repaint noise. Its transcript flush lags
  the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
  yielded the literal session id "null" and burned the whole readiness
  budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
  corrected the hooks-config comment that claimed a toggle-off sweep
  exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
  and that caseName resolves linked cases, so a generic name can land a
  worker in a real repo.

Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:30:15 +02:00
Codeman maintainer 8406c497e2 fix(terminal): stop forwarding the wheel to codex, it ignores SGR reports
DodgyBadger reported a completely dead wheel in codex tabs (#227 comment)
while the scrollbar drag worked, and the [scroll] line confirmed the
branch: forward-sgr with 967 rows of healthy local scrollback unused.

Measured against codex-cli 0.147.0 in a bare tmux: codex never enables
mouse tracking (mouse_any_flag=0), runs an inline viewport
(alternate_on=0) and pushes its transcript into the terminal's own
scrollback (history_size grows), and SGR wheel reports written to its
pane change nothing at all. Hand-encoded SGR taps are no-ops too, so
they stay (harmless), which means click-to-position is merely
unavailable there rather than damaging.

_shouldForwardWheelToApp now returns true for claude >= 2.1.187 and
nothing else; codex falls to the local-scrollback path like
shell/gemini/opencode, which is the same history the scrollbar drag was
already reaching. The claude-only PageUp fallback is untouched.

Verified in Chromium against a live codex session on an isolated
instance: routing logs local-scrollback, the viewport moves 39 -> 4 and
zero bytes go to the PTY.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:22:06 +02:00
Codeman maintainer 40b4aba043 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:35:59 +02:00
Codeman maintainer 4b44988bfc test: give daemon-control tests a unique port (3212 was already taken)
test/sse-subscription-filter.test.ts already binds 3212; sequential test
execution hid the clash. Moves the probeServer fixture to 3216 (3217 for
the nothing-listening case) per the unique-port convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:22:01 +02:00
Codeman maintainer 316d0a4c82 Merge pull request #233 from Lint111/feat/hooks-config
Conflict in refreshStaleCodemanHooks resolved by keeping every staleness
trigger: the master-side TLS-flagless curl check (hooks without -k) AND the
PR-side current-wake-marker (V3) + SubagentStop guard marker checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:21:52 +02:00
Ark0N 1184720648 Merge pull request #239 from Ark0N/feat/daemon-mode
feat(cli): codeman web -d and codeman service install (#231)
2026-08-09 02:19:50 +02:00
Ark0N b067aad9b6 Merge pull request #235 from Lint111/feat/deferred-terminal-flush
fix(terminal): drain deferred output without a wake event
2026-08-09 02:19:32 +02:00
Ark0N 19a3d7c773 Merge pull request #234 from Lint111/feat/ai-checker-stderr
fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
2026-08-09 02:19:10 +02:00
Codeman maintainer 085f4acb60 feat(cli): codeman web -d and codeman service install (#231)
Two ways to keep the server running, split by how long it should last.

`codeman web -d` relaunches the same entry script detached (setsid), with
`--stop` and `--status` alongside it. A pidfile and log live in the data
dir. `nohup` is not what makes this work: Node re-arms SIGHUP to its
default disposition even when it inherits "ignore", and cli.ts handles
SIGHUP with a graceful shutdown, so a delivered HUP still stops the
server. Removing the shell's ability to send one is the fix.

`codeman service install|uninstall|status` writes and loads the systemd
user unit or the LaunchAgent, with the installing shell's PATH baked in
(launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, which finds neither a
Homebrew/nvm node nor tmux/claude). install.sh already covers one-liner
installs; this is for npm globals.

Both refuse to start when a server is already up on the data dir, since a
second instance on the shared tmux socket attaches PTYs to the first
one's live sessions. Both poll /api/status until the child answers or
dies rather than reporting a success they have not seen. `--stop` checks
the pid still looks like a Codeman server before signalling it.

The systemd unit name and launchd label move to config/service-names.ts
so install.sh, detectSupervisor() and service install cannot drift into
supervising two copies. Instance-scoped, unchanged for the default
instance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:34:55 +02:00
Codeman maintainer d26f26fe34 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:18:03 +02:00
lior 091df2b6d8 fix(terminal): drain deferred output without a wake event 2026-08-08 23:00:36 +03:00
lior 5f775b1ab1 fix(hooks): guard subagent stops and rewake from the parent transcript
Two defects in the background-task hook scripts.

SubagentStop had no handler at all. When a subagent launched background work and
one watcher ended while others were still running, Claude could publish the
worker's last progress sentence as its final result, abandoning the live tasks.
A new guard pairs launched task IDs against completed ones and confirms liveness
by scanning /proc/<pid>/fd for an open tasks/<id>.output handle, blocking the
stop only while genuinely-live work remains. It fails open — allowing the stop —
when /proc is unavailable, nothing was launched, or everything finished.

The rewake helper watched only input.transcript_path. A subagent has its own
transcript, but Claude writes the completion queue-operation to the PARENT
transcript, so the record it waited for never appeared and the wake never fired.
It now watches both paths, but only when the relationship is provable: the
transcript's parent directory is subagents/ and its grandparent basename equals
input.session_id. It also now requires operation === 'enqueue'.

The rewake marker moves V2 -> V3; refreshStaleCodemanHooks treats absence of the
current marker as stale, so existing cases self-heal on next launch (the same
mechanism as the V1 -> V2 bump). Ownership matches on marker PREFIXES, so a
future bump still recognises older Codeman handlers and never adopts a user's.

12 tests fail on unmodified master, e.g.
  expected '[{"matcher":"Bash",…' to contain 'CODEMAN_BACKGROUND_REWAKE_V3'
  expected 'Background command bg-report-1 comple…' to contain '<codeman-background-result>'
2026-08-08 22:31:38 +03:00
lior da51193264 fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
AiCheckerBase spawned the check with `> out 2>&1`, so anything the Claude CLI
wrote to stderr landed inside the same file the verdict parser reads. A CLI that
failed to start (corrupt settings, missing auth) produced either an empty verdict
or an unparseable one, and the actual cause was destroyed on the way through —
the user saw only "Empty output from AI idle check".

stderr now goes to its own temp file. When output is empty or the verdict cannot
be parsed, the first 200 characters of stderr are appended to the error message.
The file is cleaned up alongside the existing temp files, including on the error
paths.

Two tests, both failing on master:
  expected 'export PATH="…' to contain ' 2> "'
  expected 'Empty output from AI idle check' to contain 'Claude CLI failed to load settings'
2026-08-08 22:30:39 +03:00
Codeman maintainer fa18eeef35 feat: tab action icons on the active tab only, middle-click closes tabs
Rework of the previous hover-overlay approach after feedback: sliding the
title under incoming icons made names hard to read, and icons appearing
under the cursor caused accidental gear/close clicks while switching tabs.

Now the gear/pop-out/close icons expand in flow on the ACTIVE tab only.
Selection is a deliberate click, so the strip's geometry never changes
while the pointer is aiming at a tab; hovering a background tab changes
nothing (the full title stays readable) and a stray click can only switch
sessions. Middle-click closes any tab (session tabs via the existing
close-confirm modal, web tabs via closeWebviewTab), matching browser
muscle memory so background tabs still close in one action.

The pop-out button stays opt-in via App Settings -> Tab Bar (per-device
showTabDetachButton, default off), and a detached tab keeps its icon as
the re-focus affordance. Phone layouts already used the active-only
pattern; tablets keep their always-visible touch fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer a9f26bd03a feat: fixed-width tab hover with sliding title, pop-out button now opt-in
Hovering a session tab no longer grows it. The three per-tab icons now
live in a .tab-actions wrapper that overlays the tab's right edge on
hover-capable devices: the icons slide in while the title (and any
badges) slide left by a per-tab --tab-slide distance computed in
_applyTabHoverSlide(), clipped at the left edge of .tab-info so the
readable tail (the :comment suffix) stays visible. Keyboard focus
reveals the overlay via :has(:focus-visible), so a mouse click on the
gear does not pin it open. Touch devices keep the previous in-flow
behavior (the wrapper adds no width in flow, and the legacy tap-reveal
rules are preserved under @media (hover: none)).

The open-in-a-new-window (pop-out) button is now hidden by default and
opt-in via App Settings -> Tab Bar -> "Pop-out Button on Tabs"
(showTabDetachButton, per-device, absent from SettingsUpdateSchema like
the other display keys). A tab whose session is already detached keeps
its icon as the re-focus affordance regardless of the setting.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer 8dc8b164a7 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:48:12 +02:00
Ark0N 2524759655 Merge pull request #229 from Lint111/feat/keyboard-viewport-settle
fix(mobile): coalesce keyboard viewport settling
2026-08-08 12:51:04 +02:00
Codeman maintainer 1f164bc8d2 fix(mobile): only arm the viewport settle on a real keyboard transition
A visualViewport resize event without a pending show/hide transition now
only pushes a pending settle back (_deferViewportSettle) instead of arming
fit + PTY-resize work of its own. Keyboard detection can miss a
fine-grained OS animation entirely (each step under 150px, with the
baseline chasing the animation down), while MobileDetection's own listener
still shrinks --app-height, so the per-event settle fitted xterm against a
mid-animation container with no keyboard CSS compensation and resized the
PTY to transient dims. The resulting SIGWINCH thrash (58 -> 10 -> 50 rows)
duplicated prompts and left tmux dot filler in the transcript on keyboard
close. Reproduced with a faked visualViewport driving the real handler;
master is unaffected because it never resized the PTY from this path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 12:06:08 +02:00
lior 0a1439b1e9 test(mobile): make the coalescing test actually exercise the settle path
The suite never selects a session, so initTerminal() does not run and both
`app.terminal` and `app.fitAddon` are null at rest. `_scheduleViewportSettle`
returns early on a falsy terminal, so the coalescing assertions could not
reach the behavior they claimed to cover -- the test errored on
`Cannot read properties of null` rather than measuring anything.

Installs the minimum surface the settle callback touches and restores it
afterwards, so the coalescing path executes for real.

Adds a behavioral counterpart driven through the PUBLIC entry point
(`onKeyboardShow`) instead of the internal scheduler: three viewport steps
in quick succession must produce exactly ONE refit. On master that returns
3 (each show arms its own uncoalesced 150ms timeout), so this fails by
COUNT rather than by a missing method -- which is the failure mode that
actually demonstrates the bug.

Verified: `expected 3 to be 1` on unmodified master; passes here. The
remaining 8 failures in this file are pre-existing on master and unrelated
(same null-initialization limitation of the headless harness).
2026-08-08 08:41:04 +03:00
lior 66abe6c70a fix(mobile): coalesce keyboard viewport settling 2026-08-08 08:13:58 +03:00
Codeman maintainer fa1700da5b chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 01:38:35 +02:00
Ark0N aed1e59ee3 Merge pull request #227 from Ark0N/fix/scrollback-205-round2
fix(terminal): scrollback round 2 for #205 (re-pull downgrade guard, PageUp fallback, CLI version probe retry)
2026-08-08 01:36:03 +02:00
Ark0N 7f6d18b398 Merge pull request #226 from christianhaberl/fix/input-loss-on-failed-delivery
fix(api,ws): an input whose delivery fails can be retried instead of being lost
2026-08-08 01:31:10 +02:00
Ark0N 52571c7fd4 Merge pull request #225 from christianhaberl/fix/bound-the-process-tree-walk
fix(mux): bound the process-tree walk — unbounded pgrep recursion can take a machine down
2026-08-08 01:31:00 +02:00
Ark0N cb95a8562c Merge pull request #224 from christianhaberl/fix/raw-writehead-drops-security-headers
fix(http): raw writeHead routes drop every header the security hook set
2026-08-08 01:30:47 +02:00
Codeman maintainer 3cb7e30636 fix(ui): scope the wheel-opt-out tooltip's paging fallback to Claude
The reworded tooltip promised the PageUp/PageDown fallback for Claude and
Codex alike, but _localScrollbackIsHollow() gates it to claude mode only
(codex page-key handling is unverified, as the routing tests note). A codex
user reading the old text would flip the setting expecting a rescue and get
a dead wheel instead. Say plainly that Codex has no fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 01:29:48 +02:00
Codeman maintainer 9dc4620f03 fix(terminal): stop the scroll-to-top re-pull from deleting history, page the CLI when local scrollback is hollow (#205)
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.

1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
   terminal and rewrites it from the capture, which is a win when tmux holds
   more than the browser, but for a repaint-mode pane that capture is roughly
   ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
   pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
   all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
   (escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
   rewrite when it is more than one screen short; a refused session's cooldown
   goes from 4s to 60s so a hollow pane stops re-fetching megabytes.

2. A false forwarding gate on a Claude session no longer means a dead gesture.
   Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
   travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
   SGR reports, at half a screen of travel per page key. Shift is excluded: it
   keeps meaning "local scrollback".

3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
   exception and guarded on !== undefined, so one timed-out or PATH-starved
   probe at the first Claude session start disabled wheel-forwarding for every
   Claude session until the server restarted, which fits a report of breakage on
   phone, tablet and laptop at once. Success is still cached for the process
   lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
   a pure function so the semantics are testable without spawning claude.

4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
   scoping: the setting keeps meaning exactly what it says, and fix 2 catches
   the case where "local" is empty. The App Settings tooltip now says to leave
   it off for Claude/Codex sessions.

5. _logScrollRouting() prints one line per session per distinct decision:
   forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
   mode, cliVersion, the opt-out state, mouse tracking and local scrollback
   depth. #205 ran two rounds of remote guesswork over questions that line
   answers directly.

Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:41:39 +02:00
Claudia 9d27cc0bab docs: merge the stacked doc comments the previous commits left behind
Cosmetic, but the kind that quietly costs: JSDoc tooling attaches only the
nearest block, so a stacked second block silently hides the first.

- write() had two: the original description with @param and @example, then a
  @returns-only block added on top, which dropped the params and examples from
  hover. Merged into one. The @returns wording is also honest now — write() still
  discards the data without a PTY; what changed is that it says so.
- forgetInputSeq had been inserted BETWEEN shouldApplyInput's detailed doc comment
  and its declaration, leaving that function undocumented on hover and the doc
  attached to the wrong thing. Moved below.
- The mock kept an orphaned one-line comment above failWrites' own block.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 15:29:46 +02:00
Claudia 84132d3025 fix(ws): pin the withheld ACK with a test, and correct the changeset
Two blockers from the pre-submission gate, both reproduced before fixing.

1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
   code says the opposite in as many words ("NOT an error response,
   deliberately"), the commit message says response codes are unchanged, and the
   test asserts the 200. It was a leftover sentence from an earlier iteration that
   would have shipped into the CHANGELOG announcing an API contract change that
   does not exist — and errorCode values are SemVer-relevant per
   docs/versioning-policy.md.

2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
   master left all 9 tests green, while the commit message sells "plus the whole
   WebSocket path" as part of the fix. Three tests added against the real WS
   route — ACK on delivery, ACK withheld and seq re-opened when the write did not
   land, and a deduplicated frame still ACKed so the client can drop it. Verified
   the other way round: with ws-routes.ts reverted, the middle one fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 14:52:44 +02:00
Claudia ebfcac6ad1 fix(api,ws): an input whose delivery fails can be retried instead of being lost
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.

When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.

- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
  is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
  client redelivers.
- `Session.write()` reports whether it reached a PTY.

Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.

What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.

9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:36:33 +02:00
Claudia 2e69e28e71 fix(mux): bound the process-tree walk — it can take a machine down
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Two further sites forked a `pgrep` per session on
every stats tick.

Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep`
blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned
while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in
D-state out of ~39,000 total, load average above 13,000, and a machine only
recoverable by restarting WSL — which cost every running session. Every diagnostic
command timed out too, because they read /proc as well.

- ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously
  with a single-flight guard. Async matters: under the same procfs pathology,
  `execSync`'s timeout cannot return (spawnSync waits for the unkillable child),
  which would freeze the server where a hung async poll only costs staleness.
- The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a
  visited set (a stale snapshot can contain a cycle), a depth cap and a node cap,
  both reporting when they truncate. Pure so the regression tests can exercise the
  shipped code rather than a copy of it.
- The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor
  re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there
  would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is
  bounded, so a wedged `ps` cannot stop killSession from reaching its
  process-group and tmux fallbacks.
- Any `ps` error keeps the previous snapshot instead of caching partial output as
  fresh; a truncated table would make whole subtrees invisible to the kill path.

13 tests, including one that drives TmuxManager itself — with the caps bypassed at
the call site, 3 of them fail. The snapshot refresh is stubbed there, because
otherwise the manager runs a real `ps`, replaces the fixture, and the test
silently measures the machine's own process tree instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:34:47 +02:00
Claudia 1a32e63765 fix(http): raw writeHead routes lost every header the security hook set
`reply.raw.writeHead()` writes straight to the Node response and bypasses
Fastify's header store, so everything the `onRequest` security hook granted is
silently dropped on every route that answers that way.

The visible symptom is CORS. The hook emits `Access-Control-Allow-Origin` for
localhost origins, so a page served from a local dev server may call every `/api`
endpoint cross-origin — except the four below, whose requests fail. The security
headers (`X-Content-Type-Options`, `X-Frame-Options`, CSP) were being lost the
same way.

Affected: `GET /api/events`, and `file-raw` / `tail-file` / `download` in
file-routes.ts. Each now spreads the inherited headers first and lets its own
headers win over them.

Tests drive a real WebServer and compare `/api/events` against `/api/status` for
the same Origin — the point of the fix being that the SSE route stops being the
odd one out. Verified in both directions: with the fix removed, 3 of the 5 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:31:50 +02:00
71 changed files with 12453 additions and 393 deletions
+207
View File
@@ -1,5 +1,212 @@
# aicodeman
## 1.14.1
### Patch Changes
- The Codeman agent skill is now installable, so an agent running inside a Codeman session can drive the API without you pasting docs into its prompt. Plus six fixes to the packaged skill, each found by running it live against a real instance.
## What the skill is
`skills/codeman` is a Claude Code skill that teaches an agent inside a Codeman session how to start worker sessions, send them prompts, block until they finish, read their answers and clean up. It ships in the npm package. It self-gates: outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act, so installing it globally costs unrelated sessions nothing.
## Installing it
Three ways, pick one:
```bash
codeman skill install # ~/.claude/skills/codeman, every new Claude Code session sees it
codeman skill install --case myproject # just that case; linked cases resolve by name too
codeman skill uninstall # reverses either one
```
Or turn on **App Settings > Agent Skill** (`agentSkillEnabled`, synced, default off) and Codeman injects the skill into each case when a Claude session is created there.
Installs are marker-owned: a `skills/codeman` that Codeman did not write is never touched, a stale managed copy is refreshed in place, and a symlinked skill directory is refused rather than written through. Re-run `codeman skill install` after upgrading Codeman to refresh the copy.
Note that turning `agentSkillEnabled` back off does **not** remove already-injected copies, because a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` directory. Remove them per case with `codeman skill uninstall --case <name>`.
## Using it
Once installed, just ask: "spin up three workers and have them lint, typecheck and test in parallel, then report back". The skill supplies the guard, the safety rules and the recipes. What it does under the hood:
**1. Guard.** Every Bash call re-runs a preamble that refuses outside `CODEMAN_MUX=1`, reads `CODEMAN_API_URL` and `CODEMAN_SESSION_ID`, recovers a password from the data dir `.env` or the install's service definition if one is set, and defines a fail-closed `delete_session`. It re-runs it every call because shell state does not survive between an agent's tool calls.
**2. Start a worker.**
```bash
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
```
`mode` is any of `claude`, `shell`, `opencode`, `codex`, `gemini`, `antigravity`.
**3. Wait until it is actually ready.** A new session reports `idle` before its CLI has spawned, and a brand-new case shows a trust dialog first, so the skill waits for the composer's own status bar and treats the dialog as a bounded fallback.
**4. Send a prompt and wait for the turn to end.**
```bash
BODY=$(jq -n --arg p "$PROMPT" '{input:($p+"\r"),useMux:true,clientId:"codeman-agent-1",seq:1,wait:true,waitTimeout:60000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' --data-binary "$BODY"
```
Send-and-wait registers the waiter before typing, which closes the race where a separate wait reports the previous turn's idle state as this turn's answer. For `claude` workers it resolves on the `stop` hook, typically within seconds.
**5. Read the answer.**
```bash
"${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text'
```
**6. Clean up.** `delete_session "$SID"`, for ids you created and nothing else.
Hook-less modes (`shell` and the external CLIs) have no `stop` signal and coarse lifecycle transitions, so the skill synchronizes those with a unique split marker and `wait-output ... from=buffer` instead. Worked fan-out flows, the per-mode signal table, error codes and the Docker/remote caveats live in the skill's `reference/` files, loaded on demand.
## The rules that bite
The skill documents these because each one silently wastes a run:
- **Every input must end with `\r`** or Enter is never sent and the text sits unsubmitted on the worker's prompt. `delivered:true` means "written to the pane", not "submitted".
- **Input is single-line.** Newlines are stripped.
- **A wait timeout is HTTP 200** with `wait.timedOut:true`, not an error. Loop over short waits; timeouts clamp to [1s, 600s] and the applied value comes back as `wait.timeoutMs`.
- **`stop` and `blocked` are `claude`-only.** Requesting them elsewhere is a 400.
- **Signals are edge-triggered with no history.** One that fires while no waiter is registered is unobservable afterwards, so never fire-and-forget N prompts and then gather signal-waits worker by worker.
- **Your typed command echoes into the output stream**, so a marker that appears verbatim in the input line matches before the command runs. Split it.
- **A full-screen TUI stream is space-less**, so match a single space-free token, never a phrase.
- **`pid != null` proves startup, not life.** A worker that dies inside its pane keeps `status:"idle"` and a pid. `wait?until=exit` is the death check.
## Fixes to the packaged skill
- **The self-delete guard failed open.** The old `is_self "$SID" || curl -X DELETE ...` shape meant an undefined `is_self` exited 127, the `||` branch fired, and the agent deleted its own session with the one guard bypassed. That is reachable because shell state does not survive between tool calls, so a partially re-pasted preamble was enough. The DELETE now lives inside a fail-closed `delete_session`, which also refuses an empty id and refuses when `$SELF` is unset or too short to prove the target is not the caller.
- **`clientId` was built from `$$`.** The pid changes between tool calls, so the documented "resend the identical request" loop stopped being recognized as a duplicate and retyped the prompt, submitting the turn twice. It is a fixed literal now.
- **`GET /api/v1/sessions/:id/last-response` was undocumented.** It returns the agent's final message as clean transcript text; the terminal scrape the skill previously recommended returns a wall of TUI repaint noise with the answer buried in it. It is now the documented read path for `claude` and `codex`, with the terminal buffer demoted to diagnosis and hook-less modes. Because the transcript flush lags the `stop` signal, the recipes poll it instead of reading once.
- **`quick-start` responses were never checked for `.success`.** On failure `.data.sessionId` is absent, `jq -r` prints the string `null`, and the flow burned its full readiness budget against `/api/v1/sessions/null` before reporting jq noise instead of the cause.
- **`codeman skill install --case <name>` could not resolve a linked case.** It hardcoded `~/codeman-cases/<name>` while the server resolves through `linked-cases.json` first, so it failed with "Case not found" for a case the web UI handled fine.
- **Documentation corrections**: `SESSION_BUSY` on `quick-start` is the 50-session cap rather than the waiter cap; `caseName` resolves linked cases, so a generic name can land a worker in a real repo; and the claim that a toggle-off sweep exists was wrong, so the per-case `skill uninstall` cleanup is now stated in both the README and the code.
## Also in this release
- **Terminal**: the wheel is no longer forwarded to codex, which ignores SGR mouse reports.
## 1.14.0
### Minor Changes
- Daemon mode and service install, plus subagent hook hardening and terminal/idle-checker fixes.
**New: run Codeman in the background without a terminal (#239, closes #231)**
- `codeman web -d` starts the server detached: it survives closing the shell, logs to `~/.codeman/web.log`, records a pidfile, and only reports success after the server actually answers `/api/status` (a port clash or missing dependency can never read as a clean start). `codeman web --status` and `codeman web --stop` manage it; `--stop` verifies the pid still looks like a Codeman server before signalling, so a recycled pid is never SIGTERMed.
- `codeman service install` / `status` / `uninstall`: installs a systemd user unit (Linux) or LaunchAgent (macOS) so the server comes back after reboots. The unit carries the installing shell's PATH (launchd's default PATH finds neither an nvm/Homebrew `node` nor `tmux`/`claude`), never contains `CODEMAN_PASSWORD`, and uses the same instance-scoped unit names as `install.sh` and the self-updater so no second copy can end up supervised.
- Both refuse to start a second server on one data dir (pidfile check plus a live probe): two servers on the shared tmux socket would attach to each other's sessions.
- Why `-d` exists at all: `nohup` does not protect a Node process, Node re-arms SIGHUP even when it inherits "ignore", so `nohup codeman web &` still dies on HUP. The detached relaunch (setsid) removes the controlling terminal instead.
**Subagent background-work hooks (#233, thanks @Lint111)**
- The background Bash rewake helper now also watches the top-level parent transcript when the hook fires inside a subagent: Claude records a subagent's Bash result in its own `subagents/agent-*.jsonl` but queues the completion in the lead session transcript, so subagents previously never woke. It can also inline a `CODEMAN_RESULT_BEGIN/END` marked report (up to 64 KiB) from the task output file into the wake feedback.
- New SubagentStop guard: a subagent that still owns live Monitor or background Bash processes is kept working instead of publishing an intermediate progress line as its final report. Ownership is verified against live process descriptors on `tasks/<id>.output`, so stale transcript text alone never blocks, and the guard fails open on systems without `/proc`.
- Existing cases self-heal to the new hooks on next launch.
**AI idle checker: stderr kept out of the verdict (#234, thanks @Lint111)**
The `claude -p` verdict command no longer merges stderr into the verdict file, where CLI warnings could turn a valid verdict into a parse error. On failures, the first 200 chars of stderr are attached to the diagnostic instead.
**Terminal: large final batches drain fully (#235, thanks @Lint111)**
A render-scheduling flag was cleared after the flush instead of before it, so when a large batch left a remainder behind, the remainder stayed unrendered until unrelated output arrived. This looked like truncated responses or shell commands that never finish. The flush now reschedules itself until the queue is empty.
**Docs and tests**
- README documents daemon mode and service install.
- Unique test port for the daemon-control suite.
## 1.13.0
### Minor Changes
- Agent wait primitives, the Codeman agent skill, a fix for hooks dying silently on HTTPS installs, and the tab-strip UX improvements from the previous batch.
**Agent wait primitives (new API surface, the reason this is a minor).** Three bounded long-polls let an agent driving Codeman from a shell block instead of poll:
- `GET /api/v1/sessions/:id/wait` blocks until a lifecycle signal fires (`until=stop,idle,working,blocked,exit`, `fresh=1` to require a new transition).
- `GET /api/v1/sessions/:id/wait-output` blocks until a literal substring appears in the session's output (`match=`, `nocase=`, `from=now|buffer`; never regex, by design).
- `wait`/`waitTimeout` on `POST /api/v1/sessions/:id/input` (send-and-wait) registers the waiter before typing, closing the race where a separate wait reports the previous turn's idle state as this turn's answer.
Shared semantics: a timeout is HTTP 200 with `wait.timedOut: true` (callers loop over short waits; tunnels cut idle connections), timeouts are clamped to [1s, 600s] and echoed back as `wait.timeoutMs`, all three nest the result under `data.wait`, and `status`/`limitPaused` ride along. `stop`/`blocked` exist for `claude` mode only: requesting them explicitly elsewhere is a 400, the default set silently narrows and echoes what it waited on. Capacity caps (16 waiters per session, 128 process-wide) answer 409/429, waiter slots release on client hang-up, and shutdown resolves parked waiters instead of stranding them. Bounds are operator-tunable via `CODEMAN_WAIT_*` env vars.
Reliability details that came out of three verification rounds: a worker that dies inside its tmux pane is now detected at the mux layer (pane-death probe, ~750ms cache, a 3s watcher for waits already parked), so a corpse answers `exit` instead of `idle` and send-and-wait rolls back its dedup seq when the write went nowhere; output matching normalizes charset-designation escapes (a stock bash prompt's `ESC ( B` no longer breaks `match=tnode:`) and holds back partial escapes at chunk boundaries, so matches straddling PTY chunks are found.
**Codeman agent skill (`skills/codeman`).** A packaged skill that teaches an agent running inside a Codeman session to drive the API safely: guard preamble (refuses outside `CODEMAN_MUX=1`, resolves credentials from the data dir `.env` or the install's service definition), self-protection (`is_self` prefix check in both directions), readiness for claude workers (composer-first, trust dialog as bounded fallback), send-and-wait loops that cannot report a never-submitted prompt as success, marker-synchronized shell flows, fan-out patterns, and cleanup discipline. Ships in the npm package via the `files` entry.
**Hooks were dying silently on every HTTPS install (bug fix).** The generated hook curls lacked `-k`, so on `--https` installs (self-signed cert) every hook event (`stop`, `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `teammate_idle`, `task_completed`) failed TLS verification and the failure was swallowed, taking respawn's definitive idle signals with it. Hooks are now generated with `curl -sk`, and a staleness detector regenerates the on-disk hook config of already-created cases the next time a session starts in them. Relatedly, `CODEMAN_API_URL` is no longer exported with a guessed `http://localhost:3000` fallback (wrong scheme on HTTPS installs); it is omitted unless the server has stamped the real URL, so in-session guards fail closed.
**Tab strip (from the previous batch, reported by christianhaberl):** action icons (kill/pop-out) now appear on the active tab only, middle-click closes a tab, tab hover uses a fixed width with a sliding title instead of resizing the strip, and the pop-out button is opt-in (default off).
**Docs.** `docs/api-reference.md` gained the full long-polling contract (signals by mode, readiness, what the matcher sees, response discriminators); `docs/extending-codeman.md` and the README carry verified copy-paste orchestration recipes; `docs/architecture-invariants.md` records the load-bearing ordering, liveness, and edge-triggered-signal invariants. Net +163 tests (4300 passing in the CI sweep).
## 1.12.2
### Patch Changes
- Codex input fixes: all four bugs reported by @DodgyBadger traced to one root cause (the zero-lag local-echo overlay buffering keystrokes until Enter, which starves codex's per-keystroke composer) and fixed in terminal-ui.js:
- Slash command picker never appeared in codex sessions (#222): the "/" sat in the overlay until Enter, so codex never saw it. Codex-mode sessions now use plain PTY echo (same branch as shell), so the picker pops and live-filters as you type.
- Arrow keys dead while typing, backspace dead after Ctrl+Backspace (#218): arrows were forwarded to a still-empty composer while typed text sat pending, and after a control-char flush the overlay swallowed every backspace. Codex bypasses the overlay entirely now; the shared overlay branch (claude/gemini/opencode) additionally flushes pending text on composer nav keys, then hands the session to pass-through until Enter/Ctrl+C, and forwards backspace instead of swallowing it when the overlay has no state.
- Pasting displaced the typed prompt (#219): bracketed pastes (xterm terminal.paste with DECSET 2004 active) were forwarded without flushing pending typed text, so the paste landed first. The shared branch now flushes typed text first and delays the paste sequence by 80ms, because codex's paste-burst handling drops keystrokes that arrive in the same PTY read as a bracketed paste (verified against codex 0.147.0 at the byte level).
- Long prompts overflowed the bottom of the screen (#220): long typed prompts existed only in the overlay DOM so codex never grew its composer; with plain PTY echo the composer grows and rewraps normally.
Verified end to end against a real codex 0.147.0 TUI driven by a headless browser: the pre-fix build reproduces all four bugs, the fixed build passes 17/17 assertions. New CI test file test/local-echo-codex-gating.test.ts (41 tests) pins the nav-key classifier, per-mode overlay gating, the flush helper, and pass-through routing. Known upstream limitation: Ctrl+Backspace deletes one character, not a word (xterm.js sends 0x08; word-delete needs kitty CSI-u encoding that xterm.js 6.0.0 cannot emit).
Mobile keyboard viewport settling fixes by @Lint111 (#229): coalesce keyboard viewport settling so rapid visualViewport resize events during keyboard show/hide no longer thrash the terminal fit, and only arm the settle logic on a real keyboard transition instead of every viewport resize.
## 1.12.1
### Patch Changes
- Terminal scrollback fixes, round 2 of issue #205. A Claude pane's local buffer is hollow (tmux keeps no history for a repaint-mode pane), and both retest reports traced back to that fact. The scroll-to-top full-history re-pull now refuses to rewrite the terminal when the capture holds less than the browser already does, so it can no longer delete history mid-scroll on iPhone (a refused session also re-fetches far less often). When wheel-forwarding is unavailable on a Claude session (version probe failed, CLI older than 2.1.187, or the "Wheel Scrolls Local History" opt-out) and there is no local scrollback to scroll, wheel and touch now page the CLI's own transcript via coalesced PageUp/PageDown instead of doing nothing. The `claude --version` probe no longer caches a failed run for the server's lifetime (one timed-out probe used to silently disable wheel-forwarding on every device until restart); failures retry with backoff. Every scroll gesture now logs a one-line `[scroll]` routing decision to the browser console for direct diagnosis, and the opt-out setting's tooltip explains that the paging fallback is Claude-only (Codex has none).
- 2e69e28: Bound the process-tree walk that could take a machine down.
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Across ~28 adopted tmux trees the fan-out exploded,
and because each `pgrep` blocks in the kernel while reading `/proc/<pid>/cgroup`
under WSL, none returned while the walk kept spawning more — ~13,000 `pgrep`
processes stuck in D-state out of ~39,000 total, load average above 13,000,
recoverable only by restarting WSL.
Now: one `ps` snapshot, breadth-first with a visited set, a depth cap and a node
cap, in a pure module (`proc-tree.ts`) that the regression tests exercise
directly. The snapshot is refreshed asynchronously, and the kill path forces a
fresh one so the SIGKILL escalation cannot re-read pre-SIGTERM state.
- ebfcac6: An input whose delivery fails can be retried instead of being lost for good.
Both input paths recorded the `(clientId, seq)` pair as applied and acknowledged
the frame _before_ knowing whether the write had landed — the POST route because
its mux write is fire-and-forget, the WebSocket handler because it ACKed
unconditionally. When the write then failed, the client dropped the frame from its
durable queue and the server rejected the retry as a duplicate: the reliable
delivery layer was guaranteeing exactly-once delivery of something that had never
been delivered.
The bookkeeping is now rolled back on failure and the WebSocket ACK withheld, so
the client redelivers. `Session.write()` reports whether it reached a PTY at all
instead of silently swallowing the data.
Response codes are unchanged: a session can legitimately have no PTY yet (created
but not started), so turning that into a failure status would be a contract change
of its own.
Note this does not remove the root cause: the POST still answers 200 before the
mux write is attempted, so a client that treats any 2xx as final still cannot
learn about that failure. Closing that would mean awaiting the tmux child in the
request path.
- 1a32e63: Routes that answer with `reply.raw.writeHead()` no longer drop the headers the
security hook set.
`writeHead` writes straight to the Node response and bypasses Fastify's header
store, so everything the `onRequest` hook granted was silently lost — including the
`Access-Control-Allow-Origin` it emits for localhost origins, and the
`X-Content-Type-Options` / `X-Frame-Options` / CSP headers. A localhost page could
therefore call every other `/api` endpoint cross-origin while its EventSource
failed CORS.
Affects `GET /api/events` and the three raw-writing routes in `file-routes.ts`
(`file-raw`, `tail-file`, `download`).
## 1.12.0
### Minor Changes
+12 -6
View File
@@ -74,7 +74,7 @@ When user says "COM":
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 1.12.0 (must match `package.json`)
**Version**: 1.14.1 (must match `package.json`)
## Project Overview
@@ -109,6 +109,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| CI-equivalent test sweep | `npm run test:ci` (full suite minus browser/perf — see Testing) |
| Production start | `npm run start` |
| Production logs | `journalctl --user -u codeman-web -f` |
| Detached server | `codeman web -d` (`--status`, `--stop`; pidfile+log at `dataPath('web.pid'/'web.log')`). ⚠ Refuses to start a 2nd server on one data dir — see Instance isolation |
| Install/remove the service | `codeman service install` / `status` / `uninstall` (systemd user unit on Linux, LaunchAgent on macOS; names from `config/service-names.ts`) |
**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 3 Playwright tests). Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing).
@@ -141,7 +143,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| Domain | Key files | Notes |
| ---------------- | -------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| **Entry** | `src/index.ts`, `src/cli.ts` | |
| **Entry** | `src/index.ts`, `src/cli.ts`, `daemon-control`, `service-installer`, `config/service-names` | The last three back `web -d` / `service install` |
| **Session** | `src/session.ts` ★, `session-manager`, `session-auto-ops`, `session-cli-builder`, `session-task-cache`, `session-order` (pure), `session-pty-exit-breaker`, `usage-limit-patterns`, `usage-telemetry`; `src/services/unified-session-service.ts` | Pure/unit-tested helpers are split out of `session.ts` on purpose |
| **Mux** | `src/mux-interface.ts`, `src/mux-factory.ts`, `src/tmux-manager.ts` ★ | |
| **Respawn** | `src/respawn-controller.ts` ★ + 4 helpers (`-adaptive-timing`, `-health`, `-metrics`, `-patterns`) | Read `docs/respawn-state-machine.md` first |
@@ -180,6 +182,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Input**: `session.writeViaMux()` for programmatic/curl input via tmux `send-keys -l` + `send-keys Enter`, single-line only. Interactive **browser** input goes through a durable **exactly-once** layer: a stable `clientId` + monotonic per-session `seq` persisted to localStorage until the server ACKs, so a dropped link cannot lose or double-deliver a prompt. `ws-connection-registry.ts` supersedes only same-TAB reconnects, so two tabs on one session coexist. → [architecture-invariants#input-delivery-and-ws-resilience](docs/architecture-invariants.md#input-delivery-and-ws-resilience)
**Agent wait primitives**: bounded long-polls so an agent driving Codeman from a shell can block instead of poll: `GET /api/sessions/:id/wait` (lifecycle signal), `GET /api/sessions/:id/wait-output` (literal substring, **never** regex) and `wait`/`waitTimeout` on `POST /api/sessions/:id/input`. Registry in `session-wait-registry.ts` (pure, no `Session` reference), bounds in `config/agent-wait.ts`. ⚠️ **A timeout is a 200** (`wait.timedOut`), never an error, so callers loop over short waits. ⚠️ `stop`/`blocked` come from Claude Code hooks and therefore fire for **`claude` mode ONLY** (`shell` installs none either); asking for one explicitly on another mode is a 400, the default set silently drops them. ⚠️ Send-and-wait registers the waiter BEFORE the write (a separate POST-then-wait races and reports the PREVIOUS turn), and both teardown paths must `notifySignal('exit')` BEFORE `cancelAll()`. ⚠️ Client-hangup abort listens on **`reply.raw`** guarded by `writableFinished`: on `req.raw`, `close` fires when the request BODY ends, which on a POST killed every send-and-wait instantly and no `app.inject()` test could see it. ⚠️ Worker liveness cannot come from `session.pid` — for a tmux session that is the local attach client, which outlives a worker dying inside its pane — so it is probed at the mux layer (`isPaneDead`, ~750 ms cache) on blocking waits only, never on the input hot path. ⚠️ Signals are edge-triggered with no history: one that fires with no waiter registered is unobservable afterwards, so gather fan-outs with send-and-wait or latched `wait-output` markers, never fire-and-forget-then-sequential-signal-waits. The primitives are packaged as the **`skills/codeman` agent skill**: installable via `codeman skill install [--case <name>]` / `skill uninstall`, or auto-injected into a case's `.claude/skills/` on Claude session create behind `agentSkillEnabled` (SYNCED, default OFF). Injection is ADD-ONLY at create, marker-owned (`applyAgentSkill` in `hooks-config.ts` never touches an unmarked user copy) and refuses symlinks (this repo's own `.claude/skills/codeman` is a symlink to the source, which the injector must never write through). → [architecture-invariants#agent-wait-primitives](docs/architecture-invariants.md#agent-wait-primitives), `docs/api-reference.md`
**Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
**Auto-resume on usage limit** (opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit, `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time and `SessionAutoOps` arms a timer for reset+2min, then sends Esc + `continue`. ⚠️ Respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected`), which is what prevents `/clear` from wiping the paused conversation. Claude-mode only. → [architecture-invariants#auto-resume-on-usage-limit](docs/architecture-invariants.md#auto-resume-on-usage-limit)
@@ -194,7 +198,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Docker cases**: a case can point at a **container**, with any of the five CLI backends running inside it. Like remote-SSH this is a **LOCATION OVERLAY on cases, never a sixth `SessionMode`**. Exactly one long-lived container **per case**, shared by all its sessions, so killing a session kills only that session's in-container tmux and **never** `docker stop` while siblings remain. The workspace is a real host dir bind-mounted at the **same absolute path**, which is what keeps file-routes/watchers on real host bytes and makes the in-container transcript projHash match the host. Credentials are **seeded** (RO mount, copied into the container once) rather than shared RW, so in-container CLIs never write refreshed tokens back to the host, and bind mounts are excluded from `docker commit` so exports stay secret-free. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** Config drift is detected via a label hash and a drifted launch is REFUSED rather than silently launched with stale config. ⚠️ On the loopback-only prod bind a container cannot reach 127.0.0.1, so in-container hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; otherwise idle detection falls back to output-based. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` (user guide), `docs/docker-cases-plan.md` (design)
**External CLI modes (OpenCode, Codex, Gemini, Antigravity)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. All four **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. → [architecture-invariants#external-cli-modes-opencode-codex-gemini](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini)
**External CLI modes (OpenCode, Codex, Gemini, Antigravity)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. All four **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. ⚠️ **The local-echo overlay is DISABLED for codex sessions** (`_updateLocalEchoState` in terminal-ui.js, same branch as shell): codex's composer reacts per keystroke ("/" pops a live-filtering picker, arrows edit server-side state, the composer grows as it wraps), so buffer-until-Enter starved it into issues #218/#219/#220/#222. Codex also **drops keystrokes that share a PTY read with a bracketed paste**, so flushed text and the paste sequence must go out as separate delayed writes (mirroring the Enter branch's delayed `\r`). Tests: `test/local-echo-codex-gating.test.ts`. → [architecture-invariants#external-cli-modes-opencode-codex-gemini](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini)
**Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w<n>-<case>` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization)
@@ -206,9 +210,11 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of EACH session per page load requests `full=1` (`_fullHistoryLoaded` Set); tab switches keep the cheap `?tail=` path, and scrolling up at the TOP of the buffer re-pulls `full=1` on demand (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of EACH session per page load requests `full=1` (`_fullHistoryLoaded` Set); tab switches keep the cheap `?tail=` path, and scrolling up at the TOP of the buffer re-pulls `full=1` on demand (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). ⚠️ That re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for codex/claude ≥ 2.1.187 at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
**Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for **claude ≥ 2.1.187 ONLY** at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. ⚠️ Codex was in that list and must never go back without a fresh measurement: codex-cli 0.147.0 ignores SGR wheel reports entirely (`mouse_any_flag=0`, inline viewport, transcript pushed into terminal scrollback), so forwarding produced a dead wheel (#227 follow-up). `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). ⚠️ When that gate is FALSE on a claude session whose local buffer is hollow (`baseY === 0`), the gesture becomes coalesced PageUp/PageDown key sends (`_maybePageCliTranscript`) instead of a no-op; ⚠️ and `getClaudeCliVersion()` must never cache a FAILED probe (one timeout used to disable forwarding process-wide until restart). `_logScrollRouting()` prints the routing decision and its inputs once per session — read it before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
**Detached start + service install** (issue #231): `codeman web -d` relaunches the SAME entry script with `detached:true` (setsid), so there is no controlling terminal and no shell job entry. ⚠️ `nohup` is NOT what makes this work: Node re-arms SIGHUP to its default disposition even when it inherits "ignore", and `cli.ts` handles SIGHUP with a graceful shutdown, so a delivered HUP still stops the server. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile check + `/api/status` probe): a second instance on the shared tmux socket attaches PTYs to the first one's live sessions. ⚠️ Neither may report success it has not observed — the parent polls `/api/status` until the child answers or dies, since `launchctl load` and a clean spawn are both silent about a server that starts and immediately exits. `--stop` verifies the pid still LOOKS like a Codeman server (`ps -o command=`) before signalling, because pids get recycled. Unit/label names live in `config/service-names.ts` so install.sh, `detectSupervisor()` and `service install` cannot drift into supervising two copies; they are instance-scoped, and identical to the historical names for the default instance. `service install` bakes the installing shell's PATH into the unit (launchd gives a job `/usr/bin:/bin:/usr/sbin:/sbin`, which finds neither a Homebrew/nvm `node` nor `tmux`/`claude`) and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install)
**Self-update** (App Settings → Updates): in-app updater for git-clone installs supervised by systemd/launchd (`systemd`, `launchd`, `launchd-daemon`, else `none` → "restart manually"). The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` that outlives the restart and writes progress to `update-status.json`, which the browser polls across the connection drop. `src/web/self-update.ts` splits pure helpers (unit-tested) from IO wrappers. npm installs report as non-updatable. → [architecture-invariants#self-update](docs/architecture-invariants.md#self-update)
@@ -292,7 +298,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
### API Routes
~199 handlers across 21 route files in `src/web/routes/`: system (45), sessions (32), cases (27), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
~200 handlers across 21 route files in `src/web/routes/`: system (45), sessions (34), cases (27), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
**HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`).
+104 -14
View File
@@ -85,9 +85,29 @@ codeman web --multiuser # named logins + per-user case spaces
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
<details>
<summary><strong>Run as a background service</strong></summary>
<summary><strong>Keep it running in the background</strong></summary>
The installer's final menu sets this up for you (option 2) and verifies the service actually comes up before claiming success. To configure it manually instead:
To outlive the shell you started it in, without setting anything up:
```bash
codeman web -d # detach; logs to ~/.codeman/web.log
codeman web --status # is it up, and on which pid
codeman web --stop # graceful SIGTERM; agents keep running in tmux
```
`-d` waits until the server actually answers before reporting success, and refuses to start a second one on the same data dir (two servers sharing a tmux socket attach to each other's sessions).
To have it come back after a reboot, install it as a service instead. The installer's final menu does this for you (option 2); `codeman service` is the equivalent for an `npm i -g aicodeman` install:
```bash
codeman service install # systemd user unit (Linux) or LaunchAgent (macOS)
codeman service status
codeman service uninstall
```
`service install` writes the unit with your current PATH baked in, which matters more than it sounds: launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin`, so a Homebrew or nvm `node`, `tmux` or `claude` is invisible to a hand-written plist. It never copies `CODEMAN_PASSWORD` into the unit file; add that yourself if the service needs auth.
To write the unit by hand instead:
**Linux (systemd):**
@@ -220,6 +240,8 @@ codeman web # localhost:3000 (loopback only — safe defau
codeman web --port 8080 # custom port (or set CODEMAN_PORT)
codeman web --https # self-signed TLS (only needed for remote access)
codeman web -H 0.0.0.0 # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
codeman web -d # detach: survives closing the shell (--status, --stop)
codeman service install # systemd/launchd service: comes back after reboots
```
Open the printed URL. The page is a single dashboard; everything below happens there.
@@ -268,6 +290,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
### 7. Operate & maintain
- **App Settings** — model, effort, permission startup mode, theme/skin, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
- **Run it in the background** — `codeman web -d` detaches from your shell (`--status`, `--stop`); `codeman service install` makes it a systemd user unit / macOS LaunchAgent that survives reboots. Both verify the server actually answers before reporting success, and both refuse to start a second server on one data dir. See [Keep it running in the background](#quick-start---installation).
- **Self-update** — git-clone installs update in place from **Settings → Updates**.
- **Deploy your own changes** — see [Development](#development).
@@ -403,6 +426,7 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
## More Features
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
@@ -667,6 +691,18 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
For AI agents and automation that control Codeman without a browser: an agent that spins up worker sessions, a CI bot, or **Claude Code running _inside_ a Codeman session orchestrating other sessions**. Everything the UI does is HTTP + a CLI, so an agent can do it too.
> **Shortcut: install the packaged agent skill.** Everything below (plus worked multi-worker recipes) ships as a Claude Code skill in [`skills/codeman`](skills/codeman/SKILL.md), so an agent inside a session can drive Codeman without you pasting docs into the prompt. Three ways to get it:
>
> - `npx skills add Ark0N/Codeman --skill codeman -g`: global, works for any skills-aware agent
> - `codeman skill install` (global) or `codeman skill install --case <name>`: for npm installs that never cloned the repo; `codeman skill uninstall` reverses it
> - **App Settings → Agent Skill** (`agentSkillEnabled`, default off): Codeman then injects the skill into each case on Claude session create; a user-authored `skills/codeman` in the case is never overwritten
>
> A global install (`codeman skill install`, or `npx skills add`) is picked up by **every new Claude Code session on the machine**, inside Codeman or not. The skill self-gates: outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act, so a global install costs an idle session nothing.
>
> ⚠️ Turning `agentSkillEnabled` back off **does not remove already-injected copies** (a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` dir). Remove them per case with `codeman skill uninstall --case <name>`.
### Detect that you're inside Codeman
When a CLI runs in a Codeman-managed session, these environment variables are set — read them instead of hardcoding anything:
@@ -680,15 +716,21 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
### Rules of the road (read before you POST)
1. **Single-line input only.** Programmatic input is sent as literal text **+ Enter** in one shot. Multi-line strings break the agent TUI (Ink) — send one line, or split into multiple calls.
1. **Single-line input, ending in `\r`.** Programmatic input is sent as literal text, and Enter fires **only when the input contains a carriage return**: `{"input":"run tests\r"}`. Without the `\r` the text sits on the session's prompt unsubmitted (and a combined `wait` runs its full timeout on a turn that never started). Embedded newlines are stripped rather than rejected, so `"echo A\necho B\r"` runs the joined command `echo Aecho B`: send one line per call.
2. **Make input idempotent.** Include a stable `clientId` and a monotonic per-session `seq` on `POST …/input`. The server de-duplicates, so a retry after a dropped connection can't double-deliver a prompt.
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard).
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard). ⚠️ A `401` replies with the bare string `Unauthorized`, **not** the JSON envelope, so piping it into `jq` throws a parse error instead of showing the failure: check the status before parsing.
4. **Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
5. **`/api/v1/*`** is a stable alias of `/api/*`.
6. **Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
### Recipes
```bash
# CODEMAN_API_URL is auto-set inside every Codeman session, correct scheme included.
# The fallback below fits a stock install; on a --https install set the https:// URL
# yourself and add -k to each curl (self-signed cert).
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (add -u admin:"$CODEMAN_PASSWORD" to each call if a password is set)
@@ -700,18 +742,63 @@ curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
# 2b. Wait until that worker is actually READY (see rule 8): composer marker first,
# first-run trust dialog only as the fallback. (Probing trust first and sending
# a blind Enter misfires on re-runs: the dialog text stays in the buffer forever,
# so the probe matches stale text and the Enter lands in a ready composer.)
# Match single tokens: TUI text can reach the matcher without its spaces.
until [ "$(curl -s "$API/api/sessions/$SID" | jq '.data.pid')" != null ]; do sleep 1; done
R=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=trust' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true}' # accept the first-run trust dialog
curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send a prompt into a session (exactly-once: clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures","useMux":true,"clientId":"agent-1","seq":1}'
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,"clientId":"agent-1","seq":1}'
# 4. Read the terminal back
curl -s "$API/api/sessions/$SID/output" | jq -r '.data // .'
# 4. Send a prompt and BLOCK until that turn is done (registers the wait before
# writing, so it can't answer with the previous turn's idle state)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,
"clientId":"agent-1","seq":2,"wait":"stop,exit","waitTimeout":60000}' \
| jq '.data.wait' # -> {"signal":"stop","timedOut":false,"waitedMs":41230,...}
# (`stop` is the definitive end-of-turn hook. Adding `idle` makes it resolve on a
# spinner pause too, and on anything that redraws a ❯ prompt — like a dialog.)
# 5. Stream live events (session output, agent activity, status)
# 4b. Timed out? That's a 200, not a failure. Loop over short waits.
curl -s "$API/api/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq '.data.wait'
# 4c. Or wait for a marker in the output (works for shell sessions too).
# ⚠️ Unique per call (tmux repaints replay old screen text), and SPLIT so the
# typed line never contains it: your own keystrokes echo into the output
# stream, so an unsplit marker matches before the command has run. from=buffer
# catches a marker that printed before the wait landed.
N=$RANDOM
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
curl -sG "$API/api/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
# 5. Read the terminal back. ⚠️ Use terminal?tail=, NOT /output: the latter's
# textOutput is empty for every tmux-backed (i.e. every interactive) session.
# tail counts BYTES, and what comes back is terminal data, ANSI included.
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
# 6. Stream live events (session output, agent activity, status)
curl -sN "$API/api/events" # Server-Sent Events
# 6. Schedule recurring work (cron-style job)
# 7. Schedule recurring work (cron-style job)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
@@ -719,11 +806,11 @@ curl -s -X POST "$API/api/cron/jobs" \
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
# 7. Inspect background sub-agents and their transcripts
# 8. Inspect background sub-agents and their transcripts
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
# 8. Whole-system snapshot (sessions, settings, respawn, stats)
# 9. Whole-system snapshot (sessions, settings, respawn, stats)
curl -s "$API/api/status" | jq
```
@@ -749,7 +836,7 @@ Codeman registers Claude Code hooks that `POST /api/hook-event` (`permission_pro
## API
REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
### Sessions
@@ -757,8 +844,11 @@ REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE str
| -------- | -------------------------- | ---------------------------------------------------------------------------------- |
| `GET` | `/api/sessions` | List all |
| `POST` | `/api/quick-start` | Create case + start session (`{caseName?, mode?, effort?, envOverrides?}`) |
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?}` — `clientId`+`seq` = exactly-once) |
| `GET` | `/api/sessions/:id/output` | Read terminal output |
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?, wait?, waitTimeout?}`: `clientId`+`seq` = exactly-once; `wait` blocks until the turn ends) |
| `GET` | `/api/sessions/:id/terminal` | Read terminal output (`?tail=<bytes>`, `?full=1`); the read path for interactive sessions |
| `GET` | `/api/sessions/:id/output` | Parsed one-shot output (`textOutput` is empty for tmux-backed sessions) |
| `GET` | `/api/sessions/:id/wait` | Block until a signal fires (`?until=stop,idle,exit&timeout=&fresh=`); a timeout is a `200` |
| `GET` | `/api/sessions/:id/wait-output` | Block until a literal string appears (`?match=&nocase=&from=now\|buffer&timeout=`) |
| `GET` | `/api/sessions/unified` | Unified live + history list (Session Manager) — `?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | Pin/unpin in the Session Manager (`{pinned}`) |
| `PUT` | `/api/session-order` | Sync tab order across devices (`{order: [ids]}`) |
+703
View File
@@ -0,0 +1,703 @@
# Agent Control Plan: skill packaging + wait primitives
**Status**: steps 1 to 6 IMPLEMENTED, uncommitted as of 2026-08-09. Steps 1 to 5 were
multi-round verified on 2026-08-08; step 6 (CLI install command + per-case injection +
`agentSkillEnabled`) was built 2026-08-09; see the step-6 entry at the end of
[§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
verification round found, and what is still open.
**Date**: 2026-08-08
**Scope**: Part 1 (agent skill) and Part 2 (wait primitives) were specified and built.
Parts 3 to 5 are captured so they are not lost, but remain deliberately deferred.
---
## 0. Where this came from: what herdr does
[herdr](https://github.com/herdrdev/herdr) (Rust, Apache-2.0, ~25.8k stars) is a terminal
multiplexer built around AI coding agents. Relevant findings from the research pass:
| Capability | How herdr does it |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Agent state | Four states (`idle`, `working`, `blocked`, `done`) that roll up pane to tab to workspace in a sidebar |
| Detection | Lifecycle hooks where the agent supports them (it names Pi and MastraCode), otherwise TOML manifests matched against a live bottom-buffer snapshot. Bundled manifests plus remote updates from herdr.dev, local overrides win |
| Control API | Newline-delimited JSON over a Unix socket (`~/.config/herdr/sessions/<name>/herdr.sock`), `{"id":"req_1","method":"pane.split","params":{}}`, dot-notation methods, plus long-lived event subscriptions |
| Discoverability | `herdr api schema` prints a machine-readable schema |
| Agent skill | `npx skills add herdrdev/herdr --skill herdr -g`, a SKILL.md wrapping the CLI, guarded by `test "${HERDR_ENV:-}" = 1` so an agent outside a herdr pane refuses to act |
| Persistence | Background server, detach with `ctrl+b q`, snapshot restore of workspaces/tabs/panes/cwd/layout, experimental screen-history replay, agent resume via native session ids, live PTY handoff across server replacement |
| Plugins | `herdr-plugin.toml` manifest, actions, event hooks, plugin panes, link handlers, GitHub-topic marketplace index |
The commands the skill teaches the agent:
| Group | Commands |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| workspace | `workspace list`, `workspace create` |
| tab | `tab list --workspace <id>`, `tab create` |
| pane | `pane current`, `pane list`, `pane layout`, `pane split --current --direction right --cwd <path> --no-focus`, `pane run <id> "<cmd>"`, `pane wait-output <id> --match/--regex <p> --timeout <ms>`, `pane read <id> --source visible\|recent\|detection` |
| agent | `agent list`, `agent start <name> --kind <type> --pane <id>`, `agent prompt <name> "<text>" --wait --timeout <ms>`, `agent wait <name> --until <state> --timeout <ms>`, `agent send-keys`, `agent get`, `agent read` |
### The honest comparison
herdr and Codeman are not the same product. herdr is a local, keyboard-first multiplexer with
no server, no web UI, and no autonomy layer. Codeman is a server with a browser and mobile UI,
remote and Docker cases, respawn, Ralph, cron, and the orchestrator, none of which herdr has.
What herdr genuinely does better is being **callable by the agent running inside it**. For
Codeman that is a packaging problem plus one missing primitive, not an architecture problem.
---
## 1. Gap analysis
| herdr capability | Codeman equivalent today | Gap |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- |
| `pane split` + `agent start` | `POST /api/quick-start`, `POST /api/sessions` | none, already there |
| `agent prompt` | `POST /api/sessions/:id/input` with `clientId`+`seq` exactly-once | no `--wait` |
| `pane read` | `GET /api/sessions/:id/output`, `GET /api/sessions/:id/terminal?full=1` | none |
| `agent list` / `agent get` | `GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status` | none |
| `agent wait --until <state>` | SSE only (`/api/events`) | **missing**, and SSE is impractical from a shell tool |
| `pane wait-output --match` | nothing | **missing** |
| Skill file | README section "Driving Codeman from an Agent" | **not packaged**, an agent will never find it |
| Env guard `HERDR_ENV=1` | `CODEMAN_MUX=1`, `CODEMAN_API_URL`, `CODEMAN_SESSION_ID` already exported at spawn | none, the guard variables exist |
| `blocked` state | hook events (`permission_prompt`, `elicitation_dialog`) plus CSS classes plus the phone overview NEEDS YOU section | not in the wire contract (`SessionStatus = 'idle' \| 'busy' \| 'stopped' \| 'error'`) |
| `api schema` | hand-written `docs/api-reference.md` | no machine-readable schema |
| Detection manifests | hardcoded in `usage-limit-patterns.ts`, `respawn-*-patterns`, `regex-patterns.ts` | patterns are code, not data |
| Plugin runtime | deliberately refused, see `docs/extending-codeman.md` | not a gap, a decision |
| Session handoff on restart | tmux owns the PTYs, so they already survive a Codeman restart | not a gap, solved by architecture |
**Conclusion**: roughly 90% of the capability surface already exists. Parts 1 and 2 below close
the two real gaps.
---
## 2. Part 1: the Codeman agent skill
### 2.1 Goal
An agent running inside a Codeman session can discover and correctly drive Codeman without the
user pasting API docs into the prompt, and without inventing dangerous calls.
### 2.2 Layout and distribution
The `npx skills` CLI (vercel-labs/skills) clones a GitHub repo and looks for
`skills/<name>/SKILL.md`. Claude Code natively discovers `.claude/skills/<name>/SKILL.md` in a
project and `~/.claude/skills/` globally. Both are satisfied with one source of truth plus a
symlink, which is the pattern this repo already uses for `remotion-best-practices`.
```
skills/
codeman/
SKILL.md <- single source of truth
reference/
endpoints.md <- full endpoint tables, loaded on demand
recipes.md <- worked multi-session orchestration examples
.claude/skills/codeman -> ../../skills/codeman (symlink, dogfooding in this repo)
```
Adding a `skills/` directory to the repo root costs one entry in the GitHub listing. CLAUDE.md
keeps the root short on purpose, so this needs a conscious sign-off; the alternative is
`docs/skills/codeman/` with a `--skill` path argument, which breaks the one-liner install.
**Recommendation**: accept `skills/` at the root, because the install one-liner is the whole
point of shipping a skill.
Install paths, in order of how a user gets it:
1. `npx skills add Ark0N/Codeman --skill codeman -g` (global, any agent, matches the herdr flow).
2. `codeman skill install [--global | --case <name>]`, a new CLI subcommand writing the same
file. This is the path for users who installed via npm and never cloned the repo.
3. **Automatic per-case injection**, modeled exactly on `applyStatusLineConfig(casePath, enabled)`
in `hooks-config.ts`: write `<case>/.claude/skills/codeman/SKILL.md` at case creation,
gated on a new setting. Codeman already writes `<case>/.claude/settings.local.json` hooks
through `writeHooksConfig()`, so this is the same mechanism with the same lifecycle.
Setting name: `agentSkillEnabled`. Synced (not per-device), since it changes on-disk case
content rather than display. Default: **ON after the dogfooding phase, OFF in the first
release**. Rationale for starting OFF: Claude Code loads every skill's name and description
into context on every turn, so an always-on skill has a small permanent token cost, and we
should measure that we are buying something with it first.
### 2.3 SKILL.md content
Frontmatter, per the skills convention (`name` + `description` required):
```yaml
---
name: codeman
description: >-
Control Codeman, the session manager this agent is running inside: list sessions,
start worker sessions, send prompts, read terminal output, and wait for other agents
to finish. Only usable when CODEMAN_MUX=1.
---
```
Body sections, in order:
**1. Guard (first thing, non-negotiable).**
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "not inside a Codeman session"; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set, refusing to guess}"
SELF="${CODEMAN_SESSION_ID:-}"
```
If `CODEMAN_MUX` is not `1`, the agent must stop and say it is not running inside a
Codeman-managed session. Same shape as herdr's `HERDR_ENV` guard, and the variables are
already exported by `tmux-manager.buildEnvExports()`. No fallback URL when
`CODEMAN_API_URL` is unset: any guess is the wrong scheme on an HTTPS install (prod is
HTTPS with a self-signed cert, hence `curl -sk` throughout), and a server the agent
cannot identify is not one it should be driving.
**2. Rules of the road.** Lifted and tightened from README lines 666 to 745:
- Single-line input only. Multi-line breaks the agent TUI (Ink).
- Always send `clientId` + a monotonic `seq` on `POST .../input` so a retry cannot double-deliver.
- Envelope is `{success, data}`; a few legacy GETs are bare, so read `body.data ?? body`.
- Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set. Prod is HTTPS, so `curl -sk`.
- Prefer `/api/v1/*`, the stable alias.
**3. Safety rules (the section that does not exist anywhere today).**
- Never act on `$CODEMAN_SESSION_ID`. That is you.
- Only `DELETE` sessions **you created in this conversation**, by exact id. Keep the list.
- Never bulk-delete, never loop a `DELETE` over `/api/sessions`. There is no undo.
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. Use the API.
- Creating a session consumes a slot against the 50-session cap. Clean up what you start.
**4. Recipes**, each one a single copy-pasteable curl:
| Task | Call |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| list sessions | `GET /api/v1/sessions` |
| find yourself | match ids by PREFIX of `$CODEMAN_SESSION_ID` (Docker cases truncate it to 8 chars, so an equality check never fires there) |
| start a worker | `POST /api/v1/quick-start {caseName, mode, effort}` |
| send a prompt | `POST /api/v1/sessions/:id/input {input:"…\r", useMux:true, clientId, seq}` (the trailing `\r` is what sends Enter; without it the text sits on the prompt unsubmitted) |
| send prompt and wait | `POST /api/v1/sessions/:id/input {input:"…\r", wait:"stop", waitTimeout:600000}` (Part 2) |
| wait for a worker | `GET /api/v1/sessions/:id/wait?until=stop,blocked&timeout=300000` (Part 2) |
| wait for a marker | `GET /api/v1/sessions/:id/wait-output?match=DONE_<random>&timeout=120000` (Part 2; unique per call, per §3.3's repaint rule) |
| read output | `GET /api/v1/sessions/:id/output` |
| read full scrollback | `GET /api/v1/sessions/:id/terminal?full=1` |
| watch sub-agents | `GET /api/v1/subagents` |
| schedule work | `POST /api/v1/cron/jobs` |
| clean up | `DELETE /api/v1/sessions/:id` |
**5. Pointer to `reference/endpoints.md`** for anything not in the table, so the always-loaded
part of the skill stays small.
### 2.4 An ergonomics guard worth adding server-side
The skill will tell the agent not to act on itself, but a confused agent can still try. Propose:
the skill sends `X-Codeman-Caller-Session: $CODEMAN_SESSION_ID` on every request, and the server
refuses destructive operations (`DELETE /api/sessions/:id`, kill, respawn stop) when that header
equals the target id, with a clear error.
This is a **footgun guard, not a security control**: any caller can omit the header. Document it
as such so nobody mistakes it for a boundary. It costs about 10 lines in `route-helpers.ts`.
### 2.5 Verification
Per the always-end-to-end-test rule, "the skill exists" is not done. Done is:
1. Symlink it into `.claude/skills/`, start a real throwaway Codeman session, and ask that agent
to "start a worker session that runs the test suite and tell me when it finishes".
2. Confirm from the outside that exactly one new session appeared, got the prompt, and that the
lead agent waited rather than polling in a busy loop.
3. Confirm the guard: run the same prompt in a shell with `CODEMAN_MUX` unset and confirm refusal.
4. Confirm cleanup: the worker session is deleted by exact id and no other session was touched.
Never run this against `w1`/`w2`/`w3`.
### 2.6 Files touched
- `skills/codeman/SKILL.md` (new), `skills/codeman/reference/*.md` (new)
- `.claude/skills/codeman` symlink (new)
- `src/cli.ts` (new `skill install` subcommand)
- `src/hooks-config.ts` (new `applyAgentSkill(casePath, enabled)`, mirroring `applyStatusLineConfig`)
- `src/web/schemas.ts` (`agentSkillEnabled` in `SettingsUpdateSchema`, which is `.strict()`)
- `src/web/routes/system-routes.ts` (settings PUT must resolve the flag from `merged`, never
from the raw body, per the partial-PUT invariant)
- `src/web/public/settings-ui.js` + `index.html` (checkbox)
- `package.json` `files` array, so `skills/` ships to npm
- README pointer, `docs/extending-codeman.md` seam 3 pointer
---
## 3. Part 2: wait primitives
### 3.1 Goal
Make Codeman orchestratable from a shell tool. Today the only "tell me when" channel is SSE,
which a curl-driven agent cannot practically consume: it would have to hold a streaming
connection and parse events inline. herdr solves this with blocking CLI calls. Codeman should
solve it with bounded long-poll endpoints.
All three additions are **additive**, so the versioning policy stays intact (new endpoints and
new optional fields are non-breaking).
### 3.2 The signal model
A waiter resolves on the first of a set of signals. Sources that already exist:
| Signal | Source today |
| --------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| `idle` | `Session` emits `idle` (session.ts ~1775 for Claude, ~2101 for shell), wired at `session-listener-wiring.ts:402` |
| `working` | `Session` emits `working` (session.ts ~1788), wired at `session-listener-wiring.ts:401` |
| `stop` | `POST /api/hook-event` with `event: 'stop'`, the definitive "Claude finished responding" signal already used by `controller.signalStopHook()` |
| `blocked` | `POST /api/hook-event` with `permission_prompt` or `elicitation_dialog` |
| `exit` | `Session` emits `exit` |
`stop` is the highest-quality signal for "the turn is over" and should be the documented default
for orchestration. `idle` is heuristic: output stabilization plus prompt detection, and it can
flap mid-turn when a spinner pauses. External CLI modes (`isExternalCliMode()`) have no stop
hook at all, so for opencode/codex/gemini/antigravity only `idle`, `working` and `exit` are
available. **The skill and the docs must say which signals exist per mode**, otherwise an agent
waits forever on `stop` in a codex session.
### 3.3 Endpoint specs
#### A. `GET /api/sessions/:id/wait`
| Param | Type | Default | Notes |
| --------- | ---------------------------------------------- | ---------------- | ------------------------------------------------------------ |
| `until` | comma list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on first match |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` (600000) |
| `fresh` | `0`/`1` | `0` | `1` requires a _transition_, ignoring the state at call time |
Response (always 200 unless the session is missing or a cap is hit):
```json
{
"success": true,
"data": {
"signal": "stop",
"timedOut": false,
"immediate": false,
"ended": false,
"waitedMs": 8421,
"status": "idle",
"sessionId": "...",
"until": ["stop", "idle", "exit"],
"limitPaused": false
}
}
```
`until` is echoed back because the server may narrow it: `stop`/`blocked` are dropped
from the DEFAULT set for external CLI modes (asking for them EXPLICITLY is a 400
instead, since omitting `until` must never 400). `limitPaused` tells a caller that a
timeout was expected rather than a stall worth retrying hard.
**A timeout is not an error.** `{"timedOut": true, "signal": null}` with HTTP 200, so a caller
can loop without treating every poll boundary as a failure. Errors are reserved for
`NOT_FOUND` (unknown or not-owned session) and `SESSION_BUSY` (waiter cap exceeded).
`immediate: true` means the session was already in the requested state and `fresh` was not set.
#### B. `GET /api/sessions/:id/wait-output`
| Param | Type | Default | Notes |
| --------- | ------------------------------ | -------- | --------------------------------------------------------- |
| `match` | literal string, 1 to 200 chars | required | substring match against ANSI-stripped output |
| `nocase` | `0`/`1` | `0` | case-insensitive compare |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the existing text buffer first, then waits |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` |
Response: `{ matched: true, timedOut: false, snippet: "...", waitedMs }`.
**No regex in v1, deliberately.** `search-service.ts` already avoids regex specifically so there
is no ReDoS surface, and this endpoint would be even more exposed since the pattern is attacker
supplied and the input is a live stream. herdr can offer `--regex` because Rust's regex crate is
linear-time with no backtracking; JS `RegExp` is not. If regex is wanted later, the honest
options are a length-capped subset compiled once with a match budget, or `re2`. Note it and move on.
Implementation detail that will bite if missed: a match can straddle two PTY chunks. Keep a
carry buffer of `match.length - 1` bytes from the previous chunk and test `carry + chunk`.
⚠️ **`from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, resize, or any TUI redraw, and a repaint arrives as ordinary `terminal`
data. Observed live: a marker echoed a minute earlier matched instantly on a fresh
`from=now` wait. This is inherent to a terminal multiplexer, not fixable in the registry,
so the contract is: **use a marker unique per call** (`echo DONE_$RANDOM`), never a
generic one like `BUILD OK`. The skill's recipes must show that.
The returned snippet is whitespace-collapsed (blank runs to a single newline) for
readability only; matching runs on the raw stripped text. Without it, a real pane's
`\r\n` padding between the prompt and the match fills the whole context window with
nothing, which was the first thing the live test showed.
#### C. `wait` on the existing input endpoint
`POST /api/sessions/:id/input` gains two optional fields:
```json
{ "input": "run the tests\r", "useMux": true, "clientId": "agent-1", "seq": 7, "wait": "stop", "waitTimeout": 600000 }
```
(The trailing `\r` is required on every input body: `sendInput` sends Enter only
when the input contains a carriage return.)
Response gains `"wait": { "signal": "stop", "timedOut": false, "waitedMs": 41230 }`.
This is the important one, because it closes a race the standalone `GET .../wait` cannot: between
"input delivered" and "session flips to working" there is a window where a naive
send-then-wait sees the _pre-existing_ idle state and returns instantly. The combined endpoint
**registers the waiter before writing**, so that window does not exist. This is exactly why herdr
ships `agent prompt --wait` as its own thing.
`wait` accepts `true` (the default signal set) or the same comma grammar as `until`.
Both new fields are `.nullish()`, not `.optional()`: a third-party caller building the
body with `JSON.stringify` keeps an explicit `null` on the wire, and `.optional()`
rejects that with `INVALID_INPUT`. That gotcha has shipped as a real bug twice.
Two behaviors to preserve carefully:
- **`useMux` is fire-and-forget today.** The handler responds without awaiting `writeViaMux`, on
purpose (a tmux child process must not block the HTTP response). With `wait` present the
handler already has to stay open, so it can await delivery, and a `writeViaMux` failure becomes
observable for the first time. The non-wait path must keep its current fire-and-forget shape
byte for byte.
- **Duplicate suppression.** A tagged redelivery (`clientId`+`seq` already applied) returns 200
without writing. With `wait` set it still waits, since the caller's intent is "tell me when
this settles". But it waits with `requireTransition: false`, unlike a fresh delivery: the
original turn may be long over, and requiring a new transition would block a redelivery until
timeout for no reason. Fresh delivery requires a transition, a duplicate answers from the
current state.
- **Capacity rollback.** `shouldApplyInput()` MUTATES (it records the seq), and it runs before
the waiter is registered. If registration then fails on a full pool, the handler must call
`forgetInputSeq` before returning `SESSION_BUSY`, or the caller's retry is rejected as a
duplicate and the input is lost by the very mechanism reliable delivery exists for.
### 3.4 Module design
New file `src/web/session-wait-registry.ts`, with the IO-free core unit-testable in isolation
(same split as `self-update.ts`):
```ts
type WaitSignal = 'idle' | 'working' | 'stop' | 'blocked' | 'exit';
waitForSignal(sessionId, { until: Set<WaitSignal>, timeoutMs, requireTransition }): Promise<WaitResult>
notifySignal(sessionId, signal: WaitSignal): void
waitForOutput(sessionId, { match, nocase, timeoutMs }): Promise<OutputWaitResult>
notifyOutput(sessionId, chunk: string): void
cancelAll(sessionId, reason): void
```
Wiring points, all existing:
- `src/web/session-listener-wiring.ts` around lines 190 and 200 already handles `working` and
`idle` and broadcasts them. Add a `notifySignal()` call next to each broadcast, plus `exit`.
- `src/web/routes/hook-event-routes.ts` already switches on `event` for the respawn controller.
Add `notifySignal(sessionId, 'stop' | 'blocked')` in the same switch.
- Output: `notifyOutput()` rides the ALREADY-attached `terminal` listener in
session-listener-wiring.ts. An earlier draft had the registry hand out attach/detach
callbacks so a listener could be added lazily; that was deleted once it was clear no
second listener is needed at all. The cost is one Map lookup per PTY chunk, which is why
the no-waiter check comes before the ANSI strip.
- Session deletion calls `notifySignal('exit')` then `cancelAll()`, so no promise is left
hanging. Both are required: `_doCleanupSession` detaches the session's listeners BEFORE
`session.stop()`, so on a delete the PTY exit event never reaches the registry, and an
`until=exit` caller would otherwise get a bare `ended` instead of its signal. Found by
live-testing the delete path, not by the unit tests.
Memory-leak discipline, per the 24-hour-session rules: every waiter owns a timer that is cleared
on resolve, the per-session waiter set is deleted when it empties, and the output listener is
removed with it. `test/memory-leak-prevention.test.ts` should grow a case for this.
Caps in a new `src/config/agent-wait.ts` (limits live in `src/config/`, env-overridable):
| Constant | Default | Why |
| ------------------------- | ------- | --------------------------------------- |
| `MAX_WAIT_MS` | 600000 | an unbounded long-poll is a socket leak |
| `DEFAULT_WAIT_MS` | 60000 | short enough to survive most proxies |
| `MAX_WAITERS_PER_SESSION` | 16 | |
| `MAX_WAITERS_TOTAL` | 128 | same reasoning as `MAX_SSE_CLIENTS` |
Exceeding a cap returns `SESSION_BUSY`, not a silent queue.
### 3.5 Transport concerns
Fastify is constructed with defaults in `server.ts:329-331`. `requestTimeout` defaults to 0
(disabled) and `keepAliveTimeout` (72s) applies between requests, not to an in-flight one, so a
10-minute in-process hold is fine. **Verify this on the real instance before relying on it.**
Intermediaries are the actual risk. Prod is reached through `tailscale serve`, and users also run
cloudflared tunnels; both can cut an idle connection. That is why `DEFAULT_WAIT_MS` is 60s and
why the documented pattern is a client-side loop over short waits rather than one 10-minute call.
The skill's recipes must show the loop.
### 3.6 Edge cases to get right
| Case | Behavior |
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Session already idle, `fresh=0` | return immediately, `immediate: true` |
| Session already idle, `fresh=1` | wait for the next transition into a requested state |
| Session dies mid-wait | resolve with `signal: "exit"` if `exit` was requested, otherwise resolve `timedOut:false, signal:null, ended:true`. Never hang |
| Session deleted mid-wait | same, resolve, do not throw. Verified live: `until=exit` gets `signal:"exit"`, a concurrent `until=blocked` gets `ended:true`, both in ~0ms |
| Shutdown with a wait pending | `cancelEverything()` in `stop()`. Verified live: SIGTERM with a 300s wait in flight exits in 1s |
| External CLI mode | `stop` and `blocked` never fire. Reject `until=stop` for those modes with a clear `INVALID_INPUT` rather than hanging until timeout |
| Multi-user | goes through `findSessionOrFail(ctx, id, req)`, which already enforces ownership |
| Remote / Docker cases | signals originate from the same `Session` object, so no special casing. Docker hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for `stop`/`blocked` to arrive at all; without it, only `idle` works. Document it |
| Respawn `/clear` mid-wait | a respawn cycle emits `idle`. Callers waiting on `stop` are unaffected; callers on `idle` may resolve early. Documented, not fixed |
| Limit pause | if the session is paused on a usage limit, nothing will fire until the reset. The wait times out honestly. Consider surfacing `limitPaused: true` in the response so the caller can back off |
### 3.7 Tests
- `test/session-wait-registry.test.ts` (pure): immediate resolve, transition-required, multi-signal
first-wins, timeout, cap exceeded, cancel on session end, no listener leak after resolve,
chunk-straddling output match, case-insensitive match.
- `test/routes/session-wait-routes.test.ts` (`app.inject()`, no port): all three endpoints against
a `MockSession`, including the 200-with-`timedOut` contract and the ownership 404.
- `test/routes/session-input-wait.test.ts`: the send-and-wait race, plus proof that the non-wait
path is unchanged (still returns before `writeViaMux` settles).
- Live verification on a throwaway session before COM, per the always-end-to-end-test rule.
### 3.8 Files touched
- `src/config/agent-wait.ts` (new)
- `src/web/session-wait-registry.ts` (new)
- `src/web/session-listener-wiring.ts` (notify on idle/working/exit)
- `src/web/routes/hook-event-routes.ts` (notify on stop/blocked)
- `src/web/routes/session-routes.ts` (two new routes, `wait` fields on input)
- `src/web/schemas.ts` (`SessionWaitQuerySchema`, `SessionWaitOutputQuerySchema`, extend
`SessionInputWithLimitSchema`. Note: `.optional()` rejects `null`, so the frontend and any
generated client must send `undefined`, never `null`)
- `docs/api-reference.md`, `docs/extending-codeman.md`, README API table
- `skills/codeman/SKILL.md` recipes (Part 1 depends on this)
---
## 4. Deferred: parts 3 to 5
Not in scope now, kept here so they are not lost.
### Part 3: promote `blocked` to a first-class state
`SessionStatus` is `'idle' | 'busy' | 'stopped' | 'error'`. "Needs you" exists three times over:
hook events, the `tab-alert-action` CSS class, and the phone overview NEEDS YOU section, each
re-deriving it. herdr makes `blocked` a real state that rolls up.
Add `blocked` (and possibly `done`) to `SessionStatus`, set it from the same hook events that
Part 2 uses as wait signals, and clear it on the next `working`/`stop`. Then the tab strip, the
mobile overview, the wait endpoints, and any external agent read one field.
Cost: `SessionStatus` is a widely-consumed union, so every exhaustive `switch` (the codebase has
`assertNever` and `noFallthroughCasesInSwitch`) will need a branch. That is a feature, it makes
the compiler find every site. This is a **minor** bump, not a patch: it widens a public type in
the HTTP contract.
### Part 4: `GET /api/schema`
herdr ships `herdr api schema`. Every Codeman route is already Zod-validated, so
`zod-to-json-schema` over `schemas.ts` gives a self-describing API almost free. Value: third-party
tools and the skill stop drifting from hand-written docs. Open question: whether to emit full
OpenAPI (`@fastify/swagger` would need per-route schema registration, which is a much larger
change) or just dump the Zod schemas keyed by name (cheap, 80% of the value).
### Part 5: detection manifests instead of hardcoded patterns
CLI-specific readiness, blocked and usage-limit patterns live in code across
`usage-limit-patterns.ts`, the respawn pattern helpers and `regex-patterns.ts`. Externalizing the
per-CLI ones into data files would make adding a sixth CLI a data change instead of a code change.
**Do not copy the remote-update part.** herdr auto-fetches manifest updates from herdr.dev.
Codeman auto-pulling behavioral rules from a vendor server contradicts its security posture.
Bundled manifests plus local override only, no network.
### Explicit non-goals
- **Plugin runtime and marketplace.** `docs/extending-codeman.md` already argues this: a plugin
runtime means third-party code inside a process that spawns agents with your credentials, on a
server people expose over a tunnel. The reasoning still holds. If the marketplace _pattern_ is
wanted, apply it to data (web tabs, case templates, cron recipes), never to executable code.
- **Live PTY handoff on restart.** herdr needs it because it owns the terminals. Codeman
delegates to tmux, so PTYs already survive a self-update restart.
- **Socket API.** HTTP plus SSE is the existing, documented, stable contract. A second transport
would double the surface for no capability gain.
---
## 5. Sequencing
| Step | Work | Gate |
| ---- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 ✅ | `src/config/agent-wait.ts` + `session-wait-registry.ts` + unit tests | 48 tests green |
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
| 5 ✅ | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
| 6 ✅ | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | 10 unit tests (`test/agent-skill.test.ts`) + real-server case-creation tests (`test/quick-start.test.ts`, incl. the settings PUT accepting the key) green; CLI verified live (install/uninstall, global + `--case`, foreign/symlink refusals) |
| 7 | Docs: api-reference, extending-codeman, README | |
| 8 | COM (minor bump: new endpoints, new setting, new optional fields) | both CI and Release workflows green |
Parts 1 and 2 are independent enough to land separately, but the skill is much less useful
without the wait endpoints, so the wait work goes first.
## 6. Open questions for the owner
1. ✅ `skills/` at the repo root: accepted (built that way; the install one-liner depends on it).
2. ✅ `agentSkillEnabled` default: **OFF** for the first release, per §2.2's rationale (skills
cost context on every turn; measure before defaulting on). Flip later if dogfooding earns it.
3. ✅ Both: global install via `npx skills add` / `codeman skill install`, AND per-case
auto-injection behind the (default-off) setting. Injection is add-only at session create and
marker-guarded, so a user-authored copy is never touched.
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
and not a security boundary? (Still open, not built with step 6.)
5. Regex support in `wait-output`: confirm literal-only for v1.
---
## 7. Build log: what actually happened
Written at the end of the build so the next person inherits the reasoning, not just the
diff. Process artifacts (per-agent briefs, findings, reports) live in the gitignored
`tmp/agent-wait-review/`; this section is the part worth keeping.
### What shipped
| Piece | Files |
| ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Bounds + clamping | `src/config/agent-wait.ts` (new) |
| Blocking-wait registry | `src/web/session-wait-registry.ts` (new, IO-free, unit-tested) |
| `GET .../wait`, `GET .../wait-output`, `wait`/`waitTimeout` on `POST .../input` | `src/web/routes/session-routes.ts` |
| Signal wiring | `session-listener-wiring.ts` (idle/working/exit + output), `hook-event-routes.ts` (stop/blocked), `server.ts` (teardown, shutdown) |
| Agent skill | `skills/codeman/SKILL.md` + `reference/`, `.claude/skills/codeman` symlink, `package.json` `files` |
| Docs | `api-reference.md`, `extending-codeman.md`, `architecture-invariants.md`, `README.md`, `CLAUDE.md` |
| Tests | `test/session-wait-registry.test.ts`, three `test/routes/session-*wait*.test.ts`, `http-contract.test.ts`, `mock-session.ts` |
### Bugs found in ADJACENT code, not in the new feature
These are the highest-value output of the exercise and none were on the plan:
1. **Every Codeman hook was dead on HTTPS installs.** `hooks-config.ts` built the hook
curl as `curl -s` with no `-k` while the statusline exporter 300 lines below used
`curl -sk` and documented why. Proven with the real hook command: `curl exit=60`
without the flag, success with it, and the failure swallowed by the hook's own
`2>/dev/null || true`. This silently killed `stop`, `permission_prompt`,
`elicitation_dialog`, `idle_prompt`, `teammate_idle` and `task_completed`, taking
respawn's definitive idle signals with them. Fixed, **plus** a staleness detector in
`refreshStaleCodemanHooks` that regenerates the on-disk config of already-created
cases (23 of 26 local cases carried the broken form; fixing the generator alone would
have left every one of them broken).
2. **`buildEnvExports()` exported a wrong-scheme `CODEMAN_API_URL`** (`http://` fallback
on an HTTPS install). Now omitted rather than guessed, so in-session guards fail closed.
3. **Programmatic input is only submitted when it contains `\r`.** `sendInput` sends Enter
only if the payload has a carriage return; without it the text sits in the composer
forever. Bit this build repeatedly before it was diagnosed, and had leaked into the
docs' own examples.
### Design decisions worth not re-litigating
- **A timeout is HTTP 200** with `wait.timedOut`, never a 4xx: callers loop over short
waits because tunnels cut idle connections, and every poll boundary would otherwise be
indistinguishable from failure.
- **Send-and-wait must be one endpoint.** A separate POST-then-wait races: between the
write and the flip to `working`, a wait sees the stale `idle` and reports the PREVIOUS
turn as this one. The waiter is registered before the write.
- **`stop`/`blocked` exist for `claude` mode only.** They come from Claude Code hooks;
`shell` installs none either, so keying off `isExternalCliMode()` was wrong.
- **Literal matching only, never regex.** JS `RegExp` backtracks; herdr can offer
`--regex` because Rust's regex crate is linear-time.
- **Client-hangup abort listens on `reply.raw` guarded by `writableFinished`.** On
`req.raw`, `close` fires when the request BODY ends, which on a POST killed every
send-and-wait instantly, and no `app.inject()` test can see it (inject never emits
`close`).
- **Liveness cannot come from `session.pid`.** For a tmux session that is the local
`tmux attach` client, not the worker: a worker exiting inside its pane leaves
`pane_dead=1` with the client alive, so `pid` never goes null. Liveness is probed at
the mux layer, cached (~750 ms) and only on blocking waits, never on the input hot path.
### Verification rounds
Six agents across three rounds, each verifying the previous round's work rather than its
own. Findings that mattered, in order of severity, were: the dead-pane liveness gap; the
`reply.raw` abort regression; abandoned long-polls leaking waiter slots; a crashed session
reporting `idle`; `shell` accepting `until=stop`; and a documented recipe that reported
success without running its task. Two traps recurred often enough to name:
- **Vacuous passes.** `app.inject()` never emits `close`; a latched `cancelEverything()`
in `afterEach` silently killed the registry for every later test in a file; three test
files sharing one session id against the process-wide registry let one file's leftover
waiter fail another's assertion. Any new wait test needs care on all three.
- **HTTP-only test instances.** Every isolated instance used during the build was plain
HTTP, which is exactly why the HTTPS hook bug survived so long. Test the transport the
user actually runs.
### Resolved at wrap-up (2026-08-08, conclusion pass)
- **R2-A**: the fire-and-forget-then-gather-sequentially pattern was **removed from
the skill** rather than patched. Signals are edge-triggered with no history, so a
`stop` that fires before its waiter registers is unobservable afterwards; a
`fresh=0` gather was rejected because the only `until` set that current state can
satisfy answers `idle` for a prompt that never submitted, resurrecting the exact
false-success failure R2-B had just closed. Flow 3b's pattern B now gathers on
latched `wait-output` markers (`from=buffer`), the same mechanism that makes the
shell flows reliable; the limitation is recorded in
`architecture-invariants#agent-wait-primitives` and `endpoints.md`. The durable
fix, a latched last-signal-per-turn on the server, stays with deferred Part 3.
- Docs F7/F8, F4 and the false-`idle` attribution: `api-reference.md`,
`extending-codeman.md` and `architecture-invariants.md` rewritten to the post-fix
matcher (one normalized stream, chunk-straddling found, snippet as a rendering of
the matched window), the real no-PTY answer (`ended:true`, `aborted:false`,
`delivered:false`), and the startup-idle mechanism (a session parked on the trust
dialog emits no further `idle`; the false success is the startup transition).
- Orchestrate #12, #5/R2-B, #6, and R2-C..R2-E: fire-and-forget's empty `data`
documented; every send-and-wait retry loop now treats `duplicate:true` +
`immediate:true` as "no new turn ran" and reads the terminal before believing it;
claude fan-out is pattern A (backgrounded send-and-waits) or the marker gather;
readiness budgets rebalanced (5 s stage 1, 45 s stage 3) with the virgin-case
floor named; the auth fallback now also reads the supervisor definition
(`codeman-web.service` / launchd plist) and accepts `export`-prefixed `.env`
lines; `pid != null` is documented as startup-only, never liveness.
- Both public readiness recipes (extending-codeman.md, README) are bypass-first with
the trust probe as the bounded fallback; the worked recipe carries `-k` and fails
loudly on an empty SID; the hook `-k`/self-heal fix appears in every
"hooks go missing" list; the multi-word-TUI claim is "unreliable", not "never".
### Still open
- **Release checklist**: `package.json` `files` includes `skills`, which is still
untracked. `git add skills/` must be part of the release commit, or npm publishes
a tarball without the skill (a `files` entry that does not exist is silently
ignored, so nothing fails). `test/agent-skill.test.ts` reads the packaged source,
so CI at least fails loudly if the directory goes missing from a checkout.
- The 1.13.0 changeset is written under `.changeset/`; consuming it (COM flow),
the release commit, and the deploy remain.
- Deferred with Part 3: the latched last-signal-per-turn. Nice-to-haves from the
reviews: N2 (create the death-watcher inside its `try`) and converting
timeout-shaped test detections into fast assertions.
- §2.4's `X-Codeman-Caller-Session` footgun guard: still not built (open question 4).
### Step 6 (2026-08-09): install command, per-case injection, the setting
Built to the §2.6 file list, mirroring the statusLine mechanism throughout:
| Piece | Where |
| ----- | ----- |
| `applyAgentSkill(casePath, enabled)` + `installAgentSkillInto` / `removeAgentSkillFrom` | `src/hooks-config.ts` |
| `codeman skill install` / `skill uninstall` (`--global` default, `--case <name>`) | `src/cli.ts` |
| `agentSkillEnabled` (SYNCED, default OFF) | `schemas.ts` (`SettingsUpdateSchema`), `getAgentSkillEnabled()` on `ConfigPort`/`server.ts`, checkbox in `index.html` + `settings-ui.js` |
| Injection call sites (Claude mode only) | `POST /api/sessions` next to `refreshStaleCodemanHooks`; `POST /api/quick-start` after the case-create/self-heal blocks (local + docker cases; remote skipped, its path lives on another host) |
| Tests | `test/agent-skill.test.ts` (10 unit), `test/quick-start.test.ts` (real server: default-off, PUT accepts key, injection on create, shell-mode skipped) |
Decisions worth keeping:
- **Ownership marker, prefix-matched.** The injected SKILL.md ends with
`<!-- codeman-managed-agent-skill: … -->`; install/refresh/remove all refuse a copy
without the marker (a user's own skill) and match on the PREFIX so a wording change
cannot disown older injected copies (the `BACKGROUND_WAKE_MARKER_PREFIX` pattern).
- **Symlink refusal.** This repo's own dogfooding layout
(`.claude/skills/codeman -> ../../skills/codeman`) means the injector must `lstat`
the skill dir AND its `skills/` parent and bail on a symlink, or enabling the
setting in the Codeman repo itself would overwrite the skill source through the link.
- **ADD-ONLY at session create**, same shared-`.claude` rationale as the statusLine:
a create while the setting is off must not yank the skill out from under other live
sessions in the repo. The remove path exists (CLI `skill uninstall`, tests); no
automatic sweep removes on toggle-off.
- **Removal is manifest-based, never `rm -rf`**: only files the packaged source would
have written are deleted, directories are pruned bottom-up only if they emptied, so
a user's extra notes in `reference/` survive an uninstall.
- **Source resolution**: `join(moduleDir, '..', 'skills', 'codeman')` works from
`src/` (tsx), `dist/` (tsc build), and the npm tarball alike, because all three sit
one level below the package root and `files` ships `skills/`.
- **Nothing acts on the setting at PUT time**: injection reads the merged persisted
settings at session create (`readSettings`, ~2s cache), so the partial-PUT invariant
(`toggleService` reading `merged`) is untouched by construction.
+341
View File
@@ -46,6 +46,20 @@ payload return `{ "success": true, "data": {} }`.
> `GET /api/screenshots/:name`, `GET /q/:code` (QR redirect), and the
> `GET /ws/sessions/:id/terminal` WebSocket upgrade.
> The [agent wait endpoints](#long-polling-agent-wait) use the normal envelope but
> are the only JSON endpoints that deliberately **hold the connection open**, for up
> to 600 s. Proxy operators and HTTP clients with a global read timeout need to know
> that before pointing them at Codeman.
⚠️ **A `401` is the one status that is not an envelope.** Authentication is rejected
in a request hook, before any handler runs, and it replies with the bare string
`Unauthorized` (`Unauthorized: hook secret required` on the hook path) plus
`WWW-Authenticate: Basic realm="Codeman"`. There is no `success`, no `error`, and no
`errorCode`, because the wrapping hook only wraps object payloads. So a client that
pipes every response straight into a JSON parser dies with a parse error rather than
reporting an auth failure, which is a confusing way to discover that a password is
set. Branch on the HTTP status **before** parsing.
## Error codes → HTTP status
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
@@ -66,6 +80,333 @@ the HTTP status.
Adding a new error code is non-breaking; removing or renaming one is a major change.
## Long-polling (agent wait)
Three calls block until something happens instead of answering immediately. They
exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse
events inline.
| Call | Blocks until |
|------|--------------|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
`GET .../wait`. It registers the waiter **before** writing, which closes the window
in which a separate wait sees the session still idle from the previous turn and
answers instantly with the wrong turn's result. Use it whenever you send a prompt
and want to know when that prompt is done.
### Three semantics that break callers who assume otherwise
**1. A timeout is HTTP `200`, not an error.** A wait that ends without its signal
returns `{"success":true, ...,"wait":{"timedOut":true,"signal":null}}`. The
intended pattern is a client-side loop over short waits, because `tailscale serve`
and cloudflared can both cut an idle connection, and turning every poll boundary
into a `4xx` would make that loop indistinguishable from a real failure. `408` is
auto-retried by several clients (silently doubling the polling load), `504` is what
a genuine tunnel failure looks like, and `204` cannot carry `waitedMs` / `status` /
`limitPaused`. Reserve error handling for the four codes in the table below.
**2. `stop` and `blocked` fire only for `claude` sessions.** Both come from Claude
Code hooks, and no other mode installs them: `shell` runs no agent, and the external
CLIs (`opencode`, `codex`, `gemini`, `antigravity`) render their own TUIs and post
no hooks. For every non-`claude` mode only `idle`, `working` and `exit` are
accepted, and of those only `exit` is dependable: see the caveats under
[Signals](#signals) before building on `idle`. Requesting `stop` or `blocked`
**explicitly** on such a session is a
`400`; omitting `until` never fails, the server just drops them from the default set
and echoes the narrowed set back as `wait.until`. Three more places hooks can go
missing even in `claude` mode: a **Docker case** needs
`CODEMAN_DOCKER_BRIDGE_HOOKS=1`, since a container cannot reach a loopback-bound
Codeman (without it, only `idle` / `working` / `exit` work); a **remote-SSH
case** runs the agent on another host, whose hooks may never reach this server at
all; and a case whose hook config was written by **Codeman < 1.13.0 against an
`--https` install** carries hook curls without `-k`, which TLS-fail silently (the
hook line ends in `|| true`). Codeman now writes `curl -sk` and repairs a stale
case config the next time a session starts in that case. When in doubt, ask for
`stop,idle,exit` so a session without hooks still resolves on the heuristic
signal.
**3. `from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, on resize, and on any TUI redraw, and a repaint arrives as
ordinary output, so text that was already on screen can satisfy a fresh wait. This
was observed live: a marker echoed a minute earlier matched instantly on a new
`from=now` wait. It is inherent to running the agent under a multiplexer, so the
contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
`echo $MARK`, then wait on `$MARK`), never a generic string like `BUILD OK`.
### Signals
| Signal | Source | Actually fires for |
|--------|--------|--------------------|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
is omitted is `stop,idle,exit` (`exit` is in there so a worker that crashes resolves
the wait promptly instead of burning the caller's whole timeout on something that
can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit`
once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no *further* `idle`, so it is the
startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died *inside* its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
milliseconds, and reading that as "the worker died" is wrong: it means start it, or
wait for it to come up. `status` is carried alongside so nothing is hidden. The
alternative (trusting `status`) is worse, because a dead PTY parks the session at
`status: "idle"`, which would answer the default wait with `immediate: true` for a
worker that has crashed. A worker dying while a wait is parked resolves it within
a few seconds (a background death-watcher), not at the timeout.
⚠️ **`blocked` is reachable less often than the table suggests.** It fires on two
hooks, and the default configuration suppresses one of them: Codeman spawns claude
with `--dangerously-skip-permissions`, so permission prompts do not happen unless the
instance is switched to the `auto` Claude mode (App Settings), or the caller is a
multi-user account without the bypass grant, which is forced to `--permission-mode
auto`. What does still fire under the default is `elicitation_dialog`, the agent
asking the user a question. So `until=stop,blocked,exit` is a reasonable belt on a
long turn, but a worker that never comes back is far more likely to be working than
blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
instead. The same caution applies to the external CLIs.
### Readiness is not a signal
Nothing here reports "the agent is ready for a prompt", and no combination of
`until`/`fresh` synthesizes one. A freshly created session reads as `exit` (above),
and a `claude` worker in a brand-new case comes up on the CLI's **trust dialog**,
which contains a ❯ prompt of its own. Send-and-wait posted at that moment types the
prompt into the dialog, where the `\r` never gets past it, while the session's
startup `idle` lands inside the wait window: the wait resolves on `idle` in a
couple of seconds with `timedOut: false`, which looks exactly like a completed
turn.
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
dialog only as the bounded fallback (`trust` matched → send `\r` → wait for
`bypass` again). Do not probe `trust` first and Enter blindly: the dialog text
stays in the terminal buffer for the life of the session, so a `trust` probe with
`from=buffer` keeps matching on every later run and the Enter lands in a ready
composer. A worked version is in
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
```bash
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
```
Both GET wait routes answer with `Cache-Control: no-store`, because the documented
pattern polls one identical URL in a loop and a cached `{"timedOut":true}` would
turn that loop into a busy spin. `POST .../input` sends no cache header (it is a
POST, which is not heuristically cacheable).
⚠️ **Unknown query parameters are ignored, not rejected**, with one exception
(`regex`, below). In particular `match=` on `/wait` is silently dropped and you get
a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
instead of waiting on the wrong thing. The reasoning is in
[`architecture-invariants.md`](architecture-invariants.md#agent-wait-primitives).
#### What the matcher actually sees
The matcher scans the raw PTY stream, **normalized**: ANSI escape sequences are
stripped — CSI, OSC, and the charset-designation escapes a stock bash prompt emits
on every line (`ESC ( B`), so `match=tnode:` matches a prompt that renders
`…@tnode:` — a partial escape arriving at a chunk boundary is held back until its
tail arrives, and a match may straddle PTY chunks: `printf STRAD; sleep 1; printf
DLEQQ` is matchable as `STRADDLEQQ` (all measured live). Three caveats remain:
⚠️ **It is still the byte stream, not the rendered pane.** `GET .../terminal`
answers from a tmux screen capture (`data.source: "mux-visible"`), the finished
picture; the matcher sees the stream that painted it. For linear output the two
agree once escapes are stripped, but a full-screen TUI composes its picture with
cursor positioning, so what the pane shows and what the stream carries can differ.
Seeing your string in `terminal?tail=` makes a match likely, not guaranteed.
⚠️ **A TUI's text can arrive without its spaces.** Claude Code positions words
with cursor moves rather than printing spaces, so screen text can reach the
matcher as `Quicksafetycheck:Isthisaprojectyoucreated...`. Whether a given phrase
keeps its spaces depends on how the TUI happened to draw it (measured: `I trust
this folder` matched, `Quick safety check` did not), so a multi-word `match`
against a TUI pane is unreliable rather than impossible. Match a **single
space-free token**, ideally one you printed yourself. Plain command output (a
shell worker, an `echo`) keeps its spaces.
⚠️ **The returned `snippet` is a rendering of the matched text, not a quotation of
it.** It is cut from the same normalized stream the match ran against, then
cleaned for display: remaining raw control bytes are removed (an agent pipes the
snippet into its own terminal, so a worker's bytes must not be able to reset that
display) and blank runs are collapsed. A printable needle that matched will appear
in it; a needle containing control bytes or a blank run may not survive verbatim.
```bash
MARK="DONE_$RANDOM"
curl -sG "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'timeout=120000'
```
Build the query with `-G --data-urlencode` rather than by hand: a `+` in a
hand-written query string decodes to a space.
### `POST /api/v1/sessions/:id/input` with `wait`
Two optional fields on the existing endpoint:
| Field | Type | Notes |
|-------|------|-------|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would
reject it, which has shipped as a real bug twice.
The input must end with `\r` (a real carriage return in the JSON string): Enter is
sent only when the input contains one, so text without it is typed onto the
worker's prompt but never submitted, and the wait then runs its full timeout on a
turn that never started. Verified live; this is the most common silent failure on
this endpoint.
```bash
curl -s -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true,"clientId":"agent-1","seq":1,
"wait":"stop","waitTimeout":600000}'
```
A **tagged duplicate** (a `clientId` + `seq` pair the server has already applied)
still honors `wait`, because the caller's question is unanswered, but it answers
from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
any of them:
```json
{ "success": true, "data": {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"status": "idle",
"limitPaused": false,
"wait": {
"signal": "stop", "until": ["stop", "idle", "exit"],
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
"waitedMs": 8421, "timeoutMs": 60000
}
}}
```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
`status` and `limitPaused`. `POST .../input` **without** `wait` is unchanged and
still returns `{"success": true, "data": {}}`.
⚠️ `delivered: false` has **two** meanings, and they must be told apart by
`duplicate`: with `duplicate: true` the input was suppressed as an already-applied
redelivery (harmless, the turn it refers to may be long over), while with
`duplicate: false` the **write failed** (typically no PTY behind the session). A
client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success.
| Field | Type | Meaning |
|-------|------|---------|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order:
1. `wait.signal !== null` (or `wait.matched === true`): the thing happened.
2. `wait.timedOut`: a poll boundary. Loop again.
3. `wait.ended` or `wait.aborted`: the wait answered nothing, because the session is
gone or was never running. Re-check the session instead of looping.
`wait.immediate` is not a fourth outcome: it rides along with the first one and
means the condition already held at call time, so nothing was actually waited for.
If that is not what you meant, you wanted `fresh=1` or the send-and-wait form. Note
that `{"signal":"exit","immediate":true}` on a session you just created is the
not-started-yet case, not a crash.
**The timeout is clamped, so read it back.** A request for 1800000 ms is silently
reduced to the server's ceiling (600000 ms by default, operator-tunable), and a
request for 1 ms is raised to 1000 ms. `wait.timeoutMs` is the value that was
applied. Without checking it, a caller that asked for 30 minutes and got 10 will
read the timeout as "the worker is wedged" and kill a session that was working fine.
### Errors
| `errorCode` | HTTP | When |
|-------------|------|------|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
error message names the cap that was hit.
⚠️ A `401` is **not** in this table and is not an envelope at all (see
[Response envelope](#response-envelope)). It matters most here: a polling loop that
pipes each wait straight into `jq` fails with a parse error on every iteration
against a password-protected server, which reads as "the wait endpoints are broken".
Check the status first.
The per-session cap is a **combined** budget: signal waiters and output waiters
count against the same 16, not 16 of each. An abandoned request no longer holds its
slot, because the routes release the waiter when the client disconnects, but a
client that opens many concurrent waits against one session will still hit the cap.
## Authentication
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
File diff suppressed because one or more lines are too long
+21 -3
View File
@@ -149,9 +149,17 @@ to Claude as a system reminder. This implies `"async": true`; ordinary async
hooks do not wake an idle turn, and their output waits for the next interaction.
Codeman uses this on `PostToolUse(Bash)`: a self-contained Node helper extracts
the background task ID from the Bash result, watches the session transcript for
the matching completion notification, and exits 2. It does not send terminal
input, so it cannot submit a user's partially written prompt.
the background task ID from the Bash result, watches the originating transcript
and, for subagents, the top-level parent transcript for the matching completion
notification, and exits 2. Claude records a subagent's Bash result in its
`subagents/agent-*.jsonl` file but queues completion in the lead session JSONL.
The task ID keeps each wake targeted. The helper does not send terminal input,
so it cannot submit a user's partially written prompt.
For script-dispatched Codex work, `codex-run.sh` writes the final response
between `CODEMAN_RESULT_BEGIN/END` markers in the background task output. The
rewake helper includes a maximum of 64 KiB of that report in its feedback. UI
subagent discovery and dispatcher result delivery are separate contracts.
### Notification
@@ -219,6 +227,16 @@ Or to allow exit:
**Use Cases**: Control nested loops, verify subagent output.
The hook input includes `agent_id`, `agent_transcript_path`, and
`last_assistant_message`. Like `Stop`, a command hook can return
`{"decision":"block","reason":"..."}` to keep the subagent running and feed
the reason back to it.
Codeman uses this to prevent premature reports from workers that still own live
Monitor or background-Bash processes. It derives candidate task IDs from the
subagent transcript, but requires a matching live Linux process descriptor for
`tasks/<id>.output`; historical task text by itself is not treated as active.
### TeammateIdle
**When**: When an agent-team teammate is about to go idle.
+169 -7
View File
@@ -44,6 +44,10 @@ is in [`api-reference.md`](api-reference.md).
the payload at the top level rather than under `data`. Read defensively with
`body.data ?? body`.
⚠️ A `401` is not an envelope at all: auth is rejected in a request hook that
replies with the bare string `Unauthorized`, so parsing it as JSON throws. Branch on
the status code before you parse, or a missing password looks like a broken endpoint.
**Already driving Codeman from an agent?** The README's
[Programmatic Guide](../README.md#driving-codeman-from-an-agent--programmatic-guide)
covers the in-session case: the `CODEMAN_MUX`, `CODEMAN_API_URL`,
@@ -152,10 +156,17 @@ for (;;) {
## Seam 3: HTTP API and CLI
Around 199 handlers across 21 route files cover sessions, cases, files, cron,
Around 200 handlers across 21 route files cover sessions, cases, files, cron,
respawn, Ralph, the orchestrator, search, and admin. Each route module carries an
`@fileoverview` describing its endpoints.
If the caller is an agent running _inside_ a Codeman session, install the packaged
agent skill instead of teaching it these calls by hand: `skills/codeman` in the repo
(`npx skills add Ark0N/Codeman --skill codeman -g`, or `codeman skill install
[--case <name>]`, or the synced `agentSkillEnabled` App Setting for automatic
per-case injection on Claude session create). The skill carries the guard, the
safety rules, and verified wait/orchestration recipes.
The common ones:
```bash
@@ -167,10 +178,12 @@ curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions \
-H 'Content-Type: application/json' \
-d '{"workingDir":"/home/me/project","mode":"claude"}'
# Send a prompt (single-line only)
# Send a prompt (single-line only, and it must end with \r: Enter is sent only
# when the input contains a carriage return; without it the text sits on the
# session's prompt unsubmitted)
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions/$ID/input \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests","useMux":true}'
-d '{"input":"run the tests\r","useMux":true}'
```
`POST .../input` also accepts `clientId` (stable per client, max 128 chars) and
@@ -178,6 +191,124 @@ curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions/$ID/input \
at-most-once, so retrying after a dropped connection cannot type the prompt
twice. Omit them entirely rather than sending `null`.
It also accepts `wait` and `waitTimeout`, which hold the response open until the
session finishes the turn you just started. `wait` is `true` (the default signal
set) or a comma list of `idle,working,stop,blocked,exit`; the result comes back
under `data.wait`. Sending them changes nothing for callers that do not: without
`wait` the response is still `{"success": true, "data": {}}` and the write is still
fire-and-forget. The two interact with `clientId` / `seq` in one way worth knowing:
a **tagged duplicate** (a pair the server already applied) skips the write but still
waits, answering from the session's current state rather than blocking for a
transition that already happened. It reports `"delivered": false, "duplicate": true`.
### Waiting instead of polling
Three calls block until something happens: `GET /api/v1/sessions/:id/wait` (a
lifecycle signal), `GET /api/v1/sessions/:id/wait-output` (a literal string in the
output), and the `wait` field above. Full parameter and response tables are in
[`api-reference.md`](api-reference.md#long-polling-agent-wait). Four things decide
whether your integration works, and the last one is what actually bites:
- **A timeout is a `200` with `wait.timedOut: true`**, not an error. Loop over short
waits rather than issuing one long one, because `tailscale serve` and cloudflared
both cut idle connections and a single 10-minute call is the pattern most likely
to die in the field.
- **`wait.timeoutMs`** is the timeout after server-side clamping (600 s ceiling by
default). Read it rather than assuming you got what you asked for.
- **`stop` and `blocked` only exist for `claude` sessions**, and on a `shell` session
even `idle` fires only once at startup, so send-and-wait there can only time out.
See the Gotchas below.
⚠️ **There is no readiness signal, and skipping readiness is the failure that looks
like success.** A session reports `idle` before its CLI has spawned, and a `claude`
worker in a brand-new case comes up on the CLI's **trust dialog**, which has a ❯
prompt of its own. Prompt it at that moment and the text lands in the dialog, the
`\r` does not get past it, and the session's startup `idle` lands inside the wait
window: the wait resolves on `idle` in a couple of seconds with `timedOut: false`,
indistinguishable from a finished turn. Wait for the pid, then wait for the
composer, answering the dialog only as the bounded fallback.
A worked orchestration: start a worker, get it ready, prompt it, wait, clean up.
```bash
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}" # auto-set in-session, correct scheme included
AUTH=(-u "admin:$CODEMAN_PASSWORD") # omit entirely if no password is set
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on --https installs (self-signed cert)
# 1. Start a worker session (creates the case if it does not exist yet).
# The guard matters: a TLS or auth failure otherwise leaves SID empty and every
# later step "succeeds" against nothing.
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
[ -n "$SID" ] && [ "$SID" != null ] || { echo "quick-start failed"; exit 1; }
# 2. READINESS: composer marker first, trust dialog only as the bounded fallback.
# Skip this and step 3 reports a turn that never ran. Do NOT probe trust first
# and Enter blindly: the dialog text stays in the buffer for the life of the
# session, so on every later run that probe matches stale text and the Enter
# lands in a ready composer. Match single tokens only: TUI text can arrive
# without its spaces. Stage 1 is short on purpose (an already-trusted case
# matches in <1 s; a first-run case can never pass it and pays it in full).
until [ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ]
do sleep 1; done
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=5000') # composer's status bar = ready
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' -d '{"input":"\r","useMux":true}' >/dev/null
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send the prompt AND register the wait in one call, so the answer cannot be
# the previous turn's idle state. Single line only, ending in \r (otherwise
# Enter is never sent and this wait times out on a turn that never started).
W=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize the failures\r","useMux":true,
"clientId":"orchestrator","seq":1,"wait":"stop,exit","waitTimeout":60000}' \
| jq -c '.data.wait')
# 4. That first wait probably timed out (60 s). Keep going in SHORT waits.
for _ in $(seq 1 30); do
[ "$(jq -r '.timedOut' <<<"$W")" = 'true' ] || break # signal fired, or wait ended
W=$("${CURL[@]}" \
"$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq -c '.data.wait')
done
jq -r 'if .ended or .aborted then "worker is not running"
elif .timedOut then "still working after 30 waits"
else "signal: \(.signal)" end' <<<"$W"
# 5. Read what it produced, then delete the session YOU created, by exact id.
# ⚠️ NOT /output: its textOutput is empty for every tmux-backed session.
# `tail` counts BYTES, and the payload is terminal data with ANSI in it.
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
```
Waiting on a marker instead of a signal is the form that works in **every** mode,
and the only one that works on a `shell` session:
```bash
# ⚠️ Split the marker so the typed line never contains it: your own keystrokes echo
# into the output stream, so an unsplit marker matches before the command has run.
# `from=buffer` also catches a marker that printed before the wait registered.
N=$RANDOM
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
```
For shell scripting, the `codeman` CLI is the same surface without the HTTP
plumbing:
@@ -221,10 +352,41 @@ Every one of these has cost somebody real time.
shipped bugs more than once.
- **`text/plain` bodies stay raw.** Auto-parsing them as JSON enabled
simple-request CSRF, so it is deliberate. Send `application/json`.
- **Prompts are single-line.** With `useMux: true` the server delivers your text
and then Enter as two separate writes, so you do not append `\r` yourself. A
multi-line string breaks the agent's Ink-based input handling: send one line,
or split it across calls.
- **Prompts are single-line and must end with `\r`.** The server splits your text
and Enter into two separate tmux writes (Ink needs them apart), but it sends the
Enter **only when the input contains a carriage return**. Without it your text
sits on the prompt unsubmitted, which is the single most common "the wait
endpoints don't work" report: the wait runs its full timeout on a turn that never
started. Newlines inside the string are stripped rather than rejected, so
`"echo A\necho B\r"` runs the single joined command `echo Aecho B`: send one line
per call.
- **`wait-output`'s `from=now` is not "printed after you asked".** tmux repaints
the visible screen on attach, on resize, and on any TUI redraw, and a repaint
arrives as ordinary output, so text already on screen can satisfy a fresh wait.
Observed live: a marker echoed a minute earlier matched instantly. Use a marker
unique to each call, and build it so the typed line never contains it (your own
keystrokes echo into the stream). Matching is a literal substring, so `regex=` is
rejected with a `400` rather than ignored.
- **`wait-output` matches the normalized PTY stream, not the screen.** ANSI escape
sequences are stripped (the `ESC ( B` charset escape a bash prompt emits on every
line included), a partial escape at a chunk boundary is held back until its tail
arrives, and a match may straddle PTY chunks, so text you printed yourself
matches reliably (`printf STRAD; sleep 1; printf DLEQQ` is matchable as
`STRADDLEQQ`). What can still fail is TUI output: a full-screen TUI positions
words with cursor moves, so its text can reach the matcher **without spaces** and
a multi-word match is unreliable there. Match one short space-free token, ideally
one you printed yourself, and keep it out of the typed line (your own keystrokes
echo into the stream).
- **`stop` and `blocked` never fire for `shell`, `opencode`, `codex`, `gemini` or
`antigravity` sessions.** They come from Claude Code hooks, which no other mode
installs, so only `idle`, `working` and `exit` exist there. Asking for them
explicitly is a `400`; omitting `until` is safe, since the server drops them from
the default set and echoes what it actually waited on as `wait.until`. Even in
`claude` mode, a Docker case needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for hooks to
reach the server at all, a remote-SSH case's hooks may never arrive, and a case
written by Codeman < 1.13.0 against an `--https` install carries hook curls
without `-k` that TLS-fail silently — a 1.13.0+ server rewrites them the next
time a session starts in that case.
- **Unwrap the envelope** before reading fields. `data` is not the response body.
## Publishing your integration
+189
View File
@@ -25,6 +25,195 @@ What shipped vs. what this doc proposed:
- **Finding 5 (remote probe)**: implemented (`probeRemoteCliVersion` over ssh, deferred at
session start, same login-shell wrapper as the launch).
## RETEST FAILED (2026-08-07, after v1.12.0 shipped) — analysis round 2
mtiller retested on 1.12.0 and reports it is NOT fixed (issue #205 comment, 2026-08-07 12:12 UTC;
issue reopened same day with clarifying questions: mouse vs trackpad, Shift+scroll behavior,
Claude vs shell session on the phone, and an iOS full-tab-kill to rule out stale JS). Two
failure signatures, now analyzed against the SHIPPED 1.12.0 code (not the pre-fix code):
1. **iPhone Safari (Claude session assumed)**: touch scrollback goes back only a limited
amount and sometimes REPEATS blocks of text; unreliable.
2. **Firefox on macOS (mouse)**: wheel does NOTHING at all, while Fn+Up (= PageUp) pages back
through INTACT text.
### Ruled out by code reading
- deltaMode mishandling: `_wheelScrollLinesFloat` normalizes line/page/pixel units correctly;
a Firefox line-mode notch yields ±3 lines. Not the bug.
- Ephemeral transport: `_sendInputEphemeral` (app.js) has a POST fallback when WS is down.
- Service worker: sw.js is network-first with cache fallback; it serves stale JS only when the
fetch FAILS (flaky mobile connection can do this — relevant to "unreliable" on the phone,
and the fixed `CACHE_NAME = 'codeman-v1'` never invalidates that offline copy).
### The load-bearing observation: PageUp works, the wheel does not
Fn+Up is a KEYBOARD event: xterm encodes PageUp and Claude pages its own transcript (intact
text proves Claude-side history is fine and the PTY input path is fine). The wheel path is the
capture-phase handler, and for a Claude session it has exactly two branches:
- **Forwarding branch** (`_shouldForwardWheelToApp` true): snap-to-bottom + SGR reports. If
this branch ran, the user would see the same paging motion Fn+Up produces. They see nothing.
- **Local branch** (gate false): `_smoothScrollBy` over xterm's local buffer. For a Claude
pane in repaint mode, tmux keeps `history_size≈0`, so `?full=1` returns roughly one frame:
the local buffer is structurally HOLLOW, the top-of-buffer re-pull recovers nothing, and the
wheel looks completely dead. **This matches every observed detail on Firefox.**
So the working hypothesis is that mtiller's sessions evaluate the gate FALSE. The gate
(`_shouldForwardWheelToApp`) has exactly four false-paths worth checking, in likelihood order:
1. **`terminalWheelLocalScrollback` opt-out is ON.** Plausible: a user whose scrolling was
broken on 1.11.x may well have toggled "Wheel scrolls local history" while trying to fix
it. On 1.12.0 that setting now routes the wheel to a hollow local buffer = dead wheel on
desktop AND the stale-repaint-frames experience on the phone (see below). Ask, or check
what the setting does on their export.
2. **`cliVersion` missing — CONFIRMED BUG, independent of whether it is mtiller's**:
`getClaudeCliVersion()` (utils/claude-cli-resolver.ts:124-148) caches its result
process-wide including FAILURE: on any exception it sets `_claudeVersion = null`, and the
guard is `!== undefined`, so a single failed/timed-out probe (5s `EXEC_TIMEOUT_MS`; PATH
under systemd/launchd; transient fs hiccup) at the FIRST Claude session start disables
wheel forwarding for every Claude session until the server restarts. Fix: cache success
permanently, but let failure retry (retry on next call, or a short negative-cache TTL).
Note that mtiller sees identical breakage on phone + iPad + laptop, which points at a
SERVER-side/session-side cause exactly like this (cliVersion is shared by all devices)
rather than anything browser-specific.
3. **Claude Code genuinely < 2.1.187** on their machine: gate false BY DESIGN, but the
resulting UX is a dead-end (no local history to fall back on).
4. mouseTrackingMode non-none (a DECSET leaked past the strip, e.g. emitted before attach or
split across chunks in a way the carry missed): would also kill the container handler via
the early return. Least likely, checkable via `terminal.modes.mouseTrackingMode` in console.
### The iPhone symptoms fit the same gate-false story
Touch with gate false = local `scrollLines()` over whatever repaint frames accumulated:
"repeats blocks of text" is literally what a buffer of successive overlapping repaint frames
looks like; "limited amount" is its thinness; "unreliable" is burst-dependence (finding 2)
PLUS the new re-pull being actively DESTRUCTIVE for repaint panes: `_maybeRefetchFullHistory`
does `_resetTerminalForReplay()` then writes the fetched capture, and when that capture is
one frame (Claude pane, `history_size≈0`) it REPLACES a multi-frame buffer with less than the
user had, mid-scroll. Stale pre-1.12 JS on the phone (suspended Safari tab) remains possible
until they confirm the tab kill.
### Fix directions, ranked
1. **Make the re-pull refuse downgrades** (`_maybeRefetchFullHistory`, app.js): if the fetched
capture would yield FEWER buffer rows than currently present, skip the reset+rewrite and
keep the richer buffer (optionally cache-mark the session "re-pull useless"). Small, safe,
kills the "got worse after scrolling to top" class. Consider skipping the re-pull entirely
for forwarding-capable modes where tmux keeps no history.
2. **Rescue the gate-false Claude dead-end with PageUp forwarding**: when mode is `claude`,
the gate is false, AND the local buffer has no scrollback (`baseY === 0`), translate wheel
lines into coalesced PageUp/PageDown key sends (mtiller just proved Claude pages correctly
on PageUp even on their version). Zero regression risk under that triple guard: sessions
with real local history keep local scrolling; only the currently-dead path changes.
Caveat: older Claude menus may react to PageUp; acceptable against "completely dead".
3. **Audit `getClaudeCliVersion()` failure caching** (utils/claude-cli-resolver.ts): a cached
empty probe must retry (with backoff), not poison the process.
4. **Guard the opt-out setting's footgun**: if `terminalWheelLocalScrollback` is ON for a
repaint-mode CLI session, local history is hollow; either scope the setting's effect to
modes with real local scrollback, or pair it with fix 2's PageUp fallback so it still
scrolls SOMETHING.
5. **Add a one-line gate diagnostic**: log (once per session, console) WHY the wheel chose
local vs forward: `{mode, cliVersion, optOut, trackingMode}`. The #205 thread is now two
rounds deep on guesswork a single console line would have answered.
### What shipped for round 2 (branch `fix/scrollback-205-round2`)
All five directions above, implemented as ranked:
1. **Downgrade guard** — `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the rows a
capture will occupy (ANSI stripped, `capture-pane -J` re-wrapping accounted for) and
`_maybeRefetchFullHistory` (app.js) skips the reset+rewrite when that is more than one
screen short of what xterm already holds. A refused session goes on
`_fullHistoryRepullUseless`, which raises its re-pull cooldown from 4s to 60s so a hollow
pane stops re-fetching. Measured A/B on a live Claude pane, same gesture, same buffer:
guard off → 341 rows collapse to 42 and every seeded row is gone; guard on → 341 rows
preserved. The tab-switch recovery it must not break still runs (shell buffer 401 → 44 on
a tab switch → 401 again after scrolling to the top).
2. **PageUp/PageDown fallback** — `_maybePageCliTranscript()` translates wheel/touch travel
into coalesced `\x1b[5~` / `\x1b[6~` under the triple guard (claude mode, forwarding gate
false, `baseY === 0`), through the same 40ms queue as the SGR reports. Half a screen of
travel per page: the page key always jumps a whole screen, and a 1:1 mapping was
unusably slow with a discrete wheel. Shift is excluded — it keeps meaning "local
scrollback". Verified live: opt-out ON on a Claude session sends real PageUp/PageDown to
the PTY where the wheel previously did nothing.
3. **Probe caching** — `getClaudeCliVersion()` no longer caches failure. Success is kept for
the process lifetime; a failed probe retries with a 1/2/4…15min backoff. The cache policy
is a pure function (`resolveClaudeCliVersion`) so the retry semantics are unit-testable
without spawning `claude`. The VITEST short-circuit now records nothing, where before it
wrote a permanent null.
4. **Opt-out footgun** — handled by pairing rather than by scoping: the setting keeps meaning
exactly what it says (the wheel goes local), and fix 2 catches the case where "local" is
empty. Scoping the setting away from repaint-mode CLIs would have silently overridden an
explicit user choice. The App Settings tooltip now says to leave it off for Claude/Codex.
5. **Diagnostic** — `_logScrollRouting()` prints one line per session per distinct decision:
`[scroll] <id> → forward-sgr|page-keys|local-scrollback|repull-refused-downgrade (mode=…,
cliVersion=…, localScrollbackOptOut=…, mouseTracking=…, localScrollbackRows=…)`. That
single line answers every open question in the list below.
Still unanswered by code alone: whether mtiller's Claude Code is genuinely older than
2.1.187 (false-path 3), and whether the iPhone was running stale JS. The diagnostic makes
both self-reporting, so the retest ask is now "open the console and paste the `[scroll]` line".
### What to get from mtiller (some already asked)
- Shift+scroll behavior on Firefox (distinguishes hollow-local from handler-not-firing).
- `claude --version` on the Mac (decides false-paths 2 vs 3).
- App Settings → Input → "Wheel scrolls local history" state (false-path 1).
- iPhone: Claude or shell session, and whether a full tab kill changes anything.
- Browser console: `app.terminalUi?.terminal?.modes?.mouseTrackingMode` (false-path 4).
## ROUND 3 (2026-08-09): Codex wheel dead — CONFIRMED AND FIXED
DodgyBadger (Codex latest, Chrome, Windows 11): mouse wheel does nothing in a CODEX session
while working fine in shell and web tabs; DRAGGING THE SCROLLBAR WORKS, so xterm's local
buffer demonstrably has content for their codex pane. Analysis against the shipped code:
- `_shouldForwardWheelToApp` returns true UNCONDITIONALLY for `codex` (no version gate, unlike
claude's `>= 2.1.187`), so every plain wheel tick is sent as SGR reports to Codex.
- The "verified to scroll its transcript on SGR wheel reports" claim for codex predates
current Codex builds; if Codex latest ignores SGR wheel, forwarding eats the gesture while
the healthy local scrollback (proven by the working scrollbar) sits unused.
- The #227 PageUp fallback cannot rescue this: it is gated to `claude` mode AND `baseY === 0`,
and codex here has real local scrollback. The `[scroll]` diagnostic will still say
`forward-sgr (mode=codex, ...)`, confirming the branch, worth asking the reporter to paste.
**CONFIRMED by the reporter's `[scroll]` line (2026-08-09, PR #227 comment)**:
`forward-sgr (mode=codex, cliVersion=unknown, localScrollbackOptOut=false, mouseTracking=none,
localScrollbackRows=967)`. Forwarding branch active, 967 rows of healthy local scrollback
unused, Codex ignoring the SGR reports. Environment: Codex latest, Chrome, Windows 11.
**Measured against codex-cli 0.147.0** (isolated `tmux -L codexwheel`, fake `CODEX_HOME/auth.json`,
history built with 401ing prompts), which settles it without needing a version gate at all:
| Probe | Result |
| ---------------------------------------------- | ----------------------------------------------- |
| `#{mouse_any_flag}` once the TUI is up | `0`: codex never enables mouse tracking |
| `#{alternate_on}` | `0`: inline viewport, not an alt-screen pager |
| `#{history_size}` while prompting | grows 3 → 32: the transcript goes to scrollback |
| 6 × `\x1b[<64;10;10M` written to the pane | pane capture byte-identical, nothing happens |
| control: literal `zz` | pane changes, so the probe can see changes |
| `\x1b[<0;12;5M` + release (the click-tap path) | no change either: taps are no-ops, not garbage |
Codex has no in-app pager to drive: its history lives in the terminal's own scrollback, which is
exactly what forwarding was stealing the gesture from. A version gate would be the wrong fix (and
`cliVersion=unknown` means there is no codex probe to gate on anyway).
**Fix (shipped):** `_shouldForwardWheelToApp` now returns true for `claude >= 2.1.187` and nothing
else. Codex falls to the normal local-scrollback path like shell/gemini/opencode, so wheel and touch
scroll the same history the scrollbar drag was already scrolling. The claude-only PageUp fallback is
untouched: codex never needs it, its local buffer is real. Taps stay hand-encoded for codex
(`_sessionUsesServerMouseStrip`), measured harmless, so click-to-position is merely unavailable
there rather than damaging. Lesson for the next mode added to the forward list: "it is a strip mode"
proves nothing, write a real SGR report into a live pane and diff the capture first.
Verified end-to-end in Chromium against a live codex session on an isolated instance
(`CODEMAN_INSTANCE=codexwheel`, port 5055, `envOverrides.CODEX_HOME` pointing at the fake auth
dir): trusted `page.mouse.wheel` up now logs
`[scroll] … → local-scrollback (mode=codex, …, localScrollbackRows=43)`, moves the viewport
39 → 4 (back to the Codex banner), and sends ZERO bytes to the PTY. Unit coverage:
`test/terminal-touch-tap.test.ts` ("only claude forwards — codex and gemini keep the local wheel").
Original plan follows.
## Reports
+10 -28
View File
@@ -43,38 +43,22 @@ const syncData = DEC_SYNC_START + data + DEC_SYNC_END;
this.broadcast('session:terminal', { id: sessionId, data: syncData });
```
## Client-Side Implementation (`app.js`)
## Client-Side Implementation (`terminal-ui.js`)
### `batchTerminalWrite(data)`
1. Checks if flicker filter is enabled (optional, per-session)
2. If flicker filter active: buffers screen-clear patterns (`ESC[2J`, `ESC[H ESC[J`, `ESC[nA`)
3. Accumulates data in `pendingWrites`
4. Schedules `requestAnimationFrame` if not already scheduled
5. On rAF callback: checks for incomplete sync blocks (start without end)
6. If incomplete: waits up to 50ms via `syncWaitTimeout`
7. Calls `flushPendingWrites()` when complete
### `extractSyncSegments(data)`
- Parses DEC 2026 markers, returns array of content segments
- Content before sync blocks returned as-is
- Content inside sync blocks returned without markers
- Incomplete blocks (start without end) returned with marker for next chunk
4. Calls `_scheduleTerminalWriteFlush()` if no flush is pending
5. The yielded callback clears its scheduled flag before calling `flushPendingWrites()`
6. Large batches schedule their own next chunk until the queue is empty
### `flushPendingWrites()`
```javascript
const segments = extractSyncSegments(this.pendingWrites);
this.pendingWrites = ''; // Clear before writing
for (const segment of segments) {
if (segment && !segment.startsWith(DEC_SYNC_START)) {
terminal.write(segment); // Skip incomplete blocks (start with marker)
}
}
```
Note: Segments starting with `DEC_SYNC_START` are incomplete blocks awaiting more data. These are skipped (discarded if timeout forces flush).
- Joins the queued terminal data and passes DEC 2026 markers through to xterm.js 6, which handles synchronized output natively.
- Writes at most 32KB per yield for Codex and 64KB for other modes.
- Requeues the remainder and immediately schedules another safe yield. A final large response therefore drains without waiting for another SSE event.
### `chunkedTerminalWrite(buffer, chunkSize=128KB)`
@@ -116,17 +100,15 @@ When detected, buffers 50ms of subsequent output before flushing atomically.
## Edge Cases
- **Incomplete sync blocks**: 50ms timeout forces flush (content discarded to prevent freeze)
- **Incomplete sync blocks**: xterm.js retains synchronized output until its closing marker
- **Large buffers**: Chunked writing prevents UI freeze
- **Server shutdown**: Skips batching via `_isStopping` flag
- **Session switch**: Clears flicker filter state, pending writes, and sync timeout (prevents cross-session data bleed)
- **SSE reconnect**: `handleInit()` clears all pending write state
**Trade-off:** If a sync block is split across SSE packets and the end marker doesn't arrive within 50ms, the incomplete content is discarded. This prioritizes responsiveness over completeness. In practice this is rare since the server always sends complete `SYNC_START...SYNC_END` pairs and SSE typically delivers them atomically.
## DEC Mode 2026 Compatibility
Terminals that natively support DEC 2026 will buffer and render atomically. Terminals that don't support it ignore the escape sequences harmlessly. xterm.js doesn't support DEC 2026 natively, so the client implements its own buffering by parsing the markers.
Terminals that natively support DEC 2026 buffer and render atomically. Codeman uses xterm.js 6, so the client passes the markers through instead of parsing or discarding partial blocks.
**Supporting terminals:** WezTerm, Kitty, Ghostty, iTerm2 3.5+, Windows Terminal, VSCode terminal
@@ -135,4 +117,4 @@ Terminals that natively support DEC 2026 will buffer and render atomically. Term
| File | Key Functions |
|------|---------------|
| `src/web/server.ts` | `batchTerminalData()`, `flushTerminalBatches()`, `broadcast()` |
| `src/web/public/app.js` | `batchTerminalWrite()`, `extractSyncSegments()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
| `src/web/public/terminal-ui.js` | `batchTerminalWrite()`, `_scheduleTerminalWriteFlush()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.12.0",
"version": "1.14.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.12.0",
"version": "1.14.1",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+2 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.12.0",
"version": "1.14.1",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -159,6 +159,7 @@
"dist",
"scripts/postinstall.js",
"scripts/fix-node-pty.mjs",
"skills",
"LICENSE",
"README.md"
]
+344
View File
@@ -0,0 +1,344 @@
---
name: codeman
description: >-
Drive Codeman, the session manager this agent is running inside, over its HTTP API:
list sessions, start worker sessions, send them prompts, block until they finish
(wait / wait-output / send-and-wait), read their output, and clean up. Use when asked
to orchestrate or parallelize work across Codeman sessions, watch another session, or
start and manage workers. Only usable inside a Codeman-managed session
(CODEMAN_MUX=1); refuse to act otherwise.
---
# Driving Codeman from inside a session
You are an agent running inside a Codeman-managed terminal session. Codeman is the
server that spawned you; its HTTP API can start, prompt, watch, and delete other
sessions. Every recipe below was verified live. Full endpoint tables and
troubleshooting: [reference/endpoints.md](reference/endpoints.md). Worked multi-worker
flows: [reference/recipes.md](reference/recipes.md).
## 0. Guard, and the one thing that breaks every recipe below
⚠️ **Your shell state does not survive between tool calls.** Each Bash call starts a
fresh shell, so `$API`, `$SELF`, the `CURL` array and `delete_session` are all gone by
the next call, and `$$` is a different pid. Three consequences, all of which have
teeth:
- **Re-run this entire preamble at the top of every Bash call that touches the API.**
Running it once and assuming it stuck is the single most likely way to break a run.
- **Never re-paste only half of it.** The delete guard below is written so that a
missing definition deletes nothing, but that only holds if you never hand-roll a
`DELETE` of your own.
- **Never put `$$` in a `clientId`.** It changes per call, so the "resend the identical
request" loop in §3 would stop being a duplicate and would **retype the prompt**,
submitting the turn twice. Use a fixed literal (`codeman-agent-1` below).
Only real environment variables (`CODEMAN_*`) survive, which is why this preamble
rebuilds everything else from them.
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "Not inside a Codeman-managed session; refusing to act."; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Codeman does NOT hand a session the server password. If one is set, the two
# in-reach copies are the data dir's .env (the same fallback `codeman attach`
# uses — hand-authored; nothing ever writes it) and the supervisor definition
# that install.sh wrote the password into, which is where a stock
# password-protected install actually keeps it. The data dir is wherever the
# hook-secret file lives. Values may be quoted or `export`-prefixed.
ENV_FILE="${CODEMAN_HOOK_SECRET_FILE:+${CODEMAN_HOOK_SECRET_FILE%hook-secret}.env}"
envval() { sed -n "s/^\(export \)\{0,1\}$1=//p" "$ENV_FILE" | tail -1 | sed 's/^"\(.*\)"$/\1/; s/^'\''\(.*\)'\''$/\1/'; }
if [ -z "${CODEMAN_PASSWORD:-}" ] && [ -n "$ENV_FILE" ] && [ -f "$ENV_FILE" ]; then
CODEMAN_USERNAME=$(envval CODEMAN_USERNAME)
CODEMAN_PASSWORD=$(envval CODEMAN_PASSWORD)
fi
if [ -z "${CODEMAN_PASSWORD:-}" ]; then # stock installs: install.sh puts it in the service definition
UNIT="$HOME/.config/systemd/user/codeman-web.service"
PLIST="$HOME/Library/LaunchAgents/com.codeman.web.plist"
if [ -f "$UNIT" ]; then
# install.sh backslash-escapes " and \ in the unit value; undo it or a password
# containing either recovers wrong and auth fails.
CODEMAN_PASSWORD=$(sed -n 's/^Environment="CODEMAN_PASSWORD=\(.*\)"$/\1/p' "$UNIT" | head -1 | sed 's/\\\(["\\]\)/\1/g')
elif [ -f "$PLIST" ]; then
# install.sh XML-escapes the plist value; undo it (&amp; LAST, mirroring escape order).
CODEMAN_PASSWORD=$(awk '/<key>CODEMAN_PASSWORD<\/key>/{getline; print}' "$PLIST" | sed -n 's/.*<string>\(.*\)<\/string>.*/\1/p' \
| sed -e 's/&lt;/</g' -e 's/&gt;/>/g' -e 's/&amp;/\&/g')
fi
fi
AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:$CODEMAN_PASSWORD")
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on https (self-signed cert)
# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
# `is_self "$SID" || curl -X DELETE ...` shape failed OPEN, because an undefined
# is_self exits 127 and the `||` branch then ran the delete completely unguarded.
# Undefined delete_session is "command not found", which deletes nothing.
delete_session() {
local id="${1:-}"
[ -n "$id" ] || { echo "refusing: empty session id"; return 1; }
[ "${#SELF}" -ge 8 ] || { echo "refusing: \$SELF unset or too short to prove this is not me"; return 1; }
# ids appear in full AND 8-char form (Docker exports a truncated $SELF; mux names and
# UI surfaces carry 8-char ids), so compare by prefix in BOTH directions. Equality or
# a one-directional check each miss a real combination, and the miss deletes you.
case "$id" in "$SELF"*) echo "refusing: $id is me"; return 1 ;; esac
case "$SELF" in "$id"*) echo "refusing: $id is me"; return 1 ;; esac
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$id"
}
CID=codeman-agent-1 # FIXED literal, never "agent-$$" (see §0)
```
- If `CODEMAN_MUX` is not `1`, **stop and say so**. Do not guess an API URL; a server
you are not part of is not yours to drive.
- **A 401 is plain text, not the JSON envelope**, so on a password-protected server
every `jq` in these recipes dies with `jq: parse error` instead of showing
`UNAUTHORIZED`. If that happens, check the status with `-w '%{http_code}'`; if it
is 401 and neither fallback above found a credential, **stop and tell the user
you need credentials**. The hook-secret bypass covers only `/api/hook-event` and
`/api/status-telemetry`, never session control.
- These endpoints first ship in Codeman **1.13.0**, but do not gate on the version
number: a dev build can serve them while reporting an older version. Probe
instead: `GET .../wait` on a real session id answering 404 with an `.error`
starting `Route ` means the server predates the wait endpoints (fall back to
polling `GET .../terminal?tail=` and say so); `Session ... not found` means your
session id is wrong, not the server.
## 1. Safety rules — read before any mutating call
You are yourself a session on this server, and the API has **no undo**.
- **Never act on your own session, and know that `delete_session` is the ONLY guard.**
The server has no self-protection: a session that DELETEs its own id succeeds and
dies silently (verified live). **Always delete through `delete_session "$SID"` from
§0; never write a bare `curl -X DELETE` and never reintroduce the
`is_self … || curl -X DELETE …` shape.** That older form failed open: with the
function undefined (a half-re-pasted preamble, see §0) bash returns 127, the `||`
branch fires, and the delete runs with no self-check at all. Wrapping the request
inside the guard is what makes a lost preamble delete nothing instead of deleting
you. Apply the same prefix-both-directions reasoning before any kill, respawn, or
input call you write by hand.
- **Mutating calls you may make unprompted** (this is an allowlist):
`POST /api/v1/quick-start`, `POST /api/v1/sessions/:id/input`, and
`DELETE /api/v1/sessions/:id` **only** for a session you created in this
conversation, by exact id. Keep a list of the ids you create. Everything else
mutating needs the user to have asked for it.
- **Never call these** unless the user explicitly asked, naming the target:
- `DELETE /api/cases/:name` — recursively **deletes a real directory of the user's
code** from disk. One wrong case name destroys work that was never yours.
- `DELETE /api/sessions` (no id) and `DELETE /api/subagents` (no id) — bulk kills.
- respawn / ralph / orchestrator / cron mutations — respawn runs `/clear` (wipes a
conversation), orchestrator state is a single global slot, cron jobs outlive you.
- `PUT /api/settings`, `POST /api/system/update` — global UI settings; server restart.
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. The API is the only interface.
- Sessions count against a 50-session cap and case creation is uncapped: clean up every
session you start, and don't retry `quick-start` in a loop.
## 2. Rules of the road
- **End every input with `\r`** — literally the two characters `\r` inside the JSON
string. Codeman types the text and sends Enter **only when the input contains a
carriage return**; without it your command sits unsubmitted on the worker's prompt
and everything downstream times out. `{"input":"run the tests\r",...}`. No response
field catches this: `delivered:true` means "written to the pane", **not**
"submitted" — a `\r`-less send still reports `delivered:true` and then every wait
times out, which is why the loops below are bounded and check the terminal.
- **Single-line input only.** Newlines are stripped; one line per call.
- **Build request bodies with `jq -n` for any prompt you did not author as a
literal.** The inline `-d '{"input":"'"$P"'\r"}'` pattern breaks on the first
double quote, backslash, or `$` in a real prompt:
```bash
BODY=$(jq -n --arg p "$PROMPT" '{input:($p+"\r"),useMux:true,clientId:"agent-1",seq:1,wait:true,waitTimeout:60000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' --data-binary "$BODY"
```
- **Exactly-once delivery**: always send a stable `clientId` and a monotonic
per-session `seq` on `POST .../input`. A retry after a dropped connection then
cannot double-type the prompt. Increment `seq` for each NEW input; reuse the same
pair only to re-ask about the same delivery.
- **Envelope**: success is `{"success":true,"data":…}`, errors are
`{"success":false,"error","errorCode"}`. Read `.data`. Use `/api/v1/*` paths.
- **A wait timeout is HTTP 200**, `{wait:{timedOut:true,signal:null}}` — not an error.
Loop over short waits (60 s); proxies cut long-idle connections. Timeouts are
**clamped** (ceiling 600 s): read back `wait.timeoutMs` for what was applied.
- **`stop` and `blocked` fire for `claude` sessions only** (Claude Code hooks). On
`shell`/`opencode`/`codex`/`gemini`/`antigravity`, requesting them explicitly is a
400 — and lifecycle transitions there are coarse (a short shell command may emit
**no** `idle` transition at all, verified live), so synchronize those modes with
output markers, not signals.
- **Your typed command echoes into the output stream**, so a marker that appears
verbatim in the input line matches **before the command runs**. Always split the
marker (recipe below), keep it unique per call, and use `from=buffer` so a marker
that printed before your wait landed is still found. Matching is literal — no regex.
- **Match single space-free tokens against TUI output.** A full-screen TUI (claude,
codex, …) positions text with cursor movements, not literal spaces, so the stripped
stream can read `Yes,Itrustthisfolder` and a multi-word match is unreliable there —
whether a phrase keeps its spaces depends on how the TUI happened to draw it
(observed live: some match, some never fire). Plain command output (shell workers,
`echo` lines) keeps real spaces.
## 3. Recipes (each verified live)
**List sessions / find yourself** — metadata only, safe to poll:
```bash
"${CURL[@]}" "$API/api/v1/sessions" | jq '.data[] | {id, name, mode, status}'
"${CURL[@]}" "$API/api/v1/sessions" | jq --arg s "$SELF" '.data[] | select(.id | startswith($s))'
```
**Start a claude worker and wait until it is actually ready.** A new session reports
`idle` before its CLI has spawned, and a brand-new case shows a **trust dialog**
first, so neither "wait for idle" nor "wait for ❯" means ready (the trust dialog
contains `❯` too — observed live). Codeman *can* auto-accept that dialog itself, but
the accept rides a stream match that misses on some runs (both outcomes seen live),
so wait for the composer first and handle the dialog only as the bounded fallback —
never send a blind Enter up front (if auto-accept already fired, it lands in the
composer). Stage 1 is short on purpose: an already-trusted case matches `bypass` in
under a second, while a **virgin case can never pass stage 1** (the dialog is up, so
the composer is not) and always pays it in full before the fallback runs — the long
budget belongs to stage 3, after the dialog is answered:
```bash
# ALWAYS check .success: on failure `.data.sessionId` is null, jq -r prints the string
# "null", and the flow below then burns its full readiness budget against
# /api/v1/sessions/null before reporting jq noise instead of the actual cause.
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
if [ -z "$SID" ]; then
# SESSION_BUSY here is the 50-session cap, not the waiter cap; FORBIDDEN/CONFLICT/
# OPERATION_FAILED/INVALID_INPUT are the others. None are retryable in a loop.
jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed; stopping."
exit 1
fi
for _ in $(seq 1 30); do # bounded: a bad SID would otherwise poll forever
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# ⚠️ pid != null proves STARTUP only, never life: a worker that later dies inside
# its pane keeps status "idle" and a pid (the local tmux attach client, not the
# worker). The death check is wait?until=exit, below.
SEQ=1 # $CID came from the §0 preamble; do NOT rebuild it from $$
# the composer's status bar ("bypass permissions on") is the ready marker — Codeman
# spawns claude in bypass mode. Single-token matches only: TUI text is space-less.
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# composer never appeared → the trust dialog is probably still up; accept it once
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
jq -e '.data.wait.matched' <<<"$R" >/dev/null || \
{ echo "worker $SID never became ready; inspect terminal?tail="; }
fi
```
**Send a prompt and wait for the turn to finish** (claude workers — the call to
prefer). It registers the waiter *before* typing, closing the race where a separate
wait sees the previous turn's idle state. Loop by resending the **identical** request:
the repeat is a tagged duplicate (same `clientId`+`seq`) that does not retype but
answers from the session's current state. Verified: the stop hook resolves this in
seconds; a duplicate resend answers in ~20 ms without retyping.
```bash
for TRY in $(seq 1 10); do # BOUNDED: a \r-less send never produces a signal and resends are no-op duplicates
R=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"run the tests, then summarize in one line\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ',"wait":true,"waitTimeout":60000}')
if jq -e '.data.wait.timedOut' <<<"$R" >/dev/null; then
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # two straight timeouts: prompt sitting unsubmitted?
continue
fi
# Resolved — but a duplicate answering immediately reports the session's CURRENT
# state ("it is idle now"), NOT that a new turn ran. A \r-less send lands exactly
# here on try 2 (verified live), so check the terminal before believing it:
if jq -e '.data.duplicate and .data.wait.immediate' <<<"$R" >/dev/null; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" | jq -r '.data.terminalBuffer' | tail -5
# your prompt still on the ❯ composer line = never submitted (missing \r);
# submit it with {"input":"\r"} (the only recovery), then loop again
fi
break
done
SEQ=$((SEQ+1)); jq '.data.wait.signal, .data.status' <<<"$R"
```
Read the outcome in this order: `wait.signal != null` → done (`stop` is definitive;
`idle` is heuristic) — **unless** it arrived as `duplicate:true` + `immediate:true`,
which only says the session is idle *now* and must be confirmed from the terminal
(above); `wait.timedOut` → loop again (bounded); `wait.ended` → session gone, stop.
If the loop exhausts its cap, do not keep looping: read the terminal, report what
you see, and remember that a still-typed-but-unsubmitted prompt (missing `\r`) can
only be recovered by submitting it with `{"input":"\r"}`.
**Shell worker + completion marker** — the pattern for `shell` mode (no hooks there).
The typed line must not contain the marker verbatim (the input echo would match
instantly — observed live), so build it with a variable the worker's shell expands:
```bash
N="${RANDOM}_$$"; MARK="DONE_$N" # unique per call: tmux repaints replay old text
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}'
SEQ=$((SEQ+1))
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=120000' \
| jq -r '.data.wait | {matched, snippet}'
```
The typed line shows `${M}_…`, the real output shows `DONE_… rc=<exit code>`, and the
snippet carries the exit code back to you.
**Read a worker's answer.** For `claude` and `codex` workers this is the read path:
`last-response` returns the agent's final message as clean text, taken from the
transcript rather than the screen, so it carries none of the TUI's box-drawing or
repaint noise.
```bash
for _ in $(seq 1 10); do # the transcript write LAGS the stop signal
TXT=$("${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
```
`.data` is `{text, timestamp}`. ⚠️ **Poll it, do not read it once.** `text` is written
from the transcript file, which is flushed slightly *after* the `stop` hook fires, so a
single read taken the instant send-and-wait returns comes back `""` even though the
turn finished (verified live: empty on the first call, full text seconds later). `text`
is also `""` before the worker's first completed turn, and always `""` for modes with
no transcript (`shell`, `opencode`, `gemini`, `antigravity`, verified live), which is
why the loop above is bounded rather than open-ended. Fall back to the terminal buffer there, tail in **bytes**
(`textOutput` in `GET .../output` stays empty for interactive sessions; don't use it):
```bash
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=3000" | jq -r '.data.terminalBuffer' \
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g' | grep -v '^[[:space:]]*$' | tail -30
```
⚠️ Do not use that pipeline to read a **claude/codex** answer. A full-screen TUI draws
with cursor moves, so the stripped buffer is largely one long line: `tail -30` has
almost nothing to split on and you get a wall of repaint noise with the answer buried
in it (verified live, side by side with `last-response` returning the exact prose).
The terminal buffer is for *diagnosis* (is my prompt sitting unsubmitted?), not for
reading answers. Avoid `?full=1` (entire tmux scrollback, a context bomb) unless doing
a post-mortem.
**Detect a dead worker cheaply**: `GET .../wait?until=exit&timeout=60000` answers
immediately (`signal:"exit"`, `immediate:true`) if the PTY is gone — including a
worker that exited *inside* its pane, which `GET .../sessions/:id` keeps reporting
as `status:"idle"` with a pid (that pid is the local tmux attach client, not the
worker). The wait routes are the only liveness check; a worker dying while a wait
is parked resolves it within ~3 s. A session deleted mid-wait resolves in ~1 s.
**Clean up** — only ids you created, one at a time, always through the §0 helper:
```bash
delete_session "$SID"
```
Everything else (endpoint tables, per-mode signal table, error codes, capacity
limits, Docker/remote caveats): [reference/endpoints.md](reference/endpoints.md).
Fan-out orchestration and blocked-worker handling:
[reference/recipes.md](reference/recipes.md).
+229
View File
@@ -0,0 +1,229 @@
# Codeman API reference for agents
Loaded on demand from the `codeman` skill. Assumes the guard variables from SKILL.md
(`$API`, `$SELF`, `"${CURL[@]}"`). Canonical contract: `docs/api-reference.md` in the
Codeman repo; this file is the agent-relevant subset, verified live.
## Envelope and errors
Every JSON response: `{"success":true,"data":…}` or
`{"success":false,"error":"…","errorCode":"…"}`. Branch on `errorCode`:
| `errorCode` | HTTP | Meaning |
|-------------|------|---------|
| `INVALID_INPUT` | 400 | malformed request; the message names the bad field |
| `UNAUTHORIZED` | 401 | auth required or failed (send `-u user:password`). ⚠️ The 401 body is plain text, NOT this envelope — `jq` dies with a parse error, see the guard in SKILL.md |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | on a **wait**: this session's waiter cap (16, combined signal+output) is full. On **quick-start**: the 50-session cap is full, so clean up before starting more |
| `CONFLICT` / `ALREADY_EXISTS` | 409 | conflicts with current state |
| `OPERATION_FAILED` | 422 | well-formed but could not be completed |
| `RATE_LIMITED` | 429 | per-owner or process-wide waiter pool is full — back off; switching sessions will not help |
| `INTERNAL_ERROR` | 500 | server bug |
`SESSION_BUSY` vs `RATE_LIMITED` on the wait endpoints is deliberate: the first means
"too many waiters on *this* session", the second means the *pool* is full.
## Sessions
| Task | Call |
|------|------|
| list sessions (metadata only, ~1.5 KB each, safe to poll) | `GET /api/v1/sessions` |
| one session (has `.data.pid`, `null` until the PTY spawns) | `GET /api/v1/sessions/:id` — ⚠️ **not a liveness check**: a worker that dies inside its pane keeps `status:"idle"` and a pid (the tmux attach client); `wait?until=exit` is the death check |
| unified list incl. history | `GET /api/v1/sessions/unified` → `.data.sessions[]` (NOT `.data[]`), and it folds in transcript history from the whole machine — never use it to verify cleanup; `GET /api/v1/sessions` is the cleanup check |
| start case + session in one call | `POST /api/v1/quick-start` |
| send input | `POST /api/v1/sessions/:id/input` |
| **read a worker's answer** (claude/codex) | `GET /api/v1/sessions/:id/last-response` → `.data.{text,timestamp}` — clean transcript text, no TUI noise. ⚠️ **Poll it**: the transcript flush lags the `stop` signal, so a read taken the instant send-and-wait returns is `""` (verified live). Also `""` before the first completed turn, and always `""` for `shell`/`opencode`/`gemini`/`antigravity` (no transcript) |
| read terminal (tail is in **BYTES**, raw ANSI) | `GET /api/v1/sessions/:id/terminal?tail=3000` → `.data.terminalBuffer` — for *diagnosis* (unsubmitted prompt?), not for reading answers |
| full tmux scrollback (context bomb; post-mortems only) | `GET /api/v1/sessions/:id/terminal?full=1` |
| background agents, one session | `GET /api/v1/sessions/:id/subagents` |
| background agents, global list | `GET /api/v1/subagents` (admin-only in multi-user mode) |
| server status / version | `GET /api/v1/status` → `.data.version` |
| delete one session (yours only, via `delete_session`) | `DELETE /api/v1/sessions/:id` — never call it bare; the fail-closed helper in SKILL.md §0 is the only self-protection that exists |
⚠️ `GET /api/v1/sessions/:id/output` → `.data.textOutput` looks like the obvious read
but stays **empty for interactive tmux-backed sessions** (it is fed only by the legacy
JSON-stream path). Verified empty on live claude and shell sessions. Use
`last-response` for claude/codex answers; only fall back to `terminal?tail=` for
hook-less modes, or to diagnose a prompt that was never submitted, and strip ANSI:
```bash
… | jq -r '.data.terminalBuffer' | sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g'
```
`POST /api/v1/quick-start` body (all optional):
`{"caseName":"worker-1","mode":"claude","sessionName":"w9-worker","effort":"high"}`
— `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity`; response is
`.data.{sessionId, caseName, casePath}`. Creates the case directory (a real directory
on the user's disk) if missing — do not retry it in a loop, and remember the name.
⚠️ **Branch on `.success` before reading `.data.sessionId`.** On any failure the field
is absent, `jq -r` prints the literal string `null`, and every later call then targets
`/api/v1/sessions/null`, burning the full readiness budget and reporting jq noise
instead of the real cause. Failure modes here are `SESSION_BUSY` (the **50-session
cap**, not the waiter cap), `FORBIDDEN`, `CONFLICT`, `OPERATION_FAILED` and
`INVALID_INPUT`; none of them are retryable in a loop.
⚠️ `caseName` resolves through the linked-cases registry first, so a name that happens
to match a case the user linked in lands in that **real repo**, not a fresh scratch
directory. Pick distinctive scratch names, and use a linked name deliberately when you
do want a worker in an existing checkout.
`POST /api/v1/sessions/:id/input` body:
`{"input":"one line\r","useMux":true,"clientId":"agent-1","seq":1}` plus optionally
`"wait"` / `"waitTimeout"` (below).
- ⚠️ **The input must contain `\r`** (the JSON escape, i.e. a real carriage return)
**or Enter is never sent**: the text is typed onto the worker's prompt and sits
there unsubmitted. Verified live — this is the number-one silent failure, and no
response field catches it: `delivered:true` means "written to the pane", not
"submitted". A `\r`-less send with `wait` reports `delivered:true` and then every
wait on that turn times out. Without `wait`, fire-and-forget returns an **empty**
`{"success":true,"data":{}}` — no `delivered`, no `duplicate`; those fields exist
only on the `wait` variant, so a fire-and-forget flow gets no delivery
confirmation at all.
- `input` must be single-line (newlines are stripped). To send a bare Enter (confirm
a dialog), send `{"input":"\r"}`.
- `clientId`+`seq` give exactly-once delivery: the server applies each pair at most
once. Increment `seq` per new input.
## The wait primitives
Three bounded long-polls. Shared semantics:
- **Timeout = HTTP 200** with `wait.timedOut:true`. Loop over short waits (60 s);
`tailscale serve` / cloudflared cut idle connections.
- Timeouts are **clamped** to `[1000, 600000]` ms (operator-tunable); the applied
value is echoed as `wait.timeoutMs` — read it back, never assume.
- All three nest the result under `.data.wait`, same shape, so one helper parses all.
- `.data.status` (post-wait `SessionStatus`) and `.data.limitPaused` ride along.
`limitPaused:true` means the session is paused on a usage limit and will emit
nothing until reset — a timeout is then *expected*; do not retry hard, and do not
kill the worker.
### Signals by mode
| Signal | Meaning | Available for |
|--------|---------|---------------|
| `idle` | output stabilized + prompt detected — heuristic, can flap mid-turn | every mode |
| `working` | session started producing output | every mode |
| `stop` | Claude Code `stop` hook — the definitive end-of-turn | `claude` only |
| `blocked` | `permission_prompt` / `elicitation_dialog` hook — the worker needs an answer | `claude` only |
| `exit` | PTY exited or session deleted | every mode |
Default `until` set: `stop,idle,exit`. On non-claude modes the server silently drops
`stop`/`blocked` from the *default* set (echoed back as `wait.until`, e.g.
`["idle","exit"]` on shell); requesting them *explicitly* there is a 400 naming the
mode. ⚠️ On hook-less modes the lifecycle signals are also **coarse in practice**: a
short shell command produced **no** `idle` transition within 60 s (verified live), so
a `fresh=1` / fresh-delivery wait can burn its whole timeout while the work finished
long ago. Synchronize hook-less modes with `wait-output` markers instead.
Two more places hooks go missing even in claude mode: **Docker cases** need
`CODEMAN_DOCKER_BRIDGE_HOOKS=1` on the server (without it only `idle`/`working`/
`exit` arrive), and **remote-SSH cases** run the agent on another host whose hooks may
never reach this server. When unsure, ask for `stop,idle,exit`.
⚠️ **Signals are edge-triggered with no history.** A signal that fires while no
waiter is registered is gone; no later wait can observe it (`until=stop` on a worker
whose turn already ended just times out, with or without `fresh` — verified live).
Register the waiter before the event can happen: send-and-wait does exactly that,
and `wait-output` markers with `from=buffer` are latched by construction. Never
fire-and-forget N prompts and then gather signal-waits worker by worker; every
worker that finishes before its gather is unobservable (see recipes.md Flow 3b).
### `GET /api/v1/sessions/:id/wait`
| Param | Default | Notes |
|-------|---------|-------|
| `until` | `stop,idle,exit` | comma list; unknown token → 400 naming it |
| `timeout` | 60000 | ms, clamped; applied value echoed as `wait.timeoutMs` |
| `fresh` | `0` | `1` requires an actual *transition*, ignoring the state at call time |
⚠️ A session whose PTY has not spawned (`pid:null`) or has exited counts as `exit`
**right now**: with the default set the call answers immediately
(`signal:"exit", immediate:true`). That is how you detect a dead worker cheaply — but
it also means "wait for my just-created session" needs the readiness recipe in
SKILL.md, not this endpoint.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Default | Notes |
|-------|---------|-------|
| `match` | required | literal substring, 1–200 chars, ANSI-stripped; chunk-straddling matches found; **no regex** — a `regex=` param is a 400 |
| `nocase` | `0` | case-insensitive compare; snippet keeps original casing |
| `from` | `now` | `buffer` scans the tail (~256 KB) of existing output first |
| `timeout` | 60000 | same clamp |
Four traps, all observed live:
1. **The echo of your own typed command is output.** A marker appearing verbatim in
the input line matches the moment the text is typed, before the command runs.
Split the marker with a shell variable: send `M=DONE; …; echo ${M}_1234\r`, wait
on `DONE_1234`.
2. **`from=now` misses text printed before the wait landed** — a marker echoed just
before the request registered timed out at full length. After sending a command,
always wait with `from=buffer`.
3. **`from=now` can also match too much**: tmux repaints old screen content as
ordinary output on attach/resize/redraw, so a *generic* marker (`BUILD OK`)
matches stale text. Unique-per-call markers (`DONE_$RANDOM`) make both `from`
modes safe.
4. **TUI output can be space-less in the stream.** Full-screen TUIs (claude, codex,
…) position words with cursor-movement escapes rather than literal spaces, so
the stripped stream can read `Yes,Itrustthisfolder` while the pane shows the
spaced phrase. Whether a given phrase keeps its spaces depends on how the TUI
drew it (observed live: some multi-word matches fire, some never do), so treat
multi-word matches against TUI screens as unreliable and match a **single
space-free token** (`trust`, `bypass`). Plain command output (shell workers,
`echo` lines) keeps real spaces and multi-word matches work there.
Build the query with `-G --data-urlencode` (a `+` in a hand-built query decodes to a
space). Result extras: `wait.matched`, `wait.match`, `wait.snippet` (bounded window
around the match, blank runs collapsed — the snippet is often all you need to read).
### `POST /api/v1/sessions/:id/input` with `wait`
| Field | Notes |
|-------|-------|
| `wait` | `true` (default signal set) or the same comma grammar as `until`; absent = historical fire-and-forget |
| `waitTimeout` | ms, same clamp |
Registers the waiter **before** typing, which closes the race where send-then-wait
sees the previous turn's idle state and returns instantly. Response adds `delivered`
and `duplicate` beside the standard `wait` object.
A **tagged duplicate** (same `clientId`+`seq` already applied) does not retype but
still honors `wait`, answering from the session's *current* state instead of
requiring a new transition (`delivered:false, duplicate:true` — verified: ~20 ms,
command ran exactly once). That is what makes the resend-identical-request loop in
SKILL.md correct: iteration 1 delivers and needs a transition; later iterations
resolve immediately if the turn ended in between. ⚠️ The flip side: a duplicate's
`immediate:true` answer is the current state and nothing more — an idle worker
whose prompt was never submitted (missing `\r`) produces the same
`signal:"idle", immediate:true` as one that finished the turn. Confirm from
`terminal?tail=` before reporting success; SKILL.md's loop shows where.
### Outcome parsing, in order
1. `wait.signal != null` (or `wait.matched == true`) — the thing happened.
`wait.immediate:true` rides along and means the condition already held at call
time; if that is not what you meant, you wanted `fresh=1` or send-and-wait.
2. `wait.timedOut` — poll boundary; loop again.
3. `wait.ended` — session deleted/torn down mid-wait; stop looping.
## Troubleshooting
| Symptom | Cause / fix |
|---------|-------------|
| every curl fails with a certificate error | you dropped `-k`; `CODEMAN_API_URL` is HTTPS with a self-signed cert |
| `jq: parse error` on every call | plain-text 401s: the server has a password. Check with `-w '%{http_code}'`, use the guard's `.env` fallback, and if no `.env` exists, stop and ask the user for credentials |
| input arrives but nothing happens; later waits all time out | the input had no `\r`, so Enter was never sent; the text is sitting on the worker's prompt. **Submitting it with `{"input":"\r"}` is the ONLY recovery** — Ctrl+U (0x15) and Esc do NOT clear the composer (verified live) — and the flush costs one turn in which the worker reasons about the junk; open the next real prompt with "ignore the garbled line above:" |
| `GET .../sessions/$CODEMAN_SESSION_ID` 404s | Docker case: the env id is truncated to 8 chars; find yourself with `startswith($SELF)`, and always self-compare by prefix, in both directions |
| `CODEMAN_MUX` unset but you seem to be in a session | remote-SSH case: the env vars are not exported there. Fail closed — refuse to act |
| connection refused from inside a container | a loopback-bound server is unreachable from a container, and `CODEMAN_DOCKER_BRIDGE_HOOKS=1` does **not** fix that: it opens a hooks-only listener, so hook events start flowing but `/api/v1/*` stays refused. Driving the API from inside a Docker case needs a reachable bind (an operator decision); report it, don't retry |
| wait routes 404 on a valid session id | read the `.error` text: a `Route ...` prefix means the server predates the wait endpoints (< 1.13.0; a dev build can serve them while reporting an older version, so probe, never version-compare) — poll `terminal?tail=` and say so. `Session ... not found` means your id is wrong, not the server |
| wait on `stop` never resolves | non-claude mode, or hooks not reaching the server (Docker/remote), or a case created by Codeman < 1.13.0 against an `--https` install (its hook curls lacked `-k` and TLS-failed silently; a 1.13.0+ server rewrites them the next time a session starts in that case). Use markers or `idle,exit` |
| new claude worker ignores its first prompt | it was showing the first-run trust dialog and Codeman's auto-accept missed; use the readiness recipe in SKILL.md (wait for `bypass` first, accept the dialog only as the bounded fallback) |
| `wait-output` times out although the pane shows the text | multi-word match against a TUI screen; the stream has no spaces there — match one token |
| `wait-output` matched instantly with stale text | generic marker + tmux repaint; use `DONE_$RANDOM` |
| 409 `SESSION_BUSY` on a wait | too many concurrent waiters on that session (cap 16 combined); reuse one wait per worker |
| 429 `RATE_LIMITED` on a wait | global/owner waiter pool full; back off, do not switch sessions |
+279
View File
@@ -0,0 +1,279 @@
# Worked orchestration flows
Loaded on demand from the `codeman` skill. Every flow assumes the SKILL.md §0 preamble
is in scope (`$API`, `$SELF`, `$CID`, `"${CURL[@]}"`, `delete_session`).
⚠️ **That preamble does not survive between tool calls**, so re-run it at the top of
every Bash call that uses these flows, in full. Re-pasting only part of it is the
failure mode the fail-closed `delete_session` exists to contain, and a `clientId` you
rebuild from `$$` changes per call, which turns the duplicate-resend loop in Flow 1
into a second typed prompt.
Track every session id you create; delete them (and only them) when done. The two
silent killers: **every input ends with `\r`**, and **markers must be split** so the
typed-line echo does not match them.
## Flow 1: claude worker, end to end
Start a worker, get it truly ready (trust dialog included), give it a task, wait for
the turn to finish, read the answer, clean up. Verified live: the stop hook resolves
the send-and-wait within seconds of the turn ending.
```bash
# 1. start (returns before the CLI inside is ready). ALWAYS check .success: on failure
# .data.sessionId is null, jq -r yields the string "null", and every step below
# then runs against /api/v1/sessions/null and reports jq noise, not the cause.
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-tests","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed"; exit 1; }
CREATED+=("$SID") # the cleanup list
SEQ=1 # $CID is the fixed literal from §0; never rebuild it from $$
# 2. readiness. "wait for idle" or "wait for ❯" is NOT readiness: a fresh session
# reports idle before anything spawned, and the first-run trust dialog contains ❯.
# Codeman CAN auto-accept that dialog, but the accept misses on some runs (both
# outcomes seen live), so: composer marker first, dialog only as the bounded
# fallback (a blind Enter up front would land in an already-ready composer).
# Stage 1 is SHORT on purpose: an already-trusted case matches in <1 s, while a
# virgin case can never pass it (the dialog is up) and always pays it in full —
# the long budget belongs to stage 3, after the dialog is answered.
# Single-token matches only: TUI text is space-less in the stream.
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# (pid != null proves startup only — a worker that later dies inside its pane keeps
# status "idle" and a pid. The death check is wait?until=exit.)
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
jq -e '.data.wait.matched' <<<"$R" >/dev/null || echo "worker $SID not ready; inspect terminal?tail="
fi
# 3. send-and-wait, looping on the IDENTICAL request (tagged duplicate: no retype).
# BOUNDED (a \r-less send would otherwise loop forever), body built with jq -n so
# quotes/backslashes/$ in a real prompt survive; note the appended \r.
PROMPT='run the unit tests and summarize failures in one line'
BODY=$(jq -n --arg p "$PROMPT" --arg c "$CID" --argjson s "$SEQ" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:60000}')
for TRY in $(seq 1 10); do
R=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' --data-binary "$BODY")
if jq -e '.data.wait.timedOut' <<<"$R" >/dev/null; then
jq -e '.data.limitPaused' <<<"$R" >/dev/null && sleep 60 # usage-limit pause: silence is expected
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # is the prompt sitting unsubmitted?
continue
fi
# Resolved — but duplicate + immediate is only "the session is idle NOW", which a
# never-submitted (\r-less) prompt also produces. Check before believing it:
if jq -e '.data.duplicate and .data.wait.immediate' <<<"$R" >/dev/null; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5
# prompt still on the ❯ composer line = never submitted; {"input":"\r"} is the
# only recovery, then loop again
fi
break
done
SEQ=$((SEQ+1))
# 4. interpret
case "$(jq -r '.data.wait.signal' <<<"$R")" in
stop) : ;; # definitive end of turn
idle) : ;; # heuristic — and if it rode a duplicate with
# immediate:true, it proves nothing ran (step 3)
exit) echo "worker died" ;;
null) jq -e '.data.wait.ended' <<<"$R" >/dev/null && echo "worker deleted mid-wait" ;;
esac
# 5. read the answer. For a claude worker this is last-response: clean transcript text,
# no TUI repaint noise. Do NOT scrape the terminal for this — a full-screen TUI
# draws with cursor moves, so the stripped buffer is nearly one long line and the
# answer arrives buried in redraw garbage.
# POLL it: the transcript flush lags the stop signal, so a single read taken the
# instant step 3 returned comes back "" even though the turn finished (verified live).
for _ in $(seq 1 10); do
TXT=$("${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
# (.data is {text,timestamp}; text is also "" before the first completed turn and
# always "" for shell/opencode/gemini/antigravity, which have no transcript — use
# the terminal tail there, and here only to diagnose an unsubmitted prompt.)
# 6. clean up — exact id, own list only, through the fail-closed §0 helper
delete_session "$SID"
```
Increment `SEQ` for every *new* input to the same worker. Reuse the same `SEQ` only to
re-ask about the same delivery (the duplicate-wait loop above).
## Flow 2: shell worker running a build, marker-synchronized
`shell` sessions have no hooks (`stop`/`blocked` are a 400 there), and their lifecycle
signals are coarse — a short command may emit no `idle` transition at all (verified
live), so send-and-wait can burn its whole timeout. The reliable pattern is a split,
unique marker plus `wait-output from=buffer`:
```bash
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"builder","mode":"shell"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed"; exit 1; }
CREATED+=("$SID")
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# Split marker: the typed line carries ${M}_N, only the OUTPUT carries DONE_N.
# An unsplit marker matches the echo of your own keystrokes before the build runs.
N="${RANDOM}_$$"; MARK="DONE_$N"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"codeman-build-1","seq":1}'
for TRY in $(seq 1 30); do # BOUNDED (30 min): a \r-less send makes an uncapped loop infinite
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=60000')
jq -e '.data.wait.matched' <<<"$R" >/dev/null && break
jq -e '.data.wait.ended' <<<"$R" >/dev/null && { echo "worker gone"; break; }
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # command still sitting unsubmitted?
done
jq -r '.data.wait.snippet' <<<"$R" # e.g. "DONE_123_456 rc=0" — the exit code rides the marker line
```
## Flow 3: fan out N workers, gather as each finishes
Start everything first, then gather. One in-flight wait per worker — the per-session
waiter cap is 16 and abandoned concurrent waits pile up against it.
```bash
declare -A WORKER MARKS
for task in lint typecheck unit; do
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"fan-'"$task"'","mode":"shell"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; echo "$task: spawn failed"; continue; }
WORKER[$task]=$SID; CREATED+=("$SID")
done
for task in "${!WORKER[@]}"; do
SID=${WORKER[$task]}
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
N="${task}_${RANDOM}"; MARKS[$task]="DONE_$N"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run '"$task"'; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"codeman-fan-'"$task"'","seq":1}'
done
for task in "${!WORKER[@]}"; do # sequential gather; each wait blocks until that worker is done
for TRY in $(seq 1 30); do # BOUNDED per worker, same reasoning as Flow 2
R=$("${CURL[@]}" -G "$API/api/v1/sessions/${WORKER[$task]}/wait-output" \
--data-urlencode "match=${MARKS[$task]}" --data-urlencode 'from=buffer' --data-urlencode 'timeout=60000')
jq -e '.data.wait.matched or .data.wait.ended' <<<"$R" >/dev/null && break
done
echo "$task: $(jq -r '.data.wait.snippet // "worker gone"' <<<"$R" | tail -1)"
done
```
## Flow 3b: fan out N CLAUDE workers
Send-and-wait is synchronous, so the shell-flow shape ("send everything, then
gather") does not translate directly: the send *is* the wait, and worker 2's prompt
would not go out until worker 1's turn ended. Two working patterns, both verified
live (and one anti-pattern, measured failing, replaced by B):
**A. Background the send-and-waits** (simplest; each resolved on `stop` while the
other was still running):
```bash
sendwait() { # $1=sid $2=prompt $3=seq — assumes the worker passed Flow 1's readiness
local body; body=$(jq -n --arg p "$2" --argjson s "$3" --arg c "codeman-fan-$1" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:600000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$1/input" \
-H 'Content-Type: application/json' --data-binary "$body" > "/tmp/fan-$1.json"
}
( sendwait "$SID1" 'refactor module A and reply DONE' 2 & \
sendwait "$SID2" 'write tests for module B and reply DONE' 2 & wait )
jq -c '.data.wait | {signal, waitedMs}' /tmp/fan-"$SID1".json /tmp/fan-"$SID2".json
```
One in-flight wait per worker keeps you far from the 16-per-session waiter cap.
**B. Fire-and-forget, then gather with output markers.** If you must send every
prompt before waiting on anything, do **not** gather with signal waits: signals
are edge-triggered with no history, so a `stop` that fires before the gather
reaches that worker is gone and unobservable afterwards — `fresh=1` cannot help,
and neither can omitting it (measured: worker 2's turn ended at +2 s, its
sequential `until=stop,exit&fresh=1` gather burned its full bounded 300 s and
reported nothing). Gather instead on a marker each worker prints itself, which
`from=buffer` re-finds no matter when it appeared:
```bash
# SIDS[1], SIDS[2] = worker ids that already passed Flow 1's readiness.
# The typed prompt must NOT contain the finished marker verbatim (your keystrokes
# echo into the output stream and would match instantly), so ask for it in halves:
declare -A TOK
for i in 1 2; do
TOK[$i]="${RANDOM}_$i"
BODY=$(jq -n --arg p "do task $i; when completely done print the word WORKDONE immediately followed by _${TOK[$i]}" \
--arg c "codeman-fan-$i" --argjson s 2 '{input:($p+"\r"),useMux:true,clientId:$c,seq:$s}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/${SIDS[$i]}/input" \
-H 'Content-Type: application/json' --data-binary "$BODY"
done
for i in 1 2; do # order no longer matters: the marker is latched in the buffer
"${CURL[@]}" -G "$API/api/v1/sessions/${SIDS[$i]}/wait-output" \
--data-urlencode "match=WORKDONE_${TOK[$i]}" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=600000' | jq -c '.data.wait | {matched, snippet}'
done
```
Use A unless you genuinely need to send everything before waiting on anything: A
needs no marker discipline, and resolves on the definitive `stop` instead of on
the worker remembering to print a token.
## Flow 4: watch for a worker stuck on a permission prompt
Claude workers can block on a permission dialog. `blocked` is a wait signal
(claude-mode only), so watch for it and surface the question to the user instead of
guessing an answer:
```bash
R=$("${CURL[@]}" "$API/api/v1/sessions/$SID/wait?until=stop,blocked,exit&timeout=60000")
if [ "$(jq -r '.data.wait.signal' <<<"$R")" = blocked ]; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" | jq -r '.data.terminalBuffer' \
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' | grep -v '^[[:space:]]*$' | tail -15
# show this to the user and ask how to answer; do NOT auto-confirm another
# session's permission prompt
fi
```
## Cleanup discipline
At the end of the conversation (or on abort), delete exactly what you created:
```bash
for id in "${CREATED[@]}"; do
delete_session "$id"
done
```
- Only ids from your own `CREATED` list. Never enumerate `/api/v1/sessions` and
delete by pattern; other sessions belong to the user.
- Always go through `delete_session`. It refuses an empty id, refuses when `$SELF` is
unset or too short to prove the target is not you, and prefix-checks in both
directions. A hand-written `curl -X DELETE`, or the old
`is_self "$id" || curl -X DELETE …`, has none of that: an undefined `is_self` exits
127 and the `||` branch deletes unguarded.
- If you created a *case* purely as scratch and the user confirmed it is disposable,
`DELETE /api/v1/cases/:name` removes it — but that recursively deletes the
directory from disk, so never do it without the user's explicit go-ahead for that
exact name.
+31 -3
View File
@@ -135,6 +135,7 @@ export abstract class AiCheckerBase<
// Active check state
protected checkMuxName: string | null = null;
protected checkTempFile: string | null = null;
protected checkStderrFile: string | null = null;
protected checkPromptFile: string | null = null;
protected checkPollTimer: NodeJS.Timeout | null = null;
protected checkTimeoutTimer: NodeJS.Timeout | null = null;
@@ -376,6 +377,7 @@ export abstract class AiCheckerBase<
const shortId = this.sessionId.slice(0, 8);
const timestamp = Date.now();
this.checkTempFile = join(tmpdir(), `${this.tempFilePrefix}-${shortId}-${timestamp}.txt`);
this.checkStderrFile = join(tmpdir(), `${this.tempFilePrefix}-stderr-${shortId}-${timestamp}.txt`);
this.checkPromptFile = join(tmpdir(), `${this.tempFilePrefix}-prompt-${shortId}-${timestamp}.txt`);
this.checkMuxName = `${this.muxNamePrefix}${shortId}`;
@@ -386,6 +388,7 @@ export abstract class AiCheckerBase<
// Ensure output temp file exists (empty) so we can poll it
writeFileSync(this.checkTempFile, '');
writeFileSync(this.checkStderrFile, '');
// Write prompt to file to avoid E2BIG error (argument list too long)
// The prompt can be 16KB+ which exceeds shell argument limits
@@ -396,7 +399,7 @@ export abstract class AiCheckerBase<
const modelArg = `--model "${this.config.model.replace(/"/g, '\\"')}"`;
const augmentedPath = getAugmentedPath();
const claudeCmd = `cat "${this.checkPromptFile}" | claude -p ${modelArg} --output-format text`;
const fullCmd = `export PATH="${augmentedPath}"; ${claudeCmd} > "${this.checkTempFile}" 2>&1; echo "${this.doneMarker}" >> "${this.checkTempFile}"; rm -f "${this.checkPromptFile}"`;
const fullCmd = `export PATH="${augmentedPath}"; ${claudeCmd} > "${this.checkTempFile}" 2> "${this.checkStderrFile}"; echo "${this.doneMarker}" >> "${this.checkTempFile}"; rm -f "${this.checkPromptFile}"`;
// Spawn tmux session
try {
@@ -461,18 +464,32 @@ export abstract class AiCheckerBase<
const output = content.replace(this.doneMarker, '').trim();
if (!output) {
return this.createErrorResult(`Empty output from ${this.checkDescription}`, durationMs);
const stderr = this.readStderrDiagnostic();
const detail = stderr ? `: ${stderr}` : '';
return this.createErrorResult(`Empty output from ${this.checkDescription}${detail}`, durationMs);
}
// Delegate to subclass for verdict parsing
const parsed = this.parseVerdict(output);
if (!parsed) {
return this.createErrorResult(`Could not parse verdict from: "${output.substring(0, 100)}"`, durationMs);
const stderr = this.readStderrDiagnostic();
const detail = stderr ? `; stderr: "${stderr}"` : '';
return this.createErrorResult(`Could not parse verdict from: "${output.substring(0, 100)}"${detail}`, durationMs);
}
return this.createResult(parsed.verdict, parsed.reasoning, durationMs);
}
private readStderrDiagnostic(): string {
if (!this.checkStderrFile || !existsSync(this.checkStderrFile)) return '';
try {
return readFileSync(this.checkStderrFile, 'utf-8').trim().substring(0, 200);
} catch {
return '';
}
}
private cleanupCheck(): void {
// Clear poll timer
if (this.checkPollTimer) {
@@ -509,6 +526,17 @@ export abstract class AiCheckerBase<
this.checkTempFile = null;
}
if (this.checkStderrFile) {
try {
if (existsSync(this.checkStderrFile)) {
unlinkSync(this.checkStderrFile);
}
} catch {
// Best effort cleanup
}
this.checkStderrFile = null;
}
if (this.checkPromptFile) {
try {
if (existsSync(this.checkPromptFile)) {
+322 -51
View File
@@ -12,15 +12,20 @@ import chalk from 'chalk';
import { createRequire } from 'module';
import http from 'node:http';
import https from 'node:https';
import { readFileSync } from 'node:fs';
import { isAbsolute } from 'node:path';
import { existsSync, readFileSync } from 'node:fs';
import { isAbsolute, join } from 'node:path';
import { homedir } from 'node:os';
import { dataPath } from './config/instance.js';
import { installAgentSkillInto, removeAgentSkillFrom, type AgentSkillApplyResult } from './hooks-config.js';
import { getSessionManager } from './session-manager.js';
import { getTaskQueue } from './task-queue.js';
import { getRalphLoop } from './ralph-loop.js';
import { getStore } from './state-store.js';
import { getErrorMessage } from './types.js';
import { isSupportedAttachmentExtension } from './attachment-registry.js';
import { daemonStatus, startDaemon, stopDaemon, type WebLaunchOptions } from './daemon-control.js';
import { installService, serviceStatus, uninstallService } from './service-installer.js';
import { isLoopbackBindHost, isUnauthenticatedNetworkAcknowledged } from './web/network-auth-policy.js';
const require = createRequire(import.meta.url);
const pkg = require('../package.json') as { version: string };
@@ -116,6 +121,112 @@ program
console.log(makeAttachmentMagicLink(filePath));
});
// ============ Skill Commands ============
/** Same registry the server resolves case names through (mirrors `case-routes.ts`). */
const LINKED_CASES_FILE = dataPath('linked-cases.json');
/**
* Case name to directory, checking `linked-cases.json` FIRST and falling back to the
* shared single-user cases dir. Mirrors `resolveCasePath()` in `case-routes.ts`, which
* is what the web UI and `quick-start` use. Without the linked-cases lookup this
* command rejected every case linked in from outside `~/codeman-cases` with
* "Case not found", even though the server resolved the same name fine.
*
* Sync and tolerant on purpose: a missing or malformed registry means "no linked
* cases", never a crash.
*/
function resolveCliCasePath(name: string): string {
try {
const linked = JSON.parse(readFileSync(LINKED_CASES_FILE, 'utf-8')) as Record<string, string>;
const target = linked?.[name];
if (typeof target === 'string' && target) return target;
} catch {
// no registry yet, or unreadable/invalid JSON: fall through to the cases dir
}
return join(homedir(), 'codeman-cases', name);
}
/**
* Resolve where `skill install` / `skill uninstall` operate. Global is
* `~/.claude/skills/codeman` (Claude Code's user-scope skill dir, read by every new
* session); `--case <name>` targets `<case>/.claude/skills/codeman`, resolved through
* `resolveCliCasePath()` above. The web server's automatic per-case injection
* (`agentSkillEnabled`) covers multi-user spaces; this CLI is a local operator tool
* and stays single-user.
*/
function resolveSkillTarget(options: { case?: string }): string {
if (options.case) {
const casePath = resolveCliCasePath(options.case);
if (!existsSync(casePath)) {
console.error(chalk.red(`✗ Case not found: ${casePath}`));
process.exit(1);
}
return join(casePath, '.claude', 'skills', 'codeman');
}
return join(homedir(), '.claude', 'skills', 'codeman');
}
/** Print an AgentSkillApplyResult for humans; exit non-zero when nothing was done. */
function reportSkillResult(result: AgentSkillApplyResult, target: string): void {
const messages: Record<AgentSkillApplyResult, { ok: boolean; text: string }> = {
installed: { ok: true, text: `Agent skill installed: ${target}` },
refreshed: { ok: true, text: `Agent skill refreshed (was stale): ${target}` },
unchanged: { ok: true, text: `Agent skill already up to date: ${target}` },
removed: { ok: true, text: `Agent skill removed: ${target}` },
absent: { ok: true, text: `Nothing to remove at ${target}` },
foreign: {
ok: false,
text: `${target} exists but is not Codeman-managed (no marker), refusing to touch it. Remove it yourself if you want the packaged skill there.`,
},
symlink: {
ok: false,
text: `${target} (or its parent) is a symlink, refusing to write through it.`,
},
};
const message = messages[result];
if (message.ok) {
console.log(chalk.green(`✓ ${message.text}`));
} else {
console.error(chalk.red(`✗ ${message.text}`));
process.exit(1);
}
}
const skillCmd = program
.command('skill')
.description('Manage the Codeman agent skill (lets an agent inside a session drive the API)');
skillCmd
.command('install')
.description('Install the agent skill globally (~/.claude/skills/codeman) or into one case')
.option('-g, --global', 'Install into ~/.claude/skills/codeman, picked up by every new session (the default)')
.option('-c, --case <name>', 'Install into <case>/.claude/skills/codeman instead (linked cases resolve too)')
.action(async (options: { global?: boolean; case?: string }) => {
try {
const target = resolveSkillTarget(options);
reportSkillResult(await installAgentSkillInto(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to install agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
skillCmd
.command('uninstall')
.description('Remove a Codeman-managed agent skill copy (never touches a user-authored one)')
.option('-g, --global', 'Remove from ~/.claude/skills/codeman (the default)')
.option('-c, --case <name>', 'Remove from <case>/.claude/skills/codeman instead (linked cases resolve too)')
.action(async (options: { global?: boolean; case?: string }) => {
try {
const target = resolveSkillTarget(options);
reportSkillResult(await removeAgentSkillFrom(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to remove agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
// ============ Session Commands ============
const sessionCmd = program.command('session').alias('s').description('Manage Claude sessions');
@@ -572,64 +683,224 @@ program
console.log('');
});
// ============ Web / daemon / service Commands ============
/** Shared option set for the commands that can launch a web server. */
function addWebLaunchOptions(cmd: Command): Command {
return cmd
.option('-H, --host <host>', 'Host to bind to', process.env.CODEMAN_HOST || '127.0.0.1')
.option('-p, --port <port>', 'Port to listen on (env: CODEMAN_PORT)', process.env.CODEMAN_PORT || '3000')
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
.option(
'--allow-unauthenticated-network',
'Allow non-loopback web access without CODEMAN_PASSWORD (dangerous; terminal control is exposed)'
)
.option(
'--multiuser',
'Enable opt-in multi-user mode (named users in ~/.codeman/users.json; env: CODEMAN_MULTIUSER)'
);
}
/** Normalize commander's strings into the shape daemon-control/service-installer take. */
function toWebLaunchOptions(options: {
host: string;
port: string;
https?: boolean;
titleHostname?: string;
allowUnauthenticatedNetwork?: boolean;
multiuser?: boolean;
}): WebLaunchOptions {
const port = parseInt(options.port, 10);
if (!Number.isInteger(port) || port <= 0 || port > 65535) {
console.error(chalk.red(`✗ Invalid port: ${options.port}`));
process.exit(1);
}
return {
host: options.host,
port,
https: !!options.https,
titleHostname: options.titleHostname,
allowUnauthenticatedNetwork: !!options.allowUnauthenticatedNetwork,
multiuser: !!options.multiuser,
};
}
/**
* The server prints this itself, but into a log file nobody reads when it is
* detached or supervised. Repeat it where the operator is actually looking.
*/
function warnIfUnauthenticatedNetwork(launch: WebLaunchOptions): void {
if (isLoopbackBindHost(launch.host)) return;
if (isUnauthenticatedNetworkAcknowledged(launch.allowUnauthenticatedNetwork)) return;
console.log(
chalk.yellow(
`⚠ Binding ${launch.host} without CODEMAN_PASSWORD: anyone who can reach this port gets terminal control.`
)
);
console.log(chalk.yellow(' Set CODEMAN_PASSWORD, or bind 127.0.0.1 and front it with tailscale serve.'));
}
// Web interface command
program
.command('web')
.description('Start the web interface')
.option('-H, --host <host>', 'Host to bind to', process.env.CODEMAN_HOST || '127.0.0.1')
.option('-p, --port <port>', 'Port to listen on (env: CODEMAN_PORT)', process.env.CODEMAN_PORT || '3000')
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
.option(
'--allow-unauthenticated-network',
'Allow non-loopback web access without CODEMAN_PASSWORD (dangerous; terminal control is exposed)'
)
.option('--multiuser', 'Enable opt-in multi-user mode (named users in ~/.codeman/users.json; env: CODEMAN_MULTIUSER)')
.action(async (options) => {
// The flag is surfaced to the rest of the process via the env var so
// isMultiUserMode() has a single source of truth (see config/multiuser.ts).
if (options.multiuser) process.env.CODEMAN_MULTIUSER = '1';
const { startWebServer } = await import('./web/server.js');
const host = options.host;
const port = parseInt(options.port, 10);
const https = !!options.https;
const titleHostname = options.titleHostname;
const allowUnauthenticatedNetwork = !!options.allowUnauthenticatedNetwork;
const protocol = https ? 'https' : 'http';
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
const webCmd = addWebLaunchOptions(program.command('web').description('Start the web interface'))
.option('-d, --daemon', 'Run detached in the background; survives the shell, logs to <data dir>/web.log')
.option('--stop', 'Stop a server started with --daemon')
.option('--status', 'Report whether a detached server is running');
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
webCmd.action(async (options) => {
// The flag is surfaced to the rest of the process via the env var so
// isMultiUserMode() has a single source of truth (see config/multiuser.ts).
if (options.multiuser) process.env.CODEMAN_MULTIUSER = '1';
const launch = toWebLaunchOptions(options);
try {
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
if (https) {
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
if (options.stop) {
const result = await stopDaemon(launch);
if (result.ok && result.reason === 'not-running') {
console.log(chalk.gray(`○ ${result.message}`));
return;
}
if (result.ok) {
console.log(chalk.green(`✓ ${result.message ?? `Stopped Codeman (pid ${result.pid})`}`));
console.log(chalk.gray(' Your agents keep running in tmux.'));
return;
}
console.error(chalk.red(`✗ ${result.message ?? 'Could not stop the server'}`));
process.exit(1);
}
if (options.status) {
const status = await daemonStatus(launch);
if (status.responding) {
const version = status.version ? ` (v${status.version})` : '';
console.log(chalk.green(`✓ Responding at ${status.url}${version}`));
} else {
console.log(chalk.yellow(`○ Nothing answering at ${status.url}`));
}
console.log(` Daemon pid: ${status.running ? chalk.green(String(status.pid)) : chalk.gray('not running')}`);
console.log(chalk.gray(` Pidfile: ${status.pidFile}`));
console.log(chalk.gray(` Log: ${status.logPath}`));
if (!status.running && status.responding) {
console.log(chalk.gray(' (running, but not started with --daemon: probably a service or a foreground run)'));
}
return;
}
if (options.daemon) {
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Starting Codeman in the background...'));
const result = await startDaemon(launch);
if (result.ok) {
console.log(chalk.green(`\n✓ Codeman is running at ${result.url} (pid ${result.pid})`));
console.log(chalk.gray(` Logs: ${result.logPath}`));
console.log(chalk.gray(' Stop it with: codeman web --stop'));
console.log(chalk.gray(' Want it back after a reboot? codeman service install'));
return;
}
console.error(chalk.red(`\n✗ ${result.message ?? 'Failed to start'}`));
process.exit(1);
}
const { startWebServer } = await import('./web/server.js');
const host = launch.host;
const port = launch.port;
const https = launch.https;
const titleHostname = options.titleHostname;
const allowUnauthenticatedNetwork = launch.allowUnauthenticatedNetwork ?? false;
const protocol = https ? 'https' : 'http';
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
try {
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
if (https) {
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
let shuttingDown = false;
const shutdown = async (signal: string) => {
if (shuttingDown) return;
shuttingDown = true;
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
try {
await server.stop();
} catch (err) {
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
process.exit(0);
};
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
process.on('SIGHUP', () => shutdown('SIGHUP'));
} catch (err) {
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
process.exit(1);
}
});
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
let shuttingDown = false;
const shutdown = async (signal: string) => {
if (shuttingDown) return;
shuttingDown = true;
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
try {
await server.stop();
} catch (err) {
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
}
process.exit(0);
};
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
process.on('SIGHUP', () => shutdown('SIGHUP'));
} catch (err) {
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
// Supervised service: the "still there after a reboot" answer, where `web -d` is
// the "still there after I close this shell" one (issue #231).
const serviceCmd = program
.command('service')
.description('Manage the background service (systemd user unit on Linux, LaunchAgent on macOS)');
addWebLaunchOptions(
serviceCmd.command('install').description('Install and start the service, then verify it answers')
).action(async (options) => {
const launch = toWebLaunchOptions(options);
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Installing the Codeman service...'));
const result = await installService(launch);
for (const warning of result.warnings ?? []) console.log(chalk.yellow(`⚠ ${warning}`));
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
console.log(chalk.gray(` Unit: ${result.unitPath}`));
if (process.env.CODEMAN_PASSWORD) {
console.log(
chalk.yellow(
' Note: CODEMAN_PASSWORD was NOT copied into the unit file. Add it there yourself if the service needs auth.'
)
);
}
});
serviceCmd
.command('uninstall')
.description('Stop the service and remove its unit file')
.action(() => {
const result = uninstallService();
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
});
addWebLaunchOptions(
serviceCmd.command('status').description('Show whether the service is installed and running')
).action(async (options) => {
const status = await serviceStatus(toWebLaunchOptions(options));
if (!status.kind) {
console.log(chalk.yellow(`No supported supervisor on ${process.platform}. Use \`codeman web -d\` instead.`));
return;
}
console.log(` Supervisor: ${status.kind} (${status.name})`);
console.log(` Unit file: ${status.installed ? chalk.green(status.unitPath) : chalk.gray('not installed')}`);
console.log(` Loaded: ${status.loaded ? chalk.green('yes') : chalk.gray('no')}`);
const version = status.version ? ` (v${status.version})` : '';
console.log(
` Responding: ${status.responding ? chalk.green(`yes at ${status.url}${version}`) : chalk.gray(`no at ${status.url}`)}`
);
});
// ============ Multi-user Commands ============
//
// Operate directly on ~/.codeman/users.json (via user-store) with NO running
+112
View File
@@ -0,0 +1,112 @@
/**
* @fileoverview Bounds for the agent wait primitives.
*
* These back the blocking endpoints an agent uses to orchestrate other sessions
* (`GET /api/sessions/:id/wait`, `GET /api/sessions/:id/wait-output`, and the
* `wait` field on `POST /api/sessions/:id/input`). Plan: `docs/agent-control-plan.md`.
*
* Why every value is bounded:
* - An unbounded long-poll is a socket leak. A caller that asks for a 12-hour wait
* and walks away holds a connection (and a waiter, and a timer) until the process
* restarts, so `MAX_WAIT_MS` is a hard ceiling applied server-side.
* - `DEFAULT_WAIT_MS` is deliberately short (60s). Production is reached through
* `tailscale serve` and users also run cloudflared tunnels; both can cut an idle
* connection, so the documented pattern is a client-side loop over short waits
* rather than one very long call. Fastify itself is happy to hold the request
* (`requestTimeout` defaults to 0, and `keepAliveTimeout` applies between
* requests, not to an in-flight one), the intermediaries are the constraint.
* - The waiter caps mirror `MAX_SSE_CLIENTS` in `map-limits.ts`: each pending
* waiter costs an open HTTP response plus a timer, so the pool is capped rather
* than queued. Exceeding a cap is an explicit error, never a silent wait.
* - There are THREE caps, not two, because a process-wide pool with no per-user
* dimension lets one user deny the primitive to everyone else. `middleware/auth.ts`
* already treats that shape as a bug (its `userFailures` bucket exists so "one user
* behind a NAT can't lock out everyone else"); `MAX_WAITERS_PER_OWNER` is the same
* idea for waiters. It applies only when the caller has an owner, so single-user
* mode is byte-identical to having no owner cap at all.
*
* All values are env-overridable and clamped to sane hard bounds, so a typo in an
* env var degrades to the default instead of disabling the protection.
*
* @module config/agent-wait
*/
/** Absolute floor for any wait, in ms. Sub-second waits are polling, not waiting. */
export const MIN_WAIT_MS = 1_000;
/** Ceiling the operator-configurable maximum is itself clamped to. */
const HARD_MAX_WAIT_MS = 3_600_000;
function envInt(name: string, fallback: number, min: number, max: number): number {
const raw = parseInt(process.env[name] || '', 10);
if (!Number.isFinite(raw) || raw <= 0) return fallback;
return Math.max(min, Math.min(max, raw));
}
/** Longest a single wait may block. Requests above this are clamped down, not rejected. */
export const MAX_WAIT_MS = envInt('CODEMAN_WAIT_MAX_MS', 600_000, MIN_WAIT_MS, HARD_MAX_WAIT_MS);
/** Used when the caller omits `timeout`. Never exceeds MAX_WAIT_MS. */
export const DEFAULT_WAIT_MS = Math.min(
envInt('CODEMAN_WAIT_DEFAULT_MS', 60_000, MIN_WAIT_MS, HARD_MAX_WAIT_MS),
MAX_WAIT_MS
);
/** Concurrent waiters (signal + output) allowed against one session. */
export const MAX_WAITERS_PER_SESSION = envInt('CODEMAN_WAIT_MAX_PER_SESSION', 16, 1, 256);
/**
* Ceiling the operator-configurable total is itself clamped to.
*
* 512 rather than the 4096 this started at. Every other knob in this file degrades
* safely on a bad value; a 4096 ceiling instead lets a well-meaning operator turn the
* protection into the problem, since 4096 concurrent held responses (each an open
* socket, a timer and a pending promise) exceeds the 1024 soft `RLIMIT_NOFILE` that is
* still the default on most Linux distros, before counting PTYs, SSE clients and
* WebSockets. 512 is ~5x `MAX_SSE_CLIENTS` (100, the pool this one is modelled on), so
* the knob stays useful for a busy orchestration host while the whole server still fits
* inside a default fd budget with room to spare.
*/
const HARD_MAX_WAITERS_TOTAL = 512;
/** Concurrent waiters allowed across every session in the process. */
export const MAX_WAITERS_TOTAL = envInt('CODEMAN_WAIT_MAX_TOTAL', 128, 1, HARD_MAX_WAITERS_TOTAL);
/**
* Concurrent waiters allowed for one owner (multi-user mode's `Session.owner`).
*
* Sits between the per-session cap (16) and the process-wide one (128): high enough
* that one user orchestrating several workers at once never trips it, low enough that
* a single user cannot occupy the whole pool and deny the primitive to everyone else,
* admin included. Ignored entirely when the caller has no owner, which is every
* request in single-user mode.
*/
export const MAX_WAITERS_PER_OWNER = envInt('CODEMAN_WAIT_MAX_PER_OWNER', 48, 1, HARD_MAX_WAITERS_TOTAL);
/** Bounds on the literal `match` string accepted by wait-output. */
export const MIN_MATCH_LENGTH = 1;
export const MAX_MATCH_LENGTH = 200;
/**
* Tail of the terminal buffer scanned by `wait-output?from=buffer`.
*
* The buffer itself runs to 32MB. Scanning all of it would be an ANSI strip over
* 32MB (a full second copy) on a request an agent may issue in a loop, and the
* question `from=buffer` answers is "did this appear recently", not "ever". The
* tail is continuous with the live stream, since `_terminalBuffer.append(data)`
* and `emit('terminal', data)` receive the same bytes.
*/
export const MAX_BUFFER_SCAN_BYTES = envInt('CODEMAN_WAIT_BUFFER_SCAN_BYTES', 256 * 1024, 4 * 1024, 8 * 1024 * 1024);
/** Characters of surrounding output returned either side of a wait-output match. */
export const MAX_SNIPPET_CONTEXT = 80;
/**
* Clamp a caller-supplied timeout into [MIN_WAIT_MS, MAX_WAIT_MS].
* Absent / non-numeric / non-finite input falls back to DEFAULT_WAIT_MS.
*/
export function clampWaitMs(value: unknown): number {
const n = typeof value === 'string' ? Number(value) : value;
if (typeof n !== 'number' || !Number.isFinite(n)) return DEFAULT_WAIT_MS;
return Math.max(MIN_WAIT_MS, Math.min(MAX_WAIT_MS, Math.trunc(n)));
}
+32
View File
@@ -0,0 +1,32 @@
/**
* @fileoverview Supervisor identity (systemd unit name / launchd job label).
*
* Three things now write or look for the same supervisor job: `install.sh`, the
* in-app self-updater (`web/self-update.ts` detects it to decide how to restart),
* and `codeman service install`. The names live here so they cannot drift apart,
* because a mismatch is silent in the worst way: `service install` would happily
* create a SECOND job alongside the installer's, and two servers sharing one data
* dir and one tmux socket attach PTYs to each other's live sessions
* (see config/instance.ts).
*
* The names are instance-scoped for exactly that reason: a `CODEMAN_INSTANCE=beta`
* build writing `com.codeman.web` would overwrite the production LaunchAgent. The
* DEFAULT instance keeps the historical names byte-identical, so existing installs
* and every unit install.sh has already written are unaffected.
*
* @module config/service-names
*/
import { CODEMAN_INSTANCE } from './instance.js';
/**
* Instance name reduced to characters that are safe in a filename and in a
* launchd label. `CODEMAN_INSTANCE` is arbitrary operator input.
*/
const SAFE_INSTANCE = CODEMAN_INSTANCE.replace(/[^A-Za-z0-9_-]/g, '').slice(0, 32);
/** systemd user unit: `codeman-web.service`, or `codeman-web-beta.service` for a beta. */
export const SYSTEMD_UNIT = `codeman-web${SAFE_INSTANCE ? `-${SAFE_INSTANCE}` : ''}.service`;
/** launchd job label: `com.codeman.web`, or `com.codeman.beta.web` for a beta. */
export const LAUNCHD_LABEL = SAFE_INSTANCE ? `com.codeman.${SAFE_INSTANCE}.web` : 'com.codeman.web';
+496
View File
@@ -0,0 +1,496 @@
/**
* @fileoverview Detached `codeman web` control: start (-d), stop, status.
*
* Backs `codeman web -d`, `codeman web --stop` and `codeman web --status`. The
* server itself is unchanged; this module re-launches the SAME entry script in a
* new session (`detached: true` calls setsid), so the child has no controlling
* terminal and no shell job entry. That is what actually makes it outlive the
* shell: `nohup` does not, because Node re-arms SIGHUP to its default disposition
* even when it inherits "ignore", and `cli.ts` installs a SIGHUP handler that
* shuts the server down gracefully (issue #231).
*
* Two rules shape the rest of the module:
*
* 1. **Never start a second server on one data dir.** `~/.codeman` and the
* `tmux -L codeman` socket are process-wide (config/instance.ts), so a second
* instance discovers and attaches PTYs to the first one's live sessions and
* starts resizing them. A double `-d` therefore has to be a hard error, which
* means checking both the pidfile AND the port before spawning.
* 2. **Never report success we have not seen.** The parent polls `/api/status`
* until the child answers (or dies) before printing a URL. A port clash or a
* missing dependency otherwise looks exactly like a clean start.
*
* Pure helpers (arg building, URL building, pidfile parsing, the process-identity
* check) are exported separately so they can be unit-tested without spawning.
*
* @module daemon-control
*/
import { spawn, execFileSync } from 'node:child_process';
import { appendFileSync, closeSync, existsSync, openSync, readFileSync, unlinkSync, writeFileSync } from 'node:fs';
import http from 'node:http';
import https from 'node:https';
import { dataPath } from './config/instance.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
/** How long to wait for a freshly spawned server to answer `/api/status`. */
const START_TIMEOUT_MS = 30_000;
/** How long to wait for a SIGTERM'd server to actually exit before giving up. */
const STOP_TIMEOUT_MS = 15_000;
/** Poll interval while waiting for either of the above. */
const POLL_INTERVAL_MS = 250;
/** The `web` command's options, as far as a detached relaunch cares about them. */
export interface WebLaunchOptions {
host: string;
port: number;
https: boolean;
titleHostname?: string;
allowUnauthenticatedNetwork?: boolean;
multiuser?: boolean;
}
export interface StartResult {
ok: boolean;
pid?: number;
url?: string;
/** Machine-readable failure cause; `undefined` on success. */
reason?: 'already-running' | 'exited' | 'timeout';
message?: string;
logPath: string;
}
export interface StopResult {
ok: boolean;
pid?: number;
reason?: 'not-running' | 'foreign-pid' | 'timeout' | 'no-pidfile-but-responding';
message?: string;
}
export interface DaemonStatus {
pid: number | null;
/** The pid in the pidfile is alive AND still looks like a Codeman web process. */
running: boolean;
/** Something answered `/api/status` at the expected address. */
responding: boolean;
version?: string;
url: string;
pidFile: string;
logPath: string;
}
// ─────────────────────────────────────────────────────────────────────────────
// Pure helpers
// ─────────────────────────────────────────────────────────────────────────────
/** Rebuild the `web` argv for the child, dropping the daemon flags themselves. */
export function buildWebArgs(options: WebLaunchOptions): string[] {
const args = ['web', '--host', options.host, '--port', String(options.port)];
if (options.https) args.push('--https');
if (options.titleHostname) args.push('--title-hostname', options.titleHostname);
if (options.allowUnauthenticatedNetwork) args.push('--allow-unauthenticated-network');
if (options.multiuser) args.push('--multiuser');
return args;
}
/**
* Connectable address for this bind. A wildcard bind is not itself connectable,
* so `0.0.0.0` / `::` become loopback; a bare IPv6 literal gets bracketed.
*/
export function buildBaseUrl(options: WebLaunchOptions): string {
const protocol = options.https ? 'https' : 'http';
let host = options.host.trim();
if (host === '0.0.0.0' || host === '::' || host === '') host = '127.0.0.1';
if (host.includes(':') && !host.startsWith('[')) host = `[${host}]`;
return `${protocol}://${host}:${options.port}`;
}
/** The endpoint polled for readiness. */
export function buildStatusUrl(options: WebLaunchOptions): string {
return `${buildBaseUrl(options)}/api/status`;
}
/** Parse a pidfile body. Rejects garbage, and pid 1 (init is never ours). */
export function parsePidFileContents(text: string): number | null {
const trimmed = text.trim();
if (!/^\d+$/.test(trimmed)) return null;
const pid = Number.parseInt(trimmed, 10);
if (!Number.isSafeInteger(pid) || pid <= 1) return null;
return pid;
}
/**
* Does this command line look like a Codeman web server?
*
* Pids are recycled, and a stale pidfile pointing at whatever inherited the
* number is a live footgun: `codeman web --stop` must not SIGTERM an unrelated
* process. Both the npm bin (`codeman`/`aicodeman`) and the direct entry
* (`node dist/index.js web`, `tsx src/index.ts web`) have to match.
*/
export function looksLikeCodemanWeb(command: string | null | undefined): boolean {
if (!command) return false;
if (!/(^|\s)web(\s|$)/.test(command)) return false;
return /(^|[/\s])(ai)?codeman(\s|$)/.test(command) || /index\.(js|ts)(\s|$)/.test(command);
}
// ─────────────────────────────────────────────────────────────────────────────
// Paths
// ─────────────────────────────────────────────────────────────────────────────
/**
* Resolved at call time, not module load: tests swap `HOME` per file, and the
* data dir is derived from it (see test/setup.ts).
*/
export function pidFilePath(): string {
return dataPath('web.pid');
}
/** Where a detached server's stdout/stderr is appended. */
export function logFilePath(): string {
return dataPath('web.log');
}
// ─────────────────────────────────────────────────────────────────────────────
// Process probing
// ─────────────────────────────────────────────────────────────────────────────
/** Signal 0 liveness check. EPERM means the pid exists but is not ours. */
export function isProcessAlive(pid: number): boolean {
try {
process.kill(pid, 0);
return true;
} catch (err) {
return (err as NodeJS.ErrnoException).code === 'EPERM';
}
}
/** Full command line of a pid, or null. `-o command=` is portable to macOS. */
export function readProcessCommand(pid: number): string | null {
try {
const out = execFileSync('ps', ['-o', 'command=', '-p', String(pid)], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'ignore'],
});
return out.trim() || null;
} catch {
return null;
}
}
/** Read the pidfile, returning null when it is missing, empty or malformed. */
export function readPidFile(): number | null {
const file = pidFilePath();
if (!existsSync(file)) return null;
try {
return parsePidFileContents(readFileSync(file, 'utf-8'));
} catch {
return null;
}
}
function removePidFile(): void {
try {
unlinkSync(pidFilePath());
} catch {
/* already gone */
}
}
/** Pid of a live Codeman web server recorded in the pidfile, or null. */
export function readLivePid(): number | null {
const pid = readPidFile();
if (pid === null) return null;
if (!isProcessAlive(pid)) return null;
// A recycled pid is not ours. `ps` can also legitimately fail (containers with
// no procps); treat "cannot tell" as ours rather than orphaning the pidfile.
const command = readProcessCommand(pid);
if (command !== null && !looksLikeCodemanWeb(command)) return null;
return pid;
}
// ─────────────────────────────────────────────────────────────────────────────
// HTTP readiness probe
// ─────────────────────────────────────────────────────────────────────────────
export interface ProbeResult {
/** A Codeman server answered. A 401 counts: auth is active, the server is up. */
up: boolean;
version?: string;
}
/**
* Probe `/api/status`. Self-signed certs are accepted (`--https` generates one),
* and 401 counts as up because `CODEMAN_PASSWORD` gates that route. The body is
* checked so an unrelated service squatting on the port is not read as success.
*/
export function probeServer(url: string, timeoutMs = 2000): Promise<ProbeResult> {
return new Promise((resolve) => {
let settled = false;
const done = (result: ProbeResult) => {
if (settled) return;
settled = true;
resolve(result);
};
let target: URL;
try {
target = new URL(url);
} catch {
done({ up: false });
return;
}
const transport = target.protocol === 'https:' ? https : http;
const req = transport.request(
{
protocol: target.protocol,
hostname: target.hostname,
port: target.port,
path: target.pathname,
method: 'GET',
rejectUnauthorized: false,
timeout: timeoutMs,
headers: { Accept: 'application/json' },
},
(res) => {
if (res.statusCode === 401) {
res.resume();
done({ up: true });
return;
}
let body = '';
res.setEncoding('utf-8');
res.on('data', (chunk: string) => {
if (body.length < 4096) body += chunk;
});
res.on('end', () => {
if (!body.includes('"success"')) {
done({ up: false });
return;
}
let version: string | undefined;
try {
version = (JSON.parse(body) as { data?: { version?: string } }).data?.version;
} catch {
/* body was truncated at 4KB; up is still true */
}
done({ up: true, version });
});
res.on('error', () => done({ up: false }));
}
);
req.on('timeout', () => {
req.destroy();
done({ up: false });
});
req.on('error', () => done({ up: false }));
req.end();
});
}
function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
// ─────────────────────────────────────────────────────────────────────────────
// Start / stop / status
// ─────────────────────────────────────────────────────────────────────────────
/**
* The script to relaunch. `process.execArgv` is carried over with it so a dev
* run under tsx (whose execArgv holds the tsx loader flags) re-launches through
* tsx instead of handing a `.ts` file to bare node.
*/
function entryScript(): string {
const script = process.argv[1];
if (!script) throw new Error('cannot determine the codeman entry script to relaunch');
return script;
}
/** Marks one launch in the append-only log so a tail cannot mix two runs. */
const LOG_SEPARATOR = '=== codeman web start';
/**
* Last few lines of the daemon log, for reporting a failed start. The log is
* append-only across launches, so the tail starts at the last separator when
* there is one: otherwise a crash report is padded with the previous run's
* cheerful startup banner.
*/
export function tailLog(maxLines = 15): string {
try {
const lines = readFileSync(logFilePath(), 'utf-8').trimEnd().split('\n');
const start = lines.map((line) => line.startsWith(LOG_SEPARATOR)).lastIndexOf(true);
const current = start === -1 ? lines : lines.slice(start + 1);
return current.slice(-maxLines).join('\n');
} catch {
return '';
}
}
/**
* Spawn a detached `codeman web` and wait until it answers before returning.
* Refuses when a server is already up on this data dir (see rule 1 in the module
* docblock).
*/
export async function startDaemon(options: WebLaunchOptions): Promise<StartResult> {
const logPath = logFilePath();
const url = buildBaseUrl(options);
const statusUrl = buildStatusUrl(options);
const existingPid = readLivePid();
if (existingPid !== null) {
return {
ok: false,
reason: 'already-running',
pid: existingPid,
logPath,
message: `a Codeman server is already running (pid ${existingPid}). Stop it with \`codeman web --stop\` first.`,
};
}
const alreadyServing = await probeServer(statusUrl, 1500);
if (alreadyServing.up) {
return {
ok: false,
reason: 'already-running',
logPath,
url,
message: `something is already serving ${url}. Two servers on one data dir attach to each other's tmux sessions, so refusing to start.`,
};
}
// A pidfile that survived a crash: the process is gone, so it is just litter.
if (readPidFile() !== null) removePidFile();
const args = buildWebArgs(options);
try {
appendFileSync(logPath, `\n${LOG_SEPARATOR} ${new Date().toISOString()} ===\n`, 'utf-8');
} catch {
/* the spawn below reports a genuinely unwritable log */
}
const logFd = openSync(logPath, 'a');
let child;
try {
child = spawn(process.execPath, [...process.execArgv, entryScript(), ...args], {
detached: true,
stdio: ['ignore', logFd, logFd],
env: process.env,
});
} finally {
closeSync(logFd);
}
let exited = false;
child.on('exit', () => {
exited = true;
});
child.on('error', () => {
exited = true;
});
const pid = child.pid;
if (pid === undefined) {
return { ok: false, reason: 'exited', logPath, message: 'failed to spawn the server process' };
}
writeFileSync(pidFilePath(), `${pid}\n`, 'utf-8');
const deadline = Date.now() + START_TIMEOUT_MS;
while (Date.now() < deadline) {
if (exited) {
removePidFile();
child.unref();
return {
ok: false,
reason: 'exited',
logPath,
message: `the server exited during startup. Last lines of ${logPath}:\n${tailLog()}`,
};
}
const probe = await probeServer(statusUrl, 1000);
if (probe.up) {
child.unref();
return { ok: true, pid, url, logPath };
}
await sleep(POLL_INTERVAL_MS);
}
child.unref();
return {
ok: false,
reason: 'timeout',
pid,
url,
logPath,
message: `the server did not answer ${url} within ${START_TIMEOUT_MS / 1000}s. It may still be starting; check ${logPath}.`,
};
}
/** SIGTERM the recorded server and wait for it to actually exit. */
export async function stopDaemon(options: WebLaunchOptions): Promise<StopResult> {
const pid = readPidFile();
if (pid === null) {
const probe = await probeServer(buildStatusUrl(options), 1500);
if (probe.up) {
return {
ok: false,
reason: 'no-pidfile-but-responding',
message:
'a server is responding but there is no pidfile, so it was not started with `-d`. If it is a service use `codeman service uninstall` (or stop the unit); otherwise `pkill -f "index.js web"`.',
};
}
return { ok: true, reason: 'not-running', message: 'no daemon is running; nothing to stop' };
}
if (!isProcessAlive(pid)) {
removePidFile();
return { ok: true, pid, message: `stale pidfile removed (pid ${pid} was not running)` };
}
const command = readProcessCommand(pid);
if (command !== null && !looksLikeCodemanWeb(command)) {
return {
ok: false,
reason: 'foreign-pid',
pid,
message: `pid ${pid} is not a Codeman server (${command}). Refusing to signal it; delete ${pidFilePath()} if it is stale.`,
};
}
// SIGTERM, never SIGKILL: cli.ts flushes state on the way out.
try {
process.kill(pid, 'SIGTERM');
} catch (err) {
return { ok: false, reason: 'foreign-pid', pid, message: `could not signal pid ${pid}: ${String(err)}` };
}
const deadline = Date.now() + STOP_TIMEOUT_MS;
while (Date.now() < deadline) {
if (!isProcessAlive(pid)) {
removePidFile();
return { ok: true, pid };
}
await sleep(POLL_INTERVAL_MS);
}
return {
ok: false,
reason: 'timeout',
pid,
message: `pid ${pid} did not exit within ${STOP_TIMEOUT_MS / 1000}s. Force it with \`kill -9 ${pid}\` if you are sure.`,
};
}
/** Report on both halves: the recorded process, and whether the port answers. */
export async function daemonStatus(options: WebLaunchOptions): Promise<DaemonStatus> {
const url = buildBaseUrl(options);
const pid = readPidFile();
const probe = await probeServer(buildStatusUrl(options), 2000);
return {
pid,
running: readLivePid() !== null,
responding: probe.up,
version: probe.version,
url,
pidFile: pidFilePath(),
logPath: logFilePath(),
};
}
+388 -30
View File
@@ -10,13 +10,15 @@
* Key exports:
* - `generateHooksConfig()` — returns hooks object for settings.local.json
* - `writeHooksConfig(casePath)` — writes hooks + env config to disk
* - `ensureCodemanHooks(casePath)` — safely installs/updates hooks for a managed case
* - `updateCaseEnvVars(casePath, envVars)` — merges env vars into settings
*
* Hook events generated: `idle_prompt`, `permission_prompt`, `elicitation_dialog`,
* `stop`, `teammate_idle`, `task_completed`
*
* Hook categories: `Notification` (3 matchers), `Stop` (1), `TeammateIdle` (1),
* `TaskCompleted` (1), `PostToolUse` (1 self-contained background Bash rewake)
* Hook categories: `Notification` (3 matchers), `Stop` (1), `SubagentStop` (1),
* `TeammateIdle` (1), `TaskCompleted` (1), `PostToolUse` (1 self-contained
* background Bash rewake)
*
* @dependencies types (HookEventType), config/auth-config (HOOK_TIMEOUT_SECONDS)
* @consumedby web/server (session creation), session-cli-builder (env setup)
@@ -25,8 +27,9 @@
*/
import { existsSync } from 'node:fs';
import { readFile, writeFile, mkdir } from 'node:fs/promises';
import { join } from 'node:path';
import { readFile, writeFile, mkdir, lstat, readdir, unlink, rmdir } from 'node:fs/promises';
import { join, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
import type { HookEventType } from './types.js';
import { HOOK_TIMEOUT_SECONDS } from './config/auth-config.js';
@@ -52,15 +55,19 @@ const BACKGROUND_WAKE_MARKER_PREFIX = 'CODEMAN_BACKGROUND_REWAKE_V';
* changes: `refreshStaleCodemanHooks` treats the absence of the CURRENT marker as
* stale, so healed cases pick up the new script on next launch.
*/
const BACKGROUND_WAKE_MARKER = `${BACKGROUND_WAKE_MARKER_PREFIX}2`;
const BACKGROUND_WAKE_MARKER = `${BACKGROUND_WAKE_MARKER_PREFIX}3`;
const SUBAGENT_STOP_GUARD_MARKER_PREFIX = 'CODEMAN_SUBAGENT_STOP_GUARD_V';
const SUBAGENT_STOP_GUARD_MARKER = `${SUBAGENT_STOP_GUARD_MARKER_PREFIX}1`;
const BACKGROUND_WAKE_TIMEOUT_SECONDS = 6 * 60 * 60;
/**
* Inline Node helper for Claude Code's `asyncRewake` hook.
*
* A background Bash tool returns immediately with a task ID, then Claude writes
* its completion as a queue-operation in the transcript. Watching that durable
* record avoids injecting terminal input (which could submit a user's draft).
* its completion as a queue-operation in the top-level transcript. Subagent hooks
* receive their own transcript path even though their completion is parent-owned,
* so the helper watches both paths. Watching durable records avoids injecting
* terminal input (which could submit a user's draft).
* The helper is embedded in settings via `node -e`, so it has no script path
* that can go stale after an install or plugin-cache cleanup.
*
@@ -72,8 +79,12 @@ const BACKGROUND_WAKE_TIMEOUT_SECONDS = 6 * 60 * 60;
export function generateBackgroundWakeScript(): string {
return [
"const fs = require('node:fs');",
"const path = require('node:path');",
`const ${BACKGROUND_WAKE_MARKER} = true;`,
`const deadline = Date.now() + ${BACKGROUND_WAKE_TIMEOUT_SECONDS} * 1000;`,
"const RESULT_BEGIN = '=== CODEMAN_RESULT_BEGIN ===';",
"const RESULT_END = '=== CODEMAN_RESULT_END ===';",
'const MAX_RESULT_CHARS = 65536;',
'let input = {};',
"try { input = JSON.parse(fs.readFileSync(0, 'utf8') || '{}'); } catch { process.exit(0); }",
'function findTaskId(value) {',
@@ -98,46 +109,164 @@ export function generateBackgroundWakeScript(): string {
'const taskId = findTaskId(input.tool_response);',
"const transcriptPath = typeof input.transcript_path === 'string' ? input.transcript_path : '';",
'if (!taskId || !transcriptPath) process.exit(0);',
'let position = 0;',
'try { position = Math.max(0, fs.statSync(transcriptPath).size - 262144); } catch { process.exit(0); }',
"let carry = '';",
'const transcriptPaths = [transcriptPath];',
'const sessionDir = path.dirname(path.dirname(transcriptPath));',
"if (typeof input.agent_id === 'string' && path.basename(path.dirname(transcriptPath)) === 'subagents' &&",
" typeof input.session_id === 'string' && path.basename(sessionDir) === input.session_id) {",
" transcriptPaths.push(sessionDir + '.jsonl');",
'}',
'const transcripts = [...new Set(transcriptPaths)].map((transcript) => {',
' let position = 0;',
' try { position = Math.max(0, fs.statSync(transcript).size - 262144); } catch {}',
" return { path: transcript, position, carry: '' };",
'});',
'if (!transcripts.some((transcript) => fs.existsSync(transcript.path))) process.exit(0);',
'function readMarkedResult(outputPath) {',
" if (!outputPath || !path.isAbsolute(outputPath) || path.basename(outputPath) !== taskId + '.output') return '';",
" if (path.basename(path.dirname(outputPath)) !== 'tasks') return '';",
' try {',
' const size = fs.statSync(outputPath).size;',
' const length = Math.min(size, MAX_RESULT_CHARS * 2);',
' const buffer = Buffer.allocUnsafe(length);',
" const fd = fs.openSync(outputPath, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, size - length);',
' fs.closeSync(fd);',
" const text = buffer.subarray(0, bytes).toString('utf8');",
' const begin = text.lastIndexOf(RESULT_BEGIN);',
' const end = text.indexOf(RESULT_END, begin + RESULT_BEGIN.length);',
" if (begin < 0 || end < 0) return '';",
' let result = text.slice(begin + RESULT_BEGIN.length, end).trim();',
" if (!result) return '';",
' if (result.length > MAX_RESULT_CHARS) {',
' const half = Math.floor(MAX_RESULT_CHARS / 2);',
" result = result.slice(0, half) + '\\n\\n[report truncated by Codeman]\\n\\n' + result.slice(-half);",
' }',
" return '\\n\\nCompleted task report:\\n<codeman-background-result>\\n' + result + '\\n</codeman-background-result>';",
" } catch { return ''; }",
'}',
'function inspect(text) {',
' for (const line of text.split(/\\r?\\n/)) {',
' if (!line.includes(taskId)) continue;',
' let entry;',
' try { entry = JSON.parse(line); } catch { continue; }',
" if (entry.type !== 'queue-operation' || typeof entry.content !== 'string') continue;",
" if (entry.type !== 'queue-operation' || entry.operation !== 'enqueue' || typeof entry.content !== 'string') continue;",
" if (!entry.content.includes('<task-id>' + taskId + '</task-id>')) continue;",
' const status = entry.content.match(/<status>(completed|failed|killed|error)<\\/status>/i);',
' if (!status) continue;',
' const output = entry.content.match(/<output-file>([^<]+)<\\/output-file>/i);',
" const location = output ? ' Read ' + output[1] + ' and' : '';",
" console.error('Background command ' + taskId + ' ' + status[1].toLowerCase() + '.' + location + ' continue the task.');",
" const outputPath = output ? output[1].trim() : '';",
" const location = outputPath ? ' Read ' + outputPath + ' and' : '';",
' const result = readMarkedResult(outputPath);',
" console.error('Background command ' + taskId + ' ' + status[1].toLowerCase() + '.' + location + ' continue the task.' + result);",
' process.exit(2);',
' }',
'}',
'function poll() {',
' if (Date.now() > deadline || process.ppid === 1) process.exit(0);',
'function pollTranscript(transcript) {',
' try {',
' const size = fs.statSync(transcriptPath).size;',
" if (size < position) { position = 0; carry = ''; }",
' if (size > position) {',
' const length = Math.min(size - position, 1048576);',
' const size = fs.statSync(transcript.path).size;',
" if (size < transcript.position) { transcript.position = 0; transcript.carry = ''; }",
' if (size > transcript.position) {',
' const length = Math.min(size - transcript.position, 1048576);',
' const buffer = Buffer.allocUnsafe(length);',
" const fd = fs.openSync(transcriptPath, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, position);',
" const fd = fs.openSync(transcript.path, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, transcript.position);',
' fs.closeSync(fd);',
' position += bytes;',
" carry = (carry + buffer.subarray(0, bytes).toString('utf8')).slice(-262144);",
' inspect(carry);',
' transcript.position += bytes;',
" transcript.carry = (transcript.carry + buffer.subarray(0, bytes).toString('utf8')).slice(-262144);",
' inspect(transcript.carry);',
' }',
' } catch {}',
'}',
'function poll() {',
' if (Date.now() > deadline || process.ppid === 1) process.exit(0);',
' for (const transcript of transcripts) pollTranscript(transcript);',
' setTimeout(poll, 1000);',
'}',
'poll();',
].join('\n');
}
/**
* Keep a Claude subagent alive while its Monitor or background Bash work is live.
* Claude otherwise can publish the worker's last progress sentence as an Agent
* result when one watcher ends, even if other tracked tasks are still running.
*/
export function generateSubagentStopGuardScript(): string {
return [
"const fs = require('node:fs');",
`const ${SUBAGENT_STOP_GUARD_MARKER} = true;`,
'let input = {};',
"try { input = JSON.parse(fs.readFileSync(0, 'utf8') || '{}'); } catch { process.exit(0); }",
"const transcriptPath = typeof input.agent_transcript_path === 'string' ? input.agent_transcript_path : '';",
'if (!transcriptPath) process.exit(0);',
'let text;',
'try {',
' const size = fs.statSync(transcriptPath).size;',
' const length = Math.min(size, 16 * 1024 * 1024);',
' const buffer = Buffer.allocUnsafe(length);',
" const fd = fs.openSync(transcriptPath, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, size - length);',
' fs.closeSync(fd);',
" text = buffer.subarray(0, bytes).toString('utf8');",
'} catch { process.exit(0); }',
'const launched = new Set();',
'const finished = new Set();',
'function inspectToolResult(value) {',
" const serialized = typeof value === 'string' ? value : JSON.stringify(value ?? '');",
' for (const match of serialized.matchAll(/Command running in background with ID:\\s*([A-Za-z0-9_-]+)/gi)) launched.add(match[1]);',
' for (const match of serialized.matchAll(/Monitor started \\(task ([A-Za-z0-9_-]+)/gi)) launched.add(match[1]);',
'}',
'function inspectNotifications(value) {',
" if (typeof value !== 'string' || !value.includes('<task-notification>')) return;",
' for (const match of value.matchAll(/<task-notification>([\\s\\S]*?)<\\/task-notification>/gi)) {',
' const body = match[1];',
' const id = body.match(/<task-id>([^<]+)<\\/task-id>/i);',
' const status = body.match(/<status>(completed|failed|killed|error)<\\/status>/i);',
' if (id && status) finished.add(id[1].trim());',
' }',
'}',
'for (const line of text.split(/\\r?\\n/)) {',
' let entry;',
' try { entry = JSON.parse(line); } catch { continue; }',
' const content = entry && entry.message ? entry.message.content : undefined;',
' if (Array.isArray(content)) {',
' for (const block of content) {',
" if (block && block.type === 'tool_result') inspectToolResult(block.content);",
" if (block && block.type === 'text') inspectNotifications(block.text);",
' }',
' } else {',
' inspectNotifications(content);',
' }',
' inspectNotifications(entry && entry.content);',
'}',
'function findLiveTasks(candidates) {',
' const live = new Set();',
" if (candidates.size === 0 || !fs.existsSync('/proc')) return live;",
' let processIds;',
" try { processIds = fs.readdirSync('/proc').filter((name) => /^\\d+$/.test(name)); } catch { return live; }",
' for (const processId of processIds) {',
" for (const descriptor of ['0', '1', '2']) {",
' let target;',
" try { target = fs.readlinkSync('/proc/' + processId + '/fd/' + descriptor); } catch { continue; }",
' const match = target.match(/[\\/]tasks[\\/]([A-Za-z0-9_-]+)\\.output(?: \\(deleted\\))?$/);',
' if (match && candidates.has(match[1])) live.add(match[1]);',
' }',
' if (live.size === candidates.size) break;',
' }',
' return live;',
'}',
'const unfinished = new Set([...launched].filter((taskId) => !finished.has(taskId)));',
'const active = [...findLiveTasks(unfinished)];',
'if (active.length === 0) process.exit(0);',
'const shown = active.slice(0, 8);',
"const suffix = active.length > shown.length ? ' and ' + (active.length - shown.length) + ' more' : '';",
'process.stdout.write(JSON.stringify({',
" decision: 'block',",
" reason: 'You still own active background work (' + shown.join(', ') + suffix + '). Do not return an intermediate progress message as your final report. Process the task notifications or keep actively polling until every task completes, then return one complete summary.',",
'}));',
].join('\n');
}
function withSettingsLock<T>(path: string, fn: () => Promise<T>): Promise<T> {
const prev = settingsWriteLocks.get(path) ?? Promise.resolve();
const run = prev.then(fn, fn); // run after the prior writer, regardless of its outcome
@@ -173,7 +302,11 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
const curlCmd = (event: HookEventType) =>
`HOOK_DATA=$(cat 2>/dev/null || echo '{}'); ` +
`printf '{"event":"${event}","sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$HOOK_DATA" | ` +
`curl -s -X POST "$CODEMAN_API_URL/api/hook-event" ` +
// `-k`, same as the statusline exporter: CODEMAN_API_URL is loopback HTTPS with
// a self-signed cert on --https/tailscale installs. Without it curl exits 60,
// the `|| true` swallows it, and ALL SIX hook events die silently: respawn loses
// its definitive idle signals and the wait endpoints lose stop/blocked.
`curl -sk -X POST "$CODEMAN_API_URL/api/hook-event" ` +
`-H 'Content-Type: application/json' ` +
`-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" ` +
`--data @- ` +
@@ -200,6 +333,18 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
hooks: [{ type: 'command', command: curlCmd('stop'), timeout: HOOK_TIMEOUT_SECONDS }],
},
],
SubagentStop: [
{
hooks: [
{
type: 'command',
command: 'node',
args: ['-e', generateSubagentStopGuardScript()],
timeout: HOOK_TIMEOUT_SECONDS,
},
],
},
],
TeammateIdle: [
{
hooks: [{ type: 'command', command: curlCmd('teammate_idle'), timeout: HOOK_TIMEOUT_SECONDS }],
@@ -231,8 +376,12 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
function isCodemanHookHandler(value: unknown): boolean {
try {
const serialized = JSON.stringify(value);
// Prefix, not the versioned marker: older script versions must still be ours.
return serialized.includes('/api/hook-event') || serialized.includes(BACKGROUND_WAKE_MARKER_PREFIX);
// Prefixes, not versioned markers: older script versions must still be ours.
return (
serialized.includes('/api/hook-event') ||
serialized.includes(BACKGROUND_WAKE_MARKER_PREFIX) ||
serialized.includes(SUBAGENT_STOP_GUARD_MARKER_PREFIX)
);
} catch {
return false;
}
@@ -426,6 +575,39 @@ export async function writeHooksConfig(casePath: string): Promise<void> {
});
}
/**
* Ensures an explicitly managed case has the current Codeman hooks.
*
* Unlike `refreshStaleCodemanHooks`, this may add Codeman handlers to a valid
* user-owned settings file. It is therefore reserved for case quick-starts,
* where the user has explicitly asked Codeman to manage that workspace. A
* malformed existing file is left untouched rather than replaced.
*/
export async function ensureCodemanHooks(casePath: string): Promise<void> {
const claudeDir = join(casePath, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
await withSettingsLock(settingsPath, async () => {
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
let existing: Record<string, unknown> = {};
try {
const parsed: unknown = JSON.parse(await readFile(settingsPath, 'utf-8'));
if (!parsed || typeof parsed !== 'object' || Array.isArray(parsed)) return;
existing = parsed as Record<string, unknown>;
} catch (err) {
if ((err as NodeJS.ErrnoException).code !== 'ENOENT') return;
}
const generated = generateHooksConfig();
const hooks = mergeCodemanHooks(existing.hooks, generated.hooks);
if (JSON.stringify(existing.hooks ?? {}) === JSON.stringify(hooks)) return;
await writeFile(settingsPath, JSON.stringify({ ...existing, hooks }, null, 2) + '\n');
});
}
/**
* Self-heal a case's Codeman-owned hooks block.
*
@@ -433,8 +615,11 @@ export async function writeHooksConfig(casePath: string): Promise<void> {
* X-Codeman-Hook-Secret header was added (COD-54, 2026-06-10) keep hook curls in their
* settings.local.json that POST to /api/hook-event WITHOUT the secret — which, once the
* gate requires it unconditionally (COD-91), silently 401 on a password-protected install.
* Older Codeman blocks also lack the background Bash async-rewake hook. Refresh either
* stale shape on launch so existing cases gain both current behaviors.
* Older Codeman blocks also lack the current background Bash async-rewake hook or the
* SubagentStop guard. A further stale shape: hook curls without `-k`, which exit 60 on
* every --https/tailscale install (the cert is self-signed), swallowed by the hooks'
* own `|| true` — all six hook events die silently. Refresh any of these stale shapes
* on launch so existing cases heal.
*
* Deliberately surgical: regenerates ONLY when settings.local.json already contains
* Codeman's own hook curls (they target `/api/hook-event`) and they are stale. No-op
@@ -457,7 +642,12 @@ export async function refreshStaleCodemanHooks(casePath: string): Promise<void>
// absence on our own hooks means they predate COD-54 and need regenerating.
const hasSecret = hooksJson.includes('X-Codeman-Hook-Secret');
const hasBackgroundWake = hooksJson.includes(BACKGROUND_WAKE_MARKER);
if (!isOurs || (hasSecret && hasBackgroundWake)) return;
// The pre--k curl shape: `curl -sk -X POST` does not contain `curl -s -X POST`
// as a substring, so this cleanly identifies hook curls that die with exit 60
// on a self-signed HTTPS install.
const hasTlsFlaglessCurl = hooksJson.includes('curl -s -X POST');
const hasSubagentStopGuard = hooksJson.includes(SUBAGENT_STOP_GUARD_MARKER);
if (!isOurs || (hasSecret && hasBackgroundWake && hasSubagentStopGuard && !hasTlsFlaglessCurl)) return;
const generated = generateHooksConfig();
const merged = {
...existing,
@@ -531,3 +721,171 @@ export async function applyStatusLineConfig(casePath: string, enabled: boolean):
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
});
}
// ─── Agent skill injection ───────────────────────────────────────────────────
/**
* Version-agnostic ownership prefix for the injected agent skill, same pattern as
* `BACKGROUND_WAKE_MARKER_PREFIX`: ownership is decided on the prefix so a wording
* change in the full marker cannot disown every previously injected copy.
*/
const AGENT_SKILL_MARKER_PREFIX = '<!-- codeman-managed-agent-skill';
/**
* Marker appended to the injected SKILL.md. Its presence is what makes a copy OURS:
* install/refresh/remove all refuse to touch a `skills/codeman` whose SKILL.md lacks
* it, so a user's hand-authored or hand-edited-and-de-marked skill is never clobbered.
*/
const AGENT_SKILL_MARKER = `${AGENT_SKILL_MARKER_PREFIX}: installed by Codeman; edits are overwritten while the agent-skill setting is on -->`;
/**
* Packaged source of the skill: `skills/codeman/` at the package root. Resolved
* relative to this module so it works from `src/` (tsx dev), `dist/` (tsc build),
* and an npm install (`files` includes `skills`), all of which sit one level below
* the package root.
*/
function agentSkillSourceDir(): string {
return join(dirname(fileURLToPath(import.meta.url)), '..', 'skills', 'codeman');
}
interface AgentSkillFile {
/** Path relative to the target skill dir (e.g. `reference/endpoints.md`). */
relPath: string;
content: string;
}
/**
* Read the packaged skill: SKILL.md (marker appended) plus every markdown file
* under `reference/`. Enumerated from disk rather than a hardcoded manifest so a
* new reference file ships without touching this module.
*/
async function readAgentSkillSource(): Promise<AgentSkillFile[]> {
const src = agentSkillSourceDir();
const skill = await readFile(join(src, 'SKILL.md'), 'utf-8');
const files: AgentSkillFile[] = [{ relPath: 'SKILL.md', content: `${skill.trimEnd()}\n\n${AGENT_SKILL_MARKER}\n` }];
let referenceNames: string[] = [];
try {
referenceNames = (await readdir(join(src, 'reference'))).filter((name) => name.endsWith('.md')).sort();
} catch {
// no reference dir in the source; SKILL.md alone is still a valid skill
}
for (const name of referenceNames) {
files.push({ relPath: join('reference', name), content: await readFile(join(src, 'reference', name), 'utf-8') });
}
return files;
}
async function isSymlink(path: string): Promise<boolean> {
try {
return (await lstat(path)).isSymbolicLink();
} catch {
return false;
}
}
/** What an install/remove actually did, so callers (CLI, logs) can say so. */
export type AgentSkillApplyResult =
| 'installed' // fresh copy written
| 'refreshed' // our copy was stale and got rewritten
| 'unchanged' // our copy already matches the packaged source
| 'removed' // our copy deleted
| 'absent' // nothing there to remove
| 'foreign' // a copy exists but is not ours; left untouched
| 'symlink'; // the skill dir (or its parent) is a symlink; left untouched
/**
* Install or refresh the Codeman agent skill into `skillDir` (a `.../codeman`
* directory, e.g. `<case>/.claude/skills/codeman` or `~/.claude/skills/codeman`).
*
* Refuses two shapes rather than writing through them:
* - a SYMLINK at the skill dir or its `skills/` parent: this repo's own dogfooding
* layout (`.claude/skills/codeman -> ../../skills/codeman`) would otherwise have
* the injector overwrite the repo source through the link;
* - a FOREIGN copy (SKILL.md present without our marker): that is the user's own
* skill, and per the statusLine rule we never clobber what we did not write.
*
* Idempotent and cheap: unchanged files are not rewritten, so calling on every
* session create causes no mtime churn.
*/
export async function installAgentSkillInto(skillDir: string): Promise<AgentSkillApplyResult> {
if ((await isSymlink(dirname(skillDir))) || (await isSymlink(skillDir))) return 'symlink';
let existing: string | null = null;
try {
existing = await readFile(join(skillDir, 'SKILL.md'), 'utf-8');
} catch {
// absent: fresh install
}
if (existing !== null && !existing.includes(AGENT_SKILL_MARKER_PREFIX)) return 'foreign';
const files = await readAgentSkillSource();
let changed = false;
for (const file of files) {
const target = join(skillDir, file.relPath);
let current: string | null = null;
try {
current = await readFile(target, 'utf-8');
} catch {
// missing: will be written
}
if (current === file.content) continue;
await mkdir(dirname(target), { recursive: true });
await writeFile(target, file.content);
changed = true;
}
if (!changed) return 'unchanged';
return existing === null ? 'installed' : 'refreshed';
}
/**
* Remove a Codeman-managed skill copy from `skillDir`. Same ownership and symlink
* refusals as the install path. Deletes only files the packaged source would have
* written (never `rm -rf`, so a user's extra files in the directory survive), then
* prunes the directories bottom-up if they emptied.
*/
export async function removeAgentSkillFrom(skillDir: string): Promise<AgentSkillApplyResult> {
if ((await isSymlink(dirname(skillDir))) || (await isSymlink(skillDir))) return 'symlink';
let existing: string | null = null;
try {
existing = await readFile(join(skillDir, 'SKILL.md'), 'utf-8');
} catch {
return 'absent';
}
if (!existing.includes(AGENT_SKILL_MARKER_PREFIX)) return 'foreign';
// Manifest-based, with SKILL.md as the fallback when the packaged source is
// unreadable: removal must still work on an install whose skills/ dir went missing.
const files = await readAgentSkillSource().catch((): AgentSkillFile[] => [{ relPath: 'SKILL.md', content: '' }]);
for (const file of files) {
await unlink(join(skillDir, file.relPath)).catch(() => {});
}
await rmdir(join(skillDir, 'reference')).catch(() => {}); // fails when non-empty, fine
await rmdir(skillDir).catch(() => {});
await rmdir(dirname(skillDir)).catch(() => {}); // prune `.claude/skills` if now empty
return 'removed';
}
/**
* Add or remove the Codeman agent skill in `<case>/.claude/skills/codeman`,
* mirroring `applyStatusLineConfig`'s shape. Gated by the synced `agentSkillEnabled`
* app setting (default OFF); callers gate on Claude mode, since the skill is discovered
* via `.claude/skills/`, which only Claude Code reads.
*
* Call-site policy is ADD-ONLY on session create (callers pass `enabled: true` or
* skip the call), for the statusLine reason: sessions in a repo share one `.claude/`
* dir, so a single create while the setting is off must not yank the skill out from
* under other live sessions.
*
* ⚠️ Consequence: turning `agentSkillEnabled` OFF sweeps nothing. There is deliberately
* no server-side toggle-off sweep (it would have to walk every case, including ones
* with live sessions, and would hit exactly the shared-`.claude/` hazard above), so
* already-injected copies stay on disk until removed per case with
* `codeman skill uninstall --case <name>`. The `enabled: false` branch here backs that
* CLI and the tests; it has no server call site. Keep the README's Agent Skill note in
* sync if this ever changes.
*/
export async function applyAgentSkill(casePath: string, enabled: boolean): Promise<AgentSkillApplyResult> {
const skillDir = join(casePath, '.claude', 'skills', 'codeman');
return enabled ? installAgentSkillInto(skillDir) : removeAgentSkillFrom(skillDir);
}
+87
View File
@@ -0,0 +1,87 @@
/**
* @fileoverview Bounded descendant walk over a process-tree snapshot.
*
* Split out of `tmux-manager.ts` so the traversal can be unit-tested directly. It
* previously lived as a private method, which meant the regression test had to keep
* its own copy of the algorithm — a test that passes while the shipped code rots.
*
* ## The incident this guards against
*
* On 2026-07-30 an unbounded version of this walk took a machine down. It ran
* `pgrep -P <pid>` once per node and recursed with no visited set, no depth limit and
* no node cap. Across ~28 adopted tmux trees the fan-out exploded, and because each
* `pgrep` blocks in the WSL kernel while reading `/proc/<pid>/cgroup`, none of them
* returned while the walk kept spawning more. Result: ~13,000 `pgrep` processes stuck
* in D-state out of ~39,000 total, load average above 13,000, and a machine only
* recoverable by restarting WSL — which cost every running session.
*
* Three properties make that impossible, and each has a test:
* 1. a cycle terminates instead of looping (stale snapshots can contain one),
* 2. depth is capped,
* 3. node count is capped.
*
* The fourth property — spawning nothing per node — is structural: this function
* takes a snapshot and cannot spawn anything at all.
*
* @module proc-tree
*/
/** Maximum generations to descend. Deeper than any real agent process tree. */
export const PROC_WALK_MAX_DEPTH = 10;
/** Hard ceiling on collected descendants. A backstop, not an expected limit. */
export const PROC_WALK_MAX_NODES = 500;
export interface WalkOptions {
maxDepth?: number;
maxNodes?: number;
/**
* Called once when a cap truncated the result, with which cap it was. Both are
* reported: a silent depth cap would hide a deep tree just as effectively as a
* silent node cap hides a wide one, and the whole point of this module is that
* truncation is visible rather than mysterious.
*/
onTruncated?: (pid: number, cap: number, reason: 'nodes' | 'depth') => void;
}
/**
* All descendants of `pid`, breadth-first and bounded.
*
* @param pid root of the walk; never included in the result
* @param byParent parent pid → child pids, from ONE `ps` snapshot
*/
export function collectDescendants(
pid: number,
byParent: ReadonlyMap<number, readonly number[]>,
opts: WalkOptions = {}
): number[] {
const maxDepth = opts.maxDepth ?? PROC_WALK_MAX_DEPTH;
const maxNodes = opts.maxNodes ?? PROC_WALK_MAX_NODES;
const out: number[] = [];
const visited = new Set<number>([pid]);
let frontier = [pid];
for (let depth = 0; depth < maxDepth && frontier.length; depth += 1) {
const next: number[] = [];
for (const parent of frontier) {
for (const child of byParent.get(parent) ?? []) {
if (visited.has(child)) continue; // a real tree has no cycles, a stale
visited.add(child); // snapshot can still produce one
out.push(child);
next.push(child);
if (out.length >= maxNodes) {
opts.onTruncated?.(pid, maxNodes, 'nodes');
return out;
}
}
}
frontier = next;
// Ran out of generations while descendants were still queued: the tree is
// deeper than the cap and the result is incomplete.
if (depth === maxDepth - 1 && frontier.length > 0) {
opts.onTruncated?.(pid, maxDepth, 'depth');
}
}
return out;
}
+401
View File
@@ -0,0 +1,401 @@
/**
* @fileoverview `codeman service install|uninstall|status`: write and load the
* systemd user unit (Linux) or LaunchAgent (macOS) that supervises `codeman web`.
*
* This is the "always running" half of issue #231, next to the "detached right
* now" half in daemon-control.ts. `install.sh` already does this for people who
* install with the one-liner; this exists for `npm i -g aicodeman` users, who
* otherwise have to hand-write a plist.
*
* Two details are load-bearing and easy to get wrong by hand:
*
* - **PATH.** launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` and systemd's
* user manager is nearly as bare, so a Homebrew or nvm `node`, `tmux` or
* `claude` is simply not found and sessions fail in a way that reads as a
* Codeman bug. The unit therefore carries the PATH of the shell that ran the
* install, with the running node's own directory in front.
* - **The job name.** It is the one `install.sh` and the self-updater already use
* (config/service-names.ts), so re-running install.sh later updates this unit
* instead of supervising a second copy of the server.
*
* Secrets are deliberately NOT written here. `CODEMAN_PASSWORD` in the installing
* shell is not copied into the unit; the caller is told where to add it instead,
* because a unit file is long-lived, world-readable by default, and gets copied
* into bug reports.
*
* The file writers are pure string builders so they can be unit-tested without
* touching launchctl/systemctl.
*
* @module service-installer
*/
import { execFileSync } from 'node:child_process';
import { existsSync, mkdirSync, unlinkSync, writeFileSync } from 'node:fs';
import { homedir, userInfo } from 'node:os';
import { dirname, join } from 'node:path';
import { LAUNCHD_LABEL, SYSTEMD_UNIT } from './config/service-names.js';
import { CODEMAN_INSTANCE } from './config/instance.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import {
buildBaseUrl,
buildStatusUrl,
buildWebArgs,
logFilePath,
probeServer,
type WebLaunchOptions,
} from './daemon-control.js';
export type ServiceKind = 'launchd' | 'systemd';
/** Everything a unit file needs, resolved from the environment by the caller. */
export interface ServicePlan {
kind: ServiceKind;
/** systemd unit filename or launchd label. */
name: string;
nodePath: string;
/** Runner flags carried over from the current process (tsx loader in dev). */
execArgv: string[];
scriptPath: string;
args: string[];
env: Record<string, string>;
logPath: string;
workingDir: string;
}
export interface ServiceActionResult {
ok: boolean;
message: string;
/** Path of the unit/plist that was written or removed. */
unitPath?: string;
warnings?: string[];
}
export interface ServiceStatusResult {
kind: ServiceKind | null;
name: string;
unitPath: string;
installed: boolean;
loaded: boolean;
responding: boolean;
version?: string;
url: string;
}
/** Directories worth having on PATH even when the installing shell lacked them. */
const FALLBACK_PATH_DIRS = ['/opt/homebrew/bin', '/usr/local/bin', '/usr/bin', '/bin', '/usr/sbin', '/sbin'];
// ─────────────────────────────────────────────────────────────────────────────
// Pure builders
// ─────────────────────────────────────────────────────────────────────────────
/** XML text escaping for plist `<string>` values. */
export function xmlEscape(value: string): string {
return value
.replace(/&/g, '&amp;')
.replace(/</g, '&lt;')
.replace(/>/g, '&gt;')
.replace(/"/g, '&quot;')
.replace(/'/g, '&apos;');
}
/**
* PATH for the supervised process: the running node's directory first (so an nvm
* or Homebrew node is used rather than whatever the supervisor finds), then the
* installing shell's PATH, then the fallbacks that are still missing.
*
* `node_modules/.bin` entries are dropped. npm and npx inject those for the
* lifetime of one command, and baking a project's local bin dir into a unit file
* that outlives the checkout is how a service ends up running a binary the
* operator deleted months ago.
*/
export function buildServicePath(nodeDir: string, currentPath: string, home: string): string {
const seen = new Set<string>();
const ordered: string[] = [];
const push = (dir: string) => {
const trimmed = dir.trim();
if (!trimmed || seen.has(trimmed)) return;
if (/(^|\/)node_modules\/\.bin\/?$/.test(trimmed)) return;
seen.add(trimmed);
ordered.push(trimmed);
};
push(nodeDir);
for (const dir of currentPath.split(':')) push(dir);
push(join(home, '.local', 'bin'));
for (const dir of FALLBACK_PATH_DIRS) push(dir);
return ordered.join(':');
}
/** Environment written into the unit. Never includes secrets (see module docs). */
export function buildServiceEnv(
nodeDir: string,
currentPath: string,
home: string,
lang?: string
): Record<string, string> {
const env: Record<string, string> = {
PATH: buildServicePath(nodeDir, currentPath, home),
HOME: home,
LANG: lang || 'en_US.UTF-8',
};
if (CODEMAN_INSTANCE) env.CODEMAN_INSTANCE = CODEMAN_INSTANCE;
return env;
}
/** systemd accepts double-quoted values; escape the two characters that matter. */
export function systemdQuote(value: string): string {
return `"${value.replace(/\\/g, '\\\\').replace(/"/g, '\\"')}"`;
}
export function buildLaunchAgentPlist(plan: ServicePlan): string {
const programArguments = [plan.nodePath, ...plan.execArgv, plan.scriptPath, ...plan.args]
.map((arg) => ` <string>${xmlEscape(arg)}</string>`)
.join('\n');
const environment = Object.entries(plan.env)
.map(([key, value]) => ` <key>${xmlEscape(key)}</key>\n <string>${xmlEscape(value)}</string>`)
.join('\n');
return `<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>${xmlEscape(plan.name)}</string>
<key>ProgramArguments</key>
<array>
${programArguments}
</array>
<key>EnvironmentVariables</key>
<dict>
${environment}
</dict>
<key>WorkingDirectory</key>
<string>${xmlEscape(plan.workingDir)}</string>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
<key>ThrottleInterval</key>
<integer>10</integer>
<key>StandardOutPath</key>
<string>${xmlEscape(plan.logPath)}</string>
<key>StandardErrorPath</key>
<string>${xmlEscape(plan.logPath)}</string>
</dict>
</plist>
`;
}
export function buildSystemdUnit(plan: ServicePlan): string {
const execStart = [plan.nodePath, ...plan.execArgv, plan.scriptPath, ...plan.args]
.map((arg) => (/[\s"'\\]/.test(arg) ? systemdQuote(arg) : arg))
.join(' ');
const environment = Object.entries(plan.env)
.map(([key, value]) => `Environment=${systemdQuote(`${key}=${value}`)}`)
.join('\n');
return `[Unit]
Description=Codeman Web Server
After=network.target
[Service]
Type=simple
WorkingDirectory=${plan.workingDir}
ExecStart=${execStart}
Restart=always
RestartSec=10
# Agents keep running in tmux when the server restarts, so only signal the
# server itself.
KillMode=process
${environment}
StandardOutput=journal
StandardError=journal
SyslogIdentifier=codeman
LimitNOFILE=65536
[Install]
WantedBy=default.target
`;
}
// ─────────────────────────────────────────────────────────────────────────────
// Environment resolution
// ─────────────────────────────────────────────────────────────────────────────
export function detectServiceKind(): ServiceKind | null {
if (process.platform === 'darwin') return 'launchd';
if (process.platform === 'linux') return 'systemd';
return null;
}
export function unitPathFor(kind: ServiceKind): string {
return kind === 'launchd'
? join(homedir(), 'Library', 'LaunchAgents', `${LAUNCHD_LABEL}.plist`)
: join(homedir(), '.config', 'systemd', 'user', SYSTEMD_UNIT);
}
function entryScript(): string {
const script = process.argv[1];
if (!script) throw new Error('cannot determine the codeman entry script to supervise');
return script;
}
/** Resolve a full plan from the current process and the requested web options. */
export function resolveServicePlan(kind: ServiceKind, options: WebLaunchOptions): ServicePlan {
const home = homedir();
return {
kind,
name: kind === 'launchd' ? LAUNCHD_LABEL : SYSTEMD_UNIT,
nodePath: process.execPath,
execArgv: [...process.execArgv],
scriptPath: entryScript(),
args: buildWebArgs(options),
env: buildServiceEnv(dirname(process.execPath), process.env.PATH || '', home, process.env.LANG),
logPath: logFilePath(),
workingDir: home,
};
}
function run(command: string, args: string[]): { ok: boolean; output: string } {
try {
const output = execFileSync(command, args, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'pipe'],
});
return { ok: true, output: output.trim() };
} catch (err) {
const e = err as { stderr?: Buffer | string; message?: string };
const stderr = typeof e.stderr === 'string' ? e.stderr : e.stderr?.toString('utf-8');
return { ok: false, output: (stderr || e.message || '').trim() };
}
}
// ─────────────────────────────────────────────────────────────────────────────
// Install / uninstall / status
// ─────────────────────────────────────────────────────────────────────────────
/**
* Write the unit, load it, and confirm the server actually answers before
* reporting success. `launchctl load` and `systemctl enable` are both quiet about
* a job that starts and immediately dies, which is the whole reason install.sh
* verifies too.
*/
export async function installService(options: WebLaunchOptions): Promise<ServiceActionResult> {
const kind = detectServiceKind();
if (!kind) {
return { ok: false, message: `no supported supervisor on ${process.platform}; use \`codeman web -d\` instead` };
}
const plan = resolveServicePlan(kind, options);
const unitPath = unitPathFor(kind);
const warnings: string[] = [];
mkdirSync(dirname(unitPath), { recursive: true });
if (kind === 'launchd') {
const uid = process.getuid?.() ?? 0;
// Unload any previous copy first, otherwise bootstrap fails with "service
// already loaded" and leaves the OLD job running against the NEW file.
run('launchctl', ['bootout', `gui/${uid}/${LAUNCHD_LABEL}`]);
writeFileSync(unitPath, buildLaunchAgentPlist(plan), { encoding: 'utf-8', mode: 0o600 });
const bootstrap = run('launchctl', ['bootstrap', `gui/${uid}`, unitPath]);
if (!bootstrap.ok) {
const legacy = run('launchctl', ['load', unitPath]);
if (!legacy.ok) {
return {
ok: false,
unitPath,
message: `wrote ${unitPath} but launchctl refused to load it: ${bootstrap.output}`,
};
}
}
} else {
writeFileSync(unitPath, buildSystemdUnit(plan), { encoding: 'utf-8', mode: 0o600 });
const reload = run('systemctl', ['--user', 'daemon-reload']);
if (!reload.ok) {
return {
ok: false,
unitPath,
message: `wrote ${unitPath} but \`systemctl --user daemon-reload\` failed: ${reload.output}`,
};
}
const enable = run('systemctl', ['--user', 'enable', '--now', SYSTEMD_UNIT]);
if (!enable.ok) {
return { ok: false, unitPath, message: `wrote ${unitPath} but enabling it failed: ${enable.output}` };
}
// Without lingering the unit stops at logout, which is exactly what someone
// installing a service does not want. Best effort: it needs polkit rights.
const linger = run('loginctl', ['enable-linger', userInfo().username]);
if (!linger.ok) {
warnings.push(
`could not enable lingering, so the service will stop when you log out. Run: sudo loginctl enable-linger ${userInfo().username}`
);
}
}
const url = buildBaseUrl(options);
const statusUrl = buildStatusUrl(options);
const deadline = Date.now() + 30_000;
while (Date.now() < deadline) {
const probe = await probeServer(statusUrl, 1000);
if (probe.up) {
return { ok: true, unitPath, warnings, message: `service installed and responding at ${url}` };
}
await new Promise((resolve) => setTimeout(resolve, 500));
}
const hint =
kind === 'launchd' ? `tail -20 ${plan.logPath}` : `journalctl --user -u ${SYSTEMD_UNIT} -n 20 --no-pager`;
return {
ok: false,
unitPath,
warnings,
message: `wrote and loaded ${unitPath}, but nothing answered ${url} within 30s. Check: ${hint}`,
};
}
export function uninstallService(): ServiceActionResult {
const kind = detectServiceKind();
if (!kind) return { ok: false, message: `no supported supervisor on ${process.platform}` };
const unitPath = unitPathFor(kind);
if (!existsSync(unitPath)) {
return { ok: false, unitPath, message: `no service installed at ${unitPath}` };
}
if (kind === 'launchd') {
const uid = process.getuid?.() ?? 0;
const bootout = run('launchctl', ['bootout', `gui/${uid}/${LAUNCHD_LABEL}`]);
if (!bootout.ok) run('launchctl', ['unload', unitPath]);
} else {
run('systemctl', ['--user', 'disable', '--now', SYSTEMD_UNIT]);
}
try {
unlinkSync(unitPath);
} catch (err) {
return { ok: false, unitPath, message: `stopped the service but could not remove ${unitPath}: ${String(err)}` };
}
if (kind === 'systemd') run('systemctl', ['--user', 'daemon-reload']);
return { ok: true, unitPath, message: `service stopped and ${unitPath} removed. Your tmux sessions are untouched.` };
}
export async function serviceStatus(options: WebLaunchOptions): Promise<ServiceStatusResult> {
const kind = detectServiceKind();
const url = buildBaseUrl(options);
if (!kind) {
return { kind: null, name: '', unitPath: '', installed: false, loaded: false, responding: false, url };
}
const unitPath = unitPathFor(kind);
const name = kind === 'launchd' ? LAUNCHD_LABEL : SYSTEMD_UNIT;
const installed = existsSync(unitPath);
const loaded =
kind === 'launchd'
? run('launchctl', ['list', LAUNCHD_LABEL]).ok
: run('systemctl', ['--user', 'is-active', SYSTEMD_UNIT]).output === 'active';
const probe = await probeServer(buildStatusUrl(options), 2000);
return { kind, name, unitPath, installed, loaded, responding: probe.up, version: probe.version, url };
}
+6 -2
View File
@@ -121,7 +121,10 @@ export function buildClaudeEnv(sessionId: string): Record<string, string | undef
// Inform Claude it's running within Codeman (helps prevent self-termination)
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
// CODEMAN_API_URL rides in via the process.env spread when the server has
// stamped it (WebServer.start()); no fallback: a hardcoded one was the wrong
// scheme on HTTPS installs, and a present-with-undefined key would serialize
// as the literal "CODEMAN_API_URL=undefined" (COD-115).
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
};
@@ -179,7 +182,8 @@ export function buildShellEnv(sessionId: string): Record<string, string | undefi
TERM: 'xterm-256color',
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
// CODEMAN_API_URL rides in via the process.env spread when set; no fallback
// (same reasoning as buildClaudeEnv above).
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
};
+26 -6
View File
@@ -2620,20 +2620,23 @@ export class Session extends EventEmitter {
* For interactive sessions, this is how you send user input to Claude.
* Remember to include `\r` (carriage return) to simulate pressing Enter.
*
* @param data - The input data to send (text, escape sequences, etc.)
*
* @example
* ```typescript
* session.write('hello world'); // Text only, no Enter
* session.write('\r'); // Enter key
* session.write('ls -la\r'); // Command with Enter
* ```
*
* @param data - The input data to send (text, escape sequences, etc.)
* @returns true if the data reached a PTY. A session whose PTY is gone still
* discards the data, but it used to do so with no signal at all — which is how
* input could disappear while the caller believed it had been delivered.
*/
write(data: string): void {
write(data: string): boolean {
this._trackSubmit(data);
if (this.ptyProcess) {
this.ptyProcess.write(data);
}
if (!this.ptyProcess) return false;
this.ptyProcess.write(data);
return true;
}
// ── Conversation tracking ─────────────────────────────────────────────
@@ -2691,6 +2694,23 @@ export class Session extends EventEmitter {
return true;
}
/**
* Undo the bookkeeping of {@link shouldApplyInput} for a delivery that failed.
*
* Without this, the reliable-delivery layer guarantees exactly-once delivery of
* something that may never have been delivered: the seq is recorded as applied
* BEFORE the write is attempted, so a client retry — the very mechanism the seq
* exists for — is rejected as a duplicate and the input is lost for good.
*
* Only rolls back if `seq` is still the newest recorded one; a later input has
* already superseded it and must not be re-opened.
*/
forgetInputSeq(clientId: string, seq: number): void {
if (this._appliedInputSeq.get(clientId) === seq) {
this._appliedInputSeq.set(clientId, seq - 1);
}
}
/**
* Sends input via the terminal multiplexer's direct input mechanism.
*
+122 -41
View File
@@ -22,7 +22,8 @@
*/
import { EventEmitter } from 'node:events';
import { execSync, exec } from 'node:child_process';
import { collectDescendants } from './proc-tree.js';
import { execSync, exec, execFile } from 'node:child_process';
import { promisify } from 'node:util';
const execAsync = promisify(exec);
@@ -100,6 +101,16 @@ import {
// ============================================================================
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
/** How long a cached process snapshot stays usable. */
const PROC_SNAPSHOT_TTL_MS = 2000;
/**
* How long the kill path waits for a fresh snapshot before giving up on it.
* Shorter than EXEC_TIMEOUT_MS on purpose: killSession has two further strategies
* (process group, tmux kill-session) and must reach them even when `ps` is wedged.
*/
const PROC_SNAPSHOT_WAIT_MS = 1500;
import { DEFAULT_TMUX_HISTORY_LIMIT, DEFAULT_TERMINAL_BUFFER_MAX_BYTES } from './config/terminal-history.js';
/**
@@ -1578,7 +1589,11 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
'export CODEMAN_MUX=1',
`export CODEMAN_SESSION_ID=${sessionId}`,
`export CODEMAN_MUX_NAME=${muxName}`,
`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL || 'http://localhost:3000'}`,
// Only exported when the server has stamped the real URL (scheme+host+port,
// set in WebServer.start()). A hardcoded fallback here exported the wrong
// scheme on HTTPS installs; leaving the variable unset makes in-session
// guards fail closed instead of curling a URL that was never right.
...(process.env.CODEMAN_API_URL ? [`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL}`] : []),
// Path only (not the secret value): hook curl commands cat the file at
// execution time, so the COD-54 hook secret stays off the command line.
`export CODEMAN_HOOK_SECRET_FILE="${dataPath('hook-secret')}"`,
@@ -2094,27 +2109,102 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
}
// Get all child process PIDs recursively
private getChildPids(pid: number): number[] {
const pids: number[] = [];
try {
const output = execSync(`pgrep -P ${pid}`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (output) {
for (const childPid of output
.split('\n')
.map((p) => parseInt(p, 10))
.filter((p) => !Number.isNaN(p))) {
pids.push(childPid);
pids.push(...this.getChildPids(childPid));
/** One `ps` snapshot of the whole process table, cached briefly. */
private static procSnapshot: { at: number; byParent: Map<number, number[]> } | null = null;
/** Single-flight guard so a hung `ps` cannot pile up parallel refreshes. */
private static procRefresh: { started: number; promise: Promise<Map<number, number[]>> } | null = null;
/**
* Fork ONE `ps` asynchronously and cache the parent -> children map.
*
* Async on purpose: a synchronous fork here would block the event loop on every
* stats tick, and under the procfs pathology this module exists to survive,
* `execSync`'s timeout cannot return at all (spawnSync waits for the unkillable
* child) — freezing the whole server where a hung async poll only costs staleness.
*/
private static refreshProcSnapshot(): Promise<Map<number, number[]>> {
const inFlight = TmuxManager.procRefresh;
// Reuse an in-flight refresh — unless it is old enough to be presumed stuck.
if (inFlight && Date.now() - inFlight.started < EXEC_TIMEOUT_MS * 2) return inFlight.promise;
const started = Date.now();
const promise = new Promise<Map<number, number[]>>((resolve) => {
execFile('ps', ['-eo', 'pid=,ppid='], { timeout: EXEC_TIMEOUT_MS, maxBuffer: 8 * 1024 * 1024 }, (err, out) => {
if (TmuxManager.procRefresh?.started === started) TmuxManager.procRefresh = null;
if (err) {
// ANY error, not just an empty one: a timed-out or truncated `ps` yields
// partial output, and caching that as fresh would make whole subtrees
// invisible — including to the kill path. Stale beats wrong.
console.error('[TmuxManager] process snapshot failed:', err);
resolve(TmuxManager.procSnapshot?.byParent ?? new Map());
return;
}
}
} catch {
// No children or command failed
const byParent = new Map<number, number[]>();
for (const line of String(out).split('\n')) {
const parts = line.trim().split(/\s+/);
if (parts.length < 2) continue;
const pid = parseInt(parts[0], 10);
const ppid = parseInt(parts[1], 10);
if (Number.isNaN(pid) || Number.isNaN(ppid)) continue;
const list = byParent.get(ppid);
if (list) list.push(pid);
else byParent.set(ppid, [pid]);
}
TmuxManager.procSnapshot = { at: Date.now(), byParent };
resolve(byParent);
});
});
TmuxManager.procRefresh = { started, promise };
return promise;
}
/**
* Best snapshot WITHOUT forking: returns the cache, kicking off a background
* refresh when it has gone stale, and never blocks. Stats and window-title
* consumers tolerate data one interval old; nothing that KILLS may use this.
*/
private childrenByParent(): Map<number, number[]> {
const cached = TmuxManager.procSnapshot;
if (!cached || Date.now() - cached.at >= PROC_SNAPSHOT_TTL_MS) {
void TmuxManager.refreshProcSnapshot();
}
return pids;
return cached?.byParent ?? new Map();
}
/**
* Descendants from a snapshot that is not the cached one — the kill path's variant.
*
* killSession re-scans for survivors between SIGTERM and SIGKILL, and the wait in
* between (200ms) sits far inside the cache TTL (2000ms): reading the cache there
* returns the pre-SIGTERM state verbatim, so children spawned since are invisible
* and SIGKILL aims at stale PIDs, guarded only by kill(pid, 0) — which cannot
* detect PID reuse.
*
* It forces a refresh rather than guaranteeing recency: an already-running refresh
* is reused, so the snapshot can predate this call by up to one `ps` runtime. A
* strict postdate guarantee would mean chaining a second `ps` behind every
* in-flight one, which is the fork storm this code exists to avoid.
*
* Bounded by design: waiting forever would freeze killSession before it reaches
* its process-group and tmux fallbacks.
*/
private async getChildPidsFresh(pid: number): Promise<number[]> {
let byParent: ReadonlyMap<number, readonly number[]>;
try {
byParent = await Promise.race([
TmuxManager.refreshProcSnapshot(),
new Promise<never>((_, reject) =>
setTimeout(() => reject(new Error('proc snapshot timeout')), PROC_SNAPSHOT_WAIT_MS)
),
]);
} catch {
console.warn('[TmuxManager] process snapshot did not return in time; using the cached one');
byParent = TmuxManager.procSnapshot?.byParent ?? new Map<number, number[]>();
}
return collectDescendants(pid, byParent, {
onTruncated: (root, cap, reason) =>
console.warn(`[TmuxManager] descendant walk for ${root} hit the ${cap}-${reason} cap; truncating`),
});
}
// Check if a process is still alive
@@ -2221,7 +2311,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const allPids: number[] = [currentPid];
// Strategy 1: Kill all child processes recursively
let childPids = this.getChildPids(currentPid);
let childPids = await this.getChildPidsFresh(currentPid);
if (childPids.length > 0) {
console.log(`[TmuxManager] Found ${childPids.length} child processes to kill`);
allPids.push(...childPids);
@@ -2238,7 +2328,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
await new Promise((resolve) => setTimeout(resolve, TMUX_KILL_WAIT_MS));
childPids = this.getChildPids(currentPid);
childPids = await this.getChildPidsFresh(currentPid);
for (const childPid of childPids) {
if (this.isProcessAlive(childPid)) {
try {
@@ -2452,17 +2542,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const [rss, cpu] = psOutput.split(/\s+/).map((x) => parseFloat(x) || 0);
// From the shared snapshot: this runs per session on every stats tick, and a
// pgrep per session was a fork per session per interval.
let childCount = 0;
try {
const childOutput = (
await execAsync(`pgrep -P ${session.pid} | wc -l`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
})
).stdout.trim();
childCount = parseInt(childOutput, 10) || 0;
childCount = (this.childrenByParent().get(session.pid) ?? []).length;
} catch {
// No children or command failed
// No children or snapshot unavailable
}
return {
@@ -2496,17 +2582,12 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
// Step 1: Get descendant PIDs
const descendantMap = new Map<number, number[]>();
const pgrepOutput = (
await execAsync(
`for p in ${sessionPids.join(' ')}; do children=$(pgrep -P $p 2>/dev/null | tr '\\n' ','); echo "$p:$children"; done`,
{
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}
)
).stdout.trim();
// Derived from the ONE snapshot instead of a shell loop that forks a pgrep
// per session — the shape that turned into a fork storm under load.
const byParent = this.childrenByParent();
const childLines = sessionPids.map((p) => `${p}:${(byParent.get(p) ?? []).join(',')}`).join('\n');
for (const line of pgrepOutput.split('\n')) {
for (const line of childLines.split('\n')) {
const [pidStr, childrenStr] = line.split(':');
const sessionPid = parseInt(pidStr, 10);
if (!Number.isNaN(sessionPid)) {
+95 -25
View File
@@ -108,12 +108,100 @@ export function getAugmentedPath(): string {
return _augmentedPath;
}
/** Cached `claude --version` result: string = version, null = probed but unavailable, undefined = not probed */
let _claudeVersion: string | null | undefined = undefined;
/**
* Cache state for the `claude --version` probe.
*
* `version` is only ever set from a SUCCESSFUL probe and then kept for the
* process lifetime (the binary can't change under a running server without a
* restart). Failures are tracked separately so they expire.
*/
export interface ClaudeVersionProbeState {
/** Successful probe result; `undefined` until one succeeds. */
version?: string;
/** Consecutive failed probes (drives the retry backoff). */
failures: number;
/** Timestamp of the most recent failed probe. */
lastFailureAt: number;
}
/** First retry window after a failed probe. */
const VERSION_PROBE_BASE_RETRY_MS = 60_000;
/** Ceiling for the doubling backoff, so a permanently missing binary settles down. */
const VERSION_PROBE_MAX_RETRY_MS = 15 * 60_000;
/**
* How long to wait before re-probing after `failures` consecutive failures:
* 1min, 2min, 4min… capped at 15min. Exported for tests.
*/
export function claudeVersionRetryDelayMs(failures: number): number {
if (failures <= 0) return 0;
return Math.min(VERSION_PROBE_BASE_RETRY_MS * 2 ** (failures - 1), VERSION_PROBE_MAX_RETRY_MS);
}
/**
* Cache policy for the version probe, pure apart from the `state` it mutates
* and the injected `probe` (exported so tests can drive it with a fake clock).
*
* Success is cached forever; FAILURE is not. That asymmetry is the fix for a
* real shipped bug: the old cache stored `null` on any exception and guarded on
* `!== undefined`, so a single failed probe — a 5s `EXEC_TIMEOUT_MS` timeout, a
* PATH-starved systemd/launchd environment, a transient fs hiccup — at the FIRST
* Claude session start left `cliVersion` undefined for EVERY Claude session
* until the server restarted. An undefined `cliVersion` silently disables
* wheel-forwarding to Claude's own transcript (`_shouldForwardWheelToApp`),
* which is the only route to history in repaint mode: a dead wheel on every
* device at once, matching the issue #205 retest reports.
*
* Retries back off so a genuinely absent binary still can't spawn a probe per
* session start.
*/
export function resolveClaudeCliVersion(
state: ClaudeVersionProbeState,
now: number,
probe: () => string | null
): string | null {
if (state.version !== undefined) return state.version;
if (state.failures > 0 && now - state.lastFailureAt < claudeVersionRetryDelayMs(state.failures)) return null;
let version: string | null = null;
try {
version = probe();
} catch {
version = null;
}
if (version) {
state.version = version;
state.failures = 0;
state.lastFailureAt = 0;
return version;
}
state.failures += 1;
state.lastFailureAt = now;
return null;
}
const _claudeVersionState: ClaudeVersionProbeState = { failures: 0, lastFailureAt: 0 };
/** One `claude --version` run. Throws on spawn/timeout failure. */
function probeClaudeCliVersion(): string | null {
const dir = findClaudeDir();
const bin = dir ? join(dir, 'claude') : 'claude';
// execFileSync (no shell) — the resolved path may contain spaces, and there
// is no untrusted input, but avoid a shell either way.
const out = execFileSync(bin, ['--version'], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
env: { ...process.env, PATH: getAugmentedPath() },
});
const match = out.match(/(\d+\.\d+\.\d+)/);
return match ? match[1] : null;
}
/**
* Returns the installed Claude CLI version (e.g. `"2.1.210"`), or null if it
* can't be determined. Runs `claude --version` once and caches the result.
* can't be determined. Runs `claude --version` at most once per successful
* resolution; failed probes retry with backoff (see `resolveClaudeCliVersion`).
*
* This is a deterministic alternative to scraping the interactive startup
* banner (`parseClaudeCodeInfo` in session.ts): newer Claude Code builds don't
@@ -122,28 +210,10 @@ let _claudeVersion: string | null | undefined = undefined;
* gated on it (e.g. wheel-forwarding to Claude's transcript — issue #154).
*/
export function getClaudeCliVersion(): string | null {
if (_claudeVersion !== undefined) return _claudeVersion;
// Keep the test suite hermetic — never spawn a real `claude` subprocess under
// vitest (matches IS_TEST_MODE in tmux-manager). Tests that need a version set
// it on the session directly.
if (process.env.VITEST) {
_claudeVersion = null;
return _claudeVersion;
}
try {
const dir = findClaudeDir();
const bin = dir ? join(dir, 'claude') : 'claude';
// execFileSync (no shell) — the resolved path may contain spaces, and there
// is no untrusted input, but avoid a shell either way.
const out = execFileSync(bin, ['--version'], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
env: { ...process.env, PATH: getAugmentedPath() },
});
const match = out.match(/(\d+\.\d+\.\d+)/);
_claudeVersion = match ? match[1] : null;
} catch {
_claudeVersion = null;
}
return _claudeVersion;
// it on the session directly. Deliberately does NOT touch the cache state:
// recording a phantom failure here would be the very poisoning this fixes.
if (process.env.VITEST) return null;
return resolveClaudeCliVersion(_claudeVersionState, Date.now(), probeClaudeCliVersion);
}
+2
View File
@@ -17,6 +17,8 @@ export interface ConfigPort {
getModelConfig(): Promise<{ defaultModel?: string; agentTypeOverrides?: Record<string, string> } | null>;
getClaudeModeConfig(): Promise<{ claudeMode?: ClaudeMode; allowedTools?: string }>;
getTerminalHistoryConfig(): Promise<TerminalHistoryConfig>;
/** Synced `agentSkillEnabled` app setting (default OFF); gates per-case agent-skill injection. */
getAgentSkillEnabled(): Promise<boolean>;
getDefaultClaudeMdPath(): Promise<string | undefined>;
getLightState(identity?: { username: string; role: 'admin' | 'user' }): unknown;
getLightSessionsState(): unknown[];
+45 -8
View File
@@ -519,6 +519,10 @@ class CodemanApp {
// Cooldown per session for the scroll-to-top "load more history" re-pull.
this._fullHistoryRepullAt = new Map(); // Map<sessionId, timestamp>
this._fullHistoryRepullInFlight = false;
// Sessions whose last re-pull came back THINNER than the live buffer (a
// repaint-mode CLI pane, where tmux keeps no history of its own). The pull is
// refused for those and retried far more slowly — see _maybeRefetchFullHistory.
this._fullHistoryRepullUseless = new Set();
this.terminalLoadStates = new Map(); // Map<sessionId, { generation, phase }>
this.respawnStatus = {};
this.respawnTimers = {}; // Track timed respawn timers
@@ -823,6 +827,7 @@ class CodemanApp {
this.applyLocalization();
this.applyTabWrapSettings();
this.applyMonitorVisibility();
this._setupTabMiddleClickClose();
// Must run before the first session:created can arrive: markSessionTabEntering()
// ignores ids until this sets up its state, which is what keeps the tabs
// restored on page load from animating.
@@ -3485,11 +3490,12 @@ class CodemanApp {
}
}
} else if (minimizedCount > 0 && !subagentBadgeEl) {
// Need to add badge - insert before gear icon
// Need to add badge - insert before the action-icon overlay so the
// badge stays a direct child of the tab (outside .tab-actions)
const badgeHtml = this.renderSubagentTabBadge(id, minimizedAgents);
const gearEl = tab.querySelector('.tab-gear');
if (gearEl) {
gearEl.insertAdjacentHTML('beforebegin', badgeHtml);
const actionsEl = tab.querySelector('.tab-actions');
if (actionsEl) {
actionsEl.insertAdjacentHTML('beforebegin', badgeHtml);
}
} else if (minimizedCount === 0 && subagentBadgeEl) {
// Count went to 0 - remove badge
@@ -3543,6 +3549,25 @@ class CodemanApp {
container.classList.toggle('tabs-auto-wrap', shouldWrap);
}
// Middle-click closes a tab, mirroring browser tab strips. Session tabs go
// through requestCloseSession (the same confirm modal as the x button), web
// tabs through closeWebviewTab (same as theirs). Delegated on the container:
// tabs are re-rendered wholesale, the container is stable.
_setupTabMiddleClickClose() {
const container = this.$('sessionTabs');
if (!container || this._tabAuxClickBound) return;
this._tabAuxClickBound = true;
container.addEventListener('auxclick', (e) => {
if (e.button !== 1) return;
const tab = e.target.closest?.('.session-tab');
if (!tab) return;
e.preventDefault();
e.stopPropagation();
if (tab.dataset.id) this.requestCloseSession(tab.dataset.id);
else if (tab.dataset.webviewId) this.closeWebviewTab?.(tab.dataset.webviewId);
});
}
_fullRenderSessionTabs() {
if (this._inlineRenameActive) return;
const container = this.$('sessionTabs');
@@ -3607,9 +3632,7 @@ class CodemanApp {
${hasRunningTasks ? `<span class="tab-badge" onclick="event.stopPropagation(); app.toggleTaskPanel()" aria-label="${taskStats.running} running tasks">${taskStats.running}</span>` : ''}
${subagentBadge}
${ultracodeBadge}
<span class="tab-gear" onclick="event.stopPropagation(); app.openSessionOptions(${escapeHtml(JSON.stringify(id))})" title="Session options" aria-label="Session options" tabindex="0">&#x2699;</span>
<span class="tab-detach" onclick="event.stopPropagation(); app.detachSession(${escapeHtml(JSON.stringify(id))})" title="Open in a new window" aria-label="Open session in a new window" tabindex="0">&#x29C9;</span>
<span class="tab-close" onclick="event.stopPropagation(); app.requestCloseSession(${escapeHtml(JSON.stringify(id))})" title="Close session" aria-label="Close session" tabindex="0">&times;</span>
<span class="tab-actions"><span class="tab-gear" onclick="event.stopPropagation(); app.openSessionOptions(${escapeHtml(JSON.stringify(id))})" title="Session options" aria-label="Session options" tabindex="0">&#x2699;</span><span class="tab-detach" onclick="event.stopPropagation(); app.detachSession(${escapeHtml(JSON.stringify(id))})" title="Open in a new window" aria-label="Open session in a new window" tabindex="0">&#x29C9;</span><span class="tab-close" onclick="event.stopPropagation(); app.requestCloseSession(${escapeHtml(JSON.stringify(id))})" title="Close session" aria-label="Close session" tabindex="0">&times;</span></span>
</div>`);
_tabIdx++;
}
@@ -4121,6 +4144,13 @@ class CodemanApp {
* On demand rather than automatic because that capture is unbounded-ish work: at
* the default history limit it can be megabytes, which is fine to pay when the
* user is explicitly reaching for history and not fine on every tab switch.
*
* NEVER a downgrade: for a repaint-mode CLI pane tmux keeps no history of its
* own, so the capture can be THINNER than what xterm already holds and the
* reset+rewrite below would delete history mid-scroll. `_replayWouldShrinkBuffer`
* (terminal-ui.js) is the guard, and a session that produced one useless re-pull
* gets a much longer cooldown so a hollow pane stops re-fetching megabytes on
* every scroll-up (issue #205, round 2).
*/
async _maybeRefetchFullHistory() {
const sessionId = this.activeSessionId;
@@ -4129,7 +4159,8 @@ class CodemanApp {
const now = Date.now();
// Momentum scrolling fires this dozens of times per flick, and a burst of new
// output is the normal reason to want a re-pull, so cooldown rather than latch.
if (now - (this._fullHistoryRepullAt.get(sessionId) || 0) < 4000) return;
const cooldown = this._fullHistoryRepullUseless?.has(sessionId) ? 60000 : 4000;
if (now - (this._fullHistoryRepullAt.get(sessionId) || 0) < cooldown) return;
this._fullHistoryRepullAt.set(sessionId, now);
this._fullHistoryRepullInFlight = true;
try {
@@ -4138,6 +4169,12 @@ class CodemanApp {
// Bail on a tab switch mid-fetch: writing here would paint another session's
// history into the terminal the user is now looking at.
if (!buffer || this.activeSessionId !== sessionId) return;
if (this._replayWouldShrinkBuffer(buffer)) {
(this._fullHistoryRepullUseless ||= new Set()).add(sessionId);
this._logScrollRouting?.('repull-refused-downgrade');
return;
}
this._fullHistoryRepullUseless?.delete(sessionId);
const rowsBefore = this.terminal.buffer.active.length;
this._resetTerminalForReplay();
await this.chunkedTerminalWrite(buffer, TERMINAL_CHUNK_SIZE, sessionId);
+1
View File
@@ -227,6 +227,7 @@
'Redraw Terminal Button': '重绘终端按钮',
'Tab Bar': '标签栏',
'Tall Tabs (Name + Folder)': '双行标签(名称 + 文件夹)',
'Pop-out Button on Tabs': '标签页弹出窗口按钮',
Panels: '面板',
Monitor: '监视器',
'Project Insights': '项目洞察',
+16 -1
View File
@@ -1322,7 +1322,7 @@
</div>
<!-- Input Section -->
<div class="settings-section-header">Input</div>
<div class="settings-item settings-item-multiline" title="Scroll the terminal's own local scrollback with a plain mouse wheel / two-finger swipe, instead of forwarding the wheel to the CLI's transcript. Turn on if scrolling back through history doesn't work (e.g. macOS trackpad in Claude sessions). Shift+wheel always reaches local scrollback regardless.">
<div class="settings-item settings-item-multiline" title="Scroll the terminal's own local scrollback with a plain mouse wheel / two-finger swipe, instead of forwarding the wheel to the CLI's transcript. Only Claude sessions forward, so this setting only affects them: Claude redraws the screen in place and keeps almost no local scrollback, so with this ON the wheel has little to scroll and Codeman falls back to paging Claude's own transcript. Codex, Gemini, shell and OpenCode sessions always scroll local scrollback. Shift+wheel always reaches local scrollback regardless.">
<div class="settings-item-text">
<span class="settings-item-label">Wheel Scrolls Local History</span>
<span class="settings-item-desc">Plain wheel/trackpad pages the terminal scrollback</span>
@@ -1480,6 +1480,13 @@
<span class="slider"></span>
</label>
</div>
<div class="settings-item" title="Show the open-in-new-window (pop-out) button when hovering a session tab">
<span class="settings-item-label">Pop-out Button on Tabs</span>
<label class="switch switch-sm">
<input type="checkbox" id="appSettingsShowTabDetachButton">
<span class="slider"></span>
</label>
</div>
<!-- Panels Section -->
<div class="settings-section-header">Panels</div>
@@ -1630,6 +1637,14 @@
</label>
<span class="form-hint">Enable experimental Agent Teams for all new Claude sessions (disabled by default)</span>
</div>
<div class="form-row form-row-switch">
<label>Agent Skill</label>
<label class="switch">
<input type="checkbox" id="appSettingsAgentSkill">
<span class="slider"></span>
</label>
<span class="form-hint">Give new Claude sessions the Codeman skill (start workers, send prompts, wait for results via the API)</span>
</div>
<div class="form-row">
<label>Claude Model</label>
<select id="appSettingsClaudeModel" class="form-select">
+62 -37
View File
@@ -209,9 +209,13 @@ const MobileDetection = {
* Also handles terminal scrolling and toolbar repositioning via visualViewport API.
*/
const KeyboardHandler = {
VIEWPORT_SETTLE_MS: 80,
lastViewportHeight: 0,
keyboardVisible: false,
initialViewportHeight: 0,
_viewportSettleTimer: null,
_settleScrollToBottom: false,
_settlePending: false,
/** Initialize keyboard handling */
init() {
@@ -276,6 +280,12 @@ const KeyboardHandler = {
window.removeEventListener('scroll', this._windowScrollHandler);
this._windowScrollHandler = null;
}
if (this._viewportSettleTimer) {
clearTimeout(this._viewportSettleTimer);
this._viewportSettleTimer = null;
}
this._settleScrollToBottom = false;
this._settlePending = false;
},
/** Handle viewport resize (keyboard show/hide) */
@@ -313,6 +323,7 @@ const KeyboardHandler = {
}
this.updateLayoutForKeyboard();
this._deferViewportSettle();
this.lastViewportHeight = currentHeight;
},
@@ -414,32 +425,9 @@ const KeyboardHandler = {
// iOS Safari may scroll the document to reveal xterm's hidden textarea.
window.scrollTo(0, 0);
// Refit terminal locally AND send resize to server so Claude Code (Ink)
// knows the actual terminal dimensions. Without this, Ink redraws at the
// old (larger) row count when the user types, causing content to scroll
// off the visible area with each keystroke.
// Note: the throttledResize handler still suppresses ongoing resize events
// while keyboard is up — this one-shot resize on open/close is sufficient.
setTimeout(() => {
if (typeof app !== 'undefined' && app.terminal) {
if (app.fitAddon)
try {
app.fitAddon.fit();
} catch {}
// Eliminate terminal row quantization gap: xterm can only show whole
// rows, so leftover pixels create dead space below the last row.
// Shrink .main's paddingBottom by the gap so the terminal fills flush
// to the accessory bar.
this._shrinkPaddingToFit();
app.terminal.scrollToBottom();
app._syncMobileHelperTextareaToCursor?.();
app._localEchoOverlay?.rerender?.();
// Send resize to server so PTY dimensions match xterm
this._sendTerminalResize();
}
// Reset again after fit/resize in case layout changes triggered scroll
window.scrollTo(0, 0);
}, 150);
// visualViewport emits multiple heights throughout the OS animation.
// Re-schedule on every event and fit only after the final height settles.
this._scheduleViewportSettle({ scrollToBottom: true });
// Reposition subagent windows to stack from bottom (above keyboard)
if (typeof app !== 'undefined') app.relayoutMobileSubagentWindows();
@@ -454,22 +442,59 @@ const KeyboardHandler = {
this.resetLayout();
// Refit terminal, scroll to bottom, and send resize to restore original dimensions
setTimeout(() => {
if (typeof app !== 'undefined' && app.fitAddon) {
try {
app.fitAddon.fit();
} catch {}
if (app.terminal) app.terminal.scrollToBottom();
// Send resize to server to restore full terminal size
this._sendTerminalResize();
}
}, 100);
this._scheduleViewportSettle({ scrollToBottom: true });
// Reposition subagent windows to stack from top (below header)
if (typeof app !== 'undefined') app.relayoutMobileSubagentWindows();
},
/**
* Coalesce the keyboard animation into one final xterm reflow and PTY resize.
* Only a real show/hide transition arms the settle work; ongoing viewport
* resize events merely push a pending settle back (_deferViewportSettle).
* A viewport change that never crosses the show/hide thresholds must not
* refit: keyboard detection can miss a fine-grained OS animation entirely
* (each step under 150px, with the baseline chasing the animation), and the
* container is then mid-animation with no keyboard CSS compensation, so a
* fit against it resizes the PTY to transient dims and the SIGWINCH thrash
* garbles the transcript.
*/
_scheduleViewportSettle({ scrollToBottom = false } = {}) {
this._settleScrollToBottom = this._settleScrollToBottom || scrollToBottom;
this._settlePending = true;
this._armViewportSettleTimer();
},
/** Push a pending settle back while the viewport is still animating; no-op otherwise. */
_deferViewportSettle() {
if (!this._settlePending) return;
this._armViewportSettleTimer();
},
_armViewportSettleTimer() {
if (this._viewportSettleTimer) clearTimeout(this._viewportSettleTimer);
this._viewportSettleTimer = setTimeout(() => {
this._viewportSettleTimer = null;
this._settlePending = false;
const shouldScrollToBottom = this._settleScrollToBottom;
this._settleScrollToBottom = false;
if (typeof app !== 'undefined' && app.terminal) {
if (app.fitAddon) {
try {
app.fitAddon.fit();
} catch {}
}
if (this.keyboardVisible) this._shrinkPaddingToFit();
if (shouldScrollToBottom) app.terminal.scrollToBottom();
app._syncMobileHelperTextareaToCursor?.();
app._localEchoOverlay?.rerender?.();
this._sendTerminalResize();
}
window.scrollTo(0, 0);
}, this.VIEWPORT_SETTLE_MS);
},
/** Send current terminal dimensions to the server (one-shot, for keyboard open/close) */
_sendTerminalResize() {
if (typeof app === 'undefined' || !app.activeSessionId || !app.fitAddon) return;
+13
View File
@@ -363,6 +363,7 @@ Object.assign(CodemanApp.prototype, {
document.getElementById('appSettingsCjkInput').checked = settings.cjkInputEnabled ?? defaults.cjkInputEnabled ?? false;
document.getElementById('appSettingsExtendedKeyboardBar').checked = settings.extendedKeyboardBar ?? false;
document.getElementById('appSettingsTabTwoRows').checked = settings.tabTwoRows ?? defaults.tabTwoRows ?? false;
document.getElementById('appSettingsShowTabDetachButton').checked = settings.showTabDetachButton ?? defaults.showTabDetachButton ?? false;
// Claude CLI settings
const claudeModeSelect = document.getElementById('appSettingsClaudeMode');
const allowedToolsRow = document.getElementById('allowedToolsRow');
@@ -384,6 +385,7 @@ Object.assign(CodemanApp.prototype, {
this._applyCodexSettingsVisibility();
// Claude Permissions settings
document.getElementById('appSettingsAgentTeams').checked = settings.agentTeamsEnabled ?? false;
document.getElementById('appSettingsAgentSkill').checked = settings.agentSkillEnabled ?? false;
document.getElementById('appSettingsClaudeModel').value = settings.claudeModel ?? '';
document.getElementById('appSettingsOpusContext1m').checked = settings.opusContext1mEnabled ?? false;
document.getElementById('appSettingsRemoteAutoReconnect').checked = settings.remoteAutoReconnect ?? true;
@@ -1542,6 +1544,7 @@ Object.assign(CodemanApp.prototype, {
webglRendererEnabled: document.getElementById('appSettingsWebglRenderer').checked,
extendedKeyboardBar: document.getElementById('appSettingsExtendedKeyboardBar').checked,
tabTwoRows: document.getElementById('appSettingsTabTwoRows').checked,
showTabDetachButton: document.getElementById('appSettingsShowTabDetachButton').checked,
skin: document.getElementById('appSettingsSkin').value,
// Claude CLI settings
claudeMode: document.getElementById('appSettingsClaudeMode').value,
@@ -1551,6 +1554,7 @@ Object.assign(CodemanApp.prototype, {
codexAnimationsEnabled: document.getElementById('appSettingsCodexAnimations').checked,
// Claude Permissions settings
agentTeamsEnabled: document.getElementById('appSettingsAgentTeams').checked,
agentSkillEnabled: document.getElementById('appSettingsAgentSkill').checked,
claudeModel: document.getElementById('appSettingsClaudeModel').value,
opusContext1mEnabled: document.getElementById('appSettingsOpusContext1m').checked,
remoteAutoReconnect: document.getElementById('appSettingsRemoteAutoReconnect').checked,
@@ -1726,6 +1730,7 @@ Object.assign(CodemanApp.prototype, {
showSessionButton: _ssb,
showAwayDigestButton: _adb,
showCronButton: _crb,
showTabDetachButton: _tdb,
// Phone-only home surface, and absent from SettingsUpdateSchema (.strict()).
mobileOverviewEnabled: _mov,
...serverSettings
@@ -1999,6 +2004,13 @@ Object.assign(CodemanApp.prototype, {
applyHeaderVisibilitySettings() {
const settings = this.loadAppSettingsFromStorage();
const defaults = this.getDefaultSettings();
// Tab pop-out (open-in-new-window) button: opt-in (App Settings → Tab Bar,
// default OFF, per-device). Mirrored as a class on <html>: styles.css hides
// .tab-detach without it (a tab that is already detached keeps its icon as
// the re-focus affordance for the popped-out window).
const showTabDetach = settings.showTabDetachButton ?? defaults.showTabDetachButton ?? false;
document.documentElement.classList.toggle('tabs-show-detach', showTabDetach);
const compactHeader = MobileDetection.getDeviceType() !== 'desktop';
const showFontControls = compactHeader ? false : (settings.showFontControls ?? defaults.showFontControls ?? false);
const showSystemStats = compactHeader ? false : (settings.showSystemStats ?? defaults.showSystemStats ?? true);
@@ -2355,6 +2367,7 @@ Object.assign(CodemanApp.prototype, {
'language',
'terminalWheelLocalScrollback',
'showSessionButton', 'showAwayDigestButton', 'showCronButton',
'showTabDetachButton',
'mobileOverviewEnabled',
]);
// The plan-usage chip is a PER-DEVICE display setting (desktop default ON,
+30 -3
View File
@@ -1421,7 +1421,11 @@ html[data-line-anim="packet"] .connection-line.line-enter {
transition: opacity 0.05s ease-out, width 0.05s ease-out, padding 0.05s ease-out;
}
.session-tab:hover .tab-close {
/* Icons expand on the ACTIVE tab only (in flow): selection is a deliberate
click, so the width change never happens while aiming at a tab, background
tabs keep their full title on hover, and a stray click can only switch.
Middle-click closes any tab (_setupTabMiddleClickClose in app.js). */
.session-tab.active .tab-close {
opacity: 1;
width: auto;
padding: 0.15rem 0.35rem;
@@ -1983,7 +1987,7 @@ html[data-line-anim="packet"] .connection-line.line-enter {
transition: opacity 0.15s, width 0.15s, padding 0.15s, transform 0.2s;
}
.session-tab:hover .tab-gear {
.session-tab.active .tab-gear {
opacity: 1;
width: auto;
padding: 0 0.3rem;
@@ -2008,7 +2012,7 @@ html[data-line-anim="packet"] .connection-line.line-enter {
overflow: hidden;
transition: opacity 0.15s, width 0.15s, padding 0.15s;
}
.session-tab:hover .tab-detach {
.session-tab.active .tab-detach {
opacity: 1;
width: auto;
padding: 0 0.3rem;
@@ -2047,6 +2051,29 @@ html[data-line-anim="packet"] .connection-line.line-enter {
display: inline-flex;
}
/* ===== Tab action icons: active tab only =================================
All three per-tab icons live in a .tab-actions wrapper (in flow; it adds
no width of its own while the children keep the width:0 collapse above).
They expand only on the ACTIVE tab: selection is a deliberate click, so
the tab-strip geometry never shifts while the pointer is aiming, hovering
a background tab changes nothing (full title stays readable), and a stray
click can only switch sessions. Middle-click closes any tab. The phone
layout in mobile.css follows the same active-only pattern with its own
sizing; the tablet touch fallback there keeps icons always visible. */
.session-tab .tab-actions {
display: flex;
align-items: center;
}
/* Pop-out button is opt-in (App Settings → Tab Bar, default off; per-device).
settings-ui.js mirrors the setting as the tabs-show-detach class on <html>.
A tab that is ALREADY detached keeps its icon regardless: it is the
re-focus affordance for the popped-out window. */
html:not(.tabs-show-detach) .session-tab:not(.detached) .tab-detach {
display: none;
}
/* ===== Solo (detached single-session) window chrome ===================== */
body.solo-mode .session-tabs,
body.solo-mode .header-system-stats,
+334 -45
View File
@@ -23,6 +23,44 @@
// short window, only the app's synthetic tap-to-position mouse event should
// reach xterm.
const TOUCH_COMPAT_MOUSE_SUPPRESS_MS = 450;
// Escape sequences occupy no terminal cells, so they must come out before a
// captured line's WIDTH can be measured (_estimateReplayRows). Covers OSC,
// CSI, charset designators and the short escapes tmux emits; deliberately
// approximate — this feeds a size comparison, not a renderer.
// eslint-disable-next-line no-control-regex
const REPLAY_ESCAPE_RE =
/\x1b\][^\x07\x1b]*(?:\x07|\x1b\\)|\x1b\[[0-9;?<>=!]*[ -/]*[@-~]|\x1b[()#][0-9A-Za-z]|\x1b[=>78M]/g;
// PageUp / PageDown as xterm.js encodes them. Used as the LAST-RESORT scroll
// gesture for a repaint-mode CLI whose local buffer holds no scrollback
// (_maybePageCliTranscript).
const KEY_PAGE_UP = '\x1b[5~';
const KEY_PAGE_DOWN = '\x1b[6~';
// Wheel/touch travel (in lines) that adds up to one PageUp/PageDown. Half a
// screen rather than a full one: the page key always jumps a whole screen, so
// a 1:1 mapping made the fallback feel unreachably slow with a discrete mouse
// wheel (Firefox reports 3 lines a notch → 12 notches per page). Overshooting
// the finger is the right trade against a gesture that otherwise does nothing.
const PAGE_KEY_SCREEN_FRACTION = 0.5;
// Bound on page keys emitted from one gesture batch, mirroring the SGR tick
// cap: a fling must not build a backlog that keeps paging after it stops.
const PAGE_KEY_MAX_PER_BATCH = 3;
// Composer navigation keys as xterm.js encodes user keystrokes: plain and
// modified arrows (CSI A-D, CSI 1;mA-D, SS3 A-D), Home/End (CSI H/F, SS3
// H/F, CSI 1~/4~), Insert/Delete/PgUp/PgDn (CSI 2~/3~/5~/6~, optional
// modifier). Deliberately EXCLUDES terminal query responses that also
// arrive via onData (DA `\x1b[?1;2c`, CPR `\x1b[12;34R`) and function
// keys, so only genuine cursor/editing keys trigger the local-echo flush.
// eslint-disable-next-line no-control-regex
const COMPOSER_NAV_KEY_PATTERN = /^\x1b(?:\[(?:[ABCDHF]|1;[2-8][ABCDHF]|[1-8](?:;[2-8])?~)|O[ABCDHF])$/;
// Prefix xterm.js puts on terminal.paste() payloads while the application
// has bracketed-paste mode (DECSET 2004) enabled. Codex, Claude Code and
// tmux all enable it, so browser pastes arrive as one onData chunk of
// `\x1b[200~<text>\x1b[201~`.
const BRACKETED_PASTE_START = '\x1b[200~';
function isComposerNavKey(data) {
return COMPOSER_NAV_KEY_PATTERN.test(data);
}
function isTerminalQueryResponse(data) {
return TERMINAL_QUERY_RESPONSE_PATTERN.test(data) || TERMINAL_OSC_RESPONSE_PATTERN.test(data);
@@ -60,8 +98,15 @@
global.CodemanTerminalInput = {
isTerminalQueryResponse,
shouldSuppressTerminalQueryResponse,
isComposerNavKey,
BRACKETED_PASTE_START,
USER_SCROLL_STICKY_SUPPRESS_MS,
TOUCH_COMPAT_MOUSE_SUPPRESS_MS,
REPLAY_ESCAPE_RE,
KEY_PAGE_UP,
KEY_PAGE_DOWN,
PAGE_KEY_SCREEN_FRACTION,
PAGE_KEY_MAX_PER_BATCH,
};
global.CODEMAN_XTERM_THEMES = CODEMAN_XTERM_THEMES;
global.codemanCurrentXtermTheme = currentXtermTheme;
@@ -296,7 +341,7 @@ Object.assign(CodemanApp.prototype, {
// WebGL renderer for GPU-accelerated terminal rendering.
// Previously caused "page unresponsive" crashes from synchronous GPU stalls,
// but the 48KB/frame flush cap in flushPendingWrites() now prevents
// but the mode-aware 32/64KB frame cap in flushPendingWrites() now prevents
// oversized terminal.write() calls that triggered the stalls.
// Disable with ?nowebgl URL param if GPU issues return.
// Auto-fallback: _initWebGL installs a long-task watchdog that disables
@@ -407,8 +452,8 @@ Object.assign(CodemanApp.prototype, {
this.registerFilePathLinkProvider();
// Mouse wheel: forward to the TUI only for sessions verified to handle SGR
// wheel reports (codex, and claude 2.1.187+ — see _shouldForwardWheelToApp),
// local scrollback otherwise. Claude Code 2.1.187+ scrolls its own
// wheel reports (claude 2.1.187+ — see _shouldForwardWheelToApp), local
// scrollback otherwise. Claude Code 2.1.187+ scrolls its own
// transcript on SGR wheel reports — scrolled-away tool blocks re-render
// live and stay clickable — and its select menus no longer capture wheel
// as option navigation (verified against 2.1.202: /model menu highlight
@@ -452,6 +497,7 @@ Object.assign(CodemanApp.prototype, {
ev.preventDefault();
ev.stopPropagation();
if (this._shouldForwardWheelToApp(ev)) {
this._logScrollRouting('forward-sgr');
this._forwardScrollToApp(ev.clientX, ev.clientY, this._wheelScrollLines(ev));
return;
}
@@ -459,6 +505,10 @@ Object.assign(CodemanApp.prototype, {
// a stream of tiny pixel deltas, and rounding each one to a whole line
// (the ±1 fallback) made slow drags scroll faster than the finger.
const lines = this._wheelScrollLinesFloat(ev);
// …unless there is no local scrollback to scroll, in which case page the
// CLI's own transcript instead of doing nothing (_maybePageCliTranscript).
if (this._maybePageCliTranscript(ev, lines)) return;
this._logScrollRouting('local-scrollback');
this._noteTerminalUserScroll(lines);
this._smoothScrollBy(lines);
},
@@ -476,7 +526,10 @@ Object.assign(CodemanApp.prototype, {
// drags the CLI's pinned input box off the screen (issue #205's mobile
// half). Same gate, so Shift has no touch analog but the local-scrollback
// opt-out setting and the CLI-version gate apply to touch exactly as they
// do to the wheel.
// do to the wheel — including the PageUp/PageDown fallback the wheel uses
// when that gate is false and there is no local scrollback to scroll
// (_maybePageCliTranscript), which is what keeps a swipe from being a
// complete no-op on a phone.
{
const cellHeight = () => this.terminal._core?._renderService?.dimensions?.css?.cell?.height || 13;
let touchLastX = 0;
@@ -498,7 +551,7 @@ Object.assign(CodemanApp.prototype, {
// Flick momentum keeps feeding the CLI's transcript from the last
// touch point; the 40ms coalescer batches the per-frame reports.
this._forwardScrollToApp(touchLastX, touchLastY, lines);
} else {
} else if (!this._maybePageCliTranscript({ shiftKey: false }, lines)) {
this.terminal.scrollLines(lines);
this._maybeLoadMoreHistoryOnScroll(lines);
}
@@ -566,8 +619,10 @@ Object.assign(CodemanApp.prototype, {
const lines = Math.trunc(pixelAccum / ch);
if (lines !== 0) {
if (this._shouldForwardWheelToApp({ shiftKey: false })) {
this._logScrollRouting('forward-sgr');
this._forwardScrollToApp(touchLastX, touchLastY, lines);
} else {
} else if (!this._maybePageCliTranscript({ shiftKey: false }, lines)) {
this._logScrollRouting('local-scrollback');
this._noteTerminalUserScroll(lines);
this.terminal.scrollLines(lines);
this._maybeLoadMoreHistoryOnScroll(lines);
@@ -832,11 +887,23 @@ Object.assign(CodemanApp.prototype, {
}
this._lastTerminalData = { data, time: performance.now() };
// ── Local Echo Pass-through ──
// After a composer nav key (arrow/Home/End/Delete) the real cursor may
// sit mid-text, where the overlay's append-only buffering would corrupt
// both the preview and the submitted text. Such sessions are handed
// back to plain PTY echo until Enter or Ctrl+C submits/cancels the
// composer line (see the nav-key branch below).
const echoPassthrough =
this._localEchoEnabled && this._echoPassthroughSessions?.has(this.activeSessionId);
if (echoPassthrough && (data === '\r' || data === '\x03')) {
this._echoPassthroughSessions.delete(this.activeSessionId);
}
// ── Local Echo Mode ──
// When enabled, keystrokes are buffered locally in the overlay for
// instant visual feedback. Nothing is sent to the PTY until Enter
// (or a control char) is pressed — avoids out-of-order char delivery.
if (this._localEchoEnabled) {
if (this._localEchoEnabled && !echoPassthrough) {
if (data === '\x7f') {
const source = this._localEchoOverlay?.removeChar();
if (source === 'flushed') {
@@ -853,9 +920,16 @@ Object.assign(CodemanApp.prototype, {
}
this._pendingInput += data;
flushInput();
} else if (source === false) {
// Nothing pending, nothing flushed, nothing detected. The
// composer may still hold text the overlay cannot see (buffer
// detection is suppressed after a control-char flush), so
// forward the backspace instead of swallowing it (issue #218);
// an empty composer ignores it.
this._pendingInput += data;
flushInput();
}
// 'pending' = removed unsent text (no PTY backspace needed)
// false = nothing to remove (swallow the backspace)
return;
}
if (/^[\r\n]+$/.test(data)) {
@@ -897,6 +971,41 @@ Object.assign(CodemanApp.prototype, {
// Single-byte ESC (user pressing Escape) still falls through to
// the control char handler below.
if (data.length > 1 && data.charCodeAt(0) === 27) {
// Bracketed paste (terminal.paste() while DECSET 2004 is on):
// flush typed-but-unsent overlay text FIRST so the pasted block
// lands after it in the composer, not before it (issue #219).
// The paste sequence gets its own delayed write: Codex's
// paste-burst handling drops keystrokes that arrive in the SAME
// PTY read as a bracketed paste (verified against codex 0.147),
// mirroring the delayed \r in the Enter branch above.
if (data.startsWith(window.CodemanTerminalInput.BRACKETED_PASTE_START)) {
const hadPending = !!this._localEchoOverlay?.pendingText;
this._flushLocalEchoPending();
if (hadPending) {
flushInput();
setTimeout(() => {
this._pendingInput += data;
flushInput();
}, 80);
} else {
this._pendingInput += data;
flushInput();
}
return;
}
// Composer nav keys (arrows, Home/End, Delete, PgUp/PgDn):
// flush unsent text so the key edits the real composer state,
// then hand the session to plain PTY echo until Enter/Ctrl+C.
// The cursor may now sit mid-text, where append-only buffering
// cannot track edits (issue #218).
if (window.CodemanTerminalInput.isComposerNavKey(data)) {
this._flushLocalEchoPending();
if (!this._echoPassthroughSessions) this._echoPassthroughSessions = new Set();
this._echoPassthroughSessions.add(this.activeSessionId);
this._pendingInput += data;
flushInput();
return;
}
// Multi-byte escape sequence — forward to PTY without clearing
// overlay/flushed state (terminal response, not user input)
this._pendingInput += data;
@@ -2079,6 +2188,54 @@ Object.assign(CodemanApp.prototype, {
if (this.terminal?.buffer?.active?.viewportY === 0) this._maybeRefetchFullHistory?.();
},
/**
* Rows a `?full=1` capture will occupy once written into xterm.
*
* tmux joins wrapped rows in that capture (`capture-pane -J`), so a long
* logical line re-wraps into several xterm rows on write and a bare newline
* count would undershoot; escape sequences occupy no cells and come out
* first. Approximate by construction (it ignores double-width glyphs), which
* is fine: the only consumer is a coarse size comparison
* (_replayWouldShrinkBuffer), and it runs once per cooldown-guarded re-pull.
*/
_estimateReplayRows(text, cols) {
if (typeof text !== 'string' || !text) return 0;
const width = cols > 0 ? cols : 80;
const plain = text.replace(window.CodemanTerminalInput.REPLAY_ESCAPE_RE, '');
let rows = 0;
for (const line of plain.split('\n')) {
const cells = line.endsWith('\r') ? line.length - 1 : line.length;
rows += cells > width ? Math.ceil(cells / width) : 1;
}
return rows;
},
/**
* DOWNGRADE GUARD for the scroll-to-top re-pull (issue #205, round 2).
*
* `_maybeRefetchFullHistory` resets the terminal and rewrites it from the
* capture, which is a straight win when tmux holds more than the browser —
* the burst-repaint and tab-switch losses it was built for. But a repaint-mode
* CLI pane keeps NO tmux history of its own (`history_size≈0` measured for a
* Claude pane), so there the capture is roughly ONE frame while xterm may hold
* hundreds of rows of replayed frames. Rewriting then DESTROYS history
* mid-scroll: exactly the "goes back a limited amount, repeats blocks, gets
* worse when I reach the top" report from the 1.12.0 retest.
*
* So refuse when the capture is smaller, with a one-screen tolerance because
* both sides are estimates: `buffer.active.length` includes the blank rows
* below the last line, and _estimateReplayRows can only approximate wrapping.
* Only a capture that is worse by more than a full screen counts as a
* downgrade, which leaves every genuine recovery case untouched.
*/
_replayWouldShrinkBuffer(capture) {
const term = this.terminal;
const rowsNow = term?.buffer?.active?.length || 0;
if (!rowsNow) return false;
const screen = term?.rows || 24;
return this._estimateReplayRows(capture, term?.cols) + screen < rowsNow;
},
/**
* Ease-out smooth scrolling for the local wheel path. The capture-phase
* wheel handler owns local scrolling (xterm's own smooth scroller is
@@ -2192,17 +2349,26 @@ Object.assign(CodemanApp.prototype, {
// Accumulate raw data (may contain DEC 2026 markers)
this.pendingWrites.push(data);
this._scheduleTerminalWriteFlush();
},
if (!this.writeFrameScheduled) {
this.writeFrameScheduled = true;
this._safeYield(() => {
// xterm.js 6.0 handles DEC 2026 sync markers natively — it buffers
// content between 2026h/2026l and renders atomically. No need for
// client-side incomplete-block detection; just flush every frame.
this.flushPendingWrites();
this.writeFrameScheduled = false;
});
}
/**
* Schedule one render-budgeted terminal flush.
*
* Clear the scheduled flag before flushing so flushPendingWrites() can queue
* another yield when a large final batch leaves bytes behind. Keeping the
* flag set through the flush stranded that remainder until unrelated output
* arrived, which looked like truncated responses and idle shell commands.
*/
_scheduleTerminalWriteFlush() {
if (this.writeFrameScheduled || this.pendingWrites.length === 0) return;
this.writeFrameScheduled = true;
this._safeYield(() => {
this.writeFrameScheduled = false;
// xterm.js 6.0 handles DEC 2026 sync markers natively — it buffers
// content between 2026h/2026l and renders atomically.
this.flushPendingWrites();
});
},
/**
@@ -2218,13 +2384,24 @@ Object.assign(CodemanApp.prototype, {
this.flickerFilterActive = false;
// Trigger a normal flush
if (!this.writeFrameScheduled) {
this.writeFrameScheduled = true;
this._safeYield(() => {
this.flushPendingWrites();
this.writeFrameScheduled = false;
});
}
this._scheduleTerminalWriteFlush();
},
/**
* Flush the local-echo overlay's unsent text into `_pendingInput` (no
* trailing Enter) and reset overlay + flushed-state tracking. Used before
* forwarding sequences that must arrive AFTER the typed text (bracketed
* paste, composer nav keys). The caller forwards its own sequence: nav keys
* ride the same write, pastes get a delayed second write because codex
* drops keys that share a PTY read with a bracketed paste.
*/
_flushLocalEchoPending() {
const text = this._localEchoOverlay?.pendingText || '';
this._localEchoOverlay?.clear();
this._localEchoOverlay?.suppressBufferDetection();
this._flushedOffsets?.delete(this.activeSessionId);
this._flushedTexts?.delete(this.activeSessionId);
if (text) this._pendingInput += text;
},
/**
@@ -2266,8 +2443,14 @@ Object.assign(CodemanApp.prototype, {
}
},
});
} else if (session.mode === 'shell') {
} else if (session.mode === 'shell' || session.mode === 'codex') {
// Shell mode: the shell provides its own PTY echo so the overlay isn't needed.
// Codex mode: the composer is fully interactive per keystroke. Typing
// "/" pops a live-filtering command picker (issue #222), the composer
// grows and rewraps as it fills (#220), pastes are bracketed (#219)
// and arrows/history edit server-side state (#218). Buffering
// keystrokes until Enter starves all of that, so codex sessions use
// plain PTY echo like shell.
// Disable it by clearing any pending text.
this._localEchoOverlay.clear();
this._localEchoEnabled = false;
@@ -2350,13 +2533,7 @@ Object.assign(CodemanApp.prototype, {
this.terminal.write(joined.slice(0, MAX_FRAME_BYTES));
this.pendingWrites.push(joined.slice(MAX_FRAME_BYTES));
deferred = true;
if (!this.writeFrameScheduled) {
this.writeFrameScheduled = true;
this._safeYield(() => {
this.flushPendingWrites();
this.writeFrameScheduled = false;
});
}
this._scheduleTerminalWriteFlush();
}
if (
preserveViewportY !== null &&
@@ -2671,7 +2848,11 @@ Object.assign(CodemanApp.prototype, {
/** Insert editable text at the active prompt without pressing Enter. */
insertTerminalText(text) {
if (!this.activeSessionId || !text) return;
if (this._localEchoEnabled && this._localEchoOverlay) {
if (
this._localEchoEnabled &&
this._localEchoOverlay &&
!this._echoPassthroughSessions?.has(this.activeSessionId)
) {
this._localEchoOverlay.appendText(text);
} else {
this.sendInput(text).catch(() => {});
@@ -2926,11 +3107,20 @@ Object.assign(CodemanApp.prototype, {
// Wheel forwarding gate for the container wheel handler: no Shift override,
// xterm's own encoder dormant, viewport at the bottom, and a TUI VERIFIED to
// scroll its transcript on SGR wheel reports: codex, or claude 2.1.187+
// (older Claude Code captures wheel as select-menu option navigation; an
// unknown version is treated as older). Gemini is a strip mode too but its
// wheel behavior is unverified, so it keeps the local wheel — taps/clicks
// are still forwarded for it (harmless no-ops at worst).
// scroll its transcript on SGR wheel reports — which today is claude 2.1.187+
// and nothing else (older Claude Code captures wheel as select-menu option
// navigation; an unknown version is treated as older). Gemini and codex are
// strip modes too but keep the local wheel — taps/clicks are still forwarded
// for them (harmless no-ops at worst).
//
// Codex USED to forward here and was the #227 regression (DodgyBadger, Codex
// latest / Chrome / Win11: dead wheel in codex, working scrollbar drag).
// Measured on codex-cli 0.147.0 in a bare tmux: it never enables mouse
// tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change
// NOTHING on screen — it runs an inline viewport (`alternate_on=0`) and pushes
// its transcript into the terminal's own scrollback (tmux `history_size`
// grows), so there is no in-app pager to drive and local scrollback IS the
// codex transcript. Forwarding therefore swallowed every tick.
// Wheel delta → whole scroll lines. macOS trackpads turn Shift+two-finger
// scroll into a HORIZONTAL wheel (deltaY≈0, deltaX carries the magnitude), and
// Shift routes the wheel to local scrollback (_shouldForwardWheelToApp returns
@@ -2969,16 +3159,23 @@ Object.assign(CodemanApp.prototype, {
// plain wheel to xterm's own scrollback like pre-#144, for users who prefer
// it over forwarding the wheel to the CLI's transcript (issue #154). Cheap —
// loadAppSettingsFromStorage() is cache-backed.
//
// FOOTGUN, and why it is handled downstream rather than here: for a
// repaint-mode CLI that local scrollback is EMPTY (tmux keeps no history for
// the pane), so this setting can silently convert a working wheel into a
// dead one — a plausible reading of the #205 retest, where a user whose
// scrolling was broken on 1.11.x may well have flipped it while hunting for
// a fix. Scoping the setting away from those modes would be the other
// option, but it would override an explicit user choice; instead the caller
// falls through to _maybePageCliTranscript, so the gesture still pages the
// CLI's transcript and the setting keeps meaning exactly what it says.
if (this.loadAppSettingsFromStorage?.()?.terminalWheelLocalScrollback) return false;
const mode = this.terminal?.modes?.mouseTrackingMode;
if (mode && mode !== 'none') return false;
const session = this.sessions?.get(this.activeSessionId);
const sessionMode = session?.mode || 'claude';
if (sessionMode === 'claude') {
if (!this._cliVersionAtLeast(session?.cliVersion, '2.1.187')) return false;
} else if (sessionMode !== 'codex') {
return false;
}
if (sessionMode !== 'claude') return false;
if (!this._cliVersionAtLeast(session?.cliVersion, '2.1.187')) return false;
// Deliberately NOT gated on _terminalViewportAtBottom(). It used to be, so
// that leaving the bottom handed the wheel back to local scrollback and both
// histories stayed reachable without a mode switch. In practice that inverted
@@ -3012,13 +3209,105 @@ Object.assign(CodemanApp.prototype, {
if (!pos) return;
const btn = lines < 0 ? 64 : 65;
const ticks = Math.min(Math.abs(lines), 5);
this._queueScrollBytes(`\x1b[<${btn};${pos.col};${pos.row}M`.repeat(ticks));
},
/**
* Shared 40ms coalescer for every byte a scroll gesture sends to the PTY (SGR
* wheel reports and the PageUp/PageDown fallback alike). Each flush becomes a
* tmux send-keys server-side, so per-event writes would spawn a process storm
* on a single flick; the queue is bounded so a wild scroll can't build a
* backlog that keeps scrolling after the finger stops.
*/
_queueScrollBytes(data) {
if (!data || !this.activeSessionId) return;
const queued = this._wheelSgrQueue || '';
if (queued.length > 512) return;
this._wheelSgrQueue = queued + `\x1b[<${btn};${pos.col};${pos.row}M`.repeat(ticks);
this._wheelSgrQueue = queued + data;
if (this._wheelSgrFlushTimer) return;
this._wheelSgrFlushTimer = setTimeout(() => this._flushWheelSgrQueue(), 40);
},
/**
* True when this session's LOCAL scrollback is structurally empty: a Claude
* pane in repaint mode, where tmux reports `history_size≈0` and every frame
* overwrites the last, so xterm's normal buffer never grows past one screen
* (`baseY === 0`). Scrolling that buffer is a no-op no matter how the gesture
* is routed — the "wheel does nothing at all" half of the #205 retest.
*/
_localScrollbackIsHollow() {
const mode = this.sessions?.get(this.activeSessionId)?.mode || 'claude';
if (mode !== 'claude') return false;
const buf = this.terminal?.buffer?.active;
if (!buf || buf.type === 'alternate') return false;
return (buf.baseY || 0) === 0;
},
/**
* LAST-RESORT scroll for a hollow local buffer: translate gesture lines into
* coalesced PageUp/PageDown key sends so the CLI pages its OWN transcript.
*
* The rescue path for every way `_shouldForwardWheelToApp` can come back false
* on a Claude session that has no local history to fall back on: the CLI
* version probe failed or is genuinely older than 2.1.187, or the user turned
* on "Wheel scrolls local history" (which pins the wheel to a buffer that,
* for a repaint-mode CLI, is empty — the setting's footgun). Before this, all
* of those produced a completely dead gesture; the #205 reporter proved the
* keyboard route works by paging back through intact text with Fn+Up.
*
* Triple-guarded (claude mode + gate false + `baseY === 0`), so a session with
* real local scrollback is never touched. Shift is excluded on purpose: it is
* the explicit "give me local scrollback" gesture and must keep that meaning.
*
* @returns true when the gesture was consumed here (the caller must not also
* scroll locally).
*/
_maybePageCliTranscript(ev, lines) {
if (!lines || ev?.shiftKey || !this.activeSessionId) return false;
if (!this._localScrollbackIsHollow()) return false;
// Leftover travel belongs to the tab it was made on.
if (this._pageKeySession !== this.activeSessionId) {
this._pageKeySession = this.activeSessionId;
this._pageKeyPending = 0;
}
const tuning = window.CodemanTerminalInput;
const perPage = Math.max(2, Math.round((this.terminal?.rows || 24) * tuning.PAGE_KEY_SCREEN_FRACTION));
const pending = (this._pageKeyPending || 0) + lines;
const pages = Math.trunc(pending / perPage);
this._pageKeyPending = pending - pages * perPage;
if (pages) {
const key = pages < 0 ? tuning.KEY_PAGE_UP : tuning.KEY_PAGE_DOWN;
this._queueScrollBytes(key.repeat(Math.min(Math.abs(pages), tuning.PAGE_KEY_MAX_PER_BATCH)));
}
this._logScrollRouting('page-keys');
return true;
},
/**
* One line in the console saying WHY a scroll gesture went where it went.
*
* Issue #205 ran two rounds of remote guesswork — is the CLI version probe
* empty, is the opt-out setting on, did a mouse DECSET leak past the strip? —
* that this single log answers directly. Logged once per session per distinct
* decision, so a steady gesture stays silent and a CHANGE (e.g. the version
* arriving late and flipping the route) still prints.
*/
_logScrollRouting(decision) {
const sessionId = this.activeSessionId || '(none)';
const session = this.sessions?.get(sessionId);
const optOut = !!this.loadAppSettingsFromStorage?.()?.terminalWheelLocalScrollback;
const tracking = this.terminal?.modes?.mouseTrackingMode || 'none';
const baseY = this.terminal?.buffer?.active?.baseY ?? -1;
const signature = `${decision}|${session?.mode}|${session?.cliVersion}|${optOut}|${tracking}|${baseY > 0}`;
if (!this._scrollRoutingLogged) this._scrollRoutingLogged = new Map();
if (this._scrollRoutingLogged.get(sessionId) === signature) return;
this._scrollRoutingLogged.set(sessionId, signature);
console.log(
`[scroll] ${sessionId} → ${decision} (mode=${session?.mode || '?'}, cliVersion=${session?.cliVersion || 'unknown'}, ` +
`localScrollbackOptOut=${optOut}, mouseTracking=${tracking}, localScrollbackRows=${baseY})`
);
},
_flushWheelSgrQueue() {
this._wheelSgrFlushTimer = null;
const data = this._wheelSgrQueue;
+1 -2
View File
@@ -97,8 +97,7 @@ Object.assign(CodemanApp.prototype, {
<span class="tab-name">${escapeHtml(webview.name)}</span>
</span>
</span>
<span class="tab-gear" onclick="event.stopPropagation(); app.showWebviewModal(${jsonId})" title="URL settings" aria-label="URL settings" tabindex="0">&#x2699;</span>
<span class="tab-close" onclick="event.stopPropagation(); app.closeWebviewTab(${jsonId})" title="Close tab" aria-label="Close web tab" tabindex="0">&times;</span>
<span class="tab-actions"><span class="tab-gear" onclick="event.stopPropagation(); app.showWebviewModal(${jsonId})" title="URL settings" aria-label="URL settings" tabindex="0">&#x2699;</span><span class="tab-close" onclick="event.stopPropagation(); app.closeWebviewTab(${jsonId})" title="Close tab" aria-label="Close web tab" tabindex="0">&times;</span></span>
</div>`);
idx++;
}
+22
View File
@@ -640,6 +640,25 @@ async function buildExternalAttachmentRouteItem(
}
}
/**
* Headers Fastify already put on the reply, in a shape `writeHead` accepts.
*
* `reply.raw.writeHead()` writes straight to the Node response and bypasses
* Fastify's header store, so anything the security `onRequest` hook granted — CORS
* for localhost origins, nosniff, frame-options, CSP — is silently dropped on every
* route that answers this way. Spread this first and let the route's own headers
* win over it.
*/
function inheritedHeaders(reply: {
getHeaders(): NodeJS.Dict<number | string | string[]>;
}): Record<string, number | string | string[]> {
const out: Record<string, number | string | string[]> = {};
for (const [name, value] of Object.entries(reply.getHeaders())) {
if (value !== undefined) out[name] = value;
}
return out;
}
export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & EventPort & ConfigPort): void {
// Lazy filesystem listing for the Link Existing and mobile input path pickers.
app.get('/api/filesystem/browse', async (req, reply): Promise<ApiResponse<FilesystemBrowseData>> => {
@@ -1337,6 +1356,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
const basename = rawBasename.replace(/["\\\r\n]/g, '_');
if (download === 'true' || ext === 'svg') {
reply.raw.writeHead(200, {
...inheritedHeaders(reply),
'Content-Type': ext === 'svg' ? 'application/octet-stream' : mimeTypes[ext] || 'application/octet-stream',
'Content-Disposition': `attachment; filename="${basename}"`,
'Content-Length': content.length,
@@ -1576,6 +1596,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
// Set up SSE headers
reply.raw.writeHead(200, {
...inheritedHeaders(reply),
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
Connection: 'keep-alive',
@@ -1709,6 +1730,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
const content = await fs.readFile(resolvedPath);
// Bypass Fastify compression — write directly to raw response
reply.raw.writeHead(200, {
...inheritedHeaders(reply),
'Content-Type': mimeTypes[ext] || 'application/octet-stream',
'Content-Disposition': `attachment; filename="${filename}"`,
'Content-Length': content.length,
+22
View File
@@ -10,6 +10,7 @@ import { HookEventSchema, isValidWorkingDir } from '../schemas.js';
import { sanitizeHookData, parseBody } from '../route-helpers.js';
import { persistDockerCaseClaudeSessionId } from '../../docker-hosts.js';
import { getDataDir } from '../../config/instance.js';
import { sessionWaits, hooksAvailableForMode } from '../session-wait-registry.js';
import type { SessionPort, EventPort, RespawnPort, ConfigPort, InfraPort } from '../ports/index.js';
export function registerHookEventRoutes(
@@ -22,6 +23,27 @@ export function registerHookEventRoutes(
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Session not found');
}
// Wake anything blocked on `GET /api/sessions/:id/wait`. Hooks are the only
// DEFINITIVE signals Codeman gets (`idle` is inferred from output stabilization
// and can flap mid-turn), so these two are what an orchestrating agent should
// wait on.
//
// Gated on the session's MODE, matching `resolveWaitSignals` on the read side.
// Without it the guard is one-sided: a caller cannot ASK for `stop` on a shell or
// codex session, but this endpoint would happily deliver one for it. Hook events
// carry no identity beyond a per-instance secret shared by every case, so this is
// also the cheap half of the forgery surface — a `stop` claimed for a session that
// could never legitimately emit one is now dropped instead of steering another
// agent's control flow.
const waitSession = ctx.sessions.get(sessionId);
if (waitSession && hooksAvailableForMode(waitSession.mode)) {
if (event === 'stop') {
sessionWaits.notifySignal(sessionId, 'stop');
} else if (event === 'permission_prompt' || event === 'elicitation_dialog') {
sessionWaits.notifySignal(sessionId, 'blocked');
}
}
// Signal the respawn controller based on hook event type
const controller = ctx.respawnControllers.get(sessionId);
if (controller) {
+558 -14
View File
@@ -4,7 +4,8 @@
* auto-clear, auto-compact, image watcher, flicker filter, and logout.
*/
import { FastifyInstance } from 'fastify';
import { FastifyInstance, type FastifyReply } from 'fastify';
import { z } from 'zod';
import { join, dirname, extname, basename } from 'node:path';
import { homedir } from 'node:os';
import { existsSync, statSync, mkdirSync, writeFileSync } from 'node:fs';
@@ -17,6 +18,7 @@ import {
getErrorMessage,
type ApiResponse,
type SessionColor,
type SessionStatus,
type CodexConfig,
type GeminiConfig,
type AntigravityConfig,
@@ -40,8 +42,19 @@ import {
QuickStartSchema,
InteractiveStartSchema,
SessionOrderUpdateSchema,
SessionWaitQuerySchema,
SessionWaitOutputQuerySchema,
} from '../schemas.js';
import { mergeSessionOrder } from '../../session-order.js';
import {
sessionWaits,
resolveWaitSignals,
signalForStatus,
WaitCapacityError,
type WaitSignal,
type SignalWaitResult,
} from '../session-wait-registry.js';
import { clampWaitMs, MAX_BUFFER_SCAN_BYTES } from '../../config/agent-wait.js';
import {
autoConfigureRalph,
canAccessOwned,
@@ -66,6 +79,7 @@ import {
updateCaseModel,
stripCaseEnvKeys,
applyStatusLineConfig,
applyAgentSkill,
refreshStaleCodemanHooks,
} from '../../hooks-config.js';
import { generateClaudeMd } from '../../templates/claude-md.js';
@@ -314,6 +328,234 @@ async function clampExternalCliBypassForOwner(
return { codexConfig: clampedCodex, geminiConfig: clampedGemini, antigravityConfig: clampedAntigravity };
}
// ═══════════════════════════════════════════════════════════════
// Agent wait helpers (shared by GET /wait, GET /wait-output, POST /input)
// ═══════════════════════════════════════════════════════════════
/**
* Validate a wait query WITHOUT throwing away the Zod issue.
*
* `parseBody`'s message argument REPLACES the issue text, so `?timeout=30s` came
* back as a bare "Invalid wait parameters": the caller could not tell which of
* `until`, `timeout` or `fresh` it got wrong, and its only move was to retry with
* a different guess. These endpoints are driven by an LLM with no documentation in
* context — the error message IS the documentation, which is why the signal parser
* one line later goes to the trouble of naming the bad token and listing the valid
* ones. This keeps the endpoint label AND names the offending field.
*/
function parseWaitQuery<T>(schema: z.ZodType<T>, query: unknown, label: string): T {
const result = schema.safeParse(query);
if (result.success) return result.data;
const issue = result.error.issues[0];
const field = issue && issue.path.length > 0 ? issue.path.join('.') : '';
const detail = issue?.message ?? 'validation failed';
const message = field ? `Invalid ${label} parameter '${field}': ${detail}` : `Invalid ${label} parameters: ${detail}`;
throw Object.assign(new Error(message), {
statusCode: 400,
body: createErrorResponse(ApiErrorCode.INVALID_INPUT, message),
});
}
/**
* Map a waiter-cap rejection to the code that tells the caller the truth.
*
* The two caps mean different things and warrant different recovery: `session` is
* genuinely about THIS session, while `owner` and `total` are process-wide budgets
* that say nothing about it. Reporting a global cap as `SESSION_BUSY` (409,
* documented as "Session is busy") sent an agent off to a different session to hit
* the identical error. `RATE_LIMITED` is the code whose whole meaning is "come back
* later", and clients and proxies already treat 429 that way.
*/
function waitCapacityResponse(err: WaitCapacityError): ApiResponse<never> {
const code = err.scope === 'session' ? ApiErrorCode.SESSION_BUSY : ApiErrorCode.RATE_LIMITED;
// The registry's message already names the scope and the limit; passing it through
// verbatim keeps the wording in one place.
return createErrorResponse(code, err.message);
}
/**
* The signal a session is ALREADY emitting, corrected for liveness.
*
* `signalForStatus` alone is not enough here, because `Session` parks a DEAD PTY at
* `_status = 'idle'` (both `onExit` handlers do) and the object survives in the
* session map until an explicit DELETE. Trusting the status therefore answers the
* default wait with `{signal:"idle", immediate:true}` for a worker that has
* crashed — HTTP 200, no error anywhere, and the agent types its next prompt into a
* corpse — while `until=exit` blocks for the full timeout on an event that already
* happened and can never happen again.
*
* `pid === null` means no process is behind this session: it exited, it was
* detached, or it was created and never started. All three are `exit` from a
* caller's point of view — nothing is running — and in all three the agent's
* correct next move is to (re)start the worker rather than to type at it. The
* response still carries the raw `status` alongside, so nothing is hidden.
*
* ⚠️ `pid` alone is NOT enough, and on the normal configuration it is never the
* thing that fires — see `workerIsDead()`. `dead` carries the mux layer's answer.
*
* Fixing it HERE rather than in `signalForStatus` is deliberate: liveness is not
* derivable from `SessionStatus`, and the registry holds no `Session` reference.
*/
function currentSignalFor(session: { pid: number | null; status: SessionStatus }, dead: boolean): WaitSignal | null {
if (dead || session.pid === null || session.pid === undefined) return 'exit';
return signalForStatus(session.status);
}
// ── Worker liveness for tmux-backed sessions ────────────────────────────────
//
// `session.pid` is the LOCAL `tmux attach` client, not the worker. Codeman sets
// `remain-on-exit on` for every session it creates, so when the command inside the
// pane exits, tmux keeps the pane (`pane_dead=1`), the tmux session survives, the
// attach client keeps running and `pid` never goes null — no `exit` event is emitted
// and nothing in `Session` changes. Measured on a shell worker killed with `exit 42`:
// tmux reports `pane_dead=1 status=42` while Codeman reports `pid=309406 status=idle`
// and the DEFAULT wait answers `{signal:"idle", immediate:true}` in 0 ms for a corpse.
// So the liveness check has to ask the mux layer. `pid === null` still matters: it is
// the right (and only) answer for a direct-PTY session, which has no pane to ask about.
//
// Cost control, because `isPaneDead()` is a synchronous `execSync` and `/wait` is
// polled in a loop by design:
// 1. Only mux-backed sessions are probed at all.
// 2. Only requests that actually wait probe — a plain `POST .../input` (the browser's
// hot path, thousands per session) never touches tmux.
// 3. Results are cached per pane for PANE_DEATH_TTL_MS, so a poll loop cannot turn
// into one exec per request.
// 4. The while-blocked watcher is ONE timer per session no matter how many waiters
// are parked on it, and it exists only while at least one of them is.
/** How long a pane-liveness probe is reused. Long enough to absorb a poll loop. */
const PANE_DEATH_TTL_MS = 750;
/** How often a session with a parked waiter is re-checked for a dead worker. */
const PANE_DEATH_POLL_MS = 3_000;
/** Bounded, because a 24h server churns through panes. */
const paneDeathCache = new LRUMap<string, { dead: boolean; at: number }>({ maxSize: 256 });
/** One watcher per pane, refcounted by the waits currently parked on it. */
const paneDeathWatchers = new Map<string, { timer: NodeJS.Timeout; refs: number }>();
type LivenessSession = { usesMux?: boolean; muxName?: string | null };
/**
* Whether the worker inside this session's tmux pane has exited.
*
* False for anything not tmux-backed (nothing to ask), and false when the probe is
* unavailable or throws — an unknown answer must never invent a death.
*/
function workerIsDead(mux: InfraPort['mux'], session: LivenessSession, now: number = Date.now()): boolean {
const muxName = session.usesMux === false ? null : session.muxName;
if (!muxName) return false;
// Defensive: `TerminalMultiplexer` declares it, but route-test doubles may not.
if (typeof mux?.isPaneDead !== 'function') return false;
const cached = paneDeathCache.get(muxName);
if (cached && now - cached.at < PANE_DEATH_TTL_MS) return cached.dead;
let dead = false;
try {
dead = mux.isPaneDead(muxName) === true;
} catch {
dead = false;
}
paneDeathCache.set(muxName, { dead, at: now });
return dead;
}
/**
* Release every waiter on a session whose worker has died, in the documented order.
*
* The same pair the PTY-exit listener and the delete path use, for the same reason:
* `until=exit` callers get their signal, everyone else gets `ended: true` instead of
* burning the rest of their timeout on feeds that will never produce anything.
*/
function releaseWaitersForDeadWorker(sessionId: string): void {
sessionWaits.notifySignal(sessionId, 'exit');
sessionWaits.cancelAll(sessionId);
}
/**
* While a wait is parked on a mux-backed session, poll for the worker dying.
*
* Without this, a worker that dies DURING a wait is invisible: no `exit` event fires
* (the attach client is still alive), no output arrives, and the caller blocks for its
* full timeout — the common orchestration case, "send a prompt and wait", where the
* worker crashes mid-turn.
*
* @returns a release function; call it in a `finally`, or the timer outlives the wait.
*/
function watchForDeadWorker(mux: InfraPort['mux'], session: LivenessSession, sessionId: string): () => void {
const muxName = session.usesMux === false ? null : session.muxName;
if (!muxName || typeof mux?.isPaneDead !== 'function') return () => {};
const existing = paneDeathWatchers.get(muxName);
if (existing) {
existing.refs++;
} else {
const timer = setInterval(() => {
if (!workerIsDead(mux, session)) return;
releaseWaitersForDeadWorker(sessionId);
}, PANE_DEATH_POLL_MS);
// Auxiliary to the waiter's own timer, which is deliberately NOT unref'd; this one
// must never be the reason the process stays up.
timer.unref();
paneDeathWatchers.set(muxName, { timer, refs: 1 });
}
let released = false;
return () => {
if (released) return;
released = true;
const entry = paneDeathWatchers.get(muxName);
if (!entry) return;
entry.refs--;
if (entry.refs <= 0) {
clearInterval(entry.timer);
paneDeathWatchers.delete(muxName);
}
};
}
/** Test seam: pane-liveness state is module-level, so a suite must be able to reset it. */
export function _resetPaneLivenessState(): void {
for (const entry of paneDeathWatchers.values()) clearInterval(entry.timer);
paneDeathWatchers.clear();
paneDeathCache.clear();
}
/** Test seam: how many panes are currently being watched for a dead worker. */
export function _paneDeathWatcherCount(): number {
return paneDeathWatchers.size;
}
/**
* An `AbortController` that fires when the CLIENT goes away, and only then.
*
* Freeing an abandoned waiter matters because the documented pattern is a loop of
* short waits: `curl --max-time 30 ".../wait?timeout=600000"` abandons a live waiter
* every iteration until the cap is hit and an innocent session reports busy. Same for
* any proxy that cuts the connection.
*
* ⚠️ **It must listen on the RESPONSE, not the request.** `req.raw` emits `'close'`
* as soon as the request body has finished streaming, which on a POST is BEFORE the
* handler ever blocks — measured at +1ms with `aborted: false`, indistinguishable
* from a real hang-up at +0ms. Wiring the abort there cancels every send-and-wait
* instantly and silently kills the feature (it survives on GET only because a GET has
* no body to finish). `reply.raw` emits `'close'` both when the response completes
* and when the socket dies, and `writableFinished` is what tells those apart: true
* only if the response actually went out. The guard is load-bearing, not defensive.
*
* `app.inject()` never emits `'close'` at all, so this is only observable over real
* HTTP — which is why the regression test for it binds a port.
*/
function abortOnClientHangUp(reply: FastifyReply): AbortController {
const controller = new AbortController();
reply.raw.on('close', () => {
if (!reply.raw.writableFinished) controller.abort();
});
return controller;
}
export function registerSessionRoutes(
app: FastifyInstance,
ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort
@@ -458,6 +700,13 @@ export function registerSessionRoutes(
// cases (writeHooksConfig already wrote the secret) and for non-Codeman/absent hooks.
if ((body.mode ?? 'claude') === 'claude') {
await refreshStaleCodemanHooks(workingDir).catch(() => {});
// Agent skill (docs/agent-control-plan.md §2): ADD-ONLY on create, same shared-
// .claude rationale as the statusLine above: a create must never remove the
// skill from under other live sessions in the repo. Marker-guarded, so a
// user's own skills/codeman is never touched.
if (await ctx.getAgentSkillEnabled()) {
await applyAgentSkill(workingDir, true).catch(() => {});
}
}
// Check OpenCode availability if requested
@@ -863,9 +1112,9 @@ export function registerSessionRoutes(
// ========== Send Input ==========
app.post('/api/sessions/:id/input', async (req) => {
app.post('/api/sessions/:id/input', async (req, reply) => {
const { id } = req.params as { id: string };
const { input, useMux, seq, clientId } = parseBody(SessionInputWithLimitSchema, req.body);
const { input, useMux, seq, clientId, wait, waitTimeout } = parseBody(SessionInputWithLimitSchema, req.body);
const session = findSessionOrFail(ctx, id, req);
const inputStr = String(input);
@@ -876,33 +1125,319 @@ export function registerSessionRoutes(
);
}
// Send-and-wait (agent orchestration). This has to be ONE endpoint rather than a
// POST followed by GET .../wait: between the write and the session flipping to
// `working` there is a window in which a separate wait sees the session still
// idle and returns instantly, reporting the PREVIOUS turn as this turn's answer.
// Registering the waiter before the write closes that window.
const wantsWait =
wait === true || (typeof wait === 'string' && wait.trim().length > 0) || (Array.isArray(wait) && wait.length > 0);
let until: readonly WaitSignal[] = [];
if (wantsWait) {
const resolved = resolveWaitSignals(wait === true ? undefined : wait, { mode: session.mode });
if (resolved.error) return createErrorResponse(ApiErrorCode.INVALID_INPUT, resolved.error);
until = resolved.until;
}
// Reliable delivery (POST fallback when the WebSocket is down): a 2xx IS the
// client's ACK, so a tagged duplicate redelivery must still return 200 but
// skip the write. Untagged requests (curl/legacy) always apply.
if (typeof clientId === 'string' && typeof seq === 'number' && !session.shouldApplyInput(clientId, seq)) {
const tagged = typeof clientId === 'string' && typeof seq === 'number';
const duplicate = tagged && !session.shouldApplyInput(clientId as string, seq as number);
if (duplicate && !wantsWait) {
return {};
}
// Only a waiting request pays for the tmux probe: the browser's plain input path
// (thousands of calls per session) must stay exec-free.
const workerDead = wantsWait && workerIsDead(ctx.mux, session);
const timeoutMs = clampWaitMs(waitTimeout ?? undefined);
// Same slot leak as the GET routes: a client that gives up mid-wait would
// otherwise hold a waiter for the full timeout. Response-side, always — see
// abortOnClientHangUp: on THIS route a request-side listener fires the moment the
// JSON body finishes streaming and aborts every send-and-wait before it starts.
const abort = abortOnClientHangUp(reply);
let waitPromise: Promise<SignalWaitResult> | null = null;
if (wantsWait) {
try {
waitPromise = sessionWaits.waitForSignal(id, {
until,
timeoutMs,
owner: ownerFor(req),
abortSignal: abort.signal,
// A FRESH delivery must not be satisfied by the state the session is already
// in: it is idle right now, which is precisely why we are typing at it.
// A DUPLICATE has no new turn coming, so it answers from the current state
// instead of blocking for a transition that already happened.
requireTransition: !duplicate,
currentSignal: duplicate ? currentSignalFor(session, workerDead) : undefined,
});
} catch (err) {
if (err instanceof WaitCapacityError) {
// Nothing has been written yet, but `shouldApplyInput` already consumed the
// seq. Give it back or the caller's retry is rejected as a duplicate and the
// input is lost by the very mechanism meant to make delivery reliable.
if (tagged && !duplicate) session.forgetInputSeq(clientId as string, seq as number);
return waitCapacityResponse(err);
}
throw err;
}
}
const stopDeathWatch = wantsWait ? watchForDeadWorker(ctx.mux, session, id) : () => {};
// Write input to PTY. Direct write is synchronous; writeViaMux
// (tmux send-keys) is fire-and-forget to avoid blocking the HTTP response.
if (useMux) {
// Fire-and-forget: don't block HTTP response on tmux child process.
// Fallback to direct write on failure.
//
// Because the response has already been sent by then, a failure there is the
// one case the caller can never learn about — so the dedup bookkeeping is
// rolled back. Otherwise the seq stays recorded as applied and a retry, the
// very mechanism reliable delivery exists for, is rejected as a duplicate.
const undoOnFailure = () => {
if (tagged) session.forgetInputSeq(clientId as string, seq as number);
};
// Whether the bytes actually reached a write path. Only meaningful on the wait
// path (the fire-and-forget branches return before the response is built), and
// reported there instead of the old `!duplicate`: a PTY that has exited fails
// BOTH writes, and telling the caller "delivered, but it timed out" points it at
// the wrong recovery — wait longer, when the truth is "restart the worker".
let delivered = false;
if (duplicate) {
// Redelivery of an already-applied input: skip the write, but still honor the
// wait, since the caller's question ("tell me when this settles") is unanswered.
} else if (useMux && waitPromise) {
// The response is already staying open for the wait, so the tmux write can be
// awaited here. This is the ONE path where a writeViaMux failure is observable.
const ok = await session.writeViaMux(inputStr).catch(() => false);
if (ok) {
delivered = true;
} else {
console.warn(`[Server] writeViaMux failed for session ${id}, falling back to direct write`);
delivered = session.write(inputStr);
if (!delivered) undoOnFailure();
}
} else if (useMux) {
// Fire-and-forget: don't block the HTTP response on a tmux child process.
// Fallback to a direct write on failure. Unchanged from before send-and-wait.
session
.writeViaMux(inputStr)
.then((ok) => {
if (!ok) {
console.warn(`[Server] writeViaMux failed for session ${id}, falling back to direct write`);
session.write(inputStr);
}
if (ok) return;
console.warn(`[Server] writeViaMux failed for session ${id}, falling back to direct write`);
if (!session.write(inputStr)) undoOnFailure();
})
.catch(() => {
session.write(inputStr);
if (!session.write(inputStr)) undoOnFailure();
});
} else {
session.write(inputStr);
// Same rollback. NOT an error response, deliberately: a session can
// legitimately have no PTY yet (created but not started), and callers have
// always been able to write to one without a 4xx.
delivered = session.write(inputStr);
if (!delivered && tagged) {
session.forgetInputSeq(clientId as string, seq as number);
}
}
if (!waitPromise) return {};
try {
// `send-keys` SUCCEEDS against a dead pane — tmux is happy to write into a corpse
// — so a truthful `delivered` cannot come from the write's return value alone.
// This is the case the field exists for: "delivered, but it timed out" tells an
// agent to wait longer when the truth is "restart the worker".
if (delivered && workerDead) {
delivered = false;
// The bytes went nowhere, so the seq must not be recorded as applied or the
// caller's retry against a restarted worker is refused as a duplicate.
if (!duplicate) undoOnFailure();
}
// Nothing was written and nothing will be: no turn is coming, so blocking for the
// full timeout would only delay the caller's real recovery by up to ten minutes.
// Releasing the waiter also hands its slot back immediately.
const selfReleased = !delivered && !duplicate;
if (selfReleased) abort.abort();
const result = await waitPromise;
return {
success: true,
data: {
delivered,
duplicate,
status: session.status,
limitPaused: session.isLimitPaused,
// Identical `wait` object to the two GET endpoints, so one client helper
// reads all three, `timeoutMs` (post-clamp) included.
//
// `aborted` is the CLIENT-facing "you hung up, nobody is reading this", and
// by that definition it is unobservable — which is exactly what the API
// reference promises. The abort above is the server releasing its own waiter
// on a delivery that failed, and the client IS reading this response, so
// reporting `aborted: true` there would break that promise and hand an agent
// a second, contradictory reason for the same outcome. `delivered: false`
// already says what happened; `ended` says the wait was released early.
wait: { ...result, aborted: selfReleased ? false : result.aborted, until: [...until] },
},
};
} finally {
stopDeathWatch();
}
});
// ========== Wait For A Signal (agent orchestration) ==========
//
// A bounded long-poll: block until the session hits one of `until`, then answer.
// This exists because SSE is the only "tell me when" channel Codeman has, and an
// agent driving the API from a shell tool cannot hold a stream and parse events
// inline. See docs/agent-control-plan.md.
//
// A TIMEOUT IS A 200, not an error: callers are expected to loop over short waits
// (proxies such as `tailscale serve` cut idle connections), and turning every poll
// boundary into a 4xx would make that loop indistinguishable from a real failure.
app.get('/api/sessions/:id/wait', async (req, reply) => {
const { id } = req.params as { id: string };
const query = parseWaitQuery(SessionWaitQuerySchema, req.query, 'wait');
const session = findSessionOrFail(ctx, id, req);
// An agent polls this URL in a loop with identical parameters. Any intermediary
// applying heuristic freshness to the 200 would serve the stored `timedOut:true`
// body to the next iteration instantly, turning the loop into a busy spin that
// never observes the signal.
reply.header('Cache-Control', 'no-store');
// Shared with the `wait` field on POST .../input: unknown token is a 400,
// hook-only signals are rejected explicitly but dropped from the default.
const { until, error } = resolveWaitSignals(query.until, { mode: session.mode });
if (error) return createErrorResponse(ApiErrorCode.INVALID_INPUT, error);
// The value actually applied after clamping, echoed below: a caller that asked
// for 30 minutes and silently got 10 could not otherwise tell a poll boundary
// from a wedged worker, and would kill a session that was working fine.
const timeoutMs = clampWaitMs(query.timeout);
// Free the waiter when the caller hangs up; the response can no longer be sent by
// then, so freeing the slot is the entire purpose.
const abort = abortOnClientHangUp(reply);
// A worker that dies while this request is parked emits nothing at all (the tmux
// attach client survives it), so a wait would otherwise run to its full timeout.
const stopDeathWatch = watchForDeadWorker(ctx.mux, session, id);
try {
const result = await sessionWaits.waitForSignal(id, {
until,
timeoutMs,
owner: ownerFor(req),
abortSignal: abort.signal,
requireTransition: query.fresh === '1' || query.fresh === 'true',
// Read BEFORE awaiting: this is the state the caller is asking about.
currentSignal: currentSignalFor(session, workerIsDead(ctx.mux, session)),
});
return {
success: true,
data: {
sessionId: id,
// Post-wait status, so a caller that timed out still learns where things stand.
status: session.status,
// A session paused on a usage limit emits nothing until its reset, so a
// timeout here is expected rather than a stall worth retrying hard.
limitPaused: session.isLimitPaused,
// One shape across all three endpoints, so a single `is_done(resp)` helper
// works against any of them. `result.timeoutMs` is the value actually
// applied after clamping, which is what makes the clamp observable.
wait: { ...result, until: [...until] },
},
};
} catch (err) {
if (err instanceof WaitCapacityError) return waitCapacityResponse(err);
throw err;
} finally {
stopDeathWatch();
}
});
// ========== Wait For Output (agent orchestration) ==========
//
// The companion to /wait: block until a literal string appears in this session's
// output. Same 200-on-timeout contract. Fed by the `terminal` listener in
// session-listener-wiring.ts, so what this scans is byte-for-byte what the pane
// printed, ANSI stripped.
//
// ⚠️ A tmux repaint replays text already on screen, so `from=now` can match
// something printed before the request. Callers need a marker unique per call.
app.get('/api/sessions/:id/wait-output', async (req, reply) => {
const { id } = req.params as { id: string };
// Reject `regex` loudly instead of ignoring it. Matching is deliberately literal
// (no ReDoS surface on a caller-supplied pattern over a live stream), and an agent
// that assumed otherwise would silently wait on the wrong thing.
if (req.query && typeof req.query === 'object' && 'regex' in req.query) {
return createErrorResponse(
ApiErrorCode.INVALID_INPUT,
'regex is not supported; use match=<literal substring> (optionally with nocase=1)'
);
}
const query = parseWaitQuery(SessionWaitOutputQuerySchema, req.query, 'wait-output');
const session = findSessionOrFail(ctx, id, req);
// Same reason as /wait: this URL is polled in a loop with identical parameters.
reply.header('Cache-Control', 'no-store');
const timeoutMs = clampWaitMs(query.timeout);
const abort = abortOnClientHangUp(reply);
const owner = ownerFor(req);
// Output waiters are the ones a dead worker strands hardest: the feed simply stops.
const stopDeathWatch = watchForDeadWorker(ctx.mux, session, id);
try {
// Check the cap BEFORE touching the buffer. `session.terminalBuffer` is
// `BufferAccumulator.value`, which joins the WHOLE accumulator (up to 32MB)
// before the slice below takes its tail — so a request that is going to be
// rejected anyway must not pay for a full materialization first, or the cap
// provides no backpressure at all against a `from=buffer` loop.
sessionWaits.assertCapacity(id, owner);
// `from=buffer` scans what already scrolled past before blocking. Bounded to a
// tail: the buffer runs to 32MB and this is a per-request ANSI strip.
let initialText: string | undefined;
if (query.from === 'buffer') {
const buffer = session.terminalBuffer;
initialText =
buffer.length > MAX_BUFFER_SCAN_BYTES ? buffer.slice(buffer.length - MAX_BUFFER_SCAN_BYTES) : buffer;
}
const result = await sessionWaits.waitForOutput(id, {
match: query.match,
nocase: query.nocase === '1' || query.nocase === 'true',
timeoutMs,
owner,
abortSignal: abort.signal,
initialText,
});
return {
success: true,
data: {
sessionId: id,
status: session.status,
limitPaused: session.isLimitPaused,
// Same envelope as /wait; this one carries `matched`/`snippet`/`match`
// where the signal wait carries `signal`/`until`.
wait: { ...result, match: query.match },
},
};
} catch (err) {
if (err instanceof WaitCapacityError) return waitCapacityResponse(err);
throw err;
} finally {
stopDeathWatch();
}
return {};
});
// ========== Send Named Key (tmux send-keys -H) ==========
@@ -2239,6 +2774,15 @@ export function registerSessionRoutes(
await refreshStaleCodemanHooks(resolvedCasePath).catch(() => {});
}
// Agent skill injection (docs/agent-control-plan.md §2): ADD-ONLY on create,
// marker-guarded (a user's own skills/codeman is never touched). Claude mode only
// (`.claude/skills/` is a Claude Code surface); skipped for remote cases, whose
// casePath lives on another host. Docker cases qualify: hostWorkspacePath is a
// real host dir and the skill crosses the bind mount like the rest of `.claude/`.
if (!remote && mode === 'claude' && (await ctx.getAgentSkillEnabled())) {
await applyAgentSkill(resolvedCasePath, true).catch(() => {});
}
// Docker cases: the workspace is a REAL host dir bind-mounted into the container.
// Scaffold hooks (+ a CLAUDE.md) if MISSING so in-container permission prompts and
// hook-idle detection fire (decision: wire hooks now). Never clobbers an existing
+8 -2
View File
@@ -180,13 +180,19 @@ export function registerWsRoutes(app: FastifyInstance, ctx: SessionPort, getHost
const cid = typeof msg.cid === 'string' ? msg.cid : null;
const seq = Number.isInteger(msg.seq) ? (msg.seq as number) : null;
const apply = cid && seq !== null ? session.shouldApplyInput(cid, seq) : true;
let delivered = true;
if (apply) {
// Typed input from a claim-holding desktop keeps the claim "hot"
// and re-asserts the desktop layout after a mobile override.
if (holdsDesktopClaim) session.noteDesktopActivity();
session.write(msg.d);
delivered = session.write(msg.d);
// A session whose PTY is gone swallows the write. ACKing anyway told
// the client to drop the frame from its durable queue and left the seq
// burnt, so the retry that reliable delivery exists for was rejected as
// a duplicate: the input was lost for good.
if (!delivered && cid && seq !== null) session.forgetInputSeq(cid, seq);
}
if (seq !== null && socket.readyState === 1) {
if (delivered && seq !== null && socket.readyState === 1) {
socket.send(`{"t":"ia","seq":${seq}}`);
}
} else if (
+69
View File
@@ -17,6 +17,7 @@ import {
MIN_TERMINAL_SCROLLBACK_LINES,
} from '../config/terminal-history.js';
import { MAX_EDITABLE_BYTES } from '../config/file-editing.js';
import { MIN_MATCH_LENGTH, MAX_MATCH_LENGTH } from '../config/agent-wait.js';
// ========== Path Validation ==========
@@ -759,6 +760,14 @@ export const SettingsUpdateSchema = z
/** Floating ultracode run windows w/ tab connector lines (default OFF). Also starts workflowRunWatcher. SYNCED. */
ultracodeFloatingWindows: z.boolean().optional(),
imageWatcherEnabled: z.boolean().optional(),
/**
* Inject the Codeman agent skill (`skills/codeman`) into `<case>/.claude/skills/`
* on Claude session create, so an agent inside the session can drive the API
* (see docs/agent-control-plan.md §2). SYNCED, default OFF: every skill's
* name+description costs context on every turn, so it is opt-in. Injection is
* add-only at create; a marker keeps user-authored copies untouched.
*/
agentSkillEnabled: z.boolean().optional(),
tunnelEnabled: z.boolean().optional(),
// Action field (NOT persisted): explicit per-request acknowledgment that the
// operator accepts exposing an UNAUTHENTICATED public tunnel (no CODEMAN_PASSWORD).
@@ -911,6 +920,66 @@ export const SessionInputWithLimitSchema = z.object({
// unset rather than sending null. See docs/reliable-input-delivery.md.
seq: z.number().int().nonnegative().optional(),
clientId: z.string().max(128).optional(),
// Send-and-wait (agent orchestration): `true` for the default signal set, or the
// same grammar as `GET .../wait` — a comma string or an array of signals. Absent
// means the historical fire-and-forget behavior, byte for byte.
//
// `.nullish()`, not `.optional()`: a third-party caller building the body with
// JSON.stringify keeps an explicit null on the wire, and `.optional()` rejects it
// with INVALID_INPUT. That gotcha has shipped as a real bug twice.
wait: z.union([z.boolean(), z.string().max(120), z.array(z.string().max(120)).max(8)]).nullish(),
// Unbounded above: the effective value is clamped to MAX_WAIT_MS server-side and
// returned as `data.wait.timeoutMs`, so a caller that asks for 24h sees what it
// actually got. A `.max()` here would turn the same documented clamp into a 400 for
// large-enough guesses, which is the one behaviour an agent cannot predict.
waitTimeout: z.number().int().positive().nullish(),
});
/**
* Query validation for `GET /api/sessions/:id/wait` (agent wait primitives).
*
* Everything arrives as a string. `timeout` is coerced and bounded here, then
* clamped again to the operator's ceiling by `clampWaitMs()` — the schema bound
* only keeps an absurd number out of the arithmetic. A non-numeric `timeout` is a
* 400 rather than a silent fallback, so an agent never believes it asked for a
* longer wait than it got; the value actually applied comes back as
* `data.wait.timeoutMs`, which is what makes the clamp observable. `until` is
* parsed by `parseWaitSignals()`, which reports unknown tokens instead of
* dropping them.
*
* `until` accepts an ARRAY as well as the comma string: `?until=stop&until=exit`
* is how most HTTP clients express a list, Fastify's query parser delivers a
* repeated parameter as an array, and `parseWaitSignals()` has always handled
* both. Rejecting the repeated form left that branch unreachable and 400'd the
* more natural spelling.
*/
export const SessionWaitQuerySchema = z.object({
until: z.union([z.string().max(120), z.array(z.string().max(120)).max(8)]).optional(),
// No upper bound on purpose. The contract is "clamped to [MIN_WAIT_MS, MAX_WAIT_MS]",
// and a `.max()` here contradicted it: `timeout=99999999` was a 400 mid-fan-out while
// `timeout=600001` was silently clamped, so the same documented rule produced two
// different outcomes depending on how big the caller's guess was. `clampWaitMs()`
// bounds every finite value, and `.int()` still rejects `Infinity`/`1e999` and junk.
timeout: z.coerce.number().int().positive().optional(),
fresh: z.enum(['0', '1', 'true', 'false']).optional(),
});
/**
* Query validation for `GET /api/sessions/:id/wait-output`.
*
* `match` is a LITERAL substring, never a pattern: `search-service.ts` avoids regex
* so there is no ReDoS surface, and this endpoint is more exposed still (the pattern
* would be caller-supplied and the input is a live stream). The length bound is a
* second reason the carry buffer stays small. The route separately rejects a `regex`
* parameter outright rather than ignoring it.
*/
export const SessionWaitOutputQuerySchema = z.object({
match: z.string().min(MIN_MATCH_LENGTH).max(MAX_MATCH_LENGTH),
nocase: z.enum(['0', '1', 'true', 'false']).optional(),
from: z.enum(['now', 'buffer']).optional(),
// Unbounded above for the same reason as SessionWaitQuerySchema.timeout: clamping is
// the documented contract, so a large value must clamp rather than 400.
timeout: z.coerce.number().int().positive().optional(),
});
// ========== Session Mutation Routes ==========
+4 -4
View File
@@ -30,6 +30,7 @@ import { homedir, tmpdir } from 'node:os';
import { randomUUID } from 'node:crypto';
import { createRequire } from 'node:module';
import { dataPath } from '../config/instance.js';
import { LAUNCHD_LABEL, SYSTEMD_UNIT } from '../config/service-names.js';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import type {
InstallInfo,
@@ -43,10 +44,9 @@ import type {
const require = createRequire(import.meta.url);
const { version: APP_VERSION } = require('../../package.json') as { version: string };
/** systemd unit name (matches install.sh + scripts/codeman-web.service). */
const SYSTEMD_UNIT = 'codeman-web.service';
/** launchd agent label (matches install.sh setup_launchd_service). */
const LAUNCHD_LABEL = 'com.codeman.web';
// Unit name / job label live in config/service-names.ts so install.sh, this
// detector and `codeman service install` cannot drift apart. Unchanged for the
// default instance.
/** Path to the persisted update status file. */
const STATUS_FILE = dataPath('update-status.json');
/** Network/git timeout for the "check" path (longer than EXEC_TIMEOUT_MS — ls-remote hits the network). */
+41
View File
@@ -85,6 +85,7 @@ import {
attachSessionListeners,
detachSessionListeners,
} from './session-listener-wiring.js';
import { sessionWaits } from './session-wait-registry.js';
import {
wireRespawnListeners,
setupTimedRespawn,
@@ -621,6 +622,7 @@ export class WebServer extends EventEmitter {
getModelConfig: this.getModelConfig.bind(this),
getClaudeModeConfig: this.getClaudeModeConfig.bind(this),
getTerminalHistoryConfig: this.getTerminalHistoryConfig.bind(this),
getAgentSkillEnabled: this.getAgentSkillEnabled.bind(this),
getDefaultClaudeMdPath: this.getDefaultClaudeMdPath.bind(this),
getLightState: this.getLightState.bind(this),
getLightSessionsState: this.getLightSessionsState.bind(this),
@@ -805,7 +807,24 @@ export class WebServer extends EventEmitter {
const clientId =
typeof query.clientId === 'string' && SSE_CLIENT_ID_RE.test(query.clientId) ? query.clientId : undefined;
// Carry over the headers the security hook already set on this reply.
//
// writeHead goes straight to the Node response and bypasses Fastify's header
// store, so everything the onRequest hook granted is silently dropped —
// including the Access-Control-Allow-Origin it emits for localhost origins.
// The result is an internal contradiction: a localhost page may call every
// /api endpoint cross-origin, but its EventSource fails CORS. The security
// headers (nosniff, frame-options, CSP) were lost the same way.
//
// The other raw-writeHead routes live in file-routes.ts and share a helper;
// this one keeps its own copy so the server does not import from a route
// module it registers.
const inherited: Record<string, number | string | string[]> = {};
for (const [name, value] of Object.entries(reply.getHeaders())) {
if (value !== undefined) inherited[name] = value;
}
reply.raw.writeHead(200, {
...inherited,
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
Connection: 'keep-alive',
@@ -1230,6 +1249,16 @@ export class WebServer extends EventEmitter {
}
}
// Release anything blocked on this session, in the documented order: 'exit'
// first so an until=exit caller gets its signal, then cancelAll so everyone
// else resolves with ended:true instead of timing out.
//
// The 'exit' here is NOT redundant with the PTY-exit listener: listeners are
// detached a few lines above, before `session.stop()`, so on a delete the
// session's own exit event never reaches the registry.
sessionWaits.notifySignal(sessionId, 'exit');
sessionWaits.cancelAll(sessionId);
this.broadcast(SseEvent.SessionDeleted, { id: sessionId });
}
@@ -1624,6 +1653,13 @@ export class WebServer extends EventEmitter {
return resolveTerminalHistoryConfig(settings);
}
// Whether the Codeman agent skill is injected into cases on Claude session create
// (synced `agentSkillEnabled` setting, default OFF; docs/agent-control-plan.md §2).
private async getAgentSkillEnabled(): Promise<boolean> {
const settings = await this.readSettings();
return settings.agentSkillEnabled === true;
}
// Helper to get model configuration from settings
private async getModelConfig(): Promise<{
defaultModel?: string;
@@ -2828,6 +2864,11 @@ export class WebServer extends EventEmitter {
// Gracefully close all SSE connections and clear batching state
this.sse.stop();
// Release every pending long-poll waiter. Their timers are deliberately not
// unref'd (an unref'd timer can let the process exit mid-wait and strand the
// response), so without this a 10-minute wait holds shutdown open.
sessionWaits.cancelEverything();
this.lastRecordedTokens.clear();
// Stop multiplexer and flush pending saves
+28
View File
@@ -27,6 +27,7 @@ import type { RalphStatusBlock, CircuitBreakerStatus } from '../types.js';
import { SseEvent } from './sse-events.js';
import { getLifecycleLog } from '../session-lifecycle-log.js';
import { fileStreamManager } from '../file-stream-manager.js';
import { sessionWaits } from './session-wait-registry.js';
/** Stored listener references for session cleanup (prevents memory leaks) */
export interface SessionListenerRefs {
@@ -92,6 +93,9 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
/** Batches PTY output → broadcasts `session:terminal` at 16-50ms intervals */
terminal: (data) => {
// Feeds `GET /api/sessions/:id/wait-output`. No-ops with a single Map lookup
// when nothing is waiting, which is the case on virtually every chunk.
sessionWaits.notifyOutput(session.id, data);
deps.batchTerminalData(session.id, data);
},
@@ -137,6 +141,28 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
/** Broadcasts `session:exit` + `session:updated` — PTY process exited; cleans up respawn, timers, listeners */
exit: (code) => {
// Before anything that can throw: a caller blocked on this session must learn
// the process died rather than sit until its timeout.
//
// Both halves are required, in this order — the same pair `_doCleanupSession`
// uses on the delete path, for the same reason. `notifySignal` resolves ONLY
// waiters that asked for `exit`; everyone else (`until=working`, `until=stop`,
// every wait-output) would keep a slot in the process-wide pool until their
// timeout, on a session whose feeds this very handler is about to tear down:
// `removeSessionListenerRefs` below detaches the `terminal` listener that is
// the only input to `notifyOutput`, and the `idle`/`working` listeners with it.
// Nothing can reach those waiters afterwards, so holding them is a guaranteed
// ten-minute lie. `cancelAll` answers them `ended: true`, which the plan's §3.6
// specifies for exactly this case ("Never hang").
//
// Safe against the respawn cycle: a respawn writes `/clear` + a kickstart
// prompt through the mux and never restarts the PTY, so it emits no `exit` and
// cannot cancel an orchestrating agent's wait. And for an agent driving a
// worker this is the right trade even when the PTY exit was only a tmux
// DETACH: `ended` means "re-check and re-issue", one extra round trip, versus
// burning the caller's entire timeout learning nothing.
sessionWaits.notifySignal(session.id, 'exit');
sessionWaits.cancelAll(session.id);
getLifecycleLog().log({
event: 'exit',
sessionId: session.id,
@@ -187,6 +213,7 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
/** Broadcasts `session:working` — Claude started processing */
working: () => {
sessionWaits.notifySignal(session.id, 'working');
deps.broadcast(SseEvent.SessionWorking, { id: session.id });
const tracker = deps.getRunSummaryTracker(session.id);
if (tracker) {
@@ -197,6 +224,7 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
/** Broadcasts `session:idle` — Claude finished processing, waiting for input */
idle: () => {
sessionWaits.notifySignal(session.id, 'idle');
deps.broadcast(SseEvent.SessionIdle, { id: session.id });
deps.broadcastSessionStateDebounced(session.id);
const tracker = deps.getRunSummaryTracker(session.id);
File diff suppressed because it is too large Load Diff
+122
View File
@@ -0,0 +1,122 @@
/**
* @fileoverview Unit tests for the agent-skill injection helpers in hooks-config.ts
* (`applyAgentSkill`, `installAgentSkillInto`, `removeAgentSkillFrom`).
*
* These run against the REAL packaged source (`skills/codeman/` at the repo root),
* so they double as a guard that the skill files exist and are readable: an npm
* publish without them would be caught here before the `files` entry silently
* ignores the missing directory.
*
* Pure filesystem tests in a per-test temp dir. Port: N/A.
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { mkdtemp, rm, mkdir, writeFile, readFile, symlink, readdir } from 'node:fs/promises';
import { existsSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import { applyAgentSkill, installAgentSkillInto, removeAgentSkillFrom } from '../src/hooks-config.js';
const MARKER_PREFIX = '<!-- codeman-managed-agent-skill';
let casePath: string;
const skillDir = () => join(casePath, '.claude', 'skills', 'codeman');
beforeEach(async () => {
casePath = await mkdtemp(join(tmpdir(), 'codeman-agent-skill-'));
});
afterEach(async () => {
await rm(casePath, { recursive: true, force: true });
});
describe('installAgentSkillInto / applyAgentSkill(enabled)', () => {
it('installs SKILL.md (marker appended) and the reference files from the packaged source', async () => {
const result = await applyAgentSkill(casePath, true);
expect(result).toBe('installed');
const skillMd = await readFile(join(skillDir(), 'SKILL.md'), 'utf-8');
expect(skillMd.startsWith('---\nname: codeman')).toBe(true);
expect(skillMd).toContain(MARKER_PREFIX);
// Reference files ride along byte-for-byte (no marker there).
const sourceEndpoints = await readFile(
join(process.cwd(), 'skills', 'codeman', 'reference', 'endpoints.md'),
'utf-8'
);
const injectedEndpoints = await readFile(join(skillDir(), 'reference', 'endpoints.md'), 'utf-8');
expect(injectedEndpoints).toBe(sourceEndpoints);
expect(existsSync(join(skillDir(), 'reference', 'recipes.md'))).toBe(true);
});
it('is idempotent: a second run reports unchanged', async () => {
await applyAgentSkill(casePath, true);
expect(await applyAgentSkill(casePath, true)).toBe('unchanged');
});
it('refreshes a stale Codeman-managed copy back to the packaged content', async () => {
await applyAgentSkill(casePath, true);
const original = await readFile(join(skillDir(), 'SKILL.md'), 'utf-8');
// Simulate an older injected version: content differs but the marker is intact.
await writeFile(join(skillDir(), 'SKILL.md'), `stale content\n${MARKER_PREFIX}: old -->\n`);
expect(await applyAgentSkill(casePath, true)).toBe('refreshed');
expect(await readFile(join(skillDir(), 'SKILL.md'), 'utf-8')).toBe(original);
});
it('never clobbers a user-authored skills/codeman (no marker)', async () => {
await mkdir(skillDir(), { recursive: true });
await writeFile(join(skillDir(), 'SKILL.md'), '---\nname: codeman\n---\nmy own skill\n');
expect(await applyAgentSkill(casePath, true)).toBe('foreign');
expect(await readFile(join(skillDir(), 'SKILL.md'), 'utf-8')).toContain('my own skill');
expect(existsSync(join(skillDir(), 'reference'))).toBe(false);
});
it('refuses to write through a symlinked skill dir (dogfooding layout)', async () => {
await mkdir(join(casePath, '.claude', 'skills'), { recursive: true });
await symlink(join(casePath, 'elsewhere'), skillDir());
expect(await installAgentSkillInto(skillDir())).toBe('symlink');
});
it('refuses to write through a symlinked skills/ parent', async () => {
await mkdir(join(casePath, 'real-skills'), { recursive: true });
await mkdir(join(casePath, '.claude'), { recursive: true });
await symlink(join(casePath, 'real-skills'), join(casePath, '.claude', 'skills'));
expect(await installAgentSkillInto(skillDir())).toBe('symlink');
expect(await readdir(join(casePath, 'real-skills'))).toEqual([]);
});
});
describe('removeAgentSkillFrom / applyAgentSkill(disabled)', () => {
it('removes our copy and prunes the emptied directories', async () => {
await applyAgentSkill(casePath, true);
expect(await applyAgentSkill(casePath, false)).toBe('removed');
expect(existsSync(skillDir())).toBe(false);
expect(existsSync(join(casePath, '.claude', 'skills'))).toBe(false);
// `.claude` itself is not ours to prune.
expect(existsSync(join(casePath, '.claude'))).toBe(true);
});
it('reports absent when there is nothing to remove', async () => {
expect(await applyAgentSkill(casePath, false)).toBe('absent');
});
it('leaves a user-authored copy untouched', async () => {
await mkdir(skillDir(), { recursive: true });
await writeFile(join(skillDir(), 'SKILL.md'), 'my own skill\n');
expect(await applyAgentSkill(casePath, false)).toBe('foreign');
expect(existsSync(join(skillDir(), 'SKILL.md'))).toBe(true);
});
it("preserves a user's extra files in the directory (no rm -rf)", async () => {
await applyAgentSkill(casePath, true);
await writeFile(join(skillDir(), 'reference', 'my-notes.md'), 'mine\n');
expect(await applyAgentSkill(casePath, false)).toBe('removed');
expect(existsSync(join(skillDir(), 'SKILL.md'))).toBe(false);
expect(existsSync(join(skillDir(), 'reference', 'endpoints.md'))).toBe(false);
// The user's file and the directories holding it survive.
expect(await readFile(join(skillDir(), 'reference', 'my-notes.md'), 'utf-8')).toBe('mine\n');
});
});
+54 -40
View File
@@ -71,7 +71,8 @@ describe('AiIdleChecker', () => {
describe('Output Parsing', () => {
it('should parse IDLE verdict', async () => {
// Set up mock to return IDLE result after polling
mockedReadFileSync.mockReturnValueOnce('') // writeFileSync creates empty file
mockedReadFileSync
.mockReturnValueOnce('') // writeFileSync creates empty file
.mockReturnValueOnce('IDLE\nSession shows completion message and prompt.\n__AICHECK_DONE__');
const checkPromise = checker.check('some terminal output');
@@ -87,7 +88,8 @@ describe('AiIdleChecker', () => {
});
it('should parse WORKING verdict', async () => {
mockedReadFileSync.mockReturnValueOnce('')
mockedReadFileSync
.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nSpinner characters detected, still processing.\n__AICHECK_DONE__');
const checkPromise = checker.check('some terminal output');
@@ -100,8 +102,7 @@ describe('AiIdleChecker', () => {
});
it('should handle lowercase verdict', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('idle\nDone.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('idle\nDone.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
@@ -112,7 +113,8 @@ describe('AiIdleChecker', () => {
});
it('should return ERROR for unparseable output', async () => {
mockedReadFileSync.mockReturnValueOnce('')
mockedReadFileSync
.mockReturnValueOnce('')
.mockReturnValueOnce('Something unexpected happened.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
@@ -125,8 +127,7 @@ describe('AiIdleChecker', () => {
});
it('should return ERROR for empty output', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
@@ -175,10 +176,34 @@ describe('AiIdleChecker', () => {
await vi.advanceTimersByTimeAsync(500);
await checkPromise;
expect(mockedWriteFileSync).toHaveBeenCalledWith(
expect.stringContaining('codeman-aicheck-'),
''
expect(mockedWriteFileSync).toHaveBeenCalledWith(expect.stringContaining('codeman-aicheck-'), '');
});
it('should keep Claude stderr separate from verdict output', async () => {
mockedReadFileSync.mockReturnValue('IDLE\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
await checkPromise;
const spawnArgs = mockedSpawn.mock.calls[0]?.[1];
const command = spawnArgs?.[spawnArgs.length - 1];
expect(command).toEqual(expect.any(String));
expect(command).toContain(' 2> "');
expect(command).not.toContain('2>&1');
});
it('should include Claude stderr when no verdict is produced', async () => {
mockedReadFileSync.mockImplementation((path) =>
String(path).includes('-stderr-') ? 'Claude CLI failed to load settings' : '__AICHECK_DONE__'
);
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
const result = await checkPromise;
expect(result.verdict).toBe('ERROR');
expect(result.reasoning).toContain('Claude CLI failed to load settings');
});
});
@@ -223,7 +248,7 @@ describe('AiIdleChecker', () => {
// Should have tried to kill the tmux session (initial kill + cleanup kill)
const killCalls = mockedExecSync.mock.calls.filter(
call => typeof call[0] === 'string' && call[0].includes('kill-session')
(call) => typeof call[0] === 'string' && call[0].includes('kill-session')
);
expect(killCalls.length).toBeGreaterThan(0);
});
@@ -236,8 +261,7 @@ describe('AiIdleChecker', () => {
describe('Cooldown', () => {
it('should start cooldown after WORKING verdict', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nStill processing.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nStill processing.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
@@ -250,8 +274,7 @@ describe('AiIdleChecker', () => {
});
it('should return to ready after cooldown expires', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -267,8 +290,7 @@ describe('AiIdleChecker', () => {
});
it('should not start new check during cooldown', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
const firstCheck = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -283,8 +305,7 @@ describe('AiIdleChecker', () => {
describe('Error Handling', () => {
it('should start error cooldown after parse error', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('garbage output\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage output\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -302,8 +323,7 @@ describe('AiIdleChecker', () => {
const cooldowns = [1100, 2100]; // Wait slightly longer than each cooldown
for (let i = 0; i < 3; i++) {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('garbage\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -321,8 +341,7 @@ describe('AiIdleChecker', () => {
it('should reset error counter on successful check', async () => {
// First check: error
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('garbage\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage\n__AICHECK_DONE__');
const firstCheck = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
await firstCheck;
@@ -332,8 +351,7 @@ describe('AiIdleChecker', () => {
await vi.advanceTimersByTimeAsync(1100);
// Second check: success
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('IDLE\nDone.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\nDone.\n__AICHECK_DONE__');
const secondCheck = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
await secondCheck;
@@ -352,8 +370,7 @@ describe('AiIdleChecker', () => {
describe('Buffer Handling', () => {
it('should strip ANSI codes from terminal buffer', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
const ansiBuffer = '\x1b[1mBold\x1b[0m \x1b[32mGreen\x1b[0m text';
const checkPromise = checker.check(ansiBuffer);
@@ -365,8 +382,7 @@ describe('AiIdleChecker', () => {
});
it('should trim buffer to maxContextChars', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
// Create buffer longer than maxContextChars (1000)
const longBuffer = 'x'.repeat(2000);
@@ -402,8 +418,7 @@ describe('AiIdleChecker', () => {
describe('Reset', () => {
it('should clear all state on reset', async () => {
// Trigger a WORKING verdict to set state
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -440,24 +455,24 @@ describe('AiIdleChecker', () => {
const handler = vi.fn();
checker.on('checkCompleted', handler);
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('IDLE\nAll done.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\nAll done.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
await checkPromise;
expect(handler).toHaveBeenCalledWith(expect.objectContaining({
verdict: 'IDLE',
}));
expect(handler).toHaveBeenCalledWith(
expect.objectContaining({
verdict: 'IDLE',
})
);
});
it('should emit cooldownStarted event after WORKING', async () => {
const handler = vi.fn();
checker.on('cooldownStarted', handler);
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -477,8 +492,7 @@ describe('AiIdleChecker', () => {
const cooldowns = [1100, 2100]; // Wait longer than exponential backoff
for (let i = 0; i < 3; i++) {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('garbage\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
await checkPromise;
+109
View File
@@ -0,0 +1,109 @@
/**
* Issue #205, round 2: `getClaudeCliVersion()` used to cache FAILURE forever.
*
* It stored `null` on any exception and guarded on `!== undefined`, so a single
* failed probe — the 5s exec timeout, a PATH-starved systemd/launchd
* environment, a transient fs hiccup — at the first Claude session start left
* `cliVersion` undefined for every Claude session until the server restarted.
* An undefined `cliVersion` silently disables wheel-forwarding to Claude's own
* transcript (`_shouldForwardWheelToApp`), which is the only route to history
* for a repaint-mode pane: a dead wheel on every device at once, which is what
* the reporter described (phone + iPad + laptop all broken together points at a
* SERVER-side cause, not a browser one).
*
* The probe itself can't run under vitest (it would spawn a real `claude`), so
* these drive the cache policy directly with an injected probe and clock.
*/
import { describe, expect, it, vi } from 'vitest';
import {
claudeVersionRetryDelayMs,
getClaudeCliVersion,
resolveClaudeCliVersion,
type ClaudeVersionProbeState,
} from '../src/utils/claude-cli-resolver.js';
const freshState = (): ClaudeVersionProbeState => ({ failures: 0, lastFailureAt: 0 });
describe('claude --version probe caching', () => {
it('probes once on success and never spawns again', () => {
const state = freshState();
const probe = vi.fn(() => '2.1.223');
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBe('2.1.223');
expect(resolveClaudeCliVersion(state, 2_000, probe)).toBe('2.1.223');
expect(resolveClaudeCliVersion(state, 9_999_999, probe)).toBe('2.1.223');
expect(probe).toHaveBeenCalledTimes(1);
});
it('RETRIES after a failed probe instead of poisoning the process', () => {
const state = freshState();
const probe = vi
.fn<() => string | null>()
.mockImplementationOnce(() => {
throw new Error('spawn claude ETIMEDOUT'); // the shipped failure mode
})
.mockImplementationOnce(() => '2.1.223');
// First session start: probe blows up, no version.
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBeNull();
// Immediately after, the negative cache holds — no probe storm.
expect(resolveClaudeCliVersion(state, 30_000, probe)).toBeNull();
expect(probe).toHaveBeenCalledTimes(1);
// Once the retry window elapses, the next session start probes again and
// wheel-forwarding comes back without a server restart.
expect(resolveClaudeCliVersion(state, 61_000, probe)).toBe('2.1.223');
expect(probe).toHaveBeenCalledTimes(2);
});
it('treats an unparseable version like a failure (retryable, not cached)', () => {
const state = freshState();
const probe = vi.fn<() => string | null>(() => null); // e.g. output without a x.y.z
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBeNull();
expect(resolveClaudeCliVersion(state, 61_000, probe)).toBeNull();
expect(probe).toHaveBeenCalledTimes(2);
expect(state.version).toBeUndefined(); // nothing cached as "known bad"
});
it('clears the failure streak once a probe succeeds', () => {
const state = freshState();
const probe = vi
.fn<() => string | null>()
.mockImplementationOnce(() => null)
.mockImplementationOnce(() => '2.1.223');
resolveClaudeCliVersion(state, 1_000, probe);
expect(state.failures).toBe(1);
resolveClaudeCliVersion(state, 61_000, probe);
expect(state.failures).toBe(0);
expect(state.lastFailureAt).toBe(0);
});
it('backs off so a genuinely missing binary cannot probe on every session start', () => {
expect(claudeVersionRetryDelayMs(0)).toBe(0);
expect(claudeVersionRetryDelayMs(1)).toBe(60_000);
expect(claudeVersionRetryDelayMs(2)).toBe(120_000);
expect(claudeVersionRetryDelayMs(3)).toBe(240_000);
// Capped, so it keeps retrying forever without ever spinning.
expect(claudeVersionRetryDelayMs(50)).toBe(15 * 60_000);
const state = freshState();
const probe = vi.fn<() => string | null>(() => null);
resolveClaudeCliVersion(state, 0, probe); // failure 1 → retry at 60s
resolveClaudeCliVersion(state, 30_000, probe); // still inside the window
expect(probe).toHaveBeenCalledTimes(1);
resolveClaudeCliVersion(state, 60_000, probe); // failure 2 → retry at 120s
resolveClaudeCliVersion(state, 119_000, probe);
expect(probe).toHaveBeenCalledTimes(2);
resolveClaudeCliVersion(state, 180_001, probe);
expect(probe).toHaveBeenCalledTimes(3);
});
it('stays hermetic under vitest without recording a phantom failure', () => {
// The guard returns before the probe, and — unlike the old code, which wrote
// null into the cache here — leaves the cache untouched.
expect(getClaudeCliVersion()).toBeNull();
expect(getClaudeCliVersion()).toBeNull();
});
});
+182
View File
@@ -0,0 +1,182 @@
/**
* Unit tests for the pure halves of daemon-control (issue #231): argv rebuilding,
* the readiness URL, pidfile parsing, the stale-pid identity check, and the
* `/api/status` probe against a real socket.
*/
import { describe, it, expect, afterAll, beforeAll } from 'vitest';
import http from 'node:http';
import {
buildBaseUrl,
buildStatusUrl,
buildWebArgs,
isProcessAlive,
looksLikeCodemanWeb,
parsePidFileContents,
probeServer,
} from '../src/daemon-control.js';
const PORT = 3216;
describe('buildWebArgs', () => {
it('always passes host and port through explicitly', () => {
expect(buildWebArgs({ host: '127.0.0.1', port: 3000, https: false })).toEqual([
'web',
'--host',
'127.0.0.1',
'--port',
'3000',
]);
});
it('forwards every optional flag it was given', () => {
const args = buildWebArgs({
host: '0.0.0.0',
port: 8080,
https: true,
titleHostname: 'tower',
allowUnauthenticatedNetwork: true,
multiuser: true,
});
expect(args).toEqual([
'web',
'--host',
'0.0.0.0',
'--port',
'8080',
'--https',
'--title-hostname',
'tower',
'--allow-unauthenticated-network',
'--multiuser',
]);
});
it('never re-emits the daemon flags themselves (the child must not re-fork)', () => {
const args = buildWebArgs({ host: '127.0.0.1', port: 3000, https: false });
expect(args).not.toContain('--daemon');
expect(args).not.toContain('-d');
});
});
describe('buildBaseUrl', () => {
it('is the address a browser can open, with no path on it', () => {
expect(buildBaseUrl({ host: '127.0.0.1', port: 3000, https: false })).toBe('http://127.0.0.1:3000');
expect(buildBaseUrl({ host: '0.0.0.0', port: 8443, https: true })).toBe('https://127.0.0.1:8443');
});
});
describe('buildStatusUrl', () => {
it('uses http by default and https when asked', () => {
expect(buildStatusUrl({ host: '127.0.0.1', port: 3000, https: false })).toBe('http://127.0.0.1:3000/api/status');
expect(buildStatusUrl({ host: '127.0.0.1', port: 3000, https: true })).toBe('https://127.0.0.1:3000/api/status');
});
it('rewrites wildcard binds to loopback, since they are not connectable', () => {
expect(buildStatusUrl({ host: '0.0.0.0', port: 3000, https: false })).toBe('http://127.0.0.1:3000/api/status');
expect(buildStatusUrl({ host: '::', port: 3000, https: false })).toBe('http://127.0.0.1:3000/api/status');
});
it('brackets a bare IPv6 literal', () => {
expect(buildStatusUrl({ host: '::1', port: 3000, https: false })).toBe('http://[::1]:3000/api/status');
expect(buildStatusUrl({ host: '[::1]', port: 3000, https: false })).toBe('http://[::1]:3000/api/status');
});
});
describe('parsePidFileContents', () => {
it('accepts a plain pid with surrounding whitespace', () => {
expect(parsePidFileContents('4242\n')).toBe(4242);
expect(parsePidFileContents(' 4242 ')).toBe(4242);
});
it('rejects garbage, empties and floats', () => {
expect(parsePidFileContents('')).toBeNull();
expect(parsePidFileContents('not a pid')).toBeNull();
expect(parsePidFileContents('42.5')).toBeNull();
expect(parsePidFileContents('-42')).toBeNull();
});
it('rejects pid 0 and pid 1: neither is ever our server', () => {
expect(parsePidFileContents('0')).toBeNull();
expect(parsePidFileContents('1')).toBeNull();
});
});
describe('looksLikeCodemanWeb', () => {
it('matches the ways the server is actually launched', () => {
expect(looksLikeCodemanWeb('/usr/bin/node /home/u/.codeman/app/dist/index.js web')).toBe(true);
expect(looksLikeCodemanWeb('/usr/bin/node dist/index.js web --https')).toBe(true);
expect(looksLikeCodemanWeb('node /repo/src/index.ts web --port 3000')).toBe(true);
expect(looksLikeCodemanWeb('/opt/homebrew/bin/codeman web')).toBe(true);
expect(looksLikeCodemanWeb('aicodeman web --host 0.0.0.0')).toBe(true);
});
it('rejects anything that inherited a recycled pid', () => {
expect(looksLikeCodemanWeb(null)).toBe(false);
expect(looksLikeCodemanWeb('')).toBe(false);
expect(looksLikeCodemanWeb('/usr/bin/node dist/index.js session list')).toBe(false);
expect(looksLikeCodemanWeb('vim web')).toBe(false);
expect(looksLikeCodemanWeb('/usr/lib/systemd/systemd --user')).toBe(false);
});
});
describe('isProcessAlive', () => {
it('sees this very process', () => {
expect(isProcessAlive(process.pid)).toBe(true);
});
it('does not see an unused high pid', () => {
// 2^22 is above the default pid_max on Linux and macOS.
expect(isProcessAlive(4_194_303)).toBe(false);
});
});
describe('probeServer', () => {
let server: http.Server;
beforeAll(async () => {
server = http.createServer((req, res) => {
if (req.url === '/unauthorized') {
res.writeHead(401).end('Unauthorized');
return;
}
if (req.url === '/foreign') {
res.writeHead(200, { 'Content-Type': 'text/html' }).end('<html>some other app</html>');
return;
}
res.writeHead(200, { 'Content-Type': 'application/json' });
res.end(JSON.stringify({ success: true, data: { version: '9.9.9' } }));
});
await new Promise<void>((resolve) => server.listen(PORT, '127.0.0.1', resolve));
});
afterAll(async () => {
await new Promise<void>((resolve) => server.close(() => resolve()));
});
it('reports up and reads the version back', async () => {
const result = await probeServer(`http://127.0.0.1:${PORT}/api/status`);
expect(result.up).toBe(true);
expect(result.version).toBe('9.9.9');
});
it('counts a 401 as up, because auth being active proves a server is there', async () => {
const result = await probeServer(`http://127.0.0.1:${PORT}/unauthorized`);
expect(result.up).toBe(true);
});
it('does not mistake an unrelated service squatting on the port for Codeman', async () => {
const result = await probeServer(`http://127.0.0.1:${PORT}/foreign`);
expect(result.up).toBe(false);
});
it('reports down when nothing is listening', async () => {
const result = await probeServer(`http://127.0.0.1:${PORT + 1}/api/status`, 1000);
expect(result.up).toBe(false);
});
it('reports down for a malformed url instead of throwing', async () => {
const result = await probeServer('not-a-url');
expect(result.up).toBe(false);
});
});
+20 -1
View File
@@ -127,12 +127,31 @@ describe('refreshStaleCodemanHooks', () => {
const after = JSON.parse(readFileSync(settingsPath, 'utf-8'));
expect(JSON.stringify(after.hooks)).toContain(SECRET_HEADER);
expect(JSON.stringify(after.hooks)).toContain('CODEMAN_BACKGROUND_REWAKE_V');
expect(JSON.stringify(after.hooks)).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(JSON.stringify(after.hooks.Stop)).toContain('./notify-user.sh');
expect(after.hooks.PostToolUse).toEqual(expect.arrayContaining([customPostToolUse]));
expect(after.hooks.CustomEvent).toEqual(customEvent);
});
// A case can be current on the secret AND the background-wake hook and still carry
// the `-k`-less curl shape, which exits 60 against a self-signed HTTPS API and is
// swallowed by `|| true` — every hook event dead, silently. The refresh must treat
// that as a third stale shape.
it('heals a current-looking block whose hook curls lack -k (HTTPS self-signed installs)', async () => {
const { generateHooksConfig } = await import('../src/hooks-config.js');
const flagless = JSON.parse(JSON.stringify(generateHooksConfig()).replaceAll('curl -sk ', 'curl -s '));
writeFileSync(settingsPath, JSON.stringify({ hooks: flagless.hooks }, null, 2));
await refreshStaleCodemanHooks(dir);
const after = readFileSync(settingsPath, 'utf-8');
expect(after).toContain('curl -sk -X POST');
expect(after).not.toContain('curl -s -X POST');
// and the pass is convergent: a second refresh must not rewrite
await refreshStaleCodemanHooks(dir);
expect(readFileSync(settingsPath, 'utf-8')).toBe(after);
});
it('is a no-op when settings.local.json is absent (does not create one)', async () => {
await refreshStaleCodemanHooks(dir);
expect(existsSync(settingsPath)).toBe(false);
+283 -3
View File
@@ -6,13 +6,15 @@
*/
import { describe, it, expect, beforeAll, beforeEach, afterAll, afterEach } from 'vitest';
import { existsSync, readFileSync, writeFileSync, mkdirSync, rmSync } from 'node:fs';
import { closeSync, existsSync, openSync, readFileSync, writeFileSync, mkdirSync, rmSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import { spawn } from 'node:child_process';
import {
ensureCodemanHooks,
generateBackgroundWakeScript,
generateHooksConfig,
generateSubagentStopGuardScript,
refreshStaleCodemanHooks,
writeHooksConfig,
} from '../src/hooks-config.js';
@@ -35,6 +37,20 @@ describe('generateHooksConfig', () => {
expect(config.hooks.Stop).toHaveLength(1);
});
it('should guard subagent stops while their background work is active', () => {
const config = generateHooksConfig();
const subagentHooks = config.hooks.SubagentStop as Array<{
hooks: Array<{ type: string; command: string; args: string[]; timeout: number }>;
}>;
expect(subagentHooks).toHaveLength(1);
expect(subagentHooks[0].hooks[0]).toMatchObject({
type: 'command',
command: 'node',
args: ['-e', generateSubagentStopGuardScript()],
});
});
it('should configure a self-contained Bash background-task rewake hook', () => {
const config = generateHooksConfig();
const postToolHooks = config.hooks.PostToolUse as Array<{
@@ -95,6 +111,17 @@ describe('generateHooksConfig', () => {
expect(notifHooks[0].hooks[0].command).toContain('|| true');
});
// On --https/tailscale installs CODEMAN_API_URL is HTTPS with a self-signed cert.
// A `-k`-less hook curl exits 60 there, the `|| true` swallows it, and every hook
// event (stop, permission_prompt, elicitation_dialog, idle_prompt, teammate_idle,
// task_completed) dies silently — killing respawn's idle signals and the wait
// endpoints' stop/blocked. The statusline exporter always carried -k; the hooks must too.
it('every hook curl tolerates a self-signed HTTPS API (curl -sk)', () => {
const serialized = JSON.stringify(generateHooksConfig());
expect(serialized).toContain('curl -sk -X POST');
expect(serialized).not.toContain('curl -s -X POST');
});
it('should set timeout to 10 seconds (hook timeout fields are seconds)', () => {
const config = generateHooksConfig();
const notifHooks = config.hooks.Notification as Array<{ hooks: Array<{ timeout: number }> }>;
@@ -199,7 +226,8 @@ describe('writeHooksConfig', () => {
const parsed = JSON.parse(readFileSync(settingsPath, 'utf-8'));
expect(parsed.hooks.PostToolUse).toHaveLength(1);
expect(JSON.stringify(parsed.hooks.PostToolUse)).toContain('CODEMAN_BACKGROUND_REWAKE_V');
expect(JSON.stringify(parsed.hooks.PostToolUse)).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(JSON.stringify(parsed.hooks.SubagentStop)).toContain('CODEMAN_SUBAGENT_STOP_GUARD_V1');
});
it('should replace an older rewake script version without duplicating it', async () => {
@@ -231,10 +259,29 @@ describe('writeHooksConfig', () => {
const serialized = JSON.stringify(parsed.hooks.PostToolUse);
expect(parsed.hooks.PostToolUse).toHaveLength(1);
expect(parsed.hooks.PostToolUse[0].hooks).toHaveLength(1);
expect(serialized).toContain('CODEMAN_BACKGROUND_REWAKE_V2');
expect(serialized).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(serialized).not.toContain('CODEMAN_BACKGROUND_REWAKE_V1');
});
it('replaces the V2 background hook without duplicating it', async () => {
const claudeDir = join(testDir, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
mkdirSync(claudeDir, { recursive: true });
const oldSettings = JSON.stringify({ hooks: generateHooksConfig().hooks }, null, 2).replaceAll(
'CODEMAN_BACKGROUND_REWAKE_V3',
'CODEMAN_BACKGROUND_REWAKE_V2'
);
writeFileSync(settingsPath, oldSettings);
await refreshStaleCodemanHooks(testDir);
const parsed = JSON.parse(readFileSync(settingsPath, 'utf-8'));
const postToolUse = JSON.stringify(parsed.hooks.PostToolUse);
expect(parsed.hooks.PostToolUse).toHaveLength(1);
expect(postToolUse).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(postToolUse).not.toContain('CODEMAN_BACKGROUND_REWAKE_V2');
});
it('should not add rewake hooks to a user-owned hook configuration', async () => {
const claudeDir = join(testDir, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
@@ -267,6 +314,35 @@ describe('writeHooksConfig', () => {
expect(parsed.hooks.Notification).toBeDefined();
});
it('should safely add Codeman hooks to an existing managed-case settings file', async () => {
const claudeDir = join(testDir, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
mkdirSync(claudeDir, { recursive: true });
const userHooks = {
PostToolUse: [{ matcher: 'Write', hooks: [{ type: 'command', command: './format.sh' }] }],
};
writeFileSync(settingsPath, JSON.stringify({ hooks: userHooks, permissions: { allow: ['Read'] } }, null, 2));
await ensureCodemanHooks(testDir);
const parsed = JSON.parse(readFileSync(settingsPath, 'utf-8'));
expect(parsed.permissions).toEqual({ allow: ['Read'] });
expect(parsed.hooks.PostToolUse).toEqual(expect.arrayContaining(userHooks.PostToolUse));
expect(JSON.stringify(parsed.hooks)).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(JSON.stringify(parsed.hooks)).toContain('CODEMAN_SUBAGENT_STOP_GUARD_V1');
});
it('should not replace a malformed managed-case settings file', async () => {
const claudeDir = join(testDir, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
mkdirSync(claudeDir, { recursive: true });
writeFileSync(settingsPath, '{ malformed');
await ensureCodemanHooks(testDir);
expect(readFileSync(settingsPath, 'utf-8')).toBe('{ malformed');
});
it('should handle malformed existing settings.local.json', async () => {
const claudeDir = join(testDir, '.claude');
mkdirSync(claudeDir, { recursive: true });
@@ -359,6 +435,210 @@ describe('background task rewake helper', () => {
expect(result.stderr).toContain('completed');
expect(result.stderr).toContain('/tmp/bg-test-1.output');
});
it('rewakes a subagent when Claude queues completion in the parent transcript', async () => {
const sessionId = '7148e9de-7673-48b8-bf38-6799e52c346a';
const sessionDir = join(testDir, sessionId);
const subagentDir = join(sessionDir, 'subagents');
const parentTranscriptPath = `${sessionDir}.jsonl`;
const subagentTranscriptPath = join(subagentDir, 'agent-afacts-class2.jsonl');
mkdirSync(subagentDir, { recursive: true });
writeFileSync(parentTranscriptPath, '');
writeFileSync(subagentTranscriptPath, '');
const resultPromise = runHelper({
session_id: sessionId,
agent_id: 'afacts-class2',
transcript_path: subagentTranscriptPath,
tool_response: {
backgroundTaskId: 'bg-subagent-1',
},
});
await new Promise((resolve) => setTimeout(resolve, 100));
writeFileSync(
parentTranscriptPath,
JSON.stringify({
type: 'queue-operation',
operation: 'enqueue',
content:
'<task-notification>\n<task-id>bg-subagent-1</task-id>\n<status>completed</status>\n' +
'<output-file>/tmp/bg-subagent-1.output</output-file>\n</task-notification>',
}) + '\n'
);
const result = await resultPromise;
expect(result.code).toBe(2);
expect(result.stderr).toContain('bg-subagent-1');
expect(result.stderr).toContain('/tmp/bg-subagent-1.output');
});
it('includes a marked background report in the wake feedback', async () => {
const transcriptPath = join(testDir, 'transcript.jsonl');
const tasksDir = join(testDir, 'tasks');
const outputPath = join(tasksDir, 'bg-report-1.output');
mkdirSync(tasksDir, { recursive: true });
writeFileSync(transcriptPath, '');
writeFileSync(
outputPath,
[
'launcher output',
'=== CODEMAN_RESULT_BEGIN ===',
'Summary line',
'Detail after the old 30-line preview boundary',
'=== CODEMAN_RESULT_END ===',
].join('\n')
);
const resultPromise = runHelper({
transcript_path: transcriptPath,
tool_response: {
stdout: `Command running in background with ID: bg-report-1. Output is being written to: ${outputPath}.`,
},
});
await new Promise((resolve) => setTimeout(resolve, 100));
writeFileSync(
transcriptPath,
JSON.stringify({
type: 'queue-operation',
operation: 'enqueue',
content:
'<task-notification>\n<task-id>bg-report-1</task-id>\n<status>completed</status>\n' +
`<output-file>${outputPath}</output-file>\n</task-notification>`,
}) + '\n'
);
const result = await resultPromise;
expect(result.code).toBe(2);
expect(result.stderr).toContain('<codeman-background-result>');
expect(result.stderr).toContain('Summary line');
expect(result.stderr).toContain('Detail after the old 30-line preview boundary');
});
});
describe('subagent stop guard helper', () => {
const testDir = join(tmpdir(), 'codeman-subagent-stop-guard-test-' + Date.now());
beforeEach(() => {
mkdirSync(testDir, { recursive: true });
});
afterEach(() => {
rmSync(testDir, { recursive: true, force: true });
});
function runGuard(transcriptLines: unknown[]): Promise<{ code: number | null; stdout: string; stderr: string }> {
const transcriptPath = join(testDir, 'agent-test.jsonl');
writeFileSync(transcriptPath, transcriptLines.map((line) => JSON.stringify(line)).join('\n') + '\n');
return new Promise((resolve, reject) => {
const child = spawn(process.execPath, ['-e', generateSubagentStopGuardScript()], {
stdio: ['pipe', 'pipe', 'pipe'],
});
let stdout = '';
let stderr = '';
child.stdout.setEncoding('utf8');
child.stderr.setEncoding('utf8');
child.stdout.on('data', (chunk) => {
stdout += chunk;
});
child.stderr.on('data', (chunk) => {
stderr += chunk;
});
child.on('error', reject);
child.on('close', (code) => resolve({ code, stdout, stderr }));
child.stdin.end(JSON.stringify({ agent_transcript_path: transcriptPath }));
});
}
async function withLiveTask<T>(taskId: string, action: () => Promise<T>): Promise<T> {
const tasksDir = join(testDir, 'tasks');
mkdirSync(tasksDir, { recursive: true });
const outputFd = openSync(join(tasksDir, `${taskId}.output`), 'a');
const child = spawn(process.execPath, ['-e', 'setTimeout(() => {}, 10000)'], {
stdio: ['ignore', outputFd, outputFd],
});
await new Promise<void>((resolve, reject) => {
child.once('spawn', resolve);
child.once('error', reject);
});
closeSync(outputFd);
try {
return await action();
} finally {
const closed = new Promise<void>((resolve) => child.once('close', () => resolve()));
child.kill();
await closed;
}
}
const monitorResult = (taskId: string) => ({
type: 'user',
message: {
content: [
{
type: 'tool_result',
content: `Monitor started (task ${taskId}, pid 123).`,
},
],
},
});
const completion = (taskId: string) => ({
type: 'user',
message: {
content:
`<task-notification>\n<task-id>${taskId}</task-id>\n` + '<status>completed</status>\n</task-notification>',
},
});
it('blocks an intermediate subagent stop while a sibling monitor is active', async () => {
const result = await withLiveTask('monitor-still-live', () =>
runGuard([monitorResult('monitor-first'), monitorResult('monitor-still-live'), completion('monitor-first')])
);
expect(result.code).toBe(0);
expect(result.stderr).toBe('');
expect(JSON.parse(result.stdout)).toMatchObject({ decision: 'block' });
expect(result.stdout).toContain('monitor-still-live');
expect(result.stdout).not.toContain('monitor-first,');
});
it('allows a subagent to stop after all of its monitored work finishes', async () => {
const result = await runGuard([
monitorResult('monitor-first'),
monitorResult('monitor-second'),
completion('monitor-first'),
completion('monitor-second'),
]);
expect(result.code).toBe(0);
expect(result.stdout).toBe('');
expect(result.stderr).toBe('');
});
it('also recognizes background Bash task ownership', async () => {
const result = await withLiveTask('bash-live-1', () =>
runGuard([
{
type: 'user',
message: {
content: [
{
type: 'tool_result',
content: 'Command running in background with ID: bash-live-1. Output is being written to a task file.',
},
],
},
},
])
);
expect(JSON.parse(result.stdout)).toMatchObject({ decision: 'block' });
expect(result.stdout).toContain('bash-live-1');
});
});
// ========== Hook Event API Integration Tests ==========
+93
View File
@@ -89,4 +89,97 @@ describe('Stable HTTP contract (live server)', () => {
expect(body.success).toBe(false);
expect(body.errorCode).toBe('INVALID_INPUT');
});
/**
* The agent wait primitives, through the REAL pipeline.
*
* Their own route tests hand-roll a partial copy of the preSerialization hook that
* maps errorCode to status but does NOT wrap bare payloads — so nothing there
* proves these routes emit a correct envelope, a correct status, or work through
* the /api/v1 alias, and one assertion in them pins `{}` for a response no client
* will ever receive. This is the file whose docstring already claims that scope.
*/
describe('agent wait primitives', () => {
let sessionId: string;
beforeAll(async () => {
const res = await fetch(`${base}/api/sessions`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({}),
});
sessionId = (await res.json()).data.session.id;
expect(sessionId).toBeDefined();
});
afterAll(async () => {
await fetch(`${base}/api/sessions/${sessionId}`, { method: 'DELETE' });
});
it('answers a wait timeout as a 200 inside the envelope, on the /api/v1 alias', async () => {
// A timeout is the long-poll SUCCEEDING at "did this happen within N ms?"; a
// 4xx/5xx here would make every poll boundary indistinguishable from a failure.
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait?until=working&timeout=1000`);
expect(res.status).toBe(200);
expect(res.headers.get('cache-control')).toBe('no-store');
const body = await res.json();
expect(body.success).toBe(true);
expect(body.data.sessionId).toBe(sessionId);
// The one shape all three wait endpoints share.
expect(body.data.wait.timedOut).toBe(true);
expect(body.data.wait.signal).toBeNull();
expect(body.data.wait.timeoutMs).toBe(1000);
expect(body.data.wait.until).toEqual(['working']);
});
it('returns a contract-shaped 400 for an unknown until token', async () => {
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait?until=stpo`);
expect(res.status).toBe(400);
const body = await res.json();
expect(body.success).toBe(false);
expect(body.errorCode).toBe('INVALID_INPUT');
expect(body.error).toContain('stpo');
});
it('returns a contract-shaped 400 naming the bad query parameter', async () => {
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait?timeout=30s`);
expect(res.status).toBe(400);
const body = await res.json();
expect(body.errorCode).toBe('INVALID_INPUT');
expect(body.error).toContain('timeout');
});
it('wraps the non-wait input response as { success: true, data: {} }', async () => {
// What a client actually receives on the fire-and-forget path — NOT the bare
// `{}` the handler returns and the route tests assert.
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/input`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ input: 'hello' }),
});
expect(res.status).toBe(200);
expect(await res.json()).toEqual({ success: true, data: {} });
});
it('serves wait-output through the same envelope', async () => {
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait-output?match=NEVER_APPEARS&timeout=1000`);
expect(res.status).toBe(200);
const body = await res.json();
expect(body.success).toBe(true);
expect(body.data.wait.matched).toBe(false);
expect(body.data.wait.timedOut).toBe(true);
expect(body.data.wait.match).toBe('NEVER_APPEARS');
});
it('404s an unknown session on both new routes, with the error envelope', async () => {
for (const path of ['wait?until=idle', 'wait-output?match=x']) {
const res = await fetch(`${base}/api/v1/sessions/nonexistent/${path}`);
expect(res.status).toBe(404);
const body = await res.json();
expect(body.success).toBe(false);
expect(body.errorCode).toBe('NOT_FOUND');
}
});
});
});
+240
View File
@@ -0,0 +1,240 @@
/**
* @fileoverview Local-echo gating and input-ordering helpers for codex
* sessions (issues #218/#219/#220/#222).
*
* Codex's composer is interactive per keystroke: typing "/" pops a
* live-filtering command picker (#222), the composer grows as it wraps
* (#220), pastes arrive bracketed (#219) and arrows edit server-side state
* (#218). The buffer-until-Enter local echo overlay starves all of that, so
* codex-mode sessions must use plain PTY echo like shell. The shared overlay
* branch (claude/gemini/opencode) additionally flushes typed-but-unsent text
* before forwarding bracketed pastes and composer nav keys, and hands the
* session to pass-through after a nav key.
*
* Loaded via `vm` with a stubbed context (no jsdom), mirroring
* test/input-send-order.test.ts. End-to-end behavior was verified against a
* real codex 0.147.0 TUI in tmux through a headless browser.
*/
import { readFileSync } from 'node:fs';
import { performance } from 'node:perf_hooks';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it, vi } from 'vitest';
type OverlayStub = {
pendingText: string;
cleared: number;
suppressed: number;
prompts: unknown[];
clear(): void;
suppressBufferDetection(): void;
setPrompt(p: unknown): void;
appendText: ReturnType<typeof vi.fn>;
};
type AppInstance = {
activeSessionId: string | null;
sessions: Map<string, { mode: string }>;
terminal?: { focus: () => void };
_localEchoEnabled?: boolean;
_localEchoOverlay?: OverlayStub;
_pendingInput: string;
_flushedOffsets?: Map<string, number>;
_flushedTexts?: Map<string, string>;
_echoPassthroughSessions?: Set<string>;
loadAppSettingsFromStorage: () => Record<string, unknown>;
sendInput: ReturnType<typeof vi.fn>;
_updateLocalEchoState(): void;
_flushLocalEchoPending(): void;
insertTerminalText(text: string): void;
};
function loadContext() {
const read = (f: string) => readFileSync(resolve(import.meta.dirname, `../src/web/public/${f}`), 'utf8');
const windowStub: Record<string, unknown> = {
addEventListener: vi.fn(),
removeEventListener: vi.fn(),
};
const context = vm.createContext({
console,
performance,
setInterval: vi.fn(),
clearInterval: vi.fn(),
setTimeout,
clearTimeout,
requestAnimationFrame: vi.fn(),
HTMLCanvasElement: class HTMLCanvasElement {},
WebSocket: { OPEN: 1 },
fetch: vi.fn(),
document: { addEventListener: vi.fn(), documentElement: { dataset: {} } },
localStorage: {
length: 0,
key: vi.fn(),
getItem: vi.fn(),
setItem: vi.fn(),
removeItem: vi.fn(),
},
window: windowStub,
MobileDetection: {
isTouchDevice: () => true,
isHandheldDevice: () => false,
getDeviceType: () => 'desktop',
},
});
vm.runInContext(
`${read('constants.js')}\n${read('app.js')}\n${read('terminal-ui.js')}\nglobalThis.__CodemanApp = CodemanApp;`,
context
);
const CodemanApp = (context as unknown as { __CodemanApp: { prototype: object } }).__CodemanApp;
return {
CodemanApp,
terminalInput: (windowStub as { CodemanTerminalInput?: Record<string, unknown> }).CodemanTerminalInput!,
};
}
const { CodemanApp, terminalInput } = loadContext();
const isComposerNavKey = terminalInput.isComposerNavKey as (data: string) => boolean;
function makeOverlay(pending = ''): OverlayStub {
return {
pendingText: pending,
cleared: 0,
suppressed: 0,
prompts: [],
clear() {
this.cleared++;
this.pendingText = '';
},
suppressBufferDetection() {
this.suppressed++;
},
setPrompt(p: unknown) {
this.prompts.push(p);
},
appendText: vi.fn(),
};
}
function makeApp(mode: string, overlay = makeOverlay()): AppInstance {
const app = Object.create(CodemanApp.prototype) as AppInstance;
app.activeSessionId = 's1';
app.sessions = new Map([['s1', { mode }]]);
app._localEchoOverlay = overlay;
app._pendingInput = '';
app._flushedOffsets = new Map([['s1', 3]]);
app._flushedTexts = new Map([['s1', 'abc']]);
app.loadAppSettingsFromStorage = () => ({ localEchoEnabled: true });
app.sendInput = vi.fn().mockResolvedValue(undefined);
return app;
}
describe('CodemanTerminalInput.isComposerNavKey', () => {
it.each([
'\x1b[A',
'\x1b[B',
'\x1b[C',
'\x1b[D',
'\x1b[H',
'\x1b[F',
'\x1bOA',
'\x1bOD',
'\x1bOH',
'\x1bOF',
'\x1b[1;5C', // Ctrl+Right
'\x1b[1;2A', // Shift+Up
'\x1b[3~', // Delete
'\x1b[3;5~', // Ctrl+Delete
'\x1b[5~', // PgUp
'\x1b[6~', // PgDn
'\x1b[1~', // Home variant
'\x1b[4~', // End variant
])('classifies %j as a composer nav key', (seq) => {
expect(isComposerNavKey(seq)).toBe(true);
});
it.each([
'\x1b[?1;2c', // DA1 response
'\x1b[>0;276;0c', // DA2 response
'\x1b[12;34R', // CPR response
'\x1b[1;3R', // CPR response (small coords)
'\x1b[0n', // DSR response
'\x1b[15~', // F5 (function keys stay out)
'\x1b[200~hi\x1b[201~', // bracketed paste
'\x1b[?u', // kitty keyboard query response
'\x1bOP', // F1
'\x1b',
'a',
'abc',
'\r',
])('does NOT classify %j as a composer nav key', (seq) => {
expect(isComposerNavKey(seq)).toBe(false);
});
it('exports the bracketed paste prefix xterm puts on terminal.paste()', () => {
expect(terminalInput.BRACKETED_PASTE_START).toBe('\x1b[200~');
});
});
describe('_updateLocalEchoState mode gating', () => {
it('disables the overlay for codex sessions even with the setting ON (issues #218/#219/#220/#222)', () => {
const overlay = makeOverlay('pending');
const app = makeApp('codex', overlay);
app._updateLocalEchoState();
expect(app._localEchoEnabled).toBe(false);
expect(overlay.cleared).toBeGreaterThan(0);
});
it('disables the overlay for shell sessions (PTY provides its own echo)', () => {
const app = makeApp('shell');
app._updateLocalEchoState();
expect(app._localEchoEnabled).toBe(false);
});
it.each(['claude', 'gemini', 'opencode'])('keeps the overlay enabled for %s sessions', (mode) => {
const overlay = makeOverlay();
const app = makeApp(mode, overlay);
app._updateLocalEchoState();
expect(app._localEchoEnabled).toBe(true);
expect(overlay.prompts.length).toBeGreaterThan(0);
});
});
describe('_flushLocalEchoPending', () => {
it('moves pending text into _pendingInput and resets overlay + flushed tracking', () => {
const overlay = makeOverlay('hello');
const app = makeApp('claude', overlay);
app._flushLocalEchoPending();
expect(app._pendingInput).toBe('hello');
expect(overlay.cleared).toBe(1);
expect(overlay.suppressed).toBe(1);
expect(app._flushedOffsets!.has('s1')).toBe(false);
expect(app._flushedTexts!.has('s1')).toBe(false);
});
it('appends nothing when the overlay is empty', () => {
const app = makeApp('claude', makeOverlay(''));
app._flushLocalEchoPending();
expect(app._pendingInput).toBe('');
});
});
describe('insertTerminalText pass-through routing', () => {
it('appends to the overlay while local echo is buffering', () => {
const overlay = makeOverlay();
const app = makeApp('claude', overlay);
app._localEchoEnabled = true;
app.insertTerminalText('path.txt');
expect(overlay.appendText).toHaveBeenCalledWith('path.txt');
expect(app.sendInput).not.toHaveBeenCalled();
});
it('sends directly while the session is in nav-key pass-through', () => {
const overlay = makeOverlay();
const app = makeApp('claude', overlay);
app._localEchoEnabled = true;
app._echoPassthroughSessions = new Set(['s1']);
app.insertTerminalText('path.txt');
expect(app.sendInput).toHaveBeenCalledWith('path.txt');
expect(overlay.appendText).not.toHaveBeenCalled();
});
});
+140
View File
@@ -323,6 +323,146 @@ describe('Virtual Keyboard', () => {
expect(mainPadding).toBe('');
});
it('coalesces keyboard animation frames into one final terminal fit', async () => {
const result = await page.evaluate(async () => {
// `app.terminal` and `app.fitAddon` are only assigned by initTerminal(),
// which needs a selected session this harness never creates. Both are
// null at rest, and the settle callback returns early on a falsy
// terminal — so without stand-ins this test cannot reach the behavior
// it asserts. Install the minimum surface the callback touches.
const hadTerminal = app.terminal !== null && app.terminal !== undefined;
const hadFitAddon = app.fitAddon !== null && app.fitAddon !== undefined;
if (!hadTerminal) app.terminal = { scrollToBottom() {} };
if (!hadFitAddon) app.fitAddon = { fit() {}, proposeDimensions: () => null };
const originalFit = app.fitAddon.fit.bind(app.fitAddon);
const originalSendResize = KeyboardHandler._sendTerminalResize.bind(KeyboardHandler);
const originalScrollToBottom = app.terminal.scrollToBottom.bind(app.terminal);
let fits = 0;
let resizes = 0;
let bottomRestores = 0;
app.fitAddon.fit = () => {
fits++;
};
KeyboardHandler._sendTerminalResize = () => {
resizes++;
};
app.terminal.scrollToBottom = () => {
bottomRestores++;
};
KeyboardHandler._scheduleViewportSettle({ scrollToBottom: true });
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler._scheduleViewportSettle();
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler._scheduleViewportSettle();
await new Promise((resolve) => setTimeout(resolve, 50));
const beforeFinalSettle = { fits, resizes, bottomRestores };
await new Promise((resolve) => setTimeout(resolve, KeyboardHandler.VIEWPORT_SETTLE_MS));
const afterFinalSettle = { fits, resizes, bottomRestores };
app.fitAddon.fit = originalFit;
KeyboardHandler._sendTerminalResize = originalSendResize;
app.terminal.scrollToBottom = originalScrollToBottom;
if (!hadFitAddon) app.fitAddon = null;
if (!hadTerminal) app.terminal = null;
return { beforeFinalSettle, afterFinalSettle };
});
expect(result.beforeFinalSettle).toEqual({ fits: 0, resizes: 0, bottomRestores: 0 });
expect(result.afterFinalSettle).toEqual({ fits: 1, resizes: 1, bottomRestores: 1 });
});
// Behavioral counterpart to the test above, driven through the PUBLIC entry
// point rather than the internal scheduler. Before this change each
// onKeyboardShow armed its own uncoalesced 150ms setTimeout, so a keyboard
// animation that reports several viewport steps refit the terminal once per
// step — the visible symptom being repeated reflow while the keyboard slides
// up. This asserts the observable outcome (one refit for a burst) and so
// fails on master by COUNT, not by a missing method.
it('refits once for a burst of keyboard viewport steps', async () => {
const counts = await page.evaluate(async () => {
const hadTerminal = app.terminal !== null && app.terminal !== undefined;
const hadFitAddon = app.fitAddon !== null && app.fitAddon !== undefined;
if (!hadTerminal) app.terminal = { scrollToBottom() {} };
if (!hadFitAddon) app.fitAddon = { fit() {}, proposeDimensions: () => null };
const originalFit = app.fitAddon.fit.bind(app.fitAddon);
const originalSendResize = KeyboardHandler._sendTerminalResize.bind(KeyboardHandler);
let fits = 0;
app.fitAddon.fit = () => {
fits++;
};
KeyboardHandler._sendTerminalResize = () => {};
// Three viewport steps in quick succession, as a keyboard animation
// produces on a real device.
KeyboardHandler.onKeyboardShow();
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler.onKeyboardShow();
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler.onKeyboardShow();
// Well past both the coalescing window and master's fixed 150ms timer.
await new Promise((resolve) => setTimeout(resolve, 400));
app.fitAddon.fit = originalFit;
KeyboardHandler._sendTerminalResize = originalSendResize;
if (!hadFitAddon) app.fitAddon = null;
if (!hadTerminal) app.terminal = null;
return fits;
});
// Coalesced: one refit for the whole burst. Master fires one per step.
expect(counts).toBe(1);
});
// A viewport resize with NO pending show/hide transition must not arm settle
// work of its own: keyboard detection can miss a fine-grained OS animation
// entirely (sub-150px steps with the baseline chasing the animation), and a
// fit against that mid-animation, uncompensated layout resizes the PTY to
// transient dims. The resulting SIGWINCH thrash duplicates prompts and
// garbles the transcript. Wiggles may only push a pending settle back.
it('does not refit on viewport wiggles without a keyboard transition', async () => {
const result = await page.evaluate(async () => {
const hadTerminal = app.terminal !== null && app.terminal !== undefined;
const hadFitAddon = app.fitAddon !== null && app.fitAddon !== undefined;
if (!hadTerminal) app.terminal = { scrollToBottom() {} };
if (!hadFitAddon) app.fitAddon = { fit() {}, proposeDimensions: () => null };
const originalFit = app.fitAddon.fit.bind(app.fitAddon);
const originalSendResize = KeyboardHandler._sendTerminalResize.bind(KeyboardHandler);
let fits = 0;
app.fitAddon.fit = () => {
fits++;
};
KeyboardHandler._sendTerminalResize = () => {};
// Wiggle only: nothing pending, so nothing may fire.
KeyboardHandler._deferViewportSettle();
KeyboardHandler._deferViewportSettle();
await new Promise((resolve) => setTimeout(resolve, KeyboardHandler.VIEWPORT_SETTLE_MS + 80));
const wiggleOnly = fits;
// A real transition arms the work; a following wiggle defers it but the
// settle still fires exactly once.
KeyboardHandler._scheduleViewportSettle({ scrollToBottom: true });
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler._deferViewportSettle();
await new Promise((resolve) => setTimeout(resolve, KeyboardHandler.VIEWPORT_SETTLE_MS + 80));
const afterTransition = fits;
app.fitAddon.fit = originalFit;
KeyboardHandler._sendTerminalResize = originalSendResize;
if (!hadFitAddon) app.fitAddon = null;
if (!hadTerminal) app.terminal = null;
return { wiggleOnly, afterTransition };
});
expect(result.wiggleOnly).toBe(0);
expect(result.afterTransition).toBe(1);
});
it('accessory bar has the simple-mode action buttons', async () => {
const actions = await page.evaluate(() => {
return Array.from(document.querySelectorAll('.keyboard-accessory-bar [data-action]')).map(
+1
View File
@@ -86,6 +86,7 @@ export function createMockRouteContext(options?: { sessionId?: string }) {
getModelConfig: vi.fn(async () => null),
getClaudeModeConfig: vi.fn(async () => ({})),
getTerminalHistoryConfig: vi.fn(async () => resolveTerminalHistoryConfig({})),
getAgentSkillEnabled: vi.fn(async () => false),
getDefaultClaudeMdPath: vi.fn(async () => undefined),
getLightState: vi.fn(() => ({ sessions: [], status: 'ok' })),
getLightSessionsState: vi.fn(() => {
+36 -5
View File
@@ -4,6 +4,7 @@
*/
import { EventEmitter } from 'node:events';
import { vi } from 'vitest';
import type { SessionStatus } from '../../src/types.js';
/**
* Enhanced mock session for testing RespawnController.
@@ -12,8 +13,15 @@ import { vi } from 'vitest';
export class MockSession extends EventEmitter {
id: string;
workingDir: string = '/tmp/test-workdir';
status: 'idle' | 'working' = 'idle';
pid: number = 12345;
/**
* The REAL union, deliberately. This used to be `'idle' | 'working'`, and
* `'working'` is not a `SessionStatus` at all — so `signalForStatus()` fell to its
* `default: null` branch in every route test and the busy / stopped / error halves
* of the immediate-resolve mapping had zero coverage while appearing to be tested.
*/
status: SessionStatus = 'idle';
/** `null` once the PTY is gone (or before it has ever started) — see `pid` in Session. */
pid: number | null = 12345;
isWorking: boolean = false;
private _activeChildProcesses: { pid: number; command: string }[] = [];
ralphTracker: null = null;
@@ -30,13 +38,22 @@ export class MockSession extends EventEmitter {
this._muxName = `codeman-test-${id.slice(0, 8)}`;
}
/** Direct PTY write (used by session.write()) */
write(data: string): void {
/**
* Set to simulate a session whose PTY is gone: both write paths report failure,
* which is the state in which input used to disappear silently.
*/
failWrites = false;
/** Direct PTY write (used by session.write()). Mirrors the real boolean return. */
write(data: string): boolean {
if (this.failWrites) return false;
this.writeBuffer.push(data);
return true;
}
/** Write via mux (used by respawn controller) */
async writeViaMux(data: string): Promise<boolean> {
if (this.failWrites) return false;
this.writeBuffer.push(data);
return true;
}
@@ -44,6 +61,10 @@ export class MockSession extends EventEmitter {
/** Exactly-once input dedup — mirrors Session.shouldApplyInput so route tests
* exercising the reliable-delivery path behave like production. */
private _appliedInputSeq = new Map<string, number>();
forgetInputSeq(clientId: string, seq: number): void {
if (this._appliedInputSeq.get(clientId) === seq) this._appliedInputSeq.set(clientId, seq - 1);
}
shouldApplyInput(clientId: string, seq: number): boolean {
const last = this._appliedInputSeq.get(clientId);
if (last !== undefined && seq <= last) return false;
@@ -100,7 +121,9 @@ export class MockSession extends EventEmitter {
/** Simulate working state with spinner */
simulateWorking(text: string = 'Thinking'): void {
this.simulateTerminalOutput(`${text}... \u280b`);
this.status = 'working';
// 'busy' is what the real Session sets while a turn is in flight; the old
// 'working' here was the event name, not a status value.
this.status = 'busy';
this.emit('working');
}
@@ -169,6 +192,14 @@ export class MockSession extends EventEmitter {
return this._muxName;
}
/**
* Mirrors `Session.usesMux`. True by default because that is the normal
* configuration, and it is what makes a route's pane-liveness probe reachable:
* `session.pid` is the tmux ATTACH CLIENT, so a mux-backed session's worker can be
* dead while `pid` is still a live number.
*/
usesMux: boolean = true;
/** Check for active child processes (mock returns configurable list) */
getActiveChildProcesses(): { pid: number; command: string }[] {
return this._activeChildProcesses;
+200
View File
@@ -0,0 +1,200 @@
/**
* Regression guard for the process-tree walk.
*
* On 2026-07-30 an unbounded version took a machine down: it ran `pgrep -P <pid>` per
* node and recursed with no visited set, no depth limit and no node cap. Across ~28
* adopted tmux trees the fan-out exploded, and because each `pgrep` blocks in the WSL
* kernel while reading /proc/<pid>/cgroup, none returned while the walk kept firing
* more. Result: ~13,000 `pgrep` processes stuck in D-state out of ~39,000 total, load
* average above 13,000, recoverable only by
* restarting WSL — which cost every running session.
*
* These tests exercise the SHIPPED function. An earlier version of this file carried
* its own copy of the traversal, which would have passed happily while the real code
* regressed; that is why the walk now lives in its own module.
*/
import { describe, expect, it, vi } from 'vitest';
import { collectDescendants, PROC_WALK_MAX_DEPTH, PROC_WALK_MAX_NODES } from '../src/proc-tree.js';
/** Build a parent→children map from `[parent, child]` pairs. */
function tree(pairs: [number, number][]): Map<number, number[]> {
const m = new Map<number, number[]>();
for (const [p, c] of pairs) m.set(p, [...(m.get(p) ?? []), c]);
return m;
}
/** A chain 1→2→…→n, i.e. depth n-1. */
function chain(n: number): Map<number, number[]> {
return tree(Array.from({ length: n - 1 }, (_, i) => [i + 1, i + 2] as [number, number]));
}
describe('collectDescendants', () => {
it('returns every descendant of a normal tree, root excluded', () => {
const t = tree([
[1, 2],
[1, 3],
[2, 4],
[3, 5],
]);
expect(collectDescendants(1, t).sort((a, b) => a - b)).toEqual([2, 3, 4, 5]);
});
it('terminates on a cycle instead of looping forever', () => {
// A live `ps` snapshot is not atomic; pid reuse can produce a parent loop.
const t = tree([
[1, 2],
[2, 3],
[3, 1],
[3, 2],
]);
expect(collectDescendants(1, t).sort((a, b) => a - b)).toEqual([2, 3]);
});
it('does not include the root even when something claims it as a child', () => {
expect(collectDescendants(1, tree([[1, 1]]))).toEqual([]);
});
it('caps the depth, and says so', () => {
// 40 generations available, only PROC_WALK_MAX_DEPTH may be descended. A silent
// depth cap hides a deep tree exactly as a silent node cap hides a wide one.
const onTruncated = vi.fn();
expect(collectDescendants(1, chain(40), { onTruncated })).toHaveLength(PROC_WALK_MAX_DEPTH);
expect(onTruncated).toHaveBeenCalledWith(1, PROC_WALK_MAX_DEPTH, 'depth');
});
it('stays silent about depth when the tree ends inside the cap', () => {
const onTruncated = vi.fn();
collectDescendants(1, chain(4), { onTruncated });
expect(onTruncated).not.toHaveBeenCalled();
});
it('caps the node count and reports the truncation', () => {
// One parent with far more children than the cap allows.
const wide = new Map<number, number[]>([[1, Array.from({ length: PROC_WALK_MAX_NODES * 3 }, (_, i) => i + 2)]]);
const onTruncated = vi.fn();
const out = collectDescendants(1, wide, { onTruncated });
expect(out).toHaveLength(PROC_WALK_MAX_NODES);
expect(onTruncated).toHaveBeenCalledWith(1, PROC_WALK_MAX_NODES, 'nodes');
});
it('stays silent when nothing was truncated', () => {
const onTruncated = vi.fn();
collectDescendants(1, tree([[1, 2]]), { onTruncated });
expect(onTruncated).not.toHaveBeenCalled();
});
it('survives the shape that caused the incident: many wide, deep trees', () => {
// 28 adopted trees, branching 4-wide. Depth 6 already gives 4096 nodes per tree —
// eight times the cap, which is what this asserts. (The real incident's trees were
// deeper still; building that here would mean materialising 16M map entries and
// would only test the fixture builder.)
const t = new Map<number, number[]>();
let next = 1000;
const roots: number[] = [];
for (let r = 0; r < 28; r += 1) {
const root = next++;
roots.push(root);
let frontier = [root];
for (let d = 0; d < 6; d += 1) {
const nf: number[] = [];
for (const p of frontier) {
const kids = [next++, next++, next++, next++];
t.set(p, kids);
nf.push(...kids);
}
frontier = nf;
}
}
for (const root of roots) {
const out = collectDescendants(root, t);
expect(out.length).toBeLessThanOrEqual(PROC_WALK_MAX_NODES);
}
});
it('honours explicit overrides', () => {
expect(collectDescendants(1, chain(40), { maxDepth: 3 })).toEqual([2, 3, 4]);
expect(collectDescendants(1, chain(40), { maxNodes: 2 })).toEqual([2, 3]);
});
it('returns nothing for an unknown pid or an empty snapshot', () => {
expect(collectDescendants(999, tree([[1, 2]]))).toEqual([]);
expect(collectDescendants(1, new Map())).toEqual([]);
});
});
/**
* The bound must be reachable through the code that actually kills things.
*
* The unit tests above exercise `collectDescendants` directly, which is necessary but
* not sufficient: reverting `tmux-manager.ts` to the old unbounded `pgrep -P` recursion
* left every one of them green. This asserts the wiring — that TmuxManager's descendant
* lookup goes through the bounded walk and honours its caps.
*
* The snapshot refresh is stubbed. Without that the manager runs a real `ps` and
* replaces the fixture, and the test silently measures the machine's own process tree
* instead of the tree under test — which is how the first version of this test passed
* even with both caps bypassed.
*/
describe('TmuxManager uses the bounded walk', () => {
/** Build a tree `width` wide and `depth` deep, rooted at 1. */
function bigTree(width: number, depth: number): Map<number, number[]> {
const t = new Map<number, number[]>();
let next = 2;
let frontier = [1];
for (let d = 0; d < depth; d += 1) {
const nf: number[] = [];
for (const p of frontier) {
const kids = Array.from({ length: width }, () => next++);
t.set(p, kids);
nf.push(...kids);
}
frontier = nf;
}
return t;
}
async function walkVia(fixture: Map<number, number[]>): Promise<number[]> {
const { TmuxManager } = await import('../src/tmux-manager.js');
const Klass = TmuxManager as unknown as {
refreshProcSnapshot(): Promise<Map<number, number[]>>;
procSnapshot: unknown;
};
const original = Klass.refreshProcSnapshot;
Klass.refreshProcSnapshot = () => Promise.resolve(fixture);
Klass.procSnapshot = { at: Date.now(), byParent: fixture };
try {
const mgr = new TmuxManager();
return await (mgr as unknown as { getChildPidsFresh(pid: number): Promise<number[]> }).getChildPidsFresh(1);
} finally {
Klass.refreshProcSnapshot = original;
Klass.procSnapshot = null;
}
}
it('honours the node cap on a tree far wider than it', async () => {
// 4096 descendants available; the cap is 500. With the caps bypassed — the shape
// a regression at the call site would take — this returns thousands.
const out = await walkVia(bigTree(4, 6));
expect(out.length).toBe(PROC_WALK_MAX_NODES);
});
it('honours the depth cap on a deep chain', async () => {
const chainTree = new Map<number, number[]>();
for (let i = 1; i < 40; i += 1) chainTree.set(i, [i + 1]);
const out = await walkVia(chainTree);
expect(out.length).toBe(PROC_WALK_MAX_DEPTH);
});
it('spawns nothing per node — the walk only reads the snapshot', async () => {
// A per-node spawn against this fixture would mean thousands of processes; the
// test completing at all is the assertion, plus the bound holding.
const out = await walkVia(bigTree(4, 6));
expect(out.length).toBeLessThanOrEqual(PROC_WALK_MAX_NODES);
});
});
+80
View File
@@ -332,3 +332,83 @@ describe('Case Management', () => {
});
});
});
describe('Agent skill injection (agentSkillEnabled)', () => {
let server: WebServer;
let baseUrl: string;
const createdCases: string[] = [];
beforeAll(async () => {
server = await createTestServer(TEST_PORT + 4); // 3103
await server.start();
baseUrl = `http://localhost:${TEST_PORT + 4}`;
});
afterAll(async () => {
await server.stop();
for (const caseName of createdCases) {
const casePath = join(CASES_DIR, caseName);
if (existsSync(casePath)) {
rmSync(casePath, { recursive: true, force: true });
}
}
});
it('does not inject by default, accepts the setting via PUT, then injects on quick-start', async () => {
// 1. Default OFF: a claude quick-start creates the case without the skill.
const offCase = 'test-skill-off-' + Date.now();
createdCases.push(offCase);
const offResponse = await fetch(`${baseUrl}/api/quick-start`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ caseName: offCase }),
});
const offData = await offResponse.json();
expect(offData.success).toBe(true);
expect(existsSync(join(CASES_DIR, offCase, '.claude', 'skills', 'codeman'))).toBe(false);
// 2. The `.strict()` settings schema accepts the new synced key.
const putResponse = await fetch(`${baseUrl}/api/settings`, {
method: 'PUT',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ agentSkillEnabled: true }),
});
const putData = await putResponse.json();
expect(putData.success).toBe(true);
// 3. The server's settings read is cached ~2s; outwait it so the create sees the toggle.
await new Promise((resolve) => setTimeout(resolve, 2100));
// 4. Quick-start now injects the marker-carrying skill into the new case.
const onCase = 'test-skill-on-' + Date.now();
createdCases.push(onCase);
const onResponse = await fetch(`${baseUrl}/api/quick-start`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ caseName: onCase }),
});
const onData = await onResponse.json();
expect(onData.success).toBe(true);
const skillDir = join(CASES_DIR, onCase, '.claude', 'skills', 'codeman');
const { readFileSync } = await import('node:fs');
const skillMd = readFileSync(join(skillDir, 'SKILL.md'), 'utf-8');
expect(skillMd.startsWith('---\nname: codeman')).toBe(true);
expect(skillMd).toContain('<!-- codeman-managed-agent-skill');
expect(existsSync(join(skillDir, 'reference', 'endpoints.md'))).toBe(true);
expect(existsSync(join(skillDir, 'reference', 'recipes.md'))).toBe(true);
}, 30000);
it('does not inject for shell-mode quick-start even when enabled', async () => {
const shellCase = 'test-skill-shell-' + Date.now();
createdCases.push(shellCase);
const response = await fetch(`${baseUrl}/api/quick-start`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ caseName: shellCase, mode: 'shell' }),
});
const data = await response.json();
expect(data.success).toBe(true);
expect(existsSync(join(CASES_DIR, shellCase, '.claude', 'skills', 'codeman'))).toBe(false);
});
});
+157
View File
@@ -0,0 +1,157 @@
/**
* @fileoverview A lost input must stay retryable.
*
* POST /api/sessions/:id/input answers 200 BEFORE the write is attempted — the mux
* write is fire-and-forget so the HTTP response never waits on a tmux child. The
* dedup bookkeeping, however, recorded the (clientId, seq) pair as applied at that
* same moment. A write that then failed left the client with a 200, no message in
* the pane, and a seq the server would reject as a duplicate on retry: the input was
* unrecoverable by the very mechanism meant to make delivery reliable.
*
* Observed in the wild: a prompt shown as sent in a chat client, a 200 in the proxy
* log, and an empty prompt line in the pane.
*/
import fastifyCookie from '@fastify/cookie';
import Fastify, { type FastifyInstance } from 'fastify';
import { afterEach, beforeEach, describe, expect, it } from 'vitest';
import { Session } from '../../src/session.js';
import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
async function createEnvelopeHarness(): Promise<{ app: FastifyInstance; ctx: MockRouteContext }> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
const ctx = createMockRouteContext();
registerSessionRoutes(app, ctx as never);
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
if (!req.url.startsWith('/api')) return done(null, payload);
if (payload === null || typeof payload !== 'object') return done(null, payload);
const p = payload as { success?: unknown; errorCode?: unknown };
if (p.success === false) {
if (reply.statusCode === 200 && typeof p.errorCode === 'string') {
reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
}
return done(null, payload);
}
if (p.success === true) return done(null, payload);
return done(null, { success: true, data: payload });
});
installRouteErrorHandler(app);
await app.ready();
return { app, ctx };
}
type Internals = { _appliedInputSeq: Map<string, number> };
const seqOf = (s: Session, client: string) => (s as unknown as Internals)._appliedInputSeq.get(client);
describe('input dedup bookkeeping', () => {
const make = () => new Session({ workingDir: '/tmp', mode: 'claude' });
it('accepts an increasing seq once and rejects the replay', () => {
const s = make();
expect(s.shouldApplyInput('c1', 1)).toBe(true);
expect(s.shouldApplyInput('c1', 1)).toBe(false);
expect(s.shouldApplyInput('c1', 2)).toBe(true);
});
it('forgetInputSeq re-opens a failed delivery for retry', () => {
const s = make();
expect(s.shouldApplyInput('c1', 7)).toBe(true);
s.forgetInputSeq('c1', 7); // the write failed after the 200 went out
expect(s.shouldApplyInput('c1', 7)).toBe(true);
});
it('does not re-open a seq that a later input has superseded', () => {
// Rolling back blindly would let an old, already-superseded message replay.
const s = make();
s.shouldApplyInput('c1', 7);
s.shouldApplyInput('c1', 8);
s.forgetInputSeq('c1', 7);
expect(s.shouldApplyInput('c1', 8)).toBe(false);
expect(seqOf(s, 'c1')).toBe(8);
});
it('is scoped per client', () => {
const s = make();
s.shouldApplyInput('c1', 5);
s.forgetInputSeq('c2', 5);
expect(s.shouldApplyInput('c1', 5)).toBe(false);
});
it('tolerates a rollback for a client that was never seen', () => {
const s = make();
expect(() => s.forgetInputSeq('ghost', 3)).not.toThrow();
});
});
describe('Session.write delivery signal', () => {
it('reports false when there is no PTY instead of swallowing the data', () => {
// The silent swallow was the third way input could vanish: no PTY, no error,
// no return value — the caller had no way to know.
const s = new Session({ workingDir: '/tmp', mode: 'claude' });
expect(s.write('hello\r')).toBe(false);
});
});
/**
* Wiring, not primitives.
*
* The first version of this file tested Session directly and nothing else: reverting
* the route to master — deleting the rollback call, the load-bearing half of the fix —
* left all six tests green. These drive the actual HTTP route.
*/
describe('POST /api/sessions/:id/input rollback wiring', () => {
let harness: { app: FastifyInstance; ctx: MockRouteContext };
beforeEach(async () => {
harness = await createEnvelopeHarness();
});
afterEach(async () => {
await harness.app.close();
});
const post = (body: Record<string, unknown>, id = 'test-session-1') =>
harness.app.inject({ method: 'POST', url: `/api/sessions/${id}/input`, payload: body });
it('rolls the seq back when both the mux write and the direct write fail', async () => {
const session = harness.ctx.sessions.get('test-session-1')!;
session.failWrites = true; // writeViaMux false AND write() false
await post({ input: 'lost\r', useMux: true, clientId: 'c1', seq: 1 });
await new Promise((r) => setTimeout(r, 20)); // the mux write is fire-and-forget
// The retry the client would make must be accepted, not swallowed as a duplicate.
expect(session.shouldApplyInput('c1', 1)).toBe(true);
});
it('keeps the seq burnt when delivery succeeded', async () => {
const session = harness.ctx.sessions.get('test-session-1')!;
await post({ input: 'fine\r', useMux: true, clientId: 'c1', seq: 1 });
await new Promise((r) => setTimeout(r, 20));
expect(session.shouldApplyInput('c1', 1)).toBe(false);
});
it('rolls the seq back on the non-mux path too', async () => {
// Still a 200: a session may legitimately have no PTY yet, and turning that
// into a failure status would be a contract change. Re-opening the seq is not.
const session = harness.ctx.sessions.get('test-session-1')!;
session.failWrites = true;
const res = await post({ input: 'x\r', useMux: false, clientId: 'c2', seq: 5 });
expect(res.statusCode).toBe(200);
expect(session.shouldApplyInput('c2', 5)).toBe(true);
});
});
+671
View File
@@ -0,0 +1,671 @@
/**
* @fileoverview Route tests for the `wait` field on `POST /api/sessions/:id/input`.
*
* This endpoint exists to close a race a caller cannot close from outside: between
* the write landing and the session flipping to `working`, a SEPARATE wait sees the
* session still idle and returns instantly, reporting the previous turn as this one.
* Registering the waiter before the write is the whole point, so that is what the
* first test pins.
*
* It also pins that the historical fire-and-forget path is untouched when `wait` is
* absent, that a capacity rejection gives the dedup seq back (otherwise the caller's
* retry is refused as a duplicate and the input is lost by the very mechanism
* reliable delivery exists for), and that `delivered` reports what actually happened
* to the write rather than merely "this was not a duplicate".
*
* Plan: docs/agent-control-plan.md
*/
import { describe, it, expect, afterEach, beforeAll, afterAll, vi } from 'vitest';
import fastifyCookie from '@fastify/cookie';
import Fastify, { type FastifyInstance } from 'fastify';
import type { IncomingMessage } from 'node:http';
import { registerSessionRoutes, _resetPaneLivenessState } from '../../src/web/routes/session-routes.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js';
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
import { sessionWaits } from '../../src/web/session-wait-registry.js';
import { MAX_WAIT_MS } from '../../src/config/agent-wait.js';
// Distinct per file on purpose: the three wait suites share the process-wide
// `sessionWaits` singleton, so a common id let one file's leftover waiter be counted
// by another's assertion. Failed only in a 5-file run, which is how CI runs them.
const SESSION_ID = 'input-wait-session';
const URL = `/api/sessions/${SESSION_ID}/input`;
afterEach(() => {
// Deliberately not `cancelEverything()`: it latches the registry's stopped flag,
// which would leave every later test in this file talking to a dead registry.
sessionWaits.cancelAll(SESSION_ID);
_resetPaneLivenessState();
});
async function harness(): Promise<{ app: FastifyInstance; ctx: MockRouteContext; rawRequests: IncomingMessage[] }> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
const ctx = createMockRouteContext({ sessionId: SESSION_ID });
const rawRequests: IncomingMessage[] = [];
app.addHook('onRequest', async (req) => {
rawRequests.push(req.raw);
});
registerSessionRoutes(app, ctx as never);
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
const p = payload as { success?: unknown; errorCode?: unknown } | null;
if (p && typeof p === 'object' && p.success === false && reply.statusCode === 200) {
if (typeof p.errorCode === 'string') reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
}
return done(null, payload);
});
installRouteErrorHandler(app);
await app.ready();
return { app, ctx, rawRequests };
}
const send = (app: FastifyInstance, payload: Record<string, unknown>) =>
app.inject({ method: 'POST', url: URL, payload });
describe('POST /api/sessions/:id/input without wait (unchanged behavior)', () => {
it('returns the historical bare body and registers no waiter', async () => {
const { app } = await harness();
const res = await send(app, { input: 'hello', useMux: true });
expect(res.statusCode).toBe(200);
expect(res.json()).toEqual({});
expect(sessionWaits.totalWaiterCount()).toBe(0);
});
it('still returns before the mux write settles', async () => {
// The fire-and-forget property is why the response is fast; send-and-wait must
// not have turned every input into an awaited tmux round-trip.
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
let resolveWrite: (ok: boolean) => void = () => {};
session.writeViaMux = () => new Promise<boolean>((resolve) => (resolveWrite = resolve));
const res = await send(app, { input: 'hello', useMux: true });
expect(res.json()).toEqual({});
resolveWrite(true);
});
it('a tagged duplicate still returns the bare body', async () => {
const { app } = await harness();
await send(app, { input: 'first', clientId: 'c1', seq: 1 });
const replay = await send(app, { input: 'first', clientId: 'c1', seq: 1 });
expect(replay.json()).toEqual({});
expect(sessionWaits.totalWaiterCount()).toBe(0);
});
it('a failed direct write is still not an error response', async () => {
// A session can legitimately have no PTY yet, and callers have always been able
// to write to one without a 4xx. Only the `wait` path reports delivery.
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.failWrites = true;
const res = await send(app, { input: 'x' });
expect(res.statusCode).toBe(200);
expect(res.json()).toEqual({});
});
});
describe('POST /api/sessions/:id/input with wait', () => {
it('registers the waiter BEFORE the write, so the pre-existing idle state cannot satisfy it', async () => {
// The mock session is idle. A naive send-then-wait would answer immediately with
// that stale idle; this must block until a real transition.
const { app } = await harness();
const pending = send(app, { input: 'run the tests', useMux: true, wait: true });
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
sessionWaits.notifySignal(SESSION_ID, 'stop');
const body = (await pending).json();
expect(body.success).toBe(true);
expect(body.data.delivered).toBe(true);
expect(body.data.duplicate).toBe(false);
expect(body.data.wait.signal).toBe('stop');
expect(body.data.wait.immediate).toBe(false);
expect(body.data.wait.timedOut).toBe(false);
});
it('uses the same data.wait envelope as the two GET endpoints', async () => {
const { app } = await harness();
const pending = send(app, { input: 'x', wait: 'stop' });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifySignal(SESSION_ID, 'stop');
const { data } = (await pending).json();
expect(Object.keys(data).sort()).toEqual(['delivered', 'duplicate', 'limitPaused', 'status', 'wait']);
expect(data.status).toBe('idle');
expect(data.limitPaused).toBe(false);
expect(data.wait.aborted).toBe(false);
});
it('echoes the effective timeout after clamping', async () => {
// The schema accepts up to 24h; the server caps at MAX_WAIT_MS. Without the echo
// an agent reads the cap as "my 24h wait elapsed" and kills a healthy worker.
const { app } = await harness();
const pending = send(app, { input: 'x', wait: 'stop', waitTimeout: 86_400_000 });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifySignal(SESSION_ID, 'stop');
expect((await pending).json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
});
it('delivers the input before blocking', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
const pending = send(app, { input: 'echo hi', useMux: true, wait: 'stop' });
await new Promise((resolve) => setTimeout(resolve, 20));
// The write happened while the request is still open.
expect(session.writeBuffer.join('')).toContain('echo hi');
sessionWaits.notifySignal(SESSION_ID, 'stop');
await pending;
});
it('wait: true uses the default signal set', async () => {
const { app } = await harness();
const pending = send(app, { input: 'x', wait: true });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifySignal(SESSION_ID, 'idle');
expect((await pending).json().data.wait.until).toEqual(['stop', 'idle', 'exit']);
});
it('accepts an explicit signal list', async () => {
const { app } = await harness();
const pending = send(app, { input: 'x', wait: 'exit' });
await new Promise((resolve) => setTimeout(resolve, 20));
// Not one of the requested signals: the wait must not resolve on it.
expect(sessionWaits.notifySignal(SESSION_ID, 'idle')).toBe(0);
sessionWaits.notifySignal(SESSION_ID, 'exit');
const body = (await pending).json();
expect(body.data.wait.signal).toBe('exit');
expect(body.data.wait.until).toEqual(['exit']);
});
it('accepts an array, the same grammar the query parameter takes', async () => {
const { app } = await harness();
const pending = send(app, { input: 'x', wait: ['stop', 'exit'] });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifySignal(SESSION_ID, 'exit');
const body = (await pending).json();
expect(body.data.wait.until).toEqual(['stop', 'exit']);
expect(body.data.wait.signal).toBe('exit');
});
it('rejects an unknown wait signal without writing', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
const before = session.writeBuffer.length;
const res = await send(app, { input: 'x', wait: 'stpo' });
expect(res.statusCode).toBe(400);
expect(res.json().errorCode).toBe('INVALID_INPUT');
expect(session.writeBuffer.length).toBe(before);
});
it('rejects a hook-only signal for external CLI modes', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.mode = 'codex';
const res = await send(app, { input: 'x', wait: 'stop' });
expect(res.statusCode).toBe(400);
expect(res.json().error).toContain('codex');
});
it('rejects a hook-only signal for a shell session too', async () => {
// Not an external CLI, but a plain bash PTY installs no hooks either, so `stop`
// could only ever time out.
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.mode = 'shell';
const res = await send(app, { input: 'x', wait: 'stop' });
expect(res.statusCode).toBe(400);
expect(res.json().error).toContain('shell');
});
it('times out as a 200, like the standalone wait', async () => {
const { app } = await harness();
const res = await send(app, { input: 'x', wait: 'blocked', waitTimeout: 1 });
expect(res.statusCode).toBe(200);
const body = res.json();
expect(body.data.delivered).toBe(true);
expect(body.data.wait.timedOut).toBe(true);
expect(body.data.wait.signal).toBeNull();
});
it('resolves with ended when the session goes away mid-wait', async () => {
const { app } = await harness();
const pending = send(app, { input: 'x', wait: 'stop' });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.cancelAll(SESSION_ID);
expect((await pending).json().data.wait.ended).toBe(true);
});
it('a request-body close does NOT abort the wait', async () => {
// The regression this pins: on a POST the request stream closes as soon as the
// body has been read, well before the handler blocks. Treating that as a hang-up
// aborted every send-and-wait instantly. Hang-up handling itself is proven over
// real HTTP at the bottom of this file, because inject cannot produce a socket.
const { app, rawRequests } = await harness();
const pending = send(app, { input: 'x', wait: 'stop', waitTimeout: 600_000 });
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
rawRequests[rawRequests.length - 1].emit('close');
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
sessionWaits.notifySignal(SESSION_ID, 'stop');
const body = (await pending).json();
expect(body.data.wait.aborted).toBe(false);
expect(body.data.wait.signal).toBe('stop');
});
it('a duplicate skips the write but answers from current state instead of hanging', async () => {
// The original turn is long over, so requiring a fresh transition here would
// block a redelivery until timeout for no reason.
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
await send(app, { input: 'first', clientId: 'c1', seq: 1 });
const before = session.writeBuffer.length;
const replay = await send(app, { input: 'first', clientId: 'c1', seq: 1, wait: 'idle' });
const body = replay.json();
expect(body.data.duplicate).toBe(true);
expect(body.data.delivered).toBe(false);
expect(body.data.wait.immediate).toBe(true);
expect(body.data.wait.signal).toBe('idle');
expect(session.writeBuffer.length).toBe(before);
});
it('a duplicate on a BUSY session answers working, not idle', async () => {
// Previously unreachable: MockSession's status was 'working', which is not a
// SessionStatus, so signalForStatus fell through to null and this combination
// silently proved nothing.
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
await send(app, { input: 'first', clientId: 'c2', seq: 1 });
session.status = 'busy';
const replay = await send(app, { input: 'first', clientId: 'c2', seq: 1, wait: 'working' });
const body = replay.json();
expect(body.data.duplicate).toBe(true);
expect(body.data.wait.signal).toBe('working');
expect(body.data.wait.immediate).toBe(true);
expect(body.data.status).toBe('busy');
});
it('gives the dedup seq back when a full waiter pool rejects the request', async () => {
// Otherwise the caller's retry is refused as a duplicate and the input vanishes.
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
const pendings = [];
for (let i = 0; i < 16; i++) {
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` }));
}
await new Promise((resolve) => setTimeout(resolve, 40));
expect(sessionWaits.waiterCount(SESSION_ID)).toBe(16);
const rejected = await send(app, { input: 'x', clientId: 'c9', seq: 7, wait: true });
expect(rejected.statusCode).toBe(409);
expect(rejected.json().errorCode).toBe('SESSION_BUSY');
// Nothing written...
expect(session.writeBuffer.join('')).not.toContain('x');
sessionWaits.cancelAll(SESSION_ID);
await Promise.all(pendings);
// ...and the same seq is accepted on retry rather than treated as a replay.
const retry = await send(app, { input: 'x', clientId: 'c9', seq: 7 });
expect(retry.json()).toEqual({});
expect(session.writeBuffer.join('')).toContain('x');
});
it('a null wait is treated as absent, not as a validation error', async () => {
// Zod .optional() rejects null, and a third-party caller building the body with
// JSON.stringify keeps an explicit null on the wire.
const { app } = await harness();
const res = await send(app, { input: 'x', wait: null, waitTimeout: null });
expect(res.statusCode).toBe(200);
expect(res.json()).toEqual({});
});
it('wait: false is treated as absent', async () => {
const { app } = await harness();
const res = await send(app, { input: 'x', wait: false });
expect(res.statusCode).toBe(200);
expect(res.json()).toEqual({});
expect(sessionWaits.totalWaiterCount()).toBe(0);
});
it('falls back to a direct write when the mux write fails, and still waits', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
session.writeViaMux = async () => false;
const pending = send(app, { input: 'fallback me', useMux: true, wait: 'stop' });
await new Promise((resolve) => setTimeout(resolve, 20));
expect(session.writeBuffer.join('')).toContain('fallback me');
sessionWaits.notifySignal(SESSION_ID, 'stop');
const body = (await pending).json();
expect(body.data.wait.signal).toBe('stop');
// The fallback write succeeded, so the input really was delivered.
expect(body.data.delivered).toBe(true);
});
});
describe('POST /api/sessions/:id/input: delivered reports the write, not just the dedup', () => {
it('reports delivered:false when BOTH write paths fail, instead of claiming delivery', async () => {
// A worker whose PTY has exited fails writeViaMux AND write. Reporting
// "delivered, but it timed out" points the agent at waiting longer; the truth is
// "restart the worker".
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
session.failWrites = true;
const res = await send(app, { input: 'run the tests', useMux: true, wait: 'stop', waitTimeout: 600_000 });
const body = res.json();
expect(res.statusCode).toBe(200);
expect(body.data.delivered).toBe(false);
expect(body.data.duplicate).toBe(false);
});
it('does not block for the full timeout on an input it knows never landed', async () => {
// The waiter has to be registered before the write, so it exists by the time the
// failure is known; releasing it immediately is what keeps the caller from
// waiting ten minutes for a turn that cannot start.
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.failWrites = true;
const started = Date.now();
const body = (await send(app, { input: 'x', useMux: true, wait: 'stop', waitTimeout: 600_000 })).json();
expect(Date.now() - started).toBeLessThan(2_000);
expect(body.data.delivered).toBe(false);
expect(body.data.wait.timedOut).toBe(false);
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
});
it('reports delivered:false for a failed direct (non-mux) write too', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.failWrites = true;
const body = (await send(app, { input: 'x', wait: 'stop', waitTimeout: 600_000 })).json();
expect(body.data.delivered).toBe(false);
});
it('rolls the dedup seq back when the write failed, so a retry is not a duplicate', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
session.failWrites = true;
const first = (await send(app, { input: 'x', clientId: 'c3', seq: 4, wait: 'stop', waitTimeout: 600_000 })).json();
expect(first.data.delivered).toBe(false);
session.failWrites = false;
const retry = (await send(app, { input: 'x', clientId: 'c3', seq: 4, wait: 'stop', waitTimeout: 1 })).json();
expect(retry.data.duplicate).toBe(false);
expect(retry.data.delivered).toBe(true);
expect(session.writeBuffer.join('')).toContain('x');
});
});
/**
* Client-hang-up handling, over REAL HTTP.
*
* `app.inject()` never emits a `close` event at all, so the entire abort path is
* invisible to every other test in this file — and the failure it hides is not
* subtle. On a POST, `req.raw` emits `'close'` as soon as the request BODY finishes
* streaming, which happens before the handler blocks (+1ms, `aborted: false`) and is
* indistinguishable from a real hang-up at +0ms. A request-side abort listener
* therefore cancels every send-and-wait instantly: `POST .../input {wait:"exit",
* waitTimeout:10000}` came back in 23ms with `ended:true, aborted:true, waitedMs:6`,
* i.e. the feature was dead while all 27 inject-based tests above stayed green.
*
* GET survives a request-side listener because it has no body to finish, which is
* exactly why this regression needs a POST and a real socket to catch.
*/
describe('POST /api/sessions/:id/input over real HTTP: hang-up handling', () => {
const PORT = 3181;
const base = `http://127.0.0.1:${PORT}`;
let app: FastifyInstance;
beforeAll(async () => {
app = (await harness()).app;
await app.listen({ port: PORT, host: '127.0.0.1' });
});
afterAll(async () => {
// fetch keeps its sockets alive, and `app.close()` waits for idle connections,
// so without this the teardown hook times out.
app.server.closeAllConnections();
await app.close();
});
it('a send-and-wait that is NOT aborted blocks for its full timeout', async () => {
// The regression: this returned in ~20ms with aborted:true.
const started = Date.now();
const res = await fetch(`${base}/api/sessions/${SESSION_ID}/input`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ input: 'run the tests', wait: 'exit', waitTimeout: 2000 }),
});
const elapsed = Date.now() - started;
const body = await res.json();
expect(body.data.wait.aborted).toBe(false);
expect(body.data.wait.timedOut).toBe(true);
expect(body.data.wait.waitedMs).toBeGreaterThan(1500);
expect(elapsed).toBeGreaterThan(1500);
expect(body.data.delivered).toBe(true);
});
it('a send-and-wait aborted mid-flight frees its waiter', async () => {
const controller = new AbortController();
const pending = fetch(`${base}/api/sessions/${SESSION_ID}/input`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ input: 'x', wait: 'exit', waitTimeout: 600_000 }),
signal: controller.signal,
}).catch(() => 'aborted');
await new Promise((resolve) => setTimeout(resolve, 150));
// Still parked: the body finished streaming long ago, and that must not count.
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
controller.abort();
expect(await pending).toBe('aborted');
await new Promise((resolve) => setTimeout(resolve, 100));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
});
it('a GET wait behaves the same way on both counts', async () => {
const notAborted = await fetch(`${base}/api/sessions/${SESSION_ID}/wait?until=stop&timeout=1500`);
const body = await notAborted.json();
expect(body.data.wait.aborted).toBe(false);
expect(body.data.wait.timedOut).toBe(true);
const controller = new AbortController();
const pending = fetch(`${base}/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000`, {
signal: controller.signal,
}).catch(() => 'aborted');
await new Promise((resolve) => setTimeout(resolve, 100));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
controller.abort();
expect(await pending).toBe('aborted');
await new Promise((resolve) => setTimeout(resolve, 100));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
});
it('wait-output frees its waiter on hang-up too', async () => {
const controller = new AbortController();
const pending = fetch(`${base}/api/sessions/${SESSION_ID}/wait-output?match=NEVER&timeout=600000`, {
signal: controller.signal,
}).catch(() => 'aborted');
await new Promise((resolve) => setTimeout(resolve, 100));
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
controller.abort();
expect(await pending).toBe('aborted');
await new Promise((resolve) => setTimeout(resolve, 100));
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(0);
});
});
/**
* Send-and-wait against a tmux worker that has already died.
*
* `tmux send-keys` SUCCEEDS against a dead pane, so `writeViaMux` returns true and the
* old `delivered` was true for bytes written into a corpse — with `timedOut: true`
* alongside it, which tells an agent to wait longer when the truth is "restart the
* worker". Live: `pane_dead=1 status=42`, Codeman `pid=309406 status=idle`,
* `delivered: true`.
*/
describe('POST /api/sessions/:id/input: the pane is dead', () => {
function setPaneDead(ctx: MockRouteContext, dead: boolean) {
(ctx.mux as unknown as { isPaneDead: (n: string) => boolean }).isPaneDead = () => dead;
}
it('reports delivered:false even though the mux write "succeeded"', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
setPaneDead(ctx, true);
const res = await send(app, { input: 'run the tests', useMux: true, wait: 'stop', waitTimeout: 600_000 });
const body = res.json();
// The write itself did not fail — that is the whole trap.
expect(session.writeBuffer.join('')).toContain('run the tests');
expect(body.data.delivered).toBe(false);
expect(body.data.duplicate).toBe(false);
});
it('returns at once instead of blocking on a turn that cannot start', async () => {
const { app, ctx } = await harness();
setPaneDead(ctx, true);
const started = Date.now();
const body = (await send(app, { input: 'x', useMux: true, wait: 'stop', waitTimeout: 600_000 })).json();
expect(Date.now() - started).toBeLessThan(2_000);
expect(body.data.wait.timedOut).toBe(false);
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
});
it('rolls the dedup seq back, so a retry against a restarted worker is not a duplicate', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
setPaneDead(ctx, true);
const dead = (await send(app, { input: 'x', useMux: true, clientId: 'c7', seq: 3, wait: 'stop' })).json();
expect(dead.data.delivered).toBe(false);
// Worker restarted.
setPaneDead(ctx, false);
_resetPaneLivenessState();
const retry = (
await send(app, { input: 'x', useMux: true, clientId: 'c7', seq: 3, wait: 'stop', waitTimeout: 1 })
).json();
expect(retry.data.duplicate).toBe(false);
expect(retry.data.delivered).toBe(true);
expect(session.writeBuffer.join('')).toContain('x');
});
it('a live pane still reports delivered:true', async () => {
const { app, ctx } = await harness();
setPaneDead(ctx, false);
const pending = send(app, { input: 'x', useMux: true, wait: 'stop' });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifySignal(SESSION_ID, 'stop');
expect((await pending).json().data.delivered).toBe(true);
});
it('never probes tmux on the plain (non-wait) input path', async () => {
// The browser sends thousands of these per session; they must not exec tmux.
const { app, ctx } = await harness();
const probe = vi.fn(() => false);
(ctx.mux as unknown as { isPaneDead: (n: string) => boolean }).isPaneDead = probe as never;
await send(app, { input: 'hello', useMux: true });
await send(app, { input: 'hello again', useMux: true, clientId: 'c1', seq: 1 });
expect(probe).not.toHaveBeenCalled();
});
});
describe('POST /api/sessions/:id/input: `aborted` stays a client-side fact', () => {
it('reports aborted:false when the SERVER released the waiter after a failed delivery', async () => {
// api-reference guarantees a client never sees `aborted: true`, because it means
// "you hung up, nobody is reading this". The release below is the server giving up
// on a write that failed — and the client IS reading the response, so reporting
// `aborted: true` would both break that guarantee and hand an agent a second,
// contradictory reason for an outcome `delivered: false` already explains.
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.failWrites = true;
const body = (await send(app, { input: 'x', useMux: true, wait: 'stop', waitTimeout: 600_000 })).json();
expect(body.data.delivered).toBe(false);
expect(body.data.wait.aborted).toBe(false);
expect(body.data.wait.ended).toBe(true);
expect(body.data.wait.timedOut).toBe(false);
});
it('the same holds for a dead pane', async () => {
const { app, ctx } = await harness();
(ctx.mux as unknown as { isPaneDead: () => boolean }).isPaneDead = () => true;
const body = (await send(app, { input: 'x', useMux: true, wait: 'stop', waitTimeout: 600_000 })).json();
expect(body.data.wait.aborted).toBe(false);
});
});
describe('POST /api/sessions/:id/input: an oversized waitTimeout clamps', () => {
it('accepts a value above the old schema ceiling and reports the clamp', async () => {
const { app } = await harness();
const pending = send(app, { input: 'x', wait: 'stop', waitTimeout: 99_999_999_999 });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifySignal(SESSION_ID, 'stop');
const body = (await pending).json();
expect(body.data.wait.timeoutMs).toBe(MAX_WAIT_MS);
});
it('still rejects a non-integer or negative waitTimeout', async () => {
const { app } = await harness();
for (const value of [-1, 0, 1.5]) {
const res = await send(app, { input: 'x', wait: 'stop', waitTimeout: value });
expect(res.statusCode, `waitTimeout=${value}`).toBe(400);
}
});
});
@@ -0,0 +1,432 @@
/**
* @fileoverview Route tests for `GET /api/sessions/:id/wait-output`.
*
* Same 200-on-timeout contract and same `data.wait` envelope as `/wait`. The
* additional things pinned here:
* - matching is LITERAL, and a `regex` parameter is rejected rather than ignored,
* so an agent that assumed herdr's `--regex` cannot silently wait on the wrong thing;
* - `from=buffer` scans what already scrolled past, bounded to a tail of the buffer,
* and is charged against the waiter cap BEFORE it materializes that buffer;
* - a chunk-straddling match still resolves, since PTY chunking is arbitrary;
* - a client that hangs up frees its waiter instead of holding it to the timeout.
*
* Plan: docs/agent-control-plan.md
*/
import { describe, it, expect, afterEach, vi } from 'vitest';
import fastifyCookie from '@fastify/cookie';
import Fastify, { type FastifyInstance } from 'fastify';
import type { ServerResponse } from 'node:http';
import { registerSessionRoutes, _resetPaneLivenessState } from '../../src/web/routes/session-routes.js';
import { createSessionListeners, attachSessionListeners } from '../../src/web/session-listener-wiring.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js';
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
import { sessionWaits } from '../../src/web/session-wait-registry.js';
import { MAX_MATCH_LENGTH, MAX_BUFFER_SCAN_BYTES, MAX_WAIT_MS } from '../../src/config/agent-wait.js';
// Distinct per file on purpose: the three wait suites share the process-wide
// `sessionWaits` singleton, so a common id let one file's leftover waiter be counted
// by another's assertion. Failed only in a 5-file run, which is how CI runs them.
const SESSION_ID = 'wait-output-session';
const URL = `/api/sessions/${SESSION_ID}/wait-output`;
afterEach(() => {
// Deliberately not `cancelEverything()`: it latches the registry's stopped flag,
// which would leave every later test in this file talking to a dead registry.
sessionWaits.cancelAll(SESSION_ID);
_resetPaneLivenessState();
});
/** Mirrors production's errorCode-to-status mapping; without it negative cases pass vacuously. */
async function harness(): Promise<{ app: FastifyInstance; ctx: MockRouteContext; rawReplies: ServerResponse[] }> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
const ctx = createMockRouteContext({ sessionId: SESSION_ID });
const rawReplies: ServerResponse[] = [];
app.addHook('onRequest', async (req, reply) => {
// The RESPONSE, because that is what the handler's hang-up detection listens to.
rawReplies.push(reply.raw);
});
registerSessionRoutes(app, ctx as never);
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
const p = payload as { success?: unknown; errorCode?: unknown } | null;
if (p && typeof p === 'object' && p.success === false && reply.statusCode === 200) {
if (typeof p.errorCode === 'string') reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
}
return done(null, payload);
});
installRouteErrorHandler(app);
await app.ready();
return { app, ctx, rawReplies };
}
describe('GET /api/sessions/:id/wait-output', () => {
it('resolves when the string appears on the stream', async () => {
const { app } = await harness();
const pending = app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK` });
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
sessionWaits.notifyOutput(SESSION_ID, 'running tests...\nBUILD OK\n');
const body = (await pending).json();
expect(body.success).toBe(true);
expect(body.data.wait.matched).toBe(true);
expect(body.data.wait.immediate).toBe(false);
expect(body.data.wait.snippet).toContain('BUILD OK');
expect(body.data.wait.match).toBe('BUILD OK');
expect(body.data.sessionId).toBe(SESSION_ID);
});
it('uses the same data.wait envelope as /wait, so one client helper reads both', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `${URL}?match=never&timeout=1` });
const { data } = res.json();
expect(Object.keys(data).sort()).toEqual(['limitPaused', 'sessionId', 'status', 'wait']);
expect(data.matched).toBeUndefined();
expect(data.wait.matched).toBe(false);
expect(data.wait.aborted).toBe(false);
});
it('sends Cache-Control: no-store', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `${URL}?match=never&timeout=1` });
expect(res.headers['cache-control']).toBe('no-store');
});
it('echoes the effective timeout after clamping', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.terminalBuffer = 'BUILD OK\n';
const res = await app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK&from=buffer&timeout=1800000` });
expect(res.json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
});
it('strips ANSI before matching, so colored output still matches', async () => {
const { app } = await harness();
const pending = app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK` });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifyOutput(SESSION_ID, '\x1b[32mBUILD\x1b[0m OK\n');
expect((await pending).json().data.wait.matched).toBe(true);
});
it('matches across a chunk boundary', async () => {
const { app } = await harness();
const pending = app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK` });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifyOutput(SESSION_ID, 'trailing text BUIL');
sessionWaits.notifyOutput(SESSION_ID, 'D OK done');
expect((await pending).json().data.wait.matched).toBe(true);
});
it('is case-sensitive by default and honors nocase=1', async () => {
const { app } = await harness();
const strict = app.inject({ method: 'GET', url: `${URL}?match=build%20ok&timeout=1` });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifyOutput(SESSION_ID, 'BUILD OK');
expect((await strict).json().data.wait.timedOut).toBe(true);
const loose = app.inject({ method: 'GET', url: `${URL}?match=build%20ok&nocase=1` });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifyOutput(SESSION_ID, 'BUILD OK');
const body = (await loose).json();
expect(body.data.wait.matched).toBe(true);
// Reported in the terminal's own casing, not the caller's.
expect(body.data.wait.snippet).toContain('BUILD OK');
});
it('from=buffer resolves immediately against output that already scrolled past', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.terminalBuffer = 'earlier output\n\x1b[32mBUILD OK\x1b[0m\n';
const res = await app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK&from=buffer` });
const body = res.json();
expect(body.data.wait.matched).toBe(true);
expect(body.data.wait.immediate).toBe(true);
expect(body.data.wait.waitedMs).toBe(0);
expect(sessionWaits.totalWaiterCount()).toBe(0);
});
it('from=buffer only scans a bounded tail', async () => {
const { app, ctx } = await harness();
// Old marker pushed past the scan window by newer output.
ctx.sessions.get(SESSION_ID)!.terminalBuffer = `ANCIENT${'x'.repeat(MAX_BUFFER_SCAN_BYTES + 1000)}`;
const res = await app.inject({ method: 'GET', url: `${URL}?match=ANCIENT&from=buffer&timeout=1` });
expect(res.statusCode).toBe(200);
expect(res.json().data.wait.timedOut).toBe(true);
});
it('defaults to from=now, ignoring what is already in the buffer', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.terminalBuffer = 'BUILD OK happened before you asked\n';
const res = await app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK&timeout=1` });
expect(res.json().data.wait.timedOut).toBe(true);
});
it('answers 200 with timedOut on timeout, never an error status', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `${URL}?match=never&timeout=1` });
expect(res.statusCode).toBe(200);
const body = res.json();
expect(body.success).toBe(true);
expect(body.data.wait.timedOut).toBe(true);
expect(body.data.wait.matched).toBe(false);
expect(body.data.wait.snippet).toBeNull();
});
it('resolves with ended when the session goes away mid-wait', async () => {
const { app } = await harness();
const pending = app.inject({ method: 'GET', url: `${URL}?match=never` });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.cancelAll(SESSION_ID);
const body = (await pending).json();
expect(body.success).toBe(true);
expect(body.data.wait.ended).toBe(true);
expect(body.data.wait.matched).toBe(false);
});
it('frees its waiter when the RESPONSE socket closes early', async () => {
// The injected response is genuinely destroyed by then, so the freed slot is what
// this can assert; the wire-level behaviour is pinned over real HTTP in
// session-input-wait.test.ts.
const { app, rawReplies } = await harness();
const pending = app.inject({ method: 'GET', url: `${URL}?match=never&timeout=600000` }).then(
() => 'completed',
() => 'destroyed'
);
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
rawReplies[rawReplies.length - 1].emit('close');
await new Promise((resolve) => setTimeout(resolve, 10));
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(0);
expect(await pending).toBe('destroyed');
});
it('rejects a regex parameter instead of silently ignoring it', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `${URL}?match=x&regex=%5EBUILD.*OK%24` });
expect(res.statusCode).toBe(400);
const body = res.json();
expect(body.errorCode).toBe('INVALID_INPUT');
expect(body.error).toContain('match=');
});
it('requires match', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: URL });
expect(res.statusCode).toBe(400);
expect(res.json().errorCode).toBe('INVALID_INPUT');
});
it('rejects an empty or oversized match, and says which parameter was wrong', async () => {
const { app } = await harness();
const empty = await app.inject({ method: 'GET', url: `${URL}?match=` });
expect(empty.statusCode).toBe(400);
expect(empty.json().error).toContain('match');
const huge = await app.inject({ method: 'GET', url: `${URL}?match=${'x'.repeat(MAX_MATCH_LENGTH + 1)}` });
expect(huge.statusCode).toBe(400);
expect(huge.json().error).toContain('match');
});
it('rejects a non-numeric timeout', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `${URL}?match=x&timeout=soon` });
expect(res.statusCode).toBe(400);
expect(res.json().errorCode).toBe('INVALID_INPUT');
expect(res.json().error).toContain('timeout');
});
it('rejects an unknown from value', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `${URL}?match=x&from=history` });
expect(res.statusCode).toBe(400);
});
it('404s an unknown session', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: '/api/sessions/nope/wait-output?match=x' });
expect(res.statusCode).toBe(404);
expect(res.json().success).toBe(false);
});
it('output waiters share the session waiter cap with signal waiters', async () => {
const { app } = await harness();
const pendings = [];
for (let i = 0; i < 8; i++) {
pendings.push(app.inject({ method: 'GET', url: `${URL}?match=never${i}` }));
}
for (let i = 0; i < 8; i++) {
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` }));
}
await new Promise((resolve) => setTimeout(resolve, 40));
expect(sessionWaits.waiterCount(SESSION_ID)).toBe(16);
const overflow = await app.inject({ method: 'GET', url: `${URL}?match=one-too-many` });
expect(overflow.statusCode).toBe(409);
expect(overflow.json().errorCode).toBe('SESSION_BUSY');
sessionWaits.cancelAll(SESSION_ID);
await Promise.all(pendings);
});
it('checks the cap BEFORE materializing the terminal buffer', async () => {
// `session.terminalBuffer` joins the whole 32MB accumulator. Paying that for a
// request that is about to be refused turns the cap into an amplifier: a caller
// already at the limit can loop `from=buffer` at full speed and never register a
// waiter, so nothing bounds the work.
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
const bufferReads = vi.fn(() => 'nothing to see');
Object.defineProperty(session, 'terminalBuffer', { get: bufferReads, configurable: true });
const pendings = [];
for (let i = 0; i < 16; i++) {
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` }));
}
await new Promise((resolve) => setTimeout(resolve, 40));
expect(sessionWaits.waiterCount(SESSION_ID)).toBe(16);
const overflow = await app.inject({ method: 'GET', url: `${URL}?match=x&from=buffer` });
expect(overflow.json().errorCode).toBe('SESSION_BUSY');
expect(bufferReads).not.toHaveBeenCalled();
sessionWaits.cancelAll(SESSION_ID);
await Promise.all(pendings);
});
});
/**
* The LIVE-STREAM half, wired the way production wires it.
*
* Everything above drives `sessionWaits.notifyOutput()` directly, which is the
* registry's API, not the path a real session takes. Deleting the one line that
* connects them — `sessionWaits.notifyOutput(session.id, data)` in the `terminal`
* listener — left all four wait suites green and survived a full `test:ci` sweep, so
* `wait-output?from=now` (the entire live mode, and the one the skill's recipes are
* built on) could ship severed with nothing to show for it.
*
* These go through `createSessionListeners` so the wiring itself is what is pinned.
*/
describe('GET /api/sessions/:id/wait-output: fed by the real terminal listener', () => {
function stubDeps() {
return {
broadcast: vi.fn(),
batchTerminalData: vi.fn(),
batchTaskUpdate: vi.fn(),
broadcastSessionStateDebounced: vi.fn(),
sendPushNotifications: vi.fn(),
persistSessionState: vi.fn(),
getSessionStateWithRespawn: vi.fn(() => ({})),
getRunSummaryTracker: vi.fn(() => undefined),
stopTranscriptWatcher: vi.fn(),
cleanupSessionBatches: vi.fn(),
cancelPersistDebounce: vi.fn(),
removeRunSummaryTracker: vi.fn(),
removeSessionListenerRefs: vi.fn(),
cleanupRespawnOnExit: vi.fn(),
getStore: vi.fn(() => ({ updateRalphState: vi.fn() })),
registerAttachment: vi.fn(async () => {}),
};
}
it('matches output emitted by the session, not injected into the registry', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
const deps = stubDeps();
attachSessionListeners(session as never, createSessionListeners(session as never, deps as never));
const pending = app.inject({ method: 'GET', url: `${URL}?match=LIVE_STREAM_HIT` });
await new Promise((resolve) => setTimeout(resolve, 20));
// What a PTY chunk actually does: the session emits `terminal`.
session.simulateTerminalOutput('$ echo LIVE_STREAM_HIT\r\nLIVE_STREAM_HIT\r\n');
const body = (await pending).json();
expect(body.data.wait.matched).toBe(true);
expect(body.data.wait.snippet).toContain('LIVE_STREAM_HIT');
// The listener must still forward to the SSE batcher; the wait feed is additive.
expect(deps.batchTerminalData).toHaveBeenCalled();
});
it('matches across chunk boundaries through the listener', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
const pending = app.inject({ method: 'GET', url: `${URL}?match=SPLIT_MARKER` });
await new Promise((resolve) => setTimeout(resolve, 20));
session.simulateTerminalOutput('noise SPLIT_');
session.simulateTerminalOutput('MARKER more noise');
expect((await pending).json().data.wait.matched).toBe(true);
});
it('strips ANSI on the way through the listener', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
const pending = app.inject({ method: 'GET', url: `${URL}?match=COLORED%20HIT` });
await new Promise((resolve) => setTimeout(resolve, 20));
session.simulateAnsiOutput('COLORED HIT');
expect((await pending).json().data.wait.matched).toBe(true);
});
});
describe('GET /api/sessions/:id/wait-output: a dead tmux worker', () => {
it('releases an output waiter when the worker dies while it is parked', async () => {
// The feed simply stops: no exit event, no further chunks, nothing to match.
const { app, ctx } = await harness();
let dead = false;
(ctx.mux as unknown as { isPaneDead: () => boolean }).isPaneDead = () => dead;
const pending = app.inject({ method: 'GET', url: `${URL}?match=NEVER&timeout=600000` });
await new Promise((resolve) => setTimeout(resolve, 50));
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
dead = true;
const body = (await pending).json();
expect(body.data.wait.ended).toBe(true);
expect(body.data.wait.timedOut).toBe(false);
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(0);
}, 10_000);
it('an oversized timeout clamps rather than 400ing', async () => {
const { app, ctx } = await harness();
// Resolve from the buffer so the assertion is about the clamp, not a real wait.
ctx.sessions.get(SESSION_ID)!.terminalBuffer = 'ALREADY_THERE\n';
const res = await app.inject({
method: 'GET',
url: `${URL}?match=ALREADY_THERE&from=buffer&timeout=99999999`,
});
expect(res.statusCode).toBe(200);
expect(res.json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
});
});
+811
View File
@@ -0,0 +1,811 @@
/**
* @fileoverview Route tests for `GET /api/sessions/:id/wait`.
*
* The contract this pins is the one an orchestrating agent depends on:
* - a timeout is a 200 with `wait.timedOut: true`, never a 4xx, because callers loop
* over short waits and every poll boundary would otherwise look like a failure;
* - the result is nested under `data.wait` on ALL THREE wait endpoints, so a single
* client helper reads any of them;
* - the EFFECTIVE timeout is echoed, so a caller that asked for 30 minutes and was
* clamped to 10 can tell a poll boundary from a wedged worker;
* - an unknown `until` token is a 400 rather than a silent fallback to the default,
* so a typo can never leave an agent believing it is waiting for something else,
* and a schema 400 names the parameter it rejected;
* - `stop`/`blocked` are rejected for modes that install no hooks when asked for
* EXPLICITLY, but silently dropped from the DEFAULT set, so omitting `until` never
* 400s — and `shell` counts as such a mode, even though it is not an external CLI;
* - a client that hangs up frees its waiter immediately, or a loop of
* `curl --max-time` calls wedges a process-wide cap nobody else can use;
* - a session with no PTY answers `exit`, never `idle`.
*
* Plan: docs/agent-control-plan.md
*/
import { describe, it, expect, afterEach, vi } from 'vitest';
import fastifyCookie from '@fastify/cookie';
import Fastify, { type FastifyInstance } from 'fastify';
import type { ServerResponse } from 'node:http';
import {
registerSessionRoutes,
_resetPaneLivenessState,
_paneDeathWatcherCount,
} from '../../src/web/routes/session-routes.js';
import { createSessionListeners, attachSessionListeners } from '../../src/web/session-listener-wiring.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js';
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
import { sessionWaits } from '../../src/web/session-wait-registry.js';
import { MAX_WAIT_MS, MIN_WAIT_MS, MAX_WAITERS_TOTAL, MAX_WAITERS_PER_OWNER } from '../../src/config/agent-wait.js';
// Distinct per file on purpose: the three wait suites share the process-wide
// `sessionWaits` singleton, so a common id let one file's leftover waiter be counted
// by another's assertion. Failed only in a 5-file run, which is how CI runs them.
const SESSION_ID = 'wait-routes-session';
/** Session ids a test parked filler waiters on, so cleanup can release them. */
const fillerIds = new Set<string>();
/** Park a waiter directly on the shared registry (cap tests), tracked for cleanup. */
function fillWaiter(id: string, owner?: string): Promise<unknown> {
fillerIds.add(id);
return sessionWaits.waitForSignal(id, { until: ['stop'], timeoutMs: 30_000, owner });
}
afterEach(() => {
// The routes use the process-wide registry; never leak a waiter into the next test.
// Deliberately NOT `cancelEverything()`: it latches the registry's stopped flag
// (one-way by design, so a request landing mid-shutdown cannot register a waiter
// nothing will ever cancel), which would leave every later test in this file
// talking to a dead registry and passing vacuously.
for (const id of [SESSION_ID, ...fillerIds]) sessionWaits.cancelAll(id);
fillerIds.clear();
// Pane-liveness state is module-level (one cache, one watcher per pane), so it has
// to be reset or a cached probe leaks into the next test.
_resetPaneLivenessState();
delete process.env.CODEMAN_MULTIUSER;
});
/**
* The shared route harness returns handler payloads verbatim, so a `{success:false}`
* body would still be HTTP 200. Production maps errorCode to status in server.ts, so
* mirror that here or every negative case passes vacuously.
*
* `rawReplies` collects each request's `reply.raw` — the RESPONSE, which is what the
* handler's hang-up detection listens on, and deliberately not `req.raw` (on a POST
* that one closes as soon as the body has been read, so wiring an abort to it kills
* every send-and-wait; see the real-HTTP suite in session-input-wait.test.ts).
* Emitting the event by hand is the only way to simulate a hang-up here at all:
* `app.inject()` never emits `close` on its own, verified.
*/
async function harness(options?: { authUser?: { username: string; role: 'admin' | 'user' } }): Promise<{
app: FastifyInstance;
ctx: MockRouteContext;
rawReplies: ServerResponse[];
}> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
const ctx = createMockRouteContext({ sessionId: SESSION_ID });
const rawReplies: ServerResponse[] = [];
const authUser = options?.authUser;
app.addHook('onRequest', async (req, reply) => {
// The RESPONSE, because that is what the handler's hang-up detection listens to.
rawReplies.push(reply.raw);
if (authUser) (req as unknown as { authUser: typeof authUser }).authUser = authUser;
});
registerSessionRoutes(app, ctx as never);
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
const p = payload as { success?: unknown; errorCode?: unknown } | null;
if (p && typeof p === 'object' && p.success === false && reply.statusCode === 200) {
if (typeof p.errorCode === 'string') reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
}
return done(null, payload);
});
installRouteErrorHandler(app);
await app.ready();
return { app, ctx, rawReplies };
}
describe('GET /api/sessions/:id/wait', () => {
it('resolves immediately when the session is already in a requested state', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
expect(res.statusCode).toBe(200);
const body = res.json();
expect(body.success).toBe(true);
expect(body.data.wait.signal).toBe('idle');
expect(body.data.wait.immediate).toBe(true);
expect(body.data.wait.timedOut).toBe(false);
expect(body.data.wait.aborted).toBe(false);
expect(body.data.sessionId).toBe(SESSION_ID);
expect(body.data.wait.until).toEqual(['idle']);
});
it('nests the result under data.wait, so one client helper reads all three endpoints', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
const { data } = res.json();
// Session-level facts stay at the top; everything about the wait is inside it.
expect(Object.keys(data).sort()).toEqual(['limitPaused', 'sessionId', 'status', 'wait']);
expect(data.signal).toBeUndefined();
});
it('reports the post-wait status and the limit-pause hint', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
const body = res.json();
expect(body.data.status).toBe('idle');
expect(body.data.limitPaused).toBe(false);
});
it('sends Cache-Control: no-store, so a polled long-poll cannot be served from a cache', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
expect(res.headers['cache-control']).toBe('no-store');
});
it('defaults to stop,idle,exit when until is omitted', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
const body = res.json();
expect(body.data.wait.until).toEqual(['stop', 'idle', 'exit']);
// The mock session is idle, so the default set resolves right away.
expect(body.data.wait.signal).toBe('idle');
});
it('rejects an unknown until token instead of falling back to the default', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stpo` });
expect(res.statusCode).toBe(400);
const body = res.json();
expect(body.success).toBe(false);
expect(body.errorCode).toBe('INVALID_INPUT');
expect(body.error).toContain('stpo');
});
it('rejects a non-numeric timeout', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?timeout=soon` });
expect(res.statusCode).toBe(400);
expect(res.json().errorCode).toBe('INVALID_INPUT');
});
it('names the parameter it rejected, instead of a bare "invalid parameters"', async () => {
// An agent driving this with no docs in context can only recover if the error
// says WHICH parameter was wrong; the old message named none of them.
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?timeout=30s` });
const body = res.json();
expect(body.error).toContain('timeout');
expect(body.error).toContain('wait');
});
it('404s an unknown session', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: '/api/sessions/nope/wait?until=idle' });
expect(res.statusCode).toBe(404);
expect(res.json().success).toBe(false);
});
it('resolves an in-flight wait when the signal arrives', async () => {
const { app } = await harness();
// `stop` is not the session's current signal, so this blocks.
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
sessionWaits.notifySignal(SESSION_ID, 'stop');
const body = (await pending).json();
expect(body.data.wait.signal).toBe('stop');
expect(body.data.wait.immediate).toBe(false);
expect(body.data.wait.timedOut).toBe(false);
expect(body.data.wait.waitedMs).toBeGreaterThanOrEqual(0);
});
it('fresh=1 waits for the next transition instead of answering from current state', async () => {
const { app } = await harness();
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&fresh=1` });
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
sessionWaits.notifySignal(SESSION_ID, 'idle');
const body = (await pending).json();
expect(body.data.wait.signal).toBe('idle');
expect(body.data.wait.immediate).toBe(false);
});
it('resolves with ended when the session goes away mid-wait', async () => {
const { app } = await harness();
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.cancelAll(SESSION_ID);
const body = (await pending).json();
expect(body.success).toBe(true);
expect(body.data.wait.ended).toBe(true);
expect(body.data.wait.signal).toBeNull();
expect(body.data.wait.timedOut).toBe(false);
});
it('answers 200 with timedOut on timeout, never an error status', async () => {
const { app } = await harness();
// Clamped up to the 1s floor, so this is the one deliberately slow case.
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=1` });
expect(res.statusCode).toBe(200);
const body = res.json();
expect(body.success).toBe(true);
expect(body.data.wait.timedOut).toBe(true);
expect(body.data.wait.signal).toBeNull();
expect(body.data.wait.ended).toBe(false);
});
});
describe('GET /api/sessions/:id/wait: the effective timeout is observable', () => {
it('echoes the clamped-down value when the caller asks for more than the ceiling', async () => {
// Asked for 30 minutes, got MAX_WAIT_MS. Without the echo the caller reads a
// 10-minute timeout as "30 minutes elapsed with no stop" and kills a healthy worker.
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=1800000` });
expect(res.json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
});
it('echoes the clamped-up value at the floor too', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=1` });
expect(res.json().data.wait.timeoutMs).toBe(MIN_WAIT_MS);
});
it('reports the applied default when timeout is omitted', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
expect(res.json().data.wait.timeoutMs).toBeGreaterThanOrEqual(MIN_WAIT_MS);
expect(res.json().data.wait.timeoutMs).toBeLessThanOrEqual(MAX_WAIT_MS);
});
});
describe('GET /api/sessions/:id/wait: repeated query parameters', () => {
it('accepts ?until=stop&until=exit, the way most clients express a list', async () => {
// Fastify delivers a repeated parameter as an array and parseWaitSignals has
// always handled one; only the schema was rejecting it.
const { app } = await harness();
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&until=exit` });
await new Promise((resolve) => setTimeout(resolve, 20));
sessionWaits.notifySignal(SESSION_ID, 'exit');
const body = (await pending).json();
expect(body.data.wait.until).toEqual(['stop', 'exit']);
expect(body.data.wait.signal).toBe('exit');
});
it('still reports an unknown token inside a repeated parameter', async () => {
const { app } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&until=stpo` });
expect(res.statusCode).toBe(400);
expect(res.json().error).toContain('stpo');
});
});
describe('GET /api/sessions/:id/wait: a client that hangs up frees its waiter', () => {
it('removes the waiter when the RESPONSE socket closes early', async () => {
// `curl --max-time 30 ".../wait?timeout=600000"` abandons a live waiter every
// iteration of the documented loop; sixteen of those and an innocent session
// reports busy.
//
// The response body is unreadable afterwards (the injected response really is
// destroyed, exactly as a hung-up socket would be), so the freed slot is all this
// can assert. The full behaviour, including `aborted: true` on the wire for the
// caller that did NOT hang up, is pinned over real HTTP in session-input-wait.test.ts.
const { app, rawReplies } = await harness();
const pending = app
.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` })
.then(
() => 'completed',
() => 'destroyed'
);
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
rawReplies[rawReplies.length - 1].emit('close');
await new Promise((resolve) => setTimeout(resolve, 10));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
expect(await pending).toBe('destroyed');
});
it('a close AFTER the wait resolved changes nothing', async () => {
const { app, rawReplies } = await harness();
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
expect(res.json().data.wait.aborted).toBe(false);
// Node fires `close` on every completed response too, not only on a hang-up;
// `writableFinished` is what separates them.
rawReplies[rawReplies.length - 1].emit('close');
expect(sessionWaits.totalWaiterCount()).toBe(0);
});
});
describe('GET /api/sessions/:id/wait: capacity errors name the cap that was hit', () => {
it('maps the per-session cap to SESSION_BUSY / 409', async () => {
const { app } = await harness();
const pendings = [];
for (let i = 0; i < 16; i++) {
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` }));
}
await new Promise((resolve) => setTimeout(resolve, 30));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(16);
const overflow = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
expect(overflow.statusCode).toBe(409);
expect(overflow.json().errorCode).toBe('SESSION_BUSY');
expect(overflow.json().error).toContain('session');
sessionWaits.cancelAll(SESSION_ID);
await Promise.all(pendings);
});
it('maps the process-wide cap to RATE_LIMITED / 429, because this session is not the problem', async () => {
// Reported as SESSION_BUSY, an agent concludes the session it asked about is
// busy, switches to another, and gets the identical error.
const { app } = await harness();
const others: Promise<unknown>[] = [];
for (let i = 0; i < MAX_WAITERS_TOTAL; i++) others.push(fillWaiter(`unrelated-${i}`));
expect(sessionWaits.totalWaiterCount()).toBe(MAX_WAITERS_TOTAL);
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
expect(res.statusCode).toBe(429);
expect(res.json().errorCode).toBe('RATE_LIMITED');
expect(res.json().error).toContain('total');
for (const id of fillerIds) sessionWaits.cancelAll(id);
await Promise.all(others);
});
it('maps the per-owner cap to RATE_LIMITED / 429 and charges the request to its user', async () => {
// Also proves the route passes ownerFor(req): without it the owner cap can never
// trip, and one user could hold the whole process-wide pool.
process.env.CODEMAN_MULTIUSER = '1';
const { app } = await harness({ authUser: { username: 'alice', role: 'admin' } });
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.ownerWaiterCount('alice')).toBe(1);
const others: Promise<unknown>[] = [];
while (sessionWaits.ownerWaiterCount('alice') < MAX_WAITERS_PER_OWNER) {
others.push(fillWaiter(`alice-${others.length}`, 'alice'));
}
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
expect(res.statusCode).toBe(429);
expect(res.json().error).toContain('owner');
sessionWaits.cancelAll(SESSION_ID);
for (const id of fillerIds) sessionWaits.cancelAll(id);
await Promise.all([pending, ...others]);
});
});
describe('GET /api/sessions/:id/wait: modes that install no hooks', () => {
it('rejects an explicit stop, which no external CLI ever emits', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.mode = 'codex';
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
expect(res.statusCode).toBe(400);
expect(res.json().error).toContain('codex');
});
it('rejects an explicit blocked too', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.mode = 'opencode';
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=blocked` });
expect(res.statusCode).toBe(400);
});
it('rejects stop for a SHELL session, which is not an external CLI but installs no hooks either', async () => {
// A plain bash PTY never POSTs a Stop hook, so this was a guaranteed ten-minute
// hang dressed up as a timeout — the exact failure the guard exists to prevent.
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.mode = 'shell';
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
expect(res.statusCode).toBe(400);
expect(res.json().error).toContain('shell');
});
it('silently drops hook-only signals from the DEFAULT set instead of 400ing', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.mode = 'gemini';
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
expect(res.statusCode).toBe(200);
expect(res.json().data.wait.until).toEqual(['idle', 'exit']);
});
it('drops them from the default set for shell too', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.mode = 'shell';
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
expect(res.statusCode).toBe(200);
expect(res.json().data.wait.until).toEqual(['idle', 'exit']);
});
it('still accepts idle and exit explicitly', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.mode = 'antigravity';
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle,exit` });
expect(res.statusCode).toBe(200);
expect(res.json().data.wait.until).toEqual(['idle', 'exit']);
});
});
describe('GET /api/sessions/:id/wait: liveness beats the reported status', () => {
it('answers exit for a session whose PTY is gone, not the idle its status claims', async () => {
// Session parks a dead PTY at status 'idle' and the object survives in the map,
// so the DEFAULT wait used to answer {signal:"idle", immediate:true} for a
// crashed worker — 200, success, no error, and the agent prompts a corpse.
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
session.pid = null;
session.status = 'idle';
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
const body = res.json();
expect(body.data.wait.signal).toBe('exit');
expect(body.data.wait.immediate).toBe(true);
// The raw status is still reported, so nothing is hidden from the caller.
expect(body.data.status).toBe('idle');
});
it('resolves until=exit immediately for an already-exited session instead of blocking', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.pid = null;
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit&timeout=1` });
expect(res.json().data.wait.signal).toBe('exit');
expect(res.json().data.wait.timedOut).toBe(false);
});
it('does not report idle for a dead session even when idle was asked for explicitly', async () => {
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.pid = null;
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=1` });
expect(res.json().data.wait.signal).toBeNull();
expect(res.json().data.wait.timedOut).toBe(true);
});
it('a live busy session resolves until=working immediately', async () => {
// Unreachable before: MockSession used 'working', which is not a SessionStatus,
// so signalForStatus fell through to null and this branch had no coverage.
const { app, ctx } = await harness();
ctx.sessions.get(SESSION_ID)!.status = 'busy';
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=working` });
expect(res.json().data.wait.signal).toBe('working');
expect(res.json().data.wait.immediate).toBe(true);
});
it('a live stopped/error session maps to exit', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
session.status = 'stopped';
const stopped = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit` });
expect(stopped.json().data.wait.signal).toBe('exit');
session.status = 'error';
const errored = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit` });
expect(errored.json().data.wait.signal).toBe('exit');
});
});
/**
* The other half of the wait contract, exercised through the real listener wiring
* rather than a route: a PTY that dies must RELEASE every waiter, not only the ones
* that asked for `exit`.
*
* It lives in this file because it pins the same promise the routes above make
* ("never hang"), and because the failure is only visible from the caller's side:
* the exit handler detaches the `terminal`, `idle` and `working` listeners moments
* later, so anything still registered afterwards is waiting on feeds that no longer
* exist and can only time out.
*/
describe('a PTY exit releases waiters that did not ask for exit', () => {
/** Everything the exit handler touches; the wait release must not depend on any of it. */
function stubDeps(overrides: Record<string, unknown> = {}) {
return {
broadcast: vi.fn(),
batchTerminalData: vi.fn(),
batchTaskUpdate: vi.fn(),
broadcastSessionStateDebounced: vi.fn(),
sendPushNotifications: vi.fn(),
persistSessionState: vi.fn(),
getSessionStateWithRespawn: vi.fn(() => ({})),
getRunSummaryTracker: vi.fn(() => undefined),
stopTranscriptWatcher: vi.fn(),
cleanupSessionBatches: vi.fn(),
cancelPersistDebounce: vi.fn(),
removeRunSummaryTracker: vi.fn(),
removeSessionListenerRefs: vi.fn(),
cleanupRespawnOnExit: vi.fn(),
getStore: vi.fn(() => ({ updateRalphState: vi.fn() })),
registerAttachment: vi.fn(async () => {}),
...overrides,
};
}
it('answers an until=working waiter with ended instead of leaving it to time out', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=working&timeout=600000` });
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
session.emit('exit', 1);
const body = (await pending).json();
expect(body.data.wait.ended).toBe(true);
expect(body.data.wait.timedOut).toBe(false);
expect(body.data.wait.signal).toBeNull();
});
it('still gives an until=exit waiter its signal, not a bare ended', async () => {
// Ordering matters: notifySignal('exit') must run BEFORE cancelAll, or a caller
// that asked the right question gets the generic answer.
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit&fresh=1` });
await new Promise((resolve) => setTimeout(resolve, 20));
session.emit('exit', 0);
const body = (await pending).json();
expect(body.data.wait.signal).toBe('exit');
expect(body.data.wait.ended).toBe(false);
});
it('releases output waiters too, whose only feed the exit handler is about to detach', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
const pending = app.inject({
method: 'GET',
url: `/api/sessions/${SESSION_ID}/wait-output?match=DONE&timeout=600000`,
});
await new Promise((resolve) => setTimeout(resolve, 20));
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
session.emit('exit', 1);
const body = (await pending).json();
expect(body.data.wait.ended).toBe(true);
expect(body.data.wait.matched).toBe(false);
expect(sessionWaits.waiterCount(SESSION_ID)).toBe(0);
});
it('releases them even when a later step of the exit handler throws', async () => {
// Which is why the release is the first thing in the handler.
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
const deps = stubDeps({
broadcast: vi.fn(() => {
throw new Error('SSE is down');
}),
});
attachSessionListeners(session as never, createSessionListeners(session as never, deps as never));
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` });
await new Promise((resolve) => setTimeout(resolve, 20));
session.emit('exit', 1);
expect((await pending).json().data.wait.ended).toBe(true);
});
});
/**
* Worker liveness for a tmux-backed session.
*
* `session.pid` is the local `tmux attach` CLIENT, not the worker. Codeman sets
* `remain-on-exit on`, so when the command inside the pane exits tmux keeps the pane
* (`pane_dead=1`), the tmux session survives, the attach client keeps running, `pid`
* never goes null and NO exit event fires. Reproduced live on a shell worker killed
* with `exit 42`: tmux said `pane_dead=1 status=42` while Codeman said
* `pid=309406 status=idle` and the default wait answered
* `{signal:"idle", immediate:true, waitedMs:0}` for a corpse.
*
* These cases could not exist before, because `MockSession.pid` is set by hand: the
* `pid === null` branch is the one production never reaches.
*/
describe('GET /api/sessions/:id/wait: a dead tmux worker', () => {
/** Mock ctx doubles carry no `isPaneDead`; the route treats that as "cannot tell". */
function setPaneDead(ctx: MockRouteContext, dead: boolean): ReturnType<typeof vi.fn> {
const probe = vi.fn(() => dead);
(ctx.mux as unknown as { isPaneDead: (name: string) => boolean }).isPaneDead = probe as never;
return probe;
}
it('answers exit, not the idle the session still reports', async () => {
const { app, ctx } = await harness();
const session = ctx.sessions.get(SESSION_ID)!;
setPaneDead(ctx, true);
// Exactly the live state: attach client alive, status idle, worker gone.
expect(session.pid).not.toBeNull();
expect(session.status).toBe('idle');
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
const body = res.json();
expect(body.data.wait.signal).toBe('exit');
expect(body.data.wait.immediate).toBe(true);
// The raw status is still reported, so nothing is hidden from the caller.
expect(body.data.status).toBe('idle');
});
it('resolves until=exit immediately instead of burning the whole timeout', async () => {
const { app, ctx } = await harness();
setPaneDead(ctx, true);
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit&timeout=1` });
expect(res.json().data.wait.signal).toBe('exit');
expect(res.json().data.wait.timedOut).toBe(false);
});
it('does not answer idle for a dead worker even when idle was asked for explicitly', async () => {
const { app, ctx } = await harness();
setPaneDead(ctx, true);
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=1` });
expect(res.json().data.wait.signal).toBeNull();
expect(res.json().data.wait.timedOut).toBe(true);
});
it('a live pane is unaffected', async () => {
const { app, ctx } = await harness();
setPaneDead(ctx, false);
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
expect(res.json().data.wait.signal).toBe('idle');
});
it('caches the probe, so a poll loop cannot exec tmux once per request', async () => {
const { app, ctx } = await harness();
const probe = setPaneDead(ctx, false);
for (let i = 0; i < 10; i++) {
await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
}
expect(probe.mock.calls.length).toBeLessThanOrEqual(2);
});
it('never probes a session that is not tmux-backed', async () => {
const { app, ctx } = await harness();
const probe = setPaneDead(ctx, true);
ctx.sessions.get(SESSION_ID)!.usesMux = false;
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
expect(probe).not.toHaveBeenCalled();
// Falls back to the pid rule, which is the right one for a direct PTY.
expect(res.json().data.wait.signal).toBe('idle');
});
it('releases a wait when the worker dies WHILE it is parked', async () => {
// The common orchestration case, and the one a request-time probe cannot see: no
// exit event, no output, nothing — the caller would block for its full timeout.
const { app, ctx } = await harness();
let dead = false;
(ctx.mux as unknown as { isPaneDead: () => boolean }).isPaneDead = () => dead;
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` });
await new Promise((resolve) => setTimeout(resolve, 50));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
expect(_paneDeathWatcherCount()).toBe(1);
dead = true;
const body = (await pending).json();
expect(body.data.wait.ended).toBe(true);
expect(body.data.wait.timedOut).toBe(false);
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
// ...and the watcher is torn down with the last waiter that needed it.
expect(_paneDeathWatcherCount()).toBe(0);
}, 10_000);
it('an until=exit caller parked when the worker dies gets its signal, not a bare ended', async () => {
const { app, ctx } = await harness();
let dead = false;
(ctx.mux as unknown as { isPaneDead: () => boolean }).isPaneDead = () => dead;
const pending = app.inject({
method: 'GET',
url: `/api/sessions/${SESSION_ID}/wait?until=exit&fresh=1&timeout=600000`,
});
await new Promise((resolve) => setTimeout(resolve, 50));
dead = true;
const body = (await pending).json();
expect(body.data.wait.signal).toBe('exit');
expect(body.data.wait.ended).toBe(false);
}, 10_000);
it('starts no watcher at all when the session is not tmux-backed', async () => {
const { app, ctx } = await harness();
setPaneDead(ctx, false);
ctx.sessions.get(SESSION_ID)!.usesMux = false;
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` });
await new Promise((resolve) => setTimeout(resolve, 30));
expect(_paneDeathWatcherCount()).toBe(0);
sessionWaits.cancelAll(SESSION_ID);
await pending;
});
it('shares ONE watcher across every wait parked on the same session', async () => {
const { app, ctx } = await harness();
setPaneDead(ctx, false);
const pendings = [];
for (let i = 0; i < 5; i++) {
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` }));
}
await new Promise((resolve) => setTimeout(resolve, 40));
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(5);
expect(_paneDeathWatcherCount()).toBe(1);
sessionWaits.cancelAll(SESSION_ID);
await Promise.all(pendings);
expect(_paneDeathWatcherCount()).toBe(0);
});
});
describe('GET /api/sessions/:id/wait: an oversized timeout clamps, it does not 400', () => {
it('accepts a value above the old schema ceiling and reports the clamp', async () => {
// "Clamped to [1000, 600000]" has to mean it: `timeout=600001` clamping while
// `timeout=99999999` 400s is the same documented rule producing two outcomes.
const { app } = await harness();
const res = await app.inject({
method: 'GET',
url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=99999999`,
});
expect(res.statusCode).toBe(200);
expect(res.json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
});
it('still rejects a non-finite or non-integer timeout', async () => {
const { app } = await harness();
for (const value of ['1e999', 'soon', '-1', '1.5']) {
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?timeout=${value}` });
expect(res.statusCode, `timeout=${value}`).toBe(400);
}
});
});
+49
View File
@@ -208,6 +208,55 @@ describe('ws-routes', () => {
}
});
it('ACKs a delivered input and burns its seq', async () => {
const ws = await connectWs('/ws/sessions/ws-test-session/terminal');
try {
const session = ctx._session;
ws.send(JSON.stringify({ t: 'i', d: 'ok\r', cid: 'c1', seq: 1 }));
expect(await nextMessage(ws)).toEqual({ t: 'ia', seq: 1 });
expect(session.shouldApplyInput('c1', 1)).toBe(false);
} finally {
ws.close();
}
});
it('withholds the ACK and re-opens the seq when the write did not land', async () => {
// A session whose PTY is gone swallows the write. ACKing anyway told the
// client to drop the frame from its durable queue while the seq stayed
// burnt, so the retry that reliable delivery exists for was rejected as a
// duplicate — the input was lost for good.
const ws = await connectWs('/ws/sessions/ws-test-session/terminal');
try {
const session = ctx._session;
session.failWrites = true;
ws.send(JSON.stringify({ t: 'i', d: 'lost\r', cid: 'c1', seq: 1 }));
await expect(nextMessage(ws, 600)).rejects.toThrow(/timeout/);
expect(session.shouldApplyInput('c1', 1)).toBe(true);
} finally {
ws.close();
}
});
it('still ACKs a duplicate frame the server deliberately skipped', async () => {
// Dedup must stay silent-but-acknowledged: the client has to be able to
// drop a frame it already delivered once.
const ws = await connectWs('/ws/sessions/ws-test-session/terminal');
try {
const session = ctx._session;
session.shouldApplyInput('c1', 7); // pretend seq 7 already landed
ws.send(JSON.stringify({ t: 'i', d: 'again\r', cid: 'c1', seq: 7 }));
expect(await nextMessage(ws)).toEqual({ t: 'ia', seq: 7 });
expect(session.writeBuffer).not.toContain('again\r');
} finally {
ws.close();
}
});
it('ignores input exceeding MAX_INPUT_LENGTH', async () => {
const ws = await connectWs('/ws/sessions/ws-test-session/terminal');
try {
+179
View File
@@ -0,0 +1,179 @@
/**
* Unit tests for the unit-file builders behind `codeman service install`
* (issue #231). These are the parts that must be right without launchctl or
* systemctl in the loop: PATH construction, escaping, and the file contents.
*/
import { describe, it, expect } from 'vitest';
import {
buildLaunchAgentPlist,
buildServiceEnv,
buildServicePath,
buildSystemdUnit,
detectServiceKind,
systemdQuote,
xmlEscape,
type ServicePlan,
} from '../src/service-installer.js';
function plan(overrides: Partial<ServicePlan> = {}): ServicePlan {
return {
kind: 'systemd',
name: 'codeman-web.service',
nodePath: '/usr/bin/node',
execArgv: [],
scriptPath: '/home/u/.codeman/app/dist/index.js',
args: ['web', '--host', '127.0.0.1', '--port', '3000'],
env: { PATH: '/usr/bin:/bin', HOME: '/home/u', LANG: 'en_US.UTF-8' },
logPath: '/home/u/.codeman/web.log',
workingDir: '/home/u',
...overrides,
};
}
describe('buildServicePath', () => {
it("puts the running node's directory first so nvm/homebrew node wins", () => {
const result = buildServicePath('/home/u/.nvm/versions/node/v22.0.0/bin', '/usr/bin:/bin', '/home/u');
expect(result.split(':')[0]).toBe('/home/u/.nvm/versions/node/v22.0.0/bin');
});
it('keeps the installing shell PATH, which is the whole point of the fix', () => {
const result = buildServicePath('/usr/bin', '/opt/homebrew/bin:/home/u/.bun/bin', '/home/u');
expect(result.split(':')).toContain('/home/u/.bun/bin');
expect(result.split(':')).toContain('/opt/homebrew/bin');
});
it('appends the fallbacks a bare launchd PATH would otherwise be missing', () => {
const entries = buildServicePath('/usr/bin', '/usr/bin', '/home/u').split(':');
expect(entries).toContain('/opt/homebrew/bin');
expect(entries).toContain('/home/u/.local/bin');
expect(entries).toContain('/usr/local/bin');
});
it('never repeats a directory', () => {
const entries = buildServicePath('/usr/bin', '/usr/bin:/bin:/usr/bin', '/home/u').split(':');
expect(new Set(entries).size).toBe(entries.length);
});
it('drops empty segments from a trailing-colon PATH', () => {
expect(buildServicePath('/usr/bin', '/usr/bin::/bin:', '/home/u').split(':')).not.toContain('');
});
it('drops node_modules/.bin, which npx injects for one command only', () => {
const entries = buildServicePath(
'/usr/bin',
'/repo/node_modules/.bin:/repo/node_modules/.bin/:/home/u/bin',
'/home/u'
).split(':');
expect(entries.filter((e) => e.includes('node_modules'))).toEqual([]);
expect(entries).toContain('/home/u/bin');
});
});
describe('buildServiceEnv', () => {
it('carries PATH, HOME and a LANG default', () => {
const env = buildServiceEnv('/usr/bin', '/usr/bin:/bin', '/home/u');
expect(env.HOME).toBe('/home/u');
expect(env.LANG).toBe('en_US.UTF-8');
expect(env.PATH).toContain('/usr/bin');
});
it('prefers the caller LANG when there is one', () => {
expect(buildServiceEnv('/usr/bin', '/usr/bin', '/home/u', 'de_DE.UTF-8').LANG).toBe('de_DE.UTF-8');
});
it('does not carry a password into the unit file', () => {
const env = buildServiceEnv('/usr/bin', '/usr/bin', '/home/u');
expect(Object.keys(env)).not.toContain('CODEMAN_PASSWORD');
});
});
describe('escaping', () => {
it('escapes the five XML entities', () => {
expect(xmlEscape(`a&b<c>d"e'f`)).toBe('a&amp;b&lt;c&gt;d&quot;e&apos;f');
});
it('quotes systemd values and escapes quotes and backslashes', () => {
expect(systemdQuote('plain')).toBe('"plain"');
expect(systemdQuote('with "quotes"')).toBe('"with \\"quotes\\""');
expect(systemdQuote('back\\slash')).toBe('"back\\\\slash"');
});
});
describe('buildLaunchAgentPlist', () => {
it('writes the label, the full command and the log paths', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', name: 'com.codeman.web' }));
expect(xml).toContain('<string>com.codeman.web</string>');
expect(xml).toContain('<string>/usr/bin/node</string>');
expect(xml).toContain('<string>/home/u/.codeman/app/dist/index.js</string>');
expect(xml).toContain('<string>web</string>');
expect(xml).toContain('<string>/home/u/.codeman/web.log</string>');
});
it('keeps the argument order: node, script, then the web args', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', name: 'com.codeman.web' }));
// Match whole <string> elements: the label itself contains the word "web".
const order = [
'<string>/usr/bin/node</string>',
'<string>/home/u/.codeman/app/dist/index.js</string>',
'<string>web</string>',
'<string>--port</string>',
].map((s) => xml.indexOf(s));
expect(order).toEqual([...order].sort((a, b) => a - b));
expect(order.every((i) => i > -1)).toBe(true);
});
it('carries the runner flags so a tsx dev install still boots', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', execArgv: ['--import', 'tsx'] }));
expect(xml).toContain('<string>--import</string>');
expect(xml).toContain('<string>tsx</string>');
});
it('restarts on crash and at login', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd' }));
expect(xml).toContain('<key>KeepAlive</key>');
expect(xml).toContain('<key>RunAtLoad</key>');
});
it('escapes a path with an ampersand instead of emitting broken XML', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', workingDir: '/Users/a&b' }));
expect(xml).toContain('<string>/Users/a&amp;b</string>');
expect(xml).not.toContain('<string>/Users/a&b</string>');
});
});
describe('buildSystemdUnit', () => {
it('builds ExecStart from node, script and args', () => {
expect(buildSystemdUnit(plan())).toContain(
'ExecStart=/usr/bin/node /home/u/.codeman/app/dist/index.js web --host 127.0.0.1 --port 3000'
);
});
it('quotes an argument containing spaces', () => {
const unit = buildSystemdUnit(plan({ scriptPath: '/home/my user/app/dist/index.js' }));
expect(unit).toContain('"/home/my user/app/dist/index.js"');
});
it('writes each env var as a quoted Environment line', () => {
const unit = buildSystemdUnit(plan());
expect(unit).toContain('Environment="PATH=/usr/bin:/bin"');
expect(unit).toContain('Environment="HOME=/home/u"');
});
it('keeps KillMode=process so agents survive a server restart', () => {
expect(buildSystemdUnit(plan())).toContain('KillMode=process');
});
it('is installable and restarts on failure', () => {
const unit = buildSystemdUnit(plan());
expect(unit).toContain('Restart=always');
expect(unit).toContain('WantedBy=default.target');
});
});
describe('detectServiceKind', () => {
it('maps the platform to its supervisor', () => {
const expected = process.platform === 'darwin' ? 'launchd' : process.platform === 'linux' ? 'systemd' : null;
expect(detectServiceKind()).toBe(expected);
});
});
+35
View File
@@ -64,3 +64,38 @@ describe('buildMuxAttachEnv', () => {
}
});
});
describe('spawn env CODEMAN_API_URL (no fallback)', () => {
const withApiUrl = (value: string | undefined, fn: () => void) => {
const original = process.env.CODEMAN_API_URL;
if (value === undefined) delete process.env.CODEMAN_API_URL;
else process.env.CODEMAN_API_URL = value;
try {
fn();
} finally {
if (original === undefined) delete process.env.CODEMAN_API_URL;
else process.env.CODEMAN_API_URL = original;
}
};
it('passes the server-stamped URL through verbatim', async () => {
const { buildClaudeEnv, buildShellEnv } = await import('../src/session-cli-builder.js');
withApiUrl('https://127.0.0.1:3199', () => {
expect(buildClaudeEnv('test-session').CODEMAN_API_URL).toBe('https://127.0.0.1:3199');
expect(buildShellEnv('test-session').CODEMAN_API_URL).toBe('https://127.0.0.1:3199');
});
});
// A hardcoded fallback was the wrong scheme on HTTPS installs. The key must be
// genuinely ABSENT when unset: present-with-undefined would serialize through
// node-pty as the literal string "CODEMAN_API_URL=undefined" (COD-115).
it('leaves the key absent (not undefined, not a fallback) when the server has not stamped one', async () => {
const { buildClaudeEnv, buildShellEnv } = await import('../src/session-cli-builder.js');
withApiUrl(undefined, () => {
for (const env of [buildClaudeEnv('test-session'), buildShellEnv('test-session')]) {
expect('CODEMAN_API_URL' in env).toBe(false);
expect(JSON.stringify(env)).not.toContain('localhost:3000');
}
});
});
});
File diff suppressed because it is too large Load Diff
+89
View File
@@ -0,0 +1,89 @@
/**
* @fileoverview `/api/events` must not lose the headers the security hook set.
*
* The SSE route answers with `reply.raw.writeHead()`, which writes straight to the
* Node response and bypasses Fastify's header store. Everything the `onRequest`
* security hook had granted was therefore dropped — including the
* `Access-Control-Allow-Origin` it emits for localhost origins. The contradiction is
* visible from a browser: a localhost page may call every other `/api` endpoint
* cross-origin, but its EventSource fails CORS.
*
* These tests drive a REAL WebServer. An earlier version asserted against an inline
* copy of the hook and the handler, which proved nothing: reverting the fix in
* `server.ts` left every test green.
*/
import { afterAll, beforeAll, describe, expect, it } from 'vitest';
import { WebServer } from '../src/web/server.js';
const TEST_PORT = 3119;
const LOCAL_ORIGIN = 'http://localhost:5173';
/** Open /api/events, read the response headers, then abort — it never ends on its own. */
async function eventsHeaders(baseUrl: string, origin?: string): Promise<Headers> {
const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 2000);
try {
const res = await fetch(`${baseUrl}/api/events`, {
signal: controller.signal,
headers: origin ? { Origin: origin } : undefined,
});
const headers = res.headers;
controller.abort(); // stop consuming the stream
return headers;
} finally {
clearTimeout(timeout);
}
}
describe('GET /api/events header inheritance', () => {
let server: WebServer;
let baseUrl: string;
beforeAll(async () => {
server = new WebServer(TEST_PORT, false, true);
await server.start();
baseUrl = `http://localhost:${TEST_PORT}`;
});
afterAll(async () => {
await server.stop();
}, 60000);
it('keeps the CORS header the security hook granted a localhost origin', async () => {
// The regression: this header is set on the Fastify reply and was then thrown
// away by writeHead, so an EventSource from a localhost dev server failed CORS
// while every other endpoint worked.
const headers = await eventsHeaders(baseUrl, LOCAL_ORIGIN);
expect(headers.get('access-control-allow-origin')).toBe(LOCAL_ORIGIN);
});
it('keeps the security headers the hook set', async () => {
const headers = await eventsHeaders(baseUrl);
expect(headers.get('x-content-type-options')).toBe('nosniff');
expect(headers.get('x-frame-options')).toBe('SAMEORIGIN');
expect(headers.get('content-security-policy')).toBeTruthy();
});
it('still sets the SSE headers, and they win over anything inherited', async () => {
const headers = await eventsHeaders(baseUrl);
expect(headers.get('content-type')).toBe('text/event-stream');
expect(headers.get('cache-control')).toBe('no-cache');
expect(headers.get('x-accel-buffering')).toBe('no');
});
it('grants nothing to a non-localhost origin — the hook decides, not this route', async () => {
const headers = await eventsHeaders(baseUrl, 'https://evil.example');
expect(headers.get('access-control-allow-origin')).toBeNull();
});
it('matches what a normal JSON endpoint returns for the same origin', async () => {
// The point of the fix: /api/events stops being the odd one out.
const json = await fetch(`${baseUrl}/api/status`, { headers: { Origin: LOCAL_ORIGIN } });
const sse = await eventsHeaders(baseUrl, LOCAL_ORIGIN);
expect(sse.get('access-control-allow-origin')).toBe(json.headers.get('access-control-allow-origin'));
expect(sse.get('x-content-type-options')).toBe(json.headers.get('x-content-type-options'));
});
});
+20
View File
@@ -47,6 +47,26 @@ function loadTerminalUiHarness(mode: string) {
}
describe('terminal flush budget', () => {
it('drains a large final batch without waiting for unrelated terminal output', () => {
const { app, writes } = loadTerminalUiHarness('codex');
const scheduled: Array<() => void> = [];
app._safeYield = (callback: () => void) => {
scheduled.push(callback);
};
app.isTerminalAtBottom = () => true;
app.batchTerminalWrite('x'.repeat(96 * 1024));
expect(scheduled).toHaveLength(1);
while (scheduled.length > 0) {
scheduled.shift()?.();
}
expect(writes.map((write) => write.length)).toEqual([32 * 1024, 32 * 1024, 32 * 1024]);
expect(app.pendingWrites).toEqual([]);
expect(app.writeFrameScheduled).toBe(false);
});
it('uses a smaller first-frame write budget for Codex output to reduce renderer stalls', () => {
const { app, writes } = loadTerminalUiHarness('codex');
app.pendingWrites.push('x'.repeat(96 * 1024));
+240
View File
@@ -0,0 +1,240 @@
/**
* Issue #205, round 2: the 1.12.0 retest still reported unusable scrollback —
* a completely dead wheel on Firefox/macOS (while Fn+Up paged back through
* intact text), and history on iPhone that went back a little, repeated blocks
* and got worse the further up it went.
*
* Both signatures come from a Claude pane's LOCAL buffer being hollow. tmux
* keeps no history for a repaint-mode pane (`history_size≈0`), so:
* - any gesture routed to local scrollback scrolls nothing, and
* - the scroll-to-top `?full=1` re-pull replaces a multi-frame buffer with a
* single captured frame, deleting history mid-scroll.
*
* These cover the two guards that fix it: `_replayWouldShrinkBuffer` (refuse a
* downgrading re-pull) and `_maybePageCliTranscript` (page the CLI's own
* transcript when there is nothing local to scroll), plus the diagnostic that
* makes the routing decision visible instead of guessable.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it, vi } from 'vitest';
function loadTerminalUiHarness() {
const CodemanApp = function CodemanApp(this: any) {};
const logs: string[] = [];
const context = vm.createContext({
window: {},
CodemanApp,
console: { warn: vi.fn(), log: (msg: string) => logs.push(msg) },
_crashDiag: { log: vi.fn() },
performance: { now: () => 1_000 },
requestAnimationFrame: (_fn: () => void) => 1,
setTimeout: (_fn: () => void) => 1,
Blob: function Blob() {},
URL: { createObjectURL: () => 'blob:yield', revokeObjectURL: () => {} },
Worker: function Worker(this: any) {
this.postMessage = () => {};
},
MobileDetection: { isTouchDevice: () => true },
DEC_SYNC_STRIP_RE: /\x1b\[\?2026[hl]/g,
TERMINAL_CHUNK_SIZE: 32 * 1024,
});
const code = readFileSync(resolve(import.meta.dirname, '../src/web/public/terminal-ui.js'), 'utf8');
vm.runInContext(code, context, { filename: 'terminal-ui.js' });
return { app: new (CodemanApp as any)(), logs };
}
/** A Claude session whose local buffer holds exactly one screen (baseY 0). */
function hollowClaudeApp(overrides: { cliVersion?: string; rows?: number } = {}) {
const { app, logs } = loadTerminalUiHarness();
const sent: Array<{ id: string; data: string }> = [];
app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: overrides.cliVersion }]]);
app._sendInputEphemeral = (id: string, data: string) => sent.push({ id, data });
app.terminal = {
cols: 80,
rows: overrides.rows ?? 36,
modes: { mouseTrackingMode: 'none' },
buffer: { active: { type: 'normal', viewportY: 0, baseY: 0, length: 36 } },
};
return { app, sent, logs };
}
describe('full-history re-pull downgrade guard (issue #205 round 2)', () => {
it('estimates replayed rows from wrapped, escape-laden capture text', () => {
const { app } = loadTerminalUiHarness();
expect(app._estimateReplayRows('a\r\nb\r\nc', 80)).toBe(3);
// SGR colour runs occupy no cells, so they must not inflate the estimate.
expect(app._estimateReplayRows('\x1b[38;5;196mred\x1b[0m\r\nplain', 80)).toBe(2);
// capture-pane -J joins wrapped rows, so a long logical line re-wraps on
// write — counting newlines alone would undershoot by 2 rows here.
expect(app._estimateReplayRows('x'.repeat(25), 10)).toBe(3);
expect(app._estimateReplayRows('', 80)).toBe(0);
expect(app._estimateReplayRows(undefined, 80)).toBe(0);
});
it('refuses a capture that would leave LESS history than the terminal holds', () => {
const { app } = loadTerminalUiHarness();
app.terminal = { cols: 80, rows: 36, buffer: { active: { length: 300 } } };
// Claude pane: tmux has no history, so the capture is one frame while xterm
// holds hundreds of replayed rows. Rewriting would delete them mid-scroll.
const oneFrame = Array.from({ length: 36 }, (_, i) => `frame line ${i}`).join('\r\n');
expect(app._replayWouldShrinkBuffer(oneFrame)).toBe(true);
// Shell pane after a burst/tab-switch collapse: tmux really does hold more.
const realHistory = Array.from({ length: 800 }, (_, i) => `history ${i}`).join('\r\n');
expect(app._replayWouldShrinkBuffer(realHistory)).toBe(false);
});
it('tolerates a one-screen shortfall so ordinary recoveries still replay', () => {
const { app } = loadTerminalUiHarness();
// buffer.active.length counts the blank rows under the last line and the row
// estimate can only approximate wrapping, so a near-tie must NOT read as a
// downgrade — only a capture worse by more than a full screen does.
app.terminal = { cols: 80, rows: 36, buffer: { active: { length: 120 } } };
expect(app._replayWouldShrinkBuffer(Array.from({ length: 100 }, () => 'x').join('\r\n'))).toBe(false);
expect(app._replayWouldShrinkBuffer(Array.from({ length: 40 }, () => 'x').join('\r\n'))).toBe(true);
});
it('never refuses when the terminal has no buffer to protect', () => {
const { app } = loadTerminalUiHarness();
app.terminal = { cols: 80, rows: 36, buffer: { active: { length: 0 } } };
expect(app._replayWouldShrinkBuffer('anything')).toBe(false);
});
it('is wired into _maybeRefetchFullHistory BEFORE the destructive reset', () => {
const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8');
const start = source.indexOf('async _maybeRefetchFullHistory()');
const guard = source.indexOf('this._replayWouldShrinkBuffer(buffer)', start);
const reset = source.indexOf('this._resetTerminalForReplay()', start);
expect(start).toBeGreaterThan(-1);
expect(guard).toBeGreaterThan(start);
expect(guard).toBeLessThan(reset); // refuse first, only then reset+rewrite
// A hollow pane must also stop re-fetching megabytes on every scroll-up.
expect(source).toContain('this._fullHistoryRepullUseless');
expect(source).toContain('this._fullHistoryRepullUseless?.has(sessionId) ? 60000 : 4000');
});
});
describe('PageUp/PageDown fallback for a hollow local buffer (issue #205 round 2)', () => {
it('pages the CLI transcript when the wheel gate is false and there is no scrollback', () => {
const { app, sent } = hollowClaudeApp(); // cliVersion unknown → gate false
// Half a screen of travel (rows 36 → 18 lines) buys exactly one PageUp.
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(true);
app._flushWheelSgrQueue();
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[5~' }]);
// Downward travel pages back toward the live screen.
app._maybePageCliTranscript({ shiftKey: false }, 18);
app._flushWheelSgrQueue();
expect(sent[1]).toEqual({ id: 'sess-1', data: '\x1b[6~' });
});
it('accumulates sub-page travel instead of dropping or over-sending it', () => {
const { app, sent } = hollowClaudeApp();
expect(app._maybePageCliTranscript({ shiftKey: false }, -10)).toBe(true); // consumed…
app._flushWheelSgrQueue();
expect(sent).toEqual([]); // …but below the threshold, so nothing sent yet
app._maybePageCliTranscript({ shiftKey: false }, -8); // -18 total → one page
app._flushWheelSgrQueue();
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[5~' }]);
});
it('caps the keys one gesture batch can emit', () => {
const { app, sent } = hollowClaudeApp();
app._maybePageCliTranscript({ shiftKey: false }, -1000); // 55 pages of travel
app._flushWheelSgrQueue();
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[5~'.repeat(3) }]);
});
it('leaves every session that has real local scrollback alone', () => {
const { app } = hollowClaudeApp();
// Shift is the explicit "give me local scrollback" gesture — never paged.
expect(app._maybePageCliTranscript({ shiftKey: true }, -18)).toBe(false);
// A buffer with history scrolls locally, as before.
app.terminal.buffer.active.baseY = 120;
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(false);
app.terminal.buffer.active.baseY = 0;
// Non-Claude modes keep their existing behavior (shell scrolls tmux history
// through the alt-screen strip; codex/gemini page keys are unverified).
app.sessions = new Map([['sess-1', { mode: 'shell' }]]);
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(false);
app.sessions = new Map([['sess-1', { mode: 'codex' }]]);
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(false);
// An alternate-screen pane belongs to xterm's own alt-scroll handling.
app.sessions = new Map([['sess-1', { mode: 'claude' }]]);
app.terminal.buffer.active.type = 'alternate';
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(false);
});
it('rescues the local-scrollback opt-out footgun instead of silently dying', () => {
// "Wheel scrolls local history" ON pins the wheel to a buffer that, for a
// repaint-mode CLI, is empty — a user who flipped it while hunting for a fix
// on 1.11.x would have ended up with a completely dead wheel on 1.12.0.
const { app, sent } = hollowClaudeApp({ cliVersion: '2.1.223' }); // gate would forward…
app.loadAppSettingsFromStorage = () => ({ terminalWheelLocalScrollback: true });
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); // …but the opt-out wins
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(true);
app._flushWheelSgrQueue();
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[5~' }]);
});
it('drops travel accumulated on another tab', () => {
const { app, sent } = hollowClaudeApp();
app._maybePageCliTranscript({ shiftKey: false }, -17); // just short of a page
app.activeSessionId = 'sess-2';
app.sessions.set('sess-2', { mode: 'claude' });
app._maybePageCliTranscript({ shiftKey: false }, -1); // must not complete sess-1's page
app._flushWheelSgrQueue();
expect(sent).toEqual([]);
});
it('is reachable from both the wheel and the touch paths', () => {
const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/terminal-ui.js'), 'utf8');
// Wheel: after the forwarding gate, before the local smooth scroll.
expect(source).toContain('if (this._maybePageCliTranscript(ev, lines)) return;');
// Touch: touchmove and the momentum loop both fall through to it.
expect(source.match(/else if \(!this\._maybePageCliTranscript\(\{ shiftKey: false \}, lines\)\)/g)).toHaveLength(2);
});
});
describe('scroll routing diagnostic (issue #205 round 2)', () => {
it('prints the decision and its inputs once per session, and again when it changes', () => {
const { app, logs } = hollowClaudeApp({ cliVersion: '2.1.100' });
app.loadAppSettingsFromStorage = () => ({ terminalWheelLocalScrollback: false });
app._logScrollRouting('local-scrollback');
app._logScrollRouting('local-scrollback'); // same decision → stays quiet
expect(logs).toHaveLength(1);
expect(logs[0]).toContain('sess-1 → local-scrollback');
expect(logs[0]).toContain('mode=claude');
expect(logs[0]).toContain('cliVersion=2.1.100');
expect(logs[0]).toContain('localScrollbackOptOut=false');
expect(logs[0]).toContain('mouseTracking=none');
app._logScrollRouting('page-keys'); // a changed route still prints
expect(logs).toHaveLength(2);
expect(logs[1]).toContain('page-keys');
});
it('reports an unknown CLI version, the false-path that disables forwarding', () => {
const { app, logs } = hollowClaudeApp(); // no cliVersion — the probe failed
app._logScrollRouting('page-keys');
expect(logs[0]).toContain('cliVersion=unknown');
});
});
+9 -3
View File
@@ -379,7 +379,7 @@ describe('terminal touch tap mouse guard', () => {
expect(withVersion('garbage')).toBe(false); // unparseable → assume older
});
it('wheel: codex forwards without a version; gemini never forwards', () => {
it('wheel: only claude forwards — codex and gemini keep the local wheel', () => {
const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.terminal = {
@@ -387,8 +387,14 @@ describe('terminal touch tap mouse guard', () => {
buffer: { active: { viewportY: 50, baseY: 50 } },
};
app.sessions = new Map([['sess-1', { mode: 'codex' }]]); // verified TUI, no version gate
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(true);
// Codex used to forward unconditionally, which is PR #227's regression: measured
// on codex-cli 0.147.0, it never enables mouse tracking and ignores SGR wheel
// reports outright, so forwarding ate every tick while its real local scrollback
// (the codex transcript lives there — inline viewport, no in-app pager) sat unused.
app.sessions = new Map([['sess-1', { mode: 'codex' }]]);
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
app.sessions = new Map([['sess-1', { mode: 'codex', cliVersion: '9.9.9' }]]); // no version rescues it
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
app.sessions = new Map([['sess-1', { mode: 'gemini', cliVersion: '9.9.9' }]]); // unverified TUI
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
+31 -4
View File
@@ -247,14 +247,41 @@ describe('TmuxManager (unit)', () => {
});
describe('environment exports', () => {
it('keeps COLORTERM unset for OpenCode sessions', () => {
const exports = (
const callBuildEnvExports = (mode: string) =>
(
manager as unknown as {
buildEnvExports(sessionId: string, muxName: string, mode: string): string[];
}
).buildEnvExports('session-1', 'codeman-abc12345', 'opencode');
).buildEnvExports('session-1', 'codeman-abc12345', mode);
expect(exports).toContain('unset COLORTERM');
it('keeps COLORTERM unset for OpenCode sessions', () => {
expect(callBuildEnvExports('opencode')).toContain('unset COLORTERM');
});
it('exports the server-stamped CODEMAN_API_URL verbatim', () => {
const original = process.env.CODEMAN_API_URL;
process.env.CODEMAN_API_URL = 'https://127.0.0.1:3199';
try {
expect(callBuildEnvExports('claude')).toContain('export CODEMAN_API_URL=https://127.0.0.1:3199');
} finally {
if (original === undefined) delete process.env.CODEMAN_API_URL;
else process.env.CODEMAN_API_URL = original;
}
});
// A hardcoded fallback exported the wrong scheme on HTTPS installs; unset must
// stay unset so in-session guards fail closed instead of curling a bad URL.
it('exports no CODEMAN_API_URL at all when the server has not stamped one', () => {
const original = process.env.CODEMAN_API_URL;
delete process.env.CODEMAN_API_URL;
try {
const exports = callBuildEnvExports('claude');
expect(exports.some((line) => line.startsWith('export CODEMAN_API_URL'))).toBe(false);
expect(exports.join(' ')).not.toContain('localhost:3000');
} finally {
if (original === undefined) delete process.env.CODEMAN_API_URL;
else process.env.CODEMAN_API_URL = original;
}
});
});