Compare commits

...
Author SHA1 Message Date
Codeman maintainer 94aa53c65b chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:38:40 +02:00
Codeman maintainer e88b971bb7 feat(skill): add the agent-skill install layer and harden the packaged skill
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.

Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
  Case names resolve through linked-cases.json first, mirroring the
  server's resolveCasePath(), so a case linked in from outside
  ~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
  hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
  skill is never touched, and a symlinked skill dir is refused (this
  repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
  ports/config-port.ts, server.ts, session-routes.ts (add-only injection
  on Claude session create and quick-start), plus the App Settings toggle.

Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
  Shell state does not survive between agent tool calls, and an undefined
  is_self exited 127, firing the `||` branch and deleting the caller's own
  session with the one guard bypassed. The request now lives inside the
  guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
  tool call, so the documented resend-identical-request loop stopped being
  a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
  workers. It returns clean transcript text; the terminal scrape it
  replaces returns a wall of TUI repaint noise. Its transcript flush lags
  the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
  yielded the literal session id "null" and burned the whole readiness
  budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
  corrected the hooks-config comment that claimed a toggle-off sweep
  exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
  and that caseName resolves linked cases, so a generic name can land a
  worker in a real repo.

Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:30:15 +02:00
Codeman maintainer 8406c497e2 fix(terminal): stop forwarding the wheel to codex, it ignores SGR reports
DodgyBadger reported a completely dead wheel in codex tabs (#227 comment)
while the scrollbar drag worked, and the [scroll] line confirmed the
branch: forward-sgr with 967 rows of healthy local scrollback unused.

Measured against codex-cli 0.147.0 in a bare tmux: codex never enables
mouse tracking (mouse_any_flag=0), runs an inline viewport
(alternate_on=0) and pushes its transcript into the terminal's own
scrollback (history_size grows), and SGR wheel reports written to its
pane change nothing at all. Hand-encoded SGR taps are no-ops too, so
they stay (harmless), which means click-to-position is merely
unavailable there rather than damaging.

_shouldForwardWheelToApp now returns true for claude >= 2.1.187 and
nothing else; codex falls to the local-scrollback path like
shell/gemini/opencode, which is the same history the scrollbar drag was
already reaching. The claude-only PageUp fallback is untouched.

Verified in Chromium against a live codex session on an isolated
instance: routing logs local-scrollback, the viewport moves 39 -> 4 and
zero bytes go to the PTY.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:22:06 +02:00
Codeman maintainer 40b4aba043 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:35:59 +02:00
Codeman maintainer 4b44988bfc test: give daemon-control tests a unique port (3212 was already taken)
test/sse-subscription-filter.test.ts already binds 3212; sequential test
execution hid the clash. Moves the probeServer fixture to 3216 (3217 for
the nothing-listening case) per the unique-port convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:22:01 +02:00
Codeman maintainer 316d0a4c82 Merge pull request #233 from Lint111/feat/hooks-config
Conflict in refreshStaleCodemanHooks resolved by keeping every staleness
trigger: the master-side TLS-flagless curl check (hooks without -k) AND the
PR-side current-wake-marker (V3) + SubagentStop guard marker checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:21:52 +02:00
Ark0N 1184720648 Merge pull request #239 from Ark0N/feat/daemon-mode
feat(cli): codeman web -d and codeman service install (#231)
2026-08-09 02:19:50 +02:00
Ark0N b067aad9b6 Merge pull request #235 from Lint111/feat/deferred-terminal-flush
fix(terminal): drain deferred output without a wake event
2026-08-09 02:19:32 +02:00
Ark0N 19a3d7c773 Merge pull request #234 from Lint111/feat/ai-checker-stderr
fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
2026-08-09 02:19:10 +02:00
Codeman maintainer 085f4acb60 feat(cli): codeman web -d and codeman service install (#231)
Two ways to keep the server running, split by how long it should last.

`codeman web -d` relaunches the same entry script detached (setsid), with
`--stop` and `--status` alongside it. A pidfile and log live in the data
dir. `nohup` is not what makes this work: Node re-arms SIGHUP to its
default disposition even when it inherits "ignore", and cli.ts handles
SIGHUP with a graceful shutdown, so a delivered HUP still stops the
server. Removing the shell's ability to send one is the fix.

`codeman service install|uninstall|status` writes and loads the systemd
user unit or the LaunchAgent, with the installing shell's PATH baked in
(launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, which finds neither a
Homebrew/nvm node nor tmux/claude). install.sh already covers one-liner
installs; this is for npm globals.

Both refuse to start when a server is already up on the data dir, since a
second instance on the shared tmux socket attaches PTYs to the first
one's live sessions. Both poll /api/status until the child answers or
dies rather than reporting a success they have not seen. `--stop` checks
the pid still looks like a Codeman server before signalling it.

The systemd unit name and launchd label move to config/service-names.ts
so install.sh, detectSupervisor() and service install cannot drift into
supervising two copies. Instance-scoped, unchanged for the default
instance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:34:55 +02:00
lior 091df2b6d8 fix(terminal): drain deferred output without a wake event 2026-08-08 23:00:36 +03:00
lior 5f775b1ab1 fix(hooks): guard subagent stops and rewake from the parent transcript
Two defects in the background-task hook scripts.

SubagentStop had no handler at all. When a subagent launched background work and
one watcher ended while others were still running, Claude could publish the
worker's last progress sentence as its final result, abandoning the live tasks.
A new guard pairs launched task IDs against completed ones and confirms liveness
by scanning /proc/<pid>/fd for an open tasks/<id>.output handle, blocking the
stop only while genuinely-live work remains. It fails open — allowing the stop —
when /proc is unavailable, nothing was launched, or everything finished.

The rewake helper watched only input.transcript_path. A subagent has its own
transcript, but Claude writes the completion queue-operation to the PARENT
transcript, so the record it waited for never appeared and the wake never fired.
It now watches both paths, but only when the relationship is provable: the
transcript's parent directory is subagents/ and its grandparent basename equals
input.session_id. It also now requires operation === 'enqueue'.

The rewake marker moves V2 -> V3; refreshStaleCodemanHooks treats absence of the
current marker as stale, so existing cases self-heal on next launch (the same
mechanism as the V1 -> V2 bump). Ownership matches on marker PREFIXES, so a
future bump still recognises older Codeman handlers and never adopts a user's.

12 tests fail on unmodified master, e.g.
  expected '[{"matcher":"Bash",…' to contain 'CODEMAN_BACKGROUND_REWAKE_V3'
  expected 'Background command bg-report-1 comple…' to contain '<codeman-background-result>'
2026-08-08 22:31:38 +03:00
lior da51193264 fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
AiCheckerBase spawned the check with `> out 2>&1`, so anything the Claude CLI
wrote to stderr landed inside the same file the verdict parser reads. A CLI that
failed to start (corrupt settings, missing auth) produced either an empty verdict
or an unparseable one, and the actual cause was destroyed on the way through —
the user saw only "Empty output from AI idle check".

stderr now goes to its own temp file. When output is empty or the verdict cannot
be parsed, the first 200 characters of stderr are appended to the error message.
The file is cleaned up alongside the existing temp files, including on the error
paths.

Two tests, both failing on master:
  expected 'export PATH="…' to contain ' 2> "'
  expected 'Empty output from AI idle check' to contain 'Claude CLI failed to load settings'
2026-08-08 22:30:39 +03:00
38 changed files with 3165 additions and 281 deletions
+115
View File
@@ -1,5 +1,120 @@
# aicodeman
## 1.14.1
### Patch Changes
- The Codeman agent skill is now installable, so an agent running inside a Codeman session can drive the API without you pasting docs into its prompt. Plus six fixes to the packaged skill, each found by running it live against a real instance.
## What the skill is
`skills/codeman` is a Claude Code skill that teaches an agent inside a Codeman session how to start worker sessions, send them prompts, block until they finish, read their answers and clean up. It ships in the npm package. It self-gates: outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act, so installing it globally costs unrelated sessions nothing.
## Installing it
Three ways, pick one:
```bash
codeman skill install # ~/.claude/skills/codeman, every new Claude Code session sees it
codeman skill install --case myproject # just that case; linked cases resolve by name too
codeman skill uninstall # reverses either one
```
Or turn on **App Settings > Agent Skill** (`agentSkillEnabled`, synced, default off) and Codeman injects the skill into each case when a Claude session is created there.
Installs are marker-owned: a `skills/codeman` that Codeman did not write is never touched, a stale managed copy is refreshed in place, and a symlinked skill directory is refused rather than written through. Re-run `codeman skill install` after upgrading Codeman to refresh the copy.
Note that turning `agentSkillEnabled` back off does **not** remove already-injected copies, because a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` directory. Remove them per case with `codeman skill uninstall --case <name>`.
## Using it
Once installed, just ask: "spin up three workers and have them lint, typecheck and test in parallel, then report back". The skill supplies the guard, the safety rules and the recipes. What it does under the hood:
**1. Guard.** Every Bash call re-runs a preamble that refuses outside `CODEMAN_MUX=1`, reads `CODEMAN_API_URL` and `CODEMAN_SESSION_ID`, recovers a password from the data dir `.env` or the install's service definition if one is set, and defines a fail-closed `delete_session`. It re-runs it every call because shell state does not survive between an agent's tool calls.
**2. Start a worker.**
```bash
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
```
`mode` is any of `claude`, `shell`, `opencode`, `codex`, `gemini`, `antigravity`.
**3. Wait until it is actually ready.** A new session reports `idle` before its CLI has spawned, and a brand-new case shows a trust dialog first, so the skill waits for the composer's own status bar and treats the dialog as a bounded fallback.
**4. Send a prompt and wait for the turn to end.**
```bash
BODY=$(jq -n --arg p "$PROMPT" '{input:($p+"\r"),useMux:true,clientId:"codeman-agent-1",seq:1,wait:true,waitTimeout:60000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' --data-binary "$BODY"
```
Send-and-wait registers the waiter before typing, which closes the race where a separate wait reports the previous turn's idle state as this turn's answer. For `claude` workers it resolves on the `stop` hook, typically within seconds.
**5. Read the answer.**
```bash
"${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text'
```
**6. Clean up.** `delete_session "$SID"`, for ids you created and nothing else.
Hook-less modes (`shell` and the external CLIs) have no `stop` signal and coarse lifecycle transitions, so the skill synchronizes those with a unique split marker and `wait-output ... from=buffer` instead. Worked fan-out flows, the per-mode signal table, error codes and the Docker/remote caveats live in the skill's `reference/` files, loaded on demand.
## The rules that bite
The skill documents these because each one silently wastes a run:
- **Every input must end with `\r`** or Enter is never sent and the text sits unsubmitted on the worker's prompt. `delivered:true` means "written to the pane", not "submitted".
- **Input is single-line.** Newlines are stripped.
- **A wait timeout is HTTP 200** with `wait.timedOut:true`, not an error. Loop over short waits; timeouts clamp to [1s, 600s] and the applied value comes back as `wait.timeoutMs`.
- **`stop` and `blocked` are `claude`-only.** Requesting them elsewhere is a 400.
- **Signals are edge-triggered with no history.** One that fires while no waiter is registered is unobservable afterwards, so never fire-and-forget N prompts and then gather signal-waits worker by worker.
- **Your typed command echoes into the output stream**, so a marker that appears verbatim in the input line matches before the command runs. Split it.
- **A full-screen TUI stream is space-less**, so match a single space-free token, never a phrase.
- **`pid != null` proves startup, not life.** A worker that dies inside its pane keeps `status:"idle"` and a pid. `wait?until=exit` is the death check.
## Fixes to the packaged skill
- **The self-delete guard failed open.** The old `is_self "$SID" || curl -X DELETE ...` shape meant an undefined `is_self` exited 127, the `||` branch fired, and the agent deleted its own session with the one guard bypassed. That is reachable because shell state does not survive between tool calls, so a partially re-pasted preamble was enough. The DELETE now lives inside a fail-closed `delete_session`, which also refuses an empty id and refuses when `$SELF` is unset or too short to prove the target is not the caller.
- **`clientId` was built from `$$`.** The pid changes between tool calls, so the documented "resend the identical request" loop stopped being recognized as a duplicate and retyped the prompt, submitting the turn twice. It is a fixed literal now.
- **`GET /api/v1/sessions/:id/last-response` was undocumented.** It returns the agent's final message as clean transcript text; the terminal scrape the skill previously recommended returns a wall of TUI repaint noise with the answer buried in it. It is now the documented read path for `claude` and `codex`, with the terminal buffer demoted to diagnosis and hook-less modes. Because the transcript flush lags the `stop` signal, the recipes poll it instead of reading once.
- **`quick-start` responses were never checked for `.success`.** On failure `.data.sessionId` is absent, `jq -r` prints the string `null`, and the flow burned its full readiness budget against `/api/v1/sessions/null` before reporting jq noise instead of the cause.
- **`codeman skill install --case <name>` could not resolve a linked case.** It hardcoded `~/codeman-cases/<name>` while the server resolves through `linked-cases.json` first, so it failed with "Case not found" for a case the web UI handled fine.
- **Documentation corrections**: `SESSION_BUSY` on `quick-start` is the 50-session cap rather than the waiter cap; `caseName` resolves linked cases, so a generic name can land a worker in a real repo; and the claim that a toggle-off sweep exists was wrong, so the per-case `skill uninstall` cleanup is now stated in both the README and the code.
## Also in this release
- **Terminal**: the wheel is no longer forwarded to codex, which ignores SGR mouse reports.
## 1.14.0
### Minor Changes
- Daemon mode and service install, plus subagent hook hardening and terminal/idle-checker fixes.
**New: run Codeman in the background without a terminal (#239, closes #231)**
- `codeman web -d` starts the server detached: it survives closing the shell, logs to `~/.codeman/web.log`, records a pidfile, and only reports success after the server actually answers `/api/status` (a port clash or missing dependency can never read as a clean start). `codeman web --status` and `codeman web --stop` manage it; `--stop` verifies the pid still looks like a Codeman server before signalling, so a recycled pid is never SIGTERMed.
- `codeman service install` / `status` / `uninstall`: installs a systemd user unit (Linux) or LaunchAgent (macOS) so the server comes back after reboots. The unit carries the installing shell's PATH (launchd's default PATH finds neither an nvm/Homebrew `node` nor `tmux`/`claude`), never contains `CODEMAN_PASSWORD`, and uses the same instance-scoped unit names as `install.sh` and the self-updater so no second copy can end up supervised.
- Both refuse to start a second server on one data dir (pidfile check plus a live probe): two servers on the shared tmux socket would attach to each other's sessions.
- Why `-d` exists at all: `nohup` does not protect a Node process, Node re-arms SIGHUP even when it inherits "ignore", so `nohup codeman web &` still dies on HUP. The detached relaunch (setsid) removes the controlling terminal instead.
**Subagent background-work hooks (#233, thanks @Lint111)**
- The background Bash rewake helper now also watches the top-level parent transcript when the hook fires inside a subagent: Claude records a subagent's Bash result in its own `subagents/agent-*.jsonl` but queues the completion in the lead session transcript, so subagents previously never woke. It can also inline a `CODEMAN_RESULT_BEGIN/END` marked report (up to 64 KiB) from the task output file into the wake feedback.
- New SubagentStop guard: a subagent that still owns live Monitor or background Bash processes is kept working instead of publishing an intermediate progress line as its final report. Ownership is verified against live process descriptors on `tasks/<id>.output`, so stale transcript text alone never blocks, and the guard fails open on systems without `/proc`.
- Existing cases self-heal to the new hooks on next launch.
**AI idle checker: stderr kept out of the verdict (#234, thanks @Lint111)**
The `claude -p` verdict command no longer merges stderr into the verdict file, where CLI warnings could turn a valid verdict into a parse error. On failures, the first 200 chars of stderr are attached to the diagnostic instead.
**Terminal: large final batches drain fully (#235, thanks @Lint111)**
A render-scheduling flag was cleared after the flush instead of before it, so when a large batch left a remainder behind, the remainder stayed unrendered until unrelated output arrived. This looked like truncated responses or shell commands that never finish. The flush now reschedules itself until the queue is empty.
**Docs and tests**
- README documents daemon mode and service install.
- Unique test port for the daemon-control suite.
## 1.13.0
### Minor Changes
+8 -4
View File
@@ -74,7 +74,7 @@ When user says "COM":
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 1.13.0 (must match `package.json`)
**Version**: 1.14.1 (must match `package.json`)
## Project Overview
@@ -109,6 +109,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| CI-equivalent test sweep | `npm run test:ci` (full suite minus browser/perf — see Testing) |
| Production start | `npm run start` |
| Production logs | `journalctl --user -u codeman-web -f` |
| Detached server | `codeman web -d` (`--status`, `--stop`; pidfile+log at `dataPath('web.pid'/'web.log')`). ⚠ Refuses to start a 2nd server on one data dir — see Instance isolation |
| Install/remove the service | `codeman service install` / `status` / `uninstall` (systemd user unit on Linux, LaunchAgent on macOS; names from `config/service-names.ts`) |
**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 3 Playwright tests). Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing).
@@ -141,7 +143,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| Domain | Key files | Notes |
| ---------------- | -------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| **Entry** | `src/index.ts`, `src/cli.ts` | |
| **Entry** | `src/index.ts`, `src/cli.ts`, `daemon-control`, `service-installer`, `config/service-names` | The last three back `web -d` / `service install` |
| **Session** | `src/session.ts` ★, `session-manager`, `session-auto-ops`, `session-cli-builder`, `session-task-cache`, `session-order` (pure), `session-pty-exit-breaker`, `usage-limit-patterns`, `usage-telemetry`; `src/services/unified-session-service.ts` | Pure/unit-tested helpers are split out of `session.ts` on purpose |
| **Mux** | `src/mux-interface.ts`, `src/mux-factory.ts`, `src/tmux-manager.ts` ★ | |
| **Respawn** | `src/respawn-controller.ts` ★ + 4 helpers (`-adaptive-timing`, `-health`, `-metrics`, `-patterns`) | Read `docs/respawn-state-machine.md` first |
@@ -180,7 +182,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Input**: `session.writeViaMux()` for programmatic/curl input via tmux `send-keys -l` + `send-keys Enter`, single-line only. Interactive **browser** input goes through a durable **exactly-once** layer: a stable `clientId` + monotonic per-session `seq` persisted to localStorage until the server ACKs, so a dropped link cannot lose or double-deliver a prompt. `ws-connection-registry.ts` supersedes only same-TAB reconnects, so two tabs on one session coexist. → [architecture-invariants#input-delivery-and-ws-resilience](docs/architecture-invariants.md#input-delivery-and-ws-resilience)
**Agent wait primitives**: bounded long-polls so an agent driving Codeman from a shell can block instead of poll: `GET /api/sessions/:id/wait` (lifecycle signal), `GET /api/sessions/:id/wait-output` (literal substring, **never** regex) and `wait`/`waitTimeout` on `POST /api/sessions/:id/input`. Registry in `session-wait-registry.ts` (pure, no `Session` reference), bounds in `config/agent-wait.ts`. ⚠️ **A timeout is a 200** (`wait.timedOut`), never an error, so callers loop over short waits. ⚠️ `stop`/`blocked` come from Claude Code hooks and therefore fire for **`claude` mode ONLY** (`shell` installs none either); asking for one explicitly on another mode is a 400, the default set silently drops them. ⚠️ Send-and-wait registers the waiter BEFORE the write (a separate POST-then-wait races and reports the PREVIOUS turn), and both teardown paths must `notifySignal('exit')` BEFORE `cancelAll()`. ⚠️ Client-hangup abort listens on **`reply.raw`** guarded by `writableFinished`: on `req.raw`, `close` fires when the request BODY ends, which on a POST killed every send-and-wait instantly and no `app.inject()` test could see it. ⚠️ Worker liveness cannot come from `session.pid` — for a tmux session that is the local attach client, which outlives a worker dying inside its pane — so it is probed at the mux layer (`isPaneDead`, ~750 ms cache) on blocking waits only, never on the input hot path. ⚠️ Signals are edge-triggered with no history: one that fires with no waiter registered is unobservable afterwards, so gather fan-outs with send-and-wait or latched `wait-output` markers, never fire-and-forget-then-sequential-signal-waits. → [architecture-invariants#agent-wait-primitives](docs/architecture-invariants.md#agent-wait-primitives), `docs/api-reference.md`
**Agent wait primitives**: bounded long-polls so an agent driving Codeman from a shell can block instead of poll: `GET /api/sessions/:id/wait` (lifecycle signal), `GET /api/sessions/:id/wait-output` (literal substring, **never** regex) and `wait`/`waitTimeout` on `POST /api/sessions/:id/input`. Registry in `session-wait-registry.ts` (pure, no `Session` reference), bounds in `config/agent-wait.ts`. ⚠️ **A timeout is a 200** (`wait.timedOut`), never an error, so callers loop over short waits. ⚠️ `stop`/`blocked` come from Claude Code hooks and therefore fire for **`claude` mode ONLY** (`shell` installs none either); asking for one explicitly on another mode is a 400, the default set silently drops them. ⚠️ Send-and-wait registers the waiter BEFORE the write (a separate POST-then-wait races and reports the PREVIOUS turn), and both teardown paths must `notifySignal('exit')` BEFORE `cancelAll()`. ⚠️ Client-hangup abort listens on **`reply.raw`** guarded by `writableFinished`: on `req.raw`, `close` fires when the request BODY ends, which on a POST killed every send-and-wait instantly and no `app.inject()` test could see it. ⚠️ Worker liveness cannot come from `session.pid` — for a tmux session that is the local attach client, which outlives a worker dying inside its pane — so it is probed at the mux layer (`isPaneDead`, ~750 ms cache) on blocking waits only, never on the input hot path. ⚠️ Signals are edge-triggered with no history: one that fires with no waiter registered is unobservable afterwards, so gather fan-outs with send-and-wait or latched `wait-output` markers, never fire-and-forget-then-sequential-signal-waits. The primitives are packaged as the **`skills/codeman` agent skill**: installable via `codeman skill install [--case <name>]` / `skill uninstall`, or auto-injected into a case's `.claude/skills/` on Claude session create behind `agentSkillEnabled` (SYNCED, default OFF). Injection is ADD-ONLY at create, marker-owned (`applyAgentSkill` in `hooks-config.ts` never touches an unmarked user copy) and refuses symlinks (this repo's own `.claude/skills/codeman` is a symlink to the source, which the injector must never write through). → [architecture-invariants#agent-wait-primitives](docs/architecture-invariants.md#agent-wait-primitives), `docs/api-reference.md`
**Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
@@ -210,7 +212,9 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of EACH session per page load requests `full=1` (`_fullHistoryLoaded` Set); tab switches keep the cheap `?tail=` path, and scrolling up at the TOP of the buffer re-pulls `full=1` on demand (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). ⚠️ That re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for codex/claude ≥ 2.1.187 at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). ⚠️ When that gate is FALSE on a claude session whose local buffer is hollow (`baseY === 0`), the gesture becomes coalesced PageUp/PageDown key sends (`_maybePageCliTranscript`) instead of a no-op; ⚠️ and `getClaudeCliVersion()` must never cache a FAILED probe (one timeout used to disable forwarding process-wide until restart). `_logScrollRouting()` prints the routing decision and its inputs once per session — read it before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
**Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for **claude ≥ 2.1.187 ONLY** at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. ⚠️ Codex was in that list and must never go back without a fresh measurement: codex-cli 0.147.0 ignores SGR wheel reports entirely (`mouse_any_flag=0`, inline viewport, transcript pushed into terminal scrollback), so forwarding produced a dead wheel (#227 follow-up). `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). ⚠️ When that gate is FALSE on a claude session whose local buffer is hollow (`baseY === 0`), the gesture becomes coalesced PageUp/PageDown key sends (`_maybePageCliTranscript`) instead of a no-op; ⚠️ and `getClaudeCliVersion()` must never cache a FAILED probe (one timeout used to disable forwarding process-wide until restart). `_logScrollRouting()` prints the routing decision and its inputs once per session — read it before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
**Detached start + service install** (issue #231): `codeman web -d` relaunches the SAME entry script with `detached:true` (setsid), so there is no controlling terminal and no shell job entry. ⚠️ `nohup` is NOT what makes this work: Node re-arms SIGHUP to its default disposition even when it inherits "ignore", and `cli.ts` handles SIGHUP with a graceful shutdown, so a delivered HUP still stops the server. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile check + `/api/status` probe): a second instance on the shared tmux socket attaches PTYs to the first one's live sessions. ⚠️ Neither may report success it has not observed — the parent polls `/api/status` until the child answers or dies, since `launchctl load` and a clean spawn are both silent about a server that starts and immediately exits. `--stop` verifies the pid still LOOKS like a Codeman server (`ps -o command=`) before signalling, because pids get recycled. Unit/label names live in `config/service-names.ts` so install.sh, `detectSupervisor()` and `service install` cannot drift into supervising two copies; they are instance-scoped, and identical to the historical names for the default instance. `service install` bakes the installing shell's PATH into the unit (launchd gives a job `/usr/bin:/bin:/usr/sbin:/sbin`, which finds neither a Homebrew/nvm `node` nor `tmux`/`claude`) and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install)
**Self-update** (App Settings → Updates): in-app updater for git-clone installs supervised by systemd/launchd (`systemd`, `launchd`, `launchd-daemon`, else `none` → "restart manually"). The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` that outlives the restart and writes progress to `update-status.json`, which the browser polls across the connection drop. `src/web/self-update.ts` splits pure helpers (unit-tested) from IO wrappers. npm installs report as non-updatable. → [architecture-invariants#self-update](docs/architecture-invariants.md#self-update)
+38 -2
View File
@@ -85,9 +85,29 @@ codeman web --multiuser # named logins + per-user case spaces
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
<details>
<summary><strong>Run as a background service</strong></summary>
<summary><strong>Keep it running in the background</strong></summary>
The installer's final menu sets this up for you (option 2) and verifies the service actually comes up before claiming success. To configure it manually instead:
To outlive the shell you started it in, without setting anything up:
```bash
codeman web -d # detach; logs to ~/.codeman/web.log
codeman web --status # is it up, and on which pid
codeman web --stop # graceful SIGTERM; agents keep running in tmux
```
`-d` waits until the server actually answers before reporting success, and refuses to start a second one on the same data dir (two servers sharing a tmux socket attach to each other's sessions).
To have it come back after a reboot, install it as a service instead. The installer's final menu does this for you (option 2); `codeman service` is the equivalent for an `npm i -g aicodeman` install:
```bash
codeman service install # systemd user unit (Linux) or LaunchAgent (macOS)
codeman service status
codeman service uninstall
```
`service install` writes the unit with your current PATH baked in, which matters more than it sounds: launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin`, so a Homebrew or nvm `node`, `tmux` or `claude` is invisible to a hand-written plist. It never copies `CODEMAN_PASSWORD` into the unit file; add that yourself if the service needs auth.
To write the unit by hand instead:
**Linux (systemd):**
@@ -220,6 +240,8 @@ codeman web # localhost:3000 (loopback only — safe defau
codeman web --port 8080 # custom port (or set CODEMAN_PORT)
codeman web --https # self-signed TLS (only needed for remote access)
codeman web -H 0.0.0.0 # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
codeman web -d # detach: survives closing the shell (--status, --stop)
codeman service install # systemd/launchd service: comes back after reboots
```
Open the printed URL. The page is a single dashboard; everything below happens there.
@@ -268,6 +290,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
### 7. Operate & maintain
- **App Settings** — model, effort, permission startup mode, theme/skin, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
- **Run it in the background** — `codeman web -d` detaches from your shell (`--status`, `--stop`); `codeman service install` makes it a systemd user unit / macOS LaunchAgent that survives reboots. Both verify the server actually answers before reporting success, and both refuse to start a second server on one data dir. See [Keep it running in the background](#quick-start---installation).
- **Self-update** — git-clone installs update in place from **Settings → Updates**.
- **Deploy your own changes** — see [Development](#development).
@@ -403,6 +426,7 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
## More Features
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
@@ -667,6 +691,18 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
For AI agents and automation that control Codeman without a browser: an agent that spins up worker sessions, a CI bot, or **Claude Code running _inside_ a Codeman session orchestrating other sessions**. Everything the UI does is HTTP + a CLI, so an agent can do it too.
> **Shortcut: install the packaged agent skill.** Everything below (plus worked multi-worker recipes) ships as a Claude Code skill in [`skills/codeman`](skills/codeman/SKILL.md), so an agent inside a session can drive Codeman without you pasting docs into the prompt. Three ways to get it:
>
> - `npx skills add Ark0N/Codeman --skill codeman -g`: global, works for any skills-aware agent
> - `codeman skill install` (global) or `codeman skill install --case <name>`: for npm installs that never cloned the repo; `codeman skill uninstall` reverses it
> - **App Settings → Agent Skill** (`agentSkillEnabled`, default off): Codeman then injects the skill into each case on Claude session create; a user-authored `skills/codeman` in the case is never overwritten
>
> A global install (`codeman skill install`, or `npx skills add`) is picked up by **every new Claude Code session on the machine**, inside Codeman or not. The skill self-gates: outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act, so a global install costs an idle session nothing.
>
> ⚠️ Turning `agentSkillEnabled` back off **does not remove already-injected copies** (a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` dir). Remove them per case with `codeman skill uninstall --case <name>`.
### Detect that you're inside Codeman
When a CLI runs in a Codeman-managed session, these environment variables are set — read them instead of hardcoding anything:
+52 -11
View File
@@ -1,8 +1,9 @@
# Agent Control Plan: skill packaging + wait primitives
**Status**: steps 1 to 5 IMPLEMENTED and multi-round verified, uncommitted as of 2026-08-08.
Step 6 (CLI install command + per-case injection + `agentSkillEnabled`) is not built.
See [§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
**Status**: steps 1 to 6 IMPLEMENTED, uncommitted as of 2026-08-09. Steps 1 to 5 were
multi-round verified on 2026-08-08; step 6 (CLI install command + per-case injection +
`agentSkillEnabled`) was built 2026-08-09; see the step-6 entry at the end of
[§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
verification round found, and what is still open.
**Date**: 2026-08-08
@@ -522,8 +523,8 @@ Bundled manifests plus local override only, no network.
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
| 5 | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
| 6 | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | settings partial-PUT test, case-creation test |
| 5 ✅ | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
| 6 ✅ | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | 10 unit tests (`test/agent-skill.test.ts`) + real-server case-creation tests (`test/quick-start.test.ts`, incl. the settings PUT accepting the key) green; CLI verified live (install/uninstall, global + `--case`, foreign/symlink refusals) |
| 7 | Docs: api-reference, extending-codeman, README | |
| 8 | COM (minor bump: new endpoints, new setting, new optional fields) | both CI and Release workflows green |
@@ -532,12 +533,14 @@ without the wait endpoints, so the wait work goes first.
## 6. Open questions for the owner
1. `skills/` at the repo root, accepted despite the short-root rule? (Recommended yes, the
install one-liner depends on it.)
2. `agentSkillEnabled` default: OFF for the first release then flip, or ON immediately?
3. Auto-inject the skill into every case's `.claude/skills/`, or global install only?
1. ✅ `skills/` at the repo root: accepted (built that way; the install one-liner depends on it).
2. ✅ `agentSkillEnabled` default: **OFF** for the first release, per §2.2's rationale (skills
cost context on every turn; measure before defaulting on). Flip later if dogfooding earns it.
3. ✅ Both: global install via `npx skills add` / `codeman skill install`, AND per-case
auto-injection behind the (default-off) setting. Injection is add-only at session create and
marker-guarded, so a user-authored copy is never touched.
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
and not a security boundary?
and not a security boundary? (Still open, not built with step 6.)
5. Regex support in `wait-output`: confirm literal-only for v1.
---
@@ -654,9 +657,47 @@ success without running its task. Two traps recurred often enough to name:
- **Release checklist**: `package.json` `files` includes `skills`, which is still
untracked. `git add skills/` must be part of the release commit, or npm publishes
a tarball without the skill (a `files` entry that does not exist is silently
ignored, so nothing fails).
ignored, so nothing fails). `test/agent-skill.test.ts` reads the packaged source,
so CI at least fails loudly if the directory goes missing from a checkout.
- The 1.13.0 changeset is written under `.changeset/`; consuming it (COM flow),
the release commit, and the deploy remain.
- Deferred with Part 3: the latched last-signal-per-turn. Nice-to-haves from the
reviews: N2 (create the death-watcher inside its `try`) and converting
timeout-shaped test detections into fast assertions.
- §2.4's `X-Codeman-Caller-Session` footgun guard: still not built (open question 4).
### Step 6 (2026-08-09): install command, per-case injection, the setting
Built to the §2.6 file list, mirroring the statusLine mechanism throughout:
| Piece | Where |
| ----- | ----- |
| `applyAgentSkill(casePath, enabled)` + `installAgentSkillInto` / `removeAgentSkillFrom` | `src/hooks-config.ts` |
| `codeman skill install` / `skill uninstall` (`--global` default, `--case <name>`) | `src/cli.ts` |
| `agentSkillEnabled` (SYNCED, default OFF) | `schemas.ts` (`SettingsUpdateSchema`), `getAgentSkillEnabled()` on `ConfigPort`/`server.ts`, checkbox in `index.html` + `settings-ui.js` |
| Injection call sites (Claude mode only) | `POST /api/sessions` next to `refreshStaleCodemanHooks`; `POST /api/quick-start` after the case-create/self-heal blocks (local + docker cases; remote skipped, its path lives on another host) |
| Tests | `test/agent-skill.test.ts` (10 unit), `test/quick-start.test.ts` (real server: default-off, PUT accepts key, injection on create, shell-mode skipped) |
Decisions worth keeping:
- **Ownership marker, prefix-matched.** The injected SKILL.md ends with
`<!-- codeman-managed-agent-skill: … -->`; install/refresh/remove all refuse a copy
without the marker (a user's own skill) and match on the PREFIX so a wording change
cannot disown older injected copies (the `BACKGROUND_WAKE_MARKER_PREFIX` pattern).
- **Symlink refusal.** This repo's own dogfooding layout
(`.claude/skills/codeman -> ../../skills/codeman`) means the injector must `lstat`
the skill dir AND its `skills/` parent and bail on a symlink, or enabling the
setting in the Codeman repo itself would overwrite the skill source through the link.
- **ADD-ONLY at session create**, same shared-`.claude` rationale as the statusLine:
a create while the setting is off must not yank the skill out from under other live
sessions in the repo. The remove path exists (CLI `skill uninstall`, tests); no
automatic sweep removes on toggle-off.
- **Removal is manifest-based, never `rm -rf`**: only files the packaged source would
have written are deleted, directories are pruned bottom-up only if they emptied, so
a user's extra notes in `reference/` survive an uninstall.
- **Source resolution**: `join(moduleDir, '..', 'skills', 'codeman')` works from
`src/` (tsx), `dist/` (tsc build), and the npm tarball alike, because all three sit
one level below the package root and `files` ships `skills/`.
- **Nothing acts on the setting at PUT time**: injection reads the merged persisted
settings at session create (`readSettings`, ~2s cache), so the partial-PUT invariant
(`toggleService` reading `merged`) is untouched by construction.
+19 -1
View File
@@ -80,7 +80,9 @@ Implementation detail extracted from `CLAUDE.md` so that file stays small enough
**Two strip flavors, one carry** (#205, `session.ts:_handleTerminalOutput`): the FULL strip (`isAltScreenStripMode` = codex/claude/gemini) removes alt-screen toggles, `3J`, and mouse-tracking DECSETs. Every other mode (shell/opencode/antigravity) gets the NARROW strip (`isMuxAltScreenOnlyStripMode`) — alt-screen toggles ONLY — and only when tmux-backed (`useMux`). Rationale: the tmux CLIENT emits `smcup` as its first bytes at attach, before any program runs, parking xterm in the scrollback-less alternate buffer for the whole session (touch scrolling no-ops; xterm's own wheel handler converts the wheel to Up/Down arrows = readline history cycling — both #205 symptoms). tmux never forwards a pane program's alt-screen toggles to its client (it repaints instead; measured — vim/less inside a pane emit zero to the client), so the only thing the narrow strip ever removes is tmux's own smcup. It keeps `3J` (a user's `clear` is a deliberate scrollback wipe) and the mouse DECSETs (tmux passes those through even with `mouse off`; stripping them would break htop/vim mouse support). ⚠️ The `useMux` gate is load-bearing: `startShell()`/`startInteractive()` fall back to a DIRECT PTY when mux creation fails, and there the inner program's own `?1049h` really does reach xterm — stripping it would break vim/less/htop for real. The replay path (`session-routes.ts`, via `session.usesMux`) applies the same narrow branch; the frontend `_sessionUsesServerMouseStrip()` mirror stays claude/codex/gemini because only the FULL strip touches mouse DECSETs. The chunk-boundary carry (`_altScreenSeqCarry`) runs for both flavors. Tests: `test/claude-scrollback-strip.test.ts`.
**Wheel/touch forwarding is NOT gated on viewport-at-bottom** (#205, `terminal-ui.js:_shouldForwardWheelToApp`): for sessions verified to scroll their own transcript on SGR wheel reports (codex, claude ≥ 2.1.187 — version via the local/docker/remote `--version` probes), the plain wheel AND touch drags forward as coalesced SGR reports (`_forwardScrollToApp` → `_sendSyntheticSgrWheel`, 40ms batches, 5-tick cap, 512-byte queue bound). It used to gate on the viewport being at the bottom so both scrollbacks stayed reachable, but a repaint-mode CLI keeps NO terminal scrollback of its own — xterm's buffer holds only replayed repaint frames, so local scrolling drags the CLI's pinned prompt box up the screen over stale frames; and `scrollToLastNonEmptyLine()` routinely parked the viewport off-bottom, silently pinning the wheel to local. Forwarding now snaps the viewport home first (SGR coordinates address the LIVE screen — a report computed from a scrolled-up viewport would hit-test the wrong row). Local scrollback remains on Shift+wheel and the `terminalWheelLocalScrollback` opt-out (both also cover touch via the shared gate; touch has no Shift, so the setting is its only local pin). `_wheelScrollLines()` normalizes `deltaMode` (Firefox fires LINE deltas ≈3/notch — read as pixels that rounded to 0 and fell to the ±1 fallback, ~4× too slow; PAGE deltas scale by `terminal.rows`) while keeping the #154 Shift-axis trap (macOS trackpads put Shift+scroll magnitude on deltaX). Tests: `test/terminal-touch-tap.test.ts`.
**Only claude ≥ 2.1.187 forwards the wheel; every other mode scrolls local scrollback** (#227 follow-up, `terminal-ui.js:_shouldForwardWheelToApp`). Codex was in the forward list until a reporter hit a completely dead wheel in codex tabs while the scrollbar drag worked. Measured against codex-cli 0.147.0 in a bare tmux: it never enables mouse tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change nothing on screen, because it runs an INLINE viewport (`alternate_on=0`) and pushes its transcript into the terminal's own scrollback (tmux `history_size` grows) instead of paging in-app. So for codex, local scrollback IS the transcript and forwarding swallowed every tick. ⚠️ "The TUI is a strip mode" is NOT evidence that it consumes wheel reports — verify with a real `\x1b[<64;c;rM` write into a live pane before adding a mode here. Hand-encoded SGR TAPS stay enabled for codex (`_sessionUsesServerMouseStrip`); measured, they are no-ops that insert nothing, so click-to-position is simply unavailable there rather than harmful.
**Wheel/touch forwarding is NOT gated on viewport-at-bottom** (#205, `terminal-ui.js:_shouldForwardWheelToApp`): for sessions verified to scroll their own transcript on SGR wheel reports (claude ≥ 2.1.187 — version via the local/docker/remote `--version` probes), the plain wheel AND touch drags forward as coalesced SGR reports (`_forwardScrollToApp` → `_sendSyntheticSgrWheel`, 40ms batches, 5-tick cap, 512-byte queue bound). It used to gate on the viewport being at the bottom so both scrollbacks stayed reachable, but a repaint-mode CLI keeps NO terminal scrollback of its own — xterm's buffer holds only replayed repaint frames, so local scrolling drags the CLI's pinned prompt box up the screen over stale frames; and `scrollToLastNonEmptyLine()` routinely parked the viewport off-bottom, silently pinning the wheel to local. Forwarding now snaps the viewport home first (SGR coordinates address the LIVE screen — a report computed from a scrolled-up viewport would hit-test the wrong row). Local scrollback remains on Shift+wheel and the `terminalWheelLocalScrollback` opt-out (both also cover touch via the shared gate; touch has no Shift, so the setting is its only local pin). `_wheelScrollLines()` normalizes `deltaMode` (Firefox fires LINE deltas ≈3/notch — read as pixels that rounded to 0 and fell to the ±1 fallback, ~4× too slow; PAGE deltas scale by `terminal.rows`) while keeping the #154 Shift-axis trap (macOS trackpads put Shift+scroll magnitude on deltaX). Tests: `test/terminal-touch-tap.test.ts`.
**A false gate on a Claude session must not mean a DEAD gesture** (#205 round 2, `_maybePageCliTranscript`): every way `_shouldForwardWheelToApp()` returns false leaves a repaint-mode pane scrolling a buffer that has nothing in it (`baseY === 0`) — the version probe came back empty, the CLI really is older than 2.1.187, or the user turned on `terminalWheelLocalScrollback`. The 1.12.0 retest reported exactly that: a wheel that did nothing at all while Fn+Up (PageUp) paged back through intact text, which is the proof that the CLI's own history and the PTY input path were both fine. So under the triple guard (claude mode + gate false + `baseY === 0`) wheel and touch travel is translated into coalesced `\x1b[5~` / `\x1b[6~` through the same 40ms queue as the SGR reports, at half a screen of travel per page key (the key jumps a whole screen; a 1:1 mapping was unusably slow with a discrete wheel). ⚠️ Shift is excluded on purpose — it is the explicit "give me local scrollback" gesture and must keep that meaning. ⚠️ `terminalWheelLocalScrollback` is deliberately NOT scoped away from repaint-mode CLIs even though it is a footgun there: that would silently override an explicit user choice, so the fallback catches it instead. **Server-side counterpart**: `getClaudeCliVersion()` caches SUCCESS for the process lifetime but must never cache FAILURE — it used to, so one timed-out or PATH-starved probe at the first Claude session start disabled wheel-forwarding for every Claude session until the server restarted (a dead wheel on phone, tablet and laptop at once, the signature of a server-side cause). Failures now retry with a 1/2/4…15min backoff; the policy is the pure `resolveClaudeCliVersion()`. Tests: `test/terminal-scroll-routing.test.ts`, `test/claude-cli-version-cache.test.ts`.
@@ -140,6 +142,22 @@ Tests: `test/file-editing-policy.test.ts` (pure policy), `test/routes/file-write
**Away digest** (COD-41/#136): `GET /api/away-digest?range=&since=&until=&lastViewed=` aggregates "what happened while you were away" from the lifecycle log + run-summary events + live sessions + daily token stats + recently-completed subagents into needs-attention/completed/still-running/idle/informational sections. Pure aggregator in `web/away-digest.ts` (`resolveAwayDigestRange()` validates the window — `since-last-visit`/`1h`/`today`/`24h`/`custom`, server-local TZ; `buildAwayDigest()` classifies). Header-button modal in `panels-ui.js` (button hidden on phones — regression-guarded). ⚠️ Returns `{success:true,digest}` (a legacy raw-ish shape, consistent with the other raw GET handlers in `system-routes.ts` — `{entries}`/`{config}`/`{files}`/`getSystemStats()`); frontend + tests read `.digest`. Subagent lookback is a fixed 60-min window regardless of range.
### Detached start and service install
**`codeman web -d` / `codeman service install`** (issue #231). Two answers to "keep it running", split by how long: `-d` survives the shell, the service survives a reboot. `src/daemon-control.ts` and `src/service-installer.ts`, both splitting pure builders (argv, URLs, pidfile parsing, unit-file text) from the IO.
Why a flag at all, when `nohup codeman web &` looks like it should work: it does not reliably. **Node re-arms SIGHUP to its default disposition even when it inherits "ignore" from `nohup`** (verified: `nohup node script.js &` then `kill -HUP` prints "Hangup" and dies; `/proc/<pid>/status` shows SIGHUP absent from `SigIgn`, where `nohup sleep` has it set). `cli.ts` then adds a SIGHUP handler that shuts the server down gracefully, so a delivered HUP always stops it. What actually works is removing the shell's ability to send one: `disown` in the user's shell, or `detached: true` (setsid) here. zsh HUPs running jobs on exit by default, bash does not on a clean `exit` but does when it receives SIGHUP itself, which is why "does `&` survive?" gets opposite answers on macOS and Linux.
Invariants:
- **Never a second server on one data dir.** `-d` and `service install` both check the pidfile AND probe `/api/status` first, and refuse. This is the instance-isolation hazard, not politeness: a second instance on the shared `tmux -L codeman` socket discovers the first one's live sessions, attaches PTYs to them and resizes them (see [Instance isolation](#instance-isolation-and-the-multi-instance-attach-danger)).
- **Never report success that was not observed.** The parent polls `/api/status` until the child answers or exits, then prints the URL or the tail of `web.log`. A `launchctl load`, a `systemctl enable --now` and a plain spawn are all silent about a server that starts and dies half a second later, which is why `install.sh` verifies too. The log is append-only across launches, so each start writes a separator line and the failure tail begins there.
- **A 401 counts as up.** `CODEMAN_PASSWORD` gates `/api/status`, so requiring a 200 would make readiness detection fail on exactly the installs that took security advice. The body is still checked for `"success"` so an unrelated service squatting on the port is not mistaken for Codeman.
- **`--stop` verifies identity before signalling.** Pids are recycled; a stale pidfile plus a blind `process.kill` is how a tool SIGTERMs someone's database. `ps -o command=` (portable to macOS) must still look like a Codeman web process. When `ps` itself fails, the pid is treated as ours rather than orphaning the pidfile.
- **One source of truth for the job name.** `src/config/service-names.ts` holds the systemd unit name and launchd label used by `install.sh`, `detectSupervisor()` in self-update, and `service install`. Drift here is silent and bad: `service install` would supervise a SECOND copy alongside the installer's. The names are instance-scoped (`CODEMAN_INSTANCE=beta` → `codeman-web-beta.service` / `com.codeman.beta.web`) so a beta cannot overwrite the production unit, and are byte-identical to the historical names for the default instance.
- **PATH is the reason hand-written units fail.** launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` and systemd's user manager is nearly as bare, so a Homebrew or nvm `node`, `tmux` or `claude` is simply absent. The unit therefore carries the installing shell's PATH with the running node's directory in front; `node_modules/.bin` entries are dropped, since npx injects those for one command and they would outlive the checkout.
- **No secrets in unit files.** `CODEMAN_PASSWORD` present in the installing shell is NOT copied into the plist/unit; the operator is told to add it. `CODEMAN_INSTANCE` IS copied, because without it a supervised beta would silently run against the production data dir and tmux socket.
### Self-update
**Self-update** (App Settings → Updates): in-app updater for **git-clone installs** supervised by systemd/launchd. Supervisors: `systemd` (user unit), `launchd` (GUI LaunchAgent, gui-domain kickstart), `launchd-daemon` (KeepAlive system LaunchDaemon on headless Macs — restarts rootlessly by killing the server PID and letting launchd respawn it; detected only when the daemon is bootstrapped AND KeepAlive), else `none` → "restart manually" message; on next boot a manual-restart status auto-completes when the running version matches the target. The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` (`git checkout <release tag> && npm install && npm run build && restart`) that outlives the restart; it writes progress to `dataPath('update-status.json')`, which the browser polls across the connection drop. Channel = latest `codeman@X.Y.Z` release tag; dirty trees are auto-stashed. `src/web/self-update.ts` splits PURE helpers (semver/tag parsing, reconcile decision — unit-tested) from IO wrappers (`getInstallInfo`/`checkForUpdate`/`startUpdate`/`reconcileUpdateOnBoot`). Routes: `GET /api/system/update/check`, `POST /api/system/update`, `GET /api/system/update/status`. Types: `src/types/update.ts`. npm installs report as non-updatable.
+21 -3
View File
@@ -149,9 +149,17 @@ to Claude as a system reminder. This implies `"async": true`; ordinary async
hooks do not wake an idle turn, and their output waits for the next interaction.
Codeman uses this on `PostToolUse(Bash)`: a self-contained Node helper extracts
the background task ID from the Bash result, watches the session transcript for
the matching completion notification, and exits 2. It does not send terminal
input, so it cannot submit a user's partially written prompt.
the background task ID from the Bash result, watches the originating transcript
and, for subagents, the top-level parent transcript for the matching completion
notification, and exits 2. Claude records a subagent's Bash result in its
`subagents/agent-*.jsonl` file but queues completion in the lead session JSONL.
The task ID keeps each wake targeted. The helper does not send terminal input,
so it cannot submit a user's partially written prompt.
For script-dispatched Codex work, `codex-run.sh` writes the final response
between `CODEMAN_RESULT_BEGIN/END` markers in the background task output. The
rewake helper includes a maximum of 64 KiB of that report in its feedback. UI
subagent discovery and dispatcher result delivery are separate contracts.
### Notification
@@ -219,6 +227,16 @@ Or to allow exit:
**Use Cases**: Control nested loops, verify subagent output.
The hook input includes `agent_id`, `agent_transcript_path`, and
`last_assistant_message`. Like `Stop`, a command hook can return
`{"decision":"block","reason":"..."}` to keep the subagent running and feed
the reason back to it.
Codeman uses this to prevent premature reports from workers that still own live
Monitor or background-Bash processes. It derives candidate task IDs from the
subagent transcript, but requires a matching live Linux process descriptor for
`tasks/<id>.output`; historical task text by itself is not treated as active.
### TeammateIdle
**When**: When an agent-team teammate is about to go idle.
+7
View File
@@ -160,6 +160,13 @@ Around 200 handlers across 21 route files cover sessions, cases, files, cron,
respawn, Ralph, the orchestrator, search, and admin. Each route module carries an
`@fileoverview` describing its endpoints.
If the caller is an agent running _inside_ a Codeman session, install the packaged
agent skill instead of teaching it these calls by hand: `skills/codeman` in the repo
(`npx skills add Ark0N/Codeman --skill codeman -g`, or `codeman skill install
[--case <name>]`, or the synced `agentSkillEnabled` App Setting for automatic
per-case injection on Claude session create). The skill carries the guard, the
safety rules, and verified wait/orchestration recipes.
The common ones:
```bash
+51
View File
@@ -163,6 +163,57 @@ both self-reporting, so the retest ask is now "open the console and paste the `[
- iPhone: Claude or shell session, and whether a full tab kill changes anything.
- Browser console: `app.terminalUi?.terminal?.modes?.mouseTrackingMode` (false-path 4).
## ROUND 3 (2026-08-09): Codex wheel dead — CONFIRMED AND FIXED
DodgyBadger (Codex latest, Chrome, Windows 11): mouse wheel does nothing in a CODEX session
while working fine in shell and web tabs; DRAGGING THE SCROLLBAR WORKS, so xterm's local
buffer demonstrably has content for their codex pane. Analysis against the shipped code:
- `_shouldForwardWheelToApp` returns true UNCONDITIONALLY for `codex` (no version gate, unlike
claude's `>= 2.1.187`), so every plain wheel tick is sent as SGR reports to Codex.
- The "verified to scroll its transcript on SGR wheel reports" claim for codex predates
current Codex builds; if Codex latest ignores SGR wheel, forwarding eats the gesture while
the healthy local scrollback (proven by the working scrollbar) sits unused.
- The #227 PageUp fallback cannot rescue this: it is gated to `claude` mode AND `baseY === 0`,
and codex here has real local scrollback. The `[scroll]` diagnostic will still say
`forward-sgr (mode=codex, ...)`, confirming the branch, worth asking the reporter to paste.
**CONFIRMED by the reporter's `[scroll]` line (2026-08-09, PR #227 comment)**:
`forward-sgr (mode=codex, cliVersion=unknown, localScrollbackOptOut=false, mouseTracking=none,
localScrollbackRows=967)`. Forwarding branch active, 967 rows of healthy local scrollback
unused, Codex ignoring the SGR reports. Environment: Codex latest, Chrome, Windows 11.
**Measured against codex-cli 0.147.0** (isolated `tmux -L codexwheel`, fake `CODEX_HOME/auth.json`,
history built with 401ing prompts), which settles it without needing a version gate at all:
| Probe | Result |
| ---------------------------------------------- | ----------------------------------------------- |
| `#{mouse_any_flag}` once the TUI is up | `0`: codex never enables mouse tracking |
| `#{alternate_on}` | `0`: inline viewport, not an alt-screen pager |
| `#{history_size}` while prompting | grows 3 → 32: the transcript goes to scrollback |
| 6 × `\x1b[<64;10;10M` written to the pane | pane capture byte-identical, nothing happens |
| control: literal `zz` | pane changes, so the probe can see changes |
| `\x1b[<0;12;5M` + release (the click-tap path) | no change either: taps are no-ops, not garbage |
Codex has no in-app pager to drive: its history lives in the terminal's own scrollback, which is
exactly what forwarding was stealing the gesture from. A version gate would be the wrong fix (and
`cliVersion=unknown` means there is no codex probe to gate on anyway).
**Fix (shipped):** `_shouldForwardWheelToApp` now returns true for `claude >= 2.1.187` and nothing
else. Codex falls to the normal local-scrollback path like shell/gemini/opencode, so wheel and touch
scroll the same history the scrollbar drag was already scrolling. The claude-only PageUp fallback is
untouched: codex never needs it, its local buffer is real. Taps stay hand-encoded for codex
(`_sessionUsesServerMouseStrip`), measured harmless, so click-to-position is merely unavailable
there rather than damaging. Lesson for the next mode added to the forward list: "it is a strip mode"
proves nothing, write a real SGR report into a live pane and diff the capture first.
Verified end-to-end in Chromium against a live codex session on an isolated instance
(`CODEMAN_INSTANCE=codexwheel`, port 5055, `envOverrides.CODEX_HOME` pointing at the fake auth
dir): trusted `page.mouse.wheel` up now logs
`[scroll] … → local-scrollback (mode=codex, …, localScrollbackRows=43)`, moves the viewport
39 → 4 (back to the Codex banner), and sends ZERO bytes to the PTY. Unit coverage:
`test/terminal-touch-tap.test.ts` ("only claude forwards — codex and gemini keep the local wheel").
Original plan follows.
## Reports
+10 -28
View File
@@ -43,38 +43,22 @@ const syncData = DEC_SYNC_START + data + DEC_SYNC_END;
this.broadcast('session:terminal', { id: sessionId, data: syncData });
```
## Client-Side Implementation (`app.js`)
## Client-Side Implementation (`terminal-ui.js`)
### `batchTerminalWrite(data)`
1. Checks if flicker filter is enabled (optional, per-session)
2. If flicker filter active: buffers screen-clear patterns (`ESC[2J`, `ESC[H ESC[J`, `ESC[nA`)
3. Accumulates data in `pendingWrites`
4. Schedules `requestAnimationFrame` if not already scheduled
5. On rAF callback: checks for incomplete sync blocks (start without end)
6. If incomplete: waits up to 50ms via `syncWaitTimeout`
7. Calls `flushPendingWrites()` when complete
### `extractSyncSegments(data)`
- Parses DEC 2026 markers, returns array of content segments
- Content before sync blocks returned as-is
- Content inside sync blocks returned without markers
- Incomplete blocks (start without end) returned with marker for next chunk
4. Calls `_scheduleTerminalWriteFlush()` if no flush is pending
5. The yielded callback clears its scheduled flag before calling `flushPendingWrites()`
6. Large batches schedule their own next chunk until the queue is empty
### `flushPendingWrites()`
```javascript
const segments = extractSyncSegments(this.pendingWrites);
this.pendingWrites = ''; // Clear before writing
for (const segment of segments) {
if (segment && !segment.startsWith(DEC_SYNC_START)) {
terminal.write(segment); // Skip incomplete blocks (start with marker)
}
}
```
Note: Segments starting with `DEC_SYNC_START` are incomplete blocks awaiting more data. These are skipped (discarded if timeout forces flush).
- Joins the queued terminal data and passes DEC 2026 markers through to xterm.js 6, which handles synchronized output natively.
- Writes at most 32KB per yield for Codex and 64KB for other modes.
- Requeues the remainder and immediately schedules another safe yield. A final large response therefore drains without waiting for another SSE event.
### `chunkedTerminalWrite(buffer, chunkSize=128KB)`
@@ -116,17 +100,15 @@ When detected, buffers 50ms of subsequent output before flushing atomically.
## Edge Cases
- **Incomplete sync blocks**: 50ms timeout forces flush (content discarded to prevent freeze)
- **Incomplete sync blocks**: xterm.js retains synchronized output until its closing marker
- **Large buffers**: Chunked writing prevents UI freeze
- **Server shutdown**: Skips batching via `_isStopping` flag
- **Session switch**: Clears flicker filter state, pending writes, and sync timeout (prevents cross-session data bleed)
- **SSE reconnect**: `handleInit()` clears all pending write state
**Trade-off:** If a sync block is split across SSE packets and the end marker doesn't arrive within 50ms, the incomplete content is discarded. This prioritizes responsiveness over completeness. In practice this is rare since the server always sends complete `SYNC_START...SYNC_END` pairs and SSE typically delivers them atomically.
## DEC Mode 2026 Compatibility
Terminals that natively support DEC 2026 will buffer and render atomically. Terminals that don't support it ignore the escape sequences harmlessly. xterm.js doesn't support DEC 2026 natively, so the client implements its own buffering by parsing the markers.
Terminals that natively support DEC 2026 buffer and render atomically. Codeman uses xterm.js 6, so the client passes the markers through instead of parsing or discarding partial blocks.
**Supporting terminals:** WezTerm, Kitty, Ghostty, iTerm2 3.5+, Windows Terminal, VSCode terminal
@@ -135,4 +117,4 @@ Terminals that natively support DEC 2026 will buffer and render atomically. Term
| File | Key Functions |
|------|---------------|
| `src/web/server.ts` | `batchTerminalData()`, `flushTerminalBatches()`, `broadcast()` |
| `src/web/public/app.js` | `batchTerminalWrite()`, `extractSyncSegments()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
| `src/web/public/terminal-ui.js` | `batchTerminalWrite()`, `_scheduleTerminalWriteFlush()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.13.0",
"version": "1.14.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.13.0",
"version": "1.14.1",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.13.0",
"version": "1.14.1",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
+94 -24
View File
@@ -17,7 +17,24 @@ sessions. Every recipe below was verified live. Full endpoint tables and
troubleshooting: [reference/endpoints.md](reference/endpoints.md). Worked multi-worker
flows: [reference/recipes.md](reference/recipes.md).
## 0. Guard — run this before anything else
## 0. Guard, and the one thing that breaks every recipe below
⚠️ **Your shell state does not survive between tool calls.** Each Bash call starts a
fresh shell, so `$API`, `$SELF`, the `CURL` array and `delete_session` are all gone by
the next call, and `$$` is a different pid. Three consequences, all of which have
teeth:
- **Re-run this entire preamble at the top of every Bash call that touches the API.**
Running it once and assuming it stuck is the single most likely way to break a run.
- **Never re-paste only half of it.** The delete guard below is written so that a
missing definition deletes nothing, but that only holds if you never hand-roll a
`DELETE` of your own.
- **Never put `$$` in a `clientId`.** It changes per call, so the "resend the identical
request" loop in §3 would stop being a duplicate and would **retype the prompt**,
submitting the turn twice. Use a fixed literal (`codeman-agent-1` below).
Only real environment variables (`CODEMAN_*`) survive, which is why this preamble
rebuilds everything else from them.
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "Not inside a Codeman-managed session; refusing to act."; exit 1; }
@@ -39,13 +56,35 @@ if [ -z "${CODEMAN_PASSWORD:-}" ]; then # stock installs: install.sh puts it
UNIT="$HOME/.config/systemd/user/codeman-web.service"
PLIST="$HOME/Library/LaunchAgents/com.codeman.web.plist"
if [ -f "$UNIT" ]; then
CODEMAN_PASSWORD=$(sed -n 's/^Environment="CODEMAN_PASSWORD=\(.*\)"$/\1/p' "$UNIT" | head -1)
# install.sh backslash-escapes " and \ in the unit value; undo it or a password
# containing either recovers wrong and auth fails.
CODEMAN_PASSWORD=$(sed -n 's/^Environment="CODEMAN_PASSWORD=\(.*\)"$/\1/p' "$UNIT" | head -1 | sed 's/\\\(["\\]\)/\1/g')
elif [ -f "$PLIST" ]; then
CODEMAN_PASSWORD=$(awk '/<key>CODEMAN_PASSWORD<\/key>/{getline; print}' "$PLIST" | sed -n 's/.*<string>\(.*\)<\/string>.*/\1/p')
# install.sh XML-escapes the plist value; undo it (&amp; LAST, mirroring escape order).
CODEMAN_PASSWORD=$(awk '/<key>CODEMAN_PASSWORD<\/key>/{getline; print}' "$PLIST" | sed -n 's/.*<string>\(.*\)<\/string>.*/\1/p' \
| sed -e 's/&lt;/</g' -e 's/&gt;/>/g' -e 's/&amp;/\&/g')
fi
fi
AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:$CODEMAN_PASSWORD")
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on https (self-signed cert)
# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
# `is_self "$SID" || curl -X DELETE ...` shape failed OPEN, because an undefined
# is_self exits 127 and the `||` branch then ran the delete completely unguarded.
# Undefined delete_session is "command not found", which deletes nothing.
delete_session() {
local id="${1:-}"
[ -n "$id" ] || { echo "refusing: empty session id"; return 1; }
[ "${#SELF}" -ge 8 ] || { echo "refusing: \$SELF unset or too short to prove this is not me"; return 1; }
# ids appear in full AND 8-char form (Docker exports a truncated $SELF; mux names and
# UI surfaces carry 8-char ids), so compare by prefix in BOTH directions. Equality or
# a one-directional check each miss a real combination, and the miss deletes you.
case "$id" in "$SELF"*) echo "refusing: $id is me"; return 1 ;; esac
case "$SELF" in "$id"*) echo "refusing: $id is me"; return 1 ;; esac
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$id"
}
CID=codeman-agent-1 # FIXED literal, never "agent-$$" (see §0)
```
- If `CODEMAN_MUX` is not `1`, **stop and say so**. Do not guess an API URL; a server
@@ -67,20 +106,16 @@ CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on https (self-s
You are yourself a session on this server, and the API has **no undo**.
- **Never act on your own session — and know that this check is the ONLY guard.**
- **Never act on your own session, and know that `delete_session` is the ONLY guard.**
The server has no self-protection: a session that DELETEs its own id succeeds and
dies silently (verified live). Session ids appear in both full and 8-character
forms (Docker cases export a truncated `$SELF`; mux names and UI surfaces carry
8-char ids), so compare by prefix **in both directions**, never by equality:
```bash
is_self() { case "$1" in "$SELF"*) return 0 ;; esac; case "$SELF" in "$1"*) return 0 ;; esac; return 1; }
```
One-directional or equality checks each miss a real combination (full `$SELF` vs
a target you transcribed in 8-char form, or truncated `$SELF` vs a full target)
and the miss deletes you. Check `is_self` before every `DELETE`, kill, respawn,
or input call.
dies silently (verified live). **Always delete through `delete_session "$SID"` from
§0; never write a bare `curl -X DELETE` and never reintroduce the
`is_self … || curl -X DELETE …` shape.** That older form failed open: with the
function undefined (a half-re-pasted preamble, see §0) bash returns 127, the `||`
branch fires, and the delete runs with no self-check at all. Wrapping the request
inside the guard is what makes a lost preamble delete nothing instead of deleting
you. Apply the same prefix-both-directions reasoning before any kill, respawn, or
input call you write by hand.
- **Mutating calls you may make unprompted** (this is an allowlist):
`POST /api/v1/quick-start`, `POST /api/v1/sessions/:id/input`, and
`DELETE /api/v1/sessions/:id` **only** for a session you created in this
@@ -162,15 +197,25 @@ the composer is not) and always pays it in full before the fallback runs — the
budget belongs to stage 3, after the dialog is answered:
```bash
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
# ALWAYS check .success: on failure `.data.sessionId` is null, jq -r prints the string
# "null", and the flow below then burns its full readiness budget against
# /api/v1/sessions/null before reporting jq noise instead of the actual cause.
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
if [ -z "$SID" ]; then
# SESSION_BUSY here is the 50-session cap, not the waiter cap; FORBIDDEN/CONFLICT/
# OPERATION_FAILED/INVALID_INPUT are the others. None are retryable in a loop.
jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed; stopping."
exit 1
fi
for _ in $(seq 1 30); do # bounded: a bad SID would otherwise poll forever
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# ⚠️ pid != null proves STARTUP only, never life: a worker that later dies inside
# its pane keeps status "idle" and a pid (the local tmux attach client, not the
# worker). The death check is wait?until=exit, below.
CID="agent-$$"; SEQ=1
SEQ=1 # $CID came from the §0 preamble; do NOT rebuild it from $$
# the composer's status bar ("bypass permissions on") is the ready marker — Codeman
# spawns claude in bypass mode. Single-token matches only: TUI text is space-less.
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
@@ -245,15 +290,40 @@ SEQ=$((SEQ+1))
The typed line shows `${M}_…`, the real output shows `DONE_… rc=<exit code>`, and the
snippet carries the exit code back to you.
**Read a worker's output** — the terminal buffer, tail in **bytes** (`textOutput` in
`GET .../output` stays empty for interactive sessions; don't use it):
**Read a worker's answer.** For `claude` and `codex` workers this is the read path:
`last-response` returns the agent's final message as clean text, taken from the
transcript rather than the screen, so it carries none of the TUI's box-drawing or
repaint noise.
```bash
for _ in $(seq 1 10); do # the transcript write LAGS the stop signal
TXT=$("${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
```
`.data` is `{text, timestamp}`. ⚠️ **Poll it, do not read it once.** `text` is written
from the transcript file, which is flushed slightly *after* the `stop` hook fires, so a
single read taken the instant send-and-wait returns comes back `""` even though the
turn finished (verified live: empty on the first call, full text seconds later). `text`
is also `""` before the worker's first completed turn, and always `""` for modes with
no transcript (`shell`, `opencode`, `gemini`, `antigravity`, verified live), which is
why the loop above is bounded rather than open-ended. Fall back to the terminal buffer there, tail in **bytes**
(`textOutput` in `GET .../output` stays empty for interactive sessions; don't use it):
```bash
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=3000" | jq -r '.data.terminalBuffer' \
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g' | grep -v '^[[:space:]]*$' | tail -30
```
Avoid `?full=1` (entire tmux scrollback, a context bomb) unless doing a post-mortem.
⚠️ Do not use that pipeline to read a **claude/codex** answer. A full-screen TUI draws
with cursor moves, so the stripped buffer is largely one long line: `tail -30` has
almost nothing to split on and you get a wall of repaint noise with the answer buried
in it (verified live, side by side with `last-response` returning the exact prose).
The terminal buffer is for *diagnosis* (is my prompt sitting unsubmitted?), not for
reading answers. Avoid `?full=1` (entire tmux scrollback, a context bomb) unless doing
a post-mortem.
**Detect a dead worker cheaply**: `GET .../wait?until=exit&timeout=60000` answers
immediately (`signal:"exit"`, `immediate:true`) if the PTY is gone — including a
@@ -262,10 +332,10 @@ as `status:"idle"` with a pid (that pid is the local tmux attach client, not the
worker). The wait routes are the only liveness check; a worker dying while a wait
is parked resolves it within ~3 s. A session deleted mid-wait resolves in ~1 s.
**Clean up** — only ids you created, `is_self`-checked, one at a time:
**Clean up** — only ids you created, one at a time, always through the §0 helper:
```bash
is_self "$SID" || "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
delete_session "$SID"
```
Everything else (endpoint tables, per-mode signal table, error codes, capacity
+21 -6
View File
@@ -14,7 +14,7 @@ Every JSON response: `{"success":true,"data":…}` or
| `INVALID_INPUT` | 400 | malformed request; the message names the bad field |
| `UNAUTHORIZED` | 401 | auth required or failed (send `-u user:password`). ⚠️ The 401 body is plain text, NOT this envelope — `jq` dies with a parse error, see the guard in SKILL.md |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap (16, combined signal+output) is full |
| `SESSION_BUSY` | 409 | on a **wait**: this session's waiter cap (16, combined signal+output) is full. On **quick-start**: the 50-session cap is full, so clean up before starting more |
| `CONFLICT` / `ALREADY_EXISTS` | 409 | conflicts with current state |
| `OPERATION_FAILED` | 422 | well-formed but could not be completed |
| `RATE_LIMITED` | 429 | per-owner or process-wide waiter pool is full — back off; switching sessions will not help |
@@ -32,16 +32,19 @@ Every JSON response: `{"success":true,"data":…}` or
| unified list incl. history | `GET /api/v1/sessions/unified` → `.data.sessions[]` (NOT `.data[]`), and it folds in transcript history from the whole machine — never use it to verify cleanup; `GET /api/v1/sessions` is the cleanup check |
| start case + session in one call | `POST /api/v1/quick-start` |
| send input | `POST /api/v1/sessions/:id/input` |
| read terminal (tail is in **BYTES**, raw ANSI) | `GET /api/v1/sessions/:id/terminal?tail=3000` → `.data.terminalBuffer` |
| **read a worker's answer** (claude/codex) | `GET /api/v1/sessions/:id/last-response` → `.data.{text,timestamp}` — clean transcript text, no TUI noise. ⚠️ **Poll it**: the transcript flush lags the `stop` signal, so a read taken the instant send-and-wait returns is `""` (verified live). Also `""` before the first completed turn, and always `""` for `shell`/`opencode`/`gemini`/`antigravity` (no transcript) |
| read terminal (tail is in **BYTES**, raw ANSI) | `GET /api/v1/sessions/:id/terminal?tail=3000` → `.data.terminalBuffer` — for *diagnosis* (unsubmitted prompt?), not for reading answers |
| full tmux scrollback (context bomb; post-mortems only) | `GET /api/v1/sessions/:id/terminal?full=1` |
| background agents of a session | `GET /api/v1/subagents` |
| background agents, one session | `GET /api/v1/sessions/:id/subagents` |
| background agents, global list | `GET /api/v1/subagents` (admin-only in multi-user mode) |
| server status / version | `GET /api/v1/status` → `.data.version` |
| delete one session (yours, `is_self`-checked) | `DELETE /api/v1/sessions/:id` |
| delete one session (yours only, via `delete_session`) | `DELETE /api/v1/sessions/:id` — never call it bare; the fail-closed helper in SKILL.md §0 is the only self-protection that exists |
⚠️ `GET /api/v1/sessions/:id/output` → `.data.textOutput` looks like the obvious read
but stays **empty for interactive tmux-backed sessions** (it is fed only by the legacy
JSON-stream path). Verified empty on live claude and shell sessions. Read
`terminal?tail=` instead and strip ANSI:
JSON-stream path). Verified empty on live claude and shell sessions. Use
`last-response` for claude/codex answers; only fall back to `terminal?tail=` for
hook-less modes, or to diagnose a prompt that was never submitted, and strip ANSI:
```bash
… | jq -r '.data.terminalBuffer' | sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g'
@@ -53,6 +56,18 @@ JSON-stream path). Verified empty on live claude and shell sessions. Read
`.data.{sessionId, caseName, casePath}`. Creates the case directory (a real directory
on the user's disk) if missing — do not retry it in a loop, and remember the name.
⚠️ **Branch on `.success` before reading `.data.sessionId`.** On any failure the field
is absent, `jq -r` prints the literal string `null`, and every later call then targets
`/api/v1/sessions/null`, burning the full readiness budget and reporting jq noise
instead of the real cause. Failure modes here are `SESSION_BUSY` (the **50-session
cap**, not the waiter cap), `FORBIDDEN`, `CONFLICT`, `OPERATION_FAILED` and
`INVALID_INPUT`; none of them are retryable in a loop.
⚠️ `caseName` resolves through the linked-cases registry first, so a name that happens
to match a case the user linked in lands in that **real repo**, not a fresh scratch
directory. Pick distinctive scratch names, and use a linked name deliberately when you
do want a worker in an existing checkout.
`POST /api/v1/sessions/:id/input` body:
`{"input":"one line\r","useMux":true,"clientId":"agent-1","seq":1}` plus optionally
`"wait"` / `"waitTimeout"` (below).
+55 -25
View File
@@ -1,10 +1,17 @@
# Worked orchestration flows
Loaded on demand from the `codeman` skill. Every flow assumes the guard preamble from
SKILL.md ran (`$API`, `$SELF`, `"${CURL[@]}"`, `is_self`). Track every session id you
create; delete them (and only them) when done. Remember the two silent killers:
**every input ends with `\r`**, and **markers must be split** so the typed-line echo
does not match them.
Loaded on demand from the `codeman` skill. Every flow assumes the SKILL.md §0 preamble
is in scope (`$API`, `$SELF`, `$CID`, `"${CURL[@]}"`, `delete_session`).
⚠️ **That preamble does not survive between tool calls**, so re-run it at the top of
every Bash call that uses these flows, in full. Re-pasting only part of it is the
failure mode the fail-closed `delete_session` exists to contain, and a `clientId` you
rebuild from `$$` changes per call, which turns the duplicate-resend loop in Flow 1
into a second typed prompt.
Track every session id you create; delete them (and only them) when done. The two
silent killers: **every input ends with `\r`**, and **markers must be split** so the
typed-line echo does not match them.
## Flow 1: claude worker, end to end
@@ -13,11 +20,15 @@ the turn to finish, read the answer, clean up. Verified live: the stop hook reso
the send-and-wait within seconds of the turn ending.
```bash
# 1. start (returns before the CLI inside is ready)
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-tests","mode":"claude"}' | jq -r '.data.sessionId')
# 1. start (returns before the CLI inside is ready). ALWAYS check .success: on failure
# .data.sessionId is null, jq -r yields the string "null", and every step below
# then runs against /api/v1/sessions/null and reports jq noise, not the cause.
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-tests","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed"; exit 1; }
CREATED+=("$SID") # the cleanup list
CID="agent-$$"; SEQ=1
SEQ=1 # $CID is the fixed literal from §0; never rebuild it from $$
# 2. readiness. "wait for idle" or "wait for ❯" is NOT readiness: a fresh session
# reports idle before anything spawned, and the first-run trust dialog contains ❯.
@@ -84,13 +95,23 @@ case "$(jq -r '.data.wait.signal' <<<"$R")" in
null) jq -e '.data.wait.ended' <<<"$R" >/dev/null && echo "worker deleted mid-wait" ;;
esac
# 5. read the answer: terminal tail (BYTES), ANSI-stripped. textOutput stays empty
# for interactive sessions; terminal?full=1 is a context bomb.
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=4000" | jq -r '.data.terminalBuffer' \
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g' | grep -v '^[[:space:]]*$' | tail -30
# 5. read the answer. For a claude worker this is last-response: clean transcript text,
# no TUI repaint noise. Do NOT scrape the terminal for this — a full-screen TUI
# draws with cursor moves, so the stripped buffer is nearly one long line and the
# answer arrives buried in redraw garbage.
# POLL it: the transcript flush lags the stop signal, so a single read taken the
# instant step 3 returned comes back "" even though the turn finished (verified live).
for _ in $(seq 1 10); do
TXT=$("${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
# (.data is {text,timestamp}; text is also "" before the first completed turn and
# always "" for shell/opencode/gemini/antigravity, which have no transcript — use
# the terminal tail there, and here only to diagnose an unsubmitted prompt.)
# 6. clean up — exact id, own list only, self-check
is_self "$SID" || "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
# 6. clean up — exact id, own list only, through the fail-closed §0 helper
delete_session "$SID"
```
Increment `SEQ` for every *new* input to the same worker. Reuse the same `SEQ` only to
@@ -104,8 +125,10 @@ live), so send-and-wait can burn its whole timeout. The reliable pattern is a sp
unique marker plus `wait-output from=buffer`:
```bash
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"builder","mode":"shell"}' | jq -r '.data.sessionId')
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"builder","mode":"shell"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed"; exit 1; }
CREATED+=("$SID")
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
@@ -115,7 +138,7 @@ done
# An unsplit marker matches the echo of your own keystrokes before the build runs.
N="${RANDOM}_$$"; MARK="DONE_$N"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"build-'$$'","seq":1}'
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"codeman-build-1","seq":1}'
for TRY in $(seq 1 30); do # BOUNDED (30 min): a \r-less send makes an uncapped loop infinite
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
@@ -136,8 +159,10 @@ waiter cap is 16 and abandoned concurrent waits pile up against it.
```bash
declare -A WORKER MARKS
for task in lint typecheck unit; do
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"fan-'"$task"'","mode":"shell"}' | jq -r '.data.sessionId')
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"fan-'"$task"'","mode":"shell"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; echo "$task: spawn failed"; continue; }
WORKER[$task]=$SID; CREATED+=("$SID")
done
for task in "${!WORKER[@]}"; do
@@ -147,7 +172,7 @@ for task in "${!WORKER[@]}"; do
done
N="${task}_${RANDOM}"; MARKS[$task]="DONE_$N"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run '"$task"'; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"fan-'$$'","seq":1}'
-d '{"input":"M=DONE; npm run '"$task"'; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"codeman-fan-'"$task"'","seq":1}'
done
for task in "${!WORKER[@]}"; do # sequential gather; each wait blocks until that worker is done
for TRY in $(seq 1 30); do # BOUNDED per worker, same reasoning as Flow 2
@@ -171,8 +196,8 @@ other was still running):
```bash
sendwait() { # $1=sid $2=prompt $3=seq — assumes the worker passed Flow 1's readiness
local body; body=$(jq -n --arg p "$2" --argjson s "$3" \
'{input:($p+"\r"),useMux:true,clientId:"fan-'$$'",seq:$s,wait:true,waitTimeout:600000}')
local body; body=$(jq -n --arg p "$2" --argjson s "$3" --arg c "codeman-fan-$1" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:600000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$1/input" \
-H 'Content-Type: application/json' --data-binary "$body" > "/tmp/fan-$1.json"
}
@@ -200,7 +225,7 @@ declare -A TOK
for i in 1 2; do
TOK[$i]="${RANDOM}_$i"
BODY=$(jq -n --arg p "do task $i; when completely done print the word WORKDONE immediately followed by _${TOK[$i]}" \
--arg c "fan-$$" --argjson s 2 '{input:($p+"\r"),useMux:true,clientId:$c,seq:$s}')
--arg c "codeman-fan-$i" --argjson s 2 '{input:($p+"\r"),useMux:true,clientId:$c,seq:$s}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/${SIDS[$i]}/input" \
-H 'Content-Type: application/json' --data-binary "$BODY"
done
@@ -237,12 +262,17 @@ At the end of the conversation (or on abort), delete exactly what you created:
```bash
for id in "${CREATED[@]}"; do
is_self "$id" || "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$id"
delete_session "$id"
done
```
- Only ids from your own `CREATED` list. Never enumerate `/api/v1/sessions` and
delete by pattern; other sessions belong to the user.
- Always go through `delete_session`. It refuses an empty id, refuses when `$SELF` is
unset or too short to prove the target is not you, and prefix-checks in both
directions. A hand-written `curl -X DELETE`, or the old
`is_self "$id" || curl -X DELETE …`, has none of that: an undefined `is_self` exits
127 and the `||` branch deletes unguarded.
- If you created a *case* purely as scratch and the user confirmed it is disposable,
`DELETE /api/v1/cases/:name` removes it — but that recursively deletes the
directory from disk, so never do it without the user's explicit go-ahead for that
+31 -3
View File
@@ -135,6 +135,7 @@ export abstract class AiCheckerBase<
// Active check state
protected checkMuxName: string | null = null;
protected checkTempFile: string | null = null;
protected checkStderrFile: string | null = null;
protected checkPromptFile: string | null = null;
protected checkPollTimer: NodeJS.Timeout | null = null;
protected checkTimeoutTimer: NodeJS.Timeout | null = null;
@@ -376,6 +377,7 @@ export abstract class AiCheckerBase<
const shortId = this.sessionId.slice(0, 8);
const timestamp = Date.now();
this.checkTempFile = join(tmpdir(), `${this.tempFilePrefix}-${shortId}-${timestamp}.txt`);
this.checkStderrFile = join(tmpdir(), `${this.tempFilePrefix}-stderr-${shortId}-${timestamp}.txt`);
this.checkPromptFile = join(tmpdir(), `${this.tempFilePrefix}-prompt-${shortId}-${timestamp}.txt`);
this.checkMuxName = `${this.muxNamePrefix}${shortId}`;
@@ -386,6 +388,7 @@ export abstract class AiCheckerBase<
// Ensure output temp file exists (empty) so we can poll it
writeFileSync(this.checkTempFile, '');
writeFileSync(this.checkStderrFile, '');
// Write prompt to file to avoid E2BIG error (argument list too long)
// The prompt can be 16KB+ which exceeds shell argument limits
@@ -396,7 +399,7 @@ export abstract class AiCheckerBase<
const modelArg = `--model "${this.config.model.replace(/"/g, '\\"')}"`;
const augmentedPath = getAugmentedPath();
const claudeCmd = `cat "${this.checkPromptFile}" | claude -p ${modelArg} --output-format text`;
const fullCmd = `export PATH="${augmentedPath}"; ${claudeCmd} > "${this.checkTempFile}" 2>&1; echo "${this.doneMarker}" >> "${this.checkTempFile}"; rm -f "${this.checkPromptFile}"`;
const fullCmd = `export PATH="${augmentedPath}"; ${claudeCmd} > "${this.checkTempFile}" 2> "${this.checkStderrFile}"; echo "${this.doneMarker}" >> "${this.checkTempFile}"; rm -f "${this.checkPromptFile}"`;
// Spawn tmux session
try {
@@ -461,18 +464,32 @@ export abstract class AiCheckerBase<
const output = content.replace(this.doneMarker, '').trim();
if (!output) {
return this.createErrorResult(`Empty output from ${this.checkDescription}`, durationMs);
const stderr = this.readStderrDiagnostic();
const detail = stderr ? `: ${stderr}` : '';
return this.createErrorResult(`Empty output from ${this.checkDescription}${detail}`, durationMs);
}
// Delegate to subclass for verdict parsing
const parsed = this.parseVerdict(output);
if (!parsed) {
return this.createErrorResult(`Could not parse verdict from: "${output.substring(0, 100)}"`, durationMs);
const stderr = this.readStderrDiagnostic();
const detail = stderr ? `; stderr: "${stderr}"` : '';
return this.createErrorResult(`Could not parse verdict from: "${output.substring(0, 100)}"${detail}`, durationMs);
}
return this.createResult(parsed.verdict, parsed.reasoning, durationMs);
}
private readStderrDiagnostic(): string {
if (!this.checkStderrFile || !existsSync(this.checkStderrFile)) return '';
try {
return readFileSync(this.checkStderrFile, 'utf-8').trim().substring(0, 200);
} catch {
return '';
}
}
private cleanupCheck(): void {
// Clear poll timer
if (this.checkPollTimer) {
@@ -509,6 +526,17 @@ export abstract class AiCheckerBase<
this.checkTempFile = null;
}
if (this.checkStderrFile) {
try {
if (existsSync(this.checkStderrFile)) {
unlinkSync(this.checkStderrFile);
}
} catch {
// Best effort cleanup
}
this.checkStderrFile = null;
}
if (this.checkPromptFile) {
try {
if (existsSync(this.checkPromptFile)) {
+322 -51
View File
@@ -12,15 +12,20 @@ import chalk from 'chalk';
import { createRequire } from 'module';
import http from 'node:http';
import https from 'node:https';
import { readFileSync } from 'node:fs';
import { isAbsolute } from 'node:path';
import { existsSync, readFileSync } from 'node:fs';
import { isAbsolute, join } from 'node:path';
import { homedir } from 'node:os';
import { dataPath } from './config/instance.js';
import { installAgentSkillInto, removeAgentSkillFrom, type AgentSkillApplyResult } from './hooks-config.js';
import { getSessionManager } from './session-manager.js';
import { getTaskQueue } from './task-queue.js';
import { getRalphLoop } from './ralph-loop.js';
import { getStore } from './state-store.js';
import { getErrorMessage } from './types.js';
import { isSupportedAttachmentExtension } from './attachment-registry.js';
import { daemonStatus, startDaemon, stopDaemon, type WebLaunchOptions } from './daemon-control.js';
import { installService, serviceStatus, uninstallService } from './service-installer.js';
import { isLoopbackBindHost, isUnauthenticatedNetworkAcknowledged } from './web/network-auth-policy.js';
const require = createRequire(import.meta.url);
const pkg = require('../package.json') as { version: string };
@@ -116,6 +121,112 @@ program
console.log(makeAttachmentMagicLink(filePath));
});
// ============ Skill Commands ============
/** Same registry the server resolves case names through (mirrors `case-routes.ts`). */
const LINKED_CASES_FILE = dataPath('linked-cases.json');
/**
* Case name to directory, checking `linked-cases.json` FIRST and falling back to the
* shared single-user cases dir. Mirrors `resolveCasePath()` in `case-routes.ts`, which
* is what the web UI and `quick-start` use. Without the linked-cases lookup this
* command rejected every case linked in from outside `~/codeman-cases` with
* "Case not found", even though the server resolved the same name fine.
*
* Sync and tolerant on purpose: a missing or malformed registry means "no linked
* cases", never a crash.
*/
function resolveCliCasePath(name: string): string {
try {
const linked = JSON.parse(readFileSync(LINKED_CASES_FILE, 'utf-8')) as Record<string, string>;
const target = linked?.[name];
if (typeof target === 'string' && target) return target;
} catch {
// no registry yet, or unreadable/invalid JSON: fall through to the cases dir
}
return join(homedir(), 'codeman-cases', name);
}
/**
* Resolve where `skill install` / `skill uninstall` operate. Global is
* `~/.claude/skills/codeman` (Claude Code's user-scope skill dir, read by every new
* session); `--case <name>` targets `<case>/.claude/skills/codeman`, resolved through
* `resolveCliCasePath()` above. The web server's automatic per-case injection
* (`agentSkillEnabled`) covers multi-user spaces; this CLI is a local operator tool
* and stays single-user.
*/
function resolveSkillTarget(options: { case?: string }): string {
if (options.case) {
const casePath = resolveCliCasePath(options.case);
if (!existsSync(casePath)) {
console.error(chalk.red(`✗ Case not found: ${casePath}`));
process.exit(1);
}
return join(casePath, '.claude', 'skills', 'codeman');
}
return join(homedir(), '.claude', 'skills', 'codeman');
}
/** Print an AgentSkillApplyResult for humans; exit non-zero when nothing was done. */
function reportSkillResult(result: AgentSkillApplyResult, target: string): void {
const messages: Record<AgentSkillApplyResult, { ok: boolean; text: string }> = {
installed: { ok: true, text: `Agent skill installed: ${target}` },
refreshed: { ok: true, text: `Agent skill refreshed (was stale): ${target}` },
unchanged: { ok: true, text: `Agent skill already up to date: ${target}` },
removed: { ok: true, text: `Agent skill removed: ${target}` },
absent: { ok: true, text: `Nothing to remove at ${target}` },
foreign: {
ok: false,
text: `${target} exists but is not Codeman-managed (no marker), refusing to touch it. Remove it yourself if you want the packaged skill there.`,
},
symlink: {
ok: false,
text: `${target} (or its parent) is a symlink, refusing to write through it.`,
},
};
const message = messages[result];
if (message.ok) {
console.log(chalk.green(`✓ ${message.text}`));
} else {
console.error(chalk.red(`✗ ${message.text}`));
process.exit(1);
}
}
const skillCmd = program
.command('skill')
.description('Manage the Codeman agent skill (lets an agent inside a session drive the API)');
skillCmd
.command('install')
.description('Install the agent skill globally (~/.claude/skills/codeman) or into one case')
.option('-g, --global', 'Install into ~/.claude/skills/codeman, picked up by every new session (the default)')
.option('-c, --case <name>', 'Install into <case>/.claude/skills/codeman instead (linked cases resolve too)')
.action(async (options: { global?: boolean; case?: string }) => {
try {
const target = resolveSkillTarget(options);
reportSkillResult(await installAgentSkillInto(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to install agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
skillCmd
.command('uninstall')
.description('Remove a Codeman-managed agent skill copy (never touches a user-authored one)')
.option('-g, --global', 'Remove from ~/.claude/skills/codeman (the default)')
.option('-c, --case <name>', 'Remove from <case>/.claude/skills/codeman instead (linked cases resolve too)')
.action(async (options: { global?: boolean; case?: string }) => {
try {
const target = resolveSkillTarget(options);
reportSkillResult(await removeAgentSkillFrom(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to remove agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
// ============ Session Commands ============
const sessionCmd = program.command('session').alias('s').description('Manage Claude sessions');
@@ -572,64 +683,224 @@ program
console.log('');
});
// ============ Web / daemon / service Commands ============
/** Shared option set for the commands that can launch a web server. */
function addWebLaunchOptions(cmd: Command): Command {
return cmd
.option('-H, --host <host>', 'Host to bind to', process.env.CODEMAN_HOST || '127.0.0.1')
.option('-p, --port <port>', 'Port to listen on (env: CODEMAN_PORT)', process.env.CODEMAN_PORT || '3000')
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
.option(
'--allow-unauthenticated-network',
'Allow non-loopback web access without CODEMAN_PASSWORD (dangerous; terminal control is exposed)'
)
.option(
'--multiuser',
'Enable opt-in multi-user mode (named users in ~/.codeman/users.json; env: CODEMAN_MULTIUSER)'
);
}
/** Normalize commander's strings into the shape daemon-control/service-installer take. */
function toWebLaunchOptions(options: {
host: string;
port: string;
https?: boolean;
titleHostname?: string;
allowUnauthenticatedNetwork?: boolean;
multiuser?: boolean;
}): WebLaunchOptions {
const port = parseInt(options.port, 10);
if (!Number.isInteger(port) || port <= 0 || port > 65535) {
console.error(chalk.red(`✗ Invalid port: ${options.port}`));
process.exit(1);
}
return {
host: options.host,
port,
https: !!options.https,
titleHostname: options.titleHostname,
allowUnauthenticatedNetwork: !!options.allowUnauthenticatedNetwork,
multiuser: !!options.multiuser,
};
}
/**
* The server prints this itself, but into a log file nobody reads when it is
* detached or supervised. Repeat it where the operator is actually looking.
*/
function warnIfUnauthenticatedNetwork(launch: WebLaunchOptions): void {
if (isLoopbackBindHost(launch.host)) return;
if (isUnauthenticatedNetworkAcknowledged(launch.allowUnauthenticatedNetwork)) return;
console.log(
chalk.yellow(
`⚠ Binding ${launch.host} without CODEMAN_PASSWORD: anyone who can reach this port gets terminal control.`
)
);
console.log(chalk.yellow(' Set CODEMAN_PASSWORD, or bind 127.0.0.1 and front it with tailscale serve.'));
}
// Web interface command
program
.command('web')
.description('Start the web interface')
.option('-H, --host <host>', 'Host to bind to', process.env.CODEMAN_HOST || '127.0.0.1')
.option('-p, --port <port>', 'Port to listen on (env: CODEMAN_PORT)', process.env.CODEMAN_PORT || '3000')
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
.option(
'--allow-unauthenticated-network',
'Allow non-loopback web access without CODEMAN_PASSWORD (dangerous; terminal control is exposed)'
)
.option('--multiuser', 'Enable opt-in multi-user mode (named users in ~/.codeman/users.json; env: CODEMAN_MULTIUSER)')
.action(async (options) => {
// The flag is surfaced to the rest of the process via the env var so
// isMultiUserMode() has a single source of truth (see config/multiuser.ts).
if (options.multiuser) process.env.CODEMAN_MULTIUSER = '1';
const { startWebServer } = await import('./web/server.js');
const host = options.host;
const port = parseInt(options.port, 10);
const https = !!options.https;
const titleHostname = options.titleHostname;
const allowUnauthenticatedNetwork = !!options.allowUnauthenticatedNetwork;
const protocol = https ? 'https' : 'http';
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
const webCmd = addWebLaunchOptions(program.command('web').description('Start the web interface'))
.option('-d, --daemon', 'Run detached in the background; survives the shell, logs to <data dir>/web.log')
.option('--stop', 'Stop a server started with --daemon')
.option('--status', 'Report whether a detached server is running');
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
webCmd.action(async (options) => {
// The flag is surfaced to the rest of the process via the env var so
// isMultiUserMode() has a single source of truth (see config/multiuser.ts).
if (options.multiuser) process.env.CODEMAN_MULTIUSER = '1';
const launch = toWebLaunchOptions(options);
try {
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
if (https) {
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
if (options.stop) {
const result = await stopDaemon(launch);
if (result.ok && result.reason === 'not-running') {
console.log(chalk.gray(`○ ${result.message}`));
return;
}
if (result.ok) {
console.log(chalk.green(`✓ ${result.message ?? `Stopped Codeman (pid ${result.pid})`}`));
console.log(chalk.gray(' Your agents keep running in tmux.'));
return;
}
console.error(chalk.red(`✗ ${result.message ?? 'Could not stop the server'}`));
process.exit(1);
}
if (options.status) {
const status = await daemonStatus(launch);
if (status.responding) {
const version = status.version ? ` (v${status.version})` : '';
console.log(chalk.green(`✓ Responding at ${status.url}${version}`));
} else {
console.log(chalk.yellow(`○ Nothing answering at ${status.url}`));
}
console.log(` Daemon pid: ${status.running ? chalk.green(String(status.pid)) : chalk.gray('not running')}`);
console.log(chalk.gray(` Pidfile: ${status.pidFile}`));
console.log(chalk.gray(` Log: ${status.logPath}`));
if (!status.running && status.responding) {
console.log(chalk.gray(' (running, but not started with --daemon: probably a service or a foreground run)'));
}
return;
}
if (options.daemon) {
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Starting Codeman in the background...'));
const result = await startDaemon(launch);
if (result.ok) {
console.log(chalk.green(`\n✓ Codeman is running at ${result.url} (pid ${result.pid})`));
console.log(chalk.gray(` Logs: ${result.logPath}`));
console.log(chalk.gray(' Stop it with: codeman web --stop'));
console.log(chalk.gray(' Want it back after a reboot? codeman service install'));
return;
}
console.error(chalk.red(`\n✗ ${result.message ?? 'Failed to start'}`));
process.exit(1);
}
const { startWebServer } = await import('./web/server.js');
const host = launch.host;
const port = launch.port;
const https = launch.https;
const titleHostname = options.titleHostname;
const allowUnauthenticatedNetwork = launch.allowUnauthenticatedNetwork ?? false;
const protocol = https ? 'https' : 'http';
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
try {
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
if (https) {
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
let shuttingDown = false;
const shutdown = async (signal: string) => {
if (shuttingDown) return;
shuttingDown = true;
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
try {
await server.stop();
} catch (err) {
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
process.exit(0);
};
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
process.on('SIGHUP', () => shutdown('SIGHUP'));
} catch (err) {
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
process.exit(1);
}
});
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
let shuttingDown = false;
const shutdown = async (signal: string) => {
if (shuttingDown) return;
shuttingDown = true;
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
try {
await server.stop();
} catch (err) {
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
}
process.exit(0);
};
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
process.on('SIGHUP', () => shutdown('SIGHUP'));
} catch (err) {
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
// Supervised service: the "still there after a reboot" answer, where `web -d` is
// the "still there after I close this shell" one (issue #231).
const serviceCmd = program
.command('service')
.description('Manage the background service (systemd user unit on Linux, LaunchAgent on macOS)');
addWebLaunchOptions(
serviceCmd.command('install').description('Install and start the service, then verify it answers')
).action(async (options) => {
const launch = toWebLaunchOptions(options);
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Installing the Codeman service...'));
const result = await installService(launch);
for (const warning of result.warnings ?? []) console.log(chalk.yellow(`⚠ ${warning}`));
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
console.log(chalk.gray(` Unit: ${result.unitPath}`));
if (process.env.CODEMAN_PASSWORD) {
console.log(
chalk.yellow(
' Note: CODEMAN_PASSWORD was NOT copied into the unit file. Add it there yourself if the service needs auth.'
)
);
}
});
serviceCmd
.command('uninstall')
.description('Stop the service and remove its unit file')
.action(() => {
const result = uninstallService();
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
});
addWebLaunchOptions(
serviceCmd.command('status').description('Show whether the service is installed and running')
).action(async (options) => {
const status = await serviceStatus(toWebLaunchOptions(options));
if (!status.kind) {
console.log(chalk.yellow(`No supported supervisor on ${process.platform}. Use \`codeman web -d\` instead.`));
return;
}
console.log(` Supervisor: ${status.kind} (${status.name})`);
console.log(` Unit file: ${status.installed ? chalk.green(status.unitPath) : chalk.gray('not installed')}`);
console.log(` Loaded: ${status.loaded ? chalk.green('yes') : chalk.gray('no')}`);
const version = status.version ? ` (v${status.version})` : '';
console.log(
` Responding: ${status.responding ? chalk.green(`yes at ${status.url}${version}`) : chalk.gray(`no at ${status.url}`)}`
);
});
// ============ Multi-user Commands ============
//
// Operate directly on ~/.codeman/users.json (via user-store) with NO running
+32
View File
@@ -0,0 +1,32 @@
/**
* @fileoverview Supervisor identity (systemd unit name / launchd job label).
*
* Three things now write or look for the same supervisor job: `install.sh`, the
* in-app self-updater (`web/self-update.ts` detects it to decide how to restart),
* and `codeman service install`. The names live here so they cannot drift apart,
* because a mismatch is silent in the worst way: `service install` would happily
* create a SECOND job alongside the installer's, and two servers sharing one data
* dir and one tmux socket attach PTYs to each other's live sessions
* (see config/instance.ts).
*
* The names are instance-scoped for exactly that reason: a `CODEMAN_INSTANCE=beta`
* build writing `com.codeman.web` would overwrite the production LaunchAgent. The
* DEFAULT instance keeps the historical names byte-identical, so existing installs
* and every unit install.sh has already written are unaffected.
*
* @module config/service-names
*/
import { CODEMAN_INSTANCE } from './instance.js';
/**
* Instance name reduced to characters that are safe in a filename and in a
* launchd label. `CODEMAN_INSTANCE` is arbitrary operator input.
*/
const SAFE_INSTANCE = CODEMAN_INSTANCE.replace(/[^A-Za-z0-9_-]/g, '').slice(0, 32);
/** systemd user unit: `codeman-web.service`, or `codeman-web-beta.service` for a beta. */
export const SYSTEMD_UNIT = `codeman-web${SAFE_INSTANCE ? `-${SAFE_INSTANCE}` : ''}.service`;
/** launchd job label: `com.codeman.web`, or `com.codeman.beta.web` for a beta. */
export const LAUNCHD_LABEL = SAFE_INSTANCE ? `com.codeman.${SAFE_INSTANCE}.web` : 'com.codeman.web';
+496
View File
@@ -0,0 +1,496 @@
/**
* @fileoverview Detached `codeman web` control: start (-d), stop, status.
*
* Backs `codeman web -d`, `codeman web --stop` and `codeman web --status`. The
* server itself is unchanged; this module re-launches the SAME entry script in a
* new session (`detached: true` calls setsid), so the child has no controlling
* terminal and no shell job entry. That is what actually makes it outlive the
* shell: `nohup` does not, because Node re-arms SIGHUP to its default disposition
* even when it inherits "ignore", and `cli.ts` installs a SIGHUP handler that
* shuts the server down gracefully (issue #231).
*
* Two rules shape the rest of the module:
*
* 1. **Never start a second server on one data dir.** `~/.codeman` and the
* `tmux -L codeman` socket are process-wide (config/instance.ts), so a second
* instance discovers and attaches PTYs to the first one's live sessions and
* starts resizing them. A double `-d` therefore has to be a hard error, which
* means checking both the pidfile AND the port before spawning.
* 2. **Never report success we have not seen.** The parent polls `/api/status`
* until the child answers (or dies) before printing a URL. A port clash or a
* missing dependency otherwise looks exactly like a clean start.
*
* Pure helpers (arg building, URL building, pidfile parsing, the process-identity
* check) are exported separately so they can be unit-tested without spawning.
*
* @module daemon-control
*/
import { spawn, execFileSync } from 'node:child_process';
import { appendFileSync, closeSync, existsSync, openSync, readFileSync, unlinkSync, writeFileSync } from 'node:fs';
import http from 'node:http';
import https from 'node:https';
import { dataPath } from './config/instance.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
/** How long to wait for a freshly spawned server to answer `/api/status`. */
const START_TIMEOUT_MS = 30_000;
/** How long to wait for a SIGTERM'd server to actually exit before giving up. */
const STOP_TIMEOUT_MS = 15_000;
/** Poll interval while waiting for either of the above. */
const POLL_INTERVAL_MS = 250;
/** The `web` command's options, as far as a detached relaunch cares about them. */
export interface WebLaunchOptions {
host: string;
port: number;
https: boolean;
titleHostname?: string;
allowUnauthenticatedNetwork?: boolean;
multiuser?: boolean;
}
export interface StartResult {
ok: boolean;
pid?: number;
url?: string;
/** Machine-readable failure cause; `undefined` on success. */
reason?: 'already-running' | 'exited' | 'timeout';
message?: string;
logPath: string;
}
export interface StopResult {
ok: boolean;
pid?: number;
reason?: 'not-running' | 'foreign-pid' | 'timeout' | 'no-pidfile-but-responding';
message?: string;
}
export interface DaemonStatus {
pid: number | null;
/** The pid in the pidfile is alive AND still looks like a Codeman web process. */
running: boolean;
/** Something answered `/api/status` at the expected address. */
responding: boolean;
version?: string;
url: string;
pidFile: string;
logPath: string;
}
// ─────────────────────────────────────────────────────────────────────────────
// Pure helpers
// ─────────────────────────────────────────────────────────────────────────────
/** Rebuild the `web` argv for the child, dropping the daemon flags themselves. */
export function buildWebArgs(options: WebLaunchOptions): string[] {
const args = ['web', '--host', options.host, '--port', String(options.port)];
if (options.https) args.push('--https');
if (options.titleHostname) args.push('--title-hostname', options.titleHostname);
if (options.allowUnauthenticatedNetwork) args.push('--allow-unauthenticated-network');
if (options.multiuser) args.push('--multiuser');
return args;
}
/**
* Connectable address for this bind. A wildcard bind is not itself connectable,
* so `0.0.0.0` / `::` become loopback; a bare IPv6 literal gets bracketed.
*/
export function buildBaseUrl(options: WebLaunchOptions): string {
const protocol = options.https ? 'https' : 'http';
let host = options.host.trim();
if (host === '0.0.0.0' || host === '::' || host === '') host = '127.0.0.1';
if (host.includes(':') && !host.startsWith('[')) host = `[${host}]`;
return `${protocol}://${host}:${options.port}`;
}
/** The endpoint polled for readiness. */
export function buildStatusUrl(options: WebLaunchOptions): string {
return `${buildBaseUrl(options)}/api/status`;
}
/** Parse a pidfile body. Rejects garbage, and pid 1 (init is never ours). */
export function parsePidFileContents(text: string): number | null {
const trimmed = text.trim();
if (!/^\d+$/.test(trimmed)) return null;
const pid = Number.parseInt(trimmed, 10);
if (!Number.isSafeInteger(pid) || pid <= 1) return null;
return pid;
}
/**
* Does this command line look like a Codeman web server?
*
* Pids are recycled, and a stale pidfile pointing at whatever inherited the
* number is a live footgun: `codeman web --stop` must not SIGTERM an unrelated
* process. Both the npm bin (`codeman`/`aicodeman`) and the direct entry
* (`node dist/index.js web`, `tsx src/index.ts web`) have to match.
*/
export function looksLikeCodemanWeb(command: string | null | undefined): boolean {
if (!command) return false;
if (!/(^|\s)web(\s|$)/.test(command)) return false;
return /(^|[/\s])(ai)?codeman(\s|$)/.test(command) || /index\.(js|ts)(\s|$)/.test(command);
}
// ─────────────────────────────────────────────────────────────────────────────
// Paths
// ─────────────────────────────────────────────────────────────────────────────
/**
* Resolved at call time, not module load: tests swap `HOME` per file, and the
* data dir is derived from it (see test/setup.ts).
*/
export function pidFilePath(): string {
return dataPath('web.pid');
}
/** Where a detached server's stdout/stderr is appended. */
export function logFilePath(): string {
return dataPath('web.log');
}
// ─────────────────────────────────────────────────────────────────────────────
// Process probing
// ─────────────────────────────────────────────────────────────────────────────
/** Signal 0 liveness check. EPERM means the pid exists but is not ours. */
export function isProcessAlive(pid: number): boolean {
try {
process.kill(pid, 0);
return true;
} catch (err) {
return (err as NodeJS.ErrnoException).code === 'EPERM';
}
}
/** Full command line of a pid, or null. `-o command=` is portable to macOS. */
export function readProcessCommand(pid: number): string | null {
try {
const out = execFileSync('ps', ['-o', 'command=', '-p', String(pid)], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'ignore'],
});
return out.trim() || null;
} catch {
return null;
}
}
/** Read the pidfile, returning null when it is missing, empty or malformed. */
export function readPidFile(): number | null {
const file = pidFilePath();
if (!existsSync(file)) return null;
try {
return parsePidFileContents(readFileSync(file, 'utf-8'));
} catch {
return null;
}
}
function removePidFile(): void {
try {
unlinkSync(pidFilePath());
} catch {
/* already gone */
}
}
/** Pid of a live Codeman web server recorded in the pidfile, or null. */
export function readLivePid(): number | null {
const pid = readPidFile();
if (pid === null) return null;
if (!isProcessAlive(pid)) return null;
// A recycled pid is not ours. `ps` can also legitimately fail (containers with
// no procps); treat "cannot tell" as ours rather than orphaning the pidfile.
const command = readProcessCommand(pid);
if (command !== null && !looksLikeCodemanWeb(command)) return null;
return pid;
}
// ─────────────────────────────────────────────────────────────────────────────
// HTTP readiness probe
// ─────────────────────────────────────────────────────────────────────────────
export interface ProbeResult {
/** A Codeman server answered. A 401 counts: auth is active, the server is up. */
up: boolean;
version?: string;
}
/**
* Probe `/api/status`. Self-signed certs are accepted (`--https` generates one),
* and 401 counts as up because `CODEMAN_PASSWORD` gates that route. The body is
* checked so an unrelated service squatting on the port is not read as success.
*/
export function probeServer(url: string, timeoutMs = 2000): Promise<ProbeResult> {
return new Promise((resolve) => {
let settled = false;
const done = (result: ProbeResult) => {
if (settled) return;
settled = true;
resolve(result);
};
let target: URL;
try {
target = new URL(url);
} catch {
done({ up: false });
return;
}
const transport = target.protocol === 'https:' ? https : http;
const req = transport.request(
{
protocol: target.protocol,
hostname: target.hostname,
port: target.port,
path: target.pathname,
method: 'GET',
rejectUnauthorized: false,
timeout: timeoutMs,
headers: { Accept: 'application/json' },
},
(res) => {
if (res.statusCode === 401) {
res.resume();
done({ up: true });
return;
}
let body = '';
res.setEncoding('utf-8');
res.on('data', (chunk: string) => {
if (body.length < 4096) body += chunk;
});
res.on('end', () => {
if (!body.includes('"success"')) {
done({ up: false });
return;
}
let version: string | undefined;
try {
version = (JSON.parse(body) as { data?: { version?: string } }).data?.version;
} catch {
/* body was truncated at 4KB; up is still true */
}
done({ up: true, version });
});
res.on('error', () => done({ up: false }));
}
);
req.on('timeout', () => {
req.destroy();
done({ up: false });
});
req.on('error', () => done({ up: false }));
req.end();
});
}
function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
// ─────────────────────────────────────────────────────────────────────────────
// Start / stop / status
// ─────────────────────────────────────────────────────────────────────────────
/**
* The script to relaunch. `process.execArgv` is carried over with it so a dev
* run under tsx (whose execArgv holds the tsx loader flags) re-launches through
* tsx instead of handing a `.ts` file to bare node.
*/
function entryScript(): string {
const script = process.argv[1];
if (!script) throw new Error('cannot determine the codeman entry script to relaunch');
return script;
}
/** Marks one launch in the append-only log so a tail cannot mix two runs. */
const LOG_SEPARATOR = '=== codeman web start';
/**
* Last few lines of the daemon log, for reporting a failed start. The log is
* append-only across launches, so the tail starts at the last separator when
* there is one: otherwise a crash report is padded with the previous run's
* cheerful startup banner.
*/
export function tailLog(maxLines = 15): string {
try {
const lines = readFileSync(logFilePath(), 'utf-8').trimEnd().split('\n');
const start = lines.map((line) => line.startsWith(LOG_SEPARATOR)).lastIndexOf(true);
const current = start === -1 ? lines : lines.slice(start + 1);
return current.slice(-maxLines).join('\n');
} catch {
return '';
}
}
/**
* Spawn a detached `codeman web` and wait until it answers before returning.
* Refuses when a server is already up on this data dir (see rule 1 in the module
* docblock).
*/
export async function startDaemon(options: WebLaunchOptions): Promise<StartResult> {
const logPath = logFilePath();
const url = buildBaseUrl(options);
const statusUrl = buildStatusUrl(options);
const existingPid = readLivePid();
if (existingPid !== null) {
return {
ok: false,
reason: 'already-running',
pid: existingPid,
logPath,
message: `a Codeman server is already running (pid ${existingPid}). Stop it with \`codeman web --stop\` first.`,
};
}
const alreadyServing = await probeServer(statusUrl, 1500);
if (alreadyServing.up) {
return {
ok: false,
reason: 'already-running',
logPath,
url,
message: `something is already serving ${url}. Two servers on one data dir attach to each other's tmux sessions, so refusing to start.`,
};
}
// A pidfile that survived a crash: the process is gone, so it is just litter.
if (readPidFile() !== null) removePidFile();
const args = buildWebArgs(options);
try {
appendFileSync(logPath, `\n${LOG_SEPARATOR} ${new Date().toISOString()} ===\n`, 'utf-8');
} catch {
/* the spawn below reports a genuinely unwritable log */
}
const logFd = openSync(logPath, 'a');
let child;
try {
child = spawn(process.execPath, [...process.execArgv, entryScript(), ...args], {
detached: true,
stdio: ['ignore', logFd, logFd],
env: process.env,
});
} finally {
closeSync(logFd);
}
let exited = false;
child.on('exit', () => {
exited = true;
});
child.on('error', () => {
exited = true;
});
const pid = child.pid;
if (pid === undefined) {
return { ok: false, reason: 'exited', logPath, message: 'failed to spawn the server process' };
}
writeFileSync(pidFilePath(), `${pid}\n`, 'utf-8');
const deadline = Date.now() + START_TIMEOUT_MS;
while (Date.now() < deadline) {
if (exited) {
removePidFile();
child.unref();
return {
ok: false,
reason: 'exited',
logPath,
message: `the server exited during startup. Last lines of ${logPath}:\n${tailLog()}`,
};
}
const probe = await probeServer(statusUrl, 1000);
if (probe.up) {
child.unref();
return { ok: true, pid, url, logPath };
}
await sleep(POLL_INTERVAL_MS);
}
child.unref();
return {
ok: false,
reason: 'timeout',
pid,
url,
logPath,
message: `the server did not answer ${url} within ${START_TIMEOUT_MS / 1000}s. It may still be starting; check ${logPath}.`,
};
}
/** SIGTERM the recorded server and wait for it to actually exit. */
export async function stopDaemon(options: WebLaunchOptions): Promise<StopResult> {
const pid = readPidFile();
if (pid === null) {
const probe = await probeServer(buildStatusUrl(options), 1500);
if (probe.up) {
return {
ok: false,
reason: 'no-pidfile-but-responding',
message:
'a server is responding but there is no pidfile, so it was not started with `-d`. If it is a service use `codeman service uninstall` (or stop the unit); otherwise `pkill -f "index.js web"`.',
};
}
return { ok: true, reason: 'not-running', message: 'no daemon is running; nothing to stop' };
}
if (!isProcessAlive(pid)) {
removePidFile();
return { ok: true, pid, message: `stale pidfile removed (pid ${pid} was not running)` };
}
const command = readProcessCommand(pid);
if (command !== null && !looksLikeCodemanWeb(command)) {
return {
ok: false,
reason: 'foreign-pid',
pid,
message: `pid ${pid} is not a Codeman server (${command}). Refusing to signal it; delete ${pidFilePath()} if it is stale.`,
};
}
// SIGTERM, never SIGKILL: cli.ts flushes state on the way out.
try {
process.kill(pid, 'SIGTERM');
} catch (err) {
return { ok: false, reason: 'foreign-pid', pid, message: `could not signal pid ${pid}: ${String(err)}` };
}
const deadline = Date.now() + STOP_TIMEOUT_MS;
while (Date.now() < deadline) {
if (!isProcessAlive(pid)) {
removePidFile();
return { ok: true, pid };
}
await sleep(POLL_INTERVAL_MS);
}
return {
ok: false,
reason: 'timeout',
pid,
message: `pid ${pid} did not exit within ${STOP_TIMEOUT_MS / 1000}s. Force it with \`kill -9 ${pid}\` if you are sure.`,
};
}
/** Report on both halves: the recorded process, and whether the port answers. */
export async function daemonStatus(options: WebLaunchOptions): Promise<DaemonStatus> {
const url = buildBaseUrl(options);
const pid = readPidFile();
const probe = await probeServer(buildStatusUrl(options), 2000);
return {
pid,
running: readLivePid() !== null,
responding: probe.up,
version: probe.version,
url,
pidFile: pidFilePath(),
logPath: logFilePath(),
};
}
+379 -31
View File
@@ -10,13 +10,15 @@
* Key exports:
* - `generateHooksConfig()` — returns hooks object for settings.local.json
* - `writeHooksConfig(casePath)` — writes hooks + env config to disk
* - `ensureCodemanHooks(casePath)` — safely installs/updates hooks for a managed case
* - `updateCaseEnvVars(casePath, envVars)` — merges env vars into settings
*
* Hook events generated: `idle_prompt`, `permission_prompt`, `elicitation_dialog`,
* `stop`, `teammate_idle`, `task_completed`
*
* Hook categories: `Notification` (3 matchers), `Stop` (1), `TeammateIdle` (1),
* `TaskCompleted` (1), `PostToolUse` (1 self-contained background Bash rewake)
* Hook categories: `Notification` (3 matchers), `Stop` (1), `SubagentStop` (1),
* `TeammateIdle` (1), `TaskCompleted` (1), `PostToolUse` (1 self-contained
* background Bash rewake)
*
* @dependencies types (HookEventType), config/auth-config (HOOK_TIMEOUT_SECONDS)
* @consumedby web/server (session creation), session-cli-builder (env setup)
@@ -25,8 +27,9 @@
*/
import { existsSync } from 'node:fs';
import { readFile, writeFile, mkdir } from 'node:fs/promises';
import { join } from 'node:path';
import { readFile, writeFile, mkdir, lstat, readdir, unlink, rmdir } from 'node:fs/promises';
import { join, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
import type { HookEventType } from './types.js';
import { HOOK_TIMEOUT_SECONDS } from './config/auth-config.js';
@@ -52,15 +55,19 @@ const BACKGROUND_WAKE_MARKER_PREFIX = 'CODEMAN_BACKGROUND_REWAKE_V';
* changes: `refreshStaleCodemanHooks` treats the absence of the CURRENT marker as
* stale, so healed cases pick up the new script on next launch.
*/
const BACKGROUND_WAKE_MARKER = `${BACKGROUND_WAKE_MARKER_PREFIX}2`;
const BACKGROUND_WAKE_MARKER = `${BACKGROUND_WAKE_MARKER_PREFIX}3`;
const SUBAGENT_STOP_GUARD_MARKER_PREFIX = 'CODEMAN_SUBAGENT_STOP_GUARD_V';
const SUBAGENT_STOP_GUARD_MARKER = `${SUBAGENT_STOP_GUARD_MARKER_PREFIX}1`;
const BACKGROUND_WAKE_TIMEOUT_SECONDS = 6 * 60 * 60;
/**
* Inline Node helper for Claude Code's `asyncRewake` hook.
*
* A background Bash tool returns immediately with a task ID, then Claude writes
* its completion as a queue-operation in the transcript. Watching that durable
* record avoids injecting terminal input (which could submit a user's draft).
* its completion as a queue-operation in the top-level transcript. Subagent hooks
* receive their own transcript path even though their completion is parent-owned,
* so the helper watches both paths. Watching durable records avoids injecting
* terminal input (which could submit a user's draft).
* The helper is embedded in settings via `node -e`, so it has no script path
* that can go stale after an install or plugin-cache cleanup.
*
@@ -72,8 +79,12 @@ const BACKGROUND_WAKE_TIMEOUT_SECONDS = 6 * 60 * 60;
export function generateBackgroundWakeScript(): string {
return [
"const fs = require('node:fs');",
"const path = require('node:path');",
`const ${BACKGROUND_WAKE_MARKER} = true;`,
`const deadline = Date.now() + ${BACKGROUND_WAKE_TIMEOUT_SECONDS} * 1000;`,
"const RESULT_BEGIN = '=== CODEMAN_RESULT_BEGIN ===';",
"const RESULT_END = '=== CODEMAN_RESULT_END ===';",
'const MAX_RESULT_CHARS = 65536;',
'let input = {};',
"try { input = JSON.parse(fs.readFileSync(0, 'utf8') || '{}'); } catch { process.exit(0); }",
'function findTaskId(value) {',
@@ -98,46 +109,164 @@ export function generateBackgroundWakeScript(): string {
'const taskId = findTaskId(input.tool_response);',
"const transcriptPath = typeof input.transcript_path === 'string' ? input.transcript_path : '';",
'if (!taskId || !transcriptPath) process.exit(0);',
'let position = 0;',
'try { position = Math.max(0, fs.statSync(transcriptPath).size - 262144); } catch { process.exit(0); }',
"let carry = '';",
'const transcriptPaths = [transcriptPath];',
'const sessionDir = path.dirname(path.dirname(transcriptPath));',
"if (typeof input.agent_id === 'string' && path.basename(path.dirname(transcriptPath)) === 'subagents' &&",
" typeof input.session_id === 'string' && path.basename(sessionDir) === input.session_id) {",
" transcriptPaths.push(sessionDir + '.jsonl');",
'}',
'const transcripts = [...new Set(transcriptPaths)].map((transcript) => {',
' let position = 0;',
' try { position = Math.max(0, fs.statSync(transcript).size - 262144); } catch {}',
" return { path: transcript, position, carry: '' };",
'});',
'if (!transcripts.some((transcript) => fs.existsSync(transcript.path))) process.exit(0);',
'function readMarkedResult(outputPath) {',
" if (!outputPath || !path.isAbsolute(outputPath) || path.basename(outputPath) !== taskId + '.output') return '';",
" if (path.basename(path.dirname(outputPath)) !== 'tasks') return '';",
' try {',
' const size = fs.statSync(outputPath).size;',
' const length = Math.min(size, MAX_RESULT_CHARS * 2);',
' const buffer = Buffer.allocUnsafe(length);',
" const fd = fs.openSync(outputPath, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, size - length);',
' fs.closeSync(fd);',
" const text = buffer.subarray(0, bytes).toString('utf8');",
' const begin = text.lastIndexOf(RESULT_BEGIN);',
' const end = text.indexOf(RESULT_END, begin + RESULT_BEGIN.length);',
" if (begin < 0 || end < 0) return '';",
' let result = text.slice(begin + RESULT_BEGIN.length, end).trim();',
" if (!result) return '';",
' if (result.length > MAX_RESULT_CHARS) {',
' const half = Math.floor(MAX_RESULT_CHARS / 2);',
" result = result.slice(0, half) + '\\n\\n[report truncated by Codeman]\\n\\n' + result.slice(-half);",
' }',
" return '\\n\\nCompleted task report:\\n<codeman-background-result>\\n' + result + '\\n</codeman-background-result>';",
" } catch { return ''; }",
'}',
'function inspect(text) {',
' for (const line of text.split(/\\r?\\n/)) {',
' if (!line.includes(taskId)) continue;',
' let entry;',
' try { entry = JSON.parse(line); } catch { continue; }',
" if (entry.type !== 'queue-operation' || typeof entry.content !== 'string') continue;",
" if (entry.type !== 'queue-operation' || entry.operation !== 'enqueue' || typeof entry.content !== 'string') continue;",
" if (!entry.content.includes('<task-id>' + taskId + '</task-id>')) continue;",
' const status = entry.content.match(/<status>(completed|failed|killed|error)<\\/status>/i);',
' if (!status) continue;',
' const output = entry.content.match(/<output-file>([^<]+)<\\/output-file>/i);',
" const location = output ? ' Read ' + output[1] + ' and' : '';",
" console.error('Background command ' + taskId + ' ' + status[1].toLowerCase() + '.' + location + ' continue the task.');",
" const outputPath = output ? output[1].trim() : '';",
" const location = outputPath ? ' Read ' + outputPath + ' and' : '';",
' const result = readMarkedResult(outputPath);',
" console.error('Background command ' + taskId + ' ' + status[1].toLowerCase() + '.' + location + ' continue the task.' + result);",
' process.exit(2);',
' }',
'}',
'function poll() {',
' if (Date.now() > deadline || process.ppid === 1) process.exit(0);',
'function pollTranscript(transcript) {',
' try {',
' const size = fs.statSync(transcriptPath).size;',
" if (size < position) { position = 0; carry = ''; }",
' if (size > position) {',
' const length = Math.min(size - position, 1048576);',
' const size = fs.statSync(transcript.path).size;',
" if (size < transcript.position) { transcript.position = 0; transcript.carry = ''; }",
' if (size > transcript.position) {',
' const length = Math.min(size - transcript.position, 1048576);',
' const buffer = Buffer.allocUnsafe(length);',
" const fd = fs.openSync(transcriptPath, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, position);',
" const fd = fs.openSync(transcript.path, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, transcript.position);',
' fs.closeSync(fd);',
' position += bytes;',
" carry = (carry + buffer.subarray(0, bytes).toString('utf8')).slice(-262144);",
' inspect(carry);',
' transcript.position += bytes;',
" transcript.carry = (transcript.carry + buffer.subarray(0, bytes).toString('utf8')).slice(-262144);",
' inspect(transcript.carry);',
' }',
' } catch {}',
'}',
'function poll() {',
' if (Date.now() > deadline || process.ppid === 1) process.exit(0);',
' for (const transcript of transcripts) pollTranscript(transcript);',
' setTimeout(poll, 1000);',
'}',
'poll();',
].join('\n');
}
/**
* Keep a Claude subagent alive while its Monitor or background Bash work is live.
* Claude otherwise can publish the worker's last progress sentence as an Agent
* result when one watcher ends, even if other tracked tasks are still running.
*/
export function generateSubagentStopGuardScript(): string {
return [
"const fs = require('node:fs');",
`const ${SUBAGENT_STOP_GUARD_MARKER} = true;`,
'let input = {};',
"try { input = JSON.parse(fs.readFileSync(0, 'utf8') || '{}'); } catch { process.exit(0); }",
"const transcriptPath = typeof input.agent_transcript_path === 'string' ? input.agent_transcript_path : '';",
'if (!transcriptPath) process.exit(0);',
'let text;',
'try {',
' const size = fs.statSync(transcriptPath).size;',
' const length = Math.min(size, 16 * 1024 * 1024);',
' const buffer = Buffer.allocUnsafe(length);',
" const fd = fs.openSync(transcriptPath, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, size - length);',
' fs.closeSync(fd);',
" text = buffer.subarray(0, bytes).toString('utf8');",
'} catch { process.exit(0); }',
'const launched = new Set();',
'const finished = new Set();',
'function inspectToolResult(value) {',
" const serialized = typeof value === 'string' ? value : JSON.stringify(value ?? '');",
' for (const match of serialized.matchAll(/Command running in background with ID:\\s*([A-Za-z0-9_-]+)/gi)) launched.add(match[1]);',
' for (const match of serialized.matchAll(/Monitor started \\(task ([A-Za-z0-9_-]+)/gi)) launched.add(match[1]);',
'}',
'function inspectNotifications(value) {',
" if (typeof value !== 'string' || !value.includes('<task-notification>')) return;",
' for (const match of value.matchAll(/<task-notification>([\\s\\S]*?)<\\/task-notification>/gi)) {',
' const body = match[1];',
' const id = body.match(/<task-id>([^<]+)<\\/task-id>/i);',
' const status = body.match(/<status>(completed|failed|killed|error)<\\/status>/i);',
' if (id && status) finished.add(id[1].trim());',
' }',
'}',
'for (const line of text.split(/\\r?\\n/)) {',
' let entry;',
' try { entry = JSON.parse(line); } catch { continue; }',
' const content = entry && entry.message ? entry.message.content : undefined;',
' if (Array.isArray(content)) {',
' for (const block of content) {',
" if (block && block.type === 'tool_result') inspectToolResult(block.content);",
" if (block && block.type === 'text') inspectNotifications(block.text);",
' }',
' } else {',
' inspectNotifications(content);',
' }',
' inspectNotifications(entry && entry.content);',
'}',
'function findLiveTasks(candidates) {',
' const live = new Set();',
" if (candidates.size === 0 || !fs.existsSync('/proc')) return live;",
' let processIds;',
" try { processIds = fs.readdirSync('/proc').filter((name) => /^\\d+$/.test(name)); } catch { return live; }",
' for (const processId of processIds) {',
" for (const descriptor of ['0', '1', '2']) {",
' let target;',
" try { target = fs.readlinkSync('/proc/' + processId + '/fd/' + descriptor); } catch { continue; }",
' const match = target.match(/[\\/]tasks[\\/]([A-Za-z0-9_-]+)\\.output(?: \\(deleted\\))?$/);',
' if (match && candidates.has(match[1])) live.add(match[1]);',
' }',
' if (live.size === candidates.size) break;',
' }',
' return live;',
'}',
'const unfinished = new Set([...launched].filter((taskId) => !finished.has(taskId)));',
'const active = [...findLiveTasks(unfinished)];',
'if (active.length === 0) process.exit(0);',
'const shown = active.slice(0, 8);',
"const suffix = active.length > shown.length ? ' and ' + (active.length - shown.length) + ' more' : '';",
'process.stdout.write(JSON.stringify({',
" decision: 'block',",
" reason: 'You still own active background work (' + shown.join(', ') + suffix + '). Do not return an intermediate progress message as your final report. Process the task notifications or keep actively polling until every task completes, then return one complete summary.',",
'}));',
].join('\n');
}
function withSettingsLock<T>(path: string, fn: () => Promise<T>): Promise<T> {
const prev = settingsWriteLocks.get(path) ?? Promise.resolve();
const run = prev.then(fn, fn); // run after the prior writer, regardless of its outcome
@@ -204,6 +333,18 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
hooks: [{ type: 'command', command: curlCmd('stop'), timeout: HOOK_TIMEOUT_SECONDS }],
},
],
SubagentStop: [
{
hooks: [
{
type: 'command',
command: 'node',
args: ['-e', generateSubagentStopGuardScript()],
timeout: HOOK_TIMEOUT_SECONDS,
},
],
},
],
TeammateIdle: [
{
hooks: [{ type: 'command', command: curlCmd('teammate_idle'), timeout: HOOK_TIMEOUT_SECONDS }],
@@ -235,8 +376,12 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
function isCodemanHookHandler(value: unknown): boolean {
try {
const serialized = JSON.stringify(value);
// Prefix, not the versioned marker: older script versions must still be ours.
return serialized.includes('/api/hook-event') || serialized.includes(BACKGROUND_WAKE_MARKER_PREFIX);
// Prefixes, not versioned markers: older script versions must still be ours.
return (
serialized.includes('/api/hook-event') ||
serialized.includes(BACKGROUND_WAKE_MARKER_PREFIX) ||
serialized.includes(SUBAGENT_STOP_GUARD_MARKER_PREFIX)
);
} catch {
return false;
}
@@ -430,6 +575,39 @@ export async function writeHooksConfig(casePath: string): Promise<void> {
});
}
/**
* Ensures an explicitly managed case has the current Codeman hooks.
*
* Unlike `refreshStaleCodemanHooks`, this may add Codeman handlers to a valid
* user-owned settings file. It is therefore reserved for case quick-starts,
* where the user has explicitly asked Codeman to manage that workspace. A
* malformed existing file is left untouched rather than replaced.
*/
export async function ensureCodemanHooks(casePath: string): Promise<void> {
const claudeDir = join(casePath, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
await withSettingsLock(settingsPath, async () => {
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
let existing: Record<string, unknown> = {};
try {
const parsed: unknown = JSON.parse(await readFile(settingsPath, 'utf-8'));
if (!parsed || typeof parsed !== 'object' || Array.isArray(parsed)) return;
existing = parsed as Record<string, unknown>;
} catch (err) {
if ((err as NodeJS.ErrnoException).code !== 'ENOENT') return;
}
const generated = generateHooksConfig();
const hooks = mergeCodemanHooks(existing.hooks, generated.hooks);
if (JSON.stringify(existing.hooks ?? {}) === JSON.stringify(hooks)) return;
await writeFile(settingsPath, JSON.stringify({ ...existing, hooks }, null, 2) + '\n');
});
}
/**
* Self-heal a case's Codeman-owned hooks block.
*
@@ -437,10 +615,11 @@ export async function writeHooksConfig(casePath: string): Promise<void> {
* X-Codeman-Hook-Secret header was added (COD-54, 2026-06-10) keep hook curls in their
* settings.local.json that POST to /api/hook-event WITHOUT the secret — which, once the
* gate requires it unconditionally (COD-91), silently 401 on a password-protected install.
* Older Codeman blocks also lack the background Bash async-rewake hook. A third stale
* shape: hook curls without `-k`, which exit 60 on every --https/tailscale install (the
* cert is self-signed), swallowed by the hooks' own `|| true` — all six hook events die
* silently. Refresh any of these stale shapes on launch so existing cases heal.
* Older Codeman blocks also lack the current background Bash async-rewake hook or the
* SubagentStop guard. A further stale shape: hook curls without `-k`, which exit 60 on
* every --https/tailscale install (the cert is self-signed), swallowed by the hooks'
* own `|| true` — all six hook events die silently. Refresh any of these stale shapes
* on launch so existing cases heal.
*
* Deliberately surgical: regenerates ONLY when settings.local.json already contains
* Codeman's own hook curls (they target `/api/hook-event`) and they are stale. No-op
@@ -467,7 +646,8 @@ export async function refreshStaleCodemanHooks(casePath: string): Promise<void>
// as a substring, so this cleanly identifies hook curls that die with exit 60
// on a self-signed HTTPS install.
const hasTlsFlaglessCurl = hooksJson.includes('curl -s -X POST');
if (!isOurs || (hasSecret && hasBackgroundWake && !hasTlsFlaglessCurl)) return;
const hasSubagentStopGuard = hooksJson.includes(SUBAGENT_STOP_GUARD_MARKER);
if (!isOurs || (hasSecret && hasBackgroundWake && hasSubagentStopGuard && !hasTlsFlaglessCurl)) return;
const generated = generateHooksConfig();
const merged = {
...existing,
@@ -541,3 +721,171 @@ export async function applyStatusLineConfig(casePath: string, enabled: boolean):
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
});
}
// ─── Agent skill injection ───────────────────────────────────────────────────
/**
* Version-agnostic ownership prefix for the injected agent skill, same pattern as
* `BACKGROUND_WAKE_MARKER_PREFIX`: ownership is decided on the prefix so a wording
* change in the full marker cannot disown every previously injected copy.
*/
const AGENT_SKILL_MARKER_PREFIX = '<!-- codeman-managed-agent-skill';
/**
* Marker appended to the injected SKILL.md. Its presence is what makes a copy OURS:
* install/refresh/remove all refuse to touch a `skills/codeman` whose SKILL.md lacks
* it, so a user's hand-authored or hand-edited-and-de-marked skill is never clobbered.
*/
const AGENT_SKILL_MARKER = `${AGENT_SKILL_MARKER_PREFIX}: installed by Codeman; edits are overwritten while the agent-skill setting is on -->`;
/**
* Packaged source of the skill: `skills/codeman/` at the package root. Resolved
* relative to this module so it works from `src/` (tsx dev), `dist/` (tsc build),
* and an npm install (`files` includes `skills`), all of which sit one level below
* the package root.
*/
function agentSkillSourceDir(): string {
return join(dirname(fileURLToPath(import.meta.url)), '..', 'skills', 'codeman');
}
interface AgentSkillFile {
/** Path relative to the target skill dir (e.g. `reference/endpoints.md`). */
relPath: string;
content: string;
}
/**
* Read the packaged skill: SKILL.md (marker appended) plus every markdown file
* under `reference/`. Enumerated from disk rather than a hardcoded manifest so a
* new reference file ships without touching this module.
*/
async function readAgentSkillSource(): Promise<AgentSkillFile[]> {
const src = agentSkillSourceDir();
const skill = await readFile(join(src, 'SKILL.md'), 'utf-8');
const files: AgentSkillFile[] = [{ relPath: 'SKILL.md', content: `${skill.trimEnd()}\n\n${AGENT_SKILL_MARKER}\n` }];
let referenceNames: string[] = [];
try {
referenceNames = (await readdir(join(src, 'reference'))).filter((name) => name.endsWith('.md')).sort();
} catch {
// no reference dir in the source; SKILL.md alone is still a valid skill
}
for (const name of referenceNames) {
files.push({ relPath: join('reference', name), content: await readFile(join(src, 'reference', name), 'utf-8') });
}
return files;
}
async function isSymlink(path: string): Promise<boolean> {
try {
return (await lstat(path)).isSymbolicLink();
} catch {
return false;
}
}
/** What an install/remove actually did, so callers (CLI, logs) can say so. */
export type AgentSkillApplyResult =
| 'installed' // fresh copy written
| 'refreshed' // our copy was stale and got rewritten
| 'unchanged' // our copy already matches the packaged source
| 'removed' // our copy deleted
| 'absent' // nothing there to remove
| 'foreign' // a copy exists but is not ours; left untouched
| 'symlink'; // the skill dir (or its parent) is a symlink; left untouched
/**
* Install or refresh the Codeman agent skill into `skillDir` (a `.../codeman`
* directory, e.g. `<case>/.claude/skills/codeman` or `~/.claude/skills/codeman`).
*
* Refuses two shapes rather than writing through them:
* - a SYMLINK at the skill dir or its `skills/` parent: this repo's own dogfooding
* layout (`.claude/skills/codeman -> ../../skills/codeman`) would otherwise have
* the injector overwrite the repo source through the link;
* - a FOREIGN copy (SKILL.md present without our marker): that is the user's own
* skill, and per the statusLine rule we never clobber what we did not write.
*
* Idempotent and cheap: unchanged files are not rewritten, so calling on every
* session create causes no mtime churn.
*/
export async function installAgentSkillInto(skillDir: string): Promise<AgentSkillApplyResult> {
if ((await isSymlink(dirname(skillDir))) || (await isSymlink(skillDir))) return 'symlink';
let existing: string | null = null;
try {
existing = await readFile(join(skillDir, 'SKILL.md'), 'utf-8');
} catch {
// absent: fresh install
}
if (existing !== null && !existing.includes(AGENT_SKILL_MARKER_PREFIX)) return 'foreign';
const files = await readAgentSkillSource();
let changed = false;
for (const file of files) {
const target = join(skillDir, file.relPath);
let current: string | null = null;
try {
current = await readFile(target, 'utf-8');
} catch {
// missing: will be written
}
if (current === file.content) continue;
await mkdir(dirname(target), { recursive: true });
await writeFile(target, file.content);
changed = true;
}
if (!changed) return 'unchanged';
return existing === null ? 'installed' : 'refreshed';
}
/**
* Remove a Codeman-managed skill copy from `skillDir`. Same ownership and symlink
* refusals as the install path. Deletes only files the packaged source would have
* written (never `rm -rf`, so a user's extra files in the directory survive), then
* prunes the directories bottom-up if they emptied.
*/
export async function removeAgentSkillFrom(skillDir: string): Promise<AgentSkillApplyResult> {
if ((await isSymlink(dirname(skillDir))) || (await isSymlink(skillDir))) return 'symlink';
let existing: string | null = null;
try {
existing = await readFile(join(skillDir, 'SKILL.md'), 'utf-8');
} catch {
return 'absent';
}
if (!existing.includes(AGENT_SKILL_MARKER_PREFIX)) return 'foreign';
// Manifest-based, with SKILL.md as the fallback when the packaged source is
// unreadable: removal must still work on an install whose skills/ dir went missing.
const files = await readAgentSkillSource().catch((): AgentSkillFile[] => [{ relPath: 'SKILL.md', content: '' }]);
for (const file of files) {
await unlink(join(skillDir, file.relPath)).catch(() => {});
}
await rmdir(join(skillDir, 'reference')).catch(() => {}); // fails when non-empty, fine
await rmdir(skillDir).catch(() => {});
await rmdir(dirname(skillDir)).catch(() => {}); // prune `.claude/skills` if now empty
return 'removed';
}
/**
* Add or remove the Codeman agent skill in `<case>/.claude/skills/codeman`,
* mirroring `applyStatusLineConfig`'s shape. Gated by the synced `agentSkillEnabled`
* app setting (default OFF); callers gate on Claude mode, since the skill is discovered
* via `.claude/skills/`, which only Claude Code reads.
*
* Call-site policy is ADD-ONLY on session create (callers pass `enabled: true` or
* skip the call), for the statusLine reason: sessions in a repo share one `.claude/`
* dir, so a single create while the setting is off must not yank the skill out from
* under other live sessions.
*
* ⚠️ Consequence: turning `agentSkillEnabled` OFF sweeps nothing. There is deliberately
* no server-side toggle-off sweep (it would have to walk every case, including ones
* with live sessions, and would hit exactly the shared-`.claude/` hazard above), so
* already-injected copies stay on disk until removed per case with
* `codeman skill uninstall --case <name>`. The `enabled: false` branch here backs that
* CLI and the tests; it has no server call site. Keep the README's Agent Skill note in
* sync if this ever changes.
*/
export async function applyAgentSkill(casePath: string, enabled: boolean): Promise<AgentSkillApplyResult> {
const skillDir = join(casePath, '.claude', 'skills', 'codeman');
return enabled ? installAgentSkillInto(skillDir) : removeAgentSkillFrom(skillDir);
}
+401
View File
@@ -0,0 +1,401 @@
/**
* @fileoverview `codeman service install|uninstall|status`: write and load the
* systemd user unit (Linux) or LaunchAgent (macOS) that supervises `codeman web`.
*
* This is the "always running" half of issue #231, next to the "detached right
* now" half in daemon-control.ts. `install.sh` already does this for people who
* install with the one-liner; this exists for `npm i -g aicodeman` users, who
* otherwise have to hand-write a plist.
*
* Two details are load-bearing and easy to get wrong by hand:
*
* - **PATH.** launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` and systemd's
* user manager is nearly as bare, so a Homebrew or nvm `node`, `tmux` or
* `claude` is simply not found and sessions fail in a way that reads as a
* Codeman bug. The unit therefore carries the PATH of the shell that ran the
* install, with the running node's own directory in front.
* - **The job name.** It is the one `install.sh` and the self-updater already use
* (config/service-names.ts), so re-running install.sh later updates this unit
* instead of supervising a second copy of the server.
*
* Secrets are deliberately NOT written here. `CODEMAN_PASSWORD` in the installing
* shell is not copied into the unit; the caller is told where to add it instead,
* because a unit file is long-lived, world-readable by default, and gets copied
* into bug reports.
*
* The file writers are pure string builders so they can be unit-tested without
* touching launchctl/systemctl.
*
* @module service-installer
*/
import { execFileSync } from 'node:child_process';
import { existsSync, mkdirSync, unlinkSync, writeFileSync } from 'node:fs';
import { homedir, userInfo } from 'node:os';
import { dirname, join } from 'node:path';
import { LAUNCHD_LABEL, SYSTEMD_UNIT } from './config/service-names.js';
import { CODEMAN_INSTANCE } from './config/instance.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import {
buildBaseUrl,
buildStatusUrl,
buildWebArgs,
logFilePath,
probeServer,
type WebLaunchOptions,
} from './daemon-control.js';
export type ServiceKind = 'launchd' | 'systemd';
/** Everything a unit file needs, resolved from the environment by the caller. */
export interface ServicePlan {
kind: ServiceKind;
/** systemd unit filename or launchd label. */
name: string;
nodePath: string;
/** Runner flags carried over from the current process (tsx loader in dev). */
execArgv: string[];
scriptPath: string;
args: string[];
env: Record<string, string>;
logPath: string;
workingDir: string;
}
export interface ServiceActionResult {
ok: boolean;
message: string;
/** Path of the unit/plist that was written or removed. */
unitPath?: string;
warnings?: string[];
}
export interface ServiceStatusResult {
kind: ServiceKind | null;
name: string;
unitPath: string;
installed: boolean;
loaded: boolean;
responding: boolean;
version?: string;
url: string;
}
/** Directories worth having on PATH even when the installing shell lacked them. */
const FALLBACK_PATH_DIRS = ['/opt/homebrew/bin', '/usr/local/bin', '/usr/bin', '/bin', '/usr/sbin', '/sbin'];
// ─────────────────────────────────────────────────────────────────────────────
// Pure builders
// ─────────────────────────────────────────────────────────────────────────────
/** XML text escaping for plist `<string>` values. */
export function xmlEscape(value: string): string {
return value
.replace(/&/g, '&amp;')
.replace(/</g, '&lt;')
.replace(/>/g, '&gt;')
.replace(/"/g, '&quot;')
.replace(/'/g, '&apos;');
}
/**
* PATH for the supervised process: the running node's directory first (so an nvm
* or Homebrew node is used rather than whatever the supervisor finds), then the
* installing shell's PATH, then the fallbacks that are still missing.
*
* `node_modules/.bin` entries are dropped. npm and npx inject those for the
* lifetime of one command, and baking a project's local bin dir into a unit file
* that outlives the checkout is how a service ends up running a binary the
* operator deleted months ago.
*/
export function buildServicePath(nodeDir: string, currentPath: string, home: string): string {
const seen = new Set<string>();
const ordered: string[] = [];
const push = (dir: string) => {
const trimmed = dir.trim();
if (!trimmed || seen.has(trimmed)) return;
if (/(^|\/)node_modules\/\.bin\/?$/.test(trimmed)) return;
seen.add(trimmed);
ordered.push(trimmed);
};
push(nodeDir);
for (const dir of currentPath.split(':')) push(dir);
push(join(home, '.local', 'bin'));
for (const dir of FALLBACK_PATH_DIRS) push(dir);
return ordered.join(':');
}
/** Environment written into the unit. Never includes secrets (see module docs). */
export function buildServiceEnv(
nodeDir: string,
currentPath: string,
home: string,
lang?: string
): Record<string, string> {
const env: Record<string, string> = {
PATH: buildServicePath(nodeDir, currentPath, home),
HOME: home,
LANG: lang || 'en_US.UTF-8',
};
if (CODEMAN_INSTANCE) env.CODEMAN_INSTANCE = CODEMAN_INSTANCE;
return env;
}
/** systemd accepts double-quoted values; escape the two characters that matter. */
export function systemdQuote(value: string): string {
return `"${value.replace(/\\/g, '\\\\').replace(/"/g, '\\"')}"`;
}
export function buildLaunchAgentPlist(plan: ServicePlan): string {
const programArguments = [plan.nodePath, ...plan.execArgv, plan.scriptPath, ...plan.args]
.map((arg) => ` <string>${xmlEscape(arg)}</string>`)
.join('\n');
const environment = Object.entries(plan.env)
.map(([key, value]) => ` <key>${xmlEscape(key)}</key>\n <string>${xmlEscape(value)}</string>`)
.join('\n');
return `<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>${xmlEscape(plan.name)}</string>
<key>ProgramArguments</key>
<array>
${programArguments}
</array>
<key>EnvironmentVariables</key>
<dict>
${environment}
</dict>
<key>WorkingDirectory</key>
<string>${xmlEscape(plan.workingDir)}</string>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
<key>ThrottleInterval</key>
<integer>10</integer>
<key>StandardOutPath</key>
<string>${xmlEscape(plan.logPath)}</string>
<key>StandardErrorPath</key>
<string>${xmlEscape(plan.logPath)}</string>
</dict>
</plist>
`;
}
export function buildSystemdUnit(plan: ServicePlan): string {
const execStart = [plan.nodePath, ...plan.execArgv, plan.scriptPath, ...plan.args]
.map((arg) => (/[\s"'\\]/.test(arg) ? systemdQuote(arg) : arg))
.join(' ');
const environment = Object.entries(plan.env)
.map(([key, value]) => `Environment=${systemdQuote(`${key}=${value}`)}`)
.join('\n');
return `[Unit]
Description=Codeman Web Server
After=network.target
[Service]
Type=simple
WorkingDirectory=${plan.workingDir}
ExecStart=${execStart}
Restart=always
RestartSec=10
# Agents keep running in tmux when the server restarts, so only signal the
# server itself.
KillMode=process
${environment}
StandardOutput=journal
StandardError=journal
SyslogIdentifier=codeman
LimitNOFILE=65536
[Install]
WantedBy=default.target
`;
}
// ─────────────────────────────────────────────────────────────────────────────
// Environment resolution
// ─────────────────────────────────────────────────────────────────────────────
export function detectServiceKind(): ServiceKind | null {
if (process.platform === 'darwin') return 'launchd';
if (process.platform === 'linux') return 'systemd';
return null;
}
export function unitPathFor(kind: ServiceKind): string {
return kind === 'launchd'
? join(homedir(), 'Library', 'LaunchAgents', `${LAUNCHD_LABEL}.plist`)
: join(homedir(), '.config', 'systemd', 'user', SYSTEMD_UNIT);
}
function entryScript(): string {
const script = process.argv[1];
if (!script) throw new Error('cannot determine the codeman entry script to supervise');
return script;
}
/** Resolve a full plan from the current process and the requested web options. */
export function resolveServicePlan(kind: ServiceKind, options: WebLaunchOptions): ServicePlan {
const home = homedir();
return {
kind,
name: kind === 'launchd' ? LAUNCHD_LABEL : SYSTEMD_UNIT,
nodePath: process.execPath,
execArgv: [...process.execArgv],
scriptPath: entryScript(),
args: buildWebArgs(options),
env: buildServiceEnv(dirname(process.execPath), process.env.PATH || '', home, process.env.LANG),
logPath: logFilePath(),
workingDir: home,
};
}
function run(command: string, args: string[]): { ok: boolean; output: string } {
try {
const output = execFileSync(command, args, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'pipe'],
});
return { ok: true, output: output.trim() };
} catch (err) {
const e = err as { stderr?: Buffer | string; message?: string };
const stderr = typeof e.stderr === 'string' ? e.stderr : e.stderr?.toString('utf-8');
return { ok: false, output: (stderr || e.message || '').trim() };
}
}
// ─────────────────────────────────────────────────────────────────────────────
// Install / uninstall / status
// ─────────────────────────────────────────────────────────────────────────────
/**
* Write the unit, load it, and confirm the server actually answers before
* reporting success. `launchctl load` and `systemctl enable` are both quiet about
* a job that starts and immediately dies, which is the whole reason install.sh
* verifies too.
*/
export async function installService(options: WebLaunchOptions): Promise<ServiceActionResult> {
const kind = detectServiceKind();
if (!kind) {
return { ok: false, message: `no supported supervisor on ${process.platform}; use \`codeman web -d\` instead` };
}
const plan = resolveServicePlan(kind, options);
const unitPath = unitPathFor(kind);
const warnings: string[] = [];
mkdirSync(dirname(unitPath), { recursive: true });
if (kind === 'launchd') {
const uid = process.getuid?.() ?? 0;
// Unload any previous copy first, otherwise bootstrap fails with "service
// already loaded" and leaves the OLD job running against the NEW file.
run('launchctl', ['bootout', `gui/${uid}/${LAUNCHD_LABEL}`]);
writeFileSync(unitPath, buildLaunchAgentPlist(plan), { encoding: 'utf-8', mode: 0o600 });
const bootstrap = run('launchctl', ['bootstrap', `gui/${uid}`, unitPath]);
if (!bootstrap.ok) {
const legacy = run('launchctl', ['load', unitPath]);
if (!legacy.ok) {
return {
ok: false,
unitPath,
message: `wrote ${unitPath} but launchctl refused to load it: ${bootstrap.output}`,
};
}
}
} else {
writeFileSync(unitPath, buildSystemdUnit(plan), { encoding: 'utf-8', mode: 0o600 });
const reload = run('systemctl', ['--user', 'daemon-reload']);
if (!reload.ok) {
return {
ok: false,
unitPath,
message: `wrote ${unitPath} but \`systemctl --user daemon-reload\` failed: ${reload.output}`,
};
}
const enable = run('systemctl', ['--user', 'enable', '--now', SYSTEMD_UNIT]);
if (!enable.ok) {
return { ok: false, unitPath, message: `wrote ${unitPath} but enabling it failed: ${enable.output}` };
}
// Without lingering the unit stops at logout, which is exactly what someone
// installing a service does not want. Best effort: it needs polkit rights.
const linger = run('loginctl', ['enable-linger', userInfo().username]);
if (!linger.ok) {
warnings.push(
`could not enable lingering, so the service will stop when you log out. Run: sudo loginctl enable-linger ${userInfo().username}`
);
}
}
const url = buildBaseUrl(options);
const statusUrl = buildStatusUrl(options);
const deadline = Date.now() + 30_000;
while (Date.now() < deadline) {
const probe = await probeServer(statusUrl, 1000);
if (probe.up) {
return { ok: true, unitPath, warnings, message: `service installed and responding at ${url}` };
}
await new Promise((resolve) => setTimeout(resolve, 500));
}
const hint =
kind === 'launchd' ? `tail -20 ${plan.logPath}` : `journalctl --user -u ${SYSTEMD_UNIT} -n 20 --no-pager`;
return {
ok: false,
unitPath,
warnings,
message: `wrote and loaded ${unitPath}, but nothing answered ${url} within 30s. Check: ${hint}`,
};
}
export function uninstallService(): ServiceActionResult {
const kind = detectServiceKind();
if (!kind) return { ok: false, message: `no supported supervisor on ${process.platform}` };
const unitPath = unitPathFor(kind);
if (!existsSync(unitPath)) {
return { ok: false, unitPath, message: `no service installed at ${unitPath}` };
}
if (kind === 'launchd') {
const uid = process.getuid?.() ?? 0;
const bootout = run('launchctl', ['bootout', `gui/${uid}/${LAUNCHD_LABEL}`]);
if (!bootout.ok) run('launchctl', ['unload', unitPath]);
} else {
run('systemctl', ['--user', 'disable', '--now', SYSTEMD_UNIT]);
}
try {
unlinkSync(unitPath);
} catch (err) {
return { ok: false, unitPath, message: `stopped the service but could not remove ${unitPath}: ${String(err)}` };
}
if (kind === 'systemd') run('systemctl', ['--user', 'daemon-reload']);
return { ok: true, unitPath, message: `service stopped and ${unitPath} removed. Your tmux sessions are untouched.` };
}
export async function serviceStatus(options: WebLaunchOptions): Promise<ServiceStatusResult> {
const kind = detectServiceKind();
const url = buildBaseUrl(options);
if (!kind) {
return { kind: null, name: '', unitPath: '', installed: false, loaded: false, responding: false, url };
}
const unitPath = unitPathFor(kind);
const name = kind === 'launchd' ? LAUNCHD_LABEL : SYSTEMD_UNIT;
const installed = existsSync(unitPath);
const loaded =
kind === 'launchd'
? run('launchctl', ['list', LAUNCHD_LABEL]).ok
: run('systemctl', ['--user', 'is-active', SYSTEMD_UNIT]).output === 'active';
const probe = await probeServer(buildStatusUrl(options), 2000);
return { kind, name, unitPath, installed, loaded, responding: probe.up, version: probe.version, url };
}
+2
View File
@@ -17,6 +17,8 @@ export interface ConfigPort {
getModelConfig(): Promise<{ defaultModel?: string; agentTypeOverrides?: Record<string, string> } | null>;
getClaudeModeConfig(): Promise<{ claudeMode?: ClaudeMode; allowedTools?: string }>;
getTerminalHistoryConfig(): Promise<TerminalHistoryConfig>;
/** Synced `agentSkillEnabled` app setting (default OFF); gates per-case agent-skill injection. */
getAgentSkillEnabled(): Promise<boolean>;
getDefaultClaudeMdPath(): Promise<string | undefined>;
getLightState(identity?: { username: string; role: 'admin' | 'user' }): unknown;
getLightSessionsState(): unknown[];
+9 -1
View File
@@ -1322,7 +1322,7 @@
</div>
<!-- Input Section -->
<div class="settings-section-header">Input</div>
<div class="settings-item settings-item-multiline" title="Scroll the terminal's own local scrollback with a plain mouse wheel / two-finger swipe, instead of forwarding the wheel to the CLI's transcript. Leave OFF for Claude/Codex sessions: those CLIs redraw the screen in place and keep no local scrollback, so the wheel would have almost nothing to scroll. In Claude sessions Codeman then falls back to paging the CLI's own transcript; Codex sessions have no such fallback, so the wheel goes dead. Shift+wheel always reaches local scrollback regardless.">
<div class="settings-item settings-item-multiline" title="Scroll the terminal's own local scrollback with a plain mouse wheel / two-finger swipe, instead of forwarding the wheel to the CLI's transcript. Only Claude sessions forward, so this setting only affects them: Claude redraws the screen in place and keeps almost no local scrollback, so with this ON the wheel has little to scroll and Codeman falls back to paging Claude's own transcript. Codex, Gemini, shell and OpenCode sessions always scroll local scrollback. Shift+wheel always reaches local scrollback regardless.">
<div class="settings-item-text">
<span class="settings-item-label">Wheel Scrolls Local History</span>
<span class="settings-item-desc">Plain wheel/trackpad pages the terminal scrollback</span>
@@ -1637,6 +1637,14 @@
</label>
<span class="form-hint">Enable experimental Agent Teams for all new Claude sessions (disabled by default)</span>
</div>
<div class="form-row form-row-switch">
<label>Agent Skill</label>
<label class="switch">
<input type="checkbox" id="appSettingsAgentSkill">
<span class="slider"></span>
</label>
<span class="form-hint">Give new Claude sessions the Codeman skill (start workers, send prompts, wait for results via the API)</span>
</div>
<div class="form-row">
<label>Claude Model</label>
<select id="appSettingsClaudeModel" class="form-select">
+2
View File
@@ -385,6 +385,7 @@ Object.assign(CodemanApp.prototype, {
this._applyCodexSettingsVisibility();
// Claude Permissions settings
document.getElementById('appSettingsAgentTeams').checked = settings.agentTeamsEnabled ?? false;
document.getElementById('appSettingsAgentSkill').checked = settings.agentSkillEnabled ?? false;
document.getElementById('appSettingsClaudeModel').value = settings.claudeModel ?? '';
document.getElementById('appSettingsOpusContext1m').checked = settings.opusContext1mEnabled ?? false;
document.getElementById('appSettingsRemoteAutoReconnect').checked = settings.remoteAutoReconnect ?? true;
@@ -1553,6 +1554,7 @@ Object.assign(CodemanApp.prototype, {
codexAnimationsEnabled: document.getElementById('appSettingsCodexAnimations').checked,
// Claude Permissions settings
agentTeamsEnabled: document.getElementById('appSettingsAgentTeams').checked,
agentSkillEnabled: document.getElementById('appSettingsAgentSkill').checked,
claudeModel: document.getElementById('appSettingsClaudeModel').value,
opusContext1mEnabled: document.getElementById('appSettingsOpusContext1m').checked,
remoteAutoReconnect: document.getElementById('appSettingsRemoteAutoReconnect').checked,
+40 -37
View File
@@ -341,7 +341,7 @@ Object.assign(CodemanApp.prototype, {
// WebGL renderer for GPU-accelerated terminal rendering.
// Previously caused "page unresponsive" crashes from synchronous GPU stalls,
// but the 48KB/frame flush cap in flushPendingWrites() now prevents
// but the mode-aware 32/64KB frame cap in flushPendingWrites() now prevents
// oversized terminal.write() calls that triggered the stalls.
// Disable with ?nowebgl URL param if GPU issues return.
// Auto-fallback: _initWebGL installs a long-task watchdog that disables
@@ -452,8 +452,8 @@ Object.assign(CodemanApp.prototype, {
this.registerFilePathLinkProvider();
// Mouse wheel: forward to the TUI only for sessions verified to handle SGR
// wheel reports (codex, and claude 2.1.187+ — see _shouldForwardWheelToApp),
// local scrollback otherwise. Claude Code 2.1.187+ scrolls its own
// wheel reports (claude 2.1.187+ — see _shouldForwardWheelToApp), local
// scrollback otherwise. Claude Code 2.1.187+ scrolls its own
// transcript on SGR wheel reports — scrolled-away tool blocks re-render
// live and stay clickable — and its select menus no longer capture wheel
// as option navigation (verified against 2.1.202: /model menu highlight
@@ -2349,17 +2349,26 @@ Object.assign(CodemanApp.prototype, {
// Accumulate raw data (may contain DEC 2026 markers)
this.pendingWrites.push(data);
this._scheduleTerminalWriteFlush();
},
if (!this.writeFrameScheduled) {
this.writeFrameScheduled = true;
this._safeYield(() => {
// xterm.js 6.0 handles DEC 2026 sync markers natively — it buffers
// content between 2026h/2026l and renders atomically. No need for
// client-side incomplete-block detection; just flush every frame.
this.flushPendingWrites();
this.writeFrameScheduled = false;
});
}
/**
* Schedule one render-budgeted terminal flush.
*
* Clear the scheduled flag before flushing so flushPendingWrites() can queue
* another yield when a large final batch leaves bytes behind. Keeping the
* flag set through the flush stranded that remainder until unrelated output
* arrived, which looked like truncated responses and idle shell commands.
*/
_scheduleTerminalWriteFlush() {
if (this.writeFrameScheduled || this.pendingWrites.length === 0) return;
this.writeFrameScheduled = true;
this._safeYield(() => {
this.writeFrameScheduled = false;
// xterm.js 6.0 handles DEC 2026 sync markers natively — it buffers
// content between 2026h/2026l and renders atomically.
this.flushPendingWrites();
});
},
/**
@@ -2375,13 +2384,7 @@ Object.assign(CodemanApp.prototype, {
this.flickerFilterActive = false;
// Trigger a normal flush
if (!this.writeFrameScheduled) {
this.writeFrameScheduled = true;
this._safeYield(() => {
this.flushPendingWrites();
this.writeFrameScheduled = false;
});
}
this._scheduleTerminalWriteFlush();
},
/**
@@ -2530,13 +2533,7 @@ Object.assign(CodemanApp.prototype, {
this.terminal.write(joined.slice(0, MAX_FRAME_BYTES));
this.pendingWrites.push(joined.slice(MAX_FRAME_BYTES));
deferred = true;
if (!this.writeFrameScheduled) {
this.writeFrameScheduled = true;
this._safeYield(() => {
this.flushPendingWrites();
this.writeFrameScheduled = false;
});
}
this._scheduleTerminalWriteFlush();
}
if (
preserveViewportY !== null &&
@@ -3110,11 +3107,20 @@ Object.assign(CodemanApp.prototype, {
// Wheel forwarding gate for the container wheel handler: no Shift override,
// xterm's own encoder dormant, viewport at the bottom, and a TUI VERIFIED to
// scroll its transcript on SGR wheel reports: codex, or claude 2.1.187+
// (older Claude Code captures wheel as select-menu option navigation; an
// unknown version is treated as older). Gemini is a strip mode too but its
// wheel behavior is unverified, so it keeps the local wheel — taps/clicks
// are still forwarded for it (harmless no-ops at worst).
// scroll its transcript on SGR wheel reports — which today is claude 2.1.187+
// and nothing else (older Claude Code captures wheel as select-menu option
// navigation; an unknown version is treated as older). Gemini and codex are
// strip modes too but keep the local wheel — taps/clicks are still forwarded
// for them (harmless no-ops at worst).
//
// Codex USED to forward here and was the #227 regression (DodgyBadger, Codex
// latest / Chrome / Win11: dead wheel in codex, working scrollbar drag).
// Measured on codex-cli 0.147.0 in a bare tmux: it never enables mouse
// tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change
// NOTHING on screen — it runs an inline viewport (`alternate_on=0`) and pushes
// its transcript into the terminal's own scrollback (tmux `history_size`
// grows), so there is no in-app pager to drive and local scrollback IS the
// codex transcript. Forwarding therefore swallowed every tick.
// Wheel delta → whole scroll lines. macOS trackpads turn Shift+two-finger
// scroll into a HORIZONTAL wheel (deltaY≈0, deltaX carries the magnitude), and
// Shift routes the wheel to local scrollback (_shouldForwardWheelToApp returns
@@ -3168,11 +3174,8 @@ Object.assign(CodemanApp.prototype, {
if (mode && mode !== 'none') return false;
const session = this.sessions?.get(this.activeSessionId);
const sessionMode = session?.mode || 'claude';
if (sessionMode === 'claude') {
if (!this._cliVersionAtLeast(session?.cliVersion, '2.1.187')) return false;
} else if (sessionMode !== 'codex') {
return false;
}
if (sessionMode !== 'claude') return false;
if (!this._cliVersionAtLeast(session?.cliVersion, '2.1.187')) return false;
// Deliberately NOT gated on _terminalViewportAtBottom(). It used to be, so
// that leaving the bottom handed the wheel back to local scrollback and both
// histories stayed reachable without a mode switch. In practice that inverted
+17
View File
@@ -79,6 +79,7 @@ import {
updateCaseModel,
stripCaseEnvKeys,
applyStatusLineConfig,
applyAgentSkill,
refreshStaleCodemanHooks,
} from '../../hooks-config.js';
import { generateClaudeMd } from '../../templates/claude-md.js';
@@ -699,6 +700,13 @@ export function registerSessionRoutes(
// cases (writeHooksConfig already wrote the secret) and for non-Codeman/absent hooks.
if ((body.mode ?? 'claude') === 'claude') {
await refreshStaleCodemanHooks(workingDir).catch(() => {});
// Agent skill (docs/agent-control-plan.md §2): ADD-ONLY on create, same shared-
// .claude rationale as the statusLine above: a create must never remove the
// skill from under other live sessions in the repo. Marker-guarded, so a
// user's own skills/codeman is never touched.
if (await ctx.getAgentSkillEnabled()) {
await applyAgentSkill(workingDir, true).catch(() => {});
}
}
// Check OpenCode availability if requested
@@ -2766,6 +2774,15 @@ export function registerSessionRoutes(
await refreshStaleCodemanHooks(resolvedCasePath).catch(() => {});
}
// Agent skill injection (docs/agent-control-plan.md §2): ADD-ONLY on create,
// marker-guarded (a user's own skills/codeman is never touched). Claude mode only
// (`.claude/skills/` is a Claude Code surface); skipped for remote cases, whose
// casePath lives on another host. Docker cases qualify: hostWorkspacePath is a
// real host dir and the skill crosses the bind mount like the rest of `.claude/`.
if (!remote && mode === 'claude' && (await ctx.getAgentSkillEnabled())) {
await applyAgentSkill(resolvedCasePath, true).catch(() => {});
}
// Docker cases: the workspace is a REAL host dir bind-mounted into the container.
// Scaffold hooks (+ a CLAUDE.md) if MISSING so in-container permission prompts and
// hook-idle detection fire (decision: wire hooks now). Never clobbers an existing
+8
View File
@@ -760,6 +760,14 @@ export const SettingsUpdateSchema = z
/** Floating ultracode run windows w/ tab connector lines (default OFF). Also starts workflowRunWatcher. SYNCED. */
ultracodeFloatingWindows: z.boolean().optional(),
imageWatcherEnabled: z.boolean().optional(),
/**
* Inject the Codeman agent skill (`skills/codeman`) into `<case>/.claude/skills/`
* on Claude session create, so an agent inside the session can drive the API
* (see docs/agent-control-plan.md §2). SYNCED, default OFF: every skill's
* name+description costs context on every turn, so it is opt-in. Injection is
* add-only at create; a marker keeps user-authored copies untouched.
*/
agentSkillEnabled: z.boolean().optional(),
tunnelEnabled: z.boolean().optional(),
// Action field (NOT persisted): explicit per-request acknowledgment that the
// operator accepts exposing an UNAUTHENTICATED public tunnel (no CODEMAN_PASSWORD).
+4 -4
View File
@@ -30,6 +30,7 @@ import { homedir, tmpdir } from 'node:os';
import { randomUUID } from 'node:crypto';
import { createRequire } from 'node:module';
import { dataPath } from '../config/instance.js';
import { LAUNCHD_LABEL, SYSTEMD_UNIT } from '../config/service-names.js';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import type {
InstallInfo,
@@ -43,10 +44,9 @@ import type {
const require = createRequire(import.meta.url);
const { version: APP_VERSION } = require('../../package.json') as { version: string };
/** systemd unit name (matches install.sh + scripts/codeman-web.service). */
const SYSTEMD_UNIT = 'codeman-web.service';
/** launchd agent label (matches install.sh setup_launchd_service). */
const LAUNCHD_LABEL = 'com.codeman.web';
// Unit name / job label live in config/service-names.ts so install.sh, this
// detector and `codeman service install` cannot drift apart. Unchanged for the
// default instance.
/** Path to the persisted update status file. */
const STATUS_FILE = dataPath('update-status.json');
/** Network/git timeout for the "check" path (longer than EXEC_TIMEOUT_MS — ls-remote hits the network). */
+8
View File
@@ -622,6 +622,7 @@ export class WebServer extends EventEmitter {
getModelConfig: this.getModelConfig.bind(this),
getClaudeModeConfig: this.getClaudeModeConfig.bind(this),
getTerminalHistoryConfig: this.getTerminalHistoryConfig.bind(this),
getAgentSkillEnabled: this.getAgentSkillEnabled.bind(this),
getDefaultClaudeMdPath: this.getDefaultClaudeMdPath.bind(this),
getLightState: this.getLightState.bind(this),
getLightSessionsState: this.getLightSessionsState.bind(this),
@@ -1652,6 +1653,13 @@ export class WebServer extends EventEmitter {
return resolveTerminalHistoryConfig(settings);
}
// Whether the Codeman agent skill is injected into cases on Claude session create
// (synced `agentSkillEnabled` setting, default OFF; docs/agent-control-plan.md §2).
private async getAgentSkillEnabled(): Promise<boolean> {
const settings = await this.readSettings();
return settings.agentSkillEnabled === true;
}
// Helper to get model configuration from settings
private async getModelConfig(): Promise<{
defaultModel?: string;
+122
View File
@@ -0,0 +1,122 @@
/**
* @fileoverview Unit tests for the agent-skill injection helpers in hooks-config.ts
* (`applyAgentSkill`, `installAgentSkillInto`, `removeAgentSkillFrom`).
*
* These run against the REAL packaged source (`skills/codeman/` at the repo root),
* so they double as a guard that the skill files exist and are readable: an npm
* publish without them would be caught here before the `files` entry silently
* ignores the missing directory.
*
* Pure filesystem tests in a per-test temp dir. Port: N/A.
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { mkdtemp, rm, mkdir, writeFile, readFile, symlink, readdir } from 'node:fs/promises';
import { existsSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import { applyAgentSkill, installAgentSkillInto, removeAgentSkillFrom } from '../src/hooks-config.js';
const MARKER_PREFIX = '<!-- codeman-managed-agent-skill';
let casePath: string;
const skillDir = () => join(casePath, '.claude', 'skills', 'codeman');
beforeEach(async () => {
casePath = await mkdtemp(join(tmpdir(), 'codeman-agent-skill-'));
});
afterEach(async () => {
await rm(casePath, { recursive: true, force: true });
});
describe('installAgentSkillInto / applyAgentSkill(enabled)', () => {
it('installs SKILL.md (marker appended) and the reference files from the packaged source', async () => {
const result = await applyAgentSkill(casePath, true);
expect(result).toBe('installed');
const skillMd = await readFile(join(skillDir(), 'SKILL.md'), 'utf-8');
expect(skillMd.startsWith('---\nname: codeman')).toBe(true);
expect(skillMd).toContain(MARKER_PREFIX);
// Reference files ride along byte-for-byte (no marker there).
const sourceEndpoints = await readFile(
join(process.cwd(), 'skills', 'codeman', 'reference', 'endpoints.md'),
'utf-8'
);
const injectedEndpoints = await readFile(join(skillDir(), 'reference', 'endpoints.md'), 'utf-8');
expect(injectedEndpoints).toBe(sourceEndpoints);
expect(existsSync(join(skillDir(), 'reference', 'recipes.md'))).toBe(true);
});
it('is idempotent: a second run reports unchanged', async () => {
await applyAgentSkill(casePath, true);
expect(await applyAgentSkill(casePath, true)).toBe('unchanged');
});
it('refreshes a stale Codeman-managed copy back to the packaged content', async () => {
await applyAgentSkill(casePath, true);
const original = await readFile(join(skillDir(), 'SKILL.md'), 'utf-8');
// Simulate an older injected version: content differs but the marker is intact.
await writeFile(join(skillDir(), 'SKILL.md'), `stale content\n${MARKER_PREFIX}: old -->\n`);
expect(await applyAgentSkill(casePath, true)).toBe('refreshed');
expect(await readFile(join(skillDir(), 'SKILL.md'), 'utf-8')).toBe(original);
});
it('never clobbers a user-authored skills/codeman (no marker)', async () => {
await mkdir(skillDir(), { recursive: true });
await writeFile(join(skillDir(), 'SKILL.md'), '---\nname: codeman\n---\nmy own skill\n');
expect(await applyAgentSkill(casePath, true)).toBe('foreign');
expect(await readFile(join(skillDir(), 'SKILL.md'), 'utf-8')).toContain('my own skill');
expect(existsSync(join(skillDir(), 'reference'))).toBe(false);
});
it('refuses to write through a symlinked skill dir (dogfooding layout)', async () => {
await mkdir(join(casePath, '.claude', 'skills'), { recursive: true });
await symlink(join(casePath, 'elsewhere'), skillDir());
expect(await installAgentSkillInto(skillDir())).toBe('symlink');
});
it('refuses to write through a symlinked skills/ parent', async () => {
await mkdir(join(casePath, 'real-skills'), { recursive: true });
await mkdir(join(casePath, '.claude'), { recursive: true });
await symlink(join(casePath, 'real-skills'), join(casePath, '.claude', 'skills'));
expect(await installAgentSkillInto(skillDir())).toBe('symlink');
expect(await readdir(join(casePath, 'real-skills'))).toEqual([]);
});
});
describe('removeAgentSkillFrom / applyAgentSkill(disabled)', () => {
it('removes our copy and prunes the emptied directories', async () => {
await applyAgentSkill(casePath, true);
expect(await applyAgentSkill(casePath, false)).toBe('removed');
expect(existsSync(skillDir())).toBe(false);
expect(existsSync(join(casePath, '.claude', 'skills'))).toBe(false);
// `.claude` itself is not ours to prune.
expect(existsSync(join(casePath, '.claude'))).toBe(true);
});
it('reports absent when there is nothing to remove', async () => {
expect(await applyAgentSkill(casePath, false)).toBe('absent');
});
it('leaves a user-authored copy untouched', async () => {
await mkdir(skillDir(), { recursive: true });
await writeFile(join(skillDir(), 'SKILL.md'), 'my own skill\n');
expect(await applyAgentSkill(casePath, false)).toBe('foreign');
expect(existsSync(join(skillDir(), 'SKILL.md'))).toBe(true);
});
it("preserves a user's extra files in the directory (no rm -rf)", async () => {
await applyAgentSkill(casePath, true);
await writeFile(join(skillDir(), 'reference', 'my-notes.md'), 'mine\n');
expect(await applyAgentSkill(casePath, false)).toBe('removed');
expect(existsSync(join(skillDir(), 'SKILL.md'))).toBe(false);
expect(existsSync(join(skillDir(), 'reference', 'endpoints.md'))).toBe(false);
// The user's file and the directories holding it survive.
expect(await readFile(join(skillDir(), 'reference', 'my-notes.md'), 'utf-8')).toBe('mine\n');
});
});
+54 -40
View File
@@ -71,7 +71,8 @@ describe('AiIdleChecker', () => {
describe('Output Parsing', () => {
it('should parse IDLE verdict', async () => {
// Set up mock to return IDLE result after polling
mockedReadFileSync.mockReturnValueOnce('') // writeFileSync creates empty file
mockedReadFileSync
.mockReturnValueOnce('') // writeFileSync creates empty file
.mockReturnValueOnce('IDLE\nSession shows completion message and prompt.\n__AICHECK_DONE__');
const checkPromise = checker.check('some terminal output');
@@ -87,7 +88,8 @@ describe('AiIdleChecker', () => {
});
it('should parse WORKING verdict', async () => {
mockedReadFileSync.mockReturnValueOnce('')
mockedReadFileSync
.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nSpinner characters detected, still processing.\n__AICHECK_DONE__');
const checkPromise = checker.check('some terminal output');
@@ -100,8 +102,7 @@ describe('AiIdleChecker', () => {
});
it('should handle lowercase verdict', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('idle\nDone.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('idle\nDone.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
@@ -112,7 +113,8 @@ describe('AiIdleChecker', () => {
});
it('should return ERROR for unparseable output', async () => {
mockedReadFileSync.mockReturnValueOnce('')
mockedReadFileSync
.mockReturnValueOnce('')
.mockReturnValueOnce('Something unexpected happened.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
@@ -125,8 +127,7 @@ describe('AiIdleChecker', () => {
});
it('should return ERROR for empty output', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
@@ -175,10 +176,34 @@ describe('AiIdleChecker', () => {
await vi.advanceTimersByTimeAsync(500);
await checkPromise;
expect(mockedWriteFileSync).toHaveBeenCalledWith(
expect.stringContaining('codeman-aicheck-'),
''
expect(mockedWriteFileSync).toHaveBeenCalledWith(expect.stringContaining('codeman-aicheck-'), '');
});
it('should keep Claude stderr separate from verdict output', async () => {
mockedReadFileSync.mockReturnValue('IDLE\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
await checkPromise;
const spawnArgs = mockedSpawn.mock.calls[0]?.[1];
const command = spawnArgs?.[spawnArgs.length - 1];
expect(command).toEqual(expect.any(String));
expect(command).toContain(' 2> "');
expect(command).not.toContain('2>&1');
});
it('should include Claude stderr when no verdict is produced', async () => {
mockedReadFileSync.mockImplementation((path) =>
String(path).includes('-stderr-') ? 'Claude CLI failed to load settings' : '__AICHECK_DONE__'
);
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
const result = await checkPromise;
expect(result.verdict).toBe('ERROR');
expect(result.reasoning).toContain('Claude CLI failed to load settings');
});
});
@@ -223,7 +248,7 @@ describe('AiIdleChecker', () => {
// Should have tried to kill the tmux session (initial kill + cleanup kill)
const killCalls = mockedExecSync.mock.calls.filter(
call => typeof call[0] === 'string' && call[0].includes('kill-session')
(call) => typeof call[0] === 'string' && call[0].includes('kill-session')
);
expect(killCalls.length).toBeGreaterThan(0);
});
@@ -236,8 +261,7 @@ describe('AiIdleChecker', () => {
describe('Cooldown', () => {
it('should start cooldown after WORKING verdict', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nStill processing.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nStill processing.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(500);
@@ -250,8 +274,7 @@ describe('AiIdleChecker', () => {
});
it('should return to ready after cooldown expires', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -267,8 +290,7 @@ describe('AiIdleChecker', () => {
});
it('should not start new check during cooldown', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
const firstCheck = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -283,8 +305,7 @@ describe('AiIdleChecker', () => {
describe('Error Handling', () => {
it('should start error cooldown after parse error', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('garbage output\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage output\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -302,8 +323,7 @@ describe('AiIdleChecker', () => {
const cooldowns = [1100, 2100]; // Wait slightly longer than each cooldown
for (let i = 0; i < 3; i++) {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('garbage\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -321,8 +341,7 @@ describe('AiIdleChecker', () => {
it('should reset error counter on successful check', async () => {
// First check: error
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('garbage\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage\n__AICHECK_DONE__');
const firstCheck = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
await firstCheck;
@@ -332,8 +351,7 @@ describe('AiIdleChecker', () => {
await vi.advanceTimersByTimeAsync(1100);
// Second check: success
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('IDLE\nDone.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\nDone.\n__AICHECK_DONE__');
const secondCheck = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
await secondCheck;
@@ -352,8 +370,7 @@ describe('AiIdleChecker', () => {
describe('Buffer Handling', () => {
it('should strip ANSI codes from terminal buffer', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
const ansiBuffer = '\x1b[1mBold\x1b[0m \x1b[32mGreen\x1b[0m text';
const checkPromise = checker.check(ansiBuffer);
@@ -365,8 +382,7 @@ describe('AiIdleChecker', () => {
});
it('should trim buffer to maxContextChars', async () => {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
// Create buffer longer than maxContextChars (1000)
const longBuffer = 'x'.repeat(2000);
@@ -402,8 +418,7 @@ describe('AiIdleChecker', () => {
describe('Reset', () => {
it('should clear all state on reset', async () => {
// Trigger a WORKING verdict to set state
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -440,24 +455,24 @@ describe('AiIdleChecker', () => {
const handler = vi.fn();
checker.on('checkCompleted', handler);
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('IDLE\nAll done.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\nAll done.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
await checkPromise;
expect(handler).toHaveBeenCalledWith(expect.objectContaining({
verdict: 'IDLE',
}));
expect(handler).toHaveBeenCalledWith(
expect.objectContaining({
verdict: 'IDLE',
})
);
});
it('should emit cooldownStarted event after WORKING', async () => {
const handler = vi.fn();
checker.on('cooldownStarted', handler);
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
@@ -477,8 +492,7 @@ describe('AiIdleChecker', () => {
const cooldowns = [1100, 2100]; // Wait longer than exponential backoff
for (let i = 0; i < 3; i++) {
mockedReadFileSync.mockReturnValueOnce('')
.mockReturnValueOnce('garbage\n__AICHECK_DONE__');
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage\n__AICHECK_DONE__');
const checkPromise = checker.check('output');
await vi.advanceTimersByTimeAsync(1000);
await checkPromise;
+182
View File
@@ -0,0 +1,182 @@
/**
* Unit tests for the pure halves of daemon-control (issue #231): argv rebuilding,
* the readiness URL, pidfile parsing, the stale-pid identity check, and the
* `/api/status` probe against a real socket.
*/
import { describe, it, expect, afterAll, beforeAll } from 'vitest';
import http from 'node:http';
import {
buildBaseUrl,
buildStatusUrl,
buildWebArgs,
isProcessAlive,
looksLikeCodemanWeb,
parsePidFileContents,
probeServer,
} from '../src/daemon-control.js';
const PORT = 3216;
describe('buildWebArgs', () => {
it('always passes host and port through explicitly', () => {
expect(buildWebArgs({ host: '127.0.0.1', port: 3000, https: false })).toEqual([
'web',
'--host',
'127.0.0.1',
'--port',
'3000',
]);
});
it('forwards every optional flag it was given', () => {
const args = buildWebArgs({
host: '0.0.0.0',
port: 8080,
https: true,
titleHostname: 'tower',
allowUnauthenticatedNetwork: true,
multiuser: true,
});
expect(args).toEqual([
'web',
'--host',
'0.0.0.0',
'--port',
'8080',
'--https',
'--title-hostname',
'tower',
'--allow-unauthenticated-network',
'--multiuser',
]);
});
it('never re-emits the daemon flags themselves (the child must not re-fork)', () => {
const args = buildWebArgs({ host: '127.0.0.1', port: 3000, https: false });
expect(args).not.toContain('--daemon');
expect(args).not.toContain('-d');
});
});
describe('buildBaseUrl', () => {
it('is the address a browser can open, with no path on it', () => {
expect(buildBaseUrl({ host: '127.0.0.1', port: 3000, https: false })).toBe('http://127.0.0.1:3000');
expect(buildBaseUrl({ host: '0.0.0.0', port: 8443, https: true })).toBe('https://127.0.0.1:8443');
});
});
describe('buildStatusUrl', () => {
it('uses http by default and https when asked', () => {
expect(buildStatusUrl({ host: '127.0.0.1', port: 3000, https: false })).toBe('http://127.0.0.1:3000/api/status');
expect(buildStatusUrl({ host: '127.0.0.1', port: 3000, https: true })).toBe('https://127.0.0.1:3000/api/status');
});
it('rewrites wildcard binds to loopback, since they are not connectable', () => {
expect(buildStatusUrl({ host: '0.0.0.0', port: 3000, https: false })).toBe('http://127.0.0.1:3000/api/status');
expect(buildStatusUrl({ host: '::', port: 3000, https: false })).toBe('http://127.0.0.1:3000/api/status');
});
it('brackets a bare IPv6 literal', () => {
expect(buildStatusUrl({ host: '::1', port: 3000, https: false })).toBe('http://[::1]:3000/api/status');
expect(buildStatusUrl({ host: '[::1]', port: 3000, https: false })).toBe('http://[::1]:3000/api/status');
});
});
describe('parsePidFileContents', () => {
it('accepts a plain pid with surrounding whitespace', () => {
expect(parsePidFileContents('4242\n')).toBe(4242);
expect(parsePidFileContents(' 4242 ')).toBe(4242);
});
it('rejects garbage, empties and floats', () => {
expect(parsePidFileContents('')).toBeNull();
expect(parsePidFileContents('not a pid')).toBeNull();
expect(parsePidFileContents('42.5')).toBeNull();
expect(parsePidFileContents('-42')).toBeNull();
});
it('rejects pid 0 and pid 1: neither is ever our server', () => {
expect(parsePidFileContents('0')).toBeNull();
expect(parsePidFileContents('1')).toBeNull();
});
});
describe('looksLikeCodemanWeb', () => {
it('matches the ways the server is actually launched', () => {
expect(looksLikeCodemanWeb('/usr/bin/node /home/u/.codeman/app/dist/index.js web')).toBe(true);
expect(looksLikeCodemanWeb('/usr/bin/node dist/index.js web --https')).toBe(true);
expect(looksLikeCodemanWeb('node /repo/src/index.ts web --port 3000')).toBe(true);
expect(looksLikeCodemanWeb('/opt/homebrew/bin/codeman web')).toBe(true);
expect(looksLikeCodemanWeb('aicodeman web --host 0.0.0.0')).toBe(true);
});
it('rejects anything that inherited a recycled pid', () => {
expect(looksLikeCodemanWeb(null)).toBe(false);
expect(looksLikeCodemanWeb('')).toBe(false);
expect(looksLikeCodemanWeb('/usr/bin/node dist/index.js session list')).toBe(false);
expect(looksLikeCodemanWeb('vim web')).toBe(false);
expect(looksLikeCodemanWeb('/usr/lib/systemd/systemd --user')).toBe(false);
});
});
describe('isProcessAlive', () => {
it('sees this very process', () => {
expect(isProcessAlive(process.pid)).toBe(true);
});
it('does not see an unused high pid', () => {
// 2^22 is above the default pid_max on Linux and macOS.
expect(isProcessAlive(4_194_303)).toBe(false);
});
});
describe('probeServer', () => {
let server: http.Server;
beforeAll(async () => {
server = http.createServer((req, res) => {
if (req.url === '/unauthorized') {
res.writeHead(401).end('Unauthorized');
return;
}
if (req.url === '/foreign') {
res.writeHead(200, { 'Content-Type': 'text/html' }).end('<html>some other app</html>');
return;
}
res.writeHead(200, { 'Content-Type': 'application/json' });
res.end(JSON.stringify({ success: true, data: { version: '9.9.9' } }));
});
await new Promise<void>((resolve) => server.listen(PORT, '127.0.0.1', resolve));
});
afterAll(async () => {
await new Promise<void>((resolve) => server.close(() => resolve()));
});
it('reports up and reads the version back', async () => {
const result = await probeServer(`http://127.0.0.1:${PORT}/api/status`);
expect(result.up).toBe(true);
expect(result.version).toBe('9.9.9');
});
it('counts a 401 as up, because auth being active proves a server is there', async () => {
const result = await probeServer(`http://127.0.0.1:${PORT}/unauthorized`);
expect(result.up).toBe(true);
});
it('does not mistake an unrelated service squatting on the port for Codeman', async () => {
const result = await probeServer(`http://127.0.0.1:${PORT}/foreign`);
expect(result.up).toBe(false);
});
it('reports down when nothing is listening', async () => {
const result = await probeServer(`http://127.0.0.1:${PORT + 1}/api/status`, 1000);
expect(result.up).toBe(false);
});
it('reports down for a malformed url instead of throwing', async () => {
const result = await probeServer('not-a-url');
expect(result.up).toBe(false);
});
});
+1 -1
View File
@@ -127,7 +127,7 @@ describe('refreshStaleCodemanHooks', () => {
const after = JSON.parse(readFileSync(settingsPath, 'utf-8'));
expect(JSON.stringify(after.hooks)).toContain(SECRET_HEADER);
expect(JSON.stringify(after.hooks)).toContain('CODEMAN_BACKGROUND_REWAKE_V');
expect(JSON.stringify(after.hooks)).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(JSON.stringify(after.hooks.Stop)).toContain('./notify-user.sh');
expect(after.hooks.PostToolUse).toEqual(expect.arrayContaining([customPostToolUse]));
expect(after.hooks.CustomEvent).toEqual(customEvent);
+272 -3
View File
@@ -6,13 +6,15 @@
*/
import { describe, it, expect, beforeAll, beforeEach, afterAll, afterEach } from 'vitest';
import { existsSync, readFileSync, writeFileSync, mkdirSync, rmSync } from 'node:fs';
import { closeSync, existsSync, openSync, readFileSync, writeFileSync, mkdirSync, rmSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import { spawn } from 'node:child_process';
import {
ensureCodemanHooks,
generateBackgroundWakeScript,
generateHooksConfig,
generateSubagentStopGuardScript,
refreshStaleCodemanHooks,
writeHooksConfig,
} from '../src/hooks-config.js';
@@ -35,6 +37,20 @@ describe('generateHooksConfig', () => {
expect(config.hooks.Stop).toHaveLength(1);
});
it('should guard subagent stops while their background work is active', () => {
const config = generateHooksConfig();
const subagentHooks = config.hooks.SubagentStop as Array<{
hooks: Array<{ type: string; command: string; args: string[]; timeout: number }>;
}>;
expect(subagentHooks).toHaveLength(1);
expect(subagentHooks[0].hooks[0]).toMatchObject({
type: 'command',
command: 'node',
args: ['-e', generateSubagentStopGuardScript()],
});
});
it('should configure a self-contained Bash background-task rewake hook', () => {
const config = generateHooksConfig();
const postToolHooks = config.hooks.PostToolUse as Array<{
@@ -210,7 +226,8 @@ describe('writeHooksConfig', () => {
const parsed = JSON.parse(readFileSync(settingsPath, 'utf-8'));
expect(parsed.hooks.PostToolUse).toHaveLength(1);
expect(JSON.stringify(parsed.hooks.PostToolUse)).toContain('CODEMAN_BACKGROUND_REWAKE_V');
expect(JSON.stringify(parsed.hooks.PostToolUse)).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(JSON.stringify(parsed.hooks.SubagentStop)).toContain('CODEMAN_SUBAGENT_STOP_GUARD_V1');
});
it('should replace an older rewake script version without duplicating it', async () => {
@@ -242,10 +259,29 @@ describe('writeHooksConfig', () => {
const serialized = JSON.stringify(parsed.hooks.PostToolUse);
expect(parsed.hooks.PostToolUse).toHaveLength(1);
expect(parsed.hooks.PostToolUse[0].hooks).toHaveLength(1);
expect(serialized).toContain('CODEMAN_BACKGROUND_REWAKE_V2');
expect(serialized).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(serialized).not.toContain('CODEMAN_BACKGROUND_REWAKE_V1');
});
it('replaces the V2 background hook without duplicating it', async () => {
const claudeDir = join(testDir, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
mkdirSync(claudeDir, { recursive: true });
const oldSettings = JSON.stringify({ hooks: generateHooksConfig().hooks }, null, 2).replaceAll(
'CODEMAN_BACKGROUND_REWAKE_V3',
'CODEMAN_BACKGROUND_REWAKE_V2'
);
writeFileSync(settingsPath, oldSettings);
await refreshStaleCodemanHooks(testDir);
const parsed = JSON.parse(readFileSync(settingsPath, 'utf-8'));
const postToolUse = JSON.stringify(parsed.hooks.PostToolUse);
expect(parsed.hooks.PostToolUse).toHaveLength(1);
expect(postToolUse).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(postToolUse).not.toContain('CODEMAN_BACKGROUND_REWAKE_V2');
});
it('should not add rewake hooks to a user-owned hook configuration', async () => {
const claudeDir = join(testDir, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
@@ -278,6 +314,35 @@ describe('writeHooksConfig', () => {
expect(parsed.hooks.Notification).toBeDefined();
});
it('should safely add Codeman hooks to an existing managed-case settings file', async () => {
const claudeDir = join(testDir, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
mkdirSync(claudeDir, { recursive: true });
const userHooks = {
PostToolUse: [{ matcher: 'Write', hooks: [{ type: 'command', command: './format.sh' }] }],
};
writeFileSync(settingsPath, JSON.stringify({ hooks: userHooks, permissions: { allow: ['Read'] } }, null, 2));
await ensureCodemanHooks(testDir);
const parsed = JSON.parse(readFileSync(settingsPath, 'utf-8'));
expect(parsed.permissions).toEqual({ allow: ['Read'] });
expect(parsed.hooks.PostToolUse).toEqual(expect.arrayContaining(userHooks.PostToolUse));
expect(JSON.stringify(parsed.hooks)).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
expect(JSON.stringify(parsed.hooks)).toContain('CODEMAN_SUBAGENT_STOP_GUARD_V1');
});
it('should not replace a malformed managed-case settings file', async () => {
const claudeDir = join(testDir, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
mkdirSync(claudeDir, { recursive: true });
writeFileSync(settingsPath, '{ malformed');
await ensureCodemanHooks(testDir);
expect(readFileSync(settingsPath, 'utf-8')).toBe('{ malformed');
});
it('should handle malformed existing settings.local.json', async () => {
const claudeDir = join(testDir, '.claude');
mkdirSync(claudeDir, { recursive: true });
@@ -370,6 +435,210 @@ describe('background task rewake helper', () => {
expect(result.stderr).toContain('completed');
expect(result.stderr).toContain('/tmp/bg-test-1.output');
});
it('rewakes a subagent when Claude queues completion in the parent transcript', async () => {
const sessionId = '7148e9de-7673-48b8-bf38-6799e52c346a';
const sessionDir = join(testDir, sessionId);
const subagentDir = join(sessionDir, 'subagents');
const parentTranscriptPath = `${sessionDir}.jsonl`;
const subagentTranscriptPath = join(subagentDir, 'agent-afacts-class2.jsonl');
mkdirSync(subagentDir, { recursive: true });
writeFileSync(parentTranscriptPath, '');
writeFileSync(subagentTranscriptPath, '');
const resultPromise = runHelper({
session_id: sessionId,
agent_id: 'afacts-class2',
transcript_path: subagentTranscriptPath,
tool_response: {
backgroundTaskId: 'bg-subagent-1',
},
});
await new Promise((resolve) => setTimeout(resolve, 100));
writeFileSync(
parentTranscriptPath,
JSON.stringify({
type: 'queue-operation',
operation: 'enqueue',
content:
'<task-notification>\n<task-id>bg-subagent-1</task-id>\n<status>completed</status>\n' +
'<output-file>/tmp/bg-subagent-1.output</output-file>\n</task-notification>',
}) + '\n'
);
const result = await resultPromise;
expect(result.code).toBe(2);
expect(result.stderr).toContain('bg-subagent-1');
expect(result.stderr).toContain('/tmp/bg-subagent-1.output');
});
it('includes a marked background report in the wake feedback', async () => {
const transcriptPath = join(testDir, 'transcript.jsonl');
const tasksDir = join(testDir, 'tasks');
const outputPath = join(tasksDir, 'bg-report-1.output');
mkdirSync(tasksDir, { recursive: true });
writeFileSync(transcriptPath, '');
writeFileSync(
outputPath,
[
'launcher output',
'=== CODEMAN_RESULT_BEGIN ===',
'Summary line',
'Detail after the old 30-line preview boundary',
'=== CODEMAN_RESULT_END ===',
].join('\n')
);
const resultPromise = runHelper({
transcript_path: transcriptPath,
tool_response: {
stdout: `Command running in background with ID: bg-report-1. Output is being written to: ${outputPath}.`,
},
});
await new Promise((resolve) => setTimeout(resolve, 100));
writeFileSync(
transcriptPath,
JSON.stringify({
type: 'queue-operation',
operation: 'enqueue',
content:
'<task-notification>\n<task-id>bg-report-1</task-id>\n<status>completed</status>\n' +
`<output-file>${outputPath}</output-file>\n</task-notification>`,
}) + '\n'
);
const result = await resultPromise;
expect(result.code).toBe(2);
expect(result.stderr).toContain('<codeman-background-result>');
expect(result.stderr).toContain('Summary line');
expect(result.stderr).toContain('Detail after the old 30-line preview boundary');
});
});
describe('subagent stop guard helper', () => {
const testDir = join(tmpdir(), 'codeman-subagent-stop-guard-test-' + Date.now());
beforeEach(() => {
mkdirSync(testDir, { recursive: true });
});
afterEach(() => {
rmSync(testDir, { recursive: true, force: true });
});
function runGuard(transcriptLines: unknown[]): Promise<{ code: number | null; stdout: string; stderr: string }> {
const transcriptPath = join(testDir, 'agent-test.jsonl');
writeFileSync(transcriptPath, transcriptLines.map((line) => JSON.stringify(line)).join('\n') + '\n');
return new Promise((resolve, reject) => {
const child = spawn(process.execPath, ['-e', generateSubagentStopGuardScript()], {
stdio: ['pipe', 'pipe', 'pipe'],
});
let stdout = '';
let stderr = '';
child.stdout.setEncoding('utf8');
child.stderr.setEncoding('utf8');
child.stdout.on('data', (chunk) => {
stdout += chunk;
});
child.stderr.on('data', (chunk) => {
stderr += chunk;
});
child.on('error', reject);
child.on('close', (code) => resolve({ code, stdout, stderr }));
child.stdin.end(JSON.stringify({ agent_transcript_path: transcriptPath }));
});
}
async function withLiveTask<T>(taskId: string, action: () => Promise<T>): Promise<T> {
const tasksDir = join(testDir, 'tasks');
mkdirSync(tasksDir, { recursive: true });
const outputFd = openSync(join(tasksDir, `${taskId}.output`), 'a');
const child = spawn(process.execPath, ['-e', 'setTimeout(() => {}, 10000)'], {
stdio: ['ignore', outputFd, outputFd],
});
await new Promise<void>((resolve, reject) => {
child.once('spawn', resolve);
child.once('error', reject);
});
closeSync(outputFd);
try {
return await action();
} finally {
const closed = new Promise<void>((resolve) => child.once('close', () => resolve()));
child.kill();
await closed;
}
}
const monitorResult = (taskId: string) => ({
type: 'user',
message: {
content: [
{
type: 'tool_result',
content: `Monitor started (task ${taskId}, pid 123).`,
},
],
},
});
const completion = (taskId: string) => ({
type: 'user',
message: {
content:
`<task-notification>\n<task-id>${taskId}</task-id>\n` + '<status>completed</status>\n</task-notification>',
},
});
it('blocks an intermediate subagent stop while a sibling monitor is active', async () => {
const result = await withLiveTask('monitor-still-live', () =>
runGuard([monitorResult('monitor-first'), monitorResult('monitor-still-live'), completion('monitor-first')])
);
expect(result.code).toBe(0);
expect(result.stderr).toBe('');
expect(JSON.parse(result.stdout)).toMatchObject({ decision: 'block' });
expect(result.stdout).toContain('monitor-still-live');
expect(result.stdout).not.toContain('monitor-first,');
});
it('allows a subagent to stop after all of its monitored work finishes', async () => {
const result = await runGuard([
monitorResult('monitor-first'),
monitorResult('monitor-second'),
completion('monitor-first'),
completion('monitor-second'),
]);
expect(result.code).toBe(0);
expect(result.stdout).toBe('');
expect(result.stderr).toBe('');
});
it('also recognizes background Bash task ownership', async () => {
const result = await withLiveTask('bash-live-1', () =>
runGuard([
{
type: 'user',
message: {
content: [
{
type: 'tool_result',
content: 'Command running in background with ID: bash-live-1. Output is being written to a task file.',
},
],
},
},
])
);
expect(JSON.parse(result.stdout)).toMatchObject({ decision: 'block' });
expect(result.stdout).toContain('bash-live-1');
});
});
// ========== Hook Event API Integration Tests ==========
+1
View File
@@ -86,6 +86,7 @@ export function createMockRouteContext(options?: { sessionId?: string }) {
getModelConfig: vi.fn(async () => null),
getClaudeModeConfig: vi.fn(async () => ({})),
getTerminalHistoryConfig: vi.fn(async () => resolveTerminalHistoryConfig({})),
getAgentSkillEnabled: vi.fn(async () => false),
getDefaultClaudeMdPath: vi.fn(async () => undefined),
getLightState: vi.fn(() => ({ sessions: [], status: 'ok' })),
getLightSessionsState: vi.fn(() => {
+80
View File
@@ -332,3 +332,83 @@ describe('Case Management', () => {
});
});
});
describe('Agent skill injection (agentSkillEnabled)', () => {
let server: WebServer;
let baseUrl: string;
const createdCases: string[] = [];
beforeAll(async () => {
server = await createTestServer(TEST_PORT + 4); // 3103
await server.start();
baseUrl = `http://localhost:${TEST_PORT + 4}`;
});
afterAll(async () => {
await server.stop();
for (const caseName of createdCases) {
const casePath = join(CASES_DIR, caseName);
if (existsSync(casePath)) {
rmSync(casePath, { recursive: true, force: true });
}
}
});
it('does not inject by default, accepts the setting via PUT, then injects on quick-start', async () => {
// 1. Default OFF: a claude quick-start creates the case without the skill.
const offCase = 'test-skill-off-' + Date.now();
createdCases.push(offCase);
const offResponse = await fetch(`${baseUrl}/api/quick-start`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ caseName: offCase }),
});
const offData = await offResponse.json();
expect(offData.success).toBe(true);
expect(existsSync(join(CASES_DIR, offCase, '.claude', 'skills', 'codeman'))).toBe(false);
// 2. The `.strict()` settings schema accepts the new synced key.
const putResponse = await fetch(`${baseUrl}/api/settings`, {
method: 'PUT',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ agentSkillEnabled: true }),
});
const putData = await putResponse.json();
expect(putData.success).toBe(true);
// 3. The server's settings read is cached ~2s; outwait it so the create sees the toggle.
await new Promise((resolve) => setTimeout(resolve, 2100));
// 4. Quick-start now injects the marker-carrying skill into the new case.
const onCase = 'test-skill-on-' + Date.now();
createdCases.push(onCase);
const onResponse = await fetch(`${baseUrl}/api/quick-start`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ caseName: onCase }),
});
const onData = await onResponse.json();
expect(onData.success).toBe(true);
const skillDir = join(CASES_DIR, onCase, '.claude', 'skills', 'codeman');
const { readFileSync } = await import('node:fs');
const skillMd = readFileSync(join(skillDir, 'SKILL.md'), 'utf-8');
expect(skillMd.startsWith('---\nname: codeman')).toBe(true);
expect(skillMd).toContain('<!-- codeman-managed-agent-skill');
expect(existsSync(join(skillDir, 'reference', 'endpoints.md'))).toBe(true);
expect(existsSync(join(skillDir, 'reference', 'recipes.md'))).toBe(true);
}, 30000);
it('does not inject for shell-mode quick-start even when enabled', async () => {
const shellCase = 'test-skill-shell-' + Date.now();
createdCases.push(shellCase);
const response = await fetch(`${baseUrl}/api/quick-start`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ caseName: shellCase, mode: 'shell' }),
});
const data = await response.json();
expect(data.success).toBe(true);
expect(existsSync(join(CASES_DIR, shellCase, '.claude', 'skills', 'codeman'))).toBe(false);
});
});
+179
View File
@@ -0,0 +1,179 @@
/**
* Unit tests for the unit-file builders behind `codeman service install`
* (issue #231). These are the parts that must be right without launchctl or
* systemctl in the loop: PATH construction, escaping, and the file contents.
*/
import { describe, it, expect } from 'vitest';
import {
buildLaunchAgentPlist,
buildServiceEnv,
buildServicePath,
buildSystemdUnit,
detectServiceKind,
systemdQuote,
xmlEscape,
type ServicePlan,
} from '../src/service-installer.js';
function plan(overrides: Partial<ServicePlan> = {}): ServicePlan {
return {
kind: 'systemd',
name: 'codeman-web.service',
nodePath: '/usr/bin/node',
execArgv: [],
scriptPath: '/home/u/.codeman/app/dist/index.js',
args: ['web', '--host', '127.0.0.1', '--port', '3000'],
env: { PATH: '/usr/bin:/bin', HOME: '/home/u', LANG: 'en_US.UTF-8' },
logPath: '/home/u/.codeman/web.log',
workingDir: '/home/u',
...overrides,
};
}
describe('buildServicePath', () => {
it("puts the running node's directory first so nvm/homebrew node wins", () => {
const result = buildServicePath('/home/u/.nvm/versions/node/v22.0.0/bin', '/usr/bin:/bin', '/home/u');
expect(result.split(':')[0]).toBe('/home/u/.nvm/versions/node/v22.0.0/bin');
});
it('keeps the installing shell PATH, which is the whole point of the fix', () => {
const result = buildServicePath('/usr/bin', '/opt/homebrew/bin:/home/u/.bun/bin', '/home/u');
expect(result.split(':')).toContain('/home/u/.bun/bin');
expect(result.split(':')).toContain('/opt/homebrew/bin');
});
it('appends the fallbacks a bare launchd PATH would otherwise be missing', () => {
const entries = buildServicePath('/usr/bin', '/usr/bin', '/home/u').split(':');
expect(entries).toContain('/opt/homebrew/bin');
expect(entries).toContain('/home/u/.local/bin');
expect(entries).toContain('/usr/local/bin');
});
it('never repeats a directory', () => {
const entries = buildServicePath('/usr/bin', '/usr/bin:/bin:/usr/bin', '/home/u').split(':');
expect(new Set(entries).size).toBe(entries.length);
});
it('drops empty segments from a trailing-colon PATH', () => {
expect(buildServicePath('/usr/bin', '/usr/bin::/bin:', '/home/u').split(':')).not.toContain('');
});
it('drops node_modules/.bin, which npx injects for one command only', () => {
const entries = buildServicePath(
'/usr/bin',
'/repo/node_modules/.bin:/repo/node_modules/.bin/:/home/u/bin',
'/home/u'
).split(':');
expect(entries.filter((e) => e.includes('node_modules'))).toEqual([]);
expect(entries).toContain('/home/u/bin');
});
});
describe('buildServiceEnv', () => {
it('carries PATH, HOME and a LANG default', () => {
const env = buildServiceEnv('/usr/bin', '/usr/bin:/bin', '/home/u');
expect(env.HOME).toBe('/home/u');
expect(env.LANG).toBe('en_US.UTF-8');
expect(env.PATH).toContain('/usr/bin');
});
it('prefers the caller LANG when there is one', () => {
expect(buildServiceEnv('/usr/bin', '/usr/bin', '/home/u', 'de_DE.UTF-8').LANG).toBe('de_DE.UTF-8');
});
it('does not carry a password into the unit file', () => {
const env = buildServiceEnv('/usr/bin', '/usr/bin', '/home/u');
expect(Object.keys(env)).not.toContain('CODEMAN_PASSWORD');
});
});
describe('escaping', () => {
it('escapes the five XML entities', () => {
expect(xmlEscape(`a&b<c>d"e'f`)).toBe('a&amp;b&lt;c&gt;d&quot;e&apos;f');
});
it('quotes systemd values and escapes quotes and backslashes', () => {
expect(systemdQuote('plain')).toBe('"plain"');
expect(systemdQuote('with "quotes"')).toBe('"with \\"quotes\\""');
expect(systemdQuote('back\\slash')).toBe('"back\\\\slash"');
});
});
describe('buildLaunchAgentPlist', () => {
it('writes the label, the full command and the log paths', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', name: 'com.codeman.web' }));
expect(xml).toContain('<string>com.codeman.web</string>');
expect(xml).toContain('<string>/usr/bin/node</string>');
expect(xml).toContain('<string>/home/u/.codeman/app/dist/index.js</string>');
expect(xml).toContain('<string>web</string>');
expect(xml).toContain('<string>/home/u/.codeman/web.log</string>');
});
it('keeps the argument order: node, script, then the web args', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', name: 'com.codeman.web' }));
// Match whole <string> elements: the label itself contains the word "web".
const order = [
'<string>/usr/bin/node</string>',
'<string>/home/u/.codeman/app/dist/index.js</string>',
'<string>web</string>',
'<string>--port</string>',
].map((s) => xml.indexOf(s));
expect(order).toEqual([...order].sort((a, b) => a - b));
expect(order.every((i) => i > -1)).toBe(true);
});
it('carries the runner flags so a tsx dev install still boots', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', execArgv: ['--import', 'tsx'] }));
expect(xml).toContain('<string>--import</string>');
expect(xml).toContain('<string>tsx</string>');
});
it('restarts on crash and at login', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd' }));
expect(xml).toContain('<key>KeepAlive</key>');
expect(xml).toContain('<key>RunAtLoad</key>');
});
it('escapes a path with an ampersand instead of emitting broken XML', () => {
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', workingDir: '/Users/a&b' }));
expect(xml).toContain('<string>/Users/a&amp;b</string>');
expect(xml).not.toContain('<string>/Users/a&b</string>');
});
});
describe('buildSystemdUnit', () => {
it('builds ExecStart from node, script and args', () => {
expect(buildSystemdUnit(plan())).toContain(
'ExecStart=/usr/bin/node /home/u/.codeman/app/dist/index.js web --host 127.0.0.1 --port 3000'
);
});
it('quotes an argument containing spaces', () => {
const unit = buildSystemdUnit(plan({ scriptPath: '/home/my user/app/dist/index.js' }));
expect(unit).toContain('"/home/my user/app/dist/index.js"');
});
it('writes each env var as a quoted Environment line', () => {
const unit = buildSystemdUnit(plan());
expect(unit).toContain('Environment="PATH=/usr/bin:/bin"');
expect(unit).toContain('Environment="HOME=/home/u"');
});
it('keeps KillMode=process so agents survive a server restart', () => {
expect(buildSystemdUnit(plan())).toContain('KillMode=process');
});
it('is installable and restarts on failure', () => {
const unit = buildSystemdUnit(plan());
expect(unit).toContain('Restart=always');
expect(unit).toContain('WantedBy=default.target');
});
});
describe('detectServiceKind', () => {
it('maps the platform to its supervisor', () => {
const expected = process.platform === 'darwin' ? 'launchd' : process.platform === 'linux' ? 'systemd' : null;
expect(detectServiceKind()).toBe(expected);
});
});
+20
View File
@@ -47,6 +47,26 @@ function loadTerminalUiHarness(mode: string) {
}
describe('terminal flush budget', () => {
it('drains a large final batch without waiting for unrelated terminal output', () => {
const { app, writes } = loadTerminalUiHarness('codex');
const scheduled: Array<() => void> = [];
app._safeYield = (callback: () => void) => {
scheduled.push(callback);
};
app.isTerminalAtBottom = () => true;
app.batchTerminalWrite('x'.repeat(96 * 1024));
expect(scheduled).toHaveLength(1);
while (scheduled.length > 0) {
scheduled.shift()?.();
}
expect(writes.map((write) => write.length)).toEqual([32 * 1024, 32 * 1024, 32 * 1024]);
expect(app.pendingWrites).toEqual([]);
expect(app.writeFrameScheduled).toBe(false);
});
it('uses a smaller first-frame write budget for Codex output to reduce renderer stalls', () => {
const { app, writes } = loadTerminalUiHarness('codex');
app.pendingWrites.push('x'.repeat(96 * 1024));
+9 -3
View File
@@ -379,7 +379,7 @@ describe('terminal touch tap mouse guard', () => {
expect(withVersion('garbage')).toBe(false); // unparseable → assume older
});
it('wheel: codex forwards without a version; gemini never forwards', () => {
it('wheel: only claude forwards — codex and gemini keep the local wheel', () => {
const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.terminal = {
@@ -387,8 +387,14 @@ describe('terminal touch tap mouse guard', () => {
buffer: { active: { viewportY: 50, baseY: 50 } },
};
app.sessions = new Map([['sess-1', { mode: 'codex' }]]); // verified TUI, no version gate
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(true);
// Codex used to forward unconditionally, which is PR #227's regression: measured
// on codex-cli 0.147.0, it never enables mouse tracking and ignores SGR wheel
// reports outright, so forwarding ate every tick while its real local scrollback
// (the codex transcript lives there — inline viewport, no in-app pager) sat unused.
app.sessions = new Map([['sess-1', { mode: 'codex' }]]);
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
app.sessions = new Map([['sess-1', { mode: 'codex', cliVersion: '9.9.9' }]]); // no version rescues it
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
app.sessions = new Map([['sess-1', { mode: 'gemini', cliVersion: '9.9.9' }]]); // unverified TUI
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);