mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 20:49:41 +02:00
Compare commits
50
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
40b4aba043 | ||
|
|
4b44988bfc | ||
|
|
316d0a4c82 | ||
|
|
1184720648 | ||
|
|
b067aad9b6 | ||
|
|
19a3d7c773 | ||
|
|
085f4acb60 | ||
|
|
d26f26fe34 | ||
|
|
091df2b6d8 | ||
|
|
5f775b1ab1 | ||
|
|
da51193264 | ||
|
|
fa18eeef35 | ||
|
|
a9f26bd03a | ||
|
|
8dc8b164a7 | ||
|
|
2524759655 | ||
|
|
1f164bc8d2 | ||
|
|
0a1439b1e9 | ||
|
|
66abe6c70a | ||
|
|
fa1700da5b | ||
|
|
aed1e59ee3 | ||
|
|
7f6d18b398 | ||
|
|
52571c7fd4 | ||
|
|
cb95a8562c | ||
|
|
3cb7e30636 | ||
|
|
9dc4620f03 | ||
|
|
9d27cc0bab | ||
|
|
84132d3025 | ||
|
|
cc163792e5 | ||
|
|
6f1ff17ccc | ||
|
|
f262b8cb69 | ||
|
|
5f2b491d99 | ||
|
|
c067167dbc | ||
|
|
ad2ca9b575 | ||
|
|
a7a1cef3d6 | ||
|
|
a1d7ec02e9 | ||
|
|
dfa43928af | ||
|
|
adbb74cd5a | ||
|
|
eb8d11ffc3 | ||
|
|
ebfcac6ad1 | ||
|
|
2e69e28e71 | ||
|
|
1a32e63765 | ||
|
|
d41f28bc14 | ||
|
|
322f21ef9f | ||
|
|
c2d973cb2d | ||
|
|
84e31c0ee1 | ||
|
|
0d0b772619 | ||
|
|
bfff20a093 | ||
|
|
b982c5d0e0 | ||
|
|
f50c922240 | ||
|
|
de5b048c3f |
+157
@@ -1,5 +1,162 @@
|
||||
# aicodeman
|
||||
|
||||
## 1.14.0
|
||||
|
||||
### Minor Changes
|
||||
|
||||
- Daemon mode and service install, plus subagent hook hardening and terminal/idle-checker fixes.
|
||||
|
||||
**New: run Codeman in the background without a terminal (#239, closes #231)**
|
||||
- `codeman web -d` starts the server detached: it survives closing the shell, logs to `~/.codeman/web.log`, records a pidfile, and only reports success after the server actually answers `/api/status` (a port clash or missing dependency can never read as a clean start). `codeman web --status` and `codeman web --stop` manage it; `--stop` verifies the pid still looks like a Codeman server before signalling, so a recycled pid is never SIGTERMed.
|
||||
- `codeman service install` / `status` / `uninstall`: installs a systemd user unit (Linux) or LaunchAgent (macOS) so the server comes back after reboots. The unit carries the installing shell's PATH (launchd's default PATH finds neither an nvm/Homebrew `node` nor `tmux`/`claude`), never contains `CODEMAN_PASSWORD`, and uses the same instance-scoped unit names as `install.sh` and the self-updater so no second copy can end up supervised.
|
||||
- Both refuse to start a second server on one data dir (pidfile check plus a live probe): two servers on the shared tmux socket would attach to each other's sessions.
|
||||
- Why `-d` exists at all: `nohup` does not protect a Node process, Node re-arms SIGHUP even when it inherits "ignore", so `nohup codeman web &` still dies on HUP. The detached relaunch (setsid) removes the controlling terminal instead.
|
||||
|
||||
**Subagent background-work hooks (#233, thanks @Lint111)**
|
||||
- The background Bash rewake helper now also watches the top-level parent transcript when the hook fires inside a subagent: Claude records a subagent's Bash result in its own `subagents/agent-*.jsonl` but queues the completion in the lead session transcript, so subagents previously never woke. It can also inline a `CODEMAN_RESULT_BEGIN/END` marked report (up to 64 KiB) from the task output file into the wake feedback.
|
||||
- New SubagentStop guard: a subagent that still owns live Monitor or background Bash processes is kept working instead of publishing an intermediate progress line as its final report. Ownership is verified against live process descriptors on `tasks/<id>.output`, so stale transcript text alone never blocks, and the guard fails open on systems without `/proc`.
|
||||
- Existing cases self-heal to the new hooks on next launch.
|
||||
|
||||
**AI idle checker: stderr kept out of the verdict (#234, thanks @Lint111)**
|
||||
|
||||
The `claude -p` verdict command no longer merges stderr into the verdict file, where CLI warnings could turn a valid verdict into a parse error. On failures, the first 200 chars of stderr are attached to the diagnostic instead.
|
||||
|
||||
**Terminal: large final batches drain fully (#235, thanks @Lint111)**
|
||||
|
||||
A render-scheduling flag was cleared after the flush instead of before it, so when a large batch left a remainder behind, the remainder stayed unrendered until unrelated output arrived. This looked like truncated responses or shell commands that never finish. The flush now reschedules itself until the queue is empty.
|
||||
|
||||
**Docs and tests**
|
||||
- README documents daemon mode and service install.
|
||||
- Unique test port for the daemon-control suite.
|
||||
|
||||
## 1.13.0
|
||||
|
||||
### Minor Changes
|
||||
|
||||
- Agent wait primitives, the Codeman agent skill, a fix for hooks dying silently on HTTPS installs, and the tab-strip UX improvements from the previous batch.
|
||||
|
||||
**Agent wait primitives (new API surface, the reason this is a minor).** Three bounded long-polls let an agent driving Codeman from a shell block instead of poll:
|
||||
- `GET /api/v1/sessions/:id/wait` blocks until a lifecycle signal fires (`until=stop,idle,working,blocked,exit`, `fresh=1` to require a new transition).
|
||||
- `GET /api/v1/sessions/:id/wait-output` blocks until a literal substring appears in the session's output (`match=`, `nocase=`, `from=now|buffer`; never regex, by design).
|
||||
- `wait`/`waitTimeout` on `POST /api/v1/sessions/:id/input` (send-and-wait) registers the waiter before typing, closing the race where a separate wait reports the previous turn's idle state as this turn's answer.
|
||||
|
||||
Shared semantics: a timeout is HTTP 200 with `wait.timedOut: true` (callers loop over short waits; tunnels cut idle connections), timeouts are clamped to [1s, 600s] and echoed back as `wait.timeoutMs`, all three nest the result under `data.wait`, and `status`/`limitPaused` ride along. `stop`/`blocked` exist for `claude` mode only: requesting them explicitly elsewhere is a 400, the default set silently narrows and echoes what it waited on. Capacity caps (16 waiters per session, 128 process-wide) answer 409/429, waiter slots release on client hang-up, and shutdown resolves parked waiters instead of stranding them. Bounds are operator-tunable via `CODEMAN_WAIT_*` env vars.
|
||||
|
||||
Reliability details that came out of three verification rounds: a worker that dies inside its tmux pane is now detected at the mux layer (pane-death probe, ~750ms cache, a 3s watcher for waits already parked), so a corpse answers `exit` instead of `idle` and send-and-wait rolls back its dedup seq when the write went nowhere; output matching normalizes charset-designation escapes (a stock bash prompt's `ESC ( B` no longer breaks `match=tnode:`) and holds back partial escapes at chunk boundaries, so matches straddling PTY chunks are found.
|
||||
|
||||
**Codeman agent skill (`skills/codeman`).** A packaged skill that teaches an agent running inside a Codeman session to drive the API safely: guard preamble (refuses outside `CODEMAN_MUX=1`, resolves credentials from the data dir `.env` or the install's service definition), self-protection (`is_self` prefix check in both directions), readiness for claude workers (composer-first, trust dialog as bounded fallback), send-and-wait loops that cannot report a never-submitted prompt as success, marker-synchronized shell flows, fan-out patterns, and cleanup discipline. Ships in the npm package via the `files` entry.
|
||||
|
||||
**Hooks were dying silently on every HTTPS install (bug fix).** The generated hook curls lacked `-k`, so on `--https` installs (self-signed cert) every hook event (`stop`, `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `teammate_idle`, `task_completed`) failed TLS verification and the failure was swallowed, taking respawn's definitive idle signals with it. Hooks are now generated with `curl -sk`, and a staleness detector regenerates the on-disk hook config of already-created cases the next time a session starts in them. Relatedly, `CODEMAN_API_URL` is no longer exported with a guessed `http://localhost:3000` fallback (wrong scheme on HTTPS installs); it is omitted unless the server has stamped the real URL, so in-session guards fail closed.
|
||||
|
||||
**Tab strip (from the previous batch, reported by christianhaberl):** action icons (kill/pop-out) now appear on the active tab only, middle-click closes a tab, tab hover uses a fixed width with a sliding title instead of resizing the strip, and the pop-out button is opt-in (default off).
|
||||
|
||||
**Docs.** `docs/api-reference.md` gained the full long-polling contract (signals by mode, readiness, what the matcher sees, response discriminators); `docs/extending-codeman.md` and the README carry verified copy-paste orchestration recipes; `docs/architecture-invariants.md` records the load-bearing ordering, liveness, and edge-triggered-signal invariants. Net +163 tests (4300 passing in the CI sweep).
|
||||
|
||||
## 1.12.2
|
||||
|
||||
### Patch Changes
|
||||
|
||||
- Codex input fixes: all four bugs reported by @DodgyBadger traced to one root cause (the zero-lag local-echo overlay buffering keystrokes until Enter, which starves codex's per-keystroke composer) and fixed in terminal-ui.js:
|
||||
- Slash command picker never appeared in codex sessions (#222): the "/" sat in the overlay until Enter, so codex never saw it. Codex-mode sessions now use plain PTY echo (same branch as shell), so the picker pops and live-filters as you type.
|
||||
- Arrow keys dead while typing, backspace dead after Ctrl+Backspace (#218): arrows were forwarded to a still-empty composer while typed text sat pending, and after a control-char flush the overlay swallowed every backspace. Codex bypasses the overlay entirely now; the shared overlay branch (claude/gemini/opencode) additionally flushes pending text on composer nav keys, then hands the session to pass-through until Enter/Ctrl+C, and forwards backspace instead of swallowing it when the overlay has no state.
|
||||
- Pasting displaced the typed prompt (#219): bracketed pastes (xterm terminal.paste with DECSET 2004 active) were forwarded without flushing pending typed text, so the paste landed first. The shared branch now flushes typed text first and delays the paste sequence by 80ms, because codex's paste-burst handling drops keystrokes that arrive in the same PTY read as a bracketed paste (verified against codex 0.147.0 at the byte level).
|
||||
- Long prompts overflowed the bottom of the screen (#220): long typed prompts existed only in the overlay DOM so codex never grew its composer; with plain PTY echo the composer grows and rewraps normally.
|
||||
|
||||
Verified end to end against a real codex 0.147.0 TUI driven by a headless browser: the pre-fix build reproduces all four bugs, the fixed build passes 17/17 assertions. New CI test file test/local-echo-codex-gating.test.ts (41 tests) pins the nav-key classifier, per-mode overlay gating, the flush helper, and pass-through routing. Known upstream limitation: Ctrl+Backspace deletes one character, not a word (xterm.js sends 0x08; word-delete needs kitty CSI-u encoding that xterm.js 6.0.0 cannot emit).
|
||||
|
||||
Mobile keyboard viewport settling fixes by @Lint111 (#229): coalesce keyboard viewport settling so rapid visualViewport resize events during keyboard show/hide no longer thrash the terminal fit, and only arm the settle logic on a real keyboard transition instead of every viewport resize.
|
||||
|
||||
## 1.12.1
|
||||
|
||||
### Patch Changes
|
||||
|
||||
- Terminal scrollback fixes, round 2 of issue #205. A Claude pane's local buffer is hollow (tmux keeps no history for a repaint-mode pane), and both retest reports traced back to that fact. The scroll-to-top full-history re-pull now refuses to rewrite the terminal when the capture holds less than the browser already does, so it can no longer delete history mid-scroll on iPhone (a refused session also re-fetches far less often). When wheel-forwarding is unavailable on a Claude session (version probe failed, CLI older than 2.1.187, or the "Wheel Scrolls Local History" opt-out) and there is no local scrollback to scroll, wheel and touch now page the CLI's own transcript via coalesced PageUp/PageDown instead of doing nothing. The `claude --version` probe no longer caches a failed run for the server's lifetime (one timed-out probe used to silently disable wheel-forwarding on every device until restart); failures retry with backoff. Every scroll gesture now logs a one-line `[scroll]` routing decision to the browser console for direct diagnosis, and the opt-out setting's tooltip explains that the paging fallback is Claude-only (Codex has none).
|
||||
- 2e69e28: Bound the process-tree walk that could take a machine down.
|
||||
|
||||
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
|
||||
depth limit and no node cap. Across ~28 adopted tmux trees the fan-out exploded,
|
||||
and because each `pgrep` blocks in the kernel while reading `/proc/<pid>/cgroup`
|
||||
under WSL, none returned while the walk kept spawning more — ~13,000 `pgrep`
|
||||
processes stuck in D-state out of ~39,000 total, load average above 13,000,
|
||||
recoverable only by restarting WSL.
|
||||
|
||||
Now: one `ps` snapshot, breadth-first with a visited set, a depth cap and a node
|
||||
cap, in a pure module (`proc-tree.ts`) that the regression tests exercise
|
||||
directly. The snapshot is refreshed asynchronously, and the kill path forces a
|
||||
fresh one so the SIGKILL escalation cannot re-read pre-SIGTERM state.
|
||||
|
||||
- ebfcac6: An input whose delivery fails can be retried instead of being lost for good.
|
||||
|
||||
Both input paths recorded the `(clientId, seq)` pair as applied and acknowledged
|
||||
the frame _before_ knowing whether the write had landed — the POST route because
|
||||
its mux write is fire-and-forget, the WebSocket handler because it ACKed
|
||||
unconditionally. When the write then failed, the client dropped the frame from its
|
||||
durable queue and the server rejected the retry as a duplicate: the reliable
|
||||
delivery layer was guaranteeing exactly-once delivery of something that had never
|
||||
been delivered.
|
||||
|
||||
The bookkeeping is now rolled back on failure and the WebSocket ACK withheld, so
|
||||
the client redelivers. `Session.write()` reports whether it reached a PTY at all
|
||||
instead of silently swallowing the data.
|
||||
|
||||
Response codes are unchanged: a session can legitimately have no PTY yet (created
|
||||
but not started), so turning that into a failure status would be a contract change
|
||||
of its own.
|
||||
|
||||
Note this does not remove the root cause: the POST still answers 200 before the
|
||||
mux write is attempted, so a client that treats any 2xx as final still cannot
|
||||
learn about that failure. Closing that would mean awaiting the tmux child in the
|
||||
request path.
|
||||
|
||||
- 1a32e63: Routes that answer with `reply.raw.writeHead()` no longer drop the headers the
|
||||
security hook set.
|
||||
|
||||
`writeHead` writes straight to the Node response and bypasses Fastify's header
|
||||
store, so everything the `onRequest` hook granted was silently lost — including the
|
||||
`Access-Control-Allow-Origin` it emits for localhost origins, and the
|
||||
`X-Content-Type-Options` / `X-Frame-Options` / CSP headers. A localhost page could
|
||||
therefore call every other `/api` endpoint cross-origin while its EventSource
|
||||
failed CORS.
|
||||
|
||||
Affects `GET /api/events` and the three raw-writing routes in `file-routes.ts`
|
||||
(`file-raw`, `tail-file`, `download`).
|
||||
|
||||
## 1.12.0
|
||||
|
||||
### Minor Changes
|
||||
|
||||
- Terminal scrollback overhaul (issue #205), fixing every reported scroll failure across shell and CLI sessions, desktop and mobile:
|
||||
- Shell, OpenCode and Antigravity sessions finally have working scrollback: tmux's own client-side alternate-screen switch is stripped for tmux-backed sessions (narrow strip: alt-screen toggles only, keeping `clear`'s 3J and mouse DECSETs), so xterm stays in the normal buffer instead of a scrollback-less alt buffer where the wheel turned into shell history cycling and touch scrolling did nothing. Direct-PTY fallback sessions are untouched so fullscreen apps (vim/less/htop) keep the alt screen there.
|
||||
- The wheel listener now runs in capture phase and owns the scroll: xterm's internal vscode-style viewport scroller consumed wheel events whenever local scrollback existed (and goes deaf entirely after a tab switch or replay resets the terminal), which silently killed wheel forwarding, made scrolling break after reload/tab switches, and let the CLI's input box scroll away. Local scrolling goes through buffer-level scrollLines and keeps working after resets; mouse-tracking apps and alternate-buffer sessions are passed through untouched.
|
||||
- Wheel AND touch scrolling now forward to the CLI's own transcript for Codex and Claude 2.1.187+, at any scroll position (the viewport snaps home first), so the input box stays pinned on desktop and phones alike. Shift+wheel and the "Wheel scrolls local history" setting still pin local scrollback.
|
||||
- Smooth scrolling: local wheel scrolling glides with an ease-out animation (fractional line accumulation, so slow trackpad drags track the finger instead of running ahead).
|
||||
- Full tmux history on demand: the full-scrollback replay is now per session instead of once per page load, and scrolling up at the top of the buffer re-pulls the complete tmux history, recovering everything tmux's repaint bursts or tab switches removed from the browser's copy.
|
||||
- Firefox wheel speed: wheel deltas are normalized by deltaMode (Firefox reports line units, previously read as pixels and slowed ~4x).
|
||||
- Remote SSH Claude sessions now probe the CLI version over ssh (same connection options and login-shell wrapper as the real launch), so wheel forwarding works for them too instead of silently staying off.
|
||||
|
||||
Docs: scrollback analysis and fix plan recorded in docs/, architecture invariants updated (strip flavors, capture-phase wheel ownership, per-session full-history replay); docker agent-image rebuild warning and integration-guide link fixes from the preceding docs commits.
|
||||
|
||||
## 1.11.2
|
||||
|
||||
### Patch Changes
|
||||
|
||||
- Make Antigravity (`agy`) a first-class CLI everywhere, and stop presenting Gemini CLI as a consumer product now that it is enterprise-only.
|
||||
|
||||
Antigravity was already wired into the session layer, schemas, run-mode menu and remote/Docker command maps, but the surfaces around it were never updated. Gemini keeps full support; Antigravity now sits beside it.
|
||||
|
||||
Fixes:
|
||||
- **Docker cases with `mode: 'antigravity'` were broken.** `docker/agent.Dockerfile` installs its CLIs from npm, and `agy` is not an npm package, so the binary was never in the image and the container died on command-not-found. It now gets its own installer step. The `--dir /usr/local/bin` flag is load-bearing: the installer's default `$HOME/.local/bin` resolves to root's home at build time and would be unreachable by the `agent` user the container runs as. Note the binary is roughly 190MB, making it the largest layer in the image, so rebuild with `node scripts/build-agent-image.mjs` when convenient.
|
||||
- **Welcome screen** gained a "Run Antigravity" action, gated on `agy` being present like the other CLI buttons, styled with the same cyan identity as the toolbar run button and run-mode dot.
|
||||
- **`install.sh`** now detects `agy` (search paths mirroring `antigravity-cli-resolver.ts`), counts it as a satisfying AI CLI so an Antigravity-only box is not told it has none, and recommends it instead of Gemini in the install hints.
|
||||
|
||||
Documentation corrections where it had become factually wrong: `architecture-invariants.md` described `isExternalCliMode()` as opencode/codex/gemini when the code has included antigravity for some time, said "all three modes", and omitted `ANTIGRAVITY_*` from the env-prefix allowlist row; the `agentType` enum in `cron-guide.md`, `SessionMode` in `cron-discovery.md`, and `RemoteCommandMode` in `remote-sessions.md` were all stale.
|
||||
|
||||
Also updated both READMEs (five CLIs, Gemini marked enterprise-only), the `antigravity` npm keyword, and comment drift in eight places. Test coverage added for the new welcome button.
|
||||
|
||||
Antigravity stores its state under `~/.gemini/antigravity-cli/` rather than a `~/.antigravity` directory, so the existing `.gemini` Docker credential seed already covers it. That is now recorded in a code comment so no dead configuration gets added later.
|
||||
|
||||
- b982c5d: Keep the brief Response Viewer output inside the same message card and Markdown wrapper used by the full conversation view, so opening the viewer without clicking More preserves the same readable formatting.
|
||||
|
||||
## 1.11.1
|
||||
|
||||
### Patch Changes
|
||||
|
||||
@@ -47,7 +47,7 @@ The production server caches static files for 1 year, `immutable` (`maxAge: '1y'
|
||||
|
||||
## COM Shorthand (Deployment)
|
||||
|
||||
Uses [Semantic Versioning](https://semver.org/) (`MAJOR.MINOR.PATCH`) via `@changesets/cli`. What SemVer actually covers (the CLI + documented env vars are public; the HTTP/SSE API, on-disk state, and experimental features are internal/unstable) is defined in `docs/versioning-policy.md`. Security reporting + known limitations live in `.github/SECURITY.md`.
|
||||
Uses [Semantic Versioning](https://semver.org/) (`MAJOR.MINOR.PATCH`) via `@changesets/cli`. What SemVer actually covers (the CLI, documented env vars, **and the HTTP/SSE API under `/api/v1`**: endpoint paths, response envelope, `errorCode` values and SSE event names are public/stable; on-disk state, internal TS modules, and experimental features are internal/unstable) is defined in `docs/versioning-policy.md`. Third-party integration surfaces are documented in `docs/extending-codeman.md`. Security reporting + known limitations live in `.github/SECURITY.md`.
|
||||
|
||||
When user says "COM":
|
||||
|
||||
@@ -74,7 +74,7 @@ When user says "COM":
|
||||
|
||||
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
|
||||
|
||||
**Version**: 1.11.1 (must match `package.json`)
|
||||
**Version**: 1.14.0 (must match `package.json`)
|
||||
|
||||
## Project Overview
|
||||
|
||||
@@ -102,13 +102,15 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
|
||||
| Test coverage | `npm run test:coverage` |
|
||||
| Dead-code sweep | `npm run knip` (config in `config/knip.json`, passed via `--config`) |
|
||||
| Rebuild gesture overlay | `npm run build:gesture` (esbuild `packages/gesture-control/src/codeman/entry.ts` → `src/web/public/gesture/gesture-codeman.js`; commit the result) |
|
||||
| Build the docker agent image | `node scripts/build-agent-image.mjs` (builds `codeman/agent:base` from `docker/agent.Dockerfile`; prerequisite for Docker cases; `--engine`/`--image`/`--no-cache`) |
|
||||
| Build the docker agent image | `node scripts/build-agent-image.mjs --no-cache` (builds `codeman/agent:base` from `docker/agent.Dockerfile`; prerequisite for Docker cases; `--engine`/`--image`). ⚠ **Always `--no-cache`** — a plain rebuild re-uses the cached `npm install -g` layer and silently keeps the CLIs frozen at their original versions, which once shipped a BROKEN codex while reporting success. See `docs/docker-cases.md` |
|
||||
| Gesture playground | `npm run dev` **in** `packages/gesture-control/` (standalone vite demo, fake tabs) |
|
||||
| Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) |
|
||||
| Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) |
|
||||
| CI-equivalent test sweep | `npm run test:ci` (full suite minus browser/perf — see Testing) |
|
||||
| Production start | `npm run start` |
|
||||
| Production logs | `journalctl --user -u codeman-web -f` |
|
||||
| Detached server | `codeman web -d` (`--status`, `--stop`; pidfile+log at `dataPath('web.pid'/'web.log')`). ⚠ Refuses to start a 2nd server on one data dir — see Instance isolation |
|
||||
| Install/remove the service | `codeman service install` / `status` / `uninstall` (systemd user unit on Linux, LaunchAgent on macOS; names from `config/service-names.ts`) |
|
||||
|
||||
**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 3 Playwright tests). Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing).
|
||||
|
||||
@@ -141,7 +143,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
|
||||
|
||||
| Domain | Key files | Notes |
|
||||
| ---------------- | -------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
|
||||
| **Entry** | `src/index.ts`, `src/cli.ts` | |
|
||||
| **Entry** | `src/index.ts`, `src/cli.ts`, `daemon-control`, `service-installer`, `config/service-names` | The last three back `web -d` / `service install` |
|
||||
| **Session** | `src/session.ts` ★, `session-manager`, `session-auto-ops`, `session-cli-builder`, `session-task-cache`, `session-order` (pure), `session-pty-exit-breaker`, `usage-limit-patterns`, `usage-telemetry`; `src/services/unified-session-service.ts` | Pure/unit-tested helpers are split out of `session.ts` on purpose |
|
||||
| **Mux** | `src/mux-interface.ts`, `src/mux-factory.ts`, `src/tmux-manager.ts` ★ | |
|
||||
| **Respawn** | `src/respawn-controller.ts` ★ + 4 helpers (`-adaptive-timing`, `-health`, `-metrics`, `-patterns`) | Read `docs/respawn-state-machine.md` first |
|
||||
@@ -180,6 +182,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
|
||||
|
||||
**Input**: `session.writeViaMux()` for programmatic/curl input via tmux `send-keys -l` + `send-keys Enter`, single-line only. Interactive **browser** input goes through a durable **exactly-once** layer: a stable `clientId` + monotonic per-session `seq` persisted to localStorage until the server ACKs, so a dropped link cannot lose or double-deliver a prompt. `ws-connection-registry.ts` supersedes only same-TAB reconnects, so two tabs on one session coexist. → [architecture-invariants#input-delivery-and-ws-resilience](docs/architecture-invariants.md#input-delivery-and-ws-resilience)
|
||||
|
||||
**Agent wait primitives**: bounded long-polls so an agent driving Codeman from a shell can block instead of poll: `GET /api/sessions/:id/wait` (lifecycle signal), `GET /api/sessions/:id/wait-output` (literal substring, **never** regex) and `wait`/`waitTimeout` on `POST /api/sessions/:id/input`. Registry in `session-wait-registry.ts` (pure, no `Session` reference), bounds in `config/agent-wait.ts`. ⚠️ **A timeout is a 200** (`wait.timedOut`), never an error, so callers loop over short waits. ⚠️ `stop`/`blocked` come from Claude Code hooks and therefore fire for **`claude` mode ONLY** (`shell` installs none either); asking for one explicitly on another mode is a 400, the default set silently drops them. ⚠️ Send-and-wait registers the waiter BEFORE the write (a separate POST-then-wait races and reports the PREVIOUS turn), and both teardown paths must `notifySignal('exit')` BEFORE `cancelAll()`. ⚠️ Client-hangup abort listens on **`reply.raw`** guarded by `writableFinished`: on `req.raw`, `close` fires when the request BODY ends, which on a POST killed every send-and-wait instantly and no `app.inject()` test could see it. ⚠️ Worker liveness cannot come from `session.pid` — for a tmux session that is the local attach client, which outlives a worker dying inside its pane — so it is probed at the mux layer (`isPaneDead`, ~750 ms cache) on blocking waits only, never on the input hot path. ⚠️ Signals are edge-triggered with no history: one that fires with no waiter registered is unobservable afterwards, so gather fan-outs with send-and-wait or latched `wait-output` markers, never fire-and-forget-then-sequential-signal-waits. → [architecture-invariants#agent-wait-primitives](docs/architecture-invariants.md#agent-wait-primitives), `docs/api-reference.md`
|
||||
|
||||
**Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
|
||||
|
||||
**Auto-resume on usage limit** (opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit, `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time and `SessionAutoOps` arms a timer for reset+2min, then sends Esc + `continue`. ⚠️ Respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected`), which is what prevents `/clear` from wiping the paused conversation. Claude-mode only. → [architecture-invariants#auto-resume-on-usage-limit](docs/architecture-invariants.md#auto-resume-on-usage-limit)
|
||||
@@ -194,7 +198,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
|
||||
|
||||
**Docker cases**: a case can point at a **container**, with any of the five CLI backends running inside it. Like remote-SSH this is a **LOCATION OVERLAY on cases, never a sixth `SessionMode`**. Exactly one long-lived container **per case**, shared by all its sessions, so killing a session kills only that session's in-container tmux and **never** `docker stop` while siblings remain. The workspace is a real host dir bind-mounted at the **same absolute path**, which is what keeps file-routes/watchers on real host bytes and makes the in-container transcript projHash match the host. Credentials are **seeded** (RO mount, copied into the container once) rather than shared RW, so in-container CLIs never write refreshed tokens back to the host, and bind mounts are excluded from `docker commit` so exports stay secret-free. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** Config drift is detected via a label hash and a drifted launch is REFUSED rather than silently launched with stale config. ⚠️ On the loopback-only prod bind a container cannot reach 127.0.0.1, so in-container hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; otherwise idle detection falls back to output-based. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` (user guide), `docs/docker-cases-plan.md` (design)
|
||||
|
||||
**External CLI modes (OpenCode, Codex, Gemini, Antigravity)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. All four **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. → [architecture-invariants#external-cli-modes-opencode-codex-gemini](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini)
|
||||
**External CLI modes (OpenCode, Codex, Gemini, Antigravity)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. All four **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. ⚠️ **The local-echo overlay is DISABLED for codex sessions** (`_updateLocalEchoState` in terminal-ui.js, same branch as shell): codex's composer reacts per keystroke ("/" pops a live-filtering picker, arrows edit server-side state, the composer grows as it wraps), so buffer-until-Enter starved it into issues #218/#219/#220/#222. Codex also **drops keystrokes that share a PTY read with a bracketed paste**, so flushed text and the paste sequence must go out as separate delayed writes (mirroring the Enter branch's delayed `\r`). Tests: `test/local-echo-codex-gating.test.ts`. → [architecture-invariants#external-cli-modes-opencode-codex-gemini](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini)
|
||||
|
||||
**Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w<n>-<case>` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization)
|
||||
|
||||
@@ -206,7 +210,11 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
|
||||
|
||||
**Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit)
|
||||
|
||||
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. Only the FIRST buffer load after a page load requests `full=1`; tab switches keep the cheap `?tail=` path. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
|
||||
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of EACH session per page load requests `full=1` (`_fullHistoryLoaded` Set); tab switches keep the cheap `?tail=` path, and scrolling up at the TOP of the buffer re-pulls `full=1` on demand (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). ⚠️ That re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
|
||||
|
||||
**Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for codex/claude ≥ 2.1.187 at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). ⚠️ When that gate is FALSE on a claude session whose local buffer is hollow (`baseY === 0`), the gesture becomes coalesced PageUp/PageDown key sends (`_maybePageCliTranscript`) instead of a no-op; ⚠️ and `getClaudeCliVersion()` must never cache a FAILED probe (one timeout used to disable forwarding process-wide until restart). `_logScrollRouting()` prints the routing decision and its inputs once per session — read it before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
|
||||
|
||||
**Detached start + service install** (issue #231): `codeman web -d` relaunches the SAME entry script with `detached:true` (setsid), so there is no controlling terminal and no shell job entry. ⚠️ `nohup` is NOT what makes this work: Node re-arms SIGHUP to its default disposition even when it inherits "ignore", and `cli.ts` handles SIGHUP with a graceful shutdown, so a delivered HUP still stops the server. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile check + `/api/status` probe): a second instance on the shared tmux socket attaches PTYs to the first one's live sessions. ⚠️ Neither may report success it has not observed — the parent polls `/api/status` until the child answers or dies, since `launchctl load` and a clean spawn are both silent about a server that starts and immediately exits. `--stop` verifies the pid still LOOKS like a Codeman server (`ps -o command=`) before signalling, because pids get recycled. Unit/label names live in `config/service-names.ts` so install.sh, `detectSupervisor()` and `service install` cannot drift into supervising two copies; they are instance-scoped, and identical to the historical names for the default instance. `service install` bakes the installing shell's PATH into the unit (launchd gives a job `/usr/bin:/bin:/usr/sbin:/sbin`, which finds neither a Homebrew/nvm `node` nor `tmux`/`claude`) and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install)
|
||||
|
||||
**Self-update** (App Settings → Updates): in-app updater for git-clone installs supervised by systemd/launchd (`systemd`, `launchd`, `launchd-daemon`, else `none` → "restart manually"). The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` that outlives the restart and writes progress to `update-status.json`, which the browser polls across the connection drop. `src/web/self-update.ts` splits pure helpers (unit-tested) from IO wrappers. npm installs report as non-updatable. → [architecture-invariants#self-update](docs/architecture-invariants.md#self-update)
|
||||
|
||||
@@ -290,7 +298,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
|
||||
|
||||
### API Routes
|
||||
|
||||
~199 handlers across 21 route files in `src/web/routes/`: system (45), sessions (32), cases (27), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
|
||||
~200 handlers across 21 route files in `src/web/routes/`: system (45), sessions (34), cases (27), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
|
||||
|
||||
**HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`).
|
||||
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
<h2 align="center">Mission control for AI coding agents</h2>
|
||||
|
||||
<p align="center">
|
||||
<em>Claude Code • OpenCode • Codex • Gemini • Terminal - One Dashboard • Any Device</em>
|
||||
<em>Claude Code • OpenCode • Codex • Antigravity • Gemini • Terminal - One Dashboard • Any Device</em>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
@@ -27,7 +27,7 @@
|
||||
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — parallel subagent visualization" width="900">
|
||||
</p>
|
||||
|
||||
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, or Gemini CLI inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
|
||||
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, or Gemini CLI inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
|
||||
|
||||
Get started in one line (macOS & Linux, Windows via WSL):
|
||||
|
||||
@@ -42,7 +42,7 @@ codeman web
|
||||
|
||||
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
|
||||
|
||||
- **One dashboard, four CLIs** - run [Claude Code, OpenCode, Codex, or Gemini](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
|
||||
- **One dashboard, five CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, or Gemini](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
|
||||
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
|
||||
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
|
||||
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
|
||||
@@ -68,7 +68,7 @@ This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, a
|
||||
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
|
||||
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
|
||||
|
||||
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), or [Gemini CLI](https://github.com/google-gemini/gemini-cli) (any combination works). The installer detects whichever of the four is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
|
||||
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the five is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
|
||||
|
||||
```bash
|
||||
codeman web
|
||||
@@ -85,9 +85,29 @@ codeman web --multiuser # named logins + per-user case spaces
|
||||
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
|
||||
|
||||
<details>
|
||||
<summary><strong>Run as a background service</strong></summary>
|
||||
<summary><strong>Keep it running in the background</strong></summary>
|
||||
|
||||
The installer's final menu sets this up for you (option 2) and verifies the service actually comes up before claiming success. To configure it manually instead:
|
||||
To outlive the shell you started it in, without setting anything up:
|
||||
|
||||
```bash
|
||||
codeman web -d # detach; logs to ~/.codeman/web.log
|
||||
codeman web --status # is it up, and on which pid
|
||||
codeman web --stop # graceful SIGTERM; agents keep running in tmux
|
||||
```
|
||||
|
||||
`-d` waits until the server actually answers before reporting success, and refuses to start a second one on the same data dir (two servers sharing a tmux socket attach to each other's sessions).
|
||||
|
||||
To have it come back after a reboot, install it as a service instead. The installer's final menu does this for you (option 2); `codeman service` is the equivalent for an `npm i -g aicodeman` install:
|
||||
|
||||
```bash
|
||||
codeman service install # systemd user unit (Linux) or LaunchAgent (macOS)
|
||||
codeman service status
|
||||
codeman service uninstall
|
||||
```
|
||||
|
||||
`service install` writes the unit with your current PATH baked in, which matters more than it sounds: launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin`, so a Homebrew or nvm `node`, `tmux` or `claude` is invisible to a hand-written plist. It never copies `CODEMAN_PASSWORD` into the unit file; add that yourself if the service needs auth.
|
||||
|
||||
To write the unit by hand instead:
|
||||
|
||||
**Linux (systemd):**
|
||||
|
||||
@@ -151,7 +171,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
|
||||
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
|
||||
```
|
||||
|
||||
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), or [Gemini CLI](https://github.com/google-gemini/gemini-cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
|
||||
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
|
||||
|
||||
</details>
|
||||
|
||||
@@ -220,6 +240,8 @@ codeman web # localhost:3000 (loopback only — safe defau
|
||||
codeman web --port 8080 # custom port (or set CODEMAN_PORT)
|
||||
codeman web --https # self-signed TLS (only needed for remote access)
|
||||
codeman web -H 0.0.0.0 # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
|
||||
codeman web -d # detach: survives closing the shell (--status, --stop)
|
||||
codeman service install # systemd/launchd service: comes back after reboots
|
||||
```
|
||||
|
||||
Open the printed URL. The page is a single dashboard; everything below happens there.
|
||||
@@ -231,7 +253,7 @@ Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in
|
||||
| Field | What it does |
|
||||
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
|
||||
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. |
|
||||
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Gemini`, or `Terminal` (plain shell). |
|
||||
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, or `Terminal` (plain shell). |
|
||||
| **Model** | Per-session model (App Settings → Claude Model). A soft default — `/model` still works in-session. |
|
||||
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
|
||||
|
||||
@@ -268,6 +290,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
|
||||
### 7. Operate & maintain
|
||||
|
||||
- **App Settings** — model, effort, permission startup mode, theme/skin, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
|
||||
- **Run it in the background** — `codeman web -d` detaches from your shell (`--status`, `--stop`); `codeman service install` makes it a systemd user unit / macOS LaunchAgent that survives reboots. Both verify the server actually answers before reporting success, and both refuse to start a second server on one data dir. See [Keep it running in the background](#quick-start---installation).
|
||||
- **Self-update** — git-clone installs update in place from **Settings → Updates**.
|
||||
- **Deploy your own changes** — see [Development](#development).
|
||||
|
||||
@@ -403,8 +426,9 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
|
||||
|
||||
## More Features
|
||||
|
||||
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
|
||||
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
|
||||
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
|
||||
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
|
||||
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
|
||||
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
|
||||
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
|
||||
@@ -426,7 +450,7 @@ Run a case inside its own hardened Docker container instead of directly on your
|
||||
- **Resource templates** — expand the checkbox for a **Small / Medium / Large / GPU** preset (memory, CPUs, GPU), or set your own. **Disk is elastic** — storage grows as data flows in, no fixed cap.
|
||||
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
|
||||
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
|
||||
- **Seamless auth, isolated credentials** — your host Claude / Codex / Gemini / OpenCode logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
|
||||
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
|
||||
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
|
||||
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
|
||||
|
||||
@@ -612,7 +636,7 @@ These run for **every** request — before auth, even on the default no-password
|
||||
|
||||
### Input, files & headers
|
||||
|
||||
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` env-prefix allowlist gates which settings each CLI can receive
|
||||
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` env-prefix allowlist gates which settings each CLI can receive
|
||||
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
|
||||
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
|
||||
|
||||
@@ -680,15 +704,21 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
|
||||
|
||||
### Rules of the road (read before you POST)
|
||||
|
||||
1. **Single-line input only.** Programmatic input is sent as literal text **+ Enter** in one shot. Multi-line strings break the agent TUI (Ink) — send one line, or split into multiple calls.
|
||||
1. **Single-line input, ending in `\r`.** Programmatic input is sent as literal text, and Enter fires **only when the input contains a carriage return**: `{"input":"run tests\r"}`. Without the `\r` the text sits on the session's prompt unsubmitted (and a combined `wait` runs its full timeout on a turn that never started). Embedded newlines are stripped rather than rejected, so `"echo A\necho B\r"` runs the joined command `echo Aecho B`: send one line per call.
|
||||
2. **Make input idempotent.** Include a stable `clientId` and a monotonic per-session `seq` on `POST …/input`. The server de-duplicates, so a retry after a dropped connection can't double-deliver a prompt.
|
||||
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard).
|
||||
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard). ⚠️ A `401` replies with the bare string `Unauthorized`, **not** the JSON envelope, so piping it into `jq` throws a parse error instead of showing the failure: check the status before parsing.
|
||||
4. **Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
|
||||
5. **`/api/v1/*`** is a stable alias of `/api/*`.
|
||||
6. **Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
|
||||
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
|
||||
8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
|
||||
|
||||
### Recipes
|
||||
|
||||
```bash
|
||||
# CODEMAN_API_URL is auto-set inside every Codeman session, correct scheme included.
|
||||
# The fallback below fits a stock install; on a --https install set the https:// URL
|
||||
# yourself and add -k to each curl (self-signed cert).
|
||||
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
|
||||
# (add -u admin:"$CODEMAN_PASSWORD" to each call if a password is set)
|
||||
|
||||
@@ -700,18 +730,63 @@ curl -s -X POST "$API/api/quick-start" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
|
||||
|
||||
# 2b. Wait until that worker is actually READY (see rule 8): composer marker first,
|
||||
# first-run trust dialog only as the fallback. (Probing trust first and sending
|
||||
# a blind Enter misfires on re-runs: the dialog text stays in the buffer forever,
|
||||
# so the probe matches stale text and the Enter lands in a ready composer.)
|
||||
# Match single tokens: TUI text can reach the matcher without its spaces.
|
||||
until [ "$(curl -s "$API/api/sessions/$SID" | jq '.data.pid')" != null ]; do sleep 1; done
|
||||
R=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
|
||||
--data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
|
||||
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
|
||||
T=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=trust' \
|
||||
--data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
|
||||
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
|
||||
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
|
||||
-d '{"input":"\r","useMux":true}' # accept the first-run trust dialog
|
||||
curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
|
||||
--data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' >/dev/null
|
||||
fi
|
||||
|
||||
# 3. Send a prompt into a session (exactly-once: clientId + seq)
|
||||
curl -s -X POST "$API/api/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"Run the test suite and summarize failures","useMux":true,"clientId":"agent-1","seq":1}'
|
||||
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,"clientId":"agent-1","seq":1}'
|
||||
|
||||
# 4. Read the terminal back
|
||||
curl -s "$API/api/sessions/$SID/output" | jq -r '.data // .'
|
||||
# 4. Send a prompt and BLOCK until that turn is done (registers the wait before
|
||||
# writing, so it can't answer with the previous turn's idle state)
|
||||
curl -s -X POST "$API/api/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,
|
||||
"clientId":"agent-1","seq":2,"wait":"stop,exit","waitTimeout":60000}' \
|
||||
| jq '.data.wait' # -> {"signal":"stop","timedOut":false,"waitedMs":41230,...}
|
||||
# (`stop` is the definitive end-of-turn hook. Adding `idle` makes it resolve on a
|
||||
# spinner pause too, and on anything that redraws a ❯ prompt — like a dialog.)
|
||||
|
||||
# 5. Stream live events (session output, agent activity, status)
|
||||
# 4b. Timed out? That's a 200, not a failure. Loop over short waits.
|
||||
curl -s "$API/api/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq '.data.wait'
|
||||
|
||||
# 4c. Or wait for a marker in the output (works for shell sessions too).
|
||||
# ⚠️ Unique per call (tmux repaints replay old screen text), and SPLIT so the
|
||||
# typed line never contains it: your own keystrokes echo into the output
|
||||
# stream, so an unsplit marker matches before the command has run. from=buffer
|
||||
# catches a marker that printed before the wait landed.
|
||||
N=$RANDOM
|
||||
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
|
||||
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
|
||||
curl -sG "$API/api/sessions/$SID/wait-output" \
|
||||
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
|
||||
--data-urlencode 'timeout=60000' | jq '.data.wait'
|
||||
|
||||
# 5. Read the terminal back. ⚠️ Use terminal?tail=, NOT /output: the latter's
|
||||
# textOutput is empty for every tmux-backed (i.e. every interactive) session.
|
||||
# tail counts BYTES, and what comes back is terminal data, ANSI included.
|
||||
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
|
||||
|
||||
# 6. Stream live events (session output, agent activity, status)
|
||||
curl -sN "$API/api/events" # Server-Sent Events
|
||||
|
||||
# 6. Schedule recurring work (cron-style job)
|
||||
# 7. Schedule recurring work (cron-style job)
|
||||
curl -s -X POST "$API/api/cron/jobs" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
|
||||
@@ -719,11 +794,11 @@ curl -s -X POST "$API/api/cron/jobs" \
|
||||
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
|
||||
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
|
||||
|
||||
# 7. Inspect background sub-agents and their transcripts
|
||||
# 8. Inspect background sub-agents and their transcripts
|
||||
curl -s "$API/api/subagents" | jq '.data // .'
|
||||
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
|
||||
|
||||
# 8. Whole-system snapshot (sessions, settings, respawn, stats)
|
||||
# 9. Whole-system snapshot (sessions, settings, respawn, stats)
|
||||
curl -s "$API/api/status" | jq
|
||||
```
|
||||
|
||||
@@ -749,7 +824,7 @@ Codeman registers Claude Code hooks that `POST /api/hook-event` (`permission_pro
|
||||
|
||||
## API
|
||||
|
||||
REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
|
||||
REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
|
||||
|
||||
### Sessions
|
||||
|
||||
@@ -757,8 +832,11 @@ REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE str
|
||||
| -------- | -------------------------- | ---------------------------------------------------------------------------------- |
|
||||
| `GET` | `/api/sessions` | List all |
|
||||
| `POST` | `/api/quick-start` | Create case + start session (`{caseName?, mode?, effort?, envOverrides?}`) |
|
||||
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?}` — `clientId`+`seq` = exactly-once) |
|
||||
| `GET` | `/api/sessions/:id/output` | Read terminal output |
|
||||
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?, wait?, waitTimeout?}`: `clientId`+`seq` = exactly-once; `wait` blocks until the turn ends) |
|
||||
| `GET` | `/api/sessions/:id/terminal` | Read terminal output (`?tail=<bytes>`, `?full=1`); the read path for interactive sessions |
|
||||
| `GET` | `/api/sessions/:id/output` | Parsed one-shot output (`textOutput` is empty for tmux-backed sessions) |
|
||||
| `GET` | `/api/sessions/:id/wait` | Block until a signal fires (`?until=stop,idle,exit&timeout=&fresh=`); a timeout is a `200` |
|
||||
| `GET` | `/api/sessions/:id/wait-output` | Block until a literal string appears (`?match=&nocase=&from=now\|buffer&timeout=`) |
|
||||
| `GET` | `/api/sessions/unified` | Unified live + history list (Session Manager) — `?q=&limit=` |
|
||||
| `POST` | `/api/sessions/:id/pin` | Pin/unpin in the Session Manager (`{pinned}`) |
|
||||
| `PUT` | `/api/session-order` | Sync tab order across devices (`{order: [ids]}`) |
|
||||
@@ -812,6 +890,8 @@ REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE str
|
||||
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
|
||||
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
|
||||
|
||||
> **Building something on top of Codeman?** [`docs/extending-codeman.md`](docs/extending-codeman.md) is the integration guide: render your own UI as a tab, subscribe to the SSE event stream to react when an agent needs you, drive Codeman from a script, and the traps worth knowing before you start. Codeman has no plugin runtime on purpose, so an integration is just your own process talking HTTP.
|
||||
|
||||
---
|
||||
|
||||
## Architecture
|
||||
@@ -844,7 +924,7 @@ flowchart TB
|
||||
end
|
||||
|
||||
subgraph External["External"]
|
||||
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Gemini</small>"]
|
||||
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini</small>"]
|
||||
BG["Background Agents<br/><small>(Task tool)</small>"]
|
||||
end
|
||||
end
|
||||
|
||||
+10
-8
@@ -5,7 +5,7 @@
|
||||
<h2 align="center">AI 编程智能体的任务控制中心</h2>
|
||||
|
||||
<p align="center">
|
||||
<em>Claude Code • OpenCode • Codex • Gemini • 终端 —— 统一仪表盘 • 任意设备</em>
|
||||
<em>Claude Code • OpenCode • Codex • Antigravity • Gemini • 终端 —— 统一仪表盘 • 任意设备</em>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
@@ -58,7 +58,7 @@ curl -fsSL https://getcodeman.com/install | bash
|
||||
- **重跑即更新。** 再次运行同一条命令即可原地更新已完成的安装:`~/.codeman/app` 中的本地改动会被 stash(绝不丢弃),运行中的服务会自动重启并校验。若首次安装中途失败,重跑会继续完成完整的安装流程。也可以使用 `install.sh update` 与 `install.sh uninstall`。
|
||||
- **CI / 无终端环境:** 没有终端时,涉及系统改动的步骤会带着说明中止,而不是静默执行;在自动化场景设置 `CODEMAN_NONINTERACTIVE=1` 即可批准这些步骤。
|
||||
|
||||
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli)(任意组合均可)。安装器会自动检测这四个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
|
||||
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这五个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
|
||||
|
||||
```bash
|
||||
codeman web
|
||||
@@ -141,7 +141,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
|
||||
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
|
||||
```
|
||||
|
||||
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
|
||||
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
|
||||
|
||||
</details>
|
||||
|
||||
@@ -221,7 +221,7 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
|
||||
| 字段 | 作用 |
|
||||
| ---------------------- | ------------------------------------------------------------------------------------------- |
|
||||
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。 |
|
||||
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Gemini` 或 `Terminal`(普通 shell)。 |
|
||||
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini` 或 `Terminal`(普通 shell)。 |
|
||||
| **模型** | 每会话模型(App Settings → Claude Model)。软默认值 —— 会话内 `/model` 依然有效。 |
|
||||
| **Effort / Ultracode** | 推理力度(`low`–`max`),或用 `ultracode` 开启动态多智能体工作流。随时可用 `/effort` 切换。 |
|
||||
|
||||
@@ -394,7 +394,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
|
||||
## 更多特性
|
||||
|
||||
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
|
||||
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex** 或 **Gemini**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
|
||||
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity** 或 **Gemini**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
|
||||
- **Docker 会话** —— 在隔离且加固的容器中运行案例。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一案例的多个会话共享一个容器;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
|
||||
- **远程 SSH 会话**:把案例指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
|
||||
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
|
||||
@@ -416,7 +416,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
|
||||
- **资源模板** —— 展开复选框可选 **Small / Medium / Large / GPU** 预设(内存、CPU、GPU),也可以完全自定义。**磁盘是弹性的** —— 存储随数据增长,没有固定上限。
|
||||
- **按案例共享容器** —— 多个会话可以 `docker exec` 进同一个容器;结束某个会话绝不会影响其他会话所在的容器。
|
||||
- **默认加固** —— 非 root、`--cap-drop ALL`、`no-new-privileges`、PID/内存上限,绝不使用 `--privileged` 或 docker socket;**密封(sealed)** 配置(不注入主机凭据、关闭网络)只需一个开关。
|
||||
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Gemini / OpenCode 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
|
||||
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Antigravity / Gemini / OpenCode 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
|
||||
- **迁移到另一台机器** —— 把容器的完整环境(工具链 + 工作区)导出为可移植的 `.tar.gz`,在另一台机器上导入到新案例即可继续。
|
||||
- **持久耐用** —— Codeman 重启后重连会回到同一个存活的智能体;容器停止/重启后则从绑定挂载的转录恢复对话。
|
||||
|
||||
@@ -602,7 +602,7 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
|
||||
|
||||
### 输入、文件与响应头
|
||||
|
||||
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
|
||||
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
|
||||
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 50 MB 原始与下载;`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
|
||||
- **安全响应头** —— `Content-Security-Policy`(`default-src 'self'`,每个例外都逐条列举)、`X-Content-Type-Options: nosniff`、`X-Frame-Options: SAMEORIGIN`、HTTPS 下的 HSTS,以及**仅**对 `localhost` / `127.0.0.1` / `::1` 反射的 CORS
|
||||
|
||||
@@ -802,6 +802,8 @@ Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission
|
||||
| `POST` | `/api/clipboard` | 把文本推送到所有已连接浏览器(`{text}`) |
|
||||
| `GET` | `/api/sessions/:id/run-summary` | 时间线 + 统计 |
|
||||
|
||||
> **想在 Codeman 之上做集成?**[`docs/extending-codeman.md`](docs/extending-codeman.md)(英文)是集成指南:把你自己的界面作为标签页嵌入、订阅 SSE 事件流以便在 agent 需要你时做出响应、用脚本驱动 Codeman,以及动手前值得先了解的那些坑。Codeman 刻意不提供插件运行时,所以一个集成就是你自己的进程在讲 HTTP。
|
||||
|
||||
---
|
||||
|
||||
## 架构
|
||||
@@ -834,7 +836,7 @@ flowchart TB
|
||||
end
|
||||
|
||||
subgraph External["外部"]
|
||||
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Gemini</small>"]
|
||||
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini</small>"]
|
||||
BG["后台智能体<br/><small>(Task 工具)</small>"]
|
||||
end
|
||||
end
|
||||
|
||||
+13
-3
@@ -26,8 +26,8 @@ RUN apt-get update \
|
||||
openssh-client \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# The agent CLIs (all four backends Codeman supports). Pinning is left to the
|
||||
# rebuild cadence (see docs/docker-cases-plan.md, user-decision 2).
|
||||
# The npm-published agent CLIs. Pinning is left to the rebuild cadence (see
|
||||
# docs/docker-cases-plan.md, user-decision 2).
|
||||
RUN npm install -g \
|
||||
@anthropic-ai/claude-code \
|
||||
@openai/codex \
|
||||
@@ -35,6 +35,15 @@ RUN npm install -g \
|
||||
opencode-ai \
|
||||
&& npm cache clean --force
|
||||
|
||||
# Antigravity (`agy`) is NOT on npm — Google ships a standalone binary through its
|
||||
# own installer, so it needs its own step. `--dir /usr/local/bin` is load-bearing:
|
||||
# the installer's default target is `$HOME/.local/bin`, which at build time is
|
||||
# root's home and would be unreachable by the `agent` user the container runs as.
|
||||
# ⚠️ This binary is ~190MB on its own; it is the single largest layer in the image.
|
||||
RUN curl -fsSL https://antigravity.google/cli/install.sh | bash -s -- --dir /usr/local/bin \
|
||||
&& chmod 755 /usr/local/bin/agy \
|
||||
&& agy --version
|
||||
|
||||
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
|
||||
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
|
||||
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
|
||||
@@ -50,7 +59,8 @@ ENV HOME=/home/agent
|
||||
# dirs: tokens/settings/config are seeded in as writable copies and each CLI's runtime
|
||||
# state (backups, tasks, refreshed tokens) stays container-local, while ONLY the shared
|
||||
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
|
||||
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir.)
|
||||
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
|
||||
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
|
||||
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
|
||||
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
|
||||
/home/agent/.claude/projects /home/agent/.codex/sessions \
|
||||
|
||||
@@ -0,0 +1,662 @@
|
||||
# Agent Control Plan: skill packaging + wait primitives
|
||||
|
||||
**Status**: steps 1 to 5 IMPLEMENTED and multi-round verified, uncommitted as of 2026-08-08.
|
||||
Step 6 (CLI install command + per-case injection + `agentSkillEnabled`) is not built.
|
||||
See [§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
|
||||
verification round found, and what is still open.
|
||||
|
||||
**Date**: 2026-08-08
|
||||
**Scope**: Part 1 (agent skill) and Part 2 (wait primitives) were specified and built.
|
||||
Parts 3 to 5 are captured so they are not lost, but remain deliberately deferred.
|
||||
|
||||
---
|
||||
|
||||
## 0. Where this came from: what herdr does
|
||||
|
||||
[herdr](https://github.com/herdrdev/herdr) (Rust, Apache-2.0, ~25.8k stars) is a terminal
|
||||
multiplexer built around AI coding agents. Relevant findings from the research pass:
|
||||
|
||||
| Capability | How herdr does it |
|
||||
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Agent state | Four states (`idle`, `working`, `blocked`, `done`) that roll up pane to tab to workspace in a sidebar |
|
||||
| Detection | Lifecycle hooks where the agent supports them (it names Pi and MastraCode), otherwise TOML manifests matched against a live bottom-buffer snapshot. Bundled manifests plus remote updates from herdr.dev, local overrides win |
|
||||
| Control API | Newline-delimited JSON over a Unix socket (`~/.config/herdr/sessions/<name>/herdr.sock`), `{"id":"req_1","method":"pane.split","params":{}}`, dot-notation methods, plus long-lived event subscriptions |
|
||||
| Discoverability | `herdr api schema` prints a machine-readable schema |
|
||||
| Agent skill | `npx skills add herdrdev/herdr --skill herdr -g`, a SKILL.md wrapping the CLI, guarded by `test "${HERDR_ENV:-}" = 1` so an agent outside a herdr pane refuses to act |
|
||||
| Persistence | Background server, detach with `ctrl+b q`, snapshot restore of workspaces/tabs/panes/cwd/layout, experimental screen-history replay, agent resume via native session ids, live PTY handoff across server replacement |
|
||||
| Plugins | `herdr-plugin.toml` manifest, actions, event hooks, plugin panes, link handlers, GitHub-topic marketplace index |
|
||||
|
||||
The commands the skill teaches the agent:
|
||||
|
||||
| Group | Commands |
|
||||
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| workspace | `workspace list`, `workspace create` |
|
||||
| tab | `tab list --workspace <id>`, `tab create` |
|
||||
| pane | `pane current`, `pane list`, `pane layout`, `pane split --current --direction right --cwd <path> --no-focus`, `pane run <id> "<cmd>"`, `pane wait-output <id> --match/--regex <p> --timeout <ms>`, `pane read <id> --source visible\|recent\|detection` |
|
||||
| agent | `agent list`, `agent start <name> --kind <type> --pane <id>`, `agent prompt <name> "<text>" --wait --timeout <ms>`, `agent wait <name> --until <state> --timeout <ms>`, `agent send-keys`, `agent get`, `agent read` |
|
||||
|
||||
### The honest comparison
|
||||
|
||||
herdr and Codeman are not the same product. herdr is a local, keyboard-first multiplexer with
|
||||
no server, no web UI, and no autonomy layer. Codeman is a server with a browser and mobile UI,
|
||||
remote and Docker cases, respawn, Ralph, cron, and the orchestrator, none of which herdr has.
|
||||
|
||||
What herdr genuinely does better is being **callable by the agent running inside it**. For
|
||||
Codeman that is a packaging problem plus one missing primitive, not an architecture problem.
|
||||
|
||||
---
|
||||
|
||||
## 1. Gap analysis
|
||||
|
||||
| herdr capability | Codeman equivalent today | Gap |
|
||||
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- |
|
||||
| `pane split` + `agent start` | `POST /api/quick-start`, `POST /api/sessions` | none, already there |
|
||||
| `agent prompt` | `POST /api/sessions/:id/input` with `clientId`+`seq` exactly-once | no `--wait` |
|
||||
| `pane read` | `GET /api/sessions/:id/output`, `GET /api/sessions/:id/terminal?full=1` | none |
|
||||
| `agent list` / `agent get` | `GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status` | none |
|
||||
| `agent wait --until <state>` | SSE only (`/api/events`) | **missing**, and SSE is impractical from a shell tool |
|
||||
| `pane wait-output --match` | nothing | **missing** |
|
||||
| Skill file | README section "Driving Codeman from an Agent" | **not packaged**, an agent will never find it |
|
||||
| Env guard `HERDR_ENV=1` | `CODEMAN_MUX=1`, `CODEMAN_API_URL`, `CODEMAN_SESSION_ID` already exported at spawn | none, the guard variables exist |
|
||||
| `blocked` state | hook events (`permission_prompt`, `elicitation_dialog`) plus CSS classes plus the phone overview NEEDS YOU section | not in the wire contract (`SessionStatus = 'idle' \| 'busy' \| 'stopped' \| 'error'`) |
|
||||
| `api schema` | hand-written `docs/api-reference.md` | no machine-readable schema |
|
||||
| Detection manifests | hardcoded in `usage-limit-patterns.ts`, `respawn-*-patterns`, `regex-patterns.ts` | patterns are code, not data |
|
||||
| Plugin runtime | deliberately refused, see `docs/extending-codeman.md` | not a gap, a decision |
|
||||
| Session handoff on restart | tmux owns the PTYs, so they already survive a Codeman restart | not a gap, solved by architecture |
|
||||
|
||||
**Conclusion**: roughly 90% of the capability surface already exists. Parts 1 and 2 below close
|
||||
the two real gaps.
|
||||
|
||||
---
|
||||
|
||||
## 2. Part 1: the Codeman agent skill
|
||||
|
||||
### 2.1 Goal
|
||||
|
||||
An agent running inside a Codeman session can discover and correctly drive Codeman without the
|
||||
user pasting API docs into the prompt, and without inventing dangerous calls.
|
||||
|
||||
### 2.2 Layout and distribution
|
||||
|
||||
The `npx skills` CLI (vercel-labs/skills) clones a GitHub repo and looks for
|
||||
`skills/<name>/SKILL.md`. Claude Code natively discovers `.claude/skills/<name>/SKILL.md` in a
|
||||
project and `~/.claude/skills/` globally. Both are satisfied with one source of truth plus a
|
||||
symlink, which is the pattern this repo already uses for `remotion-best-practices`.
|
||||
|
||||
```
|
||||
skills/
|
||||
codeman/
|
||||
SKILL.md <- single source of truth
|
||||
reference/
|
||||
endpoints.md <- full endpoint tables, loaded on demand
|
||||
recipes.md <- worked multi-session orchestration examples
|
||||
.claude/skills/codeman -> ../../skills/codeman (symlink, dogfooding in this repo)
|
||||
```
|
||||
|
||||
Adding a `skills/` directory to the repo root costs one entry in the GitHub listing. CLAUDE.md
|
||||
keeps the root short on purpose, so this needs a conscious sign-off; the alternative is
|
||||
`docs/skills/codeman/` with a `--skill` path argument, which breaks the one-liner install.
|
||||
**Recommendation**: accept `skills/` at the root, because the install one-liner is the whole
|
||||
point of shipping a skill.
|
||||
|
||||
Install paths, in order of how a user gets it:
|
||||
|
||||
1. `npx skills add Ark0N/Codeman --skill codeman -g` (global, any agent, matches the herdr flow).
|
||||
2. `codeman skill install [--global | --case <name>]`, a new CLI subcommand writing the same
|
||||
file. This is the path for users who installed via npm and never cloned the repo.
|
||||
3. **Automatic per-case injection**, modeled exactly on `applyStatusLineConfig(casePath, enabled)`
|
||||
in `hooks-config.ts`: write `<case>/.claude/skills/codeman/SKILL.md` at case creation,
|
||||
gated on a new setting. Codeman already writes `<case>/.claude/settings.local.json` hooks
|
||||
through `writeHooksConfig()`, so this is the same mechanism with the same lifecycle.
|
||||
|
||||
Setting name: `agentSkillEnabled`. Synced (not per-device), since it changes on-disk case
|
||||
content rather than display. Default: **ON after the dogfooding phase, OFF in the first
|
||||
release**. Rationale for starting OFF: Claude Code loads every skill's name and description
|
||||
into context on every turn, so an always-on skill has a small permanent token cost, and we
|
||||
should measure that we are buying something with it first.
|
||||
|
||||
### 2.3 SKILL.md content
|
||||
|
||||
Frontmatter, per the skills convention (`name` + `description` required):
|
||||
|
||||
```yaml
|
||||
---
|
||||
name: codeman
|
||||
description: >-
|
||||
Control Codeman, the session manager this agent is running inside: list sessions,
|
||||
start worker sessions, send prompts, read terminal output, and wait for other agents
|
||||
to finish. Only usable when CODEMAN_MUX=1.
|
||||
---
|
||||
```
|
||||
|
||||
Body sections, in order:
|
||||
|
||||
**1. Guard (first thing, non-negotiable).**
|
||||
|
||||
```bash
|
||||
test "${CODEMAN_MUX:-}" = 1 || { echo "not inside a Codeman session"; exit 1; }
|
||||
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set, refusing to guess}"
|
||||
SELF="${CODEMAN_SESSION_ID:-}"
|
||||
```
|
||||
|
||||
If `CODEMAN_MUX` is not `1`, the agent must stop and say it is not running inside a
|
||||
Codeman-managed session. Same shape as herdr's `HERDR_ENV` guard, and the variables are
|
||||
already exported by `tmux-manager.buildEnvExports()`. No fallback URL when
|
||||
`CODEMAN_API_URL` is unset: any guess is the wrong scheme on an HTTPS install (prod is
|
||||
HTTPS with a self-signed cert, hence `curl -sk` throughout), and a server the agent
|
||||
cannot identify is not one it should be driving.
|
||||
|
||||
**2. Rules of the road.** Lifted and tightened from README lines 666 to 745:
|
||||
|
||||
- Single-line input only. Multi-line breaks the agent TUI (Ink).
|
||||
- Always send `clientId` + a monotonic `seq` on `POST .../input` so a retry cannot double-deliver.
|
||||
- Envelope is `{success, data}`; a few legacy GETs are bare, so read `body.data ?? body`.
|
||||
- Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set. Prod is HTTPS, so `curl -sk`.
|
||||
- Prefer `/api/v1/*`, the stable alias.
|
||||
|
||||
**3. Safety rules (the section that does not exist anywhere today).**
|
||||
|
||||
- Never act on `$CODEMAN_SESSION_ID`. That is you.
|
||||
- Only `DELETE` sessions **you created in this conversation**, by exact id. Keep the list.
|
||||
- Never bulk-delete, never loop a `DELETE` over `/api/sessions`. There is no undo.
|
||||
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. Use the API.
|
||||
- Creating a session consumes a slot against the 50-session cap. Clean up what you start.
|
||||
|
||||
**4. Recipes**, each one a single copy-pasteable curl:
|
||||
|
||||
| Task | Call |
|
||||
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| list sessions | `GET /api/v1/sessions` |
|
||||
| find yourself | match ids by PREFIX of `$CODEMAN_SESSION_ID` (Docker cases truncate it to 8 chars, so an equality check never fires there) |
|
||||
| start a worker | `POST /api/v1/quick-start {caseName, mode, effort}` |
|
||||
| send a prompt | `POST /api/v1/sessions/:id/input {input:"…\r", useMux:true, clientId, seq}` (the trailing `\r` is what sends Enter; without it the text sits on the prompt unsubmitted) |
|
||||
| send prompt and wait | `POST /api/v1/sessions/:id/input {input:"…\r", wait:"stop", waitTimeout:600000}` (Part 2) |
|
||||
| wait for a worker | `GET /api/v1/sessions/:id/wait?until=stop,blocked&timeout=300000` (Part 2) |
|
||||
| wait for a marker | `GET /api/v1/sessions/:id/wait-output?match=DONE_<random>&timeout=120000` (Part 2; unique per call, per §3.3's repaint rule) |
|
||||
| read output | `GET /api/v1/sessions/:id/output` |
|
||||
| read full scrollback | `GET /api/v1/sessions/:id/terminal?full=1` |
|
||||
| watch sub-agents | `GET /api/v1/subagents` |
|
||||
| schedule work | `POST /api/v1/cron/jobs` |
|
||||
| clean up | `DELETE /api/v1/sessions/:id` |
|
||||
|
||||
**5. Pointer to `reference/endpoints.md`** for anything not in the table, so the always-loaded
|
||||
part of the skill stays small.
|
||||
|
||||
### 2.4 An ergonomics guard worth adding server-side
|
||||
|
||||
The skill will tell the agent not to act on itself, but a confused agent can still try. Propose:
|
||||
the skill sends `X-Codeman-Caller-Session: $CODEMAN_SESSION_ID` on every request, and the server
|
||||
refuses destructive operations (`DELETE /api/sessions/:id`, kill, respawn stop) when that header
|
||||
equals the target id, with a clear error.
|
||||
|
||||
This is a **footgun guard, not a security control**: any caller can omit the header. Document it
|
||||
as such so nobody mistakes it for a boundary. It costs about 10 lines in `route-helpers.ts`.
|
||||
|
||||
### 2.5 Verification
|
||||
|
||||
Per the always-end-to-end-test rule, "the skill exists" is not done. Done is:
|
||||
|
||||
1. Symlink it into `.claude/skills/`, start a real throwaway Codeman session, and ask that agent
|
||||
to "start a worker session that runs the test suite and tell me when it finishes".
|
||||
2. Confirm from the outside that exactly one new session appeared, got the prompt, and that the
|
||||
lead agent waited rather than polling in a busy loop.
|
||||
3. Confirm the guard: run the same prompt in a shell with `CODEMAN_MUX` unset and confirm refusal.
|
||||
4. Confirm cleanup: the worker session is deleted by exact id and no other session was touched.
|
||||
|
||||
Never run this against `w1`/`w2`/`w3`.
|
||||
|
||||
### 2.6 Files touched
|
||||
|
||||
- `skills/codeman/SKILL.md` (new), `skills/codeman/reference/*.md` (new)
|
||||
- `.claude/skills/codeman` symlink (new)
|
||||
- `src/cli.ts` (new `skill install` subcommand)
|
||||
- `src/hooks-config.ts` (new `applyAgentSkill(casePath, enabled)`, mirroring `applyStatusLineConfig`)
|
||||
- `src/web/schemas.ts` (`agentSkillEnabled` in `SettingsUpdateSchema`, which is `.strict()`)
|
||||
- `src/web/routes/system-routes.ts` (settings PUT must resolve the flag from `merged`, never
|
||||
from the raw body, per the partial-PUT invariant)
|
||||
- `src/web/public/settings-ui.js` + `index.html` (checkbox)
|
||||
- `package.json` `files` array, so `skills/` ships to npm
|
||||
- README pointer, `docs/extending-codeman.md` seam 3 pointer
|
||||
|
||||
---
|
||||
|
||||
## 3. Part 2: wait primitives
|
||||
|
||||
### 3.1 Goal
|
||||
|
||||
Make Codeman orchestratable from a shell tool. Today the only "tell me when" channel is SSE,
|
||||
which a curl-driven agent cannot practically consume: it would have to hold a streaming
|
||||
connection and parse events inline. herdr solves this with blocking CLI calls. Codeman should
|
||||
solve it with bounded long-poll endpoints.
|
||||
|
||||
All three additions are **additive**, so the versioning policy stays intact (new endpoints and
|
||||
new optional fields are non-breaking).
|
||||
|
||||
### 3.2 The signal model
|
||||
|
||||
A waiter resolves on the first of a set of signals. Sources that already exist:
|
||||
|
||||
| Signal | Source today |
|
||||
| --------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `idle` | `Session` emits `idle` (session.ts ~1775 for Claude, ~2101 for shell), wired at `session-listener-wiring.ts:402` |
|
||||
| `working` | `Session` emits `working` (session.ts ~1788), wired at `session-listener-wiring.ts:401` |
|
||||
| `stop` | `POST /api/hook-event` with `event: 'stop'`, the definitive "Claude finished responding" signal already used by `controller.signalStopHook()` |
|
||||
| `blocked` | `POST /api/hook-event` with `permission_prompt` or `elicitation_dialog` |
|
||||
| `exit` | `Session` emits `exit` |
|
||||
|
||||
`stop` is the highest-quality signal for "the turn is over" and should be the documented default
|
||||
for orchestration. `idle` is heuristic: output stabilization plus prompt detection, and it can
|
||||
flap mid-turn when a spinner pauses. External CLI modes (`isExternalCliMode()`) have no stop
|
||||
hook at all, so for opencode/codex/gemini/antigravity only `idle`, `working` and `exit` are
|
||||
available. **The skill and the docs must say which signals exist per mode**, otherwise an agent
|
||||
waits forever on `stop` in a codex session.
|
||||
|
||||
### 3.3 Endpoint specs
|
||||
|
||||
#### A. `GET /api/sessions/:id/wait`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
| --------- | ---------------------------------------------- | ---------------- | ------------------------------------------------------------ |
|
||||
| `until` | comma list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on first match |
|
||||
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` (600000) |
|
||||
| `fresh` | `0`/`1` | `0` | `1` requires a _transition_, ignoring the state at call time |
|
||||
|
||||
Response (always 200 unless the session is missing or a cap is hit):
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"data": {
|
||||
"signal": "stop",
|
||||
"timedOut": false,
|
||||
"immediate": false,
|
||||
"ended": false,
|
||||
"waitedMs": 8421,
|
||||
"status": "idle",
|
||||
"sessionId": "...",
|
||||
"until": ["stop", "idle", "exit"],
|
||||
"limitPaused": false
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`until` is echoed back because the server may narrow it: `stop`/`blocked` are dropped
|
||||
from the DEFAULT set for external CLI modes (asking for them EXPLICITLY is a 400
|
||||
instead, since omitting `until` must never 400). `limitPaused` tells a caller that a
|
||||
timeout was expected rather than a stall worth retrying hard.
|
||||
|
||||
**A timeout is not an error.** `{"timedOut": true, "signal": null}` with HTTP 200, so a caller
|
||||
can loop without treating every poll boundary as a failure. Errors are reserved for
|
||||
`NOT_FOUND` (unknown or not-owned session) and `SESSION_BUSY` (waiter cap exceeded).
|
||||
|
||||
`immediate: true` means the session was already in the requested state and `fresh` was not set.
|
||||
|
||||
#### B. `GET /api/sessions/:id/wait-output`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
| --------- | ------------------------------ | -------- | --------------------------------------------------------- |
|
||||
| `match` | literal string, 1 to 200 chars | required | substring match against ANSI-stripped output |
|
||||
| `nocase` | `0`/`1` | `0` | case-insensitive compare |
|
||||
| `from` | `now` \| `buffer` | `now` | `buffer` scans the existing text buffer first, then waits |
|
||||
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` |
|
||||
|
||||
Response: `{ matched: true, timedOut: false, snippet: "...", waitedMs }`.
|
||||
|
||||
**No regex in v1, deliberately.** `search-service.ts` already avoids regex specifically so there
|
||||
is no ReDoS surface, and this endpoint would be even more exposed since the pattern is attacker
|
||||
supplied and the input is a live stream. herdr can offer `--regex` because Rust's regex crate is
|
||||
linear-time with no backtracking; JS `RegExp` is not. If regex is wanted later, the honest
|
||||
options are a length-capped subset compiled once with a match budget, or `re2`. Note it and move on.
|
||||
|
||||
Implementation detail that will bite if missed: a match can straddle two PTY chunks. Keep a
|
||||
carry buffer of `match.length - 1` bytes from the previous chunk and test `carry + chunk`.
|
||||
|
||||
⚠️ **`from=now` does not mean "printed after you asked".** tmux repaints the visible
|
||||
screen on attach, resize, or any TUI redraw, and a repaint arrives as ordinary `terminal`
|
||||
data. Observed live: a marker echoed a minute earlier matched instantly on a fresh
|
||||
`from=now` wait. This is inherent to a terminal multiplexer, not fixable in the registry,
|
||||
so the contract is: **use a marker unique per call** (`echo DONE_$RANDOM`), never a
|
||||
generic one like `BUILD OK`. The skill's recipes must show that.
|
||||
|
||||
The returned snippet is whitespace-collapsed (blank runs to a single newline) for
|
||||
readability only; matching runs on the raw stripped text. Without it, a real pane's
|
||||
`\r\n` padding between the prompt and the match fills the whole context window with
|
||||
nothing, which was the first thing the live test showed.
|
||||
|
||||
#### C. `wait` on the existing input endpoint
|
||||
|
||||
`POST /api/sessions/:id/input` gains two optional fields:
|
||||
|
||||
```json
|
||||
{ "input": "run the tests\r", "useMux": true, "clientId": "agent-1", "seq": 7, "wait": "stop", "waitTimeout": 600000 }
|
||||
```
|
||||
|
||||
(The trailing `\r` is required on every input body: `sendInput` sends Enter only
|
||||
when the input contains a carriage return.)
|
||||
|
||||
Response gains `"wait": { "signal": "stop", "timedOut": false, "waitedMs": 41230 }`.
|
||||
|
||||
This is the important one, because it closes a race the standalone `GET .../wait` cannot: between
|
||||
"input delivered" and "session flips to working" there is a window where a naive
|
||||
send-then-wait sees the _pre-existing_ idle state and returns instantly. The combined endpoint
|
||||
**registers the waiter before writing**, so that window does not exist. This is exactly why herdr
|
||||
ships `agent prompt --wait` as its own thing.
|
||||
|
||||
`wait` accepts `true` (the default signal set) or the same comma grammar as `until`.
|
||||
Both new fields are `.nullish()`, not `.optional()`: a third-party caller building the
|
||||
body with `JSON.stringify` keeps an explicit `null` on the wire, and `.optional()`
|
||||
rejects that with `INVALID_INPUT`. That gotcha has shipped as a real bug twice.
|
||||
|
||||
Two behaviors to preserve carefully:
|
||||
|
||||
- **`useMux` is fire-and-forget today.** The handler responds without awaiting `writeViaMux`, on
|
||||
purpose (a tmux child process must not block the HTTP response). With `wait` present the
|
||||
handler already has to stay open, so it can await delivery, and a `writeViaMux` failure becomes
|
||||
observable for the first time. The non-wait path must keep its current fire-and-forget shape
|
||||
byte for byte.
|
||||
- **Duplicate suppression.** A tagged redelivery (`clientId`+`seq` already applied) returns 200
|
||||
without writing. With `wait` set it still waits, since the caller's intent is "tell me when
|
||||
this settles". But it waits with `requireTransition: false`, unlike a fresh delivery: the
|
||||
original turn may be long over, and requiring a new transition would block a redelivery until
|
||||
timeout for no reason. Fresh delivery requires a transition, a duplicate answers from the
|
||||
current state.
|
||||
- **Capacity rollback.** `shouldApplyInput()` MUTATES (it records the seq), and it runs before
|
||||
the waiter is registered. If registration then fails on a full pool, the handler must call
|
||||
`forgetInputSeq` before returning `SESSION_BUSY`, or the caller's retry is rejected as a
|
||||
duplicate and the input is lost by the very mechanism reliable delivery exists for.
|
||||
|
||||
### 3.4 Module design
|
||||
|
||||
New file `src/web/session-wait-registry.ts`, with the IO-free core unit-testable in isolation
|
||||
(same split as `self-update.ts`):
|
||||
|
||||
```ts
|
||||
type WaitSignal = 'idle' | 'working' | 'stop' | 'blocked' | 'exit';
|
||||
|
||||
waitForSignal(sessionId, { until: Set<WaitSignal>, timeoutMs, requireTransition }): Promise<WaitResult>
|
||||
notifySignal(sessionId, signal: WaitSignal): void
|
||||
waitForOutput(sessionId, { match, nocase, timeoutMs }): Promise<OutputWaitResult>
|
||||
notifyOutput(sessionId, chunk: string): void
|
||||
cancelAll(sessionId, reason): void
|
||||
```
|
||||
|
||||
Wiring points, all existing:
|
||||
|
||||
- `src/web/session-listener-wiring.ts` around lines 190 and 200 already handles `working` and
|
||||
`idle` and broadcasts them. Add a `notifySignal()` call next to each broadcast, plus `exit`.
|
||||
- `src/web/routes/hook-event-routes.ts` already switches on `event` for the respawn controller.
|
||||
Add `notifySignal(sessionId, 'stop' | 'blocked')` in the same switch.
|
||||
- Output: `notifyOutput()` rides the ALREADY-attached `terminal` listener in
|
||||
session-listener-wiring.ts. An earlier draft had the registry hand out attach/detach
|
||||
callbacks so a listener could be added lazily; that was deleted once it was clear no
|
||||
second listener is needed at all. The cost is one Map lookup per PTY chunk, which is why
|
||||
the no-waiter check comes before the ANSI strip.
|
||||
- Session deletion calls `notifySignal('exit')` then `cancelAll()`, so no promise is left
|
||||
hanging. Both are required: `_doCleanupSession` detaches the session's listeners BEFORE
|
||||
`session.stop()`, so on a delete the PTY exit event never reaches the registry, and an
|
||||
`until=exit` caller would otherwise get a bare `ended` instead of its signal. Found by
|
||||
live-testing the delete path, not by the unit tests.
|
||||
|
||||
Memory-leak discipline, per the 24-hour-session rules: every waiter owns a timer that is cleared
|
||||
on resolve, the per-session waiter set is deleted when it empties, and the output listener is
|
||||
removed with it. `test/memory-leak-prevention.test.ts` should grow a case for this.
|
||||
|
||||
Caps in a new `src/config/agent-wait.ts` (limits live in `src/config/`, env-overridable):
|
||||
|
||||
| Constant | Default | Why |
|
||||
| ------------------------- | ------- | --------------------------------------- |
|
||||
| `MAX_WAIT_MS` | 600000 | an unbounded long-poll is a socket leak |
|
||||
| `DEFAULT_WAIT_MS` | 60000 | short enough to survive most proxies |
|
||||
| `MAX_WAITERS_PER_SESSION` | 16 | |
|
||||
| `MAX_WAITERS_TOTAL` | 128 | same reasoning as `MAX_SSE_CLIENTS` |
|
||||
|
||||
Exceeding a cap returns `SESSION_BUSY`, not a silent queue.
|
||||
|
||||
### 3.5 Transport concerns
|
||||
|
||||
Fastify is constructed with defaults in `server.ts:329-331`. `requestTimeout` defaults to 0
|
||||
(disabled) and `keepAliveTimeout` (72s) applies between requests, not to an in-flight one, so a
|
||||
10-minute in-process hold is fine. **Verify this on the real instance before relying on it.**
|
||||
|
||||
Intermediaries are the actual risk. Prod is reached through `tailscale serve`, and users also run
|
||||
cloudflared tunnels; both can cut an idle connection. That is why `DEFAULT_WAIT_MS` is 60s and
|
||||
why the documented pattern is a client-side loop over short waits rather than one 10-minute call.
|
||||
The skill's recipes must show the loop.
|
||||
|
||||
### 3.6 Edge cases to get right
|
||||
|
||||
| Case | Behavior |
|
||||
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Session already idle, `fresh=0` | return immediately, `immediate: true` |
|
||||
| Session already idle, `fresh=1` | wait for the next transition into a requested state |
|
||||
| Session dies mid-wait | resolve with `signal: "exit"` if `exit` was requested, otherwise resolve `timedOut:false, signal:null, ended:true`. Never hang |
|
||||
| Session deleted mid-wait | same, resolve, do not throw. Verified live: `until=exit` gets `signal:"exit"`, a concurrent `until=blocked` gets `ended:true`, both in ~0ms |
|
||||
| Shutdown with a wait pending | `cancelEverything()` in `stop()`. Verified live: SIGTERM with a 300s wait in flight exits in 1s |
|
||||
| External CLI mode | `stop` and `blocked` never fire. Reject `until=stop` for those modes with a clear `INVALID_INPUT` rather than hanging until timeout |
|
||||
| Multi-user | goes through `findSessionOrFail(ctx, id, req)`, which already enforces ownership |
|
||||
| Remote / Docker cases | signals originate from the same `Session` object, so no special casing. Docker hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for `stop`/`blocked` to arrive at all; without it, only `idle` works. Document it |
|
||||
| Respawn `/clear` mid-wait | a respawn cycle emits `idle`. Callers waiting on `stop` are unaffected; callers on `idle` may resolve early. Documented, not fixed |
|
||||
| Limit pause | if the session is paused on a usage limit, nothing will fire until the reset. The wait times out honestly. Consider surfacing `limitPaused: true` in the response so the caller can back off |
|
||||
|
||||
### 3.7 Tests
|
||||
|
||||
- `test/session-wait-registry.test.ts` (pure): immediate resolve, transition-required, multi-signal
|
||||
first-wins, timeout, cap exceeded, cancel on session end, no listener leak after resolve,
|
||||
chunk-straddling output match, case-insensitive match.
|
||||
- `test/routes/session-wait-routes.test.ts` (`app.inject()`, no port): all three endpoints against
|
||||
a `MockSession`, including the 200-with-`timedOut` contract and the ownership 404.
|
||||
- `test/routes/session-input-wait.test.ts`: the send-and-wait race, plus proof that the non-wait
|
||||
path is unchanged (still returns before `writeViaMux` settles).
|
||||
- Live verification on a throwaway session before COM, per the always-end-to-end-test rule.
|
||||
|
||||
### 3.8 Files touched
|
||||
|
||||
- `src/config/agent-wait.ts` (new)
|
||||
- `src/web/session-wait-registry.ts` (new)
|
||||
- `src/web/session-listener-wiring.ts` (notify on idle/working/exit)
|
||||
- `src/web/routes/hook-event-routes.ts` (notify on stop/blocked)
|
||||
- `src/web/routes/session-routes.ts` (two new routes, `wait` fields on input)
|
||||
- `src/web/schemas.ts` (`SessionWaitQuerySchema`, `SessionWaitOutputQuerySchema`, extend
|
||||
`SessionInputWithLimitSchema`. Note: `.optional()` rejects `null`, so the frontend and any
|
||||
generated client must send `undefined`, never `null`)
|
||||
- `docs/api-reference.md`, `docs/extending-codeman.md`, README API table
|
||||
- `skills/codeman/SKILL.md` recipes (Part 1 depends on this)
|
||||
|
||||
---
|
||||
|
||||
## 4. Deferred: parts 3 to 5
|
||||
|
||||
Not in scope now, kept here so they are not lost.
|
||||
|
||||
### Part 3: promote `blocked` to a first-class state
|
||||
|
||||
`SessionStatus` is `'idle' | 'busy' | 'stopped' | 'error'`. "Needs you" exists three times over:
|
||||
hook events, the `tab-alert-action` CSS class, and the phone overview NEEDS YOU section, each
|
||||
re-deriving it. herdr makes `blocked` a real state that rolls up.
|
||||
|
||||
Add `blocked` (and possibly `done`) to `SessionStatus`, set it from the same hook events that
|
||||
Part 2 uses as wait signals, and clear it on the next `working`/`stop`. Then the tab strip, the
|
||||
mobile overview, the wait endpoints, and any external agent read one field.
|
||||
|
||||
Cost: `SessionStatus` is a widely-consumed union, so every exhaustive `switch` (the codebase has
|
||||
`assertNever` and `noFallthroughCasesInSwitch`) will need a branch. That is a feature, it makes
|
||||
the compiler find every site. This is a **minor** bump, not a patch: it widens a public type in
|
||||
the HTTP contract.
|
||||
|
||||
### Part 4: `GET /api/schema`
|
||||
|
||||
herdr ships `herdr api schema`. Every Codeman route is already Zod-validated, so
|
||||
`zod-to-json-schema` over `schemas.ts` gives a self-describing API almost free. Value: third-party
|
||||
tools and the skill stop drifting from hand-written docs. Open question: whether to emit full
|
||||
OpenAPI (`@fastify/swagger` would need per-route schema registration, which is a much larger
|
||||
change) or just dump the Zod schemas keyed by name (cheap, 80% of the value).
|
||||
|
||||
### Part 5: detection manifests instead of hardcoded patterns
|
||||
|
||||
CLI-specific readiness, blocked and usage-limit patterns live in code across
|
||||
`usage-limit-patterns.ts`, the respawn pattern helpers and `regex-patterns.ts`. Externalizing the
|
||||
per-CLI ones into data files would make adding a sixth CLI a data change instead of a code change.
|
||||
|
||||
**Do not copy the remote-update part.** herdr auto-fetches manifest updates from herdr.dev.
|
||||
Codeman auto-pulling behavioral rules from a vendor server contradicts its security posture.
|
||||
Bundled manifests plus local override only, no network.
|
||||
|
||||
### Explicit non-goals
|
||||
|
||||
- **Plugin runtime and marketplace.** `docs/extending-codeman.md` already argues this: a plugin
|
||||
runtime means third-party code inside a process that spawns agents with your credentials, on a
|
||||
server people expose over a tunnel. The reasoning still holds. If the marketplace _pattern_ is
|
||||
wanted, apply it to data (web tabs, case templates, cron recipes), never to executable code.
|
||||
- **Live PTY handoff on restart.** herdr needs it because it owns the terminals. Codeman
|
||||
delegates to tmux, so PTYs already survive a self-update restart.
|
||||
- **Socket API.** HTTP plus SSE is the existing, documented, stable contract. A second transport
|
||||
would double the surface for no capability gain.
|
||||
|
||||
---
|
||||
|
||||
## 5. Sequencing
|
||||
|
||||
| Step | Work | Gate |
|
||||
| ---- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| 1 ✅ | `src/config/agent-wait.ts` + `session-wait-registry.ts` + unit tests | 48 tests green |
|
||||
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
|
||||
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
|
||||
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
|
||||
| 5 | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
|
||||
| 6 | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | settings partial-PUT test, case-creation test |
|
||||
| 7 | Docs: api-reference, extending-codeman, README | |
|
||||
| 8 | COM (minor bump: new endpoints, new setting, new optional fields) | both CI and Release workflows green |
|
||||
|
||||
Parts 1 and 2 are independent enough to land separately, but the skill is much less useful
|
||||
without the wait endpoints, so the wait work goes first.
|
||||
|
||||
## 6. Open questions for the owner
|
||||
|
||||
1. `skills/` at the repo root, accepted despite the short-root rule? (Recommended yes, the
|
||||
install one-liner depends on it.)
|
||||
2. `agentSkillEnabled` default: OFF for the first release then flip, or ON immediately?
|
||||
3. Auto-inject the skill into every case's `.claude/skills/`, or global install only?
|
||||
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
|
||||
and not a security boundary?
|
||||
5. Regex support in `wait-output`: confirm literal-only for v1.
|
||||
|
||||
---
|
||||
|
||||
## 7. Build log: what actually happened
|
||||
|
||||
Written at the end of the build so the next person inherits the reasoning, not just the
|
||||
diff. Process artifacts (per-agent briefs, findings, reports) live in the gitignored
|
||||
`tmp/agent-wait-review/`; this section is the part worth keeping.
|
||||
|
||||
### What shipped
|
||||
|
||||
| Piece | Files |
|
||||
| ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Bounds + clamping | `src/config/agent-wait.ts` (new) |
|
||||
| Blocking-wait registry | `src/web/session-wait-registry.ts` (new, IO-free, unit-tested) |
|
||||
| `GET .../wait`, `GET .../wait-output`, `wait`/`waitTimeout` on `POST .../input` | `src/web/routes/session-routes.ts` |
|
||||
| Signal wiring | `session-listener-wiring.ts` (idle/working/exit + output), `hook-event-routes.ts` (stop/blocked), `server.ts` (teardown, shutdown) |
|
||||
| Agent skill | `skills/codeman/SKILL.md` + `reference/`, `.claude/skills/codeman` symlink, `package.json` `files` |
|
||||
| Docs | `api-reference.md`, `extending-codeman.md`, `architecture-invariants.md`, `README.md`, `CLAUDE.md` |
|
||||
| Tests | `test/session-wait-registry.test.ts`, three `test/routes/session-*wait*.test.ts`, `http-contract.test.ts`, `mock-session.ts` |
|
||||
|
||||
### Bugs found in ADJACENT code, not in the new feature
|
||||
|
||||
These are the highest-value output of the exercise and none were on the plan:
|
||||
|
||||
1. **Every Codeman hook was dead on HTTPS installs.** `hooks-config.ts` built the hook
|
||||
curl as `curl -s` with no `-k` while the statusline exporter 300 lines below used
|
||||
`curl -sk` and documented why. Proven with the real hook command: `curl exit=60`
|
||||
without the flag, success with it, and the failure swallowed by the hook's own
|
||||
`2>/dev/null || true`. This silently killed `stop`, `permission_prompt`,
|
||||
`elicitation_dialog`, `idle_prompt`, `teammate_idle` and `task_completed`, taking
|
||||
respawn's definitive idle signals with them. Fixed, **plus** a staleness detector in
|
||||
`refreshStaleCodemanHooks` that regenerates the on-disk config of already-created
|
||||
cases (23 of 26 local cases carried the broken form; fixing the generator alone would
|
||||
have left every one of them broken).
|
||||
2. **`buildEnvExports()` exported a wrong-scheme `CODEMAN_API_URL`** (`http://` fallback
|
||||
on an HTTPS install). Now omitted rather than guessed, so in-session guards fail closed.
|
||||
3. **Programmatic input is only submitted when it contains `\r`.** `sendInput` sends Enter
|
||||
only if the payload has a carriage return; without it the text sits in the composer
|
||||
forever. Bit this build repeatedly before it was diagnosed, and had leaked into the
|
||||
docs' own examples.
|
||||
|
||||
### Design decisions worth not re-litigating
|
||||
|
||||
- **A timeout is HTTP 200** with `wait.timedOut`, never a 4xx: callers loop over short
|
||||
waits because tunnels cut idle connections, and every poll boundary would otherwise be
|
||||
indistinguishable from failure.
|
||||
- **Send-and-wait must be one endpoint.** A separate POST-then-wait races: between the
|
||||
write and the flip to `working`, a wait sees the stale `idle` and reports the PREVIOUS
|
||||
turn as this one. The waiter is registered before the write.
|
||||
- **`stop`/`blocked` exist for `claude` mode only.** They come from Claude Code hooks;
|
||||
`shell` installs none either, so keying off `isExternalCliMode()` was wrong.
|
||||
- **Literal matching only, never regex.** JS `RegExp` backtracks; herdr can offer
|
||||
`--regex` because Rust's regex crate is linear-time.
|
||||
- **Client-hangup abort listens on `reply.raw` guarded by `writableFinished`.** On
|
||||
`req.raw`, `close` fires when the request BODY ends, which on a POST killed every
|
||||
send-and-wait instantly, and no `app.inject()` test can see it (inject never emits
|
||||
`close`).
|
||||
- **Liveness cannot come from `session.pid`.** For a tmux session that is the local
|
||||
`tmux attach` client, not the worker: a worker exiting inside its pane leaves
|
||||
`pane_dead=1` with the client alive, so `pid` never goes null. Liveness is probed at
|
||||
the mux layer, cached (~750 ms) and only on blocking waits, never on the input hot path.
|
||||
|
||||
### Verification rounds
|
||||
|
||||
Six agents across three rounds, each verifying the previous round's work rather than its
|
||||
own. Findings that mattered, in order of severity, were: the dead-pane liveness gap; the
|
||||
`reply.raw` abort regression; abandoned long-polls leaking waiter slots; a crashed session
|
||||
reporting `idle`; `shell` accepting `until=stop`; and a documented recipe that reported
|
||||
success without running its task. Two traps recurred often enough to name:
|
||||
|
||||
- **Vacuous passes.** `app.inject()` never emits `close`; a latched `cancelEverything()`
|
||||
in `afterEach` silently killed the registry for every later test in a file; three test
|
||||
files sharing one session id against the process-wide registry let one file's leftover
|
||||
waiter fail another's assertion. Any new wait test needs care on all three.
|
||||
- **HTTP-only test instances.** Every isolated instance used during the build was plain
|
||||
HTTP, which is exactly why the HTTPS hook bug survived so long. Test the transport the
|
||||
user actually runs.
|
||||
|
||||
### Resolved at wrap-up (2026-08-08, conclusion pass)
|
||||
|
||||
- **R2-A**: the fire-and-forget-then-gather-sequentially pattern was **removed from
|
||||
the skill** rather than patched. Signals are edge-triggered with no history, so a
|
||||
`stop` that fires before its waiter registers is unobservable afterwards; a
|
||||
`fresh=0` gather was rejected because the only `until` set that current state can
|
||||
satisfy answers `idle` for a prompt that never submitted, resurrecting the exact
|
||||
false-success failure R2-B had just closed. Flow 3b's pattern B now gathers on
|
||||
latched `wait-output` markers (`from=buffer`), the same mechanism that makes the
|
||||
shell flows reliable; the limitation is recorded in
|
||||
`architecture-invariants#agent-wait-primitives` and `endpoints.md`. The durable
|
||||
fix, a latched last-signal-per-turn on the server, stays with deferred Part 3.
|
||||
- Docs F7/F8, F4 and the false-`idle` attribution: `api-reference.md`,
|
||||
`extending-codeman.md` and `architecture-invariants.md` rewritten to the post-fix
|
||||
matcher (one normalized stream, chunk-straddling found, snippet as a rendering of
|
||||
the matched window), the real no-PTY answer (`ended:true`, `aborted:false`,
|
||||
`delivered:false`), and the startup-idle mechanism (a session parked on the trust
|
||||
dialog emits no further `idle`; the false success is the startup transition).
|
||||
- Orchestrate #12, #5/R2-B, #6, and R2-C..R2-E: fire-and-forget's empty `data`
|
||||
documented; every send-and-wait retry loop now treats `duplicate:true` +
|
||||
`immediate:true` as "no new turn ran" and reads the terminal before believing it;
|
||||
claude fan-out is pattern A (backgrounded send-and-waits) or the marker gather;
|
||||
readiness budgets rebalanced (5 s stage 1, 45 s stage 3) with the virgin-case
|
||||
floor named; the auth fallback now also reads the supervisor definition
|
||||
(`codeman-web.service` / launchd plist) and accepts `export`-prefixed `.env`
|
||||
lines; `pid != null` is documented as startup-only, never liveness.
|
||||
- Both public readiness recipes (extending-codeman.md, README) are bypass-first with
|
||||
the trust probe as the bounded fallback; the worked recipe carries `-k` and fails
|
||||
loudly on an empty SID; the hook `-k`/self-heal fix appears in every
|
||||
"hooks go missing" list; the multi-word-TUI claim is "unreliable", not "never".
|
||||
|
||||
### Still open
|
||||
|
||||
- **Release checklist**: `package.json` `files` includes `skills`, which is still
|
||||
untracked. `git add skills/` must be part of the release commit, or npm publishes
|
||||
a tarball without the skill (a `files` entry that does not exist is silently
|
||||
ignored, so nothing fails).
|
||||
- The 1.13.0 changeset is written under `.changeset/`; consuming it (COM flow),
|
||||
the release commit, and the deploy remain.
|
||||
- Deferred with Part 3: the latched last-signal-per-turn. Nice-to-haves from the
|
||||
reviews: N2 (create the death-watcher inside its `try`) and converting
|
||||
timeout-shaped test detections into fast assertions.
|
||||
@@ -46,6 +46,20 @@ payload return `{ "success": true, "data": {} }`.
|
||||
> `GET /api/screenshots/:name`, `GET /q/:code` (QR redirect), and the
|
||||
> `GET /ws/sessions/:id/terminal` WebSocket upgrade.
|
||||
|
||||
> The [agent wait endpoints](#long-polling-agent-wait) use the normal envelope but
|
||||
> are the only JSON endpoints that deliberately **hold the connection open**, for up
|
||||
> to 600 s. Proxy operators and HTTP clients with a global read timeout need to know
|
||||
> that before pointing them at Codeman.
|
||||
|
||||
⚠️ **A `401` is the one status that is not an envelope.** Authentication is rejected
|
||||
in a request hook, before any handler runs, and it replies with the bare string
|
||||
`Unauthorized` (`Unauthorized: hook secret required` on the hook path) plus
|
||||
`WWW-Authenticate: Basic realm="Codeman"`. There is no `success`, no `error`, and no
|
||||
`errorCode`, because the wrapping hook only wraps object payloads. So a client that
|
||||
pipes every response straight into a JSON parser dies with a parse error rather than
|
||||
reporting an auth failure, which is a confusing way to discover that a password is
|
||||
set. Branch on the HTTP status **before** parsing.
|
||||
|
||||
## Error codes → HTTP status
|
||||
|
||||
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
|
||||
@@ -66,6 +80,333 @@ the HTTP status.
|
||||
|
||||
Adding a new error code is non-breaking; removing or renaming one is a major change.
|
||||
|
||||
## Long-polling (agent wait)
|
||||
|
||||
Three calls block until something happens instead of answering immediately. They
|
||||
exist because SSE is Codeman's only other "tell me when" channel, and an agent
|
||||
driving the API from a shell tool cannot practically hold a stream and parse
|
||||
events inline.
|
||||
|
||||
| Call | Blocks until |
|
||||
|------|--------------|
|
||||
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
|
||||
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
|
||||
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
|
||||
|
||||
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
|
||||
`GET .../wait`. It registers the waiter **before** writing, which closes the window
|
||||
in which a separate wait sees the session still idle from the previous turn and
|
||||
answers instantly with the wrong turn's result. Use it whenever you send a prompt
|
||||
and want to know when that prompt is done.
|
||||
|
||||
### Three semantics that break callers who assume otherwise
|
||||
|
||||
**1. A timeout is HTTP `200`, not an error.** A wait that ends without its signal
|
||||
returns `{"success":true, ...,"wait":{"timedOut":true,"signal":null}}`. The
|
||||
intended pattern is a client-side loop over short waits, because `tailscale serve`
|
||||
and cloudflared can both cut an idle connection, and turning every poll boundary
|
||||
into a `4xx` would make that loop indistinguishable from a real failure. `408` is
|
||||
auto-retried by several clients (silently doubling the polling load), `504` is what
|
||||
a genuine tunnel failure looks like, and `204` cannot carry `waitedMs` / `status` /
|
||||
`limitPaused`. Reserve error handling for the four codes in the table below.
|
||||
|
||||
**2. `stop` and `blocked` fire only for `claude` sessions.** Both come from Claude
|
||||
Code hooks, and no other mode installs them: `shell` runs no agent, and the external
|
||||
CLIs (`opencode`, `codex`, `gemini`, `antigravity`) render their own TUIs and post
|
||||
no hooks. For every non-`claude` mode only `idle`, `working` and `exit` are
|
||||
accepted, and of those only `exit` is dependable: see the caveats under
|
||||
[Signals](#signals) before building on `idle`. Requesting `stop` or `blocked`
|
||||
**explicitly** on such a session is a
|
||||
`400`; omitting `until` never fails, the server just drops them from the default set
|
||||
and echoes the narrowed set back as `wait.until`. Three more places hooks can go
|
||||
missing even in `claude` mode: a **Docker case** needs
|
||||
`CODEMAN_DOCKER_BRIDGE_HOOKS=1`, since a container cannot reach a loopback-bound
|
||||
Codeman (without it, only `idle` / `working` / `exit` work); a **remote-SSH
|
||||
case** runs the agent on another host, whose hooks may never reach this server at
|
||||
all; and a case whose hook config was written by **Codeman < 1.13.0 against an
|
||||
`--https` install** carries hook curls without `-k`, which TLS-fail silently (the
|
||||
hook line ends in `|| true`). Codeman now writes `curl -sk` and repairs a stale
|
||||
case config the next time a session starts in that case. When in doubt, ask for
|
||||
`stop,idle,exit` so a session without hooks still resolves on the heuristic
|
||||
signal.
|
||||
|
||||
**3. `from=now` does not mean "printed after you asked".** tmux repaints the visible
|
||||
screen on attach, on resize, and on any TUI redraw, and a repaint arrives as
|
||||
ordinary output, so text that was already on screen can satisfy a fresh wait. This
|
||||
was observed live: a marker echoed a minute earlier matched instantly on a new
|
||||
`from=now` wait. It is inherent to running the agent under a multiplexer, so the
|
||||
contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
|
||||
`echo $MARK`, then wait on `$MARK`), never a generic string like `BUILD OK`.
|
||||
|
||||
### Signals
|
||||
|
||||
| Signal | Source | Actually fires for |
|
||||
|--------|--------|--------------------|
|
||||
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
|
||||
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
|
||||
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
|
||||
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
|
||||
| `exit` | no process is behind the session | every mode |
|
||||
|
||||
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
|
||||
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
|
||||
is omitted is `stop,idle,exit` (`exit` is in there so a worker that crashes resolves
|
||||
the wait promptly instead of burning the caller's whole timeout on something that
|
||||
can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit`
|
||||
once the session is up: the default set's `idle` also resolves on a spinner pause,
|
||||
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
|
||||
can land inside your first wait window and report a turn that never ran. Measured:
|
||||
a session parked on the trust dialog emits no *further* `idle`, so it is the
|
||||
startup transition, not the dialog, that produces the false success below.
|
||||
|
||||
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
|
||||
server answers from `pid === null` plus a mux-layer pane-death probe, and that
|
||||
covers a session that exited — including a worker that died *inside* its tmux pane
|
||||
while the local attach client (and therefore `pid`) lives on — one that was
|
||||
detached, and one that was **created but never started**. So the first wait
|
||||
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
|
||||
milliseconds, and reading that as "the worker died" is wrong: it means start it, or
|
||||
wait for it to come up. `status` is carried alongside so nothing is hidden. The
|
||||
alternative (trusting `status`) is worse, because a dead PTY parks the session at
|
||||
`status: "idle"`, which would answer the default wait with `immediate: true` for a
|
||||
worker that has crashed. A worker dying while a wait is parked resolves it within
|
||||
a few seconds (a background death-watcher), not at the timeout.
|
||||
|
||||
⚠️ **`blocked` is reachable less often than the table suggests.** It fires on two
|
||||
hooks, and the default configuration suppresses one of them: Codeman spawns claude
|
||||
with `--dangerously-skip-permissions`, so permission prompts do not happen unless the
|
||||
instance is switched to the `auto` Claude mode (App Settings), or the caller is a
|
||||
multi-user account without the bypass grant, which is forced to `--permission-mode
|
||||
auto`. What does still fire under the default is `elicitation_dialog`, the agent
|
||||
asking the user a question. So `until=stop,blocked,exit` is a reasonable belt on a
|
||||
long turn, but a worker that never comes back is far more likely to be working than
|
||||
blocked, and polling `blocked` alone will sit at its timeout.
|
||||
|
||||
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
|
||||
session emits its one `idle` at startup and then stays `status: "idle"` forever,
|
||||
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
|
||||
requires a transition (and so does `fresh=1`), both can only time out there:
|
||||
a documented default `wait` on a shell worker running `sleep 4` times out at the
|
||||
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
|
||||
instead. The same caution applies to the external CLIs.
|
||||
|
||||
### Readiness is not a signal
|
||||
|
||||
Nothing here reports "the agent is ready for a prompt", and no combination of
|
||||
`until`/`fresh` synthesizes one. A freshly created session reads as `exit` (above),
|
||||
and a `claude` worker in a brand-new case comes up on the CLI's **trust dialog**,
|
||||
which contains a ❯ prompt of its own. Send-and-wait posted at that moment types the
|
||||
prompt into the dialog, where the `\r` never gets past it, while the session's
|
||||
startup `idle` lands inside the wait window: the wait resolves on `idle` in a
|
||||
couple of seconds with `timedOut: false`, which looks exactly like a completed
|
||||
turn.
|
||||
|
||||
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
|
||||
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
|
||||
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
|
||||
dialog only as the bounded fallback (`trust` matched → send `\r` → wait for
|
||||
`bypass` again). Do not probe `trust` first and Enter blindly: the dialog text
|
||||
stays in the terminal buffer for the life of the session, so a `trust` probe with
|
||||
`from=buffer` keeps matching on every later run and the Enter lands in a ready
|
||||
composer. A worked version is in
|
||||
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
|
||||
|
||||
### `GET /api/v1/sessions/:id/wait`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
|-------|------|---------|-------|
|
||||
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
|
||||
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
|
||||
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
|
||||
|
||||
```bash
|
||||
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
|
||||
```
|
||||
|
||||
Both GET wait routes answer with `Cache-Control: no-store`, because the documented
|
||||
pattern polls one identical URL in a loop and a cached `{"timedOut":true}` would
|
||||
turn that loop into a busy spin. `POST .../input` sends no cache header (it is a
|
||||
POST, which is not heuristically cacheable).
|
||||
|
||||
⚠️ **Unknown query parameters are ignored, not rejected**, with one exception
|
||||
(`regex`, below). In particular `match=` on `/wait` is silently dropped and you get
|
||||
a plain signal wait, so check the endpoint path before blaming the parameters.
|
||||
|
||||
### `GET /api/v1/sessions/:id/wait-output`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
|-------|------|---------|-------|
|
||||
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
|
||||
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
|
||||
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
|
||||
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
|
||||
|
||||
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
|
||||
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
|
||||
instead of waiting on the wrong thing. The reasoning is in
|
||||
[`architecture-invariants.md`](architecture-invariants.md#agent-wait-primitives).
|
||||
|
||||
#### What the matcher actually sees
|
||||
|
||||
The matcher scans the raw PTY stream, **normalized**: ANSI escape sequences are
|
||||
stripped — CSI, OSC, and the charset-designation escapes a stock bash prompt emits
|
||||
on every line (`ESC ( B`), so `match=tnode:` matches a prompt that renders
|
||||
`…@tnode:` — a partial escape arriving at a chunk boundary is held back until its
|
||||
tail arrives, and a match may straddle PTY chunks: `printf STRAD; sleep 1; printf
|
||||
DLEQQ` is matchable as `STRADDLEQQ` (all measured live). Three caveats remain:
|
||||
|
||||
⚠️ **It is still the byte stream, not the rendered pane.** `GET .../terminal`
|
||||
answers from a tmux screen capture (`data.source: "mux-visible"`), the finished
|
||||
picture; the matcher sees the stream that painted it. For linear output the two
|
||||
agree once escapes are stripped, but a full-screen TUI composes its picture with
|
||||
cursor positioning, so what the pane shows and what the stream carries can differ.
|
||||
Seeing your string in `terminal?tail=` makes a match likely, not guaranteed.
|
||||
|
||||
⚠️ **A TUI's text can arrive without its spaces.** Claude Code positions words
|
||||
with cursor moves rather than printing spaces, so screen text can reach the
|
||||
matcher as `Quicksafetycheck:Isthisaprojectyoucreated...`. Whether a given phrase
|
||||
keeps its spaces depends on how the TUI happened to draw it (measured: `I trust
|
||||
this folder` matched, `Quick safety check` did not), so a multi-word `match`
|
||||
against a TUI pane is unreliable rather than impossible. Match a **single
|
||||
space-free token**, ideally one you printed yourself. Plain command output (a
|
||||
shell worker, an `echo`) keeps its spaces.
|
||||
|
||||
⚠️ **The returned `snippet` is a rendering of the matched text, not a quotation of
|
||||
it.** It is cut from the same normalized stream the match ran against, then
|
||||
cleaned for display: remaining raw control bytes are removed (an agent pipes the
|
||||
snippet into its own terminal, so a worker's bytes must not be able to reset that
|
||||
display) and blank runs are collapsed. A printable needle that matched will appear
|
||||
in it; a needle containing control bytes or a blank run may not survive verbatim.
|
||||
|
||||
```bash
|
||||
MARK="DONE_$RANDOM"
|
||||
curl -sG "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode "match=$MARK" --data-urlencode 'timeout=120000'
|
||||
```
|
||||
|
||||
Build the query with `-G --data-urlencode` rather than by hand: a `+` in a
|
||||
hand-written query string decodes to a space.
|
||||
|
||||
### `POST /api/v1/sessions/:id/input` with `wait`
|
||||
|
||||
Two optional fields on the existing endpoint:
|
||||
|
||||
| Field | Type | Notes |
|
||||
|-------|------|-------|
|
||||
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
|
||||
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
|
||||
|
||||
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
|
||||
"absent" rather than failing validation. That is deliberate: `.optional()` would
|
||||
reject it, which has shipped as a real bug twice.
|
||||
|
||||
The input must end with `\r` (a real carriage return in the JSON string): Enter is
|
||||
sent only when the input contains one, so text without it is typed onto the
|
||||
worker's prompt but never submitted, and the wait then runs its full timeout on a
|
||||
turn that never started. Verified live; this is the most common silent failure on
|
||||
this endpoint.
|
||||
|
||||
```bash
|
||||
curl -s -X POST "$API/api/v1/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"run the tests\r","useMux":true,"clientId":"agent-1","seq":1,
|
||||
"wait":"stop","waitTimeout":600000}'
|
||||
```
|
||||
|
||||
A **tagged duplicate** (a `clientId` + `seq` pair the server has already applied)
|
||||
still honors `wait`, because the caller's question is unanswered, but it answers
|
||||
from the session's current state rather than requiring a new transition: the
|
||||
original turn may be long over. It comes back as
|
||||
`"delivered": false, "duplicate": true`.
|
||||
|
||||
### Response
|
||||
|
||||
All three nest the wait result under `data.wait`, so one client helper works against
|
||||
any of them:
|
||||
|
||||
```json
|
||||
{ "success": true, "data": {
|
||||
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
|
||||
"status": "idle",
|
||||
"limitPaused": false,
|
||||
"wait": {
|
||||
"signal": "stop", "until": ["stop", "idle", "exit"],
|
||||
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
|
||||
"waitedMs": 8421, "timeoutMs": 60000
|
||||
}
|
||||
}}
|
||||
```
|
||||
|
||||
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
|
||||
`status` and `limitPaused`. `POST .../input` **without** `wait` is unchanged and
|
||||
still returns `{"success": true, "data": {}}`.
|
||||
|
||||
⚠️ `delivered: false` has **two** meanings, and they must be told apart by
|
||||
`duplicate`: with `duplicate: true` the input was suppressed as an already-applied
|
||||
redelivery (harmless, the turn it refers to may be long over), while with
|
||||
`duplicate: false` the **write failed** (typically no PTY behind the session). A
|
||||
client that reads `delivered === false` as "duplicate" silently treats a failed send
|
||||
as a success.
|
||||
|
||||
| Field | Type | Meaning |
|
||||
|-------|------|---------|
|
||||
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
|
||||
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
|
||||
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
|
||||
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
|
||||
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
|
||||
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
|
||||
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
|
||||
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
|
||||
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
|
||||
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
|
||||
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
|
||||
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
|
||||
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
|
||||
|
||||
Read the outcome by discriminator, in this order:
|
||||
|
||||
1. `wait.signal !== null` (or `wait.matched === true`): the thing happened.
|
||||
2. `wait.timedOut`: a poll boundary. Loop again.
|
||||
3. `wait.ended` or `wait.aborted`: the wait answered nothing, because the session is
|
||||
gone or was never running. Re-check the session instead of looping.
|
||||
|
||||
`wait.immediate` is not a fourth outcome: it rides along with the first one and
|
||||
means the condition already held at call time, so nothing was actually waited for.
|
||||
If that is not what you meant, you wanted `fresh=1` or the send-and-wait form. Note
|
||||
that `{"signal":"exit","immediate":true}` on a session you just created is the
|
||||
not-started-yet case, not a crash.
|
||||
|
||||
**The timeout is clamped, so read it back.** A request for 1800000 ms is silently
|
||||
reduced to the server's ceiling (600000 ms by default, operator-tunable), and a
|
||||
request for 1 ms is raised to 1000 ms. `wait.timeoutMs` is the value that was
|
||||
applied. Without checking it, a caller that asked for 30 minutes and got 10 will
|
||||
read the timeout as "the worker is wedged" and kill a session that was working fine.
|
||||
|
||||
### Errors
|
||||
|
||||
| `errorCode` | HTTP | When |
|
||||
|-------------|------|------|
|
||||
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
|
||||
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
|
||||
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
|
||||
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
|
||||
|
||||
The two capacity codes are deliberately different. A process-wide cap reported as
|
||||
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
|
||||
error message names the cap that was hit.
|
||||
|
||||
⚠️ A `401` is **not** in this table and is not an envelope at all (see
|
||||
[Response envelope](#response-envelope)). It matters most here: a polling loop that
|
||||
pipes each wait straight into `jq` fails with a parse error on every iteration
|
||||
against a password-protected server, which reads as "the wait endpoints are broken".
|
||||
Check the status first.
|
||||
|
||||
The per-session cap is a **combined** budget: signal waiters and output waiters
|
||||
count against the same 16, not 16 of each. An abandoned request no longer holds its
|
||||
slot, because the routes release the waiter when the client disconnects, but a
|
||||
client that opens many concurrent waits against one session will still hit the cap.
|
||||
|
||||
## Authentication
|
||||
|
||||
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -149,9 +149,17 @@ to Claude as a system reminder. This implies `"async": true`; ordinary async
|
||||
hooks do not wake an idle turn, and their output waits for the next interaction.
|
||||
|
||||
Codeman uses this on `PostToolUse(Bash)`: a self-contained Node helper extracts
|
||||
the background task ID from the Bash result, watches the session transcript for
|
||||
the matching completion notification, and exits 2. It does not send terminal
|
||||
input, so it cannot submit a user's partially written prompt.
|
||||
the background task ID from the Bash result, watches the originating transcript
|
||||
and, for subagents, the top-level parent transcript for the matching completion
|
||||
notification, and exits 2. Claude records a subagent's Bash result in its
|
||||
`subagents/agent-*.jsonl` file but queues completion in the lead session JSONL.
|
||||
The task ID keeps each wake targeted. The helper does not send terminal input,
|
||||
so it cannot submit a user's partially written prompt.
|
||||
|
||||
For script-dispatched Codex work, `codex-run.sh` writes the final response
|
||||
between `CODEMAN_RESULT_BEGIN/END` markers in the background task output. The
|
||||
rewake helper includes a maximum of 64 KiB of that report in its feedback. UI
|
||||
subagent discovery and dispatcher result delivery are separate contracts.
|
||||
|
||||
### Notification
|
||||
|
||||
@@ -219,6 +227,16 @@ Or to allow exit:
|
||||
|
||||
**Use Cases**: Control nested loops, verify subagent output.
|
||||
|
||||
The hook input includes `agent_id`, `agent_transcript_path`, and
|
||||
`last_assistant_message`. Like `Stop`, a command hook can return
|
||||
`{"decision":"block","reason":"..."}` to keep the subagent running and feed
|
||||
the reason back to it.
|
||||
|
||||
Codeman uses this to prevent premature reports from workers that still own live
|
||||
Monitor or background-Bash processes. It derives candidate task IDs from the
|
||||
subagent transcript, but requires a matching live Linux process descriptor for
|
||||
`tasks/<id>.output`; historical task text by itself is not treated as active.
|
||||
|
||||
### TeammateIdle
|
||||
|
||||
**When**: When an agent-team teammate is about to go idle.
|
||||
|
||||
@@ -44,9 +44,9 @@ records), kept distinct from the existing `ScheduledRun`.
|
||||
|
||||
## 2. Where agent/session types are defined
|
||||
|
||||
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini'`
|
||||
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity'`
|
||||
(`src/types/session.ts:43-44`). `shell` covers the brief's "Terminal/custom".
|
||||
- CLI availability resolvers in `src/utils/{claude,codex,gemini,opencode}-cli-resolver.ts`.
|
||||
- CLI availability resolvers in `src/utils/{claude,codex,gemini,antigravity,opencode}-cli-resolver.ts`.
|
||||
- **Integration point:** the job's `agentType` reuses `SessionMode` verbatim.
|
||||
|
||||
## 3. Where input is sent into a session
|
||||
|
||||
+2
-2
@@ -1,7 +1,7 @@
|
||||
# Cron Jobs — User & Operator Guide
|
||||
|
||||
Codeman's **Cron** feature lets you save named, recurring jobs that automatically
|
||||
spin up a Claude (or shell / OpenCode / Codex / Gemini) session on a schedule and
|
||||
spin up a Claude (or shell / OpenCode / Codex / Antigravity / Gemini) session on a schedule and
|
||||
feed it a prompt. Think "cron for agent sessions": _"every weekday at 3am, open a
|
||||
Claude session in `~/proj` and tell it to update dependencies and open a PR."_
|
||||
|
||||
@@ -91,7 +91,7 @@ These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
|
||||
| Field | Required | Values / limits | Notes |
|
||||
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
|
||||
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
|
||||
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
|
||||
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
|
||||
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
|
||||
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
|
||||
|
||||
+16
-1
@@ -2,7 +2,7 @@
|
||||
|
||||
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
|
||||
|
||||
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` all work inside the container.
|
||||
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` all work inside the container.
|
||||
|
||||
## One-time setup: build the base image
|
||||
|
||||
@@ -15,6 +15,21 @@ node scripts/build-agent-image.mjs # builds codeman/agent:base
|
||||
|
||||
The image is **secret-free**: credentials are delivered at runtime (bind mounts or `docker exec --env`), never baked in, so exports never leak them.
|
||||
|
||||
⚠️ **Re-build with `--no-cache`, always.** The CLIs are installed in a single `RUN npm install -g` layer, so a plain rebuild re-uses it from the Docker layer cache and the CLIs stay frozen at whatever versions the image was **first** built with, however long ago that was. Editing the Dockerfile does not help unless the edit lands at or above that line: a change appended below it leaves the npm layer cached and only runs the new step. Observed 2026-08-06: a rebuild silently kept a stale `@openai/codex@0.144.6` whose aliased platform binary had not installed, so every `codex` docker case died with `Missing optional dependency @openai/codex-linux-x64` while the build itself reported success.
|
||||
|
||||
```bash
|
||||
node scripts/build-agent-image.mjs --no-cache
|
||||
```
|
||||
|
||||
A zero exit code only proves the layers ran, not that the toolchain works. Verify by actually executing each CLI in the image, and check the build log for `Using cache` lines:
|
||||
|
||||
```bash
|
||||
docker run --rm codeman/agent:base bash -lc \
|
||||
'for c in claude codex gemini opencode agy; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
|
||||
```
|
||||
|
||||
Antigravity (`agy`) is the one CLI not installed from npm (Google ships a standalone binary), so it has its own Dockerfile step and adds roughly 190MB; a full image lands near 1.6GB.
|
||||
|
||||
## Quickest path: one-click "Run in Docker"
|
||||
|
||||
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
|
||||
|
||||
@@ -0,0 +1,410 @@
|
||||
# Extending Codeman
|
||||
|
||||
Codeman has no plugin runtime, and that is a deliberate choice rather than a
|
||||
missing feature. A plugin runtime means running third-party code inside a process
|
||||
that spawns agents with your credentials, on a server people routinely expose
|
||||
over a tunnel or Tailscale. Codeman's security model is one of its reasons to
|
||||
exist, so it does not hand that away for an extension mechanism.
|
||||
|
||||
Instead there are four seams that already work, from any language, with nothing
|
||||
installed:
|
||||
|
||||
| You want to | Use | Runs where |
|
||||
| --- | --- | --- |
|
||||
| Show your own UI inside Codeman | [Web tabs](#seam-1-web-tabs) | Your own process, rendered as a tab |
|
||||
| React when an agent needs you | [SSE events](#seam-2-sse-events) | Anywhere that can hold an HTTP connection |
|
||||
| Drive Codeman from a script | [HTTP API](#seam-3-http-api-and-cli) or the `codeman` CLI | Anywhere |
|
||||
| React inside a Claude session | [Hooks](#seam-4-hooks) | The agent's own machine |
|
||||
|
||||
Everything below is covered by the stability promise in
|
||||
[`versioning-policy.md`](versioning-policy.md): endpoint paths, the response
|
||||
envelope, `errorCode` values, and SSE event names are stable. Additive changes
|
||||
(new endpoints, new optional fields, new events) are non-breaking. Breaking
|
||||
changes ship under a new prefix (`/api/v2`).
|
||||
|
||||
## Before you start
|
||||
|
||||
**Base URL.** `http://127.0.0.1:3000` by default. Prefer the versioned prefix
|
||||
`/api/v1/...` for anything you publish; the unversioned `/api/...` is an alias.
|
||||
|
||||
**Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic on every request, or
|
||||
authenticate once and keep the `codeman_session` cookie. With no password set,
|
||||
Codeman is loopback-only and unauthenticated.
|
||||
|
||||
```bash
|
||||
curl -u admin:$CODEMAN_PASSWORD http://127.0.0.1:3000/api/v1/sessions
|
||||
```
|
||||
|
||||
**Envelope.** Every response is `{"success": true, "data": ...}` or
|
||||
`{"success": false, "error": "...", "errorCode": "..."}`. Check the HTTP status
|
||||
or `body.success`, then read `body.data`. The full `errorCode` to status mapping
|
||||
is in [`api-reference.md`](api-reference.md).
|
||||
|
||||
⚠️ A few legacy GETs (`/api/away-digest` among them) return a bare-ish body with
|
||||
the payload at the top level rather than under `data`. Read defensively with
|
||||
`body.data ?? body`.
|
||||
|
||||
⚠️ A `401` is not an envelope at all: auth is rejected in a request hook that
|
||||
replies with the bare string `Unauthorized`, so parsing it as JSON throws. Branch on
|
||||
the status code before you parse, or a missing password looks like a broken endpoint.
|
||||
|
||||
**Already driving Codeman from an agent?** The README's
|
||||
[Programmatic Guide](../README.md#driving-codeman-from-an-agent--programmatic-guide)
|
||||
covers the in-session case: the `CODEMAN_MUX`, `CODEMAN_API_URL`,
|
||||
`CODEMAN_SESSION_ID` and `CODEMAN_HOOK_SECRET_FILE` variables that let a CLI
|
||||
running inside Codeman find the API and avoid acting on itself. This page is for
|
||||
code running *outside* a session.
|
||||
|
||||
## Seam 1: Web tabs
|
||||
|
||||
The highest-leverage seam. Any web app you can serve locally becomes a tab beside
|
||||
your agent sessions. You write a normal web page; Codeman handles embedding it.
|
||||
|
||||
```bash
|
||||
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/webviews \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"name":"My Dashboard","url":"http://127.0.0.1:8787","icon":"📊"}'
|
||||
```
|
||||
|
||||
Fields: `name` (1 to 60 chars), `url`, and optionally `icon` (a single glyph, max
|
||||
8 code units), `embedMode` (`proxy` by default, or `direct`), and `trusted`.
|
||||
|
||||
Related endpoints: `GET /api/v1/webviews`, `PATCH /api/v1/webviews/:id`,
|
||||
`DELETE /api/v1/webviews/:id`, `POST /api/v1/webviews/probe` (reachability and
|
||||
framing check), `POST /api/v1/webviews/:id/open`.
|
||||
|
||||
### Why it is proxied
|
||||
|
||||
By default your page is served through Codeman's own origin at `/webview/:cap/*`
|
||||
rather than framed directly. A direct iframe fails three ways at once: production
|
||||
is HTTPS so `http://` targets are blocked as mixed content, many dashboards send
|
||||
`X-Frame-Options: DENY`, and Codeman's own `default-src 'self'` CSP blocks
|
||||
cross-origin frames. Proxying solves all three without weakening the CSP.
|
||||
|
||||
### The two things that will confuse you
|
||||
|
||||
A proxied frame is sandboxed and therefore **opaque-origin** unless you set
|
||||
`trusted: true`. Two consequences look like bugs in your own app:
|
||||
|
||||
1. **Root-absolute URLs built at runtime** (`/assets/x.png` assembled in JS)
|
||||
escape the injected `<base>` tag. Codeman injects a `runtimeUrlShim()` that
|
||||
patches the common DOM sinks, but if you construct URLs in an unusual way,
|
||||
prefer relative paths.
|
||||
2. **Same-host `fetch` and `XHR` are CORS-checked with `Origin: null`.** Codeman
|
||||
handles this with `buildProxyCorsHeaders()`, and the proxy is exempt from the
|
||||
global `OPTIONS` short-circuit. If you see "Failed to fetch" while the page
|
||||
itself renders fine, this is the area to look at.
|
||||
|
||||
⚠️ `trusted: true` opts out of the sandbox. A proxied page is served from
|
||||
Codeman's origin, so `allow-same-origin` lets it read the Codeman page and call
|
||||
the API that spawns agents. Only mark your own trusted code.
|
||||
|
||||
## Seam 2: SSE events
|
||||
|
||||
`GET /api/v1/events` is a Server-Sent Events stream. Each message is
|
||||
`event: <name>` plus `data: <json>`. There are 149 event names following a
|
||||
`domain:action` convention, registered in `src/web/sse-events.ts`.
|
||||
|
||||
The ones most integrations want:
|
||||
|
||||
| Event | Meaning |
|
||||
| --- | --- |
|
||||
| `session:created`, `session:deleted` | A session appeared or went away |
|
||||
| `session:idle` | The agent stopped working |
|
||||
| `session:completion` | A completion message was detected |
|
||||
| `session:exit`, `session:error` | The session ended or failed |
|
||||
| `hook:permission_prompt` | The agent is asking for permission |
|
||||
| `hook:idle_prompt`, `hook:stop` | The agent is waiting on you, or stopped |
|
||||
| `hook:task_completed`, `task:completed` | Work finished |
|
||||
| `subagent:discovered`, `subagent:completed` | Background agent lifecycle |
|
||||
| `mux:died` | A multiplexer session died unexpectedly |
|
||||
| `cron:runCreated`, `cron:runUpdated` | Scheduled job activity |
|
||||
|
||||
### Filtering
|
||||
|
||||
`?sessions=id1,id2` suppresses only the high-volume `session:terminal` stream for
|
||||
sessions you did not list. Lifecycle and metadata events are always delivered, so
|
||||
you cannot accidentally filter away the thing you are listening for.
|
||||
|
||||
Pass `?clientId=<uuid>` to enable live filter updates through
|
||||
`POST /api/v1/events/subscribe` without reconnecting the stream.
|
||||
|
||||
### Example: notify when any agent needs you
|
||||
|
||||
```js
|
||||
const res = await fetch('http://127.0.0.1:3000/api/v1/events', {
|
||||
headers: { Authorization: 'Basic ' + btoa(`admin:${process.env.CODEMAN_PASSWORD}`) },
|
||||
});
|
||||
const reader = res.body.getReader();
|
||||
const decoder = new TextDecoder();
|
||||
let buf = '';
|
||||
const WANTED = new Set(['hook:permission_prompt', 'hook:idle_prompt', 'session:idle']);
|
||||
|
||||
for (;;) {
|
||||
const { value, done } = await reader.read();
|
||||
if (done) break;
|
||||
buf += decoder.decode(value, { stream: true });
|
||||
const frames = buf.split('\n\n');
|
||||
buf = frames.pop() ?? '';
|
||||
for (const frame of frames) {
|
||||
const name = frame.match(/^event: (.+)$/m)?.[1];
|
||||
const data = frame.match(/^data: (.+)$/m)?.[1];
|
||||
if (name && WANTED.has(name)) notify(name, JSON.parse(data ?? '{}'));
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Seam 3: HTTP API and CLI
|
||||
|
||||
Around 200 handlers across 21 route files cover sessions, cases, files, cron,
|
||||
respawn, Ralph, the orchestrator, search, and admin. Each route module carries an
|
||||
`@fileoverview` describing its endpoints.
|
||||
|
||||
The common ones:
|
||||
|
||||
```bash
|
||||
# List sessions (live + persisted + transcript history, deduped)
|
||||
curl -u admin:$PASS http://127.0.0.1:3000/api/v1/sessions/unified
|
||||
|
||||
# Create a session
|
||||
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"workingDir":"/home/me/project","mode":"claude"}'
|
||||
|
||||
# Send a prompt (single-line only, and it must end with \r: Enter is sent only
|
||||
# when the input contains a carriage return; without it the text sits on the
|
||||
# session's prompt unsubmitted)
|
||||
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions/$ID/input \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"run the tests\r","useMux":true}'
|
||||
```
|
||||
|
||||
`POST .../input` also accepts `clientId` (stable per client, max 128 chars) and
|
||||
`seq` (monotonic per session). Send both and the server applies each pair
|
||||
at-most-once, so retrying after a dropped connection cannot type the prompt
|
||||
twice. Omit them entirely rather than sending `null`.
|
||||
|
||||
It also accepts `wait` and `waitTimeout`, which hold the response open until the
|
||||
session finishes the turn you just started. `wait` is `true` (the default signal
|
||||
set) or a comma list of `idle,working,stop,blocked,exit`; the result comes back
|
||||
under `data.wait`. Sending them changes nothing for callers that do not: without
|
||||
`wait` the response is still `{"success": true, "data": {}}` and the write is still
|
||||
fire-and-forget. The two interact with `clientId` / `seq` in one way worth knowing:
|
||||
a **tagged duplicate** (a pair the server already applied) skips the write but still
|
||||
waits, answering from the session's current state rather than blocking for a
|
||||
transition that already happened. It reports `"delivered": false, "duplicate": true`.
|
||||
|
||||
### Waiting instead of polling
|
||||
|
||||
Three calls block until something happens: `GET /api/v1/sessions/:id/wait` (a
|
||||
lifecycle signal), `GET /api/v1/sessions/:id/wait-output` (a literal string in the
|
||||
output), and the `wait` field above. Full parameter and response tables are in
|
||||
[`api-reference.md`](api-reference.md#long-polling-agent-wait). Four things decide
|
||||
whether your integration works, and the last one is what actually bites:
|
||||
|
||||
- **A timeout is a `200` with `wait.timedOut: true`**, not an error. Loop over short
|
||||
waits rather than issuing one long one, because `tailscale serve` and cloudflared
|
||||
both cut idle connections and a single 10-minute call is the pattern most likely
|
||||
to die in the field.
|
||||
- **`wait.timeoutMs`** is the timeout after server-side clamping (600 s ceiling by
|
||||
default). Read it rather than assuming you got what you asked for.
|
||||
- **`stop` and `blocked` only exist for `claude` sessions**, and on a `shell` session
|
||||
even `idle` fires only once at startup, so send-and-wait there can only time out.
|
||||
See the Gotchas below.
|
||||
|
||||
⚠️ **There is no readiness signal, and skipping readiness is the failure that looks
|
||||
like success.** A session reports `idle` before its CLI has spawned, and a `claude`
|
||||
worker in a brand-new case comes up on the CLI's **trust dialog**, which has a ❯
|
||||
prompt of its own. Prompt it at that moment and the text lands in the dialog, the
|
||||
`\r` does not get past it, and the session's startup `idle` lands inside the wait
|
||||
window: the wait resolves on `idle` in a couple of seconds with `timedOut: false`,
|
||||
indistinguishable from a finished turn. Wait for the pid, then wait for the
|
||||
composer, answering the dialog only as the bounded fallback.
|
||||
|
||||
A worked orchestration: start a worker, get it ready, prompt it, wait, clean up.
|
||||
|
||||
```bash
|
||||
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}" # auto-set in-session, correct scheme included
|
||||
AUTH=(-u "admin:$CODEMAN_PASSWORD") # omit entirely if no password is set
|
||||
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on --https installs (self-signed cert)
|
||||
|
||||
# 1. Start a worker session (creates the case if it does not exist yet).
|
||||
# The guard matters: a TLS or auth failure otherwise leaves SID empty and every
|
||||
# later step "succeeds" against nothing.
|
||||
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
|
||||
[ -n "$SID" ] && [ "$SID" != null ] || { echo "quick-start failed"; exit 1; }
|
||||
|
||||
# 2. READINESS: composer marker first, trust dialog only as the bounded fallback.
|
||||
# Skip this and step 3 reports a turn that never ran. Do NOT probe trust first
|
||||
# and Enter blindly: the dialog text stays in the buffer for the life of the
|
||||
# session, so on every later run that probe matches stale text and the Enter
|
||||
# lands in a ready composer. Match single tokens only: TUI text can arrive
|
||||
# without its spaces. Stage 1 is short on purpose (an already-trusted case
|
||||
# matches in <1 s; a first-run case can never pass it and pays it in full).
|
||||
until [ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ]
|
||||
do sleep 1; done
|
||||
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
|
||||
--data-urlencode 'timeout=5000') # composer's status bar = ready
|
||||
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
|
||||
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' \
|
||||
--data-urlencode 'timeout=2000')
|
||||
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' -d '{"input":"\r","useMux":true}' >/dev/null
|
||||
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
|
||||
--data-urlencode 'timeout=45000' >/dev/null
|
||||
fi
|
||||
|
||||
# 3. Send the prompt AND register the wait in one call, so the answer cannot be
|
||||
# the previous turn's idle state. Single line only, ending in \r (otherwise
|
||||
# Enter is never sent and this wait times out on a turn that never started).
|
||||
W=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"Run the test suite and summarize the failures\r","useMux":true,
|
||||
"clientId":"orchestrator","seq":1,"wait":"stop,exit","waitTimeout":60000}' \
|
||||
| jq -c '.data.wait')
|
||||
|
||||
# 4. That first wait probably timed out (60 s). Keep going in SHORT waits.
|
||||
for _ in $(seq 1 30); do
|
||||
[ "$(jq -r '.timedOut' <<<"$W")" = 'true' ] || break # signal fired, or wait ended
|
||||
W=$("${CURL[@]}" \
|
||||
"$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq -c '.data.wait')
|
||||
done
|
||||
jq -r 'if .ended or .aborted then "worker is not running"
|
||||
elif .timedOut then "still working after 30 waits"
|
||||
else "signal: \(.signal)" end' <<<"$W"
|
||||
|
||||
# 5. Read what it produced, then delete the session YOU created, by exact id.
|
||||
# ⚠️ NOT /output: its textOutput is empty for every tmux-backed session.
|
||||
# `tail` counts BYTES, and the payload is terminal data with ANSI in it.
|
||||
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
|
||||
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
|
||||
```
|
||||
|
||||
Waiting on a marker instead of a signal is the form that works in **every** mode,
|
||||
and the only one that works on a `shell` session:
|
||||
|
||||
```bash
|
||||
# ⚠️ Split the marker so the typed line never contains it: your own keystrokes echo
|
||||
# into the output stream, so an unsplit marker matches before the command has run.
|
||||
# `from=buffer` also catches a marker that printed before the wait registered.
|
||||
N=$RANDOM
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
|
||||
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
|
||||
--data-urlencode 'timeout=60000' | jq '.data.wait'
|
||||
```
|
||||
|
||||
For shell scripting, the `codeman` CLI is the same surface without the HTTP
|
||||
plumbing:
|
||||
|
||||
```
|
||||
codeman session start|stop|list|logs codeman task add|list|status|remove|clear
|
||||
codeman ralph start|stop|status|reset codeman users add|passwd|list
|
||||
codeman status | list | attach <path> codeman doctor
|
||||
```
|
||||
|
||||
## Seam 4: Hooks
|
||||
|
||||
Claude Code hooks post to `POST /api/v1/hook-event` from inside an agent session.
|
||||
Codeman installs its own hooks automatically, but the endpoint is open to yours.
|
||||
|
||||
```json
|
||||
{ "event": "task_completed", "sessionId": "abc123", "data": { "any": "json" } }
|
||||
```
|
||||
|
||||
`event` must be one of `permission_prompt`, `elicitation_dialog`, `idle_prompt`,
|
||||
`stop`, `teammate_idle`, `task_completed`. Each becomes the matching `hook:*` SSE
|
||||
event.
|
||||
|
||||
⚠️ This endpoint skips Basic auth so hooks keep working, but when auth is active
|
||||
the loopback bypass requires the `X-Codeman-Hook-Secret` header
|
||||
(`~/.codeman/hook-secret`) unconditionally.
|
||||
|
||||
## Gotchas
|
||||
|
||||
Every one of these has cost somebody real time.
|
||||
|
||||
- **CORS is localhost-only.** `Access-Control-Allow-Origin` is echoed only for
|
||||
`localhost`, `127.0.0.1`, and `::1`. A browser app on any other origin cannot
|
||||
call the API. Integrate server-side.
|
||||
- **A missing `Origin` header is allowed**, which is why curl, CLIs, and hooks
|
||||
work. Cross-site origins are blocked by the CSRF guard.
|
||||
- **Reverse-proxy domains are rejected** by the anti-DNS-rebinding Host allowlist
|
||||
unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix`.
|
||||
- **`null` is not `undefined`.** Request schemas use Zod `.optional()`, which
|
||||
accepts `undefined` only. `JSON.stringify({ field: null })` keeps the null on
|
||||
the wire and fails with `INVALID_INPUT`. Omit the key instead. This has caused
|
||||
shipped bugs more than once.
|
||||
- **`text/plain` bodies stay raw.** Auto-parsing them as JSON enabled
|
||||
simple-request CSRF, so it is deliberate. Send `application/json`.
|
||||
- **Prompts are single-line and must end with `\r`.** The server splits your text
|
||||
and Enter into two separate tmux writes (Ink needs them apart), but it sends the
|
||||
Enter **only when the input contains a carriage return**. Without it your text
|
||||
sits on the prompt unsubmitted, which is the single most common "the wait
|
||||
endpoints don't work" report: the wait runs its full timeout on a turn that never
|
||||
started. Newlines inside the string are stripped rather than rejected, so
|
||||
`"echo A\necho B\r"` runs the single joined command `echo Aecho B`: send one line
|
||||
per call.
|
||||
- **`wait-output`'s `from=now` is not "printed after you asked".** tmux repaints
|
||||
the visible screen on attach, on resize, and on any TUI redraw, and a repaint
|
||||
arrives as ordinary output, so text already on screen can satisfy a fresh wait.
|
||||
Observed live: a marker echoed a minute earlier matched instantly. Use a marker
|
||||
unique to each call, and build it so the typed line never contains it (your own
|
||||
keystrokes echo into the stream). Matching is a literal substring, so `regex=` is
|
||||
rejected with a `400` rather than ignored.
|
||||
- **`wait-output` matches the normalized PTY stream, not the screen.** ANSI escape
|
||||
sequences are stripped (the `ESC ( B` charset escape a bash prompt emits on every
|
||||
line included), a partial escape at a chunk boundary is held back until its tail
|
||||
arrives, and a match may straddle PTY chunks, so text you printed yourself
|
||||
matches reliably (`printf STRAD; sleep 1; printf DLEQQ` is matchable as
|
||||
`STRADDLEQQ`). What can still fail is TUI output: a full-screen TUI positions
|
||||
words with cursor moves, so its text can reach the matcher **without spaces** and
|
||||
a multi-word match is unreliable there. Match one short space-free token, ideally
|
||||
one you printed yourself, and keep it out of the typed line (your own keystrokes
|
||||
echo into the stream).
|
||||
- **`stop` and `blocked` never fire for `shell`, `opencode`, `codex`, `gemini` or
|
||||
`antigravity` sessions.** They come from Claude Code hooks, which no other mode
|
||||
installs, so only `idle`, `working` and `exit` exist there. Asking for them
|
||||
explicitly is a `400`; omitting `until` is safe, since the server drops them from
|
||||
the default set and echoes what it actually waited on as `wait.until`. Even in
|
||||
`claude` mode, a Docker case needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for hooks to
|
||||
reach the server at all, a remote-SSH case's hooks may never arrive, and a case
|
||||
written by Codeman < 1.13.0 against an `--https` install carries hook curls
|
||||
without `-k` that TLS-fail silently — a 1.13.0+ server rewrites them the next
|
||||
time a session starts in that case.
|
||||
- **Unwrap the envelope** before reading fields. `data` is not the response body.
|
||||
|
||||
## Publishing your integration
|
||||
|
||||
There is no registry and no review queue. Add the GitHub topic
|
||||
**`codeman-integration`** to your public repository so others can find it, and
|
||||
link back to Codeman in your README.
|
||||
|
||||
If a real ecosystem of these appears, a manifest format and an install command
|
||||
become worth building. Until then, these four seams are the contract, and they
|
||||
require nothing of you but HTTP.
|
||||
|
||||
## What Codeman deliberately does not have
|
||||
|
||||
- **No in-process plugin runtime.** See the reasoning at the top of this page.
|
||||
- **No build or startup hooks** for third-party code. Run your own process.
|
||||
- **No per-plugin config or state directories.** Manage your own files.
|
||||
- **No sandbox for integration code**, because Codeman never launches it. Your
|
||||
integration is your own process, started by you, with your permissions,
|
||||
talking HTTP.
|
||||
|
||||
That last point is about integration code specifically, not about Codeman.
|
||||
Sandboxing lives on a different axis here: the thing worth isolating is the
|
||||
**agent**, and you isolate it per case with
|
||||
[Docker cases](docker-cases.md), which run the agent in a hardened container with
|
||||
a bind-mounted workspace and seeded (not shared) credentials. An integration that
|
||||
creates or drives a Docker-backed session inherits that isolation for free, since
|
||||
it is a property of the session rather than of the caller.
|
||||
@@ -1,7 +1,7 @@
|
||||
# Remote Sessions (SSH)
|
||||
|
||||
Codeman can run a session's agent on a **remote host over SSH** instead of the
|
||||
local machine. The agent (Claude, OpenCode, Codex, Gemini, or a plain shell)
|
||||
local machine. The agent (Claude, OpenCode, Codex, Antigravity, Gemini, or a plain shell)
|
||||
runs inside a `tmux` server **on the remote host**, so it survives the SSH
|
||||
connection dropping; Codeman attaches to it the same way it attaches to a local
|
||||
managed session.
|
||||
@@ -30,7 +30,7 @@ Types live in `src/types/session.ts`; persistence in `src/remote-hosts.ts`.
|
||||
| `RemoteHost` (extends `RemoteSshOptions`) | A saved host: `id`, `label`, `host`, `username`, `port?`, `commands?` (per-mode launch command override). |
|
||||
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
|
||||
| `SessionRemote` (extends `RemoteSshOptions`) | The resolved bundle stamped onto a live session: host coordinates + `remotePath` + `commands`, plus **`owned?`** and **`remoteSessionName?`** (COD-105 — see [Ownership](#ownership-launched-vs-discovered-and-attached-cod-105)). Built by `toSessionRemote(host, case)` (sets `owned: true`) for the launch path, or `toAttachedSessionRemote(host, name, path)` (sets `owned: false`) for the attach path. Both copy the advanced SSH options through so every connection is identical. |
|
||||
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini'>` — the modes that can run remotely. |
|
||||
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity'>` — the modes that can run remotely. |
|
||||
| `RemoteSessionInfo` (COD-105) | One discovered remote tmux session: `name` (always `codeman-*`), `attached` (a client is connected), `created` (epoch s), `windows`. Returned by `listRemoteCodemanSessions()`. |
|
||||
|
||||
Persistence is two flat JSON arrays in the instance data dir:
|
||||
@@ -115,7 +115,7 @@ Key points:
|
||||
- **`exec <cli>`** replaces the pane shell with the agent, so the pane PID *is*
|
||||
the agent. The per-mode command comes from `remote.commands?.[mode]` or
|
||||
`defaultRemoteCommandForMode(mode)` (`exec claude` / `exec opencode` /
|
||||
`exec codex` / `exec gemini` / `exec bash -l`).
|
||||
`exec codex` / `exec gemini` / `exec agy` / `exec bash -l`).
|
||||
- The **whole tmux invocation is a single shell-quoted ssh argument**, and the
|
||||
pane command is independently quoted, so a `remotePath` with spaces is safe.
|
||||
- Connection options come from the **same `buildSshConnectionArgs(remote)`** as
|
||||
|
||||
@@ -0,0 +1,250 @@
|
||||
# Scrollback fix plan (issue #205)
|
||||
|
||||
Status: IMPLEMENTED on `fix/scrollback-shell-alt-screen` (2026-08-07), with one deliberate
|
||||
divergence from the recommendation below. Kept for the diagnosis record; the measured evidence
|
||||
behind it is `docs/scrollback-issues-analysis.md`, and the mechanisms as shipped are documented
|
||||
in `docs/architecture-invariants.md` (§ Full-scrollback replay, § Terminal scrollback: strip
|
||||
flavors and wheel/touch forwarding).
|
||||
|
||||
What shipped vs. what this doc proposed:
|
||||
|
||||
- **Bug A (deltaMode)**: implemented as specified (`_wheelScrollLines()` normalizes
|
||||
line/page/pixel units, Shift-axis trap kept).
|
||||
- **Bug B (shell scrollback)**: implemented via the NARROW alt-screen strip for tmux-backed
|
||||
shell/opencode/antigravity plus the scroll-to-top `full=1` re-pull, NOT the recommended
|
||||
approach (a) `tmux mouse on`. The measurements in the analysis doc showed the alt buffer
|
||||
comes from tmux's own client-side `smcup` at attach (tmux never forwards a pane program's
|
||||
alt-screen toggles), so stripping that one sequence fixes both symptoms with no selection
|
||||
tradeoff, keeps vim/less/htop untouched, and the re-pull also covers the repaint-burst
|
||||
history loss that `mouse on` would not have addressed.
|
||||
- **Invariant change**: the "viewport-at-bottom gate stays" invariant below was deliberately
|
||||
DROPPED for forwarding modes: a repaint-mode CLI keeps no real terminal scrollback, so the
|
||||
gate pinned users to a buffer of stale frames whenever the viewport parked off-bottom.
|
||||
Forwarding now snaps to bottom first; Shift+wheel and the opt-out setting keep local
|
||||
scrollback reachable. Touch forwards through the same gate (the mobile half of the fix).
|
||||
- **Finding 5 (remote probe)**: implemented (`probeRemoteCliVersion` over ssh, deferred at
|
||||
session start, same login-shell wrapper as the launch).
|
||||
|
||||
## RETEST FAILED (2026-08-07, after v1.12.0 shipped) — analysis round 2
|
||||
|
||||
mtiller retested on 1.12.0 and reports it is NOT fixed (issue #205 comment, 2026-08-07 12:12 UTC;
|
||||
issue reopened same day with clarifying questions: mouse vs trackpad, Shift+scroll behavior,
|
||||
Claude vs shell session on the phone, and an iOS full-tab-kill to rule out stale JS). Two
|
||||
failure signatures, now analyzed against the SHIPPED 1.12.0 code (not the pre-fix code):
|
||||
|
||||
1. **iPhone Safari (Claude session assumed)**: touch scrollback goes back only a limited
|
||||
amount and sometimes REPEATS blocks of text; unreliable.
|
||||
2. **Firefox on macOS (mouse)**: wheel does NOTHING at all, while Fn+Up (= PageUp) pages back
|
||||
through INTACT text.
|
||||
|
||||
### Ruled out by code reading
|
||||
|
||||
- deltaMode mishandling: `_wheelScrollLinesFloat` normalizes line/page/pixel units correctly;
|
||||
a Firefox line-mode notch yields ±3 lines. Not the bug.
|
||||
- Ephemeral transport: `_sendInputEphemeral` (app.js) has a POST fallback when WS is down.
|
||||
- Service worker: sw.js is network-first with cache fallback; it serves stale JS only when the
|
||||
fetch FAILS (flaky mobile connection can do this — relevant to "unreliable" on the phone,
|
||||
and the fixed `CACHE_NAME = 'codeman-v1'` never invalidates that offline copy).
|
||||
|
||||
### The load-bearing observation: PageUp works, the wheel does not
|
||||
|
||||
Fn+Up is a KEYBOARD event: xterm encodes PageUp and Claude pages its own transcript (intact
|
||||
text proves Claude-side history is fine and the PTY input path is fine). The wheel path is the
|
||||
capture-phase handler, and for a Claude session it has exactly two branches:
|
||||
|
||||
- **Forwarding branch** (`_shouldForwardWheelToApp` true): snap-to-bottom + SGR reports. If
|
||||
this branch ran, the user would see the same paging motion Fn+Up produces. They see nothing.
|
||||
- **Local branch** (gate false): `_smoothScrollBy` over xterm's local buffer. For a Claude
|
||||
pane in repaint mode, tmux keeps `history_size≈0`, so `?full=1` returns roughly one frame:
|
||||
the local buffer is structurally HOLLOW, the top-of-buffer re-pull recovers nothing, and the
|
||||
wheel looks completely dead. **This matches every observed detail on Firefox.**
|
||||
|
||||
So the working hypothesis is that mtiller's sessions evaluate the gate FALSE. The gate
|
||||
(`_shouldForwardWheelToApp`) has exactly four false-paths worth checking, in likelihood order:
|
||||
|
||||
1. **`terminalWheelLocalScrollback` opt-out is ON.** Plausible: a user whose scrolling was
|
||||
broken on 1.11.x may well have toggled "Wheel scrolls local history" while trying to fix
|
||||
it. On 1.12.0 that setting now routes the wheel to a hollow local buffer = dead wheel on
|
||||
desktop AND the stale-repaint-frames experience on the phone (see below). Ask, or check
|
||||
what the setting does on their export.
|
||||
2. **`cliVersion` missing — CONFIRMED BUG, independent of whether it is mtiller's**:
|
||||
`getClaudeCliVersion()` (utils/claude-cli-resolver.ts:124-148) caches its result
|
||||
process-wide including FAILURE: on any exception it sets `_claudeVersion = null`, and the
|
||||
guard is `!== undefined`, so a single failed/timed-out probe (5s `EXEC_TIMEOUT_MS`; PATH
|
||||
under systemd/launchd; transient fs hiccup) at the FIRST Claude session start disables
|
||||
wheel forwarding for every Claude session until the server restarts. Fix: cache success
|
||||
permanently, but let failure retry (retry on next call, or a short negative-cache TTL).
|
||||
Note that mtiller sees identical breakage on phone + iPad + laptop, which points at a
|
||||
SERVER-side/session-side cause exactly like this (cliVersion is shared by all devices)
|
||||
rather than anything browser-specific.
|
||||
3. **Claude Code genuinely < 2.1.187** on their machine: gate false BY DESIGN, but the
|
||||
resulting UX is a dead-end (no local history to fall back on).
|
||||
4. mouseTrackingMode non-none (a DECSET leaked past the strip, e.g. emitted before attach or
|
||||
split across chunks in a way the carry missed): would also kill the container handler via
|
||||
the early return. Least likely, checkable via `terminal.modes.mouseTrackingMode` in console.
|
||||
|
||||
### The iPhone symptoms fit the same gate-false story
|
||||
|
||||
Touch with gate false = local `scrollLines()` over whatever repaint frames accumulated:
|
||||
"repeats blocks of text" is literally what a buffer of successive overlapping repaint frames
|
||||
looks like; "limited amount" is its thinness; "unreliable" is burst-dependence (finding 2)
|
||||
PLUS the new re-pull being actively DESTRUCTIVE for repaint panes: `_maybeRefetchFullHistory`
|
||||
does `_resetTerminalForReplay()` then writes the fetched capture, and when that capture is
|
||||
one frame (Claude pane, `history_size≈0`) it REPLACES a multi-frame buffer with less than the
|
||||
user had, mid-scroll. Stale pre-1.12 JS on the phone (suspended Safari tab) remains possible
|
||||
until they confirm the tab kill.
|
||||
|
||||
### Fix directions, ranked
|
||||
|
||||
1. **Make the re-pull refuse downgrades** (`_maybeRefetchFullHistory`, app.js): if the fetched
|
||||
capture would yield FEWER buffer rows than currently present, skip the reset+rewrite and
|
||||
keep the richer buffer (optionally cache-mark the session "re-pull useless"). Small, safe,
|
||||
kills the "got worse after scrolling to top" class. Consider skipping the re-pull entirely
|
||||
for forwarding-capable modes where tmux keeps no history.
|
||||
2. **Rescue the gate-false Claude dead-end with PageUp forwarding**: when mode is `claude`,
|
||||
the gate is false, AND the local buffer has no scrollback (`baseY === 0`), translate wheel
|
||||
lines into coalesced PageUp/PageDown key sends (mtiller just proved Claude pages correctly
|
||||
on PageUp even on their version). Zero regression risk under that triple guard: sessions
|
||||
with real local history keep local scrolling; only the currently-dead path changes.
|
||||
Caveat: older Claude menus may react to PageUp; acceptable against "completely dead".
|
||||
3. **Audit `getClaudeCliVersion()` failure caching** (utils/claude-cli-resolver.ts): a cached
|
||||
empty probe must retry (with backoff), not poison the process.
|
||||
4. **Guard the opt-out setting's footgun**: if `terminalWheelLocalScrollback` is ON for a
|
||||
repaint-mode CLI session, local history is hollow; either scope the setting's effect to
|
||||
modes with real local scrollback, or pair it with fix 2's PageUp fallback so it still
|
||||
scrolls SOMETHING.
|
||||
5. **Add a one-line gate diagnostic**: log (once per session, console) WHY the wheel chose
|
||||
local vs forward: `{mode, cliVersion, optOut, trackingMode}`. The #205 thread is now two
|
||||
rounds deep on guesswork a single console line would have answered.
|
||||
|
||||
### What shipped for round 2 (branch `fix/scrollback-205-round2`)
|
||||
|
||||
All five directions above, implemented as ranked:
|
||||
|
||||
1. **Downgrade guard** — `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the rows a
|
||||
capture will occupy (ANSI stripped, `capture-pane -J` re-wrapping accounted for) and
|
||||
`_maybeRefetchFullHistory` (app.js) skips the reset+rewrite when that is more than one
|
||||
screen short of what xterm already holds. A refused session goes on
|
||||
`_fullHistoryRepullUseless`, which raises its re-pull cooldown from 4s to 60s so a hollow
|
||||
pane stops re-fetching. Measured A/B on a live Claude pane, same gesture, same buffer:
|
||||
guard off → 341 rows collapse to 42 and every seeded row is gone; guard on → 341 rows
|
||||
preserved. The tab-switch recovery it must not break still runs (shell buffer 401 → 44 on
|
||||
a tab switch → 401 again after scrolling to the top).
|
||||
2. **PageUp/PageDown fallback** — `_maybePageCliTranscript()` translates wheel/touch travel
|
||||
into coalesced `\x1b[5~` / `\x1b[6~` under the triple guard (claude mode, forwarding gate
|
||||
false, `baseY === 0`), through the same 40ms queue as the SGR reports. Half a screen of
|
||||
travel per page: the page key always jumps a whole screen, and a 1:1 mapping was
|
||||
unusably slow with a discrete wheel. Shift is excluded — it keeps meaning "local
|
||||
scrollback". Verified live: opt-out ON on a Claude session sends real PageUp/PageDown to
|
||||
the PTY where the wheel previously did nothing.
|
||||
3. **Probe caching** — `getClaudeCliVersion()` no longer caches failure. Success is kept for
|
||||
the process lifetime; a failed probe retries with a 1/2/4…15min backoff. The cache policy
|
||||
is a pure function (`resolveClaudeCliVersion`) so the retry semantics are unit-testable
|
||||
without spawning `claude`. The VITEST short-circuit now records nothing, where before it
|
||||
wrote a permanent null.
|
||||
4. **Opt-out footgun** — handled by pairing rather than by scoping: the setting keeps meaning
|
||||
exactly what it says (the wheel goes local), and fix 2 catches the case where "local" is
|
||||
empty. Scoping the setting away from repaint-mode CLIs would have silently overridden an
|
||||
explicit user choice. The App Settings tooltip now says to leave it off for Claude/Codex.
|
||||
5. **Diagnostic** — `_logScrollRouting()` prints one line per session per distinct decision:
|
||||
`[scroll] <id> → forward-sgr|page-keys|local-scrollback|repull-refused-downgrade (mode=…,
|
||||
cliVersion=…, localScrollbackOptOut=…, mouseTracking=…, localScrollbackRows=…)`. That
|
||||
single line answers every open question in the list below.
|
||||
|
||||
Still unanswered by code alone: whether mtiller's Claude Code is genuinely older than
|
||||
2.1.187 (false-path 3), and whether the iPhone was running stale JS. The diagnostic makes
|
||||
both self-reporting, so the retest ask is now "open the console and paste the `[scroll]` line".
|
||||
|
||||
### What to get from mtiller (some already asked)
|
||||
|
||||
- Shift+scroll behavior on Firefox (distinguishes hollow-local from handler-not-firing).
|
||||
- `claude --version` on the Mac (decides false-paths 2 vs 3).
|
||||
- App Settings → Input → "Wheel scrolls local history" state (false-path 1).
|
||||
- iPhone: Claude or shell session, and whether a full tab kill changes anything.
|
||||
- Browser console: `app.terminalUi?.terminal?.modes?.mouseTrackingMode` (false-path 4).
|
||||
|
||||
Original plan follows.
|
||||
|
||||
## Reports
|
||||
|
||||
- **Issue #205** (https://github.com/Ark0N/Codeman/issues/205), OPEN:
|
||||
- **jonocodes** (author, 2026-08-03): SHELL session. Host Mac M4, brew tmux. On Android, touch-scrolling the terminal does nothing. On desktop, the mouse wheel cycles shell command history (acts like Up/Down arrows) instead of scrolling the screen.
|
||||
- **mtiller** (comment, 2026-08-06): "similar issue just with scrolling backward to see agent output. This is with Firefox on MacOS." (Claude session implied.)
|
||||
- **Reddit r/selfhosted** comment `p21x6ts` by mmtiller (= mtiller on GitHub): scrolling broken enough across phone/iPad/laptop that they fall back to Claude's own remote-control feature. Churn-risk user who otherwise loves the product; fixing this has promo value beyond the bug itself.
|
||||
|
||||
## How scrolling works today (read this before touching anything)
|
||||
|
||||
Three independent paths, all in `src/web/public/terminal-ui.js` unless noted:
|
||||
|
||||
1. **Desktop wheel** (container `wheel` listener, ~line 421): ALWAYS `preventDefault()`s, then either
|
||||
- forwards synthetic SGR wheel reports to the app (`_sendSyntheticSgrWheel`, coalesced every 40ms, fire-and-forget) when `_shouldForwardWheelToApp(ev)` (~line 2823) passes: no Shift held, opt-out setting `terminalWheelLocalScrollback` off, xterm `mouseTrackingMode === 'none'`, session mode is `claude` with `cliVersion >= 2.1.187` or `codex`, and viewport is at bottom;
|
||||
- otherwise scrolls xterm's LOCAL scrollback via `terminal.scrollLines(lines)`.
|
||||
- `lines` comes from `_wheelScrollLines(ev)` (~line 2818): `delta / 25`, i.e. it assumes PIXEL deltas.
|
||||
- NOTE: xterm.js's own internal wheel handler sits on an element INSIDE the container, so it runs FIRST (bubble order) and is not suppressed by the container's `preventDefault`.
|
||||
2. **Touch** (touchstart/move/end, ~lines 441-585): converts touch deltas to `terminal.scrollLines()` with momentum. Touch is ALWAYS local-scrollback, never forwarded to the app. Tap-to-position (touchend, ~line 533) is separate and already handles both mouse-tracking-on and server-strip cases.
|
||||
3. **Server-side strip** (`_handleTerminalOutput`, `src/session.ts:1384`): for modes in `isAltScreenStripMode()` (`src/session.ts:179` = `codex | claude | gemini`), strips alt-screen switches (`?47/?1047/?1049`), scrollback erase (`3J`), and mouse-tracking DECSETs (`?1000-?1007` except `?1004` focus) so content stays in xterm's normal buffer with scrollback intact. Includes a chunk-boundary carry so split sequences can't leak. `shell` and `opencode` (and `antigravity`) are deliberately EXCLUDED: arbitrary shell programs (vim/less/htop) legitimately need the alt screen. There is a parity copy of this strip on the replay path (`src/web/routes/session-routes.ts`, ~line 1697) and a frontend parity check `_sessionUsesServerMouseStrip()` (terminal-ui.js ~line 2751). All three must stay in sync.
|
||||
4. Related: full-scrollback replay (`GET .../terminal?full=1` on first buffer load) fills xterm local scrollback; client scrollback is hardcoded 50k (`DEFAULT_SCROLLBACK`, constants.js) vs tmux 100k.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
### Bug A: Firefox wheel deltas (mtiller's desktop case)
|
||||
|
||||
`_wheelScrollLines()` divides by 25 assuming `WheelEvent.deltaY` is pixels (`deltaMode === 0`, Chrome/Safari behavior). Firefox commonly fires `deltaMode === 1` (LINE units, deltaY around 1-3 per notch), so `Math.round(3/25) = 0` and the `|| ±1` fallback yields 1 line per event. With a discrete mouse wheel that is 1 line per notch: scrolling feels dead/broken. This hits BOTH the local-scroll path and the forwarded path, since both use the same function.
|
||||
|
||||
**Fix**: normalize by `ev.deltaMode` in `_wheelScrollLines()`:
|
||||
- `deltaMode 0` (pixels): current behavior, `delta / 25`.
|
||||
- `deltaMode 1` (lines): use the delta directly (round, keep sign fallback).
|
||||
- `deltaMode 2` (pages): `delta * terminal.rows` (or a sane page size).
|
||||
Keep the existing Shift-axis trap intact: on macOS trackpads Shift+two-finger scroll arrives as a HORIZONTAL wheel (deltaX carries the magnitude, deltaY ~0); that's why the function reads deltaX when Shift is held (issue #154). Don't lose it.
|
||||
|
||||
**Verify**: don't trust this diagnosis blindly. First reproduce in real Firefox on macOS and log `deltaMode`/`deltaY` (Firefox trackpad input can arrive as pixels; external mouse as lines). Also confirm the session's `cliVersion` probe succeeded (a failed probe disables forwarding entirely, which would point elsewhere). Unit-test by dispatching synthetic `WheelEvent`s with explicit `deltaMode` values; a Playwright `firefox` project pass is the end-to-end check.
|
||||
|
||||
### Bug B: shell mode has NO working scrollback at all (jonocodes)
|
||||
|
||||
Chain: shell mode is excluded from the alt-screen strip (correctly) → tmux attaches on the alternate screen → xterm's alt buffer has zero scrollback. Consequences:
|
||||
- **Wheel**: xterm's own internal wheel handler runs first and, in the alt buffer, converts wheel ticks into Up/Down arrow keys (alternateScroll behavior). The shell receives arrows → command history cycles. That is jonocodes' exact desktop symptom. The container handler's `scrollLines()` afterwards is a no-op (no scrollback in alt buffer).
|
||||
- **Touch**: the touch handler's `scrollLines()` is equally a no-op → "scrolling does nothing" on Android. Exact symptom two.
|
||||
- The real history exists the whole time in tmux's 100k-line buffer; nothing exposes it.
|
||||
|
||||
**Fix, recommended approach (a): enable tmux `mouse on` for shell sessions.**
|
||||
- Server-side, set `mouse on` scoped to shell sessions' tmux sessions (`tmux set-option -t <session> mouse on` at create + on attach of recovered sessions). Do NOT set it globally on the socket: claude/codex/gemini sessions rely on the DECSET strip and must not change.
|
||||
- What this buys, all natively: tmux enables mouse tracking on the outer terminal → xterm `mouseTrackingMode` goes non-none → the container handler stands down (line ~2830 check) and xterm's own encoder forwards wheel as SGR reports → tmux scrolls its OWN copy-mode history on wheel-up, auto-exits at bottom. The alt-scroll arrow conversion disappears too (tracking mode takes precedence). Desktop is fully fixed with no new endpoints.
|
||||
- **Touch**: still needs one small client change: in the touchmove path, when the active session is `shell` AND `mouseTrackingMode !== 'none'`, convert accumulated lines to `_sendSyntheticSgrWheel(x, y, lines)` instead of `scrollLines()`. The 40ms coalescing already prevents the tmux process storm (each send is a tmux send-keys server-side; unbatched flicks would spawn dozens of processes: this constraint is documented at `_sendSyntheticSgrWheel`, do not bypass it).
|
||||
- **Selection tradeoff to verify**: with tracking on, xterm hands drag events to tmux instead of doing local browser selection. Shift+drag still does local selection (xterm shift-override). Verify this UX on desktop before shipping; if it's unacceptable, fall back to approach (b).
|
||||
- **Also verify**: vim/less/htop inside the shell still behave (they'll now receive real mouse events via tmux, generally an improvement); remote shell sessions run tmux on the REMOTE host (`tmux -L codeman-remote`) and need the same option set there if remote shells are in scope (fine to defer, note it in the changeset if skipped).
|
||||
|
||||
**Fallback approach (b), only if (a)'s selection tradeoff fails testing**: keep mouse off; when a shell session is in the alt buffer, have the client send scroll intents to a small server endpoint that drives `tmux copy-mode -e -t <pane>` + `send-keys -X -N <n> scroll-up/down`. Preserves selection semantics exactly, but needs a new endpoint, server-side batching, AND suppression of xterm's native alt-scroll arrow conversion (capture-phase wheel listener with `stopPropagation`, or `attachCustomWheelEventHandler` if the vendored xterm version has it). More moving parts; (a) should be tried first.
|
||||
|
||||
**Not acceptable**: adding `shell` to `isAltScreenStripMode()`. vim/less/htop need the alt screen; that exclusion is deliberate and documented.
|
||||
|
||||
### Bug C: mtiller's phone/iPad case — UNREPRODUCED, do not guess
|
||||
|
||||
Touch is always-local by design, and Claude sessions keep content in the normal buffer (strip), so touch scrollback "should" work there. Before coding anything: build a repro matrix (iPhone Safari / iPad Safari / Android Chrome × claude / shell) on the current release. Plausible candidates if it does reproduce: auto-scroll-to-bottom fighting user scrolls (`_noteTerminalUserScroll`, ~line 2004), or they were in shell sessions on mobile too (then Bug B covers it). Ask mtiller on #205 for session mode + Codeman version if the matrix comes up clean.
|
||||
|
||||
## Invariants the implementation MUST respect
|
||||
|
||||
- Shift+wheel always scrolls local scrollback; the trackpad Shift-axis handling from #154 stays.
|
||||
- The `terminalWheelLocalScrollback` opt-out setting keeps working (pins plain wheel to local).
|
||||
- The viewport-at-bottom gate stays: once the user scrolled up locally, wheel stays local until they return to bottom.
|
||||
- 40ms SGR coalescing: never send per-event writes to the server.
|
||||
- Strip parity triangle: `session.ts` live strip ↔ `session-routes.ts` replay strip ↔ `_sessionUsesServerMouseStrip()` in the frontend. If you touch mode lists, update all three.
|
||||
- Don't add `opencode`/`antigravity` to any strip/forward list; their TUI wheel behavior is unverified (documented at `_shouldForwardWheelToApp`).
|
||||
- The chunk-boundary sequence carry in `_handleTerminalOutput` must not be weakened.
|
||||
|
||||
## Testing (per repo rules)
|
||||
|
||||
- `npm test -- test/<file>.test.ts` only; never bare `npm test`. New test ports 3150+, never 3000.
|
||||
- Browser-test traps (documented in CLAUDE.md Testing): drive input/scroll through real events (`page.mouse.wheel`, real touch), not app internals; headless Chromium reports `isTouchDevice()` false even with `hasTouch: true`; assert on real state (xterm viewport position, `tmux -L codeman capture-pane`), not HTTP 200.
|
||||
- Shell-mode E2E: create a throwaway shell session, `seq 1 500`, then (1) wheel up on desktop shows earlier lines, not history cycling; (2) touch-scroll on a phone shows earlier lines; (3) `vim` + `less` still enter/leave the alt screen cleanly; (4) Shift+drag still selects text.
|
||||
- Firefox E2E: Playwright `firefox` project, wheel over a Claude session's finished output, assert viewport moved more than 1 line per notch.
|
||||
- End-to-end against the REAL environment before claiming done (standing user rule). w1/w2/w3 tmux sessions are the user's live sessions: never send input to them; create your own throwaway session and DELETE it by exact id when done.
|
||||
|
||||
## Related observation (not a reported bug, worth a look while in there)
|
||||
|
||||
The `claude --version` probe that feeds the forwarding gate runs only for local and docker sessions (`src/session.ts:1490` gates `!this._remote`; docker handled at :1507). Remote Claude sessions therefore never get `cliVersion` and silently keep local-only wheel. Harmless (local scrollback works) but inconsistent; cheap to fix by probing over ssh, or document as intended.
|
||||
|
||||
## Rollout
|
||||
|
||||
1. Bug A (deltaMode) is small and independent: can ship alone as a patch.
|
||||
2. Bug B (shell scrollback) is the headline fix for #205: patch or minor per COM flow.
|
||||
3. After deploy + verification: comment on #205 (what was fixed, what needs their retest), then reply to the Reddit comment `p21x6ts` with the release version. Both reporters gave environment details; address them specifically.
|
||||
@@ -0,0 +1,255 @@
|
||||
# Scrollback issues: analysis and test evidence
|
||||
|
||||
Covers GitHub issue **#205** ("Scrollback in terminal not working", jonocodes, shell mode,
|
||||
Android + macOS desktop) and the follow-up comment on it from **mtiller** (Firefox on macOS,
|
||||
"scrolling backward to see agent output"). Related closed issue: **#154** (fixed in 1.3.3).
|
||||
|
||||
Status: **analysis only, nothing implemented.** Measured against the live 1.11.2 instance on
|
||||
2026-08-06 with throwaway `zz-*` shell sessions (all deleted afterwards; the user's `w*`
|
||||
sessions were never touched).
|
||||
|
||||
---
|
||||
|
||||
## TL;DR
|
||||
|
||||
Five distinct problems, not one. #205 is fully explained by finding 1; findings 2 and 3 are
|
||||
independent and hit **every** mode including Claude, and are the likely substance of the
|
||||
"similar issue" follow-up.
|
||||
|
||||
| # | Problem | Modes affected | Severity | Confirmed |
|
||||
| - | ------- | -------------- | -------- | --------- |
|
||||
| 1 | xterm parked in the **alternate buffer** for the whole session, so there is no scrollback at all and the wheel is translated into Up/Down arrow keys | `shell`, `opencode`, `antigravity` | High | Reproduced end to end |
|
||||
| 2 | **Bursty output silently destroys a screenful** of the browser's scrollback and adds ~1 row | all | High | Measured |
|
||||
| 3 | **Tab switch collapses scrollback** to roughly one screen (`full=1` fires once per page load) | all | Medium | Measured |
|
||||
| 4 | `deltaMode` is never read, so Firefox scrolls ~4x slower per notch | all, Firefox | Low | Static, needs reporter data |
|
||||
| 5 | **Remote SSH Claude cases get no `claude --version` probe**, so wheel forwarding silently stays off (residual #154) | `claude` + remote | Medium | Static |
|
||||
|
||||
---
|
||||
|
||||
## Finding 1: shell / opencode / antigravity are stuck in xterm's alternate buffer
|
||||
|
||||
### Root cause
|
||||
|
||||
The local tmux **client** (the `tmux attach` that node-pty spawns) emits `smcup` as its very
|
||||
first bytes on attach. Captured from a real PTY:
|
||||
|
||||
```
|
||||
b'\x1b[?1049h\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2J\x1b[?12l\x1b[?25h\x1b[?1000l...'
|
||||
^^^^^^^^^^ enter alternate screen ^^^^^ application cursor keys ON
|
||||
```
|
||||
|
||||
`Session._handleTerminalOutput()` strips `\x1b[?1049h` from the live stream, but only when
|
||||
`isAltScreenStripMode(mode)` is true, and that is `claude | codex | gemini` only
|
||||
(`src/session.ts:179`). For `shell`, `opencode` and `antigravity` the sequence reaches the
|
||||
browser verbatim and xterm switches to the alternate buffer, where:
|
||||
|
||||
1. `buffer.active.type === 'alternate'` and `baseY` is pinned at 0, so there is **no
|
||||
scrollback to reach**. `terminal.scrollLines()` is a no-op, which is why touch scrolling
|
||||
on Android "does nothing".
|
||||
2. xterm's own wheel listener takes over. From the vendored bundle
|
||||
(`src/web/public/vendor/xterm.min.js`):
|
||||
|
||||
```js
|
||||
if (!this.buffer.hasScrollback) {
|
||||
if (ev.deltaY === 0) return false;
|
||||
if (coreMouseService.consumeWheelEvent(...) === 0) return this.cancel(ev, true);
|
||||
const seq = ESC + (decPrivateModes.applicationCursorKeys ? 'O' : '[') + (ev.deltaY < 0 ? 'A' : 'B');
|
||||
coreService.triggerDataEvent(seq, true);
|
||||
return this.cancel(ev, true);
|
||||
}
|
||||
```
|
||||
|
||||
tmux also set `\x1b[?1h`, so the emitted sequence is `\x1bOA`, i.e. **Up arrow**, straight
|
||||
into the shell's readline. That is exactly the reported "the mouse wheel scrolls back
|
||||
through previous commands, like pressing up".
|
||||
|
||||
3. `cancel(ev, true)` calls `preventDefault()` **and `stopPropagation()`**, and xterm's
|
||||
listener sits on `terminal.element` (a child of Codeman's container). So Codeman's own
|
||||
container wheel handler, `_shouldForwardWheelToApp` and `_wheelScrollLines` included, is
|
||||
**never reached** for these modes. That whole path is dead code for shell.
|
||||
|
||||
### Reproduction (live instance, real browser)
|
||||
|
||||
Create a shell session with the page already open, print 150 lines, then dispatch 8 wheel-up
|
||||
events over `.xterm-screen`:
|
||||
|
||||
```
|
||||
t+1500 after shell start {"type":"alternate","length":35,"baseY":0}
|
||||
t+3000 after shell start {"type":"alternate","length":35,"baseY":0}
|
||||
after 150 live lines {"type":"alternate","length":35,"baseY":0}
|
||||
WHEEL on live shell: {"ptyBytes":["OA","OA","OA","OA",
|
||||
"OA","OA","OA","OA"],
|
||||
"before":0,"after":0,"type":"alternate"}
|
||||
```
|
||||
|
||||
Both reported symptoms, one root cause.
|
||||
|
||||
### Why it looks intermittent
|
||||
|
||||
The alternate-screen sequence only ever reaches the browser through the **live stream at
|
||||
attach**. Neither replay path carries it:
|
||||
|
||||
- `?full=1` returns `capture-pane` output (`source: mux-full-history`), verified 0 hits for
|
||||
`\x1b[?1049h`.
|
||||
- `?tail=` returns the visible pane frame (`source: mux-visible`), also 0 hits; the shell byte
|
||||
buffer was empty in every probe.
|
||||
- `_resetTerminalForReplay()` calls `terminal.reset()`, which returns xterm to the normal
|
||||
buffer.
|
||||
|
||||
So: watching a shell from creation leaves you in the alternate buffer until you reload or
|
||||
switch tabs, at which point it silently starts working again. Then the next PTY attach (a
|
||||
restart, or the auto-reattach in `selectSession()`) puts you back.
|
||||
|
||||
### Is stripping safe for shell? Probably yes when tmux-backed, and the current code comment is wrong about why
|
||||
|
||||
`src/session.ts:1404` says *"shell must keep the alt screen for vim/less/htop"*. For a
|
||||
**tmux-backed** shell that reasoning does not hold: tmux is a full terminal emulator and never
|
||||
forwards a pane's alternate-screen toggles to its client, it repaints instead. Measured per
|
||||
phase on a real attach:
|
||||
|
||||
| phase | bytes | `?1049h` | `?1049l` | `?47/1047` |
|
||||
| ----- | ----: | -------: | -------: | ---------: |
|
||||
| attach | 772 | **1** | 0 | 0 |
|
||||
| `seq 1 60` echo | 1402 | 0 | 0 | 0 |
|
||||
| `less` open / end / quit | 284 / 230 / 321 | 0 | 0 | 0 |
|
||||
| `vim` open / quit | 2200 / 646 | 0 | 0 | 0 |
|
||||
|
||||
`vim` and `less` inside tmux emit **zero** alternate-screen sequences to the client.
|
||||
|
||||
The caveat that does matter: `startShell()` falls back to a **direct PTY with no tmux** when
|
||||
mux creation fails (`src/session.ts:1961`, `this._useMux = false`). In that path the inner
|
||||
app's own `?1049h` does reach xterm, and a blanket strip would break vim/less/htop for real.
|
||||
Any fix has to be conditional on `_useMux`, which is known server-side.
|
||||
|
||||
Second caveat: stripping alone buys less than it looks like, because of finding 2. It fixes
|
||||
the wheel (no more phantom Up arrows) and it makes the `full=1` replay reachable, but live
|
||||
output still will not accumulate.
|
||||
|
||||
---
|
||||
|
||||
## Finding 2: bursty output silently overwrites a screenful of browser scrollback
|
||||
|
||||
Independent of the alternate buffer, and it hits Claude sessions too.
|
||||
|
||||
tmux decides per flush whether to emit real linefeeds (which push rows into the outer
|
||||
terminal's scrollback) or to repaint the pane rectangle with cursor addressing (which
|
||||
overwrites the visible rows in place). When output outpaces its flush interval it coalesces
|
||||
into a repaint, and one screenful of the browser's history is **destroyed**.
|
||||
|
||||
Measured on one session, same page, `rows = 36`:
|
||||
|
||||
| step | `baseY` | rows containing SEED | BURST | SLOW |
|
||||
| ---- | ------: | -------------------: | ----: | ---: |
|
||||
| after `?full=1` replay (120 seeded lines) | 86 | 120 | 0 | 0 |
|
||||
| after 60 lines emitted as fast as possible | **87** (+1) | **86** (-34) | 35 | 0 |
|
||||
| after 60 lines at ~16/s (`sleep 0.06`) | **148** (+61) | 86 | 35 | 60 |
|
||||
|
||||
The burst added **one** row of scrollback and ate **34** rows of existing history. The slow
|
||||
run behaved correctly. So "I printed a bunch of lines and now I cannot scroll back" reproduces
|
||||
without the alternate buffer being involved at all, and it is rate dependent, which is exactly
|
||||
the kind of thing that reads as random flakiness.
|
||||
|
||||
Consequence: the browser's scrollback is effectively frozen at whatever the last `?full=1`
|
||||
replay produced, minus a screen per burst. tmux's own history is fine throughout
|
||||
(`history_size` kept growing, `history-limit` 2000), so the data is never actually lost
|
||||
server-side, it just never reaches the browser again until a reload.
|
||||
|
||||
---
|
||||
|
||||
## Finding 3: switching tabs collapses a session's scrollback
|
||||
|
||||
`_initialFullBufferLoad` is true for the **first buffer load after a page load only**
|
||||
(`app.js:4374`). Everything after that uses `?tail=`, which returns byte history plus the
|
||||
visible pane frame. Worse, the snapshot restore path deliberately throws away the restored
|
||||
xterm snapshot (which does carry scrollback) and replaces it with that frame
|
||||
(`app.js:4316-4328` plus `needsRewrite`).
|
||||
|
||||
Measured, switching away from session A and back:
|
||||
|
||||
```
|
||||
A: initial full=1 load {"len":152,"baseY":116,"AAA":150}
|
||||
A: after switch away and back {"len": 87,"baseY": 51,"AAA": 59}
|
||||
```
|
||||
|
||||
150 lines of history down to 59. Note also that the page's single `full=1` is consumed by
|
||||
whichever session auto-selects at load, so **every other tab starts life with one frame of
|
||||
history**.
|
||||
|
||||
---
|
||||
|
||||
## Finding 4: `deltaMode` is never read (Firefox)
|
||||
|
||||
`grep -rn "deltaMode" src/web/public packages` returns nothing. `_wheelScrollLines()`
|
||||
(`terminal-ui.js:2818`) treats `deltaY` as pixels unconditionally:
|
||||
|
||||
```js
|
||||
return Math.round(delta / 25) || (delta > 0 ? 1 : -1);
|
||||
```
|
||||
|
||||
Chrome/WebKit report `deltaMode: 0` with `deltaY` around 100 to 120 px per notch, so about 4
|
||||
to 5 lines. Firefox reports `deltaMode: 1` (`DOM_DELTA_LINE`) with `deltaY` around 3, so
|
||||
`Math.round(3/25) === 0` and the `|| ±1` fallback yields **1 line per notch**, roughly 4x
|
||||
slower. In Claude mode the same value caps the forwarded SGR report at 1 tick per event
|
||||
instead of 4, so the transcript crawls too.
|
||||
|
||||
This is sluggishness, not breakage, so it is a plausible but unproven contributor to the
|
||||
mtiller report. No Firefox build is installed under `~/.cache/ms-playwright` (chromium and
|
||||
webkit only), so this was not measured. Worth asking the reporter for `deltaMode` / `deltaY`
|
||||
from a live wheel event before acting on it.
|
||||
|
||||
---
|
||||
|
||||
## Finding 5: remote SSH Claude cases still have no version probe
|
||||
|
||||
`src/session.ts:1490` deliberately skips the deterministic `claude --version` probe for
|
||||
remote sessions and defers to the startup-banner scrape, which the same comment block
|
||||
describes as unreliable ("newer Claude Code builds don't print the banner and resumed sessions
|
||||
never show it"). That is precisely the condition #154 was filed for: `cliVersion` empty means
|
||||
`_shouldForwardWheelToApp()` returns false, wheel forwarding is off, and the user is left with
|
||||
local scrollback that (per finding 2) does not accumulate.
|
||||
|
||||
Local and Docker Claude sessions are fine; verified all 7 live sessions report
|
||||
`cliVersion=2.1.223`, so the 1.3.3 fix is still working there.
|
||||
|
||||
---
|
||||
|
||||
## Candidate directions (not decided)
|
||||
|
||||
Roughly in order of value per unit of risk.
|
||||
|
||||
1. **Extend the alternate-screen strip to tmux-backed `shell` / `opencode` / `antigravity`.**
|
||||
Gate on `_useMux` so the direct-PTY fallback keeps vim/less/htop working. Kills the phantom
|
||||
Up arrows and makes replayed history reachable. `isAltScreenStripMode()` currently takes
|
||||
only `mode`, so it would need the mux flag threaded in, and
|
||||
`test/claude-scrollback-strip.test.ts:16-17` plus `test/antigravity-mode.test.ts:116` pin
|
||||
the current answers and would need updating.
|
||||
|
||||
2. **Re-pull `?full=1` when the user scrolls to the top of the buffer.** Directly addresses
|
||||
findings 2 and 3 with machinery that already exists and is already proven to return
|
||||
complete history (200/200 lines in the probe). Needs a guard against refetch storms.
|
||||
|
||||
3. **Stop discarding the xterm snapshot on tab switch**, or request `full=1` on the first load
|
||||
per session rather than per page. Cheaper partial fix for finding 3 alone.
|
||||
|
||||
4. **Read `ev.deltaMode`** in `_wheelScrollLines()` and normalise line/page deltas to lines.
|
||||
Small, self-contained, worth doing regardless of whether it is mtiller's actual bug.
|
||||
|
||||
5. **Probe the CLI version over SSH for remote Claude cases**, mirroring the deferred
|
||||
in-container probe that Docker cases already use.
|
||||
|
||||
Option 1 alone does not fix #205's "print a bunch of lines then scroll" complaint; that needs
|
||||
2 as well.
|
||||
|
||||
## Reproduction assets
|
||||
|
||||
Scripts used, in the session scratchpad
|
||||
(`/tmp/claude-1000/-home-arkon-default-claudeman/597ffc9f-.../scratchpad/`):
|
||||
|
||||
- `ptycap.py` / `ptycap2.py`: PTY-level capture of the tmux client stream, per phase counts of
|
||||
alternate-screen and mouse-tracking sequences.
|
||||
- `sim.mjs`: replays a captured stream through `@xterm/headless` with and without the strip.
|
||||
- `browser-test*.mjs`: Playwright against the live instance, reports `buffer.active.type`,
|
||||
`baseY`, row content and the exact bytes xterm sends to the PTY on a wheel event.
|
||||
|
||||
`@xterm/headless` was installed with `npm i --no-save`, so `package.json` and the lockfile are
|
||||
untouched.
|
||||
@@ -489,7 +489,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
|
||||
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
|
||||
|
||||
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
|
||||
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini`, `~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
|
||||
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — and `~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
|
||||
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
|
||||
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
|
||||
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
|
||||
|
||||
@@ -43,38 +43,22 @@ const syncData = DEC_SYNC_START + data + DEC_SYNC_END;
|
||||
this.broadcast('session:terminal', { id: sessionId, data: syncData });
|
||||
```
|
||||
|
||||
## Client-Side Implementation (`app.js`)
|
||||
## Client-Side Implementation (`terminal-ui.js`)
|
||||
|
||||
### `batchTerminalWrite(data)`
|
||||
|
||||
1. Checks if flicker filter is enabled (optional, per-session)
|
||||
2. If flicker filter active: buffers screen-clear patterns (`ESC[2J`, `ESC[H ESC[J`, `ESC[nA`)
|
||||
3. Accumulates data in `pendingWrites`
|
||||
4. Schedules `requestAnimationFrame` if not already scheduled
|
||||
5. On rAF callback: checks for incomplete sync blocks (start without end)
|
||||
6. If incomplete: waits up to 50ms via `syncWaitTimeout`
|
||||
7. Calls `flushPendingWrites()` when complete
|
||||
|
||||
### `extractSyncSegments(data)`
|
||||
|
||||
- Parses DEC 2026 markers, returns array of content segments
|
||||
- Content before sync blocks returned as-is
|
||||
- Content inside sync blocks returned without markers
|
||||
- Incomplete blocks (start without end) returned with marker for next chunk
|
||||
4. Calls `_scheduleTerminalWriteFlush()` if no flush is pending
|
||||
5. The yielded callback clears its scheduled flag before calling `flushPendingWrites()`
|
||||
6. Large batches schedule their own next chunk until the queue is empty
|
||||
|
||||
### `flushPendingWrites()`
|
||||
|
||||
```javascript
|
||||
const segments = extractSyncSegments(this.pendingWrites);
|
||||
this.pendingWrites = ''; // Clear before writing
|
||||
for (const segment of segments) {
|
||||
if (segment && !segment.startsWith(DEC_SYNC_START)) {
|
||||
terminal.write(segment); // Skip incomplete blocks (start with marker)
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Note: Segments starting with `DEC_SYNC_START` are incomplete blocks awaiting more data. These are skipped (discarded if timeout forces flush).
|
||||
- Joins the queued terminal data and passes DEC 2026 markers through to xterm.js 6, which handles synchronized output natively.
|
||||
- Writes at most 32KB per yield for Codex and 64KB for other modes.
|
||||
- Requeues the remainder and immediately schedules another safe yield. A final large response therefore drains without waiting for another SSE event.
|
||||
|
||||
### `chunkedTerminalWrite(buffer, chunkSize=128KB)`
|
||||
|
||||
@@ -116,17 +100,15 @@ When detected, buffers 50ms of subsequent output before flushing atomically.
|
||||
|
||||
## Edge Cases
|
||||
|
||||
- **Incomplete sync blocks**: 50ms timeout forces flush (content discarded to prevent freeze)
|
||||
- **Incomplete sync blocks**: xterm.js retains synchronized output until its closing marker
|
||||
- **Large buffers**: Chunked writing prevents UI freeze
|
||||
- **Server shutdown**: Skips batching via `_isStopping` flag
|
||||
- **Session switch**: Clears flicker filter state, pending writes, and sync timeout (prevents cross-session data bleed)
|
||||
- **SSE reconnect**: `handleInit()` clears all pending write state
|
||||
|
||||
**Trade-off:** If a sync block is split across SSE packets and the end marker doesn't arrive within 50ms, the incomplete content is discarded. This prioritizes responsiveness over completeness. In practice this is rare since the server always sends complete `SYNC_START...SYNC_END` pairs and SSE typically delivers them atomically.
|
||||
|
||||
## DEC Mode 2026 Compatibility
|
||||
|
||||
Terminals that natively support DEC 2026 will buffer and render atomically. Terminals that don't support it ignore the escape sequences harmlessly. xterm.js doesn't support DEC 2026 natively, so the client implements its own buffering by parsing the markers.
|
||||
Terminals that natively support DEC 2026 buffer and render atomically. Codeman uses xterm.js 6, so the client passes the markers through instead of parsing or discarding partial blocks.
|
||||
|
||||
**Supporting terminals:** WezTerm, Kitty, Ghostty, iTerm2 3.5+, Windows Terminal, VSCode terminal
|
||||
|
||||
@@ -135,4 +117,4 @@ Terminals that natively support DEC 2026 will buffer and render atomically. Term
|
||||
| File | Key Functions |
|
||||
|------|---------------|
|
||||
| `src/web/server.ts` | `batchTerminalData()`, `flushTerminalBatches()`, `broadcast()` |
|
||||
| `src/web/public/app.js` | `batchTerminalWrite()`, `extractSyncSegments()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
|
||||
| `src/web/public/terminal-ui.js` | `batchTerminalWrite()`, `_scheduleTerminalWriteFlush()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
|
||||
|
||||
@@ -0,0 +1,201 @@
|
||||
<!-- Design doc drafted 2026-07-28 from WWDC26 session 224 research. STATUS: PLANNED, NOT IMPLEMENTED. Blocked on macOS 27 "Golden Gate" (beta now, GA expected fall 2026). -->
|
||||
|
||||
# VM Cases (macOS Virtualization framework), Implementation Plan
|
||||
|
||||
## Status
|
||||
|
||||
PLANNED, nothing implemented. This is the design + phased execution plan for a native-macOS VM isolation tier for cases ("the VM subsystem"), modeled on Docker cases (`docs/docker-cases-plan.md`). Testbed prerequisite: a macOS 27 host (see Section 8).
|
||||
|
||||
**⚠ DESIGN DIRECTION (owner, 2026-07-29): the subsystem is GUI-first.** Users want real macOS desktops, not headless SSH machines. Guests may be macOS (GUI-only in practice) or Linux (GUI or headless). Key decision 3 below carries the full consequences; anything in this doc that reads as "Linux-first / headless-first" predates this and has been revised.
|
||||
|
||||
**2026-07-29: Phase 0 substantially validated on the beta testbed; full Apple-stack reference now lives in [`docs/vm-subsystem-apple-stack.md`](vm-subsystem-apple-stack.md)** (API surfaces, beta bugs, our empirical results, and design implications). Plan-relevant corrections from that work: vmnet's topology/port-forwarding APIs are macOS 26 (only the loopback fix is 27); guest provisioning is macOS-guests-only (Linux stays cloud-init, proven working); DiskImageKit has NO flatten/merge, so the `export` subcommand ships the layer chain (or flattens in-guest) instead of flattening; seed ISOs are base-build-time only, never attached at case runtime; per-case EFI variable stores are mandatory; guest health checks read DHCP leases, never serial/ping.
|
||||
|
||||
## 1. Context and motivation
|
||||
|
||||
WWDC 2026 session 224 ("Expand the Capabilities of your Virtualization App", https://developer.apple.com/videos/play/wwdc2026/224/) shipped the missing pieces for programmatic, fleet-style VM management on macOS:
|
||||
|
||||
- **`VZMacGuestProvisioningOptions`**: automated first-boot setup of a macOS guest (user account, auto-login, SSH enabled) with zero interactive setup.
|
||||
- **DiskImageKit**: stacked disk images on the Apple Sparse Image Format (ASIF): a read-only base layer plus cheap per-VM cache/overlay layers. Direct analog of Docker image layers + writable container layer.
|
||||
- **vmnet framework**: custom network topologies and port forwarding from the host process.
|
||||
- **`VZCustomVirtioDevice`**: custom low-latency host<->guest channels (Linux guests).
|
||||
- **AccessoryAccess**: USB passthrough (not relevant to Codeman, out of scope).
|
||||
|
||||
Codeman's isolation story today is Docker cases. On macOS, Docker means Docker Desktop / a Linux VM anyway, with weaker fidelity and a heavyweight dependency. The Virtualization framework gives hardware-virtualized per-case sandboxes natively, with a layered-image story that mirrors what `scripts/build-agent-image.mjs` does for Docker. This is the premium native-macOS tier ON TOP of Docker cases, never a replacement (Docker remains the cross-platform story; the Linux prod box cannot use any of this).
|
||||
|
||||
## 2. Platform reality (hard constraints)
|
||||
|
||||
| Constraint | Detail |
|
||||
| ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Host OS | macOS 27 "Golden Gate" required for the new APIs (dev beta since 2026-06-08, public beta since 2026-07-13, GA expected fall 2026) |
|
||||
| Host hardware | Apple Silicon only (macOS 27 dropped Intel). Testbed: the owner's dedicated MacBook (Section 8); the M4 Mac mini (macOS 26.4, runs the second Codeman install) stays on stable + untouched |
|
||||
| Guest provisioning | `VZMacGuestProvisioningOptions` needs macOS 27 on BOTH host and guest. Linux guests provision via cloud-init instead |
|
||||
| macOS guest concurrency | **Hard kernel cap: 2 concurrent macOS VMs per host. MEASURED on 27 beta 4 (2026-07-29), not inferred**: the 3rd VM is refused instantly with `VZErrorDomain` code 6 while 39% of RAM is free, so more hardware does NOT raise it. Since macOS GUI guests are the headline use case, this is a real product capacity limit to schedule around and surface in the UI. Linux guests are uncapped (resource-bound only) |
|
||||
| Language | Virtualization framework is Swift/ObjC only; Node cannot call it. Requires a Swift helper binary (Key decision 2) |
|
||||
| Entitlement | Host process needs `com.apple.security.virtualization`. Fine for a locally built dev binary; distribution needs signing thought (Section 9) |
|
||||
| Nested virtualization | Linux-guest-only on M3+. A macOS 27 VM cannot dependably host its own guests, so the host-side APIs must be tested on bare-metal 27 (dual-boot) |
|
||||
| CI | Cannot run in CI (needs beta macOS on Apple Silicon). Same answer as tmux/docker: no-op all VM IO under `VITEST`, unit-test the pure parts |
|
||||
|
||||
## 3. Goal and user stories
|
||||
|
||||
Add "VM cases" to Codeman: a case can point at a per-case virtual machine on a macOS host, and any CLI backend runs inside it over the existing remote-SSH session machinery. A LOCATION OVERLAY on cases, exactly like remote-SSH and Docker cases, NEVER a sixth `SessionMode`.
|
||||
|
||||
- As a Mac user, I link a case to a VM so an autonomous run executes behind a hardware virtualization boundary (stronger than Docker's shared kernel) while file viewing, transcripts, and hooks keep working.
|
||||
- Per-case VMs are instant and cheap: a shared provisioned base image plus a per-case overlay, not a full image copy per case.
|
||||
- Killing a session kills only its in-guest tmux; the VM stays up while sibling sessions remain; case delete tears the VM down.
|
||||
- I export a case's VM overlay as a portable artifact (mirror of `docker-exports/`), secrets excluded.
|
||||
- On a non-mac host, or a Mac without the helper, the feature is invisible: zero UI, zero probes, zero errors.
|
||||
|
||||
Non-goals for the MVP: USB passthrough, custom Virtio channels (Phase 3 candidate), macOS-guest fleets (capped at 2 anyway), Kubernetes-style orchestration, Intel Macs.
|
||||
|
||||
## 4. Architecture
|
||||
|
||||
```
|
||||
Codeman (Node, unchanged session layer)
|
||||
| JSON over stdout (same pattern as shelling out to docker/tmux)
|
||||
v
|
||||
codeman-vm (Swift package: CLI + per-VM GUI runner app in the console session)
|
||||
| Virtualization / DiskImageKit / vmnet
|
||||
v
|
||||
per-case VM (macOS or Linux)
|
||||
|-- GUI mode: VZVirtualMachineView in a window --> guest screen sharing --> browser (noVNC)
|
||||
|-- shell: SSH on vmnet IP --> existing remote-SSH tmux machinery
|
||||
^ VirtioFS: host case dir mounted at the SAME absolute path
|
||||
```
|
||||
|
||||
Note the runner is a **GUI app in the console user's session**, not a detached daemon: a daemon-launched VM cannot render, which is fatal for macOS guests and for Linux desktop cases.
|
||||
|
||||
### Key decision 1: location overlay, not a mode
|
||||
|
||||
Identical reasoning to Docker/remote-SSH (see CLAUDE.md): the session layer, respawn, Ralph, recovery, and quick-start plumbing all stay untouched. `SessionMode` stays five-valued. State mirrors the Docker pair: `~/.codeman/vm-hosts.json` + `vm-cases.json`, new `src/vm-hosts.ts` with the storage + pure helpers split.
|
||||
|
||||
### Key decision 2: Swift helper CLI (`codeman-vm`)
|
||||
|
||||
The framework is Swift-only, so all VM work lives in a SwiftPM package (`packages/codeman-vm/`), a CLI with a stable JSON contract:
|
||||
|
||||
- `create-base --guest linux|macos`: build the shared base image. Linux: boot an arm64 cloud image with EFI + cloud-init, install Node 22 + tmux + the four CLIs (same inventory as `docker/agent.Dockerfile`), seal as base ASIF. macOS: IPSW restore + `VZMacGuestProvisioningOptions` (agent user, SSH on), then **desktop-readiness baking**, which is mandatory for GUI guests: suppress the per-user first-login assistant (`com.apple.SetupAssistant` keys + the User Template), enable auto-login (`autoLoginUser` + `/etc/kcpassword`), disable screensaver/lock/display-sleep, and set a static wallpaper (animated "aerials" wallpaper is unusable over remote display). ⚠ Use RAW (not ASIF) for macOS guest disks until the beta's macOS-guest space-reclamation bug is fixed.
|
||||
- `create <case>`: DiskImageKit stacked image: shared read-only base + fresh per-case overlay. Near-instant, space-efficient.
|
||||
- `start <case>` / `stop` / `status` / `ip`: lifecycle + vmnet NAT; `ip` reports the guest SSH endpoint.
|
||||
- `export <case>` / `import`: flatten overlay + workspace tar + manifest, credentials excluded (mirror of docker-export).
|
||||
|
||||
A VM dies with its owning process, so `start` spawns a DETACHED per-VM runner process (analog of the detached `scripts/self-update.sh` trick) rather than a monolithic daemon; `status` talks to it over a unix socket in the instance data dir (`dataPath()`, never a hardcoded `~/.codeman` path).
|
||||
|
||||
### Key decision 3: multi-guest, and GUI is a first-class mode (REVISED 2026-07-29 by the repo owner)
|
||||
|
||||
The subsystem supports both macOS and Linux guests, and a guest runs in one of two **display modes**:
|
||||
|
||||
| | macOS guest | Linux guest |
|
||||
| --- | --- | --- |
|
||||
| **GUI mode** | **the point of the feature**; a real macOS desktop. Mandatory: nothing renders without an attached `VZVirtualMachineView` in an unlocked host session | supported (EFI + virtio-gpu framebuffer) for desktop Linux cases |
|
||||
| **Headless mode** | not offered: a macOS guest with no view renders nothing, so a "headless macOS desktop" is a contradiction. SSH-only macOS is possible but is not what this feature is for | supported and cheap; the natural mode for agent/CI work, driven over SSH |
|
||||
|
||||
Consequences that flow from GUI being first-class:
|
||||
- VM processes are **GUI apps in the console user's session** (LaunchAgent / `launchctl asuser`), never daemons. A daemon-launched VM cannot render.
|
||||
- **The host is part of the product surface**: it must auto-login, never lock, never sleep, and keep a live WindowServer. Host lock == every VM's screen goes black, so the screen lock is effectively a global kill switch for every VM display on the machine. The product must own these host settings rather than treat them as user preference.
|
||||
- **FileVault conflicts with unattended GUI hosting** and the trade-off must be a deliberate choice: FileVault disables auto-login, so a full-disk-encrypted host needs a human at a keyboard (or a remote screen-sharing session) after every reboot before any VM can render. Options are (a) FileVault on, accept manual login per boot, (b) FileVault off on a dedicated VM host so it boots straight into a rendering session, or (c) FileVault on plus a remote-unlock runbook. Codeman should detect the state and tell the user which one they are in instead of silently serving black screens.
|
||||
- **Guests must be desktop-ready, not just booted**: auto-login, no screensaver/lock, and the per-user first-login assistant pre-suppressed at base-image time (`com.apple.SetupAssistant` keys, plus the User Template so later accounts inherit it). Otherwise the user connects to a login prompt or a setup wizard, which is exactly what happened during the first hands-on run.
|
||||
- **Capacity is capped for macOS**: at most 2 concurrent macOS VMs per host, confirmed by our own test on 27 beta 4 (3rd refused with `VZErrorDomain` 6 at 39% free RAM; it is a kernel quota, so bigger hardware does not help). Scheduling must queue or evict beyond 2, the UI must explain why, and the scheduler should tolerate the acknowledged slot-leak bug (a slot occupied with nothing running, host-reboot to clear). Linux guests are uncapped and bounded only by host resources, which is the lever for scaling case counts on one machine.
|
||||
- **Access is via the guest's own screen**, viewable in a browser through the noVNC chain (see `docs/vm-subsystem-apple-stack.md` §8), so no client-version or client-install requirements land on the user.
|
||||
|
||||
Provisioning per guest type: `VZMacGuestProvisioningOptions` for macOS (needs 27-on-27, first-boot-only, and does NOT skip the per-user wizard), cloud-init NoCloud seed ISO for Linux (proven working).
|
||||
|
||||
### Key decision 3b: the GUI VM host profile, and supervision that catches black screens
|
||||
|
||||
GUI hosting only works if the host is configured for it and supervised. This profile was derived the hard way on the testbed (prototyped there 2026-07-30) and should be what `codeman-vm` installs and verifies:
|
||||
|
||||
**Host profile** (the product should own these, not leave them to preference):
|
||||
1. **No login barrier.** Either FileVault off + auto-login (a dedicated VM host boots straight into a rendering session, fully unattended), or FileVault on and remote reboots done with `sudo fdesetup authrestart`, where the pre-boot unlock *is* the login so the machine returns already logged in with encryption intact. **`authrestart` is VERIFIED on the testbed (2026-07-30): the host rebooted remotely and came back with a live logged-in console session, FileVault still enabled, no password prompt** — this is the recommended pattern for an encrypted GUI VM host. Plain reboots on a FileVault host always need a human, so Codeman should detect that combination and warn instead of serving black screens.
|
||||
2. **Never lock**: lock policy off (needs the account password, so it is a setup step, not a scriptable one) plus `caffeinate -d -i -m -u` re-armed per session.
|
||||
3. **Never sleep**: `pmset -a sleep 0 displaysleep 0 disablesleep 1`; a physical display is NOT required (a lid-closed laptop renders fine, only an unlocked session matters). Note OS updates reset these.
|
||||
4. **Session-independent control plane**: run VPN/remote access as a system service, never a session app, and keep the access chain (forwards, VNC proxies, web endpoints) in LaunchDaemons so a session restart cannot sever operator access.
|
||||
|
||||
**Supervision** must be a **root LaunchDaemon**, not a user LaunchAgent. This is the load-bearing detail: a user agent cannot launch a GUI app into the Aqua session, so its restart attempts fail *silently* (the child dies instantly, leaving an empty log while the supervisor cheerfully reports success). A root daemon can, via `launchctl asuser <uid> sudo -u <user> …`, and those launches persist. Prototyped and verified on the testbed 2026-07-30; a working supervisor runs on a short interval and:
|
||||
|
||||
- Restarts the runner when the process is gone **or when its log shows `WindowServer event port death`**, which means it is permanently blind while still looking alive.
|
||||
- Defers restarts while the console is at the login window, and launches into whichever session actually exists (resolve the console user with `stat -f %Su /dev/console`, never a hardcoded one).
|
||||
- Re-points the guest port-forward whenever the guest's NAT lease changes, which happens on **every guest boot** under plain NAT. A vmnet DHCP reservation for a stable per-case IP is the better long-term answer.
|
||||
- **Re-applies host power settings**, because `pmset -a disablesleep 1` does NOT survive a reboot (caught on the supervisor's first run after a real reboot) and OS updates reset it too.
|
||||
- Re-arms the keep-awake helper, which dies with its session.
|
||||
- Ideally also samples the guest framebuffer for non-black content, since a black screen is the one symptom common to every failure mode here.
|
||||
|
||||
`pgrep` alone is worthless for health: every failure mode in this session presented as a healthy process.
|
||||
|
||||
### Key decision 4: sessions ride the existing remote-SSH machinery
|
||||
|
||||
A provisioned guest is literally an SSH host on a vmnet IP. Session launch = the remote-SSH flow with the host swapped in: durable remote `tmux -L codeman-remote`, session names failing `SAFE_MUX_NAME_PATTERN` on purpose, EVERY ssh command line through `buildSshConnectionArgs()` (command-injection invariant), run flows through `POST /api/quick-start` (never `POST /api/sessions`, which stat-validates `workingDir` locally). What is genuinely new is only lifecycle (create/start/stop/export) and the vm-hosts/vm-cases overlay state.
|
||||
|
||||
### Key decision 5: workspace via VirtioFS at the same absolute path
|
||||
|
||||
Mirror the Docker bind-mount invariant: the case workspace is a real host directory shared into the guest via VirtioFS and mounted at the SAME absolute path. That keeps file-routes/watchers on real host bytes and makes the in-guest transcript projHash match the host. Without this, transcripts/attachments/file viewer all silently degrade.
|
||||
|
||||
### Key decision 6: credentials seeded, hooks bridged
|
||||
|
||||
- Credentials are SEEDED (read-only share, copied into the guest once at create), never shared read-write, and excluded from exports: byte-for-byte the Docker cases rule and rationale.
|
||||
- Hooks: on the loopback-only prod bind a guest cannot reach `127.0.0.1:3000`. Mirror `CODEMAN_DOCKER_BRIDGE_HOOKS` with a `CODEMAN_VM_BRIDGE_HOOKS` opt-in listener on the vmnet gateway IP; otherwise idle detection falls back to output-based, same as Docker.
|
||||
|
||||
### Key decision 7: drift and teardown copy Docker semantics verbatim
|
||||
|
||||
Config hash label on the VM (guest type, cpu/mem, share list); a drifted launch is REFUSED, never silently launched stale. One VM per case shared by all sessions; session kill = in-guest tmux kill only; case delete = stop + remove overlay; instance-scoped boot reaper for orphaned runner processes.
|
||||
|
||||
## 5. Implementation phases
|
||||
|
||||
**Phase 0, testbed (no repo code):** dedicated MacBook on the macOS 27 beta, remotely accessible over the tailnet (setup protocol in Section 8), Xcode 27 beta, then a throwaway Swift script proving the loop: create base -> overlay -> boot -> ssh in. This validates 80% of the design before any Codeman code.
|
||||
|
||||
**Phase 1, `codeman-vm` helper:** SwiftPM package, the six subcommands above, JSON contract doc, detached runner + unix-socket status, Linux base image build. Deliverable is testable entirely without Codeman.
|
||||
|
||||
**Phase 2, Codeman integration:** types (`VmHost`/`VmCase`/`SessionVm`), `src/vm-hosts.ts` (+ pure helpers: config hash, arg building, endpoint parsing), Zod schemas, `case-routes` link/unlink + listing, `quick-start` vm branch reusing the remote-SSH launch path, `Session` threading + recovery round-trip, `VITEST` no-op layer, unit tests. Feature-detect: darwin + arm64 + helper binary present, else invisible.
|
||||
|
||||
**Phase 3, polish:** export/import UI, frontend Create Case "VM" tab + case-picker labels, SSE `vm:*` events, macOS-guest opt-in with cap surfaced, custom-Virtio input channel exploration, CLAUDE.md Key Pattern + `docs/vm-cases.md` + COM.
|
||||
|
||||
## 6. Testing
|
||||
|
||||
- Pure helpers unit-tested (ports pattern from `docker-hosts.ts`: 26 tests there, aim similar).
|
||||
- All helper-invoking IO no-ops under `VITEST` (the `IS_TEST_MODE` pattern in `tmux-manager.ts`).
|
||||
- End-to-end verification happens ON the beta MacBook, per the always-end-to-end rule: real base build, real per-case overlay boot, real quick-start into the guest, workspace round-trip through VirtioFS, session-delete keeps VM up, case-delete removes it.
|
||||
- CI never runs the real path; the static guards are type-level + unit-level only.
|
||||
|
||||
## 7. Risks
|
||||
|
||||
1. **Beta API churn**: everything here targets beta SDKs; symbol/behavior changes are likely before fall GA. Mitigation: Phase 0/1 are throwaway-tolerant; no Codeman-side commitment until the helper contract survives a beta cycle.
|
||||
2. **New artifact class**: Codeman ships pure TypeScript today; a Swift binary changes build/distribution (build-on-install via `xcrun swift build` on macs with Xcode CLT? prebuilt signed binary per release?). Needs an owner decision; local dev build is fine for the whole beta period.
|
||||
3. **Entitlement/signing**: `com.apple.security.virtualization` is trivial for local dev, real for distribution.
|
||||
4. **Adoption gating**: users need macOS 27 + Apple Silicon for months after GA. Docker cases remain the default recommendation; VM cases ship dark (feature-detected) with zero cost to everyone else.
|
||||
|
||||
## 8. Beta testbed plan: dedicated MacBook (actionable now)
|
||||
|
||||
Testbed is a dedicated MacBook the owner sacrifices to the beta (after a full backup). This supersedes the earlier dual-boot-the-Mini idea (git history has it): a dedicated machine means no OS-switching, no downtime for the Mini's live Codeman, and no FileVault pre-boot headaches.
|
||||
|
||||
**Sequencing rule that makes it headless: configure ALL remote access on the CURRENT macOS first, THEN upgrade in place.** An in-place beta upgrade preserves Remote Login, Tailscale, user accounts, and auto-login, so there is no Setup Assistant and no post-install physical step. (A fresh install would boot into GUI-only Setup Assistant with no SSH, which on a headless box is a dead end.)
|
||||
|
||||
Confirmed hardware (2026-07-28): MacBook, M3, 16 GB RAM, 256 GB disk with ~100 GB free. Verdict: green. M3 = eligible + nested-virt capable; 16 GB = host + 2-3 concurrent Linux guests (macOS guest = one at a time); 100 GB = fits with discipline: install Xcode 27 beta with the macOS platform only (skipping iOS/watchOS/tvOS simulators saves 15-20 GB), and defer any macOS guest base (~30 GB) to an external SSD or until actually needed. Linux guests + sparse ASIF overlays are the comfortable path.
|
||||
|
||||
### Pre-upgrade checklist (owner, physical, once)
|
||||
|
||||
1. Full backup (Time Machine or clone); the machine should be considered beta-only afterwards.
|
||||
2. Tailscale: install, sign into the tailnet, confirm it appears in `tailscale status` from another node.
|
||||
3. System Settings -> General -> Sharing: **Remote Login ON** (SSH) and **Screen Sharing ON** (for the rare GUI-only moments: Xcode license, Apple Account dialogs).
|
||||
4. **FileVault stays ON** (owner decision 2026-07-28, security over convenience). Consequences: auto-login is unavailable, but FileVault's pre-boot unlock doubles as login, so an unlocked boot still lands in a live GUI session; planned remote reboots go through `sudo fdesetup authrestart` (unlocks for exactly one restart); an UNPLANNED reboot (beta kernel panic, battery drain) parks the machine at the pre-boot screen, no SSH/Tailscale, until the password is typed physically. If the testbed goes silent, suspect this first. Keep it on AC so the battery absorbs power blips.
|
||||
5. Beta enrollment (manual): sign into the Apple Account in System Settings; System Settings -> General -> Software Update -> **Beta Updates** -> select the **macOS 27 Developer Beta** (preferred: framework fixes land weeks earlier than public beta; free since 2023 after accepting the agreement once at developer.apple.com; public-beta alternative: enroll at beta.apple.com). Then run the offered upgrade: plugged in, lid open, trusted network.
|
||||
6. Send over: tailnet name/IP, username, and a first-login password (key install + lockdown happens remotely right after).
|
||||
|
||||
### Post-upgrade setup (remote, over the tailnet)
|
||||
|
||||
1. Verify: `sw_vers` reports 27.x, SSH reachable.
|
||||
2. Server-ize the laptop: `sudo pmset -a sleep 0 disksleep 0 disablesleep 1` (lid-closed operation without an external display), `womp 1` (wake on network), `sudo systemsetup -setrestartpowerfailure on`. Keep on AC power.
|
||||
3. Install the controlling host's SSH key, then disable password auth.
|
||||
4. Xcode 27 beta install (the one step needing the owner's Apple Account sign-in once, doable via Screen Sharing from anywhere); `xcode-select`, license accept, verify `swift --version` + the 27 SDK (`xcrun --show-sdk-version`).
|
||||
5. Phase 0 prototype loop, all remote from here: Linux guest base image (no 27-on-27 provisioning dependency), DiskImageKit overlay, boot, vmnet NAT, ssh into the guest, run `claude --version` inside.
|
||||
6. Only after that loop works: start Phase 1 in `packages/codeman-vm/`.
|
||||
|
||||
## 9. Open decisions (owner)
|
||||
|
||||
1. Linux base distro/image for the default guest (proposal: Ubuntu 24.04 arm64 cloud image, matching the docker agent image's userland).
|
||||
2. Helper distribution for GA: build-on-install vs prebuilt signed binary vs "bring your own Xcode".
|
||||
3. Ship dark behind `CODEMAN_VM_CASES=1` for the first release, or feature-detect only?
|
||||
4. Export format parity with docker-exports (one manifest schema for both?).
|
||||
|
||||
## References
|
||||
|
||||
- Session 224: https://developer.apple.com/videos/play/wwdc2026/224/
|
||||
- Fleet-angle writeup: https://bitrise.io/blog/post/wwdc26-the-virtualization-framework-updates-that-matter-for-large-mac-fleets
|
||||
- Beta timeline: https://www.macworld.com/article/3189014/apple-july-2026-ios-ipados-macos-27-public-betas-tv-arcade-releases.html
|
||||
- Internal analogs: `docs/docker-cases-plan.md` (architecture template), `docs/remote-sessions.md` (session transport), `docs/architecture-invariants.md#docker-cases`
|
||||
@@ -0,0 +1,281 @@
|
||||
<!-- Reference doc for the VM subsystem (Codeman VM cases). Compiled 2026-07-29 from: Apple DocC JSON backend, macOS 27 beta 4 SDK on the testbed, a multi-source web research sweep, and hands-on prototyping on a MacBook Air M3 running macOS 27.0 beta (26A5388g). Companion to vm-cases-plan.md (the Codeman integration plan). -->
|
||||
|
||||
# The VM Subsystem: Apple Virtualization Stack Reference (macOS 27 "Golden Gate")
|
||||
|
||||
"VM subsystem" is the working name for Codeman's native-macOS VM isolation tier and everything under it. This document is the single place for what the Apple stack actually provides, what we have verified ourselves on the beta, and what is known-broken. The Codeman-side design lives in `docs/vm-cases-plan.md`.
|
||||
|
||||
**Research method note:** Apple's HTML doc pages are JS-rendered and come back empty to fetchers. The working route is the DocC JSON backend: `https://developer.apple.com/tutorials/data/documentation/<path>.json` (page content) and `https://developer.apple.com/tutorials/data/index/<framework>` (full symbol tree with per-symbol `beta` flags). Everything below marked "Apple docs" was parsed from that backend directly.
|
||||
|
||||
## 1. Component map and minimum OS versions
|
||||
|
||||
| Component | What it is | Min host OS | Notes |
|
||||
| --- | --- | --- | --- |
|
||||
| Virtualization.framework core | VMs, EFI/Linux boot, virtio devices, VirtioFS | macOS 11-13 era | Unchanged basics; our prototype uses nothing newer than macOS 13 APIs except the DiskImageKit bridge |
|
||||
| **DiskImageKit** | ASIF + raw disk images, layered stacks | **macOS 27** | Swift-only, no ObjC headers. Section 2 |
|
||||
| **Guest provisioning** | First-boot account/SSH setup for macOS guests | **macOS 27 host AND guest** | Mac guests only as of beta 4. Section 3 |
|
||||
| vmnet topology/port-forward/DHCP APIs | Custom networks, port forwarding | **macOS 26** (NOT 27) | 27 adds exactly one fix: loopback port forwarding. Section 4 |
|
||||
| `VZVmnetNetworkDeviceAttachment` | In-process vmnet attach | macOS 26 | |
|
||||
| **`VZCustomVirtioDevice`** family | Custom paravirt devices | **macOS 27** | Linux guests only, custom guest driver required. Section 5 |
|
||||
| AccessoryAccess (USB passthrough) | USB claim + attach to VMs | macOS 27 | Requires paid-team provisioning profile, Dock app. Out of scope for Codeman. Section 6 |
|
||||
|
||||
Corrections to the WWDC-session framing we started with: vmnet's topology family is a macOS 26 story (129 symbols, zero beta-flagged in 27); provisioning does NOT currently extend beyond macOS guests despite the generic-looking `VZGuestProvisioningOptions` base class; DiskImageKit has no attach/mount API at all (it is a file-format library that hands `DiskImage` objects to Virtualization, no `/dev/diskN`, no root needed, no entitlement documented).
|
||||
|
||||
## 2. DiskImageKit (macOS 27, Swift-only)
|
||||
|
||||
Public framework, `/System/Library/Frameworks/DiskImageKit.framework`. No ObjC headers; the API surface lives in the `.swiftinterface`. Verified present in the CLT 27 beta 4 SDK, and our prototype compiled against it with plain `swiftc` on the first attempt.
|
||||
|
||||
### API surface (complete as of beta 4)
|
||||
|
||||
```swift
|
||||
class DiskImage {
|
||||
convenience init(creating: some DiskImage.CreationConfiguration) throws
|
||||
convenience init(opening: some OpenConfigurationProtocol) throws
|
||||
func appending(any DiskImage.CreationConfiguration & DiskImage.StackableLayer) throws -> any StackedImage
|
||||
func appending(consuming DiskImage) throws -> any StackedImage // reattach an existing layer; validates parentUUID
|
||||
func truncate(blockCount: Int) throws // stacked: affects top layer; does NOT resize guest fs
|
||||
var blockCount, blockSize, format, layerType, layerUUID, parentUUID, openMode, size, url
|
||||
}
|
||||
protocol StackedImage: DiskImage { var layers: [DiskImage] }
|
||||
struct OpenConfiguration { init(url:mode:); Mode = automatic | readOnly | readWrite }
|
||||
// CreationConfiguration statics: .asif(url:blockCount:blockSize:), .asifLayer(url:type:), .raw(url:blockCount:)
|
||||
// DiskImage.LayerType: .cache | .overlay | .overlay(blockCount:)
|
||||
// DiskImage.BlockSize: .bytes512 | .bytes4096
|
||||
// Errors: CorruptedImageError, IncompatibleStackingError(reason), InvalidBlockCountError, UnsupportedFormatError
|
||||
```
|
||||
|
||||
Bridge into Virtualization is a new beta convenience init on the existing attachment class. Note there is no `readOnly:` parameter; read-only-ness comes from each layer's own `openMode`:
|
||||
|
||||
```swift
|
||||
VZDiskImageStorageDeviceAttachment(diskImage: stack, cachingMode: .automatic, synchronizationMode: .full)
|
||||
```
|
||||
|
||||
### Stacking rules (Apple docs, verbatim where quoted)
|
||||
|
||||
- ASIF works standalone or stacked. "You can only use RAW images as standalone images or as **base** images in stacked configurations." Upper layers are always ASIF.
|
||||
- **One cache layer per stack**, any number of overlays conceptually, "shallow stacks perform better" (WWDC 224). No published max-depth guidance.
|
||||
- "Layers are processed from bottom (base) to top. The **topmost layer determines the stack's size and receives all writes**." `.overlay(blockCount:)` therefore also grows the virtual disk.
|
||||
- UUID chaining: appending sets the child's `parentUUID` to the parent's `layerUUID`. Raw bases have no UUID. "The layer UUID **changes if the layer is written to**", and reattaching a mismatched layer throws `IncompatibleStackingError`. This is the mechanism that makes a shared read-only base safe.
|
||||
- Base sharing across multiple VMs is the stated design intent ("can be shared across multiple VMs"), with the WWDC caveat that per-VM auxiliary files (EFI variable store, macOS auxiliary storage) must be duplicated per VM, never shared.
|
||||
- **There is no flatten/merge.** An overlay cannot be merged back into its base (confirmed by Howard Oakley's coverage plus an independent hands-on report). Export/move flows must ship the layer chain, or flatten inside a guest (dd to a fresh attached image).
|
||||
|
||||
### Known issues and adoption
|
||||
|
||||
- **ASIF space reclamation is broken for macOS guests on the beta** (deleted files never return space, survives reboots). Linux guests reclaim correctly on both raw and ASIF via `fstrim -av`. Single detailed field report, unrefuted. Since the VM subsystem targets macOS guests, the practical rule until this is fixed is: back macOS guest disks with RAW, and revisit ASIF stacking for macOS guests each beta (stacking still works, the disks just never shrink).
|
||||
- **Zero shipping adopters anywhere.** tart has a design issue with no activity; nobody has published working DiskImageKit code. Everything must be treated as field-untested (and our own testing bears that out, Section 8).
|
||||
- Framework binary grew every beta (588 → 598 across betas 1-4); expect churn until GA.
|
||||
- Release notes list no DiskImageKit known issues in any beta, which given the above says more about the notes than the framework.
|
||||
|
||||
## 3. Guest provisioning (macOS guests only)
|
||||
|
||||
```swift
|
||||
class VZGuestProvisioningOptions: NSObject { func validate() throws } // "use one of its subclasses"
|
||||
class VZMacGuestProvisioningOptions: VZGuestProvisioningOptions {
|
||||
var fullName, username, password: String
|
||||
var logsInAutomatically: Bool
|
||||
var enablesRemoteLogin: Bool // SSH
|
||||
}
|
||||
// Wiring: VZMacOSVirtualMachineStartOptions.guestProvisioningOptions (Mac-typed)
|
||||
// .setGuestProvisioning(_:) throws (validating setter)
|
||||
```
|
||||
|
||||
- **Requires macOS 27 on host AND guest.** Older guests **silently ignore** the options (no error).
|
||||
- **First boot after restore only.** Cannot reconfigure an already-provisioned VM; property changes after start are no-ops.
|
||||
- The base class is forward-looking scaffolding; its only subclass is Mac. A Linux/cloud-init analogue may come later; do not assume it lands in 27.0. For Linux guests, cloud-init NoCloud seed ISOs remain the provisioning path (proven working, Section 8).
|
||||
- Field-verified behavior (third-party hands-on, beta 3): provisioned account gets full admin + sudo; Setup Assistant fully skipped; SSH reachable ~48 s after first boot. **Race**: the account is created late in first boot (~T+54 s), after LaunchDaemons start (~T+33 s), so anything at daemon-level must wait for the account to exist.
|
||||
- Open Apple-acknowledged bug: provisioned users are invisible to `CSIdentityQueryExecute()` (FB23716201).
|
||||
- IPSW acquisition gotcha for automation: `VZMacOSRestoreImage.latestSupported` tracks the latest *release* (returned 26.5.2), not the installed beta; beta IPSWs must be fetched from the seed CDN explicitly.
|
||||
|
||||
## 4. vmnet: a macOS 26 feature set, one macOS 27 fix
|
||||
|
||||
Everything interesting shipped in macOS 26: `vmnet_network_create`, `vmnet_network_configuration_create`, `..._add_port_forwarding_rule`, `..._add_dhcp_reservation`, subnet/prefix/MTU/external-interface setters, NAT44/NAT66/DHCP/DNS-proxy/RA disables, plus serialization (`vmnet_network_copy_serialization` / `_create_with_serialization`) for handing networks across processes. `VZVmnetNetworkDeviceAttachment` is macOS 26.
|
||||
|
||||
macOS 27's only change (beta 4 release notes, verbatim): "The vmnet port forwarding APIs now support port forwarding when communicating over loopback." That closes the old gap where the host could not reach its own forwarded ports via 127.0.0.1 (confirmed working by the original bug reporter). Directly relevant to Codeman's loopback-bound production server talking to per-case guests.
|
||||
|
||||
Gotchas:
|
||||
- vmnet networks are **not persisted**; they die with the owning process. Persist settings yourself and recreate (or serialize across processes).
|
||||
- The `com.apple.vm.networking` entitlement is still restricted ("contact your Apple representative", though DTS says most requests are approved). The plain `VZNATNetworkDeviceAttachment` needs no special entitlement and is what our prototype uses.
|
||||
- Ecosystem signal: tart's maintainer is not adopting in-process vmnet (prefers their separate-process softnet), so field testing of these APIs is thin.
|
||||
|
||||
## 5. VZCustomVirtioDevice (macOS 27, Linux guests only)
|
||||
|
||||
14 new types (`VZCustomVirtioDevice(+Configuration/Delegate/Provider)`, `VZVirtioQueue(+Element)`, `VZVirtioFeatureSet`, shared-memory-region types, `VZGuestMemoryMapping`), wired via `VZVirtualMachineConfiguration.customVirtioDevices`. Mandatory for guest discovery: `deviceID`, `pciClassID`, `pciSubclassID`, `virtioQueueCount`. You must write the Linux guest driver (Virtio spec 1.3/1.4). Threading contract: the framework calls the device/delegate on a serial queue (`deviceQueue`, defaulting to the VM's queue). Zero public adopters. For the VM subsystem this is a Phase 3+ option for a low-latency host-guest channel; SSH over NAT is proven and sufficient for now.
|
||||
|
||||
## 6. Signing and entitlements
|
||||
|
||||
- **Core loop (VZ + DiskImageKit + provisioning): ad-hoc signing with only `com.apple.security.virtualization` suffices.** Verified by us on beta 4 (plain `codesign --entitlements ... -s -` on a `swiftc` binary) and independently by third parties on beta 3. DiskImageKit documents no entitlement at all.
|
||||
- **Over-entitling is the actual trap.** Adding `com.apple.application-identifier`/team-identifier keys without an embedded provisioning profile hangs the process before `main` (watchdog kill); shipping `com.apple.vm.networking` unauthorized gets AMFI SIGKILL at exec (exit 137, no crash report, even for `--version`). Keep the entitlements plist to exactly the one key.
|
||||
- **USB passthrough breaks the ad-hoc story**: `com.apple.developer.accessory-access.usb` is profile-restricted (any paid team, no ad-hoc), additionally requires `com.apple.security.device.usb`, and `AAUSBAccessoryManager` presents UI, so it wants a Dock app, not a headless CLI. Out of scope for Codeman.
|
||||
- No Xcode required for any of the above: the CLT beta (~500 MB via `softwareupdate`) carries the full macOS 27 SDK including DiskImageKit and compiles/signs everything.
|
||||
|
||||
## 7. Ecosystem state (July 2026)
|
||||
|
||||
- **tart is now `openai/tart`** (moved from cirruslabs, mid-2026) and **relicensed to FSL-1.1-ALv2** (no longer permissive). Provisioning support shipped in 2.33.0. Old cirruslabs URLs and license assumptions are stale.
|
||||
- VirtualBuddy shipped provisioning ("Skip Setup Assistant") in 2.2 betas; had to add account-detail validation and a workaround installer for the cross-version bug below.
|
||||
- lima is deliberately waiting for GA before touching macOS 27 APIs.
|
||||
- **Code-Hex/vz (Go bindings) is dormant** (no commits since Feb 2026, no macOS 27 APIs), so the entire Go ecosystem (podman-machine, colima) currently has no path to these APIs. Swift is the only realistic binding today, which validates the VM subsystem's Swift-helper design.
|
||||
- Useful pattern if ever supporting older SDKs: resolve new classes via `NSClassFromString` at runtime (no link-time dependency), fail gracefully when absent.
|
||||
- **Cross-version restore bug**: installing a macOS 27 guest from IPSW on a macOS 26 host fails at 77-78% (`VZErrorDomain 10007`); fixed in 26.6b3 + Xcode 27b4 era, with a nasty MobileDevice.pkg trap (installing it from Xcode 27 beta on a 26 host requires a full macOS reinstall to undo). Not relevant to our 27-host testbed, very relevant to anyone on a 26 host.
|
||||
|
||||
## 8. Our empirical results (beta 4, 26A5388g, MacBook Air M3, 2026-07-29)
|
||||
|
||||
Prototype tooling, all in `~/vm-lab/` on the testbed, compiled with CLT-only `swiftc` and ad-hoc signed with the single virtualization entitlement:
|
||||
|
||||
| Tool | Purpose |
|
||||
| --- | --- |
|
||||
| `vzboot.swift` | Linux guest: EFI boot + virtio disk/net/entropy + NAT + optional cloud-init seed ISO + serial on stdio |
|
||||
| `vzstack.swift` | Same, but boots a DiskImageKit stack (read-only raw base + ASIF overlay) |
|
||||
| `vzmac.swift` | macOS guest: `install` (IPSW restore into a bundle) and `run` (boot, `--provision` for first-boot account/SSH) |
|
||||
| `vzmacgui.swift` | macOS guest in a real window via `VZVirtualMachineView` (required for the guest to render at all) |
|
||||
| `setup-seed.sh` | Builds a cloud-init NoCloud seed ISO with `hdiutil makehybrid` (volume label `cidata`) |
|
||||
| `vncproxy.py` | RFB proxy that advertises only security type 2, so version-skewed/browser clients can authenticate |
|
||||
| noVNC + `websockify` | Browser access; `websockify --web noVNC-<ver> 0.0.0.0:<port> 127.0.0.1:<proxy>` |
|
||||
| `vmwatchdog.sh` + `vmaccess.sh` | Supervision: root LaunchDaemon that restarts a blind/dead runner, re-points the forward, re-applies `pmset`, re-arms keep-awake; plus a keeper for the proxy/web endpoints |
|
||||
|
||||
Host-side diagnostics written during this work (in the session scratchpad, not on the testbed): `vnclogin.py` (Apple DH auth + session open, distinguishes "credentials rejected" from "authorized but session refused"), `vncshot.py` (decodes the raw framebuffer to PNG and reports non-black pixel counts, plus optional synthetic wake input), `relay.py` (plain TCP relay used to bridge a tailnet peer to a LAN-only host), `sshpw.py` (pty-driven password SSH for the one-time key bootstrap into a freshly provisioned guest).
|
||||
|
||||
### Proven working
|
||||
|
||||
1. **Boot**: Debian 12 arm64 cloud images (nocloud and genericcloud variants) boot under `VZEFIBootLoader` + `VZGenericPlatformConfiguration`.
|
||||
2. **Networking**: `VZNATNetworkDeviceAttachment` gives the guest a `192.168.64.x` DHCP lease from the host's bootpd (leases visible in `/var/db/dhcpd_leases`, bridge is `bridge100`).
|
||||
3. **cloud-init provisioning**: NoCloud seed ISO (built with `hdiutil makehybrid -iso -joliet -default-volume-name cidata`) created a `codeman` user with SSH key + passwordless sudo on first boot; `ssh codeman@<lease-ip>` from the host works with key auth.
|
||||
4. **DiskImageKit stack mechanics**: opening a raw base `.readOnly`, appending an ASIF overlay (`ASIFCreationConfiguration.layer(url:type:.overlay)`), attaching via `init(diskImage:)`, and booting it. The overlay received ~44 MB of boot-time writes while the **base file's SHA-256 stayed bit-identical**, which is the write-isolation property the whole per-case design rests on.
|
||||
5. **Reattach**: reopening an existing overlay and `appending(consuming:)` onto the same base passes UUID validation.
|
||||
6. **macOS guest install (added later the same day)**: `VZMacOSInstaller` restore of the 27.0 IPSW (26A5388g, fetched from the seed CDN via appledb; same build as host) into a sparse 64 GiB raw disk + auxiliary storage: INSTALL-OK on the first attempt, ~25 minutes.
|
||||
7. **Headless guest provisioning WORKS**: `VZMacGuestProvisioningOptions` via `setGuestProvisioning` (username, password, `enablesRemoteLogin`, `logsInAutomatically=false`) produced, with zero GUI interaction: an account with full admin (groups include `80(admin)`, `com.apple.access_ssh`), Remote Login on from first boot, port 22 reachable ~140 s after first-boot start, hostname auto-derived from the account ("Codemans-Virtual-Machine"). SSH password auth is on by default, so the bootstrap path is: pty-driven password login once to install `authorized_keys`, key auth thereafter. Note the provisioned account's sudo is NOT passwordless (`echo <pass> | sudo -S ...`), and provisioning is first-boot-only (later boots take no options and just boot).
|
||||
8. **Slot-leak bug NOT reproduced on 26A5388g**: a guest-initiated `shutdown -h now` fired `guestDidStop` cleanly and an immediate relaunch started fine (SSH-ready again in ~75 s), so FB22967193 (VM slot leaked on guest-initiated shutdown, host reboot to recover) did not manifest after one cycle. Either fixed in beta 4 or needs more cycles to trigger.
|
||||
|
||||
### Unstable / under investigation (beta-quality territory)
|
||||
|
||||
Boot reliability degraded over a ~15-VM session on one host boot, ending with reproducible silent hangs (VM process alive, 0% CPU, no DHCP, no ARP, nothing on serial):
|
||||
|
||||
- A genericcloud base that had been booted read-write once (cloud-init first boot) subsequently hung on every boot **with the seed ISO still attached**, while booting **without** the seed succeeded, then later runs failed in both configurations. The seed correlation is strong but was observed while host state was already suspect, so it needs a retest from a clean baseline.
|
||||
- The first stack-boot "success" that later wedged turned out (via DHCP lease timestamp arithmetic) never to have reached the network at all; its overlay growth was pre-network boot writes.
|
||||
- Working hypothesis, matching a class of acknowledged beta bugs (e.g. the VM-slot counter that leaks on guest-initiated shutdown, FB22967193, where only a host reboot recovers): accumulated hypervisor/vmnet state on the host degrades boots. Requires a host reboot + a disciplined retest matrix to confirm.
|
||||
|
||||
### Display rendering: the single most important operational finding
|
||||
|
||||
**A VZ macOS guest renders nothing unless a `VZVirtualMachineView` is attached AND the host session is actually drawing.** Verified byte-for-byte: the guest's own screen sharing serves an all-zero framebuffer (0 non-black bytes across 400 KB samples, with a sane pixel format: `rmax/gmax/bmax = 255`, shifts 16/8/0), in-guest `screencapture` fails with "could not create image from display", and no `IODisplayWrangler` shows up in the guest's `ioreg`. Three distinct states all produce black:
|
||||
|
||||
1. **Headless** (VM run with no view attached).
|
||||
2. **View attached, host session locked.** The lock screen suspends drawing and the guest's virtual GPU produces no frames.
|
||||
3. **View attached, but the app lost its WindowServer connection** (see the incident below): black permanently until the app is restarted.
|
||||
|
||||
**Consequence for the VM subsystem: rendering is a first-class requirement, not an optional extra (owner decision 2026-07-29).** The product serves GUI desktops: mandatory for macOS guests, optional-but-supported for Linux guests (which can also run headless over SSH). Any VM in GUI mode must be launched by an app that attaches a `VZVirtualMachineView`, from inside a host GUI session that is logged in and unlocked. That makes the following non-negotiable parts of the design, not workarounds:
|
||||
|
||||
- VMs run as **GUI apps in the console user's session** (launched via a LaunchAgent or `launchctl asuser`), never as daemons.
|
||||
- The **host must auto-login and never lock or sleep**; a locked host is equivalent to a powered-off display for every VM on it.
|
||||
- The **guest must auto-login, never lock, and have its first-login assistant pre-suppressed**, or the "desktop" a user connects to is a password prompt or a setup wizard.
|
||||
- A VM app that loses its WindowServer connection is **permanently blind** and must be restarted; supervision has to detect that, not just check that the process is alive.
|
||||
- The **2-concurrent-macOS-VM cap** becomes a real capacity limit for the product, so it must be surfaced in the UI and tested (still untested worldwide as of this writing).
|
||||
|
||||
### Incident 2026-07-29: `killall -HUP loginwindow` (never do this on a remote Mac)
|
||||
|
||||
Applying a wallpaper change on the testbed with `killall -HUP loginwindow` restarted the host's login session. Three consequences:
|
||||
|
||||
1. **The Mac dropped off the tailnet entirely.** Tailscale's App Store build is a GUI app living in the user session, so killing the session killed the VPN; remote access was gone until someone logged in. Recovery came from a second machine on the same LAN: it could still SSH in, and then relay ports back over the tailnet (a plain TCP relay on a tailnet-connected LAN peer is a good out-of-band path worth keeping ready).
|
||||
2. **The VM app lost its WindowServer connection** (`HIToolbox: received notification of WindowServer event port death`) while surviving as a process. Every later black screen traced to this, and nothing guest-side could fix it; only restarting the app restored rendering.
|
||||
3. The session's `caffeinate` died, so the host resumed auto-locking.
|
||||
|
||||
Rule: on a remote Mac, never run session-level commands (`killall -HUP loginwindow`, `pkill -u <user>`, logout, fast user switching). `killall WallpaperAgent` alone is session-safe. Before any such command, enumerate what depends on that session: VPN, VM processes, port forwards, keep-awake helpers.
|
||||
|
||||
### Keeping host and guest usable unattended
|
||||
|
||||
- **Host**: `caffeinate -d -i -m -u` prevents display sleep but does NOT override the lock policy. "Require password after screen saver begins or display is turned off → Never" must be set in System Settings; it needs the account password, so a passwordless-sudo shell cannot script it, and turning it off does NOT dismiss a lock that is already engaged (one more unlock is always needed). `pmset -a disablesleep 1` keeps a lid-closed laptop awake but **does not survive a reboot**, and OS updates reset it too, so a supervisor should re-apply it rather than assume it sticks.
|
||||
- **Rebooting an encrypted host**: use `sudo fdesetup authrestart`. FileVault's pre-boot unlock doubles as the login, so the machine returns with a **live logged-in console session** and encryption intact, no password prompt, and supervision can then bring the VMs back by itself. Verified 2026-07-30. A plain `reboot` parks at the lock screen and blacks out every VM until a human logs in.
|
||||
- **Guest**: set `autoLoginUser` plus a valid `/etc/kcpassword` (XOR-obfuscated password file, key `7D 89 52 23 D2 BC DE A3`, payload zero-padded to a multiple of 12). `sysadminctl -autologin` fails with `SACSetAutoLoginPassword error:22` on provisioned accounts, and a fresh guest has no Python, so generate the bytes on the controlling host and copy them in. Then `pmset -a displaysleep 0 sleep 0 disablesleep 1`, `defaults -currentHost write com.apple.screensaver idleTime 0`, `defaults write com.apple.screensaver askForPassword 0`, and `caffeinate` inside the guest. ⚠ `autoLoginUser` was observed being wiped by failed `sysadminctl -autologin` attempts; verify it after each boot until stable.
|
||||
- **Wallpaper**: animated "aerials" wallpaper is brutal over VNC. The provider lives in `~/Library/Application Support/com.apple.wallpaper/Store/Index.plist` under several keys (`AllSpacesAndDisplays:Desktop`, `:Idle`, and `SystemDefault:*` which is what the login/lock screen uses). Switch each `Provider` to `com.apple.wallpaper.choice.solid-color` with PlistBuddy and restart `WallpaperAgent`. The login-window copy is cached and only refreshes on a later login cycle.
|
||||
|
||||
### Remote GUI/SSH access to a guest (recipe, verified 2026-07-29)
|
||||
|
||||
The guest lives on the host-private NAT bridge, so remote access is guest-service + host-forward:
|
||||
|
||||
1. **In the macOS guest** (over ssh), use ONE mechanism, fully activated. The reliable form is Remote Management in a single kickstart call:
|
||||
```
|
||||
sudo .../RemoteManagement/ARDAgent.app/Contents/Resources/kickstart \
|
||||
-activate -configure -access -on \
|
||||
-clientopts -setvnclegacy -vnclegacy yes -setvncpw -vncpw <8-char-pw> \
|
||||
-allowAccessFor -allUsers -privs -all -restart -agent -menu
|
||||
```
|
||||
⚠ **Half-configured states authenticate but refuse the session.** Loading `com.apple.screensharing` while Remote Management is deactivated (or vice versa) produces an Apple-client error that names the wrong culprit: *"Screen Sharing is not permitted on <host>. Disable and re-enable Screen Sharing or Remote Management in System Settings"*. A raw-protocol client can still authenticate AND open a framebuffer in that state, so protocol-level tests pass while every Apple client fails. The remedy is exactly what the dialog says, done over ssh: `launchctl unload -w …screensharing.plist`, `kickstart -deactivate -configure -access -off`, `pkill screensharingd`, then the single activate call above.
|
||||
Notes: `launchctl enable system/com.apple.screensharing` fails with "Could not find service" on this build; `load -w` is the plain-Screen-Sharing path if you deliberately want it instead of Remote Management. Apple clients negotiate `RSA-SRP` (auth type 33) and the guest logs `Authentication: SUCCEEDED :: User Name: … :: Type: RSA-SRP` on success, which is the definitive server-side confirmation.
|
||||
2. **On the host**: a gateway port-forward makes the guest's 5900 reachable from the whole tailnet without per-client tunnels: self-authorize the host's own key, then `ssh -N -g -L 0.0.0.0:5901:<guest-ip>:5900 <user>@localhost` (nohup'd).
|
||||
⚠⚠ **NEVER forward on host port 5900.** If the host has Screen Sharing enabled (our testbed does, from the pre-upgrade checklist), launchd already owns 5900 socket-activated. The `ssh -L` bind then fails with "Address already in use" **while the tunnel process keeps running**, so every symptom of success is present (process alive, port answers, real RFB banner) yet **every connection reaches the HOST's login window, not the guest**. This cost us an hour: guest credentials failed against the host's screensharingd, which reads exactly like broken guest auth, and we chased the (real, but irrelevant) provisioned-account identity bug. Diagnostics that would have caught it instantly: `sudo lsof -nP -iTCP:5900 -sTCP:LISTEN` showing `launchd` rather than `ssh`, or the guest's own logs showing NO auth attempts during a failed login. Always use a distinct host port and verify with `lsof` that the forward owns it.
|
||||
⚠ `-g` binds all interfaces, so the forward is also visible on the host's LAN; the VNC layer still requires the account or VNC password. ⚠ The forward pins the guest IP, which changes per boot under plain NAT; re-point it after a guest reboot (the proper fix is a vmnet DHCP reservation, macOS 26 API, once we move off plain `VZNATNetworkDeviceAttachment`).
|
||||
Verified working: with the forward on 5901, both a provisioned account and a `sysadminctl`-created one authenticate successfully (RFB `SecurityResult` = 0) against the guest. The guest offers security types `[30, 33, 36, 2, 35]`, i.e. Apple DH/SRP **plus classic type 2**, so non-Apple VNC clients work with the legacy password once ARD's `-setvnclegacy` is set. (The host's screensharingd, by contrast, offered no type 2, which is itself a tell that you are talking to the wrong machine.)
|
||||
3. **SSH from any tailnet device**: `ssh -J <host-user>@<host> codeman@<guest-ip>` (jump through the host), after adding the connecting machine's key to the guest's `authorized_keys`.
|
||||
|
||||
**Client-version incompatibility (macOS 27 servers vs older Screen Sharing clients)**: an older Mac's Screen Sharing client fails Apple's `RSA-SRP` handshake against macOS 27 servers, logging `Authentication: FAILED :: User Name: <user> :: Type: RSA-SRP` server-side, while a macOS 27 client authenticates against the same servers without issue. This was verified against BOTH a macOS 27 guest and a macOS 27 host with the operator's own account, so it is a client-side version skew, not configuration, and no server-side change fixes it. Same family as the documented "macOS 26 host cannot install a 27 guest" bug. Practical workaround: bypass Apple auth entirely with classic VNC auth (security type 2), which macOS offers only when Remote Management legacy VNC is enabled. Two ways to consume it: any third-party VNC client, or a browser via noVNC.
|
||||
|
||||
**Browser-based access chain (zero client install, version-proof)**, all hosted on the Mac:
|
||||
```
|
||||
browser --HTTP/WS--> websockify (+ noVNC static files)
|
||||
--> type-2-only proxy # rewrites the server's security-type list to [2]
|
||||
--> ssh -L forward # loopback hop; see the Local Network note below
|
||||
--> guest:5900
|
||||
```
|
||||
Notes learned the hard way: (a) **never bind the forward on host port 5900** (see the launchd warning above); (b) a Python proxy cannot reach the guest subnet directly because macOS **Local Network privacy** denies headless CLI binaries, surfacing as `No route to host`, so point the proxy at a loopback `ssh -L` forward instead (Apple-signed `ssh` is unaffected); (c) noVNC needs `?resize=scale` or Scaling Mode → Local Scaling, otherwise a Retina host screen (2940x1912) is unusable in a browser window; (d) noVNC speaks security type 2 only, which is exactly why the proxy rewrite is needed.
|
||||
|
||||
**Debugging technique that settled all of this**: a ~80-line Python RFB client (scratchpad `vnclogin.py`) that implements Apple DH auth (security type 30) and continues through `ClientInit`/`ServerInit`. It reports the server's `SecurityResult` plus the framebuffer size and desktop name, which separates "credentials rejected" from "authorized but session refused" without any GUI client. Pair it with `log stream --predicate 'process == "screensharingd"'` inside the guest, and drive a REAL Apple client headlessly from the host with `sudo launchctl asuser <uid> sudo -u <user> osascript -e 'tell application "Screen Sharing" to open location "vnc://user:pass@host:port"'`, verifying the result via `lsof -nP -iTCP -a -p <pid>` (an ESTABLISHED socket to the target) since `screencapture` fails on a lid-closed laptop ("could not create image from display"). Tailscale was never implicated: both the raw client and Apple's client work over the tailnet address once the guest service is fully activated.
|
||||
|
||||
### Hard-won operational lessons (write these into any tooling)
|
||||
|
||||
- **Silent serial is normal, not failure.** Debian's GRUB/kernel log to the graphics console; nothing attaches a getty to hvc0 by default. The reliable boot signal is the DHCP lease (or passive `tcpdump -i bridge100`), never the serial port and never a quick ping (BSD ping's first packet often dies to ARP latency; passive capture showed "dead" guests alive).
|
||||
- **DHCP lease entries carry truth**: `name=` shows the guest hostname, and the lease timestamps order events; stale entries linger, so compare timestamps before attributing a lease to a boot.
|
||||
- **Never boot a base image read-write.** Every RW boot mutates it (dhclient lease cache, journal, cloud-init state) and destroys experiment reproducibility, exactly why the production design only ever boots bases under overlays. Provision INTO the base once at base-build time, or provision per-case overlays with the seed, then detach the seed.
|
||||
- **A killed SSH client does not kill a remote `nohup`'d VM**, and the survivor holds the EFI variable store lock: "The EFI variable store is already in use" (`VZErrorDomain 50002`) means a zombie VM process, `pkill` it.
|
||||
- **EFI variable stores are per-VM state.** Fresh stores boot reliably; reuse across different VM instances is at minimum suspect on this beta (Apple's own guidance for cloned VMs is one store per VM). Cheap policy: one store per case, created with the overlay, deleted with it.
|
||||
- **Downloads from cloud.debian.org mirrors truncate silently**; always verify byte count against origin `Content-Length` and resume with `curl -C -`.
|
||||
- The remote host's default shell is zsh: `=` -prefixed words (`echo ===`) explode via zsh's `=cmd` expansion; keep separators zsh-safe in automation.
|
||||
|
||||
### The 2-concurrent-macOS-VM cap: TESTED AND CONFIRMED on macOS 27 beta 4 (2026-07-29)
|
||||
|
||||
We measured it, which as far as we can tell nobody had published for macOS 27. Method: `cp -c -R` the guest bundle (APFS clonefile, instant and **zero additional disk**), regenerate the machine identifier per clone (`VZMacMachineIdentifier()` written to `machine.id`; the hardware model is reused), then launch VMs until one is refused.
|
||||
|
||||
Result: VM #1 (8 GB, GUI) and VM #2 (4 GB, headless) ran concurrently without complaint. VM #3 was refused **instantly** at `vm.start`:
|
||||
|
||||
```
|
||||
VZErrorDomain Code=6 "The maximum supported number of active virtual machines has been reached."
|
||||
NSLocalizedFailure = "The number of virtual machines exceeds the limit."
|
||||
```
|
||||
|
||||
**This is a licensing/kernel quota, not a resource limit**: the refusal came with **39% of system memory free** on a 16 GB host, and adding RAM or CPU cannot raise it. It matches the pre-27 behavior (`hv_apple_isa_vm_quota`), so nothing changed in 27 despite the framework's other additions. Linux guests are unaffected and are bounded only by host resources.
|
||||
|
||||
Design consequences: macOS-guest capacity per host is **hard-capped at 2**, so a GUI-macOS-per-case product must schedule around it (queue, evict idle VMs, or scale across hosts) and surface it in the UI. Also relevant: the acknowledged slot-leak bug (a guest-initiated shutdown failing to release a slot, recoverable only by host reboot) is far more damaging under a cap of 2 than it sounds; we did not reproduce it on beta 4, but any scheduler should treat "slot appears used but nothing is running" as a real state.
|
||||
|
||||
### Not yet tested
|
||||
- Cache layers (`LayerType.cache`), `.overlay(blockCount:)` disk growth, stack depth performance, VirtioFS + stack combination, `truncate`, ASIF disks for macOS guests (raw used so far; ASIF has the reclamation bug).
|
||||
- One more scripting lesson from this session: inner `ssh` calls inside a piped `sh -s` script MUST use `-n`, or they consume the remainder of the script from stdin and it silently never runs.
|
||||
|
||||
### Session timeline (what was actually established, 2026-07-29)
|
||||
|
||||
Linux path: base image download (with resume, mirrors truncate) → `vzboot` compiles against the beta SDK first try → EFI boot → NAT DHCP lease → cloud-init seed provisions a user with the host's SSH key → `ssh` into the guest works → DiskImageKit stack boots with an ASIF overlay taking all writes while the base stays SHA-identical. Later Linux boots became unreliable on an un-rebooted host (silent hangs, 0% CPU, no DHCP); a clean-baseline retest is still pending.
|
||||
|
||||
macOS path: seed-CDN IPSW (matched to the host build) → `VZMacOSInstaller` restore, ~25 min, first try → first boot with `VZMacGuestProvisioningOptions` creates an admin account with Remote Login on, no interaction needed, SSH reachable ~140 s later → key bootstrap over a one-time password login → guest shutdown/relaunch clean (the slot-leak bug did not reproduce) → GUI access fought through a port collision, a client-version incompatibility, the rendering dependency, and a self-inflicted session kill, ending with a browser-based path plus a guest hardened to auto-login and never lock.
|
||||
|
||||
**Lifecycle verified (stop → start), 2026-07-30**: an in-guest `shutdown -h now` fires `guestDidStop` and the runner app exits on its own; relaunching from the same bundle boots the guest in ~2 minutes straight into an auto-logged-in desktop, and the VM slot is released cleanly (an immediate restart works, so the slot-leak bug did not bite). Two operational notes: the guest takes a **new NAT lease on every boot**, so any port-forward must be re-pointed (or use a vmnet DHCP reservation), and a host reboot resets `pmset -a disablesleep`.
|
||||
|
||||
⚠ **Provisioning does NOT skip the per-user first-login assistant.** `VZMacGuestProvisioningOptions` skips the initial Setup Assistant (account creation, region, Apple Account) so the machine is immediately reachable, but the first time anyone actually logs into a desktop, macOS still presents its per-user wizard (Apple Intelligence, Siri, privacy, appearance, Touch ID). The operator hit exactly this. For a GUI-first product this MUST be pre-suppressed during base-image creation by writing `com.apple.SetupAssistant` keys for every account that will log in, and into `/System/Library/User Template/English.lproj/Library/Preferences/` so accounts created later inherit it.
|
||||
|
||||
⚠ **A partial key list is worse than none**, because the wizard simply shows the panes you missed and the operator has to click through them again after every fresh login (we hit this twice). The set that finally silenced macOS 27 beta 4: `DidSeeCloudSetup`, `DidSeeSiriSetup`, `DidSeePrivacy`, `DidSeeAppearanceSetup`, `DidSeeTouchIDSetup`, `DidSeeAvatarSetup`, `DidSeeScreenTime`, `DidSeeApplePaySetup`, `DidSeeSafariImport`, `DidSeeAccessibility`, **`DidSeeActivationLock`, `DidSeeAppStore`, `DidSeeLockdownMode`** (the three easy to miss), plus the Express-Settings flags **`SkipExpressSettingsUpdating`** and **`SkipFirstLoginOptimization`**, and the version markers `LastSeenCloudProductVersion` / `LastSeenBuddyBuildVersion` / `PreviousSystemVersion` / `PreviousBuildVersion` matching the guest build. Verify afterwards by reading the domain back and checking that no `DidSee*` key is still `0`. Note these keys change between macOS releases, so base-image creation should re-verify per OS version rather than trust a hardcoded list.
|
||||
|
||||
## 9. Design implications for Codeman's VM subsystem
|
||||
|
||||
0. **GUI is a first-class mode, and for macOS guests it is the whole point (owner decision, 2026-07-29).** The subsystem serves real desktops, not only headless SSH boxes. macOS guests are GUI-only in practice (nothing renders without an attached view). Linux guests are supported in BOTH modes: GUI when the case wants a desktop, headless-over-SSH when it wants a cheap agent sandbox. The costs of the GUI path are in §8 "Display rendering": VMs as GUI apps in a live session, a host that never locks, guests that auto-login with their first-login wizard pre-suppressed, and the macOS concurrency cap as a real capacity limit.
|
||||
1. **The macOS-specific liabilities are accepted costs, not reasons to avoid macOS guests**: provisioning is macOS-only and first-boot-only, ASIF space reclamation is broken for macOS guests on the beta (use RAW disks for macOS guests until fixed), and the 2-VM cap applies. Plan around each: RAW-backed macOS disks, provisioning baked into base-image creation, and capacity limits surfaced in the UI.
|
||||
2. **Base immutability is not just hygiene, it is load-bearing**: DiskImageKit's UUID invalidation plus our sha-stability proof make a read-only shared base per image-generation the core artifact. Bases are built once (seed attached), then only ever opened `.readOnly` under per-case overlays.
|
||||
3. **Seed ISOs are a base-build-time tool only.** Never attach a seed to a routine case boot (correlated with boot hangs on the beta, and semantically wrong anyway since cloud-init already ran).
|
||||
4. **Per-case files**: overlay ASIF + EFI variable store live and die together with the case.
|
||||
5. **Export = ship the layer chain** (base ref + overlay + manifest), not flatten; there is no flatten API. In-guest `dd` to a fresh image is the fallback for a true single-file export.
|
||||
6. **Health checking must be lease/API based**, not serial/ping based, and Codeman's `codeman-vm status` should read `/var/db/dhcpd_leases` (or use vmnet DHCP reservations for deterministic per-case IPs, a macOS 26 API).
|
||||
7. **Run `fstrim` periodically in Linux guests** (or mount with discard) so overlays stay sparse.
|
||||
8. **Entitlements plist stays minimal** (exactly `com.apple.security.virtualization`) to dodge the AMFI/watchdog traps.
|
||||
9. **Expect beta churn**: pin findings to build numbers (this doc: 26A5388g) and retest each beta; the framework binaries changed every beta so far.
|
||||
10. **A macOS guest is only "ready" when its desktop is ready**, which is a stricter bar than "the VM booted". Readiness means: VM app running with a live WindowServer connection, guest auto-logged-in (not at a login or lock screen), first-login assistant suppressed, and the guest's screen sharing serving a non-black framebuffer. Health checks should sample the framebuffer for non-black content, because every failure mode in this session (headless run, locked host, dead WindowServer, locked guest, setup wizard) presents as a perfectly healthy-looking process with a black or useless screen.
|
||||
10b. **Supervision must run as a root LaunchDaemon.** A user LaunchAgent cannot launch a GUI app into the Aqua session; its restarts fail silently (child dies instantly, empty log, supervisor reports success). Root + `launchctl asuser <uid> sudo -u <user> …` works and the launched process persists. This bit us on the first supervisor implementation and is easy to repeat.
|
||||
|
||||
11. **Remote-access plumbing belongs in the helper CLI, not in ad-hoc shell**: a `codeman-vm` implementation should own port selection (never 5900), forward lifecycle across guest IP changes (or better, vmnet DHCP reservations for stable per-case IPs), and a documented browser path, because every failure in this session came from hand-rolled plumbing rather than from the Virtualization APIs themselves.
|
||||
12. **Never let control-plane connectivity depend on a GUI session** on a remote Mac host: prefer a Tailscale system service over the App Store app, and keep a LAN-adjacent peer able to relay as an out-of-band recovery path.
|
||||
|
||||
## Sources
|
||||
|
||||
Apple DocC JSON backend (diskimagekit, virtualization, vmnet trees; macOS 27 release notes) | WWDC26 session 224 https://developer.apple.com/videos/play/wwdc2026/224/ | eclecticlight.co ASIF/virtualization coverage | developer.apple.com/forums threads 839343 (CSIdentity bug), 830118 (cross-version restore), 830119 (VM-slot leak), 830383 (VM cap), 834822 + 831902 (USB entitlements), 822658 (vmnet loopback) | openai/tart issues 1261/1263/1268/1269/1285 | Spooky-Labs provisioning design doc | VirtualBuddy 2.2 release notes | lima-vm discussions | our own test transcripts on the testbed (`~/vm-lab/*.log`, this repo's session)
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
# Web Tabs (dashboards as Codeman tabs)
|
||||
|
||||
Open any dashboard you run, Grafana, Uptime Kuma, Portainer, a status page on port
|
||||
4000, as a tab beside your Claude/Codex/Gemini sessions. Codeman becomes one mission
|
||||
4000, as a tab beside your Claude/Codex/Antigravity sessions. Codeman becomes one mission
|
||||
control instead of Codeman plus a pile of browser tabs.
|
||||
|
||||
## Using it
|
||||
|
||||
+52
-11
@@ -116,6 +116,14 @@ GEMINI_SEARCH_PATHS=(
|
||||
"$HOME/bin/gemini"
|
||||
)
|
||||
|
||||
# Antigravity CLI search paths (from src/utils/antigravity-cli-resolver.ts)
|
||||
ANTIGRAVITY_SEARCH_PATHS=(
|
||||
"$HOME/.local/bin/agy"
|
||||
"$HOME/.antigravity/bin/agy"
|
||||
"/usr/local/bin/agy"
|
||||
"$HOME/bin/agy"
|
||||
)
|
||||
|
||||
# ============================================================================
|
||||
# Color Output
|
||||
# ============================================================================
|
||||
@@ -493,6 +501,34 @@ get_gemini_path() {
|
||||
done
|
||||
}
|
||||
|
||||
check_antigravity() {
|
||||
if command -v agy &>/dev/null; then
|
||||
return 0
|
||||
fi
|
||||
|
||||
for path in "${ANTIGRAVITY_SEARCH_PATHS[@]}"; do
|
||||
if [[ -x "$path" ]]; then
|
||||
return 0
|
||||
fi
|
||||
done
|
||||
|
||||
return 1
|
||||
}
|
||||
|
||||
get_antigravity_path() {
|
||||
if command -v agy &>/dev/null; then
|
||||
command -v agy
|
||||
return
|
||||
fi
|
||||
|
||||
for path in "${ANTIGRAVITY_SEARCH_PATHS[@]}"; do
|
||||
if [[ -x "$path" ]]; then
|
||||
echo "$path"
|
||||
return
|
||||
fi
|
||||
done
|
||||
}
|
||||
|
||||
check_cloudflared() {
|
||||
# Check ~/.local/bin first (matches tunnel-manager.ts resolution order)
|
||||
if [[ -x "$HOME/.local/bin/cloudflared" ]]; then
|
||||
@@ -1993,11 +2029,12 @@ main() {
|
||||
fi
|
||||
fi
|
||||
|
||||
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini)
|
||||
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini, Antigravity)
|
||||
local has_claude=false
|
||||
local has_opencode=false
|
||||
local has_codex=false
|
||||
local has_gemini=false
|
||||
local has_antigravity=false
|
||||
|
||||
info "Checking AI CLI tools..."
|
||||
if check_claude; then
|
||||
@@ -2016,17 +2053,21 @@ main() {
|
||||
has_gemini=true
|
||||
success "Gemini CLI found at $(get_gemini_path)"
|
||||
fi
|
||||
if check_antigravity; then
|
||||
has_antigravity=true
|
||||
success "Antigravity CLI found at $(get_antigravity_path)"
|
||||
fi
|
||||
|
||||
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" ]]; then
|
||||
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" && "$has_antigravity" == "false" ]]; then
|
||||
echo ""
|
||||
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, or Gemini."
|
||||
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, Antigravity, or Gemini."
|
||||
headless_guard "install an AI CLI (curl | bash from its vendor)"
|
||||
echo ""
|
||||
echo -e " ${BOLD}Which AI CLI would you like to install?${NC}"
|
||||
echo -e " ${CYAN}1)${NC} Claude Code (Anthropic)"
|
||||
echo -e " ${CYAN}2)${NC} OpenCode (open-source)"
|
||||
echo -e " ${CYAN}3)${NC} Both"
|
||||
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex or Gemini)"
|
||||
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex or Antigravity)"
|
||||
echo ""
|
||||
|
||||
local cli_choice=""
|
||||
@@ -2071,8 +2112,8 @@ main() {
|
||||
|
||||
if [[ "$cli_choice" == "4" ]]; then
|
||||
warn "Skipping AI CLI install. Codeman will run, but sessions need a CLI to drive."
|
||||
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
|
||||
info " or: npm install -g @google/gemini-cli (Gemini)"
|
||||
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
|
||||
info " or: curl -fsSL https://antigravity.google/cli/install.sh | bash (Antigravity)"
|
||||
elif [[ "$has_claude" == "false" ]] && [[ "$has_opencode" == "false" ]]; then
|
||||
die "The selected AI CLI failed to install. Install one manually and re-run the installer."
|
||||
fi
|
||||
@@ -2372,12 +2413,12 @@ main() {
|
||||
echo -e " https://github.com/Ark0N/Codeman"
|
||||
echo ""
|
||||
|
||||
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini; then
|
||||
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini && ! check_antigravity; then
|
||||
echo -e " ${YELLOW}${BOLD}Reminder:${NC} Install at least one AI CLI to start using Codeman:"
|
||||
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
|
||||
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
|
||||
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
|
||||
echo -e " ${CYAN}npm install -g @google/gemini-cli${NC} # Gemini"
|
||||
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
|
||||
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
|
||||
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
|
||||
echo -e " ${CYAN}curl -fsSL https://antigravity.google/cli/install.sh | bash${NC} # Antigravity"
|
||||
echo ""
|
||||
fi
|
||||
|
||||
|
||||
Generated
+2
-2
@@ -1,12 +1,12 @@
|
||||
{
|
||||
"name": "aicodeman",
|
||||
"version": "1.11.1",
|
||||
"version": "1.14.0",
|
||||
"lockfileVersion": 3,
|
||||
"requires": true,
|
||||
"packages": {
|
||||
"": {
|
||||
"name": "aicodeman",
|
||||
"version": "1.11.1",
|
||||
"version": "1.14.0",
|
||||
"hasInstallScript": true,
|
||||
"license": "MIT",
|
||||
"workspaces": [
|
||||
|
||||
+3
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "aicodeman",
|
||||
"version": "1.11.1",
|
||||
"version": "1.14.0",
|
||||
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
|
||||
"type": "module",
|
||||
"main": "dist/index.js",
|
||||
@@ -55,6 +55,7 @@
|
||||
"anthropic",
|
||||
"opencode",
|
||||
"codex",
|
||||
"antigravity",
|
||||
"gemini-cli",
|
||||
"ai-agents",
|
||||
"agent",
|
||||
@@ -158,6 +159,7 @@
|
||||
"dist",
|
||||
"scripts/postinstall.js",
|
||||
"scripts/fix-node-pty.mjs",
|
||||
"skills",
|
||||
"LICENSE",
|
||||
"README.md"
|
||||
]
|
||||
|
||||
@@ -16,7 +16,7 @@
|
||||
|
||||
> ### Made for [**Codeman**](https://getcodeman.com)
|
||||
>
|
||||
> This overlay is the local echo engine of [**Codeman**](https://github.com/Ark0N/Codeman), mission control for AI coding agents: run and monitor a dozen Claude Code, Codex, OpenCode and Gemini sessions at once, watch their subagents work in live floating windows, let them run autonomously overnight, and drive all of it from your phone.
|
||||
> This overlay is the local echo engine of [**Codeman**](https://github.com/Ark0N/Codeman), mission control for AI coding agents: run and monitor a dozen Claude Code, Codex, OpenCode and Antigravity sessions at once, watch their subagents work in live floating windows, let them run autonomously overnight, and drive all of it from your phone.
|
||||
>
|
||||
> That last part is why this library exists. The demo below is a real Codeman session on two phones.
|
||||
|
||||
|
||||
@@ -0,0 +1,274 @@
|
||||
---
|
||||
name: codeman
|
||||
description: >-
|
||||
Drive Codeman, the session manager this agent is running inside, over its HTTP API:
|
||||
list sessions, start worker sessions, send them prompts, block until they finish
|
||||
(wait / wait-output / send-and-wait), read their output, and clean up. Use when asked
|
||||
to orchestrate or parallelize work across Codeman sessions, watch another session, or
|
||||
start and manage workers. Only usable inside a Codeman-managed session
|
||||
(CODEMAN_MUX=1); refuse to act otherwise.
|
||||
---
|
||||
|
||||
# Driving Codeman from inside a session
|
||||
|
||||
You are an agent running inside a Codeman-managed terminal session. Codeman is the
|
||||
server that spawned you; its HTTP API can start, prompt, watch, and delete other
|
||||
sessions. Every recipe below was verified live. Full endpoint tables and
|
||||
troubleshooting: [reference/endpoints.md](reference/endpoints.md). Worked multi-worker
|
||||
flows: [reference/recipes.md](reference/recipes.md).
|
||||
|
||||
## 0. Guard — run this before anything else
|
||||
|
||||
```bash
|
||||
test "${CODEMAN_MUX:-}" = 1 || { echo "Not inside a Codeman-managed session; refusing to act."; exit 1; }
|
||||
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
|
||||
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
|
||||
# Codeman does NOT hand a session the server password. If one is set, the two
|
||||
# in-reach copies are the data dir's .env (the same fallback `codeman attach`
|
||||
# uses — hand-authored; nothing ever writes it) and the supervisor definition
|
||||
# that install.sh wrote the password into, which is where a stock
|
||||
# password-protected install actually keeps it. The data dir is wherever the
|
||||
# hook-secret file lives. Values may be quoted or `export`-prefixed.
|
||||
ENV_FILE="${CODEMAN_HOOK_SECRET_FILE:+${CODEMAN_HOOK_SECRET_FILE%hook-secret}.env}"
|
||||
envval() { sed -n "s/^\(export \)\{0,1\}$1=//p" "$ENV_FILE" | tail -1 | sed 's/^"\(.*\)"$/\1/; s/^'\''\(.*\)'\''$/\1/'; }
|
||||
if [ -z "${CODEMAN_PASSWORD:-}" ] && [ -n "$ENV_FILE" ] && [ -f "$ENV_FILE" ]; then
|
||||
CODEMAN_USERNAME=$(envval CODEMAN_USERNAME)
|
||||
CODEMAN_PASSWORD=$(envval CODEMAN_PASSWORD)
|
||||
fi
|
||||
if [ -z "${CODEMAN_PASSWORD:-}" ]; then # stock installs: install.sh puts it in the service definition
|
||||
UNIT="$HOME/.config/systemd/user/codeman-web.service"
|
||||
PLIST="$HOME/Library/LaunchAgents/com.codeman.web.plist"
|
||||
if [ -f "$UNIT" ]; then
|
||||
CODEMAN_PASSWORD=$(sed -n 's/^Environment="CODEMAN_PASSWORD=\(.*\)"$/\1/p' "$UNIT" | head -1)
|
||||
elif [ -f "$PLIST" ]; then
|
||||
CODEMAN_PASSWORD=$(awk '/<key>CODEMAN_PASSWORD<\/key>/{getline; print}' "$PLIST" | sed -n 's/.*<string>\(.*\)<\/string>.*/\1/p')
|
||||
fi
|
||||
fi
|
||||
AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:$CODEMAN_PASSWORD")
|
||||
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on https (self-signed cert)
|
||||
```
|
||||
|
||||
- If `CODEMAN_MUX` is not `1`, **stop and say so**. Do not guess an API URL; a server
|
||||
you are not part of is not yours to drive.
|
||||
- **A 401 is plain text, not the JSON envelope**, so on a password-protected server
|
||||
every `jq` in these recipes dies with `jq: parse error` instead of showing
|
||||
`UNAUTHORIZED`. If that happens, check the status with `-w '%{http_code}'`; if it
|
||||
is 401 and neither fallback above found a credential, **stop and tell the user
|
||||
you need credentials**. The hook-secret bypass covers only `/api/hook-event` and
|
||||
`/api/status-telemetry`, never session control.
|
||||
- These endpoints first ship in Codeman **1.13.0**, but do not gate on the version
|
||||
number: a dev build can serve them while reporting an older version. Probe
|
||||
instead: `GET .../wait` on a real session id answering 404 with an `.error`
|
||||
starting `Route ` means the server predates the wait endpoints (fall back to
|
||||
polling `GET .../terminal?tail=` and say so); `Session ... not found` means your
|
||||
session id is wrong, not the server.
|
||||
|
||||
## 1. Safety rules — read before any mutating call
|
||||
|
||||
You are yourself a session on this server, and the API has **no undo**.
|
||||
|
||||
- **Never act on your own session — and know that this check is the ONLY guard.**
|
||||
The server has no self-protection: a session that DELETEs its own id succeeds and
|
||||
dies silently (verified live). Session ids appear in both full and 8-character
|
||||
forms (Docker cases export a truncated `$SELF`; mux names and UI surfaces carry
|
||||
8-char ids), so compare by prefix **in both directions**, never by equality:
|
||||
|
||||
```bash
|
||||
is_self() { case "$1" in "$SELF"*) return 0 ;; esac; case "$SELF" in "$1"*) return 0 ;; esac; return 1; }
|
||||
```
|
||||
|
||||
One-directional or equality checks each miss a real combination (full `$SELF` vs
|
||||
a target you transcribed in 8-char form, or truncated `$SELF` vs a full target)
|
||||
and the miss deletes you. Check `is_self` before every `DELETE`, kill, respawn,
|
||||
or input call.
|
||||
- **Mutating calls you may make unprompted** (this is an allowlist):
|
||||
`POST /api/v1/quick-start`, `POST /api/v1/sessions/:id/input`, and
|
||||
`DELETE /api/v1/sessions/:id` **only** for a session you created in this
|
||||
conversation, by exact id. Keep a list of the ids you create. Everything else
|
||||
mutating needs the user to have asked for it.
|
||||
- **Never call these** unless the user explicitly asked, naming the target:
|
||||
- `DELETE /api/cases/:name` — recursively **deletes a real directory of the user's
|
||||
code** from disk. One wrong case name destroys work that was never yours.
|
||||
- `DELETE /api/sessions` (no id) and `DELETE /api/subagents` (no id) — bulk kills.
|
||||
- respawn / ralph / orchestrator / cron mutations — respawn runs `/clear` (wipes a
|
||||
conversation), orchestrator state is a single global slot, cron jobs outlive you.
|
||||
- `PUT /api/settings`, `POST /api/system/update` — global UI settings; server restart.
|
||||
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. The API is the only interface.
|
||||
- Sessions count against a 50-session cap and case creation is uncapped: clean up every
|
||||
session you start, and don't retry `quick-start` in a loop.
|
||||
|
||||
## 2. Rules of the road
|
||||
|
||||
- **End every input with `\r`** — literally the two characters `\r` inside the JSON
|
||||
string. Codeman types the text and sends Enter **only when the input contains a
|
||||
carriage return**; without it your command sits unsubmitted on the worker's prompt
|
||||
and everything downstream times out. `{"input":"run the tests\r",...}`. No response
|
||||
field catches this: `delivered:true` means "written to the pane", **not**
|
||||
"submitted" — a `\r`-less send still reports `delivered:true` and then every wait
|
||||
times out, which is why the loops below are bounded and check the terminal.
|
||||
- **Single-line input only.** Newlines are stripped; one line per call.
|
||||
- **Build request bodies with `jq -n` for any prompt you did not author as a
|
||||
literal.** The inline `-d '{"input":"'"$P"'\r"}'` pattern breaks on the first
|
||||
double quote, backslash, or `$` in a real prompt:
|
||||
|
||||
```bash
|
||||
BODY=$(jq -n --arg p "$PROMPT" '{input:($p+"\r"),useMux:true,clientId:"agent-1",seq:1,wait:true,waitTimeout:60000}')
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' --data-binary "$BODY"
|
||||
```
|
||||
- **Exactly-once delivery**: always send a stable `clientId` and a monotonic
|
||||
per-session `seq` on `POST .../input`. A retry after a dropped connection then
|
||||
cannot double-type the prompt. Increment `seq` for each NEW input; reuse the same
|
||||
pair only to re-ask about the same delivery.
|
||||
- **Envelope**: success is `{"success":true,"data":…}`, errors are
|
||||
`{"success":false,"error","errorCode"}`. Read `.data`. Use `/api/v1/*` paths.
|
||||
- **A wait timeout is HTTP 200**, `{wait:{timedOut:true,signal:null}}` — not an error.
|
||||
Loop over short waits (60 s); proxies cut long-idle connections. Timeouts are
|
||||
**clamped** (ceiling 600 s): read back `wait.timeoutMs` for what was applied.
|
||||
- **`stop` and `blocked` fire for `claude` sessions only** (Claude Code hooks). On
|
||||
`shell`/`opencode`/`codex`/`gemini`/`antigravity`, requesting them explicitly is a
|
||||
400 — and lifecycle transitions there are coarse (a short shell command may emit
|
||||
**no** `idle` transition at all, verified live), so synchronize those modes with
|
||||
output markers, not signals.
|
||||
- **Your typed command echoes into the output stream**, so a marker that appears
|
||||
verbatim in the input line matches **before the command runs**. Always split the
|
||||
marker (recipe below), keep it unique per call, and use `from=buffer` so a marker
|
||||
that printed before your wait landed is still found. Matching is literal — no regex.
|
||||
- **Match single space-free tokens against TUI output.** A full-screen TUI (claude,
|
||||
codex, …) positions text with cursor movements, not literal spaces, so the stripped
|
||||
stream can read `Yes,Itrustthisfolder` and a multi-word match is unreliable there —
|
||||
whether a phrase keeps its spaces depends on how the TUI happened to draw it
|
||||
(observed live: some match, some never fire). Plain command output (shell workers,
|
||||
`echo` lines) keeps real spaces.
|
||||
|
||||
## 3. Recipes (each verified live)
|
||||
|
||||
**List sessions / find yourself** — metadata only, safe to poll:
|
||||
|
||||
```bash
|
||||
"${CURL[@]}" "$API/api/v1/sessions" | jq '.data[] | {id, name, mode, status}'
|
||||
"${CURL[@]}" "$API/api/v1/sessions" | jq --arg s "$SELF" '.data[] | select(.id | startswith($s))'
|
||||
```
|
||||
|
||||
**Start a claude worker and wait until it is actually ready.** A new session reports
|
||||
`idle` before its CLI has spawned, and a brand-new case shows a **trust dialog**
|
||||
first, so neither "wait for idle" nor "wait for ❯" means ready (the trust dialog
|
||||
contains `❯` too — observed live). Codeman *can* auto-accept that dialog itself, but
|
||||
the accept rides a stream match that misses on some runs (both outcomes seen live),
|
||||
so wait for the composer first and handle the dialog only as the bounded fallback —
|
||||
never send a blind Enter up front (if auto-accept already fired, it lands in the
|
||||
composer). Stage 1 is short on purpose: an already-trusted case matches `bypass` in
|
||||
under a second, while a **virgin case can never pass stage 1** (the dialog is up, so
|
||||
the composer is not) and always pays it in full before the fallback runs — the long
|
||||
budget belongs to stage 3, after the dialog is answered:
|
||||
|
||||
```bash
|
||||
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
|
||||
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
|
||||
for _ in $(seq 1 30); do # bounded: a bad SID would otherwise poll forever
|
||||
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
|
||||
done
|
||||
# ⚠️ pid != null proves STARTUP only, never life: a worker that later dies inside
|
||||
# its pane keeps status "idle" and a pid (the local tmux attach client, not the
|
||||
# worker). The death check is wait?until=exit, below.
|
||||
CID="agent-$$"; SEQ=1
|
||||
# the composer's status bar ("bypass permissions on") is the ready marker — Codeman
|
||||
# spawns claude in bypass mode. Single-token matches only: TUI text is space-less.
|
||||
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
|
||||
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
|
||||
# composer never appeared → the trust dialog is probably still up; accept it once
|
||||
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
|
||||
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
|
||||
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
|
||||
SEQ=$((SEQ+1))
|
||||
fi
|
||||
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
|
||||
jq -e '.data.wait.matched' <<<"$R" >/dev/null || \
|
||||
{ echo "worker $SID never became ready; inspect terminal?tail="; }
|
||||
fi
|
||||
```
|
||||
|
||||
**Send a prompt and wait for the turn to finish** (claude workers — the call to
|
||||
prefer). It registers the waiter *before* typing, closing the race where a separate
|
||||
wait sees the previous turn's idle state. Loop by resending the **identical** request:
|
||||
the repeat is a tagged duplicate (same `clientId`+`seq`) that does not retype but
|
||||
answers from the session's current state. Verified: the stop hook resolves this in
|
||||
seconds; a duplicate resend answers in ~20 ms without retyping.
|
||||
|
||||
```bash
|
||||
for TRY in $(seq 1 10); do # BOUNDED: a \r-less send never produces a signal and resends are no-op duplicates
|
||||
R=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
|
||||
-d '{"input":"run the tests, then summarize in one line\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ',"wait":true,"waitTimeout":60000}')
|
||||
if jq -e '.data.wait.timedOut' <<<"$R" >/dev/null; then
|
||||
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
|
||||
| jq -r '.data.terminalBuffer' | tail -5 # two straight timeouts: prompt sitting unsubmitted?
|
||||
continue
|
||||
fi
|
||||
# Resolved — but a duplicate answering immediately reports the session's CURRENT
|
||||
# state ("it is idle now"), NOT that a new turn ran. A \r-less send lands exactly
|
||||
# here on try 2 (verified live), so check the terminal before believing it:
|
||||
if jq -e '.data.duplicate and .data.wait.immediate' <<<"$R" >/dev/null; then
|
||||
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" | jq -r '.data.terminalBuffer' | tail -5
|
||||
# your prompt still on the ❯ composer line = never submitted (missing \r);
|
||||
# submit it with {"input":"\r"} (the only recovery), then loop again
|
||||
fi
|
||||
break
|
||||
done
|
||||
SEQ=$((SEQ+1)); jq '.data.wait.signal, .data.status' <<<"$R"
|
||||
```
|
||||
|
||||
Read the outcome in this order: `wait.signal != null` → done (`stop` is definitive;
|
||||
`idle` is heuristic) — **unless** it arrived as `duplicate:true` + `immediate:true`,
|
||||
which only says the session is idle *now* and must be confirmed from the terminal
|
||||
(above); `wait.timedOut` → loop again (bounded); `wait.ended` → session gone, stop.
|
||||
If the loop exhausts its cap, do not keep looping: read the terminal, report what
|
||||
you see, and remember that a still-typed-but-unsubmitted prompt (missing `\r`) can
|
||||
only be recovered by submitting it with `{"input":"\r"}`.
|
||||
|
||||
**Shell worker + completion marker** — the pattern for `shell` mode (no hooks there).
|
||||
The typed line must not contain the marker verbatim (the input echo would match
|
||||
instantly — observed live), so build it with a variable the worker's shell expands:
|
||||
|
||||
```bash
|
||||
N="${RANDOM}_$$"; MARK="DONE_$N" # unique per call: tmux repaints replay old text
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
|
||||
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}'
|
||||
SEQ=$((SEQ+1))
|
||||
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode "match=$MARK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=120000' \
|
||||
| jq -r '.data.wait | {matched, snippet}'
|
||||
```
|
||||
|
||||
The typed line shows `${M}_…`, the real output shows `DONE_… rc=<exit code>`, and the
|
||||
snippet carries the exit code back to you.
|
||||
|
||||
**Read a worker's output** — the terminal buffer, tail in **bytes** (`textOutput` in
|
||||
`GET .../output` stays empty for interactive sessions; don't use it):
|
||||
|
||||
```bash
|
||||
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=3000" | jq -r '.data.terminalBuffer' \
|
||||
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g' | grep -v '^[[:space:]]*$' | tail -30
|
||||
```
|
||||
|
||||
Avoid `?full=1` (entire tmux scrollback, a context bomb) unless doing a post-mortem.
|
||||
|
||||
**Detect a dead worker cheaply**: `GET .../wait?until=exit&timeout=60000` answers
|
||||
immediately (`signal:"exit"`, `immediate:true`) if the PTY is gone — including a
|
||||
worker that exited *inside* its pane, which `GET .../sessions/:id` keeps reporting
|
||||
as `status:"idle"` with a pid (that pid is the local tmux attach client, not the
|
||||
worker). The wait routes are the only liveness check; a worker dying while a wait
|
||||
is parked resolves it within ~3 s. A session deleted mid-wait resolves in ~1 s.
|
||||
|
||||
**Clean up** — only ids you created, `is_self`-checked, one at a time:
|
||||
|
||||
```bash
|
||||
is_self "$SID" || "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
|
||||
```
|
||||
|
||||
Everything else (endpoint tables, per-mode signal table, error codes, capacity
|
||||
limits, Docker/remote caveats): [reference/endpoints.md](reference/endpoints.md).
|
||||
Fan-out orchestration and blocked-worker handling:
|
||||
[reference/recipes.md](reference/recipes.md).
|
||||
@@ -0,0 +1,214 @@
|
||||
# Codeman API reference for agents
|
||||
|
||||
Loaded on demand from the `codeman` skill. Assumes the guard variables from SKILL.md
|
||||
(`$API`, `$SELF`, `"${CURL[@]}"`). Canonical contract: `docs/api-reference.md` in the
|
||||
Codeman repo; this file is the agent-relevant subset, verified live.
|
||||
|
||||
## Envelope and errors
|
||||
|
||||
Every JSON response: `{"success":true,"data":…}` or
|
||||
`{"success":false,"error":"…","errorCode":"…"}`. Branch on `errorCode`:
|
||||
|
||||
| `errorCode` | HTTP | Meaning |
|
||||
|-------------|------|---------|
|
||||
| `INVALID_INPUT` | 400 | malformed request; the message names the bad field |
|
||||
| `UNAUTHORIZED` | 401 | auth required or failed (send `-u user:password`). ⚠️ The 401 body is plain text, NOT this envelope — `jq` dies with a parse error, see the guard in SKILL.md |
|
||||
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
|
||||
| `SESSION_BUSY` | 409 | this session's waiter cap (16, combined signal+output) is full |
|
||||
| `CONFLICT` / `ALREADY_EXISTS` | 409 | conflicts with current state |
|
||||
| `OPERATION_FAILED` | 422 | well-formed but could not be completed |
|
||||
| `RATE_LIMITED` | 429 | per-owner or process-wide waiter pool is full — back off; switching sessions will not help |
|
||||
| `INTERNAL_ERROR` | 500 | server bug |
|
||||
|
||||
`SESSION_BUSY` vs `RATE_LIMITED` on the wait endpoints is deliberate: the first means
|
||||
"too many waiters on *this* session", the second means the *pool* is full.
|
||||
|
||||
## Sessions
|
||||
|
||||
| Task | Call |
|
||||
|------|------|
|
||||
| list sessions (metadata only, ~1.5 KB each, safe to poll) | `GET /api/v1/sessions` |
|
||||
| one session (has `.data.pid`, `null` until the PTY spawns) | `GET /api/v1/sessions/:id` — ⚠️ **not a liveness check**: a worker that dies inside its pane keeps `status:"idle"` and a pid (the tmux attach client); `wait?until=exit` is the death check |
|
||||
| unified list incl. history | `GET /api/v1/sessions/unified` → `.data.sessions[]` (NOT `.data[]`), and it folds in transcript history from the whole machine — never use it to verify cleanup; `GET /api/v1/sessions` is the cleanup check |
|
||||
| start case + session in one call | `POST /api/v1/quick-start` |
|
||||
| send input | `POST /api/v1/sessions/:id/input` |
|
||||
| read terminal (tail is in **BYTES**, raw ANSI) | `GET /api/v1/sessions/:id/terminal?tail=3000` → `.data.terminalBuffer` |
|
||||
| full tmux scrollback (context bomb; post-mortems only) | `GET /api/v1/sessions/:id/terminal?full=1` |
|
||||
| background agents of a session | `GET /api/v1/subagents` |
|
||||
| server status / version | `GET /api/v1/status` → `.data.version` |
|
||||
| delete one session (yours, `is_self`-checked) | `DELETE /api/v1/sessions/:id` |
|
||||
|
||||
⚠️ `GET /api/v1/sessions/:id/output` → `.data.textOutput` looks like the obvious read
|
||||
but stays **empty for interactive tmux-backed sessions** (it is fed only by the legacy
|
||||
JSON-stream path). Verified empty on live claude and shell sessions. Read
|
||||
`terminal?tail=` instead and strip ANSI:
|
||||
|
||||
```bash
|
||||
… | jq -r '.data.terminalBuffer' | sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g'
|
||||
```
|
||||
|
||||
`POST /api/v1/quick-start` body (all optional):
|
||||
`{"caseName":"worker-1","mode":"claude","sessionName":"w9-worker","effort":"high"}`
|
||||
— `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity`; response is
|
||||
`.data.{sessionId, caseName, casePath}`. Creates the case directory (a real directory
|
||||
on the user's disk) if missing — do not retry it in a loop, and remember the name.
|
||||
|
||||
`POST /api/v1/sessions/:id/input` body:
|
||||
`{"input":"one line\r","useMux":true,"clientId":"agent-1","seq":1}` plus optionally
|
||||
`"wait"` / `"waitTimeout"` (below).
|
||||
|
||||
- ⚠️ **The input must contain `\r`** (the JSON escape, i.e. a real carriage return)
|
||||
**or Enter is never sent**: the text is typed onto the worker's prompt and sits
|
||||
there unsubmitted. Verified live — this is the number-one silent failure, and no
|
||||
response field catches it: `delivered:true` means "written to the pane", not
|
||||
"submitted". A `\r`-less send with `wait` reports `delivered:true` and then every
|
||||
wait on that turn times out. Without `wait`, fire-and-forget returns an **empty**
|
||||
`{"success":true,"data":{}}` — no `delivered`, no `duplicate`; those fields exist
|
||||
only on the `wait` variant, so a fire-and-forget flow gets no delivery
|
||||
confirmation at all.
|
||||
- `input` must be single-line (newlines are stripped). To send a bare Enter (confirm
|
||||
a dialog), send `{"input":"\r"}`.
|
||||
- `clientId`+`seq` give exactly-once delivery: the server applies each pair at most
|
||||
once. Increment `seq` per new input.
|
||||
|
||||
## The wait primitives
|
||||
|
||||
Three bounded long-polls. Shared semantics:
|
||||
|
||||
- **Timeout = HTTP 200** with `wait.timedOut:true`. Loop over short waits (60 s);
|
||||
`tailscale serve` / cloudflared cut idle connections.
|
||||
- Timeouts are **clamped** to `[1000, 600000]` ms (operator-tunable); the applied
|
||||
value is echoed as `wait.timeoutMs` — read it back, never assume.
|
||||
- All three nest the result under `.data.wait`, same shape, so one helper parses all.
|
||||
- `.data.status` (post-wait `SessionStatus`) and `.data.limitPaused` ride along.
|
||||
`limitPaused:true` means the session is paused on a usage limit and will emit
|
||||
nothing until reset — a timeout is then *expected*; do not retry hard, and do not
|
||||
kill the worker.
|
||||
|
||||
### Signals by mode
|
||||
|
||||
| Signal | Meaning | Available for |
|
||||
|--------|---------|---------------|
|
||||
| `idle` | output stabilized + prompt detected — heuristic, can flap mid-turn | every mode |
|
||||
| `working` | session started producing output | every mode |
|
||||
| `stop` | Claude Code `stop` hook — the definitive end-of-turn | `claude` only |
|
||||
| `blocked` | `permission_prompt` / `elicitation_dialog` hook — the worker needs an answer | `claude` only |
|
||||
| `exit` | PTY exited or session deleted | every mode |
|
||||
|
||||
Default `until` set: `stop,idle,exit`. On non-claude modes the server silently drops
|
||||
`stop`/`blocked` from the *default* set (echoed back as `wait.until`, e.g.
|
||||
`["idle","exit"]` on shell); requesting them *explicitly* there is a 400 naming the
|
||||
mode. ⚠️ On hook-less modes the lifecycle signals are also **coarse in practice**: a
|
||||
short shell command produced **no** `idle` transition within 60 s (verified live), so
|
||||
a `fresh=1` / fresh-delivery wait can burn its whole timeout while the work finished
|
||||
long ago. Synchronize hook-less modes with `wait-output` markers instead.
|
||||
|
||||
Two more places hooks go missing even in claude mode: **Docker cases** need
|
||||
`CODEMAN_DOCKER_BRIDGE_HOOKS=1` on the server (without it only `idle`/`working`/
|
||||
`exit` arrive), and **remote-SSH cases** run the agent on another host whose hooks may
|
||||
never reach this server. When unsure, ask for `stop,idle,exit`.
|
||||
|
||||
⚠️ **Signals are edge-triggered with no history.** A signal that fires while no
|
||||
waiter is registered is gone; no later wait can observe it (`until=stop` on a worker
|
||||
whose turn already ended just times out, with or without `fresh` — verified live).
|
||||
Register the waiter before the event can happen: send-and-wait does exactly that,
|
||||
and `wait-output` markers with `from=buffer` are latched by construction. Never
|
||||
fire-and-forget N prompts and then gather signal-waits worker by worker; every
|
||||
worker that finishes before its gather is unobservable (see recipes.md Flow 3b).
|
||||
|
||||
### `GET /api/v1/sessions/:id/wait`
|
||||
|
||||
| Param | Default | Notes |
|
||||
|-------|---------|-------|
|
||||
| `until` | `stop,idle,exit` | comma list; unknown token → 400 naming it |
|
||||
| `timeout` | 60000 | ms, clamped; applied value echoed as `wait.timeoutMs` |
|
||||
| `fresh` | `0` | `1` requires an actual *transition*, ignoring the state at call time |
|
||||
|
||||
⚠️ A session whose PTY has not spawned (`pid:null`) or has exited counts as `exit`
|
||||
**right now**: with the default set the call answers immediately
|
||||
(`signal:"exit", immediate:true`). That is how you detect a dead worker cheaply — but
|
||||
it also means "wait for my just-created session" needs the readiness recipe in
|
||||
SKILL.md, not this endpoint.
|
||||
|
||||
### `GET /api/v1/sessions/:id/wait-output`
|
||||
|
||||
| Param | Default | Notes |
|
||||
|-------|---------|-------|
|
||||
| `match` | required | literal substring, 1–200 chars, ANSI-stripped; chunk-straddling matches found; **no regex** — a `regex=` param is a 400 |
|
||||
| `nocase` | `0` | case-insensitive compare; snippet keeps original casing |
|
||||
| `from` | `now` | `buffer` scans the tail (~256 KB) of existing output first |
|
||||
| `timeout` | 60000 | same clamp |
|
||||
|
||||
Four traps, all observed live:
|
||||
|
||||
1. **The echo of your own typed command is output.** A marker appearing verbatim in
|
||||
the input line matches the moment the text is typed, before the command runs.
|
||||
Split the marker with a shell variable: send `M=DONE; …; echo ${M}_1234\r`, wait
|
||||
on `DONE_1234`.
|
||||
2. **`from=now` misses text printed before the wait landed** — a marker echoed just
|
||||
before the request registered timed out at full length. After sending a command,
|
||||
always wait with `from=buffer`.
|
||||
3. **`from=now` can also match too much**: tmux repaints old screen content as
|
||||
ordinary output on attach/resize/redraw, so a *generic* marker (`BUILD OK`)
|
||||
matches stale text. Unique-per-call markers (`DONE_$RANDOM`) make both `from`
|
||||
modes safe.
|
||||
4. **TUI output can be space-less in the stream.** Full-screen TUIs (claude, codex,
|
||||
…) position words with cursor-movement escapes rather than literal spaces, so
|
||||
the stripped stream can read `Yes,Itrustthisfolder` while the pane shows the
|
||||
spaced phrase. Whether a given phrase keeps its spaces depends on how the TUI
|
||||
drew it (observed live: some multi-word matches fire, some never do), so treat
|
||||
multi-word matches against TUI screens as unreliable and match a **single
|
||||
space-free token** (`trust`, `bypass`). Plain command output (shell workers,
|
||||
`echo` lines) keeps real spaces and multi-word matches work there.
|
||||
|
||||
Build the query with `-G --data-urlencode` (a `+` in a hand-built query decodes to a
|
||||
space). Result extras: `wait.matched`, `wait.match`, `wait.snippet` (bounded window
|
||||
around the match, blank runs collapsed — the snippet is often all you need to read).
|
||||
|
||||
### `POST /api/v1/sessions/:id/input` with `wait`
|
||||
|
||||
| Field | Notes |
|
||||
|-------|-------|
|
||||
| `wait` | `true` (default signal set) or the same comma grammar as `until`; absent = historical fire-and-forget |
|
||||
| `waitTimeout` | ms, same clamp |
|
||||
|
||||
Registers the waiter **before** typing, which closes the race where send-then-wait
|
||||
sees the previous turn's idle state and returns instantly. Response adds `delivered`
|
||||
and `duplicate` beside the standard `wait` object.
|
||||
|
||||
A **tagged duplicate** (same `clientId`+`seq` already applied) does not retype but
|
||||
still honors `wait`, answering from the session's *current* state instead of
|
||||
requiring a new transition (`delivered:false, duplicate:true` — verified: ~20 ms,
|
||||
command ran exactly once). That is what makes the resend-identical-request loop in
|
||||
SKILL.md correct: iteration 1 delivers and needs a transition; later iterations
|
||||
resolve immediately if the turn ended in between. ⚠️ The flip side: a duplicate's
|
||||
`immediate:true` answer is the current state and nothing more — an idle worker
|
||||
whose prompt was never submitted (missing `\r`) produces the same
|
||||
`signal:"idle", immediate:true` as one that finished the turn. Confirm from
|
||||
`terminal?tail=` before reporting success; SKILL.md's loop shows where.
|
||||
|
||||
### Outcome parsing, in order
|
||||
|
||||
1. `wait.signal != null` (or `wait.matched == true`) — the thing happened.
|
||||
`wait.immediate:true` rides along and means the condition already held at call
|
||||
time; if that is not what you meant, you wanted `fresh=1` or send-and-wait.
|
||||
2. `wait.timedOut` — poll boundary; loop again.
|
||||
3. `wait.ended` — session deleted/torn down mid-wait; stop looping.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Symptom | Cause / fix |
|
||||
|---------|-------------|
|
||||
| every curl fails with a certificate error | you dropped `-k`; `CODEMAN_API_URL` is HTTPS with a self-signed cert |
|
||||
| `jq: parse error` on every call | plain-text 401s: the server has a password. Check with `-w '%{http_code}'`, use the guard's `.env` fallback, and if no `.env` exists, stop and ask the user for credentials |
|
||||
| input arrives but nothing happens; later waits all time out | the input had no `\r`, so Enter was never sent; the text is sitting on the worker's prompt. **Submitting it with `{"input":"\r"}` is the ONLY recovery** — Ctrl+U (0x15) and Esc do NOT clear the composer (verified live) — and the flush costs one turn in which the worker reasons about the junk; open the next real prompt with "ignore the garbled line above:" |
|
||||
| `GET .../sessions/$CODEMAN_SESSION_ID` 404s | Docker case: the env id is truncated to 8 chars; find yourself with `startswith($SELF)`, and always self-compare by prefix, in both directions |
|
||||
| `CODEMAN_MUX` unset but you seem to be in a session | remote-SSH case: the env vars are not exported there. Fail closed — refuse to act |
|
||||
| connection refused from inside a container | a loopback-bound server is unreachable from a container, and `CODEMAN_DOCKER_BRIDGE_HOOKS=1` does **not** fix that: it opens a hooks-only listener, so hook events start flowing but `/api/v1/*` stays refused. Driving the API from inside a Docker case needs a reachable bind (an operator decision); report it, don't retry |
|
||||
| wait routes 404 on a valid session id | read the `.error` text: a `Route ...` prefix means the server predates the wait endpoints (< 1.13.0; a dev build can serve them while reporting an older version, so probe, never version-compare) — poll `terminal?tail=` and say so. `Session ... not found` means your id is wrong, not the server |
|
||||
| wait on `stop` never resolves | non-claude mode, or hooks not reaching the server (Docker/remote), or a case created by Codeman < 1.13.0 against an `--https` install (its hook curls lacked `-k` and TLS-failed silently; a 1.13.0+ server rewrites them the next time a session starts in that case). Use markers or `idle,exit` |
|
||||
| new claude worker ignores its first prompt | it was showing the first-run trust dialog and Codeman's auto-accept missed; use the readiness recipe in SKILL.md (wait for `bypass` first, accept the dialog only as the bounded fallback) |
|
||||
| `wait-output` times out although the pane shows the text | multi-word match against a TUI screen; the stream has no spaces there — match one token |
|
||||
| `wait-output` matched instantly with stale text | generic marker + tmux repaint; use `DONE_$RANDOM` |
|
||||
| 409 `SESSION_BUSY` on a wait | too many concurrent waiters on that session (cap 16 combined); reuse one wait per worker |
|
||||
| 429 `RATE_LIMITED` on a wait | global/owner waiter pool full; back off, do not switch sessions |
|
||||
@@ -0,0 +1,249 @@
|
||||
# Worked orchestration flows
|
||||
|
||||
Loaded on demand from the `codeman` skill. Every flow assumes the guard preamble from
|
||||
SKILL.md ran (`$API`, `$SELF`, `"${CURL[@]}"`, `is_self`). Track every session id you
|
||||
create; delete them (and only them) when done. Remember the two silent killers:
|
||||
**every input ends with `\r`**, and **markers must be split** so the typed-line echo
|
||||
does not match them.
|
||||
|
||||
## Flow 1: claude worker, end to end
|
||||
|
||||
Start a worker, get it truly ready (trust dialog included), give it a task, wait for
|
||||
the turn to finish, read the answer, clean up. Verified live: the stop hook resolves
|
||||
the send-and-wait within seconds of the turn ending.
|
||||
|
||||
```bash
|
||||
# 1. start (returns before the CLI inside is ready)
|
||||
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
|
||||
-d '{"caseName":"worker-tests","mode":"claude"}' | jq -r '.data.sessionId')
|
||||
CREATED+=("$SID") # the cleanup list
|
||||
CID="agent-$$"; SEQ=1
|
||||
|
||||
# 2. readiness. "wait for idle" or "wait for ❯" is NOT readiness: a fresh session
|
||||
# reports idle before anything spawned, and the first-run trust dialog contains ❯.
|
||||
# Codeman CAN auto-accept that dialog, but the accept misses on some runs (both
|
||||
# outcomes seen live), so: composer marker first, dialog only as the bounded
|
||||
# fallback (a blind Enter up front would land in an already-ready composer).
|
||||
# Stage 1 is SHORT on purpose: an already-trusted case matches in <1 s, while a
|
||||
# virgin case can never pass it (the dialog is up) and always pays it in full —
|
||||
# the long budget belongs to stage 3, after the dialog is answered.
|
||||
# Single-token matches only: TUI text is space-less in the stream.
|
||||
for _ in $(seq 1 30); do
|
||||
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
|
||||
done
|
||||
# (pid != null proves startup only — a worker that later dies inside its pane keeps
|
||||
# status "idle" and a pid. The death check is wait?until=exit.)
|
||||
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
|
||||
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
|
||||
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
|
||||
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
|
||||
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
|
||||
SEQ=$((SEQ+1))
|
||||
fi
|
||||
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
|
||||
jq -e '.data.wait.matched' <<<"$R" >/dev/null || echo "worker $SID not ready; inspect terminal?tail="
|
||||
fi
|
||||
|
||||
# 3. send-and-wait, looping on the IDENTICAL request (tagged duplicate: no retype).
|
||||
# BOUNDED (a \r-less send would otherwise loop forever), body built with jq -n so
|
||||
# quotes/backslashes/$ in a real prompt survive; note the appended \r.
|
||||
PROMPT='run the unit tests and summarize failures in one line'
|
||||
BODY=$(jq -n --arg p "$PROMPT" --arg c "$CID" --argjson s "$SEQ" \
|
||||
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:60000}')
|
||||
for TRY in $(seq 1 10); do
|
||||
R=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' --data-binary "$BODY")
|
||||
if jq -e '.data.wait.timedOut' <<<"$R" >/dev/null; then
|
||||
jq -e '.data.limitPaused' <<<"$R" >/dev/null && sleep 60 # usage-limit pause: silence is expected
|
||||
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
|
||||
| jq -r '.data.terminalBuffer' | tail -5 # is the prompt sitting unsubmitted?
|
||||
continue
|
||||
fi
|
||||
# Resolved — but duplicate + immediate is only "the session is idle NOW", which a
|
||||
# never-submitted (\r-less) prompt also produces. Check before believing it:
|
||||
if jq -e '.data.duplicate and .data.wait.immediate' <<<"$R" >/dev/null; then
|
||||
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
|
||||
| jq -r '.data.terminalBuffer' | tail -5
|
||||
# prompt still on the ❯ composer line = never submitted; {"input":"\r"} is the
|
||||
# only recovery, then loop again
|
||||
fi
|
||||
break
|
||||
done
|
||||
SEQ=$((SEQ+1))
|
||||
|
||||
# 4. interpret
|
||||
case "$(jq -r '.data.wait.signal' <<<"$R")" in
|
||||
stop) : ;; # definitive end of turn
|
||||
idle) : ;; # heuristic — and if it rode a duplicate with
|
||||
# immediate:true, it proves nothing ran (step 3)
|
||||
exit) echo "worker died" ;;
|
||||
null) jq -e '.data.wait.ended' <<<"$R" >/dev/null && echo "worker deleted mid-wait" ;;
|
||||
esac
|
||||
|
||||
# 5. read the answer: terminal tail (BYTES), ANSI-stripped. textOutput stays empty
|
||||
# for interactive sessions; terminal?full=1 is a context bomb.
|
||||
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=4000" | jq -r '.data.terminalBuffer' \
|
||||
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g' | grep -v '^[[:space:]]*$' | tail -30
|
||||
|
||||
# 6. clean up — exact id, own list only, self-check
|
||||
is_self "$SID" || "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
|
||||
```
|
||||
|
||||
Increment `SEQ` for every *new* input to the same worker. Reuse the same `SEQ` only to
|
||||
re-ask about the same delivery (the duplicate-wait loop above).
|
||||
|
||||
## Flow 2: shell worker running a build, marker-synchronized
|
||||
|
||||
`shell` sessions have no hooks (`stop`/`blocked` are a 400 there), and their lifecycle
|
||||
signals are coarse — a short command may emit no `idle` transition at all (verified
|
||||
live), so send-and-wait can burn its whole timeout. The reliable pattern is a split,
|
||||
unique marker plus `wait-output from=buffer`:
|
||||
|
||||
```bash
|
||||
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
|
||||
-d '{"caseName":"builder","mode":"shell"}' | jq -r '.data.sessionId')
|
||||
CREATED+=("$SID")
|
||||
for _ in $(seq 1 30); do
|
||||
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
|
||||
done
|
||||
|
||||
# Split marker: the typed line carries ${M}_N, only the OUTPUT carries DONE_N.
|
||||
# An unsplit marker matches the echo of your own keystrokes before the build runs.
|
||||
N="${RANDOM}_$$"; MARK="DONE_$N"
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
|
||||
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"build-'$$'","seq":1}'
|
||||
|
||||
for TRY in $(seq 1 30); do # BOUNDED (30 min): a \r-less send makes an uncapped loop infinite
|
||||
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode "match=$MARK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=60000')
|
||||
jq -e '.data.wait.matched' <<<"$R" >/dev/null && break
|
||||
jq -e '.data.wait.ended' <<<"$R" >/dev/null && { echo "worker gone"; break; }
|
||||
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
|
||||
| jq -r '.data.terminalBuffer' | tail -5 # command still sitting unsubmitted?
|
||||
done
|
||||
jq -r '.data.wait.snippet' <<<"$R" # e.g. "DONE_123_456 rc=0" — the exit code rides the marker line
|
||||
```
|
||||
|
||||
## Flow 3: fan out N workers, gather as each finishes
|
||||
|
||||
Start everything first, then gather. One in-flight wait per worker — the per-session
|
||||
waiter cap is 16 and abandoned concurrent waits pile up against it.
|
||||
|
||||
```bash
|
||||
declare -A WORKER MARKS
|
||||
for task in lint typecheck unit; do
|
||||
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
|
||||
-d '{"caseName":"fan-'"$task"'","mode":"shell"}' | jq -r '.data.sessionId')
|
||||
WORKER[$task]=$SID; CREATED+=("$SID")
|
||||
done
|
||||
for task in "${!WORKER[@]}"; do
|
||||
SID=${WORKER[$task]}
|
||||
for _ in $(seq 1 30); do
|
||||
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
|
||||
done
|
||||
N="${task}_${RANDOM}"; MARKS[$task]="DONE_$N"
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
|
||||
-d '{"input":"M=DONE; npm run '"$task"'; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"fan-'$$'","seq":1}'
|
||||
done
|
||||
for task in "${!WORKER[@]}"; do # sequential gather; each wait blocks until that worker is done
|
||||
for TRY in $(seq 1 30); do # BOUNDED per worker, same reasoning as Flow 2
|
||||
R=$("${CURL[@]}" -G "$API/api/v1/sessions/${WORKER[$task]}/wait-output" \
|
||||
--data-urlencode "match=${MARKS[$task]}" --data-urlencode 'from=buffer' --data-urlencode 'timeout=60000')
|
||||
jq -e '.data.wait.matched or .data.wait.ended' <<<"$R" >/dev/null && break
|
||||
done
|
||||
echo "$task: $(jq -r '.data.wait.snippet // "worker gone"' <<<"$R" | tail -1)"
|
||||
done
|
||||
```
|
||||
|
||||
## Flow 3b: fan out N CLAUDE workers
|
||||
|
||||
Send-and-wait is synchronous, so the shell-flow shape ("send everything, then
|
||||
gather") does not translate directly: the send *is* the wait, and worker 2's prompt
|
||||
would not go out until worker 1's turn ended. Two working patterns, both verified
|
||||
live (and one anti-pattern, measured failing, replaced by B):
|
||||
|
||||
**A. Background the send-and-waits** (simplest; each resolved on `stop` while the
|
||||
other was still running):
|
||||
|
||||
```bash
|
||||
sendwait() { # $1=sid $2=prompt $3=seq — assumes the worker passed Flow 1's readiness
|
||||
local body; body=$(jq -n --arg p "$2" --argjson s "$3" \
|
||||
'{input:($p+"\r"),useMux:true,clientId:"fan-'$$'",seq:$s,wait:true,waitTimeout:600000}')
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$1/input" \
|
||||
-H 'Content-Type: application/json' --data-binary "$body" > "/tmp/fan-$1.json"
|
||||
}
|
||||
( sendwait "$SID1" 'refactor module A and reply DONE' 2 & \
|
||||
sendwait "$SID2" 'write tests for module B and reply DONE' 2 & wait )
|
||||
jq -c '.data.wait | {signal, waitedMs}' /tmp/fan-"$SID1".json /tmp/fan-"$SID2".json
|
||||
```
|
||||
|
||||
One in-flight wait per worker keeps you far from the 16-per-session waiter cap.
|
||||
|
||||
**B. Fire-and-forget, then gather with output markers.** If you must send every
|
||||
prompt before waiting on anything, do **not** gather with signal waits: signals
|
||||
are edge-triggered with no history, so a `stop` that fires before the gather
|
||||
reaches that worker is gone and unobservable afterwards — `fresh=1` cannot help,
|
||||
and neither can omitting it (measured: worker 2's turn ended at +2 s, its
|
||||
sequential `until=stop,exit&fresh=1` gather burned its full bounded 300 s and
|
||||
reported nothing). Gather instead on a marker each worker prints itself, which
|
||||
`from=buffer` re-finds no matter when it appeared:
|
||||
|
||||
```bash
|
||||
# SIDS[1], SIDS[2] = worker ids that already passed Flow 1's readiness.
|
||||
# The typed prompt must NOT contain the finished marker verbatim (your keystrokes
|
||||
# echo into the output stream and would match instantly), so ask for it in halves:
|
||||
declare -A TOK
|
||||
for i in 1 2; do
|
||||
TOK[$i]="${RANDOM}_$i"
|
||||
BODY=$(jq -n --arg p "do task $i; when completely done print the word WORKDONE immediately followed by _${TOK[$i]}" \
|
||||
--arg c "fan-$$" --argjson s 2 '{input:($p+"\r"),useMux:true,clientId:$c,seq:$s}')
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/${SIDS[$i]}/input" \
|
||||
-H 'Content-Type: application/json' --data-binary "$BODY"
|
||||
done
|
||||
for i in 1 2; do # order no longer matters: the marker is latched in the buffer
|
||||
"${CURL[@]}" -G "$API/api/v1/sessions/${SIDS[$i]}/wait-output" \
|
||||
--data-urlencode "match=WORKDONE_${TOK[$i]}" --data-urlencode 'from=buffer' \
|
||||
--data-urlencode 'timeout=600000' | jq -c '.data.wait | {matched, snippet}'
|
||||
done
|
||||
```
|
||||
|
||||
Use A unless you genuinely need to send everything before waiting on anything: A
|
||||
needs no marker discipline, and resolves on the definitive `stop` instead of on
|
||||
the worker remembering to print a token.
|
||||
|
||||
## Flow 4: watch for a worker stuck on a permission prompt
|
||||
|
||||
Claude workers can block on a permission dialog. `blocked` is a wait signal
|
||||
(claude-mode only), so watch for it and surface the question to the user instead of
|
||||
guessing an answer:
|
||||
|
||||
```bash
|
||||
R=$("${CURL[@]}" "$API/api/v1/sessions/$SID/wait?until=stop,blocked,exit&timeout=60000")
|
||||
if [ "$(jq -r '.data.wait.signal' <<<"$R")" = blocked ]; then
|
||||
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" | jq -r '.data.terminalBuffer' \
|
||||
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' | grep -v '^[[:space:]]*$' | tail -15
|
||||
# show this to the user and ask how to answer; do NOT auto-confirm another
|
||||
# session's permission prompt
|
||||
fi
|
||||
```
|
||||
|
||||
## Cleanup discipline
|
||||
|
||||
At the end of the conversation (or on abort), delete exactly what you created:
|
||||
|
||||
```bash
|
||||
for id in "${CREATED[@]}"; do
|
||||
is_self "$id" || "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$id"
|
||||
done
|
||||
```
|
||||
|
||||
- Only ids from your own `CREATED` list. Never enumerate `/api/v1/sessions` and
|
||||
delete by pattern; other sessions belong to the user.
|
||||
- If you created a *case* purely as scratch and the user confirmed it is disposable,
|
||||
`DELETE /api/v1/cases/:name` removes it — but that recursively deletes the
|
||||
directory from disk, so never do it without the user's explicit go-ahead for that
|
||||
exact name.
|
||||
+31
-3
@@ -135,6 +135,7 @@ export abstract class AiCheckerBase<
|
||||
// Active check state
|
||||
protected checkMuxName: string | null = null;
|
||||
protected checkTempFile: string | null = null;
|
||||
protected checkStderrFile: string | null = null;
|
||||
protected checkPromptFile: string | null = null;
|
||||
protected checkPollTimer: NodeJS.Timeout | null = null;
|
||||
protected checkTimeoutTimer: NodeJS.Timeout | null = null;
|
||||
@@ -376,6 +377,7 @@ export abstract class AiCheckerBase<
|
||||
const shortId = this.sessionId.slice(0, 8);
|
||||
const timestamp = Date.now();
|
||||
this.checkTempFile = join(tmpdir(), `${this.tempFilePrefix}-${shortId}-${timestamp}.txt`);
|
||||
this.checkStderrFile = join(tmpdir(), `${this.tempFilePrefix}-stderr-${shortId}-${timestamp}.txt`);
|
||||
this.checkPromptFile = join(tmpdir(), `${this.tempFilePrefix}-prompt-${shortId}-${timestamp}.txt`);
|
||||
this.checkMuxName = `${this.muxNamePrefix}${shortId}`;
|
||||
|
||||
@@ -386,6 +388,7 @@ export abstract class AiCheckerBase<
|
||||
|
||||
// Ensure output temp file exists (empty) so we can poll it
|
||||
writeFileSync(this.checkTempFile, '');
|
||||
writeFileSync(this.checkStderrFile, '');
|
||||
|
||||
// Write prompt to file to avoid E2BIG error (argument list too long)
|
||||
// The prompt can be 16KB+ which exceeds shell argument limits
|
||||
@@ -396,7 +399,7 @@ export abstract class AiCheckerBase<
|
||||
const modelArg = `--model "${this.config.model.replace(/"/g, '\\"')}"`;
|
||||
const augmentedPath = getAugmentedPath();
|
||||
const claudeCmd = `cat "${this.checkPromptFile}" | claude -p ${modelArg} --output-format text`;
|
||||
const fullCmd = `export PATH="${augmentedPath}"; ${claudeCmd} > "${this.checkTempFile}" 2>&1; echo "${this.doneMarker}" >> "${this.checkTempFile}"; rm -f "${this.checkPromptFile}"`;
|
||||
const fullCmd = `export PATH="${augmentedPath}"; ${claudeCmd} > "${this.checkTempFile}" 2> "${this.checkStderrFile}"; echo "${this.doneMarker}" >> "${this.checkTempFile}"; rm -f "${this.checkPromptFile}"`;
|
||||
|
||||
// Spawn tmux session
|
||||
try {
|
||||
@@ -461,18 +464,32 @@ export abstract class AiCheckerBase<
|
||||
const output = content.replace(this.doneMarker, '').trim();
|
||||
|
||||
if (!output) {
|
||||
return this.createErrorResult(`Empty output from ${this.checkDescription}`, durationMs);
|
||||
const stderr = this.readStderrDiagnostic();
|
||||
const detail = stderr ? `: ${stderr}` : '';
|
||||
return this.createErrorResult(`Empty output from ${this.checkDescription}${detail}`, durationMs);
|
||||
}
|
||||
|
||||
// Delegate to subclass for verdict parsing
|
||||
const parsed = this.parseVerdict(output);
|
||||
if (!parsed) {
|
||||
return this.createErrorResult(`Could not parse verdict from: "${output.substring(0, 100)}"`, durationMs);
|
||||
const stderr = this.readStderrDiagnostic();
|
||||
const detail = stderr ? `; stderr: "${stderr}"` : '';
|
||||
return this.createErrorResult(`Could not parse verdict from: "${output.substring(0, 100)}"${detail}`, durationMs);
|
||||
}
|
||||
|
||||
return this.createResult(parsed.verdict, parsed.reasoning, durationMs);
|
||||
}
|
||||
|
||||
private readStderrDiagnostic(): string {
|
||||
if (!this.checkStderrFile || !existsSync(this.checkStderrFile)) return '';
|
||||
|
||||
try {
|
||||
return readFileSync(this.checkStderrFile, 'utf-8').trim().substring(0, 200);
|
||||
} catch {
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
private cleanupCheck(): void {
|
||||
// Clear poll timer
|
||||
if (this.checkPollTimer) {
|
||||
@@ -509,6 +526,17 @@ export abstract class AiCheckerBase<
|
||||
this.checkTempFile = null;
|
||||
}
|
||||
|
||||
if (this.checkStderrFile) {
|
||||
try {
|
||||
if (existsSync(this.checkStderrFile)) {
|
||||
unlinkSync(this.checkStderrFile);
|
||||
}
|
||||
} catch {
|
||||
// Best effort cleanup
|
||||
}
|
||||
this.checkStderrFile = null;
|
||||
}
|
||||
|
||||
if (this.checkPromptFile) {
|
||||
try {
|
||||
if (existsSync(this.checkPromptFile)) {
|
||||
|
||||
+212
-49
@@ -21,6 +21,9 @@ import { getRalphLoop } from './ralph-loop.js';
|
||||
import { getStore } from './state-store.js';
|
||||
import { getErrorMessage } from './types.js';
|
||||
import { isSupportedAttachmentExtension } from './attachment-registry.js';
|
||||
import { daemonStatus, startDaemon, stopDaemon, type WebLaunchOptions } from './daemon-control.js';
|
||||
import { installService, serviceStatus, uninstallService } from './service-installer.js';
|
||||
import { isLoopbackBindHost, isUnauthenticatedNetworkAcknowledged } from './web/network-auth-policy.js';
|
||||
|
||||
const require = createRequire(import.meta.url);
|
||||
const pkg = require('../package.json') as { version: string };
|
||||
@@ -572,64 +575,224 @@ program
|
||||
console.log('');
|
||||
});
|
||||
|
||||
// ============ Web / daemon / service Commands ============
|
||||
|
||||
/** Shared option set for the commands that can launch a web server. */
|
||||
function addWebLaunchOptions(cmd: Command): Command {
|
||||
return cmd
|
||||
.option('-H, --host <host>', 'Host to bind to', process.env.CODEMAN_HOST || '127.0.0.1')
|
||||
.option('-p, --port <port>', 'Port to listen on (env: CODEMAN_PORT)', process.env.CODEMAN_PORT || '3000')
|
||||
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
|
||||
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
|
||||
.option(
|
||||
'--allow-unauthenticated-network',
|
||||
'Allow non-loopback web access without CODEMAN_PASSWORD (dangerous; terminal control is exposed)'
|
||||
)
|
||||
.option(
|
||||
'--multiuser',
|
||||
'Enable opt-in multi-user mode (named users in ~/.codeman/users.json; env: CODEMAN_MULTIUSER)'
|
||||
);
|
||||
}
|
||||
|
||||
/** Normalize commander's strings into the shape daemon-control/service-installer take. */
|
||||
function toWebLaunchOptions(options: {
|
||||
host: string;
|
||||
port: string;
|
||||
https?: boolean;
|
||||
titleHostname?: string;
|
||||
allowUnauthenticatedNetwork?: boolean;
|
||||
multiuser?: boolean;
|
||||
}): WebLaunchOptions {
|
||||
const port = parseInt(options.port, 10);
|
||||
if (!Number.isInteger(port) || port <= 0 || port > 65535) {
|
||||
console.error(chalk.red(`✗ Invalid port: ${options.port}`));
|
||||
process.exit(1);
|
||||
}
|
||||
return {
|
||||
host: options.host,
|
||||
port,
|
||||
https: !!options.https,
|
||||
titleHostname: options.titleHostname,
|
||||
allowUnauthenticatedNetwork: !!options.allowUnauthenticatedNetwork,
|
||||
multiuser: !!options.multiuser,
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* The server prints this itself, but into a log file nobody reads when it is
|
||||
* detached or supervised. Repeat it where the operator is actually looking.
|
||||
*/
|
||||
function warnIfUnauthenticatedNetwork(launch: WebLaunchOptions): void {
|
||||
if (isLoopbackBindHost(launch.host)) return;
|
||||
if (isUnauthenticatedNetworkAcknowledged(launch.allowUnauthenticatedNetwork)) return;
|
||||
console.log(
|
||||
chalk.yellow(
|
||||
`⚠ Binding ${launch.host} without CODEMAN_PASSWORD: anyone who can reach this port gets terminal control.`
|
||||
)
|
||||
);
|
||||
console.log(chalk.yellow(' Set CODEMAN_PASSWORD, or bind 127.0.0.1 and front it with tailscale serve.'));
|
||||
}
|
||||
|
||||
// Web interface command
|
||||
program
|
||||
.command('web')
|
||||
.description('Start the web interface')
|
||||
.option('-H, --host <host>', 'Host to bind to', process.env.CODEMAN_HOST || '127.0.0.1')
|
||||
.option('-p, --port <port>', 'Port to listen on (env: CODEMAN_PORT)', process.env.CODEMAN_PORT || '3000')
|
||||
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
|
||||
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
|
||||
.option(
|
||||
'--allow-unauthenticated-network',
|
||||
'Allow non-loopback web access without CODEMAN_PASSWORD (dangerous; terminal control is exposed)'
|
||||
)
|
||||
.option('--multiuser', 'Enable opt-in multi-user mode (named users in ~/.codeman/users.json; env: CODEMAN_MULTIUSER)')
|
||||
.action(async (options) => {
|
||||
// The flag is surfaced to the rest of the process via the env var so
|
||||
// isMultiUserMode() has a single source of truth (see config/multiuser.ts).
|
||||
if (options.multiuser) process.env.CODEMAN_MULTIUSER = '1';
|
||||
const { startWebServer } = await import('./web/server.js');
|
||||
const host = options.host;
|
||||
const port = parseInt(options.port, 10);
|
||||
const https = !!options.https;
|
||||
const titleHostname = options.titleHostname;
|
||||
const allowUnauthenticatedNetwork = !!options.allowUnauthenticatedNetwork;
|
||||
const protocol = https ? 'https' : 'http';
|
||||
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
|
||||
const webCmd = addWebLaunchOptions(program.command('web').description('Start the web interface'))
|
||||
.option('-d, --daemon', 'Run detached in the background; survives the shell, logs to <data dir>/web.log')
|
||||
.option('--stop', 'Stop a server started with --daemon')
|
||||
.option('--status', 'Report whether a detached server is running');
|
||||
|
||||
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
|
||||
webCmd.action(async (options) => {
|
||||
// The flag is surfaced to the rest of the process via the env var so
|
||||
// isMultiUserMode() has a single source of truth (see config/multiuser.ts).
|
||||
if (options.multiuser) process.env.CODEMAN_MULTIUSER = '1';
|
||||
const launch = toWebLaunchOptions(options);
|
||||
|
||||
try {
|
||||
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
|
||||
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
|
||||
if (https) {
|
||||
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
|
||||
if (options.stop) {
|
||||
const result = await stopDaemon(launch);
|
||||
if (result.ok && result.reason === 'not-running') {
|
||||
console.log(chalk.gray(`○ ${result.message}`));
|
||||
return;
|
||||
}
|
||||
if (result.ok) {
|
||||
console.log(chalk.green(`✓ ${result.message ?? `Stopped Codeman (pid ${result.pid})`}`));
|
||||
console.log(chalk.gray(' Your agents keep running in tmux.'));
|
||||
return;
|
||||
}
|
||||
console.error(chalk.red(`✗ ${result.message ?? 'Could not stop the server'}`));
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
if (options.status) {
|
||||
const status = await daemonStatus(launch);
|
||||
if (status.responding) {
|
||||
const version = status.version ? ` (v${status.version})` : '';
|
||||
console.log(chalk.green(`✓ Responding at ${status.url}${version}`));
|
||||
} else {
|
||||
console.log(chalk.yellow(`○ Nothing answering at ${status.url}`));
|
||||
}
|
||||
console.log(` Daemon pid: ${status.running ? chalk.green(String(status.pid)) : chalk.gray('not running')}`);
|
||||
console.log(chalk.gray(` Pidfile: ${status.pidFile}`));
|
||||
console.log(chalk.gray(` Log: ${status.logPath}`));
|
||||
if (!status.running && status.responding) {
|
||||
console.log(chalk.gray(' (running, but not started with --daemon: probably a service or a foreground run)'));
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
if (options.daemon) {
|
||||
warnIfUnauthenticatedNetwork(launch);
|
||||
console.log(chalk.cyan('Starting Codeman in the background...'));
|
||||
const result = await startDaemon(launch);
|
||||
if (result.ok) {
|
||||
console.log(chalk.green(`\n✓ Codeman is running at ${result.url} (pid ${result.pid})`));
|
||||
console.log(chalk.gray(` Logs: ${result.logPath}`));
|
||||
console.log(chalk.gray(' Stop it with: codeman web --stop'));
|
||||
console.log(chalk.gray(' Want it back after a reboot? codeman service install'));
|
||||
return;
|
||||
}
|
||||
console.error(chalk.red(`\n✗ ${result.message ?? 'Failed to start'}`));
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
const { startWebServer } = await import('./web/server.js');
|
||||
const host = launch.host;
|
||||
const port = launch.port;
|
||||
const https = launch.https;
|
||||
const titleHostname = options.titleHostname;
|
||||
const allowUnauthenticatedNetwork = launch.allowUnauthenticatedNetwork ?? false;
|
||||
const protocol = https ? 'https' : 'http';
|
||||
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
|
||||
|
||||
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
|
||||
|
||||
try {
|
||||
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
|
||||
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
|
||||
if (https) {
|
||||
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
|
||||
}
|
||||
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
|
||||
|
||||
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
|
||||
let shuttingDown = false;
|
||||
const shutdown = async (signal: string) => {
|
||||
if (shuttingDown) return;
|
||||
shuttingDown = true;
|
||||
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
|
||||
try {
|
||||
await server.stop();
|
||||
} catch (err) {
|
||||
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
|
||||
}
|
||||
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
|
||||
process.exit(0);
|
||||
};
|
||||
process.on('SIGTERM', () => shutdown('SIGTERM'));
|
||||
process.on('SIGINT', () => shutdown('SIGINT'));
|
||||
process.on('SIGHUP', () => shutdown('SIGHUP'));
|
||||
} catch (err) {
|
||||
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
|
||||
process.exit(1);
|
||||
}
|
||||
});
|
||||
|
||||
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
|
||||
let shuttingDown = false;
|
||||
const shutdown = async (signal: string) => {
|
||||
if (shuttingDown) return;
|
||||
shuttingDown = true;
|
||||
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
|
||||
try {
|
||||
await server.stop();
|
||||
} catch (err) {
|
||||
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
|
||||
}
|
||||
process.exit(0);
|
||||
};
|
||||
process.on('SIGTERM', () => shutdown('SIGTERM'));
|
||||
process.on('SIGINT', () => shutdown('SIGINT'));
|
||||
process.on('SIGHUP', () => shutdown('SIGHUP'));
|
||||
} catch (err) {
|
||||
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
|
||||
// Supervised service: the "still there after a reboot" answer, where `web -d` is
|
||||
// the "still there after I close this shell" one (issue #231).
|
||||
const serviceCmd = program
|
||||
.command('service')
|
||||
.description('Manage the background service (systemd user unit on Linux, LaunchAgent on macOS)');
|
||||
|
||||
addWebLaunchOptions(
|
||||
serviceCmd.command('install').description('Install and start the service, then verify it answers')
|
||||
).action(async (options) => {
|
||||
const launch = toWebLaunchOptions(options);
|
||||
warnIfUnauthenticatedNetwork(launch);
|
||||
console.log(chalk.cyan('Installing the Codeman service...'));
|
||||
|
||||
const result = await installService(launch);
|
||||
for (const warning of result.warnings ?? []) console.log(chalk.yellow(`⚠ ${warning}`));
|
||||
|
||||
if (!result.ok) {
|
||||
console.error(chalk.red(`✗ ${result.message}`));
|
||||
process.exit(1);
|
||||
}
|
||||
console.log(chalk.green(`✓ ${result.message}`));
|
||||
console.log(chalk.gray(` Unit: ${result.unitPath}`));
|
||||
if (process.env.CODEMAN_PASSWORD) {
|
||||
console.log(
|
||||
chalk.yellow(
|
||||
' Note: CODEMAN_PASSWORD was NOT copied into the unit file. Add it there yourself if the service needs auth.'
|
||||
)
|
||||
);
|
||||
}
|
||||
});
|
||||
|
||||
serviceCmd
|
||||
.command('uninstall')
|
||||
.description('Stop the service and remove its unit file')
|
||||
.action(() => {
|
||||
const result = uninstallService();
|
||||
if (!result.ok) {
|
||||
console.error(chalk.red(`✗ ${result.message}`));
|
||||
process.exit(1);
|
||||
}
|
||||
console.log(chalk.green(`✓ ${result.message}`));
|
||||
});
|
||||
|
||||
addWebLaunchOptions(
|
||||
serviceCmd.command('status').description('Show whether the service is installed and running')
|
||||
).action(async (options) => {
|
||||
const status = await serviceStatus(toWebLaunchOptions(options));
|
||||
if (!status.kind) {
|
||||
console.log(chalk.yellow(`No supported supervisor on ${process.platform}. Use \`codeman web -d\` instead.`));
|
||||
return;
|
||||
}
|
||||
console.log(` Supervisor: ${status.kind} (${status.name})`);
|
||||
console.log(` Unit file: ${status.installed ? chalk.green(status.unitPath) : chalk.gray('not installed')}`);
|
||||
console.log(` Loaded: ${status.loaded ? chalk.green('yes') : chalk.gray('no')}`);
|
||||
const version = status.version ? ` (v${status.version})` : '';
|
||||
console.log(
|
||||
` Responding: ${status.responding ? chalk.green(`yes at ${status.url}${version}`) : chalk.gray(`no at ${status.url}`)}`
|
||||
);
|
||||
});
|
||||
|
||||
// ============ Multi-user Commands ============
|
||||
//
|
||||
// Operate directly on ~/.codeman/users.json (via user-store) with NO running
|
||||
|
||||
@@ -0,0 +1,112 @@
|
||||
/**
|
||||
* @fileoverview Bounds for the agent wait primitives.
|
||||
*
|
||||
* These back the blocking endpoints an agent uses to orchestrate other sessions
|
||||
* (`GET /api/sessions/:id/wait`, `GET /api/sessions/:id/wait-output`, and the
|
||||
* `wait` field on `POST /api/sessions/:id/input`). Plan: `docs/agent-control-plan.md`.
|
||||
*
|
||||
* Why every value is bounded:
|
||||
* - An unbounded long-poll is a socket leak. A caller that asks for a 12-hour wait
|
||||
* and walks away holds a connection (and a waiter, and a timer) until the process
|
||||
* restarts, so `MAX_WAIT_MS` is a hard ceiling applied server-side.
|
||||
* - `DEFAULT_WAIT_MS` is deliberately short (60s). Production is reached through
|
||||
* `tailscale serve` and users also run cloudflared tunnels; both can cut an idle
|
||||
* connection, so the documented pattern is a client-side loop over short waits
|
||||
* rather than one very long call. Fastify itself is happy to hold the request
|
||||
* (`requestTimeout` defaults to 0, and `keepAliveTimeout` applies between
|
||||
* requests, not to an in-flight one), the intermediaries are the constraint.
|
||||
* - The waiter caps mirror `MAX_SSE_CLIENTS` in `map-limits.ts`: each pending
|
||||
* waiter costs an open HTTP response plus a timer, so the pool is capped rather
|
||||
* than queued. Exceeding a cap is an explicit error, never a silent wait.
|
||||
* - There are THREE caps, not two, because a process-wide pool with no per-user
|
||||
* dimension lets one user deny the primitive to everyone else. `middleware/auth.ts`
|
||||
* already treats that shape as a bug (its `userFailures` bucket exists so "one user
|
||||
* behind a NAT can't lock out everyone else"); `MAX_WAITERS_PER_OWNER` is the same
|
||||
* idea for waiters. It applies only when the caller has an owner, so single-user
|
||||
* mode is byte-identical to having no owner cap at all.
|
||||
*
|
||||
* All values are env-overridable and clamped to sane hard bounds, so a typo in an
|
||||
* env var degrades to the default instead of disabling the protection.
|
||||
*
|
||||
* @module config/agent-wait
|
||||
*/
|
||||
|
||||
/** Absolute floor for any wait, in ms. Sub-second waits are polling, not waiting. */
|
||||
export const MIN_WAIT_MS = 1_000;
|
||||
|
||||
/** Ceiling the operator-configurable maximum is itself clamped to. */
|
||||
const HARD_MAX_WAIT_MS = 3_600_000;
|
||||
|
||||
function envInt(name: string, fallback: number, min: number, max: number): number {
|
||||
const raw = parseInt(process.env[name] || '', 10);
|
||||
if (!Number.isFinite(raw) || raw <= 0) return fallback;
|
||||
return Math.max(min, Math.min(max, raw));
|
||||
}
|
||||
|
||||
/** Longest a single wait may block. Requests above this are clamped down, not rejected. */
|
||||
export const MAX_WAIT_MS = envInt('CODEMAN_WAIT_MAX_MS', 600_000, MIN_WAIT_MS, HARD_MAX_WAIT_MS);
|
||||
|
||||
/** Used when the caller omits `timeout`. Never exceeds MAX_WAIT_MS. */
|
||||
export const DEFAULT_WAIT_MS = Math.min(
|
||||
envInt('CODEMAN_WAIT_DEFAULT_MS', 60_000, MIN_WAIT_MS, HARD_MAX_WAIT_MS),
|
||||
MAX_WAIT_MS
|
||||
);
|
||||
|
||||
/** Concurrent waiters (signal + output) allowed against one session. */
|
||||
export const MAX_WAITERS_PER_SESSION = envInt('CODEMAN_WAIT_MAX_PER_SESSION', 16, 1, 256);
|
||||
|
||||
/**
|
||||
* Ceiling the operator-configurable total is itself clamped to.
|
||||
*
|
||||
* 512 rather than the 4096 this started at. Every other knob in this file degrades
|
||||
* safely on a bad value; a 4096 ceiling instead lets a well-meaning operator turn the
|
||||
* protection into the problem, since 4096 concurrent held responses (each an open
|
||||
* socket, a timer and a pending promise) exceeds the 1024 soft `RLIMIT_NOFILE` that is
|
||||
* still the default on most Linux distros, before counting PTYs, SSE clients and
|
||||
* WebSockets. 512 is ~5x `MAX_SSE_CLIENTS` (100, the pool this one is modelled on), so
|
||||
* the knob stays useful for a busy orchestration host while the whole server still fits
|
||||
* inside a default fd budget with room to spare.
|
||||
*/
|
||||
const HARD_MAX_WAITERS_TOTAL = 512;
|
||||
|
||||
/** Concurrent waiters allowed across every session in the process. */
|
||||
export const MAX_WAITERS_TOTAL = envInt('CODEMAN_WAIT_MAX_TOTAL', 128, 1, HARD_MAX_WAITERS_TOTAL);
|
||||
|
||||
/**
|
||||
* Concurrent waiters allowed for one owner (multi-user mode's `Session.owner`).
|
||||
*
|
||||
* Sits between the per-session cap (16) and the process-wide one (128): high enough
|
||||
* that one user orchestrating several workers at once never trips it, low enough that
|
||||
* a single user cannot occupy the whole pool and deny the primitive to everyone else,
|
||||
* admin included. Ignored entirely when the caller has no owner, which is every
|
||||
* request in single-user mode.
|
||||
*/
|
||||
export const MAX_WAITERS_PER_OWNER = envInt('CODEMAN_WAIT_MAX_PER_OWNER', 48, 1, HARD_MAX_WAITERS_TOTAL);
|
||||
|
||||
/** Bounds on the literal `match` string accepted by wait-output. */
|
||||
export const MIN_MATCH_LENGTH = 1;
|
||||
export const MAX_MATCH_LENGTH = 200;
|
||||
|
||||
/**
|
||||
* Tail of the terminal buffer scanned by `wait-output?from=buffer`.
|
||||
*
|
||||
* The buffer itself runs to 32MB. Scanning all of it would be an ANSI strip over
|
||||
* 32MB (a full second copy) on a request an agent may issue in a loop, and the
|
||||
* question `from=buffer` answers is "did this appear recently", not "ever". The
|
||||
* tail is continuous with the live stream, since `_terminalBuffer.append(data)`
|
||||
* and `emit('terminal', data)` receive the same bytes.
|
||||
*/
|
||||
export const MAX_BUFFER_SCAN_BYTES = envInt('CODEMAN_WAIT_BUFFER_SCAN_BYTES', 256 * 1024, 4 * 1024, 8 * 1024 * 1024);
|
||||
|
||||
/** Characters of surrounding output returned either side of a wait-output match. */
|
||||
export const MAX_SNIPPET_CONTEXT = 80;
|
||||
|
||||
/**
|
||||
* Clamp a caller-supplied timeout into [MIN_WAIT_MS, MAX_WAIT_MS].
|
||||
* Absent / non-numeric / non-finite input falls back to DEFAULT_WAIT_MS.
|
||||
*/
|
||||
export function clampWaitMs(value: unknown): number {
|
||||
const n = typeof value === 'string' ? Number(value) : value;
|
||||
if (typeof n !== 'number' || !Number.isFinite(n)) return DEFAULT_WAIT_MS;
|
||||
return Math.max(MIN_WAIT_MS, Math.min(MAX_WAIT_MS, Math.trunc(n)));
|
||||
}
|
||||
@@ -0,0 +1,32 @@
|
||||
/**
|
||||
* @fileoverview Supervisor identity (systemd unit name / launchd job label).
|
||||
*
|
||||
* Three things now write or look for the same supervisor job: `install.sh`, the
|
||||
* in-app self-updater (`web/self-update.ts` detects it to decide how to restart),
|
||||
* and `codeman service install`. The names live here so they cannot drift apart,
|
||||
* because a mismatch is silent in the worst way: `service install` would happily
|
||||
* create a SECOND job alongside the installer's, and two servers sharing one data
|
||||
* dir and one tmux socket attach PTYs to each other's live sessions
|
||||
* (see config/instance.ts).
|
||||
*
|
||||
* The names are instance-scoped for exactly that reason: a `CODEMAN_INSTANCE=beta`
|
||||
* build writing `com.codeman.web` would overwrite the production LaunchAgent. The
|
||||
* DEFAULT instance keeps the historical names byte-identical, so existing installs
|
||||
* and every unit install.sh has already written are unaffected.
|
||||
*
|
||||
* @module config/service-names
|
||||
*/
|
||||
|
||||
import { CODEMAN_INSTANCE } from './instance.js';
|
||||
|
||||
/**
|
||||
* Instance name reduced to characters that are safe in a filename and in a
|
||||
* launchd label. `CODEMAN_INSTANCE` is arbitrary operator input.
|
||||
*/
|
||||
const SAFE_INSTANCE = CODEMAN_INSTANCE.replace(/[^A-Za-z0-9_-]/g, '').slice(0, 32);
|
||||
|
||||
/** systemd user unit: `codeman-web.service`, or `codeman-web-beta.service` for a beta. */
|
||||
export const SYSTEMD_UNIT = `codeman-web${SAFE_INSTANCE ? `-${SAFE_INSTANCE}` : ''}.service`;
|
||||
|
||||
/** launchd job label: `com.codeman.web`, or `com.codeman.beta.web` for a beta. */
|
||||
export const LAUNCHD_LABEL = SAFE_INSTANCE ? `com.codeman.${SAFE_INSTANCE}.web` : 'com.codeman.web';
|
||||
@@ -0,0 +1,496 @@
|
||||
/**
|
||||
* @fileoverview Detached `codeman web` control: start (-d), stop, status.
|
||||
*
|
||||
* Backs `codeman web -d`, `codeman web --stop` and `codeman web --status`. The
|
||||
* server itself is unchanged; this module re-launches the SAME entry script in a
|
||||
* new session (`detached: true` calls setsid), so the child has no controlling
|
||||
* terminal and no shell job entry. That is what actually makes it outlive the
|
||||
* shell: `nohup` does not, because Node re-arms SIGHUP to its default disposition
|
||||
* even when it inherits "ignore", and `cli.ts` installs a SIGHUP handler that
|
||||
* shuts the server down gracefully (issue #231).
|
||||
*
|
||||
* Two rules shape the rest of the module:
|
||||
*
|
||||
* 1. **Never start a second server on one data dir.** `~/.codeman` and the
|
||||
* `tmux -L codeman` socket are process-wide (config/instance.ts), so a second
|
||||
* instance discovers and attaches PTYs to the first one's live sessions and
|
||||
* starts resizing them. A double `-d` therefore has to be a hard error, which
|
||||
* means checking both the pidfile AND the port before spawning.
|
||||
* 2. **Never report success we have not seen.** The parent polls `/api/status`
|
||||
* until the child answers (or dies) before printing a URL. A port clash or a
|
||||
* missing dependency otherwise looks exactly like a clean start.
|
||||
*
|
||||
* Pure helpers (arg building, URL building, pidfile parsing, the process-identity
|
||||
* check) are exported separately so they can be unit-tested without spawning.
|
||||
*
|
||||
* @module daemon-control
|
||||
*/
|
||||
|
||||
import { spawn, execFileSync } from 'node:child_process';
|
||||
import { appendFileSync, closeSync, existsSync, openSync, readFileSync, unlinkSync, writeFileSync } from 'node:fs';
|
||||
import http from 'node:http';
|
||||
import https from 'node:https';
|
||||
import { dataPath } from './config/instance.js';
|
||||
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
|
||||
|
||||
/** How long to wait for a freshly spawned server to answer `/api/status`. */
|
||||
const START_TIMEOUT_MS = 30_000;
|
||||
/** How long to wait for a SIGTERM'd server to actually exit before giving up. */
|
||||
const STOP_TIMEOUT_MS = 15_000;
|
||||
/** Poll interval while waiting for either of the above. */
|
||||
const POLL_INTERVAL_MS = 250;
|
||||
|
||||
/** The `web` command's options, as far as a detached relaunch cares about them. */
|
||||
export interface WebLaunchOptions {
|
||||
host: string;
|
||||
port: number;
|
||||
https: boolean;
|
||||
titleHostname?: string;
|
||||
allowUnauthenticatedNetwork?: boolean;
|
||||
multiuser?: boolean;
|
||||
}
|
||||
|
||||
export interface StartResult {
|
||||
ok: boolean;
|
||||
pid?: number;
|
||||
url?: string;
|
||||
/** Machine-readable failure cause; `undefined` on success. */
|
||||
reason?: 'already-running' | 'exited' | 'timeout';
|
||||
message?: string;
|
||||
logPath: string;
|
||||
}
|
||||
|
||||
export interface StopResult {
|
||||
ok: boolean;
|
||||
pid?: number;
|
||||
reason?: 'not-running' | 'foreign-pid' | 'timeout' | 'no-pidfile-but-responding';
|
||||
message?: string;
|
||||
}
|
||||
|
||||
export interface DaemonStatus {
|
||||
pid: number | null;
|
||||
/** The pid in the pidfile is alive AND still looks like a Codeman web process. */
|
||||
running: boolean;
|
||||
/** Something answered `/api/status` at the expected address. */
|
||||
responding: boolean;
|
||||
version?: string;
|
||||
url: string;
|
||||
pidFile: string;
|
||||
logPath: string;
|
||||
}
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
// Pure helpers
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
/** Rebuild the `web` argv for the child, dropping the daemon flags themselves. */
|
||||
export function buildWebArgs(options: WebLaunchOptions): string[] {
|
||||
const args = ['web', '--host', options.host, '--port', String(options.port)];
|
||||
if (options.https) args.push('--https');
|
||||
if (options.titleHostname) args.push('--title-hostname', options.titleHostname);
|
||||
if (options.allowUnauthenticatedNetwork) args.push('--allow-unauthenticated-network');
|
||||
if (options.multiuser) args.push('--multiuser');
|
||||
return args;
|
||||
}
|
||||
|
||||
/**
|
||||
* Connectable address for this bind. A wildcard bind is not itself connectable,
|
||||
* so `0.0.0.0` / `::` become loopback; a bare IPv6 literal gets bracketed.
|
||||
*/
|
||||
export function buildBaseUrl(options: WebLaunchOptions): string {
|
||||
const protocol = options.https ? 'https' : 'http';
|
||||
let host = options.host.trim();
|
||||
if (host === '0.0.0.0' || host === '::' || host === '') host = '127.0.0.1';
|
||||
if (host.includes(':') && !host.startsWith('[')) host = `[${host}]`;
|
||||
return `${protocol}://${host}:${options.port}`;
|
||||
}
|
||||
|
||||
/** The endpoint polled for readiness. */
|
||||
export function buildStatusUrl(options: WebLaunchOptions): string {
|
||||
return `${buildBaseUrl(options)}/api/status`;
|
||||
}
|
||||
|
||||
/** Parse a pidfile body. Rejects garbage, and pid 1 (init is never ours). */
|
||||
export function parsePidFileContents(text: string): number | null {
|
||||
const trimmed = text.trim();
|
||||
if (!/^\d+$/.test(trimmed)) return null;
|
||||
const pid = Number.parseInt(trimmed, 10);
|
||||
if (!Number.isSafeInteger(pid) || pid <= 1) return null;
|
||||
return pid;
|
||||
}
|
||||
|
||||
/**
|
||||
* Does this command line look like a Codeman web server?
|
||||
*
|
||||
* Pids are recycled, and a stale pidfile pointing at whatever inherited the
|
||||
* number is a live footgun: `codeman web --stop` must not SIGTERM an unrelated
|
||||
* process. Both the npm bin (`codeman`/`aicodeman`) and the direct entry
|
||||
* (`node dist/index.js web`, `tsx src/index.ts web`) have to match.
|
||||
*/
|
||||
export function looksLikeCodemanWeb(command: string | null | undefined): boolean {
|
||||
if (!command) return false;
|
||||
if (!/(^|\s)web(\s|$)/.test(command)) return false;
|
||||
return /(^|[/\s])(ai)?codeman(\s|$)/.test(command) || /index\.(js|ts)(\s|$)/.test(command);
|
||||
}
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
// Paths
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Resolved at call time, not module load: tests swap `HOME` per file, and the
|
||||
* data dir is derived from it (see test/setup.ts).
|
||||
*/
|
||||
export function pidFilePath(): string {
|
||||
return dataPath('web.pid');
|
||||
}
|
||||
|
||||
/** Where a detached server's stdout/stderr is appended. */
|
||||
export function logFilePath(): string {
|
||||
return dataPath('web.log');
|
||||
}
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
// Process probing
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
/** Signal 0 liveness check. EPERM means the pid exists but is not ours. */
|
||||
export function isProcessAlive(pid: number): boolean {
|
||||
try {
|
||||
process.kill(pid, 0);
|
||||
return true;
|
||||
} catch (err) {
|
||||
return (err as NodeJS.ErrnoException).code === 'EPERM';
|
||||
}
|
||||
}
|
||||
|
||||
/** Full command line of a pid, or null. `-o command=` is portable to macOS. */
|
||||
export function readProcessCommand(pid: number): string | null {
|
||||
try {
|
||||
const out = execFileSync('ps', ['-o', 'command=', '-p', String(pid)], {
|
||||
encoding: 'utf-8',
|
||||
timeout: EXEC_TIMEOUT_MS,
|
||||
stdio: ['ignore', 'pipe', 'ignore'],
|
||||
});
|
||||
return out.trim() || null;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/** Read the pidfile, returning null when it is missing, empty or malformed. */
|
||||
export function readPidFile(): number | null {
|
||||
const file = pidFilePath();
|
||||
if (!existsSync(file)) return null;
|
||||
try {
|
||||
return parsePidFileContents(readFileSync(file, 'utf-8'));
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
function removePidFile(): void {
|
||||
try {
|
||||
unlinkSync(pidFilePath());
|
||||
} catch {
|
||||
/* already gone */
|
||||
}
|
||||
}
|
||||
|
||||
/** Pid of a live Codeman web server recorded in the pidfile, or null. */
|
||||
export function readLivePid(): number | null {
|
||||
const pid = readPidFile();
|
||||
if (pid === null) return null;
|
||||
if (!isProcessAlive(pid)) return null;
|
||||
// A recycled pid is not ours. `ps` can also legitimately fail (containers with
|
||||
// no procps); treat "cannot tell" as ours rather than orphaning the pidfile.
|
||||
const command = readProcessCommand(pid);
|
||||
if (command !== null && !looksLikeCodemanWeb(command)) return null;
|
||||
return pid;
|
||||
}
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
// HTTP readiness probe
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
export interface ProbeResult {
|
||||
/** A Codeman server answered. A 401 counts: auth is active, the server is up. */
|
||||
up: boolean;
|
||||
version?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* Probe `/api/status`. Self-signed certs are accepted (`--https` generates one),
|
||||
* and 401 counts as up because `CODEMAN_PASSWORD` gates that route. The body is
|
||||
* checked so an unrelated service squatting on the port is not read as success.
|
||||
*/
|
||||
export function probeServer(url: string, timeoutMs = 2000): Promise<ProbeResult> {
|
||||
return new Promise((resolve) => {
|
||||
let settled = false;
|
||||
const done = (result: ProbeResult) => {
|
||||
if (settled) return;
|
||||
settled = true;
|
||||
resolve(result);
|
||||
};
|
||||
|
||||
let target: URL;
|
||||
try {
|
||||
target = new URL(url);
|
||||
} catch {
|
||||
done({ up: false });
|
||||
return;
|
||||
}
|
||||
|
||||
const transport = target.protocol === 'https:' ? https : http;
|
||||
const req = transport.request(
|
||||
{
|
||||
protocol: target.protocol,
|
||||
hostname: target.hostname,
|
||||
port: target.port,
|
||||
path: target.pathname,
|
||||
method: 'GET',
|
||||
rejectUnauthorized: false,
|
||||
timeout: timeoutMs,
|
||||
headers: { Accept: 'application/json' },
|
||||
},
|
||||
(res) => {
|
||||
if (res.statusCode === 401) {
|
||||
res.resume();
|
||||
done({ up: true });
|
||||
return;
|
||||
}
|
||||
let body = '';
|
||||
res.setEncoding('utf-8');
|
||||
res.on('data', (chunk: string) => {
|
||||
if (body.length < 4096) body += chunk;
|
||||
});
|
||||
res.on('end', () => {
|
||||
if (!body.includes('"success"')) {
|
||||
done({ up: false });
|
||||
return;
|
||||
}
|
||||
let version: string | undefined;
|
||||
try {
|
||||
version = (JSON.parse(body) as { data?: { version?: string } }).data?.version;
|
||||
} catch {
|
||||
/* body was truncated at 4KB; up is still true */
|
||||
}
|
||||
done({ up: true, version });
|
||||
});
|
||||
res.on('error', () => done({ up: false }));
|
||||
}
|
||||
);
|
||||
req.on('timeout', () => {
|
||||
req.destroy();
|
||||
done({ up: false });
|
||||
});
|
||||
req.on('error', () => done({ up: false }));
|
||||
req.end();
|
||||
});
|
||||
}
|
||||
|
||||
function sleep(ms: number): Promise<void> {
|
||||
return new Promise((resolve) => setTimeout(resolve, ms));
|
||||
}
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
// Start / stop / status
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* The script to relaunch. `process.execArgv` is carried over with it so a dev
|
||||
* run under tsx (whose execArgv holds the tsx loader flags) re-launches through
|
||||
* tsx instead of handing a `.ts` file to bare node.
|
||||
*/
|
||||
function entryScript(): string {
|
||||
const script = process.argv[1];
|
||||
if (!script) throw new Error('cannot determine the codeman entry script to relaunch');
|
||||
return script;
|
||||
}
|
||||
|
||||
/** Marks one launch in the append-only log so a tail cannot mix two runs. */
|
||||
const LOG_SEPARATOR = '=== codeman web start';
|
||||
|
||||
/**
|
||||
* Last few lines of the daemon log, for reporting a failed start. The log is
|
||||
* append-only across launches, so the tail starts at the last separator when
|
||||
* there is one: otherwise a crash report is padded with the previous run's
|
||||
* cheerful startup banner.
|
||||
*/
|
||||
export function tailLog(maxLines = 15): string {
|
||||
try {
|
||||
const lines = readFileSync(logFilePath(), 'utf-8').trimEnd().split('\n');
|
||||
const start = lines.map((line) => line.startsWith(LOG_SEPARATOR)).lastIndexOf(true);
|
||||
const current = start === -1 ? lines : lines.slice(start + 1);
|
||||
return current.slice(-maxLines).join('\n');
|
||||
} catch {
|
||||
return '';
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Spawn a detached `codeman web` and wait until it answers before returning.
|
||||
* Refuses when a server is already up on this data dir (see rule 1 in the module
|
||||
* docblock).
|
||||
*/
|
||||
export async function startDaemon(options: WebLaunchOptions): Promise<StartResult> {
|
||||
const logPath = logFilePath();
|
||||
const url = buildBaseUrl(options);
|
||||
const statusUrl = buildStatusUrl(options);
|
||||
|
||||
const existingPid = readLivePid();
|
||||
if (existingPid !== null) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: 'already-running',
|
||||
pid: existingPid,
|
||||
logPath,
|
||||
message: `a Codeman server is already running (pid ${existingPid}). Stop it with \`codeman web --stop\` first.`,
|
||||
};
|
||||
}
|
||||
const alreadyServing = await probeServer(statusUrl, 1500);
|
||||
if (alreadyServing.up) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: 'already-running',
|
||||
logPath,
|
||||
url,
|
||||
message: `something is already serving ${url}. Two servers on one data dir attach to each other's tmux sessions, so refusing to start.`,
|
||||
};
|
||||
}
|
||||
// A pidfile that survived a crash: the process is gone, so it is just litter.
|
||||
if (readPidFile() !== null) removePidFile();
|
||||
|
||||
const args = buildWebArgs(options);
|
||||
try {
|
||||
appendFileSync(logPath, `\n${LOG_SEPARATOR} ${new Date().toISOString()} ===\n`, 'utf-8');
|
||||
} catch {
|
||||
/* the spawn below reports a genuinely unwritable log */
|
||||
}
|
||||
const logFd = openSync(logPath, 'a');
|
||||
let child;
|
||||
try {
|
||||
child = spawn(process.execPath, [...process.execArgv, entryScript(), ...args], {
|
||||
detached: true,
|
||||
stdio: ['ignore', logFd, logFd],
|
||||
env: process.env,
|
||||
});
|
||||
} finally {
|
||||
closeSync(logFd);
|
||||
}
|
||||
|
||||
let exited = false;
|
||||
child.on('exit', () => {
|
||||
exited = true;
|
||||
});
|
||||
child.on('error', () => {
|
||||
exited = true;
|
||||
});
|
||||
|
||||
const pid = child.pid;
|
||||
if (pid === undefined) {
|
||||
return { ok: false, reason: 'exited', logPath, message: 'failed to spawn the server process' };
|
||||
}
|
||||
writeFileSync(pidFilePath(), `${pid}\n`, 'utf-8');
|
||||
|
||||
const deadline = Date.now() + START_TIMEOUT_MS;
|
||||
while (Date.now() < deadline) {
|
||||
if (exited) {
|
||||
removePidFile();
|
||||
child.unref();
|
||||
return {
|
||||
ok: false,
|
||||
reason: 'exited',
|
||||
logPath,
|
||||
message: `the server exited during startup. Last lines of ${logPath}:\n${tailLog()}`,
|
||||
};
|
||||
}
|
||||
const probe = await probeServer(statusUrl, 1000);
|
||||
if (probe.up) {
|
||||
child.unref();
|
||||
return { ok: true, pid, url, logPath };
|
||||
}
|
||||
await sleep(POLL_INTERVAL_MS);
|
||||
}
|
||||
|
||||
child.unref();
|
||||
return {
|
||||
ok: false,
|
||||
reason: 'timeout',
|
||||
pid,
|
||||
url,
|
||||
logPath,
|
||||
message: `the server did not answer ${url} within ${START_TIMEOUT_MS / 1000}s. It may still be starting; check ${logPath}.`,
|
||||
};
|
||||
}
|
||||
|
||||
/** SIGTERM the recorded server and wait for it to actually exit. */
|
||||
export async function stopDaemon(options: WebLaunchOptions): Promise<StopResult> {
|
||||
const pid = readPidFile();
|
||||
if (pid === null) {
|
||||
const probe = await probeServer(buildStatusUrl(options), 1500);
|
||||
if (probe.up) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: 'no-pidfile-but-responding',
|
||||
message:
|
||||
'a server is responding but there is no pidfile, so it was not started with `-d`. If it is a service use `codeman service uninstall` (or stop the unit); otherwise `pkill -f "index.js web"`.',
|
||||
};
|
||||
}
|
||||
return { ok: true, reason: 'not-running', message: 'no daemon is running; nothing to stop' };
|
||||
}
|
||||
|
||||
if (!isProcessAlive(pid)) {
|
||||
removePidFile();
|
||||
return { ok: true, pid, message: `stale pidfile removed (pid ${pid} was not running)` };
|
||||
}
|
||||
|
||||
const command = readProcessCommand(pid);
|
||||
if (command !== null && !looksLikeCodemanWeb(command)) {
|
||||
return {
|
||||
ok: false,
|
||||
reason: 'foreign-pid',
|
||||
pid,
|
||||
message: `pid ${pid} is not a Codeman server (${command}). Refusing to signal it; delete ${pidFilePath()} if it is stale.`,
|
||||
};
|
||||
}
|
||||
|
||||
// SIGTERM, never SIGKILL: cli.ts flushes state on the way out.
|
||||
try {
|
||||
process.kill(pid, 'SIGTERM');
|
||||
} catch (err) {
|
||||
return { ok: false, reason: 'foreign-pid', pid, message: `could not signal pid ${pid}: ${String(err)}` };
|
||||
}
|
||||
|
||||
const deadline = Date.now() + STOP_TIMEOUT_MS;
|
||||
while (Date.now() < deadline) {
|
||||
if (!isProcessAlive(pid)) {
|
||||
removePidFile();
|
||||
return { ok: true, pid };
|
||||
}
|
||||
await sleep(POLL_INTERVAL_MS);
|
||||
}
|
||||
|
||||
return {
|
||||
ok: false,
|
||||
reason: 'timeout',
|
||||
pid,
|
||||
message: `pid ${pid} did not exit within ${STOP_TIMEOUT_MS / 1000}s. Force it with \`kill -9 ${pid}\` if you are sure.`,
|
||||
};
|
||||
}
|
||||
|
||||
/** Report on both halves: the recorded process, and whether the port answers. */
|
||||
export async function daemonStatus(options: WebLaunchOptions): Promise<DaemonStatus> {
|
||||
const url = buildBaseUrl(options);
|
||||
const pid = readPidFile();
|
||||
const probe = await probeServer(buildStatusUrl(options), 2000);
|
||||
return {
|
||||
pid,
|
||||
running: readLivePid() !== null,
|
||||
responding: probe.up,
|
||||
version: probe.version,
|
||||
url,
|
||||
pidFile: pidFilePath(),
|
||||
logPath: logFilePath(),
|
||||
};
|
||||
}
|
||||
@@ -596,6 +596,9 @@ interface CredStorePolicy {
|
||||
|
||||
const CRED_STORES: CredStorePolicy[] = [
|
||||
{ rel: '.codex', shareDirs: ['sessions'], shareFiles: ['history.jsonl'], seedFiles: ['auth.json', 'config.toml'] },
|
||||
// Also covers Antigravity: `agy` nests its whole state (auth `jetski_state.pbtxt`,
|
||||
// `conversations/`, `knowledge/`) under `~/.gemini/antigravity-cli/`, so it needs no
|
||||
// entry of its own. There is no `~/.antigravity` credential dir to add.
|
||||
{ rel: '.gemini', seedWhole: true },
|
||||
{ rel: '.config/gcloud', seedWhole: true },
|
||||
{ rel: '.config/opencode', seedWhole: true },
|
||||
|
||||
+217
-28
@@ -10,13 +10,15 @@
|
||||
* Key exports:
|
||||
* - `generateHooksConfig()` — returns hooks object for settings.local.json
|
||||
* - `writeHooksConfig(casePath)` — writes hooks + env config to disk
|
||||
* - `ensureCodemanHooks(casePath)` — safely installs/updates hooks for a managed case
|
||||
* - `updateCaseEnvVars(casePath, envVars)` — merges env vars into settings
|
||||
*
|
||||
* Hook events generated: `idle_prompt`, `permission_prompt`, `elicitation_dialog`,
|
||||
* `stop`, `teammate_idle`, `task_completed`
|
||||
*
|
||||
* Hook categories: `Notification` (3 matchers), `Stop` (1), `TeammateIdle` (1),
|
||||
* `TaskCompleted` (1), `PostToolUse` (1 self-contained background Bash rewake)
|
||||
* Hook categories: `Notification` (3 matchers), `Stop` (1), `SubagentStop` (1),
|
||||
* `TeammateIdle` (1), `TaskCompleted` (1), `PostToolUse` (1 self-contained
|
||||
* background Bash rewake)
|
||||
*
|
||||
* @dependencies types (HookEventType), config/auth-config (HOOK_TIMEOUT_SECONDS)
|
||||
* @consumedby web/server (session creation), session-cli-builder (env setup)
|
||||
@@ -52,15 +54,19 @@ const BACKGROUND_WAKE_MARKER_PREFIX = 'CODEMAN_BACKGROUND_REWAKE_V';
|
||||
* changes: `refreshStaleCodemanHooks` treats the absence of the CURRENT marker as
|
||||
* stale, so healed cases pick up the new script on next launch.
|
||||
*/
|
||||
const BACKGROUND_WAKE_MARKER = `${BACKGROUND_WAKE_MARKER_PREFIX}2`;
|
||||
const BACKGROUND_WAKE_MARKER = `${BACKGROUND_WAKE_MARKER_PREFIX}3`;
|
||||
const SUBAGENT_STOP_GUARD_MARKER_PREFIX = 'CODEMAN_SUBAGENT_STOP_GUARD_V';
|
||||
const SUBAGENT_STOP_GUARD_MARKER = `${SUBAGENT_STOP_GUARD_MARKER_PREFIX}1`;
|
||||
const BACKGROUND_WAKE_TIMEOUT_SECONDS = 6 * 60 * 60;
|
||||
|
||||
/**
|
||||
* Inline Node helper for Claude Code's `asyncRewake` hook.
|
||||
*
|
||||
* A background Bash tool returns immediately with a task ID, then Claude writes
|
||||
* its completion as a queue-operation in the transcript. Watching that durable
|
||||
* record avoids injecting terminal input (which could submit a user's draft).
|
||||
* its completion as a queue-operation in the top-level transcript. Subagent hooks
|
||||
* receive their own transcript path even though their completion is parent-owned,
|
||||
* so the helper watches both paths. Watching durable records avoids injecting
|
||||
* terminal input (which could submit a user's draft).
|
||||
* The helper is embedded in settings via `node -e`, so it has no script path
|
||||
* that can go stale after an install or plugin-cache cleanup.
|
||||
*
|
||||
@@ -72,8 +78,12 @@ const BACKGROUND_WAKE_TIMEOUT_SECONDS = 6 * 60 * 60;
|
||||
export function generateBackgroundWakeScript(): string {
|
||||
return [
|
||||
"const fs = require('node:fs');",
|
||||
"const path = require('node:path');",
|
||||
`const ${BACKGROUND_WAKE_MARKER} = true;`,
|
||||
`const deadline = Date.now() + ${BACKGROUND_WAKE_TIMEOUT_SECONDS} * 1000;`,
|
||||
"const RESULT_BEGIN = '=== CODEMAN_RESULT_BEGIN ===';",
|
||||
"const RESULT_END = '=== CODEMAN_RESULT_END ===';",
|
||||
'const MAX_RESULT_CHARS = 65536;',
|
||||
'let input = {};',
|
||||
"try { input = JSON.parse(fs.readFileSync(0, 'utf8') || '{}'); } catch { process.exit(0); }",
|
||||
'function findTaskId(value) {',
|
||||
@@ -98,46 +108,164 @@ export function generateBackgroundWakeScript(): string {
|
||||
'const taskId = findTaskId(input.tool_response);',
|
||||
"const transcriptPath = typeof input.transcript_path === 'string' ? input.transcript_path : '';",
|
||||
'if (!taskId || !transcriptPath) process.exit(0);',
|
||||
'let position = 0;',
|
||||
'try { position = Math.max(0, fs.statSync(transcriptPath).size - 262144); } catch { process.exit(0); }',
|
||||
"let carry = '';",
|
||||
'const transcriptPaths = [transcriptPath];',
|
||||
'const sessionDir = path.dirname(path.dirname(transcriptPath));',
|
||||
"if (typeof input.agent_id === 'string' && path.basename(path.dirname(transcriptPath)) === 'subagents' &&",
|
||||
" typeof input.session_id === 'string' && path.basename(sessionDir) === input.session_id) {",
|
||||
" transcriptPaths.push(sessionDir + '.jsonl');",
|
||||
'}',
|
||||
'const transcripts = [...new Set(transcriptPaths)].map((transcript) => {',
|
||||
' let position = 0;',
|
||||
' try { position = Math.max(0, fs.statSync(transcript).size - 262144); } catch {}',
|
||||
" return { path: transcript, position, carry: '' };",
|
||||
'});',
|
||||
'if (!transcripts.some((transcript) => fs.existsSync(transcript.path))) process.exit(0);',
|
||||
'function readMarkedResult(outputPath) {',
|
||||
" if (!outputPath || !path.isAbsolute(outputPath) || path.basename(outputPath) !== taskId + '.output') return '';",
|
||||
" if (path.basename(path.dirname(outputPath)) !== 'tasks') return '';",
|
||||
' try {',
|
||||
' const size = fs.statSync(outputPath).size;',
|
||||
' const length = Math.min(size, MAX_RESULT_CHARS * 2);',
|
||||
' const buffer = Buffer.allocUnsafe(length);',
|
||||
" const fd = fs.openSync(outputPath, 'r');",
|
||||
' const bytes = fs.readSync(fd, buffer, 0, length, size - length);',
|
||||
' fs.closeSync(fd);',
|
||||
" const text = buffer.subarray(0, bytes).toString('utf8');",
|
||||
' const begin = text.lastIndexOf(RESULT_BEGIN);',
|
||||
' const end = text.indexOf(RESULT_END, begin + RESULT_BEGIN.length);',
|
||||
" if (begin < 0 || end < 0) return '';",
|
||||
' let result = text.slice(begin + RESULT_BEGIN.length, end).trim();',
|
||||
" if (!result) return '';",
|
||||
' if (result.length > MAX_RESULT_CHARS) {',
|
||||
' const half = Math.floor(MAX_RESULT_CHARS / 2);',
|
||||
" result = result.slice(0, half) + '\\n\\n[report truncated by Codeman]\\n\\n' + result.slice(-half);",
|
||||
' }',
|
||||
" return '\\n\\nCompleted task report:\\n<codeman-background-result>\\n' + result + '\\n</codeman-background-result>';",
|
||||
" } catch { return ''; }",
|
||||
'}',
|
||||
'function inspect(text) {',
|
||||
' for (const line of text.split(/\\r?\\n/)) {',
|
||||
' if (!line.includes(taskId)) continue;',
|
||||
' let entry;',
|
||||
' try { entry = JSON.parse(line); } catch { continue; }',
|
||||
" if (entry.type !== 'queue-operation' || typeof entry.content !== 'string') continue;",
|
||||
" if (entry.type !== 'queue-operation' || entry.operation !== 'enqueue' || typeof entry.content !== 'string') continue;",
|
||||
" if (!entry.content.includes('<task-id>' + taskId + '</task-id>')) continue;",
|
||||
' const status = entry.content.match(/<status>(completed|failed|killed|error)<\\/status>/i);',
|
||||
' if (!status) continue;',
|
||||
' const output = entry.content.match(/<output-file>([^<]+)<\\/output-file>/i);',
|
||||
" const location = output ? ' Read ' + output[1] + ' and' : '';",
|
||||
" console.error('Background command ' + taskId + ' ' + status[1].toLowerCase() + '.' + location + ' continue the task.');",
|
||||
" const outputPath = output ? output[1].trim() : '';",
|
||||
" const location = outputPath ? ' Read ' + outputPath + ' and' : '';",
|
||||
' const result = readMarkedResult(outputPath);',
|
||||
" console.error('Background command ' + taskId + ' ' + status[1].toLowerCase() + '.' + location + ' continue the task.' + result);",
|
||||
' process.exit(2);',
|
||||
' }',
|
||||
'}',
|
||||
'function poll() {',
|
||||
' if (Date.now() > deadline || process.ppid === 1) process.exit(0);',
|
||||
'function pollTranscript(transcript) {',
|
||||
' try {',
|
||||
' const size = fs.statSync(transcriptPath).size;',
|
||||
" if (size < position) { position = 0; carry = ''; }",
|
||||
' if (size > position) {',
|
||||
' const length = Math.min(size - position, 1048576);',
|
||||
' const size = fs.statSync(transcript.path).size;',
|
||||
" if (size < transcript.position) { transcript.position = 0; transcript.carry = ''; }",
|
||||
' if (size > transcript.position) {',
|
||||
' const length = Math.min(size - transcript.position, 1048576);',
|
||||
' const buffer = Buffer.allocUnsafe(length);',
|
||||
" const fd = fs.openSync(transcriptPath, 'r');",
|
||||
' const bytes = fs.readSync(fd, buffer, 0, length, position);',
|
||||
" const fd = fs.openSync(transcript.path, 'r');",
|
||||
' const bytes = fs.readSync(fd, buffer, 0, length, transcript.position);',
|
||||
' fs.closeSync(fd);',
|
||||
' position += bytes;',
|
||||
" carry = (carry + buffer.subarray(0, bytes).toString('utf8')).slice(-262144);",
|
||||
' inspect(carry);',
|
||||
' transcript.position += bytes;',
|
||||
" transcript.carry = (transcript.carry + buffer.subarray(0, bytes).toString('utf8')).slice(-262144);",
|
||||
' inspect(transcript.carry);',
|
||||
' }',
|
||||
' } catch {}',
|
||||
'}',
|
||||
'function poll() {',
|
||||
' if (Date.now() > deadline || process.ppid === 1) process.exit(0);',
|
||||
' for (const transcript of transcripts) pollTranscript(transcript);',
|
||||
' setTimeout(poll, 1000);',
|
||||
'}',
|
||||
'poll();',
|
||||
].join('\n');
|
||||
}
|
||||
|
||||
/**
|
||||
* Keep a Claude subagent alive while its Monitor or background Bash work is live.
|
||||
* Claude otherwise can publish the worker's last progress sentence as an Agent
|
||||
* result when one watcher ends, even if other tracked tasks are still running.
|
||||
*/
|
||||
export function generateSubagentStopGuardScript(): string {
|
||||
return [
|
||||
"const fs = require('node:fs');",
|
||||
`const ${SUBAGENT_STOP_GUARD_MARKER} = true;`,
|
||||
'let input = {};',
|
||||
"try { input = JSON.parse(fs.readFileSync(0, 'utf8') || '{}'); } catch { process.exit(0); }",
|
||||
"const transcriptPath = typeof input.agent_transcript_path === 'string' ? input.agent_transcript_path : '';",
|
||||
'if (!transcriptPath) process.exit(0);',
|
||||
'let text;',
|
||||
'try {',
|
||||
' const size = fs.statSync(transcriptPath).size;',
|
||||
' const length = Math.min(size, 16 * 1024 * 1024);',
|
||||
' const buffer = Buffer.allocUnsafe(length);',
|
||||
" const fd = fs.openSync(transcriptPath, 'r');",
|
||||
' const bytes = fs.readSync(fd, buffer, 0, length, size - length);',
|
||||
' fs.closeSync(fd);',
|
||||
" text = buffer.subarray(0, bytes).toString('utf8');",
|
||||
'} catch { process.exit(0); }',
|
||||
'const launched = new Set();',
|
||||
'const finished = new Set();',
|
||||
'function inspectToolResult(value) {',
|
||||
" const serialized = typeof value === 'string' ? value : JSON.stringify(value ?? '');",
|
||||
' for (const match of serialized.matchAll(/Command running in background with ID:\\s*([A-Za-z0-9_-]+)/gi)) launched.add(match[1]);',
|
||||
' for (const match of serialized.matchAll(/Monitor started \\(task ([A-Za-z0-9_-]+)/gi)) launched.add(match[1]);',
|
||||
'}',
|
||||
'function inspectNotifications(value) {',
|
||||
" if (typeof value !== 'string' || !value.includes('<task-notification>')) return;",
|
||||
' for (const match of value.matchAll(/<task-notification>([\\s\\S]*?)<\\/task-notification>/gi)) {',
|
||||
' const body = match[1];',
|
||||
' const id = body.match(/<task-id>([^<]+)<\\/task-id>/i);',
|
||||
' const status = body.match(/<status>(completed|failed|killed|error)<\\/status>/i);',
|
||||
' if (id && status) finished.add(id[1].trim());',
|
||||
' }',
|
||||
'}',
|
||||
'for (const line of text.split(/\\r?\\n/)) {',
|
||||
' let entry;',
|
||||
' try { entry = JSON.parse(line); } catch { continue; }',
|
||||
' const content = entry && entry.message ? entry.message.content : undefined;',
|
||||
' if (Array.isArray(content)) {',
|
||||
' for (const block of content) {',
|
||||
" if (block && block.type === 'tool_result') inspectToolResult(block.content);",
|
||||
" if (block && block.type === 'text') inspectNotifications(block.text);",
|
||||
' }',
|
||||
' } else {',
|
||||
' inspectNotifications(content);',
|
||||
' }',
|
||||
' inspectNotifications(entry && entry.content);',
|
||||
'}',
|
||||
'function findLiveTasks(candidates) {',
|
||||
' const live = new Set();',
|
||||
" if (candidates.size === 0 || !fs.existsSync('/proc')) return live;",
|
||||
' let processIds;',
|
||||
" try { processIds = fs.readdirSync('/proc').filter((name) => /^\\d+$/.test(name)); } catch { return live; }",
|
||||
' for (const processId of processIds) {',
|
||||
" for (const descriptor of ['0', '1', '2']) {",
|
||||
' let target;',
|
||||
" try { target = fs.readlinkSync('/proc/' + processId + '/fd/' + descriptor); } catch { continue; }",
|
||||
' const match = target.match(/[\\/]tasks[\\/]([A-Za-z0-9_-]+)\\.output(?: \\(deleted\\))?$/);',
|
||||
' if (match && candidates.has(match[1])) live.add(match[1]);',
|
||||
' }',
|
||||
' if (live.size === candidates.size) break;',
|
||||
' }',
|
||||
' return live;',
|
||||
'}',
|
||||
'const unfinished = new Set([...launched].filter((taskId) => !finished.has(taskId)));',
|
||||
'const active = [...findLiveTasks(unfinished)];',
|
||||
'if (active.length === 0) process.exit(0);',
|
||||
'const shown = active.slice(0, 8);',
|
||||
"const suffix = active.length > shown.length ? ' and ' + (active.length - shown.length) + ' more' : '';",
|
||||
'process.stdout.write(JSON.stringify({',
|
||||
" decision: 'block',",
|
||||
" reason: 'You still own active background work (' + shown.join(', ') + suffix + '). Do not return an intermediate progress message as your final report. Process the task notifications or keep actively polling until every task completes, then return one complete summary.',",
|
||||
'}));',
|
||||
].join('\n');
|
||||
}
|
||||
|
||||
function withSettingsLock<T>(path: string, fn: () => Promise<T>): Promise<T> {
|
||||
const prev = settingsWriteLocks.get(path) ?? Promise.resolve();
|
||||
const run = prev.then(fn, fn); // run after the prior writer, regardless of its outcome
|
||||
@@ -173,7 +301,11 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
|
||||
const curlCmd = (event: HookEventType) =>
|
||||
`HOOK_DATA=$(cat 2>/dev/null || echo '{}'); ` +
|
||||
`printf '{"event":"${event}","sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$HOOK_DATA" | ` +
|
||||
`curl -s -X POST "$CODEMAN_API_URL/api/hook-event" ` +
|
||||
// `-k`, same as the statusline exporter: CODEMAN_API_URL is loopback HTTPS with
|
||||
// a self-signed cert on --https/tailscale installs. Without it curl exits 60,
|
||||
// the `|| true` swallows it, and ALL SIX hook events die silently: respawn loses
|
||||
// its definitive idle signals and the wait endpoints lose stop/blocked.
|
||||
`curl -sk -X POST "$CODEMAN_API_URL/api/hook-event" ` +
|
||||
`-H 'Content-Type: application/json' ` +
|
||||
`-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" ` +
|
||||
`--data @- ` +
|
||||
@@ -200,6 +332,18 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
|
||||
hooks: [{ type: 'command', command: curlCmd('stop'), timeout: HOOK_TIMEOUT_SECONDS }],
|
||||
},
|
||||
],
|
||||
SubagentStop: [
|
||||
{
|
||||
hooks: [
|
||||
{
|
||||
type: 'command',
|
||||
command: 'node',
|
||||
args: ['-e', generateSubagentStopGuardScript()],
|
||||
timeout: HOOK_TIMEOUT_SECONDS,
|
||||
},
|
||||
],
|
||||
},
|
||||
],
|
||||
TeammateIdle: [
|
||||
{
|
||||
hooks: [{ type: 'command', command: curlCmd('teammate_idle'), timeout: HOOK_TIMEOUT_SECONDS }],
|
||||
@@ -231,8 +375,12 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
|
||||
function isCodemanHookHandler(value: unknown): boolean {
|
||||
try {
|
||||
const serialized = JSON.stringify(value);
|
||||
// Prefix, not the versioned marker: older script versions must still be ours.
|
||||
return serialized.includes('/api/hook-event') || serialized.includes(BACKGROUND_WAKE_MARKER_PREFIX);
|
||||
// Prefixes, not versioned markers: older script versions must still be ours.
|
||||
return (
|
||||
serialized.includes('/api/hook-event') ||
|
||||
serialized.includes(BACKGROUND_WAKE_MARKER_PREFIX) ||
|
||||
serialized.includes(SUBAGENT_STOP_GUARD_MARKER_PREFIX)
|
||||
);
|
||||
} catch {
|
||||
return false;
|
||||
}
|
||||
@@ -426,6 +574,39 @@ export async function writeHooksConfig(casePath: string): Promise<void> {
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Ensures an explicitly managed case has the current Codeman hooks.
|
||||
*
|
||||
* Unlike `refreshStaleCodemanHooks`, this may add Codeman handlers to a valid
|
||||
* user-owned settings file. It is therefore reserved for case quick-starts,
|
||||
* where the user has explicitly asked Codeman to manage that workspace. A
|
||||
* malformed existing file is left untouched rather than replaced.
|
||||
*/
|
||||
export async function ensureCodemanHooks(casePath: string): Promise<void> {
|
||||
const claudeDir = join(casePath, '.claude');
|
||||
const settingsPath = join(claudeDir, 'settings.local.json');
|
||||
await withSettingsLock(settingsPath, async () => {
|
||||
if (!existsSync(claudeDir)) {
|
||||
await mkdir(claudeDir, { recursive: true });
|
||||
}
|
||||
|
||||
let existing: Record<string, unknown> = {};
|
||||
try {
|
||||
const parsed: unknown = JSON.parse(await readFile(settingsPath, 'utf-8'));
|
||||
if (!parsed || typeof parsed !== 'object' || Array.isArray(parsed)) return;
|
||||
existing = parsed as Record<string, unknown>;
|
||||
} catch (err) {
|
||||
if ((err as NodeJS.ErrnoException).code !== 'ENOENT') return;
|
||||
}
|
||||
|
||||
const generated = generateHooksConfig();
|
||||
const hooks = mergeCodemanHooks(existing.hooks, generated.hooks);
|
||||
if (JSON.stringify(existing.hooks ?? {}) === JSON.stringify(hooks)) return;
|
||||
|
||||
await writeFile(settingsPath, JSON.stringify({ ...existing, hooks }, null, 2) + '\n');
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Self-heal a case's Codeman-owned hooks block.
|
||||
*
|
||||
@@ -433,8 +614,11 @@ export async function writeHooksConfig(casePath: string): Promise<void> {
|
||||
* X-Codeman-Hook-Secret header was added (COD-54, 2026-06-10) keep hook curls in their
|
||||
* settings.local.json that POST to /api/hook-event WITHOUT the secret — which, once the
|
||||
* gate requires it unconditionally (COD-91), silently 401 on a password-protected install.
|
||||
* Older Codeman blocks also lack the background Bash async-rewake hook. Refresh either
|
||||
* stale shape on launch so existing cases gain both current behaviors.
|
||||
* Older Codeman blocks also lack the current background Bash async-rewake hook or the
|
||||
* SubagentStop guard. A further stale shape: hook curls without `-k`, which exit 60 on
|
||||
* every --https/tailscale install (the cert is self-signed), swallowed by the hooks'
|
||||
* own `|| true` — all six hook events die silently. Refresh any of these stale shapes
|
||||
* on launch so existing cases heal.
|
||||
*
|
||||
* Deliberately surgical: regenerates ONLY when settings.local.json already contains
|
||||
* Codeman's own hook curls (they target `/api/hook-event`) and they are stale. No-op
|
||||
@@ -457,7 +641,12 @@ export async function refreshStaleCodemanHooks(casePath: string): Promise<void>
|
||||
// absence on our own hooks means they predate COD-54 and need regenerating.
|
||||
const hasSecret = hooksJson.includes('X-Codeman-Hook-Secret');
|
||||
const hasBackgroundWake = hooksJson.includes(BACKGROUND_WAKE_MARKER);
|
||||
if (!isOurs || (hasSecret && hasBackgroundWake)) return;
|
||||
// The pre--k curl shape: `curl -sk -X POST` does not contain `curl -s -X POST`
|
||||
// as a substring, so this cleanly identifies hook curls that die with exit 60
|
||||
// on a self-signed HTTPS install.
|
||||
const hasTlsFlaglessCurl = hooksJson.includes('curl -s -X POST');
|
||||
const hasSubagentStopGuard = hooksJson.includes(SUBAGENT_STOP_GUARD_MARKER);
|
||||
if (!isOurs || (hasSecret && hasBackgroundWake && hasSubagentStopGuard && !hasTlsFlaglessCurl)) return;
|
||||
const generated = generateHooksConfig();
|
||||
const merged = {
|
||||
...existing,
|
||||
|
||||
@@ -0,0 +1,87 @@
|
||||
/**
|
||||
* @fileoverview Bounded descendant walk over a process-tree snapshot.
|
||||
*
|
||||
* Split out of `tmux-manager.ts` so the traversal can be unit-tested directly. It
|
||||
* previously lived as a private method, which meant the regression test had to keep
|
||||
* its own copy of the algorithm — a test that passes while the shipped code rots.
|
||||
*
|
||||
* ## The incident this guards against
|
||||
*
|
||||
* On 2026-07-30 an unbounded version of this walk took a machine down. It ran
|
||||
* `pgrep -P <pid>` once per node and recursed with no visited set, no depth limit and
|
||||
* no node cap. Across ~28 adopted tmux trees the fan-out exploded, and because each
|
||||
* `pgrep` blocks in the WSL kernel while reading `/proc/<pid>/cgroup`, none of them
|
||||
* returned while the walk kept spawning more. Result: ~13,000 `pgrep` processes stuck
|
||||
* in D-state out of ~39,000 total, load average above 13,000, and a machine only
|
||||
* recoverable by restarting WSL — which cost every running session.
|
||||
*
|
||||
* Three properties make that impossible, and each has a test:
|
||||
* 1. a cycle terminates instead of looping (stale snapshots can contain one),
|
||||
* 2. depth is capped,
|
||||
* 3. node count is capped.
|
||||
*
|
||||
* The fourth property — spawning nothing per node — is structural: this function
|
||||
* takes a snapshot and cannot spawn anything at all.
|
||||
*
|
||||
* @module proc-tree
|
||||
*/
|
||||
|
||||
/** Maximum generations to descend. Deeper than any real agent process tree. */
|
||||
export const PROC_WALK_MAX_DEPTH = 10;
|
||||
|
||||
/** Hard ceiling on collected descendants. A backstop, not an expected limit. */
|
||||
export const PROC_WALK_MAX_NODES = 500;
|
||||
|
||||
export interface WalkOptions {
|
||||
maxDepth?: number;
|
||||
maxNodes?: number;
|
||||
/**
|
||||
* Called once when a cap truncated the result, with which cap it was. Both are
|
||||
* reported: a silent depth cap would hide a deep tree just as effectively as a
|
||||
* silent node cap hides a wide one, and the whole point of this module is that
|
||||
* truncation is visible rather than mysterious.
|
||||
*/
|
||||
onTruncated?: (pid: number, cap: number, reason: 'nodes' | 'depth') => void;
|
||||
}
|
||||
|
||||
/**
|
||||
* All descendants of `pid`, breadth-first and bounded.
|
||||
*
|
||||
* @param pid root of the walk; never included in the result
|
||||
* @param byParent parent pid → child pids, from ONE `ps` snapshot
|
||||
*/
|
||||
export function collectDescendants(
|
||||
pid: number,
|
||||
byParent: ReadonlyMap<number, readonly number[]>,
|
||||
opts: WalkOptions = {}
|
||||
): number[] {
|
||||
const maxDepth = opts.maxDepth ?? PROC_WALK_MAX_DEPTH;
|
||||
const maxNodes = opts.maxNodes ?? PROC_WALK_MAX_NODES;
|
||||
|
||||
const out: number[] = [];
|
||||
const visited = new Set<number>([pid]);
|
||||
let frontier = [pid];
|
||||
|
||||
for (let depth = 0; depth < maxDepth && frontier.length; depth += 1) {
|
||||
const next: number[] = [];
|
||||
for (const parent of frontier) {
|
||||
for (const child of byParent.get(parent) ?? []) {
|
||||
if (visited.has(child)) continue; // a real tree has no cycles, a stale
|
||||
visited.add(child); // snapshot can still produce one
|
||||
out.push(child);
|
||||
next.push(child);
|
||||
if (out.length >= maxNodes) {
|
||||
opts.onTruncated?.(pid, maxNodes, 'nodes');
|
||||
return out;
|
||||
}
|
||||
}
|
||||
}
|
||||
frontier = next;
|
||||
// Ran out of generations while descendants were still queued: the tree is
|
||||
// deeper than the cap and the result is incomplete.
|
||||
if (depth === maxDepth - 1 && frontier.length > 0) {
|
||||
opts.onTruncated?.(pid, maxDepth, 'depth');
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
@@ -256,6 +256,69 @@ export async function checkRemoteTmuxAvailable(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The CLI binary each session mode runs on the remote host. Antigravity's
|
||||
* binary is `agy` (the mode name is not the command); shell has no CLI to
|
||||
* probe, so it is absent.
|
||||
*/
|
||||
const REMOTE_CLI_BIN: Partial<Record<SessionMode, string>> = {
|
||||
claude: 'claude',
|
||||
opencode: 'opencode',
|
||||
codex: 'codex',
|
||||
gemini: 'gemini',
|
||||
antigravity: 'agy',
|
||||
};
|
||||
|
||||
/**
|
||||
* Build the SSH command that reads the remote CLI's version (`claude --version`
|
||||
* on the remote host). The version query is routed through
|
||||
* `remoteLoginShellCommand` (the SAME `$SHELL -i -l -c` wrapper the real
|
||||
* launch uses), because agent CLIs live on PATH only after the remote user's
|
||||
* interactive-login startup files run (see defaultRemoteCommandForMode); a bare
|
||||
* `claude --version` over ssh exits 127. Connection options come from the
|
||||
* shared `buildSshConnectionArgs`, so the probe reaches exactly the hosts the
|
||||
* launch can reach. Returns null for modes with no CLI (shell).
|
||||
*/
|
||||
export function buildRemoteCliVersionProbeCommand(
|
||||
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
|
||||
mode: SessionMode
|
||||
): string | null {
|
||||
const bin = REMOTE_CLI_BIN[mode];
|
||||
if (!bin) return null;
|
||||
return [
|
||||
...buildSshConnectionArgs(host),
|
||||
remoteSshTarget(host),
|
||||
shellescape(remoteLoginShellCommand(`${bin} --version`)),
|
||||
].join(' ');
|
||||
}
|
||||
|
||||
/**
|
||||
* Read the CLI version installed ON THE REMOTE HOST. Feeds Session.cliVersion
|
||||
* for remote sessions: the deterministic local probe deliberately skips them
|
||||
* (it would report the LOCAL host's claude), and the startup-banner scrape is
|
||||
* unreliable (newer Claude Code builds print no banner; resumed sessions never
|
||||
* do), which left cliVersion undefined and silently disabled wheel-forwarding
|
||||
* to the CLI transcript (residual #154, noted in the #205 analysis). The
|
||||
* version is parsed as the first semver in stdout, never raw output: an
|
||||
* interactive-login shell may echo rc-file noise around it. Returns undefined
|
||||
* on any failure. No-op under VITEST (mirrors checkRemoteTmuxAvailable).
|
||||
*/
|
||||
export async function probeRemoteCliVersion(
|
||||
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
|
||||
mode: SessionMode
|
||||
): Promise<string | undefined> {
|
||||
if (process.env.VITEST) return undefined;
|
||||
const command = buildRemoteCliVersionProbeCommand(host, mode);
|
||||
if (!command) return undefined;
|
||||
try {
|
||||
const { stdout } = await execAsync(command, { timeout: 15_000 });
|
||||
const match = stdout.match(/\d+\.\d+\.\d+/);
|
||||
return match ? match[0] : undefined;
|
||||
} catch {
|
||||
return undefined;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* COD-105 — build the SSH command that lists `codeman-*` tmux sessions on a
|
||||
* remote host's canonical `-L codeman` socket.
|
||||
|
||||
@@ -0,0 +1,401 @@
|
||||
/**
|
||||
* @fileoverview `codeman service install|uninstall|status`: write and load the
|
||||
* systemd user unit (Linux) or LaunchAgent (macOS) that supervises `codeman web`.
|
||||
*
|
||||
* This is the "always running" half of issue #231, next to the "detached right
|
||||
* now" half in daemon-control.ts. `install.sh` already does this for people who
|
||||
* install with the one-liner; this exists for `npm i -g aicodeman` users, who
|
||||
* otherwise have to hand-write a plist.
|
||||
*
|
||||
* Two details are load-bearing and easy to get wrong by hand:
|
||||
*
|
||||
* - **PATH.** launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` and systemd's
|
||||
* user manager is nearly as bare, so a Homebrew or nvm `node`, `tmux` or
|
||||
* `claude` is simply not found and sessions fail in a way that reads as a
|
||||
* Codeman bug. The unit therefore carries the PATH of the shell that ran the
|
||||
* install, with the running node's own directory in front.
|
||||
* - **The job name.** It is the one `install.sh` and the self-updater already use
|
||||
* (config/service-names.ts), so re-running install.sh later updates this unit
|
||||
* instead of supervising a second copy of the server.
|
||||
*
|
||||
* Secrets are deliberately NOT written here. `CODEMAN_PASSWORD` in the installing
|
||||
* shell is not copied into the unit; the caller is told where to add it instead,
|
||||
* because a unit file is long-lived, world-readable by default, and gets copied
|
||||
* into bug reports.
|
||||
*
|
||||
* The file writers are pure string builders so they can be unit-tested without
|
||||
* touching launchctl/systemctl.
|
||||
*
|
||||
* @module service-installer
|
||||
*/
|
||||
|
||||
import { execFileSync } from 'node:child_process';
|
||||
import { existsSync, mkdirSync, unlinkSync, writeFileSync } from 'node:fs';
|
||||
import { homedir, userInfo } from 'node:os';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { LAUNCHD_LABEL, SYSTEMD_UNIT } from './config/service-names.js';
|
||||
import { CODEMAN_INSTANCE } from './config/instance.js';
|
||||
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
|
||||
import {
|
||||
buildBaseUrl,
|
||||
buildStatusUrl,
|
||||
buildWebArgs,
|
||||
logFilePath,
|
||||
probeServer,
|
||||
type WebLaunchOptions,
|
||||
} from './daemon-control.js';
|
||||
|
||||
export type ServiceKind = 'launchd' | 'systemd';
|
||||
|
||||
/** Everything a unit file needs, resolved from the environment by the caller. */
|
||||
export interface ServicePlan {
|
||||
kind: ServiceKind;
|
||||
/** systemd unit filename or launchd label. */
|
||||
name: string;
|
||||
nodePath: string;
|
||||
/** Runner flags carried over from the current process (tsx loader in dev). */
|
||||
execArgv: string[];
|
||||
scriptPath: string;
|
||||
args: string[];
|
||||
env: Record<string, string>;
|
||||
logPath: string;
|
||||
workingDir: string;
|
||||
}
|
||||
|
||||
export interface ServiceActionResult {
|
||||
ok: boolean;
|
||||
message: string;
|
||||
/** Path of the unit/plist that was written or removed. */
|
||||
unitPath?: string;
|
||||
warnings?: string[];
|
||||
}
|
||||
|
||||
export interface ServiceStatusResult {
|
||||
kind: ServiceKind | null;
|
||||
name: string;
|
||||
unitPath: string;
|
||||
installed: boolean;
|
||||
loaded: boolean;
|
||||
responding: boolean;
|
||||
version?: string;
|
||||
url: string;
|
||||
}
|
||||
|
||||
/** Directories worth having on PATH even when the installing shell lacked them. */
|
||||
const FALLBACK_PATH_DIRS = ['/opt/homebrew/bin', '/usr/local/bin', '/usr/bin', '/bin', '/usr/sbin', '/sbin'];
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
// Pure builders
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
/** XML text escaping for plist `<string>` values. */
|
||||
export function xmlEscape(value: string): string {
|
||||
return value
|
||||
.replace(/&/g, '&')
|
||||
.replace(/</g, '<')
|
||||
.replace(/>/g, '>')
|
||||
.replace(/"/g, '"')
|
||||
.replace(/'/g, ''');
|
||||
}
|
||||
|
||||
/**
|
||||
* PATH for the supervised process: the running node's directory first (so an nvm
|
||||
* or Homebrew node is used rather than whatever the supervisor finds), then the
|
||||
* installing shell's PATH, then the fallbacks that are still missing.
|
||||
*
|
||||
* `node_modules/.bin` entries are dropped. npm and npx inject those for the
|
||||
* lifetime of one command, and baking a project's local bin dir into a unit file
|
||||
* that outlives the checkout is how a service ends up running a binary the
|
||||
* operator deleted months ago.
|
||||
*/
|
||||
export function buildServicePath(nodeDir: string, currentPath: string, home: string): string {
|
||||
const seen = new Set<string>();
|
||||
const ordered: string[] = [];
|
||||
const push = (dir: string) => {
|
||||
const trimmed = dir.trim();
|
||||
if (!trimmed || seen.has(trimmed)) return;
|
||||
if (/(^|\/)node_modules\/\.bin\/?$/.test(trimmed)) return;
|
||||
seen.add(trimmed);
|
||||
ordered.push(trimmed);
|
||||
};
|
||||
|
||||
push(nodeDir);
|
||||
for (const dir of currentPath.split(':')) push(dir);
|
||||
push(join(home, '.local', 'bin'));
|
||||
for (const dir of FALLBACK_PATH_DIRS) push(dir);
|
||||
return ordered.join(':');
|
||||
}
|
||||
|
||||
/** Environment written into the unit. Never includes secrets (see module docs). */
|
||||
export function buildServiceEnv(
|
||||
nodeDir: string,
|
||||
currentPath: string,
|
||||
home: string,
|
||||
lang?: string
|
||||
): Record<string, string> {
|
||||
const env: Record<string, string> = {
|
||||
PATH: buildServicePath(nodeDir, currentPath, home),
|
||||
HOME: home,
|
||||
LANG: lang || 'en_US.UTF-8',
|
||||
};
|
||||
if (CODEMAN_INSTANCE) env.CODEMAN_INSTANCE = CODEMAN_INSTANCE;
|
||||
return env;
|
||||
}
|
||||
|
||||
/** systemd accepts double-quoted values; escape the two characters that matter. */
|
||||
export function systemdQuote(value: string): string {
|
||||
return `"${value.replace(/\\/g, '\\\\').replace(/"/g, '\\"')}"`;
|
||||
}
|
||||
|
||||
export function buildLaunchAgentPlist(plan: ServicePlan): string {
|
||||
const programArguments = [plan.nodePath, ...plan.execArgv, plan.scriptPath, ...plan.args]
|
||||
.map((arg) => ` <string>${xmlEscape(arg)}</string>`)
|
||||
.join('\n');
|
||||
const environment = Object.entries(plan.env)
|
||||
.map(([key, value]) => ` <key>${xmlEscape(key)}</key>\n <string>${xmlEscape(value)}</string>`)
|
||||
.join('\n');
|
||||
|
||||
return `<?xml version="1.0" encoding="UTF-8"?>
|
||||
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
||||
<plist version="1.0">
|
||||
<dict>
|
||||
<key>Label</key>
|
||||
<string>${xmlEscape(plan.name)}</string>
|
||||
<key>ProgramArguments</key>
|
||||
<array>
|
||||
${programArguments}
|
||||
</array>
|
||||
<key>EnvironmentVariables</key>
|
||||
<dict>
|
||||
${environment}
|
||||
</dict>
|
||||
<key>WorkingDirectory</key>
|
||||
<string>${xmlEscape(plan.workingDir)}</string>
|
||||
<key>RunAtLoad</key>
|
||||
<true/>
|
||||
<key>KeepAlive</key>
|
||||
<true/>
|
||||
<key>ThrottleInterval</key>
|
||||
<integer>10</integer>
|
||||
<key>StandardOutPath</key>
|
||||
<string>${xmlEscape(plan.logPath)}</string>
|
||||
<key>StandardErrorPath</key>
|
||||
<string>${xmlEscape(plan.logPath)}</string>
|
||||
</dict>
|
||||
</plist>
|
||||
`;
|
||||
}
|
||||
|
||||
export function buildSystemdUnit(plan: ServicePlan): string {
|
||||
const execStart = [plan.nodePath, ...plan.execArgv, plan.scriptPath, ...plan.args]
|
||||
.map((arg) => (/[\s"'\\]/.test(arg) ? systemdQuote(arg) : arg))
|
||||
.join(' ');
|
||||
const environment = Object.entries(plan.env)
|
||||
.map(([key, value]) => `Environment=${systemdQuote(`${key}=${value}`)}`)
|
||||
.join('\n');
|
||||
|
||||
return `[Unit]
|
||||
Description=Codeman Web Server
|
||||
After=network.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
WorkingDirectory=${plan.workingDir}
|
||||
ExecStart=${execStart}
|
||||
Restart=always
|
||||
RestartSec=10
|
||||
# Agents keep running in tmux when the server restarts, so only signal the
|
||||
# server itself.
|
||||
KillMode=process
|
||||
${environment}
|
||||
StandardOutput=journal
|
||||
StandardError=journal
|
||||
SyslogIdentifier=codeman
|
||||
LimitNOFILE=65536
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
`;
|
||||
}
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
// Environment resolution
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
export function detectServiceKind(): ServiceKind | null {
|
||||
if (process.platform === 'darwin') return 'launchd';
|
||||
if (process.platform === 'linux') return 'systemd';
|
||||
return null;
|
||||
}
|
||||
|
||||
export function unitPathFor(kind: ServiceKind): string {
|
||||
return kind === 'launchd'
|
||||
? join(homedir(), 'Library', 'LaunchAgents', `${LAUNCHD_LABEL}.plist`)
|
||||
: join(homedir(), '.config', 'systemd', 'user', SYSTEMD_UNIT);
|
||||
}
|
||||
|
||||
function entryScript(): string {
|
||||
const script = process.argv[1];
|
||||
if (!script) throw new Error('cannot determine the codeman entry script to supervise');
|
||||
return script;
|
||||
}
|
||||
|
||||
/** Resolve a full plan from the current process and the requested web options. */
|
||||
export function resolveServicePlan(kind: ServiceKind, options: WebLaunchOptions): ServicePlan {
|
||||
const home = homedir();
|
||||
return {
|
||||
kind,
|
||||
name: kind === 'launchd' ? LAUNCHD_LABEL : SYSTEMD_UNIT,
|
||||
nodePath: process.execPath,
|
||||
execArgv: [...process.execArgv],
|
||||
scriptPath: entryScript(),
|
||||
args: buildWebArgs(options),
|
||||
env: buildServiceEnv(dirname(process.execPath), process.env.PATH || '', home, process.env.LANG),
|
||||
logPath: logFilePath(),
|
||||
workingDir: home,
|
||||
};
|
||||
}
|
||||
|
||||
function run(command: string, args: string[]): { ok: boolean; output: string } {
|
||||
try {
|
||||
const output = execFileSync(command, args, {
|
||||
encoding: 'utf-8',
|
||||
timeout: EXEC_TIMEOUT_MS,
|
||||
stdio: ['ignore', 'pipe', 'pipe'],
|
||||
});
|
||||
return { ok: true, output: output.trim() };
|
||||
} catch (err) {
|
||||
const e = err as { stderr?: Buffer | string; message?: string };
|
||||
const stderr = typeof e.stderr === 'string' ? e.stderr : e.stderr?.toString('utf-8');
|
||||
return { ok: false, output: (stderr || e.message || '').trim() };
|
||||
}
|
||||
}
|
||||
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
// Install / uninstall / status
|
||||
// ─────────────────────────────────────────────────────────────────────────────
|
||||
|
||||
/**
|
||||
* Write the unit, load it, and confirm the server actually answers before
|
||||
* reporting success. `launchctl load` and `systemctl enable` are both quiet about
|
||||
* a job that starts and immediately dies, which is the whole reason install.sh
|
||||
* verifies too.
|
||||
*/
|
||||
export async function installService(options: WebLaunchOptions): Promise<ServiceActionResult> {
|
||||
const kind = detectServiceKind();
|
||||
if (!kind) {
|
||||
return { ok: false, message: `no supported supervisor on ${process.platform}; use \`codeman web -d\` instead` };
|
||||
}
|
||||
|
||||
const plan = resolveServicePlan(kind, options);
|
||||
const unitPath = unitPathFor(kind);
|
||||
const warnings: string[] = [];
|
||||
mkdirSync(dirname(unitPath), { recursive: true });
|
||||
|
||||
if (kind === 'launchd') {
|
||||
const uid = process.getuid?.() ?? 0;
|
||||
// Unload any previous copy first, otherwise bootstrap fails with "service
|
||||
// already loaded" and leaves the OLD job running against the NEW file.
|
||||
run('launchctl', ['bootout', `gui/${uid}/${LAUNCHD_LABEL}`]);
|
||||
writeFileSync(unitPath, buildLaunchAgentPlist(plan), { encoding: 'utf-8', mode: 0o600 });
|
||||
const bootstrap = run('launchctl', ['bootstrap', `gui/${uid}`, unitPath]);
|
||||
if (!bootstrap.ok) {
|
||||
const legacy = run('launchctl', ['load', unitPath]);
|
||||
if (!legacy.ok) {
|
||||
return {
|
||||
ok: false,
|
||||
unitPath,
|
||||
message: `wrote ${unitPath} but launchctl refused to load it: ${bootstrap.output}`,
|
||||
};
|
||||
}
|
||||
}
|
||||
} else {
|
||||
writeFileSync(unitPath, buildSystemdUnit(plan), { encoding: 'utf-8', mode: 0o600 });
|
||||
const reload = run('systemctl', ['--user', 'daemon-reload']);
|
||||
if (!reload.ok) {
|
||||
return {
|
||||
ok: false,
|
||||
unitPath,
|
||||
message: `wrote ${unitPath} but \`systemctl --user daemon-reload\` failed: ${reload.output}`,
|
||||
};
|
||||
}
|
||||
const enable = run('systemctl', ['--user', 'enable', '--now', SYSTEMD_UNIT]);
|
||||
if (!enable.ok) {
|
||||
return { ok: false, unitPath, message: `wrote ${unitPath} but enabling it failed: ${enable.output}` };
|
||||
}
|
||||
// Without lingering the unit stops at logout, which is exactly what someone
|
||||
// installing a service does not want. Best effort: it needs polkit rights.
|
||||
const linger = run('loginctl', ['enable-linger', userInfo().username]);
|
||||
if (!linger.ok) {
|
||||
warnings.push(
|
||||
`could not enable lingering, so the service will stop when you log out. Run: sudo loginctl enable-linger ${userInfo().username}`
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
const url = buildBaseUrl(options);
|
||||
const statusUrl = buildStatusUrl(options);
|
||||
const deadline = Date.now() + 30_000;
|
||||
while (Date.now() < deadline) {
|
||||
const probe = await probeServer(statusUrl, 1000);
|
||||
if (probe.up) {
|
||||
return { ok: true, unitPath, warnings, message: `service installed and responding at ${url}` };
|
||||
}
|
||||
await new Promise((resolve) => setTimeout(resolve, 500));
|
||||
}
|
||||
|
||||
const hint =
|
||||
kind === 'launchd' ? `tail -20 ${plan.logPath}` : `journalctl --user -u ${SYSTEMD_UNIT} -n 20 --no-pager`;
|
||||
return {
|
||||
ok: false,
|
||||
unitPath,
|
||||
warnings,
|
||||
message: `wrote and loaded ${unitPath}, but nothing answered ${url} within 30s. Check: ${hint}`,
|
||||
};
|
||||
}
|
||||
|
||||
export function uninstallService(): ServiceActionResult {
|
||||
const kind = detectServiceKind();
|
||||
if (!kind) return { ok: false, message: `no supported supervisor on ${process.platform}` };
|
||||
|
||||
const unitPath = unitPathFor(kind);
|
||||
if (!existsSync(unitPath)) {
|
||||
return { ok: false, unitPath, message: `no service installed at ${unitPath}` };
|
||||
}
|
||||
|
||||
if (kind === 'launchd') {
|
||||
const uid = process.getuid?.() ?? 0;
|
||||
const bootout = run('launchctl', ['bootout', `gui/${uid}/${LAUNCHD_LABEL}`]);
|
||||
if (!bootout.ok) run('launchctl', ['unload', unitPath]);
|
||||
} else {
|
||||
run('systemctl', ['--user', 'disable', '--now', SYSTEMD_UNIT]);
|
||||
}
|
||||
|
||||
try {
|
||||
unlinkSync(unitPath);
|
||||
} catch (err) {
|
||||
return { ok: false, unitPath, message: `stopped the service but could not remove ${unitPath}: ${String(err)}` };
|
||||
}
|
||||
if (kind === 'systemd') run('systemctl', ['--user', 'daemon-reload']);
|
||||
|
||||
return { ok: true, unitPath, message: `service stopped and ${unitPath} removed. Your tmux sessions are untouched.` };
|
||||
}
|
||||
|
||||
export async function serviceStatus(options: WebLaunchOptions): Promise<ServiceStatusResult> {
|
||||
const kind = detectServiceKind();
|
||||
const url = buildBaseUrl(options);
|
||||
if (!kind) {
|
||||
return { kind: null, name: '', unitPath: '', installed: false, loaded: false, responding: false, url };
|
||||
}
|
||||
|
||||
const unitPath = unitPathFor(kind);
|
||||
const name = kind === 'launchd' ? LAUNCHD_LABEL : SYSTEMD_UNIT;
|
||||
const installed = existsSync(unitPath);
|
||||
const loaded =
|
||||
kind === 'launchd'
|
||||
? run('launchctl', ['list', LAUNCHD_LABEL]).ok
|
||||
: run('systemctl', ['--user', 'is-active', SYSTEMD_UNIT]).output === 'active';
|
||||
const probe = await probeServer(buildStatusUrl(options), 2000);
|
||||
|
||||
return { kind, name, unitPath, installed, loaded, responding: probe.up, version: probe.version, url };
|
||||
}
|
||||
@@ -121,7 +121,10 @@ export function buildClaudeEnv(sessionId: string): Record<string, string | undef
|
||||
// Inform Claude it's running within Codeman (helps prevent self-termination)
|
||||
CODEMAN_MUX: '1',
|
||||
CODEMAN_SESSION_ID: sessionId,
|
||||
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
|
||||
// CODEMAN_API_URL rides in via the process.env spread when the server has
|
||||
// stamped it (WebServer.start()); no fallback: a hardcoded one was the wrong
|
||||
// scheme on HTTPS installs, and a present-with-undefined key would serialize
|
||||
// as the literal "CODEMAN_API_URL=undefined" (COD-115).
|
||||
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
|
||||
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
|
||||
};
|
||||
@@ -179,7 +182,8 @@ export function buildShellEnv(sessionId: string): Record<string, string | undefi
|
||||
TERM: 'xterm-256color',
|
||||
CODEMAN_MUX: '1',
|
||||
CODEMAN_SESSION_ID: sessionId,
|
||||
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
|
||||
// CODEMAN_API_URL rides in via the process.env spread when set; no fallback
|
||||
// (same reasoning as buildClaudeEnv above).
|
||||
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
|
||||
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
|
||||
};
|
||||
|
||||
+116
-18
@@ -54,6 +54,7 @@ import {
|
||||
type SessionDocker,
|
||||
} from './types.js';
|
||||
import { probeDockerCliVersion } from './docker-hosts.js';
|
||||
import { probeRemoteCliVersion } from './remote-hosts.js';
|
||||
import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
|
||||
import { TaskTracker, type BackgroundTask } from './task-tracker.js';
|
||||
import { RalphTracker } from './ralph-tracker.js';
|
||||
@@ -180,6 +181,37 @@ export function isAltScreenStripMode(mode: SessionMode): boolean {
|
||||
return mode === 'codex' || mode === 'claude' || mode === 'gemini';
|
||||
}
|
||||
|
||||
/**
|
||||
* Modes that need the NARROW strip: alt-screen toggles only, leaving `\x1b[3J`
|
||||
* and the mouse-tracking DECSETs alone. Applies to every mode `isAltScreenStripMode`
|
||||
* excludes, but ONLY when the session is tmux-backed (`useMux`).
|
||||
*
|
||||
* The bug (issue #205): the tmux CLIENT emits `smcup` (`\x1b[?1049h`) as its first
|
||||
* bytes on attach, before any program has run. Unstripped, xterm.js parks in the
|
||||
* alternate buffer for the whole session, where `baseY` is pinned at 0 (no
|
||||
* scrollback to reach, so touch scrolling is a no-op) and xterm's own wheel handler
|
||||
* translates the wheel into `\x1bOA`/`\x1bOB` cursor keys — which readline receives
|
||||
* as shell history navigation. Both reported symptoms, one sequence.
|
||||
*
|
||||
* Why this is safe under tmux, despite the old "shell must keep the alt screen for
|
||||
* vim/less/htop" reasoning: tmux is a full terminal emulator and NEVER forwards a
|
||||
* pane's alt-screen toggles to its client, it repaints instead. Captured from a real
|
||||
* attach, `\x1b[?1049h` appears exactly once (at attach) and vim/less/htop sessions
|
||||
* inside the pane emit zero. So the only thing stripped here is tmux's own smcup.
|
||||
*
|
||||
* Why it is gated on `useMux`: `startShell()`/`startInteractive()` fall back to a
|
||||
* DIRECT PTY when mux creation fails. There the inner program's `\x1b[?1049h` really
|
||||
* does reach xterm, and stripping it would break vim/less/htop for real.
|
||||
*
|
||||
* Why it is narrower than the full strip: with tmux `mouse off`, a mouse-aware
|
||||
* program in the pane (htop, vim with `set mouse=a`) still gets its DECSETs passed
|
||||
* through to the client, so stripping those would break its mouse support. And
|
||||
* `\x1b[3J` from a user's own `clear` is a deliberate "wipe my scrollback".
|
||||
*/
|
||||
export function isMuxAltScreenOnlyStripMode(mode: SessionMode, useMux: boolean): boolean {
|
||||
return useMux && !isAltScreenStripMode(mode);
|
||||
}
|
||||
|
||||
// Note: Claude CLI PATH resolution moved to session-cli-builder.ts (buildClaudeEnv)
|
||||
|
||||
/** PTY fallback geometry when tmux can't be queried (matches pre-#80 hardcoded values). */
|
||||
@@ -195,6 +227,8 @@ const IS_TEST_MODE = !!process.env.VITEST;
|
||||
const TEST_PTY_SCRIPT = 'if (process.stdin.isTTY) process.stdin.setRawMode(true); process.stdin.pipe(process.stdout);';
|
||||
/** Delay before the in-container Claude CLI version probe (lets the container start). */
|
||||
const DOCKER_CLI_VERSION_PROBE_DELAY_MS = 3000;
|
||||
/** Delay before the over-ssh Claude CLI version probe (keeps session start off the ssh round-trip). */
|
||||
const REMOTE_CLI_VERSION_PROBE_DELAY_MS = 3000;
|
||||
|
||||
/**
|
||||
* Ask tmux for the current window geometry of `muxName` so a re-attaching PTY
|
||||
@@ -733,6 +767,15 @@ export class Session extends EventEmitter {
|
||||
return this._muxSession?.muxName ?? null;
|
||||
}
|
||||
|
||||
/**
|
||||
* True when this session's PTY is a tmux client rather than the program itself.
|
||||
* Read by the replay-side alt-screen strip, which must apply the same
|
||||
* `useMux` gate as the live strip (isMuxAltScreenOnlyStripMode).
|
||||
*/
|
||||
get usesMux(): boolean {
|
||||
return this._useMux;
|
||||
}
|
||||
|
||||
get totalCost(): number {
|
||||
return this._totalCost;
|
||||
}
|
||||
@@ -1401,9 +1444,16 @@ export class Session extends EventEmitter {
|
||||
// SSE/WS stream carries them, keeping everything in the main buffer with
|
||||
// scrollback intact. These are controlled TUIs whose cursor-positioned
|
||||
// redraws overwrite only the cells they target, so non-erased rows keep
|
||||
// their content. Gated to Codex/Claude (isAltScreenStripMode) — shell must
|
||||
// keep the alt screen for vim/less/htop.
|
||||
if (isAltScreenStripMode(this.mode)) {
|
||||
// their content. Gated to Codex/Claude/Gemini (isAltScreenStripMode).
|
||||
//
|
||||
// Every OTHER mode (shell/opencode/antigravity) gets the NARROW strip when it
|
||||
// is tmux-backed: alt-screen toggles only, because the sequence that breaks
|
||||
// scrollback there is tmux's own client-side smcup at attach, not anything the
|
||||
// program in the pane emitted (issue #205, see isMuxAltScreenOnlyStripMode).
|
||||
// 3J and the mouse DECSETs stay, so `clear` and mouse-aware TUIs keep working.
|
||||
const fullStrip = isAltScreenStripMode(this.mode);
|
||||
const altOnlyStrip = !fullStrip && isMuxAltScreenOnlyStripMode(this.mode, this._useMux);
|
||||
if (fullStrip || altOnlyStrip) {
|
||||
// Reassemble sequences split across PTY chunk boundaries first: a chunk
|
||||
// ending mid-sequence ('\x1b[?104' now, '9h' next) would slip past the
|
||||
// strip below and leave xterm stuck in the scrollback-less alt buffer
|
||||
@@ -1419,13 +1469,15 @@ export class Session extends EventEmitter {
|
||||
data = data.slice(0, -splitTail[0].length);
|
||||
if (!data) return;
|
||||
}
|
||||
data = data
|
||||
// eslint-disable-next-line no-control-regex
|
||||
.replace(/\x1b\[\?(?:47|1047|1049)[hl]/g, '')
|
||||
// eslint-disable-next-line no-control-regex
|
||||
.replace(/\x1b\[3J/g, '')
|
||||
// eslint-disable-next-line no-control-regex
|
||||
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
|
||||
// eslint-disable-next-line no-control-regex
|
||||
data = data.replace(/\x1b\[\?(?:47|1047|1049)[hl]/g, '');
|
||||
if (fullStrip) {
|
||||
data = data
|
||||
// eslint-disable-next-line no-control-regex
|
||||
.replace(/\x1b\[3J/g, '')
|
||||
// eslint-disable-next-line no-control-regex
|
||||
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
|
||||
}
|
||||
}
|
||||
|
||||
// Scan terminal output for attachment requests. `codeman://attach?...` is an
|
||||
@@ -1485,8 +1537,8 @@ export class Session extends EventEmitter {
|
||||
// never show it — which left cliVersion undefined and silently disabled
|
||||
// wheel-forwarding to Claude's own transcript (the only route to history in
|
||||
// repaint/alt-screen mode; issue #154). Remote sessions run claude on
|
||||
// another host, so a local probe wouldn't reflect their version — skip them
|
||||
// and let the banner scrape handle those. Cached process-wide, best-effort.
|
||||
// another host, so a local probe wouldn't reflect their version; they get
|
||||
// their own over-ssh probe below. Cached process-wide, best-effort.
|
||||
if (this.mode === 'claude' && !this._remote && !this._docker && !this._cliVersion) {
|
||||
const probedVersion = getClaudeCliVersion();
|
||||
if (probedVersion) {
|
||||
@@ -1525,6 +1577,32 @@ export class Session extends EventEmitter {
|
||||
}, DOCKER_CLI_VERSION_PROBE_DELAY_MS);
|
||||
}
|
||||
|
||||
// Remote sessions run claude on ANOTHER HOST, so neither the local nor the
|
||||
// docker probe applies, and the banner-scrape fallback they were left with
|
||||
// is the unreliable path #154 was filed for, so remote Claude cases silently
|
||||
// never got wheel-forwarding (noted in the #205 analysis). Probe over ssh,
|
||||
// deferred so session start never waits on the ssh round-trip.
|
||||
if (this.mode === 'claude' && this._remote && !this._cliVersion) {
|
||||
const remoteMeta = this._remote;
|
||||
setTimeout(() => {
|
||||
if (this._isStopped || this._cliVersion) return;
|
||||
void probeRemoteCliVersion(remoteMeta, this.mode)
|
||||
.then((version) => {
|
||||
if (!version || this._isStopped || this._cliVersion) return;
|
||||
this._cliVersion = version;
|
||||
this.emit('cliInfoUpdated', {
|
||||
version: this._cliVersion,
|
||||
model: this._cliModel,
|
||||
accountType: this._cliAccountType,
|
||||
latestVersion: this._cliLatestVersion,
|
||||
});
|
||||
})
|
||||
.catch(() => {
|
||||
/* best-effort */
|
||||
});
|
||||
}, REMOTE_CLI_VERSION_PROBE_DELAY_MS);
|
||||
}
|
||||
|
||||
// If mux wrapping is enabled, create or attach to a mux session
|
||||
if (this._useMux && this._mux) {
|
||||
try {
|
||||
@@ -2542,20 +2620,23 @@ export class Session extends EventEmitter {
|
||||
* For interactive sessions, this is how you send user input to Claude.
|
||||
* Remember to include `\r` (carriage return) to simulate pressing Enter.
|
||||
*
|
||||
* @param data - The input data to send (text, escape sequences, etc.)
|
||||
*
|
||||
* @example
|
||||
* ```typescript
|
||||
* session.write('hello world'); // Text only, no Enter
|
||||
* session.write('\r'); // Enter key
|
||||
* session.write('ls -la\r'); // Command with Enter
|
||||
* ```
|
||||
*
|
||||
* @param data - The input data to send (text, escape sequences, etc.)
|
||||
* @returns true if the data reached a PTY. A session whose PTY is gone still
|
||||
* discards the data, but it used to do so with no signal at all — which is how
|
||||
* input could disappear while the caller believed it had been delivered.
|
||||
*/
|
||||
write(data: string): void {
|
||||
write(data: string): boolean {
|
||||
this._trackSubmit(data);
|
||||
if (this.ptyProcess) {
|
||||
this.ptyProcess.write(data);
|
||||
}
|
||||
if (!this.ptyProcess) return false;
|
||||
this.ptyProcess.write(data);
|
||||
return true;
|
||||
}
|
||||
|
||||
// ── Conversation tracking ─────────────────────────────────────────────
|
||||
@@ -2613,6 +2694,23 @@ export class Session extends EventEmitter {
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* Undo the bookkeeping of {@link shouldApplyInput} for a delivery that failed.
|
||||
*
|
||||
* Without this, the reliable-delivery layer guarantees exactly-once delivery of
|
||||
* something that may never have been delivered: the seq is recorded as applied
|
||||
* BEFORE the write is attempted, so a client retry — the very mechanism the seq
|
||||
* exists for — is rejected as a duplicate and the input is lost for good.
|
||||
*
|
||||
* Only rolls back if `seq` is still the newest recorded one; a later input has
|
||||
* already superseded it and must not be re-opened.
|
||||
*/
|
||||
forgetInputSeq(clientId: string, seq: number): void {
|
||||
if (this._appliedInputSeq.get(clientId) === seq) {
|
||||
this._appliedInputSeq.set(clientId, seq - 1);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Sends input via the terminal multiplexer's direct input mechanism.
|
||||
*
|
||||
|
||||
+123
-42
@@ -22,7 +22,8 @@
|
||||
*/
|
||||
|
||||
import { EventEmitter } from 'node:events';
|
||||
import { execSync, exec } from 'node:child_process';
|
||||
import { collectDescendants } from './proc-tree.js';
|
||||
import { execSync, exec, execFile } from 'node:child_process';
|
||||
import { promisify } from 'node:util';
|
||||
|
||||
const execAsync = promisify(exec);
|
||||
@@ -100,6 +101,16 @@ import {
|
||||
// ============================================================================
|
||||
|
||||
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
|
||||
|
||||
/** How long a cached process snapshot stays usable. */
|
||||
const PROC_SNAPSHOT_TTL_MS = 2000;
|
||||
|
||||
/**
|
||||
* How long the kill path waits for a fresh snapshot before giving up on it.
|
||||
* Shorter than EXEC_TIMEOUT_MS on purpose: killSession has two further strategies
|
||||
* (process group, tmux kill-session) and must reach them even when `ps` is wedged.
|
||||
*/
|
||||
const PROC_SNAPSHOT_WAIT_MS = 1500;
|
||||
import { DEFAULT_TMUX_HISTORY_LIMIT, DEFAULT_TERMINAL_BUFFER_MAX_BYTES } from './config/terminal-history.js';
|
||||
|
||||
/**
|
||||
@@ -982,7 +993,7 @@ export function dockerTmuxSessionName(sessionId: string): string {
|
||||
const RESUME_ID_SAFE = /^[A-Za-z0-9._-]+$/;
|
||||
|
||||
/**
|
||||
* Append the CLI-specific resume flag to a pane command (codex/gemini). Only fires
|
||||
* Append the CLI-specific resume flag to a pane command (codex/gemini/antigravity). Only fires
|
||||
* when the in-container tmux is RE-CREATED (`new-session -A` makes the flag inert
|
||||
* on a live reattach), i.e. exactly when the previous live agent was lost and we
|
||||
* want to resume the conversation from the bind-mounted transcript. Claude mode
|
||||
@@ -1578,7 +1589,11 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
|
||||
'export CODEMAN_MUX=1',
|
||||
`export CODEMAN_SESSION_ID=${sessionId}`,
|
||||
`export CODEMAN_MUX_NAME=${muxName}`,
|
||||
`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL || 'http://localhost:3000'}`,
|
||||
// Only exported when the server has stamped the real URL (scheme+host+port,
|
||||
// set in WebServer.start()). A hardcoded fallback here exported the wrong
|
||||
// scheme on HTTPS installs; leaving the variable unset makes in-session
|
||||
// guards fail closed instead of curling a URL that was never right.
|
||||
...(process.env.CODEMAN_API_URL ? [`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL}`] : []),
|
||||
// Path only (not the secret value): hook curl commands cat the file at
|
||||
// execution time, so the COD-54 hook secret stays off the command line.
|
||||
`export CODEMAN_HOOK_SECRET_FILE="${dataPath('hook-secret')}"`,
|
||||
@@ -2094,27 +2109,102 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
|
||||
}
|
||||
}
|
||||
|
||||
// Get all child process PIDs recursively
|
||||
private getChildPids(pid: number): number[] {
|
||||
const pids: number[] = [];
|
||||
try {
|
||||
const output = execSync(`pgrep -P ${pid}`, {
|
||||
encoding: 'utf-8',
|
||||
timeout: EXEC_TIMEOUT_MS,
|
||||
}).trim();
|
||||
if (output) {
|
||||
for (const childPid of output
|
||||
.split('\n')
|
||||
.map((p) => parseInt(p, 10))
|
||||
.filter((p) => !Number.isNaN(p))) {
|
||||
pids.push(childPid);
|
||||
pids.push(...this.getChildPids(childPid));
|
||||
/** One `ps` snapshot of the whole process table, cached briefly. */
|
||||
private static procSnapshot: { at: number; byParent: Map<number, number[]> } | null = null;
|
||||
/** Single-flight guard so a hung `ps` cannot pile up parallel refreshes. */
|
||||
private static procRefresh: { started: number; promise: Promise<Map<number, number[]>> } | null = null;
|
||||
|
||||
/**
|
||||
* Fork ONE `ps` asynchronously and cache the parent -> children map.
|
||||
*
|
||||
* Async on purpose: a synchronous fork here would block the event loop on every
|
||||
* stats tick, and under the procfs pathology this module exists to survive,
|
||||
* `execSync`'s timeout cannot return at all (spawnSync waits for the unkillable
|
||||
* child) — freezing the whole server where a hung async poll only costs staleness.
|
||||
*/
|
||||
private static refreshProcSnapshot(): Promise<Map<number, number[]>> {
|
||||
const inFlight = TmuxManager.procRefresh;
|
||||
// Reuse an in-flight refresh — unless it is old enough to be presumed stuck.
|
||||
if (inFlight && Date.now() - inFlight.started < EXEC_TIMEOUT_MS * 2) return inFlight.promise;
|
||||
|
||||
const started = Date.now();
|
||||
const promise = new Promise<Map<number, number[]>>((resolve) => {
|
||||
execFile('ps', ['-eo', 'pid=,ppid='], { timeout: EXEC_TIMEOUT_MS, maxBuffer: 8 * 1024 * 1024 }, (err, out) => {
|
||||
if (TmuxManager.procRefresh?.started === started) TmuxManager.procRefresh = null;
|
||||
if (err) {
|
||||
// ANY error, not just an empty one: a timed-out or truncated `ps` yields
|
||||
// partial output, and caching that as fresh would make whole subtrees
|
||||
// invisible — including to the kill path. Stale beats wrong.
|
||||
console.error('[TmuxManager] process snapshot failed:', err);
|
||||
resolve(TmuxManager.procSnapshot?.byParent ?? new Map());
|
||||
return;
|
||||
}
|
||||
}
|
||||
} catch {
|
||||
// No children or command failed
|
||||
const byParent = new Map<number, number[]>();
|
||||
for (const line of String(out).split('\n')) {
|
||||
const parts = line.trim().split(/\s+/);
|
||||
if (parts.length < 2) continue;
|
||||
const pid = parseInt(parts[0], 10);
|
||||
const ppid = parseInt(parts[1], 10);
|
||||
if (Number.isNaN(pid) || Number.isNaN(ppid)) continue;
|
||||
const list = byParent.get(ppid);
|
||||
if (list) list.push(pid);
|
||||
else byParent.set(ppid, [pid]);
|
||||
}
|
||||
TmuxManager.procSnapshot = { at: Date.now(), byParent };
|
||||
resolve(byParent);
|
||||
});
|
||||
});
|
||||
TmuxManager.procRefresh = { started, promise };
|
||||
return promise;
|
||||
}
|
||||
|
||||
/**
|
||||
* Best snapshot WITHOUT forking: returns the cache, kicking off a background
|
||||
* refresh when it has gone stale, and never blocks. Stats and window-title
|
||||
* consumers tolerate data one interval old; nothing that KILLS may use this.
|
||||
*/
|
||||
private childrenByParent(): Map<number, number[]> {
|
||||
const cached = TmuxManager.procSnapshot;
|
||||
if (!cached || Date.now() - cached.at >= PROC_SNAPSHOT_TTL_MS) {
|
||||
void TmuxManager.refreshProcSnapshot();
|
||||
}
|
||||
return pids;
|
||||
return cached?.byParent ?? new Map();
|
||||
}
|
||||
|
||||
/**
|
||||
* Descendants from a snapshot that is not the cached one — the kill path's variant.
|
||||
*
|
||||
* killSession re-scans for survivors between SIGTERM and SIGKILL, and the wait in
|
||||
* between (200ms) sits far inside the cache TTL (2000ms): reading the cache there
|
||||
* returns the pre-SIGTERM state verbatim, so children spawned since are invisible
|
||||
* and SIGKILL aims at stale PIDs, guarded only by kill(pid, 0) — which cannot
|
||||
* detect PID reuse.
|
||||
*
|
||||
* It forces a refresh rather than guaranteeing recency: an already-running refresh
|
||||
* is reused, so the snapshot can predate this call by up to one `ps` runtime. A
|
||||
* strict postdate guarantee would mean chaining a second `ps` behind every
|
||||
* in-flight one, which is the fork storm this code exists to avoid.
|
||||
*
|
||||
* Bounded by design: waiting forever would freeze killSession before it reaches
|
||||
* its process-group and tmux fallbacks.
|
||||
*/
|
||||
private async getChildPidsFresh(pid: number): Promise<number[]> {
|
||||
let byParent: ReadonlyMap<number, readonly number[]>;
|
||||
try {
|
||||
byParent = await Promise.race([
|
||||
TmuxManager.refreshProcSnapshot(),
|
||||
new Promise<never>((_, reject) =>
|
||||
setTimeout(() => reject(new Error('proc snapshot timeout')), PROC_SNAPSHOT_WAIT_MS)
|
||||
),
|
||||
]);
|
||||
} catch {
|
||||
console.warn('[TmuxManager] process snapshot did not return in time; using the cached one');
|
||||
byParent = TmuxManager.procSnapshot?.byParent ?? new Map<number, number[]>();
|
||||
}
|
||||
return collectDescendants(pid, byParent, {
|
||||
onTruncated: (root, cap, reason) =>
|
||||
console.warn(`[TmuxManager] descendant walk for ${root} hit the ${cap}-${reason} cap; truncating`),
|
||||
});
|
||||
}
|
||||
|
||||
// Check if a process is still alive
|
||||
@@ -2221,7 +2311,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
|
||||
const allPids: number[] = [currentPid];
|
||||
|
||||
// Strategy 1: Kill all child processes recursively
|
||||
let childPids = this.getChildPids(currentPid);
|
||||
let childPids = await this.getChildPidsFresh(currentPid);
|
||||
if (childPids.length > 0) {
|
||||
console.log(`[TmuxManager] Found ${childPids.length} child processes to kill`);
|
||||
allPids.push(...childPids);
|
||||
@@ -2238,7 +2328,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, TMUX_KILL_WAIT_MS));
|
||||
|
||||
childPids = this.getChildPids(currentPid);
|
||||
childPids = await this.getChildPidsFresh(currentPid);
|
||||
for (const childPid of childPids) {
|
||||
if (this.isProcessAlive(childPid)) {
|
||||
try {
|
||||
@@ -2452,17 +2542,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
|
||||
|
||||
const [rss, cpu] = psOutput.split(/\s+/).map((x) => parseFloat(x) || 0);
|
||||
|
||||
// From the shared snapshot: this runs per session on every stats tick, and a
|
||||
// pgrep per session was a fork per session per interval.
|
||||
let childCount = 0;
|
||||
try {
|
||||
const childOutput = (
|
||||
await execAsync(`pgrep -P ${session.pid} | wc -l`, {
|
||||
encoding: 'utf-8',
|
||||
timeout: EXEC_TIMEOUT_MS,
|
||||
})
|
||||
).stdout.trim();
|
||||
childCount = parseInt(childOutput, 10) || 0;
|
||||
childCount = (this.childrenByParent().get(session.pid) ?? []).length;
|
||||
} catch {
|
||||
// No children or command failed
|
||||
// No children or snapshot unavailable
|
||||
}
|
||||
|
||||
return {
|
||||
@@ -2496,17 +2582,12 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
|
||||
// Step 1: Get descendant PIDs
|
||||
const descendantMap = new Map<number, number[]>();
|
||||
|
||||
const pgrepOutput = (
|
||||
await execAsync(
|
||||
`for p in ${sessionPids.join(' ')}; do children=$(pgrep -P $p 2>/dev/null | tr '\\n' ','); echo "$p:$children"; done`,
|
||||
{
|
||||
encoding: 'utf-8',
|
||||
timeout: EXEC_TIMEOUT_MS,
|
||||
}
|
||||
)
|
||||
).stdout.trim();
|
||||
// Derived from the ONE snapshot instead of a shell loop that forks a pgrep
|
||||
// per session — the shape that turned into a fork storm under load.
|
||||
const byParent = this.childrenByParent();
|
||||
const childLines = sessionPids.map((p) => `${p}:${(byParent.get(p) ?? []).join(',')}`).join('\n');
|
||||
|
||||
for (const line of pgrepOutput.split('\n')) {
|
||||
for (const line of childLines.split('\n')) {
|
||||
const [pidStr, childrenStr] = line.split(':');
|
||||
const sessionPid = parseInt(pidStr, 10);
|
||||
if (!Number.isNaN(sessionPid)) {
|
||||
|
||||
@@ -28,7 +28,7 @@ let _claudeDir: string | null = null;
|
||||
|
||||
/**
|
||||
* Returns true if the Claude CLI binary can be located (via `which` or one of
|
||||
* the common install directories). Mirrors `isGeminiAvailable`/`isOpenCodeAvailable`/
|
||||
* the common install directories). Mirrors `isGeminiAvailable`/`isAntigravityAvailable`/`isOpenCodeAvailable`/
|
||||
* `isCodexAvailable` in the sibling resolvers.
|
||||
*/
|
||||
export function isClaudeAvailable(): boolean {
|
||||
@@ -108,12 +108,100 @@ export function getAugmentedPath(): string {
|
||||
return _augmentedPath;
|
||||
}
|
||||
|
||||
/** Cached `claude --version` result: string = version, null = probed but unavailable, undefined = not probed */
|
||||
let _claudeVersion: string | null | undefined = undefined;
|
||||
/**
|
||||
* Cache state for the `claude --version` probe.
|
||||
*
|
||||
* `version` is only ever set from a SUCCESSFUL probe and then kept for the
|
||||
* process lifetime (the binary can't change under a running server without a
|
||||
* restart). Failures are tracked separately so they expire.
|
||||
*/
|
||||
export interface ClaudeVersionProbeState {
|
||||
/** Successful probe result; `undefined` until one succeeds. */
|
||||
version?: string;
|
||||
/** Consecutive failed probes (drives the retry backoff). */
|
||||
failures: number;
|
||||
/** Timestamp of the most recent failed probe. */
|
||||
lastFailureAt: number;
|
||||
}
|
||||
|
||||
/** First retry window after a failed probe. */
|
||||
const VERSION_PROBE_BASE_RETRY_MS = 60_000;
|
||||
/** Ceiling for the doubling backoff, so a permanently missing binary settles down. */
|
||||
const VERSION_PROBE_MAX_RETRY_MS = 15 * 60_000;
|
||||
|
||||
/**
|
||||
* How long to wait before re-probing after `failures` consecutive failures:
|
||||
* 1min, 2min, 4min… capped at 15min. Exported for tests.
|
||||
*/
|
||||
export function claudeVersionRetryDelayMs(failures: number): number {
|
||||
if (failures <= 0) return 0;
|
||||
return Math.min(VERSION_PROBE_BASE_RETRY_MS * 2 ** (failures - 1), VERSION_PROBE_MAX_RETRY_MS);
|
||||
}
|
||||
|
||||
/**
|
||||
* Cache policy for the version probe, pure apart from the `state` it mutates
|
||||
* and the injected `probe` (exported so tests can drive it with a fake clock).
|
||||
*
|
||||
* Success is cached forever; FAILURE is not. That asymmetry is the fix for a
|
||||
* real shipped bug: the old cache stored `null` on any exception and guarded on
|
||||
* `!== undefined`, so a single failed probe — a 5s `EXEC_TIMEOUT_MS` timeout, a
|
||||
* PATH-starved systemd/launchd environment, a transient fs hiccup — at the FIRST
|
||||
* Claude session start left `cliVersion` undefined for EVERY Claude session
|
||||
* until the server restarted. An undefined `cliVersion` silently disables
|
||||
* wheel-forwarding to Claude's own transcript (`_shouldForwardWheelToApp`),
|
||||
* which is the only route to history in repaint mode: a dead wheel on every
|
||||
* device at once, matching the issue #205 retest reports.
|
||||
*
|
||||
* Retries back off so a genuinely absent binary still can't spawn a probe per
|
||||
* session start.
|
||||
*/
|
||||
export function resolveClaudeCliVersion(
|
||||
state: ClaudeVersionProbeState,
|
||||
now: number,
|
||||
probe: () => string | null
|
||||
): string | null {
|
||||
if (state.version !== undefined) return state.version;
|
||||
if (state.failures > 0 && now - state.lastFailureAt < claudeVersionRetryDelayMs(state.failures)) return null;
|
||||
|
||||
let version: string | null = null;
|
||||
try {
|
||||
version = probe();
|
||||
} catch {
|
||||
version = null;
|
||||
}
|
||||
|
||||
if (version) {
|
||||
state.version = version;
|
||||
state.failures = 0;
|
||||
state.lastFailureAt = 0;
|
||||
return version;
|
||||
}
|
||||
state.failures += 1;
|
||||
state.lastFailureAt = now;
|
||||
return null;
|
||||
}
|
||||
|
||||
const _claudeVersionState: ClaudeVersionProbeState = { failures: 0, lastFailureAt: 0 };
|
||||
|
||||
/** One `claude --version` run. Throws on spawn/timeout failure. */
|
||||
function probeClaudeCliVersion(): string | null {
|
||||
const dir = findClaudeDir();
|
||||
const bin = dir ? join(dir, 'claude') : 'claude';
|
||||
// execFileSync (no shell) — the resolved path may contain spaces, and there
|
||||
// is no untrusted input, but avoid a shell either way.
|
||||
const out = execFileSync(bin, ['--version'], {
|
||||
encoding: 'utf-8',
|
||||
timeout: EXEC_TIMEOUT_MS,
|
||||
env: { ...process.env, PATH: getAugmentedPath() },
|
||||
});
|
||||
const match = out.match(/(\d+\.\d+\.\d+)/);
|
||||
return match ? match[1] : null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Returns the installed Claude CLI version (e.g. `"2.1.210"`), or null if it
|
||||
* can't be determined. Runs `claude --version` once and caches the result.
|
||||
* can't be determined. Runs `claude --version` at most once per successful
|
||||
* resolution; failed probes retry with backoff (see `resolveClaudeCliVersion`).
|
||||
*
|
||||
* This is a deterministic alternative to scraping the interactive startup
|
||||
* banner (`parseClaudeCodeInfo` in session.ts): newer Claude Code builds don't
|
||||
@@ -122,28 +210,10 @@ let _claudeVersion: string | null | undefined = undefined;
|
||||
* gated on it (e.g. wheel-forwarding to Claude's transcript — issue #154).
|
||||
*/
|
||||
export function getClaudeCliVersion(): string | null {
|
||||
if (_claudeVersion !== undefined) return _claudeVersion;
|
||||
// Keep the test suite hermetic — never spawn a real `claude` subprocess under
|
||||
// vitest (matches IS_TEST_MODE in tmux-manager). Tests that need a version set
|
||||
// it on the session directly.
|
||||
if (process.env.VITEST) {
|
||||
_claudeVersion = null;
|
||||
return _claudeVersion;
|
||||
}
|
||||
try {
|
||||
const dir = findClaudeDir();
|
||||
const bin = dir ? join(dir, 'claude') : 'claude';
|
||||
// execFileSync (no shell) — the resolved path may contain spaces, and there
|
||||
// is no untrusted input, but avoid a shell either way.
|
||||
const out = execFileSync(bin, ['--version'], {
|
||||
encoding: 'utf-8',
|
||||
timeout: EXEC_TIMEOUT_MS,
|
||||
env: { ...process.env, PATH: getAugmentedPath() },
|
||||
});
|
||||
const match = out.match(/(\d+\.\d+\.\d+)/);
|
||||
_claudeVersion = match ? match[1] : null;
|
||||
} catch {
|
||||
_claudeVersion = null;
|
||||
}
|
||||
return _claudeVersion;
|
||||
// it on the session directly. Deliberately does NOT touch the cache state:
|
||||
// recording a phantom failure here would be the very poisoning this fixes.
|
||||
if (process.env.VITEST) return null;
|
||||
return resolveClaudeCliVersion(_claudeVersionState, Date.now(), probeClaudeCliVersion);
|
||||
}
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
/**
|
||||
* @fileoverview Resolve the `cloudflared` binary across common install paths.
|
||||
*
|
||||
* Mirrors the CLI resolvers (gemini-cli-resolver.ts et al), for the same reason
|
||||
* Mirrors the CLI resolvers (antigravity-cli-resolver.ts et al), for the same reason
|
||||
* they exist: the welcome screen should not offer a button whose only possible
|
||||
* outcome is an error toast.
|
||||
*
|
||||
|
||||
+152
-34
@@ -511,7 +511,18 @@ class CodemanApp {
|
||||
this._initGeneration = 0; // dedup concurrent handleInit calls
|
||||
this._initFallbackTimer = null; // fallback timer if SSE init doesn't arrive
|
||||
this._selectGeneration = 0; // cancel stale selectSession loads
|
||||
this._initialFullBufferLoad = true; // first buffer load after a page load fetches full tmux scrollback (COD-47)
|
||||
// Sessions whose full tmux scrollback has already been replayed this page load
|
||||
// (COD-47). Tracked PER SESSION rather than as a single "first load" flag: the
|
||||
// flag was consumed by whichever session auto-selected at page load, so every
|
||||
// OTHER tab started life with one visible frame of history (issue #205).
|
||||
this._fullHistoryLoaded = new Set();
|
||||
// Cooldown per session for the scroll-to-top "load more history" re-pull.
|
||||
this._fullHistoryRepullAt = new Map(); // Map<sessionId, timestamp>
|
||||
this._fullHistoryRepullInFlight = false;
|
||||
// Sessions whose last re-pull came back THINNER than the live buffer (a
|
||||
// repaint-mode CLI pane, where tmux keeps no history of its own). The pull is
|
||||
// refused for those and retried far more slowly — see _maybeRefetchFullHistory.
|
||||
this._fullHistoryRepullUseless = new Set();
|
||||
this.terminalLoadStates = new Map(); // Map<sessionId, { generation, phase }>
|
||||
this.respawnStatus = {};
|
||||
this.respawnTimers = {}; // Track timed respawn timers
|
||||
@@ -816,6 +827,7 @@ class CodemanApp {
|
||||
this.applyLocalization();
|
||||
this.applyTabWrapSettings();
|
||||
this.applyMonitorVisibility();
|
||||
this._setupTabMiddleClickClose();
|
||||
// Must run before the first session:created can arrive: markSessionTabEntering()
|
||||
// ignores ids until this sets up its state, which is what keeps the tabs
|
||||
// restored on page load from animating.
|
||||
@@ -1923,6 +1935,37 @@ class CodemanApp {
|
||||
}
|
||||
}
|
||||
|
||||
/** Build one response-viewer message so the brief and full views share markup and CSS. */
|
||||
_buildResponseViewerMessage(text, role, agentLabel) {
|
||||
const div = document.createElement('div');
|
||||
const isUser = role === 'user';
|
||||
div.className = 'rv-message ' + (isUser ? 'rv-msg-user' : 'rv-msg-assistant');
|
||||
|
||||
const roleBadge = document.createElement('div');
|
||||
roleBadge.className = 'rv-role ' + (isUser ? 'rv-role-user' : 'rv-role-assistant');
|
||||
roleBadge.textContent = isUser ? 'You' : agentLabel;
|
||||
div.appendChild(roleBadge);
|
||||
|
||||
const renderedText = document.createElement('div');
|
||||
renderedText.className = 'rv-text';
|
||||
renderedText.innerHTML = this._renderMarkdown(text);
|
||||
div.appendChild(renderedText);
|
||||
return div;
|
||||
}
|
||||
|
||||
_getResponseViewerAgentLabel() {
|
||||
const mode = this.sessions.get(this.activeSessionId)?.mode;
|
||||
return mode === 'codex'
|
||||
? 'Codex'
|
||||
: mode === 'gemini'
|
||||
? 'Gemini'
|
||||
: mode === 'antigravity'
|
||||
? 'Antigravity'
|
||||
: mode === 'opencode'
|
||||
? 'OpenCode'
|
||||
: 'Claude';
|
||||
}
|
||||
|
||||
async toggleResponseViewer() {
|
||||
const viewer = document.getElementById('responseViewer');
|
||||
const backdrop = document.getElementById('responseViewerBackdrop');
|
||||
@@ -1945,7 +1988,7 @@ class CodemanApp {
|
||||
// Source 2: Terminal buffer fallback — strip ANSI, drop Claude CLI chrome.
|
||||
// Claude + shell only: _cleanTerminalBuffer knows Claude CLI's output, and
|
||||
// shell sessions have no transcript source at all; for TUI modes
|
||||
// (codex/opencode/gemini) it yields repaint garbage, so a clear
|
||||
// (codex/opencode/gemini/antigravity) it yields repaint garbage, so a clear
|
||||
// placeholder beats a messy screen dump there.
|
||||
const sessionMode = this.sessions.get(this.activeSessionId)?.mode || 'claude';
|
||||
if (!lastResponse && (sessionMode === 'claude' || sessionMode === 'shell')) {
|
||||
@@ -1958,7 +2001,11 @@ class CodemanApp {
|
||||
|
||||
const body = document.getElementById('responseViewerBody');
|
||||
if (lastResponse) {
|
||||
body.innerHTML = this._renderMarkdown(lastResponse);
|
||||
// Keep the brief view inside the same message wrapper as the full
|
||||
// conversation view. The wrapper supplies the card, role badge and
|
||||
// descendant markdown styles that direct body children do not get.
|
||||
body.innerHTML = '';
|
||||
body.appendChild(this._buildResponseViewerMessage(lastResponse, 'assistant', this._getResponseViewerAgentLabel()));
|
||||
this._bindResponseViewerInteractions(body);
|
||||
} else {
|
||||
body.textContent =
|
||||
@@ -1998,26 +2045,10 @@ class CodemanApp {
|
||||
}
|
||||
|
||||
// Render conversation thread
|
||||
const mode = this.sessions.get(this.activeSessionId)?.mode;
|
||||
const agentLabel =
|
||||
mode === 'codex' ? 'Codex' : mode === 'gemini' ? 'Gemini' : mode === 'antigravity' ? 'Antigravity' : mode === 'opencode' ? 'OpenCode' : 'Claude';
|
||||
const agentLabel = this._getResponseViewerAgentLabel();
|
||||
body.innerHTML = '';
|
||||
for (const msg of messages) {
|
||||
const div = document.createElement('div');
|
||||
const isUser = msg.role === 'user';
|
||||
div.className = 'rv-message ' + (isUser ? 'rv-msg-user' : 'rv-msg-assistant');
|
||||
|
||||
const role = document.createElement('div');
|
||||
role.className = 'rv-role ' + (isUser ? 'rv-role-user' : 'rv-role-assistant');
|
||||
role.textContent = isUser ? 'You' : agentLabel;
|
||||
div.appendChild(role);
|
||||
|
||||
const text = document.createElement('div');
|
||||
text.className = 'rv-text';
|
||||
text.innerHTML = this._renderMarkdown(msg.text);
|
||||
div.appendChild(text);
|
||||
|
||||
body.appendChild(div);
|
||||
body.appendChild(this._buildResponseViewerMessage(msg.text, msg.role, agentLabel));
|
||||
}
|
||||
this._bindResponseViewerInteractions(body);
|
||||
|
||||
@@ -3459,11 +3490,12 @@ class CodemanApp {
|
||||
}
|
||||
}
|
||||
} else if (minimizedCount > 0 && !subagentBadgeEl) {
|
||||
// Need to add badge - insert before gear icon
|
||||
// Need to add badge - insert before the action-icon overlay so the
|
||||
// badge stays a direct child of the tab (outside .tab-actions)
|
||||
const badgeHtml = this.renderSubagentTabBadge(id, minimizedAgents);
|
||||
const gearEl = tab.querySelector('.tab-gear');
|
||||
if (gearEl) {
|
||||
gearEl.insertAdjacentHTML('beforebegin', badgeHtml);
|
||||
const actionsEl = tab.querySelector('.tab-actions');
|
||||
if (actionsEl) {
|
||||
actionsEl.insertAdjacentHTML('beforebegin', badgeHtml);
|
||||
}
|
||||
} else if (minimizedCount === 0 && subagentBadgeEl) {
|
||||
// Count went to 0 - remove badge
|
||||
@@ -3517,6 +3549,25 @@ class CodemanApp {
|
||||
container.classList.toggle('tabs-auto-wrap', shouldWrap);
|
||||
}
|
||||
|
||||
// Middle-click closes a tab, mirroring browser tab strips. Session tabs go
|
||||
// through requestCloseSession (the same confirm modal as the x button), web
|
||||
// tabs through closeWebviewTab (same as theirs). Delegated on the container:
|
||||
// tabs are re-rendered wholesale, the container is stable.
|
||||
_setupTabMiddleClickClose() {
|
||||
const container = this.$('sessionTabs');
|
||||
if (!container || this._tabAuxClickBound) return;
|
||||
this._tabAuxClickBound = true;
|
||||
container.addEventListener('auxclick', (e) => {
|
||||
if (e.button !== 1) return;
|
||||
const tab = e.target.closest?.('.session-tab');
|
||||
if (!tab) return;
|
||||
e.preventDefault();
|
||||
e.stopPropagation();
|
||||
if (tab.dataset.id) this.requestCloseSession(tab.dataset.id);
|
||||
else if (tab.dataset.webviewId) this.closeWebviewTab?.(tab.dataset.webviewId);
|
||||
});
|
||||
}
|
||||
|
||||
_fullRenderSessionTabs() {
|
||||
if (this._inlineRenameActive) return;
|
||||
const container = this.$('sessionTabs');
|
||||
@@ -3581,9 +3632,7 @@ class CodemanApp {
|
||||
${hasRunningTasks ? `<span class="tab-badge" onclick="event.stopPropagation(); app.toggleTaskPanel()" aria-label="${taskStats.running} running tasks">${taskStats.running}</span>` : ''}
|
||||
${subagentBadge}
|
||||
${ultracodeBadge}
|
||||
<span class="tab-gear" onclick="event.stopPropagation(); app.openSessionOptions(${escapeHtml(JSON.stringify(id))})" title="Session options" aria-label="Session options" tabindex="0">⚙</span>
|
||||
<span class="tab-detach" onclick="event.stopPropagation(); app.detachSession(${escapeHtml(JSON.stringify(id))})" title="Open in a new window" aria-label="Open session in a new window" tabindex="0">⧉</span>
|
||||
<span class="tab-close" onclick="event.stopPropagation(); app.requestCloseSession(${escapeHtml(JSON.stringify(id))})" title="Close session" aria-label="Close session" tabindex="0">×</span>
|
||||
<span class="tab-actions"><span class="tab-gear" onclick="event.stopPropagation(); app.openSessionOptions(${escapeHtml(JSON.stringify(id))})" title="Session options" aria-label="Session options" tabindex="0">⚙</span><span class="tab-detach" onclick="event.stopPropagation(); app.detachSession(${escapeHtml(JSON.stringify(id))})" title="Open in a new window" aria-label="Open session in a new window" tabindex="0">⧉</span><span class="tab-close" onclick="event.stopPropagation(); app.requestCloseSession(${escapeHtml(JSON.stringify(id))})" title="Close session" aria-label="Close session" tabindex="0">×</span></span>
|
||||
</div>`);
|
||||
_tabIdx++;
|
||||
}
|
||||
@@ -4079,6 +4128,72 @@ class CodemanApp {
|
||||
this.terminal.write('\x1b[3J\x1b[H\x1b[2J');
|
||||
}
|
||||
|
||||
/**
|
||||
* "Load more history": re-pull the whole tmux scrollback when the user scrolls up
|
||||
* while already at the top of what the browser has.
|
||||
*
|
||||
* xterm's buffer is only ever a WINDOW onto tmux's real history, and two things
|
||||
* shrink it. tmux repaints the pane rectangle instead of emitting linefeeds
|
||||
* whenever output outpaces its flush interval, which OVERWRITES already-rendered
|
||||
* scrollback rather than pushing rows into it (measured: a 60-line burst added 1
|
||||
* row and destroyed 34, while the same 60 lines emitted slowly added all 60). And
|
||||
* a tab switch replays only the visible frame. Either way tmux still holds
|
||||
* everything (history-limit 100k by default), so the fix is to go ask for it with
|
||||
* the same `?full=1` capture a page reload uses (issue #205).
|
||||
*
|
||||
* On demand rather than automatic because that capture is unbounded-ish work: at
|
||||
* the default history limit it can be megabytes, which is fine to pay when the
|
||||
* user is explicitly reaching for history and not fine on every tab switch.
|
||||
*
|
||||
* NEVER a downgrade: for a repaint-mode CLI pane tmux keeps no history of its
|
||||
* own, so the capture can be THINNER than what xterm already holds and the
|
||||
* reset+rewrite below would delete history mid-scroll. `_replayWouldShrinkBuffer`
|
||||
* (terminal-ui.js) is the guard, and a session that produced one useless re-pull
|
||||
* gets a much longer cooldown so a hollow pane stops re-fetching megabytes on
|
||||
* every scroll-up (issue #205, round 2).
|
||||
*/
|
||||
async _maybeRefetchFullHistory() {
|
||||
const sessionId = this.activeSessionId;
|
||||
if (!sessionId || this._fullHistoryRepullInFlight || this._isLoadingBuffer) return;
|
||||
if (this.detachedSessions?.has(sessionId)) return;
|
||||
const now = Date.now();
|
||||
// Momentum scrolling fires this dozens of times per flick, and a burst of new
|
||||
// output is the normal reason to want a re-pull, so cooldown rather than latch.
|
||||
const cooldown = this._fullHistoryRepullUseless?.has(sessionId) ? 60000 : 4000;
|
||||
if (now - (this._fullHistoryRepullAt.get(sessionId) || 0) < cooldown) return;
|
||||
this._fullHistoryRepullAt.set(sessionId, now);
|
||||
this._fullHistoryRepullInFlight = true;
|
||||
try {
|
||||
const res = await fetch(`/api/sessions/${sessionId}/terminal?full=1`);
|
||||
const buffer = (await res.json())?.data?.terminalBuffer;
|
||||
// Bail on a tab switch mid-fetch: writing here would paint another session's
|
||||
// history into the terminal the user is now looking at.
|
||||
if (!buffer || this.activeSessionId !== sessionId) return;
|
||||
if (this._replayWouldShrinkBuffer(buffer)) {
|
||||
(this._fullHistoryRepullUseless ||= new Set()).add(sessionId);
|
||||
this._logScrollRouting?.('repull-refused-downgrade');
|
||||
return;
|
||||
}
|
||||
this._fullHistoryRepullUseless?.delete(sessionId);
|
||||
const rowsBefore = this.terminal.buffer.active.length;
|
||||
this._resetTerminalForReplay();
|
||||
await this.chunkedTerminalWrite(buffer, TERMINAL_CHUNK_SIZE, sessionId);
|
||||
if (this.activeSessionId !== sessionId) return;
|
||||
this.terminalBufferCache.set(sessionId, buffer);
|
||||
// Hold the user's place. The replay is a superset that grew the buffer
|
||||
// UPWARD, so what used to be row 0 (what they were looking at) is now `delta`
|
||||
// rows down; scrolling there reveals the recovered history above it instead
|
||||
// of teleporting them to the bottom the way a normal buffer load does.
|
||||
const delta = this.terminal.buffer.active.length - rowsBefore;
|
||||
if (delta > 0) this.terminal.scrollToLine(delta);
|
||||
else this.terminal.scrollToTop();
|
||||
} catch {
|
||||
// Transient (offline, 5xx) — the next scroll-up past the cooldown retries.
|
||||
} finally {
|
||||
this._fullHistoryRepullInFlight = false;
|
||||
}
|
||||
}
|
||||
|
||||
_shouldFocusTerminalForTabSwitch() {
|
||||
if (typeof MobileDetection === 'undefined' || !MobileDetection.isTouchDevice()) {
|
||||
return true;
|
||||
@@ -4262,7 +4377,7 @@ class CodemanApp {
|
||||
// (viewport + scrollback + colors) for an instant first paint. For codex
|
||||
// this is also a correctness fix — its byte-stream replay shows only the
|
||||
// latest TUI frame (the idle welcome banner) because codex doesn't include
|
||||
// earlier conversation in its current redraw. For claude/opencode/gemini
|
||||
// earlier conversation in its current redraw. For claude/opencode/gemini/antigravity
|
||||
// the replay is already complete, so the snapshot is purely a faster,
|
||||
// scroll-preserving first paint before the canonical fetch reconciles.
|
||||
//
|
||||
@@ -4349,11 +4464,14 @@ class CodemanApp {
|
||||
|
||||
this._setTerminalLoadState(sessionId, selectGen, 'fetching');
|
||||
_crashDiag.log('FETCH_START');
|
||||
// The FIRST buffer load after a page load requests the full tmux scrollback
|
||||
// (?full=1, COD-47) so history that scrolled off the server's byte buffer
|
||||
// comes back after a reload. Tab switches keep the fast ?tail= frame path.
|
||||
const useFullHistory = this._initialFullBufferLoad === true;
|
||||
this._initialFullBufferLoad = false;
|
||||
// The first load OF EACH SESSION this page load requests the full tmux
|
||||
// scrollback (?full=1, COD-47) so history that scrolled off the server's byte
|
||||
// buffer comes back. Later switches to an already-replayed session keep the
|
||||
// fast ?tail= frame path, which is why this is a Set and not a flag: the flag
|
||||
// version gave the full replay to the auto-selected tab and one frame of
|
||||
// history to every other one (issue #205).
|
||||
const useFullHistory = !this._fullHistoryLoaded.has(sessionId);
|
||||
if (useFullHistory) this._fullHistoryLoaded.add(sessionId);
|
||||
const res = await fetch(
|
||||
useFullHistory
|
||||
? `/api/sessions/${sessionId}/terminal?full=1`
|
||||
|
||||
@@ -227,6 +227,7 @@
|
||||
'Redraw Terminal Button': '重绘终端按钮',
|
||||
'Tab Bar': '标签栏',
|
||||
'Tall Tabs (Name + Folder)': '双行标签(名称 + 文件夹)',
|
||||
'Pop-out Button on Tabs': '标签页弹出窗口按钮',
|
||||
Panels: '面板',
|
||||
Monitor: '监视器',
|
||||
'Project Insights': '项目洞察',
|
||||
|
||||
@@ -323,6 +323,10 @@
|
||||
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
|
||||
Run OpenCode
|
||||
</button>
|
||||
<button class="welcome-btn welcome-btn-antigravity" id="welcomeAntigravityBtn" style="display: none;" onclick="app.setRunMode('antigravity'); app.runAntigravity()">
|
||||
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
|
||||
Run Antigravity
|
||||
</button>
|
||||
<button class="welcome-btn welcome-btn-gemini" id="welcomeGeminiBtn" style="display: none;" onclick="app.setRunMode('gemini'); app.runGemini()">
|
||||
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
|
||||
Run Gemini
|
||||
@@ -1318,7 +1322,7 @@
|
||||
</div>
|
||||
<!-- Input Section -->
|
||||
<div class="settings-section-header">Input</div>
|
||||
<div class="settings-item settings-item-multiline" title="Scroll the terminal's own local scrollback with a plain mouse wheel / two-finger swipe, instead of forwarding the wheel to the CLI's transcript. Turn on if scrolling back through history doesn't work (e.g. macOS trackpad in Claude sessions). Shift+wheel always reaches local scrollback regardless.">
|
||||
<div class="settings-item settings-item-multiline" title="Scroll the terminal's own local scrollback with a plain mouse wheel / two-finger swipe, instead of forwarding the wheel to the CLI's transcript. Leave OFF for Claude/Codex sessions: those CLIs redraw the screen in place and keep no local scrollback, so the wheel would have almost nothing to scroll. In Claude sessions Codeman then falls back to paging the CLI's own transcript; Codex sessions have no such fallback, so the wheel goes dead. Shift+wheel always reaches local scrollback regardless.">
|
||||
<div class="settings-item-text">
|
||||
<span class="settings-item-label">Wheel Scrolls Local History</span>
|
||||
<span class="settings-item-desc">Plain wheel/trackpad pages the terminal scrollback</span>
|
||||
@@ -1476,6 +1480,13 @@
|
||||
<span class="slider"></span>
|
||||
</label>
|
||||
</div>
|
||||
<div class="settings-item" title="Show the open-in-new-window (pop-out) button when hovering a session tab">
|
||||
<span class="settings-item-label">Pop-out Button on Tabs</span>
|
||||
<label class="switch switch-sm">
|
||||
<input type="checkbox" id="appSettingsShowTabDetachButton">
|
||||
<span class="slider"></span>
|
||||
</label>
|
||||
</div>
|
||||
|
||||
<!-- Panels Section -->
|
||||
<div class="settings-section-header">Panels</div>
|
||||
@@ -2192,7 +2203,7 @@
|
||||
<div class="form-row">
|
||||
<label>Image</label>
|
||||
<input type="text" id="dockerImage" placeholder="codeman/agent:base" autocomplete="off" autocapitalize="off" spellcheck="false">
|
||||
<span class="form-hint">Build it once with <code>node scripts/build-agent-image.mjs</code>. Contains node + claude/codex/gemini + tmux.</span>
|
||||
<span class="form-hint">Build it once with <code>node scripts/build-agent-image.mjs</code>. Contains node + claude/codex/gemini/opencode/agy + tmux.</span>
|
||||
</div>
|
||||
<div class="form-row">
|
||||
<label>Network</label>
|
||||
|
||||
@@ -209,9 +209,13 @@ const MobileDetection = {
|
||||
* Also handles terminal scrolling and toolbar repositioning via visualViewport API.
|
||||
*/
|
||||
const KeyboardHandler = {
|
||||
VIEWPORT_SETTLE_MS: 80,
|
||||
lastViewportHeight: 0,
|
||||
keyboardVisible: false,
|
||||
initialViewportHeight: 0,
|
||||
_viewportSettleTimer: null,
|
||||
_settleScrollToBottom: false,
|
||||
_settlePending: false,
|
||||
|
||||
/** Initialize keyboard handling */
|
||||
init() {
|
||||
@@ -276,6 +280,12 @@ const KeyboardHandler = {
|
||||
window.removeEventListener('scroll', this._windowScrollHandler);
|
||||
this._windowScrollHandler = null;
|
||||
}
|
||||
if (this._viewportSettleTimer) {
|
||||
clearTimeout(this._viewportSettleTimer);
|
||||
this._viewportSettleTimer = null;
|
||||
}
|
||||
this._settleScrollToBottom = false;
|
||||
this._settlePending = false;
|
||||
},
|
||||
|
||||
/** Handle viewport resize (keyboard show/hide) */
|
||||
@@ -313,6 +323,7 @@ const KeyboardHandler = {
|
||||
}
|
||||
|
||||
this.updateLayoutForKeyboard();
|
||||
this._deferViewportSettle();
|
||||
this.lastViewportHeight = currentHeight;
|
||||
},
|
||||
|
||||
@@ -414,32 +425,9 @@ const KeyboardHandler = {
|
||||
// iOS Safari may scroll the document to reveal xterm's hidden textarea.
|
||||
window.scrollTo(0, 0);
|
||||
|
||||
// Refit terminal locally AND send resize to server so Claude Code (Ink)
|
||||
// knows the actual terminal dimensions. Without this, Ink redraws at the
|
||||
// old (larger) row count when the user types, causing content to scroll
|
||||
// off the visible area with each keystroke.
|
||||
// Note: the throttledResize handler still suppresses ongoing resize events
|
||||
// while keyboard is up — this one-shot resize on open/close is sufficient.
|
||||
setTimeout(() => {
|
||||
if (typeof app !== 'undefined' && app.terminal) {
|
||||
if (app.fitAddon)
|
||||
try {
|
||||
app.fitAddon.fit();
|
||||
} catch {}
|
||||
// Eliminate terminal row quantization gap: xterm can only show whole
|
||||
// rows, so leftover pixels create dead space below the last row.
|
||||
// Shrink .main's paddingBottom by the gap so the terminal fills flush
|
||||
// to the accessory bar.
|
||||
this._shrinkPaddingToFit();
|
||||
app.terminal.scrollToBottom();
|
||||
app._syncMobileHelperTextareaToCursor?.();
|
||||
app._localEchoOverlay?.rerender?.();
|
||||
// Send resize to server so PTY dimensions match xterm
|
||||
this._sendTerminalResize();
|
||||
}
|
||||
// Reset again after fit/resize in case layout changes triggered scroll
|
||||
window.scrollTo(0, 0);
|
||||
}, 150);
|
||||
// visualViewport emits multiple heights throughout the OS animation.
|
||||
// Re-schedule on every event and fit only after the final height settles.
|
||||
this._scheduleViewportSettle({ scrollToBottom: true });
|
||||
|
||||
// Reposition subagent windows to stack from bottom (above keyboard)
|
||||
if (typeof app !== 'undefined') app.relayoutMobileSubagentWindows();
|
||||
@@ -454,22 +442,59 @@ const KeyboardHandler = {
|
||||
|
||||
this.resetLayout();
|
||||
|
||||
// Refit terminal, scroll to bottom, and send resize to restore original dimensions
|
||||
setTimeout(() => {
|
||||
if (typeof app !== 'undefined' && app.fitAddon) {
|
||||
try {
|
||||
app.fitAddon.fit();
|
||||
} catch {}
|
||||
if (app.terminal) app.terminal.scrollToBottom();
|
||||
// Send resize to server to restore full terminal size
|
||||
this._sendTerminalResize();
|
||||
}
|
||||
}, 100);
|
||||
this._scheduleViewportSettle({ scrollToBottom: true });
|
||||
|
||||
// Reposition subagent windows to stack from top (below header)
|
||||
if (typeof app !== 'undefined') app.relayoutMobileSubagentWindows();
|
||||
},
|
||||
|
||||
/**
|
||||
* Coalesce the keyboard animation into one final xterm reflow and PTY resize.
|
||||
* Only a real show/hide transition arms the settle work; ongoing viewport
|
||||
* resize events merely push a pending settle back (_deferViewportSettle).
|
||||
* A viewport change that never crosses the show/hide thresholds must not
|
||||
* refit: keyboard detection can miss a fine-grained OS animation entirely
|
||||
* (each step under 150px, with the baseline chasing the animation), and the
|
||||
* container is then mid-animation with no keyboard CSS compensation, so a
|
||||
* fit against it resizes the PTY to transient dims and the SIGWINCH thrash
|
||||
* garbles the transcript.
|
||||
*/
|
||||
_scheduleViewportSettle({ scrollToBottom = false } = {}) {
|
||||
this._settleScrollToBottom = this._settleScrollToBottom || scrollToBottom;
|
||||
this._settlePending = true;
|
||||
this._armViewportSettleTimer();
|
||||
},
|
||||
|
||||
/** Push a pending settle back while the viewport is still animating; no-op otherwise. */
|
||||
_deferViewportSettle() {
|
||||
if (!this._settlePending) return;
|
||||
this._armViewportSettleTimer();
|
||||
},
|
||||
|
||||
_armViewportSettleTimer() {
|
||||
if (this._viewportSettleTimer) clearTimeout(this._viewportSettleTimer);
|
||||
this._viewportSettleTimer = setTimeout(() => {
|
||||
this._viewportSettleTimer = null;
|
||||
this._settlePending = false;
|
||||
const shouldScrollToBottom = this._settleScrollToBottom;
|
||||
this._settleScrollToBottom = false;
|
||||
|
||||
if (typeof app !== 'undefined' && app.terminal) {
|
||||
if (app.fitAddon) {
|
||||
try {
|
||||
app.fitAddon.fit();
|
||||
} catch {}
|
||||
}
|
||||
if (this.keyboardVisible) this._shrinkPaddingToFit();
|
||||
if (shouldScrollToBottom) app.terminal.scrollToBottom();
|
||||
app._syncMobileHelperTextareaToCursor?.();
|
||||
app._localEchoOverlay?.rerender?.();
|
||||
this._sendTerminalResize();
|
||||
}
|
||||
window.scrollTo(0, 0);
|
||||
}, this.VIEWPORT_SETTLE_MS);
|
||||
},
|
||||
|
||||
/** Send current terminal dimensions to the server (one-shot, for keyboard open/close) */
|
||||
_sendTerminalResize() {
|
||||
if (typeof app === 'undefined' || !app.activeSessionId || !app.fitAddon) return;
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
/**
|
||||
* @fileoverview Quick start (case loading, session spawning for Claude/Shell/OpenCode/Codex/Gemini),
|
||||
* @fileoverview Quick start (case loading, session spawning for Claude/Shell/OpenCode/Codex/Gemini/Antigravity),
|
||||
* session options modal (per-session settings, color picker, rename),
|
||||
* session options tabs (Ralph config tab), case settings (CRUD, links),
|
||||
* create case modal, and mobile case picker.
|
||||
|
||||
@@ -363,6 +363,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
document.getElementById('appSettingsCjkInput').checked = settings.cjkInputEnabled ?? defaults.cjkInputEnabled ?? false;
|
||||
document.getElementById('appSettingsExtendedKeyboardBar').checked = settings.extendedKeyboardBar ?? false;
|
||||
document.getElementById('appSettingsTabTwoRows').checked = settings.tabTwoRows ?? defaults.tabTwoRows ?? false;
|
||||
document.getElementById('appSettingsShowTabDetachButton').checked = settings.showTabDetachButton ?? defaults.showTabDetachButton ?? false;
|
||||
// Claude CLI settings
|
||||
const claudeModeSelect = document.getElementById('appSettingsClaudeMode');
|
||||
const allowedToolsRow = document.getElementById('allowedToolsRow');
|
||||
@@ -748,6 +749,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
const buttons = [
|
||||
['welcomeClaudeBtn', 'claude'],
|
||||
['welcomeOpencodeBtn', 'opencode'],
|
||||
['welcomeAntigravityBtn', 'antigravity'],
|
||||
['welcomeGeminiBtn', 'gemini'],
|
||||
// Not a run mode, same reasoning: offering a Cloudflare Tunnel on a box
|
||||
// without cloudflared can only ever produce "cloudflared not found".
|
||||
@@ -1541,6 +1543,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
webglRendererEnabled: document.getElementById('appSettingsWebglRenderer').checked,
|
||||
extendedKeyboardBar: document.getElementById('appSettingsExtendedKeyboardBar').checked,
|
||||
tabTwoRows: document.getElementById('appSettingsTabTwoRows').checked,
|
||||
showTabDetachButton: document.getElementById('appSettingsShowTabDetachButton').checked,
|
||||
skin: document.getElementById('appSettingsSkin').value,
|
||||
// Claude CLI settings
|
||||
claudeMode: document.getElementById('appSettingsClaudeMode').value,
|
||||
@@ -1725,6 +1728,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
showSessionButton: _ssb,
|
||||
showAwayDigestButton: _adb,
|
||||
showCronButton: _crb,
|
||||
showTabDetachButton: _tdb,
|
||||
// Phone-only home surface, and absent from SettingsUpdateSchema (.strict()).
|
||||
mobileOverviewEnabled: _mov,
|
||||
...serverSettings
|
||||
@@ -1998,6 +2002,13 @@ Object.assign(CodemanApp.prototype, {
|
||||
applyHeaderVisibilitySettings() {
|
||||
const settings = this.loadAppSettingsFromStorage();
|
||||
const defaults = this.getDefaultSettings();
|
||||
|
||||
// Tab pop-out (open-in-new-window) button: opt-in (App Settings → Tab Bar,
|
||||
// default OFF, per-device). Mirrored as a class on <html>: styles.css hides
|
||||
// .tab-detach without it (a tab that is already detached keeps its icon as
|
||||
// the re-focus affordance for the popped-out window).
|
||||
const showTabDetach = settings.showTabDetachButton ?? defaults.showTabDetachButton ?? false;
|
||||
document.documentElement.classList.toggle('tabs-show-detach', showTabDetach);
|
||||
const compactHeader = MobileDetection.getDeviceType() !== 'desktop';
|
||||
const showFontControls = compactHeader ? false : (settings.showFontControls ?? defaults.showFontControls ?? false);
|
||||
const showSystemStats = compactHeader ? false : (settings.showSystemStats ?? defaults.showSystemStats ?? true);
|
||||
@@ -2354,6 +2365,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
'language',
|
||||
'terminalWheelLocalScrollback',
|
||||
'showSessionButton', 'showAwayDigestButton', 'showCronButton',
|
||||
'showTabDetachButton',
|
||||
'mobileOverviewEnabled',
|
||||
]);
|
||||
// The plan-usage chip is a PER-DEVICE display setting (desktop default ON,
|
||||
|
||||
@@ -1421,7 +1421,11 @@ html[data-line-anim="packet"] .connection-line.line-enter {
|
||||
transition: opacity 0.05s ease-out, width 0.05s ease-out, padding 0.05s ease-out;
|
||||
}
|
||||
|
||||
.session-tab:hover .tab-close {
|
||||
/* Icons expand on the ACTIVE tab only (in flow): selection is a deliberate
|
||||
click, so the width change never happens while aiming at a tab, background
|
||||
tabs keep their full title on hover, and a stray click can only switch.
|
||||
Middle-click closes any tab (_setupTabMiddleClickClose in app.js). */
|
||||
.session-tab.active .tab-close {
|
||||
opacity: 1;
|
||||
width: auto;
|
||||
padding: 0.15rem 0.35rem;
|
||||
@@ -1983,7 +1987,7 @@ html[data-line-anim="packet"] .connection-line.line-enter {
|
||||
transition: opacity 0.15s, width 0.15s, padding 0.15s, transform 0.2s;
|
||||
}
|
||||
|
||||
.session-tab:hover .tab-gear {
|
||||
.session-tab.active .tab-gear {
|
||||
opacity: 1;
|
||||
width: auto;
|
||||
padding: 0 0.3rem;
|
||||
@@ -2008,7 +2012,7 @@ html[data-line-anim="packet"] .connection-line.line-enter {
|
||||
overflow: hidden;
|
||||
transition: opacity 0.15s, width 0.15s, padding 0.15s;
|
||||
}
|
||||
.session-tab:hover .tab-detach {
|
||||
.session-tab.active .tab-detach {
|
||||
opacity: 1;
|
||||
width: auto;
|
||||
padding: 0 0.3rem;
|
||||
@@ -2047,6 +2051,29 @@ html[data-line-anim="packet"] .connection-line.line-enter {
|
||||
display: inline-flex;
|
||||
}
|
||||
|
||||
/* ===== Tab action icons: active tab only =================================
|
||||
All three per-tab icons live in a .tab-actions wrapper (in flow; it adds
|
||||
no width of its own while the children keep the width:0 collapse above).
|
||||
They expand only on the ACTIVE tab: selection is a deliberate click, so
|
||||
the tab-strip geometry never shifts while the pointer is aiming, hovering
|
||||
a background tab changes nothing (full title stays readable), and a stray
|
||||
click can only switch sessions. Middle-click closes any tab. The phone
|
||||
layout in mobile.css follows the same active-only pattern with its own
|
||||
sizing; the tablet touch fallback there keeps icons always visible. */
|
||||
|
||||
.session-tab .tab-actions {
|
||||
display: flex;
|
||||
align-items: center;
|
||||
}
|
||||
|
||||
/* Pop-out button is opt-in (App Settings → Tab Bar, default off; per-device).
|
||||
settings-ui.js mirrors the setting as the tabs-show-detach class on <html>.
|
||||
A tab that is ALREADY detached keeps its icon regardless: it is the
|
||||
re-focus affordance for the popped-out window. */
|
||||
html:not(.tabs-show-detach) .session-tab:not(.detached) .tab-detach {
|
||||
display: none;
|
||||
}
|
||||
|
||||
/* ===== Solo (detached single-session) window chrome ===================== */
|
||||
body.solo-mode .session-tabs,
|
||||
body.solo-mode .header-system-stats,
|
||||
@@ -3305,6 +3332,23 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
|
||||
transform: translateY(-1px);
|
||||
}
|
||||
|
||||
/* Antigravity: cyan identity, matching .btn-toolbar.btn-run.mode-antigravity and
|
||||
.run-mode-dot.antigravity so the welcome action reads as the same backend. */
|
||||
.welcome-btn-antigravity {
|
||||
background: linear-gradient(135deg, #0b2b33 0%, #0e7490 55%, #0891b2 100%);
|
||||
border-color: rgba(34, 211, 238, 0.4);
|
||||
color: #cffafe;
|
||||
box-shadow: 0 2px 8px rgba(34, 211, 238, 0.16), inset 0 1px 0 rgba(255, 255, 255, 0.06);
|
||||
}
|
||||
|
||||
.welcome-btn-antigravity:hover {
|
||||
background: linear-gradient(135deg, #124450 0%, #0891b2 55%, #06b6d4 100%);
|
||||
box-shadow: 0 4px 20px rgba(34, 211, 238, 0.3), 0 0 40px rgba(8, 145, 178, 0.12), inset 0 1px 0 rgba(255, 255, 255, 0.08);
|
||||
border-color: rgba(103, 232, 249, 0.5);
|
||||
color: #ecfeff;
|
||||
transform: translateY(-1px);
|
||||
}
|
||||
|
||||
.welcome-btn-gemini {
|
||||
background: linear-gradient(135deg, #10243f 0%, #174ea6 55%, #4f46e5 100%);
|
||||
border-color: rgba(96, 165, 250, 0.4);
|
||||
|
||||
+486
-44
@@ -23,6 +23,44 @@
|
||||
// short window, only the app's synthetic tap-to-position mouse event should
|
||||
// reach xterm.
|
||||
const TOUCH_COMPAT_MOUSE_SUPPRESS_MS = 450;
|
||||
// Escape sequences occupy no terminal cells, so they must come out before a
|
||||
// captured line's WIDTH can be measured (_estimateReplayRows). Covers OSC,
|
||||
// CSI, charset designators and the short escapes tmux emits; deliberately
|
||||
// approximate — this feeds a size comparison, not a renderer.
|
||||
// eslint-disable-next-line no-control-regex
|
||||
const REPLAY_ESCAPE_RE =
|
||||
/\x1b\][^\x07\x1b]*(?:\x07|\x1b\\)|\x1b\[[0-9;?<>=!]*[ -/]*[@-~]|\x1b[()#][0-9A-Za-z]|\x1b[=>78M]/g;
|
||||
// PageUp / PageDown as xterm.js encodes them. Used as the LAST-RESORT scroll
|
||||
// gesture for a repaint-mode CLI whose local buffer holds no scrollback
|
||||
// (_maybePageCliTranscript).
|
||||
const KEY_PAGE_UP = '\x1b[5~';
|
||||
const KEY_PAGE_DOWN = '\x1b[6~';
|
||||
// Wheel/touch travel (in lines) that adds up to one PageUp/PageDown. Half a
|
||||
// screen rather than a full one: the page key always jumps a whole screen, so
|
||||
// a 1:1 mapping made the fallback feel unreachably slow with a discrete mouse
|
||||
// wheel (Firefox reports 3 lines a notch → 12 notches per page). Overshooting
|
||||
// the finger is the right trade against a gesture that otherwise does nothing.
|
||||
const PAGE_KEY_SCREEN_FRACTION = 0.5;
|
||||
// Bound on page keys emitted from one gesture batch, mirroring the SGR tick
|
||||
// cap: a fling must not build a backlog that keeps paging after it stops.
|
||||
const PAGE_KEY_MAX_PER_BATCH = 3;
|
||||
// Composer navigation keys as xterm.js encodes user keystrokes: plain and
|
||||
// modified arrows (CSI A-D, CSI 1;mA-D, SS3 A-D), Home/End (CSI H/F, SS3
|
||||
// H/F, CSI 1~/4~), Insert/Delete/PgUp/PgDn (CSI 2~/3~/5~/6~, optional
|
||||
// modifier). Deliberately EXCLUDES terminal query responses that also
|
||||
// arrive via onData (DA `\x1b[?1;2c`, CPR `\x1b[12;34R`) and function
|
||||
// keys, so only genuine cursor/editing keys trigger the local-echo flush.
|
||||
// eslint-disable-next-line no-control-regex
|
||||
const COMPOSER_NAV_KEY_PATTERN = /^\x1b(?:\[(?:[ABCDHF]|1;[2-8][ABCDHF]|[1-8](?:;[2-8])?~)|O[ABCDHF])$/;
|
||||
// Prefix xterm.js puts on terminal.paste() payloads while the application
|
||||
// has bracketed-paste mode (DECSET 2004) enabled. Codex, Claude Code and
|
||||
// tmux all enable it, so browser pastes arrive as one onData chunk of
|
||||
// `\x1b[200~<text>\x1b[201~`.
|
||||
const BRACKETED_PASTE_START = '\x1b[200~';
|
||||
|
||||
function isComposerNavKey(data) {
|
||||
return COMPOSER_NAV_KEY_PATTERN.test(data);
|
||||
}
|
||||
|
||||
function isTerminalQueryResponse(data) {
|
||||
return TERMINAL_QUERY_RESPONSE_PATTERN.test(data) || TERMINAL_OSC_RESPONSE_PATTERN.test(data);
|
||||
@@ -60,8 +98,15 @@
|
||||
global.CodemanTerminalInput = {
|
||||
isTerminalQueryResponse,
|
||||
shouldSuppressTerminalQueryResponse,
|
||||
isComposerNavKey,
|
||||
BRACKETED_PASTE_START,
|
||||
USER_SCROLL_STICKY_SUPPRESS_MS,
|
||||
TOUCH_COMPAT_MOUSE_SUPPRESS_MS,
|
||||
REPLAY_ESCAPE_RE,
|
||||
KEY_PAGE_UP,
|
||||
KEY_PAGE_DOWN,
|
||||
PAGE_KEY_SCREEN_FRACTION,
|
||||
PAGE_KEY_MAX_PER_BATCH,
|
||||
};
|
||||
global.CODEMAN_XTERM_THEMES = CODEMAN_XTERM_THEMES;
|
||||
global.codemanCurrentXtermTheme = currentXtermTheme;
|
||||
@@ -296,7 +341,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
|
||||
// WebGL renderer for GPU-accelerated terminal rendering.
|
||||
// Previously caused "page unresponsive" crashes from synchronous GPU stalls,
|
||||
// but the 48KB/frame flush cap in flushPendingWrites() now prevents
|
||||
// but the mode-aware 32/64KB frame cap in flushPendingWrites() now prevents
|
||||
// oversized terminal.write() calls that triggered the stalls.
|
||||
// Disable with ?nowebgl URL param if GPU issues return.
|
||||
// Auto-fallback: _initWebGL installs a long-task watchdog that disables
|
||||
@@ -415,31 +460,79 @@ Object.assign(CodemanApp.prototype, {
|
||||
// ignores wheel reports); older versions DO capture wheel as option
|
||||
// navigation, so they keep the local wheel.
|
||||
// Shift+wheel always scrolls xterm's local scrollback (Codeman's restored
|
||||
// history lives there), and once the viewport left the bottom the wheel
|
||||
// stays local until the user scrolls back down — so both scrollbacks stay
|
||||
// reachable without a mode switch.
|
||||
// history lives there); the plain wheel stays on the CLI's transcript for
|
||||
// those modes regardless of scroll position, so the CLI's input box never
|
||||
// slides off the screen (see _shouldForwardWheelToApp).
|
||||
//
|
||||
// CAPTURE phase, deliberately, and Codeman owns the scroll. xterm's
|
||||
// viewport is a vscode-style ScrollableElement that consumes wheel events
|
||||
// itself (preventDefault + stopPropagation) whenever it believes a
|
||||
// scrollbar exists, does NOT consult attachCustomWheelEventHandler, and —
|
||||
// measured on the live instance — goes DEAF after terminal.reset(): a tab
|
||||
// switch or full-history replay leaves its scroll dimensions stale, after
|
||||
// which wheel events neither scroll nor propagate reliably. A bubble-phase
|
||||
// listener here therefore never fired once local scrollback existed
|
||||
// (measured: _shouldForwardWheelToApp call count stayed 0 while xterm
|
||||
// scrolled), and after a tab switch NOTHING scrolled at all — the "input
|
||||
// box scrolls up then it fights", "works at first, breaks after a tab
|
||||
// switch" reports on #205.
|
||||
//
|
||||
// So: capture runs ancestors-first; this handler sees every wheel first
|
||||
// and stops propagation, keeping xterm's scroller out of it entirely.
|
||||
// Local scrolling goes through terminal.scrollLines() — buffer-level, so
|
||||
// it keeps working after resets — with our own deltaMode normalization
|
||||
// (_wheelScrollLines) covering Firefox's line-unit wheels. Two cases still
|
||||
// belong to xterm and are passed through untouched:
|
||||
// - mouseTrackingMode active: xterm's own encoder forwards the wheel to
|
||||
// the PTY (htop/vim with mouse on in a shell pane);
|
||||
// - alternate buffer (direct-PTY fallback running vim/less): xterm's
|
||||
// alt-scroll handling converts the wheel to cursor keys, which is what
|
||||
// those apps expect.
|
||||
container.addEventListener(
|
||||
'wheel',
|
||||
(ev) => {
|
||||
const trackingMode = this.terminal?.modes?.mouseTrackingMode;
|
||||
if (trackingMode && trackingMode !== 'none') return;
|
||||
if (this.terminal?.buffer?.active?.type === 'alternate') return;
|
||||
ev.preventDefault();
|
||||
const lines = this._wheelScrollLines(ev);
|
||||
ev.stopPropagation();
|
||||
if (this._shouldForwardWheelToApp(ev)) {
|
||||
this._sendSyntheticSgrWheel(ev.clientX, ev.clientY, lines);
|
||||
this._logScrollRouting('forward-sgr');
|
||||
this._forwardScrollToApp(ev.clientX, ev.clientY, this._wheelScrollLines(ev));
|
||||
return;
|
||||
}
|
||||
// Local scrolling accumulates FRACTIONAL lines: a macOS trackpad emits
|
||||
// a stream of tiny pixel deltas, and rounding each one to a whole line
|
||||
// (the ±1 fallback) made slow drags scroll faster than the finger.
|
||||
const lines = this._wheelScrollLinesFloat(ev);
|
||||
// …unless there is no local scrollback to scroll, in which case page the
|
||||
// CLI's own transcript instead of doing nothing (_maybePageCliTranscript).
|
||||
if (this._maybePageCliTranscript(ev, lines)) return;
|
||||
this._logScrollRouting('local-scrollback');
|
||||
this._noteTerminalUserScroll(lines);
|
||||
this.terminal.scrollLines(lines);
|
||||
this._smoothScrollBy(lines);
|
||||
},
|
||||
{ passive: false }
|
||||
{ passive: false, capture: true }
|
||||
);
|
||||
|
||||
// Touch scrolling — use terminal.scrollLines() for all devices.
|
||||
// xterm.js DOM renderer doesn't populate xterm-viewport's scroll area,
|
||||
// so native CSS scrolling (overflow-y: scroll + touch-action: pan-y)
|
||||
// has nothing to scroll. Instead, convert touch deltas into scrollLines()
|
||||
// calls, matching the wheel handler above.
|
||||
// calls, matching the wheel handler above, including the forwarding
|
||||
// branch: for the sessions whose wheel goes to the CLI's own transcript
|
||||
// (_shouldForwardWheelToApp), a touch drag must go there too, or every
|
||||
// phone/tablet swipe scrolls the local buffer of stale repaint frames and
|
||||
// drags the CLI's pinned input box off the screen (issue #205's mobile
|
||||
// half). Same gate, so Shift has no touch analog but the local-scrollback
|
||||
// opt-out setting and the CLI-version gate apply to touch exactly as they
|
||||
// do to the wheel — including the PageUp/PageDown fallback the wheel uses
|
||||
// when that gate is false and there is no local scrollback to scroll
|
||||
// (_maybePageCliTranscript), which is what keeps a swipe from being a
|
||||
// complete no-op on a phone.
|
||||
{
|
||||
const cellHeight = () => this.terminal._core?._renderService?.dimensions?.css?.cell?.height || 13;
|
||||
let touchLastX = 0;
|
||||
let touchLastY = 0;
|
||||
let velocity = 0;
|
||||
let lastTime = 0;
|
||||
@@ -453,7 +546,16 @@ Object.assign(CodemanApp.prototype, {
|
||||
if (!isTouching && Math.abs(velocity) > 0.3) {
|
||||
// Momentum phase — convert pixel velocity to lines
|
||||
const lines = Math.round(velocity / cellHeight());
|
||||
if (lines !== 0) this.terminal.scrollLines(lines);
|
||||
if (lines !== 0) {
|
||||
if (this._shouldForwardWheelToApp({ shiftKey: false })) {
|
||||
// Flick momentum keeps feeding the CLI's transcript from the last
|
||||
// touch point; the 40ms coalescer batches the per-frame reports.
|
||||
this._forwardScrollToApp(touchLastX, touchLastY, lines);
|
||||
} else if (!this._maybePageCliTranscript({ shiftKey: false }, lines)) {
|
||||
this.terminal.scrollLines(lines);
|
||||
this._maybeLoadMoreHistoryOnScroll(lines);
|
||||
}
|
||||
}
|
||||
velocity *= 0.92;
|
||||
scrollFrame = requestAnimationFrame(scrollLoop);
|
||||
} else if (!isTouching) {
|
||||
@@ -474,6 +576,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
'touchstart',
|
||||
(ev) => {
|
||||
if (ev.touches.length === 1) {
|
||||
touchLastX = ev.touches[0].clientX;
|
||||
touchLastY = ev.touches[0].clientY;
|
||||
touchStartY = touchLastY;
|
||||
velocity = 0;
|
||||
@@ -509,13 +612,21 @@ Object.assign(CodemanApp.prototype, {
|
||||
const delta = touchLastY - touchY; // positive = scroll down
|
||||
pixelAccum += delta;
|
||||
velocity = delta * 1.2;
|
||||
touchLastX = ev.touches[0].clientX;
|
||||
touchLastY = touchY;
|
||||
// Convert accumulated pixels to whole lines
|
||||
const ch = cellHeight();
|
||||
const lines = Math.trunc(pixelAccum / ch);
|
||||
if (lines !== 0) {
|
||||
this._noteTerminalUserScroll(lines);
|
||||
this.terminal.scrollLines(lines);
|
||||
if (this._shouldForwardWheelToApp({ shiftKey: false })) {
|
||||
this._logScrollRouting('forward-sgr');
|
||||
this._forwardScrollToApp(touchLastX, touchLastY, lines);
|
||||
} else if (!this._maybePageCliTranscript({ shiftKey: false }, lines)) {
|
||||
this._logScrollRouting('local-scrollback');
|
||||
this._noteTerminalUserScroll(lines);
|
||||
this.terminal.scrollLines(lines);
|
||||
this._maybeLoadMoreHistoryOnScroll(lines);
|
||||
}
|
||||
pixelAccum -= lines * ch;
|
||||
}
|
||||
}
|
||||
@@ -776,11 +887,23 @@ Object.assign(CodemanApp.prototype, {
|
||||
}
|
||||
this._lastTerminalData = { data, time: performance.now() };
|
||||
|
||||
// ── Local Echo Pass-through ──
|
||||
// After a composer nav key (arrow/Home/End/Delete) the real cursor may
|
||||
// sit mid-text, where the overlay's append-only buffering would corrupt
|
||||
// both the preview and the submitted text. Such sessions are handed
|
||||
// back to plain PTY echo until Enter or Ctrl+C submits/cancels the
|
||||
// composer line (see the nav-key branch below).
|
||||
const echoPassthrough =
|
||||
this._localEchoEnabled && this._echoPassthroughSessions?.has(this.activeSessionId);
|
||||
if (echoPassthrough && (data === '\r' || data === '\x03')) {
|
||||
this._echoPassthroughSessions.delete(this.activeSessionId);
|
||||
}
|
||||
|
||||
// ── Local Echo Mode ──
|
||||
// When enabled, keystrokes are buffered locally in the overlay for
|
||||
// instant visual feedback. Nothing is sent to the PTY until Enter
|
||||
// (or a control char) is pressed — avoids out-of-order char delivery.
|
||||
if (this._localEchoEnabled) {
|
||||
if (this._localEchoEnabled && !echoPassthrough) {
|
||||
if (data === '\x7f') {
|
||||
const source = this._localEchoOverlay?.removeChar();
|
||||
if (source === 'flushed') {
|
||||
@@ -797,9 +920,16 @@ Object.assign(CodemanApp.prototype, {
|
||||
}
|
||||
this._pendingInput += data;
|
||||
flushInput();
|
||||
} else if (source === false) {
|
||||
// Nothing pending, nothing flushed, nothing detected. The
|
||||
// composer may still hold text the overlay cannot see (buffer
|
||||
// detection is suppressed after a control-char flush), so
|
||||
// forward the backspace instead of swallowing it (issue #218);
|
||||
// an empty composer ignores it.
|
||||
this._pendingInput += data;
|
||||
flushInput();
|
||||
}
|
||||
// 'pending' = removed unsent text (no PTY backspace needed)
|
||||
// false = nothing to remove (swallow the backspace)
|
||||
return;
|
||||
}
|
||||
if (/^[\r\n]+$/.test(data)) {
|
||||
@@ -841,6 +971,41 @@ Object.assign(CodemanApp.prototype, {
|
||||
// Single-byte ESC (user pressing Escape) still falls through to
|
||||
// the control char handler below.
|
||||
if (data.length > 1 && data.charCodeAt(0) === 27) {
|
||||
// Bracketed paste (terminal.paste() while DECSET 2004 is on):
|
||||
// flush typed-but-unsent overlay text FIRST so the pasted block
|
||||
// lands after it in the composer, not before it (issue #219).
|
||||
// The paste sequence gets its own delayed write: Codex's
|
||||
// paste-burst handling drops keystrokes that arrive in the SAME
|
||||
// PTY read as a bracketed paste (verified against codex 0.147),
|
||||
// mirroring the delayed \r in the Enter branch above.
|
||||
if (data.startsWith(window.CodemanTerminalInput.BRACKETED_PASTE_START)) {
|
||||
const hadPending = !!this._localEchoOverlay?.pendingText;
|
||||
this._flushLocalEchoPending();
|
||||
if (hadPending) {
|
||||
flushInput();
|
||||
setTimeout(() => {
|
||||
this._pendingInput += data;
|
||||
flushInput();
|
||||
}, 80);
|
||||
} else {
|
||||
this._pendingInput += data;
|
||||
flushInput();
|
||||
}
|
||||
return;
|
||||
}
|
||||
// Composer nav keys (arrows, Home/End, Delete, PgUp/PgDn):
|
||||
// flush unsent text so the key edits the real composer state,
|
||||
// then hand the session to plain PTY echo until Enter/Ctrl+C.
|
||||
// The cursor may now sit mid-text, where append-only buffering
|
||||
// cannot track edits (issue #218).
|
||||
if (window.CodemanTerminalInput.isComposerNavKey(data)) {
|
||||
this._flushLocalEchoPending();
|
||||
if (!this._echoPassthroughSessions) this._echoPassthroughSessions = new Set();
|
||||
this._echoPassthroughSessions.add(this.activeSessionId);
|
||||
this._pendingInput += data;
|
||||
flushInput();
|
||||
return;
|
||||
}
|
||||
// Multi-byte escape sequence — forward to PTY without clearing
|
||||
// overlay/flushed state (terminal response, not user input)
|
||||
this._pendingInput += data;
|
||||
@@ -1408,7 +1573,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
}
|
||||
titleSpan.appendChild(document.createTextNode(s.name || s.firstPrompt || shortDir));
|
||||
|
||||
// Badge row: mode (claude/codex/opencode/gemini/shell) + a LIVE pill.
|
||||
// Badge row: mode (claude/codex/opencode/gemini/antigravity/shell) + a LIVE pill.
|
||||
const badgeRow = document.createElement('div');
|
||||
badgeRow.className = 'history-item-badges';
|
||||
if (s.mode) {
|
||||
@@ -2009,6 +2174,121 @@ Object.assign(CodemanApp.prototype, {
|
||||
}
|
||||
},
|
||||
|
||||
/**
|
||||
* Post-scroll companion to _noteTerminalUserScroll: hitting the TOP of the
|
||||
* buffer while scrolling up is the user reaching for history the browser does
|
||||
* not have, so pull the rest of tmux's scrollback (issue #205, see
|
||||
* _maybeRefetchFullHistory). Must be called AFTER scrollLines(), since the
|
||||
* check is on the resulting position, and it is deliberately not folded into
|
||||
* _noteTerminalUserScroll for exactly that reason. Cheap: one integer compare
|
||||
* per scroll event, and the pull itself is cooldown-guarded.
|
||||
*/
|
||||
_maybeLoadMoreHistoryOnScroll(lines) {
|
||||
if (lines >= 0) return;
|
||||
if (this.terminal?.buffer?.active?.viewportY === 0) this._maybeRefetchFullHistory?.();
|
||||
},
|
||||
|
||||
/**
|
||||
* Rows a `?full=1` capture will occupy once written into xterm.
|
||||
*
|
||||
* tmux joins wrapped rows in that capture (`capture-pane -J`), so a long
|
||||
* logical line re-wraps into several xterm rows on write and a bare newline
|
||||
* count would undershoot; escape sequences occupy no cells and come out
|
||||
* first. Approximate by construction (it ignores double-width glyphs), which
|
||||
* is fine: the only consumer is a coarse size comparison
|
||||
* (_replayWouldShrinkBuffer), and it runs once per cooldown-guarded re-pull.
|
||||
*/
|
||||
_estimateReplayRows(text, cols) {
|
||||
if (typeof text !== 'string' || !text) return 0;
|
||||
const width = cols > 0 ? cols : 80;
|
||||
const plain = text.replace(window.CodemanTerminalInput.REPLAY_ESCAPE_RE, '');
|
||||
let rows = 0;
|
||||
for (const line of plain.split('\n')) {
|
||||
const cells = line.endsWith('\r') ? line.length - 1 : line.length;
|
||||
rows += cells > width ? Math.ceil(cells / width) : 1;
|
||||
}
|
||||
return rows;
|
||||
},
|
||||
|
||||
/**
|
||||
* DOWNGRADE GUARD for the scroll-to-top re-pull (issue #205, round 2).
|
||||
*
|
||||
* `_maybeRefetchFullHistory` resets the terminal and rewrites it from the
|
||||
* capture, which is a straight win when tmux holds more than the browser —
|
||||
* the burst-repaint and tab-switch losses it was built for. But a repaint-mode
|
||||
* CLI pane keeps NO tmux history of its own (`history_size≈0` measured for a
|
||||
* Claude pane), so there the capture is roughly ONE frame while xterm may hold
|
||||
* hundreds of rows of replayed frames. Rewriting then DESTROYS history
|
||||
* mid-scroll: exactly the "goes back a limited amount, repeats blocks, gets
|
||||
* worse when I reach the top" report from the 1.12.0 retest.
|
||||
*
|
||||
* So refuse when the capture is smaller, with a one-screen tolerance because
|
||||
* both sides are estimates: `buffer.active.length` includes the blank rows
|
||||
* below the last line, and _estimateReplayRows can only approximate wrapping.
|
||||
* Only a capture that is worse by more than a full screen counts as a
|
||||
* downgrade, which leaves every genuine recovery case untouched.
|
||||
*/
|
||||
_replayWouldShrinkBuffer(capture) {
|
||||
const term = this.terminal;
|
||||
const rowsNow = term?.buffer?.active?.length || 0;
|
||||
if (!rowsNow) return false;
|
||||
const screen = term?.rows || 24;
|
||||
return this._estimateReplayRows(capture, term?.cols) + screen < rowsNow;
|
||||
},
|
||||
|
||||
/**
|
||||
* Ease-out smooth scrolling for the local wheel path. The capture-phase
|
||||
* wheel handler owns local scrolling (xterm's own smooth scroller is
|
||||
* bypassed, see the listener comment), so without this every notch was an
|
||||
* instant multi-line jump. Wheel deltas accumulate into a pending line
|
||||
* count (fractional — see _wheelScrollLinesFloat) and drain ~22% per
|
||||
* animation frame with a one-line floor, so a single notch starts with a
|
||||
* gentle step and glides to an exact landing; more notches mid-glide deepen
|
||||
* the pending count, which reads as natural acceleration. A sub-line
|
||||
* residual stays pending until further input pushes it past a whole line
|
||||
* (that is what makes slow trackpad drags track the finger). Direction
|
||||
* reversals cancel arithmetically. The pending amount is dropped when the
|
||||
* active session changes mid-glide — leftover momentum must never scroll
|
||||
* the tab the user just switched to.
|
||||
*/
|
||||
_smoothScrollBy(lines) {
|
||||
if (!lines) return;
|
||||
this._smoothScrollPending = (this._smoothScrollPending || 0) + lines;
|
||||
this._smoothScrollSession = this.activeSessionId;
|
||||
if (this._smoothScrollFrame) return;
|
||||
const step = () => {
|
||||
this._smoothScrollFrame = null;
|
||||
const pending = this._smoothScrollPending || 0;
|
||||
if (!pending) return;
|
||||
if (this.activeSessionId !== this._smoothScrollSession) {
|
||||
this._smoothScrollPending = 0;
|
||||
return;
|
||||
}
|
||||
if (Math.abs(pending) < 1) return; // sub-line residual: wait for more input
|
||||
const eased = pending * 0.22;
|
||||
const move = pending > 0 ? Math.max(1, Math.floor(eased)) : Math.min(-1, Math.ceil(eased));
|
||||
this._smoothScrollPending = pending - move;
|
||||
this.terminal.scrollLines(move);
|
||||
this._maybeLoadMoreHistoryOnScroll(move);
|
||||
if (Math.abs(this._smoothScrollPending) >= 1) this._smoothScrollFrame = requestAnimationFrame(step);
|
||||
};
|
||||
this._smoothScrollFrame = requestAnimationFrame(step);
|
||||
},
|
||||
|
||||
/**
|
||||
* Hand a scroll gesture (wheel tick or touch drag, already converted to
|
||||
* lines) to the CLI as synthetic SGR wheel reports. SGR coordinates address
|
||||
* the LIVE screen (the bottom `rows` of the buffer), so a report computed
|
||||
* from a scrolled-up viewport would hit-test a different row entirely, and
|
||||
* forwarding while the user stares at stale scrollback looks like the
|
||||
* gesture is dead. Snap back first: the gesture then always acts on what the
|
||||
* CLI is drawing now.
|
||||
*/
|
||||
_forwardScrollToApp(clientX, clientY, lines) {
|
||||
if (!this._terminalViewportAtBottom()) this.terminal.scrollToBottom();
|
||||
this._sendSyntheticSgrWheel(clientX, clientY, lines);
|
||||
},
|
||||
|
||||
_hasRecentUserScrollUp() {
|
||||
if (typeof this._lastUserScrollUpAt !== 'number') return false;
|
||||
return performance.now() - this._lastUserScrollUpAt < window.CodemanTerminalInput.USER_SCROLL_STICKY_SUPPRESS_MS;
|
||||
@@ -2069,17 +2349,26 @@ Object.assign(CodemanApp.prototype, {
|
||||
|
||||
// Accumulate raw data (may contain DEC 2026 markers)
|
||||
this.pendingWrites.push(data);
|
||||
this._scheduleTerminalWriteFlush();
|
||||
},
|
||||
|
||||
if (!this.writeFrameScheduled) {
|
||||
this.writeFrameScheduled = true;
|
||||
this._safeYield(() => {
|
||||
// xterm.js 6.0 handles DEC 2026 sync markers natively — it buffers
|
||||
// content between 2026h/2026l and renders atomically. No need for
|
||||
// client-side incomplete-block detection; just flush every frame.
|
||||
this.flushPendingWrites();
|
||||
this.writeFrameScheduled = false;
|
||||
});
|
||||
}
|
||||
/**
|
||||
* Schedule one render-budgeted terminal flush.
|
||||
*
|
||||
* Clear the scheduled flag before flushing so flushPendingWrites() can queue
|
||||
* another yield when a large final batch leaves bytes behind. Keeping the
|
||||
* flag set through the flush stranded that remainder until unrelated output
|
||||
* arrived, which looked like truncated responses and idle shell commands.
|
||||
*/
|
||||
_scheduleTerminalWriteFlush() {
|
||||
if (this.writeFrameScheduled || this.pendingWrites.length === 0) return;
|
||||
this.writeFrameScheduled = true;
|
||||
this._safeYield(() => {
|
||||
this.writeFrameScheduled = false;
|
||||
// xterm.js 6.0 handles DEC 2026 sync markers natively — it buffers
|
||||
// content between 2026h/2026l and renders atomically.
|
||||
this.flushPendingWrites();
|
||||
});
|
||||
},
|
||||
|
||||
/**
|
||||
@@ -2095,13 +2384,24 @@ Object.assign(CodemanApp.prototype, {
|
||||
this.flickerFilterActive = false;
|
||||
|
||||
// Trigger a normal flush
|
||||
if (!this.writeFrameScheduled) {
|
||||
this.writeFrameScheduled = true;
|
||||
this._safeYield(() => {
|
||||
this.flushPendingWrites();
|
||||
this.writeFrameScheduled = false;
|
||||
});
|
||||
}
|
||||
this._scheduleTerminalWriteFlush();
|
||||
},
|
||||
|
||||
/**
|
||||
* Flush the local-echo overlay's unsent text into `_pendingInput` (no
|
||||
* trailing Enter) and reset overlay + flushed-state tracking. Used before
|
||||
* forwarding sequences that must arrive AFTER the typed text (bracketed
|
||||
* paste, composer nav keys). The caller forwards its own sequence: nav keys
|
||||
* ride the same write, pastes get a delayed second write because codex
|
||||
* drops keys that share a PTY read with a bracketed paste.
|
||||
*/
|
||||
_flushLocalEchoPending() {
|
||||
const text = this._localEchoOverlay?.pendingText || '';
|
||||
this._localEchoOverlay?.clear();
|
||||
this._localEchoOverlay?.suppressBufferDetection();
|
||||
this._flushedOffsets?.delete(this.activeSessionId);
|
||||
this._flushedTexts?.delete(this.activeSessionId);
|
||||
if (text) this._pendingInput += text;
|
||||
},
|
||||
|
||||
/**
|
||||
@@ -2143,8 +2443,14 @@ Object.assign(CodemanApp.prototype, {
|
||||
}
|
||||
},
|
||||
});
|
||||
} else if (session.mode === 'shell') {
|
||||
} else if (session.mode === 'shell' || session.mode === 'codex') {
|
||||
// Shell mode: the shell provides its own PTY echo so the overlay isn't needed.
|
||||
// Codex mode: the composer is fully interactive per keystroke. Typing
|
||||
// "/" pops a live-filtering command picker (issue #222), the composer
|
||||
// grows and rewraps as it fills (#220), pastes are bracketed (#219)
|
||||
// and arrows/history edit server-side state (#218). Buffering
|
||||
// keystrokes until Enter starves all of that, so codex sessions use
|
||||
// plain PTY echo like shell.
|
||||
// Disable it by clearing any pending text.
|
||||
this._localEchoOverlay.clear();
|
||||
this._localEchoEnabled = false;
|
||||
@@ -2227,13 +2533,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
this.terminal.write(joined.slice(0, MAX_FRAME_BYTES));
|
||||
this.pendingWrites.push(joined.slice(MAX_FRAME_BYTES));
|
||||
deferred = true;
|
||||
if (!this.writeFrameScheduled) {
|
||||
this.writeFrameScheduled = true;
|
||||
this._safeYield(() => {
|
||||
this.flushPendingWrites();
|
||||
this.writeFrameScheduled = false;
|
||||
});
|
||||
}
|
||||
this._scheduleTerminalWriteFlush();
|
||||
}
|
||||
if (
|
||||
preserveViewportY !== null &&
|
||||
@@ -2548,7 +2848,11 @@ Object.assign(CodemanApp.prototype, {
|
||||
/** Insert editable text at the active prompt without pressing Enter. */
|
||||
insertTerminalText(text) {
|
||||
if (!this.activeSessionId || !text) return;
|
||||
if (this._localEchoEnabled && this._localEchoOverlay) {
|
||||
if (
|
||||
this._localEchoEnabled &&
|
||||
this._localEchoOverlay &&
|
||||
!this._echoPassthroughSessions?.has(this.activeSessionId)
|
||||
) {
|
||||
this._localEchoOverlay.appendText(text);
|
||||
} else {
|
||||
this.sendInput(text).catch(() => {});
|
||||
@@ -2815,9 +3119,29 @@ Object.assign(CodemanApp.prototype, {
|
||||
// deltaY≈0 collapses to a fixed ±1 line/tick and the gesture can't page through
|
||||
// history on a trackpad (issue #154). Non-Shift and mouse-wheel paths are
|
||||
// unchanged (they carry deltaY). The `|| ±1` keeps sub-25px deltas moving.
|
||||
//
|
||||
// `deltaMode` says what UNIT the delta is in, and ignoring it made every
|
||||
// non-pixel browser scroll ~4x too slowly: Firefox reports DOM_DELTA_LINE (1)
|
||||
// with deltaY≈3 per notch, so the pixel math rounded to 0 and fell through to
|
||||
// the ±1 fallback — one line per notch, versus 4-5 for Chrome's ~110px. In
|
||||
// Claude mode the same value also capped the forwarded SGR report at one tick.
|
||||
_wheelScrollLines(ev) {
|
||||
const lines = this._wheelScrollLinesFloat(ev);
|
||||
if (!lines) return 0; // pure horizontal swipe: don't fall through to -1
|
||||
return Math.round(lines) || (lines > 0 ? 1 : -1);
|
||||
},
|
||||
|
||||
/** Unrounded variant for the smooth local-scroll path, which accumulates
|
||||
* sub-line fractions across events instead of forcing every tiny trackpad
|
||||
* delta to a whole ±1 line. Same unit handling and Shift-axis trap. */
|
||||
_wheelScrollLinesFloat(ev) {
|
||||
const delta = ev.shiftKey && Math.abs(ev.deltaX) > Math.abs(ev.deltaY) ? ev.deltaX : ev.deltaY;
|
||||
return Math.round(delta / 25) || (delta > 0 ? 1 : -1);
|
||||
if (!delta) return 0;
|
||||
return ev.deltaMode === 1 // DOM_DELTA_LINE (Firefox mouse wheel)
|
||||
? delta
|
||||
: ev.deltaMode === 2 // DOM_DELTA_PAGE
|
||||
? delta * (this.terminal?.rows || 24)
|
||||
: delta / 25; // DOM_DELTA_PIXEL (Chrome/WebKit, and every trackpad)
|
||||
},
|
||||
|
||||
_shouldForwardWheelToApp(ev) {
|
||||
@@ -2826,6 +3150,16 @@ Object.assign(CodemanApp.prototype, {
|
||||
// plain wheel to xterm's own scrollback like pre-#144, for users who prefer
|
||||
// it over forwarding the wheel to the CLI's transcript (issue #154). Cheap —
|
||||
// loadAppSettingsFromStorage() is cache-backed.
|
||||
//
|
||||
// FOOTGUN, and why it is handled downstream rather than here: for a
|
||||
// repaint-mode CLI that local scrollback is EMPTY (tmux keeps no history for
|
||||
// the pane), so this setting can silently convert a working wheel into a
|
||||
// dead one — a plausible reading of the #205 retest, where a user whose
|
||||
// scrolling was broken on 1.11.x may well have flipped it while hunting for
|
||||
// a fix. Scoping the setting away from those modes would be the other
|
||||
// option, but it would override an explicit user choice; instead the caller
|
||||
// falls through to _maybePageCliTranscript, so the gesture still pages the
|
||||
// CLI's transcript and the setting keeps meaning exactly what it says.
|
||||
if (this.loadAppSettingsFromStorage?.()?.terminalWheelLocalScrollback) return false;
|
||||
const mode = this.terminal?.modes?.mouseTrackingMode;
|
||||
if (mode && mode !== 'none') return false;
|
||||
@@ -2836,7 +3170,23 @@ Object.assign(CodemanApp.prototype, {
|
||||
} else if (sessionMode !== 'codex') {
|
||||
return false;
|
||||
}
|
||||
return this._terminalViewportAtBottom();
|
||||
// Deliberately NOT gated on _terminalViewportAtBottom(). It used to be, so
|
||||
// that leaving the bottom handed the wheel back to local scrollback and both
|
||||
// histories stayed reachable without a mode switch. In practice that inverted
|
||||
// the behavior users actually want: a repaint-mode CLI keeps NO terminal
|
||||
// scrollback of its own (tmux reports history_size=0 for a Claude pane), so
|
||||
// xterm's buffer holds only Codeman's REPLAYED repaint frames. Scrolling that
|
||||
// locally drags the CLI's own pinned furniture (the prompt box, the status
|
||||
// line) up the screen and shows stale frames underneath, which reads as "the
|
||||
// window scrolled away" rather than "I am reading history".
|
||||
//
|
||||
// And it was easy to fall into: scrollToLastNonEmptyLine() parks the viewport
|
||||
// `rows - 2` above the last non-empty row, so any tab switch onto a session
|
||||
// with trailing blank rows left the viewport off-bottom and every later wheel
|
||||
// went local. Forwarding unconditionally keeps the CLI's transcript as the
|
||||
// plain wheel's target and its input box fixed in place; local scrollback is
|
||||
// still on Shift+wheel and on the "Wheel scrolls local history" opt-out above.
|
||||
return true;
|
||||
},
|
||||
|
||||
// Encode wheel ticks as SGR reports (button 64 = up, 65 = down) at the pointer
|
||||
@@ -2853,13 +3203,105 @@ Object.assign(CodemanApp.prototype, {
|
||||
if (!pos) return;
|
||||
const btn = lines < 0 ? 64 : 65;
|
||||
const ticks = Math.min(Math.abs(lines), 5);
|
||||
this._queueScrollBytes(`\x1b[<${btn};${pos.col};${pos.row}M`.repeat(ticks));
|
||||
},
|
||||
|
||||
/**
|
||||
* Shared 40ms coalescer for every byte a scroll gesture sends to the PTY (SGR
|
||||
* wheel reports and the PageUp/PageDown fallback alike). Each flush becomes a
|
||||
* tmux send-keys server-side, so per-event writes would spawn a process storm
|
||||
* on a single flick; the queue is bounded so a wild scroll can't build a
|
||||
* backlog that keeps scrolling after the finger stops.
|
||||
*/
|
||||
_queueScrollBytes(data) {
|
||||
if (!data || !this.activeSessionId) return;
|
||||
const queued = this._wheelSgrQueue || '';
|
||||
if (queued.length > 512) return;
|
||||
this._wheelSgrQueue = queued + `\x1b[<${btn};${pos.col};${pos.row}M`.repeat(ticks);
|
||||
this._wheelSgrQueue = queued + data;
|
||||
if (this._wheelSgrFlushTimer) return;
|
||||
this._wheelSgrFlushTimer = setTimeout(() => this._flushWheelSgrQueue(), 40);
|
||||
},
|
||||
|
||||
/**
|
||||
* True when this session's LOCAL scrollback is structurally empty: a Claude
|
||||
* pane in repaint mode, where tmux reports `history_size≈0` and every frame
|
||||
* overwrites the last, so xterm's normal buffer never grows past one screen
|
||||
* (`baseY === 0`). Scrolling that buffer is a no-op no matter how the gesture
|
||||
* is routed — the "wheel does nothing at all" half of the #205 retest.
|
||||
*/
|
||||
_localScrollbackIsHollow() {
|
||||
const mode = this.sessions?.get(this.activeSessionId)?.mode || 'claude';
|
||||
if (mode !== 'claude') return false;
|
||||
const buf = this.terminal?.buffer?.active;
|
||||
if (!buf || buf.type === 'alternate') return false;
|
||||
return (buf.baseY || 0) === 0;
|
||||
},
|
||||
|
||||
/**
|
||||
* LAST-RESORT scroll for a hollow local buffer: translate gesture lines into
|
||||
* coalesced PageUp/PageDown key sends so the CLI pages its OWN transcript.
|
||||
*
|
||||
* The rescue path for every way `_shouldForwardWheelToApp` can come back false
|
||||
* on a Claude session that has no local history to fall back on: the CLI
|
||||
* version probe failed or is genuinely older than 2.1.187, or the user turned
|
||||
* on "Wheel scrolls local history" (which pins the wheel to a buffer that,
|
||||
* for a repaint-mode CLI, is empty — the setting's footgun). Before this, all
|
||||
* of those produced a completely dead gesture; the #205 reporter proved the
|
||||
* keyboard route works by paging back through intact text with Fn+Up.
|
||||
*
|
||||
* Triple-guarded (claude mode + gate false + `baseY === 0`), so a session with
|
||||
* real local scrollback is never touched. Shift is excluded on purpose: it is
|
||||
* the explicit "give me local scrollback" gesture and must keep that meaning.
|
||||
*
|
||||
* @returns true when the gesture was consumed here (the caller must not also
|
||||
* scroll locally).
|
||||
*/
|
||||
_maybePageCliTranscript(ev, lines) {
|
||||
if (!lines || ev?.shiftKey || !this.activeSessionId) return false;
|
||||
if (!this._localScrollbackIsHollow()) return false;
|
||||
// Leftover travel belongs to the tab it was made on.
|
||||
if (this._pageKeySession !== this.activeSessionId) {
|
||||
this._pageKeySession = this.activeSessionId;
|
||||
this._pageKeyPending = 0;
|
||||
}
|
||||
const tuning = window.CodemanTerminalInput;
|
||||
const perPage = Math.max(2, Math.round((this.terminal?.rows || 24) * tuning.PAGE_KEY_SCREEN_FRACTION));
|
||||
const pending = (this._pageKeyPending || 0) + lines;
|
||||
const pages = Math.trunc(pending / perPage);
|
||||
this._pageKeyPending = pending - pages * perPage;
|
||||
if (pages) {
|
||||
const key = pages < 0 ? tuning.KEY_PAGE_UP : tuning.KEY_PAGE_DOWN;
|
||||
this._queueScrollBytes(key.repeat(Math.min(Math.abs(pages), tuning.PAGE_KEY_MAX_PER_BATCH)));
|
||||
}
|
||||
this._logScrollRouting('page-keys');
|
||||
return true;
|
||||
},
|
||||
|
||||
/**
|
||||
* One line in the console saying WHY a scroll gesture went where it went.
|
||||
*
|
||||
* Issue #205 ran two rounds of remote guesswork — is the CLI version probe
|
||||
* empty, is the opt-out setting on, did a mouse DECSET leak past the strip? —
|
||||
* that this single log answers directly. Logged once per session per distinct
|
||||
* decision, so a steady gesture stays silent and a CHANGE (e.g. the version
|
||||
* arriving late and flipping the route) still prints.
|
||||
*/
|
||||
_logScrollRouting(decision) {
|
||||
const sessionId = this.activeSessionId || '(none)';
|
||||
const session = this.sessions?.get(sessionId);
|
||||
const optOut = !!this.loadAppSettingsFromStorage?.()?.terminalWheelLocalScrollback;
|
||||
const tracking = this.terminal?.modes?.mouseTrackingMode || 'none';
|
||||
const baseY = this.terminal?.buffer?.active?.baseY ?? -1;
|
||||
const signature = `${decision}|${session?.mode}|${session?.cliVersion}|${optOut}|${tracking}|${baseY > 0}`;
|
||||
if (!this._scrollRoutingLogged) this._scrollRoutingLogged = new Map();
|
||||
if (this._scrollRoutingLogged.get(sessionId) === signature) return;
|
||||
this._scrollRoutingLogged.set(sessionId, signature);
|
||||
console.log(
|
||||
`[scroll] ${sessionId} → ${decision} (mode=${session?.mode || '?'}, cliVersion=${session?.cliVersion || 'unknown'}, ` +
|
||||
`localScrollbackOptOut=${optOut}, mouseTracking=${tracking}, localScrollbackRows=${baseY})`
|
||||
);
|
||||
},
|
||||
|
||||
_flushWheelSgrQueue() {
|
||||
this._wheelSgrFlushTimer = null;
|
||||
const data = this._wheelSgrQueue;
|
||||
|
||||
@@ -97,8 +97,7 @@ Object.assign(CodemanApp.prototype, {
|
||||
<span class="tab-name">${escapeHtml(webview.name)}</span>
|
||||
</span>
|
||||
</span>
|
||||
<span class="tab-gear" onclick="event.stopPropagation(); app.showWebviewModal(${jsonId})" title="URL settings" aria-label="URL settings" tabindex="0">⚙</span>
|
||||
<span class="tab-close" onclick="event.stopPropagation(); app.closeWebviewTab(${jsonId})" title="Close tab" aria-label="Close web tab" tabindex="0">×</span>
|
||||
<span class="tab-actions"><span class="tab-gear" onclick="event.stopPropagation(); app.showWebviewModal(${jsonId})" title="URL settings" aria-label="URL settings" tabindex="0">⚙</span><span class="tab-close" onclick="event.stopPropagation(); app.closeWebviewTab(${jsonId})" title="Close tab" aria-label="Close web tab" tabindex="0">×</span></span>
|
||||
</div>`);
|
||||
idx++;
|
||||
}
|
||||
|
||||
@@ -640,6 +640,25 @@ async function buildExternalAttachmentRouteItem(
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Headers Fastify already put on the reply, in a shape `writeHead` accepts.
|
||||
*
|
||||
* `reply.raw.writeHead()` writes straight to the Node response and bypasses
|
||||
* Fastify's header store, so anything the security `onRequest` hook granted — CORS
|
||||
* for localhost origins, nosniff, frame-options, CSP — is silently dropped on every
|
||||
* route that answers this way. Spread this first and let the route's own headers
|
||||
* win over it.
|
||||
*/
|
||||
function inheritedHeaders(reply: {
|
||||
getHeaders(): NodeJS.Dict<number | string | string[]>;
|
||||
}): Record<string, number | string | string[]> {
|
||||
const out: Record<string, number | string | string[]> = {};
|
||||
for (const [name, value] of Object.entries(reply.getHeaders())) {
|
||||
if (value !== undefined) out[name] = value;
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & EventPort & ConfigPort): void {
|
||||
// Lazy filesystem listing for the Link Existing and mobile input path pickers.
|
||||
app.get('/api/filesystem/browse', async (req, reply): Promise<ApiResponse<FilesystemBrowseData>> => {
|
||||
@@ -1337,6 +1356,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
|
||||
const basename = rawBasename.replace(/["\\\r\n]/g, '_');
|
||||
if (download === 'true' || ext === 'svg') {
|
||||
reply.raw.writeHead(200, {
|
||||
...inheritedHeaders(reply),
|
||||
'Content-Type': ext === 'svg' ? 'application/octet-stream' : mimeTypes[ext] || 'application/octet-stream',
|
||||
'Content-Disposition': `attachment; filename="${basename}"`,
|
||||
'Content-Length': content.length,
|
||||
@@ -1576,6 +1596,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
|
||||
|
||||
// Set up SSE headers
|
||||
reply.raw.writeHead(200, {
|
||||
...inheritedHeaders(reply),
|
||||
'Content-Type': 'text/event-stream',
|
||||
'Cache-Control': 'no-cache',
|
||||
Connection: 'keep-alive',
|
||||
@@ -1709,6 +1730,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
|
||||
const content = await fs.readFile(resolvedPath);
|
||||
// Bypass Fastify compression — write directly to raw response
|
||||
reply.raw.writeHead(200, {
|
||||
...inheritedHeaders(reply),
|
||||
'Content-Type': mimeTypes[ext] || 'application/octet-stream',
|
||||
'Content-Disposition': `attachment; filename="${filename}"`,
|
||||
'Content-Length': content.length,
|
||||
|
||||
@@ -10,6 +10,7 @@ import { HookEventSchema, isValidWorkingDir } from '../schemas.js';
|
||||
import { sanitizeHookData, parseBody } from '../route-helpers.js';
|
||||
import { persistDockerCaseClaudeSessionId } from '../../docker-hosts.js';
|
||||
import { getDataDir } from '../../config/instance.js';
|
||||
import { sessionWaits, hooksAvailableForMode } from '../session-wait-registry.js';
|
||||
import type { SessionPort, EventPort, RespawnPort, ConfigPort, InfraPort } from '../ports/index.js';
|
||||
|
||||
export function registerHookEventRoutes(
|
||||
@@ -22,6 +23,27 @@ export function registerHookEventRoutes(
|
||||
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Session not found');
|
||||
}
|
||||
|
||||
// Wake anything blocked on `GET /api/sessions/:id/wait`. Hooks are the only
|
||||
// DEFINITIVE signals Codeman gets (`idle` is inferred from output stabilization
|
||||
// and can flap mid-turn), so these two are what an orchestrating agent should
|
||||
// wait on.
|
||||
//
|
||||
// Gated on the session's MODE, matching `resolveWaitSignals` on the read side.
|
||||
// Without it the guard is one-sided: a caller cannot ASK for `stop` on a shell or
|
||||
// codex session, but this endpoint would happily deliver one for it. Hook events
|
||||
// carry no identity beyond a per-instance secret shared by every case, so this is
|
||||
// also the cheap half of the forgery surface — a `stop` claimed for a session that
|
||||
// could never legitimately emit one is now dropped instead of steering another
|
||||
// agent's control flow.
|
||||
const waitSession = ctx.sessions.get(sessionId);
|
||||
if (waitSession && hooksAvailableForMode(waitSession.mode)) {
|
||||
if (event === 'stop') {
|
||||
sessionWaits.notifySignal(sessionId, 'stop');
|
||||
} else if (event === 'permission_prompt' || event === 'elicitation_dialog') {
|
||||
sessionWaits.notifySignal(sessionId, 'blocked');
|
||||
}
|
||||
}
|
||||
|
||||
// Signal the respawn controller based on hook event type
|
||||
const controller = ctx.respawnControllers.get(sessionId);
|
||||
if (controller) {
|
||||
|
||||
@@ -4,7 +4,8 @@
|
||||
* auto-clear, auto-compact, image watcher, flicker filter, and logout.
|
||||
*/
|
||||
|
||||
import { FastifyInstance } from 'fastify';
|
||||
import { FastifyInstance, type FastifyReply } from 'fastify';
|
||||
import { z } from 'zod';
|
||||
import { join, dirname, extname, basename } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { existsSync, statSync, mkdirSync, writeFileSync } from 'node:fs';
|
||||
@@ -17,11 +18,12 @@ import {
|
||||
getErrorMessage,
|
||||
type ApiResponse,
|
||||
type SessionColor,
|
||||
type SessionStatus,
|
||||
type CodexConfig,
|
||||
type GeminiConfig,
|
||||
type AntigravityConfig,
|
||||
} from '../../types.js';
|
||||
import { Session, isAltScreenStripMode } from '../../session.js';
|
||||
import { Session, isAltScreenStripMode, isMuxAltScreenOnlyStripMode } from '../../session.js';
|
||||
import { SseEvent } from '../sse-events.js';
|
||||
import {
|
||||
CreateSessionSchema,
|
||||
@@ -40,8 +42,19 @@ import {
|
||||
QuickStartSchema,
|
||||
InteractiveStartSchema,
|
||||
SessionOrderUpdateSchema,
|
||||
SessionWaitQuerySchema,
|
||||
SessionWaitOutputQuerySchema,
|
||||
} from '../schemas.js';
|
||||
import { mergeSessionOrder } from '../../session-order.js';
|
||||
import {
|
||||
sessionWaits,
|
||||
resolveWaitSignals,
|
||||
signalForStatus,
|
||||
WaitCapacityError,
|
||||
type WaitSignal,
|
||||
type SignalWaitResult,
|
||||
} from '../session-wait-registry.js';
|
||||
import { clampWaitMs, MAX_BUFFER_SCAN_BYTES } from '../../config/agent-wait.js';
|
||||
import {
|
||||
autoConfigureRalph,
|
||||
canAccessOwned,
|
||||
@@ -314,6 +327,234 @@ async function clampExternalCliBypassForOwner(
|
||||
return { codexConfig: clampedCodex, geminiConfig: clampedGemini, antigravityConfig: clampedAntigravity };
|
||||
}
|
||||
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
// Agent wait helpers (shared by GET /wait, GET /wait-output, POST /input)
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
|
||||
/**
|
||||
* Validate a wait query WITHOUT throwing away the Zod issue.
|
||||
*
|
||||
* `parseBody`'s message argument REPLACES the issue text, so `?timeout=30s` came
|
||||
* back as a bare "Invalid wait parameters": the caller could not tell which of
|
||||
* `until`, `timeout` or `fresh` it got wrong, and its only move was to retry with
|
||||
* a different guess. These endpoints are driven by an LLM with no documentation in
|
||||
* context — the error message IS the documentation, which is why the signal parser
|
||||
* one line later goes to the trouble of naming the bad token and listing the valid
|
||||
* ones. This keeps the endpoint label AND names the offending field.
|
||||
*/
|
||||
function parseWaitQuery<T>(schema: z.ZodType<T>, query: unknown, label: string): T {
|
||||
const result = schema.safeParse(query);
|
||||
if (result.success) return result.data;
|
||||
const issue = result.error.issues[0];
|
||||
const field = issue && issue.path.length > 0 ? issue.path.join('.') : '';
|
||||
const detail = issue?.message ?? 'validation failed';
|
||||
const message = field ? `Invalid ${label} parameter '${field}': ${detail}` : `Invalid ${label} parameters: ${detail}`;
|
||||
throw Object.assign(new Error(message), {
|
||||
statusCode: 400,
|
||||
body: createErrorResponse(ApiErrorCode.INVALID_INPUT, message),
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Map a waiter-cap rejection to the code that tells the caller the truth.
|
||||
*
|
||||
* The two caps mean different things and warrant different recovery: `session` is
|
||||
* genuinely about THIS session, while `owner` and `total` are process-wide budgets
|
||||
* that say nothing about it. Reporting a global cap as `SESSION_BUSY` (409,
|
||||
* documented as "Session is busy") sent an agent off to a different session to hit
|
||||
* the identical error. `RATE_LIMITED` is the code whose whole meaning is "come back
|
||||
* later", and clients and proxies already treat 429 that way.
|
||||
*/
|
||||
function waitCapacityResponse(err: WaitCapacityError): ApiResponse<never> {
|
||||
const code = err.scope === 'session' ? ApiErrorCode.SESSION_BUSY : ApiErrorCode.RATE_LIMITED;
|
||||
// The registry's message already names the scope and the limit; passing it through
|
||||
// verbatim keeps the wording in one place.
|
||||
return createErrorResponse(code, err.message);
|
||||
}
|
||||
|
||||
/**
|
||||
* The signal a session is ALREADY emitting, corrected for liveness.
|
||||
*
|
||||
* `signalForStatus` alone is not enough here, because `Session` parks a DEAD PTY at
|
||||
* `_status = 'idle'` (both `onExit` handlers do) and the object survives in the
|
||||
* session map until an explicit DELETE. Trusting the status therefore answers the
|
||||
* default wait with `{signal:"idle", immediate:true}` for a worker that has
|
||||
* crashed — HTTP 200, no error anywhere, and the agent types its next prompt into a
|
||||
* corpse — while `until=exit` blocks for the full timeout on an event that already
|
||||
* happened and can never happen again.
|
||||
*
|
||||
* `pid === null` means no process is behind this session: it exited, it was
|
||||
* detached, or it was created and never started. All three are `exit` from a
|
||||
* caller's point of view — nothing is running — and in all three the agent's
|
||||
* correct next move is to (re)start the worker rather than to type at it. The
|
||||
* response still carries the raw `status` alongside, so nothing is hidden.
|
||||
*
|
||||
* ⚠️ `pid` alone is NOT enough, and on the normal configuration it is never the
|
||||
* thing that fires — see `workerIsDead()`. `dead` carries the mux layer's answer.
|
||||
*
|
||||
* Fixing it HERE rather than in `signalForStatus` is deliberate: liveness is not
|
||||
* derivable from `SessionStatus`, and the registry holds no `Session` reference.
|
||||
*/
|
||||
function currentSignalFor(session: { pid: number | null; status: SessionStatus }, dead: boolean): WaitSignal | null {
|
||||
if (dead || session.pid === null || session.pid === undefined) return 'exit';
|
||||
return signalForStatus(session.status);
|
||||
}
|
||||
|
||||
// ── Worker liveness for tmux-backed sessions ────────────────────────────────
|
||||
//
|
||||
// `session.pid` is the LOCAL `tmux attach` client, not the worker. Codeman sets
|
||||
// `remain-on-exit on` for every session it creates, so when the command inside the
|
||||
// pane exits, tmux keeps the pane (`pane_dead=1`), the tmux session survives, the
|
||||
// attach client keeps running and `pid` never goes null — no `exit` event is emitted
|
||||
// and nothing in `Session` changes. Measured on a shell worker killed with `exit 42`:
|
||||
// tmux reports `pane_dead=1 status=42` while Codeman reports `pid=309406 status=idle`
|
||||
// and the DEFAULT wait answers `{signal:"idle", immediate:true}` in 0 ms for a corpse.
|
||||
// So the liveness check has to ask the mux layer. `pid === null` still matters: it is
|
||||
// the right (and only) answer for a direct-PTY session, which has no pane to ask about.
|
||||
//
|
||||
// Cost control, because `isPaneDead()` is a synchronous `execSync` and `/wait` is
|
||||
// polled in a loop by design:
|
||||
// 1. Only mux-backed sessions are probed at all.
|
||||
// 2. Only requests that actually wait probe — a plain `POST .../input` (the browser's
|
||||
// hot path, thousands per session) never touches tmux.
|
||||
// 3. Results are cached per pane for PANE_DEATH_TTL_MS, so a poll loop cannot turn
|
||||
// into one exec per request.
|
||||
// 4. The while-blocked watcher is ONE timer per session no matter how many waiters
|
||||
// are parked on it, and it exists only while at least one of them is.
|
||||
|
||||
/** How long a pane-liveness probe is reused. Long enough to absorb a poll loop. */
|
||||
const PANE_DEATH_TTL_MS = 750;
|
||||
|
||||
/** How often a session with a parked waiter is re-checked for a dead worker. */
|
||||
const PANE_DEATH_POLL_MS = 3_000;
|
||||
|
||||
/** Bounded, because a 24h server churns through panes. */
|
||||
const paneDeathCache = new LRUMap<string, { dead: boolean; at: number }>({ maxSize: 256 });
|
||||
|
||||
/** One watcher per pane, refcounted by the waits currently parked on it. */
|
||||
const paneDeathWatchers = new Map<string, { timer: NodeJS.Timeout; refs: number }>();
|
||||
|
||||
type LivenessSession = { usesMux?: boolean; muxName?: string | null };
|
||||
|
||||
/**
|
||||
* Whether the worker inside this session's tmux pane has exited.
|
||||
*
|
||||
* False for anything not tmux-backed (nothing to ask), and false when the probe is
|
||||
* unavailable or throws — an unknown answer must never invent a death.
|
||||
*/
|
||||
function workerIsDead(mux: InfraPort['mux'], session: LivenessSession, now: number = Date.now()): boolean {
|
||||
const muxName = session.usesMux === false ? null : session.muxName;
|
||||
if (!muxName) return false;
|
||||
// Defensive: `TerminalMultiplexer` declares it, but route-test doubles may not.
|
||||
if (typeof mux?.isPaneDead !== 'function') return false;
|
||||
|
||||
const cached = paneDeathCache.get(muxName);
|
||||
if (cached && now - cached.at < PANE_DEATH_TTL_MS) return cached.dead;
|
||||
|
||||
let dead = false;
|
||||
try {
|
||||
dead = mux.isPaneDead(muxName) === true;
|
||||
} catch {
|
||||
dead = false;
|
||||
}
|
||||
paneDeathCache.set(muxName, { dead, at: now });
|
||||
return dead;
|
||||
}
|
||||
|
||||
/**
|
||||
* Release every waiter on a session whose worker has died, in the documented order.
|
||||
*
|
||||
* The same pair the PTY-exit listener and the delete path use, for the same reason:
|
||||
* `until=exit` callers get their signal, everyone else gets `ended: true` instead of
|
||||
* burning the rest of their timeout on feeds that will never produce anything.
|
||||
*/
|
||||
function releaseWaitersForDeadWorker(sessionId: string): void {
|
||||
sessionWaits.notifySignal(sessionId, 'exit');
|
||||
sessionWaits.cancelAll(sessionId);
|
||||
}
|
||||
|
||||
/**
|
||||
* While a wait is parked on a mux-backed session, poll for the worker dying.
|
||||
*
|
||||
* Without this, a worker that dies DURING a wait is invisible: no `exit` event fires
|
||||
* (the attach client is still alive), no output arrives, and the caller blocks for its
|
||||
* full timeout — the common orchestration case, "send a prompt and wait", where the
|
||||
* worker crashes mid-turn.
|
||||
*
|
||||
* @returns a release function; call it in a `finally`, or the timer outlives the wait.
|
||||
*/
|
||||
function watchForDeadWorker(mux: InfraPort['mux'], session: LivenessSession, sessionId: string): () => void {
|
||||
const muxName = session.usesMux === false ? null : session.muxName;
|
||||
if (!muxName || typeof mux?.isPaneDead !== 'function') return () => {};
|
||||
|
||||
const existing = paneDeathWatchers.get(muxName);
|
||||
if (existing) {
|
||||
existing.refs++;
|
||||
} else {
|
||||
const timer = setInterval(() => {
|
||||
if (!workerIsDead(mux, session)) return;
|
||||
releaseWaitersForDeadWorker(sessionId);
|
||||
}, PANE_DEATH_POLL_MS);
|
||||
// Auxiliary to the waiter's own timer, which is deliberately NOT unref'd; this one
|
||||
// must never be the reason the process stays up.
|
||||
timer.unref();
|
||||
paneDeathWatchers.set(muxName, { timer, refs: 1 });
|
||||
}
|
||||
|
||||
let released = false;
|
||||
return () => {
|
||||
if (released) return;
|
||||
released = true;
|
||||
const entry = paneDeathWatchers.get(muxName);
|
||||
if (!entry) return;
|
||||
entry.refs--;
|
||||
if (entry.refs <= 0) {
|
||||
clearInterval(entry.timer);
|
||||
paneDeathWatchers.delete(muxName);
|
||||
}
|
||||
};
|
||||
}
|
||||
|
||||
/** Test seam: pane-liveness state is module-level, so a suite must be able to reset it. */
|
||||
export function _resetPaneLivenessState(): void {
|
||||
for (const entry of paneDeathWatchers.values()) clearInterval(entry.timer);
|
||||
paneDeathWatchers.clear();
|
||||
paneDeathCache.clear();
|
||||
}
|
||||
|
||||
/** Test seam: how many panes are currently being watched for a dead worker. */
|
||||
export function _paneDeathWatcherCount(): number {
|
||||
return paneDeathWatchers.size;
|
||||
}
|
||||
|
||||
/**
|
||||
* An `AbortController` that fires when the CLIENT goes away, and only then.
|
||||
*
|
||||
* Freeing an abandoned waiter matters because the documented pattern is a loop of
|
||||
* short waits: `curl --max-time 30 ".../wait?timeout=600000"` abandons a live waiter
|
||||
* every iteration until the cap is hit and an innocent session reports busy. Same for
|
||||
* any proxy that cuts the connection.
|
||||
*
|
||||
* ⚠️ **It must listen on the RESPONSE, not the request.** `req.raw` emits `'close'`
|
||||
* as soon as the request body has finished streaming, which on a POST is BEFORE the
|
||||
* handler ever blocks — measured at +1ms with `aborted: false`, indistinguishable
|
||||
* from a real hang-up at +0ms. Wiring the abort there cancels every send-and-wait
|
||||
* instantly and silently kills the feature (it survives on GET only because a GET has
|
||||
* no body to finish). `reply.raw` emits `'close'` both when the response completes
|
||||
* and when the socket dies, and `writableFinished` is what tells those apart: true
|
||||
* only if the response actually went out. The guard is load-bearing, not defensive.
|
||||
*
|
||||
* `app.inject()` never emits `'close'` at all, so this is only observable over real
|
||||
* HTTP — which is why the regression test for it binds a port.
|
||||
*/
|
||||
function abortOnClientHangUp(reply: FastifyReply): AbortController {
|
||||
const controller = new AbortController();
|
||||
reply.raw.on('close', () => {
|
||||
if (!reply.raw.writableFinished) controller.abort();
|
||||
});
|
||||
return controller;
|
||||
}
|
||||
|
||||
export function registerSessionRoutes(
|
||||
app: FastifyInstance,
|
||||
ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort
|
||||
@@ -415,7 +656,7 @@ export function registerSessionRoutes(
|
||||
//
|
||||
// For keys the caller is actively setting, strip any stale disk entry a prior
|
||||
// Codeman version may have written. Scope limited to:
|
||||
// - Claude mode (OpenCode/Codex/Gemini don't read .claude/settings.local.json)
|
||||
// - Claude mode (OpenCode/Codex/Gemini/Antigravity don't read .claude/settings.local.json)
|
||||
// - workingDir inside CASES_DIR / the per-user case space (Codeman's managed
|
||||
// territory — we never mutate .claude/settings.local.json in arbitrary user
|
||||
// repos that POST /api/sessions can target, as those may have hand-authored
|
||||
@@ -863,9 +1104,9 @@ export function registerSessionRoutes(
|
||||
|
||||
// ========== Send Input ==========
|
||||
|
||||
app.post('/api/sessions/:id/input', async (req) => {
|
||||
app.post('/api/sessions/:id/input', async (req, reply) => {
|
||||
const { id } = req.params as { id: string };
|
||||
const { input, useMux, seq, clientId } = parseBody(SessionInputWithLimitSchema, req.body);
|
||||
const { input, useMux, seq, clientId, wait, waitTimeout } = parseBody(SessionInputWithLimitSchema, req.body);
|
||||
const session = findSessionOrFail(ctx, id, req);
|
||||
|
||||
const inputStr = String(input);
|
||||
@@ -876,33 +1117,319 @@ export function registerSessionRoutes(
|
||||
);
|
||||
}
|
||||
|
||||
// Send-and-wait (agent orchestration). This has to be ONE endpoint rather than a
|
||||
// POST followed by GET .../wait: between the write and the session flipping to
|
||||
// `working` there is a window in which a separate wait sees the session still
|
||||
// idle and returns instantly, reporting the PREVIOUS turn as this turn's answer.
|
||||
// Registering the waiter before the write closes that window.
|
||||
const wantsWait =
|
||||
wait === true || (typeof wait === 'string' && wait.trim().length > 0) || (Array.isArray(wait) && wait.length > 0);
|
||||
let until: readonly WaitSignal[] = [];
|
||||
if (wantsWait) {
|
||||
const resolved = resolveWaitSignals(wait === true ? undefined : wait, { mode: session.mode });
|
||||
if (resolved.error) return createErrorResponse(ApiErrorCode.INVALID_INPUT, resolved.error);
|
||||
until = resolved.until;
|
||||
}
|
||||
|
||||
// Reliable delivery (POST fallback when the WebSocket is down): a 2xx IS the
|
||||
// client's ACK, so a tagged duplicate redelivery must still return 200 but
|
||||
// skip the write. Untagged requests (curl/legacy) always apply.
|
||||
if (typeof clientId === 'string' && typeof seq === 'number' && !session.shouldApplyInput(clientId, seq)) {
|
||||
const tagged = typeof clientId === 'string' && typeof seq === 'number';
|
||||
const duplicate = tagged && !session.shouldApplyInput(clientId as string, seq as number);
|
||||
if (duplicate && !wantsWait) {
|
||||
return {};
|
||||
}
|
||||
|
||||
// Only a waiting request pays for the tmux probe: the browser's plain input path
|
||||
// (thousands of calls per session) must stay exec-free.
|
||||
const workerDead = wantsWait && workerIsDead(ctx.mux, session);
|
||||
|
||||
const timeoutMs = clampWaitMs(waitTimeout ?? undefined);
|
||||
// Same slot leak as the GET routes: a client that gives up mid-wait would
|
||||
// otherwise hold a waiter for the full timeout. Response-side, always — see
|
||||
// abortOnClientHangUp: on THIS route a request-side listener fires the moment the
|
||||
// JSON body finishes streaming and aborts every send-and-wait before it starts.
|
||||
const abort = abortOnClientHangUp(reply);
|
||||
let waitPromise: Promise<SignalWaitResult> | null = null;
|
||||
if (wantsWait) {
|
||||
try {
|
||||
waitPromise = sessionWaits.waitForSignal(id, {
|
||||
until,
|
||||
timeoutMs,
|
||||
owner: ownerFor(req),
|
||||
abortSignal: abort.signal,
|
||||
// A FRESH delivery must not be satisfied by the state the session is already
|
||||
// in: it is idle right now, which is precisely why we are typing at it.
|
||||
// A DUPLICATE has no new turn coming, so it answers from the current state
|
||||
// instead of blocking for a transition that already happened.
|
||||
requireTransition: !duplicate,
|
||||
currentSignal: duplicate ? currentSignalFor(session, workerDead) : undefined,
|
||||
});
|
||||
} catch (err) {
|
||||
if (err instanceof WaitCapacityError) {
|
||||
// Nothing has been written yet, but `shouldApplyInput` already consumed the
|
||||
// seq. Give it back or the caller's retry is rejected as a duplicate and the
|
||||
// input is lost by the very mechanism meant to make delivery reliable.
|
||||
if (tagged && !duplicate) session.forgetInputSeq(clientId as string, seq as number);
|
||||
return waitCapacityResponse(err);
|
||||
}
|
||||
throw err;
|
||||
}
|
||||
}
|
||||
const stopDeathWatch = wantsWait ? watchForDeadWorker(ctx.mux, session, id) : () => {};
|
||||
|
||||
// Write input to PTY. Direct write is synchronous; writeViaMux
|
||||
// (tmux send-keys) is fire-and-forget to avoid blocking the HTTP response.
|
||||
if (useMux) {
|
||||
// Fire-and-forget: don't block HTTP response on tmux child process.
|
||||
// Fallback to direct write on failure.
|
||||
//
|
||||
// Because the response has already been sent by then, a failure there is the
|
||||
// one case the caller can never learn about — so the dedup bookkeeping is
|
||||
// rolled back. Otherwise the seq stays recorded as applied and a retry, the
|
||||
// very mechanism reliable delivery exists for, is rejected as a duplicate.
|
||||
const undoOnFailure = () => {
|
||||
if (tagged) session.forgetInputSeq(clientId as string, seq as number);
|
||||
};
|
||||
|
||||
// Whether the bytes actually reached a write path. Only meaningful on the wait
|
||||
// path (the fire-and-forget branches return before the response is built), and
|
||||
// reported there instead of the old `!duplicate`: a PTY that has exited fails
|
||||
// BOTH writes, and telling the caller "delivered, but it timed out" points it at
|
||||
// the wrong recovery — wait longer, when the truth is "restart the worker".
|
||||
let delivered = false;
|
||||
|
||||
if (duplicate) {
|
||||
// Redelivery of an already-applied input: skip the write, but still honor the
|
||||
// wait, since the caller's question ("tell me when this settles") is unanswered.
|
||||
} else if (useMux && waitPromise) {
|
||||
// The response is already staying open for the wait, so the tmux write can be
|
||||
// awaited here. This is the ONE path where a writeViaMux failure is observable.
|
||||
const ok = await session.writeViaMux(inputStr).catch(() => false);
|
||||
if (ok) {
|
||||
delivered = true;
|
||||
} else {
|
||||
console.warn(`[Server] writeViaMux failed for session ${id}, falling back to direct write`);
|
||||
delivered = session.write(inputStr);
|
||||
if (!delivered) undoOnFailure();
|
||||
}
|
||||
} else if (useMux) {
|
||||
// Fire-and-forget: don't block the HTTP response on a tmux child process.
|
||||
// Fallback to a direct write on failure. Unchanged from before send-and-wait.
|
||||
session
|
||||
.writeViaMux(inputStr)
|
||||
.then((ok) => {
|
||||
if (!ok) {
|
||||
console.warn(`[Server] writeViaMux failed for session ${id}, falling back to direct write`);
|
||||
session.write(inputStr);
|
||||
}
|
||||
if (ok) return;
|
||||
console.warn(`[Server] writeViaMux failed for session ${id}, falling back to direct write`);
|
||||
if (!session.write(inputStr)) undoOnFailure();
|
||||
})
|
||||
.catch(() => {
|
||||
session.write(inputStr);
|
||||
if (!session.write(inputStr)) undoOnFailure();
|
||||
});
|
||||
} else {
|
||||
session.write(inputStr);
|
||||
// Same rollback. NOT an error response, deliberately: a session can
|
||||
// legitimately have no PTY yet (created but not started), and callers have
|
||||
// always been able to write to one without a 4xx.
|
||||
delivered = session.write(inputStr);
|
||||
if (!delivered && tagged) {
|
||||
session.forgetInputSeq(clientId as string, seq as number);
|
||||
}
|
||||
}
|
||||
|
||||
if (!waitPromise) return {};
|
||||
|
||||
try {
|
||||
// `send-keys` SUCCEEDS against a dead pane — tmux is happy to write into a corpse
|
||||
// — so a truthful `delivered` cannot come from the write's return value alone.
|
||||
// This is the case the field exists for: "delivered, but it timed out" tells an
|
||||
// agent to wait longer when the truth is "restart the worker".
|
||||
if (delivered && workerDead) {
|
||||
delivered = false;
|
||||
// The bytes went nowhere, so the seq must not be recorded as applied or the
|
||||
// caller's retry against a restarted worker is refused as a duplicate.
|
||||
if (!duplicate) undoOnFailure();
|
||||
}
|
||||
|
||||
// Nothing was written and nothing will be: no turn is coming, so blocking for the
|
||||
// full timeout would only delay the caller's real recovery by up to ten minutes.
|
||||
// Releasing the waiter also hands its slot back immediately.
|
||||
const selfReleased = !delivered && !duplicate;
|
||||
if (selfReleased) abort.abort();
|
||||
|
||||
const result = await waitPromise;
|
||||
return {
|
||||
success: true,
|
||||
data: {
|
||||
delivered,
|
||||
duplicate,
|
||||
status: session.status,
|
||||
limitPaused: session.isLimitPaused,
|
||||
// Identical `wait` object to the two GET endpoints, so one client helper
|
||||
// reads all three, `timeoutMs` (post-clamp) included.
|
||||
//
|
||||
// `aborted` is the CLIENT-facing "you hung up, nobody is reading this", and
|
||||
// by that definition it is unobservable — which is exactly what the API
|
||||
// reference promises. The abort above is the server releasing its own waiter
|
||||
// on a delivery that failed, and the client IS reading this response, so
|
||||
// reporting `aborted: true` there would break that promise and hand an agent
|
||||
// a second, contradictory reason for the same outcome. `delivered: false`
|
||||
// already says what happened; `ended` says the wait was released early.
|
||||
wait: { ...result, aborted: selfReleased ? false : result.aborted, until: [...until] },
|
||||
},
|
||||
};
|
||||
} finally {
|
||||
stopDeathWatch();
|
||||
}
|
||||
});
|
||||
|
||||
// ========== Wait For A Signal (agent orchestration) ==========
|
||||
//
|
||||
// A bounded long-poll: block until the session hits one of `until`, then answer.
|
||||
// This exists because SSE is the only "tell me when" channel Codeman has, and an
|
||||
// agent driving the API from a shell tool cannot hold a stream and parse events
|
||||
// inline. See docs/agent-control-plan.md.
|
||||
//
|
||||
// A TIMEOUT IS A 200, not an error: callers are expected to loop over short waits
|
||||
// (proxies such as `tailscale serve` cut idle connections), and turning every poll
|
||||
// boundary into a 4xx would make that loop indistinguishable from a real failure.
|
||||
|
||||
app.get('/api/sessions/:id/wait', async (req, reply) => {
|
||||
const { id } = req.params as { id: string };
|
||||
const query = parseWaitQuery(SessionWaitQuerySchema, req.query, 'wait');
|
||||
const session = findSessionOrFail(ctx, id, req);
|
||||
|
||||
// An agent polls this URL in a loop with identical parameters. Any intermediary
|
||||
// applying heuristic freshness to the 200 would serve the stored `timedOut:true`
|
||||
// body to the next iteration instantly, turning the loop into a busy spin that
|
||||
// never observes the signal.
|
||||
reply.header('Cache-Control', 'no-store');
|
||||
|
||||
// Shared with the `wait` field on POST .../input: unknown token is a 400,
|
||||
// hook-only signals are rejected explicitly but dropped from the default.
|
||||
const { until, error } = resolveWaitSignals(query.until, { mode: session.mode });
|
||||
if (error) return createErrorResponse(ApiErrorCode.INVALID_INPUT, error);
|
||||
|
||||
// The value actually applied after clamping, echoed below: a caller that asked
|
||||
// for 30 minutes and silently got 10 could not otherwise tell a poll boundary
|
||||
// from a wedged worker, and would kill a session that was working fine.
|
||||
const timeoutMs = clampWaitMs(query.timeout);
|
||||
|
||||
// Free the waiter when the caller hangs up; the response can no longer be sent by
|
||||
// then, so freeing the slot is the entire purpose.
|
||||
const abort = abortOnClientHangUp(reply);
|
||||
// A worker that dies while this request is parked emits nothing at all (the tmux
|
||||
// attach client survives it), so a wait would otherwise run to its full timeout.
|
||||
const stopDeathWatch = watchForDeadWorker(ctx.mux, session, id);
|
||||
|
||||
try {
|
||||
const result = await sessionWaits.waitForSignal(id, {
|
||||
until,
|
||||
timeoutMs,
|
||||
owner: ownerFor(req),
|
||||
abortSignal: abort.signal,
|
||||
requireTransition: query.fresh === '1' || query.fresh === 'true',
|
||||
// Read BEFORE awaiting: this is the state the caller is asking about.
|
||||
currentSignal: currentSignalFor(session, workerIsDead(ctx.mux, session)),
|
||||
});
|
||||
|
||||
return {
|
||||
success: true,
|
||||
data: {
|
||||
sessionId: id,
|
||||
// Post-wait status, so a caller that timed out still learns where things stand.
|
||||
status: session.status,
|
||||
// A session paused on a usage limit emits nothing until its reset, so a
|
||||
// timeout here is expected rather than a stall worth retrying hard.
|
||||
limitPaused: session.isLimitPaused,
|
||||
// One shape across all three endpoints, so a single `is_done(resp)` helper
|
||||
// works against any of them. `result.timeoutMs` is the value actually
|
||||
// applied after clamping, which is what makes the clamp observable.
|
||||
wait: { ...result, until: [...until] },
|
||||
},
|
||||
};
|
||||
} catch (err) {
|
||||
if (err instanceof WaitCapacityError) return waitCapacityResponse(err);
|
||||
throw err;
|
||||
} finally {
|
||||
stopDeathWatch();
|
||||
}
|
||||
});
|
||||
|
||||
// ========== Wait For Output (agent orchestration) ==========
|
||||
//
|
||||
// The companion to /wait: block until a literal string appears in this session's
|
||||
// output. Same 200-on-timeout contract. Fed by the `terminal` listener in
|
||||
// session-listener-wiring.ts, so what this scans is byte-for-byte what the pane
|
||||
// printed, ANSI stripped.
|
||||
//
|
||||
// ⚠️ A tmux repaint replays text already on screen, so `from=now` can match
|
||||
// something printed before the request. Callers need a marker unique per call.
|
||||
|
||||
app.get('/api/sessions/:id/wait-output', async (req, reply) => {
|
||||
const { id } = req.params as { id: string };
|
||||
|
||||
// Reject `regex` loudly instead of ignoring it. Matching is deliberately literal
|
||||
// (no ReDoS surface on a caller-supplied pattern over a live stream), and an agent
|
||||
// that assumed otherwise would silently wait on the wrong thing.
|
||||
if (req.query && typeof req.query === 'object' && 'regex' in req.query) {
|
||||
return createErrorResponse(
|
||||
ApiErrorCode.INVALID_INPUT,
|
||||
'regex is not supported; use match=<literal substring> (optionally with nocase=1)'
|
||||
);
|
||||
}
|
||||
|
||||
const query = parseWaitQuery(SessionWaitOutputQuerySchema, req.query, 'wait-output');
|
||||
const session = findSessionOrFail(ctx, id, req);
|
||||
|
||||
// Same reason as /wait: this URL is polled in a loop with identical parameters.
|
||||
reply.header('Cache-Control', 'no-store');
|
||||
|
||||
const timeoutMs = clampWaitMs(query.timeout);
|
||||
const abort = abortOnClientHangUp(reply);
|
||||
const owner = ownerFor(req);
|
||||
// Output waiters are the ones a dead worker strands hardest: the feed simply stops.
|
||||
const stopDeathWatch = watchForDeadWorker(ctx.mux, session, id);
|
||||
|
||||
try {
|
||||
// Check the cap BEFORE touching the buffer. `session.terminalBuffer` is
|
||||
// `BufferAccumulator.value`, which joins the WHOLE accumulator (up to 32MB)
|
||||
// before the slice below takes its tail — so a request that is going to be
|
||||
// rejected anyway must not pay for a full materialization first, or the cap
|
||||
// provides no backpressure at all against a `from=buffer` loop.
|
||||
sessionWaits.assertCapacity(id, owner);
|
||||
|
||||
// `from=buffer` scans what already scrolled past before blocking. Bounded to a
|
||||
// tail: the buffer runs to 32MB and this is a per-request ANSI strip.
|
||||
let initialText: string | undefined;
|
||||
if (query.from === 'buffer') {
|
||||
const buffer = session.terminalBuffer;
|
||||
initialText =
|
||||
buffer.length > MAX_BUFFER_SCAN_BYTES ? buffer.slice(buffer.length - MAX_BUFFER_SCAN_BYTES) : buffer;
|
||||
}
|
||||
|
||||
const result = await sessionWaits.waitForOutput(id, {
|
||||
match: query.match,
|
||||
nocase: query.nocase === '1' || query.nocase === 'true',
|
||||
timeoutMs,
|
||||
owner,
|
||||
abortSignal: abort.signal,
|
||||
initialText,
|
||||
});
|
||||
|
||||
return {
|
||||
success: true,
|
||||
data: {
|
||||
sessionId: id,
|
||||
status: session.status,
|
||||
limitPaused: session.isLimitPaused,
|
||||
// Same envelope as /wait; this one carries `matched`/`snippet`/`match`
|
||||
// where the signal wait carries `signal`/`until`.
|
||||
wait: { ...result, match: query.match },
|
||||
},
|
||||
};
|
||||
} catch (err) {
|
||||
if (err instanceof WaitCapacityError) return waitCapacityResponse(err);
|
||||
throw err;
|
||||
} finally {
|
||||
stopDeathWatch();
|
||||
}
|
||||
return {};
|
||||
});
|
||||
|
||||
// ========== Send Named Key (tmux send-keys -H) ==========
|
||||
@@ -1703,6 +2230,11 @@ export function registerSessionRoutes(
|
||||
.replace(ALT_SCREEN_TOGGLE_PATTERN, '')
|
||||
.replace(ERASE_SCROLLBACK_PATTERN, '')
|
||||
.replace(MOUSE_TRACKING_PATTERN, '');
|
||||
} else if (isMuxAltScreenOnlyStripMode(session.mode, session.usesMux)) {
|
||||
// tmux-backed shell/opencode/antigravity: drop tmux's own client smcup only.
|
||||
// A byte buffer recorded before the live-side strip existed can still carry
|
||||
// it, and one replayed `\x1b[?1049h` re-parks xterm in the alt buffer (#205).
|
||||
strippedBuffer = strippedBuffer.replace(ALT_SCREEN_TOGGLE_PATTERN, '');
|
||||
}
|
||||
|
||||
if (tailBytes > 0 && strippedBuffer.length > tailBytes) {
|
||||
@@ -2027,7 +2559,7 @@ export function registerSessionRoutes(
|
||||
if (!host) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Remote host not found');
|
||||
|
||||
// Per-session config that is applied to the LOCAL tmux/CLI wrapper (env vars via
|
||||
// tmux setenv, effort/model CLI args, codex/gemini/opencode config) does NOT
|
||||
// tmux setenv, effort/model CLI args, codex/gemini/antigravity/opencode config) does NOT
|
||||
// cross ssh, so it would silently no-op. Reject rather than pretend it worked —
|
||||
// remote command/env customization goes through the per-host command override.
|
||||
if (
|
||||
@@ -2364,7 +2896,7 @@ export function registerSessionRoutes(
|
||||
});
|
||||
ctx.broadcast(SseEvent.SessionInteractive, { id: session.id, mode: 'shell' });
|
||||
} else {
|
||||
// 'claude', 'opencode', 'codex', and 'gemini' modes use startInteractive()
|
||||
// 'claude', 'opencode', 'codex', 'gemini', and 'antigravity' modes use startInteractive()
|
||||
await session.startInteractive();
|
||||
getLifecycleLog().log({
|
||||
event: 'started',
|
||||
|
||||
@@ -180,13 +180,19 @@ export function registerWsRoutes(app: FastifyInstance, ctx: SessionPort, getHost
|
||||
const cid = typeof msg.cid === 'string' ? msg.cid : null;
|
||||
const seq = Number.isInteger(msg.seq) ? (msg.seq as number) : null;
|
||||
const apply = cid && seq !== null ? session.shouldApplyInput(cid, seq) : true;
|
||||
let delivered = true;
|
||||
if (apply) {
|
||||
// Typed input from a claim-holding desktop keeps the claim "hot"
|
||||
// and re-asserts the desktop layout after a mobile override.
|
||||
if (holdsDesktopClaim) session.noteDesktopActivity();
|
||||
session.write(msg.d);
|
||||
delivered = session.write(msg.d);
|
||||
// A session whose PTY is gone swallows the write. ACKing anyway told
|
||||
// the client to drop the frame from its durable queue and left the seq
|
||||
// burnt, so the retry that reliable delivery exists for was rejected as
|
||||
// a duplicate: the input was lost for good.
|
||||
if (!delivered && cid && seq !== null) session.forgetInputSeq(cid, seq);
|
||||
}
|
||||
if (seq !== null && socket.readyState === 1) {
|
||||
if (delivered && seq !== null && socket.readyState === 1) {
|
||||
socket.send(`{"t":"ia","seq":${seq}}`);
|
||||
}
|
||||
} else if (
|
||||
|
||||
+62
-1
@@ -17,6 +17,7 @@ import {
|
||||
MIN_TERMINAL_SCROLLBACK_LINES,
|
||||
} from '../config/terminal-history.js';
|
||||
import { MAX_EDITABLE_BYTES } from '../config/file-editing.js';
|
||||
import { MIN_MATCH_LENGTH, MAX_MATCH_LENGTH } from '../config/agent-wait.js';
|
||||
|
||||
// ========== Path Validation ==========
|
||||
|
||||
@@ -516,7 +517,7 @@ export const DockerHostSchema = z.object({
|
||||
mountCredentials: z.boolean().optional(),
|
||||
hooksEnabled: z.boolean().optional(),
|
||||
resumeOnStart: z.boolean().optional(),
|
||||
commands: RemoteCommandOverridesSchema, // same shell/claude/opencode/codex/gemini shape
|
||||
commands: RemoteCommandOverridesSchema, // same shell/claude/opencode/codex/gemini/antigravity shape
|
||||
extraCreateArgs: z
|
||||
.array(
|
||||
z
|
||||
@@ -911,6 +912,66 @@ export const SessionInputWithLimitSchema = z.object({
|
||||
// unset rather than sending null. See docs/reliable-input-delivery.md.
|
||||
seq: z.number().int().nonnegative().optional(),
|
||||
clientId: z.string().max(128).optional(),
|
||||
// Send-and-wait (agent orchestration): `true` for the default signal set, or the
|
||||
// same grammar as `GET .../wait` — a comma string or an array of signals. Absent
|
||||
// means the historical fire-and-forget behavior, byte for byte.
|
||||
//
|
||||
// `.nullish()`, not `.optional()`: a third-party caller building the body with
|
||||
// JSON.stringify keeps an explicit null on the wire, and `.optional()` rejects it
|
||||
// with INVALID_INPUT. That gotcha has shipped as a real bug twice.
|
||||
wait: z.union([z.boolean(), z.string().max(120), z.array(z.string().max(120)).max(8)]).nullish(),
|
||||
// Unbounded above: the effective value is clamped to MAX_WAIT_MS server-side and
|
||||
// returned as `data.wait.timeoutMs`, so a caller that asks for 24h sees what it
|
||||
// actually got. A `.max()` here would turn the same documented clamp into a 400 for
|
||||
// large-enough guesses, which is the one behaviour an agent cannot predict.
|
||||
waitTimeout: z.number().int().positive().nullish(),
|
||||
});
|
||||
|
||||
/**
|
||||
* Query validation for `GET /api/sessions/:id/wait` (agent wait primitives).
|
||||
*
|
||||
* Everything arrives as a string. `timeout` is coerced and bounded here, then
|
||||
* clamped again to the operator's ceiling by `clampWaitMs()` — the schema bound
|
||||
* only keeps an absurd number out of the arithmetic. A non-numeric `timeout` is a
|
||||
* 400 rather than a silent fallback, so an agent never believes it asked for a
|
||||
* longer wait than it got; the value actually applied comes back as
|
||||
* `data.wait.timeoutMs`, which is what makes the clamp observable. `until` is
|
||||
* parsed by `parseWaitSignals()`, which reports unknown tokens instead of
|
||||
* dropping them.
|
||||
*
|
||||
* `until` accepts an ARRAY as well as the comma string: `?until=stop&until=exit`
|
||||
* is how most HTTP clients express a list, Fastify's query parser delivers a
|
||||
* repeated parameter as an array, and `parseWaitSignals()` has always handled
|
||||
* both. Rejecting the repeated form left that branch unreachable and 400'd the
|
||||
* more natural spelling.
|
||||
*/
|
||||
export const SessionWaitQuerySchema = z.object({
|
||||
until: z.union([z.string().max(120), z.array(z.string().max(120)).max(8)]).optional(),
|
||||
// No upper bound on purpose. The contract is "clamped to [MIN_WAIT_MS, MAX_WAIT_MS]",
|
||||
// and a `.max()` here contradicted it: `timeout=99999999` was a 400 mid-fan-out while
|
||||
// `timeout=600001` was silently clamped, so the same documented rule produced two
|
||||
// different outcomes depending on how big the caller's guess was. `clampWaitMs()`
|
||||
// bounds every finite value, and `.int()` still rejects `Infinity`/`1e999` and junk.
|
||||
timeout: z.coerce.number().int().positive().optional(),
|
||||
fresh: z.enum(['0', '1', 'true', 'false']).optional(),
|
||||
});
|
||||
|
||||
/**
|
||||
* Query validation for `GET /api/sessions/:id/wait-output`.
|
||||
*
|
||||
* `match` is a LITERAL substring, never a pattern: `search-service.ts` avoids regex
|
||||
* so there is no ReDoS surface, and this endpoint is more exposed still (the pattern
|
||||
* would be caller-supplied and the input is a live stream). The length bound is a
|
||||
* second reason the carry buffer stays small. The route separately rejects a `regex`
|
||||
* parameter outright rather than ignoring it.
|
||||
*/
|
||||
export const SessionWaitOutputQuerySchema = z.object({
|
||||
match: z.string().min(MIN_MATCH_LENGTH).max(MAX_MATCH_LENGTH),
|
||||
nocase: z.enum(['0', '1', 'true', 'false']).optional(),
|
||||
from: z.enum(['now', 'buffer']).optional(),
|
||||
// Unbounded above for the same reason as SessionWaitQuerySchema.timeout: clamping is
|
||||
// the documented contract, so a large value must clamp rather than 400.
|
||||
timeout: z.coerce.number().int().positive().optional(),
|
||||
});
|
||||
|
||||
// ========== Session Mutation Routes ==========
|
||||
|
||||
@@ -30,6 +30,7 @@ import { homedir, tmpdir } from 'node:os';
|
||||
import { randomUUID } from 'node:crypto';
|
||||
import { createRequire } from 'node:module';
|
||||
import { dataPath } from '../config/instance.js';
|
||||
import { LAUNCHD_LABEL, SYSTEMD_UNIT } from '../config/service-names.js';
|
||||
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
|
||||
import type {
|
||||
InstallInfo,
|
||||
@@ -43,10 +44,9 @@ import type {
|
||||
const require = createRequire(import.meta.url);
|
||||
const { version: APP_VERSION } = require('../../package.json') as { version: string };
|
||||
|
||||
/** systemd unit name (matches install.sh + scripts/codeman-web.service). */
|
||||
const SYSTEMD_UNIT = 'codeman-web.service';
|
||||
/** launchd agent label (matches install.sh setup_launchd_service). */
|
||||
const LAUNCHD_LABEL = 'com.codeman.web';
|
||||
// Unit name / job label live in config/service-names.ts so install.sh, this
|
||||
// detector and `codeman service install` cannot drift apart. Unchanged for the
|
||||
// default instance.
|
||||
/** Path to the persisted update status file. */
|
||||
const STATUS_FILE = dataPath('update-status.json');
|
||||
/** Network/git timeout for the "check" path (longer than EXEC_TIMEOUT_MS — ls-remote hits the network). */
|
||||
|
||||
@@ -85,6 +85,7 @@ import {
|
||||
attachSessionListeners,
|
||||
detachSessionListeners,
|
||||
} from './session-listener-wiring.js';
|
||||
import { sessionWaits } from './session-wait-registry.js';
|
||||
import {
|
||||
wireRespawnListeners,
|
||||
setupTimedRespawn,
|
||||
@@ -805,7 +806,24 @@ export class WebServer extends EventEmitter {
|
||||
const clientId =
|
||||
typeof query.clientId === 'string' && SSE_CLIENT_ID_RE.test(query.clientId) ? query.clientId : undefined;
|
||||
|
||||
// Carry over the headers the security hook already set on this reply.
|
||||
//
|
||||
// writeHead goes straight to the Node response and bypasses Fastify's header
|
||||
// store, so everything the onRequest hook granted is silently dropped —
|
||||
// including the Access-Control-Allow-Origin it emits for localhost origins.
|
||||
// The result is an internal contradiction: a localhost page may call every
|
||||
// /api endpoint cross-origin, but its EventSource fails CORS. The security
|
||||
// headers (nosniff, frame-options, CSP) were lost the same way.
|
||||
//
|
||||
// The other raw-writeHead routes live in file-routes.ts and share a helper;
|
||||
// this one keeps its own copy so the server does not import from a route
|
||||
// module it registers.
|
||||
const inherited: Record<string, number | string | string[]> = {};
|
||||
for (const [name, value] of Object.entries(reply.getHeaders())) {
|
||||
if (value !== undefined) inherited[name] = value;
|
||||
}
|
||||
reply.raw.writeHead(200, {
|
||||
...inherited,
|
||||
'Content-Type': 'text/event-stream',
|
||||
'Cache-Control': 'no-cache',
|
||||
Connection: 'keep-alive',
|
||||
@@ -1230,6 +1248,16 @@ export class WebServer extends EventEmitter {
|
||||
}
|
||||
}
|
||||
|
||||
// Release anything blocked on this session, in the documented order: 'exit'
|
||||
// first so an until=exit caller gets its signal, then cancelAll so everyone
|
||||
// else resolves with ended:true instead of timing out.
|
||||
//
|
||||
// The 'exit' here is NOT redundant with the PTY-exit listener: listeners are
|
||||
// detached a few lines above, before `session.stop()`, so on a delete the
|
||||
// session's own exit event never reaches the registry.
|
||||
sessionWaits.notifySignal(sessionId, 'exit');
|
||||
sessionWaits.cancelAll(sessionId);
|
||||
|
||||
this.broadcast(SseEvent.SessionDeleted, { id: sessionId });
|
||||
}
|
||||
|
||||
@@ -2828,6 +2856,11 @@ export class WebServer extends EventEmitter {
|
||||
// Gracefully close all SSE connections and clear batching state
|
||||
this.sse.stop();
|
||||
|
||||
// Release every pending long-poll waiter. Their timers are deliberately not
|
||||
// unref'd (an unref'd timer can let the process exit mid-wait and strand the
|
||||
// response), so without this a 10-minute wait holds shutdown open.
|
||||
sessionWaits.cancelEverything();
|
||||
|
||||
this.lastRecordedTokens.clear();
|
||||
|
||||
// Stop multiplexer and flush pending saves
|
||||
|
||||
@@ -27,6 +27,7 @@ import type { RalphStatusBlock, CircuitBreakerStatus } from '../types.js';
|
||||
import { SseEvent } from './sse-events.js';
|
||||
import { getLifecycleLog } from '../session-lifecycle-log.js';
|
||||
import { fileStreamManager } from '../file-stream-manager.js';
|
||||
import { sessionWaits } from './session-wait-registry.js';
|
||||
|
||||
/** Stored listener references for session cleanup (prevents memory leaks) */
|
||||
export interface SessionListenerRefs {
|
||||
@@ -92,6 +93,9 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
|
||||
|
||||
/** Batches PTY output → broadcasts `session:terminal` at 16-50ms intervals */
|
||||
terminal: (data) => {
|
||||
// Feeds `GET /api/sessions/:id/wait-output`. No-ops with a single Map lookup
|
||||
// when nothing is waiting, which is the case on virtually every chunk.
|
||||
sessionWaits.notifyOutput(session.id, data);
|
||||
deps.batchTerminalData(session.id, data);
|
||||
},
|
||||
|
||||
@@ -137,6 +141,28 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
|
||||
|
||||
/** Broadcasts `session:exit` + `session:updated` — PTY process exited; cleans up respawn, timers, listeners */
|
||||
exit: (code) => {
|
||||
// Before anything that can throw: a caller blocked on this session must learn
|
||||
// the process died rather than sit until its timeout.
|
||||
//
|
||||
// Both halves are required, in this order — the same pair `_doCleanupSession`
|
||||
// uses on the delete path, for the same reason. `notifySignal` resolves ONLY
|
||||
// waiters that asked for `exit`; everyone else (`until=working`, `until=stop`,
|
||||
// every wait-output) would keep a slot in the process-wide pool until their
|
||||
// timeout, on a session whose feeds this very handler is about to tear down:
|
||||
// `removeSessionListenerRefs` below detaches the `terminal` listener that is
|
||||
// the only input to `notifyOutput`, and the `idle`/`working` listeners with it.
|
||||
// Nothing can reach those waiters afterwards, so holding them is a guaranteed
|
||||
// ten-minute lie. `cancelAll` answers them `ended: true`, which the plan's §3.6
|
||||
// specifies for exactly this case ("Never hang").
|
||||
//
|
||||
// Safe against the respawn cycle: a respawn writes `/clear` + a kickstart
|
||||
// prompt through the mux and never restarts the PTY, so it emits no `exit` and
|
||||
// cannot cancel an orchestrating agent's wait. And for an agent driving a
|
||||
// worker this is the right trade even when the PTY exit was only a tmux
|
||||
// DETACH: `ended` means "re-check and re-issue", one extra round trip, versus
|
||||
// burning the caller's entire timeout learning nothing.
|
||||
sessionWaits.notifySignal(session.id, 'exit');
|
||||
sessionWaits.cancelAll(session.id);
|
||||
getLifecycleLog().log({
|
||||
event: 'exit',
|
||||
sessionId: session.id,
|
||||
@@ -187,6 +213,7 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
|
||||
|
||||
/** Broadcasts `session:working` — Claude started processing */
|
||||
working: () => {
|
||||
sessionWaits.notifySignal(session.id, 'working');
|
||||
deps.broadcast(SseEvent.SessionWorking, { id: session.id });
|
||||
const tracker = deps.getRunSummaryTracker(session.id);
|
||||
if (tracker) {
|
||||
@@ -197,6 +224,7 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
|
||||
|
||||
/** Broadcasts `session:idle` — Claude finished processing, waiting for input */
|
||||
idle: () => {
|
||||
sessionWaits.notifySignal(session.id, 'idle');
|
||||
deps.broadcast(SseEvent.SessionIdle, { id: session.id });
|
||||
deps.broadcastSessionStateDebounced(session.id);
|
||||
const tracker = deps.getRunSummaryTracker(session.id);
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -71,7 +71,8 @@ describe('AiIdleChecker', () => {
|
||||
describe('Output Parsing', () => {
|
||||
it('should parse IDLE verdict', async () => {
|
||||
// Set up mock to return IDLE result after polling
|
||||
mockedReadFileSync.mockReturnValueOnce('') // writeFileSync creates empty file
|
||||
mockedReadFileSync
|
||||
.mockReturnValueOnce('') // writeFileSync creates empty file
|
||||
.mockReturnValueOnce('IDLE\nSession shows completion message and prompt.\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('some terminal output');
|
||||
@@ -87,7 +88,8 @@ describe('AiIdleChecker', () => {
|
||||
});
|
||||
|
||||
it('should parse WORKING verdict', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
mockedReadFileSync
|
||||
.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('WORKING\nSpinner characters detected, still processing.\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('some terminal output');
|
||||
@@ -100,8 +102,7 @@ describe('AiIdleChecker', () => {
|
||||
});
|
||||
|
||||
it('should handle lowercase verdict', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('idle\nDone.\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('idle\nDone.\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(500);
|
||||
@@ -112,7 +113,8 @@ describe('AiIdleChecker', () => {
|
||||
});
|
||||
|
||||
it('should return ERROR for unparseable output', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
mockedReadFileSync
|
||||
.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('Something unexpected happened.\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
@@ -125,8 +127,7 @@ describe('AiIdleChecker', () => {
|
||||
});
|
||||
|
||||
it('should return ERROR for empty output', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(500);
|
||||
@@ -175,10 +176,34 @@ describe('AiIdleChecker', () => {
|
||||
await vi.advanceTimersByTimeAsync(500);
|
||||
await checkPromise;
|
||||
|
||||
expect(mockedWriteFileSync).toHaveBeenCalledWith(
|
||||
expect.stringContaining('codeman-aicheck-'),
|
||||
''
|
||||
expect(mockedWriteFileSync).toHaveBeenCalledWith(expect.stringContaining('codeman-aicheck-'), '');
|
||||
});
|
||||
|
||||
it('should keep Claude stderr separate from verdict output', async () => {
|
||||
mockedReadFileSync.mockReturnValue('IDLE\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(500);
|
||||
await checkPromise;
|
||||
|
||||
const spawnArgs = mockedSpawn.mock.calls[0]?.[1];
|
||||
const command = spawnArgs?.[spawnArgs.length - 1];
|
||||
expect(command).toEqual(expect.any(String));
|
||||
expect(command).toContain(' 2> "');
|
||||
expect(command).not.toContain('2>&1');
|
||||
});
|
||||
|
||||
it('should include Claude stderr when no verdict is produced', async () => {
|
||||
mockedReadFileSync.mockImplementation((path) =>
|
||||
String(path).includes('-stderr-') ? 'Claude CLI failed to load settings' : '__AICHECK_DONE__'
|
||||
);
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(500);
|
||||
|
||||
const result = await checkPromise;
|
||||
expect(result.verdict).toBe('ERROR');
|
||||
expect(result.reasoning).toContain('Claude CLI failed to load settings');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -223,7 +248,7 @@ describe('AiIdleChecker', () => {
|
||||
|
||||
// Should have tried to kill the tmux session (initial kill + cleanup kill)
|
||||
const killCalls = mockedExecSync.mock.calls.filter(
|
||||
call => typeof call[0] === 'string' && call[0].includes('kill-session')
|
||||
(call) => typeof call[0] === 'string' && call[0].includes('kill-session')
|
||||
);
|
||||
expect(killCalls.length).toBeGreaterThan(0);
|
||||
});
|
||||
@@ -236,8 +261,7 @@ describe('AiIdleChecker', () => {
|
||||
|
||||
describe('Cooldown', () => {
|
||||
it('should start cooldown after WORKING verdict', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('WORKING\nStill processing.\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nStill processing.\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(500);
|
||||
@@ -250,8 +274,7 @@ describe('AiIdleChecker', () => {
|
||||
});
|
||||
|
||||
it('should return to ready after cooldown expires', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
@@ -267,8 +290,7 @@ describe('AiIdleChecker', () => {
|
||||
});
|
||||
|
||||
it('should not start new check during cooldown', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
|
||||
|
||||
const firstCheck = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
@@ -283,8 +305,7 @@ describe('AiIdleChecker', () => {
|
||||
|
||||
describe('Error Handling', () => {
|
||||
it('should start error cooldown after parse error', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('garbage output\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage output\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
@@ -302,8 +323,7 @@ describe('AiIdleChecker', () => {
|
||||
const cooldowns = [1100, 2100]; // Wait slightly longer than each cooldown
|
||||
|
||||
for (let i = 0; i < 3; i++) {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('garbage\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
@@ -321,8 +341,7 @@ describe('AiIdleChecker', () => {
|
||||
|
||||
it('should reset error counter on successful check', async () => {
|
||||
// First check: error
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('garbage\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage\n__AICHECK_DONE__');
|
||||
const firstCheck = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
await firstCheck;
|
||||
@@ -332,8 +351,7 @@ describe('AiIdleChecker', () => {
|
||||
await vi.advanceTimersByTimeAsync(1100);
|
||||
|
||||
// Second check: success
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('IDLE\nDone.\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\nDone.\n__AICHECK_DONE__');
|
||||
const secondCheck = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
await secondCheck;
|
||||
@@ -352,8 +370,7 @@ describe('AiIdleChecker', () => {
|
||||
|
||||
describe('Buffer Handling', () => {
|
||||
it('should strip ANSI codes from terminal buffer', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
|
||||
|
||||
const ansiBuffer = '\x1b[1mBold\x1b[0m \x1b[32mGreen\x1b[0m text';
|
||||
const checkPromise = checker.check(ansiBuffer);
|
||||
@@ -365,8 +382,7 @@ describe('AiIdleChecker', () => {
|
||||
});
|
||||
|
||||
it('should trim buffer to maxContextChars', async () => {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\n__AICHECK_DONE__');
|
||||
|
||||
// Create buffer longer than maxContextChars (1000)
|
||||
const longBuffer = 'x'.repeat(2000);
|
||||
@@ -402,8 +418,7 @@ describe('AiIdleChecker', () => {
|
||||
describe('Reset', () => {
|
||||
it('should clear all state on reset', async () => {
|
||||
// Trigger a WORKING verdict to set state
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
@@ -440,24 +455,24 @@ describe('AiIdleChecker', () => {
|
||||
const handler = vi.fn();
|
||||
checker.on('checkCompleted', handler);
|
||||
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('IDLE\nAll done.\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('IDLE\nAll done.\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
await checkPromise;
|
||||
|
||||
expect(handler).toHaveBeenCalledWith(expect.objectContaining({
|
||||
verdict: 'IDLE',
|
||||
}));
|
||||
expect(handler).toHaveBeenCalledWith(
|
||||
expect.objectContaining({
|
||||
verdict: 'IDLE',
|
||||
})
|
||||
);
|
||||
});
|
||||
|
||||
it('should emit cooldownStarted event after WORKING', async () => {
|
||||
const handler = vi.fn();
|
||||
checker.on('cooldownStarted', handler);
|
||||
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('WORKING\nBusy.\n__AICHECK_DONE__');
|
||||
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
@@ -477,8 +492,7 @@ describe('AiIdleChecker', () => {
|
||||
const cooldowns = [1100, 2100]; // Wait longer than exponential backoff
|
||||
|
||||
for (let i = 0; i < 3; i++) {
|
||||
mockedReadFileSync.mockReturnValueOnce('')
|
||||
.mockReturnValueOnce('garbage\n__AICHECK_DONE__');
|
||||
mockedReadFileSync.mockReturnValueOnce('').mockReturnValueOnce('garbage\n__AICHECK_DONE__');
|
||||
const checkPromise = checker.check('output');
|
||||
await vi.advanceTimersByTimeAsync(1000);
|
||||
await checkPromise;
|
||||
|
||||
@@ -0,0 +1,109 @@
|
||||
/**
|
||||
* Issue #205, round 2: `getClaudeCliVersion()` used to cache FAILURE forever.
|
||||
*
|
||||
* It stored `null` on any exception and guarded on `!== undefined`, so a single
|
||||
* failed probe — the 5s exec timeout, a PATH-starved systemd/launchd
|
||||
* environment, a transient fs hiccup — at the first Claude session start left
|
||||
* `cliVersion` undefined for every Claude session until the server restarted.
|
||||
* An undefined `cliVersion` silently disables wheel-forwarding to Claude's own
|
||||
* transcript (`_shouldForwardWheelToApp`), which is the only route to history
|
||||
* for a repaint-mode pane: a dead wheel on every device at once, which is what
|
||||
* the reporter described (phone + iPad + laptop all broken together points at a
|
||||
* SERVER-side cause, not a browser one).
|
||||
*
|
||||
* The probe itself can't run under vitest (it would spawn a real `claude`), so
|
||||
* these drive the cache policy directly with an injected probe and clock.
|
||||
*/
|
||||
import { describe, expect, it, vi } from 'vitest';
|
||||
import {
|
||||
claudeVersionRetryDelayMs,
|
||||
getClaudeCliVersion,
|
||||
resolveClaudeCliVersion,
|
||||
type ClaudeVersionProbeState,
|
||||
} from '../src/utils/claude-cli-resolver.js';
|
||||
|
||||
const freshState = (): ClaudeVersionProbeState => ({ failures: 0, lastFailureAt: 0 });
|
||||
|
||||
describe('claude --version probe caching', () => {
|
||||
it('probes once on success and never spawns again', () => {
|
||||
const state = freshState();
|
||||
const probe = vi.fn(() => '2.1.223');
|
||||
|
||||
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBe('2.1.223');
|
||||
expect(resolveClaudeCliVersion(state, 2_000, probe)).toBe('2.1.223');
|
||||
expect(resolveClaudeCliVersion(state, 9_999_999, probe)).toBe('2.1.223');
|
||||
expect(probe).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
|
||||
it('RETRIES after a failed probe instead of poisoning the process', () => {
|
||||
const state = freshState();
|
||||
const probe = vi
|
||||
.fn<() => string | null>()
|
||||
.mockImplementationOnce(() => {
|
||||
throw new Error('spawn claude ETIMEDOUT'); // the shipped failure mode
|
||||
})
|
||||
.mockImplementationOnce(() => '2.1.223');
|
||||
|
||||
// First session start: probe blows up, no version.
|
||||
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBeNull();
|
||||
// Immediately after, the negative cache holds — no probe storm.
|
||||
expect(resolveClaudeCliVersion(state, 30_000, probe)).toBeNull();
|
||||
expect(probe).toHaveBeenCalledTimes(1);
|
||||
|
||||
// Once the retry window elapses, the next session start probes again and
|
||||
// wheel-forwarding comes back without a server restart.
|
||||
expect(resolveClaudeCliVersion(state, 61_000, probe)).toBe('2.1.223');
|
||||
expect(probe).toHaveBeenCalledTimes(2);
|
||||
});
|
||||
|
||||
it('treats an unparseable version like a failure (retryable, not cached)', () => {
|
||||
const state = freshState();
|
||||
const probe = vi.fn<() => string | null>(() => null); // e.g. output without a x.y.z
|
||||
|
||||
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBeNull();
|
||||
expect(resolveClaudeCliVersion(state, 61_000, probe)).toBeNull();
|
||||
expect(probe).toHaveBeenCalledTimes(2);
|
||||
expect(state.version).toBeUndefined(); // nothing cached as "known bad"
|
||||
});
|
||||
|
||||
it('clears the failure streak once a probe succeeds', () => {
|
||||
const state = freshState();
|
||||
const probe = vi
|
||||
.fn<() => string | null>()
|
||||
.mockImplementationOnce(() => null)
|
||||
.mockImplementationOnce(() => '2.1.223');
|
||||
|
||||
resolveClaudeCliVersion(state, 1_000, probe);
|
||||
expect(state.failures).toBe(1);
|
||||
resolveClaudeCliVersion(state, 61_000, probe);
|
||||
expect(state.failures).toBe(0);
|
||||
expect(state.lastFailureAt).toBe(0);
|
||||
});
|
||||
|
||||
it('backs off so a genuinely missing binary cannot probe on every session start', () => {
|
||||
expect(claudeVersionRetryDelayMs(0)).toBe(0);
|
||||
expect(claudeVersionRetryDelayMs(1)).toBe(60_000);
|
||||
expect(claudeVersionRetryDelayMs(2)).toBe(120_000);
|
||||
expect(claudeVersionRetryDelayMs(3)).toBe(240_000);
|
||||
// Capped, so it keeps retrying forever without ever spinning.
|
||||
expect(claudeVersionRetryDelayMs(50)).toBe(15 * 60_000);
|
||||
|
||||
const state = freshState();
|
||||
const probe = vi.fn<() => string | null>(() => null);
|
||||
resolveClaudeCliVersion(state, 0, probe); // failure 1 → retry at 60s
|
||||
resolveClaudeCliVersion(state, 30_000, probe); // still inside the window
|
||||
expect(probe).toHaveBeenCalledTimes(1);
|
||||
resolveClaudeCliVersion(state, 60_000, probe); // failure 2 → retry at 120s
|
||||
resolveClaudeCliVersion(state, 119_000, probe);
|
||||
expect(probe).toHaveBeenCalledTimes(2);
|
||||
resolveClaudeCliVersion(state, 180_001, probe);
|
||||
expect(probe).toHaveBeenCalledTimes(3);
|
||||
});
|
||||
|
||||
it('stays hermetic under vitest without recording a phantom failure', () => {
|
||||
// The guard returns before the probe, and — unlike the old code, which wrote
|
||||
// null into the cache here — leaves the cache untouched.
|
||||
expect(getClaudeCliVersion()).toBeNull();
|
||||
expect(getClaudeCliVersion()).toBeNull();
|
||||
});
|
||||
});
|
||||
@@ -1,5 +1,5 @@
|
||||
import { describe, expect, it } from 'vitest';
|
||||
import { Session, isAltScreenStripMode } from '../src/session.js';
|
||||
import { Session, isAltScreenStripMode, isMuxAltScreenOnlyStripMode } from '../src/session.js';
|
||||
|
||||
type SessionInternals = {
|
||||
_handleTerminalOutput(data: string): void;
|
||||
@@ -81,7 +81,7 @@ describe('Claude terminal scrollback strip', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('Shell terminal output is NOT stripped (vim/less/htop need the alt screen)', () => {
|
||||
describe('Shell terminal output on a DIRECT PTY is NOT stripped (vim/less/htop need the alt screen)', () => {
|
||||
it('leaves alt-screen toggles, scrollback-erase, and mouse-tracking intact for shell', () => {
|
||||
const session = new Session({ workingDir: '/tmp', mode: 'shell' });
|
||||
|
||||
@@ -91,3 +91,59 @@ describe('Shell terminal output is NOT stripped (vim/less/htop need the alt scre
|
||||
expect(session.terminalBuffer).toBe(vimLike);
|
||||
});
|
||||
});
|
||||
|
||||
describe('isMuxAltScreenOnlyStripMode', () => {
|
||||
it('covers exactly the modes the full strip does not, and only under tmux', () => {
|
||||
for (const mode of ['shell', 'opencode', 'antigravity'] as const) {
|
||||
expect(isMuxAltScreenOnlyStripMode(mode, true)).toBe(true);
|
||||
// Direct-PTY fallback: the program's own alt screen really does reach xterm.
|
||||
expect(isMuxAltScreenOnlyStripMode(mode, false)).toBe(false);
|
||||
}
|
||||
// The full strip already owns these; never double-gate them here.
|
||||
for (const mode of ['claude', 'codex', 'gemini'] as const) {
|
||||
expect(isMuxAltScreenOnlyStripMode(mode, true)).toBe(false);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('tmux-backed shell: strip tmux’s own client smcup, keep everything else (#205)', () => {
|
||||
it('drops alt-screen toggles so xterm keeps a scrollback buffer', () => {
|
||||
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
|
||||
|
||||
// What a real `tmux attach` emits as its first bytes.
|
||||
handleOutput(session, '\x1b[?1049h\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2Jprompt$ ');
|
||||
|
||||
expect(session.terminalBuffer).toBe('\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2Jprompt$ ');
|
||||
expect(session.terminalBuffer).not.toContain('\x1b[?1049h');
|
||||
});
|
||||
|
||||
it('KEEPS 3J and mouse-tracking, unlike the full strip', () => {
|
||||
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
|
||||
|
||||
// `clear` legitimately wipes scrollback; htop/vim mouse modes are passed
|
||||
// through by tmux even with `mouse off` and must keep working.
|
||||
handleOutput(session, '\x1b[3J\x1b[?1002h\x1b[?1006hhtop\x1b[?1006l\x1b[?1002l');
|
||||
|
||||
expect(session.terminalBuffer).toBe('\x1b[3J\x1b[?1002h\x1b[?1006hhtop\x1b[?1006l\x1b[?1002l');
|
||||
});
|
||||
|
||||
it('reassembles alt-screen sequences split across PTY chunk boundaries', () => {
|
||||
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
|
||||
const emitted: string[] = [];
|
||||
session.on('terminal', (data) => emitted.push(data));
|
||||
|
||||
handleOutput(session, 'before\x1b[?104');
|
||||
handleOutput(session, '9h after');
|
||||
|
||||
expect(session.terminalBuffer).toBe('before after');
|
||||
expect(emitted).toEqual(['before', ' after']);
|
||||
});
|
||||
|
||||
it('applies to opencode and antigravity too', () => {
|
||||
for (const mode of ['opencode', 'antigravity'] as const) {
|
||||
const session = new Session({ workingDir: '/tmp', mode, useMux: true });
|
||||
handleOutput(session, '\x1b[?1049hTUI\x1b[3J');
|
||||
expect(session.terminalBuffer).toBe('TUI\x1b[3J');
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
@@ -0,0 +1,182 @@
|
||||
/**
|
||||
* Unit tests for the pure halves of daemon-control (issue #231): argv rebuilding,
|
||||
* the readiness URL, pidfile parsing, the stale-pid identity check, and the
|
||||
* `/api/status` probe against a real socket.
|
||||
*/
|
||||
|
||||
import { describe, it, expect, afterAll, beforeAll } from 'vitest';
|
||||
import http from 'node:http';
|
||||
import {
|
||||
buildBaseUrl,
|
||||
buildStatusUrl,
|
||||
buildWebArgs,
|
||||
isProcessAlive,
|
||||
looksLikeCodemanWeb,
|
||||
parsePidFileContents,
|
||||
probeServer,
|
||||
} from '../src/daemon-control.js';
|
||||
|
||||
const PORT = 3216;
|
||||
|
||||
describe('buildWebArgs', () => {
|
||||
it('always passes host and port through explicitly', () => {
|
||||
expect(buildWebArgs({ host: '127.0.0.1', port: 3000, https: false })).toEqual([
|
||||
'web',
|
||||
'--host',
|
||||
'127.0.0.1',
|
||||
'--port',
|
||||
'3000',
|
||||
]);
|
||||
});
|
||||
|
||||
it('forwards every optional flag it was given', () => {
|
||||
const args = buildWebArgs({
|
||||
host: '0.0.0.0',
|
||||
port: 8080,
|
||||
https: true,
|
||||
titleHostname: 'tower',
|
||||
allowUnauthenticatedNetwork: true,
|
||||
multiuser: true,
|
||||
});
|
||||
expect(args).toEqual([
|
||||
'web',
|
||||
'--host',
|
||||
'0.0.0.0',
|
||||
'--port',
|
||||
'8080',
|
||||
'--https',
|
||||
'--title-hostname',
|
||||
'tower',
|
||||
'--allow-unauthenticated-network',
|
||||
'--multiuser',
|
||||
]);
|
||||
});
|
||||
|
||||
it('never re-emits the daemon flags themselves (the child must not re-fork)', () => {
|
||||
const args = buildWebArgs({ host: '127.0.0.1', port: 3000, https: false });
|
||||
expect(args).not.toContain('--daemon');
|
||||
expect(args).not.toContain('-d');
|
||||
});
|
||||
});
|
||||
|
||||
describe('buildBaseUrl', () => {
|
||||
it('is the address a browser can open, with no path on it', () => {
|
||||
expect(buildBaseUrl({ host: '127.0.0.1', port: 3000, https: false })).toBe('http://127.0.0.1:3000');
|
||||
expect(buildBaseUrl({ host: '0.0.0.0', port: 8443, https: true })).toBe('https://127.0.0.1:8443');
|
||||
});
|
||||
});
|
||||
|
||||
describe('buildStatusUrl', () => {
|
||||
it('uses http by default and https when asked', () => {
|
||||
expect(buildStatusUrl({ host: '127.0.0.1', port: 3000, https: false })).toBe('http://127.0.0.1:3000/api/status');
|
||||
expect(buildStatusUrl({ host: '127.0.0.1', port: 3000, https: true })).toBe('https://127.0.0.1:3000/api/status');
|
||||
});
|
||||
|
||||
it('rewrites wildcard binds to loopback, since they are not connectable', () => {
|
||||
expect(buildStatusUrl({ host: '0.0.0.0', port: 3000, https: false })).toBe('http://127.0.0.1:3000/api/status');
|
||||
expect(buildStatusUrl({ host: '::', port: 3000, https: false })).toBe('http://127.0.0.1:3000/api/status');
|
||||
});
|
||||
|
||||
it('brackets a bare IPv6 literal', () => {
|
||||
expect(buildStatusUrl({ host: '::1', port: 3000, https: false })).toBe('http://[::1]:3000/api/status');
|
||||
expect(buildStatusUrl({ host: '[::1]', port: 3000, https: false })).toBe('http://[::1]:3000/api/status');
|
||||
});
|
||||
});
|
||||
|
||||
describe('parsePidFileContents', () => {
|
||||
it('accepts a plain pid with surrounding whitespace', () => {
|
||||
expect(parsePidFileContents('4242\n')).toBe(4242);
|
||||
expect(parsePidFileContents(' 4242 ')).toBe(4242);
|
||||
});
|
||||
|
||||
it('rejects garbage, empties and floats', () => {
|
||||
expect(parsePidFileContents('')).toBeNull();
|
||||
expect(parsePidFileContents('not a pid')).toBeNull();
|
||||
expect(parsePidFileContents('42.5')).toBeNull();
|
||||
expect(parsePidFileContents('-42')).toBeNull();
|
||||
});
|
||||
|
||||
it('rejects pid 0 and pid 1: neither is ever our server', () => {
|
||||
expect(parsePidFileContents('0')).toBeNull();
|
||||
expect(parsePidFileContents('1')).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
describe('looksLikeCodemanWeb', () => {
|
||||
it('matches the ways the server is actually launched', () => {
|
||||
expect(looksLikeCodemanWeb('/usr/bin/node /home/u/.codeman/app/dist/index.js web')).toBe(true);
|
||||
expect(looksLikeCodemanWeb('/usr/bin/node dist/index.js web --https')).toBe(true);
|
||||
expect(looksLikeCodemanWeb('node /repo/src/index.ts web --port 3000')).toBe(true);
|
||||
expect(looksLikeCodemanWeb('/opt/homebrew/bin/codeman web')).toBe(true);
|
||||
expect(looksLikeCodemanWeb('aicodeman web --host 0.0.0.0')).toBe(true);
|
||||
});
|
||||
|
||||
it('rejects anything that inherited a recycled pid', () => {
|
||||
expect(looksLikeCodemanWeb(null)).toBe(false);
|
||||
expect(looksLikeCodemanWeb('')).toBe(false);
|
||||
expect(looksLikeCodemanWeb('/usr/bin/node dist/index.js session list')).toBe(false);
|
||||
expect(looksLikeCodemanWeb('vim web')).toBe(false);
|
||||
expect(looksLikeCodemanWeb('/usr/lib/systemd/systemd --user')).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('isProcessAlive', () => {
|
||||
it('sees this very process', () => {
|
||||
expect(isProcessAlive(process.pid)).toBe(true);
|
||||
});
|
||||
|
||||
it('does not see an unused high pid', () => {
|
||||
// 2^22 is above the default pid_max on Linux and macOS.
|
||||
expect(isProcessAlive(4_194_303)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('probeServer', () => {
|
||||
let server: http.Server;
|
||||
|
||||
beforeAll(async () => {
|
||||
server = http.createServer((req, res) => {
|
||||
if (req.url === '/unauthorized') {
|
||||
res.writeHead(401).end('Unauthorized');
|
||||
return;
|
||||
}
|
||||
if (req.url === '/foreign') {
|
||||
res.writeHead(200, { 'Content-Type': 'text/html' }).end('<html>some other app</html>');
|
||||
return;
|
||||
}
|
||||
res.writeHead(200, { 'Content-Type': 'application/json' });
|
||||
res.end(JSON.stringify({ success: true, data: { version: '9.9.9' } }));
|
||||
});
|
||||
await new Promise<void>((resolve) => server.listen(PORT, '127.0.0.1', resolve));
|
||||
});
|
||||
|
||||
afterAll(async () => {
|
||||
await new Promise<void>((resolve) => server.close(() => resolve()));
|
||||
});
|
||||
|
||||
it('reports up and reads the version back', async () => {
|
||||
const result = await probeServer(`http://127.0.0.1:${PORT}/api/status`);
|
||||
expect(result.up).toBe(true);
|
||||
expect(result.version).toBe('9.9.9');
|
||||
});
|
||||
|
||||
it('counts a 401 as up, because auth being active proves a server is there', async () => {
|
||||
const result = await probeServer(`http://127.0.0.1:${PORT}/unauthorized`);
|
||||
expect(result.up).toBe(true);
|
||||
});
|
||||
|
||||
it('does not mistake an unrelated service squatting on the port for Codeman', async () => {
|
||||
const result = await probeServer(`http://127.0.0.1:${PORT}/foreign`);
|
||||
expect(result.up).toBe(false);
|
||||
});
|
||||
|
||||
it('reports down when nothing is listening', async () => {
|
||||
const result = await probeServer(`http://127.0.0.1:${PORT + 1}/api/status`, 1000);
|
||||
expect(result.up).toBe(false);
|
||||
});
|
||||
|
||||
it('reports down for a malformed url instead of throwing', async () => {
|
||||
const result = await probeServer('not-a-url');
|
||||
expect(result.up).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -20,6 +20,15 @@ describe('frontend public asset tooling', () => {
|
||||
expect(appJs.includes(0)).toBe(false);
|
||||
});
|
||||
|
||||
it('uses the same message wrapper for brief and full response views', () => {
|
||||
const appJs = readFileSync(resolve(repoRoot, 'src/web/public/app.js'), 'utf8');
|
||||
|
||||
expect(appJs).toContain("body.appendChild(this._buildResponseViewerMessage(lastResponse, 'assistant'");
|
||||
expect(appJs).toContain('body.appendChild(this._buildResponseViewerMessage(msg.text, msg.role, agentLabel));');
|
||||
expect(appJs).toContain("div.className = 'rv-message ' + (isUser ? 'rv-msg-user' : 'rv-msg-assistant');");
|
||||
expect(appJs).toContain("renderedText.className = 'rv-text';");
|
||||
});
|
||||
|
||||
it('runs the public asset check script', () => {
|
||||
expect(() => {
|
||||
execFileSync('npm', ['run', 'check:public-assets', '--silent'], {
|
||||
|
||||
@@ -127,12 +127,31 @@ describe('refreshStaleCodemanHooks', () => {
|
||||
|
||||
const after = JSON.parse(readFileSync(settingsPath, 'utf-8'));
|
||||
expect(JSON.stringify(after.hooks)).toContain(SECRET_HEADER);
|
||||
expect(JSON.stringify(after.hooks)).toContain('CODEMAN_BACKGROUND_REWAKE_V');
|
||||
expect(JSON.stringify(after.hooks)).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
|
||||
expect(JSON.stringify(after.hooks.Stop)).toContain('./notify-user.sh');
|
||||
expect(after.hooks.PostToolUse).toEqual(expect.arrayContaining([customPostToolUse]));
|
||||
expect(after.hooks.CustomEvent).toEqual(customEvent);
|
||||
});
|
||||
|
||||
// A case can be current on the secret AND the background-wake hook and still carry
|
||||
// the `-k`-less curl shape, which exits 60 against a self-signed HTTPS API and is
|
||||
// swallowed by `|| true` — every hook event dead, silently. The refresh must treat
|
||||
// that as a third stale shape.
|
||||
it('heals a current-looking block whose hook curls lack -k (HTTPS self-signed installs)', async () => {
|
||||
const { generateHooksConfig } = await import('../src/hooks-config.js');
|
||||
const flagless = JSON.parse(JSON.stringify(generateHooksConfig()).replaceAll('curl -sk ', 'curl -s '));
|
||||
writeFileSync(settingsPath, JSON.stringify({ hooks: flagless.hooks }, null, 2));
|
||||
|
||||
await refreshStaleCodemanHooks(dir);
|
||||
|
||||
const after = readFileSync(settingsPath, 'utf-8');
|
||||
expect(after).toContain('curl -sk -X POST');
|
||||
expect(after).not.toContain('curl -s -X POST');
|
||||
// and the pass is convergent: a second refresh must not rewrite
|
||||
await refreshStaleCodemanHooks(dir);
|
||||
expect(readFileSync(settingsPath, 'utf-8')).toBe(after);
|
||||
});
|
||||
|
||||
it('is a no-op when settings.local.json is absent (does not create one)', async () => {
|
||||
await refreshStaleCodemanHooks(dir);
|
||||
expect(existsSync(settingsPath)).toBe(false);
|
||||
|
||||
+283
-3
@@ -6,13 +6,15 @@
|
||||
*/
|
||||
|
||||
import { describe, it, expect, beforeAll, beforeEach, afterAll, afterEach } from 'vitest';
|
||||
import { existsSync, readFileSync, writeFileSync, mkdirSync, rmSync } from 'node:fs';
|
||||
import { closeSync, existsSync, openSync, readFileSync, writeFileSync, mkdirSync, rmSync } from 'node:fs';
|
||||
import { join } from 'node:path';
|
||||
import { tmpdir } from 'node:os';
|
||||
import { spawn } from 'node:child_process';
|
||||
import {
|
||||
ensureCodemanHooks,
|
||||
generateBackgroundWakeScript,
|
||||
generateHooksConfig,
|
||||
generateSubagentStopGuardScript,
|
||||
refreshStaleCodemanHooks,
|
||||
writeHooksConfig,
|
||||
} from '../src/hooks-config.js';
|
||||
@@ -35,6 +37,20 @@ describe('generateHooksConfig', () => {
|
||||
expect(config.hooks.Stop).toHaveLength(1);
|
||||
});
|
||||
|
||||
it('should guard subagent stops while their background work is active', () => {
|
||||
const config = generateHooksConfig();
|
||||
const subagentHooks = config.hooks.SubagentStop as Array<{
|
||||
hooks: Array<{ type: string; command: string; args: string[]; timeout: number }>;
|
||||
}>;
|
||||
|
||||
expect(subagentHooks).toHaveLength(1);
|
||||
expect(subagentHooks[0].hooks[0]).toMatchObject({
|
||||
type: 'command',
|
||||
command: 'node',
|
||||
args: ['-e', generateSubagentStopGuardScript()],
|
||||
});
|
||||
});
|
||||
|
||||
it('should configure a self-contained Bash background-task rewake hook', () => {
|
||||
const config = generateHooksConfig();
|
||||
const postToolHooks = config.hooks.PostToolUse as Array<{
|
||||
@@ -95,6 +111,17 @@ describe('generateHooksConfig', () => {
|
||||
expect(notifHooks[0].hooks[0].command).toContain('|| true');
|
||||
});
|
||||
|
||||
// On --https/tailscale installs CODEMAN_API_URL is HTTPS with a self-signed cert.
|
||||
// A `-k`-less hook curl exits 60 there, the `|| true` swallows it, and every hook
|
||||
// event (stop, permission_prompt, elicitation_dialog, idle_prompt, teammate_idle,
|
||||
// task_completed) dies silently — killing respawn's idle signals and the wait
|
||||
// endpoints' stop/blocked. The statusline exporter always carried -k; the hooks must too.
|
||||
it('every hook curl tolerates a self-signed HTTPS API (curl -sk)', () => {
|
||||
const serialized = JSON.stringify(generateHooksConfig());
|
||||
expect(serialized).toContain('curl -sk -X POST');
|
||||
expect(serialized).not.toContain('curl -s -X POST');
|
||||
});
|
||||
|
||||
it('should set timeout to 10 seconds (hook timeout fields are seconds)', () => {
|
||||
const config = generateHooksConfig();
|
||||
const notifHooks = config.hooks.Notification as Array<{ hooks: Array<{ timeout: number }> }>;
|
||||
@@ -199,7 +226,8 @@ describe('writeHooksConfig', () => {
|
||||
|
||||
const parsed = JSON.parse(readFileSync(settingsPath, 'utf-8'));
|
||||
expect(parsed.hooks.PostToolUse).toHaveLength(1);
|
||||
expect(JSON.stringify(parsed.hooks.PostToolUse)).toContain('CODEMAN_BACKGROUND_REWAKE_V');
|
||||
expect(JSON.stringify(parsed.hooks.PostToolUse)).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
|
||||
expect(JSON.stringify(parsed.hooks.SubagentStop)).toContain('CODEMAN_SUBAGENT_STOP_GUARD_V1');
|
||||
});
|
||||
|
||||
it('should replace an older rewake script version without duplicating it', async () => {
|
||||
@@ -231,10 +259,29 @@ describe('writeHooksConfig', () => {
|
||||
const serialized = JSON.stringify(parsed.hooks.PostToolUse);
|
||||
expect(parsed.hooks.PostToolUse).toHaveLength(1);
|
||||
expect(parsed.hooks.PostToolUse[0].hooks).toHaveLength(1);
|
||||
expect(serialized).toContain('CODEMAN_BACKGROUND_REWAKE_V2');
|
||||
expect(serialized).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
|
||||
expect(serialized).not.toContain('CODEMAN_BACKGROUND_REWAKE_V1');
|
||||
});
|
||||
|
||||
it('replaces the V2 background hook without duplicating it', async () => {
|
||||
const claudeDir = join(testDir, '.claude');
|
||||
const settingsPath = join(claudeDir, 'settings.local.json');
|
||||
mkdirSync(claudeDir, { recursive: true });
|
||||
const oldSettings = JSON.stringify({ hooks: generateHooksConfig().hooks }, null, 2).replaceAll(
|
||||
'CODEMAN_BACKGROUND_REWAKE_V3',
|
||||
'CODEMAN_BACKGROUND_REWAKE_V2'
|
||||
);
|
||||
writeFileSync(settingsPath, oldSettings);
|
||||
|
||||
await refreshStaleCodemanHooks(testDir);
|
||||
|
||||
const parsed = JSON.parse(readFileSync(settingsPath, 'utf-8'));
|
||||
const postToolUse = JSON.stringify(parsed.hooks.PostToolUse);
|
||||
expect(parsed.hooks.PostToolUse).toHaveLength(1);
|
||||
expect(postToolUse).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
|
||||
expect(postToolUse).not.toContain('CODEMAN_BACKGROUND_REWAKE_V2');
|
||||
});
|
||||
|
||||
it('should not add rewake hooks to a user-owned hook configuration', async () => {
|
||||
const claudeDir = join(testDir, '.claude');
|
||||
const settingsPath = join(claudeDir, 'settings.local.json');
|
||||
@@ -267,6 +314,35 @@ describe('writeHooksConfig', () => {
|
||||
expect(parsed.hooks.Notification).toBeDefined();
|
||||
});
|
||||
|
||||
it('should safely add Codeman hooks to an existing managed-case settings file', async () => {
|
||||
const claudeDir = join(testDir, '.claude');
|
||||
const settingsPath = join(claudeDir, 'settings.local.json');
|
||||
mkdirSync(claudeDir, { recursive: true });
|
||||
const userHooks = {
|
||||
PostToolUse: [{ matcher: 'Write', hooks: [{ type: 'command', command: './format.sh' }] }],
|
||||
};
|
||||
writeFileSync(settingsPath, JSON.stringify({ hooks: userHooks, permissions: { allow: ['Read'] } }, null, 2));
|
||||
|
||||
await ensureCodemanHooks(testDir);
|
||||
|
||||
const parsed = JSON.parse(readFileSync(settingsPath, 'utf-8'));
|
||||
expect(parsed.permissions).toEqual({ allow: ['Read'] });
|
||||
expect(parsed.hooks.PostToolUse).toEqual(expect.arrayContaining(userHooks.PostToolUse));
|
||||
expect(JSON.stringify(parsed.hooks)).toContain('CODEMAN_BACKGROUND_REWAKE_V3');
|
||||
expect(JSON.stringify(parsed.hooks)).toContain('CODEMAN_SUBAGENT_STOP_GUARD_V1');
|
||||
});
|
||||
|
||||
it('should not replace a malformed managed-case settings file', async () => {
|
||||
const claudeDir = join(testDir, '.claude');
|
||||
const settingsPath = join(claudeDir, 'settings.local.json');
|
||||
mkdirSync(claudeDir, { recursive: true });
|
||||
writeFileSync(settingsPath, '{ malformed');
|
||||
|
||||
await ensureCodemanHooks(testDir);
|
||||
|
||||
expect(readFileSync(settingsPath, 'utf-8')).toBe('{ malformed');
|
||||
});
|
||||
|
||||
it('should handle malformed existing settings.local.json', async () => {
|
||||
const claudeDir = join(testDir, '.claude');
|
||||
mkdirSync(claudeDir, { recursive: true });
|
||||
@@ -359,6 +435,210 @@ describe('background task rewake helper', () => {
|
||||
expect(result.stderr).toContain('completed');
|
||||
expect(result.stderr).toContain('/tmp/bg-test-1.output');
|
||||
});
|
||||
|
||||
it('rewakes a subagent when Claude queues completion in the parent transcript', async () => {
|
||||
const sessionId = '7148e9de-7673-48b8-bf38-6799e52c346a';
|
||||
const sessionDir = join(testDir, sessionId);
|
||||
const subagentDir = join(sessionDir, 'subagents');
|
||||
const parentTranscriptPath = `${sessionDir}.jsonl`;
|
||||
const subagentTranscriptPath = join(subagentDir, 'agent-afacts-class2.jsonl');
|
||||
mkdirSync(subagentDir, { recursive: true });
|
||||
writeFileSync(parentTranscriptPath, '');
|
||||
writeFileSync(subagentTranscriptPath, '');
|
||||
|
||||
const resultPromise = runHelper({
|
||||
session_id: sessionId,
|
||||
agent_id: 'afacts-class2',
|
||||
transcript_path: subagentTranscriptPath,
|
||||
tool_response: {
|
||||
backgroundTaskId: 'bg-subagent-1',
|
||||
},
|
||||
});
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 100));
|
||||
writeFileSync(
|
||||
parentTranscriptPath,
|
||||
JSON.stringify({
|
||||
type: 'queue-operation',
|
||||
operation: 'enqueue',
|
||||
content:
|
||||
'<task-notification>\n<task-id>bg-subagent-1</task-id>\n<status>completed</status>\n' +
|
||||
'<output-file>/tmp/bg-subagent-1.output</output-file>\n</task-notification>',
|
||||
}) + '\n'
|
||||
);
|
||||
|
||||
const result = await resultPromise;
|
||||
expect(result.code).toBe(2);
|
||||
expect(result.stderr).toContain('bg-subagent-1');
|
||||
expect(result.stderr).toContain('/tmp/bg-subagent-1.output');
|
||||
});
|
||||
|
||||
it('includes a marked background report in the wake feedback', async () => {
|
||||
const transcriptPath = join(testDir, 'transcript.jsonl');
|
||||
const tasksDir = join(testDir, 'tasks');
|
||||
const outputPath = join(tasksDir, 'bg-report-1.output');
|
||||
mkdirSync(tasksDir, { recursive: true });
|
||||
writeFileSync(transcriptPath, '');
|
||||
writeFileSync(
|
||||
outputPath,
|
||||
[
|
||||
'launcher output',
|
||||
'=== CODEMAN_RESULT_BEGIN ===',
|
||||
'Summary line',
|
||||
'Detail after the old 30-line preview boundary',
|
||||
'=== CODEMAN_RESULT_END ===',
|
||||
].join('\n')
|
||||
);
|
||||
|
||||
const resultPromise = runHelper({
|
||||
transcript_path: transcriptPath,
|
||||
tool_response: {
|
||||
stdout: `Command running in background with ID: bg-report-1. Output is being written to: ${outputPath}.`,
|
||||
},
|
||||
});
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 100));
|
||||
writeFileSync(
|
||||
transcriptPath,
|
||||
JSON.stringify({
|
||||
type: 'queue-operation',
|
||||
operation: 'enqueue',
|
||||
content:
|
||||
'<task-notification>\n<task-id>bg-report-1</task-id>\n<status>completed</status>\n' +
|
||||
`<output-file>${outputPath}</output-file>\n</task-notification>`,
|
||||
}) + '\n'
|
||||
);
|
||||
|
||||
const result = await resultPromise;
|
||||
expect(result.code).toBe(2);
|
||||
expect(result.stderr).toContain('<codeman-background-result>');
|
||||
expect(result.stderr).toContain('Summary line');
|
||||
expect(result.stderr).toContain('Detail after the old 30-line preview boundary');
|
||||
});
|
||||
});
|
||||
|
||||
describe('subagent stop guard helper', () => {
|
||||
const testDir = join(tmpdir(), 'codeman-subagent-stop-guard-test-' + Date.now());
|
||||
|
||||
beforeEach(() => {
|
||||
mkdirSync(testDir, { recursive: true });
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
rmSync(testDir, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
function runGuard(transcriptLines: unknown[]): Promise<{ code: number | null; stdout: string; stderr: string }> {
|
||||
const transcriptPath = join(testDir, 'agent-test.jsonl');
|
||||
writeFileSync(transcriptPath, transcriptLines.map((line) => JSON.stringify(line)).join('\n') + '\n');
|
||||
|
||||
return new Promise((resolve, reject) => {
|
||||
const child = spawn(process.execPath, ['-e', generateSubagentStopGuardScript()], {
|
||||
stdio: ['pipe', 'pipe', 'pipe'],
|
||||
});
|
||||
let stdout = '';
|
||||
let stderr = '';
|
||||
child.stdout.setEncoding('utf8');
|
||||
child.stderr.setEncoding('utf8');
|
||||
child.stdout.on('data', (chunk) => {
|
||||
stdout += chunk;
|
||||
});
|
||||
child.stderr.on('data', (chunk) => {
|
||||
stderr += chunk;
|
||||
});
|
||||
child.on('error', reject);
|
||||
child.on('close', (code) => resolve({ code, stdout, stderr }));
|
||||
child.stdin.end(JSON.stringify({ agent_transcript_path: transcriptPath }));
|
||||
});
|
||||
}
|
||||
|
||||
async function withLiveTask<T>(taskId: string, action: () => Promise<T>): Promise<T> {
|
||||
const tasksDir = join(testDir, 'tasks');
|
||||
mkdirSync(tasksDir, { recursive: true });
|
||||
const outputFd = openSync(join(tasksDir, `${taskId}.output`), 'a');
|
||||
const child = spawn(process.execPath, ['-e', 'setTimeout(() => {}, 10000)'], {
|
||||
stdio: ['ignore', outputFd, outputFd],
|
||||
});
|
||||
await new Promise<void>((resolve, reject) => {
|
||||
child.once('spawn', resolve);
|
||||
child.once('error', reject);
|
||||
});
|
||||
closeSync(outputFd);
|
||||
|
||||
try {
|
||||
return await action();
|
||||
} finally {
|
||||
const closed = new Promise<void>((resolve) => child.once('close', () => resolve()));
|
||||
child.kill();
|
||||
await closed;
|
||||
}
|
||||
}
|
||||
|
||||
const monitorResult = (taskId: string) => ({
|
||||
type: 'user',
|
||||
message: {
|
||||
content: [
|
||||
{
|
||||
type: 'tool_result',
|
||||
content: `Monitor started (task ${taskId}, pid 123).`,
|
||||
},
|
||||
],
|
||||
},
|
||||
});
|
||||
|
||||
const completion = (taskId: string) => ({
|
||||
type: 'user',
|
||||
message: {
|
||||
content:
|
||||
`<task-notification>\n<task-id>${taskId}</task-id>\n` + '<status>completed</status>\n</task-notification>',
|
||||
},
|
||||
});
|
||||
|
||||
it('blocks an intermediate subagent stop while a sibling monitor is active', async () => {
|
||||
const result = await withLiveTask('monitor-still-live', () =>
|
||||
runGuard([monitorResult('monitor-first'), monitorResult('monitor-still-live'), completion('monitor-first')])
|
||||
);
|
||||
|
||||
expect(result.code).toBe(0);
|
||||
expect(result.stderr).toBe('');
|
||||
expect(JSON.parse(result.stdout)).toMatchObject({ decision: 'block' });
|
||||
expect(result.stdout).toContain('monitor-still-live');
|
||||
expect(result.stdout).not.toContain('monitor-first,');
|
||||
});
|
||||
|
||||
it('allows a subagent to stop after all of its monitored work finishes', async () => {
|
||||
const result = await runGuard([
|
||||
monitorResult('monitor-first'),
|
||||
monitorResult('monitor-second'),
|
||||
completion('monitor-first'),
|
||||
completion('monitor-second'),
|
||||
]);
|
||||
|
||||
expect(result.code).toBe(0);
|
||||
expect(result.stdout).toBe('');
|
||||
expect(result.stderr).toBe('');
|
||||
});
|
||||
|
||||
it('also recognizes background Bash task ownership', async () => {
|
||||
const result = await withLiveTask('bash-live-1', () =>
|
||||
runGuard([
|
||||
{
|
||||
type: 'user',
|
||||
message: {
|
||||
content: [
|
||||
{
|
||||
type: 'tool_result',
|
||||
content: 'Command running in background with ID: bash-live-1. Output is being written to a task file.',
|
||||
},
|
||||
],
|
||||
},
|
||||
},
|
||||
])
|
||||
);
|
||||
|
||||
expect(JSON.parse(result.stdout)).toMatchObject({ decision: 'block' });
|
||||
expect(result.stdout).toContain('bash-live-1');
|
||||
});
|
||||
});
|
||||
|
||||
// ========== Hook Event API Integration Tests ==========
|
||||
|
||||
@@ -89,4 +89,97 @@ describe('Stable HTTP contract (live server)', () => {
|
||||
expect(body.success).toBe(false);
|
||||
expect(body.errorCode).toBe('INVALID_INPUT');
|
||||
});
|
||||
|
||||
/**
|
||||
* The agent wait primitives, through the REAL pipeline.
|
||||
*
|
||||
* Their own route tests hand-roll a partial copy of the preSerialization hook that
|
||||
* maps errorCode to status but does NOT wrap bare payloads — so nothing there
|
||||
* proves these routes emit a correct envelope, a correct status, or work through
|
||||
* the /api/v1 alias, and one assertion in them pins `{}` for a response no client
|
||||
* will ever receive. This is the file whose docstring already claims that scope.
|
||||
*/
|
||||
describe('agent wait primitives', () => {
|
||||
let sessionId: string;
|
||||
|
||||
beforeAll(async () => {
|
||||
const res = await fetch(`${base}/api/sessions`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({}),
|
||||
});
|
||||
sessionId = (await res.json()).data.session.id;
|
||||
expect(sessionId).toBeDefined();
|
||||
});
|
||||
|
||||
afterAll(async () => {
|
||||
await fetch(`${base}/api/sessions/${sessionId}`, { method: 'DELETE' });
|
||||
});
|
||||
|
||||
it('answers a wait timeout as a 200 inside the envelope, on the /api/v1 alias', async () => {
|
||||
// A timeout is the long-poll SUCCEEDING at "did this happen within N ms?"; a
|
||||
// 4xx/5xx here would make every poll boundary indistinguishable from a failure.
|
||||
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait?until=working&timeout=1000`);
|
||||
expect(res.status).toBe(200);
|
||||
expect(res.headers.get('cache-control')).toBe('no-store');
|
||||
|
||||
const body = await res.json();
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.data.sessionId).toBe(sessionId);
|
||||
// The one shape all three wait endpoints share.
|
||||
expect(body.data.wait.timedOut).toBe(true);
|
||||
expect(body.data.wait.signal).toBeNull();
|
||||
expect(body.data.wait.timeoutMs).toBe(1000);
|
||||
expect(body.data.wait.until).toEqual(['working']);
|
||||
});
|
||||
|
||||
it('returns a contract-shaped 400 for an unknown until token', async () => {
|
||||
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait?until=stpo`);
|
||||
expect(res.status).toBe(400);
|
||||
const body = await res.json();
|
||||
expect(body.success).toBe(false);
|
||||
expect(body.errorCode).toBe('INVALID_INPUT');
|
||||
expect(body.error).toContain('stpo');
|
||||
});
|
||||
|
||||
it('returns a contract-shaped 400 naming the bad query parameter', async () => {
|
||||
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait?timeout=30s`);
|
||||
expect(res.status).toBe(400);
|
||||
const body = await res.json();
|
||||
expect(body.errorCode).toBe('INVALID_INPUT');
|
||||
expect(body.error).toContain('timeout');
|
||||
});
|
||||
|
||||
it('wraps the non-wait input response as { success: true, data: {} }', async () => {
|
||||
// What a client actually receives on the fire-and-forget path — NOT the bare
|
||||
// `{}` the handler returns and the route tests assert.
|
||||
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/input`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ input: 'hello' }),
|
||||
});
|
||||
expect(res.status).toBe(200);
|
||||
expect(await res.json()).toEqual({ success: true, data: {} });
|
||||
});
|
||||
|
||||
it('serves wait-output through the same envelope', async () => {
|
||||
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait-output?match=NEVER_APPEARS&timeout=1000`);
|
||||
expect(res.status).toBe(200);
|
||||
const body = await res.json();
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.data.wait.matched).toBe(false);
|
||||
expect(body.data.wait.timedOut).toBe(true);
|
||||
expect(body.data.wait.match).toBe('NEVER_APPEARS');
|
||||
});
|
||||
|
||||
it('404s an unknown session on both new routes, with the error envelope', async () => {
|
||||
for (const path of ['wait?until=idle', 'wait-output?match=x']) {
|
||||
const res = await fetch(`${base}/api/v1/sessions/nonexistent/${path}`);
|
||||
expect(res.status).toBe(404);
|
||||
const body = await res.json();
|
||||
expect(body.success).toBe(false);
|
||||
expect(body.errorCode).toBe('NOT_FOUND');
|
||||
}
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
@@ -0,0 +1,240 @@
|
||||
/**
|
||||
* @fileoverview Local-echo gating and input-ordering helpers for codex
|
||||
* sessions (issues #218/#219/#220/#222).
|
||||
*
|
||||
* Codex's composer is interactive per keystroke: typing "/" pops a
|
||||
* live-filtering command picker (#222), the composer grows as it wraps
|
||||
* (#220), pastes arrive bracketed (#219) and arrows edit server-side state
|
||||
* (#218). The buffer-until-Enter local echo overlay starves all of that, so
|
||||
* codex-mode sessions must use plain PTY echo like shell. The shared overlay
|
||||
* branch (claude/gemini/opencode) additionally flushes typed-but-unsent text
|
||||
* before forwarding bracketed pastes and composer nav keys, and hands the
|
||||
* session to pass-through after a nav key.
|
||||
*
|
||||
* Loaded via `vm` with a stubbed context (no jsdom), mirroring
|
||||
* test/input-send-order.test.ts. End-to-end behavior was verified against a
|
||||
* real codex 0.147.0 TUI in tmux through a headless browser.
|
||||
*/
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { performance } from 'node:perf_hooks';
|
||||
import { resolve } from 'node:path';
|
||||
import vm from 'node:vm';
|
||||
import { describe, expect, it, vi } from 'vitest';
|
||||
|
||||
type OverlayStub = {
|
||||
pendingText: string;
|
||||
cleared: number;
|
||||
suppressed: number;
|
||||
prompts: unknown[];
|
||||
clear(): void;
|
||||
suppressBufferDetection(): void;
|
||||
setPrompt(p: unknown): void;
|
||||
appendText: ReturnType<typeof vi.fn>;
|
||||
};
|
||||
|
||||
type AppInstance = {
|
||||
activeSessionId: string | null;
|
||||
sessions: Map<string, { mode: string }>;
|
||||
terminal?: { focus: () => void };
|
||||
_localEchoEnabled?: boolean;
|
||||
_localEchoOverlay?: OverlayStub;
|
||||
_pendingInput: string;
|
||||
_flushedOffsets?: Map<string, number>;
|
||||
_flushedTexts?: Map<string, string>;
|
||||
_echoPassthroughSessions?: Set<string>;
|
||||
loadAppSettingsFromStorage: () => Record<string, unknown>;
|
||||
sendInput: ReturnType<typeof vi.fn>;
|
||||
_updateLocalEchoState(): void;
|
||||
_flushLocalEchoPending(): void;
|
||||
insertTerminalText(text: string): void;
|
||||
};
|
||||
|
||||
function loadContext() {
|
||||
const read = (f: string) => readFileSync(resolve(import.meta.dirname, `../src/web/public/${f}`), 'utf8');
|
||||
const windowStub: Record<string, unknown> = {
|
||||
addEventListener: vi.fn(),
|
||||
removeEventListener: vi.fn(),
|
||||
};
|
||||
const context = vm.createContext({
|
||||
console,
|
||||
performance,
|
||||
setInterval: vi.fn(),
|
||||
clearInterval: vi.fn(),
|
||||
setTimeout,
|
||||
clearTimeout,
|
||||
requestAnimationFrame: vi.fn(),
|
||||
HTMLCanvasElement: class HTMLCanvasElement {},
|
||||
WebSocket: { OPEN: 1 },
|
||||
fetch: vi.fn(),
|
||||
document: { addEventListener: vi.fn(), documentElement: { dataset: {} } },
|
||||
localStorage: {
|
||||
length: 0,
|
||||
key: vi.fn(),
|
||||
getItem: vi.fn(),
|
||||
setItem: vi.fn(),
|
||||
removeItem: vi.fn(),
|
||||
},
|
||||
window: windowStub,
|
||||
MobileDetection: {
|
||||
isTouchDevice: () => true,
|
||||
isHandheldDevice: () => false,
|
||||
getDeviceType: () => 'desktop',
|
||||
},
|
||||
});
|
||||
vm.runInContext(
|
||||
`${read('constants.js')}\n${read('app.js')}\n${read('terminal-ui.js')}\nglobalThis.__CodemanApp = CodemanApp;`,
|
||||
context
|
||||
);
|
||||
const CodemanApp = (context as unknown as { __CodemanApp: { prototype: object } }).__CodemanApp;
|
||||
return {
|
||||
CodemanApp,
|
||||
terminalInput: (windowStub as { CodemanTerminalInput?: Record<string, unknown> }).CodemanTerminalInput!,
|
||||
};
|
||||
}
|
||||
|
||||
const { CodemanApp, terminalInput } = loadContext();
|
||||
const isComposerNavKey = terminalInput.isComposerNavKey as (data: string) => boolean;
|
||||
|
||||
function makeOverlay(pending = ''): OverlayStub {
|
||||
return {
|
||||
pendingText: pending,
|
||||
cleared: 0,
|
||||
suppressed: 0,
|
||||
prompts: [],
|
||||
clear() {
|
||||
this.cleared++;
|
||||
this.pendingText = '';
|
||||
},
|
||||
suppressBufferDetection() {
|
||||
this.suppressed++;
|
||||
},
|
||||
setPrompt(p: unknown) {
|
||||
this.prompts.push(p);
|
||||
},
|
||||
appendText: vi.fn(),
|
||||
};
|
||||
}
|
||||
|
||||
function makeApp(mode: string, overlay = makeOverlay()): AppInstance {
|
||||
const app = Object.create(CodemanApp.prototype) as AppInstance;
|
||||
app.activeSessionId = 's1';
|
||||
app.sessions = new Map([['s1', { mode }]]);
|
||||
app._localEchoOverlay = overlay;
|
||||
app._pendingInput = '';
|
||||
app._flushedOffsets = new Map([['s1', 3]]);
|
||||
app._flushedTexts = new Map([['s1', 'abc']]);
|
||||
app.loadAppSettingsFromStorage = () => ({ localEchoEnabled: true });
|
||||
app.sendInput = vi.fn().mockResolvedValue(undefined);
|
||||
return app;
|
||||
}
|
||||
|
||||
describe('CodemanTerminalInput.isComposerNavKey', () => {
|
||||
it.each([
|
||||
'\x1b[A',
|
||||
'\x1b[B',
|
||||
'\x1b[C',
|
||||
'\x1b[D',
|
||||
'\x1b[H',
|
||||
'\x1b[F',
|
||||
'\x1bOA',
|
||||
'\x1bOD',
|
||||
'\x1bOH',
|
||||
'\x1bOF',
|
||||
'\x1b[1;5C', // Ctrl+Right
|
||||
'\x1b[1;2A', // Shift+Up
|
||||
'\x1b[3~', // Delete
|
||||
'\x1b[3;5~', // Ctrl+Delete
|
||||
'\x1b[5~', // PgUp
|
||||
'\x1b[6~', // PgDn
|
||||
'\x1b[1~', // Home variant
|
||||
'\x1b[4~', // End variant
|
||||
])('classifies %j as a composer nav key', (seq) => {
|
||||
expect(isComposerNavKey(seq)).toBe(true);
|
||||
});
|
||||
|
||||
it.each([
|
||||
'\x1b[?1;2c', // DA1 response
|
||||
'\x1b[>0;276;0c', // DA2 response
|
||||
'\x1b[12;34R', // CPR response
|
||||
'\x1b[1;3R', // CPR response (small coords)
|
||||
'\x1b[0n', // DSR response
|
||||
'\x1b[15~', // F5 (function keys stay out)
|
||||
'\x1b[200~hi\x1b[201~', // bracketed paste
|
||||
'\x1b[?u', // kitty keyboard query response
|
||||
'\x1bOP', // F1
|
||||
'\x1b',
|
||||
'a',
|
||||
'abc',
|
||||
'\r',
|
||||
])('does NOT classify %j as a composer nav key', (seq) => {
|
||||
expect(isComposerNavKey(seq)).toBe(false);
|
||||
});
|
||||
|
||||
it('exports the bracketed paste prefix xterm puts on terminal.paste()', () => {
|
||||
expect(terminalInput.BRACKETED_PASTE_START).toBe('\x1b[200~');
|
||||
});
|
||||
});
|
||||
|
||||
describe('_updateLocalEchoState mode gating', () => {
|
||||
it('disables the overlay for codex sessions even with the setting ON (issues #218/#219/#220/#222)', () => {
|
||||
const overlay = makeOverlay('pending');
|
||||
const app = makeApp('codex', overlay);
|
||||
app._updateLocalEchoState();
|
||||
expect(app._localEchoEnabled).toBe(false);
|
||||
expect(overlay.cleared).toBeGreaterThan(0);
|
||||
});
|
||||
|
||||
it('disables the overlay for shell sessions (PTY provides its own echo)', () => {
|
||||
const app = makeApp('shell');
|
||||
app._updateLocalEchoState();
|
||||
expect(app._localEchoEnabled).toBe(false);
|
||||
});
|
||||
|
||||
it.each(['claude', 'gemini', 'opencode'])('keeps the overlay enabled for %s sessions', (mode) => {
|
||||
const overlay = makeOverlay();
|
||||
const app = makeApp(mode, overlay);
|
||||
app._updateLocalEchoState();
|
||||
expect(app._localEchoEnabled).toBe(true);
|
||||
expect(overlay.prompts.length).toBeGreaterThan(0);
|
||||
});
|
||||
});
|
||||
|
||||
describe('_flushLocalEchoPending', () => {
|
||||
it('moves pending text into _pendingInput and resets overlay + flushed tracking', () => {
|
||||
const overlay = makeOverlay('hello');
|
||||
const app = makeApp('claude', overlay);
|
||||
app._flushLocalEchoPending();
|
||||
expect(app._pendingInput).toBe('hello');
|
||||
expect(overlay.cleared).toBe(1);
|
||||
expect(overlay.suppressed).toBe(1);
|
||||
expect(app._flushedOffsets!.has('s1')).toBe(false);
|
||||
expect(app._flushedTexts!.has('s1')).toBe(false);
|
||||
});
|
||||
|
||||
it('appends nothing when the overlay is empty', () => {
|
||||
const app = makeApp('claude', makeOverlay(''));
|
||||
app._flushLocalEchoPending();
|
||||
expect(app._pendingInput).toBe('');
|
||||
});
|
||||
});
|
||||
|
||||
describe('insertTerminalText pass-through routing', () => {
|
||||
it('appends to the overlay while local echo is buffering', () => {
|
||||
const overlay = makeOverlay();
|
||||
const app = makeApp('claude', overlay);
|
||||
app._localEchoEnabled = true;
|
||||
app.insertTerminalText('path.txt');
|
||||
expect(overlay.appendText).toHaveBeenCalledWith('path.txt');
|
||||
expect(app.sendInput).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('sends directly while the session is in nav-key pass-through', () => {
|
||||
const overlay = makeOverlay();
|
||||
const app = makeApp('claude', overlay);
|
||||
app._localEchoEnabled = true;
|
||||
app._echoPassthroughSessions = new Set(['s1']);
|
||||
app.insertTerminalText('path.txt');
|
||||
expect(app.sendInput).toHaveBeenCalledWith('path.txt');
|
||||
expect(overlay.appendText).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
@@ -323,6 +323,146 @@ describe('Virtual Keyboard', () => {
|
||||
expect(mainPadding).toBe('');
|
||||
});
|
||||
|
||||
it('coalesces keyboard animation frames into one final terminal fit', async () => {
|
||||
const result = await page.evaluate(async () => {
|
||||
// `app.terminal` and `app.fitAddon` are only assigned by initTerminal(),
|
||||
// which needs a selected session this harness never creates. Both are
|
||||
// null at rest, and the settle callback returns early on a falsy
|
||||
// terminal — so without stand-ins this test cannot reach the behavior
|
||||
// it asserts. Install the minimum surface the callback touches.
|
||||
const hadTerminal = app.terminal !== null && app.terminal !== undefined;
|
||||
const hadFitAddon = app.fitAddon !== null && app.fitAddon !== undefined;
|
||||
if (!hadTerminal) app.terminal = { scrollToBottom() {} };
|
||||
if (!hadFitAddon) app.fitAddon = { fit() {}, proposeDimensions: () => null };
|
||||
|
||||
const originalFit = app.fitAddon.fit.bind(app.fitAddon);
|
||||
const originalSendResize = KeyboardHandler._sendTerminalResize.bind(KeyboardHandler);
|
||||
const originalScrollToBottom = app.terminal.scrollToBottom.bind(app.terminal);
|
||||
let fits = 0;
|
||||
let resizes = 0;
|
||||
let bottomRestores = 0;
|
||||
app.fitAddon.fit = () => {
|
||||
fits++;
|
||||
};
|
||||
KeyboardHandler._sendTerminalResize = () => {
|
||||
resizes++;
|
||||
};
|
||||
app.terminal.scrollToBottom = () => {
|
||||
bottomRestores++;
|
||||
};
|
||||
|
||||
KeyboardHandler._scheduleViewportSettle({ scrollToBottom: true });
|
||||
await new Promise((resolve) => setTimeout(resolve, 30));
|
||||
KeyboardHandler._scheduleViewportSettle();
|
||||
await new Promise((resolve) => setTimeout(resolve, 30));
|
||||
KeyboardHandler._scheduleViewportSettle();
|
||||
await new Promise((resolve) => setTimeout(resolve, 50));
|
||||
const beforeFinalSettle = { fits, resizes, bottomRestores };
|
||||
await new Promise((resolve) => setTimeout(resolve, KeyboardHandler.VIEWPORT_SETTLE_MS));
|
||||
const afterFinalSettle = { fits, resizes, bottomRestores };
|
||||
|
||||
app.fitAddon.fit = originalFit;
|
||||
KeyboardHandler._sendTerminalResize = originalSendResize;
|
||||
app.terminal.scrollToBottom = originalScrollToBottom;
|
||||
if (!hadFitAddon) app.fitAddon = null;
|
||||
if (!hadTerminal) app.terminal = null;
|
||||
return { beforeFinalSettle, afterFinalSettle };
|
||||
});
|
||||
|
||||
expect(result.beforeFinalSettle).toEqual({ fits: 0, resizes: 0, bottomRestores: 0 });
|
||||
expect(result.afterFinalSettle).toEqual({ fits: 1, resizes: 1, bottomRestores: 1 });
|
||||
});
|
||||
|
||||
// Behavioral counterpart to the test above, driven through the PUBLIC entry
|
||||
// point rather than the internal scheduler. Before this change each
|
||||
// onKeyboardShow armed its own uncoalesced 150ms setTimeout, so a keyboard
|
||||
// animation that reports several viewport steps refit the terminal once per
|
||||
// step — the visible symptom being repeated reflow while the keyboard slides
|
||||
// up. This asserts the observable outcome (one refit for a burst) and so
|
||||
// fails on master by COUNT, not by a missing method.
|
||||
it('refits once for a burst of keyboard viewport steps', async () => {
|
||||
const counts = await page.evaluate(async () => {
|
||||
const hadTerminal = app.terminal !== null && app.terminal !== undefined;
|
||||
const hadFitAddon = app.fitAddon !== null && app.fitAddon !== undefined;
|
||||
if (!hadTerminal) app.terminal = { scrollToBottom() {} };
|
||||
if (!hadFitAddon) app.fitAddon = { fit() {}, proposeDimensions: () => null };
|
||||
|
||||
const originalFit = app.fitAddon.fit.bind(app.fitAddon);
|
||||
const originalSendResize = KeyboardHandler._sendTerminalResize.bind(KeyboardHandler);
|
||||
let fits = 0;
|
||||
app.fitAddon.fit = () => {
|
||||
fits++;
|
||||
};
|
||||
KeyboardHandler._sendTerminalResize = () => {};
|
||||
|
||||
// Three viewport steps in quick succession, as a keyboard animation
|
||||
// produces on a real device.
|
||||
KeyboardHandler.onKeyboardShow();
|
||||
await new Promise((resolve) => setTimeout(resolve, 30));
|
||||
KeyboardHandler.onKeyboardShow();
|
||||
await new Promise((resolve) => setTimeout(resolve, 30));
|
||||
KeyboardHandler.onKeyboardShow();
|
||||
|
||||
// Well past both the coalescing window and master's fixed 150ms timer.
|
||||
await new Promise((resolve) => setTimeout(resolve, 400));
|
||||
|
||||
app.fitAddon.fit = originalFit;
|
||||
KeyboardHandler._sendTerminalResize = originalSendResize;
|
||||
if (!hadFitAddon) app.fitAddon = null;
|
||||
if (!hadTerminal) app.terminal = null;
|
||||
return fits;
|
||||
});
|
||||
|
||||
// Coalesced: one refit for the whole burst. Master fires one per step.
|
||||
expect(counts).toBe(1);
|
||||
});
|
||||
|
||||
// A viewport resize with NO pending show/hide transition must not arm settle
|
||||
// work of its own: keyboard detection can miss a fine-grained OS animation
|
||||
// entirely (sub-150px steps with the baseline chasing the animation), and a
|
||||
// fit against that mid-animation, uncompensated layout resizes the PTY to
|
||||
// transient dims. The resulting SIGWINCH thrash duplicates prompts and
|
||||
// garbles the transcript. Wiggles may only push a pending settle back.
|
||||
it('does not refit on viewport wiggles without a keyboard transition', async () => {
|
||||
const result = await page.evaluate(async () => {
|
||||
const hadTerminal = app.terminal !== null && app.terminal !== undefined;
|
||||
const hadFitAddon = app.fitAddon !== null && app.fitAddon !== undefined;
|
||||
if (!hadTerminal) app.terminal = { scrollToBottom() {} };
|
||||
if (!hadFitAddon) app.fitAddon = { fit() {}, proposeDimensions: () => null };
|
||||
|
||||
const originalFit = app.fitAddon.fit.bind(app.fitAddon);
|
||||
const originalSendResize = KeyboardHandler._sendTerminalResize.bind(KeyboardHandler);
|
||||
let fits = 0;
|
||||
app.fitAddon.fit = () => {
|
||||
fits++;
|
||||
};
|
||||
KeyboardHandler._sendTerminalResize = () => {};
|
||||
|
||||
// Wiggle only: nothing pending, so nothing may fire.
|
||||
KeyboardHandler._deferViewportSettle();
|
||||
KeyboardHandler._deferViewportSettle();
|
||||
await new Promise((resolve) => setTimeout(resolve, KeyboardHandler.VIEWPORT_SETTLE_MS + 80));
|
||||
const wiggleOnly = fits;
|
||||
|
||||
// A real transition arms the work; a following wiggle defers it but the
|
||||
// settle still fires exactly once.
|
||||
KeyboardHandler._scheduleViewportSettle({ scrollToBottom: true });
|
||||
await new Promise((resolve) => setTimeout(resolve, 30));
|
||||
KeyboardHandler._deferViewportSettle();
|
||||
await new Promise((resolve) => setTimeout(resolve, KeyboardHandler.VIEWPORT_SETTLE_MS + 80));
|
||||
const afterTransition = fits;
|
||||
|
||||
app.fitAddon.fit = originalFit;
|
||||
KeyboardHandler._sendTerminalResize = originalSendResize;
|
||||
if (!hadFitAddon) app.fitAddon = null;
|
||||
if (!hadTerminal) app.terminal = null;
|
||||
return { wiggleOnly, afterTransition };
|
||||
});
|
||||
|
||||
expect(result.wiggleOnly).toBe(0);
|
||||
expect(result.afterTransition).toBe(1);
|
||||
});
|
||||
|
||||
it('accessory bar has the simple-mode action buttons', async () => {
|
||||
const actions = await page.evaluate(() => {
|
||||
return Array.from(document.querySelectorAll('.keyboard-accessory-bar [data-action]')).map(
|
||||
|
||||
@@ -4,6 +4,7 @@
|
||||
*/
|
||||
import { EventEmitter } from 'node:events';
|
||||
import { vi } from 'vitest';
|
||||
import type { SessionStatus } from '../../src/types.js';
|
||||
|
||||
/**
|
||||
* Enhanced mock session for testing RespawnController.
|
||||
@@ -12,8 +13,15 @@ import { vi } from 'vitest';
|
||||
export class MockSession extends EventEmitter {
|
||||
id: string;
|
||||
workingDir: string = '/tmp/test-workdir';
|
||||
status: 'idle' | 'working' = 'idle';
|
||||
pid: number = 12345;
|
||||
/**
|
||||
* The REAL union, deliberately. This used to be `'idle' | 'working'`, and
|
||||
* `'working'` is not a `SessionStatus` at all — so `signalForStatus()` fell to its
|
||||
* `default: null` branch in every route test and the busy / stopped / error halves
|
||||
* of the immediate-resolve mapping had zero coverage while appearing to be tested.
|
||||
*/
|
||||
status: SessionStatus = 'idle';
|
||||
/** `null` once the PTY is gone (or before it has ever started) — see `pid` in Session. */
|
||||
pid: number | null = 12345;
|
||||
isWorking: boolean = false;
|
||||
private _activeChildProcesses: { pid: number; command: string }[] = [];
|
||||
ralphTracker: null = null;
|
||||
@@ -30,13 +38,22 @@ export class MockSession extends EventEmitter {
|
||||
this._muxName = `codeman-test-${id.slice(0, 8)}`;
|
||||
}
|
||||
|
||||
/** Direct PTY write (used by session.write()) */
|
||||
write(data: string): void {
|
||||
/**
|
||||
* Set to simulate a session whose PTY is gone: both write paths report failure,
|
||||
* which is the state in which input used to disappear silently.
|
||||
*/
|
||||
failWrites = false;
|
||||
|
||||
/** Direct PTY write (used by session.write()). Mirrors the real boolean return. */
|
||||
write(data: string): boolean {
|
||||
if (this.failWrites) return false;
|
||||
this.writeBuffer.push(data);
|
||||
return true;
|
||||
}
|
||||
|
||||
/** Write via mux (used by respawn controller) */
|
||||
async writeViaMux(data: string): Promise<boolean> {
|
||||
if (this.failWrites) return false;
|
||||
this.writeBuffer.push(data);
|
||||
return true;
|
||||
}
|
||||
@@ -44,6 +61,10 @@ export class MockSession extends EventEmitter {
|
||||
/** Exactly-once input dedup — mirrors Session.shouldApplyInput so route tests
|
||||
* exercising the reliable-delivery path behave like production. */
|
||||
private _appliedInputSeq = new Map<string, number>();
|
||||
forgetInputSeq(clientId: string, seq: number): void {
|
||||
if (this._appliedInputSeq.get(clientId) === seq) this._appliedInputSeq.set(clientId, seq - 1);
|
||||
}
|
||||
|
||||
shouldApplyInput(clientId: string, seq: number): boolean {
|
||||
const last = this._appliedInputSeq.get(clientId);
|
||||
if (last !== undefined && seq <= last) return false;
|
||||
@@ -100,7 +121,9 @@ export class MockSession extends EventEmitter {
|
||||
/** Simulate working state with spinner */
|
||||
simulateWorking(text: string = 'Thinking'): void {
|
||||
this.simulateTerminalOutput(`${text}... \u280b`);
|
||||
this.status = 'working';
|
||||
// 'busy' is what the real Session sets while a turn is in flight; the old
|
||||
// 'working' here was the event name, not a status value.
|
||||
this.status = 'busy';
|
||||
this.emit('working');
|
||||
}
|
||||
|
||||
@@ -169,6 +192,14 @@ export class MockSession extends EventEmitter {
|
||||
return this._muxName;
|
||||
}
|
||||
|
||||
/**
|
||||
* Mirrors `Session.usesMux`. True by default because that is the normal
|
||||
* configuration, and it is what makes a route's pane-liveness probe reachable:
|
||||
* `session.pid` is the tmux ATTACH CLIENT, so a mux-backed session's worker can be
|
||||
* dead while `pid` is still a live number.
|
||||
*/
|
||||
usesMux: boolean = true;
|
||||
|
||||
/** Check for active child processes (mock returns configurable list) */
|
||||
getActiveChildProcesses(): { pid: number; command: string }[] {
|
||||
return this._activeChildProcesses;
|
||||
|
||||
@@ -0,0 +1,200 @@
|
||||
/**
|
||||
* Regression guard for the process-tree walk.
|
||||
*
|
||||
* On 2026-07-30 an unbounded version took a machine down: it ran `pgrep -P <pid>` per
|
||||
* node and recursed with no visited set, no depth limit and no node cap. Across ~28
|
||||
* adopted tmux trees the fan-out exploded, and because each `pgrep` blocks in the WSL
|
||||
* kernel while reading /proc/<pid>/cgroup, none returned while the walk kept firing
|
||||
* more. Result: ~13,000 `pgrep` processes stuck in D-state out of ~39,000 total, load
|
||||
* average above 13,000, recoverable only by
|
||||
* restarting WSL — which cost every running session.
|
||||
*
|
||||
* These tests exercise the SHIPPED function. An earlier version of this file carried
|
||||
* its own copy of the traversal, which would have passed happily while the real code
|
||||
* regressed; that is why the walk now lives in its own module.
|
||||
*/
|
||||
import { describe, expect, it, vi } from 'vitest';
|
||||
|
||||
import { collectDescendants, PROC_WALK_MAX_DEPTH, PROC_WALK_MAX_NODES } from '../src/proc-tree.js';
|
||||
|
||||
/** Build a parent→children map from `[parent, child]` pairs. */
|
||||
function tree(pairs: [number, number][]): Map<number, number[]> {
|
||||
const m = new Map<number, number[]>();
|
||||
for (const [p, c] of pairs) m.set(p, [...(m.get(p) ?? []), c]);
|
||||
return m;
|
||||
}
|
||||
|
||||
/** A chain 1→2→…→n, i.e. depth n-1. */
|
||||
function chain(n: number): Map<number, number[]> {
|
||||
return tree(Array.from({ length: n - 1 }, (_, i) => [i + 1, i + 2] as [number, number]));
|
||||
}
|
||||
|
||||
describe('collectDescendants', () => {
|
||||
it('returns every descendant of a normal tree, root excluded', () => {
|
||||
const t = tree([
|
||||
[1, 2],
|
||||
[1, 3],
|
||||
[2, 4],
|
||||
[3, 5],
|
||||
]);
|
||||
expect(collectDescendants(1, t).sort((a, b) => a - b)).toEqual([2, 3, 4, 5]);
|
||||
});
|
||||
|
||||
it('terminates on a cycle instead of looping forever', () => {
|
||||
// A live `ps` snapshot is not atomic; pid reuse can produce a parent loop.
|
||||
const t = tree([
|
||||
[1, 2],
|
||||
[2, 3],
|
||||
[3, 1],
|
||||
[3, 2],
|
||||
]);
|
||||
expect(collectDescendants(1, t).sort((a, b) => a - b)).toEqual([2, 3]);
|
||||
});
|
||||
|
||||
it('does not include the root even when something claims it as a child', () => {
|
||||
expect(collectDescendants(1, tree([[1, 1]]))).toEqual([]);
|
||||
});
|
||||
|
||||
it('caps the depth, and says so', () => {
|
||||
// 40 generations available, only PROC_WALK_MAX_DEPTH may be descended. A silent
|
||||
// depth cap hides a deep tree exactly as a silent node cap hides a wide one.
|
||||
const onTruncated = vi.fn();
|
||||
expect(collectDescendants(1, chain(40), { onTruncated })).toHaveLength(PROC_WALK_MAX_DEPTH);
|
||||
expect(onTruncated).toHaveBeenCalledWith(1, PROC_WALK_MAX_DEPTH, 'depth');
|
||||
});
|
||||
|
||||
it('stays silent about depth when the tree ends inside the cap', () => {
|
||||
const onTruncated = vi.fn();
|
||||
collectDescendants(1, chain(4), { onTruncated });
|
||||
expect(onTruncated).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('caps the node count and reports the truncation', () => {
|
||||
// One parent with far more children than the cap allows.
|
||||
const wide = new Map<number, number[]>([[1, Array.from({ length: PROC_WALK_MAX_NODES * 3 }, (_, i) => i + 2)]]);
|
||||
const onTruncated = vi.fn();
|
||||
|
||||
const out = collectDescendants(1, wide, { onTruncated });
|
||||
|
||||
expect(out).toHaveLength(PROC_WALK_MAX_NODES);
|
||||
expect(onTruncated).toHaveBeenCalledWith(1, PROC_WALK_MAX_NODES, 'nodes');
|
||||
});
|
||||
|
||||
it('stays silent when nothing was truncated', () => {
|
||||
const onTruncated = vi.fn();
|
||||
collectDescendants(1, tree([[1, 2]]), { onTruncated });
|
||||
expect(onTruncated).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('survives the shape that caused the incident: many wide, deep trees', () => {
|
||||
// 28 adopted trees, branching 4-wide. Depth 6 already gives 4096 nodes per tree —
|
||||
// eight times the cap, which is what this asserts. (The real incident's trees were
|
||||
// deeper still; building that here would mean materialising 16M map entries and
|
||||
// would only test the fixture builder.)
|
||||
const t = new Map<number, number[]>();
|
||||
let next = 1000;
|
||||
const roots: number[] = [];
|
||||
for (let r = 0; r < 28; r += 1) {
|
||||
const root = next++;
|
||||
roots.push(root);
|
||||
let frontier = [root];
|
||||
for (let d = 0; d < 6; d += 1) {
|
||||
const nf: number[] = [];
|
||||
for (const p of frontier) {
|
||||
const kids = [next++, next++, next++, next++];
|
||||
t.set(p, kids);
|
||||
nf.push(...kids);
|
||||
}
|
||||
frontier = nf;
|
||||
}
|
||||
}
|
||||
|
||||
for (const root of roots) {
|
||||
const out = collectDescendants(root, t);
|
||||
expect(out.length).toBeLessThanOrEqual(PROC_WALK_MAX_NODES);
|
||||
}
|
||||
});
|
||||
|
||||
it('honours explicit overrides', () => {
|
||||
expect(collectDescendants(1, chain(40), { maxDepth: 3 })).toEqual([2, 3, 4]);
|
||||
expect(collectDescendants(1, chain(40), { maxNodes: 2 })).toEqual([2, 3]);
|
||||
});
|
||||
|
||||
it('returns nothing for an unknown pid or an empty snapshot', () => {
|
||||
expect(collectDescendants(999, tree([[1, 2]]))).toEqual([]);
|
||||
expect(collectDescendants(1, new Map())).toEqual([]);
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* The bound must be reachable through the code that actually kills things.
|
||||
*
|
||||
* The unit tests above exercise `collectDescendants` directly, which is necessary but
|
||||
* not sufficient: reverting `tmux-manager.ts` to the old unbounded `pgrep -P` recursion
|
||||
* left every one of them green. This asserts the wiring — that TmuxManager's descendant
|
||||
* lookup goes through the bounded walk and honours its caps.
|
||||
*
|
||||
* The snapshot refresh is stubbed. Without that the manager runs a real `ps` and
|
||||
* replaces the fixture, and the test silently measures the machine's own process tree
|
||||
* instead of the tree under test — which is how the first version of this test passed
|
||||
* even with both caps bypassed.
|
||||
*/
|
||||
describe('TmuxManager uses the bounded walk', () => {
|
||||
/** Build a tree `width` wide and `depth` deep, rooted at 1. */
|
||||
function bigTree(width: number, depth: number): Map<number, number[]> {
|
||||
const t = new Map<number, number[]>();
|
||||
let next = 2;
|
||||
let frontier = [1];
|
||||
for (let d = 0; d < depth; d += 1) {
|
||||
const nf: number[] = [];
|
||||
for (const p of frontier) {
|
||||
const kids = Array.from({ length: width }, () => next++);
|
||||
t.set(p, kids);
|
||||
nf.push(...kids);
|
||||
}
|
||||
frontier = nf;
|
||||
}
|
||||
return t;
|
||||
}
|
||||
|
||||
async function walkVia(fixture: Map<number, number[]>): Promise<number[]> {
|
||||
const { TmuxManager } = await import('../src/tmux-manager.js');
|
||||
const Klass = TmuxManager as unknown as {
|
||||
refreshProcSnapshot(): Promise<Map<number, number[]>>;
|
||||
procSnapshot: unknown;
|
||||
};
|
||||
const original = Klass.refreshProcSnapshot;
|
||||
Klass.refreshProcSnapshot = () => Promise.resolve(fixture);
|
||||
Klass.procSnapshot = { at: Date.now(), byParent: fixture };
|
||||
try {
|
||||
const mgr = new TmuxManager();
|
||||
return await (mgr as unknown as { getChildPidsFresh(pid: number): Promise<number[]> }).getChildPidsFresh(1);
|
||||
} finally {
|
||||
Klass.refreshProcSnapshot = original;
|
||||
Klass.procSnapshot = null;
|
||||
}
|
||||
}
|
||||
|
||||
it('honours the node cap on a tree far wider than it', async () => {
|
||||
// 4096 descendants available; the cap is 500. With the caps bypassed — the shape
|
||||
// a regression at the call site would take — this returns thousands.
|
||||
const out = await walkVia(bigTree(4, 6));
|
||||
expect(out.length).toBe(PROC_WALK_MAX_NODES);
|
||||
});
|
||||
|
||||
it('honours the depth cap on a deep chain', async () => {
|
||||
const chainTree = new Map<number, number[]>();
|
||||
for (let i = 1; i < 40; i += 1) chainTree.set(i, [i + 1]);
|
||||
|
||||
const out = await walkVia(chainTree);
|
||||
|
||||
expect(out.length).toBe(PROC_WALK_MAX_DEPTH);
|
||||
});
|
||||
|
||||
it('spawns nothing per node — the walk only reads the snapshot', async () => {
|
||||
// A per-node spawn against this fixture would mean thousands of processes; the
|
||||
// test completing at all is the assertion, plus the bound holding.
|
||||
const out = await walkVia(bigTree(4, 6));
|
||||
expect(out.length).toBeLessThanOrEqual(PROC_WALK_MAX_NODES);
|
||||
});
|
||||
});
|
||||
@@ -20,7 +20,12 @@
|
||||
|
||||
import { homedir } from 'node:os';
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import { buildSshConnectionArgs, buildRemoteTmuxCheckCommand, remoteSshTarget } from '../src/remote-hosts.js';
|
||||
import {
|
||||
buildSshConnectionArgs,
|
||||
buildRemoteTmuxCheckCommand,
|
||||
buildRemoteCliVersionProbeCommand,
|
||||
remoteSshTarget,
|
||||
} from '../src/remote-hosts.js';
|
||||
import { buildRemoteLaunchCommand } from '../src/tmux-manager.js';
|
||||
import type { SessionRemote } from '../src/types.js';
|
||||
|
||||
@@ -198,3 +203,29 @@ describe('COD-107 buildRemoteTmuxCheckCommand — same connection options as the
|
||||
expect(buildRemoteTmuxCheckCommand({ username: 'ubuntu', host: '10.0.0.42', port: 2222 })).toContain('-p 2222');
|
||||
});
|
||||
});
|
||||
|
||||
describe('buildRemoteCliVersionProbeCommand: remote CLI version over the same connection (#205)', () => {
|
||||
it('routes the version query through the interactive-login shell wrapper, like the launch', () => {
|
||||
const cmd = buildRemoteCliVersionProbeCommand(baseRemote, 'claude');
|
||||
// Same PATH-resolution wrapper as defaultRemoteCommandForMode: a bare
|
||||
// `claude --version` over ssh sees only sshd's minimal PATH (exit 127).
|
||||
expect(cmd).toBe(
|
||||
'ssh -o BatchMode=yes -o ConnectTimeout=10 ubuntu@10.0.0.42 ' +
|
||||
`'exec "\${SHELL:-/bin/sh}" -i -l -c '\\''claude --version'\\'''`
|
||||
);
|
||||
});
|
||||
|
||||
it('uses the shared connection args (proxy/identity/port), so it reaches what the launch reaches', () => {
|
||||
const cmd = buildRemoteCliVersionProbeCommand(aaDesktop, 'claude');
|
||||
expect(cmd).toContain('-o BatchMode=yes');
|
||||
expect(cmd).toContain('-p 2222');
|
||||
expect(cmd).toContain(`-i '${HOME}/.ssh/remote_ed25519'`);
|
||||
expect(cmd).toContain("-o 'ProxyCommand=nc -X 5 -x 127.0.0.1:1080 %h %p'");
|
||||
expect(cmd).toContain('aakht@192.168.55.170');
|
||||
});
|
||||
|
||||
it('maps antigravity to its real binary name and shell to no probe at all', () => {
|
||||
expect(buildRemoteCliVersionProbeCommand(baseRemote, 'antigravity')).toContain('agy --version');
|
||||
expect(buildRemoteCliVersionProbeCommand(baseRemote, 'shell')).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
@@ -0,0 +1,157 @@
|
||||
/**
|
||||
* @fileoverview A lost input must stay retryable.
|
||||
*
|
||||
* POST /api/sessions/:id/input answers 200 BEFORE the write is attempted — the mux
|
||||
* write is fire-and-forget so the HTTP response never waits on a tmux child. The
|
||||
* dedup bookkeeping, however, recorded the (clientId, seq) pair as applied at that
|
||||
* same moment. A write that then failed left the client with a 200, no message in
|
||||
* the pane, and a seq the server would reject as a duplicate on retry: the input was
|
||||
* unrecoverable by the very mechanism meant to make delivery reliable.
|
||||
*
|
||||
* Observed in the wild: a prompt shown as sent in a chat client, a 200 in the proxy
|
||||
* log, and an empty prompt line in the pane.
|
||||
*/
|
||||
|
||||
import fastifyCookie from '@fastify/cookie';
|
||||
import Fastify, { type FastifyInstance } from 'fastify';
|
||||
import { afterEach, beforeEach, describe, expect, it } from 'vitest';
|
||||
|
||||
import { Session } from '../../src/session.js';
|
||||
import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js';
|
||||
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
|
||||
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
|
||||
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
|
||||
|
||||
async function createEnvelopeHarness(): Promise<{ app: FastifyInstance; ctx: MockRouteContext }> {
|
||||
const app = Fastify({ logger: false });
|
||||
await app.register(fastifyCookie);
|
||||
const ctx = createMockRouteContext();
|
||||
registerSessionRoutes(app, ctx as never);
|
||||
|
||||
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
|
||||
if (!req.url.startsWith('/api')) return done(null, payload);
|
||||
if (payload === null || typeof payload !== 'object') return done(null, payload);
|
||||
const p = payload as { success?: unknown; errorCode?: unknown };
|
||||
if (p.success === false) {
|
||||
if (reply.statusCode === 200 && typeof p.errorCode === 'string') {
|
||||
reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
|
||||
}
|
||||
return done(null, payload);
|
||||
}
|
||||
if (p.success === true) return done(null, payload);
|
||||
return done(null, { success: true, data: payload });
|
||||
});
|
||||
|
||||
installRouteErrorHandler(app);
|
||||
await app.ready();
|
||||
return { app, ctx };
|
||||
}
|
||||
|
||||
type Internals = { _appliedInputSeq: Map<string, number> };
|
||||
const seqOf = (s: Session, client: string) => (s as unknown as Internals)._appliedInputSeq.get(client);
|
||||
|
||||
describe('input dedup bookkeeping', () => {
|
||||
const make = () => new Session({ workingDir: '/tmp', mode: 'claude' });
|
||||
|
||||
it('accepts an increasing seq once and rejects the replay', () => {
|
||||
const s = make();
|
||||
expect(s.shouldApplyInput('c1', 1)).toBe(true);
|
||||
expect(s.shouldApplyInput('c1', 1)).toBe(false);
|
||||
expect(s.shouldApplyInput('c1', 2)).toBe(true);
|
||||
});
|
||||
|
||||
it('forgetInputSeq re-opens a failed delivery for retry', () => {
|
||||
const s = make();
|
||||
expect(s.shouldApplyInput('c1', 7)).toBe(true);
|
||||
|
||||
s.forgetInputSeq('c1', 7); // the write failed after the 200 went out
|
||||
|
||||
expect(s.shouldApplyInput('c1', 7)).toBe(true);
|
||||
});
|
||||
|
||||
it('does not re-open a seq that a later input has superseded', () => {
|
||||
// Rolling back blindly would let an old, already-superseded message replay.
|
||||
const s = make();
|
||||
s.shouldApplyInput('c1', 7);
|
||||
s.shouldApplyInput('c1', 8);
|
||||
|
||||
s.forgetInputSeq('c1', 7);
|
||||
|
||||
expect(s.shouldApplyInput('c1', 8)).toBe(false);
|
||||
expect(seqOf(s, 'c1')).toBe(8);
|
||||
});
|
||||
|
||||
it('is scoped per client', () => {
|
||||
const s = make();
|
||||
s.shouldApplyInput('c1', 5);
|
||||
s.forgetInputSeq('c2', 5);
|
||||
expect(s.shouldApplyInput('c1', 5)).toBe(false);
|
||||
});
|
||||
|
||||
it('tolerates a rollback for a client that was never seen', () => {
|
||||
const s = make();
|
||||
expect(() => s.forgetInputSeq('ghost', 3)).not.toThrow();
|
||||
});
|
||||
});
|
||||
|
||||
describe('Session.write delivery signal', () => {
|
||||
it('reports false when there is no PTY instead of swallowing the data', () => {
|
||||
// The silent swallow was the third way input could vanish: no PTY, no error,
|
||||
// no return value — the caller had no way to know.
|
||||
const s = new Session({ workingDir: '/tmp', mode: 'claude' });
|
||||
expect(s.write('hello\r')).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* Wiring, not primitives.
|
||||
*
|
||||
* The first version of this file tested Session directly and nothing else: reverting
|
||||
* the route to master — deleting the rollback call, the load-bearing half of the fix —
|
||||
* left all six tests green. These drive the actual HTTP route.
|
||||
*/
|
||||
describe('POST /api/sessions/:id/input rollback wiring', () => {
|
||||
let harness: { app: FastifyInstance; ctx: MockRouteContext };
|
||||
|
||||
beforeEach(async () => {
|
||||
harness = await createEnvelopeHarness();
|
||||
});
|
||||
afterEach(async () => {
|
||||
await harness.app.close();
|
||||
});
|
||||
|
||||
const post = (body: Record<string, unknown>, id = 'test-session-1') =>
|
||||
harness.app.inject({ method: 'POST', url: `/api/sessions/${id}/input`, payload: body });
|
||||
|
||||
it('rolls the seq back when both the mux write and the direct write fail', async () => {
|
||||
const session = harness.ctx.sessions.get('test-session-1')!;
|
||||
session.failWrites = true; // writeViaMux false AND write() false
|
||||
|
||||
await post({ input: 'lost\r', useMux: true, clientId: 'c1', seq: 1 });
|
||||
await new Promise((r) => setTimeout(r, 20)); // the mux write is fire-and-forget
|
||||
|
||||
// The retry the client would make must be accepted, not swallowed as a duplicate.
|
||||
expect(session.shouldApplyInput('c1', 1)).toBe(true);
|
||||
});
|
||||
|
||||
it('keeps the seq burnt when delivery succeeded', async () => {
|
||||
const session = harness.ctx.sessions.get('test-session-1')!;
|
||||
|
||||
await post({ input: 'fine\r', useMux: true, clientId: 'c1', seq: 1 });
|
||||
await new Promise((r) => setTimeout(r, 20));
|
||||
|
||||
expect(session.shouldApplyInput('c1', 1)).toBe(false);
|
||||
});
|
||||
|
||||
it('rolls the seq back on the non-mux path too', async () => {
|
||||
// Still a 200: a session may legitimately have no PTY yet, and turning that
|
||||
// into a failure status would be a contract change. Re-opening the seq is not.
|
||||
const session = harness.ctx.sessions.get('test-session-1')!;
|
||||
session.failWrites = true;
|
||||
|
||||
const res = await post({ input: 'x\r', useMux: false, clientId: 'c2', seq: 5 });
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(session.shouldApplyInput('c2', 5)).toBe(true);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,671 @@
|
||||
/**
|
||||
* @fileoverview Route tests for the `wait` field on `POST /api/sessions/:id/input`.
|
||||
*
|
||||
* This endpoint exists to close a race a caller cannot close from outside: between
|
||||
* the write landing and the session flipping to `working`, a SEPARATE wait sees the
|
||||
* session still idle and returns instantly, reporting the previous turn as this one.
|
||||
* Registering the waiter before the write is the whole point, so that is what the
|
||||
* first test pins.
|
||||
*
|
||||
* It also pins that the historical fire-and-forget path is untouched when `wait` is
|
||||
* absent, that a capacity rejection gives the dedup seq back (otherwise the caller's
|
||||
* retry is refused as a duplicate and the input is lost by the very mechanism
|
||||
* reliable delivery exists for), and that `delivered` reports what actually happened
|
||||
* to the write rather than merely "this was not a duplicate".
|
||||
*
|
||||
* Plan: docs/agent-control-plan.md
|
||||
*/
|
||||
import { describe, it, expect, afterEach, beforeAll, afterAll, vi } from 'vitest';
|
||||
import fastifyCookie from '@fastify/cookie';
|
||||
import Fastify, { type FastifyInstance } from 'fastify';
|
||||
import type { IncomingMessage } from 'node:http';
|
||||
import { registerSessionRoutes, _resetPaneLivenessState } from '../../src/web/routes/session-routes.js';
|
||||
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
|
||||
import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js';
|
||||
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
|
||||
import { sessionWaits } from '../../src/web/session-wait-registry.js';
|
||||
import { MAX_WAIT_MS } from '../../src/config/agent-wait.js';
|
||||
|
||||
// Distinct per file on purpose: the three wait suites share the process-wide
|
||||
// `sessionWaits` singleton, so a common id let one file's leftover waiter be counted
|
||||
// by another's assertion. Failed only in a 5-file run, which is how CI runs them.
|
||||
const SESSION_ID = 'input-wait-session';
|
||||
const URL = `/api/sessions/${SESSION_ID}/input`;
|
||||
|
||||
afterEach(() => {
|
||||
// Deliberately not `cancelEverything()`: it latches the registry's stopped flag,
|
||||
// which would leave every later test in this file talking to a dead registry.
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
_resetPaneLivenessState();
|
||||
});
|
||||
|
||||
async function harness(): Promise<{ app: FastifyInstance; ctx: MockRouteContext; rawRequests: IncomingMessage[] }> {
|
||||
const app = Fastify({ logger: false });
|
||||
await app.register(fastifyCookie);
|
||||
const ctx = createMockRouteContext({ sessionId: SESSION_ID });
|
||||
const rawRequests: IncomingMessage[] = [];
|
||||
app.addHook('onRequest', async (req) => {
|
||||
rawRequests.push(req.raw);
|
||||
});
|
||||
|
||||
registerSessionRoutes(app, ctx as never);
|
||||
|
||||
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
|
||||
const p = payload as { success?: unknown; errorCode?: unknown } | null;
|
||||
if (p && typeof p === 'object' && p.success === false && reply.statusCode === 200) {
|
||||
if (typeof p.errorCode === 'string') reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
|
||||
}
|
||||
return done(null, payload);
|
||||
});
|
||||
|
||||
installRouteErrorHandler(app);
|
||||
await app.ready();
|
||||
return { app, ctx, rawRequests };
|
||||
}
|
||||
|
||||
const send = (app: FastifyInstance, payload: Record<string, unknown>) =>
|
||||
app.inject({ method: 'POST', url: URL, payload });
|
||||
|
||||
describe('POST /api/sessions/:id/input without wait (unchanged behavior)', () => {
|
||||
it('returns the historical bare body and registers no waiter', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await send(app, { input: 'hello', useMux: true });
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json()).toEqual({});
|
||||
expect(sessionWaits.totalWaiterCount()).toBe(0);
|
||||
});
|
||||
|
||||
it('still returns before the mux write settles', async () => {
|
||||
// The fire-and-forget property is why the response is fast; send-and-wait must
|
||||
// not have turned every input into an awaited tmux round-trip.
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
let resolveWrite: (ok: boolean) => void = () => {};
|
||||
session.writeViaMux = () => new Promise<boolean>((resolve) => (resolveWrite = resolve));
|
||||
|
||||
const res = await send(app, { input: 'hello', useMux: true });
|
||||
expect(res.json()).toEqual({});
|
||||
resolveWrite(true);
|
||||
});
|
||||
|
||||
it('a tagged duplicate still returns the bare body', async () => {
|
||||
const { app } = await harness();
|
||||
await send(app, { input: 'first', clientId: 'c1', seq: 1 });
|
||||
const replay = await send(app, { input: 'first', clientId: 'c1', seq: 1 });
|
||||
|
||||
expect(replay.json()).toEqual({});
|
||||
expect(sessionWaits.totalWaiterCount()).toBe(0);
|
||||
});
|
||||
|
||||
it('a failed direct write is still not an error response', async () => {
|
||||
// A session can legitimately have no PTY yet, and callers have always been able
|
||||
// to write to one without a 4xx. Only the `wait` path reports delivery.
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.failWrites = true;
|
||||
|
||||
const res = await send(app, { input: 'x' });
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json()).toEqual({});
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/input with wait', () => {
|
||||
it('registers the waiter BEFORE the write, so the pre-existing idle state cannot satisfy it', async () => {
|
||||
// The mock session is idle. A naive send-then-wait would answer immediately with
|
||||
// that stale idle; this must block until a real transition.
|
||||
const { app } = await harness();
|
||||
const pending = send(app, { input: 'run the tests', useMux: true, wait: true });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
sessionWaits.notifySignal(SESSION_ID, 'stop');
|
||||
const body = (await pending).json();
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.data.delivered).toBe(true);
|
||||
expect(body.data.duplicate).toBe(false);
|
||||
expect(body.data.wait.signal).toBe('stop');
|
||||
expect(body.data.wait.immediate).toBe(false);
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
});
|
||||
|
||||
it('uses the same data.wait envelope as the two GET endpoints', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = send(app, { input: 'x', wait: 'stop' });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifySignal(SESSION_ID, 'stop');
|
||||
const { data } = (await pending).json();
|
||||
|
||||
expect(Object.keys(data).sort()).toEqual(['delivered', 'duplicate', 'limitPaused', 'status', 'wait']);
|
||||
expect(data.status).toBe('idle');
|
||||
expect(data.limitPaused).toBe(false);
|
||||
expect(data.wait.aborted).toBe(false);
|
||||
});
|
||||
|
||||
it('echoes the effective timeout after clamping', async () => {
|
||||
// The schema accepts up to 24h; the server caps at MAX_WAIT_MS. Without the echo
|
||||
// an agent reads the cap as "my 24h wait elapsed" and kills a healthy worker.
|
||||
const { app } = await harness();
|
||||
const pending = send(app, { input: 'x', wait: 'stop', waitTimeout: 86_400_000 });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifySignal(SESSION_ID, 'stop');
|
||||
|
||||
expect((await pending).json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
|
||||
});
|
||||
|
||||
it('delivers the input before blocking', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
const pending = send(app, { input: 'echo hi', useMux: true, wait: 'stop' });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
// The write happened while the request is still open.
|
||||
expect(session.writeBuffer.join('')).toContain('echo hi');
|
||||
|
||||
sessionWaits.notifySignal(SESSION_ID, 'stop');
|
||||
await pending;
|
||||
});
|
||||
|
||||
it('wait: true uses the default signal set', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = send(app, { input: 'x', wait: true });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifySignal(SESSION_ID, 'idle');
|
||||
|
||||
expect((await pending).json().data.wait.until).toEqual(['stop', 'idle', 'exit']);
|
||||
});
|
||||
|
||||
it('accepts an explicit signal list', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = send(app, { input: 'x', wait: 'exit' });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
// Not one of the requested signals: the wait must not resolve on it.
|
||||
expect(sessionWaits.notifySignal(SESSION_ID, 'idle')).toBe(0);
|
||||
sessionWaits.notifySignal(SESSION_ID, 'exit');
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.signal).toBe('exit');
|
||||
expect(body.data.wait.until).toEqual(['exit']);
|
||||
});
|
||||
|
||||
it('accepts an array, the same grammar the query parameter takes', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = send(app, { input: 'x', wait: ['stop', 'exit'] });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifySignal(SESSION_ID, 'exit');
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.until).toEqual(['stop', 'exit']);
|
||||
expect(body.data.wait.signal).toBe('exit');
|
||||
});
|
||||
|
||||
it('rejects an unknown wait signal without writing', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
const before = session.writeBuffer.length;
|
||||
|
||||
const res = await send(app, { input: 'x', wait: 'stpo' });
|
||||
expect(res.statusCode).toBe(400);
|
||||
expect(res.json().errorCode).toBe('INVALID_INPUT');
|
||||
expect(session.writeBuffer.length).toBe(before);
|
||||
});
|
||||
|
||||
it('rejects a hook-only signal for external CLI modes', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.mode = 'codex';
|
||||
|
||||
const res = await send(app, { input: 'x', wait: 'stop' });
|
||||
expect(res.statusCode).toBe(400);
|
||||
expect(res.json().error).toContain('codex');
|
||||
});
|
||||
|
||||
it('rejects a hook-only signal for a shell session too', async () => {
|
||||
// Not an external CLI, but a plain bash PTY installs no hooks either, so `stop`
|
||||
// could only ever time out.
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.mode = 'shell';
|
||||
|
||||
const res = await send(app, { input: 'x', wait: 'stop' });
|
||||
expect(res.statusCode).toBe(400);
|
||||
expect(res.json().error).toContain('shell');
|
||||
});
|
||||
|
||||
it('times out as a 200, like the standalone wait', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await send(app, { input: 'x', wait: 'blocked', waitTimeout: 1 });
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
const body = res.json();
|
||||
expect(body.data.delivered).toBe(true);
|
||||
expect(body.data.wait.timedOut).toBe(true);
|
||||
expect(body.data.wait.signal).toBeNull();
|
||||
});
|
||||
|
||||
it('resolves with ended when the session goes away mid-wait', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = send(app, { input: 'x', wait: 'stop' });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
|
||||
expect((await pending).json().data.wait.ended).toBe(true);
|
||||
});
|
||||
|
||||
it('a request-body close does NOT abort the wait', async () => {
|
||||
// The regression this pins: on a POST the request stream closes as soon as the
|
||||
// body has been read, well before the handler blocks. Treating that as a hang-up
|
||||
// aborted every send-and-wait instantly. Hang-up handling itself is proven over
|
||||
// real HTTP at the bottom of this file, because inject cannot produce a socket.
|
||||
const { app, rawRequests } = await harness();
|
||||
const pending = send(app, { input: 'x', wait: 'stop', waitTimeout: 600_000 });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
rawRequests[rawRequests.length - 1].emit('close');
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
sessionWaits.notifySignal(SESSION_ID, 'stop');
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.aborted).toBe(false);
|
||||
expect(body.data.wait.signal).toBe('stop');
|
||||
});
|
||||
|
||||
it('a duplicate skips the write but answers from current state instead of hanging', async () => {
|
||||
// The original turn is long over, so requiring a fresh transition here would
|
||||
// block a redelivery until timeout for no reason.
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
await send(app, { input: 'first', clientId: 'c1', seq: 1 });
|
||||
const before = session.writeBuffer.length;
|
||||
|
||||
const replay = await send(app, { input: 'first', clientId: 'c1', seq: 1, wait: 'idle' });
|
||||
const body = replay.json();
|
||||
|
||||
expect(body.data.duplicate).toBe(true);
|
||||
expect(body.data.delivered).toBe(false);
|
||||
expect(body.data.wait.immediate).toBe(true);
|
||||
expect(body.data.wait.signal).toBe('idle');
|
||||
expect(session.writeBuffer.length).toBe(before);
|
||||
});
|
||||
|
||||
it('a duplicate on a BUSY session answers working, not idle', async () => {
|
||||
// Previously unreachable: MockSession's status was 'working', which is not a
|
||||
// SessionStatus, so signalForStatus fell through to null and this combination
|
||||
// silently proved nothing.
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
await send(app, { input: 'first', clientId: 'c2', seq: 1 });
|
||||
session.status = 'busy';
|
||||
|
||||
const replay = await send(app, { input: 'first', clientId: 'c2', seq: 1, wait: 'working' });
|
||||
const body = replay.json();
|
||||
|
||||
expect(body.data.duplicate).toBe(true);
|
||||
expect(body.data.wait.signal).toBe('working');
|
||||
expect(body.data.wait.immediate).toBe(true);
|
||||
expect(body.data.status).toBe('busy');
|
||||
});
|
||||
|
||||
it('gives the dedup seq back when a full waiter pool rejects the request', async () => {
|
||||
// Otherwise the caller's retry is refused as a duplicate and the input vanishes.
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
|
||||
const pendings = [];
|
||||
for (let i = 0; i < 16; i++) {
|
||||
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` }));
|
||||
}
|
||||
await new Promise((resolve) => setTimeout(resolve, 40));
|
||||
expect(sessionWaits.waiterCount(SESSION_ID)).toBe(16);
|
||||
|
||||
const rejected = await send(app, { input: 'x', clientId: 'c9', seq: 7, wait: true });
|
||||
expect(rejected.statusCode).toBe(409);
|
||||
expect(rejected.json().errorCode).toBe('SESSION_BUSY');
|
||||
// Nothing written...
|
||||
expect(session.writeBuffer.join('')).not.toContain('x');
|
||||
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
await Promise.all(pendings);
|
||||
|
||||
// ...and the same seq is accepted on retry rather than treated as a replay.
|
||||
const retry = await send(app, { input: 'x', clientId: 'c9', seq: 7 });
|
||||
expect(retry.json()).toEqual({});
|
||||
expect(session.writeBuffer.join('')).toContain('x');
|
||||
});
|
||||
|
||||
it('a null wait is treated as absent, not as a validation error', async () => {
|
||||
// Zod .optional() rejects null, and a third-party caller building the body with
|
||||
// JSON.stringify keeps an explicit null on the wire.
|
||||
const { app } = await harness();
|
||||
const res = await send(app, { input: 'x', wait: null, waitTimeout: null });
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json()).toEqual({});
|
||||
});
|
||||
|
||||
it('wait: false is treated as absent', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await send(app, { input: 'x', wait: false });
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json()).toEqual({});
|
||||
expect(sessionWaits.totalWaiterCount()).toBe(0);
|
||||
});
|
||||
|
||||
it('falls back to a direct write when the mux write fails, and still waits', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
session.writeViaMux = async () => false;
|
||||
|
||||
const pending = send(app, { input: 'fallback me', useMux: true, wait: 'stop' });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(session.writeBuffer.join('')).toContain('fallback me');
|
||||
|
||||
sessionWaits.notifySignal(SESSION_ID, 'stop');
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.signal).toBe('stop');
|
||||
// The fallback write succeeded, so the input really was delivered.
|
||||
expect(body.data.delivered).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/input: delivered reports the write, not just the dedup', () => {
|
||||
it('reports delivered:false when BOTH write paths fail, instead of claiming delivery', async () => {
|
||||
// A worker whose PTY has exited fails writeViaMux AND write. Reporting
|
||||
// "delivered, but it timed out" points the agent at waiting longer; the truth is
|
||||
// "restart the worker".
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
session.failWrites = true;
|
||||
|
||||
const res = await send(app, { input: 'run the tests', useMux: true, wait: 'stop', waitTimeout: 600_000 });
|
||||
const body = res.json();
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(body.data.delivered).toBe(false);
|
||||
expect(body.data.duplicate).toBe(false);
|
||||
});
|
||||
|
||||
it('does not block for the full timeout on an input it knows never landed', async () => {
|
||||
// The waiter has to be registered before the write, so it exists by the time the
|
||||
// failure is known; releasing it immediately is what keeps the caller from
|
||||
// waiting ten minutes for a turn that cannot start.
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.failWrites = true;
|
||||
|
||||
const started = Date.now();
|
||||
const body = (await send(app, { input: 'x', useMux: true, wait: 'stop', waitTimeout: 600_000 })).json();
|
||||
|
||||
expect(Date.now() - started).toBeLessThan(2_000);
|
||||
expect(body.data.delivered).toBe(false);
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
|
||||
});
|
||||
|
||||
it('reports delivered:false for a failed direct (non-mux) write too', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.failWrites = true;
|
||||
|
||||
const body = (await send(app, { input: 'x', wait: 'stop', waitTimeout: 600_000 })).json();
|
||||
expect(body.data.delivered).toBe(false);
|
||||
});
|
||||
|
||||
it('rolls the dedup seq back when the write failed, so a retry is not a duplicate', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
session.failWrites = true;
|
||||
|
||||
const first = (await send(app, { input: 'x', clientId: 'c3', seq: 4, wait: 'stop', waitTimeout: 600_000 })).json();
|
||||
expect(first.data.delivered).toBe(false);
|
||||
|
||||
session.failWrites = false;
|
||||
const retry = (await send(app, { input: 'x', clientId: 'c3', seq: 4, wait: 'stop', waitTimeout: 1 })).json();
|
||||
expect(retry.data.duplicate).toBe(false);
|
||||
expect(retry.data.delivered).toBe(true);
|
||||
expect(session.writeBuffer.join('')).toContain('x');
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* Client-hang-up handling, over REAL HTTP.
|
||||
*
|
||||
* `app.inject()` never emits a `close` event at all, so the entire abort path is
|
||||
* invisible to every other test in this file — and the failure it hides is not
|
||||
* subtle. On a POST, `req.raw` emits `'close'` as soon as the request BODY finishes
|
||||
* streaming, which happens before the handler blocks (+1ms, `aborted: false`) and is
|
||||
* indistinguishable from a real hang-up at +0ms. A request-side abort listener
|
||||
* therefore cancels every send-and-wait instantly: `POST .../input {wait:"exit",
|
||||
* waitTimeout:10000}` came back in 23ms with `ended:true, aborted:true, waitedMs:6`,
|
||||
* i.e. the feature was dead while all 27 inject-based tests above stayed green.
|
||||
*
|
||||
* GET survives a request-side listener because it has no body to finish, which is
|
||||
* exactly why this regression needs a POST and a real socket to catch.
|
||||
*/
|
||||
describe('POST /api/sessions/:id/input over real HTTP: hang-up handling', () => {
|
||||
const PORT = 3181;
|
||||
const base = `http://127.0.0.1:${PORT}`;
|
||||
let app: FastifyInstance;
|
||||
|
||||
beforeAll(async () => {
|
||||
app = (await harness()).app;
|
||||
await app.listen({ port: PORT, host: '127.0.0.1' });
|
||||
});
|
||||
|
||||
afterAll(async () => {
|
||||
// fetch keeps its sockets alive, and `app.close()` waits for idle connections,
|
||||
// so without this the teardown hook times out.
|
||||
app.server.closeAllConnections();
|
||||
await app.close();
|
||||
});
|
||||
|
||||
it('a send-and-wait that is NOT aborted blocks for its full timeout', async () => {
|
||||
// The regression: this returned in ~20ms with aborted:true.
|
||||
const started = Date.now();
|
||||
const res = await fetch(`${base}/api/sessions/${SESSION_ID}/input`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ input: 'run the tests', wait: 'exit', waitTimeout: 2000 }),
|
||||
});
|
||||
const elapsed = Date.now() - started;
|
||||
const body = await res.json();
|
||||
|
||||
expect(body.data.wait.aborted).toBe(false);
|
||||
expect(body.data.wait.timedOut).toBe(true);
|
||||
expect(body.data.wait.waitedMs).toBeGreaterThan(1500);
|
||||
expect(elapsed).toBeGreaterThan(1500);
|
||||
expect(body.data.delivered).toBe(true);
|
||||
});
|
||||
|
||||
it('a send-and-wait aborted mid-flight frees its waiter', async () => {
|
||||
const controller = new AbortController();
|
||||
const pending = fetch(`${base}/api/sessions/${SESSION_ID}/input`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ input: 'x', wait: 'exit', waitTimeout: 600_000 }),
|
||||
signal: controller.signal,
|
||||
}).catch(() => 'aborted');
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 150));
|
||||
// Still parked: the body finished streaming long ago, and that must not count.
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
controller.abort();
|
||||
expect(await pending).toBe('aborted');
|
||||
await new Promise((resolve) => setTimeout(resolve, 100));
|
||||
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
|
||||
});
|
||||
|
||||
it('a GET wait behaves the same way on both counts', async () => {
|
||||
const notAborted = await fetch(`${base}/api/sessions/${SESSION_ID}/wait?until=stop&timeout=1500`);
|
||||
const body = await notAborted.json();
|
||||
expect(body.data.wait.aborted).toBe(false);
|
||||
expect(body.data.wait.timedOut).toBe(true);
|
||||
|
||||
const controller = new AbortController();
|
||||
const pending = fetch(`${base}/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000`, {
|
||||
signal: controller.signal,
|
||||
}).catch(() => 'aborted');
|
||||
await new Promise((resolve) => setTimeout(resolve, 100));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
controller.abort();
|
||||
expect(await pending).toBe('aborted');
|
||||
await new Promise((resolve) => setTimeout(resolve, 100));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
|
||||
});
|
||||
|
||||
it('wait-output frees its waiter on hang-up too', async () => {
|
||||
const controller = new AbortController();
|
||||
const pending = fetch(`${base}/api/sessions/${SESSION_ID}/wait-output?match=NEVER&timeout=600000`, {
|
||||
signal: controller.signal,
|
||||
}).catch(() => 'aborted');
|
||||
await new Promise((resolve) => setTimeout(resolve, 100));
|
||||
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
controller.abort();
|
||||
expect(await pending).toBe('aborted');
|
||||
await new Promise((resolve) => setTimeout(resolve, 100));
|
||||
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(0);
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* Send-and-wait against a tmux worker that has already died.
|
||||
*
|
||||
* `tmux send-keys` SUCCEEDS against a dead pane, so `writeViaMux` returns true and the
|
||||
* old `delivered` was true for bytes written into a corpse — with `timedOut: true`
|
||||
* alongside it, which tells an agent to wait longer when the truth is "restart the
|
||||
* worker". Live: `pane_dead=1 status=42`, Codeman `pid=309406 status=idle`,
|
||||
* `delivered: true`.
|
||||
*/
|
||||
describe('POST /api/sessions/:id/input: the pane is dead', () => {
|
||||
function setPaneDead(ctx: MockRouteContext, dead: boolean) {
|
||||
(ctx.mux as unknown as { isPaneDead: (n: string) => boolean }).isPaneDead = () => dead;
|
||||
}
|
||||
|
||||
it('reports delivered:false even though the mux write "succeeded"', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
setPaneDead(ctx, true);
|
||||
|
||||
const res = await send(app, { input: 'run the tests', useMux: true, wait: 'stop', waitTimeout: 600_000 });
|
||||
const body = res.json();
|
||||
|
||||
// The write itself did not fail — that is the whole trap.
|
||||
expect(session.writeBuffer.join('')).toContain('run the tests');
|
||||
expect(body.data.delivered).toBe(false);
|
||||
expect(body.data.duplicate).toBe(false);
|
||||
});
|
||||
|
||||
it('returns at once instead of blocking on a turn that cannot start', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
setPaneDead(ctx, true);
|
||||
|
||||
const started = Date.now();
|
||||
const body = (await send(app, { input: 'x', useMux: true, wait: 'stop', waitTimeout: 600_000 })).json();
|
||||
|
||||
expect(Date.now() - started).toBeLessThan(2_000);
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
|
||||
});
|
||||
|
||||
it('rolls the dedup seq back, so a retry against a restarted worker is not a duplicate', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
setPaneDead(ctx, true);
|
||||
|
||||
const dead = (await send(app, { input: 'x', useMux: true, clientId: 'c7', seq: 3, wait: 'stop' })).json();
|
||||
expect(dead.data.delivered).toBe(false);
|
||||
|
||||
// Worker restarted.
|
||||
setPaneDead(ctx, false);
|
||||
_resetPaneLivenessState();
|
||||
const retry = (
|
||||
await send(app, { input: 'x', useMux: true, clientId: 'c7', seq: 3, wait: 'stop', waitTimeout: 1 })
|
||||
).json();
|
||||
expect(retry.data.duplicate).toBe(false);
|
||||
expect(retry.data.delivered).toBe(true);
|
||||
expect(session.writeBuffer.join('')).toContain('x');
|
||||
});
|
||||
|
||||
it('a live pane still reports delivered:true', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
setPaneDead(ctx, false);
|
||||
|
||||
const pending = send(app, { input: 'x', useMux: true, wait: 'stop' });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifySignal(SESSION_ID, 'stop');
|
||||
|
||||
expect((await pending).json().data.delivered).toBe(true);
|
||||
});
|
||||
|
||||
it('never probes tmux on the plain (non-wait) input path', async () => {
|
||||
// The browser sends thousands of these per session; they must not exec tmux.
|
||||
const { app, ctx } = await harness();
|
||||
const probe = vi.fn(() => false);
|
||||
(ctx.mux as unknown as { isPaneDead: (n: string) => boolean }).isPaneDead = probe as never;
|
||||
|
||||
await send(app, { input: 'hello', useMux: true });
|
||||
await send(app, { input: 'hello again', useMux: true, clientId: 'c1', seq: 1 });
|
||||
|
||||
expect(probe).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/input: `aborted` stays a client-side fact', () => {
|
||||
it('reports aborted:false when the SERVER released the waiter after a failed delivery', async () => {
|
||||
// api-reference guarantees a client never sees `aborted: true`, because it means
|
||||
// "you hung up, nobody is reading this". The release below is the server giving up
|
||||
// on a write that failed — and the client IS reading the response, so reporting
|
||||
// `aborted: true` would both break that guarantee and hand an agent a second,
|
||||
// contradictory reason for an outcome `delivered: false` already explains.
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.failWrites = true;
|
||||
|
||||
const body = (await send(app, { input: 'x', useMux: true, wait: 'stop', waitTimeout: 600_000 })).json();
|
||||
|
||||
expect(body.data.delivered).toBe(false);
|
||||
expect(body.data.wait.aborted).toBe(false);
|
||||
expect(body.data.wait.ended).toBe(true);
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
});
|
||||
|
||||
it('the same holds for a dead pane', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
(ctx.mux as unknown as { isPaneDead: () => boolean }).isPaneDead = () => true;
|
||||
|
||||
const body = (await send(app, { input: 'x', useMux: true, wait: 'stop', waitTimeout: 600_000 })).json();
|
||||
expect(body.data.wait.aborted).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/input: an oversized waitTimeout clamps', () => {
|
||||
it('accepts a value above the old schema ceiling and reports the clamp', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = send(app, { input: 'x', wait: 'stop', waitTimeout: 99_999_999_999 });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifySignal(SESSION_ID, 'stop');
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.timeoutMs).toBe(MAX_WAIT_MS);
|
||||
});
|
||||
|
||||
it('still rejects a non-integer or negative waitTimeout', async () => {
|
||||
const { app } = await harness();
|
||||
for (const value of [-1, 0, 1.5]) {
|
||||
const res = await send(app, { input: 'x', wait: 'stop', waitTimeout: value });
|
||||
expect(res.statusCode, `waitTimeout=${value}`).toBe(400);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,432 @@
|
||||
/**
|
||||
* @fileoverview Route tests for `GET /api/sessions/:id/wait-output`.
|
||||
*
|
||||
* Same 200-on-timeout contract and same `data.wait` envelope as `/wait`. The
|
||||
* additional things pinned here:
|
||||
* - matching is LITERAL, and a `regex` parameter is rejected rather than ignored,
|
||||
* so an agent that assumed herdr's `--regex` cannot silently wait on the wrong thing;
|
||||
* - `from=buffer` scans what already scrolled past, bounded to a tail of the buffer,
|
||||
* and is charged against the waiter cap BEFORE it materializes that buffer;
|
||||
* - a chunk-straddling match still resolves, since PTY chunking is arbitrary;
|
||||
* - a client that hangs up frees its waiter instead of holding it to the timeout.
|
||||
*
|
||||
* Plan: docs/agent-control-plan.md
|
||||
*/
|
||||
import { describe, it, expect, afterEach, vi } from 'vitest';
|
||||
import fastifyCookie from '@fastify/cookie';
|
||||
import Fastify, { type FastifyInstance } from 'fastify';
|
||||
import type { ServerResponse } from 'node:http';
|
||||
import { registerSessionRoutes, _resetPaneLivenessState } from '../../src/web/routes/session-routes.js';
|
||||
import { createSessionListeners, attachSessionListeners } from '../../src/web/session-listener-wiring.js';
|
||||
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
|
||||
import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js';
|
||||
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
|
||||
import { sessionWaits } from '../../src/web/session-wait-registry.js';
|
||||
import { MAX_MATCH_LENGTH, MAX_BUFFER_SCAN_BYTES, MAX_WAIT_MS } from '../../src/config/agent-wait.js';
|
||||
|
||||
// Distinct per file on purpose: the three wait suites share the process-wide
|
||||
// `sessionWaits` singleton, so a common id let one file's leftover waiter be counted
|
||||
// by another's assertion. Failed only in a 5-file run, which is how CI runs them.
|
||||
const SESSION_ID = 'wait-output-session';
|
||||
const URL = `/api/sessions/${SESSION_ID}/wait-output`;
|
||||
|
||||
afterEach(() => {
|
||||
// Deliberately not `cancelEverything()`: it latches the registry's stopped flag,
|
||||
// which would leave every later test in this file talking to a dead registry.
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
_resetPaneLivenessState();
|
||||
});
|
||||
|
||||
/** Mirrors production's errorCode-to-status mapping; without it negative cases pass vacuously. */
|
||||
async function harness(): Promise<{ app: FastifyInstance; ctx: MockRouteContext; rawReplies: ServerResponse[] }> {
|
||||
const app = Fastify({ logger: false });
|
||||
await app.register(fastifyCookie);
|
||||
const ctx = createMockRouteContext({ sessionId: SESSION_ID });
|
||||
const rawReplies: ServerResponse[] = [];
|
||||
app.addHook('onRequest', async (req, reply) => {
|
||||
// The RESPONSE, because that is what the handler's hang-up detection listens to.
|
||||
rawReplies.push(reply.raw);
|
||||
});
|
||||
|
||||
registerSessionRoutes(app, ctx as never);
|
||||
|
||||
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
|
||||
const p = payload as { success?: unknown; errorCode?: unknown } | null;
|
||||
if (p && typeof p === 'object' && p.success === false && reply.statusCode === 200) {
|
||||
if (typeof p.errorCode === 'string') reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
|
||||
}
|
||||
return done(null, payload);
|
||||
});
|
||||
|
||||
installRouteErrorHandler(app);
|
||||
await app.ready();
|
||||
return { app, ctx, rawReplies };
|
||||
}
|
||||
|
||||
describe('GET /api/sessions/:id/wait-output', () => {
|
||||
it('resolves when the string appears on the stream', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK` });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
|
||||
sessionWaits.notifyOutput(SESSION_ID, 'running tests...\nBUILD OK\n');
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.data.wait.matched).toBe(true);
|
||||
expect(body.data.wait.immediate).toBe(false);
|
||||
expect(body.data.wait.snippet).toContain('BUILD OK');
|
||||
expect(body.data.wait.match).toBe('BUILD OK');
|
||||
expect(body.data.sessionId).toBe(SESSION_ID);
|
||||
});
|
||||
|
||||
it('uses the same data.wait envelope as /wait, so one client helper reads both', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=never&timeout=1` });
|
||||
|
||||
const { data } = res.json();
|
||||
expect(Object.keys(data).sort()).toEqual(['limitPaused', 'sessionId', 'status', 'wait']);
|
||||
expect(data.matched).toBeUndefined();
|
||||
expect(data.wait.matched).toBe(false);
|
||||
expect(data.wait.aborted).toBe(false);
|
||||
});
|
||||
|
||||
it('sends Cache-Control: no-store', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=never&timeout=1` });
|
||||
|
||||
expect(res.headers['cache-control']).toBe('no-store');
|
||||
});
|
||||
|
||||
it('echoes the effective timeout after clamping', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.terminalBuffer = 'BUILD OK\n';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK&from=buffer&timeout=1800000` });
|
||||
expect(res.json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
|
||||
});
|
||||
|
||||
it('strips ANSI before matching, so colored output still matches', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK` });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifyOutput(SESSION_ID, '\x1b[32mBUILD\x1b[0m OK\n');
|
||||
|
||||
expect((await pending).json().data.wait.matched).toBe(true);
|
||||
});
|
||||
|
||||
it('matches across a chunk boundary', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK` });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifyOutput(SESSION_ID, 'trailing text BUIL');
|
||||
sessionWaits.notifyOutput(SESSION_ID, 'D OK done');
|
||||
|
||||
expect((await pending).json().data.wait.matched).toBe(true);
|
||||
});
|
||||
|
||||
it('is case-sensitive by default and honors nocase=1', async () => {
|
||||
const { app } = await harness();
|
||||
|
||||
const strict = app.inject({ method: 'GET', url: `${URL}?match=build%20ok&timeout=1` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifyOutput(SESSION_ID, 'BUILD OK');
|
||||
expect((await strict).json().data.wait.timedOut).toBe(true);
|
||||
|
||||
const loose = app.inject({ method: 'GET', url: `${URL}?match=build%20ok&nocase=1` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifyOutput(SESSION_ID, 'BUILD OK');
|
||||
const body = (await loose).json();
|
||||
expect(body.data.wait.matched).toBe(true);
|
||||
// Reported in the terminal's own casing, not the caller's.
|
||||
expect(body.data.wait.snippet).toContain('BUILD OK');
|
||||
});
|
||||
|
||||
it('from=buffer resolves immediately against output that already scrolled past', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.terminalBuffer = 'earlier output\n\x1b[32mBUILD OK\x1b[0m\n';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK&from=buffer` });
|
||||
const body = res.json();
|
||||
expect(body.data.wait.matched).toBe(true);
|
||||
expect(body.data.wait.immediate).toBe(true);
|
||||
expect(body.data.wait.waitedMs).toBe(0);
|
||||
expect(sessionWaits.totalWaiterCount()).toBe(0);
|
||||
});
|
||||
|
||||
it('from=buffer only scans a bounded tail', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
// Old marker pushed past the scan window by newer output.
|
||||
ctx.sessions.get(SESSION_ID)!.terminalBuffer = `ANCIENT${'x'.repeat(MAX_BUFFER_SCAN_BYTES + 1000)}`;
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=ANCIENT&from=buffer&timeout=1` });
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json().data.wait.timedOut).toBe(true);
|
||||
});
|
||||
|
||||
it('defaults to from=now, ignoring what is already in the buffer', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.terminalBuffer = 'BUILD OK happened before you asked\n';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=BUILD%20OK&timeout=1` });
|
||||
expect(res.json().data.wait.timedOut).toBe(true);
|
||||
});
|
||||
|
||||
it('answers 200 with timedOut on timeout, never an error status', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=never&timeout=1` });
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
const body = res.json();
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.data.wait.timedOut).toBe(true);
|
||||
expect(body.data.wait.matched).toBe(false);
|
||||
expect(body.data.wait.snippet).toBeNull();
|
||||
});
|
||||
|
||||
it('resolves with ended when the session goes away mid-wait', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = app.inject({ method: 'GET', url: `${URL}?match=never` });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.data.wait.ended).toBe(true);
|
||||
expect(body.data.wait.matched).toBe(false);
|
||||
});
|
||||
|
||||
it('frees its waiter when the RESPONSE socket closes early', async () => {
|
||||
// The injected response is genuinely destroyed by then, so the freed slot is what
|
||||
// this can assert; the wire-level behaviour is pinned over real HTTP in
|
||||
// session-input-wait.test.ts.
|
||||
const { app, rawReplies } = await harness();
|
||||
const pending = app.inject({ method: 'GET', url: `${URL}?match=never&timeout=600000` }).then(
|
||||
() => 'completed',
|
||||
() => 'destroyed'
|
||||
);
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
rawReplies[rawReplies.length - 1].emit('close');
|
||||
await new Promise((resolve) => setTimeout(resolve, 10));
|
||||
|
||||
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(0);
|
||||
expect(await pending).toBe('destroyed');
|
||||
});
|
||||
|
||||
it('rejects a regex parameter instead of silently ignoring it', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=x®ex=%5EBUILD.*OK%24` });
|
||||
|
||||
expect(res.statusCode).toBe(400);
|
||||
const body = res.json();
|
||||
expect(body.errorCode).toBe('INVALID_INPUT');
|
||||
expect(body.error).toContain('match=');
|
||||
});
|
||||
|
||||
it('requires match', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: URL });
|
||||
|
||||
expect(res.statusCode).toBe(400);
|
||||
expect(res.json().errorCode).toBe('INVALID_INPUT');
|
||||
});
|
||||
|
||||
it('rejects an empty or oversized match, and says which parameter was wrong', async () => {
|
||||
const { app } = await harness();
|
||||
|
||||
const empty = await app.inject({ method: 'GET', url: `${URL}?match=` });
|
||||
expect(empty.statusCode).toBe(400);
|
||||
expect(empty.json().error).toContain('match');
|
||||
|
||||
const huge = await app.inject({ method: 'GET', url: `${URL}?match=${'x'.repeat(MAX_MATCH_LENGTH + 1)}` });
|
||||
expect(huge.statusCode).toBe(400);
|
||||
expect(huge.json().error).toContain('match');
|
||||
});
|
||||
|
||||
it('rejects a non-numeric timeout', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=x&timeout=soon` });
|
||||
|
||||
expect(res.statusCode).toBe(400);
|
||||
expect(res.json().errorCode).toBe('INVALID_INPUT');
|
||||
expect(res.json().error).toContain('timeout');
|
||||
});
|
||||
|
||||
it('rejects an unknown from value', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `${URL}?match=x&from=history` });
|
||||
|
||||
expect(res.statusCode).toBe(400);
|
||||
});
|
||||
|
||||
it('404s an unknown session', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: '/api/sessions/nope/wait-output?match=x' });
|
||||
|
||||
expect(res.statusCode).toBe(404);
|
||||
expect(res.json().success).toBe(false);
|
||||
});
|
||||
|
||||
it('output waiters share the session waiter cap with signal waiters', async () => {
|
||||
const { app } = await harness();
|
||||
const pendings = [];
|
||||
for (let i = 0; i < 8; i++) {
|
||||
pendings.push(app.inject({ method: 'GET', url: `${URL}?match=never${i}` }));
|
||||
}
|
||||
for (let i = 0; i < 8; i++) {
|
||||
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` }));
|
||||
}
|
||||
await new Promise((resolve) => setTimeout(resolve, 40));
|
||||
expect(sessionWaits.waiterCount(SESSION_ID)).toBe(16);
|
||||
|
||||
const overflow = await app.inject({ method: 'GET', url: `${URL}?match=one-too-many` });
|
||||
expect(overflow.statusCode).toBe(409);
|
||||
expect(overflow.json().errorCode).toBe('SESSION_BUSY');
|
||||
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
await Promise.all(pendings);
|
||||
});
|
||||
|
||||
it('checks the cap BEFORE materializing the terminal buffer', async () => {
|
||||
// `session.terminalBuffer` joins the whole 32MB accumulator. Paying that for a
|
||||
// request that is about to be refused turns the cap into an amplifier: a caller
|
||||
// already at the limit can loop `from=buffer` at full speed and never register a
|
||||
// waiter, so nothing bounds the work.
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
const bufferReads = vi.fn(() => 'nothing to see');
|
||||
Object.defineProperty(session, 'terminalBuffer', { get: bufferReads, configurable: true });
|
||||
|
||||
const pendings = [];
|
||||
for (let i = 0; i < 16; i++) {
|
||||
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` }));
|
||||
}
|
||||
await new Promise((resolve) => setTimeout(resolve, 40));
|
||||
expect(sessionWaits.waiterCount(SESSION_ID)).toBe(16);
|
||||
|
||||
const overflow = await app.inject({ method: 'GET', url: `${URL}?match=x&from=buffer` });
|
||||
expect(overflow.json().errorCode).toBe('SESSION_BUSY');
|
||||
expect(bufferReads).not.toHaveBeenCalled();
|
||||
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
await Promise.all(pendings);
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* The LIVE-STREAM half, wired the way production wires it.
|
||||
*
|
||||
* Everything above drives `sessionWaits.notifyOutput()` directly, which is the
|
||||
* registry's API, not the path a real session takes. Deleting the one line that
|
||||
* connects them — `sessionWaits.notifyOutput(session.id, data)` in the `terminal`
|
||||
* listener — left all four wait suites green and survived a full `test:ci` sweep, so
|
||||
* `wait-output?from=now` (the entire live mode, and the one the skill's recipes are
|
||||
* built on) could ship severed with nothing to show for it.
|
||||
*
|
||||
* These go through `createSessionListeners` so the wiring itself is what is pinned.
|
||||
*/
|
||||
describe('GET /api/sessions/:id/wait-output: fed by the real terminal listener', () => {
|
||||
function stubDeps() {
|
||||
return {
|
||||
broadcast: vi.fn(),
|
||||
batchTerminalData: vi.fn(),
|
||||
batchTaskUpdate: vi.fn(),
|
||||
broadcastSessionStateDebounced: vi.fn(),
|
||||
sendPushNotifications: vi.fn(),
|
||||
persistSessionState: vi.fn(),
|
||||
getSessionStateWithRespawn: vi.fn(() => ({})),
|
||||
getRunSummaryTracker: vi.fn(() => undefined),
|
||||
stopTranscriptWatcher: vi.fn(),
|
||||
cleanupSessionBatches: vi.fn(),
|
||||
cancelPersistDebounce: vi.fn(),
|
||||
removeRunSummaryTracker: vi.fn(),
|
||||
removeSessionListenerRefs: vi.fn(),
|
||||
cleanupRespawnOnExit: vi.fn(),
|
||||
getStore: vi.fn(() => ({ updateRalphState: vi.fn() })),
|
||||
registerAttachment: vi.fn(async () => {}),
|
||||
};
|
||||
}
|
||||
|
||||
it('matches output emitted by the session, not injected into the registry', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
const deps = stubDeps();
|
||||
attachSessionListeners(session as never, createSessionListeners(session as never, deps as never));
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `${URL}?match=LIVE_STREAM_HIT` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
|
||||
// What a PTY chunk actually does: the session emits `terminal`.
|
||||
session.simulateTerminalOutput('$ echo LIVE_STREAM_HIT\r\nLIVE_STREAM_HIT\r\n');
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.matched).toBe(true);
|
||||
expect(body.data.wait.snippet).toContain('LIVE_STREAM_HIT');
|
||||
// The listener must still forward to the SSE batcher; the wait feed is additive.
|
||||
expect(deps.batchTerminalData).toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it('matches across chunk boundaries through the listener', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `${URL}?match=SPLIT_MARKER` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
session.simulateTerminalOutput('noise SPLIT_');
|
||||
session.simulateTerminalOutput('MARKER more noise');
|
||||
|
||||
expect((await pending).json().data.wait.matched).toBe(true);
|
||||
});
|
||||
|
||||
it('strips ANSI on the way through the listener', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `${URL}?match=COLORED%20HIT` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
session.simulateAnsiOutput('COLORED HIT');
|
||||
|
||||
expect((await pending).json().data.wait.matched).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id/wait-output: a dead tmux worker', () => {
|
||||
it('releases an output waiter when the worker dies while it is parked', async () => {
|
||||
// The feed simply stops: no exit event, no further chunks, nothing to match.
|
||||
const { app, ctx } = await harness();
|
||||
let dead = false;
|
||||
(ctx.mux as unknown as { isPaneDead: () => boolean }).isPaneDead = () => dead;
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `${URL}?match=NEVER&timeout=600000` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 50));
|
||||
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
dead = true;
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.ended).toBe(true);
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(0);
|
||||
}, 10_000);
|
||||
|
||||
it('an oversized timeout clamps rather than 400ing', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
// Resolve from the buffer so the assertion is about the clamp, not a real wait.
|
||||
ctx.sessions.get(SESSION_ID)!.terminalBuffer = 'ALREADY_THERE\n';
|
||||
const res = await app.inject({
|
||||
method: 'GET',
|
||||
url: `${URL}?match=ALREADY_THERE&from=buffer&timeout=99999999`,
|
||||
});
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,811 @@
|
||||
/**
|
||||
* @fileoverview Route tests for `GET /api/sessions/:id/wait`.
|
||||
*
|
||||
* The contract this pins is the one an orchestrating agent depends on:
|
||||
* - a timeout is a 200 with `wait.timedOut: true`, never a 4xx, because callers loop
|
||||
* over short waits and every poll boundary would otherwise look like a failure;
|
||||
* - the result is nested under `data.wait` on ALL THREE wait endpoints, so a single
|
||||
* client helper reads any of them;
|
||||
* - the EFFECTIVE timeout is echoed, so a caller that asked for 30 minutes and was
|
||||
* clamped to 10 can tell a poll boundary from a wedged worker;
|
||||
* - an unknown `until` token is a 400 rather than a silent fallback to the default,
|
||||
* so a typo can never leave an agent believing it is waiting for something else,
|
||||
* and a schema 400 names the parameter it rejected;
|
||||
* - `stop`/`blocked` are rejected for modes that install no hooks when asked for
|
||||
* EXPLICITLY, but silently dropped from the DEFAULT set, so omitting `until` never
|
||||
* 400s — and `shell` counts as such a mode, even though it is not an external CLI;
|
||||
* - a client that hangs up frees its waiter immediately, or a loop of
|
||||
* `curl --max-time` calls wedges a process-wide cap nobody else can use;
|
||||
* - a session with no PTY answers `exit`, never `idle`.
|
||||
*
|
||||
* Plan: docs/agent-control-plan.md
|
||||
*/
|
||||
import { describe, it, expect, afterEach, vi } from 'vitest';
|
||||
import fastifyCookie from '@fastify/cookie';
|
||||
import Fastify, { type FastifyInstance } from 'fastify';
|
||||
import type { ServerResponse } from 'node:http';
|
||||
import {
|
||||
registerSessionRoutes,
|
||||
_resetPaneLivenessState,
|
||||
_paneDeathWatcherCount,
|
||||
} from '../../src/web/routes/session-routes.js';
|
||||
import { createSessionListeners, attachSessionListeners } from '../../src/web/session-listener-wiring.js';
|
||||
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
|
||||
import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js';
|
||||
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
|
||||
import { sessionWaits } from '../../src/web/session-wait-registry.js';
|
||||
import { MAX_WAIT_MS, MIN_WAIT_MS, MAX_WAITERS_TOTAL, MAX_WAITERS_PER_OWNER } from '../../src/config/agent-wait.js';
|
||||
|
||||
// Distinct per file on purpose: the three wait suites share the process-wide
|
||||
// `sessionWaits` singleton, so a common id let one file's leftover waiter be counted
|
||||
// by another's assertion. Failed only in a 5-file run, which is how CI runs them.
|
||||
const SESSION_ID = 'wait-routes-session';
|
||||
|
||||
/** Session ids a test parked filler waiters on, so cleanup can release them. */
|
||||
const fillerIds = new Set<string>();
|
||||
|
||||
/** Park a waiter directly on the shared registry (cap tests), tracked for cleanup. */
|
||||
function fillWaiter(id: string, owner?: string): Promise<unknown> {
|
||||
fillerIds.add(id);
|
||||
return sessionWaits.waitForSignal(id, { until: ['stop'], timeoutMs: 30_000, owner });
|
||||
}
|
||||
|
||||
afterEach(() => {
|
||||
// The routes use the process-wide registry; never leak a waiter into the next test.
|
||||
// Deliberately NOT `cancelEverything()`: it latches the registry's stopped flag
|
||||
// (one-way by design, so a request landing mid-shutdown cannot register a waiter
|
||||
// nothing will ever cancel), which would leave every later test in this file
|
||||
// talking to a dead registry and passing vacuously.
|
||||
for (const id of [SESSION_ID, ...fillerIds]) sessionWaits.cancelAll(id);
|
||||
fillerIds.clear();
|
||||
// Pane-liveness state is module-level (one cache, one watcher per pane), so it has
|
||||
// to be reset or a cached probe leaks into the next test.
|
||||
_resetPaneLivenessState();
|
||||
delete process.env.CODEMAN_MULTIUSER;
|
||||
});
|
||||
|
||||
/**
|
||||
* The shared route harness returns handler payloads verbatim, so a `{success:false}`
|
||||
* body would still be HTTP 200. Production maps errorCode to status in server.ts, so
|
||||
* mirror that here or every negative case passes vacuously.
|
||||
*
|
||||
* `rawReplies` collects each request's `reply.raw` — the RESPONSE, which is what the
|
||||
* handler's hang-up detection listens on, and deliberately not `req.raw` (on a POST
|
||||
* that one closes as soon as the body has been read, so wiring an abort to it kills
|
||||
* every send-and-wait; see the real-HTTP suite in session-input-wait.test.ts).
|
||||
* Emitting the event by hand is the only way to simulate a hang-up here at all:
|
||||
* `app.inject()` never emits `close` on its own, verified.
|
||||
*/
|
||||
async function harness(options?: { authUser?: { username: string; role: 'admin' | 'user' } }): Promise<{
|
||||
app: FastifyInstance;
|
||||
ctx: MockRouteContext;
|
||||
rawReplies: ServerResponse[];
|
||||
}> {
|
||||
const app = Fastify({ logger: false });
|
||||
await app.register(fastifyCookie);
|
||||
const ctx = createMockRouteContext({ sessionId: SESSION_ID });
|
||||
const rawReplies: ServerResponse[] = [];
|
||||
|
||||
const authUser = options?.authUser;
|
||||
app.addHook('onRequest', async (req, reply) => {
|
||||
// The RESPONSE, because that is what the handler's hang-up detection listens to.
|
||||
rawReplies.push(reply.raw);
|
||||
if (authUser) (req as unknown as { authUser: typeof authUser }).authUser = authUser;
|
||||
});
|
||||
|
||||
registerSessionRoutes(app, ctx as never);
|
||||
|
||||
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
|
||||
const p = payload as { success?: unknown; errorCode?: unknown } | null;
|
||||
if (p && typeof p === 'object' && p.success === false && reply.statusCode === 200) {
|
||||
if (typeof p.errorCode === 'string') reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
|
||||
}
|
||||
return done(null, payload);
|
||||
});
|
||||
|
||||
installRouteErrorHandler(app);
|
||||
await app.ready();
|
||||
return { app, ctx, rawReplies };
|
||||
}
|
||||
|
||||
describe('GET /api/sessions/:id/wait', () => {
|
||||
it('resolves immediately when the session is already in a requested state', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
const body = res.json();
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.data.wait.signal).toBe('idle');
|
||||
expect(body.data.wait.immediate).toBe(true);
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
expect(body.data.wait.aborted).toBe(false);
|
||||
expect(body.data.sessionId).toBe(SESSION_ID);
|
||||
expect(body.data.wait.until).toEqual(['idle']);
|
||||
});
|
||||
|
||||
it('nests the result under data.wait, so one client helper reads all three endpoints', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
|
||||
|
||||
const { data } = res.json();
|
||||
// Session-level facts stay at the top; everything about the wait is inside it.
|
||||
expect(Object.keys(data).sort()).toEqual(['limitPaused', 'sessionId', 'status', 'wait']);
|
||||
expect(data.signal).toBeUndefined();
|
||||
});
|
||||
|
||||
it('reports the post-wait status and the limit-pause hint', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
|
||||
|
||||
const body = res.json();
|
||||
expect(body.data.status).toBe('idle');
|
||||
expect(body.data.limitPaused).toBe(false);
|
||||
});
|
||||
|
||||
it('sends Cache-Control: no-store, so a polled long-poll cannot be served from a cache', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
|
||||
|
||||
expect(res.headers['cache-control']).toBe('no-store');
|
||||
});
|
||||
|
||||
it('defaults to stop,idle,exit when until is omitted', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
|
||||
|
||||
const body = res.json();
|
||||
expect(body.data.wait.until).toEqual(['stop', 'idle', 'exit']);
|
||||
// The mock session is idle, so the default set resolves right away.
|
||||
expect(body.data.wait.signal).toBe('idle');
|
||||
});
|
||||
|
||||
it('rejects an unknown until token instead of falling back to the default', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stpo` });
|
||||
|
||||
expect(res.statusCode).toBe(400);
|
||||
const body = res.json();
|
||||
expect(body.success).toBe(false);
|
||||
expect(body.errorCode).toBe('INVALID_INPUT');
|
||||
expect(body.error).toContain('stpo');
|
||||
});
|
||||
|
||||
it('rejects a non-numeric timeout', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?timeout=soon` });
|
||||
|
||||
expect(res.statusCode).toBe(400);
|
||||
expect(res.json().errorCode).toBe('INVALID_INPUT');
|
||||
});
|
||||
|
||||
it('names the parameter it rejected, instead of a bare "invalid parameters"', async () => {
|
||||
// An agent driving this with no docs in context can only recover if the error
|
||||
// says WHICH parameter was wrong; the old message named none of them.
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?timeout=30s` });
|
||||
|
||||
const body = res.json();
|
||||
expect(body.error).toContain('timeout');
|
||||
expect(body.error).toContain('wait');
|
||||
});
|
||||
|
||||
it('404s an unknown session', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: '/api/sessions/nope/wait?until=idle' });
|
||||
|
||||
expect(res.statusCode).toBe(404);
|
||||
expect(res.json().success).toBe(false);
|
||||
});
|
||||
|
||||
it('resolves an in-flight wait when the signal arrives', async () => {
|
||||
const { app } = await harness();
|
||||
// `stop` is not the session's current signal, so this blocks.
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
sessionWaits.notifySignal(SESSION_ID, 'stop');
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.signal).toBe('stop');
|
||||
expect(body.data.wait.immediate).toBe(false);
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
expect(body.data.wait.waitedMs).toBeGreaterThanOrEqual(0);
|
||||
});
|
||||
|
||||
it('fresh=1 waits for the next transition instead of answering from current state', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&fresh=1` });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
sessionWaits.notifySignal(SESSION_ID, 'idle');
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.signal).toBe('idle');
|
||||
expect(body.data.wait.immediate).toBe(false);
|
||||
});
|
||||
|
||||
it('resolves with ended when the session goes away mid-wait', async () => {
|
||||
const { app } = await harness();
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.data.wait.ended).toBe(true);
|
||||
expect(body.data.wait.signal).toBeNull();
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
});
|
||||
|
||||
it('answers 200 with timedOut on timeout, never an error status', async () => {
|
||||
const { app } = await harness();
|
||||
// Clamped up to the 1s floor, so this is the one deliberately slow case.
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=1` });
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
const body = res.json();
|
||||
expect(body.success).toBe(true);
|
||||
expect(body.data.wait.timedOut).toBe(true);
|
||||
expect(body.data.wait.signal).toBeNull();
|
||||
expect(body.data.wait.ended).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id/wait: the effective timeout is observable', () => {
|
||||
it('echoes the clamped-down value when the caller asks for more than the ceiling', async () => {
|
||||
// Asked for 30 minutes, got MAX_WAIT_MS. Without the echo the caller reads a
|
||||
// 10-minute timeout as "30 minutes elapsed with no stop" and kills a healthy worker.
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=1800000` });
|
||||
|
||||
expect(res.json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
|
||||
});
|
||||
|
||||
it('echoes the clamped-up value at the floor too', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=1` });
|
||||
|
||||
expect(res.json().data.wait.timeoutMs).toBe(MIN_WAIT_MS);
|
||||
});
|
||||
|
||||
it('reports the applied default when timeout is omitted', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
|
||||
|
||||
expect(res.json().data.wait.timeoutMs).toBeGreaterThanOrEqual(MIN_WAIT_MS);
|
||||
expect(res.json().data.wait.timeoutMs).toBeLessThanOrEqual(MAX_WAIT_MS);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id/wait: repeated query parameters', () => {
|
||||
it('accepts ?until=stop&until=exit, the way most clients express a list', async () => {
|
||||
// Fastify delivers a repeated parameter as an array and parseWaitSignals has
|
||||
// always handled one; only the schema was rejecting it.
|
||||
const { app } = await harness();
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&until=exit` });
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
sessionWaits.notifySignal(SESSION_ID, 'exit');
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.until).toEqual(['stop', 'exit']);
|
||||
expect(body.data.wait.signal).toBe('exit');
|
||||
});
|
||||
|
||||
it('still reports an unknown token inside a repeated parameter', async () => {
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&until=stpo` });
|
||||
|
||||
expect(res.statusCode).toBe(400);
|
||||
expect(res.json().error).toContain('stpo');
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id/wait: a client that hangs up frees its waiter', () => {
|
||||
it('removes the waiter when the RESPONSE socket closes early', async () => {
|
||||
// `curl --max-time 30 ".../wait?timeout=600000"` abandons a live waiter every
|
||||
// iteration of the documented loop; sixteen of those and an innocent session
|
||||
// reports busy.
|
||||
//
|
||||
// The response body is unreadable afterwards (the injected response really is
|
||||
// destroyed, exactly as a hung-up socket would be), so the freed slot is all this
|
||||
// can assert. The full behaviour, including `aborted: true` on the wire for the
|
||||
// caller that did NOT hang up, is pinned over real HTTP in session-input-wait.test.ts.
|
||||
const { app, rawReplies } = await harness();
|
||||
const pending = app
|
||||
.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` })
|
||||
.then(
|
||||
() => 'completed',
|
||||
() => 'destroyed'
|
||||
);
|
||||
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
rawReplies[rawReplies.length - 1].emit('close');
|
||||
await new Promise((resolve) => setTimeout(resolve, 10));
|
||||
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
|
||||
expect(await pending).toBe('destroyed');
|
||||
});
|
||||
|
||||
it('a close AFTER the wait resolved changes nothing', async () => {
|
||||
const { app, rawReplies } = await harness();
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
|
||||
|
||||
expect(res.json().data.wait.aborted).toBe(false);
|
||||
// Node fires `close` on every completed response too, not only on a hang-up;
|
||||
// `writableFinished` is what separates them.
|
||||
rawReplies[rawReplies.length - 1].emit('close');
|
||||
expect(sessionWaits.totalWaiterCount()).toBe(0);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id/wait: capacity errors name the cap that was hit', () => {
|
||||
it('maps the per-session cap to SESSION_BUSY / 409', async () => {
|
||||
const { app } = await harness();
|
||||
const pendings = [];
|
||||
for (let i = 0; i < 16; i++) {
|
||||
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` }));
|
||||
}
|
||||
await new Promise((resolve) => setTimeout(resolve, 30));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(16);
|
||||
|
||||
const overflow = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
|
||||
expect(overflow.statusCode).toBe(409);
|
||||
expect(overflow.json().errorCode).toBe('SESSION_BUSY');
|
||||
expect(overflow.json().error).toContain('session');
|
||||
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
await Promise.all(pendings);
|
||||
});
|
||||
|
||||
it('maps the process-wide cap to RATE_LIMITED / 429, because this session is not the problem', async () => {
|
||||
// Reported as SESSION_BUSY, an agent concludes the session it asked about is
|
||||
// busy, switches to another, and gets the identical error.
|
||||
const { app } = await harness();
|
||||
const others: Promise<unknown>[] = [];
|
||||
for (let i = 0; i < MAX_WAITERS_TOTAL; i++) others.push(fillWaiter(`unrelated-${i}`));
|
||||
expect(sessionWaits.totalWaiterCount()).toBe(MAX_WAITERS_TOTAL);
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
|
||||
expect(res.statusCode).toBe(429);
|
||||
expect(res.json().errorCode).toBe('RATE_LIMITED');
|
||||
expect(res.json().error).toContain('total');
|
||||
|
||||
for (const id of fillerIds) sessionWaits.cancelAll(id);
|
||||
await Promise.all(others);
|
||||
});
|
||||
|
||||
it('maps the per-owner cap to RATE_LIMITED / 429 and charges the request to its user', async () => {
|
||||
// Also proves the route passes ownerFor(req): without it the owner cap can never
|
||||
// trip, and one user could hold the whole process-wide pool.
|
||||
process.env.CODEMAN_MULTIUSER = '1';
|
||||
const { app } = await harness({ authUser: { username: 'alice', role: 'admin' } });
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.ownerWaiterCount('alice')).toBe(1);
|
||||
|
||||
const others: Promise<unknown>[] = [];
|
||||
while (sessionWaits.ownerWaiterCount('alice') < MAX_WAITERS_PER_OWNER) {
|
||||
others.push(fillWaiter(`alice-${others.length}`, 'alice'));
|
||||
}
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
|
||||
expect(res.statusCode).toBe(429);
|
||||
expect(res.json().error).toContain('owner');
|
||||
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
for (const id of fillerIds) sessionWaits.cancelAll(id);
|
||||
await Promise.all([pending, ...others]);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id/wait: modes that install no hooks', () => {
|
||||
it('rejects an explicit stop, which no external CLI ever emits', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.mode = 'codex';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
|
||||
expect(res.statusCode).toBe(400);
|
||||
expect(res.json().error).toContain('codex');
|
||||
});
|
||||
|
||||
it('rejects an explicit blocked too', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.mode = 'opencode';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=blocked` });
|
||||
expect(res.statusCode).toBe(400);
|
||||
});
|
||||
|
||||
it('rejects stop for a SHELL session, which is not an external CLI but installs no hooks either', async () => {
|
||||
// A plain bash PTY never POSTs a Stop hook, so this was a guaranteed ten-minute
|
||||
// hang dressed up as a timeout — the exact failure the guard exists to prevent.
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.mode = 'shell';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop` });
|
||||
expect(res.statusCode).toBe(400);
|
||||
expect(res.json().error).toContain('shell');
|
||||
});
|
||||
|
||||
it('silently drops hook-only signals from the DEFAULT set instead of 400ing', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.mode = 'gemini';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json().data.wait.until).toEqual(['idle', 'exit']);
|
||||
});
|
||||
|
||||
it('drops them from the default set for shell too', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.mode = 'shell';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json().data.wait.until).toEqual(['idle', 'exit']);
|
||||
});
|
||||
|
||||
it('still accepts idle and exit explicitly', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.mode = 'antigravity';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle,exit` });
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json().data.wait.until).toEqual(['idle', 'exit']);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id/wait: liveness beats the reported status', () => {
|
||||
it('answers exit for a session whose PTY is gone, not the idle its status claims', async () => {
|
||||
// Session parks a dead PTY at status 'idle' and the object survives in the map,
|
||||
// so the DEFAULT wait used to answer {signal:"idle", immediate:true} for a
|
||||
// crashed worker — 200, success, no error, and the agent prompts a corpse.
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
session.pid = null;
|
||||
session.status = 'idle';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
|
||||
const body = res.json();
|
||||
expect(body.data.wait.signal).toBe('exit');
|
||||
expect(body.data.wait.immediate).toBe(true);
|
||||
// The raw status is still reported, so nothing is hidden from the caller.
|
||||
expect(body.data.status).toBe('idle');
|
||||
});
|
||||
|
||||
it('resolves until=exit immediately for an already-exited session instead of blocking', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.pid = null;
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit&timeout=1` });
|
||||
expect(res.json().data.wait.signal).toBe('exit');
|
||||
expect(res.json().data.wait.timedOut).toBe(false);
|
||||
});
|
||||
|
||||
it('does not report idle for a dead session even when idle was asked for explicitly', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.pid = null;
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=1` });
|
||||
expect(res.json().data.wait.signal).toBeNull();
|
||||
expect(res.json().data.wait.timedOut).toBe(true);
|
||||
});
|
||||
|
||||
it('a live busy session resolves until=working immediately', async () => {
|
||||
// Unreachable before: MockSession used 'working', which is not a SessionStatus,
|
||||
// so signalForStatus fell through to null and this branch had no coverage.
|
||||
const { app, ctx } = await harness();
|
||||
ctx.sessions.get(SESSION_ID)!.status = 'busy';
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=working` });
|
||||
expect(res.json().data.wait.signal).toBe('working');
|
||||
expect(res.json().data.wait.immediate).toBe(true);
|
||||
});
|
||||
|
||||
it('a live stopped/error session maps to exit', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
|
||||
session.status = 'stopped';
|
||||
const stopped = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit` });
|
||||
expect(stopped.json().data.wait.signal).toBe('exit');
|
||||
|
||||
session.status = 'error';
|
||||
const errored = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit` });
|
||||
expect(errored.json().data.wait.signal).toBe('exit');
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* The other half of the wait contract, exercised through the real listener wiring
|
||||
* rather than a route: a PTY that dies must RELEASE every waiter, not only the ones
|
||||
* that asked for `exit`.
|
||||
*
|
||||
* It lives in this file because it pins the same promise the routes above make
|
||||
* ("never hang"), and because the failure is only visible from the caller's side:
|
||||
* the exit handler detaches the `terminal`, `idle` and `working` listeners moments
|
||||
* later, so anything still registered afterwards is waiting on feeds that no longer
|
||||
* exist and can only time out.
|
||||
*/
|
||||
describe('a PTY exit releases waiters that did not ask for exit', () => {
|
||||
/** Everything the exit handler touches; the wait release must not depend on any of it. */
|
||||
function stubDeps(overrides: Record<string, unknown> = {}) {
|
||||
return {
|
||||
broadcast: vi.fn(),
|
||||
batchTerminalData: vi.fn(),
|
||||
batchTaskUpdate: vi.fn(),
|
||||
broadcastSessionStateDebounced: vi.fn(),
|
||||
sendPushNotifications: vi.fn(),
|
||||
persistSessionState: vi.fn(),
|
||||
getSessionStateWithRespawn: vi.fn(() => ({})),
|
||||
getRunSummaryTracker: vi.fn(() => undefined),
|
||||
stopTranscriptWatcher: vi.fn(),
|
||||
cleanupSessionBatches: vi.fn(),
|
||||
cancelPersistDebounce: vi.fn(),
|
||||
removeRunSummaryTracker: vi.fn(),
|
||||
removeSessionListenerRefs: vi.fn(),
|
||||
cleanupRespawnOnExit: vi.fn(),
|
||||
getStore: vi.fn(() => ({ updateRalphState: vi.fn() })),
|
||||
registerAttachment: vi.fn(async () => {}),
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
it('answers an until=working waiter with ended instead of leaving it to time out', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=working&timeout=600000` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
session.emit('exit', 1);
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.ended).toBe(true);
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
expect(body.data.wait.signal).toBeNull();
|
||||
});
|
||||
|
||||
it('still gives an until=exit waiter its signal, not a bare ended', async () => {
|
||||
// Ordering matters: notifySignal('exit') must run BEFORE cancelAll, or a caller
|
||||
// that asked the right question gets the generic answer.
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit&fresh=1` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
|
||||
session.emit('exit', 0);
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.signal).toBe('exit');
|
||||
expect(body.data.wait.ended).toBe(false);
|
||||
});
|
||||
|
||||
it('releases output waiters too, whose only feed the exit handler is about to detach', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
attachSessionListeners(session as never, createSessionListeners(session as never, stubDeps() as never));
|
||||
|
||||
const pending = app.inject({
|
||||
method: 'GET',
|
||||
url: `/api/sessions/${SESSION_ID}/wait-output?match=DONE&timeout=600000`,
|
||||
});
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
expect(sessionWaits.outputWaiterCount(SESSION_ID)).toBe(1);
|
||||
|
||||
session.emit('exit', 1);
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.ended).toBe(true);
|
||||
expect(body.data.wait.matched).toBe(false);
|
||||
expect(sessionWaits.waiterCount(SESSION_ID)).toBe(0);
|
||||
});
|
||||
|
||||
it('releases them even when a later step of the exit handler throws', async () => {
|
||||
// Which is why the release is the first thing in the handler.
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
const deps = stubDeps({
|
||||
broadcast: vi.fn(() => {
|
||||
throw new Error('SSE is down');
|
||||
}),
|
||||
});
|
||||
attachSessionListeners(session as never, createSessionListeners(session as never, deps as never));
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 20));
|
||||
|
||||
session.emit('exit', 1);
|
||||
|
||||
expect((await pending).json().data.wait.ended).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* Worker liveness for a tmux-backed session.
|
||||
*
|
||||
* `session.pid` is the local `tmux attach` CLIENT, not the worker. Codeman sets
|
||||
* `remain-on-exit on`, so when the command inside the pane exits tmux keeps the pane
|
||||
* (`pane_dead=1`), the tmux session survives, the attach client keeps running, `pid`
|
||||
* never goes null and NO exit event fires. Reproduced live on a shell worker killed
|
||||
* with `exit 42`: tmux said `pane_dead=1 status=42` while Codeman said
|
||||
* `pid=309406 status=idle` and the default wait answered
|
||||
* `{signal:"idle", immediate:true, waitedMs:0}` for a corpse.
|
||||
*
|
||||
* These cases could not exist before, because `MockSession.pid` is set by hand: the
|
||||
* `pid === null` branch is the one production never reaches.
|
||||
*/
|
||||
describe('GET /api/sessions/:id/wait: a dead tmux worker', () => {
|
||||
/** Mock ctx doubles carry no `isPaneDead`; the route treats that as "cannot tell". */
|
||||
function setPaneDead(ctx: MockRouteContext, dead: boolean): ReturnType<typeof vi.fn> {
|
||||
const probe = vi.fn(() => dead);
|
||||
(ctx.mux as unknown as { isPaneDead: (name: string) => boolean }).isPaneDead = probe as never;
|
||||
return probe;
|
||||
}
|
||||
|
||||
it('answers exit, not the idle the session still reports', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const session = ctx.sessions.get(SESSION_ID)!;
|
||||
setPaneDead(ctx, true);
|
||||
// Exactly the live state: attach client alive, status idle, worker gone.
|
||||
expect(session.pid).not.toBeNull();
|
||||
expect(session.status).toBe('idle');
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait` });
|
||||
const body = res.json();
|
||||
expect(body.data.wait.signal).toBe('exit');
|
||||
expect(body.data.wait.immediate).toBe(true);
|
||||
// The raw status is still reported, so nothing is hidden from the caller.
|
||||
expect(body.data.status).toBe('idle');
|
||||
});
|
||||
|
||||
it('resolves until=exit immediately instead of burning the whole timeout', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
setPaneDead(ctx, true);
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=exit&timeout=1` });
|
||||
expect(res.json().data.wait.signal).toBe('exit');
|
||||
expect(res.json().data.wait.timedOut).toBe(false);
|
||||
});
|
||||
|
||||
it('does not answer idle for a dead worker even when idle was asked for explicitly', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
setPaneDead(ctx, true);
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=1` });
|
||||
expect(res.json().data.wait.signal).toBeNull();
|
||||
expect(res.json().data.wait.timedOut).toBe(true);
|
||||
});
|
||||
|
||||
it('a live pane is unaffected', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
setPaneDead(ctx, false);
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
|
||||
expect(res.json().data.wait.signal).toBe('idle');
|
||||
});
|
||||
|
||||
it('caches the probe, so a poll loop cannot exec tmux once per request', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const probe = setPaneDead(ctx, false);
|
||||
|
||||
for (let i = 0; i < 10; i++) {
|
||||
await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
|
||||
}
|
||||
expect(probe.mock.calls.length).toBeLessThanOrEqual(2);
|
||||
});
|
||||
|
||||
it('never probes a session that is not tmux-backed', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
const probe = setPaneDead(ctx, true);
|
||||
ctx.sessions.get(SESSION_ID)!.usesMux = false;
|
||||
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=idle` });
|
||||
expect(probe).not.toHaveBeenCalled();
|
||||
// Falls back to the pid rule, which is the right one for a direct PTY.
|
||||
expect(res.json().data.wait.signal).toBe('idle');
|
||||
});
|
||||
|
||||
it('releases a wait when the worker dies WHILE it is parked', async () => {
|
||||
// The common orchestration case, and the one a request-time probe cannot see: no
|
||||
// exit event, no output, nothing — the caller would block for its full timeout.
|
||||
const { app, ctx } = await harness();
|
||||
let dead = false;
|
||||
(ctx.mux as unknown as { isPaneDead: () => boolean }).isPaneDead = () => dead;
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 50));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(1);
|
||||
expect(_paneDeathWatcherCount()).toBe(1);
|
||||
|
||||
dead = true;
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.ended).toBe(true);
|
||||
expect(body.data.wait.timedOut).toBe(false);
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(0);
|
||||
// ...and the watcher is torn down with the last waiter that needed it.
|
||||
expect(_paneDeathWatcherCount()).toBe(0);
|
||||
}, 10_000);
|
||||
|
||||
it('an until=exit caller parked when the worker dies gets its signal, not a bare ended', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
let dead = false;
|
||||
(ctx.mux as unknown as { isPaneDead: () => boolean }).isPaneDead = () => dead;
|
||||
|
||||
const pending = app.inject({
|
||||
method: 'GET',
|
||||
url: `/api/sessions/${SESSION_ID}/wait?until=exit&fresh=1&timeout=600000`,
|
||||
});
|
||||
await new Promise((resolve) => setTimeout(resolve, 50));
|
||||
dead = true;
|
||||
|
||||
const body = (await pending).json();
|
||||
expect(body.data.wait.signal).toBe('exit');
|
||||
expect(body.data.wait.ended).toBe(false);
|
||||
}, 10_000);
|
||||
|
||||
it('starts no watcher at all when the session is not tmux-backed', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
setPaneDead(ctx, false);
|
||||
ctx.sessions.get(SESSION_ID)!.usesMux = false;
|
||||
|
||||
const pending = app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` });
|
||||
await new Promise((resolve) => setTimeout(resolve, 30));
|
||||
expect(_paneDeathWatcherCount()).toBe(0);
|
||||
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
await pending;
|
||||
});
|
||||
|
||||
it('shares ONE watcher across every wait parked on the same session', async () => {
|
||||
const { app, ctx } = await harness();
|
||||
setPaneDead(ctx, false);
|
||||
|
||||
const pendings = [];
|
||||
for (let i = 0; i < 5; i++) {
|
||||
pendings.push(app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?until=stop&timeout=600000` }));
|
||||
}
|
||||
await new Promise((resolve) => setTimeout(resolve, 40));
|
||||
expect(sessionWaits.signalWaiterCount(SESSION_ID)).toBe(5);
|
||||
expect(_paneDeathWatcherCount()).toBe(1);
|
||||
|
||||
sessionWaits.cancelAll(SESSION_ID);
|
||||
await Promise.all(pendings);
|
||||
expect(_paneDeathWatcherCount()).toBe(0);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id/wait: an oversized timeout clamps, it does not 400', () => {
|
||||
it('accepts a value above the old schema ceiling and reports the clamp', async () => {
|
||||
// "Clamped to [1000, 600000]" has to mean it: `timeout=600001` clamping while
|
||||
// `timeout=99999999` 400s is the same documented rule producing two outcomes.
|
||||
const { app } = await harness();
|
||||
const res = await app.inject({
|
||||
method: 'GET',
|
||||
url: `/api/sessions/${SESSION_ID}/wait?until=idle&timeout=99999999`,
|
||||
});
|
||||
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(res.json().data.wait.timeoutMs).toBe(MAX_WAIT_MS);
|
||||
});
|
||||
|
||||
it('still rejects a non-finite or non-integer timeout', async () => {
|
||||
const { app } = await harness();
|
||||
for (const value of ['1e999', 'soon', '-1', '1.5']) {
|
||||
const res = await app.inject({ method: 'GET', url: `/api/sessions/${SESSION_ID}/wait?timeout=${value}` });
|
||||
expect(res.statusCode, `timeout=${value}`).toBe(400);
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -208,6 +208,55 @@ describe('ws-routes', () => {
|
||||
}
|
||||
});
|
||||
|
||||
it('ACKs a delivered input and burns its seq', async () => {
|
||||
const ws = await connectWs('/ws/sessions/ws-test-session/terminal');
|
||||
try {
|
||||
const session = ctx._session;
|
||||
ws.send(JSON.stringify({ t: 'i', d: 'ok\r', cid: 'c1', seq: 1 }));
|
||||
|
||||
expect(await nextMessage(ws)).toEqual({ t: 'ia', seq: 1 });
|
||||
expect(session.shouldApplyInput('c1', 1)).toBe(false);
|
||||
} finally {
|
||||
ws.close();
|
||||
}
|
||||
});
|
||||
|
||||
it('withholds the ACK and re-opens the seq when the write did not land', async () => {
|
||||
// A session whose PTY is gone swallows the write. ACKing anyway told the
|
||||
// client to drop the frame from its durable queue while the seq stayed
|
||||
// burnt, so the retry that reliable delivery exists for was rejected as a
|
||||
// duplicate — the input was lost for good.
|
||||
const ws = await connectWs('/ws/sessions/ws-test-session/terminal');
|
||||
try {
|
||||
const session = ctx._session;
|
||||
session.failWrites = true;
|
||||
|
||||
ws.send(JSON.stringify({ t: 'i', d: 'lost\r', cid: 'c1', seq: 1 }));
|
||||
|
||||
await expect(nextMessage(ws, 600)).rejects.toThrow(/timeout/);
|
||||
expect(session.shouldApplyInput('c1', 1)).toBe(true);
|
||||
} finally {
|
||||
ws.close();
|
||||
}
|
||||
});
|
||||
|
||||
it('still ACKs a duplicate frame the server deliberately skipped', async () => {
|
||||
// Dedup must stay silent-but-acknowledged: the client has to be able to
|
||||
// drop a frame it already delivered once.
|
||||
const ws = await connectWs('/ws/sessions/ws-test-session/terminal');
|
||||
try {
|
||||
const session = ctx._session;
|
||||
session.shouldApplyInput('c1', 7); // pretend seq 7 already landed
|
||||
|
||||
ws.send(JSON.stringify({ t: 'i', d: 'again\r', cid: 'c1', seq: 7 }));
|
||||
|
||||
expect(await nextMessage(ws)).toEqual({ t: 'ia', seq: 7 });
|
||||
expect(session.writeBuffer).not.toContain('again\r');
|
||||
} finally {
|
||||
ws.close();
|
||||
}
|
||||
});
|
||||
|
||||
it('ignores input exceeding MAX_INPUT_LENGTH', async () => {
|
||||
const ws = await connectWs('/ws/sessions/ws-test-session/terminal');
|
||||
try {
|
||||
|
||||
@@ -351,7 +351,13 @@ describe('Codex quick start settings', () => {
|
||||
function loadUi(flags: Record<string, boolean> | undefined) {
|
||||
const CodemanApp = function CodemanApp(this: any) {};
|
||||
const welcomeBtns: Record<string, { style: { display: string } }> = {};
|
||||
for (const id of ['welcomeClaudeBtn', 'welcomeOpencodeBtn', 'welcomeGeminiBtn', 'welcomeTunnelBtn']) {
|
||||
for (const id of [
|
||||
'welcomeClaudeBtn',
|
||||
'welcomeOpencodeBtn',
|
||||
'welcomeAntigravityBtn',
|
||||
'welcomeGeminiBtn',
|
||||
'welcomeTunnelBtn',
|
||||
]) {
|
||||
welcomeBtns[id] = { style: { display: 'PRISTINE' } };
|
||||
}
|
||||
const modeBtns: Record<string, { style: { display: string } }> = {};
|
||||
@@ -394,6 +400,7 @@ describe('Codex quick start settings', () => {
|
||||
app.applyWelcomeCliVisibility();
|
||||
expect(welcomeBtns.welcomeClaudeBtn.style.display).toBe('flex');
|
||||
expect(welcomeBtns.welcomeOpencodeBtn.style.display).toBe('none');
|
||||
expect(welcomeBtns.welcomeAntigravityBtn.style.display).toBe('none');
|
||||
expect(welcomeBtns.welcomeGeminiBtn.style.display).toBe('none');
|
||||
// #200 originally DELETED the tunnel button and its QR outright; it is gated
|
||||
// on cloudflared instead, so a box that has cloudflared keeps the feature.
|
||||
@@ -402,6 +409,12 @@ describe('Codex quick start settings', () => {
|
||||
const withTunnel = loadUi({ ...ALL_OFF, cloudflared: true });
|
||||
withTunnel.app.applyWelcomeCliVisibility();
|
||||
expect(withTunnel.welcomeBtns.welcomeTunnelBtn.style.display).toBe('flex');
|
||||
|
||||
// Antigravity is a first-class welcome action, gated on `agy` like the rest.
|
||||
const withAgy = loadUi({ ...ALL_OFF, antigravity: true });
|
||||
withAgy.app.applyWelcomeCliVisibility();
|
||||
expect(withAgy.welcomeBtns.welcomeAntigravityBtn.style.display).toBe('flex');
|
||||
expect(withAgy.welcomeBtns.welcomeClaudeBtn.style.display).toBe('none');
|
||||
});
|
||||
|
||||
it('gates every run mode in the dropdown, antigravity included, and never shell', () => {
|
||||
|
||||
@@ -0,0 +1,179 @@
|
||||
/**
|
||||
* Unit tests for the unit-file builders behind `codeman service install`
|
||||
* (issue #231). These are the parts that must be right without launchctl or
|
||||
* systemctl in the loop: PATH construction, escaping, and the file contents.
|
||||
*/
|
||||
|
||||
import { describe, it, expect } from 'vitest';
|
||||
import {
|
||||
buildLaunchAgentPlist,
|
||||
buildServiceEnv,
|
||||
buildServicePath,
|
||||
buildSystemdUnit,
|
||||
detectServiceKind,
|
||||
systemdQuote,
|
||||
xmlEscape,
|
||||
type ServicePlan,
|
||||
} from '../src/service-installer.js';
|
||||
|
||||
function plan(overrides: Partial<ServicePlan> = {}): ServicePlan {
|
||||
return {
|
||||
kind: 'systemd',
|
||||
name: 'codeman-web.service',
|
||||
nodePath: '/usr/bin/node',
|
||||
execArgv: [],
|
||||
scriptPath: '/home/u/.codeman/app/dist/index.js',
|
||||
args: ['web', '--host', '127.0.0.1', '--port', '3000'],
|
||||
env: { PATH: '/usr/bin:/bin', HOME: '/home/u', LANG: 'en_US.UTF-8' },
|
||||
logPath: '/home/u/.codeman/web.log',
|
||||
workingDir: '/home/u',
|
||||
...overrides,
|
||||
};
|
||||
}
|
||||
|
||||
describe('buildServicePath', () => {
|
||||
it("puts the running node's directory first so nvm/homebrew node wins", () => {
|
||||
const result = buildServicePath('/home/u/.nvm/versions/node/v22.0.0/bin', '/usr/bin:/bin', '/home/u');
|
||||
expect(result.split(':')[0]).toBe('/home/u/.nvm/versions/node/v22.0.0/bin');
|
||||
});
|
||||
|
||||
it('keeps the installing shell PATH, which is the whole point of the fix', () => {
|
||||
const result = buildServicePath('/usr/bin', '/opt/homebrew/bin:/home/u/.bun/bin', '/home/u');
|
||||
expect(result.split(':')).toContain('/home/u/.bun/bin');
|
||||
expect(result.split(':')).toContain('/opt/homebrew/bin');
|
||||
});
|
||||
|
||||
it('appends the fallbacks a bare launchd PATH would otherwise be missing', () => {
|
||||
const entries = buildServicePath('/usr/bin', '/usr/bin', '/home/u').split(':');
|
||||
expect(entries).toContain('/opt/homebrew/bin');
|
||||
expect(entries).toContain('/home/u/.local/bin');
|
||||
expect(entries).toContain('/usr/local/bin');
|
||||
});
|
||||
|
||||
it('never repeats a directory', () => {
|
||||
const entries = buildServicePath('/usr/bin', '/usr/bin:/bin:/usr/bin', '/home/u').split(':');
|
||||
expect(new Set(entries).size).toBe(entries.length);
|
||||
});
|
||||
|
||||
it('drops empty segments from a trailing-colon PATH', () => {
|
||||
expect(buildServicePath('/usr/bin', '/usr/bin::/bin:', '/home/u').split(':')).not.toContain('');
|
||||
});
|
||||
|
||||
it('drops node_modules/.bin, which npx injects for one command only', () => {
|
||||
const entries = buildServicePath(
|
||||
'/usr/bin',
|
||||
'/repo/node_modules/.bin:/repo/node_modules/.bin/:/home/u/bin',
|
||||
'/home/u'
|
||||
).split(':');
|
||||
expect(entries.filter((e) => e.includes('node_modules'))).toEqual([]);
|
||||
expect(entries).toContain('/home/u/bin');
|
||||
});
|
||||
});
|
||||
|
||||
describe('buildServiceEnv', () => {
|
||||
it('carries PATH, HOME and a LANG default', () => {
|
||||
const env = buildServiceEnv('/usr/bin', '/usr/bin:/bin', '/home/u');
|
||||
expect(env.HOME).toBe('/home/u');
|
||||
expect(env.LANG).toBe('en_US.UTF-8');
|
||||
expect(env.PATH).toContain('/usr/bin');
|
||||
});
|
||||
|
||||
it('prefers the caller LANG when there is one', () => {
|
||||
expect(buildServiceEnv('/usr/bin', '/usr/bin', '/home/u', 'de_DE.UTF-8').LANG).toBe('de_DE.UTF-8');
|
||||
});
|
||||
|
||||
it('does not carry a password into the unit file', () => {
|
||||
const env = buildServiceEnv('/usr/bin', '/usr/bin', '/home/u');
|
||||
expect(Object.keys(env)).not.toContain('CODEMAN_PASSWORD');
|
||||
});
|
||||
});
|
||||
|
||||
describe('escaping', () => {
|
||||
it('escapes the five XML entities', () => {
|
||||
expect(xmlEscape(`a&b<c>d"e'f`)).toBe('a&b<c>d"e'f');
|
||||
});
|
||||
|
||||
it('quotes systemd values and escapes quotes and backslashes', () => {
|
||||
expect(systemdQuote('plain')).toBe('"plain"');
|
||||
expect(systemdQuote('with "quotes"')).toBe('"with \\"quotes\\""');
|
||||
expect(systemdQuote('back\\slash')).toBe('"back\\\\slash"');
|
||||
});
|
||||
});
|
||||
|
||||
describe('buildLaunchAgentPlist', () => {
|
||||
it('writes the label, the full command and the log paths', () => {
|
||||
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', name: 'com.codeman.web' }));
|
||||
expect(xml).toContain('<string>com.codeman.web</string>');
|
||||
expect(xml).toContain('<string>/usr/bin/node</string>');
|
||||
expect(xml).toContain('<string>/home/u/.codeman/app/dist/index.js</string>');
|
||||
expect(xml).toContain('<string>web</string>');
|
||||
expect(xml).toContain('<string>/home/u/.codeman/web.log</string>');
|
||||
});
|
||||
|
||||
it('keeps the argument order: node, script, then the web args', () => {
|
||||
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', name: 'com.codeman.web' }));
|
||||
// Match whole <string> elements: the label itself contains the word "web".
|
||||
const order = [
|
||||
'<string>/usr/bin/node</string>',
|
||||
'<string>/home/u/.codeman/app/dist/index.js</string>',
|
||||
'<string>web</string>',
|
||||
'<string>--port</string>',
|
||||
].map((s) => xml.indexOf(s));
|
||||
expect(order).toEqual([...order].sort((a, b) => a - b));
|
||||
expect(order.every((i) => i > -1)).toBe(true);
|
||||
});
|
||||
|
||||
it('carries the runner flags so a tsx dev install still boots', () => {
|
||||
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', execArgv: ['--import', 'tsx'] }));
|
||||
expect(xml).toContain('<string>--import</string>');
|
||||
expect(xml).toContain('<string>tsx</string>');
|
||||
});
|
||||
|
||||
it('restarts on crash and at login', () => {
|
||||
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd' }));
|
||||
expect(xml).toContain('<key>KeepAlive</key>');
|
||||
expect(xml).toContain('<key>RunAtLoad</key>');
|
||||
});
|
||||
|
||||
it('escapes a path with an ampersand instead of emitting broken XML', () => {
|
||||
const xml = buildLaunchAgentPlist(plan({ kind: 'launchd', workingDir: '/Users/a&b' }));
|
||||
expect(xml).toContain('<string>/Users/a&b</string>');
|
||||
expect(xml).not.toContain('<string>/Users/a&b</string>');
|
||||
});
|
||||
});
|
||||
|
||||
describe('buildSystemdUnit', () => {
|
||||
it('builds ExecStart from node, script and args', () => {
|
||||
expect(buildSystemdUnit(plan())).toContain(
|
||||
'ExecStart=/usr/bin/node /home/u/.codeman/app/dist/index.js web --host 127.0.0.1 --port 3000'
|
||||
);
|
||||
});
|
||||
|
||||
it('quotes an argument containing spaces', () => {
|
||||
const unit = buildSystemdUnit(plan({ scriptPath: '/home/my user/app/dist/index.js' }));
|
||||
expect(unit).toContain('"/home/my user/app/dist/index.js"');
|
||||
});
|
||||
|
||||
it('writes each env var as a quoted Environment line', () => {
|
||||
const unit = buildSystemdUnit(plan());
|
||||
expect(unit).toContain('Environment="PATH=/usr/bin:/bin"');
|
||||
expect(unit).toContain('Environment="HOME=/home/u"');
|
||||
});
|
||||
|
||||
it('keeps KillMode=process so agents survive a server restart', () => {
|
||||
expect(buildSystemdUnit(plan())).toContain('KillMode=process');
|
||||
});
|
||||
|
||||
it('is installable and restarts on failure', () => {
|
||||
const unit = buildSystemdUnit(plan());
|
||||
expect(unit).toContain('Restart=always');
|
||||
expect(unit).toContain('WantedBy=default.target');
|
||||
});
|
||||
});
|
||||
|
||||
describe('detectServiceKind', () => {
|
||||
it('maps the platform to its supervisor', () => {
|
||||
const expected = process.platform === 'darwin' ? 'launchd' : process.platform === 'linux' ? 'systemd' : null;
|
||||
expect(detectServiceKind()).toBe(expected);
|
||||
});
|
||||
});
|
||||
@@ -64,3 +64,38 @@ describe('buildMuxAttachEnv', () => {
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('spawn env CODEMAN_API_URL (no fallback)', () => {
|
||||
const withApiUrl = (value: string | undefined, fn: () => void) => {
|
||||
const original = process.env.CODEMAN_API_URL;
|
||||
if (value === undefined) delete process.env.CODEMAN_API_URL;
|
||||
else process.env.CODEMAN_API_URL = value;
|
||||
try {
|
||||
fn();
|
||||
} finally {
|
||||
if (original === undefined) delete process.env.CODEMAN_API_URL;
|
||||
else process.env.CODEMAN_API_URL = original;
|
||||
}
|
||||
};
|
||||
|
||||
it('passes the server-stamped URL through verbatim', async () => {
|
||||
const { buildClaudeEnv, buildShellEnv } = await import('../src/session-cli-builder.js');
|
||||
withApiUrl('https://127.0.0.1:3199', () => {
|
||||
expect(buildClaudeEnv('test-session').CODEMAN_API_URL).toBe('https://127.0.0.1:3199');
|
||||
expect(buildShellEnv('test-session').CODEMAN_API_URL).toBe('https://127.0.0.1:3199');
|
||||
});
|
||||
});
|
||||
|
||||
// A hardcoded fallback was the wrong scheme on HTTPS installs. The key must be
|
||||
// genuinely ABSENT when unset: present-with-undefined would serialize through
|
||||
// node-pty as the literal string "CODEMAN_API_URL=undefined" (COD-115).
|
||||
it('leaves the key absent (not undefined, not a fallback) when the server has not stamped one', async () => {
|
||||
const { buildClaudeEnv, buildShellEnv } = await import('../src/session-cli-builder.js');
|
||||
withApiUrl(undefined, () => {
|
||||
for (const env of [buildClaudeEnv('test-session'), buildShellEnv('test-session')]) {
|
||||
expect('CODEMAN_API_URL' in env).toBe(false);
|
||||
expect(JSON.stringify(env)).not.toContain('localhost:3000');
|
||||
}
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,89 @@
|
||||
/**
|
||||
* @fileoverview `/api/events` must not lose the headers the security hook set.
|
||||
*
|
||||
* The SSE route answers with `reply.raw.writeHead()`, which writes straight to the
|
||||
* Node response and bypasses Fastify's header store. Everything the `onRequest`
|
||||
* security hook had granted was therefore dropped — including the
|
||||
* `Access-Control-Allow-Origin` it emits for localhost origins. The contradiction is
|
||||
* visible from a browser: a localhost page may call every other `/api` endpoint
|
||||
* cross-origin, but its EventSource fails CORS.
|
||||
*
|
||||
* These tests drive a REAL WebServer. An earlier version asserted against an inline
|
||||
* copy of the hook and the handler, which proved nothing: reverting the fix in
|
||||
* `server.ts` left every test green.
|
||||
*/
|
||||
|
||||
import { afterAll, beforeAll, describe, expect, it } from 'vitest';
|
||||
|
||||
import { WebServer } from '../src/web/server.js';
|
||||
|
||||
const TEST_PORT = 3119;
|
||||
const LOCAL_ORIGIN = 'http://localhost:5173';
|
||||
|
||||
/** Open /api/events, read the response headers, then abort — it never ends on its own. */
|
||||
async function eventsHeaders(baseUrl: string, origin?: string): Promise<Headers> {
|
||||
const controller = new AbortController();
|
||||
const timeout = setTimeout(() => controller.abort(), 2000);
|
||||
try {
|
||||
const res = await fetch(`${baseUrl}/api/events`, {
|
||||
signal: controller.signal,
|
||||
headers: origin ? { Origin: origin } : undefined,
|
||||
});
|
||||
const headers = res.headers;
|
||||
controller.abort(); // stop consuming the stream
|
||||
return headers;
|
||||
} finally {
|
||||
clearTimeout(timeout);
|
||||
}
|
||||
}
|
||||
|
||||
describe('GET /api/events header inheritance', () => {
|
||||
let server: WebServer;
|
||||
let baseUrl: string;
|
||||
|
||||
beforeAll(async () => {
|
||||
server = new WebServer(TEST_PORT, false, true);
|
||||
await server.start();
|
||||
baseUrl = `http://localhost:${TEST_PORT}`;
|
||||
});
|
||||
|
||||
afterAll(async () => {
|
||||
await server.stop();
|
||||
}, 60000);
|
||||
|
||||
it('keeps the CORS header the security hook granted a localhost origin', async () => {
|
||||
// The regression: this header is set on the Fastify reply and was then thrown
|
||||
// away by writeHead, so an EventSource from a localhost dev server failed CORS
|
||||
// while every other endpoint worked.
|
||||
const headers = await eventsHeaders(baseUrl, LOCAL_ORIGIN);
|
||||
expect(headers.get('access-control-allow-origin')).toBe(LOCAL_ORIGIN);
|
||||
});
|
||||
|
||||
it('keeps the security headers the hook set', async () => {
|
||||
const headers = await eventsHeaders(baseUrl);
|
||||
expect(headers.get('x-content-type-options')).toBe('nosniff');
|
||||
expect(headers.get('x-frame-options')).toBe('SAMEORIGIN');
|
||||
expect(headers.get('content-security-policy')).toBeTruthy();
|
||||
});
|
||||
|
||||
it('still sets the SSE headers, and they win over anything inherited', async () => {
|
||||
const headers = await eventsHeaders(baseUrl);
|
||||
expect(headers.get('content-type')).toBe('text/event-stream');
|
||||
expect(headers.get('cache-control')).toBe('no-cache');
|
||||
expect(headers.get('x-accel-buffering')).toBe('no');
|
||||
});
|
||||
|
||||
it('grants nothing to a non-localhost origin — the hook decides, not this route', async () => {
|
||||
const headers = await eventsHeaders(baseUrl, 'https://evil.example');
|
||||
expect(headers.get('access-control-allow-origin')).toBeNull();
|
||||
});
|
||||
|
||||
it('matches what a normal JSON endpoint returns for the same origin', async () => {
|
||||
// The point of the fix: /api/events stops being the odd one out.
|
||||
const json = await fetch(`${baseUrl}/api/status`, { headers: { Origin: LOCAL_ORIGIN } });
|
||||
const sse = await eventsHeaders(baseUrl, LOCAL_ORIGIN);
|
||||
|
||||
expect(sse.get('access-control-allow-origin')).toBe(json.headers.get('access-control-allow-origin'));
|
||||
expect(sse.get('x-content-type-options')).toBe(json.headers.get('x-content-type-options'));
|
||||
});
|
||||
});
|
||||
@@ -47,6 +47,26 @@ function loadTerminalUiHarness(mode: string) {
|
||||
}
|
||||
|
||||
describe('terminal flush budget', () => {
|
||||
it('drains a large final batch without waiting for unrelated terminal output', () => {
|
||||
const { app, writes } = loadTerminalUiHarness('codex');
|
||||
const scheduled: Array<() => void> = [];
|
||||
app._safeYield = (callback: () => void) => {
|
||||
scheduled.push(callback);
|
||||
};
|
||||
app.isTerminalAtBottom = () => true;
|
||||
|
||||
app.batchTerminalWrite('x'.repeat(96 * 1024));
|
||||
expect(scheduled).toHaveLength(1);
|
||||
|
||||
while (scheduled.length > 0) {
|
||||
scheduled.shift()?.();
|
||||
}
|
||||
|
||||
expect(writes.map((write) => write.length)).toEqual([32 * 1024, 32 * 1024, 32 * 1024]);
|
||||
expect(app.pendingWrites).toEqual([]);
|
||||
expect(app.writeFrameScheduled).toBe(false);
|
||||
});
|
||||
|
||||
it('uses a smaller first-frame write budget for Codex output to reduce renderer stalls', () => {
|
||||
const { app, writes } = loadTerminalUiHarness('codex');
|
||||
app.pendingWrites.push('x'.repeat(96 * 1024));
|
||||
|
||||
@@ -0,0 +1,240 @@
|
||||
/**
|
||||
* Issue #205, round 2: the 1.12.0 retest still reported unusable scrollback —
|
||||
* a completely dead wheel on Firefox/macOS (while Fn+Up paged back through
|
||||
* intact text), and history on iPhone that went back a little, repeated blocks
|
||||
* and got worse the further up it went.
|
||||
*
|
||||
* Both signatures come from a Claude pane's LOCAL buffer being hollow. tmux
|
||||
* keeps no history for a repaint-mode pane (`history_size≈0`), so:
|
||||
* - any gesture routed to local scrollback scrolls nothing, and
|
||||
* - the scroll-to-top `?full=1` re-pull replaces a multi-frame buffer with a
|
||||
* single captured frame, deleting history mid-scroll.
|
||||
*
|
||||
* These cover the two guards that fix it: `_replayWouldShrinkBuffer` (refuse a
|
||||
* downgrading re-pull) and `_maybePageCliTranscript` (page the CLI's own
|
||||
* transcript when there is nothing local to scroll), plus the diagnostic that
|
||||
* makes the routing decision visible instead of guessable.
|
||||
*/
|
||||
import { readFileSync } from 'node:fs';
|
||||
import { resolve } from 'node:path';
|
||||
import vm from 'node:vm';
|
||||
import { describe, expect, it, vi } from 'vitest';
|
||||
|
||||
function loadTerminalUiHarness() {
|
||||
const CodemanApp = function CodemanApp(this: any) {};
|
||||
const logs: string[] = [];
|
||||
const context = vm.createContext({
|
||||
window: {},
|
||||
CodemanApp,
|
||||
console: { warn: vi.fn(), log: (msg: string) => logs.push(msg) },
|
||||
_crashDiag: { log: vi.fn() },
|
||||
performance: { now: () => 1_000 },
|
||||
requestAnimationFrame: (_fn: () => void) => 1,
|
||||
setTimeout: (_fn: () => void) => 1,
|
||||
Blob: function Blob() {},
|
||||
URL: { createObjectURL: () => 'blob:yield', revokeObjectURL: () => {} },
|
||||
Worker: function Worker(this: any) {
|
||||
this.postMessage = () => {};
|
||||
},
|
||||
MobileDetection: { isTouchDevice: () => true },
|
||||
DEC_SYNC_STRIP_RE: /\x1b\[\?2026[hl]/g,
|
||||
TERMINAL_CHUNK_SIZE: 32 * 1024,
|
||||
});
|
||||
|
||||
const code = readFileSync(resolve(import.meta.dirname, '../src/web/public/terminal-ui.js'), 'utf8');
|
||||
vm.runInContext(code, context, { filename: 'terminal-ui.js' });
|
||||
return { app: new (CodemanApp as any)(), logs };
|
||||
}
|
||||
|
||||
/** A Claude session whose local buffer holds exactly one screen (baseY 0). */
|
||||
function hollowClaudeApp(overrides: { cliVersion?: string; rows?: number } = {}) {
|
||||
const { app, logs } = loadTerminalUiHarness();
|
||||
const sent: Array<{ id: string; data: string }> = [];
|
||||
app.activeSessionId = 'sess-1';
|
||||
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: overrides.cliVersion }]]);
|
||||
app._sendInputEphemeral = (id: string, data: string) => sent.push({ id, data });
|
||||
app.terminal = {
|
||||
cols: 80,
|
||||
rows: overrides.rows ?? 36,
|
||||
modes: { mouseTrackingMode: 'none' },
|
||||
buffer: { active: { type: 'normal', viewportY: 0, baseY: 0, length: 36 } },
|
||||
};
|
||||
return { app, sent, logs };
|
||||
}
|
||||
|
||||
describe('full-history re-pull downgrade guard (issue #205 round 2)', () => {
|
||||
it('estimates replayed rows from wrapped, escape-laden capture text', () => {
|
||||
const { app } = loadTerminalUiHarness();
|
||||
|
||||
expect(app._estimateReplayRows('a\r\nb\r\nc', 80)).toBe(3);
|
||||
// SGR colour runs occupy no cells, so they must not inflate the estimate.
|
||||
expect(app._estimateReplayRows('\x1b[38;5;196mred\x1b[0m\r\nplain', 80)).toBe(2);
|
||||
// capture-pane -J joins wrapped rows, so a long logical line re-wraps on
|
||||
// write — counting newlines alone would undershoot by 2 rows here.
|
||||
expect(app._estimateReplayRows('x'.repeat(25), 10)).toBe(3);
|
||||
expect(app._estimateReplayRows('', 80)).toBe(0);
|
||||
expect(app._estimateReplayRows(undefined, 80)).toBe(0);
|
||||
});
|
||||
|
||||
it('refuses a capture that would leave LESS history than the terminal holds', () => {
|
||||
const { app } = loadTerminalUiHarness();
|
||||
app.terminal = { cols: 80, rows: 36, buffer: { active: { length: 300 } } };
|
||||
|
||||
// Claude pane: tmux has no history, so the capture is one frame while xterm
|
||||
// holds hundreds of replayed rows. Rewriting would delete them mid-scroll.
|
||||
const oneFrame = Array.from({ length: 36 }, (_, i) => `frame line ${i}`).join('\r\n');
|
||||
expect(app._replayWouldShrinkBuffer(oneFrame)).toBe(true);
|
||||
|
||||
// Shell pane after a burst/tab-switch collapse: tmux really does hold more.
|
||||
const realHistory = Array.from({ length: 800 }, (_, i) => `history ${i}`).join('\r\n');
|
||||
expect(app._replayWouldShrinkBuffer(realHistory)).toBe(false);
|
||||
});
|
||||
|
||||
it('tolerates a one-screen shortfall so ordinary recoveries still replay', () => {
|
||||
const { app } = loadTerminalUiHarness();
|
||||
// buffer.active.length counts the blank rows under the last line and the row
|
||||
// estimate can only approximate wrapping, so a near-tie must NOT read as a
|
||||
// downgrade — only a capture worse by more than a full screen does.
|
||||
app.terminal = { cols: 80, rows: 36, buffer: { active: { length: 120 } } };
|
||||
expect(app._replayWouldShrinkBuffer(Array.from({ length: 100 }, () => 'x').join('\r\n'))).toBe(false);
|
||||
expect(app._replayWouldShrinkBuffer(Array.from({ length: 40 }, () => 'x').join('\r\n'))).toBe(true);
|
||||
});
|
||||
|
||||
it('never refuses when the terminal has no buffer to protect', () => {
|
||||
const { app } = loadTerminalUiHarness();
|
||||
app.terminal = { cols: 80, rows: 36, buffer: { active: { length: 0 } } };
|
||||
expect(app._replayWouldShrinkBuffer('anything')).toBe(false);
|
||||
});
|
||||
|
||||
it('is wired into _maybeRefetchFullHistory BEFORE the destructive reset', () => {
|
||||
const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8');
|
||||
const start = source.indexOf('async _maybeRefetchFullHistory()');
|
||||
const guard = source.indexOf('this._replayWouldShrinkBuffer(buffer)', start);
|
||||
const reset = source.indexOf('this._resetTerminalForReplay()', start);
|
||||
|
||||
expect(start).toBeGreaterThan(-1);
|
||||
expect(guard).toBeGreaterThan(start);
|
||||
expect(guard).toBeLessThan(reset); // refuse first, only then reset+rewrite
|
||||
// A hollow pane must also stop re-fetching megabytes on every scroll-up.
|
||||
expect(source).toContain('this._fullHistoryRepullUseless');
|
||||
expect(source).toContain('this._fullHistoryRepullUseless?.has(sessionId) ? 60000 : 4000');
|
||||
});
|
||||
});
|
||||
|
||||
describe('PageUp/PageDown fallback for a hollow local buffer (issue #205 round 2)', () => {
|
||||
it('pages the CLI transcript when the wheel gate is false and there is no scrollback', () => {
|
||||
const { app, sent } = hollowClaudeApp(); // cliVersion unknown → gate false
|
||||
|
||||
// Half a screen of travel (rows 36 → 18 lines) buys exactly one PageUp.
|
||||
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(true);
|
||||
app._flushWheelSgrQueue();
|
||||
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[5~' }]);
|
||||
|
||||
// Downward travel pages back toward the live screen.
|
||||
app._maybePageCliTranscript({ shiftKey: false }, 18);
|
||||
app._flushWheelSgrQueue();
|
||||
expect(sent[1]).toEqual({ id: 'sess-1', data: '\x1b[6~' });
|
||||
});
|
||||
|
||||
it('accumulates sub-page travel instead of dropping or over-sending it', () => {
|
||||
const { app, sent } = hollowClaudeApp();
|
||||
|
||||
expect(app._maybePageCliTranscript({ shiftKey: false }, -10)).toBe(true); // consumed…
|
||||
app._flushWheelSgrQueue();
|
||||
expect(sent).toEqual([]); // …but below the threshold, so nothing sent yet
|
||||
|
||||
app._maybePageCliTranscript({ shiftKey: false }, -8); // -18 total → one page
|
||||
app._flushWheelSgrQueue();
|
||||
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[5~' }]);
|
||||
});
|
||||
|
||||
it('caps the keys one gesture batch can emit', () => {
|
||||
const { app, sent } = hollowClaudeApp();
|
||||
|
||||
app._maybePageCliTranscript({ shiftKey: false }, -1000); // 55 pages of travel
|
||||
app._flushWheelSgrQueue();
|
||||
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[5~'.repeat(3) }]);
|
||||
});
|
||||
|
||||
it('leaves every session that has real local scrollback alone', () => {
|
||||
const { app } = hollowClaudeApp();
|
||||
|
||||
// Shift is the explicit "give me local scrollback" gesture — never paged.
|
||||
expect(app._maybePageCliTranscript({ shiftKey: true }, -18)).toBe(false);
|
||||
|
||||
// A buffer with history scrolls locally, as before.
|
||||
app.terminal.buffer.active.baseY = 120;
|
||||
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(false);
|
||||
app.terminal.buffer.active.baseY = 0;
|
||||
|
||||
// Non-Claude modes keep their existing behavior (shell scrolls tmux history
|
||||
// through the alt-screen strip; codex/gemini page keys are unverified).
|
||||
app.sessions = new Map([['sess-1', { mode: 'shell' }]]);
|
||||
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(false);
|
||||
app.sessions = new Map([['sess-1', { mode: 'codex' }]]);
|
||||
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(false);
|
||||
|
||||
// An alternate-screen pane belongs to xterm's own alt-scroll handling.
|
||||
app.sessions = new Map([['sess-1', { mode: 'claude' }]]);
|
||||
app.terminal.buffer.active.type = 'alternate';
|
||||
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(false);
|
||||
});
|
||||
|
||||
it('rescues the local-scrollback opt-out footgun instead of silently dying', () => {
|
||||
// "Wheel scrolls local history" ON pins the wheel to a buffer that, for a
|
||||
// repaint-mode CLI, is empty — a user who flipped it while hunting for a fix
|
||||
// on 1.11.x would have ended up with a completely dead wheel on 1.12.0.
|
||||
const { app, sent } = hollowClaudeApp({ cliVersion: '2.1.223' }); // gate would forward…
|
||||
app.loadAppSettingsFromStorage = () => ({ terminalWheelLocalScrollback: true });
|
||||
|
||||
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); // …but the opt-out wins
|
||||
expect(app._maybePageCliTranscript({ shiftKey: false }, -18)).toBe(true);
|
||||
app._flushWheelSgrQueue();
|
||||
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[5~' }]);
|
||||
});
|
||||
|
||||
it('drops travel accumulated on another tab', () => {
|
||||
const { app, sent } = hollowClaudeApp();
|
||||
|
||||
app._maybePageCliTranscript({ shiftKey: false }, -17); // just short of a page
|
||||
app.activeSessionId = 'sess-2';
|
||||
app.sessions.set('sess-2', { mode: 'claude' });
|
||||
app._maybePageCliTranscript({ shiftKey: false }, -1); // must not complete sess-1's page
|
||||
app._flushWheelSgrQueue();
|
||||
expect(sent).toEqual([]);
|
||||
});
|
||||
|
||||
it('is reachable from both the wheel and the touch paths', () => {
|
||||
const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/terminal-ui.js'), 'utf8');
|
||||
// Wheel: after the forwarding gate, before the local smooth scroll.
|
||||
expect(source).toContain('if (this._maybePageCliTranscript(ev, lines)) return;');
|
||||
// Touch: touchmove and the momentum loop both fall through to it.
|
||||
expect(source.match(/else if \(!this\._maybePageCliTranscript\(\{ shiftKey: false \}, lines\)\)/g)).toHaveLength(2);
|
||||
});
|
||||
});
|
||||
|
||||
describe('scroll routing diagnostic (issue #205 round 2)', () => {
|
||||
it('prints the decision and its inputs once per session, and again when it changes', () => {
|
||||
const { app, logs } = hollowClaudeApp({ cliVersion: '2.1.100' });
|
||||
app.loadAppSettingsFromStorage = () => ({ terminalWheelLocalScrollback: false });
|
||||
|
||||
app._logScrollRouting('local-scrollback');
|
||||
app._logScrollRouting('local-scrollback'); // same decision → stays quiet
|
||||
expect(logs).toHaveLength(1);
|
||||
expect(logs[0]).toContain('sess-1 → local-scrollback');
|
||||
expect(logs[0]).toContain('mode=claude');
|
||||
expect(logs[0]).toContain('cliVersion=2.1.100');
|
||||
expect(logs[0]).toContain('localScrollbackOptOut=false');
|
||||
expect(logs[0]).toContain('mouseTracking=none');
|
||||
|
||||
app._logScrollRouting('page-keys'); // a changed route still prints
|
||||
expect(logs).toHaveLength(2);
|
||||
expect(logs[1]).toContain('page-keys');
|
||||
});
|
||||
|
||||
it('reports an unknown CLI version, the false-path that disables forwarding', () => {
|
||||
const { app, logs } = hollowClaudeApp(); // no cliVersion — the probe failed
|
||||
app._logScrollRouting('page-keys');
|
||||
expect(logs[0]).toContain('cliVersion=unknown');
|
||||
});
|
||||
});
|
||||
@@ -311,7 +311,7 @@ describe('terminal touch tap mouse guard', () => {
|
||||
expect(sent).toEqual(['\x1b[<0;7;4M\x1b[<0;7;4m']);
|
||||
});
|
||||
|
||||
it('wheel: forwards to the app only for verified sessions at the buffer bottom without Shift', () => {
|
||||
it('wheel: forwards to the app for verified sessions without Shift, at ANY scroll position', () => {
|
||||
const { app } = loadTerminalUiHarness();
|
||||
app.activeSessionId = 'sess-1';
|
||||
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187' }]]);
|
||||
@@ -323,8 +323,14 @@ describe('terminal touch tap mouse guard', () => {
|
||||
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(true);
|
||||
expect(app._shouldForwardWheelToApp({ shiftKey: true })).toBe(false); // Shift = local scrollback
|
||||
|
||||
app.terminal.buffer.active.viewportY = 10; // browsing local scrollback
|
||||
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
|
||||
// Scrolled up into local scrollback still forwards. Gating this on the
|
||||
// viewport being at the bottom is what let a repaint-mode CLI's own prompt
|
||||
// box scroll off the screen: scrollToLastNonEmptyLine() parks the viewport
|
||||
// above the bottom, so a tab switch silently pinned the wheel to local
|
||||
// scrollback full of stale replayed frames. The wheel handler snaps the
|
||||
// viewport back to the bottom before encoding the report instead.
|
||||
app.terminal.buffer.active.viewportY = 10;
|
||||
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(true);
|
||||
app.terminal.buffer.active.viewportY = 50;
|
||||
|
||||
app.terminal.modes.mouseTrackingMode = 'vt200'; // xterm's own encoder live
|
||||
@@ -335,6 +341,24 @@ describe('terminal touch tap mouse guard', () => {
|
||||
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
|
||||
});
|
||||
|
||||
it('wheel: converts deltaMode line/page units instead of assuming pixels', () => {
|
||||
const { app } = loadTerminalUiHarness();
|
||||
app.terminal = { rows: 40 };
|
||||
|
||||
// DOM_DELTA_PIXEL (Chrome/WebKit, and every trackpad): ~110px per notch.
|
||||
expect(app._wheelScrollLines({ deltaY: 110, deltaX: 0, deltaMode: 0, shiftKey: false })).toBe(4);
|
||||
// DOM_DELTA_LINE (Firefox mouse wheel): deltaY is already lines. Read as
|
||||
// pixels this rounded to 0 and fell through to the ±1 fallback.
|
||||
expect(app._wheelScrollLines({ deltaY: 3, deltaX: 0, deltaMode: 1, shiftKey: false })).toBe(3);
|
||||
expect(app._wheelScrollLines({ deltaY: -3, deltaX: 0, deltaMode: 1, shiftKey: false })).toBe(-3);
|
||||
// DOM_DELTA_PAGE: one page is one screenful.
|
||||
expect(app._wheelScrollLines({ deltaY: 1, deltaX: 0, deltaMode: 2, shiftKey: false })).toBe(40);
|
||||
// A pure horizontal swipe must not fall through to a phantom -1.
|
||||
expect(app._wheelScrollLines({ deltaY: 0, deltaX: 90, deltaMode: 0, shiftKey: false })).toBe(0);
|
||||
// Shift + macOS trackpad reports the magnitude on deltaX (issue #154).
|
||||
expect(app._wheelScrollLines({ deltaY: 0, deltaX: -100, deltaMode: 0, shiftKey: true })).toBe(-4);
|
||||
});
|
||||
|
||||
it('wheel: gates claude forwarding on CLI version 2.1.187+ (unknown or older stays local)', () => {
|
||||
const { app } = loadTerminalUiHarness();
|
||||
app.activeSessionId = 'sess-1';
|
||||
@@ -438,6 +462,39 @@ describe('terminal touch tap mouse guard', () => {
|
||||
expect(sent).toHaveLength(1);
|
||||
});
|
||||
|
||||
it('forwarded scrolls (wheel AND touch) snap the viewport home first, then encode SGR ticks', () => {
|
||||
const { app } = loadTerminalUiHarness();
|
||||
const sent: Array<{ id: string; data: string }> = [];
|
||||
app.activeSessionId = 'sess-1';
|
||||
app.sessions = new Map([['sess-1', { mode: 'claude' }]]);
|
||||
app._sendInputEphemeral = (id: string, data: string) => sent.push({ id, data });
|
||||
const scrolledToBottom: boolean[] = [];
|
||||
app.terminal = {
|
||||
cols: 80,
|
||||
rows: 24,
|
||||
// Scrolled up into local scrollback: SGR coordinates address the LIVE
|
||||
// screen, so the report would hit-test the wrong row without the snap.
|
||||
buffer: { active: { viewportY: 10, baseY: 50 } },
|
||||
scrollToBottom: () => scrolledToBottom.push(true),
|
||||
element: {
|
||||
querySelector: () => ({ getBoundingClientRect: () => ({ left: 0, top: 0 }) }),
|
||||
},
|
||||
_core: { _renderService: { dimensions: { css: { cell: { width: 8, height: 16 } } } } },
|
||||
};
|
||||
|
||||
app._forwardScrollToApp(50, 50, -3);
|
||||
expect(scrolledToBottom).toEqual([true]);
|
||||
app._flushWheelSgrQueue();
|
||||
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[<64;7;4M'.repeat(3) }]);
|
||||
|
||||
// Already at the bottom: no snap, just the report.
|
||||
app.terminal.buffer.active.viewportY = 50;
|
||||
app._forwardScrollToApp(50, 50, 2);
|
||||
expect(scrolledToBottom).toHaveLength(1);
|
||||
app._flushWheelSgrQueue();
|
||||
expect(sent).toHaveLength(2);
|
||||
});
|
||||
|
||||
it('allows trusted mouse events after the tap window expires', () => {
|
||||
const { app, setNow } = loadTerminalUiHarness();
|
||||
const { element, dispatch } = createElementHarness();
|
||||
|
||||
@@ -247,14 +247,41 @@ describe('TmuxManager (unit)', () => {
|
||||
});
|
||||
|
||||
describe('environment exports', () => {
|
||||
it('keeps COLORTERM unset for OpenCode sessions', () => {
|
||||
const exports = (
|
||||
const callBuildEnvExports = (mode: string) =>
|
||||
(
|
||||
manager as unknown as {
|
||||
buildEnvExports(sessionId: string, muxName: string, mode: string): string[];
|
||||
}
|
||||
).buildEnvExports('session-1', 'codeman-abc12345', 'opencode');
|
||||
).buildEnvExports('session-1', 'codeman-abc12345', mode);
|
||||
|
||||
expect(exports).toContain('unset COLORTERM');
|
||||
it('keeps COLORTERM unset for OpenCode sessions', () => {
|
||||
expect(callBuildEnvExports('opencode')).toContain('unset COLORTERM');
|
||||
});
|
||||
|
||||
it('exports the server-stamped CODEMAN_API_URL verbatim', () => {
|
||||
const original = process.env.CODEMAN_API_URL;
|
||||
process.env.CODEMAN_API_URL = 'https://127.0.0.1:3199';
|
||||
try {
|
||||
expect(callBuildEnvExports('claude')).toContain('export CODEMAN_API_URL=https://127.0.0.1:3199');
|
||||
} finally {
|
||||
if (original === undefined) delete process.env.CODEMAN_API_URL;
|
||||
else process.env.CODEMAN_API_URL = original;
|
||||
}
|
||||
});
|
||||
|
||||
// A hardcoded fallback exported the wrong scheme on HTTPS installs; unset must
|
||||
// stay unset so in-session guards fail closed instead of curling a bad URL.
|
||||
it('exports no CODEMAN_API_URL at all when the server has not stamped one', () => {
|
||||
const original = process.env.CODEMAN_API_URL;
|
||||
delete process.env.CODEMAN_API_URL;
|
||||
try {
|
||||
const exports = callBuildEnvExports('claude');
|
||||
expect(exports.some((line) => line.startsWith('export CODEMAN_API_URL'))).toBe(false);
|
||||
expect(exports.join(' ')).not.toContain('localhost:3000');
|
||||
} finally {
|
||||
if (original === undefined) delete process.env.CODEMAN_API_URL;
|
||||
else process.env.CODEMAN_API_URL = original;
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
Reference in New Issue
Block a user