Compare commits

..
Author SHA1 Message Date
Codeman maintainer 915895152c phone-tab-strip: before/after for the phone header tab strip PR 2026-09-28 15:23:31 +02:00
Codeman maintainer 7d5b0bec1c design assets: tab strip redesign mockups
Screenshots for the header + tab strip redesign discussion. Orphan branch, never merged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 00:27:00 +02:00
840 changed files with 3 additions and 315278 deletions
-27
View File
@@ -1,27 +0,0 @@
# Changesets
Hello and welcome! This folder has been automatically generated by `@changesets/cli`, a build tool that works
with multi-package repos, or single-package repos to help you version and publish your code. You can find
the full documentation for it [in the repository](https://github.com/changesets/changesets).
## What is a changeset?
A changeset is a piece of information about changes made in a branch or commit. It holds three bits of information:
- What packages need to be released
- What semver bump type each package should receive (major / minor / patch)
- A summary of the changes
## How do I create a changeset?
Run `npx changeset` or create a `.md` file in this directory with the following format:
```markdown
---
"codeman": patch
---
Description of changes
```
The frontmatter specifies which package(s) to bump and the bump type. The body is the changelog entry.
-11
View File
@@ -1,11 +0,0 @@
{
"$schema": "https://unpkg.com/@changesets/config@3.1.1/schema.json",
"changelog": "@changesets/cli/changelog",
"commit": false,
"fixed": [],
"linked": [],
"access": "public",
"baseBranch": "master",
"updateInternalDependencies": "patch",
"ignore": []
}
-12
View File
@@ -1,12 +0,0 @@
root = true
[*]
indent_style = space
indent_size = 2
end_of_line = lf
charset = utf-8
trim_trailing_whitespace = true
insert_final_newline = true
[*.md]
trim_trailing_whitespace = false
-78
View File
@@ -1,78 +0,0 @@
# Security Policy
Codeman launches AI coding sessions with `--dangerously-skip-permissions`, so the
web UI is **by design a remote-code-execution surface for whoever can reach it**.
The entire security model exists to control *who* that is. Please read this before
exposing an instance beyond `localhost`. The full model lives in
[`docs/security-architecture.md`](../docs/security-architecture.md).
## Supported versions
Security fixes land on the latest published `codeman@X.Y.Z` release and `master`.
Older versions are not patched — upgrade to the latest release (App Settings →
Updates for git-clone installs, or `npm i -g aicodeman@latest`).
| Version | Supported |
| ------- | --------- |
| latest `0.9.x` / `master` | ✅ |
| anything older | ❌ (upgrade) |
## Reporting a vulnerability
**Please do not open a public issue for security problems.**
Report privately via **GitHub's private vulnerability reporting**:
the repository's **Security** tab → **Report a vulnerability**
(<https://github.com/Ark0N/Codeman/security/advisories/new>). This opens a private
advisory thread with the maintainer.
> Maintainer note: enable *Settings → Code security and analysis → Private
> vulnerability reporting* so this channel is live.
When reporting, please include: affected version/commit, the deployment shape
(loopback-only, `CODEMAN_PASSWORD` set, tunnel/`tailscale serve`, custom
reverse proxy), reproduction steps, and impact. We aim to acknowledge within a
few days. Coordinated disclosure is appreciated — we'll agree a disclosure
timeline with you once impact is confirmed.
### In scope
- Authentication / session-cookie bypass when `CODEMAN_PASSWORD` is set
- DNS-rebinding, CSRF/CSWSH, or Origin/Host-guard bypass reaching state-changing routes
- Remote code execution reachable **without** local OS access (e.g. via a browser, a tunnel, or a foreign origin)
- Path traversal / arbitrary file read or write through the HTTP API
- Supply-chain integrity of the in-app self-updater
### Out of scope (by design — see Known limitations)
- Anything requiring an already-trusted **same-machine, same-uid** process. Codeman trusts the local OS user it runs as; a peer process of that user is already inside the boundary.
- Running an authless instance bound to a non-loopback host after dismissing the startup warning (you explicitly acknowledged it).
- The default loopback + no-password posture itself (it is reachable only from the same machine).
## Trust model (summary)
- **Loopback by default.** Binds `127.0.0.1`; the no-password default is safe out of the box. Binding a non-loopback host without `CODEMAN_PASSWORD` *starts but prints a loud warning* with concrete fixes.
- **Always-on Host + Origin guards.** Block DNS-rebinding and cross-site state-changing requests even on the no-auth loopback install (a missing Origin is allowed so CLI/hooks work).
- **Optional auth.** HTTP Basic via `CODEMAN_USERNAME`/`CODEMAN_PASSWORD`; success issues an opaque server-side 256-bit cookie. Per-IP rate limiting on failures.
- **Hardened file serving, tmux launch, transport headers, and multi-instance isolation** — see the full architecture doc.
## Known limitations and accepted risk
A 1.0 release is an implicit statement that the documented model *is* the model, so
these residuals are stated explicitly. Most sit **inside the same-uid OS trust
boundary** or behind the always-on Origin guard; they matter mainly for
shared-host, multi-user, or tunneled deployments.
- **Self-update trusts an unsigned release tag.** The in-app updater does `git checkout <tag> && npm install` (lifecycle scripts run) of a tag matched only by name shape, from whatever `origin` points to — no signature/commit verification. Treat the updater as trusting your `origin` remote and your release pipeline. (Hardening tracked for 1.0.)
- **CSP ships `'unsafe-inline'`.** Inline handlers mean the Content-Security-Policy is defense-in-depth only; all AI-/file-derived sinks are escaped, but a future missed escape would be executable.
- **`workingDir` is unconstrained.** A session may be created with any absolute working directory (e.g. `/`), which becomes the file-route boundary for that session. Scope it to trusted paths on shared hosts.
- **Hook-event auth exemption is loopback-IP-based.** `POST /api/hook-event` is exempt from auth for loopback callers; because tunnels (cloudflared / `tailscale serve`) terminate at `127.0.0.1`, a loopback-terminating tunnel inherits the exemption. Set `CODEMAN_PASSWORD` and prefer a tunnel that preserves the client identity if this matters.
- **Session cookie is not bound to client IP/UA on reuse, and refreshes without an absolute cap.** A stolen cookie replays until its idle TTL elapses.
- **Multi-instance tmux socket is process-wide.** Two Codeman instances on the same `CODEMAN_INSTANCE` share a tmux socket and can attach each other's live sessions — isolate with distinct `CODEMAN_INSTANCE` values.
- **The live log-tail route reads `/var/log` and `~/logs`** in addition to the session working directory (read-only) — a deliberate choice for tailing system/app logs. On a password-protected remote deployment an authenticated user can therefore read those roots outside their session. See `docs/security-architecture.md` §5.
Recent hardening (this release): web-push subscription endpoints are restricted
to https public hosts (SSRF guard — rejects internal/metadata IPs, validated at
subscribe and send time), and tmux session names discovered on the shared socket
are validated against the safe-name pattern before reaching any shell call site.
For the detailed rationale, defenses, and recommended secure setups, see
[`docs/security-architecture.md`](../docs/security-architecture.md).
-104
View File
@@ -1,104 +0,0 @@
name: CI
on:
push:
branches: [master, main]
pull_request:
jobs:
ci:
name: Typecheck & Lint
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- name: Setup Node.js
uses: actions/setup-node@v6
with:
node-version: 22
cache: 'npm'
- name: Install dependencies
run: npm ci
- name: Check package-lock.json version sync
run: npm run check:lockfile
- name: Type check
run: npm run typecheck
- name: Lint
run: npm run lint
- name: Frontend JS syntax check
run: npm run check:frontend-syntax
- name: Format check
run: npm run format:check
- name: Server boot smoke test
run: |
set -u
if ! command -v tmux >/dev/null; then
sudo apt-get update -qq
sudo apt-get install -y tmux
fi
npx tsx src/index.ts web --port 3151 > /tmp/boot.log 2>&1 &
SERVER_PID=$!
trap "kill $SERVER_PID 2>/dev/null || true" EXIT
for i in $(seq 1 30); do
if curl -fsS http://localhost:3151/api/status -o /dev/null; then
echo "Server booted in ${i}s"
exit 0
fi
if ! kill -0 $SERVER_PID 2>/dev/null; then
echo "Server exited before becoming ready. Logs:"
cat /tmp/boot.log
exit 1
fi
sleep 1
done
echo "Server did not respond on /api/status within 30s. Logs:"
cat /tmp/boot.log
exit 1
test:
name: Unit & integration tests
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- name: Setup Node.js
uses: actions/setup-node@v6
with:
node-version: 22
cache: 'npm'
- name: Install dependencies
run: npm ci
- name: Install tmux
run: |
if ! command -v tmux >/dev/null; then
sudo apt-get update -qq
sudo apt-get install -y tmux
fi
- name: Run unit & integration tests
# Excludes the browser-driven mobile suite (test/mobile/**); see config/vitest.ci.config.ts.
# Safe in CI: TmuxManager no-ops all shell commands under VITEST (test/setup.ts).
run: npm run test:ci
- name: Run xterm-zerolag-input package tests
# Layers 1-3 of the predictive-echo suites (unit laws, fixture replay,
# seeded fuzz): deterministic, no browser, no live server. Depends on
# the ROOT `npm ci` above — workspaces hoist the package's vitest into
# the root node_modules; do not add a separate install here.
run: npx vitest run
working-directory: packages/xterm-zerolag-input
# Note: The browser-driven mobile suite (test/mobile/**) is excluded from CI —
# it needs a live server + chromium + environment-specific PNG baselines.
# Run it locally/manually. All other tests run via the `test` job above.
-74
View File
@@ -1,74 +0,0 @@
name: Release
on:
push:
branches:
- master
concurrency: ${{ github.workflow }}-${{ github.ref }}
jobs:
release:
name: Release
runs-on: ubuntu-latest
permissions:
contents: write
pull-requests: write
steps:
- name: Checkout repo
uses: actions/checkout@v6
- name: Setup Node.js
uses: actions/setup-node@v6
with:
node-version: 22
cache: npm
registry-url: https://registry.npmjs.org
- name: Install dependencies
run: npm ci
- name: Build
run: npm run build
- name: Create release PR or publish
id: changesets
uses: changesets/action@v1
with:
publish: npm run release
version: npm run version-packages
title: "chore: version packages"
commit: "chore: version packages"
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Rename release tag to codeman
if: steps.changesets.outputs.published == 'true'
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
VERSION=$(node -p "require('./package.json').version")
OLD_TAG="aicodeman@${VERSION}"
NEW_TAG="codeman@${VERSION}"
# Update the GitHub release BEFORE deleting the old tag.
# make_latest pins the "Latest" badge to the Codeman release. This repo
# publishes TWO packages (aicodeman + xterm-zerolag-input), changesets
# creates a GitHub release for each, and GitHub awards "Latest" to
# whichever was published LAST. That is a race: 1.9.2 kept the badge,
# 1.9.4 lost it to xterm-zerolag-input@0.1.7 by two seconds. All package
# releases already exist by the time this step runs, so setting it here
# is deterministic.
RELEASE_ID=$(gh release view "$OLD_TAG" --json databaseId -q .databaseId 2>/dev/null || true)
if [ -n "$RELEASE_ID" ]; then
gh api -X PATCH "repos/${{ github.repository }}/releases/${RELEASE_ID}" \
-f tag_name="$NEW_TAG" \
-f name="$NEW_TAG" \
-f make_latest=true
fi
# Retag
git tag "$NEW_TAG" "$OLD_TAG" 2>/dev/null || true
git tag -d "$OLD_TAG" 2>/dev/null || true
git push origin "$NEW_TAG" ":refs/tags/$OLD_TAG" 2>/dev/null || true
-106
View File
@@ -1,106 +0,0 @@
# Claude Code local files
.agents/
skills-lock.json
# Written by install.sh into end-user clones when setup finishes
.install-complete
# Dependencies
node_modules/
# Build output
dist/
# Vendored frontend deps (generated from node_modules by postinstall/build)
src/web/public/vendor/
# Test coverage
coverage/
# E2E test screenshots (keep baselines, ignore current/diffs)
test/e2e/screenshots/current/
test/e2e/screenshots/diffs/
# Mobile visual regression failure artifacts
test/mobile/snapshots/*.actual.png
test/mobile/snapshots/*.diff.png
# Logs
*.log
npm-debug.log*
# OS files
.DS_Store
Thumbs.db
# Editor directories
.idea/
.vscode/
*.swp
*.swo
*~
# Environment files
.env
.env.local
.env.*.local
# State files (local to each machine)
.claude/ralph-loop.local.md
# Temporary files
*.tmp
*.temp
# Generated output
out/
screenshots-echo-diag/
screenshots-readme/
screenshots-readme-real/
screenshots-real/
scripts/remotion/out/
# Local UI/README capture scratch (screenshot runs, design mockups). Not build
# output, but never meant for git — an unqualified `git add -A` during a COM has
# swept dirs like these into a release before.
design-explorations/
# Artifacts that should not be tracked
test-results/
tmp/
# Machine-local working files (never meant for git). ANCHORED so only the root
# dir matches.
/pr/
# Root `public` (a symlink to scripts/remotion/public — local artifact). ANCHORED
# with a leading slash so it does NOT also match src/web/public (a bare `public`
# would swallow the whole web UI source dir and silently un-stage any new asset
# added there). No trailing slash so it still matches the symlink, not just dirs.
/public
# Opt-in gesture overlay runtime assets: large MediaPipe wasm + model (~27 MB)
# fetched at build/install by scripts/fetch-gesture-assets.mjs, kept out of git.
# (The gesture bundle itself, gesture-codeman.js, IS tracked — built from
# packages/gesture-control source by `npm run build:gesture`.)
src/web/public/gesture/wasm/
src/web/public/gesture/*.task
# Gesture-control workspace package build outputs (source is tracked; the
# Codeman bundle is emitted to src/web/public/gesture/gesture-codeman.js instead).
packages/gesture-control/dist/
packages/gesture-control/dist-codeman/
packages/gesture-control/.vite/
# Claude Code plan tracking
plan.json
# Unfinished TUI (local development only)
src/tui/
.claude/
media-assets/
commands
todo.md
@fix_plan.md
readme-preview.mjs
# Uploaded images land here under each session working dir (runtime artifact)
.claude-images/
-1
View File
@@ -1 +0,0 @@
include=dev
-1
View File
@@ -1 +0,0 @@
22
-31
View File
@@ -1,31 +0,0 @@
dist/
coverage/
node_modules/
src/web/public/vendor/
src/web/public/gesture/
src/web/public/app.js
src/web/public/styles.css
src/web/public/mobile.css
src/web/public/index.html
# Hand-formatted public JS modules (never prettier-enforced; the new
# check-public-assets.mjs still validates NUL bytes + JS syntax on these).
src/web/public/constants.js
src/web/public/image-input.js
src/web/public/input-cjk.js
src/web/public/keyboard-accessory.js
src/web/public/notification-manager.js
src/web/public/orchestrator-panel.js
src/web/public/panels-ui.js
src/web/public/ralph-panel.js
src/web/public/ralph-wizard.js
src/web/public/respawn-ui.js
src/web/public/session-ui.js
src/web/public/settings-ui.js
src/web/public/sw.js
src/web/public/terminal-ui.js
src/web/public/voice-input.js
src/web/public/upload.html
scripts/remotion/
# Hand-maintained; Prettier escapes underscores in glob paths and corrupts paragraphs.
CLAUDE.md
-16
View File
@@ -1,16 +0,0 @@
# Repository Guidelines
Canonical agent/contributor guidance for this repository lives in [CLAUDE.md](CLAUDE.md) —
project structure, build/test/lint commands, code style, testing safety rules
(never run the full suite inside a managed tmux session), security notes, and
the deployment workflow are all maintained there. Please read it before making
changes, and keep it the single source of truth rather than duplicating
sections here.
Quick pointers:
- Type check: `tsc --noEmit` · Lint: `npm run lint` · Format: `npm run format:check`
- Targeted tests only: `npm test -- test/<file>.test.ts` (bare `npm test` is unsafe in managed sessions)
- Route tests use `app.inject()`; new tests needing ports must pick a unique `const PORT =`
- Branch off `master` for all work; Conventional Commit-style messages (`fix(mobile): ...`)
- Never commit secrets or local state from `~/.codeman/`
-2320
View File
File diff suppressed because it is too large Load Diff
-402
View File
@@ -1,402 +0,0 @@
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
> Deep implementation detail lives in [`docs/architecture-invariants.md`](docs/architecture-invariants.md). This file holds the rules that prevent mistakes; that file holds the mechanisms, file inventories, and the history behind each rule. Pointers below are written as `→ architecture-invariants#anchor`. When the goal is raw throughput, [`docs/SPEEDRUN.md`](docs/SPEEDRUN.md) is the fast-execution protocol (it removes ceremony, never the safety rules here).
>
> **This file is in `.prettierignore` on purpose.** Prettier's markdown printer escapes underscores inside the glob-heavy paths used throughout (`agent-*.jsonl` became `agent-\_.jsonl`, collapsing backtick spans and corrupting a whole paragraph). Do not remove the ignore entry, and do not run `prettier --write` on it.
>
> **Repo root is kept short on purpose** (the README sits below the file listing on GitHub). Config lives in `config/` (`eslint.config.js`, `knip.json`, the vitest configs), Prettier's config is the `"prettier"` key in `package.json`, and `SECURITY.md` is under `.github/`. Root-only files are the ones tools genuinely require there: `CLAUDE.md` + `AGENTS.md` (loaded from the root by Claude Code / Codex), `CHANGELOG.md` (changesets writes it next to `package.json`), `tsconfig.json`, `.editorconfig`, `.nvmrc`/`.npmrc`, `.prettierignore` (resolved relative to cwd), `LICENSE` (GitHub detection) and `install.sh` (its raw URL is the published install one-liner). Don't relocate those.
## Quick Reference
| Task | Command |
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Dev server | `npm run dev` (or `npx tsx src/index.ts web`) |
| Type check | `npm run typecheck` (= `tsc --noEmit`) |
| Lint | `npm run lint` (fix: `npm run lint:fix`) |
| Format | `npm run format` (check: `npm run format:check`) |
| Single test | `npm test -- test/<file>.test.ts` (or `npx vitest run --config config/vitest.config.ts test/<file>.test.ts`) — ⚠ **never** run bare `npm test`, see Testing section |
| Build | `npm run build` (esbuild via `scripts/build.mjs`, NOT tsc — `tsc --noEmit` is type-check only) |
| Production | `npm run build && systemctl --user restart codeman-web` |
## CRITICAL: Session Safety
**You may be running inside a Codeman-managed tmux session.** Before killing ANY tmux or Claude process:
1. Check: `echo $CODEMAN_MUX` - if `1`, you're in a managed session
2. **NEVER** run `tmux kill-session`, `pkill tmux`, or `pkill claude` without confirming
3. Use the web UI or `./scripts/tmux-manager.sh` instead of direct kill commands
**The working tree is shared with other agent sessions.** Several Codeman sessions run against THIS one checkout, so another session can `git checkout` a different branch, or leave half-finished untracked files, while you are mid-task.
- **Always `git branch --show-current` immediately before committing.** Observed 2026-07-27: another session ran `git checkout -b feat/web-tabs`, a commit silently landed there instead of master, and the follow-up `git push origin master` cheerfully reported "Everything up-to-date".
- To land a commit on master **without** switching branches (which would yank the tree out from under the other session): `git push origin HEAD:master` then `git branch -f master HEAD`. Never `git checkout master` to "fix" it.
- **Never `git add -A`/`git add .`** — stage explicit paths. A sweep will pick up another session's WIP.
- Another session's broken WIP can block `npm run build`, since `tsc` is the first step and the build gates on it. That is not your bug to fix. ⚠️ `tsc` still EMITS on type errors, so a failed `npm run build` leaves a rebuilt `dist/index.js` compiled from their tree; check what it pulled in before restarting the service. To deploy frontend-only changes past a blocked `tsc`, run the asset stage of `scripts/build.mjs` (everything after the `tsc`/`chmod` lines is independent of it).
## CRITICAL: Always Test Before Deploying
**NEVER COM without verifying your changes actually work.** For every fix:
1. **Backend changes**: Hit the API endpoint with `curl` and verify the response
2. **Frontend changes**: Use Playwright to load the page and assert the UI renders correctly. Use `waitUntil: 'domcontentloaded'` (not `networkidle` — SSE keeps the connection open). Wait 3-4s for polling/async data to populate, then check element visibility, text content, and CSS values
3. **Only after verification passes**, proceed with COM
The production server caches static files for 1 year, `immutable` (`maxAge: '1y'` in `server.ts`). To avoid stale frontend after a deploy, `renderIndexHtml` runs `cacheBustAssets(html)` — it appends `?v=<mtime>` to **every same-origin `.js`/`.css`** reference (mtime memoized ~1s so a burst of renders is cheap; external/already-versioned/missing refs untouched). Because `index.html` is served `no-cache`, a **normal reload now picks up edited modules/styles — no hard refresh needed** (the gesture bundle is injected separately with its own `?v=`). If you add an asset referenced by an _absolute_ URL or from JS rather than a `<script>/<link>` tag, it won't be auto-busted. ⚠️ **`index.html` itself is the exception: it is read ONCE into `indexHtmlTemplate` in the `WebServer` constructor**, so editing markup in dev needs a server restart (edited `.js`/`.css` do not) — otherwise you debug a "CSS class that doesn't apply" that is really an element still missing from the served HTML.
## COM Shorthand (Deployment)
Uses [Semantic Versioning](https://semver.org/) (`MAJOR.MINOR.PATCH`) via `@changesets/cli`. What SemVer actually covers (the CLI, documented env vars, **and the HTTP/SSE API under `/api/v1`**: endpoint paths, response envelope, `errorCode` values and SSE event names are public/stable; on-disk state, internal TS modules, and experimental features are internal/unstable) is defined in `docs/versioning-policy.md`. Third-party integration surfaces are documented in `docs/extending-codeman.md`. Security reporting + known limitations live in `.github/SECURITY.md`.
When user says "COM":
1. **Determine bump type**: `COM` = patch (default), `COM minor` = minor, `COM major` = major
2. **Create a changeset file** (no interactive prompts). Write a `.md` file in `.changeset/` with a random filename:
```bash
cat > .changeset/$(openssl rand -hex 4).md << 'CHANGESET'
---
"aicodeman": patch
---
Detailed description of ALL changes since last release (not just the most recent commit — review full git log since last version tag)
CHANGESET
```
Replace `patch` with `minor` or `major` as needed. Include `"xterm-zerolag-input": patch` on a separate line if that package changed too.
3. **Consume the changeset**: `npm run version-packages` (auto-bumps `package.json` files, updates `CHANGELOG.md`, runs `npm install --package-lock-only`, and verifies lockfile sync via `scripts/check-lockfile-sync.mjs` — all in one command; never hand-edit `CHANGELOG.md` or `package-lock.json` versions)
4. **Sync CLAUDE.md version**: Update the `**Version**` line below to match the new version from `package.json`
5. **Commit and deploy**: verify the branch first (`git branch --show-current`), then stage EXPLICIT paths — never `git add -A`, which has swept another session's WIP into a release. `git status --short` and account for every line before committing:
`git add <paths> && git commit -m "chore: version packages" && git push && npm run build && systemctl --user restart codeman-web`
6. **Wait for CI**: after `git push`, TWO workflows fire per master push — `CI` and `Release` (the npm publish + GitHub release). List both runs for the pushed commit with `gh run list --commit $(git rev-parse HEAD) --json databaseId,workflowName` and watch EACH with `gh run watch <id> --exit-status`. Confirm both pass before considering the release done (`gh run list -L 1` returns only one of the two).
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 1.18.3 (must match `package.json`)
## Project Overview
Codeman is a Claude Code session manager with web interface and autonomous Ralph Loop. Spawns Claude CLI via PTY, streams via SSE, supports respawn cycling for 24+ hour autonomous runs.
**Tech Stack**: TypeScript (ES2022/NodeNext, strict mode), Node.js, Fastify, node-pty, xterm.js. Supports Claude Code, OpenCode, Codex (OpenAI), Gemini (Google, enterprise-only since Google's June 2026 consumer cutover), Antigravity (`agy`, Google) and Pi (pi.dev) CLIs via pluggable CLI resolvers (`SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity' | 'pi'`).
**TypeScript Strictness** (see `tsconfig.json`): `noUnusedLocals`, `noUnusedParameters`, `noImplicitReturns`, `noImplicitOverride`, `noFallthroughCasesInSwitch`, `allowUnreachableCode: false`, `allowUnusedLabels: false`.
**Requirements**: Node.js 22+, Claude CLI, tmux
**Git**: Main branch is `master`. SSH session chooser: `sc` (interactive), `sc 2` (quick attach), `sc -l` (list).
## Additional Commands
`npm run dev` = dev server. Default port: `3000` (override with `--port` or the `CODEMAN_PORT` env var). To run this beta isolated alongside a prod Codeman, use `scripts/run-beta.sh` (sets `CODEMAN_INSTANCE=beta` + `CODEMAN_PORT=5000`). Commands not in Quick Reference:
| Task | Command |
| ------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Dev with TLS | `npx tsx src/index.ts web --https` |
| Override window title hostname | `npx tsx src/index.ts web --title-hostname <name>` (default: `os.hostname()` — `codeman:<name>` is used for tab title, title-flash, and OS desktop notification prefix) |
| Bind a non-loopback host | `npx tsx src/index.ts web --host 0.0.0.0` (or `-H`; env `CODEMAN_HOST`; default `127.0.0.1`). Without `CODEMAN_PASSWORD` it **starts but warns loudly** — see Common Gotchas + `docs/security-architecture.md` |
| Continuous typecheck | `tsc --noEmit --watch` |
| Watch-mode test | `npm run test:watch -- test/<file>.test.ts` (always pass a file — bare watch includes the browser suites) |
| Test coverage | `npm run test:coverage` |
| Dead-code sweep | `npm run knip` (config in `config/knip.json`, passed via `--config`) |
| Rebuild gesture overlay | `npm run build:gesture` (esbuild `packages/gesture-control/src/codeman/entry.ts` → `src/web/public/gesture/gesture-codeman.js`; commit the result) |
| Build the docker agent image | `node scripts/build-agent-image.mjs --no-cache` (builds `codeman/agent:base` from `docker/agent.Dockerfile`; prerequisite for Docker cases; `--engine`/`--image`). ⚠ **Always `--no-cache`** — a plain rebuild re-uses the cached `npm install -g` layer and silently keeps the CLIs frozen at their original versions, which once shipped a BROKEN codex while reporting success. See `docs/docker-cases.md` |
| Gesture playground | `npm run dev` **in** `packages/gesture-control/` (standalone vite demo, fake tabs) |
| Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) |
| Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) |
| CI-equivalent test sweep | `npm run test:ci` (full suite minus browser/perf — see Testing) |
| Production start | `npm run start` |
| Production logs | `journalctl --user -u codeman-web -f` |
| Detached server | `codeman web -d` (`--status`, `--stop`; pidfile+log at `dataPath('web.pid'/'web.log')`). ⚠ Refuses to start a 2nd server on one data dir — see Instance isolation |
| Install/remove the service | `codeman service install` / `status` / `uninstall` (systemd user unit on Linux, LaunchAgent on macOS; names from `config/service-names.ts`) |
**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 3 Playwright tests). Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing).
**Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`) — config lives in the **`"prettier"` key of `package.json`**, not a `.prettierrc` (keeps the repo root short; editors read it natively). `.prettierignore` stays at the root because Prettier resolves it relative to cwd. ESLint flat config (`config/eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `scripts/remotion/**`.
**Prettier scope is deliberately narrow.** `npm run format` globs only `src/**/*.ts` and `src/web/public/**`, and `.prettierignore` then exempts most of `src/web/public/*.js` (app.js, styles.css, index.html, and 14 hand-formatted modules) plus `CLAUDE.md`. Those files are hand-formatted by design; `npm run check:public-assets` and `check:frontend-syntax` are what guard them (NUL bytes + JS syntax), not Prettier. Do not "fix" a file by adding it back to Prettier's scope.
## Common Gotchas
- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink. ⚠️ **Input must END with `\r` or Enter is never sent**: `sendInput()` only issues `send-keys Enter` when the payload contains a carriage return, a `\r`-less `POST /api/sessions/:id/input` still succeeds (send-and-wait even reports `delivered:true`) while the text sits unsubmitted on the composer, and any `wait` burns its whole timeout on a turn that never started. Embedded newlines are stripped, not rejected, so `"echo A\necho B\r"` runs the joined `echo Aecho B`
- **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks
- **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly. Both `aicodeman` and `codeman` bin aliases are installed (`package.json` `bin`)
- **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically)
- **`envOverrides` flow `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` / `ANTIGRAVITY_*` / `PI_*` env vars, plus exact-key `CLAUDE_CONFIG_DIR`** — Set via `POST /api/sessions { envOverrides }`, stored on `Session._envOverrides`, exported by `tmux-manager.buildEnvExports()` at spawn time, persisted in `SessionState.envOverrides`. **Do NOT** write these to `<case>/.claude/settings.local.json` — that's the old path and creates UI/disk drift. (`GOOGLE_*` is the deliberately-broad Vertex-AI namespace for Gemini — see Multi-CLI prefix discipline.) `CLAUDE_CONFIG_DIR` (#255, exact match via `ALLOWED_ENV_KEYS` in `schemas.ts`) points a session at a separate Claude account/config dir for per-client subscriptions; it persists to state.json (a path, not a secret; losing it on restart would silently switch accounts). ⚠️ A relocated config dir writes transcripts outside `~/.claude/projects`, so the response viewer, subagent windows, ultracode panel and Read My Mind capture go blind for that session unless the user symlinks `projects` back into the shared tree (`ln -s ~/.claude/projects <configDir>/projects`). → [architecture-invariants#per-session-env-overrides-exact-key-allowlist-and-claude_config_dir](docs/architecture-invariants.md#per-session-env-overrides-exact-key-allowlist-and-claude_config_dir)
- **Effort is NOT an env var** — never carry effort as `CLAUDE_CODE_EFFORT_LEVEL`: the env var hard-locks effort and blocks in-session `/effort` switching (incl. ultracode). It flows as the dedicated `effort` payload field → `Session._effort` → `claude --effort <level>` for regular levels incl. `max` (the settings `effortLevel` key is `enum(["low","medium","high","xhigh"]).catch(undefined)` — `max` gets SILENTLY dropped there), or `claude --settings '{"ultracode":true}'` for ultracode (rejected by `--effort`). Both are soft defaults the user can override anytime. Legacy env-var entries are auto-migrated by the Session constructor and unset from tmux sessions in `applyEnvOverrides()`. See `buildEffortCliArgs()` in `session-cli-builder.ts`, tests in `test/effort-injection.test.ts`
- **Model choice flows via `settings.local.json`, NOT `--model` or env** — the App Settings **Claude Model** picker (`claudeModel` in `settings.json`) is read by `session-ui.js` at session create (wins over the legacy 1M-Opus toggles `opusContext1m`/`opusContext1mEnabled`), sent as the `modelOverride` payload field, and `updateCaseModel()` (`hooks-config.ts`) writes/deletes the `model` key in `<case>/.claude/settings.local.json`. This is the intended exception to the envOverrides rule above: model legitimately lives in `settings.local.json` (a soft default — in-session `/model` still works); env vars do not
- **Multi-CLI prefix discipline** — env-var prefix is CLI-specific (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*` vs `ANTIGRAVITY_*` vs `PI_*`) and the `ALLOWED_ENV_PREFIXES` allowlist in `schemas.ts` enforces this; non-prefix exceptions are exact keys in `ALLOWED_ENV_KEYS` (currently only `CLAUDE_CONFIG_DIR`), never a widened prefix. Gemini additionally allowlists the **broad `GOOGLE_*`** namespace (intentional: Vertex AI auth needs `GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI`; it is the loosest allowlist entry, affecting only the user's own spawned CLI). When adding a setting, decide which CLI(s) it applies to and gate the env export accordingly. Never blanket-forward all prefixes. ⚠️ Pi is the case that proves the rule: its ~34 provider keys (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `HF_TOKEN`, …) share NO prefix, and the allowlist is one GLOBAL list applied by a refine with no mode context, so admitting them for pi would widen it for every mode at once — they stay out, and pi users authenticate via `/login` or the server process's own env. Resolver design pattern: `docs/opencode-integration.md`, `docs/pi-integration.md`
- **Zod `.optional()` rejects `null`** — accepts `undefined` only. When the frontend builds a request body with `JSON.stringify`, an explicit `null` field is preserved on the wire and fails validation with `INVALID_INPUT`. Convert `null` → `undefined` before stringifying (e.g. `field: value ?? undefined`), or declare the schema `.nullish()`. This has caused real shipped bugs twice
- **`xterm-zerolag-input` is single-source** — BOTH echo addons live ONLY in `packages/xterm-zerolag-input/src/`, bundled into TWO **gitignored** vendor files: `vendor/xterm-zerolag-input.js` (buffer overlay, entry `zerolag-input-addon.ts`) and `vendor/xterm-predictive-echo.js` (codex write-through, entry `predictive-echo-addon.ts`) — dev by `scripts/postinstall.js`, prod by `scripts/build.mjs`. `app.js`/terminal-ui.js only **consume** them via `new LocalEchoOverlay(terminal)` / `new PredictiveEchoOverlay(terminal)`; there is no inline copy. So: change the package source, then rerun the bundle step (`npm install` for dev, `npm run build` for prod). **Never hand-edit `app.js` for overlay behavior, and never commit the gitignored vendor bundles.** Always test on mobile after touching it. → [architecture-invariants#xterm-zerolag-input-is-single-source](docs/architecture-invariants.md#xterm-zerolag-input-is-single-source), `docs/local-echo-overlay-plan.md`
- **Default bind is loopback-only; non-loopback without a password starts but warns** — the server defaults to `--host 127.0.0.1`. Binding non-loopback (`--host`/`-H`/`CODEMAN_HOST`) without `CODEMAN_PASSWORD` starts anyway but prints a loud warning; `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` acknowledges it. ⚠️ The production systemd unit passes no `--host`, so prod binds **localhost only**: reach it via `tailscale serve`/tunnel to `127.0.0.1`. A loopback bind is reachable through a same-host tunnel but NOT by a browser hitting the box's LAN IP. `install.sh` is separate and prompts for the binding (defaulting to LAN + a password), and preserves the existing binding on re-runs. → [architecture-invariants#default-bind-and-the-non-loopback-warning-path](docs/architecture-invariants.md#default-bind-and-the-non-loopback-warning-path), `docs/security-architecture.md`
- **Instance isolation / multi-instance attach danger** — the data dir (`~/.codeman`) and tmux socket (`tmux -L codeman`) are PROCESS-WIDE and shared by every Codeman on the machine, derived from `CODEMAN_INSTANCE` via `src/config/instance.ts`. ⚠️ A 2nd instance on the SAME socket **discovers and attaches PTYs to the first instance's live sessions**, resizing and mutating them. `$HOME` isolation is NOT enough because tmux is system-global. To run two instances, give each a distinct `CODEMAN_INSTANCE` (scopes dir + socket together), or set `CODEMAN_TMUX_SOCKET` + `CODEMAN_DATA_DIR` individually; `scripts/run-beta.sh` does this for a beta alongside prod. **Any new `~/.codeman/...` path MUST go through `dataPath()`**, never `join(homedir(), '.codeman', …)`. → [architecture-invariants#instance-isolation-and-the-multi-instance-attach-danger](docs/architecture-invariants.md#instance-isolation-and-the-multi-instance-attach-danger)
- **node-pty's macOS `spawn-helper` ships without `+x`** (issues #6, #204): `node-pty@1.1.0` publishes `prebuilds/darwin-<arch>/spawn-helper` as mode 0644, and macOS launches every PTY through it, so a stock macOS install fails every session start with `Error: posix_spawnp failed.` **Linux can never reproduce it**: `spawn-helper` is an `OS=="mac"` gyp target and node-pty ships no Linux prebuild, so node-gyp always emits an executable helper there. ⚠️ Look in **`prebuilds/<platform>-<arch>/`**, not just `build/Release/`, which does not exist on macOS. Repair is a chmod, never a mandatory rebuild (that would require Xcode CLI tools and deletes `prebuilds/` before compiling): `npm run fix:node-pty` chmods every helper then proves it by really opening a PTY. `spawnPtyWithHelperRepair()` (`utils/node-pty-repair.ts`) wraps every `pty.spawn()` in `session.ts` and self-heals a broken install on the first failure. → [architecture-invariants#node-ptys-macos-spawn-helper-must-be-executable](docs/architecture-invariants.md#node-ptys-macos-spawn-helper-must-be-executable)
- **Headless screenshots: `deviceScaleFactor` MUST be 1, and write unique filenames** — under DSF=2 xterm's WebGL renderer draws glyphs at ~2× nominal size while still *reporting* nominal cell dims, so only the pixels reveal it and only the terminal font looks wrong. And overwriting a fixed output path leaves OS image viewers showing the old render, which reads as "the fix didn't work"; `scripts/capture-real-overview.mjs` mints a timestamped filename per run. Seed the per-device `localStorage` keys (`codeman:skin`, `codeman-font-size`, `codeman-app-settings`) so the capture matches a real device. → [architecture-invariants#headless-screenshot-capture](docs/architecture-invariants.md#headless-screenshot-capture)
**Import conventions**: Utils from `./utils`, types from `./types` (barrel), config from specific `./config/*` files.
## Architecture
### Core Files (by domain)
| Domain | Key files | Notes |
| ---------------- | -------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| **Entry** | `src/index.ts`, `src/cli.ts`, `daemon-control`, `service-installer`, `config/service-names` | The last three back `web -d` / `service install` |
| **Session** | `src/session.ts` ★, `session-manager`, `session-auto-ops`, `session-cli-builder`, `session-task-cache`, `session-order` (pure), `session-pty-exit-breaker`, `usage-limit-patterns`, `usage-telemetry`; `src/services/unified-session-service.ts` | Pure/unit-tested helpers are split out of `session.ts` on purpose |
| **Mux** | `src/mux-interface.ts`, `src/mux-factory.ts`, `src/tmux-manager.ts` ★ | |
| **Respawn** | `src/respawn-controller.ts` ★ + 4 helpers (`-adaptive-timing`, `-health`, `-metrics`, `-patterns`) | Read `docs/respawn-state-machine.md` first |
| **Ralph** | `src/ralph-tracker.ts` ★, `src/ralph-loop.ts` + 5 helpers (`-config`, `-fix-plan-watcher`, `-plan-tracker`, `-stall-detector`, `-status-parser`) | Read `docs/ralph-wiggum-guide.md` first |
| **Orchestrator** | `src/orchestrator-loop.ts`, `-planner`, `-verifier` | Read `docs/orchestrator-loop-architecture.md` first |
| **Cron** | `src/cron/cron-service.ts`, `cron-time.ts` (pure next-run math), `cron-input.ts` | Read `docs/cron-discovery.md` first. Distinct from legacy `ScheduledRun` (`/api/scheduled`) |
| **Agents** | `src/subagent-watcher.ts` ★, `team-watcher`, `bash-tool-parser`, `transcript-watcher`, `workflow-run-watcher` | `workflow-run-watcher` is STANDALONE and never touches `subagent-watcher` |
| **AI** | `src/ai-checker-base.ts`, `ai-idle-checker.ts`, `ai-plan-checker.ts` | |
| **Tasks** | `src/task.ts`, `task-queue.ts`, `task-tracker.ts` | |
| **State** | `src/state-store.ts`, `run-summary.ts`, `session-lifecycle-log.ts`, `intent-store.ts` | |
| **Infra** | `src/hooks-config.ts`, `push-store`, `tunnel-manager`, `image-watcher`, `file-stream-manager`, `remote-hosts` + `remote-reconnect` (pure), `docker-hosts` + `docker-export` | Remote/docker case overlays; see Key Patterns |
| **Web tabs** | `src/webview-store.ts`, `webview-capabilities.ts`, `src/web/webview-proxy.ts` (pure), `src/web/routes/webview-routes.ts` | Dashboard URLs as tabs; NOT a SessionMode |
| **Search** | `src/search-service.ts` | Pure in-memory core for `GET /api/search` |
| **Attachments** | `src/attachment-registry.ts`, `attachment-magic`, `generated-artifact-attachments`, `session-attachment-history`, `document-preview-cache`, `document-thumbnailer`, `document-conversion-limiter`, `config/attachment-guard` | See Key Patterns |
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/` (`claude-md.ts` + `case-template.md`) | `templates/` holds the CLAUDE.md scaffold generated into new cases |
| **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (24 modules + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | |
| **Frontend** | `src/web/public/app.js` (~5K lines, core) + 29 modules + `sw.js` | See Frontend section for the load order, which is authoritative |
| **Types** | `src/types/index.ts` (barrel) → 22 domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts |
★ = Large, central file (>50KB) — read its `@fileoverview` first. All files have `@fileoverview` JSDoc — read that before diving in. Discovery aid: `grep -l '@fileoverview' src/web/routes/*.ts` lists all route modules; same grep works for `src/types/`, `src/web/public/*.js`.
**Local packages**: `packages/xterm-zerolag-input/` (local echo overlay, single-source, see Gotchas). `packages/gesture-control/` (`codeman-gesture-control`, hand-tracking overlay source, built via `npm run build:gesture`).
**Config**: `src/config/` — 20 files, no barrel (`index.ts`) exists; import from the specific file.
**Utilities**: `src/utils/` — re-exported via index. Key: `CleanupManager`, `LRUMap` (⚠ NOT in the barrel — import from `./utils/lru-map.js` directly), `StaleExpirationMap`, `BufferAccumulator`, `stripAnsi`, `Debouncer`, `KeyedDebouncer`. Also: `claude-cli-resolver`/`opencode-cli-resolver`/`codex-cli-resolver`/`gemini-cli-resolver`/`antigravity-cli-resolver`/`pi-cli-resolver` (CLI path resolution; ⚠ `pi-cli-resolver` additionally version-probes the binary, since `pi` is a generic name), `string-similarity` (fuzzy matching), `regex-patterns` (ANSI/token/spinner patterns), `assertNever` (exhaustive checks), `token-validation` (auth tokens), `nice-wrapper` (process priority).
### Data Flow
1. Session spawns `claude --dangerously-skip-permissions` via node-pty
2. PTY output buffered, ANSI stripped, parsed for JSON messages
3. WebServer broadcasts to SSE clients at `/api/events`
4. State persists to `~/.codeman/state.json` via StateStore
### Key Patterns
**Input**: `session.writeViaMux()` for programmatic/curl input via tmux `send-keys -l` + `send-keys Enter`, single-line only. Interactive **browser** input goes through a durable **exactly-once** layer: a stable `clientId` + monotonic per-session `seq` persisted to localStorage until the server ACKs, so a dropped link cannot lose or double-deliver a prompt. `ws-connection-registry.ts` supersedes only same-TAB reconnects, so two tabs on one session coexist. → [architecture-invariants#input-delivery-and-ws-resilience](docs/architecture-invariants.md#input-delivery-and-ws-resilience)
**Agent wait primitives**: bounded long-polls so an agent driving Codeman from a shell can block instead of poll: `GET /api/sessions/:id/wait` (lifecycle signal), `GET /api/sessions/:id/wait-output` (literal substring, **never** regex) and `wait`/`waitTimeout` on `POST /api/sessions/:id/input`. Registry in `session-wait-registry.ts` (pure, no `Session` reference), bounds in `config/agent-wait.ts`. ⚠️ **A timeout is a 200** (`wait.timedOut`), never an error, so callers loop over short waits. ⚠️ `stop`/`blocked` come from Claude Code hooks and therefore fire for **`claude` mode ONLY** (`shell` installs none either); asking for one explicitly on another mode is a 400, the default set silently drops them. ⚠️ Send-and-wait registers the waiter BEFORE the write (a separate POST-then-wait races and reports the PREVIOUS turn), and both teardown paths must `notifySignal('exit')` BEFORE `cancelAll()`. ⚠️ Client-hangup abort listens on **`reply.raw`** guarded by `writableFinished`: on `req.raw`, `close` fires when the request BODY ends, which on a POST killed every send-and-wait instantly and no `app.inject()` test could see it. ⚠️ Worker liveness cannot come from `session.pid` — for a tmux session that is the local attach client, which outlives a worker dying inside its pane — so it is probed at the mux layer (`isPaneDead`, ~750 ms cache) on blocking waits only, never on the input hot path. ⚠️ Signals are edge-triggered with no history: one that fires with no waiter registered is unobservable afterwards, so gather fan-outs with send-and-wait or latched `wait-output` markers, never fire-and-forget-then-sequential-signal-waits. The primitives are packaged as the **`skills/codeman` agent skill**: installable via `codeman skill install [--case <name>]` / `skill uninstall`, or auto-injected into a case's `.claude/skills/` on Claude session create behind `agentSkillEnabled` (SYNCED, default OFF). Injection is ADD-ONLY at create, marker-owned (`applyAgentSkill` in `hooks-config.ts` never touches an unmarked user copy) and refuses symlinks (this repo's own `.claude/skills/codeman` is a symlink to the source, which the injector must never write through). ⚠️ Claude Code loads a same-named USER-LEVEL skill (`~/.claude/skills/codeman`, written once by `codeman skill install` with no `--case`) over the per-case copy, and nothing used to refresh it: a stale Aug-9 user copy shadowed every fresh injection (2026-08-14: agents ran the old recipes, spawned workers serially and lost their lineage arcs), so session create now also refreshes a marker-owned user copy (`refreshUserAgentSkill`; refresh-only, never installs, foreign/symlink refused). Session create additionally pre-seeds the skill's §0 preamble cache (`seedAgentSessionPreamble` → `${XDG_CACHE_HOME:-~/.cache}/codeman-agent-<id>.sh`, local claude sessions only), single-sourced from `skills/codeman/preamble.sh` and pinned byte-identical to SKILL.md's §0 heredoc by `test/agent-skill.test.ts`, so the skill's bootstrap is a two-line loader instead of a ~150-line paste the model types out (~47 s of generation, measured live). → [architecture-invariants#agent-wait-primitives](docs/architecture-invariants.md#agent-wait-primitives), `docs/api-reference.md`
**Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer (`❯`) about once a second all through a turn, so the old "saw a ❯, wait 2s → idle" rule flipped every working session to idle two seconds in (measured: a session mid-tool-call at 17 minutes reporting `status:"idle"`). Its working indicator is `✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`: the glyph animates through `· ✢ ✳ ∗ ✻ ✽`, the gerund is randomized, and the finished line (`✻ Cooked for 2m 49s`) carries the same glyph, so neither `SPINNER_PATTERN` (braille, not what current versions draw) nor a keyword list can see it. Matching the new line in the STREAM does not work either: tmux ships partial repaints, so the whole line reaches the PTY only every few tens of seconds. So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks the SCREEN via `capturePaneText()` + `CLAUDE_WORKING_LINE_PATTERN` before believing it; a sustained run of repaints (`session-activity.ts`, pure + unit tested) is what marks a turn as started, with the same screen probe vetoing keystroke echo. Idle now lands ~3-5s after a turn ends instead of 2s into one. Claude-mode only, since an external CLI has no `❯`, so nothing would ever arm the confirmation and the session would latch busy.
**Auto-resume on usage limit** (opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit, `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time and `SessionAutoOps` arms a timer for reset+2min, then sends Esc + `continue`. ⚠️ Respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected`), which is what prevents `/clear` from wiping the paused conversation. Claude-mode only. → [architecture-invariants#auto-resume-on-usage-limit](docs/architecture-invariants.md#auto-resume-on-usage-limit)
**Plan-usage chip** (statusLine telemetry, `showPlanUsageLimits`, per-device: desktop default **ON**, handhelds OFF via the mobile block in `getDefaultSettings()`): resolve it ONLY through `planUsageChipEnabled()` in settings-ui.js, which backs all three call sites (the App Settings checkbox, the chip's visibility, and the `statusLineTelemetry` flag on session create). A chip shown without telemetry renders `—` forever. Codeman injects its own `statusLine.command` exporter which POSTs Claude's `rate_limits` blob to `POST /api/status-telemetry`. The exporter is identified by a marker, so it only ever adds/updates/removes a statusLine that is **ours**, never a user's hand-authored one, and it prints the footer through so the in-terminal statusline is not blanked. Claude-mode only; distinct from auto-resume, which reacts to the limit *message* rather than showing live %. → [architecture-invariants#plan-usage-chip-statusline-telemetry](docs/architecture-invariants.md#plan-usage-chip-statusline-telemetry), `docs/usage-limits-display-plan.md`
**Orchestrator**: State machine that turns a user goal into a phased plan and drives it to completion: `idle → planning → approval → executing → verifying → (replanning) → completed/failed`. `OrchestratorLoop` (engine) delegates plan generation to `orchestrator-planner` and per-phase verification gates to `orchestrator-verifier`, executing phases via team agents/`task-queue`. State persists under the `orchestrator` key in `state.json`. Distinct from Ralph (single-session autonomous loop) — orchestrator coordinates multi-phase, multi-agent execution. See `docs/orchestrator-loop-architecture.md`.
**Cron (`CronJob`s)**: saved, named jobs on a recurring schedule (`once`/`interval`/`daily`/`weekly`) with per-job run history. ⚠️ **Distinct from the legacy `ScheduledRun`** (`/api/scheduled`, a run-now duration-bounded loop); the two never interact and keep separate `Scheduled*` / `Cron*` names. `CronService` **reuses the existing session layer** rather than rebuilding tmux logic. Next-run math is pure and unit-tested in `cron-time.ts` (server-local timezone). The schedule is advanced BEFORE launch so a slow launch cannot re-trigger. → [architecture-invariants#cron-jobs](docs/architecture-invariants.md#cron-jobs), `docs/cron-discovery.md`
**Remote sessions + remote SSH cases**: a case can point at a remote host. The agent runs inside a durable remote `tmux -L codeman-remote` (session name `codeman-ssh-<id>`, deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so an instance on the target host never adopts it), fronted by a LOCAL tmux pane running `ssh`. Attached (`owned:false`) sessions **detach, never kill** on tab close; owned ones propagate `kill-session`. A bounded-backoff watcher auto-reconnects dropped sessions (`remoteAutoReconnect`, default ON). ⚠️ **Command-injection surface: every ssh command line must flow through `buildSshConnectionArgs()`**, which `shellescape`s every user field. Never hand-build an ssh line elsewhere. ⚠️ Run flows must route remote cases through `POST /api/quick-start`, not `POST /api/sessions` (which stat-validates `workingDir` locally and has no `caseName`). → [architecture-invariants#remote-sessions-over-ssh](docs/architecture-invariants.md#remote-sessions-over-ssh), [#remote-ssh-cases](docs/architecture-invariants.md#remote-ssh-cases), `docs/remote-sessions.md`
**Docker cases**: a case can point at a **container**, with any of the five CLI backends running inside it. Like remote-SSH this is a **LOCATION OVERLAY on cases, never a sixth `SessionMode`**. Exactly one long-lived container **per case**, shared by all its sessions, so killing a session kills only that session's in-container tmux and **never** `docker stop` while siblings remain. The workspace is a real host dir bind-mounted at the **same absolute path**, which is what keeps file-routes/watchers on real host bytes and makes the in-container transcript projHash match the host. Credentials are **seeded** (RO mount, copied into the container once) rather than shared RW, so in-container CLIs never write refreshed tokens back to the host, and bind mounts are excluded from `docker commit` so exports stay secret-free. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** Config drift is detected via a label hash and a drifted launch is REFUSED rather than silently launched with stale config. ⚠️ On the loopback-only prod bind a container cannot reach 127.0.0.1, so in-container hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; otherwise idle detection falls back to output-based. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` (user guide), `docs/docker-cases-plan.md` (design)
**External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. All five **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. ⚠️ **Codex sessions use PREDICTIVE WRITE-THROUGH echo, never the buffer overlay** (`_localEchoPolicy` in `_updateLocalEchoState`, terminal-ui.js): codex's composer reacts per keystroke ("/" pops a live-filtering picker, arrows edit server-side state, the composer grows as it wraps), so buffer-until-Enter starved it into issues #218/#219/#220/#222 and stays disabled (`_localEchoEnabled` remains false for codex). Instead, `PredictiveEchoAddon` (separate `vendor/xterm-predictive-echo.js` bundle) paints each keystroke at the predicted cell while the wire path stays BYTE-IDENTICAL: the onData hook (`_predictHookOnData`) is a plain statement with no `return`, so control always falls through into the untouched send path — pinned by vm and E2E byte-identity tests. Predictions reconcile against the parsed buffer and only while the cursor sits on the measured composer row (`isCodexComposerRow`, `/^› /`). Codex also **drops keystrokes that share a PTY read with a bracketed paste**, so flushed text and the paste sequence must go out as separate delayed writes (mirroring the Enter branch's delayed `\r`). Tests: `test/local-echo-codex-gating.test.ts`, `test/codex-predictive-echo.test.ts` (E2E vs real codex), `packages/xterm-zerolag-input/test/codex-replay.test.ts`. ⚠️ **Pi is the opposite kind of CLI and needs the opposite instincts**: it has NO permission prompts and no sandbox, so there is no bypass flag to send and Codeman must not invent one; its privileged knob is the tri-state `approveProjectTrust` (`--approve`/`--no-approve`), which makes pi EXECUTE repo-local `.pi/extensions` TypeScript, so the multi-user clamp puts pi in the **materialize** branch (an absent config still yields `--no-approve` for a non-granted owner) and `--api-key` is never wired. Pi stays OUT of `isAltScreenStripMode()` (main-screen TUI, and its 0.84.0 fullscreen mode is runtime-switchable via `/settings`, where the alt screen is load-bearing), and lands on the `'buffer'` echo policy via the `_updateLocalEchoState` fallthrough. Pi's own tests: `test/pi-mode.test.ts`, `test/routes/external-cli-bypass-clamp.test.ts`; user guide `docs/pi-integration.md`. → [architecture-invariants#external-cli-modes-opencode-codex-gemini-antigravity-pi](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi)
**Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w<n>-<case>` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization)
**Session lineage lines** (tab → tab it spawned, `sessionLineageLines`, per-device, desktop default ON): a create request may name the session that spawned it, as a `parentSessionId` body field on `POST /api/sessions` / `POST /api/quick-start` or the `X-Codeman-Parent-Session` header (the agent skill sets that once on its shared curl invocation, so every spawn recipe carries it). `resolveParentSessionId()` (route-helpers.ts) **resolves rather than trusts** it: exact id, else a UNIQUE ≥8-char prefix (ids reach agents truncated), it must be a live session the caller can see AND carry the same owner, and **anything unresolvable is DROPPED, never a 400** — a cosmetic field must not be able to fail a worker spawn. It rides `toState()` into `session_created`, so there is no new SSE event. ⚠️ Rendering is an ADDITIONAL LAYER on the existing SVG pass (`_appendLineageConnectionLines` called at the tail of `_updateConnectionLinesImmediate()`, exactly like ultracode), sharing one batched read→write reflow and the `tab:<id>` rect cache; geometry is pure in `computeLineagePath()` (constants.js). ⚠️ **ONE shape, and the second one was the bug**: every pair (flat strip or wrapped) gets a U-bridge hanging below the strip, anchored on both tabs' BOTTOM edges. A wrapped strip used to get a parent-bottom → child-TOP bezier with a ~14px row gap to bend in, which drew a flat line hidden in the gap with siblings overprinting. ⚠️ The dip is a **mis-tuned-in-both-directions corridor** (44px cap = straight thread at strip-wide spans, #285; 104px cap + full row offset = ~106px over-bow into the terminal, 2026-08-15): it now hangs from the **STRIP's bottom edge** (fallback: lower tab bottom), capped at 64px, with NO per-row offsets stacked on top — the strip-bottom baseline is also what keeps a row-1 pair's arc from drawing through row 2's tab labels. Colors cycle per CHILD in first-seen order from `CodemanLineage.COLORS` (first entry empty = the skin-tuned `--session-blue`; the rest vivid fixed hexes), set inline as `--lineage-color` so styles.css keeps owning opacity/glow/dash. ⚠️ **Desktop only**: the overlay is `z-index: 999` and the desktop header is 100 (arcs paint over it, which is what lets them touch tab bottoms), but under 1024px mobile.css makes the header `fixed; z-index: 1200` and would bury them. ⚠️ Paths carry `data-agent-id="lineage:<childId>"` because that is what `_applyLineEntrances()` queries — that one attribute is what gives them the entrance animation and its negative-`animation-delay` resume across `svg.innerHTML=''`. ⚠️ `.session-tabs` is `overflow-x: auto`, so a scrolled-out tab still HAS a rect (over the logo); edges with an endpoint outside the strip are skipped, and a passive `scroll` listener re-anchors the rest.
**Unified session list**: `GET /api/sessions/unified` merges live sessions, persisted state, lifecycle-log history, and Claude transcript files into one deduped list (pure core in `src/services/unified-session-service.ts`). Transcript rows fold into their owning session via a `claudeSessionId → Codeman id` alias map, so resumed and `/clear`-respawned sessions do not appear twice. No terminal buffers in the response, unlike `/api/sessions`. Backs the Cmd+K Session Manager, plus pinning and cross-device tab order (`PUT /api/session-order`; pure merge helpers in `src/session-order.ts`, pushing device wins and server-only ids are never dropped). → [architecture-invariants#unified-session-list-and-session-manager](docs/architecture-invariants.md#unified-session-list-and-session-manager)
**Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `elicitation_complete`, `elicitation_response`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`. See `src/hooks-config.ts`; upstream hook semantics mirrored in `docs/claude-code-hooks-reference.md`.
**Approvals Inbox** (cross-session queue of prompts waiting on a human; `approvalsInboxEnabled`, SYNCED, default OFF: every surface is opt-in; only the store and answer endpoints run regardless, so flipping it ON shows anything already pending): `web/approval-inbox.ts` is a `sessionWaits`-style singleton fed by `/api/hook-event`, holding at most ONE item per session (a new prompt supersedes), claude-mode only, in-memory. Cards are answered via `POST /api/approvals/:id/answer`, which sends a digit / Esc / idle-prompt text through `writeViaMux` (menu answers never carry `\r`). ⚠️ `option` digits are accepted ONLY when they match options parsed from the captured pane frame, and the answer path RE-CAPTURES the pane first (a dialog that no longer parses on screen means the keystroke would land in the composer, so refuse with 409). ⚠️ Resolution on the heuristic `working` signal is restricted to `idle` items; permission/question items clear only on definitive signals (`stop`, `elicitation_complete`/`elicitation_response`, exit/delete, answer, supersede, 12h TTL). The frontend seeds from `GET /api/approvals` in `handleInit` **regardless of the setting**: the seed re-arms the tab-alert state machine (`setPendingHook`) unconditionally, and only populating `this.approvals` (the inbox surfaces) is gated — seeding used to be gated wholesale, which left a reloaded page with NO red tab while a permission dialog sat blocking a session (2026-08-15); `_onApprovalResolved` clears the pending-hook alert unconditionally for the same reason. ⚠️ The red/yellow tab alert itself is a STEADY border/background/dot with a pulse on top: the original keyframes swung to transparent at 0%/100%, so half of every cycle looked like a normal tab. Push Approve/Deny buttons stay gated on the setting (`sendPushNotifications` strips `actions`/`approvalId` when OFF) and are answered from `sw.js` directly so they work with no tab open. Surfaces (all gated on the setting): header bell (marker-hidden until count > 0, phones never show it) + drawer (`approvals-ui.js`), phone overview NEEDS YOU answer strips (`mobile-overview.js`). Design: `docs/approvals-inbox-plan.md`.
**Read My Mind intent profiles** (phase 1 of `docs/readmymind-plan.md`; `readMyMindEnabled`, SYNCED, default OFF): per-CASE profiles (user-stated `goals` + the user's recent real prompts), keyed by owner + realpath(workingDir) so they survive `/clear`/respawns and multi-user scoping is structural. Capture rides the transcript (`transcript:user_prompt` from `transcript-watcher.ts`), NOT the input paths: `POST /input` sees only programmatic prompts and the WS channel is raw keystrokes. The listener lives inside `startTranscriptWatcher()`'s `if (!watcher)` block (outside it would duplicate per hook event) and is claude-only + gated on the setting per event. Store: `src/intent-store.ts` singleton, `intents.json` written 0600 tmp+rename (prompts can contain secrets; never fed to `/api/search`). Endpoints: GET/PUT/DELETE `/api/sessions/:id/intent` + POST `/api/sessions/:id/readmymind` (`readmymind-routes.ts`, ownership via `findSessionOrFail` WITH `req`; registrations stay the bare `app.<method>('path')` shape, the endpoints.md drift scanner cannot see generics). **Phase 2 (predictor + 🧠 button)**: `readmymind-context.ts` is the PURE budgeted assembler (9 ranked sources, drop order siblings→away→workspace→tools, sections 1-4 truncate only); IO lives in `readmymind-collectors.ts` (transcript TAIL read — the live watcher keeps only a 500-char snippet — + git signals, skipped for remote-SSH cases) and the route; `readmymind-predictor.ts` reuses the AiCheckerBase spawn mechanics standalone (verdict-shaped base vs freeform JSON) as a mutable singleton routes call and tests stub. Claude-mode only (400), one in flight per session (409 CONFLICT), model = `readMyMindModel` setting defaulting to `AI_CHECK_MODEL` (opus, decided). Frontend `readmymind-ui.js`: header 🧠 marker-hidden (`btn-readmymind--hidden`) until the setting is ON; phones hide it in mobile.css and get a keyboard-accessory 🧠 key instead (ships in BOTH bar templates, revealed by the `rmm-enabled` class on the BAR element — setMode() rebuilds button innerHTML, so per-key state would be wiped; synced at init + every `applyHeaderVisibilitySettings()`). Alternate suggestions render as tappable rows that swap into the editable field without losing edits; Rethink rejects the whole shown set and carries the optional steer note (`#readMyMindSteer`, sent as `steer`, shown in ready + empty-result phases, cleared on each open). Suggestions render via value/`textContent` ONLY and Send/Insert go through `POST /input` (server-side, so the sendEnterKey/local-echo trap does not apply) — nothing auto-sends, ever. User guide: `docs/readmymind.md`.
**Voice dictation via Claude** (`claudeVoiceEnabled`, SYNCED, default OFF): the mic button can transcribe through this machine's Claude Code login instead of a Deepgram key, using the same speech-to-text service the CLI's own `/voice` mode uses. ⚠️ **Claude Code's voice mode itself is unusable here**: it opens the HOST's microphone (`sox`/`arecord`), and the CLI runs in a headless tmux pane while the human is in a browser elsewhere. So Codeman captures in the browser and borrows only the backend. Audio goes browser → Codeman → Anthropic (`src/web/voice-stream.ts`): the OAuth token never reaches the page, and the browser only sends PCM and receives text. ⚠️ Credentials are **read-only** (`src/claude-credentials.ts`) and Codeman never refreshes them — a refresh rotates the refresh token and could sign the user out of their own CLI; an elapsed token reports `expired` instead. ⚠️ Capture MUST be linear16/16 kHz/mono, so it uses an **AudioWorklet**, not MediaRecorder (which cannot emit raw PCM); `voice-pcm-worklet.js` is fetched from JS, so it is invisible to `cacheBustAssets` and borrows voice-input.js's `?v=` token — **edit the two together**. ⚠️ Transcript frames carry the WHOLE running transcript, not deltas: the Claude path replaces where the Deepgram path appends. Provider choice is `voiceSettings.provider` (`auto` prefers Claude → Deepgram → Web Speech). → `docs/claude-voice-plan.md`
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`.
**Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of EACH session per page load requests `full=1` (`_fullHistoryLoaded` Set); tab switches keep the cheap `?tail=` path, and scrolling up at the TOP of the buffer re-pulls `full=1` on demand (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). ⚠️ That re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for **claude ≥ 2.1.187 ONLY** at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. ⚠️ Codex was in that list and must never go back without a fresh measurement: codex-cli 0.147.0 ignores SGR wheel reports entirely (`mouse_any_flag=0`, inline viewport, transcript pushed into terminal scrollback), so forwarding produced a dead wheel (#227 follow-up). `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). ⚠️ When that gate is FALSE on a claude session whose local buffer is hollow (`baseY === 0`), the gesture becomes coalesced PageUp/PageDown key sends (`_maybePageCliTranscript`) instead of a no-op; ⚠️ and `getClaudeCliVersion()` must never cache a FAILED probe (one timeout used to disable forwarding process-wide until restart). `_logScrollRouting()` prints the routing decision and its inputs once per session — read it before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
**Detached start + service install** (issue #231): `codeman web -d` relaunches the SAME entry script with `detached:true` (setsid), so there is no controlling terminal and no shell job entry. ⚠️ `nohup` is NOT what makes this work: Node re-arms SIGHUP to its default disposition even when it inherits "ignore", and `cli.ts` handles SIGHUP with a graceful shutdown, so a delivered HUP still stops the server. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile check + `/api/status` probe): a second instance on the shared tmux socket attaches PTYs to the first one's live sessions. ⚠️ Neither may report success it has not observed — the parent polls `/api/status` until the child answers or dies, since `launchctl load` and a clean spawn are both silent about a server that starts and immediately exits. `--stop` verifies the pid still LOOKS like a Codeman server (`ps -o command=`) before signalling, because pids get recycled. Unit/label names live in `config/service-names.ts` so install.sh, `detectSupervisor()` and `service install` cannot drift into supervising two copies; they are instance-scoped, and identical to the historical names for the default instance. `service install` bakes the installing shell's PATH into the unit (launchd gives a job `/usr/bin:/bin:/usr/sbin:/sbin`, which finds neither a Homebrew/nvm `node` nor `tmux`/`claude`) and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install)
**Self-update** (App Settings → System → Updates): in-app updater for git-clone installs supervised by systemd/launchd (`systemd`, `launchd`, `launchd-daemon`, else `none` → "restart manually"). The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` that outlives the restart and writes progress to `update-status.json`, which the browser polls across the connection drop. `src/web/self-update.ts` splits pure helpers (unit-tested) from IO wrappers. npm installs report as non-updatable. → [architecture-invariants#self-update](docs/architecture-invariants.md#self-update)
**Attachments** (live external document references; all wiring in `file-routes.ts`): a **registry** maps a stable `attachmentId` to a realpath-resolved, extension-allowlisted absolute path, so browser requests never carry arbitrary absolute paths. ⚠️ The **magic-link scanner** (`codeman://attach?...` in terminal output) is **prompt-injectable**, so its scan path is force-confined to the session workspace; a hostile prompt could otherwise exfiltrate arbitrary host files over SSE. The security gate is an extension **allowlist**, not a blocklist. `document-conversion-limiter.ts` caps converter spawns globally: without it, N large docs detected at once fork N multi-minute processes, which is a resource-exhaustion vector. → [architecture-invariants#attachments](docs/architecture-invariants.md#attachments)
**Filesystem path picker** (Link Existing "Browse" + the mobile keyboard's `📁 Path` key): lazy one-directory browsing via `GET /api/filesystem/browse`, with `GET /api/filesystem/preview` for the tapped file. Inserts the path **without** Enter, so the prompt is never submitted; the sibling `⌫ All` key clears only the unsent prompt and must never send the agent's `/clear`. ⚠️ This is a **second file-serving surface and inherits neither the attachment confinement nor its ownership scoping** — it allowlists Home, `CASES_DIR`, `/mnt/d` and `CODEMAN_FILE_PICKER_ROOTS`, blocks sensitive trees, and rejects symlink escapes **after** `realpath`. ⚠️ The optional `sessionId` is an ownership boundary that must be `canAccessOwned`-checked by hand (it does not go through `findSessionOrFail`), and in multi-user mode a non-admin gets only their own `userSpacePath` as a root: per-user spaces live INSIDE `homedir()`, so a `Home` root exposes every other user's workspace. Previews go through the same global conversion limiter, and Markdown/TXT/JSON are served as inert `text/plain`. → [architecture-invariants#filesystem-path-picker](docs/architecture-invariants.md#filesystem-path-picker)
**File Viewer edit mode** (issue #212): the file-preview overlay edits workspace text files in place — `GET .../file-content?edit=1` + `PUT /api/sessions/:id/file-content`, policy in `src/config/file-editing.ts`. This is a **third file surface and the only one that WRITES**: read-path confinement (realpath + workspace + ownership) plus sensitive/blocked/`.git` denies and an extension **allowlist**; writes are `wx`-temp + rename (no `O_CREAT` anywhere = edit-in-place is structural); optimistic concurrency via sha256 `baseHash` → 409. ⚠️ `edit=1` never truncates and the client must never save a plain-preview buffer (the 500-line truncation would silently delete the rest). ⚠️ CRLF/UTF-8 guards: EOL re-applied server-side, non-UTF-8 refused via round-trip compare. → [architecture-invariants#file-viewer-edit-mode](docs/architecture-invariants.md#file-viewer-edit-mode), `docs/file-viewer-edit-plan.md`
**Raw file bodies are streamed and range-aware**: `file-raw` and the attachments `/raw` route always advertise `Accept-Ranges: bytes` and answer a `Range` header with `206` + `Content-Range` (single-range only; parser is pure + unit-tested in `src/web/http-range.ts`, a malformed spec is ignored → 200 while an out-of-bounds one is a 416). ⚠️ A 200-only response is what made the File Viewer's `<video>` unseekable: Chrome then reports `video.seekable` as `[0, 0]`, the scrub bar is inert and `currentTime = x` silently reverts (measured on an 18MB mp4), and Safari refuses to start the media at all. ⚠️ These bodies go out through `reply.hijack()`, which bypasses Fastify's status handling — `sendRawStream` must copy the status onto `reply.raw` by hand or a partial body ships labelled `200` and the browser treats a slice as the whole file. ⚠️ Closing the preview must **pause and unload** the media (`_stopFilePreviewMedia` in panels-ui.js): dropping the overlay's `visible` class is `display:none` and nothing else, and a DETACHED `HTMLMediaElement` keeps playing, which is how the X button used to leave a video audible with no player to pause.
**Ultracode / workflow-run visualization** (opt-in, default OFF): the Workflow tool writes a completion artifact only at run *end*, so live in-flight runs exist solely as transcript dirs. `workflow-run-watcher.ts` therefore synthesizes ACTIVE runs from transcripts until the completion artifact appears and supersedes them. It is **STANDALONE** and deliberately never imports or touches `subagent-watcher.ts`, despite reading the same tree. Two independent toggles: `showUltracodeAgents` (docked panel) and `ultracodeFloatingWindows` (floating windows); the watcher starts if **either** is on. → [architecture-invariants#ultracode--workflow-run-visualization](docs/architecture-invariants.md#ultracode-and-workflow-run-visualization)
**Clone a repository as a case** (issue #236, Add Case → **Clone Repo**): `POST /api/cases/clone` clones a public repo into the caller's case space synchronously (request held open, bounded by `GIT_CLONE_TIMEOUT_MS`, no job store); `POST /api/cases/clone-preflight` reports whether the URL can be cloned anonymously plus its real branches/tags. Core in `src/git-clone.ts`. ⚠️ **The URL is a code-execution surface**: `ext::sh -c <cmd>` (and ANY `<name>::<payload>` helper) makes git run a command, so every `::` form is refused, a leading `-` is refused, and every spawn is an argv array with `--` before the operands. ⚠️ **Non-interactive or the open request hangs** — `gitNonInteractiveEnv()` closes the terminal/askpass/ssh/GCM prompt paths; `HOME`/`PATH` stay inherited, so a user's OWN credential helper may authenticate (Codeman still never collects or stores credentials, and refuses a `user:password@` URL). ⚠️ Timeout kills the process GROUP (clone fans out into child processes), the destination is removed only if this attempt created it, and repository contents win over scaffolding (existing `CLAUDE.md` kept, hooks merged, repo-shipped `.claude/settings*` reported as a warning since its hooks run locally). The **Brain** picker sets the toolbar run mode on success. → [architecture-invariants#clone-a-repository-as-a-case](docs/architecture-invariants.md#clone-a-repository-as-a-case)
**Cross-session search**: `GET /api/search` federates an in-memory search over session metadata, run-summary events, and attachment-history entries. The pure core `searchSources()` does substring matching with hard per-type caps: **no regex (so no ReDoS) and no filesystem reads (so no traversal)**. The server-private `externalPath` is never read. PAST sessions (#261) come from `session-history-index.ts`, a capped snapshot of the unified list filled **outside** the request path (`/api/sessions/unified` publishes it; a stale one is rebuilt fire-and-forget), that indirection is what keeps the no-fs property. ⚠️ The snapshot is stored UNSCOPED with a per-row owner and MUST be re-filtered through `canAccessOwned()` on read; history rows carry `jumpTo.kind:'resume-session'`, since a closed session has no tab to select. → [architecture-invariants#cross-session-search](docs/architecture-invariants.md#cross-session-search)
**Web tabs** (dashboard URLs as tabs): a saved URL renders as a tab beside agent sessions. **NOT a sixth `SessionMode`** (no PTY, no tmux, no respawn), same reasoning that keeps Docker/remote-SSH as case overlays. Dashboards are **proxied through Codeman's own origin** by default, because a direct iframe fails three ways at once: prod is HTTPS so `http://` targets are blocked as mixed content, many dashboards send `X-Frame-Options: DENY`, and our own `default-src 'self'` CSP blocks cross-origin frames. Proxying leaves the prod CSP unchanged (`/webview/...` is `'self'`). ⚠️ The proxy is **NOT an API surface**: it authenticates on an in-memory capability in the path and is correspondingly exempt from the cookie + Origin checks; that exemption is fenced to safe methods and non-`/api` paths and is pinned by `test/webview-auth-exemption.test.ts`. ⚠️ Iframes omit `allow-same-origin` unless a dashboard is explicitly marked `trusted`, and `Authorization`/`codeman_session` are stripped upstream in **both** modes so `CODEMAN_PASSWORD` cannot leak. ⚠️ A sandboxed frame is **opaque-origin**, which breaks two things `curl` can never reproduce: its runtime-built root-absolute URLs escape `<base>` (fixed by an injected `runtimeUrlShim()`), and its same-host `fetch`/XHR are CORS-checked with `Origin: null` (fixed by `buildProxyCorsHeaders()` plus exempting the proxy from the global `OPTIONS`-204 short-circuit in `registerSecurityHeaders`). Both present as the dashboard's own "Failed to fetch" while the page renders fine. → [architecture-invariants#web-tabs](docs/architecture-invariants.md#web-tabs), `docs/web-tabs.md`
**Multi-user mode** (opt-in `--multiuser` / `CODEMAN_MULTIUSER=1`, OFF by default): named users with scrypt-hashed passwords in `~/.codeman/users.json`. Gated everywhere by `isMultiUserMode()`; when OFF, behavior is byte-identical to single-user because every scoping helper short-circuits. ⚠️ **Not a security boundary at the agent layer**: every session still runs as the SAME OS account. This separates WORKSPACES; it does not sandbox users (Docker cases are the isolation story). Ownership threads through `Session.owner` and is enforced in `findSessionOrFail`, list endpoints, SSE routing (fail-closed), WS, search, and file-preview. → [architecture-invariants#multi-user-mode](docs/architecture-invariants.md#multi-user-mode), `docs/multi-user-plan.md`
**Away digest**: `GET /api/away-digest` aggregates what happened while you were away from the lifecycle log, run-summary events, live sessions, token stats, and recent subagents. Pure aggregator in `web/away-digest.ts`. ⚠️ Returns `{success:true,digest}`, a legacy raw-ish shape consistent with the other raw GET handlers in `system-routes.ts`; frontend and tests read `.digest`. → [architecture-invariants#away-digest](docs/architecture-invariants.md#away-digest)
**Ralph todo-config**: per-session `maxTodos` (FIFO-eviction cap, default 500 = `MAX_TODOS_PER_SESSION`) + `todoExpirationMinutes` (auto-expiry, default 60) set via `POST /api/sessions/:id/ralph-config` (`RalphConfigSchema`, both `.int().positive()`). Stored on the tracker (`setMaxTodos`/`setTodoExpirationMinutes`) and **persisted/read-back via `RalphTrackerState`** (surfaced in the `loopState` getter → `toState()` + SSE broadcast → modal `populateRalphForm`), mirroring how `maxIterations` round-trips. Claude-only (skipped by `isExternalCliMode`).
**Port interfaces**: Routes declare dependencies via port interfaces (`src/web/ports/`). Routes use intersection types (e.g., `SessionPort & EventPort`).
### Frontend
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `i18n.js`(1.5) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `sanitize-html.js`(5.6) → `app.js`(6) → `terminal-ui.js`(7) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `cron-ui.js`(9.7) → `settings-ui.js`(10) → `panels-ui.js`(11) → `readmymind-ui.js`(11.3) → `ultracode-panel.js`(11.5) → `approvals-ui.js`(11.6) → `admin-ui.js`(11.7) → `session-ui.js`(12) → `webview-tabs.js`(12.5) → `mobile-overview.js`(12.55) → `home-sessions.js`(12.56) → `entrance-animations.js`(12.6) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15) → `ultracode-windows.js`(15.5) → `image-input.js`(16). `i18n.js` translates static + newly inserted application DOM while skipping terminal/response/file/user-name surfaces; `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData).
**Entrance animations** (`entrance-animations.js`, all OFF by default): opt-in animations for the four things that appear when work starts, chosen per surface via `data-tab-anim` / `data-term-anim` / `data-win-anim` / `data-line-anim` on `<html>`. Defaults are the `legacy` theme, so an untouched install behaves exactly as before and every hook short-circuits on its first line. ⚠️ Tabs and connection lines are **destroyed mid-animation** on every re-render (`_fullRenderSessionTabs()` replaces the strip's innerHTML; `_updateConnectionLinesImmediate()` does `svg.innerHTML = ''`), so both are tracked by id and re-applied to the fresh element with a **negative `animation-delay`** to resume rather than restart. ⚠️ The terminal-pane styles may animate **transform / opacity / clip-path only**, xterm's FitAddon derives rows+cols from `getComputedStyle(parent).width/height`, so animating width/height/padding there would resize the PTY. ⚠️ Window styles other than `beam` transform the window, which moves the rect its connection line is aimed at; `beam` deliberately animates opacity/filter only so its line can draw toward a stable target. Persisted to its own `codeman:*Anim` localStorage keys (per-device, deliberately NOT in the `.strict()` `SettingsUpdateSchema`); picker in App Settings → Appearance, full per-surface lab at `?animlab=1`.
**Mobile tab strip scrolling** (issue #257): under 768px the tab strip is a horizontal scroller (desktop wraps to a second row instead), so the active tab can sit off-screen. Three rules keep it reachable and they only work together: `_updateActiveTabImmediate()` scrolls the selected tab into view via `computeTabScrollLeft()` (pure, in constants.js) using **rect math on the strip's own `scrollLeft`**, never `scrollIntoView()`, which would also scroll the document under a fixed header; `_fullRenderSessionTabs()` **restores `scrollLeft`** across the `innerHTML` rebuild, since ambient rebuilds (a task badge appearing, a session created elsewhere) otherwise snap a mid-swipe strip back to 0; and it re-reveals the active tab **only when it changed** (`_lastRenderedActiveTabId`), so browsing the far end of the strip is not undone by background renders. ⚠️ Mobile no longer hoists the active session to the front of the strip: that reordering ran on full renders only, so tab order flipped depending on which render path fired, and it renumbered the Alt+N badges. Scroll-into-view replaces it; do not reintroduce it.
**Phone overview home screen** (`mobile-overview.js`, phones only, per-device `mobileOverviewEnabled`, default ON): under 430px the "C" logo shows a session overview (NEEDS YOU / CURRENT SESSIONS / PAST SESSIONS) instead of the welcome overlay; tablet and desktop are unchanged. The branch lives in `showWelcome()`/`hideWelcome()` (terminal-ui.js) behind `shouldUseMobileOverview()`, which is **width-driven** (`getDeviceType() === 'mobile'`) because this is a layout decision, unlike the settings namespace which stays handheld-based. ⚠️ The container ships with the `hidden` attribute and only this module removes it: never give `.mobile-overview` a bare `display` rule, since desktop does not load `mobile.css` (`media="(max-width: 1023px)"`) and would then render it unstyled. Live re-renders ride on the tail of `_renderSessionTabsImmediate()` (every state change it needs already funnels there); PAST rows come from one `_fetchUnifiedSessions(60)` per home-screen visit and resume through the shared `resumeHistorySession()`, so they behave exactly like the welcome screen's Resume list. ⚠️ Two things must stay in lockstep with surfaces outside this module, because divergence reads as a bug rather than a style: the split Run button carries the **toolbar's own classes** (`btn-toolbar btn-run mode-<backend>` / `btn-run-gear`) so the per-backend gradient and the light-skin overrides apply unchanged (mobile.css must therefore set no `background`/`color` on it), and row status uses the **session-tab language** (green dot when fine, `pulse` while working, yellow blinking row when waiting for input, red blinking row when a question is pending, mirroring `tab-alert-idle`/`tab-alert-action`). The picker mirrors the toolbar run-mode menu (`setRunMode()` + `run()`, `openWebviewFromMenu()` for saved dashboards) and deliberately omits its Recent-Sessions block, since PAST SESSIONS is that. Status pills carry `data-i18n-skip` (generic words like "idle" collide with state strings elsewhere).
**Desktop home tab rail** (`home-sessions.js`, desktop only): the welcome overlay centers ~560px of content in a ~1400px window, so its left gutter is dead space; it carries the open tabs as a rail **docked flush to the left edge, full height** (a vertically centered card floating mid-gutter read as debris). Rows are in **tab order**, not sorted by urgency like the phone overview, because the row badges are the Alt+1..9 indices, and each carries **created / last-active** stamps. State classification is REUSED from mobile-overview.js (`_mobileOverviewState`/`_mobileOverviewCaseFor`), which is why the module loads after it. ⚠️ The rail is `position: absolute` so the centered content never moves, which is exactly why it needs a **width gate in two places** — `HOME_SESSIONS_MIN_WIDTH` (1180) in the JS plus a `max-width: 1179px` media query as the backstop for a resize that outruns the matchMedia listener; drift between them means a rail overlapping the search panel, and `test/home-sessions.test.ts` pins them equal. ⚠️ `.home-sessions` is `display: flex`, so `[hidden]` must be re-asserted as `display: none` or the module's only visibility lever does nothing. ⚠️ Size scales with the viewport off **one knob**: `width: clamp(250px, 19vw, 430px)` plus a fluid `font-size` on `.home-sessions`, with every child sized in `em` — reintroducing `rem`/px type inside the block silently breaks the scaling, and widening the clamp past the gutter reintroduces the overlap the gate exists to prevent. The age stamps are refreshed **in place** by a 20s clock (`_tickHomeSessionsTimes()`, disarmed in `hideHomeSessions()`), never by re-rendering, which would restart every row's blink and working ring. Working state is deliberately byte-identical to the phone's: pulsing green dot + the `tab-load-spin` ring reused from the tab strip + the same green halo (added to `.mobile-overview-dot--working` at the same time), so "working" reads the same on every surface; **idle** is deliberately NOT that green — dot and pill mix toward `--text-muted` so a glance separates running from sitting. Live re-renders ride the tail of `_renderSessionTabsImmediate()` alongside the phone overview.
**Welcome "Resume Conversation" list** (terminal-ui.js): `loadHistorySessions()` fetches once and caches the corpus on `_historyAll`/`_historyCases`; every subsequent view (filter box, sort select, expand, the periodic refresh in panels-ui.js) goes through `_renderHistoryList()`, so never append rows to `#historyList` directly or re-fetch to re-sort. ⚠️ The box height is **class-driven**: expanding the list without `.history-list.expanded` leaves the collapsed `max-height` in place and just deepens a scroll well, which is the bug #260 reported (35 sessions in a ~4-row box). ⚠️ The A–Z sort keys off `_historyRowLabel()`, the SAME string the row renders (`name || firstPrompt || path`), most rows are transcript-backed and have no session name, so sorting on `name` alone silently does nothing. ⚠️ A filter implies expansion, and `_renderSearch()` hides `#historyHeader` (title + controls) as one unit while a search is active. Tests: `test/history-list-controls.test.ts`.
**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)** lives in that same handler: with a selection it copies, with none it must `return true` **without** `preventDefault()` or the interrupt is lost. `copyTerminalSelection` is deliberately absent from `SHORTCUT_ACTIONS` because the generic capture loop preventDefaults every match it dispatches. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry)
**Per-device vs synced settings**: the `displayKeys` set in settings-ui.js is a **client-side merge policy**, not a wire filter. A display key seeds from the server only when localStorage has no value for it, which is what prevents one device overwriting another; `showPlanUsageLimits` is additionally `delete`d from the incoming payload outright. Separately, `SettingsUpdateSchema` is `.strict()` and simply **does not declare** `skin`, `showFileViewerButton`, `showCronButton`, `webglRendererEnabled`, `localEchoEnabled`, `cjkInputEnabled`, or `extendedKeyboardBar`, so sending one of those is a validation error. The rest (`showResponseViewer`, `showPlanUsageLimits`, `language`, and most `show*` keys) ARE in the schema and do persist server-side; they are per-device by client policy only. ⚠️ Adding a new per-device setting means deciding **both** questions: membership in `displayKeys`, and presence in the schema.
**Settings surface** (`#appSettingsModal` + `#sessionOptionsModal` + `#createCaseModal`): the `set-*` language (left rail, groups of rows, control pinned right) is shared by all three modals through ONE `:is(#appSettingsModal, #sessionOptionsModal, #createCaseModal)` scope in styles.css: an `:is()` list takes its most specific argument's specificity, so every rule keeps the id weight it had and nothing downstream shifts. **App Settings** is a rail that is a **table of contents over ONE scrolling document**, not a tab switcher: every section stays mounted (`.set-section`, ids `settings-updates|terminal|layout|appearance|models|clis|notifications|voice|shortcuts|system`, in that order, the version and the updater leading and the rest of the system settings tailing), and `switchSettingsTab(id)` keeps its historical name but SCROLLS instead of hiding. **Session Options** and **Add Case** use the same surface with a rail that really SWITCHES (`switchOptionsTab` / `switchCaseModalTab` show one `.set-section` and `.hidden` the rest, since Summary owns its own scroller, Respawn is long, and Add Case is six independent forms). ⚠️ They also take a deliberate **size-up** that App Settings does not (900px shell, 236px rail, `height:auto` between `min(560px,80vh)` and 88vh, vs App Settings' tight 760×620): they are short task panels, not a document you scan, and at scanning density they read as a few fields marooned in an empty frame. Those per-modal blocks are the design, not drift. Phones (≤860px) give App Settings the sticky `#appSettingsJump` pill and give the other two a horizontal rail strip, which neither has a pill for. ⚠️ The Session Options rail entry labelled **Session** still keys off `context` (`data-tab="context"`, `#context-tab`, `switchOptionsTab('context')`), the rename is label-only. Add Case keeps its legacy `.form-row` markup (six panels of it, every id read back by session-ui.js) and is mapped onto the look by an adapter block scoped to `#createCaseModal .set-doc`. Do not restructure those forms just to reach the row classes. ⚠️ That adapter's `summary { display:flex }` **kills the native disclosure triangle**, so every `<details>` there needs the explicit `.set-adv-chev` and both marker suppressions (`list-style` + `::-webkit-details-marker`); without it five collapsed blocks render as plain headings nobody clicks. ⚠️ **The load/save contract is `getElementById` by id**: `openAppSettings()`/`saveAppSettings()`/`openSessionOptions()` read every control by a fixed id, so moving a control between sections is free but renaming or dropping one silently stops it loading or saving. Static guards: `test/app-settings-structure.test.ts` + `test/session-options-structure.test.ts` (rail↔section pairing, one-visible-section, the `data-claude-only` entries external CLIs drop). ⚠️ Model cards (`#appSettingsModelCards`) and the effort segment are **views over hidden `<select>`s** that remain the source of truth; the cards hold the BASE model and the "1M context window" switch composes `base + [1m]` back into `claudeModel`, which is what retires the old "takes precedence over the toggle below" trap. ⚠️ `.modal-tabs`/`.modal-tab-btn`/`.modal-tab-content` are RETIRED: no modal uses them and their CSS is deleted, and a reappearance means a modal drifted off the shared surface. ⚠️ The **Header & Panels live preview** is a scale model rebuilt from the chips (`_syncLayoutPreview`); it owns NO icons, it CLONES `.set-chip-ico` out of the chip, so each icon has exactly one copy in index.html. A chip joins it via `data-preview` (slot) + `data-preview-order`, or `data-preview-text` for readouts that are not buttons. Its frame is painted from skin tokens only (hardcoded black alphas turned it into a grey slab on the light skins) and is `data-i18n-skip`. ⚠️ In Session Options → Respawn, auto-resume is a `.set-callout` whose `<label>` **wraps its own switch with no `for=`** (nesting associates them; the label+`for` pair has historically double-fired), and the cycle steps are real checkboxes (`.set-checks`), not chips. ⚠️ `admin-ui.js` injects the multi-user Users entry into `.set-rail-items` + `.set-doc`, so those hooks must survive any restructure. → [architecture-invariants#settings-surface-app-settings-session-options-add-case](docs/architecture-invariants.md#settings-surface-app-settings-session-options-add-case)
**Header button visibility**: most header controls are opt-in and hidden by a marker class (`btn-multimonitor--hidden`, `btn-response-viewer-header--hidden`, `btn-file-viewer--hidden`, `btn-cron--hidden`) that `applyHeaderVisibilitySettings()` (settings-ui.js) toggles after settings load; the multi-monitor button is instead stripped at render by `renderIndexHtml`. ⚠️ Hiding must go through the marker class: the base rules are `display:inline-flex !important`, so an inline style cannot override them. Current desktop default is WS/CPU/MEM + File Viewer + gear, with the token chip and lifecycle-log button OFF. ⚠️ New header controls must not leak onto phones; `test/mobile-header-buttons-policy.test.ts` is the static guard. → [architecture-invariants#header-button-visibility-multi-monitor-response-viewer-file-viewer-cron](docs/architecture-invariants.md#header-button-visibility-multi-monitor-response-viewer-file-viewer-cron)
**Gesture control** (camera hand-tracking overlay, opt-in, default OFF): `CODEMAN_GESTURE=1` makes the feature *available*; `gestureControlEnabled` turns it on. The bundle is injected by `renderIndexHtml` only when enabled, which is why that method is `async` and reads settings with `readSettings(true)` (a fresh read: a post-save reload lands inside the 2s cache TTL and would otherwise render the pre-toggle state). **Source lives in `packages/gesture-control/`; edit there, run `npm run build:gesture`, and commit the regenerated bundle** because dev serves the committed bundle with no runtime bundler. The MediaPipe wasm + model are fetched separately and gitignored. ⚠️ Keep `MP_VERSION` in `fetch-gesture-assets.mjs` in sync with `@mediapipe/tasks-vision`. → [architecture-invariants#gesture-control-the-source-package](docs/architecture-invariants.md#gesture-control-the-source-package)
**Theme skins / branding / i18n**: `skin` selects a palette via `data-skin` on `<html>`, applied by an **inline pre-paint script** in `index.html` reading `localStorage['codeman:skin']` to avoid a flash of wrong theme. ⚠️ A skin is **four things that must stay in sync**, and missing any one degrades silently: the `html[data-skin="…"]` token block in `styles.css`, the xterm ANSI palette in `terminal-ui.js`, the pre-paint allowlist, and the Settings picker (both in `index.html`). `test/skin-themes.test.ts` is the static guard. Light skins additionally need `color-scheme: light` and xterm `minimumContrastRatio: 4.5`, and `applyTerminalSkin()` must call the local-echo overlay's `refreshFont()` because it caches the terminal fg/bg. `displayName` changes user-facing browser branding only and must NEVER rename npm package, CLI, API, storage, CSS, or protocol identifiers. `language` (`en`/`zh-CN`) keeps English as the canonical source so live switching stays reversible. User display names flow through `textContent`/attribute APIs and the server title's HTML escaper, never `innerHTML`. → [architecture-invariants#theme-skins](docs/architecture-invariants.md#theme-skins)
**Foldable settings identity**: responsive layout is width-driven via `MobileDetection.getDeviceType()`, but the localStorage namespace uses `MobileDetection.isHandheldDevice()` so an unfolded Android foldable keeps `codeman-app-settings-mobile`. ⚠️ Do not switch per-device settings namespaces from instantaneous viewport width: a posture-triggered WebView reload would lose opt-in UI. Regression profile: `OPPO Find N5 (unfolded)` in `test/mobile/devices.ts`. → [architecture-invariants#foldable-settings-identity](docs/architecture-invariants.md#foldable-settings-identity)
**WebGL renderer toggle** (`webglRendererEnabled`, per-device): the GPU-stall watchdog's sticky `codeman-webgl-disabled` marker survives page loads and is cleared only by an explicit OFF→ON save or `?webgl=force`. `?nowebgl` forces the DOM renderer per-load. → [architecture-invariants#webgl-renderer-toggle](docs/architecture-invariants.md#webgl-renderer-toggle)
**Shell keyboard accessory bar + one-shot Ctrl** (issue #262, `keyboard-accessory.js`): a **shell**-mode session automatically swaps the mobile accessory bar for terminal controls (Ctrl, Esc, Tab, four arrows, paste, dismiss); every other mode keeps the agent bar. `setMode()` now records the user's `extendedKeyboardBar` preference as the **base** layout and `refreshForActiveSession()` (called from `selectSession`) resolves base-vs-shell, so a settings save during a shell session cannot yank the bar away and switching back restores the user's choice. ⚠️ **Ctrl is a ONE-SHOT modifier applied in `terminal.onData`, not in a keydown handler**: a virtual keyboard emits no usable key events, so the character only exists as onData text. The hook sits AFTER `shouldSuppressTerminalQueryResponse` (xterm answers DA/CPR through onData too, and one of those would silently spend the modifier) and BEFORE every send path, so the control byte follows the normal control-char route. ⚠️ **Not every onData chunk is a keystroke**, and the query filter is not enough on its own: xterm ALSO emits mouse and focus reports on its own initiative, so the hook skips them via `isTerminalFocusOrMouseReport()` (they still reach the PTY, they just don't count as the next key). The mouse half is live — a shell session keeps the NARROW strip, so mouse DECSETs reach the browser and one tap while vim/htop runs spent the armed modifier silently (measured). The focus half is defense in depth: `FOCUS_ESCAPE_FILTER` in `session.ts` strips `\x1b[?1004h` from every PTY read, so `sendFocusMode` never turns on today; if it ever did, the bar's own post-key refocus would emit `\x1b[I` and eat the modifier before the user typed. ⚠️ It must disarm on ALL of: use, second tap, any other accessory key, session switch, keyboard dismissal, and a layout swap; a modifier left armed turns the next innocent keystroke into a control byte. ⚠️ **onData is not the only input path** — with `cjkInputEnabled` on, the CJK textarea owns the keyboard (onData returns early for everything it swallows, and the focus router sends `terminal.focus()` there, which is where the bar refocuses after every key), so `_handleCjkInput()` applies the modifier too. It is that module's single choke point to the PTY, so one call covers typed characters, IME flushes, Enter, backspace and arrows. Without it an armed modifier could neither fire NOR be spent, and survived to a later keystroke. Mapping is `ctrlByteFor()` (`code & 0x1f` over @A-Z[\]^_ and a-z, plus Ctrl+Space=NUL / Ctrl+?=DEL); characters with no control equivalent pass through unchanged, like a hardware keyboard. ⚠️ The armed style is `.accessory-btn.accessory-btn-ctrl.armed` (0,3,0) in BOTH stylesheets, and it cannot outrank mobile.css's light-skin repaint at **(0,3,1)** (`:is()` inherits its most specific argument, and that list holds `.btn-toolbar.btn-shell`) — so that rule excludes the state by hand as `.accessory-btn:not(.armed)`. Without the exclusion the armed button renders identically to a resting one on all four light skins, which is worse than no armed style at all.
**Dismissing the on-screen keyboard** (PRs #279/#280, `terminal-ui.js`): the terminal parks focus on a hidden textarea that nothing used to release, so TWO gestures now blur it, and they own different regions. **(1)** `_installMobileKeyboardDismiss()` — a document-level `touchend` that fires only while the terminal input actually holds focus, **never inside `#terminalContainer`** (tap classification owns that) and **never on a control** (`MOBILE_KEYBOARD_DISMISS_EXEMPT_SELECTOR`, matched with `closest()` so an icon inside a button counts). Session tabs are covered by the selector's `[tabindex]:not([tabindex="-1"])` arm, which is what stops a tab tap from blurring and then being re-focused by `selectSession()`. **(2)** In `_handleMobileTerminalTap`, a second tap on **inert `content`** (`startedWithTerminalFocus`) blurs instead of re-focusing. ⚠️ Scoped to `content` on purpose: the prompt row (`input`) keeps focus-then-position so a second tap still places the caret, and actionable rows blur earlier via `_isActionableMobileTerminalTap`. ⚠️ **A scroll ends in `touchend` too** — dismissing there closes the keyboard and drops the composer mid-read, so travel is tracked from `touchstart` and multi-touch is never a tap. Both classifiers MUST share one threshold: `initTerminal`'s `TAP_THRESHOLD` reads `MOBILE_KEYBOARD_DISMISS_TAP_SLOP`, since a gesture the terminal calls a scroll and the dismiss handler calls a tap is exactly that bug. ⚠️ **`test:ci` excludes `test/mobile/**`, so CI cannot see the only test covering (1)** — run `npm test -- test/mobile/keyboard.test.ts` by hand and diff the FAIL list against master. That blind spot is why merging the two PRs, which conflicted semantically but not textually, produced a red suite with two green CI checks.
**Phone toolbar: Enter replaces Shell** (post-1.8.0): inside `@media (max-width: 430px)` `btn-shell` is `display:none` and `btn-enter` takes its slot (`order: 4`); starting a shell moved into the Run dropdown (`Terminal / Shell` → `setRunMode('shell')` → `run()` → `runShell()`, button label "Run SH"). `runMode` is `z.string().max(20)` server-side, so new modes need no schema change. Desktop and tablet keep the green Run Shell button unchanged.
⚠️ **`sendEnterKey()` MUST go through `terminal._core.coreService.triggerDataEvent('\r', true)`** — not `sendInput()`, and never a raw POST to `/api/sessions/:id/input`. `localEchoEnabled` defaults to `MobileDetection.isTouchDevice()`, so on every phone the characters you type are buffered in the `LocalEchoOverlay` and have **never reached the PTY**; the `onData` Enter branch in terminal-ui.js is what flushes `pendingText` first and only then sends `\r` (after an 80ms delay so text lands first). Sending a bare `\r` submits an empty line and strands the typed text on screen, so the button looks dead. Replaying the keypress reuses the overlay flush, the flushed-offset cleanup and the ordering instead of reimplementing them. `KeyboardAccessory.sendKey()` is for escape sequences (arrows/Esc) and is the WRONG template to copy for input.
⚠️ **Skin overrides outrank plain class rules.** `styles.css` nests its skin block inside `html:not([data-skin="og"]) { … }`, so a bare `.btn-toolbar` rule in there resolves to specificity **(0,2,1)** and beats a `.btn-toolbar.btn-x` rule **(0,2,0)** in `mobile.css` regardless of load order. Toolbar-button colors set from mobile.css therefore need `!important` — that is why mobile.css leans on it so heavily. Symptom: only your `!important` properties land and everything else silently renders in generic toolbar grey.
**Connection-loss UI** (`computeConnectionLossUi()` in constants.js, writer `_updateConnectionLossUi()` in app.js): the service worker serves the cached app shell, so an unreachable server (phone off the tailnet, VPN down, server stopped) used to render a normal-looking empty dashboard whose only tell was the 8px header dot, which reads as "no sessions", not "no connection". Two surfaces now: a full-screen **overlay** while no server state has loaded this page load (nothing behind it is worth preserving), and a non-blocking **banner** once it has (the terminal scrollback stays readable). ⚠️ A **2.5s grace** is load-bearing: a COM deploy restarts the server and SSE is back in ~200ms, and a banner on every deploy trains the user to ignore it. `navigator.onLine === false` skips the grace, since that is never a blip. Retry re-arms SSE **and** the terminal WS (`planWsReconnect` can 'give-up', and the SSE backoff caps at 30s).
**SSE staleness watchdog** (`computeSseStale()` in constants.js, `_checkSseStale()` + a 5s interval in app.js): an `EventSource` that stops delivering does not always error, so `onerror` never fires, the header dot stays green, and every SSE-driven surface (tab status dots, sessions created on another device, renames) freezes until the user reloads. ⚠️ The 15s server keepalive was an SSE **comment** (`:keepalive`), and comments are **invisible to `EventSource` by spec**, so there was nothing a client could observe: it is now the named `sse:heartbeat` event (`cleanupDeadClients()`, sse-stream-manager.ts), which is exactly why the frame had to change type. ⚠️ Staleness is judged **only while the status is `connected`** and the device is online; that guard is the loop breaker, since a forced `connectSSE()` leaves `connected` immediately and cannot re-fire while a reconnect is in flight. ⚠️ The liveness stamp is applied inside `addListener` itself, so every registered handler (the `_SSE_HANDLER_MAP` wrappers AND the directly-registered ones) feeds it from one place; the heartbeat's own listener is a no-op that exists **only** to be registered, since `EventSource` drops named events nobody listens for. ⚠️ The watchdog interval is cleared at the top of `connectSSE()` and nowhere else (its only teardown path); clearing it elsewhere stacks intervals. Recovery needs no new sync path: the reconnect re-runs `handleInit` → `_resetAllAppState()`. The forced reconnect logs one diagnostic line, because a middlebox that strips heartbeats presents as "silently reconnects every 45s".
**Z-index layers**: subagent windows (1000), plan agents (1100), mobile/tablet fixed header (1200, `mobile.css`), modals on ≤768px (1300 — must beat the fixed header or the modal close button is buried), log viewers (2000), connection-loss overlay (2500, above the fixed header and modals), image popups (3000), local echo overlay (7).
**Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min).
**Keyboard shortcuts**: Escape (close), Ctrl+? (shortcut overlay), Ctrl/Cmd/Alt+K (session palette), Ctrl+W (kill), Ctrl+Tab (next), Alt+[/] (prev/next tab), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter or Ctrl+Enter (newline), Ctrl+C (copy selection, else interrupt) / Ctrl+Shift+C (copy, never interrupts), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font), Shift+Wheel (local scrollback when mouse passthrough is active). Rebindable via the registry.
### Security
**Full model: [`docs/security-architecture.md`](docs/security-architecture.md)** (network binding, auth pipeline, the tunnel caveat, file-serving hardening, supply-chain, instance isolation, recommended setups). **Layer-by-layer detail with the history behind each: [architecture-invariants#security-layers](docs/architecture-invariants.md#security-layers).**
| Layer | The rule |
| ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Auth** | Optional HTTP Basic via `CODEMAN_USERNAME` (default `admin`) / `CODEMAN_PASSWORD`. Active only when `CODEMAN_PASSWORD` is set (`middleware/auth.ts`) |
| **Network bind** | Defaults to loopback. Non-loopback without a password starts but warns loudly. Classifier: `network-auth-policy.ts` |
| **Host guard** | Always-on Host-header allowlist blocking DNS rebinding. ⚠️ **Custom reverse-proxy domains are rejected** unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix` |
| **CSRF / Origin** | Always-on cross-site Origin guard on state-changing requests. **A missing Origin is allowed** so curl/CLI and hooks keep working. ⚠️ The body parser keeps `text/plain` RAW; auto-JSON-parsing it enabled simple-request CSRF |
| **QR Auth** | Single-use 6-char tokens (60s TTL) for tunnel login. See `docs/qr-auth-plan.md` |
| **Sessions** | 24h cookie (`codeman_session`), auto-extend, device context audit |
| **Rate limit** | 10 failed auth/IP → 429 (15min decay). QR and hook-secret have separate buckets, so neither can lock out login |
| **Hook bypass** | `/api/hook-event` + `/api/status-telemetry` skip Basic auth (localhost-only, schema-validated), but when auth is active the loopback bypass requires `X-Codeman-Hook-Secret` **unconditionally** (Codeman cannot detect a user's own loopback reverse proxy) |
| **Tunnel** | Enabling a tunnel **refuses** without `CODEMAN_PASSWORD` unless exposure is acknowledged via `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` or the per-request `acknowledgeUnauthTunnel:true` action field (never persisted) |
| **Validation** | Zod schemas, Unicode-aware path allowlist regex, env prefix allowlist (`CLAUDE_CODE_*`/`OPENCODE_*`/`CODEX_*`/`GEMINI_*`/`GOOGLE_*`/`ANTIGRAVITY_*`) |
| **Headers** | CORS localhost-only, CSP, X-Frame-Options, HSTS if HTTPS |
**Security-relevant env vars**: `CODEMAN_MUX` (managed session), `CODEMAN_API_URL` (auto-set for hooks), `CODEMAN_ALLOWED_HOSTS` (extra Host/Origin allowlist entries for reverse proxies; bare `.suffix` matches subdomains), `CODEMAN_DOCKER_BRIDGE_HOOKS=1` (opt-in hooks-only listener on the docker bridge gateway).
### SSE Event Registry
155 event constants in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). **Both must be kept in sync**, and `test/sse-registry-parity.test.ts` is the guard that pins it (currently exactly in sync, 155 = 155, no drift either direction). The backend file's `@fileoverview` carries the per-category breakdown.
### API Routes
~200 handlers across 24 route files in `src/web/routes/`: system (45), sessions (34), cases (29), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), approvals (3), readmymind (4), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), voice (1 + the `/ws/voice/stream` relay), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
**HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`).
## Adding Features
- **API endpoint**: Types in `src/types/` domain file, route in `src/web/routes/*-routes.ts`. Return the `ApiResponse` envelope (`{ success: true, data }`; errors via `createErrorResponse()` with proper status code). Validate with Zod schemas in `schemas.ts`.
- **SSE event**: Add to `src/web/sse-events.ts` + `SSE_EVENTS` in `constants.js`, emit via `broadcast()`, handle in `app.js` (`addListener(`)
- **Session setting**: Add to `SessionState`, include in `session.toState()`, call `persistSessionState()`
- **App setting**: decide per-device vs synced first. Per-device keys go in the `displayKeys` set in settings-ui.js and must NOT be added to `SettingsUpdateSchema` (it is `.strict()`). ⚠️ Anything in `PUT /api/settings` that acts on a setting (the `toggleService` watcher calls) must resolve from **`merged`** (persisted + incoming), never from the raw request body: a partial PUT omits keys it doesn't intend to change, and `body.x ?? default` turns every omission into "apply the default" and silently resets live services. Pinned by `test/routes/system-routes-settings-partial-put.test.ts`.
- **Hook event**: Add to `HookEventType`, add hook in `hooks-config.ts:generateHooksConfig()`, update `HookEventSchema`
- **Mobile feature**: Add to relevant singleton, guard with `MobileDetection.isMobile()`. New header buttons must stay off phones (`test/mobile-header-buttons-policy.test.ts`).
- **New test**: Pick unique port (search `const PORT =`). Route tests use `app.inject()` (no port needed) — see `test/routes/_route-test-utils.ts`.
**Validation**: Zod v4 (different API from v3). Define schemas in `schemas.ts`, use `.parse()`/`.safeParse()`.
## State Files
All in `~/.codeman/`: `state.json` (sessions, settings, respawn, orchestrator, cron jobs/runs), `mux-sessions.json` (tmux recovery), `settings.json` (user prefs), `push-keys.json` + `push-subscriptions.json`, `session-lifecycle.jsonl` (audit log), `update-status.json` (self-updater progress, polled across the service restart), `linked-cases.json`, `webviews.json` (saved web-tab dashboard URLs), `remote-hosts.json` + `remote-cases.json`, `docker-hosts.json` + `docker-cases.json` + `docker-exports/`, `subagent-window-states.json` + `subagent-parents.json` (subagent window layout, GET/PUT `/api/subagent-window-states`/`-parents`), `hook-secret` (per-instance), `users.json` (multi-user, mode 0600) + `admin-audit.jsonl`, `intents.json` (Read My Mind intent profiles, mode 0600), `certs/` (self-signed TLS for `--https`), `.env` (CODEMAN_USERNAME/PASSWORD fallback for the `codeman attach` CLI). Transient: `self-update-runner.sh`. Multi-user case spaces live OUTSIDE the data dir at `~/codeman-users/<username>/cases` (shared across instances like `~/codeman-cases`, override `CODEMAN_USER_SPACES_DIR`).
**Generated top-level dirs** (all gitignored — don't edit or commit): `dist/` (esbuild output), `out/`, `coverage/`, `test-results/`, `tmp/`, `screenshots-echo-diag/`. The committed gesture bundle (`src/web/public/gesture/gesture-codeman.js`) IS tracked, but its runtime wasm/model assets (`src/web/public/gesture/wasm/`, `*.task`) are fetched and gitignored.
## Testing
**Never run the bare full suite** (`npm test` with no file argument): the default config includes the browser-driven suites (`test/mobile/**` and 3 other Playwright tests), which need a live server + chromium + environment-specific PNG baselines and will fail/hang locally. Run individual files, or `test:ci` for a broad sweep:
```bash
npm test -- test/<specific-file>.test.ts # Single file (SAFE, uses config/vitest.config.ts)
npm test -- -t "pattern" # By name (SAFE)
npm run test:ci # Everything except browser/perf suites — what CI runs
# npm test # DON'T — includes browser/visual suites
```
Raw `npx vitest` skips `config/vitest.config.ts`; always use `npm test --` or pass `--config config/vitest.config.ts`.
**Config**: Vitest with `globals: true`, `fileParallelism: false`. Timeout 30s, teardown 60s. `config/vitest.ci.config.ts` = same minus the browser/perf excludes — keep the two configs in sync when changing shared options.
**Tmux safety**: under vitest (`VITEST` env var, set automatically), `TmuxManager` no-ops ALL shell commands and becomes a pure in-memory mock — tests physically cannot create/kill/attach real tmux sessions (`IS_TEST_MODE` in `src/tmux-manager.ts`). Every docker IO path is no-op'd the same way. `Session` is test-gated too: instead of attaching a real tmux client, it spawns a raw-mode echo PTY (`TEST_PTY_SCRIPT` in `src/session.ts`), so integration tests get a live input/output loop that echoes each byte exactly once. `test/setup.ts` gives every test file a temporary `HOME`/`USERPROFILE` (all `homedir()`-derived state, `~/.codeman` and `~/codeman-cases` included, resolves into a per-file fixture; the Playwright browser cache path is preserved), and additionally strips `CODEMAN_PASSWORD`/`CODEMAN_USERNAME` (so auth state from the running instance can't leak into tests) and `CODEMAN_GESTURE` (a shell-exported gesture flag would flip render-injection assertions). ⚠️ Raw `npx vitest` without `--config` skips `setup.ts` and with it the temp-HOME isolation.
**Ports**: Pick unique ports manually, 3150+. Search `const PORT =` before adding new tests. Never 3000 (the live instance).
⚠️ **Browser tests can pass vacuously on mobile input paths.** Two traps, both hit on 2026-07-27 while fixing the phone Enter button: **(1)** driving input with `app.sendInput('…')` writes PAST the `LocalEchoOverlay`, so `pendingText` stays empty and any overlay bug is invisible — type with `page.keyboard.type()` instead; **(2)** headless Chromium reports `MobileDetection.isTouchDevice()` **false even with `hasTouch: true`**, so `_localEchoEnabled` is off and the local-echo branch never executes. Force it (`app._localEchoEnabled = true`) or the test proves nothing. Assert on real state (`app._localEchoOverlay.pendingText`, plus `tmux -L codeman capture-pane -p -t <pane>` for what actually reached the PTY), not on HTTP 200.
**Testing against the live instance**: prod is HTTPS-only on :3000 (`curl -sk https://localhost:3000/...`). ⚠️ `w1`/`w2`/`w3` are the user's REAL sessions — never send input to them. Create your own throwaway session (`POST /api/sessions` then `POST /api/sessions/:id/shell`; creation alone leaves `pid: null` and no pane), test against that, and `DELETE` it by exact id when done.
**Respawn tests**: Use `MockSession` from `test/mocks/index.ts` (defined in `test/mocks/mock-session.ts`). **Route tests**: `app.inject({ method, url, payload })` in `test/routes/` — no live port needed. **Mobile tests**: Playwright suite in `test/mobile/` (136 device profiles). Browser-testing infra and practices: `docs/browser-testing-guide.md`.
## Debugging
```bash
tmux list-sessions # List tmux sessions
curl localhost:3000/api/sessions | jq # Check sessions
curl localhost:3000/api/status | jq # Full app state
curl localhost:3000/api/subagents | jq # Background agents
cat ~/.codeman/state.json | jq # Persisted state
```
Mobile screenshots: `~/.codeman/screenshots/`, accessed via `GET/POST /api/screenshots`.
## Performance & Limits
Target: 20 sessions, 50 agent windows at 60fps. Limits live in `src/config/` (terminal 32MB, text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100), most env-overridable.
Two constraints worth knowing before you touch them: the env-derived PTY buffer trim is **clamped to ≤75% of max**, because a trim ≥ max would disable `BufferAccumulator` trimming entirely and make memory unbounded; and browser xterm scrollback is a **separate hardcoded 50k** (`DEFAULT_SCROLLBACK` in constants.js), deliberately lower than tmux's 100k history because 100k per tab is a mobile-memory hazard. The settings keys `terminalScrollbackLines`/`terminalBufferMaxBytes`/`terminalBufferTrimBytes` are schema-validated but **inert**; only `tmuxHistoryLimit` is wired live. → [architecture-invariants#buffers-uploads-and-terminal-history](docs/architecture-invariants.md#buffers-uploads-and-terminal-history), `docs/terminal-anti-flicker.md`
**Memory leaks (24+ hour sessions)**: use `CleanupManager`, clear Maps in `stop()`, guard async with `if (this.cleanup.isStopped) return`. Frontend: store handler refs, clean in `close*()`. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Verify: `npm test -- test/memory-leak-prevention.test.ts`.
## Scripts & Tunnel
**`install.sh`** (repo root, 69KB) is the public entry point: `curl -fsSL <raw url> | bash` installs Node/tmux if missing, clones to `~/.codeman/app`, builds, and offers a systemd/launchd service. The network-access prompt is 3-way: **Tailscale** (loopback bind + guided `tailscale serve --bg <port>` HTTPS setup: install/login/operator/tailnet-HTTPS-toggle, then curl-verified end-to-end), **LAN** (0.0.0.0 + password prompt), or **local-only**; it preserves the existing binding on re-runs via `read_existing_binding()`. Tailscale state is detected dynamically from `tailscale serve status --json` (no marker files); the installer must NEVER `tailscale serve reset` or touch serve mappings other than 443→Codeman's port (users have unrelated serve config). `install.sh update`, `install.sh uninstall`, and `install.sh tailscale` (retrofit Tailscale access onto an existing install) also exist; `CODEMAN_NONINTERACTIVE=1` approves system changes for automation, `CODEMAN_TAILSCALE=1` presets the Tailscale choice (never installs Tailscale non-interactively).
Other key scripts: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh [quick|named] start|stop|status|url` (quick = random trycloudflare URL, default; `named setup|enable` = fixed-hostname tunnel via `scripts/codeman-tunnel-named.service`; bare `start|stop|url` still means quick), `scripts/run-beta.sh` (isolated beta instance), `scripts/build-agent-image.mjs` (docker base image), `scripts/self-update.sh` (detached updater). Production services: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.
-21
View File
@@ -1,21 +0,0 @@
MIT License
Copyright (c) 2024-2026 Codeman Contributors
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+3 -1098
View File
File diff suppressed because it is too large Load Diff
-992
View File
@@ -1,992 +0,0 @@
<p align="center">
<img src="docs/images/codeman-title.svg" alt="Codeman" height="60">
</p>
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
</p>
<p align="center">
<a href="README.md">English</a> &bull; <strong>简体中文</strong>
</p>
<p align="center">
<a href="https://opensource.org/licenses/MIT"><img src="https://img.shields.io/badge/License-MIT-1e3a5f?style=flat-square" alt="License: MIT"></a>
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/Node.js-22%2B-22c55e?style=flat-square&logo=node.js&logoColor=white" alt="Node.js 22+"></a>
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.9-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.9"></a>
<a href="https://fastify.dev/"><img src="https://img.shields.io/badge/Fastify-5.x-1e3a5f?style=flat-square&logo=fastify&logoColor=white" alt="Fastify"></a>
<a href="https://github.com/Ark0N/Codeman/graphs/contributors"><img src="https://img.shields.io/github/contributors/Ark0N/Codeman?style=flat-square&color=3b82f6" alt="Contributors"></a>
<a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a>
</p>
<p align="center">
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — 并行子智能体可视化" width="900">
</p>
<p align="center">
<img src="docs/images/codeman-tour-20260724.png" alt="Codeman 仪表盘导览:按项目分组的会话标签页、一键 Run 启动新智能体、页头实时用量" width="900">
</p>
> 本文档由英文版 [`README.md`](README.md) 翻译而来。如有出入,以英文版为准。
一行命令即可安装(macOS 和 Linux,Windows 通过 WSL):
```bash
curl -fsSL https://getcodeman.com/install | bash
```
```bash
codeman web
# 打开 http://localhost:3000,开启你的第一个会话
```
安装器在每次系统改动前都会先询问;重跑同一条命令即可原地更新。详见[快速开始 — 安装](#快速开始--安装)。
---
## 快速开始 — 安装
```bash
curl -fsSL https://getcodeman.com/install | bash
```
该脚本会在缺失时自动安装 Node.js 和 tmux,把 Codeman 克隆到 `~/.codeman/app` 并完成构建。几点须知:
- **先询问,后改动。** 所有系统级改动(安装软件包、下载 AI CLI)都会先征求确认;结束时的菜单可选择:直接在本终端运行、安装为后台服务(systemd/launchd,开机自启),或暂不启动。不选就不会有任何后台进程。
- **重跑即更新。** 再次运行同一条命令即可原地更新已完成的安装:`~/.codeman/app` 中的本地改动会被 stash(绝不丢弃),运行中的服务会自动重启并校验。若首次安装中途失败,重跑会继续完成完整的安装流程。也可以使用 `install.sh update` 与 `install.sh uninstall`。
- **CI / 无终端环境:** 没有终端时,涉及系统改动的步骤会带着说明中止,而不是静默执行;在自动化场景设置 `CODEMAN_NONINTERACTIVE=1` 即可批准这些步骤。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli) 或 [Pi](https://pi.dev)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这六个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
```bash
codeman web
# 打开 http://localhost:3000,开启你的第一个会话
```
**想和小团队共用一台?** 改用多用户模式启动:每人拥有自己的登录与工作空间。
```bash
codeman users add alice --admin # 创建第一个管理员账号
codeman web --multiuser # 命名登录 + 按用户隔离的案例空间
```
详见下文[多用户模式](#多用户模式可选启用)。
<details>
<summary><strong>作为后台服务运行</strong></summary>
安装器结尾的菜单(选项 2)可以帮你完成这一步,并在宣告成功前校验服务确实已启动。如需手动配置:
**Linux(systemd):**
```bash
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/codeman-web.service << EOF
[Unit]
Description=Codeman Web Server
After=network.target
[Service]
Type=simple
ExecStart=$(which node) $HOME/.codeman/app/dist/index.js web
Restart=always
RestartSec=10
[Install]
WantedBy=default.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now codeman-web
loginctl enable-linger $USER
```
**macOS(launchd):**
```bash
mkdir -p ~/Library/LaunchAgents
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.codeman.web</string>
<key>ProgramArguments</key>
<array>
<string>$(which node)</string>
<string>$HOME/.codeman/app/dist/index.js</string>
<string>web</string>
</array>
<key>RunAtLoad</key><true/>
<key>KeepAlive</key><true/>
<key>StandardOutPath</key>
<string>/tmp/codeman.log</string>
<key>StandardErrorPath</key>
<string>/tmp/codeman.log</string>
</dict>
</plist>
EOF
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
```
</details>
<details>
<summary><strong>Windows(WSL)</strong></summary>
```powershell
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli) 或 [Pi](https://pi.dev))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
</details>
---
## 移动端优化的 Web UI
在任意手机上都能获得最跟手的 AI 编程智能体体验。完整的 xterm.js 终端、本地回显、滑动导航,以及为真正的远程办公而设计的触控优化界面 —— 而不是把桌面 UI 硬塞进小屏幕。
<table>
<tr>
<td align="center" width="40%"><img src="docs/screenshots/mobile-session-keyboard-20260727.png" alt="移动端 — 通过键盘配件栏与 Enter 按钮回答智能体的方案提示" width="300"></td>
<td align="center" width="60%"><img src="docs/screenshots/mobile-toolbar-enter-20260727.png" alt="移动端工具栏:配件栏的 /init、/clear、剪贴板与 Esc,下方是 Run、案例、停止、Enter、语音与设置控件" width="440"></td>
</tr>
<tr>
<td align="center"><em>触控回答提示</em></td>
<td align="center"><em>配件栏 + 独立 Enter 按钮</em></td>
</tr>
</table>
<table>
<tr>
<th>普通终端 App</th>
<th>Codeman 移动端</th>
</tr>
<tr><td>远程输入延迟 200–300 毫秒</td><td><b>本地回显 —— 即时反馈</b></td></tr>
<tr><td>字小、无上下文</td><td>完整 xterm.js 终端</td></tr>
<tr><td>无会话管理</td><td>滑动切换会话</td></tr>
<tr><td>无通知</td><td>审批 / 空闲时推送提醒</td></tr>
<tr><td>需手动重连</td><td>tmux 持久化</td></tr>
<tr><td>看不到智能体</td><td>实时查看后台智能体</td></tr>
<tr><td>斜杠命令靠复制粘贴</td><td>一键 <code>/init</code>、<code>/clear</code>、<code>/compact</code></td></tr>
<tr><td>在手机上手打密码</td><td><b>扫二维码 —— 即时认证</b></td></tr>
</table>
- **键盘配件栏** —— 在虚拟键盘上方提供 `/init`、`/clear`、`/compact` 快捷按钮;破坏性命令需双击确认,绝不误触
- **独立的 Enter 按钮** —— 以按键方式回放,先冲刷本地回显缓冲的文本,不会让内容滞留在屏幕上
- **滑动导航与智能键盘处理** —— 左右滑动切换会话;键盘弹出时工具栏与终端整体上移(`visualViewport` API)
- **为手机而生** —— 刘海与 Home 指示条的安全区适配、44px 触控目标、底部抽屉式 case 选择器、原生惯性滚动
```bash
codeman web --https
# 在手机上打开:https://<你的IP>:3000
```
> `localhost` 走纯 HTTP 即可。从其他设备访问时请使用 `--https`,或使用 [Tailscale](https://tailscale.com/)(推荐)—— 它提供私有网络,让你无需 TLS 证书即可从手机访问 `http://<tailscale-ip>:3000`。
### 安全的二维码认证
在手机键盘上输密码太痛苦了。Codeman 用**密码学安全的一次性二维码令牌**取而代之 —— 扫描桌面上显示的二维码,手机即刻完成认证。
每个二维码编码的是一个包含 6 字符短码的 URL,该短码在服务端映射到一个 256 位密钥(`crypto.randomBytes(32)`)。令牌每 **60 秒**自动轮换,**首次扫描即原子性消费**(重放永远失败),并采用**基于哈希的 `Map.get()` 查找**,不会通过响应时延泄露任何信息。短码只是一个不透明指针 —— 真正的密钥永远不会出现在浏览器历史、`Referer` 头或 Cloudflare 边缘日志中。
该安全设计覆盖了 ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin)(USENIX Security 2025,该研究发现 Top-100 网站中有 47 个存在漏洞)所指出的全部 6 个关键二维码认证缺陷:强制一次性使用、短 TTL、密码学随机性、服务端生成、扫描时桌面实时通知(QRLjacking 检测),以及 IP + User-Agent 会话绑定与手动吊销。双层速率限制(按 IP + 全局)使得在 62^6 = 568 亿种可能短码空间内进行暴力破解变得不可行。完整安全分析见:[`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
---
## 使用 Codeman —— 人类操作指南
从头到尾走一遍如何在浏览器里驾驭 Codeman。如果你刚装好,就从这里开始。
### 1. 启动服务器
```bash
codeman web # localhost:3000(仅环回 —— 安全默认值)
codeman web --port 8080 # 自定义端口(或设置 CODEMAN_PORT)
codeman web --https # 自签名 TLS(仅远程访问时需要)
codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_PASSWORD(见「安全」)
```
打开打印出的 URL。整个页面是一个单一仪表盘;下面的一切都在这里完成。
### 2. 创建你的第一个会话
点击 **+ New Session**(或 **Quick Start**)。一个会话就是一个运行在自己 tmux 终端里的 AI CLI。你可以选择:
| 字段 | 作用 |
| ---------------------- | ------------------------------------------------------------------------------------------- |
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini`、`Pi` 或 `Terminal`(普通 shell)。 |
| **模型** | 每会话模型(App Settings → Claude Model)。软默认值 —— 会话内 `/model` 依然有效。 |
| **Effort / Ultracode** | 推理力度(`low`–`max`),或用 `ultracode` 开启动态多智能体工作流。随时可用 `/effort` 切换。 |
点击启动 —— Codeman 通过真实 PTY 拉起 CLI,并经 SSE 流式传输到你的浏览器。
### 3. 读懂仪表盘
- **标签(顶部)** —— 每个会话一个。`Alt+1`–`9` 跳转,`Ctrl+Tab` 下一个,拖拽排序(标签顺序会跨设备同步)。
- **终端(中央)** —— 真实的 `xterm.js` 终端;完整 TUI 正常渲染。直接输入并按 **Enter** 发送。`Shift+Enter` 插入换行。
- **侧边面板** —— Respawn、Orchestrator、Cron、Subagents、Settings(从工具栏切换)。
### 4. 与智能体对话
- **直接在终端输入提示** —— 即使跨越重连,输入也是精确一次送达(连接中断绝不会丢失或重复发送提示)。
- **粘贴或拖放图片**,直接进入会话。
- **语音输入** —— `Ctrl+Shift+V`(Deepgram Nova-3,自动静音停止)。
- **附件** —— 注册外部文件/文档,并内联预览 Office/PDF。
### 5. 让它自主运行
| 模式 | 用途 | 位置 |
| ---------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ |
| **Respawn** | 长时间无人值守运行 —— 空闲/限额时自动重启 CLI,带自适应时序。预设:`solo-work`、`overnight-autonomous` 等 | Respawn 标签页 |
| **Orchestrator** | 把一个目标变成分阶段计划,并跨多个智能体推动完成。 | 编排器面板 |
| **Cron** | 已保存的、命名的定时任务(`once`/`interval`/`daily`/`weekly`),到期时拉起会话并发送提示。 | ⏰ Cron 按钮(可选启用:App Settings → Display → Header Displays) |
| **Auto-resume** | 订阅限额重置后自动继续。 | Respawn 标签页(顶部) |
### 6. 随时随地访问
- **手机/平板** —— UI 完全触控优化;扫描桌面上的**二维码**即可免密码登录。
- **网络之外** —— `./scripts/tunnel.sh start` 打开一条 Cloudflare 隧道(先设置 `CODEMAN_PASSWORD`)。
- **SSH** —— `sc` 选择器可从终端附着任意会话(`sc` 交互式,`sc 2` 快速附着,`sc -l` 列表)。
### 7. 运维与维护
- **App Settings** —— 模型、effort、权限启动模式、主题/皮肤、通知、显示开关、各 CLI 的专属选项,以及跨设备同步的自定义显示名称和按设备保存的英文/简体中文界面语言。
- **自更新** —— git-clone 安装可在 **Settings → Updates** 中原地更新。
- **部署你自己的改动** —— 见[开发](#开发)。
> ⚠️ **安全提示:** 如果你正在 Codeman 受管会话*内部*工作(`echo $CODEMAN_MUX` → `1`),绝不要直接运行 `tmux kill-session` / `pkill claude` —— 请使用 Web UI 或 `./scripts/tmux-manager.sh`。
---
## 零延迟输入叠加层
<p align="center">
<img src="docs/images/zerolag-demo-20260728.gif" alt="Zerolag 演示:两台手机并排对比,即时本地回显与 600ms-2.7s 服务端回显" width="900">
</p>
远程访问你的编程智能体时(VPN、Tailscale、SSH 隧道),每次按键通常需要 200–300 毫秒往返。Codeman 实现了一套**受 Mosh 启发的本地回显系统**,无论延迟多高,打字都感觉即时。
xterm.js 内部一个像素级精准的 DOM 叠加层以 0ms 渲染按键。后台转发会以 50ms 防抖批次静默地把每个字符送往 PTY,因此 Tab 补全、`Ctrl+R` 历史搜索以及所有 shell 特性都正常工作。当服务端回显在 200–300ms 后到达时,叠加层无缝消失、真实终端文本接管 —— 整个切换过程不可见。
- **抗 Ink 架构** —— 它作为 `.xterm-screen` 内 z-index 7 的一个 `<span>` 存在,完全不受 Ink 持续重绘屏幕的影响(此前两次使用 `terminal.write()` 的尝试都失败了,因为 Ink 会破坏注入的缓冲区内容)
- **字体匹配渲染** —— 从 xterm.js 的计算样式读取 `fontFamily`、`fontSize`、`fontWeight` 与 `letterSpacing`,使叠加层文本与真实终端输出在视觉上无法区分
- **完整编辑** —— 退格、重打、粘贴(多字符)、光标跟踪,输入超过终端宽度时多行换行
- **重连后持久** —— 未发送的输入通过 localStorage 在页面刷新后保留
- **默认启用** —— 桌面端与移动端均可用,会话空闲或繁忙时都生效
> 已抽取为独立库:[`xterm-zerolag-input`](https://www.npmjs.com/package/xterm-zerolag-input) —— 见[已发布的包](#已发布的包)。
---
## 实时智能体可视化
实时观看后台智能体工作。Codeman 监控智能体活动,将每个智能体显示在一个可拖拽的浮动窗口中,并用「黑客帝国」风格的动态连接线连回父会话。
<p align="center">
<img src="docs/images/subagent-windows-20260724.png" alt="子智能体可视化 —— 三个并行 Explore 智能体的浮动窗口与实时工具调用日志" width="900">
</p>
- **浮动终端窗口** —— 每个智能体一个可拖拽、可调整大小的面板,带实时活动日志,逐条展示每一次工具调用、文件读取与进度更新
- **连接线** —— 用动态绿色线条连接父会话与其子智能体,随智能体的产生与完成实时更新
- **状态与模型徽标** —— 绿色(活动)、黄色(空闲)、蓝色(已完成)指示,并以 Haiku/Sonnet/Opus 的颜色编码区分模型
- **自动行为** —— 窗口在产生时自动打开、完成时自动最小化,标签徽标显示「AGENT」或「AGENTS (n)」计数
- **嵌套智能体** —— 支持 3 层层级(主会话 → 团队成员智能体 → 子-子智能体)
多智能体 Workflow 运行(「ultracode」)同样可视化:一个浮动运行窗口实时跟踪整个工作流,展示阶段、各智能体的 token 用量与当前工具:
<p align="center">
<img src="docs/images/ultracode-window-20260724.png" alt="Ultracode 工作流可视化 —— 实时运行窗口,含各智能体 token 与阶段" width="900">
</p>
**智能体团队(Agent Teams)** —— 一等公民式支持 Claude Code 原生的多智能体团队(`CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`)。`TeamWatcher` 轮询 `~/.claude/teams/`,将团队成员匹配到其主会话,并以实时子智能体窗口呈现,且具备**团队感知的空闲检测** —— 因此当团队成员仍在工作时,重生控制器不会被触发。详见 [`docs/agent-teams/`](docs/agent-teams/)。
---
## 重生控制器(Respawn Controller)
自主工作的核心。当智能体进入空闲,重生控制器会检测到,发送继续提示,循环执行上下文管理命令以获得全新上下文,然后恢复工作 —— 可完全无人值守运行 **24 小时以上**。
```
WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE → WATCHING
```
- **多层空闲检测** —— 完成消息、AI 驱动的空闲检查、输出静默、token 稳定性
- **用量限额自动恢复**(_可选,默认关闭_)—— 当 Claude 因订阅用量限额而停止("You've hit your limit · resets 3pm")时,Codeman 会解析重置时间,等到限额刷新(外加 2 分钟安全缓冲)后自动关闭限额对话框并发送 `continue`,让通宵任务平稳跨过 5 小时窗口而不是停摆到早晨。可识别 Claude Code 各版本的全部限额消息格式;若仍受限会自动重试;计划在 Codeman 重启后依然生效;暂停期间会阻止重生循环,避免 `/clear` 清掉等待中的对话。在会话 Respawn 标签页顶部按会话启用
- **熔断器** —— 当 Claude 卡住时防止重生抖动(CLOSED → HALF_OPEN → OPEN 状态,跟踪连续无进展与重复错误)
- **健康评分** —— 0–100 健康分,分项涵盖循环成功率、熔断器状态、迭代进展与卡死恢复
- **内置预设** —— `solo-work`(3s 空闲,60min)、`subagent-workflow`(45s,240min)、`team-lead`(90s,480min)、`ralph-todo`(8s,480min)、`overnight-autonomous`(10s,480min)
---
## 编排器循环(Orchestrator Loop)
超越单会话重生,**编排器**把一个高层目标转化为分阶段计划,并跨多个智能体推动其完成 —— 这是一个运行 `idle → planning → approval → executing → verifying → (replanning) → completed` 的状态机。
- **先规划,后执行** —— 从你的目标生成分阶段计划,并在动手前暂停等待审批;可带反馈拒绝以重新生成
- **逐阶段验证关卡** —— 每个阶段在下一阶段开始前都会被验证;失败时编排器会重新规划而非一头扎下去
- **多智能体执行** —— 将各阶段分发给团队智能体 / 任务队列,协调超出单会话能力的工作
- **崩溃安全** —— 完整状态持久化在 `state.json` 的 `orchestrator` 键下,可在重启后存续
- **可从 UI 或 API 驱动** —— 编排器面板,或 `POST /api/orchestrator/start` → `/approve` → `/status`(共 10 个端点)
> 完整设计:[`docs/orchestrator-loop-architecture.md`](docs/orchestrator-loop-architecture.md)。
---
## 多会话仪表盘
运行 **20 个并行会话**且全程可见 —— 60fps 的实时 xterm.js 终端、按会话的 token 与成本跟踪、基于标签的导航,以及一键管理。
### 持久化会话
每个会话都运行在 **tmux** 内 —— 会话可在服务器重启、网络中断与机器休眠后存续。启动时自动恢复,具备双重冗余。幽灵会话发现机制能找到孤立的 tmux 会话。受管会话带有环境标签,因此智能体不会杀掉自己的会话。
### 会话管理器与命令面板
`Ctrl/Cmd/Alt+K` 打开模糊搜索的会话面板;**Browse all sessions** 打开会话管理器:一份去重后的完整清单,涵盖 Codeman 所知的一切(活动会话、来自状态与生命周期历史的既往会话,以及 Claude 转录),每一行都显示其第一条与最近一条提示。
- **置顶(Pin)**:把会话固定到列表顶部。被置顶的会话甚至能挺过被杀掉(降级为一条轻量的已停止记录,依然可见、可恢复)。
- **名称保留**:从会话管理器恢复既往会话时保留其原有名称,而不是生成一个新名称。
- **跨设备标签顺序**:拖拽排序的标签顺序保存在服务端,你的排列会从桌面跟随到手机。
### 主机名感知的窗口标题
在多台主机上运行 Codeman(笔记本、开发机、NAS)?浏览器标签标题是 `codeman:<主机名>`,让你无需点进去就能分辨每个标签对应哪个后端:
```bash
codeman web # codeman:<os.hostname()>
codeman web --title-hostname dev-box # codeman:dev-box(用于覆盖嘈杂的主机名)
```
标题在首字节时就被模板化进所提供的 HTML 中,因此从第一帧绘制起就是正确的,且无需 JavaScript 也能工作。同样的主机名前缀也应用于标签闪烁格式(`⚠️ (N) codeman:<host>`)和操作系统级桌面通知(`codeman:<host>: <事件>`),让系统通知中心里的跨主机提醒也不再含糊。
### 智能 Token 管理
| 阈值 | 动作 | 结果 |
| --------------- | --------------- | ---------------------- |
| **110k tokens** | 自动 `/compact` | 上下文被摘要,工作继续 |
| **140k tokens** | 自动 `/clear` | 以 `/init` 全新开始 |
### 通知
当会话需要关注时实时桌面提醒 —— `permission_prompt` 与 `elicitation_dialog` 触发关键的红色标签闪烁,`idle_prompt` 触发黄色闪烁。点击任意通知即可直接跳转到相关会话。Hook 按 case 目录自动配置。
### 运行摘要(Run Summary)
点击任意会话标签上的图表图标,即可看到所发生一切的时间线 —— 重生周期、token 里程碑、自动 compact 触发、空闲/工作切换、hook 事件、错误等等。
### 零闪烁终端
基于终端的 AI 智能体(Claude Code 的 Ink、OpenCode 的 Bubble Tea)会在每次状态变更时重绘屏幕。Codeman 实现了一套 6 层抗闪烁流水线,让所有会话都获得平滑的 60fps 输出:
```
PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端 rAF → xterm.js(60fps)
```
---
## 更多特性
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity**、**Gemini** 或 **Pi**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*`、`PI_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md) 与 [`docs/pi-integration.md`](docs/pi-integration.md)
- **Docker 会话** —— 在隔离且加固的容器中运行案例。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一案例的多个会话共享一个容器;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话**:把案例指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
- **语音输入** —— 用 Deepgram Nova-3 口述提示(带 Web Speech API 回退):切换录音、自动静音停止、实时音量表(`Ctrl+Shift+V`)
- **图像输入** —— 直接把图片粘贴或拖放进会话
- **手势控制** _(可选)_ —— 一个 MediaPipe 手部追踪叠加层,可徒手抓取/拖动会话窗口并捏合按钮。用 `CODEMAN_GESTURE=1` + App Settings → Display 启用
- **多显示器横跨** _(macOS)_ —— 一键打开一个横跨所有显示器最大化的浏览器窗口,让浮动的智能体/手势面板可以跨越物理拼接缝
- **文件查看器按钮** _(可选)_ —— 头部新增一个按钮,一键切换内置文件浏览器面板;在 App Settings → Display → Header Displays 中启用
- **CJK / 输入法支持** —— 完整支持中文 / 日文 / 韩文的组合输入
- **操作系统通知与主机名感知标题** —— 桌面提醒与标签标题以 `codeman:<host>` 为前缀,使多主机配置不再含糊
---
## 隔离的 Docker 会话
让案例(case)运行在专属的加固 Docker 容器里,而不是直接跑在主机上:获得安全隔离、可复现的工具链和一键可移植性。
- **一键启动** —— 在 **New Case → Create New** 中勾选 **🐳 Run in an isolated Docker container**。Codeman 会创建案例文件夹、用默认设置启动容器,并在容器内启动智能体。无需填写任何主机/镜像/网络字段。
- **资源模板** —— 展开复选框可选 **Small / Medium / Large / GPU** 预设(内存、CPU、GPU),也可以完全自定义。**磁盘是弹性的** —— 存储随数据增长,没有固定上限。
- **按案例共享容器** —— 多个会话可以 `docker exec` 进同一个容器;结束某个会话绝不会影响其他会话所在的容器。
- **默认加固** —— 非 root、`--cap-drop ALL`、`no-new-privileges`、PID/内存上限,绝不使用 `--privileged` 或 docker socket;**密封(sealed)** 配置(不注入主机凭据、关闭网络)只需一个开关。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Antigravity / Gemini / OpenCode / Pi 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **迁移到另一台机器** —— 把容器的完整环境(工具链 + 工作区)导出为可移植的 `.tar.gz`,在另一台机器上导入到新案例即可继续。
- **持久耐用** —— Codeman 重启后重连会回到同一个存活的智能体;容器停止/重启后则从绑定挂载的转录恢复对话。
前置条件:只需 Docker(或 Podman)。智能体基础镜像会在首次使用时自动构建,构建进度实时显示在 UI 中(也可用 `node scripts/build-agent-image.mjs` 预构建)。完整指南:[`docs/docker-cases.md`](docs/docker-cases.md)。
---
## 远程 SSH 会话
把案例(case)指向另一台机器,通过 SSH 让智能体**在那台机器上**运行,同时保留同样的仪表盘、移动端 UI 与自主运行特性。你的笔记本只是一扇窗口,会话本体活在远程主机上。
- **天生持久**:智能体运行在远程主机上一个专用的 tmux 会话里,SSH 断连、网络切换或笔记本休眠都不会中断任务。重新连接后回到同一个活跃对话。
- **自动重连**:一个带上限退避的监视器发现 SSH 面板断开后,会静默重新附着到仍在运行的远程会话(设置中有总开关;主动杀掉的会话绝不会被复活)。
- **发现与附着**:列出主机上已在运行的 `codeman-*` 会话(由那台机器自己的 Codeman 或其他操作者启动)并附着其一。非你所有的已附着会话在关闭标签时**只分离,绝不杀掉**。
- **共享会话**:多个客户端可以以不同窗口尺寸同时附着同一个远程会话而互不挤压;发现列表会显示带客户端计数的「shared」徽标。
- **注入安全**:所有 ssh 命令行都经由单一的 shell 转义构建器生成,主机/路径/身份文件字段均有模式校验。
在 **New Case → Remote** 中配置(主机、用户、身份文件、可选跳板机)。完整设计:[`docs/remote-sessions.md`](docs/remote-sessions.md)。
---
## 多用户模式(可选启用)
与一个小型互信团队共享同一个 Codeman,每人拥有自己的登录与工作空间。**默认关闭**:不加该开关时,行为与单用户完全一致。
用 `codeman web --multiuser`(或 `CODEMAN_MULTIUSER=1`)启用。创建第一个管理员后,可通过 CLI 或 App Settings 中的 **Users** 标签页管理用户:
```bash
codeman users add alice --admin # 提示输入密码(或 --password-stdin)
codeman users add bob # 普通用户
codeman users list
```
- **按用户的空间**:每个用户的案例位于 `~/codeman-users/<name>/cases`;会话、案例、搜索与实时事件都按属主隔离。管理员可以看到全部。
- **可单独吊销的登录**:命名用户的密码以 scrypt 哈希保存在 `~/.codeman/users.json`;可随时禁用、重置(一次性密码)或删除账号。管理员操作审计记录在 `~/.codeman/admin-audit.jsonl`。
- **普通用户的更安全默认值**:非管理员以 `--permission-mode auto` 运行 Claude(Anthropic 的分类器护栏模式);raw shell 会话、cron `launchCommand` 与跳过权限模式需要按用户显式授权。
> ⚠️ **这只是工作空间的划分,不是用户之间的沙箱。** 所有会话都以同一个操作系统账户运行,因此有心用户的智能体依然能触及他人的文件。若需要真正的隔离,请结合 **Docker 案例**,或在不同的操作系统账户下运行独立实例。参见 [`docs/multi-user-plan.md`](docs/multi-user-plan.md) 与 [`docs/security-architecture.md`](docs/security-architecture.md) 的多用户章节。
---
## 远程访问 —— Cloudflare 隧道
使用免费的 [Cloudflare 快速隧道](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/do-more-with-tunnels/trycloudflare/),从手机或本地网络外的任意设备访问 Codeman —— 无需端口转发、无需 DNS、无需静态 IP。
```
浏览器(手机/平板)→ Cloudflare 边缘(HTTPS)→ cloudflared → localhost:3000
```
**前置条件:** 安装 [`cloudflared`](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/downloads/) 并在环境中设置 `CODEMAN_PASSWORD`。
```bash
# 快速开始
./scripts/tunnel.sh start # 启动隧道,打印公网 URL
./scripts/tunnel.sh url # 显示当前 URL
./scripts/tunnel.sh stop # 停止隧道
./scripts/tunnel.sh status # 服务状态 + URL
```
脚本会在首次运行时自动安装一个 systemd 用户服务。隧道 URL 是一个随机生成的 `*.trycloudflare.com` 地址,每次隧道重启都会改变。
<details>
<summary><strong>持久隧道(重启后存续)</strong></summary>
```bash
# 启用为持久服务
systemctl --user enable codeman-tunnel
loginctl enable-linger $USER
# 或通过 Codeman Web UI:Settings → Tunnel → 切换为开
```
</details>
<details>
<summary><strong>认证</strong></summary>
1. 首次请求 → 浏览器弹出 Basic Auth 提示(用户名:`admin` 或 `CODEMAN_USERNAME`)
2. 成功后 → 服务端签发 `codeman_session` cookie(24 小时 TTL,活动时自动延长)
3. 后续请求通过 cookie 静默认证
4. 同一 IP 失败 10 次 → 429 速率限制(15 分钟衰减)
通过隧道暴露前**务必设置 `CODEMAN_PASSWORD`** —— 否则任何拿到 URL 的人都能完全访问你的会话。
</details>
### 二维码认证
在手机键盘上输密码很糟糕。Codeman 用**短暂的一次性二维码令牌**解决这个问题 —— 扫描桌面上的二维码,手机即刻完成认证。无密码提示、无打字、无剪贴板。
```
桌面显示二维码 → 手机扫描 → GET /q/Xk9mQ3 → 服务端校验
→ 令牌原子性消费(一次性) → 签发会话 cookie → 302 跳转到 /
→ 桌面收到通知:「设备已通过二维码认证」 → 自动生成新二维码
```
只拿到裸隧道 URL(没有二维码)的人,仍会撞上标准密码提示。二维码是快速通道;密码是回退方案。
#### 工作原理
服务端维护一个轮换的、短生命周期、一次性令牌池。每个令牌由一个 256 位密钥(`crypto.randomBytes(32)`)和一个用作 URL 路径中不透明查找键的 6 字符 base62 短码配对组成。二维码编码的 URL 形如 `https://abc-xyz.trycloudflare.com/q/Xk9mQ3` —— 短码是指针,而非密钥本身,因此它绝不会通过浏览器历史、`Referer` 头或 Cloudflare 边缘日志泄露。
每 **60 秒**,服务端自动轮换到一个全新令牌。上一个令牌会保留 **90 秒的宽限期**,以处理你刚好在轮换瞬间扫描的竞争情况 —— 此后即作废。每个令牌都是**一次性**的:手机一旦成功扫描,令牌就被原子性消费,并立即为桌面显示生成一个新的。
#### 安全设计
该设计参考了 ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin)(USENIX Security 2025),该研究发现 Top-100 网站中有 47 个因横跨 42 个 CVE 的 6 个关键设计缺陷而易受二维码认证攻击。Codeman 全部六个都做了应对:
| USENIX 缺陷 | 缓解措施 |
| ---------------------------- | --------------------------------------------------------------------------------------------- |
| **缺陷 1**:缺少一次性强制 | 令牌首次扫描即原子性消费 —— 重放永远失败 |
| **缺陷 2**:长生命周期令牌 | 60s TTL + 90s 宽限,由定时器自动轮换 |
| **缺陷 3**:可预测的令牌生成 | `crypto.randomBytes(32)` —— 256 位熵。短码采用拒绝采样以消除取模偏差 |
| **缺陷 4**:客户端令牌生成 | 仅服务端 —— 令牌在嵌入二维码前绝不离开服务器 |
| **缺陷 5**:缺少状态通知 | 桌面提示:_「设备 [IP] 已通过二维码认证(Safari)。不是你?[吊销]」_ —— 实时 QRLjacking 检测 |
| **缺陷 6**:会话绑定不足 | 存储 IP + User-Agent 以供审计。通过 API 手动吊销会话。HttpOnly + Secure + SameSite=lax cookie |
#### 时序安全的查找
短码存储在 `Map<shortCode, TokenRecord>` 中。校验使用 `Map.get()` —— 一个基于哈希的 O(1) 查找,不会通过响应时延泄露目标字符串的任何信息。热路径上任何地方都没有逐字符字符串比较,彻底消除了时序侧信道攻击。
#### 速率限制(双层)
二维码认证有自己的速率限制,与密码认证完全独立:
- **按 IP**:同一 IP 失败 10 次二维码尝试即触发 429 封锁(15 分钟衰减窗口)—— 与 Basic Auth 的失败计数器分开,因此打错密码不会消耗你的二维码额度
- **全局**:所有 IP 合计每分钟 30 次二维码尝试 —— 抵御分布式暴力破解。考虑到 62^6 = 568 亿种可能短码、任意时刻仅约 2 个有效,无论如何暴力破解都在计算上不可行
#### 二维码尺寸优化
URL 被刻意保持精简(`/q/` 路径 + 6 字符码 ≈ 53–56 个字符),以瞄准 **QR 版本 4**(33×33 模块)而非版本 5(37×37)。更小的二维码在低端手机上扫描更快 —— 现代设备读取版本 4 仅需 100–300 毫秒。`/q/` 前缀相比 `/qr-auth/` 省下 7 个字节,仅此一项就足以决定二维码版本的差别。
#### 桌面体验
二维码显示每 60 秒通过 SSE 自动刷新,SVG 直接嵌入事件载荷(约 2–5KB)—— 无需额外 HTTP 请求,刷新低于 50ms。倒计时器显示剩余时间。「重新生成」按钮可即时使所有现有令牌失效并创建一个新的(在你怀疑二维码被拍照时很有用)。
当有人通过二维码认证时,桌面会弹出一个带设备 IP 与浏览器信息的通知 —— 如果不是你,一键即可吊销所有会话。
#### 威胁覆盖
| 威胁 | 为何无效 |
| ----------------------- | ------------------------------------------------------------------------------------ |
| **二维码截图被分享** | 一次性:首次扫描即消费。60s TTL:攻击者动手前已过期。桌面通知会立即提醒你。 |
| **重放攻击** | 原子性一次性消费 + 60s TTL。旧 URL 始终返回 401。 |
| **Cloudflare 边缘日志** | 短码是不透明的 6 字符查找键,而非真正的 256 位令牌。一次性意味着从日志重放永远失败。 |
| **暴力破解** | 568 亿种组合、任意时刻约 2 个有效、双层速率限制,早在统计可行性之前就已拦截。 |
| **QRLjacking** | 60s 轮换迫使实时转发。桌面提示提供即时检测。自托管单用户场景使钓鱼难以成立。 |
| **时序攻击** | 基于哈希的 Map 查找 —— 无字符串比较时序泄露。 |
| **会话 cookie 窃取** | HttpOnly + Secure + SameSite=lax + 24h TTL。可在 `POST /api/auth/revoke` 手动吊销。 |
#### 横向对比
| 平台 | 模型 | 对比 |
| ---------------- | -------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Discord** | 长生命周期令牌、无确认、[屡被利用](https://owasp.org/www-community/attacks/Qrljacking) | Codeman:一次性 + TTL + 通知 |
| **WhatsApp Web** | 手机确认「关联设备?」,约 60s 轮换 | 轮换相当;WhatsApp 额外加了显式确认(对单用户而言是可接受的取舍) |
| **Signal** | 临时公钥、端到端加密信道 | 加密更强,但 [2025 年仍被俄罗斯国家级行为者](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger)通过社会工程攻破 |
> 完整设计理由、安全分析与实现细节:[`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
---
## 安全
Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI 在设计上对任何能访问到它的人都是一个远程代码执行面 —— 整套安全模型的存在就是为了控制*谁*能访问。(启动权限模式可配置,见下文。)近期加固(v0.9.0 + v0.9.5)封堵了那些常困扰自托管开发工具的浏览器驱动攻击路径。完整模型:[`docs/security-architecture.md`](docs/security-architecture.md)。**发现了漏洞?** 私下披露方式与已知限制清单见 [`SECURITY.md`](.github/SECURITY.md)。
### 网络与访问
- **默认仅环回** —— 绑定 `127.0.0.1`,仅可从本机访问,因此「无密码」默认配置开箱即安全。在未设置 `CODEMAN_PASSWORD` 的情况下绑定非环回主机会*启动但打印一条醒目警告*,并给出三个具体修复方案(设置密码、环回 + 一个带认证的隧道,或用 `--allow-unauthenticated-network` 显式确认)
- **可选认证,真实会话** —— 通过 `CODEMAN_USERNAME`(默认 `admin`)/ `CODEMAN_PASSWORD` 的 HTTP Basic 认证。成功后签发一个不透明的 256 位 `codeman_session` cookie(`randomBytes(32)`)—— 服务端校验,而非客户端签名,因此无法离线伪造(24h TTL、自动延长、设备上下文审计日志)
- **按 IP 速率限制** —— 失败 10 次 → `429` 并带 `Retry-After`(15 分钟衰减)。即便攻击者在同一 IP 上猛攻,有效 cookie 或正确密码也能*立即*恢复 —— 这很重要,因为所有隧道流量共享同一个环回 IP。二维码认证有自己独立的限制器
- **可配置的权限模式**:`--dangerously-skip-permissions` 只是默认值。**App Settings → Claude CLI → Startup Mode** 可以把新会话切换为 Anthropic 的分类器护栏 `auto` 模式(低打扰,需要 Claude Code 2.1.207+)、`normal` 提示模式,或一份显式的允许工具列表。多用户模式下,未获授权的用户会被强制为 `auto`,shell 会话与跳过权限需要按用户显式授权
### 始终开启的浏览器加固(v0.9.5)
以下对**每个**请求都生效 —— 在认证之前,即便是默认的无密码环回安装:
- **Host 头允许列表 → 阻断 DNS 重绑定。** 一个被重绑定到 `127.0.0.1` 的自定义域名会在任何处理器运行前被 `403 host not allowed` 拒绝。允许:`localhost`、任意 IP 字面量、绑定主机、`.ts.net` / `.trycloudflare.com` / `.cfargotunnel.com`、当前受管隧道,以及 `CODEMAN_ALLOWED_HOSTS`(在此添加自定义反向代理域名 —— 逗号分隔;精确主机或前导点 `.suffix` 匹配子域名)
- **跨站 Origin / CSRF 防护。** 对变更状态的方法(`POST`/`PUT`/`PATCH`/`DELETE`),`Origin` 必须通过同一允许列表,否则返回 `403 cross-site request blocked`。*缺失*的 Origin 被允许(因此 `curl`、CLI 与 Claude Code hook 仍可工作);只有存在但外来、或不透明的 `null` origin 才会被拒绝
- **原始 `text/plain` 请求体。** 全局解析器不再对 `text/plain` 做 JSON 解析,封堵了那个跨站 `fetch` 能在无预检的情况下把 JSON 走私进写路由的 CORS「简单请求」CSRF 向量
- **WebSocket Origin 校验。** 终端 WS 升级运行同样的 Host + Origin 检查,失败时以代码 `4003` 关闭(反 CSWSH)
- **XSS 转义的智能体输出。** AI 衍生的字符串(工具名、命令参数、子智能体描述)在渲染进子智能体 / 活动面板前,于每个注入点都做 HTML 转义
### 输入、文件与响应头
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 50 MB 原始与下载;`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
- **安全响应头** —— `Content-Security-Policy`(`default-src 'self'`,每个例外都逐条列举)、`X-Content-Type-Options: nosniff`、`X-Frame-Options: SAMEORIGIN`、HTTPS 下的 HSTS,以及**仅**对 `localhost` / `127.0.0.1` / `::1` 反射的 CORS
### 供应链与隔离
- **锁定并校验的依赖** —— 安全敏感的传递依赖通过 npm `overrides` 强制为已打补丁版本;每次提交/PR 都检查锁文件完整性(所有条目都解析到 `registry.npmjs.org` 且带 `sha512` 哈希)。公共资源在 CI 中做 NUL 字节扫描与 `node --check` 校验
- **多实例隔离** —— `CODEMAN_INSTANCE` 同时限定 tmux 套接字(`-L codeman-<name>`)与数据目录(`~/.codeman-<name>`),因此两个实例绝不会互相附着对方的活动会话
> 移动端登录使用一次性、60 秒二维码令牌 —— 完整设计见上文[二维码认证](#二维码认证)(它应对了 USENIX Security 2025 二维码登录研究中的全部 6 个缺陷)。
---
## SSH 替代方案(`sc`)
如果你更喜欢 SSH(Termius、Blink 等),`sc` 命令是一个便于拇指操作的会话选择器:
```bash
sc # 交互式选择器
sc 2 # 快速附着到会话 2
sc -l # 列出会话
```
单数字选择(1–9)、颜色编码的状态、token 计数、自动刷新。用 `Ctrl+A D` 分离。
---
## 键盘快捷键
> Ctrl 绑定在 macOS 上也接受 Cmd。
| 快捷键 | 动作 |
| ------------------------------- | -------------------------------------------------------- |
| `Ctrl/Cmd+W` | 杀掉当前会话 |
| `Ctrl/Cmd/Option+K` | 查找已打开的会话或新建一个 |
| `Ctrl/Cmd+Tab` | 下一个会话 |
| `Alt/Option+[` / `Alt/Option+]` | 上一个 / 下一个会话 |
| `Alt/Option+1`–`Alt/Option+9` | 切换到第 N 个标签(按物理键位,macOS Option 布局也适用) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | 将当前标签左移 / 右移 |
| `Ctrl/Cmd+C` | 复制选中内容;未选中时中断代理 |
| `Ctrl+Shift+C` | 复制选中内容(永不中断) |
| `Ctrl/Cmd+L` | 清屏 |
| `Ctrl+Shift+R` | 恢复终端尺寸 |
| `Ctrl+Shift+V` | 切换语音输入 |
| `Ctrl/Cmd +` / `-` | 字体大小 |
| `Ctrl/Cmd+?` | 键盘帮助 |
| `Shift+Enter` | 插入换行(发送到终端) |
| `Escape` | 关闭面板与模态框 |
---
## 从智能体驱动 Codeman —— 编程指南
面向不经浏览器控制 Codeman 的 AI 智能体与自动化:一个拉起工作会话的智能体、一个 CI 机器人,或是**运行在 Codeman 会话*内部*、编排其他会话的 Claude Code**。UI 能做的一切都是 HTTP + CLI,因此智能体也能做。
> **捷径:装上打包好的智能体技能。** 下面这一整套(外加多工作会话的实战配方)已经作为 Claude Code 技能随仓库发布在 [`skills/codeman`](skills/codeman/SKILL.md),会话内部的智能体不必等你把文档粘进提示词就能驱动 Codeman。三种获取方式:
>
> - `npx skills add Ark0N/Codeman --skill codeman -g`:全局安装,任何支持技能的智能体都能用
> - `codeman skill install`(全局)或 `codeman skill install --case <name>`:给那些从 npm 安装、从未克隆过仓库的用户;`codeman skill uninstall` 可撤销
> - **App Settings → Agent Skill**(`agentSkillEnabled`,默认关闭):开启后,Codeman 会在每次于某个 case 中创建 Claude 会话时把技能注入该 case;case 里用户自己写的 `skills/codeman` 永远不会被覆盖
>
> 全局安装(`codeman skill install` 或 `npx skills add`)会被**本机每一个新建的 Claude Code 会话**读到,无论它在不在 Codeman 里。技能自带门禁:不在 Codeman 会话中(`CODEMAN_MUX` 未设置)时它拒绝动作,所以全局装上它对无关会话没有代价。
>
> ⚠️ 把 `agentSkillEnabled` 关回去**不会删掉已经注入的副本**(在创建时做清扫,会把技能从共用同一个 `.claude/` 目录的其他活动会话脚下抽走)。要删就按 case 删:`codeman skill uninstall --case <name>`。
### 检测自己身处 Codeman 内部
当 CLI 运行在 Codeman 受管会话中时,以下环境变量会被设置 —— 读取它们,别硬编码任何东西:
| 变量 | 含义 |
| -------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| `CODEMAN_MUX=1` | 你在一个受管 tmux 会话里。**绝不要** `tmux kill-session` / `pkill claude` / `pkill tmux` —— 你会杀掉自己或兄弟会话。 |
| `CODEMAN_API_URL` | API 的基础 URL(例如 `https://127.0.0.1:3000`)。下面每个调用都用它。 |
| `CODEMAN_SESSION_ID` | *你自己的*会话 id。用它避免对自己下手。 |
| `CODEMAN_HOOK_SECRET_FILE` | hook 密钥文件的路径(受管隧道开启时调用 `/api/hook-event` 必需)。 |
### 行路规则(POST 之前先读)
1. **只发单行输入,而且必须以 `\r` 结尾。** 编程输入按字面文本发送,**只有当输入里含回车符时才会触发 Enter**:`{"input":"run tests\r"}`。少了 `\r`,文本就停在会话的输入框里不被提交(同一次调用里的 `wait` 还会在一个压根没开始的回合上耗满整个超时)。内嵌的换行会被剥掉而不是报错,因此 `"echo A\necho B\r"` 执行的是拼起来的 `echo Aecho B`:一次调用只发一行。
2. **让输入幂等。** 在 `POST …/input` 上带上稳定的 `clientId` 和按会话单调递增的 `seq`。服务端会去重,因此连接中断后的重试不会重复投递提示。
3. **认证。** 若设置了 `CODEMAN_PASSWORD`,发送 HTTP Basic 认证(用户 `admin` 或 `CODEMAN_USERNAME`)或 `codeman_session` cookie。默认的环回安装无密码。缺失的 `Origin` 头被允许,因此普通 `curl` 可用;跨站的浏览器 origin 会被拒绝(CSRF 防护)。⚠️ `401` 回的是裸字符串 `Unauthorized`,**不是** JSON 信封,直接喂给 `jq` 只会抛解析错误而看不到真正的失败原因:先看状态码,再解析。
4. **响应信封。** 多数端点返回 `{ "success": true, "data": … }`(错误:`{ "success": false, "error", "errorCode" }`)。少数遗留 GET 返回裸响应体 —— **两种都要处理**(`body.data ?? body`)。
5. **`/api/v1/*`** 是 `/api/*` 的稳定别名。
6. **用等待代替轮询,别把超时当成错误。** 等待类端点在没等到事情发生时也以 HTTP `200` 加 `wait.timedOut: true` 应答,所以要循环调用短等待(默认 60 秒),而不是发一个超长的调用:隧道会掐断空闲连接。`wait.timeoutMs` 告诉你服务端钳制之后真正采用的超时(上限 600 秒)。
7. **只有 `claude` 会话会发出 `stop` 与 `blocked`。** 这两个来自 Claude Code hook;`shell` 与外部 CLI(opencode/codex/gemini/antigravity/pi)只接受 `idle`、`working` 与 `exit`。在这些模式上显式索要 `stop` 会得到 `400`;不传 `until` 则永远安全。⚠️ `shell` 会话的 `idle` 只在启动时触发**一次**,此后再也不会,所以在那里用「发送并等待」只能等到超时:没有 hook 的会话请用 `wait-output` 标记来同步。
8. **没有任何东西会报告「就绪」,得自己显式等。** 新会话在 PID 出现之前一律回答 `{"signal":"exit","immediate":true}`(意思是*还没启动*,不是*崩了*),而全新 case 里的 `claude` 工作会话接着会停在 CLI 的信任对话框上。此时给它发提示,等待会在约 2 秒后因 `idle` 解除,看上去和一个跑完的回合一模一样,而文本其实卡在对话框里。下面的配方 2b 就是避开它的顺序。
### 常用配方
```bash
# 每个 Codeman 会话里都自动设好了 CODEMAN_API_URL,协议也是对的。
# 下面的兜底值适用于标准安装;在 --https 安装上请自己写 https:// 的地址,
# 并给每个 curl 加上 -k(自签名证书)。
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (若设置了密码,给每个调用加上 -u admin:"$CODEMAN_PASSWORD")
# 1. 看看有什么在运行
curl -s "$API/api/sessions" | jq '.data // .'
# 2. 拉起一个工作会话(「case」= 命名工作目录)
curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
# 2b. 等这个工作会话真正就绪(见规则 8):先探输入框的标记,信任对话框只作兜底。
# (反过来先探信任对话框、再盲发一个 Enter,在重复运行时会误伤:对话框的文字
# 会一直留在缓冲区里,探测因此匹配到旧文本,而那个 Enter 落进了已经就绪的输入框。)
# 匹配单个词:TUI 的文字到达匹配器时可能已经丢掉了词间空格。
until [ "$(curl -s "$API/api/sessions/$SID" | jq '.data.pid')" != null ]; do sleep 1; done
R=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=trust' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true}' # 接受首次运行的信任对话框
curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' >/dev/null
fi
# 3. 向会话发送提示(精确一次:clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,"clientId":"agent-1","seq":1}'
# 4. 发送提示并阻塞到这一回合结束(先注册等待再写入,因此不会拿上一回合的状态来应答)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,
"clientId":"agent-1","seq":2,"wait":"stop,exit","waitTimeout":60000}' \
| jq '.data.wait' # -> {"signal":"stop","timedOut":false,"waitedMs":41230,...}
# (`stop` 是回合结束的权威 hook。加上 `idle` 会让它在转圈停顿时也解除,
# 任何重画出 ❯ 提示符的东西同理,比如一个对话框。)
# 4b. 超时了?那是 200,不是失败。循环调用短等待即可。
curl -s "$API/api/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq '.data.wait'
# 4c. 或者等输出里出现某个标记(shell 会话也适用)。
# ⚠️ 每次调用都要用不同的标记(tmux 重画会重放旧屏幕文字),并且把标记拆开写,
# 让敲进去的那一行本身不包含它:你自己的按键会回显进输出流,不拆开的标记会在
# 命令还没跑之前就匹配上。from=buffer 用来接住在等待落地之前就已打印的标记。
N=$RANDOM
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
curl -sG "$API/api/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
# 5. 读回答案。claude / codex 会话用 last-response:它取自 transcript 而不是屏幕,
# 因此不带 TUI 的画框与重画噪声。⚠️ 要轮询,别只读一次:transcript 落盘比 stop
# 信号稍晚,紧跟着「发送并等待」返回后立刻读,常常拿到空串。
for _ in $(seq 1 10); do
TXT=$(curl -s "$API/api/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
# 5b. 其他模式(shell/opencode/gemini/antigravity/pi)没有 transcript,读终端。
# ⚠️ 用 terminal?tail=,不要用 /output:后者的 textOutput 对每个由 tmux 承载的
# (也就是每个交互式)会话都是空的。tail 按字节计,返回的是含 ANSI 的终端数据。
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
# 6. 流式接收实时事件(会话输出、智能体活动、状态)
curl -sN "$API/api/events" # Server-Sent Events
# 7. 调度周期性工作(cron 风格任务)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
"promptMode":"inline_text","promptText":"Update dependencies and open a PR",
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
# 8. 查看后台子智能体及其活动记录
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
# 9. 全系统快照(会话、设置、重生、统计)
curl -s "$API/api/status" | jq
```
### 或使用内置 CLI
同样的操作也有命令形式(`codeman <cmd>`,括号内为别名)—— 在会话内的 shell 工具里很顺手:
```bash
codeman session start -d /path/to/repo # (s) 启动会话
codeman session list # 列出会话
codeman session logs <id> # 查看输出
codeman task add "fix the failing test" # (t) 排入任务
codeman attach <path> # 附着 Claude hook 上下文
```
### Hook(事件*回流*到 Codeman)
Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission_prompt`、`idle_prompt`、`stop`、`task_completed` 等),让仪表盘实时响应。该端点在环回上免认证,但在受管隧道下需要 `X-Codeman-Hook-Secret` 头(从 `$CODEMAN_HOOK_SECRET_FILE` 读取)。通常你不需要手动调用它 —— Codeman 会自动接好 —— 但自主层正是靠它「看见」智能体在做什么。
> 完整端点列表与请求/响应形状见下文。
---
## API
基于 Fastify 的 REST —— **21 个路由模块中约 200 个处理器**,外加一条 SSE 流和一条 WebSocket 终端通道。所有响应都使用 `ApiResponse<T>` 信封(`{success, data}` / `{success, error, errorCode}`);`/api/v1/*` 是稳定别名。以下是一个有代表性的子集:
### 会话(Sessions)
| 方法 | 端点 | 说明 |
| -------- | ------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `GET` | `/api/sessions` | 列出全部 |
| `POST` | `/api/quick-start` | 创建 case + 启动会话(`{caseName?, mode?, effort?, envOverrides?}`) |
| `POST` | `/api/sessions/:id/input` | 发送输入(`{input, useMux?, clientId?, seq?, wait?, waitTimeout?}`:`clientId`+`seq` = 精确一次;`wait` 阻塞到这一回合结束) |
| `GET` | `/api/sessions/:id/terminal` | 读取终端输出(`?tail=<bytes>`、`?full=1`):交互式会话的读取路径 |
| `GET` | `/api/sessions/:id/output` | 一次性的解析输出(tmux 承载的会话里 `textOutput` 为空) |
| `GET` | `/api/sessions/:id/wait` | 阻塞到某个信号触发(`?until=stop,idle,exit&timeout=&fresh=`);超时是 `200` |
| `GET` | `/api/sessions/:id/wait-output` | 阻塞到某个字面串出现(`?match=&nocase=&from=now\|buffer&timeout=`) |
| `GET` | `/api/sessions/unified` | 统一的活动 + 历史清单(会话管理器):`?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | 在会话管理器中置顶 / 取消置顶(`{pinned}`) |
| `PUT` | `/api/session-order` | 跨设备同步标签顺序(`{order: [ids]}`) |
| `DELETE` | `/api/sessions/:id` | 删除会话 |
### 重生(Respawn)
| 方法 | 端点 | 说明 |
| ------ | ---------------------------------- | -------------------- |
| `POST` | `/api/sessions/:id/respawn/enable` | 启用,带配置与定时器 |
| `POST` | `/api/sessions/:id/respawn/stop` | 停止控制器 |
| `PUT` | `/api/sessions/:id/respawn/config` | 更新配置 |
### 编排器(Orchestrator)
| 方法 | 端点 | 说明 |
| ------ | --------------------------- | --------------- |
| `POST` | `/api/orchestrator/start` | 从目标启动编排 |
| `POST` | `/api/orchestrator/approve` | 批准生成的计划 |
| `GET` | `/api/orchestrator/status` | 当前阶段 + 进度 |
| `POST` | `/api/orchestrator/stop` | 停止并清理 |
### Cron(定时任务)
| 方法 | 端点 | 说明 |
| ---------------- | ---------------------------- | --------------------- |
| `GET` / `POST` | `/api/cron/jobs` | 列出 / 创建 cron 任务 |
| `PUT` / `DELETE` | `/api/cron/jobs/:id` | 更新 / 删除任务 |
| `PUT` | `/api/cron/jobs/:id/enabled` | 启用 / 禁用 |
| `POST` | `/api/cron/jobs/:id/run` | 立即运行 |
| `GET` | `/api/cron/jobs/:id/runs` | 运行历史 |
### 子智能体(Subagents)
| 方法 | 端点 | 说明 |
| -------- | ------------------------------- | ------------------ |
| `GET` | `/api/subagents` | 列出所有后台智能体 |
| `GET` | `/api/subagents/:id` | 智能体信息与状态 |
| `GET` | `/api/subagents/:id/transcript` | 完整活动记录 |
| `DELETE` | `/api/subagents/:id` | 杀掉智能体进程 |
### 系统(System)
| 方法 | 端点 | 说明 |
| ------ | ------------------------------- | ---------------------------------------- |
| `GET` | `/api/events` | SSE 流 |
| `GET` | `/api/status` | 完整应用状态 |
| `POST` | `/api/hook-event` | Hook 回调 |
| `GET` | `/api/system/update/check` | 检查新发行版 |
| `POST` | `/api/system/update` | 自更新(git-clone 安装) |
| `POST` | `/api/clipboard` | 把文本推送到所有已连接浏览器(`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | 时间线 + 统计 |
> **想在 Codeman 之上做集成?**[`docs/extending-codeman.md`](docs/extending-codeman.md)(英文)是集成指南:把你自己的界面作为标签页嵌入、订阅 SSE 事件流以便在 agent 需要你时做出响应、用脚本驱动 Codeman,以及动手前值得先了解的那些坑。Codeman 刻意不提供插件运行时,所以一个集成就是你自己的进程在讲 HTTP。
---
## 架构
```mermaid
flowchart TB
subgraph Codeman["CODEMAN"]
subgraph Frontend["前端层"]
UI["Web UI<br/><small>xterm.js + 智能体窗口</small>"]
API["REST API<br/><small>Fastify</small>"]
SSE["SSE 事件<br/><small>/api/events</small>"]
end
subgraph Core["核心层"]
SM["会话管理器"]
S1["会话 (PTY)"]
S2["会话 (PTY)"]
RC["重生控制器"]
ORC["编排器循环"]
end
subgraph Detection["检测层"]
SW["子智能体监视器<br/><small>~/.claude/projects/*/subagents</small>"]
TW["团队监视器<br/><small>~/.claude/teams/*</small>"]
end
subgraph Persistence["持久化层"]
SCR["Mux 管理器<br/><small>(tmux)</small>"]
SS["状态存储<br/><small>state.json</small>"]
end
subgraph External["外部"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi</small>"]
BG["后台智能体<br/><small>(Task 工具)</small>"]
end
end
UI <--> API
API <--> SSE
API --> SM
SM --> S1
SM --> S2
SM --> RC
SM --> ORC
SM --> SS
S1 --> SCR
S2 --> SCR
RC --> SCR
ORC --> SCR
SCR --> CLI
SW --> BG
SW --> SSE
TW --> SSE
```
---
## 开发
```bash
npm install
npx tsx src/index.ts web # 开发模式
npm run build # 生产构建
npm run test:ci # 运行测试(CI 套件;浏览器套件需要额外环境)
```
完整文档见 [CLAUDE.md](./CLAUDE.md)。
---
## 代码库质量
本代码库经历了一次全面的 7 阶段重构,消除了上帝对象、集中了配置,并建立了模块化架构:
| 阶段 | 改了什么 | 影响 |
| ---------------- | ------------------------------------------------------------------------------------------------------------- | ------------------------------------------ |
| **性能** | 缓存端点、SSE 自适应批处理、缓冲区分块 | 终端延迟低于 16ms |
| **路由抽取** | `server.ts` 拆分为 15 个领域路由模块 + 认证中间件 + 端口接口 | server.ts 代码量 **−67%**(6,736 → 2,254) |
| **领域拆分** | `types.ts` → 16 个领域文件、`ralph-tracker` → 7 个文件、`respawn-controller` → 5 个文件、`session` → 6 个文件 | 不再有上帝文件 |
| **前端模块** | `app.js` → 18 个抽取模块,横跨基础设施、领域与特性层 | app.js 核心降至 **约 3.4K 行** |
| **配置合并** | 约 70 个散落的魔法数字 → 10 个领域聚焦的配置文件 | 零跨文件重复 |
| **测试基础设施** | 共享 mock 库、12 个路由测试文件、统一的 MockSession | 路由处理器可通过 `app.inject()` 测试 |
完整细节:[`docs/archive/code-structure-findings.md`](docs/archive/code-structure-findings.md)
---
## 已发布的包
### [`xterm-zerolag-input`](https://www.npmjs.com/package/xterm-zerolag-input)
[![npm](https://img.shields.io/npm/v/xterm-zerolag-input?style=flat-square&color=22c55e)](https://www.npmjs.com/package/xterm-zerolag-input)
为 xterm.js 提供即时按键反馈的叠加层。通过把输入的字符立即渲染为像素级精准的 DOM 叠加层,消除高 RTT 连接下的感知输入延迟。零依赖、可配置的提示符检测、带 78 个测试的完整状态机。
```bash
npm install xterm-zerolag-input
```
[完整文档](packages/xterm-zerolag-input/README.md)
---
## 版本策略
Codeman 遵循 [SemVer](https://semver.org/)。版本号真正承诺的内容,以及哪些算内部实现(HTTP/SSE API、磁盘上的状态、实验性特性),都写在 [`docs/versioning-policy.md`](docs/versioning-policy.md) 中。如果你的脚本依赖 HTTP API,请锁定到确切版本。
## 许可证
MIT —— 见 [LICENSE](LICENSE)
---
<p align="center">
<strong>跟踪会话。可视化智能体。掌控重生。让它在你睡觉时持续运行。</strong>
</p>
-28
View File
@@ -1,28 +0,0 @@
// @ts-check
import eslint from '@eslint/js';
import tseslint from 'typescript-eslint';
export default tseslint.config(
eslint.configs.recommended,
tseslint.configs.recommended,
{
rules: {
'no-console': 'off',
'no-debugger': 'error',
// Relax some rules that conflict with existing patterns
'@typescript-eslint/no-explicit-any': 'warn',
'@typescript-eslint/no-unused-vars': 'off', // TypeScript compiler already handles this
},
},
{
ignores: [
'dist/**',
'node_modules/**',
'coverage/**',
'src/web/public/vendor/**',
'src/web/public/app.js',
'scripts/**/*.mjs',
'scripts/remotion/**',
],
}
);
-16
View File
@@ -1,16 +0,0 @@
{
"$schema": "https://unpkg.com/knip@5/schema.json",
"entry": [
"scripts/*.mjs",
"scripts/*.js",
"scripts/watch-subagents.ts",
"scripts/remotion/Root.tsx",
"scripts/remotion/index.ts",
"test/**/*.test.ts",
"test/mobile/vitest.config.ts",
"test/**/*.mjs"
],
"project": ["src/**/*.{ts,tsx}", "scripts/**/*.{ts,tsx,mjs,js}", "test/**/*.{ts,mjs}"],
"ignoreExportsUsedInFile": true,
"ignoreDependencies": ["@remotion/cli", "@remotion/transitions", "esbuild", "agent-browser"]
}
-35
View File
@@ -1,35 +0,0 @@
import { resolve } from 'node:path';
import { defineConfig, configDefaults } from 'vitest/config';
const root = resolve(import.meta.dirname, '..');
/**
* CI test config — same as vitest.config.ts but EXCLUDES the browser-driven
* mobile suite (test/mobile/**). Those are Playwright visual-regression tests
* that need a live server + chromium + environment-specific PNG baselines, so
* they are run/maintained separately and are not part of the CI gate.
*
* Keep the rest in sync with config/vitest.config.ts.
*/
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: ['test/**/*.test.ts'],
exclude: [
...configDefaults.exclude,
'test/mobile/**', // browser/visual (Playwright + chromium)
'test/perf-*.test.ts', // timing-sensitive perf benchmarks (flaky in CI)
'test/inline-rename.test.ts', // browser (Playwright)
'test/opencode-resize.test.ts', // browser (Playwright)
'test/webgl-fallback.test.ts', // browser (Playwright)
'test/terminal-copy-shortcut.test.ts', // browser (Playwright)
'test/codex-predictive-echo.test.ts', // browser (Playwright) + real codex binary
],
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 30000,
teardownTimeout: 60000,
},
});
-26
View File
@@ -1,26 +0,0 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
const root = resolve(import.meta.dirname, '..');
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: ['test/**/*.test.ts'],
setupFiles: ['./test/setup.ts'],
// Run test files sequentially to respect mux session limits
// Individual tests within files still run in parallel where safe
fileParallelism: false,
coverage: {
provider: 'v8',
reporter: ['text', 'json', 'html'],
include: ['src/**/*.ts'],
exclude: ['src/index.ts', 'src/cli.ts'],
},
testTimeout: 30000, // 30 seconds for integration tests
// Ensure cleanup runs even on test failures
teardownTimeout: 60000,
},
});
-84
View File
@@ -1,84 +0,0 @@
# Codeman agent base image (built locally by scripts/build-agent-image.mjs).
#
# Contains the agent toolchain (node + the CLIs + git/tmux/ripgrep) but NO
# secrets: credentials are delivered at RUNTIME via bind mounts (~/.claude etc.)
# or name-only `docker exec --env`, never baked in, so `docker save` exports stay
# secret-free. tmux is a HARD prerequisite (the in-container tmux is what makes a
# reconnect durable), so it is installed here and probed before launch.
#
# HOME is made writable by an ARBITRARY host uid via the OpenShift "gid 0,
# group-writable" convention: on Linux we run `--user <hostUid>:0`, so the agent
# uid is the host uid (workspace files stay host-owned) while gid 0 keeps $HOME
# writable even though the uid is not the baked 1000.
FROM node:22-bookworm-slim
# Base toolchain. `curl` is needed for the hook callbacks (`curl -sk $CODEMAN_API_URL`),
# `procps` for `ps`, `tmux` for the durable in-container session.
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
git \
tmux \
ripgrep \
curl \
ca-certificates \
less \
procps \
openssh-client \
&& rm -rf /var/lib/apt/lists/*
# The npm-published agent CLIs. Pinning is left to the rebuild cadence (see
# docs/docker-cases-plan.md, user-decision 2).
RUN npm install -g \
@anthropic-ai/claude-code \
@openai/codex \
@google/gemini-cli \
opencode-ai \
&& npm cache clean --force
# Antigravity (`agy`) is NOT on npm — Google ships a standalone binary through its
# own installer, so it needs its own step. `--dir /usr/local/bin` is load-bearing:
# the installer's default target is `$HOME/.local/bin`, which at build time is
# root's home and would be unreachable by the `agent` user the container runs as.
# ⚠️ This binary is ~190MB on its own; it is the single largest layer in the image.
RUN curl -fsSL https://antigravity.google/cli/install.sh | bash -s -- --dir /usr/local/bin \
&& chmod 755 /usr/local/bin/agy \
&& agy --version
# Pi (pi.dev). Upstream documents --ignore-scripts (pi needs no lifecycle scripts);
# kept out of the shared npm block above so the flag cannot silently change how the
# other four CLIs install.
RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
&& npm cache clean --force \
&& pi --version
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
# only matters for a hand-run / Docker Desktop container. gid 0 + group-writable
# HOME (OpenShift arbitrary-uid convention) keeps $HOME writable for any uid.
# UTF-8 locale so tmux/Ink render Unicode box-drawing instead of VT100 ACS `q`
# glyphs (C.UTF-8 is built into glibc; no locales package needed). Codeman also
# sets these at run time so containers built before this line still get UTF-8.
ENV LANG=C.UTF-8 LC_ALL=C.UTF-8
ENV HOME=/home/agent
# `.claude` (+ `.claude/projects` mount point) and `.codex` (+ `.codex/sessions`) are
# pre-created gid-0 group-writable so the container owns its OWN credential config
# dirs: tokens/settings/config are seeded in as writable copies and each CLI's runtime
# state (backups, tasks, refreshed tokens) stays container-local, while ONLY the shared
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
# `.pi/agent` IS pre-created: pi is seeded per-FILE (auth/settings/trust/models), and a
# per-file seed copy, unlike a whole-dir one, does not create its parent directory.
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
/home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent \
&& chgrp -R 0 /home/agent \
&& chmod -R g=u /home/agent
USER agent
WORKDIR /home/agent
# Codeman overrides the command with `sleep infinity` at create time; this is the
# fallback so a hand-run container also idles rather than exiting.
CMD ["sleep", "infinity"]
-104
View File
@@ -1,104 +0,0 @@
# SPEEDRUN.md — Fast-execution protocol for Claude
Read this when the goal is **throughput**: get correct, verified work done with
minimum ceremony. This does **not** relax correctness or the safety rules in
`CLAUDE.md` — those still win. It removes _waste_, not _rigor_.
> Precedence: `CLAUDE.md` > explicit user instructions > this file. If anything
> here conflicts with `CLAUDE.md`, `CLAUDE.md` wins.
---
## The mindset
- **Act, don't announce.** No "I'm going to now…" preamble. Do the thing, report
the result.
- **Cheapest proof that the change works.** Pick the smallest check that actually
demonstrates correctness — not the biggest.
- **Batch aggressively.** Independent reads, greps, and edits go in **one**
message with parallel tool calls. Never serialize work that has no dependency.
- **Momentum over perfection.** Land a correct increment, verify it, move on.
Don't gold-plate untouched code.
---
## Loop (repeat until done)
1. **Orient once** — one parallel burst of reads/greps to load the context you
need. Don't re-read files the harness says are already current.
2. **Change** — make the edit(s). Batch independent edits.
3. **Verify cheaply** — the smallest check that proves _this_ change (see below).
4. **Advance** — next item. Only re-verify what you touched.
5. **Stop** at: list empty, a hard blocker, or a decision that's genuinely the
user's to make.
---
## Verification ladder — climb only as high as the change needs
| Change kind | Cheapest sufficient check |
|-------------|---------------------------|
| Types / signatures / imports | `tsc --noEmit` (or `--watch` already running) |
| One module's logic | `npm test -- test/<file>.test.ts` (the **one** relevant file) |
| A named behavior | `npm test -- -t "pattern"` |
| Route/handler | `app.inject()` route test, or one `curl` against the running dev server |
| Frontend render | Playwright load + assert (`waitUntil: 'domcontentloaded'`, wait 3–4s) |
| Broad / pre-merge | `npm run test:ci` (the CI-equivalent sweep) |
**Hard rules (never skip, even in a rush):**
- ⚠️ **Never run bare `npm test`** — it pulls in browser/visual suites that hang
or fail locally. Always pass a file or `-t`, or use `test:ci`.
- ⚠️ **Never COM without verifying the change actually works** first (curl the
endpoint / Playwright the UI). "Compiles" ≠ "works".
- ⚠️ **Session safety** — check `$CODEMAN_MUX`; never `tmux kill-session` /
`pkill claude` in a managed session.
- ⚠️ **Single-line prompts** for any programmatic session input.
---
## Speed tactics that pay off here
- **Parallel exploration**: dispatch `Explore` subagents (or one parallel grep
burst) instead of serial file-by-file reading when scope is uncertain.
- **`tsc --noEmit --watch`** in the background — instant type feedback, no repeat
cold starts.
- **Target one test file** — `fileParallelism: false` means the suite is serial;
running one file is dramatically faster than the sweep.
- **`curl localhost:3000/api/...`** beats spinning up a browser for backend
checks. Reserve Playwright for actual UI rendering.
- **Trust the harness** — if it says a file you just edited is current, don't
re-Read it to "confirm". The Edit already succeeded or it would have errored.
---
## Anti-patterns (these masquerade as speed, but cost time)
- Running the full test suite to check a one-file change.
- Re-reading files you already have in context.
- Narrating a plan you're about to execute anyway.
- Serial tool calls that have no dependency between them.
- Claiming "done / fixed / passing" **before** running the check that proves it.
- Deploying (COM) on green typecheck alone, without exercising the real flow.
---
## Stop-conditions (don't rush past these)
Stop and surface, don't guess, when you hit:
- A **destructive / hard-to-reverse** action (delete, overwrite, force-push).
- An **outward-facing** action (publishing, sending, deploying) not already
authorized.
- A **genuine product decision** the code can't answer.
- A **failing verification you can't explain** — debug it (see
`superpowers:systematic-debugging`), don't paper over it.
---
## Definition of done
A task is done when **all** hold:
- The change is made.
- The cheapest sufficient check **ran** and **passed** — evidence, not assertion.
- No new type errors / lint errors introduced (`tsc --noEmit`, `npm run lint`).
- You state plainly what was done and what proved it. If a step was skipped or a
test failed, say so — don't hedge, don't overclaim.
-759
View File
@@ -1,759 +0,0 @@
# Agent Control Plan: skill packaging + wait primitives
**Status**: steps 1 to 8 DONE and RELEASED. The wait primitives and the skill itself
(steps 1 to 5) shipped in **1.13.0**; the `codeman skill install` CLI, per-case injection
and `agentSkillEnabled` (step 6) shipped in **1.14.1** and were republished with fixes in
**1.14.2**. Steps 1 to 5 were multi-round verified on 2026-08-08, step 6 on 2026-08-09;
see [§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
verification round found, and the two items that genuinely remain open (§2.4's footgun
guard and the Part 3 deferrals).
**Date**: 2026-08-08
**Scope**: Part 1 (agent skill) and Part 2 (wait primitives) were specified and built.
Parts 3 to 5 are captured so they are not lost, but remain deliberately deferred.
---
## 0. Where this came from: what herdr does
[herdr](https://github.com/herdrdev/herdr) (Rust, Apache-2.0, ~25.8k stars) is a terminal
multiplexer built around AI coding agents. Relevant findings from the research pass:
| Capability | How herdr does it |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Agent state | Four states (`idle`, `working`, `blocked`, `done`) that roll up pane to tab to workspace in a sidebar |
| Detection | Lifecycle hooks where the agent supports them (it names Pi and MastraCode), otherwise TOML manifests matched against a live bottom-buffer snapshot. Bundled manifests plus remote updates from herdr.dev, local overrides win |
| Control API | Newline-delimited JSON over a Unix socket (`~/.config/herdr/sessions/<name>/herdr.sock`), `{"id":"req_1","method":"pane.split","params":{}}`, dot-notation methods, plus long-lived event subscriptions |
| Discoverability | `herdr api schema` prints a machine-readable schema |
| Agent skill | `npx skills add herdrdev/herdr --skill herdr -g`, a SKILL.md wrapping the CLI, guarded by `test "${HERDR_ENV:-}" = 1` so an agent outside a herdr pane refuses to act |
| Persistence | Background server, detach with `ctrl+b q`, snapshot restore of workspaces/tabs/panes/cwd/layout, experimental screen-history replay, agent resume via native session ids, live PTY handoff across server replacement |
| Plugins | `herdr-plugin.toml` manifest, actions, event hooks, plugin panes, link handlers, GitHub-topic marketplace index |
The commands the skill teaches the agent:
| Group | Commands |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| workspace | `workspace list`, `workspace create` |
| tab | `tab list --workspace <id>`, `tab create` |
| pane | `pane current`, `pane list`, `pane layout`, `pane split --current --direction right --cwd <path> --no-focus`, `pane run <id> "<cmd>"`, `pane wait-output <id> --match/--regex <p> --timeout <ms>`, `pane read <id> --source visible\|recent\|detection` |
| agent | `agent list`, `agent start <name> --kind <type> --pane <id>`, `agent prompt <name> "<text>" --wait --timeout <ms>`, `agent wait <name> --until <state> --timeout <ms>`, `agent send-keys`, `agent get`, `agent read` |
### The honest comparison
herdr and Codeman are not the same product. herdr is a local, keyboard-first multiplexer with
no server, no web UI, and no autonomy layer. Codeman is a server with a browser and mobile UI,
remote and Docker cases, respawn, Ralph, cron, and the orchestrator, none of which herdr has.
What herdr genuinely does better is being **callable by the agent running inside it**. For
Codeman that is a packaging problem plus one missing primitive, not an architecture problem.
---
## 1. Gap analysis
| herdr capability | Codeman equivalent today | Gap |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- |
| `pane split` + `agent start` | `POST /api/quick-start`, `POST /api/sessions` | none, already there |
| `agent prompt` | `POST /api/sessions/:id/input` with `clientId`+`seq` exactly-once | no `--wait` |
| `pane read` | `GET /api/sessions/:id/output`, `GET /api/sessions/:id/terminal?full=1` | none |
| `agent list` / `agent get` | `GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status` | none |
| `agent wait --until <state>` | SSE only (`/api/events`) | **missing**, and SSE is impractical from a shell tool |
| `pane wait-output --match` | nothing | **missing** |
| Skill file | README section "Driving Codeman from an Agent" | **not packaged**, an agent will never find it |
| Env guard `HERDR_ENV=1` | `CODEMAN_MUX=1`, `CODEMAN_API_URL`, `CODEMAN_SESSION_ID` already exported at spawn | none, the guard variables exist |
| `blocked` state | hook events (`permission_prompt`, `elicitation_dialog`) plus CSS classes plus the phone overview NEEDS YOU section | not in the wire contract (`SessionStatus = 'idle' \| 'busy' \| 'stopped' \| 'error'`) |
| `api schema` | hand-written `docs/api-reference.md` | no machine-readable schema |
| Detection manifests | hardcoded in `usage-limit-patterns.ts`, `respawn-*-patterns`, `regex-patterns.ts` | patterns are code, not data |
| Plugin runtime | deliberately refused, see `docs/extending-codeman.md` | not a gap, a decision |
| Session handoff on restart | tmux owns the PTYs, so they already survive a Codeman restart | not a gap, solved by architecture |
**Conclusion**: roughly 90% of the capability surface already exists. Parts 1 and 2 below close
the two real gaps.
The table is the 2026-08-08 snapshot that motivated the work, kept as written. The three rows
marked missing are closed since: `GET .../wait` and `GET .../wait-output` shipped in 1.13.0, and
the skill is packaged at `skills/codeman` (npm tarball included). `blocked` as a wire-contract
state, and the machine-readable schema, are still open (Parts 3 and 4).
---
## 2. Part 1: the Codeman agent skill
### 2.1 Goal
An agent running inside a Codeman session can discover and correctly drive Codeman without the
user pasting API docs into the prompt, and without inventing dangerous calls.
### 2.2 Layout and distribution
The `npx skills` CLI (vercel-labs/skills) clones a GitHub repo and looks for
`skills/<name>/SKILL.md`. Claude Code natively discovers `.claude/skills/<name>/SKILL.md` in a
project and `~/.claude/skills/` globally. Both are satisfied with one source of truth plus a
symlink, which is the pattern this repo already uses for `remotion-best-practices`.
```
skills/
codeman/
SKILL.md <- single source of truth
reference/
endpoints.md <- full endpoint tables, loaded on demand
recipes.md <- worked multi-session orchestration examples
.claude/skills/codeman -> ../../skills/codeman (symlink, dogfooding in this repo)
```
Adding a `skills/` directory to the repo root costs one entry in the GitHub listing. CLAUDE.md
keeps the root short on purpose, so this needs a conscious sign-off; the alternative is
`docs/skills/codeman/` with a `--skill` path argument, which breaks the one-liner install.
**Recommendation**: accept `skills/` at the root, because the install one-liner is the whole
point of shipping a skill.
Install paths, in order of how a user gets it:
1. `npx skills add Ark0N/Codeman --skill codeman -g` (global, any agent, matches the herdr flow).
2. `codeman skill install [--global | --case <name>]`, a new CLI subcommand writing the same
file. This is the path for users who installed via npm and never cloned the repo.
3. **Automatic per-case injection**, modeled exactly on `applyStatusLineConfig(casePath, enabled)`
in `hooks-config.ts`: write `<case>/.claude/skills/codeman/SKILL.md` at case creation,
gated on a new setting. Codeman already writes `<case>/.claude/settings.local.json` hooks
through `writeHooksConfig()`, so this is the same mechanism with the same lifecycle.
Setting name: `agentSkillEnabled`. Synced (not per-device), since it changes on-disk case
content rather than display. Default: **ON after the dogfooding phase, OFF in the first
release**. Rationale for starting OFF: Claude Code loads every skill's name and description
into context on every turn, so an always-on skill has a small permanent token cost, and we
should measure that we are buying something with it first.
### 2.3 SKILL.md content
Frontmatter, per the skills convention (`name` + `description` required):
```yaml
---
name: codeman
description: >-
Control Codeman, the session manager this agent is running inside: list sessions,
start worker sessions, send prompts, read terminal output, and wait for other agents
to finish. Only usable when CODEMAN_MUX=1.
---
```
Body sections, in order:
**1. Guard (first thing, non-negotiable).**
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "not inside a Codeman session"; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set, refusing to guess}"
SELF="${CODEMAN_SESSION_ID:-}"
```
If `CODEMAN_MUX` is not `1`, the agent must stop and say it is not running inside a
Codeman-managed session. Same shape as herdr's `HERDR_ENV` guard, and the variables are
already exported by `tmux-manager.buildEnvExports()`. No fallback URL when
`CODEMAN_API_URL` is unset: any guess is the wrong scheme on an HTTPS install (prod is
HTTPS with a self-signed cert, hence `curl -sk` throughout), and a server the agent
cannot identify is not one it should be driving.
**2. Rules of the road.** Lifted and tightened from README lines 666 to 745:
- Single-line input only. Multi-line breaks the agent TUI (Ink).
- Always send `clientId` + a monotonic `seq` on `POST .../input` so a retry cannot double-deliver.
- Envelope is `{success, data}`; a few legacy GETs are bare, so read `body.data ?? body`.
- Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set. Prod is HTTPS, so `curl -sk`.
- Prefer `/api/v1/*`, the stable alias.
**3. Safety rules (the section that does not exist anywhere today).**
- Never act on `$CODEMAN_SESSION_ID`. That is you.
- Only `DELETE` sessions **you created in this conversation**, by exact id. Keep the list.
- Never bulk-delete, never loop a `DELETE` over `/api/sessions`. There is no undo.
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. Use the API.
- Creating a session consumes a slot against the 50-session cap. Clean up what you start.
**4. Recipes**, each one a single copy-pasteable curl:
| Task | Call |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| list sessions | `GET /api/v1/sessions` |
| find yourself | match ids by PREFIX of `$CODEMAN_SESSION_ID` (Docker cases truncate it to 8 chars, so an equality check never fires there) |
| start a worker | `POST /api/v1/quick-start {caseName, mode, effort}` |
| send a prompt | `POST /api/v1/sessions/:id/input {input:"…\r", useMux:true, clientId, seq}` (the trailing `\r` is what sends Enter; without it the text sits on the prompt unsubmitted) |
| send prompt and wait | `POST /api/v1/sessions/:id/input {input:"…\r", wait:"stop", waitTimeout:600000}` (Part 2) |
| wait for a worker | `GET /api/v1/sessions/:id/wait?until=stop,blocked&timeout=300000` (Part 2) |
| wait for a marker | `GET /api/v1/sessions/:id/wait-output?match=DONE_<random>&timeout=120000` (Part 2; unique per call, per §3.3's repaint rule) |
| read output | `GET /api/v1/sessions/:id/output` |
| read full scrollback | `GET /api/v1/sessions/:id/terminal?full=1` |
| watch sub-agents | `GET /api/v1/subagents` |
| schedule work | `POST /api/v1/cron/jobs` |
| clean up | `DELETE /api/v1/sessions/:id` |
**5. Pointer to `reference/endpoints.md`** for anything not in the table, so the always-loaded
part of the skill stays small.
### 2.4 An ergonomics guard worth adding server-side
The skill will tell the agent not to act on itself, but a confused agent can still try. Propose:
the skill sends `X-Codeman-Caller-Session: $CODEMAN_SESSION_ID` on every request, and the server
refuses destructive operations (`DELETE /api/sessions/:id`, kill, respawn stop) when that header
equals the target id, with a clear error.
This is a **footgun guard, not a security control**: any caller can omit the header. Document it
as such so nobody mistakes it for a boundary. It costs about 10 lines in `route-helpers.ts`.
### 2.5 Verification
Per the always-end-to-end-test rule, "the skill exists" is not done. Done is:
1. Symlink it into `.claude/skills/`, start a real throwaway Codeman session, and ask that agent
to "start a worker session that runs the test suite and tell me when it finishes".
2. Confirm from the outside that exactly one new session appeared, got the prompt, and that the
lead agent waited rather than polling in a busy loop.
3. Confirm the guard: run the same prompt in a shell with `CODEMAN_MUX` unset and confirm refusal.
4. Confirm cleanup: the worker session is deleted by exact id and no other session was touched.
Never run this against `w1`/`w2`/`w3`.
### 2.6 Files touched
- `skills/codeman/SKILL.md` (new), `skills/codeman/reference/*.md` (new)
- `.claude/skills/codeman` symlink (new)
- `src/cli.ts` (new `skill install` subcommand)
- `src/hooks-config.ts` (new `applyAgentSkill(casePath, enabled)`, mirroring `applyStatusLineConfig`)
- `src/web/schemas.ts` (`agentSkillEnabled` in `SettingsUpdateSchema`, which is `.strict()`)
- `src/web/routes/system-routes.ts` (settings PUT must resolve the flag from `merged`, never
from the raw body, per the partial-PUT invariant)
- `src/web/public/settings-ui.js` + `index.html` (checkbox)
- `package.json` `files` array, so `skills/` ships to npm
- README pointer, `docs/extending-codeman.md` seam 3 pointer
---
## 3. Part 2: wait primitives
### 3.1 Goal
Make Codeman orchestratable from a shell tool. Today the only "tell me when" channel is SSE,
which a curl-driven agent cannot practically consume: it would have to hold a streaming
connection and parse events inline. herdr solves this with blocking CLI calls. Codeman should
solve it with bounded long-poll endpoints.
All three additions are **additive**, so the versioning policy stays intact (new endpoints and
new optional fields are non-breaking).
### 3.2 The signal model
A waiter resolves on the first of a set of signals. Sources that already exist:
| Signal | Source today |
| --------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| `idle` | `Session` emits `idle` (session.ts ~1775 for Claude, ~2101 for shell), wired at `session-listener-wiring.ts:402` |
| `working` | `Session` emits `working` (session.ts ~1788), wired at `session-listener-wiring.ts:401` |
| `stop` | `POST /api/hook-event` with `event: 'stop'`, the definitive "Claude finished responding" signal already used by `controller.signalStopHook()` |
| `blocked` | `POST /api/hook-event` with `permission_prompt` or `elicitation_dialog` |
| `exit` | `Session` emits `exit` |
`stop` is the highest-quality signal for "the turn is over" and should be the documented default
for orchestration. `idle` is heuristic: output stabilization plus prompt detection, and it can
flap mid-turn when a spinner pauses. External CLI modes (`isExternalCliMode()`) have no stop
hook at all, so for opencode/codex/gemini/antigravity only `idle`, `working` and `exit` are
available. **The skill and the docs must say which signals exist per mode**, otherwise an agent
waits forever on `stop` in a codex session.
### 3.3 Endpoint specs
#### A. `GET /api/sessions/:id/wait`
| Param | Type | Default | Notes |
| --------- | ---------------------------------------------- | ---------------- | ------------------------------------------------------------ |
| `until` | comma list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on first match |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` (600000) |
| `fresh` | `0`/`1` | `0` | `1` requires a _transition_, ignoring the state at call time |
Response (always 200 unless the session is missing or a cap is hit):
```json
{
"success": true,
"data": {
"signal": "stop",
"timedOut": false,
"immediate": false,
"ended": false,
"waitedMs": 8421,
"status": "idle",
"sessionId": "...",
"until": ["stop", "idle", "exit"],
"limitPaused": false
}
}
```
`until` is echoed back because the server may narrow it: `stop`/`blocked` are dropped
from the DEFAULT set for external CLI modes (asking for them EXPLICITLY is a 400
instead, since omitting `until` must never 400). `limitPaused` tells a caller that a
timeout was expected rather than a stall worth retrying hard.
**A timeout is not an error.** `{"timedOut": true, "signal": null}` with HTTP 200, so a caller
can loop without treating every poll boundary as a failure. Errors are reserved for
`NOT_FOUND` (unknown or not-owned session) and `SESSION_BUSY` (waiter cap exceeded).
`immediate: true` means the session was already in the requested state and `fresh` was not set.
#### B. `GET /api/sessions/:id/wait-output`
| Param | Type | Default | Notes |
| --------- | ------------------------------ | -------- | --------------------------------------------------------- |
| `match` | literal string, 1 to 200 chars | required | substring match against ANSI-stripped output |
| `nocase` | `0`/`1` | `0` | case-insensitive compare |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the existing text buffer first, then waits |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` |
Response: `{ matched: true, timedOut: false, snippet: "...", waitedMs }`.
**No regex in v1, deliberately.** `search-service.ts` already avoids regex specifically so there
is no ReDoS surface, and this endpoint would be even more exposed since the pattern is attacker
supplied and the input is a live stream. herdr can offer `--regex` because Rust's regex crate is
linear-time with no backtracking; JS `RegExp` is not. If regex is wanted later, the honest
options are a length-capped subset compiled once with a match budget, or `re2`. Note it and move on.
Implementation detail that will bite if missed: a match can straddle two PTY chunks. Keep a
carry buffer of `match.length - 1` bytes from the previous chunk and test `carry + chunk`.
⚠️ **`from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, resize, or any TUI redraw, and a repaint arrives as ordinary `terminal`
data. Observed live: a marker echoed a minute earlier matched instantly on a fresh
`from=now` wait. This is inherent to a terminal multiplexer, not fixable in the registry,
so the contract is: **use a marker unique per call** (`echo DONE_$RANDOM`), never a
generic one like `BUILD OK`. The skill's recipes must show that.
The returned snippet is whitespace-collapsed (blank runs to a single newline) for
readability only; matching runs on the raw stripped text. Without it, a real pane's
`\r\n` padding between the prompt and the match fills the whole context window with
nothing, which was the first thing the live test showed.
#### C. `wait` on the existing input endpoint
`POST /api/sessions/:id/input` gains two optional fields:
```json
{ "input": "run the tests\r", "useMux": true, "clientId": "agent-1", "seq": 7, "wait": "stop", "waitTimeout": 600000 }
```
(The trailing `\r` is required on every input body: `sendInput` sends Enter only
when the input contains a carriage return.)
Response gains `"wait": { "signal": "stop", "timedOut": false, "waitedMs": 41230 }`.
This is the important one, because it closes a race the standalone `GET .../wait` cannot: between
"input delivered" and "session flips to working" there is a window where a naive
send-then-wait sees the _pre-existing_ idle state and returns instantly. The combined endpoint
**registers the waiter before writing**, so that window does not exist. This is exactly why herdr
ships `agent prompt --wait` as its own thing.
`wait` accepts `true` (the default signal set) or the same comma grammar as `until`.
Both new fields are `.nullish()`, not `.optional()`: a third-party caller building the
body with `JSON.stringify` keeps an explicit `null` on the wire, and `.optional()`
rejects that with `INVALID_INPUT`. That gotcha has shipped as a real bug twice.
Two behaviors to preserve carefully:
- **`useMux` is fire-and-forget today.** The handler responds without awaiting `writeViaMux`, on
purpose (a tmux child process must not block the HTTP response). With `wait` present the
handler already has to stay open, so it can await delivery, and a `writeViaMux` failure becomes
observable for the first time. The non-wait path must keep its current fire-and-forget shape
byte for byte.
- **Duplicate suppression.** A tagged redelivery (`clientId`+`seq` already applied) returns 200
without writing. With `wait` set it still waits, since the caller's intent is "tell me when
this settles". But it waits with `requireTransition: false`, unlike a fresh delivery: the
original turn may be long over, and requiring a new transition would block a redelivery until
timeout for no reason. Fresh delivery requires a transition, a duplicate answers from the
current state.
- **Capacity rollback.** `shouldApplyInput()` MUTATES (it records the seq), and it runs before
the waiter is registered. If registration then fails on a full pool, the handler must call
`forgetInputSeq` before returning `SESSION_BUSY`, or the caller's retry is rejected as a
duplicate and the input is lost by the very mechanism reliable delivery exists for.
### 3.4 Module design
New file `src/web/session-wait-registry.ts`, with the IO-free core unit-testable in isolation
(same split as `self-update.ts`):
```ts
type WaitSignal = 'idle' | 'working' | 'stop' | 'blocked' | 'exit';
waitForSignal(sessionId, { until: Set<WaitSignal>, timeoutMs, requireTransition }): Promise<WaitResult>
notifySignal(sessionId, signal: WaitSignal): void
waitForOutput(sessionId, { match, nocase, timeoutMs }): Promise<OutputWaitResult>
notifyOutput(sessionId, chunk: string): void
cancelAll(sessionId, reason): void
```
Wiring points, all existing:
- `src/web/session-listener-wiring.ts` around lines 190 and 200 already handles `working` and
`idle` and broadcasts them. Add a `notifySignal()` call next to each broadcast, plus `exit`.
- `src/web/routes/hook-event-routes.ts` already switches on `event` for the respawn controller.
Add `notifySignal(sessionId, 'stop' | 'blocked')` in the same switch.
- Output: `notifyOutput()` rides the ALREADY-attached `terminal` listener in
session-listener-wiring.ts. An earlier draft had the registry hand out attach/detach
callbacks so a listener could be added lazily; that was deleted once it was clear no
second listener is needed at all. The cost is one Map lookup per PTY chunk, which is why
the no-waiter check comes before the ANSI strip.
- Session deletion calls `notifySignal('exit')` then `cancelAll()`, so no promise is left
hanging. Both are required: `_doCleanupSession` detaches the session's listeners BEFORE
`session.stop()`, so on a delete the PTY exit event never reaches the registry, and an
`until=exit` caller would otherwise get a bare `ended` instead of its signal. Found by
live-testing the delete path, not by the unit tests.
Memory-leak discipline, per the 24-hour-session rules: every waiter owns a timer that is cleared
on resolve, the per-session waiter set is deleted when it empties, and the output listener is
removed with it. `test/memory-leak-prevention.test.ts` should grow a case for this.
Caps in a new `src/config/agent-wait.ts` (limits live in `src/config/`, env-overridable):
| Constant | Default | Why |
| ------------------------- | ------- | --------------------------------------- |
| `MAX_WAIT_MS` | 600000 | an unbounded long-poll is a socket leak |
| `DEFAULT_WAIT_MS` | 60000 | short enough to survive most proxies |
| `MAX_WAITERS_PER_SESSION` | 16 | |
| `MAX_WAITERS_TOTAL` | 128 | same reasoning as `MAX_SSE_CLIENTS` |
Exceeding a cap returns `SESSION_BUSY`, not a silent queue.
### 3.5 Transport concerns
Fastify is constructed with defaults in `server.ts:329-331`. `requestTimeout` defaults to 0
(disabled) and `keepAliveTimeout` (72s) applies between requests, not to an in-flight one, so a
10-minute in-process hold is fine. **Verify this on the real instance before relying on it.**
Intermediaries are the actual risk. Prod is reached through `tailscale serve`, and users also run
cloudflared tunnels; both can cut an idle connection. That is why `DEFAULT_WAIT_MS` is 60s and
why the documented pattern is a client-side loop over short waits rather than one 10-minute call.
The skill's recipes must show the loop.
### 3.6 Edge cases to get right
| Case | Behavior |
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Session already idle, `fresh=0` | return immediately, `immediate: true` |
| Session already idle, `fresh=1` | wait for the next transition into a requested state |
| Session dies mid-wait | resolve with `signal: "exit"` if `exit` was requested, otherwise resolve `timedOut:false, signal:null, ended:true`. Never hang |
| Session deleted mid-wait | same, resolve, do not throw. Verified live: `until=exit` gets `signal:"exit"`, a concurrent `until=blocked` gets `ended:true`, both in ~0ms |
| Shutdown with a wait pending | `cancelEverything()` in `stop()`. Verified live: SIGTERM with a 300s wait in flight exits in 1s |
| External CLI mode | `stop` and `blocked` never fire. Reject `until=stop` for those modes with a clear `INVALID_INPUT` rather than hanging until timeout |
| Multi-user | goes through `findSessionOrFail(ctx, id, req)`, which already enforces ownership |
| Remote / Docker cases | signals originate from the same `Session` object, so no special casing. Docker hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for `stop`/`blocked` to arrive at all; without it, only `idle` works. Document it |
| Respawn `/clear` mid-wait | a respawn cycle emits `idle`. Callers waiting on `stop` are unaffected; callers on `idle` may resolve early. Documented, not fixed |
| Limit pause | if the session is paused on a usage limit, nothing will fire until the reset. The wait times out honestly. Consider surfacing `limitPaused: true` in the response so the caller can back off |
### 3.7 Tests
- `test/session-wait-registry.test.ts` (pure): immediate resolve, transition-required, multi-signal
first-wins, timeout, cap exceeded, cancel on session end, no listener leak after resolve,
chunk-straddling output match, case-insensitive match.
- `test/routes/session-wait-routes.test.ts` (`app.inject()`, no port): all three endpoints against
a `MockSession`, including the 200-with-`timedOut` contract and the ownership 404.
- `test/routes/session-input-wait.test.ts`: the send-and-wait race, plus proof that the non-wait
path is unchanged (still returns before `writeViaMux` settles).
- Live verification on a throwaway session before COM, per the always-end-to-end-test rule.
### 3.8 Files touched
- `src/config/agent-wait.ts` (new)
- `src/web/session-wait-registry.ts` (new)
- `src/web/session-listener-wiring.ts` (notify on idle/working/exit)
- `src/web/routes/hook-event-routes.ts` (notify on stop/blocked)
- `src/web/routes/session-routes.ts` (two new routes, `wait` fields on input)
- `src/web/schemas.ts` (`SessionWaitQuerySchema`, `SessionWaitOutputQuerySchema`, extend
`SessionInputWithLimitSchema`. Note: `.optional()` rejects `null`, so the frontend and any
generated client must send `undefined`, never `null`)
- `docs/api-reference.md`, `docs/extending-codeman.md`, README API table
- `skills/codeman/SKILL.md` recipes (Part 1 depends on this)
---
## 4. Deferred: parts 3 to 5
Not in scope now, kept here so they are not lost.
### Part 3: promote `blocked` to a first-class state
`SessionStatus` is `'idle' | 'busy' | 'stopped' | 'error'`. "Needs you" exists three times over:
hook events, the `tab-alert-action` CSS class, and the phone overview NEEDS YOU section, each
re-deriving it. herdr makes `blocked` a real state that rolls up.
Add `blocked` (and possibly `done`) to `SessionStatus`, set it from the same hook events that
Part 2 uses as wait signals, and clear it on the next `working`/`stop`. Then the tab strip, the
mobile overview, the wait endpoints, and any external agent read one field.
Cost: `SessionStatus` is a widely-consumed union, so every exhaustive `switch` (the codebase has
`assertNever` and `noFallthroughCasesInSwitch`) will need a branch. That is a feature, it makes
the compiler find every site. This is a **minor** bump, not a patch: it widens a public type in
the HTTP contract.
### Part 4: `GET /api/schema`
herdr ships `herdr api schema`. Every Codeman route is already Zod-validated, so
`zod-to-json-schema` over `schemas.ts` gives a self-describing API almost free. Value: third-party
tools and the skill stop drifting from hand-written docs. Open question: whether to emit full
OpenAPI (`@fastify/swagger` would need per-route schema registration, which is a much larger
change) or just dump the Zod schemas keyed by name (cheap, 80% of the value).
### Part 5: detection manifests instead of hardcoded patterns
CLI-specific readiness, blocked and usage-limit patterns live in code across
`usage-limit-patterns.ts`, the respawn pattern helpers and `regex-patterns.ts`. Externalizing the
per-CLI ones into data files would make adding a sixth CLI a data change instead of a code change.
**Do not copy the remote-update part.** herdr auto-fetches manifest updates from herdr.dev.
Codeman auto-pulling behavioral rules from a vendor server contradicts its security posture.
Bundled manifests plus local override only, no network.
### Explicit non-goals
- **Plugin runtime and marketplace.** `docs/extending-codeman.md` already argues this: a plugin
runtime means third-party code inside a process that spawns agents with your credentials, on a
server people expose over a tunnel. The reasoning still holds. If the marketplace _pattern_ is
wanted, apply it to data (web tabs, case templates, cron recipes), never to executable code.
- **Live PTY handoff on restart.** herdr needs it because it owns the terminals. Codeman
delegates to tmux, so PTYs already survive a self-update restart.
- **Socket API.** HTTP plus SSE is the existing, documented, stable contract. A second transport
would double the surface for no capability gain.
---
## 5. Sequencing
| Step | Work | Gate |
| ---- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 ✅ | `src/config/agent-wait.ts` + `session-wait-registry.ts` + unit tests | 48 tests green |
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
| 5 ✅ | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
| 6 ✅ | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | 10 unit tests (`test/agent-skill.test.ts`) + real-server case-creation tests (`test/quick-start.test.ts`, incl. the settings PUT accepting the key) green; CLI verified live (install/uninstall, global + `--case`, foreign/symlink refusals) |
| 7 ✅ | Docs: api-reference, extending-codeman, README | plus `architecture-invariants.md` (§agent-wait-primitives), `CLAUDE.md` and the API reference's per-mode signal table |
| 8 ✅ | COM (minor bump: new endpoints, new setting, new optional fields) | released as 1.13.0 (wait primitives + skill); step 6 followed in 1.14.1 and was republished as 1.14.2 after live-testing the packaged skill |
Parts 1 and 2 are independent enough to land separately, but the skill is much less useful
without the wait endpoints, so the wait work goes first.
## 6. Open questions for the owner
1. ✅ `skills/` at the repo root: accepted (built that way; the install one-liner depends on it).
2. ✅ `agentSkillEnabled` default: **OFF** for the first release, per §2.2's rationale (skills
cost context on every turn; measure before defaulting on). Flip later if dogfooding earns it.
3. ✅ Both: global install via `npx skills add` / `codeman skill install`, AND per-case
auto-injection behind the (default-off) setting. Injection is add-only at session create and
marker-guarded, so a user-authored copy is never touched.
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
and not a security boundary? (Still open, not built with step 6.)
5. ✅ Regex support in `wait-output`: literal-only shipped, and a `regex` query param is
rejected with a 400 rather than ignored, so an agent that assumed otherwise cannot
silently wait on the wrong thing.
---
## 7. Build log: what actually happened
Written at the end of the build so the next person inherits the reasoning, not just the
diff. Process artifacts (per-agent briefs, findings, reports) live in the gitignored
`tmp/agent-wait-review/`; this section is the part worth keeping.
### What shipped
| Piece | Files |
| ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Bounds + clamping | `src/config/agent-wait.ts` (new) |
| Blocking-wait registry | `src/web/session-wait-registry.ts` (new, IO-free, unit-tested) |
| `GET .../wait`, `GET .../wait-output`, `wait`/`waitTimeout` on `POST .../input` | `src/web/routes/session-routes.ts` |
| Signal wiring | `session-listener-wiring.ts` (idle/working/exit + output), `hook-event-routes.ts` (stop/blocked), `server.ts` (teardown, shutdown) |
| Agent skill | `skills/codeman/SKILL.md` + `reference/`, `.claude/skills/codeman` symlink, `package.json` `files` |
| Docs | `api-reference.md`, `extending-codeman.md`, `architecture-invariants.md`, `README.md`, `CLAUDE.md` |
| Tests | `test/session-wait-registry.test.ts`, three `test/routes/session-*wait*.test.ts`, `http-contract.test.ts`, `mock-session.ts` |
### Bugs found in ADJACENT code, not in the new feature
These are the highest-value output of the exercise and none were on the plan:
1. **Every Codeman hook was dead on HTTPS installs.** `hooks-config.ts` built the hook
curl as `curl -s` with no `-k` while the statusline exporter 300 lines below used
`curl -sk` and documented why. Proven with the real hook command: `curl exit=60`
without the flag, success with it, and the failure swallowed by the hook's own
`2>/dev/null || true`. This silently killed `stop`, `permission_prompt`,
`elicitation_dialog`, `idle_prompt`, `teammate_idle` and `task_completed`, taking
respawn's definitive idle signals with them. Fixed, **plus** a staleness detector in
`refreshStaleCodemanHooks` that regenerates the on-disk config of already-created
cases (23 of 26 local cases carried the broken form; fixing the generator alone would
have left every one of them broken).
2. **`buildEnvExports()` exported a wrong-scheme `CODEMAN_API_URL`** (`http://` fallback
on an HTTPS install). Now omitted rather than guessed, so in-session guards fail closed.
3. **Programmatic input is only submitted when it contains `\r`.** `sendInput` sends Enter
only if the payload has a carriage return; without it the text sits in the composer
forever. Bit this build repeatedly before it was diagnosed, and had leaked into the
docs' own examples.
### Design decisions worth not re-litigating
- **A timeout is HTTP 200** with `wait.timedOut`, never a 4xx: callers loop over short
waits because tunnels cut idle connections, and every poll boundary would otherwise be
indistinguishable from failure.
- **Send-and-wait must be one endpoint.** A separate POST-then-wait races: between the
write and the flip to `working`, a wait sees the stale `idle` and reports the PREVIOUS
turn as this one. The waiter is registered before the write.
- **`stop`/`blocked` exist for `claude` mode only.** They come from Claude Code hooks;
`shell` installs none either, so keying off `isExternalCliMode()` was wrong.
- **Literal matching only, never regex.** JS `RegExp` backtracks; herdr can offer
`--regex` because Rust's regex crate is linear-time.
- **Client-hangup abort listens on `reply.raw` guarded by `writableFinished`.** On
`req.raw`, `close` fires when the request BODY ends, which on a POST killed every
send-and-wait instantly, and no `app.inject()` test can see it (inject never emits
`close`).
- **Liveness cannot come from `session.pid`.** For a tmux session that is the local
`tmux attach` client, not the worker: a worker exiting inside its pane leaves
`pane_dead=1` with the client alive, so `pid` never goes null. Liveness is probed at
the mux layer, cached (~750 ms) and only on blocking waits, never on the input hot path.
### Verification rounds
Six agents across three rounds, each verifying the previous round's work rather than its
own. Findings that mattered, in order of severity, were: the dead-pane liveness gap; the
`reply.raw` abort regression; abandoned long-polls leaking waiter slots; a crashed session
reporting `idle`; `shell` accepting `until=stop`; and a documented recipe that reported
success without running its task. Two traps recurred often enough to name:
- **Vacuous passes.** `app.inject()` never emits `close`; a latched `cancelEverything()`
in `afterEach` silently killed the registry for every later test in a file; three test
files sharing one session id against the process-wide registry let one file's leftover
waiter fail another's assertion. Any new wait test needs care on all three.
- **HTTP-only test instances.** Every isolated instance used during the build was plain
HTTP, which is exactly why the HTTPS hook bug survived so long. Test the transport the
user actually runs.
### Resolved at wrap-up (2026-08-08, conclusion pass)
- **R2-A**: the fire-and-forget-then-gather-sequentially pattern was **removed from
the skill** rather than patched. Signals are edge-triggered with no history, so a
`stop` that fires before its waiter registers is unobservable afterwards; a
`fresh=0` gather was rejected because the only `until` set that current state can
satisfy answers `idle` for a prompt that never submitted, resurrecting the exact
false-success failure R2-B had just closed. Flow 3b's pattern B now gathers on
latched `wait-output` markers (`from=buffer`), the same mechanism that makes the
shell flows reliable; the limitation is recorded in
`architecture-invariants#agent-wait-primitives` and `endpoints.md`. The durable
fix, a latched last-signal-per-turn on the server, stays with deferred Part 3.
- Docs F7/F8, F4 and the false-`idle` attribution: `api-reference.md`,
`extending-codeman.md` and `architecture-invariants.md` rewritten to the post-fix
matcher (one normalized stream, chunk-straddling found, snippet as a rendering of
the matched window), the real no-PTY answer (`ended:true`, `aborted:false`,
`delivered:false`), and the startup-idle mechanism (a session parked on the trust
dialog emits no further `idle`; the false success is the startup transition).
- Orchestrate #12, #5/R2-B, #6, and R2-C..R2-E: fire-and-forget's empty `data`
documented; every send-and-wait retry loop now treats `duplicate:true` +
`immediate:true` as "no new turn ran" and reads the terminal before believing it;
claude fan-out is pattern A (backgrounded send-and-waits) or the marker gather;
readiness budgets rebalanced (5 s stage 1, 45 s stage 3) with the virgin-case
floor named; the auth fallback now also reads the supervisor definition
(`codeman-web.service` / launchd plist) and accepts `export`-prefixed `.env`
lines; `pid != null` is documented as startup-only, never liveness.
- Both public readiness recipes (extending-codeman.md, README) are bypass-first with
the trust probe as the bounded fallback; the worked recipe carries `-k` and fails
loudly on an empty SID; the hook `-k`/self-heal fix appears in every
"hooks go missing" list; the multi-word-TUI claim is "unreliable", not "never".
### Still open
Both release-checklist items that used to sit here are done: `skills/` is tracked and
ships through `package.json` `files` (published with 1.13.0, republished with 1.14.2),
and the changeset was consumed, committed and deployed. What is left:
- Deferred with Part 3: the latched last-signal-per-turn. Nice-to-haves from the
reviews: N2 (create the death-watcher inside its `try`, still built one line above
it in `GET .../wait`) and converting timeout-shaped test detections into fast
assertions.
- §2.4's `X-Codeman-Caller-Session` footgun guard: still not built (open question 4).
### Step 6 (2026-08-09): install command, per-case injection, the setting
Built to the §2.6 file list, mirroring the statusLine mechanism throughout:
| Piece | Where |
| ----- | ----- |
| `applyAgentSkill(casePath, enabled)` + `installAgentSkillInto` / `removeAgentSkillFrom` | `src/hooks-config.ts` |
| `codeman skill install` / `skill uninstall` (`--global` default, `--case <name>`) | `src/cli.ts` |
| `agentSkillEnabled` (SYNCED, default OFF) | `schemas.ts` (`SettingsUpdateSchema`), `getAgentSkillEnabled()` on `ConfigPort`/`server.ts`, checkbox in `index.html` + `settings-ui.js` |
| Injection call sites (Claude mode only) | `POST /api/sessions` next to `refreshStaleCodemanHooks`; `POST /api/quick-start` after the case-create/self-heal blocks (local + docker cases; remote skipped, its path lives on another host) |
| Tests | `test/agent-skill.test.ts` (10 unit), `test/quick-start.test.ts` (real server: default-off, PUT accepts key, injection on create, shell-mode skipped) |
Decisions worth keeping:
- **Ownership marker, prefix-matched.** The injected SKILL.md ends with
`<!-- codeman-managed-agent-skill: … -->`; install/refresh/remove all refuse a copy
without the marker (a user's own skill) and match on the PREFIX so a wording change
cannot disown older injected copies (the `BACKGROUND_WAKE_MARKER_PREFIX` pattern).
- **Symlink refusal.** This repo's own dogfooding layout
(`.claude/skills/codeman -> ../../skills/codeman`) means the injector must `lstat`
the skill dir AND its `skills/` parent and bail on a symlink, or enabling the
setting in the Codeman repo itself would overwrite the skill source through the link.
- **ADD-ONLY at session create**, same shared-`.claude` rationale as the statusLine:
a create while the setting is off must not yank the skill out from under other live
sessions in the repo. The remove path exists (CLI `skill uninstall`, tests); no
automatic sweep removes on toggle-off.
- **Removal is manifest-based, never `rm -rf`**: only files the packaged source would
have written are deleted, directories are pruned bottom-up only if they emptied, so
a user's extra notes in `reference/` survive an uninstall.
- **Source resolution**: `join(moduleDir, '..', 'skills', 'codeman')` works from
`src/` (tsx), `dist/` (tsc build), and the npm tarball alike, because all three sit
one level below the package root and `files` ships `skills/`.
- **Nothing acts on the setting at PUT time**: injection reads the merged persisted
settings at session create (`readSettings`, ~2s cache), so the partial-PUT invariant
(`toggleService` reading `merged`) is untouched by construction.
### 2026-08-09 addendum: cross-session messaging folded into the skill
Claude Code 2.1.224+ ships cross-session messaging: `ListAgents`/`SendMessage`
tools, a per-session Unix inbox socket, and a registry in
`~/.claude/sessions/<pid>.json`. Codeman's claude workers are ordinary local Claude
Code sessions, so the skill now routes task delivery and result collection over it
when available, while the HTTP primitives keep spawn, readiness, synchronization,
liveness and delete. New `skills/codeman/reference/messaging.md` (ships with zero
installer changes: `readAgentSkillSource()` enumerates `reference/*.md` from disk),
Flow 5 in recipes.md, and §4 in SKILL.md.
Verified live (claude-cli 2.1.226, Linux):
- A message to an idle worker starts a turn and that turn fires the normal `stop`
hook (8.3 s send-to-stop measured), so the HTTP wait primitives compose with
messaging unchanged; delivery to a busy session lands between tool calls.
- First contact needs the `name [ref]` form; the bare name errors with the exact
string to resend. The `uds:` reply address of an inbound message works as a `to`.
- The `tmux codeman-<id8>` column in `ListAgents` (and the registry's `tmux` field)
is the join key to Codeman session ids. The registry's `sessionId` field starts as
the Codeman id (we spawn `claude --session-id <id>`) but drifts after `/clear` or
resume, so it must never be the join key.
- The feature is flag-gated beyond the version: two 2.1.226 sessions on one machine,
one with an inbox socket and one without. Absence is a fallback case, not an error.
- Codeman's default `--dangerously-skip-permissions` spawn puts both ends in the
bypassing class, which delivers; mixed classes hold behind an approval dialog that
expires unattended (upstream default 5 min), which on a headless worker means the
message silently dies. The skill's backstop covers it.
Follow-up, landed in the same PR: local claude spawns now pass
`--name <session name>` so peers carry Codeman session names. The gate is
`buildNameCliArgs()` (session-cli-builder.ts), fail-closed at
`CLAUDE_NAME_FLAG_MIN_VERSION = 2.1.224`: that is the messaging release, the flag's
presence there was verified against the installed 2.1.224 binary, and the version
comes from `getClaudeCliVersion()` (null on probe failure and under vitest), so an
older or unknown CLI gets a command byte-identical to before. That matters because
claude aborts startup on an unknown option, which would kill every session spawn.
The value is allowlist-sanitized (Unicode letters/digits plus ` ._:-`, leading
dashes stripped so it cannot parse as another option, 64-char cap, empty result =
flag omitted) before the double-quoted interpolation in `buildSpawnCommand`, and
only the LOCAL command carries it: the docker/remote builders never see it, since
their CLI is not the binary the probe measured. E2E on an isolated instance
(`CODEMAN_INSTANCE`): process cmdline `claude ... --name w9-msgtest`, registry
`name: "w9-msgtest"`, `ListAgents` lists it under that name, a message round-trip
works, and its replies arrive tagged `from-name="w9-msgtest"` (a derived-name
worker's replies carry no `from-name`). A quick-start without `sessionName` has an
empty Codeman name, so the peer name stays derived: agents should name their
workers. Tests: `test/name-flag-injection.test.ts`.
-244
View File
@@ -1,244 +0,0 @@
# Claude Code Agent Teams — Reference
> Experimental feature (Feb 2026). Enable per-session via env var.
> Updated with experiment findings from 2026-02-12.
## Enabling
```bash
# Environment variable (set before starting Claude Code)
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
# In .claude/settings.local.json (case-scoped)
{
"env": {
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
}
}
# Note: "teammateMode" is NOT a valid settings key (validation rejects it).
# Display mode defaults to "in-process". For tmux, pass --teammate-mode flag via CLI.
```
## Filesystem Paths (Verified)
| Resource | Path |
|----------|------|
| Team config | `~/.claude/teams/{team-name}/config.json` |
| Teammate inboxes | `~/.claude/teams/{team-name}/inboxes/{name}.json` |
| Shared tasks | `~/.claude/tasks/{team-name}/` |
| Teammate transcripts | `~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{id}.jsonl` |
Note: Teammate transcripts appear in the **standard subagent directory** under the lead's session, NOT as separate top-level sessions.
### config.json format (verified)
```json
{
"name": "research-watchers",
"description": "Team description...",
"createdAt": 1770875105373,
"leadAgentId": "team-lead@research-watchers",
"leadSessionId": "461daa80-94ec-4e5e-a1bb-0518f78311bc",
"members": [
{
"agentId": "team-lead@research-watchers",
"name": "team-lead",
"agentType": "team-lead",
"model": "claude-opus-4-6",
"joinedAt": 1770875105373,
"tmuxPaneId": "",
"cwd": "/path/to/project",
"subscriptions": []
},
{
"agentId": "fs-researcher@research-watchers",
"name": "fs-researcher",
"agentType": "general-purpose",
"model": "claude-opus-4-6",
"prompt": "Full spawn prompt...",
"color": "blue",
"planModeRequired": false,
"joinedAt": 1770875126680,
"tmuxPaneId": "in-process",
"cwd": "/path/to/project",
"subscriptions": [],
"backendType": "in-process"
}
]
}
```
Key fields: `agentId` format is `{name}@{teamName}`, `leadSessionId` links to Codeman session, `backendType` indicates display mode, `color` for UI theming.
### Task file format (verified)
```json
{
"id": "1",
"subject": "Research Node.js fs.watch on Linux vs macOS",
"description": "Full description...",
"activeForm": "Researching Node.js fs.watch Linux vs macOS",
"status": "in_progress",
"blocks": [],
"blockedBy": [],
"owner": "fs-researcher"
}
```
Internal teammate tracking tasks have `"metadata": { "_internal": true }`.
Task states: `pending` → `in_progress` → `completed`. File locking via `.lock.lock` directory (mkdir-based atomic lock).
### Inbox message format (verified)
```json
[
{
"from": "team-lead",
"text": "{\"type\":\"task_assignment\",\"taskId\":\"1\",\"subject\":\"...\",\"assignedBy\":\"team-lead\",\"timestamp\":\"...\"}",
"timestamp": "2026-02-12T05:45:18.176Z",
"read": false
}
]
```
`text` is double-encoded JSON. Message types: `task_assignment`, `shutdown_request`, `shutdown_response`. File locking via `.json.lock` directory.
## Communication Model (CORRECTED)
**Hybrid: tool + filesystem.** The `SendMessage` tool writes to filesystem inbox files at `~/.claude/teams/{name}/inboxes/{teammate}.json`.
Each teammate AND the lead has an inbox JSON file. Messages are JSON arrays with `from`, `text` (double-encoded JSON), `timestamp`, `read` fields.
Message types observed:
- **task_assignment**: Lead assigns task to teammate
- **shutdown_request**: Lead asks teammate to shut down
- **shutdown_response**: Teammate confirms shutdown
- (Also: `message`, `broadcast`, `plan_approval_response` per docs)
**Implication:** We can intercept messages by watching inbox files AND potentially inject messages by writing to them (respecting `.json.lock` directory locking).
## Process Model (CORRECTED)
**Teammates are IN-PROCESS THREADS, not separate OS processes.**
In `in-process` mode (the default), all teammates run as threads within the single `claude` process. Only 1 claude process exists per Codeman session, regardless of team size.
This means:
- No separate PIDs to track per teammate
- All teammates share the lead's environment variables
- Lower resource overhead than separate processes
- Subagent transcript files still created (for progress tracking)
## Display Modes
| Mode | Trigger | UI | Requirement |
|------|---------|-----|------------|
| **in-process** (default) | Default | Shift+Up/Down to switch, Ctrl+T for tasks | Any terminal |
| **tmux** | `--teammate-mode tmux` | Split panes | tmux installed |
| **iTerm2** | Auto-detected | Native split panes | iTerm2 + `it2` CLI |
**For Codeman: use `in-process` only.** Codeman manages its own tmux sessions externally.
**In-process UI elements:**
- Status bar: `@main @teammate1 @teammate2 ...` with `shift+↑ to expand`
- Task list: Checkboxes with assignments `(@teammate-name)`
- Hint: `ctrl+t to show teammates`
## Hooks
Two new hook types for quality gates (verified in settings schema):
### TeammateIdle
Fires when a teammate is about to go idle.
- Exit code 0: Allow idle (normal)
- Exit code 2: Send feedback back, keep teammate working
### TaskCompleted
Fires when a task is being marked complete.
- Exit code 0: Allow completion
- Exit code 2: Prevent completion, send feedback
These are configured in `.claude/settings.local.json` alongside existing Codeman hooks.
## Subagent-Watcher Compatibility (Verified)
**Teammates appear as standard subagents.** They create transcript files at:
```
~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{id}.jsonl
```
Codeman's existing `subagent-watcher.ts` discovers them automatically. They appear in `/api/subagents` with status "active".
**Distinguishing teammates from regular subagents:**
- Description field starts with `<teammate-message teammate_id= team`
- Cross-reference with `~/.claude/teams/{name}/config.json` members
**Sub-subagents:** Teammates can spawn their own Task tool subagents, creating a 3-level hierarchy.
## Cleanup Behavior (Verified)
When the lead runs cleanup:
1. Shutdown requests sent to all teammate inboxes
2. Teammates shut down gracefully
3. ALL filesystem artifacts deleted:
- Inbox files and directory
- Config.json
- Team directory
- All task files
- Task directory
4. Cleanup is atomic — all files removed in the same second
## Comparison with Subagents (Task tool)
| Aspect | Subagents (Task tool) | Agent Teams |
|--------|----------------------|-------------|
| Spawn method | Claude's built-in Task tool | Explicit team creation |
| Process model | In-process threads | In-process threads (same!) |
| Discovery | `subagents/agent-{id}.jsonl` only | BOTH subagent dir + `~/.claude/teams/` |
| Communication | None (fire-and-forget) | Filesystem inboxes + SendMessage tool |
| Shared state | None | Shared task list + inboxes |
| Task tracking | Per-agent, no coordination | Shared with dependencies & ownership |
| Lifecycle | Auto-cleanup on completion | Lead cleanup (deletes all artifacts) |
| Sub-nesting | Can spawn sub-subagents | Teammates can spawn subagents too |
| Cost | Lower (single context) | Higher (N context windows) |
| Duration | Short-lived (seconds-minutes) | Longer-lived (minutes-hours) |
## Limitations
- No session resumption with in-process teammates (`/resume` doesn't restore them)
- One team per session, no nested teams
- Lead is fixed (cannot promote teammate)
- Permissions set at spawn (change individually after)
- Split panes require tmux or iTerm2 (not Screen)
- Task status can lag (teammates may fail to mark complete)
- Shutdown can be slow (waits for current tool call)
## Useful Commands
```bash
# Check if teams exist
ls ~/.claude/teams/
# Check team config
cat ~/.claude/teams/{name}/config.json | jq .
# Check teammate inboxes
cat ~/.claude/teams/{name}/inboxes/{teammate}.json | jq .
# Check team tasks
ls ~/.claude/tasks/{name}/
for f in ~/.claude/tasks/{name}/*.json; do cat "$f" | jq .; done
# Count Claude processes (teammates are threads, not processes)
ps aux | grep '[c]laude' | grep -v grep
# Check subagent detection of teammates
curl -s http://localhost:3000/api/subagents | jq '.data[] | select(.description | startswith("<teammate"))'
# Team interaction (in-process mode)
# Shift+Up/Down: Switch between teammates
# Enter: View teammate session
# Escape: Interrupt teammate's turn
# Ctrl+T: Toggle task list
```
-171
View File
@@ -1,171 +0,0 @@
# Codeman Agent Teams Integration — Design (Approach C: Hybrid)
> Updated 2026-02-12 with experiment findings. See `experiment-log.md` for raw data.
## Overview
Approach C combines filesystem monitoring (for team/task discovery and inbox watching) with the existing subagent-watcher (for live transcript tailing) and adjusted idle detection (to account for active teammates). The key finding from our experiment is that **teammates already appear as standard subagents**, so most infrastructure exists — we mainly need team awareness and idle detection fixes.
## Components
### 1. TeamWatcher (`src/team-watcher.ts`)
Monitors `~/.claude/teams/` for team creation/removal and tracks active teams.
**Discovery mechanism:**
- Poll `~/.claude/teams/` for directories (team names) every 3-5 seconds
- When found: parse `config.json` to get:
- `leadSessionId` → map to Codeman session
- `members` array → teammate names, agentIds, colors, models
- Watch for directory deletion (cleanup signal)
**CORRECTED from pre-experiment design:**
- ~~Each teammate has a separate Claude Code process~~ → Teammates are **in-process threads**, not separate processes
- ~~Find via `ps aux` + `/proc` PID matching~~ → Not needed, no separate PIDs
- Teammate transcripts are at `subagents/agent-{id}.jsonl` (standard subagent path), NOT separate session transcripts
**Association:**
- `config.json.leadSessionId` → Codeman session ID (direct match!)
- Each member's `agentId` (e.g., `fs-researcher@research-watchers`) → links to subagent files
- `agentType: "team-lead"` vs `"general-purpose"` distinguishes lead from teammates
**Inbox monitoring:**
- Watch `~/.claude/teams/{name}/inboxes/` for new messages
- Each teammate has a JSON file with message array
- Messages are double-encoded JSON with `from`, `text`, `timestamp`, `read` fields
- Message types: `task_assignment`, `shutdown_request`, `shutdown_response`
### 2. Team-Aware Idle Detection (HIGHEST PRIORITY)
**Problem (confirmed by experiment):** Lead session shows status "idle" in Codeman while teammates are actively working. Token count continues climbing but Codeman thinks the session is inactive.
**Solution:**
- Before declaring a session idle, check if it's a team lead
- If team lead: check `~/.claude/teams/*/config.json` for this session's `leadSessionId`
- If active team exists: check task files in `~/.claude/tasks/{team-name}/`
- Any task with `status: "in_progress"` → suppress idle detection
- All tasks `completed` AND no non-`_internal` tasks pending → allow idle
- Fallback: check subagent-watcher for active subagents on this session
**Integration points:**
- `src/ai-idle-checker.ts` — add team-awareness check before AI idle analysis
- `src/respawn-controller.ts` — consult TeamWatcher before transitioning to idle states
- `src/session.ts` — expose `hasActiveTeam()` method
**Liveness check (simplified from pre-experiment):**
- ~~Check `/proc/{pid}` existence~~ → Not needed (no separate processes)
- Check task file status instead (filesystem-based)
- Check subagent-watcher for active subagents under this session
### 3. Shared Task List UI
**Display:** New panel in web UI showing the team's shared task list.
**Data source:** Poll `~/.claude/tasks/{team-name}/` for task JSON files.
**Task file structure (verified):**
```json
{
"id": "1",
"subject": "Research Node.js fs.watch",
"description": "Full description...",
"activeForm": "Researching Node.js fs.watch",
"status": "in_progress", // pending | in_progress | completed
"blocks": [],
"blockedBy": [],
"owner": "fs-researcher" // Empty string = unassigned
}
```
Internal tracking tasks: `{ "metadata": { "_internal": true } }` — filter these from display.
**UI elements:**
- Task subject, status badge (color-coded), owner (teammate name with color)
- Dependency visualization (blockedBy indicators)
- Progress bar (completed / total non-internal tasks)
- Real-time updates via SSE
**API endpoint:** `GET /api/sessions/:id/team-tasks` → returns parsed task files
**Locking:** Respect `.lock.lock` directory lock when reading (skip if locked, retry next poll).
### 4. Teammate Display
**Decision: Option A — Enhanced subagent floating windows.**
Since teammates already appear as subagents in the existing infrastructure, we enhance rather than replace:
- **Badge:** Add "Teammate" badge to subagent windows for agents matching team config
- **Color:** Use teammate's `color` field from config.json (blue, green, yellow)
- **Name:** Show teammate name instead of agent ID
- **Persistence:** Teammate windows should stay open longer (they're longer-lived than regular subagents)
- **Status:** Show task assignment and progress from task files
**Detection logic:**
```
For each subagent detected by subagent-watcher:
1. Check if description starts with "<teammate-message"
2. OR cross-reference agentId with active team config members
3. If match → apply teammate badge, color, name
```
### 5. Inbox/Message Display
**CORRECTED: Inboxes ARE filesystem-based.**
Communication uses filesystem inbox files at `~/.claude/teams/{name}/inboxes/{teammate}.json`. We can:
1. **Watch inbox files** for real-time message monitoring
2. **Parse message types** for display:
- `task_assignment` → "Lead assigned Task #1 to fs-researcher"
- `shutdown_request` → "Lead requested shutdown"
- `shutdown_response` → "Teammate confirmed shutdown"
3. **Display as timeline** in team panel
**Potential for interaction (not tested, future work):**
- Write to teammate inbox files to inject messages
- Must respect `.json.lock` directory locking protocol
- Could enable "nudge" or "redirect" functionality from Codeman UI
## Answered Questions (from experiment)
| # | Question | Answer |
|---|----------|--------|
| 1 | Teammates in subagents dir? | **YES** — standard `subagents/agent-{id}.jsonl` path |
| 2 | subagent-watcher detects them? | **YES** — automatically, no changes needed |
| 3 | Task file structure? | Numbered JSON files with subject, status, owner, dependencies |
| 4 | Env var inheritance? | **YES** — in-process threads share parent's env |
| 5 | Processes per teammate? | **ZERO** — threads, not processes |
| 6 | config.json format? | Rich: name, agentId, agentType, model, prompt, color, backendType |
| 7 | Interact via stdin? | N/A (threads) — can interact via inbox files instead |
| 8 | In-process under Screen? | Works fine — single claude process, threads handle teammates |
| 9 | Hook events from teammates? | TeammateIdle + TaskCompleted hooks available in settings schema |
| 10 | Process tree? | Single process with threads — no child processes |
## Existing Infrastructure to Leverage
| Component | Reuse for | Status |
|-----------|-----------|--------|
| `subagent-watcher.ts` | Teammate transcript tailing | **Already works** |
| Subagent floating windows (`app.js`) | Teammate activity display | **Already works** (needs badges) |
| `task-tracker.ts` | Background task tracking patterns | Reuse patterns |
| LRUMap, StaleExpirationMap | Bounded caches for team state | Available |
| SSE broadcast | Real-time UI updates | Available |
| ~~`/proc` PID checking~~ | ~~Teammate liveness~~ | **Not needed** (threads) |
| `file-stream-manager.ts` | Watch inbox/task files | Available |
## Implementation Order (Revised)
1. **Team-aware idle detection** — prevent premature respawn/auto-compact (CRITICAL)
2. **TeamWatcher** — poll `~/.claude/teams/`, parse config.json, track active teams
3. **Teammate badge in subagent windows** — mark teammate subagents with name/color
4. **Team tasks API + UI** — `GET /api/sessions/:id/team-tasks` + task list panel
5. **Inbox monitoring** — watch inbox files, display message timeline
6. **TeammateIdle/TaskCompleted hooks** — add to Codeman's hooks config generator
## What We DON'T Need to Build
- ~~Process discovery for teammates~~ (they're threads)
- ~~Custom transcript tailing~~ (subagent-watcher handles it)
- ~~Separate teammate window infrastructure~~ (subagent windows work)
- ~~Message interception via transcript parsing~~ (inbox files are simpler)
-468
View File
@@ -1,468 +0,0 @@
# Agent Teams Experiment Log
> Experiment date: 2026-02-12
> Test case: `~/codeman-cases/agent-teams-test/`
> Team name: `research-watchers`
> Teammates: 3 (fs-researcher, perf-researcher, api-researcher)
> Lead session: `461daa80-94ec-4e5e-a1bb-0518f78311bc`
> Duration: ~3 minutes (06:45:01 → 06:48:07)
## Pre-Experiment State
```
~/.claude/teams/ — did NOT exist
~/.claude/tasks/ — 75 UUID-named directories (from regular Task tool subagents)
Claude processes — 7 (including watchers)
settings.local.json — edited to add CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
## Experiment Prompt
```
Create an agent team with 3 teammates to research the following topics in parallel:
Teammate 1 fs-researcher researches how Node.js fs.watch works on Linux vs macOS.
Teammate 2 perf-researcher researches inotify performance limits and alternatives.
Teammate 3 api-researcher researches the inotifywait command-line API.
Have each teammate write a brief summary of their findings in a separate file.
Name the team research-watchers.
```
---
## Question 1: What exact filesystem artifacts do agent teams create?
**Expected:** `~/.claude/teams/research-watchers/config.json` and `~/.claude/tasks/research-watchers/`
**Actual: CONFIRMED + SURPRISE inboxes/ directory**
```
~/.claude/teams/research-watchers/
├── config.json # Team config (members, lead, metadata)
└── inboxes/ # Filesystem-based messaging!
├── api-researcher.json # Per-teammate inbox
├── fs-researcher.json
├── perf-researcher.json
└── team-lead.json # Lead also has an inbox
~/.claude/tasks/research-watchers/
├── .lock # Empty file (presence = lock indicator?)
├── 1.json # Task: Research Node.js fs.watch
├── 2.json # Task: Research inotify performance
├── 3.json # Task: Research inotifywait CLI
├── 4.json # Internal: fs-researcher spawn tracking
├── 5.json # Internal: perf-researcher spawn tracking
└── 6.json # Internal: api-researcher spawn tracking
```
Subagent transcripts also appear in the standard subagent directory:
```
~/.claude/projects/-home-arkon-codeman-cases-agent-teams-test/
└── 461daa80.../
├── 461daa80...jsonl # Lead session transcript
└── subagents/
├── agent-ae50544.jsonl # Teammate: fs-researcher
├── agent-aa20c65.jsonl # Teammate: perf-researcher
├── agent-a29de32.jsonl # Teammate: api-researcher
├── agent-a04968e.jsonl # Sub-subagent (teammate's Task tool)
├── agent-a0d372e.jsonl # Sub-subagent
├── agent-a2ff939.jsonl # Sub-subagent
├── agent-a89ad82.jsonl # Sub-subagent
├── agent-aa1efc7.jsonl # Sub-subagent
└── agent-ab0ef07.jsonl # Sub-subagent
```
**Cleanup:** At 06:48:02, the lead deleted ALL artifacts — inboxes, config, tasks, the team directory itself. Clean removal.
---
## Question 2: Is the mailbox/communication filesystem-based or tool-based?
**Expected:** Tool-based (SendMessage tool), NOT filesystem
**Actual: BOTH! Hybrid — tool triggers filesystem writes.**
Communication uses the `SendMessage` tool internally, but the actual message delivery is via **filesystem inbox files**. Each teammate has `~/.claude/teams/{name}/inboxes/{teammate}.json` containing a JSON array of messages.
**Inbox message format:**
```json
[
{
"from": "team-lead",
"text": "{\"type\":\"task_assignment\",\"taskId\":\"1\",\"subject\":\"Research Node.js fs.watch...\",\"assignedBy\":\"team-lead\",\"timestamp\":\"...\"}",
"timestamp": "2026-02-12T05:45:18.176Z",
"read": false
}
]
```
Key observations:
- `text` field is a **JSON string** (double-encoded) containing a typed message object
- Message types observed: `task_assignment`, `shutdown_request`, `shutdown_response`
- `read` field tracks whether teammate has processed the message (false → true)
- **File locking** via `.json.lock` directories (mkdir-based atomic lock, created then deleted)
- Lead also has an inbox (`team-lead.json`) for receiving messages FROM teammates
**Implication for Codeman:** We CAN intercept messages by watching inbox JSON files! We can also potentially inject messages by writing to inbox files.
---
## Question 3: Do teammates appear in the subagents directory?
**Expected:** Unclear
**Actual: YES! Teammates appear as standard subagents.**
Teammates create transcript files at:
```
~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{agentId}.jsonl
```
This is the **exact same path pattern** that regular Task tool subagents use. The existing `subagent-watcher.ts` successfully discovers them.
Codeman's `/api/subagents` endpoint returned them with status "active":
```
Agent: ae50544 Status: active Tools: 8 Model: claude-opus-4-6
Desc: <teammate-message teammate_id= team
Agent: aa20c65 Status: active Tools: 9 Model: claude-opus-4-6
Desc: <teammate-message teammate_id= team
Agent: a29de32 Status: active Tools: 7 Model: claude-opus-4-6
Desc: <teammate-message teammate_id= team
```
**Distinguishing teammates from regular subagents:**
- Description starts with `<teammate-message teammate_id= team` (a unique marker)
- We can also cross-reference with `~/.claude/teams/{name}/config.json` members list
**Sub-subagents:** Teammates can spawn their own Task tool subagents. 3 teammates spawned 6 additional subagent files (9 total in the subagents directory).
---
## Question 4: What does config.json actually look like?
**Actual config.json (with all 3 teammates):**
```json
{
"name": "research-watchers",
"description": "Research team investigating file watching mechanisms...",
"createdAt": 1770875105373,
"leadAgentId": "team-lead@research-watchers",
"leadSessionId": "461daa80-94ec-4e5e-a1bb-0518f78311bc",
"members": [
{
"agentId": "team-lead@research-watchers",
"name": "team-lead",
"agentType": "team-lead",
"model": "claude-opus-4-6",
"joinedAt": 1770875105373,
"tmuxPaneId": "",
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
"subscriptions": []
},
{
"agentId": "fs-researcher@research-watchers",
"name": "fs-researcher",
"agentType": "general-purpose",
"model": "claude-opus-4-6",
"prompt": "You are \"fs-researcher\" on the \"research-watchers\" team...",
"color": "blue",
"planModeRequired": false,
"joinedAt": 1770875126680,
"tmuxPaneId": "in-process",
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
"subscriptions": [],
"backendType": "in-process"
},
{
"agentId": "perf-researcher@research-watchers",
"name": "perf-researcher",
"agentType": "general-purpose",
"model": "claude-opus-4-6",
"prompt": "...",
"color": "green",
"planModeRequired": false,
"joinedAt": 1770875130344,
"tmuxPaneId": "in-process",
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
"subscriptions": [],
"backendType": "in-process"
},
{
"agentId": "api-researcher@research-watchers",
"name": "api-researcher",
"agentType": "general-purpose",
"model": "claude-opus-4-6",
"prompt": "...",
"color": "yellow",
"planModeRequired": false,
"joinedAt": 1770875134997,
"tmuxPaneId": "in-process",
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
"subscriptions": [],
"backendType": "in-process"
}
]
}
```
**Key fields per member:**
- `agentId`: `{name}@{teamName}` format
- `agentType`: `"team-lead"` for lead, `"general-purpose"` for teammates
- `model`: Model used (inherits from lead)
- `prompt`: Full spawn prompt (only for teammates)
- `color`: UI color assignment (blue, green, yellow)
- `backendType`: `"in-process"` for in-process mode
- `tmuxPaneId`: `"in-process"` or actual pane ID for tmux mode
- `subscriptions`: Empty array (possibly for message routing)
**Config grows incrementally** — starts with just lead member (620 bytes), grows as teammates are added (→ 1886 → 3188 → 4551 bytes).
---
## Question 5: How do shared tasks differ from regular tasks?
**Expected:** Team name directory vs UUID, richer task format
**Actual: CONFIRMED**
**Team tasks (`~/.claude/tasks/research-watchers/`):**
```json
{
"id": "1",
"subject": "Research Node.js fs.watch on Linux vs macOS",
"description": "Research how Node.js fs.watch works differently...",
"activeForm": "Researching Node.js fs.watch Linux vs macOS",
"status": "in_progress",
"blocks": [],
"blockedBy": [],
"owner": "fs-researcher"
}
```
**Internal teammate tracking tasks (4.json, 5.json, 6.json):**
```json
{
"id": "4",
"subject": "fs-researcher",
"description": "You are \"fs-researcher\" on the \"research-watchers\" team...",
"status": "in_progress",
"blocks": [],
"blockedBy": [],
"metadata": { "_internal": true }
}
```
**Key differences from regular subagent tasks (`~/.claude/tasks/{UUID}/`):**
| Feature | Regular tasks | Team tasks |
|---------|--------------|------------|
| Directory name | UUID | Human-readable team name |
| File names | `.lock`, `.highwatermark` only | Numbered JSON files (1.json, 2.json...) |
| Content | Lock files only (no task JSON) | Full task JSON with metadata |
| Owner field | N/A | Teammate name |
| Locking | `.lock` file | `.lock.lock` directory (mkdir atomic) |
| Internal tasks | None | `_internal: true` for teammate spawn tracking |
---
## Question 6: Can we write to task/mailbox files to interact with teammates?
**Expected:** Possibly for tasks, no for messages
**Actual: LIKELY YES for both**
Evidence supporting external writes:
1. **Inbox files** are plain JSON arrays — we could append messages
2. **Task files** are plain JSON — we could modify status, add new tasks
3. **File locking** uses `.json.lock` directories — we'd need to respect the locking protocol
4. **Lock protocol**: Create directory `{file}.lock` → write → delete directory. Simple mkdir-based atomic lock.
**Not tested in this experiment** — would need a follow-up test to verify teammates actually pick up externally-added messages/tasks. But the format is clear and the locking is simple.
---
## Question 7: What happens to Codeman's idle detection with active teammates?
**Expected:** Lead may appear idle while teammates work
**Actual: Lead stays "idle" in Codeman's view, but terminal shows active status**
Observations:
- Codeman session status showed `"idle"` throughout the experiment
- The terminal output continued updating (task list checkboxes, teammate progress messages)
- Lead displayed "Befuddling..." spinner while waiting for teammates
- Token count climbed from 27k → 33k during the experiment
- The `stop` hook DID fire at the end when the team was cleaned up
**Implication:** Current idle detection may trigger prematurely if:
- It only checks Codeman's session status (which stays "idle")
- It doesn't account for active teammates
**What we need:** Check `~/.claude/teams/*/config.json` for active members before declaring idle.
---
## Question 8: How many Claude processes spawn per teammate?
**Expected:** 1 claude process per teammate
**Actual: ZERO separate processes! Teammates are in-process threads.**
```
# Only 2 claude processes (both Codeman sessions, none for teammates):
25405 claude --dangerously-skip-permissions --session-id 236f004f... (our main session)
383633 claude --dangerously-skip-permissions --session-id 461daa80... (test session + 3 teammates)
# Process tree for test session:
claude(383633)─┬─{claude}(383635)
├─{claude}(383636)
├─... (22 threads total)
└─{claude}(399252)
```
**In-process mode = threads, not processes.** All 3 teammates run as threads within the single `claude` process (PID 383633). This explains:
- No separate PIDs to track
- No `/proc/{pid}/environ` for individual teammates
- Lower resource overhead
- Shared env vars automatically
---
## Question 9: Do teammates inherit Codeman env vars (hook events)?
**Expected:** Yes, if child processes
**Actual: YES, trivially — they're in-process threads**
Since teammates are threads in the lead's process (PID 383633), they share the exact same environment:
```
CODEMAN_SCREEN=1
CODEMAN_SESSION_ID=461daa80-94ec-4e5e-a1bb-0518f78311bc
CODEMAN_SCREEN_NAME=codeman-461daa80
CODEMAN_API_URL=http://localhost:3000
```
The `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1` env var was set via `settings.local.json`'s `env` key, which Claude Code reads at startup and sets on its process.
**Hook events:** The lead session's hooks (Notification, Stop) apply to the whole process. Teammate-specific hooks (`TeammateIdle`, `TaskCompleted`) are defined in the same `settings.local.json` and would fire for the lead's session.
---
## Question 10: Does subagent-watcher pick up teammates automatically?
**Expected:** Probably not
**Actual: YES! subagent-watcher detects teammates automatically.**
Teammates create transcript files in the standard subagent path:
```
~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{id}.jsonl
```
Codeman's `/api/subagents` endpoint returned all 3 teammates as active subagents. They're indistinguishable from regular Task tool subagents except:
1. Their `description` field starts with `<teammate-message teammate_id= team`
2. They can be cross-referenced with `~/.claude/teams/{name}/config.json`
3. They tend to be longer-lived than regular subagents
**Sub-subagents:** Teammates also spawn their own Task tool subagents (6 additional agents detected), creating a 3-level hierarchy: Lead → Teammates → Sub-subagents.
---
## Filesystem Event Timeline
```
06:45:01 Session transcript created
06:45:05 ~/.claude/teams/ created
06:45:05 ~/.claude/teams/research-watchers/ created
06:45:05 config.json created (lead member only, 620 bytes)
06:45:05 ~/.claude/tasks/research-watchers/ created with .lock
06:45:11 Task 1.json created (via .lock.lock directory lock)
06:45:13 Task 2.json created
06:45:15 Task 3.json created
06:45:18 inboxes/ directory created
06:45:18 fs-researcher.json inbox created (task_assignment message)
06:45:18 perf-researcher.json inbox created
06:45:19 api-researcher.json inbox created
06:45:26 config.json updated (fs-researcher added, 1886 bytes)
06:45:26 Subagent agent-ae50544.jsonl created (fs-researcher)
06:45:26 Task 4.json created (internal: fs-researcher tracking)
06:45:30 config.json updated (perf-researcher added, 3188 bytes)
06:45:30 Subagent agent-aa20c65.jsonl created (perf-researcher)
06:45:30 Task 5.json created (internal: perf-researcher tracking)
06:45:34 config.json updated (api-researcher added, 4551 bytes)
06:45:34 Task 6.json created (internal: api-researcher tracking)
06:45:35 Subagent agent-a29de32.jsonl created (api-researcher)
06:45:35+ Teammates working, additional subagent transcripts appearing
06:47:xx Tasks completed, shutdown_requests sent to teammate inboxes
06:47:56 team-lead.json inbox created (teammates reporting back)
06:47:57 config.json updated multiple times (member removal?)
06:48:02 CLEANUP: all inbox files deleted
06:48:02 CLEANUP: inboxes/ directory deleted
06:48:02 CLEANUP: config.json deleted
06:48:02 CLEANUP: research-watchers team directory deleted
06:48:02 CLEANUP: all task files deleted (1-6.json + .lock)
06:48:02 CLEANUP: research-watchers task directory deleted
```
---
## Web UI Observations
**Terminal output:**
- Task list appears with checkboxes: `☐ Research Node.js fs.watch on Linux vs macOS`
- Checkboxes fill in as tasks complete: `☑ Research Node.js fs.watch...`
- Each task shows assigned teammate: `(@fs-researcher)`
- Spinner shows active teammate with progress
**Status bar:**
- Shows team member selector: `@main @api-researcher @fs-researcher @perf-researcher`
- Hint: `shift+↑ to expand` and `ctrl+t to show teammates`
- Standard bypass permissions and token count still visible
**Subagent floating windows:**
- Teammates DID appear as subagent floating windows in Codeman's web UI
- They show the standard subagent info (model, tool calls, description)
- Sub-subagents (teammates' own Task tool usage) also appear
**In-process mode specifics:**
- No new terminal windows or panes
- Everything renders in the single terminal session
- Shift+Up/Down would switch between teammate views (not tested interactively)
---
## Conclusions & Key Surprises
### Surprises vs expectations
1. **Inboxes ARE filesystem-based** — contrary to docs saying "SendMessage tool". It's a hybrid: the tool writes to filesystem inboxes.
2. **Teammates are threads, not processes** — no new OS processes, just threads within the lead's claude process.
3. **Teammates appear as standard subagents** — existing subagent-watcher infrastructure works out of the box!
4. **Config grows incrementally** — members are added one-by-one, not all at once.
5. **Internal tracking tasks** — tasks 4-6 with `_internal: true` track teammate spawn state.
6. **Auto-cleanup** — lead automatically cleaned up ALL artifacts after shutdown.
7. **Sub-subagents** — teammates can spawn their own Task tool subagents (3-level hierarchy).
8. **`teammateMode` is NOT a valid settings key** — display mode defaults to `in-process`.
### Design implications for Codeman
1. **TeamWatcher can be simple** — just poll `~/.claude/teams/` for directories + parse config.json
2. **Subagent-watcher already works** — no new infrastructure needed for teammate transcript tailing
3. **Idle detection needs team awareness** — check config.json members before declaring idle
4. **Message interception is possible** — watch inbox JSON files for real-time message tracking
5. **Task visualization is straightforward** — parse numbered JSON files in task directory
6. **No process tracking needed** — teammates are threads, not separate processes
7. **Distinguish teammates from subagents** — use description prefix `<teammate-message` or cross-reference config.json
### What to build first
1. **Team-aware idle detection** — highest priority, prevents premature respawn
2. **TeamWatcher** — poll `~/.claude/teams/` for team creation/removal
3. **Team tasks API** — parse task JSON files for UI display
4. **Teammate badge in subagent windows** — mark teammate subagents differently from regular ones
5. **Message timeline** — parse inbox files for inter-teammate communication display
### What we DON'T need to build
- Process discovery for teammates (they're threads)
- Custom transcript tailing (subagent-watcher handles it)
- Separate teammate window infrastructure (subagent windows work)
-565
View File
@@ -1,565 +0,0 @@
# HTTP API Reference
Codeman's HTTP API is a **stable contract** as of 1.0 — see
[`versioning-policy.md`](versioning-policy.md) for the SemVer guarantee. This page
defines the response envelope, status codes, error codes, versioning, and the SSE
event channel.
## Versioning
- The stable, public surface is served under **`/api/v1/...`**. Pin external
clients to this prefix.
- The unversioned **`/api/...`** paths are a permanent alias of the current
version (what the bundled web UI uses). They are kept working, but new external
integrations should use `/api/v1`.
- Breaking changes to the contract ship under a new prefix (`/api/v2`); `/api/v1`
keeps its semantics. Additive changes (new endpoints, new optional fields, new
error codes) are non-breaking and may appear in a minor release.
- The implementation rewrites `/api/v1/*` → `/api/*` at the server level
(`rewriteApiV1Url` in `src/web/server.ts`).
## Response envelope
Every JSON response uses one uniform envelope, applied centrally by a
`preSerialization` hook (`src/web/server.ts`) — handlers return bare data and the
hook wraps it:
**Success** — HTTP `2xx`:
```json
{ "success": true, "data": <payload> }
```
`data` is the endpoint's payload (object, array, or value). Endpoints with no
payload return `{ "success": true, "data": {} }`.
**Error** — HTTP `4xx`/`5xx`:
```json
{ "success": false, "error": "human-readable message", "errorCode": "NOT_FOUND" }
```
`ApiResponse<T>` in `src/types/api.ts` is the canonical type.
> Non-JSON endpoints are exempt from the envelope: `GET /api/sessions/:id/file-raw`,
> `GET /api/sessions/:id/tail-file` (SSE), `GET /api/download`,
> `GET /api/screenshots/:name`, `GET /q/:code` (QR redirect), and the
> `GET /ws/sessions/:id/terminal` WebSocket upgrade.
> The [agent wait endpoints](#long-polling-agent-wait) use the normal envelope but
> are the only JSON endpoints that deliberately **hold the connection open**, for up
> to 600 s. Proxy operators and HTTP clients with a global read timeout need to know
> that before pointing them at Codeman.
⚠️ **A `401` is the one status that is not an envelope.** Authentication is rejected
in a request hook, before any handler runs, and it replies with the bare string
`Unauthorized` (`Unauthorized: hook secret required` on the hook path) plus
`WWW-Authenticate: Basic realm="Codeman"`. There is no `success`, no `error`, and no
`errorCode`, because the wrapping hook only wraps object payloads. So a client that
pipes every response straight into a JSON parser dies with a parse error rather than
reporting an auth failure, which is a confusing way to discover that a password is
set. Branch on the HTTP status **before** parsing.
## Error codes → HTTP status
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
the HTTP status.
| `errorCode` | HTTP | Meaning |
|-------------|------|---------|
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
| `UNAUTHORIZED` | 401 | Authentication required or failed |
| `NOT_FOUND` | 404 | Resource does not exist |
| `SESSION_BUSY` | 409 | Session is busy |
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
| `ALREADY_EXISTS` | 409 | Resource already exists |
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
| `RATE_LIMITED` | 429 | Too many requests |
| `INTERNAL_ERROR` | 500 | Unexpected server error |
Adding a new error code is non-breaking; removing or renaming one is a major change.
## Long-polling (agent wait)
Three calls block until something happens instead of answering immediately. They
exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse
events inline.
| Call | Blocks until |
|------|--------------|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
`GET .../wait`. It registers the waiter **before** writing, which closes the window
in which a separate wait sees the session still idle from the previous turn and
answers instantly with the wrong turn's result. Use it whenever you send a prompt
and want to know when that prompt is done.
### Three semantics that break callers who assume otherwise
**1. A timeout is HTTP `200`, not an error.** A wait that ends without its signal
returns `{"success":true, ...,"wait":{"timedOut":true,"signal":null}}`. The
intended pattern is a client-side loop over short waits, because `tailscale serve`
and cloudflared can both cut an idle connection, and turning every poll boundary
into a `4xx` would make that loop indistinguishable from a real failure. `408` is
auto-retried by several clients (silently doubling the polling load), `504` is what
a genuine tunnel failure looks like, and `204` cannot carry `waitedMs` / `status` /
`limitPaused`. Reserve error handling for the four codes in the table below.
**2. `stop` and `blocked` fire only for `claude` sessions.** Both come from Claude
Code hooks, and no other mode installs them: `shell` runs no agent, and the external
CLIs (`opencode`, `codex`, `gemini`, `antigravity`, `pi`) render their own TUIs and post
no hooks. For every non-`claude` mode only `idle`, `working` and `exit` are
accepted, and of those only `exit` is dependable: see the caveats under
[Signals](#signals) before building on `idle`. Requesting `stop` or `blocked`
**explicitly** on such a session is a
`400`; omitting `until` never fails, the server just drops them from the default set
and echoes the narrowed set back as `wait.until`. Three more places hooks can go
missing even in `claude` mode: a **Docker case** needs
`CODEMAN_DOCKER_BRIDGE_HOOKS=1`, since a container cannot reach a loopback-bound
Codeman (without it, only `idle` / `working` / `exit` work); a **remote-SSH
case** runs the agent on another host, whose hooks may never reach this server at
all; and a case whose hook config was written by **Codeman < 1.13.0 against an
`--https` install** carries hook curls without `-k`, which TLS-fail silently (the
hook line ends in `|| true`). Codeman now writes `curl -sk` and repairs a stale
case config the next time a session starts in that case. When in doubt, ask for
`stop,idle,exit` so a session without hooks still resolves on the heuristic
signal.
**3. `from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, on resize, and on any TUI redraw, and a repaint arrives as
ordinary output, so text that was already on screen can satisfy a fresh wait. This
was observed live: a marker echoed a minute earlier matched instantly on a new
`from=now` wait. It is inherent to running the agent under a multiplexer, so the
contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
`echo $MARK`, then wait on `$MARK`), never a generic string like `BUILD OK`.
### Signals
| Signal | Source | Actually fires for |
|--------|--------|--------------------|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
is omitted is `stop,idle,exit` (`exit` is in there so a worker that crashes resolves
the wait promptly instead of burning the caller's whole timeout on something that
can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit`
once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no *further* `idle`, so it is the
startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died *inside* its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
milliseconds, and reading that as "the worker died" is wrong: it means start it, or
wait for it to come up. `status` is carried alongside so nothing is hidden. The
alternative (trusting `status`) is worse, because a dead PTY parks the session at
`status: "idle"`, which would answer the default wait with `immediate: true` for a
worker that has crashed. A worker dying while a wait is parked resolves it within
a few seconds (a background death-watcher), not at the timeout.
⚠️ **`blocked` is reachable less often than the table suggests.** It fires on two
hooks, and the default configuration suppresses one of them: Codeman spawns claude
with `--dangerously-skip-permissions`, so permission prompts do not happen unless the
instance is switched to the `auto` Claude mode (App Settings), or the caller is a
multi-user account without the bypass grant, which is forced to `--permission-mode
auto`. What does still fire under the default is `elicitation_dialog`, the agent
asking the user a question. So `until=stop,blocked,exit` is a reasonable belt on a
long turn, but a worker that never comes back is far more likely to be working than
blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
instead. The same caution applies to the external CLIs.
### Readiness is not a signal
Nothing here reports "the agent is ready for a prompt", and no combination of
`until`/`fresh` synthesizes one. A freshly created session reads as `exit` (above),
and a `claude` worker in a brand-new case comes up on the CLI's **trust dialog**,
which contains a ❯ prompt of its own. Send-and-wait posted at that moment types the
prompt into the dialog, where the `\r` never gets past it, while the session's
startup `idle` lands inside the wait window: the wait resolves on `idle` in a
couple of seconds with `timedOut: false`, which looks exactly like a completed
turn.
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
dialog only as the bounded fallback (`trust` matched → send `\r` → wait for
`bypass` again). Do not probe `trust` first and Enter blindly: the dialog text
stays in the terminal buffer for the life of the session, so a `trust` probe with
`from=buffer` keeps matching on every later run and the Enter lands in a ready
composer. A worked version is in
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
```bash
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
```
Both GET wait routes answer with `Cache-Control: no-store`, because the documented
pattern polls one identical URL in a loop and a cached `{"timedOut":true}` would
turn that loop into a busy spin. `POST .../input` sends no cache header (it is a
POST, which is not heuristically cacheable).
⚠️ **Unknown query parameters are ignored, not rejected**, with one exception
(`regex`, below). In particular `match=` on `/wait` is silently dropped and you get
a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
instead of waiting on the wrong thing. The reasoning is in
[`architecture-invariants.md`](architecture-invariants.md#agent-wait-primitives).
#### What the matcher actually sees
The matcher scans the raw PTY stream, **normalized**: ANSI escape sequences are
stripped — CSI, OSC, and the charset-designation escapes a stock bash prompt emits
on every line (`ESC ( B`), so `match=tnode:` matches a prompt that renders
`…@tnode:` — a partial escape arriving at a chunk boundary is held back until its
tail arrives, and a match may straddle PTY chunks: `printf STRAD; sleep 1; printf
DLEQQ` is matchable as `STRADDLEQQ` (all measured live). Three caveats remain:
⚠️ **It is still the byte stream, not the rendered pane.** `GET .../terminal`
answers from a tmux screen capture (`data.source: "mux-visible"`), the finished
picture; the matcher sees the stream that painted it. For linear output the two
agree once escapes are stripped, but a full-screen TUI composes its picture with
cursor positioning, so what the pane shows and what the stream carries can differ.
Seeing your string in `terminal?tail=` makes a match likely, not guaranteed.
⚠️ **A TUI's text can arrive without its spaces.** Claude Code positions words
with cursor moves rather than printing spaces, so screen text can reach the
matcher as `Quicksafetycheck:Isthisaprojectyoucreated...`. Whether a given phrase
keeps its spaces depends on how the TUI happened to draw it (measured: `I trust
this folder` matched, `Quick safety check` did not), so a multi-word `match`
against a TUI pane is unreliable rather than impossible. Match a **single
space-free token**, ideally one you printed yourself. Plain command output (a
shell worker, an `echo`) keeps its spaces.
⚠️ **The returned `snippet` is a rendering of the matched text, not a quotation of
it.** It is cut from the same normalized stream the match ran against, then
cleaned for display: remaining raw control bytes are removed (an agent pipes the
snippet into its own terminal, so a worker's bytes must not be able to reset that
display) and blank runs are collapsed. A printable needle that matched will appear
in it; a needle containing control bytes or a blank run may not survive verbatim.
```bash
MARK="DONE_$RANDOM"
curl -sG "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'timeout=120000'
```
Build the query with `-G --data-urlencode` rather than by hand: a `+` in a
hand-written query string decodes to a space.
### `POST /api/v1/sessions/:id/input` with `wait`
Two optional fields on the existing endpoint:
| Field | Type | Notes |
|-------|------|-------|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would
reject it, which has shipped as a real bug twice.
The input must end with `\r` (a real carriage return in the JSON string): Enter is
sent only when the input contains one, so text without it is typed onto the
worker's prompt but never submitted, and the wait then runs its full timeout on a
turn that never started. Verified live; this is the most common silent failure on
this endpoint.
```bash
curl -s -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true,"clientId":"agent-1","seq":1,
"wait":"stop","waitTimeout":600000}'
```
A **tagged duplicate** (a `clientId` + `seq` pair the server has already applied)
still honors `wait`, because the caller's question is unanswered, but it answers
from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
any of them:
```json
{ "success": true, "data": {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"status": "idle",
"limitPaused": false,
"wait": {
"signal": "stop", "until": ["stop", "idle", "exit"],
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
"waitedMs": 8421, "timeoutMs": 60000
}
}}
```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
`status` and `limitPaused`. `POST .../input` **without** `wait` is unchanged and
still returns `{"success": true, "data": {}}`.
⚠️ `delivered: false` has **two** meanings, and they must be told apart by
`duplicate`: with `duplicate: true` the input was suppressed as an already-applied
redelivery (harmless, the turn it refers to may be long over), while with
`duplicate: false` the **write failed** (typically no PTY behind the session). A
client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success.
| Field | Type | Meaning |
|-------|------|---------|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order:
1. `wait.signal !== null` (or `wait.matched === true`): the thing happened.
2. `wait.timedOut`: a poll boundary. Loop again.
3. `wait.ended` or `wait.aborted`: the wait answered nothing, because the session is
gone or was never running. Re-check the session instead of looping.
`wait.immediate` is not a fourth outcome: it rides along with the first one and
means the condition already held at call time, so nothing was actually waited for.
If that is not what you meant, you wanted `fresh=1` or the send-and-wait form. Note
that `{"signal":"exit","immediate":true}` on a session you just created is the
not-started-yet case, not a crash.
**The timeout is clamped, so read it back.** A request for 1800000 ms is silently
reduced to the server's ceiling (600000 ms by default, operator-tunable), and a
request for 1 ms is raised to 1000 ms. `wait.timeoutMs` is the value that was
applied. Without checking it, a caller that asked for 30 minutes and got 10 will
read the timeout as "the worker is wedged" and kill a session that was working fine.
### Errors
| `errorCode` | HTTP | When |
|-------------|------|------|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
error message names the cap that was hit.
⚠️ A `401` is **not** in this table and is not an envelope at all (see
[Response envelope](#response-envelope)). It matters most here: a polling loop that
pipes each wait straight into `jq` fails with a parse error on every iteration
against a password-protected server, which reads as "the wait endpoints are broken".
Check the status first.
The per-session cap is a **combined** budget: signal waiters and output waiters
count against the same 16, not 16 of each. An abandoned request no longer holds its
slot, because the routes release the waiter when the client disconnects, but a
client that opens many concurrent waits against one session will still hit the cap.
## Session lineage (`parentSessionId`)
A create request may name the session that spawned it, which the web UI draws as a
line between the two tabs. Accepted on `POST /api/v1/sessions` and
`POST /api/v1/quick-start`, either way:
```bash
# as a body field
-d '{"caseName":"worker-1","mode":"claude","parentSessionId":"'"$CODEMAN_SESSION_ID"'"}'
# or as a header, which is what an agent driving many spawns should use: set it once
# on the curl invocation and every spawn call carries it
-H "X-Codeman-Parent-Session: $CODEMAN_SESSION_ID"
```
The body field wins if both are present. The value is resolved against live sessions
(exact id, or a unique prefix of at least 8 characters) and must belong to the same
owner as the session being created.
**It cannot fail your spawn.** An unknown, stale, foreign or malformed value is
silently dropped and the session is created without lineage — never a `400`. It is
also pure decoration: it confers no permission, and a child is unaffected by its
parent exiting. It appears on session state as `parentSessionId` (absent when
unresolved) and survives a server restart.
## Approvals Inbox
Cross-session queue of prompts waiting on a human (permission dialogs,
AskUserQuestion questions, idle prompts). Claude-mode sessions only; items are
in-memory (a server restart drops them; the next prompt re-fires the hook).
Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
toolSummary?, message?, cwd?, context?, options?: {n, label}[] }`. `context`
is the ANSI-stripped visible pane frame; `options` is present only when the
dialog's numbered choices parsed confidently.
- `POST /api/v1/approvals/:id/answer` with `{ action: 'approve' }` (sends the
digit `1`), `{ action: 'deny' }` (sends Esc), `{ action: 'option', option: n }`
(sends the digit; accepted only when `n` is among the item's parsed
`options`), or `{ action: 'text', text }` (idle prompts only; submits the
line as a prompt). `404 NOT_FOUND` when the item is no longer pending,
`409 CONFLICT` when the dialog left the screen or another actor answered
first, `422 OPERATION_FAILED` when the session refused input.
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
SSE events: `approval:pending` (full item), `approval:updated` (context/options
re-captured), `approval:resolved` (`{ id, sessionId, kind, resolution }` with
`resolution` one of `answered | resolved_in_terminal | superseded |
session_ended | dismissed | expired`).
## Read My Mind intent profiles
Per-case profiles of what the user is trying to accomplish: user/agent-stated
goals plus the user's recently submitted prompts, captured from the Claude
session transcript while the opt-in `readMyMindEnabled` setting is on (default
OFF). Keyed by owner + workingDir, so the profile survives `/clear`, respawns,
and session churn. Stored in `~/.codeman/intents.json` (mode 0600); never fed
into `/api/v1/search`. Design: [`readmymind-plan.md`](readmymind-plan.md);
user guide: [`readmymind.md`](readmymind.md).
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
50, each <= 500 chars). A case with nothing recorded answers an empty
profile with `updatedAt: 0`; nothing is persisted by reads.
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
schema) replaces the goals text and answers the updated profile.
`400 INVALID_INPUT` on over-long or unknown fields.
- `DELETE /api/v1/sessions/:id/intent` -> `{ deleted: boolean }` forgets the
case's profile entirely.
- `POST /api/v1/sessions/:id/readmymind` predicts the user's next prompt:
a one-shot model call over the intent profile plus live session signals
(pending approval dialog, transcript tail, git state, run-summary events,
sibling sessions). Body is optional; the rethink flow passes
`{ steer?, rejected? }` (strict schema: `steer` <= 2000 chars, `rejected`
up to 10 strings <= 1000 chars). Answers
`{ suggestions: { prompt, why, kind }[], durationMs }` with 1-3 suggestions
(`kind`: `continue` | `verify` | `redirect`; prompts are single-line).
Claude-mode sessions only (`400 INVALID_INPUT` otherwise); one prediction in
flight per session (`409 CONFLICT`); predictor failures answer
`502 OPERATION_FAILED`. Takes 5-90 s and costs real tokens. Suggestions are
only ever returned, never sent: submitting one is the caller's explicit act.
All four enforce session ownership in multi-user mode; a foreign session id
answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the
same directory are distinct by construction.
## Voice dictation
Browser dictation transcribed through this server's Claude Code login, i.e. the
same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synced
`claudeVoiceEnabled` setting (default OFF). Design:
[`claude-voice-plan.md`](claude-voice-plan.md).
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
signed in to Claude Code on the server), `expired` (the access token elapsed;
running any Claude session refreshes it) or `malformed`. The OAuth token
itself is never returned by this or any other endpoint.
- `GET /ws/voice/stream?language=&keyterms=` (WebSocket, not under `/api`)
relays one dictation. Client sends binary frames of signed 16-bit
little-endian PCM, 16 kHz mono (<= 64 KB per frame), plus JSON control frames
`{"t":"finalize"}` (ask for the final transcript) and `{"t":"stop"}`. Server
sends `{"t":"ready"}`, `{"t":"transcript","text","final"}` (each frame is the
WHOLE running transcript, not a delta), `{"t":"error","message"}` and
`{"t":"closed"}`. Close codes: `4003` disallowed Host/Origin, `4004`
unavailable (reason in the close reason), `4008` too many concurrent streams.
Streams are capped in count and length (`src/config/voice.ts`).
## Authentication
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
`codeman_session` cookie. When enabled, unauthenticated requests get
`401 UNAUTHORIZED`; rate-limited requests get `429 RATE_LIMITED`. See
[`security-architecture.md`](security-architecture.md).
## SSE event channel
`GET /api/events` is a Server-Sent Events stream (`text/event-stream`); each
message is `event: <name>` + `data: <json>`. The event-name registry
(`src/web/sse-events.ts`, mirrored in `src/web/public/constants.js`) is part of
the stable contract — event names are not renamed without a major bump. An
optional `?sessions=<id,...>` filter suppresses only the high-volume terminal
stream; lifecycle/metadata events are delivered to all clients regardless.
### `sse:heartbeat` (liveness)
Every 15s the server writes a `sse:heartbeat` frame to every connected client:
```
event: sse:heartbeat
data: {"t":1755100000000}
```
`t` is the server's epoch-ms timestamp at write time. The frame carries no
application state and can be ignored for correctness. It exists so a client can
tell a live stream from a dead one: an `EventSource` whose connection has been
idle-closed by a proxy (or that resumed from sleep on a stale socket) keeps
delivering nothing without ever firing `onerror`. Clients that care should treat
silence longer than about three intervals as a dead stream and reconnect, which
is what the bundled frontend does.
This replaced a `:keepalive` SSE **comment**, which served the same
proxy-flushing purpose but is invisible to `EventSource` by spec and so could
never be observed by a client. Consumers written against the old behavior are
unaffected: `EventSource` dispatches only events that have a registered
listener, so an unknown event name is dropped.
## Consuming from JavaScript
The bundled frontend reads responses through `_apiJson()`
(`src/web/public/api-client.js`), which unwraps `{success:true,data}` → `data` and
returns `null` on a non-2xx / `{success:false}` response. External clients should
do the same: check the HTTP status (or `body.success`), then read `body.data`.
-106
View File
@@ -1,106 +0,0 @@
# Approvals Inbox (design)
One cross-session inbox for every prompt that is waiting on a human: permission dialogs, questions (AskUserQuestion / elicitation), and idle prompts. Cards are answerable in place (option digits, Esc, or a typed prompt) from desktop, phone overview, and push notification action buttons. Inspired by Cloudflare OS's Gatekeeper approval queue (https://github.com/cloudflare/cloudflare-os, asynchronous human-in-the-loop approvals): with a fleet of sessions the human is the bottleneck, and today answering means finding the right tab.
## Problems this fixes (all real today)
1. **No cross-session surface.** Pending prompts exist only as per-tab alert colors (`tab-alert-action`/`tab-alert-idle`) and NEEDS YOU rows on the phone overview. Answering means switching to the session and typing.
2. **Alerts die on reload.** `pendingHooks` lives only in `app.js` memory, fed by transient SSE `hook:*` events. A page reload (or a phone browser evicting the tab) silently loses every pending alert. There is no server-side record.
3. **Push Approve/Deny buttons are dead.** `PUSH_EVENT_MAP` already attaches `approve`/`deny` actions to permission pushes, and `sw.js` forwards `event.action` to the page, but the `notification-click` handler in settings-ui.js ignores it (and when no tab is open, the action is dropped entirely). The buttons render on the lock screen and do nothing.
4. **Card context is missing.** The frontend handlers read `data.question` / `data.message` / `data.tool`, but `sanitizeHookData` never forwards `message`, so notifications show generic fallback text.
## Scope
- Claude mode only (hooks fire only for `claude`; external CLIs keep their output-stabilization heuristics and get no inbox items). This mirrors the wait-primitive `stop`/`blocked` gating.
- Permission prompts occur for sessions running `ClaudeMode` `normal` / `auto` / `allowedTools` (and the trust-folder dialog even under skip-permissions). Question and idle prompts occur in every mode including `dangerously-skip-permissions`.
- In-memory store (plus the frontend seeding from it on load). Server restart drops items; hooks re-fire on the next prompt. No new state file in v1.
## Data model
At most **one active item per session**: the Claude TUI shows one dialog at a time, so a new prompt event supersedes the session's previous item (resolution `superseded`).
```ts
interface ApprovalItem {
id: string; // `${sessionId}:${seq}`
sessionId: string;
sessionName: string;
kind: 'permission' | 'question' | 'idle';
createdAt: number;
toolName?: string; // from sanitized hook data
toolSummary?: string; // command / file_path / description, already bounded
message?: string; // Notification hook `message` (newly allowlisted)
cwd?: string;
context?: string; // ANSI-stripped visible pane frame tail, ≤ 4000 chars
options?: { n: number; label: string }[]; // parsed from context when confident
}
```
Resolutions (server-emitted, item removed from pending): `answered` (via inbox), `resolved_in_terminal` (stop / elicitation_complete / elicitation_response / session went working), `superseded`, `session_ended`, `dismissed`, `expired` (12h TTL sweep).
## Backend
### Store: `src/approval-inbox.ts`
Module-level singleton in the style of `session-wait-registry.ts` (pure, no `Session` import, injected emit callback so there is no import cycle with the server):
- `notePrompt(info)` creates/supersedes the session's item; schedules ONE re-capture ~600ms later (the Notification hook can fire before the dialog finishes painting) which updates `context`/`options` and emits `approval:updated`.
- `resolveForSession(sessionId, reason)`, `dismiss(id)`, `answerable(id)`, `listPending()`, `stop()` (clears timers; tests).
- Option parsing (pure, unit-tested): consecutive `❯? N. label` lines, 2..6 options, labels ≤ 120 chars. Parsed options gate which digits the answer endpoint accepts; when parsing fails the card falls back to Approve(1)/Deny(Esc) only.
- TTL: items expire after 12h (checked on read + a lazy sweep; no standing interval).
### Wiring
- `hook-event-routes.ts`: on `permission_prompt` / `elicitation_dialog` / `idle_prompt`, call `notePrompt` with sanitized data + a pane capture callback (`mux.capturePaneBuffer(muxName)` visible frame, ANSI-stripped via existing utils; fall back to `session.terminalBuffer` tail). On `stop` / `elicitation_complete` / `elicitation_response`, `resolveForSession(id, 'resolved_in_terminal')`.
- `session-listener-wiring.ts`: `working` listener resolves **idle items only** (`working` is heuristic and can flap mid-turn, so it must never clear a pending permission/question dialog); `exit` resolves with `session_ended`. Same singleton-import pattern as `sessionWaits`.
- Session delete route: resolve with `session_ended`.
- **New hook matchers** `elicitation_complete` + `elicitation_response` added to `generateHooksConfig()`, `HookEventType`, `HookEventSchema`, and both SSE registries. `refreshStaleCodemanHooks` gets a staleness probe for them (`hooksJson.includes('elicitation_complete')`) so existing cases heal on next Claude spawn, exactly like the `-k`/secret/marker probes.
- `sanitizeHookData`: allowlist `message` (bounded 500 chars). This also un-deadens the existing notification text paths.
### Routes: `src/web/routes/approval-routes.ts`
Normal authed API (NOT the hook-secret bypass), `ApiResponse` envelope, Zod schemas in `schemas.ts`:
- `GET /api/approvals` → pending items, multi-user filtered by `canAccessOwned` (same policy as session lists).
- `POST /api/approvals/:id/answer` body `{ action: 'approve' | 'deny' | 'option' | 'text', option?, text? }`:
- `approve` → `writeViaMux('1')` (option 1 is always plain Yes; no Enter, menus react to the digit).
- `deny` → `writeViaMux('\x1b')` (Esc is the official No/cancel; precedent: auto-resume sends Esc the same way).
- `option` → digit `String(n)`; accepted only when `n` is within the item's parsed options (prevents blind digit-poking at an unparsed dialog).
- `text` → `idle` items only: single line, embedded newlines stripped, sent as `text\r` (the `\r` discipline from CLAUDE.md).
- Guards: item still pending (404 otherwise), session exists + ownership via `findSessionOrFail`, session mode installs hooks. **Answer-time re-capture**: for items whose frame parsed options, the pane is re-captured before sending; if the dialog no longer parses, the item resolves and the answer is refused with 409 (the keystroke would land in whatever now has focus). Marks `answered` BEFORE the write so a double-tap cannot double-send; rolls back to pending if the write fails.
- `POST /api/approvals/:id/dismiss` → remove without keystrokes.
### SSE
`approval:pending`, `approval:updated`, `approval:resolved` in `sse-events.ts` + `SSE_EVENTS` in constants.js (the parity test pins the sync). Broadcasts carry `sessionId`, so multi-user SSE scoping applies unchanged.
### Push
- `sendPushNotifications` payload gains `approvalId` for the three hook events. Both `approvalId` and the Approve/Deny `actions` are **gated on the opt-in setting**: with it off, permission pushes carry no buttons at all (pre-inbox they rendered and did nothing, so stripping them is the honest shape).
- `sw.js` `notificationclick`: when `event.action` is `approve`/`deny`, POST `/api/approvals/:id/answer` directly from the worker (same-origin, cookie credentials) so the buttons work **with no tab open**; on failure fall back to focusing/opening a tab. Non-action clicks keep today's behavior.
- Page-side `notification-click` handler: honor `action` instead of dropping it (also setting-gated, for stale notifications sent before the toggle flipped).
- Question/idle pushes keep no action buttons (options vary per dialog); tapping opens the inbox.
## Frontend
New module `approvals-ui.js` (@loadorder 11.2, after panels-ui.js), prettier-formatted (not added to `.prettierignore`).
- **Seed on connect**: `GET /api/approvals` on init and SSE reconnect; each pending item re-feeds `setPendingHook(...)` so tab alerts and the phone overview survive reload (fixes problem 2 with zero changes to the alert state machine).
- **Desktop**: header bell `btn-approvals` with count badge. Ships default-hidden via marker class `btn-approvals--hidden` (same policy as the attachments button, so `test/mobile-header-buttons-policy.test.ts` excludes it from the default-visible enumeration); JS shows it only while count > 0. Click toggles a drawer of cards: session name + kind, tool/message summary, mono context block, buttons rendered from parsed options (else Approve/Deny), plus Dismiss and Open session. Esc closes; existing z-index layers respected.
- **Phone**: header button stays hidden (`mobile.css`); the phone surface is the overview's NEEDS YOU section, whose rows gain inline ✓/✗ buttons for permission items (tap-through to the session remains the row's main action). Toolbar classes/status language rules from the mobile-overview section of CLAUDE.md apply.
- **i18n**: new strings registered in i18n.js (en + zh-CN); status words carry `data-i18n-skip` where they would collide (mirroring the overview pills).
- **Setting**: `approvalsInboxEnabled`, synced (in `SettingsUpdateSchema`), **default OFF** (owner decision: the entire feature is opt-in, meaning no bell, no drawer, no overview strips, no seeding, and no push action buttons until enabled in App Settings → Panels). Only the store and answer endpoints keep running regardless, so flipping the toggle ON surfaces anything already pending immediately, with no restart.
## Race honesty
The prompt can be answered in the terminal a moment before an inbox answer lands; then the keystroke would hit whatever now has focus (worst case: a digit typed into the composer, not submitted, since no `\r` is ever sent for menu answers). Mitigations, in order: answer-time re-capture (the dialog must still parse on screen or the answer is refused), answered-before-write marking, digit-only/Esc-only writes for menus, and the card's context block showing what the pane looked like when captured. This is the same class of risk `writeViaMux` automation (auto-resume, respawn) already accepts.
## Tests
- `test/approval-inbox.test.ts`: supersede per session, every resolution path, TTL, option parsing fixtures (2-option, 3-option with ❯, unparseable frame), re-capture update.
- `test/routes/approval-routes.test.ts` (`app.inject`, no port): list; hook event creates item; answer approve/deny/option writes the exact bytes (test-PTY echo asserts them); text answers restricted to idle; 404 unknown id; 409 answered twice; option out of range rejected; multi-user scoping.
- Existing suites extended: hook-event schema accepts the two new events; `sanitizeHookData` forwards bounded `message`; SSE parity + mobile-header policy pass as-is by construction.
## Docs
- CLAUDE.md: Key Patterns entry + SSE/route counts + frontend load order.
- `docs/api-reference.md`: the two endpoints + three SSE events (additive, fine under the 0.9.x contract).
File diff suppressed because one or more lines are too long
@@ -1,331 +0,0 @@
# Plan: Background Keystroke Forwarding (Local Echo Mode)
> **Supersedes**: This document merges two previous plan drafts into a single authoritative reference:
> - `docs/background-keystroke-forwarding-plan.md` (detailed design doc)
> - `.claude/plans/jazzy-bubbling-salamander.md` (Claude-generated implementation plan)
>
> The docs plan was used as the base. The Claude plan was a correct but simplified subset; its Context paragraph is incorporated below as a lead-in.
## Context
When local echo is enabled, keystrokes accumulate in the `LocalEchoOverlay.pendingText` and are only sent to the server when Enter is pressed. This means switching tabs loses the input from the actual Claude Code PTY (the overlay caches text client-side, but the PTY has nothing). If the session respawns or resets, accumulated input is lost entirely.
## Problem
When local echo is enabled, keystrokes accumulate **only** in `LocalEchoOverlay.pendingText` (a client-side string). Nothing reaches the server PTY until Enter is pressed. This creates three failure modes:
1. **Tab switch loses PTY state** — switching sessions saves overlay text to `localEchoTextCache` (a Map), but the actual Claude Code Ink process has no knowledge of what was typed. If respawn or `/clear` fires on that session, the cached text is meaningless.
2. **Session death loses input** — if the session crashes or respawns while text is pending in the overlay, that input is gone (localStorage backup `codeman_local_echo_pending` only survives page reloads, not session resets).
3. **Tab completion impossible** — pressing Tab with pending overlay text sends the raw Tab character to a PTY that has no knowledge of the typed text, so completion fails.
## Goal
Send every keystroke to the server in the background (debounced), while the overlay continues providing instant visual feedback. The overlay sits at z-index 7 with an opaque background over `.xterm-screen`, masking Ink's echo of the background-sent characters. Input persists in the actual Claude Code readline buffer across tab switches and respawns.
## Architecture
```
User keystroke
|
v
xterm.js onData(data)
|
+---> LocalEchoOverlay.addChar(data) [instant visual feedback]
|
+---> _localEchoBgBuffer += data [queue for background send]
| clearTimeout + setTimeout(50ms)
| |
| v (50ms debounce fires)
| _flushBgInput()
| |
| v
| _sendInputAsync(sessionId, buffer) [promise chain preserves order]
| |
| v
| POST /api/sessions/:id/input [{ input: "hel" }]
| |
| v
| session.write(inputStr) [direct PTY write, synchronous]
| |
| v
| Ink readline echoes "hel" [hidden behind overlay's opaque bg]
|
+--- On Enter:
1. clearTimeout(_localEchoBgTimer)
2. flush _localEchoBgBuffer via _sendInputAsync (remaining chars)
3. clear overlay
4. 120ms later: send \r via _sendInputAsync (Ink text/Enter split)
5. Ink processes "hello\r" → overlay gone, terminal visible with output
```
### Two Input Paths (important context)
The codebase has **two separate input paths** to the server:
| Path | Used by | Promise chain? | `useMux`? |
|------|---------|---------------|-----------|
| `_sendInputAsync()` (line 3626) | `onData` handler, `flushInput()` | Yes (`_inputSendChain`) | No (direct PTY write) |
| `sendInput()` (line 8755) | Mobile accessory bar, programmatic commands | **No** (raw `fetch`) | Yes (tmux `send-keys`) |
Background keystroke forwarding uses **only** the `_sendInputAsync` path, which guarantees ordering via the promise chain. The `sendInput()` path is unaffected and unmodified.
### Server-Side Input Flow
```
POST /api/sessions/:id/input { input: "hel" }
|
+-- useMux? No (default)
| session.write("hel") → ptyProcess.write("hel") [sync]
|
+-- useMux? Yes
session.writeViaMux("hel") → tmux send-keys -l "hel" [async]
```
Background sends use the default path (no `useMux`), which is a synchronous direct PTY write — faster than spawning a tmux subprocess for each character batch.
## Implementation
All changes in **one file**: `src/web/public/app.js`
### Step 1: Add background send state (in terminal setup, after line ~1999)
```js
this._localEchoBgBuffer = ''; // Characters queued for background send
this._localEchoBgTimer = null; // 50ms debounce timer ID
```
Add an atomic drain helper alongside existing `flushInput` (after line ~2008):
```js
// Atomically drain background buffer — returns contents and cancels pending timer.
// Single point of extraction prevents double-flush race conditions.
const drainBgBuffer = () => {
if (this._localEchoBgTimer) {
clearTimeout(this._localEchoBgTimer);
this._localEchoBgTimer = null;
}
const buf = this._localEchoBgBuffer;
this._localEchoBgBuffer = '';
return buf;
};
const scheduleBgFlush = () => {
if (this._localEchoBgTimer) clearTimeout(this._localEchoBgTimer);
this._localEchoBgTimer = setTimeout(() => {
this._localEchoBgTimer = null;
const buf = this._localEchoBgBuffer;
this._localEchoBgBuffer = '';
if (buf && this.activeSessionId) {
this._sendInputAsync(this.activeSessionId, buf);
}
}, 50);
};
```
**Why `drainBgBuffer` exists**: Every exit path (Enter, Ctrl+C, tab switch, echo disable) needs to flush the buffer AND cancel the timer atomically. Without a single extraction point, it's easy to forget one of the two operations, leading to double-sends when the timer fires after a manual flush.
### Step 2: Modify `onData` handler — local echo path (lines 2023–2067)
**Printable characters** (lines 2063–2067 → replace):
```js
if (data.length === 1 && data.charCodeAt(0) >= 32) {
this._localEchoOverlay?.addChar(data);
// Background: queue char for server send (50ms debounce batches rapid typing)
this._localEchoBgBuffer += data;
scheduleBgFlush();
return;
}
```
**Backspace** (lines 2024–2028 → replace):
```js
if (data === '\x7f') {
this._localEchoOverlay?.removeChar();
// Background: queue DEL for server (Ink's readline handles backspace via \x7f)
this._localEchoBgBuffer += '\x7f';
scheduleBgFlush();
return;
}
```
**Enter** (lines 2029–2050 → replace):
```js
if (/^[\r\n]+$/.test(data)) {
this._localEchoOverlay?.clear();
if (this._inputFlushTimeout) {
clearTimeout(this._inputFlushTimeout);
this._inputFlushTimeout = null;
}
// Flush any remaining background chars (e.g., last 50ms batch not yet sent)
const remaining = drainBgBuffer();
if (remaining) {
this._sendInputAsync(this.activeSessionId, remaining);
}
// Send \r after 120ms — Ink needs text and Enter as separate events.
// The promise chain in _sendInputAsync guarantees the remaining chars
// are dispatched before \r, regardless of timing.
setTimeout(() => {
this._pendingInput += '\r';
flushInput();
}, 120);
return;
}
```
**Key change from original plan**: The Enter handler no longer checks `if (text)` and branches on whether the overlay had content. With background sends, the PTY already has most/all of the text. We just flush any remainder and unconditionally send `\r` after 120ms. This simplifies the flow and handles edge cases like "user typed nothing but pressed Enter" (remainder is empty, just `\r` is sent).
**Control characters and paste** (lines 2052–2061 → replace):
```js
if (data.charCodeAt(0) < 32 || data.length > 1) {
this._localEchoOverlay?.clear();
// Flush background buffer so PTY has full text state before control char
// (critical for Tab completion — PTY needs typed text to complete against)
const remaining = drainBgBuffer();
if (remaining) {
this._sendInputAsync(this.activeSessionId, remaining);
}
// Send control char / paste text via normal path
this._pendingInput += data;
if (this._inputFlushTimeout) {
clearTimeout(this._inputFlushTimeout);
this._inputFlushTimeout = null;
}
flushInput();
return;
}
```
**Note on paste**: Desktop paste arrives via `onData` as a single multi-character string (`data.length > 1`). This falls into the control char path above, which:
1. Clears the overlay (existing behavior)
2. Flushes background buffer (new — ensures PTY has prefix text)
3. Sends paste text immediately (existing behavior)
Mobile paste via `KeyboardAccessoryBar.pasteFromClipboard()` uses `app.sendInput()` which bypasses `onData` entirely — no change needed.
### Step 3: Flush on tab switch (`selectSession()`, line ~4533)
Insert before the existing overlay save/clear block (before line 4534):
```js
// Flush background send buffer for outgoing session
if (this.activeSessionId) {
const remaining = drainBgBuffer();
if (remaining) {
this._sendInputAsync(this.activeSessionId, remaining);
}
}
```
This ensures the PTY receives all typed characters before the tab switch. When the user switches back, the terminal buffer will show the text (echoed by Ink) and the overlay will restore its cached copy on top.
### Step 4: Cleanup on local echo disable (`_updateLocalEchoState()`, lines 2362–2371)
Expand the disable transition (line 2367–2368):
```js
if (this._localEchoEnabled && !shouldEnable) {
this._localEchoOverlay?.clear();
// Flush any pending background chars before disabling
const remaining = drainBgBuffer();
if (remaining && this.activeSessionId) {
this._sendInputAsync(this.activeSessionId, remaining);
}
}
```
### Step 5: Cleanup on session delete (`deleteSession()`)
When a session is deleted, cancel any pending background timer for that session:
```js
// In deleteSession(), after removing the session from this.sessions:
drainBgBuffer(); // Discard — session is gone, nowhere to send
this.localEchoTextCache.delete(sessionId);
```
## Visual Timeline
```
t=0ms User types "h" → overlay: "h" bgBuffer: "h" timer: 50ms
t=30ms User types "e" → overlay: "he" bgBuffer: "he" timer: reset 50ms
t=60ms User types "l" → overlay: "hel" bgBuffer: "hel" timer: reset 50ms
t=110ms Debounce fires → overlay: "hel" bgBuffer: "" POST "hel" → PTY
t=115ms Ink echoes "hel" → terminal: "❯ hel" (hidden behind overlay)
t=140ms User types "l" → overlay: "hell" bgBuffer: "l" timer: 50ms
t=170ms User types "o" → overlay: "hello" bgBuffer: "lo" timer: reset 50ms
t=220ms Debounce fires → overlay: "hello" bgBuffer: "" POST "lo" → PTY
t=250ms User hits Enter → drainBgBuffer()="" overlay: cleared
t=370ms \r sent via chain → Ink processes "hello\r" → output appears
```
**Tab switch scenario:**
```
t=0ms User types "wor" → overlay: "wor" bgBuffer: "wor" timer: 50ms
t=25ms User switches tab → drainBgBuffer() sends "wor" to old session PTY
overlay text "wor" saved to localEchoTextCache
overlay cleared, new session loaded
...later...
t=5000ms User switches back → terminal shows "❯ wor" (Ink echo from background send)
overlay restores "wor" from cache, masks terminal
user continues typing seamlessly
```
## Edge Cases & Mitigations
### Confirmed Safe (JS single-threaded guarantee)
| Scenario | Why it's safe |
|----------|--------------|
| **Debounce fires during Enter handler** | Impossible. JS event loop is single-threaded — the Enter handler runs atomically. `drainBgBuffer()` cancels the timer before it can fire. |
| **Debounce fires during tab switch** | Same reason. `selectSession()` calls `drainBgBuffer()` synchronously, canceling the timer. |
| **Double-send of background buffer** | `drainBgBuffer()` atomically clears both buffer and timer. Once drained, subsequent drain returns empty string. |
| **`_pendingInput` conflict** | In local echo mode, `_pendingInput` is only used for Enter (`\r`) and control chars. Background chars use a separate `_localEchoBgBuffer`. No overlap. |
### Handled by Design
| Scenario | Handling |
|----------|---------|
| **Rapid typing / paste** | 50ms debounce batches rapid chars. At 100 WPM (~50ms/char), sends ~1 char per batch. For paste (multi-char string, `data.length > 1`), the control char path bypasses the buffer entirely and sends immediately. |
| **Network failure** | `_sendInputAsync` catches fetch failures and calls `_enqueueInput()` for retry. `_drainInputQueues()` replays on reconnect. Background chars use the same retry path. |
| **Offline mode** | `_sendInputAsync` checks `this.isOnline` and immediately enqueues if offline. Same behavior for background sends. 64KB queue cap prevents memory growth. |
| **Tab completion** | Ctrl+Tab path flushes background buffer BEFORE sending Tab char. PTY has full text for readline completion. |
| **Session respawn** | PTY already has typed text (sent in background). On respawn, Claude exits and restarts — Ink's readline buffer is lost, but the text was already processed or is no longer relevant. The overlay clears on session status change via `_updateLocalEchoState()`. |
| **SSE reconnect** | `handleInit()` saves overlay text before `selectSession()` clears it, then restores after reload (line ~3945–3966). Background buffer is cleared on reconnect since state is reset. |
### Network Ordering
**Question**: Can background sends arrive at the server out of order?
**Answer**: No, for practical purposes.
1. `_sendInputAsync` uses a **promise chain** (`_inputSendChain`) — each fetch is dispatched only after the previous one has been dispatched. This means requests are sent in order.
2. Localhost connections (HTTP/1.1) are inherently sequential on a single TCP connection.
3. Even with HTTP/2 multiplexing, Fastify (Node.js) is single-threaded — request handlers execute via the event loop in arrival order.
4. The server's `session.write()` is synchronous — it writes to the PTY immediately within the request handler.
### Known Limitations (Not Addressed)
| Limitation | Impact | Notes |
|-----------|--------|-------|
| **IME composition** | CJK input via IME would send partial composition sequences to PTY | No IME handling exists in the codebase today (line count: 0 references to `compositionstart/end/update`). Fixing this is a separate feature. |
| **`sendInput()` ordering** | Mobile accessory bar commands (`/init`, `/clear`, paste) use `sendInput()` which bypasses `_inputSendChain` — no ordering guarantee relative to background sends | Unlikely to conflict in practice: accessory bar clears the overlay first, and the commands are typically sent when no typing is in progress. |
| **localStorage stale text** | After background sends, localStorage still has overlay text. On hard reload, overlay restores text that the PTY already has → visual duplicate behind overlay | Harmless — overlay masks the terminal. On Enter, overlay clears and terminal shows correct state. Could be fixed by clearing localStorage after successful background flush, but adds complexity for minimal benefit. |
## Verification Checklist
1. **Basic typing**: Enable local echo → type "hello" → overlay shows instantly → check Network tab for batched POST requests (~50ms intervals) → press Enter → command executes
2. **Tab switch persistence**: Type "test" → switch to another tab → switch back → text visible in both overlay AND terminal prompt
3. **Backspace**: Type "helloo" → press backspace → overlay shows "hello" → check PTY received \x7f
4. **Paste**: Type "hel" → paste "lo world" → overlay clears → "lo world" sent immediately → PTY has "hello world"
5. **Tab completion**: Type "src/w" → press Tab → PTY completes to "src/web/" (background send gave PTY the prefix)
6. **Ctrl+C**: Type "hello" → press Ctrl+C → overlay clears → PTY receives pending chars + \x03
7. **Network tab**: Verify POST /api/sessions/:id/input requests appear as you type (batched ~50ms)
8. **Offline resilience**: Disconnect network → type "hello" → reconnect → verify chars are replayed via drain queue
9. **Session delete**: Type text → delete session → no console errors from orphaned timer
10. **Mobile keyboard**: Test on mobile device — typing goes through same onData path, same behavior expected
## Files Modified
| File | Changes |
|------|---------|
| `src/web/public/app.js` | ~40 lines changed across 5 locations (Steps 1–5) |
No server-side changes. No new files. No new dependencies.
@@ -1,321 +0,0 @@
# Plan: Background Keystroke Forwarding (Local Echo Mode)
## Problem
When local echo is enabled, keystrokes accumulate **only** in `LocalEchoOverlay.pendingText` (a client-side string). Nothing reaches the server PTY until Enter is pressed. This creates three failure modes:
1. **Tab switch loses PTY state** — switching sessions saves overlay text to `localEchoTextCache` (a Map), but the actual Claude Code Ink process has no knowledge of what was typed. If respawn or `/clear` fires on that session, the cached text is meaningless.
2. **Session death loses input** — if the session crashes or respawns while text is pending in the overlay, that input is gone (localStorage backup `codeman_local_echo_pending` only survives page reloads, not session resets).
3. **Tab completion impossible** — pressing Tab with pending overlay text sends the raw Tab character to a PTY that has no knowledge of the typed text, so completion fails.
## Goal
Send every keystroke to the server in the background (debounced), while the overlay continues providing instant visual feedback. The overlay sits at z-index 7 with an opaque background over `.xterm-screen`, masking Ink's echo of the background-sent characters. Input persists in the actual Claude Code readline buffer across tab switches and respawns.
## Architecture
```
User keystroke
|
v
xterm.js onData(data)
|
+---> LocalEchoOverlay.addChar(data) [instant visual feedback]
|
+---> _localEchoBgBuffer += data [queue for background send]
| clearTimeout + setTimeout(50ms)
| |
| v (50ms debounce fires)
| _flushBgInput()
| |
| v
| _sendInputAsync(sessionId, buffer) [promise chain preserves order]
| |
| v
| POST /api/sessions/:id/input [{ input: "hel" }]
| |
| v
| session.write(inputStr) [direct PTY write, synchronous]
| |
| v
| Ink readline echoes "hel" [hidden behind overlay's opaque bg]
|
+--- On Enter:
1. clearTimeout(_localEchoBgTimer)
2. flush _localEchoBgBuffer via _sendInputAsync (remaining chars)
3. clear overlay
4. 120ms later: send \r via _sendInputAsync (Ink text/Enter split)
5. Ink processes "hello\r" → overlay gone, terminal visible with output
```
### Two Input Paths (important context)
The codebase has **two separate input paths** to the server:
| Path | Used by | Promise chain? | `useMux`? |
|------|---------|---------------|-----------|
| `_sendInputAsync()` (line 3626) | `onData` handler, `flushInput()` | Yes (`_inputSendChain`) | No (direct PTY write) |
| `sendInput()` (line 8755) | Mobile accessory bar, programmatic commands | **No** (raw `fetch`) | Yes (tmux `send-keys`) |
Background keystroke forwarding uses **only** the `_sendInputAsync` path, which guarantees ordering via the promise chain. The `sendInput()` path is unaffected and unmodified.
### Server-Side Input Flow
```
POST /api/sessions/:id/input { input: "hel" }
|
+-- useMux? No (default)
| session.write("hel") → ptyProcess.write("hel") [sync]
|
+-- useMux? Yes
session.writeViaMux("hel") → tmux send-keys -l "hel" [async]
```
Background sends use the default path (no `useMux`), which is a synchronous direct PTY write — faster than spawning a tmux subprocess for each character batch.
## Implementation
All changes in **one file**: `src/web/public/app.js`
### Step 1: Add background send state (in terminal setup, after line ~1999)
```js
this._localEchoBgBuffer = ''; // Characters queued for background send
this._localEchoBgTimer = null; // 50ms debounce timer ID
```
Add an atomic drain helper alongside existing `flushInput` (after line ~2008):
```js
// Atomically drain background buffer — returns contents and cancels pending timer.
// Single point of extraction prevents double-flush race conditions.
const drainBgBuffer = () => {
if (this._localEchoBgTimer) {
clearTimeout(this._localEchoBgTimer);
this._localEchoBgTimer = null;
}
const buf = this._localEchoBgBuffer;
this._localEchoBgBuffer = '';
return buf;
};
const scheduleBgFlush = () => {
if (this._localEchoBgTimer) clearTimeout(this._localEchoBgTimer);
this._localEchoBgTimer = setTimeout(() => {
this._localEchoBgTimer = null;
const buf = this._localEchoBgBuffer;
this._localEchoBgBuffer = '';
if (buf && this.activeSessionId) {
this._sendInputAsync(this.activeSessionId, buf);
}
}, 50);
};
```
**Why `drainBgBuffer` exists**: Every exit path (Enter, Ctrl+C, tab switch, echo disable) needs to flush the buffer AND cancel the timer atomically. Without a single extraction point, it's easy to forget one of the two operations, leading to double-sends when the timer fires after a manual flush.
### Step 2: Modify `onData` handler — local echo path (lines 2023–2067)
**Printable characters** (lines 2063–2067 → replace):
```js
if (data.length === 1 && data.charCodeAt(0) >= 32) {
this._localEchoOverlay?.addChar(data);
// Background: queue char for server send (50ms debounce batches rapid typing)
this._localEchoBgBuffer += data;
scheduleBgFlush();
return;
}
```
**Backspace** (lines 2024–2028 → replace):
```js
if (data === '\x7f') {
this._localEchoOverlay?.removeChar();
// Background: queue DEL for server (Ink's readline handles backspace via \x7f)
this._localEchoBgBuffer += '\x7f';
scheduleBgFlush();
return;
}
```
**Enter** (lines 2029–2050 → replace):
```js
if (/^[\r\n]+$/.test(data)) {
this._localEchoOverlay?.clear();
if (this._inputFlushTimeout) {
clearTimeout(this._inputFlushTimeout);
this._inputFlushTimeout = null;
}
// Flush any remaining background chars (e.g., last 50ms batch not yet sent)
const remaining = drainBgBuffer();
if (remaining) {
this._sendInputAsync(this.activeSessionId, remaining);
}
// Send \r after 120ms — Ink needs text and Enter as separate events.
// The promise chain in _sendInputAsync guarantees the remaining chars
// are dispatched before \r, regardless of timing.
setTimeout(() => {
this._pendingInput += '\r';
flushInput();
}, 120);
return;
}
```
**Key change from original plan**: The Enter handler no longer checks `if (text)` and branches on whether the overlay had content. With background sends, the PTY already has most/all of the text. We just flush any remainder and unconditionally send `\r` after 120ms. This simplifies the flow and handles edge cases like "user typed nothing but pressed Enter" (remainder is empty, just `\r` is sent).
**Control characters and paste** (lines 2052–2061 → replace):
```js
if (data.charCodeAt(0) < 32 || data.length > 1) {
this._localEchoOverlay?.clear();
// Flush background buffer so PTY has full text state before control char
// (critical for Tab completion — PTY needs typed text to complete against)
const remaining = drainBgBuffer();
if (remaining) {
this._sendInputAsync(this.activeSessionId, remaining);
}
// Send control char / paste text via normal path
this._pendingInput += data;
if (this._inputFlushTimeout) {
clearTimeout(this._inputFlushTimeout);
this._inputFlushTimeout = null;
}
flushInput();
return;
}
```
**Note on paste**: Desktop paste arrives via `onData` as a single multi-character string (`data.length > 1`). This falls into the control char path above, which:
1. Clears the overlay (existing behavior)
2. Flushes background buffer (new — ensures PTY has prefix text)
3. Sends paste text immediately (existing behavior)
Mobile paste via `KeyboardAccessoryBar.pasteFromClipboard()` uses `app.sendInput()` which bypasses `onData` entirely — no change needed.
### Step 3: Flush on tab switch (`selectSession()`, line ~4533)
Insert before the existing overlay save/clear block (before line 4534):
```js
// Flush background send buffer for outgoing session
if (this.activeSessionId) {
const remaining = drainBgBuffer();
if (remaining) {
this._sendInputAsync(this.activeSessionId, remaining);
}
}
```
This ensures the PTY receives all typed characters before the tab switch. When the user switches back, the terminal buffer will show the text (echoed by Ink) and the overlay will restore its cached copy on top.
### Step 4: Cleanup on local echo disable (`_updateLocalEchoState()`, lines 2362–2371)
Expand the disable transition (line 2367–2368):
```js
if (this._localEchoEnabled && !shouldEnable) {
this._localEchoOverlay?.clear();
// Flush any pending background chars before disabling
const remaining = drainBgBuffer();
if (remaining && this.activeSessionId) {
this._sendInputAsync(this.activeSessionId, remaining);
}
}
```
### Step 5: Cleanup on session delete (`deleteSession()`)
When a session is deleted, cancel any pending background timer for that session:
```js
// In deleteSession(), after removing the session from this.sessions:
drainBgBuffer(); // Discard — session is gone, nowhere to send
this.localEchoTextCache.delete(sessionId);
```
## Visual Timeline
```
t=0ms User types "h" → overlay: "h" bgBuffer: "h" timer: 50ms
t=30ms User types "e" → overlay: "he" bgBuffer: "he" timer: reset 50ms
t=60ms User types "l" → overlay: "hel" bgBuffer: "hel" timer: reset 50ms
t=110ms Debounce fires → overlay: "hel" bgBuffer: "" POST "hel" → PTY
t=115ms Ink echoes "hel" → terminal: "❯ hel" (hidden behind overlay)
t=140ms User types "l" → overlay: "hell" bgBuffer: "l" timer: 50ms
t=170ms User types "o" → overlay: "hello" bgBuffer: "lo" timer: reset 50ms
t=220ms Debounce fires → overlay: "hello" bgBuffer: "" POST "lo" → PTY
t=250ms User hits Enter → drainBgBuffer()="" overlay: cleared
t=370ms \r sent via chain → Ink processes "hello\r" → output appears
```
**Tab switch scenario:**
```
t=0ms User types "wor" → overlay: "wor" bgBuffer: "wor" timer: 50ms
t=25ms User switches tab → drainBgBuffer() sends "wor" to old session PTY
overlay text "wor" saved to localEchoTextCache
overlay cleared, new session loaded
...later...
t=5000ms User switches back → terminal shows "❯ wor" (Ink echo from background send)
overlay restores "wor" from cache, masks terminal
user continues typing seamlessly
```
## Edge Cases & Mitigations
### Confirmed Safe (JS single-threaded guarantee)
| Scenario | Why it's safe |
|----------|--------------|
| **Debounce fires during Enter handler** | Impossible. JS event loop is single-threaded — the Enter handler runs atomically. `drainBgBuffer()` cancels the timer before it can fire. |
| **Debounce fires during tab switch** | Same reason. `selectSession()` calls `drainBgBuffer()` synchronously, canceling the timer. |
| **Double-send of background buffer** | `drainBgBuffer()` atomically clears both buffer and timer. Once drained, subsequent drain returns empty string. |
| **`_pendingInput` conflict** | In local echo mode, `_pendingInput` is only used for Enter (`\r`) and control chars. Background chars use a separate `_localEchoBgBuffer`. No overlap. |
### Handled by Design
| Scenario | Handling |
|----------|---------|
| **Rapid typing / paste** | 50ms debounce batches rapid chars. At 100 WPM (~50ms/char), sends ~1 char per batch. For paste (multi-char string, `data.length > 1`), the control char path bypasses the buffer entirely and sends immediately. |
| **Network failure** | `_sendInputAsync` catches fetch failures and calls `_enqueueInput()` for retry. `_drainInputQueues()` replays on reconnect. Background chars use the same retry path. |
| **Offline mode** | `_sendInputAsync` checks `this.isOnline` and immediately enqueues if offline. Same behavior for background sends. 64KB queue cap prevents memory growth. |
| **Tab completion** | Ctrl+Tab path flushes background buffer BEFORE sending Tab char. PTY has full text for readline completion. |
| **Session respawn** | PTY already has typed text (sent in background). On respawn, Claude exits and restarts — Ink's readline buffer is lost, but the text was already processed or is no longer relevant. The overlay clears on session status change via `_updateLocalEchoState()`. |
| **SSE reconnect** | `handleInit()` saves overlay text before `selectSession()` clears it, then restores after reload (line ~3945–3966). Background buffer is cleared on reconnect since state is reset. |
### Network Ordering
**Question**: Can background sends arrive at the server out of order?
**Answer**: No, for practical purposes.
1. `_sendInputAsync` uses a **promise chain** (`_inputSendChain`) — each fetch is dispatched only after the previous one has been dispatched. This means requests are sent in order.
2. Localhost connections (HTTP/1.1) are inherently sequential on a single TCP connection.
3. Even with HTTP/2 multiplexing, Fastify (Node.js) is single-threaded — request handlers execute via the event loop in arrival order.
4. The server's `session.write()` is synchronous — it writes to the PTY immediately within the request handler.
### Known Limitations (Not Addressed)
| Limitation | Impact | Notes |
|-----------|--------|-------|
| **IME composition** | CJK input via IME would send partial composition sequences to PTY | No IME handling exists in the codebase today (line count: 0 references to `compositionstart/end/update`). Fixing this is a separate feature. |
| **`sendInput()` ordering** | Mobile accessory bar commands (`/init`, `/clear`, paste) use `sendInput()` which bypasses `_inputSendChain` — no ordering guarantee relative to background sends | Unlikely to conflict in practice: accessory bar clears the overlay first, and the commands are typically sent when no typing is in progress. |
| **localStorage stale text** | After background sends, localStorage still has overlay text. On hard reload, overlay restores text that the PTY already has → visual duplicate behind overlay | Harmless — overlay masks the terminal. On Enter, overlay clears and terminal shows correct state. Could be fixed by clearing localStorage after successful background flush, but adds complexity for minimal benefit. |
## Verification Checklist
1. **Basic typing**: Enable local echo → type "hello" → overlay shows instantly → check Network tab for batched POST requests (~50ms intervals) → press Enter → command executes
2. **Tab switch persistence**: Type "test" → switch to another tab → switch back → text visible in both overlay AND terminal prompt
3. **Backspace**: Type "helloo" → press backspace → overlay shows "hello" → check PTY received \x7f
4. **Paste**: Type "hel" → paste "lo world" → overlay clears → "lo world" sent immediately → PTY has "hello world"
5. **Tab completion**: Type "src/w" → press Tab → PTY completes to "src/web/" (background send gave PTY the prefix)
6. **Ctrl+C**: Type "hello" → press Ctrl+C → overlay clears → PTY receives pending chars + \x03
7. **Network tab**: Verify POST /api/sessions/:id/input requests appear as you type (batched ~50ms)
8. **Offline resilience**: Disconnect network → type "hello" → reconnect → verify chars are replayed via drain queue
9. **Session delete**: Type text → delete session → no console errors from orphaned timer
10. **Mobile keyboard**: Test on mobile device — typing goes through same onData path, same behavior expected
## Files Modified
| File | Changes |
|------|---------|
| `src/web/public/app.js` | ~40 lines changed across 5 locations (Steps 1–5) |
No server-side changes. No new files. No new dependencies.
-111
View File
@@ -1,111 +0,0 @@
> **⚠️ ARCHIVED 2026-05-21 — superseded, kept for history.**
> The headline items here were verified resolved: the P0 `{WORKING_DIR}` placeholder
> is now replaced (`plan-orchestrator.ts:431`), and the "~66 dead functions in app.js"
> are gone (app.js was modularized 15K→3K LOC). A fresh `npm run knip` sweep on
> 2026-05-21 found only a handful of unused test helpers. Do not treat this as a live TODO.
# Codebase Cleanup Findings
Compiled from parallel analysis of the entire Codeman codebase by 3 research agents (2026-02-19).
## P0 — Bug Fix
### 1. `{WORKING_DIR}` placeholder never replaced in plan-orchestrator.ts
- **File:** `src/plan-orchestrator.ts:409`
- `RESEARCH_AGENT_PROMPT` has `{WORKING_DIR}` placeholder but only `{TASK}` is replaced
- The literal string `{WORKING_DIR}` gets sent to the AI model
- **Fix:** Add `.replace('{WORKING_DIR}', this.workingDir)` after the `{TASK}` replacement
## P1 — Dead Code Removal (High Impact)
### 2. ~66 dead functions in app.js
- Functions never called: `clearAll()`, `toggleSubagentDropdown()`, `goHome()`, `showRalphWizard()`, `minimizeRalphWizard()`, `restoreRalphWizard()`, `ralphWizardNext()`, `ralphWizardBack()`, `skipPlanGeneration()`, `regeneratePlan()`, `incrementTabCount()`, `decrementTabCount()`, `incrementShellCount()`, `decrementShellCount()`, `stopClaude()`, and ~50 more
- Many are remnants of abandoned features (Ralph wizard, plan version history)
- **Estimated savings:** 300-500 lines
### 3. ~74 dead CSS selectors in styles.css
- Major dead blocks: Task Panel System (`.task-panel`), Process Panel System (`.process-panel`), Monitor Tabs (`.monitor-tabs`), Ralph Metadata (`.ralph-progress-section`, `.ralph-meta`), Plan Editor Toolbar, Plan Version History
- Plus ~30 minor unused utility/component selectors
- **Estimated savings:** ~400 lines
### 4. 13 dead type definitions in types.ts (~150 lines)
- Dead request interfaces (superseded by Zod schemas): `CreateSessionRequest`, `RunPromptRequest`, `SessionInputRequest`, `ResizeRequest`, `CreateCaseRequest`, `QuickStartRequest`, `CreateScheduledRunRequest`, `QuickRunRequest`, `HookEventRequest`
- Other dead types: `TaskAssignment`, `MemoryMetrics`, `RalphStateRecord`
- Dead function: `createSuccessResponse` (exported, never imported)
- **Estimated savings:** ~150 lines
### 5. 9 unused constants in map-limits.ts
- `MAX_PENDING_HOOKS`, `MAX_SESSION_HISTORY`, `MAX_SSE_CLIENTS_PER_SESSION`, `MAX_TOTAL_SSE_CLIENTS`, `FILE_WATCHER_WARNING_THRESHOLD`, `MAX_QUEUED_TASKS`, `MAX_COMPLETED_TASKS_HISTORY`, `COMPLETED_TODO_TTL_MS`, `MAX_CONCURRENT_SESSIONS`
- 9 of 14 exports are dead — only 5 are actually imported
### 6. Dead `SessionInputSchema` in schemas.ts
- `SessionInputSchema` (line 87) is defined/exported but never imported
- `SessionInputWithLimitSchema` is the one actually used
### 7. Dead `code-reviewer.ts` prompt file
- `src/prompts/code-reviewer.ts` — entire file is dead, `CODE_REVIEWER_PROMPT` never imported
- Re-exported in `src/prompts/index.ts` but no consumer
### 8. Dead utility exports
- **Default exports** (4 files): `lru-map.ts`, `cleanup-manager.ts`, `stale-expiration-map.ts`, `buffer-accumulator.ts` — all have `export default` that's never used
- **`stripAnsiSimple`** in `regex-patterns.ts` — exported, never imported (only `stripAnsi` used)
- **String similarity**: `isSimilar`, `isSimilarByDistance`, `stringSimilarity`, `levenshteinDistance` — none imported externally
- **LRUMap methods**: `oldest()`, `newest()`, `peek()`, `expireOlderThan()`, `valuesInOrder()`, `maxEntries`, `freeSlots` — never called
- **StaleExpirationMap methods**: `touch()`, `getAge()`, `getRemainingTtl()`, `peek()` — never called
- **CleanupManager methods**: `registerWatcher()`, `registerListener()`, `registerStream()`, `getRegistrations()`, `resourceCounts` — never called
### 9. Dead backend functions
- `resetSessionManager()` in session-manager.ts:300 — never imported
- `getStoredTasks()` in task-queue.ts:264 — never called
- `start()` in session.ts:1918 — no-op legacy method
- Empty `updateStatsFromEvent()` in run-summary.ts:397 — called every event, does nothing
### 10. Dead TS type exports
- `AiCheckerEvents<R>`, `AiIdleCheckerEvents`, `AiPlanCheckerEvents` — never imported
- `AiCheckStatus`, `AiPlanCheckStatus` — backwards compat aliases, never imported
## P2 — Performance & Efficiency
### 11. task-queue.ts `getCount()` iterates all tasks 5 times
- Called every Ralph Loop tick — creates array from Map, then filters 4 times
- **Fix:** Single-pass counting like `TaskTracker.getStats()` does
### 12. transcript-watcher.ts double file read
- `readNewEntries()` reads the file twice: once for CRLF detection, once for parsing
- `crlfDelay: Infinity` already handles both line endings
- **Fix:** Remove the raw buffer CRLF check, read once
### 13. tmux-manager.ts `saveSessions()` no debounce
- Rapid calls can overlap; no in-flight guard unlike `StateStore`
- **Fix:** Add debouncing or in-flight tracking
## P3 — Consolidation & Consistency
### 14. Duplicate `SAFE_PATH_PATTERN` regex
- `schemas.ts:15` and `tmux-manager.ts:81` — identical regex
- **Fix:** Share from one location
### 15. Duplicate `MAX_CONCURRENT_SESSIONS`
- `map-limits.ts:57` (dead) vs `server.ts:131` (used, hardcoded)
- **Fix:** server.ts should import from map-limits
### 16. Duplicate cache TTLs in server.ts
- `SESSIONS_LIST_CACHE_TTL` and `LIGHT_STATE_CACHE_TTL_MS` — both 1000ms
- **Fix:** Consolidate into one constant
### 17. Inconsistent path import in server.ts
- Imports both `path` default and destructured `{ join, dirname, resolve, relative, isAbsolute }`
- 3 lines use `path.join()` while everywhere else uses `join()`
- **Fix:** Remove default import, use `join()` consistently
### 18. Re-export indirection for `getAugmentedPath`
- `session.ts:89` re-exports from `claude-cli-resolver.ts` for backwards compat
- `ai-checker-base.ts` should import directly from source
### 19. `cliInfoUpdated` event missing from SessionEvents interface
- Emitted in `session.ts:1742`, handled in `server.ts:4214`, but not in the interface
- Type safety gap — handlers aren't type-checked
### 20. Array instead of Set for `_childAgentIds` in session.ts
- Uses `includes()`/`indexOf()` for lookups (O(n))
- Small lists in practice, but Set is more appropriate
-983
View File
@@ -1,983 +0,0 @@
> **⚠️ ARCHIVED 2026-05-21 — superseded, kept for history.**
> The "Critical" structural items here are done: `server.ts` 6,736→2,065 LOC,
> `app.js` 15,196→3,083 LOC, `types.ts` 1,443→12 LOC (now a barrel → `src/types/`).
> The phase plans that executed this work are in `docs/archive/phase*-plan.md`.
> Do not treat this as a live TODO; see CLAUDE.md for current architecture.
# Code Structure & Quality Findings
**Date**: 2026-02-28
**Scope**: Full codebase analysis across 5 dimensions: frontend, backend, TypeScript, testing, and utilities/config.
This document contains detailed findings for agent teams to write implementation plans and execute improvements. Each section includes severity, specific locations, and recommended fixes.
---
## Table of Contents
1. [Critical: server.ts God Object (6,736 LOC)](#1-critical-serverts-god-object)
2. [Critical: app.js Monolith (15,196 LOC)](#2-critical-appjs-monolith)
3. [Critical: CleanupManager Unused Despite Existing](#3-critical-cleanupmanager-unused)
4. [High: Duplicated Debounce/Timer Patterns](#4-high-duplicated-debouncetimer-patterns)
5. [High: Large Domain Files Need Splitting](#5-high-large-domain-files-need-splitting)
6. [High: types.ts God File (1,443 LOC)](#6-high-typests-god-file)
7. [High: Zod Schemas Duplicate TypeScript Types](#7-high-zod-schemas-duplicate-typescript-types)
8. [High: Test Coverage Gaps](#8-high-test-coverage-gaps)
9. [High: Duplicated Test Mocks](#9-high-duplicated-test-mocks)
10. [Medium: Hardcoded Magic Values](#10-medium-hardcoded-magic-values)
11. [Medium: Frontend Global State Monolith](#11-medium-frontend-global-state-monolith)
12. [Medium: Frontend Code Duplication](#12-medium-frontend-code-duplication)
13. [Medium: Inconsistent Logging](#13-medium-inconsistent-logging)
14. [Medium: Utils Barrel Export Gaps](#14-medium-utils-barrel-export-gaps)
15. [Medium: Non-Null Assertion Risks](#15-medium-non-null-assertion-risks)
16. [Low: Dead Utility Functions](#16-low-dead-utility-functions)
17. [Low: No Dependency Injection for File I/O](#17-low-no-dependency-injection-for-file-io)
18. [Scorecard & Prioritized Roadmap](#18-scorecard--prioritized-roadmap)
---
## 1. Critical: server.ts God Object
**File**: `src/web/server.ts` (6,736 lines)
**Severity**: CRITICAL
**Impact**: Hardest file to maintain, test, and extend. Imports 38 modules.
### Problem
The `WebServer` class handles everything: HTTP routing (~110 routes), authentication, SSE broadcasting, terminal data batching, state persistence, session lifecycle, respawn orchestration, file serving, tunnel management, plan orchestration, and subagent coordination.
**Key metrics**:
- 40+ private properties (Maps, timers, caches)
- 70+ methods
- `setupRoutes()` is 2,000+ LOC of inline route handlers
- Zero test coverage
### Current Structure (Bad)
```
WebServer class (6,736 LOC)
├── Auth session management (lines 469, 668-698)
├── SSE client management (lines 407-408, 5843-5880)
├── Terminal data batching (lines 414-416, 5909-5966)
├── Task update batching (line 426, 5995-6028)
├── State persistence batching (lines 429-430, 6028-6061)
├── Respawn lifecycle (lines 445-451, 5425-5534)
├── Session cleanup (lines 4769-4961)
├── Listener setup (lines 544-643)
└── setupRoutes() (lines 645+, 2000+ LOC)
├── /api/sessions/* (30+ routes inline)
├── /api/respawn/* (7 routes inline)
├── /api/subagents/* (7 routes inline)
├── /api/plan/* (5 routes inline)
├── /api/push/* (4 routes inline)
└── ... 60+ more inline
```
### Recommended Structure
```
src/web/
├── server.ts (~500 LOC - HTTP setup, route registration only)
├── routes/
│ ├── session-routes.ts (session CRUD, input, resize)
│ ├── respawn-routes.ts (respawn control endpoints)
│ ├── subagent-routes.ts (background agent tracking)
│ ├── plan-routes.ts (plan generation & management)
│ ├── push-routes.ts (web push subscriptions)
│ ├── mux-routes.ts (tmux management)
│ ├── case-routes.ts (case management)
│ ├── file-routes.ts (file browsing/serving)
│ └── system-routes.ts (status, stats, config, settings)
├── middleware/
│ ├── auth.ts (Basic Auth + session cookies)
│ └── error-handler.ts (centralized error responses)
└── services/
├── sse-manager.ts (SSE client + broadcast)
├── terminal-batcher.ts (60fps terminal batching)
└── session-lifecycle.ts (listener setup/teardown)
```
### Duplication in server.ts
**Error response pattern** repeated 189 times:
```typescript
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Session not found');
```
**Fix**: Extract `findSessionOrFail()` middleware:
```typescript
const findSessionOrFail = (sessionId: string) => {
const session = this.sessions.get(sessionId);
if (!session) throw new NotFoundError('Session not found');
return session;
};
```
**Event listener setup** copy-pasted for subagent watcher, image watcher, and team watcher (lines 544-643). Same attach/detach pattern duplicated 3 times.
---
## 2. Critical: app.js Monolith
**File**: `src/web/public/app.js` (15,196 lines)
**Severity**: CRITICAL
**Impact**: Untestable, hard to navigate, tightly coupled systems.
### Extractable Modules (by priority)
| Module | Lines | Current Location | Impact |
|--------|-------|------------------|--------|
| Mobile handlers (MobileDetection, KeyboardHandler, SwipeHandler) | ~300 | lines 168-620 | High |
| Voice input (DeepgramProvider, VoiceInput) | ~830 | lines 631-1471 | High |
| NotificationManager | ~450 | lines 2218-2663 | High |
| xterm-zerolag-input (inlined copy from packages/) | ~400 | lines 1756-2153 | High |
| KeyboardAccessoryBar | ~195 | lines 1480-1680 | Medium |
| FocusTrap | ~60 | lines 1690-1748 | Medium |
### CodemanApp Class (12,000+ LOC)
The main `CodemanApp` class starting at line 2665 has:
- **60+ Maps/Sets** in the constructor (lines 2667-2805)
- **18 Map instances** with complex cross-references (subagents, parents, teams, windows)
- **10+ monolithic methods** exceeding 100 lines each
**Largest methods**:
| Method | Lines | Size |
|--------|-------|------|
| `renderAppSettings()` | 14400-14700 | ~300 LOC |
| `selectSession()` | 6028-6250 | ~220 LOC |
| `batchTerminalWrite()` | 7482-7700 | ~200 LOC |
| `renderSessionTabs()` | 5814-6000 | ~180 LOC |
| `openSubagentWindow()` | 11927-12100 | ~170 LOC |
| `handleInit()` | 5183-5350 | ~170 LOC |
### Recommended Split
```
src/web/public/
├── app.js (~4000 LOC - core app, session mgmt, SSE)
├── mobile.js (~300 LOC - MobileDetection, KeyboardHandler, SwipeHandler)
├── voice.js (~830 LOC - DeepgramProvider, VoiceInput)
├── notifications.js (~450 LOC - NotificationManager)
├── keyboard-accessory.js (~200 LOC - KeyboardAccessoryBar)
├── api-client.js (~100 LOC - fetch wrapper with error handling)
└── config.js (~50 LOC - magic numbers, z-index layers)
```
---
## 3. Critical: CleanupManager Unused
**File**: `src/utils/cleanup-manager.ts` (320 lines)
**Severity**: CRITICAL
**Impact**: Memory leak risk. Well-designed utility exists but is never used. Every file manages cleanup manually.
### Current State
`CleanupManager` is exported from the utils barrel but has **0 instantiations** in production code. Instead, every file implements manual cleanup:
**respawn-controller.ts** (worst offender):
```typescript
// 11 timer properties, manually cleared in stop()
private stepTimer: NodeJS.Timeout | null = null;
private completionConfirmTimer: NodeJS.Timeout | null = null;
private noOutputTimer: NodeJS.Timeout | null = null;
// ... 8 more
stop() {
if (this.stepTimer) clearTimeout(this.stepTimer);
if (this.completionConfirmTimer) clearTimeout(this.completionConfirmTimer);
// ... 9 more clearTimeout/clearInterval calls
}
```
**Files that should use CleanupManager**:
| File | Timer/Listener Count | Current Cleanup |
|------|---------------------|-----------------|
| `respawn-controller.ts` | 11 timers + intervals | 11 manual clearTimeout/clearInterval |
| `web/server.ts` | 6+ timers, debounce map | Manual in stop(), some may leak |
| `state-store.ts` | 2 debounce timers | Manual clearTimeout |
| `push-store.ts` | 1 save timer | Manual clearTimeout |
| `subagent-watcher.ts` | debounce map + watchers | Manual clear + close |
| `ralph-tracker.ts` | 3 debounce timers | Manual clear |
| `bash-tool-parser.ts` | 1 debounce timer | Manual clear |
| `image-watcher.ts` | 1 debounce map | Manual clear |
### Fix
Migrate all timer management to use `CleanupManager`. Example for respawn-controller.ts:
```typescript
// Before: 11 fields + 11 clearTimeout calls
private stepTimer: NodeJS.Timeout | null = null;
// ...
// After: 1 field, auto-cleanup
private cleanup = new CleanupManager();
startStep() {
this.cleanup.setTimeout(() => { ... }, 5000, 'step');
}
stop() {
this.cleanup.dispose(); // Clears everything
}
```
---
## 4. High: Duplicated Debounce/Timer Patterns
**Severity**: HIGH
**Impact**: 8+ files implement debounce independently. Bug fixes need to be applied everywhere.
### Pattern Inventory
```typescript
// Pattern 1: Manual timer ref (used in 6 files)
private saveTimer: NodeJS.Timeout | null = null;
debouncedSave() {
if (this.saveTimer) clearTimeout(this.saveTimer);
this.saveTimer = setTimeout(() => this.save(), 500);
}
// Pattern 2: Timer Map (used in 3 files)
private fileDebouncers = new Map<string, NodeJS.Timeout>();
debounce(key: string) {
const existing = this.fileDebouncers.get(key);
if (existing) clearTimeout(existing);
this.fileDebouncers.set(key, setTimeout(() => { ... }, 100));
}
// Pattern 3: State flag (used in 2 files)
private isSaving = false;
```
### Locations
| File | Debounce Vars | Delay (ms) |
|------|---------------|------------|
| `state-store.ts` | `saveTimeout`, `ralphStateSaveTimeout` | 500 |
| `push-store.ts` | `saveTimer` | 500 |
| `web/server.ts` | `persistDebounceTimers` (Map) | 500 |
| `subagent-watcher.ts` | `fileDebouncers` (Map) | 100 |
| `ralph-tracker.ts` | 3 debounce timers | 50, 30000 |
| `bash-tool-parser.ts` | `EVENT_DEBOUNCE_MS` | 50 |
| `image-watcher.ts` | debounce map | 200 |
| `respawn-controller.ts` | 11 timer fields | various |
### Fix
Create a `Debouncer` utility:
```typescript
// src/utils/debouncer.ts
export class Debouncer {
private timer: NodeJS.Timeout | null = null;
constructor(private readonly delayMs: number) {}
run(fn: () => void): void {
if (this.timer) clearTimeout(this.timer);
this.timer = setTimeout(fn, this.delayMs);
}
cancel(): void {
if (this.timer) clearTimeout(this.timer);
this.timer = null;
}
}
// Usage:
private saveDeb = new Debouncer(500);
this.saveDeb.run(() => this.save());
// cleanup: this.saveDeb.cancel();
```
---
## 5. High: Large Domain Files Need Splitting
**Severity**: HIGH
**Impact**: Complex state machines spanning 3,000+ lines are hard to understand and test.
### ralph-tracker.ts (3,905 LOC)
**5 responsibilities mixed**:
1. Output Parsing (~900 LOC) - Line-by-line parsing, state extraction
2. Todo Management (~700 LOC) - Parsing, dedup, expiry
3. Plan Tracking (~800 LOC) - Enhanced plan tasks, checkpoints
4. Circuit Breaker (~400 LOC) - State machine for stuck detection
5. File Watching (~300 LOC) - Monitor external state files
**Recommended split**:
```
ralph-tracker.ts (core output parsing, ~1200 LOC)
ralph-todo-manager.ts (todo parsing + management, ~700 LOC)
ralph-plan-tracker.ts (plan tasks + checkpoints, ~800 LOC)
ralph-circuit-breaker.ts (circuit breaker logic, ~400 LOC)
```
### respawn-controller.ts (3,611 LOC)
**6 responsibilities mixed**:
1. State Machine (~1,000 LOC) - 6+ states, transitions
2. Idle Detection (~800 LOC) - 5 layers + multi-signal combining
3. AI Checkers (~600 LOC) - Idle + plan checkers integration
4. Health Scoring (~500 LOC) - Metrics, circuit breaker, scoring
5. Action Logging (~300 LOC) - Timeline, detection status
6. Stuck-State Detection (~250 LOC) - Timeout tracking
**Recommended split**:
```
respawn-controller.ts (state machine core, ~1000 LOC)
respawn-idle-detection.ts (all 5 idle detection layers, ~800 LOC)
respawn-health-scorer.ts (metrics & health scoring, ~500 LOC)
```
### session.ts (2,418 LOC)
**8 responsibilities mixed**:
1. PTY Management (~600 LOC)
2. Terminal I/O (~400 LOC)
3. Token Tracking (~200 LOC)
4. Task Tracking (~250 LOC)
5. Ralph Integration (~200 LOC)
6. Auto-Clear/Compact (~300 LOC)
7. Image Watching (~100 LOC)
8. CLI Detection (~150 LOC)
**Recommended split**:
```
session.ts (PTY + terminal I/O core, ~1000 LOC)
session-tracking.ts (token + task + Ralph, ~500 LOC)
session-auto-ops.ts (auto-clear/compact + image, ~300 LOC)
```
---
## 6. High: types.ts God File
**File**: `src/types.ts` (1,443 lines, 72 exported definitions)
**Severity**: HIGH
**Impact**: Every file imports from types.ts. Hard to find relevant types.
### Current Contents
- 46 interfaces
- 25 types
- 1 enum (ApiErrorCode)
- 9 factory functions (createInitialState, etc.)
### Recommended Split
```
src/types/
├── index.ts (barrel export - transparent migration)
├── session.ts (SessionState, SessionConfig, SessionMode, SessionColor)
├── task.ts (TaskState, TaskDefinition, TaskStatus)
├── respawn.ts (RespawnConfig, RespawnState, CircuitBreakerStatus)
├── ralph.ts (RalphLoopState, RalphTrackerState, RalphTodoItem)
├── api.ts (ApiResponse, ApiErrorCode, HookEventType, all route types)
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
└── common.ts (Disposable, BufferConfig, CleanupResourceType)
```
The barrel export makes this a transparent refactor - existing `import from './types'` continues to work.
---
## 7. High: Zod Schemas Duplicate TypeScript Types
**File**: `src/web/schemas.ts` (508 lines)
**Severity**: HIGH
**Impact**: When a type changes, the Zod schema must be manually updated too. Source of bugs.
### Problem
Zod schemas manually duplicate TypeScript interfaces. **Zero `z.infer` usage found.**
```typescript
// types.ts (manual interface)
export interface CreateSessionRequest {
workingDir?: string;
mode?: SessionMode;
name?: string;
}
// schemas.ts (manual Zod schema - duplicated!)
export const CreateSessionSchema = z.object({
workingDir: safePathSchema.optional(),
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
name: z.string().max(100).optional(),
});
```
### Fix
Use `z.infer` to derive TypeScript types from Zod schemas (single source of truth):
```typescript
// schemas.ts
export const CreateSessionSchema = z.object({
workingDir: safePathSchema.optional(),
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
name: z.string().max(100).optional(),
});
// types.ts (auto-derived)
export type CreateSessionRequest = z.infer<typeof CreateSessionSchema>;
```
**Affected schemas** (~10):
- CreateSessionSchema
- RunPromptSchema
- ResizeSchema
- CreateCaseSchema
- QuickStartSchema
- HookEventSchema
- RespawnConfigSchema
- ConfigUpdateSchema
- SettingsUpdateSchema
---
## 8. High: Test Coverage Gaps
**Severity**: HIGH
**Impact**: Critical code paths untested. Regressions go unnoticed.
### Untested Source Files
| File | Lines | Risk |
|------|-------|------|
| `src/web/server.ts` | 6,736 | CRITICAL - Core REST API, 280+ routes |
| `src/plan-orchestrator.ts` | ~500 | HIGH - Multi-agent plan generation |
| `src/tunnel-manager.ts` | ~200 | MEDIUM - Cloudflare tunnel |
| `src/session-lifecycle-log.ts` | ~150 | MEDIUM - JSONL audit log |
| `src/ai-plan-checker.ts` | ~300 | MEDIUM - Plan completion detection |
| `src/templates/claude-md.ts` | ~200 | LOW - CLAUDE.md generation |
| `src/utils/claude-cli-resolver.ts` | ~100 | LOW - CLI path resolution |
| `src/utils/opencode-cli-resolver.ts` | ~100 | LOW - OpenCode CLI support |
| `src/utils/regex-patterns.ts` | ~100 | LOW - Used everywhere! |
| `src/utils/token-validation.ts` | ~50 | LOW - Token counting |
### Test Quality Issues
**10 "not.toThrow()" tests without behavior verification**:
```typescript
// BAD: Only checks it doesn't crash
expect(() => tracker.processMessage(null)).not.toThrow();
// GOOD: Also verify defensive behavior
expect(() => tracker.processMessage(null)).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
```
Locations:
- `task-tracker.test.ts` - 5 instances
- `image-watcher.test.ts` - 1 instance
- `task-queue.test.ts` - 1 instance
- Others scattered
---
## 9. High: Duplicated Test Mocks
**Severity**: HIGH
**Impact**: Mock changes need updating in 4 places. Inconsistent mock behavior.
### MockSession Defined 4 Times
| File | Usage |
|------|-------|
| `test/respawn-controller.test.ts` | Full mock with event emitter |
| `test/session-manager.test.ts` | Simpler mock |
| `test/respawn-team-awareness.test.ts` | Copy of respawn-controller mock |
| `test/respawn-test-utils.ts` | **Comprehensive mock - UNUSED!** |
### MockStateStore Defined 2 Times
| File | Usage |
|------|-------|
| `test/session-manager.test.ts` | Basic mock |
| `test/ralph-loop.test.ts` | Separate implementation |
### Unused Test Utilities
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
- `createTimeController()` - Abstraction over vitest fake timers
- `MockAiIdleChecker` - Fully mocked AI idle checker
- `MockAiPlanChecker` - Fully mocked plan checker
- Factory functions for pre-configured controllers
### Fix
Create `test/mocks/` directory:
```
test/
├── mocks/
│ ├── mock-session.ts (single MockSession, used everywhere)
│ ├── mock-state-store.ts (single MockStateStore)
│ └── index.ts (barrel export)
├── utils/
│ └── time-controller.ts (from respawn-test-utils.ts)
└── ... test files
```
---
## 10. Medium: Hardcoded Magic Values
**Severity**: MEDIUM
**Impact**: Hard to tune, inconsistent when same value appears in multiple places.
### Already Centralized (Good)
- `src/config/buffer-limits.ts` - All buffer sizes
- `src/config/map-limits.ts` - All collection limits
### NOT Centralized (40+ values scattered)
**In server.ts** (lines 145-194):
```typescript
const TASK_UPDATE_BATCH_INTERVAL = 100;
const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
const SESSIONS_LIST_CACHE_TTL = 1000;
const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
const MAX_TERMINAL_COLS = 500;
const MAX_TERMINAL_ROWS = 200;
const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
const MAX_AUTH_SESSIONS = 100;
const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
const STATS_COLLECTION_INTERVAL_MS = 2000;
const MAX_INPUT_LENGTH = 64 * 1024;
```
**In hooks-config.ts**: `timeout: 10000` hardcoded 6 times.
**In respawn-controller.ts** (lines 538-565): 10 timing constants.
**In utils**: `EXEC_TIMEOUT_MS = 5000` duplicated in both `claude-cli-resolver.ts` and `opencode-cli-resolver.ts`.
**In app.js**:
```javascript
// line 27: 600000 - stuck detection threshold
// line 24: 5000 - default scrollback
// lines 34-35: 128*1024, 256*1024 - chunk sizes
// lines 152-155: 150, 100 - keyboard detection thresholds
// lines 573-575: 80, 300, 100 - swipe detection params
```
### Fix
Create additional config files:
```
src/config/
├── buffer-limits.ts (existing)
├── map-limits.ts (existing)
├── server-config.ts (NEW - web server intervals, auth, caching)
├── timing-config.ts (NEW - debounce delays, check intervals)
└── terminal-config.ts (NEW - max cols/rows, batch intervals)
```
---
## 11. Medium: Frontend Global State Monolith
**Severity**: MEDIUM
**Impact**: All state in single CodemanApp class. Tight coupling between unrelated systems.
### 60+ State Variables in CodemanApp Constructor (lines 2667-2805)
```javascript
this.sessions = new Map(); // Session data
this.subagents = new Map(); // Agent tracking
this.subagentActivity = new Map(); // Tool call tracking
this.subagentToolResults = new Map(); // Result caching
this.subagentParentMap = new Map(); // Agent-to-session mapping
this.teams = new Map(); // Team tracking
this.teamTasks = new Map(); // Team task state
this.planSubagents = new Map(); // Plan agent tracking
this.pendingWrites = []; // Terminal write queue
this.terminalBufferCache = new Map(); // Buffer caching (unbounded!)
this.projectInsights = new Map(); // Bash tool insights
// ... 40+ more
```
### Problems
1. **18 Map instances** with complex cross-references (no garbage collection strategy)
2. **No domain separation**: Session, subagent, notification, UI, and network state mixed
3. **Implicit dependencies**: `selectSession()` requires 5+ Maps to be in consistent state
4. **`terminalBufferCache`** has no max size - can grow unbounded with many sessions
### Recommended Domain Split
```javascript
// Instead of 60+ flat properties:
class SessionState {
sessions = new Map();
sessionOrder = [];
terminalBuffers = new Map();
tabAlerts = new Map();
}
class SubagentState {
subagents = new Map();
activity = new Map();
parentMap = new Map();
windows = new Map();
minimized = new Map();
}
class TeamState {
teams = new Map();
tasks = new Map();
teammates = new Map();
}
class UIState {
activeSessionId = null;
draggedTabId = null;
isLoadingBuffer = false;
}
```
---
## 12. Medium: Frontend Code Duplication
**Severity**: MEDIUM
**Impact**: Repeated patterns increase maintenance burden and inconsistency risk.
### Duplicated Patterns
**API fetch calls** (~50 instances):
```javascript
// Repeated everywhere:
fetch(`/api/sessions/${sessionId}/...`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({...})
}).catch(() => {})
```
**Fix**: Extract `ApiClient` class.
**`innerHTML` usage** (104 instances):
- Mix of template strings, createElement chains, and direct innerHTML
- Some with manual XSS escaping (`text.replace(/</g, '&lt;')`), some without
- No consistent DOM creation pattern
**`typeof app !== 'undefined'` checks** (20+ instances):
- Lines 458, 467, 481, 614, 617, 1549, etc.
- **Fix**: Ensure `app` is always defined as global singleton.
**Element visibility toggling** (212+ occurrences):
```javascript
element.classList.add('active')
element.classList.remove('active')
```
**Fix**: Create `toggleClass(el, className, condition)` utility.
### Event Listener Issues
- **152 `addEventListener` calls** with fragile cleanup
- **Mix of inline (`onclick="app.method()"`) and addEventListener** - hard to track
- **Element cache (`_elemCache`) never invalidated** if DOM elements are recreated (line 2808)
- **Tab drag-and-drop listeners** may not clean up if user switches tabs mid-drag
---
## 13. Medium: Inconsistent Logging
**Severity**: MEDIUM
**Impact**: Hard to debug in production. Can't filter by severity or component.
### Current State
- **345 console calls** across source files
- **No structured logging** - all `console.log/error` directly
- **No log levels** (DEBUG, INFO, WARN, ERROR)
### Inconsistent Prefixes
```typescript
// Some files use brackets:
console.log('[Session] Starting interactive...');
console.log('[RalphLoop] Task assigned...');
console.log('[TunnelManager] Tunnel started');
// Others use no prefix:
console.error('Failed to spawn PTY:', err);
console.log('Server listening on port', port);
```
### Positive: CleanupManager Has Debug Mode
`src/utils/cleanup-manager.ts` has a `debugMode` flag for conditional debug logging - good pattern not replicated elsewhere.
### Fix
Either:
1. Enforce consistent `[ComponentName]` prefixes via lint rule
2. Create lightweight logger abstraction (not a heavy framework)
---
## 14. Medium: Utils Barrel Export Gaps
**File**: `src/utils/index.ts`
**Severity**: MEDIUM
**Impact**: Forces deep imports, unclear public API.
### Missing Exports
These functions are defined but NOT exported from the barrel:
- `createAnsiPatternFull()` and `createAnsiPatternSimple()` (factory functions from `regex-patterns.ts`)
- `SAFE_PATH_PATTERN` (from `regex-patterns.ts`)
- `validateTokenCounts()` and `validateTokensAndCost()` (from `token-validation.ts`)
- `isSimilar()`, `isSimilarByDistance()`, `levenshteinDistance()`, `normalizePhrase()` (from `string-similarity.ts` - though some are dead code, see finding #16)
### Deep Import Anti-Pattern (16 instances)
Some files bypass the barrel unnecessarily:
```typescript
// Could use barrel:
import { BufferAccumulator } from './utils/buffer-accumulator.js';
import { LRUMap } from './utils/lru-map.js';
// Must deep import (not in barrel):
import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';
```
### Fix
Add missing exports to `src/utils/index.ts` and update import sites.
---
## 15. Medium: Non-Null Assertion Risks
**Severity**: MEDIUM
**Impact**: Runtime crashes if assumptions violated. 37 instances found.
### Distribution
| File | Count | Risk Level |
|------|-------|------------|
| `src/web/server.ts` | 10 | Low (auth flow verified) |
| `src/session.ts` | 6 | **High** (mux/terminal refs) |
| `src/respawn-controller.ts` | 4 | Low (config validated) |
| `src/lru-map.ts` | 3 | Low (checked lookups) |
| `src/subagent-watcher.ts` | 2 | Low (pending tool calls) |
| Others | 12 | Low |
### High-Risk Examples (session.ts)
```typescript
// Line 915 - _mux could be null if startInteractive called during cleanup
`[Session] Starting interactive (with ${this._mux!.backend})`
// Line 954 - _muxSession could be null in race condition
this._muxSession!.muxName
```
### Fix
Add null guards before assertions, or document invariants:
```typescript
// Before:
this._mux!.backend
// After:
if (!this._mux) throw new Error('Invariant: _mux must be initialized before startInteractive');
this._mux.backend
```
### Positive Notes
- **0 instances of `as any`**
- **0 instances of `@ts-ignore` or `@ts-expect-error`**
- TypeScript overall score: 8.5/10
---
## 16. Low: Dead Utility Functions
**File**: `src/utils/string-similarity.ts`
**Severity**: LOW
**Impact**: Code clutter, confusion about what's actually used.
### Unused Functions
These are defined and exported but **never imported anywhere**:
- `isSimilar(a, b, threshold)` - similarity check with threshold
- `isSimilarByDistance(a, b, maxDistance)` - Levenshtein-based check
- `levenshteinDistance(a, b)` - raw edit distance
- `normalizePhrase(phrase)` - phrase normalization
### Actually Used
Only these are imported from the barrel:
- `stringSimilarity()` - used in ralph-tracker.ts
- `fuzzyPhraseMatch()` - used in ralph-tracker.ts
- `todoContentHash()` - used in ralph-tracker.ts
### Fix
Delete unused functions or mark as `@internal` if kept for future use.
---
## 17. Low: No Dependency Injection for File I/O
**Severity**: LOW (practical impact limited at current scale)
**Impact**: Can't mock filesystem for unit tests. 68+ hard-coded filesystem calls.
### Examples
```typescript
// state-store.ts - directly imports and uses fs
import { readFileSync, writeFileSync, existsSync, mkdirSync } from 'node:fs';
// push-store.ts - hard-coded paths
const KEYS_FILE = join(DATA_DIR, 'push-keys.json');
const SUBS_FILE = join(DATA_DIR, 'push-subscriptions.json');
// ai-checker-base.ts - direct execSync
execSync(`tmux kill-session -t "${this.checkMuxName}"`, { timeout: 3000 });
```
### Why This Is Lower Priority
- The codebase uses integration tests (spawning real processes/tmux sessions) rather than unit tests
- Most filesystem operations are in infrastructure code, not business logic
- Adding DI would be a large refactor with limited near-term benefit
---
## 18. Scorecard & Prioritized Roadmap
### Overall Scores (Post-Implementation)
| Category | Before | After | Notes |
|----------|--------|-------|-------|
| TypeScript Safety | 8.5/10 | 9/10 | 0 `any`, 0 `@ts-ignore`, Zod `z.infer` eliminates type drift |
| Error Handling | 8/10 | 8/10 | Unchanged — already strong |
| Async/Promise Safety | 9.5/10 | 9.5/10 | Unchanged — already strong |
| Resource Cleanup | 7/10 | 8/10 | CleanupManager adopted in server.ts, subagent-watcher, bash-tool-parser; Debouncer in 6 files. **Gaps**: respawn-controller (10+ manual timers) and ralph-tracker (2 manual timers) not migrated |
| Module Organization | 5/10 | 8/10 | Routes extracted (12 modules), types split (14 domain files), domain files split (ralph: 7, respawn: 5, session: 6) |
| Test Coverage | 6/10 | 7.5/10 | Shared mock infrastructure, 12 route test files, MockSession/MockStateStore consolidated |
| Config Centralization | 6/10 | 9/10 | 9 config files, ~65 constants centralized, 0 cross-file duplicates |
| Frontend Architecture | 4/10 | 7/10 | 8 extracted modules (3,453 LOC), app.js reduced 24% (15.2K → 11.5K), xterm-zerolag-input vendor build |
| Code Duplication | 5/10 | 8/10 | Debouncer utility, shared test mocks, barrel exports, config consolidation |
### Implementation Phases
**Phase 1 - Quick Wins (1-2 days)** ✅ COMPLETE
1. ✅ Export missing functions from utils barrel (~30 min) — `createAnsiPatternFull`, `createAnsiPatternSimple`, `SAFE_PATH_PATTERN`, `validateTokenCounts`, `validateTokensAndCost` all now exported from `src/utils/index.ts`
2. ✅ Delete dead utility functions (~15 min) — `isSimilar()` removed from `string-similarity.ts`; `levenshteinDistance()`, `isSimilarByDistance()`, `normalizePhrase()` made private (used internally by `fuzzyPhraseMatch`/`stringSimilarity`)
3. ✅ Consolidate duplicated `EXEC_TIMEOUT_MS` constant (~15 min) — Created `src/config/exec-timeout.ts` as single source of truth; `claude-cli-resolver.ts`, `opencode-cli-resolver.ts`, and `tmux-manager.ts` all import from it
4. ✅ Add `z.infer` to Zod schemas (~2 hours) — `src/web/schemas.ts` now has 36 `z.infer` type exports (lines 512-547) covering all schemas
5. ✅ Fix 10 weak "not.toThrow()" tests (~1 hour) — All `not.toThrow()` calls now have behavior assertions: `task-tracker.test.ts` (6 instances all followed by state checks), `image-watcher.test.ts` (1 instance followed by length check), `session-manager.test.ts` (1 instance followed by count check)
**Phase 2 - CleanupManager & Debounce (2-3 days)** ✅ COMPLETE
1. ✅ Create `Debouncer` utility class (~1 hour) — Created `src/utils/debouncer.ts` with `Debouncer` and `KeyedDebouncer` classes; exported from `src/utils/index.ts`
2. ✅ Migrate all 8 files from manual debounce to Debouncer — `state-store.ts` (2 Debouncers), `push-store.ts` (1 Debouncer), `bash-tool-parser.ts` (1 Debouncer), `image-watcher.ts` (1 KeyedDebouncer), `subagent-watcher.ts` (2 KeyedDebouncers), `server.ts` (1 KeyedDebouncer for persist timers), `ralph-tracker.ts` (2 Debouncers replacing 4 manual fields: `_todoUpdateTimer`, `_loopUpdateTimer`, `_todoUpdatePending`, `_loopUpdatePending`)
3. ✅ Migrate respawn-controller to CleanupManager — 10 manual timer fields replaced with single `CleanupManager` instance + `timerIds` Map. `startTrackedTimer()`/`cancelTrackedTimer()` preserved as wrappers for UI countdown display and timer events. `clearTimers()` uses dispose-and-recreate pattern for state transitions.
4. ✅ Migrate server.ts timer cleanup to CleanupManager (~2 hours) — `private cleanup = new CleanupManager()` present; terminal batch timers and pending respawn starts left as manual Maps (complex lifecycle)
5. ✅ Migrate remaining files — `bash-tool-parser.ts` (CleanupManager ✅), `subagent-watcher.ts` (CleanupManager ✅), `ralph-tracker.ts` (Debouncer ✅)
**Phase 3 - server.ts Route Extraction (3-4 days)** ✅ COMPLETE
1. ✅ Created `src/web/routes/` with 12 domain route modules + index barrel (4,090 LOC total): session (909), system (768), ralph (533), plan (459), respawn (315), case, file, hook-event, mux, push, scheduled, team
2. ✅ Created `src/web/middleware/auth.ts` (193 LOC) — Basic Auth, session cookies, rate limiting, security headers, CORS
3. ✅ Created `src/web/ports/` with 7 typed port interfaces (142 LOC) — SessionPort, EventPort, RespawnPort, ConfigPort, InfraPort, AuthPort; routes declare dependencies via intersection types
4. ✅ Created `src/web/route-helpers.ts` (154 LOC) — `findSessionOrFail()`, `formatUptime()`, `sanitizeHookData()`, `autoConfigureRalph()`
5. ✅ Reduced `server.ts` from 6,736 → 2,697 LOC (60% reduction). Remaining LOC is justified infrastructure: session lifecycle, SSE broadcast engine, terminal batching, respawn integration, resource cleanup
**Phase 4 - Domain File Splitting (2-3 days)** ✅ COMPLETE
1. ✅ Split `types.ts` into `src/types/` directory — 14 domain files (1,469 LOC total): common, session, task, app-state, respawn, ralph, api, lifecycle, run-summary, tools, teams, push, plan + index barrel. Original `types.ts` is now a 1-line re-export
2. ✅ Split `ralph-tracker.ts` into 7 files (exceeded plan of 4) — ralph-tracker (2,391), ralph-plan-tracker (477), ralph-status-parser (552), ralph-fix-plan-watcher (366), ralph-stall-detector (166), ralph-config (153), ralph-loop (522)
3. ✅ Split `respawn-controller.ts` into 5 files (exceeded plan of 3) — respawn-controller (3,228), respawn-health (229), respawn-metrics (229), respawn-patterns (131), respawn-adaptive-timing (134)
4. ✅ Split `session.ts` into 6 files (exceeded plan of 3) — session (2,168), session-manager (298), session-auto-ops (284), session-cli-builder (132), session-task-cache (101), session-lifecycle-log (114)
**Phase 5 - Frontend Modularization (3-4 days)** ✅ COMPLETE
1. ✅ Extracted `constants.js` (238 LOC) — shared constants, timing values, Z-index layers, `escapeHtml()`, `extractSyncSegments()`
2. ✅ Extracted `mobile-handlers.js` (449 LOC) — `MobileDetection`, `KeyboardHandler`, `SwipeHandler`
3. ✅ Extracted `voice-input.js` (853 LOC) — `DeepgramProvider`, `VoiceInput`
4. ✅ Extracted `notification-manager.js` (445 LOC) — `NotificationManager` class (5-layer system)
5. ✅ Extracted `keyboard-accessory.js` (279 LOC) — `KeyboardAccessoryBar`, `FocusTrap`
6. ✅ Extracted `api-client.js` (70 LOC) — `_api()`, `_apiJson()`, `_apiPost()`, `_apiPut()`
7. ✅ Extracted `subagent-windows.js` (1,119 LOC) — 13 subagent window methods
8. ✅ Removed inlined xterm-zerolag-input copy → built to `vendor/xterm-zerolag-input.js` from `packages/xterm-zerolag-input/`
9. ✅ Reduced `app.js` from ~15,200 → 11,473 LOC (24% reduction). All scripts loaded in correct dependency order in `index.html`
**Phase 6 - Config Consolidation (1 day)** ✅ COMPLETE
1. ✅ Created 6 new domain-focused config files (better than plan's 2 generic files): `server-timing.ts` (13 constants), `auth-config.ts` (5 constants), `tunnel-config.ts` (8 constants), `terminal-limits.ts` (4 constants), `ai-defaults.ts` (3 constants), `team-config.ts` (3 constants)
2. ✅ Total: 9 config files in `src/config/`, ~65 constants centralized
3. ✅ Eliminated all cross-file duplicates: `STATS_COLLECTION_INTERVAL_MS` (was in 2 files), `timeout: 10000` (was 6× inline in hooks-config.ts → `HOOK_TIMEOUT_MS`), AI model string (was in 5 files → `AI_CHECK_MODEL`), `MAX_TRACKED_AGENTS` (was shadowed in subagent-watcher.ts)
4. ✅ CLAUDE.md updated with config files table, import conventions, resource limits references
**Phase 7 - Test Infrastructure (2-3 days)** ✅ COMPLETE
1. ✅ Created `test/mocks/` directory with 5 files (541 LOC): `mock-session.ts` (312), `mock-state-store.ts` (60), `mock-route-context.ts` (121), `test-helpers.ts` (37), `index.ts` (11 — barrel export)
2. ✅ Consolidated MockSession into single shared definition — no duplicate class definitions remain (2 `vi.mock()`-based copies intentionally left in session-manager.test.ts and ralph-loop.test.ts)
3. ✅ `respawn-test-utils.ts` converted to backward-compatibility shim — re-exports from `test/mocks/`, retains respawn-specific utilities (MockAiIdleChecker, TimeController, etc.)
4. ✅ Created initial 3 route test files with 58 total tests: `session-routes.test.ts` (34 tests), `respawn-routes.test.ts` (13 tests), `system-routes.test.ts` (11 tests). Route test harness uses `app.inject()` — no real ports needed
5. ✅ All 12 route modules now have dedicated test files in `test/routes/`: session, respawn, system, ralph, plan, push, team, mux, file, scheduled, hook-event, case
---
## Appendix: File Size Inventory (Post-Implementation)
### Before vs After
| File | Before | After | Change |
|------|--------|-------|--------|
| `src/web/server.ts` | 6,736 | 2,697 | **−60%** (routes, auth, ports extracted) |
| `src/web/public/app.js` | 15,196 | 11,473 | **−24%** (8 modules extracted) |
| `src/ralph-tracker.ts` | 3,905 | 2,391 | **−39%** (6 companion files extracted) |
| `src/respawn-controller.ts` | 3,611 | 3,228 | **−11%** (4 companion files extracted) |
| `src/session.ts` | 2,418 | 2,168 | **−10%** (5 companion files extracted) |
| `src/types.ts` | 1,443 | 1 | **−99%** (14 domain files in `src/types/`) |
### New Infrastructure Created
| Directory | Files | Total LOC | Purpose |
|-----------|-------|-----------|---------|
| `src/web/routes/` | 13 | 4,090 | Domain route modules |
| `src/web/ports/` | 7 | 142 | Port interfaces for DI |
| `src/web/middleware/` | 1 | 193 | Auth middleware |
| `src/types/` | 14 | 1,469 | Domain type files |
| `src/config/` | 9 | ~450 | Centralized config |
| `test/mocks/` | 5 | 541 | Shared test mocks |
| `test/routes/` | 4 | ~500 | Route handler tests |
### Extracted Frontend Modules
| Module | Lines | Purpose |
|--------|-------|---------|
| `subagent-windows.js` | 1,119 | Subagent window management |
| `voice-input.js` | 853 | DeepgramProvider, VoiceInput |
| `mobile-handlers.js` | 449 | MobileDetection, KeyboardHandler, SwipeHandler |
| `notification-manager.js` | 445 | 5-layer notification system |
| `keyboard-accessory.js` | 279 | KeyboardAccessoryBar, FocusTrap |
| `constants.js` | 238 | Shared constants, timing, Z-index |
| `api-client.js` | 70 | API fetch wrapper |
### What's Working Well
These patterns should be **preserved, not refactored**:
- Clean one-way dependency graph (no circular deps)
- EventEmitter-based decoupling between domain models
- Proper `import type` usage (19 files, consistent)
- Utility type adoption (101 instances of Record, Partial, Omit, etc.)
- `assertNever()` for exhaustive switch checking
- `StaleExpirationMap` and `LRUMap` for bounded collections
- State persistence circuit breaker pattern
- TypeScript strict mode with all safety flags enabled
- `CleanupManager` for centralized timer/watcher disposal
- `Debouncer`/`KeyedDebouncer` for consistent debounce patterns
- Port interfaces for route module dependency injection
- `Object.assign(CodemanApp.prototype, ...)` for frontend module composition
@@ -1,409 +0,0 @@
# First-Load Performance Optimization Plan
**Date**: 2026-02-18
**Audit by**: 4-agent team (css-analyst, js-analyst, server-analyst, deps-analyst)
**Scope**: First browser load of Codeman web UI at `/`
---
## Current State (Baseline)
### Payload
| Asset | Raw | Compressed | Render-Blocking? |
|-------|-----|-----------|-----------------|
| `index.html` | 82 KB | ~15 KB | N/A (document) |
| `styles.css` | 154 KB | ~25 KB | **YES** |
| `mobile.css` | 34 KB | ~7 KB | **YES** (missing media query) |
| `xterm.css` (CDN) | 2 KB | ~2 KB | No (preload pattern) |
| `xterm.min.js` (CDN) | 67 KB | ~65 KB | No (defer) |
| `xterm-addon-fit` (CDN) | 1 KB | ~1 KB | No (defer) |
| `app.js` | 563 KB | ~126 KB | No (defer) |
| **Total** | **903 KB** | **~241 KB** | |
### Request Waterfall (13 requests on first load)
```
T=0 GET / (82KB doc)
T+20ms ├── styles.css?v=0.1536 (154KB — BLOCKS RENDER)
├── mobile.css?v=0.1536 (34KB — BLOCKS RENDER on all viewports!)
├── xterm.css (CDN, preloaded) (2KB — non-blocking, already async)
├── xterm.min.js (CDN, defer) (67KB)
├── xterm-addon-fit.min.js (CDN) (1KB)
└── app.js?v=0.1536 (defer) (563KB)
[FIRST PAINT blocked by: styles.css + mobile.css]
T+200ms JS execution starts
├── new Terminal() + terminal.open() ← HEAVY sync (canvas creation)
├── connectSSE() → /api/events ← SSE stream
├── loadState() → /api/status ← DUPLICATE of SSE init!
├── loadQuickStartCases()
│ ├── /api/settings ← fetched TWICE
│ └── /api/cases?_t=<timestamp> ← cache-busted unnecessarily
├── startSystemStatsPolling()
│ └── /api/system/stats ← starts immediately, every 2s
└── loadAppSettingsFromServer()
└── /api/settings ← DUPLICATE #2
T+500ms First Meaningful Paint (terminal + header visible)
```
### Problems
1. **2 render-blocking CSS files** — mobile.css blocks desktop for no reason
2. **Double handleInit()** — SSE init + /api/status both call full state reset
3. **Duplicate /api/settings** — fetched twice in init chain
4. **563KB unminified JS monolith** — no build minification at all
5. **154KB unminified CSS** — 70% is for modals/wizards (below-the-fold)
6. **Sync terminal.open()** — heaviest single call, blocks before first paint
7. **12 modals pre-rendered** — ~600+ hidden DOM nodes, ~60KB HTML
8. **Stats polling starts immediately** — even with 0 sessions
9. **No loading skeleton** — blank black screen until all CSS+JS loads
10. **CDN dependency** — 3 xterm files from jsdelivr (DNS+TLS latency)
11. **1h cache for versioned assets** — could be 1yr+immutable with ?v= busting
12. **No HTTP/2** — 6-connection limit queues some requests
13. **On-the-fly compression** — no pre-compressed .gz/.br files
### What's Already Good (don't touch)
- Single shared Terminal instance (buffer swapping)
- Teammate terminals created lazily on window open
- Subagent windows use HTML logs, not Terminal instances
- `getLightState()` has 1s TTL cache
- SSE init sends lightweight state (no terminal buffers)
- Buffer hydration uses chunked writes (128KB via rAF)
- `selectSession()` defers secondary panels via requestIdleCallback
- System fonts only — zero web font loading
- xterm.css already uses async preload pattern
- Proper SSE reconnection with exponential backoff
- CSS `contain` on header/tabs for layout isolation
---
## Implementation Plan (15 steps, ordered by impact/effort)
### Phase 1: Quick Wins (1-line to 15-min changes)
#### Step 1: Add media attribute to mobile.css
**Impact**: HIGH — 34KB stops blocking render on desktop
**File**: `src/web/public/index.html:14`
```html
<!-- BEFORE -->
<link rel="stylesheet" href="mobile.css?v=0.1536">
<!-- AFTER -->
<link rel="stylesheet" href="mobile.css?v=0.1536" media="(max-width: 1023px)">
```
Browser still downloads it (for potential resize) but won't block rendering on desktop. The mobile.css file header says this was intended but never implemented.
---
#### Step 2: Remove duplicate /api/status + double handleInit()
**Impact**: HIGH — eliminates redundant API call + double state reset (clears 15+ Maps, 7+ timers, runs cleanupAllFloatingWindows(), double renderSessionTabs())
**Files**: `src/web/public/app.js`
The SSE `init` event (server.ts:618) sends `getLightState()`. The `loadState()` in `init()` at `app.js:1554` fetches identical data from `/api/status`. Both call `handleInit()` which wipes state. The `_initGeneration` guard only protects session-restore, NOT the expensive cleanup (lines 3389-3503).
```js
// In init() — REMOVE this.loadState(), add SSE fallback:
this.connectSSE();
// Remove: this.loadState();
this._initFallbackTimer = setTimeout(() => {
if (this._initGeneration === 0) this.loadState();
}, 3000);
// In handleInit() — clear fallback timer:
handleInit(data) {
if (this._initFallbackTimer) {
clearTimeout(this._initFallbackTimer);
this._initFallbackTimer = null;
}
// ... rest of handleInit
}
```
---
#### Step 3: Deduplicate /api/settings fetch
**Impact**: MEDIUM — removes 1 redundant API call
**Files**: `src/web/public/app.js:7341` (loadQuickStartCases), `app.js:9964` (loadAppSettingsFromServer)
```js
// In init() — fetch settings once, share the promise:
const settingsPromise = fetch('/api/settings').then(r => r.json());
this.loadQuickStartCases(null, settingsPromise);
this.loadAppSettingsFromServer(settingsPromise);
```
Both functions need to accept an optional pre-fetched settings promise parameter.
---
#### Step 4: Remove cache-busting from /api/cases
**Impact**: LOW — allows browser caching
**File**: `src/web/public/app.js:7351`
```js
// BEFORE
const res = await fetch('/api/cases?_t=' + Date.now());
// AFTER
const res = await fetch('/api/cases');
```
---
#### Step 5: Defer system stats polling
**Impact**: MEDIUM — removes API call every 2s when idle
**Files**: `src/web/public/app.js:1567`, `app.js:15261`
Move `startSystemStatsPolling()` out of `init()`. Start it in `handleInit()` only when `data.sessions.length > 0`.
---
### Phase 2: Build Pipeline (30-min changes, highest payload impact)
#### Step 6: Self-host xterm.js assets
**Impact**: MEDIUM-HIGH — eliminates CDN DNS/TLS latency (~100ms even with preconnect)
**Files**: `src/web/public/index.html`, `package.json` build script
xterm is NOT in package.json — add it:
```bash
npm install xterm@5.3.0 @xterm/addon-fit@0.8.0 --save
```
Build script addition:
```bash
mkdir -p dist/web/public/vendor
cp node_modules/xterm/css/xterm.css dist/web/public/vendor/
cp node_modules/xterm/lib/xterm.min.js dist/web/public/vendor/
cp node_modules/@xterm/addon-fit/lib/xterm-addon-fit.min.js dist/web/public/vendor/
```
Update index.html CDN URLs to `/vendor/xterm.min.js` etc. Remove preconnect/dns-prefetch for jsdelivr.
---
#### Step 7: Add esbuild minification to build
**Impact**: HIGH — biggest single optimization for payload size
**File**: `package.json` build script
Current build just does `cp -r src/web/public dist/web/`. No minification.
app.js stats: 1,525 comment lines (10%), 89 console.* statements, 23% whitespace.
```bash
# Add to build script after cp:
npx esbuild dist/web/public/app.js --minify --drop:console --outfile=dist/web/public/app.js --allow-overwrite
npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite
npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite
```
Expected savings:
| File | Before (gzip) | After (gzip) | Saved |
|------|---------------|-------------|-------|
| app.js | ~126 KB | ~85 KB | ~41 KB (33%) |
| styles.css | ~25 KB | ~18 KB | ~7 KB (28%) |
| mobile.css | ~7 KB | ~5 KB | ~2 KB (29%) |
| **Total** | **~158 KB** | **~108 KB** | **~50 KB** |
---
#### Step 8: Pre-compress static assets at build time
**Impact**: MEDIUM — eliminates per-request CPU compression
**Files**: `package.json` build script, potentially `src/web/server.ts`
```bash
# Add to build script after minification:
for f in dist/web/public/*.{js,css,html}; do
gzip -9 -k "$f"
brotli -9 -k "$f"
done
```
Check if `@fastify/static` supports `preCompressed: true` option. If not, serve pre-compressed files via custom Accept-Encoding check.
---
#### Step 9: Extend cache duration for versioned assets
**Impact**: LOW (first load) / HIGH (repeat visits)
**File**: `src/web/server.ts:601`
```js
// BEFORE
maxAge: '1h'
// AFTER
maxAge: '1y',
immutable: true
```
Safe because all assets use `?v=0.1536` cache-busting. First-load unaffected, but all repeat visits serve from disk cache instantly.
---
### Phase 3: Perceived Performance (30-60min, user experience)
#### Step 10: Add loading skeleton
**Impact**: MEDIUM-HIGH — instant visual structure instead of black screen
**File**: `src/web/public/index.html`
Add minimal inline `<style>` + skeleton HTML in `<body>`:
```html
<style>
.skeleton { display: flex; flex-direction: column; height: 100vh; background: #0a0a0a; }
.skeleton-header { height: 40px; background: #111; border-bottom: 1px solid #222; }
.skeleton-terminal { flex: 1; background: #0d0d0d; }
.app-loaded .skeleton { display: none; }
</style>
<div class="skeleton">
<div class="skeleton-header"></div>
<div class="skeleton-terminal"></div>
</div>
```
In `app.js` init() end: `document.body.classList.add('app-loaded');`
---
#### Step 11: Defer terminal creation to after first paint
**Impact**: MEDIUM-HIGH — terminal.open() is heaviest sync call
**File**: `src/web/public/app.js:1545`
```js
init() {
// ... mobile detection, visibility settings ...
document.documentElement.classList.remove('mobile-init');
// Show skeleton immediately, defer heavy terminal init
requestAnimationFrame(() => {
this.initTerminal();
this.connectSSE();
// ... rest of init
});
}
```
Lets browser paint header/tabs/skeleton before canvas creation.
---
### Phase 4: DOM + Payload Reduction (1-3 hours)
#### Step 12: Lazy-create modals on first open
**Impact**: HIGH — removes ~600+ DOM nodes, ~60KB hidden HTML
**Files**: `src/web/public/index.html`, `src/web/public/app.js`
12 modals pre-rendered in index.html:
- `helpModal` (lines 227-447)
- `sessionOptionsModal` (lines 448-714) — 266 lines
- `appSettingsModal` (lines 715-900+)
- `createCaseModal`, `mobileCasePickerModal`, `ralphWizardModal`, `killAllModal`, `closeConfirmModal`, `savePresetModal`, `tokenStatsModal`, `filePreviewModal`, notification drawer
Replace each modal's HTML with `<div id="helpModal" class="modal"></div>`. On first open, inject full HTML via template function. Cache after creation.
---
#### Step 13: Batch initial API calls into one endpoint
**Impact**: MEDIUM — reduces 4+ API calls to 1
**Files**: `src/web/server.ts`, `src/web/public/app.js`
Create `GET /api/init-bundle`:
```json
{
"status": { /* getLightState() */ },
"cases": [ /* case list */ ],
"settings": { /* user settings */ }
}
```
Use as SSE init fallback (step 2's timeout). Saves HTTP round trips.
---
#### Step 14: Trim SSE init payload
**Impact**: LOW-MEDIUM — reduces init payload by removing data not needed for first paint
**File**: `src/web/server.ts`
Remove from SSE init event: `taskTree`, `ralphTodos`, `ralphTodoStats` per session. These can be fetched on-demand when user opens a session's details panel.
---
#### Step 15: Enable HTTP/2
**Impact**: MEDIUM — multiplexed loading over single connection
**File**: `src/web/server.ts`
```js
// BEFORE
const server = Fastify({ logger: false });
// AFTER (when HTTPS is enabled)
const server = Fastify({
logger: false,
http2: true // Only works with HTTPS
});
```
Only applicable for `--https` mode. HTTP/2 multiplexing eliminates the 6-connection limit queuing.
---
## Expected Combined Impact
| Metric | Before | After | Improvement |
|--------|--------|-------|-------------|
| First Paint | ~300ms | ~100ms | **-200ms** (skeleton visible instantly) |
| First Contentful Paint | ~400ms | ~200ms | **-200ms** (no mobile.css blocking desktop) |
| Time to Interactive | ~600ms | ~350ms | **-250ms** (fewer API calls, deferred terminal) |
| Total compressed payload | ~241 KB | ~191 KB | **-50 KB (21%)** via minification |
| Init API calls | 6-7 (2 dupes) | 2-3 | **-60%** fewer requests |
| Initial DOM nodes | ~1800+ | ~1200 | **-600** (lazy modals) |
---
## Verification
After each step, verify with Playwright:
```js
const { chromium } = require('playwright');
const browser = await chromium.launch();
const page = await browser.newPage();
// Measure first paint
await page.goto('http://localhost:3000', { waitUntil: 'domcontentloaded' });
await page.waitForTimeout(4000); // Wait for async data
// Check UI renders correctly
const header = await page.locator('.header').isVisible();
const tabs = await page.locator('.session-tabs').isVisible();
const terminal = await page.locator('.terminal-container').isVisible();
console.log({ header, tabs, terminal });
await browser.close();
```
---
## Files Changed Per Step (for implementation agent)
| Step | Files Modified |
|------|---------------|
| 1 | `index.html` |
| 2 | `app.js` |
| 3 | `app.js` |
| 4 | `app.js` |
| 5 | `app.js` |
| 6 | `index.html`, `package.json` |
| 7 | `package.json` |
| 8 | `package.json`, optionally `server.ts` |
| 9 | `server.ts` |
| 10 | `index.html`, `app.js` |
| 11 | `app.js` |
| 12 | `index.html`, `app.js` |
| 13 | `server.ts`, `app.js` |
| 14 | `server.ts` |
| 15 | `server.ts` |
-388
View File
@@ -1,388 +0,0 @@
# Performance Audit: First Page Load
**Date**: 2026-02-18
**Scope**: Browser first-load of Codeman web UI (`/`)
**Method**: Static analysis by 4 parallel audit agents (server, frontend, SSE/xterm, asset pipeline)
---
## Current State Summary
### Payload Sizes (measured from live server, port 3000)
| Asset | Raw Size | Gzip | Brotli | Lines | Render-Blocking? |
|-------|----------|------|--------|-------|-----------------|
| `index.html` | 82 KB | 15 KB | 15 KB | 1,479 | N/A (document) |
| `app.js` | 562 KB | 126 KB | 125 KB | 15,354 | No (`defer`) |
| `styles.css` | 154 KB | 25 KB | 27 KB | 8,199 | **YES** |
| `mobile.css` | 34 KB | 7 KB | 7 KB | 1,493 | **YES** (no media query!) |
| `xterm.css` (CDN) | 2 KB | 2 KB | — | — | **YES** (external CDN) |
| `xterm.min.js` (CDN) | 67 KB | 65 KB | — | — | No (`defer`) |
| `xterm-addon-fit` (CDN) | 1 KB | 1 KB | — | — | No (`defer`) |
| **Total local** | **832 KB** | **173 KB** | **174 KB** | | |
| **Total w/ CDN** | **~902 KB** | **~241 KB** | | | |
**Server compression**: Brotli preferred (`Content-Encoding: br`), via `@fastify/compress` with threshold 1024. Compression is **on-the-fly per request** — no pre-compressed files exist.
**HTTP headers verified**: `Cache-Control: public, max-age=3600`, weak ETags auto-generated by `@fastify/static`, `Vary: accept-encoding`, CSP + security headers present.
### Request Waterfall on First Load (6-7 API calls!)
```
Browser hits /
├── index.html ............................ (82 KB document)
├── styles.css?v=0.1533 .................. (render-blocking CSS, 154 KB)
├── mobile.css?v=0.1533 .................. (render-blocking CSS, 34 KB — wasted on desktop!)
├── xterm.css (CDN) ...................... (render-blocking CSS — external!)
├── xterm.min.js (CDN, defer) ........... (67 KB, parallel download)
├── xterm-addon-fit.min.js (CDN, defer) .. (1 KB, parallel download)
├── app.js?v=0.1533 (defer) ............. (562 KB, parallel download)
│
│ [FIRST PAINT blocked until ALL CSS downloaded + parsed]
│
├── JS executes: new CodemanApp().init()
│ ├── initTerminal() ................... (SYNC: new Terminal() + terminal.open() → canvas creation)
│ ├── connectSSE() → /api/events ....... (SSE → fires 'init' with getLightState())
│ ├── loadState() → /api/status ........ (DUPLICATE #1: same data as SSE init!)
│ ├── loadQuickStartCases()
│ │ ├── /api/settings ................ (settings fetch #1)
│ │ └── /api/cases?_t=<timestamp> ... (case list, cache-busted!)
│ ├── startSystemStatsPolling() → /api/system/stats (every 2s, starts immediately)
│ └── loadAppSettingsFromServer() → /api/settings (DUPLICATE #2: settings fetched again!)
```
**Total init API calls**: 6-7 requests, with **2 duplicates** (`/api/status` = SSE init, `/api/settings` fetched twice).
### Critical Path Bottlenecks
1. **3 render-blocking CSS files** (one from CDN, one wasted on desktop)
2. **Synchronous `terminal.open()`** blocks main thread during init (canvas creation)
3. **Double `handleInit()` execution** — SSE init + `/api/status` both call it, causing full state reset + cleanup twice within ~100ms
4. **`/api/settings` fetched twice** — once in `loadQuickStartCases()`, once in `loadAppSettingsFromServer()`
5. **No loading skeleton** — blank `#0a0a0a` screen until CSS+JS fully loaded
6. **12 modals pre-rendered** in HTML — ~600+ DOM elements, ~60KB of invisible HTML
7. **562KB monolith `app.js`** unminified — 1,525 comment lines (10%), 89 `console.*` statements, 23% whitespace
8. **No minification in build** — `cp -r` copies raw source to dist
9. **Stats polling starts immediately** — 2s interval even with no sessions
10. **Version query strings stale** — HTML has `?v=0.1533`, package.json is `0.1534`
### What's Already Good
- Only **1 xterm Terminal instance** shared across all sessions (buffer swapping on tab switch)
- Teammate terminals created **lazily** on window open (with `requestAnimationFrame` defer)
- Subagent windows use **HTML activity logs**, not additional Terminal instances
- `getLightState()` has a **1-second TTL cache** — no duplicate server-side computation
- SSE init sends **lightweight state** (no terminal buffers) — buffers fetched on-demand per tab
- Buffer hydration uses **chunked writes** (128KB chunks via `requestAnimationFrame`) — no UI jank
- `selectSession()` defers secondary panels via **`requestIdleCallback`**
- Buffer fetch is **tail-mode** (last 256KB only, not full 2MB)
- **System fonts only** — no web font downloads blocking paint
- All JS scripts use **`defer`**
- SSE reconnection has **proper exponential backoff** with timeout cleanup
---
## Optimization Plan
### Phase 1: Quick Wins (High Impact, Low Effort)
#### 1.1 Add `media` attribute to mobile.css
**Impact**: HIGH — 34KB CSS stops blocking render on desktop
**Effort**: 1 line change
**File**: `src/web/public/index.html:14`
```html
<!-- Before -->
<link rel="stylesheet" href="mobile.css?v=...">
<!-- After -->
<link rel="stylesheet" href="mobile.css?v=..." media="(max-width: 1023px)">
```
The browser still downloads it (for potential resize) but won't block rendering on desktop. The `mobile.css` comment on line 4 says this was *intended* but never implemented.
#### 1.2 Eliminate duplicate `/api/status` fetch + double `handleInit()`
**Impact**: HIGH — removes 1 redundant API call + eliminates double state reset (clearing 15+ Maps, 7+ timers, `cleanupAllFloatingWindows()`, double `renderSessionTabs()`, double async subagent restore chain)
**Effort**: Small
**Files**: `src/web/public/app.js:1554`, `app.js:3566-3574`
The SSE `init` event (`server.ts:618`) already sends `getLightState()`. The `loadState()` at `app.js:1554` fetches identical data from `/api/status`. Both call `handleInit()` which does a full state reset — whichever arrives second **wipes all state from the first** and rebuilds from scratch.
The `_initGeneration` guard (line 3373/3549) only protects the session-restore at the end, NOT the expensive full cleanup (lines 3389-3503).
**Approach**: Remove `this.loadState()` from `init()`. Add a fallback timeout:
```js
// In init():
this.connectSSE();
// Remove: this.loadState();
this._initFallbackTimer = setTimeout(() => {
if (this._initGeneration === 0) this.loadState();
}, 3000);
```
Clear the timer in `handleInit()`:
```js
handleInit(data) {
if (this._initFallbackTimer) {
clearTimeout(this._initFallbackTimer);
this._initFallbackTimer = null;
}
// ... rest of handleInit
}
```
#### 1.3 Deduplicate `/api/settings` fetch
**Impact**: MEDIUM — removes 1 redundant API call
**Effort**: Small
**Files**: `src/web/public/app.js:7341` (in `loadQuickStartCases`), `app.js:9964` (in `loadAppSettingsFromServer`)
Both fetch `/api/settings`. Fetch it once, pass the result to both consumers:
```js
// In init():
const settingsPromise = fetch('/api/settings').then(r => r.json());
this.loadQuickStartCases(null, settingsPromise);
this.loadAppSettingsFromServer(settingsPromise);
```
#### 1.4 Defer system stats polling
**Impact**: MEDIUM — removes 1 API call every 2s when idle
**Effort**: Small
**Files**: `src/web/public/app.js:1567`, `app.js:15261-15271`
`fetchSystemStats()` already has a visibility guard (line 15282: skips if `#headerSystemStats` is `display: none`), but the interval still ticks. Move `startSystemStatsPolling()` out of `init()` — start it in `handleInit()` only when `data.sessions.length > 0`.
#### 1.5 Preload xterm.css to unblock render
**Impact**: MEDIUM — external CDN CSS currently blocks first paint
**Effort**: 2 line change
**File**: `src/web/public/index.html:15`
```html
<!-- Before -->
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css">
<!-- After -->
<link rel="preload" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css" as="style" onload="this.onload=null;this.rel='stylesheet'">
<noscript><link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css"></noscript>
```
Terminal won't display until xterm.js executes anyway, so the CSS doesn't need to block initial paint.
#### 1.6 Fix stale version query strings
**Impact**: LOW — prevents serving cached stale assets after deploy
**Effort**: Small
**File**: COM script in CLAUDE.md
The HTML references `?v=0.1533` while package.json is already at `0.1534`. The COM workflow should auto-update HTML version strings. Add to the COM script:
```bash
# After incrementing version in package.json + CLAUDE.md:
sed -i "s/?v=[0-9.]*/?v=$NEW_VERSION/g" src/web/public/index.html
```
#### 1.7 Remove cache-busting from `/api/cases`
**Impact**: LOW — allows HTTP caching of case list
**Effort**: 1 line change
**File**: `src/web/public/app.js:7351`
```js
// Before:
const res = await fetch('/api/cases?_t=' + Date.now());
// After:
const res = await fetch('/api/cases');
```
The case list rarely changes during a session. Let the browser cache it.
---
### Phase 2: Medium Effort (High Impact)
#### 2.1 Add loading skeleton
**Impact**: MEDIUM-HIGH — perceived performance improvement (instant visual structure)
**Effort**: Small-Medium
**File**: `src/web/public/index.html`
Add minimal inline `<style>` + skeleton HTML in `<body>` showing a dark header bar + terminal placeholder. Hidden by `app.js` once init completes:
```html
<style>
.skeleton { display: flex; flex-direction: column; height: 100vh; }
.skeleton-header { height: 40px; background: #111; border-bottom: 1px solid #222; }
.skeleton-terminal { flex: 1; background: #0d0d0d; }
.app-loaded .skeleton { display: none; }
</style>
<div class="skeleton">
<div class="skeleton-header"></div>
<div class="skeleton-terminal"></div>
</div>
```
In `app.js` init(), add `document.body.classList.add('app-loaded')` at the end.
#### 2.2 Defer xterm.js terminal creation to after first paint
**Impact**: MEDIUM-HIGH — `terminal.open()` is the heaviest synchronous call in init
**Effort**: Medium
**Files**: `src/web/public/app.js:1545`, `app.js:1578-1639`
```js
init() {
// ... mobile detection, visibility settings ...
document.documentElement.classList.remove('mobile-init');
// Show skeleton/header immediately, defer heavy terminal init
requestAnimationFrame(() => {
this.initTerminal();
this.connectSSE();
// ... rest of init
});
}
```
Lets the browser paint the header/tabs before the terminal canvas is created.
#### 2.3 Batch initial API calls into one endpoint
**Impact**: MEDIUM — reduces 4+ API calls to 1
**Effort**: Medium
**Files**: `src/web/server.ts`, `src/web/public/app.js`
Create `/api/init-bundle`:
```json
{
"status": { /* getLightState() */ },
"cases": [ /* case list */ ],
"settings": { /* user settings */ }
}
```
Replaces `/api/status` (fallback), `/api/cases`, `/api/settings`. Saves HTTP round trips and server-side work.
#### 2.4 Lazy-create modals on first open
**Impact**: HIGH — removes ~600+ DOM elements from initial parse (~60KB of HTML)
**Effort**: Medium-High
**Files**: `src/web/public/index.html`, `src/web/public/app.js`
12 modals pre-rendered in `index.html`:
- `helpModal` (lines 227-447)
- `sessionOptionsModal` (lines 448-714) — **266 lines alone**
- `appSettingsModal` (lines 715-900+)
- `createCaseModal`, `mobileCasePickerModal`, `ralphWizardModal`, `killAllModal`, `closeConfirmModal`, `savePresetModal`, `tokenStatsModal`, `filePreviewModal`, notification drawer
**Approach**: Replace each modal's HTML with `<div id="helpModal" class="modal"></div>`. On first open, inject full HTML via `createModalContent()`. Cache after creation.
---
### Phase 3: Build Pipeline (Highest Impact)
#### 3.1 Self-host xterm.js assets
**Impact**: MEDIUM — eliminates CDN dependency + latency, enables local caching
**Effort**: Low-Medium
**Files**: `src/web/public/index.html`, `package.json` build script
```bash
# Build script addition:
mkdir -p dist/web/public/vendor
cp node_modules/xterm/css/xterm.css dist/web/public/vendor/
cp node_modules/xterm/lib/xterm.min.js dist/web/public/vendor/
cp node_modules/@xterm/addon-fit/lib/xterm-addon-fit.min.js dist/web/public/vendor/
```
Update HTML to reference `/vendor/xterm.min.js` etc. Removes render-blocking CDN CSS entirely.
#### 3.2 Add esbuild minification to build
**Impact**: HIGH — ~38 KB compressed savings (16% of local payload)
**Effort**: Medium
**Files**: `package.json` (build script)
Current build just does `cp -r src/web/public dist/web/`. No minification at all.
**app.js specifics**: 1,525 comment lines (10%), 89 `console.*` statements, 23% whitespace.
```bash
# Add to build script:
npx esbuild dist/web/public/app.js --minify --drop:console --outfile=dist/web/public/app.js --allow-overwrite
npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite
npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite
```
Expected: `app.js` 562KB → ~350KB minified → ~90KB gzip (from 126KB). `--drop:console` removes all 89 debug statements.
Note: `app.js` is vanilla JS (not modules), so esbuild works directly as a minifier.
#### 3.3 Pre-compress static assets at build time
**Impact**: MEDIUM — eliminates per-request CPU compression work
**Effort**: Low
**Files**: `package.json` build script, `src/web/server.ts`
Currently `@fastify/compress` compresses on-the-fly for every request. Pre-compress at build time:
```bash
# Build script:
for f in dist/web/public/*.{js,css,html}; do
gzip -9 -k "$f"
brotli -9 -k "$f"
done
```
Then configure `@fastify/static` with `preCompressed: true` (if supported) or serve pre-compressed files via custom logic.
#### 3.4 Extract critical CSS inline
**Impact**: MEDIUM — eliminates render-blocking `styles.css` for first paint
**Effort**: Medium-High
**Files**: `src/web/public/styles.css`, `src/web/public/index.html`
Identify ~2-3KB of CSS needed for first paint (body, header, tab bar, terminal container) and inline it in `<head>`. Load full `styles.css` asynchronously:
```html
<style>/* ~50 lines of critical CSS */</style>
<link rel="preload" href="styles.css?v=..." as="style" onload="this.onload=null;this.rel='stylesheet'">
```
---
## Impact Estimates
| # | Optimization | First Paint | TTI | Effort |
|---|-------------|-------------|-----|--------|
| 1.1 | mobile.css media query | -50ms | — | 1 min |
| 1.2 | Remove duplicate fetch + double handleInit | — | -100-200ms | 15 min |
| 1.3 | Deduplicate settings fetch | — | -50ms | 10 min |
| 1.4 | Defer stats polling | — | -20ms | 10 min |
| 1.5 | Preload xterm.css | -100-300ms | — | 5 min |
| 1.6 | Fix stale version strings | cache correctness | — | 5 min |
| 1.7 | Remove cases cache-bust | — | -10ms | 1 min |
| 2.1 | Loading skeleton | perceived -500ms | — | 30 min |
| 2.2 | Defer terminal init | -50-100ms | -50ms | 30 min |
| 2.3 | Batch API endpoint | — | -100-200ms | 1 hr |
| 2.4 | Lazy modals | -30-50ms parse | -50ms | 2-3 hrs |
| 3.1 | Self-host xterm | -100-300ms | — | 20 min |
| 3.2 | Minify JS/CSS | -50-100ms parse | — | 30 min |
| 3.3 | Pre-compress assets | -10-30ms TTFB | — | 20 min |
| 3.4 | Critical CSS inline | -200-400ms | — | 2 hrs |
**Combined estimate**: First paint **300-800ms faster**, TTI **200-500ms faster**.
---
## Implementation Order (for implementation agent)
Do these in order — each step is independently testable:
1. **1.1** — mobile.css media query (1 line, instant win)
2. **1.5** — Preload xterm.css (2 lines, big render-blocking fix)
3. **1.2** — Remove duplicate `/api/status` + double handleInit
4. **1.3** — Deduplicate `/api/settings` fetch
5. **1.7** — Remove cache-busting from `/api/cases`
6. **3.1** — Self-host xterm.js (removes CDN dependency entirely)
7. **3.2** — Add esbuild minification to build
8. **1.4** — Defer stats polling
9. **1.6** — Fix stale version strings in COM workflow
10. **2.1** — Loading skeleton
11. **2.2** — Defer terminal init after first paint
12. **2.3** — Batch init API endpoint
13. **2.4** — Lazy modals (biggest refactor, do last)
14. **3.3** — Pre-compress assets (nice-to-have)
15. **3.4** — Critical CSS extraction (only if still needed after above)
**Verification after each step**: Use Playwright to load the page with `waitUntil: 'domcontentloaded'`, measure first paint timing, check that the UI renders correctly with 3-4s wait for async data.
-423
View File
@@ -1,423 +0,0 @@
# Performance Analysis & Optimization Opportunities
**Date**: 2026-03-07
**Scope**: Full-stack performance analysis — backend PTY handling, SSE broadcasting, frontend terminal rendering, local echo overlay, DOM updates, config/scaling limits.
**Constraint**: All recommendations preserve existing functionality including local echo, backpressure, anti-flicker pipeline, and mobile support.
---
## Executive Summary
The codebase is already well-optimized in critical paths. The multi-layer backpressure system, adaptive terminal batching, DEC 2026 sync markers, and incremental state serialization are strong. The main opportunities are in **reducing unnecessary work** (SSE filtering, DOM rebuilds, lazy terminal init) rather than algorithmic changes.
**Top 5 high-impact opportunities:**
| # | Optimization | Impact | Risk | Effort |
|---|-------------|--------|------|--------|
| 1 | Session-scoped SSE subscriptions | Bandwidth -60-80%, CPU -40% | Medium | Medium |
| 2 | Lazy xterm.js for minimized subagent windows | Memory -3.5MB at 50 agents | Low | Low |
| 3 | Targeted badge update (skip full tab rebuild) | Eliminates O(n) reflow on badge change | Low | Low |
| 4 | Conditional SSE padding (tunnel-only, terminal-only) | Bandwidth -70% when tunneled | Low | Low |
| 5 | Canvas renderer on mobile | GPU pressure reduction, battery savings | Low | Low |
---
## 1. SSE Broadcasting
### Current State
- **92 event types** broadcast to all connected clients (max 100)
- Single `JSON.stringify()` per event, shared across all clients (efficient)
- **No per-client filtering** — every client receives every event regardless of which session they're viewing
- 8KB padding appended to **every** event when tunnel is active (forces Cloudflare proxy flush)
- Backpressure: clients marked as backpressured if `reply.raw.write()` returns false; recovery via `session:needsRefresh`
### Bottlenecks
**B1: No session-scoped SSE subscriptions** (`server.ts:1986`)
- Client viewing session A still receives all events for sessions B through T
- With 20 active sessions, ~95% of terminal events are irrelevant to any given client
- Cost: wasted bandwidth, CPU for JSON parsing, and event handler dispatch on client
**B2: Unconditional 8KB padding** (`server.ts:1977`)
- Every event gets 8KB comment padding when tunnel is active
- A `task:updated` event (~200 bytes payload) becomes ~8.2KB
- High-frequency events like `session:terminal` need the padding; low-frequency events like `session:created` don't
### Recommendations
**R1: Session-scoped SSE subscriptions** (High impact)
- Add `?sessions=id1,id2` query param to `/api/events` SSE endpoint
- Server filters events by session ID before broadcasting
- Client subscribes to active session + "global" events (session lifecycle, system)
- Re-subscribes on tab switch (or subscribe to all with client-side filter as fallback)
- **Savings**: ~80% bandwidth reduction for single-session viewers; ~60% for multi-session dashboards
**R2: Tiered SSE padding** (Medium impact)
- Only pad `session:terminal` events and SSE heartbeats (the two that need proxy flush)
- Skip padding for low-frequency structural events (`session:created`, `task:updated`, etc.)
- **Savings**: ~70% padding overhead reduction; terminal events already large enough to flush
---
## 2. Terminal Rendering
### Current State (Well-Optimized)
- **6-layer anti-flicker pipeline**: Server batching (adaptive 16-50ms) → DEC 2026 sync wrap → single JSON serialize → client rAF batching → sync segment parser → chunked buffer loading (32KB/frame)
- **64KB/frame write budget** with DEC 2026 sync-segment awareness (prevents 141KB single-frame freezes)
- **3-layer backpressure**: SSE cap (128KB queued → drop + refresh), frame budget (64KB/frame), chunked restore (32KB/frame)
- WebGL renderer enabled by default with canvas fallback on context loss
- Typical latency: 16-32ms; worst case: ~115ms (50ms server batch + 50ms sync wait + 16ms rAF)
### Bottlenecks
**B3: WebGL on mobile** (`app.js:627-637`)
- Mobile GPUs are weaker; WebGL context loss more likely on low-end devices
- Canvas renderer is sufficient for mobile (typically 1 session, smaller viewport)
**B4: Static scrollback for all sessions** (`app.js:572`)
- Default 5000 lines scrollback for all sessions regardless of activity level
- Heavy output sessions (build logs, test runners) accumulate large scroll buffers
**B5: No addon lazy loading**
- FitAddon, Unicode11Addon, and WebGLAddon all loaded at terminal init
- Unicode11Addon only needed for CJK content; WebGLAddon is large
### Recommendations
**R3: Force canvas renderer on mobile** (Low risk)
- Detect `MobileDetection.isMobile()` and skip WebGL addon loading
- Reduces GPU memory pressure, prevents context loss crashes
- Mobile typically has 1-2 sessions — canvas performance is more than adequate
**R4: Dynamic scrollback based on session activity** (Low risk)
- Active sessions (working state): 5000 lines (current default)
- Inactive/idle sessions: reduce to 2000 lines
- Restore on session select (fetch from server buffer)
- **Savings**: ~60% scrollback memory for idle sessions
**R5: Lazy-load Unicode11Addon** (Low risk)
- Only load when CJK content is detected in terminal output
- Detection: check for characters in CJK Unicode ranges during ANSI stripping (already iterating)
- Most sessions never need it
---
## 3. DOM & Session Tab Rendering
### Current State
- Session tabs use **intelligent incremental updates** with debounced 100ms rendering
- Incremental path: only updates changed properties (classes, textContent, badges) when session list is stable
- Full rebuild path: triggered when sessions added/removed **or badge count changes**
- Subagent windows: per-window xterm.js instances, even when minimized
### Bottlenecks
**B6: Badge count change triggers full tab rebuild** (`app.js:3207-3209`)
- A single subagent badge increment on one tab triggers `_fullRenderSessionTabs()` — rebuilds entire sidebar HTML via `innerHTML =`
- With 20 sessions, this is an O(n) reflow for a single badge number change
- Badge changes are frequent during active subagent work
**B7: Minimized subagent windows retain xterm.js instances** (`subagent-windows.js`)
- 50 subagent windows × ~75KB per xterm.js instance = ~3.75MB DOM memory
- Minimized windows are invisible but their terminals remain in DOM
- xterm.js instances continue processing resize events even when hidden
**B8: `backdrop-filter: blur()` on overlays** (`styles.css:2246-2247, 3098`)
- Forces new stacking context, disables browser compositing optimizations
- 50-100ms layout thrashing on modal open/close
- Only 2 uses, but they're on frequently toggled overlays
### Recommendations
**R6: Targeted badge update without full rebuild** (Low risk)
- When badge count changes but session list is stable, update only the badge `<span>` textContent
- Keep incremental path for badge changes; only use full rebuild for structural changes (add/remove sessions)
- **Savings**: Eliminates O(n) reflow per badge change; reduces to O(1) targeted update
**R7: Lazy xterm.js initialization for subagent windows** (Medium impact)
- Only create xterm.js Terminal instance when window is restored/maximized
- On minimize: serialize terminal buffer, dispose Terminal instance, keep buffer in memory
- On restore: create new Terminal, write buffer back
- **Savings**: ~3.5MB DOM reduction at 50 minimized agents; eliminates hidden resize processing
- **Trade-off**: ~200-500ms restore delay (buffer write), mitigated by chunked loading
**R8: Replace `backdrop-filter: blur()` with `background: rgba()`** (Low risk)
- Use semi-transparent background instead of blur effect
- Or use `will-change: transform` hint if blur is kept
- **Savings**: Eliminates forced recomposition layer; 50-100ms faster overlay open
---
## 4. Backend PTY & State Management
### Current State (Excellent)
- **BufferAccumulator**: Array-based chunking with lazy join on read — avoids O(n) string concatenation
- **ANSI stripping**: Throttled at 150ms intervals with lazy evaluation (not per-chunk)
- **State persistence**: 500ms debounce + incremental JSON caching per session (only dirty sessions re-serialized)
- **Expensive parsers**: Throttled to 150ms window, accumulated data capped at 64KB
- **Memory**: All buffers have hard limits (2MB terminal, 1MB text, 1000 messages, 64KB line buffer)
### Bottlenecks
**B9: Pending clean data cap at 64KB** (`session.ts:1097-1133`)
- Between 150ms processing windows, raw PTY data accumulates in `_pendingCleanData`
- Capped at 64KB — excess data rolls off (old data discarded)
- During heavy output (large build logs), this means parsers may miss content
- Acceptable trade-off for performance, but worth documenting
**B10: `LRUMap.delete()` is O(n) worst case** (`utils/lru-map.ts:137-138`)
- When deleting the newest entry, iterates all keys to find new newest
- Rare in practice (delete is uncommon; set/get are hot paths)
- Could matter during mass cleanup of 500 agents
### Recommendations
**R9: Consider adaptive pending data cap** (Low priority)
- During idle detection (critical to get right), increase cap to 128KB
- During active working state, keep at 64KB (parsers less critical)
- **Benefit**: More accurate idle detection during heavy output
**R10: Track second-newest in LRUMap** (Low priority)
- Maintain a `_secondNewestKey` alongside `_newestKey`
- On delete of newest, promote second-newest without iteration
- Only matters at scale (500+ agents with frequent eviction)
---
## 5. Local Echo & Input Path
### Current State (Well-Designed)
- **DOM overlay approach** — `<span>` elements in `.xterm-screen` at z-index 7, completely independent of `terminal.write()`
- **Render caching**: `_lastRenderKey` includes text, position, column offsets — skips redundant re-renders
- **Input flow**: Char accumulation → Enter triggers flush → 80ms delay before `\r` (ensures text reaches PTY first)
- **Tab completion**: Baseline snapshot → detect buffer change → 300ms fallback timer
- **CJK support**: Per-character width detection with `terminal.unicode.getStringCellWidth()` preferred, manual fallback
- **Prompt detection**: Bottom-up line scan, O(rows) — cached position, column-lock prevents jitter
### Bottlenecks
**B11: tmux send-keys latency** (~50-100ms per input)
- Each `writeViaMux()` spawns a child process (`tmux send-keys`)
- Text and Enter sent separately with 50ms delay between
- For rapid typing: characters batch before Enter, so overhead is per-command not per-keystroke
- **Acceptable trade-off** for session persistence (tmux survives server restarts)
**B12: 80ms delay between text flush and Enter** (`app.js:872-875`)
- Intentional: ensures text reaches PTY before Enter, preventing Ink from processing empty input
- Adds 80ms to perceived Enter-to-response latency
- Could potentially be reduced with acknowledgment-based approach
**B13: Scroll listener on terminal viewport** (`zerolag-input-addon.ts:139`)
- 50ms debounced re-render on scroll — acceptable but fires frequently during heavy output
- Overlay hidden when scrolled up (correct behavior), shown when at bottom
### Recommendations
**R11: Reduce Enter delay from 80ms to 50ms** (Low risk, test carefully)
- The tmux `send-keys` already has 50ms internal delay
- Combined with network latency, 80ms client-side may be excessive
- Test with Ink-heavy sessions (Claude Code's status bar) — if text arrives before Enter at 50ms, reduce
- **Savings**: 30ms perceived latency reduction per command
**R12: Batch tmux send-keys via stdin pipe** (Medium effort, high impact for rapid input)
- Instead of spawning `tmux send-keys` per input, maintain a persistent connection
- Use `tmux -C` (control mode) for programmatic interaction without child process spawning
- **Savings**: Eliminate ~50-100ms process spawn overhead per input
- **Risk**: Control mode has different semantics; needs careful testing with session persistence
**R13: Skip overlay re-render during heavy output scroll** (Low risk)
- When terminal is receiving >10KB/s output, hide overlay entirely (user isn't typing during heavy output)
- Re-show overlay after 500ms of output silence
- **Savings**: Eliminates unnecessary DOM overlay re-renders during build logs / test output
---
## 6. Polling & File Watchers
### Current State
- **SubagentWatcher**: 1s base poll, full scan throttled to every 5s, fs.watch() on known directories
- **TranscriptWatcher**: 1 per session, fs.watch() primary with 1s poll fallback
- **ImageWatcher**: chokidar per session with 100ms stability poll, burst limit 20/10s
- **TeamWatcher**: chokidar primary with 30s poll fallback, LRU caches (50 teams, 200 tasks)
- **RalphTracker**: Todo cleanup every 5 minutes
### Scaling Profile (20 sessions)
| Component | Instances | Frequency | Total ops/sec |
|-----------|-----------|-----------|---------------|
| SubagentWatcher | 1 (global) | Full scan every 5s | 0.2/s |
| TranscriptWatcher | 20 | 1s poll (fallback) | 20/s max |
| ImageWatcher | 20 | 100ms poll (during writes only) | 200/s burst |
| TeamWatcher | 1 (global) | 30s poll (fallback) | 0.03/s |
| SSE heartbeat | 1 (global) | 15s | 0.07/s |
| SSE dead client check | 1 (global) | 30s | 0.03/s |
| Mux stats collection | 1 (global) | 2s | 0.5/s |
| **Total steady-state** | | | **~21/s** |
### Recommendations
**R14: Increase TranscriptWatcher poll interval to 2s** (Low risk)
- Transcript changes are infrequent (new messages every few seconds at most)
- fs.watch() is the primary mechanism; polling is fallback
- **Savings**: Halves fallback filesystem checks (20/s → 10/s for 20 sessions)
**R15: Share chokidar instances for co-located session directories** (Medium effort)
- Sessions in the same parent directory could share a single chokidar watcher with depth:3
- Common case: multiple sessions in `~/projects/foo/` — one watcher covers all
- **Savings**: Reduce chokidar instances from 20 to ~5-10 for typical workloads
---
## 7. Frontend Asset Delivery
### Current State
- **app.js**: 12,027 lines (source) → esbuild minified → gzip/brotli compressed (~30-40KB gzipped)
- **Static caching**: `maxAge: '1y'` via `@fastify/static`
- **Service worker**: Push notification handler only — no asset caching
- **No code splitting**: Single monolithic app.js bundle
### Bottlenecks
**B14: No cache-busting mechanism**
- `maxAge: '1y'` means browsers cache aggressively
- After deployment, users need `Ctrl+Shift+R` to see updates
- No content hash in filenames or ETags for automatic invalidation
**B15: Monolithic app.js**
- All 12K lines loaded on initial page load regardless of which features are used
- Ralph wizard, plan orchestrator UI, team management — all loaded upfront
- Mobile loads the same bundle as desktop
### Recommendations
**R16: Add content hash to asset filenames** (Medium impact)
- Build step: rename `app.js` → `app.[hash].js`
- Generate a manifest or inject hash into HTML template
- Keep `maxAge: '1y'` — cache invalidation happens via filename change
- **Savings**: Eliminates stale cache issues after deployment; removes need for manual hard refresh
**R17: Code-split app.js into core + feature modules** (High effort, medium impact)
- Core (~4K lines): terminal, SSE, session management, tabs, input handling
- Deferred (~8K lines): Ralph wizard, plan UI, team management, subagent windows, image viewer
- Load deferred modules on first use via dynamic `import()` or lazy `<script>` injection
- **Savings**: ~60% reduction in initial load size; faster time-to-interactive
- **Risk**: Complexity increase; need to handle loading states for deferred features
- **Note**: May not be worth the effort given the app is already gzipped to ~30-40KB
---
## 8. CSS Performance
### Current State
- **styles.css**: 7,153 lines with ~45 box-shadow uses, 2 backdrop-filter uses
- Animations: GPU-accelerated keyframes for pulsing alerts, loading spinners
- Z-index layering: well-organized (subagent 1000, plan 1100, log 2000, image 3000, overlay 7)
### Recommendations
**R18: Replace backdrop-filter with opaque overlay** (Low risk, covered in R8)
**R19: Use `contain: content` on subagent windows** (Low risk)
- Add CSS containment to subagent window containers
- Prevents layout changes inside windows from triggering reflow on parent
- Especially valuable with 50 windows: changes in one window won't invalidate others
- ```css
.subagent-window { contain: content; }
```
- **Savings**: Reduces layout recalculation scope from global to per-window
**R20: Use `content-visibility: auto` on off-screen subagent windows** (Low risk)
- Browser skips rendering of off-screen windows entirely
- Combined with `contain-intrinsic-size` to prevent layout shift
- ```css
.subagent-window.minimized { content-visibility: hidden; }
```
- **Savings**: Browser skips paint/layout for minimized windows; complements R7
---
## 9. Memory & Scaling Limits
### Current Budget (20 sessions)
| Component | Per Session | Total | Status |
|-----------|-----------|-------|--------|
| Terminal buffer | 2MB | 40MB | Hard-limited, auto-trim |
| Text output | 1MB | 20MB | Hard-limited, auto-trim |
| Messages | ~1MB | 20MB | Capped at 1000, trims to 800 |
| Respawn buffer | 1MB | 20MB | Hard-limited |
| **Buffers total** | | **100MB** | Acceptable |
| TranscriptWatcher | ~100KB | 2MB | |
| ImageWatcher | ~50KB | 1MB | |
| SubagentWatcher | ~500KB | 500KB | Global |
| Frontend terminal cache | ~256KB | 5MB | LRU, max 20 entries |
| **Total estimated** | | **~110MB** | Comfortable |
### At Max Scale (50 sessions)
- Buffers: ~250MB
- Watchers: ~5MB
- **Total: ~255MB** + Node.js overhead — acceptable on modern hardware
### Potential Leak Vectors (All Mitigated)
- `_shortIdCache` in server — unbounded Map, but entries are tiny (string→string); grows at O(sessions created), not O(events)
- All CleanupManager-registered resources tracked and disposed on session stop
- `isStopped` guard prevents new timers after session cleanup
---
## 10. Implementation Priority Matrix
### Phase 1 — Quick Wins (1-2 hours each, low risk)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R6 | Targeted badge update | `app.js` (3207-3209) |
| R3 | Canvas renderer on mobile | `app.js` (627-637) |
| R8 | Replace backdrop-filter blur | `styles.css` (2246, 3098) |
| R19 | CSS containment on subagent windows | `styles.css` |
| R20 | `content-visibility: hidden` on minimized windows | `styles.css` |
### Phase 2 — Medium Effort (half-day each)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R2 | Tiered SSE padding | `server.ts` (broadcast function) |
| R7 | Lazy xterm.js for minimized subagents | `subagent-windows.js` |
| R11 | Reduce Enter delay to 50ms | `app.js` (872-875), test with Ink |
| R14 | TranscriptWatcher 2s poll | `transcript-watcher.ts` |
| R16 | Content-hash asset filenames | `build.mjs`, `server.ts` |
### Phase 3 — Larger Initiatives (1-2 days each)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R1 | Session-scoped SSE subscriptions | `server.ts`, `app.js` (SSE connect) |
| R5 | Lazy Unicode11Addon loading | `app.js`, build pipeline |
| R12 | Persistent tmux control mode | `tmux-manager.ts` |
| R17 | Code-split app.js | `app.js`, `build.mjs`, HTML template |
### Not Recommended (Low ROI or High Risk)
| # | Why Not |
|---|---------|
| R4 | Dynamic scrollback adds complexity; memory savings marginal vs total budget |
| R9 | Adaptive pending data cap adds state; current 64KB cap rarely matters |
| R10 | LRUMap.delete() O(n) is theoretical; never triggered at current scale |
| R15 | Shared chokidar instances add directory-matching complexity for minimal gain |
---
## Appendix: Key File Locations
| Area | File | Key Lines |
|------|------|-----------|
| SSE broadcast | `src/web/server.ts` | 1961-1989 (broadcast), 1934-1959 (backpressure) |
| Terminal batching | `src/web/server.ts` | 1994-2048 (per-session adaptive batching) |
| Frame budget | `src/web/public/app.js` | 1370-1478 (flushPendingWrites, 64KB cap) |
| Flicker filter | `src/web/public/app.js` | 1176-1255 (50ms sync wait, 256KB safety) |
| Tab rendering | `src/web/public/app.js` | 3108-3357 (incremental + full rebuild) |
| Tab switching | `src/web/public/app.js` | 3560-3760 (cache + chunked load + deferred UI) |
| Local echo | `packages/xterm-zerolag-input/src/` | All files (overlay, prompt, CJK) |
| Local echo integration | `src/web/public/app.js` | 640, 815-988 (input flow) |
| Subagent windows | `src/web/public/subagent-windows.js` | Full file (window mgmt, drag, minimize) |
| State persistence | `src/state-store.ts` | 161-250 (debounced save, incremental JSON) |
| Buffer accumulator | `src/utils/buffer-accumulator.ts` | Full file (array chunks, lazy join) |
| PTY handling | `src/session.ts` | 1046-1133 (data flow), 1173-1230 (parsing) |
| Config limits | `src/config/` | 9 files (buffer, map, timing, auth, etc.) |
| Anti-flicker docs | `docs/terminal-anti-flicker.md` | Architecture reference |
| CSS | `src/web/public/styles.css` | 2246 (backdrop-filter), full file |
| Build pipeline | `scripts/build.mjs` | 59-68 (minify + compress) |
@@ -1,266 +0,0 @@
# Codeman Performance Investigation Report
**Date**: 2026-02-20
**Scope**: Why Codeman feels sluggish when multiple Claude tabs are very busy
**Method**: 4-agent parallel analysis of server, PTY pipeline, frontend, and background systems
---
## Executive Summary
When multiple Claude sessions are actively producing heavy terminal output (e.g., building, writing files, running tests), Codeman's UI becomes sluggish. This investigation identified **14 bottlenecks** across 4 layers of the stack. The root cause is **cumulative event loop blocking** — no single operation is catastrophically slow, but dozens of small synchronous operations run on every PTY data chunk, and with N busy sessions producing chunks every few milliseconds, the event loop gets saturated.
The most impactful findings are ranked by severity below.
---
## Critical Findings (Event Loop Blockers)
### 1. PTY Data Handler Chain — O(output_volume) per session, synchronous
**File**: `src/session.ts:986-1086`
**Severity**: CRITICAL
Every chunk of PTY output from a busy Claude session runs through this synchronous chain on the Node.js event loop:
```
PTY onData → ANSI strip regex → ralph-tracker → bash-tool-parser →
token parser → CLI info parser → task description parser →
idle/working detection → emit('terminal') → emit('output')
```
**Key costs per chunk:**
- `ANSI_ESCAPE_PATTERN_FULL` regex (line 999): Complex regex with alternation, runs on every chunk where any consumer needs clean data
- `ralphTracker.processCleanData()` (line 1014): Splits into lines, runs regex per line, checks multi-line patterns
- `bashToolParser.processCleanData()` (line 1020): Similar line-by-line regex processing
- `parseTaskDescriptionsFromTerminalData()` (line 1038): Regex scan for parenthesized descriptions
- Working/idle detection (lines 1043-1085): Multiple `includes()` checks plus `getCleanData()` calls
**The lazy `getCleanData()` pattern (line 997-1002)** was a good optimization — it avoids ANSI stripping when no consumer needs it. But when Ralph tracking is enabled (common during active work), `getCleanData()` is called on every chunk, negating the optimization.
**With 5 busy sessions** producing 50+ chunks/second each, this means 250+ synchronous processing chains per second on the event loop. Each chain involves string allocation, regex matching, and line splitting.
### 2. Broadcast Serialization — JSON.stringify on every flush
**File**: `src/web/server.ts:4941-4967`
**Severity**: CRITICAL
The `broadcast()` method calls `JSON.stringify(data)` synchronously for every event. Terminal data is the highest-frequency event. During `flushTerminalBatches()` (line 5030), broadcast is called once per session with pending data. With 10 busy sessions flushing every 16-50ms, that's 200-625 `JSON.stringify` calls per second on terminal data alone.
The terminal data payload is a string that gets double-encoded: the raw terminal string is embedded inside a JSON object `{id, data}`, then that object is JSON.stringify'd. For large chunks (up to 32KB per the `BATCH_FLUSH_THRESHOLD`), this creates significant garbage collection pressure.
**Additionally**, the `session:updated` broadcast includes `toLightDetailedState()` which serializes `taskTree`, `tokens`, `bufferStats`, and `respawnConfig` — this is called on many state changes, not just terminal data.
### 3. Single-Timer Batching — All sessions share one setTimeout
**File**: `src/web/server.ts:5017-5027`
**Severity**: HIGH
The `batchTerminalData()` method uses a **single shared timer** (`this.terminalBatchTimer`) for all sessions. When the timer fires, `flushTerminalBatches()` iterates ALL pending sessions and broadcasts each one. This means:
- One extremely busy session's rapid data can force the timer to fire at the minimum interval (16ms), flushing ALL sessions at that rate
- The flush itself iterates all pending sessions synchronously
- The `_minBatchInterval` optimization (line 5003) means the fastest session dictates the timer for everyone
This creates a **thundering herd** effect: all session flushes happen in a single synchronous burst rather than being staggered.
### 4. State Persistence Storms
**File**: `src/web/server.ts:3879-3917`
**Severity**: HIGH
`persistSessionState()` is called from **28+ locations** in server.ts. Each call sets a 100ms debounce timer per session. During heavy activity, this means:
- Frequent timer creation/cancellation (GC pressure)
- The actual persist (`_persistSessionStateNow`) calls `session.toState()` which creates a new object, then `store.setSession()` which triggers `JSON.stringify` of the entire state store and `writeFileSync` to disk
The `StateStore` (via `state-store.ts`) debounces its own write, but the overhead is in the per-session `toState()` serialization and object creation, not just the disk write.
---
## High-Severity Findings
### 5. Ralph Tracker Line Processing — O(lines) per chunk
**File**: `src/ralph-tracker.ts:1337-1375`
**Severity**: HIGH (when Ralph tracking is enabled)
When enabled, `processCleanData()`:
1. Appends to a line buffer (string concatenation)
2. Splits on `\n` (creates array)
3. Calls `processLine()` on each line (regex matching per line)
4. Calls `checkMultiLinePatterns()` (additional regex on full chunk)
5. Calls `maybeCleanupExpiredTodos()` (iterates todos Map)
For a busy session producing 100+ lines/second, this is significant. The line buffer can grow up to `MAX_LINE_BUFFER_SIZE` before being truncated, and the split/iterate pattern creates garbage on every chunk.
### 6. Subagent Watcher Polling — O(agents) every 1-10 seconds
**File**: `src/subagent-watcher.ts:225-274`
**Severity**: MEDIUM-HIGH
Three periodic operations:
- **Poll interval** (1s): Lightweight check, but full directory scan every 5th poll (5s)
- **Liveness check** (10s): Runs `pgrep` (child process spawn), then iterates ALL tracked agents to check if alive. With 50+ subagents (common with agent teams), this is a non-trivial burst.
- **File watchers**: One `chokidar` watcher per tracked agent directory, plus transcript file watchers. With many agents, this means many active file watchers consuming kernel inotify resources.
The `getClaudePids()` call spawns a child process (`pgrep`) every 10 seconds. Under heavy load, child process spawning competes with the event loop.
### 7. SSE Client Iteration — O(clients) per broadcast
**File**: `src/web/server.ts:4964-4966`
**Severity**: MEDIUM
Every `broadcast()` iterates all SSE clients to send the pre-formatted message. With multiple browser tabs or mobile clients, each flush sends data to every client. The `reply.raw.write()` call goes through Node's HTTP stream, which is generally non-blocking but can cause backpressure cascades.
The backpressure handling (line 4916-4938) correctly skips backpressured clients, but the `once('drain')` handler sends a `session:needsRefresh` event, which the client responds to by fetching the full buffer — potentially a 2MB request — amplifying the problem.
### 8. Event Emitter Fan-Out in Session
**File**: `src/session.ts:1008-1009`
**Severity**: MEDIUM
Every PTY data chunk emits TWO events: `terminal` and `output`. The `terminal` event triggers `batchTerminalData()` in server.ts. The `output` event may trigger additional handlers. EventEmitter dispatch is synchronous — all listeners run before the next operation in the PTY handler continues.
With busy sessions, this means every chunk blocks the event loop for: PTY processing + all terminal listeners + all output listeners.
---
## Medium-Severity Findings
### 9. Respawn Controller Timer Accumulation
**File**: `src/respawn-controller.ts` (various)
**Severity**: MEDIUM
Each session with respawn enabled runs multiple timers:
- Idle detection timeout
- AI checker interval (when active)
- Output silence detection interval
- Token stability interval
- Circuit breaker state timeouts
With 10 sessions with respawn, that's 50+ active timers. While individual timers are cheap, the cumulative effect on the event loop's timer queue is non-trivial — the libuv timer heap has O(log n) insertion but all callbacks run synchronously.
### 10. Team Watcher Polling
**File**: `src/team-watcher.ts`
**Severity**: MEDIUM (when agent teams are active)
Polls `~/.claude/teams/` directory every few seconds. Each poll reads config.json files and task files. With active teams, this adds filesystem reads to the event loop's I/O budget.
### 11. Frontend Terminal Write Batching
**File**: `src/web/public/app.js` (batchTerminalWrite/flushPendingWrites)
**Severity**: MEDIUM
The frontend batches terminal writes at `requestAnimationFrame` rate (16ms). When receiving SSE events from multiple busy sessions:
- `batchTerminalWrite()` is called for EVERY session's data, even sessions not currently displayed
- Terminal instances exist for all sessions (not just the active tab)
- Each `flushPendingWrites()` calls `terminal.write()` which triggers xterm.js rendering
Hidden tabs still process terminal writes, consuming CPU for rendering that's never displayed.
### 12. Frontend Connection Line Rendering
**File**: `src/web/public/app.js` (updateConnectionLines)
**Severity**: LOW-MEDIUM
Connection lines between parent/child agent windows are recalculated on window moves, resizes, and potentially on terminal writes. With many subagent windows open, this involves DOM reads (getBoundingClientRect) that force layout recalculation.
### 13. Image Watcher File System Events
**File**: `src/image-watcher.ts`
**Severity**: LOW
Uses chokidar to watch for image files in session working directories. With many sessions in the same or overlapping directories, watchers may generate redundant events. The `awaitWriteFinish` and burst throttling mitigate this, but the kernel inotify resources add up.
### 14. ANSI Escape Regex Complexity
**File**: `src/session.ts:999`
**Severity**: LOW (but cumulative)
`ANSI_ESCAPE_PATTERN_FULL` is a complex regex with multiple alternation branches. While V8's regex engine handles this well for typical terminal data, adversarial input (deeply nested escape sequences) could cause superlinear matching time. The `FOCUS_ESCAPE_FILTER` regex runs first on every chunk.
---
## Scaling Analysis
| Resource | Per Session | 10 Sessions | 20 Sessions |
|----------|-------------|-------------|-------------|
| PTY data handlers | 1 synchronous chain | 10 chains competing for event loop | 20 chains — event loop saturation likely |
| Broadcast calls (terminal only) | 20-60/sec | 200-600/sec | 400-1200/sec |
| JSON.stringify (terminal) | 20-60/sec | 200-600/sec | 400-1200/sec |
| Active timers | ~5 | ~50 | ~100 |
| File watchers (subagents) | 2-5 | 20-50 | 40-100 |
| SSE writes per flush | N clients | N clients x 10 sessions | N clients x 20 sessions |
| Ralph line processing | O(lines/sec) | O(10 x lines/sec) | O(20 x lines/sec) |
**The critical threshold appears to be 5-8 simultaneously busy sessions**, where the cumulative PTY processing + broadcast serialization + timer callbacks start to exceed the event loop's capacity for responsive handling.
---
## Root Cause Architecture Diagram
```
Busy Claude Session 1 ─┐
Busy Claude Session 2 ─┤ ┌──────────────────────┐
Busy Claude Session 3 ─┼───→│ Node.js Event Loop │
Busy Claude Session 4 ─┤ │ (SINGLE THREAD) │
Busy Claude Session 5 ─┘ │ │
│ PTY handlers (sync) │◄── BOTTLENECK 1
│ ANSI strip regex │
│ Ralph tracker │
│ Bash tool parser │
│ Idle detection │
│ │ │
│ ▼ │
│ EventEmitter.emit() │◄── BOTTLENECK 2
│ │ │
│ ▼ │
│ batchTerminalData() │
│ (shared timer) │◄── BOTTLENECK 3
│ │ │
│ ▼ │
│ flushTerminalBatches() │
│ broadcast() per session│
│ JSON.stringify() each │◄── BOTTLENECK 4
│ write() to N clients │
│ │
│ + persistSessionState │◄── BOTTLENECK 5
│ + respawn timers │
│ + subagent polling │
│ + team watcher │
└────────────────────────┘
```
---
## Recommendations (Not Implemented — For Discussion)
### Tier 1: Highest Impact, Lowest Risk
1. **Disable processing for non-visible sessions**: Skip Ralph tracking, bash tool parsing, and task description parsing for sessions that no active SSE client is viewing. Only buffer terminal data.
2. **Per-session flush staggering**: Instead of one shared timer flushing all sessions, use individual timers offset by `index * (interval/N)` to spread flushes across the batch window.
3. **Skip hidden tab terminal writes on frontend**: Don't call `terminal.write()` for terminals not in the active tab. Lazy-load on tab switch.
### Tier 2: Medium Impact
4. **Worker thread for ANSI stripping and parsing**: Move the regex-heavy ANSI strip + Ralph parsing to a worker thread pool. PTY data → worker → clean data back to main thread.
5. **Pre-formatted SSE messages for terminal data**: Since terminal events are just `{id, data}`, build the SSE message string directly without `JSON.stringify`.
6. **Adaptive processing based on load**: When event loop lag exceeds a threshold (measured via `setTimeout(0)` drift), reduce processing — skip Ralph, increase batch intervals, reduce subagent poll frequency.
### Tier 3: Longer-Term Architectural
7. **Process-per-session or cluster mode**: Move each session's PTY handling to a separate Node.js worker or process, communicating to the main server via IPC.
8. **Binary protocol for terminal data**: Replace JSON-encoded SSE terminal events with binary frames (e.g., MessagePack or raw binary WebSocket frames) to eliminate double-encoding.
9. **Selective SSE subscriptions**: Clients subscribe to specific sessions instead of receiving all events. The server only broadcasts to interested clients.
---
## How to Validate
To confirm these findings, instrument with:
```typescript
// Add to event loop — measures how long synchronous work takes
let lastCheck = Date.now();
setInterval(() => {
const now = Date.now();
const lag = now - lastCheck - 100; // 100ms interval
if (lag > 10) console.log(`[PERF] Event loop lag: ${lag}ms`);
lastCheck = now;
}, 100);
```
And in `flushTerminalBatches()`:
```typescript
const start = performance.now();
// ... existing flush logic ...
const elapsed = performance.now() - start;
if (elapsed > 5) console.log(`[PERF] Flush took ${elapsed.toFixed(1)}ms for ${this.terminalBatches.size} sessions`);
```
This will show exactly when and how much the event loop is being blocked during heavy session activity.
@@ -1,168 +0,0 @@
# Performance & Responsiveness Optimization Plan
**Date**: 2026-02-28
**Status**: Phases 1–4 Complete. Phase 5 optional/deferred.
---
## Executive Summary
Three independent research passes analyzed the Codeman codebase for performance bottlenecks across frontend rendering, backend hot paths, and system-level resource usage. The codebase already has strong foundational optimizations (per-session adaptive batching, rAF terminal writes, DEC 2026 sync markers, backpressure handling). This plan targets the remaining high-impact opportunities.
**Key finding**: The biggest wins come from **skipping unnecessary work** — serializing unchanged state, processing output nobody is watching, and reducing broadcast volume.
---
## Phase 1: Quick Wins — COMPLETE
All Phase 1 items were found to already exist in the codebase during verification:
| # | Item | Status | Evidence |
|---|------|--------|----------|
| 1.1 | Skip terminal writes for hidden tabs | Done | SSE handler filters by `activeSessionId` (app.js:4076) |
| 1.2 | mobile.css media query | Done | `media="(max-width: 1023px)"` on link tag (index.html:13) |
| 1.3 | Deduplicate init API calls | Done | `_initGeneration` dedup + 3s fallback timer (app.js:2901-2904) |
| 1.4 | Remove cache-busting timestamps | Done | No `?_t=` patterns found anywhere |
| 1.5 | JS/CSS minification + compression | Done | esbuild minify + gzip + brotli in build.mjs (lines 42-51) |
---
## Phase 2: Frontend Responsiveness — COMPLETE
### 2.1 Batch `getBoundingClientRect()` in connection lines — DONE
- **Files**: `src/web/public/app.js` (`_updateConnectionLinesImmediate()`)
- **Change**: Refactored to batch all layout reads into Phase 1 (collect all rects into a Map), then perform all SVG writes in Phase 2 using cached values. Classic read-then-write pattern prevents interleaved forced reflows.
### 2.2 Clean up ResizeObservers — Already implemented
- `forceCloseSubagentWindow()` disconnects observers (app.js:12618-12620)
- `cleanupAllFloatingWindows()` disconnects all on reconnect (app.js:12649-12653)
- Observer refs stored on `windowData.resizeObserver` (app.js:12492)
### 2.3 Drag handler cleanup — Already implemented
- `makeWindowDraggable()` returns listener refs, stored in `windowData.dragListeners`
- `forceCloseSubagentWindow()` removes all document-level drag listeners (app.js:12622-12630)
- Panel drags add listeners on mousedown, remove on mouseup (app.js:10253-10284)
### 2.4 Mobile window position cache — Skipped
- O(n) loop over max ~20 windows; complexity of cached counter not justified
### 2.5 Lazy modal DOM — Skipped
- Large effort, marginal benefit for a vanilla JS app with fast DOM construction
---
## Phase 3: Backend Hot Paths — COMPLETE
### 3.1 State diff broadcasts — ALREADY OPTIMIZED
- `broadcastSessionStateDebounced()` already batches at 500ms intervals
- `toLightDetailedState()` excludes heavy buffers (textOutput, terminalBuffer)
- Per-session serialization is <1ms; with debouncing, only 1-3 sessions serialize per flush
- JSON.stringify happens once per broadcast (not per client) — serialization cost is negligible
- Full state diffs would add significant frontend complexity for marginal gain
### 3.2 Improve session list cache hit rate — DONE
- **Files**: `src/web/server.ts` (`broadcast()` method)
- **Change**: Cache now only invalidated on truly structural events (`session:created`, `session:deleted`, `session:updated`) instead of on every `session:*` and `respawn:*` event. High-frequency events like `session:working`, `session:idle`, `session:completion`, `respawn:stateChanged` no longer defeat the 1s TTL cache.
- **Impact**: Cache hit ratio from ~0% to ~80%+ during active sessions. The debounced `session:updated` still refreshes the cache within 500ms of any state change.
### 3.3 Skip PTY processing — ALREADY OPTIMIZED
- `_processExpensiveParsers()` is already throttled to every 150ms (not per-chunk)
- Lazy ANSI stripping via `getCleanData()` closure — only computed when a consumer needs it
- Quick pre-checks skip parsers when content is irrelevant (e.g., token parser only runs if data contains "token")
- OpenCode sessions skip all Claude-specific parsers entirely
- Further optimization would require visibility-aware processing, adding complexity for marginal gain
### 3.4 Batch subagent liveness checks — Deferred
- `/proc/{pid}` stat calls are ~0.1ms each; even with 500 agents, total is 50ms every 10s
- Current approach is simple and reliable; batching adds race condition risk
- Consider only if profiling shows this as a bottleneck
### 3.5 Deduplicate detection update emissions — DONE
- **Files**: `src/respawn-controller.ts` (`startDetectionUpdates()`)
- **Change**: Detection status now only emitted when key fields (confidenceLevel, statusText, controller state) actually change. Previously emitted every 2s regardless, broadcasting identical status to all SSE clients.
- **Impact**: For stable/idle sessions, eliminates ~100% of redundant detection broadcasts. For active sessions, reduces broadcasts to only meaningful state transitions.
---
## Phase 4: System-Level Improvements — COMPLETE
### 4.1 Incremental state persistence — DONE
- **Files**: `src/state-store.ts` (`assembleStateJson()`, `setSession()`)
- **Change**: Added `dirtySessions` Set and `cachedSessionJsons` Map. On persist, only dirty sessions are re-serialized; clean sessions reuse cached JSON fragments. `setSession()` marks sessions dirty; `assembleStateJson()` rebuilds only changed fragments.
- **Impact**: Serialization cost reduced from O(all sessions) to O(dirty sessions). Typical steady-state: 1-2 dirty sessions instead of 50.
### 4.2 Replace polling with fs watchers for team watcher — DONE
- **Files**: `src/team-watcher.ts` (`setupFsWatchers()`)
- **Change**: Added chokidar watchers on both `~/.claude/teams/` and `~/.claude/tasks/` directories for instant event-driven detection. Lock files ignored via chokidar config. Mtime-based dedup skips unchanged files. Polling interval relaxed from 5s to 30s as a fallback.
- **Impact**: Near-instant team detection; polling overhead eliminated for normal operation.
### 4.3 Consolidate subagent file watchers — DONE
- **Files**: `src/subagent-watcher.ts` (`setupDirectoryWatcher()`)
- **Change**: Replaced per-agent chokidar watchers with one `fs.watch()` per session subagent directory. Events are routed to the correct agent via filename. Per-file debouncing (100ms) prevents hammering on bulk discovery.
- **Impact**: Inotify watchers reduced from potentially 500 (one per agent) to ~50 (one per session directory).
### 4.4 Stream transcript files instead of full reads — DONE
- **Files**: `src/subagent-watcher.ts` (`tailFile()`, `findDescriptionInAgentFile()`, parent transcript lookup)
- **Change**: Multiple streaming strategies implemented:
- **Live monitoring**: Position-based `tailFile()` with `createReadStream({ start: fromPosition })` — only reads new content
- **Parent transcript lookup**: Streams only last 16KB (`createReadStream({ start: offset })`)
- **Description extraction**: Streams only first 8KB, exits early after 5 lines
- **Full read**: Only for on-demand transcript review panel (with optional `limit` parameter)
- **Impact**: File I/O for bulk agent discovery reduced from ~50MB to ~5MB.
---
## Phase 5: Long-Term Architectural (Optional) — NOT STARTED
These items are deferred until scaling demands justify the complexity.
### 5.1 Worker thread for PTY processing
- **Files**: `src/session.ts`
- **Problem**: ANSI stripping, Ralph tracking, and bash tool parsing all run on the main event loop. At scale (50 busy sessions), this consumes 300-500ms CPU/sec.
- **Fix**: Offload ANSI strip + line processing to a worker thread pool. Main thread receives clean text + parsed events.
- **Impact**: Frees event loop for I/O operations. Most impactful at 10+ concurrent busy sessions.
### 5.2 Per-session SSE subscriptions
- **Files**: `src/web/server.ts`
- **Problem**: Every SSE event is broadcast to all connected clients. A client watching session A still receives events for sessions B through Z.
- **Fix**: Clients subscribe to specific session IDs. Server only sends events to interested clients.
- **Impact**: Reduces SSE broadcast fan-out from N clients to ~1-2 per event. Major improvement at 100 SSE clients.
### 5.3 O(1) LRUMap via doubly-linked list
- **Files**: `src/utils/lru-map.ts` (~lines 98-110)
- **Problem**: `get()` uses delete + re-insert to refresh position — O(n) on Map iteration for delete.
- **Fix**: Implement classic LRU with doubly-linked list + Map for O(1) get/put/evict.
- **Impact**: Low — current sizes (max 500) make this barely measurable. Only worthwhile if LRUMap is used on hot paths.
---
## Completion Summary
| Phase | Scope | Status | Items |
|-------|-------|--------|-------|
| 1 | Quick Wins | **Complete** | 5/5 (all pre-existing) |
| 2 | Frontend Responsiveness | **Complete** | 3/3 actionable done, 2 skipped |
| 3 | Backend Hot Paths | **Complete** | 4/4 actionable done, 1 deferred |
| 4 | System-Level | **Complete** | 4/4 done |
| 5 | Long-Term Architectural | **Not started** | 0/3 — deferred until needed |
**Overall**: 16/16 actionable items complete. 3 optional items deferred.
---
## Measurement
Before starting Phase 5, establish baselines:
1. **Frontend**: Record Chrome DevTools Performance trace with 10 sessions open. Measure:
- Frame rate during rapid terminal output
- Long tasks (>50ms) count per 30s
- Heap size after 1h session
2. **Backend**: Add `performance.now()` instrumentation around:
- `flushSessionTerminalBatch()` — time per flush
- `broadcastSessionStateDebounced()` — serialization time
- `StateStore.save()` — persist time
- Event loop lag via `monitorEventLoopDelay()`
3. **First load**: Lighthouse score on desktop and mobile (simulated 3G)
-74
View File
@@ -1,74 +0,0 @@
# Codeman Performance Optimization Plan
## Current State
The backend is **already production-grade** — SSE broadcasting, state persistence, terminal batching, buffer management, and memory patterns are all well-optimized. The biggest gains are on the **frontend delivery** side.
## Implemented Optimizations
### 1. V8 Compile Cache (10-20% faster cold start)
**Files:** `scripts/codeman-web.service`, `package.json`
Node.js re-parses and compiles all JS on every cold start. `NODE_COMPILE_CACHE` caches V8 compiled bytecode to disk, reusing it on subsequent starts.
- Added `Environment=NODE_COMPILE_CACHE=/home/arkon/.codeman/compile-cache` to systemd service
- Added to `npm start` script for non-systemd usage
- Zero code changes, immediate win on every restart
### 2. WebGL Addon Lazy-Loading (244KB saved on mobile, non-blocking on desktop)
**Files:** `src/web/public/index.html`, `src/web/public/app.js`
`xterm-addon-webgl.min.js` (244KB) was loaded eagerly for all users via `<script defer>`, but only used on desktop with WebGL2 support.
- Removed `<script defer>` from `index.html`
- Added dynamic script loading in `app.js` — only downloads on desktop when WebGL is needed
- Mobile users never download the file at all (244KB saved)
- Desktop: loads in parallel with page rendering, addon initializes when ready
- Graceful fallback: canvas renderer used if WebGL unavailable or script fails
### 3. Preload Hints (~50-100ms faster perceived load)
**Files:** `src/web/public/index.html`
Browser discovers `<script defer>` tags only when the parser reaches them at the bottom of `<body>`. By then, the HTML parse has blocked for hundreds of lines.
- Added `<link rel="preload" as="script">` in `<head>` for `vendor/xterm.min.js`, `constants.js`, `app.js`
- Browser starts fetching critical scripts immediately during HTML parse (before reaching `<body>`)
- Zero runtime overhead — just hints for the browser's preload scanner
### 4. Batch Tmux Reconciliation (N subprocess calls → 1)
**Files:** `src/tmux-manager.ts`
`reconcileSessions()` previously called `tmux has-session` + `tmux display-message` per known session, plus `tmux list-sessions` for discovery, plus `tmux display-message` per discovered session. With 20 sessions: 41+ subprocess calls.
- Replaced with single `tmux list-panes -a -F '#{session_name}\t#{pane_pid}'` call
- Builds a Map from the result, then does O(1) lookups for both known and discovered sessions
- Also replaced inner O(n) `isKnown` scan with a Set lookup
- 20 sessions: 41 subprocess calls → 1, with faster lookups
### 5. Asset Hashing / Cache Busting (already implemented)
**Files:** `scripts/build.mjs` (pre-existing)
Content-hash cache busting was already implemented in the build script:
- All app JS/CSS files get content hashes (`app.abc123.js`)
- `index.html` rewritten to reference hashed filenames
- Pre-compressed with gzip + Brotli
- 1-year immutable cache works correctly — new deploys get new filenames
## Already Optimized (No Action Needed)
| Area | Why It's Fine |
|------|---------------|
| **SSE Broadcasting** | Single serialization per broadcast, preformatted frames, backpressure handling, session subscription filtering |
| **State Persistence** | 500ms debounce, incremental per-session JSON caching, async atomic writes, circuit breaker on failures |
| **Terminal Batching** | Adaptive intervals (16-50ms), per-session queues, immediate flush at 32KB, array-based accumulation |
| **Buffer Management** | BufferAccumulator (array-push, lazy join), auto-trim at 2MB/1MB, no string concatenation in hot paths |
| **ANSI Stripping** | Pre-compiled regex via factory functions, single-pass processing |
| **Static File Serving** | @fastify/static with 1-year cache, pre-compressed Brotli/gzip, no-cache for HTML |
| **Memory Management** | CleanupManager, LRUMap, StaleExpirationMap, bounded buffers, explicit listener cleanup |
| **Import Patterns** | Pure ESM, lazy web server import, no circular deps, no dynamic imports in hot paths |
| **Config Loading** | Small constant files, no I/O at import time, specific imports (no barrel) |
@@ -1,788 +0,0 @@
# Phase 4: Domain File Splitting — Implementation Plan
**Date**: 2026-03-01
**Prerequisites**: Phase 1-3 complete (utils cleanup, CleanupManager/Debouncer migration, route extraction)
**Goal**: Split 4 god files into focused modules with barrel exports for transparent migration.
---
## Table of Contents
1. [Split types.ts into types/ directory](#1-split-typests-into-types-directory)
2. [Split ralph-tracker.ts into focused modules](#2-split-ralph-trackerts-into-focused-modules)
3. [Split respawn-controller.ts into focused modules](#3-split-respawn-controllerts-into-focused-modules)
4. [Split session.ts into focused modules](#4-split-sessionts-into-focused-modules)
5. [Execution Order & Dependencies](#5-execution-order--dependencies)
6. [Validation Checklist](#6-validation-checklist)
---
## 1. Split types.ts into types/ directory
**Current**: 1,443 lines, 71 exports, imported by 36 files.
**Risk**: LOW — pure type refactor, no runtime behavior change.
### Target Structure
```
src/types/
├── index.ts (barrel re-export — transparent migration)
├── common.ts (Disposable, BufferConfig, CleanupResourceType, CleanupRegistration)
├── session.ts (SessionStatus, SessionMode, ClaudeMode, SessionConfig, SessionColor,
│ SessionState, OpenCodeConfig, SessionOutput)
├── task.ts (TaskStatus, TaskDefinition, TaskState)
├── app-state.ts (AppState, AppConfig, GlobalStats, TokenUsageEntry, TokenStats,
│ DEFAULT_CONFIG, createInitialState, createInitialGlobalStats)
├── respawn.ts (RespawnConfig, PersistedRespawnConfig, CycleOutcome,
│ RespawnCycleMetrics, RespawnAggregateMetrics, HealthStatus,
│ RalphLoopHealthScore, TimingHistory, RespawnPreset)
├── ralph.ts (RalphLoopStatus, RalphLoopState, RalphTodoStatus, RalphTodoPriority,
│ RalphTodoItem, RalphTodoProgress, RalphSessionState,
│ RalphStatusValue, RalphTestsStatus, RalphWorkType, RalphStatusBlock,
│ CompletionConfidence, RalphTrackerState,
│ CircuitBreakerState, CircuitBreakerReason, CircuitBreakerStatus,
│ createInitialCircuitBreakerStatus, createInitialRalphTrackerState,
│ createInitialRalphSessionState)
├── api.ts (ApiErrorCode, ApiResponse, HookEventType, QuickStartResponse,
│ CaseInfo, createErrorResponse, isError, getErrorMessage)
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
├── run-summary.ts (RunSummaryEventType, RunSummaryEventSeverity, RunSummaryEvent,
│ RunSummaryStats, RunSummary, createInitialRunSummaryStats)
├── tools.ts (ActiveBashToolStatus, ActiveBashTool, ImageDetectedEvent)
├── teams.ts (TeamConfig, TeamMember, TeamTask, InboxMessage, PaneInfo)
├── push.ts (PushSubscriptionRecord, VapidKeys)
└── plan.ts (PlanTaskStatus, TddPhase, PlanItem re-export, NiceConfig,
DEFAULT_NICE_CONFIG, ProcessStats)
```
### Steps
1. **Create `src/types/` directory** and each domain file above.
2. **Move types** from `src/types.ts` into their domain files. Preserve all JSDoc comments. Each file should import from siblings as needed (e.g., `ralph.ts` imports `CircuitBreakerState` within itself — no cross-file deps needed since they're in the same file).
3. **Create barrel `src/types/index.ts`** that re-exports everything:
```typescript
export * from './common.js';
export * from './session.js';
export * from './task.js';
export * from './app-state.js';
export * from './respawn.js';
export * from './ralph.js';
export * from './api.js';
export * from './lifecycle.js';
export * from './run-summary.js';
export * from './tools.js';
export * from './teams.js';
export * from './push.js';
export * from './plan.js';
```
4. **Delete old `src/types.ts`** and replace with a single-line re-export barrel:
```typescript
export * from './types/index.js';
```
This ensures `import { ... } from './types.js'` continues to work everywhere — zero changes to 36 import sites.
5. **Verify**: `tsc --noEmit` and `npm run lint` must pass. No runtime changes.
### Internal Dependencies Between Domain Files
Some types reference others across domains. Handle with imports:
| File | Imports From |
|------|-------------|
| `app-state.ts` | `session.ts` (SessionState), `task.ts` (TaskState), `ralph.ts` (RalphLoopState, RalphSessionState) |
| `respawn.ts` | None (self-contained) |
| `ralph.ts` | None (self-contained) |
| `run-summary.ts` | None (self-contained) |
| `api.ts` | None (self-contained) |
| `session.ts` | `respawn.ts` (RespawnConfig), `ralph.ts` (RalphTrackerState, RalphTodoItem, CircuitBreakerStatus, RalphSessionState, RunSummaryEvent) |
Wait — `SessionState` references `RespawnConfig`, `RalphTrackerState`, `CircuitBreakerStatus`, and `RunSummaryEvent`. This creates imports from `session.ts` → `respawn.ts`, `ralph.ts`, `run-summary.ts`. This is fine (one-way deps, no cycles).
---
## 2. Split ralph-tracker.ts into focused modules
**Current**: 3,868 lines, single `RalphTracker` class with 5 responsibilities.
**Risk**: MEDIUM — class has shared mutable state, but extractable modules are well-isolated.
### Coupling Analysis Summary
| Module | Coupling | Extractability |
|--------|----------|----------------|
| Plan task tracking | LOW | HIGH — only reads `cycleCount` |
| Fix-plan file watching | LOW | HIGH — callback-based todo replacement |
| Iteration stall detection | LOW | HIGH — notification-based |
| RALPH_STATUS block parsing + circuit breaker | MEDIUM | MEDIUM — callback for circuit breaker updates |
| Todo parsing, loop detection, completion | HIGH | LOW — deeply entangled shared state |
### Target Structure
```
src/
├── ralph-tracker.ts (~1,800 LOC — core: output parsing, loop state,
│ todo management, completion detection)
├── ralph-plan-tracker.ts (~600 LOC — plan tasks, checkpoints, history, rollback)
├── ralph-status-parser.ts (~300 LOC — RALPH_STATUS block parsing, circuit breaker)
├── ralph-fix-plan-watcher.ts (~150 LOC — @fix_plan.md file watching)
└── ralph-stall-detector.ts (~80 LOC — iteration stall detection)
```
### Step 2a: Extract `RalphPlanTracker` (~600 LOC)
**Why first**: Lowest coupling. Only dependency is `cycleCount` for checkpoint detection.
**Extract these from `RalphTracker`**:
Types to export:
- `EnhancedPlanTask` (interface, currently lines 56-87)
- `CheckpointReview` (interface, currently lines 90-139)
Properties to move:
- `_planVersion: number`
- `_planHistory: Array<{version, timestamp, tasks, summary}>`
- `_planTasks: Map<string, EnhancedPlanTask>`
- `_checkpointIterations: number[]`
- `_lastCheckpointIteration: number`
Methods to move:
- `initializePlanTasks(items)`
- `updatePlanTask(taskId, update)`
- `addPlanTask(params)`
- `getPlanTasks()`
- `generateCheckpointReview()`
- `getPlanHistory()`
- `rollbackToVersion(version)`
- `isCheckpointDue()`
- `planVersion` getter
- `_savePlanToHistory()` (private)
- `_unblockDependentTasks()` (private)
- `_checkForCheckpoint()` (private)
Events emitted (define in new class):
- `planInitialized`
- `planTaskUpdate`
- `taskBlocked`
- `taskUnblocked`
- `planCheckpoint`
**Interface with parent**:
```typescript
export class RalphPlanTracker extends EventEmitter {
constructor() { ... }
// Parent calls this when iteration changes (for checkpoint detection)
notifyCycleCount(cycleCount: number): void { ... }
// Full public API moves here unchanged
initializePlanTasks(items: PlanItem[]): void { ... }
updatePlanTask(taskId: string, update: { ... }): { ... } | null { ... }
// ...etc
}
```
**In `RalphTracker`**: Replace plan methods with delegation:
```typescript
readonly planTracker = new RalphPlanTracker();
// Forward plan events
this.planTracker.on('planInitialized', (...args) => this.emit('planInitialized', ...args));
// ...etc
// In detectLoopStatus(), when cycleCount changes:
this.planTracker.notifyCycleCount(this._loopState.cycleCount);
```
### Step 2b: Extract `RalphFixPlanWatcher` (~150 LOC)
**Extract these**:
Properties:
- `_workingDir: string | null`
- `_fixPlanPath: string | null`
- `_fixPlanWatcher: FSWatcher | null`
- `_fixPlanWatcherErrorHandler`
- `_fixPlanReloadDeb`
Methods:
- `setWorkingDir(workingDir)`
- `loadFixPlanFromDisk()`
- `startWatchingFixPlan()`
- `stopWatchingFixPlan()`
- `handleFixPlanChange()`
- `isFileAuthoritative` getter
**Interface with parent**:
```typescript
export class RalphFixPlanWatcher extends EventEmitter {
get isFileAuthoritative(): boolean { ... }
setWorkingDir(workingDir: string): void { ... }
stop(): void { ... }
}
// Events:
// 'todosLoaded' → (todos: Array<{id, content, status, priority}>) — parent replaces _todos
```
**In `RalphTracker`**:
```typescript
readonly fixPlanWatcher = new RalphFixPlanWatcher();
constructor() {
this.fixPlanWatcher.on('todosLoaded', (items) => {
// Replace _todos with file-based items
this._todos.clear();
for (const item of items) {
this.addOrUpdateTodo(item.id, item.content, item.status, item.priority);
}
});
}
// Delegate isFileAuthoritative
get isFileAuthoritative(): boolean {
return this.fixPlanWatcher.isFileAuthoritative;
}
```
### Step 2c: Extract `RalphStallDetector` (~80 LOC)
**Extract these**:
Properties:
- `_lastIterationChangeTime`
- `_lastObservedIteration`
- `_iterationStallTimerId`
- `_iterationStallWarningMs`
- `_iterationStallCriticalMs`
- `_iterationStallWarned`
Methods:
- `startIterationStallDetection()`
- `stopIterationStallDetection()`
- `checkIterationStall()`
- `getIterationStallMetrics()`
- `configureIterationStallThresholds(warningMs, criticalMs)`
**Interface with parent**:
```typescript
export class RalphStallDetector extends EventEmitter {
constructor(private cleanup: CleanupManager) { ... }
start(): void { ... }
stop(): void { ... }
// Parent calls when iteration changes
notifyIterationChanged(iteration: number): void {
this._lastIterationChangeTime = Date.now();
this._lastObservedIteration = iteration;
this._iterationStallWarned = false;
}
// Parent calls to check if loop is active
setLoopActive(active: boolean): void { ... }
getIterationStallMetrics(): { ... } { ... }
}
// Events: 'iterationStallWarning', 'iterationStallCritical'
```
### Step 2d: Extract `RalphStatusParser` (~300 LOC)
**Extract these**:
Properties:
- `_circuitBreaker: CircuitBreakerStatus`
- `_statusBlockBuffer: string[]`
- `_inStatusBlock: boolean`
- `_lastStatusBlock: RalphStatusBlock | null`
- `_completionIndicators: number`
- `_exitGateMet: boolean`
- `_totalFilesModified: number`
- `_totalTasksCompleted: number`
Methods:
- `processStatusBlockLine(line)`
- `parseStatusBlock(lines)`
- `detectCompletionIndicators(line)`
- `updateCircuitBreaker(hasProgress, testsStatus, status)`
- `resetCircuitBreaker()`
- `circuitBreakerStatus` getter
- `lastStatusBlock` getter
- `cumulativeStats` getter
- `exitGateMet` getter
Regex patterns to move:
- `RALPH_STATUS_START_PATTERN` through `RALPH_RECOMMENDATION_PATTERN`
- `COMPLETION_INDICATOR_PATTERNS`
**Interface with parent**:
```typescript
export class RalphStatusParser extends EventEmitter {
processLine(line: string): void { ... } // calls processStatusBlockLine + detectCompletionIndicators
get circuitBreakerStatus(): CircuitBreakerStatus { ... }
get lastStatusBlock(): RalphStatusBlock | null { ... }
get exitGateMet(): boolean { ... }
get cumulativeStats(): { ... } { ... }
resetCircuitBreaker(): void { ... }
reset(): void { ... }
}
// Events: 'statusBlockDetected', 'circuitBreakerUpdate', 'exitGateMet'
```
**In `RalphTracker.processLine()`**:
```typescript
// Replace inline status block handling with delegation
this.statusParser.processLine(line);
```
### Step 2e: Keep in `ralph-tracker.ts` (~1,800 LOC)
The core remains tightly coupled and stays together:
- Output parsing pipeline (`processTerminalData`, `processCleanData`, `processLine`)
- Loop state management (`_loopState`, `detectLoopStatus`, `enable/disable/startLoop/stopLoop`)
- Todo management (`_todos`, `detectTodoItems`, `addOrUpdateTodo`, `updateTodoStatus`, `getTodoStats`)
- Completion detection (`detectCompletionPhrase`, `handleCompletionPhrase`, `calculateCompletionConfidence`)
- All-tasks-complete detection (`detectAllTasksComplete`)
- Auto-enable logic (`shouldAutoEnable`)
- Lifecycle (`reset`, `fullReset`, `clear`, `restoreState`, `destroy`)
- Event debouncing and buffering
The class coordinates the extracted modules via composition:
```typescript
export class RalphTracker extends EventEmitter {
readonly planTracker = new RalphPlanTracker();
readonly fixPlanWatcher = new RalphFixPlanWatcher();
readonly stallDetector: RalphStallDetector;
readonly statusParser = new RalphStatusParser();
constructor() {
super();
this.stallDetector = new RalphStallDetector(this.cleanup);
this._wireSubModuleEvents();
}
private _wireSubModuleEvents(): void {
// Forward all sub-module events through RalphTracker
// so external consumers don't need to know about the split
for (const event of ['planInitialized', 'planTaskUpdate', ...]) {
this.planTracker.on(event, (...args) => this.emit(event, ...args));
}
// ...same for statusParser, stallDetector, fixPlanWatcher
}
}
```
### Migration Safety
- All events continue to be emitted from `RalphTracker` (forwarded from sub-modules)
- All public methods stay on `RalphTracker` (delegated to sub-modules)
- External consumers (`session.ts`, `case-routes.ts`) see zero API changes
- New sub-modules are exposed as `readonly` properties for direct access where needed
---
## 3. Split respawn-controller.ts into focused modules
**Current**: 3,611 lines, single `RespawnController` class with 6 responsibilities.
**Risk**: MEDIUM — health scoring and metrics are cleanly decoupled; detection is tightly coupled.
### Coupling Analysis Summary
| Module | Coupling | Extractability |
|--------|----------|----------------|
| Health scoring | NONE | HIGH — pure calculations from metrics |
| Cycle metrics | LOW | HIGH — standalone tracking |
| Adaptive timing | LOW | HIGH — standalone timing adjustments |
| Stuck-state detection | LOW | MEDIUM — needs state + config refs |
| Pattern detection utilities | NONE | HIGH — pure functions |
| State machine + idle detection + AI checkers | HIGH | LOW — deeply entangled |
### Target Structure
```
src/
├── respawn-controller.ts (~2,200 LOC — state machine, idle detection,
│ AI checkers, terminal handling, hook signals,
│ auto-accept, step execution)
├── respawn-health.ts (~250 LOC — health scoring + recommendations)
├── respawn-metrics.ts (~200 LOC — cycle metrics + aggregate stats)
├── respawn-adaptive-timing.ts (~100 LOC — adaptive timing with percentile calc)
└── respawn-patterns.ts (~50 LOC — terminal pattern detection utilities)
```
### Step 3a: Extract `RespawnPatterns` (~50 LOC)
**Pure utility functions, zero coupling**.
Move:
- `isCompletionMessage(data): boolean`
- `hasWorkingPattern(data, window): boolean`
- `extractTokenCount(data): number | null`
- `PROMPT_PATTERNS` array
- `WORKING_PATTERNS` array
```typescript
// src/respawn-patterns.ts
import { TOKEN_PATTERN, SPINNER_PATTERN } from './utils/index.js';
export const PROMPT_PATTERNS = ['❯', '>', '$', '%', '#'];
export const WORKING_PATTERNS = [/* 70+ patterns */];
export function isCompletionMessage(data: string): boolean { ... }
export function hasWorkingPattern(data: string, window: string): boolean { ... }
export function extractTokenCount(data: string): number | null { ... }
```
**In `RespawnController`**: Import and call:
```typescript
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
```
### Step 3b: Extract `RespawnAdaptiveTiming` (~100 LOC)
**Self-contained timing controller**.
Move properties:
- `timingHistory: TimingHistory`
Move methods:
- `recordTimingData(idleDetectionMs, cycleDurationMs)`
- `updateAdaptiveTiming()`
- `getTimingHistory()`
- `getAdaptiveCompletionConfirmMs()`
```typescript
export class RespawnAdaptiveTiming {
private timingHistory: TimingHistory;
constructor(private config: { adaptiveMinConfirmMs: number; adaptiveMaxConfirmMs: number }) {
this.timingHistory = { recentIdleDetectionMs: [], recentCycleDurationMs: [], ... };
}
recordTimingData(idleDetectionMs: number, cycleDurationMs: number): void { ... }
getAdaptiveCompletionConfirmMs(): number { ... }
getTimingHistory(): TimingHistory { ... }
reset(): void { ... }
}
```
### Step 3c: Extract `RespawnCycleMetrics` (~200 LOC)
**Standalone metrics tracker**.
Move properties:
- `currentCycleMetrics`
- `recentCycleMetrics[]`
- `aggregateMetrics`
- `MAX_CYCLE_METRICS_IN_MEMORY`
Move methods:
- `startCycleMetrics(idleReason)`
- `recordCycleStep(step)`
- `completeCycleMetrics(outcome, errorMessage?)`
- `updateAggregateMetrics(metrics)`
- `getAggregateMetrics()`
- `getRecentCycleMetrics(limit?)`
```typescript
export class RespawnCycleMetricsTracker {
private currentCycleMetrics: Partial<RespawnCycleMetrics> | null = null;
private recentCycleMetrics: RespawnCycleMetrics[] = [];
private aggregateMetrics: RespawnAggregateMetrics;
startCycle(sessionId: string, cycleNumber: number, idleReason: string): void { ... }
recordStep(step: string): void { ... }
completeCycle(outcome: CycleOutcome, errorMessage?: string): RespawnCycleMetrics | null { ... }
getAggregate(): RespawnAggregateMetrics { ... }
getRecent(limit?: number): RespawnCycleMetrics[] { ... }
reset(): void { ... }
}
```
**Callback**: `completeCycle()` returns the completed metrics so the controller can pass them to `adaptiveTiming.recordTimingData()`.
### Step 3d: Extract `RespawnHealthCalculator` (~250 LOC)
**Pure calculation — no state of its own**.
Move methods:
- `calculateHealthScore()`
- `calculateCycleSuccessScore()`
- `calculateCircuitBreakerScore()`
- `calculateIterationProgressScore()`
- `calculateAiCheckerScore()`
- `calculateStuckRecoveryScore()`
- `generateHealthRecommendations(components)`
- `generateHealthSummary(score, status, components)`
- `shouldSkipClear()` (belongs here since it's a pure calculation on token/config)
```typescript
export interface HealthInputs {
aggregateMetrics: RespawnAggregateMetrics;
circuitBreakerStatus: CircuitBreakerStatus;
iterationStallMetrics: { stallDurationMs: number; warningMs: number; criticalMs: number } | null;
aiCheckerState: { disabled: boolean; inCooldown: boolean; hasErrors: boolean };
stuckRecoveryCount: number;
maxStuckRecoveries: number;
}
export function calculateHealthScore(inputs: HealthInputs): RalphLoopHealthScore { ... }
export function shouldSkipClear(
lastTokenCount: number,
skipClearThresholdPercent: number,
maxContextTokens: number
): boolean { ... }
```
**Made as pure functions** (not a class) since they hold no state.
### Step 3e: Keep in `respawn-controller.ts` (~2,200 LOC)
The core state machine, idle detection, and AI checker integration stays:
- State machine transitions (`setState`, `start`, `stop`, `pause`, `resume`)
- Terminal data handling (`handleTerminalData`)
- All 5 idle detection layers + hook signals
- AI checker integration (`tryStartAiCheck`, `startAiCheck`, `startPlanCheck`)
- Auto-accept logic
- Step execution (`sendUpdateDocs`, `sendClear`, `sendInit`, `sendKickstart`)
- Timer management (`startTrackedTimer`, `cancelTrackedTimer`)
- Stuck-state detection and recovery
- Action logging
The class composes extracted modules:
```typescript
import { RespawnAdaptiveTiming } from './respawn-adaptive-timing.js';
import { RespawnCycleMetricsTracker } from './respawn-metrics.js';
import { calculateHealthScore, shouldSkipClear } from './respawn-health.js';
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
export class RespawnController extends EventEmitter {
private adaptiveTiming: RespawnAdaptiveTiming;
private cycleMetrics: RespawnCycleMetricsTracker;
calculateHealthScore(): RalphLoopHealthScore {
return calculateHealthScore({
aggregateMetrics: this.cycleMetrics.getAggregate(),
circuitBreakerStatus: this.session.ralphTracker.circuitBreakerStatus,
iterationStallMetrics: this.session.ralphTracker.getIterationStallMetrics(),
aiCheckerState: { ... },
stuckRecoveryCount: this.stuckRecoveryCount,
maxStuckRecoveries: this.config.maxStuckRecoveries ?? 3,
});
}
}
```
---
## 4. Split session.ts into focused modules
**Current**: 2,418 lines, single `Session` class.
**Risk**: LOW-MEDIUM — extractable pieces are utility-like with clear boundaries.
### Coupling Analysis Summary
| Module | Coupling | Extractability |
|--------|----------|----------------|
| CLI arg builder | NONE | HIGH — pure functions used at spawn time |
| Auto-compact/clear | LOW | HIGH — self-contained automation with config |
| Token tracking | LOW | MEDIUM — reads PTY output, writes state |
| Task description cache | LOW | HIGH — separate LRU cache |
| PTY + mux lifecycle | HIGH | KEEP — core of the class |
| Tracker integration | HIGH | KEEP — event forwarding plumbing |
### Target Structure
```
src/
├── session.ts (~1,600 LOC — PTY lifecycle, terminal I/O,
│ tracker integration, output processing,
│ token tracking, state management)
├── session-cli-builder.ts (~250 LOC — Claude/OpenCode CLI arg construction)
├── session-auto-ops.ts (~300 LOC — auto-compact, auto-clear automation)
└── session-task-cache.ts (~100 LOC — task description LRU cache)
```
### Step 4a: Extract `SessionCliBuilder` (~250 LOC)
**Pure functions — zero coupling to Session instance**.
Move:
- `buildClaudeArgs()` logic (currently inlined in `startInteractive` and `runPrompt`)
- `buildOpenCodeArgs()` logic
- Model mapping constants
- Claude mode to flag mapping
- Environment variable construction
```typescript
// src/session-cli-builder.ts
export interface CliBuilderConfig {
claudeMode: ClaudeMode;
model?: string;
workingDir: string;
sessionId: string;
niceConfig?: NiceConfig;
isOpenCode?: boolean;
openCodeConfig?: OpenCodeConfig;
}
export function buildInteractiveArgs(config: CliBuilderConfig): string[] { ... }
export function buildPromptArgs(config: CliBuilderConfig, prompt: string): string[] { ... }
export function buildShellArgs(shell?: string): string[] { ... }
export function buildClaudeEnv(config: CliBuilderConfig): Record<string, string> { ... }
```
### Step 4b: Extract `SessionAutoOps` (~300 LOC)
**Self-contained automation with config-based thresholds**.
Move properties:
- `_autoCompactThreshold`
- `_autoClearThreshold`
- `_isAutoCompacting`
- `_isAutoClearing`
- `_autoCompactCount`
- `_autoClearCount`
- `_lastAutoCompactTime`
- `_lastAutoClearTime`
Move methods:
- `checkAutoCompact(tokenCount)`
- `performAutoCompact()`
- `checkAutoClear(tokenCount)`
- `performAutoClear()`
- Auto-compact/clear threshold configuration
```typescript
export class SessionAutoOps extends EventEmitter {
constructor(
private writeCommand: (command: string) => Promise<void>,
private getTokenCount: () => number,
config: { compactThreshold: number; clearThreshold: number }
) { ... }
/** Called after token count updates. Checks thresholds and triggers if needed. */
checkThresholds(tokenCount: number): void { ... }
updateConfig(config: { compactThreshold?: number; clearThreshold?: number }): void { ... }
getStats(): { autoCompactCount: number; autoClearCount: number; ... } { ... }
}
// Events: 'autoCompact', 'autoClear'
```
**In `Session`**: Compose and wire:
```typescript
private autoOps = new SessionAutoOps(
(cmd) => this.writeViaMux(cmd),
() => this._state.tokenCount,
{ compactThreshold: 110_000, clearThreshold: 140_000 }
);
```
### Step 4c: Extract `SessionTaskCache` (~100 LOC)
**Isolated LRU cache for task descriptions**.
Move:
- `_taskDescriptionCache: LRUMap<number, { description: string; timestamp: number }>`
- `_taskDescriptionMaxAge`
- `findTaskDescriptionNear(lineNumber)`
- `cacheTaskDescription(lineNumber, description)`
```typescript
export class SessionTaskCache {
private cache: LRUMap<number, { description: string; timestamp: number }>;
private maxAgeMs: number;
constructor(maxSize: number = 50, maxAgeMs: number = 30_000) { ... }
find(lineNumber: number, searchRadius: number = 50): string | null { ... }
add(lineNumber: number, description: string): void { ... }
clear(): void { ... }
}
```
### Step 4d: Keep in `session.ts` (~1,600 LOC)
The core stays together:
- PTY process management (`spawn`, `kill`, `resize`, `writeViaMux`)
- Data streaming pipeline (PTY → buffer → ANSI strip → JSON parse → events)
- Tracker initialization and event forwarding (RalphTracker, BashToolParser, TaskTracker)
- Output processing (message extraction, completion detection)
- Token tracking (status line parsing)
- State management (`toState()`, `updateState()`)
- Session lifecycle (`startInteractive`, `startShell`, `runPrompt`)
- CLI info detection (version, model, account)
---
## 5. Execution Order & Dependencies
Execute in this order to minimize risk. Each step is independently deployable.
```
Step 1: types.ts split
↓ (no runtime change, just file reorganization)
Step 2a: RalphPlanTracker extraction
↓ (independent of types split)
Step 2b: RalphFixPlanWatcher extraction
Step 2c: RalphStallDetector extraction
Step 2d: RalphStatusParser extraction
↓ (ralph-tracker.ts now ~1,800 LOC)
Step 3a: RespawnPatterns extraction
Step 3b: RespawnAdaptiveTiming extraction
Step 3c: RespawnCycleMetrics extraction
Step 3d: RespawnHealthCalculator extraction
↓ (respawn-controller.ts now ~2,200 LOC)
Step 4a: SessionCliBuilder extraction
Step 4b: SessionAutoOps extraction
Step 4c: SessionTaskCache extraction
↓ (session.ts now ~1,600 LOC)
```
**Parallelization**: Steps 1, 2a-2d, 3a-3d, and 4a-4c can be done by separate agents in parallel since they touch different files. However, within each group, sequential execution is safer.
### Risk Mitigation
- **Barrel exports**: Every split uses delegation + barrel re-export so external consumers see zero API changes
- **Event forwarding**: Sub-modules emit events, parent class forwards them — no event contract changes
- **Incremental**: Each step can be verified independently with `tsc --noEmit` + `npm run lint`
- **No test changes needed**: External API stays identical; existing tests continue to pass
---
## 6. Validation Checklist
After each step, verify:
- [ ] `tsc --noEmit` passes (no type errors)
- [ ] `npm run lint` passes (no unused imports, etc.)
- [ ] `npm run format:check` passes
- [ ] `npx vitest run test/respawn-controller.test.ts` passes (for respawn splits)
- [ ] `npx vitest run test/ralph-tracker.test.ts` passes (for ralph splits)
- [ ] `npx vitest run test/session-manager.test.ts` passes (for session splits)
- [ ] Dev server starts: `npx tsx src/index.ts web`
- [ ] Existing sessions work (create, interact, delete)
- [ ] Respawn cycle works (enable respawn, verify idle detection fires)
- [ ] No new circular dependencies: `npx madge --circular src/`
### Size Targets
| File | Before | After |
|------|--------|-------|
| `src/types.ts` | 1,443 LOC | 1 LOC (re-export barrel) |
| `src/ralph-tracker.ts` | 3,868 LOC | ~1,800 LOC |
| `src/respawn-controller.ts` | 3,611 LOC | ~2,200 LOC |
| `src/session.ts` | 2,418 LOC | ~1,600 LOC |
| **Total new files** | — | 12 files |
| **Net LOC change** | — | ~0 (refactor only) |
-738
View File
@@ -1,738 +0,0 @@
# Phase 1 Implementation Plan: Quick Wins
**Source**: `docs/code-structure-findings.md` (Phase 1 - Quick Wins section)
**Estimated effort**: 1-2 days
**Tasks**: 5 independent tasks (can be done in parallel unless noted)
---
## Safety Constraints
Before starting ANY work, read and follow these rules:
1. **Never run `npx vitest run`** (full suite) -- it kills tmux sessions. You are running inside a Codeman-managed tmux session.
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
3. **Never test on port 3000** -- the live dev server runs there. Tests use ports 3150+.
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
6. **Never kill tmux sessions** -- check `echo $CODEMAN_MUX` first.
---
## Task Dependencies
All 5 tasks are independent and can be done in parallel. However:
- Task 1 (barrel exports) is a prerequisite if you want to update import sites to use the barrel after Task 3 (consolidate EXEC_TIMEOUT_MS). The EXEC_TIMEOUT_MS consolidation creates a new export that should be added to the barrel.
- Task 2 (delete dead functions) removes functions that Task 1 would otherwise need to add to the barrel. Do Task 2 first or simultaneously with Task 1 to avoid adding exports for dead code.
**Recommended order**: Task 2 -> Task 1 -> Task 3 -> Task 4 -> Task 5
---
## Task 1: Export Missing Functions from Utils Barrel
**File**: `src/utils/index.ts`
**Time**: ~30 minutes
### Problem
The barrel file (`src/utils/index.ts`) is missing exports for several functions that are defined in util modules, forcing consumers to use deep imports or preventing usage entirely.
### Missing Exports
From `src/utils/regex-patterns.ts`:
- `createAnsiPatternFull()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
- `createAnsiPatternSimple()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
- `stripAnsi()` -- ANSI stripping utility
- `SAFE_PATH_PATTERN` -- regex for safe file paths (currently deep-imported by `schemas.ts` and `tmux-manager.ts`)
From `src/utils/token-validation.ts`:
- `validateTokenCounts()` -- token count validation (documented in CLAUDE.md)
- `validateTokensAndCost()` -- token + cost validation (documented in CLAUDE.md)
**Note**: Do NOT export `isSimilar`, `isSimilarByDistance`, `levenshteinDistance`, or `normalizePhrase` from `string-similarity.ts` -- these are dead code (see Task 2).
### Edit 1: Add missing regex-patterns exports
**File**: `src/utils/index.ts`
**Old code** (lines 13-18):
```typescript
export {
ANSI_ESCAPE_PATTERN_FULL,
ANSI_ESCAPE_PATTERN_SIMPLE,
TOKEN_PATTERN,
SPINNER_PATTERN,
} from './regex-patterns.js';
```
**New code**:
```typescript
export {
ANSI_ESCAPE_PATTERN_FULL,
ANSI_ESCAPE_PATTERN_SIMPLE,
TOKEN_PATTERN,
SPINNER_PATTERN,
createAnsiPatternFull,
createAnsiPatternSimple,
stripAnsi,
SAFE_PATH_PATTERN,
} from './regex-patterns.js';
```
### Edit 2: Add missing token-validation exports
**File**: `src/utils/index.ts`
**Old code** (line 19):
```typescript
export { MAX_SESSION_TOKENS } from './token-validation.js';
```
**New code**:
```typescript
export { MAX_SESSION_TOKENS, validateTokenCounts, validateTokensAndCost } from './token-validation.js';
```
### Optional follow-up: Update deep imports to use barrel
These files currently deep-import `SAFE_PATH_PATTERN` and could be updated to use the barrel instead:
- `src/web/schemas.ts` line 11: `import { SAFE_PATH_PATTERN } from '../utils/regex-patterns.js';` could become `import { SAFE_PATH_PATTERN } from '../utils/index.js';`
- `src/tmux-manager.ts` line 44: `import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';` could become part of existing barrel import
This is a low-priority cosmetic change. The barrel export itself is the important fix.
### Verification
```bash
tsc --noEmit
npm run lint
```
---
## Task 2: Delete Dead Utility Functions
**File**: `src/utils/string-similarity.ts`
**Time**: ~15 minutes
### Problem
Four exported functions in `string-similarity.ts` are never imported anywhere in the codebase:
- `levenshteinDistance()` (lines 27-69)
- `isSimilar()` (lines 106-108)
- `isSimilarByDistance()` (lines 123-125)
- `normalizePhrase()` (lines 139-144)
Only three functions are actually used (all by `ralph-tracker.ts` via the barrel):
- `stringSimilarity()` -- uses `levenshteinDistance()` internally
- `fuzzyPhraseMatch()` -- uses `normalizePhrase()` and `isSimilarByDistance()` internally
- `todoContentHash()`
### Strategy
`levenshteinDistance()` is called by `stringSimilarity()`, and `normalizePhrase()` and `isSimilarByDistance()` are called by `fuzzyPhraseMatch()`. So they cannot be deleted -- they just need to be un-exported (made private to the module).
`isSimilar()` is truly dead -- not called by anything. Delete it entirely.
### Edit 1: Remove `export` from `levenshteinDistance`
**File**: `src/utils/string-similarity.ts`
**Old code** (line 27):
```typescript
export function levenshteinDistance(a: string, b: string): number {
```
**New code**:
```typescript
function levenshteinDistance(a: string, b: string): number {
```
### Edit 2: Delete `isSimilar` function entirely
**File**: `src/utils/string-similarity.ts`
**Old code** (lines 94-108):
```typescript
/**
* Check if two strings are similar within a given threshold.
*
* @param a - First string
* @param b - Second string
* @param threshold - Minimum similarity ratio (default: 0.85 = 85% similar)
* @returns True if similarity >= threshold
*
* @example
* isSimilar('COMPLETE', 'COMPLET', 0.85) // true (87.5% similar)
* isSimilar('COMPLETE', 'DONE', 0.85) // false (0% similar)
*/
export function isSimilar(a: string, b: string, threshold = 0.85): boolean {
return stringSimilarity(a, b) >= threshold;
}
```
**New code**: (delete entirely -- replace with empty string)
### Edit 3: Remove `export` from `isSimilarByDistance`
**File**: `src/utils/string-similarity.ts`
**Old code** (line 123):
```typescript
export function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
```
**New code**:
```typescript
function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
```
### Edit 4: Remove `export` from `normalizePhrase`
**File**: `src/utils/string-similarity.ts`
**Old code** (line 139):
```typescript
export function normalizePhrase(phrase: string): string {
```
**New code**:
```typescript
function normalizePhrase(phrase: string): string {
```
### Verification
```bash
tsc --noEmit
npx vitest run test/string-utilities.test.ts
npm run lint
```
Note: If `test/string-utilities.test.ts` imports any of the now-unexported functions, those test imports will fail. Check the test file and remove tests for `isSimilar` (deleted) and update any direct tests for `levenshteinDistance`, `isSimilarByDistance`, `normalizePhrase` to test them indirectly through the public API (`stringSimilarity`, `fuzzyPhraseMatch`), or remove those tests.
---
## Task 3: Consolidate Duplicated `EXEC_TIMEOUT_MS` Constant
**Files**:
- `src/utils/claude-cli-resolver.ts` (line 17)
- `src/utils/opencode-cli-resolver.ts` (line 16)
- `src/tmux-manager.ts` (line 63) -- also has its own copy
**Time**: ~15 minutes
### Problem
`EXEC_TIMEOUT_MS = 5000` is defined identically in three files. Changes need to happen in all three places.
### Strategy
Create a shared constant and export it. The natural home is a new config file since the existing config files (`buffer-limits.ts`, `map-limits.ts`) follow this pattern. However, to keep it minimal, we can add it to an existing config file or create a small one.
**Recommended approach**: Add to `src/config/timing-config.ts` (new file) as a single constant. This file can grow later in Phase 6 to hold other timing constants.
Alternatively, the simplest approach: export from one of the existing utils and import in the others. Since both CLI resolvers are in `src/utils/`, the cleanest approach is to put it in a shared location.
### Option A: Add to existing config (simpler)
Create `src/config/exec-timeout.ts`:
**New file**: `src/config/exec-timeout.ts`
```typescript
/**
* Timeout for child process exec commands (e.g., `which claude`, `which opencode`, tmux commands).
* Used across CLI resolvers and tmux manager.
*/
export const EXEC_TIMEOUT_MS = 5000;
```
### Edit 1: Update `claude-cli-resolver.ts`
**File**: `src/utils/claude-cli-resolver.ts`
**Old code** (lines 11-17):
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { delimiter, dirname, join } from 'node:path';
import { homedir } from 'node:os';
/** Timeout for exec commands (5 seconds) */
const EXEC_TIMEOUT_MS = 5000;
```
**New code**:
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { delimiter, dirname, join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
```
### Edit 2: Update `opencode-cli-resolver.ts`
**File**: `src/utils/opencode-cli-resolver.ts`
**Old code** (lines 10-16):
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { homedir } from 'node:os';
/** Timeout for exec commands (5 seconds) */
const EXEC_TIMEOUT_MS = 5000;
```
**New code**:
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
```
### Edit 3: Update `tmux-manager.ts`
**File**: `src/tmux-manager.ts`
**Old code** (line 63):
```typescript
const EXEC_TIMEOUT_MS = 5000;
```
**New code**:
```typescript
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
```
Note: `tmux-manager.ts` already has many imports at the top of the file. Add this import near the other local imports (around lines 43-56). The `const EXEC_TIMEOUT_MS = 5000;` on line 63 should be deleted entirely (replaced with the import).
### Verification
```bash
tsc --noEmit
npm run lint
```
---
## Task 4: Add `z.infer` to Zod Schemas
**Files**:
- `src/web/schemas.ts` (add type exports)
- `src/types.ts` (replace manual interfaces with `z.infer` re-exports where applicable)
**Time**: ~2 hours
### Problem
All 30+ Zod schemas in `schemas.ts` define validation rules, but zero use `z.infer` to derive TypeScript types. Instead, `types.ts` manually duplicates interfaces that match the schemas. When a schema changes, the type must be manually updated too.
### Strategy
Add `z.infer` type exports to `schemas.ts` for each exported schema. This creates derived types as the single source of truth. For schemas that have corresponding manual interfaces in `types.ts`, the manual interface can be replaced with a re-export of the inferred type.
**Important**: Not all schemas have matching interfaces in `types.ts`. The `RespawnConfig` interface in `types.ts` (line 395) has all required fields, while `RespawnConfigSchema` has all optional fields (it's for partial updates). These are NOT the same type and should NOT be unified.
### Edit 1: Add inferred type exports to `schemas.ts`
**File**: `src/web/schemas.ts`
After each schema definition, add a corresponding type export. Add the following lines at the **end of the file** (after line 509):
**Old code** (end of file, lines 506-509):
```typescript
.optional(),
});
```
Wait -- the end of file is actually at line 509 after the `RalphLoopStartSchema`. Add the type exports after the last schema:
**Append to end of file** `src/web/schemas.ts`:
```typescript
// ========== Inferred Types ==========
// Derive TypeScript types from Zod schemas (single source of truth)
export type CreateSessionInput = z.infer<typeof CreateSessionSchema>;
export type RunPromptInput = z.infer<typeof RunPromptSchema>;
export type ResizeInput = z.infer<typeof ResizeSchema>;
export type CreateCaseInput = z.infer<typeof CreateCaseSchema>;
export type QuickStartInput = z.infer<typeof QuickStartSchema>;
export type HookEventInput = z.infer<typeof HookEventSchema>;
export type RespawnConfigInput = z.infer<typeof RespawnConfigSchema>;
export type ConfigUpdateInput = z.infer<typeof ConfigUpdateSchema>;
export type SettingsUpdateInput = z.infer<typeof SettingsUpdateSchema>;
export type SessionInputWithLimitInput = z.infer<typeof SessionInputWithLimitSchema>;
export type SessionNameInput = z.infer<typeof SessionNameSchema>;
export type SessionColorInput = z.infer<typeof SessionColorSchema>;
export type RalphConfigInput = z.infer<typeof RalphConfigSchema>;
export type FixPlanImportInput = z.infer<typeof FixPlanImportSchema>;
export type RalphPromptWriteInput = z.infer<typeof RalphPromptWriteSchema>;
export type AutoClearInput = z.infer<typeof AutoClearSchema>;
export type AutoCompactInput = z.infer<typeof AutoCompactSchema>;
export type ImageWatcherInput = z.infer<typeof ImageWatcherSchema>;
export type FlickerFilterInput = z.infer<typeof FlickerFilterSchema>;
export type QuickRunInput = z.infer<typeof QuickRunSchema>;
export type ScheduledRunInput = z.infer<typeof ScheduledRunSchema>;
export type LinkCaseInput = z.infer<typeof LinkCaseSchema>;
export type GeneratePlanInput = z.infer<typeof GeneratePlanSchema>;
export type GeneratePlanDetailedInput = z.infer<typeof GeneratePlanDetailedSchema>;
export type CancelPlanInput = z.infer<typeof CancelPlanSchema>;
export type PlanTaskUpdateInput = z.infer<typeof PlanTaskUpdateSchema>;
export type PlanTaskAddInput = z.infer<typeof PlanTaskAddSchema>;
export type CpuLimitInput = z.infer<typeof CpuLimitSchema>;
export type SubagentWindowStatesInput = z.infer<typeof SubagentWindowStatesSchema>;
export type SubagentParentMapInput = z.infer<typeof SubagentParentMapSchema>;
export type InteractiveRespawnInput = z.infer<typeof InteractiveRespawnSchema>;
export type RespawnEnableInput = z.infer<typeof RespawnEnableSchema>;
export type PushSubscribeInput = z.infer<typeof PushSubscribeSchema>;
export type PushPreferencesUpdateInput = z.infer<typeof PushPreferencesUpdateSchema>;
export type RalphLoopStartInput = z.infer<typeof RalphLoopStartSchema>;
```
### What NOT to do
Do NOT replace the `RespawnConfig` interface in `types.ts` with `z.infer<typeof RespawnConfigSchema>`. The schema has all optional fields (for partial config updates), but the interface has required fields (for the full config object). These are intentionally different shapes.
Similarly, do NOT try to unify every interface in `types.ts` with a schema -- most interfaces in `types.ts` represent internal domain objects (SessionState, TaskState, etc.) that have no corresponding Zod schema. The schemas only exist for API request validation.
### Future opportunity
In a future phase, route handlers in `server.ts` can use these inferred types for request body typing:
```typescript
const body = CreateSessionSchema.parse(request.body) as CreateSessionInput;
```
This task only adds the type exports. Migrating route handlers to use them is out of scope.
### Verification
```bash
tsc --noEmit
npm run lint
npm run format:check
```
---
## Task 5: Fix Weak `not.toThrow()` Tests with Behavioral Assertions
**Files**:
- `test/task-tracker.test.ts` -- 6 instances
- `test/image-watcher.test.ts` -- 1 instance
- `test/task-queue.test.ts` -- 1 instance
- `test/hooks-config.test.ts` -- 1 instance
- `test/session-manager.test.ts` -- 1 instance
**Time**: ~1 hour
### Problem
10 tests only assert `not.toThrow()` without verifying the actual defensive behavior. These tests prove the code doesn't crash but don't verify it does the right thing.
### Fix Strategy
After each `not.toThrow()`, add a behavioral assertion that verifies the state is correct (e.g., no tasks were created, no side effects occurred).
### Edit 1: `task-tracker.test.ts` -- null message (line 566)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle null message', () => {
expect(() => tracker.processMessage(null)).not.toThrow();
});
```
**New code**:
```typescript
it('should handle null message', () => {
expect(() => tracker.processMessage(null)).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
expect(tracker.getRunningCount()).toBe(0);
});
```
### Edit 2: `task-tracker.test.ts` -- message without content (line 569-571)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle message without content', () => {
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
});
```
**New code**:
```typescript
it('should handle message without content', () => {
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
});
```
### Edit 3: `task-tracker.test.ts` -- empty content array (line 573-575)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle empty content array', () => {
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
});
```
**New code**:
```typescript
it('should handle empty content array', () => {
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
});
```
### Edit 4: `task-tracker.test.ts` -- tool_result for unknown task (lines 577-590)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle tool_result for unknown task', () => {
expect(() => {
tracker.processMessage({
message: {
content: [{
type: 'tool_result',
tool_use_id: 'unknown-task',
is_error: false,
content: 'Done',
}],
},
});
}).not.toThrow();
});
```
**New code**:
```typescript
it('should handle tool_result for unknown task', () => {
expect(() => {
tracker.processMessage({
message: {
content: [{
type: 'tool_result',
tool_use_id: 'unknown-task',
is_error: false,
content: 'Done',
}],
},
});
}).not.toThrow();
expect(tracker.getTask('unknown-task')).toBeUndefined();
expect(tracker.getAllTasks().size).toBe(0);
});
```
### Edit 5: `task-tracker.test.ts` -- empty terminal output (lines 592-595)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle empty terminal output', () => {
expect(() => tracker.processTerminalOutput('')).not.toThrow();
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
});
```
**New code**:
```typescript
it('should handle empty terminal output', () => {
expect(() => tracker.processTerminalOutput('')).not.toThrow();
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
expect(tracker.getRunningCount()).toBe(0);
});
```
### Edit 6: `image-watcher.test.ts` -- unwatchSession for non-watched session (line 123)
**File**: `test/image-watcher.test.ts`
**Old code**:
```typescript
it('should be safe to call for non-watched session', () => {
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
});
```
**New code**:
```typescript
it('should be safe to call for non-watched session', () => {
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
expect(watcher.getWatchedSessions()).toHaveLength(0);
});
```
### Edit 7: `task-queue.test.ts` -- dependencies on non-existent tasks (lines 538-542)
**File**: `test/task-queue.test.ts`
**Old code**:
```typescript
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
// Dependencies on non-existent tasks are valid - they just won't be satisfied
expect(() => {
queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
}).not.toThrow();
});
```
**New code**:
```typescript
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
// Dependencies on non-existent tasks are valid - they just won't be satisfied
let task: ReturnType<typeof queue.addTask> | undefined;
expect(() => {
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
}).not.toThrow();
expect(task).toBeDefined();
expect(task!.dependencies).toEqual(['non-existent-id']);
// Task should be pending but blocked (dependency unsatisfied)
expect(queue.next()?.prompt).toBeUndefined();
});
```
Wait -- `queue.next()` returns `null` when no next task is available (all blocked). Let me adjust:
**New code** (corrected):
```typescript
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
// Dependencies on non-existent tasks are valid - they just won't be satisfied
let task: ReturnType<typeof queue.addTask> | undefined;
expect(() => {
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
}).not.toThrow();
expect(task).toBeDefined();
expect(task!.dependencies).toEqual(['non-existent-id']);
// Task exists but is blocked (dependency unsatisfied), so next() skips it
expect(queue.getAllTasks()).toHaveLength(1);
expect(queue.next()).toBeNull();
});
```
### Edit 8: `hooks-config.test.ts` -- valid JSON check (line 129)
**File**: `test/hooks-config.test.ts`
**Old code**:
```typescript
it('should write valid JSON', () => {
writeHooksConfig(testDir);
const settingsPath = join(testDir, '.claude', 'settings.local.json');
const content = readFileSync(settingsPath, 'utf-8');
expect(() => JSON.parse(content)).not.toThrow();
});
```
**New code**:
```typescript
it('should write valid JSON', () => {
writeHooksConfig(testDir);
const settingsPath = join(testDir, '.claude', 'settings.local.json');
const content = readFileSync(settingsPath, 'utf-8');
const parsed = JSON.parse(content);
expect(parsed).toBeDefined();
expect(typeof parsed).toBe('object');
expect(parsed.hooks).toBeDefined();
});
```
### Edit 9: `session-manager.test.ts` -- stopSession for non-existent (line 216)
**File**: `test/session-manager.test.ts`
**Old code**:
```typescript
it('should handle non-existent session gracefully', async () => {
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
});
```
**New code**:
```typescript
it('should handle non-existent session gracefully', async () => {
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
expect(manager.getSessionCount()).toBe(0);
});
```
### Verification
Run each test file individually:
```bash
npx vitest run test/task-tracker.test.ts
npx vitest run test/image-watcher.test.ts
npx vitest run test/task-queue.test.ts
npx vitest run test/hooks-config.test.ts
npx vitest run test/session-manager.test.ts
```
**Important**: `hooks-config.test.ts` and `session-manager.test.ts` spawn real servers on ports 3130-3131. Only run them if you are NOT running other tests that use those ports.
---
## Final Verification Checklist
After all 5 tasks are complete, run the following in order:
```bash
# 1. TypeScript type checking
tsc --noEmit
# 2. Linting
npm run lint
# 3. Formatting
npm run format:check
# 4. Run affected test files individually (NOT the full suite)
npx vitest run test/string-utilities.test.ts
npx vitest run test/task-tracker.test.ts
npx vitest run test/image-watcher.test.ts
npx vitest run test/task-queue.test.ts
npx vitest run test/session-manager.test.ts
npx vitest run test/hooks-config.test.ts
```
If any formatting issues arise, fix with:
```bash
npm run format
```
If any lint issues arise, fix with:
```bash
npm run lint:fix
```
### Summary of Changes
| Task | Files Modified | Files Created |
|------|---------------|---------------|
| 1. Barrel exports | `src/utils/index.ts` | -- |
| 2. Dead functions | `src/utils/string-similarity.ts` | -- |
| 3. EXEC_TIMEOUT_MS | `src/utils/claude-cli-resolver.ts`, `src/utils/opencode-cli-resolver.ts`, `src/tmux-manager.ts` | `src/config/exec-timeout.ts` |
| 4. z.infer types | `src/web/schemas.ts` | -- |
| 5. Weak tests | `test/task-tracker.test.ts`, `test/image-watcher.test.ts`, `test/task-queue.test.ts`, `test/hooks-config.test.ts`, `test/session-manager.test.ts` | -- |
**Total files modified**: 10
**Total files created**: 1
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -1,689 +0,0 @@
# Phase 6 Implementation Plan: Config Consolidation
**Source**: `docs/code-structure-findings.md` (Phase 6 — Config Consolidation)
**Estimated effort**: 1 day
**Tasks**: 8 tasks with dependencies (see dependency graph below)
---
## Safety Constraints
Before starting ANY work, read and follow these rules:
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
7. **Verify the dev server starts**: After each task, run `npx tsx src/index.ts web --port 3099 &` on a non-production port, confirm `curl -s http://localhost:3099/api/status | jq .status` returns `"ok"`, then kill the background process.
---
## Goal
Consolidate ~70 scattered numeric constants from 15+ source files into 6 new domain-focused config files, eliminating cross-file duplicates (including a 5x-duplicated AI model string) and making all tuning knobs discoverable in `src/config/`.
**Non-goal**: Moving every constant. Module-internal implementation details (like regex patterns, algorithm-specific magic numbers, or constants only used once in deeply coupled logic) stay where they are. The goal is discoverability of operational tuning knobs, not mechanical relocation.
---
## Design Decisions
### What gets centralized (and why)
Constants are candidates for centralization when they meet **any** of these criteria:
1. **Duplicated across files** — DRY violation (e.g., `STATS_COLLECTION_INTERVAL_MS` in `server.ts` and `mux-routes.ts`, AI model string in 5 files)
2. **Operational tuning knobs** — values an operator might want to adjust for performance, security, or behavior without understanding the implementation (e.g., SSE health check interval, auth session TTL, rate limits)
3. **Cross-cutting concerns** — values that establish system-wide contracts (e.g., max terminal dimensions used by both server routes and frontend)
### What stays in place (and why)
Constants that are **internal implementation details** of a single module stay where they are:
- **Algorithm parameters** — `TODO_SIMILARITY_THRESHOLD`, `adaptiveCompletionConfirmMs`, confidence weights. These are meaningless without understanding the algorithm.
- **Display/UI formatting** — `TEXT_PREVIEW_LENGTH`, `SMART_TITLE_MAX_LENGTH`, `COMMAND_DISPLAY_LENGTH` in `subagent-watcher.ts`. Only used locally, tightly coupled to rendering logic.
- **Module-internal timing** — `LINE_BUFFER_FLUSH_INTERVAL` in `session.ts`, `AI_CHECK_POLL_INTERVAL` in `ai-checker-base.ts`. Internal implementation of specific features.
- **Frontend constants** — `constants.js` already centralizes frontend values well. Don't mix frontend and backend config.
- **Respawn `DEFAULT_CONFIG`** — these are user-configurable defaults for the respawn config interface, not system constants. They live properly in `respawn-controller.ts`. The AI model/context defaults within it are replaced with imports from the new `ai-defaults.ts` (Task 5).
- **Session auto-ops thresholds** — `AUTO_RETRY_DELAY_MS`, `COMPACT_COOLDOWN_MS`, etc. in `session-auto-ops.ts` are internal to that module's retry logic and already well-documented in place.
### File organization: domain-based, not category-based
A single `timing-config.ts` with 70 unrelated timing values would be worse than the current state — developers would need to grep it just like they grep the whole codebase now. Instead, constants are grouped by **the system they configure**:
| New File | Domain | Developer Question It Answers |
|----------|--------|-------------------------------|
| `server-timing.ts` | Web server performance | "How do I tune SSE batching / terminal throughput?" |
| `auth-config.ts` | Authentication & security | "What are the rate limits and session TTLs?" |
| `tunnel-config.ts` | QR auth & Cloudflare tunnel | "What are the QR token rotation parameters?" |
| `terminal-limits.ts` | Terminal dimensions & input | "What are the max cols/rows/input size?" |
| `ai-defaults.ts` | AI checker model & context | "What model do the AI checkers use? What's the context limit?" |
| `team-config.ts` | Agent Teams polling & caching | "How often does team polling run? What are the cache limits?" |
---
## Task Dependencies
```
Task 1 (server-timing.ts)
Task 2 (auth-config.ts)
Task 3 (tunnel-config.ts)
Task 4 (terminal-limits.ts)
Task 5 (ai-defaults.ts)
Task 6 (team-config.ts)
└──> Task 7 (Fix remaining duplicates)
└──> Task 8 (Update CLAUDE.md + final verification)
```
**Tasks 1–6** are independent and can run in parallel.
**Task 7** depends on Tasks 1–6 (needs the new config files to exist).
**Task 8** depends on Task 7.
---
## Task 1: Create `src/config/server-timing.ts`
**Estimated effort**: 30 minutes
**Files created**: `src/config/server-timing.ts`
**Files modified**: `src/web/server.ts`, `src/web/routes/mux-routes.ts`
### Constants to extract from `src/web/server.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `TERMINAL_BATCH_INTERVAL` | `16` | Terminal data batching interval (60fps) |
| `TASK_UPDATE_BATCH_INTERVAL` | `100` | Task event batching interval (ms) |
| `STATE_UPDATE_DEBOUNCE_INTERVAL` | `500` | State persistence debounce (ms) |
| `SESSIONS_LIST_CACHE_TTL` | `1000` | Sessions list cache TTL (ms) |
| `SCHEDULED_CLEANUP_INTERVAL` | `300000` | Scheduled runs cleanup check (5 min) |
| `SCHEDULED_RUN_MAX_AGE` | `3600000` | Completed scheduled run max age (1 hour) |
| `SSE_HEALTH_CHECK_INTERVAL` | `30000` | SSE client health check (30s) |
| `SESSION_LIMIT_WAIT_MS` | `5000` | Session limit retry wait (5s) |
| `ITERATION_PAUSE_MS` | `2000` | Scheduled run iteration pause (2s) |
| `BATCH_FLUSH_THRESHOLD` | `32768` | Terminal batch immediate flush threshold (32KB) |
| `STATS_COLLECTION_INTERVAL_MS` | `2000` | Mux stats collection interval (2s) |
### Implementation
1. Create `src/config/server-timing.ts` with all 11 constants, preserving existing JSDoc comments.
2. In `src/web/server.ts`: Remove the 11 local constant declarations (lines ~92–121). Add `import { TERMINAL_BATCH_INTERVAL, ... } from '../config/server-timing.js'`.
3. In `src/web/routes/mux-routes.ts`: Remove the duplicate `STATS_COLLECTION_INTERVAL_MS` (line 10) and its comment. Add `import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js'`. This fixes a **duplicate constant** (finding #10).
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Web server performance and scheduling constants.
*
* Controls terminal batching throughput, SSE health checking,
* state persistence debouncing, and scheduled run timing.
*
* @module config/server-timing
*/
// ============================================================================
// Terminal & SSE Performance
// ============================================================================
/** Terminal data batching interval — targets 60fps (ms) */
export const TERMINAL_BATCH_INTERVAL = 16;
/** Immediate flush threshold for terminal batches (bytes).
* Set high (32KB) to allow effective batching; avg Ink events are ~14KB. */
export const BATCH_FLUSH_THRESHOLD = 32 * 1024;
/** Task event batching interval (ms) */
export const TASK_UPDATE_BATCH_INTERVAL = 100;
/** SSE client health check interval (ms) */
export const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
// ============================================================================
// State Persistence
// ============================================================================
/** State update debounce — batches expensive toDetailedState() calls (ms) */
export const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
/** Sessions list cache TTL — avoids re-serializing on every SSE init (ms) */
export const SESSIONS_LIST_CACHE_TTL = 1000;
// ============================================================================
// Scheduled Runs
// ============================================================================
/** Scheduled runs cleanup check interval (ms) */
export const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
/** Completed scheduled run max age before cleanup (ms) */
export const SCHEDULED_RUN_MAX_AGE = 60 * 60 * 1000;
/** Session limit retry wait before retrying (ms) */
export const SESSION_LIMIT_WAIT_MS = 5000;
/** Pause between scheduled run iterations (ms) */
export const ITERATION_PAUSE_MS = 2000;
// ============================================================================
// Mux Stats
// ============================================================================
/** Mux stats collection interval (ms) */
export const STATS_COLLECTION_INTERVAL_MS = 2000;
```
### Verification
```bash
tsc --noEmit
npx tsx src/index.ts web --port 3099 &
curl -s http://localhost:3099/api/status | jq .status # "ok"
kill %1
```
---
## Task 2: Create `src/config/auth-config.ts`
**Estimated effort**: 20 minutes
**Files created**: `src/config/auth-config.ts`
**Files modified**: `src/web/middleware/auth.ts`, `src/hooks-config.ts`
### Constants to extract from `src/web/middleware/auth.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `AUTH_SESSION_TTL_MS` | `86400000` | Auth session cookie TTL (24h) |
| `MAX_AUTH_SESSIONS` | `100` | Max concurrent auth sessions |
| `AUTH_FAILURE_MAX` | `10` | Max failed auth attempts per IP |
| `AUTH_FAILURE_WINDOW_MS` | `900000` | Failed auth tracking window (15 min) |
### Constants to extract from `src/hooks-config.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `HOOK_TIMEOUT_MS` | `10000` | Timeout for Claude Code hook commands |
The `timeout: 10000` value is hardcoded 6 times in `hooks-config.ts` as inline literals. Extract to a single named constant.
### Implementation
1. Create `src/config/auth-config.ts` with the 5 constants.
2. In `src/web/middleware/auth.ts`: Remove the 4 local constant declarations (lines 17–25). Add import from `../../config/auth-config.js`. Keep `AUTH_COOKIE_NAME` in place — it's a string identifier, not a tunable numeric constant.
3. In `src/hooks-config.ts`: Replace all 6 inline `timeout: 10000` occurrences with `timeout: HOOK_TIMEOUT_MS`. Add import from `./config/auth-config.js`.
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Authentication, rate limiting, and hook security constants.
*
* Controls auth session lifecycle, brute-force protection,
* and Claude Code hook timeouts.
*
* @module config/auth-config
*/
// ============================================================================
// Session Cookies
// ============================================================================
/** Auth session cookie TTL — matches autonomous run length (ms) */
export const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
/** Max concurrent auth sessions per server */
export const MAX_AUTH_SESSIONS = 100;
// ============================================================================
// Rate Limiting
// ============================================================================
/** Max failed auth attempts per IP before 429 rejection */
export const AUTH_FAILURE_MAX = 10;
/** Failed auth attempt tracking window (ms) */
export const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
// ============================================================================
// Hooks
// ============================================================================
/** Timeout for Claude Code hook curl commands (ms) */
export const HOOK_TIMEOUT_MS = 10000;
```
### Verification
```bash
tsc --noEmit
npm run lint
```
---
## Task 3: Create `src/config/tunnel-config.ts`
**Estimated effort**: 20 minutes
**Files created**: `src/config/tunnel-config.ts`
**Files modified**: `src/tunnel-manager.ts`
### Constants to extract from `src/tunnel-manager.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `QR_TOKEN_TTL_MS` | `60000` | QR token auto-rotation interval (60s) |
| `QR_TOKEN_GRACE_MS` | `90000` | Grace period for previous token (90s) |
| `SHORT_CODE_LENGTH` | `6` | Length of QR short code |
| `QR_RATE_LIMIT_MAX` | `30` | Global QR attempt rate limit |
| `QR_RATE_LIMIT_WINDOW_MS` | `60000` | QR rate limit reset window (60s) |
| `URL_TIMEOUT_MS` | `30000` | Cloudflared URL fetch timeout (30s) |
| `RESTART_DELAY_MS` | `5000` | Tunnel restart delay after crash (5s) |
| `FORCE_KILL_MS` | `5000` | SIGTERM → SIGKILL escalation timeout (5s) |
### Implementation
1. Create `src/config/tunnel-config.ts` with all 8 constants.
2. In `src/tunnel-manager.ts`: Remove the 8 local constant declarations (lines ~39–75). Add `import { QR_TOKEN_TTL_MS, ... } from './config/tunnel-config.js'`.
3. Keep the `TUNNEL_URL_REGEX` in `tunnel-manager.ts` — it's a parsing detail, not a tuning knob.
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Cloudflare tunnel and QR authentication constants.
*
* Controls QR token rotation timing, rate limiting,
* and tunnel process lifecycle.
*
* @module config/tunnel-config
*/
// ============================================================================
// QR Token Rotation
// ============================================================================
/** QR token auto-rotation interval (ms) */
export const QR_TOKEN_TTL_MS = 60_000;
/** Grace period — previous token still valid during rotation (ms) */
export const QR_TOKEN_GRACE_MS = 90_000;
/** Length of the short code in QR URL path (chars) */
export const SHORT_CODE_LENGTH = 6;
// ============================================================================
// QR Rate Limiting
// ============================================================================
/** Global rate limit for QR auth attempts across all IPs */
export const QR_RATE_LIMIT_MAX = 30;
/** QR rate limit reset window (ms) */
export const QR_RATE_LIMIT_WINDOW_MS = 60_000;
// ============================================================================
// Tunnel Process Lifecycle
// ============================================================================
/** Max time to wait for cloudflared URL before timeout (ms) */
export const URL_TIMEOUT_MS = 30_000;
/** Restart delay after unexpected tunnel exit (ms) */
export const RESTART_DELAY_MS = 5_000;
/** SIGTERM → SIGKILL escalation timeout (ms) */
export const FORCE_KILL_MS = 5_000;
```
### Verification
```bash
tsc --noEmit
```
---
## Task 4: Create `src/config/terminal-limits.ts`
**Estimated effort**: 20 minutes
**Files created**: `src/config/terminal-limits.ts`
**Files modified**: `src/web/routes/session-routes.ts`
### Constants to extract from `src/web/routes/session-routes.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `MAX_INPUT_LENGTH` | `65536` | Max input length per request (64KB) |
| `MAX_TERMINAL_COLS` | `500` | Max terminal columns |
| `MAX_TERMINAL_ROWS` | `200` | Max terminal rows |
| `MAX_SESSION_NAME_LENGTH` | `128` | Max session name length (chars) |
### Why a separate file instead of adding to `buffer-limits.ts`
`buffer-limits.ts` covers memory buffer sizes (2MB terminal, 1MB text). These constants are **validation limits** for API inputs — different concern. A terminal resize request must not exceed `MAX_TERMINAL_COLS`; this has nothing to do with buffer trimming.
### Implementation
1. Create `src/config/terminal-limits.ts` with all 4 constants.
2. In `src/web/routes/session-routes.ts`: Remove the 4 local constant declarations (lines 45–48). Add `import { MAX_INPUT_LENGTH, MAX_TERMINAL_COLS, MAX_TERMINAL_ROWS, MAX_SESSION_NAME_LENGTH } from '../../config/terminal-limits.js'`.
3. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Terminal dimension and input validation limits.
*
* Used by API routes to validate resize, input, and session
* creation requests. Separate from buffer-limits.ts which
* controls memory buffer sizes.
*
* @module config/terminal-limits
*/
/** Max input length per API request (bytes) */
export const MAX_INPUT_LENGTH = 64 * 1024;
/** Max terminal columns for resize requests */
export const MAX_TERMINAL_COLS = 500;
/** Max terminal rows for resize requests */
export const MAX_TERMINAL_ROWS = 200;
/** Max session name length (chars) */
export const MAX_SESSION_NAME_LENGTH = 128;
```
### Verification
```bash
tsc --noEmit
```
---
## Task 5: Create `src/config/ai-defaults.ts`
**Estimated effort**: 30 minutes
**Files created**: `src/config/ai-defaults.ts`
**Files modified**: `src/respawn-controller.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts`, `src/web/routes/respawn-routes.ts`
### Problem: AI model string duplicated 5 times
The model identifier `'claude-opus-4-5-20251101'` appears in 5 places across 4 files. When the model changes, all 5 must be updated — a guaranteed source of bugs. The context limits (`16000`, `8000`) are similarly scattered across 3 files each.
| Constant | Current Value | Duplicated In |
|----------|---------------|---------------|
| `AI_CHECK_MODEL` | `'claude-opus-4-5-20251101'` | `respawn-controller.ts` (×2: idle + plan), `ai-idle-checker.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` (×2: idle + plan) |
| `AI_IDLE_CHECK_MAX_CONTEXT` | `16000` | `respawn-controller.ts`, `ai-idle-checker.ts`, `respawn-routes.ts` |
| `AI_PLAN_CHECK_MAX_CONTEXT` | `8000` | `respawn-controller.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` |
### Implementation
1. Create `src/config/ai-defaults.ts` with the 3 constants.
2. In `src/respawn-controller.ts` `DEFAULT_CONFIG` (line 538): Replace `aiIdleCheckModel: 'claude-opus-4-5-20251101'` with `aiIdleCheckModel: AI_CHECK_MODEL`, `aiIdleCheckMaxContext: 16000` with `aiIdleCheckMaxContext: AI_IDLE_CHECK_MAX_CONTEXT`, `aiPlanCheckModel: 'claude-opus-4-5-20251101'` with `aiPlanCheckModel: AI_CHECK_MODEL`, `aiPlanCheckMaxContext: 8000` with `aiPlanCheckMaxContext: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
3. In `src/ai-idle-checker.ts` `DEFAULT_AI_CHECK_CONFIG` (line 46): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 16000` with `maxContextChars: AI_IDLE_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
4. In `src/ai-plan-checker.ts` `DEFAULT_PLAN_CHECK_CONFIG` (line 45): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 8000` with `maxContextChars: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
5. In `src/web/routes/respawn-routes.ts` config merge block (lines 173–179): Replace all 4 inline fallback values with imports from `../../config/ai-defaults.js`.
6. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Default model and context limits for AI-powered checkers.
*
* Centralizes the AI model identifier and context window sizes used by
* the idle checker, plan checker, respawn controller defaults, and
* respawn route fallbacks. Change the model here when upgrading.
*
* @module config/ai-defaults
*/
/** Default model for AI idle and plan checkers */
export const AI_CHECK_MODEL = 'claude-opus-4-5-20251101';
/** Max context chars for idle checker (~4k tokens) */
export const AI_IDLE_CHECK_MAX_CONTEXT = 16000;
/** Max context chars for plan checker (~2k tokens, plan mode UI is compact) */
export const AI_PLAN_CHECK_MAX_CONTEXT = 8000;
```
### Verification
```bash
tsc --noEmit
# Verify no remaining hardcoded model strings
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
```
---
## Task 6: Create `src/config/team-config.ts`
**Estimated effort**: 15 minutes
**Files created**: `src/config/team-config.ts`
**Files modified**: `src/team-watcher.ts`
### Constants to extract from `src/team-watcher.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `TEAM_POLL_INTERVAL_MS` | `30000` | Team directory poll interval (30s) |
| `MAX_CACHED_TEAMS` | `50` | LRU cache size for team configs |
| `MAX_CACHED_TASKS` | `200` | LRU cache size for team tasks + inboxes |
### Why centralize these
Team polling frequency and cache sizes are operational knobs that affect both performance (polling too often wastes CPU) and responsiveness (polling too rarely means stale team state in the UI). They're also the kind of values a developer tuning for a large team deployment would want to find quickly. `MAX_CACHED_TASKS` is used for both the task cache and inbox cache — worth documenting.
### Implementation
1. Create `src/config/team-config.ts` with the 3 constants.
2. In `src/team-watcher.ts`: Remove the 3 local constants (lines 23–25). Add `import { TEAM_POLL_INTERVAL_MS, MAX_CACHED_TEAMS, MAX_CACHED_TASKS } from './config/team-config.js'`. Note: rename `POLL_INTERVAL_MS` → `TEAM_POLL_INTERVAL_MS` to avoid ambiguity with the identically-named constant in `subagent-watcher.ts`.
3. Update the usage site: `setInterval(... POLL_INTERVAL_MS)` → `setInterval(... TEAM_POLL_INTERVAL_MS)`.
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Agent Teams polling and cache configuration.
*
* Controls how frequently TeamWatcher polls ~/.claude/teams/
* and how many teams/tasks are cached in memory.
*
* @module config/team-config
*/
/** Team directory poll interval (ms) */
export const TEAM_POLL_INTERVAL_MS = 30_000;
/** Max cached team configs (LRU eviction) */
export const MAX_CACHED_TEAMS = 50;
/** Max cached team tasks and inbox messages (LRU eviction).
* Used for both teamTasks and inboxCache maps. */
export const MAX_CACHED_TASKS = 200;
```
### Verification
```bash
tsc --noEmit
```
---
## Task 7: Fix remaining cross-file duplicates
**Estimated effort**: 30 minutes
**Files modified**: `src/index.ts`, `src/subagent-watcher.ts`
### Duplicate 1: `STATS_COLLECTION_INTERVAL_MS`
Already fixed in Task 1 — both `server.ts` and `mux-routes.ts` now import from `server-timing.ts`.
### Duplicate 2: AI model string
Already fixed in Task 5 — all 5 occurrences now import from `ai-defaults.ts`.
### Duplicate 3: `MAX_SCREENSHOT_SIZE` / `MAX_TEXT_FILE_SIZE` / `MAX_RAW_FILE_SIZE`
These file size limits in `file-routes.ts` and `system-routes.ts` are **API-specific validation limits**. They're only used in their respective route files and aren't duplicated. **Leave in place** — they're local to their route module and well-commented.
### Action A: Move `MAX_CONSECUTIVE_ERRORS` and `ERROR_RESET_MS` to config
`src/index.ts` has two process-level constants that are operational tuning knobs:
| Constant | Value | Purpose |
|----------|-------|---------|
| `MAX_CONSECUTIVE_ERRORS` | `5` | Max consecutive unhandled errors before process exit |
| `ERROR_RESET_MS` | `60000` | Error counter reset interval (1 min) |
These belong in a config file since they control server reliability behavior. Add them to `src/config/server-timing.ts` (they're server operational constants).
1. Add to `src/config/server-timing.ts`:
```typescript
// ============================================================================
// Process Error Recovery
// ============================================================================
/** Max consecutive unhandled errors before auto-restart */
export const MAX_CONSECUTIVE_ERRORS = 5;
/** Error counter reset interval — forgives errors after quiet period (ms) */
export const ERROR_RESET_MS = 60_000;
```
2. In `src/index.ts`: Remove lines 19–20, add import from `./config/server-timing.js`.
3. Run `tsc --noEmit`.
### Action B: Fix `MAX_TRACKED_AGENTS` shadow in `subagent-watcher.ts`
`subagent-watcher.ts` defines its own `MAX_TRACKED_AGENTS = 500` locally instead of importing the identical value from `config/map-limits.ts`. This is a latent bug — if someone changes the config value, the subagent watcher's copy stays stale.
1. In `src/subagent-watcher.ts`: Remove the local `MAX_TRACKED_AGENTS` constant. Add `import { MAX_TRACKED_AGENTS } from './config/map-limits.js'` (the value there is `MAX_TODOS_PER_SESSION = 500` — **verify** the map-limits constant is actually named `MAX_TRACKED_AGENTS` or if it needs to be added). If the constant doesn't exist in `map-limits.ts` under that name, add it.
2. Run `tsc --noEmit`.
### Verification
```bash
tsc --noEmit
npm run lint
npm run format:check
```
---
## Task 8: Update CLAUDE.md and final verification
**Estimated effort**: 20 minutes
**Files modified**: `CLAUDE.md`
### Updates to CLAUDE.md
1. **Config Files table** (`src/config/`): Add the 6 new files:
| File | Purpose |
|------|---------|
| `buffer-limits.ts` | Terminal/text buffer size limits |
| `map-limits.ts` | Global limits for Maps, sessions, watchers |
| `exec-timeout.ts` | Execution timeout configuration |
| `server-timing.ts` | Web server batching, SSE, scheduled run timing |
| `auth-config.ts` | Auth session TTL, rate limits, hook timeout |
| `tunnel-config.ts` | QR token rotation, tunnel process lifecycle |
| `terminal-limits.ts` | Terminal dimension and input validation limits |
| `ai-defaults.ts` | AI checker model and context limits |
| `team-config.ts` | Agent Teams polling and cache sizes |
2. **Import Conventions** section: Add:
```
- **Config**: Import from specific files: `import { MAX_TERMINAL_COLS } from './config/terminal-limits'`
```
3. **Phase 6 status** in `docs/code-structure-findings.md`: Mark as COMPLETE with summary of what was done.
### Final verification checklist
```bash
# Type checking
tsc --noEmit
# Linting
npm run lint
# Formatting
npm run format:check
# Dev server starts
npx tsx src/index.ts web --port 3099 &
curl -s http://localhost:3099/api/status | jq .status # "ok"
kill %1
# Verify no remaining duplicates
grep -rn 'STATS_COLLECTION_INTERVAL_MS' src/ # Should only appear in config + import sites
grep -rn 'timeout: 10000' src/hooks-config.ts # Should be 0 — all replaced with HOOK_TIMEOUT_MS
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
```
---
## What is NOT in scope (and why)
These constants were considered but deliberately left in their current files:
### Respawn controller defaults (`src/respawn-controller.ts`)
The `DEFAULT_CONFIG` object (lines 538–578) contains ~30 default values for the `RespawnConfig` interface. These are **user-facing configuration defaults**, not system constants — they're the starting values for a config object that users can modify via the API and UI. Centralizing them would break the locality between the config interface definition and its defaults. They already have excellent JSDoc with `@default` tags. The only values extracted are the AI model/context constants (Task 5) which are duplicated in other files.
### Subagent watcher timing (`src/subagent-watcher.ts`)
The 18 constants at lines 129–158 are all internal to the subagent watcher's polling/lifecycle algorithm. Moving them to a config file would force developers to context-switch between two files to understand the polling logic. They're already grouped with clear comments. Exception: `MAX_TRACKED_AGENTS` is consolidated with `map-limits.ts` (Task 7B) since it duplicates a global limit.
### Session auto-ops timing (`src/session-auto-ops.ts`)
The 8 constants at lines 19–40 are internal to the auto-compact/clear retry state machine. They form a coherent group that's meaningless without the surrounding implementation context.
### Run summary constants (`src/run-summary.ts`)
`MAX_EVENTS`, `TRIM_TO_EVENTS`, `TOKEN_MILESTONE_INTERVAL`, `STATE_STUCK_WARNING_MS`, `STATE_STUCK_CHECK_INTERVAL` — all module-internal. The buffer-style limits (`MAX_EVENTS`/`TRIM_TO_EVENTS`) follow the same pattern as `buffer-limits.ts` but are only used in this one file.
### Frontend (`src/web/public/constants.js`)
Already well-centralized. Frontend and backend run in different environments — mixing them in TypeScript config files would create import problems. If frontend constants need expansion, do it in `constants.js`. Note: `app.js` has 2 inline uses of `256 * 1024` that should use the existing `TERMINAL_TAIL_SIZE` from `constants.js` — a minor cleanup that can be done opportunistically but is not worth a task here.
### Tmux manager timing (`src/tmux-manager.ts`)
The 6 constants (lines 65–78) are internal to tmux process lifecycle management. They're low-level retry/wait values that are meaningless without understanding the tmux spawn sequence.
### Process-internal constants
`image-watcher.ts`, `bash-tool-parser.ts`, `transcript-watcher.ts`, `ralph-tracker.ts`, `task-tracker.ts`, `file-stream-manager.ts`, `session-lifecycle-log.ts`, `session-task-cache.ts`, `respawn-metrics.ts`, `respawn-adaptive-timing.ts`, `ai-checker-base.ts` — all have module-local constants that are internal implementation details.
### `localhost:3000` default URL
The string `'http://localhost:3000'` or port `3000` appears as a fallback default in ~5 files (`session-cli-builder.ts`, `tmux-manager.ts`, `tunnel-manager.ts`, `server.ts`, CLI). While technically duplicated, extracting it provides little value — each usage has a different fallback chain (env var → config → hardcoded) and the port is also baked into systemd service files and documentation. The risk of a missed update is low since port 3000 is deeply conventional.
### `SAVE_DEBOUNCE_MS = 500` in `state-store.ts` / `push-store.ts`
Same value (500ms), but they debounce different persistence targets (state.json vs push-subscriptions.json). If one needed faster/slower debouncing, they'd diverge. Coupling them would be misleading.
---
## Summary
| Metric | Before | After |
|--------|--------|-------|
| Config files in `src/config/` | 3 | 9 |
| Constants centralized | ~25 | ~65 |
| Cross-file duplicates | 9+ (`STATS_COLLECTION_INTERVAL_MS`, `timeout: 10000` ×6, AI model ×5, context limits ×3 each, `MAX_TRACKED_AGENTS`) | 0 |
| Files with `timeout: 10000` inline | 1 (6 occurrences) | 0 |
| Files with hardcoded AI model string | 4 (5 occurrences) | 1 (config only) |
| Files modified | — | 11 |
| Files created | — | 6 |
@@ -1,953 +0,0 @@
# Phase 7 Implementation Plan: Test Infrastructure
**Source**: `docs/code-structure-findings.md` (Phase 7 — Test Infrastructure)
**Estimated effort**: 2–3 days
**Tasks**: 11 tasks with dependencies (see dependency graph below)
---
## Safety Constraints
Before starting ANY work, read and follow these rules:
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
7. **Port assignments for this phase**: New tests use ports 3220–3229 (see individual tasks for assignments).
---
## Goal
Eliminate duplicated test mocks, activate the unused `respawn-test-utils.ts` utilities, and add route-level test coverage for the server's 12 route modules — the single largest untested area in the codebase (162 route handlers, 0 dedicated tests).
**Non-goals**:
- Full end-to-end integration tests (those require real Claude CLI / tmux sessions)
- 100% route coverage in this phase — focus on the highest-value route modules first
- Refactoring test patterns in existing passing tests that don't use shared mocks
- Migrating `vi.mock()`-based module replacement mocks (different pattern, see Task 6/7)
---
## Current State
### Mock Duplication (Finding #9)
`MockSession` is defined **4 times** across test files with varying levels of completeness:
| File | Properties | Methods | EventEmitter | Notes |
|------|-----------|---------|-------------|-------|
| `test/respawn-test-utils.ts` | 6 | 20+ | Yes | **Most complete**. Includes terminal simulation, token count, ANSI output, plan mode prompts. **Never imported by any test.** |
| `test/respawn-controller.test.ts` | 6 | 9 | Yes | Subset of respawn-test-utils. Missing token simulation, ANSI helpers. |
| `test/respawn-team-awareness.test.ts` | ~6 | ~9 | Yes | Near-copy of respawn-controller.test.ts version. |
| `test/session-manager.test.ts` | 4 | 8 | Yes | **Inside `vi.mock()` factory** — replaces `../src/session.js` module. Different shape: `start()`/`stop()`/`toState()`/`sendInput()` for lifecycle testing. |
`MockStateStore` is defined **2 times** (both inside `vi.mock()` factories):
| File | Shape | Methods | Mock Pattern |
|------|-------|---------|-------------|
| `test/session-manager.test.ts` | `{ sessions, config }` | `getConfig`, `getSessions`, `getSession`, `setSession`, `removeSession` | `vi.mock('../src/state-store.js')` |
| `test/ralph-loop.test.ts` | `{ ralphLoop, tasks, config }` | `getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask` | `vi.mock('../src/state-store.js')` |
### Important: Two distinct mocking patterns
The codebase uses two different mocking patterns that require different migration strategies:
1. **Direct instantiation** (respawn-controller, respawn-team-awareness): `MockSession` is defined at file scope and instantiated directly in tests. These can be migrated to shared mocks via simple import replacement.
2. **Module replacement** (session-manager, ralph-loop): Mocks are defined inside `vi.mock()` factories that replace entire modules (`../src/session.js`, `../src/state-store.js`). These factories run in an isolated scope and return `{ Session: MockClass }` or `{ getStore: vi.fn(() => instance) }`. Migrating these requires either `vi.hoisted()` or restructuring the test's module mocking — higher risk for limited benefit.
### Unused Test Utilities
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
- `TimeController` / `createTimeController()` — abstraction over vitest fake timers
- `MockAiIdleChecker` / `MockAiPlanChecker` — fully mocked AI checkers with result queueing
- `createStateTracker()` / `createEventRecorder()` — state transition and event recording
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — pre-configured RespawnConfig objects
- `waitForState()` / `waitForEvent()` / `createDeferred()` — async test helpers
- `terminalOutputs` — factory object for common terminal output patterns
### Route Test Coverage
Currently **zero** dedicated tests for the 12 route modules in `src/web/routes/`. The existing test files that touch API endpoints:
| Test File | What It Tests | Approach |
|-----------|--------------|----------|
| `test/api-responses.test.ts` | Response structure validation | Imports types, no HTTP calls |
| `test/api-generate-plan.test.ts` | Plan generation API | Mocks validation logic, Port 3191 declared |
| `test/auth-security.test.ts` | Auth middleware | Integration tests with WebServer, Ports 3160/3161 |
| `test/qr-auth.test.ts` | QR authentication | Integration + unit tests, Port 3162 |
None of these test the route handlers themselves with real HTTP requests against a running Fastify instance.
---
## Design Decisions
### Shared mocks: Superset strategy
Rather than creating a lowest-common-denominator mock, `MockSession` in `test/mocks/` will be the **superset** from `respawn-test-utils.ts` (the most complete version). Test files that need a simpler mock can just ignore the extra methods — having unused methods costs nothing, but missing methods forces local re-definition.
### vi.mock() tests: Don't migrate
The `session-manager.test.ts` and `ralph-loop.test.ts` tests define mocks inside `vi.mock()` factories. These use **module-level replacement** (replacing `../src/session.js` and `../src/state-store.js` entirely), which is fundamentally different from the direct-instantiation pattern. Migrating them would require `vi.hoisted()` or factory restructuring — high complexity for limited benefit since these mocks are already working. We leave these as-is and create the shared mocks for **new** tests and for the two direct-instantiation tests (Tasks 4–5).
### MockStateStore: Union of both shapes
The shared `MockStateStore` in `test/mocks/` will include methods from both existing definitions (session management + Ralph loop), so any **new** test can use it. Methods default to no-ops via `vi.fn()`. Existing `vi.mock()`-based tests are not migrated.
### Route testing strategy: Lightweight Fastify instances
Each route test file will:
1. Create a minimal `Fastify` instance
2. Register **only** the route module under test
3. Provide a mock context object satisfying the port interfaces
4. Use `app.inject()` (Fastify's built-in test helper) — no real HTTP, no port needed
This avoids port conflicts entirely and runs fast. Only tests that need SSE or WebSocket behavior will use a real listening server with assigned ports.
### Port assignments (for tests needing real servers)
| Port | Test File | Purpose |
|------|-----------|---------|
| 3220 | `test/routes/session-routes.test.ts` | SSE integration (if needed) |
| 3221 | `test/routes/system-routes.test.ts` | Status/stats endpoints |
| 3222 | `test/routes/respawn-routes.test.ts` | Respawn API |
| 3223 | `test/routes/ralph-routes.test.ts` | Ralph API |
| 3224–3229 | Reserved | Future route tests |
Most tests should NOT need real ports — `app.inject()` is preferred. Verified: ports 3220–3229 are completely unused by existing tests (highest used port is 3211 in `opencode-resize.test.ts`).
---
## Task Dependencies
```
Task 1 (Consolidate MockSession)
Task 2 (Consolidate MockStateStore)
└──> Task 3 (Create test/mocks/ barrel)
├──> Task 4 (Migrate respawn-controller.test.ts)
├──> Task 5 (Migrate respawn-team-awareness.test.ts)
└──> Task 6 (Route test scaffold + helpers)
├──> Task 7 (Session routes tests)
└──> Task 8 (System + respawn routes tests)
Task 9 (Slim down respawn-test-utils.ts) — depends on Tasks 4, 5
```
**Tasks 1–2** are independent and can run in parallel.
**Task 3** depends on Tasks 1–2.
**Tasks 4–6** depend on Task 3 and can run in parallel.
**Tasks 7–8** depend on Task 6 and can run in parallel.
**Task 9** depends on Tasks 4, 5 (must verify migrations work before removing duplicates from source).
---
## Task 1: Consolidate MockSession into `test/mocks/mock-session.ts`
**Estimated effort**: 2 hours
**Files created**: `test/mocks/mock-session.ts`
**Files modified**: None yet (consumers migrate in Tasks 4–5)
### Source
The canonical MockSession comes from `test/respawn-test-utils.ts` (lines 89–241). It is the most complete version with:
- All properties needed by `RespawnController`: `id`, `workingDir`, `status`, `writeBuffer`, `terminalBuffer`, `muxName`
- `write()` / `writeViaMux()` for input simulation
- Buffer inspection: `lastWrite`, `hasWritten(pattern)`, `clearWriteBuffer()`
- Terminal simulation: `simulateTerminalOutput()`, `simulatePrompt()`, `simulateReady()`, `simulateCompletionMessage()`, `simulateWorking()`, `simulateClearComplete()`, `simulateInitComplete()`, `simulatePlanModePrompt()`, `simulateElicitationDialog()`, `simulateTokenCount()`, `simulateAnsiOutput()`
- Lifecycle: `close()`
### Implementation
1. Create `test/mocks/` directory.
2. Create `test/mocks/mock-session.ts`:
- Copy the `MockSession` class **exactly** from `test/respawn-test-utils.ts` (lines 89–241)
- Copy `terminalOutputs` helper object (tightly coupled to mock)
- Copy `createMockSession()` factory function
- Export all three: `export { MockSession, createMockSession, terminalOutputs }`
- Ensure all `vi` imports come from `vitest`
**CRITICAL**: Copy the source verbatim — do NOT rewrite the simulation methods. The respawn controller's detection logic matches specific output patterns (e.g., `'\u276f '` for prompt, `'\u273b Worked for'` for completion). Using different patterns would cause test failures.
### Template
```typescript
/**
* Shared MockSession for tests that need terminal simulation.
*
* Copied from test/respawn-test-utils.ts (the canonical, most complete version).
* Used by respawn, route, and subagent tests.
*/
import { EventEmitter } from 'node:events';
// Copy MockSession class exactly from test/respawn-test-utils.ts lines 89–241
export class MockSession extends EventEmitter {
// ... (copy verbatim from respawn-test-utils.ts)
}
/**
* Factory for common terminal output strings.
* Must match the patterns used in MockSession's simulate* methods.
*/
export const terminalOutputs = {
// ... (copy verbatim from respawn-test-utils.ts)
};
/**
* Convenience factory.
*/
export function createMockSession(id?: string): MockSession {
return new MockSession(id);
}
```
### Verification
```bash
tsc --noEmit # Ensure file compiles
```
---
## Task 2: Consolidate MockStateStore into `test/mocks/mock-state-store.ts`
**Estimated effort**: 1 hour
**Files created**: `test/mocks/mock-state-store.ts`
**Files modified**: None (existing vi.mock()-based tests are NOT migrated; this is for new route tests)
### Source
Union of both existing definitions:
- From `test/session-manager.test.ts`: session CRUD methods (`getConfig`, `getSession`, `setSession`, `removeSession`, `getSessions`)
- From `test/ralph-loop.test.ts`: Ralph state methods (`getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask`)
### Template
```typescript
/**
* Shared MockStateStore for tests.
*
* Includes methods for both session management and Ralph loop testing.
* All methods are vi.fn() spies — tests can override return values as needed.
*
* NOTE: This is for direct instantiation in new tests. Existing tests that
* use vi.mock('../src/state-store.js') keep their inline definitions.
*/
import { vi } from 'vitest';
export class MockStateStore {
state: Record<string, unknown> = {
sessions: {} as Record<string, unknown>,
config: { maxConcurrentSessions: 5 },
ralphLoop: { status: 'stopped' },
tasks: {} as Record<string, unknown>,
};
// Session methods
getConfig = vi.fn(() => this.state.config);
getSessions = vi.fn(() => this.state.sessions as Record<string, unknown>);
getSession = vi.fn((id: string) => (this.state.sessions as Record<string, unknown>)[id]);
setSession = vi.fn((id: string, state: unknown) => {
(this.state.sessions as Record<string, unknown>)[id] = state;
});
removeSession = vi.fn((id: string) => {
delete (this.state.sessions as Record<string, unknown>)[id];
});
// Ralph state methods
getRalphLoopState = vi.fn(() => this.state.ralphLoop);
setRalphLoopState = vi.fn((update: Record<string, unknown>) => {
this.state.ralphLoop = { ...(this.state.ralphLoop as Record<string, unknown>), ...update };
});
// Task methods
getTasks = vi.fn(() => this.state.tasks);
setTask = vi.fn();
removeTask = vi.fn();
// Settings methods
getSettings = vi.fn(() => ({}));
setSettings = vi.fn();
// Generic persistence
save = vi.fn();
load = vi.fn();
/** Reset all state and mocks for clean test isolation */
reset(): void {
this.state = {
sessions: {},
config: { maxConcurrentSessions: 5 },
ralphLoop: { status: 'stopped' },
tasks: {},
};
vi.clearAllMocks();
}
}
```
### Verification
```bash
tsc --noEmit
```
---
## Task 3: Create `test/mocks/index.ts` barrel export
**Estimated effort**: 30 minutes
**Depends on**: Tasks 1, 2
**Files created**: `test/mocks/index.ts`, `test/mocks/test-helpers.ts`
**Files modified**: None
### Implementation
1. Create `test/mocks/test-helpers.ts` with the async utilities from `respawn-test-utils.ts`:
```typescript
/**
* Reusable async test helpers.
* Extracted from respawn-test-utils.ts.
*/
/** Wait for an EventEmitter to emit a specific event, with timeout */
export function waitForEvent(
emitter: { once: (event: string, listener: (...args: unknown[]) => void) => void },
event: string,
timeoutMs = 5000,
): Promise<unknown> {
return new Promise((resolve, reject) => {
const timer = setTimeout(
() => reject(new Error(`Timed out waiting for event "${event}" after ${timeoutMs}ms`)),
timeoutMs,
);
emitter.once(event, (...args: unknown[]) => {
clearTimeout(timer);
resolve(args.length === 1 ? args[0] : args);
});
});
}
/** Create a deferred promise with external resolve/reject */
export function createDeferred<T = void>(): {
promise: Promise<T>;
resolve: (value: T) => void;
reject: (reason?: unknown) => void;
} {
let resolve!: (value: T) => void;
let reject!: (reason?: unknown) => void;
const promise = new Promise<T>((res, rej) => {
resolve = res;
reject = rej;
});
return { promise, resolve, reject };
}
```
2. Create `test/mocks/index.ts` barrel:
```typescript
/**
* Shared test mocks — import from here instead of defining inline.
*
* @example
* import { MockSession, MockStateStore, terminalOutputs } from './mocks/index.js';
*/
export { MockSession, createMockSession, terminalOutputs } from './mock-session.js';
export { MockStateStore } from './mock-state-store.js';
export { waitForEvent, createDeferred } from './test-helpers.js';
```
### Verification
```bash
tsc --noEmit
```
---
## Task 4: Migrate `respawn-controller.test.ts` to shared mocks
**Estimated effort**: 30 minutes
**Depends on**: Task 3
**Files modified**: `test/respawn-controller.test.ts`
### Steps
1. Remove the local `MockSession` class definition (approx. 50 lines).
2. Add: `import { MockSession } from './mocks/index.js';`
3. Verify all test methods still exist on the shared mock. The shared mock is a superset, so all existing usage should work.
4. If the local mock had any test-specific customizations (e.g., extra properties added in `beforeEach`), keep those in the test file as inline assignments on the shared instance.
5. Run the test to confirm it passes.
### Potential issues
- The local mock's `simulateCompletionMessage()` may have a slightly different output format than the shared mock's (from respawn-test-utils.ts). Verify the respawn controller's completion detection regex matches the shared mock's output pattern (`'\u273b Worked for ...'`).
- If the local mock adds `pid` or `isWorking` properties that the shared mock doesn't have, add inline assignments in `beforeEach`.
### Verification
```bash
npx vitest run test/respawn-controller.test.ts
```
---
## Task 5: Migrate `respawn-team-awareness.test.ts` to shared mocks
**Estimated effort**: 30 minutes
**Depends on**: Task 3
**Files modified**: `test/respawn-team-awareness.test.ts`
### Steps
1. Remove the local `MockSession` class definition.
2. Add: `import { MockSession } from './mocks/index.js';`
3. Keep `MockTeamWatcher` in this file — it's test-specific and extends the real `TeamWatcher`, not a general-purpose mock.
4. Run the test to confirm it passes.
### Verification
```bash
npx vitest run test/respawn-team-awareness.test.ts
```
---
## Task 6: Create route test scaffold and helpers
**Estimated effort**: 2 hours
**Depends on**: Task 3
**Files created**: `test/mocks/mock-route-context.ts`, `test/routes/` directory, `test/routes/_route-test-utils.ts`
### Problem
The 12 route modules in `src/web/routes/` have zero dedicated test coverage. Each route module takes `(app: FastifyInstance, ctx: PortIntersection)` — we need a reusable way to create mock context objects that satisfy the port interfaces.
### Design
Create a `MockRouteContext` factory that builds a mock object satisfying all port interfaces. Each port's methods are `vi.fn()` stubs. Tests can override specific methods as needed.
### Route registration signatures (verified)
Each route module requires a specific port intersection. The mock must satisfy all of them:
| Route Module | Required Ports |
|-------------|----------------|
| `registerSessionRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
| `registerSystemRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
| `registerRespawnRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
| `registerRalphRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
| `registerPlanRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort` |
| `registerCaseRoutes` | `EventPort & ConfigPort` |
| `registerScheduledRoutes` | `SessionPort & EventPort & InfraPort` |
| `registerFileRoutes` | `SessionPort` |
| `registerMuxRoutes` | `InfraPort` |
| `registerPushRoutes` | `InfraPort` |
| `registerTeamRoutes` | `InfraPort` |
| `registerHookEventRoutes` | `EventPort & AuthPort` |
### Implementation
1. Create `test/mocks/mock-route-context.ts`:
```typescript
/**
* Mock context for route handler testing.
*
* Satisfies ALL port interfaces (SessionPort, EventPort, RespawnPort,
* ConfigPort, InfraPort, AuthPort) so any route module can be tested.
* Override specific methods in individual tests as needed.
*
* Verified against actual port interfaces in src/web/ports/:
* - SessionPort: 6 methods (sessions, addSession, cleanupSession,
* setupSessionListeners, persistSessionState, persistSessionStateNow,
* getSessionStateWithRespawn)
* - EventPort: 5 methods (broadcast, sendPushNotifications, batchTerminalData,
* broadcastSessionStateDebounced, batchTaskUpdate)
* - RespawnPort: 2 maps + 4 methods
* - ConfigPort: 5 readonly + 7 methods (incl getDefaultClaudeMdPath,
* getLightState, getLightSessionsState, stopTranscriptWatcher)
* - InfraPort: 7 readonly + 2 methods (startScheduledRun, stopScheduledRun)
* - AuthPort: 3 readonly (authSessions, qrAuthFailures, https)
*/
import { vi } from 'vitest';
import { MockSession, createMockSession } from './mock-session.js';
/**
* Creates a mock context that satisfies all port interfaces.
* Pre-populated with one session for convenience.
*/
export function createMockRouteContext(options?: { sessionId?: string }) {
const sessionId = options?.sessionId ?? 'test-session-1';
const session = createMockSession(sessionId);
const sessions = new Map<string, MockSession>();
sessions.set(sessionId, session);
return {
// -- SessionPort --
sessions,
addSession: vi.fn(),
cleanupSession: vi.fn(),
setupSessionListeners: vi.fn(),
persistSessionState: vi.fn(),
persistSessionStateNow: vi.fn(),
getSessionStateWithRespawn: vi.fn((s: unknown) => s),
// -- EventPort --
broadcast: vi.fn(),
sendPushNotifications: vi.fn(),
batchTerminalData: vi.fn(),
broadcastSessionStateDebounced: vi.fn(),
batchTaskUpdate: vi.fn(),
// -- RespawnPort --
respawnControllers: new Map(),
respawnTimers: new Map(),
setupRespawnListeners: vi.fn(),
setupTimedRespawn: vi.fn(),
restoreRespawnController: vi.fn(),
saveRespawnConfig: vi.fn(),
// -- ConfigPort --
store: {
getConfig: vi.fn(() => ({})),
getSessions: vi.fn(() => ({})),
getSession: vi.fn(),
setSession: vi.fn(),
removeSession: vi.fn(),
getSettings: vi.fn(() => ({})),
setSettings: vi.fn(),
getRalphLoopState: vi.fn(() => ({})),
setRalphLoopState: vi.fn(),
getTasks: vi.fn(() => ({})),
save: vi.fn(),
load: vi.fn(),
},
port: 3000,
https: false,
testMode: true,
serverStartTime: Date.now(),
getGlobalNiceConfig: vi.fn(async () => undefined),
getModelConfig: vi.fn(async () => null),
getClaudeModeConfig: vi.fn(async () => ({})),
getDefaultClaudeMdPath: vi.fn(async () => undefined),
getLightState: vi.fn(() => ({ sessions: [], status: 'ok' })),
getLightSessionsState: vi.fn(() => []),
startTranscriptWatcher: vi.fn(),
stopTranscriptWatcher: vi.fn(),
// -- InfraPort --
mux: {
createSession: vi.fn(),
killSession: vi.fn(),
listSessions: vi.fn(() => []),
getStats: vi.fn(() => ({})),
},
runSummaryTrackers: new Map(),
activePlanOrchestrators: new Map(),
scheduledRuns: new Map(),
teamWatcher: { getTeams: vi.fn(() => []), hasActiveTeammates: vi.fn(() => false) },
tunnelManager: null,
pushStore: null,
startScheduledRun: vi.fn(),
stopScheduledRun: vi.fn(),
// -- AuthPort --
authSessions: null,
qrAuthFailures: null,
// https already declared above in ConfigPort (shared property)
// Convenience accessors (not part of any port interface)
_session: session,
_sessionId: sessionId,
};
}
export type MockRouteContext = ReturnType<typeof createMockRouteContext>;
```
2. Add to `test/mocks/index.ts` barrel:
```typescript
export { createMockRouteContext, type MockRouteContext } from './mock-route-context.js';
```
3. Create `test/routes/` directory for route test files.
4. Create `test/routes/_route-test-utils.ts` with Fastify test helpers:
```typescript
/**
* Shared utilities for route testing.
*
* Creates minimal Fastify instances with just the route module under test
* and a mock context. Uses app.inject() for HTTP testing without real ports.
*/
import Fastify, { type FastifyInstance } from 'fastify';
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
export interface RouteTestHarness {
app: FastifyInstance;
ctx: MockRouteContext;
}
/**
* Creates a Fastify instance with a route module registered against a mock context.
*
* @param registerFn - The route registration function (e.g., registerSessionRoutes).
* Uses `any` for ctx parameter because route functions expect typed port intersections
* that MockRouteContext satisfies structurally but not nominally.
* @param ctxOptions - Optional overrides for the mock context
*/
export async function createRouteTestHarness(
// eslint-disable-next-line @typescript-eslint/no-explicit-any
registerFn: (app: FastifyInstance, ctx: any) => void,
ctxOptions?: { sessionId?: string },
): Promise<RouteTestHarness> {
const app = Fastify({ logger: false });
const ctx = createMockRouteContext(ctxOptions);
registerFn(app, ctx);
await app.ready();
return { app, ctx };
}
```
### Why `ctx: any` in the harness
Route registration functions like `registerSessionRoutes(app, ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort)` expect specific port intersection types. TypeScript won't accept `unknown` here because it's not assignable to the port types. The `MockRouteContext` satisfies the interfaces structurally (it has all the required properties and methods), but since it's not declared as implementing them, we need `any` at the call site. This is the standard pattern for test mocks in TypeScript.
### Verification
```bash
tsc --noEmit
```
---
## Task 7: Add session routes tests
**Estimated effort**: 4 hours
**Depends on**: Task 6
**Files created**: `test/routes/session-routes.test.ts`
**Port**: 3220 (only if SSE tests needed; prefer `app.inject()`)
### Coverage targets
`src/web/routes/session-routes.ts` is the largest route module (43 handlers). Focus on the most critical endpoints first:
#### Priority 1: Session CRUD (must test)
| Method | Path | What to test |
|--------|------|-------------|
| `GET` | `/api/sessions` | Returns session list; empty when no sessions |
| `GET` | `/api/sessions/:id` | Returns session state; 404 for unknown ID |
| `POST` | `/api/sessions` | Creates session; validates workingDir; rejects invalid paths |
| `DELETE` | `/api/sessions/:id` | Calls cleanupSession; 404 for unknown ID |
#### Priority 2: Session I/O
| Method | Path | What to test |
|--------|------|-------------|
| `POST` | `/api/sessions/:id/input` | Sends input to session; validates input length; 404 for unknown |
| `POST` | `/api/sessions/:id/resize` | Validates cols/rows bounds; 404 for unknown |
| `GET` | `/api/sessions/:id/buffer` | Returns terminal buffer; 404 for unknown |
#### Priority 3: Session actions
| Method | Path | What to test |
|--------|------|-------------|
| `POST` | `/api/sessions/:id/run` | Runs prompt on session |
| `POST` | `/api/sessions/:id/clear` | Clears session |
| `POST` | `/api/sessions/:id/compact` | Compacts session |
| `POST` | `/api/sessions/:id/interactive` | Starts interactive mode |
| `POST` | `/api/sessions/:id/quick-start` | Quick start flow |
### Test pattern
```typescript
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { createRouteTestHarness, type RouteTestHarness } from './_route-test-utils.js';
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
describe('session-routes', () => {
let harness: RouteTestHarness;
beforeEach(async () => {
harness = await createRouteTestHarness(registerSessionRoutes);
});
afterEach(async () => {
await harness.app.close();
});
describe('GET /api/sessions', () => {
it('returns empty array when no sessions', async () => {
harness.ctx.sessions.clear();
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
expect(res.statusCode).toBe(200);
expect(JSON.parse(res.body)).toEqual([]);
});
it('returns session list with one session', async () => {
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
expect(res.statusCode).toBe(200);
const sessions = JSON.parse(res.body);
expect(sessions).toHaveLength(1);
});
});
describe('GET /api/sessions/:id', () => {
it('returns 404 for unknown session', async () => {
const res = await harness.app.inject({
method: 'GET',
url: '/api/sessions/nonexistent',
});
expect(res.statusCode).toBe(404);
});
});
describe('POST /api/sessions/:id/input', () => {
it('rejects input exceeding max length', async () => {
const res = await harness.app.inject({
method: 'POST',
url: `/api/sessions/${harness.ctx._sessionId}/input`,
payload: { input: 'x'.repeat(65537) },
});
expect(res.statusCode).toBe(400);
});
});
describe('POST /api/sessions/:id/resize', () => {
it('rejects cols exceeding max', async () => {
const res = await harness.app.inject({
method: 'POST',
url: `/api/sessions/${harness.ctx._sessionId}/resize`,
payload: { cols: 501, rows: 24 },
});
expect(res.statusCode).toBe(400);
});
});
});
```
### Key assertions to include
- **404 for unknown sessions**: Every `:id` endpoint must return 404 for nonexistent IDs
- **Input validation**: Bad paths, oversized inputs, invalid resize dimensions
- **Side effects**: Verify `ctx.broadcast()` was called with correct event type after mutations
- **Response shape**: Verify response bodies match expected API types
### Verification
```bash
npx vitest run test/routes/session-routes.test.ts
```
---
## Task 8: Add system + respawn routes tests
**Estimated effort**: 4 hours
**Depends on**: Task 6
**Files created**: `test/routes/system-routes.test.ts`, `test/routes/respawn-routes.test.ts`
### System routes (`src/web/routes/system-routes.ts`)
Focus on status and configuration endpoints:
| Method | Path | What to test |
|--------|------|-------------|
| `GET` | `/api/status` | Returns server status with uptime, session count |
| `GET` | `/api/stats` | Returns mux stats |
| `GET` | `/api/config` | Returns current config |
| `PUT` | `/api/config` | Updates config; validates input |
| `GET` | `/api/settings` | Returns user settings |
| `PUT` | `/api/settings` | Updates settings; validates input |
| `GET` | `/api/subagents` | Returns subagent list |
| `GET` | `/api/screenshots` | Returns screenshot list |
### Respawn routes (`src/web/routes/respawn-routes.ts`)
| Method | Path | What to test |
|--------|------|-------------|
| `GET` | `/api/sessions/:id/respawn` | Returns respawn status; null when not configured |
| `POST` | `/api/sessions/:id/respawn/start` | Starts respawn; 404 for unknown session |
| `POST` | `/api/sessions/:id/respawn/stop` | Stops respawn; 404 for unknown session |
| `PUT` | `/api/sessions/:id/respawn/config` | Updates respawn config; validates |
| `POST` | `/api/sessions/:id/respawn/enable` | Enables respawn loop |
| `POST` | `/api/sessions/:id/respawn/disable` | Disables respawn loop |
### Test patterns
Same pattern as Task 7 — `createRouteTestHarness` with `registerSystemRoutes` / `registerRespawnRoutes`.
For respawn tests, pre-populate `ctx.respawnControllers` with a mock controller in `beforeEach`:
```typescript
beforeEach(async () => {
harness = await createRouteTestHarness(registerRespawnRoutes);
// Add a mock respawn controller for the default session
harness.ctx.respawnControllers.set(harness.ctx._sessionId, {
getState: vi.fn(() => 'idle'),
getConfig: vi.fn(() => ({})),
getStatus: vi.fn(() => ({ state: 'idle', health: 100 })),
start: vi.fn(),
stop: vi.fn(),
updateConfig: vi.fn(),
enable: vi.fn(),
disable: vi.fn(),
});
});
```
### Verification
```bash
npx vitest run test/routes/system-routes.test.ts
npx vitest run test/routes/respawn-routes.test.ts
```
---
## Task 9: Slim down `respawn-test-utils.ts`
**Estimated effort**: 30 minutes
**Depends on**: Tasks 4, 5
**Files modified**: `test/respawn-test-utils.ts`
After Tasks 4–5 are verified passing with shared mocks, slim down `respawn-test-utils.ts` to remove duplicates.
### Steps
1. **Remove** from `respawn-test-utils.ts` what has been moved to shared mocks:
- `MockSession` class → now in `test/mocks/mock-session.ts`
- `createMockSession()` → now in `test/mocks/mock-session.ts`
- `terminalOutputs` → now in `test/mocks/mock-session.ts`
- `waitForEvent()` / `createDeferred()` → now in `test/mocks/test-helpers.ts`
2. **Keep** respawn-specific utilities that don't belong in the general mocks:
- `TimeController` / `createTimeController()` — respawn-specific timer control
- `MockAiIdleChecker` / `MockAiPlanChecker` — respawn-specific AI mocks
- `createStateTracker()` / `createEventRecorder()` — respawn state tracking
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — respawn config presets
- `waitForState()` — respawn state machine waiter
3. **Update imports** in `respawn-test-utils.ts` to re-use shared mocks:
```typescript
import { MockSession, createMockSession, terminalOutputs } from './mocks/index.js';
import { waitForEvent, createDeferred } from './mocks/index.js';
export { MockSession, createMockSession, terminalOutputs, waitForEvent, createDeferred };
```
This preserves backward compatibility for any future tests that import from `respawn-test-utils.ts` directly while eliminating the duplication.
### Verification
```bash
tsc --noEmit
npx vitest run test/respawn-controller.test.ts
npx vitest run test/respawn-team-awareness.test.ts
```
---
## What is NOT in scope (and why)
### Migrating `session-manager.test.ts` and `ralph-loop.test.ts` mocks
Both files define mocks inside `vi.mock()` factories that replace entire modules:
```typescript
// session-manager.test.ts — mock replaces ../src/session.js
vi.mock('../src/session.js', () => {
class MockSession extends EventEmitter { ... }
return { Session: MockSession };
});
// ralph-loop.test.ts — mock replaces ../src/state-store.js
vi.mock('../src/state-store.js', () => {
class MockStateStore { ... }
return { getStore: vi.fn(() => instance), StateStore: MockStateStore };
});
```
These are fundamentally different from the direct-instantiation pattern:
- The `vi.mock()` factory runs in an isolated scope — outer imports are not available
- The mock class must be returned with the exact export names (`Session`, `getStore`, `StateStore`)
- The `session-manager.test.ts` MockSession auto-registers into a shared `mockState.sessions` Map (tight coupling with test setup)
Migrating would require `vi.hoisted()` to share the class between factory and test scope, plus restructuring the test's module-mocking setup. This is high-complexity, high-risk refactoring with limited benefit since these tests already work. The shared `MockStateStore` in `test/mocks/` is available for **new** tests (like route tests) that use direct instantiation instead.
### Full integration tests with real Fastify server
Route tests use `app.inject()` which simulates HTTP without opening ports. Full integration tests that spin up `WebServer`, create real sessions, and stream SSE would be valuable but are a separate effort requiring:
- A test WebServer factory
- Session lifecycle management in tests
- SSE client test utilities
- Significantly more setup/teardown complexity
### Testing auth middleware in route tests
Route tests bypass authentication (no auth middleware registered on the test Fastify instance). Auth middleware has its own dedicated tests in `auth-security.test.ts` and `qr-auth.test.ts`. Testing auth + routes together is a future integration test concern.
### Testing SSE event streaming
SSE integration requires a running server with `EventSource` client. This is significantly more complex than `app.inject()` tests and is deferred. The existing `sse-events.test.ts` covers SSE patterns.
### Complete route coverage for all 12 modules
This phase covers the 3 highest-value route modules (session, system, respawn — 98 of 162 handlers). The remaining 9 modules (ralph, plan, push, team, mux, file, scheduled, hook-event, case) should be added incrementally in follow-up work.
---
## Summary
| Metric | Before | After |
|--------|--------|-------|
| MockSession definitions | 4 (across 4 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
| MockStateStore definitions | 2 (across 2 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
| Files importing from `respawn-test-utils.ts` | 0 | Utilities split into `test/mocks/` |
| Route test files | 0 | 3 (session, system, respawn) |
| Route handlers with dedicated tests | 0 | ~30 (highest-priority endpoints) |
| Shared mock directory | None | `test/mocks/` with 5 files + barrel |
### Final verification checklist
```bash
# Type checking
tsc --noEmit
# Linting
npm run lint
# Formatting
npm run format:check
# Run all affected tests individually
npx vitest run test/respawn-controller.test.ts
npx vitest run test/respawn-team-awareness.test.ts
npx vitest run test/routes/session-routes.test.ts
npx vitest run test/routes/system-routes.test.ts
npx vitest run test/routes/respawn-routes.test.ts
# Verify unchanged tests still pass
npx vitest run test/session-manager.test.ts
npx vitest run test/ralph-loop.test.ts
# Dev server still starts
npx tsx src/index.ts web --port 3099 &
curl -s http://localhost:3099/api/status | jq .status # "ok"
kill %1
```
-247
View File
@@ -1,247 +0,0 @@
# Ralph Loop Plan Improvement Roadmap
> Research-backed improvements for rock-solid AI planning with auto-improvement capabilities.
**Created**: 2026-01-27
**Status**: Implementation in Progress
---
## Table of Contents
1. [Research Summary](#research-summary)
2. [Current State Analysis](#current-state-analysis)
3. [Proposed Improvements](#proposed-improvements)
4. [Implementation Plan](#implementation-plan)
5. [Sources](#sources)
---
## Research Summary
### Key Insights from Industry Best Practices
#### 1. Self-Verification is Critical
> "Claude performs dramatically better when it can verify its own work, like run tests, compare screenshots, and validate outputs. Without clear success criteria, it might produce something that looks right but actually doesn't work."
> — [Anthropic Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
#### 2. Iterative Refinement Patterns (AWS)
> "A generator agent produces output, an evaluator agent reviews using evaluation rubric, and based on feedback, an optimizer agent revises the output. Loop repeats until criteria met."
> — [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
#### 3. Dynamic Task Decomposition (TDAG Framework)
> "Dynamically decomposes complex tasks into smaller subtasks and assigns each to a specifically generated subagent, enhancing adaptability in diverse and unpredictable real-world tasks."
> — [TDAG Framework - arXiv](https://arxiv.org/abs/2402.10178)
#### 4. Multi-Stage Verification Workflow
> "o3: Generate plan → Sonnet: Verify and create task list → Sonnet: Execute → Sonnet: Verify against plan → o3: Final verification → Issues bake back into plan"
> — [Claude Code Best Practices Community](https://rosmur.github.io/claudecode-best-practices/)
#### 5. Self-Improving Agents
> "Through an iterative refinement process (analyze outcome → adjust approach → try again), the agent becomes more adept at handling tasks over time. It effectively builds a growing knowledge base of what strategies work best."
> — [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
#### 6. Memory Architecture for Planning
> "Agents use three memory layers: working memory for short-lived calculations, episodic memory for step-by-step histories, and semantic memory for long-term knowledge."
> — [LLM Agent Research](https://www.promptingguide.ai/research/llm-agents)
---
## Current State Analysis
### What We Have
The current plan generation system (`/api/generate-plan` and `/api/generate-plan-detailed`):
1. **Standard Mode**: Single Opus 4.5 call with TDD-focused prompt
2. **Enhanced Mode**: 4 parallel subagents (Requirements, Architecture, Testing, Risks) + Verification
### Current Plan Item Structure
```json
{
"content": "Implement login endpoint",
"priority": "P0"
}
```
### Limitations
| Issue | Impact |
|-------|--------|
| No verification criteria | Can't automatically validate completion |
| No test pairing | TDD not enforced structurally |
| Static plans | No adaptation during execution |
| No dependencies | Can't track blocking relationships |
| No failure tracking | Same errors repeat |
| No checkpoints | Plans run until completion or failure |
---
## Proposed Improvements
### Enhanced Plan Item Structure
```typescript
interface EnhancedPlanItem {
id: string; // Unique identifier (e.g., "P0-001")
content: string; // Task description
priority: 'P0' | 'P1' | 'P2'; // Criticality
phase: 'setup' | 'test' | 'impl' | 'verify'; // Development phase
// NEW: Verification
verificationCriteria: string; // How to know it's done
testCommand?: string; // Command to run for verification
// NEW: Dependencies
dependencies: string[]; // IDs of tasks that must complete first
blockedBy?: string[]; // Runtime: tasks blocking this one
// NEW: Execution tracking
status: 'pending' | 'in_progress' | 'completed' | 'failed' | 'blocked';
attempts: number; // How many times attempted
lastError?: string; // Most recent failure reason
completedAt?: number; // Timestamp of completion
// NEW: Metadata
estimatedComplexity: 'low' | 'medium' | 'high';
rollbackStrategy?: string; // How to undo if needed
version: number; // Plan version this belongs to
}
```
### Runtime Plan Adaptation Flow
```
┌─────────────────────────────────────────────────────────────────┐
│ RUNTIME PLAN LOOP │
├─────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ Execute │──▶│ Verify │──▶│ Success? │──▶│ Mark │ │
│ │ Task │ │ Output │ │ │ │ Complete │ │
│ └──────────┘ └──────────┘ └────┬─────┘ └──────────┘ │
│ │ No │
│ ▼ │
│ ┌──────────┐ │
│ │ Analyze │ │
│ │ Failure │ │
│ └────┬─────┘ │
│ │ │
│ ┌──────────────┼──────────────┐ │
│ ▼ ▼ ▼ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ Retry │ │ Add Fix │ │ Escalate │ │
│ │ (< 3x) │ │ Sub-Task │ │ BLOCKED │ │
│ └──────────┘ └──────────┘ └──────────┘ │
│ │
└───────────────────────────────────────────────────────────────┘
```
### Checkpoint Review System
At iterations 5, 10, 20, 30, 50:
1. Pause execution
2. Summarize progress (completed/failed/pending)
3. Identify stuck items (3+ failures)
4. Generate alternative approaches for stuck items
5. Update plan with new strategies
6. Continue with refined plan
---
## Implementation Plan
### Phase 1: Quick Wins (Implementing Now)
#### 1.1 Add Verification Criteria to Plan Items
- Modify plan generation prompts to require `verificationCriteria`
- Update `PlanItem` interface in `types.ts`
- Update plan orchestrator prompts
#### 1.2 Pair Test/Implementation Steps
- Ensure every implementation step has a corresponding test step
- Group items: test → implement → verify
- Add phase field to track TDD cycle
#### 1.3 Checkpoint Review Prompts
- Add checkpoint logic to Ralph tracker
- At iterations 5, 10, 20: inject review prompt
- Generate progress summary and stuck item analysis
### Phase 2: Medium Effort (Implementing Now)
#### 2.1 Failure Tracking
- Track `attempts` and `lastError` per task
- After 3 failures, auto-generate debug sub-task
- Record failure patterns in plan history
#### 2.2 Plan Versioning
- Add `version` field to plans
- Keep history in `@fix_plan.md` with version markers
- Allow rollback to previous versions
- Track which version each task belongs to
#### 2.3 Dependency Tracking
- Add `dependencies` field to plan items
- Validate dependency graph (no cycles)
- Block tasks until dependencies complete
- Show dependency status in UI
### Phase 3: Future Enhancements
#### 3.1 Full Runtime Adaptation
- TDAG-style dynamic decomposition
- Auto-generate sub-tasks for complex items
- Learning from failure patterns
#### 3.2 Multi-Model Verification
- Haiku: Fast initial generation
- Sonnet: Verification and refinement
- Opus: Final quality check
#### 3.3 Plan Memory System
- Episodic memory: What worked/failed in this session
- Semantic memory: Patterns across projects
- Use for future plan generation
---
## File Changes Required
### New/Modified Files
| File | Changes |
|------|---------|
| `src/types.ts` | Add `EnhancedPlanItem` interface |
| `src/plan-orchestrator.ts` | Update prompts, add versioning |
| `src/ralph-tracker.ts` | Add checkpoint logic, failure tracking |
| `src/web/server.ts` | New endpoints for plan updates |
| `src/web/public/app.js` | UI for enhanced plan display |
### New Endpoints
| Method | Endpoint | Purpose |
|--------|----------|---------|
| PATCH | `/api/sessions/:id/plan/task/:taskId` | Update task status |
| POST | `/api/sessions/:id/plan/checkpoint` | Trigger checkpoint review |
| GET | `/api/sessions/:id/plan/history` | Get plan version history |
| POST | `/api/sessions/:id/plan/rollback/:version` | Rollback to version |
---
## Sources
- [Anthropic Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
- [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
- [TDAG: Multi-Agent Task Decomposition Framework](https://arxiv.org/abs/2402.10178)
- [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
- [OpenAI Self-Evolving Agents Cookbook](https://cookbook.openai.com/examples/partners/self_evolving_agents/autonomous_agent_retraining)
- [Task Decomposition for Coding Agents](https://mgx.dev/insights/task-decomposition-for-coding-agents-architectures-advancements-and-future-directions/)
- [Claude Code Best Practices Community Guide](https://rosmur.github.io/claudecode-best-practices/)
- [LLM Agents Prompt Engineering Guide](https://www.promptingguide.ai/research/llm-agents)
- [Agentic AI Implementation Guide](https://www.sketchdev.io/blog/agentic-ai-implementation-guide)
---
*This document is part of the Codeman project. See [CLAUDE.md](../CLAUDE.md) for main documentation.*
-251
View File
@@ -1,251 +0,0 @@
# Ralph Loop Improvements Plan
## Overview
This plan details improvements to Codeman's Ralph Loop system based on best practices from the Ralph Claude Code repository (https://github.com/frankbria/ralph-claude-code).
## Key Concepts to Implement
### RALPH_STATUS Block Format
Claude outputs this structured block at the end of every response for better tracking:
```
---RALPH_STATUS---
STATUS: IN_PROGRESS | COMPLETE | BLOCKED
TASKS_COMPLETED_THIS_LOOP: <number>
FILES_MODIFIED: <number>
TESTS_STATUS: PASSING | FAILING | NOT_RUN
WORK_TYPE: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING
EXIT_SIGNAL: false | true
RECOMMENDATION: <one line summary of what to do next>
---END_RALPH_STATUS---
```
### Dual-Condition Exit Gate
Exit requires BOTH conditions:
1. `completion_indicators >= 2` (heuristic detection from natural language patterns)
2. Claude's explicit `EXIT_SIGNAL: true` in the RALPH_STATUS block
### Circuit Breaker Pattern
Three states: CLOSED → HALF_OPEN → OPEN
| From State | Condition | To State |
|------------|-----------|----------|
| CLOSED | consecutive_no_progress >= 2 | HALF_OPEN |
| CLOSED | consecutive_no_progress >= 3 | OPEN |
| CLOSED | consecutive_same_error >= 5 | OPEN |
| HALF_OPEN | progress detected | CLOSED |
| HALF_OPEN | consecutive_no_progress >= 3 | OPEN |
| OPEN | Manual reset | CLOSED |
### @fix_plan.md Structure
```markdown
# Fix Plan
## High Priority (P0)
- [ ] Critical: Fix authentication bug
- [ ] Blocker: Database connection timeout
## Standard (P1)
- [ ] Feature: Add user profile page
## Nice to Have (P2)
- [ ] Improvement: Add dark mode
## Completed
- [x] Setup: Initialize project structure
```
---
## Phase 1: Quick Wins (1-2 days)
### 1.1 RALPH_STATUS Block Parsing
**What**: Add parsing support for the structured RALPH_STATUS block format in RalphTracker.
**Implementation**:
- Add regex pattern to detect `---RALPH_STATUS---` blocks
- Parse fields: STATUS, TASKS_COMPLETED_THIS_LOOP, FILES_MODIFIED, TESTS_STATUS, WORK_TYPE, EXIT_SIGNAL, RECOMMENDATION
- Store in extended `RalphTrackerState` type
- Emit new events: `ralphStatusUpdate`
**Files**: `ralph-tracker.ts`, `types.ts`
### 1.2 Enhanced Status Display in UI
**What**: Display RALPH_STATUS fields in the Ralph State Panel.
**Implementation**:
- Add UI elements: WORK_TYPE indicator, TESTS_STATUS badge, FILES_MODIFIED count
- Show RECOMMENDATION text in expanded view
- Color-code status (IN_PROGRESS=blue, COMPLETE=green, BLOCKED=red)
**Files**: `app.js`, `styles.css`, `index.html`
### 1.3 Prompt Template Improvements
**What**: Add specification-by-example exit scenarios to prompts.
**Implementation**:
- Add "Exit Scenarios" section to case-template.md
- Document when to continue vs. when to output completion
- Include testing limits guidance (max 20% effort on tests)
- Add RALPH_STATUS block instructions
**Files**: `case-template.md`
### 1.4 Better Wizard Validation
**What**: Add client-side validation and helpful warnings.
**Implementation**:
- Warn if task description < 50 chars
- Warn if no success criteria mentioned
- Suggest adding test requirements if none detected
- Validate completion phrase is uppercase alphanumeric
**Files**: `app.js`
---
## Phase 2: Core Improvements (3-5 days)
### 2.1 Circuit Breaker Pattern
**What**: Implement three-state circuit breaker to detect stuck loops.
**Implementation**:
- Create `CircuitBreaker` class with CLOSED, HALF_OPEN, OPEN states
- Track: files_modified, tasks_completed, error_patterns per iteration
- Triggers: N consecutive no-progress, same error M times, tests failing K iterations
- Emit events: `circuitBreakerStateChange`
**Files**: New `circuit-breaker.ts`, integrate into `ralph-tracker.ts`
### 2.2 Circuit Breaker UI
**What**: Visual indicator in Ralph panel.
**Implementation**:
- Badge: green (CLOSED), yellow (HALF_OPEN), red (OPEN)
- Warning before tripping
- Notification when circuit opens
- Manual reset button
**Files**: `app.js`, `styles.css`, `index.html`
### 2.3 @fix_plan.md Integration
**What**: Generate and track structured task plan file.
**Implementation**:
- Generate `@fix_plan.md` in working directory when loop starts
- Watch file for changes and sync with RalphTracker todos
- Parse priority levels (P0, P1, P2)
- Show priority in UI
**Files**: New `fix-plan.ts`, `ralph-tracker.ts`, `server.ts`
### 2.4 Wizard Plan Generation Step
**What**: Add third wizard step for AI-assisted plan generation.
**Implementation**:
- Step 2: "Plan Generation" between Task Setup and Launch
- Use Claude to break down task into fix plan items
- Allow edit/reorder before launch
- Generate @fix_plan.md with selected items
**Files**: `app.js`, `index.html`, `server.ts`
### 2.5 Smart Respawn Integration
**What**: Use RALPH_STATUS for respawn decisions.
**Implementation**:
- Use EXIT_SIGNAL field for respawn decisions
- If STATUS=BLOCKED, trigger circuit breaker instead of respawn
- Pass RECOMMENDATION to respawn update prompt
**Files**: `respawn-controller.ts`, `ralph-tracker.ts`
---
## Phase 3: Advanced Features (5+ days)
### 3.1 Template Library
- Bug Fix, Feature, Refactoring, Test Coverage, Documentation templates
- Template selector in wizard
- Custom templates in `~/.codeman/templates/`
### 3.2 Tool Permissions
- Configure allowed Claude tools per loop
- Generate hook configuration
- Store in session config
### 3.3 Per-Iteration Timeout
- Max time per iteration (5-60 min)
- Auto-continue on timeout
- Log timeout events
### 3.4 Rate Limiting
- Max tokens per iteration
- Max API calls per minute
- Cooldown between iterations
### 3.5 Metrics Dashboard
- Time-series charts (files modified, tasks completed, tokens)
- Aggregate statistics
- Export to JSON/CSV
---
## Priority Matrix
| Item | Effort | Impact | Priority |
|------|--------|--------|----------|
| 1.1 RALPH_STATUS Parsing | Low | High | P0 |
| 1.2 Status Display UI | Low | Medium | P0 |
| 1.3 Prompt Templates | Low | High | P0 |
| 1.4 Wizard Validation | Low | Medium | P1 |
| 2.1 Circuit Breaker | Medium | High | P1 |
| 2.2 Circuit Breaker UI | Medium | Medium | P1 |
| 2.3 Fix Plan Integration | Medium | High | P1 |
| 2.4 Plan Generation Step | Medium | Medium | P2 |
| 2.5 Respawn Integration | Medium | High | P1 |
| 3.1 Template Selection | High | Medium | P2 |
| 3.2 Tool Permissions | High | Medium | P3 |
| 3.3 Per-Iteration Timeout | High | Medium | P2 |
| 3.4 Rate Limiting | High | Low | P3 |
| 3.5 Metrics Dashboard | High | Medium | P3 |
---
## Reference: Ralph Claude Code Best Practices
### Testing Guidelines
- LIMIT testing to ~20% of total effort per loop
- PRIORITIZE: Implementation > Documentation > Tests
- Only write tests for NEW functionality
- Do NOT refactor existing tests unless broken
### What NOT to Do
- Do NOT continue with busy work when EXIT_SIGNAL should be true
- Do NOT run tests repeatedly without implementing new features
- Do NOT refactor code that is already working
- Do NOT add features not in specifications
- Do NOT forget the status block
### Exit Scenarios (Specification by Example)
1. **Successful Completion**: All tasks done → EXIT_SIGNAL=true
2. **Test-Only Loop**: No implementation, only testing → continue but warn
3. **Stuck on Error**: Same error 5 times → circuit breaker opens
4. **No Work Remaining**: All specs done → EXIT_SIGNAL=true
5. **Making Progress**: Normal flow → continue
6. **Blocked**: Needs human intervention → STATUS=BLOCKED
-434
View File
@@ -1,434 +0,0 @@
# Ralph Tracker Phase 1 Implementation Plan
## Overview
This plan details how to enhance the existing RalphTracker with RALPH_STATUS block parsing, circuit breaker pattern, and dual-condition exit gate.
---
## 1. Current State Analysis
### What RalphTracker Already Does Well
- **Todo Detection**: Supports 5 formats (checkboxes, indicators, status in parentheses, native TodoWrite, checkmark-based)
- **Completion Phrases**: Detects `<promise>PHRASE</promise>` with occurrence-based logic (1st = store, 2nd = complete)
- **Loop State Tracking**: Tracks active/inactive, iteration counts, max iterations, elapsed hours, cycle counts
- **Auto-Enable**: Disabled by default, auto-enables when Ralph patterns detected
- **Event System**: Emits `loopUpdate`, `todoUpdate`, `completionDetected`, `enabled` events
- **SSE Integration**: Events forwarded via `session:ralphLoopUpdate`, `session:ralphTodoUpdate`, `session:ralphCompletionDetected`
- **Debouncing**: EVENT_DEBOUNCE_MS (50ms) for rapid updates to prevent UI jitter
- **Cleanup**: MAX_TODO_ITEMS (50), TODO_EXPIRY_MS (1 hour), throttled cleanup
### Current Limitations
| Feature | Status |
|---------|--------|
| RALPH_STATUS block parsing | Missing |
| Circuit breaker pattern | Missing |
| Priority-based todos (P0/P1/P2) | Missing |
| Dual-condition exit gate | Missing |
| Files modified tracking | Missing |
| Tests status tracking | Missing |
| Work type classification | Missing |
---
## 2. New Type Definitions (types.ts)
```typescript
// ========== RALPH_STATUS Block Types ==========
export type RalphStatusValue = 'IN_PROGRESS' | 'COMPLETE' | 'BLOCKED';
export type RalphTestsStatus = 'PASSING' | 'FAILING' | 'NOT_RUN';
export type RalphWorkType = 'IMPLEMENTATION' | 'TESTING' | 'DOCUMENTATION' | 'REFACTORING';
/**
* Parsed RALPH_STATUS block from Claude output.
*/
export interface RalphStatusBlock {
status: RalphStatusValue;
tasksCompletedThisLoop: number;
filesModified: number;
testsStatus: RalphTestsStatus;
workType: RalphWorkType;
exitSignal: boolean;
recommendation: string;
parsedAt: number;
}
// ========== Circuit Breaker Types ==========
export type CircuitBreakerState = 'CLOSED' | 'HALF_OPEN' | 'OPEN';
export type CircuitBreakerReason =
| 'normal_operation'
| 'no_progress_warning'
| 'no_progress_open'
| 'same_error_repeated'
| 'tests_failing_too_long'
| 'progress_detected'
| 'manual_reset';
export interface CircuitBreakerStatus {
state: CircuitBreakerState;
consecutiveNoProgress: number;
consecutiveSameError: number;
consecutiveTestsFailure: number;
lastProgressIteration: number;
reason: string;
reasonCode: CircuitBreakerReason;
lastTransitionAt: number;
lastErrorMessage: string | null;
}
// ========== Priority Todo Types ==========
export type RalphTodoPriority = 'P0' | 'P1' | 'P2' | null;
// ========== Helper Functions ==========
export function createInitialCircuitBreakerStatus(): CircuitBreakerStatus {
return {
state: 'CLOSED',
consecutiveNoProgress: 0,
consecutiveSameError: 0,
consecutiveTestsFailure: 0,
lastProgressIteration: 0,
reason: 'Initial state',
reasonCode: 'normal_operation',
lastTransitionAt: Date.now(),
lastErrorMessage: null,
};
}
```
---
## 3. New Regex Patterns (ralph-tracker.ts)
```typescript
// ---------- RALPH_STATUS Block Patterns ----------
const RALPH_STATUS_START_PATTERN = /^---RALPH_STATUS---\s*$/;
const RALPH_STATUS_END_PATTERN = /^---END_RALPH_STATUS---\s*$/;
const RALPH_STATUS_FIELD_PATTERN = /^STATUS:\s*(IN_PROGRESS|COMPLETE|BLOCKED)\s*$/i;
const RALPH_TASKS_COMPLETED_PATTERN = /^TASKS_COMPLETED_THIS_LOOP:\s*(\d+)\s*$/i;
const RALPH_FILES_MODIFIED_PATTERN = /^FILES_MODIFIED:\s*(\d+)\s*$/i;
const RALPH_TESTS_STATUS_PATTERN = /^TESTS_STATUS:\s*(PASSING|FAILING|NOT_RUN)\s*$/i;
const RALPH_WORK_TYPE_PATTERN = /^WORK_TYPE:\s*(IMPLEMENTATION|TESTING|DOCUMENTATION|REFACTORING)\s*$/i;
const RALPH_EXIT_SIGNAL_PATTERN = /^EXIT_SIGNAL:\s*(true|false)\s*$/i;
const RALPH_RECOMMENDATION_PATTERN = /^RECOMMENDATION:\s*(.+)$/i;
// ---------- Completion Indicator Patterns ----------
const COMPLETION_INDICATOR_PATTERNS = [
/all\s+(?:tasks?|items?|work)\s+(?:are\s+)?(?:completed?|done|finished)/i,
/(?:completed?|finished)\s+all\s+(?:tasks?|items?|work)/i,
/nothing\s+(?:left|remaining)\s+to\s+do/i,
/no\s+more\s+(?:tasks?|items?|work)/i,
/everything\s+(?:is\s+)?(?:completed?|done)/i,
];
// ---------- Priority Pattern ----------
const TODO_PRIORITY_PATTERN = /^\s*(?:\[.\])?\s*(?:Critical:|Blocker:|Feature:|Improvement:)?\s*\(?(P[012])\)?:?\s*/i;
```
---
## 4. New State Properties (ralph-tracker.ts)
```typescript
// Add to RalphTracker class
// Circuit breaker state tracking
private _circuitBreaker: CircuitBreakerStatus;
// RALPH_STATUS block parsing state
private _statusBlockBuffer: string[] = [];
private _inStatusBlock: boolean = false;
private _lastStatusBlock: RalphStatusBlock | null = null;
// Dual-condition exit tracking
private _completionIndicators: number = 0;
private _exitGateMet: boolean = false;
// Cumulative tracking
private _totalFilesModified: number = 0;
private _totalTasksCompleted: number = 0;
```
---
## 5. New Methods to Implement
### 5.1 RALPH_STATUS Block Parsing
```typescript
private processStatusBlockLine(line: string): void {
const trimmed = line.trim();
if (RALPH_STATUS_START_PATTERN.test(trimmed)) {
this._inStatusBlock = true;
this._statusBlockBuffer = [];
return;
}
if (this._inStatusBlock && RALPH_STATUS_END_PATTERN.test(trimmed)) {
this._inStatusBlock = false;
this.parseStatusBlock(this._statusBlockBuffer);
this._statusBlockBuffer = [];
return;
}
if (this._inStatusBlock) {
this._statusBlockBuffer.push(trimmed);
}
}
private parseStatusBlock(lines: string[]): void {
const block: Partial<RalphStatusBlock> = { parsedAt: Date.now() };
for (const line of lines) {
// Parse each field...
}
if (block.status !== undefined) {
this._lastStatusBlock = fullBlock;
this.handleStatusBlock(fullBlock);
}
}
private handleStatusBlock(block: RalphStatusBlock): void {
this._totalFilesModified += block.filesModified;
this._totalTasksCompleted += block.tasksCompletedThisLoop;
const hasProgress = block.filesModified > 0 || block.tasksCompletedThisLoop > 0;
this.updateCircuitBreaker(hasProgress, block.testsStatus, block.status);
if (block.status === 'COMPLETE') {
this._completionIndicators++;
}
if (block.exitSignal && this._completionIndicators >= 2) {
this._exitGateMet = true;
this.emit('exitGateMet', { completionIndicators: this._completionIndicators, exitSignal: true });
}
this.emit('statusBlockDetected', block);
}
```
### 5.2 Circuit Breaker Logic
```typescript
private updateCircuitBreaker(
hasProgress: boolean,
testsStatus: RalphTestsStatus,
status: RalphStatusValue
): void {
const prevState = this._circuitBreaker.state;
if (hasProgress) {
this._circuitBreaker.consecutiveNoProgress = 0;
this._circuitBreaker.lastProgressIteration = this._loopState.cycleCount;
if (this._circuitBreaker.state === 'HALF_OPEN') {
this._circuitBreaker.state = 'CLOSED';
this._circuitBreaker.reasonCode = 'progress_detected';
}
} else {
this._circuitBreaker.consecutiveNoProgress++;
if (this._circuitBreaker.state === 'CLOSED') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reasonCode = 'no_progress_open';
} else if (this._circuitBreaker.consecutiveNoProgress >= 2) {
this._circuitBreaker.state = 'HALF_OPEN';
this._circuitBreaker.reasonCode = 'no_progress_warning';
}
}
}
if (prevState !== this._circuitBreaker.state) {
this._circuitBreaker.lastTransitionAt = Date.now();
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
}
}
resetCircuitBreaker(): void {
this._circuitBreaker = createInitialCircuitBreakerStatus();
this._circuitBreaker.reasonCode = 'manual_reset';
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
}
```
### 5.3 Update processLine Method
```typescript
private processLine(line: string): void {
const trimmed = line.trim();
if (!trimmed) return;
// NEW: Check for RALPH_STATUS block
this.processStatusBlockLine(trimmed);
// NEW: Check for completion indicators
this.detectCompletionIndicators(trimmed);
// EXISTING: Rest of the detection methods...
this.detectCompletionPhrase(trimmed);
this.detectAllTasksComplete(trimmed);
this.detectTaskCompletion(trimmed);
this.detectLoopStatus(trimmed);
this.detectTodoItems(trimmed);
}
```
---
## 6. New Events to Add
```typescript
export interface RalphTrackerEvents {
// Existing events
loopUpdate: (state: RalphTrackerState) => void;
todoUpdate: (todos: RalphTodoItem[]) => void;
completionDetected: (phrase: string) => void;
enabled: () => void;
// New events
statusBlockDetected: (block: RalphStatusBlock) => void;
circuitBreakerUpdate: (status: CircuitBreakerStatus) => void;
exitGateMet: (data: { completionIndicators: number; exitSignal: boolean }) => void;
}
```
---
## 7. Server Integration (server.ts)
```typescript
// Add new SSE event handlers in setupSessionListeners()
session.on('ralphStatusBlockDetected', (block: RalphStatusBlock) => {
this.broadcast('session:ralphStatusUpdate', { sessionId: session.id, block });
});
session.on('ralphCircuitBreakerUpdate', (status: CircuitBreakerStatus) => {
this.broadcast('session:circuitBreakerUpdate', { sessionId: session.id, status });
});
session.on('ralphExitGateMet', (data) => {
this.broadcast('session:exitGateMet', { sessionId: session.id, ...data });
});
// Add API endpoint for circuit breaker reset
this.app.post('/api/sessions/:id/ralph-circuit-breaker/reset', async (req) => {
const session = this.sessions.get(req.params.id);
if (!session) return { success: false, error: 'Session not found' };
session.ralphTracker?.resetCircuitBreaker();
return { success: true };
});
```
---
## 8. Frontend Changes (app.js)
### New SSE Event Listeners
```javascript
this.eventSource.addEventListener('session:ralphStatusUpdate', (e) => {
const data = JSON.parse(e.data);
this.updateRalphStatusBlock(data.sessionId, data.block);
});
this.eventSource.addEventListener('session:circuitBreakerUpdate', (e) => {
const data = JSON.parse(e.data);
this.updateCircuitBreaker(data.sessionId, data.status);
});
```
### New Rendering Methods
```javascript
updateRalphStatusBlock(sessionId, block) {
// Store and render status block
}
renderRalphStatusBlock(block) {
// Render STATUS, WORK_TYPE, TESTS_STATUS, RECOMMENDATION
}
updateCircuitBreaker(sessionId, status) {
// Store and render circuit breaker state
}
renderCircuitBreaker(status) {
// Render badge: green (CLOSED), yellow (HALF_OPEN), red (OPEN)
}
```
---
## 9. Implementation Order
| Step | Task | Time |
|------|------|------|
| 1 | Add type definitions to `types.ts` | 30 min |
| 2 | Add regex patterns to `ralph-tracker.ts` | 30 min |
| 3 | Add state properties to RalphTracker class | 15 min |
| 4 | Implement RALPH_STATUS parsing methods | 1.5 hr |
| 5 | Implement circuit breaker logic | 1 hr |
| 6 | Implement completion indicators | 30 min |
| 7 | Update events interface | 15 min |
| 8 | Add server SSE handlers and API endpoint | 45 min |
| 9 | Add frontend event listeners and rendering | 1 hr |
| 10 | Add CSS styles | 30 min |
| 11 | Update HTML structure | 15 min |
| 12 | Write unit tests | 1.5 hr |
**Total: ~8 hours**
---
## 10. Files to Modify
| File | Changes |
|------|---------|
| `src/types.ts` | Add RalphStatusBlock, CircuitBreakerStatus, helper functions |
| `src/ralph-tracker.ts` | Add patterns, state, parsing methods, circuit breaker |
| `src/web/server.ts` | Add SSE handlers, circuit breaker reset endpoint |
| `src/web/public/app.js` | Add event listeners, rendering methods |
| `src/web/public/styles.css` | Add status block and circuit breaker styles |
| `src/web/public/index.html` | Add UI elements to Ralph panel |
| `test/ralph-tracker.test.ts` | Add tests for new functionality |
---
## 11. Test Cases to Add
1. **RALPH_STATUS Parsing**
- Parse valid status block with all fields
- Parse block with missing optional fields
- Ignore malformed blocks
- Handle multiple blocks in sequence
2. **Circuit Breaker State Transitions**
- CLOSED → HALF_OPEN on 2 no-progress
- HALF_OPEN → OPEN on 3 no-progress
- HALF_OPEN → CLOSED on progress
- Manual reset from OPEN
3. **Dual-Condition Exit Gate**
- Exit when indicators >= 2 AND exitSignal = true
- No exit when indicators >= 2 but exitSignal = false
- No exit when exitSignal = true but indicators < 2
4. **Integration Tests**
- SSE events broadcast correctly
- UI updates on status block detection
- Circuit breaker badge updates
-385
View File
@@ -1,385 +0,0 @@
# Respawn Controller Idle Detection Improvement Plan
## Executive Summary
The current respawn controller relies primarily on **parsing terminal output** to detect idle states. This approach is fragile and leads to false positives/negatives (e.g., the w3-reddit-analyse session).
**Key insight**: Claude Code provides **direct, authoritative signals** via hooks and files that definitively indicate session state. We're receiving some of these signals but not using them for idle detection!
---
## Current Detection Layers (What We Have)
| Layer | Signal | Source | Reliability |
|-------|--------|--------|-------------|
| 1 | Completion message ("Worked for Xm Xs") | Terminal parsing | Medium - can miss edge cases |
| 2 | Output silence (configurable duration) | Terminal activity | Low - Claude can be processing silently |
| 3 | Token stability | Terminal parsing | Low - tokens don't change during I/O waits |
| 4 | Working pattern absence | Terminal parsing | Medium - patterns can be missed |
| 5 | AI idle check | Spawned Claude CLI | High but slow (90s timeout) |
**Problem**: All layers depend on **parsing terminal output**, which is inherently unreliable.
---
## Available Claude Code Signals (Not Fully Utilized)
### 1. `Stop` Hook ⭐ CRITICAL - DEFINITIVE SIGNAL
**What it is**: Fires when the main Claude Code agent **finishes responding**.
**From docs**: "Runs when the main Claude Code agent has finished responding. Does not run if the stoppage occurred due to a user interrupt."
**Current status**: We receive it via `/api/hook-event` but **don't use it for idle detection**!
**Input received**:
```json
{
"session_id": "abc123",
"transcript_path": "~/.claude/projects/.../00893aaf.jsonl",
"hook_event_name": "Stop",
"stop_hook_active": true // Important for preventing loops
}
```
**Action needed**: The `Stop` hook should be the **PRIMARY** idle detection signal. When Claude fires Stop, the agent has definitively finished its response cycle.
### 2. `idle_prompt` Notification ⭐ HIGH VALUE
**What it is**: Fires after **60+ seconds of idle time** when Claude is waiting for user input.
**From docs**: "When Claude is waiting for user input (after 60+ seconds of idle time)"
**Current status**: We receive it but only forward it to the UI for notification display.
**Action needed**: Use `idle_prompt` as a **definitive confirmation** that Claude is idle. If we receive this, there's no need for AI idle checks or output silence timers.
### 3. Transcript JSONL File ⭐ HIGH VALUE
**What it is**: Complete conversation history at `~/.claude/projects/{project-hash}/{session-id}.jsonl`
**Current status**: We already watch subagent transcripts but **not the main session transcript**.
**Data available**:
- Every message (user, assistant, system)
- Every tool call with inputs/outputs
- Progress events
- Structured, parseable JSON
**Action needed**:
- Monitor the main transcript file (path provided in every hook input)
- Parse the last few entries to detect:
- Tool completion
- Assistant message completion
- Error states
- Plan mode prompts
### 4. `PostToolUse` Hook - Tool Completion Tracking
**What it is**: Fires immediately after any tool completes successfully.
**Use case**: Track exactly when tools finish to understand execution flow.
**Current status**: Not implemented.
**Action needed**: Add PostToolUse hooks to track tool completion events.
### 5. `SubagentStop` Hook - Background Agent Completion
**What it is**: Fires when a subagent (Task tool) finishes responding.
**Current status**: Not implemented in hooks config (we watch JSONL files separately).
**Action needed**: Add to hooks config for redundant detection.
### 6. `permission_prompt` and `elicitation_dialog` - Blocking State Detection
**What it is**: Fires when Claude needs user input (permission or question).
**Current status**: We receive and use for auto-accept blocking.
**Enhancement**: Use as definitive "Claude is NOT idle - it's waiting for user action".
---
## Proposed Architecture: Multi-Signal Idle Detection
### New Detection Hierarchy
```
Priority 1 (Definitive):
└── Stop hook received → CONFIRMED IDLE
└── idle_prompt received → CONFIRMED IDLE (60s+ idle)
Priority 2 (Blocking):
└── permission_prompt received → NOT IDLE (waiting for permission)
└── elicitation_dialog received → NOT IDLE (waiting for answer)
└── Working patterns in terminal → NOT IDLE
Priority 3 (Supporting):
└── Transcript analysis → Check last entries for completion
└── Output silence + token stability → Weak idle signal
Priority 4 (Fallback):
└── AI idle check → Only if no definitive signals after timeout
```
### State Machine Changes
```
┌─────────────────────────────────────┐
│ │
▼ │
┌─────────────────────┐ │
│ WATCHING │◄──────────────────────────────┤
└─────────────────────┘ │
│ │ │
│ │ Stop hook or idle_prompt │
│ └────────────────────────┐ │
│ ▼ │
│ Output silence ┌────────────┐ │
│ (no definitive signals) │ HOOK_IDLE │───────┤
│ └────────────┘ │
▼ (skip AI check) │
┌────────────────────┐ │
│ CONFIRMING_IDLE │ │
└────────────────────┘ │
│ │
│ Silence confirmed │
▼ │
┌────────────────────┐ │
│ AI_CHECKING │──── IDLE verdict ─────────────┤
└────────────────────┘ │
│ │
│ WORKING verdict │
└─────────────────────────────────────────────┘
```
### New State: `hook_idle`
When a definitive hook signal is received:
1. Skip AI idle check entirely (saves time and API calls)
2. Short confirmation period (2-3s) to handle race conditions
3. Proceed directly to respawn sequence
---
## Implementation Plan
### Phase 1: Use Stop Hook for Idle Detection ✅ COMPLETED
**Files modified**:
- `src/respawn-controller.ts`
- `src/web/server.ts`
- `test/respawn-controller.test.ts`
**Changes implemented**:
1. Added `stopHookReceived`, `stopHookTime`, `idlePromptReceived`, `idlePromptTime` fields to `DetectionStatus`
2. Added `hookConfirmTimer` for short confirmation after hook signal (3s)
3. Added `signalStopHook()` method:
- Sets `stopHookReceived = true` and timestamp
- Cancels any running AI check (hook is definitive)
- Starts 3s confirmation timer
- If no new output during confirmation → triggers respawn cycle
4. Added `signalIdlePrompt()` method:
- Sets `idlePromptReceived = true` and timestamp
- Immediately confirms idle (skips confirmation timer - 60s+ already proven)
5. Added `resetHookState()` to clear hook flags on:
- Controller start
- Working patterns detected
- Cycle completion
6. Updated server.ts `/api/hook-event` endpoint to call:
- `controller.signalStopHook()` for `stop` events
- `controller.signalIdlePrompt()` for `idle_prompt` events
7. Updated `getDetectionStatus()`:
- Returns hook signal states
- Sets confidence to 100% when hook received
- Updates statusText to show hook status
**Tests added** (9 new tests in `RespawnController Hook-Based Idle Detection` describe block):
- `should expose signalStopHook method`
- `should expose signalIdlePrompt method`
- `should set stopHookReceived in detection status when Stop hook signaled`
- `should include hook status in statusText when Stop hook received`
- `should trigger respawn cycle after Stop hook confirmation`
- `should immediately confirm idle when idle_prompt signaled (skip confirmation)`
- `should cancel Stop hook confirmation if working patterns detected`
- `should ignore Stop hook when not in watching state`
- `should have 100% confidence when hook signal is received`
**Detection status update** (implemented):
```typescript
interface DetectionStatus {
/** Layer 0: Stop hook received (highest priority - definitive signal) */
stopHookReceived: boolean;
stopHookTime: number | null;
/** Layer 0: idle_prompt notification received (definitive signal) */
idlePromptReceived: boolean;
idlePromptTime: number | null;
// Existing fields...
}
```
### Phase 2: Use idle_prompt for Definitive Idle ✅ COMPLETED (in Phase 1)
**Already implemented in Phase 1**:
1. `signalIdlePrompt()` method sets `idlePromptReceived = true`
2. Immediately calls `onIdleConfirmed()` - skips all other detection
3. Server.ts calls `controller.signalIdlePrompt()` when `idle_prompt` event received
4. 60s+ of Claude waiting = definitive idle signal
### Phase 3: Transcript File Monitoring ✅ COMPLETED
**New file**: `src/transcript-watcher.ts`
**Functionality implemented**:
1. Watch the session's transcript JSONL file using `fs.watch()`
2. Parse new entries as they're appended (incremental reading from last position)
3. Detect:
- `result` entry → `transcript:complete` event (isComplete = true)
- `tool_use` content block → `transcript:tool_start` event
- `tool_result` content block → `transcript:tool_end` event
- `AskUserQuestion` or `ExitPlanMode` tools → `transcript:plan_mode` event
- Error conditions in result entries
4. Emit structured events consumed by respawn controller
**Integration implemented**:
- `transcript_path` added to allowed hook data fields in `sanitizeHookData()`
- `transcriptWatchers` Map added to WebServer for per-session watchers
- `startTranscriptWatcher()` creates watcher and wires up events:
- `transcript:complete` → `controller.signalTranscriptComplete()`
- `transcript:plan_mode` → `controller.signalTranscriptPlanMode()`
- `stopTranscriptWatcher()` cleans up on session cleanup
- Hook events with `transcript_path` automatically start watching
**RespawnController methods added**:
- `signalTranscriptComplete()` - Supporting signal that can accelerate idle detection
- `signalTranscriptPlanMode()` - Cancels auto-accept timer (like elicitation)
**Tests added** (13 tests in `test/transcript-watcher.test.ts`):
- Initialization tests
- File watching tests (existing file, non-existent file, stop, updatePath)
- Entry processing tests (user entry, result entry, tool execution, plan mode, errors)
- State management tests
### Phase 4: Enhanced Hook Configuration
**Update `src/hooks-config.ts`**:
```typescript
export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
return {
hooks: {
Notification: [
{ matcher: 'idle_prompt', hooks: [...] },
{ matcher: 'permission_prompt', hooks: [...] },
{ matcher: 'elicitation_dialog', hooks: [...] },
],
Stop: [{ hooks: [...] }],
// NEW: Add these
PostToolUse: [
{ matcher: '*', hooks: [...] } // Track all tool completions
],
SubagentStop: [{ hooks: [...] }],
PreCompact: [
{ matcher: '*', hooks: [...] } // Track compaction
],
},
};
}
```
### Phase 5: Confidence Scoring Overhaul
Replace current confidence calculation with weighted signals:
```typescript
function calculateConfidence(): number {
let confidence = 0;
// Definitive signals (100% confidence)
if (this.stopHookReceived) confidence = 100;
if (this.idlePromptReceived) confidence = 100;
// Blocking signals (0% confidence)
if (this.permissionPromptReceived) return 0;
if (this.elicitationReceived) return 0;
if (this.workingPatternRecent) return 0;
// Supporting signals (build up to ~80%)
if (confidence < 100) {
if (this.outputSilent) confidence += 30;
if (this.tokensStable) confidence += 20;
if (this.transcriptShowsCompletion) confidence += 30;
}
return Math.min(100, confidence);
}
```
---
## Expected Benefits
| Metric | Current | After Implementation |
|--------|---------|---------------------|
| False positive rate | ~15-20% | <5% |
| Detection latency | 10-90s (AI check) | 3-5s (hook-based) |
| API calls for AI check | Every idle detection | Only when hooks unavailable |
| Reliability | Medium | High (definitive signals) |
---
## Testing Strategy
### Unit Tests
1. `Stop` hook triggers immediate idle confirmation
2. `idle_prompt` skips all other detection
3. `permission_prompt` blocks idle detection
4. Transcript parsing correctly identifies completion
5. Fallback to AI check when no hooks received
### Integration Tests
1. End-to-end with real Claude session
2. Hook event delivery and handling
3. Transcript file monitoring
4. Race condition handling
### Scenarios to Test
1. Normal completion → Stop hook → respawn
2. Long-running task → idle_prompt → respawn
3. Permission needed → wait for user action
4. AskUserQuestion → wait for user answer
5. Plan mode → auto-accept → continue
6. Hooks disabled/unavailable → fallback to AI check
---
## Migration Path
1. **Implement Phase 1** - Stop hook detection (low risk, high value)
2. **Deploy and monitor** - Verify Stop hooks are reliable
3. **Implement Phase 2** - idle_prompt (simple addition)
4. **Implement Phase 3** - Transcript monitoring (more complex)
5. **Implement Phase 4** - Enhanced hooks (optional, for completeness)
6. **Implement Phase 5** - Refactor confidence scoring
---
## Open Questions
1. **Stop hook reliability**: Does it fire 100% of the time? Edge cases?
2. **Transcript file location**: Always at the path in hook input?
3. **Hook delivery latency**: How quickly do hooks fire after state change?
4. **Race conditions**: What if Stop hook and new work happen simultaneously?
---
## References
- [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)
- [Agent SDK Documentation](https://platform.claude.com/docs/en/agent-sdk/overview)
- Current implementation: `src/respawn-controller.ts`
- Hooks config: `src/hooks-config.ts`
- Subagent watcher: `src/subagent-watcher.ts`
-172
View File
@@ -1,172 +0,0 @@
# Run Summary Feature - Implementation Plan
## Overview
The Run Summary feature provides users with a consolidated view of what happened in their session while they were away. It tracks significant events, issues, and statistics, presenting them in an easy-to-digest format.
## Data Structures
### RunSummaryEventType
```typescript
type RunSummaryEventType =
| 'session_started'
| 'session_stopped'
| 'respawn_cycle_started'
| 'respawn_cycle_completed'
| 'respawn_state_change'
| 'error'
| 'warning'
| 'token_milestone'
| 'auto_compact'
| 'auto_clear'
| 'idle_detected'
| 'working_detected'
| 'ralph_completion'
| 'ai_check_result'
| 'hook_event'
| 'state_stuck';
```
### RunSummaryEvent
```typescript
interface RunSummaryEvent {
id: string;
timestamp: number;
type: RunSummaryEventType;
severity: 'info' | 'warning' | 'error' | 'success';
title: string;
details?: string;
metadata?: Record<string, unknown>;
}
```
### RunSummary
```typescript
interface RunSummary {
sessionId: string;
sessionName: string;
startedAt: number;
lastUpdatedAt: number;
events: RunSummaryEvent[];
stats: {
totalRespawnCycles: number;
totalTokensUsed: number;
peakTokens: number;
totalTimeActiveMs: number;
totalTimeIdleMs: number;
errorCount: number;
warningCount: number;
aiCheckCount: number;
lastIdleAt: number | null;
lastWorkingAt: number | null;
stateTransitions: number;
};
}
```
## Files to Create/Modify
### 1. `src/run-summary.ts` (NEW)
- `RunSummaryTracker` class
- Event tracking and aggregation
- Statistics calculation
- Max 1000 events per session (FIFO trimming)
### 2. `src/types.ts` (MODIFY)
- Add `RunSummaryEvent`, `RunSummaryEventType`, `RunSummary` interfaces
- Add `RunSummaryEventSeverity` type
### 3. `src/web/server.ts` (MODIFY)
- Create `RunSummaryTracker` per session
- Subscribe to session events and forward to tracker
- Subscribe to respawn controller events
- Add API endpoint: `GET /api/sessions/:id/run-summary`
- Broadcast `session:runSummaryUpdate` SSE event
### 4. `src/web/public/app.js` (MODIFY)
- Add "Run Summary" button to session header
- Create modal to display summary
- Handle `session:runSummaryUpdate` SSE event
- Timeline view for events
- Stats cards at top
### 5. `src/web/public/index.html` (MODIFY)
- Add modal HTML structure for run summary
### 6. `src/web/public/styles.css` (MODIFY)
- Styles for run summary modal and timeline
## Event Sources
| Event Type | Source | Trigger |
|------------|--------|---------|
| session_started | Session | `startInteractive()` / `startShell()` |
| session_stopped | Session | `stop()` |
| respawn_cycle_started | RespawnController | State → `sending_update` |
| respawn_cycle_completed | RespawnController | State → `watching` (after cycle) |
| respawn_state_change | RespawnController | Any state transition |
| error | Various | Errors caught in try/catch |
| warning | RunSummaryTracker | State stuck > 5min, high tokens |
| token_milestone | Session | Every 50k tokens |
| auto_compact | Session | `autoCompact` event |
| auto_clear | Session | `autoClear` event |
| idle_detected | Session | `idle` event |
| working_detected | Session | `working` event |
| ralph_completion | RalphTracker | `completionDetected` event |
| ai_check_result | RespawnController | AI check completes |
| hook_event | Server | `/api/hook-event` endpoint |
| state_stuck | RunSummaryTracker | Same state > 10min |
## API Endpoint
### GET /api/sessions/:id/run-summary
Returns the full run summary for a session.
Response:
```json
{
"success": true,
"summary": {
"sessionId": "...",
"sessionName": "...",
"startedAt": 1234567890,
"lastUpdatedAt": 1234567890,
"events": [...],
"stats": {...}
}
}
```
## UI Design
### Summary Modal
- Header: Session name, duration, status indicator
- Stats Cards Row:
- Respawn Cycles: count
- Tokens Used: peak / current
- Active Time: formatted duration
- Issues: errors + warnings count
- Timeline:
- Vertical timeline of events
- Color-coded by severity (green=success, blue=info, yellow=warning, red=error)
- Expandable details
- Filter by event type
- Footer: "Close" button
## Implementation Steps
1. Add types to `types.ts`
2. Create `run-summary.ts` with RunSummaryTracker class
3. Integrate tracker with server.ts (create per session, wire events)
4. Add API endpoint
5. Add frontend modal and button
6. Test with live session
## Storage
Run summaries are kept in memory only (not persisted to disk) since:
- They're session-specific and regenerated on session start
- Persisting thousands of events would bloat state.json
- Server restart = fresh session anyway
If persistence is needed later, could add to `state-inner.json` with per-session limits.
@@ -1,387 +0,0 @@
# Codeman TypeScript Improvement Suggestions
**Generated**: February 2026
**Based on**: Research into TypeScript best practices (2024-2025) and codebase analysis
---
## 🔴 High Priority (Low effort, high impact)
### 1. Use the Already-Installed Zod for API Validation
Zod v4.3.6 is in `package.json` but **never imported**. API routes use unsafe type assertions:
```typescript
// Current (unsafe)
const body = req.body as CreateSessionRequest;
// Recommended
const result = CreateSessionSchema.safeParse(req.body);
if (!result.success) return createErrorResponse(ApiErrorCode.INVALID_INPUT, ...);
```
**Impact**: Prevents runtime errors from malformed client requests.
**Files to update**: `src/web/server.ts` (all POST/PUT routes)
---
### 2. Add `assertNever` for Exhaustive Switch Checking
Switch statements on union types (e.g., `respawn-controller.ts:1072`, `ralph-tracker.ts:2088`) lack exhaustive checking. Adding new union members won't cause compile errors.
```typescript
// Add to src/utils/type-safety.ts
export function assertNever(x: never, message?: string): never {
throw new Error(message ?? `Unexpected value: ${JSON.stringify(x)}`);
}
// Usage in switch statements
switch (status) {
case 'idle': return handleIdle();
case 'busy': return handleBusy();
case 'stopped': return handleStopped();
case 'error': return handleError();
default: return assertNever(status);
}
```
**Impact**: Compile-time guarantee all cases are handled.
**Files affected**: `respawn-controller.ts`, `ralph-tracker.ts`, any file with switch on union types
---
### 3. Standardize `createErrorResponse` Usage
Currently only used in 2 files despite being a good pattern. Many routes still use ad-hoc error responses.
**Impact**: Consistent API error format across all endpoints.
---
## 🟡 Medium Priority (Medium effort, significant benefit)
### 4. Convert `ApiResponse<T>` to Discriminated Union
Current interface has optional properties; discriminated union enables better narrowing:
```typescript
// Current (types.ts)
interface ApiResponse<T> { success: boolean; error?: string; data?: T; }
// Better
type ApiResponse<T> =
| { success: true; data: T }
| { success: false; error: string; errorCode: ApiErrorCode };
// Usage with exhaustive checking
function handleResponse<T>(response: ApiResponse<T>): T {
if (response.success) {
return response.data; // TypeScript knows data exists
} else {
throw new Error(response.error); // TypeScript knows error exists
}
}
```
---
### 5. Add Branded Types for Token Counts
Prevents mixing input/output tokens in calculations:
```typescript
// src/types/branded.ts
type Brand<K, T extends string> = K & { readonly __brand: T };
export type InputTokens = Brand<number, 'InputTokens'>;
export type OutputTokens = Brand<number, 'OutputTokens'>;
export type TokenCount = Brand<number, 'TokenCount'>;
export type Milliseconds = Brand<number, 'Milliseconds'>;
// Constructor functions
export function inputTokens(value: number): InputTokens {
if (value < 0) throw new Error('Token count cannot be negative');
return value as InputTokens;
}
```
**Use cases**:
- Token counts (`_totalInputTokens`, `_totalOutputTokens`)
- Timeout values (`idleTimeoutMs`, `completionConfirmMs`, `noOutputTimeoutMs`)
- IDs (`SessionId`, `TaskId`, `CycleId`)
---
### 6. Dependency Injection for Core Services
Replace hidden singleton dependencies with constructor injection for better testability:
```typescript
// Current: Hidden dependencies
export class RalphLoop extends EventEmitter {
constructor() {
this.sessionManager = getSessionManager();
this.store = getStore();
}
}
// Better: Explicit dependencies
export interface RalphLoopDeps {
sessionManager: SessionManager;
taskQueue: TaskQueue;
store: StateStore;
}
export class RalphLoop extends EventEmitter {
constructor(deps: RalphLoopDeps, options?: RalphLoopOptions) {
this.sessionManager = deps.sessionManager;
// ...
}
}
// Production factory
export function createRalphLoop(options?: RalphLoopOptions): RalphLoop {
return new RalphLoop({
sessionManager: getSessionManager(),
taskQueue: getTaskQueue(),
store: getStore(),
}, options);
}
```
**Start with**: `RalphLoop` (has the most dependencies)
**Benefits**: Easier testing, explicit dependencies, SOLID compliance
---
### 7. Enforce Consistent `import type` Usage
Mixed usage across codebase. Add ESLint rule:
```json
{
"rules": {
"@typescript-eslint/consistent-type-imports": ["error", {
"prefer": "type-imports",
"fixStyle": "separate-type-imports"
}]
}
}
```
**Benefits**: Reduced bundle size, better tree-shaking, cleaner separation
---
### 8. Add Circular Dependency Detection
```bash
npm install -D dpdm
```
Add to `package.json`:
```json
{
"scripts": {
"check:circular": "dpdm --no-warning --no-tree src/index.ts"
}
}
```
**Potential risk areas identified**:
- `ralph-loop.ts` → `session-manager.ts` → `session.ts`
- `respawn-controller.ts` → `session.ts` → `ai-idle-checker.ts`
---
## 🟢 Lower Priority (Higher effort, situational benefit)
### 9. Apply `as const satisfies` to Default Configs
Preserves literal types while validating structure:
```typescript
// Current
export const DEFAULT_NICE_CONFIG: NiceConfig = {
enabled: false,
niceValue: 10,
};
// niceValue is type: number
// Better
export const DEFAULT_NICE_CONFIG = {
enabled: false,
niceValue: 10,
} as const satisfies NiceConfig;
// niceValue is type: 10 (literal)
```
**Files**: `types.ts`, `respawn-controller.ts` (DEFAULT_CONFIG)
---
### 10. Create Custom Error Class Hierarchy
Replace string-based errors with typed errors:
```typescript
// src/errors.ts
export class CodemanError extends Error {
constructor(
message: string,
public code: string,
public context?: Record<string, unknown>
) {
super(message);
Object.setPrototypeOf(this, CodemanError.prototype);
this.name = 'CodemanError';
}
}
export class SessionError extends CodemanError {
constructor(message: string, code: string, public sessionId: string) {
super(message, code, { sessionId });
this.name = 'SessionError';
}
}
export class ValidationError extends CodemanError {
constructor(message: string, public field: string, public value: unknown) {
super(message, 'VALIDATION_ERROR', { field, value });
this.name = 'ValidationError';
}
}
export class ScreenError extends CodemanError {
constructor(message: string, public screenName: string, public operation: string) {
super(message, 'SCREEN_ERROR', { screenName, operation });
this.name = 'ScreenError';
}
}
```
---
### 11. Split Large Files
**`types.ts` (~1500 lines)**:
```
src/types/
index.ts # Re-exports all
session.types.ts # Session-related types
task.types.ts # Task-related types
ralph.types.ts # Ralph loop types
api.types.ts # API request/response types
config.types.ts # Configuration types
factories.ts # createInitialState(), etc.
```
**`server.ts`**:
```
src/web/
server.ts # Main Fastify setup
routes/
sessions.ts # Session management routes
respawn.ts # Respawn control routes
scheduled.ts # Scheduled run routes
system.ts # System status routes
sse/
manager.ts # SSE client management
```
---
### 12. Formalize Result Pattern
Existing `validateTokenCounts` returns `{ isValid, reason }` which is essentially a Result.
**Option A: Simple Result type (no dependency)**:
```typescript
// src/utils/result.ts
export type Result<T, E = Error> =
| { success: true; data: T }
| { success: false; error: E };
export const ok = <T>(data: T): Result<T, never> => ({ success: true, data });
export const err = <E>(error: E): Result<never, E> => ({ success: false, error });
```
**Option B: Install neverthrow**:
```bash
npm install neverthrow
```
Provides chaining (`map`, `andThen`, `match`) and `ResultAsync` for async operations.
---
### 13. Template Literal Types for IDs
Enforce ID formats at compile time:
```typescript
type CycleIdFormat = `${string}:cycle-${number}`;
type ScreenSessionName = `codeman-${string}`;
interface RespawnCycleMetrics {
cycleId: CycleIdFormat; // Enforces format at compile time
}
```
---
## Summary Table
| # | Suggestion | Category | Effort | Impact |
|---|------------|----------|--------|--------|
| 1 | Use Zod for API validation | Error Handling | Low | High |
| 2 | Add `assertNever` utility | Type Safety | Low | High |
| 3 | Standardize `createErrorResponse` | Error Handling | Low | Medium |
| 4 | Discriminated union for `ApiResponse` | Type Safety | Medium | High |
| 5 | Branded types for tokens | Type Safety | Medium | Medium |
| 6 | Dependency injection for services | Architecture | Medium | High |
| 7 | Enforce `import type` | Architecture | Low | Medium |
| 8 | Circular dependency detection | Architecture | Low | Medium |
| 9 | `as const satisfies` for configs | Type Safety | Low | Low |
| 10 | Custom error classes | Error Handling | Medium | Medium |
| 11 | Split large files | Architecture | High | Medium |
| 12 | Formalize Result pattern | Error Handling | Medium | Medium |
| 13 | Template literal types for IDs | Type Safety | Low | Low |
---
## Notable Strengths to Keep
These patterns are already well-implemented and should be preserved:
- **Circuit breaker pattern** in `state-store.ts` and `ai-checker-base.ts` (excellent resilience)
- **`getErrorMessage()` utility** (solid, used in 8 files)
- **Barrel files for `utils/` and `prompts/`** (appropriate size, good organization)
- **Strict TypeScript config** (comprehensive strictness settings)
- **Well-documented configuration** in `src/config/`
- **Extensive union types** for status tracking (18+ well-defined types)
- **Type guards** like `isError()` for runtime narrowing
---
## References
### Type Safety
- [TypeScript Handbook: Narrowing](https://www.typescriptlang.org/docs/handbook/2/narrowing.html)
- [Fullstory: Discriminated Unions](https://www.fullstory.com/blog/discriminated-unions-and-exhaustiveness-checking-in-typescript/)
- [Learning TypeScript: Branded Types](https://www.learningtypescript.com/articles/branded-types)
- [Total TypeScript: satisfies Operator](https://www.totaltypescript.com/how-to-use-satisfies-operator)
### Error Handling
- [neverthrow GitHub](https://github.com/supermacro/neverthrow)
- [Zod Documentation](https://zod.dev/)
- [Custom Errors in TypeScript](https://medium.com/@Nelsonalfonso/understanding-custom-errors-in-typescript-a-complete-guide-f47a1df9354c)
### Architecture
- [Please Stop Using Barrel Files - TkDodo](https://tkdodo.eu/blog/please-stop-using-barrel-files)
- [TypeScript Dependency Injection](https://softwarepatternslexicon.com/js/typescript-and-javascript-design-patterns/dependency-injection-with-typescript/)
- [dpdm - Circular Dependency Detector](https://github.com/acrazing/dpdm)
- [Consistent Type Imports - typescript-eslint](https://typescript-eslint.io/blog/consistent-type-imports-and-exports-why-and-how/)
-155
View File
@@ -1,155 +0,0 @@
# Voice Input V2 — Implementation Plan
## Executive Summary
Fix and improve the existing VoiceInput implementation. The core class is solid but has **critical integration bugs** that prevent it from working on mobile, plus several UX improvements needed to make it feel fast and polished.
---
## Current State: What Exists
The `VoiceInput` singleton (app.js:602-830) is already committed and uses the Web Speech API with:
- Toggle mode (tap start/stop), 5s silence auto-stop
- `interimResults: true` for streaming transcription preview
- iOS Safari `isFinal` workaround (750ms stability timer)
- Desktop button in `toolbar-right`, mobile button in `KeyboardAccessoryBar`
- `voice-pulse` CSS animation, `.voice-preview` overlay
- Cleanup on SSE reconnect, haptic feedback on mobile
## Critical Bugs Found (Must Fix)
### Bug 1: Mobile button NEVER shows (CRITICAL)
`KeyboardAccessoryBar.init()` runs at line 2239, BEFORE `VoiceInput.init()` at line 2240. The accessory bar template checks `VoiceInput.supported` at render time — but `init()` hasn't run yet, so `supported` is still `false`. The inline `style="${VoiceInput.supported ? '' : 'display:none'}"` always resolves to `display:none`.
**Fix:** Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`, OR remove the inline style check and have `VoiceInput.init()` show/hide the mobile button after the fact (like it does for desktop).
### Bug 2: `_showButtons()` ignores mobile button
`_showButtons()` only targets `#voiceInputBtn` (desktop). It never removes `display:none` from the mobile `[data-action="voice"]` button.
**Fix:** Add mobile button selector to `_showButtons()`.
### Bug 3: Recognition instance leak on cleanup
`cleanup()` stops recording and removes the preview element, but doesn't null out `this.recognition`. After `cleanup()` + `init()` on SSE reconnect, the old `SpeechRecognition` instance with its handlers is orphaned.
**Fix:** Add `this.recognition = null` in `cleanup()`.
## UX Improvements (Should Fix)
### Improvement 1: Consider auto-sending after voice
Currently, voice text is inserted but the user must press Enter. This is safe but adds friction. Two options:
- **Option A (safe, current):** Insert text, user presses Enter — good for a terminal where wrong commands matter
- **Option B (fast):** Insert text + auto-send `\r` after a brief 500ms delay — feels more "voice assistant"-like
- **Recommendation:** Keep Option A as default, but add an optional setting for auto-send
### Improvement 2: Shorter silence timeout for commands
5 seconds of silence before auto-stop feels slow for short terminal commands. Consider:
- 3 seconds for auto-stop (still generous for natural pauses)
- Or make it configurable via settings
### Improvement 3: Better visual state on mobile
The blue-tinted voice button in the accessory bar is distinctive but subtle. When recording:
- The `.recording` class turns it red with pulse — good
- But the button is small among other buttons — easy to miss the state change
- Consider: also show a small red dot indicator in the header or terminal area during recording
## Architecture Decision: Keep Web Speech API
Confirmed by research: Web Speech API is the right choice.
- **Free, fast (150-300ms interim), trivial complexity**
- Chrome + Safari = ~70% of users, ~95% of Codeman's target audience (devs on Chrome)
- Works on localhost without HTTPS
- Accuracy is adequate for English command dictation
- Deepgram streaming (Phase 2 optional) only if accuracy complaints arise
- Skip Whisper batch entirely (too slow for interactive voice input)
## Implementation Plan
### Phase 1: Fix Critical Bugs (Priority)
**File: `src/web/public/app.js`**
1. **Fix init order** — Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`:
```
// Current (broken):
KeyboardAccessoryBar.init();
VoiceInput.init();
// Fixed:
VoiceInput.init();
KeyboardAccessoryBar.init();
```
2. **Fix `_showButtons()` to handle mobile** — Add mobile button selector:
```javascript
_showButtons() {
const desktopBtn = document.getElementById('voiceInputBtn');
if (desktopBtn) desktopBtn.style.display = '';
// Also show mobile button (may not exist yet if KeyboardAccessoryBar hasn't init'd)
const mobileBtn = document.querySelector('[data-action="voice"]');
if (mobileBtn) mobileBtn.style.display = '';
}
```
3. **Fix cleanup leak** — Null out recognition instance:
```javascript
cleanup() {
if (this.isRecording) this.stop();
if (this.previewEl) {
this.previewEl.remove();
this.previewEl = null;
}
this.recognition = null; // <-- add this
clearTimeout(this.silenceTimeout);
clearTimeout(this._stabilityTimer);
// ... rest
}
```
4. **Remove inline style from mobile button template** — Since `_showButtons()` will handle visibility, the template should always render the button visible and let `init()` hide it if unsupported:
```
// Current (broken):
style="${VoiceInput.supported ? '' : 'display:none'}"
// Fixed: remove the style attr entirely, let _showButtons/_hideButtons manage it
```
Actually better: **always show the button** if we init VoiceInput before KeyboardAccessoryBar. The `VoiceInput.supported` will be set correctly by then.
### Phase 2: UX Polish
5. **Reduce silence timeout** from 5s to 3s for snappier feel
6. **Add recording indicator** — When recording, add a subtle pulsing red dot to the session header or status area so the recording state is visible even if the button is off-screen
7. **Voice input setting** — Add a toggle in App Settings to enable/disable voice input (some users may not want the button). Default: enabled on supported browsers.
### Phase 3: Future Enhancements (Not in this PR)
- Language selector (currently hardcoded `en-US`)
- Auto-send option (insert text + `\r` automatically)
- Deepgram WebSocket fallback for Firefox/Edge
- Waveform visualization during recording
- Voice command recognition ("clear", "compact", "new session")
## Files to Modify
| File | Changes |
|------|---------|
| `src/web/public/app.js` | Fix init order, fix `_showButtons()`, fix `cleanup()`, remove inline style, reduce silence timeout |
| `src/web/public/mobile.css` | (optional) Adjust voice preview positioning if needed |
## Testing Plan
1. **Desktop Chrome:** Verify mic button visible in toolbar-right, click toggles recording state, interim text shows in preview, final text inserted at prompt
2. **Mobile Chrome (emulated):** Verify mic button visible in accessory bar, tap toggles recording, pulse animation plays
3. **Firefox:** Verify mic button is hidden (no SpeechRecognition support)
4. **SSE reconnect:** Verify cleanup stops recording and re-init works
5. **No active session:** Verify toast "No active session" shows when tapping mic with no session
## Risk Assessment
| Risk | Impact | Mitigation |
|------|--------|------------|
| iOS Safari isFinal bug | Medium | Already handled by 750ms stability timer |
| Chrome auto-stops after 60s | Low | Prompts are short; 3s silence timeout covers this |
| Mic permission denied | Low | Error toast with clear message |
| Init order regression | High | Integration test to verify button visibility |
-343
View File
@@ -1,343 +0,0 @@
# Browser Testing Guide for Codeman
This guide documents the browser testing infrastructure, framework comparison results, and best practices for testing the Codeman web UI.
## Quick Start
```bash
# Run standalone benchmark (recommended - avoids vitest hook issues)
npx tsx scripts/browser-comparison.mjs
# Run existing browser E2E tests
npm test -- test/browser-e2e.test.ts
```
## Framework Comparison Results
We tested three browser automation frameworks against the Codeman web UI:
| Framework | Avg Duration | Best For |
|-----------|--------------|----------|
| **Puppeteer** | 1223ms | Simple operations, Chrome-specific features |
| **Playwright** | 1373ms | Complex interactions, cross-browser, debugging |
| **Agent-Browser** | N/A (timeout) | AI agent navigation with semantic locators |
### Detailed Benchmarks
| Scenario | Playwright | Puppeteer |
|----------|------------|-----------|
| Page load | 1445ms | 433ms |
| Element selection | 442ms | 373ms |
| Modal interaction | 1487ms | 1605ms |
| Rapid operations (5 cycles) | 2119ms | 2482ms |
**Key findings:**
- Puppeteer is faster for simple page loads and element selection
- Playwright handles rapid/complex interactions better (auto-waiting)
- Agent-browser CLI has startup overhead issues in this environment
## Known Issues
### Vitest Hook Timeouts
**Problem:** Browser tests using vitest's `beforeAll`/`afterAll` hooks consistently timeout, even when the tests actually complete successfully.
**Symptoms:**
- Tests show as "skipped"
- Error: "Hook timed out in 60000ms"
- But cleanup messages appear (indicating tests ran)
**Root cause:** Unclear - possibly related to:
- vitest's module isolation with async browser launches
- Interaction between global setup.ts hooks and test-level hooks
- Multiple test file imports causing duplicate hook execution
**Workarounds:**
1. **Use standalone scripts** (recommended):
```bash
npx tsx scripts/browser-comparison.mjs
```
2. **Run browser code directly in tests** (not in hooks):
```typescript
it('should test something', async () => {
const browser = await chromium.launch();
// ... test code ...
await browser.close();
});
```
3. **Use the existing browser-e2e.test.ts pattern** which uses agent-browser CLI commands via `execSync` (avoids async hook issues)
## Test File Structure
### Port Allocation
| Port Range | Test File |
|------------|-----------|
| 3150-3153 | browser-e2e.test.ts (existing) |
| 3154 | file-link-click.test.ts |
| 3155 | browser-playwright.test.ts |
| 3156 | browser-puppeteer.test.ts |
| 3157 | browser-agent.test.ts |
| 3158-3160 | browser-comparison.test.ts |
| 3180-3182 | scripts/browser-comparison.mjs |
### File Purposes
| File | Framework | Status |
|------|-----------|--------|
| `test/browser-e2e.test.ts` | agent-browser | ✅ Working |
| `test/browser-playwright.test.ts` | Playwright | ⚠️ Vitest hook issues |
| `test/browser-puppeteer.test.ts` | Puppeteer | ⚠️ Vitest hook issues |
| `test/browser-agent.test.ts` | agent-browser | ⚠️ Vitest hook issues |
| `scripts/browser-comparison.mjs` | All three | ✅ Working (standalone) |
## Framework-Specific Patterns
### Playwright
```typescript
import { chromium } from 'playwright';
const browser = await chromium.launch({
headless: true,
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage'],
});
const page = await browser.newPage();
await page.goto('http://localhost:3000');
// Auto-waiting selectors
await page.click('.btn-claude');
await page.waitForSelector('.session-tab', { state: 'visible' });
// Assertions with expect
await expect(page.locator('.header')).toBeVisible();
await expect(page).toHaveTitle('Codeman');
await browser.close();
```
**Pros:**
- Built-in auto-waiting
- Excellent trace viewer for debugging
- Cross-browser support (Chromium, Firefox, WebKit)
- Native `expect` assertions
**Cons:**
- Slightly slower page loads
- Larger dependency
### Puppeteer
```typescript
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({
headless: true,
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage'],
});
const page = await browser.newPage();
await page.goto('http://localhost:3000');
// Manual waiting often needed
await page.click('.btn-claude');
await page.waitForSelector('.session-tab', { visible: true });
// Element queries
const title = await page.title();
const text = await page.$eval('.logo', el => el.textContent);
// CDP access for advanced features
const client = await page.target().createCDPSession();
await client.send('Performance.enable');
await browser.close();
```
**Pros:**
- Faster for simple operations
- Direct Chrome DevTools Protocol access
- Smaller dependency
- Good for Chrome-specific testing
**Cons:**
- Chrome/Chromium only
- Manual waiting required
- Less robust for complex interactions
### Agent-Browser (CLI)
```typescript
import { execSync } from 'node:child_process';
function agentBrowser(cmd: string): string {
return execSync(`npx agent-browser ${cmd}`, {
timeout: 30000,
encoding: 'utf-8',
}).trim();
}
function agentBrowserJson<T>(cmd: string): T {
const result = agentBrowser(`${cmd} --json`);
return JSON.parse(result).data;
}
// Usage
agentBrowser('open http://localhost:3000');
agentBrowser('click ".btn-claude"');
const title = agentBrowserJson<{title: string}>('get title');
// Semantic locators (AI-friendly)
agentBrowser('find role button click --name "Submit"');
agentBrowser('find text "Settings" click');
// Accessibility snapshot
const snapshot = agentBrowser('snapshot');
agentBrowser('close');
```
**Pros:**
- AI-agent friendly (semantic locators)
- Accessibility tree snapshots
- Simple CLI interface
- Reference-based selection (@e1, @e2)
**Cons:**
- CLI overhead (spawn process per command)
- Slower for rapid operations
- Less programmatic control
## Best Practices
### 1. Use Standalone Scripts for Benchmarks
Vitest has issues with browser hooks. For reliable benchmarking:
```bash
# Create a standalone .mjs script
npx tsx scripts/browser-comparison.mjs
```
### 2. Browser Launch Arguments
Always include these args for headless environments:
```typescript
{
headless: true,
args: [
'--no-sandbox', // Required for Docker/CI
'--disable-setuid-sandbox',
'--disable-dev-shm-usage', // Prevents /dev/shm issues
],
}
```
### 3. Install Playwright Browsers
```bash
npx playwright install chromium
```
### 4. Wait for Server Startup
```typescript
const server = new WebServer(PORT);
await server.start();
await new Promise(r => setTimeout(r, 1000)); // Allow server to stabilize
```
### 5. Clean Up Sessions
Track created sessions for cleanup:
```typescript
const createdSessions: string[] = [];
// In test
const response = await fetch(`${BASE_URL}/api/sessions`);
const data = await response.json();
createdSessions.push(data.sessions[0].id);
// In cleanup
for (const id of createdSessions) {
await fetch(`${BASE_URL}/api/sessions/${id}`, { method: 'DELETE' });
}
```
### 6. Handle Modal Timing
Modals have animation delays:
```typescript
// Playwright (auto-waits)
await page.click('.help-btn');
await page.waitForSelector('#helpModal', { state: 'visible' });
// Puppeteer (manual wait)
await page.click('.help-btn');
await page.waitForSelector('#helpModal', { visible: true });
// Agent-browser (explicit delay)
agentBrowser('click ".help-btn"');
await new Promise(r => setTimeout(r, 500));
```
## Key DOM Selectors
For reference when writing browser tests:
```
.btn-claude // Create Claude session button
.btn-settings // Settings button
.help-btn // Help button
.session-tab // Session tabs
.session-tab.active // Active session tab
.xterm // Terminal container
#helpModal // Help modal
#appSettingsModal // Settings modal
#sessionOptionsModal // Session Options (same set-* surface)
#createCaseModal // Add Case (same set-* surface)
.set-rail-item // Rail entry: scrolls in App Settings, switches in the other two
.set-section // A settings section (`.hidden` on the inactive ones outside App Settings)
.set-row // One setting: label + description left, control right
.modal-content // Modal content
.modal-close // Modal close button
.header-brand .logo // Logo text
#versionDisplay // Version display
#quickStartCase // Quick start dropdown
```
## Recommendations by Use Case
| Use Case | Recommended Framework |
|----------|----------------------|
| CI/CD testing | Playwright |
| Chrome-specific features | Puppeteer |
| AI agent development | Agent-Browser |
| Visual regression | Playwright |
| Performance testing | Puppeteer |
| Accessibility testing | Agent-Browser |
| Cross-browser testing | Playwright |
| Quick prototyping | Agent-Browser CLI |
## Dependencies
```json
{
"devDependencies": {
"playwright": "^1.58.0",
"puppeteer": "^24.36.0",
"agent-browser": "^0.6.0"
}
}
```
Install browsers after npm install:
```bash
npx playwright install chromium
```
-676
View File
@@ -1,676 +0,0 @@
# Claude Code Hooks Reference
> Official documentation for Claude Code hooks system, extracted from [code.claude.com](https://code.claude.com/docs/en/hooks).
**Last Updated**: 2026-07-25
**Source**: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)
> This is a maintained summary, not an exhaustive copy of the upstream reference.
> Check the source link for event-specific schemas before adding a new hook.
---
## Overview
Hooks are automated scripts that execute at specific events during your Claude Code session. They allow you to:
- Validate, modify, or block tool usage
- Add context to prompts
- Implement custom workflows
- Control agent behavior
---
## Configuration
Hooks are configured in settings files:
| File | Scope |
| ----------------------------- | -------------------------- |
| `~/.claude/settings.json` | User (global) |
| `.claude/settings.json` | Project |
| `.claude/settings.local.json` | Local project (gitignored) |
| Plugin hook files | Plugin-specific |
### Basic Structure
```json
{
"hooks": {
"EventName": [
{
"matcher": "ToolPattern",
"hooks": [
{
"type": "command",
"command": "your-command-here"
}
]
}
]
}
}
```
**Key Fields**:
- `matcher`: Pattern to match tool names (case-sensitive, supports regex like `Edit|Write` or `*` for all)
- `type`: `"command"`, `"http"`, `"mcp_tool"`, `"prompt"`, or `"agent"` where the event supports it
- `command`: Bash command to execute
- `prompt`: LLM prompt for evaluation (prompt-based hooks only)
- `timeout`: Optional timeout in seconds (default: 60)
---
## Hook Events
Claude Code's current event surface is broader than the detailed subset below. In
particular, `TeammateIdle` and `TaskCompleted` are supported lifecycle events used
by Codeman; they are not stale or plugin-defined event names.
### PreToolUse
**When**: After Claude creates tool parameters, before processing the tool call.
**Use Cases**: Approval, denial, or modification of tool calls.
**Common Matchers**:
- `Bash` - Shell commands
- `Write` - File writing
- `Edit` - File editing
- `Read` - File reading
- `Agent` - Subagent tasks
- `WebFetch`, `WebSearch` - Web operations
- `mcp__<server>__<tool>` - MCP tools
**Output Control**:
```json
{
"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "allow|deny|ask",
"permissionDecisionReason": "string",
"updatedInput": {
"field_to_modify": "new value"
},
"additionalContext": "Context for Claude"
}
}
```
### PermissionRequest
**When**: When the user is shown a permission dialog.
**Use Cases**: Auto-approve or deny permissions.
**Output Control**:
```json
{
"hookSpecificOutput": {
"hookEventName": "PermissionRequest",
"decision": {
"behavior": "allow|deny",
"updatedInput": {},
"message": "deny reason",
"interrupt": false
}
}
}
```
### PostToolUse
**When**: Immediately after a tool completes successfully.
**Use Cases**: Provide feedback, run formatters/linters, log operations.
**Output Control**:
```json
{
"decision": "block",
"reason": "Explanation",
"hookSpecificOutput": {
"hookEventName": "PostToolUse",
"additionalContext": "Additional information"
}
}
```
#### Asynchronous Rewake
Command hooks can set `"asyncRewake": true` to run asynchronously and wake an
idle Claude turn when the hook exits with code 2. The hook's stderr is delivered
to Claude as a system reminder. This implies `"async": true`; ordinary async
hooks do not wake an idle turn, and their output waits for the next interaction.
Codeman uses this on `PostToolUse(Bash)`: a self-contained Node helper extracts
the background task ID from the Bash result, watches the originating transcript
and, for subagents, the top-level parent transcript for the matching completion
notification, and exits 2. Claude records a subagent's Bash result in its
`subagents/agent-*.jsonl` file but queues completion in the lead session JSONL.
The task ID keeps each wake targeted. The helper does not send terminal input,
so it cannot submit a user's partially written prompt.
For script-dispatched Codex work, `codex-run.sh` writes the final response
between `CODEMAN_RESULT_BEGIN/END` markers in the background task output. The
rewake helper includes a maximum of 64 KiB of that report in its feedback. UI
subagent discovery and dispatcher result delivery are separate contracts.
### Notification
**When**: When Claude Code sends notifications.
**Matchers**:
- `permission_prompt`
- `idle_prompt`
- `auth_success`
- `elicitation_dialog`
- `elicitation_complete`
- `elicitation_response`
### UserPromptSubmit
**When**: When the user submits a prompt, before Claude processes it.
**Use Cases**: Add context, validate, or block prompts.
**Output Control**:
```json
{
"decision": "block",
"reason": "Explanation",
"hookSpecificOutput": {
"hookEventName": "UserPromptSubmit",
"additionalContext": "My additional context"
}
}
```
### Stop
**When**: When the main Claude Code agent finishes responding.
**Important**: Does NOT run on user interrupt.
**Use Cases**: **Ralph Wiggum loops** - block exit and refeed prompt.
**Output Control**:
```json
{
"decision": "block",
"reason": "Must provide when blocking"
}
```
Or to allow exit:
```json
{
"continue": true,
"stopReason": "optional message"
}
```
**Note**: For Stop events, `"continue": false` takes precedence over `"decision": "block"`.
### SubagentStop
**When**: When a subagent (Agent tool call) finishes responding.
**Use Cases**: Control nested loops, verify subagent output.
The hook input includes `agent_id`, `agent_transcript_path`, and
`last_assistant_message`. Like `Stop`, a command hook can return
`{"decision":"block","reason":"..."}` to keep the subagent running and feed
the reason back to it.
Codeman uses this to prevent premature reports from workers that still own live
Monitor or background-Bash processes. It derives candidate task IDs from the
subagent transcript, but requires a matching live Linux process descriptor for
`tasks/<id>.output`; historical task text by itself is not treated as active.
### TeammateIdle
**When**: When an agent-team teammate is about to go idle.
**Use Cases**: Reassign work, continue a teammate loop, or notify an orchestrator.
**Matcher Support**: None. The hook fires for every occurrence.
### TaskCompleted
**When**: When a task is about to be marked completed.
**Use Cases**: Validate completion or forward team progress to an external UI.
**Matcher Support**: None. The hook fires for every occurrence.
### PreCompact
**When**: Before a compact operation.
**Matchers**:
- `manual` - Invoked from `/compact`
- `auto` - Invoked from auto-compact
### SessionStart
**When**: When Claude Code starts or resumes a session.
**Matchers**:
- `startup` - Fresh start
- `resume` - From `--resume`, `--continue`, or `/resume`
- `clear` - From `/clear`
- `compact` - From auto or manual compact
**Use Cases**: Load development context, set environment variables.
**Persisting Environment Variables**:
```bash
#!/bin/bash
if [ -n "$CLAUDE_ENV_FILE" ]; then
echo 'export NODE_ENV=production' >> "$CLAUDE_ENV_FILE"
echo 'export API_KEY=your-api-key' >> "$CLAUDE_ENV_FILE"
fi
exit 0
```
**Output Control**:
```json
{
"hookSpecificOutput": {
"hookEventName": "SessionStart",
"additionalContext": "Context to load"
}
}
```
### SessionEnd
**When**: When a session ends.
**Reason Values**:
- `clear`
- `logout`
- `prompt_input_exit`
- `other`
**Use Cases**: Cleanup tasks, logging.
---
## Hook Input
Hooks receive JSON via stdin with common fields:
```json
{
"session_id": "abc123",
"transcript_path": "/path/to/transcript.jsonl",
"cwd": "/current/directory",
"permission_mode": "default",
"hook_event_name": "PreToolUse",
"tool_name": "Bash",
"tool_input": {},
"tool_use_id": "toolu_01ABC123..."
}
```
### Tool-Specific Input
**Bash**:
```json
{
"tool_name": "Bash",
"tool_input": {
"command": "psql -c 'SELECT * FROM users'",
"description": "Query the users table",
"timeout": 120000
}
}
```
**Write**:
```json
{
"tool_name": "Write",
"tool_input": {
"file_path": "/path/to/file.txt",
"content": "file content"
}
}
```
**Edit**:
```json
{
"tool_name": "Edit",
"tool_input": {
"file_path": "/path/to/file.txt",
"old_string": "original text",
"new_string": "replacement text"
}
}
```
---
## Hook Output
### Exit Codes
| Code | Behavior |
| ----- | --------------------------------------------------------------------- |
| 0 | Success. `stdout` processed (shown in verbose or added as context) |
| 2 | Blocking error. Only `stderr` used. Blocks tool/prompt based on event |
| Other | Non-blocking error. `stderr` shown in verbose, execution continues |
### JSON Output (Exit Code 0)
```json
{
"continue": true,
"stopReason": "optional message",
"suppressOutput": true,
"systemMessage": "optional warning"
}
```
---
## Prompt-Based Hooks
Prompt and agent handlers are supported by decision-oriented events including
`PreToolUse`, `PermissionRequest`, `PostToolUse`, `PostToolUseFailure`,
`PostToolBatch`, `UserPromptSubmit`, `Stop`, `SubagentStop`, `TaskCreated`, and
`TaskCompleted`. Check the upstream reference before choosing a handler type.
For example, a Stop event can use LLM-based evaluation:
```json
{
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "prompt",
"prompt": "Should Claude stop? Context: $ARGUMENTS\n\nCheck if all tasks are complete.",
"timeout": 30
}
]
}
]
}
}
```
**LLM Response Format**:
```json
{
"ok": true,
"reason": "Explanation when ok is false"
}
```
---
## Component-Scoped Hooks
Hooks can be defined in Skills, Agents, and Slash Commands using frontmatter:
```markdown
---
name: secure-operations
hooks:
PreToolUse:
- matcher: 'Bash'
hooks:
- type: command
command: './scripts/security-check.sh'
---
```
These hooks:
- Are scoped to the component's lifecycle
- Only run when that component is active
- Support all hook events; a subagent-scoped `Stop` is converted to `SubagentStop`
---
## MCP Tools
MCP tools follow the pattern `mcp__<server>__<tool>`:
```json
{
"hooks": {
"PreToolUse": [
{
"matcher": "mcp__memory__.*",
"hooks": [
{
"type": "command",
"command": "echo 'Memory operation' >> ~/mcp.log"
}
]
},
{
"matcher": "mcp__.*__write.*",
"hooks": [
{
"type": "command",
"command": "/home/user/scripts/validate-mcp-write.py"
}
]
}
]
}
}
```
---
## Examples
### Bash Command Validation
```python
#!/usr/bin/env python3
import json
import re
import sys
VALIDATION_RULES = [
(r"\bgrep\b(?!.*\|)", "Use 'rg' instead of 'grep'"),
(r"\bfind\s+\S+\s+-name\b", "Use 'rg --files' instead of 'find -name'"),
]
try:
input_data = json.load(sys.stdin)
except json.JSONDecodeError as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
tool_name = input_data.get("tool_name", "")
tool_input = input_data.get("tool_input", {})
command = tool_input.get("command", "")
if tool_name != "Bash" or not command:
sys.exit(1)
issues = []
for pattern, message in VALIDATION_RULES:
if re.search(pattern, command):
issues.append(message)
if issues:
for message in issues:
print(f"- {message}", file=sys.stderr)
sys.exit(2)
```
### Auto-Approve Documentation Reads
```python
#!/usr/bin/env python3
import json
import sys
try:
input_data = json.load(sys.stdin)
except json.JSONDecodeError as e:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
tool_name = input_data.get("tool_name", "")
tool_input = input_data.get("tool_input", {})
if tool_name == "Read":
file_path = tool_input.get("file_path", "")
if file_path.endswith((".md", ".mdx", ".txt", ".json")):
output = {
"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "allow",
"permissionDecisionReason": "Documentation file auto-approved"
},
"suppressOutput": True
}
print(json.dumps(output))
sys.exit(0)
sys.exit(0)
```
### Post-Write Formatter
```json
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "npx prettier --write \"$TOOL_INPUT_FILE_PATH\" 2>/dev/null || true"
}
]
}
]
}
}
```
### Ralph Wiggum Stop Hook
```bash
#!/bin/bash
# ralph-stop-hook.sh
STATE_FILE=".claude/ralph-loop.local.md"
# Check if state file exists
if [ ! -f "$STATE_FILE" ]; then
exit 0 # No active loop, allow exit
fi
# Read state from YAML frontmatter
ENABLED=$(grep -m1 "^enabled:" "$STATE_FILE" | cut -d' ' -f2)
ITERATION=$(grep -m1 "^iteration:" "$STATE_FILE" | cut -d' ' -f2)
MAX_ITER=$(grep -m1 "^max-iterations:" "$STATE_FILE" | cut -d' ' -f2)
PROMISE=$(grep -m1 "^completion-promise:" "$STATE_FILE" | cut -d' ' -f2-)
# Check if disabled
if [ "$ENABLED" = "false" ]; then
exit 0
fi
# Check max iterations
if [ -n "$MAX_ITER" ] && [ "$ITERATION" -ge "$MAX_ITER" ]; then
exit 0
fi
# Check for completion promise in output
if [ -n "$PROMISE" ]; then
if echo "$CLAUDE_OUTPUT" | grep -q "<promise>$PROMISE</promise>"; then
exit 0
fi
fi
# Block exit, increment iteration
NEW_ITER=$((ITERATION + 1))
sed -i "s/^iteration:.*/iteration: $NEW_ITER/" "$STATE_FILE"
# Output block decision
echo '{"decision": "block", "reason": "Completion promise not found. Iteration '"$NEW_ITER"'."}'
exit 0
```
---
## Environment Variables
| Variable | Description |
| -------------------- | ------------------------------------------------ |
| `CLAUDE_PROJECT_DIR` | Project root directory |
| `CLAUDE_CODE_REMOTE` | `"true"` for web, empty for CLI |
| `CLAUDE_ENV_FILE` | Path to write persistent env vars (SessionStart) |
---
## Debugging
Use `claude --debug` to see detailed hook execution:
```
[DEBUG] Executing hooks for PostToolUse:Write
[DEBUG] Found 1 hook matchers in settings
[DEBUG] Matched 1 hooks for query "Write"
[DEBUG] Executing hook command: <command> with timeout 60000ms
[DEBUG] Hook command completed with status 0: <stdout>
```
Use `/hooks` command to view registered hooks and make changes.
---
## Execution Details
- **Timeout**: 60-second default per hook, configurable
- **Parallelization**: All matching hooks run in parallel
- **Deduplication**: Identical commands deduplicated automatically
- **Matchers**: Only apply to tool-based hooks (PreToolUse, PostToolUse, PostToolUseFailure, PermissionRequest)
---
## Security Best Practices
1. **Validate and sanitize inputs** - Never trust input data blindly
2. **Always quote shell variables** - Use `"$VAR"` not `$VAR`
3. **Block path traversal** - Check for `..` in file paths
4. **Use absolute paths** - Specify full paths for scripts (use `$CLAUDE_PROJECT_DIR`)
5. **Skip sensitive files** - Avoid `.env`, `.git/`, keys, etc.
---
_Source: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)_
-121
View File
@@ -1,121 +0,0 @@
# Claude voice dictation in Codeman
Wire Codeman's existing mic button to the same speech-to-text service Claude Code's own
`/voice` mode uses, so dictation works with **no third-party API key** for anyone already
signed in to Claude Code on the server.
## Why the CLI's own voice mode cannot be reused directly
Claude Code 2.1.x ships voice input: `/voice hold|tap|off` arms it, the CLI opens the
**host's** microphone (native `audio-capture-napi`, falling back to `sox`/`arecord` on Linux
after probing `/proc/asound/cards`), streams PCM upstream and types the transcript into its
own composer.
Every part of that is on the wrong machine for Codeman. The CLI runs inside a tmux pane on
the server, which is typically headless and has no sound card at all, while the human is in
a browser on a phone somewhere else. Toggling `/voice` in the pane from Codeman would arm a
microphone nobody is sitting in front of. So Codeman keeps capturing audio in the browser,
where the user actually is, and only borrows the CLI's **transcription backend**.
## The backend, as the CLI uses it
Extracted from the 2.1.226 binary (`connectVoiceStream`):
| | |
| --- | --- |
| URL | `wss://api.anthropic.com/api/ws/speech_to_text/voice_stream` |
| Query | `encoding=linear16`, `sample_rate=16000`, `channels=1`, `endpointing_ms=300`, `utterance_end_ms=1000`, `language=<lang>`, `use_conversation_engine=true`, `stt_provider=deepgram-nova3` |
| Headers | `Authorization: Bearer <Claude Code OAuth access token>`, `User-Agent`, `x-app: cli`, `anthropic-client-platform`, optional `x-config-keyterms` |
| Audio | raw binary frames, PCM signed 16-bit little-endian, 16 kHz, mono |
| Keepalive | `{"type":"KeepAlive"}` on open, then every 8 s |
| Finalize | `{"type":"CloseStream"}`, then wait for the endpoint frame |
| Downstream | `{"type":"TranscriptText"\|"TranscriptInterim","data":"…"}` (running interim), `{"type":"TranscriptEndpoint"}` (promotes the pending interim to final), `{"type":"TranscriptError",…}`, `{"type":"error","message":…}` |
Deepgram Nova-3 runs server-side, so the Deepgram-quality result arrives without a Deepgram
account. Verified against the live endpoint before this design was written: connect, stream
PCM, receive interims and an endpoint frame.
## Architecture
The browser cannot call that endpoint itself: it would need the OAuth bearer token in page
JavaScript (and CORS would refuse anyway). So the audio goes browser → Codeman → Anthropic,
and Codeman is the only thing that ever touches the token.
```
mic → AudioWorklet (Float32 → PCM16 @16 kHz)
→ wss://<codeman>/ws/voice/stream [cookie/basic auth, Origin+Host guarded]
→ VoiceStreamRelay (reads ~/.claude/.credentials.json per connect)
→ wss://api.anthropic.com/api/ws/speech_to_text/voice_stream
← {"t":"transcript","text":…,"final":…} → existing _insertText() path
```
Nothing about the insert path changes: the transcript lands in the same preview overlay,
the same direct/compose insert modes, the same green Send button.
### Server pieces
- **`src/claude-credentials.ts`** — locate and parse the Claude Code OAuth credentials.
`parseClaudeCredentials()` is pure (JSON string + `now` → status) and unit-tested;
`readClaudeOAuthToken()` wraps it with IO: `$CLAUDE_CONFIG_DIR/.credentials.json` or
`~/.claude/.credentials.json`, and on macOS the login keychain
(`security find-generic-password -s "Claude Code-credentials"`).
**Read-only, always.** Codeman never writes credentials and never refreshes the token: a
refresh rotates the refresh token, and racing Claude Code's own refresh could sign the
user out of their CLI. An expired token surfaces as a plain "run a Claude session to
refresh" error instead.
The token is never logged, never returned by any endpoint, and never sent to the browser.
- **`src/web/voice-stream.ts`** — pure `buildVoiceStreamUrl()` / `buildVoiceStreamHeaders()` /
`sanitizeKeyterms()` (ASCII-only, deduped, 1024-char cap, mirroring the CLI), plus
`VoiceStreamRelay`, which owns one upstream socket: keepalive timer, audio passthrough,
transcript translation, finalize, and the caps below.
- **`src/web/routes/voice-routes.ts`**
- `GET /api/voice/status` → `{ available, reason, subscriptionType?, expiresAt? }`. Never
the token. `available:false` with a machine-readable `reason` (`disabled`, `no-credentials`,
`expired`) is what the settings row and the provider resolver read.
- `GET /ws/voice/stream?language=&keyterms=` → the relay. Same upgrade guard as
`/ws/sessions/:id/terminal`: allowed Host, same-site Origin, and the global auth hook has
already run on the handshake.
Caps, because an open mic is an open pipe: one stream per connection, `MAX_VOICE_STREAMS`
concurrent server-wide, a hard `MAX_STREAM_MS` per stream, and a per-frame size cap. A tab
left recording cannot bill an unbounded amount of upstream audio.
### Frontend pieces
- **`voice-pcm-worklet.js`** — an `AudioWorkletProcessor` converting Float32 blocks to PCM16
and posting ~256 ms frames back. `MediaRecorder` cannot produce raw PCM, which is why the
existing Deepgram path (container audio, auto-detected) cannot be reused as-is. Falls back
to `ScriptProcessorNode` where AudioWorklet is unavailable.
- **`ClaudeVoiceProvider`** in `voice-input.js` — mirrors `DeepgramProvider`'s shape
(`start({language, keyterms, onStream, onResult, onError, onEnd})`) so `VoiceInput` treats
the three providers uniformly.
- **Provider resolution** — new `voiceSettings.provider`: `auto` (default) | `claude` |
`deepgram` | `webspeech`. `auto` picks Claude when `/api/voice/status` reports it
available, else Deepgram when a key is set, else Web Speech. Pinning a provider always
wins, so an existing Deepgram user can keep exactly what they have.
### Settings
- `claudeVoiceEnabled` — synced, **default OFF**, gating the whole server side. Off is the
honest default: turning it on means this machine's Claude subscription starts paying for
transcription for whoever can reach the UI, and the audio goes to Anthropic rather than to
wherever it went before. One switch in Settings → Voice, and the mic works with no key.
- `voiceSettings.provider` — per the resolution table above; joins the existing synced
`voiceSettings` object.
## Things worth knowing
- **This uses an undocumented endpoint with subscription credentials.** It is the user's own
token, on the user's own machine, driving the user's own Claude Code install, but it is not
a published API and Anthropic can change or restrict it. Default-OFF is deliberate; the
Deepgram and Web Speech paths stay untouched as the supported fallbacks.
- **Multi-user mode**: every user's dictation would run on the server owner's Claude
credentials, exactly as every user's *sessions* already run on them. Consistent, but worth
stating out loud in the settings copy.
- **Token lifetime** is about 8 hours, refreshed by Claude Code itself whenever it runs. The
relay re-reads the file on every connect rather than caching, so a refresh is picked up on
the next press of the mic.
- **HTTPS or localhost**: `getUserMedia` needs a secure context. Prod is HTTPS behind
`tailscale serve`, so this is already satisfied; the existing error copy covers the rest.
-588
View File
@@ -1,588 +0,0 @@
# Claude Code Build Brief: Add Scheduling to Codeman
## 0. Purpose of This Brief
You are Claude Code working inside the Codeman repository.
Your task is to add a **small, reliable scheduling layer** to Codeman while preserving Codeman's existing architecture and session-management behavior.
This is not a greenfield rewrite. This is not a full product rebuild. This is a focused extension.
The target user wants Codeman-like tmux/web/session management, but with first-class scheduled jobs for Claude, Codex, OpenCode, Terminal, or any other configurable coding-agent harness.
---
## 1. Non-Negotiable Goal
Add scheduling to Codeman so a user can define a scheduled coding-agent job that:
1. Has a name.
2. Uses an existing Codeman-supported agent/session type where possible.
3. Has a working directory.
4. Has a prompt or prompt file.
5. Has a schedule.
6. Can be enabled or disabled.
7. Can be manually run now.
8. When due, creates a Codeman/tmux session.
9. Sends the configured prompt into that session.
10. Records last run, next run, status, and run history.
The first working version should prioritize **scheduling correctness and reuse of Codeman's existing tmux/session system** over UI polish.
---
## 2. Core Architectural Rule
Do **not** rebuild Codeman's session layer.
Reuse existing Codeman functionality for:
- Creating sessions.
- Naming sessions.
- Launching Claude/Codex/OpenCode/Terminal sessions.
- Sending input into sessions.
- Displaying sessions in the web UI.
- Killing sessions.
- Tracking session status if already supported.
If an internal API/service/function already exists, reuse it.
If no reusable function exists, create a thin wrapper around the existing implementation rather than duplicating logic.
---
## 3. Product Boundary
This build is **Codeman + Scheduler**.
It is not yet:
- A full quota engine.
- A full lock manager.
- A replacement for Codeman's terminal UI.
- A new FastAPI application.
- A multi-tenant SaaS platform.
- A complex cron-management product.
- A full agent autonomy framework.
Keep the build small and shippable.
---
## 4. Required Working Scope for v0.1
Implement the following minimum features.
### 4.1 Scheduled Jobs List
Create a UI page showing all scheduled jobs.
Each row/card should show:
- Job name.
- Agent/session type.
- Working directory.
- Schedule type.
- Enabled/disabled state.
- Last run time.
- Next run time.
- Last run status.
- Actions:
- Run Now.
- Enable/Disable.
- Edit.
- Delete.
### 4.2 Create/Edit Scheduled Job
Create a form for scheduled jobs with these fields:
- `name`
- `agent_type`
- Reuse Codeman's existing session/agent types where possible.
- Include at least Terminal/custom command if supported.
- `working_directory`
- `launch_command` if needed by Codeman's model.
- `prompt_mode`
- `inline_text`
- `prompt_file_path`
- `prompt_text`
- `prompt_file_path`
- `input_mode`
- `paste`
- `typed`
- `schedule_type`
- `once`
- `interval_minutes`
- `daily_time`
- `weekly_time`
- `run_at` for one-time jobs.
- `interval_minutes` for interval jobs.
- `daily_time` for daily jobs.
- `weekly_days` and `weekly_time` for weekly jobs.
- `enabled`
- `notes` optional.
Do not build a complex visual cron editor in v0.1.
### 4.3 Run Now
Every scheduled job must support a `Run Now` action.
Run Now should:
1. Create a new session through Codeman's existing session creation logic.
2. Send the configured prompt into the session using Codeman's existing input mechanism.
3. Create a run-history record.
4. Update last-run fields.
5. Redirect or link the user to the created Codeman session.
### 4.4 Background Scheduler Loop
Add a small background scheduler loop that runs inside the Codeman backend process.
The loop should:
1. Wake every 15-60 seconds.
2. Load enabled schedules.
3. Find schedules where `next_run_at <= now`.
4. Create a scheduled run.
5. Launch the session using existing Codeman session logic.
6. Send the prompt.
7. Record run history.
8. Compute the next run time.
9. Avoid duplicate launches if the loop overlaps or restarts.
Keep this simple and robust.
### 4.5 Run History
Every scheduled execution should create a run-history record.
Track:
- `id`
- `scheduled_job_id`
- `session_id` or Codeman session reference.
- `session_name` if applicable.
- `started_at`
- `finished_at` optional.
- `status`
- `created`
- `session_started`
- `prompt_sent`
- `failed`
- `error_message` optional.
- `trigger_type`
- `scheduled`
- `manual_run_now`
- `created_session_url` or route reference if easy.
---
## 5. Scheduling Rules
### 5.1 Once
Run at a specific date/time.
After successful launch:
- Set `enabled = false`, or mark as completed.
### 5.2 Interval
Run every N minutes.
Example:
- Every 60 minutes.
- Every 240 minutes.
After launch:
- `next_run_at = now + interval_minutes`.
### 5.3 Daily
Run every day at HH:MM.
After launch:
- Compute the next occurrence of HH:MM after now.
### 5.4 Weekly
Run on selected weekdays at HH:MM.
After launch:
- Compute the next selected weekday/time after now.
### 5.5 Timezone
Use the server's local timezone for v0.1 unless Codeman already has timezone handling.
Add a visible note in the UI:
> Times use the server's local timezone.
Do not overbuild timezone support in v0.1.
---
## 6. Data Storage Decision
First inspect Codeman's existing persistence model.
If Codeman already has a database or persistence layer:
- Reuse it.
- Add scheduled job and scheduled run models/tables/records using the existing pattern.
If Codeman uses files or JSON state:
- Use the same style for v0.1.
- Prefer simple persistence over introducing a heavy new dependency.
If there is no appropriate persistence layer:
- Add SQLite only if it fits the codebase cleanly.
- Otherwise use a JSON file store for the first version.
Do not introduce Postgres, Redis, Celery, or a separate scheduler service.
---
## 7. Concurrency and Duplicate-Run Guard
Implement a basic duplicate-run guard.
A schedule should not launch twice for the same due time.
Minimum acceptable approach:
- Before launching, create/update a run record with a `created` or `launching` state.
- Use a schedule-level `last_triggered_at` or `last_due_key` to avoid double launching.
- If launch fails, record failure clearly.
Do not build distributed locks. Codeman is expected to be local/single-instance for v0.1.
---
## 8. Multi-Session Warning
When the user clicks `Run Now`, show a warning if there are already active sessions for the same agent type.
Minimum behavior:
- If active sessions exist, show a confirmation warning.
- User can continue anyway.
For scheduled automatic runs:
- Add a setting on the scheduled job:
- `warn_only`
- `skip_if_same_agent_running`
Default:
- `warn_only` for manual runs.
- `skip_if_same_agent_running = false` for automatic runs unless easy to implement.
Do not build a complete quota engine in v0.1.
---
## 9. Prompt Sending Rules
The scheduler must support sending the configured prompt into the created session.
Prompt source:
1. Inline prompt text.
2. Prompt file path.
Input mode:
1. Paste mode.
2. Typed mode.
If only one input mode is easy with Codeman's current internals, implement that first and structure the code so the other can be added later.
Important:
- Do not send prompts to a session if session creation failed.
- Record prompt-send success/failure in run history.
- Save enough metadata to understand what prompt was used.
---
## 10. UI Bifurcation
Keep UI changes cleanly separated.
Add scheduler UI under a clear navigation item:
- `Scheduled Jobs`
Do not clutter the existing session dashboard.
The existing session dashboard may show sessions created by scheduled jobs, but the scheduling controls should live in their own section.
Recommended pages/routes:
- `/schedules`
- `/schedules/new`
- `/schedules/:id`
- `/schedules/:id/edit`
- `/schedules/:id/run-now`
- `/schedules/:id/enable`
- `/schedules/:id/disable`
- `/schedules/:id/delete`
Use Codeman's existing frontend conventions and routing style.
---
## 11. Backend Bifurcation
Keep scheduler code separate from existing session code.
Recommended logical modules, adapted to Codeman's actual structure:
- `scheduler/model` or equivalent.
- `scheduler/store` or equivalent.
- `scheduler/service` for schedule calculations and launch logic.
- `scheduler/loop` for the background due-job checker.
- `scheduler/routes` for API/UI endpoints.
- `scheduler/time` for next-run calculations.
Do not mix scheduling logic directly into terminal rendering, xterm handling, or low-level tmux code.
The scheduler service should call session services; it should not own tmux directly unless Codeman has no session abstraction.
---
## 12. Required Discovery Phase Before Coding
Before implementing, inspect the Codeman repo and produce a short architecture note in the terminal or in a file called:
`docs/cron-discovery.md`
This note must identify:
1. Where session creation happens.
2. Where agent/session types are defined.
3. Where input is sent into a session.
4. Where active sessions are listed.
5. Where session kill/delete is handled.
6. How session state is stored.
7. Whether there is existing persistence.
8. Where backend routes live.
9. Where frontend pages/components live.
10. The smallest integration points for scheduling.
Do not start coding until this discovery is complete.
---
## 13. Implementation Phases
### Phase 1: Discovery
Deliverable:
- `docs/cron-discovery.md`
Must answer the 10 discovery questions above.
### Phase 2: Data Model / Persistence
Deliverable:
- Scheduled job persistence.
- Scheduled run history persistence.
- Basic create/read/update/delete operations.
### Phase 3: Scheduler Calculation Logic
Deliverable:
- Functions to compute `next_run_at` for:
- once
- interval
- daily
- weekly
Add tests if the repo has an existing test setup.
### Phase 4: Manual Run Now
Deliverable:
- Create scheduled job.
- Click Run Now.
- Codeman session is created.
- Prompt is sent.
- Run history is recorded.
- UI links to the session.
This is the most important milestone.
### Phase 5: Background Scheduler Loop
Deliverable:
- Enabled schedules launch automatically when due.
- Run history is recorded.
- `last_run_at` and `next_run_at` update.
- Duplicate launch guard exists.
### Phase 6: UI Polish Only After Functionality
Deliverable:
- Scheduled jobs list is readable.
- Create/edit form is usable.
- Status labels are clear.
- Errors are visible.
Do not polish before Phase 4 works.
---
## 14. Acceptance Criteria
The build is acceptable when all these pass.
### Manual Run
1. Create a schedule/job with inline prompt.
2. Click Run Now.
3. A new Codeman/tmux session starts.
4. Prompt is sent into that session.
5. The created session is visible in Codeman's normal session UI.
6. Run history shows success or failure.
### One-Time Schedule
1. Create a one-time schedule 2 minutes in the future.
2. Wait for it to become due.
3. Scheduler launches a session.
4. Prompt is sent.
5. Schedule does not repeatedly launch forever.
### Interval Schedule
1. Create interval schedule every 2 minutes.
2. It launches once when due.
3. It computes the next due time.
4. It does not launch duplicates for the same due time.
### Daily Schedule
1. Create daily schedule at a time a few minutes ahead.
2. It launches when due.
3. Next run becomes tomorrow at the same time.
### Disable Schedule
1. Disable a schedule.
2. It does not launch even when due.
### Error Handling
1. Invalid working directory produces visible error.
2. Invalid prompt file produces visible error.
3. Failed session launch creates failed run-history entry.
---
## 15. Explicitly Out of Scope for v0.1
Do not implement these unless all required scope is already working:
- Full quota engine.
- Advanced lock manager.
- Post-run git inspection reports.
- Complex recurring calendar UI.
- User accounts / RBAC.
- External distributed workers.
- Redis.
- Postgres.
- Celery.
- Kubernetes.
- A separate Python service.
- Full visual cron editor.
- AI-generated follow-up prompts.
- Automatic continuation after idle.
- Any attempt to bypass agent quotas or platform limits.
---
## 16. Quality Rules
Follow these rules while coding:
1. Reuse existing Codeman services and conventions.
2. Keep scheduler code isolated.
3. Prefer boring, readable code over clever abstractions.
4. Add error messages that a human can understand.
5. Do not break existing Codeman sessions.
6. Do not rename existing core concepts unnecessarily.
7. Do not introduce large dependencies without strong reason.
8. Keep v0.1 local-first and single-instance.
9. Commit in logical chunks if git is available.
10. After coding, provide a final implementation summary.
---
## 17. Final Response Required from Claude Code
At the end, report:
1. Files changed.
2. New routes/pages added.
3. New data structures added.
4. How the scheduler loop works.
5. How to run the app.
6. How to test manual Run Now.
7. How to test scheduled execution.
8. Known limitations.
9. Suggested v0.2 improvements.
---
## 18. v0.2 Ideas, Not for Current Build
Keep these in mind but do not build unless v0.1 is complete:
- Quota-aware scheduling.
- Manual takeover locks.
- Post-idle inspection.
- Git diff reports.
- Schedule groups.
- Prompt templates.
- Agent-specific concurrency rules.
- Better timezone support.
- Audit events.
- More advanced cron expressions.
---
## 19. Final Reminder
The goal is to add **scheduling** to Codeman quickly and cleanly.
Do not drift into building a new platform.
The highest-priority path is:
1. Discover existing Codeman integration points.
2. Add scheduled job persistence.
3. Add Run Now.
4. Add background due-job loop.
5. Add minimal UI.
6. Verify that scheduled jobs create real Codeman/tmux sessions and send prompts.
-142
View File
@@ -1,142 +0,0 @@
# CRON_DISCOVERY.md
Phase 1 deliverable for the "Add Scheduling to Codeman" build brief.
This documents the existing Codeman architecture and the smallest integration
points for a cron. **No session/tmux logic will be rebuilt** —
the new code is purely a trigger + persistence + history layer on top of the
existing primitives.
Stack: `aicodeman` v1.2.1 — Fastify 5 backend, `node-pty` + tmux sessions,
vanilla-JS SPA frontend served as static assets, JSON file state store, zod
validation, ports-based dependency injection.
---
## 0. Critical finding: an existing `ScheduledRun` is NOT a cron
Codeman already has a `ScheduledRun` concept (`/api/scheduled`,
`src/web/ports/infra-port.ts:14-26`, `src/web/server.ts:1480-1605`). It is a
**run-now, duration-bounded autonomous loop**: given `{prompt, workingDir,
durationMinutes}` it immediately spawns/kills throwaway sessions in a loop until
the duration elapses. It has **no** time-based triggering, recurrence
(once/interval/daily/weekly), enable/disable, next-run calculation, run history,
or persistence across restarts.
Therefore the brief's core (the calendar/cron trigger layer) does **not** exist
and must be built. The execution primitives it sits on top of **do** exist and
will be reused. To honor brief §16 ("do not rename existing core concepts"), the
new feature is named **`CronJob`** (with **`CronJobRun`** history
records), kept distinct from the existing `ScheduledRun`.
---
## 1. Where session creation happens
- Canonical create flow: `POST /api/sessions`,
`src/web/routes/session-routes.ts:262-438`.
- `new Session({ workingDir, mode, ... })` (`src/session.ts:421-570`)
- `ctx.addSession(session)` → `ctx.setupSessionListeners(session)` →
`ctx.persistSessionState(session)` (all via `SessionPort`).
- `SessionPort` interface: `src/web/ports/session-port.ts:8-16`.
- **Integration point:** the cron service will mirror this exact sequence
(create → addSession → setupSessionListeners → start) via `SessionPort`,
not reimplement it.
## 2. Where agent/session types are defined
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity' | 'pi'`
(`src/types/session.ts:43-44`). `shell` covers the brief's "Terminal/custom".
- CLI availability resolvers in `src/utils/{claude,codex,gemini,antigravity,opencode,pi}-cli-resolver.ts`.
- **Integration point:** the job's `agentType` reuses `SessionMode` verbatim.
## 3. Where input is sent into a session
- Raw / paste: `session.write(data)` (`src/session.ts:2243-2247`) — direct PTY write.
- Typed (recommended): `session.writeViaMux(data)` (`src/session.ts:2301-2311`)
— tmux `send-keys`, falls back to PTY. Submit requires trailing `\r`.
- **Integration point:** prompt delivery uses `writeViaMux` (typed) by default,
`write` (paste) as the alternate `input_mode`.
## 4. Where active sessions are listed
- `ctx.sessions: ReadonlyMap<string, Session>` (`SessionPort`).
- Filters: `Array.from(ctx.sessions.values()).filter(s => s.mode === X)` and
`.isBusy()` / `.isIdle()` (`src/session-manager.ts:220-247`).
- **Integration point:** the §8 multi-session warning queries this map.
## 5. Where session kill/delete is handled
- `ctx.cleanupSession(sessionId, killMux?, reason?)`
(`SessionPort`; impl `src/web/server.ts:997-1152`). Underlying
`session.stop(killMux)` at `src/session.ts:2498-2585`.
- The cron does **not** kill sessions it launches (the brief wants them
visible in the normal session UI); cleanup stays user-driven.
_Superseded post-review:_ recurring jobs now default to
`autoClosePreviousSession: true` — the previous run's still-open session is
closed via `cleanupSession` when the next run fires (see
`docs/cron-guide.md` §8); opt out per job for fully user-driven cleanup.
## 6. How session state is stored / 7. Existing persistence
- JSON file store: `~/.codeman/state.json` (+ `state-inner.json` for Ralph).
`StateStore` class `src/state-store.ts:71`; `AppState` interface
`src/types/app-state.ts:99-114`.
- Pattern: declare a field on `AppState`, add typed get/set methods on
`StateStore` that mutate in-memory state and call the debounced `save()`
(500ms debounce, atomic temp-file+rename, `.bak` backup, circuit breaker).
- **Integration point:** add `cronJobs?: Record<string, CronJob>` and
`cronJobRuns?: Record<string, CronJobRun>` to `AppState`, with
matching `StateStore` accessors. No new DB (brief §6 forbids Postgres/Redis).
## 8. Where backend routes live
- Route modules: `src/web/routes/*.ts`; barrel `src/web/routes/index.ts`;
registered in `WebServer.setupRoutes()` `src/web/server.ts:858-876` with a
single `ctx` object from `createRouteContext()` (`src/web/server.ts:553-613`)
that satisfies all port interfaces.
- Validation: zod schemas in `src/web/schemas.ts`, applied via
`parseBody(Schema, req.body)` (`src/web/route-helpers.ts:101-111`).
- Errors: `createErrorResponse(ApiErrorCode.X, msg)` / `ApiResponse`
(`src/types/api.ts`), auto-mapped to HTTP status by a `preSerialization` hook
(`src/web/server.ts:644-659`).
- SSE: `ctx.broadcast(SseEvent.X, data)` (`EventPort`,
`src/web/sse-events.ts`); frontend mirror in `src/web/public/constants.js`.
- **Integration point:** new `cron-routes.ts` registered alongside the
others; new zod schema; new `SseEvent` constants for job list/run changes.
## 9. Where frontend pages/components live
- Vanilla-JS SPA: single `src/web/public/index.html` + feature mixin files
(`Object.assign(CodemanApp.prototype, {...})`). API via `api-client.js`
(`_apiJson/_apiPost/_apiDelete`). Build = esbuild minify + content-hash, no
bundler (`scripts/build.mjs`).
- UI is panels/modals toggled by JS classes; forms use `.form-row` / `.modal`
conventions (`styles.css`). SSE handler map in `app.js`.
- **Integration point:** add a new `cron-ui.js` mixin + a panel/modal in
`index.html` + nav entry, following the orchestrator/respawn panel pattern.
## 10. Background-loop pattern (for the due-checker)
- Established pattern: `this.cleanup.setInterval(fn, intervalMs, {description})`
in `WebServer.start()` (`src/web/server.ts:~1942-1966`), auto-disposed in
`WebServer.stop()` via `this.cleanup.dispose()` (`src/web/server.ts:2336`).
RalphLoop (`src/ralph-loop.ts:268-286`) shows the self-rescheduling guard idiom.
- **Integration point:** register a 30s cron tick via `cleanup.setInterval`;
no manual shutdown wiring needed.
---
## Smallest integration points (summary)
| New piece | Reuses | Location |
| --- | --- | --- |
| `CronJob` / `CronJobRun` types | — (new) | `src/types/cron.ts` |
| Persistence | `StateStore` / `AppState` | `src/types/app-state.ts`, `src/state-store.ts` |
| Next-run time math | — (new, pure, unit-tested) | `src/cron/cron-time.ts` |
| Launch + send prompt | `SessionPort` (`addSession`/listeners/`writeViaMux`) | `src/cron/cron-service.ts` |
| Background due loop | `cleanup.setInterval` pattern | `src/cron/cron-loop.ts` |
| Routes + schema | route/ports/zod/SSE patterns | `src/web/routes/cron-routes.ts`, `src/web/schemas.ts`, `src/web/sse-events.ts` |
| UI | panel/modal/mixin conventions | `src/web/public/cron-ui.js`, `index.html` |
Nothing in the session, tmux, persistence, routing, or SSE subsystems is
rewritten — the cron is additive and calls existing services.
-426
View File
@@ -1,426 +0,0 @@
# Cron Jobs — User & Operator Guide
Codeman's **Cron** feature lets you save named, recurring jobs that automatically
spin up a Claude (or shell / OpenCode / Codex / Antigravity / Gemini / Pi) session on a schedule and
feed it a prompt. Think "cron for agent sessions": _"every weekday at 3am, open a
Claude session in `~/proj` and tell it to update dependencies and open a PR."_
- **UI**: the **⏰ Cron** button in the header → the Cron Jobs modal (`#cronModal`).
- **API**: `/api/cron/jobs*` and `/api/cron/runs`.
- **Code**: `src/cron/cron-service.ts`, `src/cron/cron-time.ts`, `src/cron/cron-input.ts`,
types in `src/types/cron.ts`, routes in `src/web/routes/cron-routes.ts`,
frontend in `src/web/public/cron-ui.js`.
> **Not to be confused with `ScheduledRun` (`/api/scheduled`).** That older,
> deliberately-separate concept is a _run-now, duration-bounded autonomous loop_
> (`{prompt, workingDir, durationMinutes}` → spawn/kill throwaway sessions until
> the duration elapses). It has no recurrence, no saved jobs, and no next-run
> calculation. The two systems never interact. This guide is only about **Cron
> jobs** (`Cron*`). See `docs/cron-discovery.md` §0.
---
## 1. Quick start
### In the browser
1. Click **⏰ Cron** in the header.
2. Click **+ New Job**.
3. Fill in a **name**, pick an **agent type** and **working directory**, choose a
**prompt** (inline text or a file path), pick a **schedule**, and leave
**Enabled** on.
4. **Save**. The job appears in the list with its computed **next run**.
5. Use **Run Now** to fire it immediately without waiting for the schedule.
### With curl
```bash
API=http://localhost:3000
# Create a daily job (03:00 server-local time)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{
"name": "nightly-deps",
"agentType": "claude",
"workingDir": "/home/me/proj",
"promptMode": "inline_text",
"promptText": "Update dependencies and open a PR",
"inputMode": "typed",
"scheduleType": "daily",
"dailyTime": "03:00",
"enabled": true,
"concurrencyPolicy": "warn_only"
}' | jq
# List jobs
curl -s "$API/api/cron/jobs" | jq
# Run one immediately
curl -s -X POST "$API/api/cron/jobs/<jobId>/run" | jq
# See a job's run history
curl -s "$API/api/cron/jobs/<jobId>/runs" | jq
```
---
## 2. Concepts
| Term | Meaning |
| -------------------------- | ------------------------------------------------------------------------------------------------ |
| **Cron job** (`CronJob`) | A saved, named definition: what agent to launch, where, with what prompt, on what schedule. |
| **Run** (`CronJobRun`) | One execution of a job — a history record with a status and a link to the session it created. |
| **Schedule type** | How fire times are computed: `once`, `interval`, `daily`, or `weekly`. |
| **Next run** (`nextRunAt`) | Server-computed epoch-ms of the next fire. `null` when the job is disabled or has no future run. |
| **Due tick** | A background loop (every 30s) that launches any enabled job whose `nextRunAt` has passed. |
A job is essentially a **trigger + persistence + history layer on top of the
existing session primitives**. When a job fires, the cron service does exactly
what the "quick start" route does — `new Session(...)` → `addSession` →
`setupSessionListeners` → `startInteractive()`/`startShell()` → deliver the
prompt. It does **not** reimplement any tmux/PTY logic.
---
## 3. The job form — every field
These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
(`src/types/cron.ts`).
| Field | Required | Values / limits | Notes |
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` \| `pi` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. ⚠️ A `pi` job's readiness poll looks for `❯`/a token count, neither of which pi prints, so it burns the poll budget and then sends the prompt anyway (slower start, still works). |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
| `promptText` | conditional | ≤ 100000 chars, **single line** | Required when `promptMode = inline_text`. Newlines are rejected (see §6). |
| `promptFilePath` | conditional | valid path | Required when `promptMode = prompt_file_path`. Confined to `workingDir` (see §5). |
| `inputMode` | ✅ | `paste` \| `typed` | How the prompt is delivered. See §6. |
| `scheduleType` | ✅ | `once` \| `interval` \| `daily` \| `weekly` | See §4. |
| `runAt` | conditional | epoch-ms (positive int) | Required for `once`. |
| `intervalMinutes` | conditional | 1–525600 (≤ 1 year) | Required for `interval`. |
| `dailyTime` | conditional | `HH:MM` (24h) | Required for `daily`. Server-local time. |
| `weeklyDays` | conditional | array of 1–7 ints, each 0–6 (0 = Sunday) | Required for `weekly`. |
| `weeklyTime` | conditional | `HH:MM` (24h) | Required for `weekly`. Server-local time. |
| `enabled` | ✅ | boolean | Disabled jobs never auto-fire (but **Run Now** still works). |
| `notes` | — | ≤ 2000 chars | Free-form. |
| `concurrencyPolicy` | ✅ | `warn_only` \| `skip_if_same_agent_running` | Applies to **automatic** runs only. See §7. |
| `autoClosePreviousSession` | — | boolean (default **true**) | Recurring schedules only (ignored for `once`): when the next run fires, the still-open session created by this job's **previous** run is closed first via the normal cleanup path. See §8. |
**Cross-field validation** (`refineCronJob` in `schemas.ts`): the conditional
fields above are enforced by a Zod `superRefine` on create. A missing dependent
field (e.g. `scheduleType: "once"` with no `runAt`) is rejected with
`INVALID_INPUT` and a field-specific message.
> ⚠️ **Update caveat.** `PUT /api/cron/jobs/:id` uses a `.partial()` schema that
> does **not** re-run the cross-field `superRefine`. To keep partial edits safe,
> `updateJob()` re-validates the **merged** job against the full `CronJobSchema`
> and throws `400` if the result is inconsistent (e.g. switching to `once`
> without a `runAt`). So the store is never left with a half-valid job.
---
## 4. Schedule types
Next-run math lives in `src/cron/cron-time.ts` (pure, unit-tested in
`test/cron-time.test.ts`). **All wall-clock times use the server's local
timezone** (v0.1 decision).
### `once`
- Fires a single time at the absolute `runAt` epoch-ms.
- A **missed** one-time job (server was down at `runAt`) **still fires once** on
the next tick — `computeNextRunAt` returns `runAt` even if it's in the past,
until the job has fired.
- After firing, the job **self-disables**: `completedOnce = true`, `enabled =
false`, `nextRunAt = null`.
### `interval`
- Fires every `intervalMinutes`, computed as `fireTime + intervalMinutes`.
- ⚠️ **Drift**: the next run re-anchors to the actual fire time, not to an ideal
cadence — a slow tick or restart shifts subsequent runs slightly later. This is
an accepted limitation.
### `daily`
- Fires at `dailyTime` (`HH:MM`) every day, server-local.
- If today's time has already passed, the next run is tomorrow at that time.
### `weekly`
- Fires at `weeklyTime` on each weekday in `weeklyDays` (0 = Sunday … 6 =
Saturday), server-local.
- The next run is the soonest upcoming matching weekday/time within the next 7
days.
---
## 5. Prompt source (`promptMode`)
### `inline_text`
The prompt is the literal `promptText`. Simplest option.
### `prompt_file_path`
The prompt is read from a file at fire time. **This path is security-hardened**
because a job config is attacker-controllable and the file's contents are
injected into an agent session (an exfiltration sink over SSE/terminal).
`resolveSafePromptPath()` enforces, in order:
1. **`realpath` resolution** — symlinks are resolved to their true target, for
the prompt file **and for `workingDir` itself**.
2. **`workingDir` is not a trust boundary** — because it is user-supplied, the
resolved `workingDir` is itself rejected if it is `/` or resolves into a
blocked tree (`/etc`, `/root`, operator extras) or a pseudo-filesystem
(`/proc`, `/sys`, `/dev`). This closes the `workingDir: '/proc'` +
`promptFilePath: '/proc/self/environ'` env-exfil trick. The same rule is
enforced earlier, at job create/update.
3. **Blocklist** (defense-in-depth) — sensitive trees (`/etc`, `/root`,
`/proc`, `/sys`, `/dev`, known secret locations) are rejected for the
resolved prompt file.
4. **Allowlist (primary gate)** — the resolved path **must live inside the job's
(resolved) `workingDir`** (`validateSessionFilePath`). A symlink escaping the
workspace fails here.
5. **Regular-file check** — directories, FIFOs, and `/dev/*` character devices
are rejected (they would hang or OOM an unbounded read).
6. **Size cap** — files larger than **1 MiB** (`MAX_PROMPT_FILE_BYTES`) are
rejected.
7. **Single-line check** — after trailing newlines are stripped, the file
content must be a single line (see §6).
If any check fails, the run is recorded as **`failed`** with the reason; no
session is created.
---
## 6. Prompt delivery (`inputMode`)
Once the CLI is ready (see §8), the prompt is written to the session with a
trailing carriage return:
| Mode | Mechanism | Use when |
| ------- | --------------------------------------------------------------- | ------------------------------------------------ |
| `typed` | `session.writeViaMux()` — tmux `send-keys -l` (literal) + Enter | Default; behaves like a human typing the prompt. |
| `paste` | `session.write()` — writes directly to the PTY/mux | Bulk paste-style delivery. |
> ⚠️ **Single-line only — enforced.** Like all programmatic input in Codeman,
> multi-line delivery would be silently corrupted (Ink-based TUIs treat a
> newline as submit; typed mode fuses lines). So newlines are **rejected**: the
> schema and the form refuse a multi-line `promptText`, and at fire time a
> prompt file whose content is multi-line (after stripping trailing newlines)
> fails the run with a clear `errorMessage`. Put multi-line instructions in a
> file the agent is told to read itself (e.g. "read TASKS.md and do it").
---
## 7. Concurrency policy (automatic runs)
`concurrencyPolicy` governs what happens when a **scheduled** run is due and
sessions of the same `agentType` already exist:
| Policy | Behavior |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `warn_only` | Always launch. (The count is surfaced but not blocking.) |
| `skip_if_same_agent_running` | If ≥ 1 **other, live** session of that mode is active, **skip** this fire — record a `skipped` run and (for recurring schedules) advance the schedule without launching. |
Notes on `skip_if_same_agent_running`:
- Only **live** sessions block: a tab whose CLI already exited (status
`stopped`/`error`) does not count.
- Sessions created by **this job's own previous runs never block it** —
otherwise a recurring job would deadlock on the session it created last time
and fire exactly once.
- A skipped **`once`** job is **not consumed**: it stays armed and retries on
the next tick until the blocking session goes away, then fires its single run.
- A skip is **not** a run: it sets `lastStatus = 'skipped'` but does **not**
advance `lastRunAt`.
- Consecutive skips are **coalesced** — a perpetually-skipped interval job writes
**one** skip record per streak, not one every tick, so it can't bloat
`state.json`.
**Run Now ignores this policy on the server.** The browser shows a `confirm()`
warning if same-type sessions are active, but if you proceed (or call the API
directly), the job launches unconditionally.
---
## 8. What happens when a job fires
Sequence in `CronService.launch()`:
1. A `CronJobRun` is created with status **`created`** and broadcast
(`cron:runCreated`).
2. The prompt is resolved (inline or file, single-line enforced). Failure →
**`failed`**.
3. `workingDir` is checked (`statSync().isDirectory()`). Missing/not-a-dir →
**`failed`**.
4. **Auto-close previous session** (recurring schedules, unless
`autoClosePreviousSession: false`): any still-open session created by this
job's previous runs is closed via the normal session-cleanup path.
5. The global session cap is checked (`MAX_CONCURRENT_SESSIONS = 50`). At cap →
**`failed`**.
6. A `Session` is created **with `useMux: true`** (so it runs inside tmux),
registered, listeners attached, and started via `startInteractive()`
(`startShell()` for `shell` mode). Model/claudeMode come from global config.
Run status → **`session_started`**.
7. **Readiness wait** (async, non-blocking): for non-shell agents the service
polls the terminal buffer up to **60 × 500ms** for a `❯` prompt or the string
`tokens`, then settles **2000ms** (`CRON_READY_SETTLE_MS`). Shell mode waits
1000ms, then sends the optional `launchCommand` as the first input line
(+1000ms settle).
8. The prompt is delivered (`typed`/`paste`, trailing `\r`). Run status →
**`prompt_sent`**; `finishedAt` stamped. Delivery failure (e.g. the mux
session is gone) → **`failed`**.
The created session is a **normal, persistent interactive session** — it appears
as its own tab and keeps running after the prompt is sent. The run's
`createdSessionUrl` is a deep link (`/?session=<id>`); the UI focuses it
automatically after **Run Now**.
> ⚠️ **Session-cap math if you disable auto-close.** With
> `autoClosePreviousSession: false`, nothing ever closes the sessions a
> recurring job creates — an interval job every 30 min creates 48 tabs/day and
> hits the global 50-session cap in ~25 hours (sooner with existing tabs), after
> which **every** fire of **every** job fails with "Maximum concurrent sessions
> reached" until you delete tabs by hand. Leave auto-close on for unattended
> recurring jobs, or clean up sessions yourself.
### The background tick
`tickDueJobs()` runs every **30s** (`CRON_TICK_INTERVAL`, registered in
`server.ts`). For each enabled job whose `nextRunAt ≤ now`:
- **Duplicate-launch guard**: `lastDueKey = jobId:fireTime`. If this due time was
already consumed (overlap/restart), the job is just advanced, not relaunched.
- The schedule is **advanced _before_ launching** so a slow launch can't be
re-triggered by the next tick.
- On boot, `init()` recomputes `nextRunAt` for loaded jobs (dead `once` jobs stay
dead).
---
## 9. Run history & statuses
Each job keeps a history of `CronJobRun` records. Statuses (`CronJobRunStatus`):
| Status | Meaning |
| ----------------- | ------------------------------------------------------------- |
| `created` | Run record created; prompt/session not yet started. |
| `session_started` | Session launched successfully. |
| `prompt_sent` | Prompt delivered — the happy-path terminal state. |
| `failed` | Something went wrong (see `errorMessage`). |
| `skipped` | A scheduled fire was skipped by `skip_if_same_agent_running`. |
Each run also records `triggerType` (`scheduled` or `manual_run_now`),
`sessionId`/`sessionName`, timestamps, and `createdSessionUrl`.
**History is capped globally** at **500 records** (`MAX_CRON_RUN_HISTORY`); the
oldest are pruned first. Deleting a job also deletes its run records.
---
## 10. API reference
All responses use the standard `ApiResponse<T>` envelope (`{success, data}` /
`{success, error, errorCode}`). `/api/v1/*` is a stable alias.
| Method | Endpoint | Body | Returns |
| -------- | ---------------------------- | ---------------------- | --------------------------------- |
| `GET` | `/api/cron/jobs` | — | `CronJob[]` |
| `POST` | `/api/cron/jobs` | `CronJobSchema` | `{ job }` |
| `GET` | `/api/cron/jobs/:id` | — | `CronJob` (404 if missing) |
| `PUT` | `/api/cron/jobs/:id` | partial `CronJob` | `{ job }` (400 if merge invalid) |
| `DELETE` | `/api/cron/jobs/:id` | — | `{}` |
| `PUT` | `/api/cron/jobs/:id/enabled` | `{ enabled: boolean }` | `{ job }` |
| `POST` | `/api/cron/jobs/:id/run` | — | `{ run, activeAgents }` |
| `GET` | `/api/cron/jobs/:id/runs` | — | `CronJobRun[]` (newest first) |
| `GET` | `/api/cron/runs` | — | all `CronJobRun[]` (newest first) |
---
## 11. SSE events
Emitted on `/api/events`, mirrored in `SSE_EVENTS` (`constants.js`):
| Event | Payload | When |
| ------------------ | ------------ | -------------------------------------------------------------------- |
| `cron:jobsChanged` | `{ jobs }` | Any job created / updated / enabled / status change. |
| `cron:jobDeleted` | `{ id }` | A job was deleted. |
| `cron:runCreated` | `CronJobRun` | A run (incl. skips) started. |
| `cron:runUpdated` | `CronJobRun` | A run advanced state (`session_started` / `prompt_sent` / `failed`). |
---
## 12. State & persistence
Persisted in `~/.codeman/state.json` via `StateStore`:
- `AppState.cronJobs` — map of `id → CronJob`.
- `AppState.cronJobRuns` — map of `id → CronJobRun`.
Jobs and their schedules survive restarts; `init()` recomputes `nextRunAt` on
boot. Sessions the jobs create persist through the normal session-recovery path.
---
## 13. Limits & constants
| Constant | Value | Source |
| ------------------------ | --------------------- | ------------------------------------------------ |
| Due-tick interval | 30s | `CRON_TICK_INTERVAL` (`config/server-timing.ts`) |
| Readiness poll | 60 × 500ms | `CRON_READY_MAX_ATTEMPTS` |
| Readiness settle | 2000ms | `CRON_READY_SETTLE_MS` |
| Run-history cap (global) | 500 | `MAX_CRON_RUN_HISTORY` (`config/map-limits.ts`) |
| Saved-jobs cap | 100 | `MAX_CRON_JOBS` (`config/map-limits.ts`) |
| Concurrent-session cap | 50 | `MAX_CONCURRENT_SESSIONS` |
| Prompt-file size cap | 1 MiB | `MAX_PROMPT_FILE_BYTES` (`cron-service.ts`) |
| `name` length | 1–200 | `CronJobSchema` |
| `promptText` length | ≤ 100000 | `CronJobSchema` |
| `intervalMinutes` | 1–525600 | `CronJobSchema` |
| `weeklyDays` | 1–7 entries, each 0–6 | `CronJobSchema` |
---
## 14. Known limitations
- **Server-local timezone only** — `daily`/`weekly` times are interpreted in the
host's local time; there is no per-job timezone.
- **Interval drift** — `interval` re-anchors to the actual fire time; long-running
intervals slowly shift.
- **Single-line prompts** — multi-line prompts are rejected (schema, form, and
at fire time for prompt files); tell the agent to read a file itself for
multi-line instructions.
- **`runNow` / tick race** — a manual Run Now firing at the same instant as a
scheduled tick is theoretically possible; benign (you may get two sessions).
- **`{enabled:true}` on a dead `once` job** — re-enabling a fired one-time job
without changing its schedule leaves it enabled-but-dead (won't fire); change
the schedule to re-arm.
---
## 15. Troubleshooting
| Symptom | Likely cause | Fix |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| Job never fires | Disabled, or `nextRunAt: null` | Check **Enabled**; verify the schedule fields are complete. |
| Run shows `failed` immediately | Bad `workingDir`, prompt-file rejected, or session cap hit | Read `errorMessage` on the run; confirm the dir exists and the prompt file is inside it and < 1 MiB. |
| Run shows `skipped` | `skip_if_same_agent_running` + another live same-type session (this job's own sessions and dead tabs don't count) | Switch to `warn_only`, or wait for the other session to end. |
| Run fails with "single line" | Multi-line prompt text / prompt file | Keep the prompt to one line; point the agent at a file to read for long instructions. |
| Sessions pile up between runs | `autoClosePreviousSession: false` | Re-enable auto-close, or delete old tabs before the 50-session cap bites (see §8). |
| Wrong fire time | Timezone assumption | Times are **server-local** — check the host clock/TZ. |
| One-time job won't re-fire | `completedOnce` set | Edit the schedule (any real schedule change re-arms it). |
---
## 16. Related docs
- `docs/cron-discovery.md` — architecture / integration-point analysis (why the
feature reuses the session layer and stays distinct from `ScheduledRun`).
- `docs/cron-build-brief.md` — the original build brief / requirements.
- `CLAUDE.md` → **Key Patterns → Cron** — the one-paragraph engineering summary.
- Tests: `test/cron-time.test.ts` (schedule math), `test/cron-service.test.ts`
(CRUD, tick, concurrency, security).
-433
View File
@@ -1,433 +0,0 @@
<!-- Design doc generated via ultracode multi-agent workflow (wf_e3a7498b-26f): 3 architecture proposals -> judge panel -> synthesis -> completeness critic. -->
# Docker Session Mode, Implementation Plan
## Decisions (locked 2026-07-19, by repo owner)
1. **Isolation posture**: CONVENIENT default (bind-mount host `~/.claude` etc. read-write so the existing login just works; network on; still hardened non-root + cap-drop + resource caps). SEALED profile (`mountCredentials:false` + `network:none`) is a per-case opt-in.
2. **Export**: offer BOTH full-image (`commit`+`save`+workspace tar) AND workspace-only, side by side, no default (ask each time).
3. **Base image**: BUILD LOCALLY on first use via `scripts/build-agent-image.mjs` from a repo `docker/agent.Dockerfile`. No registry required. (GHCR pull can be added later.)
4. **Hooks**: WIRE HOOKS NOW. Codeman scaffolds `.claude/settings.local.json` + CLAUDE.md into the linked host workspace dir (same as local cases), enabling in-container permission prompts, hook-idle detection, and the Claude Model picker.
Adopted defaults for the remaining open items (Section 10): resume-on-restart ON; container is per-CASE and shared by multiple sessions (killing one session only kills its in-container tmux session, never `docker stop` while siblings remain; stop/remove only on explicit teardown or case-delete); rootless caps = ship-with-warning (`capsEnforced` surfaced); remote docker daemon = local-first; podman = docker-first best-effort.
## Implementation status (branch `feat/docker-session-mode`)
DONE and END-TO-END VERIFIED against a real docker daemon (create host, link case, quick-start shell in a real container, workspace bind-mount round-trip, hook scaffolding, session-delete keeps the shared container up, case-delete `docker rm`s it):
- Phase 0-1: types (`DockerHost`/`DockerCase`/`SessionDocker`), `src/docker-hosts.ts` (storage, pure `buildDockerBaseArgs`/`buildDockerCreateArgs`, `containerApiUrl`, `hostGatewayAlias`, config-hash, credential-mount resolution, daemon probes), `DockerHostSchema`/`DockerCaseLinkSchema`. 26 unit tests.
- Phase 2: `tmux-manager` `buildDockerLaunchCommand` (image-check -> ensure -> start -> exec, resume-aware), `buildDockerKillCommand` (in-container tmux only, multi-session safe), stop/remove; wired into `createSession`/`respawnPane`/`killSession`. 14 unit tests.
- Phase 3: `Session` threading (`_docker`, toState, option builders, in-container cliVersion probe, `resolveMuxAttachCwd`), `server.ts` recovery round-trip.
- Phase 4: `case-routes` `/api/docker-hosts` CRUD + `/api/cases/docker-link` + listing + docker-unlink; `session-routes` `/api/quick-start` docker branch (rejects per-session config, probes availability + tmux, scaffolds hooks, seeds resume id).
- Phase 5 (partial): `docker/agent.Dockerfile` + `scripts/build-agent-image.mjs` (built + verified: node 22, tmux, claude/codex/gemini/opencode, arbitrary-uid HOME). Host-guard allowlists `host.docker.internal`/`host.containers.internal` for in-container hooks.
- Full CI green (3445 tests).
REMAINING:
- Phase 6: export / import (`docker commit` + `save | gzip` + workspace tar + manifest; `load` + quarantined re-tag), GC / boot reaper, disk-safety prechecks, drift-recreate route, SSE `docker:*` events. THE "move to a new machine" feature.
- Phase 7: frontend Create Case "Docker" tab + `linkDockerCase` + run wiring + case-picker labels + export/import UI.
- Phase 8: CLAUDE.md "Docker cases" Key Pattern + `docs/docker-cases.md` + COM.
- Deferred refinements: in-container model-picker via `settings.local.json`; live mid-run resume-id capture into `DockerCase.lastClaudeSessionId`; rootless/Desktop uid probe (currently a platform heuristic).
## 1. Goal & user stories
Add "Docker cases" to Codeman: a case can point at a container instead of a local or remote-SSH path, and any of the five CLI backends (`claude` / `shell` / `opencode` / `codex` / `gemini`) runs inside that container. It is modeled as a LOCATION OVERLAY on cases, exactly like the remote-SSH feature (COD-94/#145), never as a sixth `SessionMode`.
User stories:
- As the repo owner, I link a case to a per-project container so an autonomous Claude/Ralph run executes in a hardened sandbox (cap-drop, non-root, resource caps) instead of directly on my host, while keeping my existing OAuth login and transcript history working with zero extra setup.
- I set default, per-case-changeable container settings (image, network mode, memory/cpu/pids caps) at link time and edit them later, and edits actually take effect through a recreate-on-drift path (see Section 4).
- I reconnect after a Codeman restart and land back in the SAME running agent with the conversation intact. When the CONTAINER itself was stopped/rebooted/OOM-killed (which destroys the in-container tmux), the next launch RESUMES the last conversation from the bind-mounted transcript rather than starting fresh (durability model in Section 2, Key decision 1).
- I export a finished run's whole environment (toolchain plus workspace) to a portable, secret-free `.tar.gz`, move it to another machine, and import it back into a fresh case in one click.
- The container never accumulates: killing the session stops it, deleting the case removes it, and an instance-scoped boot reaper reaps containers whose case is gone.
Non-goals for the MVP: multi-tenant untrusted-code isolation guarantees (Codeman is loopback-default and single-operator, and the agent already runs `--dangerously-skip-permissions` on the host today), Kubernetes/compose orchestration, and per-command ephemeral containers.
## 2. Chosen architecture and why
The design grafts the strongest idea from each of the three proposals:
- Overlay-not-a-mode + faithful remote-SSH mirror (from "Docker Cases as a Location Overlay"): lowest churn, rides the existing quick-start / mux-sessions / state / recovery plumbing.
- Convenient-but-hardened default with an opt-in sealed profile, plus exec-time name-only secret env (from "Sealed Sandbox"): a strict security improvement over today's on-host execution without the UX tax of forcing an in-container re-login.
- One-artifact export + in-app import route (from "Container-as-Cargo"): the genuinely new, high-value capability Codeman lacks.
### Key decision 1: persistent per-CASE container, durable in-container tmux, AND resume-on-restart (the two-layer durability model)
Exactly one long-lived container per Docker case, named as a pure slug function `codeman-case-<slug>` (Docker charset `^[a-zA-Z0-9][a-zA-Z0-9_.-]+$`; Codeman already slugs case names for tmux), so create-if-missing and boot recovery are idempotent. PID1 is `sleep infinity` under `--init` (tini reaps zombies and forwards `docker stop`'s SIGTERM); the CLI is NOT the container command. The CLI runs inside a DURABLE in-container tmux on a dedicated socket `-L codeman-docker`, session `codeman-dkr-<id8>`, the direct analog of remote's `-L codeman-remote` / `codeman-ssh-<id8>`.
Two DIFFERENT failure surfaces need two DIFFERENT recovery layers, and conflating them is the central flaw the critic caught:
1. Codeman-PROCESS restart while the container stays up: the in-container tmux is still alive, so `tmux new-session -A` (attach-or-create) reattaches the SAME live agent and the paneCommand is ignored. This is the remote-SSH durability idiom and it works unchanged.
2. CONTAINER stop / daemon restart / host reboot / OOM-kill: the in-container tmux is GONE (fresh PID1). `new-session -A` will now CREATE a fresh session and run the paneCommand, which would start a brand-new conversation. This is the case the raw plan silently lost. Because the transcript directory is bind-mounted from the host (Key decision 3), the fix is to launch with RESUME: the paneCommand becomes `exec claude --dangerously-skip-permissions --resume <claudeSessionId>` (codex uses `resume <id>`, gemini `--resume <id>`) whenever a captured `claudeSessionId` exists. The `-A` semantics make this self-selecting: the resume flag only ever executes when tmux is actually re-created, which is exactly when the live session was lost. When tmux is still alive (case 1), attach wins and the flag is inert.
Capturing / persisting / reusing the resume id (the missing mechanism the critic flagged): Codeman already learns `Session.claudeSessionId` from transcript correlation (which works here because projHash matches, Key decision 3) and persists it in `SessionState`. We thread that value into `createSessionOptions` / `respawnPaneOptions` for docker so `buildDockerLaunchCommand` can inject the resume flag on any relaunch. To make a NEW Codeman session (new `id8`) re-launched against the same case resume its predecessor's conversation, we ALSO persist `lastClaudeSessionId` on the `DockerCase` record; the quick-start docker branch seeds the new `Session` with it when the `dockerResumeOnStart` setting is on. First-ever launch has no id, so it starts fresh. This is user-decision 7 (default resume behavior).
Reconciling with stop-on-kill and with the `--restart` policy (the internal inconsistency the critic found): the container is created with `--restart no` uniformly (Codeman's idempotent create-if-missing plus boot recovery is the single recovery mechanism; a restart policy would not preserve the conversation anyway because a restarted container gets a fresh PID1/tmux). Boot recovery re-runs `buildDockerLaunchCommand` from the restored `MuxSession.docker` (`docker inspect || docker create; docker start`, then exec with resume), so a host reboot or daemon restart recreates+starts the container and resumes the conversation instead of the session vanishing. `reconcileSessions` (tmux-manager.ts ~1800-1815) must NOT hard-delete a docker session merely because no LOCAL pane exists after the local `-L codeman` server died; docker (like remote) sessions are restored from `mux-sessions.json` and relaunched. This relaunch path is explicitly part of Phase 4/Phase 3 recovery work, not assumed.
Why this over the alternatives: `docker exec` gets SIGHUP and dies when its client TTY closes, so a bare `docker exec claude` restarts the CLI on every reconnect/respawn. The inner tmux plus resume is what makes reconnect idempotent across BOTH failure surfaces. Because this durability is the single most important design point, tmux-in-image is a HARD gated prerequisite (`checkDockerTmuxAvailable`), never a silent fallback to bare exec. Rejected alternatives: ephemeral-per-run or bare-exec containers (no reattach durability); a literal `'docker'` `SessionMode` (touches dozens of switch/enum sites and diverges from the remote overlay precedent, since Docker is a LOCATION orthogonal to the 5 CLI backends).
### Key decision 2: CLI + auth delivery
One prebuilt base image (built once, contains NO secrets): `node:22-bookworm-slim` + `git tmux ripgrep ca-certificates`, `npm i -g @anthropic-ai/claude-code @openai/codex @google/gemini-cli opencode-ai`, an `agent` user, HOME dirs made writable by an arbitrary host uid via the OpenShift "gid 0, group-writable" convention (Key decision 6). Because the toolchain is baked, export is reproducible and needs no network at import time. The image name/namespace/registry and its refresh cadence are user-decision 2 (the `codeman/agent:base` placeholder implies a Docker Hub org the project may not own).
Credentials are delivered ONLY at runtime, two commit-safe channels, default convenient:
- OAuth/config-file CLIs (Claude Max/Pro, gcloud, opencode): bind-mount the host credential dirs read-write (`~/.claude`, `~/.codex`, `~/.gemini` + `~/.config/gcloud`, `~/.config/opencode`) so the common user "just works" with no in-container login. Because these are bind mounts, `docker commit` (which captures only the container's own writable layer, never bind mounts) physically cannot capture them, so exports stay secret-free.
- API-key CLIs (codex/gemini): exec-time NAME-ONLY `docker exec --env OPENAI_API_KEY --env GEMINI_API_KEY ...` (no `=value`), sourced from Codeman's own process env. Only the key NAME appears in argv (no `ps` leak), and per-exec env is never captured by `docker commit`. This is the technique Codeman already uses via `tmux setenv` for the local Codex/Gemini panes, so it composes with existing machinery.
Per-host `DockerHost.mountCredentials` defaults `true` (convenient); setting it `false` yields a SEALED profile (no host cred mounts, in-container login only) for genuinely untrusted work. CRITICAL sealed-mode export rule (the leak the critic caught): in sealed mode the in-container login writes tokens into the container's OWN writable layer, which `docker commit` DOES capture, so a full-image export of a sealed container would ship credentials. Therefore full-image export is REFUSED for `mountCredentials:false` containers by default; the user may either take a workspace-only export (always safe) or opt into a pre-commit scrub that `docker exec`s `rm -rf ~/.claude ~/.codex ~/.gemini ~/.config/gcloud ~/.config/opencode` inside the container before commit (destructive to the in-container login, which is the point). This is enforced in the export route, not left to a manifest assertion.
Per-session `envOverrides` / `effort` / `codexConfig` / `geminiConfig` / `openCodeConfig` are REJECTED at quick-start exactly like the remote branch (session-routes.ts ~1698-1710). `modelOverride` is the one deliberate difference from remote: because the docker workspace is a REAL bind-mounted host dir that Codeman scaffolds (Key decision 5 and Section 6), `updateCaseModel()` can write the `model` key into `<workspace>/.claude/settings.local.json` and the in-container `claude` reads it, so the App Settings Claude Model picker works for docker cases. `effort` is a `--effort` CLI arg applied only by the local-spawn path we bypass, so it stays rejected (surfaced honestly in the UI, not silently inert). Per-mode command customization goes through `DockerHost.commands.<mode>` (`defaultDockerCommandForMode`, mirror of `defaultRemoteCommandForMode` at remote-hosts.ts:60). NEVER bake secrets into an image layer and NEVER pass a secret via create-time `-e` (both are committed).
Rejected alternative: sealed-by-default. For a single-operator loopback tool where the agent already runs skip-permissions on the host, forcing an in-container OAuth re-login is a UX regression with little real gain. We keep sealed as an opt-in. Rejected alternative: baking a login into the image, which leaks the instant you `docker save`.
### Key decision 3: workspace mount, container CWD, and transcript correlation
Bind-mount the host workspace dir into the container at the SAME absolute path (`dst == src`, mirror the host path), and set both `Session.workingDir` and the container workdir to that host path.
Two problems this solves that the raw proposals got wrong:
- File features: `DockerCase.hostWorkspacePath` is a REAL host directory, so `Session.workingDir = hostWorkspacePath` keeps file-routes, attachments, image-watcher, and previews working on real host bytes (unlike remote, where the path is remote-only and those features no-op). All three proposals wired `casePath = <container path>`; we deliberately diverge and use the host path.
- Transcript correlation: Claude writes transcripts under `~/.claude/projects/<hash-of-CWD>/`. By mirroring the host path as the container CWD, the projHash computed inside the container equals the host-side hash Codeman's transcript/subagent/workflow watchers expect, so correlation keeps working (and, in turn, feeds the resume-id capture in Key decision 1). A `/workspace`-style fixed dst would break it. Mirror-vs-fixed is user-decision 3.
`resolveMuxAttachCwd` still returns `/tmp` for docker sessions (the LOCAL bash pane only runs `docker exec`; it never needs the workspace as its cwd), mirroring remote.
### Key decision 4: network default and the engine-specific host gateway
Default `bridge` (own netns, NAT egress, no inbound), per-case changeable to `none` (offline shell sandbox; warned because it breaks the API CLIs) or `custom` (a user-defined bridge `codeman-net-<slug>`, the chokepoint for a future egress allowlist). `host` networking and any `-p` inbound publish are structurally unrepresentable in the flag builder and schema. Rationale: every API-backed CLI (Claude, Codex, Gemini) plus npm/git needs egress, so `bridge` is the only sane functional default; `none` is reserved for `shell`.
The host-callback gateway alias is ENGINE-SPECIFIC (the critic's podman finding): Docker uses `host.docker.internal`, Podman uses `host.containers.internal` (Docker's alias only exists on recent podman). A helper `hostGatewayAlias(engine)` returns the right name; Section 2.5, the create args, the `CODEMAN_API_URL` rewrite, and the host-guard allowlist all consume it, and BOTH aliases are added to the allowlist so a mixed fleet keeps working.
### Key decision 5: hooks actually reach the host AND are actually installed
Two independent things must both be true for a hook to fire, and the raw plan wired only the first:
1. Network reachability. Claude Code hooks POST to `$CODEMAN_API_URL` (`curl -sk`). Inside a bridge container `localhost` is the container and prod binds `127.0.0.1`, so we set `--add-host <gatewayAlias>:host-gateway` on create (skipped on Docker Desktop, where the alias is native), add the gateway alias to the host guard, and provide `CODEMAN_API_URL` and the hook secret (below).
2. Hook INSTALLATION. Hooks live in `<workspace>/.claude/settings.local.json`, written by the quick-start scaffolding block (around session-routes.ts ~1776) that calls `writeHooksConfig()` / `updateCaseModel()`. The raw plan extended the `!remote` guard to `!remote && !docker`, which would SKIP that block and silently disable ALL hooks regardless of networking. For docker the workspace is a REAL bind-mounted host dir, so the scaffolding block MUST run. Precise fix: extend to `!remote && !docker` ONLY the LOCAL-CLI-availability and local-spawn guards (the ones that stat the local binary or build the local spawn command); leave the workspace-scaffolding guard at `!remote` so it runs for docker. This same decision is what makes `modelOverride` work (Key decision 2). Consequence, surfaced as user-decision 4: linking a docker case now WRITES `.claude/settings.local.json` (and the CLAUDE.md scaffold, matching local-case behavior) into the user's real host directory, a behavioral shift from "link a dir" to "link and scaffold a dir."
`CODEMAN_API_URL` derivation (the wrong-scheme bug the critic caught): prod is HTTPS-only on 3000, and `server.ts` (~2000) auto-sets `process.env.CODEMAN_API_URL = ${protocol}://${apiHost}:${port}`. Hardcoding `http://host.docker.internal:3000` fails every hook. Instead a pure helper `containerApiUrl(process.env.CODEMAN_API_URL, engine)` parses the running URL and substitutes ONLY the hostname with `hostGatewayAlias(engine)`, preserving scheme and port (`https://host.docker.internal:3000`). Unit-tested against http, https, non-default ports, and both engines. Passed as create-time `--env CODEMAN_API_URL=<derived>` (case-stable, non-secret).
Hook secret and session attribution:
- `~/.codeman/hook-secret` is bind-mounted read-only to a container path; `--env CODEMAN_HOOK_SECRET_FILE=<that path>` is create-time (a path is non-secret; the bytes ride the bind mount and are never committed).
- `CODEMAN_SESSION_ID` (which the generated hooks reference at hooks-config.ts:78-80 to attribute events) plus `CODEMAN_MUX=1` are SESSION-scoped, so they are passed at EXEC time via `docker exec --env CODEMAN_SESSION_ID=<id> --env CODEMAN_MUX=1` (non-secret, value inline is fine, and exec env is not committed). Because a `tmux` session started fresh only inherits the invoking env when it starts the SERVER, the launch chain ALSO runs `tmux -L codeman-docker setenv -g CODEMAN_SESSION_ID <id>` (and `CODEMAN_MUX`) so reattaches and newly created panes see the same values. This mirrors how Codeman already injects per-session env into tmux for the external CLIs.
Hooks-in-MVP-vs-deferred stays user-decision 4; if deferred, docker ships as explicitly hook-degraded and we lean on output-based idle detection through the docker-exec PTY.
### Key decision 6: uid / HOME / rootless enforcement / macOS Docker Desktop
The raw plan showed `--user 1000:1000` in one place and `--user "$(id -u):$(id -g)"` in another and never resolved HOME writability; this section fixes all of it.
- Linux native (docker rootful or rootless): run `--user <hostUid>:0` (host uid, GID 0). The image follows the OpenShift arbitrary-uid convention: `HOME=/home/agent`, and `/home/agent` plus the tool cache dirs (`~/.npm`, `~/.cache`, `~/.config`) are owned `root:0` and group-writable (`chmod -R g+w`, `g+s` on dirs) so a process with GID 0 can write HOME even though its UID is not 1000. This keeps workspace files host-owned (the agent's UID is the host UID) AND keeps HOME writable, so the CLIs actually start.
- Podman rootless: use `--userns=keep-id` (maps the host uid to the image's `agent` uid inside the container) instead of `--user`, so `/home/agent` is owned by the running user and workspace files are host-owned. This is a real per-engine branch in `buildDockerCreateArgs`.
- macOS Docker Desktop: `--user <macUid>` (e.g. 501) does not own the image's `/home/agent`, so non-bind HOME writes fail EACCES and the CLIs may not start; Desktop also does its own bind-mount uid translation, provides `host.docker.internal` natively (no `--add-host`), and its VM memory ceiling can cap `--memory`. Detect Desktop via `docker info` (Server OS `linuxkit` / `OperatingString` contains "Docker Desktop") and take a dedicated path: do NOT pass `--user` (run as the image's baked `agent` uid and rely on Desktop's translation for workspace access), skip `--add-host`, and note in the UI that memory caps are subject to the VM ceiling.
Rootless resource-cap enforcement (the silently-inert risk): rootless Docker without cgroup-v2 systemd delegation (`Delegate=yes`) silently IGNORES `--memory`/`--cpus`/`--pids-limit`. The probe checks `docker info` for `CgroupVersion=2` plus rootless plus delegation; if caps cannot be enforced, `checkDockerAvailable` returns `capsEnforced:false` and the link/probe surfaces "resource caps are advisory on this engine." Whether to REQUIRE delegation or ship-with-warning is user-decision 6.
## 3. Data model
New TypeScript types in `src/types/session.ts`, added right after the remote types (lines 46-99). SessionMode (line 44) is UNCHANGED.
```ts
export type DockerCommandMode = Extract<SessionMode, 'shell' | 'claude' | 'opencode' | 'codex' | 'gemini'>;
export type DockerEngine = 'docker' | 'podman';
export type DockerNetworkMode = 'bridge' | 'none' | 'custom'; // never 'host'
export interface DockerResourceLimits {
memory?: string; // '4g' -> --memory 4g --memory-swap 4g (swap==memory: real OOM cap)
cpus?: string; // '2'
pidsLimit?: number; // 512 (fork-bomb guard)
nofile?: string; // '4096:8192'
shmSize?: string; // optional; only when a tool needs /dev/shm
}
export interface DockerHost {
id: string;
label: string;
engine?: DockerEngine; // default resolved by probe (docker, else podman)
image: string; // default resolved image ref (see user-decision 2)
daemonHost?: string; // advanced: -H ssh://user@host / DOCKER_HOST
context?: string; // advanced: --context <ctx>
network?: DockerNetworkMode; // default 'bridge'
networkName?: string; // when network === 'custom'
resources?: DockerResourceLimits;
mountCredentials?: boolean; // default true (false = sealed; blocks full-image export)
hooksEnabled?: boolean; // default true (host-gateway callback wiring)
resumeOnStart?: boolean; // default true (see Key decision 1 / user-decision 7)
commands?: Partial<Record<DockerCommandMode, string>>;
extraCreateArgs?: string[]; // validated like extraSshOptions
extraExecArgs?: string[];
}
export interface DockerCase {
name: string;
type: 'docker';
hostId: string;
hostWorkspacePath: string; // absolute HOST dir: bind src + Session.workingDir
containerWorkdir?: string; // container path; default = hostWorkspacePath (mirror -> projHash match)
container?: string; // default codeman-case-<slug>
lastClaudeSessionId?: string; // captured resume id (Key decision 1)
}
export interface SessionDocker { // flattened, round-trips through mux/state (mirror SessionRemote at 91)
hostId: string;
label: string;
engine: DockerEngine;
image: string;
containerName: string;
hostWorkspacePath: string;
containerWorkdir: string;
network: DockerNetworkMode;
networkName?: string;
resources?: DockerResourceLimits;
mountCredentials: boolean;
hooksEnabled: boolean;
resumeOnStart: boolean;
daemonHost?: string;
context?: string;
commands?: Partial<Record<DockerCommandMode, string>>;
extraCreateArgs?: string[];
extraExecArgs?: string[];
configHash?: string; // drift detection (Key decision, Section 4)
}
```
- `SessionState` gains `docker?: SessionDocker` immediately after `remote?` (line 219). It persists automatically because `SessionState` is structural and `state-store.ts` stores `toState()` verbatim.
- `src/mux-interface.ts`: add `docker?: SessionDocker` to `MuxSession` (after line 38), `CreateSessionOptions` (after 81), `RespawnPaneOptions` (after 105). `MuxSession.docker` round-trips through `mux-sessions.json` automatically.
- `src/types/api.ts` `CaseInfo`: add `'docker'` to the `location` union and a `docker?: { hostId; container; image?; path; network }` display block.
- `src/services/unified-session-service.ts`: add a boolean `docker?` flag on `UnifiedSessionItem` and source rows, set from `MuxSession.docker` presence (mirror the `remote` flag at ~line 200 and the harvest at session-routes.ts:2313).
New state files (all via `dataPath()`, mirroring `remote-hosts.json` / `remote-cases.json`):
- `~/.codeman/docker-hosts.json` (reusable engine/image/network/resource profiles).
- `~/.codeman/docker-cases.json` (`name -> DockerCase`, including `lastClaudeSessionId`).
- `~/.codeman/docker-exports/` (dedicated dir for `.image.tar.gz` + `.workspace.tar.gz` + `manifest.json`; never inline in state.json; retention/pruning per Section 5).
No new `state.json` / `mux-sessions.json` files: `SessionState.docker` and `MuxSession.docker` ride the existing serialization.
## 4. Container lifecycle (exact command shapes)
All builders are PURE string functions (directly unit-testable). Host values interpolated into the outer `bash -c "..."` layer (container name, image, workdir, host paths) are `shellescape()`'d and, for user-supplied fields, schema-rejected for `$`/backtick via `NO_SHELL_META`. The escaping chain here is DEEPER than remote's single `ssh '<tmux ...>'`: the whole `docker inspect || docker create <dozens of --mount/--env/shellescaped host paths>` is interpolated into `bash -c "..."` then `JSON.stringify`'d into respawn-pane. This is a known place to get stuck, so it is covered by concrete escaping tests (Section 9), including host workspace paths containing spaces, not just a "we call shellescape" claim.
New in `src/tmux-manager.ts`:
```ts
const DOCKER_TMUX_SOCKET = 'codeman-docker';
// 'dkr' letters deliberately FAIL SAFE_MUX_NAME_PATTERN (^codeman-[a-f0-9-]+$),
// so a Codeman running INSIDE the container never adopts/resizes/respawns our session.
export function dockerTmuxSessionName(id: string): string { return `codeman-dkr-${id.slice(0, 8)}`; }
```
`buildDockerBaseArgs(docker)` (pure, in `docker-hosts.ts`, mirror of `buildSshConnectionArgs`) emits the engine prefix tokens: `docker` (or `podman`) + optional `--context <ctx>` or `-H <daemonHost>`. `buildDockerCreateArgs(docker, sessionId)` emits the `docker create` flag array (with the per-engine uid/userns branch from Key decision 6).
IMAGE PRESENCE (before any create, the auto-pull footgun the critic caught): the launch chain runs `docker image inspect <image> >/dev/null 2>&1` first; on miss it exits with a distinct message ("base image <ref> not present: build with scripts/build-agent-image.mjs or pull it") rather than triggering a blocking multi-GB auto-pull inside the tmux pane. `docker create` carries `--pull=never`. The tmux-availability probe likewise uses `docker run --rm --pull=never <image> sh -lc 'command -v tmux'` and reports the same build/pull hint if the image is absent, so the 15s-bounded probe never hangs on a pull.
CREATE (the ensure step, embedded in the launch string):
```
docker create \
--name codeman-case-myproj --hostname myproj \
--label codeman.managed=1 --label codeman.instance=<CODEMAN_INSTANCE> \
--label codeman.case=myproj --label codeman.session=<id8> \
--label codeman.confighash=<hash> \
--pull=never --init --restart no \
--user 1000:0 \
--workdir '/home/arkon/cases/myproj' \
--mount type=bind,src='/home/arkon/cases/myproj',dst='/home/arkon/cases/myproj' \
--mount type=bind,src='/home/arkon/.claude',dst='/home/agent/.claude' \
--mount type=bind,src='/home/arkon/.codeman/hook-secret',dst='/home/agent/.codeman/hook-secret',readonly \
--add-host host.docker.internal:host-gateway \
--memory 4g --memory-swap 4g --cpus 2 --pids-limit 512 --ulimit nofile=4096:8192 \
--cap-drop ALL --security-opt no-new-privileges \
--network bridge \
--env HOME=/home/agent --env TERM=xterm-256color --env COLORTERM=truecolor \
--env CODEMAN_API_URL=https://host.docker.internal:3000 \
--env CODEMAN_HOOK_SECRET_FILE=/home/agent/.codeman/hook-secret \
codeman/agent:base \
sleep infinity
```
- `--user 1000:0` shown is the Linux-native form with GID 0 (Key decision 6); it is actually `--user <hostUid>:0`, or `--userns=keep-id` for podman rootless, or omitted on Docker Desktop. The literal is illustrative only.
- Create-time `--env` carries only NON-SESSION, non-secret, case-stable values (safe to be committed): the DERIVED `CODEMAN_API_URL` (https-preserving, Key decision 5) and the hook-secret FILE PATH. `CODEMAN_SESSION_ID`/`CODEMAN_MUX` and the codex/gemini key NAMES are exec-time only.
- `codeman.instance=<CODEMAN_INSTANCE>` is REQUIRED on the label set so the boot reaper is instance-scoped (a beta/second instance must never reap prod's containers).
- `codeman.confighash` is a stable hash of the drift-relevant create args (image, resources, network, mounts, non-session env). Drift detection (user story 2, the config-never-takes-effect gap): on launch the ensure block compares the desired hash to the existing container's label; on mismatch the launch does NOT silently reuse the stale container. Instead the docker route returns a "container config changed, recreate?" action (SSE + UI confirm), and on confirm Codeman `docker rm`'s and recreates. rm destroys in-image (non-bind) state, but the workspace and transcripts survive on their bind mounts and the conversation is restored via `--resume`, so the recreate is safe. Auto-recreate-vs-prompt is a UI choice; the MVP prompts.
- `--restart no` (resolved consistently with Key decision 1; recovery is Codeman's idempotent create-if-missing, not an engine restart policy, which also matters for Podman which has no daemon).
EXEC (`buildDockerLaunchCommand`, the docker analog of `buildRemoteLaunchCommand`, TTY-correct, resume-aware). The whole thing is ONE `bash -c` string that image-checks, ensures, starts, primes tmux env, then execs:
```
docker image inspect codeman/agent:base >/dev/null 2>&1 || { echo 'Codeman: base image codeman/agent:base not present (build or pull it)'; exit 1; } ; \
docker inspect codeman-case-myproj >/dev/null 2>&1 || docker create <all create args above> ; \
docker start codeman-case-myproj >/dev/null 2>&1 || { echo 'Codeman: container codeman-case-myproj failed to start (daemon down?)'; exit 1; } ; \
exec docker exec -it \
--workdir '/home/arkon/cases/myproj' \
--env TERM=xterm-256color --env COLORTERM=truecolor \
--env CODEMAN_SESSION_ID=1a2b3c4d --env CODEMAN_MUX=1 \
--env OPENAI_API_KEY --env GEMINI_API_KEY \
codeman-case-myproj \
sh -lc 'tmux -L codeman-docker setenv -g CODEMAN_SESSION_ID 1a2b3c4d \; setenv -g CODEMAN_MUX 1 \; new-session -A -s codeman-dkr-1a2b3c4d -c '\''/home/arkon/cases/myproj'\'' '\''cd /home/arkon/cases/myproj && exec claude --dangerously-skip-permissions --resume <claudeSessionId>'\'' \; set -t codeman-dkr-1a2b3c4d status off \; set -t codeman-dkr-1a2b3c4d mouse off \; set -t codeman-dkr-1a2b3c4d prefix C-q \; set -s escape-time 0'
```
- `docker exec -it`: `-t` allocates a PTY and forwards SIGWINCH into the container so the Ink TUI re-lays-out on pane resize; `TERM`/`COLORTERM` prevent degraded rendering. `--env OPENAI_API_KEY` (name only) is present only for codex/gemini and is exec-time (never committed). `CODEMAN_SESSION_ID`/`CODEMAN_MUX` are exec-time values plus a `tmux setenv -g` prime so reattaches and new panes inherit them (Key decision 5).
- `--resume <claudeSessionId>` (codex `resume <id>`, gemini `--resume <id>`) is appended to `modeCommand` ONLY when a captured id exists; on first launch it is omitted. `new-session -A` makes the flag inert on a live-tmux reattach and effective only when tmux is re-created (Key decision 1).
- `modeCommand = docker.commands?.[mode] || defaultDockerCommandForMode(mode)` (`exec claude --dangerously-skip-permissions`, `exec bash -l`, etc.), with the resume suffix injected by the builder.
- Escaping survives every layer identically to remote in shape but deeper in nesting: `paneCommand` (`cd ... && exec ...`) is one shellescaped tmux arg, the whole `tmuxInvocation` is one shellescaped `sh -lc` arg, and the outer string is `JSON.stringify()`'d into `bash -c` by respawn-pane (tmux-manager.ts:1329).
Wire-up (extend the two existing seams to 3-way):
- createSession (tmux-manager.ts:1276): `const fullCmd = docker ? buildDockerLaunchCommand({ mode, docker, sessionId, resumeSessionId }) : remote ? buildRemoteLaunchCommand({ mode, remote, sessionId }) : localFullCmd;`
- launchCmd cd-skip (tmux-manager.ts:1327): `const launchCmd = (remote || docker) ? fullCmd : \`cd ${JSON.stringify(workingDir)} && ${fullCmd}\`;`
- respawnPane: same two edits at lines 1524 and 1542.
START / reattach-after-reboot: the ensure block (image-check, `docker inspect || docker create`, `docker start`) is fully idempotent, so boot recovery just re-runs `buildDockerLaunchCommand` from the restored `MuxSession.docker` with the persisted resume id. A rebooted host recreates the container and resumes the conversation.
DOCKER-DOWN surfacing (the PTY-exit-breaker false-trip risk): if `docker start` or `docker exec` cannot attach (daemon down, container missing), the launch prints a docker-specific message and exits, which alone would still count toward `session-pty-exit-breaker` and show a generic "respawn breaker tripped" push. To avoid masking the cause, the docker reattach path runs a fast `checkDockerAvailable` pre-flight: if the daemon/container is unreachable, Codeman broadcasts a docker-specific error (SSE + push, "container <name> is not running / daemon down") and SKIPS the auto-reattach that would trip the breaker, rather than fast-looping `docker exec`.
STOP / KILL (`killSession` Strategy 3c, right after remote's Strategy 3b at tmux-manager.ts:1719, guarded by `IS_TEST_MODE`):
```ts
if (session.docker) {
// best-effort, fire-and-forget, timeout-bounded so it never blocks the local kill
execAsync(buildDockerKillCommand({ docker: session.docker, sessionId }), { timeout: EXEC_TIMEOUT_MS }).catch(() => {});
}
```
`buildDockerKillCommand` emits: `docker exec codeman-case-<slug> tmux -L codeman-docker kill-session -t codeman-dkr-<id8> ; docker stop -t 10 codeman-case-<slug>`. Stopping frees CPU/RAM and, per Key decision 1, is safe for conversation continuity because the NEXT launch resumes from the bind-mounted transcript via `--resume`. Whether to stop at all (RAM vs instant live-agent reattach) is user-decision 6/1 (reframed honestly). The bind-mounted workspace and transcripts always survive on the host.
REMOVE: only on explicit case delete (`docker rm -f codeman-case-<slug>`), gated behind an "export first?" UI prompt because rm destroys any in-image (non-bind) state. Instance-scoped boot reaper (fixing the racy/cross-instance reaper): after `docker-cases.json` is loaded AND after `restoreMuxSessions` has run, enumerate `docker ps -a --filter label=codeman.managed=1 --filter label=codeman.instance=<CODEMAN_INSTANCE> --format '{{.Names}}\t{{index .Labels "codeman.case"}}'` and `docker rm -f` only containers whose case is gone from THIS instance's `docker-cases.json`. The instance filter is what stops a beta reaping prod's containers (the exact cross-instance hazard the project memory warns about).
AVAILABILITY PROBE (`docker-hosts.ts`, timeout-bounded like `checkRemoteTmuxAvailable`'s 15s, `IS_TEST_MODE` no-op):
```
docker info --format '{{json .}}' # server up, CgroupVersion, rootless, OS (Desktop detect), cap-delegation
docker image inspect <image> --format '{{.Id}}' # image PRESENT (no auto-pull)
docker run --rm --pull=never <image> sh -lc 'command -v tmux' # tmux-in-image gate (hard prerequisite), only if image present
```
`checkDockerAvailable()` returns `{ ok, engine, rootless, isDesktop, cgroupV2, capsEnforced }` (parse `SecurityOptions` for `name=rootless`, `CgroupVersion`, delegation, and Server OS for Desktop). `checkDockerTmuxAvailable(host)` returns a structured result with a user-facing error and correct install hint (NOT `npm install -g`; the hint is "build/pull the base image" for a missing image and "install docker or podman" for a missing engine).
IN-CONTAINER CLI VERSION (fixing the #154 wheel-forwarding regression): the raw plan skipped the LOCAL `cliVersion` probe for docker (correct, since it reports the HOST claude) but left `cliVersion` undefined, which disables trackpad wheel-forwarding. Instead, for docker sessions Codeman runs an IN-CONTAINER probe `docker exec <container> claude --version` (bounded, `IS_TEST_MODE` no-op) and feeds THAT into `cliVersion`. This also means a stale baked CLI is visible; combined with the rebuild-cadence in user-decision 2, agents are not silently pinned to an old claude.
## 5. Export / Import
EXPORT is a concurrency-bounded job (reuse `runWithConversionLimit` from `document-conversion-limiter.ts` so N simultaneous exports cannot fork-bomb the host). Route `POST /api/docker-cases/:name/export`.
Preconditions (the consistency and leak risks the critic caught):
- Sealed guard: if `mountCredentials:false`, full-image export is REFUSED unless the caller explicitly opts into the pre-commit scrub (Key decision 2). Workspace-only export is always allowed.
- Quiesce + free-space: require the session idle, then `docker pause` the container spanning BOTH the workspace tar AND the commit so the two artifacts are mutually consistent (the raw plan paused only the commit, leaving the bind-mount tar to run against a mid-write agent). Before any heavy step, precheck free space in the exports dir and in `/var/lib/docker`; if below `DOCKER_EXPORT_MIN_FREE_BYTES`, refuse with a clear error (a full `/var/lib/docker` wedges the daemon and breaks EVERY session on the host).
Steps (all cleanup in try/finally so a mid-way failure never orphans an intermediate image or leaves the container paused):
1. `docker commit -c 'LABEL codeman.exported=1' codeman-case-<slug> codeman/export-<slug>:<ts>` (unique tag per export defeats the stale-image trap). Optional pre-commit scrub in sealed mode as above; also blank instance-specific committed env (`-c 'ENV CODEMAN_API_URL='` etc.) so the image carries no stale host references.
2. `docker save codeman/export-<slug>:<ts> | gzip` streamed in fixed 8192-byte chunks to `~/.codeman/docker-exports/<slug>-<ts>.image.tar.gz`. Uses `docker save` (layers + repo:tag + CMD), never `docker export` (flat rootfs), so restore is a trivial `docker load`.
3. `tar --numeric-owner -C <hostWorkspacePath> -czf <slug>-<ts>.workspace.tar.gz .` while paused (the bind-mounted workspace is NOT in the image, so it travels separately and consistently).
4. Write `manifest.json`: schema version, caseName, image tag, engine, containerWorkdir, resource/network config, codeman version, base-image digest, createdAt, per-member sha256, `mountCredentials`, and `secretFree` (true only for convenient-mode or scrubbed-sealed exports).
5. `docker rmi codeman/export-<slug>:<ts>` in the `finally` (delete the intermediate committed image regardless of success), then `docker unpause`.
The three files are wrapped in one bundle `<slug>-<ts>.codeman-container.tgz` and offered as a downloadable artifact through the existing file-routes streaming + attachment-registry handoff.
Retention / disk budget (user-decision 3): `docker-exports/` is capped at `DOCKER_EXPORT_KEEP` most-recent bundles with an auto-prune on each new export, plus the free-space precheck above. Workspace scrub: the WORKSPACE tar gets a scan/warn pass for agent-created `.env` / `.git/credentials` (a distinct leak channel from container creds). A lighter "workspace-only" export (just the workspace tar, no commit/save) is the fast default for 24h+ runs; full-image is the explicit heavier option (user-decision 7 in the original list, now decision on the default button below).
What travels: the baked toolchain image plus any in-image writes, and the workspace tar. What does NOT travel: bind-mounted credentials (physically excluded from commit) and anything that lived only in a bind mount. Secret-free by construction in convenient mode, and enforced (refuse-or-scrub) in sealed mode.
IMPORT `POST /api/docker-cases/import` (untrusted-bundle containment, the traversal/overwrite risk): stream the uploaded bundle, validate every manifest checksum BEFORE any extraction or load. Extract the workspace tar with `tar --no-absolute-names -C <fresh dir>` PLUS per-entry validation rejecting any member whose normalized path escapes the destination (leading `/` or `..` components). `gunzip | docker load` the image, then RE-TAG the loaded image id into a quarantined namespace `codeman/imported-<slug>:<ts>` and NEVER allow the load to overwrite `codeman/agent:base` or any pre-existing tag (capture the loaded id, ignore the bundle's repo:tag). Create a NEW `DockerCase` pointing at the quarantined image with THIS host's mounts/creds and the manifest's resource/network config, and recreate the container hardened (cap-drop ALL, no-new-privileges, non-root, `--pull=never`, CMD overridden to `sleep infinity`). The destination supplies its own login, so credentials never cross machines. Plus `GET /api/docker-exports` (list) and `DELETE /api/docker-exports/:filename`, all behind Codeman's existing auth / loopback-default / host-guard / Origin-CSRF stack.
## 6. Codeman integration (file-by-file, mirroring the remote-SSH feature)
- `src/types/session.ts`: add `DockerCommandMode`, `DockerEngine`, `DockerNetworkMode`, `DockerResourceLimits`, `DockerHost`, `DockerCase`, `SessionDocker` (Section 3). Add `docker?: SessionDocker` to `SessionState` after line 219. SessionMode (line 44) UNCHANGED.
- `src/mux-interface.ts`: add `docker?: SessionDocker` to `MuxSession` (38), `CreateSessionOptions` (81), `RespawnPaneOptions` (105).
- `src/docker-hosts.ts` (NEW, direct mirror of `src/remote-hosts.ts`): `readDockerHosts`/`writeDockerHosts`/`readDockerCases`/`writeDockerCases` (via `dataPath`, including `lastClaudeSessionId` read/write), `defaultDockerCommandForMode` (mirror line 60), `dockerDisplayPath` (`container:/path`, mirror `remoteDisplayPath` at 205), `toSessionDocker(host, case)` (mirror `toSessionRemote` at 212), `buildDockerBaseArgs`/`buildDockerCreateArgs` (per-engine uid/userns branch), `hostGatewayAlias(engine)`, `containerApiUrl(processApiUrl, engine)` (scheme+port-preserving, unit-tested), `checkDockerAvailable`/`checkDockerTmuxAvailable`/`probeDockerCliVersion` (15s-bounded, `IS_TEST_MODE` no-op), a config-hash helper for drift, its own POSIX `shellescape` copy (mirror line 83). `const IS_TEST_MODE = !!process.env.VITEST;` gates every real `docker` invocation.
- `src/tmux-manager.ts`: add `DOCKER_TMUX_SOCKET`, `dockerTmuxSessionName`, `buildDockerLaunchCommand` (resume-aware, image-check, env-prime), `buildDockerKillCommand` (Section 4). Extend the two `fullCmd` ternaries (1276, 1524) and the two `launchCmd` cd-skips (1327, 1542). Add `killSession` Strategy 3c after 1719. Ensure `reconcileSessions` (~1800-1815) does NOT hard-delete docker sessions on local-tmux death (recovery relaunch path).
- `src/session.ts`: add `_docker?: SessionDocker` field (mirror `_remote` at 403), constructor arg (477), assignment (550). Thread `docker: this._docker` and `resumeSessionId: this._claudeSessionId` into BOTH `createSessionOptions` and `respawnPaneOptions` in `startInteractive` (1352/1370) and the second path (1740/1750). Emit `docker: this._docker` in `toState()` (1010). Replace the LOCAL cliVersion probe at 1320 for docker with the IN-CONTAINER `probeDockerCliVersion` (do not merely skip it). Extend `resolveMuxAttachCwd(workingDir, remote, docker)` (215) to return `/tmp` when `docker` is set. On claudeSessionId capture, persist it to the owning `DockerCase.lastClaudeSessionId`.
- `src/web/server.ts`: in `restoreMuxSessions` (2160), add `docker: muxSession.docker ?? savedState?.docker` to the `new Session({...})` call (2195-2216), and skip docker in the same `isExternalCliMode`/Ralph recovery guards as remote. Register the instance-scoped boot reaper to run AFTER docker-cases load and AFTER `restoreMuxSessions`. Ensure `CODEMAN_API_URL` derivation reads the SAME `process.env.CODEMAN_API_URL` the server sets at ~2000.
- `src/web/schemas.ts`: add `DockerHostSchema` and `DockerCaseLinkSchema` (below). The three mode enums (177/373/705) and `QuickStartSchema` (368) UNCHANGED (docker resolves by `caseName` lookup like remote).
- `src/web/routes/session-routes.ts`: import the docker helpers from `../../docker-hosts.js`. Add a docker branch in `/api/quick-start` parallel to the remote branch (1686-1720): `readDockerCases` -> find by `caseName` -> `readDockerHosts` -> find by `hostId`; reject `envOverrides`/`effort`/`codexConfig`/`geminiConfig`/`openCodeConfig` (but ACCEPT `modelOverride`, which flows via scaffolded `settings.local.json`); run `checkDockerAvailable` + `checkDockerTmuxAvailable` (image-present, engine, caps-enforced); surface `capsEnforced:false` and Desktop notes; set `casePath = dockerCase.hostWorkspacePath` (REAL host dir), `docker = toSessionDocker(host, dockerCase)`, and seed `resumeSessionId` from `dockerCase.lastClaudeSessionId` when `resumeOnStart`. Extend the LOCAL-availability and local-spawn guards (around 1796/1810) to `!remote && !docker`, but DO NOT extend the workspace-scaffolding guard (~1776, `writeHooksConfig`/`updateCaseModel`), which MUST run for docker. Pass `docker` into `new Session` (1847); `autoConfigureRalph` (1853) gated on `!docker`. Add `docker: m.docker !== undefined ? true : undefined` to the unified harvest (2313).
- `src/web/routes/case-routes.ts`: import the docker read/write/check helpers + schemas. Add a docker listing loop in `GET /api/cases` (mirror 94-119, `location: 'docker'`, `docker: {...}` via `dockerDisplayPath`). Add `/api/docker-hosts` GET/POST/PUT/DELETE (mirror 168-204) and `POST /api/cases/docker-link` (mirror 206-232; run `checkDockerAvailable`/`checkDockerTmuxAvailable` at link time; broadcast `CaseLinked` with `type: 'docker'`). Add a docker-unlink branch to `DELETE /api/cases/:name` (mirror 288-296; `docker rm -f`; broadcast `CaseDeleted` `type: 'docker-unlinked'`). Add the docker branch to single-case `GET` (mirror 358-368). Add `POST /api/docker-cases/:name/export`, `/import`, `GET/DELETE /api/docker-exports`, and a `POST /api/docker-cases/:name/recreate` (drift confirm) per Sections 4 and 5.
- `src/web/sse-events.ts` + `src/web/public/constants.js`: reuse `CaseLinked`/`CaseDeleted` for CRUD. Add `docker:exportProgress`, `docker:exportComplete`, `docker:importComplete`, `docker:configDrift`, and `docker:containerError` to BOTH registries (kept in sync per CLAUDE.md).
- Frontend `src/web/public/index.html` (~1831): add a Docker `modal-tab-btn` next to Remote; add a `#case-docker` panel mirroring `#case-remote` with `dockerCaseName`, `dockerHostWorkspacePath`, `dockerContainer`, `dockerImage`, `dockerHostId`, and an Advanced `<details>` for network mode, resource caps, `mountCredentials`, `resumeOnStart`, and remote daemon. Surface a "scaffolds .claude into this host dir" note (user-decision 4) and a "resource caps advisory on this engine" warning when `capsEnforced:false`.
- Frontend `src/web/public/session-ui.js`: `formatCasePickerLabel` (48) + `buildCasePickerOptions` (71-73) handle `location === 'docker'` (`name @ container`, add container/image to the search haystack); `resetCaseModalFields` (~1514) add a `dockerFields` array; `switchCaseModalTab` (1573/1580/1597) handle `'case-docker'`; `submitCaseModal` add the docker branch; new `linkDockerCase()` (mirror `linkRemoteCase` at 1689) POSTing `/api/docker-hosts` then `/api/cases/docker-link`, sending omitted optionals as `undefined` (spread `...(x ? {x} : {})`, never `null`, per the Zod `.optional()`-rejects-null gotcha); `runClaude` (520) / `runShell` (702) extend the `location === 'remote'` routing to also match `'docker'`; `runOpenCode`/`runCodex`/`runGemini` (792/846/900) make the `isRemote` checks `isRemoteOrDocker` so local status probes are skipped. In the session-options Summary tab, note that `effort` is inert for docker (rejected) while `model` IS honored via `settings.local.json`.
- Frontend `src/web/public/panels-ui.js` (425-426): add `caseItem?.docker?.path`/`container` to the case-search fields.
Schemas (`src/web/schemas.ts`), mirroring `RemoteHostSchema` (299) / `RemoteCaseLinkSchema` (351):
```ts
export const DockerHostSchema = z.object({
id: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid docker host id'),
label: z.string().min(1).max(100),
engine: z.enum(['docker', 'podman']).optional(),
image: z.string().min(1).max(512).regex(/^[a-zA-Z0-9][\w./:@-]*$/, 'Invalid image ref').regex(NO_SHELL_META),
daemonHost: z.string().max(512).regex(NO_SHELL_META, 'Invalid daemon host').optional(),
context: z.string().max(128).regex(/^[a-zA-Z0-9._-]+$/, 'Invalid context').optional(),
network: z.enum(['bridge', 'none', 'custom']).optional(),
networkName: z.string().max(128).regex(/^[a-zA-Z0-9][a-zA-Z0-9_.-]+$/).optional(),
resources: z.object({
memory: z.string().regex(/^\d+[bkmg]?$/i).optional(),
cpus: z.string().regex(/^\d+(\.\d+)?$/).optional(),
pidsLimit: z.number().int().positive().max(100000).optional(),
nofile: z.string().regex(/^\d+:\d+$/).optional(),
shmSize: z.string().regex(/^\d+[bkmg]?$/i).optional(),
}).strict().optional(),
mountCredentials: z.boolean().optional(),
hooksEnabled: z.boolean().optional(),
resumeOnStart: z.boolean().optional(),
commands: RemoteCommandOverridesSchema, // reuse the shared shape
extraCreateArgs: z.array(z.string().min(1).max(1024).regex(NO_SHELL_INJECTION).refine(noCommandSubstitution)).max(32).optional(),
extraExecArgs: z.array(z.string().min(1).max(1024).regex(NO_SHELL_INJECTION).refine(noCommandSubstitution)).max(32).optional(),
});
export const DockerCaseLinkSchema = z.object({
name: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid case name format'),
hostId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid docker host id'),
hostWorkspacePath: z.string().min(1).max(2000).regex(/^\//, 'Path must be absolute').regex(NO_SHELL_META, 'Invalid characters in workspace path'),
containerWorkdir: z.string().min(1).max(2000).regex(/^\//).regex(NO_SHELL_META).optional(),
container: z.string().min(2).max(128).regex(/^[a-zA-Z0-9][a-zA-Z0-9_.-]+$/, 'Invalid container name').optional(),
});
```
`NO_SHELL_META` (rejects `$`/backtick, schemas.ts:297) is REQUIRED on `image`, `hostWorkspacePath`, `containerWorkdir`, and `container`, because all four reach the outer `bash -c "..."` double-quote layer where `$(...)`/backtick re-expose, exactly the reason `remotePath`/`identityFile` use it. `--privileged` and any `-v /var/run/docker.sock` are structurally unrepresentable (never emitted by the builder, never accepted by the schema).
## 7. Security model
- Hardening flags on every create: `--cap-drop ALL`, `--security-opt no-new-privileges` (NOT auto-set by rootless Docker or Podman, so always explicit), the uid/userns branch of Key decision 6 (never container-root; workspace files stay host-owned and HOME stays writable via GID 0), `--pids-limit` (fork-bomb guard), `--memory` with `--memory-swap == --memory` (real OOM cap), `--ulimit nofile`, `--init`, `--pull=never`. NEVER `--privileged`, NEVER mount the docker socket into the agent container. `--storage-opt size=` is emitted ONLY after the probe confirms overlay2-on-xfs-pquota or btrfs (the AICE-class silently-ignored trap); otherwise it is omitted and the UI does not advertise a size cap. Resource caps are advertised as ENFORCED only when the probe reports `capsEnforced:true`; under non-delegated rootless they are labeled advisory (user-decision 6).
- Engine: prefer whichever the probe finds, Podman-rootless first for security (a container-root breakout lands as an unprivileged host user). Rootless bind-mount ownership uses `--userns=keep-id` (Podman) vs `--user <hostUid>:0` (Docker), so real per-engine branching lives in `buildDockerCreateArgs`. Docker Desktop takes its own uid path (Key decision 6).
- Blast radius (the combined-posture the critic asked to surface, user-decision 5): the default convenient profile mounts an arbitrary host workspace dir RW (host-owned, mirrored path) AND host `~/.claude`/`~/.codex`/`~/.gemini`/`~/.config/gcloud`/`~/.config/opencode` RW into a NETWORK-ENABLED container. Container-run agent code can therefore read/modify those host trees and reach the network simultaneously. This is still a strict improvement over today's on-host skip-permissions execution, but the user must accept the combined posture explicitly; the sealed profile plus `network:none` is the mitigation for genuinely untrusted work.
- Secret handling: creds arrive ONLY as bind-mounted files (default) or exec-time NAME-ONLY `--env` (codex/gemini keys), NEVER as create-time `-e` and NEVER as an image layer. Sealed-mode export is refuse-or-scrub (Section 5), closing the sealed-leak inversion.
- CLAUDE.md "Multi-CLI prefix discipline": the exec-time name-only env is restricted to the CLI-specific keys per mode (Claude: none with OAuth mount; Codex: `OPENAI_API_KEY`/`CODEX_API_KEY`; Gemini: `GEMINI_API_KEY`/`GOOGLE_*`), never a blanket forward. `envOverrides` is rejected for docker, so the `ALLOWED_ENV_PREFIXES` allowlist is not widened.
- hook-secret: bind-mounted read-only, referenced via `CODEMAN_HOOK_SECRET_FILE` (a path, non-secret); the secret bytes never enter env or the image. Both `host.docker.internal` and `host.containers.internal` are added to the host-guard allowlist so the in-container hook curl's Host header passes on either engine.
- Host guard / instance isolation: the in-container tmux socket (`codeman-docker`) and name (`codeman-dkr-<id8>`) deliberately FAIL a container-internal Codeman's `SAFE_MUX_NAME_PATTERN`, so a nested Codeman never adopts our session (unit-asserted). The boot reaper is instance-scoped by the `codeman.instance` label so a beta never reaps prod. Any remote-daemon (`-H`/`--context`) mode is host-root-equivalent and stays strictly behind the existing auth/loopback/host-guard/Origin-CSRF stack.
- Import containment: untrusted bundles are checksum-validated, extracted with traversal guards, and loaded into a quarantined image namespace (never overwriting the base image), then run with the same hardening.
## 8. Phased implementation (branch: `feat/docker-session-mode`)
Each phase is independently testable; per CLAUDE.md, end-to-end test in the real env before COM. All new docker IO paths carry `const IS_TEST_MODE = !!process.env.VITEST;` and no-op under it; the pure command builders are tested directly.
- Phase 0: base image + engine probe. Author `docker/agent.Dockerfile` (OpenShift arbitrary-uid HOME) and `scripts/build-agent-image.mjs` (build or pull the base image; digest recorded). Add `checkDockerAvailable`/`checkDockerTmuxAvailable`/`containerApiUrl`/`hostGatewayAlias` (IS_TEST_MODE no-op) and `GET /api/docker/status`. Test: probe stub returns available/caps/Desktop flags under VITEST; `containerApiUrl` preserves scheme+port and swaps host per engine; status route returns the envelope.
- Phase 1: types + storage + schemas. Add all types (Section 3), `src/docker-hosts.ts`, `DockerHostSchema`/`DockerCaseLinkSchema`. Test: `docker-hosts.test.ts` (round-trip incl. `lastClaudeSessionId`, display path, config-hash stability); `docker-exec-options.test.ts` (schema rejects `$`/backtick in image/workdir/container).
- Phase 2: tmux-manager builders. Add `DOCKER_TMUX_SOCKET`, `dockerTmuxSessionName`, `buildDockerLaunchCommand` (resume-aware, image-check, env-prime), `buildDockerKillCommand`; wire the two ternaries + two cd-skips + Strategy 3c; harden `reconcileSessions` against docker hard-delete. Test (pure strings): adopt-proof name fails `SAFE_MUX_NAME_PATTERN`; image-check precedes create; `new-session -A` idempotent; resume flag present only when a resume id is passed; `--pull=never` present; instance label present; escaping survives `bash -c` -> `docker exec` -> `sh -lc` -> tmux WITH a host workspace path containing spaces.
- Phase 3: session.ts + mux + recovery. Add `_docker` + `resumeSessionId` threading, in-container cliVersion probe, `resolveMuxAttachCwd`, mux-interface fields, `restoreMuxSessions` passthrough, instance-scoped reaper wiring, claudeSessionId -> `DockerCase.lastClaudeSessionId` persistence, unified flag. Test: `toState()` emits docker; a persisted docker session round-trips through mux/state; a relaunch injects the persisted resume id (mock mux); reaper only targets this instance's orphaned containers.
- Phase 4: routes + first real e2e. case-routes CRUD + listing + drift-recreate; session-routes quick-start branch (scaffolding RUNS, local-availability guards skip, model accepted, effort/config rejected). Manual e2e on a real docker host: docker-host create -> docker-link -> quick-start; confirm the pane runs `claude` in the container, files land host-owned, a Codeman restart reattaches the SAME live agent, and a `docker stop` followed by relaunch RESUMES the conversation.
- Phase 5: hooks connectivity + installation. host-gateway (per engine), derived `CODEMAN_API_URL`, hook-secret mount, `CODEMAN_SESSION_ID`/`CODEMAN_MUX` exec-env + tmux setenv, host-guard allowlist, and the scaffolding write into the real workspace. Manual e2e: trigger a permission prompt from inside the container and confirm it surfaces; verify hook payloads carry the right session id. If deferred, ship docker as explicitly hook-degraded and verify output-based idle detection through the docker-exec PTY.
- Phase 6: export/import + GC + disk safety. quiesce+pause span, free-space precheck, commit+save+gzip + workspace tar + manifest + streaming download; sealed-mode refuse-or-scrub; retention/auto-prune; import with checksum validation + traversal guard + quarantined re-tag; drift-recreate; boot reaper; `runWithConversionLimit` cap; `docker rmi` in finally. Manual e2e: export, `docker load` on a second machine (or fresh case), import, confirm toolchain + workspace restored and NO creds present; attempt a sealed full-image export and confirm it is refused-or-scrubbed; attempt a `../` bundle and confirm it is rejected.
- Phase 7: frontend. Docker tab, `linkDockerCase`, run wiring, case-picker labels, panels search, caps-advisory + scaffold-warning + effort-inert notes. Verify with Playwright (`waitUntil: 'domcontentloaded'`, 3-4s settle) that the Docker tab renders and a linked docker case appears in the picker.
- Phase 8: docs + COM. Update CLAUDE.md (a "Docker cases" Key Pattern paragraph mirroring remote-SSH, plus the new state files, routes counts, and the resume/durability model), `docs/docker-cases.md`, then COM per the standard flow.
## 9. Test plan
- Unit (pure, CI-safe, mirror `test/remote-hosts.test.ts` / `test/remote-ssh-options.test.ts`):
- `test/docker-hosts.test.ts`: storage round-trip (incl. `lastClaudeSessionId`), `dockerDisplayPath`, `defaultDockerCommandForMode`, `toSessionDocker`, `containerApiUrl` (http/https, custom port, docker vs podman gateway), config-hash stability/drift, `buildDockerCreateArgs` flag ordering (cap-drop/no-new-privileges/memory==memory-swap/instance-label/`--pull=never` present; host/privileged/socket absent; per-engine uid vs `--userns=keep-id`).
- `test/docker-exec-options.test.ts`: `buildDockerLaunchCommand`/`buildDockerKillCommand` string shape and escaping through `bash -c` -> `docker exec` -> `sh -lc` -> tmux, including a workspace path with spaces; resume flag present only with a resume id; image-presence check precedes create; `dockerTmuxSessionName` fails `SAFE_MUX_NAME_PATTERN`; schema rejects `$`/backtick in image/workdir/container/name; `linkDockerCase`-shaped bodies with omitted optionals validate (no `null` on the wire).
- Probe no-op: `checkDockerAvailable`/`checkDockerTmuxAvailable`/`probeDockerCliVersion` return canned values under VITEST and never spawn.
- Integration (route tests via `app.inject()`, docker no-op'd): `/api/docker-hosts` CRUD; `/api/cases/docker-link` dup-check + broadcast; `GET /api/cases` includes the docker case with `location: 'docker'`; `/api/quick-start` docker branch rejects `envOverrides`/`effort`/config but ACCEPTS `modelOverride`, runs the workspace-scaffolding path, and constructs a session with `docker` set + seeded resume id; `DELETE /api/cases/:name` docker-unlink; export refuse-or-scrub for sealed; import traversal rejection; reaper instance-scoping (label filter). Pick a unique port only if a live-server test is added (search `const PORT =`; 3150+).
- Manual end-to-end (real docker daemon, the mandatory "always end-to-end test" gate): build the base image; link a docker case; quick-start `claude`; verify OAuth via the mounted `~/.claude`, transcript correlation (subagent/workflow watchers show the session), host-owned files, and a working permission-prompt hook; reattach after a Codeman PROCESS restart (SAME live agent); `docker stop` then relaunch and confirm conversation RESUME; reboot-equivalent (daemon restart) and confirm boot recovery recreates+resumes; change the host's memory/image and confirm the drift-recreate prompt fires; export (convenient) and confirm the tar `docker load`s with no creds; attempt a sealed full-image export and confirm refuse-or-scrub; import into a fresh case; delete the case and confirm `docker rm -f` plus instance-scoped reaper GC; confirm a docker-down state surfaces a docker-specific error and does NOT trip the generic PTY-exit breaker.
## 10. Open decisions for the user
1. Credential + blast-radius posture (combined). Convenient default bind-mounts host `~/.claude` etc. RW AND an arbitrary host workspace RW into a network-enabled container, so container-run agent code can read/modify those host trees and reach the network at the same time. Recommended: convenient default plus a per-host SEALED opt-in (`mountCredentials:false` + `network:none`) for untrusted work. Please confirm you accept the combined arbitrary-workspace-plus-egress-plus-host-creds posture for the default profile (it is still a net improvement over today's on-host skip-permissions execution).
2. Base image ownership, registry, and freshness. The `codeman/agent:base` placeholder implies a Docker Hub org the project may not own. Pick the real registry/namespace (GHCR under the repo is the natural fit), decide digest pinning, and set a REBUILD CADENCE so agents are not stuck on a stale baked `claude` (the in-container version probe surfaces staleness, but something must trigger rebuilds). Choose: pull a pinned published image, build locally on first use via `scripts/build-agent-image.mjs`, or both.
3. Container CWD strategy. Mirror the host workspace path inside the container (recommended: makes transcript projHash correlate, file features and resume capture work) vs a fixed `/workspace` (simpler mount, breaks watcher correlation). Please confirm the mirror approach.
4. Hooks in the MVP AND workspace scaffolding. Making docker hooks fire requires WRITING `.claude/settings.local.json` (and the CLAUDE.md scaffold) into the user's REAL linked host directory, a behavioral shift from "link a dir" to "link and scaffold a dir." Choose: wire hooks + scaffolding now (Phase 5, recommended, and it also enables the model picker), or ship docker as explicitly hook-degraded (no permission prompts / hook-idle) for v1 and add later. Confirm you are OK with Codeman mutating the linked host workspace.
5. Session-kill teardown and RESUME (reframed honestly). `docker stop` on session kill is not merely "free RAM vs instant reattach": it destroys the in-container live agent, and the conversation survives ONLY because the next launch runs `--resume` from the bind-mounted transcript. Choose: keep the container running (costs RAM, preserves the exact live in-flight agent) vs stop and rely on `--resume` (frees RAM, may lose uncommitted in-flight tool state). Case-delete always `docker rm -f`.
6. Rootless enforcement posture. Under rootless without cgroup-v2 systemd delegation, `--memory`/`--cpus`/`--pids-limit` are SILENTLY ignored. Choose: REQUIRE delegation (refuse to link a host that cannot enforce caps) or ship-with-warning ("resource caps are advisory on your engine"). The probe reports `capsEnforced` either way.
7. Default resume behavior. Should a re-linked or re-run docker case default to resuming its last conversation (`resumeOnStart:true`, using `DockerCase.lastClaudeSessionId`) rather than starting clean? This is the crux of making the durability story real and is the recommended default, but it changes user-visible behavior (a new session in an existing case continues the prior conversation).
8. Export defaults and disk budget. Default export button: workspace-only (fast, small, files-only, recommended for 24h+ runs) vs full-image (reproducible env, multi-GB). Also set the retention cap (max retained exports), the auto-prune policy, and the free-space threshold below which export is refused (a full `/var/lib/docker` breaks EVERY session on the host, not just docker ones).
9. Remote docker daemon (`-H ssh://...` / `--context`). Support in the MVP (composes with remote hosts, adds host-root trust surface) or local-daemon-only first.
10. Podman parity depth. Full `--userns=keep-id` plus Quadlet boot-persistence, or Docker-first with Podman as best-effort and boot-persistence via Codeman's idempotent create-if-missing only. Note the podman host alias is `host.containers.internal`, already handled per engine.
-112
View File
@@ -1,112 +0,0 @@
# Docker cases
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` / `pi` all work inside the container.
## One-time setup: build the base image
The container needs a base image with the agent toolchain (node, the CLIs, git, tmux). Build it locally once:
```bash
node scripts/build-agent-image.mjs # builds codeman/agent:base
# options: --engine docker|podman --image <ref> --no-cache
```
The image is **secret-free**: credentials are delivered at runtime (bind mounts or `docker exec --env`), never baked in, so exports never leak them.
⚠️ **Re-build with `--no-cache`, always.** The CLIs are installed in a single `RUN npm install -g` layer, so a plain rebuild re-uses it from the Docker layer cache and the CLIs stay frozen at whatever versions the image was **first** built with, however long ago that was. Editing the Dockerfile does not help unless the edit lands at or above that line: a change appended below it leaves the npm layer cached and only runs the new step. Observed 2026-08-06: a rebuild silently kept a stale `@openai/codex@0.144.6` whose aliased platform binary had not installed, so every `codex` docker case died with `Missing optional dependency @openai/codex-linux-x64` while the build itself reported success.
```bash
node scripts/build-agent-image.mjs --no-cache
```
A zero exit code only proves the layers ran, not that the toolchain works. Verify by actually executing each CLI in the image, and check the build log for `Using cache` lines:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy pi; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
Antigravity (`agy`) is the one CLI not installed from npm (Google ships a standalone binary), so it has its own Dockerfile step and adds roughly 190MB; a full image lands near 1.6GB. Pi also gets its own step, because upstream documents installing it with `--ignore-scripts` and that flag must not silently change how the other four npm CLIs install.
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md).
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
Click the checkbox's **Container settings** to optionally tweak the predefined defaults, including a **Template** picker:
| Template | Memory | CPUs | GPUs |
|----------|--------|------|------|
| Small | 2 GB | 1 | none |
| Medium (default) | 4 GB | 2 | none |
| Large | 8 GB | 4 | none |
| GPU | 8 GB | 4 | all (needs the NVIDIA container toolkit) |
**Disk is elastic** — the container's storage grows automatically as data flows in; there is no fixed cap (bounded only by host disk). Any tweaked setting creates a dedicated per-case host so it never changes the shared `default`.
## Create a docker case (full control)
App → **New case → Docker** tab:
- **Case Name** / **Workspace Path**: the workspace is a real HOST directory bind-mounted into the container at the same path. Codeman scaffolds `CLAUDE.md` + `.claude/settings.local.json` (hooks) into it, and file previews / attachments work on the real bytes.
- **Host ID**: a reusable docker host profile (image, network, resources). Reuse the same ID across cases to share settings.
- **Network**: `bridge` (internet on, default), `none` (fully isolated), or a `custom` bridge.
- **Advanced**: memory / CPU caps, **Mount host credentials** (on = your existing `~/.claude` login just works; off = a sealed sandbox you log into inside the container), **Resume last conversation on relaunch**.
Then run it like any case (Run Claude / Run Shell / …). The first launch creates the container (`codeman-case-<name>`); subsequent sessions attach to the same one.
Equivalent API:
```bash
curl -X POST localhost:3000/api/docker-hosts -d '{"id":"local","label":"Local","image":"codeman/agent:base"}'
curl -X POST localhost:3000/api/cases/docker-link -d '{"name":"sandbox","hostId":"local","hostWorkspacePath":"/home/you/projects/sandbox"}'
curl -X POST localhost:3000/api/quick-start -d '{"caseName":"sandbox","mode":"claude"}'
```
## Lifecycle
- **Reconnect after a Codeman restart** lands back in the same live agent (the in-container tmux survives).
- **Container stop / host reboot** restarts the container and **resumes** the last conversation from the bind-mounted transcript. Claude sessions launch with a pinned conversation id (`--session-id <sessionId>`, with a `--resume` fallback when the transcript already exists), and the case remembers its last conversation (`lastClaudeSessionId`), so a relaunch after the container was stopped, rebooted, or recreated continues where it left off.
- **Killing one session** only kills that session's in-container tmux session; the shared container stays up for sibling sessions.
- **Editing the docker host config** (image, memory, network, ...) is detected on the next launch: the desired config hash is compared against the container's `codeman.confighash` label, and a mismatch refuses the launch with a "config changed, recreate?" confirm. Confirming calls `POST /api/docker-cases/:name/recreate` (refused while sessions of the case are live), which removes the container so the next launch recreates it with the new config; the workspace and the conversation survive.
- **Deleting the case** `docker rm -f`s the container (the bind-mounted workspace on the host survives). An instance-scoped boot reaper removes containers whose case is gone.
## Isolation & security
Every container runs hardened: `--cap-drop ALL`, `--security-opt no-new-privileges`, non-root (`--user <hostUid>:0` so workspace files stay host-owned), `--pids-limit`, `--memory` == `--memory-swap`, `--init`. Never `--privileged`, never the docker socket. The default **convenient** profile bind-mounts host credential dirs read-write so the common login just works (creds stay on the host, never captured by `docker commit`); the **sealed** profile (`mountCredentials:false` + `network:none`) is the opt-in for genuinely untrusted work.
Rootless engines without cgroup-v2 systemd delegation cannot enforce resource caps; linking such a host warns that caps are advisory.
## Export / Import (move to another machine)
**Export** (from the Docker tab, or `POST /api/docker-cases/:name/export`): choose
- **Full image + workspace**: `docker commit` the container to an image, `docker save` it, tar the workspace, and a manifest, all into one portable `<case>-<ts>.codeman-container.tgz` (the whole toolchain, installed packages, and files). Runs in the background; you are notified when the bundle is ready.
- **Workspace only**: just the project files (fast, small).
The container is paused across the capture so the image and workspace are consistent; a full `/var/lib/docker` is guarded against with a free-space precheck; the intermediate image is always cleaned up.
**Import** (`POST /api/docker-cases/import`, or the Manage tab): copy the `.tgz` onto the new machine's `~/.codeman/docker-exports/`, then import it into a new case. The manifest and per-member SHA-256 checksums are validated, the workspace tar is extracted with a path-traversal guard, and the image is `docker load`ed and **re-tagged into a quarantined namespace** (`codeman/imported-<case>:<ts>`) so it never overwrites a local tag. The destination supplies its own credentials, so nothing secret crosses machines.
`GET /api/docker-exports` lists bundles; `GET /api/docker-exports/:filename` downloads one; `DELETE` removes one.
## Hooks require the server to be reachable from the container
In-container hooks (permission events, hook-based idle/stop/task notifications) POST to `CODEMAN_API_URL`, which is derived as `https://host.docker.internal:<port>` (`host.docker.internal` → the docker bridge gateway, e.g. `172.17.0.1`, via `--add-host …:host-gateway`). For that callback to succeed, the Codeman server must be **listening on an interface the container can reach**.
- If Codeman binds **loopback-only** (`127.0.0.1`, the default and the production systemd config), a container reaching `172.17.0.1:<port>` cannot connect, so by default **in-container hooks do not fire**. The session still works fully: idle/stop detection falls back to **output-based** detection through the `docker exec` PTY (which always works), and claude runs with `--dangerously-skip-permissions` so there are no permission prompts to forward anyway.
- **To enable in-container hooks on a loopback-only server, set `CODEMAN_DOCKER_BRIDGE_HOOKS=1`** (env). Codeman then starts a SECOND listener bound to the docker bridge gateway (`172.17.0.1`, auto-detected; override with `CODEMAN_DOCKER_BRIDGE_HOST`) that serves **only the hook endpoints** (`/api/hook-event`, `/api/status-telemetry`) and delegates them into the same secret-gated pipeline. The bridge is host-internal (containers + host, not the LAN), and every other path returns `403`, so this does not widen your network exposure. Add `Environment=CODEMAN_DOCKER_BRIDGE_HOOKS=1` to the systemd unit and restart.
- Alternatively, bind `0.0.0.0` **with `CODEMAN_PASSWORD` set** (exposes on the LAN too).
The host-gateway mapping, `CODEMAN_API_URL` derivation, host-guard allowlist, and hook-secret mount are all wired correctly; `CODEMAN_DOCKER_BRIDGE_HOOKS` closes the last gap for loopback-only servers.
## Notes & limits
- Requires Docker (or Podman) with a reachable daemon; tmux must be present in the base image (a hard prerequisite, probed at link time).
- Per-session `envOverrides` / `effort` / per-CLI config are rejected for docker cases (they do not cross into the container); configure the container via the docker host's per-mode command override instead.
- macOS Docker Desktop takes a dedicated uid path (the baked image uid; memory caps are subject to the VM ceiling).
Design + rationale: [`docker-cases-plan.md`](./docker-cases-plan.md).
-417
View File
@@ -1,417 +0,0 @@
# Extending Codeman
Codeman has no plugin runtime, and that is a deliberate choice rather than a
missing feature. A plugin runtime means running third-party code inside a process
that spawns agents with your credentials, on a server people routinely expose
over a tunnel or Tailscale. Codeman's security model is one of its reasons to
exist, so it does not hand that away for an extension mechanism.
Instead there are four seams that already work, from any language, with nothing
installed:
| You want to | Use | Runs where |
| --- | --- | --- |
| Show your own UI inside Codeman | [Web tabs](#seam-1-web-tabs) | Your own process, rendered as a tab |
| React when an agent needs you | [SSE events](#seam-2-sse-events) | Anywhere that can hold an HTTP connection |
| Drive Codeman from a script | [HTTP API](#seam-3-http-api-and-cli) or the `codeman` CLI | Anywhere |
| React inside a Claude session | [Hooks](#seam-4-hooks) | The agent's own machine |
Everything below is covered by the stability promise in
[`versioning-policy.md`](versioning-policy.md): endpoint paths, the response
envelope, `errorCode` values, and SSE event names are stable. Additive changes
(new endpoints, new optional fields, new events) are non-breaking. Breaking
changes ship under a new prefix (`/api/v2`).
## Before you start
**Base URL.** `http://127.0.0.1:3000` by default. Prefer the versioned prefix
`/api/v1/...` for anything you publish; the unversioned `/api/...` is an alias.
**Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic on every request, or
authenticate once and keep the `codeman_session` cookie. With no password set,
Codeman is loopback-only and unauthenticated.
```bash
curl -u admin:$CODEMAN_PASSWORD http://127.0.0.1:3000/api/v1/sessions
```
**Envelope.** Every response is `{"success": true, "data": ...}` or
`{"success": false, "error": "...", "errorCode": "..."}`. Check the HTTP status
or `body.success`, then read `body.data`. The full `errorCode` to status mapping
is in [`api-reference.md`](api-reference.md).
⚠️ A few legacy GETs (`/api/away-digest` among them) return a bare-ish body with
the payload at the top level rather than under `data`. Read defensively with
`body.data ?? body`.
⚠️ A `401` is not an envelope at all: auth is rejected in a request hook that
replies with the bare string `Unauthorized`, so parsing it as JSON throws. Branch on
the status code before you parse, or a missing password looks like a broken endpoint.
**Already driving Codeman from an agent?** The README's
[Programmatic Guide](../README.md#driving-codeman-from-an-agent--programmatic-guide)
covers the in-session case: the `CODEMAN_MUX`, `CODEMAN_API_URL`,
`CODEMAN_SESSION_ID` and `CODEMAN_HOOK_SECRET_FILE` variables that let a CLI
running inside Codeman find the API and avoid acting on itself. This page is for
code running *outside* a session.
## Seam 1: Web tabs
The highest-leverage seam. Any web app you can serve locally becomes a tab beside
your agent sessions. You write a normal web page; Codeman handles embedding it.
```bash
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/webviews \
-H 'Content-Type: application/json' \
-d '{"name":"My Dashboard","url":"http://127.0.0.1:8787","icon":"📊"}'
```
Fields: `name` (1 to 60 chars), `url`, and optionally `icon` (a single glyph, max
8 code units), `embedMode` (`proxy` by default, or `direct`), and `trusted`.
Related endpoints: `GET /api/v1/webviews`, `PATCH /api/v1/webviews/:id`,
`DELETE /api/v1/webviews/:id`, `POST /api/v1/webviews/probe` (reachability and
framing check), `POST /api/v1/webviews/:id/open`.
### Why it is proxied
By default your page is served through Codeman's own origin at `/webview/:cap/*`
rather than framed directly. A direct iframe fails three ways at once: production
is HTTPS so `http://` targets are blocked as mixed content, many dashboards send
`X-Frame-Options: DENY`, and Codeman's own `default-src 'self'` CSP blocks
cross-origin frames. Proxying solves all three without weakening the CSP.
### The two things that will confuse you
A proxied frame is sandboxed and therefore **opaque-origin** unless you set
`trusted: true`. Two consequences look like bugs in your own app:
1. **Root-absolute URLs built at runtime** (`/assets/x.png` assembled in JS)
escape the injected `<base>` tag. Codeman injects a `runtimeUrlShim()` that
patches the common DOM sinks, but if you construct URLs in an unusual way,
prefer relative paths.
2. **Same-host `fetch` and `XHR` are CORS-checked with `Origin: null`.** Codeman
handles this with `buildProxyCorsHeaders()`, and the proxy is exempt from the
global `OPTIONS` short-circuit. If you see "Failed to fetch" while the page
itself renders fine, this is the area to look at.
⚠️ `trusted: true` opts out of the sandbox. A proxied page is served from
Codeman's origin, so `allow-same-origin` lets it read the Codeman page and call
the API that spawns agents. Only mark your own trusted code.
## Seam 2: SSE events
`GET /api/v1/events` is a Server-Sent Events stream. Each message is
`event: <name>` plus `data: <json>`. There are 149 event names following a
`domain:action` convention, registered in `src/web/sse-events.ts`.
The ones most integrations want:
| Event | Meaning |
| --- | --- |
| `session:created`, `session:deleted` | A session appeared or went away |
| `session:idle` | The agent stopped working |
| `session:completion` | A completion message was detected |
| `session:exit`, `session:error` | The session ended or failed |
| `hook:permission_prompt` | The agent is asking for permission |
| `hook:idle_prompt`, `hook:stop` | The agent is waiting on you, or stopped |
| `hook:task_completed`, `task:completed` | Work finished |
| `subagent:discovered`, `subagent:completed` | Background agent lifecycle |
| `mux:died` | A multiplexer session died unexpectedly |
| `cron:runCreated`, `cron:runUpdated` | Scheduled job activity |
### Filtering
`?sessions=id1,id2` suppresses only the high-volume `session:terminal` stream for
sessions you did not list. Lifecycle and metadata events are always delivered, so
you cannot accidentally filter away the thing you are listening for.
Pass `?clientId=<uuid>` to enable live filter updates through
`POST /api/v1/events/subscribe` without reconnecting the stream.
### Example: notify when any agent needs you
```js
const res = await fetch('http://127.0.0.1:3000/api/v1/events', {
headers: { Authorization: 'Basic ' + btoa(`admin:${process.env.CODEMAN_PASSWORD}`) },
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buf = '';
const WANTED = new Set(['hook:permission_prompt', 'hook:idle_prompt', 'session:idle']);
for (;;) {
const { value, done } = await reader.read();
if (done) break;
buf += decoder.decode(value, { stream: true });
const frames = buf.split('\n\n');
buf = frames.pop() ?? '';
for (const frame of frames) {
const name = frame.match(/^event: (.+)$/m)?.[1];
const data = frame.match(/^data: (.+)$/m)?.[1];
if (name && WANTED.has(name)) notify(name, JSON.parse(data ?? '{}'));
}
}
```
## Seam 3: HTTP API and CLI
Around 200 handlers across 21 route files cover sessions, cases, files, cron,
respawn, Ralph, the orchestrator, search, and admin. Each route module carries an
`@fileoverview` describing its endpoints.
If the caller is an agent running _inside_ a Codeman session, install the packaged
agent skill instead of teaching it these calls by hand: `skills/codeman` in the repo
(`npx skills add Ark0N/Codeman --skill codeman -g`, or `codeman skill install
[--case <name>]`, or the synced `agentSkillEnabled` App Setting for automatic
per-case injection on Claude session create). The skill carries the guard, the
safety rules, and verified wait/orchestration recipes.
The common ones:
```bash
# List sessions (live + persisted + transcript history, deduped)
curl -u admin:$PASS http://127.0.0.1:3000/api/v1/sessions/unified
# Create a session
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions \
-H 'Content-Type: application/json' \
-d '{"workingDir":"/home/me/project","mode":"claude"}'
# Send a prompt (single-line only, and it must end with \r: Enter is sent only
# when the input contains a carriage return; without it the text sits on the
# session's prompt unsubmitted)
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions/$ID/input \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true}'
```
`POST .../input` also accepts `clientId` (stable per client, max 128 chars) and
`seq` (monotonic per session). Send both and the server applies each pair
at-most-once, so retrying after a dropped connection cannot type the prompt
twice. Omit them entirely rather than sending `null`.
It also accepts `wait` and `waitTimeout`, which hold the response open until the
session finishes the turn you just started. `wait` is `true` (the default signal
set) or a comma list of `idle,working,stop,blocked,exit`; the result comes back
under `data.wait`. Sending them changes nothing for callers that do not: without
`wait` the response is still `{"success": true, "data": {}}` and the write is still
fire-and-forget. The two interact with `clientId` / `seq` in one way worth knowing:
a **tagged duplicate** (a pair the server already applied) skips the write but still
waits, answering from the session's current state rather than blocking for a
transition that already happened. It reports `"delivered": false, "duplicate": true`.
### Waiting instead of polling
Three calls block until something happens: `GET /api/v1/sessions/:id/wait` (a
lifecycle signal), `GET /api/v1/sessions/:id/wait-output` (a literal string in the
output), and the `wait` field above. Full parameter and response tables are in
[`api-reference.md`](api-reference.md#long-polling-agent-wait). Four things decide
whether your integration works, and the last one is what actually bites:
- **A timeout is a `200` with `wait.timedOut: true`**, not an error. Loop over short
waits rather than issuing one long one, because `tailscale serve` and cloudflared
both cut idle connections and a single 10-minute call is the pattern most likely
to die in the field.
- **`wait.timeoutMs`** is the timeout after server-side clamping (600 s ceiling by
default). Read it rather than assuming you got what you asked for.
- **`stop` and `blocked` only exist for `claude` sessions**, and on a `shell` session
even `idle` fires only once at startup, so send-and-wait there can only time out.
See the Gotchas below.
⚠️ **There is no readiness signal, and skipping readiness is the failure that looks
like success.** A session reports `idle` before its CLI has spawned, and a `claude`
worker in a brand-new case comes up on the CLI's **trust dialog**, which has a ❯
prompt of its own. Prompt it at that moment and the text lands in the dialog, the
`\r` does not get past it, and the session's startup `idle` lands inside the wait
window: the wait resolves on `idle` in a couple of seconds with `timedOut: false`,
indistinguishable from a finished turn. Wait for the pid, then wait for the
composer, answering the dialog only as the bounded fallback.
A worked orchestration: start a worker, get it ready, prompt it, wait, clean up.
```bash
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}" # auto-set in-session, correct scheme included
AUTH=(-u "admin:$CODEMAN_PASSWORD") # omit entirely if no password is set
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on --https installs (self-signed cert)
# 1. Start a worker session (creates the case if it does not exist yet).
# The guard matters: a TLS or auth failure otherwise leaves SID empty and every
# later step "succeeds" against nothing.
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
[ -n "$SID" ] && [ "$SID" != null ] || { echo "quick-start failed"; exit 1; }
# 2. READINESS: composer marker first, trust dialog only as the bounded fallback.
# Skip this and step 3 reports a turn that never ran. Do NOT probe trust first
# and Enter blindly: the dialog text stays in the buffer for the life of the
# session, so on every later run that probe matches stale text and the Enter
# lands in a ready composer. Match single tokens only: TUI text can arrive
# without its spaces. Stage 1 is short on purpose (an already-trusted case
# matches in <1 s; a first-run case can never pass it and pays it in full).
until [ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ]
do sleep 1; done
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=5000') # composer's status bar = ready
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' -d '{"input":"\r","useMux":true}' >/dev/null
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send the prompt AND register the wait in one call, so the answer cannot be
# the previous turn's idle state. Single line only, ending in \r (otherwise
# Enter is never sent and this wait times out on a turn that never started).
W=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize the failures\r","useMux":true,
"clientId":"orchestrator","seq":1,"wait":"stop,exit","waitTimeout":60000}' \
| jq -c '.data.wait')
# 4. That first wait probably timed out (60 s). Keep going in SHORT waits.
for _ in $(seq 1 30); do
[ "$(jq -r '.timedOut' <<<"$W")" = 'true' ] || break # signal fired, or wait ended
W=$("${CURL[@]}" \
"$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq -c '.data.wait')
done
jq -r 'if .ended or .aborted then "worker is not running"
elif .timedOut then "still working after 30 waits"
else "signal: \(.signal)" end' <<<"$W"
# 5. Read what it produced, then delete the session YOU created, by exact id.
# ⚠️ NOT /output: its textOutput is empty for every tmux-backed session.
# `tail` counts BYTES, and the payload is terminal data with ANSI in it.
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
```
Waiting on a marker instead of a signal is the form that works in **every** mode,
and the only one that works on a `shell` session:
```bash
# ⚠️ Split the marker so the typed line never contains it: your own keystrokes echo
# into the output stream, so an unsplit marker matches before the command has run.
# `from=buffer` also catches a marker that printed before the wait registered.
N=$RANDOM
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
```
For shell scripting, the `codeman` CLI is the same surface without the HTTP
plumbing:
```
codeman session start|stop|list|logs codeman task add|list|status|remove|clear
codeman ralph start|stop|status|reset codeman users add|passwd|list
codeman status | list | attach <path> codeman doctor
```
## Seam 4: Hooks
Claude Code hooks post to `POST /api/v1/hook-event` from inside an agent session.
Codeman installs its own hooks automatically, but the endpoint is open to yours.
```json
{ "event": "task_completed", "sessionId": "abc123", "data": { "any": "json" } }
```
`event` must be one of `permission_prompt`, `elicitation_dialog`, `idle_prompt`,
`stop`, `teammate_idle`, `task_completed`. Each becomes the matching `hook:*` SSE
event.
⚠️ This endpoint skips Basic auth so hooks keep working, but when auth is active
the loopback bypass requires the `X-Codeman-Hook-Secret` header
(`~/.codeman/hook-secret`) unconditionally.
## Gotchas
Every one of these has cost somebody real time.
- **CORS is localhost-only.** `Access-Control-Allow-Origin` is echoed only for
`localhost`, `127.0.0.1`, and `::1`. A browser app on any other origin cannot
call the API. Integrate server-side.
- **A missing `Origin` header is allowed**, which is why curl, CLIs, and hooks
work. Cross-site origins are blocked by the CSRF guard.
- **Reverse-proxy domains are rejected** by the anti-DNS-rebinding Host allowlist
unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix`.
- **`null` is not `undefined`.** Request schemas use Zod `.optional()`, which
accepts `undefined` only. `JSON.stringify({ field: null })` keeps the null on
the wire and fails with `INVALID_INPUT`. Omit the key instead. This has caused
shipped bugs more than once.
- **`text/plain` bodies stay raw.** Auto-parsing them as JSON enabled
simple-request CSRF, so it is deliberate. Send `application/json`.
- **Prompts are single-line and must end with `\r`.** The server splits your text
and Enter into two separate tmux writes (Ink needs them apart), but it sends the
Enter **only when the input contains a carriage return**. Without it your text
sits on the prompt unsubmitted, which is the single most common "the wait
endpoints don't work" report: the wait runs its full timeout on a turn that never
started. Newlines inside the string are stripped rather than rejected, so
`"echo A\necho B\r"` runs the single joined command `echo Aecho B`: send one line
per call.
- **`wait-output`'s `from=now` is not "printed after you asked".** tmux repaints
the visible screen on attach, on resize, and on any TUI redraw, and a repaint
arrives as ordinary output, so text already on screen can satisfy a fresh wait.
Observed live: a marker echoed a minute earlier matched instantly. Use a marker
unique to each call, and build it so the typed line never contains it (your own
keystrokes echo into the stream). Matching is a literal substring, so `regex=` is
rejected with a `400` rather than ignored.
- **`wait-output` matches the normalized PTY stream, not the screen.** ANSI escape
sequences are stripped (the `ESC ( B` charset escape a bash prompt emits on every
line included), a partial escape at a chunk boundary is held back until its tail
arrives, and a match may straddle PTY chunks, so text you printed yourself
matches reliably (`printf STRAD; sleep 1; printf DLEQQ` is matchable as
`STRADDLEQQ`). What can still fail is TUI output: a full-screen TUI positions
words with cursor moves, so its text can reach the matcher **without spaces** and
a multi-word match is unreliable there. Match one short space-free token, ideally
one you printed yourself, and keep it out of the typed line (your own keystrokes
echo into the stream).
- **`stop` and `blocked` never fire for `shell`, `opencode`, `codex`, `gemini`,
`antigravity` or `pi` sessions.** They come from Claude Code hooks, which no other mode
installs, so only `idle`, `working` and `exit` exist there. Asking for them
explicitly is a `400`; omitting `until` is safe, since the server drops them from
the default set and echoes what it actually waited on as `wait.until`. Even in
`claude` mode, a Docker case needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for hooks to
reach the server at all, a remote-SSH case's hooks may never arrive, and a case
written by Codeman < 1.13.0 against an `--https` install carries hook curls
without `-k` that TLS-fail silently — a 1.13.0+ server rewrites them the next
time a session starts in that case.
- **Unwrap the envelope** before reading fields. `data` is not the response body.
## Publishing your integration
There is no registry and no review queue. Add the GitHub topic
**`codeman-integration`** to your public repository so others can find it, and
link back to Codeman in your README.
If a real ecosystem of these appears, a manifest format and an install command
become worth building. Until then, these four seams are the contract, and they
require nothing of you but HTTP.
## What Codeman deliberately does not have
- **No in-process plugin runtime.** See the reasoning at the top of this page.
- **No build or startup hooks** for third-party code. Run your own process.
- **No per-plugin config or state directories.** Manage your own files.
- **No sandbox for integration code**, because Codeman never launches it. Your
integration is your own process, started by you, with your permissions,
talking HTTP.
That last point is about integration code specifically, not about Codeman.
Sandboxing lives on a different axis here: the thing worth isolating is the
**agent**, and you isolate it per case with
[Docker cases](docker-cases.md), which run the agent in a hardened container with
a bind-mounted workspace and seeded (not shared) credentials. An integration that
creates or drives a Docker-backed session inherits that isolation for free, since
it is a property of the session rather than of the caller.
-432
View File
@@ -1,432 +0,0 @@
# File Viewer edit mode (issue #212)
Plan only. No implementation yet.
Goal: close the loop "agent writes a file, you review it in the viewer, tweak two lines, save, tell the
agent to continue" without hopping into the terminal, with the phone as the primary target.
Scope from the issue: an Edit toggle on text previews, a write endpoint that inherits the read path's
confinement, text-only, edit-in-place (no create, no delete, no rename), no editing through the
Docker/remote overlays.
---
## 1. What exists today
**Read path (backend), all in `src/web/routes/file-routes.ts`:**
| Route | Line | Notes |
| ------------------------------------ | ------ | ------------------------------------------------------------------ |
| `GET /api/sessions/:id/files` | `741` | Tree scan of `session.workingDir`, hidden files off by default |
| `GET /api/sessions/:id/file-content` | `865` | The text/preview classifier. `findSessionOrFail` + `validateSessionFilePath` |
| `GET /api/sessions/:id/file-raw` | `1018` | Bytes, 50MB cap |
| `GET /api/sessions/:id/file-preview` | `1254` | DOCX/PPTX to PDF, everything else redirects to `file-raw` |
| `GET /api/download` | `1384` | The only read route that also runs `isSensitivePath()` |
`file-content` classification order (`file-routes.ts:881-1011`): extension buckets (image / video / audio /
known-binary) return metadata only; otherwise the bytes are read, sniffed for a NUL in the first 8KB, and
either reported as `type:'binary'` or decoded as UTF-8 and **truncated to `lines` (default 500, hard cap
10000)**. Caps: `MAX_TEXT_FILE_SIZE` 10MB.
Confinement is `validateSessionFilePath()` (`src/web/route-helpers.ts:67`): `resolve()` then `realpathSync()`
then reject if the result is not under `workingDir`. Because it realpaths the *full* path, a symlink whose
target escapes the workspace is already rejected. Ownership is `findSessionOrFail()` which runs
`canAccessOwned()` (`route-helpers.ts:102`), a no-op outside multi-user mode.
**Read path (frontend), `src/web/public/panels-ui.js`:**
- `loadFileBrowser()` `2947`, `renderFileBrowserTree()` `2978`, click to `openFilePreview()` `3056`.
- `openFilePreview(filePath, sessionId, attachmentId)` `3193`: attachment-id branch, then docx/pptx, pdf,
svg branches, then the generic `file-content` fetch at `3274` with **`&lines=500` hardcoded**, rendering
text as `<pre><code>${escapeHtml(...)}</code></pre>` at `3298` and stashing `this.filePreviewContent`.
- `closeFilePreview()` `3308`, `copyFilePreviewContent()` `3751`.
- Markup: `src/web/public/index.html:420-432` (`filePreviewOverlay` / `-Title` / `-Body` / `-Footer`, two
header buttons: copy and close).
- CSS: `src/web/public/styles.css:9320-9430`. Overlay `z-index: 2000`, window `80vw/80vh`, capped
`900x700`. There are **no `.file-preview-*` rules in `mobile.css` at all**.
**Reachability on phones.** The header File Viewer button is hidden below 430px
(`mobile.css:482`, locked by `KNOWN_PHONE_HIDDEN` in `test/mobile-header-buttons-policy.test.ts`), so on a
phone the preview overlay is reached through:
1. an attachment card's **Preview** button (`panels-ui.js:3451`), which is exactly the "agent just wrote a
file" path the issue describes,
2. the attachment-history drawer (`panels-ui.js:3709`),
3. App Settings to Panels to **File Browser** (`showFileBrowser`, applied in `settings-ui.js:2202`; the
panel is mobile-styled at `mobile.css:1868`).
So edit mode is reachable on a phone today via (1) and (2) without touching the header policy. Improving
the entry point is listed as an open decision in section 10, not assumed.
---
## 2. Threat model, stated honestly
Anyone who can call this API can already reach `POST /api/sessions/:id/input` and type an arbitrary prompt
into an agent running with `--dangerously-skip-permissions`. A workspace-confined write endpoint therefore
does not create a new privilege tier for an authenticated caller.
What it *would* create if built carelessly is a **new host-write primitive reachable by path**, so the
things this plan actually defends against are:
1. **Path traversal / symlink escape** writing outside the workspace.
2. **TOCTOU**: a path component that becomes a symlink between validation and write.
3. **Cross-user writes** in multi-user mode (`canAccessOwned`).
4. **Silent data loss**, which is the highest-probability real-world failure here and gets its own section.
CSRF is already covered: `registerHostGuard()` (`src/web/middleware/auth.ts:555-578`) rejects any
non-safe-method request whose `Origin` is cross-site. The webview-capability exemption at that gate is
fenced to `GET`/`HEAD` for the Referer form (`auth.ts:161`) and to `/webview/:cap/*` paths for the path
form, so a proxied dashboard cannot reach a new `PUT /api/...`. Using `PUT` + `application/json` also
forces a preflight for any cross-origin attempt.
---
## 3. Backend design
### 3.1 New policy module: `src/config/file-editing.ts`
Pure, unit-testable, no IO (config lives in `src/config/`, no barrel, import the file directly).
```ts
export const MAX_EDITABLE_BYTES = 512 * 1024; // content cap, both directions
export const EDITABLE_EXTENSIONS: ReadonlySet<string>; // ts,tsx,js,jsx,mjs,cjs,json,jsonc,md,mdx,txt,
// css,scss,less,html,htm,xml,svg?,yml,yaml,toml,
// ini,cfg,conf,env?,sh,bash,zsh,fish,py,rb,go,rs,
// java,kt,swift,c,h,cpp,hpp,cs,php,sql,graphql,
// proto,lua,pl,r,jl,tf,gradle,csv,tsv,log,diff,patch
export const EDITABLE_BASENAMES: ReadonlySet<string>; // Dockerfile, Makefile, LICENSE, .gitignore,
// .prettierignore, .editorconfig, .nvmrc, ...
export function isEditableFileName(fileName: string): boolean;
export function isDeniedEditRelativePath(rel: string): boolean; // `.git/` subtree
export function detectEol(text: string): 'lf' | 'crlf';
export function applyEol(text: string, eol: 'lf' | 'crlf'): string;
```
Decisions baked in:
- **Allowlist, not blocklist**, per the issue and per the existing attachment-guard precedent.
- `svg` and `env` are deliberately marked with `?` above: `svg` is served as an untrusted octet-stream on
the read side (`file-routes.ts:118`) so allowing an edit is defensible, but I recommend **excluding
both** in v1. `.env` files are matched by `isSensitivePath()` anyway and would be rejected downstream;
excluding them at the allowlist keeps a single obvious refusal.
- `isDeniedEditRelativePath` blocks the `.git/` subtree: `.git/hooks/*` is code execution and a corrupt
index is unrecoverable-looking to a user who only wanted to fix a typo. Other dotfiles stay allowed but
are not reachable from the tree UI anyway (`showHidden=false`).
### 3.2 Read-for-edit: extend the existing GET
`GET /api/sessions/:id/file-content?path=<rel>&edit=1`
When `edit=1`:
- skip line truncation entirely (a truncated buffer must never become an edit buffer, see section 4.1),
- enforce `MAX_EDITABLE_BYTES` instead of `MAX_TEXT_FILE_SIZE` and answer 413 over it (as a structured
throw with `statusCode: 413`, the `throwFilesystemPickerError` pattern, since the central errorCode-to-
status map has no 413 entry; see the error-mechanics note in 3.3),
- run the editability gate (`isEditableFileName`, `isDeniedEditRelativePath`, `isSensitivePath`,
`isBlockedAttachmentPath`) and the content gate (NUL sniff plus UTF-8 round-trip, see 4.3),
- return `{ content, size, mtimeMs, totalLines, truncated: false, extension, editable: true, hash, eol }`.
`hash` is `sha256` hex of the exact on-disk bytes.
Non-`edit` responses gain **only** `editable: boolean` (additive, no shape change for existing consumers),
which is all the UI needs to decide whether to show the Edit button. No `hash` on plain reads: the Edit
action re-fetches with `edit=1` anyway (section 4.1), which is where the hash comes from, and hashing every
casual 10MB preview would be pure waste.
### 3.3 Write: `PUT /api/sessions/:id/file-content`
Body (new `FileWriteSchema` in `src/web/schemas.ts`, Zod v4):
```ts
{ path: string, content: string, baseHash: string, eol?: 'lf'|'crlf', force?: boolean }
```
Registered with an explicit route option `{ bodyLimit: 4 * 1024 * 1024 }`. **Fastify's default `bodyLimit`
is 1MB and this repo configures none**, and JSON escaping expands content: 2x for a file full of quotes or
backslashes, up to 6x for control characters (each serialized as a `\uXXXX` escape), so 512KB of content
can legitimately exceed 1MB on the wire; blowing the limit produces a raw `FST_ERR_CTP_BODY_TOO_LARGE`, not an `ApiResponse` envelope. Two
related sizing notes: `z.string().max()` counts **UTF-16 code units, not bytes**, so the schema's `.max()`
is only a coarse pre-filter and the real cap is an explicit `Buffer.byteLength(content, 'utf8')` check in
the handler (step 7a below); and 4MB comfortably bounds the worst-case expansion of a 512KB file without
inviting multi-MB bodies elsewhere.
**Error mechanics** (matters for both prod behavior and testability): a handler that *returns* a
`{success:false, errorCode}` envelope gets its HTTP status assigned centrally by the preSerialization hook
in `server.ts` (`httpStatusForErrorCode()`, `src/types/api.ts`), but the route-test harness
(`test/routes/_route-test-utils.ts`) installs only `installRouteErrorHandler`, **not** that hook, so
returned envelopes surface as HTTP 200 in tests. The PUT handler should therefore use the same
structured-**throw** pattern as the filesystem picker (`throwFilesystemPickerError`, `file-routes.ts:411`):
thrown `{statusCode, body}` errors are rendered identically in prod and in the harness, and they allow the
one status the code map cannot express (413). The error envelope itself is strictly
`{success:false, error, errorCode}`, **it has no data arm**, so no error response may carry extra payload.
Handler order (each step is a test case):
1. `findSessionOrFail(ctx, id, req)` (live sessions only, matching the read route, and it carries the
multi-user ownership check).
2. `parseBody(FileWriteSchema, req.body)`, then `Buffer.byteLength(content, 'utf8') <= MAX_EDITABLE_BYTES`
or 413 (the schema `.max()` alone cannot enforce a byte cap, see the sizing note above).
3. `validateSessionFilePath(session.workingDir, path)` or 404 (do not distinguish "outside workspace" from
"missing", matching the read route).
4. `isSensitivePath(resolvedPath) || isBlockedAttachmentPath(resolvedPath, guard.blockedTrees)` or 403.
5. `isDeniedEditRelativePath(relativePath)` or 403.
6. `isEditableFileName(basename(resolvedPath))` or 400.
7. `stat`: must be `isFile()`, size within `MAX_EDITABLE_BYTES`, else 400/413. **No `O_CREAT` anywhere in
this handler**, which is what enforces edit-in-place.
8. Read current bytes, compute `hash`, run the NUL sniff and the UTF-8 round-trip check, else 400.
9. `hash !== baseHash && !force` gives **409 CONFLICT** (`ApiErrorCode.CONFLICT`, plain envelope; the error
arm carries no data, see the error-mechanics note). The client's conflict dialog gets fresh state by
re-fetching `edit=1`, which it needs for its Reload action anyway.
10. Build the output buffer: `applyEol(content, eol ?? detected-from-original)`; re-check
`Buffer.byteLength` against the cap.
11. Write atomically in the resolved parent directory:
`fs.open(<dir>/.<name>.codeman-tmp-<rand>, 'wx', stat.mode & 0o777)`, then `fchmod(stat.mode & 0o777)`
(open's mode argument is masked by the process umask, so the chmod is what actually preserves an
unusual mode), write, `fsync`, close, `fs.rename(tmp, resolvedPath)`, unlink the temp on any failure.
12. Re-stat, return `{ success: true, data: { path, size, mtimeMs, hash, totalLines } }`.
Why `O_EXCL` temp plus rename rather than truncate-in-place:
- `wx` cannot follow a pre-existing symlink, which closes the TOCTOU window from step 3 to step 11 without
needing `O_NOFOLLOW` gymnastics.
- `rename()` does not follow a symlink in the final component, so even if `resolvedPath` were swapped for a
symlink after validation, the symlink itself is replaced and the swap target is untouched.
- A crash mid-write leaves the original intact.
Caveat to document in the code comment: rename replaces the inode, so hardlinks to the file keep the old
content. That is the same trade-off vim makes by default and is preferable to a truncate window here.
No SSE event in v1. Nothing else in the app needs to know: `image-watcher.ts` only reacts to
`.png/.jpg/.jpeg/.gif/.webp/.bmp/.svg/.pdf/.docx/.pptx` adds (`image-watcher.ts:23-25`), none of which are
editable text, and the temp filename does not match either.
---
## 4. The five traps
These are the parts that turn a "small write endpoint" into a bug report.
### 4.1 Truncation (the data-loss trap)
The frontend fetches `&lines=500` (`panels-ui.js:3274`). Saving that buffer back would **delete every line
past 500**. Worse, the content hash of the full file would still match, so an optimistic-concurrency check
cannot catch it.
Mitigations, all three:
- The Edit affordance is only offered when the loaded payload came from `edit=1` (which never truncates).
Tapping Edit on an already-rendered preview **re-fetches** with `edit=1` before swapping in the editor.
- The read-for-edit path 413s above `MAX_EDITABLE_BYTES` rather than truncating, so "too big to edit here"
is an explicit refusal with a message, never a silent partial buffer.
- A test asserts `edit=1` never returns `truncated: true`.
### 4.2 Line endings
A `<textarea>`'s `.value` normalizes to LF. Saving a CRLF file naively rewrites every line, producing a
whole-file diff for a two-line change. So: the read returns the detected `eol`, the client echoes it back
unchanged, and the server re-applies it. Mixed-EOL files use the dominant style, which is lossy for the
minority lines; call that out in the response and accept it in v1.
### 4.3 Encoding
`buf.toString('utf-8')` on a latin-1 or otherwise non-UTF-8 file yields U+FFFD replacement characters, and
writing that back **corrupts the file**. The check is a round-trip:
`Buffer.from(decoded, 'utf8').equals(buf)`. If it fails, `editable: false` and the write is refused. This
also catches binary content that the NUL sniff misses. A UTF-8 BOM survives because it round-trips as a
leading U+FEFF; do not strip it.
### 4.4 Concurrency with the agent
The whole use case is editing a file the agent just wrote and may write again. `baseHash` plus 409 is the
guard. Do not use mtime alone: agents rewrite files within a single filesystem timestamp tick, and an
identical rewrite should not be reported as a conflict.
### 4.5 Symlinks and TOCTOU
Covered by `validateSessionFilePath` (escape) plus `wx` temp and `rename` (post-validation swap). One
intentional allowance: a symlink whose target is *inside* the workspace is edited through to its target,
because `validateSessionFilePath` returns the realpath. That matches what a user tapping the file expects.
---
## 5. Frontend design
All in `panels-ui.js` (prettier-exempt, hand-formatted; match the surrounding style), `index.html`,
`styles.css`, `mobile.css`.
### 5.1 State
```js
filePreviewEdit = { active, sessionId, path, baseHash, eol, original, dirty }
```
Reset in `closeFilePreview()` and on every `openFilePreview()` entry.
### 5.2 Markup (`index.html:420-432`)
Add one header button (pencil, `btn-icon-sm`, `id="filePreviewEditBtn"`, hidden by default) next to the
copy button, and an edit bar inside the footer region holding Save / Cancel / a dirty dot. Keep the
existing footer text element; the edit bar is a sibling toggled by class so the read-mode footer is
untouched.
### 5.3 Behavior
- `openFilePreview()` shows the Edit button only when the response has `editable: true` and the render took
the text branch. Attachment-id previews, media, binary, pdf, docx/pptx and svg all leave it hidden.
- **Enter edit**: re-fetch with `edit=1`; on 413 or `editable:false`, toast the reason and stay in read
mode. This fetch must **parse the error envelope on non-ok responses**: the existing generic
`if (!res.ok) throw new Error('Failed to load file')` pattern (`panels-ui.js:3275`) would swallow the
specific "too large to edit here" message, since error envelopes arrive with real 4xx statuses in prod. On success replace the body with `<textarea class="file-preview-editor" spellcheck="false"
autocapitalize="off" autocorrect="off" autocomplete="off" wrap="off">` and assign `.value = content`
(never `innerHTML`, so no escaping question arises). Do **not** autofocus: on a phone that opens the
keyboard before the user has picked a line.
- `input` sets `dirty` and enables Save.
- **Save**: `PUT` with `baseHash`, `eol`, and `content`. On success update `baseHash`/`original` from the
response, leave edit mode, re-render the read view from the local editor value (the response carries
metadata only, not content), toast "Saved". On **409** offer `Reload (discard mine)` / `Overwrite`:
Reload re-fetches `edit=1` and replaces the buffer; Overwrite re-sends with `force: true`. The 409 body
itself carries no state (section 3.3, step 9).
- **Cancel / close / Escape while dirty**: `confirm('Discard unsaved changes?')`, consistent with the
existing `window.confirm` usage in this codebase (`panels-ui.js:4323`, `app.js:4176`). Note the global
Escape handler (`app.js:999-1007`) closes other panels via `closeAllPanels()` but does not touch this
overlay today; if Escape-to-close is wired up as part of this work it must go through the same dirty
guard.
- `copyFilePreviewContent()` copies the live editor value while editing.
⚠️ Repo gotcha to respect at the fetch call: **Zod `.optional()` rejects `null`**. Build the body with
`eol: eol ?? undefined` (or declare `.nullish()`), or the PUT fails `INVALID_INPUT`. This has shipped as a
real bug twice.
### 5.4 Mobile
- **Sizing.** The window is `80vw/80vh` centered with no mobile override, so when the keyboard opens on iOS
the lower half sits behind it. Add a `@media (max-width: 430px)` block using
`height: var(--app-height, 100vh)`, full width, no border radius. `--app-height` is already maintained
against `visualViewport` by `KeyboardHandler.handleViewportResize()` (`mobile-handlers.js:283-317`), so
the editor tracks the keyboard for free.
- **iOS zoom.** The editor font must be >= 16px on phones; there is an existing zoom-prevention block at
`mobile.css` under `@media (max-width: 768px)`. Verify it covers `textarea` and do not override it with a
smaller `rem` value.
- **Accessory bar.** Focusing any input fires `KeyboardHandler.onKeyboardShow()`, which calls
`KeyboardAccessoryBar.show()` and refits/resizes the terminal (`mobile-handlers.js:407+`). The bar's keys
target the **terminal**, not the editor, so an Esc or clear-input tap while editing goes to the agent.
The overlay's `z-index: 2000` covers the bar's `51`, so it is not visible, but confirm it is not
interactive underneath and consider an explicit `KeyboardAccessoryBar.hide()` while the editor holds
focus. This is the item most likely to look "fine on desktop, wrong on the phone".
- No header-policy change is needed (section 1), so
`test/mobile-header-buttons-policy.test.ts` stays untouched.
### 5.5 i18n
`i18n.js` already skips `textarea`, `pre`, `code` and `.file-preview-content` in its `SKIP_SELECTOR`
(`i18n.js:20-38`), so file content is never translated. Add zh-CN entries for the new chrome: Edit, Save,
Cancel, Unsaved changes, Discard unsaved changes?, File changed on disk, Reload, Overwrite, Saved,
Too large to edit here.
---
## 6. Docker and remote cases
Out of scope per the issue, and the current behavior already degrades correctly:
- **Docker cases**: the workspace is a host directory bind-mounted at the same absolute path, so a host-side
write is visible in the container immediately. Edit mode works and needs nothing special. Worth one line
in the docs.
- **Remote SSH cases**: `workingDir` is a path on the remote host. `validateSessionFilePath` realpaths it
locally, which fails, so the write returns 404 exactly like the read routes do today. Confirm the viewer
shows a clean empty/error state rather than an unexplained failure, and do not attempt an SFTP path.
---
## 7. Tests
| File | Kind | Covers |
| ------------------------------------------- | ----------- | ---------------------------------------------------------------------- |
| `test/file-editing-policy.test.ts` | pure unit | `isEditableFileName` (allow + deny + basenames), `isDeniedEditRelativePath`, `detectEol`/`applyEol` round-trip incl. mixed EOL, BOM preservation |
| `test/routes/file-write-routes.test.ts` | `app.inject` | The handler order in 3.3, against a **real temp dir** (do not `vi.mock('node:fs')` in this file; set `MockSession.workingDir`, `test/mocks/mock-session.ts:14`) |
| extend `test/routes/file-routes.test.ts` | `app.inject` | `edit=1` never truncates; `editable` present on the plain read |
Status-code caveat for all of these: the route-test harness does not install the server's preSerialization
envelope hook, so a handler that *returns* an error envelope answers 200 in tests. The statuses below are
only assertable because the plan has the handler **throw** structured errors (section 3.3, error
mechanics), which `installRouteErrorHandler` renders identically in prod and in the harness.
Route cases to assert explicitly:
1. happy path writes the bytes and returns a new hash
2. `../` and absolute paths give 404
3. symlink pointing outside the workspace gives 404
4. symlink pointing inside is written through to the target
5. non-allowlisted extension gives 400
6. `.git/config` gives 403
7. a `.env` in the workspace gives 403 (sensitive-path)
8. a file with a NUL byte gives 400
9. a latin-1 file that fails the UTF-8 round-trip gives 400
10. stale `baseHash` gives 409 (`CONFLICT` envelope, no data); `force:true` then succeeds
11. over `MAX_EDITABLE_BYTES` gives 413
12. a path that does not exist gives 404 and creates nothing (no `O_CREAT`)
13. multi-user: `authUser: {role:'user'}` against another user's session gives 404 (pass `authUser` to
`createRouteTestHarness`, otherwise the synthetic admin makes the test pass vacuously)
14. CRLF file edited and saved stays CRLF
15. file mode is preserved across the temp-plus-rename
Run with `npm test -- test/routes/file-write-routes.test.ts`, never bare `npm test`.
**End-to-end verification before any deploy** (unit tests passing is not sufficient here):
- `curl -sk https://localhost:3000/...` against a **throwaway** session created for the purpose, never
`w1`/`w2`/`w3`; delete it by exact id afterwards.
- Playwright on a phone profile: open a preview, tap Edit, type with `page.keyboard.type()`, Save, then
assert the bytes on disk changed. Assert real state, not HTTP 200.
---
## 8. Docs and release
- This plan lives at `docs/file-viewer-edit-plan.md`.
- `docs/architecture-invariants.md`: new anchor `#file-viewer-edit-mode` covering the write confinement
chain, the truncation invariant, and why temp-plus-rename.
- `CLAUDE.md`: one line under the **Filesystem path picker** neighborhood noting that the File Viewer now
has a **third** file surface and that it is the only one that writes, plus its confinement rules.
Remember `CLAUDE.md` is prettier-ignored on purpose.
- `docs/api-reference.md`: the new `PUT` and the `edit=1` query.
- Release: a normal COM applies (the 1.10.0 batch hold is over). This is a new user-facing feature plus an
additive API surface, so **COM minor** when it ships.
Formatting note: `panels-ui.js`, `styles.css`, `mobile.css`, `index.html` are all in `.prettierignore` and
are hand-formatted; new TypeScript (`src/config/file-editing.ts`, route + schema edits) is prettier-enforced
and must pass `npm run format:check`.
---
## 9. Implementation order
Each phase is independently reviewable and leaves the tree working.
1. **Policy module + tests.** `src/config/file-editing.ts` and `test/file-editing-policy.test.ts`. Pure, no
route wiring. (Small.)
2. **Read-for-edit.** `edit=1` (returning `hash`/`eol`) plus the additive `editable` flag on plain reads,
tests. Nothing consumes it yet. (Small.)
3. **Write endpoint.** `FileWriteSchema`, `PUT` handler, `test/routes/file-write-routes.test.ts`. Fully
testable by curl before any UI exists. (Medium, the security-relevant part.)
4. **Desktop UI.** Edit button, textarea swap, Save/Cancel, dirty guard, 409 flow. (Medium.)
5. **Mobile pass.** `mobile.css` sizing against `--app-height`, font size, accessory-bar interaction,
real-device check. (Small but the part that decides whether the feature is actually usable.)
6. **Docs, i18n strings, changeset.**
---
## 10. Open decisions
1. **Editor widget.** Recommend a plain `<textarea>` for v1: zero dependencies, no CSP question, no bundle
growth, and it is the only thing guaranteed to behave with the iOS keyboard. CodeMirror-light with
syntax highlighting is a clean follow-up once the write path is proven. The issue allows either.
2. **Phone entry point.** Edit mode is reachable on a phone through attachment cards and the history
drawer without changing anything. A dedicated toolbar or overview affordance for "browse this session's
files" would make it discoverable, but it is a separate UX change and would need a decision against the
deliberately minimal phone header policy. Recommend deferring it and revisiting after the feature ships.
3. **`svg` editability.** Recommend excluded in v1 (it is deliberately treated as untrusted on the read
side). Easy to add later.
4. **Create / delete / rename.** Explicitly out of scope per the issue. Note that keeping `O_CREAT` out of
the handler is what makes that a structural property rather than a convention.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 357 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 941 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 1.0 MiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 56 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 859 KiB

-3
View File
@@ -1,3 +0,0 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 320 60">
<text x="160" y="48" font-family="system-ui, -apple-system, 'Segoe UI', Roboto, sans-serif" font-size="52" font-weight="700" fill="#60a5fa" text-anchor="middle">Codeman</text>
</svg>

Before

Width:  |  Height:  |  Size: 247 B

Binary file not shown.

Before

Width:  |  Height:  |  Size: 357 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 96 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 87 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 84 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 195 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 50 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 3.0 MiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 583 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 4.3 MiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 226 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 138 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 537 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 34 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 207 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 332 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 808 KiB

-233
View File
@@ -1,233 +0,0 @@
# Local Echo Overlay — Implementation Plan
> **Status: SHIPPED.** Implementation lives in `packages/xterm-zerolag-input/src/` (overlay-renderer.ts, prompt-finder.ts, cell-dimensions.ts, zerolag-input-addon.ts) with the embedded copy in `src/web/public/app.js`. This document is retained as historical design context.
## Context
User accesses Codeman remotely from Thailand to Switzerland over Tailscale (~200-300ms RTT).
Every keystroke is invisible for 200-300ms before the server echoes it back. This makes typing
painfully slow on mobile. Previous attempts to write directly to xterm.js buffer failed because
Ink (Claude Code's terminal framework) does full-screen redraws that corrupt injected characters.
## Approach: DOM Overlay (Mosh-inspired)
A single absolutely-positioned `<span>` inside xterm.js's `.xterm-screen` element that shows
typed characters at the cursor position. This completely avoids buffer conflicts with Ink because
we never write to xterm.js's buffer — the overlay is a pure DOM element sitting on top.
**Why this works when buffer writes don't:** Ink owns the terminal buffer and does full-line
redraws. A DOM overlay sits in a separate rendering layer (z-index 7) and doesn't interfere
with Ink's cursor management or screen redraws at all. When Ink redraws (server output arrives),
we simply hide the overlay.
**Why it will look indistinguishable:** We use the DOM renderer (not canvas/WebGL), so both
terminal text and overlay text are rendered by the same browser font engine with identical
sub-pixel rendering. (Originally designed against xterm.js v5.3.0; project now on `@xterm/xterm` ^6.0.0 — the internal `_core._renderService.dimensions` access path still works in v6.)
## Key Technical Details (from research)
### Pixel Positioning Formula
```js
// Same formula used by BufferDecorationRenderer, CompositionHelper, Terminal._syncTextArea
const dims = terminal._core._renderService.dimensions;
const left = cursorX * dims.css.cell.width; // CSS pixels, relative to .xterm-screen
const top = cursorY * dims.css.cell.height; // CSS pixels, relative to .xterm-screen
```
- `cursorX` = `terminal.buffer.active.cursorX` (0 to terminal.cols)
- `cursorY` = `terminal.buffer.active.cursorY` (0 to terminal.rows-1, ALREADY viewport-relative)
- No scroll offset math needed
### Cell Dimensions (no public API in v5/v6 — use internal; public in v7+)
```js
const dims = terminal._core._renderService.dimensions;
dims.css.cell.width // e.g., 8.4px
dims.css.cell.height // e.g., 17px
```
Public `terminal.dimensions` only available in v7.0.0+.
### xterm.js DOM Structure
```
div.terminal.xterm
├── div.xterm-viewport (overflow-y: scroll)
└── div.xterm-screen (position: relative) ← INSERT OVERLAY HERE
├── div.xterm-helpers (z-index: 5)
├── div.xterm-rows (the actual text) (z-index: auto/0)
├── div.xterm-selection (z-index: 1)
└── div.xterm-decoration-container (z-index: 6-7)
```
### Z-Index Layers
| Layer | Z-Index |
|-------|---------|
| textarea | -5 |
| row content (DOM renderer) | auto (0) |
| selection | 1 |
| composition (IME) | 1 |
| helpers | 5 |
| decorations | 6 |
| decorations (top layer) | 7 ← OUR OVERLAY |
| overview ruler | 8 |
| accessibility | 10 |
### Font Matching CSS
```css
.local-echo-overlay {
position: absolute;
z-index: 7;
pointer-events: none;
white-space: pre;
font-kerning: none;
overflow: hidden;
display: none;
/* Set dynamically: left, top, height, line-height, font-family, font-size, color, letter-spacing */
}
```
Critical: match `letter-spacing` from `.xterm-rows` container (DPR rounding compensation).
### Font Properties from Terminal
```js
terminal.options.fontFamily // '"Fira Code", "Cascadia Code", ...'
terminal.options.fontSize // 14 (10 on mobile)
terminal.options.fontWeight // 'normal'
terminal.options.letterSpacing // 0
terminal.options.lineHeight // 1.2
```
Use actual `dims.css.cell.height` for line-height (not the multiplier).
## Files to Modify
### `src/web/public/app.js` — All logic
1. **Constructor** (~line 1455): Initialize overlay state variables
2. **After terminal creation** (in `setupTerminal` or similar): Create overlay DOM element
3. **`terminal.onData` handler** (~line 1801): Echo printable chars to overlay when idle
4. **`flushPendingWrites`** (~line 2083): Hide overlay when server output arrives
5. **SSE event handlers**: Update overlay state on session:idle/working/exit
6. **`selectSession`**: Clear overlay on tab switch
7. **`handleInit`**: Clear overlay on SSE reconnect
8. **Settings load/save** (`openAppSettings`/`saveAppSettings`): Toggle checkbox
### `src/web/public/index.html` — Settings toggle
After Image Watcher section (~line 878), add "Input" section with checkbox.
## Implementation Details
### Overlay Class (inline in app.js, near extractSyncSegments)
```js
class LocalEchoOverlay {
constructor(terminal) {
this.terminal = terminal;
this.overlay = document.createElement('span');
// ... CSS setup ...
const screen = terminal.element.querySelector('.xterm-screen');
screen.appendChild(this.overlay);
this.pendingText = '';
this.timeout = null;
}
addChar(char) {
this.pendingText += char;
this._render();
this._resetTimeout();
}
removeChar() {
if (this.pendingText.length > 0) {
this.pendingText = this.pendingText.slice(0, -1);
this._render();
if (this.pendingText.length > 0) this._resetTimeout();
else this._clearTimeout();
}
}
clear() {
this.pendingText = '';
this.overlay.textContent = '';
this.overlay.style.display = 'none';
this._clearTimeout();
}
_render() {
if (!this.pendingText) { this.clear(); return; }
const dims = this.terminal._core._renderService.dimensions;
const cellW = dims.css.cell.width;
const cellH = dims.css.cell.height;
const cursorX = this.terminal.buffer.active.cursorX;
const cursorY = this.terminal.buffer.active.cursorY;
this.overlay.style.left = (cursorX * cellW) + 'px';
this.overlay.style.top = (cursorY * cellH) + 'px';
this.overlay.style.height = cellH + 'px';
this.overlay.style.lineHeight = cellH + 'px';
this.overlay.textContent = this.pendingText;
this.overlay.style.display = '';
}
_resetTimeout() {
this._clearTimeout();
this.timeout = setTimeout(() => this.clear(), 2000);
}
_clearTimeout() {
if (this.timeout) { clearTimeout(this.timeout); this.timeout = null; }
}
get hasPending() { return this.pendingText.length > 0; }
dispose() {
this.clear();
this.overlay.remove();
}
}
```
### Integration Points
**Input handler (`terminal.onData`):**
- Backspace (`\x7f`): if overlay has pending + echo enabled → `overlay.removeChar()`
- Enter (`\r`/`\n`): `overlay.clear()`, disable echo (session goes busy)
- Other control chars / multi-char (paste): `overlay.clear()`
- Single printable char (charCode >= 32, length === 1): if echo enabled → `overlay.addChar(data)`
**Output handler (`flushPendingWrites`):**
- After writing segments: if overlay has pending text → `overlay.clear()` (server confirmed)
**State management:**
- `_localEchoEnabled` boolean, updated on session status change + settings change
- Only enabled when: setting on + active session is idle
- On idle→busy transition: clear overlay
- On tab switch: clear overlay
- On SSE reconnect: clear overlay
### Settings
**index.html:** Checkbox `appSettingsLocalEcho` under "Input" section header
**openAppSettings:** Load `settings.localEchoEnabled ?? false`
**saveAppSettings:** Save checkbox + call `_updateLocalEchoState()`
Default: **disabled** (opt-in)
## Edge Cases
| Case | Handling |
|---|---|
| Paste (multi-char onData) | data.length > 1 → NOT echoed. Server echoes it. |
| Misprediction | Server output arrives → overlay cleared → server redraws correctly |
| Idle→busy race | _updateLocalEchoState() disables + clears overlay |
| Server unresponsive | 2s timeout → overlay cleared |
| Tab switch | selectSession() clears overlay |
| SSE reconnect | handleInit() clears overlay |
| Terminal resize | Overlay position recalculated on next _render() |
| Scrolled back | cursorY is viewport-relative, position stays correct |
| Unicode/emoji | data.length > 1 → not echoed (ASCII-only) |
## What NOT to Do
- Do NOT write to `terminal.write()` — Ink conflicts
- Do NOT use `registerDecoration` — requires markers, can't follow cursor smoothly
- Do NOT try to match predictions against server output — Ink's full-line redraws make this impossible
- Do NOT use `stripAnsiForMatch` / `findEscapeEnd` — removed, not needed for overlay approach
-228
View File
@@ -1,228 +0,0 @@
# Mobile E2E Testing Report
**Date**: 2026-01-31
**Status**: All 32 tests passing
## Overview
Comprehensive mobile E2E testing was performed using Playwright with Chromium in mobile emulation mode. Tests validate touch interactions, responsive design, mobile-specific UI behaviors, and edge cases across various device viewports.
## Test Coverage Summary
| Test File | Tests | Description |
|-----------|-------|-------------|
| `mobile-safari.e2e.ts` | 6 | Core mobile Safari/iPhone tests |
| `mobile-comprehensive.e2e.ts` | 13 | UI components, modals, interactions |
| `mobile-edge-cases.e2e.ts` | 13 | Edge cases: orientation, narrow screens, safe areas |
## Bugs Found and Fixed
### 1. Monitor Panel Overlapping Toolbar on Mobile
**File**: `src/web/public/styles.css` (lines 7969-7982)
**Problem**: The monitor panel was positioned at `bottom: var(--toolbar-height)` (40px), but the mobile toolbar has `height: auto` with `flex-wrap: wrap`, causing it to be taller than 40px. This resulted in the monitor panel header intercepting tap events on the "Run Claude" button.
**Error message**:
```
<div class="monitor-panel-title">Monitor</div> from <div id="monitorPanel" class="monitor-panel">…</div> subtree intercepts pointer events
```
**Fix**: Hide monitor and subagents panels on phones by default:
```css
@media (max-width: 430px) {
.monitor-panel,
.subagents-panel {
display: none !important;
}
}
```
**Rationale**: On phone screens (<430px), there isn't enough space for these panels anyway. Users can still access session info via the header and session options modal.
---
### 2. WebKit Browser Missing System Dependencies
**File**: `test/e2e/fixtures/mobile-browser.fixture.ts`
**Problem**: WebKit requires system libraries (libgtk-4, libgstreamer, etc.) that may not be installed on all systems, causing mobile tests to fail.
**Fix**: Added fallback to Chromium with mobile emulation:
```typescript
try {
browser = await webkit.launch({ headless: true });
userAgent = 'Mozilla/5.0 (iPhone; CPU iPhone OS 18_0 like Mac OS X)...';
} catch {
// WebKit failed, use Chromium with mobile emulation
browser = await chromium.launch({ headless: true, args: [...] });
userAgent = 'Mozilla/5.0 (Linux; Android 14; Pixel 8)...';
}
```
---
### 3. Race Condition in Session Tab Detection
**File**: `test/e2e/workflows/mobile-safari.e2e.ts`
**Problem**: Test waited for `.session-tab` selector but then checked `.session-tab.active`, causing timing issues where the tab existed but wasn't yet marked as active.
**Fix**: Wait for the active tab directly:
```typescript
// Before (race condition)
await page.waitForSelector('.session-tab', { timeout: ... });
const tabVisible = await page.isVisible('.session-tab.active');
// After (correct)
await page.waitForSelector('.session-tab.active', { timeout: ... });
const tabVisible = await page.isVisible('.session-tab.active');
```
---
## Known Limitations
### No Kill All Button on Mobile
**Status**: By design (not a bug)
The "Kill All" button is located in the Monitor panel, which is hidden on mobile devices (<430px). Users can close sessions individually via the close button on each session tab.
**Consideration for future**: Could add a "Kill All" option in the app settings modal or a long-press context menu on session tabs.
### No Help Button on Mobile
**Status**: By design
There is no dedicated help button in the mobile UI. Help is accessible via:
- Keyboard shortcut (`?` key)
- App settings modal
---
## Test Coverage
### mobile-safari.e2e.ts (Port 3191)
| Test | Description |
|------|-------------|
| Touch-friendly UI rendering | Verifies `touch-device` and `device-mobile` body classes |
| 44px minimum touch targets | Ensures buttons meet WCAG AA touch target requirements |
| Tap gestures for session creation | Creates session via tap on Run Claude button |
| Always-visible close buttons | Verifies opacity:1 on touch devices (no hover dependency) |
| Header hiding on small screens | Brand, stats, font controls hidden on phones |
| Tablet viewport rendering | iPad Pro 11" (834x1194) renders with `device-desktop` + `touch-device` |
### mobile-comprehensive.e2e.ts (Port 3192)
| Test | Description |
|------|-------------|
| Welcome overlay buttons | Touch-friendly welcome overlay with 44px+ button height |
| Run Claude button prominence | Button visible with `flex: 1` on mobile |
| Case dropdown visibility | Dropdown accessible and functional |
| Version display hiding | `.toolbar-center` hidden on phones |
| Horizontal tab scrolling | Session tabs allow `overflow-x: auto` scrolling |
| Tab switching on tap | Tapping tabs switches active session |
| Full-screen modals | Modals use 100% width/height on phones |
| Create case modal | Case creation modal accessible via + button |
| Notification button | Notification bell visible and tappable |
| Settings button | Settings gear has adequate touch target |
| Close confirmation modal | Close button triggers confirmation dialog |
| Token count display | Token counter visible in header |
| Ralph wizard full-screen | Wizard modal renders full-screen |
### mobile-edge-cases.e2e.ts (Port 3193)
| Test | Description |
|------|-------------|
| Landscape orientation handling | 874x402 landscape mode with proper classes |
| Terminal in landscape | Terminal renders with adequate height |
| Very narrow viewport (280px) | Galaxy Fold folded state usable |
| Narrow screen toolbar | Toolbar doesn't overflow on 280px |
| Session options via gear icon | Gear icon visible, modal opens on tap |
| Modal tab switching | Session options modal tabs work on touch |
| Terminal tap interactions | Terminal responds to touch events |
| Primary touch targets | Main buttons meet 44px height requirement |
| iOS safe area CSS variables | `--safe-area-*` variables defined |
| Double-tap zoom prevention | touch-action styles applied |
| Modal body scrolling | `overflow-y: auto` for touch scrolling |
| Viewport meta tag | Proper mobile viewport configuration |
| Android Pixel viewport | 412x915 Pixel 7a renders correctly |
---
## Mobile CSS Breakpoints
| Breakpoint | Class | Description |
|------------|-------|-------------|
| < 430px | `device-mobile` | Phone - most features hidden/simplified |
| 430-768px | `device-tablet` | Tablet - intermediate layout |
| > 768px | `device-desktop` | Desktop - full features |
Touch devices also get `touch-device` class regardless of screen size.
---
## Viewports Tested
| Device | Width | Height | Scale | Notes |
|--------|-------|--------|-------|-------|
| iPhone 17 Pro | 402 | 874 | 3x | Primary phone test |
| iPhone 17 Pro Landscape | 874 | 402 | 3x | Orientation testing |
| iPhone 17 Pro Max | 440 | 956 | 3x | Larger phone |
| iPad Pro 11" | 834 | 1194 | 2x | Tablet testing |
| Galaxy Fold (folded) | 280 | 653 | 3x | Extreme narrow test |
| Pixel 7a | 412 | 915 | 2.625x | Android testing |
---
## Running Mobile Tests
```bash
# Install Playwright browsers (Chromium is required, WebKit optional)
npx playwright install chromium
# Run individual test files
npx vitest run test/e2e/workflows/mobile-safari.e2e.ts
npx vitest run test/e2e/workflows/mobile-comprehensive.e2e.ts
npx vitest run test/e2e/workflows/mobile-edge-cases.e2e.ts
# Run all mobile tests together
npx vitest run test/e2e/workflows/mobile-safari.e2e.ts test/e2e/workflows/mobile-comprehensive.e2e.ts test/e2e/workflows/mobile-edge-cases.e2e.ts
```
---
## Port Allocations
| Port | Test File |
|------|-----------|
| 3191 | mobile-safari.e2e.ts |
| 3192 | mobile-comprehensive.e2e.ts |
| 3193 | mobile-edge-cases.e2e.ts |
---
## Key Mobile UI Behaviors
1. **Monitor/Subagents panels**: Hidden on phones (<430px)
2. **Toolbar**: Wraps content with `flex-wrap: wrap`, variable height
3. **Session tabs**: Horizontal scroll with hidden scrollbar
4. **Modals**: Full-screen on phones (100% width/height)
5. **Touch targets**: Minimum 44px height for WCAG compliance
6. **Close buttons**: Always visible (opacity: 1) on touch devices
7. **Header**: Brand, stats, font controls hidden on phones
8. **Safe areas**: CSS variables for iOS notch handling
---
## Future Improvements
1. Add swipe gesture tests for tab navigation
2. Add virtual keyboard handling tests (show/hide behavior)
3. Add orientation change tests (dynamic portrait/landscape switching)
4. Add safe area inset tests for iOS notch handling with actual device values
5. Consider showing a condensed monitor indicator on mobile
6. Add "Kill All" option accessible from mobile UI
7. Test pull-to-refresh prevention on iOS Safari
-282
View File
@@ -1,282 +0,0 @@
# Multi-User Mode: Design Plan
Status: **IMPLEMENTED on `feat/multiuser-mode`** (phases 1-5; opt-in, off by default). Target: opt-in multi-user support behind a `--multiuser` flag, with per-user case spaces and an admin panel for user management.
Shipped by phase:
- **Phase 1** (user store + mode plumbing + CLI): `src/user-store.ts` (scrypt, atomic 0600 writes, last-admin invariants, serialized read-modify-write), `src/config/multiuser.ts`, `codeman users add|passwd|list|rm`, `--multiuser` flag, bootstrap-on-first-boot. Tests: `test/user-store.test.ts`.
- **Phase 2** (multi-user auth): parallel async auth branch (`src/web/middleware/auth.ts`), `req.authUser`, per-username rate bucket, `mustChangePassword` lockbox, `GET /api/me` + `POST /api/me/password`, QR identity-bound minting, network-bind + tunnel exemptions, new error codes. Tests: `test/multiuser-auth.test.ts`.
- **Phase 3** (ownership threading): `Session.owner` at every create path + recovery mirror; `findSessionOrFail` owner check + list filtering; §6.3 permission policy (`resolveClaudeModeForUser` at all spawn sites incl. one-shots via `buildPromptArgs`; shell/launchCommand grant); per-user case spaces (`resolveCasesDir`) + owner-scoped case list + admin-only host CRUD; `workingDir` confinement; `sessionCapacityState` per-user cap. Tests: `test/ownership-scoping.test.ts`.
- **Phase 4** (event fan-out): WS owner gate; SSE per-client identity + `broadcast`/terminal-batch routing (`deriveSseHint`, fail-closed); `getLightState` per-identity filtering; file-route preview/thumbnail/history + `GET /api/search` scoping.
- **Phase 5** (admin API + frontend): `src/web/routes/admin-routes.ts` (user CRUD, one-time passwords, last-admin guards, session revoke/kill) + `src/web/admin-audit.ts`; `public/admin-ui.js` (identity boot, change-password modal + interceptor, admin Users tab). Tests: `test/admin-routes.test.ts`, `test/admin-ui.test.ts`.
Deferred follow-ups (documented, non-blocking): away-digest + subagent/workflow REST-list scoping, push-subscription identity/routing, per-user screenshot subdirs, `linked-cases.json` v2 owner field, `ScheduledRun.owner`, plan-orchestrator internal one-shot mode resolution, and a Playwright browser pass. Phase 6 (login form replacing Basic) remains out of scope.
## 1. Summary
Today Codeman is strictly single-user: one optional credential pair (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`), one shared `~/codeman-cases` folder, one global session list, and a global SSE/WS fan-out. This plan adds an opt-in **multi-user mode**:
- **Off by default.** Without the flag, behavior stays byte-identical to today (same auth path, same paths, same payloads). All new code is gated behind `isMultiUserMode()`.
- **`codeman web --multiuser`** (or `CODEMAN_MULTIUSER=1`) enables named users with individually hashed passwords stored in `~/.codeman/users.json`.
- **Each user gets their own space**: `~/codeman-users/<username>/cases/<case>` replaces the shared `~/codeman-cases` for that user. Sessions, cases, attachments, search, digests, and SSE events are scoped to their owner.
- **Admin panel** (App Settings, admin-only "Users" tab): create/delete users, change/reset passwords, enable/disable accounts, delete a user's space, see per-user live sessions and disk usage, force logout.
## 2. Threat Model (read first, be honest about this)
Multi-user mode is **workspace separation for a trusted team, NOT security isolation between mutually distrusting users**:
- Every session still runs as the **same OS account** with `claude --dangerously-skip-permissions`. Any user can ask their agent to `cat /home/<host>/codeman-users/otheruser/...`. The web layer enforces scoping; the agent layer cannot.
- **Shell sessions and custom launch commands are the bluntest holes**: `SessionMode = 'shell'` hands out a raw shell as the host account, and a cron job's `launchCommand` runs an arbitrary command; no Claude permission classifier is involved in either. These must be gated behind the same grant as bypass (section 6.3), otherwise the `auto`-mode mitigation below is theater.
- All sessions share one tmux socket (`-L codeman`), one `~/.claude` (transcripts, credentials, plan usage), one Claude subscription.
- Mitigation for stronger isolation: pair a user's cases with **Docker cases** (container per case, `docs/docker-cases.md`), or run separate Codeman instances per user (`CODEMAN_INSTANCE`, separate OS accounts). True per-user OS isolation is explicitly **out of scope** for this feature.
- Partial mitigation at the agent layer: non-admin users default to Claude's `auto` permission mode (section 6.3), whose safety classifier blocks destructive actions and credential exfiltration. That reduces, but does not eliminate, cross-user snooping; the `canBypassPermissions` grant reopens it and should be given deliberately.
This must be stated loudly in `docs/security-architecture.md`, the README section, and the admin panel UI ("Users share the host account; this separates workspaces, it does not sandbox users from each other").
Also note the flip side: multi-user mode strictly _improves_ today's network posture, because it removes the single shared password and gives every person their own revocable credential.
## 3. Activation and Mode Rules
| Condition | Behavior |
| ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| No flag (default) | Exactly today's behavior. `users.json` is never read. Single-user auth via `CODEMAN_PASSWORD` if set. |
| `--multiuser` / `CODEMAN_MULTIUSER=1`, `users.json` has users | Multi-user auth active. `CODEMAN_PASSWORD` is ignored for login (warn if set). |
| `--multiuser`, no `users.json` (first boot) | Bootstrap: if `CODEMAN_USERNAME`/`CODEMAN_PASSWORD` are set, create that user as the initial admin and continue. Otherwise refuse to start with instructions to run `codeman users add <name> --admin`. Never start multi-user with zero users (there would be no way in). |
| `--multiuser` on a non-loopback bind | Allowed without `CODEMAN_PASSWORD`: `server.ts start()` treats "multi-user with >= 1 enabled user" as satisfying the auth requirement in the loud-warning check (wire into the existing `isLoopbackBindHost()` branch). |
| Flag later removed | Single-user mode again. Sessions/state that carry `owner` fields keep working (owner is simply ignored); user spaces remain on disk untouched. |
Plumbing: flag in `src/cli.ts` (web command), env in a new `src/config/multiuser.ts` exporting `isMultiUserMode()`. Per-instance like everything else: a beta instance (`CODEMAN_INSTANCE=beta`) has its own `users.json` via `dataPath()`.
## 4. Data Model and Disk Layout
### 4.1 `~/.codeman/users.json` (via `dataPath('users.json')`, mode 0600, atomic write: tmp + rename)
```jsonc
{
"version": 1,
"users": [
{
"username": "alice", // canonical lowercase slug
"role": "admin", // "admin" | "user"
"password": {
"algo": "scrypt", // node:crypto scrypt, no new deps
"N": 16384,
"r": 8,
"p": 1,
"salt": "<hex 32B>",
"hash": "<hex 64B>",
},
"disabled": false,
"mustChangePassword": false, // set by admin reset; gates all API access until changed
"canBypassPermissions": false, // permission-mode grant, see section 6.3; false for new users
"createdAt": 1752900000000,
"lastLoginAt": 1752900000000,
},
],
}
```
- **Username rules**: `^[a-z0-9][a-z0-9_-]{1,31}$` (it becomes a folder name), stored lowercase, unique case-insensitively. Reserve `admin`? No: any name can be admin; role is a field, not a name.
- **Hashing**: `scrypt` from `node:crypto` with per-user salt, compared via `timingSafeEqual`. Params stored per record so they can be raised later; verify tolerates old params and rehashes on next successful login.
- New module `src/user-store.ts` (mirrors the `remote-hosts.ts` / `docker-hosts.ts` pattern): `readUsers()`, `writeUsers()`, `verifyPassword()`, `createUser()`, `setPassword()`, `deleteUser()`, plus pure helpers (`isValidUsername`, `hashPassword`) that are unit-testable without IO. In-process cache with short TTL like `readSettings`, invalidated on every write; the short TTL also covers the CLI (section 10) editing `users.json` while the server runs (cross-process changes picked up within the TTL).
### 4.2 User spaces
```
~/codeman-users/
alice/
cases/
my-project/ <- same layout as today's ~/codeman-cases/<case>
bob/
cases/
```
- New helper in `route-helpers.ts`:
`resolveCasesDir(user?: AuthUser): string`
single-user mode: returns `CASES_DIR` (today's `~/codeman-cases`); multi-user: returns `join(USER_SPACES_DIR, user.username, 'cases')`, creating it lazily on first use.
- `CASES_DIR` stays exported for single-user code paths, but every route usage (see 6) switches to the resolver.
- The **user folder** (`~/codeman-users/<username>/`) is the deletion unit for "delete user + space" and leaves room for future per-user extras (uploads, exports) beside `cases/`.
- Legacy `~/codeman-cases` in multi-user mode: surfaces to admins only, as a read-only "Unassigned (legacy)" group in the case list, with an admin action `POST /api/admin/cases/assign { case, username }` that `fs.rename`s the folder into a user's space (same-filesystem move, cheap). No automatic migration.
## 5. Auth Pipeline Changes (`src/web/middleware/auth.ts`)
Keep the existing single-user branch untouched. Add a parallel multi-user branch selected once at registration time:
1. **Credential check**: Basic header parsed into `username:password`, verified against the user store (scrypt + `timingSafeEqual`). Disabled users fail closed.
2. **Cookie sessions**: same `codeman_session` cookie and `StaleExpirationMap`, but `AuthSessionRecord` gains `username` and `role`. All existing TTL/sliding/eviction logic reused. Eviction cap becomes per-user aware (evict oldest _of that user_ first) so one user cannot flush everyone's sessions by logging in 100 times.
3. **Request identity**: decorate `req.authUser = { username, role }` (Fastify decorateRequest). In single-user mode `req.authUser` is `{ username: 'admin', role: 'admin' }` when auth is on, and a synthetic admin when auth is off, so downstream code has ONE code path.
4. **Rate limiting**: keep the per-IP bucket; add a per-username failure bucket (same `StaleExpirationMap` pattern) so a botnet cannot brute-force one account across IPs, and one flaky user behind a NAT cannot lock out the rest.
5. **`mustChangePassword` gate**: when set, every API request except `GET /api/me`, `POST /api/me/password`, and static assets returns 403 with `errorCode: 'PASSWORD_CHANGE_REQUIRED'`; the frontend intercepts that code and shows the change-password modal.
6. **Password change vs Basic-auth caching**: browsers cache Basic credentials. After a password change we revoke all of that user's cookie sessions; the next request falls to Basic with stale creds, gets 401, and the browser re-prompts. Acceptable for v1; a proper login form is Phase 6 (see 15).
7. **Unchanged**: hook-secret loopback bypass (hooks authenticate the _instance_, not a user; the event maps to a session which has an owner), host guard, Origin/CSRF guard, security headers.
8. **WS upgrade identity** (`ws-routes.ts`): the global auth `onRequest` hook does run on the upgrade request (`@fastify/websocket` v11 runs hooks before the handshake; browsers send the session cookie), but the route handler itself only checks Host/Origin and never learns WHO authenticated. Multi-user: the handler reads the decorated `req.authUser` and closes 4003 unless owner or admin (section 6.4; identity plumbing lands in Phase 2, the owner check in Phase 4 once sessions have owners). Add a regression test that an upgrade with no credentials is rejected while auth is active: the handler-level Host/Origin gate alone must never be mistaken for auth.
9. **QR auth** (`/q/:code` redemption in `system-routes.ts`, minting in `tunnel-manager.ts`): today there is ONE global token, auto-rotated every 60s with a 90s grace window. A globally-rotating token cannot carry an identity (every logged-in user sees the same code), so multi-user mode replaces rotation with **on-demand minting**: an authenticated `POST /api/tunnel/qr` mints a single-use, short-TTL token bound to `req.authUser.username` (field on `QrTokenRecord`); redemption creates a cookie session for that user. Existing rate-limit buckets (`qrAuthFailures`, global `QR_RATE_LIMIT_MAX`) apply unchanged. Single-user mode keeps the rotating token.
New error codes in `src/types/api.ts`: `FORBIDDEN`, `PASSWORD_CHANGE_REQUIRED`, `USER_EXISTS`, `USER_NOT_FOUND`, `LAST_ADMIN`.
Role guard helper in `route-helpers.ts`: `requireAdmin(req, reply): boolean` used as the first line of every admin handler (403 `FORBIDDEN`), plus `requireOwnerOrAdmin(req, session)`.
## 6. Ownership Threading (the big refactor)
### 6.1 Sessions
- `Session` gains `owner?: string` (constructor option), persisted in `SessionState.owner`, included in `toState()`, round-tripped through recovery (`mux-sessions.json` entries carry it, `restoreMuxSessions` passes it back, exactly like `remote`/`docker`).
- Every session-creating path stamps the owner from `req.authUser`. Verified inventory of `new Session(...)` call sites: `POST /api/sessions` (session-routes.ts:444), `POST /api/quick-start` (:1956), `POST /api/run` one-shot (:1652), Ralph start (ralph-routes.ts:327), **cron** (cron-service.ts:352; `CronJob` gains `owner`, stamped at job create, launched as the job's owner), legacy `ScheduledRun` loop (server.ts:1603), plan generation + plan-orchestrator agents (plan-routes.ts:128, plan-orchestrator.ts:422/578; owner = requesting user), and recovery (server.ts:2225, next bullet). Two non-paths, also verified: **respawn never constructs a new Session** (it re-spawns the PTY on the same object, so `owner` survives automatically; no inheritance logic needed), and **orchestrator-loop creates no sessions** (it schedules work onto existing idle sessions via the task queue; its scoping requirement is different: it must only pick idle sessions owned by the goal's creator).
- Recovery: `owner` must ALSO be mirrored on `MuxSession` (mux-sessions.json) and read back mux-first like `remote`/`docker` (`muxSession.owner ?? savedState?.owner`, the server.ts:2246-2250 pattern), or a reboot erases ownership on the next persist.
- Every session-reading/mutating route filters: non-admin users only see and act on `session.owner === req.authUser.username`. Centralize in `findSessionOrFail` (route-helpers.ts:87; the owner check there covers the 6 route files that use it: system/session/respawn/ralph/file/plan-routes) and in the list endpoints (`GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status`). The Phase 3 audit must grep for BOTH `sessionManager.getSession` AND direct map access (`ctx.sessions.get(` / `.has(`): ws-routes and hook-event-routes reach sessions that way and bypass `findSessionOrFail`.
- Admins see everything; every session row carries `owner` so the UI can badge it.
### 6.2 Cases
- All `CASES_DIR` call sites switch to `resolveCasesDir(req.authUser)`: `case-routes.ts` (list/create/delete/CLAUDE.md scaffolding, name-collision checks, docker quickcreate), `session-routes.ts` (quick-start case resolution, the workingDir-inside-cases env-strip check), `ralph-routes.ts` (case path resolution), and `plan-routes.ts:231` (easy to miss). Case-name-to-path resolution is currently DUPLICATED (`resolveCasePath` in case-routes.ts:82 and an inline copy in quick-start, session-routes.ts:1846-1863); consolidate into one owner-aware resolver as part of this refactor instead of patching both copies.
- Registries that map case names to metadata become owner-scoped. `remote-cases.json`/`docker-cases.json` are arrays of objects, so entries simply gain `owner?: string` (absent = legacy: admin-only). `linked-cases.json` is a flat `Record<caseName, path>` with no room for a field: it needs a v2 shape (`{ "version": 2, "cases": { "<name>": { "path": "...", "owner": "..." } } }`) with read-time migration of the v1 form; it is read in two places (case-routes AND inline in quick-start), both must move to the new reader. Case names only need to be unique per user.
- **Remote hosts and Docker hosts are machine-level resources**: CRUD on `/api/docker-hosts` and remote-host endpoints becomes admin-only in multi-user mode; regular users can _use_ hosts on their own cases but not define them. (Docker containers exec as the host account; letting any user define arbitrary `docker run` args is admin-equivalent.)
- Case deletion, exports (`docker-exports/`), and imports check ownership; export filenames get an owner prefix to avoid collisions (fits the existing `^[a-zA-Z0-9._-]+\.tgz$` download guard).
- **Workspace confinement for non-admins (the linchpin, do not skip)**: today `POST /api/sessions` accepts ANY host directory as `workingDir` (the only check is `statSync().isDirectory()`, session-routes.ts:305-318), and file-routes/attachments confine reads to `session.workingDir`. Without a new rule the whole scoping story is circular: a user points a session at `~/codeman-users/bob` (or `/home`) and the web layer itself serves that subtree, no agent needed. Rule: in multi-user mode a non-admin's `workingDir` must realpath-resolve inside their own space, enforced at `POST /api/sessions`, `POST /api/run`, cron job create AND fire time (the dir can change owners between the two), and Ralph auto-configure. Admins are unrestricted. This one rule is what makes the section 6.4 file-route line ("own space or own sessions' workingDirs") meaningful.
### 6.3 Per-user Claude permission-mode policy
Codeman now ships a global **Startup Mode** picker (App Settings, Claude CLI tab: `settings.claudeMode`, values `dangerously-skip-permissions` (default) | `auto` | `normal` | `allowedTools`; `auto` emits `--permission-mode auto`, Anthropic's classifier-guarded low-prompt mode). Multi-user mode layers a per-user policy on top of it:
- **Default for regular users: `auto` only.** A non-admin's Claude sessions are forced to `--permission-mode auto` regardless of the global `claudeMode` setting. `normal` and `allowedTools` are also permitted (they are strictly more restrictive than auto), but `dangerously-skip-permissions` is NOT.
- **Bypass is an explicit admin grant**: `canBypassPermissions: true` on the user record (default `false`, section 4.1). Only with that grant does the global skip-permissions default (or a future per-user choice) apply to their sessions.
- **Admins** are unrestricted; the global setting applies to them as-is.
- **Single enforcement point**: a pure `resolveClaudeModeForUser(globalMode, user)` in `user-store.ts`, applied server-side at option-resolution time, BEFORE the Session constructor, so both downstream arg builders inherit it for free (`buildPermissionArgs` in session-cli-builder.ts for the direct-PTY path AND `buildClaudePermissionFlags` in tmux-manager.ts for tmux panes; there are two builders, not one). Call sites where `getClaudeModeConfig()` feeds a spawn: session-routes.ts:452/1964, ralph-routes.ts:334, cron-service.ts:360, and recovery (server.ts:2214/2233). Recovery re-reads the GLOBAL setting on reboot, so the resolver must run there with the RECOVERED owner, or a restart silently un-downgrades every restored session. Never resolved in the frontend, so it cannot be bypassed via payload.
- **Downgrade, don't error**: a non-granted user whose effective mode would be bypass gets `auto` silently (logged + surfaced as a badge on the session), so shared presets keep working.
- **Other CLIs' bypass equivalents** follow the same grant: Codex `--dangerously-bypass-approvals-and-sandbox` (`codexDangerouslyBypassApprovals`) and Gemini `--approval-mode yolo` are refused for non-granted users (Gemini falls back to `auto_edit`, Codex to its default sandbox). Whether this stays one grant or splits per-CLI is an open question (section 15).
- **Shell mode and custom launch commands follow the grant too**: `mode: 'shell'` sessions and cron `launchCommand` are arbitrary command execution as the host account, strictly stronger than any bypass flag, and no permission-mode downgrade applies to them. Non-granted users get 403 `FORBIDDEN` on shell session/quick-start creation and on cron jobs carrying `launchCommand` (checked at create AND at fire time). Folding them under `canBypassPermissions` keeps the model one-bit; section 15 asks whether it should split.
- **Admin UI**: a "Can skip permissions" toggle per user in the Users tab (PATCH field, section 8), with a warning echoing the section 2 threat model.
- Revoking the grant takes effect on the user's NEXT session start; live sessions are listed so the admin can restart them.
### 6.4 Everything else that lists or streams
| Surface | Scoping rule |
| -------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| SSE `/api/events` | Per-connection filter (see 7) |
| WS terminal (`ws-routes.ts`) | Handler reads `req.authUser` (section 5.8) and closes 4003 unless owner or admin; today it checks Host/Origin only and has no identity |
| `GET /api/search` | `harvestSources()` only over owned sessions |
| `GET /api/away-digest` | Aggregate only owned sessions/events |
| `GET /api/subagents`, workflow runs | Filter by owning session (`claudeSessionId -> session -> owner`); agents not attributable to any session: admin-only |
| Push (`push-routes.ts`) | Subscription records currently carry NO identity (keyed by endpoint only): `subscribe` stamps `username`. All 8 `PUSH_EVENT_MAP` events are session-scoped, so routing = resolve owner from `data.sessionId`, deliver to that owner's (plus admins') subscriptions. Legacy identity-less subscriptions: admin-only delivery |
| Screenshots `/api/screenshots` | Per-user subdir `~/.codeman/screenshots/<username>/` in multi-user mode. Note: `GET /:name` deliberately rejects `/` in names as traversal, so derive the subdir server-side from `req.authUser` and keep client-visible names flat |
| Attachments | Already session-scoped; inherits the session owner check. `attachmentConfineToWorkspace` is a global, default-OFF setting today: in multi-user mode it is FORCED ON for non-admins regardless of the setting (their attachments must resolve inside their own space); the setting keeps meaning what it means for admins |
| File routes (browse/preview) | Path allowlist adds: non-admin paths must resolve (realpath) inside their own space or their own sessions' workingDirs |
| Settings (`settings.json`) | Global, admin-only writes in multi-user mode; reads allowed (per-device display keys stay in localStorage as today). Per-user server settings: out of scope v1 |
| System ops (self-update, tunnel toggle, span-displays, docker image build) | Admin-only |
| `getLightState` init snapshot | Filtered per connection. Actual contents to filter (verified): `sessions`, `scheduledRuns`, `respawnStatus`, `subagents`, `workflowRuns`, `planUsage` (host-plan telemetry: admin-only); `globalStats` stays coarse-global. Cron jobs are NOT in the snapshot (they have their own REST route; filter there). The snapshot is cached process-wide (`LIGHT_STATE_CACHE_TTL_MS`): either key the cache per role/user or filter AFTER the cache on each send |
## 7. SSE Event Filtering
`/api/events` currently broadcasts everything to everyone. Ground truth first (verified): `broadcast()` lives in `SseStreamManager` (`sse-stream-manager.ts`), not server.ts; clients are keyed by the raw Fastify reply (`sseClients: Map<FastifyReply, Set<string> | null>`, plus `sseClientsById` for live filter updates); the existing `?sessions=` filter is a bandwidth optimization applied ONLY to `session:terminal` batches in `flushSessionTerminalBatch()`, while `broadcast()` itself loops ALL clients unconditionally. The single-client delivery primitive already exists (`sendSSE`, used for the per-connection init snapshot). Plan:
- At connection time, resolve `req.authUser` and store `{ username, role }` with the client. Concretely: extend `addClient(reply, sessionFilter, isRemote, clientId)` to take the identity and change the `sseClients` map value to `{ filter, identity }` (or add a parallel `Map<reply, identity>`); there is no per-client record object today to hang it on.
- `broadcast()` gains an optional routing hint: `broadcast(event, data, { sessionId?, adminOnly?, username? })`. Resolution order per client: admin sees all; `username` targets one user; `sessionId` resolves owner via SessionManager; `adminOnly` for machine-level events (docker image builds, tunnel, self-update); no hint = broadcast to all (connection status etc.).
- **Enforce the identity check in BOTH `broadcast()` AND `flushSessionTerminalBatch()`**: the terminal batch path does not go through `broadcast()`, and it carries the highest-value payload (raw terminal bytes).
- Sweep of the ~120 backend event constants in `sse-events.ts`: mechanically, everything `session:*`, `ralph:*`, `respawn:*`, `subagent:*`, `workflow:*`, `attachment:*`, `cron:*` (job owner) carries or can resolve a sessionId/owner; `docker:*`, `system:*`, tunnel and update events are adminOnly; a short tail needs case-by-case decisions during implementation.
- The existing `?sessions=` filter and `/api/events/subscribe` compose with (never override) the ownership filter: the subscription filter can only narrow within what the identity allows.
## 8. Admin API (`src/web/routes/admin-routes.ts`, new module + `AdminPort`)
All handlers: multi-user mode only (404 otherwise), `requireAdmin`, Zod schemas in `schemas.ts`, `ApiResponse` envelope, audit-logged.
| Endpoint | Behavior |
| ------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `GET /api/admin/users` | List users + stats: role, disabled, createdAt, lastLoginAt, live session count, case count, space disk usage (best-effort async walk, cached 60s), active cookie-session count |
| `POST /api/admin/users` | Create: `{ username, role, password? }`. No password given: generate a one-time password, return it ONCE in the response, set `mustChangePassword` |
| `PATCH /api/admin/users/:username` | `{ role?, disabled?, canBypassPermissions? }`. Demoting/disabling the last enabled admin: 409 `LAST_ADMIN`. Disable also revokes cookie sessions. `canBypassPermissions` is the section 6.3 grant (default false) |
| `POST /api/admin/users/:username/reset-password` | Generates one-time password (returned once), sets `mustChangePassword`, revokes cookie sessions |
| `POST /api/admin/users/:username/logout` | Revoke all cookie sessions for that user. Honest limit under Basic auth: the browser silently re-sends cached credentials and gets a fresh cookie on the next request, so logout only truly ends QR-issued sessions; to actually lock someone out, disable the account or reset the password. Say so in the panel tooltip until Phase 6 |
| `DELETE /api/admin/users/:username` | `{ deleteSpace?: boolean }` (default false). Refuses last admin. Kills the user's live sessions first (normal kill flow, incl. docker/remote teardown per case), revokes cookies, removes from store. With `deleteSpace`: guarded recursive delete of `~/codeman-users/<username>` (realpath must be inside `USER_SPACES_DIR`, top-level dir must not be a symlink), plus their registry entries and push subscriptions |
| `POST /api/admin/cases/assign` | Move a legacy `~/codeman-cases/<case>` into a user's space (`fs.rename`) |
| Self-service `GET /api/me` | `{ username, role, mustChangePassword }` (works in single-user mode too: synthetic admin; the frontend uses it to decide whether to render admin UI) |
| Self-service `POST /api/me/password` | `{ currentPassword, newPassword }`, verifies current, min length 8, revokes other sessions, clears `mustChangePassword` |
**Audit log**: append-only `~/.codeman/admin-audit.jsonl` (same idiom as `session-lifecycle.jsonl`): timestamp, acting admin, action, target, request IP. User management without an audit trail is not acceptable even for a homelab tool.
SSE additions (both `sse-events.ts` and `constants.js`): `admin:usersChanged` (adminOnly; the panel re-fetches) and `auth:passwordChangeRequired` (targeted to the user).
## 9. Frontend
- **`GET /api/me` on boot** (app.js init): stores `window.__codemanUser`; everything below keys off it. Single-user mode returns the synthetic admin, so the UI needs no mode awareness beyond "am I admin".
- **Admin panel**: new tab "Users" in the App Settings modal (settings-ui.js), rendered only for admins in multi-user mode. Table of users with actions (create, reset password showing the one-time password in a copy-to-clipboard reveal, enable/disable, role toggle, logout, delete with a typed-username confirm for the delete-space variant). No new header button (mobile header policy test stays green; the settings modal is already reachable everywhere).
- **Change-password modal**: shown on `PASSWORD_CHANGE_REQUIRED` (fetch interceptor in api-client.js) and reachable from settings for self-service.
- **Owner badges**: admin's session tabs and the session palette/manager show `owner` on foreign sessions; regular users see no change.
- New module `admin-ui.js` if the settings-ui.js addition gets large (load order after settings-ui, before session-ui), else keep inside settings-ui.js. Follow the `@fileoverview` + `@loadorder` convention either way.
## 10. CLI Additions (`src/cli.ts`)
Headless bootstrap and recovery must not require the web UI:
```
codeman users add <name> [--admin] # prompts for password (hidden input), or --password-stdin
codeman users passwd <name> # reset password
codeman users list
codeman users rm <name> [--delete-space]
```
These operate directly on `users.json` via `user-store.ts` (no server needed), honoring `CODEMAN_INSTANCE`. This is also the answer to "locked out: last admin forgot password".
## 11. Limits and Config
- New `src/config/multiuser.ts`: `isMultiUserMode()`, `USER_SPACES_DIR` (`~/codeman-users`, overridable via `CODEMAN_USER_SPACES_DIR` for tests), `MAX_USERS` (default 25), per-user session cap (default: global cap / 2, env `CODEMAN_MAX_SESSIONS_PER_USER`).
- Cap enforcement is currently COPY-PASTED: the global `MAX_CONCURRENT_SESSIONS` (50, `config/map-limits.ts:25`) check appears at 6 independent sites (session-routes.ts:298/1622/1683, ralph-routes.ts:275, cron-service.ts:340, server.ts:1595). Do not add a 7th copy per site: extract one `assertSessionCapacity(ctx, owner?)` helper doing the global + per-user checks and use it everywhere, or the per-user cap WILL miss a path.
- Global limits (50 sessions, SSE clients 100, terminal buffers) are unchanged and shared; the per-user session cap is the fairness lever.
## 12. Compatibility Matrix
| Concern | Guarantee |
| ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Default (no flag) | No behavior change. No new file reads on the hot path. All new fields optional in state |
| State round-trip | `SessionState.owner`, `MuxSession.owner`, `CronJob.owner`, registry `owner` fields are optional; old state loads clean; new state loaded by an old build ignores unknown fields (existing tolerant parsing) |
| Instance isolation | `users.json`, audit log, screenshots subdirs all via `dataPath()`; user spaces dir is shared across instances like `~/codeman-cases` is today (documented) |
| API versioning | HTTP API is internal per `docs/versioning-policy.md`; still, all changes are additive. Ship as a **minor** version |
| Hooks | Unchanged (instance-level hook secret; owner resolved from the session) |
## 13. Implementation Phases
Each phase is independently shippable behind the flag and ends with its tests green.
**Phase 1: user store + mode plumbing** (no behavior change yet)
`src/user-store.ts`, `src/config/multiuser.ts`, CLI `users` subcommands, bootstrap-on-first-boot logic, `users.json` schema + atomic writes.
Tests: `test/user-store.test.ts` (hashing, verify, params upgrade, username validation, atomic write, last-admin invariants; pure, no server).
**Phase 2: multi-user auth**
Auth middleware branch, `req.authUser` decoration, cookie records with username/role, per-username rate bucket, `mustChangePassword` gate, WS upgrade identity plumbing + unauthenticated-upgrade regression test (section 5.8), QR on-demand minting + identity binding (section 5.9), `GET /api/me`, `POST /api/me/password`, error codes, network-bind check integration.
Tests: `test/multiuser-auth.test.ts` (live server, unique port 3170+; wrong password, disabled user, cookie carries identity, per-user rate limit isolation, mustChangePassword lockbox, QR redemption identity). Reuse the `delete process.env.CODEMAN_PASSWORD` idiom from `test/setup.ts`.
**Phase 3: ownership threading**
Session `owner` + persistence + `MuxSession` mirror + recovery; `resolveCasesDir()` refactor across case/session/ralph/plan routes (consolidating the duplicated case-path resolution); registry owner fields incl. the linked-cases v2 shape; `findSessionOrFail` owner check + the direct-`sessions.get` audit; list filtering; owner stamping across ALL create paths from 6.1; **non-admin workingDir confinement** (6.2); permission-mode/shell/launchCommand policy (6.3); `assertSessionCapacity` helper + per-user cap.
Tests: `test/routes/ownership-scoping.test.ts` (inject-based: user A cannot read/kill/input user B's session, case lists are disjoint, admin sees both), extend `test/cron-service.test.ts` for owner stamping, recovery round-trip in the existing mux-recovery tests.
**Phase 4: event fan-out + remaining surfaces**
SSE routing hints + client identity (enforced in BOTH `broadcast()` and the terminal-batch flush), WS owner gate (identity landed in Phase 2), search/digest/subagent/workflow scoping, push subscription identity + owner routing, screenshot subdirs, file-route scoping, `getLightState` filtering + per-identity caching, admin-only system ops.
Tests: `test/sse-ownership.test.ts` (two SSE clients, event for A's session reaches only A + admin), WS upgrade rejection test, search/digest scoping tests.
**Phase 5: admin API + frontend**
`admin-routes.ts` + `AdminPort` + schemas + audit log + `admin:usersChanged`; settings-ui Users tab, change-password modal, owner badges, api-client interceptor.
Tests: `test/routes/admin-routes.test.ts` (CRUD, last-admin 409, one-time password flow, delete-space guard rails incl. symlink refusal), frontend vm-sandbox test following `test/run-mode-ui.test.ts` pattern, Playwright pass per the always-end-to-end rule before calling it done.
**Phase 6 (optional, later): login page**
Replace Basic with a form + `POST /api/login` in multi-user mode only (fixes browser credential caching UX, enables logout button). Explicitly deferred; Basic works for v1.
**Docs**: update `docs/security-architecture.md` (new section: multi-user model + threat model from section 2), `README.md` (short opt-in section), `CLAUDE.md` (Key Patterns entry + State Files + route/SSE counts), this file gets a "shipped" status stamp per phase.
## 14. Key Risks / Decisions Made
1. **Not a security boundary at the agent layer** (section 2). Decided: ship with loud documentation; Docker cases are the isolation story.
2. **`findSessionOrFail` as the single enforcement point** for ~30 session routes: any route that fetches sessions another way must be audited in Phase 3 (grep for `sessionManager.getSession` outside route-helpers).
3. **SSE sweep is the riskiest surface**: a missed event leaks metadata (not terminal content, which is session-scoped, but names/paths). Phase 4 includes a checklist pass over all ~138 events with the default flipped to "owner-scoped unless explicitly global": fail closed.
4. **Basic-auth password-change UX** is mediocre (browser re-prompt). Accepted for v1; Phase 6 fixes it properly.
5. **Legacy case migration** is manual (admin assigns). No silent moves of user data.
6. **Case-name uniqueness becomes per-user**; tmux session names already include the session id so no collision, but the `w<n>-<case>` tab naming and lifecycle-log rows should include the owner for disambiguation in admin views.
7. **`workingDir` confinement (6.2) is the single most load-bearing rule**: every file-serving and agent-spawning surface downstream trusts `session.workingDir`. Review and test it as carefully as the auth branch (foreign-space path, symlink into a foreign space, `..` traversal, cron fire-time re-check).
8. **The WS handler never sees identity today** (auth happens only in the global hook): the 5.8 wiring is new code on a security-sensitive path; cover unauthenticated, foreign-user, and admin upgrades with tests.
## 15. Open Questions (answer before Phase 3)
1. Should admins' own cases live in `~/codeman-users/<admin>/cases` (symmetric, proposed) or keep using legacy `~/codeman-cases`? Proposed: symmetric; legacy dir is a migration source only.
2. Per-user settings (respawn presets, notification prefs): global-only in v1. Worth a `users/<name>/settings.json` overlay later?
3. Should regular users be allowed to create Docker cases on admin-defined hosts (proposed: yes) or is Docker entirely admin-only?
4. Session handoff: does an admin need "reassign session/case to another user"? (Cheap to add next to `cases/assign`; not in v1 scope.)
5. Permission-mode grants (section 6.3): one `canBypassPermissions` flag covering Claude/Codex/Gemini bypass equivalents PLUS shell mode and cron `launchCommand` (proposed: one flag, keep it one-bit), or split into `canBypassPermissions` + `canRunArbitraryCommands`? And should admins be able to set a per-user DEFAULT mode (for example force `normal` for an intern) rather than just gating bypass?
6. OpenCode has no single bypass flag (its permission config rides `OPENCODE_CONFIG_CONTENT`): decide what the grant means there before Phase 3, or exclude OpenCode mode for non-granted users in v1.
File diff suppressed because it is too large Load Diff
-367
View File
@@ -1,367 +0,0 @@
# Orchestrator Loop — Architecture & Data Flow
> Technical architecture document. Not for GitHub.
## System Overview
```
┌─────────────────────────────────────────────────────────────────────┐
│ CODEMAN WEB UI │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Orchestrator Dashboard │ │
│ │ [Goal Input] [Plan View] [Phase Progress] [Agent Activity] │ │
│ └───────────────────────────┬──────────────────────────────────┘ │
│ │ SSE Events │
│ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Orchestrator API Routes (/api/orchestrator/*) │ │
│ └───────────────────────────┬──────────────────────────────────┘ │
└───────────────────────────────┼─────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────────────┐
│ ORCHESTRATOR LOOP │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────────────┐ │
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
│ │ Planner │ │ Loop (state │ │ Verifier │ │
│ │ │ │ machine) │ │ │ │
│ │ • Research │◄──►│ • Phase mgmt │◄──►│ • Test runner │ │
│ │ • Plan gen │ │ • Task queue │ │ • AI review │ │
│ │ • Phasing │ │ • Event loop │ │ • Output checks │ │
│ └──────┬───────┘ └──────┬───────┘ └──────────┬───────────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ EXISTING CODEMAN INFRASTRUCTURE │ │
│ │ │ │
│ │ SessionManager ←→ Sessions ←→ PTY (Claude CLI) │ │
│ │ ↑ ↑ ↑ │ │
│ │ │ │ │ │ │
│ │ TaskQueue RalphTracker RespawnController │ │
│ │ StateStore HooksConfig TeamWatcher │ │
│ │ Auto-Ops SubagentWatcher SSE Broadcast │ │
│ └──────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
```
## Data Flow: Complete Lifecycle
### 1. User Submits Goal
```
User → POST /api/orchestrator/start { goal: "Build a REST API...", config: {...} }
→ OrchestratorLoop.start(goal)
→ state = PLANNING
→ emit('stateChanged', 'planning')
→ SSE: orchestrator:stateChanged
```
### 2. Planning Phase
```
OrchestratorPlanner.generatePlan(goal)
→ PlanOrchestrator.generateDetailedPlan(goal)
→ [Research Agent] → enriched task description
→ [Planner Agent] → PlanItem[]
→ groupIntoPhases(planItems)
→ topological sort by dependencies
→ group into layers
→ assign team strategies
→ OrchestratorPlan { phases: [...] }
→ state = APPROVAL
→ emit('planReady', plan)
→ SSE: orchestrator:planReady
```
### 3. User Approves Plan
```
User → POST /api/orchestrator/approve
→ OrchestratorLoop.approvePlan()
→ state = EXECUTING
→ executePhase(phases[0])
```
### 4. Phase Execution
```
executePhase(phase)
→ For each task in phase:
→ Convert to CreateTaskOptions
→ Add to TaskQueue with completion phrase "PHASE_{N}_TASK_{M}_DONE"
→ If phase.teamStrategy.type === 'team':
→ Start session with AGENT_TEAMS enabled
→ Send team orchestration prompt to lead
→ Else:
→ Assign tasks to available sessions (same as RalphLoop)
→ Listen for task completion events:
→ TaskQueue emits taskCompleted
→ Check: all phase tasks done?
→ Yes → state = VERIFYING → verifyPhase(phase)
→ No → wait for more completions
```
### 5. Verification
```
verifyPhase(phase)
→ OrchestratorVerifier.verify(phase, session)
→ Run test commands via session
→ Check file existence
→ AI review (optional)
→ If passed:
→ phase.status = 'passed'
→ emit('phaseCompleted', phase)
→ If more phases: executePhase(nextPhase)
→ If last phase: state = COMPLETED
→ If failed:
→ phase.attempts++
→ If attempts < maxAttempts:
→ state = REPLANNING
→ Generate recovery tasks
→ state = EXECUTING (retry)
→ Else:
→ state = FAILED
→ emit('phaseFailed', phase, reason)
```
### 6. Context Management Between Phases
```
After phase completion:
→ If config.compactBetweenPhases:
→ session.sendInput('/compact')
→ Wait for compact to complete
→ If config.respawnBetweenMilestones && phase is a milestone:
→ Save orchestrator state to StateStore
→ Respawn session (kill + recreate)
→ Send resume prompt with phase context
```
## File Layout
```
src/
├── orchestrator-loop.ts # Main state machine (~400 lines)
├── orchestrator-planner.ts # Plan generation + phase grouping (~300 lines)
├── orchestrator-verifier.ts # Phase verification (~200 lines)
├── types/
│ └── orchestrator.ts # All orchestrator types (~150 lines)
├── prompts/
│ └── orchestrator.ts # Prompt templates (~200 lines)
├── web/
│ ├── routes/
│ │ └── orchestrator-routes.ts # API endpoints (~250 lines)
│ └── public/
│ └── orchestrator-ui.js # Frontend panel (~500 lines)
```
## Integration Points with Existing Code
### StateStore (`src/state-store.ts`)
```typescript
// Add to AppState interface
orchestrator?: OrchestratorPersistState;
// Add methods
getOrchestratorState(): OrchestratorPersistState;
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
```
### SSE Events (`src/web/sse-events.ts`)
```typescript
// Add ~8 new events
export const SseEvent = {
// ... existing
ORCHESTRATOR_STATE_CHANGED: 'orchestrator:stateChanged',
ORCHESTRATOR_PLAN_READY: 'orchestrator:planReady',
ORCHESTRATOR_PHASE_STARTED: 'orchestrator:phaseStarted',
ORCHESTRATOR_PHASE_COMPLETED: 'orchestrator:phaseCompleted',
ORCHESTRATOR_PHASE_FAILED: 'orchestrator:phaseFailed',
ORCHESTRATOR_VERIFICATION: 'orchestrator:verificationResult',
ORCHESTRATOR_COMPLETED: 'orchestrator:completed',
ORCHESTRATOR_ERROR: 'orchestrator:error',
} as const;
```
### Frontend Constants (`src/web/public/constants.js`)
```javascript
// Mirror SSE events
SSE_EVENTS.ORCHESTRATOR_STATE_CHANGED = 'orchestrator:stateChanged';
// ... etc
```
### Route Registration (`src/web/routes/index.ts`)
```typescript
import { registerOrchestratorRoutes } from './orchestrator-routes.js';
// Add to barrel export
```
### Server (`src/web/server.ts`)
```typescript
// Initialize OrchestratorLoop alongside RalphLoop
const orchestratorLoop = new OrchestratorLoop(config);
// Register routes
registerOrchestratorRoutes(app, { ...ctx, orchestrator: orchestratorLoop });
```
### Port Interface (`src/web/ports/`)
```typescript
// New port
export interface OrchestratorPort {
orchestrator: OrchestratorLoop;
}
```
## Prompt Flow Through System
The key insight is how prompts flow from Orchestrator → Session → Claude:
```
OrchestratorLoop decides to execute Phase 3, Task 2
│
▼
Converts OrchestratorTask to CreateTaskOptions:
{
prompt: "Implement the rate limiter middleware. Read src/middleware/auth.ts
for the pattern. Add to src/middleware/rate-limiter.ts. Must export
a Fastify plugin. When done: <promise>PHASE_3_TASK_2_DONE</promise>",
priority: 100,
dependencies: ["phase-3-task-1"], // Must finish auth middleware first
completionPhrase: "PHASE_3_TASK_2_DONE",
timeoutMs: 600000 // 10 minutes
}
│
▼
TaskQueue.addTask(options)
│
▼
RalphLoop.tick() → assignTasks() // OR OrchestratorLoop does its own assignment
│
▼
session.sendInput(task.prompt)
│
▼
writeViaMux() → tmux send-keys -l "prompt..." + Enter
│
▼
Claude CLI receives prompt, executes, outputs results
│
▼
RalphTracker.processData() → detects "PHASE_3_TASK_2_DONE"
│
▼
emit('completionDetected') → OrchestratorLoop.handleTaskCompleted()
│
▼
Check: all tasks in Phase 3 done? → If yes → verifyPhase(phase3)
```
## Team Agent Flow (When Enabled)
```
Phase has teamStrategy.type === 'team'
│
▼
OrchestratorLoop creates/reuses a session with:
env: { CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS: '1' }
│
▼
Sends team orchestration prompt:
"You're the team lead for Phase 3: Core Implementation.
Your team should work on these tasks in parallel:
1. Rate limiter middleware (teammate 1)
2. Error handling middleware (teammate 2)
3. Validation layer (teammate 3)
Context files to read first: [...]
Each teammate should output their task's completion phrase when done.
When ALL tasks are complete, output: <promise>PHASE_3_COMPLETE</promise>"
│
▼
Claude Code team-lead spawns teammates
│
▼
TeamWatcher detects new team in ~/.claude/teams/
→ Matches to session via leadSessionId
→ Tracks teammate activity
│
▼
Teammates work in parallel (in-process threads)
│
▼
hook: teammate_idle → POST /api/hook-event
→ OrchestratorLoop notes teammate finished
│
▼
hook: task_completed → POST /api/hook-event
→ Or: RalphTracker detects PHASE_3_COMPLETE
→ OrchestratorLoop → phase complete → verify
```
## Error Recovery Strategy
```
Task fails (timeout, error, session crash)
│
├─ Task-level retry (up to 2 retries per task)
│ → Reset task to pending
│ → Re-queue with modified prompt: "Previous attempt failed: {error}. Try again..."
│
├─ Phase-level retry (up to 3 retries per phase)
│ → Respawn session (fresh context)
│ → Re-execute entire phase with learnings from failure
│ → Modified prompt includes what went wrong
│
└─ Orchestration-level failure
→ All retries exhausted
→ state = FAILED
→ Notify user with detailed failure report
→ User can: modify plan → retry, skip phase → continue, or stop
```
## Interaction with Ralph Loop
Ralph Loop and Orchestrator Loop are **mutually exclusive** on the same sessions:
```
if (orchestratorLoop.isRunning()) {
// Orchestrator controls task assignment
// Ralph Loop should not interfere
// Respawn Controller uses 'orchestrator' preset
}
if (ralphLoop.isRunning()) {
// Ralph controls task assignment
// Orchestrator should not start
}
```
The Orchestrator can optionally USE the Ralph Loop internally for phase execution (delegate phase tasks to Ralph's queue), or manage task assignment directly. Decision: **manage directly** — gives more control over phase boundaries and verification timing.
## Summary of What Touches What
| Existing File | Change |
|---|---|
| `src/types/index.ts` | Export orchestrator types |
| `src/state-store.ts` | Add orchestrator state persistence |
| `src/web/sse-events.ts` | Add ~8 orchestrator events |
| `src/web/routes/index.ts` | Register orchestrator routes |
| `src/web/server.ts` | Initialize OrchestratorLoop |
| `src/web/public/constants.js` | Mirror SSE events |
| `src/web/public/app.js` | Add orchestrator event listeners, panel toggle |
| `src/web/route-helpers.ts` | Add 'orchestrator' respawn preset |
| New File | Purpose |
|---|---|
| `src/orchestrator-loop.ts` | Core state machine |
| `src/orchestrator-planner.ts` | Plan generation + phasing |
| `src/orchestrator-verifier.ts` | Phase verification |
| `src/types/orchestrator.ts` | Type definitions |
| `src/prompts/orchestrator.ts` | Prompt templates |
| `src/web/routes/orchestrator-routes.ts` | API endpoints |
| `src/web/public/orchestrator-ui.js` | Frontend panel |
| `src/web/ports/orchestrator-port.ts` | Port interface |
-633
View File
@@ -1,633 +0,0 @@
# Orchestrator Loop — Detailed Implementation Plan (v2)
> Internal research/planning document. Not for GitHub.
## Vision
The **Orchestrator Loop** is a new autonomous execution mode that transforms high-level user goals into phased, verified, team-coordinated implementations. Unlike Ralph Loop (flat task queue → idle sessions), the Orchestrator manages the full lifecycle: **plan → approve → execute → verify → adapt → complete**.
```
USER: "Add OAuth2 login with Google/GitHub, role-based access control, and API key management"
ORCHESTRATOR:
Phase 1: Research & Setup ✅ (3m) — scaffold, deps, config
Phase 2: Auth Core ✅ (8m) — OAuth2 flow, session mgmt
Phase 3: Provider Integration 🔄 (12m) — Google + GitHub (parallel via team agents)
Phase 4: RBAC ⏳ — roles, permissions, middleware
Phase 5: API Keys ⏳ — generation, validation, rate limits
Phase 6: Testing & Review ⏳ — integration tests, security review
Progress: ━━━━━━━━━━━━━━━━━━━━ 40% | Agents: 3 active | Time: 23m
```
## Architecture
```
┌─────────────────────────────────────────────────────────────────┐
│ OrchestratorLoop │
│ │
│ ┌────────────────┐ ┌────────────────┐ ┌──────────────────┐ │
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
│ │ Planner │ │ Executor │ │ Verifier │ │
│ │ │ │ │ │ │ │
│ │ PlanOrchestrator│ │ TaskQueue │ │ AI review │ │
│ │ + phase grouper│ │ SessionManager │ │ Test commands │ │
│ │ + team strategy│ │ Team prompts │ │ File checks │ │
│ └───────┬────────┘ └───────┬────────┘ └─────────┬────────┘ │
│ │ │ │ │
│ └───────────────────┼──────────────────────┘ │
│ │ │
│ ┌─────────▼─────────┐ │
│ │ Existing Codeman │ │
│ │ Infrastructure │ │
│ │ │ │
│ │ SessionManager │ │
│ │ TaskQueue │ │
│ │ RespawnController │ │
│ │ TeamWatcher │ │
│ │ PlanOrchestrator │ │
│ │ StateStore │ │
│ │ Hooks + SSE │ │
│ └────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
```
## State Machine
```
┌─────────┐
│ IDLE │
└────┬────┘
│ start(goal)
▼
┌─────────┐
┌────────│PLANNING │────────┐
│ fail └────┬────┘ │
▼ │ plan ready │ user cancels
┌────────┐ ▼ ▼
│ FAILED │ ┌─────────┐ ┌────────┐
└────────┘ │APPROVAL │ │ IDLE │
▲ └────┬────┘ └────────┘
│ │ approve
│ ▼
│ ┌──────────┐
│ ┌───►│EXECUTING │◄────────────────────┐
│ │ └────┬─────┘ │
│ │ │ all tasks in phase done │
│ │ ▼ │
│ │ ┌──────────┐ │
│ │ │VERIFYING │ │
│ │ └────┬─────┘ │
│ │ pass │ │ fail │
│ │ ▼ ▼ │
│ │ more ┌──────────┐ │
│ │ phases?│REPLANNING│── retry ────────┘
│ │ │ └────┬─────┘
│ │ │ │ max retries
│ │ │ ▼
│ │ │ ┌────────┐
│ └────┘ │ FAILED │
│ next └────────┘
│ phase
│ │
│ ▼
│ ┌───────────┐
└─│ COMPLETED │
└───────────┘
```
**States:** `idle` | `planning` | `approval` | `executing` | `verifying` | `replanning` | `completed` | `failed` | `paused`
Transitions are event-driven. The state machine is the single source of truth — all methods check `this.state` before acting.
## Type Definitions
### `src/types/orchestrator.ts`
```typescript
// ═══════════════════════════════════════════════════════════════
// State Machine
// ═══════════════════════════════════════════════════════════════
export type OrchestratorState =
| 'idle'
| 'planning'
| 'approval'
| 'executing'
| 'verifying'
| 'replanning'
| 'completed'
| 'failed'
| 'paused';
// ═══════════════════════════════════════════════════════════════
// Plan Structure
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorPlan {
id: string;
goal: string;
createdAt: number;
phases: OrchestratorPhase[];
metadata: {
totalTasks: number;
estimatedComplexity: 'low' | 'medium' | 'high';
modelUsed: string;
planDurationMs: number;
};
}
export interface OrchestratorPhase {
id: string; // "phase-1", "phase-2"
name: string; // Human-readable name
description: string;
order: number;
status: PhaseStatus;
tasks: OrchestratorTask[];
verificationCriteria: string[];
testCommands: string[];
maxAttempts: number; // Default: 3
attempts: number; // Current attempt count
startedAt: number | null;
completedAt: number | null;
durationMs: number | null;
teamStrategy: TeamStrategy;
}
export type PhaseStatus =
| 'pending'
| 'executing'
| 'verifying'
| 'passed'
| 'failed'
| 'skipped';
export interface OrchestratorTask {
id: string; // "phase-1-task-1"
phaseId: string;
prompt: string; // Single-line prompt for Claude
status: 'pending' | 'running' | 'completed' | 'failed';
assignedSessionId: string | null;
queueTaskId: string | null; // Links to TaskQueue task
parallel: boolean; // Can run in parallel with sibling tasks
completionPhrase: string; // Unique phrase for completion detection
timeoutMs: number;
startedAt: number | null;
completedAt: number | null;
error: string | null;
retries: number;
}
// ═══════════════════════════════════════════════════════════════
// Team Strategy
// ═══════════════════════════════════════════════════════════════
export type TeamStrategy =
| { type: 'single' } // One session handles all
| { type: 'parallel'; maxSessions: number } // Multiple sessions
| { type: 'team'; config: TeamSetup } // Agent teams
export interface TeamSetup {
leadPrompt: string;
suggestedTeammates: string[]; // Role descriptions
maxTeammates: number;
}
// ═══════════════════════════════════════════════════════════════
// Verification
// ═══════════════════════════════════════════════════════════════
export interface VerificationResult {
passed: boolean;
checks: VerificationCheck[];
summary: string;
suggestions: string[]; // Recovery hints for replanning
}
export interface VerificationCheck {
type: 'test_command' | 'ai_review' | 'file_check';
description: string;
passed: boolean;
output?: string;
}
// ═══════════════════════════════════════════════════════════════
// Configuration
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorConfig {
plannerModel: string; // Default: 'opus'
researchEnabled: boolean; // Default: true
autoApprove: boolean; // Default: false
maxPhaseRetries: number; // Default: 3
phaseTimeoutMs: number; // Default: 1800000 (30min)
enableTeamAgents: boolean; // Default: true
maxParallelSessions: number; // Default: 3
verificationMode: 'strict' | 'moderate' | 'lenient';
compactBetweenPhases: boolean; // Default: true
}
// ═══════════════════════════════════════════════════════════════
// Persistence (saved to ~/.codeman/state.json)
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorPersistState {
state: OrchestratorState;
plan: OrchestratorPlan | null;
currentPhaseIndex: number;
startedAt: number | null;
completedAt: number | null;
config: OrchestratorConfig;
stats: OrchestratorStats;
}
export interface OrchestratorStats {
phasesCompleted: number;
phasesFailed: number;
totalTasksCompleted: number;
totalTasksFailed: number;
totalDurationMs: number;
replanCount: number;
}
```
## New Files (Implementation Order)
### Step 1: `src/types/orchestrator.ts` — Type definitions
All interfaces above. No dependencies. ~120 lines.
### Step 2: `src/orchestrator-planner.ts` — Plan generation + phase grouping
~300 lines. Wraps existing PlanOrchestrator.
```typescript
/**
* @fileoverview Orchestrator plan generation — converts goals into phased plans.
*
* Uses PlanOrchestrator for AI plan generation, then groups PlanItems into
* sequential phases with team strategies and verification criteria.
*
* @module orchestrator-planner
*/
export class OrchestratorPlanner {
constructor(mux: TerminalMultiplexer, workingDir: string, config: OrchestratorConfig);
/** Generate plan from goal. Uses PlanOrchestrator internally. */
async generatePlan(goal: string, onProgress?: ProgressCallback): Promise<OrchestratorPlan>;
/** Cancel in-progress plan generation. */
async cancel(): Promise<void>;
// Internal
private groupIntoPhases(items: PlanItem[], goal: string): OrchestratorPhase[];
private assignTeamStrategies(phases: OrchestratorPhase[]): void;
private generateCompletionPhrases(plan: OrchestratorPlan): void;
}
```
**Phase grouping algorithm:**
1. Topological sort by `PlanItem.dependencies`
2. Group into dependency layers (Kahn's algorithm)
3. Within each layer, sub-group by `tddPhase` (setup → test → impl → verify → review)
4. Merge adjacent small phases (< 2 tasks) if they share the same tddPhase
5. Assign team strategies:
- 1-2 tasks → `{ type: 'single' }`
- 3+ independent tasks → `{ type: 'parallel', maxSessions: Math.min(taskCount, config.maxParallelSessions) }`
- 4+ tasks with high complexity → `{ type: 'team', config: { ... } }`
6. Generate unique completion phrases per task: `ORCH_P{phaseOrder}_T{taskIndex}`
### Step 3: `src/orchestrator-verifier.ts` — Phase verification
~200 lines.
```typescript
/**
* @fileoverview Orchestrator phase verification.
*
* Runs verification checks after each phase completes:
* test commands, AI review, and file existence checks.
*
* @module orchestrator-verifier
*/
export class OrchestratorVerifier {
constructor(config: OrchestratorConfig);
/** Run all verification checks for a completed phase. */
async verifyPhase(
phase: OrchestratorPhase,
session: Session,
mode: 'strict' | 'moderate' | 'lenient'
): Promise<VerificationResult>;
// Verification strategies
private async runTestCommands(commands: string[], session: Session): Promise<VerificationCheck[]>;
private async aiReview(phase: OrchestratorPhase, session: Session): Promise<VerificationCheck>;
}
```
**Verification modes:**
- `strict`: ALL test commands must pass AND AI review must approve
- `moderate`: Test commands must pass, AI review is advisory
- `lenient`: At least one test command passes, AI review skipped
**AI review prompt (sent as a task to the session):**
```
Review Phase "{phase.name}" completion. Check:
1. Expected functionality works
2. No obvious regressions
3. Code quality is acceptable
Criteria: {phase.verificationCriteria.join('\n')}
If ALL criteria are met, respond: ORCH_VERIFY_PASS
If ANY criteria fail, respond: ORCH_VERIFY_FAIL and explain what failed.
```
### Step 4: `src/orchestrator-loop.ts` — Core state machine
~500 lines. Main orchestrator engine.
```typescript
/**
* @fileoverview Orchestrator Loop — phased plan execution with team agents.
*
* State machine that generates plans from user goals, executes them
* phase-by-phase with verification gates, and adapts on failure.
*
* @module orchestrator-loop
*/
export interface OrchestratorLoopEvents {
stateChanged: (state: OrchestratorState, prevState: OrchestratorState) => void;
planReady: (plan: OrchestratorPlan) => void;
phaseStarted: (phase: OrchestratorPhase) => void;
phaseCompleted: (phase: OrchestratorPhase) => void;
phaseFailed: (phase: OrchestratorPhase, reason: string) => void;
taskAssigned: (task: OrchestratorTask, sessionId: string) => void;
taskCompleted: (task: OrchestratorTask) => void;
taskFailed: (task: OrchestratorTask, error: string) => void;
verificationResult: (phase: OrchestratorPhase, result: VerificationResult) => void;
completed: (stats: OrchestratorStats) => void;
error: (error: Error) => void;
}
export class OrchestratorLoop extends EventEmitter {
private state: OrchestratorState = 'idle';
private plan: OrchestratorPlan | null = null;
private currentPhaseIndex = 0;
private config: OrchestratorConfig;
private planner: OrchestratorPlanner;
private verifier: OrchestratorVerifier;
private sessionManager: SessionManager;
private taskQueue: TaskQueue;
private store: StateStore;
private stats: OrchestratorStats;
private cleanup: CleanupManager;
private pausedState: OrchestratorState | null = null; // State before pause
// ── Lifecycle ──────────────────────────────────────────────
constructor(mux: TerminalMultiplexer, workingDir: string, config?: Partial<OrchestratorConfig>);
/** Start orchestration with a goal. Transitions: idle → planning */
async start(goal: string): Promise<void>;
/** Approve the generated plan. Transitions: approval → executing */
async approve(): Promise<void>;
/** Reject plan with feedback. Transitions: approval → planning (regenerate) */
async reject(feedback: string): Promise<void>;
/** Pause execution. Saves current state. */
pause(): void;
/** Resume from pause. */
resume(): void;
/** Stop everything and clean up. → idle */
async stop(): Promise<void>;
/** Skip current phase. → executing (next phase) or completed */
async skipPhase(phaseId: string): Promise<void>;
/** Retry a failed phase. → executing */
async retryPhase(phaseId: string): Promise<void>;
// ── Getters ────────────────────────────────────────────────
getState(): OrchestratorState;
getPlan(): OrchestratorPlan | null;
getCurrentPhase(): OrchestratorPhase | null;
getStats(): OrchestratorStats;
getStatus(): OrchestratorPersistState;
// ── Internal: Phase Execution ──────────────────────────────
private async executeCurrentPhase(): Promise<void>;
private async executePhase(phase: OrchestratorPhase): Promise<void>;
private async assignPhaseTasks(phase: OrchestratorPhase): Promise<void>;
private handleTaskCompleted(taskId: string): void;
private handleTaskFailed(taskId: string, error: string): void;
private async onPhaseTasksComplete(phase: OrchestratorPhase): Promise<void>;
// ── Internal: Verification ─────────────────────────────────
private async verifyCurrentPhase(): Promise<void>;
private async handleVerificationResult(phase: OrchestratorPhase, result: VerificationResult): Promise<void>;
// ── Internal: Replanning ───────────────────────────────────
private async replanPhase(phase: OrchestratorPhase, failures: string[]): Promise<void>;
// ── Internal: State Machine ────────────────────────────────
private setState(newState: OrchestratorState): void;
private advanceToNextPhase(): Promise<void>;
private persist(): void;
private restore(): void;
}
```
**Key execution flow in `executePhase()`:**
1. Mark phase as `executing`, emit `phaseStarted`
2. For each task in phase:
- Create a `CreateTaskOptions` from `OrchestratorTask`
- Add to `TaskQueue` with proper dependencies + completion phrase
- Store the TaskQueue task ID in `OrchestratorTask.queueTaskId`
3. Poll task completion (listen to TaskQueue events)
4. When all tasks complete → call `onPhaseTasksComplete()`
5. `onPhaseTasksComplete()` triggers verification
**How tasks get assigned to sessions:**
The OrchestratorLoop does NOT manage session assignment directly. It adds tasks to the existing TaskQueue and starts a mini poll loop that assigns pending tasks to idle sessions — the same pattern as RalphLoop's `assignTasks()`. This reuses existing session management.
**Team agent flow:**
For phases with `teamStrategy.type === 'team'`:
- Start a single session with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
- Instead of adding individual tasks to TaskQueue, send ONE comprehensive prompt to the lead
- The prompt instructs the lead to create teammates and delegate
- Monitor via TeamWatcher for team task completion + hook events
- Phase completion is detected via the lead's completion phrase
### Step 5: `src/web/routes/orchestrator-routes.ts` — API endpoints
~300 lines.
```
POST /api/orchestrator/start — { goal, config? } → start planning
POST /api/orchestrator/approve — approve generated plan
POST /api/orchestrator/reject — { feedback } → reject + replan
POST /api/orchestrator/pause — pause execution
POST /api/orchestrator/resume — resume execution
POST /api/orchestrator/stop — stop orchestration
GET /api/orchestrator/status — full state + plan + stats
GET /api/orchestrator/plan — plan details only
POST /api/orchestrator/phase/:id/skip — skip a phase
POST /api/orchestrator/phase/:id/retry — retry a failed phase
```
Port dependency: `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort`
The route module receives the OrchestratorLoop instance via the InfraPort (added to `createRouteContext()`).
### Step 6: SSE Events — `src/web/sse-events.ts` additions
```typescript
// ─── Orchestrator ────────────────────────────────────────────────────────────
/** Orchestrator state machine transitioned. */
export const OrchestratorStateChanged = 'orchestrator:stateChanged' as const;
/** Orchestrator plan generated and ready for approval. */
export const OrchestratorPlanReady = 'orchestrator:planReady' as const;
/** Orchestrator phase started executing. */
export const OrchestratorPhaseStarted = 'orchestrator:phaseStarted' as const;
/** Orchestrator phase completed successfully. */
export const OrchestratorPhaseCompleted = 'orchestrator:phaseCompleted' as const;
/** Orchestrator phase failed. */
export const OrchestratorPhaseFailed = 'orchestrator:phaseFailed' as const;
/** Orchestrator verification result for a phase. */
export const OrchestratorVerification = 'orchestrator:verification' as const;
/** Orchestrator task assigned to session. */
export const OrchestratorTaskAssigned = 'orchestrator:taskAssigned' as const;
/** Orchestrator task completed. */
export const OrchestratorTaskCompleted = 'orchestrator:taskCompleted' as const;
/** Orchestrator task failed. */
export const OrchestratorTaskFailed = 'orchestrator:taskFailed' as const;
/** All phases completed successfully. */
export const OrchestratorCompleted = 'orchestrator:completed' as const;
/** Orchestrator error. */
export const OrchestratorError = 'orchestrator:error' as const;
```
11 new events. Add to `SseEvent` namespace object + mirror in `constants.js`.
### Step 7: State persistence — `src/state-store.ts` additions
Add to `AppState`:
```typescript
orchestrator?: OrchestratorPersistState;
```
Add methods:
```typescript
getOrchestratorState(): OrchestratorPersistState | null;
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
clearOrchestratorState(): void;
```
### Step 8: Server integration — `src/web/server.ts` modifications
1. Import `OrchestratorLoop` and `registerOrchestratorRoutes`
2. Add `private orchestratorLoop: OrchestratorLoop` field
3. Initialize in constructor (lazy — created on first start, not at boot)
4. Add to `createRouteContext()` InfraPort: `orchestratorLoop: this.orchestratorLoop`
5. Wire up OrchestratorLoop events → SSE broadcasts
6. Register routes: `registerOrchestratorRoutes(this.app, ctx)`
7. Clean up in `stop()`
### Step 9: `src/web/public/orchestrator-ui.js` — Frontend panel
~500 lines. New frontend module.
**Load order**: After `panels-ui.js` (11), before `ralph-wizard.js` (13). So load order = 11.5.
**UI elements:**
- Goal input form (text area + config toggles)
- Plan approval view (phase list, task details, approve/reject buttons)
- Execution dashboard (progress bar, phase cards, task status indicators)
- Agent activity panel (session count, team status)
- Controls (pause, resume, stop, skip phase, retry phase)
**SSE listeners:**
- All 11 orchestrator events → update UI state
- Reuses existing session/respawn/team event handlers for agent monitoring
### Step 10: `src/prompts/orchestrator.ts` — Prompt templates
~200 lines.
Templates for:
- Phase execution prompt (tells Claude what to do in this phase)
- Team lead delegation prompt (instructs lead to create and coordinate teammates)
- Verification prompt (asks Claude to verify phase output)
- Replan prompt (gives failure context, asks for recovery steps)
### Step 11: Constants, schemas, route barrel updates
- `src/web/public/constants.js` — Add 11 SSE event mirrors
- `src/web/schemas.ts` — Add Zod schemas for orchestrator API input validation
- `src/web/routes/index.ts` — Export `registerOrchestratorRoutes`
- `src/web/ports/infra-port.ts` — Add `orchestratorLoop` to InfraPort
- `src/types/index.ts` — Export orchestrator types
## Existing File Modifications Summary
| File | Change | Lines |
|------|--------|-------|
| `src/types/index.ts` | Add orchestrator barrel export | +1 |
| `src/web/sse-events.ts` | Add 11 orchestrator events + SseEvent entries | +30 |
| `src/web/public/constants.js` | Mirror 11 SSE events | +15 |
| `src/web/routes/index.ts` | Export registerOrchestratorRoutes | +1 |
| `src/web/ports/infra-port.ts` | Add orchestratorLoop to InfraPort | +3 |
| `src/web/server.ts` | Initialize OrchestratorLoop, wire events, register routes | +40 |
| `src/web/schemas.ts` | Add orchestrator Zod schemas | +20 |
| `src/state-store.ts` | Add orchestrator state persistence | +20 |
| `src/web/public/app.js` | Add orchestrator SSE listeners + panel toggle | +30 |
| `src/web/public/index.html` | Add orchestrator-ui.js script tag | +1 |
**Total new code**: ~2,300 lines across 6 new files
**Total modifications**: ~160 lines across 10 existing files
## Implementation Execution Order
This is the actual build order — each step is a commit checkpoint:
1. **Types** — `src/types/orchestrator.ts` + barrel export. Zero risk, pure types.
2. **SSE events** — Add all 11 events to both `sse-events.ts` and `constants.js`. Wire in SseEvent namespace.
3. **State persistence** — Add orchestrator state to StateStore. Small, isolated change.
4. **Schemas** — Add Zod validation schemas for API input.
5. **Planner** — `src/orchestrator-planner.ts`. Can test in isolation.
6. **Verifier** — `src/orchestrator-verifier.ts`. Can test in isolation.
7. **Core loop** — `src/orchestrator-loop.ts`. The big one. Depends on planner + verifier.
8. **Prompts** — `src/prompts/orchestrator.ts`. Templates used by core loop.
9. **Port + routes** — `src/web/ports/infra-port.ts` update + `src/web/routes/orchestrator-routes.ts`.
10. **Server integration** — Wire OrchestratorLoop into WebServer. Routes become live.
11. **Frontend** — `src/web/public/orchestrator-ui.js` + app.js listeners + index.html script tag.
12. **Tests** — `test/orchestrator-*.test.ts`.
13. **Typecheck + lint** — Fix all issues, ensure CI passes.
## Edge Cases & Error Handling
- **Session limit reached**: Queue tasks and wait for sessions to free up (existing SessionManager handles this)
- **All sessions crash during phase**: Mark phase as failed, attempt replan
- **Verification flaky**: `moderate` mode allows test retries; `lenient` skips AI review
- **Plan too large**: Cap at 10 phases, 50 total tasks. Warn user.
- **Context overflow**: Auto-compact between phases. Respawn if needed (orchestrator state is external).
- **User pauses mid-phase**: Pause task assignment, don't cancel running tasks. Resume picks up where it left off.
- **Network/API errors during planning**: Retry plan generation up to 2 times, then fail with clear message.
- **Orchestrator vs Ralph conflict**: Mutually exclusive. Starting orchestrator stops Ralph if running. Starting Ralph stops orchestrator.
## Testing Strategy
- **Unit tests**: `test/orchestrator-planner.test.ts` — phase grouping algorithm, team strategy assignment
- **Unit tests**: `test/orchestrator-verifier.test.ts` — verification logic with mocked sessions
- **Integration tests**: `test/orchestrator-loop.test.ts` — state machine transitions, task lifecycle
- **Route tests**: `test/routes/orchestrator-routes.test.ts` — API validation, status responses
All tests use `MockSession` pattern from existing test infrastructure. No real tmux needed.
-157
View File
@@ -1,157 +0,0 @@
# Orchestrator Loop — Research Findings
> Research doc for the new "Orchestrator Loop" feature. Not for GitHub.
## What We're Building
A new autonomous loop variant — **Orchestrator Loop** — that takes high-level user tasks, decomposes them into a detailed plan using team agents, and executes the plan step-by-step with quality gates. Unlike Ralph Loop (which executes a flat task queue), the Orchestrator coordinates **planning, delegation, and verification** as a continuous cycle.
**Core idea**: User inputs a goal → Orchestrator creates a detailed plan → spins up team agents for parallel execution → validates each step → adapts the plan based on results → delivers polished output.
## Existing Infrastructure Analysis
### What We Can Reuse
#### 1. Ralph Loop (`src/ralph-loop.ts`)
- **Pattern**: Poll loop with `start() → tick() → stop()` lifecycle
- **Reusable**: Event-driven task assignment, session completion handling, timeout management
- **Limitation**: Flat task queue — no concept of phases, dependencies between task groups, or adaptive replanning
- **Key insight**: `assignTaskToSession()` uses `session.sendInput(task.prompt)` — simple prompt injection into PTY
#### 2. Task Queue (`src/task-queue.ts`) + Task (`src/task.ts`)
- **Already has**: Priority ordering, dependency tracking between tasks, completion phrase detection
- **Limitation**: No task *groups* or *phases*. Dependencies are task-to-task, not phase-to-phase
- **Key insight**: Tasks support `completionPhrase` — a string the task watches for in output. This is how Ralph knows a task is done
#### 3. Plan Orchestrator (`src/plan-orchestrator.ts`)
- **Already has**: 2-agent plan generation (Research Agent → Planner Agent), TDD-aware plan items with P0/P1/P2 priorities
- **Output**: `PlanItem[]` with dependencies, verification criteria, TDD phases, complexity ratings
- **Limitation**: Plan generation only — no execution. Plans are generated then sit in state/UI for human review
- **Key insight**: Uses `Session` directly to run Claude subagent instances for research and planning. Returns structured JSON
#### 4. Team Agents (`src/team-watcher.ts`, `~/.claude/teams/`)
- **Already has**: Team creation, member tracking, filesystem inbox messaging, task management via `~/.claude/tasks/{team-name}/`
- **Limitation**: Codeman can only *observe* teams (TeamWatcher is read-only polling), not *create* or *orchestrate* them
- **Key insight**: Teams are a Claude Code feature. Codeman monitors them but doesn't control them. We can't programmatically create teammates — Claude Code does that when you use `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
#### 5. Respawn Controller (`src/respawn-controller.ts`)
- **Already has**: Preset-based automation (ralph-todo, overnight-autonomous), circuit breaker, health scoring
- **Key insight**: The `ralph-todo` preset (8s idle, 480min max) is designed for autonomous task execution. We'd need a new preset or make Orchestrator Loop set its own timing
#### 6. Session Auto-Ops (`src/session-auto-ops.ts`)
- **Already has**: Auto-compact at token thresholds, auto-clear for context management
- **Key insight**: Critical for long Orchestrator runs — prevents context overflow during multi-step execution
#### 7. Hooks (`src/hooks-config.ts`)
- **Already has**: `idle_prompt`, `stop`, `teammate_idle`, `task_completed` hook events
- **Key insight**: Hooks fire POST to `/api/hook-event` — this is how Codeman knows when Claude is idle, stopped, or completed a task. The Orchestrator Loop can listen to these same events
### What We Need to Build New
1. **Plan → Task decomposition**: Convert PlanOrchestrator output (PlanItem[]) into executable task groups with phase ordering
2. **Multi-phase execution engine**: Execute plan phases sequentially, tasks within phases in parallel
3. **Verification gates**: After each phase, run verification (test commands, AI review) before proceeding
4. **Adaptive replanning**: When a task fails or verification fails, generate a recovery plan
5. **Team agent orchestration**: Leverage Claude Code's agent teams for parallel execution within phases
6. **Progress tracking & UI**: Real-time dashboard showing plan progress, phase status, agent activity
## How Teams Actually Work (Important Constraint)
After deep research, here's the reality of agent teams:
```
User starts session with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
→ Claude Code creates a team-lead
→ Team-lead spawns teammates (in-process threads)
→ Teammates appear as subagents (detected by SubagentWatcher)
→ Communication via ~/.claude/teams/{name}/inboxes/{member}.json
→ Tasks tracked in ~/.claude/tasks/{team-name}/{N}.json
```
**Codeman cannot programmatically create team members.** This is a Claude Code internal feature. However, Codeman CAN:
- Start a session that has teams enabled
- Send a prompt to the lead that instructs it to use agent teams
- Monitor team activity via TeamWatcher
- React to teammate_idle and task_completed hook events
- Read team task status from the filesystem
**This means**: The Orchestrator Loop orchestrates at the *session prompt* level, not the *team member* level. We tell the lead what to do, and the lead decides how to use its team.
## Architecture Decision: Prompt-Level Orchestration
Given the team constraint, the Orchestrator Loop works by:
1. **Planning phase**: Use PlanOrchestrator to generate a detailed plan from user input
2. **Execution phase**: Feed plan steps as prompts to sessions, one phase at a time
3. **Verification phase**: After each phase, run verification prompts and check results
4. **Adaptation phase**: If verification fails, generate recovery prompts
The "team agents" aspect works by:
- Starting sessions with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
- Crafting prompts that *instruct the lead to delegate* to teammates
- Monitoring team activity to track parallel progress
- The lead agent is smart enough to decompose work across its team
## Key Technical Findings
### Session Input Mechanics
```typescript
// From session.ts - how we send prompts
await session.sendInput(task.prompt); // Uses writeViaMux() internally
// writeViaMux() does: tmux send-keys -l "prompt text" + tmux send-keys Enter
// CRITICAL: Single-line only! Multi-line breaks Ink rendering
```
### Completion Detection Chain
```
PTY output → RalphTracker.processData() → completion phrase fuzzy match
→ CompletionConfidence scoring (multi-signal: promise tag + todos + exit signal)
→ If confident → emit 'completionDetected'
→ RalphLoop listens → marks task complete → assigns next
```
### How Plan Items Map to Tasks
```typescript
// PlanItem has:
interface PlanItem {
id: string; // "P0-001"
content: string; // "Implement error handling for API endpoints"
priority: 'P0' | 'P1' | 'P2';
dependencies: string[]; // ["P0-000"] — other PlanItem IDs
verificationCriteria: string;
testCommand: string;
tddPhase: 'setup' | 'test' | 'impl' | 'verify' | 'review';
complexity: 'low' | 'medium' | 'high';
}
// Task has:
interface CreateTaskOptions {
prompt: string;
priority: number;
dependencies: string[]; // Task IDs
completionPhrase: string;
timeoutMs: number;
}
// Natural mapping: PlanItem.content → Task.prompt
// PlanItem.dependencies → Task.dependencies
// PlanItem.priority → Task.priority (P0=100, P1=50, P2=10)
// PlanItem.verificationCriteria → verification task prompt
```
### Context Management for Long Runs
- Auto-compact at ~110k tokens (configurable)
- Auto-clear at ~140k tokens (configurable)
- Respawn cycling: kill + restart session to reset context entirely
- For Orchestrator: we want compact between phases, respawn between major milestones
## Risk Assessment
| Risk | Severity | Mitigation |
|------|----------|------------|
| Context overflow during complex phases | High | Auto-compact between tasks, respawn between phases |
| Team agents not predictable | Medium | Orchestrate at session level, let Claude decide team delegation |
| Plan too ambitious → infinite loop | High | Phase budgets (max attempts per phase), circuit breaker |
| Verification too strict → blocks progress | Medium | Configurable strictness, human override via UI |
| Single-line prompt limit | Medium | Use CLAUDE.md file for complex instructions, prompt references file |
| Long planning phase delays execution | Low | Show plan for approval before execution |
-681
View File
@@ -1,681 +0,0 @@
# Pi (pi.dev) Run Mode: Implementation Plan
Tracking issue: [#206 "Plans to support pi.dev?"](https://github.com/Ark0N/Codeman/issues/206)
Status: **IMPLEMENTED 2026-08-13** (see `docs/pi-integration.md` for the user-facing
guide). Everything below is the design record; the open questions were resolved
empirically against pi 0.84.1 and the answers are recorded inline as **RESULT**
notes. Originally reworked 2026-08-06; **rechecked 2026-08-13 against master @
`f39beb3` (v1.17.0)**, and every line anchor below was re-verified at that commit (the 1.11.2-era
anchors drifted heavily: six releases landed in between, including the settings-surface overhaul and
the codex predictive-echo work, both of which added new pi touchpoints, §2.10 and the Brain picker in
Phase 3). Upstream facts verified against `@earendil-works/pi-coding-agent` **v0.84.1** (npm latest,
published 2026-08-07) and the [`earendil-works/pi`](https://github.com/earendil-works/pi) repo (cite
that name: upstream docs still contain stale `pi-mono` links from a repo rename). Line numbers are
anchors for orientation, not contracts; they drift.
---
## 1. What Pi is
[Pi](https://pi.dev) (MIT) is a minimal, extensible coding-agent harness. Facts below are verified
against the upstream docs in `packages/coding-agent/docs/`.
| Property | Value |
| ---------------- | -------------------------------------------------------------------------------------------------- |
| Binary | `pi` (`bin: { pi: 'dist/cli.js' }`) |
| npm package | `@earendil-works/pi-coding-agent`, latest **0.84.1** (2026-08-07; 0.84.0 was 2026-08-06); `legacy-node20` dist-tag at 0.74.2 |
| Install | `npm install -g --ignore-scripts @earendil-works/pi-coding-agent`, or `curl -fsSL https://pi.dev/install.sh \| sh` (the curl installer also goes through global npm, so both uninstall via npm) |
| Config dir | `~/.pi/agent` (override: `PI_CODING_AGENT_DIR`). Holds `auth.json`, `trust.json`, `settings.json`, `models.json` (user-defined providers), `models-store.json` (cached catalogs), `keybindings.json`, `extensions/`, `skills/`, `prompts/`, `themes/`, `AGENTS.md`, `SYSTEM.md`, and the package trees `npm/` + `git/` |
| Sessions | `~/.pi/agent/sessions/--<cwd with / replaced by ->--/<timestamp>_<uuid>.jsonl`, tree-structured (`id`/`parentId`), format v3. Overrides: `PI_CODING_AGENT_SESSION_DIR`, `--session-dir` |
| Credentials | `~/.pi/agent/auth.json` (OAuth subscriptions + API keys, auto-refresh), plus ~34 provider env vars with **no common prefix**. 0.84.1 adds `pi auth check` (auth preflight with optional credential output) |
| TUI | Default: **main screen with terminal-owned scrollback**. Since **0.84.0** an experimental fullscreen mode exists, selectable via `--tui-mode fullscreen` **or at runtime through `/settings`**; the default remains the main-screen mode |
| Providers | 15+ (Anthropic, OpenAI, Google, Azure, Bedrock, Mistral, Groq, xAI, OpenRouter, Copilot, Baseten since 0.84.0, ...). OAuth subscription login via `/login` for six: ChatGPT Plus/Pro, Claude Pro/Max, GitHub Copilot, xAI, OpenRouter, Radius |
| Permission model | **No permission prompts at all.** No built-in sandbox, no MCP (none planned), no sub-agents, no plan mode, no to-dos, no background bash. Tools run with the user's own permissions |
| Trust model | "Project trust" gates **loading** of project-local `.pi/` config/extensions/skills and **installing missing project packages**, not tool execution. Triggered only when the cwd (or an ancestor) contains `.pi/settings.json`, `.pi/extensions\|skills\|prompts\|themes`, `.pi/SYSTEM.md`/`.pi/APPEND_SYSTEM.md`, or `.agents/skills`; a bare `.pi/` directory does NOT prompt. Global `defaultProjectTrust`: `ask` (default) / `always` / `never` |
Three consequences shape the whole integration:
1. **There is no `--dangerously-skip-permissions` analog and none is needed.** Pi never prompts for
tool approval. The Claude/Codex/Gemini/Antigravity pattern of "send the bypass flag so the session
is not stuck on a modal" does not apply. Codeman must not invent a flag here.
2. **The one privileged knob is `--approve` / `-a`** (trust project-local files for this run), which
makes pi load and execute project `.pi/extensions` TypeScript **and run an npm install of missing
project packages**. That is the field the multi-user clamp has to cover. Its explicit inverse
`-na` / `--no-approve` exists, which lets the clamp force-deny rather than merely omit (§3, §5.2).
3. **Provider keys cannot ride the env allowlist.** Pi's provider key vars (`ANTHROPIC_API_KEY`,
`OPENAI_API_KEY`, `DEEPSEEK_API_KEY`, `HF_TOKEN`, `BASETEN_API_KEY`, ...) share no prefix, so
there is no way to admit them through `ALLOWED_ENV_PREFIXES` without widening the list for every
mode (§2.4).
---
## 2. Design decisions
### 2.1 Mode identity
`SessionMode` gains `'pi'`. Not a location overlay (unlike Docker/remote-SSH cases), not a web tab:
a real sixth CLI backend with its own PTY, tmux session and respawn behaviour, exactly like
`antigravity`. Append `pi` after `antigravity` in every enum/list to keep ordering consistent.
| Surface | Value |
| ---------------- | --------------------------------------------------------------------- |
| `SessionMode` | `'pi'` |
| Display label | `Pi` |
| Tab badge | `pi` (two-letter lowercase, like `sh`/`oc`/`cx`/`gm`/`ag`) |
| Run button label | `Run PI` (short-label ternary in `_applyRunMode`, pattern `Run AG`) |
| Kill-menu label | `Kill Tmux & Pi` |
| Identity color | **`#f472b6` (rose-400)**. Verified free: live computed values on the default skin are claude `#38b6f0`, opencode `#44b993`, codex `#2b8fd9`, gemini `#8ab4f8`, antigravity `#22d3ee`, shell `#98a2b1`, web `#38bdf8`; purple is codex's base hex and amber reads as the shell tab badge, so pink/rose (or orange `#fb923c`) are the only genuinely free hues. No `pi` CSS identifier collides anywhere (`mode-pi`, `.tab-mode.pi`, `.run-mode-dot.pi` all grep clean, re-checked at f39beb3) |
| Env prefix | `PI_` |
| Dependency id | `pi` |
| Status endpoint | `GET /api/pi/status` |
### 2.2 `isExternalCliMode()` yes, `isAltScreenStripMode()` no
Pi joins `isExternalCliMode()` (`session.ts:164-167`): its own TUI, its own output format, so the
Ralph tracker, `BashToolParser`, token/CLI-info scraping and the `❯` readiness probe all stay off
(gates at `session.ts:1100`, `:1701`, `:2000`, `:2103`), and readiness falls back to the output
stabilization used by the other external CLIs.
Pi stays **out** of `isAltScreenStripMode()` (`session.ts:197-199`, currently codex/claude/gemini;
antigravity and opencode are deliberately excluded). Pi's default TUI renders into the main screen
with terminal-owned scrollback, so there is nothing to strip. The fullscreen mode **shipped in
0.84.0 and is runtime-switchable via `/settings`**, so Codeman cannot assume a pi session stays
main-screen for its lifetime; staying out of the strip list is exactly what makes that safe (the alt
screen is load-bearing when the user flips to fullscreen, as it is for `opencode`). Putting pi IN
the strip list would corrupt fullscreen sessions. Three mirrors must stay consistent (all unchanged
for pi, i.e. pi appears in none of them): the replay-side strip in `session-routes.ts:2275`, the
live-stream twin in `session.ts`, and the frontend `_sessionUsesServerMouseStrip()` in
`terminal-ui.js` (usages `:3432`, `:3697`).
### 2.3 tmux required, no direct-PTY fallback, no per-mode configurator
Same rule as the other external CLIs: `pi` mode throws if tmux is unavailable. Add a fourth block to
the guard chain at `session.ts:1751-1768` (antigravity's is `:1765-1768`).
**No `_configurePi()` is needed.** Opencode/codex/gemini each have a tmux-`setenv` configurator
(`tmux-manager.ts:1709-1727`), but antigravity has none: it relies entirely on the generic
`applyEnvOverrides()` (`tmux-manager.ts:1643`, `VALID_KEY = /^[A-Z_][A-Z0-9_]*$/`), which runs for
every mode in both create (`:1880`) and respawn (`:2107`) and injects via socket-scoped
`tmux setenv`, never the spawn command line. Pi follows the antigravity precedent: `PI_*` overrides
flow through `applyEnvOverrides()` and nothing else.
Pi joins the truecolor branches: `buildEnvExports()` (`tmux-manager.ts:1604-1609`,
`export COLORTERM=truecolor` + `unset NO_COLOR` for codex/gemini/antigravity) and the attach-env
condition at `session.ts:1400-1402` (`buildMuxAttachEnv(...)`, whose comment says it must mirror
`buildEnvExports`). Add `|| mode === 'pi'` to both, or the tmux session and the attach client
disagree about color depth.
### 2.4 Env prefix: `PI_` only
Add `'PI_'` to `ALLOWED_ENV_PREFIXES` (`schemas.ts:125`) and to the prose error message at `:163`
(two edits: the message hardcodes the list, and since 1.12+ it also names the exact-key allowlist,
currently `...ANTIGRAVITY_* keys and CLAUDE_CONFIG_DIR are allowed.`; there is now a separate
`ALLOWED_ENV_KEYS` exact-key set alongside the prefix list, which pi does not need to touch). That
covers every documented variable pi reads: `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`,
`PI_PACKAGE_DIR`, `PI_OFFLINE`, `PI_SKIP_VERSION_CHECK`, `PI_TELEMETRY`, `PI_CACHE_RETENTION`,
`PI_SHARE_VIEWER_URL`, `PI_HARDWARE_CURSOR`, `PI_EXPERIMENTAL` (whose meaning 0.84.0 extended to
strict JSON-schema tool sampling). (Pi also *sets* `PI_CODING_AGENT=true` and `AI_AGENT=pi` in child
processes; those are output markers, not inputs, and need nothing from us.)
**Deliberately not added:** `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, `XAI_API_KEY`,
`GROQ_API_KEY`, `MISTRAL_API_KEY` and the other ~28 provider keys. `ALLOWED_ENV_PREFIXES` is a
single global list applied by one Zod refine with no mode context (`safeEnvOverridesSchema`,
`schemas.ts:153-165`), so allowlisting bare provider keys for pi would widen the allowlist for
**every** mode at once, violating the multi-CLI prefix discipline in CLAUDE.md. Users authenticate
pi through `/login` (stored in `~/.pi/agent/auth.json`, auto-refreshed) or by exporting the key in
the Codeman server process's own environment.
Making the allowlist mode-aware is the clean fix, listed as a follow-up in §9. Do not smuggle it
into this change.
### 2.5 Docker credential policy: seed files, not the whole dir
`CRED_STORES` (`docker-hosts.ts:597-605`; file unchanged since the 2026-08-06 verification) gets a
`.pi/agent` entry. Nested `rel` paths already work (`.config/gcloud` maps to seed name
`.config-gcloud` via the `replace(/\//g, '-')` at `:620`). Unlike antigravity, which needed **no**
entry (`agy` nests all state under `~/.gemini/antigravity-cli/`, already covered by the `.gemini`
policy, per the comment at `:599-602`), pi has its own top-level dir and needs its own entry. Use
`seedFiles`, **not** `seedWhole`:
```ts
{ rel: '.pi/agent', seedFiles: ['auth.json', 'settings.json', 'trust.json', 'models.json', 'models-store.json'] },
```
Rationale: `~/.pi/agent` also contains `sessions/`, `extensions/`, `skills/` and the installed
package trees (`npm/`, `git/`), which on an active host is easily gigabytes; `seedWhole` would
`cp -a` all of it into every container start. The five seeded files are what pi needs to
authenticate and behave consistently: `models.json` is in the list because it holds user-defined
custom providers, and omitting it would silently strip those inside containers. Seeding (RO mount
then copy) also means the in-container pi never writes refreshed OAuth tokens back to the host,
which is the whole point of the seeding policy, and bind mounts stay excluded from `docker commit`
so exports remain secret-free.
Trade-off to accept and document: in-container pi sessions are not visible host-side, so `pi -c`
inside a Docker case only sees that container's own history. Codex shares `sessions/` RW precisely
because Codeman reads it host-side for the response viewer; there is no such reader for pi yet
(the response-viewer follow-up in §9 would justify flipping this).
### 2.6 The `pi` binary name is generic
Unlike `agy`/`codex`/`gemini`, `pi` is a short, common name (Raspberry Pi tooling, personal scripts,
`$PATH` accidents). The resolver must not blindly trust a hit. None of the existing external-CLI
resolvers execute their binary (only `claude-cli-resolver.ts` does, via the cached
`getClaudeCliVersion()`, skipped under vitest), so the sanity check is new ground: model it on
`getClaudeCliVersion()`. Run `pi --version` once via `execFileSync`, cache the result module-level,
skip under `VITEST`, and require output matching `/^\d+\.\d+\.\d+/`; on mismatch treat the binary as
unavailable and log the rejected path. Surface `{ available, path, version }` from
`GET /api/pi/status` so a misresolution is diagnosable from the UI (additive relative to the sibling
endpoints' `{ available, path }`). The `dependency-registry` entry carries `versionArg: '--version'`
for `codeman doctor`.
### 2.7 tmux extended keys (a real pi-specific footgun)
Pi documents (`docs/tmux.md`, verified verbatim) that without
```tmux
set -g extended-keys on
set -g extended-keys-format csi-u
```
tmux collapses `Shift+Enter` and `Ctrl+Enter` into a plain `\r` (and `Alt+Enter` into `\x1b\r`), and
pi's editor uses those for newline vs submit. `extended-keys-format` requires tmux 3.5+; tmux
3.2-3.4 works with `extended-keys on` alone (pi then falls back to xterm `modifyOtherKeys`).
Codeman's own browser input path sends `\r` for submit, so basic use works unconfigured, but
newline-in-editor is degraded both for a user typing in an attached terminal (`sc`) and potentially
for the browser Shift+Enter path.
Upstream recommends `~/.tmux.conf` and notes the setting may need a full `tmux kill-server` restart
to take effect. **Codeman must NEVER run `kill-server` on its socket** (it would kill every live
session, including `w1`/`w2`/`w3`). Action: attempt to set both options **server-scoped on
Codeman's own socket only** (`tmux -L codeman set -s ...`, never `-g` on the user's default socket)
at the point the tmux server is first started, verify with `tmux -L codeman show-options -s` and an
empirical Shift+Enter test which scope actually takes for the installed tmux version, and fall back
to a documented manual step in `docs/pi-integration.md` (a `~/.tmux.conf` snippet plus the
kill-server caveat) if it cannot be applied safely to an already-running server. Upstream does not
discuss socket- or server-scoped configuration at all, so this verification is original work, not a
doc lookup.
**RESULT (measured, tmux 3.4 + pi 0.84.1):** `tmux -L <socket> set -s extended-keys on` takes effect
on an **already-running** server with **no `kill-server`** — pi's own startup warning
(`Warning: tmux extended-keys is off…`, a convenient in-band probe) disappears for the next session
started afterwards. `extended-keys-format` does **not exist on tmux 3.4** and errors with
`invalid option: extended-keys-format`, so the two options must be issued independently rather than
chained. Decision: Codeman does **not** set this itself — it is a server-wide tmux option affecting
every session of every backend, so silently changing key encoding is not Codeman's call. It is
documented as a user step in `docs/pi-integration.md` instead, carrying the measured facts.
### 2.8 The completeness trap: which mode tables fail loud vs silent
Adding `'pi'` to the `SessionMode` union makes some omissions compile errors and leaves others
silent. The plan calls this out so review can focus on the silent ones.
**Loud (typecheck fails until edited):** `getModeLabel()` (`session.ts:168-183`, exhaustive switch
with no default), `defaultDockerCommandForMode` and `defaultRemoteCommandForMode` (both typed
`Record<...CommandMode, string>`), **but only after** `RemoteCommandMode` (`types/session.ts:48-51`)
and `DockerCommandMode` (`:157-161`) are widened: both are `Extract<SessionMode, '...'>` with every
member spelled out, so forgetting the `Extract` lists keeps `tsc` green while docker/remote pi cases
silently fall back to `exec bash -l` via the `|| commands.shell` on the lookup. Edit union + both
`Extract` lists + both `Record` literals together.
**Silent (compiles clean, mode just doesn't work):**
- `appendResumeFlag()` (`tmux-manager.ts:1030-1042`) has a `default:` arm; a missing `case 'pi'`
silently drops docker resume.
- `buildSpawnCommand()` (`:770-825`) and `buildPathExport()` (`:1680-1707`) are if-chains with
fallthrough returns; a missing branch spawns pi as a login shell / with no PATH augmentation.
- `isExternalCliMode()` / `isAltScreenStripMode()` are boolean chains.
- The `runMode` accessor's **setter whitelist** (`session-ui.js:2949-2960`) coerces any unknown mode
to `'claude'`. Omitting `pi` there makes the mode **unselectable while every other edit appears to
work**: this is the single most deceptive omission in the frontend.
- `window.__codemanCliAvailable` (injected by `renderIndexHtml`, `server.ts:1375-1407`): the client
treats a **missing key as available** (`isCliAvailable` in settings-ui.js), so forgetting the
injection un-gates pi on boxes without the CLI instead of hiding it.
### 2.9 The Daylight skin cascade eats per-mode run-button colors
A finding that changes the CSS work (verified empirically with computed styles on the live
instance, re-confirmed at f39beb3): `styles.css:13681` opens a nested skin block,
`html:not([data-skin="og"]) { ... }`, and the **default skin is `daylight-blue`, not `og`**, so the
block is live for every default-skin user. Inside it, `.btn-toolbar.btn-run` is re-declared
generically and per-mode only for claude/opencode/codex (codex at `:13787`). CSS nesting adds the
wrapper's specificity (the nested rules resolve to (0,3,1) vs (0,3,0) for
`.btn-toolbar.btn-run.mode-X`), so **gemini's and antigravity's toolbar gradients are dead on the
default skin**: both render the generic claude gradient today, still unfixed as of f39beb3. The
base-sheet rules (gemini/antigravity at `:4406`/`:4420`) only ever render on the `og` skin. Since
1.12+ styles.css itself documents this trap in comments (`:9214`, `:11091`), which confirms the
mechanism.
Consequences for pi:
- The toolbar gradient needs **two** rules: one in the base sheet (`:4420` area, for `og`), and one
**inside** the `13681` block next to codex's (`:13787` area), using the block's own idiom
(or the color is invisible to the average user).
- `mobile.css` phone-toolbar colors need `!important` on `background`/`border-color`/`color`,
exactly as the CLAUDE.md gotcha prescribes. Antigravity's phone block (`mobile.css:895-910`,
inside the `@media (max-width: 430px)` opened at `:338`) has no `!important` and is dead on the
default skin; do not copy that mistake.
- Three surfaces work from base rules alone (verified): run-mode **dots** (list at `:4506-4516`;
the skin block overrides only claude/opencode/codex/shell dots, so a base-sheet
`.run-mode-dot.pi` renders as authored), **tab badges**, and the **welcome button** (the skin
block overrides only claude/opencode/tunnel welcome buttons).
- Optional, separate cleanup (not this change): gemini/antigravity could get the same in-block
treatment to resurrect their colors.
### 2.10 Local-echo policy: pi lands on the buffer overlay by default
New since the first draft of this plan: the codex predictive-echo work (1.13+) introduced a
per-session echo policy in `_updateLocalEchoState()` (terminal-ui.js, `_localEchoPolicy` set at
`:2837`): `codex → 'predict'` (write-through predictive echo), `shell → 'off'`, **everything else
→ 'buffer'** (the `LocalEchoOverlay` that buffers typed text until Enter). Pi therefore gets the
buffer overlay on touch devices with zero edits, via the fallthrough.
That default is a real open question, not a freebie: the codex history (issues #218/#219/#220/#222)
shows that a composer which re-renders per keystroke (live-filtering slash picker, server-side
cursor movement, wrap-as-you-type) is starved by buffer-until-Enter, and pi's editor is exactly
such a composer. Decision for v1: ship with the default `'buffer'` policy but make phone-profile
typing an explicit E2E gate (§7 step 4); if pi's editor mis-renders under the overlay, the cheap
fallback is forcing `'off'` for pi (one branch in `_updateLocalEchoState`), and teaching the
predict path pi's composer row is a follow-up, not a v1 requirement.
`test/local-echo-codex-gating.test.ts` pins the per-mode policy via
`it.each(['claude', 'gemini', 'opencode'])` lists (`:193`, `:376`); add `'pi'` to those lists once
the buffer decision is confirmed (or pin the `'off'` branch if that is the outcome).
**RESULT (measured, pi 0.84.1, iPhone 14 Pro profile + a PTY-level A/B):** the buffer policy
**holds**; codex's failure mode does **not** reproduce. Pi's slash picker re-filters on the **whole
composer content**, not on per-keystroke deltas: a one-shot literal write of `/set` (what the overlay
flush does) filters the picker to `settings` **identically** to sending `/ s e t` as five separate
keystrokes, and the delayed `\r` then selects it and opens the settings menu. Prose prompts buffer
correctly (`pendingText` right, nothing on the PTY before Enter), flush on Enter, and are accepted as
a single prompt. `'pi'` was added to both `it.each` lists. The `'off'` fallback stays documented but
unused.
---
## 3. Config surface: `PiConfig` to CLI flags
```ts
/** Pi CLI session configuration */
export interface PiConfig {
/** Model pattern or ID. Supports `provider/id` and a `:<thinking>` suffix (e.g. `sonnet:high`). Passed via --model. */
model?: string;
/** Provider name (anthropic, openai, google, ...). Passed via --provider. */
provider?: string;
/** Reasoning level. Passed via --thinking. */
thinking?: 'off' | 'minimal' | 'low' | 'medium' | 'high' | 'xhigh' | 'max';
/** Continue the most recent session (-c). Per-cwd scoping is strongly implied upstream but not documented; treat as probable. */
continueSession?: boolean;
/** Resume a specific session by ID or partial UUID (--session). Codeman deliberately accepts ids only, never paths. */
resumeSessionId?: string;
/**
* Tri-state project trust (repo-local `.pi/` settings/extensions/skills, plus installing
* missing project packages):
* true -> --approve (trust for this run; loads and EXECUTES repository TypeScript)
* false -> --no-approve (force-deny; the trust prompt never appears)
* absent -> pi's own defaultProjectTrust (ask).
* Multi-user: MATERIALIZED to false for non-granted owners (§5.2).
*/
approveProjectTrust?: boolean;
}
```
Flag mapping in `buildPiCommand()` (new, `tmux-manager.ts`, directly after `buildAntigravityCommand`
at `:718-736`; every builder there regex-allowlists each user value and silently drops failures
because the result lands in a `bash -c "..."` string):
| Field | Flag | Validation |
| --------------------- | ------------------------------- | --------------------------------------------------------------------------------- |
| `approveProjectTrust` | `--approve` / `--no-approve` / nothing | tri-state boolean, clamped (§5.2) |
| `model` | `--model <v>` | `/^[a-zA-Z0-9._\-/:]+$/` (`:` for `sonnet:high`, `/` for `openai/gpt-4o`) |
| `provider` | `--provider <v>` | `/^[a-z0-9-]+$/` |
| `thinking` | `--thinking <v>` | runtime allowlist of the 7 enum values (defense in depth beyond Zod) |
| `resumeSessionId` | `--session <v>` | `/^[a-zA-Z0-9._-]+$/` (same shape as `RESUME_ID_SAFE`, `:1021`; excludes paths on purpose) |
| `continueSession` | `-c` | boolean; **skipped when a valid `resumeSessionId` is present** (the two conflict) |
**Not** wired in v1, with reasons:
- `--api-key <key>`: ⚠️ **never wire this.** It puts a provider secret on the spawn command line,
which is exactly what the socket-scoped `tmux setenv` discipline exists to prevent (visible in
`ps`, tmux server state, and logs). Listed here so nobody "helpfully" adds it later.
- `--tui-mode` (released in 0.84.0): never passed by Codeman. The main-screen default is the
friendly case for the browser terminal, and fullscreen remains the user's own runtime choice via
`/settings` (§2.2 is designed for that). `--use-theme` (still unreleased) likewise.
- `--name <name>` (`-n`): nice for `/resume` readability, but names contain spaces and would be the
first user-controlled value needing real shell quoting in `buildSpawnCommand`. Defer.
- `--no-session`: ephemeral mode fights respawn/resume. Defer.
- `-p`/`--print`, `--mode json`, `--mode rpc`: non-interactive transports, a different product shape
(§9). Note upstream already shipped a breaking change to JSON-mode `message_update` framing, so
any future consumer must assemble deltas.
- `--tools` / `--exclude-tools` / `--no-tools` / `--no-builtin-tools` (`-t`/`-xt`/`-nt`/`-nbt`): a
genuinely useful "read-only session" affordance (0.84.0 also added a `defaultTools` setting), but
it needs UI design. Follow-up.
- `-r`/`--resume` (interactive picker), `--fork`, `-e`/`--extension`, `--skill`, `--system-prompt`,
`--append-system-prompt`, `--export`, `--models`, `--list-models`: not session-manager concerns in
v1. (`-e` matters later: §9's extension follow-up notes CLI extensions load before trust
resolution.)
---
## 4. Implementation phases
### Phase 1: Backend core
| File | Change |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| `src/utils/pi-cli-resolver.ts` | **New**, mirror `antigravity-cli-resolver.ts` (65 lines: search-dir list, module-level cache with `''` negative sentinel, `which pi` first). Search dirs: `~/.local/bin`, `/usr/local/bin`, `~/.bun/bin`, `~/.npm-global/bin`, `~/bin`. Add the `pi --version` sanity probe from §2.6 (execFileSync, cached, vitest-skipped). Export `resolvePiDir()`, `isPiAvailable()`, `getPiCliVersion()` |
| `src/utils/index.ts` | Re-export the three (resolver block `:30-36`) |
| `src/types/session.ts` | `SessionMode` union `:46`; **both `Extract` lists**: `RemoteCommandMode` `:48-51`, `DockerCommandMode` `:157-161` (§2.8); new `PiConfig` after `AntigravityConfig` (`:325-333`); `SessionState.piConfig` after `:486`; `@fileoverview` mode list `:11` + config list `:17` |
| `src/mux-interface.ts` | `piConfig?: PiConfig` on `CreateSessionOptions` (config block ends `:78`) and `RespawnPaneOptions` (ends `:109`) |
| `src/session.ts` | `isExternalCliMode()` `:164-167` (+pi); `getModeLabel()` `:168-183` (+`'Pi'`); `_piConfig` field decl `:466-470`; ctor option `:556-563` + apply `:652-654`; `toState()` `:1227-1230`; `_buildRespawnPaneOptions()` `:1466-1469` (single source of truth shared by `startInteractive` and `reattachRemote`); `startInteractive()` createSessionOptions `:1680-1683`; COLORTERM attach-env condition `:1400-1402` (+pi); requires-tmux guard chain `:1751-1768` (new block: "Pi sessions require tmux for env override injection via setenv") |
| `src/tmux-manager.ts` | `buildPiCommand()` after `:736` per §3; `buildSpawnCommand()` signature `:770-779` + dispatch branch after `:822-825`; `appendResumeFlag()` `:1030-1042` (`case 'pi': return \`${modeCommand} --session ${resumeId}\`;`); `buildEnvExports()` truecolor branches `:1604-1609` (+pi); `buildPathExport()` `:1680-1707` (+pi branch calling `resolvePiDir()`); missing-CLI error chain in `createSession` `:1788-1806` (+pi, install hint `npm install -g --ignore-scripts @earendil-works/pi-coding-agent`; note `respawnPane` deliberately has no such check); `piConfig` threading at the four sites `:1748`, `:1817`, `:2041`, `:2080`. **No `_configurePi`** (§2.3) |
| `src/config/dependency-registry.ts` | New entry after antigravity's (`:101-108`; file unchanged since 2026-08-06): `{ id: 'pi', label: 'Pi CLI', category: 'core', required: false, usedBy: ['Pi sessions'], resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['pi'], versionArg: '--version' } }] }` |
| `src/docker-hosts.ts` | `defaultDockerCommandForMode` `:138-149`: `pi: 'exec pi'`. `CRED_STORES` `:597-605`: the `.pi/agent` seedFiles entry per §2.5 (nested `rel` already handled at `:613-645`). File unchanged since 2026-08-06 |
| `src/remote-hosts.ts` | `defaultRemoteCommandForMode` `:92-118`: `pi: remoteLoginShellCommand('pi')` (`remoteLoginShellCommand` at `:88-90`). Login-shell routing is mandatory (the #209/e803186 lesson: ssh remote-command exec sees only sshd's minimal PATH, and npm's global bin is usually only on PATH via rc files) |
### Phase 2: Web layer
| File | Change |
| ---------------------------------- | ----------------------------------------------------------------------------------------------- |
| `src/web/schemas.ts` | `'PI_'` in `ALLOWED_ENV_PREFIXES` `:125` **and** the prose error message `:163` (which now also names `CLAUDE_CONFIG_DIR`; the `ALLOWED_ENV_KEYS` exact-key set needs no change); new `PiConfigSchema` after `AntigravityConfigSchema` (`:256-271`), mirroring §3's regexes, `.optional()`, not `.strict()`; `piConfig` on `CreateSessionSchema` (`:299` area) and `QuickStartSchema` (`:712` area); `'pi'` in all three mode enums (`:285`, `:708`, cron `agentType` `:1214`; they are byte-identical and there is no fourth); `pi` key in `RemoteCommandOverridesSchema` `:426-436` (it is `.strict()`, so an unknown key is a hard error today; one edit covers both remote `:501` and docker `:577` reuse) |
| `src/web/routes/session-routes.ts` | Thread `piConfig` through create (`POST /api/sessions`): disk-strip exclusion chain `:705-712`, availability gate `:782-790` (+`isPiAvailable` with install-hint error), model resolution `:825-838` (`mode === 'pi' ? body.piConfig?.model : ...`), clamp call `:845`, Session ctor `:860` (`piConfig: mode === 'pi' ? gatedPiConfig : undefined`). Quick-start (`POST /api/quick-start`, handler `:2559`): remote-case config rejection `:2614-2621` and docker-case `:2645-2652` (+`piConfig`: per-CLI config does not cross ssh or the bind mount), hooks-scaffold exclusions `:2801`/`:2809`, availability gate `:2744-2752` (local-case branch only), env-strip chains `:2833`/`:2863`, model resolution `:2885`, clamp `:2897`, ctor `:2913`. **Extend `clampExternalCliBypassForOwner()`** (`:305-336`, doc comment above): fifth param + return field; pi joins the **materialize** branch per §5.2. Alt-screen replay-strip at `:2275` unchanged (pi not in it, §2.2) |
| `src/web/routes/system-routes.ts` | `GET /api/pi/status` after the antigravity handler (`:418-426`; file unchanged since 2026-08-06), same shape plus `version` (§2.6); update the "CLI Integrations" prose comment `:377` |
| `src/web/server.ts` | Restore path: `piConfig: muxSession.mode === 'pi' ? savedState?.piConfig : undefined` after `:2636`. **`renderIndexHtml` CLI-availability injection `:1375-1407`**: add `isPiAvailable` to the dynamic-import tuple (`:1382`) and a `pi` key to the injected object (`:1399`). Per §2.8 a missing key reads as *available*, so this is a correctness edit, not polish |
### Phase 3: Frontend
The antigravity touchpoints are the template. Since the first draft, the settings-surface overhaul
moved most anchors and added one **new touchpoint** (the clone-repo Brain picker below).
`constants.js`, `api-client.js`, `ralph-wizard.js`, `cron-ui.js`, `webview-tabs.js` and `sw.js`
still need **no** changes (re-verified zero mode coupling at f39beb3; cron-ui reads the `<select>`
generically and special-cases only `shell`).
| File | Change |
| ------------------- | ------------------------------------------------------------------------------------------------------ |
| `index.html` | Welcome button `welcomePiBtn` after Gemini's (antigravity's is `:347`; there is deliberately no codex welcome button), `display:none` default, `onclick="app.setRunMode('pi'); app.runPi()"`, text `Run Pi`; run-mode-option row with `.run-mode-dot.pi` after antigravity's (`:526-528`), before the `.run-mode-sep` `:529`; cron `<option value="pi">Pi</option>` after `:803`; **NEW: the clone-repo "Brain" picker** (`cloneCaseBrain`, `:2476-2486`): add `<option value="pi" data-cli="pi">Pi</option>` after the antigravity option `:2483` (gating is automatic: session-ui.js `:2107-2115` hides options whose `data-cli` fails `isCliAvailable`, and `:2250` reads the value at clone time); docker image hint `:2624` (`claude/codex/gemini/opencode/agy` + pi). No per-CLI remote-command override field needed (only codex has one, `:2559`) |
| `session-ui.js` | `@fileoverview` mode list `:2`; `run()` dispatch branch after `:400-402`; `_refreshRunModeAvailability` list `:468` (+`'pi'` as a quoted literal, the static test in §6 demands it); short-label ternary `:565` (+`'Run PI'`); **the `runMode` setter whitelist `:2949-2960`** (§2.8, the deceptive one); new `runPi()` modeled on `runAntigravity()` `:1170-1219`: same remote/docker skip, same `_beginSessionLaunchStatus` frame, probes `/api/pi/status` reading `(await res.json()).data.available` (envelope!), **sends no `piConfig` at all** (no bypass exists and trust defaults are pi's own; envOverrides still sent for local cases), install-hint error text matching Phase 1's; `isAltMode` `:1233` and `isExternalCli` `:1263` four-way comparisons (+pi) |
| `settings-ui.js` | `applyWelcomeCliVisibility()` `:1176-1191`: add `['welcomePiBtn', 'pi']` |
| `app.js` | Response-viewer agent label `:1998-2009` (+pi -> `'Pi'`); tab badge ternary `:3884` (`<span class="tab-mode pi" aria-hidden="true">pi</span>`; claude stays badge-less); kill-title ternary `:5046-5057` (`Kill Tmux & Pi`) |
| `panels-ui.js` | Command-palette `labels` map `:430` (+`pi: 'Pi'`; the `\|\| mode` fallback means this is cosmetic, not load-bearing) |
| `mobile-overview.js`| `MOBILE_OVERVIEW_RUN_MODES` `:55-62`: `{ mode: 'pi', label: 'Pi', short: 'Pi' }` after antigravity `:60`, before the shell entry. Nothing else: the Run-button badge (`:499`) and menu builder (`:554-556`) consume the list generically, and the buttons carry `btn-toolbar btn-run mode-pi`, which is exactly why they inherit the §2.9 cascade problem and its fix |
| `terminal-ui.js` | Badge-row comment `:1750` only (the badge itself is a raw `s.mode` passthrough, no list to extend). `_sessionUsesServerMouseStrip` unchanged (§2.2). `_updateLocalEchoState` unchanged for v1 (§2.10: pi lands on `'buffer'` via the fallthrough; only touch it if E2E forces the `'off'` fallback) |
| `i18n.js` | `'Run Pi': '运行 Pi'` in the zh-CN table (`:102-107`, matches the welcome-button text; short labels like `Run PI` are deliberately untranslated, as are the other modes') |
| `styles.css` | Tab badge `.session-tab .tab-mode.pi` after `:2157` (`background: rgba(244,114,182,0.2); color: #f472b6;`); add `.session-tab .tab-mode.pi` to the light-skin ink list `:325-336` (gemini + antigravity are its precedent, `:332`); welcome `.welcome-btn-pi` + `:hover` after antigravity's `:3366` block, rose family (e.g. base `linear-gradient(135deg, #33121f 0%, #9d174d 55%, #be185d 100%)`, border `rgba(244,114,182,0.4)`, text `#fce7f3`); toolbar gradient pair `.btn-toolbar.btn-run.mode-pi, .btn-toolbar.btn-run-gear.mode-pi` + `:hover` after `:4420`'s antigravity block; `.run-mode-dot.pi { background: #f472b6; }` in the dot list `:4506-4516`; **and the §2.9 rule inside the Daylight block** next to codex's `:13787` (e.g. `background: linear-gradient(135deg, #be185d, #f472b6); border-color: #be185d; color: #fff1f7;`). The dot needs no skin-block entry (the block overrides only claude/opencode/codex/shell dots; gemini/antigravity dots already fall through correctly) |
| `mobile.css` | Phone toolbar block after `:910` inside the `@media (max-width: 430px)` opened at `:338`: `mode-pi` base + `:active`, **with `!important` on background/border-color/color** (§2.9; antigravity's block `:895-910` omits it and is dead); light-skin override entry after `:2985` with the same four-skin `html:is(...)` prefix as its siblings |
### Phase 4: Docker image and installer
Both files are unchanged since the 2026-08-06 verification; all anchors stand.
- `docker/agent.Dockerfile`: a **separate** `RUN` step after the antigravity block (`:38-45`), not a
fifth line in the shared npm block (`:31-36`), because pi documents `--ignore-scripts` and that
flag must not silently change how the other four install:
```dockerfile
# Pi (pi.dev). Upstream documents --ignore-scripts (pi needs no lifecycle scripts);
# kept out of the shared npm block above so the flag cannot affect the other CLIs.
RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
&& npm cache clean --force \
&& pi --version
```
Implementation checklist item: the gid-0 pre-created dirs at `:64-68` include `.claude/projects`
and `.codex/sessions`; verify whether the cred-seed copy into `~/.pi/agent` creates its target
dir in a fresh container or whether `.pi/agent` must join that `mkdir` line. Rebuild with
`node scripts/build-agent-image.mjs --no-cache` (the script itself needs no change; nothing in it
is CLI-specific). The cached npm layer has silently frozen a CLI at a broken version before; see
`docs/docker-cases.md`.
- `install.sh` (six edit sites, all verified): `PI_SEARCH_PATHS` block after `:125` (mirror the
resolver's dirs); `check_pi` / `get_pi_path` pair inserted at `:531` (antigravity's pair spans
`:504-530`); the satisfying-AI-CLI chain `:2032-2063` (`has_pi` local at `:2037` area, detect
block after `:2059`, widen the five-way test at `:2061` and the warn text at `:2063`); the menu
option-4 text `:2070`; the skip-path hints `:2115-2116` (add
`npm install -g --ignore-scripts @earendil-works/pi-coding-agent (Pi)`); the final no-CLI
reminder `:2416-2423` (add `check_pi` to the condition and a pi line to the echo block).
Detection plus a hint only; do **not** add an auto-install path in this change.
### Phase 5: Docs
- `docs/pi-integration.md` (**new**, user-facing): install (both installers uninstall via npm), auth
(`/login` OAuth for six providers vs API keys; `pi auth check` for preflight; Claude Pro/Max
third-party harness usage bills as Anthropic "extra usage" per token, not plan limits; OpenRouter
login supports pasting the redirect URL, which matters over remote SSH), what Codeman wires up
and deliberately does not (§3, incl. never passing `--tui-mode`), the tmux extended-keys note
from §2.7 with the manual `~/.tmux.conf` fallback, Docker/remote behaviour (in-container sessions
invisible host-side), the trust model in §1 words, known gaps.
- `CLAUDE.md`: tech-stack line (six CLIs + `SessionMode` union), the env-prefix gotcha bullet, the
multi-CLI prefix-discipline bullet, the "External CLI modes" key-pattern paragraph (note it now
also carries the codex predictive-echo block; pi's echo-policy decision from §2.10 belongs in the
same paragraph), the `src/utils/` resolver list.
- `docs/architecture-invariants.md`: the external-CLI-modes section. ⚠️ Its anchor was already
renamed once to `#external-cli-modes-opencode-codex-gemini-antigravity` while CLAUDE.md's link
text still shows the old name; when renaming again for pi, update every inbound link (CLAUDE.md
and this file).
- `docs/docker-cases.md` (cred-seeding table + supported modes + image contents),
`docs/remote-sessions.md` (`RemoteCommandMode`), `docs/cron-guide.md` + `docs/cron-discovery.md`
(`agentType` enum; note the readiness caveat from §6's cron paragraph),
`docs/security-architecture.md` (env prefix allowlist row).
- `README.md` + `README.zh-CN.md`: six CLIs.
- `package.json` keywords: `pi`.
- Update the issue #206 thread when it ships.
---
## 5. Security checklist
1. **Command injection.** Every `PiConfig` value is regex-validated in `buildPiCommand()` before
entering the `bash -c "..."` string; anything failing validation is dropped, not escaped
(matching the four existing builders). No user string reaches the spawn line unvalidated. Pinned
by a "rejects unsafe values" test per field.
2. **Multi-user clamp, materialize branch.** `approveProjectTrust` is the privilege-shaped field: it
makes pi execute repository-supplied TypeScript and install project packages.
`clampExternalCliBypassForOwner()` (`session-routes.ts:305-336`) has two branches, and pi belongs
in the **gemini-style materialize branch**, not the codex/antigravity only-if-sent branch:
pi's absent-config default is an *interactive trust prompt the session user can answer
themselves in the terminal*, so merely omitting `--approve` is not a clamp. For a non-granted
owner, materialize `{ ...(piConfig ?? {}), approveProjectTrust: false }` so `buildPiCommand`
always emits `--no-approve` and the prompt never appears. Both call sites (`:845`, `:2897`)
widen. This helper still has **zero test coverage** (re-confirmed at f39beb3); §6 adds the first
tests.
3. **Secrets stay off the command line.** `PI_*` overrides flow through `applyEnvOverrides()` /
socket-scoped `tmux setenv`, never inlined into the spawn string. No `-e` at container create
time. And `--api-key` is never wired (§3): it would put a provider secret into `ps`/tmux state.
4. **Env allowlist not widened.** Only the `PI_` prefix is added; the provider keys stay out (§2.4)
and `ALLOWED_ENV_KEYS` is untouched. Pinned by a test that `PI_OFFLINE` passes and
`ANTHROPIC_API_KEY` still fails validation.
5. **Docker seeding, not sharing.** Per §2.5: RO mount then copy, so refreshed OAuth tokens never
write back to the host; bind mounts stay excluded from `docker commit` so exports remain
secret-free.
6. **Remote SSH.** `pi` mode goes through `defaultRemoteCommandForMode` and therefore
`buildSshConnectionArgs()`. No hand-built ssh line anywhere.
7. **No sandbox claims.** Pi documents that it has no sandbox and no permission prompts, and that
extensions run with the user's full permissions. Codeman docs must say plainly that a pi session
can read, write and execute anything the Codeman user can, and point at Docker cases as the
isolation story. Do not imply the trust prompt is a safety boundary (upstream itself says it is
not). Worth one doc sentence: `pi auth print-api-key` / `print-bearer-token` (0.83.0) and
`pi auth check` (0.84.1) mean a pi session can print its own provider credentials by design;
isolation, again, is Docker.
8. **Loud-vs-silent audit.** Before review, walk §2.8's silent list and confirm each site has its
pi branch; the loud ones the compiler already caught.
---
## 6. Test plan
- `test/pi-mode.test.ts` (**new**, modeled on `test/antigravity-mode.test.ts`, 125 lines, no port;
file unchanged since 2026-08-06 so its structure remains the template):
`CreateSessionSchema`/`QuickStartSchema` accept a pi config; unsafe `model`/`provider`/
`resumeSessionId` values are rejected (`'pi; rm -rf /'` shapes); `buildSpawnCommand({ mode: 'pi', ... })`
emits expected flags, drops invalid ones, emits `--no-approve` for `approveProjectTrust: false`
and `--approve` for `true`, and skips `-c` when a `resumeSessionId` is present;
`defaultDockerCommandForMode('pi') === 'exec pi'` and
`defaultRemoteCommandForMode('pi') === 'exec "${SHELL:-/bin/sh}" -i -l -c \'pi\''`;
`isExternalCliMode('pi') === true`, `isAltScreenStripMode('pi') === false`; the env pair
(`PI_OFFLINE` accepted, `ANTHROPIC_API_KEY` rejected), mirroring antigravity-mode `:49-63`.
- **First-ever coverage for `clampExternalCliBypassForOwner`** (still nothing in `test/` touches
it): cover pi's materialize branch (absent config still yields `approveProjectTrust: false` for a
non-granted owner; a sent `true` is forced to `false`; granted owner passes through) and, while
there, pin the three existing modes' behavior. Prefer exporting the helper for direct unit tests
over a heavier multi-user route fixture; either way it lives under `test/routes/`.
- `test/run-mode-ui.test.ts`: extend `loadUi()`'s stub lists (welcome-button ids, mode buttons,
`ALL_OFF`) and add pi welcome/dropdown gating cases; note the static parser test
`'gates every mode the run-mode menu actually offers'` (`:433-456`) picks up the new
`data-mode="pi"` from index.html automatically and **fails until** `_refreshRunModeAvailability`
contains a quoted `'pi'`, which is exactly the regression it exists for. Add a
`describe('Pi quick start')` modeled on the antigravity one (`:840`) driving `runPi()` against a
stubbed `/api/pi/status` + `/api/quick-start`, asserting the posted body has `mode: 'pi'` and
**no `piConfig`**, and that the envelope is unwrapped. (The short-label assertion pattern is at
`:82`, `'Run AG'`.)
- `test/render-index-html.test.ts` `:141`: the injected `window.__codemanCliAvailable` is asserted
with an exact `toEqual` and now carries **seven** keys (claude, opencode, codex, gemini,
antigravity, cloudflared, and since 1.12+ `git`), so it **must** gain the `pi` key (and the
resolver mock an `isPiAvailable`); its comment explains why: a dropped key silently un-gates
(§2.8).
- `test/routes/system-routes.test.ts`: `GET /api/pi/status` shape, modeled on the antigravity
describe (`:816-838`) + resolver mock (`:84-87`); file unchanged since 2026-08-06.
- `test/mobile-overview.test.ts`: `:375` is an exact-array `toEqual` over the run-menu modes and
**will fail until updated** to include `'pi'` (the second exact-array at `:366`,
`['claude', 'shell']`, is a gating case and stays as-is); the sibling static parser then covers
the new entry automatically. The no-hex-literals guard only scans `.mobile-overview*` rules, so
pi's `mode-pi` colors in mobile.css do not trip it.
- `test/local-echo-codex-gating.test.ts` (§2.10): once the buffer-policy decision is confirmed in
E2E, add `'pi'` to the `it.each(['claude', 'gemini', 'opencode'])` lists (`:193`, `:376`) so the
chosen policy is pinned.
- `test/skin-themes.test.ts`: will NOT trip (it enumerates skins, not modes); run it anyway since
styles.css is touched. `test/mobile-header-buttons-policy.test.ts`: trips only if a header
button is added; pi adds none (welcome button and run-menu rows are outside `header-right`).
- Cron: schema-level acceptance of `agentType: 'pi'` (the service consumes `SessionMode`
generically; `src/cron/` is unchanged since the first draft). Known, documented degradation: the
readiness poll (`cron-service.ts:515`) looks for `❯`/`tokens`, which pi never prints, so cron pi
jobs burn the ready-poll attempts and then send anyway. Acceptable for v1; note it in
`docs/cron-guide.md`.
- Sweep with `npm run test:ci`. Never bare `npm test`. No new ports needed (all new/extended suites
are portless).
---
## 7. End-to-end verification (required before COM)
Unit tests passing is not evidence the mode works (pi is not currently installed on the dev box, so
step 1 is a real step). Before shipping:
1. Install pi (`npm install -g --ignore-scripts @earendil-works/pi-coding-agent`), authenticate once
with `/login`.
2. `curl -sk https://localhost:3000/api/pi/status | jq` reports `available: true`, the right path,
and a sane `version`.
3. Create a **throwaway** case, launch a pi session from the Run dropdown, send a prompt from the
browser, confirm the reply renders and scrollback survives a tab switch. Do not touch
`w1`/`w2`/`w3`.
4. **Local-echo policy gate (§2.10):** on a phone profile, type into the pi editor through the
buffer overlay (drive with `page.keyboard.type()`, never `app.sendInput()`, and force
`app._localEchoEnabled = true`; headless Chromium reports touch as false) and confirm pi's
composer renders the flushed text correctly on Enter. If it mis-renders, flip pi to the `'off'`
branch in `_updateLocalEchoState` and pin that instead.
5. Visual pass on the **default skin** (the §2.9 finding makes this the load-bearing check, not a
formality): run-button gradient actually renders rose (not generic claude blue), dot, tab badge,
welcome button, kill-menu label; then a phone profile (toolbar `!important` colors and light-skin
overrides are the usual regressions).
6. Kill and respawn the session; confirm `piConfig` round-trips through `state.json` and the pane
comes back with the same flags. Then `/clear`-style respawn via the Respawn tab.
7. Extended keys (§2.7): in an attached terminal, verify whether Shift+Enter inserts a newline in
pi's editor with and without the socket-scoped options; record the outcome in
`docs/pi-integration.md` either way. While attached, also flip `/settings` to the fullscreen TUI
and back to confirm the no-strip decision holds (§2.2).
8. Trust model: point a throwaway case at a repo containing `.pi/extensions`, confirm the trust
prompt appears interactively and that a multi-user non-granted session instead launches with
`--no-approve` (prompt never shown, extensions not loaded).
9. **NOT RUN in this pass — an honest gap.** Docker case with `mode: 'pi'`: rebuild the agent image with `--no-cache`, confirm `pi --version`
inside the container **as the `agent` user**, confirm seeded auth works and a session starts
(this is exactly where the antigravity Docker path broke in 1.11.2: the CLI was never installed
in the image).
10. **NOT RUN in this pass — the other gap.** Remote SSH case with `mode: 'pi'`: confirm the
login-shell wrapper resolves the npm global bin.
11. Only then: changeset, `COM minor` (new capability, additive to the API surface).
**Verification actually performed** (2026-08-13, pi 0.84.1, isolated `CODEMAN_INSTANCE=pi-beta`
server on :5055 with its own tmux socket and data dir): steps 1-8 pass. Highlights:
`/api/pi/status` resolved through the **search-dir fallback** (pi installed to `~/.npm-global/bin`,
deliberately not on PATH) and reported
`{available:true, path:'/home/arkon/.npm-global/bin', version:'0.84.1'}`; the real spawn line came
out as `… COLORTERM=truecolor … && pi --approve --provider anthropic --thinking high`; `piConfig`
round-tripped through `state.json` across a **full server restart**; the trust prompt appeared for a
case containing `.pi/extensions` + `.pi/settings.json`, and `--no-approve` suppressed it
(`This project is not trusted. Project .pi resources and packages are ignored.`); on the **default
`daylight-blue` skin** the toolbar Run button computed to
`linear-gradient(135deg, rgb(190,24,93), rgb(244,114,182))` — genuinely rose and **distinct from
claude's blue**, so the §2.9 cascade trap is avoided; and flipping `/settings` to the fullscreen TUI
put the pane into the alt screen (`alternate_on=1`), **empirically confirming §2.2**: had pi been in
the strip list, Codeman would have stripped that switch and corrupted the session. Steps 9-10 need a
Docker daemon and a remote host respectively.
---
## 8. Effort estimate
Calibrated against the real antigravity history, which is the honest baseline: the feature commit
`26cbbe0` was 24 files, +638/-63, and it then took **four follow-up commits** (`e803186` login-shell
routing, `292ba2c` ownership helpers, `5d28999` CLI gating incl. tests, `0d0b772` docs/installer/UI
propagation) totaling roughly +600/-170 across ~43 file-touches to make the mode actually
first-class. Budgeting only the feature-commit shape under-scopes by ~40%. This plan folds all four
follow-up surfaces in from the start (login-shell routing in Phase 1, availability gating in Phases
2-3, installer/docs propagation in Phases 4-5), so expect the full footprint in one pass:
| Phase | Size |
| --------------------- | -------------------------------------------------------------------------- |
| 1. Backend core | ~260 lines across 9 files, one new file (resolver incl. version probe) |
| 2. Web layer | ~110 lines across 4 files (incl. the clamp widening + availability inject) |
| 3. Frontend | ~175 lines across 10 files (enumerations + CSS in two sheets + skin block + the Brain picker option) |
| 4. Docker + installer | ~45 lines, plus one `--no-cache` image rebuild |
| 5. Docs | one new doc, ~10 files touched |
| 6. Tests | one new test file, 6 extended (2 of which fail loudly until updated), plus the first clamp coverage |
---
## 9. Out of scope, tracked as follow-ups
- **A Codeman pi extension for real idle/completion events (highest value, now fully de-risked).**
Pi extensions are TypeScript modules with Node built-ins and npm deps available, so an HTTP POST
to `/api/hook-event` is trivial. The **`agent_settled`** event **shipped in 0.84.0** and is
documented for exactly this use case (fires only when pi will not continue on its own: after
auto-retries, auto-compaction and queued follow-ups; `ctx.isIdle()` is true inside the handler).
That is a genuine idle signal replacing output-silence heuristics, i.e. the same class of upgrade
hooks give Claude sessions. The bash tool exposes five env vars (`PI_SESSION_ID`,
`PI_SESSION_FILE`, `PI_PROVIDER`, `PI_MODEL`, `PI_REASONING_LEVEL`), injected per command. Bonus:
an extension can own the **`project_trust`** event (first yes/no wins, and CLI `-e` extensions
load *before* trust resolution), so Codeman could answer the trust prompt programmatically, a
cleaner mechanism than the `--approve` flag for both the single-user convenience case and the
multi-user deny case.
- **Response viewer for pi.** Sessions are JSONL v3 under
`~/.pi/agent/sessions/--<cwd-dashed>--/<timestamp>_<uuid>.jsonl` with an `id`/`parentId` tree and
typed content blocks (text, image, thinking, toolCall); the cwd-derived dir name is trivially
computable host-side. Feasible, and it would justify flipping the Docker cred policy to share
`sessions/` RW like Codex.
- **Mode-aware env allowlist.** Would let pi sessions accept provider keys without widening the
global list. Needs `ALLOWED_ENV_PREFIXES` to become a per-mode map plus mode context inside the
Zod refine.
- **`--tools` / `--exclude-tools` / `--no-tools` / `--no-builtin-tools` read-only sessions** (plus
the 0.84.0 `defaultTools` setting). Real product value, needs UI.
- **Predictive echo for pi's composer** if the §2.10 buffer decision does not hold up in practice:
teach `PredictiveEchoAddon` pi's composer row the way `isCodexComposerRow` handles codex's.
- **`--mode json` / `--mode rpc`, and upstream's experimental remote-session client APIs**
(transport-neutral `PiClient`, CBOR protocol, Unix-socket transport, `RemoteSession` controller,
still unreleased as of 0.84.1). A potential non-PTY integration path, a different architecture
from the tmux+PTY model. Note the already-shipped breaking change to `message_update` framing
(delta-only): any consumer must assemble deltas between `message_start`/`message_end`.
- **`--name` for session labels.** Blocked on shell-quoting a user string in `buildSpawnCommand`.
---
## 10. Risks
| Risk | Mitigation |
| ------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
| `pi` resolves to an unrelated binary | `pi --version` + semver-shape check in the resolver (§2.6); path and version shown in `/api/pi/status` |
| Pi's TUI repaints in a way the browser terminal handles badly | Test scrollback and repaint early (step 3 of §7); pi's default is main-screen with terminal-owned scrollback, which is the friendly case |
| Fullscreen TUI mode (shipped 0.84.0, runtime-switchable) | Already designed for: pi stays OUT of the strip list, so a user flipping `/settings` to fullscreen gets opencode-like alt-screen behavior, not corruption. §7 step 7 tests the flip explicitly |
| The buffer local-echo overlay fights pi's live composer | §2.10: explicit E2E gate (§7 step 4) with the one-line `'off'` fallback; predictive echo for pi is a tracked follow-up, not a v1 blocker |
| Pi moves fast (pre-1.0; 9 releases in the 7 weeks before 0.84.1) | Keep the flag surface small; every flag validated and droppable; nothing pinned in the Dockerfile beyond the `--no-cache` rebuild cadence. Live example of the hazard: `--tui-mode` went from main-only docs to released between the two drafts of this plan |
| Docker image grows | Pi is an npm package; the layer is modest next to the ~190MB `agy` binary |
| Trust prompt blocks a session | Narrower than feared: only fires when `.pi/settings.json`, `.pi/extensions\|skills\|prompts\|themes`, `.pi/SYSTEM.md`/`APPEND_SYSTEM.md` or `.agents/skills` exists (bare `.pi/` does not). Documented; `approveProjectTrust` is the opt-in escape hatch; multi-user forces `--no-approve` (§5.2); the `project_trust` extension follow-up removes the prompt entirely |
| Interactive `/login` OAuth can't complete headlessly | Document: authenticate once interactively (or seed `auth.json`); `pi auth check` verifies credentials preflight; OpenRouter's paste-the-redirect-URL flow covers remote SSH |
| Provider auth is awkward without key prefixes in the allowlist | `/login` writes `~/.pi/agent/auth.json` once and Docker seeds it; the mode-aware allowlist follow-up removes the friction |
| Cron pi jobs mis-detect readiness | Known degradation, documented in §6; readiness falls through after the poll budget and the prompt still sends |
-235
View File
@@ -1,235 +0,0 @@
# Pi (pi.dev) sessions
Codeman can drive [Pi](https://pi.dev) (`@earendil-works/pi-coding-agent`, MIT) as a
session backend, alongside Claude Code, OpenCode, Codex, Gemini and Antigravity.
`pi` is a sixth **run mode**: its own PTY, its own tmux session, its own tab colour
(rose). It is not a location overlay like Docker or remote-SSH cases, and it is not
a web tab.
Tracking issue: [#206](https://github.com/Ark0N/Codeman/issues/206). The design
rationale behind each decision below lives in `docs/pi-integration-plan.md`.
## Install
```bash
npm install -g --ignore-scripts @earendil-works/pi-coding-agent
# or
curl -fsSL https://pi.dev/install.sh | sh
```
Both installers end up going through global npm, so either one uninstalls with
`npm uninstall -g @earendil-works/pi-coding-agent`.
Codeman finds the binary via `which pi` and then the usual global-bin locations
(`~/.local/bin`, `/usr/local/bin`, `~/.bun/bin`, `~/.npm-global/bin`, `~/bin`).
**`pi` is a short, generic name**, so unlike the other CLI resolvers Codeman does
not trust a `which` hit on its own: it runs `pi --version` once and requires
semver-shaped output. Anything else is rejected as "not installed" and the
rejected path is logged. Check what it resolved:
```bash
curl -s localhost:3000/api/pi/status | jq
# { "available": true, "path": "/home/you/.local/bin", "version": "0.84.1" }
```
That endpoint carries `version` on top of the shape the sibling `/api/*/status`
endpoints return, precisely so a misresolution is visible rather than presenting
as "the mode just doesn't work".
## Authenticate
Pi supports 15+ providers. Two ways in:
- **OAuth subscription login** — run `/login` inside a pi session. Six providers
support it: ChatGPT Plus/Pro, Claude Pro/Max, GitHub Copilot, xAI, OpenRouter
and Radius. Credentials land in `~/.pi/agent/auth.json` and pi refreshes them
itself. OpenRouter's flow accepts a pasted redirect URL, which is what makes it
workable over remote SSH.
- **API keys** — exported in the environment of the **Codeman server process**.
⚠️ **Provider API keys cannot be sent as per-session `envOverrides`.** Pi reads
about 34 provider variables (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`,
`DEEPSEEK_API_KEY`, `HF_TOKEN`, `BASETEN_API_KEY`, …) that share no common prefix.
Codeman's env allowlist is a single global list applied to every mode at once, so
admitting bare provider keys for pi would widen the allowlist for Claude, Codex,
Gemini and everything else too. Only the **`PI_*`** prefix was added, which covers
every documented pi input: `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`,
`PI_PACKAGE_DIR`, `PI_OFFLINE`, `PI_SKIP_VERSION_CHECK`, `PI_TELEMETRY`,
`PI_CACHE_RETENTION`, `PI_SHARE_VIEWER_URL`, `PI_HARDWARE_CURSOR`,
`PI_EXPERIMENTAL`.
`pi auth check` verifies credentials before you start a long run.
Note if you authenticate with a Claude Pro/Max subscription: third-party harness
usage bills as Anthropic "extra usage" per token rather than against plan limits.
## What Codeman wires up
`PiConfig` (per session, persisted in `state.json`, round-trips through respawn):
| Field | Flag | Notes |
| --------------------- | -------------------------------------- | ---------------------------------------------------------------- |
| `model` | `--model <v>` | Accepts `provider/id` and a `:<thinking>` suffix (`sonnet:high`) |
| `provider` | `--provider <v>` | `anthropic`, `openai`, `google`, … |
| `thinking` | `--thinking <v>` | `off`/`minimal`/`low`/`medium`/`high`/`xhigh`/`max` |
| `continueSession` | `-c` | Skipped when `resumeSessionId` is set (the two conflict) |
| `resumeSessionId` | `--session <v>` | Ids only, never paths |
| `approveProjectTrust` | `--approve` / `--no-approve` / nothing | Tri-state, see below |
Every value is regex-validated and **dropped** (not escaped) if it fails, because
the result is interpolated into the pane's `bash -c "…"` command.
The Run button sends **no `PiConfig` at all**: pi has no permission prompts to
bypass, and project trust is a decision the person at the terminal makes.
## What Codeman deliberately does NOT wire up
- **`--api-key`.** Never. It would put a provider secret on the spawn command
line, visible in `ps`, tmux server state and logs. `PI_*` overrides go through
socket-scoped `tmux setenv` for exactly this reason.
- **`--tui-mode`.** Pi's default main-screen TUI is the friendly case for a
browser terminal. The fullscreen mode (0.84.0) stays your own runtime choice via
`/settings`.
- **`--name`, `--no-session`, `-p`/`--print`, `--mode json`, `--mode rpc`,
`--tools`/`--exclude-tools`, `-e`/`--extension`, `--skill`,
`--system-prompt`.** Tracked as follow-ups in the plan doc.
## Permission and trust model — read this
**Pi has no permission prompts and no sandbox.** There is no
`--dangerously-skip-permissions` analog and none is needed: tools run with the
user's own permissions, always. A pi session can read, write and execute anything
the Codeman user can. If you need isolation, use a **Docker case** — that is the
isolation story, here as everywhere else in Codeman.
Pi's "project trust" prompt is **not** a safety boundary (upstream says so too).
It gates *loading* repo-local `.pi/` config, extensions and skills, and
*installing* missing project packages. It only appears when the cwd or an ancestor
contains `.pi/settings.json`, `.pi/extensions|skills|prompts|themes`,
`.pi/SYSTEM.md`/`.pi/APPEND_SYSTEM.md`, or `.agents/skills`. A bare `.pi/`
directory does not trigger it.
`approveProjectTrust: true` answers it with `--approve`, which means pi **loads
and executes repository-supplied TypeScript** and runs an npm install for missing
project packages. Treat it exactly as seriously as that sounds.
**Multi-user mode:** for an owner without the privileged-command grant, Codeman
materializes `approveProjectTrust: false` so the pane launches with
`--no-approve` and the prompt never appears. Merely *omitting* `--approve` would
not be a clamp, since pi's own default is to ask and the session user could just
answer yes.
Also worth knowing: `pi auth print-api-key` / `print-bearer-token` and
`pi auth check` mean a pi session can print its own provider credentials by
design. Isolation is Docker.
## tmux extended keys (Shift+Enter)
Pi's editor uses `Shift+Enter` / `Ctrl+Enter` for newline-vs-submit. Without
extended keys, tmux collapses both into a plain `\r`. Upstream recommends:
```tmux
set -g extended-keys on
set -g extended-keys-format csi-u
```
`extended-keys-format` needs tmux 3.5+; on 3.2–3.4 `extended-keys on` alone works
(pi falls back to xterm `modifyOtherKeys`).
Codeman's browser input path sends `\r` for submit, so basic use works
unconfigured — what degrades is newline-in-editor, mostly when you attach to the
pane directly (`sc`).
⚠️ Upstream notes the setting may need a full `tmux kill-server` to take effect.
**Never run `tmux kill-server` on Codeman's socket** — it would kill every live
session, `w1`/`w2`/`w3` included.
**Measured (tmux 3.4, pi 0.84.1): no `kill-server` is needed.** Setting the option
server-scoped on Codeman's own socket takes effect on the ALREADY-RUNNING server;
the next pi session starts without the warning. Existing sessions keep the old
setting until they respawn.
```bash
tmux -L codeman set -s extended-keys on
tmux -L codeman set -s extended-keys-format csi-u # tmux 3.5+ only, see below
tmux -L codeman show-options -s | grep extended # verify
```
On **tmux 3.4 and older, `extended-keys-format` does not exist** and the second
line fails with `invalid option: extended-keys-format`. That is harmless — pi
falls back to xterm `modifyOtherKeys` and `extended-keys on` alone silences the
warning. Run the two lines independently rather than chained.
Pi tells you which state it is in: an unconfigured session prints
`Warning: tmux extended-keys is off. Modified Enter keys may not work.` in its
startup banner, so you can verify the change by starting a new pi session.
⚠️ Use `-L <socket>` and `-s`, never `-g` on your default socket, and never
`kill-server`. Codeman does not set this for you: it is a server-wide tmux option
and silently changing key encoding for every session of every backend is not
Codeman's call to make.
## Typing from the browser (local echo)
On touch devices Codeman buffers typed characters in the `LocalEchoOverlay` and
flushes them to the PTY on Enter. Pi gets that `'buffer'` policy, the same as
Claude, Gemini and OpenCode.
This was an explicit open question, because that policy is exactly what broke
Codex (issues #218/#219/#220/#222): Codex's composer reacts per keystroke, so
buffer-until-Enter starved it. **Measured against pi 0.84.1: it does not
reproduce.** Pi's slash-command picker re-filters on the whole composer content
rather than on per-keystroke deltas, so a one-shot flush of `/set` filters the
picker down to `settings` identically to typing it character by character, and
the delayed `\r` then selects it. Prose prompts flush and submit correctly too.
If a future pi release changes that, the cheap fallback is one `'off'` branch in
`_updateLocalEchoState` (terminal-ui.js); teaching `PredictiveEchoAddon` pi's
composer row is the larger follow-up.
## Docker cases
The agent image (`docker/agent.Dockerfile`) installs pi in its own `RUN` step with
`--ignore-scripts`, kept out of the shared npm block so the flag cannot change how
the other four CLIs install. Rebuild with:
```bash
node scripts/build-agent-image.mjs --no-cache # --no-cache is mandatory
```
Credentials are **seeded**, not shared: `~/.pi/agent/auth.json`, `settings.json`,
`trust.json`, `models.json` and `models-store.json` are mounted read-only and
copied into the container's own `~/.pi/agent`. So an in-container pi never writes
refreshed OAuth tokens back to the host, and `docker commit` exports stay
secret-free. `models.json` is in the list because it holds user-defined custom
providers, which would otherwise silently vanish inside containers.
Only those five files are seeded because `~/.pi/agent` also holds `sessions/`,
`extensions/`, `skills/` and the installed package trees (`npm/`, `git/`), which
on an active host is easily gigabytes.
**Trade-off:** in-container pi sessions are invisible host-side, so `pi -c` inside
a Docker case only sees that container's own history.
## Remote SSH cases
`pi` mode is routed through an interactive login shell
(`exec "$SHELL" -i -l -c 'pi'`), because sshd's remote-command PATH does not
include npm's global bin on most hosts. Per-session config and `envOverrides` do
not cross ssh and are rejected rather than silently ignored; use the per-host
command override instead.
## Known gaps
- **No idle/completion hook.** Pi has no hook system Codeman can install into, so
idle detection falls back to output-stabilization like the other external CLIs.
Pi 0.84.0 shipped an `agent_settled` extension event that is a genuine idle
signal; a Codeman pi extension using it is the highest-value follow-up.
- **No response viewer.** Pi writes JSONL v3 session files under
`~/.pi/agent/sessions/`; nothing reads them yet.
- **Cron jobs mis-detect readiness.** The cron readiness poll looks for `❯` or a
token count, neither of which pi prints, so a pi cron job burns its poll budget
and then sends the prompt anyway. It works; it is just slower to start.
- **Ralph, respawn heuristics, token/CLI-info parsing and the `❯` readiness probe
are off** for pi, as for every external CLI.
-143
View File
@@ -1,143 +0,0 @@
# Predictive write-through echo for codex
Zero-lag local echo for codex sessions via a second, mosh-style mode in the
`xterm-zerolag-input` package: every keystroke goes to the PTY exactly as the
1.12.2 overlay-disabled path did (byte-identical wire behavior), while a
`PredictiveEchoAddon` simultaneously paints the predicted glyph at the predicted
cell. When the real echo lands, the prediction is confirmed and its span removed
(invisible swap: identical glyph beneath). Mispredictions drop via a mismatch
cascade + TTL. Visual-only, self-healing.
## Why this exists
Issues #218/#219/#220/#222 (one root cause) forced 1.12.2 to disable the
LocalEchoOverlay for codex: buffer-until-Enter starves codex's per-keystroke TUI
(live slash picker, arrows editing server-side composer state, composer
rewrap/growth, paste_burst classification). Buffer mode is structurally
incompatible with codex; write-through prediction is the only echo mode that
can coexist with it.
## The reconciliation lesson (do not regress this)
`docs/local-echo-overlay-plan.md` ("What NOT to Do") documented that matching
predictions against the raw output STREAM fails against Ink/TUI full-line
redraws. This design reads the parsed terminal BUFFER instead (cells after
xterm's parser ran), which converges to the same cells no matter how the bytes
arrived. The Phase 0 recordings prove the point twice over: tmux converts
codex's full-line redraws into minimal in-place deltas (an echo arrives as
`e\x1b[K\x1b[20;80H...`), and codex itself paints word gaps with ECH+cursor-forward
instead of spaces. Stream matching can never survive that; buffer diffing does
not care.
## Phase 0 measurements (codex-cli 0.147.0 via tmux, 100x30, 2026-08-09)
Recorded with `scripts/dev/record-codex-frames.mjs` (production pipeline:
codex inside tmux `status off`, chunks passed through the same full strip
`session.ts _handleTerminalOutput()` applies to codex mode). Fixtures in
`packages/xterm-zerolag-input/test/fixtures/codex/`; replay/measure with
`scripts/dev/analyze-codex-frames.mjs <fixture>`.
| Question | Measured answer |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Composer signature | Cursor row starts `"› "` (U+203A + space), text begins col 2. Present when empty (placeholder), while typing, and while the slash picker filters. `CODEX_COMPOSER_ROW_RE = /^› /` |
| Composer text color | Plain default foreground, zero SGR around echoed chars. Span `foregroundColor` default (theme fg) is an exact match |
| Placeholder | Cycling hint text ("Use /skills...", "Improve documentation in @filename", ...) rendered AT the cursor cell. First prediction lands over placeholder glyphs: covered by the snapshot + cursor-advance rules |
| Wrap | Word-wrap near `cols - 2`; continuation rows are indented 2 spaces WITHOUT `› `. The gate therefore suppresses predictions on wrapped lines: deliberate fallback to real echo, wrap was the #220 ghost zone. `edgeMarginCells = 4` |
| Modal (trust dialog) | Cursor parks on `" Press enter to continue"`: no `› ` prefix, gate false, zero predictions painted while keystrokes still reach the PTY (the ghost eliminator) |
| Streaming | Error/reconnect bursts render above a re-rendered composer that keeps the `› ` signature; end-of-frame cursor parks at the insertion point (col 2 of the composer row). Confirms the cursor-advance confirm rule and the no-drop-on-baseY rule |
| Echo shape under tmux | tmux emits minimal deltas for simple echoes and full repaints for busy frames; both converge in the parsed buffer |
| Slash picker | Picker rows render below; the cursor row keeps the composer signature and advances per filter char, so predictions stay active while filtering (#222 surface) |
Constants decided at the Phase 0 gate: `CODEX_COMPOSER_ROW_RE = /^› /`,
`ttlMs = 1000`, `maxPending = 32`, `cursorGraceMs = 150`, `edgeMarginCells = 4`,
span colors = theme defaults, `underlinePredictions = false`.
## Algorithm
See `PredictiveEchoAddon` in
`packages/xterm-zerolag-input/src/predictive-echo-addon.ts`. Summary of the
rules and why each exists:
- **State**: ordered `PredictionRecord[]` (`seq`, `char`, `width`, cumulative
`offsetCells`, `snapshot` of the cell at predict time, `sentAt`,
`mismatches`), plus a run `_anchor {row, col}` captured when the outstanding
count goes 0 -> 1. Positions are FIXED at predict time; confirmation deletes
spans and never re-lays-out, so partial confirmation causes zero jitter.
- **predictChar(ch)** runs an inline reconcile first and re-anchors whenever
outstanding drains to zero (absorbs the echo-landed-between-keystrokes race).
Guards: dims present, cursor numbers present, `viewportY === baseY`,
`predictWhen` gate, single codepoint >= 0x20 (not 0x7f), width <= 2,
`maxPending`, edge margin. Returns false = suppressed; the consumer sends the
keystroke regardless.
- **Coordinate base is `baseY`**: xterm's `cursorY` is baseY-relative, so
absolute buffer line = `baseY + row`. `viewportY` would only coincide while
the scrolled-to-bottom guards hold; the addon never relies on that.
- **reconcile()** (debounced `onWriteParsed` microtask, inline in predictChar,
TTL timer): clears everything when scrolled up; off-anchor-row cursor
tolerated for `cursorGraceMs` then clears; PREFIX-ONLY confirm loop requiring
cell match AND cursor advanced past the record (prevents false confirms
against placeholder glyphs and makes identical in-place tmux repaints a
no-op); TWO-PASS mismatch rule (a cell that is neither snapshot nor predicted
char must persist across two passes before cascading the drop: a half-parsed
row on pass N is fully redrawn a few ms later); TTL drop of the stale suffix.
- **No drop on baseY change**: codex streams push lines to history while the
composer stays viewport-pinned; predictions are row-relative to the pinned
composer and remain valid (measured above).
- **Anchor hold** (added by the independent post-build review): after any wire
input whose cursor effect the display has not shown yet (backspace with
nothing outstanding = deleting echoed text, every 'clear'-classified input,
an IME/plain-paste 'text' commit, and the bypass send paths), new
predictions are suppressed until the next PARSED write. Anchoring on the
stale cursor painted ghosts one cell off ("tehh" on backspace-then-retype
within RTT), blank-neutral and therefore TTL-lived. Worst case is exactly
one unpredicted keystroke: its own echo is a write, which releases the hold.
- **predictBackspace()** pops the newest outstanding record (informational
return; the consumer forwards `\x7f` unconditionally). Deleting already-echoed
text renders at RTT in v1.
- **CJK/wide**: 2-cell spans, stacking by cumulative visual width, leading-cell
confirm. In Codeman, IME input never reaches the hook (`window.cjkActive`
returns from onData first); package support exists for other consumers.
## Integration map (Codeman)
- Policy: `_localEchoPolicy` (`'buffer' | 'predict' | 'off'`) computed at the
end of `_updateLocalEchoState()`; codex + `localEchoEnabled` -> `'predict'`
while `_localEchoEnabled` stays false (every 1.12.2 consumer unchanged).
- onData hook sits between the buffer block and Normal Mode, classifies via
`classifyPredictInput()` (pure, on `window.CodemanTerminalInput`), never
returns, try/catch-wrapped: the wire path below is byte-identical with the
predictor active, absent, or throwing.
- Composer gate: `isCodexComposerRow()` set via `setPredictWhen()` at
construction (the vendor footer stays package-agnostic).
- Second vendor bundle `vendor/xterm-predictive-echo.js` (postinstall + build);
the zerolag bundle build command is untouched and its output byte-identical.
Missing/broken bundle = plain 1.12.2 echo (`typeof PredictiveEchoOverlay ===
'undefined'` guard).
- Prediction clears on: tab switch, SSE reconnect init, `insertTerminalText`,
`clearTerminalInput`, voice send, keyboard-accessory `sendKey`, resize, skin
and font changes re-read style via `refreshFont()`.
## Risk register
Eliminated structurally: other-mode regression (zero edits to buffer
addon/branches, byte-identical existing bundle, policy-matrix + byte-identity
tests); bundle breakage (separate bundle, graceful degradation); wire
corruption (no-return fall-through + try/catch + byte-identity pins at vm and
E2E level); modal ghosts (measured predictWhen gate); false confirms
(cursor-advance rule); mid-parse flicker drops (two-pass rule); wrap
misplacement (edge margin + continuation-row gate fallback + off-row grace).
Accepted residuals (visual-only, self-healing <= ttlMs, kill-switchable via
`localEchoEnabled` per device): no predictions on wrapped continuation lines
(gate false there, deliberate); brief dropout during composer growth; DOM-span
vs WebGL glyph rendering can differ subtly (same trade-off as the buffer
overlay, same font recipe); typing during an unsynchronized half-frame can
mis-anchor one run (mismatch/TTL cleans within 1s).
## Future work
RTT-adaptive TTL; mosh-style confidence gating (paint only after the link
proves laggy); predicted backspace into echoed text; predict mode for shell
prompts; unifying the small font/container duplication between the two addons
once predict mode has proven out; continuation-line prediction behind a
smarter composer-extent detector.
-723
View File
@@ -1,723 +0,0 @@
# QR Code Authentication Plan
> Ephemeral, single-use auth tokens embedded in the tunnel QR code — scan to auto-authenticate, while the bare tunnel URL stays password-protected.
## Problem
When the Cloudflare tunnel is active, anyone who discovers the `*.trycloudflare.com` URL can access Codeman (they just need the Basic Auth password, or if no password is set, full open access). The QR code currently encodes the raw tunnel URL — it provides no additional security. We want:
1. **Scanning the QR code** → seamless, instant access (no password prompt)
2. **Having only the URL** → blocked by Basic Auth (no access without credentials)
## Design
### Core Concept: Ephemeral Single-Use QR Tokens
The server maintains a rotating pool of short-lived, single-use tokens. The QR code encodes a short URL containing a lookup code that maps to the real token server-side. When scanned, the server validates the token, atomically consumes it, issues a session cookie, and redirects to `/`. The token is **not** the password — it's a separate, independent, ephemeral authentication pathway.
```
Desktop → displays QR (auto-refreshes every 60s via SSE)
QR Code → https://abc-xyz.trycloudflare.com/q/Xk9mQ3
Phone → scans, GET /q/Xk9mQ3
Server → looks up short code via Map (hash-based, timing-safe)
→ finds token record → validates TTL
→ atomically consumes token (single-use)
→ issues codeman_session cookie
→ 302 redirect to /
→ SSE push: new QR with embedded SVG for desktop display
→ desktop toast: "Device [IP] authenticated via QR"
→ audit log entry to session-lifecycle.jsonl
User → lands on app, fully authenticated
```
Someone who only has `https://abc-xyz.trycloudflare.com/` gets the standard Basic Auth prompt.
### Token Properties
| Property | Value |
|----------|-------|
| Length | 32 bytes (256 bits entropy) |
| Generation | `crypto.randomBytes(32).toString('hex')` |
| Short code | 6 chars base62, rejection-sampled (no modulo bias) |
| Short code derivation | Independent random generation (not derived from token) |
| Storage | In-memory `Map<shortCode, QrTokenRecord>` (no disk persistence) |
| TTL | 60 seconds (auto-rotation via timer), 90s grace for previous token |
| Effective window | Up to 90 seconds for the previous token (documented, not hidden) |
| Usage | **Single-use** — atomically consumed on first valid scan |
| URL format | Short code in path (`/q/Xk9mQ3`), not query params |
| URL length | ~53-56 chars total — targets QR Version 4 (33x33) for fast scanning |
| Scope | Only valid when `CODEMAN_PASSWORD` is set (no point without auth) |
| Lookup | `Map.get()` — hash-based O(1), no timing side-channel |
### Why This Design?
**Why not embed the password directly?**
- Password would appear in browser history, Cloudflare edge logs, and URL bars
- Password can't be rotated independently from QR access
**Why not a long-lived multi-use token? (original design)**
- A static token is functionally a second password — if the QR image leaks (screenshot shared, shoulder surfing, Cloudflare logs), the attacker has permanent access
- The USENIX Security 2025 paper ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) found 47 of the top-100 websites vulnerable due to exactly this pattern — missing single-use enforcement and long-lived tokens were 2 of the 6 critical design flaws identified
**Why short codes in the URL path instead of query params?**
- Query params (`?t=TOKEN`) leak into browser history, address bar, `Referer` headers, and Cloudflare edge logs
- Path-based short codes (`/q/Xk9mQ3`) are opaque references — the real token never appears in URLs
- Short codes are 6-char base62 (62^6 = 56.8 billion combinations), sufficient for lookup since they're backed by the full 256-bit token for validation and rate-limited to 10 attempts/IP
- The short `/q/` path (vs `/qr-auth/`) saves 7 bytes, helping keep the QR at Version 4 (33x33 modules) instead of Version 5 (37x37) — faster scanning on budget phones
## Auth Flow Diagram
```
┌─────────────┐ scan QR ┌──────────────────────────────────────┐
│ Mobile │ ────────────→ │ GET /q/Xk9mQ3 │
│ Device │ │ │
└─────────────┘ │ 1. Auth middleware sees /q/ │
│ → skips Basic Auth check │
│ 2. Route handler: Map.get(shortCode) │
│ → hash-based lookup (timing-safe) │
│ 3. Checks TTL (90s grace for prev) │
│ → token not expired? │
│ 4. Checks consumed flag │
│ → not already used? │
│ 5. Atomically marks token consumed │
│ 6. Issues codeman_session cookie │
│ 7. 302 redirect to / │
│ 8. Audit log → session-lifecycle.jsonl│
│ 9. SSE push: tunnel:qrRegenerated │
│ → desktop refreshes QR (SVG inline)│
│ 10. Desktop toast: "Device auth'd" │
└──────────────────────────────────────┘
┌─────────────┐ replay URL ┌──────────────────────────────────────┐
│ Attacker │ ────────────→ │ GET /q/Xk9mQ3 │
│ (stale code) │ │ │
└─────────────┘ │ 1. Map.get(shortCode) → not found │
│ OR token consumed OR expired │
│ 2. Increment QR rate limit counter │
│ (separate from Basic Auth counter) │
│ 3. 401 Unauthorized │
└──────────────────────────────────────┘
┌─────────────┐ URL only ┌──────────────────────────────────────┐
│ Attacker │ ────────────→ │ GET / │
│ (no token) │ │ │
└─────────────┘ │ 1. Auth middleware checks cookie │
│ → no cookie │
│ 2. Checks Basic Auth header │
│ → no header │
│ 3. Returns 401 + WWW-Authenticate │
│ → Browser shows password popup │
└──────────────────────────────────────┘
```
## Implementation
### 1. Token Manager — `src/tunnel-manager.ts`
Add a `QrTokenRecord` type and token rotation logic to `TunnelManager`. The token rotates every 60 seconds. A consumed token is immediately replaced. Up to 2 tokens can be valid simultaneously (current + previous, to handle the race where someone scans right as rotation happens). The previous token has a 90s grace period (not a full extra 60s — only enough to cover the scan-during-rotation race).
**Design decisions from security review:**
- **Map-based lookup** (not array scan) — `Map.get()` uses hash-based O(1) lookup, eliminating timing side-channels from string comparison
- **Rejection sampling** for short codes — avoids modulo bias (`256 % 62 != 0` gives 25% overrepresentation for first 6 charset chars)
- **SVG cache** — stores generated QR SVG per rotation cycle to avoid regenerating on every `/api/tunnel/qr` poll
- **Separate rate limit counter** — QR auth failures tracked independently from Basic Auth failures
```typescript
import { randomBytes } from 'node:crypto';
interface QrTokenRecord {
token: string; // 64 hex chars (256 bits)
shortCode: string; // 6 chars base62 (for URL path)
createdAt: number; // Date.now()
consumed: boolean; // single-use flag
}
const QR_TOKEN_TTL_MS = 60_000; // 60 seconds
const QR_TOKEN_GRACE_MS = 90_000; // 90s grace for previous token (scan-during-rotation)
const SHORT_CODE_LENGTH = 6;
const QR_RATE_LIMIT_MAX = 30; // global rate limit across all IPs
const QR_RATE_LIMIT_WINDOW_MS = 60_000; // 1 minute window
/** Rejection-sampled short code generation — no modulo bias */
function generateShortCode(): string {
const chars = 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789';
const maxUnbiased = 248; // largest multiple of 62 that fits in a byte (248 = 62 * 4)
const result: string[] = [];
while (result.length < SHORT_CODE_LENGTH) {
const [byte] = randomBytes(1);
if (byte < maxUnbiased) result.push(chars[byte % 62]);
// else: discard and re-draw (rejection sampling)
}
return result.join('');
}
export class TunnelManager extends EventEmitter {
// Map-based lookup: shortCode → QrTokenRecord (timing-safe, no string comparison)
private qrTokensByCode = new Map<string, QrTokenRecord>();
private currentShortCode: string | null = null;
private rotationTimer: ReturnType<typeof setInterval> | null = null;
// SVG cache — regenerated only on token rotation, not per request
private cachedQrSvg: { shortCode: string; svg: string } | null = null;
// Global rate limit counter (separate from Basic Auth rate limiting)
private qrAttemptCount = 0;
private qrRateLimitResetTimer: ReturnType<typeof setInterval> | null = null;
constructor() {
super();
this.rotateToken();
this.rotationTimer = setInterval(() => this.rotateToken(), QR_TOKEN_TTL_MS);
this.qrRateLimitResetTimer = setInterval(() => { this.qrAttemptCount = 0; }, QR_RATE_LIMIT_WINDOW_MS);
}
private rotateToken(): void {
const record: QrTokenRecord = {
token: randomBytes(32).toString('hex'),
shortCode: generateShortCode(),
createdAt: Date.now(),
consumed: false,
};
// Evict expired tokens from the Map
const now = Date.now();
for (const [code, rec] of this.qrTokensByCode) {
if (now - rec.createdAt > QR_TOKEN_GRACE_MS || rec.consumed) {
this.qrTokensByCode.delete(code);
}
}
this.qrTokensByCode.set(record.shortCode, record);
this.currentShortCode = record.shortCode;
this.cachedQrSvg = null; // invalidate SVG cache
this.emit('qrTokenRotated');
}
/** Get the current (newest) token's short code for QR URL */
getCurrentShortCode(): string | undefined {
return this.currentShortCode ?? undefined;
}
/** Get cached QR SVG, regenerating only if the short code changed */
async getQrSvg(tunnelUrl: string): Promise<string> {
const code = this.currentShortCode;
if (!code) throw new Error('No QR token available');
if (this.cachedQrSvg?.shortCode === code) return this.cachedQrSvg.svg;
const QRCode = require('qrcode');
const svg = await QRCode.toString(`${tunnelUrl}/q/${code}`, { type: 'svg', margin: 2, width: 256 });
this.cachedQrSvg = { shortCode: code, svg };
return svg;
}
/**
* Validate and atomically consume a token by short code.
* Returns { success, ip?, ua? } for audit logging on success.
* Map.get() is hash-based — no timing side-channel from string comparison.
*/
consumeToken(shortCode: string): boolean {
// Global rate limit (across all IPs)
if (this.qrAttemptCount >= QR_RATE_LIMIT_MAX) return false;
this.qrAttemptCount++;
const record = this.qrTokensByCode.get(shortCode);
if (!record) return false;
if (record.consumed) return false;
const now = Date.now();
if (now - record.createdAt > QR_TOKEN_GRACE_MS) return false;
// Atomic consume (single-threaded JS = no race)
record.consumed = true;
// Immediately rotate so desktop gets a fresh QR
this.rotateToken();
this.emit('qrTokenRegenerated');
return true;
}
/** Force-regenerate (manual revocation via API) */
regenerateQrToken(): void {
// Invalidate all existing tokens
this.qrTokensByCode.clear();
this.currentShortCode = null;
this.rotateToken();
this.emit('qrTokenRegenerated');
}
stopRotation(): void {
if (this.rotationTimer) {
clearInterval(this.rotationTimer);
this.rotationTimer = null;
}
if (this.qrRateLimitResetTimer) {
clearInterval(this.qrRateLimitResetTimer);
this.qrRateLimitResetTimer = null;
}
}
}
```
### 2. Auth Middleware Bypass — `src/web/middleware/auth.ts`
Add `/q/` to the bypass list (same pattern as `/api/hook-event`). The route handler itself handles token validation and rate limiting.
```typescript
// In the onRequest hook, add before Basic Auth check:
if (req.url.startsWith('/q/')) {
done(); // Let the route handler deal with token validation
return;
}
```
**Important**: Unlike `/api/hook-event` (localhost-only), `/q/` must be reachable from any IP (remote devices scan the QR). Rate limiting is handled by two independent mechanisms:
1. **Per-IP rate limit** — reuses the `authFailures` StaleExpirationMap (10 attempts/IP/15min), but tracked via a **separate counter** from Basic Auth failures (so a user who fat-fingers their password doesn't burn their QR attempts)
2. **Global path rate limit** — `TunnelManager.qrAttemptCount` caps total QR attempts to 30/minute across all IPs, defending against distributed brute force
### 3. Auto-Auth Route — `src/web/routes/system-routes.ts`
Add `GET /q/:code` as a top-level route (not under `/api/`):
```typescript
app.get('/q/:code', async (req, reply) => {
const shortCode = (req.params as { code: string }).code;
const authPassword = process.env.CODEMAN_PASSWORD;
// No point if auth isn't enabled
if (!authPassword) {
return reply.redirect('/');
}
// Per-IP rate limit (separate counter from Basic Auth failures)
const clientIp = req.ip;
const qrFailures = ctx.authState.qrAuthFailures?.get(clientIp) ?? 0;
if (qrFailures >= 10) {
return reply.code(429).send('Too Many Requests');
}
// Validate and atomically consume the token
// consumeToken() also checks the global rate limit (30/min across all IPs)
if (!shortCode || !ctx.tunnelManager.consumeToken(shortCode)) {
ctx.authState.qrAuthFailures?.set(clientIp, qrFailures + 1);
return reply.code(401).send('Invalid or expired QR code');
}
// Issue session cookie (same as Basic Auth success path)
const sessionToken = randomBytes(32).toString('hex');
const clientUA = req.headers['user-agent'] ?? '';
ctx.authState.authSessions?.set(sessionToken, {
ip: clientIp,
ua: clientUA,
createdAt: Date.now(),
});
ctx.authState.qrAuthFailures?.delete(clientIp);
// Audit log — write to session-lifecycle.jsonl for forensic analysis
ctx.lifecycleLog?.append({
event: 'qr_auth',
ip: clientIp,
ua: clientUA,
timestamp: Date.now(),
shortCodePrefix: shortCode.slice(0, 3) + '***', // partial for privacy
});
reply.setCookie(AUTH_COOKIE_NAME, sessionToken, {
httpOnly: true,
secure: ctx.https,
sameSite: 'lax',
maxAge: 86400, // 24h
path: '/',
});
// Broadcast auth notification — desktop sees who authenticated (QRLjacking detection)
broadcast('tunnel:qrAuthUsed', {
ip: clientIp,
ua: clientUA,
timestamp: Date.now(),
});
return reply.redirect('/');
});
```
### 4. Update QR Code URL — `src/web/routes/system-routes.ts`
Modify `/api/tunnel/qr` to encode the short-code URL. Uses the `TunnelManager.getQrSvg()` cache — SVG is regenerated only when the token rotates, not on every request.
```typescript
app.get('/api/tunnel/qr', async (_req, reply) => {
const url = ctx.tunnelManager.getUrl();
if (!url) {
return reply.code(404).send(createErrorResponse(ApiErrorCode.NOT_FOUND, 'Tunnel not running'));
}
const authPassword = process.env.CODEMAN_PASSWORD;
// If auth is enabled, use the cached SVG with embedded short code
if (authPassword) {
const svg = await ctx.tunnelManager.getQrSvg(url);
return { svg, authEnabled: true };
}
// No auth — just encode the raw tunnel URL
const QRCode = require('qrcode');
const svg = await QRCode.toString(url, { type: 'svg', margin: 2, width: 256 });
return { svg, authEnabled: false };
});
```
### 5. Token Regeneration Endpoint — `src/web/routes/system-routes.ts`
Manual revocation — invalidates ALL existing tokens and creates a fresh one:
```typescript
app.post('/api/tunnel/qr/regenerate', async () => {
ctx.tunnelManager.regenerateQrToken();
return { success: true };
});
```
### 6. Frontend Updates — `src/web/public/app.js`
#### QR Overlay Changes
- **Auto-refresh via inline SVG**: Listen for `tunnel:qrRotated` SSE events which now include the SVG directly in the payload — no extra HTTP fetch needed, sub-50ms refresh on desktop.
- **Countdown indicator**: Small "expires in Xs" text under the QR that counts down from 60. Reassures the user the QR is live and not stale.
- **Regenerate button**: "Regenerate QR" button. Calls `POST /api/tunnel/qr/regenerate` — SSE event delivers the new SVG.
- **Auth badge**: Lock icon or "Single-use auth" label when auth is active.
- **URL display**: Show the raw tunnel URL (not the auth URL) for manual copy — users who copy the URL authenticate via Basic Auth. The QR is the fast path.
- **Auth notification toast**: When `tunnel:qrAuthUsed` fires, show a 10-second toast: "Device [IP] authenticated via QR (Safari). Not you? [Revoke]". This is the primary QRLjacking detection mechanism (USENIX Flaw-5).
```javascript
// Auto-refresh QR on rotation — SVG is inline in the event payload
addListener('tunnel:qrRotated', (data) => {
if (data.svg) {
updateQrDisplay(data.svg); // direct DOM update, no fetch
} else {
refreshTunnelQR(); // fallback: fetch from API
}
});
// Also refresh on manual regeneration
addListener('tunnel:qrRegenerated', (data) => {
if (data.svg) {
updateQrDisplay(data.svg);
} else {
refreshTunnelQR();
}
});
// QRLjacking detection — notify desktop user when QR is consumed
addListener('tunnel:qrAuthUsed', (data) => {
showNotificationToast(
`Device authenticated via QR (${parseUAFamily(data.ua)}, ${data.ip}). Not you?`,
{
duration: 10000,
action: { label: 'Revoke', onClick: () => revokeAllSessions() },
}
);
});
// In showTunnelQR(), after fetching /api/tunnel/qr:
if (data.authEnabled) {
const badge = document.createElement('div');
badge.textContent = 'Single-use auth \u00b7 refreshes every 60s';
badge.style.cssText = 'margin-top:8px;font-size:11px;color:var(--text-secondary)';
container.parentElement.appendChild(badge);
}
```
#### Welcome Screen QR
Same auto-refresh behavior applies to `_updateWelcomeTunnelBtn()` — the QR is fetched from `/api/tunnel/qr` so token embedding happens automatically.
### 7. SSE Events
Three events for the frontend. QR rotation events embed the SVG directly in the payload to eliminate an extra HTTP fetch — the desktop gets the new QR in a single SSE push (~2-5KB SVG, well within SSE limits).
```typescript
// In server.ts, listen for tunnelManager events:
// Auto-rotation every 60s — desktop refreshes QR silently (SVG inline)
tunnelManager.on('qrTokenRotated', async () => {
const url = tunnelManager.getUrl();
if (url && process.env.CODEMAN_PASSWORD) {
const svg = await tunnelManager.getQrSvg(url);
broadcast('tunnel:qrRotated', { svg });
} else {
broadcast('tunnel:qrRotated', {});
}
});
// Manual regeneration or post-consumption — desktop refreshes QR (SVG inline)
tunnelManager.on('qrTokenRegenerated', async () => {
const url = tunnelManager.getUrl();
if (url && process.env.CODEMAN_PASSWORD) {
const svg = await tunnelManager.getQrSvg(url);
broadcast('tunnel:qrRegenerated', { svg });
} else {
broadcast('tunnel:qrRegenerated', {});
}
});
// QR auth consumed — desktop shows notification toast (QRLjacking detection)
// Note: this is broadcast from the route handler, not tunnelManager
// Event: tunnel:qrAuthUsed { ip, ua, timestamp }
```
### 8. Session Cookie Binding & Revocation
Enhance session records to include device context for audit purposes. The UA is stored for **logging only** — not for blocking.
**Why no UA-family blocking (`majorUAChanged`)?** Security review found this is security theater:
- UA strings are trivially spoofable by any attacker who can steal a cookie
- Chrome UA reduction (2022+) makes family detection unreliable
- Mobile WebView → browser switches trigger false positives on the same device
- HttpOnly + Secure + SameSite=lax + 24h TTL already protect against cookie theft
- The attacker who can exfiltrate a cookie can also replay the exact UA
Instead, provide **manual session revocation** as the active defense:
```typescript
// Session record stores device context for audit logging (not blocking):
ctx.authState.authSessions?.set(sessionToken, {
ip: clientIp,
ua: req.headers['user-agent'] ?? '',
createdAt: Date.now(),
method: 'qr', // 'qr' | 'basic' — tracks how session was created
});
// Manual revocation endpoint — kill specific session or all sessions
app.post('/api/auth/revoke', async (req, reply) => {
const { sessionToken: target } = req.body as { sessionToken?: string };
if (target) {
ctx.authState.authSessions?.delete(target);
} else {
// Revoke all sessions (nuclear option)
ctx.authState.authSessions?.clear();
}
return { success: true };
});
```
**Note**: This is a breaking type change. The `AuthState` interface must be updated from `StaleExpirationMap<string, string>` (token → clientIp) to `StaleExpirationMap<string, { ip, ua, createdAt, method }>`. All session validation code in `auth.ts` must be updated simultaneously.
### 9. Cleanup — `src/tunnel-manager.ts`
Stop the rotation timer in the `stop()` method:
```typescript
stop(): void {
this.stopRotation();
// ... existing cleanup
}
```
## Security Analysis
### Threat Model
| Threat | Attack Vector | Mitigation | Residual Risk |
|--------|--------------|------------|---------------|
| **QR screenshot shared** | Attacker gets image of QR code | Single-use: token consumed on first scan. 60s TTL: expired by the time attacker tries. Desktop toast notification alerts user if someone else scans. | If attacker scans faster than legitimate user (~seconds), they win the race. Low risk: requires physical proximity + speed. User sees notification and can revoke. |
| **Cloudflare edge logs** | Cloudflare logs the full URL path | Short code is opaque (6-char lookup key), not the real token. Single-use: replaying from logs always fails. 60s TTL (90s grace): expired before log review. `trycloudflare.com` quick tunnels have no customer-accessible logging controls — the privacy implications are inherent to using free quick tunnels. | Cloudflare has TLS termination access regardless. Ephemeral short codes are far less valuable than a permanent token. |
| **Brute force short code** | Attacker guesses `/q/XXXXXX` | Per-IP rate limiting (10/IP/15min) + global path rate limit (30/min across all IPs). 62^6 = 56.8 billion combinations. Only ~2 valid codes at any time. | Infeasible: expected guesses to hit = ~2.8×10^10, rate limits block well before. |
| **Replay attack** | Reuse a previously valid URL | Single-use consumption + 60s TTL (90s grace). Old codes always 401. | None — replay is impossible by design. |
| **QRLjacking** | Attacker displays your QR on phishing site | No companion app = limited mitigation. However: 60s rotation means attacker must relay in real-time. Desktop toast notification ("Device [IP] authenticated via QR. Not you? [Revoke]") provides real-time detection. Self-hosted single-user context makes phishing implausible. | Theoretical risk for multi-user deployments. Mitigated by notification toast for single-user. Note: Signal's linked-device QR flow was exploited by Russian state actors (UNC5792/Sandworm) via quishing in 2025 — but that targeted a multi-user messaging platform, not a self-hosted dev tool. |
| **Session cookie theft** | XSS or network sniffing steals cookie | HttpOnly + Secure flags. SameSite=lax prevents CSRF. 24h TTL limits exposure window. Manual revocation via `/api/auth/revoke`. | Standard web cookie risks apply. Mitigated by security headers (CSP, etc.). |
| **Token in server logs** | Access log captures URL path | Log `/q/*` with short code masked or omitted. Configure Fastify logger to redact `/q/` paths. | Path still appears in server access logs (mitigated by masking). |
| **Timing attack** | Measure response time to leak short code | Map-based lookup (`Map.get()`) — hash-based O(1), no character-by-character timing leak. No string comparison in the hot path. | None — timing side channel eliminated by design. |
| **Token not in query params** | N/A (this is a mitigation) | Short code in URL path avoids browser history, Referer headers, and address bar exposure. | Path still appears in server access logs (mitigated by masking). |
| **Distributed brute force** | Multiple IPs guess codes simultaneously | Global rate limit (30/min total across all IPs) in addition to per-IP limit. | Infeasible given keyspace. Global limit prevents botnet-scale attempts. |
| **CSRF on regenerate** | Cross-origin POST to `/api/tunnel/qr/regenerate` | SameSite=lax cookies are NOT sent with cross-origin POST requests, providing CSRF protection. Endpoint requires authenticated session. | Verify SameSite=lax behavior through cloudflared tunnel. |
### USENIX Security 2025 Flaw Coverage
The [Zhang et al. paper](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) (USENIX Security 2025, 47 of top-100 websites vulnerable, 42 CVEs) identified 6 critical design flaws. Coverage:
| USENIX Flaw | Status | Implementation |
|-------------|--------|----------------|
| Flaw-1: Missing single-use enforcement | **Fixed** | Atomic `consumed` flag, Map-based lookup |
| Flaw-2: Long-lived tokens | **Fixed** | 60s TTL, 90s grace, auto-rotation |
| Flaw-3: Predictable QrId generation | **Fixed** | `crypto.randomBytes(32)` — 256-bit entropy, rejection-sampled short codes |
| Flaw-4: Client-side QrId generation | **Fixed** | Server-side generation only |
| Flaw-5: Missing status notification | **Fixed** | Desktop toast notification via `tunnel:qrAuthUsed` SSE event. Shows device IP/UA with [Revoke] button. |
| Flaw-6: Inadequate session binding | **Partial** | IP + UA stored for audit. No cryptographic channel binding (requires companion app / FIDO2 — overkill for single-user). Manual revocation as active defense. |
### Industry Comparison
| Platform | Model | How This Plan Compares |
|----------|-------|----------------------|
| **Discord** | Long-lived session token, no confirmation, repeatedly exploited via QRLjacking | **Better** — single-use + TTL + notification toast |
| **WhatsApp Web** | Pre-authenticated phone confirms "Link device?", ~60s rotation | **Comparable** rotation model; missing WhatsApp's explicit confirmation prompt (acceptable: single-user, no account selection) |
| **Signal** | Ephemeral public key in QR, E2E encrypted channel via Signal protocol | **Below** — no cryptographic channel binding. Note: Signal's QR flow was exploited by state actors in 2025 despite stronger crypto, showing that protocol strength alone doesn't prevent social engineering. |
| **1Password** | Noise framework E2E channel, post-quantum pre-shared keys, confirmation codes | **Below** — but 1Password is a credential manager with different threat model. Overkill for a dev tool. |
| **FIDO2 CTAP 2.2** | BLE proximity + cryptographic binding + biometric verification | **Below** — but requires BLE stack, FIDO server, and companion authenticator. Completely inappropriate here. |
### Comparison to Prior Design
| Property | Original Plan | Current Plan |
|----------|--------------|--------------|
| Token TTL | Infinite (until restart) | 60 seconds (90s grace for previous token) |
| Reuse | Multi-use (same QR works forever) | Single-use (consumed atomically on first scan) |
| Secret in URL | Query param (`?t=64-char-hex`) | Opaque short code in path (`/q/Xk9mQ3`) |
| Leak impact | Permanent access until manual revoke | Worthless after first use or 90s, whichever comes first |
| Desktop QR refresh | Manual only | Auto-refresh every 60s via SSE with inline SVG |
| Session binding | IP only | IP + UA stored for audit (not blocking). Manual revocation endpoint. |
| Auth notification | None | Desktop toast: "Device [IP] authenticated via QR. Not you? [Revoke]" |
| Audit logging | None | `session-lifecycle.jsonl` entry on every QR auth event |
| Rate limiting | Per-IP only, shared with Basic Auth | Per-IP (separate counter) + global path limit (30/min) |
| Short code generation | Modulo-biased | Rejection-sampled (no bias) |
| Short code lookup | Array scan (timing leak) | Map-based O(1) (timing-safe) |
| Connect latency | ~50ms (localhost only) | ~150-300ms through Cloudflare tunnel (honest estimate) |
### What This Does NOT Protect Against
- **FIDO2/passkey-level phishing resistance**: Would require BLE proximity verification and cryptographic channel binding. Overkill for a self-hosted single-user dev tool. The FIDO2 CTAP 2.2 hybrid transport is the gold standard but requires BLE hardware and a companion authenticator.
- **Compromised phone**: If the attacker has physical access to the phone that scans, no QR scheme helps.
- **Compromised Cloudflare tunnel**: Cloudflare terminates TLS and can inspect all traffic. This is inherent to using `trycloudflare.com` quick tunnels — use `--https` for end-to-end encryption if this matters.
- **State-sponsored quishing**: Sophisticated attackers could create convincing phishing pages that relay the QR in real-time. The 60s rotation and desktop notification toast mitigate this for the single-user case, but a dedicated attacker with social engineering could theoretically succeed within the TTL window.
### Standards Compliance Note
This design is **inspired by but does not conform to** [OASIS SQRAP v1.0](https://docs.oasis-open.org/esat/sqrap/v1.0/cs01/sqrap-v1.0-cs01.html). SQRAP's architecture requires a companion mobile app with stored identity keys, public key channel binding, back-channel authentication, and user presence verification (biometric/PIN). These are fundamentally incompatible with a browser-scan-to-authenticate flow. SQRAP is referenced for awareness of formal QR auth standards, not as a compliance target.
## Performance
The design prioritizes speed on connect. Latency depends on whether the request goes through a Cloudflare tunnel or is localhost:
### Localhost (no tunnel)
| Step | Latency |
|------|---------|
| QR scan (physical) | ~1-2s (user action) |
| `GET /q/:code` → Map.get() lookup + consume | <1ms |
| Cookie set + 302 redirect | <1ms |
| Browser follows redirect to `/` | <5ms |
| **Total (after scan)** | **<10ms** |
### Through Cloudflare Tunnel (typical mobile use case)
Each request traverses: phone → Cloudflare edge (TLS termination) → cloudflared → localhost. The 302 redirect means **two full round trips** through the tunnel.
| Step | Latency |
|------|---------|
| QR scan (physical) | ~1-2s (user action) |
| DNS resolution for `*.trycloudflare.com` | 20-80ms (first request, cached after) |
| TLS handshake to Cloudflare edge | 50-100ms (first request, 0 with TLS resumption) |
| `GET /q/:code` through tunnel (request + response) | 30-90ms |
| Browser follows 302 redirect: `GET /` through tunnel | 30-90ms |
| **Total first connection (cold)** | **~200-400ms** |
| **Total subsequent (TLS/DNS cached)** | **~100-200ms** |
This is still fast — **imperceptible after the 1-2s physical QR scan action**. For comparison, VS Code Remote Tunnels (through Azure) adds 20-100ms per hop.
### Why Not Eliminate the Redirect?
The 302 means two round trips. Alternatives considered:
- **200 + serve `index.html` directly**: URL bar shows `/q/Xk9mQ3`, relative paths break, couples auth to static serving. Not worth the complexity.
- **200 + `<meta http-equiv="refresh">`**: Still two requests, plus HTML parse delay. Actually slower.
- **200 + JavaScript redirect**: Same problem, plus fails if JS disabled.
The 302 is clean, universally supported, and the extra 30-90ms is invisible to users.
### QR Code Size Optimization
The URL `https://xxx-yyy.trycloudflare.com/q/Xk9mQ3` is ~53-56 characters. At QR Error Correction Level M:
| QR Version | Grid Size | Byte Capacity | Fits? |
|------------|-----------|---------------|-------|
| Version 3 | 29x29 | 42 bytes | No |
| Version 4 | 33x33 | 62 bytes | Yes (comfortably) |
| Version 5 | 37x37 | 84 bytes | Yes |
The shortened `/q/` path (vs `/qr-auth/`) and 6-char code (vs 8-char) save 9 bytes, targeting Version 4 (33x33) for faster scanning on budget Android phones. Modern phones scan Version 4 QR codes in 100-300ms — the user action of pointing the camera dominates.
### Desktop QR Refresh
Token rotation SSE events now embed the SVG directly in the payload (~2-5KB). The desktop gets the new QR in a single SSE push — no extra HTTP fetch needed. Refresh latency: **sub-50ms** (SSE adaptive batching at 16-50ms).
### SVG Caching
QR SVG is cached per rotation cycle on `TunnelManager.cachedQrSvg`. The SVG is regenerated only when the token rotates (every 60s), not on every `/api/tunnel/qr` request. SVG format is optimal: resolution-independent (retina-safe), inline-able (no extra HTTP request), ~2-5KB, renders in <1ms.
## Edge Cases
1. **Scan during rotation**: The server keeps 2 tokens (current + previous). If the user scans right as rotation happens, the previous token is still valid for up to 60s more. Seamless.
2. **Server restart**: All tokens cleared (in-memory). New token generated immediately. Tunnel URL also changes (trycloudflare gives a new subdomain), so old QR codes are doubly dead.
3. **Multiple devices**: Each scan consumes the token and triggers a fresh one. To auth a second device, wait for the QR to refresh (≤60s) or hit "Regenerate QR" on the desktop, then scan the new code.
4. **Token without tunnel**: `/qr-auth/:code` works even on localhost. If you have the code and it's valid, you get authenticated regardless of access method.
5. **Tunnel restart (same server)**: Tokens survive tunnel restarts (stored on `TunnelManager` instance). But new tunnel URL = new QR code generated. Short code stays valid until consumed or expired.
6. **Desktop browser closed during scan**: Token is consumed server-side. The scanning phone gets authenticated. When the desktop reopens, SSE reconnects and shows a fresh QR. No state corruption.
7. **Race condition: two phones scan same QR**: First scanner wins (atomic `consumed = true`). Second scanner gets 401. This is correct behavior — single-use by design.
## Files to Modify
| File | Changes |
|------|---------|
| `src/tunnel-manager.ts` | `QrTokenRecord` type, `Map<shortCode, record>` token pool, rejection-sampled `generateShortCode()`, rotation timer, `consumeToken()`, `getCurrentShortCode()`, `getQrSvg()` (cached), `regenerateQrToken()`, global rate limit counter, cleanup in `stop()` |
| `src/web/middleware/auth.ts` | Add `/q/` bypass in `onRequest` hook. Enhance session record type from `string` to `{ ip, ua, createdAt, method }` (**breaking type change** — all consumers must update). Add `qrAuthFailures` StaleExpirationMap (separate from Basic Auth `authFailures`). |
| `src/web/routes/system-routes.ts` | Modify `/api/tunnel/qr` to use `getQrSvg()` cache. Add `GET /q/:code` with atomic consume, audit log, and `tunnel:qrAuthUsed` broadcast. Add `POST /api/tunnel/qr/regenerate`. Add `POST /api/auth/revoke`. |
| `src/web/server.ts` | Pass `authState` + `lifecycleLog` to route context. Listen for `qrTokenRotated` and `qrTokenRegenerated` events → broadcast SSE with inline SVG. |
| `src/web/public/app.js` | Auto-refresh QR from inline SSE SVG payload (no extra fetch). Countdown timer. Regenerate button. Auth badge. Auth notification toast on `tunnel:qrAuthUsed` with [Revoke] action. |
| `src/session-lifecycle-log.ts` | Add `qr_auth` event type to lifecycle log schema |
| `src/types/api.ts` | Update `AuthState` interface: `authSessions` value type, add `qrAuthFailures` map |
## Complexity Estimate
Medium change. Core logic (Map-based token pool, rejection-sampled short codes, SVG cache, atomic consumption, cookie issuance, audit logging) is ~120 lines. Rate limiting (separate QR counter + global path limit) adds ~20 lines. SSE plumbing with inline SVG adds ~30 lines. Frontend (inline SVG refresh, auth notification toast with revoke, countdown) is ~40 lines. Auth type migration (session record type change) touches ~10 lines across middleware. No new dependencies — `crypto` and `qrcode` are already available.
## Testing
### Automated
```bash
# Unit test for token manager
npx vitest run test/qr-auth.test.ts
```
Test cases:
- Token rotation generates unique short codes (6-char, base62)
- Short codes have uniform character distribution (no modulo bias — verify with chi-squared test over 10K samples)
- `consumeToken()` returns true on first use, false on second
- Expired tokens (>90s old) return false
- Previous token still works during 90s grace period
- Token at exactly 60s still valid (within grace), token at 91s rejected
- `regenerateQrToken()` invalidates all existing tokens (Map cleared)
- Short code lookup is case-sensitive
- Per-IP rate limiting increments on invalid codes (separate from Basic Auth counter)
- Global rate limit (30/min) blocks attempts across all IPs
- SVG cache returns same string for same short code, regenerates on rotation
- Audit log entry written on successful QR auth
- `tunnel:qrAuthUsed` SSE event broadcast on successful QR auth
- `tunnel:qrRotated` SSE event includes inline SVG payload
- Map-based lookup does not leak timing information (no string comparison in hot path)
### Manual
1. Start server with `CODEMAN_PASSWORD=test`
2. Enable tunnel
3. Verify `/api/tunnel/qr` returns QR encoding `https://...trycloudflare.com/q/Xk9mQ3`
4. Open the QR URL in incognito → should auto-redirect to `/` with session cookie
5. Verify desktop shows notification toast: "Device [IP] authenticated via QR"
6. Open the **same** URL again → should get 401 (single-use consumed)
7. Wait 60s → verify QR display auto-updated (new short code, inline SVG via SSE)
8. Open just the tunnel URL → should get Basic Auth prompt
9. Call `POST /api/tunnel/qr/regenerate` → old QR URL returns 401, new QR appears
10. Verify per-IP rate limiting: 10+ failed `/q/badcode` → 429
11. Verify Basic Auth failures don't consume QR rate limit budget (and vice versa)
12. Check `~/.codeman/session-lifecycle.jsonl` for `qr_auth` entries after successful scan
13. Click [Revoke] on the notification toast → verify session is invalidated
## References
- [USENIX Security 2025: "Demystifying the (In)Security of QR Code-based Login in Real-world Deployments"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) — 6 design flaws, 5 attack types, 42 CVEs across 47 of top-100 websites. Primary design reference for this plan.
- [OWASP QRLJacking](https://owasp.org/www-community/attacks/Qrljacking) — canonical QR session hijacking reference
- [OASIS SQRAP v1.0 Standard](https://docs.oasis-open.org/esat/sqrap/v1.0/cs01/sqrap-v1.0-cs01.html) — formal standard for secure QR authentication. **Not a compliance target** for this plan (requires companion app + PKI). Referenced for awareness only.
- [FIDO2 CTAP 2.2 Hybrid Transport](https://fidoalliance.org/specs/fido-v2.2-rd-20230321/fido-client-to-authenticator-protocol-v2.2-rd-20230321.html) — gold standard for cross-device auth (overkill for this use case)
- [Google GTIG: Signal QR quishing by Russian state actors (2025)](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger) — UNC5792/Sandworm exploited Signal's linked-device QR flow via phishing. Demonstrates that even cryptographically strong QR auth can be defeated by social engineering.
- [CVE-2026-2144: Magic Login QR Code Plugin race condition](https://www.cvedetails.com/cve/CVE-2026-2144/) — QR token stored as predictable static file, race window between creation and deletion. Validates this plan's in-memory-only approach.
-846
View File
@@ -1,846 +0,0 @@
# Ralph Wiggum Loop: Complete Guide
> This document consolidates official Anthropic documentation, community best practices, and implementation details for autonomous Claude Code loops.
**Last Updated**: 2026-01-24
**Sources**: [Official Anthropic Plugin](https://github.com/anthropics/claude-code/tree/main/plugins/ralph-wiggum), [Claude Code Docs](https://code.claude.com/docs/en/hooks), [Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
---
## Table of Contents
1. [Overview](#overview)
2. [Core Concept](#core-concept)
3. [Official Plugin Reference](#official-plugin-reference)
4. [The Promise Tag Contract](#the-promise-tag-contract)
5. [TodoWrite Tool Integration](#todowrite-tool-integration)
6. [Hooks System](#hooks-system)
7. [Best Practices](#best-practices)
8. [Prompt Templates](#prompt-templates)
9. [When to Use (and Not Use)](#when-to-use-and-not-use)
10. [Real-World Examples](#real-world-examples)
11. [Codeman Implementation](#codeman-implementation)
12. [Troubleshooting](#troubleshooting)
---
## Overview
Ralph Wiggum is an autonomous loop technique for Claude Code, named after The Simpsons character. It enables Claude to work iteratively on tasks for hours without human intervention, self-correcting until completion criteria are met.
**Core Philosophy**:
- **Iteration > Perfection**: Don't aim for perfect on first try; let the loop refine
- **Failures Are Data**: "Deterministically bad" means failures are predictable and informative
- **Operator Skill Matters**: Success depends on writing good prompts, not just having a good model
- **Persistence Wins**: Keep trying until success; the loop handles retry logic
**Origin**: Created by Geoffrey Huntley, formalized into an official Anthropic plugin by Boris Cherny (Head of Claude Code) in late 2025.
---
## Core Concept
The simplest form of a Ralph loop:
```bash
while :; do cat PROMPT.md | claude ; done
```
**How It Works**:
1. Claude processes a task prompt
2. Attempts to exit when "done"
3. A **Stop hook** intercepts the exit
4. Checks for **completion promise** (e.g., `<promise>COMPLETE</promise>`)
5. If not found, re-feeds the same prompt
6. Files from previous iteration persist, so Claude sees its own work
7. Cycle repeats until completion or max iterations reached
**Key insight**: The prompt never changes between iterations, but Claude's previous work persists in files, allowing autonomous improvement by reading past work.
---
## Official Plugin Reference
### Installation
```bash
# Add Anthropic's official plugin marketplace
/plugin marketplace add anthropics/claude-plugins-official
# Install Ralph Wiggum plugin
/plugin install ralph-wiggum@claude-plugins-official
```
### Commands
#### `/ralph-loop:ralph-loop`
Start an autonomous loop in the current session.
```bash
/ralph-loop:ralph-loop
```
When invoked, this skill prompts you to configure:
- **Task prompt**: The work to be done (persists across iterations)
- **Max iterations**: Safety limit on iterations (recommended: always set this)
- **Completion promise**: The phrase that signals completion (e.g., `COMPLETE`)
#### `/ralph-loop:cancel-ralph`
Cancel the active Ralph loop.
```bash
/ralph-loop:cancel-ralph
```
#### `/ralph-loop:help`
Show help and usage information.
```bash
/ralph-loop:help
```
### State File
The plugin persists state to `.claude/ralph-loop.local.md`:
```yaml
---
enabled: true
iteration: 5
max-iterations: 50
completion-promise: "COMPLETE"
---
# Original Prompt
Build a REST API for todos...
```
**YAML Fields**:
- `enabled` (boolean): Controls hook activation
- `iteration` (integer): Current iteration count (0-indexed)
- `max-iterations` (integer): Optional maximum
- `completion-promise` (string): Optional completion text
---
## The Promise Tag Contract
The completion phrase pattern is the core contract between Claude and the loop system:
```
<promise>PHRASE</promise>
```
**Examples**:
- `<promise>COMPLETE</promise>` - Generic completion
- `<promise>TESTS_PASS</promise>` - Test-specific completion
- `<promise>TIME_COMPLETE</promise>` - Time-aware loop completion
- `<promise>FIXED</promise>` - Bug fix completion
### How Completion Detection Works
1. **Exact String Matching**: The `--completion-promise` uses case-sensitive exact matching
2. **Output Scanning**: The Stop hook scans Claude's final output for the promise tag
3. **Exit Control**: If found, exit is allowed. If not, loop continues.
### False Positive Prevention
The official implementation (and Codeman) prevents false positives when completion phrases appear in:
- Initial prompts
- Documentation or examples
- Comments
**Solution**: Codeman uses **occurrence-based detection** to distinguish prompts from actual completions:
- **1st occurrence**: Store as expected phrase (likely in the prompt)
- **2nd occurrence**: Emit `completionDetected` (actual completion)
- **If loop already active**: Emit immediately (explicit loop start via `/ralph-loop:ralph-loop`)
```typescript
// From codeman/src/ralph-tracker.ts
private handleCompletionPhrase(phrase: string): void {
const count = (this._completionPhraseCount.get(phrase) || 0) + 1;
this._completionPhraseCount.set(phrase, count);
// Store phrase on first occurrence
if (!this._loopState.completionPhrase) {
this._loopState.completionPhrase = phrase;
this._loopState.lastActivity = Date.now();
this.emit('loopUpdate', this.loopState);
}
// Emit completion if loop is active OR this is 2nd+ occurrence
if (this._loopState.active || count >= 2) {
this._loopState.active = false;
this._loopState.lastActivity = Date.now();
this.emit('completionDetected', phrase);
this.emit('loopUpdate', this.loopState);
}
}
```
This approach handles both scenarios:
1. **Explicit loop start**: User runs `/ralph-loop:ralph-loop`, loop is active, first completion phrase triggers
2. **Implicit completion**: Phrase appears in prompt (1st), then Claude outputs it on completion (2nd)
---
## TodoWrite Tool Integration
The **TodoWrite tool** is Claude Code's built-in task management system that integrates with Ralph loops.
### How It Works
Claude uses TodoWrite to:
1. Break complex tasks into subtasks
2. Track progress through iterations
3. Provide visibility into current state
4. Resume work after context resets
### Todo Formats Detected
**Format 1: Markdown Checkboxes**
```markdown
- [ ] Pending task
- [x] Completed task
- [X] Completed task (uppercase)
```
**Format 2: Status Indicators**
```
Todo: ☐ Pending task
Todo: ◐ In progress task
Todo: ✓ Completed task
```
**Format 3: Parenthetical Status**
```
- Task name (pending)
- Task name (in_progress)
- Task name (completed)
```
**Format 4: Native Checkboxes (without "Todo:" prefix)**
```
☐ Pending task
◐ In progress task
☒ Completed task
```
**Format 5: Claude Code Checkmark-Based TodoWrite Output**
```
✔ Task #1 created: Fix the authentication bug
✔ #1 Fix the authentication bug
✔ Task #1 updated: status → in progress
✔ Task #1 updated: status → completed
```
This is the primary output format used by Claude Code's TodoWrite tool in CLI sessions. The tracker maps task numbers to content, allowing status updates to reference tasks by number.
### System Reminder Integration
From official Claude Code documentation:
> After commands like an ls -la run via bash tool, system-reminder tags are injected to remind the model to use the TodoWrite tool if it hasn't been using it so far.
The system prompt includes:
> "IMPORTANT: Always use the TodoWrite tool to plan and track tasks throughout the conversation."
### Checklists for Complex Workflows
From [Anthropic Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices):
> For large tasks with multiple steps or requiring exhaustive solutions—like code migrations, fixing numerous lint errors, or running complex build scripts—improve performance by having Claude use a Markdown file (or even a GitHub issue!) as a checklist and working scratchpad.
---
## Hooks System
Ralph loops are powered by Claude Code's hooks system. Understanding hooks is essential for customization.
### Hook Events Reference
| Event | When | Use Case |
|-------|------|----------|
| `PreToolUse` | Before tool execution | Validate, modify, or block tool calls |
| `PostToolUse` | After tool completes | Provide feedback, run formatters/linters |
| `Stop` | When Claude finishes | **Ralph loop control** - block exit, refeed prompt |
| `SubagentStop` | When subagent finishes | Control nested loops |
| `UserPromptSubmit` | User submits prompt | Add context, validate input |
| `SessionStart` | Session begins | Load environment, context |
| `SessionEnd` | Session ends | Cleanup, logging |
| `PermissionRequest` | Permission dialog shown | Auto-approve/deny |
| `PreCompact` | Before compact | Backup, preprocessing |
### Stop Hook for Ralph Loops
The Stop hook is the key mechanism:
```json
{
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "./scripts/ralph-stop-hook.sh"
}
]
}
]
}
}
```
**Stop Hook Logic**:
1. Check if `.claude/ralph-loop.local.md` exists
2. Read `enabled` flag from YAML frontmatter
3. Check for `completion-promise` in output
4. Check if `iteration >= max-iterations`
5. If none match, block exit and refeed prompt
### Hook Output for Stop Events
```json
{
"decision": "block",
"reason": "Completion promise not found. Restarting iteration."
}
```
Or to allow exit:
```json
{
"continue": true,
"stopReason": "Completion promise detected"
}
```
### Prompt-Based Hooks
For more sophisticated evaluation, use LLM-based hooks:
```json
{
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "prompt",
"prompt": "Check if the task is complete. Context: $ARGUMENTS\n\nRespond with {\"ok\": true} if done, {\"ok\": false, \"reason\": \"...\"} if not.",
"timeout": 30
}
]
}
]
}
}
```
---
## Best Practices
### 1. Always Set `--max-iterations`
> This cannot be overstated: always set `--max-iterations`. Autonomous loops consume tokens rapidly. A typical 50-iteration loop on a medium-sized codebase can cost $50-100+ in API usage.
```bash
/ralph-loop:ralph-loop
# Then configure: max-iterations=30, completion-promise="DONE"
```
### 2. Define Clear, Measurable Success Criteria
**Bad**:
```
Build a todo API and make it good.
```
**Good**:
```
Build a REST API for todos.
Completion criteria:
- All CRUD endpoints working (GET, POST, PUT, DELETE)
- Input validation with error messages
- Tests passing with >80% coverage
- README with API documentation
Output <promise>COMPLETE</promise> when ALL criteria are met.
```
### 3. Use Test-Driven Verification
> The most effective Ralph Loop tasks include built-in verification. This creates a natural feedback loop within the loop.
```
Implement user authentication using TDD:
1. Write failing tests for each requirement
2. Implement feature to make tests pass
3. Run tests after each change
4. If any fail, debug and fix
5. Refactor if needed
6. Output <promise>TESTS_PASS</promise> when all tests green
```
### 4. Include Escape Hatches
```
Primary task: Implement feature X
If stuck after 10 iterations:
- Document what's blocking progress
- List approaches that were attempted
- Suggest alternative approaches
- Output <promise>BLOCKED</promise>
```
### 5. Incremental Goals for Large Tasks
**Bad**:
```
Create a complete e-commerce platform.
```
**Good**:
```
Build e-commerce platform in phases:
Phase 1: User authentication
- JWT-based auth
- Tests passing
- Commit: "feat: add user auth"
Phase 2: Product catalog
- CRUD for products
- Search functionality
- Tests passing
- Commit: "feat: add product catalog"
Phase 3: Shopping cart
- Add/remove items
- Persist cart state
- Tests passing
- Commit: "feat: add shopping cart"
Output <promise>COMPLETE</promise> when all phases done.
```
### 6. Commit Frequently
```
After each meaningful completion:
1. git add .
2. git commit -m "descriptive message"
This creates recovery points and shows progress in git history.
```
### 7. Test Before Long Runs
> Pro tip: Test manually with one iteration before running 50-iteration loops.
```bash
# Test with 1 iteration first
/ralph-loop:ralph-loop
# Configure: max-iterations=1
# Then run full loop
/ralph-loop:ralph-loop
# Configure: max-iterations=50
```
### 8. Use Git for Safety
> Always run Ralph loops in a git-tracked directory. If something goes wrong, you can revert. Each iteration adds to git history, giving you a clear trail of what changed.
---
## Prompt Templates
### Template 1: Test-Driven Development
```markdown
# Task: [FEATURE_NAME]
## Requirements
- [Requirement 1]
- [Requirement 2]
- [Requirement 3]
## Approach
Follow TDD methodology:
1. Write failing tests for each requirement
2. Implement minimal code to pass tests
3. Run tests: `npm test`
4. If tests fail, read error, fix, repeat
5. When all tests pass, refactor if needed
6. Commit: `git add . && git commit -m "feat: [feature]"`
## Completion
Output <promise>TESTS_PASS</promise> when:
- All tests pass
- Code is committed
- No lint errors
```
### Template 2: Migration/Refactor
```markdown
# Task: Migrate from [OLD] to [NEW]
## Scope
Files to migrate: `src/**/*.ts`
## Migration Steps
For each file:
1. Update imports
2. Replace deprecated patterns
3. Run type check: `npx tsc --noEmit`
4. If errors, fix them
5. Run tests: `npm test`
6. Commit: `git commit -m "refactor: migrate [file]"`
## Completion
Output <promise>MIGRATION_COMPLETE</promise> when:
- All files migrated
- Type check passes
- All tests pass
- All changes committed
```
### Template 3: Bug Fix
```markdown
# Bug: [BUG_DESCRIPTION]
## Reproduction
[Steps to reproduce]
## Investigation
1. Find the root cause
2. Document findings
## Fix
1. Write a failing test that reproduces the bug
2. Implement the fix
3. Verify test passes
4. Check for regressions: `npm test`
5. Commit: `git commit -m "fix: [description]"`
## Completion
Output <promise>FIXED</promise> when:
- Bug is fixed
- Test added to prevent regression
- All tests pass
```
### Template 4: Time-Aware Loop
```markdown
# Task: Optimize API performance
## Primary Goals
1. Profile existing endpoints
2. Identify bottlenecks
3. Implement optimizations
4. Verify improvements
## Duration
Minimum runtime: 4 hours
## Self-Generated Tasks
If primary goals complete before 4 hours:
- Add caching layers
- Optimize database queries
- Add request batching
- Improve error handling
- Add performance tests
## Completion
Output <promise>TIME_COMPLETE</promise> when:
- All primary goals achieved
- Minimum 4 hours elapsed
- All tests pass
```
---
## When to Use (and Not Use)
### Good Use Cases
| Use Case | Why It Works |
|----------|--------------|
| **Large refactors** | Clear mechanical steps, verifiable via tests |
| **Framework migrations** | Repetitive patterns, type checking validates |
| **Test coverage** | "Add tests for uncovered functions" is measurable |
| **Greenfield projects** | Can run overnight, tests verify correctness |
| **Batch operations** | Same operation across many files |
| **Dependency upgrades** | API changes are well-documented |
### Poor Use Cases
| Use Case | Why It Fails |
|----------|--------------|
| **Ambiguous requirements** | Can't define success criteria |
| **Architectural decisions** | Requires human judgment |
| **Security-critical code** | Needs human review |
| **Production debugging** | Often requires context not in code |
| **UX/design decisions** | Subjective, not automatable |
| **Exploratory work** | "Figure out why it's slow" has no clear endpoint |
### Decision Framework
Ask yourself:
1. **Can I define "done" objectively?** (tests pass, lint clean, etc.)
2. **Is there automatic verification?** (tests, type checking, linting)
3. **Is the task mechanical or creative?** (mechanical = good for Ralph)
4. **What's the cost of failure?** (high cost = needs human review)
---
## Real-World Examples
### Example 1: Y Combinator Hackathon
- **Task**: Generate multiple repositories overnight
- **Result**: 6 repositories generated autonomously
- **Key**: Each repo had clear completion criteria
### Example 2: $50K Contract
- **Task**: Large codebase migration
- **Result**: Completed for $297 in API costs
- **Key**: Well-defined migration patterns, comprehensive tests
### Example 3: Programming Language (Cursed)
- **Task**: "Make me a programming language like Golang but with Gen Z slang keywords"
- **Result**: Functional compiler with LLVM backend, standard library, editor support
- **Duration**: 3 months of autonomous iteration
- **Keywords**: `slay` (function), `sus` (variable), `based` (true)
### Example 4: React Migration
- **Task**: Upgrade from React v16 to v19
- **Result**: 14-hour autonomous session, complete migration
- **Key**: Clear deprecation warnings, comprehensive test suite
---
## Codeman Implementation
Codeman implements Ralph Wiggum tracking via the `RalphTracker` class in `src/ralph-tracker.ts`.
### Auto-Detection Patterns
The tracker automatically enables when detecting:
| Pattern | Example | Regex |
|---------|---------|-------|
| Ralph command | `/ralph-loop:ralph-loop` | `/\/ralph-loop\|starting ralph/i` |
| Promise tag | `<promise>COMPLETE</promise>` | `/<promise>([^<]+)<\/promise>/` |
| TodoWrite | `Todos have been modified` | `/TodoWrite\|todos?\s*(?:updated\|written)/i` |
| Iteration | `Iteration 5/50` or `[5/50]` | `/(?:iteration)\s*#?(\d+)(?:\s*[\/of]\s*(\d+))?/i` |
| Todo checkbox | `- [ ] Task` | `/^[-*]\s*\[([xX ])\]\s+(.+)$/gm` |
| Todo indicator | `Todo: ☐ Task` | `/Todo:\s*(☐\|◐\|✓)/g` |
| All complete | `All tasks completed` | `/all\s+tasks?\s+completed?\|all\s+done/i` |
| Task done | `Task 8 is done` | `/task\s*#?\d+\s*(?:is\s+)?done/i` |
### Completion Detection
Multi-strategy detection to catch various completion signals:
1. **Tagged phrase**: `<promise>PHRASE</promise>` - First occurrence stores phrase, second triggers completion
2. **Bare phrase**: Detects phrase without tags once expected phrase is known (e.g., Claude outputs `COMPLETE` instead of `<promise>COMPLETE</promise>`)
3. **All complete signals**: Detects "All X files/tasks created/completed" messages, marks all todos complete and emits completion
4. **Explicit task completion**: Matches "Task N is done" patterns
### Session Lifecycle
Each session has its **own independent tracker**:
| Action | Result |
|--------|--------|
| New session opened | Fresh tracker, no carryover |
| Tab closed | Tracker state cleared, UI panel hides |
| Switch tabs | Panel shows tracker for active session |
| `tracker.reset()` | Clears todos/state, keeps enabled status |
| `tracker.fullReset()` | Complete reset to initial state |
| `tracker.configure({...})` | Partial config update (enabled, completionPhrase, maxIterations) |
### State Structure
```typescript
interface RalphLoopState {
enabled: boolean; // Tracker active?
active: boolean; // Loop running?
completionPhrase: string | null;
startedAt: number | null;
cycleCount: number;
maxIterations: number | null;
lastActivity: number;
elapsedHours: number | null;
}
interface RalphTodoItem {
id: string;
content: string;
status: 'pending' | 'in_progress' | 'completed';
detectedAt: number;
}
```
### API Endpoints
| Method | Endpoint | Description |
|--------|----------|-------------|
| GET | `/api/sessions/:id/ralph-state` | Get loop state and todos |
| POST | `/api/sessions/:id/ralph-config` | Configure tracker settings |
**POST `/ralph-config` Options**:
```json
{
"enabled": true, // Enable/disable tracker
"reset": true, // Soft reset (clears state, keeps enabled)
"reset": "full", // Full reset (clears everything)
"completionPhrase": "DONE" // Set expected completion phrase
}
```
**GET Response**:
```json
{
"success": true,
"data": {
"loop": {
"enabled": true,
"active": true,
"completionPhrase": "COMPLETE",
"cycleCount": 5,
"maxIterations": 50,
"elapsedHours": 2.5
},
"todos": [
{ "id": "todo-abc", "content": "Fix auth", "status": "completed" },
{ "id": "todo-def", "content": "Add tests", "status": "in_progress" }
],
"todoStats": { "total": 5, "pending": 2, "inProgress": 1, "completed": 2 }
}
}
```
### SSE Events
| Event | Data | When |
|-------|------|------|
| `session:ralphLoopUpdate` | `RalphLoopState` | Loop state changes |
| `session:ralphTodoUpdate` | `RalphTodoItem[]` | Todos detected/updated |
| `session:ralphCompletionDetected` | `{ phrase: string }` | Completion phrase found |
### Skill Commands
```bash
/ralph-loop:ralph-loop # Start Ralph Loop in current session
/ralph-loop:cancel-ralph # Cancel active Ralph Loop
/ralph-loop:help # Show help and usage
```
---
## Troubleshooting
### Loop Never Completes
**Cause**: Completion criteria aren't clear enough.
**Solution**: Be more specific about what "done" means. Include testable criteria:
```
Output <promise>DONE</promise> when:
- `npm test` exits with code 0
- `npm run lint` exits with code 0
- All files committed
```
### Same Error Every Iteration
**Cause**: Claude is stuck in a failure loop.
**Solution**: Add escape hatch to prompt:
```
If stuck after 10 iterations with the same error:
1. Document the error and what was tried
2. Suggest alternative approaches
3. Output <promise>STUCK</promise>
```
### High API Costs
**Cause**: Too many iterations, large context.
**Solutions**:
1. Always set `--max-iterations`
2. Use `/clear` between major phases
3. Keep files small and focused
4. Test with 1 iteration first
### False Completion Detection
**Cause**: Completion phrase appears in prompt or documentation.
**Solution**: Use unique, unlikely phrases:
```
# Bad (might appear in docs)
<promise>COMPLETE</promise>
# Good (unique)
<promise>TASK_XYZ_VERIFIED_DONE</promise>
```
### Tracker Not Enabling
**Cause**: No Ralph patterns detected in output.
**Solution**:
1. Manually enable: `POST /api/sessions/:id/ralph-config { "enabled": true }`
2. Or ensure Claude outputs recognizable patterns
### Context Window Exhaustion
**Cause**: Long-running loops accumulate context.
**Solution**: Configure auto-clear:
```bash
POST /api/sessions/:id/auto-clear
{ "enabled": true, "threshold": 140000 }
```
---
## References
### Official Documentation
- [Anthropic Ralph Wiggum Plugin](https://github.com/anthropics/claude-code/tree/main/plugins/ralph-wiggum)
- [Claude Code Hooks Reference](https://code.claude.com/docs/en/hooks)
- [Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
- [Claude Code Overview](https://code.claude.com/docs/en/overview)
### Community Resources
- [Awesome Claude - Ralph Wiggum](https://awesomeclaude.ai/ralph-wiggum)
- [Claude Fast - Autonomous Agent Loops](https://claudefa.st/blog/guide/mechanics/autonomous-agent-loops)
- [DeepWiki - Ralph Loop](https://deepwiki.com/anthropics/claude-plugins-official/5.2.2-ralph-loop)
### Related Codeman Files
- `src/ralph-tracker.ts` - Core detection engine
- `src/ralph-loop.ts` - Task orchestration
- `src/respawn-controller.ts` - Session cycling
- `src/spawn-orchestrator.ts` - Autonomous agent lifecycle (uses RalphTracker for completion)
- `src/spawn-detector.ts` - Detects `<spawn1337>` tags in terminal output
- `src/types.ts` - Type definitions
---
*This documentation is maintained as part of the Codeman project. For updates, see the main [CLAUDE.md](../CLAUDE.md).*
-140
View File
@@ -1,140 +0,0 @@
# Read My Mind (design)
A 🧠 button that predicts the prompt you were about to type. Codeman keeps a per-case **intent profile** (your stated goals plus the real prompts you recently sent), feeds it and the live pane tail to a one-shot `claude -p`, and shows the predicted next prompt in a plan-mode-style approval dialog: **Send** / **Rethink** (with an optional steer note) / **Insert** (drop it on the composer to edit) / **Dismiss**. It is also a skill surface: the agent can read the intent profile, record intentions, and request a prediction over the HTTP API. Suggestions are **never auto-sent**; the human click is the boundary.
## UX flow
1. User hits 🧠 (desktop header button; phone: keyboard-accessory key).
2. Modal opens with a spinner, then the top suggestion in an editable single-line field, rationale below it, up to 2 alternates as tappable rows.
3. Buttons: **Send** (submits with `\r`), **Insert** (sends without `\r`, so the text sits unsubmitted on the CLI composer for editing, a documented mechanism), **Rethink** (optional free-text steer, e.g. "no, I meant the mobile bug", re-runs with the rejected suggestions included), **Dismiss**.
4. Accepted prompts flow back into the intent history like any other sent prompt, so the profile self-corrects.
## Scope (v1)
- Claude mode only (capture rides Claude transcripts; external CLIs have no transcript watcher). Mirrors the approvals-inbox scoping.
- Opt-in: `readMyMindEnabled`, synced, default **OFF**. While OFF: no capture, no UI surfaces. Privacy first, and every press costs real tokens.
- One prediction in flight per session; the button disables while checking.
- Sync request/response (the predictor takes 5-30s; agent-wait long-polls already hold requests longer). No new SSE events in v1.
## Data model
Per case, not per session: intentions outlive `/clear` and respawns.
```ts
interface IntentProfile {
key: string; // sha256(owner + ':' + realpath(workingDir)).slice(0, 16)
workingDir: string;
updatedAt: number;
goals: string; // freeform markdown, user/agent editable, ≤ 8 KB
recentPrompts: { ts: number; sessionId: string; text: string }[]; // FIFO cap 50, each ≤ 500 chars
}
```
Storage: `dataPath('intents.json')`, written mode 0600 (prompts can contain secrets; same posture as `users.json`). Never enters the `/api/search` index. Add to the CLAUDE.md State Files list.
## Intent capture
**Source: the session transcript, not the input paths.** `POST /api/sessions/:id/input` sees only programmatic input, and the WS channel delivers raw keystrokes (`session.write(msg.d)`), so neither yields clean submitted prompts. Claude's own JSONL transcript records every user turn as structured text, and `transcript-watcher.ts` already tails it. Add a `userPrompt` event there:
- Emit for `type: 'user'` entries whose content is a string or contains a text block; skip entries that are only `tool_result` blocks (tool results are wrapped as user messages).
- Skip `<command-name>` / `<local-command-stdout>` tagged entries (local slash-command echo, not intent).
- Skip texts < 3 chars (menu digits, Esc artifacts), truncate to 500, drop consecutive duplicates ("continue" spam from auto-resume stays but dedupes).
`IntentStore` (new `src/intent-store.ts`, pure core + IO wrapper, in the style of `session-order.ts`) subscribes via session wiring, gated on the setting resolved from **merged** settings per the partial-PUT rule.
## Context assembly (how the mind reading actually works)
The quality of the suggestion is decided before the model ever runs, by what we put in front of it. A new pure function `buildPredictionContext()` (in `src/readmymind-context.ts`, unit-testable with fixtures, no IO of its own; collectors inject their data) assembles a budgeted, priority-ordered prompt from every signal Codeman already has:
| # | Source | What it contributes | Cap |
| - | ------ | ------------------- | --- |
| 1 | **Pending dialog** (approvals-inbox store, when present) | If the session is sitting on an AskUserQuestion / permission / idle prompt, the honest "next prompt" is an *answer*. The dialog text + parsed options go in first and the model is told to answer it. | 2 KB |
| 2 | **User goals** (`goals` from the intent profile) | The only fully-trusted statement of what the user wants. Highest authority in the trust ranking below. | 8 KB |
| 3 | **Last assistant turn** (transcript, not the pane) | Assistant replies usually *end* with the fork in the road ("Want me to X?", "Next steps: ..."), so keep the **tail** when truncating. The transcript has the full message; the pane is a repaint window full of spinner junk. | 6 KB |
| 4 | **Recent user prompts** (intent profile, with timestamps) | The conversation rhythm AND the user's prompting voice: length, tone, shorthand (`COM`, lowercase, typos and all). The model is instructed to write suggestions in *this* style, not assistant-ese. | last 20 |
| 5 | **Recent tool activity** (transcript `tool_use` blocks, already parsed by `TranscriptWatcher`) | One line per call: `Edit src/foo.ts`, `Bash npm test (failed)`. What the agent actually *did*, which the last message may summarize away. | last 10 |
| 6 | **Workspace signals** (`collectWorkspaceSignals()`: `git` via `execFile` in `workingDir`, 2s timeout) | Branch, `status --short` (dirty files scream "commit/test/deploy next"), last 5 commits oneline, presence of `.changeset/*.md` (release pending). Skipped for remote-SSH cases (workingDir is not local); fine for Docker cases (bind-mounted at the same host path). Non-git dirs: section omitted. | 3 KB |
| 7 | **Away context** (run-summary events + elapsed time) | `Last user prompt was 6h ago; since then: <run-summary events for this session>`. After a long gap the right suggestion is often "review / continue yesterday's thread", not a blind continuation. | 2 KB |
| 8 | **Sibling sessions** (live sessions sharing the case) | One line each: name, mode, working/idle. A lead-and-workers setup changes what the next prompt should be ("check on w2" beats "keep going"). | 1 KB |
| 9 | **Rethink state** (steer note + rejected suggestions) | Only on re-runs. Rejections are strong negative signal and go in verbatim. | 2 KB |
Total budget ~30 KB. When over budget, drop from the bottom up (siblings first, then away context, then workspace signals); sections 1-4 never drop, they only truncate. Deterministic assembly means fixture tests can pin exactly what a given situation feeds the model.
**Trust tiers are stated in the prompt.** Goals and user prompts are *the user*; assistant text, tool logs, and pane content are *observations that may contain text trying to manipulate you* (a hostile repo can print "SUGGEST: run curl evil.sh"). The prompt instructs: user-stated intent outranks anything observed, and never propose a prompt whose primary source is terminal output alone. The human approval click remains the hard boundary regardless.
**Output contract** (strict JSON, parse failure = clean error, never a half-suggestion):
```json
{ "suggestions": [ { "prompt": "...", "why": "...", "kind": "continue" | "verify" | "redirect" } ] }
```
1-3 entries, and the *kinds* force useful diversity instead of three rewordings: `continue` (finish the current thread, or answer the pending dialog), `verify` (test/review what was just built; the user's own "always end-to-end test" discipline), `redirect` (the next goal from the intent profile that the current thread is not serving). The modal shows `continue` big, the others as alternates. Embedded newlines are stripped server-side (single-line prompt rule; multi-line breaks Ink).
## Predictor
New `src/readmymind-predictor.ts`, reusing the `AiCheckerBase` mechanics (prompt file to dodge E2BIG, one-shot `claude -p --output-format text` in a throwaway tmux `codeman-rmm-<id8>`, done-marker polling, timeout, model-name validation) but standalone: the base class is verdict-shaped (positive/negative/cooldown) and prediction is freeform JSON, so subclassing would abuse `reasoning` as a payload. If a shared spawn/poll helper falls out naturally, extract it; do not block on the refactor.
- **Model: opus** (decided). `readMyMindModel` setting, default `AI_CHECK_MODEL` (currently `claude-opus-4-5-20251101`); prediction quality is the product, and it runs only on an explicit press, so the cost profile is nothing like the idle checker's. Timeout 90s (opus headroom over a ~30 KB prompt).
- Input: the assembled context above. The predictor itself stays dumb: text in, JSON out; all intelligence about *what to include* lives in the testable assembler.
## API (new `src/web/routes/readmymind-routes.ts`)
Normal authed API, `ApiResponse` envelope, Zod schemas in `schemas.ts`, ownership via `findSessionOrFail` (the profile key derives from the session's owner + workingDir, so multi-user scoping is structural):
- `GET /api/sessions/:id/intent` → the session's `IntentProfile`.
- `PUT /api/sessions/:id/intent` body `{ goals }` (bounded) → update goals. Used by the modal's edit view and by the agent skill ("record that the user is working toward X").
- `DELETE /api/sessions/:id/intent` → forget everything for this case (the modal's "Forget" affordance).
- `POST /api/sessions/:id/readmymind` body `{ steer?, rejected? }` → `{ suggestions }`. 409 `INVALID_STATE` while a prediction is already running for the session; claude-mode sessions only (400 otherwise, mirroring wait-signal gating).
## Frontend
New module `readmymind-ui.js` (@loadorder 11.3, after panels-ui.js), prettier-formatted.
- **Desktop**: header button `btn-readmymind`, default-hidden via marker class `btn-readmymind--hidden` (the `!important` display rules require the marker-class pattern), shown by `applyHeaderVisibilitySettings()` when the setting is ON. Off phones per `test/mobile-header-buttons-policy.test.ts`.
- **Phone**: a 🧠 key on the keyboard accessory bar (that bar is where input helpers live, and phones are where typing hurts most). Opens the same modal. Modal z-index respects the ≤768px layer rules (1300+).
- **Send** goes server-side: `POST /api/sessions/:id/input` with `\r` appended. Deliberately NOT the browser keystroke path, so the `sendEnterKey` / local-echo-overlay trap never applies (the modal is UI chrome, not terminal typing). **Insert** is the same POST without `\r`.
- i18n strings registered (en + zh-CN); suggestion text itself carries `data-i18n-skip`.
## Skill integration
The user-facing promise: the button is also a skill. Extend `skills/codeman`:
- New section "Read My Mind: intent + prediction" with the three intent verbs (read profile, append/replace goals, predict) and the guard notes (single-line prompts, never auto-send to another session without the user asking).
- Update `reference/endpoints.md` (the endpoints.md drift test pins this).
- The auto-injected case copy heals via the existing marker-owned `applyAgentSkill` mechanism; nothing new needed there.
Agent use cases this unlocks: a lead session records intentions as the user states them ("remember: shipping 1.16 is the goal"), and a returning user gets a prediction grounded in what the agent knew, not just raw prompt history.
## Security / privacy
- **The human gate is the injection mitigation**: pane output (attacker-influenceable) flows into the predictor, so its output is only ever *proposed*, rendered as text (`textContent`), and sent solely by an explicit user click. No auto-send path exists, including for the skill.
- Intent data: 0600 file, bounded fields, per-owner keys, endpoints ownership-checked, excluded from search, cleared via DELETE.
- Predictor spawns with the user's own credentials exactly like the AI idle/plan checkers; model name shell-validated the same way.
- Setting OFF stops capture immediately; existing data stays until DELETE (explicit, not silent).
## Tests
- `test/intent-store.test.ts`: key derivation, caps/FIFO, consecutive-dupe skip, tag/tool_result filtering fixtures, 0600 mode, multi-user key separation.
- `test/readmymind-context.test.ts`: fixture scenarios pinning the assembled prompt: pending-dialog-first ordering, tail-keeping truncation of the assistant turn, budget drop order (siblings before workspace signals), remote-case git skip, trust-tier framing present, rejected suggestions included only on rethink.
- `test/readmymind-predictor.test.ts`: strict JSON parse, garbage output → error result, newline stripping, `kind` validation, rejected-suggestions threading into the prompt.
- `test/routes/readmymind-routes.test.ts` (`app.inject`): CRUD round-trip, predict with a stubbed predictor, 409 while in flight, non-claude 400, ownership 404, Send/Insert byte assertions via the test-PTY echo (`\r` present vs absent).
- Transcript capture: extend the transcript-watcher fixtures with user-turn entries.
## Phases
1. **Intent store + capture + intent endpoints + skill docs.** Immediately useful to agents even before any UI exists.
2. **Context assembler + predictor + predict endpoint + desktop button/modal.** The feature as pitched. The assembler ships with all collectors it can serve from day one (transcript, intent, git, run-summary, siblings); the approvals collector activates when PR #245 lands.
3. **Phone accessory key, rethink steering, alternates row.** Part 1 (shipped): the alternates row (tappable, swap into the field without losing edits; Rethink rejects the whole shown set), the phone 🧠 keyboard-accessory key (both bar templates, `rmm-enabled` marker class on the bar), and a phone-sized modal (small dialog, not full-screen). Part 2 (shipped): rethink steering, the free-text steer note under the suggestions, sent as `steer`, visible whenever Rethink is live (ready and empty-result phases), cleared on each open; the empty-result copy points at the note, and the footer buttons moved to the styled `btn-toolbar` convention (the bare `btn btn-*` classes they shipped with match no CSS in this codebase and rendered as unstyled UA buttons).
4. Explicitly later: proactive predict-on-idle (ghost suggestion chip), auto-compaction of `recentPrompts` into `goals` via a cheap model, codex/gemini capture, cross-case "global" intent.
## Open questions
- Should Rethink's rejected-suggestion memory persist across modal closes, or reset each open?
- Is a composer-adjacent placement (next to the toolbar Run controls) better than the header for discoverability?
- Pending-dialog input (source #1) consumes the approvals-inbox store (PR #245, merged): the phase-2 collector reads pending items directly from `src/approval-inbox.ts`.
## Docs
- CLAUDE.md: Key Patterns entry, State Files (`intents.json`), frontend load order, route count.
- `docs/api-reference.md`: four endpoints (additive under the 0.9.x contract).
- `skills/codeman/reference/endpoints.md`: new rows (drift-test enforced).
-108
View File
@@ -1,108 +0,0 @@
# Read My Mind
Codeman's per-case memory of what you are trying to accomplish, and the 🧠 button that turns it into a predicted next prompt. Each case gets an **intent profile**: a freeform `goals` text (written by you or your agent) plus the prompts you actually submitted, captured automatically while the feature is on. Pressing 🧠 feeds that profile and the live session signals to a one-shot model call and shows the predicted prompt for you to send, edit, or rethink. Nothing is ever sent to a session automatically. Design doc: [`readmymind-plan.md`](readmymind-plan.md).
## What it does
- Captures the prompts you submit in Claude sessions into a per-case history (50 most recent, bounded).
- Lets you (or your agent) record explicit goals per case.
- Predicts your next prompt on demand (the 🧠 header button, or `POST .../readmymind` for agents): the suggestion arrives in a modal with Send / Insert / Rethink / Dismiss.
- Exposes the profile over the HTTP API, and to agents through the `codeman` skill, so an agent can ground its work in what you actually want instead of guessing from the last screenful.
## Turning it on
App Settings → Header & Panels → Cross-session features → **Read My Mind** (synced setting `readMyMindEnabled`, default **OFF**). It gates everything: capture, the header button, and nothing shows anywhere while it is off. The API equivalent:
```bash
curl -sk -X PUT https://localhost:3000/api/settings \
-H 'Content-Type: application/json' \
-d '{"readMyMindEnabled": true}'
```
Add `-u user:password` if your install has `CODEMAN_PASSWORD` set, and drop `-k`/use `http://` for a plain-HTTP dev server. Turning it OFF stops capture immediately; existing profiles stay until you delete them (below).
## The 🧠 button
On a Claude session, press the brain button in the header (desktop) or the 🧠 key on the keyboard accessory bar (phones and tablets; it appears when the setting is on). Codeman assembles everything it already knows: your goals, your recent prompts (with your voice: length, tone, shorthand), the tail of the last assistant reply, recent tool activity, git state (branch, dirty files, pending changesets), how long you have been away and what happened meanwhile, sibling sessions in the same case, and any dialog the session is currently waiting on. A one-shot model call (opus by default, `readMyMindModel` to override) turns that into 1-3 suggestions; the top one lands in an editable field with its rationale, and the others render as tappable alternate rows: tap one to swap it into the field (edits you already made are kept on the row you leave).
- **Send** submits it to the session (with Enter).
- **Insert** drops it on the CLI composer *without* Enter, so you can edit it in the terminal before sending.
- **Rethink** re-runs with everything shown (the field and the alternates) recorded as rejected. An optional steer note below the suggestions ("no, I meant the mobile bug") rides along as your own words, the highest-authority signal the predictor gets; it stays in the field across re-runs until you clear it or reopen the modal.
- **Dismiss** closes; nothing happens.
A prediction takes 5-90 seconds and costs real tokens; one runs per session at a time. If the session is sitting on a permission/question dialog, the suggestion is usually an answer to that dialog: that is intentional.
**Security note**: the prediction reads observable content (assistant output, tool logs, git output) which a hostile repo could try to steer. The predictor is told user-stated intent outranks anything observed, and, more importantly, a suggestion is only ever *proposed*: your click is the boundary. No auto-send path exists, including for agents.
## What gets captured, exactly
Capture reads the Claude session transcript, not your keystrokes: when a user turn lands in the transcript, its text is folded into the case's profile. Filters applied on the way in:
- **Claude-mode sessions only.** Shell, OpenCode, Codex, Gemini, Antigravity, and Pi sessions are never captured (they have no transcript watcher).
- Tool results, local slash-command echo (`/model` and friends), system wrappers, and interrupt markers are skipped.
- Entries shorter than 3 characters are skipped (menu digits, Esc artifacts).
- Consecutive duplicates collapse (auto-resume's "continue" spam counts once per run).
- Each prompt is stored as one line, truncated to 500 characters; the history caps at 50 prompts FIFO.
Because the transcript path arrives via Claude Code hooks, capture needs hooks to reach the server, the same condition as hook-based idle detection. Docker cases against a loopback-only server need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; remote-SSH cases do not capture.
## What is never captured
- Anything while `readMyMindEnabled` is OFF (capture is not retroactive).
- Terminal output, keystrokes, passwords typed into shells: only submitted Claude prompts are read.
- Nothing leaves the machine beyond the model call you explicitly trigger, and profiles are never fed into `/api/search`.
## Where it lives, and how to wipe it
Profiles live in `~/.codeman/intents.json`, written atomically at mode 0600 (captured prompts can contain secrets). The file is per Codeman instance. Keys derive from owner + the case's resolved working directory, so profiles survive `/clear`, respawn cycles, and session churn, and in multi-user mode two owners of the same directory get separate profiles.
Forget one case: `DELETE /api/sessions/:id/intent` (below). Forget everything: stop the server and delete `~/.codeman/intents.json`.
## The API
Four endpoints, session-scoped so ownership is enforced by the session itself (`/api/v1/` aliases work too; full spec in [`api-reference.md`](api-reference.md)):
```bash
# Read the profile for a session's case
curl -sk https://localhost:3000/api/sessions/$SID/intent | jq '.data.intent'
# Record goals (REPLACES the text: read + merge if you want to append)
curl -sk -X PUT https://localhost:3000/api/sessions/$SID/intent \
-H 'Content-Type: application/json' \
-d '{"goals":"ship 1.17; then mobile polish"}'
# Forget the case
curl -sk -X DELETE https://localhost:3000/api/sessions/$SID/intent
# Predict the next prompt (claude-mode only; takes 5-90 s)
curl -sk -X POST https://localhost:3000/api/sessions/$SID/readmymind \
-H 'Content-Type: application/json' -d '{}' | jq '.data.suggestions'
```
A case with nothing recorded answers an empty profile with `updatedAt: 0`; reads never persist anything. Goals cap at 8192 characters and the schema is strict, so unknown fields or over-long goals answer `400 INVALID_INPUT`. A session you do not own answers `404 NOT_FOUND`, indistinguishable from a nonexistent one. Predict answers `{ suggestions: [{ prompt, why, kind }], durationMs }` (`kind`: `continue` / `verify` / `redirect`), `409 CONFLICT` while one is already running, `400 INVALID_INPUT` on non-claude sessions, and `502 OPERATION_FAILED` when the model produced no usable JSON. The rethink flow passes `{"steer":"…","rejected":["…"]}`.
## For agents (the skill)
The `codeman` agent skill documents the same verbs (SKILL.md §3 plus `reference/endpoints.md`), with the ground rules: read the profile to understand what the user wants, record goals the user actually stated, merge instead of blind-writing (PUT replaces), never delete a profile unprompted, and never send a predicted suggestion into a session unless the user asked. It is the user's memory, not the agent's.
## What comes next
Explicitly later: proactive predict-on-idle, auto-compaction of the prompt history into goals, non-Claude capture. See the phases section of [`readmymind-plan.md`](readmymind-plan.md).
## Troubleshooting
| Symptom | Cause / fix |
| ------- | ----------- |
| No 🧠 button in the header | `readMyMindEnabled` is OFF (App Settings → Header & Panels → Cross-session features), you are on a phone (there it is a key on the keyboard accessory bar instead, visible while typing), or the active session is not claude-mode |
| Prediction feels generic | The profile is thin: record goals (PUT or ask your agent to), and let capture accumulate a few real prompts first |
| "A prediction is already running" (409) | One per session at a time; wait for the current one (up to 90 s) |
| Prediction fails (502) | The model returned no usable JSON, or the CLI could not start; retry. Check `readMyMindModel` if you overrode it |
| Profile stays empty although I am prompting | `readMyMindEnabled` was OFF at the time (capture is not retroactive), the session is not claude-mode, or hooks are not reaching the server (Docker case on a loopback bind without `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, or a remote-SSH case) |
| Short answers I typed are missing | Entries under 3 characters are filtered by design (menu digits, Esc artifacts) |
| My goals text vanished after an agent wrote to it | PUT replaces the whole text; the skill tells agents to read + merge, but a blind write wins. Re-state the goals; consider phrasing them in the session so capture keeps the evidence |
| Two profiles for what I think is one case | Different owners in multi-user mode, or genuinely different directories; paths are realpath-resolved, so symlink spellings converge but distinct checkouts do not |
| `400 INVALID_INPUT` on PUT | Goals over 8192 chars, or an extra field in the body (strict schema) |
## Where the code lives
`src/intent-store.ts` (store + pure helpers, singleton), the `transcript:user_prompt` event in `src/transcript-watcher.ts`, capture wiring in `src/web/server.ts` (`captureIntentPrompt`), context assembly in `src/readmymind-context.ts` (pure) + `src/readmymind-collectors.ts` (transcript tail + git IO), the predictor in `src/readmymind-predictor.ts`, routes in `src/web/routes/readmymind-routes.ts`, schemas in `src/web/schemas.ts`, frontend in `src/web/public/readmymind-ui.js`. Tests: `test/intent-store.test.ts`, `test/readmymind-context.test.ts`, `test/readmymind-collectors.test.ts`, `test/readmymind-predictor.test.ts`, `test/routes/readmymind-routes.test.ts`, and the capture cases in `test/transcript-watcher.test.ts`.
-72
View File
@@ -1,72 +0,0 @@
# Reliable input delivery (exactly-once, durable)
## The bug this fixes
With local echo on, pressing Enter cleared the overlay and then sent the prompt
over the WebSocket **fire-and-forget** (`ws.send({t:'i',d})`). On a flaky link
(e.g. a moving train) the socket is frequently *half-open*: `readyState === OPEN`
so `ws.send()` does **not** throw, but the underlying TCP is dead, so the frame is
silently discarded. Nothing was enqueued (the send "succeeded"), the on-screen
prompt was already wiped, and `navigator.onLine` stays `true` — so a long typed
prompt vanished with no trace and no resend.
## The guarantee
Every byte of user input is **recorded durably before delivery** and **only
dropped once the server ACKs it** — so a half-open socket, a reconnect, or a page
reload can never lose input. Redelivery is **exactly-once**: the server applies
each `(clientId, seq)` at most once, so a resend can't type the prompt twice.
## How it works
### Client (`app.js`)
- A stable **`clientId`** (`localStorage['codeman:clientId']`) identifies this
browser to the server's dedup across reconnects and reloads.
- Each input frame gets a **monotonic per-session `seq`**. Frame records
(`{seq,data,useMux,ts,tries,sentAt}`) live in `_pendingDeliveries`
(`Map<sessionId, record[]>`), persisted (debounced, + flushed on `pagehide`/
`visibilitychange`) to `localStorage['codeman:pendingInput']`. The seq counters
persist too, so seqs stay monotonic across reloads (never reset — a reset would
let the server treat fresh input as an already-applied duplicate).
- **Delivery** (`_drainSession`):
- **WS path** — when the socket is `OPEN` for the session, send each not-yet-sent
record (`sentAt === 0`) in seq order over the single ordered stream. Records
stay pending until the server's `{t:'ia',seq}` ACK removes them.
- **POST path** — when no WS, POST records in order, awaiting each (the HTTP 2xx
*is* the ACK). A 404/410 (session gone) drops the record rather than retry
forever.
- **Half-open recovery** (`_redeliverSweep`, every 2s): if the active WS session's
oldest record is unacked past `_reliableAckTimeoutMs` (4s), the socket is assumed
dead — `ws.close()` forces a fast reconnect; `onopen` (`_onWsReady`) resets
`sentAt = 0` and re-sends everything pending. Also re-drains background sessions
over POST, and fires on SSE-reconnect / `online`.
- The connection indicator shows pending count/bytes (`_pendingBytes`).
### Server
- **`Session.shouldApplyInput(clientId, seq)`** — returns `true` exactly once per
`(clientId, seq)`: the first time a seq strictly greater than that client's
last-applied is seen. A replayed/lower seq returns `false`. Bounded MRU map
(`MAX_INPUT_DEDUP_CLIENTS = 256`).
- **WS route** (`ws-routes.ts`) — parses optional `cid`/`seq` on `{t:'i'}`; applies
via `shouldApplyInput` (skips a duplicate, still ACKs with `{t:'ia',seq}` so the
client drops it). Untagged frames apply unconditionally (no behavior change).
- **POST route** (`/api/sessions/:id/input`) — optional `seq`/`clientId` in
`SessionInputWithLimitSchema`; a deduped duplicate returns 200 without writing
(the 200 is the client's ACK). `curl`/legacy callers omit the fields and always
apply.
## Known limitation
Dedup state is in-memory on the server. A **server restart** between a write and
the client's redelivery of that same seq could re-apply it (a rare duplicate).
This is a deliberate trade-off: favor *never losing input* over a rare duplicate
across the narrow restart window.
## Tests
- `test/reliable-input-dedup.test.ts` — `Session.shouldApplyInput` exactly-once
semantics (monotonic, per-client, gap-tolerant, eviction-safe).
- `test/routes/session-routes.test.ts` — POST `/input` applies a tagged
`(clientId, seq)` once on redelivery; untagged input always applies.

Some files were not shown because too many files have changed in this diff Show More