Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
915895152c | ||
|
|
7d5b0bec1c |
@@ -1,27 +0,0 @@
|
||||
# Changesets
|
||||
|
||||
Hello and welcome! This folder has been automatically generated by `@changesets/cli`, a build tool that works
|
||||
with multi-package repos, or single-package repos to help you version and publish your code. You can find
|
||||
the full documentation for it [in the repository](https://github.com/changesets/changesets).
|
||||
|
||||
## What is a changeset?
|
||||
|
||||
A changeset is a piece of information about changes made in a branch or commit. It holds three bits of information:
|
||||
|
||||
- What packages need to be released
|
||||
- What semver bump type each package should receive (major / minor / patch)
|
||||
- A summary of the changes
|
||||
|
||||
## How do I create a changeset?
|
||||
|
||||
Run `npx changeset` or create a `.md` file in this directory with the following format:
|
||||
|
||||
```markdown
|
||||
---
|
||||
"codeman": patch
|
||||
---
|
||||
|
||||
Description of changes
|
||||
```
|
||||
|
||||
The frontmatter specifies which package(s) to bump and the bump type. The body is the changelog entry.
|
||||
@@ -1,11 +0,0 @@
|
||||
{
|
||||
"$schema": "https://unpkg.com/@changesets/config@3.1.1/schema.json",
|
||||
"changelog": "@changesets/cli/changelog",
|
||||
"commit": false,
|
||||
"fixed": [],
|
||||
"linked": [],
|
||||
"access": "public",
|
||||
"baseBranch": "master",
|
||||
"updateInternalDependencies": "patch",
|
||||
"ignore": []
|
||||
}
|
||||
@@ -1,30 +0,0 @@
|
||||
{
|
||||
"name": "codeman",
|
||||
"owner": {
|
||||
"name": "Ark0N",
|
||||
"url": "https://github.com/Ark0N"
|
||||
},
|
||||
"description": "Codeman, self-hosted mission control for AI coding agents. Ships the codeman agent skill: let one Claude Code session spawn, prompt, wait on and read other sessions.",
|
||||
"plugins": [
|
||||
{
|
||||
"name": "codeman",
|
||||
"source": "./plugins/codeman",
|
||||
"description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
|
||||
"version": "1.33.1",
|
||||
"author": {
|
||||
"name": "Ark0N",
|
||||
"url": "https://github.com/Ark0N"
|
||||
},
|
||||
"homepage": "https://getcodeman.com",
|
||||
"category": "productivity",
|
||||
"keywords": [
|
||||
"codeman",
|
||||
"orchestration",
|
||||
"multi-agent",
|
||||
"session-manager",
|
||||
"tmux",
|
||||
"claude-code"
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1,21 +0,0 @@
|
||||
.git
|
||||
.agents
|
||||
.claude
|
||||
.codex
|
||||
# `**/` matters: a .dockerignore pattern is matched against the WHOLE
|
||||
# context-relative path, so a bare `.env` excludes ONLY the root file and
|
||||
# `COPY . .` would bake docker/.env -- CODEMAN_PASSWORD and any provider API
|
||||
# keys -- into the published image at /opt/codeman/docker/.env (verified).
|
||||
**/.env
|
||||
**/.env.*
|
||||
!**/.env.example
|
||||
# Same shape: docker/docker-compose.override.yml is the documented home for
|
||||
# host-specific settings, so it must not ride COPY . . into the image either.
|
||||
**/docker-compose.override.*
|
||||
node_modules
|
||||
dist
|
||||
coverage
|
||||
out
|
||||
test-results
|
||||
tmp
|
||||
*.log
|
||||
@@ -1,12 +0,0 @@
|
||||
root = true
|
||||
|
||||
[*]
|
||||
indent_style = space
|
||||
indent_size = 2
|
||||
end_of_line = lf
|
||||
charset = utf-8
|
||||
trim_trailing_whitespace = true
|
||||
insert_final_newline = true
|
||||
|
||||
[*.md]
|
||||
trim_trailing_whitespace = false
|
||||
@@ -1,89 +0,0 @@
|
||||
# Contributing to Codeman
|
||||
|
||||
Thanks for wanting to help! Codeman is a small project with a fast loop: issues usually get a response within a day, good PRs get reviewed quickly, and every release credits its contributors and bug reporters by name in the release notes. This guide gets you from clone to merged PR without stepping on the traps.
|
||||
|
||||
## The short version
|
||||
|
||||
1. **Bugs**: open an issue with your OS, install method (installer / npm / git clone), browser, and which CLI + version the session was running.
|
||||
2. **Questions and ideas**: use [Discussions](https://github.com/Ark0N/Codeman/discussions), not issues.
|
||||
3. **Small fixes** (docs, typos, a new skin, a translation): just send the PR.
|
||||
4. **Anything bigger**: open an issue or Discussion first and get a nod before building. Codeman has strong architectural invariants, and a design chat up front is what turns a big idea into a merged PR instead of a stalled one. This flow works: features like Clone Repo (#236) went idea, then design discussion, then review, then shipped.
|
||||
5. **Security issues**: never a public issue. See [SECURITY.md](SECURITY.md).
|
||||
|
||||
## Dev setup
|
||||
|
||||
Requirements: Node.js 22+ (see `.nvmrc`), tmux, and at least one supported agent CLI on your PATH (Claude Code is the primary one).
|
||||
|
||||
```bash
|
||||
git clone https://github.com/Ark0N/Codeman.git
|
||||
cd Codeman
|
||||
npm install # postinstall builds the vendored xterm addon bundles
|
||||
npm run dev # dev server on http://localhost:3000
|
||||
```
|
||||
|
||||
The frontend is plain JS served from `src/web/public/` with no bundler in dev: edit a `.js`/`.css` file and reload the page. The one exception is `index.html`, which is read once at server start, so markup changes need a server restart.
|
||||
|
||||
## Before you push
|
||||
|
||||
CI runs all of these, so save yourself a round trip:
|
||||
|
||||
```bash
|
||||
npm run typecheck # tsc --noEmit, strict mode
|
||||
npm run lint
|
||||
npm run format:check
|
||||
npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules
|
||||
```
|
||||
|
||||
### Tests
|
||||
|
||||
```bash
|
||||
npm test # the gate — exactly what CI runs
|
||||
npm test -- test/<file>.test.ts # one file
|
||||
```
|
||||
|
||||
`npm test` is the same suite CI runs, so a green run locally means a green run there. It leaves out three suites that cannot pass on an arbitrary machine, each with its own command:
|
||||
|
||||
```bash
|
||||
npm run test:browser # Playwright + chromium (+ a live server; codex-predictive-echo needs a real codex binary)
|
||||
npm run test:mobile # the above plus environment-specific PNG baselines
|
||||
npm run test:perf # wall-clock benchmarks — run on an otherwise idle machine
|
||||
npm run test:all # literally everything, environmental failures included
|
||||
```
|
||||
|
||||
Expect `test:browser`/`test:mobile`/`test:perf` to fail where the machine cannot provide what they need; read that as "not runnable here", not as a regression. `config/test-suites.ts` holds the globs, and both configs derive from it, so the exclusions and those runners cannot drift apart.
|
||||
|
||||
If you add a test that binds a port, pick a unique one at 3150 or above (search the repo for `const PORT =` first). Never 3000.
|
||||
|
||||
Tests are tmux-safe by design: under vitest, the tmux layer becomes an in-memory mock, so tests cannot touch real sessions.
|
||||
|
||||
## Finding your way around
|
||||
|
||||
- Every source file starts with a `@fileoverview` JSDoc block. Read it before diving into the file, it is the map.
|
||||
- [`CLAUDE.md`](../CLAUDE.md) at the repo root is the densest architecture primer in the repo. It is written for AI coding agents, but the invariants and gotchas in it apply to humans exactly the same, and most review feedback on PRs traces back to something already written there.
|
||||
- Deep mechanisms and the history behind each rule live in [`docs/architecture-invariants.md`](../docs/architecture-invariants.md).
|
||||
- Third-party extension surfaces are documented in [`docs/extending-codeman.md`](../docs/extending-codeman.md).
|
||||
|
||||
## Great first contributions
|
||||
|
||||
These are well-fenced areas where a first PR is genuinely easy to get right:
|
||||
|
||||
- **A new theme skin.** A skin is four things kept in sync: the `html[data-skin="…"]` token block in `styles.css`, the xterm ANSI palette in `terminal-ui.js`, the pre-paint allowlist and the settings picker (both in `index.html`). `test/skin-themes.test.ts` statically checks the sync, so if the test passes, your skin works.
|
||||
- **A new language.** `src/web/public/i18n.js` is dependency-free, English is the canonical source, and `zh-CN` is a complete example to copy. Add your language's entries and register it in `SUPPORTED_LANGUAGES`.
|
||||
- **Docs.** If you got stuck on something and then figured it out, the sentence that would have unstuck you is a PR.
|
||||
- Anything labeled [`good first issue`](https://github.com/Ark0N/Codeman/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22).
|
||||
|
||||
Bigger extension points worth discussing first: new CLI backends (the pluggable resolver pattern has absorbed six CLIs so far; `docs/extending-codeman.md` and `docs/opencode-integration.md` show the shape), and real-device testing reports, especially mobile, which always find things emulation cannot.
|
||||
|
||||
## PR expectations
|
||||
|
||||
- **One change per PR.** Small and focused reviews fast; a grab-bag stalls.
|
||||
- Target the `master` branch.
|
||||
- **Keep your branch mergeable.** A PR with conflicts silently gets no CI runs at all (GitHub quirk), so rebase or merge master when conflicts appear.
|
||||
- Include or update tests when you change behavior. Route handlers have a lightweight pattern in `test/routes/` using `app.inject()` (no live server needed).
|
||||
- Formatting is Prettier with a deliberately narrow scope (`npm run format`), several frontend files are hand-formatted on purpose and excluded via `.prettierignore`. Don't "fix" a file by adding it back into Prettier's scope.
|
||||
- Don't bump versions or touch `CHANGELOG.md`; releases are handled by the maintainer via changesets after merge.
|
||||
- AI-assisted contributions are welcome (much of Codeman is built that way), with one condition: you must understand what you're submitting and have actually run it. "The model said it works" is not a test.
|
||||
|
||||
## Conduct
|
||||
|
||||
Be kind, be direct, assume good faith. Report unacceptable behavior privately via the contact in [SECURITY.md](SECURITY.md).
|
||||
@@ -1,86 +0,0 @@
|
||||
# Security Policy
|
||||
|
||||
Codeman launches AI coding sessions with `--dangerously-skip-permissions`, so the
|
||||
web UI is **by design a remote-code-execution surface for whoever can reach it**.
|
||||
The entire security model exists to control *who* that is. Please read this before
|
||||
exposing an instance beyond `localhost`. The full model lives in
|
||||
[`docs/security-architecture.md`](../docs/security-architecture.md).
|
||||
|
||||
## Supported versions
|
||||
|
||||
Security fixes land on the latest published `codeman@X.Y.Z` release and `master`.
|
||||
Older versions are not patched — upgrade to the latest release (App Settings →
|
||||
Updates for git-clone installs, or `npm i -g aicodeman@latest`).
|
||||
|
||||
| Version | Supported |
|
||||
| ------- | --------- |
|
||||
| latest `0.9.x` / `master` | ✅ |
|
||||
| anything older | ❌ (upgrade) |
|
||||
|
||||
## Reporting a vulnerability
|
||||
|
||||
**Please do not open a public issue for security problems.**
|
||||
|
||||
Report privately via **GitHub's private vulnerability reporting**:
|
||||
the repository's **Security** tab → **Report a vulnerability**
|
||||
(<https://github.com/Ark0N/Codeman/security/advisories/new>). This opens a private
|
||||
advisory thread with the maintainer.
|
||||
|
||||
> Maintainer note: enable *Settings → Code security and analysis → Private
|
||||
> vulnerability reporting* so this channel is live.
|
||||
|
||||
When reporting, please include: affected version/commit, the deployment shape
|
||||
(loopback-only, `CODEMAN_PASSWORD` set, tunnel/`tailscale serve`, custom
|
||||
reverse proxy), reproduction steps, and impact. We aim to acknowledge within a
|
||||
few days. Coordinated disclosure is appreciated — we'll agree a disclosure
|
||||
timeline with you once impact is confirmed.
|
||||
|
||||
### In scope
|
||||
- Authentication / session-cookie bypass when `CODEMAN_PASSWORD` is set
|
||||
- DNS-rebinding, CSRF/CSWSH, or Origin/Host-guard bypass reaching state-changing routes
|
||||
- Remote code execution reachable **without** local OS access (e.g. via a browser, a tunnel, or a foreign origin)
|
||||
- Path traversal / arbitrary file read or write through the HTTP API
|
||||
- Supply-chain integrity of the in-app self-updater
|
||||
|
||||
### Out of scope (by design — see Known limitations)
|
||||
- Anything requiring an already-trusted **same-machine, same-uid** process. Codeman trusts the local OS user it runs as; a peer process of that user is already inside the boundary.
|
||||
- Running an authless instance bound to a non-loopback host after dismissing the startup warning (you explicitly acknowledged it).
|
||||
- The default loopback + no-password posture itself (it is reachable only from the same machine).
|
||||
|
||||
## Trust model (summary)
|
||||
|
||||
- **Loopback by default.** Binds `127.0.0.1`; the no-password default is safe out of the box. Binding a non-loopback host without `CODEMAN_PASSWORD` *starts but prints a loud warning* with concrete fixes.
|
||||
- **Always-on Host + Origin guards.** Block DNS-rebinding and cross-site state-changing requests even on the no-auth loopback install (a missing Origin is allowed so CLI/hooks work).
|
||||
- **Optional auth.** HTTP Basic via `CODEMAN_USERNAME`/`CODEMAN_PASSWORD`; success issues an opaque server-side 256-bit cookie. Per-IP rate limiting on failures.
|
||||
- **Hardened file serving, tmux launch, transport headers, and multi-instance isolation** — see the full architecture doc.
|
||||
|
||||
## Known limitations and accepted risk
|
||||
|
||||
A 1.0 release is an implicit statement that the documented model *is* the model, so
|
||||
these residuals are stated explicitly. Most sit **inside the same-uid OS trust
|
||||
boundary** or behind the always-on Origin guard; they matter mainly for
|
||||
shared-host, multi-user, or tunneled deployments.
|
||||
|
||||
- **Self-update trusts an unsigned release tag.** The in-app updater does `git checkout <tag> && npm install` (lifecycle scripts run) of a tag matched only by name shape, from whatever `origin` points to — no signature/commit verification. Treat the updater as trusting your `origin` remote and your release pipeline. (Hardening tracked for 1.0.)
|
||||
- **CSP ships `'unsafe-inline'`.** Inline handlers mean the Content-Security-Policy is defense-in-depth only; all AI-/file-derived sinks are escaped, but a future missed escape would be executable.
|
||||
- **`workingDir` is unconstrained.** A session may be created with any absolute working directory (e.g. `/`), which becomes the file-route boundary for that session. Scope it to trusted paths on shared hosts.
|
||||
- **Hook-event auth exemption is loopback-IP-based.** `POST /api/hook-event` is exempt from auth for loopback callers; because tunnels (cloudflared / `tailscale serve`) terminate at `127.0.0.1`, a loopback-terminating tunnel inherits the exemption. Set `CODEMAN_PASSWORD` and prefer a tunnel that preserves the client identity if this matters.
|
||||
- **Session cookie is not bound to client IP/UA on reuse, and refreshes without an absolute cap.** A stolen cookie replays until its idle TTL elapses.
|
||||
- **Multi-instance tmux socket is process-wide.** Two Codeman instances on the same `CODEMAN_INSTANCE` share a tmux socket and can attach each other's live sessions — isolate with distinct `CODEMAN_INSTANCE` values.
|
||||
- **The live log-tail route reads `/var/log` and `~/logs`** in addition to the session working directory (read-only) — a deliberate choice for tailing system/app logs. On a password-protected remote deployment an authenticated user can therefore read those roots outside their session. See `docs/security-architecture.md` §5.
|
||||
|
||||
- **The web-tab proxy fetches from the server's network position.** Any authenticated user can save a dashboard URL on loopback or a private range and have Codeman relay to it; that is the feature. Link-local and cloud-metadata addresses are the only refused targets (see below). On a shared host, restrict who holds an account.
|
||||
|
||||
Recent hardening (2026-09-04): the web-tab proxy, its "Test" probe and its
|
||||
WebSocket relay refuse link-local and cloud-metadata targets (`169.254.0.0/16`,
|
||||
`fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`,
|
||||
`metadata.google.internal`), judged on the RESOLVED address so a DNS name pointing
|
||||
there is refused too; proxy capabilities are revoked on logout, admin logout and
|
||||
user deletion; proxied responses carry `Referrer-Policy: same-origin`. Earlier:
|
||||
web-push subscription endpoints are restricted to https public hosts (SSRF guard,
|
||||
rejects internal/metadata IP literals, validated at subscribe and send time), and
|
||||
tmux session names discovered on the shared socket are validated against the
|
||||
safe-name pattern before reaching any shell call site.
|
||||
|
||||
For the detailed rationale, defenses, and recommended secure setups, see
|
||||
[`docs/security-architecture.md`](../docs/security-architecture.md).
|
||||
@@ -1,222 +0,0 @@
|
||||
name: CI
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [master, main]
|
||||
pull_request:
|
||||
|
||||
jobs:
|
||||
ci:
|
||||
name: Typecheck & Lint
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: 22
|
||||
cache: 'npm'
|
||||
|
||||
- name: Install dependencies
|
||||
run: npm ci
|
||||
|
||||
- name: Check package-lock.json version sync
|
||||
run: npm run check:lockfile
|
||||
|
||||
- name: Type check
|
||||
run: npm run typecheck
|
||||
|
||||
- name: Lint
|
||||
run: npm run lint
|
||||
|
||||
- name: Frontend JS syntax check
|
||||
run: npm run check:frontend-syntax
|
||||
|
||||
- name: Format check
|
||||
run: npm run format:check
|
||||
|
||||
# install.sh reaches users through `curl | bash` with nothing between it and
|
||||
# them, and until now nothing in this repo checked it at all: no shellcheck,
|
||||
# no bats, and the vitest gate is Node-only.
|
||||
- name: install.sh syntax
|
||||
run: bash -n install.sh
|
||||
|
||||
# macOS ships bash 3.2 and this runner has bash 5, so the constructs that
|
||||
# actually break a Mac install are invisible here without a container. This
|
||||
# step is what catches them — in particular expanding an EMPTY array under
|
||||
# `set -u`, which bash 3.2 treats as an unbound variable and `bash -n`
|
||||
# cannot see because it is a runtime error, not a syntax one.
|
||||
- name: install.sh runs on bash 3.2 (macOS's version)
|
||||
run: |
|
||||
set -euo pipefail
|
||||
docker run --rm -v "$PWD":/w -w /w bash:3.2 bash -n /w/install.sh
|
||||
docker run --rm -v "$PWD":/w -w /w -e CODEMAN_INSTALL_SH_LIB=1 bash:3.2 bash -c '
|
||||
set -euo pipefail
|
||||
. /w/install.sh
|
||||
detect_all_clis
|
||||
# `shell` declares no binaries, so its offset/length window is length 0.
|
||||
# Iterating it is the empty-array case; reaching here means it did not abort.
|
||||
echo "bash $BASH_VERSION: ${#CLI_IDS[@]} CLIs, $CLI_FOUND_COUNT found"
|
||||
cli_catalog_names >/dev/null
|
||||
cli_catalog_print_install_hints >/dev/null
|
||||
# The install menu with nothing installed and the user answering "s":
|
||||
# skipping must warn and continue, never trip the "failed to install"
|
||||
# gate (it did once, aborting the install before the clone).
|
||||
has_tty() { return 0; }
|
||||
headless_guard() { return 0; }
|
||||
read_reply() { eval "$1=s"; }
|
||||
NONINTERACTIVE=0
|
||||
k=0; while [[ $k -lt ${#CLI_ALL_BINS[@]} ]]; do CLI_ALL_BINS[$k]="no-such-cli-$k"; k=$((k + 1)); done
|
||||
k=0; while [[ $k -lt ${#CLI_ALL_PATHS[@]} ]]; do CLI_ALL_PATHS[$k]="/nonexistent/$k"; k=$((k + 1)); done
|
||||
CLI_DETECT_DONE=""; detect_all_clis
|
||||
offer_ai_cli_install >/dev/null 2>&1
|
||||
echo "bash $BASH_VERSION: skipping the AI CLI install menu continues"
|
||||
'
|
||||
# Issue #382: the dsh identity probe builds an OPTIONAL `timeout` prefix as an
|
||||
# array, and on stock macOS there is no `timeout`, so the array is empty and the
|
||||
# expansion aborts the whole installer under `set -u`. The step above cannot
|
||||
# reach that branch: this image HAS `timeout`, and with no `dsh` on PATH the
|
||||
# probe is never called at all. So hide `timeout` and call it directly.
|
||||
docker run --rm -v "$PWD":/w -w /w -e CODEMAN_INSTALL_SH_LIB=1 bash:3.2 bash -c '
|
||||
set -euo pipefail
|
||||
. /w/install.sh
|
||||
printf "#!/bin/sh\necho \"DeepSeek Harness 0.1\"\n" > /tmp/dsh
|
||||
printf "#!/bin/sh\necho \"dancer shell (Debian dsh)\"\n" > /tmp/not-dsh
|
||||
chmod 755 /tmp/dsh /tmp/not-dsh
|
||||
# A PATH the probe can still work on, minus the binary under test.
|
||||
mkdir -p /tmp/nobin
|
||||
for b in grep sh; do ln -sf "$(command -v $b)" "/tmp/nobin/$b"; done
|
||||
export PATH=/tmp/nobin
|
||||
if command -v timeout >/dev/null 2>&1; then
|
||||
echo "timeout is still on PATH, so this is NOT exercising the empty-array branch" >&2
|
||||
exit 1
|
||||
fi
|
||||
dsh_banner_probe /tmp/dsh
|
||||
if dsh_banner_probe /tmp/not-dsh; then
|
||||
echo "identity probe accepted a foreign dsh" >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "bash $BASH_VERSION: dsh identity probe survives a missing timeout"
|
||||
'
|
||||
# Installer v2: the question phase runs before the build, and every decision it
|
||||
# takes is bash logic over stubbed tailscale state. Drive the flags, the launch
|
||||
# default, the occupied-:443 menu and the rename question with canned answers,
|
||||
# so a bash-4 construct or a flipped default in any of them fails here, not on a
|
||||
# Mac. The JSON parsers need node (absent in this image) and are stubbed; their
|
||||
# own coverage is test/install-sh-invariants.test.ts plus the vitest gate.
|
||||
docker run --rm -v "$PWD":/w -w /w -e CODEMAN_INSTALL_SH_LIB=1 -e HOME=/tmp/h bash:3.2 bash -c '
|
||||
set -euo pipefail
|
||||
mkdir -p /tmp/h
|
||||
. /w/install.sh
|
||||
parse_flags --tailscale --service --name Build-Box --port 4000
|
||||
[[ "$CODEMAN_TAILSCALE" == "1" && "$LAUNCH_PRESET" == "2" && "$TS_NAME" == "Build-Box" && "$CODEMAN_PORT" == "4000" ]]
|
||||
[[ "$(ts_sanitize_name "$TS_NAME")" == "build-box" ]]
|
||||
has_tty() { return 0; }
|
||||
ANSWER=""; read_reply() { eval "$1=\"\$ANSWER\""; }
|
||||
systemctl() { return 0; }
|
||||
LAUNCH_PRESET=""; NONINTERACTIVE=0
|
||||
choose_launch_mode linux >/dev/null 2>&1
|
||||
[[ "$LAUNCH_CHOICE" == "2" ]]
|
||||
check_tailscale() { return 0; }
|
||||
ts_status_field() { case "$1" in "s.BackendState") printf Running ;; "s.Self && s.Self.DNSName") printf "box.tail.ts.net." ;; esac; }
|
||||
ts_backend_state() { printf Running; }
|
||||
ts_dns_name() { printf box.tail.ts.net; }
|
||||
ts_serve_443_target_port() { printf 8080; }
|
||||
ts_serve_find_port_mapping() { :; }
|
||||
ts_serve_port_used() { return 1; }
|
||||
detect_tailscale_serve_url() { :; }
|
||||
tailscale_choose_mapping >/dev/null 2>&1
|
||||
[[ "$TS_SERVE_MODE" == "path" && "$BIND_BASE_URL" == "/codeman" ]]
|
||||
RENAMED=""; tailscale_rename_node() { RENAMED="$1"; }
|
||||
TS_NAME=""; tailscale_choose_name >/dev/null 2>&1
|
||||
[[ -z "$RENAMED" ]]
|
||||
# A flag re-run keeps the password the unit already carries (and so
|
||||
# never writes the unauthenticated ack), and the hand-start line the
|
||||
# done screen prints carries every non-default value.
|
||||
read_existing_binding() { EXISTING_FOUND=1; EXISTING_HOST=0.0.0.0; EXISTING_PASSWORD=s3cret; EXISTING_ACK=0; EXISTING_BASE_URL=""; }
|
||||
CODEMAN_HOST=0.0.0.0; CODEMAN_TAILSCALE=0; unset CODEMAN_PASSWORD; BIND_ACK=0
|
||||
choose_network_binding >/dev/null 2>&1
|
||||
[[ "$BIND_PASSWORD" == "s3cret" && "$BIND_ACK" == "0" ]]
|
||||
BIND_HOST=0.0.0.0; BIND_PASSWORD=x; BIND_ACK=0; BIND_BASE_URL=/codeman; CODEMAN_PORT=4000
|
||||
[[ "$(start_command_hint)" == "CODEMAN_HOST=0.0.0.0 CODEMAN_PASSWORD="*" CODEMAN_BASE_URL=/codeman CODEMAN_PORT=4000 codeman web" ]]
|
||||
RECONFIGURE=0; parse_flags --port 4001; [[ "$RECONFIGURE" == "1" ]]
|
||||
echo "bash $BASH_VERSION: question phase (flags, launch default, occupied :443, rename opt-in, kept password, start line) ok"
|
||||
'
|
||||
|
||||
- name: CLI catalogue artifacts are in sync with stock.ts
|
||||
run: npm run generate:cli-catalog -- --check
|
||||
|
||||
- name: Server boot smoke test
|
||||
run: |
|
||||
set -u
|
||||
if ! command -v tmux >/dev/null; then
|
||||
sudo apt-get update -qq
|
||||
sudo apt-get install -y tmux
|
||||
fi
|
||||
npx tsx src/index.ts web --port 3151 > /tmp/boot.log 2>&1 &
|
||||
SERVER_PID=$!
|
||||
trap "kill $SERVER_PID 2>/dev/null || true" EXIT
|
||||
for i in $(seq 1 30); do
|
||||
if curl -fsS http://localhost:3151/api/status -o /dev/null; then
|
||||
echo "Server booted in ${i}s"
|
||||
exit 0
|
||||
fi
|
||||
if ! kill -0 $SERVER_PID 2>/dev/null; then
|
||||
echo "Server exited before becoming ready. Logs:"
|
||||
cat /tmp/boot.log
|
||||
exit 1
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
echo "Server did not respond on /api/status within 30s. Logs:"
|
||||
cat /tmp/boot.log
|
||||
exit 1
|
||||
|
||||
test:
|
||||
name: Unit & integration tests
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: 22
|
||||
cache: 'npm'
|
||||
|
||||
- name: Install dependencies
|
||||
run: npm ci
|
||||
|
||||
- name: Install tmux
|
||||
run: |
|
||||
if ! command -v tmux >/dev/null; then
|
||||
sudo apt-get update -qq
|
||||
sudo apt-get install -y tmux
|
||||
fi
|
||||
|
||||
- name: Run unit & integration tests
|
||||
# Excludes the suites that need chromium, per-machine PNG baselines or a
|
||||
# quiet machine — see config/test-suites.ts for the list and the reason
|
||||
# behind each entry. Identical to what `npm test` runs locally.
|
||||
# Safe in CI: TmuxManager no-ops all shell commands under VITEST (test/setup.ts).
|
||||
run: npm run test:ci
|
||||
|
||||
- name: Run xterm-zerolag-input package tests
|
||||
# Layers 1-3 of the predictive-echo suites (unit laws, fixture replay,
|
||||
# seeded fuzz): deterministic, no browser, no live server. Depends on
|
||||
# the ROOT `npm ci` above — workspaces hoist the package's vitest into
|
||||
# the root node_modules; do not add a separate install here.
|
||||
run: npx vitest run
|
||||
working-directory: packages/xterm-zerolag-input
|
||||
|
||||
# Note: three suites are excluded from CI, each with its own local runner:
|
||||
# npm run test:browser Playwright + chromium (+ a live server, and a real
|
||||
# codex binary for codex-predictive-echo)
|
||||
# npm run test:mobile the above plus environment-specific PNG baselines
|
||||
# npm run test:perf wall-clock benchmarks; need an otherwise idle machine
|
||||
# config/test-suites.ts holds the globs; the configs derive from it so the
|
||||
# exclusions here and those runners cannot drift apart. Everything else runs in
|
||||
# the `test` job above, which is the same thing `npm test` runs.
|
||||
@@ -1,74 +0,0 @@
|
||||
name: Release
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
|
||||
concurrency: ${{ github.workflow }}-${{ github.ref }}
|
||||
|
||||
jobs:
|
||||
release:
|
||||
name: Release
|
||||
runs-on: ubuntu-latest
|
||||
permissions:
|
||||
contents: write
|
||||
pull-requests: write
|
||||
steps:
|
||||
- name: Checkout repo
|
||||
uses: actions/checkout@v6
|
||||
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: 22
|
||||
cache: npm
|
||||
registry-url: https://registry.npmjs.org
|
||||
|
||||
- name: Install dependencies
|
||||
run: npm ci
|
||||
|
||||
- name: Build
|
||||
run: npm run build
|
||||
|
||||
- name: Create release PR or publish
|
||||
id: changesets
|
||||
uses: changesets/action@v1
|
||||
with:
|
||||
publish: npm run release
|
||||
version: npm run version-packages
|
||||
title: "chore: version packages"
|
||||
commit: "chore: version packages"
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
|
||||
|
||||
- name: Rename release tag to codeman
|
||||
if: steps.changesets.outputs.published == 'true'
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
VERSION=$(node -p "require('./package.json').version")
|
||||
OLD_TAG="aicodeman@${VERSION}"
|
||||
NEW_TAG="codeman@${VERSION}"
|
||||
|
||||
# Update the GitHub release BEFORE deleting the old tag.
|
||||
# make_latest pins the "Latest" badge to the Codeman release. This repo
|
||||
# publishes TWO packages (aicodeman + xterm-zerolag-input), changesets
|
||||
# creates a GitHub release for each, and GitHub awards "Latest" to
|
||||
# whichever was published LAST. That is a race: 1.9.2 kept the badge,
|
||||
# 1.9.4 lost it to xterm-zerolag-input@0.1.7 by two seconds. All package
|
||||
# releases already exist by the time this step runs, so setting it here
|
||||
# is deterministic.
|
||||
RELEASE_ID=$(gh release view "$OLD_TAG" --json databaseId -q .databaseId 2>/dev/null || true)
|
||||
if [ -n "$RELEASE_ID" ]; then
|
||||
gh api -X PATCH "repos/${{ github.repository }}/releases/${RELEASE_ID}" \
|
||||
-f tag_name="$NEW_TAG" \
|
||||
-f name="$NEW_TAG" \
|
||||
-f make_latest=true
|
||||
fi
|
||||
|
||||
# Retag
|
||||
git tag "$NEW_TAG" "$OLD_TAG" 2>/dev/null || true
|
||||
git tag -d "$OLD_TAG" 2>/dev/null || true
|
||||
git push origin "$NEW_TAG" ":refs/tags/$OLD_TAG" 2>/dev/null || true
|
||||
@@ -1,109 +0,0 @@
|
||||
name: Sync Wiki
|
||||
|
||||
# Publishes docs/wiki/ to the repository's GitHub wiki.
|
||||
#
|
||||
# The wiki is a separate git repo with no CI and no review, so the source of truth
|
||||
# lives in docs/wiki/ and this workflow mirrors it. Browser edits to the wiki are
|
||||
# overwritten by the next sync; fix pages with a PR against docs/wiki/ instead.
|
||||
#
|
||||
# One-time setup: GitHub only creates <repo>.wiki.git once the first page has been
|
||||
# saved in the browser. Save a stub page at /wiki/_new before the first run.
|
||||
#
|
||||
# Token: GITHUB_TOKEN can push to the wiki on most repos but not all. If a run fails
|
||||
# with 403, add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret;
|
||||
# it is preferred automatically when present. Note the 403 usually surfaces on the
|
||||
# PUSH, not the clone: this repo is public, so a read-only token still clones the
|
||||
# wiki fine. Both steps carry the hint.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [master]
|
||||
paths:
|
||||
- 'docs/wiki/**'
|
||||
- '.github/workflows/wiki-sync.yml'
|
||||
workflow_dispatch:
|
||||
|
||||
concurrency: ${{ github.workflow }}
|
||||
|
||||
jobs:
|
||||
sync:
|
||||
name: Push docs/wiki to the wiki
|
||||
runs-on: ubuntu-latest
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
steps:
|
||||
- name: Checkout repo
|
||||
uses: actions/checkout@v6
|
||||
|
||||
- name: Clone wiki
|
||||
env:
|
||||
WIKI_TOKEN: ${{ secrets.WIKI_TOKEN || secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
if ! git clone "https://x-access-token:${WIKI_TOKEN}@github.com/${GITHUB_REPOSITORY}.wiki.git" wiki 2>"${RUNNER_TEMP}/clone-err.txt"; then
|
||||
cat "${RUNNER_TEMP}/clone-err.txt"
|
||||
echo "::error::Could not clone ${GITHUB_REPOSITORY}.wiki.git. If this says 'Repository not found', the wiki has never had a page: save one at https://github.com/${GITHUB_REPOSITORY}/wiki/_new and re-run. If it says 403, add a WIKI_TOKEN secret."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
- name: Mirror pages
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
||||
# The mirror deletes before it copies, so an empty source would wipe
|
||||
# every published page and the commit step would happily push that. A
|
||||
# MISSING directory already fails safely (cp aborts under set -e); an
|
||||
# empty one does not, so check explicitly. This is the one failure mode
|
||||
# here that destroys something a browser edit cannot get back.
|
||||
if [ ! -d docs/wiki ]; then
|
||||
echo "::error::docs/wiki does not exist. Refusing to mirror, which would delete the entire published wiki."
|
||||
exit 1
|
||||
fi
|
||||
pages=$(find docs/wiki -maxdepth 1 -name '*.md' | wc -l)
|
||||
if [ "$pages" -eq 0 ]; then
|
||||
echo "::error::docs/wiki contains no .md pages. Refusing to mirror, which would delete the entire published wiki."
|
||||
exit 1
|
||||
fi
|
||||
echo "Mirroring ${pages} pages."
|
||||
|
||||
find wiki -mindepth 1 -maxdepth 1 ! -name '.git' -exec rm -rf {} +
|
||||
cp -R docs/wiki/. wiki/
|
||||
|
||||
- name: Stamp the documented version
|
||||
run: |
|
||||
set -euo pipefail
|
||||
# _Footer.md renders on every page and used to carry a hand-written
|
||||
# version, which went stale on every release because nothing refreshed
|
||||
# it. It carries {{VERSION}} instead and the series is stamped here.
|
||||
series="$(node -p "require('./package.json').version.split('.').slice(0,2).join('.') + '.x'")"
|
||||
# grep exits 1 when it matches nothing, which under `set -o pipefail`
|
||||
# would fail the step instead of warning, so test before substituting.
|
||||
if grep -rlq '{{VERSION}}' wiki/; then
|
||||
grep -rlZ '{{VERSION}}' wiki/ | xargs -0 -r sed -i "s/{{VERSION}}/${series}/g"
|
||||
else
|
||||
echo "::warning::No {{VERSION}} placeholder found in docs/wiki. The published version line can no longer be refreshed automatically."
|
||||
fi
|
||||
if grep -rq '{{VERSION}}' wiki/; then
|
||||
echo "::error::A {{VERSION}} placeholder survived substitution and would be published verbatim."
|
||||
exit 1
|
||||
fi
|
||||
echo "Stamped version ${series}."
|
||||
|
||||
- name: Commit and push
|
||||
run: |
|
||||
set -euo pipefail
|
||||
cd wiki
|
||||
git config user.name 'github-actions[bot]'
|
||||
git config user.email '41898282+github-actions[bot]@users.noreply.github.com'
|
||||
git add -A
|
||||
if git diff --quiet --cached; then
|
||||
echo "Wiki already up to date."
|
||||
exit 0
|
||||
fi
|
||||
git commit -m "docs: sync wiki from docs/wiki @ ${GITHUB_SHA:0:7}"
|
||||
if ! git push 2>"${RUNNER_TEMP}/push-err.txt"; then
|
||||
cat "${RUNNER_TEMP}/push-err.txt"
|
||||
echo "::error::Could not push to ${GITHUB_REPOSITORY}.wiki.git. A 403 here means the token can read the wiki but not write it, which is the usual GITHUB_TOKEN case: add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret."
|
||||
exit 1
|
||||
fi
|
||||
@@ -1,115 +0,0 @@
|
||||
# Claude Code local files
|
||||
.agents/
|
||||
skills-lock.json
|
||||
|
||||
|
||||
# In-session decision scratchpad (context-survival mechanism, not a deliverable)
|
||||
DECISIONS.md
|
||||
# Written by install.sh into end-user clones when setup finishes
|
||||
.install-complete
|
||||
|
||||
# Dependencies
|
||||
node_modules/
|
||||
|
||||
# Build output
|
||||
dist/
|
||||
|
||||
# Vendored frontend deps (generated from node_modules by postinstall/build)
|
||||
src/web/public/vendor/
|
||||
|
||||
# Test coverage
|
||||
coverage/
|
||||
|
||||
# E2E test screenshots (keep baselines, ignore current/diffs)
|
||||
test/e2e/screenshots/current/
|
||||
test/e2e/screenshots/diffs/
|
||||
|
||||
# Mobile visual regression failure artifacts
|
||||
test/mobile/snapshots/*.actual.png
|
||||
test/mobile/snapshots/*.diff.png
|
||||
|
||||
# Logs
|
||||
*.log
|
||||
npm-debug.log*
|
||||
|
||||
# OS files
|
||||
.DS_Store
|
||||
Thumbs.db
|
||||
|
||||
# Editor directories
|
||||
.idea/
|
||||
.vscode/
|
||||
*.swp
|
||||
*.swo
|
||||
*~
|
||||
|
||||
# Environment files
|
||||
.env
|
||||
.env.local
|
||||
.env.*.local
|
||||
|
||||
# Local Compose customisation (host-specific, not part of the project)
|
||||
docker-compose.override.yml
|
||||
docker-compose.override.yaml
|
||||
|
||||
# State files (local to each machine)
|
||||
.claude/ralph-loop.local.md
|
||||
|
||||
# Temporary files
|
||||
*.tmp
|
||||
*.temp
|
||||
|
||||
# Generated output
|
||||
out/
|
||||
screenshots-echo-diag/
|
||||
screenshots-readme/
|
||||
screenshots-readme-real/
|
||||
screenshots-real/
|
||||
scripts/remotion/out/
|
||||
|
||||
# Local UI/README capture scratch (screenshot runs, design mockups). Not build
|
||||
# output, but never meant for git — an unqualified `git add -A` during a COM has
|
||||
# swept dirs like these into a release before.
|
||||
design-explorations/
|
||||
|
||||
# Artifacts that should not be tracked
|
||||
test-results/
|
||||
tmp/
|
||||
# Machine-local working files (never meant for git). ANCHORED so only the root
|
||||
# dir matches.
|
||||
/pr/
|
||||
# Root `public` (a symlink to scripts/remotion/public — local artifact). ANCHORED
|
||||
# with a leading slash so it does NOT also match src/web/public (a bare `public`
|
||||
# would swallow the whole web UI source dir and silently un-stage any new asset
|
||||
# added there). No trailing slash so it still matches the symlink, not just dirs.
|
||||
/public
|
||||
|
||||
# Opt-in gesture overlay runtime assets: large MediaPipe wasm + model (~27 MB)
|
||||
# fetched at build/install by scripts/fetch-gesture-assets.mjs, kept out of git.
|
||||
# (The gesture bundle itself, gesture-codeman.js, IS tracked — built from
|
||||
# packages/gesture-control source by `npm run build:gesture`.)
|
||||
src/web/public/gesture/wasm/
|
||||
src/web/public/gesture/*.task
|
||||
|
||||
# Gesture-control workspace package build outputs (source is tracked; the
|
||||
# Codeman bundle is emitted to src/web/public/gesture/gesture-codeman.js instead).
|
||||
packages/gesture-control/dist/
|
||||
packages/gesture-control/dist-codeman/
|
||||
packages/gesture-control/.vite/
|
||||
|
||||
# Claude Code plan tracking
|
||||
plan.json
|
||||
|
||||
.claude/
|
||||
media-assets/
|
||||
commands
|
||||
todo.md
|
||||
@fix_plan.md
|
||||
readme-preview.mjs
|
||||
|
||||
# Uploaded images land here under each session working dir (runtime artifact)
|
||||
.claude-images/
|
||||
|
||||
# Local-LLM harness smoke-test config (real IPs/keys) — see the .example.json
|
||||
# alongside it in scripts/, which IS tracked as the template.
|
||||
scripts/local-llm-test.config.json
|
||||
@@ -1,31 +0,0 @@
|
||||
dist/
|
||||
coverage/
|
||||
node_modules/
|
||||
src/web/public/vendor/
|
||||
src/web/public/gesture/
|
||||
src/web/public/app.js
|
||||
src/web/public/styles.css
|
||||
src/web/public/mobile.css
|
||||
src/web/public/index.html
|
||||
# Hand-formatted public JS modules (never prettier-enforced; the new
|
||||
# check-public-assets.mjs still validates NUL bytes + JS syntax on these).
|
||||
src/web/public/constants.js
|
||||
src/web/public/image-input.js
|
||||
src/web/public/input-cjk.js
|
||||
src/web/public/keyboard-accessory.js
|
||||
src/web/public/notification-manager.js
|
||||
src/web/public/orchestrator-panel.js
|
||||
src/web/public/panels-ui.js
|
||||
src/web/public/ralph-panel.js
|
||||
src/web/public/ralph-wizard.js
|
||||
src/web/public/respawn-ui.js
|
||||
src/web/public/session-ui.js
|
||||
src/web/public/settings-ui.js
|
||||
src/web/public/sw.js
|
||||
src/web/public/terminal-ui.js
|
||||
src/web/public/voice-input.js
|
||||
src/web/public/upload.html
|
||||
scripts/remotion/
|
||||
|
||||
# Hand-maintained; Prettier escapes underscores in glob paths and corrupts paragraphs.
|
||||
CLAUDE.md
|
||||
@@ -1,17 +0,0 @@
|
||||
# Repository Guidelines
|
||||
|
||||
Canonical agent/contributor guidance for this repository lives in [CLAUDE.md](CLAUDE.md) —
|
||||
project structure, build/test/lint commands, code style, testing safety rules
|
||||
(`npm test` is the CI gate and is safe to run bare; the three excluded suites
|
||||
have their own runners), security notes, and
|
||||
the deployment workflow are all maintained there. Please read it before making
|
||||
changes, and keep it the single source of truth rather than duplicating
|
||||
sections here.
|
||||
|
||||
Quick pointers:
|
||||
|
||||
- Type check: `tsc --noEmit` · Lint: `npm run lint` · Format: `npm run format:check`
|
||||
- Tests: `npm test` (the CI gate, safe to run bare) or `npm test -- test/<file>.test.ts` for one file
|
||||
- Route tests use `app.inject()`; new tests needing ports must pick a unique `const PORT =`
|
||||
- Branch off `master` for all work; Conventional Commit-style messages (`fix(mobile): ...`)
|
||||
- Never commit secrets or local state from `~/.codeman/`
|
||||
@@ -1,491 +0,0 @@
|
||||
# CLAUDE.md
|
||||
|
||||
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
||||
|
||||
> Deep implementation detail lives in [`docs/architecture-invariants.md`](docs/architecture-invariants.md). This file holds the rules that prevent mistakes; that file holds the mechanisms, file inventories, and the history behind each rule. Pointers below are written as `→ architecture-invariants#anchor`. When the goal is raw throughput, [`docs/SPEEDRUN.md`](docs/SPEEDRUN.md) is the fast-execution protocol (it removes ceremony, never the safety rules here). The user-facing manual is `docs/wiki/` (mirrored to the GitHub wiki by CI; see the CI note under Additional Commands), and `AGENTS.md` deliberately just points here.
|
||||
>
|
||||
> **This file is in `.prettierignore` on purpose.** Prettier's markdown printer escapes underscores inside the glob-heavy paths used throughout (`agent-*.jsonl` became `agent-\_.jsonl`, collapsing backtick spans and corrupting a whole paragraph). Do not remove the ignore entry, and do not run `prettier --write` on it.
|
||||
>
|
||||
> **Repo root is kept short on purpose** (the README sits below the file listing on GitHub). Config lives in `config/` (`eslint.config.js`, `knip.json`, the vitest configs), Prettier's config is the `"prettier"` key in `package.json`, and `SECURITY.md` is under `.github/`. Root-only files are the ones tools genuinely require there: `CLAUDE.md` + `AGENTS.md` (loaded from the root by Claude Code / Codex), `CHANGELOG.md` (changesets writes it next to `package.json`), `tsconfig.json`, `.editorconfig`, `.nvmrc`/`.npmrc`, `.prettierignore` (resolved relative to cwd), `LICENSE` (GitHub detection), `.dockerignore` (the build context is the repo root, so Docker resolves it there and nowhere else) and `install.sh` (its raw URL is the published install one-liner). `.claude-plugin/marketplace.json` is root-only for the same reason: `/plugin marketplace add Ark0N/Codeman` reads it from the repo root and nowhere else, which makes the repo its own plugin marketplace. The one plugin it lists is `plugins/codeman/` (manifest + README + a MIRROR of `skills/codeman/`), and ⚠️ the plugin is a small separate directory on purpose: `claude plugin install` copies the plugin root into its cache, and a plugin root that carries a `package.json` gets an **npm install** at install time (measured with the repo root as plugin root: 832 MB, 511 packages and this repo's postinstall on every installer's machine), while a symlink to `skills/codeman` would dangle in the copy. `skills/codeman/` stays the single source; `scripts/sync-plugin.mjs` mirrors it and writes `package.json`'s version into both manifests inside `version-packages`, and `test/plugin-manifest.test.ts` pins the byte-identity, the versions, the absence of a `package.json` in the plugin root and that no other component (`commands/`, `agents/`, `hooks/`, `.mcp.json`, `settings.json`) rides along. `npm run check:plugin` runs that drift check plus both strict validations with EXPLICIT paths (`plugins/codeman`, `.claude-plugin/marketplace.json`): a bare `.` argument copied out of prose reads as a full stop and gets dropped, which surfaces as `missing required argument 'path'`. Needs the `claude` CLI, so it is a local check, not a CI step. Don't relocate those.
|
||||
|
||||
## Quick Reference
|
||||
|
||||
| Task | Command |
|
||||
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Dev server | `npm run dev` (or `npx tsx src/index.ts web`) |
|
||||
| Type check | `npm run typecheck` (= `tsc --noEmit`) |
|
||||
| Lint | `npm run lint` (fix: `npm run lint:fix`) |
|
||||
| Format | `npm run format` (check: `npm run format:check`) |
|
||||
| Tests | `npm test` (the CI gate — safe to run bare) · one file: `npm test -- test/<file>.test.ts` · see Testing for the excluded suites |
|
||||
| Build | `npm run build` (esbuild via `scripts/build.mjs`, NOT tsc — `tsc --noEmit` is type-check only) |
|
||||
| Production | `npm run build && systemctl --user restart codeman-web` |
|
||||
|
||||
## CRITICAL: Session Safety
|
||||
|
||||
**You may be running inside a Codeman-managed tmux session.** Before killing ANY tmux or Claude process:
|
||||
|
||||
1. Check: `echo $CODEMAN_MUX` - if `1`, you're in a managed session
|
||||
2. **NEVER** run `tmux kill-session`, `pkill tmux`, or `pkill claude` without confirming
|
||||
3. Use the web UI or `./scripts/tmux-manager.sh` instead of direct kill commands
|
||||
|
||||
**The working tree is shared with other agent sessions.** Several Codeman sessions run against THIS one checkout, so another session can `git checkout` a different branch, or leave half-finished untracked files, while you are mid-task.
|
||||
|
||||
- **Always `git branch --show-current` immediately before committing.** Observed 2026-07-27: another session ran `git checkout -b feat/web-tabs`, a commit silently landed there instead of master, and the follow-up `git push origin master` cheerfully reported "Everything up-to-date".
|
||||
- To land a commit on master **without** switching branches (which would yank the tree out from under the other session): `git push origin HEAD:master` then `git branch -f master HEAD`. Never `git checkout master` to "fix" it.
|
||||
- **Never `git add -A`/`git add .`** — stage explicit paths. A sweep will pick up another session's WIP.
|
||||
- Another session's broken WIP can block `npm run build`, since `tsc` is the first step and the build gates on it. That is not your bug to fix. ⚠️ `tsc` still EMITS on type errors, so a failed `npm run build` leaves a rebuilt `dist/index.js` compiled from their tree; check what it pulled in before restarting the service. To deploy frontend-only changes past a blocked `tsc`, run the asset stage of `scripts/build.mjs` (everything after the `tsc`/`chmod` lines is independent of it).
|
||||
|
||||
## CRITICAL: Always Test Before Deploying
|
||||
|
||||
**NEVER COM without verifying your changes actually work.** For every fix:
|
||||
|
||||
1. **Backend changes**: Hit the API endpoint with `curl` and verify the response
|
||||
2. **Frontend changes**: Use Playwright to load the page and assert the UI renders correctly. Use `waitUntil: 'domcontentloaded'` (not `networkidle` — SSE keeps the connection open). Wait 3-4s for polling/async data to populate, then check element visibility, text content, and CSS values
|
||||
3. **Only after verification passes**, proceed with COM
|
||||
|
||||
The production server caches static files for 1 year, `immutable` (`maxAge: '1y'` in `server.ts`). To avoid stale frontend after a deploy, `renderIndexHtml` runs `cacheBustAssets(html)` — it appends `?v=<mtime>` to **every same-origin `.js`/`.css`** reference (mtime memoized ~1s so a burst of renders is cheap; external/already-versioned/missing refs untouched). Because `index.html` is served `no-cache`, a **normal reload now picks up edited modules/styles — no hard refresh needed** (the gesture bundle is injected separately with its own `?v=`). If you add an asset referenced by an _absolute_ URL or from JS rather than a `<script>/<link>` tag, it won't be auto-busted. ⚠️ **`index.html` itself is the exception: it is read ONCE into `indexHtmlTemplate` in the `WebServer` constructor**, so editing markup in dev needs a server restart (edited `.js`/`.css` do not) — otherwise you debug a "CSS class that doesn't apply" that is really an element still missing from the served HTML.
|
||||
|
||||
## COM Shorthand (Deployment)
|
||||
|
||||
Uses [Semantic Versioning](https://semver.org/) (`MAJOR.MINOR.PATCH`) via `@changesets/cli`. What SemVer actually covers (the CLI, documented env vars, **and the HTTP/SSE API under `/api/v1`**: endpoint paths, response envelope, `errorCode` values and SSE event names are public/stable; on-disk state, internal TS modules, and experimental features are internal/unstable) is defined in `docs/versioning-policy.md`. Third-party integration surfaces are documented in `docs/extending-codeman.md`. Security reporting + known limitations live in `.github/SECURITY.md`.
|
||||
|
||||
When user says "COM":
|
||||
|
||||
1. **Determine bump type**: `COM` = patch (default), `COM minor` = minor, `COM major` = major
|
||||
2. **Create a changeset file** (no interactive prompts). Write a `.md` file in `.changeset/` with a random filename:
|
||||
|
||||
```bash
|
||||
cat > .changeset/$(openssl rand -hex 4).md << 'CHANGESET'
|
||||
---
|
||||
"aicodeman": patch
|
||||
---
|
||||
|
||||
Detailed description of ALL changes since last release (not just the most recent commit — review full git log since last version tag)
|
||||
CHANGESET
|
||||
```
|
||||
|
||||
Replace `patch` with `minor` or `major` as needed. Include `"xterm-zerolag-input": patch` on a separate line if that package changed too.
|
||||
|
||||
3. **Consume the changeset**: `npm run version-packages` (auto-bumps `package.json` files, updates `CHANGELOG.md`, runs `npm install --package-lock-only`, and verifies lockfile sync via `scripts/check-lockfile-sync.mjs` — all in one command; never hand-edit `CHANGELOG.md` or `package-lock.json` versions)
|
||||
4. **Sync CLAUDE.md version**: Update the `**Version**` line below to match the new version from `package.json`
|
||||
5. **Commit and deploy**: verify the branch first (`git branch --show-current`), then stage EXPLICIT paths — never `git add -A`, which has swept another session's WIP into a release. `git status --short` and account for every line before committing:
|
||||
`git add <paths> && git commit -m "chore: version packages" && git push && npm run build && systemctl --user restart codeman-web`
|
||||
6. **Refresh the getcodeman.com version badge**: the landing page's status bar carries the release version (`v<x.y.z> · getcodeman.com · MIT`), so it goes stale on every release if nobody bumps it. The site source and its deploy script are maintained outside this repository, on the maintainer's machine only; follow the local site handbook there, which also covers the numbers strip and `sitemap.xml` refresh that belong in the same pass. Poll production (`curl -s https://getcodeman.com/ | grep v<x.y.z>`) before calling it done, since the edge lags a deploy by up to a minute. Not applicable to contributor clones — skip it and say so.
|
||||
7. **Wait for CI**: after `git push`, TWO workflows fire per master push — `CI` and `Release` (the npm publish + GitHub release). List both runs for the pushed commit with `gh run list --commit $(git rev-parse HEAD) --json databaseId,workflowName` and watch EACH with `gh run watch <id> --exit-status`. Confirm both pass before considering the release done (`gh run list -L 1` returns only one of the two).
|
||||
|
||||
8. **Announce it in Discussions**: every release gets a post in the [Announcements](https://github.com/Ark0N/Codeman/discussions/categories/announcements) category, shaped like #418 and #302: the short version first (one bold lead-in per theme, features before fixes, what it does for the user rather than how it works), contributor @-mentions inline where their work is described (a mention notifies them, and a contributor reposting is the cheapest reach this project has), a link to the release, and the Thanks names at the end. Casual first-person voice, no em-dashes, humanizer pass when the skill is available. A same-day follow-on patch (1.28.1 after 1.28.0) is folded into the previous post as an edit, never a second thread. Post it with `gh api graphql -F body=@<file> -f title='Codeman <x.y.z>: <hook>' -f repo=R_kgDOQ-SMDg -f cat=DIC_kwDOQ-SMDs4DCHZE -f query='mutation($repo:ID!,$cat:ID!,$title:String!,$body:String!){createDiscussion(input:{repositoryId:$repo,categoryId:$cat,title:$title,body:$body}){discussion{number url}}}'` (the repo's node id and its Announcements category id; `pinDiscussion` does not exist in the API, so pinning stays a click in the UI). This step exists because announcements stopped at 1.18 (#302) while ten releases shipped with nobody notified; #418 is the backfill covering 1.19.0 to 1.28.1, and release notes only count as content once they reach a surface people are subscribed to.
|
||||
|
||||
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
|
||||
|
||||
**Version**: 1.33.1 (must match `package.json`)
|
||||
|
||||
## Project Overview
|
||||
|
||||
Codeman is a Claude Code session manager with web interface and autonomous Ralph Loop. Spawns Claude CLI via PTY, streams via SSE, supports respawn cycling for 24+ hour autonomous runs.
|
||||
|
||||
**Tech Stack**: TypeScript (ES2022/NodeNext, strict mode), Node.js, Fastify, node-pty, xterm.js. Supports Claude Code, OpenCode, Codex (OpenAI), Gemini (Google, enterprise-only since Google's June 2026 consumer cutover), Antigravity (`agy`, Google), Pi (pi.dev), Grok Build (`grok`, xAI), DeepSeek Harness (`dsh`) and OMP (`omp`) CLIs via pluggable CLI resolvers (`SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity' | 'pi' | 'grok' | 'deepseek' | 'omp'`).
|
||||
|
||||
**TypeScript Strictness** (see `tsconfig.json`): `noUnusedLocals`, `noUnusedParameters`, `noImplicitReturns`, `noImplicitOverride`, `noFallthroughCasesInSwitch`, `allowUnreachableCode: false`, `allowUnusedLabels: false`.
|
||||
|
||||
**Requirements**: Node.js 22+, Claude CLI, tmux
|
||||
|
||||
**Git**: Main branch is `master`. Terminal session dashboard: `codeman tui` (`--list` to list, `codeman tui <n>` to attach).
|
||||
|
||||
## Additional Commands
|
||||
|
||||
`npm run dev` = dev server. Default port: `3000` (override with `--port` or the `CODEMAN_PORT` env var). To run this beta isolated alongside a prod Codeman, use `scripts/run-beta.sh` (sets `CODEMAN_INSTANCE=beta` + `CODEMAN_PORT=5000`). Commands not in Quick Reference:
|
||||
|
||||
| Task | Command |
|
||||
| ------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Terminal dashboard | `codeman tui` (`--list` prints the numbered list and exits, `codeman tui <n>` attaches to row n; both short-circuit before any screen setup). Needs a TTY; without a server it starts attach-only. `docs/tui.md` |
|
||||
| Dev with TLS | `npx tsx src/index.ts web --https` |
|
||||
| Override window title hostname | `npx tsx src/index.ts web --title-hostname <name>` (default: `os.hostname()` — `codeman:<name>` is used for tab title, title-flash, and OS desktop notification prefix) |
|
||||
| Bind a non-loopback host | `npx tsx src/index.ts web --host 0.0.0.0` (or `-H`; env `CODEMAN_HOST`; default `127.0.0.1`). Without `CODEMAN_PASSWORD` it **starts but warns loudly** — see Common Gotchas + `docs/security-architecture.md` |
|
||||
| Mount under a reverse-proxy sub-path | `npx tsx src/index.ts web --base-url /codeman` (env `CODEMAN_BASE_URL`; default `/`). Normalized in `src/config/base-path.ts` (`''` = root). See Reverse-proxy base path below + `docs/wiki/Remote-Access.md` |
|
||||
| Continuous typecheck | `tsc --noEmit --watch` |
|
||||
| Watch-mode test | `npm run test:watch -- test/<file>.test.ts` (runs the CI gate's config; pass a file to narrow it) |
|
||||
| Test coverage | `npm run test:coverage` |
|
||||
| Dead-code sweep | `npm run knip` (config in `config/knip.json`, passed via `--config`) |
|
||||
| Rebuild gesture overlay | `npm run build:gesture` (esbuild `packages/gesture-control/src/codeman/entry.ts` → `src/web/public/gesture/gesture-codeman.js`; commit the result) |
|
||||
| Regenerate the CLI catalogue | `npm run generate:cli-catalog` (`--check` to fail on drift). Rewrites `config/clis.stock.json` **and** the marked block in `install.sh` from `stock.ts`. ⚠ **Commit both.** They are what `install.sh` and the Docker agent image read, since neither can import TypeScript; `test/cli-catalog-sync.test.ts` and a CI `--check` step fail if either goes stale. See `docs/cli-registry.md` |
|
||||
| Build the docker agent image | `node scripts/build-agent-image.mjs --no-cache` (builds `codeman/agent:base` from `docker/agent.Dockerfile`; prerequisite for Docker cases; `--engine`/`--image`). ⚠ **Always `--no-cache`** — a plain rebuild re-uses the cached `npm install -g` layer and silently keeps the CLIs frozen at their original versions, which once shipped a BROKEN codex while reporting success. See `docs/docker-cases.md` |
|
||||
| Gesture playground | `npm run dev` **in** `packages/gesture-control/` (standalone vite demo, fake tabs) |
|
||||
| Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) |
|
||||
| Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) |
|
||||
| Excluded-suite runners | `npm run test:browser` · `npm run test:mobile` · `npm run test:perf` · `npm run test:all` (everything, environmental failures included) — see Testing |
|
||||
| Production start | `npm run start` |
|
||||
| Production logs | `journalctl --user -u codeman-web -f` |
|
||||
| Detached server | `codeman web -d` (`--status`, `--stop`; pidfile+log at `dataPath('web.pid'/'web.log')`). ⚠ Refuses to start a 2nd server on one data dir — see Instance isolation |
|
||||
| Install/remove the service | `codeman service install` / `status` / `uninstall` (systemd user unit on Linux, LaunchAgent on macOS; names from `config/service-names.ts`) |
|
||||
| Dependency doctor | `codeman doctor` (alias `check-deps`; `--json`, `--category core\|office\|other`). Probes Node/Claude CLI/tmux/LibreOffice/MS Office against `config/dependency-registry.ts`; engine is pure given an injectable `ProbeHost` |
|
||||
| Multi-user accounts | `codeman users add <name>` / `passwd <name>` / `list` / `rm <name>` (writes `~/.codeman/users.json`, mode 0600; see Multi-user mode) |
|
||||
|
||||
**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 9 Playwright tests; globs live in `config/test-suites.ts`), followed by the **`packages/xterm-zerolag-input` package tests** (a bare `npx vitest run` in that directory; its vitest is hoisted by the root `npm ci`, so no separate install, and `npm test` at the root does NOT run them). `npm test` runs this same config, so local green == CI green. Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing). A third workflow, `wiki-sync.yml`, fires only on master pushes touching `docs/wiki/**` and mirrors that directory to the GitHub wiki (browser edits to the wiki are overwritten by the next sync, so fix pages via `docs/wiki/`).
|
||||
|
||||
**Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`) — config lives in the **`"prettier"` key of `package.json`**, not a `.prettierrc` (keeps the repo root short; editors read it natively). `.prettierignore` stays at the root because Prettier resolves it relative to cwd. ESLint flat config (`config/eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `scripts/remotion/**`.
|
||||
|
||||
**Prettier scope is deliberately narrow.** `npm run format` globs only `src/**/*.ts` and `src/web/public/**` (`lint` only `src/**/*.ts`), and `.prettierignore` then exempts most of `src/web/public/*.js` (app.js, styles.css, **mobile.css**, index.html, upload.html, and 15 hand-formatted modules) plus `CLAUDE.md`. Those files are hand-formatted by design; `npm run check:public-assets` and `check:frontend-syntax` are what guard them (NUL bytes + JS syntax), not Prettier. Do not "fix" a file by adding it back to Prettier's scope.
|
||||
|
||||
## Common Gotchas
|
||||
|
||||
- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink. ⚠️ **Input must END with `\r` or Enter is never sent**: `sendInput()` only issues `send-keys Enter` when the payload contains a carriage return, a `\r`-less `POST /api/sessions/:id/input` still succeeds (send-and-wait even reports `delivered:true`) while the text sits unsubmitted on the composer, and any `wait` burns its whole timeout on a turn that never started. Embedded newlines are stripped, not rejected, so `"echo A\necho B\r"` runs the joined `echo Aecho B`. ⚠️ **Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints** while still taking the typed text (measured 2026-09-19: an Enter at 28 s stranded the prompt, one at 51 s submitted it), so text+`\r` sent at readiness sits unsent with `0 tokens` and a `wait` burns its timeout. So the SERVER verifies every programmatic write that carried a `\r`: `SubmitVerifier` (`session-submit-verifier.ts`, armed from `writeViaMux`) reads the pane on a 2 s to 60 s schedule and re-sends Enter only while the LAST composer line (the CLI's own `promptGlyph`) verifiably still holds the head of what was sent; an empty composer, other text, or no composer line at all (a shell, a direct-PTY session) ends it, and a newer write replaces the schedule. The skill's `sendwait` keeps its own copy of the loop (`_composer_text` in `skills/codeman/preamble.sh`) for servers that predate this. The `shift+tab` footer only means the composer painted, never that Enter is accepted
|
||||
- **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks
|
||||
- **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly. Both `aicodeman` and `codeman` bin aliases are installed (`package.json` `bin`)
|
||||
- **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically)
|
||||
- **`envOverrides` flow `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` / `ANTIGRAVITY_*` / `PI_*` / `GROK_*` / `XAI_*` / `DSH_*` / `DEEPSEEK_*` env vars, plus exact-key `CLAUDE_CONFIG_DIR`** — Set via `POST /api/sessions { envOverrides }`, stored on `Session._envOverrides`, exported by `tmux-manager.buildEnvExports()` at spawn time, persisted in `SessionState.envOverrides`. **Do NOT** write these to `<case>/.claude/settings.local.json` — that's the old path and creates UI/disk drift. (`GOOGLE_*` is the deliberately-broad Vertex-AI namespace for Gemini — see Multi-CLI prefix discipline.) `CLAUDE_CONFIG_DIR` (#255, exact match via `ALLOWED_ENV_KEYS` in `schemas.ts`) points a session at a separate Claude account/config dir for per-client subscriptions; it persists to state.json (a path, not a secret; losing it on restart would silently switch accounts). ⚠️ A relocated config dir writes transcripts outside `~/.claude/projects`, so the response viewer, subagent windows, ultracode panel and Read My Mind capture go blind for that session unless the user symlinks `projects` back into the shared tree (`ln -s ~/.claude/projects <configDir>/projects`). ⚠️ It is also one of claude's `privilegedEnvKeys` (Custom Model Endpoint Profiles, since it can redirect a session's traffic same as any other injected var), so in multi-user mode setting it via `envOverrides` is admin-only, and a non-granted owner's already-persisted `CLAUDE_CONFIG_DIR` is stripped on reboot-restore — silently returning that session to the default Claude account rather than the one it was pointed at (see `session-env-clamp.ts`). → [architecture-invariants#per-session-env-overrides-exact-key-allowlist-and-claude_config_dir](docs/architecture-invariants.md#per-session-env-overrides-exact-key-allowlist-and-claude_config_dir)
|
||||
- **Effort is NOT an env var** — never carry effort as `CLAUDE_CODE_EFFORT_LEVEL`: the env var hard-locks effort and blocks in-session `/effort` switching (incl. ultracode). It flows as the dedicated `effort` payload field → `Session._effort` → `claude --effort <level>` for regular levels incl. `max` (the settings `effortLevel` key is `enum(["low","medium","high","xhigh"]).catch(undefined)` — `max` gets SILENTLY dropped there), or `claude --settings '{"ultracode":true}'` for ultracode (rejected by `--effort`). Both are soft defaults the user can override anytime. Legacy env-var entries are auto-migrated by the Session constructor and unset from tmux sessions in `applyEnvOverrides()`. See `buildEffortCliArgs()` in `session-cli-builder.ts`, tests in `test/effort-injection.test.ts`
|
||||
- **Model choice flows via `settings.local.json`, NOT `--model` or env** — the App Settings **Claude Model** picker (`claudeModel` in `settings.json`) is read by `session-ui.js` at session create (wins over the legacy 1M-Opus toggles `opusContext1m`/`opusContext1mEnabled`), sent as the `modelOverride` payload field, and `updateCaseModel()` (`hooks-config.ts`) writes/deletes the `model` key in `<case>/.claude/settings.local.json`. This is the intended exception to the envOverrides rule above: model legitimately lives in `settings.local.json` (a soft default — in-session `/model` still works); env vars do not
|
||||
- **Multi-CLI prefix discipline** — env-var prefix is CLI-specific (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*` vs `ANTIGRAVITY_*` vs `PI_*` vs `GROK_*` vs `DSH_*`) and the `ALLOWED_ENV_PREFIXES` allowlist in `schemas.ts` enforces this; non-prefix exceptions are exact keys in `ALLOWED_ENV_KEYS` (currently only `CLAUDE_CONFIG_DIR`), never a widened prefix. Gemini additionally allowlists the **broad `GOOGLE_*`** namespace (intentional: Vertex AI auth needs `GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI`; it is the loosest allowlist entry, affecting only the user's own spawned CLI), and Grok allowlists **`XAI_*`** for the same vendor-namespace reason (`XAI_API_KEY` is grok's documented auth var). When adding a setting, decide which CLI(s) it applies to and gate the env export accordingly. Never blanket-forward all prefixes. ⚠️ Pi is the case that proves the rule: its ~34 provider keys (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `HF_TOKEN`, …) share NO prefix, and the allowlist is one GLOBAL list applied by a refine with no mode context, so admitting them for pi would widen it for every mode at once — they stay out, and pi users authenticate via `/login` or the server process's own env. ⚠️ DeepSeek repeats pi's lesson exactly: a dsh `settings.yaml` can nominate ANY env var as a provider credential (`apiKeyEnv`), so only the vendor namespaces `DSH_*` (launcher inputs incl. `DSH_PERMISSION_MODE`) and `DEEPSEEK_*` (`DEEPSEEK_API_KEY`/`DEEPSEEK_BASE_URL`) are admitted; foreign provider keys authenticate from dsh's own files or the server env. Resolver design pattern: `docs/opencode-integration.md`, `docs/pi-integration.md`, `docs/grok-integration.md`, `docs/deepseek-integration.md`
|
||||
- **Zod `.optional()` rejects `null`** — accepts `undefined` only. When the frontend builds a request body with `JSON.stringify`, an explicit `null` field is preserved on the wire and fails validation with `INVALID_INPUT`. Convert `null` → `undefined` before stringifying (e.g. `field: value ?? undefined`), or declare the schema `.nullish()`. This has caused real shipped bugs twice
|
||||
- **Local-echo overlay stays on screen**: the overlay lays its wrapped lines out DOWNWARD from the prompt row, and the text has not reached the PTY yet, so the CLI never learns the prompt is long and nothing scrolls to make room. With the keyboard up only a handful of rows are visible, so a long prompt used to run off the bottom and the user typed blind. The block now grows UPWARD once it would pass the last visible row (optional `totalRows` in `RenderParams`; the line divs are opaque, so they cover transcript above), and a prompt taller than the viewport keeps its TAIL. ⚠️ Separately, `_shrinkPaddingToFit()` (mobile-handlers.js) must never shrink `main`'s padding-bottom below the MEASURED height of the fixed bars: on phones the toolbar and accessory bar are `position: fixed`, so that padding is the only thing reserving room for them, and taking it pulled the terminal's bottom row behind them. Tests: `packages/xterm-zerolag-input/test/overlay-renderer.test.ts`, `test/mobile-keyboard-bottom-padding.test.ts`.
|
||||
- **`xterm-zerolag-input` is single-source** — BOTH echo addons live ONLY in `packages/xterm-zerolag-input/src/`, bundled into TWO **gitignored** vendor files: `vendor/xterm-zerolag-input.js` (buffer overlay, entry `zerolag-input-addon.ts`) and `vendor/xterm-predictive-echo.js` (codex write-through, entry `predictive-echo-addon.ts`) — dev by `scripts/postinstall.js`, prod by `scripts/build.mjs`. `app.js`/terminal-ui.js only **consume** them via `new LocalEchoOverlay(terminal)` / `new PredictiveEchoOverlay(terminal)`; there is no inline copy. So: change the package source, then rerun the bundle step (`npm install` for dev, `npm run build` for prod). **Never hand-edit `app.js` for overlay behavior, and never commit the gitignored vendor bundles.** Always test on mobile after touching it. → [architecture-invariants#xterm-zerolag-input-is-single-source](docs/architecture-invariants.md#xterm-zerolag-input-is-single-source), `docs/local-echo-overlay-plan.md`
|
||||
- **Default bind is loopback-only; non-loopback without a password starts but warns** — the server defaults to `--host 127.0.0.1`. Binding non-loopback (`--host`/`-H`/`CODEMAN_HOST`) without `CODEMAN_PASSWORD` starts anyway but prints a loud warning; `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` acknowledges it. ⚠️ The production systemd unit passes no `--host`, so prod binds **localhost only**: reach it via `tailscale serve`/tunnel to `127.0.0.1`. A loopback bind is reachable through a same-host tunnel but NOT by a browser hitting the box's LAN IP. `install.sh` is separate and prompts for the binding (defaulting to LAN + a password), and preserves the existing binding on re-runs. → [architecture-invariants#default-bind-and-the-non-loopback-warning-path](docs/architecture-invariants.md#default-bind-and-the-non-loopback-warning-path), `docs/security-architecture.md`
|
||||
- **Instance isolation / multi-instance attach danger** — the data dir (`~/.codeman`) and tmux socket (`tmux -L codeman`) are PROCESS-WIDE and shared by every Codeman on the machine, derived from `CODEMAN_INSTANCE` via `src/config/instance.ts`. ⚠️ A 2nd instance on the SAME socket **discovers and attaches PTYs to the first instance's live sessions**, resizing and mutating them. `$HOME` isolation is NOT enough because tmux is system-global. To run two instances, give each a distinct `CODEMAN_INSTANCE` (scopes dir + socket together), or set `CODEMAN_TMUX_SOCKET` + `CODEMAN_DATA_DIR` individually; `scripts/run-beta.sh` does this for a beta alongside prod. **Any new `~/.codeman/...` path MUST go through `dataPath()`**, never `join(homedir(), '.codeman', …)`, and **any new `tmux -L` caller through `resolveTmuxSocketName()`** (both in `config/instance.ts`): the TUI shells out to tmux from a second process, and a hardcoded `codeman` there would point a beta instance at prod's panes. → [architecture-invariants#instance-isolation-and-the-multi-instance-attach-danger](docs/architecture-invariants.md#instance-isolation-and-the-multi-instance-attach-danger)
|
||||
- **node-pty's macOS `spawn-helper` ships without `+x`** (issues #6, #204): `node-pty@1.1.0` publishes `prebuilds/darwin-<arch>/spawn-helper` as mode 0644, and macOS launches every PTY through it, so a stock macOS install fails every session start with `Error: posix_spawnp failed.` **Linux can never reproduce it**: `spawn-helper` is an `OS=="mac"` gyp target and node-pty ships no Linux prebuild, so node-gyp always emits an executable helper there. ⚠️ The flip side of that: since Linux has no prebuild, `npm install` **needs a C/C++ toolchain there** (`make`, `g++`, `python3`), so `install.sh` checks for and installs one alongside Node/tmux/git — a stock Ubuntu 24 server has none and died inside node-gyp with `not found: make`. Do not drop that step. ⚠️ Look in **`prebuilds/<platform>-<arch>/`**, not just `build/Release/`, which does not exist on macOS. Repair is a chmod, never a mandatory rebuild (that would require Xcode CLI tools and deletes `prebuilds/` before compiling): `npm run fix:node-pty` chmods every helper then proves it by really opening a PTY. `spawnPtyWithHelperRepair()` (`utils/node-pty-repair.ts`) wraps every `pty.spawn()` in `session.ts` and self-heals a broken install on the first failure. → [architecture-invariants#node-ptys-macos-spawn-helper-must-be-executable](docs/architecture-invariants.md#node-ptys-macos-spawn-helper-must-be-executable)
|
||||
- **Headless screenshots: `deviceScaleFactor` MUST be 1, and write unique filenames** — under DSF=2 xterm's WebGL renderer draws glyphs at ~2× nominal size while still *reporting* nominal cell dims, so only the pixels reveal it and only the terminal font looks wrong. And overwriting a fixed output path leaves OS image viewers showing the old render, which reads as "the fix didn't work"; `scripts/capture-real-overview.mjs` mints a timestamped filename per run. Seed the per-device `localStorage` keys (`codeman:skin`, `codeman-font-size`, `codeman-app-settings`) so the capture matches a real device. → [architecture-invariants#headless-screenshot-capture](docs/architecture-invariants.md#headless-screenshot-capture)
|
||||
|
||||
**Import conventions**: Utils from `./utils`, types from `./types` (barrel), config from specific `./config/*` files.
|
||||
|
||||
## Architecture
|
||||
|
||||
### Core Files (by domain)
|
||||
|
||||
| Domain | Key files | Notes |
|
||||
| ---------------- | -------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
|
||||
| **Entry** | `src/index.ts`, `src/cli.ts`, `daemon-control`, `service-installer`, `config/service-names`, `cli-style` | The last three back `web -d` / `service install`; `cli-style` is the shared palette/table/spinner/confirm kit |
|
||||
| **TUI** | `src/tui/`: `tui-app` ★ + `tui-client` (the only IO) over a pure core (`-model`, `-layout`, `-render`, `-keys`, `-ansi`, `-composer`, `-approvals`, `-digest`, `-sse`, `-types`) | `codeman tui`, a CLIENT of the server, never a second brain. Design doc: `docs/tui-plan.md`; user guide `docs/tui.md` |
|
||||
| **DeepSeek** | `src/utils/deepseek-cli-resolver.ts`, `src/deepseek-status-shim.ts`, `src/deepseek-web-server.ts` (background `dsh web`, not a session) | `dsh` is a PROFILE LAUNCHER, not an agent; read `docs/deepseek-integration.md` first |
|
||||
| **Session** | `src/session.ts` ★, `session-manager`, `session-auto-ops`, `session-cli-builder`, `session-task-cache`, `session-order` (pure), `session-pty-exit-breaker`, `session-trust-dialog` (pure), `usage-limit-patterns`, `usage-telemetry`; `src/services/unified-session-service.ts` | Pure/unit-tested helpers are split out of `session.ts` on purpose |
|
||||
| **Mux** | `src/mux-interface.ts`, `src/mux-factory.ts`, `src/tmux-manager.ts` ★, `src/proc-tree.ts` (pure, bounded descendant walk) | |
|
||||
| **Respawn** | `src/respawn-controller.ts` ★ + 4 helpers (`-adaptive-timing`, `-health`, `-metrics`, `-patterns`) | Read `docs/respawn-state-machine.md` first |
|
||||
| **Ralph** | `src/ralph-tracker.ts` ★, `src/ralph-loop.ts` + 5 helpers (`-config`, `-fix-plan-watcher`, `-plan-tracker`, `-stall-detector`, `-status-parser`) | Read `docs/ralph-wiggum-guide.md` first |
|
||||
| **Orchestrator** | `src/orchestrator-loop.ts`, `-planner`, `-verifier` | Read `docs/orchestrator-loop-architecture.md` first |
|
||||
| **Cron** | `src/cron/cron-service.ts`, `cron-time.ts` (pure next-run math), `cron-input.ts` | Read `docs/cron-discovery.md` first. Distinct from legacy `ScheduledRun` (`/api/scheduled`) |
|
||||
| **Agents** | `src/subagent-watcher.ts` ★, `team-watcher`, `bash-tool-parser`, `transcript-watcher`, `workflow-run-watcher` | `workflow-run-watcher` is STANDALONE and never touches `subagent-watcher` |
|
||||
| **AI** | `src/ai-checker-base.ts`, `ai-idle-checker.ts`, `ai-plan-checker.ts` | |
|
||||
| **Tasks** | `src/task.ts`, `task-queue.ts`, `task-tracker.ts` | |
|
||||
| **State** | `src/state-store.ts`, `run-summary.ts`, `session-lifecycle-log.ts`, `intent-store.ts`, `tab-layout.ts` (pure model) + `-service` (sole mutation boundary) + `-persistence` + `-legacy-order` | |
|
||||
| **Infra** | `src/hooks-config.ts`, `push-store`, `tunnel-manager`, `image-watcher`, `file-stream-manager`, `remote-hosts` + `remote-reconnect` + `remote-wake` (IO: `dgram`/`net`/`child_process`), `docker-hosts` + `docker-export` | Remote/docker case overlays; see Key Patterns |
|
||||
| **Web tabs** | `src/webview-store.ts`, `webview-capabilities.ts`, `src/web/webview-proxy.ts` (pure), `src/web/routes/webview-routes.ts` | Dashboard URLs as tabs; NOT a SessionMode |
|
||||
| **Search** | `src/search-service.ts` | Pure in-memory core for `GET /api/search` |
|
||||
| **Attachments** | `src/attachment-registry.ts`, `attachment-magic`, `generated-artifact-attachments`, `session-attachment-history`, `document-preview-cache`, `document-thumbnailer`, `document-conversion-limiter`, `config/attachment-guard` | See Key Patterns |
|
||||
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/` (`claude-md.ts` + `case-template.md`) | `templates/` holds the CLAUDE.md scaffold generated into new cases |
|
||||
| **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (one module per domain + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | |
|
||||
| **Frontend** | `src/web/public/app.js` (core) + the modules listed in the Frontend load order + `sw.js` (+ `voice-pcm-worklet.js`, fetched from JS, not in the load order) | See Frontend section for the load order, which is authoritative |
|
||||
| **Types** | `src/types/index.ts` (barrel) → domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts |
|
||||
|
||||
★ = Large, central file (>50KB) — read its `@fileoverview` first. All files have `@fileoverview` JSDoc — read that before diving in. Discovery aid: `grep -l '@fileoverview' src/web/routes/*.ts` lists all route modules; same grep works for `src/types/`, `src/web/public/*.js`.
|
||||
|
||||
**Local packages**: `packages/xterm-zerolag-input/` (local echo overlay, single-source, see Gotchas). `packages/gesture-control/` (`codeman-gesture-control`, hand-tracking overlay source, built via `npm run build:gesture`).
|
||||
|
||||
**Config**: `src/config/` — flat files plus the `cli-registry/` subdir, no barrel (`index.ts`) exists; import from the specific file. ⚠️ There are TWO `config/` directories: the repo-root `config/` holds tooling only (ESLint, knip, the vitest configs, `test-suites.ts`), while runtime config lives in `src/config/`. Throughout this file a bare `config/<name>.ts` in a code context means `src/config/<name>.ts`.
|
||||
|
||||
**Utilities**: `src/utils/` — re-exported via index. Key: `CleanupManager`, `LRUMap` (⚠ NOT in the barrel — import from `./utils/lru-map.js` directly), `StaleExpirationMap`, `BufferAccumulator`, `stripAnsi`, `Debouncer`, `KeyedDebouncer`. Also: `claude-cli-resolver`/`opencode-cli-resolver`/`codex-cli-resolver`/`gemini-cli-resolver`/`antigravity-cli-resolver`/`pi-cli-resolver`/`grok-cli-resolver`/`deepseek-cli-resolver`/`omp-cli-resolver` (CLI path resolution, one per `SessionMode`, all nine sharing the lookup chain in `cli-executable-resolver`: server PATH, then that CLI's install dirs, then an interactive login shell LAST, since it is the only step that spawns anything and it is what finds nvm/Homebrew installs under a service manager's minimal PATH; ⚠ `pi-`, `grok-` and `deepseek-cli-resolver` additionally probe the binary's identity, since `pi` is a generic name, `grok` has npm squatters, and Debian ships an unrelated `dsh`), `file-query` (⚠ Files-panel search matcher, glob-by-two-pointer, never RegExp), `string-similarity` (fuzzy matching), `regex-patterns` (ANSI/token/spinner patterns), `assertNever` (exhaustive checks), `token-validation` (auth tokens), `nice-wrapper` (process priority), `shell-resolver` (⚠ resolves a real login shell for `mode: 'shell'`; the literal string `$SHELL` used to be expanded by the SERVER's shell, which is empty in a container), `event-loop-monitor` (a sync `execSync` freezes the port while the process stays alive, leaving no trace), `dependency-checker` + `dependency-report` (the `codeman doctor` probe engine, registry in `config/dependency-registry.ts`).
|
||||
|
||||
### Data Flow
|
||||
|
||||
1. Session spawns `claude --dangerously-skip-permissions` via node-pty
|
||||
2. PTY output buffered, ANSI stripped, parsed for JSON messages
|
||||
3. WebServer broadcasts to SSE clients at `/api/events`
|
||||
4. State persists to `~/.codeman/state.json` via StateStore
|
||||
|
||||
### Key Patterns
|
||||
|
||||
**Input**: `session.writeViaMux()` for programmatic/curl input via tmux `send-keys -l` + `send-keys Enter`, single-line only. Interactive **browser** input goes through a durable **exactly-once** layer: a stable `clientId` + monotonic per-session `seq` persisted to localStorage until the server ACKs, so a dropped link cannot lose or double-deliver a prompt. `ws-connection-registry.ts` supersedes only same-TAB reconnects, so two tabs on one session coexist. → [architecture-invariants#input-delivery-and-ws-resilience](docs/architecture-invariants.md#input-delivery-and-ws-resilience)
|
||||
|
||||
**Agent wait primitives**: bounded long-polls: `GET /api/sessions/:id/wait`, `GET .../wait-output` (literal substring, **never** regex) and `wait`/`waitTimeout` on `POST .../input`. Registry `session-wait-registry.ts` (pure), bounds `config/agent-wait.ts`. ⚠️ A timeout is a 200 (`wait.timedOut`). ⚠️ `stop`/`blocked` exist for `claude` and `deepseek` ONLY (rule lives in `hooksAvailableForMode()`): explicit request elsewhere is a 400. ⚠️ Send-and-wait registers the waiter BEFORE the write; teardown must `notifySignal('exit')` BEFORE `cancelAll()`; hangup abort listens on `reply.raw` (guarded by `writableFinished`), never `req.raw`; liveness comes from `isPaneDead`, never `session.pid`. ⚠️ Signals are edge-triggered with no history: gather fan-outs via send-and-wait or `wait-output` markers. Packaged as the `skills/codeman` skill (`codeman skill install`, plugin marketplace, or injection behind `agentSkillEnabled`, SYNCED, default OFF): injection is add-only, marker-owned (`applyAgentSkill`), refuses symlinks, and refreshes a marker-owned user-level copy (`refreshUserAgentSkill`). → [architecture-invariants#agent-wait-primitives](docs/architecture-invariants.md#agent-wait-primitives), `docs/api-reference.md`
|
||||
|
||||
**Agent-created case marker** (`src/agent-case-marker.ts`): a case dir that `POST /api/quick-start` CREATES for an agent-driven spawn (signal: the preamble's `X-Codeman-Agent-Origin` header / `agentOrigin` field, else a resolved `parentSessionId`) gets `.codeman-agent-case.json`, published as `agentCreated` on `GET /api/cases`; `GET /api/cases/agent-created` is the cleanup listing (`inUse`, `modifiedAt`) behind Add Case → Manage. ⚠️ Only the branch that CREATES the directory may write it: never label a linked case, cloned repo or pre-existing path (it drives a recursive delete). ⚠️ Reading is total: anything but a well-formed v1 marker reads as not agent-created. ⚠️ Removal stays on `DELETE /api/cases/:name` (the ONE recursive-delete path), and the sweep excludes `inUse` cases. ⚠️ Changing the preamble's headers requires bumping `CODEMAN_PREAMBLE`. → [architecture-invariants#agent-created-case-marker](docs/architecture-invariants.md#agent-created-case-marker)
|
||||
|
||||
**Agent preamble cache GC**: the §0 preamble seeded per claude session (`$XDG_CACHE_HOME/codeman-agent-<id>.sh`) is now REMOVED with the session (`removeAgentSessionPreamble` from `_doCleanupSession`, `killMux` only — a detach leaves the session recoverable and its agent would come back to a loader whose file we deleted) and swept at boot (`pruneAgentSessionPreambles(this.sessions.keys())`, once, after restore, so every session this instance owns is in the keep set). Nothing removed them before: 236 leftovers measured on a working machine, the oldest three weeks old. ⚠️ The sweep needs BOTH guards — never a live session's file at any age (the two-line loader reads it mid-run), and `AGENT_PREAMBLE_MAX_AGE_MS` (7d) of age on top, which is what keeps ANOTHER instance's sessions (whose ids this process cannot see) out of the blast radius. Losing one is degradation, not breakage: the §0 fallback block rewrites it. Tests live with the seed's in `test/agent-skill.test.ts`.
|
||||
|
||||
**Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
|
||||
|
||||
⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer all through a turn, and its working line (`✻ Actualizing… (13m 23s · …)`) is invisible to `SPINNER_PATTERN`, keyword lists and the raw stream. `_confirmIdle()` (session.ts) requires the pane to go quiet AND the SCREEN (`capturePaneText()` + the working-line pattern) to agree; a sustained run of repaints (`session-activity.ts`) marks a turn as started. ⚠️ The composer glyph and working line are per-CLI registry DATA (`capabilities.workDetect`), never Claude constants; a CLI declaring neither falls back to Claude's pair. ⚠️ `workingLine` is config-supplied and runs on the PTY hot path, so it must compile through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()` (ReDoS guard; null, not throw). → [architecture-invariants#idle-detection-composer-glyph-and-working-line](docs/architecture-invariants.md#idle-detection-composer-glyph-and-working-line)
|
||||
|
||||
⚠️ **A quiet pane is not always a pane that wants you.** A CLI can declare an optional `capabilities.workDetect.watchingLine` (a monitor, background shell or cloud hand-off it is still running); the idle probe reads it into `Session.watching` and `notePrompt()` opens that idle item ALREADY acknowledged, so no surface alerts. Only `idle` is eligible, and the label is pane-derived and prompt-injectable, so a pattern must anchor on chrome only that CLI draws. → [architecture-invariants#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you](docs/architecture-invariants.md#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you). Tests: `test/session-watching.test.ts`, `test/watching-no-alert.test.ts`.
|
||||
|
||||
**An exited agent in a live pane** (`paneExit`, #446): panes use `remain-on-exit on`, so `/exit` leaves a pane, session and pid that look alive; `TmuxManager.startPaneExitWatcher()` publishes `SessionState.paneExit` via `session:updated`. ⚠️ Never set `status: 'error'` or null the `pid` for it; the field is TRI-STATE (absent = UNKNOWN, never alive, scoped by `Session.paneExitApplies`); an absent `#{pane_dead_status}` is not 0; a path that starts a command in a pane must clear the record AND persist. A clean exit is CLOSED via `cleanupSession()` (`pane-exit-sweep.ts`): only an explicit numeric status 0 with no signal, confirmed by 2 reads, with no start/attach in flight (`paneLifecycleInFlight`) and not within 10 s of one (a startup error keeps its row); a crashed agent keeps its row. → [architecture-invariants#an-exited-agent-in-a-live-pane-paneexit](docs/architecture-invariants.md#an-exited-agent-in-a-live-pane-paneexit)
|
||||
|
||||
**Dead-pane respawn resume pin** (`_buildRespawnPaneOptionsWithResumePin()`, session.ts): recovering a dead pane, like a custom-model `restartCli()`, must pin the conversation or claude refuses the reused `--session-id`. The pin takes the first transcript-backed candidate (chain tail, launch seed, own id), never `_claudeSessionId`, adds nothing when none is backed, and is never applied to remote or docker sessions. → [architecture-invariants#dead-pane-respawn-the-resume-pin](docs/architecture-invariants.md#dead-pane-respawn-the-resume-pin)
|
||||
|
||||
**Workspace-trust dialog auto-accept** (`session-trust-dialog.ts`, pure): Claude Code's per-directory trust dialog is always answered yes, or the session is stuck. ⚠️ Match the compacted SCREEN (`compactScreenText()`, all whitespace removed), never the stream (tmux sends words joined by cursor-forwards, not spaces). ⚠️ Never answer with a blind `\r` (newer versions highlight "No, exit" first): `trustDialogNextKey()` returns ONE key per re-read frame (arrow, then Enter only once `❯` is on the trust option), and the LAST marked option wins. ⚠️ All three guards must hold: startup window `TRUST_DIALOG_WINDOW_MS` (90s), two-marker match (`isTrustDialogScreen`), attempt cap `TRUST_DIALOG_MAX_ATTEMPTS` (6). ⚠️ Read `capturePaneText()`; only a direct-PTY session falls back to a SHORT buffer tail. ⚠️ The scan must schedule its own next read (`_trustDialogTimer`, cleared in `_clearAllTimers()`), not rely on PTY output. → [architecture-invariants#workspace-trust-dialog-auto-accept](docs/architecture-invariants.md#workspace-trust-dialog-auto-accept)
|
||||
|
||||
**Process-tree walks are bounded** (`proc-tree.ts`, pure): `collectDescendants(pid, byParent)` is the ONE descendant traversal, fed by one cached `ps -eo pid=,ppid=` snapshot (`refreshProcSnapshot()` in tmux-manager.ts: async, in-flight-shared, and ANY error discards the result rather than caching a truncated `ps`). ⚠️ Never walk a process tree with per-node `pgrep` or unbounded recursion (the unbounded version took a machine down): the walk must terminate on cycles, cap depth (`PROC_WALK_MAX_DEPTH`) and node count (`PROC_WALK_MAX_NODES`), and never spawn anything. ⚠️ Keep it in its own module so the test exercises the shipped code, and report truncation through `onTruncated` naming both caps, never silently. → [architecture-invariants#process-tree-walks-are-bounded](docs/architecture-invariants.md#process-tree-walks-are-bounded)
|
||||
|
||||
**Auto-resume on usage limit** (opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit, `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time and `SessionAutoOps` arms a timer for reset+2min, then sends Esc + `continue`. ⚠️ Respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected`), which is what prevents `/clear` from wiping the paused conversation. Claude-mode only. → [architecture-invariants#auto-resume-on-usage-limit](docs/architecture-invariants.md#auto-resume-on-usage-limit)
|
||||
|
||||
**Plan-usage chip** (`showPlanUsageLimits`, per-device: desktop default **ON**, handhelds OFF): resolve DISPLAY only through `planUsageChipEnabled()` (settings-ui.js). The same setting is the server-side COLLECTION switch, read fresh by `readPlanUsageTelemetryEnabled()` (hooks-config.ts) at every claude create/respawn. ⚠️ An ABSENT key reads as ON in the reader; `GET /api/settings` must never write. ⚠️ A save sends `showPlanUsageLimits` ONLY when it flips the chip on that device (`planUsageCollectionFlip()`), or a phone switches collection off for every desktop. Claude data comes from the statusLine exporter, injected as an EPHEMERAL `claude --settings` flag (`resolveStatusLineCliCommand`), never written to disk, WRAPPING a user's own statusLine, posting to `POST /api/status-telemetry`. Codex comes from a read-only `account/rateLimits/read` poll (main bucket only). → [architecture-invariants#plan-usage-chip-statusline-telemetry](docs/architecture-invariants.md#plan-usage-chip-statusline-telemetry), `docs/usage-limits-display-plan.md`
|
||||
|
||||
**Orchestrator**: State machine that turns a user goal into a phased plan and drives it to completion: `idle → planning → approval → executing → verifying → (replanning) → completed/failed`. `OrchestratorLoop` (engine) delegates plan generation to `orchestrator-planner` and per-phase verification gates to `orchestrator-verifier`, executing phases via team agents/`task-queue`. State persists under the `orchestrator` key in `state.json`. Distinct from Ralph (single-session autonomous loop) — orchestrator coordinates multi-phase, multi-agent execution. See `docs/orchestrator-loop-architecture.md`.
|
||||
|
||||
**Cron (`CronJob`s)**: saved, named jobs on a recurring schedule (`once`/`interval`/`daily`/`weekly`) with per-job run history. ⚠️ **Distinct from the legacy `ScheduledRun`** (`/api/scheduled`, a run-now duration-bounded loop); the two never interact and keep separate `Scheduled*` / `Cron*` names. `CronService` **reuses the existing session layer** rather than rebuilding tmux logic. Next-run math is pure and unit-tested in `cron-time.ts` (server-local timezone). The schedule is advanced BEFORE launch so a slow launch cannot re-trigger. → [architecture-invariants#cron-jobs](docs/architecture-invariants.md#cron-jobs), `docs/cron-discovery.md`
|
||||
|
||||
**Remote sessions + remote SSH cases**: a case can point at a remote host. The agent runs in a durable remote `tmux -L codeman-remote`, fronted by a LOCAL tmux pane running `ssh`. Attached (`owned:false`) sessions **detach, never kill**; owned ones propagate `kill-session`. Auto-reconnect (`remoteAutoReconnect`, default ON) revives ONLY when `remoteTmuxSessionAlive()` proves the remote session alive. ⚠️ Classify that probe by EXIT STATUS (`classifyRemoteAliveExit`), never stdout. ⚠️ **Command-injection surface: every ssh command line must flow through `buildSshConnectionArgs()`**; never hand-build one. ⚠️ Remote file reads (`src/remote-files.ts`, attachment routes too) take browser paths only as `shellescape`d tokens, resolve symlinks fail-closed, cap on the REMOTE size, never copy to local disk, are bounded by `src/remote-ssh-limiter.ts`, and pick the host from the SESSION, never the path; no writes over ssh (the `PUT` guard must precede local path validation). ⚠️ Route remote cases through `POST /api/quick-start`, not `POST /api/sessions`. → [architecture-invariants#remote-sessions-over-ssh](docs/architecture-invariants.md#remote-sessions-over-ssh), [#remote-ssh-cases](docs/architecture-invariants.md#remote-ssh-cases), `docs/remote-sessions.md`
|
||||
|
||||
**Wake-on-LAN (`remote-wake.ts`)**: optional `RemoteHost.wakeMac` (magic packet) or `RemoteHost.wakeCommand` (single executable, no shell, takes precedence) lets HTTP input, `POST /api/sessions/:id/wake` and the user's create/attach (`ensureHostAwake`) wake a sleeping host. ⚠️ Only an explicit user request may wake: never give the registry to the auto-reconnect watcher, `handleRemoteSessionDropped`, boot recovery or `cron-service.ts`, and `GET /api/sessions/:id/reachability` must never wake. ⚠️ Detection is a throttled bare TCP probe; never add `ServerAliveInterval`, and a `jumpHost`/`socksProxy`/`ProxyCommand` host is reachability-UNKNOWN (`isProbeable()`): never buffer, gate or banner on it. ⚠️ In multi-user mode a non-admin attach 403s BEFORE host lookup. ⚠️ WS keystrokes bypass the registry, so the banner (`host-wake-ui.js`) must not promise queued input. Waiting requests use the 40 s `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS`. → [architecture-invariants#remote-ssh-cases](docs/architecture-invariants.md#remote-ssh-cases)
|
||||
|
||||
**Docker cases**: a case can point at a **container** running any CLI mode inside it, a **LOCATION OVERLAY on cases, never a `SessionMode`**. One long-lived container **per case**, shared by its sessions: killing a session kills only its in-container tmux, **never** `docker stop` while siblings remain. The workspace is bind-mounted at the **same absolute path**. Credentials are **seeded**, never shared RW. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** A drifted config (label hash) REFUSES the launch. ⚠️ An **adopted** container (`DockerCase.owned === false`) is only `exec`ed into: never create, start, stop, restart, remove, `docker commit` or `docker pause` it; fail closed. Test `owned === false`, never truthiness. ⚠️ Apply `owned` AFTER `dockerConfigHash`. ⚠️ Run modes come from the CONTAINER (`availableModes`), and a failed probe is normal for an OWNED case. ⚠️ Root exec user drops the bypass flag via the registry's `overlays.docker.rootCommand`, never a branch. ⚠️ Adoption is admin-only in multi-user mode. ⚠️ Loopback prod needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for in-container hooks. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md`
|
||||
|
||||
**Docker Compose deployment** (`docker/`): Codeman runs in a container and spawns Docker cases as **SIBLING** containers via the host socket, never nested. `resolveDockerDaemonMountSource()` maps HOME bind sources into the daemon's namespace (`CODEMAN_DOCKER_HOST_HOME`); `CODEMAN_CASES_PATH` makes workspaces resolve to the same absolute path on both sides. ⚠️ `CODEMAN_CASES_PATH` must move every consumer: resolve it only via `config/cases-dir.ts`. ⚠️ `.dockerignore` matches whole paths: keep `**/.env` or `docker/.env` secrets ship in the image. ⚠️ Long-form binds create missing sources ROOT-OWNED: `Start-Codeman.sh` pre-creates them, and `docker/entrypoint.sh` (root, `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against `cap_drop: ALL`, `KILL` for tini; pinned by the test) fixes ownership then drops to `PUID:PGID` via `setpriv`, never re-owning foreign dirs. ⚠️ Append `/opt/codeman-cli` to `PATH`, never prepend. ⚠️ `server.Dockerfile`, the compose file and `.env.example` feed the self-updater's environment gate (`docs/docker-self-update.md`). → [architecture-invariants#docker-compose-deployment](docs/architecture-invariants.md#docker-compose-deployment), `docs/docker-compose.md`
|
||||
|
||||
**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` (discovery, launch argv template, env handling, `capabilities`, and the `overlays` behind remote/docker pane commands). **No code outside `stock.ts` may branch on a CLI id**: use a capability field or a NAMED PROFILE (`profiles.ts`); `test/cli-registry-no-id-branching.test.ts` and `test/frontend-cli-no-id-branching.test.ts` enforce it. ⚠️ Config holds typed argv tokens, never shell text; literals are validated at LOAD time and a bad one rejects the whole entry. ⚠️ Keep `external`, `hooks` and `altScreen` independent; never derive one from another. ⚠️ Config regexes (`discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine`) must compile through `compileVersionRegex()`. ⚠️ `privilegedParams[].param` names a LAUNCH PARAM, not the legacy `<Mode>Config` field (bridged only by `launch.legacyConfigAliases`); a wrong name silently clamps nothing. ⚠️ Resolve the registry AT CALL TIME, never in a module-level const. ⚠️ Remote claude/omp arms of `buildRemoteLaunchCommand` are not covered by the pane-command golden. `~/.codeman/clis.json` overrides entries. It is WRITTEN only by the opt-in CLI management routes (`cliManagementEnabled`, default OFF; `/api/clis`, `cli-registry-routes.ts`), and only through `mutateRegistryFile()` in `registry-writer.ts`, which serializes mutations and refuses (409) a file that does not parse or has group/world permission bits rather than overwriting it. Importing the registry still writes nothing. → [architecture-invariants#cli-registry](docs/architecture-invariants.md#cli-registry), `docs/cli-registry.md`
|
||||
|
||||
**External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token parsing, ❯ readiness; readiness is output stabilization); work detection is per-CLI `capabilities.workDetect` data, not this gate. All eight **require tmux, no direct PTY fallback** (secrets go via socket-scoped `tmux setenv`, never the command line). ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope. ⚠️ **Codex uses predictive write-through echo, never the buffer overlay**: `_predictHookOnData` must never `return` (wire path stays byte-identical), and flushed text and a bracketed paste must go out as separate delayed writes. ⚠️ **Pi**: no bypass flag, never invent one; `approveProjectTrust` executes repo code, so it is in the clamp's **materialize** branch; never wire `--api-key`. ⚠️ **Grok**: `alwaysApprove` is stripped for non-granted owners (only-if-sent). ⚠️ **DeepSeek**: the agent is a PROFILE (Run gates on `isDeepSeekRunnable()`); the permission switch is the `DSH_PERMISSION_MODE` env var, so `clampEnvOverridesForOwner()` must DROP `DSH_PERMISSION_MODE`, `DSH_HOME` and `DEEPSEEK_BASE_URL` for non-granted owners; `hooksAvailableForMode()` is per-SESSION for it (pass `sessionHookOptions(session)`) and is never a stand-in for `mode === 'claude'`; answers come from `deepseek-transcript.ts`, paired by header `cwd` + boot window, never newest-mtime. ⚠️ **OMP**: `OMP_AUTH_BROKER_URL`/`_TOKEN` are clamped the same way. → [architecture-invariants#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp)
|
||||
|
||||
**DeepSeek web UI** (`POST`/`GET`/`DELETE /api/deepseek/web`, `deepseek-web-server.ts`): the Run menu's "DeepSeek web UI..." entry supervises ONE background `dsh web` child process, deliberately **NOT a shell session**. ⚠️ Every piece is load-bearing: exactly one server (a second click REUSES it), restarted when the browser authority changes (`--trusted-host`, last asker wins), killed on shutdown via `stopDeepSeekWeb()` (the detached child would otherwise outlive Codeman and hold its port), and failures returned to the caller. ⚠️ Never hardcode the port: search from 3080 across 40, detect free ports by BINDING, then wait for the server to really answer. ⚠️ Both `POST` and `DELETE` must stay behind `canUsernameRunPrivilegedCommands` (booting a profile runs its plugin code; the server is shared). → [architecture-invariants#deepseek-web-ui](docs/architecture-invariants.md#deepseek-web-ui)
|
||||
|
||||
**Custom Model Endpoint Profiles** (opt-in, `customModelEndpointsEnabled`, SYNCED, default OFF; `docs/custom-model-endpoints.md`): points a session at a user-configured OpenAI-compatible endpoint (store `custom-model-hosts.ts`, `~/.codeman/custom-model-hosts.json`, 0600). `authStyle` is `bearer` or `api-key`, never both headers (hangs the server). The per-CLI redirect is registry data, `capabilities.customModelInjection` (`env` / `configContentEnv` / `configDir` / `unsupported`), computed by the pure `custom-model-injection.ts`; ⚠️ `configDir` writes an isolated per-session config, NEVER the user's real `~/.codex`/`~/.pi`/`~/.omp`/grok config. ⚠️ Claude applies via `Session.restartCli()` (a relaunch in the existing pane, so the conversation is pinned as `resumeSessionId`), the rest one-shot at launch; retired env keys must also be `setenv -u`'d (`_pendingEnvUnsets`), since tmux env survives `respawn-pane`. ⚠️ Remote/Docker sessions are refused (400). ⚠️ Persist only KEYS (`__customModel`), never values (they carry the API key). ⚠️ Every redirectable var must be in that CLI's `privilegedEnvKeys`, and `ANTHROPIC_*` stays out of claude's `allowedPrefixes`. ⚠️ The Run-menu picker builds entries from `window.__codemanCustomModelClis` (escaped via `escapeScriptJson()`), never a hardcoded CLI id list, and launches through `run()` via a temporary `_runMode` swap, never `setRunMode()`. → [architecture-invariants#custom-model-endpoint-profiles](docs/architecture-invariants.md#custom-model-endpoint-profiles)
|
||||
|
||||
⚠️ **llama-swap endpoints** (one model at a time): the apply routes check `GET /running` and return `requiresConfirmation` before evicting a model another live session uses; `confirmedSwap` and `confirmedContext` are SEPARATE flags and must stay so. Claude alone gets a context floor (`CLAUDE_MIN_SAFE_CONTEXT_TOKENS`); context is parsed from `/running`'s `cmd`, never trusted from `/props`. Backend log lines come from llama-swap's `/api/events` `upstream` source, never `/logs`. → [architecture-invariants#custom-model-endpoint-profiles](docs/architecture-invariants.md#custom-model-endpoint-profiles)
|
||||
|
||||
**Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms) so a double click cannot create duplicate `w<n>-<case>` sessions; `_ensureCreatedSessionVisible()` runs before `selectSession()` and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first both render exactly one tab. ⚠️ **Closing has the mirror-image race**: `closeSession()` must read `wasActive` BEFORE its `await` and announce the delete via `_closingSessions`, and `_onSessionDeleted` skips the active-session handoff for ids in that set; never read `activeSessionId` after the fact. The fallback picks the first `sessionOrder` entry still in `sessions`. Tests: `test/session-close-fallback.test.ts`. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization)
|
||||
|
||||
**Session lineage lines** (tab → tab it spawned, `sessionLineageLines`, per-device, desktop default ON): a create request may name its spawner via a `parentSessionId` body field or the `X-Codeman-Parent-Session` header; `resolveParentSessionId()` (route-helpers.ts) resolves it (exact id or unique ≥8-char prefix, live, visible, same owner) and ⚠️ anything unresolvable is DROPPED, never a 400. Rides `toState()`, no new SSE event. ⚠️ Rendering is a LAYER on the existing SVG pass (`_appendLineageConnectionLines` at the tail of `_updateConnectionLinesImmediate()`), geometry pure in `computeLineagePath()`: one U-bridge shape hanging from the strip bottom, colors keyed on the SPAWNING tab and memoized (never by draw index). ⚠️ Desktop only (z-index vs the fixed mobile header). ⚠️ Paths must keep `data-agent-id="lineage:<childId>"` (the entrance animation queries it); skip edges whose endpoint is scrolled out of the strip. → [architecture-invariants#session-lineage-lines-tab--tab-it-spawned](docs/architecture-invariants.md#session-lineage-lines-tab--tab-it-spawned)
|
||||
|
||||
**Auto-named sessions** (`autoNameSessions`, SYNCED, default OFF): a placeholder tab (`w3-myapp`) takes its first real prompt as a title in the `<prefix>: <title>` form, so the case identity and `w<n>` counter survive. Ownership is `SessionState.nameSource` (`placeholder` | `auto` | `manual`; the `name` setter / `PUT /api/sessions/:id/name` makes it `manual`, never touched again). ⚠️ `applyAutoName()` flips to `auto` even if the string is unchanged, so only the FIRST titled prompt names the tab. ⚠️ Only user input counts: `SessionWriteOptions.fromUser` is set by the browser WS path and `POST /api/sessions/:id/input` ONLY; any new user-input path must set it (and the send-key Shift+Enter path must call `trackUserInput()`). ⚠️ The pure tracker (`session-auto-name.ts`) sits on the raw keystroke stream with an explicit rule per key; add a rule for any new key class. ⚠️ `nameSource` also decides `--name`: only a `manual` name is pinned on the claude CLI (`Session.cliPinnedName`), since `--name` is also the `/resume` title; a rename appends a `custom-title` row to a LOCAL, non-docker transcript, and a same-name PUT is a no-op (never flips to `manual`). Tests: `test/session-auto-name.test.ts`. → [architecture-invariants#auto-named-sessions-first-prompt--tab-title](docs/architecture-invariants.md#auto-named-sessions-first-prompt--tab-title)
|
||||
|
||||
**Maintainer bot (external)**: the Telegram bot that reviews open PRs and triages discussion threads in Codeman sessions used to live at `scripts/pr-bot/`. It moved OUT of this repository on 2026-09-14, to `~/codeman-cases/prbot/` (its own private git repo, systemd unit `codeman-pr-bot`, guide + agent rules in its own `README.md` and `CLAUDE.md`). It is a CLIENT of Codeman's HTTP API like any other, so nothing here depends on it and it is not part of the server, the CLI or the npm package. ⚠️ It spawns real sessions named `prbot-<n>` / `dscbot-<n>` on the local Codeman and holds clones under `~/.codeman/pr-bot/`, so those session names and that data dir are taken; it also fetches PR heads into `refs/pr-bot/*` of this checkout and must never check out, reset or clean it. The CHANGELOG entries for 1.25.0 and earlier still describe it, which is history rather than drift.
|
||||
|
||||
**Unified session list**: `GET /api/sessions/unified` merges live sessions, persisted state, lifecycle-log history and transcript files into one deduped list (pure core `src/services/unified-session-service.ts`), backing the Cmd+K Session Manager, pinning and cross-device tab order (`PUT /api/session-order`, `src/session-order.ts`). ⚠️ Transcript history is THREE stores (`~/.claude/projects`, `~/.omp/agent/sessions`, `~/.codex/sessions`), folded via the `claudeSessionId → Codeman id` alias map (not Claude-only despite the name). ⚠️ `resumeId` is set by a SCANNER row only, never a live session; every surface that re-projects these rows (phone overview included) must carry it through, or a tap silently starts a second conversation. → [architecture-invariants#unified-session-list-and-session-manager](docs/architecture-invariants.md#unified-session-list-and-session-manager)
|
||||
|
||||
**Owner tab layouts** (`tab-layout*.ts` + `GET`/`PUT /api/tab-layout`): named tab GROUPS over the flat strip, scoped per owner (`@single` when multi-user is off), persisted as `tabLayouts` in state.json. BACKEND ONLY: no frontend calls these routes yet. ⚠️ `TabLayoutService` is the single mutation boundary (one completed server action = at most one versioned write); never write layout state from a route or manager directly. ⚠️ The layout PROJECTS onto `PUT /api/session-order` via `tab-layout-legacy-order.ts`; change both sides together. ⚠️ Reconciliation is gated on a SUCCESSFUL restore (`markRestorationComplete`/`assertDeletionReady()`): a failed restore must leave the layout untouched or live tabs get pruned. → [architecture-invariants#owner-tab-layouts](docs/architecture-invariants.md#owner-tab-layouts)
|
||||
|
||||
**Hook events**: Claude Code hooks trigger via `/api/hook-event` (`permission_prompt`, `elicitation_dialog`, `elicitation_complete`, `elicitation_response`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`, `prompt_submitted`); see `src/hooks-config.ts` and `docs/claude-code-hooks-reference.md`. ⚠️ Every claude session installs the hooks block into its workspace (add-only merge) from every create path and from `restoreMuxSessions()`, gated by `workspaceHooksEnabled` (SYNCED, default ON). ⚠️ Route that decision through `applyWorkspaceHooks`, never call `ensureCodemanHooks` at a new site, or the setting silently stops applying. ⚠️ An AskUserQuestion / plan-selection dialog arrives as `permission_prompt` (RED alert), not `elicitation_dialog` (MCP elicitation). → [architecture-invariants#hook-events-and-workspace-hook-installation](docs/architecture-invariants.md#hook-events-and-workspace-hook-installation)
|
||||
|
||||
**Reboot restore** (`src/reboot-restore.ts` pure + `web/reboot-restore-registry.ts` + `routes/reboot-restore-routes.ts` + `reboot-restore-ui.js`): after a host reboot kills every pane, Codeman holds an IN-MEMORY plan of the destroyed sessions and a banner offers to rebuild them. ⚠️ The heuristic only decides whether to ASK. ⚠️ Rebuild is TAKE-then-build (entries leave the plan before the first `await`, single-flighted per owner). ⚠️ Re-check grant, workspace and already-live at click time, reading already-live FRESH per entry; key confinement on the entry's OWNER, never the caller. ⚠️ Rebuilt sessions come back disarmed: no respawn/Ralph, `rearmAutoResumeSchedule: false`, and pass `nameSource` through. ⚠️ Undo a failed rebuild with `discardPartiallyBuiltSession()`, NEVER `cleanupSession()`. Claude-mode only, never remote/docker. Tests: `test/reboot-restore.test.ts`. → [architecture-invariants#reboot-restore](docs/architecture-invariants.md#reboot-restore)
|
||||
|
||||
**Approvals Inbox** (`approvalsInboxEnabled`, SYNCED, default OFF; the store and answer endpoints run regardless): `web/approval-inbox.ts` is an in-memory, claude-only queue fed by `/api/hook-event`, at most ONE item per session, answered via `POST /api/approvals/:id/answer` through `writeViaMux` (menu answers never carry `\r`). ⚠️ Accept `option` digits ONLY if they match options parsed from a fresh RE-CAPTURE of the pane; a dialog no longer on screen is a 409. ⚠️ Resolve permission/question items only via the pane-verified `verifyStillAnswerable()`; the heuristic `working` signal alone may resolve `idle` items only. ⚠️ `applyCapture()` is ADD-ONLY for `options`. ⚠️ Viewing ACKNOWLEDGES an idle item (never resolves it) and only a human selection does: app-made selections pass `selectSession(id, { auto: true })`; `_ackDelivery` spends the IDLE alert only. ⚠️ `handleInit` seeds tab alerts from `GET /api/approvals` REGARDLESS of the setting. → [architecture-invariants#approvals-inbox](docs/architecture-invariants.md#approvals-inbox)
|
||||
|
||||
**Read My Mind intent profiles** (`readMyMindEnabled`, SYNCED, default OFF; `docs/readmymind-plan.md`): per-CASE profiles (goals + recent prompts) keyed by owner + realpath(workingDir), captured from the transcript (`transcript:user_prompt`), never the input paths; the listener must stay inside `startTranscriptWatcher()`'s `if (!watcher)` block. Store `src/intent-store.ts` → `intents.json`, ⚠️ written 0600 tmp+rename and never fed to `/api/search` (prompts carry secrets). Routes in `readmymind-routes.ts` (ownership via `findSessionOrFail` WITH `req`); predictor = pure `readmymind-context.ts` + IO in `readmymind-collectors.ts` + `readmymind-predictor.ts`, claude-only, one in flight per session (409). ⚠️ Suggestions render via value/`textContent` ONLY and nothing auto-sends, ever. Frontend `readmymind-ui.js`. → [architecture-invariants#read-my-mind-intent-profiles](docs/architecture-invariants.md#read-my-mind-intent-profiles)
|
||||
|
||||
**Voice dictation via Claude** (`claudeVoiceEnabled`, SYNCED, default OFF): the mic transcribes through this machine's Claude Code login (the CLI `/voice` backend) instead of Deepgram; the browser captures, `src/web/voice-stream.ts` relays to Anthropic. ⚠️ The OAuth token never reaches the page. ⚠️ Credentials are READ-ONLY (`src/claude-credentials.ts`): never refresh them (it rotates the refresh token and can sign the user out of their CLI). ⚠️ Capture must be linear16/16 kHz/mono via an AudioWorklet; `voice-pcm-worklet.js` borrows voice-input.js's `?v=` token, so edit the two together. ⚠️ Claude transcript frames are cumulative: replace, never append. → [architecture-invariants#voice-dictation-via-claude](docs/architecture-invariants.md#voice-dictation-via-claude)
|
||||
|
||||
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`.
|
||||
|
||||
**Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit)
|
||||
|
||||
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and Shell loads the rest only via **Load full history**, never on ordinary scroll. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
|
||||
|
||||
**Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions)
|
||||
|
||||
**Terminal touch gestures: link taps and text selection**: on touch devices xterm's linkifier and SelectionService never see the gesture, so both are driven explicitly (terminal-ui.js). ⚠️ A tap activates the link under it through the SAME provider as the hover linkifier (`_terminalLinkAtPoint`), synchronously inside `touchend` (keeps the user gesture `window.open` needs) and BEFORE any mouse report; the caret's logical line (`_tapIsOnCaretLine`) and TUI-owned rows (`_isActionableMobileTerminalTap`) keep their meaning. ⚠️ Gate on the caret line, never on tap intent (a shell calls every tap `'input'`). ⚠️ Long-press selects via xterm's public `select()`; keep the three guards: suppress the compat mouse pair after `touchend`, the bounded focus guard + `contextmenu` suppression for the platform long-press, and no closing `terminal.focus()` on phones. Tests: `test/terminal-touch-tap.test.ts`. → [architecture-invariants#terminal-touch-gestures-link-taps-and-text-selection](docs/architecture-invariants.md#terminal-touch-gestures-link-taps-and-text-selection)
|
||||
|
||||
**Auto Copy (copy-on-select)** (`autoCopySelection`, per-device, default OFF): a finished terminal selection lands on the clipboard with no keystroke. ⚠️ Copy at the END of a gesture, never in `onSelectionChange` (per-cell); it only arms `_autoCopyPending` and a document-level `mouseup` flushes. ⚠️ The flush must be SYNCHRONOUS in the handler (both clipboard paths need user activation); never defer it to a timer. ⚠️ Touch needs its own calls from `_endTouchSelectionGesture()`/`_selectTouchSelectionLine()` (no mouseup arrives). ⚠️ Unlike `copyTerminalSelection()`, never clear the selection or focus the terminal; restore prior focus. Guards are pure in `decideAutoCopy()` (constants.js, 1M-char cap, refused not truncated). Tests: `test/terminal-auto-copy.test.ts`. → [architecture-invariants#auto-copy-copy-on-select](docs/architecture-invariants.md#auto-copy-copy-on-select)
|
||||
|
||||
**Ctrl+V paste trap** (`image-input.js`): `Ctrl+V` routes through `_handleImagePaste()`, which focuses a hidden `contenteditable` trap and reads the clipboard from the paste event landing there; images upload and their paths are typed in, text goes through `terminal.paste()` so bracketed-paste markers survive. ⚠️ **The trap must consume exactly ONE paste event** (Firefox delivers two per keypress: the `execCommand('paste')` event and the keydown's default action); the one-shot flag lives on the trap, never on a browser check. ⚠️ Do not remove the `execCommand('paste')` call: on some mobile engines it is the only route into the trap, and the trap is the only place image blobs are read. Tests: `test/image-paste-trap.test.ts`. → [architecture-invariants#terminal-paste-ctrlv](docs/architecture-invariants.md#terminal-paste-ctrlv)
|
||||
|
||||
**Terminal scrollback strip + wheel/touch forwarding**: codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity/omp get a NARROW strip (alt-screen toggles only). ⚠️ Gated on `useMux`: direct-PTY sessions must keep the alt screen. Wheel and touch forward to the CLI for **claude ≥ 2.1.187 ONLY**; ⚠️ never re-add codex without a fresh measurement (it ignores SGR wheel reports). ⚠️ `getClaudeCliVersion()` must never cache a FAILED probe. ⚠️ Hand-report clicks only while the CLI has mouse tracking on: `_shouldReportMouseToCli()` gates all three report sites on `cliMouseTracking` (from `_recordStrippedMouseMode()`, session.ts), or a plain shell prints the reports as literal text. Read `_logScrollRouting()` before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
|
||||
**Detached start + service install**: `codeman web -d` relaunches the same entry script `detached:true` (setsid); `nohup` is not what makes it survive. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile + `/api/status` probe), or a second instance attaches to the first one's live sessions. ⚠️ Never report success not observed: poll `/api/status` until the child answers or dies. `--stop` must verify the pid still looks like Codeman (`ps -o command=`) before signalling. Unit/label names live only in `config/service-names.ts`. `service install` bakes the installing shell's PATH into the unit and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install)
|
||||
|
||||
**Self-update** (App Settings → System → Updates): in-app updater for git-clone installs under a supervisor (`systemd`, `launchd`, `launchd-daemon`, `docker-compose`, else `none`). The work runs in a DETACHED `scripts/self-update.sh` writing `update-status.json`, polled across the restart; pure helpers in `src/web/self-update.ts`. ⚠️ Compose: the restart kills the script, so nothing may be appended after the `restarting` marker; the repo must stay a host bind mount over `/opt/codeman` and the image must keep devDependencies + toolchain. ⚠️ `evaluateEnvironmentGate()` refuses releases that change `server.Dockerfile`/`docker-compose.yaml` or add `.env.example` keys, re-evaluated on `POST /api/system/update`; unknowns fail OPEN, but the exit-to-restart needs `--restart-by-exit 1` (`CODEMAN_RESTART_BY_EXIT=1` only in the Compose file). ⚠️ Keep the agent CLIs in `server.Dockerfile` pinned. → [docs/docker-self-update.md](docs/docker-self-update.md), [architecture-invariants#self-update](docs/architecture-invariants.md#self-update)
|
||||
|
||||
**Reverse-proxy base path** (`--base-url` / `CODEMAN_BASE_URL`, default `/`; pure single source `src/config/base-path.ts`, normalized to `''` or `/foo`): mounts Codeman under a sub-path behind a proxy that forwards the prefix unchanged. Few choke points: `stripBasePath()` in Fastify's `rewriteUrl` (routes stay prefix-agnostic; unprefixed requests still answer), one `onSend` hook rebasing `Location`, `renderIndexHtml` rewriting `<base href>` + injecting `window.__CODEMAN_BASE__`, and `CodemanBase.url()` (constants.js) for runtime URLs. ⚠️ Keep template asset refs RELATIVE, and route every root-absolute frontend URL (EventSource/WebSocket/`window.open`/src) through `CodemanBase.url()`. ⚠️ Web-tab proxy egress goes through `proxyPrefixFor(cap, basePath)`; ingress parsers stay base-agnostic. ⚠️ `--base-url` must ride `buildWebArgs` and `resolveServicePlan`. Tests: `test/base-path.test.ts`. → [architecture-invariants#reverse-proxy-base-path](docs/architecture-invariants.md#reverse-proxy-base-path)
|
||||
|
||||
**Attachments** (live external document references; all wiring in `file-routes.ts`): a **registry** maps a stable `attachmentId` to a realpath-resolved, extension-allowlisted absolute path, so browser requests never carry arbitrary absolute paths. ⚠️ The **magic-link scanner** (`codeman://attach?...` in terminal output) is **prompt-injectable**, so its scan path is force-confined to the session workspace; a hostile prompt could otherwise exfiltrate arbitrary host files over SSE. The security gate is an extension **allowlist**, not a blocklist. `document-conversion-limiter.ts` caps converter spawns globally: without it, N large docs detected at once fork N multi-minute processes, which is a resource-exhaustion vector. → [architecture-invariants#attachments](docs/architecture-invariants.md#attachments)
|
||||
|
||||
**File-path links (terminal + chat)**: a path an agent prints is clickable on BOTH surfaces and opens the file-preview overlay. ⚠️ ONE pattern (`FILE_PATH_LINK_PATTERN` / `absoluteFilePathPattern()` in constants.js) feeds the xterm link provider AND `_linkifyFilePaths()`, a fresh instance per call (`lastIndex`). The chat linkifier walks TEXT NODES with DOM APIs, never rebuilds sanitized markup as a string. ⚠️ An out-of-workspace path goes through the ATTACHMENT routes (`POST /api/sessions/:id/attachments` with `notify: false`), never by widening `file-content`/`file-raw` or `file-stream-manager`'s `tail -f` allowlist. ⚠️ `TEXT_ATTACHMENT_EXTENSIONS` IS `EDITABLE_EXTENSIONS` (never a second list), and widening READ must never widen RUN: `html`/`htm`/`svg` stay download-only, other text is inert `text/plain`+`nosniff`. Media extensions are single-sourced in `attachment-registry.ts`. → [architecture-invariants#file-path-links-terminal--response-viewer](docs/architecture-invariants.md#file-path-links-terminal--response-viewer)
|
||||
|
||||
**Filesystem path picker** (Link Existing "Browse" + the mobile keyboard's `📁 Path` key): lazy one-directory browsing via `GET /api/filesystem/browse`, with `GET /api/filesystem/preview` for the tapped file. Inserts the path **without** Enter, so the prompt is never submitted; the sibling `⌫ All` key clears only the unsent prompt and must never send the agent's `/clear`. ⚠️ This is a **second file-serving surface and inherits neither the attachment confinement nor its ownership scoping** — it allowlists Home, `CASES_DIR`, `/mnt/d` and `CODEMAN_FILE_PICKER_ROOTS`, blocks sensitive trees, and rejects symlink escapes **after** `realpath`. ⚠️ The optional `sessionId` is an ownership boundary that must be `canAccessOwned`-checked by hand (it does not go through `findSessionOrFail`), and in multi-user mode a non-admin gets only their own `userSpacePath` as a root: per-user spaces live INSIDE `homedir()`, so a `Home` root exposes every other user's workspace. Previews go through the same global conversion limiter, and Markdown/TXT/JSON are served as inert `text/plain`. → [architecture-invariants#filesystem-path-picker](docs/architecture-invariants.md#filesystem-path-picker)
|
||||
|
||||
**File Viewer edit mode** (issue #212): the file-preview overlay edits workspace text files in place — `GET .../file-content?edit=1` + `PUT /api/sessions/:id/file-content`, policy in `src/config/file-editing.ts`. This is a **third file surface and the only one that WRITES**: read-path confinement (realpath + workspace + ownership) plus sensitive/blocked/`.git` denies and an extension **allowlist**; writes are `wx`-temp + rename (no `O_CREAT` anywhere = edit-in-place is structural); optimistic concurrency via sha256 `baseHash` → 409. ⚠️ `edit=1` never truncates and the client must never save a plain-preview buffer (the 500-line truncation would silently delete the rest). ⚠️ CRLF/UTF-8 guards: EOL re-applied server-side, non-UTF-8 refused via round-trip compare. → [architecture-invariants#file-viewer-edit-mode](docs/architecture-invariants.md#file-viewer-edit-mode), `docs/file-viewer-edit-plan.md`
|
||||
|
||||
**Files panel search** (COD-236, the `q` param on `GET /api/sessions/:id/files`): `compileFileQuery()` (`utils/file-query.ts`, pure) compiles the query into a predicate the server-side walk prunes with; a query returns a FLAT match list and the walk recurses past non-matching directories. An empty, whitespace-only or overlong (`MAX_QUERY_LENGTH`, 256) query compiles to `null`, keeping the default tree response byte-identical. ⚠️ **Never compile a glob into a RegExp** (`*a*a*a…` backtracks and freezes the event loop for the whole server): `globMatch()` is a two-pointer wildcard walk. → [architecture-invariants#files-panel-search](docs/architecture-invariants.md#files-panel-search)
|
||||
|
||||
**Raw file bodies are streamed and range-aware**: `file-raw`, the attachments `/raw` route and `GET /api/download` share `sendFileBody()`, advertise `Accept-Ranges: bytes` and answer `Range` with `206` + `Content-Range` (single-range, parser in `src/web/http-range.ts`); without it `<video>` cannot seek. The size cap (`MAX_FILE_DOWNLOAD_BYTES`, default 2GB, env `CODEMAN_MAX_DOWNLOAD_BYTES`, `0` = unlimited) is a sanity bound, not memory protection; never reintroduce a whole-file buffer. ⚠️ Bodies go out via `reply.hijack()`, so `sendRawStream` must copy the status onto `reply.raw` by hand or a partial body ships as `200`. ⚠️ Closing the preview must pause and unload media (`_stopFilePreviewMedia`), since a detached `HTMLMediaElement` keeps playing. → [architecture-invariants#raw-file-bodies-streamed-and-range-aware](docs/architecture-invariants.md#raw-file-bodies-streamed-and-range-aware)
|
||||
|
||||
**Ultracode / workflow-run visualization** (opt-in, default OFF): the Workflow tool writes a completion artifact only at run *end*, so live in-flight runs exist solely as transcript dirs. `workflow-run-watcher.ts` therefore synthesizes ACTIVE runs from transcripts until the completion artifact appears and supersedes them. It is **STANDALONE** and deliberately never imports or touches `subagent-watcher.ts`, despite reading the same tree. Two independent toggles: `showUltracodeAgents` (docked panel) and `ultracodeFloatingWindows` (floating windows); the watcher starts if **either** is on. → [architecture-invariants#ultracode--workflow-run-visualization](docs/architecture-invariants.md#ultracode-and-workflow-run-visualization)
|
||||
|
||||
**Clone a repository as a case** (issue #236, Add Case → **Clone Repo**): `POST /api/cases/clone` clones synchronously into the caller's case space (bounded by `GIT_CLONE_TIMEOUT_MS`, no job store); `POST /api/cases/clone-preflight` checks anonymous cloneability and lists refs. Core in `src/git-clone.ts`. ⚠️ **The URL is a code-execution surface**: refuse every `::` form and a leading `-`, spawn only argv arrays with `--` before operands. ⚠️ Stay non-interactive (`gitNonInteractiveEnv()`) or the open request hangs; never collect credentials, refuse `user:password@` URLs. ⚠️ Timeout kills the process GROUP, remove the destination only if this attempt created it, and repo contents win over scaffolding (existing `CLAUDE.md` kept, hooks merged, repo `.claude/settings*` warned about). → [architecture-invariants#clone-a-repository-as-a-case](docs/architecture-invariants.md#clone-a-repository-as-a-case)
|
||||
|
||||
**Cross-session search**: `GET /api/search` federates an in-memory search over session metadata, run-summary events, and attachment-history entries. The pure core `searchSources()` does substring matching with hard per-type caps: **no regex (so no ReDoS) and no filesystem reads (so no traversal)**. The server-private `externalPath` is never read. PAST sessions (#261) come from `session-history-index.ts`, a capped snapshot of the unified list filled **outside** the request path (`/api/sessions/unified` publishes it; a stale one is rebuilt fire-and-forget), that indirection is what keeps the no-fs property. ⚠️ The snapshot is stored UNSCOPED with a per-row owner and MUST be re-filtered through `canAccessOwned()` on read; history rows carry `jumpTo.kind:'resume-session'`, since a closed session has no tab to select. → [architecture-invariants#cross-session-search](docs/architecture-invariants.md#cross-session-search)
|
||||
|
||||
**Web tabs** (dashboard URLs as tabs): a saved URL renders as a tab beside sessions, **NOT a `SessionMode`**, and is proxied through Codeman's own origin (`/webview/<cap>/...`). ⚠️ The proxy is NOT an API surface: its capability-based auth exemption stays fenced to safe methods and non-route paths (`test/webview-auth-exemption.test.ts`). ⚠️ Iframes omit `allow-same-origin` unless `trusted`, and `Authorization`/`codeman_session` are stripped upstream in both modes. ⚠️ **Egress guard**: link-local and cloud-metadata targets are refused at save time, by a sync hostname check at each connect (IP literals skip DNS), AND on the resolved address (`webview-egress.ts`); use the `undici` package's own `fetch` + `Agent`, never Node's global fetch. ⚠️ Loopback links in agent output auto-open as proxied web tabs (`openLinkThroughWebTabIfLoopback`), but never auto-route `*.localhost` (prompt-injectable DNS). Capabilities are revoked on logout (`revokeOwner`). → [architecture-invariants#web-tabs](docs/architecture-invariants.md#web-tabs), `docs/web-tabs.md`
|
||||
|
||||
**Multi-user mode** (opt-in `--multiuser` / `CODEMAN_MULTIUSER=1`, OFF by default): named users with scrypt-hashed passwords in `~/.codeman/users.json`. Gated everywhere by `isMultiUserMode()`; when OFF, behavior is byte-identical to single-user because every scoping helper short-circuits. ⚠️ **Not a security boundary at the agent layer**: every session still runs as the SAME OS account. This separates WORKSPACES; it does not sandbox users (Docker cases are the isolation story). Ownership threads through `Session.owner` and is enforced in `findSessionOrFail`, list endpoints, SSE routing (fail-closed), WS, search, and file-preview. → [architecture-invariants#multi-user-mode](docs/architecture-invariants.md#multi-user-mode), `docs/multi-user-plan.md`
|
||||
|
||||
**Away digest**: `GET /api/away-digest` aggregates what happened while you were away from the lifecycle log, run-summary events, live sessions, token stats, and recent subagents. Pure aggregator in `web/away-digest.ts`. ⚠️ Returns `{success:true,digest}`, a legacy raw-ish shape consistent with the other raw GET handlers in `system-routes.ts`; frontend and tests read `.digest`. → [architecture-invariants#away-digest](docs/architecture-invariants.md#away-digest)
|
||||
|
||||
**Ralph todo-config**: per-session `maxTodos` (FIFO-eviction cap, default 500 = `MAX_TODOS_PER_SESSION`) + `todoExpirationMinutes` (auto-expiry, default 60) set via `POST /api/sessions/:id/ralph-config` (`RalphConfigSchema`, both `.int().positive()`). Stored on the tracker (`setMaxTodos`/`setTodoExpirationMinutes`) and **persisted/read-back via `RalphTrackerState`** (surfaced in the `loopState` getter → `toState()` + SSE broadcast → modal `populateRalphForm`), mirroring how `maxIterations` round-trips. Claude-only (skipped by `isExternalCliMode`).
|
||||
|
||||
**Port interfaces**: Routes declare dependencies via port interfaces (`src/web/ports/`). Routes use intersection types (e.g., `SessionPort & EventPort`).
|
||||
|
||||
### Frontend
|
||||
|
||||
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `i18n.js`(1.5) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `terminal-keycode229-recovery.js`(5.55) → `sanitize-html.js`(5.6) → `app.js`(6) → `tab-rail-resize.js`(6.5) → `terminal-ui.js`(7) → `terminal-split.js`(7.5) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `cron-ui.js`(9.7) → `settings-ui.js`(10) → `panels-ui.js`(11) → `readmymind-ui.js`(11.3) → `ultracode-panel.js`(11.5) → `approvals-ui.js`(11.6) → `reboot-restore-ui.js`(11.65) → `admin-ui.js`(11.7) → `session-ui.js`(12) → `host-wake-ui.js`(12.2) → `webview-tabs.js`(12.5) → `mobile-overview.js`(12.55) → `home-sessions.js`(12.56) → `entrance-animations.js`(12.6) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15) → `ultracode-windows.js`(15.5) → `session-lineage.js`(15.6) → `image-input.js`(16). `i18n.js` translates static + newly inserted application DOM while skipping terminal/response/file/user-name surfaces; `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData). `terminal-keycode229-recovery.js` forwards a committed `input` event that xterm's `_inputEvent` guard drops (Chrome-on-Android soft keyboards send `composed: true` after a keydown), and only when xterm emitted no canonical data for that keystroke. ⚠️ **That decision is settled at the NEXT keydown as well as on its own zero-delay timer** (#441): the drain runs from xterm's custom key handler, which fires BEFORE xterm processes that key, so a soft keyboard that commits the last character and sends Enter in one InputConnection transaction puts the character on the wire ahead of the `\r`. On the timer alone that character is not merely late, it is LOST: xterm emits the `\r` first and bumps the canonical counter past the candidate's snapshot, so the candidate stands down (measured, `hell\r` where the user typed `hello`). The trade is that a keydown decides with less evidence than the timer did, since xterm's own keyCode-229 rescue has not run yet; that is safe for Enter, which clears the textarea so the pending diff emits nothing. Ordering is pinned by `test/terminal-keycode229-recovery.browser.test.ts`, which the CI gate does NOT run.
|
||||
|
||||
**Entrance animations** (`entrance-animations.js`, all OFF by default): opt-in animations for tabs, terminal, windows and connection lines, chosen via `data-tab-anim` / `data-term-anim` / `data-win-anim` / `data-line-anim` on `<html>`; the default `legacy` theme short-circuits every hook. ⚠️ Tabs and lines are destroyed mid-animation on re-render, so re-apply to the fresh element by id with a negative `animation-delay` (resume, never restart). ⚠️ Terminal-pane styles may animate only transform / opacity / clip-path (anything else resizes the PTY via FitAddon); `blur` is the ONE sanctioned `filter` exception, do not generalise it. ⚠️ Line glow lives in `--line-glow` so blur keyframes interpolate. Persisted per-device in `codeman:*Anim` localStorage keys, never in `SettingsUpdateSchema`; lab at `?animlab=1`. Test: `test/entrance-animations.test.ts`. → [architecture-invariants#entrance-animations](docs/architecture-invariants.md#entrance-animations)
|
||||
|
||||
**Mobile tab strip scrolling** (issue #257): under 768px the tab strip scrolls horizontally, so the active tab must be kept reachable. `_updateActiveTabImmediate()` reveals it via `computeTabScrollLeft()` (constants.js, rect math on the strip's own `scrollLeft`, never `scrollIntoView()`, which scrolls the document under the fixed header); `_fullRenderSessionTabs()` must restore `scrollLeft` across rebuilds and re-reveal only when the active tab changed (`_lastRenderedActiveTabId`). ⚠️ The phone-block `min-width` on `.session-tab.active .tab-name` keeps the tab's centre off the gear/close icons, sized for numberless tabs 10+ (floor 40px); do not shrink it. ⚠️ Never reintroduce hoisting the active session to the front of the strip. Test: `test/mobile-tab-tap-zones.test.ts`. → [architecture-invariants#mobile-tab-strip-scrolling](docs/architecture-invariants.md#mobile-tab-strip-scrolling)
|
||||
|
||||
**Session list layout: header strip or left sidebar** (`sessionListLayout`, default `header`; per-device via `displayKeys`, also in `SettingsUpdateSchema`): the list can move into a collapsible `<aside>` (Alt+B, `toggleSessionSidebar`) or, via `tabOrientation`, a resizable vertical `#tabRail` (desktop/tablet only). ⚠️ There is ONE `#sessionTabs`, MOVED between hosts, never a second list: `applySessionListLayout()` runs first, then `applyTabOrientation()`, both BEFORE `applyTabWrapSettings()`, and both arm/disarm `_startSidebarRichClock()`. ⚠️ Axis decisions use `_isVerticalTabList()`, never `isSessionSidebarActive()` alone. ⚠️ Rich rows share one gate, `isRichTabRows()`; rich CSS pairs sidebar+rail with comma-grouped selectors, never `:is()`, and card rules stay rail-scoped. ⚠️ Rail sort (`tabRailSort`, default `activity`) is the flex `order` property only, never a DOM reorder; the arrow-key walk alone follows computed `order`. ⚠️ Leaving sidebar mode clears `_sidebarFilter`; the handheld overlay drawer is `inert` when closed, the docked rail never. → [architecture-invariants#session-list-layout-header-strip-vs-left-sidebar](docs/architecture-invariants.md#session-list-layout-header-strip-vs-left-sidebar)
|
||||
|
||||
**Phone overview home screen** (`mobile-overview.js`, per-device `mobileOverviewEnabled`, default ON): under 600px the "C" logo shows NEEDS YOU / CURRENT / PAST SESSIONS instead of the welcome overlay, branched in `showWelcome()`/`hideWelcome()` via width-driven `shouldUseMobileOverview()`. ⚠️ The container ships `hidden` and only this module removes it: never give `.mobile-overview` a bare `display` rule (desktop does not load `mobile.css`). ⚠️ The split Run button must carry the toolbar's own classes (`btn-toolbar btn-run mode-<backend>` / `btn-run-gear`) and mobile.css must set no `background`/`color` on it; row status must mirror the session-tab alert language. PAST rows resume through the shared `resumeHistorySession()`. Status pills carry `data-i18n-skip`. → [architecture-invariants#phone-overview-home-screen](docs/architecture-invariants.md#phone-overview-home-screen)
|
||||
|
||||
**Desktop home tab rail** (`home-sessions.js`, desktop only): the welcome overlay's left gutter carries the open tabs as a rail docked flush left, full height, in overview order, each row showing `created … · <state> <duration>` from `_mobileOverviewSince()`; state classification is reused from mobile-overview.js (so it loads after it). ⚠️ The number badge is the Alt+1..9 tab-strip index, never renumber it to row position. ⚠️ The width gate lives in two places that must stay equal: `HOME_SESSIONS_MIN_WIDTH` (1180) and a `max-width: 1179px` media query. ⚠️ `.home-sessions[hidden]` must re-assert `display: none`. ⚠️ Size all children in `em` off the one `clamp()` knob, never `rem`/px. Age stamps tick in place (`_tickHomeSessionsTimes()`), never by re-render. Test: `test/home-sessions.test.ts`. → [architecture-invariants#desktop-home-tab-rail](docs/architecture-invariants.md#desktop-home-tab-rail)
|
||||
|
||||
**Home-screen session order** (`CodemanSessionOrder` in constants.js, pure): BOTH home screens (phone overview, desktop rail) must order rows through this ONE comparator. Rank `needs` → `error` → `waiting` → `working` → `idle` → `done`. ⚠️ The tiebreak flips: states a session is still IN sort oldest-first, states it has STOPPED sort newest-first. ⚠️ The running group keys off `lastSubmitAt`, never `lastActivityAt` (a working pane repaints constantly). ⚠️ A 0 stamp means unknown and sorts last within its state. Final tiebreak is `orderIndex`, so the list never shuffles. The tab strip itself is NOT sorted by this. Test: `test/session-overview-order.test.ts`. → [architecture-invariants#home-screen-session-order](docs/architecture-invariants.md#home-screen-session-order)
|
||||
|
||||
**Welcome "Resume Conversation" list** (terminal-ui.js): `loadHistorySessions()` fetches once and caches the corpus on `_historyAll`/`_historyCases`; every subsequent view (filter box, sort select, expand, the periodic refresh in panels-ui.js) goes through `_renderHistoryList()`, so never append rows to `#historyList` directly or re-fetch to re-sort. ⚠️ The box height is **class-driven**: expanding the list without `.history-list.expanded` leaves the collapsed `max-height` in place and just deepens a scroll well, which is the bug #260 reported (35 sessions in a ~4-row box). ⚠️ The A–Z sort keys off `_historyRowLabel()`, the SAME string the row renders (`name || firstPrompt || path`), most rows are transcript-backed and have no session name, so sorting on `name` alone silently does nothing. ⚠️ A filter implies expansion, and `_renderSearch()` hides `#historyHeader` (title + controls) as one unit while a search is active. Tests: `test/history-list-controls.test.ts`.
|
||||
|
||||
**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)**: with no selection it must `return true` without `preventDefault()` or the interrupt is lost; keep `copyTerminalSelection` out of `SHORTCUT_ACTIONS`. The gate tests the CLEANED selection (`CodemanCopySelection.clean`: trailing padding, plus a LEADING margin only up to the width the CLI declares in `capabilities.transcriptGutter`, never one derived from the pane); the strip is not idempotent, so clean once and pass the RAW selection on, and leave Alt+drag column selections untouched. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry)
|
||||
|
||||
**Per-device vs synced settings**: the `displayKeys` set in settings-ui.js is a **client-side merge policy**, not a wire filter. A display key seeds from the server only when localStorage has no value for it, which is what prevents one device overwriting another; `showPlanUsageLimits` is additionally `delete`d from the incoming payload outright. Separately, `SettingsUpdateSchema` is `.strict()` and simply **does not declare** `skin`, `showFileViewerButton`, `showCronButton`, `webglRendererEnabled`, `localEchoEnabled`, `cjkInputEnabled`, or `extendedKeyboardBar`, so sending one of those is a validation error. The rest (`showResponseViewer`, `showPlanUsageLimits`, `language`, and most `show*` keys) ARE in the schema and do persist server-side; they are per-device by client policy only. ⚠️ Adding a new per-device setting means deciding **both** questions: membership in `displayKeys`, and presence in the schema.
|
||||
|
||||
**Settings surface** (`#appSettingsModal` + `#sessionOptionsModal` + `#createCaseModal`): one `set-*` language shared through a single `:is(...)` id scope in styles.css. App Settings' rail is a table of contents over ONE scrolling document (`switchSettingsTab` scrolls); Session Options and Add Case really switch (`switchOptionsTab` / `switchCaseModalTab`), and their larger per-modal size blocks are the design, not drift. ⚠️ **The load/save contract is `getElementById` by id**: renaming or dropping a control id silently stops it loading or saving. ⚠️ The Session Options "Session" entry still keys off `context` (label-only rename). ⚠️ Add Case keeps its legacy `.form-row` markup via an adapter; every `<details>` there needs `.set-adv-chev` plus both marker suppressions. ⚠️ Model cards and the effort segment are views over hidden `<select>`s, which stay the source of truth. ⚠️ `.modal-tabs*` classes are retired; `admin-ui.js` needs `.set-rail-items` + `.set-doc` to survive any restructure. Guard: `test/app-settings-structure.test.ts`. → [architecture-invariants#settings-surface-app-settings-session-options-add-case](docs/architecture-invariants.md#settings-surface-app-settings-session-options-add-case)
|
||||
|
||||
**Header button visibility**: most header controls are opt-in and hidden by a marker class (`btn-multimonitor--hidden`, `btn-response-viewer-header--hidden`, `btn-file-viewer--hidden`, `btn-cron--hidden`) that `applyHeaderVisibilitySettings()` (settings-ui.js) toggles after settings load; the multi-monitor button is instead stripped at render by `renderIndexHtml`. ⚠️ Hiding must go through the marker class: the base rules are `display:inline-flex !important`, so an inline style cannot override them. Current desktop default is WS/CPU/MEM + File Viewer + gear, with the token chip and lifecycle-log button OFF. ⚠️ New header controls must not leak onto phones; `test/mobile-header-buttons-policy.test.ts` is the static guard. → [architecture-invariants#header-button-visibility-multi-monitor-response-viewer-file-viewer-cron](docs/architecture-invariants.md#header-button-visibility-multi-monitor-response-viewer-file-viewer-cron)
|
||||
|
||||
**Gesture control** (camera hand-tracking overlay, opt-in, default OFF): `CODEMAN_GESTURE=1` makes the feature *available*; `gestureControlEnabled` turns it on. The bundle is injected by `renderIndexHtml` only when enabled, which is why that method is `async` and reads settings with `readSettings(true)` (a fresh read: a post-save reload lands inside the 2s cache TTL and would otherwise render the pre-toggle state). **Source lives in `packages/gesture-control/`; edit there, run `npm run build:gesture`, and commit the regenerated bundle** because dev serves the committed bundle with no runtime bundler. The MediaPipe wasm + model are fetched separately and gitignored. ⚠️ Keep `MP_VERSION` in `fetch-gesture-assets.mjs` in sync with `@mediapipe/tasks-vision`. → [architecture-invariants#gesture-control-the-source-package](docs/architecture-invariants.md#gesture-control-the-source-package)
|
||||
|
||||
**Terminal font weight** (`terminalFontWeight` / `terminalFontWeightBold`, per-device, default = xterm's own `normal`/`bold`): Claude Code's markdown bold is a bare `ESC[1m`, so the weight step is its only cue. `CodemanTerminalFont.resolveWeights()` (constants.js, pure) resolves each slot against **its own** xterm default. ⚠️ The `@font-face` for `fonts/jetbrains-mono-variable.woff2` must stay declared `100 800` (the browser synthesizes from the descriptor, not the file); narrowing it silently makes the setting a no-op. ⚠️ A live save must reach both echo overlays (`refreshFont()`) and open Agent Teams panes. ⚠️ Leave `_awaitTerminalFont()` untouched. Test: `test/terminal-font-weight.test.ts`. → [architecture-invariants#terminal-font-weight](docs/architecture-invariants.md#terminal-font-weight)
|
||||
|
||||
**Theme skins / branding / i18n**: `skin` selects a palette via `data-skin` on `<html>`, applied by an **inline pre-paint script** in `index.html` reading `localStorage['codeman:skin']` to avoid a flash of wrong theme. ⚠️ A skin is **four things that must stay in sync**, and missing any one degrades silently: the `html[data-skin="…"]` token block in `styles.css`, the xterm ANSI palette in `terminal-ui.js`, the pre-paint allowlist, and the Settings picker (both in `index.html`). `test/skin-themes.test.ts` is the static guard. Light skins additionally need `color-scheme: light` and xterm `minimumContrastRatio: 4.5`, and `applyTerminalSkin()` must call the local-echo overlay's `refreshFont()` because it caches the terminal fg/bg. `displayName` changes user-facing browser branding only and must NEVER rename npm package, CLI, API, storage, CSS, or protocol identifiers. `language` (`en`/`zh-CN`) keeps English as the canonical source so live switching stays reversible. User display names flow through `textContent`/attribute APIs and the server title's HTML escaper, never `innerHTML`. → [architecture-invariants#theme-skins](docs/architecture-invariants.md#theme-skins)
|
||||
|
||||
**Foldable settings identity**: responsive layout is width-driven via `MobileDetection.getDeviceType()`, but the localStorage namespace uses `MobileDetection.isHandheldDevice()` so an unfolded Android foldable keeps `codeman-app-settings-mobile`. ⚠️ Do not switch per-device settings namespaces from instantaneous viewport width: a posture-triggered WebView reload would lose opt-in UI. Regression profile: `OPPO Find N5 (unfolded)` in `test/mobile/devices.ts`. → [architecture-invariants#foldable-settings-identity](docs/architecture-invariants.md#foldable-settings-identity)
|
||||
|
||||
**Folding devices: a fold is not a keyboard, and dialogs avoid the hinge**: ⚠️ in `handleViewportResize()`, a visual-viewport resize that changes the WIDTH is a shape change (rotation, fold) and must never be read as the keyboard; it re-baselines instead, or `keyboardVisible` latches with no keyboard. ⚠️ `init()` must seed `lastViewportWidth`. ⚠️ With the keyboard up, a shape change baselines to `window.innerHeight`, never the shrunk visual height. ⚠️ The hinge is reserved via `--fold-inline-end`/`--fold-block-end` (0px when unfolded): each overlay fold rule must re-state its own gutter, a base gutter overridden by a later `@media` block needs its own fold restatement there on a zero base, and dialogs use physical sides (left/top segment) in every language. Guard: `test/foldable-layout.test.ts`. → [architecture-invariants#folding-devices](docs/architecture-invariants.md#folding-devices)
|
||||
|
||||
**WebGL renderer toggle** (`webglRendererEnabled`, per-device): the GPU-stall watchdog's sticky `codeman-webgl-disabled` marker survives page loads and is cleared only by an explicit OFF→ON save or `?webgl=force`. `?nowebgl` forces the DOM renderer per-load. → [architecture-invariants#webgl-renderer-toggle](docs/architecture-invariants.md#webgl-renderer-toggle)
|
||||
|
||||
**Shell keyboard accessory bar + one-shot Ctrl** (`keyboard-accessory.js`): a shell-mode session swaps the mobile accessory bar for terminal controls; `setMode()` records `extendedKeyboardBar` as the base layout and `refreshForActiveSession()` resolves base-vs-shell. ⚠️ Ctrl is a one-shot modifier applied in `terminal.onData` (after `shouldSuppressTerminalQueryResponse`, before every send path), and must skip `isTerminalFocusOrMouseReport()` chunks. ⚠️ It must disarm on use, second tap, any other accessory key, session switch, keyboard dismissal and layout swap. ⚠️ `_handleCjkInput()` must apply it too (the CJK textarea bypasses onData). Mapping: `ctrlByteFor()`. ⚠️ mobile.css's light-skin repaint must keep excluding `.accessory-btn:not(.armed)` or the armed state is invisible. → [architecture-invariants#shell-keyboard-accessory-bar-and-one-shot-ctrl](docs/architecture-invariants.md#shell-keyboard-accessory-bar-and-one-shot-ctrl)
|
||||
|
||||
**Mobile prompt composer** (`keyboard-accessory.js`): the agent bars' Paste key is **Compose**, a native multiline dialog where only **Send** submits (the shell bar keeps plain Paste). Opening it adopts the whole terminal prompt (`_takePendingLocalEcho`), erasing the flushed prefix with backspaces counted in code points. ⚠️ Drafts are per-session and in memory only (`_composerDrafts`), never persisted (prompts carry secrets). ⚠️ Delivery is a hand-built bracketed-paste frame via `_sendInputAsync` WITHOUT `useMux`, then a separate delayed Enter WITH it; never `terminal.paste()`, and the frame must never take the mux fallback (it strips newlines). ⚠️ `_composerMaxLength` must stay derived from `MAX_INPUT_LENGTH` minus the markers, or an oversized frame wedges the durable queue. ⚠️ The composer overlay needs its own gutter restatement after the fold rules. Test: `test/mobile-prompt-composer.test.ts`. → [architecture-invariants#mobile-prompt-composer](docs/architecture-invariants.md#mobile-prompt-composer)
|
||||
|
||||
**PTY and browser terminal geometry** (#464): a browser terminal whose width differs from the PTY's garbles Claude's redraws, so `syncTerminalGeometry()` (terminal-ui.js) is the ONE function that may resize the main terminal (never a bare `fitAddon.fit()`), a font change is a geometry change, and the fit is withheld wherever the SIGWINCH is. ⚠️ Resize is answered with `Session.ptyGeometry`, and a client adopts its COLUMNS only (never rows); `ptyGeometry` is null without a live pane. → [architecture-invariants#pty-and-browser-terminal-geometry](docs/architecture-invariants.md#pty-and-browser-terminal-geometry)
|
||||
|
||||
**Terminal resilience**: a replay clear is the queued in-stream `\x1bc` in `_resetTerminalForReplay()`, never `reset()`/`clear()` (queued bytes fuse into the snapshot); the renderer watchdog `_kickRenderer()` reads xterm privates, pinned by `test/xterm-private-api.test.ts` against the resolved lockfile version; every terminal capture fetch has a deadline that covers the BODY (`_fetchTerminalCapture`). → [architecture-invariants#terminal-resilience-replay-clears-renderer-liveness-fetch-deadlines](docs/architecture-invariants.md#terminal-resilience-replay-clears-renderer-liveness-fetch-deadlines)
|
||||
|
||||
**WebSocket output-gap reconcile** (`_wsOutputGapSession`, app.js): output frames carry no sequence number, so an unintentional WS close while SSE stays up marks the session and the next open reconciles. ⚠️ The marker is cleared only once a repaint actually happened (`_markTerminalBufferReconciled()`, never in a `finally`). → [architecture-invariants#websocket-output-gap-reconcile](docs/architecture-invariants.md#websocket-output-gap-reconcile)
|
||||
|
||||
**Service worker precache** (`sw.js` + `scripts/build.mjs`): `BUILD_ID` and `HASHED_ASSETS` are build-generated and the build THROWS unless each declaration appears exactly once; `caches.match` must pass `ignoreSearch: true` because `cacheBustAssets` appends `?v=` to hashed names. → [architecture-invariants#service-worker-precache-and-cache-key](docs/architecture-invariants.md#service-worker-precache-and-cache-key)
|
||||
|
||||
**Dismissing the on-screen keyboard** (`terminal-ui.js`): two gestures blur the terminal's hidden textarea. (1) `_installMobileKeyboardDismiss()`, a document `touchend` that must never fire inside `#terminalContainer` or on a control (`MOBILE_KEYBOARD_DISMISS_EXEMPT_SELECTOR`, via `closest()`). (2) In `_handleMobileTerminalTap`, a second tap on inert `content` blurs; the prompt row keeps focus-then-position. ⚠️ A scroll also ends in `touchend`: both classifiers must share one threshold (`TAP_THRESHOLD` reads `MOBILE_KEYBOARD_DISMISS_TAP_SLOP`), and multi-touch is never a tap. ⚠️ CI cannot see the only test for (1): run `npm run test:mobile -- test/mobile/keyboard.test.ts` by hand and diff the FAIL list against master. → [architecture-invariants#dismissing-the-on-screen-keyboard](docs/architecture-invariants.md#dismissing-the-on-screen-keyboard)
|
||||
|
||||
**Phone toolbar: Enter replaces Shell** (post-1.8.0): inside `@media (max-width: 599px)` `btn-shell` is `display:none` and `btn-enter` takes its slot (`order: 4`); starting a shell moved into the Run dropdown (`Terminal / Shell` → `setRunMode('shell')` → `run()` → `runShell()`, button label "Run SH"). `runMode` is `z.string().max(20)` server-side, so new modes need no schema change. Desktop and tablet keep the green Run Shell button unchanged.
|
||||
|
||||
⚠️ **`sendEnterKey()` MUST go through `terminal._core.coreService.triggerDataEvent('\r', true)`** — not `sendInput()`, and never a raw POST to `/api/sessions/:id/input`. `localEchoEnabled` defaults to `MobileDetection.isTouchDevice()`, so on every phone the characters you type are buffered in the `LocalEchoOverlay` and have **never reached the PTY**; the `onData` Enter branch in terminal-ui.js is what flushes `pendingText` first and only then sends `\r` (after an 80ms delay so text lands first). Sending a bare `\r` submits an empty line and strands the typed text on screen, so the button looks dead. Replaying the keypress reuses the overlay flush, the flushed-offset cleanup and the ordering instead of reimplementing them. `KeyboardAccessory.sendKey()` is for escape sequences (arrows/Esc) and is the WRONG template to copy for input.
|
||||
|
||||
⚠️ **Skin overrides outrank plain class rules.** `styles.css` nests its skin block inside `html:not([data-skin="og"]) { … }`, so a bare `.btn-toolbar` rule in there resolves to specificity **(0,2,1)** and beats a `.btn-toolbar.btn-x` rule **(0,2,0)** in `mobile.css` regardless of load order. Toolbar-button colors set from mobile.css therefore need `!important` — that is why mobile.css leans on it so heavily. Symptom: only your `!important` properties land and everything else silently renders in generic toolbar grey.
|
||||
|
||||
**Connection-loss UI** (`computeConnectionLossUi()` in constants.js, writer `_updateConnectionLossUi()` in app.js): the service worker serves the cached app shell, so an unreachable server (phone off the tailnet, VPN down, server stopped) used to render a normal-looking empty dashboard whose only tell was the 8px header dot, which reads as "no sessions", not "no connection". Two surfaces now: a full-screen **overlay** while no server state has loaded this page load (nothing behind it is worth preserving), and a non-blocking **banner** once it has (the terminal scrollback stays readable). ⚠️ A **2.5s grace** is load-bearing: a COM deploy restarts the server and SSE is back in ~200ms, and a banner on every deploy trains the user to ignore it. `navigator.onLine === false` skips the grace, since that is never a blip. Retry re-arms SSE **and** the terminal WS (`planWsReconnect` can 'give-up', and the SSE backoff caps at 30s).
|
||||
|
||||
**SSE staleness watchdog** (`computeSseStale()` in constants.js, `_checkSseStale()` + a 5s interval in app.js): an `EventSource` can stop delivering without erroring, so the client forces a reconnect when nothing arrives. ⚠️ The server keepalive must stay the named `sse:heartbeat` event (`cleanupDeadClients()`, sse-stream-manager.ts), never an SSE comment, which `EventSource` cannot observe; its no-op client listener must stay registered. ⚠️ Judge staleness only while `connected` and online (the loop breaker). ⚠️ The liveness stamp lives inside `addListener`. ⚠️ Clear the interval only at the top of `connectSSE()`, or intervals stack. → [architecture-invariants#sse-staleness-watchdog](docs/architecture-invariants.md#sse-staleness-watchdog)
|
||||
|
||||
**Z-index layers** (keep new overlays consistent with this stack): local echo overlay (7), terminal touch-selection bar (900, below floating agent windows), subagent windows + split picker menu (1000), plan agents (1100), mobile/tablet fixed header (1200), modals on ≤768px (1300, must beat the fixed header), log viewers (2000), connection-loss overlay (2500), image popups (3000), response viewer (5000, backdrop 4999), file-preview overlay (5100, must outrank the response viewer that launches it), toasts/path picker (10000+), custom-model center-status banner (10001; its `[hidden]` must re-assert `display: none` or `dismiss()` leaves an invisible click-blocker), custom-model swap-confirm/context-warning modals (10010). → [architecture-invariants#z-index-layers](docs/architecture-invariants.md#z-index-layers)
|
||||
|
||||
**Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min).
|
||||
|
||||
**Keyboard shortcuts**: Escape (close), Ctrl+? (shortcut overlay), Ctrl/Cmd/Alt+K (session palette), Ctrl+W (kill), Ctrl+Tab (next), Alt+[/] (prev/next tab), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter or Ctrl+Enter (newline), Ctrl+C (copy selection, else interrupt) / Ctrl+Shift+C (copy, never interrupts), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font), Shift+Wheel (local scrollback when mouse passthrough is active), Shift+drag (start a selection in a stripped-DECSET pane, where xterm's own Shift branch is unreachable and a Shift+drag used to select nothing; `_installShiftDragSelection`), right-click (copy the selection, the mintty/PuTTY convention, since xterm paints into a canvas and the native menu has no Copy for it; with nothing selected the native menu is left alone). Rebindable via the registry.
|
||||
|
||||
### Security
|
||||
|
||||
**Full model: [`docs/security-architecture.md`](docs/security-architecture.md)** (network binding, auth pipeline, the tunnel caveat, file-serving hardening, supply-chain, instance isolation, recommended setups). **Layer-by-layer detail with the history behind each: [architecture-invariants#security-layers](docs/architecture-invariants.md#security-layers).**
|
||||
|
||||
| Layer | The rule |
|
||||
| ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| **Auth** | Optional HTTP Basic via `CODEMAN_USERNAME` (default `admin`) / `CODEMAN_PASSWORD`. Active only when `CODEMAN_PASSWORD` is set (`middleware/auth.ts`) |
|
||||
| **Network bind** | Defaults to loopback. Non-loopback without a password starts but warns loudly. Classifier: `network-auth-policy.ts` |
|
||||
| **Host guard** | Always-on Host-header allowlist blocking DNS rebinding. ⚠️ **Custom reverse-proxy domains are rejected** unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix` |
|
||||
| **CSRF / Origin** | Always-on cross-site Origin guard on state-changing requests. **A missing Origin is allowed** so curl/CLI and hooks keep working. ⚠️ The body parser keeps `text/plain` RAW; auto-JSON-parsing it enabled simple-request CSRF |
|
||||
| **QR Auth** | Single-use 6-char tokens (60s TTL) for tunnel login. See `docs/qr-auth-plan.md` |
|
||||
| **Sessions** | 24h cookie (`codeman_session`), auto-extend, device context audit |
|
||||
| **Rate limit** | 10 failed auth/IP → 429 (15min decay). QR and hook-secret have separate buckets, so neither can lock out login |
|
||||
| **Hook bypass** | `/api/hook-event` + `/api/status-telemetry` skip Basic auth (localhost-only, schema-validated), but when auth is active the loopback bypass requires `X-Codeman-Hook-Secret` **unconditionally** (Codeman cannot detect a user's own loopback reverse proxy) |
|
||||
| **Lost-frame page** | The THIRD unauthenticated 200, beside the two hook routes, and the only one decided by request headers alone: a `GET`/`HEAD` carrying `Sec-Fetch-Dest: iframe\|frame`, `Accept: text/html` and mode `navigate` (or none), for a path that is NOT a registered route (never `/api/`, `/ws/`, `/q/`), is answered BEFORE the credential checks with the static web-tab recovery page (`lostWebviewFramePage`: no reflected input, `default-src 'none'` plus its own script hash, `no-store`). `/` is the one registered route also admitted, only when the request carries neither `codeman_session` nor `Authorization` (nothing in Codeman frames its own root; a sandboxed frame has neither), since the landing page masks to exactly `/` and its reload otherwise rendered Codeman inside the web tab. ⚠️ A non-browser client can set those headers, so an unauthenticated caller can tell a registered route (401) from a non-route (200) and enumerate the route table; accepted, the routes are public in `docs/api-reference.md`. Pinned by `test/webview-auth-exemption.test.ts` + `test/webview-lost-root-frame.test.ts` |
|
||||
| **Tunnel** | Enabling a tunnel **refuses** without `CODEMAN_PASSWORD` unless exposure is acknowledged via `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` or the per-request `acknowledgeUnauthTunnel:true` action field (never persisted) |
|
||||
| **Validation** | Zod schemas, Unicode-aware path allowlist regex, env prefix allowlist (`CLAUDE_CODE_*`/`OPENCODE_*`/`CODEX_*`/`GEMINI_*`/`GOOGLE_*`/`ANTIGRAVITY_*`/`PI_*`/`GROK_*`/`XAI_*`/`DSH_*`/`DEEPSEEK_*`) |
|
||||
| **Headers** | CORS localhost-only, CSP, X-Frame-Options, HSTS if HTTPS |
|
||||
|
||||
**Security-relevant env vars**: `CODEMAN_MUX` (managed session), `CODEMAN_API_URL` (auto-set for hooks), `CODEMAN_ALLOWED_HOSTS` (extra Host/Origin allowlist entries for reverse proxies; bare `.suffix` matches subdomains), `CODEMAN_DOCKER_BRIDGE_HOOKS=1` (opt-in hooks-only listener on the docker bridge gateway).
|
||||
|
||||
### SSE Event Registry
|
||||
|
||||
Event constants live in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). **Both must be kept in sync**; `test/sse-registry-parity.test.ts` pins it. ⚠️ `hook:agent_working` is the one hook event with no Claude Code hook behind it — the DeepSeek status bridge reports it (see External CLI modes). The backend file's `@fileoverview` carries the per-category breakdown, including the two Web tab events.
|
||||
|
||||
### API Routes
|
||||
|
||||
One module per domain in `src/web/routes/` (plus a barrel; `ls src/web/routes/` for the current list). Beyond the `/api` routes: the `/webview/:cap/*` proxy, the `/ws/voice/stream` relay and the terminal WebSocket. Each file has `@fileoverview` with endpoint details.
|
||||
|
||||
**HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`).
|
||||
|
||||
## Adding Features
|
||||
|
||||
- **API endpoint**: Types in `src/types/` domain file, route in `src/web/routes/*-routes.ts`. Return the `ApiResponse` envelope (`{ success: true, data }`; errors via `createErrorResponse()` with proper status code). Validate with Zod schemas in `schemas.ts`.
|
||||
- **SSE event**: Add to `src/web/sse-events.ts` + `SSE_EVENTS` in `constants.js`, emit via `broadcast()`, handle in `app.js` (`addListener(`)
|
||||
- **Session setting**: Add to `SessionState`, include in `session.toState()`, call `persistSessionState()`
|
||||
- **App setting**: decide per-device vs synced first. Per-device keys go in the `displayKeys` set in settings-ui.js and must NOT be added to `SettingsUpdateSchema` (it is `.strict()`). ⚠️ Anything in `PUT /api/settings` that acts on a setting (the `toggleService` watcher calls) must resolve from **`merged`** (persisted + incoming), never from the raw request body: a partial PUT omits keys it doesn't intend to change, and `body.x ?? default` turns every omission into "apply the default" and silently resets live services. Pinned by `test/routes/system-routes-settings-partial-put.test.ts`.
|
||||
- **Hook event**: Add to `HookEventType`, add hook in `hooks-config.ts:generateHooksConfig()`, update `HookEventSchema`
|
||||
- **Mobile feature**: Add to relevant singleton, guard with `MobileDetection.isMobile()`. New header buttons must stay off phones (`test/mobile-header-buttons-policy.test.ts`).
|
||||
- **New test**: Pick unique port (search `const PORT =`). Route tests use `app.inject()` (no port needed) — see `test/routes/_route-test-utils.ts`.
|
||||
|
||||
**Validation**: Zod v4 (different API from v3). Define schemas in `schemas.ts`, use `.parse()`/`.safeParse()`.
|
||||
|
||||
## State Files
|
||||
|
||||
All in `~/.codeman/`: `state.json` (sessions, settings, respawn, orchestrator, cron jobs/runs, owner tab layouts), `mux-sessions.json` (tmux recovery), `settings.json` (user prefs), `push-keys.json` + `push-subscriptions.json`, `session-lifecycle.jsonl` (audit log), `update-status.json` (self-updater progress, polled across the service restart), `docker-env-applied.json` (Compose deployment only: sha256 of the Dockerfile + compose file the running container was built from, written by `Start-Codeman.sh`, read by the self-updater's environment gate), `docker-build-source.json` (Compose deployment only: the checkout's HEAD commit and `package-lock.json` hash the `codeman-node-modules`/`codeman-dist` volumes currently reflect, written by both `Start-Codeman.sh` and a successful in-place self-update, compared to detect and refresh a volume left stale by an externally-triggered rebuild), `linked-cases.json`, `webviews.json` (saved web-tab dashboard URLs), `remote-hosts.json` + `remote-cases.json`, `docker-hosts.json` + `docker-cases.json` + `docker-exports/`, `subagent-window-states.json` + `subagent-parents.json` (subagent window layout, GET/PUT `/api/subagent-window-states`/`-parents`), `hook-secret` (per-instance), `users.json` (multi-user, mode 0600) + `admin-audit.jsonl`, `intents.json` (Read My Mind intent profiles, mode 0600), `certs/` (self-signed TLS for `--https`), `.env` (CODEMAN_USERNAME/PASSWORD fallback for the `codeman attach` CLI), `install.log` (installer step output, written by `install.sh`'s `run_step`) and `tailscale-rename` (the node name before `install.sh` renamed it, so uninstall can offer it back; both installer-route only). Transient: `self-update-runner.sh`. Multi-user case spaces live OUTSIDE the data dir at `~/codeman-users/<username>/cases` (shared across instances like `~/codeman-cases`, override `CODEMAN_USER_SPACES_DIR`).
|
||||
|
||||
**Generated top-level dirs** (all gitignored — don't edit or commit): `dist/` (esbuild output), `out/`, `coverage/`, `test-results/`, `tmp/`, `screenshots-echo-diag/`. The committed gesture bundle (`src/web/public/gesture/gesture-codeman.js`) IS tracked, but its runtime wasm/model assets (`src/web/public/gesture/wasm/`, `*.task`) are fetched and gitignored.
|
||||
|
||||
## Testing
|
||||
|
||||
**`npm test` is the gate and is safe to run bare** — it runs `config/vitest.ci.config.ts`, exactly what CI runs, so local green means CI green.
|
||||
|
||||
```bash
|
||||
npm test # The gate — what CI runs
|
||||
npm test -- test/<specific-file>.test.ts # Single file
|
||||
npm test -- -t "pattern" # By name
|
||||
```
|
||||
|
||||
Three suites are deliberately left out, because they cannot pass on an arbitrary machine. Each has its own runner, and a failure there means "not runnable here", not a regression:
|
||||
|
||||
```bash
|
||||
npm run test:browser # Playwright + chromium, live server; codex-predictive-echo also needs a real codex binary
|
||||
npm run test:mobile # the above plus environment-specific PNG baselines (own config, own pretest vendor step)
|
||||
npm run test:perf # wall-clock benchmarks — need an otherwise idle machine
|
||||
npm run test:all # literally everything; fails ~87 tests on a clean master here, which is why it is not the default
|
||||
```
|
||||
|
||||
⚠️ **`npm test` cannot see those suites**, so a change touching mobile/gesture/terminal-render behaviour needs the matching runner by hand — diff its FAIL list against master rather than reading it as pass/fail. That blind spot is what let two semantically-conflicting PRs merge green (see the on-screen-keyboard note above).
|
||||
|
||||
⚠️ **A file filter must match the runner.** `npm test -- test/mobile/keyboard.test.ts` matches nothing and exits GREEN having run zero tests, because the gate's config excludes that path — an excluded file needs its own runner (`npm run test:mobile -- <file>`, `npm run test:browser -- <file>`, `npm run test:perf -- <file>`). Vitest treats "no files matched a filter" as success, so read the file count, not just the colour.
|
||||
|
||||
Raw `npx vitest` skips the config (and with it `setup.ts`); always use `npm test --` or pass `--config`.
|
||||
|
||||
**Config**: Vitest with `globals: true`, `fileParallelism: false`. Timeout 30s, teardown 60s. `config/vitest.config.ts` is the everything-config behind `test:all`; `config/vitest.ci.config.ts` is the gate and derives its excludes from `config/test-suites.ts`, which is also what `vitest.browser.config.ts` and `vitest.perf.config.ts` derive their includes from — so the exclusions and the runners cannot drift apart. Keep shared options in sync across them.
|
||||
|
||||
**Tmux safety**: under vitest (`VITEST`), `TmuxManager` no-ops ALL shell commands (`IS_TEST_MODE` in `src/tmux-manager.ts`), docker IO is no-op'd likewise, and `Session` spawns an echo PTY (`TEST_PTY_SCRIPT`) instead of attaching tmux. `test/setup.ts` gives each file a temp `HOME`/`USERPROFILE` and strips `CODEMAN_PASSWORD`/`CODEMAN_USERNAME`, `CODEMAN_GESTURE` and `CODEMAN_INSTANCE`/`CODEMAN_DATA_DIR`/`CODEMAN_TMUX_SOCKET` (pinned by `test/test-env-isolation.test.ts`). ⚠️ `CODEMAN_DATA_DIR` overrides the temp HOME, so never drop its strip; strip `CODEMAN_INSTANCE` in the setup file, never in a hook (captured at first import). ⚠️ Delete case trees only via `safeRmHomeTree()`. ⚠️ Raw `npx vitest` without `--config` skips `setup.ts` and its isolation. → [architecture-invariants#test-isolation-tmux-docker-and-home](docs/architecture-invariants.md#test-isolation-tmux-docker-and-home)
|
||||
|
||||
**Ports**: Pick unique ports manually, 3150+. Search `const PORT =` before adding new tests. Never 3000 (the live instance).
|
||||
|
||||
⚠️ **Browser tests can pass vacuously on mobile input paths.** Two traps, both hit on 2026-07-27 while fixing the phone Enter button: **(1)** driving input with `app.sendInput('…')` writes PAST the `LocalEchoOverlay`, so `pendingText` stays empty and any overlay bug is invisible — type with `page.keyboard.type()` instead; **(2)** headless Chromium reports `MobileDetection.isTouchDevice()` **false even with `hasTouch: true`**, so `_localEchoEnabled` is off and the local-echo branch never executes. Force it (`app._localEchoEnabled = true`) or the test proves nothing. Assert on real state (`app._localEchoOverlay.pendingText`, plus `tmux -L codeman capture-pane -p -t <pane>` for what actually reached the PTY), not on HTTP 200.
|
||||
|
||||
**Testing against the live instance**: prod is HTTPS-only on :3000 (`curl -sk https://localhost:3000/...`). ⚠️ `w1`/`w2`/`w3` are the user's REAL sessions — never send input to them. Create your own throwaway session (`POST /api/sessions` then `POST /api/sessions/:id/shell`; creation alone leaves `pid: null` and no pane), test against that, and `DELETE` it by exact id when done.
|
||||
|
||||
**Respawn tests**: Use `MockSession` from `test/mocks/index.ts` (defined in `test/mocks/mock-session.ts`). **Route tests**: `app.inject({ method, url, payload })` in `test/routes/` — no live port needed. **Mobile tests**: Playwright suite in `test/mobile/` (device profiles in `test/mobile/devices.ts`). Browser-testing infra and practices: `docs/browser-testing-guide.md`.
|
||||
|
||||
## Debugging
|
||||
|
||||
```bash
|
||||
tmux -L codeman list-sessions # Codeman's own socket (bare `tmux` shows the default one)
|
||||
curl -sk https://localhost:3000/api/sessions | jq # Check sessions (prod is HTTPS-only; dev on :3000 is plain http)
|
||||
curl -sk https://localhost:3000/api/status | jq # Full app state
|
||||
curl -sk https://localhost:3000/api/subagents | jq # Background agents
|
||||
cat ~/.codeman/state.json | jq # Persisted state
|
||||
```
|
||||
|
||||
Mobile screenshots: `~/.codeman/screenshots/`, accessed via `GET/POST /api/screenshots`.
|
||||
|
||||
## Performance & Limits
|
||||
|
||||
Target: 20 sessions, 50 agent windows at 60fps. Limits live in `src/config/` (terminal 32MB, text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100), most env-overridable.
|
||||
|
||||
Two constraints worth knowing before you touch them: the env-derived PTY buffer trim is **clamped to ≤75% of max**, because a trim ≥ max would disable `BufferAccumulator` trimming entirely and make memory unbounded; and browser xterm scrollback is a **separate hardcoded 50k** (`DEFAULT_SCROLLBACK` in constants.js), deliberately lower than tmux's 100k history because 100k per tab is a mobile-memory hazard. tmux <3.7 allocates `history-limit` at pane creation, while tmux 3.7+ can resize live panes (lowering the value can discard retained lines); already-evicted lines never return. The settings keys `terminalScrollbackLines`/`terminalBufferMaxBytes`/`terminalBufferTrimBytes` are schema-validated but **inert**; only `tmuxHistoryLimit` is wired. → [architecture-invariants#buffers-uploads-and-terminal-history](docs/architecture-invariants.md#buffers-uploads-and-terminal-history), `docs/terminal-anti-flicker.md`
|
||||
|
||||
**Memory leaks (24+ hour sessions)**: use `CleanupManager`, clear Maps in `stop()`, guard async with `if (this.cleanup.isStopped) return`. Frontend: store handler refs, clean in `close*()`. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Verify: `npm test -- test/memory-leak-prevention.test.ts`.
|
||||
|
||||
## Scripts & Tunnel
|
||||
|
||||
**`install.sh`** (repo root) is the public `curl | bash` installer: it installs Node/tmux/git/build tools, clones to `~/.codeman/app`, builds, and offers a systemd/launchd service, asking every question BEFORE the unattended build (log in `~/.codeman/install.log`). Network access is Tailscale / LAN / local-only; re-runs preserve the existing binding AND password (`read_existing_binding()`). ⚠️ Compose the hand-start env only in `start_command_hint`/`export_bind_env`, and call `stop_background_helpers` before `exec`. ⚠️ Tailscale rename is opt-in (default NO, never under `--yes`/non-interactive); NEVER `tailscale serve reset`, touch a mapping it did not create, run `tailscale funnel` or advertise a Service (`test/install-sh-invariants.test.ts`). ⚠️ Stay **bash 3.2** clean (no `declare -A`, `mapfile`, namerefs, `${x,,}`, here-strings, empty-array expansion under `set -u`). ⚠️ Execute only commands from the generated CLI block (`CLI_INSTALL_CMD_TRUSTED`, `npm run generate:cli-catalog`). → [architecture-invariants#installsh-the-public-installer](docs/architecture-invariants.md#installsh-the-public-installer)
|
||||
|
||||
Other key scripts: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh [quick|named] start|stop|status|url` (quick = random trycloudflare URL, default; `named setup|enable` = fixed-hostname tunnel via `scripts/codeman-tunnel-named.service`; bare `start|stop|url` still means quick), `scripts/run-beta.sh` (isolated beta instance), `scripts/build-agent-image.mjs` (docker base image), `scripts/self-update.sh` (detached updater). Production services: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.
|
||||
@@ -1,21 +0,0 @@
|
||||
MIT License
|
||||
|
||||
Copyright (c) 2024-2026 Codeman Contributors
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||
of this software and associated documentation files (the "Software"), to deal
|
||||
in the Software without restriction, including without limitation the rights
|
||||
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||
copies of the Software, and to permit persons to whom the Software is
|
||||
furnished to do so, subject to the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be included in all
|
||||
copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||
SOFTWARE.
|
||||
@@ -1,281 +0,0 @@
|
||||
[
|
||||
{
|
||||
"id": "claude",
|
||||
"label": "Claude Code",
|
||||
"shortBadge": "CC",
|
||||
"enabled": true,
|
||||
"order": 0,
|
||||
"kind": "agent",
|
||||
"discovery": {
|
||||
"binaries": [
|
||||
"claude"
|
||||
],
|
||||
"searchDirs": [
|
||||
"~/.local/bin",
|
||||
"~/.claude/local",
|
||||
"/usr/local/bin",
|
||||
"~/.npm-global/bin",
|
||||
"~/bin"
|
||||
],
|
||||
"install": {
|
||||
"command": {
|
||||
"linux": "curl -fsSL https://claude.ai/install.sh | bash",
|
||||
"darwin": "curl -fsSL https://claude.ai/install.sh | bash",
|
||||
"wsl": "curl -fsSL https://claude.ai/install.sh | bash"
|
||||
},
|
||||
"npmPackage": "@anthropic-ai/claude-code",
|
||||
"docsUrl": "https://docs.claude.com/claude-code"
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "shell",
|
||||
"label": "Shell",
|
||||
"shortBadge": "SH",
|
||||
"enabled": true,
|
||||
"order": 1,
|
||||
"kind": "shell",
|
||||
"discovery": {
|
||||
"binaries": [],
|
||||
"searchDirs": [],
|
||||
"install": {
|
||||
"command": {}
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "opencode",
|
||||
"label": "OpenCode",
|
||||
"shortBadge": "OC",
|
||||
"enabled": true,
|
||||
"order": 10,
|
||||
"kind": "agent",
|
||||
"discovery": {
|
||||
"binaries": [
|
||||
"opencode"
|
||||
],
|
||||
"searchDirs": [
|
||||
"~/.opencode/bin",
|
||||
"~/.local/bin",
|
||||
"/usr/local/bin",
|
||||
"~/go/bin",
|
||||
"~/.bun/bin",
|
||||
"~/.npm-global/bin",
|
||||
"~/bin"
|
||||
],
|
||||
"install": {
|
||||
"command": {
|
||||
"linux": "curl -fsSL https://opencode.ai/install | bash",
|
||||
"darwin": "curl -fsSL https://opencode.ai/install | bash"
|
||||
},
|
||||
"npmPackage": "opencode-ai",
|
||||
"docsUrl": "https://opencode.ai/docs"
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "codex",
|
||||
"label": "Codex",
|
||||
"shortBadge": "CX",
|
||||
"enabled": true,
|
||||
"order": 20,
|
||||
"kind": "agent",
|
||||
"discovery": {
|
||||
"binaries": [
|
||||
"codex"
|
||||
],
|
||||
"searchDirs": [
|
||||
"~/.codex/bin",
|
||||
"~/.local/bin",
|
||||
"/usr/local/bin",
|
||||
"~/.bun/bin",
|
||||
"~/.npm-global/bin",
|
||||
"~/bin"
|
||||
],
|
||||
"install": {
|
||||
"command": {
|
||||
"linux": "npm install -g @openai/codex",
|
||||
"darwin": "npm install -g @openai/codex"
|
||||
},
|
||||
"npmPackage": "@openai/codex",
|
||||
"docsUrl": "https://developers.openai.com/codex/cli"
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "gemini",
|
||||
"label": "Gemini",
|
||||
"shortBadge": "GM",
|
||||
"enabled": true,
|
||||
"order": 30,
|
||||
"kind": "agent",
|
||||
"discovery": {
|
||||
"binaries": [
|
||||
"gemini"
|
||||
],
|
||||
"searchDirs": [
|
||||
"~/.gemini/bin",
|
||||
"~/.local/bin",
|
||||
"/usr/local/bin",
|
||||
"~/.bun/bin",
|
||||
"~/.npm-global/bin",
|
||||
"~/bin"
|
||||
],
|
||||
"install": {
|
||||
"command": {
|
||||
"linux": "npm install -g @google/gemini-cli",
|
||||
"darwin": "npm install -g @google/gemini-cli"
|
||||
},
|
||||
"npmPackage": "@google/gemini-cli",
|
||||
"docsUrl": "https://github.com/google-gemini/gemini-cli"
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "antigravity",
|
||||
"label": "Antigravity",
|
||||
"shortBadge": "AG",
|
||||
"enabled": true,
|
||||
"order": 40,
|
||||
"kind": "agent",
|
||||
"discovery": {
|
||||
"binaries": [
|
||||
"agy"
|
||||
],
|
||||
"searchDirs": [
|
||||
"~/.local/bin",
|
||||
"~/.antigravity/bin",
|
||||
"/usr/local/bin",
|
||||
"~/bin"
|
||||
],
|
||||
"install": {
|
||||
"command": {
|
||||
"linux": "curl -fsSL https://antigravity.google/cli/install.sh | bash",
|
||||
"darwin": "curl -fsSL https://antigravity.google/cli/install.sh | bash"
|
||||
},
|
||||
"docsUrl": "https://antigravity.google/cli"
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "pi",
|
||||
"label": "Pi",
|
||||
"shortBadge": "PI",
|
||||
"enabled": true,
|
||||
"order": 50,
|
||||
"kind": "agent",
|
||||
"discovery": {
|
||||
"binaries": [
|
||||
"pi"
|
||||
],
|
||||
"searchDirs": [
|
||||
"~/.local/bin",
|
||||
"/usr/local/bin",
|
||||
"~/.bun/bin",
|
||||
"~/.npm-global/bin",
|
||||
"~/bin"
|
||||
],
|
||||
"install": {
|
||||
"command": {
|
||||
"linux": "npm install -g --ignore-scripts @earendil-works/pi-coding-agent",
|
||||
"darwin": "npm install -g --ignore-scripts @earendil-works/pi-coding-agent"
|
||||
},
|
||||
"npmPackage": "@earendil-works/pi-coding-agent",
|
||||
"docsUrl": "https://pi.dev",
|
||||
"agentImageLayer": {
|
||||
"kind": "dedicated",
|
||||
"reason": "installed with --ignore-scripts in its own layer, so the flag cannot leak to the shared block"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "grok",
|
||||
"label": "Grok",
|
||||
"shortBadge": "GK",
|
||||
"enabled": true,
|
||||
"order": 70,
|
||||
"kind": "agent",
|
||||
"discovery": {
|
||||
"binaries": [
|
||||
"grok"
|
||||
],
|
||||
"searchDirs": [
|
||||
"~/.grok/bin",
|
||||
"~/.local/bin",
|
||||
"/usr/local/bin",
|
||||
"~/bin"
|
||||
],
|
||||
"install": {
|
||||
"command": {
|
||||
"linux": "curl -fsSL https://x.ai/cli/install.sh | bash",
|
||||
"darwin": "curl -fsSL https://x.ai/cli/install.sh | bash"
|
||||
},
|
||||
"docsUrl": "https://github.com/xai-org/grok-build"
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "deepseek",
|
||||
"label": "DeepSeek",
|
||||
"shortBadge": "DS",
|
||||
"enabled": true,
|
||||
"order": 80,
|
||||
"kind": "agent",
|
||||
"discovery": {
|
||||
"binaries": [
|
||||
"dsh"
|
||||
],
|
||||
"searchDirs": [
|
||||
"~/.local/bin",
|
||||
"/usr/local/bin",
|
||||
"~/.npm-global/bin",
|
||||
"~/bin"
|
||||
],
|
||||
"identity": {
|
||||
"arg": "--help",
|
||||
"regex": "DeepSeek\\s+Harness"
|
||||
},
|
||||
"install": {
|
||||
"command": {
|
||||
"linux": "npm install -g @deepseek-ai/dsh",
|
||||
"darwin": "npm install -g @deepseek-ai/dsh"
|
||||
},
|
||||
"npmPackage": "@deepseek-ai/dsh",
|
||||
"docsUrl": "https://github.com/deepseek-ai/deepseek-harness",
|
||||
"agentImageLayer": {
|
||||
"kind": "dedicated",
|
||||
"reason": "needs pnpm alongside it (dsh plugin, issue #352) and a dsh-tui profile install"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "omp",
|
||||
"label": "OMP",
|
||||
"shortBadge": "OM",
|
||||
"enabled": true,
|
||||
"order": 90,
|
||||
"kind": "agent",
|
||||
"discovery": {
|
||||
"binaries": [
|
||||
"omp"
|
||||
],
|
||||
"searchDirs": [
|
||||
"~/.local/bin",
|
||||
"~/.omp/bin",
|
||||
"/usr/local/bin",
|
||||
"~/.bun/bin",
|
||||
"~/.npm-global/bin",
|
||||
"~/bin"
|
||||
],
|
||||
"install": {
|
||||
"command": {
|
||||
"linux": "curl -fsSL https://omp.sh/install | sh",
|
||||
"darwin": "brew install can1357/tap/omp"
|
||||
},
|
||||
"docsUrl": "https://omp.sh"
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
@@ -1,28 +0,0 @@
|
||||
// @ts-check
|
||||
import eslint from '@eslint/js';
|
||||
import tseslint from 'typescript-eslint';
|
||||
|
||||
export default tseslint.config(
|
||||
eslint.configs.recommended,
|
||||
tseslint.configs.recommended,
|
||||
{
|
||||
rules: {
|
||||
'no-console': 'off',
|
||||
'no-debugger': 'error',
|
||||
// Relax some rules that conflict with existing patterns
|
||||
'@typescript-eslint/no-explicit-any': 'warn',
|
||||
'@typescript-eslint/no-unused-vars': 'off', // TypeScript compiler already handles this
|
||||
},
|
||||
},
|
||||
{
|
||||
ignores: [
|
||||
'dist/**',
|
||||
'node_modules/**',
|
||||
'coverage/**',
|
||||
'src/web/public/vendor/**',
|
||||
'src/web/public/app.js',
|
||||
'scripts/**/*.mjs',
|
||||
'scripts/remotion/**',
|
||||
],
|
||||
}
|
||||
);
|
||||
@@ -1,16 +0,0 @@
|
||||
{
|
||||
"$schema": "https://unpkg.com/knip@5/schema.json",
|
||||
"entry": [
|
||||
"scripts/*.mjs",
|
||||
"scripts/*.js",
|
||||
"scripts/watch-subagents.ts",
|
||||
"scripts/remotion/Root.tsx",
|
||||
"scripts/remotion/index.ts",
|
||||
"test/**/*.test.ts",
|
||||
"test/mobile/vitest.config.ts",
|
||||
"test/**/*.mjs"
|
||||
],
|
||||
"project": ["src/**/*.{ts,tsx}", "scripts/**/*.{ts,tsx,mjs,js}", "test/**/*.{ts,mjs}"],
|
||||
"ignoreExportsUsedInFile": true,
|
||||
"ignoreDependencies": ["@remotion/cli", "@remotion/transitions", "esbuild", "agent-browser"]
|
||||
}
|
||||
@@ -1,54 +0,0 @@
|
||||
/**
|
||||
* The test suites that `npm test` deliberately does NOT run, in one place.
|
||||
*
|
||||
* Why this file exists: the exclusion list used to live only in
|
||||
* config/vitest.ci.config.ts, as literals. Anything excluded there was
|
||||
* therefore reachable only by running the everything-config by hand and reading
|
||||
* past its failures — and a newly excluded file was reachable by nothing at
|
||||
* all, silently, because nothing pointed at it. Both configs now derive their
|
||||
* globs from the arrays below, so adding a suite here puts it in exactly one
|
||||
* runner and takes it out of exactly one gate.
|
||||
*
|
||||
* Adding a new test that cannot run in CI: put its glob in the array that
|
||||
* describes WHY it cannot, not in whichever one is shortest.
|
||||
*/
|
||||
|
||||
/**
|
||||
* Playwright-driven: needs chromium and, in most cases, a live Codeman server
|
||||
* on a real port. Deterministic where the environment provides both, which is
|
||||
* why these are a runnable suite (`npm run test:browser`) rather than skipped.
|
||||
*/
|
||||
export const BROWSER_TEST_GLOBS = [
|
||||
'test/tab-rail-resize.browser.test.ts',
|
||||
'test/session-sidebar-ux.browser.test.ts',
|
||||
'test/session-options-responsive.browser.test.ts',
|
||||
'test/inline-rename.test.ts',
|
||||
'test/opencode-resize.test.ts',
|
||||
'test/webgl-fallback.test.ts',
|
||||
'test/terminal-copy-shortcut.test.ts',
|
||||
'test/terminal-keycode229-recovery.browser.test.ts',
|
||||
'test/capture-load-window.browser.test.ts',
|
||||
'test/capture-geometry-retry.browser.test.ts',
|
||||
'test/codex-predictive-echo.test.ts', // also needs a real codex binary
|
||||
'test/split-pane-terminal.browser.test.ts',
|
||||
'test/split-pane-orchestration.browser.test.ts',
|
||||
'test/split-pane-auto-collapse.browser.test.ts',
|
||||
];
|
||||
|
||||
/**
|
||||
* Wall-clock benchmarks. They assert on durations, so a loaded shared runner
|
||||
* fails them for reasons that have nothing to do with the diff under test.
|
||||
*/
|
||||
export const PERF_TEST_GLOBS = ['test/perf-*.test.ts'];
|
||||
|
||||
/**
|
||||
* Browser + visual regression: chromium AND environment-specific PNG baselines
|
||||
* that are generated per machine. Has its own config
|
||||
* (test/mobile/vitest.config.ts) because it needs serial execution, a longer
|
||||
* timeout and the `pretest:mobile` vendor step — run it with
|
||||
* `npm run test:mobile`, not through the configs here.
|
||||
*/
|
||||
export const MOBILE_TEST_GLOBS = ['test/mobile/**'];
|
||||
|
||||
/** Everything `npm test` skips. */
|
||||
export const NON_CI_TEST_GLOBS = [...MOBILE_TEST_GLOBS, ...PERF_TEST_GLOBS, ...BROWSER_TEST_GLOBS];
|
||||
@@ -1,11 +0,0 @@
|
||||
{
|
||||
"extends": "../tsconfig.json",
|
||||
"compilerOptions": {
|
||||
"rootDir": "..",
|
||||
"noEmit": true,
|
||||
"declaration": false,
|
||||
"declarationMap": false,
|
||||
"sourceMap": false
|
||||
},
|
||||
"include": ["../scripts/test-local-llm-harnesses.ts"]
|
||||
}
|
||||
@@ -1,34 +0,0 @@
|
||||
import { resolve } from 'node:path';
|
||||
import { defineConfig } from 'vitest/config';
|
||||
import { BROWSER_TEST_GLOBS } from './test-suites';
|
||||
|
||||
const root = resolve(import.meta.dirname, '..');
|
||||
|
||||
/**
|
||||
* The Playwright-driven suite `npm test` skips — `npm run test:browser`.
|
||||
*
|
||||
* Needs chromium and, for most of these, a live Codeman server on a real port;
|
||||
* codex-predictive-echo also needs a real codex binary. Expect failures where
|
||||
* the machine cannot provide those, and read them as "not runnable here", not
|
||||
* as a regression.
|
||||
*
|
||||
* The mobile suite is NOT here: it needs per-machine PNG baselines, serial
|
||||
* execution and the `pretest:mobile` vendor step, so it keeps its own config
|
||||
* (test/mobile/vitest.config.ts) behind `npm run test:mobile`.
|
||||
*
|
||||
* fileParallelism stays off for the same reason as every other config in this
|
||||
* directory: these bind real ports and drive real tmux sessions, and two files
|
||||
* doing that at once fail each other rather than the code.
|
||||
*/
|
||||
export default defineConfig({
|
||||
test: {
|
||||
root,
|
||||
globals: true,
|
||||
environment: 'node',
|
||||
include: BROWSER_TEST_GLOBS,
|
||||
setupFiles: ['./test/setup.ts'],
|
||||
fileParallelism: false,
|
||||
testTimeout: 60000,
|
||||
teardownTimeout: 60000,
|
||||
},
|
||||
});
|
||||
@@ -1,30 +0,0 @@
|
||||
import { resolve } from 'node:path';
|
||||
import { defineConfig, configDefaults } from 'vitest/config';
|
||||
import { NON_CI_TEST_GLOBS } from './test-suites';
|
||||
|
||||
const root = resolve(import.meta.dirname, '..');
|
||||
|
||||
/**
|
||||
* The default gate — what `npm test` and CI both run.
|
||||
*
|
||||
* Same as vitest.config.ts but EXCLUDES the suites that cannot pass on an
|
||||
* arbitrary machine: browser-driven (Playwright + chromium), visual-regression
|
||||
* (per-machine PNG baselines) and wall-clock perf. Those are not unmaintained;
|
||||
* they have their own runners (`test:browser`, `test:mobile`, `test:perf`).
|
||||
* See config/test-suites.ts for the list and the reason behind each entry.
|
||||
*
|
||||
* Keep the rest in sync with config/vitest.config.ts.
|
||||
*/
|
||||
export default defineConfig({
|
||||
test: {
|
||||
root,
|
||||
globals: true,
|
||||
environment: 'node',
|
||||
include: ['test/**/*.test.ts'],
|
||||
exclude: [...configDefaults.exclude, ...NON_CI_TEST_GLOBS],
|
||||
setupFiles: ['./test/setup.ts'],
|
||||
fileParallelism: false,
|
||||
testTimeout: 30000,
|
||||
teardownTimeout: 60000,
|
||||
},
|
||||
});
|
||||
@@ -1,37 +0,0 @@
|
||||
import { resolve } from 'node:path';
|
||||
import { defineConfig } from 'vitest/config';
|
||||
|
||||
const root = resolve(import.meta.dirname, '..');
|
||||
|
||||
/**
|
||||
* EVERY test in the repo, including the ones that cannot pass on an arbitrary
|
||||
* machine — `npm run test:all`. Reach for it when you want the complete picture
|
||||
* and are prepared to read past environmental failures.
|
||||
*
|
||||
* This is NOT what `npm test` runs. On a machine without chromium, a free port
|
||||
* or per-machine PNG baselines this config fails ~87 tests on a clean master,
|
||||
* which makes it useless as a pass/fail signal: the default gate is
|
||||
* config/vitest.ci.config.ts, and the suites it leaves out each have their own
|
||||
* runner (`test:browser`, `test:perf`, `test:mobile`). See config/test-suites.ts.
|
||||
*/
|
||||
export default defineConfig({
|
||||
test: {
|
||||
root,
|
||||
globals: true,
|
||||
environment: 'node',
|
||||
include: ['test/**/*.test.ts'],
|
||||
setupFiles: ['./test/setup.ts'],
|
||||
// Run test files sequentially to respect mux session limits
|
||||
// Individual tests within files still run in parallel where safe
|
||||
fileParallelism: false,
|
||||
coverage: {
|
||||
provider: 'v8',
|
||||
reporter: ['text', 'json', 'html'],
|
||||
include: ['src/**/*.ts'],
|
||||
exclude: ['src/index.ts', 'src/cli.ts'],
|
||||
},
|
||||
testTimeout: 30000, // 30 seconds for integration tests
|
||||
// Ensure cleanup runs even on test failures
|
||||
teardownTimeout: 60000,
|
||||
},
|
||||
});
|
||||
@@ -1,25 +0,0 @@
|
||||
import { resolve } from 'node:path';
|
||||
import { defineConfig } from 'vitest/config';
|
||||
import { PERF_TEST_GLOBS } from './test-suites';
|
||||
|
||||
const root = resolve(import.meta.dirname, '..');
|
||||
|
||||
/**
|
||||
* The wall-clock benchmarks `npm test` skips — `npm run test:perf`.
|
||||
*
|
||||
* These assert on durations, so run them on an otherwise idle machine: a loaded
|
||||
* runner fails them for reasons that have nothing to do with the diff under
|
||||
* test, which is exactly why they are not part of the default gate.
|
||||
*/
|
||||
export default defineConfig({
|
||||
test: {
|
||||
root,
|
||||
globals: true,
|
||||
environment: 'node',
|
||||
include: PERF_TEST_GLOBS,
|
||||
setupFiles: ['./test/setup.ts'],
|
||||
fileParallelism: false,
|
||||
testTimeout: 60000,
|
||||
teardownTimeout: 60000,
|
||||
},
|
||||
});
|
||||
@@ -1,90 +0,0 @@
|
||||
# =============================================================================
|
||||
# Codeman Docker Compose environment template
|
||||
# Copy this file to .env and set the values for the Docker host.
|
||||
# =============================================================================
|
||||
|
||||
TZ=Australia/Perth
|
||||
|
||||
# Optional overrides for direct `docker compose` use. The Bash start script
|
||||
# detects these values from CODEMAN_APPDATA_PATH automatically. Compose uses
|
||||
# 1000:1000 when the variables are omitted.
|
||||
# PUID=1000
|
||||
# PGID=1000
|
||||
|
||||
# Name of the account that runs Codeman and all local CLI sessions. Changing
|
||||
# this value rebuilds the image with a matching account.
|
||||
CODEMAN_RUNTIME_USER=codeman
|
||||
|
||||
# Required. Persistent Codeman application data, CLI credentials, and session
|
||||
# state are stored here on the host and mounted at the runtime account's home
|
||||
# directory in the container.
|
||||
CODEMAN_APPDATA_PATH=/mnt/user/appdata/codeman
|
||||
|
||||
# Optional. Absolute host path of this Codeman checkout, mounted at
|
||||
# /opt/codeman so App Settings -> Updates can update Codeman in place. The Bash
|
||||
# start script detects it from the compose file's own location, so it only needs
|
||||
# setting for direct `docker compose` use or a checkout kept elsewhere. Point it
|
||||
# at a directory that is not a git checkout and in-app updates are unavailable.
|
||||
# CODEMAN_REPO_PATH=/mnt/user/appdata/codeman/app
|
||||
|
||||
# Required for Docker cases. This must be an absolute path on the Docker host.
|
||||
# Codeman and each isolated case use this same path, so it cannot be a
|
||||
# container-only path such as /home/codeman/codeman-cases.
|
||||
CODEMAN_CASES_PATH=/mnt/user/appdata/codeman/codeman-cases
|
||||
|
||||
# Required. Network bind address, host port, and local image tag.
|
||||
CODEMAN_HOST=0.0.0.0
|
||||
CODEMAN_PORT=3000
|
||||
CODEMAN_IMAGE=codeman:local
|
||||
|
||||
# Required for any network-accessible Codeman instance. Use a unique, strong
|
||||
# password. This file is safe to commit; copy it to .env and set the value.
|
||||
CODEMAN_PASSWORD=changeme
|
||||
|
||||
# Required. Username for Codeman HTTP Basic authentication.
|
||||
CODEMAN_USERNAME=admin
|
||||
|
||||
# Optional. Extra Host-header allowlist entries for a reverse-proxied domain
|
||||
# (comma-separated; a bare `.suffix` matches every subdomain). Without it a
|
||||
# proxied request is rejected with `403 Forbidden: host not allowed`. See
|
||||
# README.md, "Reverse-proxy host allowlist".
|
||||
# CODEMAN_ALLOWED_HOSTS=codeman.example.com,.internal.example.com
|
||||
|
||||
# The GitHub CLI (gh) and the Azure CLI (az, with the azure-devops extension)
|
||||
# can be built into the images as git credential helpers, so Codeman can clone
|
||||
# private GitHub and Azure DevOps repositories. Both are OFF by default and are
|
||||
# NOT set here: turn them on in docker-compose.override.yml with the build args
|
||||
# CODEMAN_INSTALL_GH / CODEMAN_INSTALL_AZ and, for the Docker-case agent image,
|
||||
# the environment variables CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ. See
|
||||
# README.md, "Private repositories".
|
||||
|
||||
# Optional: authenticate Gemini CLI without an interactive login.
|
||||
GEMINI_API_KEY=
|
||||
|
||||
# Linux default. On Docker Desktop, use the socket path supported by your
|
||||
# Docker installation when it differs from /var/run/docker.sock.
|
||||
DOCKER_SOCKET=/var/run/docker.sock
|
||||
|
||||
# Optional override for direct `docker compose` use. The Bash start script
|
||||
# detects this from DOCKER_SOCKET automatically. The direct Compose default is
|
||||
# 999, but the correct value depends on the Docker host.
|
||||
# DOCKER_SOCKET_GID=999
|
||||
|
||||
# Set to 1 only when Docker-case hook callbacks are required.
|
||||
CODEMAN_DOCKER_BRIDGE_HOOKS=0
|
||||
|
||||
# Set to 1 when `docker info` reports `SwapLimit=false`. The case memory limit
|
||||
# remains active; Codeman omits --memory-swap and filters the daemon's exact
|
||||
# unsupported-swap warning while preserving all other Docker create errors.
|
||||
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=0
|
||||
|
||||
# Required only when applying the macvlan example in README.md.
|
||||
CODEMAN_MACVLAN_NETWORK=br0.11
|
||||
CODEMAN_IPV4_ADDRESS=10.10.11.236
|
||||
CODEMAN_MAC_ADDRESS=02:10:11:00:00:EC
|
||||
|
||||
# Required only when creating a new managed macvlan network, rather than using
|
||||
# the external-network macvlan example.
|
||||
CODEMAN_MACVLAN_PARENT=br0.11
|
||||
CODEMAN_MACVLAN_SUBNET=10.10.11.0/24
|
||||
CODEMAN_MACVLAN_GATEWAY=10.10.11.1
|
||||
@@ -1,236 +0,0 @@
|
||||
# Codeman Docker deployment
|
||||
|
||||
This folder contains the Compose configuration, server image Dockerfile, and environment template for a locally built Codeman server.
|
||||
|
||||
## Start
|
||||
|
||||
From the repository root, create the runtime environment file and set the required values, especially `CODEMAN_PASSWORD`.
|
||||
|
||||
```sh
|
||||
cp docker/.env.example docker/.env
|
||||
bash docker/Start-Codeman.sh
|
||||
```
|
||||
|
||||
On PowerShell, use the following commands instead. Running Compose from inside `docker/` with no `-f` lets it discover `docker-compose.override.yml` on its own (see [Local customisation](#local-customisation)); naming the file with `-f docker/docker-compose.yaml` from the repository root silently drops the override unless it is named too.
|
||||
|
||||
```powershell
|
||||
Copy-Item docker/.env.example docker/.env
|
||||
Set-Location docker
|
||||
docker compose --env-file .env up --build -d
|
||||
```
|
||||
|
||||
Every required value is defined and explained in `.env.example`. `GEMINI_API_KEY` is intentionally optional and may remain blank.
|
||||
|
||||
The container starts as root so `entrypoint.sh` can correct the ownership of a bind source the Docker daemon created (it creates a missing one as `root:root`), then drops to `PUID:PGID` with `setpriv` before the server starts, so Codeman itself never runs privileged. That drop needs `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the file's `cap_drop: ALL`; a compose file written elsewhere (Unraid's Compose Manager, a hand-written unit) must carry the same additions, and the entrypoint names them when they are missing. A directory owned by neither root nor `PUID:PGID` is never re-owned: it is probed for writability as the runtime account and refused with a clear message if that fails. Setting `user:` in Compose skips the whole step.
|
||||
|
||||
On Linux, `Start-Codeman.sh` stops with an error when required paths are missing. It creates the application-data directory when safe, detects its numeric owner as `PUID:PGID`, and detects `DOCKER_SOCKET_GID` from the configured Docker socket. It rejects a root-owned application-data directory because Codeman and its local CLI sessions must remain unprivileged.
|
||||
|
||||
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `codeman`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
|
||||
|
||||
To retain Docker-case support without root when running Compose directly, set `DOCKER_SOCKET_GID` to the numeric group ID of the host socket. On a standard Linux Docker host, obtain it with `stat -c '%g' /var/run/docker.sock`. The Bash start script detects it automatically.
|
||||
|
||||
## Updating
|
||||
|
||||
Use **App Settings → Updates** in the web UI. The checkout Compose builds from is
|
||||
also mounted at `/opt/codeman`, so an update's `git checkout` and rebuild persist
|
||||
on the host, and the server exiting is what restarts the container onto the new
|
||||
build.
|
||||
|
||||
Releases that change `server.Dockerfile`, `docker-compose.yaml`, or add a key to
|
||||
`.env.example` cannot be applied that way — the updater detects them, names what
|
||||
changed, and asks you to run `Start-Codeman.sh` here on the host instead. Details:
|
||||
[`../docs/docker-self-update.md`](../docs/docker-self-update.md).
|
||||
|
||||
### Major updates
|
||||
|
||||
`Start-Codeman.sh` rebuilds the image on every start, but with the layer cache,
|
||||
and it refreshes the build-artefact volumes selectively: `codeman-dist` when
|
||||
the checkout's HEAD moved, `codeman-node-modules` only when `package-lock.json`
|
||||
changed. That is right for an ordinary `git pull`. It is not enough when a
|
||||
`server.Dockerfile` change bumps the Node base image without touching the
|
||||
lockfile: `node-pty` is compiled from source (there is no Linux prebuild), so
|
||||
the old `codeman-node-modules` volume would keep a build made for the previous
|
||||
Node version. For that case, or whenever you want to be certain of what ships,
|
||||
`docker/Update-Codeman.sh` force-rebuilds the image with no layer cache, stops
|
||||
the stack, removes the `codeman-node-modules` and `codeman-dist` volumes, then
|
||||
hands off to `Start-Codeman.sh` for the usual start:
|
||||
|
||||
```sh
|
||||
bash docker/Update-Codeman.sh
|
||||
```
|
||||
|
||||
Pass `--keep-volumes` to skip clearing them (safe only if you know the
|
||||
rebuilt image's `node_modules`/`dist` did not change). The scripted default
|
||||
is the "Resetting the build artefacts" procedure in
|
||||
[`../docs/docker-self-update.md`](../docs/docker-self-update.md). Only those
|
||||
two volumes are removed, by name within this Compose project; any volume a
|
||||
`docker-compose.override.yml` adds is left alone, and application data and
|
||||
case workspaces are host bind mounts, never touched either way.
|
||||
|
||||
## Private repositories (GitHub and Azure DevOps)
|
||||
|
||||
The images can include the GitHub CLI (`gh`) and the Azure CLI (`az`, with the `azure-devops` extension), wired into the system Git configuration as credential helpers, so Codeman can clone private repositories. Both are **opt-in and off by default**, and are turned on per host in `docker-compose.override.yml`.
|
||||
|
||||
### Turning them on
|
||||
|
||||
Add the build arguments to `docker-compose.override.yml` (see [Local customisation](#local-customisation)), then rebuild with `Start-Codeman.sh`. Set only the one you need:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
codeman:
|
||||
build:
|
||||
args:
|
||||
CODEMAN_INSTALL_GH: '1'
|
||||
CODEMAN_INSTALL_AZ: '1'
|
||||
environment:
|
||||
# The same two switches for the Docker-case agent image Codeman builds.
|
||||
CODEMAN_AGENT_IMAGE_INSTALL_GH: '1'
|
||||
CODEMAN_AGENT_IMAGE_INSTALL_AZ: '1'
|
||||
```
|
||||
|
||||
The `build: args:` pair controls the Codeman server image. The `environment:` pair controls the agent image for [Docker cases](../docs/docker-cases.md), which Codeman builds on the first Docker case; an agent image that already exists is not rebuilt by this, so run `node scripts/build-agent-image.mjs --no-cache` inside the container afterwards. The same variables work in front of that command when building it by hand. Values must be `0` or `1`; anything else stops the build with an error naming the argument.
|
||||
|
||||
They are not `.env` settings: turning a CLI on is a per-host choice, which is what the override file is for, and a new `.env.example` key makes the in-app updater refuse to update every existing installation until its `.env` gains the key.
|
||||
|
||||
The Azure CLI is the large one, about 600 MB of the roughly 670 MB the pair adds. A CLI left off leaves nothing functional behind: no apt repository, no package, no `azure-devops` extension and no credential-helper entry, so git for that host behaves exactly as it does without this feature. With both off the image is functionally unchanged; it still carries the `AZURE_EXTENSION_DIR` variable, an empty extensions directory and one small layer that copies and then removes the helper script.
|
||||
|
||||
### Signing in
|
||||
|
||||
With a CLI on, the system Git configuration routes credentials through it:
|
||||
|
||||
| Host | Credential helper | Sign in with |
|
||||
| ----------------------------------------------------- | ----------------------------------------- | ---------------------------- |
|
||||
| `https://github.com`, `https://gist.github.com` | `gh auth git-credential` | `gh auth login` |
|
||||
| `https://dev.azure.com`, `https://*.visualstudio.com` | `/usr/local/bin/git-credential-azure-cli` | `az login --use-device-code` |
|
||||
|
||||
Codeman itself still collects no Git credentials. Sign the container in once from a **Terminal / Shell** session (Run menu). The session runs as the runtime account, so the sign-in is stored under `CODEMAN_APPDATA_PATH` (`~/.config/gh`, `~/.azure`) and survives rebuilds and container recreation:
|
||||
|
||||
```sh
|
||||
gh auth login # GitHub.com -> HTTPS -> "Login with a web browser" (device code)
|
||||
az login --use-device-code # then: az devops configure --defaults organization=https://dev.azure.com/<org>
|
||||
```
|
||||
|
||||
After that, **Add Case → Clone Repo** accepts private `https://` URLs on those hosts, and `git clone` works from any session. Until a CLI is signed in its helper prints nothing, so a private clone fails immediately with the usual authentication error rather than waiting on a prompt.
|
||||
|
||||
**Multi-user mode:** every Codeman user's git runs as the same server account, so these sign-ins would otherwise be shared. Clone Repo therefore runs a **non-admin**'s clone and preflight with every git credential helper cleared (`git -c credential.helper=`): a non-admin can clone public repositories and anything their own SSH setup allows, but not a private https repository through the admin's `gh`/`az` sign-in. Admins, and single-user mode, keep the helpers. A non-admin's own agent sessions still run as that same account, and with the agent-image `gh`/`az` switches on, a non-admin's Docker case with credential seeding on also receives the server account's `gh`/`az` sign-in, the same as the Claude and Codex credentials; see `docs/security-architecture.md`, multi-user mode.
|
||||
|
||||
Azure DevOps is authenticated with an Entra ID access token that the helper requests from `az` for each Git operation, so nothing is written to disk beyond `az`'s own sign-in. An account that has to use a personal access token can set `AZURE_DEVOPS_EXT_PAT` for the container instead (for example under `environment:` in `docker-compose.override.yml`); the helper prefers it when present. SSH remotes are unaffected by any of this and keep using the account's own keys.
|
||||
|
||||
Docker cases copy these sign-ins into a case container only when the matching agent-image switch is on (`CODEMAN_AGENT_IMAGE_INSTALL_GH=1` for `~/.config/gh/hosts.yml` and `config.yml`, `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` for the sign-in files from `~/.azure`) and the case has credential seeding on. With a switch off they are never copied, even when the files exist, because a GitHub token or an Azure refresh token is usable by anything in the container. The copies are made when the container is **created**, so an existing case container never picks them up: after turning a switch on, signing in, or rebuilding the agent image, **recreate the case container** (remove it; the next session in that case creates a fresh one).
|
||||
|
||||
The GitHub agent skill for `gh` installs into the runtime account's home in the same session:
|
||||
|
||||
```sh
|
||||
gh skill install cli/cli gh --scope user
|
||||
gh skill update gh # after a later gh release
|
||||
```
|
||||
|
||||
### Versions
|
||||
|
||||
Both CLIs, and the extension, are installed from their vendors' repositories with no version pinned, so they arrive at whatever is current when that build step runs. Docker caches the step, though: `Start-Codeman.sh` rebuilds with the cache, which keeps the versions from the first build until the Dockerfile changes at or above that step or the image is rebuilt with `--no-cache`. They are apt packages owned by root, so they cannot be upgraded from a session; `az extension update --name azure-devops` is the exception and works without a rebuild.
|
||||
|
||||
## Local customisation
|
||||
|
||||
Compose merges `docker-compose.override.yml` on top of `docker-compose.yaml`. Keep host-specific changes there rather than editing `docker-compose.yaml`, so this repository can be updated without losing them. Both `docker-compose.override.yml` and `docker-compose.override.yaml` are ignored by Git.
|
||||
|
||||
`Start-Codeman.sh` names the Compose file explicitly, which disables Compose's automatic discovery of the override file, so the script adds it back when one is present and prints the file it used. Running `docker compose` from this folder without any `-f` option finds it automatically. When passing `-f docker/docker-compose.yaml` from the repository root, add `-f docker/docker-compose.override.yml` as well, or the override is silently ignored.
|
||||
|
||||
An override file adds to and replaces individual settings. It cannot delete a key from `docker-compose.yaml`, and Compose concatenates rather than replaces `ports`, so removing a published port still requires editing `docker-compose.yaml`. The example below replaces the restart policy and adds a mount, leaving every other setting in place:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
codeman:
|
||||
restart: always
|
||||
volumes:
|
||||
- /srv/projects:/srv/projects
|
||||
```
|
||||
|
||||
### Reverse-proxy host allowlist
|
||||
|
||||
Codeman rejects any request whose `Host` header is not on its own allowlist - a
|
||||
DNS-rebinding guard, not a Compose or Docker concern. Loopback, any IP literal,
|
||||
the configured `--host`, and a few tunnel-provider suffixes are allowed by
|
||||
default; a reverse-proxied domain is not, and is rejected with
|
||||
`403 Forbidden: host not allowed` before the request reaches any handler.
|
||||
|
||||
Add the domain with `CODEMAN_ALLOWED_HOSTS` in `.env`:
|
||||
|
||||
```sh
|
||||
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
|
||||
```
|
||||
|
||||
`docker-compose.yaml` forwards it into the container (Compose only passes
|
||||
through the environment keys it explicitly lists, and this is one of them, with
|
||||
an empty default so the line is optional in `.env`).
|
||||
|
||||
See the application's own `docs/wiki/Remote-Access.md` for the full allowlist
|
||||
format and the tunnel providers it accepts by default.
|
||||
|
||||
## Application data storage
|
||||
|
||||
The default configuration uses a host-folder bind mount:
|
||||
|
||||
```yaml
|
||||
volumes:
|
||||
- type: bind
|
||||
source: ${CODEMAN_APPDATA_PATH}
|
||||
target: /home/${CODEMAN_RUNTIME_USER}
|
||||
```
|
||||
|
||||
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/codeman`.
|
||||
|
||||
`CODEMAN_CASES_PATH` is the separate host directory for managed case workspaces. It is mounted into Codeman at the same absolute path, allowing the host Docker daemon to bind it into an isolated case container. Set it to a child directory of `CODEMAN_APPDATA_PATH` unless you deliberately store workspaces elsewhere.
|
||||
|
||||
Compose also exposes `CODEMAN_APPDATA_PATH` to Codeman as `CODEMAN_DOCKER_HOST_HOME`. This lets Docker case seed files, CLI credentials and the hook secret be mounted using paths that exist in the host daemon's filesystem. Direct host installations do not set this variable and retain their existing behaviour.
|
||||
|
||||
Set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` when `docker info` reports `SwapLimit=false`. Codeman continues to apply the configured case memory limit, omits Docker's unsupported `--memory-swap` option, and filters only the daemon's exact swap-capability warning. Every other Docker create error and its exit status remain visible.
|
||||
|
||||
For an existing installation created by a root-running image, change ownership of the application-data directory before upgrading so the configured `PUID` and `PGID` can read the saved credentials and state:
|
||||
|
||||
```sh
|
||||
chown -R 99:100 /mnt/user/appdata/codeman
|
||||
```
|
||||
|
||||
Replace `99:100` and the path with the values from your `.env` file.
|
||||
|
||||
Do not replace this bind mount with a Docker-managed named volume when Docker cases are enabled. Codeman passes seed, credential, transcript and hook-secret bind sources to the host Docker daemon, so their source files must have stable paths in the daemon's filesystem. A named volume does not provide the required host path mapping.
|
||||
|
||||
## Static macvlan networking
|
||||
|
||||
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section from `docker-compose.yaml` and add the following to the `codeman` service. The service and network additions can instead be placed in `docker-compose.override.yml`, but the `ports:` removal cannot, as described under [Local customisation](#local-customisation):
|
||||
|
||||
```yaml
|
||||
mac_address: ${CODEMAN_MAC_ADDRESS}
|
||||
networks:
|
||||
codeman_lan:
|
||||
ipv4_address: ${CODEMAN_IPV4_ADDRESS}
|
||||
```
|
||||
|
||||
Then add this top-level network declaration:
|
||||
|
||||
```yaml
|
||||
networks:
|
||||
codeman_lan:
|
||||
external: true
|
||||
name: ${CODEMAN_MACVLAN_NETWORK}
|
||||
```
|
||||
|
||||
Set `CODEMAN_MACVLAN_NETWORK`, `CODEMAN_IPV4_ADDRESS`, and `CODEMAN_MAC_ADDRESS` in `.env`. The values in `.env.example` match the supplied Unraid example network and should be changed for other hosts.
|
||||
|
||||
### Create a managed macvlan network
|
||||
|
||||
If an external macvlan network does not already exist, use this top-level declaration instead. Do not use it together with the external-network declaration.
|
||||
|
||||
```yaml
|
||||
networks:
|
||||
codeman_lan:
|
||||
driver: macvlan
|
||||
driver_opts:
|
||||
parent: ${CODEMAN_MACVLAN_PARENT}
|
||||
ipam:
|
||||
config:
|
||||
- subnet: ${CODEMAN_MACVLAN_SUBNET}
|
||||
gateway: ${CODEMAN_MACVLAN_GATEWAY}
|
||||
```
|
||||
|
||||
Macvlan containers are ordinarily not reachable from their Docker host without additional host-network routing. Confirm the selected address, MAC address, parent interface, and subnet are reserved and valid for the target network before starting the stack.
|
||||
@@ -1,309 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
script_dir=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||
env_file="$script_dir/.env"
|
||||
compose_file="$script_dir/docker-compose.yaml"
|
||||
|
||||
if [[ ! -f "$env_file" ]]; then
|
||||
printf 'Error: Docker environment file is missing: %s\n' "$env_file" >&2
|
||||
printf 'Create it from %s/.env.example before starting Codeman.\n' "$script_dir" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Naming a Compose file explicitly disables Compose's automatic discovery of
|
||||
# the override file, so it has to be added back by hand. Without this, local
|
||||
# customisation in docker-compose.override.yml is silently ignored. The
|
||||
# candidates are checked in Compose's own precedence order - measured on
|
||||
# Compose v5.5.0 with both present: it uses `.yml` and ignores `.yaml`.
|
||||
override_yml="$script_dir/docker-compose.override.yml"
|
||||
override_yaml="$script_dir/docker-compose.override.yaml"
|
||||
if [[ -f "$override_yml" && -f "$override_yaml" ]]; then
|
||||
printf 'Warning: both %s and %s exist; Compose uses .yml and ignores .yaml.\n' \
|
||||
"$override_yml" "$override_yaml" >&2
|
||||
fi
|
||||
compose_files=(-f "$compose_file")
|
||||
for override_file in "$override_yml" "$override_yaml"; do
|
||||
if [[ -f "$override_file" ]]; then
|
||||
compose_files+=(-f "$override_file")
|
||||
printf 'Using Compose override file: %s\n' "$override_file"
|
||||
break
|
||||
fi
|
||||
done
|
||||
compose_command=(docker compose --env-file "$env_file" "${compose_files[@]}")
|
||||
appdata_path=$(
|
||||
"${compose_command[@]}" config --environment |
|
||||
awk -F= '$1 == "CODEMAN_APPDATA_PATH" { sub(/^[^=]*=/, ""); print; exit }'
|
||||
)
|
||||
cases_path=$(
|
||||
"${compose_command[@]}" config --environment |
|
||||
awk -F= '$1 == "CODEMAN_CASES_PATH" { sub(/^[^=]*=/, ""); print; exit }'
|
||||
)
|
||||
docker_socket=$(
|
||||
"${compose_command[@]}" config --environment |
|
||||
awk -F= '$1 == "DOCKER_SOCKET" { sub(/^[^=]*=/, ""); print; exit }'
|
||||
)
|
||||
|
||||
if [[ -z "$appdata_path" ]]; then
|
||||
printf 'Error: CODEMAN_APPDATA_PATH is not set in %s\n' "$env_file" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [[ ! -d "$appdata_path" ]]; then
|
||||
if [[ "$EUID" == '0' ]]; then
|
||||
printf 'Error: Refusing to create CODEMAN_APPDATA_PATH as root: %s\n' "$appdata_path" >&2
|
||||
printf 'Create it as the unprivileged account that should run Codeman, then retry.\n' >&2
|
||||
exit 1
|
||||
fi
|
||||
mkdir -p -- "$appdata_path"
|
||||
fi
|
||||
|
||||
if [[ -z "$cases_path" ]]; then
|
||||
printf 'Error: CODEMAN_CASES_PATH is not set in %s\n' "$env_file" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# `stat -c` is GNU, `stat -f` is BSD/macOS; the bind sources live on the Docker
|
||||
# host, so both need to work.
|
||||
owner_of() {
|
||||
stat -c '%u:%g' -- "$1" 2>/dev/null || stat -f '%u:%g' "$1" 2>/dev/null
|
||||
}
|
||||
|
||||
if ! owner_ids=$(owner_of "$appdata_path"); then
|
||||
printf 'Error: Cannot determine the owner of CODEMAN_APPDATA_PATH: %s\n' "$appdata_path" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
export PUID=${owner_ids%%:*}
|
||||
export PGID=${owner_ids##*:}
|
||||
|
||||
if [[ "$PUID" == '0' ]]; then
|
||||
printf 'Error: CODEMAN_APPDATA_PATH is owned by root: %s\n' "$appdata_path" >&2
|
||||
printf 'Change the directory ownership to the unprivileged account that should run Codeman.\n' >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Pre-creating this here, exactly like CODEMAN_APPDATA_PATH above, means Compose
|
||||
# never has to materialise a missing bind source itself - which it does as
|
||||
# root:root - so the in-container entrypoint's chown never has to run for this
|
||||
# path at all. It happens AFTER PUID/PGID are known (they come from the appdata
|
||||
# directory just above) so the new directory can be given that exact owner: a
|
||||
# plain `mkdir -p` lands as the invoking user's uid and PRIMARY gid, and on a
|
||||
# host set up the way the README suggests (`chown -R 99:100 <appdata>`) that gid
|
||||
# is not PGID, which the container would then refuse to run on. Unlike appdata,
|
||||
# an EXISTING cases directory is left exactly as it is: the README explicitly
|
||||
# allows pointing this at a normal projects directory the host account already
|
||||
# owns, and the container checks that it is WRITABLE as PUID:PGID rather than
|
||||
# who owns it.
|
||||
if [[ ! -d "$cases_path" ]]; then
|
||||
mkdir -p -- "$cases_path"
|
||||
if [[ "$(owner_of "$cases_path")" != "$PUID:$PGID" ]]; then
|
||||
# As root this always succeeds; as a member of PGID a chgrp does; anyone
|
||||
# else gets the clear error here, where the fix is obvious, rather than a
|
||||
# restart loop from the container.
|
||||
if ! chown -- "$PUID:$PGID" "$cases_path" 2>/dev/null; then
|
||||
printf 'Error: created CODEMAN_CASES_PATH (%s) but could not make it %s:%s (the owner of CODEMAN_APPDATA_PATH).\n' \
|
||||
"$cases_path" "$PUID" "$PGID" >&2
|
||||
printf 'Run `chown %s:%s %s` as root, or create the directory as that account, then retry.\n' \
|
||||
"$PUID" "$PGID" "$cases_path" >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ -z "$docker_socket" || ! -S "$docker_socket" ]]; then
|
||||
printf 'Error: DOCKER_SOCKET is not a Unix socket: %s\n' "${docker_socket:-<unset>}" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if socket_ids=$(stat -c '%u:%g' -- "$docker_socket" 2>/dev/null); then
|
||||
:
|
||||
elif socket_ids=$(stat -f '%u:%g' "$docker_socket" 2>/dev/null); then
|
||||
:
|
||||
else
|
||||
printf 'Error: Cannot determine the owner of DOCKER_SOCKET: %s\n' "$docker_socket" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
export DOCKER_SOCKET_GID=${socket_ids##*:}
|
||||
|
||||
repo_path=${CODEMAN_REPO_PATH:-$(cd -- "$script_dir/.." && pwd)}
|
||||
if [[ ! -d "$repo_path" ]]; then
|
||||
printf 'Error: CODEMAN_REPO_PATH is not a directory: %s\n' "$repo_path" >&2
|
||||
exit 1
|
||||
fi
|
||||
export CODEMAN_REPO_PATH="$repo_path"
|
||||
|
||||
# The in-app updater runs `git checkout` and `npm install` against this checkout
|
||||
# as PUID:PGID. If the directory belongs to someone else, git refuses outright
|
||||
# ("detected dubious ownership") and the update fails at the first step — so warn
|
||||
# here, where the fix is obvious, rather than in a failed update hours later.
|
||||
if repo_owner=$(stat -c '%u' -- "$repo_path" 2>/dev/null || stat -f '%u' "$repo_path" 2>/dev/null); then
|
||||
if [[ "$repo_owner" != "$PUID" ]]; then
|
||||
printf 'Warning: %s is owned by UID %s but Codeman runs as UID %s.\n' "$repo_path" "$repo_owner" "$PUID" >&2
|
||||
printf 'In-app updates will fail until the ownership matches. Codeman itself still starts.\n' >&2
|
||||
fi
|
||||
fi
|
||||
|
||||
if [[ ! -d "$repo_path/.git" ]]; then
|
||||
printf 'Note: %s is not a git checkout, so in-app updates are unavailable.\n' "$repo_path" >&2
|
||||
fi
|
||||
|
||||
# Reads HEAD without requiring a `git` binary on the host — this script
|
||||
# otherwise checks the checkout only by testing for `.git` as a directory, and
|
||||
# resolving refs by hand keeps that the same "no host git needed" guarantee.
|
||||
# ⚠️ A worktree checkout has `.git` as a FILE (`gitdir: <path>`), not a
|
||||
# directory, so this returns nothing there and the volume-refresh check below
|
||||
# silently no-ops — consistent with the `-d .git` test used everywhere else in
|
||||
# this script, not a special case, but worth knowing if a worktree checkout
|
||||
# stops picking up a stale-volume refresh it should have caught.
|
||||
git_head_commit() {
|
||||
local git_dir="$1/.git" head_ref ref_path
|
||||
[[ -d "$git_dir" ]] || return 1
|
||||
head_ref=$(cat -- "$git_dir/HEAD" 2>/dev/null) || return 1
|
||||
if [[ "$head_ref" == ref:* ]]; then
|
||||
ref_path="${head_ref#ref: }"
|
||||
if [[ -f "$git_dir/$ref_path" ]]; then
|
||||
cat -- "$git_dir/$ref_path"
|
||||
else
|
||||
# Packed after a `git gc`; the loose ref file above is gone.
|
||||
awk -v ref="$ref_path" '$2 == ref { print $1; exit }' "$git_dir/packed-refs" 2>/dev/null
|
||||
fi
|
||||
else
|
||||
printf '%s' "$head_ref"
|
||||
fi
|
||||
}
|
||||
|
||||
# Record what the container is about to be built and created FROM. The in-app
|
||||
# updater compares these against the release it wants to apply: a release that
|
||||
# changes either file cannot be applied by the container restarting itself (a
|
||||
# restart reuses the existing image and config), so it is refused and the user
|
||||
# is sent back here. Written on every start, so the baseline always describes
|
||||
# the container that is actually running. See docs/docker-self-update.md.
|
||||
if command -v sha256sum >/dev/null 2>&1; then
|
||||
sha256_of() { sha256sum -- "$1" | cut -d' ' -f1; }
|
||||
elif command -v shasum >/dev/null 2>&1; then
|
||||
sha256_of() { shasum -a 256 -- "$1" | cut -d' ' -f1; }
|
||||
else
|
||||
sha256_of() { printf ''; }
|
||||
fi
|
||||
|
||||
dockerfile_sha=$(sha256_of "$script_dir/server.Dockerfile")
|
||||
compose_sha=$(sha256_of "$compose_file")
|
||||
if [[ -n "$dockerfile_sha" && -n "$compose_sha" ]]; then
|
||||
# $CODEMAN_APPDATA_PATH is mounted at the runtime account's home, so this is
|
||||
# dataPath('docker-env-applied.json') as the server inside the container sees it.
|
||||
state_dir="$appdata_path/.codeman"
|
||||
mkdir -p -- "$state_dir"
|
||||
printf '{\n "dockerfileSha256": "%s",\n "composeSha256": "%s"\n}\n' \
|
||||
"$dockerfile_sha" "$compose_sha" >"$state_dir/docker-env-applied.json.tmp"
|
||||
mv -- "$state_dir/docker-env-applied.json.tmp" "$state_dir/docker-env-applied.json"
|
||||
# A root-run start (common on Unraid) would otherwise leave a root-owned
|
||||
# `.codeman` on a FIRST start, before the container has created it as PUID,
|
||||
# and the unprivileged server could then never write its own state there.
|
||||
if [[ "$EUID" == '0' ]]; then
|
||||
chown -- "$PUID:$PGID" "$state_dir" "$state_dir/docker-env-applied.json"
|
||||
fi
|
||||
else
|
||||
printf 'Warning: no sha256 tool found; in-app updates will not detect environment changes.\n' >&2
|
||||
fi
|
||||
|
||||
# codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded from
|
||||
# the image only while EMPTY, so a rebuilt image's fresh output sits unused
|
||||
# behind old volume content until something clears it. The in-app self-updater
|
||||
# never hits this — it rebuilds INSIDE the running container, into the very
|
||||
# volume already in use — but a `docker compose build` triggered from outside
|
||||
# it (this script, after a `git pull`) does: the container comes back up
|
||||
# looking unchanged. Detect that here and clear just the affected volume(s) so
|
||||
# the build below actually takes effect. Best-effort: with no sha256 tool this
|
||||
# quietly does nothing, same as the environment-gate block above.
|
||||
volumes_to_refresh=()
|
||||
if [[ -n "$dockerfile_sha" ]]; then
|
||||
repo_head=$(git_head_commit "$repo_path" || true)
|
||||
lockfile_sha=$(sha256_of "$repo_path/package-lock.json" 2>/dev/null || true)
|
||||
source_state_file="$state_dir/docker-build-source.json"
|
||||
prev_head=''
|
||||
prev_lockfile_sha=''
|
||||
if [[ -f "$source_state_file" ]]; then
|
||||
prev_head=$(sed -n 's/.*"headCommit": *"\([^"]*\)".*/\1/p' "$source_state_file")
|
||||
prev_lockfile_sha=$(sed -n 's/.*"lockfileSha256": *"\([^"]*\)".*/\1/p' "$source_state_file")
|
||||
fi
|
||||
|
||||
[[ -n "$repo_head" && "$repo_head" != "$prev_head" ]] && volumes_to_refresh+=('codeman-dist')
|
||||
[[ -n "$lockfile_sha" && "$lockfile_sha" != "$prev_lockfile_sha" ]] && volumes_to_refresh+=('codeman-node-modules')
|
||||
fi
|
||||
|
||||
if [[ ${#volumes_to_refresh[@]} -eq 0 ]]; then
|
||||
exec "${compose_command[@]}" up --build -d
|
||||
fi
|
||||
|
||||
# Runs even on this script's very first invocation against an EXISTING
|
||||
# deployment, deliberately: that deployment's volumes may already be stale
|
||||
# (there was no earlier version of this check to have caught it), and clearing
|
||||
# an already-empty or nonexistent volume is a harmless no-op, so there is no
|
||||
# fresh-install case this needs to avoid.
|
||||
printf 'Source changed since the last start; refreshing: %s\n' "${volumes_to_refresh[*]}"
|
||||
|
||||
# Build BEFORE taking the stack down: the image build is the slow part and needs
|
||||
# no container stopped, so the deployment is offline only for the recreate.
|
||||
"${compose_command[@]}" build
|
||||
|
||||
# `com.docker.compose.volume` is the volume KEY, not a project-qualified name -
|
||||
# a second stack on the same host (a beta instance started with a different
|
||||
# COMPOSE_PROJECT_NAME, say) that also declares a volume keyed `codeman-dist`
|
||||
# shares that label, and `head -n1` would pick whichever the daemon happens to
|
||||
# list first. Scope the lookup to THIS stack's own resolved project name so it
|
||||
# can only ever match this stack's volume. The name is read from the resolved
|
||||
# config's top-level `name` key, indentation-agnostic (the formatting is not a
|
||||
# contract), and the FIRST `name` in the output is the project's: nested ones
|
||||
# (a network's `name:`) come later. `--format json` needs Compose v2.3+.
|
||||
project_name=$(
|
||||
"${compose_command[@]}" config --format json 2>/dev/null |
|
||||
sed -n 's/^[[:space:]]*"name":[[:space:]]*"\([^"]*\)".*$/\1/p' | head -n1
|
||||
)
|
||||
|
||||
"${compose_command[@]}" down
|
||||
|
||||
# Track whether the volumes were actually cleared. The marker below is written
|
||||
# ONLY on success: with an unresolvable project name the label filter would
|
||||
# match nothing, nothing would be removed, and a marker recording the new HEAD
|
||||
# would stop this check from ever firing again while the stale volume kept
|
||||
# serving old code. A failed removal likewise leaves the marker alone, so the
|
||||
# next start retries, and the stack is brought back up regardless rather than
|
||||
# left down.
|
||||
refreshed=1
|
||||
if [[ -z "$project_name" ]]; then
|
||||
# The documented reset (docs/docker-self-update.md): both volumes re-seed from
|
||||
# the image by a plain copy, so clearing the extra one costs a copy, not data.
|
||||
printf 'Warning: could not resolve the Compose project name; clearing both build-artefact volumes with `down --volumes` instead.\n' >&2
|
||||
"${compose_command[@]}" down --volumes || refreshed=0
|
||||
else
|
||||
for key in "${volumes_to_refresh[@]}"; do
|
||||
volume_name=$(
|
||||
docker volume ls -q \
|
||||
--filter "label=com.docker.compose.volume=$key" \
|
||||
--filter "label=com.docker.compose.project=$project_name" |
|
||||
head -n1
|
||||
)
|
||||
if [[ -n "$volume_name" ]] && ! docker volume rm -- "$volume_name"; then
|
||||
printf 'Warning: could not remove volume %s; it will be retried on the next start.\n' "$volume_name" >&2
|
||||
refreshed=0
|
||||
fi
|
||||
done
|
||||
fi
|
||||
|
||||
if [[ "$refreshed" == '1' ]]; then
|
||||
printf '{\n "headCommit": "%s",\n "lockfileSha256": "%s"\n}\n' \
|
||||
"$repo_head" "$lockfile_sha" >"$source_state_file.tmp"
|
||||
mv -- "$source_state_file.tmp" "$source_state_file"
|
||||
if [[ "$EUID" == '0' ]]; then
|
||||
chown -- "$PUID:$PGID" "$source_state_file"
|
||||
fi
|
||||
else
|
||||
printf 'Warning: the build-artefact volumes were NOT refreshed; the container may serve stale code until the next successful start.\n' >&2
|
||||
fi
|
||||
|
||||
# Already built above, so no --build here: a second build would only re-check
|
||||
# the cache.
|
||||
exec "${compose_command[@]}" up -d
|
||||
@@ -1,255 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# The scripted major-update path for the Docker Compose deployment.
|
||||
#
|
||||
# docker/README.md and docs/docker-self-update.md both point operators here for
|
||||
# anything the in-app updater itself refuses to apply: a changed
|
||||
# `server.Dockerfile`, a changed `docker-compose.yaml`, or a new required
|
||||
# `.env.example` key. None of those can be applied by a container restarting
|
||||
# itself — a restart reuses the existing image and configuration (see "The
|
||||
# environment gate" in docs/docker-self-update.md) — so this script does the
|
||||
# three things an in-place update cannot: force a real image rebuild with no
|
||||
# layer cache, stop the stack, then hand off to Start-Codeman.sh for the same
|
||||
# careful PUID/PGID, override-file and fingerprint handling every other start
|
||||
# goes through.
|
||||
#
|
||||
# ⚠️ Build BEFORE stopping the stack, deliberately, same reasoning as
|
||||
# Start-Codeman.sh's own build-then-down ordering: the build needs nothing
|
||||
# stopped, so a slow --no-cache rebuild costs no downtime, and a build failure
|
||||
# (a bad Dockerfile edit, a network blip pulling a base image) leaves the
|
||||
# ALREADY-RUNNING stack untouched instead of stopped with nothing to bring it
|
||||
# back.
|
||||
#
|
||||
# ⚠️ Clears the codeman-node-modules/codeman-dist named volumes by DEFAULT.
|
||||
# Docker seeds a named volume from the image only while that volume is EMPTY,
|
||||
# so a rebuilt image's fresh node_modules/dist otherwise sit unused behind a
|
||||
# volume's old content and the container comes back up looking unchanged —
|
||||
# exactly wrong for a script whose whole point is "be certain of what ships".
|
||||
# Start-Codeman.sh clears codeman-dist when the checkout's HEAD moved and
|
||||
# codeman-node-modules only when `package-lock.json` changed. A released
|
||||
# server.Dockerfile change arrives through `git pull`, so HEAD moves and dist
|
||||
# is refreshed, but a Dockerfile change that bumps the Node base image leaves
|
||||
# the lockfile untouched while every native module (node-pty is compiled from
|
||||
# source, there is no Linux prebuild) has to be rebuilt against the new Node
|
||||
# ABI. Start-Codeman.sh would keep the old codeman-node-modules volume, and it
|
||||
# never builds with --no-cache. This script clears BOTH volumes, and ONLY
|
||||
# those two (targeted `docker volume rm` by Compose label, never
|
||||
# `down --volumes`, which would also take any volume an override file adds).
|
||||
# Pass --keep-volumes to opt out and reuse whatever is already in them.
|
||||
#
|
||||
# Usage: docker/Update-Codeman.sh [--keep-volumes]
|
||||
# --keep-volumes Do not clear codeman-node-modules/codeman-dist. Safe to
|
||||
# combine with a source change Start-Codeman.sh's own
|
||||
# detection would have cleared anyway; unsafe if the reason
|
||||
# you are here is a change to server.Dockerfile alone.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
script_dir=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
|
||||
env_file="$script_dir/.env"
|
||||
compose_file="$script_dir/docker-compose.yaml"
|
||||
|
||||
keep_volumes=0
|
||||
for arg in "$@"; do
|
||||
case "$arg" in
|
||||
--keep-volumes)
|
||||
keep_volumes=1
|
||||
;;
|
||||
--help | -h)
|
||||
printf 'Usage: bash %s [--keep-volumes]\n' "$0"
|
||||
exit 0
|
||||
;;
|
||||
*)
|
||||
printf 'Error: unrecognised argument: %s\n' "$arg" >&2
|
||||
printf 'Usage: bash %s [--keep-volumes]\n' "$0" >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [[ ! -f "$env_file" ]]; then
|
||||
printf 'Error: Docker environment file is missing: %s\n' "$env_file" >&2
|
||||
printf 'Create it from %s/.env.example before running this script.\n' "$script_dir" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Same override-file discovery as Start-Codeman.sh, and deliberately kept in
|
||||
# step with it: a stack built here and started there must resolve to the exact
|
||||
# same Compose files, or this script's build could target a configuration the
|
||||
# handoff's own `up` never actually uses. Compose's own precedence (measured on
|
||||
# v5.5.0 with both present: it uses .yml and ignores .yaml).
|
||||
override_yml="$script_dir/docker-compose.override.yml"
|
||||
override_yaml="$script_dir/docker-compose.override.yaml"
|
||||
if [[ -f "$override_yml" && -f "$override_yaml" ]]; then
|
||||
printf 'Warning: both %s and %s exist; Compose uses .yml and ignores .yaml.\n' \
|
||||
"$override_yml" "$override_yaml" >&2
|
||||
fi
|
||||
compose_files=(-f "$compose_file")
|
||||
for override_file in "$override_yml" "$override_yaml"; do
|
||||
if [[ -f "$override_file" ]]; then
|
||||
compose_files+=(-f "$override_file")
|
||||
printf 'Using Compose override file: %s\n' "$override_file"
|
||||
break
|
||||
fi
|
||||
done
|
||||
compose_command=(docker compose --env-file "$env_file" "${compose_files[@]}")
|
||||
|
||||
# Collision guard. Start-Codeman.sh has no equivalent; this is the only one,
|
||||
# and it has to run before this script's own --no-cache build, `down` and
|
||||
# volume removal below. docker-compose.yaml hard-codes `name: codeman`, so a
|
||||
# second checkout run without COMPOSE_PROJECT_NAME resolves to the SAME Compose
|
||||
# project as any other checkout on the host and would operate on ITS
|
||||
# containers and volumes.
|
||||
#
|
||||
# The project name is read from the resolved config's top-level `name` key
|
||||
# (the first `name` in the output; nested ones come later), the same parse
|
||||
# Start-Codeman.sh uses. `--format json` needs Compose v2.3+. This is the first
|
||||
# `docker` call the script makes, so its failure is reported here rather than
|
||||
# left to `set -e`, which would exit with no output at all.
|
||||
if ! project_config=$("${compose_command[@]}" config --format json); then
|
||||
printf 'Error: `docker compose config --format json` failed (see the message above, if any).\n' >&2
|
||||
printf 'Check that Docker and Compose v2.3+ are installed and on PATH, and that\n' >&2
|
||||
printf '%s and the Compose files in %s are valid.\n' "$env_file" "$script_dir" >&2
|
||||
exit 1
|
||||
fi
|
||||
project_name=$(
|
||||
printf '%s\n' "$project_config" |
|
||||
sed -n 's/^[[:space:]]*"name":[[:space:]]*"\([^"]*\)".*$/\1/p' | head -n1
|
||||
)
|
||||
if [[ -n "$project_name" ]]; then
|
||||
# `|| true` on the pipeline's LAST command: under `set -o pipefail`, `grep -v`
|
||||
# exits 1 when nothing survives the filter — the ordinary, no-collision case,
|
||||
# since `docker ps` finds nothing at all on a first-ever deployment or a
|
||||
# single matching (own) working_dir gets filtered out. Without it, that exit
|
||||
# status propagates through the command substitution and `set -e` aborts the
|
||||
# WHOLE script right here, every time, regardless of whether a collision
|
||||
# actually exists — caught only by actually running this end-to-end (a
|
||||
# static text/regex check on the source cannot see it). The empty-line
|
||||
# filter keeps a container with no working_dir label from winning head -n1
|
||||
# and hiding a real collision behind it.
|
||||
other_working_dir=$(
|
||||
docker ps -a --filter "label=com.docker.compose.project=$project_name" \
|
||||
--format '{{.Label "com.docker.compose.project.working_dir"}}' 2>/dev/null |
|
||||
grep -v -F -x -- "$script_dir" | grep -v '^$' | head -n1 || true
|
||||
)
|
||||
if [[ -n "$other_working_dir" ]]; then
|
||||
printf 'Error: Compose project "%s" is already in use by a DIFFERENT checkout:\n' "$project_name" >&2
|
||||
printf ' %s\n' "$other_working_dir" >&2
|
||||
printf 'This checkout is:\n' >&2
|
||||
printf ' %s\n' "$script_dir" >&2
|
||||
printf '\n' >&2
|
||||
printf 'docker-compose.yaml hard-codes `name: %s`, so two checkouts on the same host\n' "$project_name" >&2
|
||||
printf 'collide unless each one sets a distinct COMPOSE_PROJECT_NAME. Continuing would\n' >&2
|
||||
printf 'rebuild and stop the OTHER checkout'"'"'s running container and, by default,\n' >&2
|
||||
printf 'delete its codeman-node-modules/codeman-dist volumes.\n' >&2
|
||||
printf '\n' >&2
|
||||
printf 'Fix: export COMPOSE_PROJECT_NAME=<something-unique-to-this-checkout> before\n' >&2
|
||||
printf 'running this script, then retry.\n' >&2
|
||||
printf '\n' >&2
|
||||
printf 'If instead THIS checkout was moved or renamed after its container was created,\n' >&2
|
||||
printf 'the path above is its own old location: remove the old container (for example\n' >&2
|
||||
printf '`docker rm -f <container>` for the codeman container) and retry, rather than\n' >&2
|
||||
printf 'setting COMPOSE_PROJECT_NAME, which would start a second project beside it.\n' >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# Same owner-detection Start-Codeman.sh uses to derive PUID/PGID for its own
|
||||
# build — without it, the --no-cache build below gets Compose's untouched
|
||||
# default of 1000:1000, and on any host whose appdata owner differs (99:100 on
|
||||
# Unraid, per docker/README.md's chown example), Start-Codeman.sh's own
|
||||
# correctly-PUID'd build during the handoff then rebuilds those layers with the
|
||||
# right values anyway — so the "no cache, certain of what ships" image this
|
||||
# script produces is not the one that actually ends up running.
|
||||
#
|
||||
# Deliberately NOT the same as Start-Codeman.sh's own handling of a MISSING
|
||||
# appdata directory (which creates it): this script updates an EXISTING
|
||||
# deployment, so a missing appdata path means there is nothing here yet to
|
||||
# update, and creating one would just be this script quietly doing
|
||||
# Start-Codeman.sh's first-run job worse.
|
||||
appdata_path=$(
|
||||
"${compose_command[@]}" config --environment |
|
||||
awk -F= '$1 == "CODEMAN_APPDATA_PATH" { sub(/^[^=]*=/, ""); print; exit }'
|
||||
)
|
||||
if [[ -z "$appdata_path" || ! -d "$appdata_path" ]]; then
|
||||
printf 'Error: CODEMAN_APPDATA_PATH is not set or does not exist: %s\n' "${appdata_path:-<unset>}" >&2
|
||||
printf 'Run docker/Start-Codeman.sh first to set up a new deployment.\n' >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# `stat -c` is GNU, `stat -f` is BSD/macOS; the bind source lives on the Docker
|
||||
# host, so both need to work. Identical to Start-Codeman.sh's own helper.
|
||||
owner_of() {
|
||||
stat -c '%u:%g' -- "$1" 2>/dev/null || stat -f '%u:%g' "$1" 2>/dev/null
|
||||
}
|
||||
|
||||
if ! owner_ids=$(owner_of "$appdata_path"); then
|
||||
printf 'Error: Cannot determine the owner of CODEMAN_APPDATA_PATH: %s\n' "$appdata_path" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
export PUID=${owner_ids%%:*}
|
||||
export PGID=${owner_ids##*:}
|
||||
|
||||
if [[ "$PUID" == '0' ]]; then
|
||||
printf 'Error: CODEMAN_APPDATA_PATH is owned by root: %s\n' "$appdata_path" >&2
|
||||
printf 'Change the directory ownership to the unprivileged account that should run Codeman.\n' >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# --no-cache, always: a plain `build` reuses cached layers (npm install, apt
|
||||
# packages, the CLI installs baked into the image) and can silently keep them
|
||||
# frozen at whatever they were the day the cache was populated — exactly wrong
|
||||
# for a major update, whose whole point is being certain of what actually
|
||||
# ships. `scripts/build-agent-image.mjs` makes the same call for the same
|
||||
# reason (see its entry in CLAUDE.md's Additional Commands table). Runs BEFORE
|
||||
# the stack is stopped — see the header comment for why.
|
||||
printf 'Building a fresh image (--no-cache)...\n'
|
||||
"${compose_command[@]}" build --no-cache
|
||||
|
||||
printf 'Stopping the stack...\n'
|
||||
if [[ "$keep_volumes" == '1' || -n "$project_name" ]]; then
|
||||
"${compose_command[@]}" down
|
||||
else
|
||||
# No resolvable project name means the label filter below could match
|
||||
# nothing, so fall back to Compose's own removal, and say what it really does.
|
||||
printf 'Warning: could not resolve the Compose project name; clearing EVERY named volume\n' >&2
|
||||
printf 'in this Compose project (override file included) with `down --volumes` instead.\n' >&2
|
||||
"${compose_command[@]}" down --volumes
|
||||
fi
|
||||
|
||||
# Targeted removal of exactly the two build-artefact volumes, scoped by label to
|
||||
# THIS project (the volume key alone is shared by any other stack declaring the
|
||||
# same key). Same lookup as Start-Codeman.sh's refresh. A failure is reported,
|
||||
# not fatal: the stack is already down, and the handoff below is what brings
|
||||
# it back up.
|
||||
if [[ "$keep_volumes" != '1' && -n "$project_name" ]]; then
|
||||
printf 'Clearing the codeman-node-modules/codeman-dist volumes (pass --keep-volumes to skip).\n'
|
||||
for key in codeman-node-modules codeman-dist; do
|
||||
volume_name=$(
|
||||
docker volume ls -q \
|
||||
--filter "label=com.docker.compose.volume=$key" \
|
||||
--filter "label=com.docker.compose.project=$project_name" |
|
||||
head -n1
|
||||
) || volume_name=''
|
||||
if [[ -n "$volume_name" ]] && ! docker volume rm -- "$volume_name"; then
|
||||
printf 'Warning: could not remove volume %s; the container may keep serving the\n' "$volume_name" >&2
|
||||
printf 'previous build from it. Remove it by hand and rerun this script.\n' >&2
|
||||
fi
|
||||
done
|
||||
fi
|
||||
|
||||
# Start-Codeman.sh does everything a plain `up -d` does not: re-derives
|
||||
# PUID/PGID, pre-creates CODEMAN_CASES_PATH with the right ownership, resolves
|
||||
# DOCKER_SOCKET_GID, records the server.Dockerfile/docker-compose.yaml
|
||||
# fingerprint the in-app updater's gate reads on every future update, and
|
||||
# starts the (already freshly built) image. Reimplementing any of that here
|
||||
# would only risk drifting out of step with it — hand off instead, exactly as
|
||||
# docs/docker-self-update.md's own reset procedure does.
|
||||
#
|
||||
# ⚠️ `bash`, not a bare exec of the path: Start-Codeman.sh is committed
|
||||
# non-executable (100644), the same as this script, and is documented
|
||||
# everywhere as `bash docker/Start-Codeman.sh` rather than
|
||||
# `./docker/Start-Codeman.sh` — execing the bare path fails with EACCES.
|
||||
printf 'Handing off to Start-Codeman.sh...\n'
|
||||
exec bash "$script_dir/Start-Codeman.sh"
|
||||
@@ -1,262 +0,0 @@
|
||||
# Codeman agent base image (built locally by scripts/build-agent-image.mjs).
|
||||
#
|
||||
# Contains the agent toolchain (node + the CLIs + git/tmux/ripgrep) but NO
|
||||
# secrets: credentials are delivered at RUNTIME via bind mounts (~/.claude etc.)
|
||||
# or name-only `docker exec --env`, never baked in, so `docker save` exports stay
|
||||
# secret-free. tmux is a HARD prerequisite (the in-container tmux is what makes a
|
||||
# reconnect durable), so it is installed here and probed before launch.
|
||||
#
|
||||
# HOME is made writable by an ARBITRARY host uid via the OpenShift "gid 0,
|
||||
# group-writable" convention: on Linux we run `--user <hostUid>:0`, so the agent
|
||||
# uid is the host uid (workspace files stay host-owned) while gid 0 keeps $HOME
|
||||
# writable even though the uid is not the baked 1000.
|
||||
FROM node:22-bookworm-slim
|
||||
|
||||
# Base toolchain. `curl` is needed for the hook callbacks (`curl -sk $CODEMAN_API_URL`),
|
||||
# `procps` for `ps`, `tmux` for the durable in-container session.
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends \
|
||||
git \
|
||||
libsecret-1-0 \
|
||||
tmux \
|
||||
ripgrep \
|
||||
curl \
|
||||
ca-certificates \
|
||||
less \
|
||||
procps \
|
||||
openssh-client \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# GitHub CLI and Azure CLI (+ the azure-devops extension) with the same system
|
||||
# git credential helpers as docker/server.Dockerfile, so an agent in a Docker
|
||||
# case can clone and push to private GitHub / Azure DevOps repositories. The
|
||||
# sign-ins themselves are NOT baked in: `~/.config/gh` and `~/.azure` are seeded
|
||||
# per container at launch like every other CLI's credentials (CRED_STORES in
|
||||
# src/docker-hosts.ts), and a helper whose CLI is not signed in prints nothing,
|
||||
# so git fails fast instead of prompting. See server.Dockerfile for why the
|
||||
# vendor apt repositories are configured here rather than via deb_install.sh.
|
||||
#
|
||||
# Each is OPT-IN and OFF by default, like the server image: CODEMAN_INSTALL_GH=1
|
||||
# / CODEMAN_INSTALL_AZ=1 turn one on; off leaves no repository, package,
|
||||
# extension or helper entry. scripts/build-agent-image.mjs and the in-app
|
||||
# auto-build pass them from CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ in their own
|
||||
# environment (for the Compose deployment: `environment:` in
|
||||
# docker-compose.override.yml), and pass nothing when those are unset, so
|
||||
# these defaults (off) apply.
|
||||
ARG CODEMAN_INSTALL_GH=0
|
||||
ARG CODEMAN_INSTALL_AZ=0
|
||||
RUN set -eux; \
|
||||
for flag in "CODEMAN_INSTALL_GH=${CODEMAN_INSTALL_GH}" "CODEMAN_INSTALL_AZ=${CODEMAN_INSTALL_AZ}"; do \
|
||||
case "${flag#*=}" in 0|1) ;; *) echo "${flag%%=*} must be 0 or 1, got '${flag#*=}'" >&2; exit 1;; esac; \
|
||||
done; \
|
||||
codename="$(. /etc/os-release && echo "${VERSION_CODENAME}")"; \
|
||||
arch="$(dpkg --print-architecture)"; \
|
||||
pkgs=""; \
|
||||
install -d -m 0755 /etc/apt/keyrings; \
|
||||
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then \
|
||||
curl -fsSL -o /etc/apt/keyrings/githubcli-archive-keyring.gpg \
|
||||
https://cli.github.com/packages/githubcli-archive-keyring.gpg; \
|
||||
chmod go+r /etc/apt/keyrings/githubcli-archive-keyring.gpg; \
|
||||
echo "deb [arch=${arch} signed-by=/etc/apt/keyrings/githubcli-archive-keyring.gpg] https://cli.github.com/packages stable main" \
|
||||
> /etc/apt/sources.list.d/github-cli.list; \
|
||||
pkgs="${pkgs} gh"; \
|
||||
fi; \
|
||||
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
|
||||
curl -fsSL -o /etc/apt/keyrings/microsoft.asc \
|
||||
https://packages.microsoft.com/keys/microsoft.asc; \
|
||||
chmod go+r /etc/apt/keyrings/microsoft.asc; \
|
||||
echo "deb [arch=${arch} signed-by=/etc/apt/keyrings/microsoft.asc] https://packages.microsoft.com/repos/azure-cli/ ${codename} main" \
|
||||
> /etc/apt/sources.list.d/azure-cli.list; \
|
||||
pkgs="${pkgs} azure-cli"; \
|
||||
fi; \
|
||||
if [ -n "${pkgs}" ]; then \
|
||||
apt-get update; \
|
||||
apt-get install -y --no-install-recommends ${pkgs}; \
|
||||
rm -rf /var/lib/apt/lists/*; \
|
||||
fi; \
|
||||
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then gh --version; fi; \
|
||||
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then az version --output none; fi
|
||||
|
||||
# Outside HOME so the seeded `~/.azure` (auth files only) never has to carry
|
||||
# extensions. gid 0 + group-writable, the same arbitrary-uid convention as HOME
|
||||
# below, so `az extension update` works as whatever uid the container runs as.
|
||||
# Created even without az; an empty directory costs nothing.
|
||||
ENV AZURE_EXTENSION_DIR=/opt/az-extensions
|
||||
RUN set -eux; \
|
||||
install -d -m 0755 "${AZURE_EXTENSION_DIR}"; \
|
||||
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
|
||||
az extension add --name azure-devops --only-show-errors; \
|
||||
rm -rf /root/.azure; \
|
||||
fi; \
|
||||
chgrp -R 0 "${AZURE_EXTENSION_DIR}"; \
|
||||
chmod -R g=u "${AZURE_EXTENSION_DIR}"
|
||||
|
||||
# Only an installed CLI gets a helper entry (see server.Dockerfile).
|
||||
COPY docker/git-credential-azure-cli /usr/local/bin/git-credential-azure-cli
|
||||
RUN set -eux; \
|
||||
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then \
|
||||
for host in https://github.com https://gist.github.com; do \
|
||||
git config --system "credential.${host}.helper" '!/usr/bin/gh auth git-credential'; \
|
||||
done; \
|
||||
fi; \
|
||||
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
|
||||
chmod 0755 /usr/local/bin/git-credential-azure-cli; \
|
||||
for host in https://dev.azure.com 'https://*.visualstudio.com'; do \
|
||||
git config --system "credential.${host}.helper" /usr/local/bin/git-credential-azure-cli; \
|
||||
git config --system "credential.${host}.useHttpPath" true; \
|
||||
done; \
|
||||
else \
|
||||
rm -f /usr/local/bin/git-credential-azure-cli; \
|
||||
fi
|
||||
|
||||
# The npm-published agent CLIs, supplied by scripts/build-agent-image.mjs from
|
||||
# config/clis.stock.json so a new stock CLI needs no edit here. The default is
|
||||
# today's literal list, so a bare `docker build` still produces the same image.
|
||||
#
|
||||
# ⚠️ Expanded UNQUOTED on purpose: word splitting is what turns the list into
|
||||
# several arguments. Every token is validated against
|
||||
# ^[@A-Za-z0-9][@A-Za-z0-9/._-]*$ on the producing side
|
||||
# (scripts/lib/cli-catalog.mjs) precisely because of that.
|
||||
#
|
||||
# ⚠️ Filtered on each entry's `enabled` flag, so a CLI that ships disabled is
|
||||
# never baked into every image.
|
||||
#
|
||||
# Pinning is left to the rebuild cadence (see docs/docker-cases-plan.md,
|
||||
# user-decision 2).
|
||||
# ⚠️ The default is in REGISTRY order, byte-identical to what the generator emits.
|
||||
# A different order is a different RUN string, which is a different layer hash and
|
||||
# so a needless cache miss between a bare `docker build` and a scripted one.
|
||||
ARG CLI_NPM_PACKAGES="@anthropic-ai/claude-code opencode-ai @openai/codex @google/gemini-cli"
|
||||
# uv/uvx: MCP servers are commonly launched with `uvx <package>` (e.g. the Nginx
|
||||
# Proxy Manager MCP), and Codex failed to enable them with "uvx not found". Copied
|
||||
# from the pinned upstream image into root-owned /usr/local/bin, never pip-installed.
|
||||
COPY --from=ghcr.io/astral-sh/uv:0.9 /uv /uvx /usr/local/bin/
|
||||
RUN npm install -g ${CLI_NPM_PACKAGES} \
|
||||
&& npm cache clean --force
|
||||
|
||||
# Antigravity (`agy`) is NOT on npm — Google ships a standalone binary through its
|
||||
# own installer, so it needs its own step. `--dir /usr/local/bin` is load-bearing:
|
||||
# the installer's default target is `$HOME/.local/bin`, which at build time is
|
||||
# root's home and would be unreachable by the `agent` user the container runs as.
|
||||
# ⚠️ This binary is ~190MB on its own; it is the single largest layer in the image.
|
||||
RUN curl -fsSL https://antigravity.google/cli/install.sh | bash -s -- --dir /usr/local/bin \
|
||||
&& chmod 755 /usr/local/bin/agy \
|
||||
&& agy --version
|
||||
|
||||
# Pi (pi.dev). Upstream documents --ignore-scripts (pi needs no lifecycle scripts);
|
||||
# kept out of the shared npm block above so the flag cannot silently change how the
|
||||
# rest of that block's CLIs install — a fixed count would go stale here since
|
||||
# CLI_NPM_PACKAGES (above) is now a generated, dynamic list rather than a hand-kept one.
|
||||
RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
|
||||
&& npm cache clean --force \
|
||||
&& pi --version
|
||||
|
||||
# Grok Build (`grok`, xAI) is NOT on npm: a standalone ~160MB Rust binary through
|
||||
# xAI's installer, which targets $HOME/.grok/bin with no --dir override. At build
|
||||
# time that is root's home and unreachable by the `agent` user, so copy the binary
|
||||
# into /usr/local/bin and drop root's ~/.grok in the same layer so the image does
|
||||
# not carry the download twice. The staging cp -T is what makes this survive the
|
||||
# installer's own behavior EITHER way: newer installers already symlink
|
||||
# /usr/local/bin/grok -> /root/.grok/bin/grok, and a direct `cp -L` onto that
|
||||
# symlink fails with "same file" (2026-08-24 rebuild), while removing the link
|
||||
# first and copying fresh works for both old and new installers.
|
||||
RUN curl -fsSL https://x.ai/cli/install.sh | bash \
|
||||
&& cp -L /root/.grok/bin/grok /usr/local/bin/grok.real \
|
||||
&& rm -f /usr/local/bin/grok \
|
||||
&& mv /usr/local/bin/grok.real /usr/local/bin/grok \
|
||||
&& chmod 755 /usr/local/bin/grok \
|
||||
&& rm -rf /root/.grok /root/.local/bin/grok /root/.local/bin/agent \
|
||||
&& grok --version
|
||||
|
||||
# DeepSeek Harness (`dsh`). A normal npm package, but the ONLY entry here whose
|
||||
# binary runs nothing on its own: `dsh` is a profile launcher, and DeepSeek ships
|
||||
# only `web` and `headless`, so without an interactive profile a
|
||||
# `mode: 'deepseek'` container would start a pane that dies on arrival. The
|
||||
# profile itself is installed further down, into the `agent` HOME, because
|
||||
# Codeman deliberately does NOT seed `profiles/` from the host: it is a
|
||||
# per-profile node_modules tree, host-arch-specific and far too large to copy on
|
||||
# every container start.
|
||||
# ⚠️ `pnpm` is a HARD dependency of `dsh plugin`, not optional tooling: the
|
||||
# subcommand is a thin forwarder that `spawnSync`s a literal `pnpm` with no
|
||||
# fallback to npm, so on an image without it the profile install below dies
|
||||
# with `dsh: pnpm not found on PATH` / exit 127 and takes the whole build with
|
||||
# it (issue #352). It stays on PATH at runtime too, so a container user can run
|
||||
# `dsh plugin add` themselves.
|
||||
RUN npm install -g @deepseek-ai/dsh pnpm \
|
||||
&& npm cache clean --force \
|
||||
&& dsh --version \
|
||||
&& pnpm --version
|
||||
|
||||
# OMP (Oh My Pi) is NOT on npm: a standalone binary via omp.sh's installer, which
|
||||
# targets $HOME/.local/bin with no --dir override (verified 2026-08-27 — the
|
||||
# resolver's OMP_SEARCH_DIRS lists ~/.omp/bin first, which turned out to be the
|
||||
# WRONG guess for the installer's actual target; build this step for real
|
||||
# rather than trust that ordering). At build time $HOME is root's home and
|
||||
# unreachable by the `agent` user, so copy the binary into /usr/local/bin and
|
||||
# drop root's ~/.local/bin/omp in the same layer so the image does not carry
|
||||
# the download twice.
|
||||
RUN curl -fsSL https://omp.sh/install | sh \
|
||||
&& cp -L /root/.local/bin/omp /usr/local/bin/omp.real \
|
||||
&& rm -f /usr/local/bin/omp \
|
||||
&& mv /usr/local/bin/omp.real /usr/local/bin/omp \
|
||||
&& chmod 755 /usr/local/bin/omp \
|
||||
&& rm -f /root/.local/bin/omp \
|
||||
&& omp --version
|
||||
|
||||
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
|
||||
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
|
||||
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
|
||||
# only matters for a hand-run / Docker Desktop container. gid 0 + group-writable
|
||||
# HOME (OpenShift arbitrary-uid convention) keeps $HOME writable for any uid.
|
||||
# UTF-8 locale so tmux/Ink render Unicode box-drawing instead of VT100 ACS `q`
|
||||
# glyphs (C.UTF-8 is built into glibc; no locales package needed). Codeman also
|
||||
# sets these at run time so containers built before this line still get UTF-8.
|
||||
ENV LANG=C.UTF-8 LC_ALL=C.UTF-8
|
||||
ENV HOME=/home/agent
|
||||
# `.claude` (+ `.claude/projects` mount point) and `.codex` (+ `.codex/sessions`) are
|
||||
# pre-created gid-0 group-writable so the container owns its OWN credential config
|
||||
# dirs: tokens/settings/config are seeded in as writable copies and each CLI's runtime
|
||||
# state (backups, tasks, refreshed tokens) stays container-local, while ONLY the shared
|
||||
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
|
||||
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
|
||||
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
|
||||
# `.pi/agent` and `.grok` ARE pre-created: both are seeded per-FILE (pi:
|
||||
# auth/settings/trust/models; grok: auth.json/config.toml/pager.toml), and a
|
||||
# per-file seed copy, unlike a whole-dir one, does not create its parent directory.
|
||||
# `.dsh` is pre-created for the same per-file reason (.env/settings.yaml/
|
||||
# cordis.patch.yml), and the interactive profile is built into it HERE rather than
|
||||
# after `USER agent`: this layer's closing chgrp/chmod is what makes the whole tree
|
||||
# writable by the arbitrary uid the container actually runs as, and a profile
|
||||
# installed after it would miss that fixup. DSH_HOME points the launcher at the
|
||||
# agent's dir while this still runs as root.
|
||||
# ⚠️ `dangerouslyAllowAllBuilds` is what keeps that profile install from becoming
|
||||
# the next #352. pnpm (unlike npm) blocks dependency lifecycle scripts by default
|
||||
# and FAILS the install over it — `ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on
|
||||
# pnpm 11.24 — so any package in the tui's tree that ships one stops the build
|
||||
# dead. An allowlist of the offenders rots: `@deepseek-harness-tui/dsh-tui` is
|
||||
# resolved by dist-tag, not pinned, and 0.9.3 pulled `@google/genai` (a
|
||||
# `preinstall: no-op`) where 0.10.0-beta.x does not, so the names to allow move
|
||||
# under us between rebuilds. Allowing them wholesale is also the SAME exposure
|
||||
# this image already accepts three layers up: `npm install -g` runs the install
|
||||
# scripts of every transitive dep of the five CLIs above it, with no gate at all.
|
||||
# `.omp/agent` is pre-created for the same reason `.codex` is: it is a MIXED
|
||||
# store (per-file config seeds PLUS a shared `sessions/` RW bind mount for
|
||||
# Codeman's own host-side history/resume reads), and neither kind of artifact
|
||||
# creates its own parent directory.
|
||||
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
|
||||
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
|
||||
/home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent /home/agent/.grok \
|
||||
/home/agent/.dsh /home/agent/.omp/agent \
|
||||
&& DSH_HOME=/home/agent/.dsh HOME=/home/agent \
|
||||
dsh plugin --profile dsh-tui add --config.dangerouslyAllowAllBuilds=true \
|
||||
@deepseek-harness-tui/dsh-tui \
|
||||
&& test -f /home/agent/.dsh/profiles/dsh-tui/package.json \
|
||||
&& chgrp -R 0 /home/agent \
|
||||
&& chmod -R g=u /home/agent
|
||||
|
||||
USER agent
|
||||
WORKDIR /home/agent
|
||||
|
||||
# Codeman overrides the command with `sleep infinity` at create time; this is the
|
||||
# fallback so a hand-run container also idles rather than exiting.
|
||||
CMD ["sleep", "infinity"]
|
||||
@@ -1,132 +0,0 @@
|
||||
name: codeman
|
||||
|
||||
services:
|
||||
codeman:
|
||||
build:
|
||||
context: ..
|
||||
dockerfile: docker/server.Dockerfile
|
||||
args:
|
||||
CODEMAN_RUNTIME_USER: ${CODEMAN_RUNTIME_USER}
|
||||
PGID: ${PGID:-1000}
|
||||
PUID: ${PUID:-1000}
|
||||
image: ${CODEMAN_IMAGE}
|
||||
init: true
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "${CODEMAN_PORT}:${CODEMAN_PORT}"
|
||||
environment:
|
||||
# Tells the self-updater to restart by exiting (the restart policy below
|
||||
# relaunches it) rather than by looking for an init system that is not
|
||||
# here. Also set in the image; repeated so a container started without the
|
||||
# image default still self-identifies.
|
||||
CODEMAN_IN_CONTAINER: "1"
|
||||
# This file sets `restart: unless-stopped` below, so the updater may restart
|
||||
# the server by EXITING. Declared here and only here, never in the image: a
|
||||
# container started by plain `docker run` has no restart policy unless the
|
||||
# operator gave it one, and there the updater asks the daemon instead and
|
||||
# stages the update for a manual restart when it cannot get an answer.
|
||||
CODEMAN_RESTART_BY_EXIT: "1"
|
||||
CODEMAN_DOCKER_BRIDGE_HOOKS: ${CODEMAN_DOCKER_BRIDGE_HOOKS}
|
||||
# Host-side equivalent of the runtime user's HOME. Docker case seed,
|
||||
# credential and hook mounts are translated into the daemon namespace.
|
||||
CODEMAN_DOCKER_HOST_HOME: ${CODEMAN_APPDATA_PATH}
|
||||
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT: ${CODEMAN_DOCKER_DISABLE_SWAP_LIMIT}
|
||||
CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH}
|
||||
# Extra Host-header allowlist entries for a reverse-proxied deployment
|
||||
# (docker/README.md, "Reverse-proxy host allowlist"). Optional, so it
|
||||
# defaults to empty rather than requiring a line in every .env.
|
||||
CODEMAN_ALLOWED_HOSTS: ${CODEMAN_ALLOWED_HOSTS:-}
|
||||
CODEMAN_HOST: ${CODEMAN_HOST}
|
||||
CODEMAN_PASSWORD: ${CODEMAN_PASSWORD}
|
||||
CODEMAN_PORT: ${CODEMAN_PORT}
|
||||
CODEMAN_USERNAME: ${CODEMAN_USERNAME}
|
||||
GEMINI_API_KEY: ${GEMINI_API_KEY}
|
||||
PGID: ${PGID:-1000}
|
||||
PUID: ${PUID:-1000}
|
||||
TZ: ${TZ}
|
||||
group_add:
|
||||
# Retain access to the host Docker socket without running as root.
|
||||
- ${DOCKER_SOCKET_GID:-999}
|
||||
volumes:
|
||||
# Application data and CLI credentials persist on the configured host
|
||||
# path, rather than in a Docker-managed volume.
|
||||
- type: bind
|
||||
source: ${CODEMAN_APPDATA_PATH}
|
||||
target: /home/${CODEMAN_RUNTIME_USER}
|
||||
# Docker cases are sibling containers on the host daemon. Their workspace
|
||||
# must be visible to Codeman at the same absolute path used by that daemon.
|
||||
- type: bind
|
||||
source: ${CODEMAN_CASES_PATH}
|
||||
target: ${CODEMAN_CASES_PATH}
|
||||
# Codeman uses the host daemon to create isolated Docker cases. This is
|
||||
# Docker-outside-of-Docker, not Docker-in-Docker.
|
||||
- type: bind
|
||||
source: ${DOCKER_SOCKET}
|
||||
target: /var/run/docker.sock
|
||||
# The application source, so App Settings -> Updates can update in place.
|
||||
# This is the SAME checkout used as the build context above, mounted over
|
||||
# the image's baked copy: a `git checkout` performed inside the container
|
||||
# then lands on the host and survives the container being recreated.
|
||||
# Without it the pull would go to the container's writable layer and be
|
||||
# silently discarded by the next `up`. See docs/docker-self-update.md.
|
||||
# Defaults to `..` — the build context above — which Compose resolves
|
||||
# against the project directory, so plain `docker compose up` works with
|
||||
# no extra configuration. Set CODEMAN_REPO_PATH only to point elsewhere.
|
||||
- type: bind
|
||||
source: ${CODEMAN_REPO_PATH:-..}
|
||||
target: /opt/codeman
|
||||
# Build artefacts live in named volumes layered OVER the repo bind mount,
|
||||
# so `npm install` and `npm run build` inside the container never write
|
||||
# into the host checkout. That keeps container-compiled native modules
|
||||
# (node-pty is built from source here) out of a checkout that may also be
|
||||
# used to run Codeman natively, and keeps `git status` clean. Docker seeds
|
||||
# an EMPTY named volume from the image, so the first start inherits the
|
||||
# image's already-built node_modules and dist rather than paying for a
|
||||
# bootstrap build.
|
||||
- type: volume
|
||||
source: codeman-node-modules
|
||||
target: /opt/codeman/node_modules
|
||||
- type: volume
|
||||
source: codeman-dist
|
||||
target: /opt/codeman/dist
|
||||
extra_hosts:
|
||||
- "host.docker.internal:host-gateway"
|
||||
security_opt:
|
||||
- no-new-privileges:true
|
||||
cap_drop:
|
||||
- ALL
|
||||
cap_add:
|
||||
# The entrypoint corrects bind-mount ownership as root before dropping to
|
||||
# PUID:PGID. Everything not listed here remains dropped by cap_drop above.
|
||||
# test/docker-entrypoint.test.ts pins this list against what the
|
||||
# entrypoint and `init: true` actually need, so a capability cannot go
|
||||
# missing silently again.
|
||||
- CHOWN
|
||||
- DAC_OVERRIDE
|
||||
# `init: true` makes tini PID 1, and tini stays ROOT while the entrypoint
|
||||
# drops the server to PUID. Signalling a process of a different uid needs
|
||||
# CAP_KILL; without it tini's SIGTERM forward fails ("Unexpected error
|
||||
# when forwarding signal: 'Operation not permitted'"), tini dies, and the
|
||||
# PID namespace teardown SIGKILLs the server instead of letting
|
||||
# `server.stop()` flush state on every `docker compose down`/`restart`.
|
||||
- KILL
|
||||
- SETGID
|
||||
- SETUID
|
||||
healthcheck:
|
||||
test:
|
||||
- CMD-SHELL
|
||||
- >-
|
||||
node -e "fetch('http://127.0.0.1:${CODEMAN_PORT}/api/status').then((response) => process.exit(response.status < 500 ? 0 : 1)).catch(() => process.exit(1))"
|
||||
interval: 30s
|
||||
timeout: 5s
|
||||
retries: 3
|
||||
start_period: 30s
|
||||
|
||||
volumes:
|
||||
# Container-owned build artefacts. They persist across container recreation,
|
||||
# so an in-app update's `npm install` output is not thrown away by the next
|
||||
# `up`, and they are seeded from the image on first use. Removing them (or
|
||||
# `docker compose down -v`) is the supported reset: the next start rebuilds
|
||||
# from the image.
|
||||
codeman-node-modules:
|
||||
codeman-dist:
|
||||
@@ -1,165 +0,0 @@
|
||||
#!/bin/sh
|
||||
# Corrects ownership - host bind mounts, and the image-baked CLI prefix -
|
||||
# then drops to PUID:PGID.
|
||||
#
|
||||
# Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When
|
||||
# either path does not exist yet - a first run, a cleared application-data
|
||||
# directory, a restored backup - the Docker daemon creates it owned by root,
|
||||
# and an unprivileged server cannot then create its own state directory. The
|
||||
# result is a container that restarts forever on:
|
||||
#
|
||||
# Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman'
|
||||
#
|
||||
# Running this as root and dropping afterwards removes that failure mode without
|
||||
# leaving the server privileged. The same root start also lets it re-assert
|
||||
# /opt/codeman-cli's ownership on every start, not just at image build time -
|
||||
# see the comment at that chown below for why that matters for anyone who
|
||||
# runs the compose file directly rather than through Start-Codeman.sh.
|
||||
#
|
||||
# Capabilities this script needs against the compose file's `cap_drop: ALL`
|
||||
# (test/docker-entrypoint.test.ts pins the list against docker-compose.yaml):
|
||||
# CHOWN + DAC_OVERRIDE the chown of a root-owned bind source below
|
||||
# SETUID + SETGID the setpriv drop itself
|
||||
# KILL NOT used here, but required by the container: with
|
||||
# `init: true` tini is PID 1 and runs as root while the
|
||||
# server runs as PUID, and signalling a process of a
|
||||
# different uid needs CAP_KILL. Without it every
|
||||
# `docker compose down`/`restart` ends in tini dying with
|
||||
# "Unexpected error when forwarding signal" and the
|
||||
# server being SIGKILLed instead of stopping cleanly.
|
||||
|
||||
set -eu
|
||||
|
||||
# Honour an explicit `user:` in Compose: when the container was not started as
|
||||
# root there is nothing to correct and no privilege to drop.
|
||||
if [ "$(id -u)" -ne 0 ]; then
|
||||
exec "$@"
|
||||
fi
|
||||
|
||||
# Everything below runs as root and calls stat, chown, id, setpriv and friends
|
||||
# by bare name, so the lookup path must not contain a directory the runtime
|
||||
# account can write to. /opt/codeman-cli/bin is exactly that (it is chowned to
|
||||
# PUID:PGID so sessions can update the agent CLIs in place), and the image
|
||||
# appends it to PATH for the server's sake. Resolve root's commands through the
|
||||
# system directories only, and hand the image's full PATH back to the server at
|
||||
# the exec below, since Codeman resolves the agent CLIs through it.
|
||||
runtime_path=$PATH
|
||||
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
|
||||
export PATH
|
||||
|
||||
: "${PUID:=1000}"
|
||||
: "${PGID:=1000}"
|
||||
|
||||
# The capabilities the compose file must grant, named in the diagnosis below so
|
||||
# an out-of-tree compose file (Unraid's Compose Manager, a hand-written unit)
|
||||
# fails with a one-line fix instead of a restart loop.
|
||||
required_caps='CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID'
|
||||
|
||||
# Pre-flight the drop itself before touching anything. A container started with
|
||||
# `cap_drop: ALL` and none of the additions above fails here, and would otherwise
|
||||
# die at the final exec with a bare "setpriv: setresuid failed: Operation not
|
||||
# permitted" after chown had already failed, or worse, misreport a perfectly
|
||||
# writable directory as unwritable because the probe below could not drop
|
||||
# privileges to test it.
|
||||
if ! setpriv --reuid "$PUID" --regid "$PGID" --clear-groups true 2>/dev/null; then
|
||||
printf 'entrypoint: cannot drop privileges to PUID:PGID (%s:%s).\n' "$PUID" "$PGID" >&2
|
||||
printf 'entrypoint: this image starts as root and drops with setpriv, which needs\n' >&2
|
||||
printf 'entrypoint: cap_add: [%s]\n' "$required_caps" >&2
|
||||
printf 'entrypoint: on top of cap_drop: ALL (see docker/docker-compose.yaml). Add them to the\n' >&2
|
||||
printf 'entrypoint: compose file that started this container, or set `user:` to skip the drop entirely.\n' >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Preserve the supplementary groups Compose granted through group_add - that is
|
||||
# how the Docker socket stays reachable - while discarding root's own group.
|
||||
supplementary=$(id -G | tr ' ' '\n' | grep -vx 0 | paste -sd, -)
|
||||
[ -n "$supplementary" ] || supplementary="$PGID"
|
||||
|
||||
# Writable as the account the server is about to become? A real probe, run as
|
||||
# exactly the identity the final exec below produces (PUID, PGID, the same
|
||||
# supplementary groups, capabilities dropped), rather than a comparison of
|
||||
# owners: ownership is not writability. A group-writable tree owned by another
|
||||
# account, an ACL, or a CIFS/NFS mount that reports some unrelated uid are all
|
||||
# fine to run on and would all fail an owner check.
|
||||
writable_as_runtime() {
|
||||
setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" test -w "$1" 2>/dev/null
|
||||
}
|
||||
|
||||
for target in "${HOME:-}" "${CODEMAN_CASES_PATH:-}"; do
|
||||
[ -n "$target" ] && [ -d "$target" ] || continue
|
||||
owner=$(stat -c '%u:%g' "$target")
|
||||
[ "$owner" = "${PUID}:${PGID}" ] && continue
|
||||
|
||||
# Only ever correct a directory the DAEMON created: root-owned, because
|
||||
# neither PUID nor PGID existed yet when it materialised the missing bind
|
||||
# source. Anything else - a host tree that legitimately belongs to some
|
||||
# OTHER account, such as an existing CODEMAN_CASES_PATH the README already
|
||||
# allows pointing at a normal project directory - is not this container's
|
||||
# to reassign; recursively chowning it on every mismatch silently rewrote
|
||||
# a credentials tree or a projects directory to PUID:PGID with one log
|
||||
# line to explain it. Such a directory is left alone and only PROBED below.
|
||||
#
|
||||
# The chown is deliberately not fatal. A bind mount backed by NFS, CIFS or a
|
||||
# rootless daemon can refuse chown while still being perfectly writable, and
|
||||
# the probe below is what decides whether the server can run on it.
|
||||
if [ "${owner%%:*}" = '0' ]; then
|
||||
if chown -R "${PUID}:${PGID}" "$target" 2>/dev/null; then
|
||||
printf 'entrypoint: corrected ownership of %s to %s:%s\n' "$target" "$PUID" "$PGID"
|
||||
else
|
||||
printf 'entrypoint: warning: cannot change ownership of %s to %s:%s; checking whether it is writable anyway\n' \
|
||||
"$target" "$PUID" "$PGID" >&2
|
||||
fi
|
||||
fi
|
||||
|
||||
if writable_as_runtime "$target"; then
|
||||
if [ "${owner%%:*}" != '0' ]; then
|
||||
printf 'entrypoint: %s is owned by %s, not %s:%s, but is writable as the runtime account; leaving its ownership alone\n' \
|
||||
"$target" "$owner" "$PUID" "$PGID"
|
||||
fi
|
||||
continue
|
||||
fi
|
||||
|
||||
printf 'entrypoint: %s is not writable as PUID:PGID (%s:%s); it is owned by %s.\n' \
|
||||
"$target" "$PUID" "$PGID" "$owner" >&2
|
||||
printf 'entrypoint: refusing to change ownership of a directory this container did not create.\n' >&2
|
||||
printf 'entrypoint: either chown it on the host, make it writable to %s:%s, or set PUID/PGID to match its owner.\n' \
|
||||
"$PUID" "$PGID" >&2
|
||||
exit 1
|
||||
done
|
||||
|
||||
# /opt/codeman-cli (the four agent CLIs) is chowned to PUID:PGID once, at
|
||||
# image BUILD time, from the PUID/PGID build args - server.Dockerfile's own
|
||||
# comment on that RUN step explains why it lives in its own prefix rather than
|
||||
# /usr/local. Unlike HOME/CODEMAN_CASES_PATH above, that bake happens only
|
||||
# when the image is actually rebuilt (`docker compose up --build`, which
|
||||
# Start-Codeman.sh always does) - a deployment that instead runs the compose
|
||||
# file directly (Unraid's Compose Manager, a native Debian systemd unit, any
|
||||
# `docker compose up`/`restart` with no --build) can change PUID/PGID in .env
|
||||
# and restart without ever rebuilding, at which point the container runs as
|
||||
# the NEW uid while the CLI directory is still owned by the OLD one baked into
|
||||
# the image layer - silently breaking the very "self-update a CLI in place"
|
||||
# fix this directory exists for. Re-assert it here, every start, unconditionally:
|
||||
# unlike the host bind mounts above, this is pure image content Codeman itself
|
||||
# populated, never host data that might legitimately belong to someone else,
|
||||
# so there is no ownership to be careful about - it is always correct for it
|
||||
# to be owned by whoever this container is about to run as.
|
||||
if [ -d /opt/codeman-cli ] && [ "$(stat -c '%u:%g' /opt/codeman-cli)" != "${PUID}:${PGID}" ]; then
|
||||
chown -R "${PUID}:${PGID}" /opt/codeman-cli
|
||||
fi
|
||||
|
||||
# Discarding group 0 is right for root's own group, but it also discards a
|
||||
# `group_add: 0` that was there to reach a Docker socket owned by root:root.
|
||||
# The previous image ran as PUID with that group kept, so say so rather than
|
||||
# letting Docker-case support vanish silently on such a host.
|
||||
if [ -S /var/run/docker.sock ] && [ "$(stat -c '%g' /var/run/docker.sock)" = '0' ]; then
|
||||
printf 'entrypoint: warning: /var/run/docker.sock is owned by group 0, which is dropped along with root;\n' >&2
|
||||
printf 'entrypoint: warning: Docker cases will not work from this container. Give the socket a dedicated\n' >&2
|
||||
printf 'entrypoint: warning: group on the host and set DOCKER_SOCKET_GID to it.\n' >&2
|
||||
fi
|
||||
|
||||
# No `--bounding-set -all` here: it is a silent no-op without CAP_SETPCAP, which
|
||||
# the compose file deliberately does not grant, and `no-new-privileges` already
|
||||
# makes the bounding set moot. The reuid/regid drop leaves CapPrm/CapEff empty.
|
||||
# The image's full PATH goes back to the server here; see the top of the file.
|
||||
exec setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" \
|
||||
env PATH="$runtime_path" "$@"
|
||||
@@ -1,34 +0,0 @@
|
||||
#!/bin/sh
|
||||
# Git credential helper for Azure DevOps, backed by the signed-in Azure CLI.
|
||||
#
|
||||
# Configured in the image's system gitconfig for https://dev.azure.com and
|
||||
# https://*.visualstudio.com (see server.Dockerfile). On `get` it answers with
|
||||
# an Entra ID access token for the Azure DevOps resource as the password, the
|
||||
# same token type Git Credential Manager uses for Azure Repos. It never prompts:
|
||||
# when `az` is not signed in it prints nothing, so git fails fast with its own
|
||||
# authentication error instead of hanging a request that has no terminal.
|
||||
#
|
||||
# AZURE_DEVOPS_EXT_PAT, the azure-devops extension's own PAT variable, is used
|
||||
# instead when it is set, for accounts that authenticate with a PAT.
|
||||
|
||||
# `store` and `erase` are no-ops: the token belongs to az, which refreshes it.
|
||||
[ "$1" = "get" ] || exit 0
|
||||
|
||||
# Drain the request git writes on stdin; the host scoping is in gitconfig.
|
||||
cat >/dev/null
|
||||
|
||||
if [ -n "${AZURE_DEVOPS_EXT_PAT:-}" ]; then
|
||||
printf 'username=pat\npassword=%s\n' "$AZURE_DEVOPS_EXT_PAT"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
command -v az >/dev/null 2>&1 || exit 0
|
||||
|
||||
# 499b84ac-1321-427f-aa17-267ca6975798 is the fixed application ID of Azure
|
||||
# DevOps: https://learn.microsoft.com/azure/devops/integrate/get-started/authentication/service-principal-managed-identity
|
||||
token="$(az account get-access-token \
|
||||
--resource 499b84ac-1321-427f-aa17-267ca6975798 \
|
||||
--query accessToken --output tsv 2>/dev/null)" || exit 0
|
||||
[ -n "$token" ] || exit 0
|
||||
|
||||
printf 'username=azure-cli\npassword=%s\n' "$token"
|
||||
@@ -1,304 +0,0 @@
|
||||
# syntax=docker/dockerfile:1
|
||||
|
||||
# Build the application from the checkout supplied as the Docker build context.
|
||||
# No published Codeman application image is required.
|
||||
FROM node:22-bookworm-slim AS build
|
||||
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends python3 make g++ \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
WORKDIR /opt/codeman
|
||||
|
||||
COPY . .
|
||||
|
||||
# devDependencies are deliberately KEPT (no `npm prune --omit=dev`). The in-app
|
||||
# updater rebuilds from inside this container, and `npm run build` is tsc +
|
||||
# esbuild — both devDependencies. Pruning them saves image size and takes the
|
||||
# self-updater with it. See docs/docker-self-update.md.
|
||||
RUN npm ci \
|
||||
&& npm run build \
|
||||
&& npm cache clean --force
|
||||
|
||||
# The Docker CLI talks to the host daemon through the socket mounted by
|
||||
# docker/docker-compose.yaml. It does not run a Docker daemon in this container.
|
||||
FROM node:22-bookworm-slim
|
||||
|
||||
ARG CODEMAN_RUNTIME_USER=codeman
|
||||
ARG PUID=1000
|
||||
ARG PGID=1000
|
||||
|
||||
# python3/make/g++ are here for the SELF-UPDATER, not for this build. An update
|
||||
# runs `npm install` inside the running container, and node-pty ships no Linux
|
||||
# prebuild, so a release that bumps it compiles from source right here. Without
|
||||
# a toolchain that install fails and the update rolls back — every time, on the
|
||||
# releases that need it most. Same reason install.sh installs one on bare hosts.
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends \
|
||||
ca-certificates \
|
||||
curl \
|
||||
g++ \
|
||||
git \
|
||||
libsecret-1-0 \
|
||||
make \
|
||||
openssh-client \
|
||||
procps \
|
||||
python3 \
|
||||
ripgrep \
|
||||
tmux \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# The Docker CLI, taken from the official image rather than Debian's `docker.io`.
|
||||
# That package is the full ENGINE: with --no-install-recommends it still pulls 15
|
||||
# packages including containerd, runc, dmsetup and iptables, none of which a
|
||||
# client that only talks to a mounted socket can use. Measured on top of this
|
||||
# base image: `docker.io` costs 266 MB and ships Docker 20.10.24 (2023), while
|
||||
# these two files cost 108 MB and ship the current CLI (493 MB vs 335 MB total).
|
||||
#
|
||||
# The binaries are STATIC Go builds, so they run on this glibc image even though
|
||||
# the image they come from is Alpine (verified: `docker --version`, `docker ps`
|
||||
# and `docker build` all work here against a mounted host socket).
|
||||
#
|
||||
# buildx is copied on purpose. `scripts/build-agent-image.mjs` shells out to
|
||||
# `docker build` — Codeman auto-builds the agent image on the first Docker case —
|
||||
# and without the plugin that silently falls back to the CLASSIC builder, which
|
||||
# Docker has deprecated and will eventually drop. `docker-compose` is NOT copied:
|
||||
# Codeman never shells out to it.
|
||||
COPY --from=docker:29-cli /usr/local/bin/docker /usr/local/bin/docker
|
||||
COPY --from=docker:29-cli \
|
||||
/usr/local/libexec/docker/cli-plugins/docker-buildx \
|
||||
/usr/local/libexec/docker/cli-plugins/docker-buildx
|
||||
|
||||
# GitHub CLI and Azure CLI (with the azure-devops extension), so a user can sign
|
||||
# this container in to GitHub and Azure DevOps from a Codeman shell session and
|
||||
# then clone PRIVATE repositories, both from that session and through Add Case
|
||||
# -> Clone Repo. Codeman still collects no Git credentials itself: the clone
|
||||
# path (src/git-clone.ts) only inherits HOME and git's config, so whatever the
|
||||
# user signs in to here is what authenticates, and nothing when they have not
|
||||
# (the clone then fails fast with AUTH_REQUIRED, exactly as before).
|
||||
#
|
||||
# Each is OPT-IN and OFF by default: the image is functionally unchanged
|
||||
# unless the build gets CODEMAN_INSTALL_GH=1 and/or CODEMAN_INSTALL_AZ=1, which
|
||||
# a deployment sets under `build: args:` in docker-compose.override.yml
|
||||
# (docker/README.md, "Private repositories"). Off installs no apt repository,
|
||||
# package, extension or credential-helper entry; all that remains is the
|
||||
# AZURE_EXTENSION_DIR variable, its empty directory and one layer that copies
|
||||
# and then removes the helper script. The Azure CLI is the heavy one (~600 MB,
|
||||
# mostly its bundled Python). The base docker-compose.yaml
|
||||
# and .env deliberately do not carry them: turning a CLI on is a per-host
|
||||
# choice, which is what the override file is for, and a new .env.example key
|
||||
# would make the self-updater refuse existing installs until their .env gained
|
||||
# it (docs/docker-self-update.md).
|
||||
#
|
||||
# Both come from their vendors' own apt repositories, the same ones the
|
||||
# documented one-liners configure (https://github.com/cli/cli/blob/trunk/docs/install_linux.md
|
||||
# and https://learn.microsoft.com/cli/azure/install-azure-cli-linux?pivots=apt).
|
||||
# Microsoft's `deb_install.sh` is deliberately not piped into the build: it does
|
||||
# exactly this plus a `gnupg` install, and a remote script run at build time is
|
||||
# the one step a reviewer cannot read in this file. apt reads an ASCII-armoured
|
||||
# `.asc` key directly, which is what keeps `gnupg` out of the image.
|
||||
#
|
||||
# Not pinned, unlike the agent CLIs below: nothing in Codeman depends on a
|
||||
# particular gh or az behaviour, so the pinning argument there does not apply.
|
||||
# The layer cache still keeps whatever version the first build fetched until a
|
||||
# --no-cache rebuild.
|
||||
ARG CODEMAN_INSTALL_GH=0
|
||||
ARG CODEMAN_INSTALL_AZ=0
|
||||
RUN set -eux; \
|
||||
for flag in "CODEMAN_INSTALL_GH=${CODEMAN_INSTALL_GH}" "CODEMAN_INSTALL_AZ=${CODEMAN_INSTALL_AZ}"; do \
|
||||
case "${flag#*=}" in 0|1) ;; *) echo "${flag%%=*} must be 0 or 1, got '${flag#*=}'" >&2; exit 1;; esac; \
|
||||
done; \
|
||||
codename="$(. /etc/os-release && echo "${VERSION_CODENAME}")"; \
|
||||
arch="$(dpkg --print-architecture)"; \
|
||||
pkgs=""; \
|
||||
install -d -m 0755 /etc/apt/keyrings; \
|
||||
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then \
|
||||
curl -fsSL -o /etc/apt/keyrings/githubcli-archive-keyring.gpg \
|
||||
https://cli.github.com/packages/githubcli-archive-keyring.gpg; \
|
||||
chmod go+r /etc/apt/keyrings/githubcli-archive-keyring.gpg; \
|
||||
echo "deb [arch=${arch} signed-by=/etc/apt/keyrings/githubcli-archive-keyring.gpg] https://cli.github.com/packages stable main" \
|
||||
> /etc/apt/sources.list.d/github-cli.list; \
|
||||
pkgs="${pkgs} gh"; \
|
||||
fi; \
|
||||
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
|
||||
curl -fsSL -o /etc/apt/keyrings/microsoft.asc \
|
||||
https://packages.microsoft.com/keys/microsoft.asc; \
|
||||
chmod go+r /etc/apt/keyrings/microsoft.asc; \
|
||||
echo "deb [arch=${arch} signed-by=/etc/apt/keyrings/microsoft.asc] https://packages.microsoft.com/repos/azure-cli/ ${codename} main" \
|
||||
> /etc/apt/sources.list.d/azure-cli.list; \
|
||||
pkgs="${pkgs} azure-cli"; \
|
||||
fi; \
|
||||
if [ -n "${pkgs}" ]; then \
|
||||
apt-get update; \
|
||||
apt-get install -y --no-install-recommends ${pkgs}; \
|
||||
rm -rf /var/lib/apt/lists/*; \
|
||||
fi
|
||||
|
||||
# The azure-devops extension goes into a SYSTEM directory rather than the
|
||||
# default ~/.azure/cliextensions: HOME is the application-data bind mount, which
|
||||
# hides anything installed there at build time. The directory is handed to the
|
||||
# runtime account below (next to /opt/codeman-cli) so `az extension update`
|
||||
# works from a session. Nothing that runs as root executes from it. It is
|
||||
# created even without az, so the chown below does not have to know.
|
||||
ENV AZURE_EXTENSION_DIR=/opt/codeman-az-extensions
|
||||
RUN set -eux; \
|
||||
install -d -m 0755 "${AZURE_EXTENSION_DIR}"; \
|
||||
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
|
||||
az extension add --name azure-devops --only-show-errors; \
|
||||
rm -rf /root/.azure; \
|
||||
fi
|
||||
|
||||
# Git credential helpers, in the SYSTEM gitconfig so they apply to every
|
||||
# account and survive a fresh application-data directory. Each one answers only
|
||||
# for its own host and prints nothing when its CLI is not signed in, so git
|
||||
# falls through to its normal non-interactive failure. Only an installed CLI
|
||||
# gets an entry: a helper naming a missing binary would print an error on every
|
||||
# clone from that host.
|
||||
# github.com `gh auth git-credential`, what `gh auth setup-git` configures.
|
||||
# Azure DevOps an Entra ID token from `az login` (git-credential-azure-cli),
|
||||
# for both dev.azure.com and the legacy *.visualstudio.com hosts.
|
||||
COPY docker/git-credential-azure-cli /usr/local/bin/git-credential-azure-cli
|
||||
RUN set -eux; \
|
||||
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then \
|
||||
for host in https://github.com https://gist.github.com; do \
|
||||
git config --system "credential.${host}.helper" '!/usr/bin/gh auth git-credential'; \
|
||||
done; \
|
||||
fi; \
|
||||
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
|
||||
chmod 0755 /usr/local/bin/git-credential-azure-cli; \
|
||||
for host in https://dev.azure.com 'https://*.visualstudio.com'; do \
|
||||
git config --system "credential.${host}.helper" /usr/local/bin/git-credential-azure-cli; \
|
||||
git config --system "credential.${host}.useHttpPath" true; \
|
||||
done; \
|
||||
else \
|
||||
rm -f /usr/local/bin/git-credential-azure-cli; \
|
||||
fi
|
||||
|
||||
# Keep credentials out of the image. Users authenticate these CLIs at runtime
|
||||
# through Codeman sessions, and the configured host bind mount retains state.
|
||||
#
|
||||
# Installed into a DEDICATED prefix, /opt/codeman-cli, not the base image's
|
||||
# default /usr/local. A session needs write access to wherever these CLIs live
|
||||
# so it can self-update one in place (observed via Codex's own
|
||||
# `npm install -g @openai/codex`, which renames the old package directory
|
||||
# aside before installing the new one — a rename needs write access to the
|
||||
# PARENT directory, not just the target, so the runtime account needs that
|
||||
# access at the directory level). Chowning /usr/local/bin and
|
||||
# /usr/local/lib/node_modules directly to get it would ALSO hand away
|
||||
# entrypoint.sh (COPY'd to /usr/local/bin below, root-owned, executed as root
|
||||
# on every container start with CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node
|
||||
# binary: owning the DIRECTORY is enough to rename it aside and drop a
|
||||
# replacement, even though the file itself stays root-owned, which would let a
|
||||
# compromised session arrange for its own script to run as root at the next
|
||||
# restart — undoing the "the server itself never runs privileged" guarantee
|
||||
# the entrypoint exists to provide. /opt/codeman-cli holds nothing else to
|
||||
# escalate through, so owning it is exactly the CLI-update access it needs and
|
||||
# no more.
|
||||
#
|
||||
# ⚠️ PINNED ON PURPOSE. Unpinned, the agent CLI versions a user ends up with are
|
||||
# a function of WHEN their image was built, not of any commit — so a Codeman
|
||||
# release that depends on newer CLI behaviour (the trust-dialog handling is
|
||||
# pinned to Claude Code 2.1.252's layout; wheel forwarding to >= 2.1.187) breaks
|
||||
# on an older image with no diff anywhere to explain why. In-app updates make
|
||||
# rebuilds RARER, which makes that drift worse. Pinning turns "this release needs
|
||||
# a newer CLI" into a Dockerfile change, which the updater's environment gate
|
||||
# already detects and refuses (docs/docker-self-update.md).
|
||||
#
|
||||
# Bump these deliberately, in a release. `--no-cache` is still needed to rebuild
|
||||
# this layer when only the pins change upstream.
|
||||
# The prefix is APPENDED to PATH, never prepended: it is chowned to the runtime
|
||||
# account below, and entrypoint.sh runs as root calling stat/chown/setpriv by
|
||||
# bare name. A prefix ahead of /usr/bin would let a session drop a `setpriv`
|
||||
# there and have it run as root at the next container start (measured with a
|
||||
# minimal image of this exact shape). The four CLIs live only in this prefix,
|
||||
# so they still resolve; entrypoint.sh additionally pins its own PATH to the
|
||||
# system directories for the root part of the start.
|
||||
# uv/uvx: MCP servers are commonly launched with `uvx <package>` (e.g. the Nginx
|
||||
# Proxy Manager MCP), and Codex failed to enable them with "uvx not found". Copied
|
||||
# from the pinned upstream image into root-owned /usr/local/bin, never pip-installed.
|
||||
COPY --from=ghcr.io/astral-sh/uv:0.9 /uv /uvx /usr/local/bin/
|
||||
ENV NPM_CONFIG_PREFIX=/opt/codeman-cli
|
||||
ENV PATH=$PATH:/opt/codeman-cli/bin
|
||||
# pnpm is not an agent CLI: it is here because `dsh plugin` (DeepSeek Harness, which
|
||||
# this image leaves to be installed at runtime, see SERVER_INTENTIONAL_OMISSIONS in
|
||||
# test/docker-agent-image-coverage.test.ts) spawns a literal `pnpm` with no npm
|
||||
# fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed
|
||||
# with `dsh: pnpm not found on PATH` (exit 127) on this image. The agent image
|
||||
# already carries it for the same reason (#352). It lives in the same
|
||||
# runtime-writable prefix as the CLIs, so a session can update it in place.
|
||||
RUN npm install --global \
|
||||
@anthropic-ai/claude-code@2.1.258 \
|
||||
@google/gemini-cli@0.58.0 \
|
||||
@openai/codex@0.152.1 \
|
||||
opencode-ai@1.18.26 \
|
||||
pnpm@12.6.0 \
|
||||
&& npm cache clean --force
|
||||
|
||||
# Keep the web server and every local Codeman session unprivileged. PUID and
|
||||
# PGID match the host-owned application-data directory mounted by Compose. The
|
||||
# requested GID may not exist in the base image, and a host UID such as 1000 may
|
||||
# already belong to the baked `node` account, so handle both cases explicitly.
|
||||
#
|
||||
# The trailing chown hands the CLI prefix (/opt/codeman-cli, populated above)
|
||||
# to that same account, so a session can self-update one of the CLIs in place.
|
||||
# /usr/local stays root-owned throughout — see the comment on the npm install
|
||||
# above for why that boundary matters.
|
||||
RUN set -eux; \
|
||||
case "${PUID}" in ''|*[!0-9]*) echo "PUID must be numeric" >&2; exit 1;; esac; \
|
||||
case "${PGID}" in ''|*[!0-9]*) echo "PGID must be numeric" >&2; exit 1;; esac; \
|
||||
if [ "${PUID}" -eq 0 ]; then \
|
||||
echo "PUID must identify an unprivileged account, not root" >&2; \
|
||||
exit 1; \
|
||||
fi; \
|
||||
if ! getent group "${PGID}" >/dev/null; then \
|
||||
groupadd --gid "${PGID}" codeman-runtime; \
|
||||
fi; \
|
||||
existing_user="$(getent passwd "${PUID}" | cut -d: -f1 || true)"; \
|
||||
if [ -n "${existing_user}" ]; then \
|
||||
usermod \
|
||||
--login "${CODEMAN_RUNTIME_USER}" \
|
||||
--gid "${PGID}" \
|
||||
--home "/home/${CODEMAN_RUNTIME_USER}" \
|
||||
--move-home \
|
||||
--shell /bin/bash \
|
||||
"${existing_user}"; \
|
||||
else \
|
||||
useradd \
|
||||
--uid "${PUID}" \
|
||||
--gid "${PGID}" \
|
||||
--create-home \
|
||||
--home-dir "/home/${CODEMAN_RUNTIME_USER}" \
|
||||
--shell /bin/bash \
|
||||
"${CODEMAN_RUNTIME_USER}"; \
|
||||
fi; \
|
||||
chown -R "${PUID}:${PGID}" /opt/codeman-cli /opt/codeman-az-extensions
|
||||
|
||||
WORKDIR /opt/codeman
|
||||
|
||||
COPY --from=build /opt/codeman /opt/codeman
|
||||
|
||||
# CODEMAN_IN_CONTAINER tells the self-updater it must restart by exiting rather
|
||||
# than by asking an init system that is not here (src/web/self-update.ts).
|
||||
# NODE_ENV stays `production`; the updater passes `npm install --include=dev`
|
||||
# explicitly, since that value would otherwise omit the build toolchain.
|
||||
ENV CODEMAN_IN_CONTAINER=1 \
|
||||
CODEMAN_PORT=3000 \
|
||||
HOME=/home/${CODEMAN_RUNTIME_USER} \
|
||||
NODE_ENV=production
|
||||
|
||||
# Runtime defaults for the entrypoint, matching the account created above.
|
||||
ENV PGID=${PGID} PUID=${PUID}
|
||||
|
||||
EXPOSE 3000
|
||||
|
||||
# The container starts as root so the entrypoint can correct the ownership of
|
||||
# the host bind mounts, which the daemon creates as root whenever they do not
|
||||
# already exist. The entrypoint then drops to PUID:PGID with setpriv, so the
|
||||
# server itself never runs privileged. Setting `user:` in Compose bypasses both
|
||||
# steps, leaving the caller in full control.
|
||||
COPY docker/entrypoint.sh /usr/local/bin/entrypoint.sh
|
||||
RUN chmod 0755 /usr/local/bin/entrypoint.sh
|
||||
|
||||
ENTRYPOINT ["/usr/local/bin/entrypoint.sh"]
|
||||
|
||||
CMD ["node", "dist/index.js", "web"]
|
||||
@@ -1,104 +0,0 @@
|
||||
# SPEEDRUN.md — Fast-execution protocol for Claude
|
||||
|
||||
Read this when the goal is **throughput**: get correct, verified work done with
|
||||
minimum ceremony. This does **not** relax correctness or the safety rules in
|
||||
`CLAUDE.md` — those still win. It removes _waste_, not _rigor_.
|
||||
|
||||
> Precedence: `CLAUDE.md` > explicit user instructions > this file. If anything
|
||||
> here conflicts with `CLAUDE.md`, `CLAUDE.md` wins.
|
||||
|
||||
---
|
||||
|
||||
## The mindset
|
||||
|
||||
- **Act, don't announce.** No "I'm going to now…" preamble. Do the thing, report
|
||||
the result.
|
||||
- **Cheapest proof that the change works.** Pick the smallest check that actually
|
||||
demonstrates correctness — not the biggest.
|
||||
- **Batch aggressively.** Independent reads, greps, and edits go in **one**
|
||||
message with parallel tool calls. Never serialize work that has no dependency.
|
||||
- **Momentum over perfection.** Land a correct increment, verify it, move on.
|
||||
Don't gold-plate untouched code.
|
||||
|
||||
---
|
||||
|
||||
## Loop (repeat until done)
|
||||
|
||||
1. **Orient once** — one parallel burst of reads/greps to load the context you
|
||||
need. Don't re-read files the harness says are already current.
|
||||
2. **Change** — make the edit(s). Batch independent edits.
|
||||
3. **Verify cheaply** — the smallest check that proves _this_ change (see below).
|
||||
4. **Advance** — next item. Only re-verify what you touched.
|
||||
5. **Stop** at: list empty, a hard blocker, or a decision that's genuinely the
|
||||
user's to make.
|
||||
|
||||
---
|
||||
|
||||
## Verification ladder — climb only as high as the change needs
|
||||
|
||||
| Change kind | Cheapest sufficient check |
|
||||
|-------------|---------------------------|
|
||||
| Types / signatures / imports | `tsc --noEmit` (or `--watch` already running) |
|
||||
| One module's logic | `npm test -- test/<file>.test.ts` (the **one** relevant file) |
|
||||
| A named behavior | `npm test -- -t "pattern"` |
|
||||
| Route/handler | `app.inject()` route test, or one `curl` against the running dev server |
|
||||
| Frontend render | Playwright load + assert (`waitUntil: 'domcontentloaded'`, wait 3–4s) |
|
||||
| Broad / pre-merge | `npm run test:ci` (the CI-equivalent sweep) |
|
||||
|
||||
**Hard rules (never skip, even in a rush):**
|
||||
- ⚠️ **Never run bare `npm test`** — it pulls in browser/visual suites that hang
|
||||
or fail locally. Always pass a file or `-t`, or use `test:ci`.
|
||||
- ⚠️ **Never COM without verifying the change actually works** first (curl the
|
||||
endpoint / Playwright the UI). "Compiles" ≠ "works".
|
||||
- ⚠️ **Session safety** — check `$CODEMAN_MUX`; never `tmux kill-session` /
|
||||
`pkill claude` in a managed session.
|
||||
- ⚠️ **Single-line prompts** for any programmatic session input.
|
||||
|
||||
---
|
||||
|
||||
## Speed tactics that pay off here
|
||||
|
||||
- **Parallel exploration**: dispatch `Explore` subagents (or one parallel grep
|
||||
burst) instead of serial file-by-file reading when scope is uncertain.
|
||||
- **`tsc --noEmit --watch`** in the background — instant type feedback, no repeat
|
||||
cold starts.
|
||||
- **Target one test file** — `fileParallelism: false` means the suite is serial;
|
||||
running one file is dramatically faster than the sweep.
|
||||
- **`curl localhost:3000/api/...`** beats spinning up a browser for backend
|
||||
checks. Reserve Playwright for actual UI rendering.
|
||||
- **Trust the harness** — if it says a file you just edited is current, don't
|
||||
re-Read it to "confirm". The Edit already succeeded or it would have errored.
|
||||
|
||||
---
|
||||
|
||||
## Anti-patterns (these masquerade as speed, but cost time)
|
||||
|
||||
- Running the full test suite to check a one-file change.
|
||||
- Re-reading files you already have in context.
|
||||
- Narrating a plan you're about to execute anyway.
|
||||
- Serial tool calls that have no dependency between them.
|
||||
- Claiming "done / fixed / passing" **before** running the check that proves it.
|
||||
- Deploying (COM) on green typecheck alone, without exercising the real flow.
|
||||
|
||||
---
|
||||
|
||||
## Stop-conditions (don't rush past these)
|
||||
|
||||
Stop and surface, don't guess, when you hit:
|
||||
- A **destructive / hard-to-reverse** action (delete, overwrite, force-push).
|
||||
- An **outward-facing** action (publishing, sending, deploying) not already
|
||||
authorized.
|
||||
- A **genuine product decision** the code can't answer.
|
||||
- A **failing verification you can't explain** — debug it (see
|
||||
`superpowers:systematic-debugging`), don't paper over it.
|
||||
|
||||
---
|
||||
|
||||
## Definition of done
|
||||
|
||||
A task is done when **all** hold:
|
||||
- The change is made.
|
||||
- The cheapest sufficient check **ran** and **passed** — evidence, not assertion.
|
||||
- No new type errors / lint errors introduced (`tsc --noEmit`, `npm run lint`).
|
||||
- You state plainly what was done and what proved it. If a step was skipped or a
|
||||
test failed, say so — don't hedge, don't overclaim.
|
||||
@@ -1,769 +0,0 @@
|
||||
# Agent Control Plan: skill packaging + wait primitives
|
||||
|
||||
**Status**: steps 1 to 8 DONE and RELEASED. The wait primitives and the skill itself
|
||||
(steps 1 to 5) shipped in **1.13.0**; the `codeman skill install` CLI, per-case injection
|
||||
and `agentSkillEnabled` (step 6) shipped in **1.14.1** and were republished with fixes in
|
||||
**1.14.2**. Steps 1 to 5 were multi-round verified on 2026-08-08, step 6 on 2026-08-09;
|
||||
see [§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
|
||||
verification round found, and the two items that genuinely remain open (§2.4's footgun
|
||||
guard and the Part 3 deferrals).
|
||||
|
||||
**Date**: 2026-08-08
|
||||
**Scope**: Part 1 (agent skill) and Part 2 (wait primitives) were specified and built.
|
||||
Parts 3 to 5 are captured so they are not lost, but remain deliberately deferred.
|
||||
|
||||
---
|
||||
|
||||
## 0. Where this came from: what herdr does
|
||||
|
||||
[herdr](https://github.com/herdrdev/herdr) (Rust, Apache-2.0, ~25.8k stars) is a terminal
|
||||
multiplexer built around AI coding agents. Relevant findings from the research pass:
|
||||
|
||||
| Capability | How herdr does it |
|
||||
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Agent state | Four states (`idle`, `working`, `blocked`, `done`) that roll up pane to tab to workspace in a sidebar |
|
||||
| Detection | Lifecycle hooks where the agent supports them (it names Pi and MastraCode), otherwise TOML manifests matched against a live bottom-buffer snapshot. Bundled manifests plus remote updates from herdr.dev, local overrides win |
|
||||
| Control API | Newline-delimited JSON over a Unix socket (`~/.config/herdr/sessions/<name>/herdr.sock`), `{"id":"req_1","method":"pane.split","params":{}}`, dot-notation methods, plus long-lived event subscriptions |
|
||||
| Discoverability | `herdr api schema` prints a machine-readable schema |
|
||||
| Agent skill | `npx skills add herdrdev/herdr --skill herdr -g`, a SKILL.md wrapping the CLI, guarded by `test "${HERDR_ENV:-}" = 1` so an agent outside a herdr pane refuses to act |
|
||||
| Persistence | Background server, detach with `ctrl+b q`, snapshot restore of workspaces/tabs/panes/cwd/layout, experimental screen-history replay, agent resume via native session ids, live PTY handoff across server replacement |
|
||||
| Plugins | `herdr-plugin.toml` manifest, actions, event hooks, plugin panes, link handlers, GitHub-topic marketplace index |
|
||||
|
||||
The commands the skill teaches the agent:
|
||||
|
||||
| Group | Commands |
|
||||
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| workspace | `workspace list`, `workspace create` |
|
||||
| tab | `tab list --workspace <id>`, `tab create` |
|
||||
| pane | `pane current`, `pane list`, `pane layout`, `pane split --current --direction right --cwd <path> --no-focus`, `pane run <id> "<cmd>"`, `pane wait-output <id> --match/--regex <p> --timeout <ms>`, `pane read <id> --source visible\|recent\|detection` |
|
||||
| agent | `agent list`, `agent start <name> --kind <type> --pane <id>`, `agent prompt <name> "<text>" --wait --timeout <ms>`, `agent wait <name> --until <state> --timeout <ms>`, `agent send-keys`, `agent get`, `agent read` |
|
||||
|
||||
### The honest comparison
|
||||
|
||||
herdr and Codeman are not the same product. herdr is a local, keyboard-first multiplexer with
|
||||
no server, no web UI, and no autonomy layer. Codeman is a server with a browser and mobile UI,
|
||||
remote and Docker cases, respawn, Ralph, cron, and the orchestrator, none of which herdr has.
|
||||
|
||||
What herdr genuinely does better is being **callable by the agent running inside it**. For
|
||||
Codeman that is a packaging problem plus one missing primitive, not an architecture problem.
|
||||
|
||||
---
|
||||
|
||||
## 1. Gap analysis
|
||||
|
||||
| herdr capability | Codeman equivalent today | Gap |
|
||||
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- |
|
||||
| `pane split` + `agent start` | `POST /api/quick-start`, `POST /api/sessions` | none, already there |
|
||||
| `agent prompt` | `POST /api/sessions/:id/input` with `clientId`+`seq` exactly-once | no `--wait` |
|
||||
| `pane read` | `GET /api/sessions/:id/output`, `GET /api/sessions/:id/terminal?full=1` | none |
|
||||
| `agent list` / `agent get` | `GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status` | none |
|
||||
| `agent wait --until <state>` | SSE only (`/api/events`) | **missing**, and SSE is impractical from a shell tool |
|
||||
| `pane wait-output --match` | nothing | **missing** |
|
||||
| Skill file | README section "Driving Codeman from an Agent" | **not packaged**, an agent will never find it |
|
||||
| Env guard `HERDR_ENV=1` | `CODEMAN_MUX=1`, `CODEMAN_API_URL`, `CODEMAN_SESSION_ID` already exported at spawn | none, the guard variables exist |
|
||||
| `blocked` state | hook events (`permission_prompt`, `elicitation_dialog`) plus CSS classes plus the phone overview NEEDS YOU section | not in the wire contract (`SessionStatus = 'idle' \| 'busy' \| 'stopped' \| 'error'`) |
|
||||
| `api schema` | hand-written `docs/api-reference.md` | no machine-readable schema |
|
||||
| Detection manifests | hardcoded in `usage-limit-patterns.ts`, `respawn-*-patterns`, `regex-patterns.ts` | patterns are code, not data |
|
||||
| Plugin runtime | deliberately refused, see `docs/extending-codeman.md` | not a gap, a decision |
|
||||
| Session handoff on restart | tmux owns the PTYs, so they already survive a Codeman restart | not a gap, solved by architecture |
|
||||
|
||||
**Conclusion**: roughly 90% of the capability surface already exists. Parts 1 and 2 below close
|
||||
the two real gaps.
|
||||
|
||||
The table is the 2026-08-08 snapshot that motivated the work, kept as written. The three rows
|
||||
marked missing are closed since: `GET .../wait` and `GET .../wait-output` shipped in 1.13.0, and
|
||||
the skill is packaged at `skills/codeman` (npm tarball included). `blocked` as a wire-contract
|
||||
state, and the machine-readable schema, are still open (Parts 3 and 4).
|
||||
|
||||
---
|
||||
|
||||
## 2. Part 1: the Codeman agent skill
|
||||
|
||||
### 2.1 Goal
|
||||
|
||||
An agent running inside a Codeman session can discover and correctly drive Codeman without the
|
||||
user pasting API docs into the prompt, and without inventing dangerous calls.
|
||||
|
||||
### 2.2 Layout and distribution
|
||||
|
||||
The `npx skills` CLI (vercel-labs/skills) clones a GitHub repo and looks for
|
||||
`skills/<name>/SKILL.md`. Claude Code natively discovers `.claude/skills/<name>/SKILL.md` in a
|
||||
project and `~/.claude/skills/` globally. Both are satisfied with one source of truth plus a
|
||||
symlink, which is the pattern this repo already uses for `remotion-best-practices`.
|
||||
|
||||
```
|
||||
skills/
|
||||
codeman/
|
||||
SKILL.md <- single source of truth
|
||||
reference/
|
||||
endpoints.md <- full endpoint tables, loaded on demand
|
||||
recipes.md <- worked multi-session orchestration examples
|
||||
.claude/skills/codeman -> ../../skills/codeman (symlink, dogfooding in this repo)
|
||||
```
|
||||
|
||||
Adding a `skills/` directory to the repo root costs one entry in the GitHub listing. CLAUDE.md
|
||||
keeps the root short on purpose, so this needs a conscious sign-off; the alternative is
|
||||
`docs/skills/codeman/` with a `--skill` path argument, which breaks the one-liner install.
|
||||
**Recommendation**: accept `skills/` at the root, because the install one-liner is the whole
|
||||
point of shipping a skill.
|
||||
|
||||
Install paths, in order of how a user gets it:
|
||||
|
||||
1. `npx skills add Ark0N/Codeman --skill codeman -g` (global, any agent, matches the herdr flow).
|
||||
2. `codeman skill install [--global | --case <name>]`, a new CLI subcommand writing the same
|
||||
file. This is the path for users who installed via npm and never cloned the repo.
|
||||
3. **Automatic per-case injection**, modeled exactly on `applyStatusLineConfig(casePath, enabled)`
|
||||
in `hooks-config.ts`: write `<case>/.claude/skills/codeman/SKILL.md` at case creation,
|
||||
gated on a new setting. Codeman already writes `<case>/.claude/settings.local.json` hooks
|
||||
through `writeHooksConfig()`, so this is the same mechanism with the same lifecycle.
|
||||
|
||||
Setting name: `agentSkillEnabled`. Synced (not per-device), since it changes on-disk case
|
||||
content rather than display. Default: **ON after the dogfooding phase, OFF in the first
|
||||
release**. Rationale for starting OFF: Claude Code loads every skill's name and description
|
||||
into context on every turn, so an always-on skill has a small permanent token cost, and we
|
||||
should measure that we are buying something with it first.
|
||||
|
||||
### 2.3 SKILL.md content
|
||||
|
||||
Frontmatter, per the skills convention (`name` + `description` required):
|
||||
|
||||
```yaml
|
||||
---
|
||||
name: codeman
|
||||
description: >-
|
||||
Control Codeman, the session manager this agent is running inside: list sessions,
|
||||
start worker sessions, send prompts, read terminal output, and wait for other agents
|
||||
to finish. Only usable when CODEMAN_MUX=1.
|
||||
---
|
||||
```
|
||||
|
||||
Body sections, in order:
|
||||
|
||||
**1. Guard (first thing, non-negotiable).**
|
||||
|
||||
```bash
|
||||
test "${CODEMAN_MUX:-}" = 1 || { echo "not inside a Codeman session"; exit 1; }
|
||||
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set, refusing to guess}"
|
||||
SELF="${CODEMAN_SESSION_ID:-}"
|
||||
```
|
||||
|
||||
If `CODEMAN_MUX` is not `1`, the agent must stop and say it is not running inside a
|
||||
Codeman-managed session. Same shape as herdr's `HERDR_ENV` guard, and the variables are
|
||||
already exported by `tmux-manager.buildEnvExports()`. No fallback URL when
|
||||
`CODEMAN_API_URL` is unset: any guess is the wrong scheme on an HTTPS install (prod is
|
||||
HTTPS with a self-signed cert, hence `curl -sk` throughout), and a server the agent
|
||||
cannot identify is not one it should be driving.
|
||||
|
||||
**2. Rules of the road.** Lifted and tightened from README lines 666 to 745:
|
||||
|
||||
- Single-line input only. Multi-line breaks the agent TUI (Ink).
|
||||
- Always send `clientId` + a monotonic `seq` on `POST .../input` so a retry cannot double-deliver.
|
||||
- Envelope is `{success, data}`; a few legacy GETs are bare, so read `body.data ?? body`.
|
||||
- Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set. Prod is HTTPS, so `curl -sk`.
|
||||
- Prefer `/api/v1/*`, the stable alias.
|
||||
|
||||
**3. Safety rules (the section that does not exist anywhere today).**
|
||||
|
||||
- Never act on `$CODEMAN_SESSION_ID`. That is you.
|
||||
- Only `DELETE` sessions **you created in this conversation**, by exact id. Keep the list.
|
||||
- Never bulk-delete, never loop a `DELETE` over `/api/sessions`. There is no undo.
|
||||
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. Use the API.
|
||||
- Creating a session consumes a slot against the 50-session cap. Clean up what you start.
|
||||
|
||||
**4. Recipes**, each one a single copy-pasteable curl:
|
||||
|
||||
| Task | Call |
|
||||
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| list sessions | `GET /api/v1/sessions` |
|
||||
| find yourself | match ids by PREFIX of `$CODEMAN_SESSION_ID` (Docker cases truncate it to 8 chars, so an equality check never fires there) |
|
||||
| start a worker | `POST /api/v1/quick-start {caseName, mode, effort}` |
|
||||
| send a prompt | `POST /api/v1/sessions/:id/input {input:"…\r", useMux:true, clientId, seq}` (the trailing `\r` is what sends Enter; without it the text sits on the prompt unsubmitted) |
|
||||
| send prompt and wait | `POST /api/v1/sessions/:id/input {input:"…\r", wait:"stop", waitTimeout:600000}` (Part 2) |
|
||||
| wait for a worker | `GET /api/v1/sessions/:id/wait?until=stop,blocked&timeout=300000` (Part 2) |
|
||||
| wait for a marker | `GET /api/v1/sessions/:id/wait-output?match=DONE_<random>&timeout=120000` (Part 2; unique per call, per §3.3's repaint rule) |
|
||||
| read output | `GET /api/v1/sessions/:id/output` |
|
||||
| read full scrollback | `GET /api/v1/sessions/:id/terminal?full=1` |
|
||||
| watch sub-agents | `GET /api/v1/subagents` |
|
||||
| schedule work | `POST /api/v1/cron/jobs` |
|
||||
| clean up | `DELETE /api/v1/sessions/:id` |
|
||||
|
||||
**5. Pointer to `reference/endpoints.md`** for anything not in the table, so the always-loaded
|
||||
part of the skill stays small.
|
||||
|
||||
### 2.4 An ergonomics guard worth adding server-side
|
||||
|
||||
The skill will tell the agent not to act on itself, but a confused agent can still try. Propose:
|
||||
the skill sends `X-Codeman-Caller-Session: $CODEMAN_SESSION_ID` on every request, and the server
|
||||
refuses destructive operations (`DELETE /api/sessions/:id`, kill, respawn stop) when that header
|
||||
equals the target id, with a clear error.
|
||||
|
||||
This is a **footgun guard, not a security control**: any caller can omit the header. Document it
|
||||
as such so nobody mistakes it for a boundary. It costs about 10 lines in `route-helpers.ts`.
|
||||
|
||||
### 2.5 Verification
|
||||
|
||||
Per the always-end-to-end-test rule, "the skill exists" is not done. Done is:
|
||||
|
||||
1. Symlink it into `.claude/skills/`, start a real throwaway Codeman session, and ask that agent
|
||||
to "start a worker session that runs the test suite and tell me when it finishes".
|
||||
2. Confirm from the outside that exactly one new session appeared, got the prompt, and that the
|
||||
lead agent waited rather than polling in a busy loop.
|
||||
3. Confirm the guard: run the same prompt in a shell with `CODEMAN_MUX` unset and confirm refusal.
|
||||
4. Confirm cleanup: the worker session is deleted by exact id and no other session was touched.
|
||||
|
||||
Never run this against `w1`/`w2`/`w3`.
|
||||
|
||||
### 2.6 Files touched
|
||||
|
||||
- `skills/codeman/SKILL.md` (new), `skills/codeman/reference/*.md` (new)
|
||||
- `.claude/skills/codeman` symlink (new)
|
||||
- `src/cli.ts` (new `skill install` subcommand)
|
||||
- `src/hooks-config.ts` (new `applyAgentSkill(casePath, enabled)`, mirroring `applyStatusLineConfig`)
|
||||
- `src/web/schemas.ts` (`agentSkillEnabled` in `SettingsUpdateSchema`, which is `.strict()`)
|
||||
- `src/web/routes/system-routes.ts` (settings PUT must resolve the flag from `merged`, never
|
||||
from the raw body, per the partial-PUT invariant)
|
||||
- `src/web/public/settings-ui.js` + `index.html` (checkbox)
|
||||
- `package.json` `files` array, so `skills/` ships to npm
|
||||
- README pointer, `docs/extending-codeman.md` seam 3 pointer
|
||||
|
||||
---
|
||||
|
||||
## 3. Part 2: wait primitives
|
||||
|
||||
### 3.1 Goal
|
||||
|
||||
Make Codeman orchestratable from a shell tool. Today the only "tell me when" channel is SSE,
|
||||
which a curl-driven agent cannot practically consume: it would have to hold a streaming
|
||||
connection and parse events inline. herdr solves this with blocking CLI calls. Codeman should
|
||||
solve it with bounded long-poll endpoints.
|
||||
|
||||
All three additions are **additive**, so the versioning policy stays intact (new endpoints and
|
||||
new optional fields are non-breaking).
|
||||
|
||||
### 3.2 The signal model
|
||||
|
||||
A waiter resolves on the first of a set of signals. Sources that already exist:
|
||||
|
||||
| Signal | Source today |
|
||||
| --------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `idle` | `Session` emits `idle` (session.ts ~1775 for Claude, ~2101 for shell), wired at `session-listener-wiring.ts:402` |
|
||||
| `working` | `Session` emits `working` (session.ts ~1788), wired at `session-listener-wiring.ts:401` |
|
||||
| `stop` | `POST /api/hook-event` with `event: 'stop'`, the definitive "Claude finished responding" signal already used by `controller.signalStopHook()` |
|
||||
| `blocked` | `POST /api/hook-event` with `permission_prompt` or `elicitation_dialog` |
|
||||
| `exit` | `Session` emits `exit` |
|
||||
|
||||
`stop` is the highest-quality signal for "the turn is over" and should be the documented default
|
||||
for orchestration. `idle` is heuristic: output stabilization plus prompt detection, and it can
|
||||
flap mid-turn when a spinner pauses. External CLI modes (`isExternalCliMode()`) have no stop
|
||||
hook at all, so for opencode/codex/gemini/antigravity only `idle`, `working` and `exit` are
|
||||
available. **The skill and the docs must say which signals exist per mode**, otherwise an agent
|
||||
waits forever on `stop` in a codex session.
|
||||
|
||||
### 3.3 Endpoint specs
|
||||
|
||||
#### A. `GET /api/sessions/:id/wait`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
| --------- | ---------------------------------------------- | ---------------- | ------------------------------------------------------------ |
|
||||
| `until` | comma list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on first match |
|
||||
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` (600000) |
|
||||
| `fresh` | `0`/`1` | `0` | `1` requires a _transition_, ignoring the state at call time |
|
||||
|
||||
Response (always 200 unless the session is missing or a cap is hit):
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"data": {
|
||||
"signal": "stop",
|
||||
"timedOut": false,
|
||||
"immediate": false,
|
||||
"ended": false,
|
||||
"waitedMs": 8421,
|
||||
"status": "idle",
|
||||
"sessionId": "...",
|
||||
"until": ["stop", "idle", "exit"],
|
||||
"limitPaused": false
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`until` is echoed back because the server may narrow it: `stop`/`blocked` are dropped
|
||||
from the DEFAULT set for external CLI modes (asking for them EXPLICITLY is a 400
|
||||
instead, since omitting `until` must never 400). `limitPaused` tells a caller that a
|
||||
timeout was expected rather than a stall worth retrying hard.
|
||||
|
||||
**A timeout is not an error.** `{"timedOut": true, "signal": null}` with HTTP 200, so a caller
|
||||
can loop without treating every poll boundary as a failure. Errors are reserved for
|
||||
`NOT_FOUND` (unknown or not-owned session) and `SESSION_BUSY` (waiter cap exceeded).
|
||||
|
||||
`immediate: true` means the session was already in the requested state and `fresh` was not set.
|
||||
|
||||
#### B. `GET /api/sessions/:id/wait-output`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
| --------- | ------------------------------ | -------- | --------------------------------------------------------- |
|
||||
| `match` | literal string, 1 to 200 chars | required | substring match against ANSI-stripped output |
|
||||
| `nocase` | `0`/`1` | `0` | case-insensitive compare |
|
||||
| `from` | `now` \| `buffer` | `now` | `buffer` scans the existing text buffer first, then waits |
|
||||
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` |
|
||||
|
||||
Response: `{ matched: true, timedOut: false, snippet: "...", waitedMs }`.
|
||||
|
||||
**No regex in v1, deliberately.** `search-service.ts` already avoids regex specifically so there
|
||||
is no ReDoS surface, and this endpoint would be even more exposed since the pattern is attacker
|
||||
supplied and the input is a live stream. herdr can offer `--regex` because Rust's regex crate is
|
||||
linear-time with no backtracking; JS `RegExp` is not. If regex is wanted later, the honest
|
||||
options are a length-capped subset compiled once with a match budget, or `re2`. Note it and move on.
|
||||
|
||||
Implementation detail that will bite if missed: a match can straddle two PTY chunks. Keep a
|
||||
carry buffer of `match.length - 1` bytes from the previous chunk and test `carry + chunk`.
|
||||
|
||||
⚠️ **`from=now` does not mean "printed after you asked".** tmux repaints the visible
|
||||
screen on attach, resize, or any TUI redraw, and a repaint arrives as ordinary `terminal`
|
||||
data. Observed live: a marker echoed a minute earlier matched instantly on a fresh
|
||||
`from=now` wait. This is inherent to a terminal multiplexer, not fixable in the registry,
|
||||
so the contract is: **use a marker unique per call** (`echo DONE_$RANDOM`), never a
|
||||
generic one like `BUILD OK`. The skill's recipes must show that.
|
||||
|
||||
The returned snippet is whitespace-collapsed (blank runs to a single newline) for
|
||||
readability only; matching runs on the raw stripped text. Without it, a real pane's
|
||||
`\r\n` padding between the prompt and the match fills the whole context window with
|
||||
nothing, which was the first thing the live test showed.
|
||||
|
||||
#### C. `wait` on the existing input endpoint
|
||||
|
||||
`POST /api/sessions/:id/input` gains two optional fields:
|
||||
|
||||
```json
|
||||
{ "input": "run the tests\r", "useMux": true, "clientId": "agent-1", "seq": 7, "wait": "stop", "waitTimeout": 600000 }
|
||||
```
|
||||
|
||||
(The trailing `\r` is required on every input body: `sendInput` sends Enter only
|
||||
when the input contains a carriage return.)
|
||||
|
||||
Response gains `"wait": { "signal": "stop", "timedOut": false, "waitedMs": 41230 }`.
|
||||
|
||||
This is the important one, because it closes a race the standalone `GET .../wait` cannot: between
|
||||
"input delivered" and "session flips to working" there is a window where a naive
|
||||
send-then-wait sees the _pre-existing_ idle state and returns instantly. The combined endpoint
|
||||
**registers the waiter before writing**, so that window does not exist. This is exactly why herdr
|
||||
ships `agent prompt --wait` as its own thing.
|
||||
|
||||
`wait` accepts `true` (the default signal set) or the same comma grammar as `until`.
|
||||
Both new fields are `.nullish()`, not `.optional()`: a third-party caller building the
|
||||
body with `JSON.stringify` keeps an explicit `null` on the wire, and `.optional()`
|
||||
rejects that with `INVALID_INPUT`. That gotcha has shipped as a real bug twice.
|
||||
|
||||
Two behaviors to preserve carefully:
|
||||
|
||||
- **`useMux` is fire-and-forget today.** The handler responds without awaiting `writeViaMux`, on
|
||||
purpose (a tmux child process must not block the HTTP response). With `wait` present the
|
||||
handler already has to stay open, so it can await delivery, and a `writeViaMux` failure becomes
|
||||
observable for the first time. The non-wait path must keep its current fire-and-forget shape
|
||||
byte for byte.
|
||||
- **Duplicate suppression.** A tagged redelivery (`clientId`+`seq` already applied) returns 200
|
||||
without writing. With `wait` set it still waits, since the caller's intent is "tell me when
|
||||
this settles". But it waits with `requireTransition: false`, unlike a fresh delivery: the
|
||||
original turn may be long over, and requiring a new transition would block a redelivery until
|
||||
timeout for no reason. Fresh delivery requires a transition, a duplicate answers from the
|
||||
current state.
|
||||
- **Capacity rollback.** `shouldApplyInput()` MUTATES (it records the seq), and it runs before
|
||||
the waiter is registered. If registration then fails on a full pool, the handler must call
|
||||
`forgetInputSeq` before returning `SESSION_BUSY`, or the caller's retry is rejected as a
|
||||
duplicate and the input is lost by the very mechanism reliable delivery exists for.
|
||||
|
||||
### 3.4 Module design
|
||||
|
||||
New file `src/web/session-wait-registry.ts`, with the IO-free core unit-testable in isolation
|
||||
(same split as `self-update.ts`):
|
||||
|
||||
```ts
|
||||
type WaitSignal = 'idle' | 'working' | 'stop' | 'blocked' | 'exit';
|
||||
|
||||
waitForSignal(sessionId, { until: Set<WaitSignal>, timeoutMs, requireTransition }): Promise<WaitResult>
|
||||
notifySignal(sessionId, signal: WaitSignal): void
|
||||
waitForOutput(sessionId, { match, nocase, timeoutMs }): Promise<OutputWaitResult>
|
||||
notifyOutput(sessionId, chunk: string): void
|
||||
cancelAll(sessionId, reason): void
|
||||
```
|
||||
|
||||
Wiring points, all existing:
|
||||
|
||||
- `src/web/session-listener-wiring.ts` around lines 190 and 200 already handles `working` and
|
||||
`idle` and broadcasts them. Add a `notifySignal()` call next to each broadcast, plus `exit`.
|
||||
- `src/web/routes/hook-event-routes.ts` already switches on `event` for the respawn controller.
|
||||
Add `notifySignal(sessionId, 'stop' | 'blocked')` in the same switch.
|
||||
- Output: `notifyOutput()` rides the ALREADY-attached `terminal` listener in
|
||||
session-listener-wiring.ts. An earlier draft had the registry hand out attach/detach
|
||||
callbacks so a listener could be added lazily; that was deleted once it was clear no
|
||||
second listener is needed at all. The cost is one Map lookup per PTY chunk, which is why
|
||||
the no-waiter check comes before the ANSI strip.
|
||||
- Session deletion calls `notifySignal('exit')` then `cancelAll()`, so no promise is left
|
||||
hanging. Both are required: `_doCleanupSession` detaches the session's listeners BEFORE
|
||||
`session.stop()`, so on a delete the PTY exit event never reaches the registry, and an
|
||||
`until=exit` caller would otherwise get a bare `ended` instead of its signal. Found by
|
||||
live-testing the delete path, not by the unit tests.
|
||||
|
||||
Memory-leak discipline, per the 24-hour-session rules: every waiter owns a timer that is cleared
|
||||
on resolve, the per-session waiter set is deleted when it empties, and the output listener is
|
||||
removed with it. `test/memory-leak-prevention.test.ts` should grow a case for this.
|
||||
|
||||
Caps in a new `src/config/agent-wait.ts` (limits live in `src/config/`, env-overridable):
|
||||
|
||||
| Constant | Default | Why |
|
||||
| ------------------------- | ------- | --------------------------------------- |
|
||||
| `MAX_WAIT_MS` | 600000 | an unbounded long-poll is a socket leak |
|
||||
| `DEFAULT_WAIT_MS` | 60000 | short enough to survive most proxies |
|
||||
| `MAX_WAITERS_PER_SESSION` | 16 | |
|
||||
| `MAX_WAITERS_TOTAL` | 128 | same reasoning as `MAX_SSE_CLIENTS` |
|
||||
|
||||
Exceeding a cap returns `SESSION_BUSY`, not a silent queue.
|
||||
|
||||
### 3.5 Transport concerns
|
||||
|
||||
Fastify is constructed with defaults in `server.ts:329-331`. `requestTimeout` defaults to 0
|
||||
(disabled) and `keepAliveTimeout` (72s) applies between requests, not to an in-flight one, so a
|
||||
10-minute in-process hold is fine. **Verify this on the real instance before relying on it.**
|
||||
|
||||
Intermediaries are the actual risk. Prod is reached through `tailscale serve`, and users also run
|
||||
cloudflared tunnels; both can cut an idle connection. That is why `DEFAULT_WAIT_MS` is 60s and
|
||||
why the documented pattern is a client-side loop over short waits rather than one 10-minute call.
|
||||
The skill's recipes must show the loop.
|
||||
|
||||
### 3.6 Edge cases to get right
|
||||
|
||||
| Case | Behavior |
|
||||
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Session already idle, `fresh=0` | return immediately, `immediate: true` |
|
||||
| Session already idle, `fresh=1` | wait for the next transition into a requested state |
|
||||
| Session dies mid-wait | resolve with `signal: "exit"` if `exit` was requested, otherwise resolve `timedOut:false, signal:null, ended:true`. Never hang |
|
||||
| Session deleted mid-wait | same, resolve, do not throw. Verified live: `until=exit` gets `signal:"exit"`, a concurrent `until=blocked` gets `ended:true`, both in ~0ms |
|
||||
| Shutdown with a wait pending | `cancelEverything()` in `stop()`. Verified live: SIGTERM with a 300s wait in flight exits in 1s |
|
||||
| External CLI mode | `stop` and `blocked` never fire. Reject `until=stop` for those modes with a clear `INVALID_INPUT` rather than hanging until timeout |
|
||||
| Multi-user | goes through `findSessionOrFail(ctx, id, req)`, which already enforces ownership |
|
||||
| Remote / Docker cases | signals originate from the same `Session` object, so no special casing. Docker hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for `stop`/`blocked` to arrive at all; without it, only `idle` works. Document it |
|
||||
| Respawn `/clear` mid-wait | a respawn cycle emits `idle`. Callers waiting on `stop` are unaffected; callers on `idle` may resolve early. Documented, not fixed |
|
||||
| Limit pause | if the session is paused on a usage limit, nothing will fire until the reset. The wait times out honestly. Consider surfacing `limitPaused: true` in the response so the caller can back off |
|
||||
|
||||
### 3.7 Tests
|
||||
|
||||
- `test/session-wait-registry.test.ts` (pure): immediate resolve, transition-required, multi-signal
|
||||
first-wins, timeout, cap exceeded, cancel on session end, no listener leak after resolve,
|
||||
chunk-straddling output match, case-insensitive match.
|
||||
- `test/routes/session-wait-routes.test.ts` (`app.inject()`, no port): all three endpoints against
|
||||
a `MockSession`, including the 200-with-`timedOut` contract and the ownership 404.
|
||||
- `test/routes/session-input-wait.test.ts`: the send-and-wait race, plus proof that the non-wait
|
||||
path is unchanged (still returns before `writeViaMux` settles).
|
||||
- Live verification on a throwaway session before COM, per the always-end-to-end-test rule.
|
||||
|
||||
### 3.8 Files touched
|
||||
|
||||
- `src/config/agent-wait.ts` (new)
|
||||
- `src/web/session-wait-registry.ts` (new)
|
||||
- `src/web/session-listener-wiring.ts` (notify on idle/working/exit)
|
||||
- `src/web/routes/hook-event-routes.ts` (notify on stop/blocked)
|
||||
- `src/web/routes/session-routes.ts` (two new routes, `wait` fields on input)
|
||||
- `src/web/schemas.ts` (`SessionWaitQuerySchema`, `SessionWaitOutputQuerySchema`, extend
|
||||
`SessionInputWithLimitSchema`. Note: `.optional()` rejects `null`, so the frontend and any
|
||||
generated client must send `undefined`, never `null`)
|
||||
- `docs/api-reference.md`, `docs/extending-codeman.md`, README API table
|
||||
- `skills/codeman/SKILL.md` recipes (Part 1 depends on this)
|
||||
|
||||
---
|
||||
|
||||
## 4. Deferred: parts 3 to 5
|
||||
|
||||
Not in scope now, kept here so they are not lost.
|
||||
|
||||
### Part 3: promote `blocked` to a first-class state
|
||||
|
||||
`SessionStatus` is `'idle' | 'busy' | 'stopped' | 'error'`. "Needs you" exists three times over:
|
||||
hook events, the `tab-alert-action` CSS class, and the phone overview NEEDS YOU section, each
|
||||
re-deriving it. herdr makes `blocked` a real state that rolls up.
|
||||
|
||||
Add `blocked` (and possibly `done`) to `SessionStatus`, set it from the same hook events that
|
||||
Part 2 uses as wait signals, and clear it on the next `working`/`stop`. Then the tab strip, the
|
||||
mobile overview, the wait endpoints, and any external agent read one field.
|
||||
|
||||
Cost: `SessionStatus` is a widely-consumed union, so every exhaustive `switch` (the codebase has
|
||||
`assertNever` and `noFallthroughCasesInSwitch`) will need a branch. That is a feature, it makes
|
||||
the compiler find every site. This is a **minor** bump, not a patch: it widens a public type in
|
||||
the HTTP contract.
|
||||
|
||||
### Part 4: `GET /api/schema`
|
||||
|
||||
herdr ships `herdr api schema`. Every Codeman route is already Zod-validated, so
|
||||
`zod-to-json-schema` over `schemas.ts` gives a self-describing API almost free. Value: third-party
|
||||
tools and the skill stop drifting from hand-written docs. Open question: whether to emit full
|
||||
OpenAPI (`@fastify/swagger` would need per-route schema registration, which is a much larger
|
||||
change) or just dump the Zod schemas keyed by name (cheap, 80% of the value).
|
||||
|
||||
### Part 5: detection manifests instead of hardcoded patterns
|
||||
|
||||
CLI-specific readiness, blocked and usage-limit patterns live in code across
|
||||
`usage-limit-patterns.ts`, the respawn pattern helpers and `regex-patterns.ts`. Externalizing the
|
||||
per-CLI ones into data files would make adding a sixth CLI a data change instead of a code change.
|
||||
|
||||
**Do not copy the remote-update part.** herdr auto-fetches manifest updates from herdr.dev.
|
||||
Codeman auto-pulling behavioral rules from a vendor server contradicts its security posture.
|
||||
Bundled manifests plus local override only, no network.
|
||||
|
||||
### Explicit non-goals
|
||||
|
||||
- **Plugin runtime and marketplace.** `docs/extending-codeman.md` already argues this: a plugin
|
||||
runtime means third-party code inside a process that spawns agents with your credentials, on a
|
||||
server people expose over a tunnel. The reasoning still holds. If the marketplace _pattern_ is
|
||||
wanted, apply it to data (web tabs, case templates, cron recipes), never to executable code.
|
||||
- **Live PTY handoff on restart.** herdr needs it because it owns the terminals. Codeman
|
||||
delegates to tmux, so PTYs already survive a self-update restart.
|
||||
- **Socket API.** HTTP plus SSE is the existing, documented, stable contract. A second transport
|
||||
would double the surface for no capability gain.
|
||||
|
||||
---
|
||||
|
||||
## 5. Sequencing
|
||||
|
||||
| Step | Work | Gate |
|
||||
| ---- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| 1 ✅ | `src/config/agent-wait.ts` + `session-wait-registry.ts` + unit tests | 48 tests green |
|
||||
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
|
||||
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
|
||||
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
|
||||
| 5 ✅ | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
|
||||
| 6 ✅ | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | 10 unit tests (`test/agent-skill.test.ts`) + real-server case-creation tests (`test/quick-start.test.ts`, incl. the settings PUT accepting the key) green; CLI verified live (install/uninstall, global + `--case`, foreign/symlink refusals) |
|
||||
| 7 ✅ | Docs: api-reference, extending-codeman, README | plus `architecture-invariants.md` (§agent-wait-primitives), `CLAUDE.md` and the API reference's per-mode signal table |
|
||||
| 8 ✅ | COM (minor bump: new endpoints, new setting, new optional fields) | released as 1.13.0 (wait primitives + skill); step 6 followed in 1.14.1 and was republished as 1.14.2 after live-testing the packaged skill |
|
||||
|
||||
Parts 1 and 2 are independent enough to land separately, but the skill is much less useful
|
||||
without the wait endpoints, so the wait work goes first.
|
||||
|
||||
## 6. Open questions for the owner
|
||||
|
||||
1. ✅ `skills/` at the repo root: accepted (built that way; the install one-liner depends on it).
|
||||
2. ✅ `agentSkillEnabled` default: **OFF** for the first release, per §2.2's rationale (skills
|
||||
cost context on every turn; measure before defaulting on). Flip later if dogfooding earns it.
|
||||
3. ✅ Both: global install via `npx skills add` / `codeman skill install`, AND per-case
|
||||
auto-injection behind the (default-off) setting. Injection is add-only at session create and
|
||||
marker-guarded, so a user-authored copy is never touched.
|
||||
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
|
||||
and not a security boundary? (Still open, not built with step 6.)
|
||||
5. ✅ Regex support in `wait-output`: literal-only shipped, and a `regex` query param is
|
||||
rejected with a 400 rather than ignored, so an agent that assumed otherwise cannot
|
||||
silently wait on the wrong thing.
|
||||
|
||||
---
|
||||
|
||||
## 7. Build log: what actually happened
|
||||
|
||||
Written at the end of the build so the next person inherits the reasoning, not just the
|
||||
diff. Process artifacts (per-agent briefs, findings, reports) live in the gitignored
|
||||
`tmp/agent-wait-review/`; this section is the part worth keeping.
|
||||
|
||||
### What shipped
|
||||
|
||||
| Piece | Files |
|
||||
| ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Bounds + clamping | `src/config/agent-wait.ts` (new) |
|
||||
| Blocking-wait registry | `src/web/session-wait-registry.ts` (new, IO-free, unit-tested) |
|
||||
| `GET .../wait`, `GET .../wait-output`, `wait`/`waitTimeout` on `POST .../input` | `src/web/routes/session-routes.ts` |
|
||||
| Signal wiring | `session-listener-wiring.ts` (idle/working/exit + output), `hook-event-routes.ts` (stop/blocked), `server.ts` (teardown, shutdown) |
|
||||
| Agent skill | `skills/codeman/SKILL.md` + `reference/`, `.claude/skills/codeman` symlink, `package.json` `files` |
|
||||
| Docs | `api-reference.md`, `extending-codeman.md`, `architecture-invariants.md`, `README.md`, `CLAUDE.md` |
|
||||
| Tests | `test/session-wait-registry.test.ts`, three `test/routes/session-*wait*.test.ts`, `http-contract.test.ts`, `mock-session.ts` |
|
||||
|
||||
### Bugs found in ADJACENT code, not in the new feature
|
||||
|
||||
These are the highest-value output of the exercise and none were on the plan:
|
||||
|
||||
1. **Every Codeman hook was dead on HTTPS installs.** `hooks-config.ts` built the hook
|
||||
curl as `curl -s` with no `-k` while the statusline exporter 300 lines below used
|
||||
`curl -sk` and documented why. Proven with the real hook command: `curl exit=60`
|
||||
without the flag, success with it, and the failure swallowed by the hook's own
|
||||
`2>/dev/null || true`. This silently killed `stop`, `permission_prompt`,
|
||||
`elicitation_dialog`, `idle_prompt`, `teammate_idle` and `task_completed`, taking
|
||||
respawn's definitive idle signals with them. Fixed, **plus** a staleness detector in
|
||||
`refreshStaleCodemanHooks` that regenerates the on-disk config of already-created
|
||||
cases (23 of 26 local cases carried the broken form; fixing the generator alone would
|
||||
have left every one of them broken).
|
||||
2. **`buildEnvExports()` exported a wrong-scheme `CODEMAN_API_URL`** (`http://` fallback
|
||||
on an HTTPS install). Now omitted rather than guessed, so in-session guards fail closed.
|
||||
3. **Programmatic input is only submitted when it contains `\r`.** `sendInput` sends Enter
|
||||
only if the payload has a carriage return; without it the text sits in the composer
|
||||
forever. Bit this build repeatedly before it was diagnosed, and had leaked into the
|
||||
docs' own examples.
|
||||
|
||||
### Design decisions worth not re-litigating
|
||||
|
||||
- **A timeout is HTTP 200** with `wait.timedOut`, never a 4xx: callers loop over short
|
||||
waits because tunnels cut idle connections, and every poll boundary would otherwise be
|
||||
indistinguishable from failure.
|
||||
- **Send-and-wait must be one endpoint.** A separate POST-then-wait races: between the
|
||||
write and the flip to `working`, a wait sees the stale `idle` and reports the PREVIOUS
|
||||
turn as this one. The waiter is registered before the write.
|
||||
- **`stop`/`blocked` exist for `claude` mode only.** They come from Claude Code hooks;
|
||||
`shell` installs none either, so keying off `isExternalCliMode()` was wrong.
|
||||
- **Literal matching only, never regex.** JS `RegExp` backtracks; herdr can offer
|
||||
`--regex` because Rust's regex crate is linear-time.
|
||||
- **Client-hangup abort listens on `reply.raw` guarded by `writableFinished`.** On
|
||||
`req.raw`, `close` fires when the request BODY ends, which on a POST killed every
|
||||
send-and-wait instantly, and no `app.inject()` test can see it (inject never emits
|
||||
`close`).
|
||||
- **Liveness cannot come from `session.pid`.** For a tmux session that is the local
|
||||
`tmux attach` client, not the worker: a worker exiting inside its pane leaves
|
||||
`pane_dead=1` with the client alive, so `pid` never goes null. Liveness is probed at
|
||||
the mux layer, cached (~750 ms) and only on blocking waits, never on the input hot path.
|
||||
|
||||
### Verification rounds
|
||||
|
||||
Six agents across three rounds, each verifying the previous round's work rather than its
|
||||
own. Findings that mattered, in order of severity, were: the dead-pane liveness gap; the
|
||||
`reply.raw` abort regression; abandoned long-polls leaking waiter slots; a crashed session
|
||||
reporting `idle`; `shell` accepting `until=stop`; and a documented recipe that reported
|
||||
success without running its task. Two traps recurred often enough to name:
|
||||
|
||||
- **Vacuous passes.** `app.inject()` never emits `close`; a latched `cancelEverything()`
|
||||
in `afterEach` silently killed the registry for every later test in a file; three test
|
||||
files sharing one session id against the process-wide registry let one file's leftover
|
||||
waiter fail another's assertion. Any new wait test needs care on all three.
|
||||
- **HTTP-only test instances.** Every isolated instance used during the build was plain
|
||||
HTTP, which is exactly why the HTTPS hook bug survived so long. Test the transport the
|
||||
user actually runs.
|
||||
|
||||
### Resolved at wrap-up (2026-08-08, conclusion pass)
|
||||
|
||||
- **R2-A**: the fire-and-forget-then-gather-sequentially pattern was **removed from
|
||||
the skill** rather than patched. Signals are edge-triggered with no history, so a
|
||||
`stop` that fires before its waiter registers is unobservable afterwards; a
|
||||
`fresh=0` gather was rejected because the only `until` set that current state can
|
||||
satisfy answers `idle` for a prompt that never submitted, resurrecting the exact
|
||||
false-success failure R2-B had just closed. Flow 3b's pattern B now gathers on
|
||||
latched `wait-output` markers (`from=buffer`), the same mechanism that makes the
|
||||
shell flows reliable; the limitation is recorded in
|
||||
`architecture-invariants#agent-wait-primitives` and `endpoints.md`. The durable
|
||||
fix, a latched last-signal-per-turn on the server, stays with deferred Part 3.
|
||||
- Docs F7/F8, F4 and the false-`idle` attribution: `api-reference.md`,
|
||||
`extending-codeman.md` and `architecture-invariants.md` rewritten to the post-fix
|
||||
matcher (one normalized stream, chunk-straddling found, snippet as a rendering of
|
||||
the matched window), the real no-PTY answer (`ended:true`, `aborted:false`,
|
||||
`delivered:false`), and the startup-idle mechanism (a session parked on the trust
|
||||
dialog emits no further `idle`; the false success is the startup transition).
|
||||
- Orchestrate #12, #5/R2-B, #6, and R2-C..R2-E: fire-and-forget's empty `data`
|
||||
documented; every send-and-wait retry loop now treats `duplicate:true` +
|
||||
`immediate:true` as "no new turn ran" and reads the terminal before believing it;
|
||||
claude fan-out is pattern A (backgrounded send-and-waits) or the marker gather;
|
||||
readiness budgets rebalanced (5 s stage 1, 45 s stage 3) with the virgin-case
|
||||
floor named; the auth fallback now also reads the supervisor definition
|
||||
(`codeman-web.service` / launchd plist) and accepts `export`-prefixed `.env`
|
||||
lines; `pid != null` is documented as startup-only, never liveness.
|
||||
- Both public readiness recipes (extending-codeman.md, README) are bypass-first with
|
||||
the trust probe as the bounded fallback; the worked recipe carries `-k` and fails
|
||||
loudly on an empty SID; the hook `-k`/self-heal fix appears in every
|
||||
"hooks go missing" list; the multi-word-TUI claim is "unreliable", not "never".
|
||||
|
||||
### Still open
|
||||
|
||||
Both release-checklist items that used to sit here are done: `skills/` is tracked and
|
||||
ships through `package.json` `files` (published with 1.13.0, republished with 1.14.2),
|
||||
and the changeset was consumed, committed and deployed. What is left:
|
||||
|
||||
- Deferred with Part 3: the latched last-signal-per-turn. Nice-to-haves from the
|
||||
reviews: N2 (create the death-watcher inside its `try`, still built one line above
|
||||
it in `GET .../wait`) and converting timeout-shaped test detections into fast
|
||||
assertions.
|
||||
- §2.4's `X-Codeman-Caller-Session` footgun guard: still not built (open question 4).
|
||||
|
||||
### Step 6 (2026-08-09): install command, per-case injection, the setting
|
||||
|
||||
Built to the §2.6 file list, mirroring the statusLine mechanism throughout:
|
||||
|
||||
| Piece | Where |
|
||||
| ----- | ----- |
|
||||
| `applyAgentSkill(casePath, enabled)` + `installAgentSkillInto` / `removeAgentSkillFrom` | `src/hooks-config.ts` |
|
||||
| `codeman skill install` / `skill uninstall` (`--global` default, `--case <name>`) | `src/cli.ts` |
|
||||
| `agentSkillEnabled` (SYNCED, default OFF) | `schemas.ts` (`SettingsUpdateSchema`), `getAgentSkillEnabled()` on `ConfigPort`/`server.ts`, checkbox in `index.html` + `settings-ui.js` |
|
||||
| Injection call sites (Claude mode only) | `POST /api/sessions` next to `refreshStaleCodemanHooks`; `POST /api/quick-start` after the case-create/self-heal blocks (local + docker cases; remote skipped, its path lives on another host) |
|
||||
| Tests | `test/agent-skill.test.ts` (10 unit), `test/quick-start.test.ts` (real server: default-off, PUT accepts key, injection on create, shell-mode skipped) |
|
||||
|
||||
Decisions worth keeping:
|
||||
|
||||
- **Ownership marker, prefix-matched.** The injected SKILL.md ends with
|
||||
`<!-- codeman-managed-agent-skill: … -->`; install/refresh/remove all refuse a copy
|
||||
without the marker (a user's own skill) and match on the PREFIX so a wording change
|
||||
cannot disown older injected copies (the `BACKGROUND_WAKE_MARKER_PREFIX` pattern).
|
||||
- **Symlink refusal.** This repo's own dogfooding layout
|
||||
(`.claude/skills/codeman -> ../../skills/codeman`) means the injector must `lstat`
|
||||
the skill dir AND its `skills/` parent and bail on a symlink, or enabling the
|
||||
setting in the Codeman repo itself would overwrite the skill source through the link.
|
||||
- **ADD-ONLY at session create**, same shared-`.claude` rationale as the statusLine:
|
||||
a create while the setting is off must not yank the skill out from under other live
|
||||
sessions in the repo. The remove path exists (CLI `skill uninstall`, tests); no
|
||||
automatic sweep removes on toggle-off.
|
||||
- **Removal is manifest-based, never `rm -rf`**: only files the packaged source would
|
||||
have written are deleted, directories are pruned bottom-up only if they emptied, so
|
||||
a user's extra notes in `reference/` survive an uninstall.
|
||||
- **Source resolution**: `join(moduleDir, '..', 'skills', 'codeman')` works from
|
||||
`src/` (tsx), `dist/` (tsc build), and the npm tarball alike, because all three sit
|
||||
one level below the package root and `files` ships `skills/`.
|
||||
- **Nothing acts on the setting at PUT time**: injection reads the merged persisted
|
||||
settings at session create (`readSettings`, ~2s cache), so the partial-PUT invariant
|
||||
(`toggleService` reading `merged`) is untouched by construction.
|
||||
|
||||
### 2026-08-09 addendum: cross-session messaging folded into the skill
|
||||
|
||||
Claude Code 2.1.224+ ships cross-session messaging: `ListAgents`/`SendMessage`
|
||||
tools, a per-session Unix inbox socket, and a registry in
|
||||
`~/.claude/sessions/<pid>.json`. Codeman's claude workers are ordinary local Claude
|
||||
Code sessions, so the skill now routes task delivery and result collection over it
|
||||
when available, while the HTTP primitives keep spawn, readiness, synchronization,
|
||||
liveness and delete. New `skills/codeman/reference/messaging.md` (ships with zero
|
||||
installer changes: `readAgentSkillSource()` enumerates `reference/*.md` from disk),
|
||||
Flow 5 in recipes.md, and §4 in SKILL.md.
|
||||
|
||||
Verified live (claude-cli 2.1.226, Linux):
|
||||
|
||||
- A message to an idle worker starts a turn and that turn fires the normal `stop`
|
||||
hook (8.3 s send-to-stop measured), so the HTTP wait primitives compose with
|
||||
messaging unchanged; delivery to a busy session lands between tool calls.
|
||||
- First contact needs the `name [ref]` form; the bare name errors with the exact
|
||||
string to resend. The `uds:` reply address of an inbound message works as a `to`.
|
||||
- The `tmux codeman-<id8>` column in `ListAgents` (and the registry's `tmux` field)
|
||||
is the join key to Codeman session ids. The registry's `sessionId` field starts as
|
||||
the Codeman id (we spawn `claude --session-id <id>`) but drifts after `/clear` or
|
||||
resume, so it must never be the join key.
|
||||
- The feature is flag-gated beyond the version: two 2.1.226 sessions on one machine,
|
||||
one with an inbox socket and one without. Absence is a fallback case, not an error.
|
||||
- Codeman's default `--dangerously-skip-permissions` spawn puts both ends in the
|
||||
bypassing class, which delivers; mixed classes hold behind an approval dialog that
|
||||
expires unattended (upstream default 5 min), which on a headless worker means the
|
||||
message silently dies. The skill's backstop covers it.
|
||||
|
||||
Follow-up, landed in the same PR: local claude spawns now pass
|
||||
`--name <session name>` so peers carry Codeman session names. The gate is
|
||||
`buildNameCliArgs()` (session-cli-builder.ts), fail-closed at
|
||||
`CLAUDE_NAME_FLAG_MIN_VERSION = 2.1.224`: that is the messaging release, the flag's
|
||||
presence there was verified against the installed 2.1.224 binary, and the version
|
||||
comes from `getClaudeCliVersion()` (null on probe failure and under vitest), so an
|
||||
older or unknown CLI gets a command byte-identical to before. That matters because
|
||||
claude aborts startup on an unknown option, which would kill every session spawn.
|
||||
The value is allowlist-sanitized (Unicode letters/digits plus ` ._:-`, leading
|
||||
dashes stripped so it cannot parse as another option, 64-char cap, empty result =
|
||||
flag omitted) before the double-quoted interpolation in `buildSpawnCommand`, and
|
||||
only the LOCAL command carries it: the docker/remote builders never see it, since
|
||||
their CLI is not the binary the probe measured. E2E on an isolated instance
|
||||
(`CODEMAN_INSTANCE`): process cmdline `claude ... --name w9-msgtest`, registry
|
||||
`name: "w9-msgtest"`, `ListAgents` lists it under that name, a message round-trip
|
||||
works, and its replies arrive tagged `from-name="w9-msgtest"` (a derived-name
|
||||
worker's replies carry no `from-name`). A quick-start without `sessionName` has an
|
||||
empty Codeman name, so the peer name stays derived: agents should name their
|
||||
workers. Tests: `test/name-flag-injection.test.ts`.
|
||||
|
||||
Later narrowing: `--name` is not only the peer name but also the `/resume` picker
|
||||
entry and the terminal title, and a pinned title stops Claude generating its own, so
|
||||
pinning the `w1-myapp` placeholder listed every conversation of a case under the same
|
||||
name in `/resume`. Only a manual name is pinned now (`Session.cliPinnedName`,
|
||||
`nameSource === 'manual'`, carried to the builders as `cliName`); placeholder and auto
|
||||
names leave Claude to title the conversation. A rename in Codeman appends a
|
||||
`custom-title` row to the conversation's transcript (`claude-session-title.ts`), the
|
||||
row `/rename` writes. Tests: `test/claude-resume-title.test.ts`,
|
||||
`test/routes/session-name-routes.test.ts`.
|
||||
@@ -1,244 +0,0 @@
|
||||
# Claude Code Agent Teams — Reference
|
||||
|
||||
> Experimental feature (Feb 2026). Enable per-session via env var.
|
||||
> Updated with experiment findings from 2026-02-12.
|
||||
|
||||
## Enabling
|
||||
|
||||
```bash
|
||||
# Environment variable (set before starting Claude Code)
|
||||
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
|
||||
|
||||
# In .claude/settings.local.json (case-scoped)
|
||||
{
|
||||
"env": {
|
||||
"CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1"
|
||||
}
|
||||
}
|
||||
# Note: "teammateMode" is NOT a valid settings key (validation rejects it).
|
||||
# Display mode defaults to "in-process". For tmux, pass --teammate-mode flag via CLI.
|
||||
```
|
||||
|
||||
## Filesystem Paths (Verified)
|
||||
|
||||
| Resource | Path |
|
||||
|----------|------|
|
||||
| Team config | `~/.claude/teams/{team-name}/config.json` |
|
||||
| Teammate inboxes | `~/.claude/teams/{team-name}/inboxes/{name}.json` |
|
||||
| Shared tasks | `~/.claude/tasks/{team-name}/` |
|
||||
| Teammate transcripts | `~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{id}.jsonl` |
|
||||
|
||||
Note: Teammate transcripts appear in the **standard subagent directory** under the lead's session, NOT as separate top-level sessions.
|
||||
|
||||
### config.json format (verified)
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "research-watchers",
|
||||
"description": "Team description...",
|
||||
"createdAt": 1770875105373,
|
||||
"leadAgentId": "team-lead@research-watchers",
|
||||
"leadSessionId": "461daa80-94ec-4e5e-a1bb-0518f78311bc",
|
||||
"members": [
|
||||
{
|
||||
"agentId": "team-lead@research-watchers",
|
||||
"name": "team-lead",
|
||||
"agentType": "team-lead",
|
||||
"model": "claude-opus-4-6",
|
||||
"joinedAt": 1770875105373,
|
||||
"tmuxPaneId": "",
|
||||
"cwd": "/path/to/project",
|
||||
"subscriptions": []
|
||||
},
|
||||
{
|
||||
"agentId": "fs-researcher@research-watchers",
|
||||
"name": "fs-researcher",
|
||||
"agentType": "general-purpose",
|
||||
"model": "claude-opus-4-6",
|
||||
"prompt": "Full spawn prompt...",
|
||||
"color": "blue",
|
||||
"planModeRequired": false,
|
||||
"joinedAt": 1770875126680,
|
||||
"tmuxPaneId": "in-process",
|
||||
"cwd": "/path/to/project",
|
||||
"subscriptions": [],
|
||||
"backendType": "in-process"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Key fields: `agentId` format is `{name}@{teamName}`, `leadSessionId` links to Codeman session, `backendType` indicates display mode, `color` for UI theming.
|
||||
|
||||
### Task file format (verified)
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "1",
|
||||
"subject": "Research Node.js fs.watch on Linux vs macOS",
|
||||
"description": "Full description...",
|
||||
"activeForm": "Researching Node.js fs.watch Linux vs macOS",
|
||||
"status": "in_progress",
|
||||
"blocks": [],
|
||||
"blockedBy": [],
|
||||
"owner": "fs-researcher"
|
||||
}
|
||||
```
|
||||
|
||||
Internal teammate tracking tasks have `"metadata": { "_internal": true }`.
|
||||
|
||||
Task states: `pending` → `in_progress` → `completed`. File locking via `.lock.lock` directory (mkdir-based atomic lock).
|
||||
|
||||
### Inbox message format (verified)
|
||||
|
||||
```json
|
||||
[
|
||||
{
|
||||
"from": "team-lead",
|
||||
"text": "{\"type\":\"task_assignment\",\"taskId\":\"1\",\"subject\":\"...\",\"assignedBy\":\"team-lead\",\"timestamp\":\"...\"}",
|
||||
"timestamp": "2026-02-12T05:45:18.176Z",
|
||||
"read": false
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
`text` is double-encoded JSON. Message types: `task_assignment`, `shutdown_request`, `shutdown_response`. File locking via `.json.lock` directory.
|
||||
|
||||
## Communication Model (CORRECTED)
|
||||
|
||||
**Hybrid: tool + filesystem.** The `SendMessage` tool writes to filesystem inbox files at `~/.claude/teams/{name}/inboxes/{teammate}.json`.
|
||||
|
||||
Each teammate AND the lead has an inbox JSON file. Messages are JSON arrays with `from`, `text` (double-encoded JSON), `timestamp`, `read` fields.
|
||||
|
||||
Message types observed:
|
||||
- **task_assignment**: Lead assigns task to teammate
|
||||
- **shutdown_request**: Lead asks teammate to shut down
|
||||
- **shutdown_response**: Teammate confirms shutdown
|
||||
- (Also: `message`, `broadcast`, `plan_approval_response` per docs)
|
||||
|
||||
**Implication:** We can intercept messages by watching inbox files AND potentially inject messages by writing to them (respecting `.json.lock` directory locking).
|
||||
|
||||
## Process Model (CORRECTED)
|
||||
|
||||
**Teammates are IN-PROCESS THREADS, not separate OS processes.**
|
||||
|
||||
In `in-process` mode (the default), all teammates run as threads within the single `claude` process. Only 1 claude process exists per Codeman session, regardless of team size.
|
||||
|
||||
This means:
|
||||
- No separate PIDs to track per teammate
|
||||
- All teammates share the lead's environment variables
|
||||
- Lower resource overhead than separate processes
|
||||
- Subagent transcript files still created (for progress tracking)
|
||||
|
||||
## Display Modes
|
||||
|
||||
| Mode | Trigger | UI | Requirement |
|
||||
|------|---------|-----|------------|
|
||||
| **in-process** (default) | Default | Shift+Up/Down to switch, Ctrl+T for tasks | Any terminal |
|
||||
| **tmux** | `--teammate-mode tmux` | Split panes | tmux installed |
|
||||
| **iTerm2** | Auto-detected | Native split panes | iTerm2 + `it2` CLI |
|
||||
|
||||
**For Codeman: use `in-process` only.** Codeman manages its own tmux sessions externally.
|
||||
|
||||
**In-process UI elements:**
|
||||
- Status bar: `@main @teammate1 @teammate2 ...` with `shift+↑ to expand`
|
||||
- Task list: Checkboxes with assignments `(@teammate-name)`
|
||||
- Hint: `ctrl+t to show teammates`
|
||||
|
||||
## Hooks
|
||||
|
||||
Two new hook types for quality gates (verified in settings schema):
|
||||
|
||||
### TeammateIdle
|
||||
Fires when a teammate is about to go idle.
|
||||
- Exit code 0: Allow idle (normal)
|
||||
- Exit code 2: Send feedback back, keep teammate working
|
||||
|
||||
### TaskCompleted
|
||||
Fires when a task is being marked complete.
|
||||
- Exit code 0: Allow completion
|
||||
- Exit code 2: Prevent completion, send feedback
|
||||
|
||||
These are configured in `.claude/settings.local.json` alongside existing Codeman hooks.
|
||||
|
||||
## Subagent-Watcher Compatibility (Verified)
|
||||
|
||||
**Teammates appear as standard subagents.** They create transcript files at:
|
||||
```
|
||||
~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{id}.jsonl
|
||||
```
|
||||
|
||||
Codeman's existing `subagent-watcher.ts` discovers them automatically. They appear in `/api/subagents` with status "active".
|
||||
|
||||
**Distinguishing teammates from regular subagents:**
|
||||
- Description field starts with `<teammate-message teammate_id= team`
|
||||
- Cross-reference with `~/.claude/teams/{name}/config.json` members
|
||||
|
||||
**Sub-subagents:** Teammates can spawn their own Task tool subagents, creating a 3-level hierarchy.
|
||||
|
||||
## Cleanup Behavior (Verified)
|
||||
|
||||
When the lead runs cleanup:
|
||||
1. Shutdown requests sent to all teammate inboxes
|
||||
2. Teammates shut down gracefully
|
||||
3. ALL filesystem artifacts deleted:
|
||||
- Inbox files and directory
|
||||
- Config.json
|
||||
- Team directory
|
||||
- All task files
|
||||
- Task directory
|
||||
4. Cleanup is atomic — all files removed in the same second
|
||||
|
||||
## Comparison with Subagents (Task tool)
|
||||
|
||||
| Aspect | Subagents (Task tool) | Agent Teams |
|
||||
|--------|----------------------|-------------|
|
||||
| Spawn method | Claude's built-in Task tool | Explicit team creation |
|
||||
| Process model | In-process threads | In-process threads (same!) |
|
||||
| Discovery | `subagents/agent-{id}.jsonl` only | BOTH subagent dir + `~/.claude/teams/` |
|
||||
| Communication | None (fire-and-forget) | Filesystem inboxes + SendMessage tool |
|
||||
| Shared state | None | Shared task list + inboxes |
|
||||
| Task tracking | Per-agent, no coordination | Shared with dependencies & ownership |
|
||||
| Lifecycle | Auto-cleanup on completion | Lead cleanup (deletes all artifacts) |
|
||||
| Sub-nesting | Can spawn sub-subagents | Teammates can spawn subagents too |
|
||||
| Cost | Lower (single context) | Higher (N context windows) |
|
||||
| Duration | Short-lived (seconds-minutes) | Longer-lived (minutes-hours) |
|
||||
|
||||
## Limitations
|
||||
|
||||
- No session resumption with in-process teammates (`/resume` doesn't restore them)
|
||||
- One team per session, no nested teams
|
||||
- Lead is fixed (cannot promote teammate)
|
||||
- Permissions set at spawn (change individually after)
|
||||
- Split panes require tmux or iTerm2 (not Screen)
|
||||
- Task status can lag (teammates may fail to mark complete)
|
||||
- Shutdown can be slow (waits for current tool call)
|
||||
|
||||
## Useful Commands
|
||||
|
||||
```bash
|
||||
# Check if teams exist
|
||||
ls ~/.claude/teams/
|
||||
|
||||
# Check team config
|
||||
cat ~/.claude/teams/{name}/config.json | jq .
|
||||
|
||||
# Check teammate inboxes
|
||||
cat ~/.claude/teams/{name}/inboxes/{teammate}.json | jq .
|
||||
|
||||
# Check team tasks
|
||||
ls ~/.claude/tasks/{name}/
|
||||
for f in ~/.claude/tasks/{name}/*.json; do cat "$f" | jq .; done
|
||||
|
||||
# Count Claude processes (teammates are threads, not processes)
|
||||
ps aux | grep '[c]laude' | grep -v grep
|
||||
|
||||
# Check subagent detection of teammates
|
||||
curl -s http://localhost:3000/api/subagents | jq '.data[] | select(.description | startswith("<teammate"))'
|
||||
|
||||
# Team interaction (in-process mode)
|
||||
# Shift+Up/Down: Switch between teammates
|
||||
# Enter: View teammate session
|
||||
# Escape: Interrupt teammate's turn
|
||||
# Ctrl+T: Toggle task list
|
||||
```
|
||||
@@ -1,171 +0,0 @@
|
||||
# Codeman Agent Teams Integration — Design (Approach C: Hybrid)
|
||||
|
||||
> Updated 2026-02-12 with experiment findings. See `experiment-log.md` for raw data.
|
||||
|
||||
## Overview
|
||||
|
||||
Approach C combines filesystem monitoring (for team/task discovery and inbox watching) with the existing subagent-watcher (for live transcript tailing) and adjusted idle detection (to account for active teammates). The key finding from our experiment is that **teammates already appear as standard subagents**, so most infrastructure exists — we mainly need team awareness and idle detection fixes.
|
||||
|
||||
## Components
|
||||
|
||||
### 1. TeamWatcher (`src/team-watcher.ts`)
|
||||
|
||||
Monitors `~/.claude/teams/` for team creation/removal and tracks active teams.
|
||||
|
||||
**Discovery mechanism:**
|
||||
- Poll `~/.claude/teams/` for directories (team names) every 3-5 seconds
|
||||
- When found: parse `config.json` to get:
|
||||
- `leadSessionId` → map to Codeman session
|
||||
- `members` array → teammate names, agentIds, colors, models
|
||||
- Watch for directory deletion (cleanup signal)
|
||||
|
||||
**CORRECTED from pre-experiment design:**
|
||||
- ~~Each teammate has a separate Claude Code process~~ → Teammates are **in-process threads**, not separate processes
|
||||
- ~~Find via `ps aux` + `/proc` PID matching~~ → Not needed, no separate PIDs
|
||||
- Teammate transcripts are at `subagents/agent-{id}.jsonl` (standard subagent path), NOT separate session transcripts
|
||||
|
||||
**Association:**
|
||||
- `config.json.leadSessionId` → Codeman session ID (direct match!)
|
||||
- Each member's `agentId` (e.g., `fs-researcher@research-watchers`) → links to subagent files
|
||||
- `agentType: "team-lead"` vs `"general-purpose"` distinguishes lead from teammates
|
||||
|
||||
**Inbox monitoring:**
|
||||
- Watch `~/.claude/teams/{name}/inboxes/` for new messages
|
||||
- Each teammate has a JSON file with message array
|
||||
- Messages are double-encoded JSON with `from`, `text`, `timestamp`, `read` fields
|
||||
- Message types: `task_assignment`, `shutdown_request`, `shutdown_response`
|
||||
|
||||
### 2. Team-Aware Idle Detection (HIGHEST PRIORITY)
|
||||
|
||||
**Problem (confirmed by experiment):** Lead session shows status "idle" in Codeman while teammates are actively working. Token count continues climbing but Codeman thinks the session is inactive.
|
||||
|
||||
**Solution:**
|
||||
- Before declaring a session idle, check if it's a team lead
|
||||
- If team lead: check `~/.claude/teams/*/config.json` for this session's `leadSessionId`
|
||||
- If active team exists: check task files in `~/.claude/tasks/{team-name}/`
|
||||
- Any task with `status: "in_progress"` → suppress idle detection
|
||||
- All tasks `completed` AND no non-`_internal` tasks pending → allow idle
|
||||
- Fallback: check subagent-watcher for active subagents on this session
|
||||
|
||||
**Integration points:**
|
||||
- `src/ai-idle-checker.ts` — add team-awareness check before AI idle analysis
|
||||
- `src/respawn-controller.ts` — consult TeamWatcher before transitioning to idle states
|
||||
- `src/session.ts` — expose `hasActiveTeam()` method
|
||||
|
||||
**Liveness check (simplified from pre-experiment):**
|
||||
- ~~Check `/proc/{pid}` existence~~ → Not needed (no separate processes)
|
||||
- Check task file status instead (filesystem-based)
|
||||
- Check subagent-watcher for active subagents under this session
|
||||
|
||||
### 3. Shared Task List UI
|
||||
|
||||
**Display:** New panel in web UI showing the team's shared task list.
|
||||
|
||||
**Data source:** Poll `~/.claude/tasks/{team-name}/` for task JSON files.
|
||||
|
||||
**Task file structure (verified):**
|
||||
```json
|
||||
{
|
||||
"id": "1",
|
||||
"subject": "Research Node.js fs.watch",
|
||||
"description": "Full description...",
|
||||
"activeForm": "Researching Node.js fs.watch",
|
||||
"status": "in_progress", // pending | in_progress | completed
|
||||
"blocks": [],
|
||||
"blockedBy": [],
|
||||
"owner": "fs-researcher" // Empty string = unassigned
|
||||
}
|
||||
```
|
||||
|
||||
Internal tracking tasks: `{ "metadata": { "_internal": true } }` — filter these from display.
|
||||
|
||||
**UI elements:**
|
||||
- Task subject, status badge (color-coded), owner (teammate name with color)
|
||||
- Dependency visualization (blockedBy indicators)
|
||||
- Progress bar (completed / total non-internal tasks)
|
||||
- Real-time updates via SSE
|
||||
|
||||
**API endpoint:** `GET /api/sessions/:id/team-tasks` → returns parsed task files
|
||||
|
||||
**Locking:** Respect `.lock.lock` directory lock when reading (skip if locked, retry next poll).
|
||||
|
||||
### 4. Teammate Display
|
||||
|
||||
**Decision: Option A — Enhanced subagent floating windows.**
|
||||
|
||||
Since teammates already appear as subagents in the existing infrastructure, we enhance rather than replace:
|
||||
|
||||
- **Badge:** Add "Teammate" badge to subagent windows for agents matching team config
|
||||
- **Color:** Use teammate's `color` field from config.json (blue, green, yellow)
|
||||
- **Name:** Show teammate name instead of agent ID
|
||||
- **Persistence:** Teammate windows should stay open longer (they're longer-lived than regular subagents)
|
||||
- **Status:** Show task assignment and progress from task files
|
||||
|
||||
**Detection logic:**
|
||||
```
|
||||
For each subagent detected by subagent-watcher:
|
||||
1. Check if description starts with "<teammate-message"
|
||||
2. OR cross-reference agentId with active team config members
|
||||
3. If match → apply teammate badge, color, name
|
||||
```
|
||||
|
||||
### 5. Inbox/Message Display
|
||||
|
||||
**CORRECTED: Inboxes ARE filesystem-based.**
|
||||
|
||||
Communication uses filesystem inbox files at `~/.claude/teams/{name}/inboxes/{teammate}.json`. We can:
|
||||
|
||||
1. **Watch inbox files** for real-time message monitoring
|
||||
2. **Parse message types** for display:
|
||||
- `task_assignment` → "Lead assigned Task #1 to fs-researcher"
|
||||
- `shutdown_request` → "Lead requested shutdown"
|
||||
- `shutdown_response` → "Teammate confirmed shutdown"
|
||||
3. **Display as timeline** in team panel
|
||||
|
||||
**Potential for interaction (not tested, future work):**
|
||||
- Write to teammate inbox files to inject messages
|
||||
- Must respect `.json.lock` directory locking protocol
|
||||
- Could enable "nudge" or "redirect" functionality from Codeman UI
|
||||
|
||||
## Answered Questions (from experiment)
|
||||
|
||||
| # | Question | Answer |
|
||||
|---|----------|--------|
|
||||
| 1 | Teammates in subagents dir? | **YES** — standard `subagents/agent-{id}.jsonl` path |
|
||||
| 2 | subagent-watcher detects them? | **YES** — automatically, no changes needed |
|
||||
| 3 | Task file structure? | Numbered JSON files with subject, status, owner, dependencies |
|
||||
| 4 | Env var inheritance? | **YES** — in-process threads share parent's env |
|
||||
| 5 | Processes per teammate? | **ZERO** — threads, not processes |
|
||||
| 6 | config.json format? | Rich: name, agentId, agentType, model, prompt, color, backendType |
|
||||
| 7 | Interact via stdin? | N/A (threads) — can interact via inbox files instead |
|
||||
| 8 | In-process under Screen? | Works fine — single claude process, threads handle teammates |
|
||||
| 9 | Hook events from teammates? | TeammateIdle + TaskCompleted hooks available in settings schema |
|
||||
| 10 | Process tree? | Single process with threads — no child processes |
|
||||
|
||||
## Existing Infrastructure to Leverage
|
||||
|
||||
| Component | Reuse for | Status |
|
||||
|-----------|-----------|--------|
|
||||
| `subagent-watcher.ts` | Teammate transcript tailing | **Already works** |
|
||||
| Subagent floating windows (`app.js`) | Teammate activity display | **Already works** (needs badges) |
|
||||
| `task-tracker.ts` | Background task tracking patterns | Reuse patterns |
|
||||
| LRUMap, StaleExpirationMap | Bounded caches for team state | Available |
|
||||
| SSE broadcast | Real-time UI updates | Available |
|
||||
| ~~`/proc` PID checking~~ | ~~Teammate liveness~~ | **Not needed** (threads) |
|
||||
| `file-stream-manager.ts` | Watch inbox/task files | Available |
|
||||
|
||||
## Implementation Order (Revised)
|
||||
|
||||
1. **Team-aware idle detection** — prevent premature respawn/auto-compact (CRITICAL)
|
||||
2. **TeamWatcher** — poll `~/.claude/teams/`, parse config.json, track active teams
|
||||
3. **Teammate badge in subagent windows** — mark teammate subagents with name/color
|
||||
4. **Team tasks API + UI** — `GET /api/sessions/:id/team-tasks` + task list panel
|
||||
5. **Inbox monitoring** — watch inbox files, display message timeline
|
||||
6. **TeammateIdle/TaskCompleted hooks** — add to Codeman's hooks config generator
|
||||
|
||||
## What We DON'T Need to Build
|
||||
|
||||
- ~~Process discovery for teammates~~ (they're threads)
|
||||
- ~~Custom transcript tailing~~ (subagent-watcher handles it)
|
||||
- ~~Separate teammate window infrastructure~~ (subagent windows work)
|
||||
- ~~Message interception via transcript parsing~~ (inbox files are simpler)
|
||||
@@ -1,468 +0,0 @@
|
||||
# Agent Teams Experiment Log
|
||||
|
||||
> Experiment date: 2026-02-12
|
||||
> Test case: `~/codeman-cases/agent-teams-test/`
|
||||
> Team name: `research-watchers`
|
||||
> Teammates: 3 (fs-researcher, perf-researcher, api-researcher)
|
||||
> Lead session: `461daa80-94ec-4e5e-a1bb-0518f78311bc`
|
||||
> Duration: ~3 minutes (06:45:01 → 06:48:07)
|
||||
|
||||
## Pre-Experiment State
|
||||
|
||||
```
|
||||
~/.claude/teams/ — did NOT exist
|
||||
~/.claude/tasks/ — 75 UUID-named directories (from regular Task tool subagents)
|
||||
Claude processes — 7 (including watchers)
|
||||
settings.local.json — edited to add CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
|
||||
```
|
||||
|
||||
## Experiment Prompt
|
||||
|
||||
```
|
||||
Create an agent team with 3 teammates to research the following topics in parallel:
|
||||
Teammate 1 fs-researcher researches how Node.js fs.watch works on Linux vs macOS.
|
||||
Teammate 2 perf-researcher researches inotify performance limits and alternatives.
|
||||
Teammate 3 api-researcher researches the inotifywait command-line API.
|
||||
Have each teammate write a brief summary of their findings in a separate file.
|
||||
Name the team research-watchers.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Question 1: What exact filesystem artifacts do agent teams create?
|
||||
|
||||
**Expected:** `~/.claude/teams/research-watchers/config.json` and `~/.claude/tasks/research-watchers/`
|
||||
|
||||
**Actual: CONFIRMED + SURPRISE inboxes/ directory**
|
||||
|
||||
```
|
||||
~/.claude/teams/research-watchers/
|
||||
├── config.json # Team config (members, lead, metadata)
|
||||
└── inboxes/ # Filesystem-based messaging!
|
||||
├── api-researcher.json # Per-teammate inbox
|
||||
├── fs-researcher.json
|
||||
├── perf-researcher.json
|
||||
└── team-lead.json # Lead also has an inbox
|
||||
|
||||
~/.claude/tasks/research-watchers/
|
||||
├── .lock # Empty file (presence = lock indicator?)
|
||||
├── 1.json # Task: Research Node.js fs.watch
|
||||
├── 2.json # Task: Research inotify performance
|
||||
├── 3.json # Task: Research inotifywait CLI
|
||||
├── 4.json # Internal: fs-researcher spawn tracking
|
||||
├── 5.json # Internal: perf-researcher spawn tracking
|
||||
└── 6.json # Internal: api-researcher spawn tracking
|
||||
```
|
||||
|
||||
Subagent transcripts also appear in the standard subagent directory:
|
||||
```
|
||||
~/.claude/projects/-home-arkon-codeman-cases-agent-teams-test/
|
||||
└── 461daa80.../
|
||||
├── 461daa80...jsonl # Lead session transcript
|
||||
└── subagents/
|
||||
├── agent-ae50544.jsonl # Teammate: fs-researcher
|
||||
├── agent-aa20c65.jsonl # Teammate: perf-researcher
|
||||
├── agent-a29de32.jsonl # Teammate: api-researcher
|
||||
├── agent-a04968e.jsonl # Sub-subagent (teammate's Task tool)
|
||||
├── agent-a0d372e.jsonl # Sub-subagent
|
||||
├── agent-a2ff939.jsonl # Sub-subagent
|
||||
├── agent-a89ad82.jsonl # Sub-subagent
|
||||
├── agent-aa1efc7.jsonl # Sub-subagent
|
||||
└── agent-ab0ef07.jsonl # Sub-subagent
|
||||
```
|
||||
|
||||
**Cleanup:** At 06:48:02, the lead deleted ALL artifacts — inboxes, config, tasks, the team directory itself. Clean removal.
|
||||
|
||||
---
|
||||
|
||||
## Question 2: Is the mailbox/communication filesystem-based or tool-based?
|
||||
|
||||
**Expected:** Tool-based (SendMessage tool), NOT filesystem
|
||||
|
||||
**Actual: BOTH! Hybrid — tool triggers filesystem writes.**
|
||||
|
||||
Communication uses the `SendMessage` tool internally, but the actual message delivery is via **filesystem inbox files**. Each teammate has `~/.claude/teams/{name}/inboxes/{teammate}.json` containing a JSON array of messages.
|
||||
|
||||
**Inbox message format:**
|
||||
```json
|
||||
[
|
||||
{
|
||||
"from": "team-lead",
|
||||
"text": "{\"type\":\"task_assignment\",\"taskId\":\"1\",\"subject\":\"Research Node.js fs.watch...\",\"assignedBy\":\"team-lead\",\"timestamp\":\"...\"}",
|
||||
"timestamp": "2026-02-12T05:45:18.176Z",
|
||||
"read": false
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
Key observations:
|
||||
- `text` field is a **JSON string** (double-encoded) containing a typed message object
|
||||
- Message types observed: `task_assignment`, `shutdown_request`, `shutdown_response`
|
||||
- `read` field tracks whether teammate has processed the message (false → true)
|
||||
- **File locking** via `.json.lock` directories (mkdir-based atomic lock, created then deleted)
|
||||
- Lead also has an inbox (`team-lead.json`) for receiving messages FROM teammates
|
||||
|
||||
**Implication for Codeman:** We CAN intercept messages by watching inbox JSON files! We can also potentially inject messages by writing to inbox files.
|
||||
|
||||
---
|
||||
|
||||
## Question 3: Do teammates appear in the subagents directory?
|
||||
|
||||
**Expected:** Unclear
|
||||
|
||||
**Actual: YES! Teammates appear as standard subagents.**
|
||||
|
||||
Teammates create transcript files at:
|
||||
```
|
||||
~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{agentId}.jsonl
|
||||
```
|
||||
|
||||
This is the **exact same path pattern** that regular Task tool subagents use. The existing `subagent-watcher.ts` successfully discovers them.
|
||||
|
||||
Codeman's `/api/subagents` endpoint returned them with status "active":
|
||||
```
|
||||
Agent: ae50544 Status: active Tools: 8 Model: claude-opus-4-6
|
||||
Desc: <teammate-message teammate_id= team
|
||||
Agent: aa20c65 Status: active Tools: 9 Model: claude-opus-4-6
|
||||
Desc: <teammate-message teammate_id= team
|
||||
Agent: a29de32 Status: active Tools: 7 Model: claude-opus-4-6
|
||||
Desc: <teammate-message teammate_id= team
|
||||
```
|
||||
|
||||
**Distinguishing teammates from regular subagents:**
|
||||
- Description starts with `<teammate-message teammate_id= team` (a unique marker)
|
||||
- We can also cross-reference with `~/.claude/teams/{name}/config.json` members list
|
||||
|
||||
**Sub-subagents:** Teammates can spawn their own Task tool subagents. 3 teammates spawned 6 additional subagent files (9 total in the subagents directory).
|
||||
|
||||
---
|
||||
|
||||
## Question 4: What does config.json actually look like?
|
||||
|
||||
**Actual config.json (with all 3 teammates):**
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "research-watchers",
|
||||
"description": "Research team investigating file watching mechanisms...",
|
||||
"createdAt": 1770875105373,
|
||||
"leadAgentId": "team-lead@research-watchers",
|
||||
"leadSessionId": "461daa80-94ec-4e5e-a1bb-0518f78311bc",
|
||||
"members": [
|
||||
{
|
||||
"agentId": "team-lead@research-watchers",
|
||||
"name": "team-lead",
|
||||
"agentType": "team-lead",
|
||||
"model": "claude-opus-4-6",
|
||||
"joinedAt": 1770875105373,
|
||||
"tmuxPaneId": "",
|
||||
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
|
||||
"subscriptions": []
|
||||
},
|
||||
{
|
||||
"agentId": "fs-researcher@research-watchers",
|
||||
"name": "fs-researcher",
|
||||
"agentType": "general-purpose",
|
||||
"model": "claude-opus-4-6",
|
||||
"prompt": "You are \"fs-researcher\" on the \"research-watchers\" team...",
|
||||
"color": "blue",
|
||||
"planModeRequired": false,
|
||||
"joinedAt": 1770875126680,
|
||||
"tmuxPaneId": "in-process",
|
||||
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
|
||||
"subscriptions": [],
|
||||
"backendType": "in-process"
|
||||
},
|
||||
{
|
||||
"agentId": "perf-researcher@research-watchers",
|
||||
"name": "perf-researcher",
|
||||
"agentType": "general-purpose",
|
||||
"model": "claude-opus-4-6",
|
||||
"prompt": "...",
|
||||
"color": "green",
|
||||
"planModeRequired": false,
|
||||
"joinedAt": 1770875130344,
|
||||
"tmuxPaneId": "in-process",
|
||||
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
|
||||
"subscriptions": [],
|
||||
"backendType": "in-process"
|
||||
},
|
||||
{
|
||||
"agentId": "api-researcher@research-watchers",
|
||||
"name": "api-researcher",
|
||||
"agentType": "general-purpose",
|
||||
"model": "claude-opus-4-6",
|
||||
"prompt": "...",
|
||||
"color": "yellow",
|
||||
"planModeRequired": false,
|
||||
"joinedAt": 1770875134997,
|
||||
"tmuxPaneId": "in-process",
|
||||
"cwd": "/home/arkon/codeman-cases/agent-teams-test",
|
||||
"subscriptions": [],
|
||||
"backendType": "in-process"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Key fields per member:**
|
||||
- `agentId`: `{name}@{teamName}` format
|
||||
- `agentType`: `"team-lead"` for lead, `"general-purpose"` for teammates
|
||||
- `model`: Model used (inherits from lead)
|
||||
- `prompt`: Full spawn prompt (only for teammates)
|
||||
- `color`: UI color assignment (blue, green, yellow)
|
||||
- `backendType`: `"in-process"` for in-process mode
|
||||
- `tmuxPaneId`: `"in-process"` or actual pane ID for tmux mode
|
||||
- `subscriptions`: Empty array (possibly for message routing)
|
||||
|
||||
**Config grows incrementally** — starts with just lead member (620 bytes), grows as teammates are added (→ 1886 → 3188 → 4551 bytes).
|
||||
|
||||
---
|
||||
|
||||
## Question 5: How do shared tasks differ from regular tasks?
|
||||
|
||||
**Expected:** Team name directory vs UUID, richer task format
|
||||
|
||||
**Actual: CONFIRMED**
|
||||
|
||||
**Team tasks (`~/.claude/tasks/research-watchers/`):**
|
||||
```json
|
||||
{
|
||||
"id": "1",
|
||||
"subject": "Research Node.js fs.watch on Linux vs macOS",
|
||||
"description": "Research how Node.js fs.watch works differently...",
|
||||
"activeForm": "Researching Node.js fs.watch Linux vs macOS",
|
||||
"status": "in_progress",
|
||||
"blocks": [],
|
||||
"blockedBy": [],
|
||||
"owner": "fs-researcher"
|
||||
}
|
||||
```
|
||||
|
||||
**Internal teammate tracking tasks (4.json, 5.json, 6.json):**
|
||||
```json
|
||||
{
|
||||
"id": "4",
|
||||
"subject": "fs-researcher",
|
||||
"description": "You are \"fs-researcher\" on the \"research-watchers\" team...",
|
||||
"status": "in_progress",
|
||||
"blocks": [],
|
||||
"blockedBy": [],
|
||||
"metadata": { "_internal": true }
|
||||
}
|
||||
```
|
||||
|
||||
**Key differences from regular subagent tasks (`~/.claude/tasks/{UUID}/`):**
|
||||
| Feature | Regular tasks | Team tasks |
|
||||
|---------|--------------|------------|
|
||||
| Directory name | UUID | Human-readable team name |
|
||||
| File names | `.lock`, `.highwatermark` only | Numbered JSON files (1.json, 2.json...) |
|
||||
| Content | Lock files only (no task JSON) | Full task JSON with metadata |
|
||||
| Owner field | N/A | Teammate name |
|
||||
| Locking | `.lock` file | `.lock.lock` directory (mkdir atomic) |
|
||||
| Internal tasks | None | `_internal: true` for teammate spawn tracking |
|
||||
|
||||
---
|
||||
|
||||
## Question 6: Can we write to task/mailbox files to interact with teammates?
|
||||
|
||||
**Expected:** Possibly for tasks, no for messages
|
||||
|
||||
**Actual: LIKELY YES for both**
|
||||
|
||||
Evidence supporting external writes:
|
||||
1. **Inbox files** are plain JSON arrays — we could append messages
|
||||
2. **Task files** are plain JSON — we could modify status, add new tasks
|
||||
3. **File locking** uses `.json.lock` directories — we'd need to respect the locking protocol
|
||||
4. **Lock protocol**: Create directory `{file}.lock` → write → delete directory. Simple mkdir-based atomic lock.
|
||||
|
||||
**Not tested in this experiment** — would need a follow-up test to verify teammates actually pick up externally-added messages/tasks. But the format is clear and the locking is simple.
|
||||
|
||||
---
|
||||
|
||||
## Question 7: What happens to Codeman's idle detection with active teammates?
|
||||
|
||||
**Expected:** Lead may appear idle while teammates work
|
||||
|
||||
**Actual: Lead stays "idle" in Codeman's view, but terminal shows active status**
|
||||
|
||||
Observations:
|
||||
- Codeman session status showed `"idle"` throughout the experiment
|
||||
- The terminal output continued updating (task list checkboxes, teammate progress messages)
|
||||
- Lead displayed "Befuddling..." spinner while waiting for teammates
|
||||
- Token count climbed from 27k → 33k during the experiment
|
||||
- The `stop` hook DID fire at the end when the team was cleaned up
|
||||
|
||||
**Implication:** Current idle detection may trigger prematurely if:
|
||||
- It only checks Codeman's session status (which stays "idle")
|
||||
- It doesn't account for active teammates
|
||||
|
||||
**What we need:** Check `~/.claude/teams/*/config.json` for active members before declaring idle.
|
||||
|
||||
---
|
||||
|
||||
## Question 8: How many Claude processes spawn per teammate?
|
||||
|
||||
**Expected:** 1 claude process per teammate
|
||||
|
||||
**Actual: ZERO separate processes! Teammates are in-process threads.**
|
||||
|
||||
```
|
||||
# Only 2 claude processes (both Codeman sessions, none for teammates):
|
||||
25405 claude --dangerously-skip-permissions --session-id 236f004f... (our main session)
|
||||
383633 claude --dangerously-skip-permissions --session-id 461daa80... (test session + 3 teammates)
|
||||
|
||||
# Process tree for test session:
|
||||
claude(383633)─┬─{claude}(383635)
|
||||
├─{claude}(383636)
|
||||
├─... (22 threads total)
|
||||
└─{claude}(399252)
|
||||
```
|
||||
|
||||
**In-process mode = threads, not processes.** All 3 teammates run as threads within the single `claude` process (PID 383633). This explains:
|
||||
- No separate PIDs to track
|
||||
- No `/proc/{pid}/environ` for individual teammates
|
||||
- Lower resource overhead
|
||||
- Shared env vars automatically
|
||||
|
||||
---
|
||||
|
||||
## Question 9: Do teammates inherit Codeman env vars (hook events)?
|
||||
|
||||
**Expected:** Yes, if child processes
|
||||
|
||||
**Actual: YES, trivially — they're in-process threads**
|
||||
|
||||
Since teammates are threads in the lead's process (PID 383633), they share the exact same environment:
|
||||
```
|
||||
CODEMAN_SCREEN=1
|
||||
CODEMAN_SESSION_ID=461daa80-94ec-4e5e-a1bb-0518f78311bc
|
||||
CODEMAN_SCREEN_NAME=codeman-461daa80
|
||||
CODEMAN_API_URL=http://localhost:3000
|
||||
```
|
||||
|
||||
The `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1` env var was set via `settings.local.json`'s `env` key, which Claude Code reads at startup and sets on its process.
|
||||
|
||||
**Hook events:** The lead session's hooks (Notification, Stop) apply to the whole process. Teammate-specific hooks (`TeammateIdle`, `TaskCompleted`) are defined in the same `settings.local.json` and would fire for the lead's session.
|
||||
|
||||
---
|
||||
|
||||
## Question 10: Does subagent-watcher pick up teammates automatically?
|
||||
|
||||
**Expected:** Probably not
|
||||
|
||||
**Actual: YES! subagent-watcher detects teammates automatically.**
|
||||
|
||||
Teammates create transcript files in the standard subagent path:
|
||||
```
|
||||
~/.claude/projects/{hash}/{leadSessionId}/subagents/agent-{id}.jsonl
|
||||
```
|
||||
|
||||
Codeman's `/api/subagents` endpoint returned all 3 teammates as active subagents. They're indistinguishable from regular Task tool subagents except:
|
||||
1. Their `description` field starts with `<teammate-message teammate_id= team`
|
||||
2. They can be cross-referenced with `~/.claude/teams/{name}/config.json`
|
||||
3. They tend to be longer-lived than regular subagents
|
||||
|
||||
**Sub-subagents:** Teammates also spawn their own Task tool subagents (6 additional agents detected), creating a 3-level hierarchy: Lead → Teammates → Sub-subagents.
|
||||
|
||||
---
|
||||
|
||||
## Filesystem Event Timeline
|
||||
|
||||
```
|
||||
06:45:01 Session transcript created
|
||||
06:45:05 ~/.claude/teams/ created
|
||||
06:45:05 ~/.claude/teams/research-watchers/ created
|
||||
06:45:05 config.json created (lead member only, 620 bytes)
|
||||
06:45:05 ~/.claude/tasks/research-watchers/ created with .lock
|
||||
06:45:11 Task 1.json created (via .lock.lock directory lock)
|
||||
06:45:13 Task 2.json created
|
||||
06:45:15 Task 3.json created
|
||||
06:45:18 inboxes/ directory created
|
||||
06:45:18 fs-researcher.json inbox created (task_assignment message)
|
||||
06:45:18 perf-researcher.json inbox created
|
||||
06:45:19 api-researcher.json inbox created
|
||||
06:45:26 config.json updated (fs-researcher added, 1886 bytes)
|
||||
06:45:26 Subagent agent-ae50544.jsonl created (fs-researcher)
|
||||
06:45:26 Task 4.json created (internal: fs-researcher tracking)
|
||||
06:45:30 config.json updated (perf-researcher added, 3188 bytes)
|
||||
06:45:30 Subagent agent-aa20c65.jsonl created (perf-researcher)
|
||||
06:45:30 Task 5.json created (internal: perf-researcher tracking)
|
||||
06:45:34 config.json updated (api-researcher added, 4551 bytes)
|
||||
06:45:34 Task 6.json created (internal: api-researcher tracking)
|
||||
06:45:35 Subagent agent-a29de32.jsonl created (api-researcher)
|
||||
06:45:35+ Teammates working, additional subagent transcripts appearing
|
||||
06:47:xx Tasks completed, shutdown_requests sent to teammate inboxes
|
||||
06:47:56 team-lead.json inbox created (teammates reporting back)
|
||||
06:47:57 config.json updated multiple times (member removal?)
|
||||
06:48:02 CLEANUP: all inbox files deleted
|
||||
06:48:02 CLEANUP: inboxes/ directory deleted
|
||||
06:48:02 CLEANUP: config.json deleted
|
||||
06:48:02 CLEANUP: research-watchers team directory deleted
|
||||
06:48:02 CLEANUP: all task files deleted (1-6.json + .lock)
|
||||
06:48:02 CLEANUP: research-watchers task directory deleted
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Web UI Observations
|
||||
|
||||
**Terminal output:**
|
||||
- Task list appears with checkboxes: `☐ Research Node.js fs.watch on Linux vs macOS`
|
||||
- Checkboxes fill in as tasks complete: `☑ Research Node.js fs.watch...`
|
||||
- Each task shows assigned teammate: `(@fs-researcher)`
|
||||
- Spinner shows active teammate with progress
|
||||
|
||||
**Status bar:**
|
||||
- Shows team member selector: `@main @api-researcher @fs-researcher @perf-researcher`
|
||||
- Hint: `shift+↑ to expand` and `ctrl+t to show teammates`
|
||||
- Standard bypass permissions and token count still visible
|
||||
|
||||
**Subagent floating windows:**
|
||||
- Teammates DID appear as subagent floating windows in Codeman's web UI
|
||||
- They show the standard subagent info (model, tool calls, description)
|
||||
- Sub-subagents (teammates' own Task tool usage) also appear
|
||||
|
||||
**In-process mode specifics:**
|
||||
- No new terminal windows or panes
|
||||
- Everything renders in the single terminal session
|
||||
- Shift+Up/Down would switch between teammate views (not tested interactively)
|
||||
|
||||
---
|
||||
|
||||
## Conclusions & Key Surprises
|
||||
|
||||
### Surprises vs expectations
|
||||
|
||||
1. **Inboxes ARE filesystem-based** — contrary to docs saying "SendMessage tool". It's a hybrid: the tool writes to filesystem inboxes.
|
||||
2. **Teammates are threads, not processes** — no new OS processes, just threads within the lead's claude process.
|
||||
3. **Teammates appear as standard subagents** — existing subagent-watcher infrastructure works out of the box!
|
||||
4. **Config grows incrementally** — members are added one-by-one, not all at once.
|
||||
5. **Internal tracking tasks** — tasks 4-6 with `_internal: true` track teammate spawn state.
|
||||
6. **Auto-cleanup** — lead automatically cleaned up ALL artifacts after shutdown.
|
||||
7. **Sub-subagents** — teammates can spawn their own Task tool subagents (3-level hierarchy).
|
||||
8. **`teammateMode` is NOT a valid settings key** — display mode defaults to `in-process`.
|
||||
|
||||
### Design implications for Codeman
|
||||
|
||||
1. **TeamWatcher can be simple** — just poll `~/.claude/teams/` for directories + parse config.json
|
||||
2. **Subagent-watcher already works** — no new infrastructure needed for teammate transcript tailing
|
||||
3. **Idle detection needs team awareness** — check config.json members before declaring idle
|
||||
4. **Message interception is possible** — watch inbox JSON files for real-time message tracking
|
||||
5. **Task visualization is straightforward** — parse numbered JSON files in task directory
|
||||
6. **No process tracking needed** — teammates are threads, not separate processes
|
||||
7. **Distinguish teammates from subagents** — use description prefix `<teammate-message` or cross-reference config.json
|
||||
|
||||
### What to build first
|
||||
|
||||
1. **Team-aware idle detection** — highest priority, prevents premature respawn
|
||||
2. **TeamWatcher** — poll `~/.claude/teams/` for team creation/removal
|
||||
3. **Team tasks API** — parse task JSON files for UI display
|
||||
4. **Teammate badge in subagent windows** — mark teammate subagents differently from regular ones
|
||||
5. **Message timeline** — parse inbox files for inter-teammate communication display
|
||||
|
||||
### What we DON'T need to build
|
||||
|
||||
- Process discovery for teammates (they're threads)
|
||||
- Custom transcript tailing (subagent-watcher handles it)
|
||||
- Separate teammate window infrastructure (subagent windows work)
|
||||
@@ -1,773 +0,0 @@
|
||||
# HTTP API Reference
|
||||
|
||||
Codeman's HTTP API is a **stable contract** as of 1.0 — see
|
||||
[`versioning-policy.md`](versioning-policy.md) for the SemVer guarantee. This page
|
||||
defines the response envelope, status codes, error codes, versioning, and the SSE
|
||||
event channel.
|
||||
|
||||
## Versioning
|
||||
|
||||
- The stable, public surface is served under **`/api/v1/...`**. Pin external
|
||||
clients to this prefix.
|
||||
- The unversioned **`/api/...`** paths are a permanent alias of the current
|
||||
version (what the bundled web UI uses). They are kept working, but new external
|
||||
integrations should use `/api/v1`.
|
||||
- Breaking changes to the contract ship under a new prefix (`/api/v2`); `/api/v1`
|
||||
keeps its semantics. Additive changes (new endpoints, new optional fields, new
|
||||
error codes) are non-breaking and may appear in a minor release.
|
||||
- The implementation rewrites `/api/v1/*` → `/api/*` at the server level
|
||||
(`rewriteApiV1Url` in `src/web/server.ts`).
|
||||
|
||||
## Response envelope
|
||||
|
||||
Every JSON response uses one uniform envelope, applied centrally by a
|
||||
`preSerialization` hook (`src/web/server.ts`) — handlers return bare data and the
|
||||
hook wraps it:
|
||||
|
||||
**Success** — HTTP `2xx`:
|
||||
|
||||
```json
|
||||
{ "success": true, "data": <payload> }
|
||||
```
|
||||
|
||||
`data` is the endpoint's payload (object, array, or value). Endpoints with no
|
||||
payload return `{ "success": true, "data": {} }`.
|
||||
|
||||
**Error** — HTTP `4xx`/`5xx`:
|
||||
|
||||
```json
|
||||
{ "success": false, "error": "human-readable message", "errorCode": "NOT_FOUND" }
|
||||
```
|
||||
|
||||
`ApiResponse<T>` in `src/types/api.ts` is the canonical type.
|
||||
|
||||
> Non-JSON endpoints are exempt from the envelope: `GET /api/sessions/:id/file-raw`,
|
||||
> `GET /api/sessions/:id/tail-file` (SSE), `GET /api/download`,
|
||||
> `GET /api/screenshots/:name`, `GET /q/:code` (QR redirect), and the
|
||||
> `GET /ws/sessions/:id/terminal` WebSocket upgrade.
|
||||
|
||||
> The [agent wait endpoints](#long-polling-agent-wait) use the normal envelope but
|
||||
> are the only JSON endpoints that deliberately **hold the connection open**, for up
|
||||
> to 600 s. Proxy operators and HTTP clients with a global read timeout need to know
|
||||
> that before pointing them at Codeman.
|
||||
|
||||
⚠️ **A `401` is the one status that is not an envelope.** Authentication is rejected
|
||||
in a request hook, before any handler runs, and it replies with the bare string
|
||||
`Unauthorized` (`Unauthorized: hook secret required` on the hook path) plus
|
||||
`WWW-Authenticate: Basic realm="Codeman"`. There is no `success`, no `error`, and no
|
||||
`errorCode`, because the wrapping hook only wraps object payloads. So a client that
|
||||
pipes every response straight into a JSON parser dies with a parse error rather than
|
||||
reporting an auth failure, which is a confusing way to discover that a password is
|
||||
set. Branch on the HTTP status **before** parsing.
|
||||
|
||||
## Error codes → HTTP status
|
||||
|
||||
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
|
||||
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
|
||||
the HTTP status.
|
||||
|
||||
| `errorCode` | HTTP | Meaning |
|
||||
|-------------|------|---------|
|
||||
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
|
||||
| `UNAUTHORIZED` | 401 | Authentication required or failed |
|
||||
| `NOT_FOUND` | 404 | Resource does not exist |
|
||||
| `SESSION_BUSY` | 409 | Session is busy |
|
||||
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
|
||||
| `ALREADY_EXISTS` | 409 | Resource already exists |
|
||||
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
|
||||
| `RATE_LIMITED` | 429 | Too many requests |
|
||||
| `INTERNAL_ERROR` | 500 | Unexpected server error |
|
||||
|
||||
Adding a new error code is non-breaking; removing or renaming one is a major change.
|
||||
|
||||
## Long-polling (agent wait)
|
||||
|
||||
Three calls block until something happens instead of answering immediately. They
|
||||
exist because SSE is Codeman's only other "tell me when" channel, and an agent
|
||||
driving the API from a shell tool cannot practically hold a stream and parse
|
||||
events inline.
|
||||
|
||||
| Call | Blocks until |
|
||||
|------|--------------|
|
||||
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
|
||||
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
|
||||
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
|
||||
|
||||
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
|
||||
`GET .../wait`. It registers the waiter **before** writing, which closes the window
|
||||
in which a separate wait sees the session still idle from the previous turn and
|
||||
answers instantly with the wrong turn's result. Use it whenever you send a prompt
|
||||
and want to know when that prompt is done.
|
||||
|
||||
### Three semantics that break callers who assume otherwise
|
||||
|
||||
**1. A timeout is HTTP `200`, not an error.** A wait that ends without its signal
|
||||
returns `{"success":true, ...,"wait":{"timedOut":true,"signal":null}}`. The
|
||||
intended pattern is a client-side loop over short waits, because `tailscale serve`
|
||||
and cloudflared can both cut an idle connection, and turning every poll boundary
|
||||
into a `4xx` would make that loop indistinguishable from a real failure. `408` is
|
||||
auto-retried by several clients (silently doubling the polling load), `504` is what
|
||||
a genuine tunnel failure looks like, and `204` cannot carry `waitedMs` / `status` /
|
||||
`limitPaused`. Reserve error handling for the four codes in the table below.
|
||||
|
||||
**2. `stop` and `blocked` fire only for `claude` sessions.** Both come from Claude
|
||||
Code hooks, and no other mode installs them: `shell` runs no agent, and the external
|
||||
CLIs (`opencode`, `codex`, `gemini`, `antigravity`, `pi`) render their own TUIs and post
|
||||
no hooks. For every non-`claude` mode only `idle`, `working` and `exit` are
|
||||
accepted, and of those only `exit` is dependable: see the caveats under
|
||||
[Signals](#signals) before building on `idle`. Requesting `stop` or `blocked`
|
||||
**explicitly** on such a session is a
|
||||
`400`; omitting `until` never fails, the server just drops them from the default set
|
||||
and echoes the narrowed set back as `wait.until`. Three more places hooks can go
|
||||
missing even in `claude` mode: a **Docker case** needs
|
||||
`CODEMAN_DOCKER_BRIDGE_HOOKS=1`, since a container cannot reach a loopback-bound
|
||||
Codeman (without it, only `idle` / `working` / `exit` work); a **remote-SSH
|
||||
case** runs the agent on another host, whose hooks may never reach this server at
|
||||
all; and a case whose hook config was written by **Codeman < 1.13.0 against an
|
||||
`--https` install** carries hook curls without `-k`, which TLS-fail silently (the
|
||||
hook line ends in `|| true`). Codeman now writes `curl -sk` and repairs a stale
|
||||
case config the next time a session starts in that case. When in doubt, ask for
|
||||
`stop,idle,exit` so a session without hooks still resolves on the heuristic
|
||||
signal.
|
||||
|
||||
**3. `from=now` does not mean "printed after you asked".** tmux repaints the visible
|
||||
screen on attach, on resize, and on any TUI redraw, and a repaint arrives as
|
||||
ordinary output, so text that was already on screen can satisfy a fresh wait. This
|
||||
was observed live: a marker echoed a minute earlier matched instantly on a new
|
||||
`from=now` wait. It is inherent to running the agent under a multiplexer, so the
|
||||
contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
|
||||
`echo $MARK`, then wait on `$MARK`), never a generic string like `BUILD OK`.
|
||||
|
||||
### Signals
|
||||
|
||||
| Signal | Source | Actually fires for |
|
||||
|--------|--------|--------------------|
|
||||
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
|
||||
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
|
||||
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
|
||||
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
|
||||
| `exit` | no process is behind the session | every mode |
|
||||
|
||||
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
|
||||
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
|
||||
is omitted is `stop,idle,exit` (`exit` is in there so a worker that crashes resolves
|
||||
the wait promptly instead of burning the caller's whole timeout on something that
|
||||
can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit`
|
||||
once the session is up: the default set's `idle` also resolves on a spinner pause,
|
||||
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
|
||||
can land inside your first wait window and report a turn that never ran. Measured:
|
||||
a session parked on the trust dialog emits no *further* `idle`, so it is the
|
||||
startup transition, not the dialog, that produces the false success below.
|
||||
|
||||
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
|
||||
server answers from `pid === null` plus a mux-layer pane-death probe, and that
|
||||
covers a session that exited — including a worker that died *inside* its tmux pane
|
||||
while the local attach client (and therefore `pid`) lives on — one that was
|
||||
detached, and one that was **created but never started**. So the first wait
|
||||
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
|
||||
milliseconds, and reading that as "the worker died" is wrong: it means start it, or
|
||||
wait for it to come up. `status` is carried alongside so nothing is hidden. The
|
||||
alternative (trusting `status`) is worse, because a dead PTY parks the session at
|
||||
`status: "idle"`, which would answer the default wait with `immediate: true` for a
|
||||
worker that has crashed. A worker dying while a wait is parked resolves it within
|
||||
a few seconds (a background death-watcher), not at the timeout.
|
||||
|
||||
⚠️ **`blocked` is reachable less often than the table suggests.** It fires on two
|
||||
hooks, and the default configuration suppresses one of them: Codeman spawns claude
|
||||
with `--dangerously-skip-permissions`, so permission prompts do not happen unless the
|
||||
instance is switched to the `auto` Claude mode (App Settings), or the caller is a
|
||||
multi-user account without the bypass grant, which is forced to `--permission-mode
|
||||
auto`. What does still fire under the default is `elicitation_dialog`, the agent
|
||||
asking the user a question. So `until=stop,blocked,exit` is a reasonable belt on a
|
||||
long turn, but a worker that never comes back is far more likely to be working than
|
||||
blocked, and polling `blocked` alone will sit at its timeout.
|
||||
|
||||
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
|
||||
session emits its one `idle` at startup and then stays `status: "idle"` forever,
|
||||
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
|
||||
requires a transition (and so does `fresh=1`), both can only time out there:
|
||||
a documented default `wait` on a shell worker running `sleep 4` times out at the
|
||||
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
|
||||
instead. The same caution applies to the external CLIs.
|
||||
|
||||
### Readiness is not a signal
|
||||
|
||||
Nothing here reports "the agent is ready for a prompt", and no combination of
|
||||
`until`/`fresh` synthesizes one. A freshly created session reads as `exit` (above),
|
||||
and a `claude` worker in a brand-new case comes up on the CLI's **trust dialog**,
|
||||
which contains a ❯ prompt of its own. Send-and-wait posted at that moment types the
|
||||
prompt into the dialog, where the `\r` never gets past it, while the session's
|
||||
startup `idle` lands inside the wait window: the wait resolves on `idle` in a
|
||||
couple of seconds with `timedOut: false`, which looks exactly like a completed
|
||||
turn.
|
||||
|
||||
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
|
||||
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
|
||||
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
|
||||
dialog only as the bounded fallback.
|
||||
|
||||
⚠️ **The fallback is not a bare `\r`.** Claude Code 2.1.252 unnumbered the dialog's
|
||||
options, reversed them and highlights `No, exit`, so an Enter sent blind quits the
|
||||
CLI and the pane dies seconds after the spawn. Read the `❯` marker off the current
|
||||
frame (`GET /api/v1/sessions/:id/terminal?full=1`), send `ESC [ B` while it is on
|
||||
`No, exit`, re-read, and confirm only once it is on `Yes, I trust this folder`.
|
||||
Reading the current frame is also what keeps this correct on later runs: the dialog
|
||||
text stays in the terminal buffer for the life of the session, so a `trust` probe
|
||||
with `from=buffer` keeps matching long after the dialog is gone. A worked version is in
|
||||
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
|
||||
|
||||
### `GET /api/v1/sessions/:id/wait`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
|-------|------|---------|-------|
|
||||
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
|
||||
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
|
||||
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
|
||||
|
||||
```bash
|
||||
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
|
||||
```
|
||||
|
||||
Both GET wait routes answer with `Cache-Control: no-store`, because the documented
|
||||
pattern polls one identical URL in a loop and a cached `{"timedOut":true}` would
|
||||
turn that loop into a busy spin. `POST .../input` sends no cache header (it is a
|
||||
POST, which is not heuristically cacheable).
|
||||
|
||||
⚠️ **Unknown query parameters are ignored, not rejected**, with one exception
|
||||
(`regex`, below). In particular `match=` on `/wait` is silently dropped and you get
|
||||
a plain signal wait, so check the endpoint path before blaming the parameters.
|
||||
|
||||
### `GET /api/v1/sessions/:id/wait-output`
|
||||
|
||||
| Param | Type | Default | Notes |
|
||||
|-------|------|---------|-------|
|
||||
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
|
||||
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
|
||||
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
|
||||
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
|
||||
|
||||
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
|
||||
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
|
||||
instead of waiting on the wrong thing. The reasoning is in
|
||||
[`architecture-invariants.md`](architecture-invariants.md#agent-wait-primitives).
|
||||
|
||||
#### What the matcher actually sees
|
||||
|
||||
The matcher scans the raw PTY stream, **normalized**: ANSI escape sequences are
|
||||
stripped — CSI, OSC, and the charset-designation escapes a stock bash prompt emits
|
||||
on every line (`ESC ( B`), so `match=tnode:` matches a prompt that renders
|
||||
`…@tnode:` — a partial escape arriving at a chunk boundary is held back until its
|
||||
tail arrives, and a match may straddle PTY chunks: `printf STRAD; sleep 1; printf
|
||||
DLEQQ` is matchable as `STRADDLEQQ` (all measured live). Three caveats remain:
|
||||
|
||||
⚠️ **It is still the byte stream, not the rendered pane.** `GET .../terminal`
|
||||
answers from a tmux screen capture (`data.source: "mux-visible"`), the finished
|
||||
picture; the matcher sees the stream that painted it. For linear output the two
|
||||
agree once escapes are stripped, but a full-screen TUI composes its picture with
|
||||
cursor positioning, so what the pane shows and what the stream carries can differ.
|
||||
Seeing your string in `terminal?tail=` makes a match likely, not guaranteed.
|
||||
|
||||
⚠️ **A TUI's text can arrive without its spaces.** Claude Code positions words
|
||||
with cursor moves rather than printing spaces, so screen text can reach the
|
||||
matcher as `Quicksafetycheck:Isthisaprojectyoucreated...`. Whether a given phrase
|
||||
keeps its spaces depends on how the TUI happened to draw it (measured: `I trust
|
||||
this folder` matched, `Quick safety check` did not), so a multi-word `match`
|
||||
against a TUI pane is unreliable rather than impossible. Match a **single
|
||||
space-free token**, ideally one you printed yourself. Plain command output (a
|
||||
shell worker, an `echo`) keeps its spaces.
|
||||
|
||||
⚠️ **The returned `snippet` is a rendering of the matched text, not a quotation of
|
||||
it.** It is cut from the same normalized stream the match ran against, then
|
||||
cleaned for display: remaining raw control bytes are removed (an agent pipes the
|
||||
snippet into its own terminal, so a worker's bytes must not be able to reset that
|
||||
display) and blank runs are collapsed. A printable needle that matched will appear
|
||||
in it; a needle containing control bytes or a blank run may not survive verbatim.
|
||||
|
||||
```bash
|
||||
MARK="DONE_$RANDOM"
|
||||
curl -sG "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode "match=$MARK" --data-urlencode 'timeout=120000'
|
||||
```
|
||||
|
||||
Build the query with `-G --data-urlencode` rather than by hand: a `+` in a
|
||||
hand-written query string decodes to a space.
|
||||
|
||||
### `POST /api/v1/sessions/:id/input` with `wait`
|
||||
|
||||
Two optional fields on the existing endpoint:
|
||||
|
||||
| Field | Type | Notes |
|
||||
|-------|------|-------|
|
||||
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
|
||||
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
|
||||
|
||||
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
|
||||
"absent" rather than failing validation. That is deliberate: `.optional()` would
|
||||
reject it, which has shipped as a real bug twice.
|
||||
|
||||
The input must end with `\r` (a real carriage return in the JSON string): Enter is
|
||||
sent only when the input contains one, so text without it is typed onto the
|
||||
worker's prompt but never submitted, and the wait then runs its full timeout on a
|
||||
turn that never started. Verified live; this is the most common silent failure on
|
||||
this endpoint.
|
||||
|
||||
```bash
|
||||
curl -s -X POST "$API/api/v1/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"run the tests\r","useMux":true,"clientId":"agent-1","seq":1,
|
||||
"wait":"stop","waitTimeout":600000}'
|
||||
```
|
||||
|
||||
A **tagged duplicate** (a `clientId` + `seq` pair the server has already applied)
|
||||
still honors `wait`, because the caller's question is unanswered, but it answers
|
||||
from the session's current state rather than requiring a new transition: the
|
||||
original turn may be long over. It comes back as
|
||||
`"delivered": false, "duplicate": true`.
|
||||
|
||||
**Wake-on-LAN hosts** (`docs/remote-sessions.md` §Wake-on-LAN): when the session's
|
||||
remote host has a wake target and is asleep, the non-wait form answers `200` with
|
||||
`{"buffered": true}` — the bytes are held and flushed after the host is back — or
|
||||
`{"buffered": true, "dropped": true}` for a chunk over the 4 KB wake buffer, which
|
||||
is gone (never delivered as a fragment). Both fields are additive to the historical
|
||||
bare `{}`. With `wait`, the route blocks on the wake instead and answers
|
||||
`422 OPERATION_FAILED` ("did not come back after a wake-on-LAN request — nothing was
|
||||
sent") when the host never returns, rather than writing into the stalled pane and
|
||||
reporting `delivered:true` plus a timeout.
|
||||
|
||||
Two endpoints back that flow directly, both scoped to one session's remote host and
|
||||
both refusing a session that is not remote (`400 INVALID_INPUT`):
|
||||
|
||||
| Method | Path | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `GET` | `/api/sessions/:id/reachability` | Whether the session's remote host answers SSH right now, plus whether a wake target is configured. Read-only: it never wakes. `{"reachable": true\|false\|null, "wakeConfigured": "mac"\|"command"\|"none"}`, where `null` means the answer is unknown (a proxied host, where a TCP probe proves nothing). |
|
||||
| `POST` | `/api/sessions/:id/wake` | Wake the host and wait for it to accept SSH again, bounded by the request budget. `422 OPERATION_FAILED` when it does not come back; `400 INVALID_INPUT` with "No wake-on-LAN target configured for this host" when nothing is set. |
|
||||
|
||||
⚠️ Waking is deliberately reachable only from an explicit user action (this route, a
|
||||
session create/attach, or typing into a sleeping session). No watcher, dropped-session
|
||||
handler or boot-recovery path may wake a host, or a suspended machine would be woken
|
||||
again seconds after every suspend; `test/remote-wake.test.ts` pins that as an import
|
||||
fence around `src/remote-wake.ts`.
|
||||
|
||||
### Response
|
||||
|
||||
All three nest the wait result under `data.wait`, so one client helper works against
|
||||
any of them:
|
||||
|
||||
```json
|
||||
{ "success": true, "data": {
|
||||
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
|
||||
"status": "idle",
|
||||
"limitPaused": false,
|
||||
"wait": {
|
||||
"signal": "stop", "until": ["stop", "idle", "exit"],
|
||||
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
|
||||
"waitedMs": 8421, "timeoutMs": 60000
|
||||
}
|
||||
}}
|
||||
```
|
||||
|
||||
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
|
||||
`status` and `limitPaused`. `POST .../input` **without** `wait` is unchanged and
|
||||
still returns `{"success": true, "data": {}}`.
|
||||
|
||||
⚠️ `delivered: false` has **two** meanings, and they must be told apart by
|
||||
`duplicate`: with `duplicate: true` the input was suppressed as an already-applied
|
||||
redelivery (harmless, the turn it refers to may be long over), while with
|
||||
`duplicate: false` the **write failed** (typically no PTY behind the session). A
|
||||
client that reads `delivered === false` as "duplicate" silently treats a failed send
|
||||
as a success.
|
||||
|
||||
| Field | Type | Meaning |
|
||||
|-------|------|---------|
|
||||
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
|
||||
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
|
||||
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
|
||||
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
|
||||
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
|
||||
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
|
||||
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
|
||||
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
|
||||
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
|
||||
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
|
||||
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
|
||||
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
|
||||
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
|
||||
|
||||
Read the outcome by discriminator, in this order:
|
||||
|
||||
1. `wait.signal !== null` (or `wait.matched === true`): the thing happened.
|
||||
2. `wait.timedOut`: a poll boundary. Loop again.
|
||||
3. `wait.ended` or `wait.aborted`: the wait answered nothing, because the session is
|
||||
gone or was never running. Re-check the session instead of looping.
|
||||
|
||||
`wait.immediate` is not a fourth outcome: it rides along with the first one and
|
||||
means the condition already held at call time, so nothing was actually waited for.
|
||||
If that is not what you meant, you wanted `fresh=1` or the send-and-wait form. Note
|
||||
that `{"signal":"exit","immediate":true}` on a session you just created is the
|
||||
not-started-yet case, not a crash.
|
||||
|
||||
**The timeout is clamped, so read it back.** A request for 1800000 ms is silently
|
||||
reduced to the server's ceiling (600000 ms by default, operator-tunable), and a
|
||||
request for 1 ms is raised to 1000 ms. `wait.timeoutMs` is the value that was
|
||||
applied. Without checking it, a caller that asked for 30 minutes and got 10 will
|
||||
read the timeout as "the worker is wedged" and kill a session that was working fine.
|
||||
|
||||
### Errors
|
||||
|
||||
| `errorCode` | HTTP | When |
|
||||
|-------------|------|------|
|
||||
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
|
||||
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
|
||||
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
|
||||
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
|
||||
|
||||
The two capacity codes are deliberately different. A process-wide cap reported as
|
||||
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
|
||||
error message names the cap that was hit.
|
||||
|
||||
⚠️ A `401` is **not** in this table and is not an envelope at all (see
|
||||
[Response envelope](#response-envelope)). It matters most here: a polling loop that
|
||||
pipes each wait straight into `jq` fails with a parse error on every iteration
|
||||
against a password-protected server, which reads as "the wait endpoints are broken".
|
||||
Check the status first.
|
||||
|
||||
The per-session cap is a **combined** budget: signal waiters and output waiters
|
||||
count against the same 16, not 16 of each. An abandoned request no longer holds its
|
||||
slot, because the routes release the waiter when the client disconnects, but a
|
||||
client that opens many concurrent waits against one session will still hit the cap.
|
||||
|
||||
## Session lineage (`parentSessionId`)
|
||||
|
||||
A create request may name the session that spawned it, which the web UI draws as a
|
||||
line between the two tabs. Accepted on `POST /api/v1/sessions` and
|
||||
`POST /api/v1/quick-start`, either way:
|
||||
|
||||
```bash
|
||||
# as a body field
|
||||
-d '{"caseName":"worker-1","mode":"claude","parentSessionId":"'"$CODEMAN_SESSION_ID"'"}'
|
||||
|
||||
# or as a header, which is what an agent driving many spawns should use: set it once
|
||||
# on the curl invocation and every spawn call carries it
|
||||
-H "X-Codeman-Parent-Session: $CODEMAN_SESSION_ID"
|
||||
```
|
||||
|
||||
The body field wins if both are present. The value is resolved against live sessions
|
||||
(exact id, or a unique prefix of at least 8 characters) and must belong to the same
|
||||
owner as the session being created.
|
||||
|
||||
**It cannot fail your spawn.** An unknown, stale, foreign or malformed value is
|
||||
silently dropped and the session is created without lineage — never a `400`. It is
|
||||
also pure decoration: it confers no permission, and a child is unaffected by its
|
||||
parent exiting. It appears on session state as `parentSessionId` (absent when
|
||||
unresolved) and survives a server restart.
|
||||
|
||||
## Approvals Inbox
|
||||
|
||||
Cross-session queue of prompts waiting on a human (permission dialogs,
|
||||
AskUserQuestion questions, idle prompts). Claude-mode sessions only; items are
|
||||
in-memory (a server restart drops them; the next prompt re-fires the hook).
|
||||
Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
|
||||
|
||||
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
|
||||
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
|
||||
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
|
||||
toolSummary?, message?, cwd?, context?, options?: {n, label}[],
|
||||
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
|
||||
`options` is present only when the dialog's numbered choices parsed
|
||||
confidently; `acknowledgedAt` marks an item a human has already looked at
|
||||
(see `/viewed` below) and tells clients not to re-arm its tab alert. Listing
|
||||
also runs a staleness sweep over the caller's own items: the pane is
|
||||
re-captured, and an item whose dialog no longer parses is resolved as
|
||||
`resolved_in_terminal` instead of being returned (only items whose original
|
||||
frame parsed `options` can be dropped this way, so an unreadable capture
|
||||
keeps the item).
|
||||
- `POST /api/v1/approvals/:id/answer` with `{ action: 'approve' }` (sends the
|
||||
digit `1`), `{ action: 'deny' }` (sends Esc), `{ action: 'option', option: n }`
|
||||
(sends the digit; accepted only when `n` is among the item's parsed
|
||||
`options`), or `{ action: 'text', text }` (idle prompts only; submits the
|
||||
line as a prompt). `404 NOT_FOUND` when the item is no longer pending,
|
||||
`409 CONFLICT` when the dialog left the screen or another actor answered
|
||||
first, `422 OPERATION_FAILED` when the session refused input.
|
||||
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
|
||||
- `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId,
|
||||
acknowledged: itemId | null }`. Marks the session's pending **idle** item as
|
||||
seen by a human (the web UI calls it when you open the session's tab): the
|
||||
item stays pending and answerable, but stops arming the yellow tab alert on
|
||||
every client, including after a reload. Permission/question items are never
|
||||
acknowledged this way, since looking at a dialog does not answer it. `404`
|
||||
for an unknown or inaccessible session; acknowledging twice is a no-op
|
||||
(`acknowledged: null`).
|
||||
|
||||
SSE events: `approval:pending` (full item), `approval:updated` (context/options
|
||||
re-captured, or the item acknowledged), `approval:resolved` (`{ id, sessionId, kind, resolution }` with
|
||||
`resolution` one of `answered | resolved_in_terminal | superseded |
|
||||
session_ended | dismissed | expired`).
|
||||
|
||||
## Reboot restore
|
||||
|
||||
A host reboot takes the tmux server down with it, so every pane dies and the
|
||||
board comes up empty. At boot Codeman works out which sessions the reboot
|
||||
destroyed and holds that plan in memory, and these endpoints let a client offer
|
||||
it to the user. Nothing creates a pane until the user asks: the boot-time reboot
|
||||
heuristic decides whether to ASK, never whether to act.
|
||||
|
||||
Claude-mode sessions only (others carry their conversation id in their own
|
||||
config object); remote and docker sessions are never offered, because both need
|
||||
another host or container to be up. The plan is in-memory, so a server restart
|
||||
drops it and the offer is gone; the conversations themselves are unaffected,
|
||||
since they live in the CLI's own transcript store and stay reachable from the
|
||||
Resume list. A plan nobody spends expires after 24 hours.
|
||||
|
||||
- `GET /api/v1/reboot-restore` → `{ sessions: RestorableSession[],
|
||||
scrollbackRestored: false }`, ownership-scoped in multi-user mode.
|
||||
`RestorableSession`: `{ id, name?, workingDir, mode, owner? }`. The persisted
|
||||
record itself is never sent. `scrollbackRestored` is always `false` and exists
|
||||
so a client states it: a restored session is a NEW pane, so the conversation
|
||||
continues and the terminal history does not.
|
||||
- `POST /api/v1/reboot-restore/restore` with `{ sessionIds?: string[] }` (omit
|
||||
to restore everything the caller can see) → `{ restored: RestorableSession[],
|
||||
skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing`
|
||||
(the directory is gone), `workspace-forbidden` (in multi-user mode it is
|
||||
outside the workspace of the user the session belongs to, re-checked against
|
||||
that owner's current grant rather than the caller's), `already-live` (the conversation is already
|
||||
open, typically resumed by hand from the Resume list), `capacity-reached`
|
||||
(the global or per-user session cap), or `rebuild-failed` (the agent would not
|
||||
start, most often a CLI binary missing from the server's PATH).
|
||||
`409 CONFLICT` when that caller already has a restore running. Entries are
|
||||
removed from the plan before any pane is built, so a double-click cannot put
|
||||
two panes on one conversation; anything that never became a pane goes back on
|
||||
offer, except `already-live`, which cannot stop being true. A restored session
|
||||
comes back attached, idle and disarmed: respawn controllers and Ralph loops
|
||||
are never re-armed automatically.
|
||||
- `POST /api/v1/reboot-restore/dismiss` → `{ dismissed: n }`. Drops the offer
|
||||
for everything the caller can see.
|
||||
|
||||
Each rebuilt session also emits the ordinary `session:created` SSE event, so
|
||||
clients other than the one that clicked pick it up without refetching.
|
||||
|
||||
## Read My Mind intent profiles
|
||||
|
||||
Per-case profiles of what the user is trying to accomplish: user/agent-stated
|
||||
goals plus the user's recently submitted prompts, captured from the Claude
|
||||
session transcript while the opt-in `readMyMindEnabled` setting is on (default
|
||||
OFF). Keyed by owner + workingDir, so the profile survives `/clear`, respawns,
|
||||
and session churn. Stored in `~/.codeman/intents.json` (mode 0600); never fed
|
||||
into `/api/v1/search`. Design: [`readmymind-plan.md`](readmymind-plan.md);
|
||||
user guide: [`readmymind.md`](readmymind.md).
|
||||
|
||||
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
|
||||
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
|
||||
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
|
||||
50, each <= 500 chars). A case with nothing recorded answers an empty
|
||||
profile with `updatedAt: 0`; nothing is persisted by reads.
|
||||
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
|
||||
schema) replaces the goals text and answers the updated profile.
|
||||
`400 INVALID_INPUT` on over-long or unknown fields.
|
||||
- `DELETE /api/v1/sessions/:id/intent` -> `{ deleted: boolean }` forgets the
|
||||
case's profile entirely.
|
||||
- `POST /api/v1/sessions/:id/readmymind` predicts the user's next prompt:
|
||||
a one-shot model call over the intent profile plus live session signals
|
||||
(pending approval dialog, transcript tail, git state, run-summary events,
|
||||
sibling sessions). Body is optional; the rethink flow passes
|
||||
`{ steer?, rejected? }` (strict schema: `steer` <= 2000 chars, `rejected`
|
||||
up to 10 strings <= 1000 chars). Answers
|
||||
`{ suggestions: { prompt, why, kind }[], durationMs }` with 1-3 suggestions
|
||||
(`kind`: `continue` | `verify` | `redirect`; prompts are single-line).
|
||||
Claude-mode sessions only (`400 INVALID_INPUT` otherwise); one prediction in
|
||||
flight per session (`409 CONFLICT`); predictor failures answer
|
||||
`502 OPERATION_FAILED`. Takes 5-90 s and costs real tokens. Suggestions are
|
||||
only ever returned, never sent: submitting one is the caller's explicit act.
|
||||
|
||||
All four enforce session ownership in multi-user mode; a foreign session id
|
||||
answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the
|
||||
same directory are distinct by construction.
|
||||
|
||||
## Custom Model Endpoints
|
||||
|
||||
Points a session's harness at a user-configured OpenAI-compatible endpoint —
|
||||
local (llama.cpp, vLLM, DGX Spark) or cloud (Azure AI Foundry, OpenRouter) —
|
||||
instead of its native cloud backend, gated by the opt-in
|
||||
`customModelEndpointsEnabled` setting (default OFF). Endpoints are
|
||||
machine-level infra, like remote/docker hosts: writes are admin-only in
|
||||
multi-user mode. Design: [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md);
|
||||
user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
|
||||
|
||||
- `GET /api/v1/model-endpoints` -> `CustomModelHost[]`, an unwrapped bare
|
||||
array like every other list route (still riding the standard `{success,
|
||||
data}` envelope on the wire — unwrap it the same way). Answers `[]` for a
|
||||
non-admin in multi-user mode. `apiKey` is never returned; `apiKeySet:
|
||||
boolean` reports whether one is stored, so a client can render "unchanged
|
||||
if left blank" without ever holding the real value.
|
||||
- `POST /api/v1/model-endpoints` with `{ id, label, baseUrl, apiKey?,
|
||||
authStyle?, defaultModelId? }` creates one. `id` must match
|
||||
`^[a-zA-Z0-9_-]+$`; `authStyle` is `bearer` (default) or `api-key`, never
|
||||
both (a real server hung indefinitely when sent both headers on one
|
||||
request); `baseUrl` must be `http(s)`, carry no embedded credentials, and
|
||||
is refused if it points at (or resolves to) a link-local or
|
||||
cloud-metadata address. `409 ALREADY_EXISTS` on a duplicate id.
|
||||
- `PUT /api/v1/model-endpoints/:id` updates one. An **absent** `apiKey`
|
||||
keeps the stored one rather than clearing it — the client never receives
|
||||
the real value to resend deliberately unchanged, so omission is the only
|
||||
way to say "leave it alone"; there is no way to clear a key back to unset
|
||||
this way. `defaultModelId`, when set, must be one of that endpoint's own
|
||||
`models` (`400 INVALID_INPUT` otherwise).
|
||||
- `DELETE /api/v1/model-endpoints/:id` removes one.
|
||||
- `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's
|
||||
own `GET /v1/models` and stores the result as `models`, updating
|
||||
`lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own
|
||||
response already reports loaded) `modelContextLengths` and `modelSizesGB`.
|
||||
A `defaultModelId` that no longer appears in the fresh list is dropped
|
||||
rather than carried forward invalid. Failures answer `422 OPERATION_FAILED`
|
||||
with the underlying connection error, or a named egress refusal if the
|
||||
resolved address turned out to be blocked. The same refresh also runs
|
||||
automatically for every saved endpoint every 5 minutes in the background
|
||||
(`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, started from
|
||||
`server.ts`), so there is no route for triggering "refresh all" — one
|
||||
endpoint being unreachable on a cycle never blocks the others.
|
||||
- `GET /api/v1/model-endpoints/:id/running-status` -> `{ isLlamaSwap,
|
||||
running: [{model, state}], logLine? }`, read-only, no admin gate
|
||||
(any session owner who could already point a session at this endpoint can
|
||||
equally ask what it currently has loaded). `isLlamaSwap` is
|
||||
feature-detected via the endpoint's own `GET /running` — a plain
|
||||
llama.cpp/OpenAI-compatible server has none and always answers `false`.
|
||||
`logLine`, present only when `isLlamaSwap` is true, is the most recent
|
||||
REAL backend `llama-server` process log line (`load_model: ...`,
|
||||
`llama_server: model loaded`, etc.), sourced from the endpoint's own
|
||||
`GET /api/events` SSE stream and filtered to `source: "upstream"` frames
|
||||
only (never llama-swap's own `source: "proxy"` request-access log) — one
|
||||
connection is held open per endpoint and reused across every poller,
|
||||
idle-closed after 30s of nobody asking. This is what the Run-menu
|
||||
picker's loading banner polls once a second while a model is loading.
|
||||
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId,
|
||||
confirmed? } | { clear: true }` applies (or clears) the session's
|
||||
selection and **restarts the session's CLI process in place** — every
|
||||
supported harness reads its endpoint config at process start, never per
|
||||
turn, so there is no live hot-swap. (`POST /api/v1/quick-start`'s own
|
||||
`customModel: { endpointId, modelId, confirmed? }` field is the
|
||||
no-restart equivalent for a session that doesn't exist yet — see below.)
|
||||
A Claude session resumes its existing conversation across the restart;
|
||||
pi/omp/grok additionally get a forced `--model`/`-m` value, since for
|
||||
those three the config file alone does not select it. `400 INVALID_INPUT`
|
||||
for a remote (SSH) or Docker session — both restart their agent
|
||||
differently under the hood, and applying to one would report success
|
||||
while changing nothing. Two more responses replace the normal
|
||||
`{customModel, restarted}` shape, neither an error, and neither restarts
|
||||
or creates anything on the first ask. ⚠️ **Each is answered by its OWN
|
||||
flag on the retry, and answering one is not consent to the other**: they
|
||||
are questions about different people, and while they shared a single flag
|
||||
a caller who confirmed the context warning silently agreed to evict
|
||||
another session's model as well. Send `confirmedContext: true` to proceed
|
||||
past the context warning, `confirmedSwap: true` past the swap conflict,
|
||||
and both when both were asked (they accumulate, so the second retry still
|
||||
carries the first answer). The original `confirmed: true` still means
|
||||
BOTH and is still accepted, because it shipped in this feature's
|
||||
HTTP-API-only cut; new callers should send the specific one:
|
||||
- `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` —
|
||||
llama.cpp/llama-swap only runs one model at a time, and switching would
|
||||
unload a model another **live session's own selection** is actively
|
||||
using. Never returned for a plain (non-llama-swap) server, and never
|
||||
just because a swap is needed at all — only when it would disrupt
|
||||
someone else.
|
||||
- `{requiresContextWarning: true, modelId, contextLength,
|
||||
minSafeContextTokens}` — Claude Code's own fixed per-turn overhead
|
||||
(system prompt + tool schemas) can exceed a small model's entire
|
||||
discovered context on its own, before any conversation history exists
|
||||
to compact, guaranteeing the very first message fails regardless of
|
||||
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`. Gated on the CLI registry declaring a
|
||||
`contextLengthVar` (claude only today), so it never fires for another
|
||||
harness.
|
||||
- `POST /api/v1/quick-start`'s `customModel: { endpointId, modelId,
|
||||
confirmed?, confirmedContext?, confirmedSwap? }` field (alongside its
|
||||
normal `caseName`/`mode`/etc. body)
|
||||
computes the same injection **before** the session exists and launches
|
||||
directly on the endpoint — no restart, because there was never a
|
||||
native-backend boot to restart away from. Runs the identical checks as
|
||||
the dedicated route above (`requiresConfirmation`/`requiresContextWarning`,
|
||||
same shapes, same per-question `confirmedContext`/`confirmedSwap` retry),
|
||||
and is refused the same way
|
||||
for a remote or Docker case. This is what the Run-menu picker uses for
|
||||
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the
|
||||
dedicated restart route above (its `--resume`-based restart is far less
|
||||
jarring than a full relaunch, and folding it into the one-shot path is
|
||||
separate work — see `docs/custom-model-endpoints-plan.md`).
|
||||
|
||||
## CLI management
|
||||
|
||||
Read and write the CLI registry (`docs/cli-registry.md`). Every **write** route answers `403 FORBIDDEN` while `cliManagementEnabled` is off (the default), and for a non-admin in multi-user mode. A write that would overwrite a `clis.json` which does not parse, or which has group/world permission bits, is refused with `409 CONFLICT` and a message naming the fix; the file is left untouched.
|
||||
|
||||
| Method | Path | Body | Notes |
|
||||
| -------- | ----------------------------- | ------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
|
||||
| `GET` | `/api/clis` | none | Every entry, disabled ones included: `id`, `label`, `shortBadge`, `order`, `kind`, `enabled`, `stock`, `installed`, and `installCommand` for a stock entry. Not gated; a non-admin in multi-user mode gets `[]`. |
|
||||
| `PUT` | `/api/clis/:id` | `{ enabled }` | Toggle an existing entry, stock or custom. `404` for an unknown id; `400 INVALID_INPUT` when disabling a `kind: 'shell'` entry. |
|
||||
| `POST` | `/api/clis/:id/install` | none | Run a **stock** entry's install command (never a custom one: `400`). `409 CONFLICT` while an install for the same id is running; `422 OPERATION_FAILED` with the output tail when it fails. Never enables the entry. |
|
||||
| `POST` | `/api/clis` | `{ id, label, shortBadge, binaries, argv, enabled? }` | Create a custom entry. `409 ALREADY_EXISTS` for a stock id or an existing custom id. `enabled` defaults to `true`. |
|
||||
| `PUT` | `/api/clis/custom/:id` | `{ label, shortBadge, binaries, argv, enabled? }` | Replace an existing custom entry. An absent `enabled` keeps the entry's current state. `400` for a stock id, `404` for an unknown one. |
|
||||
| `DELETE` | `/api/clis/:id` | none | Delete a custom entry. `400` for a stock id, `404` for an unknown one. |
|
||||
|
||||
## Voice dictation
|
||||
|
||||
Browser dictation transcribed through this server's Claude Code login, i.e. the
|
||||
same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synced
|
||||
`claudeVoiceEnabled` setting (default OFF). Design:
|
||||
[`claude-voice-plan.md`](claude-voice-plan.md).
|
||||
|
||||
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
|
||||
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
|
||||
signed in to Claude Code on the server), `expired` (the access token elapsed;
|
||||
running any Claude session refreshes it) or `malformed`. The OAuth token
|
||||
itself is never returned by this or any other endpoint.
|
||||
- `GET /ws/voice/stream?language=&keyterms=` (WebSocket, not under `/api`)
|
||||
relays one dictation. Client sends binary frames of signed 16-bit
|
||||
little-endian PCM, 16 kHz mono (<= 64 KB per frame), plus JSON control frames
|
||||
`{"t":"finalize"}` (ask for the final transcript) and `{"t":"stop"}`. Server
|
||||
sends `{"t":"ready"}`, `{"t":"transcript","text","final"}` (each frame is the
|
||||
WHOLE running transcript, not a delta), `{"t":"error","message"}` and
|
||||
`{"t":"closed"}`. Close codes: `4003` disallowed Host/Origin, `4004`
|
||||
unavailable (reason in the close reason), `4008` too many concurrent streams.
|
||||
Streams are capped in count and length (`src/config/voice.ts`).
|
||||
|
||||
## Authentication
|
||||
|
||||
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
|
||||
`codeman_session` cookie. When enabled, unauthenticated requests get
|
||||
`401 UNAUTHORIZED`; rate-limited requests get `429 RATE_LIMITED`. See
|
||||
[`security-architecture.md`](security-architecture.md).
|
||||
|
||||
## SSE event channel
|
||||
|
||||
`GET /api/events` is a Server-Sent Events stream (`text/event-stream`); each
|
||||
message is `event: <name>` + `data: <json>`. The event-name registry
|
||||
(`src/web/sse-events.ts`, mirrored in `src/web/public/constants.js`) is part of
|
||||
the stable contract — event names are not renamed without a major bump. An
|
||||
optional `?sessions=<id,...>` filter suppresses only the high-volume terminal
|
||||
stream; lifecycle/metadata events are delivered to all clients regardless.
|
||||
|
||||
### `sse:heartbeat` (liveness)
|
||||
|
||||
Every 15s the server writes a `sse:heartbeat` frame to every connected client:
|
||||
|
||||
```
|
||||
event: sse:heartbeat
|
||||
data: {"t":1755100000000}
|
||||
```
|
||||
|
||||
`t` is the server's epoch-ms timestamp at write time. The frame carries no
|
||||
application state and can be ignored for correctness. It exists so a client can
|
||||
tell a live stream from a dead one: an `EventSource` whose connection has been
|
||||
idle-closed by a proxy (or that resumed from sleep on a stale socket) keeps
|
||||
delivering nothing without ever firing `onerror`. Clients that care should treat
|
||||
silence longer than about three intervals as a dead stream and reconnect, which
|
||||
is what the bundled frontend does.
|
||||
|
||||
This replaced a `:keepalive` SSE **comment**, which served the same
|
||||
proxy-flushing purpose but is invisible to `EventSource` by spec and so could
|
||||
never be observed by a client. Consumers written against the old behavior are
|
||||
unaffected: `EventSource` dispatches only events that have a registered
|
||||
listener, so an unknown event name is dropped.
|
||||
|
||||
## Consuming from JavaScript
|
||||
|
||||
The bundled frontend reads responses through `_apiJson()`
|
||||
(`src/web/public/api-client.js`), which unwraps `{success:true,data}` → `data` and
|
||||
returns `null` on a non-2xx / `{success:false}` response. External clients should
|
||||
do the same: check the HTTP status (or `body.success`), then read `body.data`.
|
||||
@@ -1,107 +0,0 @@
|
||||
# Approvals Inbox (design)
|
||||
|
||||
One cross-session inbox for every prompt that is waiting on a human: permission dialogs, questions (AskUserQuestion / elicitation), and idle prompts. Cards are answerable in place (option digits, Esc, or a typed prompt) from desktop, phone overview, and push notification action buttons. Inspired by Cloudflare OS's Gatekeeper approval queue (https://github.com/cloudflare/cloudflare-os, asynchronous human-in-the-loop approvals): with a fleet of sessions the human is the bottleneck, and today answering means finding the right tab.
|
||||
|
||||
## Problems this fixes (all real today)
|
||||
|
||||
1. **No cross-session surface.** Pending prompts exist only as per-tab alert colors (`tab-alert-action`/`tab-alert-idle`) and NEEDS YOU rows on the phone overview. Answering means switching to the session and typing.
|
||||
2. **Alerts die on reload.** `pendingHooks` lives only in `app.js` memory, fed by transient SSE `hook:*` events. A page reload (or a phone browser evicting the tab) silently loses every pending alert. There is no server-side record.
|
||||
3. **Push Approve/Deny buttons are dead.** `PUSH_EVENT_MAP` already attaches `approve`/`deny` actions to permission pushes, and `sw.js` forwards `event.action` to the page, but the `notification-click` handler in settings-ui.js ignores it (and when no tab is open, the action is dropped entirely). The buttons render on the lock screen and do nothing.
|
||||
4. **Card context is missing.** The frontend handlers read `data.question` / `data.message` / `data.tool`, but `sanitizeHookData` never forwards `message`, so notifications show generic fallback text.
|
||||
|
||||
## Scope
|
||||
|
||||
- Claude mode only (hooks fire only for `claude`; external CLIs keep their output-stabilization heuristics and get no inbox items). This mirrors the wait-primitive `stop`/`blocked` gating.
|
||||
- Permission prompts occur for sessions running `ClaudeMode` `normal` / `auto` / `allowedTools` (and the trust-folder dialog even under skip-permissions). Question and idle prompts occur in every mode including `dangerously-skip-permissions`.
|
||||
- In-memory store (plus the frontend seeding from it on load). Server restart drops items; hooks re-fire on the next prompt. No new state file in v1.
|
||||
|
||||
## Data model
|
||||
|
||||
At most **one active item per session**: the Claude TUI shows one dialog at a time, so a new prompt event supersedes the session's previous item (resolution `superseded`).
|
||||
|
||||
```ts
|
||||
interface ApprovalItem {
|
||||
id: string; // `${sessionId}:${seq}`
|
||||
sessionId: string;
|
||||
sessionName: string;
|
||||
kind: 'permission' | 'question' | 'idle';
|
||||
createdAt: number;
|
||||
toolName?: string; // from sanitized hook data
|
||||
toolSummary?: string; // command / file_path / description, already bounded
|
||||
message?: string; // Notification hook `message` (newly allowlisted)
|
||||
cwd?: string;
|
||||
context?: string; // ANSI-stripped visible pane frame tail, ≤ 4000 chars
|
||||
options?: { n: number; label: string }[]; // parsed from context when confident
|
||||
}
|
||||
```
|
||||
|
||||
Resolutions (server-emitted, item removed from pending): `answered` (via inbox), `resolved_in_terminal` (stop / elicitation_complete / elicitation_response / session went working), `superseded`, `session_ended`, `dismissed`, `expired` (12h TTL sweep).
|
||||
|
||||
## Backend
|
||||
|
||||
### Store: `src/approval-inbox.ts`
|
||||
|
||||
Module-level singleton in the style of `session-wait-registry.ts` (pure, no `Session` import, injected emit callback so there is no import cycle with the server):
|
||||
|
||||
- `notePrompt(info)` creates/supersedes the session's item; schedules ONE re-capture ~600ms later (the Notification hook can fire before the dialog finishes painting) which updates `context`/`options` and emits `approval:updated`.
|
||||
- `resolveForSession(sessionId, reason)`, `dismiss(id)`, `answerable(id)`, `listPending()`, `stop()` (clears timers; tests).
|
||||
- Option parsing (pure, unit-tested): consecutive `❯? N. label` lines, 2..6 options, labels ≤ 120 chars. Parsed options gate which digits the answer endpoint accepts; when parsing fails the card falls back to Approve(1)/Deny(Esc) only.
|
||||
- TTL: items expire after 12h (checked on read + a lazy sweep; no standing interval).
|
||||
|
||||
### Wiring
|
||||
|
||||
- `hook-event-routes.ts`: on `permission_prompt` / `elicitation_dialog` / `idle_prompt`, call `notePrompt` with sanitized data + a pane capture callback (`mux.capturePaneBuffer(muxName)` visible frame, ANSI-stripped via existing utils; fall back to `session.terminalBuffer` tail). On `stop` / `elicitation_complete` / `elicitation_response`, `resolveForSession(id, 'resolved_in_terminal')`.
|
||||
- `session-listener-wiring.ts`: `working` listener resolves **idle items only** (`working` is heuristic and can flap mid-turn, so it must never clear a pending permission/question dialog); `exit` resolves with `session_ended`. Same singleton-import pattern as `sessionWaits`.
|
||||
- Session delete route: resolve with `session_ended`.
|
||||
- **New hook matchers** `elicitation_complete` + `elicitation_response` added to `generateHooksConfig()`, `HookEventType`, `HookEventSchema`, and both SSE registries. `refreshStaleCodemanHooks` gets a staleness probe for them (`hooksJson.includes('elicitation_complete')`) so existing cases heal on next Claude spawn, exactly like the `-k`/secret/marker probes.
|
||||
- `sanitizeHookData`: allowlist `message` (bounded 500 chars). This also un-deadens the existing notification text paths.
|
||||
|
||||
### Routes: `src/web/routes/approval-routes.ts`
|
||||
|
||||
Normal authed API (NOT the hook-secret bypass), `ApiResponse` envelope, Zod schemas in `schemas.ts`:
|
||||
|
||||
- `GET /api/approvals` → pending items, multi-user filtered by `canAccessOwned` (same policy as session lists). Also sweeps the caller's own items for staleness through `verifyStillAnswerable()`: Claude Code fires no "permission answered" hook, so a dialog answered in the terminal used to sit pending until `stop` and re-arm a red tab alert on the next page load. Only items whose original frame parsed options can be dropped this way, so an unreadable capture keeps the alert.
|
||||
- `POST /api/approvals/:id/answer` body `{ action: 'approve' | 'deny' | 'option' | 'text', option?, text? }`:
|
||||
- `approve` → `writeViaMux('1')` (option 1 is always plain Yes; no Enter, menus react to the digit).
|
||||
- `deny` → `writeViaMux('\x1b')` (Esc is the official No/cancel; precedent: auto-resume sends Esc the same way).
|
||||
- `option` → digit `String(n)`; accepted only when `n` is within the item's parsed options (prevents blind digit-poking at an unparsed dialog).
|
||||
- `text` → `idle` items only: single line, embedded newlines stripped, sent as `text\r` (the `\r` discipline from CLAUDE.md).
|
||||
- Guards: item still pending (404 otherwise), session exists + ownership via `findSessionOrFail`, session mode installs hooks. **Answer-time re-capture**: for items whose frame parsed options, the pane is re-captured before sending; if the dialog no longer parses, the item resolves and the answer is refused with 409 (the keystroke would land in whatever now has focus). Marks `answered` BEFORE the write so a double-tap cannot double-send; rolls back to pending if the write fails.
|
||||
- `POST /api/approvals/:id/dismiss` → remove without keystrokes.
|
||||
- `POST /api/approvals/session/:sessionId/viewed` → acknowledge the session's pending **idle** item (`acknowledgedAt`, emitted as `approval:updated`). Added after the owner reported that a yellow tab clicked and checked went yellow again on reload: the view-clears-idle rule lived in one browser's memory, so the seed re-armed it and other devices never saw the clear. Acknowledgement is deliberately **not** resolution (the prompt is still unanswered, so it stays in the inbox and stays available as Read My Mind context), and deliberately **idle-only** (looking at a permission/question dialog does not answer it, so the red alert survives being viewed).
|
||||
|
||||
### SSE
|
||||
|
||||
`approval:pending`, `approval:updated`, `approval:resolved` in `sse-events.ts` + `SSE_EVENTS` in constants.js (the parity test pins the sync). Broadcasts carry `sessionId`, so multi-user SSE scoping applies unchanged.
|
||||
|
||||
### Push
|
||||
|
||||
- `sendPushNotifications` payload gains `approvalId` for the three hook events. Both `approvalId` and the Approve/Deny `actions` are **gated on the opt-in setting**: with it off, permission pushes carry no buttons at all (pre-inbox they rendered and did nothing, so stripping them is the honest shape).
|
||||
- `sw.js` `notificationclick`: when `event.action` is `approve`/`deny`, POST `/api/approvals/:id/answer` directly from the worker (same-origin, cookie credentials) so the buttons work **with no tab open**; on failure fall back to focusing/opening a tab. Non-action clicks keep today's behavior.
|
||||
- Page-side `notification-click` handler: honor `action` instead of dropping it (also setting-gated, for stale notifications sent before the toggle flipped).
|
||||
- Question/idle pushes keep no action buttons (options vary per dialog); tapping opens the inbox.
|
||||
|
||||
## Frontend
|
||||
|
||||
New module `approvals-ui.js` (@loadorder 11.2, after panels-ui.js), prettier-formatted (not added to `.prettierignore`).
|
||||
|
||||
- **Seed on connect**: `GET /api/approvals` on init and SSE reconnect; each pending item re-feeds `setPendingHook(...)` so tab alerts and the phone overview survive reload (fixes problem 2 with zero changes to the alert state machine). Items carrying `acknowledgedAt` are skipped, and `markIdleAlertSeen()` (app.js) is what sets it: viewing a session clears its yellow locally and POSTs `.../viewed`, so "I checked it" survives the reload and reaches the user's other devices through `approval:updated`.
|
||||
- **Desktop**: header bell `btn-approvals` with count badge. Ships default-hidden via marker class `btn-approvals--hidden` (same policy as the attachments button, so `test/mobile-header-buttons-policy.test.ts` excludes it from the default-visible enumeration); JS shows it only while count > 0. Click toggles a drawer of cards: session name + kind, tool/message summary, mono context block, buttons rendered from parsed options (else Approve/Deny), plus Dismiss and Open session. Esc closes; existing z-index layers respected.
|
||||
- **Phone**: header button stays hidden (`mobile.css`); the phone surface is the overview's NEEDS YOU section, whose rows gain inline ✓/✗ buttons for permission items (tap-through to the session remains the row's main action). Toolbar classes/status language rules from the mobile-overview section of CLAUDE.md apply.
|
||||
- **i18n**: new strings registered in i18n.js (en + zh-CN); status words carry `data-i18n-skip` where they would collide (mirroring the overview pills).
|
||||
- **Setting**: `approvalsInboxEnabled`, synced (in `SettingsUpdateSchema`), **default OFF** (owner decision: the entire feature is opt-in, meaning no bell, no drawer, no overview strips, no seeding, and no push action buttons until enabled in App Settings → Panels). Only the store and answer endpoints keep running regardless, so flipping the toggle ON surfaces anything already pending immediately, with no restart.
|
||||
|
||||
## Race honesty
|
||||
|
||||
The prompt can be answered in the terminal a moment before an inbox answer lands; then the keystroke would hit whatever now has focus (worst case: a digit typed into the composer, not submitted, since no `\r` is ever sent for menu answers). Mitigations, in order: answer-time re-capture (the dialog must still parse on screen or the answer is refused), answered-before-write marking, digit-only/Esc-only writes for menus, and the card's context block showing what the pane looked like when captured. This is the same class of risk `writeViaMux` automation (auto-resume, respawn) already accepts.
|
||||
|
||||
## Tests
|
||||
|
||||
- `test/approval-inbox.test.ts`: supersede per session, every resolution path, TTL, option parsing fixtures (2-option, 3-option with ❯, unparseable frame), re-capture update.
|
||||
- `test/routes/approval-routes.test.ts` (`app.inject`, no port): list; hook event creates item; answer approve/deny/option writes the exact bytes (test-PTY echo asserts them); text answers restricted to idle; 404 unknown id; 409 answered twice; option out of range rejected; multi-user scoping.
|
||||
- Existing suites extended: hook-event schema accepts the two new events; `sanitizeHookData` forwards bounded `message`; SSE parity + mobile-header policy pass as-is by construction.
|
||||
|
||||
## Docs
|
||||
|
||||
- CLAUDE.md: Key Patterns entry + SSE/route counts + frontend load order.
|
||||
- `docs/api-reference.md`: the two endpoints + three SSE events (additive, fine under the 0.9.x contract).
|
||||
@@ -1,331 +0,0 @@
|
||||
# Plan: Background Keystroke Forwarding (Local Echo Mode)
|
||||
|
||||
> **Supersedes**: This document merges two previous plan drafts into a single authoritative reference:
|
||||
> - `docs/background-keystroke-forwarding-plan.md` (detailed design doc)
|
||||
> - `.claude/plans/jazzy-bubbling-salamander.md` (Claude-generated implementation plan)
|
||||
>
|
||||
> The docs plan was used as the base. The Claude plan was a correct but simplified subset; its Context paragraph is incorporated below as a lead-in.
|
||||
|
||||
## Context
|
||||
|
||||
When local echo is enabled, keystrokes accumulate in the `LocalEchoOverlay.pendingText` and are only sent to the server when Enter is pressed. This means switching tabs loses the input from the actual Claude Code PTY (the overlay caches text client-side, but the PTY has nothing). If the session respawns or resets, accumulated input is lost entirely.
|
||||
|
||||
## Problem
|
||||
|
||||
When local echo is enabled, keystrokes accumulate **only** in `LocalEchoOverlay.pendingText` (a client-side string). Nothing reaches the server PTY until Enter is pressed. This creates three failure modes:
|
||||
|
||||
1. **Tab switch loses PTY state** — switching sessions saves overlay text to `localEchoTextCache` (a Map), but the actual Claude Code Ink process has no knowledge of what was typed. If respawn or `/clear` fires on that session, the cached text is meaningless.
|
||||
2. **Session death loses input** — if the session crashes or respawns while text is pending in the overlay, that input is gone (localStorage backup `codeman_local_echo_pending` only survives page reloads, not session resets).
|
||||
3. **Tab completion impossible** — pressing Tab with pending overlay text sends the raw Tab character to a PTY that has no knowledge of the typed text, so completion fails.
|
||||
|
||||
## Goal
|
||||
|
||||
Send every keystroke to the server in the background (debounced), while the overlay continues providing instant visual feedback. The overlay sits at z-index 7 with an opaque background over `.xterm-screen`, masking Ink's echo of the background-sent characters. Input persists in the actual Claude Code readline buffer across tab switches and respawns.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
User keystroke
|
||||
|
|
||||
v
|
||||
xterm.js onData(data)
|
||||
|
|
||||
+---> LocalEchoOverlay.addChar(data) [instant visual feedback]
|
||||
|
|
||||
+---> _localEchoBgBuffer += data [queue for background send]
|
||||
| clearTimeout + setTimeout(50ms)
|
||||
| |
|
||||
| v (50ms debounce fires)
|
||||
| _flushBgInput()
|
||||
| |
|
||||
| v
|
||||
| _sendInputAsync(sessionId, buffer) [promise chain preserves order]
|
||||
| |
|
||||
| v
|
||||
| POST /api/sessions/:id/input [{ input: "hel" }]
|
||||
| |
|
||||
| v
|
||||
| session.write(inputStr) [direct PTY write, synchronous]
|
||||
| |
|
||||
| v
|
||||
| Ink readline echoes "hel" [hidden behind overlay's opaque bg]
|
||||
|
|
||||
+--- On Enter:
|
||||
1. clearTimeout(_localEchoBgTimer)
|
||||
2. flush _localEchoBgBuffer via _sendInputAsync (remaining chars)
|
||||
3. clear overlay
|
||||
4. 120ms later: send \r via _sendInputAsync (Ink text/Enter split)
|
||||
5. Ink processes "hello\r" → overlay gone, terminal visible with output
|
||||
```
|
||||
|
||||
### Two Input Paths (important context)
|
||||
|
||||
The codebase has **two separate input paths** to the server:
|
||||
|
||||
| Path | Used by | Promise chain? | `useMux`? |
|
||||
|------|---------|---------------|-----------|
|
||||
| `_sendInputAsync()` (line 3626) | `onData` handler, `flushInput()` | Yes (`_inputSendChain`) | No (direct PTY write) |
|
||||
| `sendInput()` (line 8755) | Mobile accessory bar, programmatic commands | **No** (raw `fetch`) | Yes (tmux `send-keys`) |
|
||||
|
||||
Background keystroke forwarding uses **only** the `_sendInputAsync` path, which guarantees ordering via the promise chain. The `sendInput()` path is unaffected and unmodified.
|
||||
|
||||
### Server-Side Input Flow
|
||||
|
||||
```
|
||||
POST /api/sessions/:id/input { input: "hel" }
|
||||
|
|
||||
+-- useMux? No (default)
|
||||
| session.write("hel") → ptyProcess.write("hel") [sync]
|
||||
|
|
||||
+-- useMux? Yes
|
||||
session.writeViaMux("hel") → tmux send-keys -l "hel" [async]
|
||||
```
|
||||
|
||||
Background sends use the default path (no `useMux`), which is a synchronous direct PTY write — faster than spawning a tmux subprocess for each character batch.
|
||||
|
||||
## Implementation
|
||||
|
||||
All changes in **one file**: `src/web/public/app.js`
|
||||
|
||||
### Step 1: Add background send state (in terminal setup, after line ~1999)
|
||||
|
||||
```js
|
||||
this._localEchoBgBuffer = ''; // Characters queued for background send
|
||||
this._localEchoBgTimer = null; // 50ms debounce timer ID
|
||||
```
|
||||
|
||||
Add an atomic drain helper alongside existing `flushInput` (after line ~2008):
|
||||
|
||||
```js
|
||||
// Atomically drain background buffer — returns contents and cancels pending timer.
|
||||
// Single point of extraction prevents double-flush race conditions.
|
||||
const drainBgBuffer = () => {
|
||||
if (this._localEchoBgTimer) {
|
||||
clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = null;
|
||||
}
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
return buf;
|
||||
};
|
||||
|
||||
const scheduleBgFlush = () => {
|
||||
if (this._localEchoBgTimer) clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = setTimeout(() => {
|
||||
this._localEchoBgTimer = null;
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
if (buf && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, buf);
|
||||
}
|
||||
}, 50);
|
||||
};
|
||||
```
|
||||
|
||||
**Why `drainBgBuffer` exists**: Every exit path (Enter, Ctrl+C, tab switch, echo disable) needs to flush the buffer AND cancel the timer atomically. Without a single extraction point, it's easy to forget one of the two operations, leading to double-sends when the timer fires after a manual flush.
|
||||
|
||||
### Step 2: Modify `onData` handler — local echo path (lines 2023–2067)
|
||||
|
||||
**Printable characters** (lines 2063–2067 → replace):
|
||||
```js
|
||||
if (data.length === 1 && data.charCodeAt(0) >= 32) {
|
||||
this._localEchoOverlay?.addChar(data);
|
||||
// Background: queue char for server send (50ms debounce batches rapid typing)
|
||||
this._localEchoBgBuffer += data;
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Backspace** (lines 2024–2028 → replace):
|
||||
```js
|
||||
if (data === '\x7f') {
|
||||
this._localEchoOverlay?.removeChar();
|
||||
// Background: queue DEL for server (Ink's readline handles backspace via \x7f)
|
||||
this._localEchoBgBuffer += '\x7f';
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Enter** (lines 2029–2050 → replace):
|
||||
```js
|
||||
if (/^[\r\n]+$/.test(data)) {
|
||||
this._localEchoOverlay?.clear();
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
// Flush any remaining background chars (e.g., last 50ms batch not yet sent)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send \r after 120ms — Ink needs text and Enter as separate events.
|
||||
// The promise chain in _sendInputAsync guarantees the remaining chars
|
||||
// are dispatched before \r, regardless of timing.
|
||||
setTimeout(() => {
|
||||
this._pendingInput += '\r';
|
||||
flushInput();
|
||||
}, 120);
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Key change from original plan**: The Enter handler no longer checks `if (text)` and branches on whether the overlay had content. With background sends, the PTY already has most/all of the text. We just flush any remainder and unconditionally send `\r` after 120ms. This simplifies the flow and handles edge cases like "user typed nothing but pressed Enter" (remainder is empty, just `\r` is sent).
|
||||
|
||||
**Control characters and paste** (lines 2052–2061 → replace):
|
||||
```js
|
||||
if (data.charCodeAt(0) < 32 || data.length > 1) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush background buffer so PTY has full text state before control char
|
||||
// (critical for Tab completion — PTY needs typed text to complete against)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send control char / paste text via normal path
|
||||
this._pendingInput += data;
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
flushInput();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Note on paste**: Desktop paste arrives via `onData` as a single multi-character string (`data.length > 1`). This falls into the control char path above, which:
|
||||
1. Clears the overlay (existing behavior)
|
||||
2. Flushes background buffer (new — ensures PTY has prefix text)
|
||||
3. Sends paste text immediately (existing behavior)
|
||||
|
||||
Mobile paste via `KeyboardAccessoryBar.pasteFromClipboard()` uses `app.sendInput()` which bypasses `onData` entirely — no change needed.
|
||||
|
||||
### Step 3: Flush on tab switch (`selectSession()`, line ~4533)
|
||||
|
||||
Insert before the existing overlay save/clear block (before line 4534):
|
||||
|
||||
```js
|
||||
// Flush background send buffer for outgoing session
|
||||
if (this.activeSessionId) {
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This ensures the PTY receives all typed characters before the tab switch. When the user switches back, the terminal buffer will show the text (echoed by Ink) and the overlay will restore its cached copy on top.
|
||||
|
||||
### Step 4: Cleanup on local echo disable (`_updateLocalEchoState()`, lines 2362–2371)
|
||||
|
||||
Expand the disable transition (line 2367–2368):
|
||||
```js
|
||||
if (this._localEchoEnabled && !shouldEnable) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush any pending background chars before disabling
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Step 5: Cleanup on session delete (`deleteSession()`)
|
||||
|
||||
When a session is deleted, cancel any pending background timer for that session:
|
||||
```js
|
||||
// In deleteSession(), after removing the session from this.sessions:
|
||||
drainBgBuffer(); // Discard — session is gone, nowhere to send
|
||||
this.localEchoTextCache.delete(sessionId);
|
||||
```
|
||||
|
||||
## Visual Timeline
|
||||
|
||||
```
|
||||
t=0ms User types "h" → overlay: "h" bgBuffer: "h" timer: 50ms
|
||||
t=30ms User types "e" → overlay: "he" bgBuffer: "he" timer: reset 50ms
|
||||
t=60ms User types "l" → overlay: "hel" bgBuffer: "hel" timer: reset 50ms
|
||||
t=110ms Debounce fires → overlay: "hel" bgBuffer: "" POST "hel" → PTY
|
||||
t=115ms Ink echoes "hel" → terminal: "❯ hel" (hidden behind overlay)
|
||||
t=140ms User types "l" → overlay: "hell" bgBuffer: "l" timer: 50ms
|
||||
t=170ms User types "o" → overlay: "hello" bgBuffer: "lo" timer: reset 50ms
|
||||
t=220ms Debounce fires → overlay: "hello" bgBuffer: "" POST "lo" → PTY
|
||||
t=250ms User hits Enter → drainBgBuffer()="" overlay: cleared
|
||||
t=370ms \r sent via chain → Ink processes "hello\r" → output appears
|
||||
```
|
||||
|
||||
**Tab switch scenario:**
|
||||
```
|
||||
t=0ms User types "wor" → overlay: "wor" bgBuffer: "wor" timer: 50ms
|
||||
t=25ms User switches tab → drainBgBuffer() sends "wor" to old session PTY
|
||||
overlay text "wor" saved to localEchoTextCache
|
||||
overlay cleared, new session loaded
|
||||
...later...
|
||||
t=5000ms User switches back → terminal shows "❯ wor" (Ink echo from background send)
|
||||
overlay restores "wor" from cache, masks terminal
|
||||
user continues typing seamlessly
|
||||
```
|
||||
|
||||
## Edge Cases & Mitigations
|
||||
|
||||
### Confirmed Safe (JS single-threaded guarantee)
|
||||
|
||||
| Scenario | Why it's safe |
|
||||
|----------|--------------|
|
||||
| **Debounce fires during Enter handler** | Impossible. JS event loop is single-threaded — the Enter handler runs atomically. `drainBgBuffer()` cancels the timer before it can fire. |
|
||||
| **Debounce fires during tab switch** | Same reason. `selectSession()` calls `drainBgBuffer()` synchronously, canceling the timer. |
|
||||
| **Double-send of background buffer** | `drainBgBuffer()` atomically clears both buffer and timer. Once drained, subsequent drain returns empty string. |
|
||||
| **`_pendingInput` conflict** | In local echo mode, `_pendingInput` is only used for Enter (`\r`) and control chars. Background chars use a separate `_localEchoBgBuffer`. No overlap. |
|
||||
|
||||
### Handled by Design
|
||||
|
||||
| Scenario | Handling |
|
||||
|----------|---------|
|
||||
| **Rapid typing / paste** | 50ms debounce batches rapid chars. At 100 WPM (~50ms/char), sends ~1 char per batch. For paste (multi-char string, `data.length > 1`), the control char path bypasses the buffer entirely and sends immediately. |
|
||||
| **Network failure** | `_sendInputAsync` catches fetch failures and calls `_enqueueInput()` for retry. `_drainInputQueues()` replays on reconnect. Background chars use the same retry path. |
|
||||
| **Offline mode** | `_sendInputAsync` checks `this.isOnline` and immediately enqueues if offline. Same behavior for background sends. 64KB queue cap prevents memory growth. |
|
||||
| **Tab completion** | Ctrl+Tab path flushes background buffer BEFORE sending Tab char. PTY has full text for readline completion. |
|
||||
| **Session respawn** | PTY already has typed text (sent in background). On respawn, Claude exits and restarts — Ink's readline buffer is lost, but the text was already processed or is no longer relevant. The overlay clears on session status change via `_updateLocalEchoState()`. |
|
||||
| **SSE reconnect** | `handleInit()` saves overlay text before `selectSession()` clears it, then restores after reload (line ~3945–3966). Background buffer is cleared on reconnect since state is reset. |
|
||||
|
||||
### Network Ordering
|
||||
|
||||
**Question**: Can background sends arrive at the server out of order?
|
||||
|
||||
**Answer**: No, for practical purposes.
|
||||
|
||||
1. `_sendInputAsync` uses a **promise chain** (`_inputSendChain`) — each fetch is dispatched only after the previous one has been dispatched. This means requests are sent in order.
|
||||
2. Localhost connections (HTTP/1.1) are inherently sequential on a single TCP connection.
|
||||
3. Even with HTTP/2 multiplexing, Fastify (Node.js) is single-threaded — request handlers execute via the event loop in arrival order.
|
||||
4. The server's `session.write()` is synchronous — it writes to the PTY immediately within the request handler.
|
||||
|
||||
### Known Limitations (Not Addressed)
|
||||
|
||||
| Limitation | Impact | Notes |
|
||||
|-----------|--------|-------|
|
||||
| **IME composition** | CJK input via IME would send partial composition sequences to PTY | No IME handling exists in the codebase today (line count: 0 references to `compositionstart/end/update`). Fixing this is a separate feature. |
|
||||
| **`sendInput()` ordering** | Mobile accessory bar commands (`/init`, `/clear`, paste) use `sendInput()` which bypasses `_inputSendChain` — no ordering guarantee relative to background sends | Unlikely to conflict in practice: accessory bar clears the overlay first, and the commands are typically sent when no typing is in progress. |
|
||||
| **localStorage stale text** | After background sends, localStorage still has overlay text. On hard reload, overlay restores text that the PTY already has → visual duplicate behind overlay | Harmless — overlay masks the terminal. On Enter, overlay clears and terminal shows correct state. Could be fixed by clearing localStorage after successful background flush, but adds complexity for minimal benefit. |
|
||||
|
||||
## Verification Checklist
|
||||
|
||||
1. **Basic typing**: Enable local echo → type "hello" → overlay shows instantly → check Network tab for batched POST requests (~50ms intervals) → press Enter → command executes
|
||||
2. **Tab switch persistence**: Type "test" → switch to another tab → switch back → text visible in both overlay AND terminal prompt
|
||||
3. **Backspace**: Type "helloo" → press backspace → overlay shows "hello" → check PTY received \x7f
|
||||
4. **Paste**: Type "hel" → paste "lo world" → overlay clears → "lo world" sent immediately → PTY has "hello world"
|
||||
5. **Tab completion**: Type "src/w" → press Tab → PTY completes to "src/web/" (background send gave PTY the prefix)
|
||||
6. **Ctrl+C**: Type "hello" → press Ctrl+C → overlay clears → PTY receives pending chars + \x03
|
||||
7. **Network tab**: Verify POST /api/sessions/:id/input requests appear as you type (batched ~50ms)
|
||||
8. **Offline resilience**: Disconnect network → type "hello" → reconnect → verify chars are replayed via drain queue
|
||||
9. **Session delete**: Type text → delete session → no console errors from orphaned timer
|
||||
10. **Mobile keyboard**: Test on mobile device — typing goes through same onData path, same behavior expected
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/web/public/app.js` | ~40 lines changed across 5 locations (Steps 1–5) |
|
||||
|
||||
No server-side changes. No new files. No new dependencies.
|
||||
@@ -1,321 +0,0 @@
|
||||
# Plan: Background Keystroke Forwarding (Local Echo Mode)
|
||||
|
||||
## Problem
|
||||
|
||||
When local echo is enabled, keystrokes accumulate **only** in `LocalEchoOverlay.pendingText` (a client-side string). Nothing reaches the server PTY until Enter is pressed. This creates three failure modes:
|
||||
|
||||
1. **Tab switch loses PTY state** — switching sessions saves overlay text to `localEchoTextCache` (a Map), but the actual Claude Code Ink process has no knowledge of what was typed. If respawn or `/clear` fires on that session, the cached text is meaningless.
|
||||
2. **Session death loses input** — if the session crashes or respawns while text is pending in the overlay, that input is gone (localStorage backup `codeman_local_echo_pending` only survives page reloads, not session resets).
|
||||
3. **Tab completion impossible** — pressing Tab with pending overlay text sends the raw Tab character to a PTY that has no knowledge of the typed text, so completion fails.
|
||||
|
||||
## Goal
|
||||
|
||||
Send every keystroke to the server in the background (debounced), while the overlay continues providing instant visual feedback. The overlay sits at z-index 7 with an opaque background over `.xterm-screen`, masking Ink's echo of the background-sent characters. Input persists in the actual Claude Code readline buffer across tab switches and respawns.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
User keystroke
|
||||
|
|
||||
v
|
||||
xterm.js onData(data)
|
||||
|
|
||||
+---> LocalEchoOverlay.addChar(data) [instant visual feedback]
|
||||
|
|
||||
+---> _localEchoBgBuffer += data [queue for background send]
|
||||
| clearTimeout + setTimeout(50ms)
|
||||
| |
|
||||
| v (50ms debounce fires)
|
||||
| _flushBgInput()
|
||||
| |
|
||||
| v
|
||||
| _sendInputAsync(sessionId, buffer) [promise chain preserves order]
|
||||
| |
|
||||
| v
|
||||
| POST /api/sessions/:id/input [{ input: "hel" }]
|
||||
| |
|
||||
| v
|
||||
| session.write(inputStr) [direct PTY write, synchronous]
|
||||
| |
|
||||
| v
|
||||
| Ink readline echoes "hel" [hidden behind overlay's opaque bg]
|
||||
|
|
||||
+--- On Enter:
|
||||
1. clearTimeout(_localEchoBgTimer)
|
||||
2. flush _localEchoBgBuffer via _sendInputAsync (remaining chars)
|
||||
3. clear overlay
|
||||
4. 120ms later: send \r via _sendInputAsync (Ink text/Enter split)
|
||||
5. Ink processes "hello\r" → overlay gone, terminal visible with output
|
||||
```
|
||||
|
||||
### Two Input Paths (important context)
|
||||
|
||||
The codebase has **two separate input paths** to the server:
|
||||
|
||||
| Path | Used by | Promise chain? | `useMux`? |
|
||||
|------|---------|---------------|-----------|
|
||||
| `_sendInputAsync()` (line 3626) | `onData` handler, `flushInput()` | Yes (`_inputSendChain`) | No (direct PTY write) |
|
||||
| `sendInput()` (line 8755) | Mobile accessory bar, programmatic commands | **No** (raw `fetch`) | Yes (tmux `send-keys`) |
|
||||
|
||||
Background keystroke forwarding uses **only** the `_sendInputAsync` path, which guarantees ordering via the promise chain. The `sendInput()` path is unaffected and unmodified.
|
||||
|
||||
### Server-Side Input Flow
|
||||
|
||||
```
|
||||
POST /api/sessions/:id/input { input: "hel" }
|
||||
|
|
||||
+-- useMux? No (default)
|
||||
| session.write("hel") → ptyProcess.write("hel") [sync]
|
||||
|
|
||||
+-- useMux? Yes
|
||||
session.writeViaMux("hel") → tmux send-keys -l "hel" [async]
|
||||
```
|
||||
|
||||
Background sends use the default path (no `useMux`), which is a synchronous direct PTY write — faster than spawning a tmux subprocess for each character batch.
|
||||
|
||||
## Implementation
|
||||
|
||||
All changes in **one file**: `src/web/public/app.js`
|
||||
|
||||
### Step 1: Add background send state (in terminal setup, after line ~1999)
|
||||
|
||||
```js
|
||||
this._localEchoBgBuffer = ''; // Characters queued for background send
|
||||
this._localEchoBgTimer = null; // 50ms debounce timer ID
|
||||
```
|
||||
|
||||
Add an atomic drain helper alongside existing `flushInput` (after line ~2008):
|
||||
|
||||
```js
|
||||
// Atomically drain background buffer — returns contents and cancels pending timer.
|
||||
// Single point of extraction prevents double-flush race conditions.
|
||||
const drainBgBuffer = () => {
|
||||
if (this._localEchoBgTimer) {
|
||||
clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = null;
|
||||
}
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
return buf;
|
||||
};
|
||||
|
||||
const scheduleBgFlush = () => {
|
||||
if (this._localEchoBgTimer) clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = setTimeout(() => {
|
||||
this._localEchoBgTimer = null;
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
if (buf && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, buf);
|
||||
}
|
||||
}, 50);
|
||||
};
|
||||
```
|
||||
|
||||
**Why `drainBgBuffer` exists**: Every exit path (Enter, Ctrl+C, tab switch, echo disable) needs to flush the buffer AND cancel the timer atomically. Without a single extraction point, it's easy to forget one of the two operations, leading to double-sends when the timer fires after a manual flush.
|
||||
|
||||
### Step 2: Modify `onData` handler — local echo path (lines 2023–2067)
|
||||
|
||||
**Printable characters** (lines 2063–2067 → replace):
|
||||
```js
|
||||
if (data.length === 1 && data.charCodeAt(0) >= 32) {
|
||||
this._localEchoOverlay?.addChar(data);
|
||||
// Background: queue char for server send (50ms debounce batches rapid typing)
|
||||
this._localEchoBgBuffer += data;
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Backspace** (lines 2024–2028 → replace):
|
||||
```js
|
||||
if (data === '\x7f') {
|
||||
this._localEchoOverlay?.removeChar();
|
||||
// Background: queue DEL for server (Ink's readline handles backspace via \x7f)
|
||||
this._localEchoBgBuffer += '\x7f';
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Enter** (lines 2029–2050 → replace):
|
||||
```js
|
||||
if (/^[\r\n]+$/.test(data)) {
|
||||
this._localEchoOverlay?.clear();
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
// Flush any remaining background chars (e.g., last 50ms batch not yet sent)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send \r after 120ms — Ink needs text and Enter as separate events.
|
||||
// The promise chain in _sendInputAsync guarantees the remaining chars
|
||||
// are dispatched before \r, regardless of timing.
|
||||
setTimeout(() => {
|
||||
this._pendingInput += '\r';
|
||||
flushInput();
|
||||
}, 120);
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Key change from original plan**: The Enter handler no longer checks `if (text)` and branches on whether the overlay had content. With background sends, the PTY already has most/all of the text. We just flush any remainder and unconditionally send `\r` after 120ms. This simplifies the flow and handles edge cases like "user typed nothing but pressed Enter" (remainder is empty, just `\r` is sent).
|
||||
|
||||
**Control characters and paste** (lines 2052–2061 → replace):
|
||||
```js
|
||||
if (data.charCodeAt(0) < 32 || data.length > 1) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush background buffer so PTY has full text state before control char
|
||||
// (critical for Tab completion — PTY needs typed text to complete against)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send control char / paste text via normal path
|
||||
this._pendingInput += data;
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
flushInput();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Note on paste**: Desktop paste arrives via `onData` as a single multi-character string (`data.length > 1`). This falls into the control char path above, which:
|
||||
1. Clears the overlay (existing behavior)
|
||||
2. Flushes background buffer (new — ensures PTY has prefix text)
|
||||
3. Sends paste text immediately (existing behavior)
|
||||
|
||||
Mobile paste via `KeyboardAccessoryBar.pasteFromClipboard()` uses `app.sendInput()` which bypasses `onData` entirely — no change needed.
|
||||
|
||||
### Step 3: Flush on tab switch (`selectSession()`, line ~4533)
|
||||
|
||||
Insert before the existing overlay save/clear block (before line 4534):
|
||||
|
||||
```js
|
||||
// Flush background send buffer for outgoing session
|
||||
if (this.activeSessionId) {
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This ensures the PTY receives all typed characters before the tab switch. When the user switches back, the terminal buffer will show the text (echoed by Ink) and the overlay will restore its cached copy on top.
|
||||
|
||||
### Step 4: Cleanup on local echo disable (`_updateLocalEchoState()`, lines 2362–2371)
|
||||
|
||||
Expand the disable transition (line 2367–2368):
|
||||
```js
|
||||
if (this._localEchoEnabled && !shouldEnable) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush any pending background chars before disabling
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Step 5: Cleanup on session delete (`deleteSession()`)
|
||||
|
||||
When a session is deleted, cancel any pending background timer for that session:
|
||||
```js
|
||||
// In deleteSession(), after removing the session from this.sessions:
|
||||
drainBgBuffer(); // Discard — session is gone, nowhere to send
|
||||
this.localEchoTextCache.delete(sessionId);
|
||||
```
|
||||
|
||||
## Visual Timeline
|
||||
|
||||
```
|
||||
t=0ms User types "h" → overlay: "h" bgBuffer: "h" timer: 50ms
|
||||
t=30ms User types "e" → overlay: "he" bgBuffer: "he" timer: reset 50ms
|
||||
t=60ms User types "l" → overlay: "hel" bgBuffer: "hel" timer: reset 50ms
|
||||
t=110ms Debounce fires → overlay: "hel" bgBuffer: "" POST "hel" → PTY
|
||||
t=115ms Ink echoes "hel" → terminal: "❯ hel" (hidden behind overlay)
|
||||
t=140ms User types "l" → overlay: "hell" bgBuffer: "l" timer: 50ms
|
||||
t=170ms User types "o" → overlay: "hello" bgBuffer: "lo" timer: reset 50ms
|
||||
t=220ms Debounce fires → overlay: "hello" bgBuffer: "" POST "lo" → PTY
|
||||
t=250ms User hits Enter → drainBgBuffer()="" overlay: cleared
|
||||
t=370ms \r sent via chain → Ink processes "hello\r" → output appears
|
||||
```
|
||||
|
||||
**Tab switch scenario:**
|
||||
```
|
||||
t=0ms User types "wor" → overlay: "wor" bgBuffer: "wor" timer: 50ms
|
||||
t=25ms User switches tab → drainBgBuffer() sends "wor" to old session PTY
|
||||
overlay text "wor" saved to localEchoTextCache
|
||||
overlay cleared, new session loaded
|
||||
...later...
|
||||
t=5000ms User switches back → terminal shows "❯ wor" (Ink echo from background send)
|
||||
overlay restores "wor" from cache, masks terminal
|
||||
user continues typing seamlessly
|
||||
```
|
||||
|
||||
## Edge Cases & Mitigations
|
||||
|
||||
### Confirmed Safe (JS single-threaded guarantee)
|
||||
|
||||
| Scenario | Why it's safe |
|
||||
|----------|--------------|
|
||||
| **Debounce fires during Enter handler** | Impossible. JS event loop is single-threaded — the Enter handler runs atomically. `drainBgBuffer()` cancels the timer before it can fire. |
|
||||
| **Debounce fires during tab switch** | Same reason. `selectSession()` calls `drainBgBuffer()` synchronously, canceling the timer. |
|
||||
| **Double-send of background buffer** | `drainBgBuffer()` atomically clears both buffer and timer. Once drained, subsequent drain returns empty string. |
|
||||
| **`_pendingInput` conflict** | In local echo mode, `_pendingInput` is only used for Enter (`\r`) and control chars. Background chars use a separate `_localEchoBgBuffer`. No overlap. |
|
||||
|
||||
### Handled by Design
|
||||
|
||||
| Scenario | Handling |
|
||||
|----------|---------|
|
||||
| **Rapid typing / paste** | 50ms debounce batches rapid chars. At 100 WPM (~50ms/char), sends ~1 char per batch. For paste (multi-char string, `data.length > 1`), the control char path bypasses the buffer entirely and sends immediately. |
|
||||
| **Network failure** | `_sendInputAsync` catches fetch failures and calls `_enqueueInput()` for retry. `_drainInputQueues()` replays on reconnect. Background chars use the same retry path. |
|
||||
| **Offline mode** | `_sendInputAsync` checks `this.isOnline` and immediately enqueues if offline. Same behavior for background sends. 64KB queue cap prevents memory growth. |
|
||||
| **Tab completion** | Ctrl+Tab path flushes background buffer BEFORE sending Tab char. PTY has full text for readline completion. |
|
||||
| **Session respawn** | PTY already has typed text (sent in background). On respawn, Claude exits and restarts — Ink's readline buffer is lost, but the text was already processed or is no longer relevant. The overlay clears on session status change via `_updateLocalEchoState()`. |
|
||||
| **SSE reconnect** | `handleInit()` saves overlay text before `selectSession()` clears it, then restores after reload (line ~3945–3966). Background buffer is cleared on reconnect since state is reset. |
|
||||
|
||||
### Network Ordering
|
||||
|
||||
**Question**: Can background sends arrive at the server out of order?
|
||||
|
||||
**Answer**: No, for practical purposes.
|
||||
|
||||
1. `_sendInputAsync` uses a **promise chain** (`_inputSendChain`) — each fetch is dispatched only after the previous one has been dispatched. This means requests are sent in order.
|
||||
2. Localhost connections (HTTP/1.1) are inherently sequential on a single TCP connection.
|
||||
3. Even with HTTP/2 multiplexing, Fastify (Node.js) is single-threaded — request handlers execute via the event loop in arrival order.
|
||||
4. The server's `session.write()` is synchronous — it writes to the PTY immediately within the request handler.
|
||||
|
||||
### Known Limitations (Not Addressed)
|
||||
|
||||
| Limitation | Impact | Notes |
|
||||
|-----------|--------|-------|
|
||||
| **IME composition** | CJK input via IME would send partial composition sequences to PTY | No IME handling exists in the codebase today (line count: 0 references to `compositionstart/end/update`). Fixing this is a separate feature. |
|
||||
| **`sendInput()` ordering** | Mobile accessory bar commands (`/init`, `/clear`, paste) use `sendInput()` which bypasses `_inputSendChain` — no ordering guarantee relative to background sends | Unlikely to conflict in practice: accessory bar clears the overlay first, and the commands are typically sent when no typing is in progress. |
|
||||
| **localStorage stale text** | After background sends, localStorage still has overlay text. On hard reload, overlay restores text that the PTY already has → visual duplicate behind overlay | Harmless — overlay masks the terminal. On Enter, overlay clears and terminal shows correct state. Could be fixed by clearing localStorage after successful background flush, but adds complexity for minimal benefit. |
|
||||
|
||||
## Verification Checklist
|
||||
|
||||
1. **Basic typing**: Enable local echo → type "hello" → overlay shows instantly → check Network tab for batched POST requests (~50ms intervals) → press Enter → command executes
|
||||
2. **Tab switch persistence**: Type "test" → switch to another tab → switch back → text visible in both overlay AND terminal prompt
|
||||
3. **Backspace**: Type "helloo" → press backspace → overlay shows "hello" → check PTY received \x7f
|
||||
4. **Paste**: Type "hel" → paste "lo world" → overlay clears → "lo world" sent immediately → PTY has "hello world"
|
||||
5. **Tab completion**: Type "src/w" → press Tab → PTY completes to "src/web/" (background send gave PTY the prefix)
|
||||
6. **Ctrl+C**: Type "hello" → press Ctrl+C → overlay clears → PTY receives pending chars + \x03
|
||||
7. **Network tab**: Verify POST /api/sessions/:id/input requests appear as you type (batched ~50ms)
|
||||
8. **Offline resilience**: Disconnect network → type "hello" → reconnect → verify chars are replayed via drain queue
|
||||
9. **Session delete**: Type text → delete session → no console errors from orphaned timer
|
||||
10. **Mobile keyboard**: Test on mobile device — typing goes through same onData path, same behavior expected
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/web/public/app.js` | ~40 lines changed across 5 locations (Steps 1–5) |
|
||||
|
||||
No server-side changes. No new files. No new dependencies.
|
||||
@@ -1,111 +0,0 @@
|
||||
> **⚠️ ARCHIVED 2026-05-21 — superseded, kept for history.**
|
||||
> The headline items here were verified resolved: the P0 `{WORKING_DIR}` placeholder
|
||||
> is now replaced (`plan-orchestrator.ts:431`), and the "~66 dead functions in app.js"
|
||||
> are gone (app.js was modularized 15K→3K LOC). A fresh `npm run knip` sweep on
|
||||
> 2026-05-21 found only a handful of unused test helpers. Do not treat this as a live TODO.
|
||||
|
||||
# Codebase Cleanup Findings
|
||||
|
||||
Compiled from parallel analysis of the entire Codeman codebase by 3 research agents (2026-02-19).
|
||||
|
||||
## P0 — Bug Fix
|
||||
|
||||
### 1. `{WORKING_DIR}` placeholder never replaced in plan-orchestrator.ts
|
||||
- **File:** `src/plan-orchestrator.ts:409`
|
||||
- `RESEARCH_AGENT_PROMPT` has `{WORKING_DIR}` placeholder but only `{TASK}` is replaced
|
||||
- The literal string `{WORKING_DIR}` gets sent to the AI model
|
||||
- **Fix:** Add `.replace('{WORKING_DIR}', this.workingDir)` after the `{TASK}` replacement
|
||||
|
||||
## P1 — Dead Code Removal (High Impact)
|
||||
|
||||
### 2. ~66 dead functions in app.js
|
||||
- Functions never called: `clearAll()`, `toggleSubagentDropdown()`, `goHome()`, `showRalphWizard()`, `minimizeRalphWizard()`, `restoreRalphWizard()`, `ralphWizardNext()`, `ralphWizardBack()`, `skipPlanGeneration()`, `regeneratePlan()`, `incrementTabCount()`, `decrementTabCount()`, `incrementShellCount()`, `decrementShellCount()`, `stopClaude()`, and ~50 more
|
||||
- Many are remnants of abandoned features (Ralph wizard, plan version history)
|
||||
- **Estimated savings:** 300-500 lines
|
||||
|
||||
### 3. ~74 dead CSS selectors in styles.css
|
||||
- Major dead blocks: Task Panel System (`.task-panel`), Process Panel System (`.process-panel`), Monitor Tabs (`.monitor-tabs`), Ralph Metadata (`.ralph-progress-section`, `.ralph-meta`), Plan Editor Toolbar, Plan Version History
|
||||
- Plus ~30 minor unused utility/component selectors
|
||||
- **Estimated savings:** ~400 lines
|
||||
|
||||
### 4. 13 dead type definitions in types.ts (~150 lines)
|
||||
- Dead request interfaces (superseded by Zod schemas): `CreateSessionRequest`, `RunPromptRequest`, `SessionInputRequest`, `ResizeRequest`, `CreateCaseRequest`, `QuickStartRequest`, `CreateScheduledRunRequest`, `QuickRunRequest`, `HookEventRequest`
|
||||
- Other dead types: `TaskAssignment`, `MemoryMetrics`, `RalphStateRecord`
|
||||
- Dead function: `createSuccessResponse` (exported, never imported)
|
||||
- **Estimated savings:** ~150 lines
|
||||
|
||||
### 5. 9 unused constants in map-limits.ts
|
||||
- `MAX_PENDING_HOOKS`, `MAX_SESSION_HISTORY`, `MAX_SSE_CLIENTS_PER_SESSION`, `MAX_TOTAL_SSE_CLIENTS`, `FILE_WATCHER_WARNING_THRESHOLD`, `MAX_QUEUED_TASKS`, `MAX_COMPLETED_TASKS_HISTORY`, `COMPLETED_TODO_TTL_MS`, `MAX_CONCURRENT_SESSIONS`
|
||||
- 9 of 14 exports are dead — only 5 are actually imported
|
||||
|
||||
### 6. Dead `SessionInputSchema` in schemas.ts
|
||||
- `SessionInputSchema` (line 87) is defined/exported but never imported
|
||||
- `SessionInputWithLimitSchema` is the one actually used
|
||||
|
||||
### 7. Dead `code-reviewer.ts` prompt file
|
||||
- `src/prompts/code-reviewer.ts` — entire file is dead, `CODE_REVIEWER_PROMPT` never imported
|
||||
- Re-exported in `src/prompts/index.ts` but no consumer
|
||||
|
||||
### 8. Dead utility exports
|
||||
- **Default exports** (4 files): `lru-map.ts`, `cleanup-manager.ts`, `stale-expiration-map.ts`, `buffer-accumulator.ts` — all have `export default` that's never used
|
||||
- **`stripAnsiSimple`** in `regex-patterns.ts` — exported, never imported (only `stripAnsi` used)
|
||||
- **String similarity**: `isSimilar`, `isSimilarByDistance`, `stringSimilarity`, `levenshteinDistance` — none imported externally
|
||||
- **LRUMap methods**: `oldest()`, `newest()`, `peek()`, `expireOlderThan()`, `valuesInOrder()`, `maxEntries`, `freeSlots` — never called
|
||||
- **StaleExpirationMap methods**: `touch()`, `getAge()`, `getRemainingTtl()`, `peek()` — never called
|
||||
- **CleanupManager methods**: `registerWatcher()`, `registerListener()`, `registerStream()`, `getRegistrations()`, `resourceCounts` — never called
|
||||
|
||||
### 9. Dead backend functions
|
||||
- `resetSessionManager()` in session-manager.ts:300 — never imported
|
||||
- `getStoredTasks()` in task-queue.ts:264 — never called
|
||||
- `start()` in session.ts:1918 — no-op legacy method
|
||||
- Empty `updateStatsFromEvent()` in run-summary.ts:397 — called every event, does nothing
|
||||
|
||||
### 10. Dead TS type exports
|
||||
- `AiCheckerEvents<R>`, `AiIdleCheckerEvents`, `AiPlanCheckerEvents` — never imported
|
||||
- `AiCheckStatus`, `AiPlanCheckStatus` — backwards compat aliases, never imported
|
||||
|
||||
## P2 — Performance & Efficiency
|
||||
|
||||
### 11. task-queue.ts `getCount()` iterates all tasks 5 times
|
||||
- Called every Ralph Loop tick — creates array from Map, then filters 4 times
|
||||
- **Fix:** Single-pass counting like `TaskTracker.getStats()` does
|
||||
|
||||
### 12. transcript-watcher.ts double file read
|
||||
- `readNewEntries()` reads the file twice: once for CRLF detection, once for parsing
|
||||
- `crlfDelay: Infinity` already handles both line endings
|
||||
- **Fix:** Remove the raw buffer CRLF check, read once
|
||||
|
||||
### 13. tmux-manager.ts `saveSessions()` no debounce
|
||||
- Rapid calls can overlap; no in-flight guard unlike `StateStore`
|
||||
- **Fix:** Add debouncing or in-flight tracking
|
||||
|
||||
## P3 — Consolidation & Consistency
|
||||
|
||||
### 14. Duplicate `SAFE_PATH_PATTERN` regex
|
||||
- `schemas.ts:15` and `tmux-manager.ts:81` — identical regex
|
||||
- **Fix:** Share from one location
|
||||
|
||||
### 15. Duplicate `MAX_CONCURRENT_SESSIONS`
|
||||
- `map-limits.ts:57` (dead) vs `server.ts:131` (used, hardcoded)
|
||||
- **Fix:** server.ts should import from map-limits
|
||||
|
||||
### 16. Duplicate cache TTLs in server.ts
|
||||
- `SESSIONS_LIST_CACHE_TTL` and `LIGHT_STATE_CACHE_TTL_MS` — both 1000ms
|
||||
- **Fix:** Consolidate into one constant
|
||||
|
||||
### 17. Inconsistent path import in server.ts
|
||||
- Imports both `path` default and destructured `{ join, dirname, resolve, relative, isAbsolute }`
|
||||
- 3 lines use `path.join()` while everywhere else uses `join()`
|
||||
- **Fix:** Remove default import, use `join()` consistently
|
||||
|
||||
### 18. Re-export indirection for `getAugmentedPath`
|
||||
- `session.ts:89` re-exports from `claude-cli-resolver.ts` for backwards compat
|
||||
- `ai-checker-base.ts` should import directly from source
|
||||
|
||||
### 19. `cliInfoUpdated` event missing from SessionEvents interface
|
||||
- Emitted in `session.ts:1742`, handled in `server.ts:4214`, but not in the interface
|
||||
- Type safety gap — handlers aren't type-checked
|
||||
|
||||
### 20. Array instead of Set for `_childAgentIds` in session.ts
|
||||
- Uses `includes()`/`indexOf()` for lookups (O(n))
|
||||
- Small lists in practice, but Set is more appropriate
|
||||
@@ -1,983 +0,0 @@
|
||||
> **⚠️ ARCHIVED 2026-05-21 — superseded, kept for history.**
|
||||
> The "Critical" structural items here are done: `server.ts` 6,736→2,065 LOC,
|
||||
> `app.js` 15,196→3,083 LOC, `types.ts` 1,443→12 LOC (now a barrel → `src/types/`).
|
||||
> The phase plans that executed this work are in `docs/archive/phase*-plan.md`.
|
||||
> Do not treat this as a live TODO; see CLAUDE.md for current architecture.
|
||||
|
||||
# Code Structure & Quality Findings
|
||||
|
||||
**Date**: 2026-02-28
|
||||
**Scope**: Full codebase analysis across 5 dimensions: frontend, backend, TypeScript, testing, and utilities/config.
|
||||
|
||||
This document contains detailed findings for agent teams to write implementation plans and execute improvements. Each section includes severity, specific locations, and recommended fixes.
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Critical: server.ts God Object (6,736 LOC)](#1-critical-serverts-god-object)
|
||||
2. [Critical: app.js Monolith (15,196 LOC)](#2-critical-appjs-monolith)
|
||||
3. [Critical: CleanupManager Unused Despite Existing](#3-critical-cleanupmanager-unused)
|
||||
4. [High: Duplicated Debounce/Timer Patterns](#4-high-duplicated-debouncetimer-patterns)
|
||||
5. [High: Large Domain Files Need Splitting](#5-high-large-domain-files-need-splitting)
|
||||
6. [High: types.ts God File (1,443 LOC)](#6-high-typests-god-file)
|
||||
7. [High: Zod Schemas Duplicate TypeScript Types](#7-high-zod-schemas-duplicate-typescript-types)
|
||||
8. [High: Test Coverage Gaps](#8-high-test-coverage-gaps)
|
||||
9. [High: Duplicated Test Mocks](#9-high-duplicated-test-mocks)
|
||||
10. [Medium: Hardcoded Magic Values](#10-medium-hardcoded-magic-values)
|
||||
11. [Medium: Frontend Global State Monolith](#11-medium-frontend-global-state-monolith)
|
||||
12. [Medium: Frontend Code Duplication](#12-medium-frontend-code-duplication)
|
||||
13. [Medium: Inconsistent Logging](#13-medium-inconsistent-logging)
|
||||
14. [Medium: Utils Barrel Export Gaps](#14-medium-utils-barrel-export-gaps)
|
||||
15. [Medium: Non-Null Assertion Risks](#15-medium-non-null-assertion-risks)
|
||||
16. [Low: Dead Utility Functions](#16-low-dead-utility-functions)
|
||||
17. [Low: No Dependency Injection for File I/O](#17-low-no-dependency-injection-for-file-io)
|
||||
18. [Scorecard & Prioritized Roadmap](#18-scorecard--prioritized-roadmap)
|
||||
|
||||
---
|
||||
|
||||
## 1. Critical: server.ts God Object
|
||||
|
||||
**File**: `src/web/server.ts` (6,736 lines)
|
||||
**Severity**: CRITICAL
|
||||
**Impact**: Hardest file to maintain, test, and extend. Imports 38 modules.
|
||||
|
||||
### Problem
|
||||
|
||||
The `WebServer` class handles everything: HTTP routing (~110 routes), authentication, SSE broadcasting, terminal data batching, state persistence, session lifecycle, respawn orchestration, file serving, tunnel management, plan orchestration, and subagent coordination.
|
||||
|
||||
**Key metrics**:
|
||||
- 40+ private properties (Maps, timers, caches)
|
||||
- 70+ methods
|
||||
- `setupRoutes()` is 2,000+ LOC of inline route handlers
|
||||
- Zero test coverage
|
||||
|
||||
### Current Structure (Bad)
|
||||
|
||||
```
|
||||
WebServer class (6,736 LOC)
|
||||
├── Auth session management (lines 469, 668-698)
|
||||
├── SSE client management (lines 407-408, 5843-5880)
|
||||
├── Terminal data batching (lines 414-416, 5909-5966)
|
||||
├── Task update batching (line 426, 5995-6028)
|
||||
├── State persistence batching (lines 429-430, 6028-6061)
|
||||
├── Respawn lifecycle (lines 445-451, 5425-5534)
|
||||
├── Session cleanup (lines 4769-4961)
|
||||
├── Listener setup (lines 544-643)
|
||||
└── setupRoutes() (lines 645+, 2000+ LOC)
|
||||
├── /api/sessions/* (30+ routes inline)
|
||||
├── /api/respawn/* (7 routes inline)
|
||||
├── /api/subagents/* (7 routes inline)
|
||||
├── /api/plan/* (5 routes inline)
|
||||
├── /api/push/* (4 routes inline)
|
||||
└── ... 60+ more inline
|
||||
```
|
||||
|
||||
### Recommended Structure
|
||||
|
||||
```
|
||||
src/web/
|
||||
├── server.ts (~500 LOC - HTTP setup, route registration only)
|
||||
├── routes/
|
||||
│ ├── session-routes.ts (session CRUD, input, resize)
|
||||
│ ├── respawn-routes.ts (respawn control endpoints)
|
||||
│ ├── subagent-routes.ts (background agent tracking)
|
||||
│ ├── plan-routes.ts (plan generation & management)
|
||||
│ ├── push-routes.ts (web push subscriptions)
|
||||
│ ├── mux-routes.ts (tmux management)
|
||||
│ ├── case-routes.ts (case management)
|
||||
│ ├── file-routes.ts (file browsing/serving)
|
||||
│ └── system-routes.ts (status, stats, config, settings)
|
||||
├── middleware/
|
||||
│ ├── auth.ts (Basic Auth + session cookies)
|
||||
│ └── error-handler.ts (centralized error responses)
|
||||
└── services/
|
||||
├── sse-manager.ts (SSE client + broadcast)
|
||||
├── terminal-batcher.ts (60fps terminal batching)
|
||||
└── session-lifecycle.ts (listener setup/teardown)
|
||||
```
|
||||
|
||||
### Duplication in server.ts
|
||||
|
||||
**Error response pattern** repeated 189 times:
|
||||
```typescript
|
||||
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Session not found');
|
||||
```
|
||||
|
||||
**Fix**: Extract `findSessionOrFail()` middleware:
|
||||
```typescript
|
||||
const findSessionOrFail = (sessionId: string) => {
|
||||
const session = this.sessions.get(sessionId);
|
||||
if (!session) throw new NotFoundError('Session not found');
|
||||
return session;
|
||||
};
|
||||
```
|
||||
|
||||
**Event listener setup** copy-pasted for subagent watcher, image watcher, and team watcher (lines 544-643). Same attach/detach pattern duplicated 3 times.
|
||||
|
||||
---
|
||||
|
||||
## 2. Critical: app.js Monolith
|
||||
|
||||
**File**: `src/web/public/app.js` (15,196 lines)
|
||||
**Severity**: CRITICAL
|
||||
**Impact**: Untestable, hard to navigate, tightly coupled systems.
|
||||
|
||||
### Extractable Modules (by priority)
|
||||
|
||||
| Module | Lines | Current Location | Impact |
|
||||
|--------|-------|------------------|--------|
|
||||
| Mobile handlers (MobileDetection, KeyboardHandler, SwipeHandler) | ~300 | lines 168-620 | High |
|
||||
| Voice input (DeepgramProvider, VoiceInput) | ~830 | lines 631-1471 | High |
|
||||
| NotificationManager | ~450 | lines 2218-2663 | High |
|
||||
| xterm-zerolag-input (inlined copy from packages/) | ~400 | lines 1756-2153 | High |
|
||||
| KeyboardAccessoryBar | ~195 | lines 1480-1680 | Medium |
|
||||
| FocusTrap | ~60 | lines 1690-1748 | Medium |
|
||||
|
||||
### CodemanApp Class (12,000+ LOC)
|
||||
|
||||
The main `CodemanApp` class starting at line 2665 has:
|
||||
- **60+ Maps/Sets** in the constructor (lines 2667-2805)
|
||||
- **18 Map instances** with complex cross-references (subagents, parents, teams, windows)
|
||||
- **10+ monolithic methods** exceeding 100 lines each
|
||||
|
||||
**Largest methods**:
|
||||
| Method | Lines | Size |
|
||||
|--------|-------|------|
|
||||
| `renderAppSettings()` | 14400-14700 | ~300 LOC |
|
||||
| `selectSession()` | 6028-6250 | ~220 LOC |
|
||||
| `batchTerminalWrite()` | 7482-7700 | ~200 LOC |
|
||||
| `renderSessionTabs()` | 5814-6000 | ~180 LOC |
|
||||
| `openSubagentWindow()` | 11927-12100 | ~170 LOC |
|
||||
| `handleInit()` | 5183-5350 | ~170 LOC |
|
||||
|
||||
### Recommended Split
|
||||
|
||||
```
|
||||
src/web/public/
|
||||
├── app.js (~4000 LOC - core app, session mgmt, SSE)
|
||||
├── mobile.js (~300 LOC - MobileDetection, KeyboardHandler, SwipeHandler)
|
||||
├── voice.js (~830 LOC - DeepgramProvider, VoiceInput)
|
||||
├── notifications.js (~450 LOC - NotificationManager)
|
||||
├── keyboard-accessory.js (~200 LOC - KeyboardAccessoryBar)
|
||||
├── api-client.js (~100 LOC - fetch wrapper with error handling)
|
||||
└── config.js (~50 LOC - magic numbers, z-index layers)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Critical: CleanupManager Unused
|
||||
|
||||
**File**: `src/utils/cleanup-manager.ts` (320 lines)
|
||||
**Severity**: CRITICAL
|
||||
**Impact**: Memory leak risk. Well-designed utility exists but is never used. Every file manages cleanup manually.
|
||||
|
||||
### Current State
|
||||
|
||||
`CleanupManager` is exported from the utils barrel but has **0 instantiations** in production code. Instead, every file implements manual cleanup:
|
||||
|
||||
**respawn-controller.ts** (worst offender):
|
||||
```typescript
|
||||
// 11 timer properties, manually cleared in stop()
|
||||
private stepTimer: NodeJS.Timeout | null = null;
|
||||
private completionConfirmTimer: NodeJS.Timeout | null = null;
|
||||
private noOutputTimer: NodeJS.Timeout | null = null;
|
||||
// ... 8 more
|
||||
|
||||
stop() {
|
||||
if (this.stepTimer) clearTimeout(this.stepTimer);
|
||||
if (this.completionConfirmTimer) clearTimeout(this.completionConfirmTimer);
|
||||
// ... 9 more clearTimeout/clearInterval calls
|
||||
}
|
||||
```
|
||||
|
||||
**Files that should use CleanupManager**:
|
||||
| File | Timer/Listener Count | Current Cleanup |
|
||||
|------|---------------------|-----------------|
|
||||
| `respawn-controller.ts` | 11 timers + intervals | 11 manual clearTimeout/clearInterval |
|
||||
| `web/server.ts` | 6+ timers, debounce map | Manual in stop(), some may leak |
|
||||
| `state-store.ts` | 2 debounce timers | Manual clearTimeout |
|
||||
| `push-store.ts` | 1 save timer | Manual clearTimeout |
|
||||
| `subagent-watcher.ts` | debounce map + watchers | Manual clear + close |
|
||||
| `ralph-tracker.ts` | 3 debounce timers | Manual clear |
|
||||
| `bash-tool-parser.ts` | 1 debounce timer | Manual clear |
|
||||
| `image-watcher.ts` | 1 debounce map | Manual clear |
|
||||
|
||||
### Fix
|
||||
|
||||
Migrate all timer management to use `CleanupManager`. Example for respawn-controller.ts:
|
||||
|
||||
```typescript
|
||||
// Before: 11 fields + 11 clearTimeout calls
|
||||
private stepTimer: NodeJS.Timeout | null = null;
|
||||
// ...
|
||||
|
||||
// After: 1 field, auto-cleanup
|
||||
private cleanup = new CleanupManager();
|
||||
|
||||
startStep() {
|
||||
this.cleanup.setTimeout(() => { ... }, 5000, 'step');
|
||||
}
|
||||
|
||||
stop() {
|
||||
this.cleanup.dispose(); // Clears everything
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. High: Duplicated Debounce/Timer Patterns
|
||||
|
||||
**Severity**: HIGH
|
||||
**Impact**: 8+ files implement debounce independently. Bug fixes need to be applied everywhere.
|
||||
|
||||
### Pattern Inventory
|
||||
|
||||
```typescript
|
||||
// Pattern 1: Manual timer ref (used in 6 files)
|
||||
private saveTimer: NodeJS.Timeout | null = null;
|
||||
debouncedSave() {
|
||||
if (this.saveTimer) clearTimeout(this.saveTimer);
|
||||
this.saveTimer = setTimeout(() => this.save(), 500);
|
||||
}
|
||||
|
||||
// Pattern 2: Timer Map (used in 3 files)
|
||||
private fileDebouncers = new Map<string, NodeJS.Timeout>();
|
||||
debounce(key: string) {
|
||||
const existing = this.fileDebouncers.get(key);
|
||||
if (existing) clearTimeout(existing);
|
||||
this.fileDebouncers.set(key, setTimeout(() => { ... }, 100));
|
||||
}
|
||||
|
||||
// Pattern 3: State flag (used in 2 files)
|
||||
private isSaving = false;
|
||||
```
|
||||
|
||||
### Locations
|
||||
|
||||
| File | Debounce Vars | Delay (ms) |
|
||||
|------|---------------|------------|
|
||||
| `state-store.ts` | `saveTimeout`, `ralphStateSaveTimeout` | 500 |
|
||||
| `push-store.ts` | `saveTimer` | 500 |
|
||||
| `web/server.ts` | `persistDebounceTimers` (Map) | 500 |
|
||||
| `subagent-watcher.ts` | `fileDebouncers` (Map) | 100 |
|
||||
| `ralph-tracker.ts` | 3 debounce timers | 50, 30000 |
|
||||
| `bash-tool-parser.ts` | `EVENT_DEBOUNCE_MS` | 50 |
|
||||
| `image-watcher.ts` | debounce map | 200 |
|
||||
| `respawn-controller.ts` | 11 timer fields | various |
|
||||
|
||||
### Fix
|
||||
|
||||
Create a `Debouncer` utility:
|
||||
|
||||
```typescript
|
||||
// src/utils/debouncer.ts
|
||||
export class Debouncer {
|
||||
private timer: NodeJS.Timeout | null = null;
|
||||
|
||||
constructor(private readonly delayMs: number) {}
|
||||
|
||||
run(fn: () => void): void {
|
||||
if (this.timer) clearTimeout(this.timer);
|
||||
this.timer = setTimeout(fn, this.delayMs);
|
||||
}
|
||||
|
||||
cancel(): void {
|
||||
if (this.timer) clearTimeout(this.timer);
|
||||
this.timer = null;
|
||||
}
|
||||
}
|
||||
|
||||
// Usage:
|
||||
private saveDeb = new Debouncer(500);
|
||||
this.saveDeb.run(() => this.save());
|
||||
// cleanup: this.saveDeb.cancel();
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. High: Large Domain Files Need Splitting
|
||||
|
||||
**Severity**: HIGH
|
||||
**Impact**: Complex state machines spanning 3,000+ lines are hard to understand and test.
|
||||
|
||||
### ralph-tracker.ts (3,905 LOC)
|
||||
|
||||
**5 responsibilities mixed**:
|
||||
1. Output Parsing (~900 LOC) - Line-by-line parsing, state extraction
|
||||
2. Todo Management (~700 LOC) - Parsing, dedup, expiry
|
||||
3. Plan Tracking (~800 LOC) - Enhanced plan tasks, checkpoints
|
||||
4. Circuit Breaker (~400 LOC) - State machine for stuck detection
|
||||
5. File Watching (~300 LOC) - Monitor external state files
|
||||
|
||||
**Recommended split**:
|
||||
```
|
||||
ralph-tracker.ts (core output parsing, ~1200 LOC)
|
||||
ralph-todo-manager.ts (todo parsing + management, ~700 LOC)
|
||||
ralph-plan-tracker.ts (plan tasks + checkpoints, ~800 LOC)
|
||||
ralph-circuit-breaker.ts (circuit breaker logic, ~400 LOC)
|
||||
```
|
||||
|
||||
### respawn-controller.ts (3,611 LOC)
|
||||
|
||||
**6 responsibilities mixed**:
|
||||
1. State Machine (~1,000 LOC) - 6+ states, transitions
|
||||
2. Idle Detection (~800 LOC) - 5 layers + multi-signal combining
|
||||
3. AI Checkers (~600 LOC) - Idle + plan checkers integration
|
||||
4. Health Scoring (~500 LOC) - Metrics, circuit breaker, scoring
|
||||
5. Action Logging (~300 LOC) - Timeline, detection status
|
||||
6. Stuck-State Detection (~250 LOC) - Timeout tracking
|
||||
|
||||
**Recommended split**:
|
||||
```
|
||||
respawn-controller.ts (state machine core, ~1000 LOC)
|
||||
respawn-idle-detection.ts (all 5 idle detection layers, ~800 LOC)
|
||||
respawn-health-scorer.ts (metrics & health scoring, ~500 LOC)
|
||||
```
|
||||
|
||||
### session.ts (2,418 LOC)
|
||||
|
||||
**8 responsibilities mixed**:
|
||||
1. PTY Management (~600 LOC)
|
||||
2. Terminal I/O (~400 LOC)
|
||||
3. Token Tracking (~200 LOC)
|
||||
4. Task Tracking (~250 LOC)
|
||||
5. Ralph Integration (~200 LOC)
|
||||
6. Auto-Clear/Compact (~300 LOC)
|
||||
7. Image Watching (~100 LOC)
|
||||
8. CLI Detection (~150 LOC)
|
||||
|
||||
**Recommended split**:
|
||||
```
|
||||
session.ts (PTY + terminal I/O core, ~1000 LOC)
|
||||
session-tracking.ts (token + task + Ralph, ~500 LOC)
|
||||
session-auto-ops.ts (auto-clear/compact + image, ~300 LOC)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. High: types.ts God File
|
||||
|
||||
**File**: `src/types.ts` (1,443 lines, 72 exported definitions)
|
||||
**Severity**: HIGH
|
||||
**Impact**: Every file imports from types.ts. Hard to find relevant types.
|
||||
|
||||
### Current Contents
|
||||
|
||||
- 46 interfaces
|
||||
- 25 types
|
||||
- 1 enum (ApiErrorCode)
|
||||
- 9 factory functions (createInitialState, etc.)
|
||||
|
||||
### Recommended Split
|
||||
|
||||
```
|
||||
src/types/
|
||||
├── index.ts (barrel export - transparent migration)
|
||||
├── session.ts (SessionState, SessionConfig, SessionMode, SessionColor)
|
||||
├── task.ts (TaskState, TaskDefinition, TaskStatus)
|
||||
├── respawn.ts (RespawnConfig, RespawnState, CircuitBreakerStatus)
|
||||
├── ralph.ts (RalphLoopState, RalphTrackerState, RalphTodoItem)
|
||||
├── api.ts (ApiResponse, ApiErrorCode, HookEventType, all route types)
|
||||
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
|
||||
└── common.ts (Disposable, BufferConfig, CleanupResourceType)
|
||||
```
|
||||
|
||||
The barrel export makes this a transparent refactor - existing `import from './types'` continues to work.
|
||||
|
||||
---
|
||||
|
||||
## 7. High: Zod Schemas Duplicate TypeScript Types
|
||||
|
||||
**File**: `src/web/schemas.ts` (508 lines)
|
||||
**Severity**: HIGH
|
||||
**Impact**: When a type changes, the Zod schema must be manually updated too. Source of bugs.
|
||||
|
||||
### Problem
|
||||
|
||||
Zod schemas manually duplicate TypeScript interfaces. **Zero `z.infer` usage found.**
|
||||
|
||||
```typescript
|
||||
// types.ts (manual interface)
|
||||
export interface CreateSessionRequest {
|
||||
workingDir?: string;
|
||||
mode?: SessionMode;
|
||||
name?: string;
|
||||
}
|
||||
|
||||
// schemas.ts (manual Zod schema - duplicated!)
|
||||
export const CreateSessionSchema = z.object({
|
||||
workingDir: safePathSchema.optional(),
|
||||
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
|
||||
name: z.string().max(100).optional(),
|
||||
});
|
||||
```
|
||||
|
||||
### Fix
|
||||
|
||||
Use `z.infer` to derive TypeScript types from Zod schemas (single source of truth):
|
||||
|
||||
```typescript
|
||||
// schemas.ts
|
||||
export const CreateSessionSchema = z.object({
|
||||
workingDir: safePathSchema.optional(),
|
||||
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
|
||||
name: z.string().max(100).optional(),
|
||||
});
|
||||
|
||||
// types.ts (auto-derived)
|
||||
export type CreateSessionRequest = z.infer<typeof CreateSessionSchema>;
|
||||
```
|
||||
|
||||
**Affected schemas** (~10):
|
||||
- CreateSessionSchema
|
||||
- RunPromptSchema
|
||||
- ResizeSchema
|
||||
- CreateCaseSchema
|
||||
- QuickStartSchema
|
||||
- HookEventSchema
|
||||
- RespawnConfigSchema
|
||||
- ConfigUpdateSchema
|
||||
- SettingsUpdateSchema
|
||||
|
||||
---
|
||||
|
||||
## 8. High: Test Coverage Gaps
|
||||
|
||||
**Severity**: HIGH
|
||||
**Impact**: Critical code paths untested. Regressions go unnoticed.
|
||||
|
||||
### Untested Source Files
|
||||
|
||||
| File | Lines | Risk |
|
||||
|------|-------|------|
|
||||
| `src/web/server.ts` | 6,736 | CRITICAL - Core REST API, 280+ routes |
|
||||
| `src/plan-orchestrator.ts` | ~500 | HIGH - Multi-agent plan generation |
|
||||
| `src/tunnel-manager.ts` | ~200 | MEDIUM - Cloudflare tunnel |
|
||||
| `src/session-lifecycle-log.ts` | ~150 | MEDIUM - JSONL audit log |
|
||||
| `src/ai-plan-checker.ts` | ~300 | MEDIUM - Plan completion detection |
|
||||
| `src/templates/claude-md.ts` | ~200 | LOW - CLAUDE.md generation |
|
||||
| `src/utils/claude-cli-resolver.ts` | ~100 | LOW - CLI path resolution |
|
||||
| `src/utils/opencode-cli-resolver.ts` | ~100 | LOW - OpenCode CLI support |
|
||||
| `src/utils/regex-patterns.ts` | ~100 | LOW - Used everywhere! |
|
||||
| `src/utils/token-validation.ts` | ~50 | LOW - Token counting |
|
||||
|
||||
### Test Quality Issues
|
||||
|
||||
**10 "not.toThrow()" tests without behavior verification**:
|
||||
```typescript
|
||||
// BAD: Only checks it doesn't crash
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
|
||||
// GOOD: Also verify defensive behavior
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
```
|
||||
|
||||
Locations:
|
||||
- `task-tracker.test.ts` - 5 instances
|
||||
- `image-watcher.test.ts` - 1 instance
|
||||
- `task-queue.test.ts` - 1 instance
|
||||
- Others scattered
|
||||
|
||||
---
|
||||
|
||||
## 9. High: Duplicated Test Mocks
|
||||
|
||||
**Severity**: HIGH
|
||||
**Impact**: Mock changes need updating in 4 places. Inconsistent mock behavior.
|
||||
|
||||
### MockSession Defined 4 Times
|
||||
|
||||
| File | Usage |
|
||||
|------|-------|
|
||||
| `test/respawn-controller.test.ts` | Full mock with event emitter |
|
||||
| `test/session-manager.test.ts` | Simpler mock |
|
||||
| `test/respawn-team-awareness.test.ts` | Copy of respawn-controller mock |
|
||||
| `test/respawn-test-utils.ts` | **Comprehensive mock - UNUSED!** |
|
||||
|
||||
### MockStateStore Defined 2 Times
|
||||
|
||||
| File | Usage |
|
||||
|------|-------|
|
||||
| `test/session-manager.test.ts` | Basic mock |
|
||||
| `test/ralph-loop.test.ts` | Separate implementation |
|
||||
|
||||
### Unused Test Utilities
|
||||
|
||||
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
|
||||
- `createTimeController()` - Abstraction over vitest fake timers
|
||||
- `MockAiIdleChecker` - Fully mocked AI idle checker
|
||||
- `MockAiPlanChecker` - Fully mocked plan checker
|
||||
- Factory functions for pre-configured controllers
|
||||
|
||||
### Fix
|
||||
|
||||
Create `test/mocks/` directory:
|
||||
```
|
||||
test/
|
||||
├── mocks/
|
||||
│ ├── mock-session.ts (single MockSession, used everywhere)
|
||||
│ ├── mock-state-store.ts (single MockStateStore)
|
||||
│ └── index.ts (barrel export)
|
||||
├── utils/
|
||||
│ └── time-controller.ts (from respawn-test-utils.ts)
|
||||
└── ... test files
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 10. Medium: Hardcoded Magic Values
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Hard to tune, inconsistent when same value appears in multiple places.
|
||||
|
||||
### Already Centralized (Good)
|
||||
|
||||
- `src/config/buffer-limits.ts` - All buffer sizes
|
||||
- `src/config/map-limits.ts` - All collection limits
|
||||
|
||||
### NOT Centralized (40+ values scattered)
|
||||
|
||||
**In server.ts** (lines 145-194):
|
||||
```typescript
|
||||
const TASK_UPDATE_BATCH_INTERVAL = 100;
|
||||
const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
|
||||
const SESSIONS_LIST_CACHE_TTL = 1000;
|
||||
const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
|
||||
const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
|
||||
const MAX_TERMINAL_COLS = 500;
|
||||
const MAX_TERMINAL_ROWS = 200;
|
||||
const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
|
||||
const MAX_AUTH_SESSIONS = 100;
|
||||
const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
|
||||
const STATS_COLLECTION_INTERVAL_MS = 2000;
|
||||
const MAX_INPUT_LENGTH = 64 * 1024;
|
||||
```
|
||||
|
||||
**In hooks-config.ts**: `timeout: 10000` hardcoded 6 times.
|
||||
|
||||
**In respawn-controller.ts** (lines 538-565): 10 timing constants.
|
||||
|
||||
**In utils**: `EXEC_TIMEOUT_MS = 5000` duplicated in both `claude-cli-resolver.ts` and `opencode-cli-resolver.ts`.
|
||||
|
||||
**In app.js**:
|
||||
```javascript
|
||||
// line 27: 600000 - stuck detection threshold
|
||||
// line 24: 5000 - default scrollback
|
||||
// lines 34-35: 128*1024, 256*1024 - chunk sizes
|
||||
// lines 152-155: 150, 100 - keyboard detection thresholds
|
||||
// lines 573-575: 80, 300, 100 - swipe detection params
|
||||
```
|
||||
|
||||
### Fix
|
||||
|
||||
Create additional config files:
|
||||
```
|
||||
src/config/
|
||||
├── buffer-limits.ts (existing)
|
||||
├── map-limits.ts (existing)
|
||||
├── server-config.ts (NEW - web server intervals, auth, caching)
|
||||
├── timing-config.ts (NEW - debounce delays, check intervals)
|
||||
└── terminal-config.ts (NEW - max cols/rows, batch intervals)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 11. Medium: Frontend Global State Monolith
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: All state in single CodemanApp class. Tight coupling between unrelated systems.
|
||||
|
||||
### 60+ State Variables in CodemanApp Constructor (lines 2667-2805)
|
||||
|
||||
```javascript
|
||||
this.sessions = new Map(); // Session data
|
||||
this.subagents = new Map(); // Agent tracking
|
||||
this.subagentActivity = new Map(); // Tool call tracking
|
||||
this.subagentToolResults = new Map(); // Result caching
|
||||
this.subagentParentMap = new Map(); // Agent-to-session mapping
|
||||
this.teams = new Map(); // Team tracking
|
||||
this.teamTasks = new Map(); // Team task state
|
||||
this.planSubagents = new Map(); // Plan agent tracking
|
||||
this.pendingWrites = []; // Terminal write queue
|
||||
this.terminalBufferCache = new Map(); // Buffer caching (unbounded!)
|
||||
this.projectInsights = new Map(); // Bash tool insights
|
||||
// ... 40+ more
|
||||
```
|
||||
|
||||
### Problems
|
||||
|
||||
1. **18 Map instances** with complex cross-references (no garbage collection strategy)
|
||||
2. **No domain separation**: Session, subagent, notification, UI, and network state mixed
|
||||
3. **Implicit dependencies**: `selectSession()` requires 5+ Maps to be in consistent state
|
||||
4. **`terminalBufferCache`** has no max size - can grow unbounded with many sessions
|
||||
|
||||
### Recommended Domain Split
|
||||
|
||||
```javascript
|
||||
// Instead of 60+ flat properties:
|
||||
class SessionState {
|
||||
sessions = new Map();
|
||||
sessionOrder = [];
|
||||
terminalBuffers = new Map();
|
||||
tabAlerts = new Map();
|
||||
}
|
||||
|
||||
class SubagentState {
|
||||
subagents = new Map();
|
||||
activity = new Map();
|
||||
parentMap = new Map();
|
||||
windows = new Map();
|
||||
minimized = new Map();
|
||||
}
|
||||
|
||||
class TeamState {
|
||||
teams = new Map();
|
||||
tasks = new Map();
|
||||
teammates = new Map();
|
||||
}
|
||||
|
||||
class UIState {
|
||||
activeSessionId = null;
|
||||
draggedTabId = null;
|
||||
isLoadingBuffer = false;
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 12. Medium: Frontend Code Duplication
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Repeated patterns increase maintenance burden and inconsistency risk.
|
||||
|
||||
### Duplicated Patterns
|
||||
|
||||
**API fetch calls** (~50 instances):
|
||||
```javascript
|
||||
// Repeated everywhere:
|
||||
fetch(`/api/sessions/${sessionId}/...`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({...})
|
||||
}).catch(() => {})
|
||||
```
|
||||
**Fix**: Extract `ApiClient` class.
|
||||
|
||||
**`innerHTML` usage** (104 instances):
|
||||
- Mix of template strings, createElement chains, and direct innerHTML
|
||||
- Some with manual XSS escaping (`text.replace(/</g, '<')`), some without
|
||||
- No consistent DOM creation pattern
|
||||
|
||||
**`typeof app !== 'undefined'` checks** (20+ instances):
|
||||
- Lines 458, 467, 481, 614, 617, 1549, etc.
|
||||
- **Fix**: Ensure `app` is always defined as global singleton.
|
||||
|
||||
**Element visibility toggling** (212+ occurrences):
|
||||
```javascript
|
||||
element.classList.add('active')
|
||||
element.classList.remove('active')
|
||||
```
|
||||
**Fix**: Create `toggleClass(el, className, condition)` utility.
|
||||
|
||||
### Event Listener Issues
|
||||
|
||||
- **152 `addEventListener` calls** with fragile cleanup
|
||||
- **Mix of inline (`onclick="app.method()"`) and addEventListener** - hard to track
|
||||
- **Element cache (`_elemCache`) never invalidated** if DOM elements are recreated (line 2808)
|
||||
- **Tab drag-and-drop listeners** may not clean up if user switches tabs mid-drag
|
||||
|
||||
---
|
||||
|
||||
## 13. Medium: Inconsistent Logging
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Hard to debug in production. Can't filter by severity or component.
|
||||
|
||||
### Current State
|
||||
|
||||
- **345 console calls** across source files
|
||||
- **No structured logging** - all `console.log/error` directly
|
||||
- **No log levels** (DEBUG, INFO, WARN, ERROR)
|
||||
|
||||
### Inconsistent Prefixes
|
||||
|
||||
```typescript
|
||||
// Some files use brackets:
|
||||
console.log('[Session] Starting interactive...');
|
||||
console.log('[RalphLoop] Task assigned...');
|
||||
console.log('[TunnelManager] Tunnel started');
|
||||
|
||||
// Others use no prefix:
|
||||
console.error('Failed to spawn PTY:', err);
|
||||
console.log('Server listening on port', port);
|
||||
```
|
||||
|
||||
### Positive: CleanupManager Has Debug Mode
|
||||
|
||||
`src/utils/cleanup-manager.ts` has a `debugMode` flag for conditional debug logging - good pattern not replicated elsewhere.
|
||||
|
||||
### Fix
|
||||
|
||||
Either:
|
||||
1. Enforce consistent `[ComponentName]` prefixes via lint rule
|
||||
2. Create lightweight logger abstraction (not a heavy framework)
|
||||
|
||||
---
|
||||
|
||||
## 14. Medium: Utils Barrel Export Gaps
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Forces deep imports, unclear public API.
|
||||
|
||||
### Missing Exports
|
||||
|
||||
These functions are defined but NOT exported from the barrel:
|
||||
- `createAnsiPatternFull()` and `createAnsiPatternSimple()` (factory functions from `regex-patterns.ts`)
|
||||
- `SAFE_PATH_PATTERN` (from `regex-patterns.ts`)
|
||||
- `validateTokenCounts()` and `validateTokensAndCost()` (from `token-validation.ts`)
|
||||
- `isSimilar()`, `isSimilarByDistance()`, `levenshteinDistance()`, `normalizePhrase()` (from `string-similarity.ts` - though some are dead code, see finding #16)
|
||||
|
||||
### Deep Import Anti-Pattern (16 instances)
|
||||
|
||||
Some files bypass the barrel unnecessarily:
|
||||
```typescript
|
||||
// Could use barrel:
|
||||
import { BufferAccumulator } from './utils/buffer-accumulator.js';
|
||||
import { LRUMap } from './utils/lru-map.js';
|
||||
|
||||
// Must deep import (not in barrel):
|
||||
import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';
|
||||
```
|
||||
|
||||
### Fix
|
||||
|
||||
Add missing exports to `src/utils/index.ts` and update import sites.
|
||||
|
||||
---
|
||||
|
||||
## 15. Medium: Non-Null Assertion Risks
|
||||
|
||||
**Severity**: MEDIUM
|
||||
**Impact**: Runtime crashes if assumptions violated. 37 instances found.
|
||||
|
||||
### Distribution
|
||||
|
||||
| File | Count | Risk Level |
|
||||
|------|-------|------------|
|
||||
| `src/web/server.ts` | 10 | Low (auth flow verified) |
|
||||
| `src/session.ts` | 6 | **High** (mux/terminal refs) |
|
||||
| `src/respawn-controller.ts` | 4 | Low (config validated) |
|
||||
| `src/lru-map.ts` | 3 | Low (checked lookups) |
|
||||
| `src/subagent-watcher.ts` | 2 | Low (pending tool calls) |
|
||||
| Others | 12 | Low |
|
||||
|
||||
### High-Risk Examples (session.ts)
|
||||
|
||||
```typescript
|
||||
// Line 915 - _mux could be null if startInteractive called during cleanup
|
||||
`[Session] Starting interactive (with ${this._mux!.backend})`
|
||||
|
||||
// Line 954 - _muxSession could be null in race condition
|
||||
this._muxSession!.muxName
|
||||
```
|
||||
|
||||
### Fix
|
||||
|
||||
Add null guards before assertions, or document invariants:
|
||||
```typescript
|
||||
// Before:
|
||||
this._mux!.backend
|
||||
|
||||
// After:
|
||||
if (!this._mux) throw new Error('Invariant: _mux must be initialized before startInteractive');
|
||||
this._mux.backend
|
||||
```
|
||||
|
||||
### Positive Notes
|
||||
|
||||
- **0 instances of `as any`**
|
||||
- **0 instances of `@ts-ignore` or `@ts-expect-error`**
|
||||
- TypeScript overall score: 8.5/10
|
||||
|
||||
---
|
||||
|
||||
## 16. Low: Dead Utility Functions
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
**Severity**: LOW
|
||||
**Impact**: Code clutter, confusion about what's actually used.
|
||||
|
||||
### Unused Functions
|
||||
|
||||
These are defined and exported but **never imported anywhere**:
|
||||
- `isSimilar(a, b, threshold)` - similarity check with threshold
|
||||
- `isSimilarByDistance(a, b, maxDistance)` - Levenshtein-based check
|
||||
- `levenshteinDistance(a, b)` - raw edit distance
|
||||
- `normalizePhrase(phrase)` - phrase normalization
|
||||
|
||||
### Actually Used
|
||||
|
||||
Only these are imported from the barrel:
|
||||
- `stringSimilarity()` - used in ralph-tracker.ts
|
||||
- `fuzzyPhraseMatch()` - used in ralph-tracker.ts
|
||||
- `todoContentHash()` - used in ralph-tracker.ts
|
||||
|
||||
### Fix
|
||||
|
||||
Delete unused functions or mark as `@internal` if kept for future use.
|
||||
|
||||
---
|
||||
|
||||
## 17. Low: No Dependency Injection for File I/O
|
||||
|
||||
**Severity**: LOW (practical impact limited at current scale)
|
||||
**Impact**: Can't mock filesystem for unit tests. 68+ hard-coded filesystem calls.
|
||||
|
||||
### Examples
|
||||
|
||||
```typescript
|
||||
// state-store.ts - directly imports and uses fs
|
||||
import { readFileSync, writeFileSync, existsSync, mkdirSync } from 'node:fs';
|
||||
|
||||
// push-store.ts - hard-coded paths
|
||||
const KEYS_FILE = join(DATA_DIR, 'push-keys.json');
|
||||
const SUBS_FILE = join(DATA_DIR, 'push-subscriptions.json');
|
||||
|
||||
// ai-checker-base.ts - direct execSync
|
||||
execSync(`tmux kill-session -t "${this.checkMuxName}"`, { timeout: 3000 });
|
||||
```
|
||||
|
||||
### Why This Is Lower Priority
|
||||
|
||||
- The codebase uses integration tests (spawning real processes/tmux sessions) rather than unit tests
|
||||
- Most filesystem operations are in infrastructure code, not business logic
|
||||
- Adding DI would be a large refactor with limited near-term benefit
|
||||
|
||||
---
|
||||
|
||||
## 18. Scorecard & Prioritized Roadmap
|
||||
|
||||
### Overall Scores (Post-Implementation)
|
||||
|
||||
| Category | Before | After | Notes |
|
||||
|----------|--------|-------|-------|
|
||||
| TypeScript Safety | 8.5/10 | 9/10 | 0 `any`, 0 `@ts-ignore`, Zod `z.infer` eliminates type drift |
|
||||
| Error Handling | 8/10 | 8/10 | Unchanged — already strong |
|
||||
| Async/Promise Safety | 9.5/10 | 9.5/10 | Unchanged — already strong |
|
||||
| Resource Cleanup | 7/10 | 8/10 | CleanupManager adopted in server.ts, subagent-watcher, bash-tool-parser; Debouncer in 6 files. **Gaps**: respawn-controller (10+ manual timers) and ralph-tracker (2 manual timers) not migrated |
|
||||
| Module Organization | 5/10 | 8/10 | Routes extracted (12 modules), types split (14 domain files), domain files split (ralph: 7, respawn: 5, session: 6) |
|
||||
| Test Coverage | 6/10 | 7.5/10 | Shared mock infrastructure, 12 route test files, MockSession/MockStateStore consolidated |
|
||||
| Config Centralization | 6/10 | 9/10 | 9 config files, ~65 constants centralized, 0 cross-file duplicates |
|
||||
| Frontend Architecture | 4/10 | 7/10 | 8 extracted modules (3,453 LOC), app.js reduced 24% (15.2K → 11.5K), xterm-zerolag-input vendor build |
|
||||
| Code Duplication | 5/10 | 8/10 | Debouncer utility, shared test mocks, barrel exports, config consolidation |
|
||||
|
||||
### Implementation Phases
|
||||
|
||||
**Phase 1 - Quick Wins (1-2 days)** ✅ COMPLETE
|
||||
1. ✅ Export missing functions from utils barrel (~30 min) — `createAnsiPatternFull`, `createAnsiPatternSimple`, `SAFE_PATH_PATTERN`, `validateTokenCounts`, `validateTokensAndCost` all now exported from `src/utils/index.ts`
|
||||
2. ✅ Delete dead utility functions (~15 min) — `isSimilar()` removed from `string-similarity.ts`; `levenshteinDistance()`, `isSimilarByDistance()`, `normalizePhrase()` made private (used internally by `fuzzyPhraseMatch`/`stringSimilarity`)
|
||||
3. ✅ Consolidate duplicated `EXEC_TIMEOUT_MS` constant (~15 min) — Created `src/config/exec-timeout.ts` as single source of truth; `claude-cli-resolver.ts`, `opencode-cli-resolver.ts`, and `tmux-manager.ts` all import from it
|
||||
4. ✅ Add `z.infer` to Zod schemas (~2 hours) — `src/web/schemas.ts` now has 36 `z.infer` type exports (lines 512-547) covering all schemas
|
||||
5. ✅ Fix 10 weak "not.toThrow()" tests (~1 hour) — All `not.toThrow()` calls now have behavior assertions: `task-tracker.test.ts` (6 instances all followed by state checks), `image-watcher.test.ts` (1 instance followed by length check), `session-manager.test.ts` (1 instance followed by count check)
|
||||
|
||||
**Phase 2 - CleanupManager & Debounce (2-3 days)** ✅ COMPLETE
|
||||
1. ✅ Create `Debouncer` utility class (~1 hour) — Created `src/utils/debouncer.ts` with `Debouncer` and `KeyedDebouncer` classes; exported from `src/utils/index.ts`
|
||||
2. ✅ Migrate all 8 files from manual debounce to Debouncer — `state-store.ts` (2 Debouncers), `push-store.ts` (1 Debouncer), `bash-tool-parser.ts` (1 Debouncer), `image-watcher.ts` (1 KeyedDebouncer), `subagent-watcher.ts` (2 KeyedDebouncers), `server.ts` (1 KeyedDebouncer for persist timers), `ralph-tracker.ts` (2 Debouncers replacing 4 manual fields: `_todoUpdateTimer`, `_loopUpdateTimer`, `_todoUpdatePending`, `_loopUpdatePending`)
|
||||
3. ✅ Migrate respawn-controller to CleanupManager — 10 manual timer fields replaced with single `CleanupManager` instance + `timerIds` Map. `startTrackedTimer()`/`cancelTrackedTimer()` preserved as wrappers for UI countdown display and timer events. `clearTimers()` uses dispose-and-recreate pattern for state transitions.
|
||||
4. ✅ Migrate server.ts timer cleanup to CleanupManager (~2 hours) — `private cleanup = new CleanupManager()` present; terminal batch timers and pending respawn starts left as manual Maps (complex lifecycle)
|
||||
5. ✅ Migrate remaining files — `bash-tool-parser.ts` (CleanupManager ✅), `subagent-watcher.ts` (CleanupManager ✅), `ralph-tracker.ts` (Debouncer ✅)
|
||||
|
||||
**Phase 3 - server.ts Route Extraction (3-4 days)** ✅ COMPLETE
|
||||
1. ✅ Created `src/web/routes/` with 12 domain route modules + index barrel (4,090 LOC total): session (909), system (768), ralph (533), plan (459), respawn (315), case, file, hook-event, mux, push, scheduled, team
|
||||
2. ✅ Created `src/web/middleware/auth.ts` (193 LOC) — Basic Auth, session cookies, rate limiting, security headers, CORS
|
||||
3. ✅ Created `src/web/ports/` with 7 typed port interfaces (142 LOC) — SessionPort, EventPort, RespawnPort, ConfigPort, InfraPort, AuthPort; routes declare dependencies via intersection types
|
||||
4. ✅ Created `src/web/route-helpers.ts` (154 LOC) — `findSessionOrFail()`, `formatUptime()`, `sanitizeHookData()`, `autoConfigureRalph()`
|
||||
5. ✅ Reduced `server.ts` from 6,736 → 2,697 LOC (60% reduction). Remaining LOC is justified infrastructure: session lifecycle, SSE broadcast engine, terminal batching, respawn integration, resource cleanup
|
||||
|
||||
**Phase 4 - Domain File Splitting (2-3 days)** ✅ COMPLETE
|
||||
1. ✅ Split `types.ts` into `src/types/` directory — 14 domain files (1,469 LOC total): common, session, task, app-state, respawn, ralph, api, lifecycle, run-summary, tools, teams, push, plan + index barrel. Original `types.ts` is now a 1-line re-export
|
||||
2. ✅ Split `ralph-tracker.ts` into 7 files (exceeded plan of 4) — ralph-tracker (2,391), ralph-plan-tracker (477), ralph-status-parser (552), ralph-fix-plan-watcher (366), ralph-stall-detector (166), ralph-config (153), ralph-loop (522)
|
||||
3. ✅ Split `respawn-controller.ts` into 5 files (exceeded plan of 3) — respawn-controller (3,228), respawn-health (229), respawn-metrics (229), respawn-patterns (131), respawn-adaptive-timing (134)
|
||||
4. ✅ Split `session.ts` into 6 files (exceeded plan of 3) — session (2,168), session-manager (298), session-auto-ops (284), session-cli-builder (132), session-task-cache (101), session-lifecycle-log (114)
|
||||
|
||||
**Phase 5 - Frontend Modularization (3-4 days)** ✅ COMPLETE
|
||||
1. ✅ Extracted `constants.js` (238 LOC) — shared constants, timing values, Z-index layers, `escapeHtml()`, `extractSyncSegments()`
|
||||
2. ✅ Extracted `mobile-handlers.js` (449 LOC) — `MobileDetection`, `KeyboardHandler`, `SwipeHandler`
|
||||
3. ✅ Extracted `voice-input.js` (853 LOC) — `DeepgramProvider`, `VoiceInput`
|
||||
4. ✅ Extracted `notification-manager.js` (445 LOC) — `NotificationManager` class (5-layer system)
|
||||
5. ✅ Extracted `keyboard-accessory.js` (279 LOC) — `KeyboardAccessoryBar`, `FocusTrap`
|
||||
6. ✅ Extracted `api-client.js` (70 LOC) — `_api()`, `_apiJson()`, `_apiPost()`, `_apiPut()`
|
||||
7. ✅ Extracted `subagent-windows.js` (1,119 LOC) — 13 subagent window methods
|
||||
8. ✅ Removed inlined xterm-zerolag-input copy → built to `vendor/xterm-zerolag-input.js` from `packages/xterm-zerolag-input/`
|
||||
9. ✅ Reduced `app.js` from ~15,200 → 11,473 LOC (24% reduction). All scripts loaded in correct dependency order in `index.html`
|
||||
|
||||
**Phase 6 - Config Consolidation (1 day)** ✅ COMPLETE
|
||||
1. ✅ Created 6 new domain-focused config files (better than plan's 2 generic files): `server-timing.ts` (13 constants), `auth-config.ts` (5 constants), `tunnel-config.ts` (8 constants), `terminal-limits.ts` (4 constants), `ai-defaults.ts` (3 constants), `team-config.ts` (3 constants)
|
||||
2. ✅ Total: 9 config files in `src/config/`, ~65 constants centralized
|
||||
3. ✅ Eliminated all cross-file duplicates: `STATS_COLLECTION_INTERVAL_MS` (was in 2 files), `timeout: 10000` (was 6× inline in hooks-config.ts → `HOOK_TIMEOUT_MS`), AI model string (was in 5 files → `AI_CHECK_MODEL`), `MAX_TRACKED_AGENTS` (was shadowed in subagent-watcher.ts)
|
||||
4. ✅ CLAUDE.md updated with config files table, import conventions, resource limits references
|
||||
|
||||
**Phase 7 - Test Infrastructure (2-3 days)** ✅ COMPLETE
|
||||
1. ✅ Created `test/mocks/` directory with 5 files (541 LOC): `mock-session.ts` (312), `mock-state-store.ts` (60), `mock-route-context.ts` (121), `test-helpers.ts` (37), `index.ts` (11 — barrel export)
|
||||
2. ✅ Consolidated MockSession into single shared definition — no duplicate class definitions remain (2 `vi.mock()`-based copies intentionally left in session-manager.test.ts and ralph-loop.test.ts)
|
||||
3. ✅ `respawn-test-utils.ts` converted to backward-compatibility shim — re-exports from `test/mocks/`, retains respawn-specific utilities (MockAiIdleChecker, TimeController, etc.)
|
||||
4. ✅ Created initial 3 route test files with 58 total tests: `session-routes.test.ts` (34 tests), `respawn-routes.test.ts` (13 tests), `system-routes.test.ts` (11 tests). Route test harness uses `app.inject()` — no real ports needed
|
||||
5. ✅ All 12 route modules now have dedicated test files in `test/routes/`: session, respawn, system, ralph, plan, push, team, mux, file, scheduled, hook-event, case
|
||||
|
||||
---
|
||||
|
||||
## Appendix: File Size Inventory (Post-Implementation)
|
||||
|
||||
### Before vs After
|
||||
|
||||
| File | Before | After | Change |
|
||||
|------|--------|-------|--------|
|
||||
| `src/web/server.ts` | 6,736 | 2,697 | **−60%** (routes, auth, ports extracted) |
|
||||
| `src/web/public/app.js` | 15,196 | 11,473 | **−24%** (8 modules extracted) |
|
||||
| `src/ralph-tracker.ts` | 3,905 | 2,391 | **−39%** (6 companion files extracted) |
|
||||
| `src/respawn-controller.ts` | 3,611 | 3,228 | **−11%** (4 companion files extracted) |
|
||||
| `src/session.ts` | 2,418 | 2,168 | **−10%** (5 companion files extracted) |
|
||||
| `src/types.ts` | 1,443 | 1 | **−99%** (14 domain files in `src/types/`) |
|
||||
|
||||
### New Infrastructure Created
|
||||
|
||||
| Directory | Files | Total LOC | Purpose |
|
||||
|-----------|-------|-----------|---------|
|
||||
| `src/web/routes/` | 13 | 4,090 | Domain route modules |
|
||||
| `src/web/ports/` | 7 | 142 | Port interfaces for DI |
|
||||
| `src/web/middleware/` | 1 | 193 | Auth middleware |
|
||||
| `src/types/` | 14 | 1,469 | Domain type files |
|
||||
| `src/config/` | 9 | ~450 | Centralized config |
|
||||
| `test/mocks/` | 5 | 541 | Shared test mocks |
|
||||
| `test/routes/` | 4 | ~500 | Route handler tests |
|
||||
|
||||
### Extracted Frontend Modules
|
||||
|
||||
| Module | Lines | Purpose |
|
||||
|--------|-------|---------|
|
||||
| `subagent-windows.js` | 1,119 | Subagent window management |
|
||||
| `voice-input.js` | 853 | DeepgramProvider, VoiceInput |
|
||||
| `mobile-handlers.js` | 449 | MobileDetection, KeyboardHandler, SwipeHandler |
|
||||
| `notification-manager.js` | 445 | 5-layer notification system |
|
||||
| `keyboard-accessory.js` | 279 | KeyboardAccessoryBar, FocusTrap |
|
||||
| `constants.js` | 238 | Shared constants, timing, Z-index |
|
||||
| `api-client.js` | 70 | API fetch wrapper |
|
||||
|
||||
### What's Working Well
|
||||
|
||||
These patterns should be **preserved, not refactored**:
|
||||
- Clean one-way dependency graph (no circular deps)
|
||||
- EventEmitter-based decoupling between domain models
|
||||
- Proper `import type` usage (19 files, consistent)
|
||||
- Utility type adoption (101 instances of Record, Partial, Omit, etc.)
|
||||
- `assertNever()` for exhaustive switch checking
|
||||
- `StaleExpirationMap` and `LRUMap` for bounded collections
|
||||
- State persistence circuit breaker pattern
|
||||
- TypeScript strict mode with all safety flags enabled
|
||||
- `CleanupManager` for centralized timer/watcher disposal
|
||||
- `Debouncer`/`KeyedDebouncer` for consistent debounce patterns
|
||||
- Port interfaces for route module dependency injection
|
||||
- `Object.assign(CodemanApp.prototype, ...)` for frontend module composition
|
||||
@@ -1,409 +0,0 @@
|
||||
# First-Load Performance Optimization Plan
|
||||
|
||||
**Date**: 2026-02-18
|
||||
**Audit by**: 4-agent team (css-analyst, js-analyst, server-analyst, deps-analyst)
|
||||
**Scope**: First browser load of Codeman web UI at `/`
|
||||
|
||||
---
|
||||
|
||||
## Current State (Baseline)
|
||||
|
||||
### Payload
|
||||
|
||||
| Asset | Raw | Compressed | Render-Blocking? |
|
||||
|-------|-----|-----------|-----------------|
|
||||
| `index.html` | 82 KB | ~15 KB | N/A (document) |
|
||||
| `styles.css` | 154 KB | ~25 KB | **YES** |
|
||||
| `mobile.css` | 34 KB | ~7 KB | **YES** (missing media query) |
|
||||
| `xterm.css` (CDN) | 2 KB | ~2 KB | No (preload pattern) |
|
||||
| `xterm.min.js` (CDN) | 67 KB | ~65 KB | No (defer) |
|
||||
| `xterm-addon-fit` (CDN) | 1 KB | ~1 KB | No (defer) |
|
||||
| `app.js` | 563 KB | ~126 KB | No (defer) |
|
||||
| **Total** | **903 KB** | **~241 KB** | |
|
||||
|
||||
### Request Waterfall (13 requests on first load)
|
||||
|
||||
```
|
||||
T=0 GET / (82KB doc)
|
||||
T+20ms ├── styles.css?v=0.1536 (154KB — BLOCKS RENDER)
|
||||
├── mobile.css?v=0.1536 (34KB — BLOCKS RENDER on all viewports!)
|
||||
├── xterm.css (CDN, preloaded) (2KB — non-blocking, already async)
|
||||
├── xterm.min.js (CDN, defer) (67KB)
|
||||
├── xterm-addon-fit.min.js (CDN) (1KB)
|
||||
└── app.js?v=0.1536 (defer) (563KB)
|
||||
|
||||
[FIRST PAINT blocked by: styles.css + mobile.css]
|
||||
|
||||
T+200ms JS execution starts
|
||||
├── new Terminal() + terminal.open() ← HEAVY sync (canvas creation)
|
||||
├── connectSSE() → /api/events ← SSE stream
|
||||
├── loadState() → /api/status ← DUPLICATE of SSE init!
|
||||
├── loadQuickStartCases()
|
||||
│ ├── /api/settings ← fetched TWICE
|
||||
│ └── /api/cases?_t=<timestamp> ← cache-busted unnecessarily
|
||||
├── startSystemStatsPolling()
|
||||
│ └── /api/system/stats ← starts immediately, every 2s
|
||||
└── loadAppSettingsFromServer()
|
||||
└── /api/settings ← DUPLICATE #2
|
||||
|
||||
T+500ms First Meaningful Paint (terminal + header visible)
|
||||
```
|
||||
|
||||
### Problems
|
||||
|
||||
1. **2 render-blocking CSS files** — mobile.css blocks desktop for no reason
|
||||
2. **Double handleInit()** — SSE init + /api/status both call full state reset
|
||||
3. **Duplicate /api/settings** — fetched twice in init chain
|
||||
4. **563KB unminified JS monolith** — no build minification at all
|
||||
5. **154KB unminified CSS** — 70% is for modals/wizards (below-the-fold)
|
||||
6. **Sync terminal.open()** — heaviest single call, blocks before first paint
|
||||
7. **12 modals pre-rendered** — ~600+ hidden DOM nodes, ~60KB HTML
|
||||
8. **Stats polling starts immediately** — even with 0 sessions
|
||||
9. **No loading skeleton** — blank black screen until all CSS+JS loads
|
||||
10. **CDN dependency** — 3 xterm files from jsdelivr (DNS+TLS latency)
|
||||
11. **1h cache for versioned assets** — could be 1yr+immutable with ?v= busting
|
||||
12. **No HTTP/2** — 6-connection limit queues some requests
|
||||
13. **On-the-fly compression** — no pre-compressed .gz/.br files
|
||||
|
||||
### What's Already Good (don't touch)
|
||||
|
||||
- Single shared Terminal instance (buffer swapping)
|
||||
- Teammate terminals created lazily on window open
|
||||
- Subagent windows use HTML logs, not Terminal instances
|
||||
- `getLightState()` has 1s TTL cache
|
||||
- SSE init sends lightweight state (no terminal buffers)
|
||||
- Buffer hydration uses chunked writes (128KB via rAF)
|
||||
- `selectSession()` defers secondary panels via requestIdleCallback
|
||||
- System fonts only — zero web font loading
|
||||
- xterm.css already uses async preload pattern
|
||||
- Proper SSE reconnection with exponential backoff
|
||||
- CSS `contain` on header/tabs for layout isolation
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan (15 steps, ordered by impact/effort)
|
||||
|
||||
### Phase 1: Quick Wins (1-line to 15-min changes)
|
||||
|
||||
#### Step 1: Add media attribute to mobile.css
|
||||
**Impact**: HIGH — 34KB stops blocking render on desktop
|
||||
**File**: `src/web/public/index.html:14`
|
||||
|
||||
```html
|
||||
<!-- BEFORE -->
|
||||
<link rel="stylesheet" href="mobile.css?v=0.1536">
|
||||
|
||||
<!-- AFTER -->
|
||||
<link rel="stylesheet" href="mobile.css?v=0.1536" media="(max-width: 1023px)">
|
||||
```
|
||||
|
||||
Browser still downloads it (for potential resize) but won't block rendering on desktop. The mobile.css file header says this was intended but never implemented.
|
||||
|
||||
---
|
||||
|
||||
#### Step 2: Remove duplicate /api/status + double handleInit()
|
||||
**Impact**: HIGH — eliminates redundant API call + double state reset (clears 15+ Maps, 7+ timers, runs cleanupAllFloatingWindows(), double renderSessionTabs())
|
||||
**Files**: `src/web/public/app.js`
|
||||
|
||||
The SSE `init` event (server.ts:618) sends `getLightState()`. The `loadState()` in `init()` at `app.js:1554` fetches identical data from `/api/status`. Both call `handleInit()` which wipes state. The `_initGeneration` guard only protects session-restore, NOT the expensive cleanup (lines 3389-3503).
|
||||
|
||||
```js
|
||||
// In init() — REMOVE this.loadState(), add SSE fallback:
|
||||
this.connectSSE();
|
||||
// Remove: this.loadState();
|
||||
this._initFallbackTimer = setTimeout(() => {
|
||||
if (this._initGeneration === 0) this.loadState();
|
||||
}, 3000);
|
||||
|
||||
// In handleInit() — clear fallback timer:
|
||||
handleInit(data) {
|
||||
if (this._initFallbackTimer) {
|
||||
clearTimeout(this._initFallbackTimer);
|
||||
this._initFallbackTimer = null;
|
||||
}
|
||||
// ... rest of handleInit
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### Step 3: Deduplicate /api/settings fetch
|
||||
**Impact**: MEDIUM — removes 1 redundant API call
|
||||
**Files**: `src/web/public/app.js:7341` (loadQuickStartCases), `app.js:9964` (loadAppSettingsFromServer)
|
||||
|
||||
```js
|
||||
// In init() — fetch settings once, share the promise:
|
||||
const settingsPromise = fetch('/api/settings').then(r => r.json());
|
||||
this.loadQuickStartCases(null, settingsPromise);
|
||||
this.loadAppSettingsFromServer(settingsPromise);
|
||||
```
|
||||
|
||||
Both functions need to accept an optional pre-fetched settings promise parameter.
|
||||
|
||||
---
|
||||
|
||||
#### Step 4: Remove cache-busting from /api/cases
|
||||
**Impact**: LOW — allows browser caching
|
||||
**File**: `src/web/public/app.js:7351`
|
||||
|
||||
```js
|
||||
// BEFORE
|
||||
const res = await fetch('/api/cases?_t=' + Date.now());
|
||||
// AFTER
|
||||
const res = await fetch('/api/cases');
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### Step 5: Defer system stats polling
|
||||
**Impact**: MEDIUM — removes API call every 2s when idle
|
||||
**Files**: `src/web/public/app.js:1567`, `app.js:15261`
|
||||
|
||||
Move `startSystemStatsPolling()` out of `init()`. Start it in `handleInit()` only when `data.sessions.length > 0`.
|
||||
|
||||
---
|
||||
|
||||
### Phase 2: Build Pipeline (30-min changes, highest payload impact)
|
||||
|
||||
#### Step 6: Self-host xterm.js assets
|
||||
**Impact**: MEDIUM-HIGH — eliminates CDN DNS/TLS latency (~100ms even with preconnect)
|
||||
**Files**: `src/web/public/index.html`, `package.json` build script
|
||||
|
||||
xterm is NOT in package.json — add it:
|
||||
```bash
|
||||
npm install xterm@5.3.0 @xterm/addon-fit@0.8.0 --save
|
||||
```
|
||||
|
||||
Build script addition:
|
||||
```bash
|
||||
mkdir -p dist/web/public/vendor
|
||||
cp node_modules/xterm/css/xterm.css dist/web/public/vendor/
|
||||
cp node_modules/xterm/lib/xterm.min.js dist/web/public/vendor/
|
||||
cp node_modules/@xterm/addon-fit/lib/xterm-addon-fit.min.js dist/web/public/vendor/
|
||||
```
|
||||
|
||||
Update index.html CDN URLs to `/vendor/xterm.min.js` etc. Remove preconnect/dns-prefetch for jsdelivr.
|
||||
|
||||
---
|
||||
|
||||
#### Step 7: Add esbuild minification to build
|
||||
**Impact**: HIGH — biggest single optimization for payload size
|
||||
**File**: `package.json` build script
|
||||
|
||||
Current build just does `cp -r src/web/public dist/web/`. No minification.
|
||||
|
||||
app.js stats: 1,525 comment lines (10%), 89 console.* statements, 23% whitespace.
|
||||
|
||||
```bash
|
||||
# Add to build script after cp:
|
||||
npx esbuild dist/web/public/app.js --minify --drop:console --outfile=dist/web/public/app.js --allow-overwrite
|
||||
npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite
|
||||
npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite
|
||||
```
|
||||
|
||||
Expected savings:
|
||||
| File | Before (gzip) | After (gzip) | Saved |
|
||||
|------|---------------|-------------|-------|
|
||||
| app.js | ~126 KB | ~85 KB | ~41 KB (33%) |
|
||||
| styles.css | ~25 KB | ~18 KB | ~7 KB (28%) |
|
||||
| mobile.css | ~7 KB | ~5 KB | ~2 KB (29%) |
|
||||
| **Total** | **~158 KB** | **~108 KB** | **~50 KB** |
|
||||
|
||||
---
|
||||
|
||||
#### Step 8: Pre-compress static assets at build time
|
||||
**Impact**: MEDIUM — eliminates per-request CPU compression
|
||||
**Files**: `package.json` build script, potentially `src/web/server.ts`
|
||||
|
||||
```bash
|
||||
# Add to build script after minification:
|
||||
for f in dist/web/public/*.{js,css,html}; do
|
||||
gzip -9 -k "$f"
|
||||
brotli -9 -k "$f"
|
||||
done
|
||||
```
|
||||
|
||||
Check if `@fastify/static` supports `preCompressed: true` option. If not, serve pre-compressed files via custom Accept-Encoding check.
|
||||
|
||||
---
|
||||
|
||||
#### Step 9: Extend cache duration for versioned assets
|
||||
**Impact**: LOW (first load) / HIGH (repeat visits)
|
||||
**File**: `src/web/server.ts:601`
|
||||
|
||||
```js
|
||||
// BEFORE
|
||||
maxAge: '1h'
|
||||
|
||||
// AFTER
|
||||
maxAge: '1y',
|
||||
immutable: true
|
||||
```
|
||||
|
||||
Safe because all assets use `?v=0.1536` cache-busting. First-load unaffected, but all repeat visits serve from disk cache instantly.
|
||||
|
||||
---
|
||||
|
||||
### Phase 3: Perceived Performance (30-60min, user experience)
|
||||
|
||||
#### Step 10: Add loading skeleton
|
||||
**Impact**: MEDIUM-HIGH — instant visual structure instead of black screen
|
||||
**File**: `src/web/public/index.html`
|
||||
|
||||
Add minimal inline `<style>` + skeleton HTML in `<body>`:
|
||||
|
||||
```html
|
||||
<style>
|
||||
.skeleton { display: flex; flex-direction: column; height: 100vh; background: #0a0a0a; }
|
||||
.skeleton-header { height: 40px; background: #111; border-bottom: 1px solid #222; }
|
||||
.skeleton-terminal { flex: 1; background: #0d0d0d; }
|
||||
.app-loaded .skeleton { display: none; }
|
||||
</style>
|
||||
<div class="skeleton">
|
||||
<div class="skeleton-header"></div>
|
||||
<div class="skeleton-terminal"></div>
|
||||
</div>
|
||||
```
|
||||
|
||||
In `app.js` init() end: `document.body.classList.add('app-loaded');`
|
||||
|
||||
---
|
||||
|
||||
#### Step 11: Defer terminal creation to after first paint
|
||||
**Impact**: MEDIUM-HIGH — terminal.open() is heaviest sync call
|
||||
**File**: `src/web/public/app.js:1545`
|
||||
|
||||
```js
|
||||
init() {
|
||||
// ... mobile detection, visibility settings ...
|
||||
document.documentElement.classList.remove('mobile-init');
|
||||
|
||||
// Show skeleton immediately, defer heavy terminal init
|
||||
requestAnimationFrame(() => {
|
||||
this.initTerminal();
|
||||
this.connectSSE();
|
||||
// ... rest of init
|
||||
});
|
||||
}
|
||||
```
|
||||
|
||||
Lets browser paint header/tabs/skeleton before canvas creation.
|
||||
|
||||
---
|
||||
|
||||
### Phase 4: DOM + Payload Reduction (1-3 hours)
|
||||
|
||||
#### Step 12: Lazy-create modals on first open
|
||||
**Impact**: HIGH — removes ~600+ DOM nodes, ~60KB hidden HTML
|
||||
**Files**: `src/web/public/index.html`, `src/web/public/app.js`
|
||||
|
||||
12 modals pre-rendered in index.html:
|
||||
- `helpModal` (lines 227-447)
|
||||
- `sessionOptionsModal` (lines 448-714) — 266 lines
|
||||
- `appSettingsModal` (lines 715-900+)
|
||||
- `createCaseModal`, `mobileCasePickerModal`, `ralphWizardModal`, `killAllModal`, `closeConfirmModal`, `savePresetModal`, `tokenStatsModal`, `filePreviewModal`, notification drawer
|
||||
|
||||
Replace each modal's HTML with `<div id="helpModal" class="modal"></div>`. On first open, inject full HTML via template function. Cache after creation.
|
||||
|
||||
---
|
||||
|
||||
#### Step 13: Batch initial API calls into one endpoint
|
||||
**Impact**: MEDIUM — reduces 4+ API calls to 1
|
||||
**Files**: `src/web/server.ts`, `src/web/public/app.js`
|
||||
|
||||
Create `GET /api/init-bundle`:
|
||||
```json
|
||||
{
|
||||
"status": { /* getLightState() */ },
|
||||
"cases": [ /* case list */ ],
|
||||
"settings": { /* user settings */ }
|
||||
}
|
||||
```
|
||||
|
||||
Use as SSE init fallback (step 2's timeout). Saves HTTP round trips.
|
||||
|
||||
---
|
||||
|
||||
#### Step 14: Trim SSE init payload
|
||||
**Impact**: LOW-MEDIUM — reduces init payload by removing data not needed for first paint
|
||||
**File**: `src/web/server.ts`
|
||||
|
||||
Remove from SSE init event: `taskTree`, `ralphTodos`, `ralphTodoStats` per session. These can be fetched on-demand when user opens a session's details panel.
|
||||
|
||||
---
|
||||
|
||||
#### Step 15: Enable HTTP/2
|
||||
**Impact**: MEDIUM — multiplexed loading over single connection
|
||||
**File**: `src/web/server.ts`
|
||||
|
||||
```js
|
||||
// BEFORE
|
||||
const server = Fastify({ logger: false });
|
||||
|
||||
// AFTER (when HTTPS is enabled)
|
||||
const server = Fastify({
|
||||
logger: false,
|
||||
http2: true // Only works with HTTPS
|
||||
});
|
||||
```
|
||||
|
||||
Only applicable for `--https` mode. HTTP/2 multiplexing eliminates the 6-connection limit queuing.
|
||||
|
||||
---
|
||||
|
||||
## Expected Combined Impact
|
||||
|
||||
| Metric | Before | After | Improvement |
|
||||
|--------|--------|-------|-------------|
|
||||
| First Paint | ~300ms | ~100ms | **-200ms** (skeleton visible instantly) |
|
||||
| First Contentful Paint | ~400ms | ~200ms | **-200ms** (no mobile.css blocking desktop) |
|
||||
| Time to Interactive | ~600ms | ~350ms | **-250ms** (fewer API calls, deferred terminal) |
|
||||
| Total compressed payload | ~241 KB | ~191 KB | **-50 KB (21%)** via minification |
|
||||
| Init API calls | 6-7 (2 dupes) | 2-3 | **-60%** fewer requests |
|
||||
| Initial DOM nodes | ~1800+ | ~1200 | **-600** (lazy modals) |
|
||||
|
||||
---
|
||||
|
||||
## Verification
|
||||
|
||||
After each step, verify with Playwright:
|
||||
|
||||
```js
|
||||
const { chromium } = require('playwright');
|
||||
const browser = await chromium.launch();
|
||||
const page = await browser.newPage();
|
||||
|
||||
// Measure first paint
|
||||
await page.goto('http://localhost:3000', { waitUntil: 'domcontentloaded' });
|
||||
await page.waitForTimeout(4000); // Wait for async data
|
||||
|
||||
// Check UI renders correctly
|
||||
const header = await page.locator('.header').isVisible();
|
||||
const tabs = await page.locator('.session-tabs').isVisible();
|
||||
const terminal = await page.locator('.terminal-container').isVisible();
|
||||
console.log({ header, tabs, terminal });
|
||||
|
||||
await browser.close();
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Files Changed Per Step (for implementation agent)
|
||||
|
||||
| Step | Files Modified |
|
||||
|------|---------------|
|
||||
| 1 | `index.html` |
|
||||
| 2 | `app.js` |
|
||||
| 3 | `app.js` |
|
||||
| 4 | `app.js` |
|
||||
| 5 | `app.js` |
|
||||
| 6 | `index.html`, `package.json` |
|
||||
| 7 | `package.json` |
|
||||
| 8 | `package.json`, optionally `server.ts` |
|
||||
| 9 | `server.ts` |
|
||||
| 10 | `index.html`, `app.js` |
|
||||
| 11 | `app.js` |
|
||||
| 12 | `index.html`, `app.js` |
|
||||
| 13 | `server.ts`, `app.js` |
|
||||
| 14 | `server.ts` |
|
||||
| 15 | `server.ts` |
|
||||
@@ -1,388 +0,0 @@
|
||||
# Performance Audit: First Page Load
|
||||
|
||||
**Date**: 2026-02-18
|
||||
**Scope**: Browser first-load of Codeman web UI (`/`)
|
||||
**Method**: Static analysis by 4 parallel audit agents (server, frontend, SSE/xterm, asset pipeline)
|
||||
|
||||
---
|
||||
|
||||
## Current State Summary
|
||||
|
||||
### Payload Sizes (measured from live server, port 3000)
|
||||
|
||||
| Asset | Raw Size | Gzip | Brotli | Lines | Render-Blocking? |
|
||||
|-------|----------|------|--------|-------|-----------------|
|
||||
| `index.html` | 82 KB | 15 KB | 15 KB | 1,479 | N/A (document) |
|
||||
| `app.js` | 562 KB | 126 KB | 125 KB | 15,354 | No (`defer`) |
|
||||
| `styles.css` | 154 KB | 25 KB | 27 KB | 8,199 | **YES** |
|
||||
| `mobile.css` | 34 KB | 7 KB | 7 KB | 1,493 | **YES** (no media query!) |
|
||||
| `xterm.css` (CDN) | 2 KB | 2 KB | — | — | **YES** (external CDN) |
|
||||
| `xterm.min.js` (CDN) | 67 KB | 65 KB | — | — | No (`defer`) |
|
||||
| `xterm-addon-fit` (CDN) | 1 KB | 1 KB | — | — | No (`defer`) |
|
||||
| **Total local** | **832 KB** | **173 KB** | **174 KB** | | |
|
||||
| **Total w/ CDN** | **~902 KB** | **~241 KB** | | | |
|
||||
|
||||
**Server compression**: Brotli preferred (`Content-Encoding: br`), via `@fastify/compress` with threshold 1024. Compression is **on-the-fly per request** — no pre-compressed files exist.
|
||||
|
||||
**HTTP headers verified**: `Cache-Control: public, max-age=3600`, weak ETags auto-generated by `@fastify/static`, `Vary: accept-encoding`, CSP + security headers present.
|
||||
|
||||
### Request Waterfall on First Load (6-7 API calls!)
|
||||
|
||||
```
|
||||
Browser hits /
|
||||
├── index.html ............................ (82 KB document)
|
||||
├── styles.css?v=0.1533 .................. (render-blocking CSS, 154 KB)
|
||||
├── mobile.css?v=0.1533 .................. (render-blocking CSS, 34 KB — wasted on desktop!)
|
||||
├── xterm.css (CDN) ...................... (render-blocking CSS — external!)
|
||||
├── xterm.min.js (CDN, defer) ........... (67 KB, parallel download)
|
||||
├── xterm-addon-fit.min.js (CDN, defer) .. (1 KB, parallel download)
|
||||
├── app.js?v=0.1533 (defer) ............. (562 KB, parallel download)
|
||||
│
|
||||
│ [FIRST PAINT blocked until ALL CSS downloaded + parsed]
|
||||
│
|
||||
├── JS executes: new CodemanApp().init()
|
||||
│ ├── initTerminal() ................... (SYNC: new Terminal() + terminal.open() → canvas creation)
|
||||
│ ├── connectSSE() → /api/events ....... (SSE → fires 'init' with getLightState())
|
||||
│ ├── loadState() → /api/status ........ (DUPLICATE #1: same data as SSE init!)
|
||||
│ ├── loadQuickStartCases()
|
||||
│ │ ├── /api/settings ................ (settings fetch #1)
|
||||
│ │ └── /api/cases?_t=<timestamp> ... (case list, cache-busted!)
|
||||
│ ├── startSystemStatsPolling() → /api/system/stats (every 2s, starts immediately)
|
||||
│ └── loadAppSettingsFromServer() → /api/settings (DUPLICATE #2: settings fetched again!)
|
||||
```
|
||||
|
||||
**Total init API calls**: 6-7 requests, with **2 duplicates** (`/api/status` = SSE init, `/api/settings` fetched twice).
|
||||
|
||||
### Critical Path Bottlenecks
|
||||
|
||||
1. **3 render-blocking CSS files** (one from CDN, one wasted on desktop)
|
||||
2. **Synchronous `terminal.open()`** blocks main thread during init (canvas creation)
|
||||
3. **Double `handleInit()` execution** — SSE init + `/api/status` both call it, causing full state reset + cleanup twice within ~100ms
|
||||
4. **`/api/settings` fetched twice** — once in `loadQuickStartCases()`, once in `loadAppSettingsFromServer()`
|
||||
5. **No loading skeleton** — blank `#0a0a0a` screen until CSS+JS fully loaded
|
||||
6. **12 modals pre-rendered** in HTML — ~600+ DOM elements, ~60KB of invisible HTML
|
||||
7. **562KB monolith `app.js`** unminified — 1,525 comment lines (10%), 89 `console.*` statements, 23% whitespace
|
||||
8. **No minification in build** — `cp -r` copies raw source to dist
|
||||
9. **Stats polling starts immediately** — 2s interval even with no sessions
|
||||
10. **Version query strings stale** — HTML has `?v=0.1533`, package.json is `0.1534`
|
||||
|
||||
### What's Already Good
|
||||
|
||||
- Only **1 xterm Terminal instance** shared across all sessions (buffer swapping on tab switch)
|
||||
- Teammate terminals created **lazily** on window open (with `requestAnimationFrame` defer)
|
||||
- Subagent windows use **HTML activity logs**, not additional Terminal instances
|
||||
- `getLightState()` has a **1-second TTL cache** — no duplicate server-side computation
|
||||
- SSE init sends **lightweight state** (no terminal buffers) — buffers fetched on-demand per tab
|
||||
- Buffer hydration uses **chunked writes** (128KB chunks via `requestAnimationFrame`) — no UI jank
|
||||
- `selectSession()` defers secondary panels via **`requestIdleCallback`**
|
||||
- Buffer fetch is **tail-mode** (last 256KB only, not full 2MB)
|
||||
- **System fonts only** — no web font downloads blocking paint
|
||||
- All JS scripts use **`defer`**
|
||||
- SSE reconnection has **proper exponential backoff** with timeout cleanup
|
||||
|
||||
---
|
||||
|
||||
## Optimization Plan
|
||||
|
||||
### Phase 1: Quick Wins (High Impact, Low Effort)
|
||||
|
||||
#### 1.1 Add `media` attribute to mobile.css
|
||||
**Impact**: HIGH — 34KB CSS stops blocking render on desktop
|
||||
**Effort**: 1 line change
|
||||
**File**: `src/web/public/index.html:14`
|
||||
|
||||
```html
|
||||
<!-- Before -->
|
||||
<link rel="stylesheet" href="mobile.css?v=...">
|
||||
|
||||
<!-- After -->
|
||||
<link rel="stylesheet" href="mobile.css?v=..." media="(max-width: 1023px)">
|
||||
```
|
||||
|
||||
The browser still downloads it (for potential resize) but won't block rendering on desktop. The `mobile.css` comment on line 4 says this was *intended* but never implemented.
|
||||
|
||||
#### 1.2 Eliminate duplicate `/api/status` fetch + double `handleInit()`
|
||||
**Impact**: HIGH — removes 1 redundant API call + eliminates double state reset (clearing 15+ Maps, 7+ timers, `cleanupAllFloatingWindows()`, double `renderSessionTabs()`, double async subagent restore chain)
|
||||
**Effort**: Small
|
||||
**Files**: `src/web/public/app.js:1554`, `app.js:3566-3574`
|
||||
|
||||
The SSE `init` event (`server.ts:618`) already sends `getLightState()`. The `loadState()` at `app.js:1554` fetches identical data from `/api/status`. Both call `handleInit()` which does a full state reset — whichever arrives second **wipes all state from the first** and rebuilds from scratch.
|
||||
|
||||
The `_initGeneration` guard (line 3373/3549) only protects the session-restore at the end, NOT the expensive full cleanup (lines 3389-3503).
|
||||
|
||||
**Approach**: Remove `this.loadState()` from `init()`. Add a fallback timeout:
|
||||
|
||||
```js
|
||||
// In init():
|
||||
this.connectSSE();
|
||||
// Remove: this.loadState();
|
||||
this._initFallbackTimer = setTimeout(() => {
|
||||
if (this._initGeneration === 0) this.loadState();
|
||||
}, 3000);
|
||||
```
|
||||
|
||||
Clear the timer in `handleInit()`:
|
||||
```js
|
||||
handleInit(data) {
|
||||
if (this._initFallbackTimer) {
|
||||
clearTimeout(this._initFallbackTimer);
|
||||
this._initFallbackTimer = null;
|
||||
}
|
||||
// ... rest of handleInit
|
||||
}
|
||||
```
|
||||
|
||||
#### 1.3 Deduplicate `/api/settings` fetch
|
||||
**Impact**: MEDIUM — removes 1 redundant API call
|
||||
**Effort**: Small
|
||||
**Files**: `src/web/public/app.js:7341` (in `loadQuickStartCases`), `app.js:9964` (in `loadAppSettingsFromServer`)
|
||||
|
||||
Both fetch `/api/settings`. Fetch it once, pass the result to both consumers:
|
||||
|
||||
```js
|
||||
// In init():
|
||||
const settingsPromise = fetch('/api/settings').then(r => r.json());
|
||||
this.loadQuickStartCases(null, settingsPromise);
|
||||
this.loadAppSettingsFromServer(settingsPromise);
|
||||
```
|
||||
|
||||
#### 1.4 Defer system stats polling
|
||||
**Impact**: MEDIUM — removes 1 API call every 2s when idle
|
||||
**Effort**: Small
|
||||
**Files**: `src/web/public/app.js:1567`, `app.js:15261-15271`
|
||||
|
||||
`fetchSystemStats()` already has a visibility guard (line 15282: skips if `#headerSystemStats` is `display: none`), but the interval still ticks. Move `startSystemStatsPolling()` out of `init()` — start it in `handleInit()` only when `data.sessions.length > 0`.
|
||||
|
||||
#### 1.5 Preload xterm.css to unblock render
|
||||
**Impact**: MEDIUM — external CDN CSS currently blocks first paint
|
||||
**Effort**: 2 line change
|
||||
**File**: `src/web/public/index.html:15`
|
||||
|
||||
```html
|
||||
<!-- Before -->
|
||||
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css">
|
||||
|
||||
<!-- After -->
|
||||
<link rel="preload" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css" as="style" onload="this.onload=null;this.rel='stylesheet'">
|
||||
<noscript><link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css"></noscript>
|
||||
```
|
||||
|
||||
Terminal won't display until xterm.js executes anyway, so the CSS doesn't need to block initial paint.
|
||||
|
||||
#### 1.6 Fix stale version query strings
|
||||
**Impact**: LOW — prevents serving cached stale assets after deploy
|
||||
**Effort**: Small
|
||||
**File**: COM script in CLAUDE.md
|
||||
|
||||
The HTML references `?v=0.1533` while package.json is already at `0.1534`. The COM workflow should auto-update HTML version strings. Add to the COM script:
|
||||
|
||||
```bash
|
||||
# After incrementing version in package.json + CLAUDE.md:
|
||||
sed -i "s/?v=[0-9.]*/?v=$NEW_VERSION/g" src/web/public/index.html
|
||||
```
|
||||
|
||||
#### 1.7 Remove cache-busting from `/api/cases`
|
||||
**Impact**: LOW — allows HTTP caching of case list
|
||||
**Effort**: 1 line change
|
||||
**File**: `src/web/public/app.js:7351`
|
||||
|
||||
```js
|
||||
// Before:
|
||||
const res = await fetch('/api/cases?_t=' + Date.now());
|
||||
// After:
|
||||
const res = await fetch('/api/cases');
|
||||
```
|
||||
|
||||
The case list rarely changes during a session. Let the browser cache it.
|
||||
|
||||
---
|
||||
|
||||
### Phase 2: Medium Effort (High Impact)
|
||||
|
||||
#### 2.1 Add loading skeleton
|
||||
**Impact**: MEDIUM-HIGH — perceived performance improvement (instant visual structure)
|
||||
**Effort**: Small-Medium
|
||||
**File**: `src/web/public/index.html`
|
||||
|
||||
Add minimal inline `<style>` + skeleton HTML in `<body>` showing a dark header bar + terminal placeholder. Hidden by `app.js` once init completes:
|
||||
|
||||
```html
|
||||
<style>
|
||||
.skeleton { display: flex; flex-direction: column; height: 100vh; }
|
||||
.skeleton-header { height: 40px; background: #111; border-bottom: 1px solid #222; }
|
||||
.skeleton-terminal { flex: 1; background: #0d0d0d; }
|
||||
.app-loaded .skeleton { display: none; }
|
||||
</style>
|
||||
<div class="skeleton">
|
||||
<div class="skeleton-header"></div>
|
||||
<div class="skeleton-terminal"></div>
|
||||
</div>
|
||||
```
|
||||
|
||||
In `app.js` init(), add `document.body.classList.add('app-loaded')` at the end.
|
||||
|
||||
#### 2.2 Defer xterm.js terminal creation to after first paint
|
||||
**Impact**: MEDIUM-HIGH — `terminal.open()` is the heaviest synchronous call in init
|
||||
**Effort**: Medium
|
||||
**Files**: `src/web/public/app.js:1545`, `app.js:1578-1639`
|
||||
|
||||
```js
|
||||
init() {
|
||||
// ... mobile detection, visibility settings ...
|
||||
document.documentElement.classList.remove('mobile-init');
|
||||
|
||||
// Show skeleton/header immediately, defer heavy terminal init
|
||||
requestAnimationFrame(() => {
|
||||
this.initTerminal();
|
||||
this.connectSSE();
|
||||
// ... rest of init
|
||||
});
|
||||
}
|
||||
```
|
||||
|
||||
Lets the browser paint the header/tabs before the terminal canvas is created.
|
||||
|
||||
#### 2.3 Batch initial API calls into one endpoint
|
||||
**Impact**: MEDIUM — reduces 4+ API calls to 1
|
||||
**Effort**: Medium
|
||||
**Files**: `src/web/server.ts`, `src/web/public/app.js`
|
||||
|
||||
Create `/api/init-bundle`:
|
||||
```json
|
||||
{
|
||||
"status": { /* getLightState() */ },
|
||||
"cases": [ /* case list */ ],
|
||||
"settings": { /* user settings */ }
|
||||
}
|
||||
```
|
||||
|
||||
Replaces `/api/status` (fallback), `/api/cases`, `/api/settings`. Saves HTTP round trips and server-side work.
|
||||
|
||||
#### 2.4 Lazy-create modals on first open
|
||||
**Impact**: HIGH — removes ~600+ DOM elements from initial parse (~60KB of HTML)
|
||||
**Effort**: Medium-High
|
||||
**Files**: `src/web/public/index.html`, `src/web/public/app.js`
|
||||
|
||||
12 modals pre-rendered in `index.html`:
|
||||
- `helpModal` (lines 227-447)
|
||||
- `sessionOptionsModal` (lines 448-714) — **266 lines alone**
|
||||
- `appSettingsModal` (lines 715-900+)
|
||||
- `createCaseModal`, `mobileCasePickerModal`, `ralphWizardModal`, `killAllModal`, `closeConfirmModal`, `savePresetModal`, `tokenStatsModal`, `filePreviewModal`, notification drawer
|
||||
|
||||
**Approach**: Replace each modal's HTML with `<div id="helpModal" class="modal"></div>`. On first open, inject full HTML via `createModalContent()`. Cache after creation.
|
||||
|
||||
---
|
||||
|
||||
### Phase 3: Build Pipeline (Highest Impact)
|
||||
|
||||
#### 3.1 Self-host xterm.js assets
|
||||
**Impact**: MEDIUM — eliminates CDN dependency + latency, enables local caching
|
||||
**Effort**: Low-Medium
|
||||
**Files**: `src/web/public/index.html`, `package.json` build script
|
||||
|
||||
```bash
|
||||
# Build script addition:
|
||||
mkdir -p dist/web/public/vendor
|
||||
cp node_modules/xterm/css/xterm.css dist/web/public/vendor/
|
||||
cp node_modules/xterm/lib/xterm.min.js dist/web/public/vendor/
|
||||
cp node_modules/@xterm/addon-fit/lib/xterm-addon-fit.min.js dist/web/public/vendor/
|
||||
```
|
||||
|
||||
Update HTML to reference `/vendor/xterm.min.js` etc. Removes render-blocking CDN CSS entirely.
|
||||
|
||||
#### 3.2 Add esbuild minification to build
|
||||
**Impact**: HIGH — ~38 KB compressed savings (16% of local payload)
|
||||
**Effort**: Medium
|
||||
**Files**: `package.json` (build script)
|
||||
|
||||
Current build just does `cp -r src/web/public dist/web/`. No minification at all.
|
||||
|
||||
**app.js specifics**: 1,525 comment lines (10%), 89 `console.*` statements, 23% whitespace.
|
||||
|
||||
```bash
|
||||
# Add to build script:
|
||||
npx esbuild dist/web/public/app.js --minify --drop:console --outfile=dist/web/public/app.js --allow-overwrite
|
||||
npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite
|
||||
npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite
|
||||
```
|
||||
|
||||
Expected: `app.js` 562KB → ~350KB minified → ~90KB gzip (from 126KB). `--drop:console` removes all 89 debug statements.
|
||||
|
||||
Note: `app.js` is vanilla JS (not modules), so esbuild works directly as a minifier.
|
||||
|
||||
#### 3.3 Pre-compress static assets at build time
|
||||
**Impact**: MEDIUM — eliminates per-request CPU compression work
|
||||
**Effort**: Low
|
||||
**Files**: `package.json` build script, `src/web/server.ts`
|
||||
|
||||
Currently `@fastify/compress` compresses on-the-fly for every request. Pre-compress at build time:
|
||||
|
||||
```bash
|
||||
# Build script:
|
||||
for f in dist/web/public/*.{js,css,html}; do
|
||||
gzip -9 -k "$f"
|
||||
brotli -9 -k "$f"
|
||||
done
|
||||
```
|
||||
|
||||
Then configure `@fastify/static` with `preCompressed: true` (if supported) or serve pre-compressed files via custom logic.
|
||||
|
||||
#### 3.4 Extract critical CSS inline
|
||||
**Impact**: MEDIUM — eliminates render-blocking `styles.css` for first paint
|
||||
**Effort**: Medium-High
|
||||
**Files**: `src/web/public/styles.css`, `src/web/public/index.html`
|
||||
|
||||
Identify ~2-3KB of CSS needed for first paint (body, header, tab bar, terminal container) and inline it in `<head>`. Load full `styles.css` asynchronously:
|
||||
|
||||
```html
|
||||
<style>/* ~50 lines of critical CSS */</style>
|
||||
<link rel="preload" href="styles.css?v=..." as="style" onload="this.onload=null;this.rel='stylesheet'">
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Impact Estimates
|
||||
|
||||
| # | Optimization | First Paint | TTI | Effort |
|
||||
|---|-------------|-------------|-----|--------|
|
||||
| 1.1 | mobile.css media query | -50ms | — | 1 min |
|
||||
| 1.2 | Remove duplicate fetch + double handleInit | — | -100-200ms | 15 min |
|
||||
| 1.3 | Deduplicate settings fetch | — | -50ms | 10 min |
|
||||
| 1.4 | Defer stats polling | — | -20ms | 10 min |
|
||||
| 1.5 | Preload xterm.css | -100-300ms | — | 5 min |
|
||||
| 1.6 | Fix stale version strings | cache correctness | — | 5 min |
|
||||
| 1.7 | Remove cases cache-bust | — | -10ms | 1 min |
|
||||
| 2.1 | Loading skeleton | perceived -500ms | — | 30 min |
|
||||
| 2.2 | Defer terminal init | -50-100ms | -50ms | 30 min |
|
||||
| 2.3 | Batch API endpoint | — | -100-200ms | 1 hr |
|
||||
| 2.4 | Lazy modals | -30-50ms parse | -50ms | 2-3 hrs |
|
||||
| 3.1 | Self-host xterm | -100-300ms | — | 20 min |
|
||||
| 3.2 | Minify JS/CSS | -50-100ms parse | — | 30 min |
|
||||
| 3.3 | Pre-compress assets | -10-30ms TTFB | — | 20 min |
|
||||
| 3.4 | Critical CSS inline | -200-400ms | — | 2 hrs |
|
||||
|
||||
**Combined estimate**: First paint **300-800ms faster**, TTI **200-500ms faster**.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Order (for implementation agent)
|
||||
|
||||
Do these in order — each step is independently testable:
|
||||
|
||||
1. **1.1** — mobile.css media query (1 line, instant win)
|
||||
2. **1.5** — Preload xterm.css (2 lines, big render-blocking fix)
|
||||
3. **1.2** — Remove duplicate `/api/status` + double handleInit
|
||||
4. **1.3** — Deduplicate `/api/settings` fetch
|
||||
5. **1.7** — Remove cache-busting from `/api/cases`
|
||||
6. **3.1** — Self-host xterm.js (removes CDN dependency entirely)
|
||||
7. **3.2** — Add esbuild minification to build
|
||||
8. **1.4** — Defer stats polling
|
||||
9. **1.6** — Fix stale version strings in COM workflow
|
||||
10. **2.1** — Loading skeleton
|
||||
11. **2.2** — Defer terminal init after first paint
|
||||
12. **2.3** — Batch init API endpoint
|
||||
13. **2.4** — Lazy modals (biggest refactor, do last)
|
||||
14. **3.3** — Pre-compress assets (nice-to-have)
|
||||
15. **3.4** — Critical CSS extraction (only if still needed after above)
|
||||
|
||||
**Verification after each step**: Use Playwright to load the page with `waitUntil: 'domcontentloaded'`, measure first paint timing, check that the UI renders correctly with 3-4s wait for async data.
|
||||
@@ -1,423 +0,0 @@
|
||||
# Performance Analysis & Optimization Opportunities
|
||||
|
||||
**Date**: 2026-03-07
|
||||
**Scope**: Full-stack performance analysis — backend PTY handling, SSE broadcasting, frontend terminal rendering, local echo overlay, DOM updates, config/scaling limits.
|
||||
**Constraint**: All recommendations preserve existing functionality including local echo, backpressure, anti-flicker pipeline, and mobile support.
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
The codebase is already well-optimized in critical paths. The multi-layer backpressure system, adaptive terminal batching, DEC 2026 sync markers, and incremental state serialization are strong. The main opportunities are in **reducing unnecessary work** (SSE filtering, DOM rebuilds, lazy terminal init) rather than algorithmic changes.
|
||||
|
||||
**Top 5 high-impact opportunities:**
|
||||
|
||||
| # | Optimization | Impact | Risk | Effort |
|
||||
|---|-------------|--------|------|--------|
|
||||
| 1 | Session-scoped SSE subscriptions | Bandwidth -60-80%, CPU -40% | Medium | Medium |
|
||||
| 2 | Lazy xterm.js for minimized subagent windows | Memory -3.5MB at 50 agents | Low | Low |
|
||||
| 3 | Targeted badge update (skip full tab rebuild) | Eliminates O(n) reflow on badge change | Low | Low |
|
||||
| 4 | Conditional SSE padding (tunnel-only, terminal-only) | Bandwidth -70% when tunneled | Low | Low |
|
||||
| 5 | Canvas renderer on mobile | GPU pressure reduction, battery savings | Low | Low |
|
||||
|
||||
---
|
||||
|
||||
## 1. SSE Broadcasting
|
||||
|
||||
### Current State
|
||||
- **92 event types** broadcast to all connected clients (max 100)
|
||||
- Single `JSON.stringify()` per event, shared across all clients (efficient)
|
||||
- **No per-client filtering** — every client receives every event regardless of which session they're viewing
|
||||
- 8KB padding appended to **every** event when tunnel is active (forces Cloudflare proxy flush)
|
||||
- Backpressure: clients marked as backpressured if `reply.raw.write()` returns false; recovery via `session:needsRefresh`
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B1: No session-scoped SSE subscriptions** (`server.ts:1986`)
|
||||
- Client viewing session A still receives all events for sessions B through T
|
||||
- With 20 active sessions, ~95% of terminal events are irrelevant to any given client
|
||||
- Cost: wasted bandwidth, CPU for JSON parsing, and event handler dispatch on client
|
||||
|
||||
**B2: Unconditional 8KB padding** (`server.ts:1977`)
|
||||
- Every event gets 8KB comment padding when tunnel is active
|
||||
- A `task:updated` event (~200 bytes payload) becomes ~8.2KB
|
||||
- High-frequency events like `session:terminal` need the padding; low-frequency events like `session:created` don't
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R1: Session-scoped SSE subscriptions** (High impact)
|
||||
- Add `?sessions=id1,id2` query param to `/api/events` SSE endpoint
|
||||
- Server filters events by session ID before broadcasting
|
||||
- Client subscribes to active session + "global" events (session lifecycle, system)
|
||||
- Re-subscribes on tab switch (or subscribe to all with client-side filter as fallback)
|
||||
- **Savings**: ~80% bandwidth reduction for single-session viewers; ~60% for multi-session dashboards
|
||||
|
||||
**R2: Tiered SSE padding** (Medium impact)
|
||||
- Only pad `session:terminal` events and SSE heartbeats (the two that need proxy flush)
|
||||
- Skip padding for low-frequency structural events (`session:created`, `task:updated`, etc.)
|
||||
- **Savings**: ~70% padding overhead reduction; terminal events already large enough to flush
|
||||
|
||||
---
|
||||
|
||||
## 2. Terminal Rendering
|
||||
|
||||
### Current State (Well-Optimized)
|
||||
- **6-layer anti-flicker pipeline**: Server batching (adaptive 16-50ms) → DEC 2026 sync wrap → single JSON serialize → client rAF batching → sync segment parser → chunked buffer loading (32KB/frame)
|
||||
- **64KB/frame write budget** with DEC 2026 sync-segment awareness (prevents 141KB single-frame freezes)
|
||||
- **3-layer backpressure**: SSE cap (128KB queued → drop + refresh), frame budget (64KB/frame), chunked restore (32KB/frame)
|
||||
- WebGL renderer enabled by default with canvas fallback on context loss
|
||||
- Typical latency: 16-32ms; worst case: ~115ms (50ms server batch + 50ms sync wait + 16ms rAF)
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B3: WebGL on mobile** (`app.js:627-637`)
|
||||
- Mobile GPUs are weaker; WebGL context loss more likely on low-end devices
|
||||
- Canvas renderer is sufficient for mobile (typically 1 session, smaller viewport)
|
||||
|
||||
**B4: Static scrollback for all sessions** (`app.js:572`)
|
||||
- Default 5000 lines scrollback for all sessions regardless of activity level
|
||||
- Heavy output sessions (build logs, test runners) accumulate large scroll buffers
|
||||
|
||||
**B5: No addon lazy loading**
|
||||
- FitAddon, Unicode11Addon, and WebGLAddon all loaded at terminal init
|
||||
- Unicode11Addon only needed for CJK content; WebGLAddon is large
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R3: Force canvas renderer on mobile** (Low risk)
|
||||
- Detect `MobileDetection.isMobile()` and skip WebGL addon loading
|
||||
- Reduces GPU memory pressure, prevents context loss crashes
|
||||
- Mobile typically has 1-2 sessions — canvas performance is more than adequate
|
||||
|
||||
**R4: Dynamic scrollback based on session activity** (Low risk)
|
||||
- Active sessions (working state): 5000 lines (current default)
|
||||
- Inactive/idle sessions: reduce to 2000 lines
|
||||
- Restore on session select (fetch from server buffer)
|
||||
- **Savings**: ~60% scrollback memory for idle sessions
|
||||
|
||||
**R5: Lazy-load Unicode11Addon** (Low risk)
|
||||
- Only load when CJK content is detected in terminal output
|
||||
- Detection: check for characters in CJK Unicode ranges during ANSI stripping (already iterating)
|
||||
- Most sessions never need it
|
||||
|
||||
---
|
||||
|
||||
## 3. DOM & Session Tab Rendering
|
||||
|
||||
### Current State
|
||||
- Session tabs use **intelligent incremental updates** with debounced 100ms rendering
|
||||
- Incremental path: only updates changed properties (classes, textContent, badges) when session list is stable
|
||||
- Full rebuild path: triggered when sessions added/removed **or badge count changes**
|
||||
- Subagent windows: per-window xterm.js instances, even when minimized
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B6: Badge count change triggers full tab rebuild** (`app.js:3207-3209`)
|
||||
- A single subagent badge increment on one tab triggers `_fullRenderSessionTabs()` — rebuilds entire sidebar HTML via `innerHTML =`
|
||||
- With 20 sessions, this is an O(n) reflow for a single badge number change
|
||||
- Badge changes are frequent during active subagent work
|
||||
|
||||
**B7: Minimized subagent windows retain xterm.js instances** (`subagent-windows.js`)
|
||||
- 50 subagent windows × ~75KB per xterm.js instance = ~3.75MB DOM memory
|
||||
- Minimized windows are invisible but their terminals remain in DOM
|
||||
- xterm.js instances continue processing resize events even when hidden
|
||||
|
||||
**B8: `backdrop-filter: blur()` on overlays** (`styles.css:2246-2247, 3098`)
|
||||
- Forces new stacking context, disables browser compositing optimizations
|
||||
- 50-100ms layout thrashing on modal open/close
|
||||
- Only 2 uses, but they're on frequently toggled overlays
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R6: Targeted badge update without full rebuild** (Low risk)
|
||||
- When badge count changes but session list is stable, update only the badge `<span>` textContent
|
||||
- Keep incremental path for badge changes; only use full rebuild for structural changes (add/remove sessions)
|
||||
- **Savings**: Eliminates O(n) reflow per badge change; reduces to O(1) targeted update
|
||||
|
||||
**R7: Lazy xterm.js initialization for subagent windows** (Medium impact)
|
||||
- Only create xterm.js Terminal instance when window is restored/maximized
|
||||
- On minimize: serialize terminal buffer, dispose Terminal instance, keep buffer in memory
|
||||
- On restore: create new Terminal, write buffer back
|
||||
- **Savings**: ~3.5MB DOM reduction at 50 minimized agents; eliminates hidden resize processing
|
||||
- **Trade-off**: ~200-500ms restore delay (buffer write), mitigated by chunked loading
|
||||
|
||||
**R8: Replace `backdrop-filter: blur()` with `background: rgba()`** (Low risk)
|
||||
- Use semi-transparent background instead of blur effect
|
||||
- Or use `will-change: transform` hint if blur is kept
|
||||
- **Savings**: Eliminates forced recomposition layer; 50-100ms faster overlay open
|
||||
|
||||
---
|
||||
|
||||
## 4. Backend PTY & State Management
|
||||
|
||||
### Current State (Excellent)
|
||||
- **BufferAccumulator**: Array-based chunking with lazy join on read — avoids O(n) string concatenation
|
||||
- **ANSI stripping**: Throttled at 150ms intervals with lazy evaluation (not per-chunk)
|
||||
- **State persistence**: 500ms debounce + incremental JSON caching per session (only dirty sessions re-serialized)
|
||||
- **Expensive parsers**: Throttled to 150ms window, accumulated data capped at 64KB
|
||||
- **Memory**: All buffers have hard limits (2MB terminal, 1MB text, 1000 messages, 64KB line buffer)
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B9: Pending clean data cap at 64KB** (`session.ts:1097-1133`)
|
||||
- Between 150ms processing windows, raw PTY data accumulates in `_pendingCleanData`
|
||||
- Capped at 64KB — excess data rolls off (old data discarded)
|
||||
- During heavy output (large build logs), this means parsers may miss content
|
||||
- Acceptable trade-off for performance, but worth documenting
|
||||
|
||||
**B10: `LRUMap.delete()` is O(n) worst case** (`utils/lru-map.ts:137-138`)
|
||||
- When deleting the newest entry, iterates all keys to find new newest
|
||||
- Rare in practice (delete is uncommon; set/get are hot paths)
|
||||
- Could matter during mass cleanup of 500 agents
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R9: Consider adaptive pending data cap** (Low priority)
|
||||
- During idle detection (critical to get right), increase cap to 128KB
|
||||
- During active working state, keep at 64KB (parsers less critical)
|
||||
- **Benefit**: More accurate idle detection during heavy output
|
||||
|
||||
**R10: Track second-newest in LRUMap** (Low priority)
|
||||
- Maintain a `_secondNewestKey` alongside `_newestKey`
|
||||
- On delete of newest, promote second-newest without iteration
|
||||
- Only matters at scale (500+ agents with frequent eviction)
|
||||
|
||||
---
|
||||
|
||||
## 5. Local Echo & Input Path
|
||||
|
||||
### Current State (Well-Designed)
|
||||
- **DOM overlay approach** — `<span>` elements in `.xterm-screen` at z-index 7, completely independent of `terminal.write()`
|
||||
- **Render caching**: `_lastRenderKey` includes text, position, column offsets — skips redundant re-renders
|
||||
- **Input flow**: Char accumulation → Enter triggers flush → 80ms delay before `\r` (ensures text reaches PTY first)
|
||||
- **Tab completion**: Baseline snapshot → detect buffer change → 300ms fallback timer
|
||||
- **CJK support**: Per-character width detection with `terminal.unicode.getStringCellWidth()` preferred, manual fallback
|
||||
- **Prompt detection**: Bottom-up line scan, O(rows) — cached position, column-lock prevents jitter
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B11: tmux send-keys latency** (~50-100ms per input)
|
||||
- Each `writeViaMux()` spawns a child process (`tmux send-keys`)
|
||||
- Text and Enter sent separately with 50ms delay between
|
||||
- For rapid typing: characters batch before Enter, so overhead is per-command not per-keystroke
|
||||
- **Acceptable trade-off** for session persistence (tmux survives server restarts)
|
||||
|
||||
**B12: 80ms delay between text flush and Enter** (`app.js:872-875`)
|
||||
- Intentional: ensures text reaches PTY before Enter, preventing Ink from processing empty input
|
||||
- Adds 80ms to perceived Enter-to-response latency
|
||||
- Could potentially be reduced with acknowledgment-based approach
|
||||
|
||||
**B13: Scroll listener on terminal viewport** (`zerolag-input-addon.ts:139`)
|
||||
- 50ms debounced re-render on scroll — acceptable but fires frequently during heavy output
|
||||
- Overlay hidden when scrolled up (correct behavior), shown when at bottom
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R11: Reduce Enter delay from 80ms to 50ms** (Low risk, test carefully)
|
||||
- The tmux `send-keys` already has 50ms internal delay
|
||||
- Combined with network latency, 80ms client-side may be excessive
|
||||
- Test with Ink-heavy sessions (Claude Code's status bar) — if text arrives before Enter at 50ms, reduce
|
||||
- **Savings**: 30ms perceived latency reduction per command
|
||||
|
||||
**R12: Batch tmux send-keys via stdin pipe** (Medium effort, high impact for rapid input)
|
||||
- Instead of spawning `tmux send-keys` per input, maintain a persistent connection
|
||||
- Use `tmux -C` (control mode) for programmatic interaction without child process spawning
|
||||
- **Savings**: Eliminate ~50-100ms process spawn overhead per input
|
||||
- **Risk**: Control mode has different semantics; needs careful testing with session persistence
|
||||
|
||||
**R13: Skip overlay re-render during heavy output scroll** (Low risk)
|
||||
- When terminal is receiving >10KB/s output, hide overlay entirely (user isn't typing during heavy output)
|
||||
- Re-show overlay after 500ms of output silence
|
||||
- **Savings**: Eliminates unnecessary DOM overlay re-renders during build logs / test output
|
||||
|
||||
---
|
||||
|
||||
## 6. Polling & File Watchers
|
||||
|
||||
### Current State
|
||||
- **SubagentWatcher**: 1s base poll, full scan throttled to every 5s, fs.watch() on known directories
|
||||
- **TranscriptWatcher**: 1 per session, fs.watch() primary with 1s poll fallback
|
||||
- **ImageWatcher**: chokidar per session with 100ms stability poll, burst limit 20/10s
|
||||
- **TeamWatcher**: chokidar primary with 30s poll fallback, LRU caches (50 teams, 200 tasks)
|
||||
- **RalphTracker**: Todo cleanup every 5 minutes
|
||||
|
||||
### Scaling Profile (20 sessions)
|
||||
| Component | Instances | Frequency | Total ops/sec |
|
||||
|-----------|-----------|-----------|---------------|
|
||||
| SubagentWatcher | 1 (global) | Full scan every 5s | 0.2/s |
|
||||
| TranscriptWatcher | 20 | 1s poll (fallback) | 20/s max |
|
||||
| ImageWatcher | 20 | 100ms poll (during writes only) | 200/s burst |
|
||||
| TeamWatcher | 1 (global) | 30s poll (fallback) | 0.03/s |
|
||||
| SSE heartbeat | 1 (global) | 15s | 0.07/s |
|
||||
| SSE dead client check | 1 (global) | 30s | 0.03/s |
|
||||
| Mux stats collection | 1 (global) | 2s | 0.5/s |
|
||||
| **Total steady-state** | | | **~21/s** |
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R14: Increase TranscriptWatcher poll interval to 2s** (Low risk)
|
||||
- Transcript changes are infrequent (new messages every few seconds at most)
|
||||
- fs.watch() is the primary mechanism; polling is fallback
|
||||
- **Savings**: Halves fallback filesystem checks (20/s → 10/s for 20 sessions)
|
||||
|
||||
**R15: Share chokidar instances for co-located session directories** (Medium effort)
|
||||
- Sessions in the same parent directory could share a single chokidar watcher with depth:3
|
||||
- Common case: multiple sessions in `~/projects/foo/` — one watcher covers all
|
||||
- **Savings**: Reduce chokidar instances from 20 to ~5-10 for typical workloads
|
||||
|
||||
---
|
||||
|
||||
## 7. Frontend Asset Delivery
|
||||
|
||||
### Current State
|
||||
- **app.js**: 12,027 lines (source) → esbuild minified → gzip/brotli compressed (~30-40KB gzipped)
|
||||
- **Static caching**: `maxAge: '1y'` via `@fastify/static`
|
||||
- **Service worker**: Push notification handler only — no asset caching
|
||||
- **No code splitting**: Single monolithic app.js bundle
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B14: No cache-busting mechanism**
|
||||
- `maxAge: '1y'` means browsers cache aggressively
|
||||
- After deployment, users need `Ctrl+Shift+R` to see updates
|
||||
- No content hash in filenames or ETags for automatic invalidation
|
||||
|
||||
**B15: Monolithic app.js**
|
||||
- All 12K lines loaded on initial page load regardless of which features are used
|
||||
- Ralph wizard, plan orchestrator UI, team management — all loaded upfront
|
||||
- Mobile loads the same bundle as desktop
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R16: Add content hash to asset filenames** (Medium impact)
|
||||
- Build step: rename `app.js` → `app.[hash].js`
|
||||
- Generate a manifest or inject hash into HTML template
|
||||
- Keep `maxAge: '1y'` — cache invalidation happens via filename change
|
||||
- **Savings**: Eliminates stale cache issues after deployment; removes need for manual hard refresh
|
||||
|
||||
**R17: Code-split app.js into core + feature modules** (High effort, medium impact)
|
||||
- Core (~4K lines): terminal, SSE, session management, tabs, input handling
|
||||
- Deferred (~8K lines): Ralph wizard, plan UI, team management, subagent windows, image viewer
|
||||
- Load deferred modules on first use via dynamic `import()` or lazy `<script>` injection
|
||||
- **Savings**: ~60% reduction in initial load size; faster time-to-interactive
|
||||
- **Risk**: Complexity increase; need to handle loading states for deferred features
|
||||
- **Note**: May not be worth the effort given the app is already gzipped to ~30-40KB
|
||||
|
||||
---
|
||||
|
||||
## 8. CSS Performance
|
||||
|
||||
### Current State
|
||||
- **styles.css**: 7,153 lines with ~45 box-shadow uses, 2 backdrop-filter uses
|
||||
- Animations: GPU-accelerated keyframes for pulsing alerts, loading spinners
|
||||
- Z-index layering: well-organized (subagent 1000, plan 1100, log 2000, image 3000, overlay 7)
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R18: Replace backdrop-filter with opaque overlay** (Low risk, covered in R8)
|
||||
|
||||
**R19: Use `contain: content` on subagent windows** (Low risk)
|
||||
- Add CSS containment to subagent window containers
|
||||
- Prevents layout changes inside windows from triggering reflow on parent
|
||||
- Especially valuable with 50 windows: changes in one window won't invalidate others
|
||||
- ```css
|
||||
.subagent-window { contain: content; }
|
||||
```
|
||||
- **Savings**: Reduces layout recalculation scope from global to per-window
|
||||
|
||||
**R20: Use `content-visibility: auto` on off-screen subagent windows** (Low risk)
|
||||
- Browser skips rendering of off-screen windows entirely
|
||||
- Combined with `contain-intrinsic-size` to prevent layout shift
|
||||
- ```css
|
||||
.subagent-window.minimized { content-visibility: hidden; }
|
||||
```
|
||||
- **Savings**: Browser skips paint/layout for minimized windows; complements R7
|
||||
|
||||
---
|
||||
|
||||
## 9. Memory & Scaling Limits
|
||||
|
||||
### Current Budget (20 sessions)
|
||||
| Component | Per Session | Total | Status |
|
||||
|-----------|-----------|-------|--------|
|
||||
| Terminal buffer | 2MB | 40MB | Hard-limited, auto-trim |
|
||||
| Text output | 1MB | 20MB | Hard-limited, auto-trim |
|
||||
| Messages | ~1MB | 20MB | Capped at 1000, trims to 800 |
|
||||
| Respawn buffer | 1MB | 20MB | Hard-limited |
|
||||
| **Buffers total** | | **100MB** | Acceptable |
|
||||
| TranscriptWatcher | ~100KB | 2MB | |
|
||||
| ImageWatcher | ~50KB | 1MB | |
|
||||
| SubagentWatcher | ~500KB | 500KB | Global |
|
||||
| Frontend terminal cache | ~256KB | 5MB | LRU, max 20 entries |
|
||||
| **Total estimated** | | **~110MB** | Comfortable |
|
||||
|
||||
### At Max Scale (50 sessions)
|
||||
- Buffers: ~250MB
|
||||
- Watchers: ~5MB
|
||||
- **Total: ~255MB** + Node.js overhead — acceptable on modern hardware
|
||||
|
||||
### Potential Leak Vectors (All Mitigated)
|
||||
- `_shortIdCache` in server — unbounded Map, but entries are tiny (string→string); grows at O(sessions created), not O(events)
|
||||
- All CleanupManager-registered resources tracked and disposed on session stop
|
||||
- `isStopped` guard prevents new timers after session cleanup
|
||||
|
||||
---
|
||||
|
||||
## 10. Implementation Priority Matrix
|
||||
|
||||
### Phase 1 — Quick Wins (1-2 hours each, low risk)
|
||||
| # | Optimization | Files to Change |
|
||||
|---|-------------|-----------------|
|
||||
| R6 | Targeted badge update | `app.js` (3207-3209) |
|
||||
| R3 | Canvas renderer on mobile | `app.js` (627-637) |
|
||||
| R8 | Replace backdrop-filter blur | `styles.css` (2246, 3098) |
|
||||
| R19 | CSS containment on subagent windows | `styles.css` |
|
||||
| R20 | `content-visibility: hidden` on minimized windows | `styles.css` |
|
||||
|
||||
### Phase 2 — Medium Effort (half-day each)
|
||||
| # | Optimization | Files to Change |
|
||||
|---|-------------|-----------------|
|
||||
| R2 | Tiered SSE padding | `server.ts` (broadcast function) |
|
||||
| R7 | Lazy xterm.js for minimized subagents | `subagent-windows.js` |
|
||||
| R11 | Reduce Enter delay to 50ms | `app.js` (872-875), test with Ink |
|
||||
| R14 | TranscriptWatcher 2s poll | `transcript-watcher.ts` |
|
||||
| R16 | Content-hash asset filenames | `build.mjs`, `server.ts` |
|
||||
|
||||
### Phase 3 — Larger Initiatives (1-2 days each)
|
||||
| # | Optimization | Files to Change |
|
||||
|---|-------------|-----------------|
|
||||
| R1 | Session-scoped SSE subscriptions | `server.ts`, `app.js` (SSE connect) |
|
||||
| R5 | Lazy Unicode11Addon loading | `app.js`, build pipeline |
|
||||
| R12 | Persistent tmux control mode | `tmux-manager.ts` |
|
||||
| R17 | Code-split app.js | `app.js`, `build.mjs`, HTML template |
|
||||
|
||||
### Not Recommended (Low ROI or High Risk)
|
||||
| # | Why Not |
|
||||
|---|---------|
|
||||
| R4 | Dynamic scrollback adds complexity; memory savings marginal vs total budget |
|
||||
| R9 | Adaptive pending data cap adds state; current 64KB cap rarely matters |
|
||||
| R10 | LRUMap.delete() O(n) is theoretical; never triggered at current scale |
|
||||
| R15 | Shared chokidar instances add directory-matching complexity for minimal gain |
|
||||
|
||||
---
|
||||
|
||||
## Appendix: Key File Locations
|
||||
|
||||
| Area | File | Key Lines |
|
||||
|------|------|-----------|
|
||||
| SSE broadcast | `src/web/server.ts` | 1961-1989 (broadcast), 1934-1959 (backpressure) |
|
||||
| Terminal batching | `src/web/server.ts` | 1994-2048 (per-session adaptive batching) |
|
||||
| Frame budget | `src/web/public/app.js` | 1370-1478 (flushPendingWrites, 64KB cap) |
|
||||
| Flicker filter | `src/web/public/app.js` | 1176-1255 (50ms sync wait, 256KB safety) |
|
||||
| Tab rendering | `src/web/public/app.js` | 3108-3357 (incremental + full rebuild) |
|
||||
| Tab switching | `src/web/public/app.js` | 3560-3760 (cache + chunked load + deferred UI) |
|
||||
| Local echo | `packages/xterm-zerolag-input/src/` | All files (overlay, prompt, CJK) |
|
||||
| Local echo integration | `src/web/public/app.js` | 640, 815-988 (input flow) |
|
||||
| Subagent windows | `src/web/public/subagent-windows.js` | Full file (window mgmt, drag, minimize) |
|
||||
| State persistence | `src/state-store.ts` | 161-250 (debounced save, incremental JSON) |
|
||||
| Buffer accumulator | `src/utils/buffer-accumulator.ts` | Full file (array chunks, lazy join) |
|
||||
| PTY handling | `src/session.ts` | 1046-1133 (data flow), 1173-1230 (parsing) |
|
||||
| Config limits | `src/config/` | 9 files (buffer, map, timing, auth, etc.) |
|
||||
| Anti-flicker docs | `docs/terminal-anti-flicker.md` | Architecture reference |
|
||||
| CSS | `src/web/public/styles.css` | 2246 (backdrop-filter), full file |
|
||||
| Build pipeline | `scripts/build.mjs` | 59-68 (minify + compress) |
|
||||
@@ -1,266 +0,0 @@
|
||||
# Codeman Performance Investigation Report
|
||||
|
||||
**Date**: 2026-02-20
|
||||
**Scope**: Why Codeman feels sluggish when multiple Claude tabs are very busy
|
||||
**Method**: 4-agent parallel analysis of server, PTY pipeline, frontend, and background systems
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
When multiple Claude sessions are actively producing heavy terminal output (e.g., building, writing files, running tests), Codeman's UI becomes sluggish. This investigation identified **14 bottlenecks** across 4 layers of the stack. The root cause is **cumulative event loop blocking** — no single operation is catastrophically slow, but dozens of small synchronous operations run on every PTY data chunk, and with N busy sessions producing chunks every few milliseconds, the event loop gets saturated.
|
||||
|
||||
The most impactful findings are ranked by severity below.
|
||||
|
||||
---
|
||||
|
||||
## Critical Findings (Event Loop Blockers)
|
||||
|
||||
### 1. PTY Data Handler Chain — O(output_volume) per session, synchronous
|
||||
**File**: `src/session.ts:986-1086`
|
||||
**Severity**: CRITICAL
|
||||
|
||||
Every chunk of PTY output from a busy Claude session runs through this synchronous chain on the Node.js event loop:
|
||||
|
||||
```
|
||||
PTY onData → ANSI strip regex → ralph-tracker → bash-tool-parser →
|
||||
token parser → CLI info parser → task description parser →
|
||||
idle/working detection → emit('terminal') → emit('output')
|
||||
```
|
||||
|
||||
**Key costs per chunk:**
|
||||
- `ANSI_ESCAPE_PATTERN_FULL` regex (line 999): Complex regex with alternation, runs on every chunk where any consumer needs clean data
|
||||
- `ralphTracker.processCleanData()` (line 1014): Splits into lines, runs regex per line, checks multi-line patterns
|
||||
- `bashToolParser.processCleanData()` (line 1020): Similar line-by-line regex processing
|
||||
- `parseTaskDescriptionsFromTerminalData()` (line 1038): Regex scan for parenthesized descriptions
|
||||
- Working/idle detection (lines 1043-1085): Multiple `includes()` checks plus `getCleanData()` calls
|
||||
|
||||
**The lazy `getCleanData()` pattern (line 997-1002)** was a good optimization — it avoids ANSI stripping when no consumer needs it. But when Ralph tracking is enabled (common during active work), `getCleanData()` is called on every chunk, negating the optimization.
|
||||
|
||||
**With 5 busy sessions** producing 50+ chunks/second each, this means 250+ synchronous processing chains per second on the event loop. Each chain involves string allocation, regex matching, and line splitting.
|
||||
|
||||
### 2. Broadcast Serialization — JSON.stringify on every flush
|
||||
**File**: `src/web/server.ts:4941-4967`
|
||||
**Severity**: CRITICAL
|
||||
|
||||
The `broadcast()` method calls `JSON.stringify(data)` synchronously for every event. Terminal data is the highest-frequency event. During `flushTerminalBatches()` (line 5030), broadcast is called once per session with pending data. With 10 busy sessions flushing every 16-50ms, that's 200-625 `JSON.stringify` calls per second on terminal data alone.
|
||||
|
||||
The terminal data payload is a string that gets double-encoded: the raw terminal string is embedded inside a JSON object `{id, data}`, then that object is JSON.stringify'd. For large chunks (up to 32KB per the `BATCH_FLUSH_THRESHOLD`), this creates significant garbage collection pressure.
|
||||
|
||||
**Additionally**, the `session:updated` broadcast includes `toLightDetailedState()` which serializes `taskTree`, `tokens`, `bufferStats`, and `respawnConfig` — this is called on many state changes, not just terminal data.
|
||||
|
||||
### 3. Single-Timer Batching — All sessions share one setTimeout
|
||||
**File**: `src/web/server.ts:5017-5027`
|
||||
**Severity**: HIGH
|
||||
|
||||
The `batchTerminalData()` method uses a **single shared timer** (`this.terminalBatchTimer`) for all sessions. When the timer fires, `flushTerminalBatches()` iterates ALL pending sessions and broadcasts each one. This means:
|
||||
|
||||
- One extremely busy session's rapid data can force the timer to fire at the minimum interval (16ms), flushing ALL sessions at that rate
|
||||
- The flush itself iterates all pending sessions synchronously
|
||||
- The `_minBatchInterval` optimization (line 5003) means the fastest session dictates the timer for everyone
|
||||
|
||||
This creates a **thundering herd** effect: all session flushes happen in a single synchronous burst rather than being staggered.
|
||||
|
||||
### 4. State Persistence Storms
|
||||
**File**: `src/web/server.ts:3879-3917`
|
||||
**Severity**: HIGH
|
||||
|
||||
`persistSessionState()` is called from **28+ locations** in server.ts. Each call sets a 100ms debounce timer per session. During heavy activity, this means:
|
||||
|
||||
- Frequent timer creation/cancellation (GC pressure)
|
||||
- The actual persist (`_persistSessionStateNow`) calls `session.toState()` which creates a new object, then `store.setSession()` which triggers `JSON.stringify` of the entire state store and `writeFileSync` to disk
|
||||
|
||||
The `StateStore` (via `state-store.ts`) debounces its own write, but the overhead is in the per-session `toState()` serialization and object creation, not just the disk write.
|
||||
|
||||
---
|
||||
|
||||
## High-Severity Findings
|
||||
|
||||
### 5. Ralph Tracker Line Processing — O(lines) per chunk
|
||||
**File**: `src/ralph-tracker.ts:1337-1375`
|
||||
**Severity**: HIGH (when Ralph tracking is enabled)
|
||||
|
||||
When enabled, `processCleanData()`:
|
||||
1. Appends to a line buffer (string concatenation)
|
||||
2. Splits on `\n` (creates array)
|
||||
3. Calls `processLine()` on each line (regex matching per line)
|
||||
4. Calls `checkMultiLinePatterns()` (additional regex on full chunk)
|
||||
5. Calls `maybeCleanupExpiredTodos()` (iterates todos Map)
|
||||
|
||||
For a busy session producing 100+ lines/second, this is significant. The line buffer can grow up to `MAX_LINE_BUFFER_SIZE` before being truncated, and the split/iterate pattern creates garbage on every chunk.
|
||||
|
||||
### 6. Subagent Watcher Polling — O(agents) every 1-10 seconds
|
||||
**File**: `src/subagent-watcher.ts:225-274`
|
||||
**Severity**: MEDIUM-HIGH
|
||||
|
||||
Three periodic operations:
|
||||
- **Poll interval** (1s): Lightweight check, but full directory scan every 5th poll (5s)
|
||||
- **Liveness check** (10s): Runs `pgrep` (child process spawn), then iterates ALL tracked agents to check if alive. With 50+ subagents (common with agent teams), this is a non-trivial burst.
|
||||
- **File watchers**: One `chokidar` watcher per tracked agent directory, plus transcript file watchers. With many agents, this means many active file watchers consuming kernel inotify resources.
|
||||
|
||||
The `getClaudePids()` call spawns a child process (`pgrep`) every 10 seconds. Under heavy load, child process spawning competes with the event loop.
|
||||
|
||||
### 7. SSE Client Iteration — O(clients) per broadcast
|
||||
**File**: `src/web/server.ts:4964-4966`
|
||||
**Severity**: MEDIUM
|
||||
|
||||
Every `broadcast()` iterates all SSE clients to send the pre-formatted message. With multiple browser tabs or mobile clients, each flush sends data to every client. The `reply.raw.write()` call goes through Node's HTTP stream, which is generally non-blocking but can cause backpressure cascades.
|
||||
|
||||
The backpressure handling (line 4916-4938) correctly skips backpressured clients, but the `once('drain')` handler sends a `session:needsRefresh` event, which the client responds to by fetching the full buffer — potentially a 2MB request — amplifying the problem.
|
||||
|
||||
### 8. Event Emitter Fan-Out in Session
|
||||
**File**: `src/session.ts:1008-1009`
|
||||
**Severity**: MEDIUM
|
||||
|
||||
Every PTY data chunk emits TWO events: `terminal` and `output`. The `terminal` event triggers `batchTerminalData()` in server.ts. The `output` event may trigger additional handlers. EventEmitter dispatch is synchronous — all listeners run before the next operation in the PTY handler continues.
|
||||
|
||||
With busy sessions, this means every chunk blocks the event loop for: PTY processing + all terminal listeners + all output listeners.
|
||||
|
||||
---
|
||||
|
||||
## Medium-Severity Findings
|
||||
|
||||
### 9. Respawn Controller Timer Accumulation
|
||||
**File**: `src/respawn-controller.ts` (various)
|
||||
**Severity**: MEDIUM
|
||||
|
||||
Each session with respawn enabled runs multiple timers:
|
||||
- Idle detection timeout
|
||||
- AI checker interval (when active)
|
||||
- Output silence detection interval
|
||||
- Token stability interval
|
||||
- Circuit breaker state timeouts
|
||||
|
||||
With 10 sessions with respawn, that's 50+ active timers. While individual timers are cheap, the cumulative effect on the event loop's timer queue is non-trivial — the libuv timer heap has O(log n) insertion but all callbacks run synchronously.
|
||||
|
||||
### 10. Team Watcher Polling
|
||||
**File**: `src/team-watcher.ts`
|
||||
**Severity**: MEDIUM (when agent teams are active)
|
||||
|
||||
Polls `~/.claude/teams/` directory every few seconds. Each poll reads config.json files and task files. With active teams, this adds filesystem reads to the event loop's I/O budget.
|
||||
|
||||
### 11. Frontend Terminal Write Batching
|
||||
**File**: `src/web/public/app.js` (batchTerminalWrite/flushPendingWrites)
|
||||
**Severity**: MEDIUM
|
||||
|
||||
The frontend batches terminal writes at `requestAnimationFrame` rate (16ms). When receiving SSE events from multiple busy sessions:
|
||||
- `batchTerminalWrite()` is called for EVERY session's data, even sessions not currently displayed
|
||||
- Terminal instances exist for all sessions (not just the active tab)
|
||||
- Each `flushPendingWrites()` calls `terminal.write()` which triggers xterm.js rendering
|
||||
|
||||
Hidden tabs still process terminal writes, consuming CPU for rendering that's never displayed.
|
||||
|
||||
### 12. Frontend Connection Line Rendering
|
||||
**File**: `src/web/public/app.js` (updateConnectionLines)
|
||||
**Severity**: LOW-MEDIUM
|
||||
|
||||
Connection lines between parent/child agent windows are recalculated on window moves, resizes, and potentially on terminal writes. With many subagent windows open, this involves DOM reads (getBoundingClientRect) that force layout recalculation.
|
||||
|
||||
### 13. Image Watcher File System Events
|
||||
**File**: `src/image-watcher.ts`
|
||||
**Severity**: LOW
|
||||
|
||||
Uses chokidar to watch for image files in session working directories. With many sessions in the same or overlapping directories, watchers may generate redundant events. The `awaitWriteFinish` and burst throttling mitigate this, but the kernel inotify resources add up.
|
||||
|
||||
### 14. ANSI Escape Regex Complexity
|
||||
**File**: `src/session.ts:999`
|
||||
**Severity**: LOW (but cumulative)
|
||||
|
||||
`ANSI_ESCAPE_PATTERN_FULL` is a complex regex with multiple alternation branches. While V8's regex engine handles this well for typical terminal data, adversarial input (deeply nested escape sequences) could cause superlinear matching time. The `FOCUS_ESCAPE_FILTER` regex runs first on every chunk.
|
||||
|
||||
---
|
||||
|
||||
## Scaling Analysis
|
||||
|
||||
| Resource | Per Session | 10 Sessions | 20 Sessions |
|
||||
|----------|-------------|-------------|-------------|
|
||||
| PTY data handlers | 1 synchronous chain | 10 chains competing for event loop | 20 chains — event loop saturation likely |
|
||||
| Broadcast calls (terminal only) | 20-60/sec | 200-600/sec | 400-1200/sec |
|
||||
| JSON.stringify (terminal) | 20-60/sec | 200-600/sec | 400-1200/sec |
|
||||
| Active timers | ~5 | ~50 | ~100 |
|
||||
| File watchers (subagents) | 2-5 | 20-50 | 40-100 |
|
||||
| SSE writes per flush | N clients | N clients x 10 sessions | N clients x 20 sessions |
|
||||
| Ralph line processing | O(lines/sec) | O(10 x lines/sec) | O(20 x lines/sec) |
|
||||
|
||||
**The critical threshold appears to be 5-8 simultaneously busy sessions**, where the cumulative PTY processing + broadcast serialization + timer callbacks start to exceed the event loop's capacity for responsive handling.
|
||||
|
||||
---
|
||||
|
||||
## Root Cause Architecture Diagram
|
||||
|
||||
```
|
||||
Busy Claude Session 1 ─┐
|
||||
Busy Claude Session 2 ─┤ ┌──────────────────────┐
|
||||
Busy Claude Session 3 ─┼───→│ Node.js Event Loop │
|
||||
Busy Claude Session 4 ─┤ │ (SINGLE THREAD) │
|
||||
Busy Claude Session 5 ─┘ │ │
|
||||
│ PTY handlers (sync) │◄── BOTTLENECK 1
|
||||
│ ANSI strip regex │
|
||||
│ Ralph tracker │
|
||||
│ Bash tool parser │
|
||||
│ Idle detection │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ EventEmitter.emit() │◄── BOTTLENECK 2
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ batchTerminalData() │
|
||||
│ (shared timer) │◄── BOTTLENECK 3
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ flushTerminalBatches() │
|
||||
│ broadcast() per session│
|
||||
│ JSON.stringify() each │◄── BOTTLENECK 4
|
||||
│ write() to N clients │
|
||||
│ │
|
||||
│ + persistSessionState │◄── BOTTLENECK 5
|
||||
│ + respawn timers │
|
||||
│ + subagent polling │
|
||||
│ + team watcher │
|
||||
└────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Recommendations (Not Implemented — For Discussion)
|
||||
|
||||
### Tier 1: Highest Impact, Lowest Risk
|
||||
1. **Disable processing for non-visible sessions**: Skip Ralph tracking, bash tool parsing, and task description parsing for sessions that no active SSE client is viewing. Only buffer terminal data.
|
||||
2. **Per-session flush staggering**: Instead of one shared timer flushing all sessions, use individual timers offset by `index * (interval/N)` to spread flushes across the batch window.
|
||||
3. **Skip hidden tab terminal writes on frontend**: Don't call `terminal.write()` for terminals not in the active tab. Lazy-load on tab switch.
|
||||
|
||||
### Tier 2: Medium Impact
|
||||
4. **Worker thread for ANSI stripping and parsing**: Move the regex-heavy ANSI strip + Ralph parsing to a worker thread pool. PTY data → worker → clean data back to main thread.
|
||||
5. **Pre-formatted SSE messages for terminal data**: Since terminal events are just `{id, data}`, build the SSE message string directly without `JSON.stringify`.
|
||||
6. **Adaptive processing based on load**: When event loop lag exceeds a threshold (measured via `setTimeout(0)` drift), reduce processing — skip Ralph, increase batch intervals, reduce subagent poll frequency.
|
||||
|
||||
### Tier 3: Longer-Term Architectural
|
||||
7. **Process-per-session or cluster mode**: Move each session's PTY handling to a separate Node.js worker or process, communicating to the main server via IPC.
|
||||
8. **Binary protocol for terminal data**: Replace JSON-encoded SSE terminal events with binary frames (e.g., MessagePack or raw binary WebSocket frames) to eliminate double-encoding.
|
||||
9. **Selective SSE subscriptions**: Clients subscribe to specific sessions instead of receiving all events. The server only broadcasts to interested clients.
|
||||
|
||||
---
|
||||
|
||||
## How to Validate
|
||||
|
||||
To confirm these findings, instrument with:
|
||||
```typescript
|
||||
// Add to event loop — measures how long synchronous work takes
|
||||
let lastCheck = Date.now();
|
||||
setInterval(() => {
|
||||
const now = Date.now();
|
||||
const lag = now - lastCheck - 100; // 100ms interval
|
||||
if (lag > 10) console.log(`[PERF] Event loop lag: ${lag}ms`);
|
||||
lastCheck = now;
|
||||
}, 100);
|
||||
```
|
||||
|
||||
And in `flushTerminalBatches()`:
|
||||
```typescript
|
||||
const start = performance.now();
|
||||
// ... existing flush logic ...
|
||||
const elapsed = performance.now() - start;
|
||||
if (elapsed > 5) console.log(`[PERF] Flush took ${elapsed.toFixed(1)}ms for ${this.terminalBatches.size} sessions`);
|
||||
```
|
||||
|
||||
This will show exactly when and how much the event loop is being blocked during heavy session activity.
|
||||
@@ -1,168 +0,0 @@
|
||||
# Performance & Responsiveness Optimization Plan
|
||||
|
||||
**Date**: 2026-02-28
|
||||
**Status**: Phases 1–4 Complete. Phase 5 optional/deferred.
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Three independent research passes analyzed the Codeman codebase for performance bottlenecks across frontend rendering, backend hot paths, and system-level resource usage. The codebase already has strong foundational optimizations (per-session adaptive batching, rAF terminal writes, DEC 2026 sync markers, backpressure handling). This plan targets the remaining high-impact opportunities.
|
||||
|
||||
**Key finding**: The biggest wins come from **skipping unnecessary work** — serializing unchanged state, processing output nobody is watching, and reducing broadcast volume.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Quick Wins — COMPLETE
|
||||
|
||||
All Phase 1 items were found to already exist in the codebase during verification:
|
||||
|
||||
| # | Item | Status | Evidence |
|
||||
|---|------|--------|----------|
|
||||
| 1.1 | Skip terminal writes for hidden tabs | Done | SSE handler filters by `activeSessionId` (app.js:4076) |
|
||||
| 1.2 | mobile.css media query | Done | `media="(max-width: 1023px)"` on link tag (index.html:13) |
|
||||
| 1.3 | Deduplicate init API calls | Done | `_initGeneration` dedup + 3s fallback timer (app.js:2901-2904) |
|
||||
| 1.4 | Remove cache-busting timestamps | Done | No `?_t=` patterns found anywhere |
|
||||
| 1.5 | JS/CSS minification + compression | Done | esbuild minify + gzip + brotli in build.mjs (lines 42-51) |
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Frontend Responsiveness — COMPLETE
|
||||
|
||||
### 2.1 Batch `getBoundingClientRect()` in connection lines — DONE
|
||||
- **Files**: `src/web/public/app.js` (`_updateConnectionLinesImmediate()`)
|
||||
- **Change**: Refactored to batch all layout reads into Phase 1 (collect all rects into a Map), then perform all SVG writes in Phase 2 using cached values. Classic read-then-write pattern prevents interleaved forced reflows.
|
||||
|
||||
### 2.2 Clean up ResizeObservers — Already implemented
|
||||
- `forceCloseSubagentWindow()` disconnects observers (app.js:12618-12620)
|
||||
- `cleanupAllFloatingWindows()` disconnects all on reconnect (app.js:12649-12653)
|
||||
- Observer refs stored on `windowData.resizeObserver` (app.js:12492)
|
||||
|
||||
### 2.3 Drag handler cleanup — Already implemented
|
||||
- `makeWindowDraggable()` returns listener refs, stored in `windowData.dragListeners`
|
||||
- `forceCloseSubagentWindow()` removes all document-level drag listeners (app.js:12622-12630)
|
||||
- Panel drags add listeners on mousedown, remove on mouseup (app.js:10253-10284)
|
||||
|
||||
### 2.4 Mobile window position cache — Skipped
|
||||
- O(n) loop over max ~20 windows; complexity of cached counter not justified
|
||||
|
||||
### 2.5 Lazy modal DOM — Skipped
|
||||
- Large effort, marginal benefit for a vanilla JS app with fast DOM construction
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Backend Hot Paths — COMPLETE
|
||||
|
||||
### 3.1 State diff broadcasts — ALREADY OPTIMIZED
|
||||
- `broadcastSessionStateDebounced()` already batches at 500ms intervals
|
||||
- `toLightDetailedState()` excludes heavy buffers (textOutput, terminalBuffer)
|
||||
- Per-session serialization is <1ms; with debouncing, only 1-3 sessions serialize per flush
|
||||
- JSON.stringify happens once per broadcast (not per client) — serialization cost is negligible
|
||||
- Full state diffs would add significant frontend complexity for marginal gain
|
||||
|
||||
### 3.2 Improve session list cache hit rate — DONE
|
||||
- **Files**: `src/web/server.ts` (`broadcast()` method)
|
||||
- **Change**: Cache now only invalidated on truly structural events (`session:created`, `session:deleted`, `session:updated`) instead of on every `session:*` and `respawn:*` event. High-frequency events like `session:working`, `session:idle`, `session:completion`, `respawn:stateChanged` no longer defeat the 1s TTL cache.
|
||||
- **Impact**: Cache hit ratio from ~0% to ~80%+ during active sessions. The debounced `session:updated` still refreshes the cache within 500ms of any state change.
|
||||
|
||||
### 3.3 Skip PTY processing — ALREADY OPTIMIZED
|
||||
- `_processExpensiveParsers()` is already throttled to every 150ms (not per-chunk)
|
||||
- Lazy ANSI stripping via `getCleanData()` closure — only computed when a consumer needs it
|
||||
- Quick pre-checks skip parsers when content is irrelevant (e.g., token parser only runs if data contains "token")
|
||||
- OpenCode sessions skip all Claude-specific parsers entirely
|
||||
- Further optimization would require visibility-aware processing, adding complexity for marginal gain
|
||||
|
||||
### 3.4 Batch subagent liveness checks — Deferred
|
||||
- `/proc/{pid}` stat calls are ~0.1ms each; even with 500 agents, total is 50ms every 10s
|
||||
- Current approach is simple and reliable; batching adds race condition risk
|
||||
- Consider only if profiling shows this as a bottleneck
|
||||
|
||||
### 3.5 Deduplicate detection update emissions — DONE
|
||||
- **Files**: `src/respawn-controller.ts` (`startDetectionUpdates()`)
|
||||
- **Change**: Detection status now only emitted when key fields (confidenceLevel, statusText, controller state) actually change. Previously emitted every 2s regardless, broadcasting identical status to all SSE clients.
|
||||
- **Impact**: For stable/idle sessions, eliminates ~100% of redundant detection broadcasts. For active sessions, reduces broadcasts to only meaningful state transitions.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: System-Level Improvements — COMPLETE
|
||||
|
||||
### 4.1 Incremental state persistence — DONE
|
||||
- **Files**: `src/state-store.ts` (`assembleStateJson()`, `setSession()`)
|
||||
- **Change**: Added `dirtySessions` Set and `cachedSessionJsons` Map. On persist, only dirty sessions are re-serialized; clean sessions reuse cached JSON fragments. `setSession()` marks sessions dirty; `assembleStateJson()` rebuilds only changed fragments.
|
||||
- **Impact**: Serialization cost reduced from O(all sessions) to O(dirty sessions). Typical steady-state: 1-2 dirty sessions instead of 50.
|
||||
|
||||
### 4.2 Replace polling with fs watchers for team watcher — DONE
|
||||
- **Files**: `src/team-watcher.ts` (`setupFsWatchers()`)
|
||||
- **Change**: Added chokidar watchers on both `~/.claude/teams/` and `~/.claude/tasks/` directories for instant event-driven detection. Lock files ignored via chokidar config. Mtime-based dedup skips unchanged files. Polling interval relaxed from 5s to 30s as a fallback.
|
||||
- **Impact**: Near-instant team detection; polling overhead eliminated for normal operation.
|
||||
|
||||
### 4.3 Consolidate subagent file watchers — DONE
|
||||
- **Files**: `src/subagent-watcher.ts` (`setupDirectoryWatcher()`)
|
||||
- **Change**: Replaced per-agent chokidar watchers with one `fs.watch()` per session subagent directory. Events are routed to the correct agent via filename. Per-file debouncing (100ms) prevents hammering on bulk discovery.
|
||||
- **Impact**: Inotify watchers reduced from potentially 500 (one per agent) to ~50 (one per session directory).
|
||||
|
||||
### 4.4 Stream transcript files instead of full reads — DONE
|
||||
- **Files**: `src/subagent-watcher.ts` (`tailFile()`, `findDescriptionInAgentFile()`, parent transcript lookup)
|
||||
- **Change**: Multiple streaming strategies implemented:
|
||||
- **Live monitoring**: Position-based `tailFile()` with `createReadStream({ start: fromPosition })` — only reads new content
|
||||
- **Parent transcript lookup**: Streams only last 16KB (`createReadStream({ start: offset })`)
|
||||
- **Description extraction**: Streams only first 8KB, exits early after 5 lines
|
||||
- **Full read**: Only for on-demand transcript review panel (with optional `limit` parameter)
|
||||
- **Impact**: File I/O for bulk agent discovery reduced from ~50MB to ~5MB.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: Long-Term Architectural (Optional) — NOT STARTED
|
||||
|
||||
These items are deferred until scaling demands justify the complexity.
|
||||
|
||||
### 5.1 Worker thread for PTY processing
|
||||
- **Files**: `src/session.ts`
|
||||
- **Problem**: ANSI stripping, Ralph tracking, and bash tool parsing all run on the main event loop. At scale (50 busy sessions), this consumes 300-500ms CPU/sec.
|
||||
- **Fix**: Offload ANSI strip + line processing to a worker thread pool. Main thread receives clean text + parsed events.
|
||||
- **Impact**: Frees event loop for I/O operations. Most impactful at 10+ concurrent busy sessions.
|
||||
|
||||
### 5.2 Per-session SSE subscriptions
|
||||
- **Files**: `src/web/server.ts`
|
||||
- **Problem**: Every SSE event is broadcast to all connected clients. A client watching session A still receives events for sessions B through Z.
|
||||
- **Fix**: Clients subscribe to specific session IDs. Server only sends events to interested clients.
|
||||
- **Impact**: Reduces SSE broadcast fan-out from N clients to ~1-2 per event. Major improvement at 100 SSE clients.
|
||||
|
||||
### 5.3 O(1) LRUMap via doubly-linked list
|
||||
- **Files**: `src/utils/lru-map.ts` (~lines 98-110)
|
||||
- **Problem**: `get()` uses delete + re-insert to refresh position — O(n) on Map iteration for delete.
|
||||
- **Fix**: Implement classic LRU with doubly-linked list + Map for O(1) get/put/evict.
|
||||
- **Impact**: Low — current sizes (max 500) make this barely measurable. Only worthwhile if LRUMap is used on hot paths.
|
||||
|
||||
---
|
||||
|
||||
## Completion Summary
|
||||
|
||||
| Phase | Scope | Status | Items |
|
||||
|-------|-------|--------|-------|
|
||||
| 1 | Quick Wins | **Complete** | 5/5 (all pre-existing) |
|
||||
| 2 | Frontend Responsiveness | **Complete** | 3/3 actionable done, 2 skipped |
|
||||
| 3 | Backend Hot Paths | **Complete** | 4/4 actionable done, 1 deferred |
|
||||
| 4 | System-Level | **Complete** | 4/4 done |
|
||||
| 5 | Long-Term Architectural | **Not started** | 0/3 — deferred until needed |
|
||||
|
||||
**Overall**: 16/16 actionable items complete. 3 optional items deferred.
|
||||
|
||||
---
|
||||
|
||||
## Measurement
|
||||
|
||||
Before starting Phase 5, establish baselines:
|
||||
|
||||
1. **Frontend**: Record Chrome DevTools Performance trace with 10 sessions open. Measure:
|
||||
- Frame rate during rapid terminal output
|
||||
- Long tasks (>50ms) count per 30s
|
||||
- Heap size after 1h session
|
||||
|
||||
2. **Backend**: Add `performance.now()` instrumentation around:
|
||||
- `flushSessionTerminalBatch()` — time per flush
|
||||
- `broadcastSessionStateDebounced()` — serialization time
|
||||
- `StateStore.save()` — persist time
|
||||
- Event loop lag via `monitorEventLoopDelay()`
|
||||
|
||||
3. **First load**: Lighthouse score on desktop and mobile (simulated 3G)
|
||||
@@ -1,74 +0,0 @@
|
||||
# Codeman Performance Optimization Plan
|
||||
|
||||
## Current State
|
||||
|
||||
The backend is **already production-grade** — SSE broadcasting, state persistence, terminal batching, buffer management, and memory patterns are all well-optimized. The biggest gains are on the **frontend delivery** side.
|
||||
|
||||
## Implemented Optimizations
|
||||
|
||||
### 1. V8 Compile Cache (10-20% faster cold start)
|
||||
|
||||
**Files:** `scripts/codeman-web.service`, `package.json`
|
||||
|
||||
Node.js re-parses and compiles all JS on every cold start. `NODE_COMPILE_CACHE` caches V8 compiled bytecode to disk, reusing it on subsequent starts.
|
||||
|
||||
- Added `Environment=NODE_COMPILE_CACHE=/home/arkon/.codeman/compile-cache` to systemd service
|
||||
- Added to `npm start` script for non-systemd usage
|
||||
- Zero code changes, immediate win on every restart
|
||||
|
||||
### 2. WebGL Addon Lazy-Loading (244KB saved on mobile, non-blocking on desktop)
|
||||
|
||||
**Files:** `src/web/public/index.html`, `src/web/public/app.js`
|
||||
|
||||
`xterm-addon-webgl.min.js` (244KB) was loaded eagerly for all users via `<script defer>`, but only used on desktop with WebGL2 support.
|
||||
|
||||
- Removed `<script defer>` from `index.html`
|
||||
- Added dynamic script loading in `app.js` — only downloads on desktop when WebGL is needed
|
||||
- Mobile users never download the file at all (244KB saved)
|
||||
- Desktop: loads in parallel with page rendering, addon initializes when ready
|
||||
- Graceful fallback: canvas renderer used if WebGL unavailable or script fails
|
||||
|
||||
### 3. Preload Hints (~50-100ms faster perceived load)
|
||||
|
||||
**Files:** `src/web/public/index.html`
|
||||
|
||||
Browser discovers `<script defer>` tags only when the parser reaches them at the bottom of `<body>`. By then, the HTML parse has blocked for hundreds of lines.
|
||||
|
||||
- Added `<link rel="preload" as="script">` in `<head>` for `vendor/xterm.min.js`, `constants.js`, `app.js`
|
||||
- Browser starts fetching critical scripts immediately during HTML parse (before reaching `<body>`)
|
||||
- Zero runtime overhead — just hints for the browser's preload scanner
|
||||
|
||||
### 4. Batch Tmux Reconciliation (N subprocess calls → 1)
|
||||
|
||||
**Files:** `src/tmux-manager.ts`
|
||||
|
||||
`reconcileSessions()` previously called `tmux has-session` + `tmux display-message` per known session, plus `tmux list-sessions` for discovery, plus `tmux display-message` per discovered session. With 20 sessions: 41+ subprocess calls.
|
||||
|
||||
- Replaced with single `tmux list-panes -a -F '#{session_name}\t#{pane_pid}'` call
|
||||
- Builds a Map from the result, then does O(1) lookups for both known and discovered sessions
|
||||
- Also replaced inner O(n) `isKnown` scan with a Set lookup
|
||||
- 20 sessions: 41 subprocess calls → 1, with faster lookups
|
||||
|
||||
### 5. Asset Hashing / Cache Busting (already implemented)
|
||||
|
||||
**Files:** `scripts/build.mjs` (pre-existing)
|
||||
|
||||
Content-hash cache busting was already implemented in the build script:
|
||||
- All app JS/CSS files get content hashes (`app.abc123.js`)
|
||||
- `index.html` rewritten to reference hashed filenames
|
||||
- Pre-compressed with gzip + Brotli
|
||||
- 1-year immutable cache works correctly — new deploys get new filenames
|
||||
|
||||
## Already Optimized (No Action Needed)
|
||||
|
||||
| Area | Why It's Fine |
|
||||
|------|---------------|
|
||||
| **SSE Broadcasting** | Single serialization per broadcast, preformatted frames, backpressure handling, session subscription filtering |
|
||||
| **State Persistence** | 500ms debounce, incremental per-session JSON caching, async atomic writes, circuit breaker on failures |
|
||||
| **Terminal Batching** | Adaptive intervals (16-50ms), per-session queues, immediate flush at 32KB, array-based accumulation |
|
||||
| **Buffer Management** | BufferAccumulator (array-push, lazy join), auto-trim at 2MB/1MB, no string concatenation in hot paths |
|
||||
| **ANSI Stripping** | Pre-compiled regex via factory functions, single-pass processing |
|
||||
| **Static File Serving** | @fastify/static with 1-year cache, pre-compressed Brotli/gzip, no-cache for HTML |
|
||||
| **Memory Management** | CleanupManager, LRUMap, StaleExpirationMap, bounded buffers, explicit listener cleanup |
|
||||
| **Import Patterns** | Pure ESM, lazy web server import, no circular deps, no dynamic imports in hot paths |
|
||||
| **Config Loading** | Small constant files, no I/O at import time, specific imports (no barrel) |
|
||||
@@ -1,788 +0,0 @@
|
||||
# Phase 4: Domain File Splitting — Implementation Plan
|
||||
|
||||
**Date**: 2026-03-01
|
||||
**Prerequisites**: Phase 1-3 complete (utils cleanup, CleanupManager/Debouncer migration, route extraction)
|
||||
**Goal**: Split 4 god files into focused modules with barrel exports for transparent migration.
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Split types.ts into types/ directory](#1-split-typests-into-types-directory)
|
||||
2. [Split ralph-tracker.ts into focused modules](#2-split-ralph-trackerts-into-focused-modules)
|
||||
3. [Split respawn-controller.ts into focused modules](#3-split-respawn-controllerts-into-focused-modules)
|
||||
4. [Split session.ts into focused modules](#4-split-sessionts-into-focused-modules)
|
||||
5. [Execution Order & Dependencies](#5-execution-order--dependencies)
|
||||
6. [Validation Checklist](#6-validation-checklist)
|
||||
|
||||
---
|
||||
|
||||
## 1. Split types.ts into types/ directory
|
||||
|
||||
**Current**: 1,443 lines, 71 exports, imported by 36 files.
|
||||
**Risk**: LOW — pure type refactor, no runtime behavior change.
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/types/
|
||||
├── index.ts (barrel re-export — transparent migration)
|
||||
├── common.ts (Disposable, BufferConfig, CleanupResourceType, CleanupRegistration)
|
||||
├── session.ts (SessionStatus, SessionMode, ClaudeMode, SessionConfig, SessionColor,
|
||||
│ SessionState, OpenCodeConfig, SessionOutput)
|
||||
├── task.ts (TaskStatus, TaskDefinition, TaskState)
|
||||
├── app-state.ts (AppState, AppConfig, GlobalStats, TokenUsageEntry, TokenStats,
|
||||
│ DEFAULT_CONFIG, createInitialState, createInitialGlobalStats)
|
||||
├── respawn.ts (RespawnConfig, PersistedRespawnConfig, CycleOutcome,
|
||||
│ RespawnCycleMetrics, RespawnAggregateMetrics, HealthStatus,
|
||||
│ RalphLoopHealthScore, TimingHistory, RespawnPreset)
|
||||
├── ralph.ts (RalphLoopStatus, RalphLoopState, RalphTodoStatus, RalphTodoPriority,
|
||||
│ RalphTodoItem, RalphTodoProgress, RalphSessionState,
|
||||
│ RalphStatusValue, RalphTestsStatus, RalphWorkType, RalphStatusBlock,
|
||||
│ CompletionConfidence, RalphTrackerState,
|
||||
│ CircuitBreakerState, CircuitBreakerReason, CircuitBreakerStatus,
|
||||
│ createInitialCircuitBreakerStatus, createInitialRalphTrackerState,
|
||||
│ createInitialRalphSessionState)
|
||||
├── api.ts (ApiErrorCode, ApiResponse, HookEventType, QuickStartResponse,
|
||||
│ CaseInfo, createErrorResponse, isError, getErrorMessage)
|
||||
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
|
||||
├── run-summary.ts (RunSummaryEventType, RunSummaryEventSeverity, RunSummaryEvent,
|
||||
│ RunSummaryStats, RunSummary, createInitialRunSummaryStats)
|
||||
├── tools.ts (ActiveBashToolStatus, ActiveBashTool, ImageDetectedEvent)
|
||||
├── teams.ts (TeamConfig, TeamMember, TeamTask, InboxMessage, PaneInfo)
|
||||
├── push.ts (PushSubscriptionRecord, VapidKeys)
|
||||
└── plan.ts (PlanTaskStatus, TddPhase, PlanItem re-export, NiceConfig,
|
||||
DEFAULT_NICE_CONFIG, ProcessStats)
|
||||
```
|
||||
|
||||
### Steps
|
||||
|
||||
1. **Create `src/types/` directory** and each domain file above.
|
||||
|
||||
2. **Move types** from `src/types.ts` into their domain files. Preserve all JSDoc comments. Each file should import from siblings as needed (e.g., `ralph.ts` imports `CircuitBreakerState` within itself — no cross-file deps needed since they're in the same file).
|
||||
|
||||
3. **Create barrel `src/types/index.ts`** that re-exports everything:
|
||||
```typescript
|
||||
export * from './common.js';
|
||||
export * from './session.js';
|
||||
export * from './task.js';
|
||||
export * from './app-state.js';
|
||||
export * from './respawn.js';
|
||||
export * from './ralph.js';
|
||||
export * from './api.js';
|
||||
export * from './lifecycle.js';
|
||||
export * from './run-summary.js';
|
||||
export * from './tools.js';
|
||||
export * from './teams.js';
|
||||
export * from './push.js';
|
||||
export * from './plan.js';
|
||||
```
|
||||
|
||||
4. **Delete old `src/types.ts`** and replace with a single-line re-export barrel:
|
||||
```typescript
|
||||
export * from './types/index.js';
|
||||
```
|
||||
This ensures `import { ... } from './types.js'` continues to work everywhere — zero changes to 36 import sites.
|
||||
|
||||
5. **Verify**: `tsc --noEmit` and `npm run lint` must pass. No runtime changes.
|
||||
|
||||
### Internal Dependencies Between Domain Files
|
||||
|
||||
Some types reference others across domains. Handle with imports:
|
||||
|
||||
| File | Imports From |
|
||||
|------|-------------|
|
||||
| `app-state.ts` | `session.ts` (SessionState), `task.ts` (TaskState), `ralph.ts` (RalphLoopState, RalphSessionState) |
|
||||
| `respawn.ts` | None (self-contained) |
|
||||
| `ralph.ts` | None (self-contained) |
|
||||
| `run-summary.ts` | None (self-contained) |
|
||||
| `api.ts` | None (self-contained) |
|
||||
| `session.ts` | `respawn.ts` (RespawnConfig), `ralph.ts` (RalphTrackerState, RalphTodoItem, CircuitBreakerStatus, RalphSessionState, RunSummaryEvent) |
|
||||
|
||||
Wait — `SessionState` references `RespawnConfig`, `RalphTrackerState`, `CircuitBreakerStatus`, and `RunSummaryEvent`. This creates imports from `session.ts` → `respawn.ts`, `ralph.ts`, `run-summary.ts`. This is fine (one-way deps, no cycles).
|
||||
|
||||
---
|
||||
|
||||
## 2. Split ralph-tracker.ts into focused modules
|
||||
|
||||
**Current**: 3,868 lines, single `RalphTracker` class with 5 responsibilities.
|
||||
**Risk**: MEDIUM — class has shared mutable state, but extractable modules are well-isolated.
|
||||
|
||||
### Coupling Analysis Summary
|
||||
|
||||
| Module | Coupling | Extractability |
|
||||
|--------|----------|----------------|
|
||||
| Plan task tracking | LOW | HIGH — only reads `cycleCount` |
|
||||
| Fix-plan file watching | LOW | HIGH — callback-based todo replacement |
|
||||
| Iteration stall detection | LOW | HIGH — notification-based |
|
||||
| RALPH_STATUS block parsing + circuit breaker | MEDIUM | MEDIUM — callback for circuit breaker updates |
|
||||
| Todo parsing, loop detection, completion | HIGH | LOW — deeply entangled shared state |
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/
|
||||
├── ralph-tracker.ts (~1,800 LOC — core: output parsing, loop state,
|
||||
│ todo management, completion detection)
|
||||
├── ralph-plan-tracker.ts (~600 LOC — plan tasks, checkpoints, history, rollback)
|
||||
├── ralph-status-parser.ts (~300 LOC — RALPH_STATUS block parsing, circuit breaker)
|
||||
├── ralph-fix-plan-watcher.ts (~150 LOC — @fix_plan.md file watching)
|
||||
└── ralph-stall-detector.ts (~80 LOC — iteration stall detection)
|
||||
```
|
||||
|
||||
### Step 2a: Extract `RalphPlanTracker` (~600 LOC)
|
||||
|
||||
**Why first**: Lowest coupling. Only dependency is `cycleCount` for checkpoint detection.
|
||||
|
||||
**Extract these from `RalphTracker`**:
|
||||
|
||||
Types to export:
|
||||
- `EnhancedPlanTask` (interface, currently lines 56-87)
|
||||
- `CheckpointReview` (interface, currently lines 90-139)
|
||||
|
||||
Properties to move:
|
||||
- `_planVersion: number`
|
||||
- `_planHistory: Array<{version, timestamp, tasks, summary}>`
|
||||
- `_planTasks: Map<string, EnhancedPlanTask>`
|
||||
- `_checkpointIterations: number[]`
|
||||
- `_lastCheckpointIteration: number`
|
||||
|
||||
Methods to move:
|
||||
- `initializePlanTasks(items)`
|
||||
- `updatePlanTask(taskId, update)`
|
||||
- `addPlanTask(params)`
|
||||
- `getPlanTasks()`
|
||||
- `generateCheckpointReview()`
|
||||
- `getPlanHistory()`
|
||||
- `rollbackToVersion(version)`
|
||||
- `isCheckpointDue()`
|
||||
- `planVersion` getter
|
||||
- `_savePlanToHistory()` (private)
|
||||
- `_unblockDependentTasks()` (private)
|
||||
- `_checkForCheckpoint()` (private)
|
||||
|
||||
Events emitted (define in new class):
|
||||
- `planInitialized`
|
||||
- `planTaskUpdate`
|
||||
- `taskBlocked`
|
||||
- `taskUnblocked`
|
||||
- `planCheckpoint`
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphPlanTracker extends EventEmitter {
|
||||
constructor() { ... }
|
||||
|
||||
// Parent calls this when iteration changes (for checkpoint detection)
|
||||
notifyCycleCount(cycleCount: number): void { ... }
|
||||
|
||||
// Full public API moves here unchanged
|
||||
initializePlanTasks(items: PlanItem[]): void { ... }
|
||||
updatePlanTask(taskId: string, update: { ... }): { ... } | null { ... }
|
||||
// ...etc
|
||||
}
|
||||
```
|
||||
|
||||
**In `RalphTracker`**: Replace plan methods with delegation:
|
||||
```typescript
|
||||
readonly planTracker = new RalphPlanTracker();
|
||||
|
||||
// Forward plan events
|
||||
this.planTracker.on('planInitialized', (...args) => this.emit('planInitialized', ...args));
|
||||
// ...etc
|
||||
|
||||
// In detectLoopStatus(), when cycleCount changes:
|
||||
this.planTracker.notifyCycleCount(this._loopState.cycleCount);
|
||||
```
|
||||
|
||||
### Step 2b: Extract `RalphFixPlanWatcher` (~150 LOC)
|
||||
|
||||
**Extract these**:
|
||||
|
||||
Properties:
|
||||
- `_workingDir: string | null`
|
||||
- `_fixPlanPath: string | null`
|
||||
- `_fixPlanWatcher: FSWatcher | null`
|
||||
- `_fixPlanWatcherErrorHandler`
|
||||
- `_fixPlanReloadDeb`
|
||||
|
||||
Methods:
|
||||
- `setWorkingDir(workingDir)`
|
||||
- `loadFixPlanFromDisk()`
|
||||
- `startWatchingFixPlan()`
|
||||
- `stopWatchingFixPlan()`
|
||||
- `handleFixPlanChange()`
|
||||
- `isFileAuthoritative` getter
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphFixPlanWatcher extends EventEmitter {
|
||||
get isFileAuthoritative(): boolean { ... }
|
||||
|
||||
setWorkingDir(workingDir: string): void { ... }
|
||||
stop(): void { ... }
|
||||
}
|
||||
|
||||
// Events:
|
||||
// 'todosLoaded' → (todos: Array<{id, content, status, priority}>) — parent replaces _todos
|
||||
```
|
||||
|
||||
**In `RalphTracker`**:
|
||||
```typescript
|
||||
readonly fixPlanWatcher = new RalphFixPlanWatcher();
|
||||
|
||||
constructor() {
|
||||
this.fixPlanWatcher.on('todosLoaded', (items) => {
|
||||
// Replace _todos with file-based items
|
||||
this._todos.clear();
|
||||
for (const item of items) {
|
||||
this.addOrUpdateTodo(item.id, item.content, item.status, item.priority);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
// Delegate isFileAuthoritative
|
||||
get isFileAuthoritative(): boolean {
|
||||
return this.fixPlanWatcher.isFileAuthoritative;
|
||||
}
|
||||
```
|
||||
|
||||
### Step 2c: Extract `RalphStallDetector` (~80 LOC)
|
||||
|
||||
**Extract these**:
|
||||
|
||||
Properties:
|
||||
- `_lastIterationChangeTime`
|
||||
- `_lastObservedIteration`
|
||||
- `_iterationStallTimerId`
|
||||
- `_iterationStallWarningMs`
|
||||
- `_iterationStallCriticalMs`
|
||||
- `_iterationStallWarned`
|
||||
|
||||
Methods:
|
||||
- `startIterationStallDetection()`
|
||||
- `stopIterationStallDetection()`
|
||||
- `checkIterationStall()`
|
||||
- `getIterationStallMetrics()`
|
||||
- `configureIterationStallThresholds(warningMs, criticalMs)`
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphStallDetector extends EventEmitter {
|
||||
constructor(private cleanup: CleanupManager) { ... }
|
||||
|
||||
start(): void { ... }
|
||||
stop(): void { ... }
|
||||
|
||||
// Parent calls when iteration changes
|
||||
notifyIterationChanged(iteration: number): void {
|
||||
this._lastIterationChangeTime = Date.now();
|
||||
this._lastObservedIteration = iteration;
|
||||
this._iterationStallWarned = false;
|
||||
}
|
||||
|
||||
// Parent calls to check if loop is active
|
||||
setLoopActive(active: boolean): void { ... }
|
||||
|
||||
getIterationStallMetrics(): { ... } { ... }
|
||||
}
|
||||
|
||||
// Events: 'iterationStallWarning', 'iterationStallCritical'
|
||||
```
|
||||
|
||||
### Step 2d: Extract `RalphStatusParser` (~300 LOC)
|
||||
|
||||
**Extract these**:
|
||||
|
||||
Properties:
|
||||
- `_circuitBreaker: CircuitBreakerStatus`
|
||||
- `_statusBlockBuffer: string[]`
|
||||
- `_inStatusBlock: boolean`
|
||||
- `_lastStatusBlock: RalphStatusBlock | null`
|
||||
- `_completionIndicators: number`
|
||||
- `_exitGateMet: boolean`
|
||||
- `_totalFilesModified: number`
|
||||
- `_totalTasksCompleted: number`
|
||||
|
||||
Methods:
|
||||
- `processStatusBlockLine(line)`
|
||||
- `parseStatusBlock(lines)`
|
||||
- `detectCompletionIndicators(line)`
|
||||
- `updateCircuitBreaker(hasProgress, testsStatus, status)`
|
||||
- `resetCircuitBreaker()`
|
||||
- `circuitBreakerStatus` getter
|
||||
- `lastStatusBlock` getter
|
||||
- `cumulativeStats` getter
|
||||
- `exitGateMet` getter
|
||||
|
||||
Regex patterns to move:
|
||||
- `RALPH_STATUS_START_PATTERN` through `RALPH_RECOMMENDATION_PATTERN`
|
||||
- `COMPLETION_INDICATOR_PATTERNS`
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphStatusParser extends EventEmitter {
|
||||
processLine(line: string): void { ... } // calls processStatusBlockLine + detectCompletionIndicators
|
||||
|
||||
get circuitBreakerStatus(): CircuitBreakerStatus { ... }
|
||||
get lastStatusBlock(): RalphStatusBlock | null { ... }
|
||||
get exitGateMet(): boolean { ... }
|
||||
get cumulativeStats(): { ... } { ... }
|
||||
|
||||
resetCircuitBreaker(): void { ... }
|
||||
reset(): void { ... }
|
||||
}
|
||||
|
||||
// Events: 'statusBlockDetected', 'circuitBreakerUpdate', 'exitGateMet'
|
||||
```
|
||||
|
||||
**In `RalphTracker.processLine()`**:
|
||||
```typescript
|
||||
// Replace inline status block handling with delegation
|
||||
this.statusParser.processLine(line);
|
||||
```
|
||||
|
||||
### Step 2e: Keep in `ralph-tracker.ts` (~1,800 LOC)
|
||||
|
||||
The core remains tightly coupled and stays together:
|
||||
- Output parsing pipeline (`processTerminalData`, `processCleanData`, `processLine`)
|
||||
- Loop state management (`_loopState`, `detectLoopStatus`, `enable/disable/startLoop/stopLoop`)
|
||||
- Todo management (`_todos`, `detectTodoItems`, `addOrUpdateTodo`, `updateTodoStatus`, `getTodoStats`)
|
||||
- Completion detection (`detectCompletionPhrase`, `handleCompletionPhrase`, `calculateCompletionConfidence`)
|
||||
- All-tasks-complete detection (`detectAllTasksComplete`)
|
||||
- Auto-enable logic (`shouldAutoEnable`)
|
||||
- Lifecycle (`reset`, `fullReset`, `clear`, `restoreState`, `destroy`)
|
||||
- Event debouncing and buffering
|
||||
|
||||
The class coordinates the extracted modules via composition:
|
||||
```typescript
|
||||
export class RalphTracker extends EventEmitter {
|
||||
readonly planTracker = new RalphPlanTracker();
|
||||
readonly fixPlanWatcher = new RalphFixPlanWatcher();
|
||||
readonly stallDetector: RalphStallDetector;
|
||||
readonly statusParser = new RalphStatusParser();
|
||||
|
||||
constructor() {
|
||||
super();
|
||||
this.stallDetector = new RalphStallDetector(this.cleanup);
|
||||
this._wireSubModuleEvents();
|
||||
}
|
||||
|
||||
private _wireSubModuleEvents(): void {
|
||||
// Forward all sub-module events through RalphTracker
|
||||
// so external consumers don't need to know about the split
|
||||
for (const event of ['planInitialized', 'planTaskUpdate', ...]) {
|
||||
this.planTracker.on(event, (...args) => this.emit(event, ...args));
|
||||
}
|
||||
// ...same for statusParser, stallDetector, fixPlanWatcher
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Migration Safety
|
||||
|
||||
- All events continue to be emitted from `RalphTracker` (forwarded from sub-modules)
|
||||
- All public methods stay on `RalphTracker` (delegated to sub-modules)
|
||||
- External consumers (`session.ts`, `case-routes.ts`) see zero API changes
|
||||
- New sub-modules are exposed as `readonly` properties for direct access where needed
|
||||
|
||||
---
|
||||
|
||||
## 3. Split respawn-controller.ts into focused modules
|
||||
|
||||
**Current**: 3,611 lines, single `RespawnController` class with 6 responsibilities.
|
||||
**Risk**: MEDIUM — health scoring and metrics are cleanly decoupled; detection is tightly coupled.
|
||||
|
||||
### Coupling Analysis Summary
|
||||
|
||||
| Module | Coupling | Extractability |
|
||||
|--------|----------|----------------|
|
||||
| Health scoring | NONE | HIGH — pure calculations from metrics |
|
||||
| Cycle metrics | LOW | HIGH — standalone tracking |
|
||||
| Adaptive timing | LOW | HIGH — standalone timing adjustments |
|
||||
| Stuck-state detection | LOW | MEDIUM — needs state + config refs |
|
||||
| Pattern detection utilities | NONE | HIGH — pure functions |
|
||||
| State machine + idle detection + AI checkers | HIGH | LOW — deeply entangled |
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/
|
||||
├── respawn-controller.ts (~2,200 LOC — state machine, idle detection,
|
||||
│ AI checkers, terminal handling, hook signals,
|
||||
│ auto-accept, step execution)
|
||||
├── respawn-health.ts (~250 LOC — health scoring + recommendations)
|
||||
├── respawn-metrics.ts (~200 LOC — cycle metrics + aggregate stats)
|
||||
├── respawn-adaptive-timing.ts (~100 LOC — adaptive timing with percentile calc)
|
||||
└── respawn-patterns.ts (~50 LOC — terminal pattern detection utilities)
|
||||
```
|
||||
|
||||
### Step 3a: Extract `RespawnPatterns` (~50 LOC)
|
||||
|
||||
**Pure utility functions, zero coupling**.
|
||||
|
||||
Move:
|
||||
- `isCompletionMessage(data): boolean`
|
||||
- `hasWorkingPattern(data, window): boolean`
|
||||
- `extractTokenCount(data): number | null`
|
||||
- `PROMPT_PATTERNS` array
|
||||
- `WORKING_PATTERNS` array
|
||||
|
||||
```typescript
|
||||
// src/respawn-patterns.ts
|
||||
import { TOKEN_PATTERN, SPINNER_PATTERN } from './utils/index.js';
|
||||
|
||||
export const PROMPT_PATTERNS = ['❯', '>', '$', '%', '#'];
|
||||
|
||||
export const WORKING_PATTERNS = [/* 70+ patterns */];
|
||||
|
||||
export function isCompletionMessage(data: string): boolean { ... }
|
||||
export function hasWorkingPattern(data: string, window: string): boolean { ... }
|
||||
export function extractTokenCount(data: string): number | null { ... }
|
||||
```
|
||||
|
||||
**In `RespawnController`**: Import and call:
|
||||
```typescript
|
||||
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
|
||||
```
|
||||
|
||||
### Step 3b: Extract `RespawnAdaptiveTiming` (~100 LOC)
|
||||
|
||||
**Self-contained timing controller**.
|
||||
|
||||
Move properties:
|
||||
- `timingHistory: TimingHistory`
|
||||
|
||||
Move methods:
|
||||
- `recordTimingData(idleDetectionMs, cycleDurationMs)`
|
||||
- `updateAdaptiveTiming()`
|
||||
- `getTimingHistory()`
|
||||
- `getAdaptiveCompletionConfirmMs()`
|
||||
|
||||
```typescript
|
||||
export class RespawnAdaptiveTiming {
|
||||
private timingHistory: TimingHistory;
|
||||
|
||||
constructor(private config: { adaptiveMinConfirmMs: number; adaptiveMaxConfirmMs: number }) {
|
||||
this.timingHistory = { recentIdleDetectionMs: [], recentCycleDurationMs: [], ... };
|
||||
}
|
||||
|
||||
recordTimingData(idleDetectionMs: number, cycleDurationMs: number): void { ... }
|
||||
getAdaptiveCompletionConfirmMs(): number { ... }
|
||||
getTimingHistory(): TimingHistory { ... }
|
||||
reset(): void { ... }
|
||||
}
|
||||
```
|
||||
|
||||
### Step 3c: Extract `RespawnCycleMetrics` (~200 LOC)
|
||||
|
||||
**Standalone metrics tracker**.
|
||||
|
||||
Move properties:
|
||||
- `currentCycleMetrics`
|
||||
- `recentCycleMetrics[]`
|
||||
- `aggregateMetrics`
|
||||
- `MAX_CYCLE_METRICS_IN_MEMORY`
|
||||
|
||||
Move methods:
|
||||
- `startCycleMetrics(idleReason)`
|
||||
- `recordCycleStep(step)`
|
||||
- `completeCycleMetrics(outcome, errorMessage?)`
|
||||
- `updateAggregateMetrics(metrics)`
|
||||
- `getAggregateMetrics()`
|
||||
- `getRecentCycleMetrics(limit?)`
|
||||
|
||||
```typescript
|
||||
export class RespawnCycleMetricsTracker {
|
||||
private currentCycleMetrics: Partial<RespawnCycleMetrics> | null = null;
|
||||
private recentCycleMetrics: RespawnCycleMetrics[] = [];
|
||||
private aggregateMetrics: RespawnAggregateMetrics;
|
||||
|
||||
startCycle(sessionId: string, cycleNumber: number, idleReason: string): void { ... }
|
||||
recordStep(step: string): void { ... }
|
||||
completeCycle(outcome: CycleOutcome, errorMessage?: string): RespawnCycleMetrics | null { ... }
|
||||
getAggregate(): RespawnAggregateMetrics { ... }
|
||||
getRecent(limit?: number): RespawnCycleMetrics[] { ... }
|
||||
reset(): void { ... }
|
||||
}
|
||||
```
|
||||
|
||||
**Callback**: `completeCycle()` returns the completed metrics so the controller can pass them to `adaptiveTiming.recordTimingData()`.
|
||||
|
||||
### Step 3d: Extract `RespawnHealthCalculator` (~250 LOC)
|
||||
|
||||
**Pure calculation — no state of its own**.
|
||||
|
||||
Move methods:
|
||||
- `calculateHealthScore()`
|
||||
- `calculateCycleSuccessScore()`
|
||||
- `calculateCircuitBreakerScore()`
|
||||
- `calculateIterationProgressScore()`
|
||||
- `calculateAiCheckerScore()`
|
||||
- `calculateStuckRecoveryScore()`
|
||||
- `generateHealthRecommendations(components)`
|
||||
- `generateHealthSummary(score, status, components)`
|
||||
- `shouldSkipClear()` (belongs here since it's a pure calculation on token/config)
|
||||
|
||||
```typescript
|
||||
export interface HealthInputs {
|
||||
aggregateMetrics: RespawnAggregateMetrics;
|
||||
circuitBreakerStatus: CircuitBreakerStatus;
|
||||
iterationStallMetrics: { stallDurationMs: number; warningMs: number; criticalMs: number } | null;
|
||||
aiCheckerState: { disabled: boolean; inCooldown: boolean; hasErrors: boolean };
|
||||
stuckRecoveryCount: number;
|
||||
maxStuckRecoveries: number;
|
||||
}
|
||||
|
||||
export function calculateHealthScore(inputs: HealthInputs): RalphLoopHealthScore { ... }
|
||||
|
||||
export function shouldSkipClear(
|
||||
lastTokenCount: number,
|
||||
skipClearThresholdPercent: number,
|
||||
maxContextTokens: number
|
||||
): boolean { ... }
|
||||
```
|
||||
|
||||
**Made as pure functions** (not a class) since they hold no state.
|
||||
|
||||
### Step 3e: Keep in `respawn-controller.ts` (~2,200 LOC)
|
||||
|
||||
The core state machine, idle detection, and AI checker integration stays:
|
||||
- State machine transitions (`setState`, `start`, `stop`, `pause`, `resume`)
|
||||
- Terminal data handling (`handleTerminalData`)
|
||||
- All 5 idle detection layers + hook signals
|
||||
- AI checker integration (`tryStartAiCheck`, `startAiCheck`, `startPlanCheck`)
|
||||
- Auto-accept logic
|
||||
- Step execution (`sendUpdateDocs`, `sendClear`, `sendInit`, `sendKickstart`)
|
||||
- Timer management (`startTrackedTimer`, `cancelTrackedTimer`)
|
||||
- Stuck-state detection and recovery
|
||||
- Action logging
|
||||
|
||||
The class composes extracted modules:
|
||||
```typescript
|
||||
import { RespawnAdaptiveTiming } from './respawn-adaptive-timing.js';
|
||||
import { RespawnCycleMetricsTracker } from './respawn-metrics.js';
|
||||
import { calculateHealthScore, shouldSkipClear } from './respawn-health.js';
|
||||
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
|
||||
|
||||
export class RespawnController extends EventEmitter {
|
||||
private adaptiveTiming: RespawnAdaptiveTiming;
|
||||
private cycleMetrics: RespawnCycleMetricsTracker;
|
||||
|
||||
calculateHealthScore(): RalphLoopHealthScore {
|
||||
return calculateHealthScore({
|
||||
aggregateMetrics: this.cycleMetrics.getAggregate(),
|
||||
circuitBreakerStatus: this.session.ralphTracker.circuitBreakerStatus,
|
||||
iterationStallMetrics: this.session.ralphTracker.getIterationStallMetrics(),
|
||||
aiCheckerState: { ... },
|
||||
stuckRecoveryCount: this.stuckRecoveryCount,
|
||||
maxStuckRecoveries: this.config.maxStuckRecoveries ?? 3,
|
||||
});
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Split session.ts into focused modules
|
||||
|
||||
**Current**: 2,418 lines, single `Session` class.
|
||||
**Risk**: LOW-MEDIUM — extractable pieces are utility-like with clear boundaries.
|
||||
|
||||
### Coupling Analysis Summary
|
||||
|
||||
| Module | Coupling | Extractability |
|
||||
|--------|----------|----------------|
|
||||
| CLI arg builder | NONE | HIGH — pure functions used at spawn time |
|
||||
| Auto-compact/clear | LOW | HIGH — self-contained automation with config |
|
||||
| Token tracking | LOW | MEDIUM — reads PTY output, writes state |
|
||||
| Task description cache | LOW | HIGH — separate LRU cache |
|
||||
| PTY + mux lifecycle | HIGH | KEEP — core of the class |
|
||||
| Tracker integration | HIGH | KEEP — event forwarding plumbing |
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/
|
||||
├── session.ts (~1,600 LOC — PTY lifecycle, terminal I/O,
|
||||
│ tracker integration, output processing,
|
||||
│ token tracking, state management)
|
||||
├── session-cli-builder.ts (~250 LOC — Claude/OpenCode CLI arg construction)
|
||||
├── session-auto-ops.ts (~300 LOC — auto-compact, auto-clear automation)
|
||||
└── session-task-cache.ts (~100 LOC — task description LRU cache)
|
||||
```
|
||||
|
||||
### Step 4a: Extract `SessionCliBuilder` (~250 LOC)
|
||||
|
||||
**Pure functions — zero coupling to Session instance**.
|
||||
|
||||
Move:
|
||||
- `buildClaudeArgs()` logic (currently inlined in `startInteractive` and `runPrompt`)
|
||||
- `buildOpenCodeArgs()` logic
|
||||
- Model mapping constants
|
||||
- Claude mode to flag mapping
|
||||
- Environment variable construction
|
||||
|
||||
```typescript
|
||||
// src/session-cli-builder.ts
|
||||
export interface CliBuilderConfig {
|
||||
claudeMode: ClaudeMode;
|
||||
model?: string;
|
||||
workingDir: string;
|
||||
sessionId: string;
|
||||
niceConfig?: NiceConfig;
|
||||
isOpenCode?: boolean;
|
||||
openCodeConfig?: OpenCodeConfig;
|
||||
}
|
||||
|
||||
export function buildInteractiveArgs(config: CliBuilderConfig): string[] { ... }
|
||||
export function buildPromptArgs(config: CliBuilderConfig, prompt: string): string[] { ... }
|
||||
export function buildShellArgs(shell?: string): string[] { ... }
|
||||
export function buildClaudeEnv(config: CliBuilderConfig): Record<string, string> { ... }
|
||||
```
|
||||
|
||||
### Step 4b: Extract `SessionAutoOps` (~300 LOC)
|
||||
|
||||
**Self-contained automation with config-based thresholds**.
|
||||
|
||||
Move properties:
|
||||
- `_autoCompactThreshold`
|
||||
- `_autoClearThreshold`
|
||||
- `_isAutoCompacting`
|
||||
- `_isAutoClearing`
|
||||
- `_autoCompactCount`
|
||||
- `_autoClearCount`
|
||||
- `_lastAutoCompactTime`
|
||||
- `_lastAutoClearTime`
|
||||
|
||||
Move methods:
|
||||
- `checkAutoCompact(tokenCount)`
|
||||
- `performAutoCompact()`
|
||||
- `checkAutoClear(tokenCount)`
|
||||
- `performAutoClear()`
|
||||
- Auto-compact/clear threshold configuration
|
||||
|
||||
```typescript
|
||||
export class SessionAutoOps extends EventEmitter {
|
||||
constructor(
|
||||
private writeCommand: (command: string) => Promise<void>,
|
||||
private getTokenCount: () => number,
|
||||
config: { compactThreshold: number; clearThreshold: number }
|
||||
) { ... }
|
||||
|
||||
/** Called after token count updates. Checks thresholds and triggers if needed. */
|
||||
checkThresholds(tokenCount: number): void { ... }
|
||||
|
||||
updateConfig(config: { compactThreshold?: number; clearThreshold?: number }): void { ... }
|
||||
getStats(): { autoCompactCount: number; autoClearCount: number; ... } { ... }
|
||||
}
|
||||
|
||||
// Events: 'autoCompact', 'autoClear'
|
||||
```
|
||||
|
||||
**In `Session`**: Compose and wire:
|
||||
```typescript
|
||||
private autoOps = new SessionAutoOps(
|
||||
(cmd) => this.writeViaMux(cmd),
|
||||
() => this._state.tokenCount,
|
||||
{ compactThreshold: 110_000, clearThreshold: 140_000 }
|
||||
);
|
||||
```
|
||||
|
||||
### Step 4c: Extract `SessionTaskCache` (~100 LOC)
|
||||
|
||||
**Isolated LRU cache for task descriptions**.
|
||||
|
||||
Move:
|
||||
- `_taskDescriptionCache: LRUMap<number, { description: string; timestamp: number }>`
|
||||
- `_taskDescriptionMaxAge`
|
||||
- `findTaskDescriptionNear(lineNumber)`
|
||||
- `cacheTaskDescription(lineNumber, description)`
|
||||
|
||||
```typescript
|
||||
export class SessionTaskCache {
|
||||
private cache: LRUMap<number, { description: string; timestamp: number }>;
|
||||
private maxAgeMs: number;
|
||||
|
||||
constructor(maxSize: number = 50, maxAgeMs: number = 30_000) { ... }
|
||||
|
||||
find(lineNumber: number, searchRadius: number = 50): string | null { ... }
|
||||
add(lineNumber: number, description: string): void { ... }
|
||||
clear(): void { ... }
|
||||
}
|
||||
```
|
||||
|
||||
### Step 4d: Keep in `session.ts` (~1,600 LOC)
|
||||
|
||||
The core stays together:
|
||||
- PTY process management (`spawn`, `kill`, `resize`, `writeViaMux`)
|
||||
- Data streaming pipeline (PTY → buffer → ANSI strip → JSON parse → events)
|
||||
- Tracker initialization and event forwarding (RalphTracker, BashToolParser, TaskTracker)
|
||||
- Output processing (message extraction, completion detection)
|
||||
- Token tracking (status line parsing)
|
||||
- State management (`toState()`, `updateState()`)
|
||||
- Session lifecycle (`startInteractive`, `startShell`, `runPrompt`)
|
||||
- CLI info detection (version, model, account)
|
||||
|
||||
---
|
||||
|
||||
## 5. Execution Order & Dependencies
|
||||
|
||||
Execute in this order to minimize risk. Each step is independently deployable.
|
||||
|
||||
```
|
||||
Step 1: types.ts split
|
||||
↓ (no runtime change, just file reorganization)
|
||||
Step 2a: RalphPlanTracker extraction
|
||||
↓ (independent of types split)
|
||||
Step 2b: RalphFixPlanWatcher extraction
|
||||
Step 2c: RalphStallDetector extraction
|
||||
Step 2d: RalphStatusParser extraction
|
||||
↓ (ralph-tracker.ts now ~1,800 LOC)
|
||||
Step 3a: RespawnPatterns extraction
|
||||
Step 3b: RespawnAdaptiveTiming extraction
|
||||
Step 3c: RespawnCycleMetrics extraction
|
||||
Step 3d: RespawnHealthCalculator extraction
|
||||
↓ (respawn-controller.ts now ~2,200 LOC)
|
||||
Step 4a: SessionCliBuilder extraction
|
||||
Step 4b: SessionAutoOps extraction
|
||||
Step 4c: SessionTaskCache extraction
|
||||
↓ (session.ts now ~1,600 LOC)
|
||||
```
|
||||
|
||||
**Parallelization**: Steps 1, 2a-2d, 3a-3d, and 4a-4c can be done by separate agents in parallel since they touch different files. However, within each group, sequential execution is safer.
|
||||
|
||||
### Risk Mitigation
|
||||
|
||||
- **Barrel exports**: Every split uses delegation + barrel re-export so external consumers see zero API changes
|
||||
- **Event forwarding**: Sub-modules emit events, parent class forwards them — no event contract changes
|
||||
- **Incremental**: Each step can be verified independently with `tsc --noEmit` + `npm run lint`
|
||||
- **No test changes needed**: External API stays identical; existing tests continue to pass
|
||||
|
||||
---
|
||||
|
||||
## 6. Validation Checklist
|
||||
|
||||
After each step, verify:
|
||||
|
||||
- [ ] `tsc --noEmit` passes (no type errors)
|
||||
- [ ] `npm run lint` passes (no unused imports, etc.)
|
||||
- [ ] `npm run format:check` passes
|
||||
- [ ] `npx vitest run test/respawn-controller.test.ts` passes (for respawn splits)
|
||||
- [ ] `npx vitest run test/ralph-tracker.test.ts` passes (for ralph splits)
|
||||
- [ ] `npx vitest run test/session-manager.test.ts` passes (for session splits)
|
||||
- [ ] Dev server starts: `npx tsx src/index.ts web`
|
||||
- [ ] Existing sessions work (create, interact, delete)
|
||||
- [ ] Respawn cycle works (enable respawn, verify idle detection fires)
|
||||
- [ ] No new circular dependencies: `npx madge --circular src/`
|
||||
|
||||
### Size Targets
|
||||
|
||||
| File | Before | After |
|
||||
|------|--------|-------|
|
||||
| `src/types.ts` | 1,443 LOC | 1 LOC (re-export barrel) |
|
||||
| `src/ralph-tracker.ts` | 3,868 LOC | ~1,800 LOC |
|
||||
| `src/respawn-controller.ts` | 3,611 LOC | ~2,200 LOC |
|
||||
| `src/session.ts` | 2,418 LOC | ~1,600 LOC |
|
||||
| **Total new files** | — | 12 files |
|
||||
| **Net LOC change** | — | ~0 (refactor only) |
|
||||
@@ -1,738 +0,0 @@
|
||||
# Phase 1 Implementation Plan: Quick Wins
|
||||
|
||||
**Source**: `docs/code-structure-findings.md` (Phase 1 - Quick Wins section)
|
||||
**Estimated effort**: 1-2 days
|
||||
**Tasks**: 5 independent tasks (can be done in parallel unless noted)
|
||||
|
||||
---
|
||||
|
||||
## Safety Constraints
|
||||
|
||||
Before starting ANY work, read and follow these rules:
|
||||
|
||||
1. **Never run `npx vitest run`** (full suite) -- it kills tmux sessions. You are running inside a Codeman-managed tmux session.
|
||||
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
|
||||
3. **Never test on port 3000** -- the live dev server runs there. Tests use ports 3150+.
|
||||
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
|
||||
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
|
||||
6. **Never kill tmux sessions** -- check `echo $CODEMAN_MUX` first.
|
||||
|
||||
---
|
||||
|
||||
## Task Dependencies
|
||||
|
||||
All 5 tasks are independent and can be done in parallel. However:
|
||||
- Task 1 (barrel exports) is a prerequisite if you want to update import sites to use the barrel after Task 3 (consolidate EXEC_TIMEOUT_MS). The EXEC_TIMEOUT_MS consolidation creates a new export that should be added to the barrel.
|
||||
- Task 2 (delete dead functions) removes functions that Task 1 would otherwise need to add to the barrel. Do Task 2 first or simultaneously with Task 1 to avoid adding exports for dead code.
|
||||
|
||||
**Recommended order**: Task 2 -> Task 1 -> Task 3 -> Task 4 -> Task 5
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Export Missing Functions from Utils Barrel
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
**Time**: ~30 minutes
|
||||
|
||||
### Problem
|
||||
|
||||
The barrel file (`src/utils/index.ts`) is missing exports for several functions that are defined in util modules, forcing consumers to use deep imports or preventing usage entirely.
|
||||
|
||||
### Missing Exports
|
||||
|
||||
From `src/utils/regex-patterns.ts`:
|
||||
- `createAnsiPatternFull()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
|
||||
- `createAnsiPatternSimple()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
|
||||
- `stripAnsi()` -- ANSI stripping utility
|
||||
- `SAFE_PATH_PATTERN` -- regex for safe file paths (currently deep-imported by `schemas.ts` and `tmux-manager.ts`)
|
||||
|
||||
From `src/utils/token-validation.ts`:
|
||||
- `validateTokenCounts()` -- token count validation (documented in CLAUDE.md)
|
||||
- `validateTokensAndCost()` -- token + cost validation (documented in CLAUDE.md)
|
||||
|
||||
**Note**: Do NOT export `isSimilar`, `isSimilarByDistance`, `levenshteinDistance`, or `normalizePhrase` from `string-similarity.ts` -- these are dead code (see Task 2).
|
||||
|
||||
### Edit 1: Add missing regex-patterns exports
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
|
||||
**Old code** (lines 13-18):
|
||||
```typescript
|
||||
export {
|
||||
ANSI_ESCAPE_PATTERN_FULL,
|
||||
ANSI_ESCAPE_PATTERN_SIMPLE,
|
||||
TOKEN_PATTERN,
|
||||
SPINNER_PATTERN,
|
||||
} from './regex-patterns.js';
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
export {
|
||||
ANSI_ESCAPE_PATTERN_FULL,
|
||||
ANSI_ESCAPE_PATTERN_SIMPLE,
|
||||
TOKEN_PATTERN,
|
||||
SPINNER_PATTERN,
|
||||
createAnsiPatternFull,
|
||||
createAnsiPatternSimple,
|
||||
stripAnsi,
|
||||
SAFE_PATH_PATTERN,
|
||||
} from './regex-patterns.js';
|
||||
```
|
||||
|
||||
### Edit 2: Add missing token-validation exports
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
|
||||
**Old code** (line 19):
|
||||
```typescript
|
||||
export { MAX_SESSION_TOKENS } from './token-validation.js';
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
export { MAX_SESSION_TOKENS, validateTokenCounts, validateTokensAndCost } from './token-validation.js';
|
||||
```
|
||||
|
||||
### Optional follow-up: Update deep imports to use barrel
|
||||
|
||||
These files currently deep-import `SAFE_PATH_PATTERN` and could be updated to use the barrel instead:
|
||||
|
||||
- `src/web/schemas.ts` line 11: `import { SAFE_PATH_PATTERN } from '../utils/regex-patterns.js';` could become `import { SAFE_PATH_PATTERN } from '../utils/index.js';`
|
||||
- `src/tmux-manager.ts` line 44: `import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';` could become part of existing barrel import
|
||||
|
||||
This is a low-priority cosmetic change. The barrel export itself is the important fix.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Delete Dead Utility Functions
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
**Time**: ~15 minutes
|
||||
|
||||
### Problem
|
||||
|
||||
Four exported functions in `string-similarity.ts` are never imported anywhere in the codebase:
|
||||
- `levenshteinDistance()` (lines 27-69)
|
||||
- `isSimilar()` (lines 106-108)
|
||||
- `isSimilarByDistance()` (lines 123-125)
|
||||
- `normalizePhrase()` (lines 139-144)
|
||||
|
||||
Only three functions are actually used (all by `ralph-tracker.ts` via the barrel):
|
||||
- `stringSimilarity()` -- uses `levenshteinDistance()` internally
|
||||
- `fuzzyPhraseMatch()` -- uses `normalizePhrase()` and `isSimilarByDistance()` internally
|
||||
- `todoContentHash()`
|
||||
|
||||
### Strategy
|
||||
|
||||
`levenshteinDistance()` is called by `stringSimilarity()`, and `normalizePhrase()` and `isSimilarByDistance()` are called by `fuzzyPhraseMatch()`. So they cannot be deleted -- they just need to be un-exported (made private to the module).
|
||||
|
||||
`isSimilar()` is truly dead -- not called by anything. Delete it entirely.
|
||||
|
||||
### Edit 1: Remove `export` from `levenshteinDistance`
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (line 27):
|
||||
```typescript
|
||||
export function levenshteinDistance(a: string, b: string): number {
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
function levenshteinDistance(a: string, b: string): number {
|
||||
```
|
||||
|
||||
### Edit 2: Delete `isSimilar` function entirely
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (lines 94-108):
|
||||
```typescript
|
||||
/**
|
||||
* Check if two strings are similar within a given threshold.
|
||||
*
|
||||
* @param a - First string
|
||||
* @param b - Second string
|
||||
* @param threshold - Minimum similarity ratio (default: 0.85 = 85% similar)
|
||||
* @returns True if similarity >= threshold
|
||||
*
|
||||
* @example
|
||||
* isSimilar('COMPLETE', 'COMPLET', 0.85) // true (87.5% similar)
|
||||
* isSimilar('COMPLETE', 'DONE', 0.85) // false (0% similar)
|
||||
*/
|
||||
export function isSimilar(a: string, b: string, threshold = 0.85): boolean {
|
||||
return stringSimilarity(a, b) >= threshold;
|
||||
}
|
||||
```
|
||||
|
||||
**New code**: (delete entirely -- replace with empty string)
|
||||
|
||||
### Edit 3: Remove `export` from `isSimilarByDistance`
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (line 123):
|
||||
```typescript
|
||||
export function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
|
||||
```
|
||||
|
||||
### Edit 4: Remove `export` from `normalizePhrase`
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (line 139):
|
||||
```typescript
|
||||
export function normalizePhrase(phrase: string): string {
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
function normalizePhrase(phrase: string): string {
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npx vitest run test/string-utilities.test.ts
|
||||
npm run lint
|
||||
```
|
||||
|
||||
Note: If `test/string-utilities.test.ts` imports any of the now-unexported functions, those test imports will fail. Check the test file and remove tests for `isSimilar` (deleted) and update any direct tests for `levenshteinDistance`, `isSimilarByDistance`, `normalizePhrase` to test them indirectly through the public API (`stringSimilarity`, `fuzzyPhraseMatch`), or remove those tests.
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Consolidate Duplicated `EXEC_TIMEOUT_MS` Constant
|
||||
|
||||
**Files**:
|
||||
- `src/utils/claude-cli-resolver.ts` (line 17)
|
||||
- `src/utils/opencode-cli-resolver.ts` (line 16)
|
||||
- `src/tmux-manager.ts` (line 63) -- also has its own copy
|
||||
|
||||
**Time**: ~15 minutes
|
||||
|
||||
### Problem
|
||||
|
||||
`EXEC_TIMEOUT_MS = 5000` is defined identically in three files. Changes need to happen in all three places.
|
||||
|
||||
### Strategy
|
||||
|
||||
Create a shared constant and export it. The natural home is a new config file since the existing config files (`buffer-limits.ts`, `map-limits.ts`) follow this pattern. However, to keep it minimal, we can add it to an existing config file or create a small one.
|
||||
|
||||
**Recommended approach**: Add to `src/config/timing-config.ts` (new file) as a single constant. This file can grow later in Phase 6 to hold other timing constants.
|
||||
|
||||
Alternatively, the simplest approach: export from one of the existing utils and import in the others. Since both CLI resolvers are in `src/utils/`, the cleanest approach is to put it in a shared location.
|
||||
|
||||
### Option A: Add to existing config (simpler)
|
||||
|
||||
Create `src/config/exec-timeout.ts`:
|
||||
|
||||
**New file**: `src/config/exec-timeout.ts`
|
||||
```typescript
|
||||
/**
|
||||
* Timeout for child process exec commands (e.g., `which claude`, `which opencode`, tmux commands).
|
||||
* Used across CLI resolvers and tmux manager.
|
||||
*/
|
||||
export const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
### Edit 1: Update `claude-cli-resolver.ts`
|
||||
|
||||
**File**: `src/utils/claude-cli-resolver.ts`
|
||||
|
||||
**Old code** (lines 11-17):
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { delimiter, dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
|
||||
/** Timeout for exec commands (5 seconds) */
|
||||
const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { delimiter, dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
|
||||
```
|
||||
|
||||
### Edit 2: Update `opencode-cli-resolver.ts`
|
||||
|
||||
**File**: `src/utils/opencode-cli-resolver.ts`
|
||||
|
||||
**Old code** (lines 10-16):
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
|
||||
/** Timeout for exec commands (5 seconds) */
|
||||
const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
|
||||
```
|
||||
|
||||
### Edit 3: Update `tmux-manager.ts`
|
||||
|
||||
**File**: `src/tmux-manager.ts`
|
||||
|
||||
**Old code** (line 63):
|
||||
```typescript
|
||||
const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
|
||||
```
|
||||
|
||||
Note: `tmux-manager.ts` already has many imports at the top of the file. Add this import near the other local imports (around lines 43-56). The `const EXEC_TIMEOUT_MS = 5000;` on line 63 should be deleted entirely (replaced with the import).
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Add `z.infer` to Zod Schemas
|
||||
|
||||
**Files**:
|
||||
- `src/web/schemas.ts` (add type exports)
|
||||
- `src/types.ts` (replace manual interfaces with `z.infer` re-exports where applicable)
|
||||
|
||||
**Time**: ~2 hours
|
||||
|
||||
### Problem
|
||||
|
||||
All 30+ Zod schemas in `schemas.ts` define validation rules, but zero use `z.infer` to derive TypeScript types. Instead, `types.ts` manually duplicates interfaces that match the schemas. When a schema changes, the type must be manually updated too.
|
||||
|
||||
### Strategy
|
||||
|
||||
Add `z.infer` type exports to `schemas.ts` for each exported schema. This creates derived types as the single source of truth. For schemas that have corresponding manual interfaces in `types.ts`, the manual interface can be replaced with a re-export of the inferred type.
|
||||
|
||||
**Important**: Not all schemas have matching interfaces in `types.ts`. The `RespawnConfig` interface in `types.ts` (line 395) has all required fields, while `RespawnConfigSchema` has all optional fields (it's for partial updates). These are NOT the same type and should NOT be unified.
|
||||
|
||||
### Edit 1: Add inferred type exports to `schemas.ts`
|
||||
|
||||
**File**: `src/web/schemas.ts`
|
||||
|
||||
After each schema definition, add a corresponding type export. Add the following lines at the **end of the file** (after line 509):
|
||||
|
||||
**Old code** (end of file, lines 506-509):
|
||||
```typescript
|
||||
.optional(),
|
||||
});
|
||||
```
|
||||
|
||||
Wait -- the end of file is actually at line 509 after the `RalphLoopStartSchema`. Add the type exports after the last schema:
|
||||
|
||||
**Append to end of file** `src/web/schemas.ts`:
|
||||
|
||||
```typescript
|
||||
|
||||
// ========== Inferred Types ==========
|
||||
// Derive TypeScript types from Zod schemas (single source of truth)
|
||||
|
||||
export type CreateSessionInput = z.infer<typeof CreateSessionSchema>;
|
||||
export type RunPromptInput = z.infer<typeof RunPromptSchema>;
|
||||
export type ResizeInput = z.infer<typeof ResizeSchema>;
|
||||
export type CreateCaseInput = z.infer<typeof CreateCaseSchema>;
|
||||
export type QuickStartInput = z.infer<typeof QuickStartSchema>;
|
||||
export type HookEventInput = z.infer<typeof HookEventSchema>;
|
||||
export type RespawnConfigInput = z.infer<typeof RespawnConfigSchema>;
|
||||
export type ConfigUpdateInput = z.infer<typeof ConfigUpdateSchema>;
|
||||
export type SettingsUpdateInput = z.infer<typeof SettingsUpdateSchema>;
|
||||
export type SessionInputWithLimitInput = z.infer<typeof SessionInputWithLimitSchema>;
|
||||
export type SessionNameInput = z.infer<typeof SessionNameSchema>;
|
||||
export type SessionColorInput = z.infer<typeof SessionColorSchema>;
|
||||
export type RalphConfigInput = z.infer<typeof RalphConfigSchema>;
|
||||
export type FixPlanImportInput = z.infer<typeof FixPlanImportSchema>;
|
||||
export type RalphPromptWriteInput = z.infer<typeof RalphPromptWriteSchema>;
|
||||
export type AutoClearInput = z.infer<typeof AutoClearSchema>;
|
||||
export type AutoCompactInput = z.infer<typeof AutoCompactSchema>;
|
||||
export type ImageWatcherInput = z.infer<typeof ImageWatcherSchema>;
|
||||
export type FlickerFilterInput = z.infer<typeof FlickerFilterSchema>;
|
||||
export type QuickRunInput = z.infer<typeof QuickRunSchema>;
|
||||
export type ScheduledRunInput = z.infer<typeof ScheduledRunSchema>;
|
||||
export type LinkCaseInput = z.infer<typeof LinkCaseSchema>;
|
||||
export type GeneratePlanInput = z.infer<typeof GeneratePlanSchema>;
|
||||
export type GeneratePlanDetailedInput = z.infer<typeof GeneratePlanDetailedSchema>;
|
||||
export type CancelPlanInput = z.infer<typeof CancelPlanSchema>;
|
||||
export type PlanTaskUpdateInput = z.infer<typeof PlanTaskUpdateSchema>;
|
||||
export type PlanTaskAddInput = z.infer<typeof PlanTaskAddSchema>;
|
||||
export type CpuLimitInput = z.infer<typeof CpuLimitSchema>;
|
||||
export type SubagentWindowStatesInput = z.infer<typeof SubagentWindowStatesSchema>;
|
||||
export type SubagentParentMapInput = z.infer<typeof SubagentParentMapSchema>;
|
||||
export type InteractiveRespawnInput = z.infer<typeof InteractiveRespawnSchema>;
|
||||
export type RespawnEnableInput = z.infer<typeof RespawnEnableSchema>;
|
||||
export type PushSubscribeInput = z.infer<typeof PushSubscribeSchema>;
|
||||
export type PushPreferencesUpdateInput = z.infer<typeof PushPreferencesUpdateSchema>;
|
||||
export type RalphLoopStartInput = z.infer<typeof RalphLoopStartSchema>;
|
||||
```
|
||||
|
||||
### What NOT to do
|
||||
|
||||
Do NOT replace the `RespawnConfig` interface in `types.ts` with `z.infer<typeof RespawnConfigSchema>`. The schema has all optional fields (for partial config updates), but the interface has required fields (for the full config object). These are intentionally different shapes.
|
||||
|
||||
Similarly, do NOT try to unify every interface in `types.ts` with a schema -- most interfaces in `types.ts` represent internal domain objects (SessionState, TaskState, etc.) that have no corresponding Zod schema. The schemas only exist for API request validation.
|
||||
|
||||
### Future opportunity
|
||||
|
||||
In a future phase, route handlers in `server.ts` can use these inferred types for request body typing:
|
||||
```typescript
|
||||
const body = CreateSessionSchema.parse(request.body) as CreateSessionInput;
|
||||
```
|
||||
This task only adds the type exports. Migrating route handlers to use them is out of scope.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
npm run format:check
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Fix Weak `not.toThrow()` Tests with Behavioral Assertions
|
||||
|
||||
**Files**:
|
||||
- `test/task-tracker.test.ts` -- 6 instances
|
||||
- `test/image-watcher.test.ts` -- 1 instance
|
||||
- `test/task-queue.test.ts` -- 1 instance
|
||||
- `test/hooks-config.test.ts` -- 1 instance
|
||||
- `test/session-manager.test.ts` -- 1 instance
|
||||
|
||||
**Time**: ~1 hour
|
||||
|
||||
### Problem
|
||||
|
||||
10 tests only assert `not.toThrow()` without verifying the actual defensive behavior. These tests prove the code doesn't crash but don't verify it does the right thing.
|
||||
|
||||
### Fix Strategy
|
||||
|
||||
After each `not.toThrow()`, add a behavioral assertion that verifies the state is correct (e.g., no tasks were created, no side effects occurred).
|
||||
|
||||
### Edit 1: `task-tracker.test.ts` -- null message (line 566)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle null message', () => {
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle null message', () => {
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
expect(tracker.getRunningCount()).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 2: `task-tracker.test.ts` -- message without content (line 569-571)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle message without content', () => {
|
||||
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle message without content', () => {
|
||||
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 3: `task-tracker.test.ts` -- empty content array (line 573-575)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle empty content array', () => {
|
||||
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle empty content array', () => {
|
||||
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 4: `task-tracker.test.ts` -- tool_result for unknown task (lines 577-590)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle tool_result for unknown task', () => {
|
||||
expect(() => {
|
||||
tracker.processMessage({
|
||||
message: {
|
||||
content: [{
|
||||
type: 'tool_result',
|
||||
tool_use_id: 'unknown-task',
|
||||
is_error: false,
|
||||
content: 'Done',
|
||||
}],
|
||||
},
|
||||
});
|
||||
}).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle tool_result for unknown task', () => {
|
||||
expect(() => {
|
||||
tracker.processMessage({
|
||||
message: {
|
||||
content: [{
|
||||
type: 'tool_result',
|
||||
tool_use_id: 'unknown-task',
|
||||
is_error: false,
|
||||
content: 'Done',
|
||||
}],
|
||||
},
|
||||
});
|
||||
}).not.toThrow();
|
||||
expect(tracker.getTask('unknown-task')).toBeUndefined();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 5: `task-tracker.test.ts` -- empty terminal output (lines 592-595)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle empty terminal output', () => {
|
||||
expect(() => tracker.processTerminalOutput('')).not.toThrow();
|
||||
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle empty terminal output', () => {
|
||||
expect(() => tracker.processTerminalOutput('')).not.toThrow();
|
||||
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
expect(tracker.getRunningCount()).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 6: `image-watcher.test.ts` -- unwatchSession for non-watched session (line 123)
|
||||
|
||||
**File**: `test/image-watcher.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should be safe to call for non-watched session', () => {
|
||||
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should be safe to call for non-watched session', () => {
|
||||
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
|
||||
expect(watcher.getWatchedSessions()).toHaveLength(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 7: `task-queue.test.ts` -- dependencies on non-existent tasks (lines 538-542)
|
||||
|
||||
**File**: `test/task-queue.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
|
||||
// Dependencies on non-existent tasks are valid - they just won't be satisfied
|
||||
expect(() => {
|
||||
queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
|
||||
}).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
|
||||
// Dependencies on non-existent tasks are valid - they just won't be satisfied
|
||||
let task: ReturnType<typeof queue.addTask> | undefined;
|
||||
expect(() => {
|
||||
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
|
||||
}).not.toThrow();
|
||||
expect(task).toBeDefined();
|
||||
expect(task!.dependencies).toEqual(['non-existent-id']);
|
||||
// Task should be pending but blocked (dependency unsatisfied)
|
||||
expect(queue.next()?.prompt).toBeUndefined();
|
||||
});
|
||||
```
|
||||
|
||||
Wait -- `queue.next()` returns `null` when no next task is available (all blocked). Let me adjust:
|
||||
|
||||
**New code** (corrected):
|
||||
```typescript
|
||||
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
|
||||
// Dependencies on non-existent tasks are valid - they just won't be satisfied
|
||||
let task: ReturnType<typeof queue.addTask> | undefined;
|
||||
expect(() => {
|
||||
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
|
||||
}).not.toThrow();
|
||||
expect(task).toBeDefined();
|
||||
expect(task!.dependencies).toEqual(['non-existent-id']);
|
||||
// Task exists but is blocked (dependency unsatisfied), so next() skips it
|
||||
expect(queue.getAllTasks()).toHaveLength(1);
|
||||
expect(queue.next()).toBeNull();
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 8: `hooks-config.test.ts` -- valid JSON check (line 129)
|
||||
|
||||
**File**: `test/hooks-config.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should write valid JSON', () => {
|
||||
writeHooksConfig(testDir);
|
||||
const settingsPath = join(testDir, '.claude', 'settings.local.json');
|
||||
const content = readFileSync(settingsPath, 'utf-8');
|
||||
expect(() => JSON.parse(content)).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should write valid JSON', () => {
|
||||
writeHooksConfig(testDir);
|
||||
const settingsPath = join(testDir, '.claude', 'settings.local.json');
|
||||
const content = readFileSync(settingsPath, 'utf-8');
|
||||
const parsed = JSON.parse(content);
|
||||
expect(parsed).toBeDefined();
|
||||
expect(typeof parsed).toBe('object');
|
||||
expect(parsed.hooks).toBeDefined();
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 9: `session-manager.test.ts` -- stopSession for non-existent (line 216)
|
||||
|
||||
**File**: `test/session-manager.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle non-existent session gracefully', async () => {
|
||||
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle non-existent session gracefully', async () => {
|
||||
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
|
||||
expect(manager.getSessionCount()).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
Run each test file individually:
|
||||
|
||||
```bash
|
||||
npx vitest run test/task-tracker.test.ts
|
||||
npx vitest run test/image-watcher.test.ts
|
||||
npx vitest run test/task-queue.test.ts
|
||||
npx vitest run test/hooks-config.test.ts
|
||||
npx vitest run test/session-manager.test.ts
|
||||
```
|
||||
|
||||
**Important**: `hooks-config.test.ts` and `session-manager.test.ts` spawn real servers on ports 3130-3131. Only run them if you are NOT running other tests that use those ports.
|
||||
|
||||
---
|
||||
|
||||
## Final Verification Checklist
|
||||
|
||||
After all 5 tasks are complete, run the following in order:
|
||||
|
||||
```bash
|
||||
# 1. TypeScript type checking
|
||||
tsc --noEmit
|
||||
|
||||
# 2. Linting
|
||||
npm run lint
|
||||
|
||||
# 3. Formatting
|
||||
npm run format:check
|
||||
|
||||
# 4. Run affected test files individually (NOT the full suite)
|
||||
npx vitest run test/string-utilities.test.ts
|
||||
npx vitest run test/task-tracker.test.ts
|
||||
npx vitest run test/image-watcher.test.ts
|
||||
npx vitest run test/task-queue.test.ts
|
||||
npx vitest run test/session-manager.test.ts
|
||||
npx vitest run test/hooks-config.test.ts
|
||||
```
|
||||
|
||||
If any formatting issues arise, fix with:
|
||||
```bash
|
||||
npm run format
|
||||
```
|
||||
|
||||
If any lint issues arise, fix with:
|
||||
```bash
|
||||
npm run lint:fix
|
||||
```
|
||||
|
||||
### Summary of Changes
|
||||
|
||||
| Task | Files Modified | Files Created |
|
||||
|------|---------------|---------------|
|
||||
| 1. Barrel exports | `src/utils/index.ts` | -- |
|
||||
| 2. Dead functions | `src/utils/string-similarity.ts` | -- |
|
||||
| 3. EXEC_TIMEOUT_MS | `src/utils/claude-cli-resolver.ts`, `src/utils/opencode-cli-resolver.ts`, `src/tmux-manager.ts` | `src/config/exec-timeout.ts` |
|
||||
| 4. z.infer types | `src/web/schemas.ts` | -- |
|
||||
| 5. Weak tests | `test/task-tracker.test.ts`, `test/image-watcher.test.ts`, `test/task-queue.test.ts`, `test/hooks-config.test.ts`, `test/session-manager.test.ts` | -- |
|
||||
|
||||
**Total files modified**: 10
|
||||
**Total files created**: 1
|
||||
@@ -1,689 +0,0 @@
|
||||
# Phase 6 Implementation Plan: Config Consolidation
|
||||
|
||||
**Source**: `docs/code-structure-findings.md` (Phase 6 — Config Consolidation)
|
||||
**Estimated effort**: 1 day
|
||||
**Tasks**: 8 tasks with dependencies (see dependency graph below)
|
||||
|
||||
---
|
||||
|
||||
## Safety Constraints
|
||||
|
||||
Before starting ANY work, read and follow these rules:
|
||||
|
||||
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
|
||||
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
|
||||
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
|
||||
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
|
||||
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
|
||||
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
|
||||
7. **Verify the dev server starts**: After each task, run `npx tsx src/index.ts web --port 3099 &` on a non-production port, confirm `curl -s http://localhost:3099/api/status | jq .status` returns `"ok"`, then kill the background process.
|
||||
|
||||
---
|
||||
|
||||
## Goal
|
||||
|
||||
Consolidate ~70 scattered numeric constants from 15+ source files into 6 new domain-focused config files, eliminating cross-file duplicates (including a 5x-duplicated AI model string) and making all tuning knobs discoverable in `src/config/`.
|
||||
|
||||
**Non-goal**: Moving every constant. Module-internal implementation details (like regex patterns, algorithm-specific magic numbers, or constants only used once in deeply coupled logic) stay where they are. The goal is discoverability of operational tuning knobs, not mechanical relocation.
|
||||
|
||||
---
|
||||
|
||||
## Design Decisions
|
||||
|
||||
### What gets centralized (and why)
|
||||
|
||||
Constants are candidates for centralization when they meet **any** of these criteria:
|
||||
|
||||
1. **Duplicated across files** — DRY violation (e.g., `STATS_COLLECTION_INTERVAL_MS` in `server.ts` and `mux-routes.ts`, AI model string in 5 files)
|
||||
2. **Operational tuning knobs** — values an operator might want to adjust for performance, security, or behavior without understanding the implementation (e.g., SSE health check interval, auth session TTL, rate limits)
|
||||
3. **Cross-cutting concerns** — values that establish system-wide contracts (e.g., max terminal dimensions used by both server routes and frontend)
|
||||
|
||||
### What stays in place (and why)
|
||||
|
||||
Constants that are **internal implementation details** of a single module stay where they are:
|
||||
|
||||
- **Algorithm parameters** — `TODO_SIMILARITY_THRESHOLD`, `adaptiveCompletionConfirmMs`, confidence weights. These are meaningless without understanding the algorithm.
|
||||
- **Display/UI formatting** — `TEXT_PREVIEW_LENGTH`, `SMART_TITLE_MAX_LENGTH`, `COMMAND_DISPLAY_LENGTH` in `subagent-watcher.ts`. Only used locally, tightly coupled to rendering logic.
|
||||
- **Module-internal timing** — `LINE_BUFFER_FLUSH_INTERVAL` in `session.ts`, `AI_CHECK_POLL_INTERVAL` in `ai-checker-base.ts`. Internal implementation of specific features.
|
||||
- **Frontend constants** — `constants.js` already centralizes frontend values well. Don't mix frontend and backend config.
|
||||
- **Respawn `DEFAULT_CONFIG`** — these are user-configurable defaults for the respawn config interface, not system constants. They live properly in `respawn-controller.ts`. The AI model/context defaults within it are replaced with imports from the new `ai-defaults.ts` (Task 5).
|
||||
- **Session auto-ops thresholds** — `AUTO_RETRY_DELAY_MS`, `COMPACT_COOLDOWN_MS`, etc. in `session-auto-ops.ts` are internal to that module's retry logic and already well-documented in place.
|
||||
|
||||
### File organization: domain-based, not category-based
|
||||
|
||||
A single `timing-config.ts` with 70 unrelated timing values would be worse than the current state — developers would need to grep it just like they grep the whole codebase now. Instead, constants are grouped by **the system they configure**:
|
||||
|
||||
| New File | Domain | Developer Question It Answers |
|
||||
|----------|--------|-------------------------------|
|
||||
| `server-timing.ts` | Web server performance | "How do I tune SSE batching / terminal throughput?" |
|
||||
| `auth-config.ts` | Authentication & security | "What are the rate limits and session TTLs?" |
|
||||
| `tunnel-config.ts` | QR auth & Cloudflare tunnel | "What are the QR token rotation parameters?" |
|
||||
| `terminal-limits.ts` | Terminal dimensions & input | "What are the max cols/rows/input size?" |
|
||||
| `ai-defaults.ts` | AI checker model & context | "What model do the AI checkers use? What's the context limit?" |
|
||||
| `team-config.ts` | Agent Teams polling & caching | "How often does team polling run? What are the cache limits?" |
|
||||
|
||||
---
|
||||
|
||||
## Task Dependencies
|
||||
|
||||
```
|
||||
Task 1 (server-timing.ts)
|
||||
Task 2 (auth-config.ts)
|
||||
Task 3 (tunnel-config.ts)
|
||||
Task 4 (terminal-limits.ts)
|
||||
Task 5 (ai-defaults.ts)
|
||||
Task 6 (team-config.ts)
|
||||
└──> Task 7 (Fix remaining duplicates)
|
||||
└──> Task 8 (Update CLAUDE.md + final verification)
|
||||
```
|
||||
|
||||
**Tasks 1–6** are independent and can run in parallel.
|
||||
**Task 7** depends on Tasks 1–6 (needs the new config files to exist).
|
||||
**Task 8** depends on Task 7.
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Create `src/config/server-timing.ts`
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Files created**: `src/config/server-timing.ts`
|
||||
**Files modified**: `src/web/server.ts`, `src/web/routes/mux-routes.ts`
|
||||
|
||||
### Constants to extract from `src/web/server.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `TERMINAL_BATCH_INTERVAL` | `16` | Terminal data batching interval (60fps) |
|
||||
| `TASK_UPDATE_BATCH_INTERVAL` | `100` | Task event batching interval (ms) |
|
||||
| `STATE_UPDATE_DEBOUNCE_INTERVAL` | `500` | State persistence debounce (ms) |
|
||||
| `SESSIONS_LIST_CACHE_TTL` | `1000` | Sessions list cache TTL (ms) |
|
||||
| `SCHEDULED_CLEANUP_INTERVAL` | `300000` | Scheduled runs cleanup check (5 min) |
|
||||
| `SCHEDULED_RUN_MAX_AGE` | `3600000` | Completed scheduled run max age (1 hour) |
|
||||
| `SSE_HEALTH_CHECK_INTERVAL` | `30000` | SSE client health check (30s) |
|
||||
| `SESSION_LIMIT_WAIT_MS` | `5000` | Session limit retry wait (5s) |
|
||||
| `ITERATION_PAUSE_MS` | `2000` | Scheduled run iteration pause (2s) |
|
||||
| `BATCH_FLUSH_THRESHOLD` | `32768` | Terminal batch immediate flush threshold (32KB) |
|
||||
| `STATS_COLLECTION_INTERVAL_MS` | `2000` | Mux stats collection interval (2s) |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/server-timing.ts` with all 11 constants, preserving existing JSDoc comments.
|
||||
2. In `src/web/server.ts`: Remove the 11 local constant declarations (lines ~92–121). Add `import { TERMINAL_BATCH_INTERVAL, ... } from '../config/server-timing.js'`.
|
||||
3. In `src/web/routes/mux-routes.ts`: Remove the duplicate `STATS_COLLECTION_INTERVAL_MS` (line 10) and its comment. Add `import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js'`. This fixes a **duplicate constant** (finding #10).
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Web server performance and scheduling constants.
|
||||
*
|
||||
* Controls terminal batching throughput, SSE health checking,
|
||||
* state persistence debouncing, and scheduled run timing.
|
||||
*
|
||||
* @module config/server-timing
|
||||
*/
|
||||
|
||||
// ============================================================================
|
||||
// Terminal & SSE Performance
|
||||
// ============================================================================
|
||||
|
||||
/** Terminal data batching interval — targets 60fps (ms) */
|
||||
export const TERMINAL_BATCH_INTERVAL = 16;
|
||||
|
||||
/** Immediate flush threshold for terminal batches (bytes).
|
||||
* Set high (32KB) to allow effective batching; avg Ink events are ~14KB. */
|
||||
export const BATCH_FLUSH_THRESHOLD = 32 * 1024;
|
||||
|
||||
/** Task event batching interval (ms) */
|
||||
export const TASK_UPDATE_BATCH_INTERVAL = 100;
|
||||
|
||||
/** SSE client health check interval (ms) */
|
||||
export const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
|
||||
|
||||
// ============================================================================
|
||||
// State Persistence
|
||||
// ============================================================================
|
||||
|
||||
/** State update debounce — batches expensive toDetailedState() calls (ms) */
|
||||
export const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
|
||||
|
||||
/** Sessions list cache TTL — avoids re-serializing on every SSE init (ms) */
|
||||
export const SESSIONS_LIST_CACHE_TTL = 1000;
|
||||
|
||||
// ============================================================================
|
||||
// Scheduled Runs
|
||||
// ============================================================================
|
||||
|
||||
/** Scheduled runs cleanup check interval (ms) */
|
||||
export const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
|
||||
|
||||
/** Completed scheduled run max age before cleanup (ms) */
|
||||
export const SCHEDULED_RUN_MAX_AGE = 60 * 60 * 1000;
|
||||
|
||||
/** Session limit retry wait before retrying (ms) */
|
||||
export const SESSION_LIMIT_WAIT_MS = 5000;
|
||||
|
||||
/** Pause between scheduled run iterations (ms) */
|
||||
export const ITERATION_PAUSE_MS = 2000;
|
||||
|
||||
// ============================================================================
|
||||
// Mux Stats
|
||||
// ============================================================================
|
||||
|
||||
/** Mux stats collection interval (ms) */
|
||||
export const STATS_COLLECTION_INTERVAL_MS = 2000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npx tsx src/index.ts web --port 3099 &
|
||||
curl -s http://localhost:3099/api/status | jq .status # "ok"
|
||||
kill %1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Create `src/config/auth-config.ts`
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files created**: `src/config/auth-config.ts`
|
||||
**Files modified**: `src/web/middleware/auth.ts`, `src/hooks-config.ts`
|
||||
|
||||
### Constants to extract from `src/web/middleware/auth.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `AUTH_SESSION_TTL_MS` | `86400000` | Auth session cookie TTL (24h) |
|
||||
| `MAX_AUTH_SESSIONS` | `100` | Max concurrent auth sessions |
|
||||
| `AUTH_FAILURE_MAX` | `10` | Max failed auth attempts per IP |
|
||||
| `AUTH_FAILURE_WINDOW_MS` | `900000` | Failed auth tracking window (15 min) |
|
||||
|
||||
### Constants to extract from `src/hooks-config.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `HOOK_TIMEOUT_MS` | `10000` | Timeout for Claude Code hook commands |
|
||||
|
||||
The `timeout: 10000` value is hardcoded 6 times in `hooks-config.ts` as inline literals. Extract to a single named constant.
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/auth-config.ts` with the 5 constants.
|
||||
2. In `src/web/middleware/auth.ts`: Remove the 4 local constant declarations (lines 17–25). Add import from `../../config/auth-config.js`. Keep `AUTH_COOKIE_NAME` in place — it's a string identifier, not a tunable numeric constant.
|
||||
3. In `src/hooks-config.ts`: Replace all 6 inline `timeout: 10000` occurrences with `timeout: HOOK_TIMEOUT_MS`. Add import from `./config/auth-config.js`.
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Authentication, rate limiting, and hook security constants.
|
||||
*
|
||||
* Controls auth session lifecycle, brute-force protection,
|
||||
* and Claude Code hook timeouts.
|
||||
*
|
||||
* @module config/auth-config
|
||||
*/
|
||||
|
||||
// ============================================================================
|
||||
// Session Cookies
|
||||
// ============================================================================
|
||||
|
||||
/** Auth session cookie TTL — matches autonomous run length (ms) */
|
||||
export const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
|
||||
|
||||
/** Max concurrent auth sessions per server */
|
||||
export const MAX_AUTH_SESSIONS = 100;
|
||||
|
||||
// ============================================================================
|
||||
// Rate Limiting
|
||||
// ============================================================================
|
||||
|
||||
/** Max failed auth attempts per IP before 429 rejection */
|
||||
export const AUTH_FAILURE_MAX = 10;
|
||||
|
||||
/** Failed auth attempt tracking window (ms) */
|
||||
export const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
|
||||
|
||||
// ============================================================================
|
||||
// Hooks
|
||||
// ============================================================================
|
||||
|
||||
/** Timeout for Claude Code hook curl commands (ms) */
|
||||
export const HOOK_TIMEOUT_MS = 10000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Create `src/config/tunnel-config.ts`
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files created**: `src/config/tunnel-config.ts`
|
||||
**Files modified**: `src/tunnel-manager.ts`
|
||||
|
||||
### Constants to extract from `src/tunnel-manager.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `QR_TOKEN_TTL_MS` | `60000` | QR token auto-rotation interval (60s) |
|
||||
| `QR_TOKEN_GRACE_MS` | `90000` | Grace period for previous token (90s) |
|
||||
| `SHORT_CODE_LENGTH` | `6` | Length of QR short code |
|
||||
| `QR_RATE_LIMIT_MAX` | `30` | Global QR attempt rate limit |
|
||||
| `QR_RATE_LIMIT_WINDOW_MS` | `60000` | QR rate limit reset window (60s) |
|
||||
| `URL_TIMEOUT_MS` | `30000` | Cloudflared URL fetch timeout (30s) |
|
||||
| `RESTART_DELAY_MS` | `5000` | Tunnel restart delay after crash (5s) |
|
||||
| `FORCE_KILL_MS` | `5000` | SIGTERM → SIGKILL escalation timeout (5s) |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/tunnel-config.ts` with all 8 constants.
|
||||
2. In `src/tunnel-manager.ts`: Remove the 8 local constant declarations (lines ~39–75). Add `import { QR_TOKEN_TTL_MS, ... } from './config/tunnel-config.js'`.
|
||||
3. Keep the `TUNNEL_URL_REGEX` in `tunnel-manager.ts` — it's a parsing detail, not a tuning knob.
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Cloudflare tunnel and QR authentication constants.
|
||||
*
|
||||
* Controls QR token rotation timing, rate limiting,
|
||||
* and tunnel process lifecycle.
|
||||
*
|
||||
* @module config/tunnel-config
|
||||
*/
|
||||
|
||||
// ============================================================================
|
||||
// QR Token Rotation
|
||||
// ============================================================================
|
||||
|
||||
/** QR token auto-rotation interval (ms) */
|
||||
export const QR_TOKEN_TTL_MS = 60_000;
|
||||
|
||||
/** Grace period — previous token still valid during rotation (ms) */
|
||||
export const QR_TOKEN_GRACE_MS = 90_000;
|
||||
|
||||
/** Length of the short code in QR URL path (chars) */
|
||||
export const SHORT_CODE_LENGTH = 6;
|
||||
|
||||
// ============================================================================
|
||||
// QR Rate Limiting
|
||||
// ============================================================================
|
||||
|
||||
/** Global rate limit for QR auth attempts across all IPs */
|
||||
export const QR_RATE_LIMIT_MAX = 30;
|
||||
|
||||
/** QR rate limit reset window (ms) */
|
||||
export const QR_RATE_LIMIT_WINDOW_MS = 60_000;
|
||||
|
||||
// ============================================================================
|
||||
// Tunnel Process Lifecycle
|
||||
// ============================================================================
|
||||
|
||||
/** Max time to wait for cloudflared URL before timeout (ms) */
|
||||
export const URL_TIMEOUT_MS = 30_000;
|
||||
|
||||
/** Restart delay after unexpected tunnel exit (ms) */
|
||||
export const RESTART_DELAY_MS = 5_000;
|
||||
|
||||
/** SIGTERM → SIGKILL escalation timeout (ms) */
|
||||
export const FORCE_KILL_MS = 5_000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Create `src/config/terminal-limits.ts`
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files created**: `src/config/terminal-limits.ts`
|
||||
**Files modified**: `src/web/routes/session-routes.ts`
|
||||
|
||||
### Constants to extract from `src/web/routes/session-routes.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `MAX_INPUT_LENGTH` | `65536` | Max input length per request (64KB) |
|
||||
| `MAX_TERMINAL_COLS` | `500` | Max terminal columns |
|
||||
| `MAX_TERMINAL_ROWS` | `200` | Max terminal rows |
|
||||
| `MAX_SESSION_NAME_LENGTH` | `128` | Max session name length (chars) |
|
||||
|
||||
### Why a separate file instead of adding to `buffer-limits.ts`
|
||||
|
||||
`buffer-limits.ts` covers memory buffer sizes (2MB terminal, 1MB text). These constants are **validation limits** for API inputs — different concern. A terminal resize request must not exceed `MAX_TERMINAL_COLS`; this has nothing to do with buffer trimming.
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/terminal-limits.ts` with all 4 constants.
|
||||
2. In `src/web/routes/session-routes.ts`: Remove the 4 local constant declarations (lines 45–48). Add `import { MAX_INPUT_LENGTH, MAX_TERMINAL_COLS, MAX_TERMINAL_ROWS, MAX_SESSION_NAME_LENGTH } from '../../config/terminal-limits.js'`.
|
||||
3. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Terminal dimension and input validation limits.
|
||||
*
|
||||
* Used by API routes to validate resize, input, and session
|
||||
* creation requests. Separate from buffer-limits.ts which
|
||||
* controls memory buffer sizes.
|
||||
*
|
||||
* @module config/terminal-limits
|
||||
*/
|
||||
|
||||
/** Max input length per API request (bytes) */
|
||||
export const MAX_INPUT_LENGTH = 64 * 1024;
|
||||
|
||||
/** Max terminal columns for resize requests */
|
||||
export const MAX_TERMINAL_COLS = 500;
|
||||
|
||||
/** Max terminal rows for resize requests */
|
||||
export const MAX_TERMINAL_ROWS = 200;
|
||||
|
||||
/** Max session name length (chars) */
|
||||
export const MAX_SESSION_NAME_LENGTH = 128;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Create `src/config/ai-defaults.ts`
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Files created**: `src/config/ai-defaults.ts`
|
||||
**Files modified**: `src/respawn-controller.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts`, `src/web/routes/respawn-routes.ts`
|
||||
|
||||
### Problem: AI model string duplicated 5 times
|
||||
|
||||
The model identifier `'claude-opus-4-5-20251101'` appears in 5 places across 4 files. When the model changes, all 5 must be updated — a guaranteed source of bugs. The context limits (`16000`, `8000`) are similarly scattered across 3 files each.
|
||||
|
||||
| Constant | Current Value | Duplicated In |
|
||||
|----------|---------------|---------------|
|
||||
| `AI_CHECK_MODEL` | `'claude-opus-4-5-20251101'` | `respawn-controller.ts` (×2: idle + plan), `ai-idle-checker.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` (×2: idle + plan) |
|
||||
| `AI_IDLE_CHECK_MAX_CONTEXT` | `16000` | `respawn-controller.ts`, `ai-idle-checker.ts`, `respawn-routes.ts` |
|
||||
| `AI_PLAN_CHECK_MAX_CONTEXT` | `8000` | `respawn-controller.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/ai-defaults.ts` with the 3 constants.
|
||||
2. In `src/respawn-controller.ts` `DEFAULT_CONFIG` (line 538): Replace `aiIdleCheckModel: 'claude-opus-4-5-20251101'` with `aiIdleCheckModel: AI_CHECK_MODEL`, `aiIdleCheckMaxContext: 16000` with `aiIdleCheckMaxContext: AI_IDLE_CHECK_MAX_CONTEXT`, `aiPlanCheckModel: 'claude-opus-4-5-20251101'` with `aiPlanCheckModel: AI_CHECK_MODEL`, `aiPlanCheckMaxContext: 8000` with `aiPlanCheckMaxContext: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
|
||||
3. In `src/ai-idle-checker.ts` `DEFAULT_AI_CHECK_CONFIG` (line 46): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 16000` with `maxContextChars: AI_IDLE_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
|
||||
4. In `src/ai-plan-checker.ts` `DEFAULT_PLAN_CHECK_CONFIG` (line 45): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 8000` with `maxContextChars: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
|
||||
5. In `src/web/routes/respawn-routes.ts` config merge block (lines 173–179): Replace all 4 inline fallback values with imports from `../../config/ai-defaults.js`.
|
||||
6. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Default model and context limits for AI-powered checkers.
|
||||
*
|
||||
* Centralizes the AI model identifier and context window sizes used by
|
||||
* the idle checker, plan checker, respawn controller defaults, and
|
||||
* respawn route fallbacks. Change the model here when upgrading.
|
||||
*
|
||||
* @module config/ai-defaults
|
||||
*/
|
||||
|
||||
/** Default model for AI idle and plan checkers */
|
||||
export const AI_CHECK_MODEL = 'claude-opus-4-5-20251101';
|
||||
|
||||
/** Max context chars for idle checker (~4k tokens) */
|
||||
export const AI_IDLE_CHECK_MAX_CONTEXT = 16000;
|
||||
|
||||
/** Max context chars for plan checker (~2k tokens, plan mode UI is compact) */
|
||||
export const AI_PLAN_CHECK_MAX_CONTEXT = 8000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
# Verify no remaining hardcoded model strings
|
||||
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 6: Create `src/config/team-config.ts`
|
||||
|
||||
**Estimated effort**: 15 minutes
|
||||
**Files created**: `src/config/team-config.ts`
|
||||
**Files modified**: `src/team-watcher.ts`
|
||||
|
||||
### Constants to extract from `src/team-watcher.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `TEAM_POLL_INTERVAL_MS` | `30000` | Team directory poll interval (30s) |
|
||||
| `MAX_CACHED_TEAMS` | `50` | LRU cache size for team configs |
|
||||
| `MAX_CACHED_TASKS` | `200` | LRU cache size for team tasks + inboxes |
|
||||
|
||||
### Why centralize these
|
||||
|
||||
Team polling frequency and cache sizes are operational knobs that affect both performance (polling too often wastes CPU) and responsiveness (polling too rarely means stale team state in the UI). They're also the kind of values a developer tuning for a large team deployment would want to find quickly. `MAX_CACHED_TASKS` is used for both the task cache and inbox cache — worth documenting.
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/team-config.ts` with the 3 constants.
|
||||
2. In `src/team-watcher.ts`: Remove the 3 local constants (lines 23–25). Add `import { TEAM_POLL_INTERVAL_MS, MAX_CACHED_TEAMS, MAX_CACHED_TASKS } from './config/team-config.js'`. Note: rename `POLL_INTERVAL_MS` → `TEAM_POLL_INTERVAL_MS` to avoid ambiguity with the identically-named constant in `subagent-watcher.ts`.
|
||||
3. Update the usage site: `setInterval(... POLL_INTERVAL_MS)` → `setInterval(... TEAM_POLL_INTERVAL_MS)`.
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Agent Teams polling and cache configuration.
|
||||
*
|
||||
* Controls how frequently TeamWatcher polls ~/.claude/teams/
|
||||
* and how many teams/tasks are cached in memory.
|
||||
*
|
||||
* @module config/team-config
|
||||
*/
|
||||
|
||||
/** Team directory poll interval (ms) */
|
||||
export const TEAM_POLL_INTERVAL_MS = 30_000;
|
||||
|
||||
/** Max cached team configs (LRU eviction) */
|
||||
export const MAX_CACHED_TEAMS = 50;
|
||||
|
||||
/** Max cached team tasks and inbox messages (LRU eviction).
|
||||
* Used for both teamTasks and inboxCache maps. */
|
||||
export const MAX_CACHED_TASKS = 200;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 7: Fix remaining cross-file duplicates
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Files modified**: `src/index.ts`, `src/subagent-watcher.ts`
|
||||
|
||||
### Duplicate 1: `STATS_COLLECTION_INTERVAL_MS`
|
||||
|
||||
Already fixed in Task 1 — both `server.ts` and `mux-routes.ts` now import from `server-timing.ts`.
|
||||
|
||||
### Duplicate 2: AI model string
|
||||
|
||||
Already fixed in Task 5 — all 5 occurrences now import from `ai-defaults.ts`.
|
||||
|
||||
### Duplicate 3: `MAX_SCREENSHOT_SIZE` / `MAX_TEXT_FILE_SIZE` / `MAX_RAW_FILE_SIZE`
|
||||
|
||||
These file size limits in `file-routes.ts` and `system-routes.ts` are **API-specific validation limits**. They're only used in their respective route files and aren't duplicated. **Leave in place** — they're local to their route module and well-commented.
|
||||
|
||||
### Action A: Move `MAX_CONSECUTIVE_ERRORS` and `ERROR_RESET_MS` to config
|
||||
|
||||
`src/index.ts` has two process-level constants that are operational tuning knobs:
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `MAX_CONSECUTIVE_ERRORS` | `5` | Max consecutive unhandled errors before process exit |
|
||||
| `ERROR_RESET_MS` | `60000` | Error counter reset interval (1 min) |
|
||||
|
||||
These belong in a config file since they control server reliability behavior. Add them to `src/config/server-timing.ts` (they're server operational constants).
|
||||
|
||||
1. Add to `src/config/server-timing.ts`:
|
||||
```typescript
|
||||
// ============================================================================
|
||||
// Process Error Recovery
|
||||
// ============================================================================
|
||||
|
||||
/** Max consecutive unhandled errors before auto-restart */
|
||||
export const MAX_CONSECUTIVE_ERRORS = 5;
|
||||
|
||||
/** Error counter reset interval — forgives errors after quiet period (ms) */
|
||||
export const ERROR_RESET_MS = 60_000;
|
||||
```
|
||||
2. In `src/index.ts`: Remove lines 19–20, add import from `./config/server-timing.js`.
|
||||
3. Run `tsc --noEmit`.
|
||||
|
||||
### Action B: Fix `MAX_TRACKED_AGENTS` shadow in `subagent-watcher.ts`
|
||||
|
||||
`subagent-watcher.ts` defines its own `MAX_TRACKED_AGENTS = 500` locally instead of importing the identical value from `config/map-limits.ts`. This is a latent bug — if someone changes the config value, the subagent watcher's copy stays stale.
|
||||
|
||||
1. In `src/subagent-watcher.ts`: Remove the local `MAX_TRACKED_AGENTS` constant. Add `import { MAX_TRACKED_AGENTS } from './config/map-limits.js'` (the value there is `MAX_TODOS_PER_SESSION = 500` — **verify** the map-limits constant is actually named `MAX_TRACKED_AGENTS` or if it needs to be added). If the constant doesn't exist in `map-limits.ts` under that name, add it.
|
||||
2. Run `tsc --noEmit`.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
npm run format:check
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 8: Update CLAUDE.md and final verification
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files modified**: `CLAUDE.md`
|
||||
|
||||
### Updates to CLAUDE.md
|
||||
|
||||
1. **Config Files table** (`src/config/`): Add the 6 new files:
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `buffer-limits.ts` | Terminal/text buffer size limits |
|
||||
| `map-limits.ts` | Global limits for Maps, sessions, watchers |
|
||||
| `exec-timeout.ts` | Execution timeout configuration |
|
||||
| `server-timing.ts` | Web server batching, SSE, scheduled run timing |
|
||||
| `auth-config.ts` | Auth session TTL, rate limits, hook timeout |
|
||||
| `tunnel-config.ts` | QR token rotation, tunnel process lifecycle |
|
||||
| `terminal-limits.ts` | Terminal dimension and input validation limits |
|
||||
| `ai-defaults.ts` | AI checker model and context limits |
|
||||
| `team-config.ts` | Agent Teams polling and cache sizes |
|
||||
|
||||
2. **Import Conventions** section: Add:
|
||||
```
|
||||
- **Config**: Import from specific files: `import { MAX_TERMINAL_COLS } from './config/terminal-limits'`
|
||||
```
|
||||
|
||||
3. **Phase 6 status** in `docs/code-structure-findings.md`: Mark as COMPLETE with summary of what was done.
|
||||
|
||||
### Final verification checklist
|
||||
|
||||
```bash
|
||||
# Type checking
|
||||
tsc --noEmit
|
||||
|
||||
# Linting
|
||||
npm run lint
|
||||
|
||||
# Formatting
|
||||
npm run format:check
|
||||
|
||||
# Dev server starts
|
||||
npx tsx src/index.ts web --port 3099 &
|
||||
curl -s http://localhost:3099/api/status | jq .status # "ok"
|
||||
kill %1
|
||||
|
||||
# Verify no remaining duplicates
|
||||
grep -rn 'STATS_COLLECTION_INTERVAL_MS' src/ # Should only appear in config + import sites
|
||||
grep -rn 'timeout: 10000' src/hooks-config.ts # Should be 0 — all replaced with HOOK_TIMEOUT_MS
|
||||
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## What is NOT in scope (and why)
|
||||
|
||||
These constants were considered but deliberately left in their current files:
|
||||
|
||||
### Respawn controller defaults (`src/respawn-controller.ts`)
|
||||
|
||||
The `DEFAULT_CONFIG` object (lines 538–578) contains ~30 default values for the `RespawnConfig` interface. These are **user-facing configuration defaults**, not system constants — they're the starting values for a config object that users can modify via the API and UI. Centralizing them would break the locality between the config interface definition and its defaults. They already have excellent JSDoc with `@default` tags. The only values extracted are the AI model/context constants (Task 5) which are duplicated in other files.
|
||||
|
||||
### Subagent watcher timing (`src/subagent-watcher.ts`)
|
||||
|
||||
The 18 constants at lines 129–158 are all internal to the subagent watcher's polling/lifecycle algorithm. Moving them to a config file would force developers to context-switch between two files to understand the polling logic. They're already grouped with clear comments. Exception: `MAX_TRACKED_AGENTS` is consolidated with `map-limits.ts` (Task 7B) since it duplicates a global limit.
|
||||
|
||||
### Session auto-ops timing (`src/session-auto-ops.ts`)
|
||||
|
||||
The 8 constants at lines 19–40 are internal to the auto-compact/clear retry state machine. They form a coherent group that's meaningless without the surrounding implementation context.
|
||||
|
||||
### Run summary constants (`src/run-summary.ts`)
|
||||
|
||||
`MAX_EVENTS`, `TRIM_TO_EVENTS`, `TOKEN_MILESTONE_INTERVAL`, `STATE_STUCK_WARNING_MS`, `STATE_STUCK_CHECK_INTERVAL` — all module-internal. The buffer-style limits (`MAX_EVENTS`/`TRIM_TO_EVENTS`) follow the same pattern as `buffer-limits.ts` but are only used in this one file.
|
||||
|
||||
### Frontend (`src/web/public/constants.js`)
|
||||
|
||||
Already well-centralized. Frontend and backend run in different environments — mixing them in TypeScript config files would create import problems. If frontend constants need expansion, do it in `constants.js`. Note: `app.js` has 2 inline uses of `256 * 1024` that should use the existing `TERMINAL_TAIL_SIZE` from `constants.js` — a minor cleanup that can be done opportunistically but is not worth a task here.
|
||||
|
||||
### Tmux manager timing (`src/tmux-manager.ts`)
|
||||
|
||||
The 6 constants (lines 65–78) are internal to tmux process lifecycle management. They're low-level retry/wait values that are meaningless without understanding the tmux spawn sequence.
|
||||
|
||||
### Process-internal constants
|
||||
|
||||
`image-watcher.ts`, `bash-tool-parser.ts`, `transcript-watcher.ts`, `ralph-tracker.ts`, `task-tracker.ts`, `file-stream-manager.ts`, `session-lifecycle-log.ts`, `session-task-cache.ts`, `respawn-metrics.ts`, `respawn-adaptive-timing.ts`, `ai-checker-base.ts` — all have module-local constants that are internal implementation details.
|
||||
|
||||
### `localhost:3000` default URL
|
||||
|
||||
The string `'http://localhost:3000'` or port `3000` appears as a fallback default in ~5 files (`session-cli-builder.ts`, `tmux-manager.ts`, `tunnel-manager.ts`, `server.ts`, CLI). While technically duplicated, extracting it provides little value — each usage has a different fallback chain (env var → config → hardcoded) and the port is also baked into systemd service files and documentation. The risk of a missed update is low since port 3000 is deeply conventional.
|
||||
|
||||
### `SAVE_DEBOUNCE_MS = 500` in `state-store.ts` / `push-store.ts`
|
||||
|
||||
Same value (500ms), but they debounce different persistence targets (state.json vs push-subscriptions.json). If one needed faster/slower debouncing, they'd diverge. Coupling them would be misleading.
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
| Metric | Before | After |
|
||||
|--------|--------|-------|
|
||||
| Config files in `src/config/` | 3 | 9 |
|
||||
| Constants centralized | ~25 | ~65 |
|
||||
| Cross-file duplicates | 9+ (`STATS_COLLECTION_INTERVAL_MS`, `timeout: 10000` ×6, AI model ×5, context limits ×3 each, `MAX_TRACKED_AGENTS`) | 0 |
|
||||
| Files with `timeout: 10000` inline | 1 (6 occurrences) | 0 |
|
||||
| Files with hardcoded AI model string | 4 (5 occurrences) | 1 (config only) |
|
||||
| Files modified | — | 11 |
|
||||
| Files created | — | 6 |
|
||||
@@ -1,953 +0,0 @@
|
||||
# Phase 7 Implementation Plan: Test Infrastructure
|
||||
|
||||
**Source**: `docs/code-structure-findings.md` (Phase 7 — Test Infrastructure)
|
||||
**Estimated effort**: 2–3 days
|
||||
**Tasks**: 11 tasks with dependencies (see dependency graph below)
|
||||
|
||||
---
|
||||
|
||||
## Safety Constraints
|
||||
|
||||
Before starting ANY work, read and follow these rules:
|
||||
|
||||
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
|
||||
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
|
||||
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
|
||||
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
|
||||
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
|
||||
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
|
||||
7. **Port assignments for this phase**: New tests use ports 3220–3229 (see individual tasks for assignments).
|
||||
|
||||
---
|
||||
|
||||
## Goal
|
||||
|
||||
Eliminate duplicated test mocks, activate the unused `respawn-test-utils.ts` utilities, and add route-level test coverage for the server's 12 route modules — the single largest untested area in the codebase (162 route handlers, 0 dedicated tests).
|
||||
|
||||
**Non-goals**:
|
||||
- Full end-to-end integration tests (those require real Claude CLI / tmux sessions)
|
||||
- 100% route coverage in this phase — focus on the highest-value route modules first
|
||||
- Refactoring test patterns in existing passing tests that don't use shared mocks
|
||||
- Migrating `vi.mock()`-based module replacement mocks (different pattern, see Task 6/7)
|
||||
|
||||
---
|
||||
|
||||
## Current State
|
||||
|
||||
### Mock Duplication (Finding #9)
|
||||
|
||||
`MockSession` is defined **4 times** across test files with varying levels of completeness:
|
||||
|
||||
| File | Properties | Methods | EventEmitter | Notes |
|
||||
|------|-----------|---------|-------------|-------|
|
||||
| `test/respawn-test-utils.ts` | 6 | 20+ | Yes | **Most complete**. Includes terminal simulation, token count, ANSI output, plan mode prompts. **Never imported by any test.** |
|
||||
| `test/respawn-controller.test.ts` | 6 | 9 | Yes | Subset of respawn-test-utils. Missing token simulation, ANSI helpers. |
|
||||
| `test/respawn-team-awareness.test.ts` | ~6 | ~9 | Yes | Near-copy of respawn-controller.test.ts version. |
|
||||
| `test/session-manager.test.ts` | 4 | 8 | Yes | **Inside `vi.mock()` factory** — replaces `../src/session.js` module. Different shape: `start()`/`stop()`/`toState()`/`sendInput()` for lifecycle testing. |
|
||||
|
||||
`MockStateStore` is defined **2 times** (both inside `vi.mock()` factories):
|
||||
|
||||
| File | Shape | Methods | Mock Pattern |
|
||||
|------|-------|---------|-------------|
|
||||
| `test/session-manager.test.ts` | `{ sessions, config }` | `getConfig`, `getSessions`, `getSession`, `setSession`, `removeSession` | `vi.mock('../src/state-store.js')` |
|
||||
| `test/ralph-loop.test.ts` | `{ ralphLoop, tasks, config }` | `getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask` | `vi.mock('../src/state-store.js')` |
|
||||
|
||||
### Important: Two distinct mocking patterns
|
||||
|
||||
The codebase uses two different mocking patterns that require different migration strategies:
|
||||
|
||||
1. **Direct instantiation** (respawn-controller, respawn-team-awareness): `MockSession` is defined at file scope and instantiated directly in tests. These can be migrated to shared mocks via simple import replacement.
|
||||
|
||||
2. **Module replacement** (session-manager, ralph-loop): Mocks are defined inside `vi.mock()` factories that replace entire modules (`../src/session.js`, `../src/state-store.js`). These factories run in an isolated scope and return `{ Session: MockClass }` or `{ getStore: vi.fn(() => instance) }`. Migrating these requires either `vi.hoisted()` or restructuring the test's module mocking — higher risk for limited benefit.
|
||||
|
||||
### Unused Test Utilities
|
||||
|
||||
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
|
||||
|
||||
- `TimeController` / `createTimeController()` — abstraction over vitest fake timers
|
||||
- `MockAiIdleChecker` / `MockAiPlanChecker` — fully mocked AI checkers with result queueing
|
||||
- `createStateTracker()` / `createEventRecorder()` — state transition and event recording
|
||||
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — pre-configured RespawnConfig objects
|
||||
- `waitForState()` / `waitForEvent()` / `createDeferred()` — async test helpers
|
||||
- `terminalOutputs` — factory object for common terminal output patterns
|
||||
|
||||
### Route Test Coverage
|
||||
|
||||
Currently **zero** dedicated tests for the 12 route modules in `src/web/routes/`. The existing test files that touch API endpoints:
|
||||
|
||||
| Test File | What It Tests | Approach |
|
||||
|-----------|--------------|----------|
|
||||
| `test/api-responses.test.ts` | Response structure validation | Imports types, no HTTP calls |
|
||||
| `test/api-generate-plan.test.ts` | Plan generation API | Mocks validation logic, Port 3191 declared |
|
||||
| `test/auth-security.test.ts` | Auth middleware | Integration tests with WebServer, Ports 3160/3161 |
|
||||
| `test/qr-auth.test.ts` | QR authentication | Integration + unit tests, Port 3162 |
|
||||
|
||||
None of these test the route handlers themselves with real HTTP requests against a running Fastify instance.
|
||||
|
||||
---
|
||||
|
||||
## Design Decisions
|
||||
|
||||
### Shared mocks: Superset strategy
|
||||
|
||||
Rather than creating a lowest-common-denominator mock, `MockSession` in `test/mocks/` will be the **superset** from `respawn-test-utils.ts` (the most complete version). Test files that need a simpler mock can just ignore the extra methods — having unused methods costs nothing, but missing methods forces local re-definition.
|
||||
|
||||
### vi.mock() tests: Don't migrate
|
||||
|
||||
The `session-manager.test.ts` and `ralph-loop.test.ts` tests define mocks inside `vi.mock()` factories. These use **module-level replacement** (replacing `../src/session.js` and `../src/state-store.js` entirely), which is fundamentally different from the direct-instantiation pattern. Migrating them would require `vi.hoisted()` or factory restructuring — high complexity for limited benefit since these mocks are already working. We leave these as-is and create the shared mocks for **new** tests and for the two direct-instantiation tests (Tasks 4–5).
|
||||
|
||||
### MockStateStore: Union of both shapes
|
||||
|
||||
The shared `MockStateStore` in `test/mocks/` will include methods from both existing definitions (session management + Ralph loop), so any **new** test can use it. Methods default to no-ops via `vi.fn()`. Existing `vi.mock()`-based tests are not migrated.
|
||||
|
||||
### Route testing strategy: Lightweight Fastify instances
|
||||
|
||||
Each route test file will:
|
||||
1. Create a minimal `Fastify` instance
|
||||
2. Register **only** the route module under test
|
||||
3. Provide a mock context object satisfying the port interfaces
|
||||
4. Use `app.inject()` (Fastify's built-in test helper) — no real HTTP, no port needed
|
||||
|
||||
This avoids port conflicts entirely and runs fast. Only tests that need SSE or WebSocket behavior will use a real listening server with assigned ports.
|
||||
|
||||
### Port assignments (for tests needing real servers)
|
||||
|
||||
| Port | Test File | Purpose |
|
||||
|------|-----------|---------|
|
||||
| 3220 | `test/routes/session-routes.test.ts` | SSE integration (if needed) |
|
||||
| 3221 | `test/routes/system-routes.test.ts` | Status/stats endpoints |
|
||||
| 3222 | `test/routes/respawn-routes.test.ts` | Respawn API |
|
||||
| 3223 | `test/routes/ralph-routes.test.ts` | Ralph API |
|
||||
| 3224–3229 | Reserved | Future route tests |
|
||||
|
||||
Most tests should NOT need real ports — `app.inject()` is preferred. Verified: ports 3220–3229 are completely unused by existing tests (highest used port is 3211 in `opencode-resize.test.ts`).
|
||||
|
||||
---
|
||||
|
||||
## Task Dependencies
|
||||
|
||||
```
|
||||
Task 1 (Consolidate MockSession)
|
||||
Task 2 (Consolidate MockStateStore)
|
||||
└──> Task 3 (Create test/mocks/ barrel)
|
||||
├──> Task 4 (Migrate respawn-controller.test.ts)
|
||||
├──> Task 5 (Migrate respawn-team-awareness.test.ts)
|
||||
└──> Task 6 (Route test scaffold + helpers)
|
||||
├──> Task 7 (Session routes tests)
|
||||
└──> Task 8 (System + respawn routes tests)
|
||||
|
||||
Task 9 (Slim down respawn-test-utils.ts) — depends on Tasks 4, 5
|
||||
```
|
||||
|
||||
**Tasks 1–2** are independent and can run in parallel.
|
||||
**Task 3** depends on Tasks 1–2.
|
||||
**Tasks 4–6** depend on Task 3 and can run in parallel.
|
||||
**Tasks 7–8** depend on Task 6 and can run in parallel.
|
||||
**Task 9** depends on Tasks 4, 5 (must verify migrations work before removing duplicates from source).
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Consolidate MockSession into `test/mocks/mock-session.ts`
|
||||
|
||||
**Estimated effort**: 2 hours
|
||||
**Files created**: `test/mocks/mock-session.ts`
|
||||
**Files modified**: None yet (consumers migrate in Tasks 4–5)
|
||||
|
||||
### Source
|
||||
|
||||
The canonical MockSession comes from `test/respawn-test-utils.ts` (lines 89–241). It is the most complete version with:
|
||||
|
||||
- All properties needed by `RespawnController`: `id`, `workingDir`, `status`, `writeBuffer`, `terminalBuffer`, `muxName`
|
||||
- `write()` / `writeViaMux()` for input simulation
|
||||
- Buffer inspection: `lastWrite`, `hasWritten(pattern)`, `clearWriteBuffer()`
|
||||
- Terminal simulation: `simulateTerminalOutput()`, `simulatePrompt()`, `simulateReady()`, `simulateCompletionMessage()`, `simulateWorking()`, `simulateClearComplete()`, `simulateInitComplete()`, `simulatePlanModePrompt()`, `simulateElicitationDialog()`, `simulateTokenCount()`, `simulateAnsiOutput()`
|
||||
- Lifecycle: `close()`
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `test/mocks/` directory.
|
||||
2. Create `test/mocks/mock-session.ts`:
|
||||
- Copy the `MockSession` class **exactly** from `test/respawn-test-utils.ts` (lines 89–241)
|
||||
- Copy `terminalOutputs` helper object (tightly coupled to mock)
|
||||
- Copy `createMockSession()` factory function
|
||||
- Export all three: `export { MockSession, createMockSession, terminalOutputs }`
|
||||
- Ensure all `vi` imports come from `vitest`
|
||||
|
||||
**CRITICAL**: Copy the source verbatim — do NOT rewrite the simulation methods. The respawn controller's detection logic matches specific output patterns (e.g., `'\u276f '` for prompt, `'\u273b Worked for'` for completion). Using different patterns would cause test failures.
|
||||
|
||||
### Template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared MockSession for tests that need terminal simulation.
|
||||
*
|
||||
* Copied from test/respawn-test-utils.ts (the canonical, most complete version).
|
||||
* Used by respawn, route, and subagent tests.
|
||||
*/
|
||||
import { EventEmitter } from 'node:events';
|
||||
|
||||
// Copy MockSession class exactly from test/respawn-test-utils.ts lines 89–241
|
||||
export class MockSession extends EventEmitter {
|
||||
// ... (copy verbatim from respawn-test-utils.ts)
|
||||
}
|
||||
|
||||
/**
|
||||
* Factory for common terminal output strings.
|
||||
* Must match the patterns used in MockSession's simulate* methods.
|
||||
*/
|
||||
export const terminalOutputs = {
|
||||
// ... (copy verbatim from respawn-test-utils.ts)
|
||||
};
|
||||
|
||||
/**
|
||||
* Convenience factory.
|
||||
*/
|
||||
export function createMockSession(id?: string): MockSession {
|
||||
return new MockSession(id);
|
||||
}
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit # Ensure file compiles
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Consolidate MockStateStore into `test/mocks/mock-state-store.ts`
|
||||
|
||||
**Estimated effort**: 1 hour
|
||||
**Files created**: `test/mocks/mock-state-store.ts`
|
||||
**Files modified**: None (existing vi.mock()-based tests are NOT migrated; this is for new route tests)
|
||||
|
||||
### Source
|
||||
|
||||
Union of both existing definitions:
|
||||
|
||||
- From `test/session-manager.test.ts`: session CRUD methods (`getConfig`, `getSession`, `setSession`, `removeSession`, `getSessions`)
|
||||
- From `test/ralph-loop.test.ts`: Ralph state methods (`getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask`)
|
||||
|
||||
### Template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared MockStateStore for tests.
|
||||
*
|
||||
* Includes methods for both session management and Ralph loop testing.
|
||||
* All methods are vi.fn() spies — tests can override return values as needed.
|
||||
*
|
||||
* NOTE: This is for direct instantiation in new tests. Existing tests that
|
||||
* use vi.mock('../src/state-store.js') keep their inline definitions.
|
||||
*/
|
||||
import { vi } from 'vitest';
|
||||
|
||||
export class MockStateStore {
|
||||
state: Record<string, unknown> = {
|
||||
sessions: {} as Record<string, unknown>,
|
||||
config: { maxConcurrentSessions: 5 },
|
||||
ralphLoop: { status: 'stopped' },
|
||||
tasks: {} as Record<string, unknown>,
|
||||
};
|
||||
|
||||
// Session methods
|
||||
getConfig = vi.fn(() => this.state.config);
|
||||
getSessions = vi.fn(() => this.state.sessions as Record<string, unknown>);
|
||||
getSession = vi.fn((id: string) => (this.state.sessions as Record<string, unknown>)[id]);
|
||||
setSession = vi.fn((id: string, state: unknown) => {
|
||||
(this.state.sessions as Record<string, unknown>)[id] = state;
|
||||
});
|
||||
removeSession = vi.fn((id: string) => {
|
||||
delete (this.state.sessions as Record<string, unknown>)[id];
|
||||
});
|
||||
|
||||
// Ralph state methods
|
||||
getRalphLoopState = vi.fn(() => this.state.ralphLoop);
|
||||
setRalphLoopState = vi.fn((update: Record<string, unknown>) => {
|
||||
this.state.ralphLoop = { ...(this.state.ralphLoop as Record<string, unknown>), ...update };
|
||||
});
|
||||
|
||||
// Task methods
|
||||
getTasks = vi.fn(() => this.state.tasks);
|
||||
setTask = vi.fn();
|
||||
removeTask = vi.fn();
|
||||
|
||||
// Settings methods
|
||||
getSettings = vi.fn(() => ({}));
|
||||
setSettings = vi.fn();
|
||||
|
||||
// Generic persistence
|
||||
save = vi.fn();
|
||||
load = vi.fn();
|
||||
|
||||
/** Reset all state and mocks for clean test isolation */
|
||||
reset(): void {
|
||||
this.state = {
|
||||
sessions: {},
|
||||
config: { maxConcurrentSessions: 5 },
|
||||
ralphLoop: { status: 'stopped' },
|
||||
tasks: {},
|
||||
};
|
||||
vi.clearAllMocks();
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Create `test/mocks/index.ts` barrel export
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Tasks 1, 2
|
||||
**Files created**: `test/mocks/index.ts`, `test/mocks/test-helpers.ts`
|
||||
**Files modified**: None
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `test/mocks/test-helpers.ts` with the async utilities from `respawn-test-utils.ts`:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Reusable async test helpers.
|
||||
* Extracted from respawn-test-utils.ts.
|
||||
*/
|
||||
|
||||
/** Wait for an EventEmitter to emit a specific event, with timeout */
|
||||
export function waitForEvent(
|
||||
emitter: { once: (event: string, listener: (...args: unknown[]) => void) => void },
|
||||
event: string,
|
||||
timeoutMs = 5000,
|
||||
): Promise<unknown> {
|
||||
return new Promise((resolve, reject) => {
|
||||
const timer = setTimeout(
|
||||
() => reject(new Error(`Timed out waiting for event "${event}" after ${timeoutMs}ms`)),
|
||||
timeoutMs,
|
||||
);
|
||||
emitter.once(event, (...args: unknown[]) => {
|
||||
clearTimeout(timer);
|
||||
resolve(args.length === 1 ? args[0] : args);
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
/** Create a deferred promise with external resolve/reject */
|
||||
export function createDeferred<T = void>(): {
|
||||
promise: Promise<T>;
|
||||
resolve: (value: T) => void;
|
||||
reject: (reason?: unknown) => void;
|
||||
} {
|
||||
let resolve!: (value: T) => void;
|
||||
let reject!: (reason?: unknown) => void;
|
||||
const promise = new Promise<T>((res, rej) => {
|
||||
resolve = res;
|
||||
reject = rej;
|
||||
});
|
||||
return { promise, resolve, reject };
|
||||
}
|
||||
```
|
||||
|
||||
2. Create `test/mocks/index.ts` barrel:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared test mocks — import from here instead of defining inline.
|
||||
*
|
||||
* @example
|
||||
* import { MockSession, MockStateStore, terminalOutputs } from './mocks/index.js';
|
||||
*/
|
||||
|
||||
export { MockSession, createMockSession, terminalOutputs } from './mock-session.js';
|
||||
export { MockStateStore } from './mock-state-store.js';
|
||||
export { waitForEvent, createDeferred } from './test-helpers.js';
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Migrate `respawn-controller.test.ts` to shared mocks
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Task 3
|
||||
**Files modified**: `test/respawn-controller.test.ts`
|
||||
|
||||
### Steps
|
||||
|
||||
1. Remove the local `MockSession` class definition (approx. 50 lines).
|
||||
2. Add: `import { MockSession } from './mocks/index.js';`
|
||||
3. Verify all test methods still exist on the shared mock. The shared mock is a superset, so all existing usage should work.
|
||||
4. If the local mock had any test-specific customizations (e.g., extra properties added in `beforeEach`), keep those in the test file as inline assignments on the shared instance.
|
||||
5. Run the test to confirm it passes.
|
||||
|
||||
### Potential issues
|
||||
|
||||
- The local mock's `simulateCompletionMessage()` may have a slightly different output format than the shared mock's (from respawn-test-utils.ts). Verify the respawn controller's completion detection regex matches the shared mock's output pattern (`'\u273b Worked for ...'`).
|
||||
- If the local mock adds `pid` or `isWorking` properties that the shared mock doesn't have, add inline assignments in `beforeEach`.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/respawn-controller.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Migrate `respawn-team-awareness.test.ts` to shared mocks
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Task 3
|
||||
**Files modified**: `test/respawn-team-awareness.test.ts`
|
||||
|
||||
### Steps
|
||||
|
||||
1. Remove the local `MockSession` class definition.
|
||||
2. Add: `import { MockSession } from './mocks/index.js';`
|
||||
3. Keep `MockTeamWatcher` in this file — it's test-specific and extends the real `TeamWatcher`, not a general-purpose mock.
|
||||
4. Run the test to confirm it passes.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/respawn-team-awareness.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 6: Create route test scaffold and helpers
|
||||
|
||||
**Estimated effort**: 2 hours
|
||||
**Depends on**: Task 3
|
||||
**Files created**: `test/mocks/mock-route-context.ts`, `test/routes/` directory, `test/routes/_route-test-utils.ts`
|
||||
|
||||
### Problem
|
||||
|
||||
The 12 route modules in `src/web/routes/` have zero dedicated test coverage. Each route module takes `(app: FastifyInstance, ctx: PortIntersection)` — we need a reusable way to create mock context objects that satisfy the port interfaces.
|
||||
|
||||
### Design
|
||||
|
||||
Create a `MockRouteContext` factory that builds a mock object satisfying all port interfaces. Each port's methods are `vi.fn()` stubs. Tests can override specific methods as needed.
|
||||
|
||||
### Route registration signatures (verified)
|
||||
|
||||
Each route module requires a specific port intersection. The mock must satisfy all of them:
|
||||
|
||||
| Route Module | Required Ports |
|
||||
|-------------|----------------|
|
||||
| `registerSessionRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
|
||||
| `registerSystemRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
|
||||
| `registerRespawnRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
|
||||
| `registerRalphRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
|
||||
| `registerPlanRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort` |
|
||||
| `registerCaseRoutes` | `EventPort & ConfigPort` |
|
||||
| `registerScheduledRoutes` | `SessionPort & EventPort & InfraPort` |
|
||||
| `registerFileRoutes` | `SessionPort` |
|
||||
| `registerMuxRoutes` | `InfraPort` |
|
||||
| `registerPushRoutes` | `InfraPort` |
|
||||
| `registerTeamRoutes` | `InfraPort` |
|
||||
| `registerHookEventRoutes` | `EventPort & AuthPort` |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `test/mocks/mock-route-context.ts`:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Mock context for route handler testing.
|
||||
*
|
||||
* Satisfies ALL port interfaces (SessionPort, EventPort, RespawnPort,
|
||||
* ConfigPort, InfraPort, AuthPort) so any route module can be tested.
|
||||
* Override specific methods in individual tests as needed.
|
||||
*
|
||||
* Verified against actual port interfaces in src/web/ports/:
|
||||
* - SessionPort: 6 methods (sessions, addSession, cleanupSession,
|
||||
* setupSessionListeners, persistSessionState, persistSessionStateNow,
|
||||
* getSessionStateWithRespawn)
|
||||
* - EventPort: 5 methods (broadcast, sendPushNotifications, batchTerminalData,
|
||||
* broadcastSessionStateDebounced, batchTaskUpdate)
|
||||
* - RespawnPort: 2 maps + 4 methods
|
||||
* - ConfigPort: 5 readonly + 7 methods (incl getDefaultClaudeMdPath,
|
||||
* getLightState, getLightSessionsState, stopTranscriptWatcher)
|
||||
* - InfraPort: 7 readonly + 2 methods (startScheduledRun, stopScheduledRun)
|
||||
* - AuthPort: 3 readonly (authSessions, qrAuthFailures, https)
|
||||
*/
|
||||
import { vi } from 'vitest';
|
||||
import { MockSession, createMockSession } from './mock-session.js';
|
||||
|
||||
/**
|
||||
* Creates a mock context that satisfies all port interfaces.
|
||||
* Pre-populated with one session for convenience.
|
||||
*/
|
||||
export function createMockRouteContext(options?: { sessionId?: string }) {
|
||||
const sessionId = options?.sessionId ?? 'test-session-1';
|
||||
const session = createMockSession(sessionId);
|
||||
const sessions = new Map<string, MockSession>();
|
||||
sessions.set(sessionId, session);
|
||||
|
||||
return {
|
||||
// -- SessionPort --
|
||||
sessions,
|
||||
addSession: vi.fn(),
|
||||
cleanupSession: vi.fn(),
|
||||
setupSessionListeners: vi.fn(),
|
||||
persistSessionState: vi.fn(),
|
||||
persistSessionStateNow: vi.fn(),
|
||||
getSessionStateWithRespawn: vi.fn((s: unknown) => s),
|
||||
|
||||
// -- EventPort --
|
||||
broadcast: vi.fn(),
|
||||
sendPushNotifications: vi.fn(),
|
||||
batchTerminalData: vi.fn(),
|
||||
broadcastSessionStateDebounced: vi.fn(),
|
||||
batchTaskUpdate: vi.fn(),
|
||||
|
||||
// -- RespawnPort --
|
||||
respawnControllers: new Map(),
|
||||
respawnTimers: new Map(),
|
||||
setupRespawnListeners: vi.fn(),
|
||||
setupTimedRespawn: vi.fn(),
|
||||
restoreRespawnController: vi.fn(),
|
||||
saveRespawnConfig: vi.fn(),
|
||||
|
||||
// -- ConfigPort --
|
||||
store: {
|
||||
getConfig: vi.fn(() => ({})),
|
||||
getSessions: vi.fn(() => ({})),
|
||||
getSession: vi.fn(),
|
||||
setSession: vi.fn(),
|
||||
removeSession: vi.fn(),
|
||||
getSettings: vi.fn(() => ({})),
|
||||
setSettings: vi.fn(),
|
||||
getRalphLoopState: vi.fn(() => ({})),
|
||||
setRalphLoopState: vi.fn(),
|
||||
getTasks: vi.fn(() => ({})),
|
||||
save: vi.fn(),
|
||||
load: vi.fn(),
|
||||
},
|
||||
port: 3000,
|
||||
https: false,
|
||||
testMode: true,
|
||||
serverStartTime: Date.now(),
|
||||
getGlobalNiceConfig: vi.fn(async () => undefined),
|
||||
getModelConfig: vi.fn(async () => null),
|
||||
getClaudeModeConfig: vi.fn(async () => ({})),
|
||||
getDefaultClaudeMdPath: vi.fn(async () => undefined),
|
||||
getLightState: vi.fn(() => ({ sessions: [], status: 'ok' })),
|
||||
getLightSessionsState: vi.fn(() => []),
|
||||
startTranscriptWatcher: vi.fn(),
|
||||
stopTranscriptWatcher: vi.fn(),
|
||||
|
||||
// -- InfraPort --
|
||||
mux: {
|
||||
createSession: vi.fn(),
|
||||
killSession: vi.fn(),
|
||||
listSessions: vi.fn(() => []),
|
||||
getStats: vi.fn(() => ({})),
|
||||
},
|
||||
runSummaryTrackers: new Map(),
|
||||
activePlanOrchestrators: new Map(),
|
||||
scheduledRuns: new Map(),
|
||||
teamWatcher: { getTeams: vi.fn(() => []), hasActiveTeammates: vi.fn(() => false) },
|
||||
tunnelManager: null,
|
||||
pushStore: null,
|
||||
startScheduledRun: vi.fn(),
|
||||
stopScheduledRun: vi.fn(),
|
||||
|
||||
// -- AuthPort --
|
||||
authSessions: null,
|
||||
qrAuthFailures: null,
|
||||
// https already declared above in ConfigPort (shared property)
|
||||
|
||||
// Convenience accessors (not part of any port interface)
|
||||
_session: session,
|
||||
_sessionId: sessionId,
|
||||
};
|
||||
}
|
||||
|
||||
export type MockRouteContext = ReturnType<typeof createMockRouteContext>;
|
||||
```
|
||||
|
||||
2. Add to `test/mocks/index.ts` barrel:
|
||||
|
||||
```typescript
|
||||
export { createMockRouteContext, type MockRouteContext } from './mock-route-context.js';
|
||||
```
|
||||
|
||||
3. Create `test/routes/` directory for route test files.
|
||||
|
||||
4. Create `test/routes/_route-test-utils.ts` with Fastify test helpers:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared utilities for route testing.
|
||||
*
|
||||
* Creates minimal Fastify instances with just the route module under test
|
||||
* and a mock context. Uses app.inject() for HTTP testing without real ports.
|
||||
*/
|
||||
import Fastify, { type FastifyInstance } from 'fastify';
|
||||
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
|
||||
|
||||
export interface RouteTestHarness {
|
||||
app: FastifyInstance;
|
||||
ctx: MockRouteContext;
|
||||
}
|
||||
|
||||
/**
|
||||
* Creates a Fastify instance with a route module registered against a mock context.
|
||||
*
|
||||
* @param registerFn - The route registration function (e.g., registerSessionRoutes).
|
||||
* Uses `any` for ctx parameter because route functions expect typed port intersections
|
||||
* that MockRouteContext satisfies structurally but not nominally.
|
||||
* @param ctxOptions - Optional overrides for the mock context
|
||||
*/
|
||||
export async function createRouteTestHarness(
|
||||
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
||||
registerFn: (app: FastifyInstance, ctx: any) => void,
|
||||
ctxOptions?: { sessionId?: string },
|
||||
): Promise<RouteTestHarness> {
|
||||
const app = Fastify({ logger: false });
|
||||
const ctx = createMockRouteContext(ctxOptions);
|
||||
|
||||
registerFn(app, ctx);
|
||||
await app.ready();
|
||||
|
||||
return { app, ctx };
|
||||
}
|
||||
```
|
||||
|
||||
### Why `ctx: any` in the harness
|
||||
|
||||
Route registration functions like `registerSessionRoutes(app, ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort)` expect specific port intersection types. TypeScript won't accept `unknown` here because it's not assignable to the port types. The `MockRouteContext` satisfies the interfaces structurally (it has all the required properties and methods), but since it's not declared as implementing them, we need `any` at the call site. This is the standard pattern for test mocks in TypeScript.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 7: Add session routes tests
|
||||
|
||||
**Estimated effort**: 4 hours
|
||||
**Depends on**: Task 6
|
||||
**Files created**: `test/routes/session-routes.test.ts`
|
||||
**Port**: 3220 (only if SSE tests needed; prefer `app.inject()`)
|
||||
|
||||
### Coverage targets
|
||||
|
||||
`src/web/routes/session-routes.ts` is the largest route module (43 handlers). Focus on the most critical endpoints first:
|
||||
|
||||
#### Priority 1: Session CRUD (must test)
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `GET` | `/api/sessions` | Returns session list; empty when no sessions |
|
||||
| `GET` | `/api/sessions/:id` | Returns session state; 404 for unknown ID |
|
||||
| `POST` | `/api/sessions` | Creates session; validates workingDir; rejects invalid paths |
|
||||
| `DELETE` | `/api/sessions/:id` | Calls cleanupSession; 404 for unknown ID |
|
||||
|
||||
#### Priority 2: Session I/O
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `POST` | `/api/sessions/:id/input` | Sends input to session; validates input length; 404 for unknown |
|
||||
| `POST` | `/api/sessions/:id/resize` | Validates cols/rows bounds; 404 for unknown |
|
||||
| `GET` | `/api/sessions/:id/buffer` | Returns terminal buffer; 404 for unknown |
|
||||
|
||||
#### Priority 3: Session actions
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `POST` | `/api/sessions/:id/run` | Runs prompt on session |
|
||||
| `POST` | `/api/sessions/:id/clear` | Clears session |
|
||||
| `POST` | `/api/sessions/:id/compact` | Compacts session |
|
||||
| `POST` | `/api/sessions/:id/interactive` | Starts interactive mode |
|
||||
| `POST` | `/api/sessions/:id/quick-start` | Quick start flow |
|
||||
|
||||
### Test pattern
|
||||
|
||||
```typescript
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
|
||||
import { createRouteTestHarness, type RouteTestHarness } from './_route-test-utils.js';
|
||||
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
|
||||
|
||||
describe('session-routes', () => {
|
||||
let harness: RouteTestHarness;
|
||||
|
||||
beforeEach(async () => {
|
||||
harness = await createRouteTestHarness(registerSessionRoutes);
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
await harness.app.close();
|
||||
});
|
||||
|
||||
describe('GET /api/sessions', () => {
|
||||
it('returns empty array when no sessions', async () => {
|
||||
harness.ctx.sessions.clear();
|
||||
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(JSON.parse(res.body)).toEqual([]);
|
||||
});
|
||||
|
||||
it('returns session list with one session', async () => {
|
||||
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
|
||||
expect(res.statusCode).toBe(200);
|
||||
const sessions = JSON.parse(res.body);
|
||||
expect(sessions).toHaveLength(1);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id', () => {
|
||||
it('returns 404 for unknown session', async () => {
|
||||
const res = await harness.app.inject({
|
||||
method: 'GET',
|
||||
url: '/api/sessions/nonexistent',
|
||||
});
|
||||
expect(res.statusCode).toBe(404);
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/input', () => {
|
||||
it('rejects input exceeding max length', async () => {
|
||||
const res = await harness.app.inject({
|
||||
method: 'POST',
|
||||
url: `/api/sessions/${harness.ctx._sessionId}/input`,
|
||||
payload: { input: 'x'.repeat(65537) },
|
||||
});
|
||||
expect(res.statusCode).toBe(400);
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/resize', () => {
|
||||
it('rejects cols exceeding max', async () => {
|
||||
const res = await harness.app.inject({
|
||||
method: 'POST',
|
||||
url: `/api/sessions/${harness.ctx._sessionId}/resize`,
|
||||
payload: { cols: 501, rows: 24 },
|
||||
});
|
||||
expect(res.statusCode).toBe(400);
|
||||
});
|
||||
});
|
||||
});
|
||||
```
|
||||
|
||||
### Key assertions to include
|
||||
|
||||
- **404 for unknown sessions**: Every `:id` endpoint must return 404 for nonexistent IDs
|
||||
- **Input validation**: Bad paths, oversized inputs, invalid resize dimensions
|
||||
- **Side effects**: Verify `ctx.broadcast()` was called with correct event type after mutations
|
||||
- **Response shape**: Verify response bodies match expected API types
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/routes/session-routes.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 8: Add system + respawn routes tests
|
||||
|
||||
**Estimated effort**: 4 hours
|
||||
**Depends on**: Task 6
|
||||
**Files created**: `test/routes/system-routes.test.ts`, `test/routes/respawn-routes.test.ts`
|
||||
|
||||
### System routes (`src/web/routes/system-routes.ts`)
|
||||
|
||||
Focus on status and configuration endpoints:
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `GET` | `/api/status` | Returns server status with uptime, session count |
|
||||
| `GET` | `/api/stats` | Returns mux stats |
|
||||
| `GET` | `/api/config` | Returns current config |
|
||||
| `PUT` | `/api/config` | Updates config; validates input |
|
||||
| `GET` | `/api/settings` | Returns user settings |
|
||||
| `PUT` | `/api/settings` | Updates settings; validates input |
|
||||
| `GET` | `/api/subagents` | Returns subagent list |
|
||||
| `GET` | `/api/screenshots` | Returns screenshot list |
|
||||
|
||||
### Respawn routes (`src/web/routes/respawn-routes.ts`)
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `GET` | `/api/sessions/:id/respawn` | Returns respawn status; null when not configured |
|
||||
| `POST` | `/api/sessions/:id/respawn/start` | Starts respawn; 404 for unknown session |
|
||||
| `POST` | `/api/sessions/:id/respawn/stop` | Stops respawn; 404 for unknown session |
|
||||
| `PUT` | `/api/sessions/:id/respawn/config` | Updates respawn config; validates |
|
||||
| `POST` | `/api/sessions/:id/respawn/enable` | Enables respawn loop |
|
||||
| `POST` | `/api/sessions/:id/respawn/disable` | Disables respawn loop |
|
||||
|
||||
### Test patterns
|
||||
|
||||
Same pattern as Task 7 — `createRouteTestHarness` with `registerSystemRoutes` / `registerRespawnRoutes`.
|
||||
|
||||
For respawn tests, pre-populate `ctx.respawnControllers` with a mock controller in `beforeEach`:
|
||||
|
||||
```typescript
|
||||
beforeEach(async () => {
|
||||
harness = await createRouteTestHarness(registerRespawnRoutes);
|
||||
// Add a mock respawn controller for the default session
|
||||
harness.ctx.respawnControllers.set(harness.ctx._sessionId, {
|
||||
getState: vi.fn(() => 'idle'),
|
||||
getConfig: vi.fn(() => ({})),
|
||||
getStatus: vi.fn(() => ({ state: 'idle', health: 100 })),
|
||||
start: vi.fn(),
|
||||
stop: vi.fn(),
|
||||
updateConfig: vi.fn(),
|
||||
enable: vi.fn(),
|
||||
disable: vi.fn(),
|
||||
});
|
||||
});
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/routes/system-routes.test.ts
|
||||
npx vitest run test/routes/respawn-routes.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 9: Slim down `respawn-test-utils.ts`
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Tasks 4, 5
|
||||
**Files modified**: `test/respawn-test-utils.ts`
|
||||
|
||||
After Tasks 4–5 are verified passing with shared mocks, slim down `respawn-test-utils.ts` to remove duplicates.
|
||||
|
||||
### Steps
|
||||
|
||||
1. **Remove** from `respawn-test-utils.ts` what has been moved to shared mocks:
|
||||
- `MockSession` class → now in `test/mocks/mock-session.ts`
|
||||
- `createMockSession()` → now in `test/mocks/mock-session.ts`
|
||||
- `terminalOutputs` → now in `test/mocks/mock-session.ts`
|
||||
- `waitForEvent()` / `createDeferred()` → now in `test/mocks/test-helpers.ts`
|
||||
|
||||
2. **Keep** respawn-specific utilities that don't belong in the general mocks:
|
||||
- `TimeController` / `createTimeController()` — respawn-specific timer control
|
||||
- `MockAiIdleChecker` / `MockAiPlanChecker` — respawn-specific AI mocks
|
||||
- `createStateTracker()` / `createEventRecorder()` — respawn state tracking
|
||||
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — respawn config presets
|
||||
- `waitForState()` — respawn state machine waiter
|
||||
|
||||
3. **Update imports** in `respawn-test-utils.ts` to re-use shared mocks:
|
||||
```typescript
|
||||
import { MockSession, createMockSession, terminalOutputs } from './mocks/index.js';
|
||||
import { waitForEvent, createDeferred } from './mocks/index.js';
|
||||
export { MockSession, createMockSession, terminalOutputs, waitForEvent, createDeferred };
|
||||
```
|
||||
|
||||
This preserves backward compatibility for any future tests that import from `respawn-test-utils.ts` directly while eliminating the duplication.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npx vitest run test/respawn-controller.test.ts
|
||||
npx vitest run test/respawn-team-awareness.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## What is NOT in scope (and why)
|
||||
|
||||
### Migrating `session-manager.test.ts` and `ralph-loop.test.ts` mocks
|
||||
|
||||
Both files define mocks inside `vi.mock()` factories that replace entire modules:
|
||||
|
||||
```typescript
|
||||
// session-manager.test.ts — mock replaces ../src/session.js
|
||||
vi.mock('../src/session.js', () => {
|
||||
class MockSession extends EventEmitter { ... }
|
||||
return { Session: MockSession };
|
||||
});
|
||||
|
||||
// ralph-loop.test.ts — mock replaces ../src/state-store.js
|
||||
vi.mock('../src/state-store.js', () => {
|
||||
class MockStateStore { ... }
|
||||
return { getStore: vi.fn(() => instance), StateStore: MockStateStore };
|
||||
});
|
||||
```
|
||||
|
||||
These are fundamentally different from the direct-instantiation pattern:
|
||||
- The `vi.mock()` factory runs in an isolated scope — outer imports are not available
|
||||
- The mock class must be returned with the exact export names (`Session`, `getStore`, `StateStore`)
|
||||
- The `session-manager.test.ts` MockSession auto-registers into a shared `mockState.sessions` Map (tight coupling with test setup)
|
||||
|
||||
Migrating would require `vi.hoisted()` to share the class between factory and test scope, plus restructuring the test's module-mocking setup. This is high-complexity, high-risk refactoring with limited benefit since these tests already work. The shared `MockStateStore` in `test/mocks/` is available for **new** tests (like route tests) that use direct instantiation instead.
|
||||
|
||||
### Full integration tests with real Fastify server
|
||||
|
||||
Route tests use `app.inject()` which simulates HTTP without opening ports. Full integration tests that spin up `WebServer`, create real sessions, and stream SSE would be valuable but are a separate effort requiring:
|
||||
- A test WebServer factory
|
||||
- Session lifecycle management in tests
|
||||
- SSE client test utilities
|
||||
- Significantly more setup/teardown complexity
|
||||
|
||||
### Testing auth middleware in route tests
|
||||
|
||||
Route tests bypass authentication (no auth middleware registered on the test Fastify instance). Auth middleware has its own dedicated tests in `auth-security.test.ts` and `qr-auth.test.ts`. Testing auth + routes together is a future integration test concern.
|
||||
|
||||
### Testing SSE event streaming
|
||||
|
||||
SSE integration requires a running server with `EventSource` client. This is significantly more complex than `app.inject()` tests and is deferred. The existing `sse-events.test.ts` covers SSE patterns.
|
||||
|
||||
### Complete route coverage for all 12 modules
|
||||
|
||||
This phase covers the 3 highest-value route modules (session, system, respawn — 98 of 162 handlers). The remaining 9 modules (ralph, plan, push, team, mux, file, scheduled, hook-event, case) should be added incrementally in follow-up work.
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
| Metric | Before | After |
|
||||
|--------|--------|-------|
|
||||
| MockSession definitions | 4 (across 4 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
|
||||
| MockStateStore definitions | 2 (across 2 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
|
||||
| Files importing from `respawn-test-utils.ts` | 0 | Utilities split into `test/mocks/` |
|
||||
| Route test files | 0 | 3 (session, system, respawn) |
|
||||
| Route handlers with dedicated tests | 0 | ~30 (highest-priority endpoints) |
|
||||
| Shared mock directory | None | `test/mocks/` with 5 files + barrel |
|
||||
|
||||
### Final verification checklist
|
||||
|
||||
```bash
|
||||
# Type checking
|
||||
tsc --noEmit
|
||||
|
||||
# Linting
|
||||
npm run lint
|
||||
|
||||
# Formatting
|
||||
npm run format:check
|
||||
|
||||
# Run all affected tests individually
|
||||
npx vitest run test/respawn-controller.test.ts
|
||||
npx vitest run test/respawn-team-awareness.test.ts
|
||||
npx vitest run test/routes/session-routes.test.ts
|
||||
npx vitest run test/routes/system-routes.test.ts
|
||||
npx vitest run test/routes/respawn-routes.test.ts
|
||||
|
||||
# Verify unchanged tests still pass
|
||||
npx vitest run test/session-manager.test.ts
|
||||
npx vitest run test/ralph-loop.test.ts
|
||||
|
||||
# Dev server still starts
|
||||
npx tsx src/index.ts web --port 3099 &
|
||||
curl -s http://localhost:3099/api/status | jq .status # "ok"
|
||||
kill %1
|
||||
```
|
||||
@@ -1,247 +0,0 @@
|
||||
# Ralph Loop Plan Improvement Roadmap
|
||||
|
||||
> Research-backed improvements for rock-solid AI planning with auto-improvement capabilities.
|
||||
|
||||
**Created**: 2026-01-27
|
||||
**Status**: Implementation in Progress
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Research Summary](#research-summary)
|
||||
2. [Current State Analysis](#current-state-analysis)
|
||||
3. [Proposed Improvements](#proposed-improvements)
|
||||
4. [Implementation Plan](#implementation-plan)
|
||||
5. [Sources](#sources)
|
||||
|
||||
---
|
||||
|
||||
## Research Summary
|
||||
|
||||
### Key Insights from Industry Best Practices
|
||||
|
||||
#### 1. Self-Verification is Critical
|
||||
> "Claude performs dramatically better when it can verify its own work, like run tests, compare screenshots, and validate outputs. Without clear success criteria, it might produce something that looks right but actually doesn't work."
|
||||
> — [Anthropic Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
|
||||
|
||||
#### 2. Iterative Refinement Patterns (AWS)
|
||||
> "A generator agent produces output, an evaluator agent reviews using evaluation rubric, and based on feedback, an optimizer agent revises the output. Loop repeats until criteria met."
|
||||
> — [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
|
||||
|
||||
#### 3. Dynamic Task Decomposition (TDAG Framework)
|
||||
> "Dynamically decomposes complex tasks into smaller subtasks and assigns each to a specifically generated subagent, enhancing adaptability in diverse and unpredictable real-world tasks."
|
||||
> — [TDAG Framework - arXiv](https://arxiv.org/abs/2402.10178)
|
||||
|
||||
#### 4. Multi-Stage Verification Workflow
|
||||
> "o3: Generate plan → Sonnet: Verify and create task list → Sonnet: Execute → Sonnet: Verify against plan → o3: Final verification → Issues bake back into plan"
|
||||
> — [Claude Code Best Practices Community](https://rosmur.github.io/claudecode-best-practices/)
|
||||
|
||||
#### 5. Self-Improving Agents
|
||||
> "Through an iterative refinement process (analyze outcome → adjust approach → try again), the agent becomes more adept at handling tasks over time. It effectively builds a growing knowledge base of what strategies work best."
|
||||
> — [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
|
||||
|
||||
#### 6. Memory Architecture for Planning
|
||||
> "Agents use three memory layers: working memory for short-lived calculations, episodic memory for step-by-step histories, and semantic memory for long-term knowledge."
|
||||
> — [LLM Agent Research](https://www.promptingguide.ai/research/llm-agents)
|
||||
|
||||
---
|
||||
|
||||
## Current State Analysis
|
||||
|
||||
### What We Have
|
||||
|
||||
The current plan generation system (`/api/generate-plan` and `/api/generate-plan-detailed`):
|
||||
|
||||
1. **Standard Mode**: Single Opus 4.5 call with TDD-focused prompt
|
||||
2. **Enhanced Mode**: 4 parallel subagents (Requirements, Architecture, Testing, Risks) + Verification
|
||||
|
||||
### Current Plan Item Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"content": "Implement login endpoint",
|
||||
"priority": "P0"
|
||||
}
|
||||
```
|
||||
|
||||
### Limitations
|
||||
|
||||
| Issue | Impact |
|
||||
|-------|--------|
|
||||
| No verification criteria | Can't automatically validate completion |
|
||||
| No test pairing | TDD not enforced structurally |
|
||||
| Static plans | No adaptation during execution |
|
||||
| No dependencies | Can't track blocking relationships |
|
||||
| No failure tracking | Same errors repeat |
|
||||
| No checkpoints | Plans run until completion or failure |
|
||||
|
||||
---
|
||||
|
||||
## Proposed Improvements
|
||||
|
||||
### Enhanced Plan Item Structure
|
||||
|
||||
```typescript
|
||||
interface EnhancedPlanItem {
|
||||
id: string; // Unique identifier (e.g., "P0-001")
|
||||
content: string; // Task description
|
||||
priority: 'P0' | 'P1' | 'P2'; // Criticality
|
||||
phase: 'setup' | 'test' | 'impl' | 'verify'; // Development phase
|
||||
|
||||
// NEW: Verification
|
||||
verificationCriteria: string; // How to know it's done
|
||||
testCommand?: string; // Command to run for verification
|
||||
|
||||
// NEW: Dependencies
|
||||
dependencies: string[]; // IDs of tasks that must complete first
|
||||
blockedBy?: string[]; // Runtime: tasks blocking this one
|
||||
|
||||
// NEW: Execution tracking
|
||||
status: 'pending' | 'in_progress' | 'completed' | 'failed' | 'blocked';
|
||||
attempts: number; // How many times attempted
|
||||
lastError?: string; // Most recent failure reason
|
||||
completedAt?: number; // Timestamp of completion
|
||||
|
||||
// NEW: Metadata
|
||||
estimatedComplexity: 'low' | 'medium' | 'high';
|
||||
rollbackStrategy?: string; // How to undo if needed
|
||||
version: number; // Plan version this belongs to
|
||||
}
|
||||
```
|
||||
|
||||
### Runtime Plan Adaptation Flow
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ RUNTIME PLAN LOOP │
|
||||
├─────────────────────────────────────────────────────────────────┤
|
||||
│ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ Execute │──▶│ Verify │──▶│ Success? │──▶│ Mark │ │
|
||||
│ │ Task │ │ Output │ │ │ │ Complete │ │
|
||||
│ └──────────┘ └──────────┘ └────┬─────┘ └──────────┘ │
|
||||
│ │ No │
|
||||
│ ▼ │
|
||||
│ ┌──────────┐ │
|
||||
│ │ Analyze │ │
|
||||
│ │ Failure │ │
|
||||
│ └────┬─────┘ │
|
||||
│ │ │
|
||||
│ ┌──────────────┼──────────────┐ │
|
||||
│ ▼ ▼ ▼ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ Retry │ │ Add Fix │ │ Escalate │ │
|
||||
│ │ (< 3x) │ │ Sub-Task │ │ BLOCKED │ │
|
||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||
│ │
|
||||
└───────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### Checkpoint Review System
|
||||
|
||||
At iterations 5, 10, 20, 30, 50:
|
||||
1. Pause execution
|
||||
2. Summarize progress (completed/failed/pending)
|
||||
3. Identify stuck items (3+ failures)
|
||||
4. Generate alternative approaches for stuck items
|
||||
5. Update plan with new strategies
|
||||
6. Continue with refined plan
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Phase 1: Quick Wins (Implementing Now)
|
||||
|
||||
#### 1.1 Add Verification Criteria to Plan Items
|
||||
- Modify plan generation prompts to require `verificationCriteria`
|
||||
- Update `PlanItem` interface in `types.ts`
|
||||
- Update plan orchestrator prompts
|
||||
|
||||
#### 1.2 Pair Test/Implementation Steps
|
||||
- Ensure every implementation step has a corresponding test step
|
||||
- Group items: test → implement → verify
|
||||
- Add phase field to track TDD cycle
|
||||
|
||||
#### 1.3 Checkpoint Review Prompts
|
||||
- Add checkpoint logic to Ralph tracker
|
||||
- At iterations 5, 10, 20: inject review prompt
|
||||
- Generate progress summary and stuck item analysis
|
||||
|
||||
### Phase 2: Medium Effort (Implementing Now)
|
||||
|
||||
#### 2.1 Failure Tracking
|
||||
- Track `attempts` and `lastError` per task
|
||||
- After 3 failures, auto-generate debug sub-task
|
||||
- Record failure patterns in plan history
|
||||
|
||||
#### 2.2 Plan Versioning
|
||||
- Add `version` field to plans
|
||||
- Keep history in `@fix_plan.md` with version markers
|
||||
- Allow rollback to previous versions
|
||||
- Track which version each task belongs to
|
||||
|
||||
#### 2.3 Dependency Tracking
|
||||
- Add `dependencies` field to plan items
|
||||
- Validate dependency graph (no cycles)
|
||||
- Block tasks until dependencies complete
|
||||
- Show dependency status in UI
|
||||
|
||||
### Phase 3: Future Enhancements
|
||||
|
||||
#### 3.1 Full Runtime Adaptation
|
||||
- TDAG-style dynamic decomposition
|
||||
- Auto-generate sub-tasks for complex items
|
||||
- Learning from failure patterns
|
||||
|
||||
#### 3.2 Multi-Model Verification
|
||||
- Haiku: Fast initial generation
|
||||
- Sonnet: Verification and refinement
|
||||
- Opus: Final quality check
|
||||
|
||||
#### 3.3 Plan Memory System
|
||||
- Episodic memory: What worked/failed in this session
|
||||
- Semantic memory: Patterns across projects
|
||||
- Use for future plan generation
|
||||
|
||||
---
|
||||
|
||||
## File Changes Required
|
||||
|
||||
### New/Modified Files
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/types.ts` | Add `EnhancedPlanItem` interface |
|
||||
| `src/plan-orchestrator.ts` | Update prompts, add versioning |
|
||||
| `src/ralph-tracker.ts` | Add checkpoint logic, failure tracking |
|
||||
| `src/web/server.ts` | New endpoints for plan updates |
|
||||
| `src/web/public/app.js` | UI for enhanced plan display |
|
||||
|
||||
### New Endpoints
|
||||
|
||||
| Method | Endpoint | Purpose |
|
||||
|--------|----------|---------|
|
||||
| PATCH | `/api/sessions/:id/plan/task/:taskId` | Update task status |
|
||||
| POST | `/api/sessions/:id/plan/checkpoint` | Trigger checkpoint review |
|
||||
| GET | `/api/sessions/:id/plan/history` | Get plan version history |
|
||||
| POST | `/api/sessions/:id/plan/rollback/:version` | Rollback to version |
|
||||
|
||||
---
|
||||
|
||||
## Sources
|
||||
|
||||
- [Anthropic Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
|
||||
- [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
|
||||
- [TDAG: Multi-Agent Task Decomposition Framework](https://arxiv.org/abs/2402.10178)
|
||||
- [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
|
||||
- [OpenAI Self-Evolving Agents Cookbook](https://cookbook.openai.com/examples/partners/self_evolving_agents/autonomous_agent_retraining)
|
||||
- [Task Decomposition for Coding Agents](https://mgx.dev/insights/task-decomposition-for-coding-agents-architectures-advancements-and-future-directions/)
|
||||
- [Claude Code Best Practices Community Guide](https://rosmur.github.io/claudecode-best-practices/)
|
||||
- [LLM Agents Prompt Engineering Guide](https://www.promptingguide.ai/research/llm-agents)
|
||||
- [Agentic AI Implementation Guide](https://www.sketchdev.io/blog/agentic-ai-implementation-guide)
|
||||
|
||||
---
|
||||
|
||||
*This document is part of the Codeman project. See [CLAUDE.md](../CLAUDE.md) for main documentation.*
|
||||
@@ -1,251 +0,0 @@
|
||||
# Ralph Loop Improvements Plan
|
||||
|
||||
## Overview
|
||||
|
||||
This plan details improvements to Codeman's Ralph Loop system based on best practices from the Ralph Claude Code repository (https://github.com/frankbria/ralph-claude-code).
|
||||
|
||||
## Key Concepts to Implement
|
||||
|
||||
### RALPH_STATUS Block Format
|
||||
|
||||
Claude outputs this structured block at the end of every response for better tracking:
|
||||
|
||||
```
|
||||
---RALPH_STATUS---
|
||||
STATUS: IN_PROGRESS | COMPLETE | BLOCKED
|
||||
TASKS_COMPLETED_THIS_LOOP: <number>
|
||||
FILES_MODIFIED: <number>
|
||||
TESTS_STATUS: PASSING | FAILING | NOT_RUN
|
||||
WORK_TYPE: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING
|
||||
EXIT_SIGNAL: false | true
|
||||
RECOMMENDATION: <one line summary of what to do next>
|
||||
---END_RALPH_STATUS---
|
||||
```
|
||||
|
||||
### Dual-Condition Exit Gate
|
||||
|
||||
Exit requires BOTH conditions:
|
||||
1. `completion_indicators >= 2` (heuristic detection from natural language patterns)
|
||||
2. Claude's explicit `EXIT_SIGNAL: true` in the RALPH_STATUS block
|
||||
|
||||
### Circuit Breaker Pattern
|
||||
|
||||
Three states: CLOSED → HALF_OPEN → OPEN
|
||||
|
||||
| From State | Condition | To State |
|
||||
|------------|-----------|----------|
|
||||
| CLOSED | consecutive_no_progress >= 2 | HALF_OPEN |
|
||||
| CLOSED | consecutive_no_progress >= 3 | OPEN |
|
||||
| CLOSED | consecutive_same_error >= 5 | OPEN |
|
||||
| HALF_OPEN | progress detected | CLOSED |
|
||||
| HALF_OPEN | consecutive_no_progress >= 3 | OPEN |
|
||||
| OPEN | Manual reset | CLOSED |
|
||||
|
||||
### @fix_plan.md Structure
|
||||
|
||||
```markdown
|
||||
# Fix Plan
|
||||
|
||||
## High Priority (P0)
|
||||
- [ ] Critical: Fix authentication bug
|
||||
- [ ] Blocker: Database connection timeout
|
||||
|
||||
## Standard (P1)
|
||||
- [ ] Feature: Add user profile page
|
||||
|
||||
## Nice to Have (P2)
|
||||
- [ ] Improvement: Add dark mode
|
||||
|
||||
## Completed
|
||||
- [x] Setup: Initialize project structure
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Quick Wins (1-2 days)
|
||||
|
||||
### 1.1 RALPH_STATUS Block Parsing
|
||||
|
||||
**What**: Add parsing support for the structured RALPH_STATUS block format in RalphTracker.
|
||||
|
||||
**Implementation**:
|
||||
- Add regex pattern to detect `---RALPH_STATUS---` blocks
|
||||
- Parse fields: STATUS, TASKS_COMPLETED_THIS_LOOP, FILES_MODIFIED, TESTS_STATUS, WORK_TYPE, EXIT_SIGNAL, RECOMMENDATION
|
||||
- Store in extended `RalphTrackerState` type
|
||||
- Emit new events: `ralphStatusUpdate`
|
||||
|
||||
**Files**: `ralph-tracker.ts`, `types.ts`
|
||||
|
||||
### 1.2 Enhanced Status Display in UI
|
||||
|
||||
**What**: Display RALPH_STATUS fields in the Ralph State Panel.
|
||||
|
||||
**Implementation**:
|
||||
- Add UI elements: WORK_TYPE indicator, TESTS_STATUS badge, FILES_MODIFIED count
|
||||
- Show RECOMMENDATION text in expanded view
|
||||
- Color-code status (IN_PROGRESS=blue, COMPLETE=green, BLOCKED=red)
|
||||
|
||||
**Files**: `app.js`, `styles.css`, `index.html`
|
||||
|
||||
### 1.3 Prompt Template Improvements
|
||||
|
||||
**What**: Add specification-by-example exit scenarios to prompts.
|
||||
|
||||
**Implementation**:
|
||||
- Add "Exit Scenarios" section to case-template.md
|
||||
- Document when to continue vs. when to output completion
|
||||
- Include testing limits guidance (max 20% effort on tests)
|
||||
- Add RALPH_STATUS block instructions
|
||||
|
||||
**Files**: `case-template.md`
|
||||
|
||||
### 1.4 Better Wizard Validation
|
||||
|
||||
**What**: Add client-side validation and helpful warnings.
|
||||
|
||||
**Implementation**:
|
||||
- Warn if task description < 50 chars
|
||||
- Warn if no success criteria mentioned
|
||||
- Suggest adding test requirements if none detected
|
||||
- Validate completion phrase is uppercase alphanumeric
|
||||
|
||||
**Files**: `app.js`
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Core Improvements (3-5 days)
|
||||
|
||||
### 2.1 Circuit Breaker Pattern
|
||||
|
||||
**What**: Implement three-state circuit breaker to detect stuck loops.
|
||||
|
||||
**Implementation**:
|
||||
- Create `CircuitBreaker` class with CLOSED, HALF_OPEN, OPEN states
|
||||
- Track: files_modified, tasks_completed, error_patterns per iteration
|
||||
- Triggers: N consecutive no-progress, same error M times, tests failing K iterations
|
||||
- Emit events: `circuitBreakerStateChange`
|
||||
|
||||
**Files**: New `circuit-breaker.ts`, integrate into `ralph-tracker.ts`
|
||||
|
||||
### 2.2 Circuit Breaker UI
|
||||
|
||||
**What**: Visual indicator in Ralph panel.
|
||||
|
||||
**Implementation**:
|
||||
- Badge: green (CLOSED), yellow (HALF_OPEN), red (OPEN)
|
||||
- Warning before tripping
|
||||
- Notification when circuit opens
|
||||
- Manual reset button
|
||||
|
||||
**Files**: `app.js`, `styles.css`, `index.html`
|
||||
|
||||
### 2.3 @fix_plan.md Integration
|
||||
|
||||
**What**: Generate and track structured task plan file.
|
||||
|
||||
**Implementation**:
|
||||
- Generate `@fix_plan.md` in working directory when loop starts
|
||||
- Watch file for changes and sync with RalphTracker todos
|
||||
- Parse priority levels (P0, P1, P2)
|
||||
- Show priority in UI
|
||||
|
||||
**Files**: New `fix-plan.ts`, `ralph-tracker.ts`, `server.ts`
|
||||
|
||||
### 2.4 Wizard Plan Generation Step
|
||||
|
||||
**What**: Add third wizard step for AI-assisted plan generation.
|
||||
|
||||
**Implementation**:
|
||||
- Step 2: "Plan Generation" between Task Setup and Launch
|
||||
- Use Claude to break down task into fix plan items
|
||||
- Allow edit/reorder before launch
|
||||
- Generate @fix_plan.md with selected items
|
||||
|
||||
**Files**: `app.js`, `index.html`, `server.ts`
|
||||
|
||||
### 2.5 Smart Respawn Integration
|
||||
|
||||
**What**: Use RALPH_STATUS for respawn decisions.
|
||||
|
||||
**Implementation**:
|
||||
- Use EXIT_SIGNAL field for respawn decisions
|
||||
- If STATUS=BLOCKED, trigger circuit breaker instead of respawn
|
||||
- Pass RECOMMENDATION to respawn update prompt
|
||||
|
||||
**Files**: `respawn-controller.ts`, `ralph-tracker.ts`
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Advanced Features (5+ days)
|
||||
|
||||
### 3.1 Template Library
|
||||
- Bug Fix, Feature, Refactoring, Test Coverage, Documentation templates
|
||||
- Template selector in wizard
|
||||
- Custom templates in `~/.codeman/templates/`
|
||||
|
||||
### 3.2 Tool Permissions
|
||||
- Configure allowed Claude tools per loop
|
||||
- Generate hook configuration
|
||||
- Store in session config
|
||||
|
||||
### 3.3 Per-Iteration Timeout
|
||||
- Max time per iteration (5-60 min)
|
||||
- Auto-continue on timeout
|
||||
- Log timeout events
|
||||
|
||||
### 3.4 Rate Limiting
|
||||
- Max tokens per iteration
|
||||
- Max API calls per minute
|
||||
- Cooldown between iterations
|
||||
|
||||
### 3.5 Metrics Dashboard
|
||||
- Time-series charts (files modified, tasks completed, tokens)
|
||||
- Aggregate statistics
|
||||
- Export to JSON/CSV
|
||||
|
||||
---
|
||||
|
||||
## Priority Matrix
|
||||
|
||||
| Item | Effort | Impact | Priority |
|
||||
|------|--------|--------|----------|
|
||||
| 1.1 RALPH_STATUS Parsing | Low | High | P0 |
|
||||
| 1.2 Status Display UI | Low | Medium | P0 |
|
||||
| 1.3 Prompt Templates | Low | High | P0 |
|
||||
| 1.4 Wizard Validation | Low | Medium | P1 |
|
||||
| 2.1 Circuit Breaker | Medium | High | P1 |
|
||||
| 2.2 Circuit Breaker UI | Medium | Medium | P1 |
|
||||
| 2.3 Fix Plan Integration | Medium | High | P1 |
|
||||
| 2.4 Plan Generation Step | Medium | Medium | P2 |
|
||||
| 2.5 Respawn Integration | Medium | High | P1 |
|
||||
| 3.1 Template Selection | High | Medium | P2 |
|
||||
| 3.2 Tool Permissions | High | Medium | P3 |
|
||||
| 3.3 Per-Iteration Timeout | High | Medium | P2 |
|
||||
| 3.4 Rate Limiting | High | Low | P3 |
|
||||
| 3.5 Metrics Dashboard | High | Medium | P3 |
|
||||
|
||||
---
|
||||
|
||||
## Reference: Ralph Claude Code Best Practices
|
||||
|
||||
### Testing Guidelines
|
||||
- LIMIT testing to ~20% of total effort per loop
|
||||
- PRIORITIZE: Implementation > Documentation > Tests
|
||||
- Only write tests for NEW functionality
|
||||
- Do NOT refactor existing tests unless broken
|
||||
|
||||
### What NOT to Do
|
||||
- Do NOT continue with busy work when EXIT_SIGNAL should be true
|
||||
- Do NOT run tests repeatedly without implementing new features
|
||||
- Do NOT refactor code that is already working
|
||||
- Do NOT add features not in specifications
|
||||
- Do NOT forget the status block
|
||||
|
||||
### Exit Scenarios (Specification by Example)
|
||||
|
||||
1. **Successful Completion**: All tasks done → EXIT_SIGNAL=true
|
||||
2. **Test-Only Loop**: No implementation, only testing → continue but warn
|
||||
3. **Stuck on Error**: Same error 5 times → circuit breaker opens
|
||||
4. **No Work Remaining**: All specs done → EXIT_SIGNAL=true
|
||||
5. **Making Progress**: Normal flow → continue
|
||||
6. **Blocked**: Needs human intervention → STATUS=BLOCKED
|
||||
@@ -1,434 +0,0 @@
|
||||
# Ralph Tracker Phase 1 Implementation Plan
|
||||
|
||||
## Overview
|
||||
|
||||
This plan details how to enhance the existing RalphTracker with RALPH_STATUS block parsing, circuit breaker pattern, and dual-condition exit gate.
|
||||
|
||||
---
|
||||
|
||||
## 1. Current State Analysis
|
||||
|
||||
### What RalphTracker Already Does Well
|
||||
|
||||
- **Todo Detection**: Supports 5 formats (checkboxes, indicators, status in parentheses, native TodoWrite, checkmark-based)
|
||||
- **Completion Phrases**: Detects `<promise>PHRASE</promise>` with occurrence-based logic (1st = store, 2nd = complete)
|
||||
- **Loop State Tracking**: Tracks active/inactive, iteration counts, max iterations, elapsed hours, cycle counts
|
||||
- **Auto-Enable**: Disabled by default, auto-enables when Ralph patterns detected
|
||||
- **Event System**: Emits `loopUpdate`, `todoUpdate`, `completionDetected`, `enabled` events
|
||||
- **SSE Integration**: Events forwarded via `session:ralphLoopUpdate`, `session:ralphTodoUpdate`, `session:ralphCompletionDetected`
|
||||
- **Debouncing**: EVENT_DEBOUNCE_MS (50ms) for rapid updates to prevent UI jitter
|
||||
- **Cleanup**: MAX_TODO_ITEMS (50), TODO_EXPIRY_MS (1 hour), throttled cleanup
|
||||
|
||||
### Current Limitations
|
||||
|
||||
| Feature | Status |
|
||||
|---------|--------|
|
||||
| RALPH_STATUS block parsing | Missing |
|
||||
| Circuit breaker pattern | Missing |
|
||||
| Priority-based todos (P0/P1/P2) | Missing |
|
||||
| Dual-condition exit gate | Missing |
|
||||
| Files modified tracking | Missing |
|
||||
| Tests status tracking | Missing |
|
||||
| Work type classification | Missing |
|
||||
|
||||
---
|
||||
|
||||
## 2. New Type Definitions (types.ts)
|
||||
|
||||
```typescript
|
||||
// ========== RALPH_STATUS Block Types ==========
|
||||
|
||||
export type RalphStatusValue = 'IN_PROGRESS' | 'COMPLETE' | 'BLOCKED';
|
||||
export type RalphTestsStatus = 'PASSING' | 'FAILING' | 'NOT_RUN';
|
||||
export type RalphWorkType = 'IMPLEMENTATION' | 'TESTING' | 'DOCUMENTATION' | 'REFACTORING';
|
||||
|
||||
/**
|
||||
* Parsed RALPH_STATUS block from Claude output.
|
||||
*/
|
||||
export interface RalphStatusBlock {
|
||||
status: RalphStatusValue;
|
||||
tasksCompletedThisLoop: number;
|
||||
filesModified: number;
|
||||
testsStatus: RalphTestsStatus;
|
||||
workType: RalphWorkType;
|
||||
exitSignal: boolean;
|
||||
recommendation: string;
|
||||
parsedAt: number;
|
||||
}
|
||||
|
||||
// ========== Circuit Breaker Types ==========
|
||||
|
||||
export type CircuitBreakerState = 'CLOSED' | 'HALF_OPEN' | 'OPEN';
|
||||
|
||||
export type CircuitBreakerReason =
|
||||
| 'normal_operation'
|
||||
| 'no_progress_warning'
|
||||
| 'no_progress_open'
|
||||
| 'same_error_repeated'
|
||||
| 'tests_failing_too_long'
|
||||
| 'progress_detected'
|
||||
| 'manual_reset';
|
||||
|
||||
export interface CircuitBreakerStatus {
|
||||
state: CircuitBreakerState;
|
||||
consecutiveNoProgress: number;
|
||||
consecutiveSameError: number;
|
||||
consecutiveTestsFailure: number;
|
||||
lastProgressIteration: number;
|
||||
reason: string;
|
||||
reasonCode: CircuitBreakerReason;
|
||||
lastTransitionAt: number;
|
||||
lastErrorMessage: string | null;
|
||||
}
|
||||
|
||||
// ========== Priority Todo Types ==========
|
||||
|
||||
export type RalphTodoPriority = 'P0' | 'P1' | 'P2' | null;
|
||||
|
||||
// ========== Helper Functions ==========
|
||||
|
||||
export function createInitialCircuitBreakerStatus(): CircuitBreakerStatus {
|
||||
return {
|
||||
state: 'CLOSED',
|
||||
consecutiveNoProgress: 0,
|
||||
consecutiveSameError: 0,
|
||||
consecutiveTestsFailure: 0,
|
||||
lastProgressIteration: 0,
|
||||
reason: 'Initial state',
|
||||
reasonCode: 'normal_operation',
|
||||
lastTransitionAt: Date.now(),
|
||||
lastErrorMessage: null,
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. New Regex Patterns (ralph-tracker.ts)
|
||||
|
||||
```typescript
|
||||
// ---------- RALPH_STATUS Block Patterns ----------
|
||||
|
||||
const RALPH_STATUS_START_PATTERN = /^---RALPH_STATUS---\s*$/;
|
||||
const RALPH_STATUS_END_PATTERN = /^---END_RALPH_STATUS---\s*$/;
|
||||
const RALPH_STATUS_FIELD_PATTERN = /^STATUS:\s*(IN_PROGRESS|COMPLETE|BLOCKED)\s*$/i;
|
||||
const RALPH_TASKS_COMPLETED_PATTERN = /^TASKS_COMPLETED_THIS_LOOP:\s*(\d+)\s*$/i;
|
||||
const RALPH_FILES_MODIFIED_PATTERN = /^FILES_MODIFIED:\s*(\d+)\s*$/i;
|
||||
const RALPH_TESTS_STATUS_PATTERN = /^TESTS_STATUS:\s*(PASSING|FAILING|NOT_RUN)\s*$/i;
|
||||
const RALPH_WORK_TYPE_PATTERN = /^WORK_TYPE:\s*(IMPLEMENTATION|TESTING|DOCUMENTATION|REFACTORING)\s*$/i;
|
||||
const RALPH_EXIT_SIGNAL_PATTERN = /^EXIT_SIGNAL:\s*(true|false)\s*$/i;
|
||||
const RALPH_RECOMMENDATION_PATTERN = /^RECOMMENDATION:\s*(.+)$/i;
|
||||
|
||||
// ---------- Completion Indicator Patterns ----------
|
||||
|
||||
const COMPLETION_INDICATOR_PATTERNS = [
|
||||
/all\s+(?:tasks?|items?|work)\s+(?:are\s+)?(?:completed?|done|finished)/i,
|
||||
/(?:completed?|finished)\s+all\s+(?:tasks?|items?|work)/i,
|
||||
/nothing\s+(?:left|remaining)\s+to\s+do/i,
|
||||
/no\s+more\s+(?:tasks?|items?|work)/i,
|
||||
/everything\s+(?:is\s+)?(?:completed?|done)/i,
|
||||
];
|
||||
|
||||
// ---------- Priority Pattern ----------
|
||||
|
||||
const TODO_PRIORITY_PATTERN = /^\s*(?:\[.\])?\s*(?:Critical:|Blocker:|Feature:|Improvement:)?\s*\(?(P[012])\)?:?\s*/i;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. New State Properties (ralph-tracker.ts)
|
||||
|
||||
```typescript
|
||||
// Add to RalphTracker class
|
||||
|
||||
// Circuit breaker state tracking
|
||||
private _circuitBreaker: CircuitBreakerStatus;
|
||||
|
||||
// RALPH_STATUS block parsing state
|
||||
private _statusBlockBuffer: string[] = [];
|
||||
private _inStatusBlock: boolean = false;
|
||||
private _lastStatusBlock: RalphStatusBlock | null = null;
|
||||
|
||||
// Dual-condition exit tracking
|
||||
private _completionIndicators: number = 0;
|
||||
private _exitGateMet: boolean = false;
|
||||
|
||||
// Cumulative tracking
|
||||
private _totalFilesModified: number = 0;
|
||||
private _totalTasksCompleted: number = 0;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. New Methods to Implement
|
||||
|
||||
### 5.1 RALPH_STATUS Block Parsing
|
||||
|
||||
```typescript
|
||||
private processStatusBlockLine(line: string): void {
|
||||
const trimmed = line.trim();
|
||||
|
||||
if (RALPH_STATUS_START_PATTERN.test(trimmed)) {
|
||||
this._inStatusBlock = true;
|
||||
this._statusBlockBuffer = [];
|
||||
return;
|
||||
}
|
||||
|
||||
if (this._inStatusBlock && RALPH_STATUS_END_PATTERN.test(trimmed)) {
|
||||
this._inStatusBlock = false;
|
||||
this.parseStatusBlock(this._statusBlockBuffer);
|
||||
this._statusBlockBuffer = [];
|
||||
return;
|
||||
}
|
||||
|
||||
if (this._inStatusBlock) {
|
||||
this._statusBlockBuffer.push(trimmed);
|
||||
}
|
||||
}
|
||||
|
||||
private parseStatusBlock(lines: string[]): void {
|
||||
const block: Partial<RalphStatusBlock> = { parsedAt: Date.now() };
|
||||
|
||||
for (const line of lines) {
|
||||
// Parse each field...
|
||||
}
|
||||
|
||||
if (block.status !== undefined) {
|
||||
this._lastStatusBlock = fullBlock;
|
||||
this.handleStatusBlock(fullBlock);
|
||||
}
|
||||
}
|
||||
|
||||
private handleStatusBlock(block: RalphStatusBlock): void {
|
||||
this._totalFilesModified += block.filesModified;
|
||||
this._totalTasksCompleted += block.tasksCompletedThisLoop;
|
||||
|
||||
const hasProgress = block.filesModified > 0 || block.tasksCompletedThisLoop > 0;
|
||||
this.updateCircuitBreaker(hasProgress, block.testsStatus, block.status);
|
||||
|
||||
if (block.status === 'COMPLETE') {
|
||||
this._completionIndicators++;
|
||||
}
|
||||
|
||||
if (block.exitSignal && this._completionIndicators >= 2) {
|
||||
this._exitGateMet = true;
|
||||
this.emit('exitGateMet', { completionIndicators: this._completionIndicators, exitSignal: true });
|
||||
}
|
||||
|
||||
this.emit('statusBlockDetected', block);
|
||||
}
|
||||
```
|
||||
|
||||
### 5.2 Circuit Breaker Logic
|
||||
|
||||
```typescript
|
||||
private updateCircuitBreaker(
|
||||
hasProgress: boolean,
|
||||
testsStatus: RalphTestsStatus,
|
||||
status: RalphStatusValue
|
||||
): void {
|
||||
const prevState = this._circuitBreaker.state;
|
||||
|
||||
if (hasProgress) {
|
||||
this._circuitBreaker.consecutiveNoProgress = 0;
|
||||
this._circuitBreaker.lastProgressIteration = this._loopState.cycleCount;
|
||||
|
||||
if (this._circuitBreaker.state === 'HALF_OPEN') {
|
||||
this._circuitBreaker.state = 'CLOSED';
|
||||
this._circuitBreaker.reasonCode = 'progress_detected';
|
||||
}
|
||||
} else {
|
||||
this._circuitBreaker.consecutiveNoProgress++;
|
||||
|
||||
if (this._circuitBreaker.state === 'CLOSED') {
|
||||
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
|
||||
this._circuitBreaker.state = 'OPEN';
|
||||
this._circuitBreaker.reasonCode = 'no_progress_open';
|
||||
} else if (this._circuitBreaker.consecutiveNoProgress >= 2) {
|
||||
this._circuitBreaker.state = 'HALF_OPEN';
|
||||
this._circuitBreaker.reasonCode = 'no_progress_warning';
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (prevState !== this._circuitBreaker.state) {
|
||||
this._circuitBreaker.lastTransitionAt = Date.now();
|
||||
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
|
||||
}
|
||||
}
|
||||
|
||||
resetCircuitBreaker(): void {
|
||||
this._circuitBreaker = createInitialCircuitBreakerStatus();
|
||||
this._circuitBreaker.reasonCode = 'manual_reset';
|
||||
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
|
||||
}
|
||||
```
|
||||
|
||||
### 5.3 Update processLine Method
|
||||
|
||||
```typescript
|
||||
private processLine(line: string): void {
|
||||
const trimmed = line.trim();
|
||||
if (!trimmed) return;
|
||||
|
||||
// NEW: Check for RALPH_STATUS block
|
||||
this.processStatusBlockLine(trimmed);
|
||||
|
||||
// NEW: Check for completion indicators
|
||||
this.detectCompletionIndicators(trimmed);
|
||||
|
||||
// EXISTING: Rest of the detection methods...
|
||||
this.detectCompletionPhrase(trimmed);
|
||||
this.detectAllTasksComplete(trimmed);
|
||||
this.detectTaskCompletion(trimmed);
|
||||
this.detectLoopStatus(trimmed);
|
||||
this.detectTodoItems(trimmed);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. New Events to Add
|
||||
|
||||
```typescript
|
||||
export interface RalphTrackerEvents {
|
||||
// Existing events
|
||||
loopUpdate: (state: RalphTrackerState) => void;
|
||||
todoUpdate: (todos: RalphTodoItem[]) => void;
|
||||
completionDetected: (phrase: string) => void;
|
||||
enabled: () => void;
|
||||
|
||||
// New events
|
||||
statusBlockDetected: (block: RalphStatusBlock) => void;
|
||||
circuitBreakerUpdate: (status: CircuitBreakerStatus) => void;
|
||||
exitGateMet: (data: { completionIndicators: number; exitSignal: boolean }) => void;
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Server Integration (server.ts)
|
||||
|
||||
```typescript
|
||||
// Add new SSE event handlers in setupSessionListeners()
|
||||
|
||||
session.on('ralphStatusBlockDetected', (block: RalphStatusBlock) => {
|
||||
this.broadcast('session:ralphStatusUpdate', { sessionId: session.id, block });
|
||||
});
|
||||
|
||||
session.on('ralphCircuitBreakerUpdate', (status: CircuitBreakerStatus) => {
|
||||
this.broadcast('session:circuitBreakerUpdate', { sessionId: session.id, status });
|
||||
});
|
||||
|
||||
session.on('ralphExitGateMet', (data) => {
|
||||
this.broadcast('session:exitGateMet', { sessionId: session.id, ...data });
|
||||
});
|
||||
|
||||
// Add API endpoint for circuit breaker reset
|
||||
this.app.post('/api/sessions/:id/ralph-circuit-breaker/reset', async (req) => {
|
||||
const session = this.sessions.get(req.params.id);
|
||||
if (!session) return { success: false, error: 'Session not found' };
|
||||
|
||||
session.ralphTracker?.resetCircuitBreaker();
|
||||
return { success: true };
|
||||
});
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. Frontend Changes (app.js)
|
||||
|
||||
### New SSE Event Listeners
|
||||
|
||||
```javascript
|
||||
this.eventSource.addEventListener('session:ralphStatusUpdate', (e) => {
|
||||
const data = JSON.parse(e.data);
|
||||
this.updateRalphStatusBlock(data.sessionId, data.block);
|
||||
});
|
||||
|
||||
this.eventSource.addEventListener('session:circuitBreakerUpdate', (e) => {
|
||||
const data = JSON.parse(e.data);
|
||||
this.updateCircuitBreaker(data.sessionId, data.status);
|
||||
});
|
||||
```
|
||||
|
||||
### New Rendering Methods
|
||||
|
||||
```javascript
|
||||
updateRalphStatusBlock(sessionId, block) {
|
||||
// Store and render status block
|
||||
}
|
||||
|
||||
renderRalphStatusBlock(block) {
|
||||
// Render STATUS, WORK_TYPE, TESTS_STATUS, RECOMMENDATION
|
||||
}
|
||||
|
||||
updateCircuitBreaker(sessionId, status) {
|
||||
// Store and render circuit breaker state
|
||||
}
|
||||
|
||||
renderCircuitBreaker(status) {
|
||||
// Render badge: green (CLOSED), yellow (HALF_OPEN), red (OPEN)
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 9. Implementation Order
|
||||
|
||||
| Step | Task | Time |
|
||||
|------|------|------|
|
||||
| 1 | Add type definitions to `types.ts` | 30 min |
|
||||
| 2 | Add regex patterns to `ralph-tracker.ts` | 30 min |
|
||||
| 3 | Add state properties to RalphTracker class | 15 min |
|
||||
| 4 | Implement RALPH_STATUS parsing methods | 1.5 hr |
|
||||
| 5 | Implement circuit breaker logic | 1 hr |
|
||||
| 6 | Implement completion indicators | 30 min |
|
||||
| 7 | Update events interface | 15 min |
|
||||
| 8 | Add server SSE handlers and API endpoint | 45 min |
|
||||
| 9 | Add frontend event listeners and rendering | 1 hr |
|
||||
| 10 | Add CSS styles | 30 min |
|
||||
| 11 | Update HTML structure | 15 min |
|
||||
| 12 | Write unit tests | 1.5 hr |
|
||||
|
||||
**Total: ~8 hours**
|
||||
|
||||
---
|
||||
|
||||
## 10. Files to Modify
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/types.ts` | Add RalphStatusBlock, CircuitBreakerStatus, helper functions |
|
||||
| `src/ralph-tracker.ts` | Add patterns, state, parsing methods, circuit breaker |
|
||||
| `src/web/server.ts` | Add SSE handlers, circuit breaker reset endpoint |
|
||||
| `src/web/public/app.js` | Add event listeners, rendering methods |
|
||||
| `src/web/public/styles.css` | Add status block and circuit breaker styles |
|
||||
| `src/web/public/index.html` | Add UI elements to Ralph panel |
|
||||
| `test/ralph-tracker.test.ts` | Add tests for new functionality |
|
||||
|
||||
---
|
||||
|
||||
## 11. Test Cases to Add
|
||||
|
||||
1. **RALPH_STATUS Parsing**
|
||||
- Parse valid status block with all fields
|
||||
- Parse block with missing optional fields
|
||||
- Ignore malformed blocks
|
||||
- Handle multiple blocks in sequence
|
||||
|
||||
2. **Circuit Breaker State Transitions**
|
||||
- CLOSED → HALF_OPEN on 2 no-progress
|
||||
- HALF_OPEN → OPEN on 3 no-progress
|
||||
- HALF_OPEN → CLOSED on progress
|
||||
- Manual reset from OPEN
|
||||
|
||||
3. **Dual-Condition Exit Gate**
|
||||
- Exit when indicators >= 2 AND exitSignal = true
|
||||
- No exit when indicators >= 2 but exitSignal = false
|
||||
- No exit when exitSignal = true but indicators < 2
|
||||
|
||||
4. **Integration Tests**
|
||||
- SSE events broadcast correctly
|
||||
- UI updates on status block detection
|
||||
- Circuit breaker badge updates
|
||||
@@ -1,385 +0,0 @@
|
||||
# Respawn Controller Idle Detection Improvement Plan
|
||||
|
||||
## Executive Summary
|
||||
|
||||
The current respawn controller relies primarily on **parsing terminal output** to detect idle states. This approach is fragile and leads to false positives/negatives (e.g., the w3-reddit-analyse session).
|
||||
|
||||
**Key insight**: Claude Code provides **direct, authoritative signals** via hooks and files that definitively indicate session state. We're receiving some of these signals but not using them for idle detection!
|
||||
|
||||
---
|
||||
|
||||
## Current Detection Layers (What We Have)
|
||||
|
||||
| Layer | Signal | Source | Reliability |
|
||||
|-------|--------|--------|-------------|
|
||||
| 1 | Completion message ("Worked for Xm Xs") | Terminal parsing | Medium - can miss edge cases |
|
||||
| 2 | Output silence (configurable duration) | Terminal activity | Low - Claude can be processing silently |
|
||||
| 3 | Token stability | Terminal parsing | Low - tokens don't change during I/O waits |
|
||||
| 4 | Working pattern absence | Terminal parsing | Medium - patterns can be missed |
|
||||
| 5 | AI idle check | Spawned Claude CLI | High but slow (90s timeout) |
|
||||
|
||||
**Problem**: All layers depend on **parsing terminal output**, which is inherently unreliable.
|
||||
|
||||
---
|
||||
|
||||
## Available Claude Code Signals (Not Fully Utilized)
|
||||
|
||||
### 1. `Stop` Hook ⭐ CRITICAL - DEFINITIVE SIGNAL
|
||||
|
||||
**What it is**: Fires when the main Claude Code agent **finishes responding**.
|
||||
|
||||
**From docs**: "Runs when the main Claude Code agent has finished responding. Does not run if the stoppage occurred due to a user interrupt."
|
||||
|
||||
**Current status**: We receive it via `/api/hook-event` but **don't use it for idle detection**!
|
||||
|
||||
**Input received**:
|
||||
```json
|
||||
{
|
||||
"session_id": "abc123",
|
||||
"transcript_path": "~/.claude/projects/.../00893aaf.jsonl",
|
||||
"hook_event_name": "Stop",
|
||||
"stop_hook_active": true // Important for preventing loops
|
||||
}
|
||||
```
|
||||
|
||||
**Action needed**: The `Stop` hook should be the **PRIMARY** idle detection signal. When Claude fires Stop, the agent has definitively finished its response cycle.
|
||||
|
||||
### 2. `idle_prompt` Notification ⭐ HIGH VALUE
|
||||
|
||||
**What it is**: Fires after **60+ seconds of idle time** when Claude is waiting for user input.
|
||||
|
||||
**From docs**: "When Claude is waiting for user input (after 60+ seconds of idle time)"
|
||||
|
||||
**Current status**: We receive it but only forward it to the UI for notification display.
|
||||
|
||||
**Action needed**: Use `idle_prompt` as a **definitive confirmation** that Claude is idle. If we receive this, there's no need for AI idle checks or output silence timers.
|
||||
|
||||
### 3. Transcript JSONL File ⭐ HIGH VALUE
|
||||
|
||||
**What it is**: Complete conversation history at `~/.claude/projects/{project-hash}/{session-id}.jsonl`
|
||||
|
||||
**Current status**: We already watch subagent transcripts but **not the main session transcript**.
|
||||
|
||||
**Data available**:
|
||||
- Every message (user, assistant, system)
|
||||
- Every tool call with inputs/outputs
|
||||
- Progress events
|
||||
- Structured, parseable JSON
|
||||
|
||||
**Action needed**:
|
||||
- Monitor the main transcript file (path provided in every hook input)
|
||||
- Parse the last few entries to detect:
|
||||
- Tool completion
|
||||
- Assistant message completion
|
||||
- Error states
|
||||
- Plan mode prompts
|
||||
|
||||
### 4. `PostToolUse` Hook - Tool Completion Tracking
|
||||
|
||||
**What it is**: Fires immediately after any tool completes successfully.
|
||||
|
||||
**Use case**: Track exactly when tools finish to understand execution flow.
|
||||
|
||||
**Current status**: Not implemented.
|
||||
|
||||
**Action needed**: Add PostToolUse hooks to track tool completion events.
|
||||
|
||||
### 5. `SubagentStop` Hook - Background Agent Completion
|
||||
|
||||
**What it is**: Fires when a subagent (Task tool) finishes responding.
|
||||
|
||||
**Current status**: Not implemented in hooks config (we watch JSONL files separately).
|
||||
|
||||
**Action needed**: Add to hooks config for redundant detection.
|
||||
|
||||
### 6. `permission_prompt` and `elicitation_dialog` - Blocking State Detection
|
||||
|
||||
**What it is**: Fires when Claude needs user input (permission or question).
|
||||
|
||||
**Current status**: We receive and use for auto-accept blocking.
|
||||
|
||||
**Enhancement**: Use as definitive "Claude is NOT idle - it's waiting for user action".
|
||||
|
||||
---
|
||||
|
||||
## Proposed Architecture: Multi-Signal Idle Detection
|
||||
|
||||
### New Detection Hierarchy
|
||||
|
||||
```
|
||||
Priority 1 (Definitive):
|
||||
└── Stop hook received → CONFIRMED IDLE
|
||||
└── idle_prompt received → CONFIRMED IDLE (60s+ idle)
|
||||
|
||||
Priority 2 (Blocking):
|
||||
└── permission_prompt received → NOT IDLE (waiting for permission)
|
||||
└── elicitation_dialog received → NOT IDLE (waiting for answer)
|
||||
└── Working patterns in terminal → NOT IDLE
|
||||
|
||||
Priority 3 (Supporting):
|
||||
└── Transcript analysis → Check last entries for completion
|
||||
└── Output silence + token stability → Weak idle signal
|
||||
|
||||
Priority 4 (Fallback):
|
||||
└── AI idle check → Only if no definitive signals after timeout
|
||||
```
|
||||
|
||||
### State Machine Changes
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────┐
|
||||
│ │
|
||||
▼ │
|
||||
┌─────────────────────┐ │
|
||||
│ WATCHING │◄──────────────────────────────┤
|
||||
└─────────────────────┘ │
|
||||
│ │ │
|
||||
│ │ Stop hook or idle_prompt │
|
||||
│ └────────────────────────┐ │
|
||||
│ ▼ │
|
||||
│ Output silence ┌────────────┐ │
|
||||
│ (no definitive signals) │ HOOK_IDLE │───────┤
|
||||
│ └────────────┘ │
|
||||
▼ (skip AI check) │
|
||||
┌────────────────────┐ │
|
||||
│ CONFIRMING_IDLE │ │
|
||||
└────────────────────┘ │
|
||||
│ │
|
||||
│ Silence confirmed │
|
||||
▼ │
|
||||
┌────────────────────┐ │
|
||||
│ AI_CHECKING │──── IDLE verdict ─────────────┤
|
||||
└────────────────────┘ │
|
||||
│ │
|
||||
│ WORKING verdict │
|
||||
└─────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### New State: `hook_idle`
|
||||
|
||||
When a definitive hook signal is received:
|
||||
1. Skip AI idle check entirely (saves time and API calls)
|
||||
2. Short confirmation period (2-3s) to handle race conditions
|
||||
3. Proceed directly to respawn sequence
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Phase 1: Use Stop Hook for Idle Detection ✅ COMPLETED
|
||||
|
||||
**Files modified**:
|
||||
- `src/respawn-controller.ts`
|
||||
- `src/web/server.ts`
|
||||
- `test/respawn-controller.test.ts`
|
||||
|
||||
**Changes implemented**:
|
||||
1. Added `stopHookReceived`, `stopHookTime`, `idlePromptReceived`, `idlePromptTime` fields to `DetectionStatus`
|
||||
2. Added `hookConfirmTimer` for short confirmation after hook signal (3s)
|
||||
3. Added `signalStopHook()` method:
|
||||
- Sets `stopHookReceived = true` and timestamp
|
||||
- Cancels any running AI check (hook is definitive)
|
||||
- Starts 3s confirmation timer
|
||||
- If no new output during confirmation → triggers respawn cycle
|
||||
4. Added `signalIdlePrompt()` method:
|
||||
- Sets `idlePromptReceived = true` and timestamp
|
||||
- Immediately confirms idle (skips confirmation timer - 60s+ already proven)
|
||||
5. Added `resetHookState()` to clear hook flags on:
|
||||
- Controller start
|
||||
- Working patterns detected
|
||||
- Cycle completion
|
||||
6. Updated server.ts `/api/hook-event` endpoint to call:
|
||||
- `controller.signalStopHook()` for `stop` events
|
||||
- `controller.signalIdlePrompt()` for `idle_prompt` events
|
||||
7. Updated `getDetectionStatus()`:
|
||||
- Returns hook signal states
|
||||
- Sets confidence to 100% when hook received
|
||||
- Updates statusText to show hook status
|
||||
|
||||
**Tests added** (9 new tests in `RespawnController Hook-Based Idle Detection` describe block):
|
||||
- `should expose signalStopHook method`
|
||||
- `should expose signalIdlePrompt method`
|
||||
- `should set stopHookReceived in detection status when Stop hook signaled`
|
||||
- `should include hook status in statusText when Stop hook received`
|
||||
- `should trigger respawn cycle after Stop hook confirmation`
|
||||
- `should immediately confirm idle when idle_prompt signaled (skip confirmation)`
|
||||
- `should cancel Stop hook confirmation if working patterns detected`
|
||||
- `should ignore Stop hook when not in watching state`
|
||||
- `should have 100% confidence when hook signal is received`
|
||||
|
||||
**Detection status update** (implemented):
|
||||
```typescript
|
||||
interface DetectionStatus {
|
||||
/** Layer 0: Stop hook received (highest priority - definitive signal) */
|
||||
stopHookReceived: boolean;
|
||||
stopHookTime: number | null;
|
||||
/** Layer 0: idle_prompt notification received (definitive signal) */
|
||||
idlePromptReceived: boolean;
|
||||
idlePromptTime: number | null;
|
||||
|
||||
// Existing fields...
|
||||
}
|
||||
```
|
||||
|
||||
### Phase 2: Use idle_prompt for Definitive Idle ✅ COMPLETED (in Phase 1)
|
||||
|
||||
**Already implemented in Phase 1**:
|
||||
1. `signalIdlePrompt()` method sets `idlePromptReceived = true`
|
||||
2. Immediately calls `onIdleConfirmed()` - skips all other detection
|
||||
3. Server.ts calls `controller.signalIdlePrompt()` when `idle_prompt` event received
|
||||
4. 60s+ of Claude waiting = definitive idle signal
|
||||
|
||||
### Phase 3: Transcript File Monitoring ✅ COMPLETED
|
||||
|
||||
**New file**: `src/transcript-watcher.ts`
|
||||
|
||||
**Functionality implemented**:
|
||||
1. Watch the session's transcript JSONL file using `fs.watch()`
|
||||
2. Parse new entries as they're appended (incremental reading from last position)
|
||||
3. Detect:
|
||||
- `result` entry → `transcript:complete` event (isComplete = true)
|
||||
- `tool_use` content block → `transcript:tool_start` event
|
||||
- `tool_result` content block → `transcript:tool_end` event
|
||||
- `AskUserQuestion` or `ExitPlanMode` tools → `transcript:plan_mode` event
|
||||
- Error conditions in result entries
|
||||
4. Emit structured events consumed by respawn controller
|
||||
|
||||
**Integration implemented**:
|
||||
- `transcript_path` added to allowed hook data fields in `sanitizeHookData()`
|
||||
- `transcriptWatchers` Map added to WebServer for per-session watchers
|
||||
- `startTranscriptWatcher()` creates watcher and wires up events:
|
||||
- `transcript:complete` → `controller.signalTranscriptComplete()`
|
||||
- `transcript:plan_mode` → `controller.signalTranscriptPlanMode()`
|
||||
- `stopTranscriptWatcher()` cleans up on session cleanup
|
||||
- Hook events with `transcript_path` automatically start watching
|
||||
|
||||
**RespawnController methods added**:
|
||||
- `signalTranscriptComplete()` - Supporting signal that can accelerate idle detection
|
||||
- `signalTranscriptPlanMode()` - Cancels auto-accept timer (like elicitation)
|
||||
|
||||
**Tests added** (13 tests in `test/transcript-watcher.test.ts`):
|
||||
- Initialization tests
|
||||
- File watching tests (existing file, non-existent file, stop, updatePath)
|
||||
- Entry processing tests (user entry, result entry, tool execution, plan mode, errors)
|
||||
- State management tests
|
||||
|
||||
### Phase 4: Enhanced Hook Configuration
|
||||
|
||||
**Update `src/hooks-config.ts`**:
|
||||
|
||||
```typescript
|
||||
export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
|
||||
return {
|
||||
hooks: {
|
||||
Notification: [
|
||||
{ matcher: 'idle_prompt', hooks: [...] },
|
||||
{ matcher: 'permission_prompt', hooks: [...] },
|
||||
{ matcher: 'elicitation_dialog', hooks: [...] },
|
||||
],
|
||||
Stop: [{ hooks: [...] }],
|
||||
// NEW: Add these
|
||||
PostToolUse: [
|
||||
{ matcher: '*', hooks: [...] } // Track all tool completions
|
||||
],
|
||||
SubagentStop: [{ hooks: [...] }],
|
||||
PreCompact: [
|
||||
{ matcher: '*', hooks: [...] } // Track compaction
|
||||
],
|
||||
},
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
### Phase 5: Confidence Scoring Overhaul
|
||||
|
||||
Replace current confidence calculation with weighted signals:
|
||||
|
||||
```typescript
|
||||
function calculateConfidence(): number {
|
||||
let confidence = 0;
|
||||
|
||||
// Definitive signals (100% confidence)
|
||||
if (this.stopHookReceived) confidence = 100;
|
||||
if (this.idlePromptReceived) confidence = 100;
|
||||
|
||||
// Blocking signals (0% confidence)
|
||||
if (this.permissionPromptReceived) return 0;
|
||||
if (this.elicitationReceived) return 0;
|
||||
if (this.workingPatternRecent) return 0;
|
||||
|
||||
// Supporting signals (build up to ~80%)
|
||||
if (confidence < 100) {
|
||||
if (this.outputSilent) confidence += 30;
|
||||
if (this.tokensStable) confidence += 20;
|
||||
if (this.transcriptShowsCompletion) confidence += 30;
|
||||
}
|
||||
|
||||
return Math.min(100, confidence);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Expected Benefits
|
||||
|
||||
| Metric | Current | After Implementation |
|
||||
|--------|---------|---------------------|
|
||||
| False positive rate | ~15-20% | <5% |
|
||||
| Detection latency | 10-90s (AI check) | 3-5s (hook-based) |
|
||||
| API calls for AI check | Every idle detection | Only when hooks unavailable |
|
||||
| Reliability | Medium | High (definitive signals) |
|
||||
|
||||
---
|
||||
|
||||
## Testing Strategy
|
||||
|
||||
### Unit Tests
|
||||
1. `Stop` hook triggers immediate idle confirmation
|
||||
2. `idle_prompt` skips all other detection
|
||||
3. `permission_prompt` blocks idle detection
|
||||
4. Transcript parsing correctly identifies completion
|
||||
5. Fallback to AI check when no hooks received
|
||||
|
||||
### Integration Tests
|
||||
1. End-to-end with real Claude session
|
||||
2. Hook event delivery and handling
|
||||
3. Transcript file monitoring
|
||||
4. Race condition handling
|
||||
|
||||
### Scenarios to Test
|
||||
1. Normal completion → Stop hook → respawn
|
||||
2. Long-running task → idle_prompt → respawn
|
||||
3. Permission needed → wait for user action
|
||||
4. AskUserQuestion → wait for user answer
|
||||
5. Plan mode → auto-accept → continue
|
||||
6. Hooks disabled/unavailable → fallback to AI check
|
||||
|
||||
---
|
||||
|
||||
## Migration Path
|
||||
|
||||
1. **Implement Phase 1** - Stop hook detection (low risk, high value)
|
||||
2. **Deploy and monitor** - Verify Stop hooks are reliable
|
||||
3. **Implement Phase 2** - idle_prompt (simple addition)
|
||||
4. **Implement Phase 3** - Transcript monitoring (more complex)
|
||||
5. **Implement Phase 4** - Enhanced hooks (optional, for completeness)
|
||||
6. **Implement Phase 5** - Refactor confidence scoring
|
||||
|
||||
---
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. **Stop hook reliability**: Does it fire 100% of the time? Edge cases?
|
||||
2. **Transcript file location**: Always at the path in hook input?
|
||||
3. **Hook delivery latency**: How quickly do hooks fire after state change?
|
||||
4. **Race conditions**: What if Stop hook and new work happen simultaneously?
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)
|
||||
- [Agent SDK Documentation](https://platform.claude.com/docs/en/agent-sdk/overview)
|
||||
- Current implementation: `src/respawn-controller.ts`
|
||||
- Hooks config: `src/hooks-config.ts`
|
||||
- Subagent watcher: `src/subagent-watcher.ts`
|
||||
@@ -1,172 +0,0 @@
|
||||
# Run Summary Feature - Implementation Plan
|
||||
|
||||
## Overview
|
||||
|
||||
The Run Summary feature provides users with a consolidated view of what happened in their session while they were away. It tracks significant events, issues, and statistics, presenting them in an easy-to-digest format.
|
||||
|
||||
## Data Structures
|
||||
|
||||
### RunSummaryEventType
|
||||
```typescript
|
||||
type RunSummaryEventType =
|
||||
| 'session_started'
|
||||
| 'session_stopped'
|
||||
| 'respawn_cycle_started'
|
||||
| 'respawn_cycle_completed'
|
||||
| 'respawn_state_change'
|
||||
| 'error'
|
||||
| 'warning'
|
||||
| 'token_milestone'
|
||||
| 'auto_compact'
|
||||
| 'auto_clear'
|
||||
| 'idle_detected'
|
||||
| 'working_detected'
|
||||
| 'ralph_completion'
|
||||
| 'ai_check_result'
|
||||
| 'hook_event'
|
||||
| 'state_stuck';
|
||||
```
|
||||
|
||||
### RunSummaryEvent
|
||||
```typescript
|
||||
interface RunSummaryEvent {
|
||||
id: string;
|
||||
timestamp: number;
|
||||
type: RunSummaryEventType;
|
||||
severity: 'info' | 'warning' | 'error' | 'success';
|
||||
title: string;
|
||||
details?: string;
|
||||
metadata?: Record<string, unknown>;
|
||||
}
|
||||
```
|
||||
|
||||
### RunSummary
|
||||
```typescript
|
||||
interface RunSummary {
|
||||
sessionId: string;
|
||||
sessionName: string;
|
||||
startedAt: number;
|
||||
lastUpdatedAt: number;
|
||||
events: RunSummaryEvent[];
|
||||
stats: {
|
||||
totalRespawnCycles: number;
|
||||
totalTokensUsed: number;
|
||||
peakTokens: number;
|
||||
totalTimeActiveMs: number;
|
||||
totalTimeIdleMs: number;
|
||||
errorCount: number;
|
||||
warningCount: number;
|
||||
aiCheckCount: number;
|
||||
lastIdleAt: number | null;
|
||||
lastWorkingAt: number | null;
|
||||
stateTransitions: number;
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
## Files to Create/Modify
|
||||
|
||||
### 1. `src/run-summary.ts` (NEW)
|
||||
- `RunSummaryTracker` class
|
||||
- Event tracking and aggregation
|
||||
- Statistics calculation
|
||||
- Max 1000 events per session (FIFO trimming)
|
||||
|
||||
### 2. `src/types.ts` (MODIFY)
|
||||
- Add `RunSummaryEvent`, `RunSummaryEventType`, `RunSummary` interfaces
|
||||
- Add `RunSummaryEventSeverity` type
|
||||
|
||||
### 3. `src/web/server.ts` (MODIFY)
|
||||
- Create `RunSummaryTracker` per session
|
||||
- Subscribe to session events and forward to tracker
|
||||
- Subscribe to respawn controller events
|
||||
- Add API endpoint: `GET /api/sessions/:id/run-summary`
|
||||
- Broadcast `session:runSummaryUpdate` SSE event
|
||||
|
||||
### 4. `src/web/public/app.js` (MODIFY)
|
||||
- Add "Run Summary" button to session header
|
||||
- Create modal to display summary
|
||||
- Handle `session:runSummaryUpdate` SSE event
|
||||
- Timeline view for events
|
||||
- Stats cards at top
|
||||
|
||||
### 5. `src/web/public/index.html` (MODIFY)
|
||||
- Add modal HTML structure for run summary
|
||||
|
||||
### 6. `src/web/public/styles.css` (MODIFY)
|
||||
- Styles for run summary modal and timeline
|
||||
|
||||
## Event Sources
|
||||
|
||||
| Event Type | Source | Trigger |
|
||||
|------------|--------|---------|
|
||||
| session_started | Session | `startInteractive()` / `startShell()` |
|
||||
| session_stopped | Session | `stop()` |
|
||||
| respawn_cycle_started | RespawnController | State → `sending_update` |
|
||||
| respawn_cycle_completed | RespawnController | State → `watching` (after cycle) |
|
||||
| respawn_state_change | RespawnController | Any state transition |
|
||||
| error | Various | Errors caught in try/catch |
|
||||
| warning | RunSummaryTracker | State stuck > 5min, high tokens |
|
||||
| token_milestone | Session | Every 50k tokens |
|
||||
| auto_compact | Session | `autoCompact` event |
|
||||
| auto_clear | Session | `autoClear` event |
|
||||
| idle_detected | Session | `idle` event |
|
||||
| working_detected | Session | `working` event |
|
||||
| ralph_completion | RalphTracker | `completionDetected` event |
|
||||
| ai_check_result | RespawnController | AI check completes |
|
||||
| hook_event | Server | `/api/hook-event` endpoint |
|
||||
| state_stuck | RunSummaryTracker | Same state > 10min |
|
||||
|
||||
## API Endpoint
|
||||
|
||||
### GET /api/sessions/:id/run-summary
|
||||
Returns the full run summary for a session.
|
||||
|
||||
Response:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"summary": {
|
||||
"sessionId": "...",
|
||||
"sessionName": "...",
|
||||
"startedAt": 1234567890,
|
||||
"lastUpdatedAt": 1234567890,
|
||||
"events": [...],
|
||||
"stats": {...}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## UI Design
|
||||
|
||||
### Summary Modal
|
||||
- Header: Session name, duration, status indicator
|
||||
- Stats Cards Row:
|
||||
- Respawn Cycles: count
|
||||
- Tokens Used: peak / current
|
||||
- Active Time: formatted duration
|
||||
- Issues: errors + warnings count
|
||||
- Timeline:
|
||||
- Vertical timeline of events
|
||||
- Color-coded by severity (green=success, blue=info, yellow=warning, red=error)
|
||||
- Expandable details
|
||||
- Filter by event type
|
||||
- Footer: "Close" button
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Add types to `types.ts`
|
||||
2. Create `run-summary.ts` with RunSummaryTracker class
|
||||
3. Integrate tracker with server.ts (create per session, wire events)
|
||||
4. Add API endpoint
|
||||
5. Add frontend modal and button
|
||||
6. Test with live session
|
||||
|
||||
## Storage
|
||||
|
||||
Run summaries are kept in memory only (not persisted to disk) since:
|
||||
- They're session-specific and regenerated on session start
|
||||
- Persisting thousands of events would bloat state.json
|
||||
- Server restart = fresh session anyway
|
||||
|
||||
If persistence is needed later, could add to `state-inner.json` with per-session limits.
|
||||
@@ -1,387 +0,0 @@
|
||||
# Codeman TypeScript Improvement Suggestions
|
||||
|
||||
**Generated**: February 2026
|
||||
**Based on**: Research into TypeScript best practices (2024-2025) and codebase analysis
|
||||
|
||||
---
|
||||
|
||||
## 🔴 High Priority (Low effort, high impact)
|
||||
|
||||
### 1. Use the Already-Installed Zod for API Validation
|
||||
|
||||
Zod v4.3.6 is in `package.json` but **never imported**. API routes use unsafe type assertions:
|
||||
|
||||
```typescript
|
||||
// Current (unsafe)
|
||||
const body = req.body as CreateSessionRequest;
|
||||
|
||||
// Recommended
|
||||
const result = CreateSessionSchema.safeParse(req.body);
|
||||
if (!result.success) return createErrorResponse(ApiErrorCode.INVALID_INPUT, ...);
|
||||
```
|
||||
|
||||
**Impact**: Prevents runtime errors from malformed client requests.
|
||||
|
||||
**Files to update**: `src/web/server.ts` (all POST/PUT routes)
|
||||
|
||||
---
|
||||
|
||||
### 2. Add `assertNever` for Exhaustive Switch Checking
|
||||
|
||||
Switch statements on union types (e.g., `respawn-controller.ts:1072`, `ralph-tracker.ts:2088`) lack exhaustive checking. Adding new union members won't cause compile errors.
|
||||
|
||||
```typescript
|
||||
// Add to src/utils/type-safety.ts
|
||||
export function assertNever(x: never, message?: string): never {
|
||||
throw new Error(message ?? `Unexpected value: ${JSON.stringify(x)}`);
|
||||
}
|
||||
|
||||
// Usage in switch statements
|
||||
switch (status) {
|
||||
case 'idle': return handleIdle();
|
||||
case 'busy': return handleBusy();
|
||||
case 'stopped': return handleStopped();
|
||||
case 'error': return handleError();
|
||||
default: return assertNever(status);
|
||||
}
|
||||
```
|
||||
|
||||
**Impact**: Compile-time guarantee all cases are handled.
|
||||
|
||||
**Files affected**: `respawn-controller.ts`, `ralph-tracker.ts`, any file with switch on union types
|
||||
|
||||
---
|
||||
|
||||
### 3. Standardize `createErrorResponse` Usage
|
||||
|
||||
Currently only used in 2 files despite being a good pattern. Many routes still use ad-hoc error responses.
|
||||
|
||||
**Impact**: Consistent API error format across all endpoints.
|
||||
|
||||
---
|
||||
|
||||
## 🟡 Medium Priority (Medium effort, significant benefit)
|
||||
|
||||
### 4. Convert `ApiResponse<T>` to Discriminated Union
|
||||
|
||||
Current interface has optional properties; discriminated union enables better narrowing:
|
||||
|
||||
```typescript
|
||||
// Current (types.ts)
|
||||
interface ApiResponse<T> { success: boolean; error?: string; data?: T; }
|
||||
|
||||
// Better
|
||||
type ApiResponse<T> =
|
||||
| { success: true; data: T }
|
||||
| { success: false; error: string; errorCode: ApiErrorCode };
|
||||
|
||||
// Usage with exhaustive checking
|
||||
function handleResponse<T>(response: ApiResponse<T>): T {
|
||||
if (response.success) {
|
||||
return response.data; // TypeScript knows data exists
|
||||
} else {
|
||||
throw new Error(response.error); // TypeScript knows error exists
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 5. Add Branded Types for Token Counts
|
||||
|
||||
Prevents mixing input/output tokens in calculations:
|
||||
|
||||
```typescript
|
||||
// src/types/branded.ts
|
||||
type Brand<K, T extends string> = K & { readonly __brand: T };
|
||||
|
||||
export type InputTokens = Brand<number, 'InputTokens'>;
|
||||
export type OutputTokens = Brand<number, 'OutputTokens'>;
|
||||
export type TokenCount = Brand<number, 'TokenCount'>;
|
||||
export type Milliseconds = Brand<number, 'Milliseconds'>;
|
||||
|
||||
// Constructor functions
|
||||
export function inputTokens(value: number): InputTokens {
|
||||
if (value < 0) throw new Error('Token count cannot be negative');
|
||||
return value as InputTokens;
|
||||
}
|
||||
```
|
||||
|
||||
**Use cases**:
|
||||
- Token counts (`_totalInputTokens`, `_totalOutputTokens`)
|
||||
- Timeout values (`idleTimeoutMs`, `completionConfirmMs`, `noOutputTimeoutMs`)
|
||||
- IDs (`SessionId`, `TaskId`, `CycleId`)
|
||||
|
||||
---
|
||||
|
||||
### 6. Dependency Injection for Core Services
|
||||
|
||||
Replace hidden singleton dependencies with constructor injection for better testability:
|
||||
|
||||
```typescript
|
||||
// Current: Hidden dependencies
|
||||
export class RalphLoop extends EventEmitter {
|
||||
constructor() {
|
||||
this.sessionManager = getSessionManager();
|
||||
this.store = getStore();
|
||||
}
|
||||
}
|
||||
|
||||
// Better: Explicit dependencies
|
||||
export interface RalphLoopDeps {
|
||||
sessionManager: SessionManager;
|
||||
taskQueue: TaskQueue;
|
||||
store: StateStore;
|
||||
}
|
||||
|
||||
export class RalphLoop extends EventEmitter {
|
||||
constructor(deps: RalphLoopDeps, options?: RalphLoopOptions) {
|
||||
this.sessionManager = deps.sessionManager;
|
||||
// ...
|
||||
}
|
||||
}
|
||||
|
||||
// Production factory
|
||||
export function createRalphLoop(options?: RalphLoopOptions): RalphLoop {
|
||||
return new RalphLoop({
|
||||
sessionManager: getSessionManager(),
|
||||
taskQueue: getTaskQueue(),
|
||||
store: getStore(),
|
||||
}, options);
|
||||
}
|
||||
```
|
||||
|
||||
**Start with**: `RalphLoop` (has the most dependencies)
|
||||
|
||||
**Benefits**: Easier testing, explicit dependencies, SOLID compliance
|
||||
|
||||
---
|
||||
|
||||
### 7. Enforce Consistent `import type` Usage
|
||||
|
||||
Mixed usage across codebase. Add ESLint rule:
|
||||
|
||||
```json
|
||||
{
|
||||
"rules": {
|
||||
"@typescript-eslint/consistent-type-imports": ["error", {
|
||||
"prefer": "type-imports",
|
||||
"fixStyle": "separate-type-imports"
|
||||
}]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Benefits**: Reduced bundle size, better tree-shaking, cleaner separation
|
||||
|
||||
---
|
||||
|
||||
### 8. Add Circular Dependency Detection
|
||||
|
||||
```bash
|
||||
npm install -D dpdm
|
||||
```
|
||||
|
||||
Add to `package.json`:
|
||||
```json
|
||||
{
|
||||
"scripts": {
|
||||
"check:circular": "dpdm --no-warning --no-tree src/index.ts"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Potential risk areas identified**:
|
||||
- `ralph-loop.ts` → `session-manager.ts` → `session.ts`
|
||||
- `respawn-controller.ts` → `session.ts` → `ai-idle-checker.ts`
|
||||
|
||||
---
|
||||
|
||||
## 🟢 Lower Priority (Higher effort, situational benefit)
|
||||
|
||||
### 9. Apply `as const satisfies` to Default Configs
|
||||
|
||||
Preserves literal types while validating structure:
|
||||
|
||||
```typescript
|
||||
// Current
|
||||
export const DEFAULT_NICE_CONFIG: NiceConfig = {
|
||||
enabled: false,
|
||||
niceValue: 10,
|
||||
};
|
||||
// niceValue is type: number
|
||||
|
||||
// Better
|
||||
export const DEFAULT_NICE_CONFIG = {
|
||||
enabled: false,
|
||||
niceValue: 10,
|
||||
} as const satisfies NiceConfig;
|
||||
// niceValue is type: 10 (literal)
|
||||
```
|
||||
|
||||
**Files**: `types.ts`, `respawn-controller.ts` (DEFAULT_CONFIG)
|
||||
|
||||
---
|
||||
|
||||
### 10. Create Custom Error Class Hierarchy
|
||||
|
||||
Replace string-based errors with typed errors:
|
||||
|
||||
```typescript
|
||||
// src/errors.ts
|
||||
export class CodemanError extends Error {
|
||||
constructor(
|
||||
message: string,
|
||||
public code: string,
|
||||
public context?: Record<string, unknown>
|
||||
) {
|
||||
super(message);
|
||||
Object.setPrototypeOf(this, CodemanError.prototype);
|
||||
this.name = 'CodemanError';
|
||||
}
|
||||
}
|
||||
|
||||
export class SessionError extends CodemanError {
|
||||
constructor(message: string, code: string, public sessionId: string) {
|
||||
super(message, code, { sessionId });
|
||||
this.name = 'SessionError';
|
||||
}
|
||||
}
|
||||
|
||||
export class ValidationError extends CodemanError {
|
||||
constructor(message: string, public field: string, public value: unknown) {
|
||||
super(message, 'VALIDATION_ERROR', { field, value });
|
||||
this.name = 'ValidationError';
|
||||
}
|
||||
}
|
||||
|
||||
export class ScreenError extends CodemanError {
|
||||
constructor(message: string, public screenName: string, public operation: string) {
|
||||
super(message, 'SCREEN_ERROR', { screenName, operation });
|
||||
this.name = 'ScreenError';
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 11. Split Large Files
|
||||
|
||||
**`types.ts` (~1500 lines)**:
|
||||
```
|
||||
src/types/
|
||||
index.ts # Re-exports all
|
||||
session.types.ts # Session-related types
|
||||
task.types.ts # Task-related types
|
||||
ralph.types.ts # Ralph loop types
|
||||
api.types.ts # API request/response types
|
||||
config.types.ts # Configuration types
|
||||
factories.ts # createInitialState(), etc.
|
||||
```
|
||||
|
||||
**`server.ts`**:
|
||||
```
|
||||
src/web/
|
||||
server.ts # Main Fastify setup
|
||||
routes/
|
||||
sessions.ts # Session management routes
|
||||
respawn.ts # Respawn control routes
|
||||
scheduled.ts # Scheduled run routes
|
||||
system.ts # System status routes
|
||||
sse/
|
||||
manager.ts # SSE client management
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 12. Formalize Result Pattern
|
||||
|
||||
Existing `validateTokenCounts` returns `{ isValid, reason }` which is essentially a Result.
|
||||
|
||||
**Option A: Simple Result type (no dependency)**:
|
||||
```typescript
|
||||
// src/utils/result.ts
|
||||
export type Result<T, E = Error> =
|
||||
| { success: true; data: T }
|
||||
| { success: false; error: E };
|
||||
|
||||
export const ok = <T>(data: T): Result<T, never> => ({ success: true, data });
|
||||
export const err = <E>(error: E): Result<never, E> => ({ success: false, error });
|
||||
```
|
||||
|
||||
**Option B: Install neverthrow**:
|
||||
```bash
|
||||
npm install neverthrow
|
||||
```
|
||||
|
||||
Provides chaining (`map`, `andThen`, `match`) and `ResultAsync` for async operations.
|
||||
|
||||
---
|
||||
|
||||
### 13. Template Literal Types for IDs
|
||||
|
||||
Enforce ID formats at compile time:
|
||||
|
||||
```typescript
|
||||
type CycleIdFormat = `${string}:cycle-${number}`;
|
||||
type ScreenSessionName = `codeman-${string}`;
|
||||
|
||||
interface RespawnCycleMetrics {
|
||||
cycleId: CycleIdFormat; // Enforces format at compile time
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Summary Table
|
||||
|
||||
| # | Suggestion | Category | Effort | Impact |
|
||||
|---|------------|----------|--------|--------|
|
||||
| 1 | Use Zod for API validation | Error Handling | Low | High |
|
||||
| 2 | Add `assertNever` utility | Type Safety | Low | High |
|
||||
| 3 | Standardize `createErrorResponse` | Error Handling | Low | Medium |
|
||||
| 4 | Discriminated union for `ApiResponse` | Type Safety | Medium | High |
|
||||
| 5 | Branded types for tokens | Type Safety | Medium | Medium |
|
||||
| 6 | Dependency injection for services | Architecture | Medium | High |
|
||||
| 7 | Enforce `import type` | Architecture | Low | Medium |
|
||||
| 8 | Circular dependency detection | Architecture | Low | Medium |
|
||||
| 9 | `as const satisfies` for configs | Type Safety | Low | Low |
|
||||
| 10 | Custom error classes | Error Handling | Medium | Medium |
|
||||
| 11 | Split large files | Architecture | High | Medium |
|
||||
| 12 | Formalize Result pattern | Error Handling | Medium | Medium |
|
||||
| 13 | Template literal types for IDs | Type Safety | Low | Low |
|
||||
|
||||
---
|
||||
|
||||
## Notable Strengths to Keep
|
||||
|
||||
These patterns are already well-implemented and should be preserved:
|
||||
|
||||
- **Circuit breaker pattern** in `state-store.ts` and `ai-checker-base.ts` (excellent resilience)
|
||||
- **`getErrorMessage()` utility** (solid, used in 8 files)
|
||||
- **Barrel files for `utils/` and `prompts/`** (appropriate size, good organization)
|
||||
- **Strict TypeScript config** (comprehensive strictness settings)
|
||||
- **Well-documented configuration** in `src/config/`
|
||||
- **Extensive union types** for status tracking (18+ well-defined types)
|
||||
- **Type guards** like `isError()` for runtime narrowing
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
### Type Safety
|
||||
- [TypeScript Handbook: Narrowing](https://www.typescriptlang.org/docs/handbook/2/narrowing.html)
|
||||
- [Fullstory: Discriminated Unions](https://www.fullstory.com/blog/discriminated-unions-and-exhaustiveness-checking-in-typescript/)
|
||||
- [Learning TypeScript: Branded Types](https://www.learningtypescript.com/articles/branded-types)
|
||||
- [Total TypeScript: satisfies Operator](https://www.totaltypescript.com/how-to-use-satisfies-operator)
|
||||
|
||||
### Error Handling
|
||||
- [neverthrow GitHub](https://github.com/supermacro/neverthrow)
|
||||
- [Zod Documentation](https://zod.dev/)
|
||||
- [Custom Errors in TypeScript](https://medium.com/@Nelsonalfonso/understanding-custom-errors-in-typescript-a-complete-guide-f47a1df9354c)
|
||||
|
||||
### Architecture
|
||||
- [Please Stop Using Barrel Files - TkDodo](https://tkdodo.eu/blog/please-stop-using-barrel-files)
|
||||
- [TypeScript Dependency Injection](https://softwarepatternslexicon.com/js/typescript-and-javascript-design-patterns/dependency-injection-with-typescript/)
|
||||
- [dpdm - Circular Dependency Detector](https://github.com/acrazing/dpdm)
|
||||
- [Consistent Type Imports - typescript-eslint](https://typescript-eslint.io/blog/consistent-type-imports-and-exports-why-and-how/)
|
||||
@@ -1,155 +0,0 @@
|
||||
# Voice Input V2 — Implementation Plan
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Fix and improve the existing VoiceInput implementation. The core class is solid but has **critical integration bugs** that prevent it from working on mobile, plus several UX improvements needed to make it feel fast and polished.
|
||||
|
||||
---
|
||||
|
||||
## Current State: What Exists
|
||||
|
||||
The `VoiceInput` singleton (app.js:602-830) is already committed and uses the Web Speech API with:
|
||||
- Toggle mode (tap start/stop), 5s silence auto-stop
|
||||
- `interimResults: true` for streaming transcription preview
|
||||
- iOS Safari `isFinal` workaround (750ms stability timer)
|
||||
- Desktop button in `toolbar-right`, mobile button in `KeyboardAccessoryBar`
|
||||
- `voice-pulse` CSS animation, `.voice-preview` overlay
|
||||
- Cleanup on SSE reconnect, haptic feedback on mobile
|
||||
|
||||
## Critical Bugs Found (Must Fix)
|
||||
|
||||
### Bug 1: Mobile button NEVER shows (CRITICAL)
|
||||
`KeyboardAccessoryBar.init()` runs at line 2239, BEFORE `VoiceInput.init()` at line 2240. The accessory bar template checks `VoiceInput.supported` at render time — but `init()` hasn't run yet, so `supported` is still `false`. The inline `style="${VoiceInput.supported ? '' : 'display:none'}"` always resolves to `display:none`.
|
||||
|
||||
**Fix:** Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`, OR remove the inline style check and have `VoiceInput.init()` show/hide the mobile button after the fact (like it does for desktop).
|
||||
|
||||
### Bug 2: `_showButtons()` ignores mobile button
|
||||
`_showButtons()` only targets `#voiceInputBtn` (desktop). It never removes `display:none` from the mobile `[data-action="voice"]` button.
|
||||
|
||||
**Fix:** Add mobile button selector to `_showButtons()`.
|
||||
|
||||
### Bug 3: Recognition instance leak on cleanup
|
||||
`cleanup()` stops recording and removes the preview element, but doesn't null out `this.recognition`. After `cleanup()` + `init()` on SSE reconnect, the old `SpeechRecognition` instance with its handlers is orphaned.
|
||||
|
||||
**Fix:** Add `this.recognition = null` in `cleanup()`.
|
||||
|
||||
## UX Improvements (Should Fix)
|
||||
|
||||
### Improvement 1: Consider auto-sending after voice
|
||||
Currently, voice text is inserted but the user must press Enter. This is safe but adds friction. Two options:
|
||||
- **Option A (safe, current):** Insert text, user presses Enter — good for a terminal where wrong commands matter
|
||||
- **Option B (fast):** Insert text + auto-send `\r` after a brief 500ms delay — feels more "voice assistant"-like
|
||||
- **Recommendation:** Keep Option A as default, but add an optional setting for auto-send
|
||||
|
||||
### Improvement 2: Shorter silence timeout for commands
|
||||
5 seconds of silence before auto-stop feels slow for short terminal commands. Consider:
|
||||
- 3 seconds for auto-stop (still generous for natural pauses)
|
||||
- Or make it configurable via settings
|
||||
|
||||
### Improvement 3: Better visual state on mobile
|
||||
The blue-tinted voice button in the accessory bar is distinctive but subtle. When recording:
|
||||
- The `.recording` class turns it red with pulse — good
|
||||
- But the button is small among other buttons — easy to miss the state change
|
||||
- Consider: also show a small red dot indicator in the header or terminal area during recording
|
||||
|
||||
## Architecture Decision: Keep Web Speech API
|
||||
|
||||
Confirmed by research: Web Speech API is the right choice.
|
||||
- **Free, fast (150-300ms interim), trivial complexity**
|
||||
- Chrome + Safari = ~70% of users, ~95% of Codeman's target audience (devs on Chrome)
|
||||
- Works on localhost without HTTPS
|
||||
- Accuracy is adequate for English command dictation
|
||||
- Deepgram streaming (Phase 2 optional) only if accuracy complaints arise
|
||||
- Skip Whisper batch entirely (too slow for interactive voice input)
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Phase 1: Fix Critical Bugs (Priority)
|
||||
|
||||
**File: `src/web/public/app.js`**
|
||||
|
||||
1. **Fix init order** — Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`:
|
||||
```
|
||||
// Current (broken):
|
||||
KeyboardAccessoryBar.init();
|
||||
VoiceInput.init();
|
||||
|
||||
// Fixed:
|
||||
VoiceInput.init();
|
||||
KeyboardAccessoryBar.init();
|
||||
```
|
||||
|
||||
2. **Fix `_showButtons()` to handle mobile** — Add mobile button selector:
|
||||
```javascript
|
||||
_showButtons() {
|
||||
const desktopBtn = document.getElementById('voiceInputBtn');
|
||||
if (desktopBtn) desktopBtn.style.display = '';
|
||||
// Also show mobile button (may not exist yet if KeyboardAccessoryBar hasn't init'd)
|
||||
const mobileBtn = document.querySelector('[data-action="voice"]');
|
||||
if (mobileBtn) mobileBtn.style.display = '';
|
||||
}
|
||||
```
|
||||
|
||||
3. **Fix cleanup leak** — Null out recognition instance:
|
||||
```javascript
|
||||
cleanup() {
|
||||
if (this.isRecording) this.stop();
|
||||
if (this.previewEl) {
|
||||
this.previewEl.remove();
|
||||
this.previewEl = null;
|
||||
}
|
||||
this.recognition = null; // <-- add this
|
||||
clearTimeout(this.silenceTimeout);
|
||||
clearTimeout(this._stabilityTimer);
|
||||
// ... rest
|
||||
}
|
||||
```
|
||||
|
||||
4. **Remove inline style from mobile button template** — Since `_showButtons()` will handle visibility, the template should always render the button visible and let `init()` hide it if unsupported:
|
||||
```
|
||||
// Current (broken):
|
||||
style="${VoiceInput.supported ? '' : 'display:none'}"
|
||||
|
||||
// Fixed: remove the style attr entirely, let _showButtons/_hideButtons manage it
|
||||
```
|
||||
Actually better: **always show the button** if we init VoiceInput before KeyboardAccessoryBar. The `VoiceInput.supported` will be set correctly by then.
|
||||
|
||||
### Phase 2: UX Polish
|
||||
|
||||
5. **Reduce silence timeout** from 5s to 3s for snappier feel
|
||||
|
||||
6. **Add recording indicator** — When recording, add a subtle pulsing red dot to the session header or status area so the recording state is visible even if the button is off-screen
|
||||
|
||||
7. **Voice input setting** — Add a toggle in App Settings to enable/disable voice input (some users may not want the button). Default: enabled on supported browsers.
|
||||
|
||||
### Phase 3: Future Enhancements (Not in this PR)
|
||||
|
||||
- Language selector (currently hardcoded `en-US`)
|
||||
- Auto-send option (insert text + `\r` automatically)
|
||||
- Deepgram WebSocket fallback for Firefox/Edge
|
||||
- Waveform visualization during recording
|
||||
- Voice command recognition ("clear", "compact", "new session")
|
||||
|
||||
## Files to Modify
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/web/public/app.js` | Fix init order, fix `_showButtons()`, fix `cleanup()`, remove inline style, reduce silence timeout |
|
||||
| `src/web/public/mobile.css` | (optional) Adjust voice preview positioning if needed |
|
||||
|
||||
## Testing Plan
|
||||
|
||||
1. **Desktop Chrome:** Verify mic button visible in toolbar-right, click toggles recording state, interim text shows in preview, final text inserted at prompt
|
||||
2. **Mobile Chrome (emulated):** Verify mic button visible in accessory bar, tap toggles recording, pulse animation plays
|
||||
3. **Firefox:** Verify mic button is hidden (no SpeechRecognition support)
|
||||
4. **SSE reconnect:** Verify cleanup stops recording and re-init works
|
||||
5. **No active session:** Verify toast "No active session" shows when tapping mic with no session
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Impact | Mitigation |
|
||||
|------|--------|------------|
|
||||
| iOS Safari isFinal bug | Medium | Already handled by 750ms stability timer |
|
||||
| Chrome auto-stops after 60s | Low | Prompts are short; 3s silence timeout covers this |
|
||||
| Mic permission denied | Low | Error toast with clear message |
|
||||
| Init order regression | High | Integration test to verify button visibility |
|
||||
@@ -1,343 +0,0 @@
|
||||
# Browser Testing Guide for Codeman
|
||||
|
||||
This guide documents the browser testing infrastructure, framework comparison results, and best practices for testing the Codeman web UI.
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# Run standalone benchmark (recommended - avoids vitest hook issues)
|
||||
npx tsx scripts/browser-comparison.mjs
|
||||
|
||||
# Run existing browser E2E tests
|
||||
npm test -- test/browser-e2e.test.ts
|
||||
```
|
||||
|
||||
## Framework Comparison Results
|
||||
|
||||
We tested three browser automation frameworks against the Codeman web UI:
|
||||
|
||||
| Framework | Avg Duration | Best For |
|
||||
|-----------|--------------|----------|
|
||||
| **Puppeteer** | 1223ms | Simple operations, Chrome-specific features |
|
||||
| **Playwright** | 1373ms | Complex interactions, cross-browser, debugging |
|
||||
| **Agent-Browser** | N/A (timeout) | AI agent navigation with semantic locators |
|
||||
|
||||
### Detailed Benchmarks
|
||||
|
||||
| Scenario | Playwright | Puppeteer |
|
||||
|----------|------------|-----------|
|
||||
| Page load | 1445ms | 433ms |
|
||||
| Element selection | 442ms | 373ms |
|
||||
| Modal interaction | 1487ms | 1605ms |
|
||||
| Rapid operations (5 cycles) | 2119ms | 2482ms |
|
||||
|
||||
**Key findings:**
|
||||
- Puppeteer is faster for simple page loads and element selection
|
||||
- Playwright handles rapid/complex interactions better (auto-waiting)
|
||||
- Agent-browser CLI has startup overhead issues in this environment
|
||||
|
||||
## Known Issues
|
||||
|
||||
### Vitest Hook Timeouts
|
||||
|
||||
**Problem:** Browser tests using vitest's `beforeAll`/`afterAll` hooks consistently timeout, even when the tests actually complete successfully.
|
||||
|
||||
**Symptoms:**
|
||||
- Tests show as "skipped"
|
||||
- Error: "Hook timed out in 60000ms"
|
||||
- But cleanup messages appear (indicating tests ran)
|
||||
|
||||
**Root cause:** Unclear - possibly related to:
|
||||
- vitest's module isolation with async browser launches
|
||||
- Interaction between global setup.ts hooks and test-level hooks
|
||||
- Multiple test file imports causing duplicate hook execution
|
||||
|
||||
**Workarounds:**
|
||||
1. **Use standalone scripts** (recommended):
|
||||
```bash
|
||||
npx tsx scripts/browser-comparison.mjs
|
||||
```
|
||||
|
||||
2. **Run browser code directly in tests** (not in hooks):
|
||||
```typescript
|
||||
it('should test something', async () => {
|
||||
const browser = await chromium.launch();
|
||||
// ... test code ...
|
||||
await browser.close();
|
||||
});
|
||||
```
|
||||
|
||||
3. **Use the existing browser-e2e.test.ts pattern** which uses agent-browser CLI commands via `execSync` (avoids async hook issues)
|
||||
|
||||
## Test File Structure
|
||||
|
||||
### Port Allocation
|
||||
|
||||
| Port Range | Test File |
|
||||
|------------|-----------|
|
||||
| 3150-3153 | browser-e2e.test.ts (existing) |
|
||||
| 3154 | file-link-click.test.ts |
|
||||
| 3155 | browser-playwright.test.ts |
|
||||
| 3156 | browser-puppeteer.test.ts |
|
||||
| 3157 | browser-agent.test.ts |
|
||||
| 3158-3160 | browser-comparison.test.ts |
|
||||
| 3180-3182 | scripts/browser-comparison.mjs |
|
||||
|
||||
### File Purposes
|
||||
|
||||
| File | Framework | Status |
|
||||
|------|-----------|--------|
|
||||
| `test/browser-e2e.test.ts` | agent-browser | ✅ Working |
|
||||
| `test/browser-playwright.test.ts` | Playwright | ⚠️ Vitest hook issues |
|
||||
| `test/browser-puppeteer.test.ts` | Puppeteer | ⚠️ Vitest hook issues |
|
||||
| `test/browser-agent.test.ts` | agent-browser | ⚠️ Vitest hook issues |
|
||||
| `scripts/browser-comparison.mjs` | All three | ✅ Working (standalone) |
|
||||
|
||||
## Framework-Specific Patterns
|
||||
|
||||
### Playwright
|
||||
|
||||
```typescript
|
||||
import { chromium } from 'playwright';
|
||||
|
||||
const browser = await chromium.launch({
|
||||
headless: true,
|
||||
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage'],
|
||||
});
|
||||
|
||||
const page = await browser.newPage();
|
||||
await page.goto('http://localhost:3000');
|
||||
|
||||
// Auto-waiting selectors
|
||||
await page.click('.btn-claude');
|
||||
await page.waitForSelector('.session-tab', { state: 'visible' });
|
||||
|
||||
// Assertions with expect
|
||||
await expect(page.locator('.header')).toBeVisible();
|
||||
await expect(page).toHaveTitle('Codeman');
|
||||
|
||||
await browser.close();
|
||||
```
|
||||
|
||||
**Pros:**
|
||||
- Built-in auto-waiting
|
||||
- Excellent trace viewer for debugging
|
||||
- Cross-browser support (Chromium, Firefox, WebKit)
|
||||
- Native `expect` assertions
|
||||
|
||||
**Cons:**
|
||||
- Slightly slower page loads
|
||||
- Larger dependency
|
||||
|
||||
### Puppeteer
|
||||
|
||||
```typescript
|
||||
import puppeteer from 'puppeteer';
|
||||
|
||||
const browser = await puppeteer.launch({
|
||||
headless: true,
|
||||
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage'],
|
||||
});
|
||||
|
||||
const page = await browser.newPage();
|
||||
await page.goto('http://localhost:3000');
|
||||
|
||||
// Manual waiting often needed
|
||||
await page.click('.btn-claude');
|
||||
await page.waitForSelector('.session-tab', { visible: true });
|
||||
|
||||
// Element queries
|
||||
const title = await page.title();
|
||||
const text = await page.$eval('.logo', el => el.textContent);
|
||||
|
||||
// CDP access for advanced features
|
||||
const client = await page.target().createCDPSession();
|
||||
await client.send('Performance.enable');
|
||||
|
||||
await browser.close();
|
||||
```
|
||||
|
||||
**Pros:**
|
||||
- Faster for simple operations
|
||||
- Direct Chrome DevTools Protocol access
|
||||
- Smaller dependency
|
||||
- Good for Chrome-specific testing
|
||||
|
||||
**Cons:**
|
||||
- Chrome/Chromium only
|
||||
- Manual waiting required
|
||||
- Less robust for complex interactions
|
||||
|
||||
### Agent-Browser (CLI)
|
||||
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
|
||||
function agentBrowser(cmd: string): string {
|
||||
return execSync(`npx agent-browser ${cmd}`, {
|
||||
timeout: 30000,
|
||||
encoding: 'utf-8',
|
||||
}).trim();
|
||||
}
|
||||
|
||||
function agentBrowserJson<T>(cmd: string): T {
|
||||
const result = agentBrowser(`${cmd} --json`);
|
||||
return JSON.parse(result).data;
|
||||
}
|
||||
|
||||
// Usage
|
||||
agentBrowser('open http://localhost:3000');
|
||||
agentBrowser('click ".btn-claude"');
|
||||
const title = agentBrowserJson<{title: string}>('get title');
|
||||
|
||||
// Semantic locators (AI-friendly)
|
||||
agentBrowser('find role button click --name "Submit"');
|
||||
agentBrowser('find text "Settings" click');
|
||||
|
||||
// Accessibility snapshot
|
||||
const snapshot = agentBrowser('snapshot');
|
||||
|
||||
agentBrowser('close');
|
||||
```
|
||||
|
||||
**Pros:**
|
||||
- AI-agent friendly (semantic locators)
|
||||
- Accessibility tree snapshots
|
||||
- Simple CLI interface
|
||||
- Reference-based selection (@e1, @e2)
|
||||
|
||||
**Cons:**
|
||||
- CLI overhead (spawn process per command)
|
||||
- Slower for rapid operations
|
||||
- Less programmatic control
|
||||
|
||||
## Best Practices
|
||||
|
||||
### 1. Use Standalone Scripts for Benchmarks
|
||||
|
||||
Vitest has issues with browser hooks. For reliable benchmarking:
|
||||
|
||||
```bash
|
||||
# Create a standalone .mjs script
|
||||
npx tsx scripts/browser-comparison.mjs
|
||||
```
|
||||
|
||||
### 2. Browser Launch Arguments
|
||||
|
||||
Always include these args for headless environments:
|
||||
|
||||
```typescript
|
||||
{
|
||||
headless: true,
|
||||
args: [
|
||||
'--no-sandbox', // Required for Docker/CI
|
||||
'--disable-setuid-sandbox',
|
||||
'--disable-dev-shm-usage', // Prevents /dev/shm issues
|
||||
],
|
||||
}
|
||||
```
|
||||
|
||||
### 3. Install Playwright Browsers
|
||||
|
||||
```bash
|
||||
npx playwright install chromium
|
||||
```
|
||||
|
||||
### 4. Wait for Server Startup
|
||||
|
||||
```typescript
|
||||
const server = new WebServer(PORT);
|
||||
await server.start();
|
||||
await new Promise(r => setTimeout(r, 1000)); // Allow server to stabilize
|
||||
```
|
||||
|
||||
### 5. Clean Up Sessions
|
||||
|
||||
Track created sessions for cleanup:
|
||||
|
||||
```typescript
|
||||
const createdSessions: string[] = [];
|
||||
|
||||
// In test
|
||||
const response = await fetch(`${BASE_URL}/api/sessions`);
|
||||
const data = await response.json();
|
||||
createdSessions.push(data.sessions[0].id);
|
||||
|
||||
// In cleanup
|
||||
for (const id of createdSessions) {
|
||||
await fetch(`${BASE_URL}/api/sessions/${id}`, { method: 'DELETE' });
|
||||
}
|
||||
```
|
||||
|
||||
### 6. Handle Modal Timing
|
||||
|
||||
Modals have animation delays:
|
||||
|
||||
```typescript
|
||||
// Playwright (auto-waits)
|
||||
await page.click('.help-btn');
|
||||
await page.waitForSelector('#helpModal', { state: 'visible' });
|
||||
|
||||
// Puppeteer (manual wait)
|
||||
await page.click('.help-btn');
|
||||
await page.waitForSelector('#helpModal', { visible: true });
|
||||
|
||||
// Agent-browser (explicit delay)
|
||||
agentBrowser('click ".help-btn"');
|
||||
await new Promise(r => setTimeout(r, 500));
|
||||
```
|
||||
|
||||
## Key DOM Selectors
|
||||
|
||||
For reference when writing browser tests:
|
||||
|
||||
```
|
||||
.btn-claude // Create Claude session button
|
||||
.btn-settings // Settings button
|
||||
.help-btn // Help button
|
||||
.session-tab // Session tabs
|
||||
.session-tab.active // Active session tab
|
||||
.xterm // Terminal container
|
||||
#helpModal // Help modal
|
||||
#appSettingsModal // Settings modal
|
||||
#sessionOptionsModal // Session Options (same set-* surface)
|
||||
#createCaseModal // Add Case (same set-* surface)
|
||||
.set-rail-item // Rail entry: scrolls in App Settings, switches in the other two
|
||||
.set-section // A settings section (`.hidden` on the inactive ones outside App Settings)
|
||||
.set-row // One setting: label + description left, control right
|
||||
.modal-content // Modal content
|
||||
.modal-close // Modal close button
|
||||
.header-brand .logo // Logo text
|
||||
#versionDisplay // Version display
|
||||
#quickStartCase // Quick start dropdown
|
||||
```
|
||||
|
||||
## Recommendations by Use Case
|
||||
|
||||
| Use Case | Recommended Framework |
|
||||
|----------|----------------------|
|
||||
| CI/CD testing | Playwright |
|
||||
| Chrome-specific features | Puppeteer |
|
||||
| AI agent development | Agent-Browser |
|
||||
| Visual regression | Playwright |
|
||||
| Performance testing | Puppeteer |
|
||||
| Accessibility testing | Agent-Browser |
|
||||
| Cross-browser testing | Playwright |
|
||||
| Quick prototyping | Agent-Browser CLI |
|
||||
|
||||
## Dependencies
|
||||
|
||||
```json
|
||||
{
|
||||
"devDependencies": {
|
||||
"playwright": "^1.58.0",
|
||||
"puppeteer": "^24.36.0",
|
||||
"agent-browser": "^0.6.0"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Install browsers after npm install:
|
||||
```bash
|
||||
npx playwright install chromium
|
||||
```
|
||||
@@ -1,676 +0,0 @@
|
||||
# Claude Code Hooks Reference
|
||||
|
||||
> Official documentation for Claude Code hooks system, extracted from [code.claude.com](https://code.claude.com/docs/en/hooks).
|
||||
|
||||
**Last Updated**: 2026-07-25
|
||||
**Source**: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)
|
||||
|
||||
> This is a maintained summary, not an exhaustive copy of the upstream reference.
|
||||
> Check the source link for event-specific schemas before adding a new hook.
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
Hooks are automated scripts that execute at specific events during your Claude Code session. They allow you to:
|
||||
|
||||
- Validate, modify, or block tool usage
|
||||
- Add context to prompts
|
||||
- Implement custom workflows
|
||||
- Control agent behavior
|
||||
|
||||
---
|
||||
|
||||
## Configuration
|
||||
|
||||
Hooks are configured in settings files:
|
||||
|
||||
| File | Scope |
|
||||
| ----------------------------- | -------------------------- |
|
||||
| `~/.claude/settings.json` | User (global) |
|
||||
| `.claude/settings.json` | Project |
|
||||
| `.claude/settings.local.json` | Local project (gitignored) |
|
||||
| Plugin hook files | Plugin-specific |
|
||||
|
||||
### Basic Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"EventName": [
|
||||
{
|
||||
"matcher": "ToolPattern",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "your-command-here"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Key Fields**:
|
||||
|
||||
- `matcher`: Pattern to match tool names (case-sensitive, supports regex like `Edit|Write` or `*` for all)
|
||||
- `type`: `"command"`, `"http"`, `"mcp_tool"`, `"prompt"`, or `"agent"` where the event supports it
|
||||
- `command`: Bash command to execute
|
||||
- `prompt`: LLM prompt for evaluation (prompt-based hooks only)
|
||||
- `timeout`: Optional timeout in seconds (default: 60)
|
||||
|
||||
---
|
||||
|
||||
## Hook Events
|
||||
|
||||
Claude Code's current event surface is broader than the detailed subset below. In
|
||||
particular, `TeammateIdle` and `TaskCompleted` are supported lifecycle events used
|
||||
by Codeman; they are not stale or plugin-defined event names.
|
||||
|
||||
### PreToolUse
|
||||
|
||||
**When**: After Claude creates tool parameters, before processing the tool call.
|
||||
|
||||
**Use Cases**: Approval, denial, or modification of tool calls.
|
||||
|
||||
**Common Matchers**:
|
||||
|
||||
- `Bash` - Shell commands
|
||||
- `Write` - File writing
|
||||
- `Edit` - File editing
|
||||
- `Read` - File reading
|
||||
- `Agent` - Subagent tasks
|
||||
- `WebFetch`, `WebSearch` - Web operations
|
||||
- `mcp__<server>__<tool>` - MCP tools
|
||||
|
||||
**Output Control**:
|
||||
|
||||
```json
|
||||
{
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PreToolUse",
|
||||
"permissionDecision": "allow|deny|ask",
|
||||
"permissionDecisionReason": "string",
|
||||
"updatedInput": {
|
||||
"field_to_modify": "new value"
|
||||
},
|
||||
"additionalContext": "Context for Claude"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### PermissionRequest
|
||||
|
||||
**When**: When the user is shown a permission dialog.
|
||||
|
||||
**Use Cases**: Auto-approve or deny permissions.
|
||||
|
||||
**Output Control**:
|
||||
|
||||
```json
|
||||
{
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PermissionRequest",
|
||||
"decision": {
|
||||
"behavior": "allow|deny",
|
||||
"updatedInput": {},
|
||||
"message": "deny reason",
|
||||
"interrupt": false
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### PostToolUse
|
||||
|
||||
**When**: Immediately after a tool completes successfully.
|
||||
|
||||
**Use Cases**: Provide feedback, run formatters/linters, log operations.
|
||||
|
||||
**Output Control**:
|
||||
|
||||
```json
|
||||
{
|
||||
"decision": "block",
|
||||
"reason": "Explanation",
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PostToolUse",
|
||||
"additionalContext": "Additional information"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Asynchronous Rewake
|
||||
|
||||
Command hooks can set `"asyncRewake": true` to run asynchronously and wake an
|
||||
idle Claude turn when the hook exits with code 2. The hook's stderr is delivered
|
||||
to Claude as a system reminder. This implies `"async": true`; ordinary async
|
||||
hooks do not wake an idle turn, and their output waits for the next interaction.
|
||||
|
||||
Codeman uses this on `PostToolUse(Bash)`: a self-contained Node helper extracts
|
||||
the background task ID from the Bash result, watches the originating transcript
|
||||
and, for subagents, the top-level parent transcript for the matching completion
|
||||
notification, and exits 2. Claude records a subagent's Bash result in its
|
||||
`subagents/agent-*.jsonl` file but queues completion in the lead session JSONL.
|
||||
The task ID keeps each wake targeted. The helper does not send terminal input,
|
||||
so it cannot submit a user's partially written prompt.
|
||||
|
||||
For script-dispatched Codex work, `codex-run.sh` writes the final response
|
||||
between `CODEMAN_RESULT_BEGIN/END` markers in the background task output. The
|
||||
rewake helper includes a maximum of 64 KiB of that report in its feedback. UI
|
||||
subagent discovery and dispatcher result delivery are separate contracts.
|
||||
|
||||
### Notification
|
||||
|
||||
**When**: When Claude Code sends notifications.
|
||||
|
||||
**Matchers**:
|
||||
|
||||
- `permission_prompt`
|
||||
- `idle_prompt`
|
||||
- `auth_success`
|
||||
- `elicitation_dialog`
|
||||
- `elicitation_complete`
|
||||
- `elicitation_response`
|
||||
|
||||
### UserPromptSubmit
|
||||
|
||||
**When**: When the user submits a prompt, before Claude processes it.
|
||||
|
||||
**Use Cases**: Add context, validate, or block prompts.
|
||||
|
||||
**Output Control**:
|
||||
|
||||
```json
|
||||
{
|
||||
"decision": "block",
|
||||
"reason": "Explanation",
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "UserPromptSubmit",
|
||||
"additionalContext": "My additional context"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Stop
|
||||
|
||||
**When**: When the main Claude Code agent finishes responding.
|
||||
|
||||
**Important**: Does NOT run on user interrupt.
|
||||
|
||||
**Use Cases**: **Ralph Wiggum loops** - block exit and refeed prompt.
|
||||
|
||||
**Output Control**:
|
||||
|
||||
```json
|
||||
{
|
||||
"decision": "block",
|
||||
"reason": "Must provide when blocking"
|
||||
}
|
||||
```
|
||||
|
||||
Or to allow exit:
|
||||
|
||||
```json
|
||||
{
|
||||
"continue": true,
|
||||
"stopReason": "optional message"
|
||||
}
|
||||
```
|
||||
|
||||
**Note**: For Stop events, `"continue": false` takes precedence over `"decision": "block"`.
|
||||
|
||||
### SubagentStop
|
||||
|
||||
**When**: When a subagent (Agent tool call) finishes responding.
|
||||
|
||||
**Use Cases**: Control nested loops, verify subagent output.
|
||||
|
||||
The hook input includes `agent_id`, `agent_transcript_path`, and
|
||||
`last_assistant_message`. Like `Stop`, a command hook can return
|
||||
`{"decision":"block","reason":"..."}` to keep the subagent running and feed
|
||||
the reason back to it.
|
||||
|
||||
Codeman uses this to prevent premature reports from workers that still own live
|
||||
Monitor or background-Bash processes. It derives candidate task IDs from the
|
||||
subagent transcript, but requires a matching live Linux process descriptor for
|
||||
`tasks/<id>.output`; historical task text by itself is not treated as active.
|
||||
|
||||
### TeammateIdle
|
||||
|
||||
**When**: When an agent-team teammate is about to go idle.
|
||||
|
||||
**Use Cases**: Reassign work, continue a teammate loop, or notify an orchestrator.
|
||||
|
||||
**Matcher Support**: None. The hook fires for every occurrence.
|
||||
|
||||
### TaskCompleted
|
||||
|
||||
**When**: When a task is about to be marked completed.
|
||||
|
||||
**Use Cases**: Validate completion or forward team progress to an external UI.
|
||||
|
||||
**Matcher Support**: None. The hook fires for every occurrence.
|
||||
|
||||
### PreCompact
|
||||
|
||||
**When**: Before a compact operation.
|
||||
|
||||
**Matchers**:
|
||||
|
||||
- `manual` - Invoked from `/compact`
|
||||
- `auto` - Invoked from auto-compact
|
||||
|
||||
### SessionStart
|
||||
|
||||
**When**: When Claude Code starts or resumes a session.
|
||||
|
||||
**Matchers**:
|
||||
|
||||
- `startup` - Fresh start
|
||||
- `resume` - From `--resume`, `--continue`, or `/resume`
|
||||
- `clear` - From `/clear`
|
||||
- `compact` - From auto or manual compact
|
||||
|
||||
**Use Cases**: Load development context, set environment variables.
|
||||
|
||||
**Persisting Environment Variables**:
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
if [ -n "$CLAUDE_ENV_FILE" ]; then
|
||||
echo 'export NODE_ENV=production' >> "$CLAUDE_ENV_FILE"
|
||||
echo 'export API_KEY=your-api-key' >> "$CLAUDE_ENV_FILE"
|
||||
fi
|
||||
exit 0
|
||||
```
|
||||
|
||||
**Output Control**:
|
||||
|
||||
```json
|
||||
{
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "SessionStart",
|
||||
"additionalContext": "Context to load"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### SessionEnd
|
||||
|
||||
**When**: When a session ends.
|
||||
|
||||
**Reason Values**:
|
||||
|
||||
- `clear`
|
||||
- `logout`
|
||||
- `prompt_input_exit`
|
||||
- `other`
|
||||
|
||||
**Use Cases**: Cleanup tasks, logging.
|
||||
|
||||
---
|
||||
|
||||
## Hook Input
|
||||
|
||||
Hooks receive JSON via stdin with common fields:
|
||||
|
||||
```json
|
||||
{
|
||||
"session_id": "abc123",
|
||||
"transcript_path": "/path/to/transcript.jsonl",
|
||||
"cwd": "/current/directory",
|
||||
"permission_mode": "default",
|
||||
"hook_event_name": "PreToolUse",
|
||||
"tool_name": "Bash",
|
||||
"tool_input": {},
|
||||
"tool_use_id": "toolu_01ABC123..."
|
||||
}
|
||||
```
|
||||
|
||||
### Tool-Specific Input
|
||||
|
||||
**Bash**:
|
||||
|
||||
```json
|
||||
{
|
||||
"tool_name": "Bash",
|
||||
"tool_input": {
|
||||
"command": "psql -c 'SELECT * FROM users'",
|
||||
"description": "Query the users table",
|
||||
"timeout": 120000
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Write**:
|
||||
|
||||
```json
|
||||
{
|
||||
"tool_name": "Write",
|
||||
"tool_input": {
|
||||
"file_path": "/path/to/file.txt",
|
||||
"content": "file content"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Edit**:
|
||||
|
||||
```json
|
||||
{
|
||||
"tool_name": "Edit",
|
||||
"tool_input": {
|
||||
"file_path": "/path/to/file.txt",
|
||||
"old_string": "original text",
|
||||
"new_string": "replacement text"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Hook Output
|
||||
|
||||
### Exit Codes
|
||||
|
||||
| Code | Behavior |
|
||||
| ----- | --------------------------------------------------------------------- |
|
||||
| 0 | Success. `stdout` processed (shown in verbose or added as context) |
|
||||
| 2 | Blocking error. Only `stderr` used. Blocks tool/prompt based on event |
|
||||
| Other | Non-blocking error. `stderr` shown in verbose, execution continues |
|
||||
|
||||
### JSON Output (Exit Code 0)
|
||||
|
||||
```json
|
||||
{
|
||||
"continue": true,
|
||||
"stopReason": "optional message",
|
||||
"suppressOutput": true,
|
||||
"systemMessage": "optional warning"
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Prompt-Based Hooks
|
||||
|
||||
Prompt and agent handlers are supported by decision-oriented events including
|
||||
`PreToolUse`, `PermissionRequest`, `PostToolUse`, `PostToolUseFailure`,
|
||||
`PostToolBatch`, `UserPromptSubmit`, `Stop`, `SubagentStop`, `TaskCreated`, and
|
||||
`TaskCompleted`. Check the upstream reference before choosing a handler type.
|
||||
|
||||
For example, a Stop event can use LLM-based evaluation:
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"Stop": [
|
||||
{
|
||||
"hooks": [
|
||||
{
|
||||
"type": "prompt",
|
||||
"prompt": "Should Claude stop? Context: $ARGUMENTS\n\nCheck if all tasks are complete.",
|
||||
"timeout": 30
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**LLM Response Format**:
|
||||
|
||||
```json
|
||||
{
|
||||
"ok": true,
|
||||
"reason": "Explanation when ok is false"
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Component-Scoped Hooks
|
||||
|
||||
Hooks can be defined in Skills, Agents, and Slash Commands using frontmatter:
|
||||
|
||||
```markdown
|
||||
---
|
||||
name: secure-operations
|
||||
hooks:
|
||||
PreToolUse:
|
||||
- matcher: 'Bash'
|
||||
hooks:
|
||||
- type: command
|
||||
command: './scripts/security-check.sh'
|
||||
---
|
||||
```
|
||||
|
||||
These hooks:
|
||||
|
||||
- Are scoped to the component's lifecycle
|
||||
- Only run when that component is active
|
||||
- Support all hook events; a subagent-scoped `Stop` is converted to `SubagentStop`
|
||||
|
||||
---
|
||||
|
||||
## MCP Tools
|
||||
|
||||
MCP tools follow the pattern `mcp__<server>__<tool>`:
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"PreToolUse": [
|
||||
{
|
||||
"matcher": "mcp__memory__.*",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "echo 'Memory operation' >> ~/mcp.log"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"matcher": "mcp__.*__write.*",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "/home/user/scripts/validate-mcp-write.py"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Examples
|
||||
|
||||
### Bash Command Validation
|
||||
|
||||
```python
|
||||
#!/usr/bin/env python3
|
||||
import json
|
||||
import re
|
||||
import sys
|
||||
|
||||
VALIDATION_RULES = [
|
||||
(r"\bgrep\b(?!.*\|)", "Use 'rg' instead of 'grep'"),
|
||||
(r"\bfind\s+\S+\s+-name\b", "Use 'rg --files' instead of 'find -name'"),
|
||||
]
|
||||
|
||||
try:
|
||||
input_data = json.load(sys.stdin)
|
||||
except json.JSONDecodeError as e:
|
||||
print(f"Error: {e}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
tool_name = input_data.get("tool_name", "")
|
||||
tool_input = input_data.get("tool_input", {})
|
||||
command = tool_input.get("command", "")
|
||||
|
||||
if tool_name != "Bash" or not command:
|
||||
sys.exit(1)
|
||||
|
||||
issues = []
|
||||
for pattern, message in VALIDATION_RULES:
|
||||
if re.search(pattern, command):
|
||||
issues.append(message)
|
||||
|
||||
if issues:
|
||||
for message in issues:
|
||||
print(f"- {message}", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
```
|
||||
|
||||
### Auto-Approve Documentation Reads
|
||||
|
||||
```python
|
||||
#!/usr/bin/env python3
|
||||
import json
|
||||
import sys
|
||||
|
||||
try:
|
||||
input_data = json.load(sys.stdin)
|
||||
except json.JSONDecodeError as e:
|
||||
print(f"Error: {e}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
tool_name = input_data.get("tool_name", "")
|
||||
tool_input = input_data.get("tool_input", {})
|
||||
|
||||
if tool_name == "Read":
|
||||
file_path = tool_input.get("file_path", "")
|
||||
if file_path.endswith((".md", ".mdx", ".txt", ".json")):
|
||||
output = {
|
||||
"hookSpecificOutput": {
|
||||
"hookEventName": "PreToolUse",
|
||||
"permissionDecision": "allow",
|
||||
"permissionDecisionReason": "Documentation file auto-approved"
|
||||
},
|
||||
"suppressOutput": True
|
||||
}
|
||||
print(json.dumps(output))
|
||||
sys.exit(0)
|
||||
|
||||
sys.exit(0)
|
||||
```
|
||||
|
||||
### Post-Write Formatter
|
||||
|
||||
```json
|
||||
{
|
||||
"hooks": {
|
||||
"PostToolUse": [
|
||||
{
|
||||
"matcher": "Edit|Write",
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "npx prettier --write \"$TOOL_INPUT_FILE_PATH\" 2>/dev/null || true"
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Ralph Wiggum Stop Hook
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
# ralph-stop-hook.sh
|
||||
|
||||
STATE_FILE=".claude/ralph-loop.local.md"
|
||||
|
||||
# Check if state file exists
|
||||
if [ ! -f "$STATE_FILE" ]; then
|
||||
exit 0 # No active loop, allow exit
|
||||
fi
|
||||
|
||||
# Read state from YAML frontmatter
|
||||
ENABLED=$(grep -m1 "^enabled:" "$STATE_FILE" | cut -d' ' -f2)
|
||||
ITERATION=$(grep -m1 "^iteration:" "$STATE_FILE" | cut -d' ' -f2)
|
||||
MAX_ITER=$(grep -m1 "^max-iterations:" "$STATE_FILE" | cut -d' ' -f2)
|
||||
PROMISE=$(grep -m1 "^completion-promise:" "$STATE_FILE" | cut -d' ' -f2-)
|
||||
|
||||
# Check if disabled
|
||||
if [ "$ENABLED" = "false" ]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Check max iterations
|
||||
if [ -n "$MAX_ITER" ] && [ "$ITERATION" -ge "$MAX_ITER" ]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# Check for completion promise in output
|
||||
if [ -n "$PROMISE" ]; then
|
||||
if echo "$CLAUDE_OUTPUT" | grep -q "<promise>$PROMISE</promise>"; then
|
||||
exit 0
|
||||
fi
|
||||
fi
|
||||
|
||||
# Block exit, increment iteration
|
||||
NEW_ITER=$((ITERATION + 1))
|
||||
sed -i "s/^iteration:.*/iteration: $NEW_ITER/" "$STATE_FILE"
|
||||
|
||||
# Output block decision
|
||||
echo '{"decision": "block", "reason": "Completion promise not found. Iteration '"$NEW_ITER"'."}'
|
||||
exit 0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Environment Variables
|
||||
|
||||
| Variable | Description |
|
||||
| -------------------- | ------------------------------------------------ |
|
||||
| `CLAUDE_PROJECT_DIR` | Project root directory |
|
||||
| `CLAUDE_CODE_REMOTE` | `"true"` for web, empty for CLI |
|
||||
| `CLAUDE_ENV_FILE` | Path to write persistent env vars (SessionStart) |
|
||||
|
||||
---
|
||||
|
||||
## Debugging
|
||||
|
||||
Use `claude --debug` to see detailed hook execution:
|
||||
|
||||
```
|
||||
[DEBUG] Executing hooks for PostToolUse:Write
|
||||
[DEBUG] Found 1 hook matchers in settings
|
||||
[DEBUG] Matched 1 hooks for query "Write"
|
||||
[DEBUG] Executing hook command: <command> with timeout 60000ms
|
||||
[DEBUG] Hook command completed with status 0: <stdout>
|
||||
```
|
||||
|
||||
Use `/hooks` command to view registered hooks and make changes.
|
||||
|
||||
---
|
||||
|
||||
## Execution Details
|
||||
|
||||
- **Timeout**: 60-second default per hook, configurable
|
||||
- **Parallelization**: All matching hooks run in parallel
|
||||
- **Deduplication**: Identical commands deduplicated automatically
|
||||
- **Matchers**: Only apply to tool-based hooks (PreToolUse, PostToolUse, PostToolUseFailure, PermissionRequest)
|
||||
|
||||
---
|
||||
|
||||
## Security Best Practices
|
||||
|
||||
1. **Validate and sanitize inputs** - Never trust input data blindly
|
||||
2. **Always quote shell variables** - Use `"$VAR"` not `$VAR`
|
||||
3. **Block path traversal** - Check for `..` in file paths
|
||||
4. **Use absolute paths** - Specify full paths for scripts (use `$CLAUDE_PROJECT_DIR`)
|
||||
5. **Skip sensitive files** - Avoid `.env`, `.git/`, keys, etc.
|
||||
|
||||
---
|
||||
|
||||
_Source: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)_
|
||||
@@ -1,121 +0,0 @@
|
||||
# Claude voice dictation in Codeman
|
||||
|
||||
Wire Codeman's existing mic button to the same speech-to-text service Claude Code's own
|
||||
`/voice` mode uses, so dictation works with **no third-party API key** for anyone already
|
||||
signed in to Claude Code on the server.
|
||||
|
||||
## Why the CLI's own voice mode cannot be reused directly
|
||||
|
||||
Claude Code 2.1.x ships voice input: `/voice hold|tap|off` arms it, the CLI opens the
|
||||
**host's** microphone (native `audio-capture-napi`, falling back to `sox`/`arecord` on Linux
|
||||
after probing `/proc/asound/cards`), streams PCM upstream and types the transcript into its
|
||||
own composer.
|
||||
|
||||
Every part of that is on the wrong machine for Codeman. The CLI runs inside a tmux pane on
|
||||
the server, which is typically headless and has no sound card at all, while the human is in
|
||||
a browser on a phone somewhere else. Toggling `/voice` in the pane from Codeman would arm a
|
||||
microphone nobody is sitting in front of. So Codeman keeps capturing audio in the browser,
|
||||
where the user actually is, and only borrows the CLI's **transcription backend**.
|
||||
|
||||
## The backend, as the CLI uses it
|
||||
|
||||
Extracted from the 2.1.226 binary (`connectVoiceStream`):
|
||||
|
||||
| | |
|
||||
| --- | --- |
|
||||
| URL | `wss://api.anthropic.com/api/ws/speech_to_text/voice_stream` |
|
||||
| Query | `encoding=linear16`, `sample_rate=16000`, `channels=1`, `endpointing_ms=300`, `utterance_end_ms=1000`, `language=<lang>`, `use_conversation_engine=true`, `stt_provider=deepgram-nova3` |
|
||||
| Headers | `Authorization: Bearer <Claude Code OAuth access token>`, `User-Agent`, `x-app: cli`, `anthropic-client-platform`, optional `x-config-keyterms` |
|
||||
| Audio | raw binary frames, PCM signed 16-bit little-endian, 16 kHz, mono |
|
||||
| Keepalive | `{"type":"KeepAlive"}` on open, then every 8 s |
|
||||
| Finalize | `{"type":"CloseStream"}`, then wait for the endpoint frame |
|
||||
| Downstream | `{"type":"TranscriptText"\|"TranscriptInterim","data":"…"}` (running interim), `{"type":"TranscriptEndpoint"}` (promotes the pending interim to final), `{"type":"TranscriptError",…}`, `{"type":"error","message":…}` |
|
||||
|
||||
Deepgram Nova-3 runs server-side, so the Deepgram-quality result arrives without a Deepgram
|
||||
account. Verified against the live endpoint before this design was written: connect, stream
|
||||
PCM, receive interims and an endpoint frame.
|
||||
|
||||
## Architecture
|
||||
|
||||
The browser cannot call that endpoint itself: it would need the OAuth bearer token in page
|
||||
JavaScript (and CORS would refuse anyway). So the audio goes browser → Codeman → Anthropic,
|
||||
and Codeman is the only thing that ever touches the token.
|
||||
|
||||
```
|
||||
mic → AudioWorklet (Float32 → PCM16 @16 kHz)
|
||||
→ wss://<codeman>/ws/voice/stream [cookie/basic auth, Origin+Host guarded]
|
||||
→ VoiceStreamRelay (reads ~/.claude/.credentials.json per connect)
|
||||
→ wss://api.anthropic.com/api/ws/speech_to_text/voice_stream
|
||||
← {"t":"transcript","text":…,"final":…} → existing _insertText() path
|
||||
```
|
||||
|
||||
Nothing about the insert path changes: the transcript lands in the same preview overlay,
|
||||
the same direct/compose insert modes, the same green Send button.
|
||||
|
||||
### Server pieces
|
||||
|
||||
- **`src/claude-credentials.ts`** — locate and parse the Claude Code OAuth credentials.
|
||||
`parseClaudeCredentials()` is pure (JSON string + `now` → status) and unit-tested;
|
||||
`readClaudeOAuthToken()` wraps it with IO: `$CLAUDE_CONFIG_DIR/.credentials.json` or
|
||||
`~/.claude/.credentials.json`, and on macOS the login keychain
|
||||
(`security find-generic-password -s "Claude Code-credentials"`).
|
||||
**Read-only, always.** Codeman never writes credentials and never refreshes the token: a
|
||||
refresh rotates the refresh token, and racing Claude Code's own refresh could sign the
|
||||
user out of their CLI. An expired token surfaces as a plain "run a Claude session to
|
||||
refresh" error instead.
|
||||
The token is never logged, never returned by any endpoint, and never sent to the browser.
|
||||
|
||||
- **`src/web/voice-stream.ts`** — pure `buildVoiceStreamUrl()` / `buildVoiceStreamHeaders()` /
|
||||
`sanitizeKeyterms()` (ASCII-only, deduped, 1024-char cap, mirroring the CLI), plus
|
||||
`VoiceStreamRelay`, which owns one upstream socket: keepalive timer, audio passthrough,
|
||||
transcript translation, finalize, and the caps below.
|
||||
|
||||
- **`src/web/routes/voice-routes.ts`**
|
||||
- `GET /api/voice/status` → `{ available, reason, subscriptionType?, expiresAt? }`. Never
|
||||
the token. `available:false` with a machine-readable `reason` (`disabled`, `no-credentials`,
|
||||
`expired`) is what the settings row and the provider resolver read.
|
||||
- `GET /ws/voice/stream?language=&keyterms=` → the relay. Same upgrade guard as
|
||||
`/ws/sessions/:id/terminal`: allowed Host, same-site Origin, and the global auth hook has
|
||||
already run on the handshake.
|
||||
|
||||
Caps, because an open mic is an open pipe: one stream per connection, `MAX_VOICE_STREAMS`
|
||||
concurrent server-wide, a hard `MAX_STREAM_MS` per stream, and a per-frame size cap. A tab
|
||||
left recording cannot bill an unbounded amount of upstream audio.
|
||||
|
||||
### Frontend pieces
|
||||
|
||||
- **`voice-pcm-worklet.js`** — an `AudioWorkletProcessor` converting Float32 blocks to PCM16
|
||||
and posting ~256 ms frames back. `MediaRecorder` cannot produce raw PCM, which is why the
|
||||
existing Deepgram path (container audio, auto-detected) cannot be reused as-is. Falls back
|
||||
to `ScriptProcessorNode` where AudioWorklet is unavailable.
|
||||
- **`ClaudeVoiceProvider`** in `voice-input.js` — mirrors `DeepgramProvider`'s shape
|
||||
(`start({language, keyterms, onStream, onResult, onError, onEnd})`) so `VoiceInput` treats
|
||||
the three providers uniformly.
|
||||
- **Provider resolution** — new `voiceSettings.provider`: `auto` (default) | `claude` |
|
||||
`deepgram` | `webspeech`. `auto` picks Claude when `/api/voice/status` reports it
|
||||
available, else Deepgram when a key is set, else Web Speech. Pinning a provider always
|
||||
wins, so an existing Deepgram user can keep exactly what they have.
|
||||
|
||||
### Settings
|
||||
|
||||
- `claudeVoiceEnabled` — synced, **default OFF**, gating the whole server side. Off is the
|
||||
honest default: turning it on means this machine's Claude subscription starts paying for
|
||||
transcription for whoever can reach the UI, and the audio goes to Anthropic rather than to
|
||||
wherever it went before. One switch in Settings → Voice, and the mic works with no key.
|
||||
- `voiceSettings.provider` — per the resolution table above; joins the existing synced
|
||||
`voiceSettings` object.
|
||||
|
||||
## Things worth knowing
|
||||
|
||||
- **This uses an undocumented endpoint with subscription credentials.** It is the user's own
|
||||
token, on the user's own machine, driving the user's own Claude Code install, but it is not
|
||||
a published API and Anthropic can change or restrict it. Default-OFF is deliberate; the
|
||||
Deepgram and Web Speech paths stay untouched as the supported fallbacks.
|
||||
- **Multi-user mode**: every user's dictation would run on the server owner's Claude
|
||||
credentials, exactly as every user's *sessions* already run on them. Consistent, but worth
|
||||
stating out loud in the settings copy.
|
||||
- **Token lifetime** is about 8 hours, refreshed by Claude Code itself whenever it runs. The
|
||||
relay re-reads the file on every connect rather than caching, so a refresh is picked up on
|
||||
the next press of the mic.
|
||||
- **HTTPS or localhost**: `getUserMedia` needs a secure context. Prod is HTTPS behind
|
||||
`tailscale serve`, so this is already satisfied; the existing error copy covers the rest.
|
||||
@@ -1,314 +0,0 @@
|
||||
# CLI management Settings UI + write API — plan
|
||||
|
||||
> Tracked separately from `DEPLOYMENT_PLAN.md` (PR B2, merged) and `docs/copilot-integration-plan.md`
|
||||
> (parked). This is "PR C" from the original #343 review: *"settings UI + write endpoints +
|
||||
> auto-install, once we've settled the trust model... I want to make that call on its own, not
|
||||
> inside a 100-file diff."*
|
||||
>
|
||||
> **Phase 0 is CLOSED as of 2026-09-21** — all three original pieces are IN SCOPE (expanded from
|
||||
> this plan's first draft, which recommended #2/#3 as separate/out-of-scope; the user chose full
|
||||
> scope instead, with the risk called out explicitly for #3 before confirming). See "Decisions"
|
||||
> below for the full record.
|
||||
|
||||
## Status as of 2026-09-22
|
||||
|
||||
**Phases 1–6 are ALL IMPLEMENTED** (commits `da07b38c` "add cliManagementEnabled flag and GET
|
||||
/api/clis" and `db4557d9` "Phases 3-6 - write API + custom entries + Settings UI", both on this
|
||||
branch, `feat/cli-management`). Confirmed present in the tree: `cliManagementEnabled` in
|
||||
`SettingsUpdateSchema`; `GET /api/clis`, `PUT /api/clis/:id`, `POST /api/clis/:id/install`,
|
||||
`POST /api/clis`, `PUT /api/clis/custom/:id`, `DELETE /api/clis/:id` in
|
||||
`src/web/routes/cli-registry-routes.ts`; the `shell`/`claude` `UNDISABLEABLE_IDS` backend guard;
|
||||
`isAdmin(req)` gating on both the list and write routes; `appendAdminAudit` wired into the install
|
||||
route; tmp+rename+`0o600` writes in `registry-writer.ts`; the full Settings UI (row list, toggle,
|
||||
Install button, custom-entry create/edit/delete form) in `settings-ui.js` + `index.html`.
|
||||
`test/routes/cli-registry-routes.test.ts` (425 lines) and `test/cli-registry-no-id-branching.test.ts`
|
||||
cover it. This status section, plus the fix and gap below, is the one piece of that work done in
|
||||
a *different* session from the one that wrote Phases 1–6 — reviewed by reading the diff and
|
||||
verifying each claim against the actual routes/tests, not by re-implementing anything.
|
||||
|
||||
### Gotcha found and fixed (commit `0c77dd0a`)
|
||||
|
||||
**Toggling a CLI off in Settings had no effect anywhere except the Settings row itself.**
|
||||
`window.__codemanCliAvailable` — the flag `isCliAvailable()` reads client-side to gate the
|
||||
welcome-screen buttons, the Run-menu dropdown and the mobile overview — is injected **once**, at
|
||||
initial page render (`server.ts`), built purely from each CLI's own installed-on-PATH resolver
|
||||
(`isClaudeAvailable()` etc.), with **no reference to the registry's `enabled` flag at all**. So
|
||||
disabling a CLI here updated its own row and nothing else — every launch surface kept offering it,
|
||||
both live and after a full page reload, since even a *fresh* render never consulted the registry.
|
||||
Root-caused and reported by the user testing the live feature ("toggle those off, they still
|
||||
appear in that menu and on the front main screen").
|
||||
|
||||
Fixed two places:
|
||||
- `server.ts`: after building `available`, intersect the nine real `SessionMode` ids against
|
||||
`enabledClis()`. `git`/`cloudflared` (utility binaries, not CLI registry entries) and
|
||||
`deepseekBinary` (a secondary installed-only flag for the "add a profile" affordance) are
|
||||
deliberately left alone — they were never registry-gated to begin with.
|
||||
- `settings-ui.js`: `toggleCliEnabled()` now patches `window.__codemanCliAvailable` in place and
|
||||
refreshes the welcome screen, the mobile overview and an already-open Run menu, mirroring the
|
||||
existing `installDeepSeekProfile()` pattern for the same "injected once, needs an explicit
|
||||
patch" reason — the server-side fix alone still left every surface stale until the next reload.
|
||||
|
||||
New test in `test/render-index-html.test.ts`: an installed-but-disabled CLI (codex, forced via
|
||||
`clis.json` + `reloadCliRegistry()`) reads as unavailable, while an installed-and-enabled one
|
||||
(claude) is unaffected by the override.
|
||||
|
||||
**Verified on the Debian devbox** (`codeman-devbox`, real tmux — this sandbox has none and
|
||||
`WebServer`'s constructor hard-requires it): typecheck clean, the new test passes (17/17 in
|
||||
`render-index-html.test.ts`), the CLI-registry suites pass (86/86), and the **full CI gate is
|
||||
green — 415 test files, 7855 tests, 0 failures**.
|
||||
|
||||
### Launch-surface registry integration — completed
|
||||
|
||||
The welcome screen, desktop Run menu and mobile Run picker now use the same injected CLI catalog.
|
||||
Every enabled registry entry is rendered; unavailable binaries remain hidden as before. Settings
|
||||
updates the catalog and availability flags in place after enable/disable, create, edit or delete,
|
||||
so the launch surfaces update without a page reload. A custom entry uses the generic quick-start
|
||||
path, while stock entries retain their existing per-CLI launch settings.
|
||||
|
||||
Not otherwise re-verified line-by-line against every Phase 1–6 checklist item below (e.g. the
|
||||
exact wording of toasts, the "same PR" sequencing notes) — the checklists are left as originally
|
||||
written; treat the **Status** section above as authoritative for what exists.
|
||||
|
||||
---
|
||||
|
||||
## Background
|
||||
|
||||
`src/config/cli-registry/registry.ts` is READ-ONLY today, and says so in its own header comment:
|
||||
|
||||
> "⚠️ READ-ONLY. Nothing in this module writes, creates or migrates the file... there is no
|
||||
> settings UI and no write API yet... A `seededStockIds` ratchet belongs with the write API that
|
||||
> needs it."
|
||||
|
||||
Confirmed on `master` (2026-09-21): no `/api/clis` route exists at all (read or write);
|
||||
`~/.codeman/clis.json` is hand-edit-only; `resolveInstallCommandForPlatform()` is documented
|
||||
"Display text only — never executed" — nothing runs an install command server-side today. The
|
||||
original #343 review flagged the opposite (`spawn(command, {shell: true})`, `env.allowedPrefixes`
|
||||
contributed from a write) as needing its own trust-model decision; that decision was never made
|
||||
after the split, just dropped. This plan makes it.
|
||||
|
||||
**Closest existing precedent, and the template this plan follows for the read/write API**:
|
||||
`src/web/routes/custom-model-routes.ts` + `src/custom-model-hosts.ts` (#393/#430/#459) — a small
|
||||
per-item JSON store, Settings-UI-driven, admin-gated in multi-user mode, tmp+rename+0600 writes.
|
||||
|
||||
**Precedent for the new master feature flag (Phase 1)**: `customModelEndpointsEnabled` —
|
||||
`z.boolean().optional()` in `SettingsUpdateSchema` (`schemas.ts:1319`), a checkbox read/written by
|
||||
id in `openAppSettings()`/`saveAppSettings()` (`settings-ui.js:401`/`:2120`). SYNCED, not
|
||||
per-device (present in the schema, absent from `displayKeys`), default OFF.
|
||||
|
||||
**Spec refs for the whole plan:**
|
||||
- `src/config/cli-registry/registry.ts` — the read path; `resolveRegistry()`'s merge semantics
|
||||
(`deepMerge`, `UNMERGEABLE_KEYS`) apply unchanged to whatever this plan writes
|
||||
- `docs/cli-registry.md` — registry shape, "The override file", "Arg-template safety" (the four
|
||||
layers Phase 5's custom-entry validation must not weaken), "Adding a CLI" (the 5-step recipe a
|
||||
custom entry does NOT get to skip just because it arrives via UI instead of a stock.ts edit)
|
||||
- `src/web/routes/custom-model-routes.ts` + `src/custom-model-hosts.ts` — read/write API template
|
||||
- `docs/multi-user-plan.md`, `docs/security-architecture.md` — admin-gating conventions
|
||||
- `CLAUDE.md` §Multi-user mode, §"Settings surface", §"Per-device vs synced settings"
|
||||
|
||||
---
|
||||
|
||||
## Decisions (Phase 0, closed 2026-09-21)
|
||||
|
||||
1. **Enable/disable a stock CLI's `enabled` flag** — IN SCOPE. Plus a **master feature flag**
|
||||
(`cliManagementEnabled`, synced, default OFF) gating the whole Settings UI section's visibility,
|
||||
matching this codebase's standing convention for new admin-facing surfaces.
|
||||
2. **Auto-install** (stock CLIs' already-shipped, already-vetted install commands) — IN SCOPE,
|
||||
same PR.
|
||||
3. **Custom CLI entries via the UI** — IN SCOPE, **typed-argv only**: a custom entry goes through
|
||||
the exact same schema/argv-safety path stock entries do (named token patterns, no raw shell-text
|
||||
field). Its install command stays **display-only text**, same as every stock entry today — Phase
|
||||
4's auto-install NEVER executes a custom entry's install command, only a stock one's. This is
|
||||
the one place scope was deliberately narrowed relative to what was agreed in principle, because
|
||||
`docs/cli-registry.md`'s arg-template-safety section exists specifically to keep config free of
|
||||
shell text, and a free-text install command for a user-defined entry would reopen exactly that.
|
||||
4. **`shell`/`claude` un-disableable** — enforced at the **backend**, not just the UI (a
|
||||
frontend-only guard is bypassable with curl).
|
||||
5. **Non-admin visibility in multi-user mode** — the CLI-management Settings section is **hidden
|
||||
entirely** for a non-admin, not shown-empty.
|
||||
6. **`seededStockIds` ratchet** — not needed. `deepMerge()` only overrides a key the file actually
|
||||
sets, so a CLI absent from `clis.json.clis` always falls through to its stock `enabled` value
|
||||
with no special-casing. (Carried over from the first draft, not re-litigated.)
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — Master feature flag: `cliManagementEnabled`
|
||||
|
||||
**Status:** DONE (commit `da07b38c`) — verified present in `SettingsUpdateSchema`, `index.html`,
|
||||
`openAppSettings()`/`saveAppSettings()`.
|
||||
|
||||
**Spec refs:**
|
||||
- `schemas.ts:1319` (`customModelEndpointsEnabled`) — the exact pattern to mirror: `z.boolean().optional()`
|
||||
in `SettingsUpdateSchema`
|
||||
- `settings-ui.js:401`/`:2120` — checkbox read/write by id in `openAppSettings()`/`saveAppSettings()`
|
||||
- `CLAUDE.md` §"Adding Features" → "App setting" — decide per-device vs synced FIRST (this one is
|
||||
synced: a feature toggle, not a display preference) and add to `displayKeys` NEVER for a synced
|
||||
setting
|
||||
|
||||
**Checklist:**
|
||||
- [x] Add `cliManagementEnabled: z.boolean().optional()` to `SettingsUpdateSchema`
|
||||
- [x] Add the checkbox to `index.html`'s `#settings-clis` section, above where Phase 6's per-CLI
|
||||
list will render — reads/writes via `openAppSettings()`/`saveAppSettings()` by id, same as
|
||||
`customModelEndpointsEnabled`
|
||||
- [x] `readCliManagementEnabled()` helper (mirrors `readCustomModelEndpointsEnabled()` in
|
||||
`custom-model-routes.ts:609`) for the route file(s) in Phases 2-5 to gate on
|
||||
- [x] When OFF: `GET /api/clis` still exists but the Settings UI section stays hidden
|
||||
(`applyCliManagementVisibility()`); the write endpoints reject (see Phase 3)
|
||||
|
||||
**Verify:** `npm run typecheck` passes; a unit test confirms `SettingsUpdateSchema` accepts/rejects
|
||||
the field correctly; toggling it in a fresh browser profile shows/hides the Settings section with
|
||||
no server restart.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — Read endpoint: `GET /api/clis`
|
||||
|
||||
**Status:** DONE (commit `da07b38c`) — verified present in `src/web/routes/cli-registry-routes.ts`.
|
||||
|
||||
**Spec refs:**
|
||||
- `src/web/routes/custom-model-routes.ts:730` (`GET /api/model-endpoints`) — multi-user read
|
||||
gating: empty list for a non-admin, never a 403
|
||||
- `src/config/cli-registry/registry.ts` — `listClis()` (every entry, including disabled stock
|
||||
ones — this is an admin/settings surface, unlike `enabledClis()`)
|
||||
- `window.__codemanCliAvailable`'s resolvers (`isClaudeAvailable()` etc.) — candidate `installed`
|
||||
source; confirm whether to reuse directly or the response needs its own probe (Open Question 4,
|
||||
carried from the first draft — still genuinely open, decide during this phase not before)
|
||||
|
||||
**Checklist:**
|
||||
- [x] New route file `cli-registry-routes.ts`
|
||||
- [x] Response excludes `launch`/`env`/`capabilities`/`overlays`/`discovery`
|
||||
- [x] `isMultiUserMode() && !isAdmin(req)` → `[]`
|
||||
- [x] Unit tests in `test/routes/cli-registry-routes.test.ts` (admin/non-admin/single-user,
|
||||
disabled stock CLI still present)
|
||||
|
||||
**Verify:** `npm test -- test/routes/cli-registry-routes.test.ts` passes; `curl localhost:3000/api/clis | jq`
|
||||
shows every stock CLI including disabled ones.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3 — Write endpoint: `PUT /api/clis/:id` (stock enable/disable)
|
||||
|
||||
**Status:** DONE (commit `db4557d9`) — `UNDISABLEABLE_IDS`, admin gate, tmp+rename+0600 all
|
||||
confirmed present.
|
||||
|
||||
**Spec refs:**
|
||||
- `src/web/routes/custom-model-routes.ts:753` + `src/custom-model-hosts.ts:91` — write-path
|
||||
template: `adminOnly` gate, read-modify-write the WHOLE file, tmp+rename+0600
|
||||
- `registry.ts:47` (`filePath()` = `dataPath(...)`) and `reloadCliRegistry()` — write to the same
|
||||
resolved path, invalidate the cache on every successful write or the change is invisible until
|
||||
restart
|
||||
|
||||
**Checklist:**
|
||||
- [x] Body: `{ enabled: boolean }`. Zod schema in `schemas.ts`
|
||||
- [x] Gate order: `cliManagementEnabled` → `adminOnly` → shell/claude guard → stock-only guard
|
||||
- [x] Rejects disabling `shell` or `claude` (`UNDISABLEABLE_IDS`)
|
||||
- [x] Rejects a write for an id that isn't a stock CLI
|
||||
- [x] Deep-merges `{ clis: { [id]: { enabled } } }`, preserving other override keys
|
||||
- [x] tmp+rename+0600 write, `reloadCliRegistry()` on success
|
||||
- [x] Unit tests (`test/routes/cli-registry-routes.test.ts`)
|
||||
|
||||
**Verify:** `npm test` full gate green; `curl -X PUT localhost:3000/api/clis/grok -d '{"enabled":false}'`
|
||||
then `GET /api/clis` shows the change with no restart; same against `shell`/`claude` returns an
|
||||
error and changes nothing; `ls -la ~/.codeman/clis.json` shows mode 0600.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4 — Auto-install: `POST /api/clis/:id/install` (stock CLIs only)
|
||||
|
||||
**Status:** DONE (commit `db4557d9`) — route present, `appendAdminAudit` wired in.
|
||||
|
||||
**Spec refs:**
|
||||
- `registry.ts:231` (`resolveInstallCommandForPlatform`) — currently "Display text only — never
|
||||
executed"; this phase is what changes that, for stock entries only, with Decision 2's sign-off
|
||||
- Original #343 review's exact concern re: `env.allowedPrefixes` contributed from a write — stays
|
||||
out of scope; this phase only ever runs a command, never touches the env allowlist
|
||||
|
||||
**Checklist:**
|
||||
- [x] Separate endpoint from Phase 3's toggle
|
||||
- [x] Gate order: `cliManagementEnabled` → `adminOnly` → stock-entry-only guard
|
||||
- [x] `resolveInstallCommandForPlatform(entry)` for the target
|
||||
- [x] Bounded execution (timeout, captured stdout/stderr)
|
||||
- [x] Does NOT auto-enable on successful install
|
||||
- [x] Audit-logged via `appendAdminAudit`
|
||||
- [x] Unit tests
|
||||
|
||||
**Verify:** a real install triggered via the endpoint against a CLI not currently installed,
|
||||
`GET /api/clis`'s `installed` field flips true with no restart; audit log entry present; attempting
|
||||
install against a custom entry's id fails with a clear error; full CI gate green.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5 — Custom CLI entries: create / update / delete via API
|
||||
|
||||
**Status:** DONE (commit `db4557d9`) — `POST /api/clis`, `PUT /api/clis/custom/:id`,
|
||||
`DELETE /api/clis/:id` all present. Open Question 2 resolved: a **separate** endpoint
|
||||
(`PUT /api/clis/custom/:id`), not Phase 3's `PUT /api/clis/:id` widened.
|
||||
|
||||
**Spec refs:**
|
||||
- `docs/cli-registry.md` §"Arg-template safety" (all four layers), §"Adding a CLI" (the 5-step
|
||||
recipe) — a custom entry created via this API must satisfy the SAME schema (`CliEntrySchema`)
|
||||
every stock entry does; there is no relaxed path for UI-originated entries
|
||||
- `registry.ts`'s `resolveRegistry()` — the custom-entry branch (`stock: false`, dropped with a
|
||||
warning on validation failure, never falls back silently) already exists and is unchanged by
|
||||
this phase; this phase only adds a way to WRITE what that branch reads
|
||||
|
||||
**Checklist:**
|
||||
- [x] `POST /api/clis` (create), full `CliEntrySchema` validation
|
||||
- [x] `PUT /api/clis/custom/:id` (update) — separate endpoint from Phase 3's stock toggle
|
||||
- [x] `DELETE /api/clis/:id` refuses for any stock id
|
||||
- [x] `id` collision check against existing stock ids
|
||||
- [x] `discovery.install.command` on a custom entry stays DISPLAY-ONLY
|
||||
- [x] Same tmp+rename+0600 write pattern, `reloadCliRegistry()` on every successful mutation
|
||||
- [x] Unit tests
|
||||
|
||||
**Verify:** `npm test` full gate green; create a custom entry via curl, confirm it appears in
|
||||
`GET /api/clis` — **confirm it appears in the Run menu is UNVERIFIED and currently FALSE, see
|
||||
"Outstanding" above**; delete it, confirm it's gone and `clis.json` no longer references it.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6 — Settings UI
|
||||
|
||||
**Status:** DONE (commit `db4557d9`) — `#cliListGroup`, row rendering, toggle, Install button,
|
||||
custom-entry create/edit/delete form all present in `settings-ui.js`/`index.html`. Manual browser
|
||||
verification per the phase's own "Verify" step (flag on/off, non-admin hidden, toggle stops the
|
||||
Run menu offering a CLI, create/enable/launch a custom entry, delete it, shell/claude undisableable)
|
||||
has **not** been re-run in this session — the toggle→Run-menu leg specifically was BROKEN until the
|
||||
gotcha fix above, and the create→launch leg for a custom entry is the confirmed gap in
|
||||
"Outstanding".
|
||||
|
||||
**Spec refs:**
|
||||
- `index.html:2357` (`#settings-clis`) — the existing home; Phase 1's master toggle at the top,
|
||||
then the per-CLI list, then (if `cliManagementEnabled`) a "custom CLI" creation form, all above
|
||||
the existing Codex-only groups
|
||||
- `CLAUDE.md` §"Settings surface" — App Settings scrolls, it does not tab-switch
|
||||
- `admin-ui.js` — pattern for an admin-only-VISIBLE section (not just admin-only-writable),
|
||||
needed here per Decision 5
|
||||
|
||||
**Checklist:**
|
||||
- [x] Whole section hidden when `cliManagementEnabled` is OFF, and separately hidden for a
|
||||
non-admin in multi-user mode (`_applyCliManagementAdminGate`)
|
||||
- [x] Fetches `GET /api/clis` when the section becomes visible; renders one row per CLI
|
||||
- [x] Stock rows: enabled toggle only; `shell`/`claude` rows show the toggle disabled/greyed
|
||||
- [x] Custom rows: enabled toggle plus edit/delete affordances
|
||||
- [x] "Add custom CLI" form (id/label/badge/binary/argv)
|
||||
- [x] Toggle/edit/delete update the row in place
|
||||
|
||||
**Verify:** manual browser test per `CLAUDE.md`'s "Always Test Before Deploying" rule — **not yet
|
||||
re-run end-to-end in this session**; do this before considering the feature ready to ship, and
|
||||
expect the custom-entry-launch step to fail until the Outstanding gap above is closed.
|
||||
|
||||
---
|
||||
|
||||
## Remaining Open Questions
|
||||
|
||||
1. **Phase 2's `installed` source** — resolved: reuses `window.__codemanCliAvailable`'s existing
|
||||
resolvers via `GET /api/clis`'s own probe (confirmed by reading the route).
|
||||
2. **Phase 5's `PUT` endpoint shape** — resolved: a **separate** endpoint
|
||||
(`PUT /api/clis/custom/:id`), not Phase 3's toggle route widened.
|
||||
3. **Sequencing against the parked Copilot plan** — unchanged, still not blocking.
|
||||
4. **NEW: custom-CLI Run-menu integration** — see "Outstanding" above. Not decided or started.
|
||||
|
||||
---
|
||||
|
||||
Implementation is underway (see Status above); this line is left for history rather than removed —
|
||||
the plan was originally approved before Phases 1–6 landed.
|
||||
@@ -1,240 +0,0 @@
|
||||
# The CLI registry
|
||||
|
||||
Every run mode Codeman can launch — Claude Code, Terminal/Shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek Harness and OMP — is a `CliEntry`: a data record describing how to find the binary, how to build its command line, what environment it needs, and what it can do. Code that used to ask "which CLI is this?" asks the entry instead.
|
||||
|
||||
## Where it lives
|
||||
|
||||
| File | What it holds |
|
||||
| ------------- | ------------------------------------------------------------------------------------------------- |
|
||||
| `types.ts` | The `CliEntry` interface and everything under it. Read this first. |
|
||||
| `stock.ts` | The shipped catalog. **The only file allowed to name a CLI id.** |
|
||||
| `schema.ts` | Zod validation, including the cross-field checks that reject an incoherent entry at LOAD time. |
|
||||
| `argv.ts` | The argv engine: the only code that turns typed tokens into a command string. |
|
||||
| `patterns.ts` | The NAMED value patterns (`model`, `uuid`, `path-segment`, …) and the regex-compilation guard. |
|
||||
| `profiles.ts` | The names of behaviours that genuinely need code, kept import-free so `schema.ts` can validate one. |
|
||||
| `registry.ts` | Loading, merging `~/.codeman/clis.json`, and the accessors (`getCli`, `enabledClis`). |
|
||||
|
||||
`src/session-cli-registry-bridge.ts` maps the legacy per-mode option bag onto the engine, and `src/utils/cli-resolver.ts` / `src/utils/cli-launcher.ts` do registry-driven binary resolution and launcher-profile dispatch.
|
||||
|
||||
## The override file
|
||||
|
||||
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read after a change made through CLI management (below).
|
||||
|
||||
## Managing CLIs from Settings
|
||||
|
||||
App Settings → Agents & CLIs → **CLI management** (`cliManagementEnabled`, default OFF; admin-only in multi-user mode) lists every entry with an installed/not-installed badge and:
|
||||
|
||||
- toggles any entry on or off. A `kind: 'shell'` entry cannot be disabled, and the row shows no switch for it. A disabled CLI disappears from the Run menu, the welcome screen and the phone overview, and new session requests for it are rejected.
|
||||
- installs a missing **stock** CLI by running its shipped install command, after a confirm that names the exact command. Only one install per CLI runs at a time, and the command runs without any `CODEMAN_*` variable in its environment. A custom entry's install command is never executed.
|
||||
- adds, edits and deletes **custom** entries (id, label, badge, binaries, launch argv). The server re-validates the whole assembled entry through `CliEntrySchema`, so the form cannot bypass the load-time rules.
|
||||
|
||||
These are the only writes to `clis.json`. They are serialized, and a file that does not parse or has unsafe permissions is refused rather than overwritten; fix it (or `chmod 600` it) and retry. The HTTP routes are listed in `docs/api-reference.md` under *CLI management*.
|
||||
|
||||
## The shape of an entry
|
||||
|
||||
```ts
|
||||
interface CliEntry {
|
||||
id: CliId; // 'codex'
|
||||
label: string; // 'Codex' — shown in menus
|
||||
shortBadge: string; // tab badge, e.g. 'CX'
|
||||
accent: string; // single hex colour
|
||||
enabled: boolean;
|
||||
stock: boolean; // set by the loader; a custom entry can never claim it
|
||||
order: number;
|
||||
kind: 'agent' | 'shell';
|
||||
discovery: CliDiscovery; // how to find and prove the binary
|
||||
launch: CliLaunch; // the structured argv template
|
||||
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
|
||||
capabilities: CliCapabilities; // what every call site reads instead of the id
|
||||
// .workDetect?: { promptGlyph, workingLine, watchingLine?, watchingLines? } — how
|
||||
// this CLI's pane shows work, and how it shows work it started in the background
|
||||
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
|
||||
}
|
||||
```
|
||||
|
||||
`capabilities` is the important part. It is what `isExternalCliMode()`, `isAltScreenStripMode()`, `hooksAvailableForMode()` and every other former per-mode branch actually read.
|
||||
|
||||
### Regexes that come from config
|
||||
|
||||
Three capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`. All three go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
|
||||
|
||||
`workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
|
||||
|
||||
`watchingLine` reads a different row of the same screen. A CLI draws it while work the agent
|
||||
itself started is still running — Claude prints `⏵⏵ bypass permissions on · 1 monitor · ← for
|
||||
agents` while a monitor, a backgrounded shell or a cloud session is live. Codeman turns that
|
||||
into `Session.watching`, and an idle prompt from such a session opens already acknowledged,
|
||||
so a pane waiting for its own background work never raises an alert a human cannot answer.
|
||||
Group 1 is the label, and a CLI that declares no pattern reports no background work.
|
||||
|
||||
Two CLIs declare such a row today, and they put it in different places. Claude writes its
|
||||
chip on the last row of the screen, so it keeps the default one-row window and anchors on
|
||||
the `·` its footer joins items with. Codex pins
|
||||
`1 background terminal running · /ps to view · /stop to close` ABOVE its composer, which
|
||||
puts the row third from the bottom once the status line and the composer are counted, so its
|
||||
entry declares `watchingLines: 3` and matches that row end to end. Both were measured
|
||||
against live panes rather than read out of a binary, which is the standard for adding a
|
||||
third.
|
||||
|
||||
That label is the one value in the registry that an AGENT can influence, because it comes off
|
||||
the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern
|
||||
for a new CLI. `watchingLabel()` in `session-activity.ts` searches only the last few
|
||||
non-blank rows, which should be the part of the screen the CLI draws rather than the agent,
|
||||
and the pattern should anchor on chrome only that CLI can produce. Keep the window as small
|
||||
as the layout allows, since every row it adds is another row the agent may be able to write.
|
||||
The label is also ANSI-stripped and length-capped at the source, and every interpolation of
|
||||
it into markup goes through `escapeHtml()`, since it ends up on a badge and in an approval
|
||||
card.
|
||||
|
||||
The two shipped entries do not sit equally well behind that rule, and the difference decides
|
||||
what a pattern is allowed to do. Claude's chip is the last row, so its one-row window holds
|
||||
nothing the agent can write — not even the status line above it, whose command a session
|
||||
running with permissions bypassed can write into its own `.claude/settings.json`. Codex's row
|
||||
shares its slot with the last row of the transcript whenever no terminal is running, so a
|
||||
message ending in that exact line is matched. What keeps that harmless is `hooks: 'none'`: no
|
||||
hook event from a codex session reaches the approvals inbox, so a forged label costs a wrong
|
||||
badge and cannot silence an alert. Before giving a CLI both hook signals and a pattern, make
|
||||
sure its row is one the agent cannot write.
|
||||
|
||||
### Three capabilities that must stay independent
|
||||
|
||||
`external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks.
|
||||
|
||||
`test/cli-capability-predicates.test.ts` asserts that no two of the three are equivalent across the catalog, so collapsing them fails the build rather than a user's session.
|
||||
|
||||
## Arg-template safety
|
||||
|
||||
The composed command line is interpolated into `bash -c "…"` inside tmux, which makes command construction a security boundary. Four independent layers keep config out of it:
|
||||
|
||||
1. **Config contains no shell text.** There is no `command: "..."` field anywhere in the schema. An entry declares a sequence of typed tokens; `argv.ts` is the only place that turns them into a string, and it owns every separator itself — one space between tokens, ` || ` between fallback variants. Neither can originate from config, because config has no field that could hold either.
|
||||
2. **Every literal is validated at LOAD time** against a safe-word pattern (no space, quote, backtick, `$`, `;`, `&`, `|`, redirection, parens, braces, newline or backslash). A bad literal **rejects the whole entry** rather than being dropped, because a silently dropped flag would change security-relevant behaviour — losing `--no-approve` is not a cosmetic difference.
|
||||
3. **Values resolve through NAMED patterns.** A value placeholder selects a `TokenPattern` (`model`, `uuid`, `slug`, `path-segment`, `tool-list`, …) from `patterns.ts`; config can never supply its own regex for a value, so a `clis.json` structurally cannot widen its own validation. A value that fails its pattern drops the whole argument, exactly as the hand-written builders did: an invalid `--model` omits `--model`, it never substitutes something else.
|
||||
4. **Escaping is independent of validation.** `renderToken()` re-checks the resolved value before emitting it unquoted, and single-quotes anything else — so even a value that somehow bypassed validation is quoted, never concatenated raw.
|
||||
|
||||
The only config-supplied regexes are `discovery.version.regex` and `discovery.identity.regex`. Both run against **command output** rather than a shell token, both are compiled through `compileVersionRegex()` (length cap, nested-quantifier rejection, never the `g` flag), and the output they see is truncated first.
|
||||
|
||||
## Named profiles: the escape hatch
|
||||
|
||||
Some differences genuinely need to run code rather than be described. Those are **named profiles**: a capability field holds a profile NAME, and the implementation lives in one place keyed by that name — never by CLI id.
|
||||
|
||||
- `discovery.launcherProfile` — for a CLI whose binary is not the agent. `dsh` boots `$DSH_HOME/profiles/<name>`, so "installed" and "runnable" have different answers; the profile answers both, plus why a specifically-named target will not work. Implemented in `utils/cli-launcher.ts`.
|
||||
- `env.setenvProfile` — per-CLI environment setup that is more than a list of keys, such as DeepSeek's status bridge.
|
||||
- `capabilities.transcript` — which on-disk history reader understands this CLI (`claude-jsonl`, `codex-rollout`, `deepseek-zstd`, `omp-jsonl`, `none`).
|
||||
- `capabilities.echo.predictProfile` — the predictive-echo model a composer needs.
|
||||
|
||||
The names live in `profiles.ts`, which is kept free of imports so `schema.ts` can validate a name at load time. A profile this build does not implement is a load-time error naming the field, rather than a CLI that silently looks permanently uninstalled.
|
||||
|
||||
## DeepSeek: the four assumptions it breaks
|
||||
|
||||
DeepSeek is worth reading before assuming an entry looks like its siblings — the schema carries four extensions because of it.
|
||||
|
||||
| What it breaks | How the registry expresses it |
|
||||
| ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `dsh` is a profile LAUNCHER, not the agent, so "installed" is not "runnable". | `discovery.launcherProfile` + `discovery.launcherTargetParam`. |
|
||||
| Its permission switch is the **`DSH_PERMISSION_MODE` env var**, not a flag — the harness has none. | `env.configSetenv` (so the ordinary `privilegedParams` clamp still reaches it) **and** `capabilities.privilegedEnvKeys`. |
|
||||
| It is the only non-claude mode with real hook signals, and for it that is a per-SESSION question. | `capabilities.hooks: 'supervised'` — a third state, not a boolean. |
|
||||
| Its transcript is zstd session files, one frame per write. | `capabilities.transcript: 'deepseek-zstd'`. |
|
||||
|
||||
## Identity probes
|
||||
|
||||
`discovery.identity` asks the binary whether it is the program we meant, and it runs **before** the version probe, because a version probe cannot tell an impostor from the real thing. Debian ships an unrelated `dsh` (dancer's shell) that answers `--version` perfectly happily, and npm carries squatters for both `pi` and `grok`.
|
||||
|
||||
`discovery.version.requireVersionMatch` is the weaker companion: a binary whose version output has the wrong shape counts as ABSENT rather than present-with-unknown-version. That is what a short, generic binary name needs, and it is what keeps `codeman doctor` and the run mode from telling the user opposite things about the same binary — both read the same regex off the same entry.
|
||||
|
||||
## The no-id-branching rule
|
||||
|
||||
`test/cli-registry-no-id-branching.test.ts` fails the build if a CLI id comparison appears outside the stock catalog. It builds its id list from the live catalog, blanks comment lines before scanning (comments legitimately quote the pattern to explain why a branch was removed, and blanking rather than dropping is what keeps reported line numbers pointing at the real file), and keeps an allowlist in which **every entry carries its reason**.
|
||||
|
||||
It matches four shapes, not one: `mode === '<id>'`, `mode !== '<id>'`, `case '<id>':`, and `['<id>', …].includes(mode)`. The first version matched `===` only, and that gap was not academic — the refactor it guards converted the `===` sites and left the negated ones, so 36 `!==` branches survived it, including a seven-mode chain auto-enabling Ralph under a comment asking the next person to keep it in step with a predicate by hand while the sibling code path already read the capability. A guard that sees half the shapes reports a count measured over the half it happens to catch.
|
||||
|
||||
The allowlist is not a formality. If a branch is about what a CLI can DO it belongs in `CliCapabilities`; the entries that remain are things that are not CLI-behaviour branches at all — chiefly the legacy per-mode `<Mode>Config` objects on `POST /api/sessions`, which are a fact about the public HTTP API rather than about any CLI, plus a few documented cases where `mode === 'claude'` is genuinely the right question (Read My Mind reads Claude's _own_ transcript, so a capability there would be actively wrong).
|
||||
|
||||
`test/frontend-cli-no-id-branching.test.ts` is the same guard for the two frontend files the CLI registry's Run-menu consolidation touches, `session-ui.js` and `mobile-overview.js` — deliberately not the rest of `src/web/public/`, whose per-CLI rules stay out of scope for now (see "Fields declared for later" below). Its allowlist keys on `<file>::<expression>` with no line number, since a single unrelated edit to a contended file would otherwise shift every subsequent line and make every entry go stale at once, and each entry additionally carries the exact number of approved call sites — a bare key would let a brand-new branch reusing an already-approved expression land unreviewed. Its comparison shape differs from the backend guard's in one respect: the left-hand side may be any identifier, not only one named `mode`, `id` or `agentType`, because the review of #458 found `const m = this._runMode; if (m === 'codex')` slipping past the named form while the scanned file already filters with `(m) => m !== 'shell'`.
|
||||
|
||||
## Two namespaces called `param`
|
||||
|
||||
`launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a **launch param**. The **legacy wire field** a param arrives as is a separate namespace, and `launch.legacyConfigAliases` is the only bridge between the two.
|
||||
|
||||
This matters because it is invisible when it is wrong. `capabilities.privilegedParams[].param` is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a name from the wrong namespace clamps **nothing**: no load error, no failing test, the clamp simply stops running. Codex is the entry where the two names differ (`bypassApprovals` as the param, `dangerouslyBypassApprovals` on the wire), so it is the one that catches a regression. `schema.ts` rejects any entry naming a param it never declared, on both `configSetenv.fromParam` and `privilegedParams.param`.
|
||||
|
||||
## Fields declared for later
|
||||
|
||||
`accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. (`shortBadge` was on this list until the CLI management list in Settings started showing it.) They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
|
||||
|
||||
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, so re-measure before wiring one up. `accent` is the one exception: it was measured against styles.css on 2026-09-21 (method in the comment above `CLAUDE` in `stock.ts`), though nothing keeps it in step with the CSS either. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
|
||||
|
||||
`overlays.credStore` is in the same category, for a sharper reason: the Docker credential-seeding path still reads its own `CRED_STORES` table, because this shape allows ONE store per CLI and the live table needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek declares none here even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is simultaneously the worst thing here to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest.
|
||||
|
||||
Everything else in the interface is live, including `overlays.remote` / `overlays.docker`, which back `defaultRemoteCommandForMode()` and `defaultDockerCommandForMode()` directly. Those two used to be hardcoded `Record<…CommandMode, string>` tables duplicating the registry with nothing keeping the two in step; `test/location-overlay-commands.test.ts` pins every resulting command as a literal string.
|
||||
|
||||
## Consumers outside the server
|
||||
|
||||
Two things need the catalogue but cannot import TypeScript, so `npm run generate:cli-catalog`
|
||||
(`scripts/generate-cli-catalog.mts`) emits two artifacts from `stock.ts`. Both are committed,
|
||||
and `test/cli-catalog-sync.test.ts` fails if either drifts from a fresh generation.
|
||||
|
||||
| Artifact | Consumer | Why it exists |
|
||||
| ------------------------------------ | ---------------------------------- | ---------------------------------------------------------------------------------- |
|
||||
| `config/clis.stock.json` | `scripts/lib/cli-catalog.mjs` (Docker build args), tests | A `.mjs` cannot import the registry. |
|
||||
| a marked block inside `install.sh` | the installer itself | It runs via `curl \| bash` before any checkout exists, so it can read neither. |
|
||||
|
||||
Only `id`, `label`, `shortBadge`, `enabled`, `order`, `kind` and `discovery` are exported.
|
||||
`launch`, `env`, `capabilities` and `overlays` are spawn-time concerns the server alone
|
||||
interprets, and a test asserts they never leak into the artifact — a second reading of the
|
||||
launch model in a consumer that cannot be tested against a real spawn is exactly what this
|
||||
registry exists to prevent.
|
||||
|
||||
The install.sh copy is **embedded, not fetched**, and is the FULL catalogue. An earlier design
|
||||
fetched it and fell back to a hardcoded two-CLI list, which degraded silently on an empty
|
||||
response; there is no degraded mode to fall into now, and no network fetch either — a `curl |
|
||||
bash` from master already carries a catalogue exactly as fresh as the script itself, so there is
|
||||
nothing a refresh would buy that isn't already true. An earlier draft added an opt-in refresh
|
||||
with a `TRUSTED`/`DISPLAY` array split to keep it from ever writing the executed command; it was
|
||||
dropped before merge rather than shipped half-verified — the split's only actual write was the
|
||||
label, `DISPLAY` never diverged from `TRUSTED` in practice, and the added surface (a second
|
||||
array, a fetch path, three failure shapes to warn on) bought nothing the embedded copy didn't
|
||||
already have.
|
||||
|
||||
### The install-command trust boundary
|
||||
|
||||
Three rules, and the middle one is why the embed matters:
|
||||
|
||||
1. **The server never executes an entry's `install.command`.** Unchanged, and still enforced by nothing executing it: the field is display text (`CliDiscovery.install.command`).
|
||||
2. **`install.sh` executes only commands embedded in itself.** Those arrive in the same file, over the same TLS fetch, in the same commit as the `curl \| bash` line that fetched the script — identical trust to the hardcoded vendor one-liners it replaces.
|
||||
3. **Nothing fetched at install time is ever executed.** There is no second code path that fetches anything after the script itself has been fetched.
|
||||
|
||||
That is mechanical rather than a promise. `CLI_INSTALL_CMD_TRUSTED` is written only from the
|
||||
generated block and is the only array the installer ever runs or displays — there is no second
|
||||
array a refresh could rewrite, because there is no refresh. `test/cli-catalog-sync.test.ts`
|
||||
asserts that the embedded commands are exactly the registry's, and
|
||||
`test/install-sh-invariants.test.ts` that nothing in `install.sh` `eval`s.
|
||||
|
||||
### bash 3.2
|
||||
|
||||
macOS ships bash 3.2 and the documented install is `curl -fsSL <url> | bash` under
|
||||
`set -euo pipefail`, so a bash-4 construct is not a warning there — it kills the install. The
|
||||
generated block therefore uses parallel indexed arrays with **offset/length windows** into one
|
||||
flat array instead of delimiters (a `$HOME` containing a space needs no `IFS` handling, and an
|
||||
entry with nothing to contribute gets length 0 and is never iterated). CI runs `bash -n` and
|
||||
executes the script inside a real `bash:3.2` container, because the empty-window case is a
|
||||
runtime `set -u` abort that `bash -n` cannot see.
|
||||
|
||||
## Resolve at call time, never at import
|
||||
|
||||
Anything reading the registry must resolve it when it is asked, not when its module is first imported. `sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()` and each resolver's `searchDirs` thunk all re-read the catalog per call.
|
||||
|
||||
A module-level const freezes at first import, and the failure is asymmetric: a CLI enabled while the server is running moved the run menu but not the frozen surface, so validation rejected a mode the menu offered, or `codeman doctor` reported a catalog nobody had any more.
|
||||
|
||||
## Adding a CLI
|
||||
|
||||
1. Add a `CliEntry` to `stock.ts`.
|
||||
2. Run `npm run generate:cli-catalog` and commit **both** artifacts (`config/clis.stock.json` and `install.sh`). The installer's detection, its install menu, its reminder text and the Docker agent image all follow from that one step — this is what makes upstream `b6d0f1fa` ("wire OMP into install.sh's CLI detection, it had none") impossible rather than merely fixed.
|
||||
3. Add a golden spawn-command pin to `test/cli-registry-spawn-golden.test.ts`, a row to `test/cli-capability-predicates.test.ts`, its remote/docker commands to `test/location-overlay-commands.test.ts`, and its search paths to `test/install-sh-detection-parity.test.ts`.
|
||||
4. Only if it cannot install with a plain `npm install -g <pkg>`: give it a layer in `docker/agent.Dockerfile` and set `discovery.install.agentImageLayer: { kind: 'dedicated', reason }` on its entry in `stock.ts`. `test/docker-agent-image-coverage.test.ts` requires both, so an exclusion cannot quietly become an omission. An entry with no `npmPackage` needs only the Dockerfile layer, since it never enters the shared npm layer in the first place.
|
||||
5. That is usually all. If you find yourself wanting to add an `if` somewhere, the guard test will tell you — and the answer is a capability field, or a named profile if it genuinely needs to run code.
|
||||
|
||||
## See also
|
||||
|
||||
- [Agent CLIs](wiki/Agent-CLIs.md) — the user-facing per-CLI guide.
|
||||
- `docs/architecture-invariants.md` — the mechanics and the history behind the rules above.
|
||||
- `docs/deepseek-integration.md` — why DeepSeek is shaped the way it is.
|
||||
@@ -1,588 +0,0 @@
|
||||
# Claude Code Build Brief: Add Scheduling to Codeman
|
||||
|
||||
## 0. Purpose of This Brief
|
||||
|
||||
You are Claude Code working inside the Codeman repository.
|
||||
|
||||
Your task is to add a **small, reliable scheduling layer** to Codeman while preserving Codeman's existing architecture and session-management behavior.
|
||||
|
||||
This is not a greenfield rewrite. This is not a full product rebuild. This is a focused extension.
|
||||
|
||||
The target user wants Codeman-like tmux/web/session management, but with first-class scheduled jobs for Claude, Codex, OpenCode, Terminal, or any other configurable coding-agent harness.
|
||||
|
||||
---
|
||||
|
||||
## 1. Non-Negotiable Goal
|
||||
|
||||
Add scheduling to Codeman so a user can define a scheduled coding-agent job that:
|
||||
|
||||
1. Has a name.
|
||||
2. Uses an existing Codeman-supported agent/session type where possible.
|
||||
3. Has a working directory.
|
||||
4. Has a prompt or prompt file.
|
||||
5. Has a schedule.
|
||||
6. Can be enabled or disabled.
|
||||
7. Can be manually run now.
|
||||
8. When due, creates a Codeman/tmux session.
|
||||
9. Sends the configured prompt into that session.
|
||||
10. Records last run, next run, status, and run history.
|
||||
|
||||
The first working version should prioritize **scheduling correctness and reuse of Codeman's existing tmux/session system** over UI polish.
|
||||
|
||||
---
|
||||
|
||||
## 2. Core Architectural Rule
|
||||
|
||||
Do **not** rebuild Codeman's session layer.
|
||||
|
||||
Reuse existing Codeman functionality for:
|
||||
|
||||
- Creating sessions.
|
||||
- Naming sessions.
|
||||
- Launching Claude/Codex/OpenCode/Terminal sessions.
|
||||
- Sending input into sessions.
|
||||
- Displaying sessions in the web UI.
|
||||
- Killing sessions.
|
||||
- Tracking session status if already supported.
|
||||
|
||||
If an internal API/service/function already exists, reuse it.
|
||||
|
||||
If no reusable function exists, create a thin wrapper around the existing implementation rather than duplicating logic.
|
||||
|
||||
---
|
||||
|
||||
## 3. Product Boundary
|
||||
|
||||
This build is **Codeman + Scheduler**.
|
||||
|
||||
It is not yet:
|
||||
|
||||
- A full quota engine.
|
||||
- A full lock manager.
|
||||
- A replacement for Codeman's terminal UI.
|
||||
- A new FastAPI application.
|
||||
- A multi-tenant SaaS platform.
|
||||
- A complex cron-management product.
|
||||
- A full agent autonomy framework.
|
||||
|
||||
Keep the build small and shippable.
|
||||
|
||||
---
|
||||
|
||||
## 4. Required Working Scope for v0.1
|
||||
|
||||
Implement the following minimum features.
|
||||
|
||||
### 4.1 Scheduled Jobs List
|
||||
|
||||
Create a UI page showing all scheduled jobs.
|
||||
|
||||
Each row/card should show:
|
||||
|
||||
- Job name.
|
||||
- Agent/session type.
|
||||
- Working directory.
|
||||
- Schedule type.
|
||||
- Enabled/disabled state.
|
||||
- Last run time.
|
||||
- Next run time.
|
||||
- Last run status.
|
||||
- Actions:
|
||||
- Run Now.
|
||||
- Enable/Disable.
|
||||
- Edit.
|
||||
- Delete.
|
||||
|
||||
### 4.2 Create/Edit Scheduled Job
|
||||
|
||||
Create a form for scheduled jobs with these fields:
|
||||
|
||||
- `name`
|
||||
- `agent_type`
|
||||
- Reuse Codeman's existing session/agent types where possible.
|
||||
- Include at least Terminal/custom command if supported.
|
||||
- `working_directory`
|
||||
- `launch_command` if needed by Codeman's model.
|
||||
- `prompt_mode`
|
||||
- `inline_text`
|
||||
- `prompt_file_path`
|
||||
- `prompt_text`
|
||||
- `prompt_file_path`
|
||||
- `input_mode`
|
||||
- `paste`
|
||||
- `typed`
|
||||
- `schedule_type`
|
||||
- `once`
|
||||
- `interval_minutes`
|
||||
- `daily_time`
|
||||
- `weekly_time`
|
||||
- `run_at` for one-time jobs.
|
||||
- `interval_minutes` for interval jobs.
|
||||
- `daily_time` for daily jobs.
|
||||
- `weekly_days` and `weekly_time` for weekly jobs.
|
||||
- `enabled`
|
||||
- `notes` optional.
|
||||
|
||||
Do not build a complex visual cron editor in v0.1.
|
||||
|
||||
### 4.3 Run Now
|
||||
|
||||
Every scheduled job must support a `Run Now` action.
|
||||
|
||||
Run Now should:
|
||||
|
||||
1. Create a new session through Codeman's existing session creation logic.
|
||||
2. Send the configured prompt into the session using Codeman's existing input mechanism.
|
||||
3. Create a run-history record.
|
||||
4. Update last-run fields.
|
||||
5. Redirect or link the user to the created Codeman session.
|
||||
|
||||
### 4.4 Background Scheduler Loop
|
||||
|
||||
Add a small background scheduler loop that runs inside the Codeman backend process.
|
||||
|
||||
The loop should:
|
||||
|
||||
1. Wake every 15-60 seconds.
|
||||
2. Load enabled schedules.
|
||||
3. Find schedules where `next_run_at <= now`.
|
||||
4. Create a scheduled run.
|
||||
5. Launch the session using existing Codeman session logic.
|
||||
6. Send the prompt.
|
||||
7. Record run history.
|
||||
8. Compute the next run time.
|
||||
9. Avoid duplicate launches if the loop overlaps or restarts.
|
||||
|
||||
Keep this simple and robust.
|
||||
|
||||
### 4.5 Run History
|
||||
|
||||
Every scheduled execution should create a run-history record.
|
||||
|
||||
Track:
|
||||
|
||||
- `id`
|
||||
- `scheduled_job_id`
|
||||
- `session_id` or Codeman session reference.
|
||||
- `session_name` if applicable.
|
||||
- `started_at`
|
||||
- `finished_at` optional.
|
||||
- `status`
|
||||
- `created`
|
||||
- `session_started`
|
||||
- `prompt_sent`
|
||||
- `failed`
|
||||
- `error_message` optional.
|
||||
- `trigger_type`
|
||||
- `scheduled`
|
||||
- `manual_run_now`
|
||||
- `created_session_url` or route reference if easy.
|
||||
|
||||
---
|
||||
|
||||
## 5. Scheduling Rules
|
||||
|
||||
### 5.1 Once
|
||||
|
||||
Run at a specific date/time.
|
||||
|
||||
After successful launch:
|
||||
|
||||
- Set `enabled = false`, or mark as completed.
|
||||
|
||||
### 5.2 Interval
|
||||
|
||||
Run every N minutes.
|
||||
|
||||
Example:
|
||||
|
||||
- Every 60 minutes.
|
||||
- Every 240 minutes.
|
||||
|
||||
After launch:
|
||||
|
||||
- `next_run_at = now + interval_minutes`.
|
||||
|
||||
### 5.3 Daily
|
||||
|
||||
Run every day at HH:MM.
|
||||
|
||||
After launch:
|
||||
|
||||
- Compute the next occurrence of HH:MM after now.
|
||||
|
||||
### 5.4 Weekly
|
||||
|
||||
Run on selected weekdays at HH:MM.
|
||||
|
||||
After launch:
|
||||
|
||||
- Compute the next selected weekday/time after now.
|
||||
|
||||
### 5.5 Timezone
|
||||
|
||||
Use the server's local timezone for v0.1 unless Codeman already has timezone handling.
|
||||
|
||||
Add a visible note in the UI:
|
||||
|
||||
> Times use the server's local timezone.
|
||||
|
||||
Do not overbuild timezone support in v0.1.
|
||||
|
||||
---
|
||||
|
||||
## 6. Data Storage Decision
|
||||
|
||||
First inspect Codeman's existing persistence model.
|
||||
|
||||
If Codeman already has a database or persistence layer:
|
||||
|
||||
- Reuse it.
|
||||
- Add scheduled job and scheduled run models/tables/records using the existing pattern.
|
||||
|
||||
If Codeman uses files or JSON state:
|
||||
|
||||
- Use the same style for v0.1.
|
||||
- Prefer simple persistence over introducing a heavy new dependency.
|
||||
|
||||
If there is no appropriate persistence layer:
|
||||
|
||||
- Add SQLite only if it fits the codebase cleanly.
|
||||
- Otherwise use a JSON file store for the first version.
|
||||
|
||||
Do not introduce Postgres, Redis, Celery, or a separate scheduler service.
|
||||
|
||||
---
|
||||
|
||||
## 7. Concurrency and Duplicate-Run Guard
|
||||
|
||||
Implement a basic duplicate-run guard.
|
||||
|
||||
A schedule should not launch twice for the same due time.
|
||||
|
||||
Minimum acceptable approach:
|
||||
|
||||
- Before launching, create/update a run record with a `created` or `launching` state.
|
||||
- Use a schedule-level `last_triggered_at` or `last_due_key` to avoid double launching.
|
||||
- If launch fails, record failure clearly.
|
||||
|
||||
Do not build distributed locks. Codeman is expected to be local/single-instance for v0.1.
|
||||
|
||||
---
|
||||
|
||||
## 8. Multi-Session Warning
|
||||
|
||||
When the user clicks `Run Now`, show a warning if there are already active sessions for the same agent type.
|
||||
|
||||
Minimum behavior:
|
||||
|
||||
- If active sessions exist, show a confirmation warning.
|
||||
- User can continue anyway.
|
||||
|
||||
For scheduled automatic runs:
|
||||
|
||||
- Add a setting on the scheduled job:
|
||||
- `warn_only`
|
||||
- `skip_if_same_agent_running`
|
||||
|
||||
Default:
|
||||
|
||||
- `warn_only` for manual runs.
|
||||
- `skip_if_same_agent_running = false` for automatic runs unless easy to implement.
|
||||
|
||||
Do not build a complete quota engine in v0.1.
|
||||
|
||||
---
|
||||
|
||||
## 9. Prompt Sending Rules
|
||||
|
||||
The scheduler must support sending the configured prompt into the created session.
|
||||
|
||||
Prompt source:
|
||||
|
||||
1. Inline prompt text.
|
||||
2. Prompt file path.
|
||||
|
||||
Input mode:
|
||||
|
||||
1. Paste mode.
|
||||
2. Typed mode.
|
||||
|
||||
If only one input mode is easy with Codeman's current internals, implement that first and structure the code so the other can be added later.
|
||||
|
||||
Important:
|
||||
|
||||
- Do not send prompts to a session if session creation failed.
|
||||
- Record prompt-send success/failure in run history.
|
||||
- Save enough metadata to understand what prompt was used.
|
||||
|
||||
---
|
||||
|
||||
## 10. UI Bifurcation
|
||||
|
||||
Keep UI changes cleanly separated.
|
||||
|
||||
Add scheduler UI under a clear navigation item:
|
||||
|
||||
- `Scheduled Jobs`
|
||||
|
||||
Do not clutter the existing session dashboard.
|
||||
|
||||
The existing session dashboard may show sessions created by scheduled jobs, but the scheduling controls should live in their own section.
|
||||
|
||||
Recommended pages/routes:
|
||||
|
||||
- `/schedules`
|
||||
- `/schedules/new`
|
||||
- `/schedules/:id`
|
||||
- `/schedules/:id/edit`
|
||||
- `/schedules/:id/run-now`
|
||||
- `/schedules/:id/enable`
|
||||
- `/schedules/:id/disable`
|
||||
- `/schedules/:id/delete`
|
||||
|
||||
Use Codeman's existing frontend conventions and routing style.
|
||||
|
||||
---
|
||||
|
||||
## 11. Backend Bifurcation
|
||||
|
||||
Keep scheduler code separate from existing session code.
|
||||
|
||||
Recommended logical modules, adapted to Codeman's actual structure:
|
||||
|
||||
- `scheduler/model` or equivalent.
|
||||
- `scheduler/store` or equivalent.
|
||||
- `scheduler/service` for schedule calculations and launch logic.
|
||||
- `scheduler/loop` for the background due-job checker.
|
||||
- `scheduler/routes` for API/UI endpoints.
|
||||
- `scheduler/time` for next-run calculations.
|
||||
|
||||
Do not mix scheduling logic directly into terminal rendering, xterm handling, or low-level tmux code.
|
||||
|
||||
The scheduler service should call session services; it should not own tmux directly unless Codeman has no session abstraction.
|
||||
|
||||
---
|
||||
|
||||
## 12. Required Discovery Phase Before Coding
|
||||
|
||||
Before implementing, inspect the Codeman repo and produce a short architecture note in the terminal or in a file called:
|
||||
|
||||
`docs/cron-discovery.md`
|
||||
|
||||
This note must identify:
|
||||
|
||||
1. Where session creation happens.
|
||||
2. Where agent/session types are defined.
|
||||
3. Where input is sent into a session.
|
||||
4. Where active sessions are listed.
|
||||
5. Where session kill/delete is handled.
|
||||
6. How session state is stored.
|
||||
7. Whether there is existing persistence.
|
||||
8. Where backend routes live.
|
||||
9. Where frontend pages/components live.
|
||||
10. The smallest integration points for scheduling.
|
||||
|
||||
Do not start coding until this discovery is complete.
|
||||
|
||||
---
|
||||
|
||||
## 13. Implementation Phases
|
||||
|
||||
### Phase 1: Discovery
|
||||
|
||||
Deliverable:
|
||||
|
||||
- `docs/cron-discovery.md`
|
||||
|
||||
Must answer the 10 discovery questions above.
|
||||
|
||||
### Phase 2: Data Model / Persistence
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Scheduled job persistence.
|
||||
- Scheduled run history persistence.
|
||||
- Basic create/read/update/delete operations.
|
||||
|
||||
### Phase 3: Scheduler Calculation Logic
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Functions to compute `next_run_at` for:
|
||||
- once
|
||||
- interval
|
||||
- daily
|
||||
- weekly
|
||||
|
||||
Add tests if the repo has an existing test setup.
|
||||
|
||||
### Phase 4: Manual Run Now
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Create scheduled job.
|
||||
- Click Run Now.
|
||||
- Codeman session is created.
|
||||
- Prompt is sent.
|
||||
- Run history is recorded.
|
||||
- UI links to the session.
|
||||
|
||||
This is the most important milestone.
|
||||
|
||||
### Phase 5: Background Scheduler Loop
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Enabled schedules launch automatically when due.
|
||||
- Run history is recorded.
|
||||
- `last_run_at` and `next_run_at` update.
|
||||
- Duplicate launch guard exists.
|
||||
|
||||
### Phase 6: UI Polish Only After Functionality
|
||||
|
||||
Deliverable:
|
||||
|
||||
- Scheduled jobs list is readable.
|
||||
- Create/edit form is usable.
|
||||
- Status labels are clear.
|
||||
- Errors are visible.
|
||||
|
||||
Do not polish before Phase 4 works.
|
||||
|
||||
---
|
||||
|
||||
## 14. Acceptance Criteria
|
||||
|
||||
The build is acceptable when all these pass.
|
||||
|
||||
### Manual Run
|
||||
|
||||
1. Create a schedule/job with inline prompt.
|
||||
2. Click Run Now.
|
||||
3. A new Codeman/tmux session starts.
|
||||
4. Prompt is sent into that session.
|
||||
5. The created session is visible in Codeman's normal session UI.
|
||||
6. Run history shows success or failure.
|
||||
|
||||
### One-Time Schedule
|
||||
|
||||
1. Create a one-time schedule 2 minutes in the future.
|
||||
2. Wait for it to become due.
|
||||
3. Scheduler launches a session.
|
||||
4. Prompt is sent.
|
||||
5. Schedule does not repeatedly launch forever.
|
||||
|
||||
### Interval Schedule
|
||||
|
||||
1. Create interval schedule every 2 minutes.
|
||||
2. It launches once when due.
|
||||
3. It computes the next due time.
|
||||
4. It does not launch duplicates for the same due time.
|
||||
|
||||
### Daily Schedule
|
||||
|
||||
1. Create daily schedule at a time a few minutes ahead.
|
||||
2. It launches when due.
|
||||
3. Next run becomes tomorrow at the same time.
|
||||
|
||||
### Disable Schedule
|
||||
|
||||
1. Disable a schedule.
|
||||
2. It does not launch even when due.
|
||||
|
||||
### Error Handling
|
||||
|
||||
1. Invalid working directory produces visible error.
|
||||
2. Invalid prompt file produces visible error.
|
||||
3. Failed session launch creates failed run-history entry.
|
||||
|
||||
---
|
||||
|
||||
## 15. Explicitly Out of Scope for v0.1
|
||||
|
||||
Do not implement these unless all required scope is already working:
|
||||
|
||||
- Full quota engine.
|
||||
- Advanced lock manager.
|
||||
- Post-run git inspection reports.
|
||||
- Complex recurring calendar UI.
|
||||
- User accounts / RBAC.
|
||||
- External distributed workers.
|
||||
- Redis.
|
||||
- Postgres.
|
||||
- Celery.
|
||||
- Kubernetes.
|
||||
- A separate Python service.
|
||||
- Full visual cron editor.
|
||||
- AI-generated follow-up prompts.
|
||||
- Automatic continuation after idle.
|
||||
- Any attempt to bypass agent quotas or platform limits.
|
||||
|
||||
---
|
||||
|
||||
## 16. Quality Rules
|
||||
|
||||
Follow these rules while coding:
|
||||
|
||||
1. Reuse existing Codeman services and conventions.
|
||||
2. Keep scheduler code isolated.
|
||||
3. Prefer boring, readable code over clever abstractions.
|
||||
4. Add error messages that a human can understand.
|
||||
5. Do not break existing Codeman sessions.
|
||||
6. Do not rename existing core concepts unnecessarily.
|
||||
7. Do not introduce large dependencies without strong reason.
|
||||
8. Keep v0.1 local-first and single-instance.
|
||||
9. Commit in logical chunks if git is available.
|
||||
10. After coding, provide a final implementation summary.
|
||||
|
||||
---
|
||||
|
||||
## 17. Final Response Required from Claude Code
|
||||
|
||||
At the end, report:
|
||||
|
||||
1. Files changed.
|
||||
2. New routes/pages added.
|
||||
3. New data structures added.
|
||||
4. How the scheduler loop works.
|
||||
5. How to run the app.
|
||||
6. How to test manual Run Now.
|
||||
7. How to test scheduled execution.
|
||||
8. Known limitations.
|
||||
9. Suggested v0.2 improvements.
|
||||
|
||||
---
|
||||
|
||||
## 18. v0.2 Ideas, Not for Current Build
|
||||
|
||||
Keep these in mind but do not build unless v0.1 is complete:
|
||||
|
||||
- Quota-aware scheduling.
|
||||
- Manual takeover locks.
|
||||
- Post-idle inspection.
|
||||
- Git diff reports.
|
||||
- Schedule groups.
|
||||
- Prompt templates.
|
||||
- Agent-specific concurrency rules.
|
||||
- Better timezone support.
|
||||
- Audit events.
|
||||
- More advanced cron expressions.
|
||||
|
||||
---
|
||||
|
||||
## 19. Final Reminder
|
||||
|
||||
The goal is to add **scheduling** to Codeman quickly and cleanly.
|
||||
|
||||
Do not drift into building a new platform.
|
||||
|
||||
The highest-priority path is:
|
||||
|
||||
1. Discover existing Codeman integration points.
|
||||
2. Add scheduled job persistence.
|
||||
3. Add Run Now.
|
||||
4. Add background due-job loop.
|
||||
5. Add minimal UI.
|
||||
6. Verify that scheduled jobs create real Codeman/tmux sessions and send prompts.
|
||||
|
||||
@@ -1,142 +0,0 @@
|
||||
# CRON_DISCOVERY.md
|
||||
|
||||
Phase 1 deliverable for the "Add Scheduling to Codeman" build brief.
|
||||
This documents the existing Codeman architecture and the smallest integration
|
||||
points for a cron. **No session/tmux logic will be rebuilt** —
|
||||
the new code is purely a trigger + persistence + history layer on top of the
|
||||
existing primitives.
|
||||
|
||||
Stack: `aicodeman` v1.2.1 — Fastify 5 backend, `node-pty` + tmux sessions,
|
||||
vanilla-JS SPA frontend served as static assets, JSON file state store, zod
|
||||
validation, ports-based dependency injection.
|
||||
|
||||
---
|
||||
|
||||
## 0. Critical finding: an existing `ScheduledRun` is NOT a cron
|
||||
|
||||
Codeman already has a `ScheduledRun` concept (`/api/scheduled`,
|
||||
`src/web/ports/infra-port.ts:14-26`, `src/web/server.ts:1480-1605`). It is a
|
||||
**run-now, duration-bounded autonomous loop**: given `{prompt, workingDir,
|
||||
durationMinutes}` it immediately spawns/kills throwaway sessions in a loop until
|
||||
the duration elapses. It has **no** time-based triggering, recurrence
|
||||
(once/interval/daily/weekly), enable/disable, next-run calculation, run history,
|
||||
or persistence across restarts.
|
||||
|
||||
Therefore the brief's core (the calendar/cron trigger layer) does **not** exist
|
||||
and must be built. The execution primitives it sits on top of **do** exist and
|
||||
will be reused. To honor brief §16 ("do not rename existing core concepts"), the
|
||||
new feature is named **`CronJob`** (with **`CronJobRun`** history
|
||||
records), kept distinct from the existing `ScheduledRun`.
|
||||
|
||||
---
|
||||
|
||||
## 1. Where session creation happens
|
||||
|
||||
- Canonical create flow: `POST /api/sessions`,
|
||||
`src/web/routes/session-routes.ts:262-438`.
|
||||
- `new Session({ workingDir, mode, ... })` (`src/session.ts:421-570`)
|
||||
- `ctx.addSession(session)` → `ctx.setupSessionListeners(session)` →
|
||||
`ctx.persistSessionState(session)` (all via `SessionPort`).
|
||||
- `SessionPort` interface: `src/web/ports/session-port.ts:8-16`.
|
||||
- **Integration point:** the cron service will mirror this exact sequence
|
||||
(create → addSession → setupSessionListeners → start) via `SessionPort`,
|
||||
not reimplement it.
|
||||
|
||||
## 2. Where agent/session types are defined
|
||||
|
||||
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity' | 'pi'`
|
||||
(`src/types/session.ts:43-44`). `shell` covers the brief's "Terminal/custom".
|
||||
- CLI availability resolvers in `src/utils/{claude,codex,gemini,antigravity,opencode,pi}-cli-resolver.ts`.
|
||||
- **Integration point:** the job's `agentType` reuses `SessionMode` verbatim.
|
||||
|
||||
## 3. Where input is sent into a session
|
||||
|
||||
- Raw / paste: `session.write(data)` (`src/session.ts:2243-2247`) — direct PTY write.
|
||||
- Typed (recommended): `session.writeViaMux(data)` (`src/session.ts:2301-2311`)
|
||||
— tmux `send-keys`, falls back to PTY. Submit requires trailing `\r`.
|
||||
- **Integration point:** prompt delivery uses `writeViaMux` (typed) by default,
|
||||
`write` (paste) as the alternate `input_mode`.
|
||||
|
||||
## 4. Where active sessions are listed
|
||||
|
||||
- `ctx.sessions: ReadonlyMap<string, Session>` (`SessionPort`).
|
||||
- Filters: `Array.from(ctx.sessions.values()).filter(s => s.mode === X)` and
|
||||
`.isBusy()` / `.isIdle()` (`src/session-manager.ts:220-247`).
|
||||
- **Integration point:** the §8 multi-session warning queries this map.
|
||||
|
||||
## 5. Where session kill/delete is handled
|
||||
|
||||
- `ctx.cleanupSession(sessionId, killMux?, reason?)`
|
||||
(`SessionPort`; impl `src/web/server.ts:997-1152`). Underlying
|
||||
`session.stop(killMux)` at `src/session.ts:2498-2585`.
|
||||
- The cron does **not** kill sessions it launches (the brief wants them
|
||||
visible in the normal session UI); cleanup stays user-driven.
|
||||
_Superseded post-review:_ recurring jobs now default to
|
||||
`autoClosePreviousSession: true` — the previous run's still-open session is
|
||||
closed via `cleanupSession` when the next run fires (see
|
||||
`docs/cron-guide.md` §8); opt out per job for fully user-driven cleanup.
|
||||
|
||||
## 6. How session state is stored / 7. Existing persistence
|
||||
|
||||
- JSON file store: `~/.codeman/state.json` (+ `state-inner.json` for Ralph).
|
||||
`StateStore` class `src/state-store.ts:71`; `AppState` interface
|
||||
`src/types/app-state.ts:99-114`.
|
||||
- Pattern: declare a field on `AppState`, add typed get/set methods on
|
||||
`StateStore` that mutate in-memory state and call the debounced `save()`
|
||||
(500ms debounce, atomic temp-file+rename, `.bak` backup, circuit breaker).
|
||||
- **Integration point:** add `cronJobs?: Record<string, CronJob>` and
|
||||
`cronJobRuns?: Record<string, CronJobRun>` to `AppState`, with
|
||||
matching `StateStore` accessors. No new DB (brief §6 forbids Postgres/Redis).
|
||||
|
||||
## 8. Where backend routes live
|
||||
|
||||
- Route modules: `src/web/routes/*.ts`; barrel `src/web/routes/index.ts`;
|
||||
registered in `WebServer.setupRoutes()` `src/web/server.ts:858-876` with a
|
||||
single `ctx` object from `createRouteContext()` (`src/web/server.ts:553-613`)
|
||||
that satisfies all port interfaces.
|
||||
- Validation: zod schemas in `src/web/schemas.ts`, applied via
|
||||
`parseBody(Schema, req.body)` (`src/web/route-helpers.ts:101-111`).
|
||||
- Errors: `createErrorResponse(ApiErrorCode.X, msg)` / `ApiResponse`
|
||||
(`src/types/api.ts`), auto-mapped to HTTP status by a `preSerialization` hook
|
||||
(`src/web/server.ts:644-659`).
|
||||
- SSE: `ctx.broadcast(SseEvent.X, data)` (`EventPort`,
|
||||
`src/web/sse-events.ts`); frontend mirror in `src/web/public/constants.js`.
|
||||
- **Integration point:** new `cron-routes.ts` registered alongside the
|
||||
others; new zod schema; new `SseEvent` constants for job list/run changes.
|
||||
|
||||
## 9. Where frontend pages/components live
|
||||
|
||||
- Vanilla-JS SPA: single `src/web/public/index.html` + feature mixin files
|
||||
(`Object.assign(CodemanApp.prototype, {...})`). API via `api-client.js`
|
||||
(`_apiJson/_apiPost/_apiDelete`). Build = esbuild minify + content-hash, no
|
||||
bundler (`scripts/build.mjs`).
|
||||
- UI is panels/modals toggled by JS classes; forms use `.form-row` / `.modal`
|
||||
conventions (`styles.css`). SSE handler map in `app.js`.
|
||||
- **Integration point:** add a new `cron-ui.js` mixin + a panel/modal in
|
||||
`index.html` + nav entry, following the orchestrator/respawn panel pattern.
|
||||
|
||||
## 10. Background-loop pattern (for the due-checker)
|
||||
|
||||
- Established pattern: `this.cleanup.setInterval(fn, intervalMs, {description})`
|
||||
in `WebServer.start()` (`src/web/server.ts:~1942-1966`), auto-disposed in
|
||||
`WebServer.stop()` via `this.cleanup.dispose()` (`src/web/server.ts:2336`).
|
||||
RalphLoop (`src/ralph-loop.ts:268-286`) shows the self-rescheduling guard idiom.
|
||||
- **Integration point:** register a 30s cron tick via `cleanup.setInterval`;
|
||||
no manual shutdown wiring needed.
|
||||
|
||||
---
|
||||
|
||||
## Smallest integration points (summary)
|
||||
|
||||
| New piece | Reuses | Location |
|
||||
| --- | --- | --- |
|
||||
| `CronJob` / `CronJobRun` types | — (new) | `src/types/cron.ts` |
|
||||
| Persistence | `StateStore` / `AppState` | `src/types/app-state.ts`, `src/state-store.ts` |
|
||||
| Next-run time math | — (new, pure, unit-tested) | `src/cron/cron-time.ts` |
|
||||
| Launch + send prompt | `SessionPort` (`addSession`/listeners/`writeViaMux`) | `src/cron/cron-service.ts` |
|
||||
| Background due loop | `cleanup.setInterval` pattern | `src/cron/cron-loop.ts` |
|
||||
| Routes + schema | route/ports/zod/SSE patterns | `src/web/routes/cron-routes.ts`, `src/web/schemas.ts`, `src/web/sse-events.ts` |
|
||||
| UI | panel/modal/mixin conventions | `src/web/public/cron-ui.js`, `index.html` |
|
||||
|
||||
Nothing in the session, tmux, persistence, routing, or SSE subsystems is
|
||||
rewritten — the cron is additive and calls existing services.
|
||||
@@ -1,426 +0,0 @@
|
||||
# Cron Jobs — User & Operator Guide
|
||||
|
||||
Codeman's **Cron** feature lets you save named, recurring jobs that automatically
|
||||
spin up a Claude (or shell / OpenCode / Codex / Antigravity / Gemini / Pi) session on a schedule and
|
||||
feed it a prompt. Think "cron for agent sessions": _"every weekday at 3am, open a
|
||||
Claude session in `~/proj` and tell it to update dependencies and open a PR."_
|
||||
|
||||
- **UI**: the **⏰ Cron** button in the header → the Cron Jobs modal (`#cronModal`).
|
||||
- **API**: `/api/cron/jobs*` and `/api/cron/runs`.
|
||||
- **Code**: `src/cron/cron-service.ts`, `src/cron/cron-time.ts`, `src/cron/cron-input.ts`,
|
||||
types in `src/types/cron.ts`, routes in `src/web/routes/cron-routes.ts`,
|
||||
frontend in `src/web/public/cron-ui.js`.
|
||||
|
||||
> **Not to be confused with `ScheduledRun` (`/api/scheduled`).** That older,
|
||||
> deliberately-separate concept is a _run-now, duration-bounded autonomous loop_
|
||||
> (`{prompt, workingDir, durationMinutes}` → spawn/kill throwaway sessions until
|
||||
> the duration elapses). It has no recurrence, no saved jobs, and no next-run
|
||||
> calculation. The two systems never interact. This guide is only about **Cron
|
||||
> jobs** (`Cron*`). See `docs/cron-discovery.md` §0.
|
||||
|
||||
---
|
||||
|
||||
## 1. Quick start
|
||||
|
||||
### In the browser
|
||||
|
||||
1. Click **⏰ Cron** in the header.
|
||||
2. Click **+ New Job**.
|
||||
3. Fill in a **name**, pick an **agent type** and **working directory**, choose a
|
||||
**prompt** (inline text or a file path), pick a **schedule**, and leave
|
||||
**Enabled** on.
|
||||
4. **Save**. The job appears in the list with its computed **next run**.
|
||||
5. Use **Run Now** to fire it immediately without waiting for the schedule.
|
||||
|
||||
### With curl
|
||||
|
||||
```bash
|
||||
API=http://localhost:3000
|
||||
|
||||
# Create a daily job (03:00 server-local time)
|
||||
curl -s -X POST "$API/api/cron/jobs" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"name": "nightly-deps",
|
||||
"agentType": "claude",
|
||||
"workingDir": "/home/me/proj",
|
||||
"promptMode": "inline_text",
|
||||
"promptText": "Update dependencies and open a PR",
|
||||
"inputMode": "typed",
|
||||
"scheduleType": "daily",
|
||||
"dailyTime": "03:00",
|
||||
"enabled": true,
|
||||
"concurrencyPolicy": "warn_only"
|
||||
}' | jq
|
||||
|
||||
# List jobs
|
||||
curl -s "$API/api/cron/jobs" | jq
|
||||
|
||||
# Run one immediately
|
||||
curl -s -X POST "$API/api/cron/jobs/<jobId>/run" | jq
|
||||
|
||||
# See a job's run history
|
||||
curl -s "$API/api/cron/jobs/<jobId>/runs" | jq
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Concepts
|
||||
|
||||
| Term | Meaning |
|
||||
| -------------------------- | ------------------------------------------------------------------------------------------------ |
|
||||
| **Cron job** (`CronJob`) | A saved, named definition: what agent to launch, where, with what prompt, on what schedule. |
|
||||
| **Run** (`CronJobRun`) | One execution of a job — a history record with a status and a link to the session it created. |
|
||||
| **Schedule type** | How fire times are computed: `once`, `interval`, `daily`, or `weekly`. |
|
||||
| **Next run** (`nextRunAt`) | Server-computed epoch-ms of the next fire. `null` when the job is disabled or has no future run. |
|
||||
| **Due tick** | A background loop (every 30s) that launches any enabled job whose `nextRunAt` has passed. |
|
||||
|
||||
A job is essentially a **trigger + persistence + history layer on top of the
|
||||
existing session primitives**. When a job fires, the cron service does exactly
|
||||
what the "quick start" route does — `new Session(...)` → `addSession` →
|
||||
`setupSessionListeners` → `startInteractive()`/`startShell()` → deliver the
|
||||
prompt. It does **not** reimplement any tmux/PTY logic.
|
||||
|
||||
---
|
||||
|
||||
## 3. The job form — every field
|
||||
|
||||
These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
|
||||
(`src/types/cron.ts`).
|
||||
|
||||
| Field | Required | Values / limits | Notes |
|
||||
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
|
||||
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` \| `pi` \| `grok` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. ⚠️ A `pi` or `grok` job's readiness poll looks for `❯`/a token count, which neither CLI prints, so it burns the poll budget and then sends the prompt anyway (slower start, still works). |
|
||||
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
|
||||
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
|
||||
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
|
||||
| `promptText` | conditional | ≤ 100000 chars, **single line** | Required when `promptMode = inline_text`. Newlines are rejected (see §6). |
|
||||
| `promptFilePath` | conditional | valid path | Required when `promptMode = prompt_file_path`. Confined to `workingDir` (see §5). |
|
||||
| `inputMode` | ✅ | `paste` \| `typed` | How the prompt is delivered. See §6. |
|
||||
| `scheduleType` | ✅ | `once` \| `interval` \| `daily` \| `weekly` | See §4. |
|
||||
| `runAt` | conditional | epoch-ms (positive int) | Required for `once`. |
|
||||
| `intervalMinutes` | conditional | 1–525600 (≤ 1 year) | Required for `interval`. |
|
||||
| `dailyTime` | conditional | `HH:MM` (24h) | Required for `daily`. Server-local time. |
|
||||
| `weeklyDays` | conditional | array of 1–7 ints, each 0–6 (0 = Sunday) | Required for `weekly`. |
|
||||
| `weeklyTime` | conditional | `HH:MM` (24h) | Required for `weekly`. Server-local time. |
|
||||
| `enabled` | ✅ | boolean | Disabled jobs never auto-fire (but **Run Now** still works). |
|
||||
| `notes` | — | ≤ 2000 chars | Free-form. |
|
||||
| `concurrencyPolicy` | ✅ | `warn_only` \| `skip_if_same_agent_running` | Applies to **automatic** runs only. See §7. |
|
||||
| `autoClosePreviousSession` | — | boolean (default **true**) | Recurring schedules only (ignored for `once`): when the next run fires, the still-open session created by this job's **previous** run is closed first via the normal cleanup path. See §8. |
|
||||
|
||||
**Cross-field validation** (`refineCronJob` in `schemas.ts`): the conditional
|
||||
fields above are enforced by a Zod `superRefine` on create. A missing dependent
|
||||
field (e.g. `scheduleType: "once"` with no `runAt`) is rejected with
|
||||
`INVALID_INPUT` and a field-specific message.
|
||||
|
||||
> ⚠️ **Update caveat.** `PUT /api/cron/jobs/:id` uses a `.partial()` schema that
|
||||
> does **not** re-run the cross-field `superRefine`. To keep partial edits safe,
|
||||
> `updateJob()` re-validates the **merged** job against the full `CronJobSchema`
|
||||
> and throws `400` if the result is inconsistent (e.g. switching to `once`
|
||||
> without a `runAt`). So the store is never left with a half-valid job.
|
||||
|
||||
---
|
||||
|
||||
## 4. Schedule types
|
||||
|
||||
Next-run math lives in `src/cron/cron-time.ts` (pure, unit-tested in
|
||||
`test/cron-time.test.ts`). **All wall-clock times use the server's local
|
||||
timezone** (v0.1 decision).
|
||||
|
||||
### `once`
|
||||
|
||||
- Fires a single time at the absolute `runAt` epoch-ms.
|
||||
- A **missed** one-time job (server was down at `runAt`) **still fires once** on
|
||||
the next tick — `computeNextRunAt` returns `runAt` even if it's in the past,
|
||||
until the job has fired.
|
||||
- After firing, the job **self-disables**: `completedOnce = true`, `enabled =
|
||||
false`, `nextRunAt = null`.
|
||||
|
||||
### `interval`
|
||||
|
||||
- Fires every `intervalMinutes`, computed as `fireTime + intervalMinutes`.
|
||||
- ⚠️ **Drift**: the next run re-anchors to the actual fire time, not to an ideal
|
||||
cadence — a slow tick or restart shifts subsequent runs slightly later. This is
|
||||
an accepted limitation.
|
||||
|
||||
### `daily`
|
||||
|
||||
- Fires at `dailyTime` (`HH:MM`) every day, server-local.
|
||||
- If today's time has already passed, the next run is tomorrow at that time.
|
||||
|
||||
### `weekly`
|
||||
|
||||
- Fires at `weeklyTime` on each weekday in `weeklyDays` (0 = Sunday … 6 =
|
||||
Saturday), server-local.
|
||||
- The next run is the soonest upcoming matching weekday/time within the next 7
|
||||
days.
|
||||
|
||||
---
|
||||
|
||||
## 5. Prompt source (`promptMode`)
|
||||
|
||||
### `inline_text`
|
||||
|
||||
The prompt is the literal `promptText`. Simplest option.
|
||||
|
||||
### `prompt_file_path`
|
||||
|
||||
The prompt is read from a file at fire time. **This path is security-hardened**
|
||||
because a job config is attacker-controllable and the file's contents are
|
||||
injected into an agent session (an exfiltration sink over SSE/terminal).
|
||||
`resolveSafePromptPath()` enforces, in order:
|
||||
|
||||
1. **`realpath` resolution** — symlinks are resolved to their true target, for
|
||||
the prompt file **and for `workingDir` itself**.
|
||||
2. **`workingDir` is not a trust boundary** — because it is user-supplied, the
|
||||
resolved `workingDir` is itself rejected if it is `/` or resolves into a
|
||||
blocked tree (`/etc`, `/root`, operator extras) or a pseudo-filesystem
|
||||
(`/proc`, `/sys`, `/dev`). This closes the `workingDir: '/proc'` +
|
||||
`promptFilePath: '/proc/self/environ'` env-exfil trick. The same rule is
|
||||
enforced earlier, at job create/update.
|
||||
3. **Blocklist** (defense-in-depth) — sensitive trees (`/etc`, `/root`,
|
||||
`/proc`, `/sys`, `/dev`, known secret locations) are rejected for the
|
||||
resolved prompt file.
|
||||
4. **Allowlist (primary gate)** — the resolved path **must live inside the job's
|
||||
(resolved) `workingDir`** (`validateSessionFilePath`). A symlink escaping the
|
||||
workspace fails here.
|
||||
5. **Regular-file check** — directories, FIFOs, and `/dev/*` character devices
|
||||
are rejected (they would hang or OOM an unbounded read).
|
||||
6. **Size cap** — files larger than **1 MiB** (`MAX_PROMPT_FILE_BYTES`) are
|
||||
rejected.
|
||||
7. **Single-line check** — after trailing newlines are stripped, the file
|
||||
content must be a single line (see §6).
|
||||
|
||||
If any check fails, the run is recorded as **`failed`** with the reason; no
|
||||
session is created.
|
||||
|
||||
---
|
||||
|
||||
## 6. Prompt delivery (`inputMode`)
|
||||
|
||||
Once the CLI is ready (see §8), the prompt is written to the session with a
|
||||
trailing carriage return:
|
||||
|
||||
| Mode | Mechanism | Use when |
|
||||
| ------- | --------------------------------------------------------------- | ------------------------------------------------ |
|
||||
| `typed` | `session.writeViaMux()` — tmux `send-keys -l` (literal) + Enter | Default; behaves like a human typing the prompt. |
|
||||
| `paste` | `session.write()` — writes directly to the PTY/mux | Bulk paste-style delivery. |
|
||||
|
||||
> ⚠️ **Single-line only — enforced.** Like all programmatic input in Codeman,
|
||||
> multi-line delivery would be silently corrupted (Ink-based TUIs treat a
|
||||
> newline as submit; typed mode fuses lines). So newlines are **rejected**: the
|
||||
> schema and the form refuse a multi-line `promptText`, and at fire time a
|
||||
> prompt file whose content is multi-line (after stripping trailing newlines)
|
||||
> fails the run with a clear `errorMessage`. Put multi-line instructions in a
|
||||
> file the agent is told to read itself (e.g. "read TASKS.md and do it").
|
||||
|
||||
---
|
||||
|
||||
## 7. Concurrency policy (automatic runs)
|
||||
|
||||
`concurrencyPolicy` governs what happens when a **scheduled** run is due and
|
||||
sessions of the same `agentType` already exist:
|
||||
|
||||
| Policy | Behavior |
|
||||
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `warn_only` | Always launch. (The count is surfaced but not blocking.) |
|
||||
| `skip_if_same_agent_running` | If ≥ 1 **other, live** session of that mode is active, **skip** this fire — record a `skipped` run and (for recurring schedules) advance the schedule without launching. |
|
||||
|
||||
Notes on `skip_if_same_agent_running`:
|
||||
|
||||
- Only **live** sessions block: a tab whose CLI already exited (status
|
||||
`stopped`/`error`) does not count.
|
||||
- Sessions created by **this job's own previous runs never block it** —
|
||||
otherwise a recurring job would deadlock on the session it created last time
|
||||
and fire exactly once.
|
||||
- A skipped **`once`** job is **not consumed**: it stays armed and retries on
|
||||
the next tick until the blocking session goes away, then fires its single run.
|
||||
- A skip is **not** a run: it sets `lastStatus = 'skipped'` but does **not**
|
||||
advance `lastRunAt`.
|
||||
- Consecutive skips are **coalesced** — a perpetually-skipped interval job writes
|
||||
**one** skip record per streak, not one every tick, so it can't bloat
|
||||
`state.json`.
|
||||
|
||||
**Run Now ignores this policy on the server.** The browser shows a `confirm()`
|
||||
warning if same-type sessions are active, but if you proceed (or call the API
|
||||
directly), the job launches unconditionally.
|
||||
|
||||
---
|
||||
|
||||
## 8. What happens when a job fires
|
||||
|
||||
Sequence in `CronService.launch()`:
|
||||
|
||||
1. A `CronJobRun` is created with status **`created`** and broadcast
|
||||
(`cron:runCreated`).
|
||||
2. The prompt is resolved (inline or file, single-line enforced). Failure →
|
||||
**`failed`**.
|
||||
3. `workingDir` is checked (`statSync().isDirectory()`). Missing/not-a-dir →
|
||||
**`failed`**.
|
||||
4. **Auto-close previous session** (recurring schedules, unless
|
||||
`autoClosePreviousSession: false`): any still-open session created by this
|
||||
job's previous runs is closed via the normal session-cleanup path.
|
||||
5. The global session cap is checked (`MAX_CONCURRENT_SESSIONS = 50`). At cap →
|
||||
**`failed`**.
|
||||
6. A `Session` is created **with `useMux: true`** (so it runs inside tmux),
|
||||
registered, listeners attached, and started via `startInteractive()`
|
||||
(`startShell()` for `shell` mode). Model/claudeMode come from global config.
|
||||
Run status → **`session_started`**.
|
||||
7. **Readiness wait** (async, non-blocking): for non-shell agents the service
|
||||
polls the terminal buffer up to **60 × 500ms** for a `❯` prompt or the string
|
||||
`tokens`, then settles **2000ms** (`CRON_READY_SETTLE_MS`). Shell mode waits
|
||||
1000ms, then sends the optional `launchCommand` as the first input line
|
||||
(+1000ms settle).
|
||||
8. The prompt is delivered (`typed`/`paste`, trailing `\r`). Run status →
|
||||
**`prompt_sent`**; `finishedAt` stamped. Delivery failure (e.g. the mux
|
||||
session is gone) → **`failed`**.
|
||||
|
||||
The created session is a **normal, persistent interactive session** — it appears
|
||||
as its own tab and keeps running after the prompt is sent. The run's
|
||||
`createdSessionUrl` is a deep link (`/?session=<id>`); the UI focuses it
|
||||
automatically after **Run Now**.
|
||||
|
||||
> ⚠️ **Session-cap math if you disable auto-close.** With
|
||||
> `autoClosePreviousSession: false`, nothing ever closes the sessions a
|
||||
> recurring job creates — an interval job every 30 min creates 48 tabs/day and
|
||||
> hits the global 50-session cap in ~25 hours (sooner with existing tabs), after
|
||||
> which **every** fire of **every** job fails with "Maximum concurrent sessions
|
||||
> reached" until you delete tabs by hand. Leave auto-close on for unattended
|
||||
> recurring jobs, or clean up sessions yourself.
|
||||
|
||||
### The background tick
|
||||
|
||||
`tickDueJobs()` runs every **30s** (`CRON_TICK_INTERVAL`, registered in
|
||||
`server.ts`). For each enabled job whose `nextRunAt ≤ now`:
|
||||
|
||||
- **Duplicate-launch guard**: `lastDueKey = jobId:fireTime`. If this due time was
|
||||
already consumed (overlap/restart), the job is just advanced, not relaunched.
|
||||
- The schedule is **advanced _before_ launching** so a slow launch can't be
|
||||
re-triggered by the next tick.
|
||||
- On boot, `init()` recomputes `nextRunAt` for loaded jobs (dead `once` jobs stay
|
||||
dead).
|
||||
|
||||
---
|
||||
|
||||
## 9. Run history & statuses
|
||||
|
||||
Each job keeps a history of `CronJobRun` records. Statuses (`CronJobRunStatus`):
|
||||
|
||||
| Status | Meaning |
|
||||
| ----------------- | ------------------------------------------------------------- |
|
||||
| `created` | Run record created; prompt/session not yet started. |
|
||||
| `session_started` | Session launched successfully. |
|
||||
| `prompt_sent` | Prompt delivered — the happy-path terminal state. |
|
||||
| `failed` | Something went wrong (see `errorMessage`). |
|
||||
| `skipped` | A scheduled fire was skipped by `skip_if_same_agent_running`. |
|
||||
|
||||
Each run also records `triggerType` (`scheduled` or `manual_run_now`),
|
||||
`sessionId`/`sessionName`, timestamps, and `createdSessionUrl`.
|
||||
|
||||
**History is capped globally** at **500 records** (`MAX_CRON_RUN_HISTORY`); the
|
||||
oldest are pruned first. Deleting a job also deletes its run records.
|
||||
|
||||
---
|
||||
|
||||
## 10. API reference
|
||||
|
||||
All responses use the standard `ApiResponse<T>` envelope (`{success, data}` /
|
||||
`{success, error, errorCode}`). `/api/v1/*` is a stable alias.
|
||||
|
||||
| Method | Endpoint | Body | Returns |
|
||||
| -------- | ---------------------------- | ---------------------- | --------------------------------- |
|
||||
| `GET` | `/api/cron/jobs` | — | `CronJob[]` |
|
||||
| `POST` | `/api/cron/jobs` | `CronJobSchema` | `{ job }` |
|
||||
| `GET` | `/api/cron/jobs/:id` | — | `CronJob` (404 if missing) |
|
||||
| `PUT` | `/api/cron/jobs/:id` | partial `CronJob` | `{ job }` (400 if merge invalid) |
|
||||
| `DELETE` | `/api/cron/jobs/:id` | — | `{}` |
|
||||
| `PUT` | `/api/cron/jobs/:id/enabled` | `{ enabled: boolean }` | `{ job }` |
|
||||
| `POST` | `/api/cron/jobs/:id/run` | — | `{ run, activeAgents }` |
|
||||
| `GET` | `/api/cron/jobs/:id/runs` | — | `CronJobRun[]` (newest first) |
|
||||
| `GET` | `/api/cron/runs` | — | all `CronJobRun[]` (newest first) |
|
||||
|
||||
---
|
||||
|
||||
## 11. SSE events
|
||||
|
||||
Emitted on `/api/events`, mirrored in `SSE_EVENTS` (`constants.js`):
|
||||
|
||||
| Event | Payload | When |
|
||||
| ------------------ | ------------ | -------------------------------------------------------------------- |
|
||||
| `cron:jobsChanged` | `{ jobs }` | Any job created / updated / enabled / status change. |
|
||||
| `cron:jobDeleted` | `{ id }` | A job was deleted. |
|
||||
| `cron:runCreated` | `CronJobRun` | A run (incl. skips) started. |
|
||||
| `cron:runUpdated` | `CronJobRun` | A run advanced state (`session_started` / `prompt_sent` / `failed`). |
|
||||
|
||||
---
|
||||
|
||||
## 12. State & persistence
|
||||
|
||||
Persisted in `~/.codeman/state.json` via `StateStore`:
|
||||
|
||||
- `AppState.cronJobs` — map of `id → CronJob`.
|
||||
- `AppState.cronJobRuns` — map of `id → CronJobRun`.
|
||||
|
||||
Jobs and their schedules survive restarts; `init()` recomputes `nextRunAt` on
|
||||
boot. Sessions the jobs create persist through the normal session-recovery path.
|
||||
|
||||
---
|
||||
|
||||
## 13. Limits & constants
|
||||
|
||||
| Constant | Value | Source |
|
||||
| ------------------------ | --------------------- | ------------------------------------------------ |
|
||||
| Due-tick interval | 30s | `CRON_TICK_INTERVAL` (`config/server-timing.ts`) |
|
||||
| Readiness poll | 60 × 500ms | `CRON_READY_MAX_ATTEMPTS` |
|
||||
| Readiness settle | 2000ms | `CRON_READY_SETTLE_MS` |
|
||||
| Run-history cap (global) | 500 | `MAX_CRON_RUN_HISTORY` (`config/map-limits.ts`) |
|
||||
| Saved-jobs cap | 100 | `MAX_CRON_JOBS` (`config/map-limits.ts`) |
|
||||
| Concurrent-session cap | 50 | `MAX_CONCURRENT_SESSIONS` |
|
||||
| Prompt-file size cap | 1 MiB | `MAX_PROMPT_FILE_BYTES` (`cron-service.ts`) |
|
||||
| `name` length | 1–200 | `CronJobSchema` |
|
||||
| `promptText` length | ≤ 100000 | `CronJobSchema` |
|
||||
| `intervalMinutes` | 1–525600 | `CronJobSchema` |
|
||||
| `weeklyDays` | 1–7 entries, each 0–6 | `CronJobSchema` |
|
||||
|
||||
---
|
||||
|
||||
## 14. Known limitations
|
||||
|
||||
- **Server-local timezone only** — `daily`/`weekly` times are interpreted in the
|
||||
host's local time; there is no per-job timezone.
|
||||
- **Interval drift** — `interval` re-anchors to the actual fire time; long-running
|
||||
intervals slowly shift.
|
||||
- **Single-line prompts** — multi-line prompts are rejected (schema, form, and
|
||||
at fire time for prompt files); tell the agent to read a file itself for
|
||||
multi-line instructions.
|
||||
- **`runNow` / tick race** — a manual Run Now firing at the same instant as a
|
||||
scheduled tick is theoretically possible; benign (you may get two sessions).
|
||||
- **`{enabled:true}` on a dead `once` job** — re-enabling a fired one-time job
|
||||
without changing its schedule leaves it enabled-but-dead (won't fire); change
|
||||
the schedule to re-arm.
|
||||
|
||||
---
|
||||
|
||||
## 15. Troubleshooting
|
||||
|
||||
| Symptom | Likely cause | Fix |
|
||||
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
|
||||
| Job never fires | Disabled, or `nextRunAt: null` | Check **Enabled**; verify the schedule fields are complete. |
|
||||
| Run shows `failed` immediately | Bad `workingDir`, prompt-file rejected, or session cap hit | Read `errorMessage` on the run; confirm the dir exists and the prompt file is inside it and < 1 MiB. |
|
||||
| Run shows `skipped` | `skip_if_same_agent_running` + another live same-type session (this job's own sessions and dead tabs don't count) | Switch to `warn_only`, or wait for the other session to end. |
|
||||
| Run fails with "single line" | Multi-line prompt text / prompt file | Keep the prompt to one line; point the agent at a file to read for long instructions. |
|
||||
| Sessions pile up between runs | `autoClosePreviousSession: false` | Re-enable auto-close, or delete old tabs before the 50-session cap bites (see §8). |
|
||||
| Wrong fire time | Timezone assumption | Times are **server-local** — check the host clock/TZ. |
|
||||
| One-time job won't re-fire | `completedOnce` set | Edit the schedule (any real schedule change re-arms it). |
|
||||
|
||||
---
|
||||
|
||||
## 16. Related docs
|
||||
|
||||
- `docs/cron-discovery.md` — architecture / integration-point analysis (why the
|
||||
feature reuses the session layer and stays distinct from `ScheduledRun`).
|
||||
- `docs/cron-build-brief.md` — the original build brief / requirements.
|
||||
- `CLAUDE.md` → **Key Patterns → Cron** — the one-paragraph engineering summary.
|
||||
- Tests: `test/cron-time.test.ts` (schedule math), `test/cron-service.test.ts`
|
||||
(CRUD, tick, concurrency, security).
|
||||
@@ -1,375 +0,0 @@
|
||||
# Custom Model Endpoint Profiles (all harnesses, local or cloud)
|
||||
|
||||
## Context
|
||||
|
||||
The author pays for Claude Code but also runs a capable local model behind an
|
||||
OpenAI-compatible server (llama.cpp) — and wants the same mechanism to work
|
||||
against a **cloud** OpenAI-compatible endpoint too (e.g. Azure AI Foundry's
|
||||
OpenAI-compatible inference endpoint, OpenRouter, a self-hosted gateway).
|
||||
Right now every Codeman session mode defaults to its native cloud backend
|
||||
with no way to redirect a session at any other endpoint from the UI — the
|
||||
closest existing precedent is DeepSeek's server-env-sourced
|
||||
`DEEPSEEK_BASE_URL`, which isn't user-facing.
|
||||
|
||||
**Scope note**: this plan originally said "local LLM." It now covers any
|
||||
OpenAI-compatible endpoint the user configures — local (llama.cpp, Ollama,
|
||||
vLLM) or cloud (Azure AI Foundry, OpenRouter, a company gateway). The
|
||||
mechanism is identical (a base URL Codeman probes via `GET /v1/models`); the
|
||||
only real differences are auth-header convention (cloud endpoints often want
|
||||
an `api-key` header, e.g. Azure, rather than `Authorization: Bearer`) and
|
||||
that a cloud "model" may actually be a deployment name distinct from the
|
||||
underlying model family (Azure AI Foundry deployments) — both are called out
|
||||
where they matter below. Naming throughout this plan is **"custom model
|
||||
endpoint,"** not "local model," to keep that scope explicit.
|
||||
|
||||
### Additional use case: on-premises AI hardware
|
||||
|
||||
"Local" isn't limited to a desktop running llama.cpp — a growing category of
|
||||
purpose-built, on-premises AI hardware exists specifically to run a serious
|
||||
model on-site with an OpenAI-compatible server, and this feature is exactly
|
||||
the on-ramp for pointing Codeman at one:
|
||||
|
||||
- **NVIDIA DGX Spark** (and the DGX Spark-class "Spark" mini-supercomputer
|
||||
line) — a compact on-prem inference/training box aimed at running large
|
||||
local models with an OpenAI-compatible API surface.
|
||||
- **AMD "Strix Halo" (Ryzen AI Max)** on-prem AI mini-PCs — unified-memory
|
||||
APU hardware marketed for local LLM inference, typically fronted by
|
||||
llama.cpp/Ollama/vLLM the same way a home server would be.
|
||||
|
||||
Neither needs anything new from this design: both present a standard
|
||||
`/v1/models` + `/v1/chat/completions` OpenAI-compatible surface once the
|
||||
inference server is running, so they're just another `baseUrl` entry in the
|
||||
custom-model-hosts store, same as llama.cpp or a cloud endpoint. The
|
||||
justification for building this generically (rather than hardcoding "point
|
||||
Claude at my llama.cpp box") is precisely this: **the same endpoint registry
|
||||
and per-CLI injection mechanism should work unmodified for any current or
|
||||
future OpenAI-compatible box or service** — a home GPU rig today, a Spark or
|
||||
Strix Halo appliance tomorrow, a company's on-prem inference cluster after
|
||||
that — without Codeman needing to know or care what's actually serving the
|
||||
model on the other end of that URL.
|
||||
|
||||
A concrete example worth naming: **[Ark0N/Qwen5090](https://github.com/Ark0N/Qwen5090)**
|
||||
(from the same GitHub account as this project's owner) is a one-click
|
||||
Windows / one-command Linux installer that stands up Qwen3.8-27B locally on
|
||||
an RTX 5090 (or another RTX 50-series card with ≥24GB) behind an
|
||||
OpenAI-compatible API, served by any of vLLM, NInfer, or llama.cpp — MIT-
|
||||
licensed tooling over Apache-2.0 Qwen weights. It's a direct, ready-made
|
||||
target for this feature: point a custom-model-hosts entry at whichever
|
||||
backend it's running, and it needs nothing further from Codeman's side. It's
|
||||
also notable for already wiring up DeepSeek Harness and Claude Code as
|
||||
coding agents against that local server itself, which is effectively the
|
||||
same "point a Codeman-supported harness at a local endpoint" idea this
|
||||
feature is generalizing — worth using as a real-world reference/test target
|
||||
once chunk 5 (session integration) exists, alongside the author's own llama.cpp
|
||||
box.
|
||||
|
||||
Each harness has its own (different-shaped) mechanism for pointing at a
|
||||
custom OpenAI-compatible base URL + model — env vars for Claude, a JSON
|
||||
config blob for opencode, a TOML file for Codex, etc. The author gave the
|
||||
starting recipes for those three; the rest (Gemini, Pi, Grok, DeepSeek, OMP,
|
||||
Antigravity) were researched for this plan and are flagged by confidence
|
||||
below. A real end-to-end pass against the author's own llama-swap server
|
||||
(`scripts/test-local-llm-harnesses.ts`, inside a `codeman/agent:llm-test`
|
||||
Docker image with all 9 CLIs installed) then confirmed **claude and
|
||||
opencode work end-to-end**, corrected a real Codex config.toml schema bug
|
||||
the given recipe had (see the Codex row below), and surfaced that Codex's
|
||||
_protocol_ — not just its config shape — does not work against a plain
|
||||
OpenAI-Chat-Completions server like llama.cpp/llama-swap at all. Confidence
|
||||
below reflects what was actually observed, not just what was planned.
|
||||
|
||||
The feature must be:
|
||||
|
||||
- **Off by default**, one settings toggle turns it on.
|
||||
- Endpoint entry: user gives a base URL — a LAN address or a cloud URL —
|
||||
plus an optional API key, and Codeman calls `GET <baseUrl>/v1/models` to
|
||||
discover and store the available model (or deployment) list.
|
||||
- A **new toolbar selector** (separate from the existing Run-mode menu, since
|
||||
it's a modifier on top of whichever harness is already selected/running)
|
||||
lets the user pick "Cloud (default)" — the harness's own native backend —
|
||||
or a model discovered from one of the configured custom endpoints.
|
||||
- Picking a custom-endpoint model for an **already-running session restarts
|
||||
that session's CLI process** with the injected env/config pointed at that
|
||||
endpoint (confirmed with the maintainer — these harnesses read endpoint config at
|
||||
process start, not per-turn, so a live hot-swap isn't possible).
|
||||
- **New sessions always default back to the harness's native cloud backend.**
|
||||
A custom-endpoint selection is a per-session override, not a sticky global
|
||||
default — starting a fresh CLI (any mode) always launches against its
|
||||
native backend unless the user explicitly picks a custom endpoint for that
|
||||
new session too. The toolbar selector is scoped to "this session," never
|
||||
carried forward as the default for future sessions.
|
||||
|
||||
This follows the repo's existing data-driven CLI-registry philosophy
|
||||
(`test/cli-registry-no-id-branching.test.ts`): per-CLI behavior is a
|
||||
declared capability, never an `if (mode === 'claude')` branch.
|
||||
|
||||
## Per-CLI injection recipes (confidence-ranked)
|
||||
|
||||
| CLI | Mechanism | Confidence |
|
||||
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
|
||||
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
|
||||
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol picture more nuanced than a flat break, re-verified live twice on 2026-09-17 against a llama-swap deployment that DOES answer `/v1/responses`** (an earlier test's `Reconnecting...`/`high demand` failure does not reproduce against every llama-swap setup): a plain, no-tool-call chat turn (`codex exec 'reply with just OK'`) returned a real reply. But a real tool-call attempt (`run the shell command: echo hello`) came back as an `agent_message` TEXT item — the tool-call JSON printed as the model's answer, not a `function_call` item codex would actually execute (confirmed via `codex exec --json`'s raw event stream: `item.completed`/`agent_message`, never `function_call`). Since tool execution is what makes codex a coding agent at all, this remains **not usable for real work**, just with a different, more specific failure mode than previously documented — still do not present this as working. Separately, EVERY custom-endpoint codex session also prints `warning: Model metadata for '<id>' not found. Defaulting to fallback metadata...` on launch (confirmed harmless — the successful plain-text reply above still had it): codex's per-model metadata (reasoning tiers, system-prompt templates, context-window figures) comes from `models_cache.json`, a LOCAL CACHE of OpenAI's own hosted model catalog that a custom model can never appear in by construction. No config.toml override exists for it, and the isolated `CODEX_HOME` never gets a `models_cache.json` written into it at all (confirmed: inspected a live, actively-used isolated dir — codex evidently can't reach OpenAI's catalog endpoint for this session and just falls back silently every time, with no file left behind to fix or clean up). Fabricating a fake catalog entry to suppress the warning would mean copying the _shape_ of OpenAI's own proprietary schema — including their real per-model system-prompt content, visible in a genuine `models_cache.json` — for a warning confirmed to have no effect on the actual (broken) tool-calling outcome; not worth building |
|
||||
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
|
||||
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
|
||||
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
|
||||
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`), now with `appendV1Suffix: true` (see confidence). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Root cause of the original `HTTP_404` found and fixed, by reading dsh's own bundled source — the same bar pi/grok's fixes were held to.** Installed `@deepseek-ai/dsh` (all its real published dependencies) into a scratch directory purely to read `@deepseek-ai/dsh-llm-deepseek/lib/index.js`: it builds its request as `fetch(\`${connection.baseURL}/chat/completions\`, ...)`with`baseURL`read straight from`DEEPSEEK_BASE_URL`(or defaulting to DeepSeek's real public API root,`https://api.deepseek.com`, which also carries no `/v1`) — no `/v1` insertion of dsh's own, unlike the OpenAI-SDK convention this recipe originally assumed. llama-swap/llama.cpp only ever serves the OpenAI-conventional `/v1/chat/completions`. Confirmed live: `POST <baseUrl>/chat/completions` → `404`, `POST <baseUrl>/v1/chat/completions` → `200`, on the exact same endpoint — and dsh's own error-message template, `DeepSeek API error (HTTP ${status})`, reproduces the originally reported `dsh: HTTP_404: DeepSeek API error (HTTP 404)` precisely. Fixed by adding `appendV1Suffix` (env kind only, deepseek's entry alone — claude/gemini must NOT get it, since claude was already confirmed working against the unmodified `baseUrl`), which runs `endpoint.baseUrl` through the same `withV1Suffix()` helper `configDir`-kind CLIs already use. ⚠️ Not yet re-run end-to-end with a real `dsh` binary — no install available in this environment (no npm-installed CLI binary in `PATH`, and the `codeman-test-picker` container doesn't bundle it either); the fix is source-confirmed and live-verified at the HTTP level, but a genuine "hello world" reply through `dsh` itself is the remaining step before promoting this to **verified** alongside claude/opencode/pi/grok/omp |
|
||||
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
|
||||
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
|
||||
|
||||
Everything web-researched-but-unverified gets implemented but must be
|
||||
smoke-tested against real installs of those CLIs before being called done —
|
||||
call this out explicitly when implementing, don't just ship on faith.
|
||||
|
||||
**Cloud-endpoint specifics** to keep in mind per recipe above: an Azure AI
|
||||
Foundry-style endpoint typically wants the API key in an `api-key` header
|
||||
rather than (or in addition to) `Authorization: Bearer`, and its "model" is
|
||||
often a deployment name rather than the underlying model family name — the
|
||||
discovery step (`GET /v1/models`) still works the same way against Azure AI
|
||||
Foundry's OpenAI-compatible endpoint shape, but a user may need to type the
|
||||
deployment name manually if it isn't returned as expected.
|
||||
|
||||
## Architecture
|
||||
|
||||
### 1. Registry: new `capabilities.customModelInjection` field
|
||||
|
||||
Extend `src/config/cli-registry/types.ts` / `schema.ts` with a discriminated
|
||||
union on each `CliEntry.capabilities`:
|
||||
|
||||
```ts
|
||||
type CustomModelInjection =
|
||||
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] }
|
||||
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' }
|
||||
| {
|
||||
kind: 'configDir';
|
||||
dirEnvVar: string;
|
||||
fileName: string;
|
||||
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml';
|
||||
}
|
||||
| { kind: 'unsupported' };
|
||||
```
|
||||
|
||||
Declared per stock.ts entry per the table above. A pure function in a new
|
||||
`src/custom-model-injection.ts` (`buildCustomModelInjection(entry, endpoint, modelId)`)
|
||||
turns `(CliEntry, endpoint, modelId)` into either an `envOverrides` object
|
||||
(kind `env`/`configContentEnv`) or a `{ dirEnvVar, files: [{path, content}] }`
|
||||
descriptor (kind `configDir`) — unit-testable with no IO, mirroring how
|
||||
`session-cli-builder.ts` is pure. The `configDir` kind additionally needs an
|
||||
IO wrapper that writes those files under
|
||||
`dataPath('custom-model-configs/<sessionId>/')` (new dir, cleaned up on
|
||||
session delete — same lifecycle as other per-session generated state).
|
||||
|
||||
### 2. Endpoint registry: `src/custom-model-hosts.ts`
|
||||
|
||||
Same read-array/write-array shape as `src/remote-hosts.ts` /
|
||||
`src/webview-store.ts`: `~/.codeman/custom-model-hosts.json` holding
|
||||
`CustomModelEndpoint[] = { id, label, baseUrl, apiKey?, authStyle?: 'bearer'|'api-key'|'both', models?: string[], lastDiscoveredAt? }`.
|
||||
`authStyle` defaults to `'both'` (send both header conventions on the
|
||||
discovery probe, same approach the smoke-test script below uses) so one
|
||||
endpoint entry works whether it's llama.cpp or Azure without the user having
|
||||
to know which header their box wants in advance.
|
||||
|
||||
New route file `src/web/routes/custom-model-routes.ts` (registered in the
|
||||
routes barrel), mirroring `case-routes.ts`'s remote/docker-host CRUD
|
||||
(`GET/POST/PUT/DELETE /api/model-endpoints`, admin-gated in multi-user mode
|
||||
the same way) plus:
|
||||
|
||||
- `POST /api/model-endpoints/:id/discover-models` — fetches
|
||||
`${baseUrl}/v1/models`, stores the `data[].id` list, returns it. Bounded
|
||||
timeout, and run the target through the **same SSRF egress guard already
|
||||
used for web tabs** (`webview-egress-policy.ts` — reject link-local/cloud
|
||||
metadata addresses) — this still matters for a cloud URL too, since the
|
||||
guard is about preventing a redirect to internal infra, not about
|
||||
local-vs-cloud.
|
||||
|
||||
**Why discovery rather than a free-text model field**: it removes the one
|
||||
piece of configuration most likely to trip a user up — hand-typing the
|
||||
exact model identifier a given inference server expects, which varies by
|
||||
server and is an easy source of a silent "model not found" failure with no
|
||||
useful error surfaced back through a CLI's own startup. Discovery also
|
||||
means this design is not limited to a single-model box: a **multi-model
|
||||
gateway** such as **[llama-swap](https://github.com/mostlygeek/llama-swap)**
|
||||
(hot-swaps between several loaded llama.cpp model configs behind one
|
||||
OpenAI-compatible endpoint) or a vLLM/LiteLLM/Ollama instance serving
|
||||
several models advertises ALL of them through the same `/v1/models` call —
|
||||
so one endpoint entry surfaces every model that gateway can serve, with no
|
||||
extra per-model configuration on Codeman's side at all.
|
||||
|
||||
### 3. Settings
|
||||
|
||||
- New synced boolean `customModelEndpointsEnabled` in `SettingsUpdateSchema`
|
||||
(`src/web/schemas.ts`), default `false`, documented inline like
|
||||
`readMyMindEnabled`/`workspaceHooksEnabled`.
|
||||
- New `.set-group` "Custom Model Endpoints" inside the **Agents & CLIs**
|
||||
section (`settings-clis`, `index.html:2150+`) with the enable toggle plus
|
||||
a list-editor (add/refresh-models/delete rows) for endpoints — closest
|
||||
existing precedent is the respawn-presets array editor
|
||||
(`schemas.ts:1285-1305`, `index.html:1243-1244`) for add/apply/delete-by-id
|
||||
semantics, backed by the new CRUD routes above.
|
||||
|
||||
### 4. Toolbar UI
|
||||
|
||||
> **Superseded.** This section describes the toolbar-button design as originally
|
||||
> planned. What actually shipped is a Run-menu picker instead: one generated entry
|
||||
> per (capable harness, saved endpoint) pair directly in the existing `#runModeMenu`
|
||||
> dropdown, rather than a separate `#customModelBtn`/`#customModelMenu` surface. See
|
||||
> [`docs/custom-model-endpoints.md`](custom-model-endpoints.md#the-run-menu-picker)
|
||||
> for the current design; the sections below (session-restart mechanics, security)
|
||||
> remain accurate regardless of which UI calls the underlying route.
|
||||
|
||||
- New header/toolbar button (e.g. `#customModelBtn`, `btn-toolbar
|
||||
btn-custom-model`), marker-hidden by default (`btn-custom-model--hidden`)
|
||||
and revealed by `applyHeaderVisibilitySettings()` only when
|
||||
`customModelEndpointsEnabled` is on — same pattern as the File
|
||||
Viewer/Cron buttons.
|
||||
- Clicking opens a dropdown (`#customModelMenu`, same `.run-mode-menu`-style
|
||||
markup as the existing Run-mode gear menu) listing "Cloud (default)" plus
|
||||
every discovered model, grouped by endpoint. An entry is disabled with a
|
||||
tooltip when the active session's CLI has `customModelInjection.kind ===
|
||||
'unsupported'` (Antigravity) or none declared.
|
||||
- Selecting an entry calls a new route:
|
||||
`POST /api/sessions/:id/custom-model { endpointId, modelId } | { clear: true }`.
|
||||
Server: resolve the CLI entry for `session.mode`, build the injection via
|
||||
§1, persist it as a new `session.customModel` state field (surfaced in
|
||||
`toState()`/SSE so the tab can show a small badge, e.g. "🖥 qwen3 (local)"
|
||||
or "☁ gpt-4o-mini (azure)", and the choice survives reload), merge into
|
||||
the session's `envOverrides`, and **respawn the pane's CLI process**
|
||||
through the same respawn/interactive-restart path
|
||||
`session.ts`/`tmux-manager.ts` already use for effort/model changes
|
||||
(`_configureCliEnv()` + `applyEnvOverrides()` at spawn time) — reuse,
|
||||
don't reinvent, the existing kill-and-relaunch-in-pane machinery.
|
||||
- New-session creation deliberately does **not** inherit a prior custom-
|
||||
endpoint choice: `buildEnvOverrides()` (session-ui.js) never carries the
|
||||
toolbar selection forward to the next `run()` call. Every new session
|
||||
starts on its native backend; picking a custom endpoint in the toolbar for
|
||||
a session applies only to that session (and, if done before Run is
|
||||
clicked, to the one session about to be created — not to sessions created
|
||||
afterward).
|
||||
|
||||
### 5. Multi-user security clamp
|
||||
|
||||
Every new env var this feature introduces that can redirect a session's
|
||||
traffic (and thus wherever its credentials go) — `ANTHROPIC_BASE_URL`,
|
||||
`GOOGLE_GEMINI_BASE_URL`, `GROK_BASE_URL`, the `CODEX_HOME`/`PI_CONFIG_DIR`
|
||||
dir-redirects, plus the already-privileged `DEEPSEEK_BASE_URL` — must be
|
||||
added to each CLI's `capabilities.privilegedEnvKeys` so
|
||||
`clampEnvOverridesForOwner()` strips them for a non-granted multi-user
|
||||
owner, exactly the precedent already documented for `DEEPSEEK_BASE_URL`/
|
||||
`OMP_AUTH_BROKER_URL`. This matters _more_, not less, now that endpoints can
|
||||
be cloud URLs: redirecting a non-granted user's session to an attacker's
|
||||
cloud endpoint is a credential-exfiltration path, not just a mischief
|
||||
redirect to a LAN box. Endpoint CRUD itself stays admin-only in multi-user
|
||||
mode, same as remote/docker hosts.
|
||||
|
||||
## Files touched (representative, not exhaustive)
|
||||
|
||||
- `src/config/cli-registry/types.ts`, `schema.ts`, `stock.ts` — new capability + per-entry declarations
|
||||
- `src/custom-model-injection.ts` (new) — pure per-CLI descriptor builder + unit tests
|
||||
- `src/custom-model-hosts.ts` (new) — endpoint store
|
||||
- `src/web/routes/custom-model-routes.ts` (new) — CRUD + discovery route
|
||||
- `src/web/routes/session-routes.ts` — `POST /api/sessions/:id/custom-model`, clamp wiring
|
||||
- `src/web/schemas.ts` — `customModelEndpointsEnabled`, endpoint/discover payload schemas, privileged-key updates
|
||||
- `src/session.ts` — `customModel` state field, `toState()` surface
|
||||
- `src/web/public/index.html`, `settings-ui.js`, `session-ui.js`, `styles.css` — settings group, toolbar button/menu, badge, accent CSS
|
||||
- `src/web/sse-events.ts` + `constants.js` — if a dedicated SSE event is warranted for the badge (or just ride existing session-update broadcasts)
|
||||
- `test/fixtures/mock-openai-server.ts` (new) + `test/custom-model-injection-contract.test.ts` (new) — see Mock-server validation below
|
||||
- `scripts/test-local-llm-harnesses.ts` (already added, this branch; run via `npx tsx`) — the standalone real-CLI-and-real-endpoint smoke test, supporting any `--base-url` (local or cloud). Dynamic: derives its harness list and every env var/config it injects from the live CLI registry + `buildCustomModelInjection()` rather than a second hand-maintained copy — only the one-shot invocation flags (`ONE_SHOT` table) are CLI-specific info the registry doesn't model and stay hand-maintained
|
||||
- `docs/custom-model-endpoints.md` (new) + a CLAUDE.md pointer bullet under External CLI modes / envOverrides
|
||||
|
||||
## Mock-server validation strategy (CI-runnable, no real CLI binaries needed)
|
||||
|
||||
Spawning nine real CLI binaries in CI isn't realistic, and neither the author's
|
||||
llama.cpp box nor a real cloud subscription can be a CI dependency. So the
|
||||
injection _logic_ gets a tier of automated coverage that sits between the
|
||||
pure unit tests and the live manual checks in Verification:
|
||||
|
||||
1. **`test/fixtures/mock-openai-server.ts`** — a small in-process HTTP
|
||||
server (plain `http.createServer`, no external deps, port picked per the
|
||||
existing `const PORT = 3150+` convention) that:
|
||||
- Serves `GET /v1/models` → a fixed fake model list (`{data:[{id:'qwen3'},...]}`),
|
||||
for testing the discovery route.
|
||||
- Serves `POST /v1/chat/completions` (OpenAI shape) **and**
|
||||
`POST /v1/messages` (Anthropic Messages-API shape, since that's what
|
||||
`ANTHROPIC_BASE_URL` traffic looks like) and records every request it
|
||||
receives (headers, body, path) into an array the test can assert on —
|
||||
including which auth header style it saw, so the `authStyle: 'both'`
|
||||
default and Azure's `api-key` convention both get real coverage.
|
||||
- Returns a minimal valid completion so a client library doesn't choke
|
||||
on the response shape.
|
||||
|
||||
2. **`test/custom-model-injection-contract.test.ts`** — for every CLI with a
|
||||
`customModelInjection` capability (i.e. every row in the table above
|
||||
except `antigravity`):
|
||||
- Point a fixture `CustomModelEndpoint` at the mock server's URL.
|
||||
- Call `buildCustomModelInjection(entry, endpoint, modelId)` (the pure
|
||||
function from §1) to get the real env vars / config-file content that
|
||||
would be injected into that CLI's session.
|
||||
- Replay those exact values through a minimal HTTP request shaped the
|
||||
way that CLI is documented to send it (Anthropic Messages shape for
|
||||
claude; OpenAI chat-completions shape for opencode/codex/pi/grok/omp;
|
||||
`GOOGLE_GEMINI_BASE_URL`'s OpenAI-compat shape for gemini; dsh's
|
||||
provider call for deepseek) against the mock server.
|
||||
- Assert the mock server received the request **at the injected
|
||||
`baseUrl`**, with **the injected API key** in the expected header, and
|
||||
**the injected model id** in the body/path — i.e. prove the values
|
||||
Codeman computes are internally consistent and would reach the right
|
||||
place with the right identifiers, end to end, in CI, on every push.
|
||||
- Also cover the `configDir` kind (codex/pi/omp): assert the written
|
||||
`config.toml`/`models.json`/`models.yml` file parses and contains the
|
||||
same base URL/key/model, and that it's written under the isolated
|
||||
per-session dir rather than the user's real config path.
|
||||
|
||||
3. **Explicit, stated limitation** (goes in the test file's `@fileoverview`
|
||||
and in this doc, not left implicit): this proves _"if the CLI honors its
|
||||
documented env/config contract, it will hit the right endpoint with the
|
||||
right model."_ It does **not** prove the real CLI binary actually reads
|
||||
that env var / config file the way its docs say — that's still the job
|
||||
of the live manual checks in Verification step 4-5 below, and is exactly
|
||||
why the confidence table above did not stop at "researched" — every CLI
|
||||
except antigravity (no mechanism at all) has since been run against a
|
||||
real llama-swap server via `scripts/test-local-llm-harnesses.ts`:
|
||||
claude/opencode/pi/grok/omp are confirmed PASS end-to-end, codex is
|
||||
confirmed FAIL for a real documented protocol reason (Responses-API-only
|
||||
since Feb 2026), and gemini/deepseek are confirmed reaching the server
|
||||
but failing for reasons not yet root-caused (see their table rows). The
|
||||
mock-server suite catches regressions in Codeman's own logic; it cannot
|
||||
catch a CLI changing its env-var name in a future release, or a real
|
||||
cloud endpoint behaving differently from a local llama.cpp box.
|
||||
|
||||
## Verification
|
||||
|
||||
1. `npm run typecheck && npm test` after each slice — this now includes the
|
||||
mock-server contract suite from above, so injection-logic regressions
|
||||
are caught automatically without touching real infrastructure.
|
||||
2. Unit tests for `buildCustomModelInjection()` per CLI kind (pure, no IO).
|
||||
3. Route tests (`app.inject`) for the new CRUD + discover-models endpoint
|
||||
(mock `fetch` for `/v1/models`), and for the multi-user clamp on the new
|
||||
privileged keys (mirror `test/routes/external-cli-bypass-clamp.test.ts`).
|
||||
4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.ts`
|
||||
exercises every harness the CLI registry declares `customModelInjection`
|
||||
support for against a real `--base-url` — local or cloud — outside of
|
||||
Codeman's UI entirely, and is DYNAMIC (reads `enabledClis()` + calls the
|
||||
real `buildCustomModelInjection()`, so a future registry change is picked
|
||||
up automatically with zero edits to the script). Already run to
|
||||
completion against the author's llama-swap server (a LAN address,
|
||||
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
|
||||
claude/opencode/pi/grok/omp **PASS**, codex **partially works and still
|
||||
isn't usable** (plain chat succeeds against a llama-swap deployment that
|
||||
answers `/v1/responses`, but a real tool-call attempt comes back as
|
||||
inert text rather than an executable `function_call` — see the
|
||||
confidence table row for the full, re-verified picture), gemini/deepseek
|
||||
**UNCONFIRMED**
|
||||
(reach the server, fail for undiagnosed reasons — see their table rows),
|
||||
antigravity **SKIP** (no mechanism). Re-run this against a real cloud
|
||||
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
|
||||
prove the `authStyle`/deployment-name handling holds up outside llama.cpp.
|
||||
5. Once the full feature (not just the standalone script) is built: add an
|
||||
endpoint via the real UI, hit discover-models, confirm the returned model
|
||||
list, pick Claude + the model on a real session, confirm via
|
||||
`tmux -L codeman capture-pane`/`tmux showenv -t <pane>` that
|
||||
`ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`/`ANTHROPIC_DEFAULT_*_MODEL` are
|
||||
set post-restart, and confirm the endpoint's own logs show the next
|
||||
prompt actually landing there. Repeat for opencode and Codex at minimum
|
||||
before considering this shippable; spot-check the web-researched CLIs
|
||||
and correct the plan's confidence table with what's actually observed.
|
||||
6. `npm run lint && npm run format:check`.
|
||||
7. Update `CHANGELOG.md`/changeset per the COM workflow when shipping.
|
||||
@@ -1,563 +0,0 @@
|
||||
# Custom Model Endpoint Profiles
|
||||
|
||||
Point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi,
|
||||
Grok, DeepSeek, or OMP — at a custom OpenAI-compatible endpoint instead of
|
||||
its native cloud backend, for a given session. "Custom endpoint" covers both
|
||||
**local** hardware (llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built
|
||||
boxes like NVIDIA DGX Spark or AMD Strix Halo mini-PCs) and **cloud**
|
||||
services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
|
||||
company gateway) — anything answering `GET /v1/models` and
|
||||
`POST /v1/chat/completions` in the standard shape. Design doc, per-CLI
|
||||
recipe confidence table, and security reasoning:
|
||||
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
|
||||
|
||||
> **Status**: fully wired end to end — registry capability, the injection
|
||||
> engine, the endpoint store + discovery route, both the restart-in-place
|
||||
> apply route (Claude) and the one-shot quick-start launch path (every
|
||||
> other supported harness), a settings-panel CRUD surface, and the Run-menu
|
||||
> picker described below. Antigravity has no known custom-endpoint
|
||||
> mechanism and is not supported. The HTTP API (examples below) still works
|
||||
> directly and is what the picker itself calls under the hood.
|
||||
|
||||
## Turning it on
|
||||
|
||||
App Settings → Models → **Custom model endpoints** (synced setting
|
||||
`customModelEndpointsEnabled`, default **OFF**). Turning it on does two
|
||||
things: it reveals the endpoint list/add/edit/discover panel in that same
|
||||
settings section, and it makes the Run menu offer a generated entry per
|
||||
(harness, endpoint) pair — see "The Run-menu picker" below. The API
|
||||
equivalent:
|
||||
|
||||
```bash
|
||||
curl -sk -X PUT https://localhost:3000/api/settings \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"customModelEndpointsEnabled": true}'
|
||||
```
|
||||
|
||||
## Adding an endpoint
|
||||
|
||||
Via App Settings → Models → Custom model endpoints → **+ Add endpoint**, or
|
||||
directly:
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/model-endpoints \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"id": "llama-box", "label": "Home llama.cpp", "baseUrl": "http://192.168.1.50:8080"}'
|
||||
```
|
||||
|
||||
`apiKey` is optional (most local servers don't check it). `authStyle`
|
||||
(`bearer` | `api-key`, default `bearer`) controls which auth header
|
||||
convention discovery uses: `bearer` is `Authorization: Bearer <key>`
|
||||
(llama.cpp, OpenAI-compatible servers, most gateways), `api-key` is the
|
||||
`api-key: <key>` header Azure AI Foundry wants. There is deliberately no
|
||||
"send both" option: measured against a real llama-swap server, a request
|
||||
carrying both headers hung indefinitely. `baseUrl` must be `http(s)`, carry
|
||||
no embedded credentials, and may not point at a link-local or cloud-metadata
|
||||
address; discovery re-checks the address the name actually resolves to.
|
||||
|
||||
Discover its available models:
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/model-endpoints/llama-box/discover-models
|
||||
```
|
||||
|
||||
This calls the endpoint's own `GET /v1/models` and stores the returned list
|
||||
on the endpoint record; `GET /api/model-endpoints` lists everything
|
||||
configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
|
||||
Endpoint management is admin-only in multi-user mode, same as remote/docker
|
||||
hosts — these are machine-level infra, not per-user settings.
|
||||
|
||||
**Context length is discovered too, opportunistically and safely.** The plain
|
||||
`GET /v1/models` response has no context-window field. Discovery only ever
|
||||
looks for one for a model llama-swap's own response already reports
|
||||
`status.value === "loaded"` for — never for an unloaded one, because
|
||||
llama-swap treats `?model=` as a routing hint and asking about a model that
|
||||
isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side
|
||||
effect of what should be read-only discovery. A server with no `status` field
|
||||
on any entry at all (not llama-swap) gets no context-length enrichment,
|
||||
rather than guessing. A model's previously-learned context length survives a
|
||||
later cycle where it wasn't the loaded one; it's dropped only once the model
|
||||
disappears from the endpoint's list entirely. Stored per model in
|
||||
`modelContextLengths` and applied automatically (see "Applying a model to a
|
||||
session" below) so a CLI that would otherwise assume a large default context
|
||||
window for an unrecognized model id stops silently overflowing a much
|
||||
smaller real one.
|
||||
|
||||
**Where that number actually comes from matters, and got this wrong once
|
||||
already.** The first cut read it from llama.cpp's own
|
||||
`GET /props?model=<id>` (`n_ctx`) — plausible, and it worked in testing, but
|
||||
confirmed live to be actively WRONG for a `--fit-ctx`-launched llama-swap
|
||||
backend: `/props` reported `n_ctx: 154112` for a model llama-swap itself had
|
||||
launched with `--fit-ctx 16384`, and the real server then refused a request
|
||||
right at that real 16384-token limit — `/props`'s `n_ctx` appears to report
|
||||
the model's theoretical/trained maximum there, not the runtime-configured
|
||||
one. Discovery now parses the REAL configured size straight out of
|
||||
llama-swap's own launch command instead (`GET /running`'s `cmd` field —
|
||||
`--fit-ctx <N>` first, then the plain llama.cpp `-c`/`--ctx-size` a
|
||||
hand-written command might use), and only falls back to the `/props` probe
|
||||
when `cmd` states no recognizable flag at all.
|
||||
|
||||
**File size is discovered too, when the server states one.** llama-swap
|
||||
writes a GB figure into an auto-discovered model's own `description`
|
||||
(`"Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"`), parsed
|
||||
into `modelSizesGB` — unlike context length, this needs no `/props` probe
|
||||
(the figure is right there in the `/v1/models` response) and so is populated
|
||||
for every model regardless of loaded state. A hand-configured profile's own
|
||||
description has no such figure and correctly gets no entry, never a guess.
|
||||
Used only to label the Run-menu picker's "loading model" banner (e.g.
|
||||
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB) on llama-swap..."); never anything
|
||||
a server-side check relies on.
|
||||
|
||||
**The loading banner is unbounded by design, and says so — no countdown, no
|
||||
automatic give-up.** An earlier version scaled an expected-time estimate and
|
||||
a timeout off the model's file size and auto-closed the session once that
|
||||
elapsed, but a real load's actual duration depends on hardware this feature
|
||||
has no way to know (VRAM, storage speed, whatever else is contending for the
|
||||
GPU) — any fixed number was a guess dressed up as a fact, and a model that
|
||||
genuinely takes 10+ minutes on slower hardware would just get killed
|
||||
mid-load by its own display. The banner now says outright that it can take a
|
||||
while depending on hardware and model size, polls
|
||||
`GET /api/model-endpoints/:id/running-status` every second for as long as it
|
||||
takes, and carries a **Cancel** button (rendered on the banner itself) that
|
||||
ends the wait and closes the session the load was for — the user's own call
|
||||
on when it's taking too long, not a fixed number baked into the client.
|
||||
|
||||
**The banner's second line is the real backend log line, not a guess.**
|
||||
llama-swap's `GET /api/events` SSE stream carries the actual `llama-server`
|
||||
process's own stdout — `load_model: loading model '<path>'`,
|
||||
`llama_server: model loaded`, tokenizer warnings, all of it — tagged
|
||||
`source: "upstream"`, distinct from llama-swap's own `source: "proxy"`
|
||||
request-access lines. `running-status`'s response now includes `logLine`
|
||||
(via `getLatestLlamaSwapLogLine`), and the banner shows it on its own line
|
||||
under the disclaimer, e.g. "llama.cpp: load_model: loading model '...'" —
|
||||
confirmed live end-to-end through a real forced swap, sequentially showing
|
||||
the model path, a tokenizer warning, then staying on whatever llama.cpp last
|
||||
printed once the load goes quiet (never cleared back to blank). ⚠️
|
||||
**`GET /logs` — the endpoint this feature's own first cut was built
|
||||
against — turns out to carry ONLY llama-swap's own proxy request-access
|
||||
log.** Confirmed live it never showed a single backend line, even seconds
|
||||
after a real, verified model swap; `/api/events`'s `logData` frames are the
|
||||
only source that actually has it, and its own `source` field (`upstream` vs
|
||||
`proxy`) is what `getLatestLlamaSwapLogLine` filters on. One `/api/events`
|
||||
connection is held open per endpoint and reused across every session
|
||||
watching a load on it (confirmed live to stay open indefinitely, unlike
|
||||
`/logs`, which closes after a fixed ~100KB), idle-closed after 30s of nobody
|
||||
polling it (`pruneIdleLlamaSwapLogTails`, same 20s sweep as the
|
||||
swap-displacement check below).
|
||||
|
||||
`defaultModelId` names which discovered model the picker pre-marks for that
|
||||
endpoint — the settings panel's Edit form exposes it as a select populated
|
||||
from the endpoint's own discovered `models`, and the route refuses a value
|
||||
that isn't one of them. It is applied automatically only when the endpoint
|
||||
has exactly one discovered model (nothing to choose); with two or more it
|
||||
is a pre-selection in the model-picker dialog below, never a silent default.
|
||||
Re-discovering drops a default that no longer appears in the fresh list
|
||||
rather than carrying an invalid one forward.
|
||||
|
||||
**Model lists refresh themselves.** A background sweep (`server.ts`,
|
||||
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
|
||||
saved endpoint the same way the manual `POST .../discover-models` route
|
||||
does, best-effort per endpoint — one being unreachable on a given cycle
|
||||
never blocks the others. Off under `npm test`, same reasoning as the Codex
|
||||
plan-usage poll it sits beside: no real network to hit, no server instance
|
||||
to keep the timer alive for.
|
||||
|
||||
## The Run-menu picker
|
||||
|
||||
With the setting on and at least one endpoint carrying a discovered model,
|
||||
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
|
||||
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
|
||||
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
|
||||
registry's own `capabilities.customModelInjection` at page render
|
||||
(`window.__codemanCustomModelClis`, `server.ts`) — never a hardcoded id list
|
||||
in the frontend — so a CLI whose injection recipe lands later shows up with
|
||||
no frontend change, and Antigravity (`unsupported`) never does.
|
||||
|
||||
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
|
||||
`session-ui.js`) rather than trusting anything cached from the dropdown's
|
||||
own render — the model list can have changed via the 5-minute sweep above
|
||||
or a settings-panel edit since the menu opened. With exactly one discovered
|
||||
model it runs straight away; with two or more, a small modal
|
||||
(`#customModelPickModal`) lists them and asks which one to use for this
|
||||
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
|
||||
the point of asking is letting one launch deliberately differ from the
|
||||
saved default, not just confirming it.
|
||||
|
||||
The modal promotes exactly one row to the top of the list rather than
|
||||
always showing raw discovery order, so the zero-wait choice is the one
|
||||
under your thumb:
|
||||
|
||||
- **"Currently loaded"** — a model from this host's own list that
|
||||
llama-swap reports `ready` right now, queried via
|
||||
`GET /api/model-endpoints/:id/running-status`. Bounded client-side to
|
||||
~800ms (`Promise.race`), on top of the route's own 5s server-side
|
||||
timeout, so an endpoint that is asleep or firewalled cannot leave the
|
||||
modal invisible for the full 5s after the Run menu has already closed.
|
||||
- **"Last used"** — shown only when nothing is currently loaded: the model
|
||||
actually launched last for this exact (harness, endpoint) pair, read
|
||||
from the per-device `codeman:customModelLastUsed:<mode>:<endpointId>`
|
||||
localStorage key. Written by `_runCustomModelEntryViaRestart` (claude)
|
||||
and `_quickStartWithCustomModelConfirm` (every one-shot launch; the
|
||||
`runCustomModelEntry` entry point itself only dispatches between the
|
||||
two) only once the model is actually applied, never on the mere click —
|
||||
declining the context-window warning means this exact model cannot work
|
||||
with this CLI at all, so promoting it next time would be actively wrong,
|
||||
not just premature.
|
||||
|
||||
Neither tag reorders anything past that one promoted row. The "Default"
|
||||
pill is a separate span, not a third value of the same slot: a promoted
|
||||
row that is also the endpoint's `defaultModelId` shows both tags (on a
|
||||
single-purpose GPU box that is the common case, and an exclusive slot
|
||||
silently dropped the Default marking for exactly that row), and a row
|
||||
with neither promotion nor default shows no tag at all.
|
||||
|
||||
**How the launch itself applies the endpoint depends on the harness.** For
|
||||
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP (`runCustomModelEntry` →
|
||||
`_runCustomModelEntryOneShot`), the endpoint/model is folded into the SAME
|
||||
`POST /api/quick-start` call that creates the session (`customModel` field),
|
||||
so the session launches directly on the endpoint — no restart, no visible
|
||||
relaunch. Claude (`_runCustomModelEntryViaRestart`) still uses the original
|
||||
two-step design: the launch runs a single native session exactly the way its
|
||||
own Run-menu entry would, then **waits for the new session to go idle**
|
||||
(`GET .../wait?until=idle`, bounded at 20s — a normal 200 either way, never
|
||||
an error, per the wait endpoint's own contract) before applying the endpoint
|
||||
via the restart route below. That wait exists because a freshly launched CLI
|
||||
reports itself as `busy` for its own startup (a boot spinner, a
|
||||
workspace-trust check) well before the apply call would otherwise reach it,
|
||||
and the apply route correctly refuses to restart a session mid-turn — a
|
||||
fresh boot looks exactly like one from the outside. A session still busy
|
||||
after the wait reaches the apply call anyway and gets that route's own
|
||||
honest `SESSION_BUSY` error, now visible as a sticky toast with a close
|
||||
button rather than a generic message that vanished in three seconds. Claude
|
||||
stays on this path because its own restart (`--resume`-based, keeping the
|
||||
conversation) is far less jarring than the other seven's, and `runClaude()`'s
|
||||
multi-tab launch and docker-config-drift confirm/retry loop make folding it
|
||||
into the one-shot path separate work. It is a
|
||||
one-off "try this endpoint" action, not a sticky mode: the plain Run button
|
||||
still means "this harness, native cloud" afterward. Entries are hidden
|
||||
entirely for a remote or Docker active case, since the apply route refuses
|
||||
both (see the next section).
|
||||
|
||||
## Launching directly on an endpoint (no restart)
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/quick-start \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"caseName": "myapp", "mode": "codex", "customModel": {"endpointId": "llama-box", "modelId": "qwen3"}}'
|
||||
```
|
||||
|
||||
`POST /api/quick-start`'s `customModel` field (`{endpointId, modelId,
|
||||
confirmed?}`) computes the same injection the restart route below does, but
|
||||
BEFORE the session exists — the session is minted its own id up front
|
||||
(`crypto.randomUUID()`), the injection (env vars, and for a `configDir`-kind
|
||||
CLI, the written config file) targets that real id, and the session launches
|
||||
already pointed at the endpoint. No restart, because there was never a
|
||||
native-backend launch to restart away from. Runs the same llama-swap
|
||||
conflict check as the restart route (below) — a `409`-shaped
|
||||
`{requiresConfirmation, currentlyLoadedModel, affectedSessions}` response
|
||||
with no session created, resolved by retrying with `confirmedSwap: true` — and
|
||||
is refused the same way for a remote or Docker case. This is what the
|
||||
Run-menu picker uses for opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP;
|
||||
Claude still uses the restart route below (see "The Run-menu picker" above
|
||||
for why).
|
||||
|
||||
## Applying a model to an ALREADY-RUNNING session
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"endpointId": "llama-box", "modelId": "qwen3"}'
|
||||
```
|
||||
|
||||
This computes the CLI-specific env vars / config for that session's mode
|
||||
(see the recipe table in `custom-model-endpoints-plan.md`) and **restarts the session's
|
||||
CLI process in place** — same pane, same tmux session, fresh env. That
|
||||
restart is necessary, not incidental: every supported harness reads its
|
||||
endpoint config at process start, not per-turn, so there is no live
|
||||
hot-swap. A Claude session is relaunched with `--resume <conversation> ||
|
||||
--session-id <id>`, so it continues the conversation it was on; pi, omp and
|
||||
grok are relaunched with the `--model` value that selects the injected
|
||||
provider (`custom/<modelId>` for pi and omp, `codeman-custom` for grok),
|
||||
since for those three the config file alone does not switch the model.
|
||||
**Remote (SSH) and Docker sessions are refused** (400) for now: their restart
|
||||
reattaches the durable remote/in-container tmux rather than relaunching the
|
||||
agent, so the selection would report success and change nothing.
|
||||
|
||||
**Claude gets two more env vars when known/applicable, both declared on its
|
||||
registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
|
||||
|
||||
- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context
|
||||
length (see the discovery section above) whenever one is known. Without
|
||||
it, Claude Code assumes a large (200k) window for any unrecognized custom
|
||||
model id and never compacts, which reliably overflows a much smaller real
|
||||
local context — confirmed live: a stock ~33.7K-token system prompt against
|
||||
a 16384-token llama-swap model failed with `exceeds the available context
|
||||
size`. No entry for the model in `modelContextLengths` means the var is
|
||||
simply omitted, never a guess. ⚠️ **This var only affects when Claude
|
||||
Code compacts conversation _history_ — it cannot fix a model whose real
|
||||
context is smaller than Claude Code's own fixed per-turn overhead**
|
||||
(system prompt + tool schemas, empirically ~36.4K tokens, confirmed live
|
||||
via an `in:0 out:0` failure on the very first message, before any
|
||||
history exists to compact). No context-length declaration changes that
|
||||
fixed overhead, so a model below the safe floor fails outright on
|
||||
message one regardless of what this var says. See "Context-window floor
|
||||
warning" below for how Codeman catches this case before launching
|
||||
instead of after.
|
||||
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
|
||||
the `configDir`-kind CLIs use (empty, no files written into it), so the
|
||||
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
|
||||
claude.ai OAuth login. Claude Code still prints "Both claude.ai and
|
||||
ANTHROPIC_API_KEY set" when the two coexist in the same config directory —
|
||||
cosmetic (confirmed live: the API key wins for actual requests either way,
|
||||
visible in the terminal's own `API Usage Billing` line) but worth
|
||||
eliminating rather than living with. The directory's `projects`
|
||||
subdirectory is symlinked (a junction on Windows) back to the real
|
||||
`~/.claude/projects` so the response viewer, subagent windows and Read My
|
||||
Mind keep working for that session — the same trade-off and fix documented
|
||||
for a manually-set `CLAUDE_CONFIG_DIR` in
|
||||
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
|
||||
automatically here. Best-effort: a platform that refuses the symlink keeps
|
||||
the pre-existing blind-response-viewer side effect rather than failing the
|
||||
whole custom-model apply over it. ⚠️ **This relocates the whole `.claude`
|
||||
tree, not just transcripts**: a custom-model Claude session also loses the
|
||||
user's global `settings.json`, user-level skills (the codeman agent skill
|
||||
included), user-level agents and commands, and the MCP servers configured
|
||||
in `~/.claude.json` — none of those are symlinked back, only `projects` is.
|
||||
A fine trade for "point this session at my local llama.cpp," but worth
|
||||
knowing before it surprises you mid-session.
|
||||
|
||||
**That isolated directory needed one more fix to actually be usable
|
||||
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
|
||||
real profile's prior "Detected a custom API key — use it?" approvals, so
|
||||
without more, Claude Code stops and asks that on _every single launch_ —
|
||||
confirmed live, and with nobody at a TTY to answer, its own default answer
|
||||
("No") silently refuses the very key this feature just injected, which
|
||||
looks like the endpoint being ignored entirely. `customModelInjection`'s
|
||||
`apiKeyTrustFile` (`{ relPath: '.claude.json', shape:
|
||||
'claude-api-key-responses' }` on claude's entry) pre-seeds that exact
|
||||
approval: the apply step merges `customApiKeyResponses.approved: [apiKey]`
|
||||
into `<configDir>/.claude.json`, the same field a real answered prompt
|
||||
itself writes to (confirmed against a real file after answering by hand
|
||||
once) — this answers the prompt in advance rather than bypassing it. The
|
||||
merge preserves whatever else the CLI already wrote into that file on an
|
||||
earlier launch in the same isolated directory (`userID`, `numStartups`,
|
||||
earlier approved keys), and a missing or corrupt file is treated as empty
|
||||
rather than failing the apply.
|
||||
|
||||
**A fresh `CLAUDE_CONFIG_DIR` isn't just missing that one approval — Claude
|
||||
Code treats it as a brand-new profile and replays its ENTIRE first-run
|
||||
sequence on every launch: the theme picker, the security-notes screen, the
|
||||
per-project "trust this folder?" dialog, and (running with
|
||||
`--dangerously-skip-permissions`) a one-time warning about bypassing
|
||||
permissions.** Confirmed live: none of these show up again for a real,
|
||||
already-onboarded profile, but every custom-model session gets a fresh,
|
||||
otherwise-empty isolated directory, so it saw all four every single time.
|
||||
`customModelInjection`'s `skipFirstRunPrompts` (`true` on claude's entry,
|
||||
requires `apiKeyTrustFile` since it reuses the same file) pre-seeds the
|
||||
state a real profile accumulates from answering all of that once:
|
||||
`hasCompletedOnboarding: true` and the launching session's own
|
||||
`projects[workingDir].hasTrustDialogAccepted: true` go into the same
|
||||
`<configDir>/.claude.json` the API-key approval above already merges into
|
||||
(other projects, and other fields on this session's own project entry, are
|
||||
left untouched), and `skipDangerousModePermissionPrompt: true` goes into
|
||||
`<configDir>/settings.json` — a different file, merged the same
|
||||
corrupt-tolerant way. `workingDir` is used exactly as the session was
|
||||
launched with as its cwd, never realpath'd or slash-normalized, since
|
||||
that's the literal string Claude Code itself uses as the project key.
|
||||
|
||||
**llama-swap gets two more fixes on top of the context-length/config-dir
|
||||
ones above, both from watching a real switch live.** llama.cpp only ever
|
||||
runs one model at a time; llama-swap swaps the backing process on demand,
|
||||
which can take anywhere from a few seconds to well over a minute:
|
||||
|
||||
- **The conflict check.** Both apply routes (the restart one here and the
|
||||
one-shot `POST /api/quick-start` above) call llama-swap's own
|
||||
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
|
||||
compatible server (no such endpoint) is simply never checked. If a
|
||||
_different_ model is currently loaded and ready, and another **live
|
||||
session's own selection** is using it, the apply returns
|
||||
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
|
||||
instead of silently switching — nothing is applied or created yet.
|
||||
Retrying with `confirmedSwap: true` skips the check (the legacy `confirmed: true`
|
||||
still means both questions). Switching with nothing
|
||||
else affected proceeds immediately; this is a warning about disrupting
|
||||
another session, never a gate on the switch itself.
|
||||
- **Actually starting the load.** llama-swap has no "switch model" admin
|
||||
call — the only thing that starts a swap is a real inference request
|
||||
naming the model, and confirmed live: applying a selection alone never
|
||||
reached llama-swap at all (nothing in its own server logs), since nothing
|
||||
had actually asked it to load anything yet. Both apply routes now also
|
||||
send the smallest real request that will —
|
||||
`POST <baseUrl>/v1/chat/completions` with `max_tokens: 1` and one
|
||||
throwaway message — whenever the
|
||||
target model isn't already the one loaded and ready, fire-and-forget (its
|
||||
response is never read; `GET /api/model-endpoints/:id/running-status`,
|
||||
polled client-side, is what actually confirms readiness). The response
|
||||
also carries `modelSwapInProgress: true` in that case, which is what
|
||||
drives the Run-menu picker's own "loading model" status banner.
|
||||
|
||||
## Catching a swap after the fact
|
||||
|
||||
The conflict check above only runs at the moment a session is created or a
|
||||
model is applied — it has no way to catch a swap that happens **later**.
|
||||
Confirmed live: a session created while nothing else conflicted at that
|
||||
exact instant can still get silently displaced afterward, once a
|
||||
_different_ session's own normal use (or its own create-time load trigger)
|
||||
asks llama-swap to load something else. llama-swap has no push
|
||||
notification of its own for this, so a background sweep
|
||||
(`detectCustomModelSwapDisplacements`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS`
|
||||
= 20s in `server.ts`) polls `GET /running` once per distinct endpoint that
|
||||
has at least one live custom-model session, and compares each such
|
||||
session's own `modelId` against what is actually loaded. A session whose
|
||||
model is no longer in that list gets a `custom-model:swapped-out` SSE event
|
||||
(`{sessionId, sessionName, endpointId, previousModel, currentlyLoadedModel}`),
|
||||
shown as a global toast — global rather than tied to that session's tab,
|
||||
since the whole point is telling the user before they type into it
|
||||
expecting the model they picked. Notifies **once per displacement**: the
|
||||
same de-dupe `Set` clears a session's flag once its own model is loaded and
|
||||
ready again, so a later, genuinely new displacement notifies again rather
|
||||
than the session staying silently un-notified forever after the first one.
|
||||
|
||||
## Context-window floor warning
|
||||
|
||||
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
|
||||
empirically ~36.4K tokens) can exceed a small local model's _entire_ real
|
||||
context on its own, before any conversation history exists to fill it —
|
||||
confirmed live twice, both as an `in:0 out:0` failure on the very first
|
||||
message sent. `CLAUDE_CODE_MAX_CONTEXT_TOKENS` (above) cannot fix this: it
|
||||
only governs when Claude Code compacts conversation history, and there is
|
||||
no history yet on message one. Applying such a model would look like the
|
||||
endpoint being ignored, or the wrong model being used, when in fact the
|
||||
endpoint applied correctly and the model is simply too small for this CLI.
|
||||
|
||||
Both apply routes (the restart route and the one-shot `POST
|
||||
/api/quick-start`) now check for this **before** launching or restarting
|
||||
anything, gated on the CLI's registry entry declaring a `contextLengthVar`
|
||||
(currently only claude — the check is a no-op for every other CLI by
|
||||
construction, never a hardcoded mode check). If the model's discovered
|
||||
context (`modelContextLengths`, from discovery above) is below
|
||||
`CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000, comfortably above the measured
|
||||
~36.4K overhead), the response is `{requiresContextWarning: true, modelId,
|
||||
contextLength, minSafeContextTokens}` instead of applying — nothing is
|
||||
restarted or created yet. A context length that was never discovered at
|
||||
all skips the check entirely (nothing to compare, so it fails open rather
|
||||
than warning on every model an endpoint hasn't reported a size for).
|
||||
Retrying with `confirmedContext: true` launches anyway (the legacy `confirmed: true` still means both questions).
|
||||
|
||||
The Run-menu picker shows this as an in-app modal
|
||||
(`#customModelContextWarningModal`, matching the llama-swap conflict
|
||||
modal's look) naming the model, its discovered context, and the safe
|
||||
floor, and explaining the fix: reconfigure llama-swap to give that model
|
||||
(or a smaller one) an explicit larger context instead of relying on
|
||||
auto-fit (`--fit-ctx`), which optimizes for the biggest _model_ that fits
|
||||
rather than the biggest _context_ — e.g. adding `-c 65536` (or as large a
|
||||
`--ctx-size` as the hardware holds) to that model's llama-swap config
|
||||
entry. A smaller model at a much larger explicit context often fits in
|
||||
the same VRAM a bigger model's auto-fit context gets shrunk to make room
|
||||
for.
|
||||
|
||||
Clear back to the harness's native cloud default with:
|
||||
|
||||
```bash
|
||||
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
|
||||
-H 'Content-Type: application/json' -d '{"clear": true}'
|
||||
```
|
||||
|
||||
Clearing also removes the env vars the selection injected from the tmux
|
||||
session (they persist there and would otherwise be inherited by the
|
||||
relaunched CLI) and deletes the per-session config directory
|
||||
(`~/.codeman/custom-model-configs/<sessionId>`, written 0600 because pi and
|
||||
omp embed the API key in it). That directory is also removed when the
|
||||
session is deleted. The selection survives a Codeman restart: the endpoint
|
||||
id, model and injected key NAMES are persisted, the values are re-derived
|
||||
from the endpoint store on recovery, and the pane keeps running against the
|
||||
endpoint in between because tmux retains its environment.
|
||||
|
||||
⚠️ Clearing removes injected keys **by name**, and `CLAUDE_CONFIG_DIR` is one
|
||||
of the names claude's selection injects — so a session that ALSO had
|
||||
`CLAUDE_CONFIG_DIR` set through the generic `envOverrides` field (the
|
||||
per-client-account case) loses that override on clear too, and silently
|
||||
falls back to the server's default Claude account. If you route a session
|
||||
to a specific account this way, re-apply the override after clearing a
|
||||
custom-model selection from it.
|
||||
|
||||
**New sessions always default back to the harness's native backend.** A
|
||||
custom-endpoint selection is a per-session choice, never a sticky global
|
||||
default — starting a fresh session doesn't inherit whatever the last one was
|
||||
pointed at.
|
||||
|
||||
## Confidence per harness
|
||||
|
||||
Every harness except Antigravity has now been run end-to-end against a real
|
||||
llama-swap server via `scripts/test-local-llm-harnesses.ts` (a dynamic
|
||||
script that reads the live CLI registry, so a registry change is picked up
|
||||
automatically). Results:
|
||||
|
||||
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
|
||||
came back through the endpoint.
|
||||
- **Codex** — the config is structurally correct, and against a llama-swap
|
||||
server that DOES answer `/v1/responses` (confirmed live: a plain,
|
||||
no-tool-call chat turn returned a real reply), the picture is more
|
||||
nuanced than a flat failure. A real tool-call attempt (`run the shell
|
||||
command: echo hello`) came back as `agent_message` TEXT — literally the
|
||||
tool-call JSON printed as the model's answer — instead of a
|
||||
`function_call` item Codex would actually execute (confirmed via `codex
|
||||
exec --json`'s raw event stream). So plain chat can work while the thing
|
||||
that makes Codex a coding agent — actually running commands and editing
|
||||
files — does not; treat Codex as still unreliable for real work against a
|
||||
llama.cpp/llama-swap endpoint, tool-calling gap included, not just the
|
||||
earlier-documented `wire_api` mismatch (which not every deployment hits
|
||||
the same way — some legitimately have no `/v1/responses` route at all).
|
||||
Separately, EVERY custom-endpoint Codex session prints `Model metadata
|
||||
for '<id>' not found. Defaulting to fallback metadata...` on launch —
|
||||
confirmed harmless (the reply above still came back correctly): Codex's
|
||||
model metadata (reasoning-tier options, per-model system-prompt
|
||||
templates, context-window figures) comes from `models_cache.json`, a
|
||||
local cache of OpenAI's own hosted model catalog that a custom local
|
||||
model can never appear in by construction, since it isn't one of
|
||||
OpenAI's models. There's no config.toml override for a model's metadata,
|
||||
and fabricating a fake catalog entry would mean copying the _shape_ of
|
||||
OpenAI's own proprietary schema (their per-model system-prompt content
|
||||
included) for a warning that doesn't otherwise affect behavior — not
|
||||
something to build into discovery.
|
||||
- **Gemini** — fails with `Invalid auth method selected`, traced to an
|
||||
undocumented `GATEWAY` auth path gemini-cli selects once
|
||||
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
|
||||
(several auth workarounds were tried and ruled out); do not rely on
|
||||
Gemini support yet.
|
||||
- **DeepSeek** — root cause of the `HTTP_404` found and fixed. DeepSeek
|
||||
Harness's own bundled provider module (`@deepseek-ai/dsh-llm-deepseek`)
|
||||
builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with no
|
||||
`/v1` insertion of its own (its real public API, `https://api.deepseek.com`,
|
||||
expects the caller's base URL to already carry any needed prefix) —
|
||||
confirmed by reading its own source and, live, that
|
||||
`POST <baseUrl>/chat/completions` 404s against llama-swap while
|
||||
`POST <baseUrl>/v1/chat/completions` succeeds; the harness's own error
|
||||
template (`DeepSeek API error (HTTP ${status})`) matches the originally
|
||||
reported symptom exactly. `customModelInjection`'s new `appendV1Suffix`
|
||||
(deepseek's entry only — claude/gemini must NOT get it, since claude was
|
||||
already confirmed working against the raw `baseUrl`) fixes it by writing
|
||||
`DEEPSEEK_BASE_URL` with `/v1` appended. Not yet re-run end-to-end with a
|
||||
real `dsh` binary (no install available in this environment) — the fix
|
||||
is source-confirmed and live-verified at the HTTP level, but a real
|
||||
"hello world" reply through `dsh` itself is still outstanding before
|
||||
calling this fully verified like the harnesses above.
|
||||
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
|
||||
|
||||
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
|
||||
each result. `scripts/test-local-llm-harnesses.ts` is the standalone script
|
||||
used to check a harness against a real endpoint outside the web UI
|
||||
entirely; see its own `--help` for usage.
|
||||
|
||||
## Security note
|
||||
|
||||
Every env var this feature can set that redirects a session's traffic
|
||||
(`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, etc.) is
|
||||
listed in that CLI's `privilegedEnvKeys` in the CLI registry, so a
|
||||
non-granted multi-user owner cannot set one directly via the generic
|
||||
`envOverrides` API field — only through this feature's own route, which
|
||||
computes the value from an admin-configured, SSRF-guarded endpoint rather
|
||||
than trusting arbitrary client input. See the "Multi-user security
|
||||
hardening" section of `custom-model-endpoints-plan.md` for the full reasoning; several
|
||||
of these were reachable via the generic `envOverrides` field even before
|
||||
this feature existed, and building this surfaced and closed that gap.
|
||||
@@ -1,178 +0,0 @@
|
||||
# DeepSeek Harness (`dsh`) integration plan
|
||||
|
||||
> **Status**: Executed. This document records the plan, the decision behind each
|
||||
> wiring point, and what was and was not verified. The user-facing guide is
|
||||
> [`deepseek-integration.md`](./deepseek-integration.md); the per-decision
|
||||
> invariants live in
|
||||
> [`architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek`](./architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek).
|
||||
> Template: the grok integration ([`grok-integration-plan.md`](./grok-integration-plan.md)),
|
||||
> itself calibrated against pi. Every fact below was measured against a live
|
||||
> **dsh 0.1.1-rc.2** install and **@deepseek-harness-tui/dsh-tui 0.9.0**, not read
|
||||
> from documentation.
|
||||
|
||||
## 1. What the DeepSeek Harness is
|
||||
|
||||
[deepseek-ai/deepseek-harness](https://github.com/deepseek-ai/deepseek-harness)
|
||||
(open-sourced 2026-08-13, MIT) is a plugin-native agent framework: tools, skills,
|
||||
sessions, sandboxes and whole APPS are Cordis plugins composed into *profiles*.
|
||||
`dsh` is the launcher — `dsh --profile <name>` boots
|
||||
`$DSH_HOME/profiles/<name>`, an ordered stack of plugin-bundle patch layers under
|
||||
the user's own overrides. State lives in `~/.dsh` (`.env` 0600, `settings.yaml`,
|
||||
`cordis.patch.yml`, `profiles/`, `sessions/`, `storages/`).
|
||||
|
||||
## 2. Shape decisions (why DeepSeek is wired the way it is)
|
||||
|
||||
DeepSeek is a ninth run mode. Never a location overlay, never a web tab (the
|
||||
browser UI is handled separately, §3). Three of its decisions have no precedent
|
||||
in the six external CLIs before it.
|
||||
|
||||
| Question | Decision | Why |
|
||||
| --- | --- | --- |
|
||||
| What does a pane run? | `dsh --profile <name>`, profile discovered | **The decision that shapes everything else.** DeepSeek ships `web`, `headless` and `base` — no terminal agent. The interactive front door is always a third-party plugin, so Codeman resolves a binary AND a profile inventory, and "available" means both. `resolveDefaultDeepSeekProfile()` prefers a recognized TUI, then an UNRECOGNIZED profile (anyone can publish an app bundle; a classifier that has not heard of one must not hide it), and refuses `web`/`headless`, which cannot occupy a pane. |
|
||||
| Which TUI? | none blessed; default for BOOTSTRAP only | `POST /api/deepseek/install-profile` defaults to `@deepseek-harness-tui/dsh-tui` (~27.5k weekly downloads, ~4x the next, MIT, and it speaks the status contract in §2.3), but accepts any npm name and the resolver never assumes that profile exists. Codeman offers a default; it does not pick a winner. |
|
||||
| Permission bypass | `DSH_PERMISSION_MODE` env export, no flag | The harness has NO command-line permission option; its sandbox/approval rows read one env var with three presets (`read-only` / `workspace-write` / `danger-full-access`, read off `dsh --dump-default-config`). This is the one legitimate exception to the `CLAUDE_CODE_EFFORT_LEVEL` ban: that var hard-locks in-session switching, whereas the harness reads this with `??` as a boot-time DEFAULT, so it stays soft. Exported via `tmux setenv`, never on the command line. The Run button sends `danger-full-access`, matching every sibling Run button. |
|
||||
| Multi-user clamp branch | only-if-sent, clamped to `workspace-write`, **plus an env-var half** | Omitting the export leaves the harness on `workspace-write`, which still ASKS, so an absent config is already safe (the codex/antigravity/grok shape, not pi's materialize). Clamping to `workspace-write` rather than `read-only` is deliberate: the clamp removes privilege, it must not break a session's ability to edit its own workspace. ⚠️ Unlike every sibling, clamping the CONFIG is only half the gate: the switch is an env var, `DSH_*` is an allowlisted `envOverrides` prefix, and `applyEnvOverrides()` runs AFTER `_configureDeepSeek()`, so `envOverrides: {DSH_PERMISSION_MODE: 'danger-full-access'}` on the same request would land last and win. `clampEnvOverridesForOwner()` drops `DSH_PERMISSION_MODE` and `DSH_HOME` for a non-granted owner (dropping falls through to the clamped export). `DSH_HOME` because it aims the launcher at a profile tree whose plugin code runs at BOOT, before any approval row. |
|
||||
| `hooksAvailableForMode()` granularity | per SESSION for deepseek, per mode for everything else | `deepSeekConfig.statusReporting: false` disarms the `HERDR_*` export, and the triple is the only reason a dsh session posts anything, so a mode-only answer would accept `until=stop` where nothing can send one — the infinite-wait the predicate exists to prevent. Call sites pass `sessionHookOptions(session)`; the default stays permissive so a forgotten one degrades to the old behaviour. ⚠️ Profile conformance stays unknowable at request time (an unrecognized profile is deliberately launchable), so a non-conforming TUI still times out on an explicit `stop`; the default set keeps `idle`/`exit` for that. ⚠️ The predicate is NOT "is this claude": Read My Mind and intent capture read Claude's transcript and were silently widened by this change, so they compare `mode === 'claude'` directly now. |
|
||||
| Profile install spawn | own process group, hand-rolled timeout | `dsh plugin add` fans out into package-manager children, and spawn's built-in `timeout` signals only the direct child: survivors keep the inherited stdio pipes open, `close` never fires, and the held-open request leaks with no route-level deadline. `detached: true` + negative-pid SIGTERM→SIGKILL, the same escalation `runGit()` uses for the same reason, plus a last-resort reap for a grandchild that escaped the group. |
|
||||
| Idle detection | **real hook events via a status shim** | The standout decision. The TUI already reports its lifecycle to a supervising process through a generic env-gated contract inherited from Herdr: `HERDR_ENV=1` + `HERDR_BIN_PATH` + `HERDR_PANE_ID` make it run `<bin> pane report-agent <id> --state idle\|working\|blocked …` on every state change, exit 0 = delivered. `deepseek-status-shim.ts` generates a script into the data dir and points `HERDR_BIN_PATH` at it. So deepseek is the only non-claude mode that passes `hooksAvailableForMode()` — earned by emitting definitive signals, not granted. An interface implementation, not an impersonation: no real `herdr` binary is ever executed, and a TUI that ignores the contract simply falls back to output stabilization. |
|
||||
| `agent_working` event | new, 157th SSE constant | The one hook event with no Claude Code hook behind it. A harness turn cannot run while its own modal approval is on screen, so "started working" proves a dialog was answered in the terminal. Without it a dsh red alert would survive until the next `stop` — the exact stuck-alert bug the claude path already fixed once, and its pane-capture staleness sweep is Claude-dialog-shaped and cannot help here. |
|
||||
| Resolver | identity probe THEN version probe | Strictest of the family, and not by preference. `dsh` is not merely a squattable npm name: Debian ships an unrelated `dsh` (dancer's shell, `apt install dsh`) which would answer a version probe convincingly and then be handed a spawn line. `dsh --help` must match `DeepSeek Harness` first. `DEEPSEEK_VERSION_REGEX` keeps the prerelease tail (`0.1.1-rc.2`), since truncating it would report an rc as a release. |
|
||||
| Env allowlist | `DSH_*` + `DEEPSEEK_*` | `DSH_*` covers the launcher's documented inputs (`DSH_HOME`, `DSH_PERMISSION_MODE`, `DSH_TELEMETRY_MODE`, the `DSH_TUI_*` knobs); `DEEPSEEK_*` is the vendor namespace holding `DEEPSEEK_API_KEY`/`DEEPSEEK_BASE_URL`, same reasoning that admitted `XAI_*` for grok. ⚠️ Pi's lesson repeats exactly: a dsh `settings.yaml` can nominate ANY env var as a provider credential (`apiKeyEnv`), and the allowlist is one GLOBAL list, so admitting those would widen every mode at once. They stay out. |
|
||||
| Model | NOT a session field | The model is a composition entry (`agent-default-model`) in the profile's config tree, set in `~/.dsh/settings.yaml` + `cordis.patch.yml`. Both create paths deliberately resolve no model for this mode rather than inventing a flag. |
|
||||
| Alt-screen strip | OUT of `isAltScreenStripMode()` | Third-party fullscreen TUIs with their own scrollback and mouse handling — the opencode case, not the Ink case. |
|
||||
| Local echo | `'buffer'` via the `_updateLocalEchoState` fallthrough | UNMEASURED against a live authenticated session (see §5), same honest gap grok shipped with. The leading TUI's composer supports `@` completion and history search, which *may* make it per-keystroke reactive like codex; if so the fallback is the `'off'` branch. |
|
||||
| Docker | image installs dsh AND a profile | Profiles are deliberately NOT seeded from the host: each is a per-profile `node_modules` tree, host-arch-specific and far too large to copy per container start. Only `~/.dsh/.env`, `settings.yaml`, `cordis.patch.yml` are seeded (auth + model composition). The profile install rides the `useradd` layer so the closing `chgrp`/`chmod g=u` covers it, which is what keeps it usable under the arbitrary uid the container runs as. |
|
||||
| Remote SSH | `exec "$SHELL" -i -l -c 'dsh'` | Boots the remote box's default profile; a remote with several needs the per-host `commands.deepseek` override, since `deepSeekConfig` does not cross ssh. |
|
||||
|
||||
## 3. The web profile
|
||||
|
||||
The browser UI is the only interactive surface DeepSeek ships itself, so it gets
|
||||
a **shortcut, not a run mode**: `Run ▸ DeepSeek web UI…` starts
|
||||
`dsh web --no-open --host 127.0.0.1 --port <free> --trusted-host <codeman-authority>`
|
||||
as a background process and opens the URL as an ordinary web tab.
|
||||
|
||||
The server was a **shell session** first, on the reasoning that Codeman already
|
||||
supervises those (visible, scrollable, killable, dies with its tab) so nothing
|
||||
new had to own a long-lived HTTP server. That version worked and was still
|
||||
wrong in use: clicking "open the DeepSeek web UI" put a terminal tab on screen
|
||||
next to the web tab actually asked for, every single time, and after the first
|
||||
launch the terminal was pure noise. Opening a dashboard should open one tab.
|
||||
|
||||
So `POST /api/deepseek/web` owns it instead (`src/deepseek-web-server.ts`), and
|
||||
what the session gave away for free is now explicit: exactly one server, reused
|
||||
rather than raced on a second click; restarted when the requested authority
|
||||
changes; killed on Codeman shutdown (a detached child would otherwise hold its
|
||||
port against the next start — the very EADDRINUSE this feature already got
|
||||
wrong once); and boot output captured, since with no shell tab there is nowhere
|
||||
else for a stack trace to land. It is fenced at the same bar as the profile
|
||||
installer: booting a dsh profile executes the plugin code in it, so it requires
|
||||
the privileged grant in multi-user mode.
|
||||
|
||||
`--trusted-host` is load-bearing — dsh fences its `/api` behind a browser-trust
|
||||
check on the request authority, and a Codeman web tab reaches it through
|
||||
Codeman's own origin via the webview proxy, not directly. The authority comes
|
||||
from the CLIENT (`location.host`) because only the browser knows which of a
|
||||
multi-homed Codeman's origins is actually in play.
|
||||
|
||||
Three things about this shortcut are load-bearing and each came from it failing
|
||||
in exactly that way against a real install:
|
||||
|
||||
- **The port is chosen, never hardcoded.** `GET /api/deepseek/web-port` walks
|
||||
3080..3119 for a free loopback port. 3080 is dsh's own default, which makes it
|
||||
precisely the port a DeepSeek user is most likely to already be serving on:
|
||||
binding it unconditionally killed the launch with `EADDRINUSE` against the
|
||||
user's own `dsh web`.
|
||||
- **The tab is opened only after the server answers.** The launch polls
|
||||
`POST /api/webviews/probe` until the URL responds, so a server that dies on
|
||||
startup reports the failure and points at its shell tab, instead of silently
|
||||
persisting a dashboard aimed at nothing.
|
||||
- **The saved tab is `trusted: true`, and must be.** An untrusted webview is
|
||||
sandboxed without `allow-same-origin`, which breaks this dashboard twice: the
|
||||
dsh client-runtime reads `localStorage` while loading plugins and dies there,
|
||||
and an opaque-origin frame sends `Origin: null`, so dsh's trust check 403s
|
||||
every `/api` call regardless of what `--trusted-host` names. Passing
|
||||
`location.host` only means anything once the frame actually carries that
|
||||
origin. The trade is real — a trusted proxied frame is same-origin with
|
||||
Codeman and can reach Codeman's API — and is defensible only because this
|
||||
particular dashboard is an agent harness Codeman just started itself on
|
||||
loopback, which can already run code as the user. It is not a precedent for
|
||||
trusting third-party dashboards generally.
|
||||
|
||||
The record is marked `managed: 'deepseek-web'`, which keeps it out of the
|
||||
saved-dashboard list: the shortcut that maintains it is already a menu entry, so
|
||||
listing both showed the same dashboard twice. Being managed is also what lets a
|
||||
relaunch repoint the existing row instead of stacking one dead dashboard per
|
||||
restart, since the port is now chosen per launch.
|
||||
|
||||
The authority baked into `--trusted-host` is the one the launch was clicked
|
||||
from, and reuse is conditional on it: a running server fenced for a *different*
|
||||
origin is stopped and restarted rather than reused, because reusing it renders a
|
||||
page whose every API call 403s — which reads as a broken dashboard rather than a
|
||||
misconfigured one.
|
||||
|
||||
## 4. Touch points (the checklist)
|
||||
|
||||
Backend: `types/session.ts` (SessionMode + `DeepSeekConfig` + SessionState),
|
||||
`utils/deepseek-cli-resolver.ts` (new) + barrel, `deepseek-status-shim.ts` (new),
|
||||
`tmux-manager.ts` (`buildDeepSeekCommand`, dispatch, resume flag, PATH export,
|
||||
truecolor, `_configureDeepSeek`, availability error, plumbing), `session.ts`
|
||||
(external-mode gate, label, config plumbing, tmux-required error, attach env),
|
||||
`mux-interface.ts`, `schemas.ts` (prefixes, `DeepSeekConfigSchema`,
|
||||
`DeepSeekInstallProfileSchema`, both mode enums, remote overrides, cron agentType,
|
||||
`agent_working`), `session-wait-registry.ts` (`hooksAvailableForMode`),
|
||||
`hook-event-routes.ts` (`APPROVAL_RESOLVING_EVENTS`), `session-routes.ts` (clamp +
|
||||
both create paths + `resolveDeepSeekLaunchError`), `system-routes.ts`
|
||||
(`GET /api/deepseek/status`, `POST /api/deepseek/install-profile`), `server.ts`
|
||||
(availability inject + mux restore), `sse-events.ts`, `docker-hosts.ts`,
|
||||
`remote-hosts.ts`, `config/dependency-registry.ts`,
|
||||
`response-viewer-transcript.ts`, `cron/cron-service.ts` (comment),
|
||||
`tui/tui-client.ts` + `tui-app.ts`.
|
||||
|
||||
Frontend: `index.html` (welcome button, run-mode entry, install affordance, web-UI
|
||||
shortcut, cron option, clone Brain option), `session-ui.js` (`runDeepSeek()`,
|
||||
`runDeepSeekWeb()`, `installDeepSeekProfile()`, dispatch, availability, "Run DS"
|
||||
label, external-CLI gates), `app.js` (label, `ds` tab badge, kill-menu, SSE map),
|
||||
`settings-ui.js` (welcome gate + `_onHookAgentWorking`), `constants.js`,
|
||||
`mobile-overview.js`, `home-sessions.js`, `panels-ui.js`, `i18n.js`,
|
||||
`terminal-ui.js`, `styles.css` + `mobile.css` (brand-indigo identity; the non-og
|
||||
skin block and the mobile `!important` pair are both load-bearing).
|
||||
|
||||
Meta: `docker/agent.Dockerfile`, `install.sh`, `package.json` keyword,
|
||||
`skills/codeman/reference/*`, CLAUDE.md, `architecture-invariants.md`.
|
||||
|
||||
Tests: `test/deepseek-mode.test.ts` + `test/deepseek-cli-resolver.test.ts` (new);
|
||||
`run-mode-ui`, `render-index-html`, `mobile-overview`, `agent-skill-mode-lists`
|
||||
(extended).
|
||||
|
||||
## 5. Verification performed
|
||||
|
||||
See the summary at the end of the implementing session for the live run. In
|
||||
short: the CI gate green; the resolver, profile inventory, spawn-line and clamp
|
||||
behaviour covered by 31 new unit tests; and an isolated instance used to exercise
|
||||
`GET /api/deepseek/status` and a real session against the live dsh install.
|
||||
|
||||
**Not verified (honest gaps):**
|
||||
|
||||
- The local-echo `'buffer'` policy against the TUI's real composer (§2). If it
|
||||
turns out per-keystroke reactive like codex's, flip it to the `'off'` branch;
|
||||
teaching `PredictiveEchoAddon` its composer row is the larger follow-up.
|
||||
- Scrollback/repaint behaviour of a third-party fullscreen TUI under the narrow
|
||||
strip during a long session.
|
||||
- A Docker case with `mode: 'deepseek'` (needs a `--no-cache` agent-image
|
||||
rebuild — see the `--no-cache` rule in CLAUDE.md).
|
||||
- A remote-SSH deepseek case.
|
||||
- The web-UI shortcut against a tunnel authority. Loopback and a tailnet name are
|
||||
both verified end to end through the webview proxy (dashboard renders, its
|
||||
`/api` calls succeed, no shell session created).
|
||||
|
||||
## 6. Follow-ups
|
||||
|
||||
- **Response viewer**: read `~/.dsh/sessions/**` (JSONL) the way codex rollouts
|
||||
are read back. Highest-value follow-up, and very achievable.
|
||||
- **`headless` as an execution backend** for Codeman's own AI checks
|
||||
(`ai-idle-checker`, `ai-plan-checker`), today Claude-only.
|
||||
- **Profile/model picker in Session Options**, reading `GET /api/deepseek/status`
|
||||
`.profiles`.
|
||||
- **`--patch` overlays per session**, which is the harness-native way to change
|
||||
agent composition without touching the user's profile.
|
||||
- Measure the local-echo policy and pin the result the way pi did.
|
||||
@@ -1,307 +0,0 @@
|
||||
# DeepSeek Harness (`dsh`) in Codeman
|
||||
|
||||
Codeman can run [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)
|
||||
as a session backend, alongside Claude Code, OpenCode, Codex, Gemini,
|
||||
Antigravity, Pi and Grok. It is the ninth run mode, and the one that is wired
|
||||
least like the others, for two reasons worth understanding before you use it.
|
||||
|
||||
## 1. The agent is a profile, not the binary
|
||||
|
||||
`dsh` is a **launcher**, not an agent. It boots a *profile*: an ordered stack of
|
||||
plugin-bundle patch layers under `$DSH_HOME/profiles/<name>` (`$DSH_HOME`
|
||||
defaults to `~/.dsh`). DeepSeek ships three bundles and none of them is a
|
||||
terminal agent:
|
||||
|
||||
| Profile | What it is | Can Codeman run it in a tab? |
|
||||
| ------------ | --------------------------------- | ---------------------------- |
|
||||
| `web` | the browser UI, served on :3080 | no — but see §6 |
|
||||
| `headless` | answers one task and exits | no |
|
||||
| (`base`) | the shared core, no app at all | no |
|
||||
|
||||
The interactive terminal front door is **always a third-party plugin**. So
|
||||
"DeepSeek is installed" and "Codeman can start a DeepSeek session" are different
|
||||
questions, and Codeman answers both separately:
|
||||
|
||||
```bash
|
||||
curl -s localhost:3000/api/deepseek/status | jq
|
||||
{
|
||||
"available": true, # the `dsh` binary resolved and proved its identity
|
||||
"runnable": false, # ...but nothing installed can drive a pane
|
||||
"path": "/home/you/.local/bin",
|
||||
"version": "0.1.1-rc.2",
|
||||
"dshHome": "/home/you/.dsh",
|
||||
"defaultProfile": null,
|
||||
"profiles": [ { "name": "web", "kind": "web", "bundles": [...] } ]
|
||||
}
|
||||
```
|
||||
|
||||
### Installing a terminal profile
|
||||
|
||||
From the UI: open the **Run** dropdown. When `dsh` is installed but no
|
||||
pane-capable profile is, the menu shows **DeepSeek — add a terminal profile…**.
|
||||
One click installs one and the normal DeepSeek entry appears.
|
||||
|
||||
By hand, or to pick a different front door:
|
||||
|
||||
```bash
|
||||
dsh plugin --profile dsh-tui add @deepseek-harness-tui/dsh-tui
|
||||
```
|
||||
|
||||
⚠️ **`pnpm` has to be on PATH for either route.** `dsh plugin` is a thin forwarder
|
||||
that spawns a literal `pnpm` with no npm fallback, so without one it exits 127 with
|
||||
`dsh: pnpm not found on PATH` — both by hand and behind the UI button, which
|
||||
surfaces that same line as the install error. `npm install -g pnpm` (or
|
||||
`corepack enable pnpm`) is the fix. This is what broke the Docker agent image in
|
||||
[#352](https://github.com/Ark0N/Codeman/issues/352); the image now installs pnpm
|
||||
alongside `dsh`. The Compose server image (`docker/server.Dockerfile`) does not
|
||||
ship `dsh`, since it is installed at runtime, but it does ship pnpm so the UI
|
||||
button works there too.
|
||||
|
||||
Codeman's default is `@deepseek-harness-tui/dsh-tui` because it is by a wide
|
||||
margin the most used community TUI, it is MIT, and it implements the status
|
||||
contract described in §3. It is a **default, not a requirement**: any profile
|
||||
under `$DSH_HOME/profiles` that is not `web` or `headless` shows up in the
|
||||
inventory and can be launched, including one you compose yourself. The endpoint
|
||||
accepts any npm package name:
|
||||
|
||||
```bash
|
||||
curl -sX POST localhost:3000/api/deepseek/install-profile \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"profile":"my-tui","package":"@someone/dsh-tui"}'
|
||||
```
|
||||
|
||||
Installing a plugin is arbitrary code execution on the host, so in multi-user
|
||||
mode this endpoint requires the can-bypass-permissions grant (the same bar as a
|
||||
`shell` session). The request is held open while the package manager runs and is
|
||||
bounded at five minutes; the install runs in its own process group, so hitting
|
||||
that bound kills the whole tree rather than just the launcher.
|
||||
|
||||
> **`dsh` is also a Debian program.** `apt install dsh` gives you "dancer's
|
||||
> shell", a distributed shell, which would answer `--version` convincingly.
|
||||
> Codeman's resolver therefore demands the harness's own help banner before it
|
||||
> will point a spawn line at a candidate, and `GET /api/deepseek/status` reports
|
||||
> `path` and `version` so a misresolution is diagnosable rather than presenting
|
||||
> as "the mode just doesn't work".
|
||||
|
||||
## 2. Permissions are an env var, not a flag
|
||||
|
||||
The harness has **no `--dangerously-skip-permissions` equivalent**. Its sandbox
|
||||
and approval rows are configuration, driven by one documented input,
|
||||
`DSH_PERMISSION_MODE`, with three presets (read off `dsh --dump-default-config`):
|
||||
|
||||
| `DSH_PERMISSION_MODE` | sandbox | approvals | notes |
|
||||
| --------------------- | -------------------- | --------- | ------------------------- |
|
||||
| `read-only` | `read-only` | ask | |
|
||||
| `workspace-write` | `workspace-write` | ask | the harness's own default |
|
||||
| `danger-full-access` | `danger-full-access` | **never** | what the Run button sends |
|
||||
|
||||
Codeman exports it via `tmux setenv`, never on the command line. Because the
|
||||
harness reads it with `??`, it is a **soft default**: it sets the boot-time
|
||||
preset and you can still change permission mode inside the session.
|
||||
|
||||
Omitting it entirely leaves the harness on `workspace-write`, which still asks —
|
||||
which is why the multi-user clamp only needs to force a *sent* value down. A
|
||||
non-granted owner's `danger-full-access` becomes `workspace-write`, not
|
||||
`read-only`: the clamp removes privilege without breaking the session's ability
|
||||
to edit its own workspace.
|
||||
|
||||
Because the switch is an env var rather than a flag, that clamp has a second half
|
||||
no other CLI needs. `DSH_*` is an allowlisted `envOverrides` prefix (it has to be:
|
||||
that is also how you set the harness's ordinary knobs), and env overrides are
|
||||
applied *after* the permission export, so in multi-user mode a non-granted owner
|
||||
sending
|
||||
|
||||
```json
|
||||
{ "mode": "deepseek", "envOverrides": { "DSH_PERMISSION_MODE": "danger-full-access" } }
|
||||
```
|
||||
|
||||
would otherwise hand back the privilege the config clamp just removed. For a
|
||||
non-granted owner Codeman therefore **drops `DSH_PERMISSION_MODE` and `DSH_HOME`
|
||||
from `envOverrides`**; dropping them falls through to the clamped config and the
|
||||
server's own `DSH_HOME`. `DSH_HOME` is in that list because it points the
|
||||
launcher at a profile tree, and a profile's plugin code runs at boot, before any
|
||||
approval row can apply. Single-user installs and granted owners are unaffected.
|
||||
|
||||
## 3. Real idle detection (the interesting part)
|
||||
|
||||
Every other external CLI mode in Codeman is **readiness-guessed**: Codeman
|
||||
watches the PTY go quiet and infers that a turn ended. Claude is the exception,
|
||||
because Claude Code fires hooks.
|
||||
|
||||
DeepSeek is the second exception. The community terminal front door already
|
||||
reports its own lifecycle to a supervising process through a generic,
|
||||
env-var-gated contract (inherited from [Herdr](https://herdr.dev)): when
|
||||
`HERDR_ENV=1`, `HERDR_BIN_PATH` and `HERDR_PANE_ID` are set, it shells out on
|
||||
every state change with
|
||||
|
||||
```
|
||||
"$HERDR_BIN_PATH" pane report-agent "$HERDR_PANE_ID" \
|
||||
--source custom:dsh-tui --agent dsh-tui \
|
||||
--state idle|working|blocked [--message ...] --seq N
|
||||
```
|
||||
|
||||
Codeman points `HERDR_BIN_PATH` at a small generated shim
|
||||
(`~/.codeman/dsh-status-shim.mjs`, written at session create) which forwards each
|
||||
report to `POST /api/hook-event`. The mapping:
|
||||
|
||||
| Harness state | Codeman hook event | What you get |
|
||||
| ------------- | ------------------ | -------------------------------------------------------- |
|
||||
| `blocked` | `permission_prompt`| red "needs you" tab alert + an Approvals Inbox item |
|
||||
| `idle` | `stop` | definitive end-of-turn: respawn triggers, `wait` returns |
|
||||
| `working` | `agent_working` | clears an alert answered in the terminal, at once |
|
||||
|
||||
So a DeepSeek session gets Claude-grade signals: `GET /api/sessions/:id/wait`
|
||||
really can block on `stop` and `blocked` for it, and it is the only non-Claude
|
||||
mode for which that is true (`hooksAvailableForMode`).
|
||||
|
||||
That is a per-*session* answer, not a per-mode one. Turning the bridge off with
|
||||
`deepSeekConfig.statusReporting: false` means nothing will ever post a hook event
|
||||
for that session, so an explicit `until=stop` is refused up front (with a message
|
||||
naming the setting) rather than blocking for your whole timeout. Omitting `until`
|
||||
never fails: the hook-only signals are dropped from the default set and you still
|
||||
get `idle` and `exit`.
|
||||
|
||||
One limit worth knowing: whether the *profile* implements the contract cannot be
|
||||
known at request time (Codeman deliberately treats an unrecognized profile as
|
||||
launchable). A dsh session running a non-conforming TUI therefore still accepts
|
||||
`until=stop` and will time out on it. `idle`/`exit` are the reliable pair there.
|
||||
|
||||
This is an interface implementation, not an impersonation — nothing on your
|
||||
machine executes a real `herdr` binary. If you use a terminal profile that does
|
||||
*not* implement the contract, the shim is simply never called and the mode falls
|
||||
back to output-stabilization readiness like its siblings. Turn it off per session
|
||||
with `deepSeekConfig.statusReporting: false`.
|
||||
|
||||
## 4. Starting a session
|
||||
|
||||
From the UI, pick **DeepSeek** in the Run dropdown (or the **Run DeepSeek**
|
||||
welcome button) and press Run. Over the API:
|
||||
|
||||
```bash
|
||||
curl -sX POST localhost:3000/api/quick-start \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"caseName": "myproject",
|
||||
"mode": "deepseek",
|
||||
"deepSeekConfig": {
|
||||
"profile": "dsh-tui",
|
||||
"permissionMode": "danger-full-access"
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
`deepSeekConfig` fields: `profile`, `permissionMode`, `resumeSession`,
|
||||
`resumeSessionId`, `statusReporting`. Resume prefers an explicit id over the
|
||||
most-recent form, and both are passed through to the profile's app, which is
|
||||
where `--resume` is understood.
|
||||
|
||||
**Models are not a session field.** The model is a composition entry in the
|
||||
profile's config tree (`agent-default-model`), not a CLI flag, so Codeman does
|
||||
not try to set one. Configure it where the harness does: `~/.dsh/settings.yaml`
|
||||
plus a home-level `~/.dsh/cordis.patch.yml`, or a `--patch` overlay on the
|
||||
profile. That is also how you point dsh at a local or third-party provider.
|
||||
|
||||
**Environment.** `DSH_*` and `DEEPSEEK_*` are allowlisted for `envOverrides`
|
||||
(so `DSH_HOME`, `DSH_PERMISSION_MODE`, `DEEPSEEK_API_KEY`, `DEEPSEEK_BASE_URL`
|
||||
all flow through). Provider keys with *other* names are deliberately not: a dsh
|
||||
`settings.yaml` can nominate any env var as a credential via `apiKeyEnv`, and
|
||||
Codeman's allowlist is global, so admitting them would widen it for every mode at
|
||||
once. Authenticate those the way dsh does, from the file or the server's own
|
||||
environment.
|
||||
|
||||
## 5. Reading a session back, and driving one as a worker
|
||||
|
||||
dsh writes a real transcript — `$DSH_HOME/sessions/<mangled-cwd>/<id>/session.jsonl.zstd`
|
||||
— so `GET /api/sessions/:id/last-response` reads that rather than segmenting the
|
||||
pane, and the Response Viewer shows a dsh conversation the way it shows a claude
|
||||
or codex one (`?context=full` returns prompt / response / tool blocks).
|
||||
|
||||
Reading the pane instead is not merely coarse for this mode, it is wrong: dsh-TUI
|
||||
paints a full-screen splash, so the segmenter answered a `last-response` call for
|
||||
a fresh dsh session with its ASCII-art logo — which anything polling for a
|
||||
worker's first answer reads as an answer. Three things about the file shaped the
|
||||
reader (`src/deepseek-transcript.ts`):
|
||||
|
||||
- **It is one zstd FRAME per append, not one zstd stream.** `zstd -dc` decodes all
|
||||
of them, Node's `zlib` zstd decoder stops at the first: a real 56-line
|
||||
transcript came back as 1 line. The reader walks frame headers itself. On a Node
|
||||
older than 22.15 (no zstd at all) the mode falls back to the pane, as before.
|
||||
- **Not every `user/message` is the user.** Each turn also records a
|
||||
plugin-sourced runtime-context snapshot; only `source.kind === 'user'` is a
|
||||
prompt.
|
||||
- **A failed turn is not an empty one.** `turn/end` carries the provider's error,
|
||||
which is returned as `Turn error: …` (and an early stop such as `max-tokens` as
|
||||
`Turn ended: …`) instead of an empty string that reads as "still thinking".
|
||||
|
||||
The transcript reader applies to **local** dsh sessions only. A Docker case's
|
||||
harness writes its transcript inside the container's own `~/.dsh` (the workspace
|
||||
bind mount does not cover it), and a remote-SSH case's lives on the remote host,
|
||||
so the local reader could never find those files — such sessions keep the pane
|
||||
segmenter, coarse but real. The splash caveat above applies to them accordingly.
|
||||
|
||||
### As an agent worker
|
||||
|
||||
Because dsh has both halves — a real end-of-turn signal and a real transcript — an
|
||||
agent can drive a dsh session the same way it drives a claude one, and the bundled
|
||||
`codeman` agent skill does. Spawning `beta:deepseek` in its worker list gives a
|
||||
worker that is tasked, waited on and read with the same calls as its claude
|
||||
siblings; no other external CLI mode qualifies. Two edges are worth repeating here:
|
||||
|
||||
- **Readiness is not the stop signal.** The harness reports `idle` at boot roughly
|
||||
300 ms *before* the composer paints (measured 2.26 s vs 2.56 s after spawn), so a
|
||||
send-and-wait fired immediately after create resolves on that boot report,
|
||||
reports a turn that never ran, and leaves the prompt in a pane that was not yet
|
||||
accepting input. Wait for the composer (`❯`) instead.
|
||||
- **Wait on `stop`, not on the default signal set.** That set also carries `idle`,
|
||||
which for every external CLI is inferred from output stabilization; a dsh TUI
|
||||
that repaints rarely reads as idle mid-turn.
|
||||
|
||||
## 6. The web UI as a tab
|
||||
|
||||
The browser UI is the one interactive surface DeepSeek ships itself, so it gets a
|
||||
shortcut rather than a run mode: **Run ▸ DeepSeek web UI…** starts
|
||||
`dsh web --no-open --host 127.0.0.1 --port <free> --trusted-host <codeman-host>`
|
||||
as a background child process (`src/deepseek-web-server.ts`, behind
|
||||
`POST/GET/DELETE /api/deepseek/web`) and opens it as a Codeman web tab once the
|
||||
server actually answers.
|
||||
|
||||
It is a child process rather than a shell session because the session version
|
||||
opened a terminal tab nobody asked for on every click. What the session gave for
|
||||
free is therefore explicit here: one instance with reuse, a restart when the
|
||||
requested `--trusted-host` authority differs from the running one, a kill on
|
||||
server stop, and captured boot output. The `--trusted-host` flag is load-bearing —
|
||||
dsh fences its `/api` behind a browser-trust check on the request authority, and a
|
||||
Codeman web tab reaches it through Codeman's own origin via the webview proxy, not
|
||||
directly. Without it the page renders and every API call fails.
|
||||
|
||||
## 7. Docker and remote cases
|
||||
|
||||
Docker cases work: the agent image installs `dsh` and bootstraps a `dsh-tui`
|
||||
profile into the container. Profiles are deliberately **not** seeded from the
|
||||
host (each is a per-profile `node_modules` tree, host-arch-specific and far too
|
||||
large to copy on every container start); only `~/.dsh/.env`, `settings.yaml` and
|
||||
`cordis.patch.yml` are seeded, which is what carries auth and model composition
|
||||
in. As with pi and grok, in-container sessions are invisible host-side:
|
||||
`~/.dsh/sessions` inside a container is that container's own.
|
||||
|
||||
Remote SSH cases default to `dsh` through a login shell, which boots the remote
|
||||
box's default profile. If the remote has several, name one with the per-host
|
||||
`commands.deepseek` override — the local `deepSeekConfig` does not cross ssh.
|
||||
|
||||
## 8. What is not wired
|
||||
|
||||
Deliberately minimal, on the same reasoning as the grok integration: the harness
|
||||
is a fast-moving developer preview and every flag added is a flag validated
|
||||
forever.
|
||||
|
||||
- `--patch` overlays per session (the profile's own layers apply as normal).
|
||||
- `dsh plugin` management beyond first-time profile install.
|
||||
- The `headless` profile as a one-shot execution backend for Codeman's own
|
||||
internal AI checks (today those are Claude-only).
|
||||
- Model/provider selection from Session Options.
|
||||
|
||||
## Verified against
|
||||
|
||||
`dsh 0.1.1-rc.2` and `@deepseek-harness-tui/dsh-tui 0.9.0`. The permission
|
||||
presets, the profile layout, and the supervisor contract above were all read off
|
||||
the live install rather than from documentation.
|
||||
@@ -1,433 +0,0 @@
|
||||
<!-- Design doc generated via ultracode multi-agent workflow (wf_e3a7498b-26f): 3 architecture proposals -> judge panel -> synthesis -> completeness critic. -->
|
||||
|
||||
# Docker Session Mode, Implementation Plan
|
||||
|
||||
## Decisions (locked 2026-07-19, by repo owner)
|
||||
|
||||
1. **Isolation posture**: CONVENIENT default (bind-mount host `~/.claude` etc. read-write so the existing login just works; network on; still hardened non-root + cap-drop + resource caps). SEALED profile (`mountCredentials:false` + `network:none`) is a per-case opt-in.
|
||||
2. **Export**: offer BOTH full-image (`commit`+`save`+workspace tar) AND workspace-only, side by side, no default (ask each time).
|
||||
3. **Base image**: BUILD LOCALLY on first use via `scripts/build-agent-image.mjs` from a repo `docker/agent.Dockerfile`. No registry required. (GHCR pull can be added later.)
|
||||
4. **Hooks**: WIRE HOOKS NOW. Codeman scaffolds `.claude/settings.local.json` + CLAUDE.md into the linked host workspace dir (same as local cases), enabling in-container permission prompts, hook-idle detection, and the Claude Model picker.
|
||||
|
||||
Adopted defaults for the remaining open items (Section 10): resume-on-restart ON; container is per-CASE and shared by multiple sessions (killing one session only kills its in-container tmux session, never `docker stop` while siblings remain; stop/remove only on explicit teardown or case-delete); rootless caps = ship-with-warning (`capsEnforced` surfaced); remote docker daemon = local-first; podman = docker-first best-effort.
|
||||
|
||||
## Implementation status (branch `feat/docker-session-mode`)
|
||||
|
||||
DONE and END-TO-END VERIFIED against a real docker daemon (create host, link case, quick-start shell in a real container, workspace bind-mount round-trip, hook scaffolding, session-delete keeps the shared container up, case-delete `docker rm`s it):
|
||||
|
||||
- Phase 0-1: types (`DockerHost`/`DockerCase`/`SessionDocker`), `src/docker-hosts.ts` (storage, pure `buildDockerBaseArgs`/`buildDockerCreateArgs`, `containerApiUrl`, `hostGatewayAlias`, config-hash, credential-mount resolution, daemon probes), `DockerHostSchema`/`DockerCaseLinkSchema`. 26 unit tests.
|
||||
- Phase 2: `tmux-manager` `buildDockerLaunchCommand` (image-check -> ensure -> start -> exec, resume-aware), `buildDockerKillCommand` (in-container tmux only, multi-session safe), stop/remove; wired into `createSession`/`respawnPane`/`killSession`. 14 unit tests.
|
||||
- Phase 3: `Session` threading (`_docker`, toState, option builders, in-container cliVersion probe, `resolveMuxAttachCwd`), `server.ts` recovery round-trip.
|
||||
- Phase 4: `case-routes` `/api/docker-hosts` CRUD + `/api/cases/docker-link` + listing + docker-unlink; `session-routes` `/api/quick-start` docker branch (rejects per-session config, probes availability + tmux, scaffolds hooks, seeds resume id).
|
||||
- Phase 5 (partial): `docker/agent.Dockerfile` + `scripts/build-agent-image.mjs` (built + verified: node 22, tmux, claude/codex/gemini/opencode, arbitrary-uid HOME). Host-guard allowlists `host.docker.internal`/`host.containers.internal` for in-container hooks.
|
||||
- Full CI green (3445 tests).
|
||||
|
||||
REMAINING:
|
||||
|
||||
- Phase 6: export / import (`docker commit` + `save | gzip` + workspace tar + manifest; `load` + quarantined re-tag), GC / boot reaper, disk-safety prechecks, drift-recreate route, SSE `docker:*` events. THE "move to a new machine" feature.
|
||||
- Phase 7: frontend Create Case "Docker" tab + `linkDockerCase` + run wiring + case-picker labels + export/import UI.
|
||||
- Phase 8: CLAUDE.md "Docker cases" Key Pattern + `docs/docker-cases.md` + COM.
|
||||
- Deferred refinements: in-container model-picker via `settings.local.json`; live mid-run resume-id capture into `DockerCase.lastClaudeSessionId`; rootless/Desktop uid probe (currently a platform heuristic).
|
||||
|
||||
## 1. Goal & user stories
|
||||
|
||||
Add "Docker cases" to Codeman: a case can point at a container instead of a local or remote-SSH path, and any of the five CLI backends (`claude` / `shell` / `opencode` / `codex` / `gemini`) runs inside that container. It is modeled as a LOCATION OVERLAY on cases, exactly like the remote-SSH feature (COD-94/#145), never as a sixth `SessionMode`.
|
||||
|
||||
User stories:
|
||||
|
||||
- As the repo owner, I link a case to a per-project container so an autonomous Claude/Ralph run executes in a hardened sandbox (cap-drop, non-root, resource caps) instead of directly on my host, while keeping my existing OAuth login and transcript history working with zero extra setup.
|
||||
- I set default, per-case-changeable container settings (image, network mode, memory/cpu/pids caps) at link time and edit them later, and edits actually take effect through a recreate-on-drift path (see Section 4).
|
||||
- I reconnect after a Codeman restart and land back in the SAME running agent with the conversation intact. When the CONTAINER itself was stopped/rebooted/OOM-killed (which destroys the in-container tmux), the next launch RESUMES the last conversation from the bind-mounted transcript rather than starting fresh (durability model in Section 2, Key decision 1).
|
||||
- I export a finished run's whole environment (toolchain plus workspace) to a portable, secret-free `.tar.gz`, move it to another machine, and import it back into a fresh case in one click.
|
||||
- The container never accumulates: killing the session stops it, deleting the case removes it, and an instance-scoped boot reaper reaps containers whose case is gone.
|
||||
|
||||
Non-goals for the MVP: multi-tenant untrusted-code isolation guarantees (Codeman is loopback-default and single-operator, and the agent already runs `--dangerously-skip-permissions` on the host today), Kubernetes/compose orchestration, and per-command ephemeral containers.
|
||||
|
||||
## 2. Chosen architecture and why
|
||||
|
||||
The design grafts the strongest idea from each of the three proposals:
|
||||
|
||||
- Overlay-not-a-mode + faithful remote-SSH mirror (from "Docker Cases as a Location Overlay"): lowest churn, rides the existing quick-start / mux-sessions / state / recovery plumbing.
|
||||
- Convenient-but-hardened default with an opt-in sealed profile, plus exec-time name-only secret env (from "Sealed Sandbox"): a strict security improvement over today's on-host execution without the UX tax of forcing an in-container re-login.
|
||||
- One-artifact export + in-app import route (from "Container-as-Cargo"): the genuinely new, high-value capability Codeman lacks.
|
||||
|
||||
### Key decision 1: persistent per-CASE container, durable in-container tmux, AND resume-on-restart (the two-layer durability model)
|
||||
|
||||
Exactly one long-lived container per Docker case, named as a pure slug function `codeman-case-<slug>` (Docker charset `^[a-zA-Z0-9][a-zA-Z0-9_.-]+$`; Codeman already slugs case names for tmux), so create-if-missing and boot recovery are idempotent. PID1 is `sleep infinity` under `--init` (tini reaps zombies and forwards `docker stop`'s SIGTERM); the CLI is NOT the container command. The CLI runs inside a DURABLE in-container tmux on a dedicated socket `-L codeman-docker`, session `codeman-dkr-<id8>`, the direct analog of remote's `-L codeman-remote` / `codeman-ssh-<id8>`.
|
||||
|
||||
Two DIFFERENT failure surfaces need two DIFFERENT recovery layers, and conflating them is the central flaw the critic caught:
|
||||
|
||||
1. Codeman-PROCESS restart while the container stays up: the in-container tmux is still alive, so `tmux new-session -A` (attach-or-create) reattaches the SAME live agent and the paneCommand is ignored. This is the remote-SSH durability idiom and it works unchanged.
|
||||
2. CONTAINER stop / daemon restart / host reboot / OOM-kill: the in-container tmux is GONE (fresh PID1). `new-session -A` will now CREATE a fresh session and run the paneCommand, which would start a brand-new conversation. This is the case the raw plan silently lost. Because the transcript directory is bind-mounted from the host (Key decision 3), the fix is to launch with RESUME: the paneCommand becomes `exec claude --dangerously-skip-permissions --resume <claudeSessionId>` (codex uses `resume <id>`, gemini `--resume <id>`) whenever a captured `claudeSessionId` exists. The `-A` semantics make this self-selecting: the resume flag only ever executes when tmux is actually re-created, which is exactly when the live session was lost. When tmux is still alive (case 1), attach wins and the flag is inert.
|
||||
|
||||
Capturing / persisting / reusing the resume id (the missing mechanism the critic flagged): Codeman already learns `Session.claudeSessionId` from transcript correlation (which works here because projHash matches, Key decision 3) and persists it in `SessionState`. We thread that value into `createSessionOptions` / `respawnPaneOptions` for docker so `buildDockerLaunchCommand` can inject the resume flag on any relaunch. To make a NEW Codeman session (new `id8`) re-launched against the same case resume its predecessor's conversation, we ALSO persist `lastClaudeSessionId` on the `DockerCase` record; the quick-start docker branch seeds the new `Session` with it when the `dockerResumeOnStart` setting is on. First-ever launch has no id, so it starts fresh. This is user-decision 7 (default resume behavior).
|
||||
|
||||
Reconciling with stop-on-kill and with the `--restart` policy (the internal inconsistency the critic found): the container is created with `--restart no` uniformly (Codeman's idempotent create-if-missing plus boot recovery is the single recovery mechanism; a restart policy would not preserve the conversation anyway because a restarted container gets a fresh PID1/tmux). Boot recovery re-runs `buildDockerLaunchCommand` from the restored `MuxSession.docker` (`docker inspect || docker create; docker start`, then exec with resume), so a host reboot or daemon restart recreates+starts the container and resumes the conversation instead of the session vanishing. `reconcileSessions` (tmux-manager.ts ~1800-1815) must NOT hard-delete a docker session merely because no LOCAL pane exists after the local `-L codeman` server died; docker (like remote) sessions are restored from `mux-sessions.json` and relaunched. This relaunch path is explicitly part of Phase 4/Phase 3 recovery work, not assumed.
|
||||
|
||||
Why this over the alternatives: `docker exec` gets SIGHUP and dies when its client TTY closes, so a bare `docker exec claude` restarts the CLI on every reconnect/respawn. The inner tmux plus resume is what makes reconnect idempotent across BOTH failure surfaces. Because this durability is the single most important design point, tmux-in-image is a HARD gated prerequisite (`checkDockerTmuxAvailable`), never a silent fallback to bare exec. Rejected alternatives: ephemeral-per-run or bare-exec containers (no reattach durability); a literal `'docker'` `SessionMode` (touches dozens of switch/enum sites and diverges from the remote overlay precedent, since Docker is a LOCATION orthogonal to the 5 CLI backends).
|
||||
|
||||
### Key decision 2: CLI + auth delivery
|
||||
|
||||
One prebuilt base image (built once, contains NO secrets): `node:22-bookworm-slim` + `git tmux ripgrep ca-certificates`, `npm i -g @anthropic-ai/claude-code @openai/codex @google/gemini-cli opencode-ai`, an `agent` user, HOME dirs made writable by an arbitrary host uid via the OpenShift "gid 0, group-writable" convention (Key decision 6). Because the toolchain is baked, export is reproducible and needs no network at import time. The image name/namespace/registry and its refresh cadence are user-decision 2 (the `codeman/agent:base` placeholder implies a Docker Hub org the project may not own).
|
||||
|
||||
Credentials are delivered ONLY at runtime, two commit-safe channels, default convenient:
|
||||
|
||||
- OAuth/config-file CLIs (Claude Max/Pro, gcloud, opencode): bind-mount the host credential dirs read-write (`~/.claude`, `~/.codex`, `~/.gemini` + `~/.config/gcloud`, `~/.config/opencode`) so the common user "just works" with no in-container login. Because these are bind mounts, `docker commit` (which captures only the container's own writable layer, never bind mounts) physically cannot capture them, so exports stay secret-free.
|
||||
- API-key CLIs (codex/gemini): exec-time NAME-ONLY `docker exec --env OPENAI_API_KEY --env GEMINI_API_KEY ...` (no `=value`), sourced from Codeman's own process env. Only the key NAME appears in argv (no `ps` leak), and per-exec env is never captured by `docker commit`. This is the technique Codeman already uses via `tmux setenv` for the local Codex/Gemini panes, so it composes with existing machinery.
|
||||
|
||||
Per-host `DockerHost.mountCredentials` defaults `true` (convenient); setting it `false` yields a SEALED profile (no host cred mounts, in-container login only) for genuinely untrusted work. CRITICAL sealed-mode export rule (the leak the critic caught): in sealed mode the in-container login writes tokens into the container's OWN writable layer, which `docker commit` DOES capture, so a full-image export of a sealed container would ship credentials. Therefore full-image export is REFUSED for `mountCredentials:false` containers by default; the user may either take a workspace-only export (always safe) or opt into a pre-commit scrub that `docker exec`s `rm -rf ~/.claude ~/.codex ~/.gemini ~/.config/gcloud ~/.config/opencode` inside the container before commit (destructive to the in-container login, which is the point). This is enforced in the export route, not left to a manifest assertion.
|
||||
|
||||
Per-session `envOverrides` / `effort` / `codexConfig` / `geminiConfig` / `openCodeConfig` are REJECTED at quick-start exactly like the remote branch (session-routes.ts ~1698-1710). `modelOverride` is the one deliberate difference from remote: because the docker workspace is a REAL bind-mounted host dir that Codeman scaffolds (Key decision 5 and Section 6), `updateCaseModel()` can write the `model` key into `<workspace>/.claude/settings.local.json` and the in-container `claude` reads it, so the App Settings Claude Model picker works for docker cases. `effort` is a `--effort` CLI arg applied only by the local-spawn path we bypass, so it stays rejected (surfaced honestly in the UI, not silently inert). Per-mode command customization goes through `DockerHost.commands.<mode>` (`defaultDockerCommandForMode`, mirror of `defaultRemoteCommandForMode` at remote-hosts.ts:60). NEVER bake secrets into an image layer and NEVER pass a secret via create-time `-e` (both are committed).
|
||||
|
||||
Rejected alternative: sealed-by-default. For a single-operator loopback tool where the agent already runs skip-permissions on the host, forcing an in-container OAuth re-login is a UX regression with little real gain. We keep sealed as an opt-in. Rejected alternative: baking a login into the image, which leaks the instant you `docker save`.
|
||||
|
||||
### Key decision 3: workspace mount, container CWD, and transcript correlation
|
||||
|
||||
Bind-mount the host workspace dir into the container at the SAME absolute path (`dst == src`, mirror the host path), and set both `Session.workingDir` and the container workdir to that host path.
|
||||
|
||||
Two problems this solves that the raw proposals got wrong:
|
||||
|
||||
- File features: `DockerCase.hostWorkspacePath` is a REAL host directory, so `Session.workingDir = hostWorkspacePath` keeps file-routes, attachments, image-watcher, and previews working on real host bytes (unlike remote, where the path is remote-only and those features no-op). All three proposals wired `casePath = <container path>`; we deliberately diverge and use the host path.
|
||||
- Transcript correlation: Claude writes transcripts under `~/.claude/projects/<hash-of-CWD>/`. By mirroring the host path as the container CWD, the projHash computed inside the container equals the host-side hash Codeman's transcript/subagent/workflow watchers expect, so correlation keeps working (and, in turn, feeds the resume-id capture in Key decision 1). A `/workspace`-style fixed dst would break it. Mirror-vs-fixed is user-decision 3.
|
||||
|
||||
`resolveMuxAttachCwd` still returns `/tmp` for docker sessions (the LOCAL bash pane only runs `docker exec`; it never needs the workspace as its cwd), mirroring remote.
|
||||
|
||||
### Key decision 4: network default and the engine-specific host gateway
|
||||
|
||||
Default `bridge` (own netns, NAT egress, no inbound), per-case changeable to `none` (offline shell sandbox; warned because it breaks the API CLIs) or `custom` (a user-defined bridge `codeman-net-<slug>`, the chokepoint for a future egress allowlist). `host` networking and any `-p` inbound publish are structurally unrepresentable in the flag builder and schema. Rationale: every API-backed CLI (Claude, Codex, Gemini) plus npm/git needs egress, so `bridge` is the only sane functional default; `none` is reserved for `shell`.
|
||||
|
||||
The host-callback gateway alias is ENGINE-SPECIFIC (the critic's podman finding): Docker uses `host.docker.internal`, Podman uses `host.containers.internal` (Docker's alias only exists on recent podman). A helper `hostGatewayAlias(engine)` returns the right name; Section 2.5, the create args, the `CODEMAN_API_URL` rewrite, and the host-guard allowlist all consume it, and BOTH aliases are added to the allowlist so a mixed fleet keeps working.
|
||||
|
||||
### Key decision 5: hooks actually reach the host AND are actually installed
|
||||
|
||||
Two independent things must both be true for a hook to fire, and the raw plan wired only the first:
|
||||
|
||||
1. Network reachability. Claude Code hooks POST to `$CODEMAN_API_URL` (`curl -sk`). Inside a bridge container `localhost` is the container and prod binds `127.0.0.1`, so we set `--add-host <gatewayAlias>:host-gateway` on create (skipped on Docker Desktop, where the alias is native), add the gateway alias to the host guard, and provide `CODEMAN_API_URL` and the hook secret (below).
|
||||
2. Hook INSTALLATION. Hooks live in `<workspace>/.claude/settings.local.json`, written by the quick-start scaffolding block (around session-routes.ts ~1776) that calls `writeHooksConfig()` / `updateCaseModel()`. The raw plan extended the `!remote` guard to `!remote && !docker`, which would SKIP that block and silently disable ALL hooks regardless of networking. For docker the workspace is a REAL bind-mounted host dir, so the scaffolding block MUST run. Precise fix: extend to `!remote && !docker` ONLY the LOCAL-CLI-availability and local-spawn guards (the ones that stat the local binary or build the local spawn command); leave the workspace-scaffolding guard at `!remote` so it runs for docker. This same decision is what makes `modelOverride` work (Key decision 2). Consequence, surfaced as user-decision 4: linking a docker case now WRITES `.claude/settings.local.json` (and the CLAUDE.md scaffold, matching local-case behavior) into the user's real host directory, a behavioral shift from "link a dir" to "link and scaffold a dir."
|
||||
|
||||
`CODEMAN_API_URL` derivation (the wrong-scheme bug the critic caught): prod is HTTPS-only on 3000, and `server.ts` (~2000) auto-sets `process.env.CODEMAN_API_URL = ${protocol}://${apiHost}:${port}`. Hardcoding `http://host.docker.internal:3000` fails every hook. Instead a pure helper `containerApiUrl(process.env.CODEMAN_API_URL, engine)` parses the running URL and substitutes ONLY the hostname with `hostGatewayAlias(engine)`, preserving scheme and port (`https://host.docker.internal:3000`). Unit-tested against http, https, non-default ports, and both engines. Passed as create-time `--env CODEMAN_API_URL=<derived>` (case-stable, non-secret).
|
||||
|
||||
Hook secret and session attribution:
|
||||
- `~/.codeman/hook-secret` is bind-mounted read-only to a container path; `--env CODEMAN_HOOK_SECRET_FILE=<that path>` is create-time (a path is non-secret; the bytes ride the bind mount and are never committed).
|
||||
- `CODEMAN_SESSION_ID` (which the generated hooks reference at hooks-config.ts:78-80 to attribute events) plus `CODEMAN_MUX=1` are SESSION-scoped, so they are passed at EXEC time via `docker exec --env CODEMAN_SESSION_ID=<id> --env CODEMAN_MUX=1` (non-secret, value inline is fine, and exec env is not committed). Because a `tmux` session started fresh only inherits the invoking env when it starts the SERVER, the launch chain ALSO runs `tmux -L codeman-docker setenv -g CODEMAN_SESSION_ID <id>` (and `CODEMAN_MUX`) so reattaches and newly created panes see the same values. This mirrors how Codeman already injects per-session env into tmux for the external CLIs.
|
||||
|
||||
Hooks-in-MVP-vs-deferred stays user-decision 4; if deferred, docker ships as explicitly hook-degraded and we lean on output-based idle detection through the docker-exec PTY.
|
||||
|
||||
### Key decision 6: uid / HOME / rootless enforcement / macOS Docker Desktop
|
||||
|
||||
The raw plan showed `--user 1000:1000` in one place and `--user "$(id -u):$(id -g)"` in another and never resolved HOME writability; this section fixes all of it.
|
||||
|
||||
- Linux native (docker rootful or rootless): run `--user <hostUid>:0` (host uid, GID 0). The image follows the OpenShift arbitrary-uid convention: `HOME=/home/agent`, and `/home/agent` plus the tool cache dirs (`~/.npm`, `~/.cache`, `~/.config`) are owned `root:0` and group-writable (`chmod -R g+w`, `g+s` on dirs) so a process with GID 0 can write HOME even though its UID is not 1000. This keeps workspace files host-owned (the agent's UID is the host UID) AND keeps HOME writable, so the CLIs actually start.
|
||||
- Podman rootless: use `--userns=keep-id` (maps the host uid to the image's `agent` uid inside the container) instead of `--user`, so `/home/agent` is owned by the running user and workspace files are host-owned. This is a real per-engine branch in `buildDockerCreateArgs`.
|
||||
- macOS Docker Desktop: `--user <macUid>` (e.g. 501) does not own the image's `/home/agent`, so non-bind HOME writes fail EACCES and the CLIs may not start; Desktop also does its own bind-mount uid translation, provides `host.docker.internal` natively (no `--add-host`), and its VM memory ceiling can cap `--memory`. Detect Desktop via `docker info` (Server OS `linuxkit` / `OperatingString` contains "Docker Desktop") and take a dedicated path: do NOT pass `--user` (run as the image's baked `agent` uid and rely on Desktop's translation for workspace access), skip `--add-host`, and note in the UI that memory caps are subject to the VM ceiling.
|
||||
|
||||
Rootless resource-cap enforcement (the silently-inert risk): rootless Docker without cgroup-v2 systemd delegation (`Delegate=yes`) silently IGNORES `--memory`/`--cpus`/`--pids-limit`. The probe checks `docker info` for `CgroupVersion=2` plus rootless plus delegation; if caps cannot be enforced, `checkDockerAvailable` returns `capsEnforced:false` and the link/probe surfaces "resource caps are advisory on this engine." Whether to REQUIRE delegation or ship-with-warning is user-decision 6.
|
||||
|
||||
## 3. Data model
|
||||
|
||||
New TypeScript types in `src/types/session.ts`, added right after the remote types (lines 46-99). SessionMode (line 44) is UNCHANGED.
|
||||
|
||||
```ts
|
||||
export type DockerCommandMode = Extract<SessionMode, 'shell' | 'claude' | 'opencode' | 'codex' | 'gemini'>;
|
||||
export type DockerEngine = 'docker' | 'podman';
|
||||
export type DockerNetworkMode = 'bridge' | 'none' | 'custom'; // never 'host'
|
||||
|
||||
export interface DockerResourceLimits {
|
||||
memory?: string; // '4g' -> --memory 4g --memory-swap 4g (swap==memory: real OOM cap)
|
||||
cpus?: string; // '2'
|
||||
pidsLimit?: number; // 512 (fork-bomb guard)
|
||||
nofile?: string; // '4096:8192'
|
||||
shmSize?: string; // optional; only when a tool needs /dev/shm
|
||||
}
|
||||
|
||||
export interface DockerHost {
|
||||
id: string;
|
||||
label: string;
|
||||
engine?: DockerEngine; // default resolved by probe (docker, else podman)
|
||||
image: string; // default resolved image ref (see user-decision 2)
|
||||
daemonHost?: string; // advanced: -H ssh://user@host / DOCKER_HOST
|
||||
context?: string; // advanced: --context <ctx>
|
||||
network?: DockerNetworkMode; // default 'bridge'
|
||||
networkName?: string; // when network === 'custom'
|
||||
resources?: DockerResourceLimits;
|
||||
mountCredentials?: boolean; // default true (false = sealed; blocks full-image export)
|
||||
hooksEnabled?: boolean; // default true (host-gateway callback wiring)
|
||||
resumeOnStart?: boolean; // default true (see Key decision 1 / user-decision 7)
|
||||
commands?: Partial<Record<DockerCommandMode, string>>;
|
||||
extraCreateArgs?: string[]; // validated like extraSshOptions
|
||||
extraExecArgs?: string[];
|
||||
}
|
||||
|
||||
export interface DockerCase {
|
||||
name: string;
|
||||
type: 'docker';
|
||||
hostId: string;
|
||||
hostWorkspacePath: string; // absolute HOST dir: bind src + Session.workingDir
|
||||
containerWorkdir?: string; // container path; default = hostWorkspacePath (mirror -> projHash match)
|
||||
container?: string; // default codeman-case-<slug>
|
||||
lastClaudeSessionId?: string; // captured resume id (Key decision 1)
|
||||
}
|
||||
|
||||
export interface SessionDocker { // flattened, round-trips through mux/state (mirror SessionRemote at 91)
|
||||
hostId: string;
|
||||
label: string;
|
||||
engine: DockerEngine;
|
||||
image: string;
|
||||
containerName: string;
|
||||
hostWorkspacePath: string;
|
||||
containerWorkdir: string;
|
||||
network: DockerNetworkMode;
|
||||
networkName?: string;
|
||||
resources?: DockerResourceLimits;
|
||||
mountCredentials: boolean;
|
||||
hooksEnabled: boolean;
|
||||
resumeOnStart: boolean;
|
||||
daemonHost?: string;
|
||||
context?: string;
|
||||
commands?: Partial<Record<DockerCommandMode, string>>;
|
||||
extraCreateArgs?: string[];
|
||||
extraExecArgs?: string[];
|
||||
configHash?: string; // drift detection (Key decision, Section 4)
|
||||
}
|
||||
```
|
||||
|
||||
- `SessionState` gains `docker?: SessionDocker` immediately after `remote?` (line 219). It persists automatically because `SessionState` is structural and `state-store.ts` stores `toState()` verbatim.
|
||||
- `src/mux-interface.ts`: add `docker?: SessionDocker` to `MuxSession` (after line 38), `CreateSessionOptions` (after 81), `RespawnPaneOptions` (after 105). `MuxSession.docker` round-trips through `mux-sessions.json` automatically.
|
||||
- `src/types/api.ts` `CaseInfo`: add `'docker'` to the `location` union and a `docker?: { hostId; container; image?; path; network }` display block.
|
||||
- `src/services/unified-session-service.ts`: add a boolean `docker?` flag on `UnifiedSessionItem` and source rows, set from `MuxSession.docker` presence (mirror the `remote` flag at ~line 200 and the harvest at session-routes.ts:2313).
|
||||
|
||||
New state files (all via `dataPath()`, mirroring `remote-hosts.json` / `remote-cases.json`):
|
||||
|
||||
- `~/.codeman/docker-hosts.json` (reusable engine/image/network/resource profiles).
|
||||
- `~/.codeman/docker-cases.json` (`name -> DockerCase`, including `lastClaudeSessionId`).
|
||||
- `~/.codeman/docker-exports/` (dedicated dir for `.image.tar.gz` + `.workspace.tar.gz` + `manifest.json`; never inline in state.json; retention/pruning per Section 5).
|
||||
|
||||
No new `state.json` / `mux-sessions.json` files: `SessionState.docker` and `MuxSession.docker` ride the existing serialization.
|
||||
|
||||
## 4. Container lifecycle (exact command shapes)
|
||||
|
||||
All builders are PURE string functions (directly unit-testable). Host values interpolated into the outer `bash -c "..."` layer (container name, image, workdir, host paths) are `shellescape()`'d and, for user-supplied fields, schema-rejected for `$`/backtick via `NO_SHELL_META`. The escaping chain here is DEEPER than remote's single `ssh '<tmux ...>'`: the whole `docker inspect || docker create <dozens of --mount/--env/shellescaped host paths>` is interpolated into `bash -c "..."` then `JSON.stringify`'d into respawn-pane. This is a known place to get stuck, so it is covered by concrete escaping tests (Section 9), including host workspace paths containing spaces, not just a "we call shellescape" claim.
|
||||
|
||||
New in `src/tmux-manager.ts`:
|
||||
|
||||
```ts
|
||||
const DOCKER_TMUX_SOCKET = 'codeman-docker';
|
||||
// 'dkr' letters deliberately FAIL SAFE_MUX_NAME_PATTERN (^codeman-[a-f0-9-]+$),
|
||||
// so a Codeman running INSIDE the container never adopts/resizes/respawns our session.
|
||||
export function dockerTmuxSessionName(id: string): string { return `codeman-dkr-${id.slice(0, 8)}`; }
|
||||
```
|
||||
|
||||
`buildDockerBaseArgs(docker)` (pure, in `docker-hosts.ts`, mirror of `buildSshConnectionArgs`) emits the engine prefix tokens: `docker` (or `podman`) + optional `--context <ctx>` or `-H <daemonHost>`. `buildDockerCreateArgs(docker, sessionId)` emits the `docker create` flag array (with the per-engine uid/userns branch from Key decision 6).
|
||||
|
||||
IMAGE PRESENCE (before any create, the auto-pull footgun the critic caught): the launch chain runs `docker image inspect <image> >/dev/null 2>&1` first; on miss it exits with a distinct message ("base image <ref> not present: build with scripts/build-agent-image.mjs or pull it") rather than triggering a blocking multi-GB auto-pull inside the tmux pane. `docker create` carries `--pull=never`. The tmux-availability probe likewise uses `docker run --rm --pull=never <image> sh -lc 'command -v tmux'` and reports the same build/pull hint if the image is absent, so the 15s-bounded probe never hangs on a pull.
|
||||
|
||||
CREATE (the ensure step, embedded in the launch string):
|
||||
|
||||
```
|
||||
docker create \
|
||||
--name codeman-case-myproj --hostname myproj \
|
||||
--label codeman.managed=1 --label codeman.instance=<CODEMAN_INSTANCE> \
|
||||
--label codeman.case=myproj --label codeman.session=<id8> \
|
||||
--label codeman.confighash=<hash> \
|
||||
--pull=never --init --restart no \
|
||||
--user 1000:0 \
|
||||
--workdir '/home/arkon/cases/myproj' \
|
||||
--mount type=bind,src='/home/arkon/cases/myproj',dst='/home/arkon/cases/myproj' \
|
||||
--mount type=bind,src='/home/arkon/.claude',dst='/home/agent/.claude' \
|
||||
--mount type=bind,src='/home/arkon/.codeman/hook-secret',dst='/home/agent/.codeman/hook-secret',readonly \
|
||||
--add-host host.docker.internal:host-gateway \
|
||||
--memory 4g --memory-swap 4g --cpus 2 --pids-limit 512 --ulimit nofile=4096:8192 \
|
||||
--cap-drop ALL --security-opt no-new-privileges \
|
||||
--network bridge \
|
||||
--env HOME=/home/agent --env TERM=xterm-256color --env COLORTERM=truecolor \
|
||||
--env CODEMAN_API_URL=https://host.docker.internal:3000 \
|
||||
--env CODEMAN_HOOK_SECRET_FILE=/home/agent/.codeman/hook-secret \
|
||||
codeman/agent:base \
|
||||
sleep infinity
|
||||
```
|
||||
|
||||
- `--user 1000:0` shown is the Linux-native form with GID 0 (Key decision 6); it is actually `--user <hostUid>:0`, or `--userns=keep-id` for podman rootless, or omitted on Docker Desktop. The literal is illustrative only.
|
||||
- Create-time `--env` carries only NON-SESSION, non-secret, case-stable values (safe to be committed): the DERIVED `CODEMAN_API_URL` (https-preserving, Key decision 5) and the hook-secret FILE PATH. `CODEMAN_SESSION_ID`/`CODEMAN_MUX` and the codex/gemini key NAMES are exec-time only.
|
||||
- `codeman.instance=<CODEMAN_INSTANCE>` is REQUIRED on the label set so the boot reaper is instance-scoped (a beta/second instance must never reap prod's containers).
|
||||
- `codeman.confighash` is a stable hash of the drift-relevant create args (image, resources, network, mounts, non-session env). Drift detection (user story 2, the config-never-takes-effect gap): on launch the ensure block compares the desired hash to the existing container's label; on mismatch the launch does NOT silently reuse the stale container. Instead the docker route returns a "container config changed, recreate?" action (SSE + UI confirm), and on confirm Codeman `docker rm`'s and recreates. rm destroys in-image (non-bind) state, but the workspace and transcripts survive on their bind mounts and the conversation is restored via `--resume`, so the recreate is safe. Auto-recreate-vs-prompt is a UI choice; the MVP prompts.
|
||||
- `--restart no` (resolved consistently with Key decision 1; recovery is Codeman's idempotent create-if-missing, not an engine restart policy, which also matters for Podman which has no daemon).
|
||||
|
||||
EXEC (`buildDockerLaunchCommand`, the docker analog of `buildRemoteLaunchCommand`, TTY-correct, resume-aware). The whole thing is ONE `bash -c` string that image-checks, ensures, starts, primes tmux env, then execs:
|
||||
|
||||
```
|
||||
docker image inspect codeman/agent:base >/dev/null 2>&1 || { echo 'Codeman: base image codeman/agent:base not present (build or pull it)'; exit 1; } ; \
|
||||
docker inspect codeman-case-myproj >/dev/null 2>&1 || docker create <all create args above> ; \
|
||||
docker start codeman-case-myproj >/dev/null 2>&1 || { echo 'Codeman: container codeman-case-myproj failed to start (daemon down?)'; exit 1; } ; \
|
||||
exec docker exec -it \
|
||||
--workdir '/home/arkon/cases/myproj' \
|
||||
--env TERM=xterm-256color --env COLORTERM=truecolor \
|
||||
--env CODEMAN_SESSION_ID=1a2b3c4d --env CODEMAN_MUX=1 \
|
||||
--env OPENAI_API_KEY --env GEMINI_API_KEY \
|
||||
codeman-case-myproj \
|
||||
sh -lc 'tmux -L codeman-docker setenv -g CODEMAN_SESSION_ID 1a2b3c4d \; setenv -g CODEMAN_MUX 1 \; new-session -A -s codeman-dkr-1a2b3c4d -c '\''/home/arkon/cases/myproj'\'' '\''cd /home/arkon/cases/myproj && exec claude --dangerously-skip-permissions --resume <claudeSessionId>'\'' \; set -t codeman-dkr-1a2b3c4d status off \; set -t codeman-dkr-1a2b3c4d mouse off \; set -t codeman-dkr-1a2b3c4d prefix C-q \; set -s escape-time 0'
|
||||
```
|
||||
|
||||
- `docker exec -it`: `-t` allocates a PTY and forwards SIGWINCH into the container so the Ink TUI re-lays-out on pane resize; `TERM`/`COLORTERM` prevent degraded rendering. `--env OPENAI_API_KEY` (name only) is present only for codex/gemini and is exec-time (never committed). `CODEMAN_SESSION_ID`/`CODEMAN_MUX` are exec-time values plus a `tmux setenv -g` prime so reattaches and new panes inherit them (Key decision 5).
|
||||
- `--resume <claudeSessionId>` (codex `resume <id>`, gemini `--resume <id>`) is appended to `modeCommand` ONLY when a captured id exists; on first launch it is omitted. `new-session -A` makes the flag inert on a live-tmux reattach and effective only when tmux is re-created (Key decision 1).
|
||||
- `modeCommand = docker.commands?.[mode] || defaultDockerCommandForMode(mode)` (`exec claude --dangerously-skip-permissions`, `exec bash -l`, etc.), with the resume suffix injected by the builder.
|
||||
- Escaping survives every layer identically to remote in shape but deeper in nesting: `paneCommand` (`cd ... && exec ...`) is one shellescaped tmux arg, the whole `tmuxInvocation` is one shellescaped `sh -lc` arg, and the outer string is `JSON.stringify()`'d into `bash -c` by respawn-pane (tmux-manager.ts:1329).
|
||||
|
||||
Wire-up (extend the two existing seams to 3-way):
|
||||
|
||||
- createSession (tmux-manager.ts:1276): `const fullCmd = docker ? buildDockerLaunchCommand({ mode, docker, sessionId, resumeSessionId }) : remote ? buildRemoteLaunchCommand({ mode, remote, sessionId }) : localFullCmd;`
|
||||
- launchCmd cd-skip (tmux-manager.ts:1327): `const launchCmd = (remote || docker) ? fullCmd : \`cd ${JSON.stringify(workingDir)} && ${fullCmd}\`;`
|
||||
- respawnPane: same two edits at lines 1524 and 1542.
|
||||
|
||||
START / reattach-after-reboot: the ensure block (image-check, `docker inspect || docker create`, `docker start`) is fully idempotent, so boot recovery just re-runs `buildDockerLaunchCommand` from the restored `MuxSession.docker` with the persisted resume id. A rebooted host recreates the container and resumes the conversation.
|
||||
|
||||
DOCKER-DOWN surfacing (the PTY-exit-breaker false-trip risk): if `docker start` or `docker exec` cannot attach (daemon down, container missing), the launch prints a docker-specific message and exits, which alone would still count toward `session-pty-exit-breaker` and show a generic "respawn breaker tripped" push. To avoid masking the cause, the docker reattach path runs a fast `checkDockerAvailable` pre-flight: if the daemon/container is unreachable, Codeman broadcasts a docker-specific error (SSE + push, "container <name> is not running / daemon down") and SKIPS the auto-reattach that would trip the breaker, rather than fast-looping `docker exec`.
|
||||
|
||||
STOP / KILL (`killSession` Strategy 3c, right after remote's Strategy 3b at tmux-manager.ts:1719, guarded by `IS_TEST_MODE`):
|
||||
|
||||
```ts
|
||||
if (session.docker) {
|
||||
// best-effort, fire-and-forget, timeout-bounded so it never blocks the local kill
|
||||
execAsync(buildDockerKillCommand({ docker: session.docker, sessionId }), { timeout: EXEC_TIMEOUT_MS }).catch(() => {});
|
||||
}
|
||||
```
|
||||
|
||||
`buildDockerKillCommand` emits: `docker exec codeman-case-<slug> tmux -L codeman-docker kill-session -t codeman-dkr-<id8> ; docker stop -t 10 codeman-case-<slug>`. Stopping frees CPU/RAM and, per Key decision 1, is safe for conversation continuity because the NEXT launch resumes from the bind-mounted transcript via `--resume`. Whether to stop at all (RAM vs instant live-agent reattach) is user-decision 6/1 (reframed honestly). The bind-mounted workspace and transcripts always survive on the host.
|
||||
|
||||
REMOVE: only on explicit case delete (`docker rm -f codeman-case-<slug>`), gated behind an "export first?" UI prompt because rm destroys any in-image (non-bind) state. Instance-scoped boot reaper (fixing the racy/cross-instance reaper): after `docker-cases.json` is loaded AND after `restoreMuxSessions` has run, enumerate `docker ps -a --filter label=codeman.managed=1 --filter label=codeman.instance=<CODEMAN_INSTANCE> --format '{{.Names}}\t{{index .Labels "codeman.case"}}'` and `docker rm -f` only containers whose case is gone from THIS instance's `docker-cases.json`. The instance filter is what stops a beta reaping prod's containers (the exact cross-instance hazard the project memory warns about).
|
||||
|
||||
AVAILABILITY PROBE (`docker-hosts.ts`, timeout-bounded like `checkRemoteTmuxAvailable`'s 15s, `IS_TEST_MODE` no-op):
|
||||
|
||||
```
|
||||
docker info --format '{{json .}}' # server up, CgroupVersion, rootless, OS (Desktop detect), cap-delegation
|
||||
docker image inspect <image> --format '{{.Id}}' # image PRESENT (no auto-pull)
|
||||
docker run --rm --pull=never <image> sh -lc 'command -v tmux' # tmux-in-image gate (hard prerequisite), only if image present
|
||||
```
|
||||
|
||||
`checkDockerAvailable()` returns `{ ok, engine, rootless, isDesktop, cgroupV2, capsEnforced }` (parse `SecurityOptions` for `name=rootless`, `CgroupVersion`, delegation, and Server OS for Desktop). `checkDockerTmuxAvailable(host)` returns a structured result with a user-facing error and correct install hint (NOT `npm install -g`; the hint is "build/pull the base image" for a missing image and "install docker or podman" for a missing engine).
|
||||
|
||||
IN-CONTAINER CLI VERSION (fixing the #154 wheel-forwarding regression): the raw plan skipped the LOCAL `cliVersion` probe for docker (correct, since it reports the HOST claude) but left `cliVersion` undefined, which disables trackpad wheel-forwarding. Instead, for docker sessions Codeman runs an IN-CONTAINER probe `docker exec <container> claude --version` (bounded, `IS_TEST_MODE` no-op) and feeds THAT into `cliVersion`. This also means a stale baked CLI is visible; combined with the rebuild-cadence in user-decision 2, agents are not silently pinned to an old claude.
|
||||
|
||||
## 5. Export / Import
|
||||
|
||||
EXPORT is a concurrency-bounded job (reuse `runWithConversionLimit` from `document-conversion-limiter.ts` so N simultaneous exports cannot fork-bomb the host). Route `POST /api/docker-cases/:name/export`.
|
||||
|
||||
Preconditions (the consistency and leak risks the critic caught):
|
||||
- Sealed guard: if `mountCredentials:false`, full-image export is REFUSED unless the caller explicitly opts into the pre-commit scrub (Key decision 2). Workspace-only export is always allowed.
|
||||
- Quiesce + free-space: require the session idle, then `docker pause` the container spanning BOTH the workspace tar AND the commit so the two artifacts are mutually consistent (the raw plan paused only the commit, leaving the bind-mount tar to run against a mid-write agent). Before any heavy step, precheck free space in the exports dir and in `/var/lib/docker`; if below `DOCKER_EXPORT_MIN_FREE_BYTES`, refuse with a clear error (a full `/var/lib/docker` wedges the daemon and breaks EVERY session on the host).
|
||||
|
||||
Steps (all cleanup in try/finally so a mid-way failure never orphans an intermediate image or leaves the container paused):
|
||||
|
||||
1. `docker commit -c 'LABEL codeman.exported=1' codeman-case-<slug> codeman/export-<slug>:<ts>` (unique tag per export defeats the stale-image trap). Optional pre-commit scrub in sealed mode as above; also blank instance-specific committed env (`-c 'ENV CODEMAN_API_URL='` etc.) so the image carries no stale host references.
|
||||
2. `docker save codeman/export-<slug>:<ts> | gzip` streamed in fixed 8192-byte chunks to `~/.codeman/docker-exports/<slug>-<ts>.image.tar.gz`. Uses `docker save` (layers + repo:tag + CMD), never `docker export` (flat rootfs), so restore is a trivial `docker load`.
|
||||
3. `tar --numeric-owner -C <hostWorkspacePath> -czf <slug>-<ts>.workspace.tar.gz .` while paused (the bind-mounted workspace is NOT in the image, so it travels separately and consistently).
|
||||
4. Write `manifest.json`: schema version, caseName, image tag, engine, containerWorkdir, resource/network config, codeman version, base-image digest, createdAt, per-member sha256, `mountCredentials`, and `secretFree` (true only for convenient-mode or scrubbed-sealed exports).
|
||||
5. `docker rmi codeman/export-<slug>:<ts>` in the `finally` (delete the intermediate committed image regardless of success), then `docker unpause`.
|
||||
|
||||
The three files are wrapped in one bundle `<slug>-<ts>.codeman-container.tgz` and offered as a downloadable artifact through the existing file-routes streaming + attachment-registry handoff.
|
||||
|
||||
Retention / disk budget (user-decision 3): `docker-exports/` is capped at `DOCKER_EXPORT_KEEP` most-recent bundles with an auto-prune on each new export, plus the free-space precheck above. Workspace scrub: the WORKSPACE tar gets a scan/warn pass for agent-created `.env` / `.git/credentials` (a distinct leak channel from container creds). A lighter "workspace-only" export (just the workspace tar, no commit/save) is the fast default for 24h+ runs; full-image is the explicit heavier option (user-decision 7 in the original list, now decision on the default button below).
|
||||
|
||||
What travels: the baked toolchain image plus any in-image writes, and the workspace tar. What does NOT travel: bind-mounted credentials (physically excluded from commit) and anything that lived only in a bind mount. Secret-free by construction in convenient mode, and enforced (refuse-or-scrub) in sealed mode.
|
||||
|
||||
IMPORT `POST /api/docker-cases/import` (untrusted-bundle containment, the traversal/overwrite risk): stream the uploaded bundle, validate every manifest checksum BEFORE any extraction or load. Extract the workspace tar with `tar --no-absolute-names -C <fresh dir>` PLUS per-entry validation rejecting any member whose normalized path escapes the destination (leading `/` or `..` components). `gunzip | docker load` the image, then RE-TAG the loaded image id into a quarantined namespace `codeman/imported-<slug>:<ts>` and NEVER allow the load to overwrite `codeman/agent:base` or any pre-existing tag (capture the loaded id, ignore the bundle's repo:tag). Create a NEW `DockerCase` pointing at the quarantined image with THIS host's mounts/creds and the manifest's resource/network config, and recreate the container hardened (cap-drop ALL, no-new-privileges, non-root, `--pull=never`, CMD overridden to `sleep infinity`). The destination supplies its own login, so credentials never cross machines. Plus `GET /api/docker-exports` (list) and `DELETE /api/docker-exports/:filename`, all behind Codeman's existing auth / loopback-default / host-guard / Origin-CSRF stack.
|
||||
|
||||
## 6. Codeman integration (file-by-file, mirroring the remote-SSH feature)
|
||||
|
||||
- `src/types/session.ts`: add `DockerCommandMode`, `DockerEngine`, `DockerNetworkMode`, `DockerResourceLimits`, `DockerHost`, `DockerCase`, `SessionDocker` (Section 3). Add `docker?: SessionDocker` to `SessionState` after line 219. SessionMode (line 44) UNCHANGED.
|
||||
- `src/mux-interface.ts`: add `docker?: SessionDocker` to `MuxSession` (38), `CreateSessionOptions` (81), `RespawnPaneOptions` (105).
|
||||
- `src/docker-hosts.ts` (NEW, direct mirror of `src/remote-hosts.ts`): `readDockerHosts`/`writeDockerHosts`/`readDockerCases`/`writeDockerCases` (via `dataPath`, including `lastClaudeSessionId` read/write), `defaultDockerCommandForMode` (mirror line 60), `dockerDisplayPath` (`container:/path`, mirror `remoteDisplayPath` at 205), `toSessionDocker(host, case)` (mirror `toSessionRemote` at 212), `buildDockerBaseArgs`/`buildDockerCreateArgs` (per-engine uid/userns branch), `hostGatewayAlias(engine)`, `containerApiUrl(processApiUrl, engine)` (scheme+port-preserving, unit-tested), `checkDockerAvailable`/`checkDockerTmuxAvailable`/`probeDockerCliVersion` (15s-bounded, `IS_TEST_MODE` no-op), a config-hash helper for drift, its own POSIX `shellescape` copy (mirror line 83). `const IS_TEST_MODE = !!process.env.VITEST;` gates every real `docker` invocation.
|
||||
- `src/tmux-manager.ts`: add `DOCKER_TMUX_SOCKET`, `dockerTmuxSessionName`, `buildDockerLaunchCommand` (resume-aware, image-check, env-prime), `buildDockerKillCommand` (Section 4). Extend the two `fullCmd` ternaries (1276, 1524) and the two `launchCmd` cd-skips (1327, 1542). Add `killSession` Strategy 3c after 1719. Ensure `reconcileSessions` (~1800-1815) does NOT hard-delete docker sessions on local-tmux death (recovery relaunch path).
|
||||
- `src/session.ts`: add `_docker?: SessionDocker` field (mirror `_remote` at 403), constructor arg (477), assignment (550). Thread `docker: this._docker` and `resumeSessionId: this._claudeSessionId` into BOTH `createSessionOptions` and `respawnPaneOptions` in `startInteractive` (1352/1370) and the second path (1740/1750). Emit `docker: this._docker` in `toState()` (1010). Replace the LOCAL cliVersion probe at 1320 for docker with the IN-CONTAINER `probeDockerCliVersion` (do not merely skip it). Extend `resolveMuxAttachCwd(workingDir, remote, docker)` (215) to return `/tmp` when `docker` is set. On claudeSessionId capture, persist it to the owning `DockerCase.lastClaudeSessionId`.
|
||||
- `src/web/server.ts`: in `restoreMuxSessions` (2160), add `docker: muxSession.docker ?? savedState?.docker` to the `new Session({...})` call (2195-2216), and skip docker in the same `isExternalCliMode`/Ralph recovery guards as remote. Register the instance-scoped boot reaper to run AFTER docker-cases load and AFTER `restoreMuxSessions`. Ensure `CODEMAN_API_URL` derivation reads the SAME `process.env.CODEMAN_API_URL` the server sets at ~2000.
|
||||
- `src/web/schemas.ts`: add `DockerHostSchema` and `DockerCaseLinkSchema` (below). The three mode enums (177/373/705) and `QuickStartSchema` (368) UNCHANGED (docker resolves by `caseName` lookup like remote).
|
||||
- `src/web/routes/session-routes.ts`: import the docker helpers from `../../docker-hosts.js`. Add a docker branch in `/api/quick-start` parallel to the remote branch (1686-1720): `readDockerCases` -> find by `caseName` -> `readDockerHosts` -> find by `hostId`; reject `envOverrides`/`effort`/`codexConfig`/`geminiConfig`/`openCodeConfig` (but ACCEPT `modelOverride`, which flows via scaffolded `settings.local.json`); run `checkDockerAvailable` + `checkDockerTmuxAvailable` (image-present, engine, caps-enforced); surface `capsEnforced:false` and Desktop notes; set `casePath = dockerCase.hostWorkspacePath` (REAL host dir), `docker = toSessionDocker(host, dockerCase)`, and seed `resumeSessionId` from `dockerCase.lastClaudeSessionId` when `resumeOnStart`. Extend the LOCAL-availability and local-spawn guards (around 1796/1810) to `!remote && !docker`, but DO NOT extend the workspace-scaffolding guard (~1776, `writeHooksConfig`/`updateCaseModel`), which MUST run for docker. Pass `docker` into `new Session` (1847); `autoConfigureRalph` (1853) gated on `!docker`. Add `docker: m.docker !== undefined ? true : undefined` to the unified harvest (2313).
|
||||
- `src/web/routes/case-routes.ts`: import the docker read/write/check helpers + schemas. Add a docker listing loop in `GET /api/cases` (mirror 94-119, `location: 'docker'`, `docker: {...}` via `dockerDisplayPath`). Add `/api/docker-hosts` GET/POST/PUT/DELETE (mirror 168-204) and `POST /api/cases/docker-link` (mirror 206-232; run `checkDockerAvailable`/`checkDockerTmuxAvailable` at link time; broadcast `CaseLinked` with `type: 'docker'`). Add a docker-unlink branch to `DELETE /api/cases/:name` (mirror 288-296; `docker rm -f`; broadcast `CaseDeleted` `type: 'docker-unlinked'`). Add the docker branch to single-case `GET` (mirror 358-368). Add `POST /api/docker-cases/:name/export`, `/import`, `GET/DELETE /api/docker-exports`, and a `POST /api/docker-cases/:name/recreate` (drift confirm) per Sections 4 and 5.
|
||||
- `src/web/sse-events.ts` + `src/web/public/constants.js`: reuse `CaseLinked`/`CaseDeleted` for CRUD. Add `docker:exportProgress`, `docker:exportComplete`, `docker:importComplete`, `docker:configDrift`, and `docker:containerError` to BOTH registries (kept in sync per CLAUDE.md).
|
||||
- Frontend `src/web/public/index.html` (~1831): add a Docker `modal-tab-btn` next to Remote; add a `#case-docker` panel mirroring `#case-remote` with `dockerCaseName`, `dockerHostWorkspacePath`, `dockerContainer`, `dockerImage`, `dockerHostId`, and an Advanced `<details>` for network mode, resource caps, `mountCredentials`, `resumeOnStart`, and remote daemon. Surface a "scaffolds .claude into this host dir" note (user-decision 4) and a "resource caps advisory on this engine" warning when `capsEnforced:false`.
|
||||
- Frontend `src/web/public/session-ui.js`: `formatCasePickerLabel` (48) + `buildCasePickerOptions` (71-73) handle `location === 'docker'` (`name @ container`, add container/image to the search haystack); `resetCaseModalFields` (~1514) add a `dockerFields` array; `switchCaseModalTab` (1573/1580/1597) handle `'case-docker'`; `submitCaseModal` add the docker branch; new `linkDockerCase()` (mirror `linkRemoteCase` at 1689) POSTing `/api/docker-hosts` then `/api/cases/docker-link`, sending omitted optionals as `undefined` (spread `...(x ? {x} : {})`, never `null`, per the Zod `.optional()`-rejects-null gotcha); `runClaude` (520) / `runShell` (702) extend the `location === 'remote'` routing to also match `'docker'`; `runOpenCode`/`runCodex`/`runGemini` (792/846/900) make the `isRemote` checks `isRemoteOrDocker` so local status probes are skipped. In the session-options Summary tab, note that `effort` is inert for docker (rejected) while `model` IS honored via `settings.local.json`.
|
||||
- Frontend `src/web/public/panels-ui.js` (425-426): add `caseItem?.docker?.path`/`container` to the case-search fields.
|
||||
|
||||
Schemas (`src/web/schemas.ts`), mirroring `RemoteHostSchema` (299) / `RemoteCaseLinkSchema` (351):
|
||||
|
||||
```ts
|
||||
export const DockerHostSchema = z.object({
|
||||
id: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid docker host id'),
|
||||
label: z.string().min(1).max(100),
|
||||
engine: z.enum(['docker', 'podman']).optional(),
|
||||
image: z.string().min(1).max(512).regex(/^[a-zA-Z0-9][\w./:@-]*$/, 'Invalid image ref').regex(NO_SHELL_META),
|
||||
daemonHost: z.string().max(512).regex(NO_SHELL_META, 'Invalid daemon host').optional(),
|
||||
context: z.string().max(128).regex(/^[a-zA-Z0-9._-]+$/, 'Invalid context').optional(),
|
||||
network: z.enum(['bridge', 'none', 'custom']).optional(),
|
||||
networkName: z.string().max(128).regex(/^[a-zA-Z0-9][a-zA-Z0-9_.-]+$/).optional(),
|
||||
resources: z.object({
|
||||
memory: z.string().regex(/^\d+[bkmg]?$/i).optional(),
|
||||
cpus: z.string().regex(/^\d+(\.\d+)?$/).optional(),
|
||||
pidsLimit: z.number().int().positive().max(100000).optional(),
|
||||
nofile: z.string().regex(/^\d+:\d+$/).optional(),
|
||||
shmSize: z.string().regex(/^\d+[bkmg]?$/i).optional(),
|
||||
}).strict().optional(),
|
||||
mountCredentials: z.boolean().optional(),
|
||||
hooksEnabled: z.boolean().optional(),
|
||||
resumeOnStart: z.boolean().optional(),
|
||||
commands: RemoteCommandOverridesSchema, // reuse the shared shape
|
||||
extraCreateArgs: z.array(z.string().min(1).max(1024).regex(NO_SHELL_INJECTION).refine(noCommandSubstitution)).max(32).optional(),
|
||||
extraExecArgs: z.array(z.string().min(1).max(1024).regex(NO_SHELL_INJECTION).refine(noCommandSubstitution)).max(32).optional(),
|
||||
});
|
||||
|
||||
export const DockerCaseLinkSchema = z.object({
|
||||
name: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid case name format'),
|
||||
hostId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid docker host id'),
|
||||
hostWorkspacePath: z.string().min(1).max(2000).regex(/^\//, 'Path must be absolute').regex(NO_SHELL_META, 'Invalid characters in workspace path'),
|
||||
containerWorkdir: z.string().min(1).max(2000).regex(/^\//).regex(NO_SHELL_META).optional(),
|
||||
container: z.string().min(2).max(128).regex(/^[a-zA-Z0-9][a-zA-Z0-9_.-]+$/, 'Invalid container name').optional(),
|
||||
});
|
||||
```
|
||||
|
||||
`NO_SHELL_META` (rejects `$`/backtick, schemas.ts:297) is REQUIRED on `image`, `hostWorkspacePath`, `containerWorkdir`, and `container`, because all four reach the outer `bash -c "..."` double-quote layer where `$(...)`/backtick re-expose, exactly the reason `remotePath`/`identityFile` use it. `--privileged` and any `-v /var/run/docker.sock` are structurally unrepresentable (never emitted by the builder, never accepted by the schema).
|
||||
|
||||
## 7. Security model
|
||||
|
||||
- Hardening flags on every create: `--cap-drop ALL`, `--security-opt no-new-privileges` (NOT auto-set by rootless Docker or Podman, so always explicit), the uid/userns branch of Key decision 6 (never container-root; workspace files stay host-owned and HOME stays writable via GID 0), `--pids-limit` (fork-bomb guard), `--memory` with `--memory-swap == --memory` (real OOM cap), `--ulimit nofile`, `--init`, `--pull=never`. NEVER `--privileged`, NEVER mount the docker socket into the agent container. `--storage-opt size=` is emitted ONLY after the probe confirms overlay2-on-xfs-pquota or btrfs (the AICE-class silently-ignored trap); otherwise it is omitted and the UI does not advertise a size cap. Resource caps are advertised as ENFORCED only when the probe reports `capsEnforced:true`; under non-delegated rootless they are labeled advisory (user-decision 6).
|
||||
- Engine: prefer whichever the probe finds, Podman-rootless first for security (a container-root breakout lands as an unprivileged host user). Rootless bind-mount ownership uses `--userns=keep-id` (Podman) vs `--user <hostUid>:0` (Docker), so real per-engine branching lives in `buildDockerCreateArgs`. Docker Desktop takes its own uid path (Key decision 6).
|
||||
- Blast radius (the combined-posture the critic asked to surface, user-decision 5): the default convenient profile mounts an arbitrary host workspace dir RW (host-owned, mirrored path) AND host `~/.claude`/`~/.codex`/`~/.gemini`/`~/.config/gcloud`/`~/.config/opencode` RW into a NETWORK-ENABLED container. Container-run agent code can therefore read/modify those host trees and reach the network simultaneously. This is still a strict improvement over today's on-host skip-permissions execution, but the user must accept the combined posture explicitly; the sealed profile plus `network:none` is the mitigation for genuinely untrusted work.
|
||||
- Secret handling: creds arrive ONLY as bind-mounted files (default) or exec-time NAME-ONLY `--env` (codex/gemini keys), NEVER as create-time `-e` and NEVER as an image layer. Sealed-mode export is refuse-or-scrub (Section 5), closing the sealed-leak inversion.
|
||||
- CLAUDE.md "Multi-CLI prefix discipline": the exec-time name-only env is restricted to the CLI-specific keys per mode (Claude: none with OAuth mount; Codex: `OPENAI_API_KEY`/`CODEX_API_KEY`; Gemini: `GEMINI_API_KEY`/`GOOGLE_*`), never a blanket forward. `envOverrides` is rejected for docker, so the `ALLOWED_ENV_PREFIXES` allowlist is not widened.
|
||||
- hook-secret: bind-mounted read-only, referenced via `CODEMAN_HOOK_SECRET_FILE` (a path, non-secret); the secret bytes never enter env or the image. Both `host.docker.internal` and `host.containers.internal` are added to the host-guard allowlist so the in-container hook curl's Host header passes on either engine.
|
||||
- Host guard / instance isolation: the in-container tmux socket (`codeman-docker`) and name (`codeman-dkr-<id8>`) deliberately FAIL a container-internal Codeman's `SAFE_MUX_NAME_PATTERN`, so a nested Codeman never adopts our session (unit-asserted). The boot reaper is instance-scoped by the `codeman.instance` label so a beta never reaps prod. Any remote-daemon (`-H`/`--context`) mode is host-root-equivalent and stays strictly behind the existing auth/loopback/host-guard/Origin-CSRF stack.
|
||||
- Import containment: untrusted bundles are checksum-validated, extracted with traversal guards, and loaded into a quarantined image namespace (never overwriting the base image), then run with the same hardening.
|
||||
|
||||
## 8. Phased implementation (branch: `feat/docker-session-mode`)
|
||||
|
||||
Each phase is independently testable; per CLAUDE.md, end-to-end test in the real env before COM. All new docker IO paths carry `const IS_TEST_MODE = !!process.env.VITEST;` and no-op under it; the pure command builders are tested directly.
|
||||
|
||||
- Phase 0: base image + engine probe. Author `docker/agent.Dockerfile` (OpenShift arbitrary-uid HOME) and `scripts/build-agent-image.mjs` (build or pull the base image; digest recorded). Add `checkDockerAvailable`/`checkDockerTmuxAvailable`/`containerApiUrl`/`hostGatewayAlias` (IS_TEST_MODE no-op) and `GET /api/docker/status`. Test: probe stub returns available/caps/Desktop flags under VITEST; `containerApiUrl` preserves scheme+port and swaps host per engine; status route returns the envelope.
|
||||
- Phase 1: types + storage + schemas. Add all types (Section 3), `src/docker-hosts.ts`, `DockerHostSchema`/`DockerCaseLinkSchema`. Test: `docker-hosts.test.ts` (round-trip incl. `lastClaudeSessionId`, display path, config-hash stability); `docker-exec-options.test.ts` (schema rejects `$`/backtick in image/workdir/container).
|
||||
- Phase 2: tmux-manager builders. Add `DOCKER_TMUX_SOCKET`, `dockerTmuxSessionName`, `buildDockerLaunchCommand` (resume-aware, image-check, env-prime), `buildDockerKillCommand`; wire the two ternaries + two cd-skips + Strategy 3c; harden `reconcileSessions` against docker hard-delete. Test (pure strings): adopt-proof name fails `SAFE_MUX_NAME_PATTERN`; image-check precedes create; `new-session -A` idempotent; resume flag present only when a resume id is passed; `--pull=never` present; instance label present; escaping survives `bash -c` -> `docker exec` -> `sh -lc` -> tmux WITH a host workspace path containing spaces.
|
||||
- Phase 3: session.ts + mux + recovery. Add `_docker` + `resumeSessionId` threading, in-container cliVersion probe, `resolveMuxAttachCwd`, mux-interface fields, `restoreMuxSessions` passthrough, instance-scoped reaper wiring, claudeSessionId -> `DockerCase.lastClaudeSessionId` persistence, unified flag. Test: `toState()` emits docker; a persisted docker session round-trips through mux/state; a relaunch injects the persisted resume id (mock mux); reaper only targets this instance's orphaned containers.
|
||||
- Phase 4: routes + first real e2e. case-routes CRUD + listing + drift-recreate; session-routes quick-start branch (scaffolding RUNS, local-availability guards skip, model accepted, effort/config rejected). Manual e2e on a real docker host: docker-host create -> docker-link -> quick-start; confirm the pane runs `claude` in the container, files land host-owned, a Codeman restart reattaches the SAME live agent, and a `docker stop` followed by relaunch RESUMES the conversation.
|
||||
- Phase 5: hooks connectivity + installation. host-gateway (per engine), derived `CODEMAN_API_URL`, hook-secret mount, `CODEMAN_SESSION_ID`/`CODEMAN_MUX` exec-env + tmux setenv, host-guard allowlist, and the scaffolding write into the real workspace. Manual e2e: trigger a permission prompt from inside the container and confirm it surfaces; verify hook payloads carry the right session id. If deferred, ship docker as explicitly hook-degraded and verify output-based idle detection through the docker-exec PTY.
|
||||
- Phase 6: export/import + GC + disk safety. quiesce+pause span, free-space precheck, commit+save+gzip + workspace tar + manifest + streaming download; sealed-mode refuse-or-scrub; retention/auto-prune; import with checksum validation + traversal guard + quarantined re-tag; drift-recreate; boot reaper; `runWithConversionLimit` cap; `docker rmi` in finally. Manual e2e: export, `docker load` on a second machine (or fresh case), import, confirm toolchain + workspace restored and NO creds present; attempt a sealed full-image export and confirm it is refused-or-scrubbed; attempt a `../` bundle and confirm it is rejected.
|
||||
- Phase 7: frontend. Docker tab, `linkDockerCase`, run wiring, case-picker labels, panels search, caps-advisory + scaffold-warning + effort-inert notes. Verify with Playwright (`waitUntil: 'domcontentloaded'`, 3-4s settle) that the Docker tab renders and a linked docker case appears in the picker.
|
||||
- Phase 8: docs + COM. Update CLAUDE.md (a "Docker cases" Key Pattern paragraph mirroring remote-SSH, plus the new state files, routes counts, and the resume/durability model), `docs/docker-cases.md`, then COM per the standard flow.
|
||||
|
||||
## 9. Test plan
|
||||
|
||||
- Unit (pure, CI-safe, mirror `test/remote-hosts.test.ts` / `test/remote-ssh-options.test.ts`):
|
||||
- `test/docker-hosts.test.ts`: storage round-trip (incl. `lastClaudeSessionId`), `dockerDisplayPath`, `defaultDockerCommandForMode`, `toSessionDocker`, `containerApiUrl` (http/https, custom port, docker vs podman gateway), config-hash stability/drift, `buildDockerCreateArgs` flag ordering (cap-drop/no-new-privileges/memory==memory-swap/instance-label/`--pull=never` present; host/privileged/socket absent; per-engine uid vs `--userns=keep-id`).
|
||||
- `test/docker-exec-options.test.ts`: `buildDockerLaunchCommand`/`buildDockerKillCommand` string shape and escaping through `bash -c` -> `docker exec` -> `sh -lc` -> tmux, including a workspace path with spaces; resume flag present only with a resume id; image-presence check precedes create; `dockerTmuxSessionName` fails `SAFE_MUX_NAME_PATTERN`; schema rejects `$`/backtick in image/workdir/container/name; `linkDockerCase`-shaped bodies with omitted optionals validate (no `null` on the wire).
|
||||
- Probe no-op: `checkDockerAvailable`/`checkDockerTmuxAvailable`/`probeDockerCliVersion` return canned values under VITEST and never spawn.
|
||||
- Integration (route tests via `app.inject()`, docker no-op'd): `/api/docker-hosts` CRUD; `/api/cases/docker-link` dup-check + broadcast; `GET /api/cases` includes the docker case with `location: 'docker'`; `/api/quick-start` docker branch rejects `envOverrides`/`effort`/config but ACCEPTS `modelOverride`, runs the workspace-scaffolding path, and constructs a session with `docker` set + seeded resume id; `DELETE /api/cases/:name` docker-unlink; export refuse-or-scrub for sealed; import traversal rejection; reaper instance-scoping (label filter). Pick a unique port only if a live-server test is added (search `const PORT =`; 3150+).
|
||||
- Manual end-to-end (real docker daemon, the mandatory "always end-to-end test" gate): build the base image; link a docker case; quick-start `claude`; verify OAuth via the mounted `~/.claude`, transcript correlation (subagent/workflow watchers show the session), host-owned files, and a working permission-prompt hook; reattach after a Codeman PROCESS restart (SAME live agent); `docker stop` then relaunch and confirm conversation RESUME; reboot-equivalent (daemon restart) and confirm boot recovery recreates+resumes; change the host's memory/image and confirm the drift-recreate prompt fires; export (convenient) and confirm the tar `docker load`s with no creds; attempt a sealed full-image export and confirm refuse-or-scrub; import into a fresh case; delete the case and confirm `docker rm -f` plus instance-scoped reaper GC; confirm a docker-down state surfaces a docker-specific error and does NOT trip the generic PTY-exit breaker.
|
||||
|
||||
## 10. Open decisions for the user
|
||||
|
||||
1. Credential + blast-radius posture (combined). Convenient default bind-mounts host `~/.claude` etc. RW AND an arbitrary host workspace RW into a network-enabled container, so container-run agent code can read/modify those host trees and reach the network at the same time. Recommended: convenient default plus a per-host SEALED opt-in (`mountCredentials:false` + `network:none`) for untrusted work. Please confirm you accept the combined arbitrary-workspace-plus-egress-plus-host-creds posture for the default profile (it is still a net improvement over today's on-host skip-permissions execution).
|
||||
2. Base image ownership, registry, and freshness. The `codeman/agent:base` placeholder implies a Docker Hub org the project may not own. Pick the real registry/namespace (GHCR under the repo is the natural fit), decide digest pinning, and set a REBUILD CADENCE so agents are not stuck on a stale baked `claude` (the in-container version probe surfaces staleness, but something must trigger rebuilds). Choose: pull a pinned published image, build locally on first use via `scripts/build-agent-image.mjs`, or both.
|
||||
3. Container CWD strategy. Mirror the host workspace path inside the container (recommended: makes transcript projHash correlate, file features and resume capture work) vs a fixed `/workspace` (simpler mount, breaks watcher correlation). Please confirm the mirror approach.
|
||||
4. Hooks in the MVP AND workspace scaffolding. Making docker hooks fire requires WRITING `.claude/settings.local.json` (and the CLAUDE.md scaffold) into the user's REAL linked host directory, a behavioral shift from "link a dir" to "link and scaffold a dir." Choose: wire hooks + scaffolding now (Phase 5, recommended, and it also enables the model picker), or ship docker as explicitly hook-degraded (no permission prompts / hook-idle) for v1 and add later. Confirm you are OK with Codeman mutating the linked host workspace.
|
||||
5. Session-kill teardown and RESUME (reframed honestly). `docker stop` on session kill is not merely "free RAM vs instant reattach": it destroys the in-container live agent, and the conversation survives ONLY because the next launch runs `--resume` from the bind-mounted transcript. Choose: keep the container running (costs RAM, preserves the exact live in-flight agent) vs stop and rely on `--resume` (frees RAM, may lose uncommitted in-flight tool state). Case-delete always `docker rm -f`.
|
||||
6. Rootless enforcement posture. Under rootless without cgroup-v2 systemd delegation, `--memory`/`--cpus`/`--pids-limit` are SILENTLY ignored. Choose: REQUIRE delegation (refuse to link a host that cannot enforce caps) or ship-with-warning ("resource caps are advisory on your engine"). The probe reports `capsEnforced` either way.
|
||||
7. Default resume behavior. Should a re-linked or re-run docker case default to resuming its last conversation (`resumeOnStart:true`, using `DockerCase.lastClaudeSessionId`) rather than starting clean? This is the crux of making the durability story real and is the recommended default, but it changes user-visible behavior (a new session in an existing case continues the prior conversation).
|
||||
8. Export defaults and disk budget. Default export button: workspace-only (fast, small, files-only, recommended for 24h+ runs) vs full-image (reproducible env, multi-GB). Also set the retention cap (max retained exports), the auto-prune policy, and the free-space threshold below which export is refused (a full `/var/lib/docker` breaks EVERY session on the host, not just docker ones).
|
||||
9. Remote docker daemon (`-H ssh://...` / `--context`). Support in the MVP (composes with remote hosts, adds host-root trust surface) or local-daemon-only first.
|
||||
10. Podman parity depth. Full `--userns=keep-id` plus Quadlet boot-persistence, or Docker-first with Podman as best-effort and boot-persistence via Codeman's idempotent create-if-missing only. Note the podman host alias is `host.containers.internal`, already handled per engine.
|
||||
@@ -1,206 +0,0 @@
|
||||
# Docker cases
|
||||
|
||||
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
|
||||
|
||||
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` / `pi` / `grok` / `deepseek` / `omp` all work inside the container.
|
||||
|
||||
## One-time setup: build the base image
|
||||
|
||||
The container needs a base image with the agent toolchain (node, the CLIs, git, tmux). Build it locally once:
|
||||
|
||||
```bash
|
||||
node scripts/build-agent-image.mjs # builds codeman/agent:base
|
||||
# options: --engine docker|podman --image <ref> --no-cache
|
||||
```
|
||||
|
||||
The image is **secret-free**: credentials are delivered at runtime (bind mounts or `docker exec --env`), never baked in, so exports never leak them.
|
||||
|
||||
⚠️ **Re-build with `--no-cache`, always.** The CLIs are installed in a single `RUN npm install -g` layer, so a plain rebuild re-uses it from the Docker layer cache and the CLIs stay frozen at whatever versions the image was **first** built with, however long ago that was. Editing the Dockerfile does not help unless the edit lands at or above that line: a change appended below it leaves the npm layer cached and only runs the new step. Observed 2026-08-06: a rebuild silently kept a stale `@openai/codex@0.144.6` whose aliased platform binary had not installed, so every `codex` docker case died with `Missing optional dependency @openai/codex-linux-x64` while the build itself reported success.
|
||||
|
||||
```bash
|
||||
node scripts/build-agent-image.mjs --no-cache
|
||||
```
|
||||
|
||||
### Which CLIs the image contains
|
||||
|
||||
The npm-published CLIs come from `ARG CLI_NPM_PACKAGES`, which `scripts/build-agent-image.mjs`
|
||||
fills from `config/clis.stock.json` (generated from `src/config/cli-registry/stock.ts`). Adding
|
||||
a stock CLI that installs with a plain `npm install -g` needs no Dockerfile edit. The ARG
|
||||
defaults to the same list in the same order, so a bare `docker build` produces a byte-identical
|
||||
layer — a different order would be a different `RUN` string and so a needless cache miss.
|
||||
|
||||
⚠️ It reads the **stock** catalogue, never the merged registry. A user's `~/.codeman/clis.json`
|
||||
must not change what is inside an image tagged `codeman/agent:base`, or two machines holding
|
||||
that tag hold different images and every cache decision downstream is a lie. Each entry's
|
||||
`enabled` flag IS honoured, so a CLI that ships disabled is never baked in.
|
||||
|
||||
Five CLIs keep hand-written layers, for two different reasons that are easy to conflate.
|
||||
`antigravity`, `grok` and `omp` declare no `npmPackage` at all, so they never enter the shared
|
||||
npm layer and each gets a vendor-installer layer instead. `pi` and `deepseek` ARE on npm but
|
||||
carry `discovery.install.agentImageLayer` in `stock.ts` (a REGISTRY field, rather than an
|
||||
id-keyed table duplicated between the two producers of the image's build args), which pulls
|
||||
them out of the shared layer because a plain `npm install -g` is not enough for them:
|
||||
|
||||
| CLI | Why it is not in the shared npm layer |
|
||||
| ------------- | ------------------------------------------------------------------------------------- |
|
||||
| `pi` | Installs with `--ignore-scripts`, kept in its own layer so the flag cannot leak to the others. |
|
||||
| `deepseek` | Needs `pnpm` alongside it (`dsh plugin`, issue #352) plus a `dsh-tui` profile install. |
|
||||
| `antigravity` | Not on npm — Google ships a standalone binary (~190MB, the largest layer). |
|
||||
| `grok`, `omp` | Not on npm — standalone vendor installers. |
|
||||
|
||||
`test/docker-agent-image-coverage.test.ts` requires every special case to carry a written
|
||||
reason AND still be present in the Dockerfile, so an exclusion cannot silently become an
|
||||
omission — which is the same failure upstream `b6d0f1fa` hit in `install.sh`.
|
||||
|
||||
Two things build this image: `scripts/build-agent-image.mjs` (a human) and
|
||||
`ensureAgentBaseImage()` in `src/docker-hosts.ts` (the app, on the first Docker case). They
|
||||
assemble the argv independently, because a `.mjs` cannot import TypeScript, so
|
||||
`test/agent-image-build-args-parity.test.ts` pins them together. Without it, an image built by
|
||||
hand and one built by the app could hold different CLIs under the same tag.
|
||||
|
||||
A zero exit code only proves the layers ran, not that the toolchain works. Verify by actually executing each CLI in the image, and check the build log for `Using cache` lines:
|
||||
|
||||
```bash
|
||||
docker run --rm codeman/agent:base bash -lc \
|
||||
'for c in claude codex gemini opencode agy pi grok dsh omp; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
|
||||
```
|
||||
|
||||
⚠️ `dsh --version` is the one line above that answers a different question than the
|
||||
others: `dsh` is a profile launcher, so a working binary says nothing about whether
|
||||
the image can actually run a DeepSeek session. Check the profile the Dockerfile
|
||||
installs into the agent's HOME as well, or a `mode: 'deepseek'` case starts a pane
|
||||
that dies on arrival:
|
||||
|
||||
```bash
|
||||
docker run --rm codeman/agent:base ls ~/.dsh/profiles/dsh-tui/package.json
|
||||
```
|
||||
|
||||
Building that profile is also why `pnpm` is in the image: `dsh plugin` forwards straight to a literal `pnpm` and exits 127 without it (issue #352), and pnpm — unlike npm — blocks dependency lifecycle scripts by default and fails the install over it, so the profile step passes `--config.dangerouslyAllowAllBuilds=true`.
|
||||
|
||||
Antigravity (`agy`) and Grok (`grok`) are the two CLIs not installed from npm (Google and xAI ship standalone binaries), so each has its own Dockerfile step, adding roughly 190MB and 160MB respectively. Pi also gets its own step, because upstream documents installing it with `--ignore-scripts` and that flag must not silently change how the other npm CLIs install.
|
||||
|
||||
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md). Grok is seeded per-file for the same reason (`auth.json`, `config.toml`, `pager.toml` out of `~/.grok`, which also holds `sessions/`, `memory/` and the ~160MB binary under `downloads/`), with the same consequence for `grok -c`. See [`grok-integration.md`](./grok-integration.md). OMP is the one CLI in this family where `sessions/` is the EXCEPTION rather than the rule: `~/.omp/agent/{config.yml,mcp.json,models.yml,settings.yml}` are seeded per-file (the dir also holds SQLite caches and `terminal-sessions/`), but `~/.omp/agent/sessions/` is shared RW like codex's, not seeded, because Codeman reads it host-side for history recovery and `--resume` pinning. See [`omp-integration.md`](./omp-integration.md).
|
||||
|
||||
The image can also carry the GitHub CLI (`gh`) and the Azure CLI (`az` + the `azure-devops` extension, in `AZURE_EXTENSION_DIR=/opt/az-extensions` so it stays out of the seeded HOME), wired into the system git config as credential helpers for github.com and dev.azure.com / *.visualstudio.com, exactly as in `docker/server.Dockerfile`. Their sign-ins are seeded per-FILE like pi's: `~/.config/gh/{hosts.yml,config.yml}` and `~/.azure/{azureProfile.json,msal_token_cache.json,service_principal_entries.json,clouds.config,config}`, never `~/.azure`'s logs, command index or extensions. A token kept in a desktop keyring, or in the encrypted MSAL cache az uses on Windows/macOS, is not in those files and does not carry. None of the three is version-pinned; the `--no-cache` rebuild recommended above is also what refreshes them. Both CLIs are opt-in and OFF by default: `CODEMAN_AGENT_IMAGE_INSTALL_GH=1` / `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` in the environment of `scripts/build-agent-image.mjs`, or of the Codeman server for its own auto-build (in the Compose deployment, `environment:` in `docker-compose.override.yml`), become the `CODEMAN_INSTALL_GH` / `CODEMAN_INSTALL_AZ` build args and put that CLI, its extension and its helper entry into the image. Unset passes nothing, so a default build's argv is unchanged and the image has neither. The sign-in seeds follow the same switches, read when a case container is created: `.config/gh` only with `CODEMAN_AGENT_IMAGE_INSTALL_GH=1`, `.azure` only with `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` (`enabledByEnv` in `CRED_STORES`), never merely because the files exist. Seeds are create-time mounts and deliberately not part of the config hash (hashing them would trip the drift gate for every case), so an existing case container picks them up only when it is recreated.
|
||||
|
||||
## Quickest path: one-click "Run in Docker"
|
||||
|
||||
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
|
||||
|
||||
Click the checkbox's **Container settings** to optionally tweak the predefined defaults, including a **Template** picker:
|
||||
|
||||
| Template | Memory | CPUs | GPUs |
|
||||
|----------|--------|------|------|
|
||||
| Small | 2 GB | 1 | none |
|
||||
| Medium (default) | 4 GB | 2 | none |
|
||||
| Large | 8 GB | 4 | none |
|
||||
| GPU | 8 GB | 4 | all (needs the NVIDIA container toolkit) |
|
||||
|
||||
**Disk is elastic** — the container's storage grows automatically as data flows in; there is no fixed cap (bounded only by host disk). Any tweaked setting creates a dedicated per-case host so it never changes the shared `default`.
|
||||
|
||||
## Create a docker case (full control)
|
||||
|
||||
App → **New case → Docker** tab:
|
||||
|
||||
- **Case Name** / **Workspace Path**: the workspace is a real HOST directory bind-mounted into the container at the same path. Codeman scaffolds `CLAUDE.md` + `.claude/settings.local.json` (hooks) into it, and file previews / attachments work on the real bytes.
|
||||
- **Host ID**: a reusable docker host profile (image, network, resources). Reuse the same ID across cases to share settings.
|
||||
- **Network**: `bridge` (internet on, default), `none` (fully isolated), or a `custom` bridge.
|
||||
- **Advanced**: memory / CPU caps, **Mount host credentials** (on = your existing `~/.claude` login just works; off = a sealed sandbox you log into inside the container), **Resume last conversation on relaunch**.
|
||||
|
||||
Then run it like any case (Run Claude / Run Shell / …). The first launch creates the container (`codeman-case-<name>`); subsequent sessions attach to the same one.
|
||||
|
||||
Equivalent API:
|
||||
|
||||
```bash
|
||||
curl -X POST localhost:3000/api/docker-hosts -d '{"id":"local","label":"Local","image":"codeman/agent:base"}'
|
||||
curl -X POST localhost:3000/api/cases/docker-link -d '{"name":"sandbox","hostId":"local","hostWorkspacePath":"/home/you/projects/sandbox"}'
|
||||
curl -X POST localhost:3000/api/quick-start -d '{"caseName":"sandbox","mode":"claude"}'
|
||||
```
|
||||
|
||||
## Attach to a container you already run
|
||||
|
||||
The tab's **Attach to an existing container** toggle points a case at a container **you**
|
||||
built and run. Codeman only ever `docker exec`s into it: it never creates, starts, stops,
|
||||
restarts or removes it, and it seeds no credentials into it, so the CLIs inside must already
|
||||
be installed and logged in. A missing or stopped container is an error to report, not a state
|
||||
to fix — start it yourself and reopen the session.
|
||||
|
||||
- **Container Name** is a picker over the engine's containers that you can also type into
|
||||
(the engine may be remote, or the container may not exist yet when you fill the form).
|
||||
Stopped containers are listed too, sorted last and labelled, so "mine isn't here" is never
|
||||
a dead end.
|
||||
- **Container Workdir** is a path that must already exist **inside** the container. Adoption
|
||||
mounts nothing, so it need not match the host workspace path; **Browse** lists directories
|
||||
inside the container itself. Without this check, a wrong path fails at launch as a bare
|
||||
`execvp failed` inside the pane.
|
||||
- **Workspace Path** is still a real host directory. It backs file previews, attachments and
|
||||
watchers exactly as it does for an owned case, but here it is only a mirror: nothing is
|
||||
bind-mounted, so point it at whatever host directory your container already exposes.
|
||||
- **Check container** runs a read-only preflight and reports what is inside before you commit
|
||||
to a case name (running or not, tmux present, which CLIs resolved).
|
||||
- **Run modes come from the container**, not the host: a host with no `claude` still offers
|
||||
Claude if the container ships it, and a mode the container lacks is hidden.
|
||||
- Claude is launched **without** `--dangerously-skip-permissions` when the container's exec
|
||||
user is root, because Claude Code refuses that flag as root and the refusal is only visible
|
||||
inside the container.
|
||||
- Image, network and resource settings disappear from the form: they describe a
|
||||
`docker create` that adoption never runs.
|
||||
|
||||
Recreate is refused for an adopted case, full-image export is refused (it would commit a
|
||||
container that is not ours), unlinking the case leaves the container running, and the boot
|
||||
reaper skips it. Workspace-only export still works and never pauses the container.
|
||||
|
||||
Equivalent API:
|
||||
|
||||
```bash
|
||||
curl -X POST localhost:3000/api/docker-cases/adopt-preflight -d '{"hostId":"local","container":"my-dev-box","containerWorkdir":"/workspace"}'
|
||||
curl -X POST localhost:3000/api/cases/docker-adopt -d '{"name":"devbox","hostId":"local","container":"my-dev-box","hostWorkspacePath":"/home/you/projects/devbox","containerWorkdir":"/workspace"}'
|
||||
```
|
||||
|
||||
In multi-user mode adoption is **admin-only**, unlike `docker-link`: an adopted container's
|
||||
mounts belong to whoever built it, so one mounting `/` would hand the adopter the whole host.
|
||||
|
||||
## Lifecycle
|
||||
|
||||
- **Reconnect after a Codeman restart** lands back in the same live agent (the in-container tmux survives).
|
||||
- **Container stop / host reboot** restarts the container and **resumes** the last conversation from the bind-mounted transcript. Claude sessions launch with a pinned conversation id (`--session-id <sessionId>`, with a `--resume` fallback when the transcript already exists), and the case remembers its last conversation (`lastClaudeSessionId`), so a relaunch after the container was stopped, rebooted, or recreated continues where it left off.
|
||||
- **Killing one session** only kills that session's in-container tmux session; the shared container stays up for sibling sessions.
|
||||
- **Editing the docker host config** (image, memory, network, ...) is detected on the next launch: the desired config hash is compared against the container's `codeman.confighash` label, and a mismatch refuses the launch with a "config changed, recreate?" confirm. Confirming calls `POST /api/docker-cases/:name/recreate` (refused while sessions of the case are live), which removes the container so the next launch recreates it with the new config; the workspace and the conversation survive.
|
||||
- **Deleting the case** `docker rm -f`s the container (the bind-mounted workspace on the host survives). An instance-scoped boot reaper removes containers whose case is gone.
|
||||
|
||||
## Isolation & security
|
||||
|
||||
Every container runs hardened: `--cap-drop ALL`, `--security-opt no-new-privileges`, non-root (`--user <hostUid>:0` so workspace files stay host-owned), `--pids-limit`, `--memory` == `--memory-swap`, `--init`. Never `--privileged`, never the docker socket. The default **convenient** profile bind-mounts host credential dirs read-write so the common login just works (creds stay on the host, never captured by `docker commit`); the **sealed** profile (`mountCredentials:false` + `network:none`) is the opt-in for genuinely untrusted work.
|
||||
|
||||
Rootless engines without cgroup-v2 systemd delegation cannot enforce resource caps; linking such a host warns that caps are advisory.
|
||||
|
||||
## Export / Import (move to another machine)
|
||||
|
||||
**Export** (from the Docker tab, or `POST /api/docker-cases/:name/export`): choose
|
||||
|
||||
- **Full image + workspace**: `docker commit` the container to an image, `docker save` it, tar the workspace, and a manifest, all into one portable `<case>-<ts>.codeman-container.tgz` (the whole toolchain, installed packages, and files). Runs in the background; you are notified when the bundle is ready.
|
||||
- **Workspace only**: just the project files (fast, small).
|
||||
|
||||
The container is paused across the capture so the image and workspace are consistent; a full `/var/lib/docker` is guarded against with a free-space precheck; the intermediate image is always cleaned up.
|
||||
|
||||
**Import** (`POST /api/docker-cases/import`, or the Manage tab): copy the `.tgz` onto the new machine's `~/.codeman/docker-exports/`, then import it into a new case. The manifest and per-member SHA-256 checksums are validated, the workspace tar is extracted with a path-traversal guard, and the image is `docker load`ed and **re-tagged into a quarantined namespace** (`codeman/imported-<case>:<ts>`) so it never overwrites a local tag. The destination supplies its own credentials, so nothing secret crosses machines.
|
||||
|
||||
`GET /api/docker-exports` lists bundles; `GET /api/docker-exports/:filename` downloads one; `DELETE` removes one.
|
||||
|
||||
## Hooks require the server to be reachable from the container
|
||||
|
||||
In-container hooks (permission events, hook-based idle/stop/task notifications) POST to `CODEMAN_API_URL`, which is derived as `https://host.docker.internal:<port>` (`host.docker.internal` → the docker bridge gateway, e.g. `172.17.0.1`, via `--add-host …:host-gateway`). For that callback to succeed, the Codeman server must be **listening on an interface the container can reach**.
|
||||
|
||||
- If Codeman binds **loopback-only** (`127.0.0.1`, the default and the production systemd config), a container reaching `172.17.0.1:<port>` cannot connect, so by default **in-container hooks do not fire**. The session still works fully: idle/stop detection falls back to **output-based** detection through the `docker exec` PTY (which always works), and claude runs with `--dangerously-skip-permissions` so there are no permission prompts to forward anyway.
|
||||
- **To enable in-container hooks on a loopback-only server, set `CODEMAN_DOCKER_BRIDGE_HOOKS=1`** (env). Codeman then starts a SECOND listener bound to the docker bridge gateway (`172.17.0.1`, auto-detected; override with `CODEMAN_DOCKER_BRIDGE_HOST`) that serves **only the hook endpoints** (`/api/hook-event`, `/api/status-telemetry`) and delegates them into the same secret-gated pipeline. The bridge is host-internal (containers + host, not the LAN), and every other path returns `403`, so this does not widen your network exposure. Add `Environment=CODEMAN_DOCKER_BRIDGE_HOOKS=1` to the systemd unit and restart.
|
||||
- Alternatively, bind `0.0.0.0` **with `CODEMAN_PASSWORD` set** (exposes on the LAN too).
|
||||
|
||||
The host-gateway mapping, `CODEMAN_API_URL` derivation, host-guard allowlist, and hook-secret mount are all wired correctly; `CODEMAN_DOCKER_BRIDGE_HOOKS` closes the last gap for loopback-only servers.
|
||||
|
||||
## Notes & limits
|
||||
|
||||
- Requires Docker (or Podman) with a reachable daemon; tmux must be present in the base image (a hard prerequisite, probed at link time).
|
||||
- Per-session `envOverrides` / `effort` / per-CLI config are rejected for docker cases (they do not cross into the container); configure the container via the docker host's per-mode command override instead.
|
||||
- macOS Docker Desktop takes a dedicated uid path (the baked image uid; memory caps are subject to the VM ceiling).
|
||||
|
||||
Design + rationale: [`docker-cases-plan.md`](./docker-cases-plan.md).
|
||||
@@ -1,82 +0,0 @@
|
||||
# Docker Compose deployment
|
||||
|
||||
This configuration builds the Codeman application image locally from this checkout. It does not download or depend on a pre-built Codeman image.
|
||||
|
||||
For the Compose configuration, environment settings, storage migration, and macvlan networking examples, see the [Docker deployment guide](../docker/README.md).
|
||||
|
||||
The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image.
|
||||
|
||||
It can also include the GitHub CLI (`gh`) and the Azure CLI (`az`) with the `azure-devops` extension, wired in as Git credential helpers, so Clone Repo and `git clone` reach private GitHub and Azure DevOps repositories once they are signed in. Both are off by default; [Turning them on](../docker/README.md#turning-them-on) shows the `docker-compose.override.yml` settings.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Docker Engine or Docker Desktop with Docker Compose v2
|
||||
- A reachable Docker daemon
|
||||
|
||||
The application container mounts the Docker daemon socket so Codeman can create and manage its isolated Docker cases. Treat anyone who can administer this Compose project as having Docker-host-equivalent access.
|
||||
|
||||
## Start
|
||||
|
||||
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
|
||||
|
||||
```sh
|
||||
cp docker/.env.example docker/.env
|
||||
```
|
||||
|
||||
On PowerShell, use the following command instead.
|
||||
|
||||
```powershell
|
||||
Copy-Item docker/.env.example docker/.env
|
||||
```
|
||||
|
||||
On Linux, run the stack with the start script. It determines `PUID` and `PGID` from the owner of `CODEMAN_APPDATA_PATH`, and `DOCKER_SOCKET_GID` from the configured Docker socket, before invoking Compose. A root-owned application-data directory is rejected so the runtime account cannot become UID 0.
|
||||
|
||||
```sh
|
||||
bash docker/Start-Codeman.sh
|
||||
```
|
||||
|
||||
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner. Naming the file with `-f` disables Compose's own discovery of `docker/docker-compose.override.yml`, so add a second `-f` for it when you keep one (see `docker/README.md`, Local customisation).
|
||||
|
||||
```sh
|
||||
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
|
||||
```
|
||||
|
||||
The container starts as root, corrects the ownership of a bind source the daemon had to create, and drops to `PUID:PGID` with `setpriv` before Codeman starts; the capabilities that needs are declared in `docker/docker-compose.yaml` and named by the entrypoint when a compose file written elsewhere lacks them.
|
||||
|
||||
Open `http://localhost:3000` and sign in with the username and password from `docker/.env`.
|
||||
|
||||
## Operations
|
||||
|
||||
The local image is tagged `codeman:local` by default. Change `CODEMAN_IMAGE` in `docker/.env` if a different local tag suits your environment.
|
||||
|
||||
```sh
|
||||
docker compose --env-file docker/.env -f docker/docker-compose.yaml logs -f codeman
|
||||
bash docker/Start-Codeman.sh
|
||||
docker compose --env-file docker/.env -f docker/docker-compose.yaml down
|
||||
```
|
||||
|
||||
`CODEMAN_APPDATA_PATH` holds Codeman state and survives container recreation. Remove that host directory only when deliberately resetting the installation.
|
||||
|
||||
`CODEMAN_CASES_PATH` must be an absolute path on the Docker host. Compose mounts it at the same path inside Codeman, so the host daemon can bind the managed workspace into isolated Docker cases. Do not set it to `/home/${CODEMAN_RUNTIME_USER}/codeman-cases`.
|
||||
|
||||
Compose passes `CODEMAN_APPDATA_PATH` into Codeman as `CODEMAN_DOCKER_HOST_HOME`. Codeman uses that value to translate generated Docker seed, credential and hook-secret bind sources from the container's home path into paths visible to the host Docker daemon.
|
||||
|
||||
If `docker info` reports `SwapLimit=false`, set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1`. Isolated cases retain their configured memory limit. Codeman omits the unsupported swap-limit option and filters only the daemon's exact swap-capability warning while retaining every other Docker create error.
|
||||
|
||||
If that directory was created by an earlier root-running image, change its ownership to the configured `PUID:PGID` before starting this version. This preserves existing CLI credentials and session state while allowing the unprivileged runtime account to use them.
|
||||
|
||||
## Updating
|
||||
|
||||
Codeman updates itself from **App Settings → Updates**, as it does on a bare host. The checkout mounted at `/opt/codeman` is the same directory Compose builds from, so the update's `git checkout` and rebuild land on the host and survive container recreation; the restart is the server exiting, which `restart: unless-stopped` turns into a relaunch on the new build.
|
||||
|
||||
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those. For a major update, or a base-image change `Start-Codeman.sh` does not fully pick up, `docker/Update-Codeman.sh` rebuilds with no layer cache and clears the two build-artefact volumes before handing off to it (see "Major updates" in `docker/README.md`).
|
||||
|
||||
`CODEMAN_REPO_PATH` overrides which checkout is mounted. It defaults to the compose project's parent directory, so it normally needs no setting. Point it at a directory that is not a git checkout and in-app updates are reported as unavailable.
|
||||
|
||||
Full detail, including the fingerprint baseline and the troubleshooting table: [`docker-self-update.md`](docker-self-update.md).
|
||||
|
||||
## Docker cases
|
||||
|
||||
The default socket path is `/var/run/docker.sock`, which works with a standard Linux Docker Engine. The Bash start script detects its numeric group ID. When running Compose directly, set `DOCKER_SOCKET_GID`, for example using `stat -c '%g' /var/run/docker.sock`, so the unprivileged `CODEMAN_RUNTIME_USER` account can create Docker cases. Docker Desktop users should set `DOCKER_SOCKET` in `docker/.env` only when their Docker installation exposes a different compatible socket path.
|
||||
|
||||
Codeman Docker cases are sibling containers on the host daemon, not children of the application container. The Compose configuration handles their workspace bind mount through `CODEMAN_CASES_PATH`; the `/home/${CODEMAN_RUNTIME_USER}` application-data mapping is for Codeman state and ordinary in-container sessions, not sibling-case workspaces.
|
||||
@@ -1,239 +0,0 @@
|
||||
# Self-update in the Docker Compose deployment
|
||||
|
||||
Codeman running as a container updates itself from **App Settings → Updates**, the
|
||||
same place and the same button as a bare-host install. This document explains how
|
||||
that works, what it deliberately refuses to do, and how to recover when it stops.
|
||||
|
||||
The bare-host updater is documented in
|
||||
[`architecture-invariants.md#self-update`](architecture-invariants.md#self-update);
|
||||
this file covers only what the container changes.
|
||||
|
||||
## The short version
|
||||
|
||||
| Change in the release | Applied by |
|
||||
| ----------------------------------------- | ------------------------------------------------------------------------------- |
|
||||
| Application code | The in-app updater |
|
||||
| `docker/server.Dockerfile` | `docker/Start-Codeman.sh` on the host |
|
||||
| `docker/docker-compose.yaml` | `docker/Start-Codeman.sh` on the host |
|
||||
| New key in `docker/.env.example` | Add it to `docker/.env`, then `Start-Codeman.sh` |
|
||||
| A major update, or a Node base-image bump | `docker/Update-Codeman.sh` on the host (no-cache rebuild + fresh build volumes) |
|
||||
|
||||
The in-app updater detects the three middle rows itself and refuses with a
|
||||
message naming what changed, so you never have to work out which case you are in.
|
||||
`Update-Codeman.sh` is the heavier option for when `Start-Codeman.sh` is not
|
||||
enough: see "Major updates" in `docker/README.md`.
|
||||
|
||||
## Why the container needs its own path
|
||||
|
||||
The bare-host updater does `git checkout <tag> && npm install && npm run build`,
|
||||
then asks systemd or launchd to restart the service. Two of those assumptions are
|
||||
false in a container:
|
||||
|
||||
1. **There is no init system.** A container's supervisor is the Docker daemon,
|
||||
which acts on the container, not on processes inside it.
|
||||
2. **The image is immutable.** A `git pull` into the image's baked `/opt/codeman`
|
||||
would land in the container's writable layer, survive `docker restart`, and be
|
||||
silently discarded by the next `docker compose up`.
|
||||
|
||||
Both are solved by configuration rather than by a second updater:
|
||||
|
||||
- **The checkout is a host bind mount.** `docker-compose.yaml` mounts the repo
|
||||
(the same directory used as the build context) over `/opt/codeman`, so the
|
||||
updater's `git checkout` writes to the host filesystem and survives the
|
||||
container being recreated.
|
||||
- **The restart is the server exiting.** `restart: unless-stopped` relaunches the
|
||||
container whenever its main process ends, including on a clean exit — so the
|
||||
updater's final step is to signal the server, and Docker starts it again on the
|
||||
freshly built `dist/`.
|
||||
|
||||
Everything else — the release-tag channel, the auto-stash, the atomic
|
||||
`update-status.json` the browser polls across the connection drop, the boot-time
|
||||
reconcile that flips `restarting` to `completed` — is the existing machinery,
|
||||
unchanged. The container path is a new `SupervisorKind`, not a new updater.
|
||||
|
||||
## What the pieces are
|
||||
|
||||
| Piece | Role |
|
||||
| ---------------------------------------------- | ------------------------------------------------------------------- |
|
||||
| Repo bind mount at `/opt/codeman` | Makes the pull persistent. Without it, self-update is unavailable. |
|
||||
| `codeman-node-modules`, `codeman-dist` volumes | Container-owned build artefacts, layered over the bind mount. |
|
||||
| `CODEMAN_IN_CONTAINER=1` | Tells `detectSupervisor()` to restart by exiting. |
|
||||
| `restart: unless-stopped` | Turns that exit into a restart. Verified before every update. |
|
||||
| `CODEMAN_RESTART_BY_EXIT=1` | The Compose file's declaration of that policy, so the updater may exit even with no Docker socket. |
|
||||
| Toolchain + devDependencies in the image | Lets `npm install` and `npm run build` run inside the container. |
|
||||
| `docker-env-applied.json` | Fingerprint baseline, written by `Start-Codeman.sh` on every start. |
|
||||
| `docker-build-source.json` | What HEAD/`package-lock.json` the build artefact volumes currently reflect. Written by both `Start-Codeman.sh` and this in-place update, so the two agree on whether those volumes are stale. |
|
||||
|
||||
### Why build artefacts are in named volumes
|
||||
|
||||
`node_modules` and `dist` are mounted as named volumes **on top of** the repo bind
|
||||
mount. Without that, an update's `npm install` would write into the host checkout,
|
||||
leaving container-compiled native modules (node-pty builds from source here) in a
|
||||
directory that may also be used to run Codeman natively, and leaving `git status`
|
||||
permanently noisy.
|
||||
|
||||
Docker seeds an empty named volume from the image, so the first start inherits the
|
||||
image's already-built `node_modules` and `dist` and pays no bootstrap cost.
|
||||
`docker compose down -v` is the supported reset: the next start re-seeds them.
|
||||
|
||||
That seeding-only-while-empty behaviour has a second, less obvious edge: it also
|
||||
means a plain `docker compose build` triggered from OUTSIDE the container (for
|
||||
example `Start-Codeman.sh`, after a `git pull` done by hand rather than through
|
||||
this in-app updater) produces a fresh image whose freshly-built `dist`/
|
||||
`node_modules` then sit unused behind the volumes' OLD content — the container
|
||||
comes back up looking unchanged. `Start-Codeman.sh` detects this by comparing the
|
||||
checkout's current HEAD and `package-lock.json` hash against `docker-build-source.json`,
|
||||
and clears just the affected volume(s) before its own `--build` if they moved.
|
||||
This in-place update writes that same file after a successful build precisely so
|
||||
that comparison does not fire on stale information: without it, the next plain
|
||||
`Start-Codeman.sh` run would see the HEAD this update just checked out, not
|
||||
recognise it as already accounted for, and wipe the volumes this update just
|
||||
correctly rebuilt right back to the OLDER image.
|
||||
|
||||
### Why the runtime image carries a build toolchain
|
||||
|
||||
`npm run build` is `tsc` plus `esbuild`, both devDependencies, so the image no
|
||||
longer runs `npm prune --omit=dev`. And `npm install` may rebuild node-pty, which
|
||||
ships no Linux prebuild, so `python3`, `make` and `g++` are installed as well.
|
||||
|
||||
This is the real cost of in-place updates: a noticeably larger image than a
|
||||
runtime-only one. It buys an update that takes about a minute instead of a full
|
||||
image rebuild, and it is why `NODE_ENV=production` is paired with an explicit
|
||||
`npm install --include=dev` in the updater.
|
||||
|
||||
## The environment gate
|
||||
|
||||
An in-place update applies **code only**. A restarted container reuses its existing
|
||||
image and configuration, so a release that changes the environment cannot take
|
||||
effect that way — and would half-apply: new code against an old environment. The
|
||||
updater therefore checks the **target release's own files**, read straight out of
|
||||
git with `git show <tag>:<path>` before anything is checked out.
|
||||
|
||||
### 1. `server.Dockerfile` changed, so the image must be rebuilt
|
||||
|
||||
Compared by sha256 against the fingerprint `Start-Codeman.sh` recorded when the
|
||||
running container was built.
|
||||
|
||||
### 2. `docker-compose.yaml` changed, so the container must be recreated
|
||||
|
||||
Same mechanism. A restart cannot pick up a new mount, port or environment
|
||||
variable; only recreating the container can.
|
||||
|
||||
### 3. `.env.example` gained keys your `.env` has no value for
|
||||
|
||||
The check that matters most, because **Compose will not tell you**. An unset
|
||||
`${VAR}` interpolates to the empty string; Compose prints a warning to a terminal
|
||||
nobody is watching and starts anyway. A new required setting therefore arrives as
|
||||
a silently blank environment variable and misbehaves later, far from the cause.
|
||||
The updater names the missing keys instead.
|
||||
|
||||
Commented-out lines in `.env.example` are deliberately *not* keys — that is how
|
||||
the file marks optional overrides such as `# PUID=1000`, and counting them would
|
||||
block updates on settings you are meant to leave alone.
|
||||
|
||||
### 4. A restart policy that would not bring the container back
|
||||
|
||||
Before signalling the server, the updater asks the Docker daemon for its own
|
||||
container's restart policy. If it is `no`, the update is refused: applying it
|
||||
would take Codeman down and leave no UI to recover from.
|
||||
|
||||
If the policy cannot be read at all (no Docker socket mounted) the update is
|
||||
still allowed, but the final step changes: the server exits only when the
|
||||
Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (the shipped one does, because
|
||||
it is the file that sets `restart: unless-stopped`) or the daemon confirmed an
|
||||
auto-restart policy. Otherwise the build completes and the panel asks you to
|
||||
restart the container by hand. A container started by plain `docker run` with no
|
||||
restart policy therefore gets a staged update, never an outage.
|
||||
|
||||
### What the gate deliberately does not do
|
||||
|
||||
Every unknown fails **open**:
|
||||
|
||||
- A missing fingerprint baseline (a container started before this feature existed)
|
||||
is not treated as a change, or those installs could never update at all.
|
||||
- An unreadable `.env`, an unreachable Docker socket, or a target tag whose files
|
||||
cannot be read all yield "no blocker" rather than a refusal.
|
||||
|
||||
The one place an unknown does NOT fail open is the kill itself: with neither the
|
||||
Compose declaration nor a daemon answer, the updater stages the build and asks
|
||||
for a manual restart rather than exiting a server nothing may bring back.
|
||||
|
||||
The gate catches a specific, detectable class of mistake; it is not a last line of
|
||||
defence. It is also re-evaluated server-side on `POST /api/system/update`, so
|
||||
hiding the button in the UI is a courtesy rather than the control.
|
||||
|
||||
## The one residual risk
|
||||
|
||||
The gate is derived from the diff, so it cannot see a release that needs a newer
|
||||
environment **without changing any of those files** — for example, code that
|
||||
depends on newer agent-CLI behaviour.
|
||||
|
||||
That is why the four global CLIs in `server.Dockerfile` are **pinned**. Unpinned,
|
||||
the versions a user ends up with are a function of when their image was built
|
||||
rather than of any commit, and in-app updates make rebuilds rarer, which makes
|
||||
that drift worse over time. Pinned, "this release needs a newer CLI" becomes a
|
||||
Dockerfile change, which check 1 already detects. Bump them deliberately, as part
|
||||
of a release.
|
||||
|
||||
The complementary merge-side guard is `test/docker-compose-env-parity.test.ts`,
|
||||
which fails CI when a variable is added to `docker-compose.yaml` without an entry
|
||||
in `.env.example`, or the reverse.
|
||||
|
||||
## Sequence of an in-place update
|
||||
|
||||
1. **Check** — `GET /api/system/update/check` finds the latest release tag, fetches
|
||||
that one ref so the gate can read the target's files, and returns any blockers.
|
||||
2. **Start** — `POST /api/system/update` re-evaluates the gate, writes `queued` to
|
||||
`update-status.json`, stages `self-update.sh` outside the repo and runs it.
|
||||
3. **Apply** — stash if dirty, fetch the tag, check it out, `npm install
|
||||
--include=dev`, `npm run build`. A failure at any step rolls back to the
|
||||
previous commit, rebuilds it and reports `failed`; the server is never
|
||||
restarted into a broken build.
|
||||
4. **Restart** — write the terminal `restarting` marker, then signal the server.
|
||||
The container exits and Docker restarts it.
|
||||
5. **Reconcile** — the rebooted server compares its own version against the target
|
||||
and flips the status to `completed` or `failed`. The browser, still polling,
|
||||
picks that up.
|
||||
|
||||
Step 4 kills the updater script along with the container — unlike the systemd
|
||||
path, it does not outlive the restart. That is safe only because the terminal
|
||||
marker is written first, which is why nothing may be appended after the kill.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**"This install can't update itself (unknown)"** — the repo bind mount is missing,
|
||||
so the container is running the baked image copy. Check `CODEMAN_REPO_PATH` and
|
||||
confirm the mounted directory really contains `.git`.
|
||||
|
||||
**The update fails immediately with a git ownership or permission error** — the
|
||||
mounted checkout belongs to a different user than the one Codeman runs as
|
||||
(`PUID`), so git refuses it as "dubious ownership". `Start-Codeman.sh` warns
|
||||
about this at start; fix it by chowning the checkout to the same account that
|
||||
owns `CODEMAN_APPDATA_PATH`.
|
||||
|
||||
**A rebuild is reported as required every time** — the fingerprint baseline does
|
||||
not match the checkout. `Start-Codeman.sh` writes it on every start, so start
|
||||
through that script rather than a bare `docker compose up` after either file
|
||||
changes.
|
||||
|
||||
**Codeman does not come back after an update** — the build succeeded, since the
|
||||
updater gates the restart on it, so read the container logs with `docker compose
|
||||
logs codeman`. To roll back, check out the previous tag in the host checkout and
|
||||
run `docker/Start-Codeman.sh`.
|
||||
|
||||
**The update failed during `npm install`** — most likely a native rebuild with no
|
||||
toolchain, meaning the image predates the toolchain being added. Rebuild once from
|
||||
the host and the in-app path works from then on.
|
||||
|
||||
**Resetting the build artefacts** — `docker compose down -v`, then
|
||||
`Start-Codeman.sh`. This discards the named volumes and re-seeds them from a fresh
|
||||
image. `docker/Update-Codeman.sh` scripts the same reset by default for the two
|
||||
build-artefact volumes (`codeman-node-modules`, `codeman-dist`) only, plus an
|
||||
unconditional `--no-cache` rebuild, which a plain `Start-Codeman.sh` run does not
|
||||
force on its own. See "Major updates" in `docker/README.md`.
|
||||
|
||||
## Disabling it
|
||||
|
||||
Set `CODEMAN_DISABLE_SELF_UPDATE=1` in `docker/.env` and pass it through in the
|
||||
compose file's `environment:` block. The Updates panel then reports that in-app
|
||||
updates are disabled, and the host-side script is the only way to update.
|
||||
@@ -1,437 +0,0 @@
|
||||
# Extending Codeman
|
||||
|
||||
Codeman has no plugin runtime, and that is a deliberate choice rather than a
|
||||
missing feature. A plugin runtime means running third-party code inside a process
|
||||
that spawns agents with your credentials, on a server people routinely expose
|
||||
over a tunnel or Tailscale. Codeman's security model is one of its reasons to
|
||||
exist, so it does not hand that away for an extension mechanism.
|
||||
|
||||
Instead there are four seams that already work, from any language, with nothing
|
||||
installed:
|
||||
|
||||
| You want to | Use | Runs where |
|
||||
| --- | --- | --- |
|
||||
| Show your own UI inside Codeman | [Web tabs](#seam-1-web-tabs) | Your own process, rendered as a tab |
|
||||
| React when an agent needs you | [SSE events](#seam-2-sse-events) | Anywhere that can hold an HTTP connection |
|
||||
| Drive Codeman from a script | [HTTP API](#seam-3-http-api-and-cli) or the `codeman` CLI | Anywhere |
|
||||
| React inside a Claude session | [Hooks](#seam-4-hooks) | The agent's own machine |
|
||||
|
||||
Everything below is covered by the stability promise in
|
||||
[`versioning-policy.md`](versioning-policy.md): endpoint paths, the response
|
||||
envelope, `errorCode` values, and SSE event names are stable. Additive changes
|
||||
(new endpoints, new optional fields, new events) are non-breaking. Breaking
|
||||
changes ship under a new prefix (`/api/v2`).
|
||||
|
||||
## Before you start
|
||||
|
||||
**Base URL.** `http://127.0.0.1:3000` by default. Prefer the versioned prefix
|
||||
`/api/v1/...` for anything you publish; the unversioned `/api/...` is an alias.
|
||||
|
||||
**Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic on every request, or
|
||||
authenticate once and keep the `codeman_session` cookie. With no password set,
|
||||
Codeman is loopback-only and unauthenticated.
|
||||
|
||||
```bash
|
||||
curl -u admin:$CODEMAN_PASSWORD http://127.0.0.1:3000/api/v1/sessions
|
||||
```
|
||||
|
||||
**Envelope.** Every response is `{"success": true, "data": ...}` or
|
||||
`{"success": false, "error": "...", "errorCode": "..."}`. Check the HTTP status
|
||||
or `body.success`, then read `body.data`. The full `errorCode` to status mapping
|
||||
is in [`api-reference.md`](api-reference.md).
|
||||
|
||||
⚠️ A few legacy GETs (`/api/away-digest` among them) return a bare-ish body with
|
||||
the payload at the top level rather than under `data`. Read defensively with
|
||||
`body.data ?? body`.
|
||||
|
||||
⚠️ A `401` is not an envelope at all: auth is rejected in a request hook that
|
||||
replies with the bare string `Unauthorized`, so parsing it as JSON throws. Branch on
|
||||
the status code before you parse, or a missing password looks like a broken endpoint.
|
||||
|
||||
**Already driving Codeman from an agent?** The README's
|
||||
[Programmatic Guide](../README.md#driving-codeman-from-an-agent--programmatic-guide)
|
||||
covers the in-session case: the `CODEMAN_MUX`, `CODEMAN_API_URL`,
|
||||
`CODEMAN_SESSION_ID` and `CODEMAN_HOOK_SECRET_FILE` variables that let a CLI
|
||||
running inside Codeman find the API and avoid acting on itself. This page is for
|
||||
code running *outside* a session.
|
||||
|
||||
## Seam 1: Web tabs
|
||||
|
||||
The highest-leverage seam. Any web app you can serve locally becomes a tab beside
|
||||
your agent sessions. You write a normal web page; Codeman handles embedding it.
|
||||
|
||||
```bash
|
||||
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/webviews \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"name":"My Dashboard","url":"http://127.0.0.1:8787","icon":"📊"}'
|
||||
```
|
||||
|
||||
Fields: `name` (1 to 60 chars), `url`, and optionally `icon` (a single glyph, max
|
||||
8 code units), `embedMode` (`proxy` by default, or `direct`), and `trusted`.
|
||||
|
||||
Related endpoints: `GET /api/v1/webviews`, `PATCH /api/v1/webviews/:id`,
|
||||
`DELETE /api/v1/webviews/:id`, `POST /api/v1/webviews/probe` (reachability and
|
||||
framing check), `POST /api/v1/webviews/:id/open`.
|
||||
|
||||
### Why it is proxied
|
||||
|
||||
By default your page is served through Codeman's own origin at `/webview/:cap/*`
|
||||
rather than framed directly. A direct iframe fails three ways at once: production
|
||||
is HTTPS so `http://` targets are blocked as mixed content, many dashboards send
|
||||
`X-Frame-Options: DENY`, and Codeman's own `default-src 'self'` CSP blocks
|
||||
cross-origin frames. Proxying solves all three without weakening the CSP.
|
||||
|
||||
### The two things that will confuse you
|
||||
|
||||
A proxied frame is sandboxed and therefore **opaque-origin** unless you set
|
||||
`trusted: true`. Two consequences look like bugs in your own app:
|
||||
|
||||
1. **Root-absolute URLs built at runtime** (`/assets/x.png` assembled in JS)
|
||||
escape the injected `<base>` tag. Codeman injects a `runtimeUrlShim()` that
|
||||
patches the common DOM sinks, but if you construct URLs in an unusual way,
|
||||
prefer relative paths.
|
||||
2. **Same-host `fetch` and `XHR` are CORS-checked with `Origin: null`.** Codeman
|
||||
handles this with `buildProxyCorsHeaders()`, and the proxy is exempt from the
|
||||
global `OPTIONS` short-circuit. If you see "Failed to fetch" while the page
|
||||
itself renders fine, this is the area to look at.
|
||||
|
||||
⚠️ `trusted: true` opts out of the sandbox. A proxied page is served from
|
||||
Codeman's origin, so `allow-same-origin` lets it read the Codeman page and call
|
||||
the API that spawns agents. Only mark your own trusted code.
|
||||
|
||||
## Seam 2: SSE events
|
||||
|
||||
`GET /api/v1/events` is a Server-Sent Events stream. Each message is
|
||||
`event: <name>` plus `data: <json>`. There are 149 event names following a
|
||||
`domain:action` convention, registered in `src/web/sse-events.ts`.
|
||||
|
||||
The ones most integrations want:
|
||||
|
||||
| Event | Meaning |
|
||||
| --- | --- |
|
||||
| `session:created`, `session:deleted` | A session appeared or went away |
|
||||
| `session:idle` | The agent stopped working |
|
||||
| `session:completion` | A completion message was detected |
|
||||
| `session:exit`, `session:error` | The session ended or failed |
|
||||
| `hook:permission_prompt` | The agent is asking for permission |
|
||||
| `hook:idle_prompt`, `hook:stop` | The agent is waiting on you, or stopped |
|
||||
| `hook:task_completed`, `task:completed` | Work finished |
|
||||
| `subagent:discovered`, `subagent:completed` | Background agent lifecycle |
|
||||
| `mux:died` | A multiplexer session died unexpectedly |
|
||||
| `cron:runCreated`, `cron:runUpdated` | Scheduled job activity |
|
||||
|
||||
### Filtering
|
||||
|
||||
`?sessions=id1,id2` suppresses only the high-volume `session:terminal` stream for
|
||||
sessions you did not list. Lifecycle and metadata events are always delivered, so
|
||||
you cannot accidentally filter away the thing you are listening for.
|
||||
|
||||
Pass `?clientId=<uuid>` to enable live filter updates through
|
||||
`POST /api/v1/events/subscribe` without reconnecting the stream.
|
||||
|
||||
### Example: notify when any agent needs you
|
||||
|
||||
```js
|
||||
const res = await fetch('http://127.0.0.1:3000/api/v1/events', {
|
||||
headers: { Authorization: 'Basic ' + btoa(`admin:${process.env.CODEMAN_PASSWORD}`) },
|
||||
});
|
||||
const reader = res.body.getReader();
|
||||
const decoder = new TextDecoder();
|
||||
let buf = '';
|
||||
const WANTED = new Set(['hook:permission_prompt', 'hook:idle_prompt', 'session:idle']);
|
||||
|
||||
for (;;) {
|
||||
const { value, done } = await reader.read();
|
||||
if (done) break;
|
||||
buf += decoder.decode(value, { stream: true });
|
||||
const frames = buf.split('\n\n');
|
||||
buf = frames.pop() ?? '';
|
||||
for (const frame of frames) {
|
||||
const name = frame.match(/^event: (.+)$/m)?.[1];
|
||||
const data = frame.match(/^data: (.+)$/m)?.[1];
|
||||
if (name && WANTED.has(name)) notify(name, JSON.parse(data ?? '{}'));
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Seam 3: HTTP API and CLI
|
||||
|
||||
Around 200 handlers across 21 route files cover sessions, cases, files, cron,
|
||||
respawn, Ralph, the orchestrator, search, and admin. Each route module carries an
|
||||
`@fileoverview` describing its endpoints.
|
||||
|
||||
If the caller is an agent running _inside_ a Codeman session, install the packaged
|
||||
agent skill instead of teaching it these calls by hand: `skills/codeman` in the repo
|
||||
(`npx skills add Ark0N/Codeman --skill codeman -g`, or `codeman skill install
|
||||
[--case <name>]`, or the synced `agentSkillEnabled` App Setting for automatic
|
||||
per-case injection on Claude session create). The skill carries the guard, the
|
||||
safety rules, and verified wait/orchestration recipes.
|
||||
|
||||
The common ones:
|
||||
|
||||
```bash
|
||||
# List sessions (live + persisted + transcript history, deduped)
|
||||
curl -u admin:$PASS http://127.0.0.1:3000/api/v1/sessions/unified
|
||||
|
||||
# Create a session
|
||||
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"workingDir":"/home/me/project","mode":"claude"}'
|
||||
|
||||
# Send a prompt (single-line only, and it must end with \r: Enter is sent only
|
||||
# when the input contains a carriage return; without it the text sits on the
|
||||
# session's prompt unsubmitted)
|
||||
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions/$ID/input \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"run the tests\r","useMux":true}'
|
||||
```
|
||||
|
||||
`POST .../input` also accepts `clientId` (stable per client, max 128 chars) and
|
||||
`seq` (monotonic per session). Send both and the server applies each pair
|
||||
at-most-once, so retrying after a dropped connection cannot type the prompt
|
||||
twice. Omit them entirely rather than sending `null`.
|
||||
|
||||
It also accepts `wait` and `waitTimeout`, which hold the response open until the
|
||||
session finishes the turn you just started. `wait` is `true` (the default signal
|
||||
set) or a comma list of `idle,working,stop,blocked,exit`; the result comes back
|
||||
under `data.wait`. Sending them changes nothing for callers that do not: without
|
||||
`wait` the response is still `{"success": true, "data": {}}` and the write is still
|
||||
fire-and-forget. The two interact with `clientId` / `seq` in one way worth knowing:
|
||||
a **tagged duplicate** (a pair the server already applied) skips the write but still
|
||||
waits, answering from the session's current state rather than blocking for a
|
||||
transition that already happened. It reports `"delivered": false, "duplicate": true`.
|
||||
|
||||
### Waiting instead of polling
|
||||
|
||||
Three calls block until something happens: `GET /api/v1/sessions/:id/wait` (a
|
||||
lifecycle signal), `GET /api/v1/sessions/:id/wait-output` (a literal string in the
|
||||
output), and the `wait` field above. Full parameter and response tables are in
|
||||
[`api-reference.md`](api-reference.md#long-polling-agent-wait). Four things decide
|
||||
whether your integration works, and the last one is what actually bites:
|
||||
|
||||
- **A timeout is a `200` with `wait.timedOut: true`**, not an error. Loop over short
|
||||
waits rather than issuing one long one, because `tailscale serve` and cloudflared
|
||||
both cut idle connections and a single 10-minute call is the pattern most likely
|
||||
to die in the field.
|
||||
- **`wait.timeoutMs`** is the timeout after server-side clamping (600 s ceiling by
|
||||
default). Read it rather than assuming you got what you asked for.
|
||||
- **`stop` and `blocked` only exist for `claude` sessions**, and on a `shell` session
|
||||
even `idle` fires only once at startup, so send-and-wait there can only time out.
|
||||
See the Gotchas below.
|
||||
|
||||
⚠️ **There is no readiness signal, and skipping readiness is the failure that looks
|
||||
like success.** A session reports `idle` before its CLI has spawned, and a `claude`
|
||||
worker in a brand-new case comes up on the CLI's **trust dialog**, which has a ❯
|
||||
prompt of its own. Prompt it at that moment and the text lands in the dialog, the
|
||||
`\r` does not get past it, and the session's startup `idle` lands inside the wait
|
||||
window: the wait resolves on `idle` in a couple of seconds with `timedOut: false`,
|
||||
indistinguishable from a finished turn. Wait for the pid, then wait for the
|
||||
composer, answering the dialog only as the bounded fallback.
|
||||
|
||||
⚠️ **Answering it is not "press Enter".** Claude Code 2.1.252 dropped the options'
|
||||
numbers, reversed them, and highlights `No, exit` by default, so a blind `\r` quits
|
||||
the CLI and the pane is dead seconds after the spawn. Read the `❯` marker off the
|
||||
rendered pane (`GET .../terminal?full=1`), send `ESC [ B` while it sits on `No, exit`,
|
||||
re-read, and confirm only once the marker is on `Yes, I trust this folder`. Codeman's
|
||||
own auto-accept (`trustDialogNextKey()` in `src/session-trust-dialog.ts`) does exactly
|
||||
this, inside a 90 s startup window and a 6-keystroke cap.
|
||||
|
||||
A worked orchestration: start a worker, get it ready, prompt it, wait, clean up.
|
||||
|
||||
```bash
|
||||
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}" # auto-set in-session, correct scheme included
|
||||
AUTH=(-u "admin:$CODEMAN_PASSWORD") # omit entirely if no password is set
|
||||
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on --https installs (self-signed cert)
|
||||
|
||||
# 1. Start a worker session (creates the case if it does not exist yet).
|
||||
# The guard matters: a TLS or auth failure otherwise leaves SID empty and every
|
||||
# later step "succeeds" against nothing.
|
||||
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
|
||||
[ -n "$SID" ] && [ "$SID" != null ] || { echo "quick-start failed"; exit 1; }
|
||||
|
||||
# 2. READINESS: composer marker first, trust dialog only as the bounded fallback.
|
||||
# Skip this and step 3 reports a turn that never ran. Match single tokens only:
|
||||
# TUI text can arrive without its spaces. Stage 1 is short on purpose (an
|
||||
# already-trusted case matches in <1 s; a first-run case can never pass it and
|
||||
# pays it in full).
|
||||
# ⚠️ NEVER answer the dialog with a bare \r. Its highlighted option is `No, exit`
|
||||
# (claude-cli 2.1.252), so a blind Enter quits the CLI; and the dialog text stays
|
||||
# in the buffer for the life of the session, so a `from=buffer` probe for `trust`
|
||||
# keeps matching long after it is gone. Read the CURRENT pane instead and steer.
|
||||
until [ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ]
|
||||
do sleep 1; done
|
||||
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
|
||||
--data-urlencode 'timeout=5000') # composer's status bar = ready
|
||||
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
|
||||
ESC=$(printf '\033') # \x1b is GNU-sed only; this form also works on macOS
|
||||
for _ in 1 2 3 4 5 6; do
|
||||
# Which option the ❯ marker sits on, read off the CURRENT frame.
|
||||
K=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/terminal" --data-urlencode 'full=1' \
|
||||
| jq -r '.data.terminalBuffer // empty' \
|
||||
| sed -e "s/$ESC\[[0-9;?]*[a-zA-Z]//g" -e "s/$ESC[()][AB0]//g" | tr -d ' \t' \
|
||||
| grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
|
||||
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/')
|
||||
[ -n "$K" ] || break # no dialog on screen: nothing to answer
|
||||
[ "$K" = confirm ] && IN="\r" || IN="$ESC[B"
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d "$(jq -nc --arg i "$IN" '{input:$i,useMux:true}')" >/dev/null
|
||||
[ "$K" = confirm ] && break
|
||||
sleep 1 # re-read: confirm the arrow landed
|
||||
done
|
||||
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
|
||||
--data-urlencode 'timeout=45000' >/dev/null
|
||||
fi
|
||||
|
||||
# 3. Send the prompt AND register the wait in one call, so the answer cannot be
|
||||
# the previous turn's idle state. Single line only, ending in \r (otherwise
|
||||
# Enter is never sent and this wait times out on a turn that never started).
|
||||
W=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"input":"Run the test suite and summarize the failures\r","useMux":true,
|
||||
"clientId":"orchestrator","seq":1,"wait":"stop,exit","waitTimeout":60000}' \
|
||||
| jq -c '.data.wait')
|
||||
|
||||
# 4. That first wait probably timed out (60 s). Keep going in SHORT waits.
|
||||
for _ in $(seq 1 30); do
|
||||
[ "$(jq -r '.timedOut' <<<"$W")" = 'true' ] || break # signal fired, or wait ended
|
||||
W=$("${CURL[@]}" \
|
||||
"$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq -c '.data.wait')
|
||||
done
|
||||
jq -r 'if .ended or .aborted then "worker is not running"
|
||||
elif .timedOut then "still working after 30 waits"
|
||||
else "signal: \(.signal)" end' <<<"$W"
|
||||
|
||||
# 5. Read what it produced, then delete the session YOU created, by exact id.
|
||||
# ⚠️ NOT /output: its textOutput is empty for every tmux-backed session.
|
||||
# `tail` counts BYTES, and the payload is terminal data with ANSI in it.
|
||||
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
|
||||
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
|
||||
```
|
||||
|
||||
Waiting on a marker instead of a signal is the form that works in **every** mode,
|
||||
and the only one that works on a `shell` session:
|
||||
|
||||
```bash
|
||||
# ⚠️ Split the marker so the typed line never contains it: your own keystrokes echo
|
||||
# into the output stream, so an unsplit marker matches before the command has run.
|
||||
# `from=buffer` also catches a marker that printed before the wait registered.
|
||||
N=$RANDOM
|
||||
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
|
||||
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
|
||||
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
|
||||
--data-urlencode 'timeout=60000' | jq '.data.wait'
|
||||
```
|
||||
|
||||
For shell scripting, the `codeman` CLI is the same surface without the HTTP
|
||||
plumbing:
|
||||
|
||||
```
|
||||
codeman session start|stop|list|logs codeman task add|list|status|remove|clear
|
||||
codeman ralph start|stop|status|reset codeman users add|passwd|list
|
||||
codeman status | list | attach <path> codeman doctor
|
||||
```
|
||||
|
||||
## Seam 4: Hooks
|
||||
|
||||
Claude Code hooks post to `POST /api/v1/hook-event` from inside an agent session.
|
||||
Codeman installs its own hooks automatically, but the endpoint is open to yours.
|
||||
|
||||
```json
|
||||
{ "event": "task_completed", "sessionId": "abc123", "data": { "any": "json" } }
|
||||
```
|
||||
|
||||
`event` must be one of `permission_prompt`, `elicitation_dialog`, `idle_prompt`,
|
||||
`stop`, `teammate_idle`, `task_completed`. Each becomes the matching `hook:*` SSE
|
||||
event.
|
||||
|
||||
⚠️ This endpoint skips Basic auth so hooks keep working, but when auth is active
|
||||
the loopback bypass requires the `X-Codeman-Hook-Secret` header
|
||||
(`~/.codeman/hook-secret`) unconditionally.
|
||||
|
||||
## Gotchas
|
||||
|
||||
Every one of these has cost somebody real time.
|
||||
|
||||
- **CORS is localhost-only.** `Access-Control-Allow-Origin` is echoed only for
|
||||
`localhost`, `127.0.0.1`, and `::1`. A browser app on any other origin cannot
|
||||
call the API. Integrate server-side.
|
||||
- **A missing `Origin` header is allowed**, which is why curl, CLIs, and hooks
|
||||
work. Cross-site origins are blocked by the CSRF guard.
|
||||
- **Reverse-proxy domains are rejected** by the anti-DNS-rebinding Host allowlist
|
||||
unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix`.
|
||||
- **`null` is not `undefined`.** Request schemas use Zod `.optional()`, which
|
||||
accepts `undefined` only. `JSON.stringify({ field: null })` keeps the null on
|
||||
the wire and fails with `INVALID_INPUT`. Omit the key instead. This has caused
|
||||
shipped bugs more than once.
|
||||
- **`text/plain` bodies stay raw.** Auto-parsing them as JSON enabled
|
||||
simple-request CSRF, so it is deliberate. Send `application/json`.
|
||||
- **Prompts are single-line and must end with `\r`.** The server splits your text
|
||||
and Enter into two separate tmux writes (Ink needs them apart), but it sends the
|
||||
Enter **only when the input contains a carriage return**. Without it your text
|
||||
sits on the prompt unsubmitted, which is the single most common "the wait
|
||||
endpoints don't work" report: the wait runs its full timeout on a turn that never
|
||||
started. Newlines inside the string are stripped rather than rejected, so
|
||||
`"echo A\necho B\r"` runs the single joined command `echo Aecho B`: send one line
|
||||
per call.
|
||||
- **`wait-output`'s `from=now` is not "printed after you asked".** tmux repaints
|
||||
the visible screen on attach, on resize, and on any TUI redraw, and a repaint
|
||||
arrives as ordinary output, so text already on screen can satisfy a fresh wait.
|
||||
Observed live: a marker echoed a minute earlier matched instantly. Use a marker
|
||||
unique to each call, and build it so the typed line never contains it (your own
|
||||
keystrokes echo into the stream). Matching is a literal substring, so `regex=` is
|
||||
rejected with a `400` rather than ignored.
|
||||
- **`wait-output` matches the normalized PTY stream, not the screen.** ANSI escape
|
||||
sequences are stripped (the `ESC ( B` charset escape a bash prompt emits on every
|
||||
line included), a partial escape at a chunk boundary is held back until its tail
|
||||
arrives, and a match may straddle PTY chunks, so text you printed yourself
|
||||
matches reliably (`printf STRAD; sleep 1; printf DLEQQ` is matchable as
|
||||
`STRADDLEQQ`). What can still fail is TUI output: a full-screen TUI positions
|
||||
words with cursor moves, so its text can reach the matcher **without spaces** and
|
||||
a multi-word match is unreliable there. Match one short space-free token, ideally
|
||||
one you printed yourself, and keep it out of the typed line (your own keystrokes
|
||||
echo into the stream).
|
||||
- **`stop` and `blocked` never fire for `shell`, `opencode`, `codex`, `gemini`,
|
||||
`antigravity` or `pi` sessions.** They come from Claude Code hooks, which no other mode
|
||||
installs, so only `idle`, `working` and `exit` exist there. Asking for them
|
||||
explicitly is a `400`; omitting `until` is safe, since the server drops them from
|
||||
the default set and echoes what it actually waited on as `wait.until`. Even in
|
||||
`claude` mode, a Docker case needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for hooks to
|
||||
reach the server at all, a remote-SSH case's hooks may never arrive, and a case
|
||||
written by Codeman < 1.13.0 against an `--https` install carries hook curls
|
||||
without `-k` that TLS-fail silently — a 1.13.0+ server rewrites them the next
|
||||
time a session starts in that case.
|
||||
- **Unwrap the envelope** before reading fields. `data` is not the response body.
|
||||
|
||||
## Publishing your integration
|
||||
|
||||
There is no registry and no review queue. Add the GitHub topic
|
||||
**`codeman-integration`** to your public repository so others can find it, and
|
||||
link back to Codeman in your README.
|
||||
|
||||
If a real ecosystem of these appears, a manifest format and an install command
|
||||
become worth building. Until then, these four seams are the contract, and they
|
||||
require nothing of you but HTTP.
|
||||
|
||||
## What Codeman deliberately does not have
|
||||
|
||||
- **No in-process plugin runtime.** See the reasoning at the top of this page.
|
||||
- **No build or startup hooks** for third-party code. Run your own process.
|
||||
- **No per-plugin config or state directories.** Manage your own files.
|
||||
- **No sandbox for integration code**, because Codeman never launches it. Your
|
||||
integration is your own process, started by you, with your permissions,
|
||||
talking HTTP.
|
||||
|
||||
That last point is about integration code specifically, not about Codeman.
|
||||
Sandboxing lives on a different axis here: the thing worth isolating is the
|
||||
**agent**, and you isolate it per case with
|
||||
[Docker cases](docker-cases.md), which run the agent in a hardened container with
|
||||
a bind-mounted workspace and seeded (not shared) credentials. An integration that
|
||||
creates or drives a Docker-backed session inherits that isolation for free, since
|
||||
it is a property of the session rather than of the caller.
|
||||
@@ -1,435 +0,0 @@
|
||||
# File Viewer edit mode (issue #212)
|
||||
|
||||
Plan only. No implementation yet.
|
||||
|
||||
Goal: close the loop "agent writes a file, you review it in the viewer, tweak two lines, save, tell the
|
||||
agent to continue" without hopping into the terminal, with the phone as the primary target.
|
||||
|
||||
Scope from the issue: an Edit toggle on text previews, a write endpoint that inherits the read path's
|
||||
confinement, text-only, edit-in-place (no create, no delete, no rename), no editing through the
|
||||
Docker/remote overlays.
|
||||
|
||||
---
|
||||
|
||||
## 1. What exists today
|
||||
|
||||
**Read path (backend), all in `src/web/routes/file-routes.ts`:**
|
||||
|
||||
| Route | Line | Notes |
|
||||
| ------------------------------------ | ------ | ------------------------------------------------------------------ |
|
||||
| `GET /api/sessions/:id/files` | `741` | Tree scan of `session.workingDir`, hidden files off by default |
|
||||
| `GET /api/sessions/:id/file-content` | `865` | The text/preview classifier. `findSessionOrFail` + `validateSessionFilePath` |
|
||||
| `GET /api/sessions/:id/file-raw` | `1018` | Bytes, 50MB cap |
|
||||
| `GET /api/sessions/:id/file-preview` | `1254` | DOCX/PPTX to PDF, everything else redirects to `file-raw` |
|
||||
| `GET /api/download` | `1384` | The only read route that also runs `isSensitivePath()` |
|
||||
|
||||
`file-content` classification order (`file-routes.ts:881-1011`): extension buckets (image / video / audio /
|
||||
known-binary) return metadata only; otherwise the bytes are read, sniffed for a NUL in the first 8KB, and
|
||||
either reported as `type:'binary'` or decoded as UTF-8 and **truncated to `lines` (default 500, hard cap
|
||||
10000)**. Caps: `MAX_TEXT_FILE_SIZE` 10MB.
|
||||
|
||||
Confinement is `validateSessionFilePath()` (`src/web/route-helpers.ts:67`): `resolve()` then `realpathSync()`
|
||||
then reject if the result is not under `workingDir`. Because it realpaths the *full* path, a symlink whose
|
||||
target escapes the workspace is already rejected. Ownership is `findSessionOrFail()` which runs
|
||||
`canAccessOwned()` (`route-helpers.ts:102`), a no-op outside multi-user mode.
|
||||
|
||||
**Read path (frontend), `src/web/public/panels-ui.js`:**
|
||||
|
||||
- `loadFileBrowser()` `2947`, `renderFileBrowserTree()` `2978`, click to `openFilePreview()` `3056`.
|
||||
- `openFilePreview(filePath, sessionId, attachmentId)` `3193`: attachment-id branch, then docx/pptx, pdf,
|
||||
svg branches, then the generic `file-content` fetch at `3274` with **`&lines=500` hardcoded**, rendering
|
||||
text as `<pre><code>${escapeHtml(...)}</code></pre>` at `3298` and stashing `this.filePreviewContent`.
|
||||
- `closeFilePreview()` `3308`, `copyFilePreviewContent()` `3751`.
|
||||
- Markup: `src/web/public/index.html:420-432` (`filePreviewOverlay` / `-Title` / `-Body` / `-Footer`, two
|
||||
header buttons: copy and close).
|
||||
- CSS: `src/web/public/styles.css:9320-9430`. Overlay `z-index: 2000`, window `80vw/80vh`, capped
|
||||
`900x700`. There are **no `.file-preview-*` rules in `mobile.css` at all**.
|
||||
|
||||
**Reachability on phones.** The header File Viewer button is hidden below 430px
|
||||
(`mobile.css:482`, locked by `KNOWN_PHONE_HIDDEN` in `test/mobile-header-buttons-policy.test.ts`), so on a
|
||||
phone the preview overlay is reached through:
|
||||
|
||||
1. an attachment card's **Preview** button (`panels-ui.js:3451`), which is exactly the "agent just wrote a
|
||||
file" path the issue describes,
|
||||
2. the attachment-history drawer (`panels-ui.js:3709`),
|
||||
3. App Settings to Panels to **File Browser** (`showFileBrowser`, applied in `settings-ui.js:2202`; the
|
||||
panel is mobile-styled at `mobile.css:1868`).
|
||||
|
||||
So edit mode is reachable on a phone today via (1) and (2) without touching the header policy. Improving
|
||||
the entry point is listed as an open decision in section 10, not assumed.
|
||||
|
||||
---
|
||||
|
||||
## 2. Threat model, stated honestly
|
||||
|
||||
Anyone who can call this API can already reach `POST /api/sessions/:id/input` and type an arbitrary prompt
|
||||
into an agent running with `--dangerously-skip-permissions`. A workspace-confined write endpoint therefore
|
||||
does not create a new privilege tier for an authenticated caller.
|
||||
|
||||
What it *would* create if built carelessly is a **new host-write primitive reachable by path**, so the
|
||||
things this plan actually defends against are:
|
||||
|
||||
1. **Path traversal / symlink escape** writing outside the workspace.
|
||||
2. **TOCTOU**: a path component that becomes a symlink between validation and write.
|
||||
3. **Cross-user writes** in multi-user mode (`canAccessOwned`).
|
||||
4. **Silent data loss**, which is the highest-probability real-world failure here and gets its own section.
|
||||
|
||||
CSRF is already covered: `registerHostGuard()` (`src/web/middleware/auth.ts:555-578`) rejects any
|
||||
non-safe-method request whose `Origin` is cross-site. The webview-capability exemption at that gate is
|
||||
fenced to `GET`/`HEAD` for the Referer form (`auth.ts:161`) and to `/webview/:cap/*` paths for the path
|
||||
form, so a proxied dashboard cannot reach a new `PUT /api/...`. Using `PUT` + `application/json` also
|
||||
forces a preflight for any cross-origin attempt.
|
||||
|
||||
---
|
||||
|
||||
## 3. Backend design
|
||||
|
||||
### 3.1 New policy module: `src/config/file-editing.ts`
|
||||
|
||||
Pure, unit-testable, no IO (config lives in `src/config/`, no barrel, import the file directly).
|
||||
|
||||
```ts
|
||||
export const MAX_EDITABLE_BYTES = 512 * 1024; // content cap, both directions
|
||||
export const EDITABLE_EXTENSIONS: ReadonlySet<string>; // ts,tsx,js,jsx,mjs,cjs,json,jsonc,md,mdx,txt,
|
||||
// css,scss,less,html,htm,xml,svg?,yml,yaml,toml,
|
||||
// ini,cfg,conf,env?,sh,bash,zsh,fish,py,rb,go,rs,
|
||||
// java,kt,swift,c,h,cpp,hpp,cs,php,sql,graphql,
|
||||
// proto,lua,pl,r,jl,tf,gradle,csv,tsv,log,diff,patch
|
||||
export const EDITABLE_BASENAMES: ReadonlySet<string>; // Dockerfile, Makefile, LICENSE, .gitignore,
|
||||
// .prettierignore, .editorconfig, .nvmrc, ...
|
||||
export function isEditableFileName(fileName: string): boolean;
|
||||
export function isDeniedEditRelativePath(rel: string): boolean; // `.git/` subtree
|
||||
export function detectEol(text: string): 'lf' | 'crlf';
|
||||
export function applyEol(text: string, eol: 'lf' | 'crlf'): string;
|
||||
```
|
||||
|
||||
Decisions baked in:
|
||||
|
||||
- **Allowlist, not blocklist**, per the issue and per the existing attachment-guard precedent.
|
||||
- `svg` and `env` are deliberately marked with `?` above: `svg` is served as an untrusted octet-stream on
|
||||
the read side (`file-routes.ts:118`) so allowing an edit is defensible, but I recommend **excluding
|
||||
both** in v1. `.env` files are matched by `isSensitivePath()` anyway and would be rejected downstream;
|
||||
excluding them at the allowlist keeps a single obvious refusal.
|
||||
- `isDeniedEditRelativePath` blocks the `.git/` subtree: `.git/hooks/*` is code execution and a corrupt
|
||||
index is unrecoverable-looking to a user who only wanted to fix a typo. Other dotfiles stay allowed but
|
||||
are not reachable from the tree UI anyway (`showHidden=false`).
|
||||
|
||||
### 3.2 Read-for-edit: extend the existing GET
|
||||
|
||||
`GET /api/sessions/:id/file-content?path=<rel>&edit=1`
|
||||
|
||||
When `edit=1`:
|
||||
|
||||
- skip line truncation entirely (a truncated buffer must never become an edit buffer, see section 4.1),
|
||||
- enforce `MAX_EDITABLE_BYTES` instead of `MAX_TEXT_FILE_SIZE` and answer 413 over it (as a structured
|
||||
throw with `statusCode: 413`, the `throwFilesystemPickerError` pattern, since the central errorCode-to-
|
||||
status map has no 413 entry; see the error-mechanics note in 3.3),
|
||||
- run the editability gate (`isEditableFileName`, `isDeniedEditRelativePath`, `isSensitivePath`,
|
||||
`isBlockedAttachmentPath`) and the content gate (NUL sniff plus UTF-8 round-trip, see 4.3),
|
||||
- return `{ content, size, mtimeMs, totalLines, truncated: false, extension, editable: true, hash, eol }`.
|
||||
`hash` is `sha256` hex of the exact on-disk bytes.
|
||||
|
||||
Non-`edit` responses gain **only** `editable: boolean` (additive, no shape change for existing consumers),
|
||||
which is all the UI needs to decide whether to show the Edit button. No `hash` on plain reads: the Edit
|
||||
action re-fetches with `edit=1` anyway (section 4.1), which is where the hash comes from, and hashing every
|
||||
casual 10MB preview would be pure waste.
|
||||
|
||||
### 3.3 Write: `PUT /api/sessions/:id/file-content`
|
||||
|
||||
Body (new `FileWriteSchema` in `src/web/schemas.ts`, Zod v4):
|
||||
|
||||
```ts
|
||||
{ path: string, content: string, baseHash: string, eol?: 'lf'|'crlf', force?: boolean }
|
||||
```
|
||||
|
||||
Registered with an explicit route option `{ bodyLimit: 4 * 1024 * 1024 }`. **Fastify's default `bodyLimit`
|
||||
is 1MB and this repo configures none**, and JSON escaping expands content: 2x for a file full of quotes or
|
||||
backslashes, up to 6x for control characters (each serialized as a `\uXXXX` escape), so 512KB of content
|
||||
can legitimately exceed 1MB on the wire; blowing the limit produces a raw `FST_ERR_CTP_BODY_TOO_LARGE`, not an `ApiResponse` envelope. Two
|
||||
related sizing notes: `z.string().max()` counts **UTF-16 code units, not bytes**, so the schema's `.max()`
|
||||
is only a coarse pre-filter and the real cap is an explicit `Buffer.byteLength(content, 'utf8')` check in
|
||||
the handler (step 7a below); and 4MB comfortably bounds the worst-case expansion of a 512KB file without
|
||||
inviting multi-MB bodies elsewhere.
|
||||
|
||||
**Error mechanics** (matters for both prod behavior and testability): a handler that *returns* a
|
||||
`{success:false, errorCode}` envelope gets its HTTP status assigned centrally by the preSerialization hook
|
||||
in `server.ts` (`httpStatusForErrorCode()`, `src/types/api.ts`), but the route-test harness
|
||||
(`test/routes/_route-test-utils.ts`) installs only `installRouteErrorHandler`, **not** that hook, so
|
||||
returned envelopes surface as HTTP 200 in tests. The PUT handler should therefore use the same
|
||||
structured-**throw** pattern as the filesystem picker (`throwFilesystemPickerError`, `file-routes.ts:411`):
|
||||
thrown `{statusCode, body}` errors are rendered identically in prod and in the harness, and they allow the
|
||||
one status the code map cannot express (413). The error envelope itself is strictly
|
||||
`{success:false, error, errorCode}`, **it has no data arm**, so no error response may carry extra payload.
|
||||
|
||||
Handler order (each step is a test case):
|
||||
|
||||
1. `findSessionOrFail(ctx, id, req)` (live sessions only, matching the read route, and it carries the
|
||||
multi-user ownership check).
|
||||
2. `parseBody(FileWriteSchema, req.body)`, then `Buffer.byteLength(content, 'utf8') <= MAX_EDITABLE_BYTES`
|
||||
or 413 (the schema `.max()` alone cannot enforce a byte cap, see the sizing note above).
|
||||
3. `validateSessionFilePath(session.workingDir, path)` or 404 (do not distinguish "outside workspace" from
|
||||
"missing", matching the read route).
|
||||
4. `isSensitivePath(resolvedPath) || isBlockedAttachmentPath(resolvedPath, guard.blockedTrees)` or 403.
|
||||
5. `isDeniedEditRelativePath(relativePath)` or 403.
|
||||
6. `isEditableFileName(basename(resolvedPath))` or 400.
|
||||
7. `stat`: must be `isFile()`, size within `MAX_EDITABLE_BYTES`, else 400/413. **No `O_CREAT` anywhere in
|
||||
this handler**, which is what enforces edit-in-place.
|
||||
8. Read current bytes, compute `hash`, run the NUL sniff and the UTF-8 round-trip check, else 400.
|
||||
9. `hash !== baseHash && !force` gives **409 CONFLICT** (`ApiErrorCode.CONFLICT`, plain envelope; the error
|
||||
arm carries no data, see the error-mechanics note). The client's conflict dialog gets fresh state by
|
||||
re-fetching `edit=1`, which it needs for its Reload action anyway.
|
||||
10. Build the output buffer: `applyEol(content, eol ?? detected-from-original)`; re-check
|
||||
`Buffer.byteLength` against the cap.
|
||||
11. Write atomically in the resolved parent directory:
|
||||
`fs.open(<dir>/.<name>.codeman-tmp-<rand>, 'wx', stat.mode & 0o777)`, then `fchmod(stat.mode & 0o777)`
|
||||
(open's mode argument is masked by the process umask, so the chmod is what actually preserves an
|
||||
unusual mode), write, `fsync`, close, `fs.rename(tmp, resolvedPath)`, unlink the temp on any failure.
|
||||
12. Re-stat, return `{ success: true, data: { path, size, mtimeMs, hash, totalLines } }`.
|
||||
|
||||
Why `O_EXCL` temp plus rename rather than truncate-in-place:
|
||||
|
||||
- `wx` cannot follow a pre-existing symlink, which closes the TOCTOU window from step 3 to step 11 without
|
||||
needing `O_NOFOLLOW` gymnastics.
|
||||
- `rename()` does not follow a symlink in the final component, so even if `resolvedPath` were swapped for a
|
||||
symlink after validation, the symlink itself is replaced and the swap target is untouched.
|
||||
- A crash mid-write leaves the original intact.
|
||||
|
||||
Caveat to document in the code comment: rename replaces the inode, so hardlinks to the file keep the old
|
||||
content. That is the same trade-off vim makes by default and is preferable to a truncate window here.
|
||||
|
||||
No SSE event in v1. Nothing else in the app needs to know: `image-watcher.ts` only reacts to
|
||||
`.png/.jpg/.jpeg/.gif/.webp/.bmp/.svg/.pdf/.docx/.pptx` adds (`image-watcher.ts:23-25`), none of which are
|
||||
editable text, and the temp filename does not match either.
|
||||
|
||||
---
|
||||
|
||||
## 4. The five traps
|
||||
|
||||
These are the parts that turn a "small write endpoint" into a bug report.
|
||||
|
||||
### 4.1 Truncation (the data-loss trap)
|
||||
|
||||
The frontend fetches `&lines=500` (`panels-ui.js:3274`). Saving that buffer back would **delete every line
|
||||
past 500**. Worse, the content hash of the full file would still match, so an optimistic-concurrency check
|
||||
cannot catch it.
|
||||
|
||||
Mitigations, all three:
|
||||
|
||||
- The Edit affordance is only offered when the loaded payload came from `edit=1` (which never truncates).
|
||||
Tapping Edit on an already-rendered preview **re-fetches** with `edit=1` before swapping in the editor.
|
||||
- The read-for-edit path 413s above `MAX_EDITABLE_BYTES` rather than truncating, so "too big to edit here"
|
||||
is an explicit refusal with a message, never a silent partial buffer.
|
||||
- A test asserts `edit=1` never returns `truncated: true`.
|
||||
|
||||
### 4.2 Line endings
|
||||
|
||||
A `<textarea>`'s `.value` normalizes to LF. Saving a CRLF file naively rewrites every line, producing a
|
||||
whole-file diff for a two-line change. So: the read returns the detected `eol`, the client echoes it back
|
||||
unchanged, and the server re-applies it. Mixed-EOL files use the dominant style, which is lossy for the
|
||||
minority lines; call that out in the response and accept it in v1.
|
||||
|
||||
### 4.3 Encoding
|
||||
|
||||
`buf.toString('utf-8')` on a latin-1 or otherwise non-UTF-8 file yields U+FFFD replacement characters, and
|
||||
writing that back **corrupts the file**. The check is a round-trip:
|
||||
`Buffer.from(decoded, 'utf8').equals(buf)`. If it fails, `editable: false` and the write is refused. This
|
||||
also catches binary content that the NUL sniff misses. A UTF-8 BOM survives because it round-trips as a
|
||||
leading U+FEFF; do not strip it.
|
||||
|
||||
### 4.4 Concurrency with the agent
|
||||
|
||||
The whole use case is editing a file the agent just wrote and may write again. `baseHash` plus 409 is the
|
||||
guard. Do not use mtime alone: agents rewrite files within a single filesystem timestamp tick, and an
|
||||
identical rewrite should not be reported as a conflict.
|
||||
|
||||
### 4.5 Symlinks and TOCTOU
|
||||
|
||||
Covered by `validateSessionFilePath` (escape) plus `wx` temp and `rename` (post-validation swap). One
|
||||
intentional allowance: a symlink whose target is *inside* the workspace is edited through to its target,
|
||||
because `validateSessionFilePath` returns the realpath. That matches what a user tapping the file expects.
|
||||
|
||||
---
|
||||
|
||||
## 5. Frontend design
|
||||
|
||||
All in `panels-ui.js` (prettier-exempt, hand-formatted; match the surrounding style), `index.html`,
|
||||
`styles.css`, `mobile.css`.
|
||||
|
||||
### 5.1 State
|
||||
|
||||
```js
|
||||
filePreviewEdit = { active, sessionId, path, baseHash, eol, original, dirty }
|
||||
```
|
||||
|
||||
Reset in `closeFilePreview()` and on every `openFilePreview()` entry.
|
||||
|
||||
### 5.2 Markup (`index.html:420-432`)
|
||||
|
||||
Add one header button (pencil, `btn-icon-sm`, `id="filePreviewEditBtn"`, hidden by default) next to the
|
||||
copy button, and an edit bar inside the footer region holding Save / Cancel / a dirty dot. Keep the
|
||||
existing footer text element; the edit bar is a sibling toggled by class so the read-mode footer is
|
||||
untouched.
|
||||
|
||||
### 5.3 Behavior
|
||||
|
||||
- `openFilePreview()` shows the Edit button only when the response has `editable: true` and the render took
|
||||
the text branch. Attachment-id previews, media, binary, pdf, docx/pptx and svg all leave it hidden.
|
||||
- **Enter edit**: re-fetch with `edit=1`; on 413 or `editable:false`, toast the reason and stay in read
|
||||
mode. This fetch must **parse the error envelope on non-ok responses**: the existing generic
|
||||
`if (!res.ok) throw new Error('Failed to load file')` pattern (`panels-ui.js:3275`) would swallow the
|
||||
specific "too large to edit here" message, since error envelopes arrive with real 4xx statuses in prod. On success replace the body with `<textarea class="file-preview-editor" spellcheck="false"
|
||||
autocapitalize="off" autocorrect="off" autocomplete="off" wrap="off">` and assign `.value = content`
|
||||
(never `innerHTML`, so no escaping question arises). Do **not** autofocus: on a phone that opens the
|
||||
keyboard before the user has picked a line.
|
||||
- `input` sets `dirty` and enables Save.
|
||||
- **Save**: `PUT` with `baseHash`, `eol`, and `content`. On success update `baseHash`/`original` from the
|
||||
response, leave edit mode, re-render the read view from the local editor value (the response carries
|
||||
metadata only, not content), toast "Saved". On **409** offer `Reload (discard mine)` / `Overwrite`:
|
||||
Reload re-fetches `edit=1` and replaces the buffer; Overwrite re-sends with `force: true`. The 409 body
|
||||
itself carries no state (section 3.3, step 9).
|
||||
- **Cancel / close / Escape while dirty**: `confirm('Discard unsaved changes?')`, consistent with the
|
||||
existing `window.confirm` usage in this codebase (`panels-ui.js:4323`, `app.js:4176`). Note the global
|
||||
Escape handler (`app.js:999-1007`) closes other panels via `closeAllPanels()` but does not touch this
|
||||
overlay today; if Escape-to-close is wired up as part of this work it must go through the same dirty
|
||||
guard.
|
||||
- `copyFilePreviewContent()` copies the live editor value while editing.
|
||||
|
||||
⚠️ Repo gotcha to respect at the fetch call: **Zod `.optional()` rejects `null`**. Build the body with
|
||||
`eol: eol ?? undefined` (or declare `.nullish()`), or the PUT fails `INVALID_INPUT`. This has shipped as a
|
||||
real bug twice.
|
||||
|
||||
### 5.4 Mobile
|
||||
|
||||
- **Sizing.** The window is `80vw/80vh` centered with no mobile override, so when the keyboard opens on iOS
|
||||
the lower half sits behind it. Add a `@media (max-width: 430px)` block using
|
||||
`height: var(--app-height, 100vh)`, full width, no border radius. `--app-height` is already maintained
|
||||
against `visualViewport` by `KeyboardHandler.handleViewportResize()` (`mobile-handlers.js:283-317`), so
|
||||
the editor tracks the keyboard for free.
|
||||
- **iOS zoom.** The editor font must be >= 16px on phones; there is an existing zoom-prevention block at
|
||||
`mobile.css` under `@media (max-width: 768px)`. Verify it covers `textarea` and do not override it with a
|
||||
smaller `rem` value.
|
||||
- **Accessory bar.** Focusing any input fires `KeyboardHandler.onKeyboardShow()`, which calls
|
||||
`KeyboardAccessoryBar.show()` and refits/resizes the terminal (`mobile-handlers.js:407+`). The bar's keys
|
||||
target the **terminal**, not the editor, so an Esc or clear-input tap while editing goes to the agent.
|
||||
The overlay's `z-index: 2000` covers the bar's `51`, so it is not visible, but confirm it is not
|
||||
interactive underneath and consider an explicit `KeyboardAccessoryBar.hide()` while the editor holds
|
||||
focus. This is the item most likely to look "fine on desktop, wrong on the phone".
|
||||
- No header-policy change is needed (section 1), so
|
||||
`test/mobile-header-buttons-policy.test.ts` stays untouched.
|
||||
|
||||
### 5.5 i18n
|
||||
|
||||
`i18n.js` already skips `textarea`, `pre`, `code` and `.file-preview-content` in its `SKIP_SELECTOR`
|
||||
(`i18n.js:20-38`), so file content is never translated. Add zh-CN entries for the new chrome: Edit, Save,
|
||||
Cancel, Unsaved changes, Discard unsaved changes?, File changed on disk, Reload, Overwrite, Saved,
|
||||
Too large to edit here.
|
||||
|
||||
---
|
||||
|
||||
## 6. Docker and remote cases
|
||||
|
||||
Out of scope per the issue, and the current behavior already degrades correctly:
|
||||
|
||||
- **Docker cases**: the workspace is a host directory bind-mounted at the same absolute path, so a host-side
|
||||
write is visible in the container immediately. Edit mode works and needs nothing special. Worth one line
|
||||
in the docs.
|
||||
- **Remote SSH cases**: `workingDir` is a path on the remote host, and the READ routes now
|
||||
resolve it over ssh (`src/remote-files.ts`, same `buildSshConnectionArgs` discipline as the
|
||||
launch path — #415). What stays unsupported is the WRITE side: an `edit=1` / `PUT` answers
|
||||
`400` "editing is not supported for files in a remote (SSH) case", `editable` is always
|
||||
`false`, office previews and generated thumbnails answer `400`, and no remote file is ever
|
||||
copied to the server's disk. Do not attempt an SFTP write path.
|
||||
|
||||
---
|
||||
|
||||
## 7. Tests
|
||||
|
||||
| File | Kind | Covers |
|
||||
| ------------------------------------------- | ----------- | ---------------------------------------------------------------------- |
|
||||
| `test/file-editing-policy.test.ts` | pure unit | `isEditableFileName` (allow + deny + basenames), `isDeniedEditRelativePath`, `detectEol`/`applyEol` round-trip incl. mixed EOL, BOM preservation |
|
||||
| `test/routes/file-write-routes.test.ts` | `app.inject` | The handler order in 3.3, against a **real temp dir** (do not `vi.mock('node:fs')` in this file; set `MockSession.workingDir`, `test/mocks/mock-session.ts:14`) |
|
||||
| extend `test/routes/file-routes.test.ts` | `app.inject` | `edit=1` never truncates; `editable` present on the plain read |
|
||||
|
||||
Status-code caveat for all of these: the route-test harness does not install the server's preSerialization
|
||||
envelope hook, so a handler that *returns* an error envelope answers 200 in tests. The statuses below are
|
||||
only assertable because the plan has the handler **throw** structured errors (section 3.3, error
|
||||
mechanics), which `installRouteErrorHandler` renders identically in prod and in the harness.
|
||||
|
||||
Route cases to assert explicitly:
|
||||
|
||||
1. happy path writes the bytes and returns a new hash
|
||||
2. `../` and absolute paths give 404
|
||||
3. symlink pointing outside the workspace gives 404
|
||||
4. symlink pointing inside is written through to the target
|
||||
5. non-allowlisted extension gives 400
|
||||
6. `.git/config` gives 403
|
||||
7. a `.env` in the workspace gives 403 (sensitive-path)
|
||||
8. a file with a NUL byte gives 400
|
||||
9. a latin-1 file that fails the UTF-8 round-trip gives 400
|
||||
10. stale `baseHash` gives 409 (`CONFLICT` envelope, no data); `force:true` then succeeds
|
||||
11. over `MAX_EDITABLE_BYTES` gives 413
|
||||
12. a path that does not exist gives 404 and creates nothing (no `O_CREAT`)
|
||||
13. multi-user: `authUser: {role:'user'}` against another user's session gives 404 (pass `authUser` to
|
||||
`createRouteTestHarness`, otherwise the synthetic admin makes the test pass vacuously)
|
||||
14. CRLF file edited and saved stays CRLF
|
||||
15. file mode is preserved across the temp-plus-rename
|
||||
|
||||
Run with `npm test -- test/routes/file-write-routes.test.ts`, never bare `npm test`.
|
||||
|
||||
**End-to-end verification before any deploy** (unit tests passing is not sufficient here):
|
||||
|
||||
- `curl -sk https://localhost:3000/...` against a **throwaway** session created for the purpose, never
|
||||
`w1`/`w2`/`w3`; delete it by exact id afterwards.
|
||||
- Playwright on a phone profile: open a preview, tap Edit, type with `page.keyboard.type()`, Save, then
|
||||
assert the bytes on disk changed. Assert real state, not HTTP 200.
|
||||
|
||||
---
|
||||
|
||||
## 8. Docs and release
|
||||
|
||||
- This plan lives at `docs/file-viewer-edit-plan.md`.
|
||||
- `docs/architecture-invariants.md`: new anchor `#file-viewer-edit-mode` covering the write confinement
|
||||
chain, the truncation invariant, and why temp-plus-rename.
|
||||
- `CLAUDE.md`: one line under the **Filesystem path picker** neighborhood noting that the File Viewer now
|
||||
has a **third** file surface and that it is the only one that writes, plus its confinement rules.
|
||||
Remember `CLAUDE.md` is prettier-ignored on purpose.
|
||||
- `docs/api-reference.md`: the new `PUT` and the `edit=1` query.
|
||||
- Release: a normal COM applies (the 1.10.0 batch hold is over). This is a new user-facing feature plus an
|
||||
additive API surface, so **COM minor** when it ships.
|
||||
|
||||
Formatting note: `panels-ui.js`, `styles.css`, `mobile.css`, `index.html` are all in `.prettierignore` and
|
||||
are hand-formatted; new TypeScript (`src/config/file-editing.ts`, route + schema edits) is prettier-enforced
|
||||
and must pass `npm run format:check`.
|
||||
|
||||
---
|
||||
|
||||
## 9. Implementation order
|
||||
|
||||
Each phase is independently reviewable and leaves the tree working.
|
||||
|
||||
1. **Policy module + tests.** `src/config/file-editing.ts` and `test/file-editing-policy.test.ts`. Pure, no
|
||||
route wiring. (Small.)
|
||||
2. **Read-for-edit.** `edit=1` (returning `hash`/`eol`) plus the additive `editable` flag on plain reads,
|
||||
tests. Nothing consumes it yet. (Small.)
|
||||
3. **Write endpoint.** `FileWriteSchema`, `PUT` handler, `test/routes/file-write-routes.test.ts`. Fully
|
||||
testable by curl before any UI exists. (Medium, the security-relevant part.)
|
||||
4. **Desktop UI.** Edit button, textarea swap, Save/Cancel, dirty guard, 409 flow. (Medium.)
|
||||
5. **Mobile pass.** `mobile.css` sizing against `--app-height`, font size, accessory-bar interaction,
|
||||
real-device check. (Small but the part that decides whether the feature is actually usable.)
|
||||
6. **Docs, i18n strings, changeset.**
|
||||
|
||||
---
|
||||
|
||||
## 10. Open decisions
|
||||
|
||||
1. **Editor widget.** Recommend a plain `<textarea>` for v1: zero dependencies, no CSP question, no bundle
|
||||
growth, and it is the only thing guaranteed to behave with the iOS keyboard. CodeMirror-light with
|
||||
syntax highlighting is a clean follow-up once the write path is proven. The issue allows either.
|
||||
2. **Phone entry point.** Edit mode is reachable on a phone through attachment cards and the history
|
||||
drawer without changing anything. A dedicated toolbar or overview affordance for "browse this session's
|
||||
files" would make it discoverable, but it is a separate UX change and would need a decision against the
|
||||
deliberately minimal phone header policy. Recommend deferring it and revisiting after the feature ships.
|
||||
3. **`svg` editability.** Recommend excluded in v1 (it is deliberately treated as untrusted on the read
|
||||
side). Easy to add later.
|
||||
4. **Create / delete / rename.** Explicitly out of scope per the issue. Note that keeping `O_CREAT` out of
|
||||
the handler is what makes that a structural property rather than a convention.
|
||||
@@ -1,106 +0,0 @@
|
||||
# Grok Build (xAI) integration plan
|
||||
|
||||
> **Status**: Executed. This document records the plan, the decision behind each wiring
|
||||
> point, and what was and was not verified. The user-facing guide is
|
||||
> [`grok-integration.md`](./grok-integration.md); the per-decision invariants live in
|
||||
> [`architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok`](./architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok).
|
||||
> Template: the pi integration (`c5b5963`, [`pi-integration-plan.md`](./pi-integration-plan.md)),
|
||||
> which was itself calibrated against the four follow-up commits the antigravity
|
||||
> integration needed. All of grok's facts below were verified against **grok 1.0.5**
|
||||
> (`grok 1.0.5 (5115b46bc9)`), installed live during the work.
|
||||
|
||||
## 1. What Grok Build is
|
||||
|
||||
[xai-org/grok-build](https://github.com/xai-org/grok-build) is xAI's coding agent: a
|
||||
Rust fullscreen-TUI binary named `grok`, installed by
|
||||
`curl -fsSL https://x.ai/cli/install.sh | bash` into `~/.grok/bin` (with symlinks into
|
||||
`~/.local/bin`; the installer also ships an `agent` alias). Config lives in
|
||||
`~/.grok/config.toml`, TUI appearance in `~/.grok/pager.toml`, credentials in
|
||||
`~/.grok/auth.json` (0600), sessions under `~/.grok/sessions/`. Auth is browser OAuth
|
||||
on first launch, `grok login --device-auth` for SSH boxes, or `XAI_API_KEY` for
|
||||
headless use. It has Claude-style permission modes (`default`/`acceptEdits`/`auto`/
|
||||
`dontAsk`/`bypassPermissions`/`plan`), allow/deny rules, hooks, MCP, subagents, and a
|
||||
headless `-p` mode.
|
||||
|
||||
## 2. Shape decisions (why grok is wired the way it is)
|
||||
|
||||
Grok is a seventh run mode, alongside Claude Code, shell, OpenCode, Codex, Gemini,
|
||||
Antigravity and Pi. Never a location overlay, never a web tab. Its wiring mixes two
|
||||
existing shapes:
|
||||
|
||||
| Question | Decision | Why |
|
||||
| --- | --- | --- |
|
||||
| Permission bypass | `GrokConfig.alwaysApprove` -> `--always-approve` | Grok's real flag (verified via `--help`): "Auto-approve all tool executions", i.e. its `bypassPermissions` mode. Config-level deny rules still apply on top. The Run button sends `true`, matching `runAntigravity()` and Claude's own `--dangerously-skip-permissions` default: Codeman sessions exist for autonomous work. |
|
||||
| Multi-user clamp branch | only-if-sent (codex/antigravity branch) | A bare `grok` spawn is grok's own ask-mode default, which is already safe, so the clamp only needs to force a SENT `alwaysApprove` off. Contrast pi, whose absent default is an answerable prompt and therefore needs the materialize branch. Cron needs nothing for grok for the same reason (`clampCronExternalCliConfigs`). |
|
||||
| Alt-screen strip | OUT of `isAltScreenStripMode()` | Grok is a fullscreen alternate-screen TUI with mouse support (its own scrollback pane, `pager.toml [terminal] alt_screen`), i.e. the opencode case, not the Ink repaint case. It falls through to the narrow tmux-attach strip like opencode/antigravity/pi. |
|
||||
| Resolver | version probe, like pi | `grok` has npm squatters (the unrelated `@vibe-kit/grok-cli` installs a `grok` bin). Candidates must pass `grok --version`; `GROK_VERSION_REGEX` is exported and shared with the dependency registry so doctor and run mode cannot disagree. The probe cannot tell two version-printing `grok`s apart, so `GET /api/grok/status` surfaces path AND version. Search dirs: `~/.grok/bin` first (installer target), then `~/.local/bin`, `/usr/local/bin`, `~/bin`. |
|
||||
| Env allowlist | `GROK_*` + `XAI_*` prefixes | `GROK_*` covers grok's documented inputs (`GROK_HOME`, `GROK_CONFIG`/`GROK_CONFIG_PATH`, `GROK_MEMORY`, `GROK_WORKFLOWS`, `GROK_SANDBOX`, `GROK_OIDC_*`, `GROK_AUTH_PROVIDER_COMMAND`). `XAI_*` is xAI's vendor namespace and carries `XAI_API_KEY`, grok's documented headless auth var: the same narrow-vendor-namespace reasoning that admitted `GOOGLE_*` for gemini. Foreign provider keys stay out, as always. |
|
||||
| Resume | `--resume <id>` / `--continue`, id-regexed | Grok's `--resume` also matches session TITLES (arbitrary user strings, case-insensitive). The `^[a-zA-Z0-9._-]+$` regex doubles as the no-titles rule, so nothing free-form can reach the `bash -c` spawn line. A valid explicit id wins over `-c`, mirroring pi. |
|
||||
| Local echo | `'buffer'` via the `_updateLocalEchoState` fallthrough | UNMEASURED against an authenticated session (see §4). If grok's composer turns out per-keystroke reactive like codex's, the fallback is one `'off'` branch; teaching `PredictiveEchoAddon` grok's composer row is the larger follow-up. |
|
||||
| Truecolor | `COLORTERM=truecolor` + `unset NO_COLOR` | Rust TUI with themes; joins the codex/gemini/antigravity/pi list in `buildEnvExports()` and `buildMuxAttachEnv()`. |
|
||||
| Docker credentials | per-file seed: `auth.json`, `config.toml`, `pager.toml` | `~/.grok` also holds `sessions/`, `memory/`, `completions/`, `docs/` and the ~160MB binary under `downloads/`; a whole-dir seed would copy all of it on every container start. Same trade-off as pi: in-container sessions are invisible host-side, so `grok -c` in a Docker case sees only that container's history. |
|
||||
| Docker install | own Dockerfile step | Not an npm package. xAI's installer has no `--dir` override, so the step copies `/root/.grok/bin/grok` (through the symlink, `cp -L`) into `/usr/local/bin` and removes root's `~/.grok` in the same layer. |
|
||||
| Remote SSH | `exec "$SHELL" -i -l -c 'grok'` | sshd's remote-command PATH does not include `~/.grok/bin`; same login-shell fix as every other agent CLI. |
|
||||
| What is NOT wired | `--permission-mode`, `--allow`/`--deny`, `-p` headless, `--worktree`, `--sandbox`, `--reasoning-effort`, `-s/--session-id`, `--fork-session`, `--agent`, `--output-format` | Follow-ups. The flag surface is kept minimal on purpose; grok is pre-1.0-style fast-moving and every flag added is a flag validated forever. |
|
||||
|
||||
## 3. Touch points (the checklist)
|
||||
|
||||
Backend: `types/session.ts` (SessionMode + GrokConfig + SessionState), `utils/grok-cli-resolver.ts` (new)
|
||||
+ barrel, `tmux-manager.ts` (`buildGrokCommand`, dispatch, resume flag, PATH export, truecolor,
|
||||
availability error, plumbing), `session.ts` (external-mode gate, label, config plumbing,
|
||||
tmux-required error, attach env), `mux-interface.ts`, `schemas.ts` (prefixes, `GrokConfigSchema`,
|
||||
both mode enums, remote command overrides, cron agentType), `session-routes.ts` (clamp + both
|
||||
create paths), `system-routes.ts` (`GET /api/grok/status`), `server.ts` (availability inject +
|
||||
mux restore), `docker-hosts.ts`, `remote-hosts.ts`, `config/dependency-registry.ts`,
|
||||
`cron/cron-service.ts` (comment), `response-viewer-transcript.ts`, `tui/tui-client.ts` + `tui-app.ts`.
|
||||
|
||||
Frontend: `index.html` (welcome button, run-mode entry, cron option, clone Brain option),
|
||||
`session-ui.js` (`runGrok()`, dispatch, availability, "Run GK" label, external-CLI gates,
|
||||
runMode setter), `app.js` (label, `gk` tab badge, kill-menu), `settings-ui.js`,
|
||||
`mobile-overview.js`, `home-sessions.js`, `panels-ui.js`, `i18n.js`, `styles.css` +
|
||||
`mobile.css` (charcoal monochrome identity; the non-og skin block and the mobile
|
||||
`!important` pair are both load-bearing, see the pi plan's §2.9 cascade trap).
|
||||
|
||||
Meta: `docker/agent.Dockerfile`, `install.sh`, `package.json` keyword, changeset,
|
||||
`skills/codeman/reference/*`, CLAUDE.md, READMEs, `architecture-invariants.md`,
|
||||
`remote-sessions.md`, `security-architecture.md`, `docker-cases.md`, `cron-guide.md`.
|
||||
|
||||
Tests: `test/grok-mode.test.ts` + `test/grok-cli-resolver.test.ts` (new);
|
||||
`external-cli-bypass-clamp`, `system-routes`, `render-index-html`, `run-mode-ui`,
|
||||
`mobile-overview`, `local-echo-codex-gating` (extended).
|
||||
|
||||
## 4. Verification performed
|
||||
|
||||
On this box, with grok 1.0.5 really installed and an isolated
|
||||
`CODEMAN_INSTANCE=grokwt` server (own data dir, own tmux socket, port 5077):
|
||||
|
||||
1. `npm test` (the CI gate): green, 5900+ tests. `typecheck`, `lint`, `format:check`,
|
||||
`check:frontend-syntax`, `check:public-assets`, `check:lockfile`: green.
|
||||
2. `GET /api/grok/status` -> `{available: true, path: "/home/arkon/.local/bin", version: "1.0.5"}`
|
||||
through the real resolver and probe.
|
||||
3. `POST /api/quick-start {mode: "grok", grokConfig: {alwaysApprove: true}}` -> session
|
||||
created, tmux pane spawned, real spawn line verified to end in `grok --always-approve`,
|
||||
and the actual grok TUI rendered its OAuth device-approval screen in the pane
|
||||
(unauthenticated box, so sign-in is exactly where a first run lands).
|
||||
4. `grokConfig` persisted into the instance's `state.json`.
|
||||
5. Session deleted by exact id; instance data dir and throwaway case removed.
|
||||
|
||||
**Not verified (honest gaps, all requiring an xAI account or more hardware):**
|
||||
an authenticated conversation end to end; the local-echo buffer policy against grok's
|
||||
real composer (§2); scrollback/repaint behavior of the fullscreen TUI under the narrow
|
||||
strip during a long session; a Docker case with `mode: 'grok'` (needs a `--no-cache`
|
||||
agent-image rebuild); a remote-SSH grok case; cron readiness degradation (expected:
|
||||
same slow-start-then-send as pi, documented in `cron-guide.md`).
|
||||
|
||||
## 5. Follow-ups
|
||||
|
||||
- Idle/completion signal: grok has a hooks system (user-guide `10-hooks.md`); a hook
|
||||
POSTing to `/api/hook-event` could give grok sessions real idle detection instead of
|
||||
output-stabilization. Highest-value follow-up, same slot as pi's `agent_settled` idea.
|
||||
- Response viewer: sessions are ACP JSONL under `~/.grok/sessions/<encoded-cwd>/<id>/updates.jsonl`;
|
||||
`grok -p ... --output-format json | jq -r '.sessionId'` exists for correlation.
|
||||
- Permission-mode picker (`--permission-mode`, `--allow`/`--deny`) in Session Options.
|
||||
- Measure the local-echo policy and the fullscreen-TUI scrollback behavior against an
|
||||
authenticated session; pin the result in `local-echo-codex-gating` the way pi did.
|
||||
- `grok doctor` is a built-in terminal-support check worth pointing users at when a
|
||||
pane renders oddly.
|
||||
@@ -1,133 +0,0 @@
|
||||
# Grok Build (xAI) sessions
|
||||
|
||||
Codeman can drive [Grok Build](https://github.com/xai-org/grok-build) (xAI's `grok`
|
||||
CLI, the agent behind docs.x.ai/build) as a session backend, alongside Claude Code,
|
||||
OpenCode, Codex, Gemini, Antigravity and Pi. `grok` is a seventh **run mode**: its own
|
||||
PTY, its own tmux session, its own tab identity (monochrome charcoal, `gk` badge). It
|
||||
is not a location overlay like Docker or remote-SSH cases, and it is not a web tab.
|
||||
|
||||
The design rationale behind each decision below lives in
|
||||
[`grok-integration-plan.md`](./grok-integration-plan.md). Everything here was verified
|
||||
against grok 1.0.5.
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
curl -fsSL https://x.ai/cli/install.sh | bash
|
||||
```
|
||||
|
||||
The installer places the binary in `~/.grok/bin` and symlinks it into `~/.local/bin`
|
||||
(it also installs an `agent` alias Codeman ignores). `grok update` self-updates.
|
||||
|
||||
Codeman resolves the binary via the server PATH and then the usual install locations,
|
||||
`~/.grok/bin` first. **`grok` is a name with known squatters** (the unrelated
|
||||
`@vibe-kit/grok-cli` npm package also installs a `grok` bin), so like `pi` the
|
||||
resolver does not trust a PATH hit on its own: it runs `grok --version` once and
|
||||
requires version-shaped output (`grok 1.0.5 (5115b46bc9)`). Check what it resolved:
|
||||
|
||||
```bash
|
||||
curl -s localhost:3000/api/grok/status | jq
|
||||
# { "available": true, "path": "/home/you/.grok/bin", "version": "1.0.5" }
|
||||
```
|
||||
|
||||
The endpoint carries `version` on top of the sibling `/api/*/status` shape precisely
|
||||
so a misresolution is visible rather than presenting as "the mode just doesn't work".
|
||||
|
||||
## Authenticate
|
||||
|
||||
- **Browser OAuth (default)**: the first `grok` run opens a sign-in flow; in a
|
||||
Codeman pane you get the device-code screen with a URL to open elsewhere.
|
||||
Credentials land in `~/.grok/auth.json` (0600) and refresh automatically.
|
||||
- **Device code**: `grok login --device-auth`, made for SSH boxes and headless hosts.
|
||||
- **API key**: `export XAI_API_KEY="xai-..."` (console.x.ai). Used as a fallback when
|
||||
no session token exists. As a per-session Codeman `envOverride` it flows through
|
||||
socket-scoped `tmux setenv`, never the spawn command line.
|
||||
- **Enterprise OIDC**: `GROK_OIDC_ISSUER` / `GROK_OIDC_CLIENT_ID`.
|
||||
|
||||
## What Codeman wires up
|
||||
|
||||
`GrokConfig` (per session, persisted in `state.json`, round-trips through respawn):
|
||||
|
||||
| Field | Flag | Notes |
|
||||
| ----------------- | --------------------------- | --------------------------------------------------------------------- |
|
||||
| `model` | `--model <v>` | e.g. `grok-4.5`, or a custom `[model.<name>]` from `config.toml` |
|
||||
| `alwaysApprove` | `--always-approve` | Grok's `bypassPermissions` mode; deny rules still apply on top |
|
||||
| `continueSession` | `--continue` | Most recent session for the working directory; skipped when resuming |
|
||||
| `resumeSessionId` | `--resume <v>` | Ids only, never titles (grok's own `--resume` also matches titles) |
|
||||
|
||||
Every value is regex-validated and **dropped** (not escaped) if it fails, because the
|
||||
result is interpolated into the pane's `bash -c "..."` command.
|
||||
|
||||
The Run button sends `grokConfig: { alwaysApprove: true }`, the same product decision
|
||||
as Claude's `--dangerously-skip-permissions` default and Antigravity's
|
||||
`--dangerously-skip-permissions`: Codeman sessions exist for autonomous work. Keep
|
||||
hard limits as `deny` rules in `~/.grok/config.toml` (they apply in every mode), and
|
||||
in **multi-user mode** a non-granted owner's `alwaysApprove` is forced off
|
||||
server-side; a bare `grok` spawn is grok's own ask-mode default.
|
||||
|
||||
Env overrides: the `GROK_*` prefix (`GROK_HOME`, `GROK_CONFIG`, `GROK_MEMORY`,
|
||||
`GROK_WORKFLOWS`, `GROK_SANDBOX`, `GROK_OIDC_*`, ...) plus the `XAI_*` vendor
|
||||
namespace (`XAI_API_KEY`) are allowlisted. Foreign provider keys are not, as ever.
|
||||
|
||||
## What Codeman deliberately does NOT wire up
|
||||
|
||||
- **`--permission-mode`, `--allow`/`--deny`.** The boolean covers the autonomous
|
||||
case; the full rule surface is a follow-up with UI.
|
||||
- **`-p`/headless, `--output-format`, `--json-schema`.** Codeman drives the TUI.
|
||||
- **`--worktree`, `--sandbox`, `--reasoning-effort`, `-s/--session-id`,
|
||||
`--fork-session`, `--agent`/`--agents`.** Tracked as follow-ups in the plan doc.
|
||||
|
||||
## Terminal behavior
|
||||
|
||||
Grok renders a **fullscreen alternate-screen TUI** (scrollback pane + prompt, mouse
|
||||
supported). Under Codeman it runs inside tmux like every external CLI, so the
|
||||
fullscreen rendering stays inside the pane and the browser terminal shows tmux's
|
||||
repaints; grok stays out of the alt-screen strip list on purpose (the opencode case,
|
||||
not the Ink case). If a pane renders oddly, `grok doctor` checks terminal, color and
|
||||
input support without starting a session, and `~/.grok/pager.toml` can force
|
||||
`alt_screen = "inline"`.
|
||||
|
||||
On touch devices grok currently gets the buffered local-echo overlay like Claude,
|
||||
Gemini, OpenCode and Pi. This is the fallthrough default and has not been measured
|
||||
against an authenticated grok composer; if grok turns out per-keystroke reactive the
|
||||
way codex was (issues #218/#219/#220/#222), the fix is the `'off'` branch in
|
||||
`_updateLocalEchoState` (terminal-ui.js).
|
||||
|
||||
## Docker cases
|
||||
|
||||
The agent image installs grok in its own Dockerfile step (not npm; xAI's installer
|
||||
targets `$HOME/.grok/bin` with no `--dir` override, so the binary is copied to
|
||||
`/usr/local/bin`). Rebuild with the mandatory `--no-cache`:
|
||||
|
||||
```bash
|
||||
node scripts/build-agent-image.mjs --no-cache
|
||||
```
|
||||
|
||||
Credentials are **seeded**, not shared: `auth.json`, `config.toml` and `pager.toml`
|
||||
are copied into the container's own `~/.grok`, so an in-container grok never writes
|
||||
refreshed OAuth tokens back to the host and `docker commit` exports stay secret-free.
|
||||
Only those three files, because `~/.grok` also holds `sessions/`, `memory/` and the
|
||||
~160MB binary under `downloads/`. Trade-off, same as pi: in-container sessions are
|
||||
invisible host-side, so `grok -c` inside a Docker case only sees that container's own
|
||||
history.
|
||||
|
||||
## Remote SSH cases
|
||||
|
||||
`grok` mode is routed through an interactive login shell
|
||||
(`exec "$SHELL" -i -l -c 'grok'`), because sshd's remote-command PATH does not include
|
||||
`~/.grok/bin`. Per-session config and `envOverrides` do not cross ssh and are rejected
|
||||
rather than silently ignored; use the per-host command override instead. For auth on
|
||||
the remote host, `grok login --device-auth` exists for exactly this.
|
||||
|
||||
## Known gaps
|
||||
|
||||
- **No idle/completion hook yet.** Idle detection falls back to output-stabilization
|
||||
like the other external CLIs. Grok has a hooks system, so a Codeman hook POSTing to
|
||||
`/api/hook-event` is the highest-value follow-up.
|
||||
- **No response viewer.** Grok writes ACP JSONL sessions under
|
||||
`~/.grok/sessions/<encoded-cwd>/<session-id>/updates.jsonl`; nothing reads them yet.
|
||||
- **Cron jobs mis-detect readiness.** The readiness poll looks for `❯` or a token
|
||||
count, neither of which grok prints, so a grok cron job burns its poll budget and
|
||||
then sends the prompt anyway. It works; it is just slower to start.
|
||||
- **Ralph, respawn heuristics, token/CLI-info parsing and the `❯` readiness probe are
|
||||
off** for grok, as for every external CLI.
|
||||
|
Before Width: | Height: | Size: 357 KiB |
|
Before Width: | Height: | Size: 941 KiB |
|
Before Width: | Height: | Size: 1.0 MiB |
|
Before Width: | Height: | Size: 56 KiB |
|
Before Width: | Height: | Size: 859 KiB |
@@ -1,3 +0,0 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 320 60">
|
||||
<text x="160" y="48" font-family="system-ui, -apple-system, 'Segoe UI', Roboto, sans-serif" font-size="52" font-weight="700" fill="#60a5fa" text-anchor="middle">Codeman</text>
|
||||
</svg>
|
||||
|
Before Width: | Height: | Size: 247 B |
|
Before Width: | Height: | Size: 357 KiB |
|
Before Width: | Height: | Size: 96 KiB |
|
Before Width: | Height: | Size: 87 KiB |
|
Before Width: | Height: | Size: 84 KiB |